Staff Software Engineer

Crusoe
San Francisco, CA

Crusoe is on a mission to accelerate the abundance of energy and intelligence . As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.

We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that — with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.

We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.

If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.

Crusoe’s Data Center Infrastructure Engineering (DCIE) team is fundamental to our mission of providing AI hardware and infrastructure as a service. The team provides infrastructure for Crusoe’s fleet GPU’s and data center. The team sits at the nexus of high performance computing and AI infrastructure as a service.

The DCIE team owns, deployment maintenance, observability, critical environment, and automation. The team builds, and maintains GPU clusters, develops automation for logical and physical maintenance, provision systems, and observability tooling.

About the Role:

We are seeking a highly skilled and motivated Software Engineer to join Crusoe’s Data Center Infrastructure Engineering team. This position is focused on the development of software for the management of a fleet of GPU servers as well as the data centers that house those systems. The role focuses on the developing and implementing advanced diagnostic, observability, automation and repair tooling for high-performance GPU compute clusters.

The ideal new team member will be a hands-on problem solver who is comfortable working independently. The new team member will play a critical role in maintaining the health and scalability of Crusoe’s rapidly growing GPU fleet.

What You’ll Be Doing:

  • Developing and implementing deep-level diagnostics and troubleshooting of hardware faults within GPU racks and high-density compute systems.

  • Developing troubleshooting and automation tooling for GPU platforms including NVIDIA A100, H200, GB200, B200 and AMD 350X / 355X.

  • Developing automation and AI agents for executing component-level diagnosis and remediation for failed or degraded hardware.

  • In conjunction with data center operations develop innovative tooling and AI agents for managing the critical environment.

  • Developing tooling for post-repair validation and testing tools such as burn-in, Pytorch, and NVIDIA NCCL to ensure system stability and performance.

  • Own the deployment, monitoring, and operational support of developed tooling, ensuring solutions maximize GPU fleet availability and performance to drive customer success.

  • Developing automation and operational tooling for facilities management power as well as direct liquid cooling hardware systems

What You’ll Bring to the Team:

  • Software engineering experience.

  • The ability to identify a problem, rapidly develop a scalable solution and ship it.

  • Ability to lean in and assist team members working on critical or complex technical initiatives.

  • Ability to set the technical direction for a specific project and execute.

  • Expertise in distributed systems, reliability, and cloud platforms (Kubernetes, IaC, GCP etc.)

  • Strength in at least one programming language - Go, Python, Java, Rust.

  • Strong analytical and problem-solving skills.

  • Excellent communication and collaboration skills.

  • Ability to work independently and within a team

Nice to Have:

  • Experience with Temporal and Kubernetes.

  • Experience working directly with hardware vendors.

  • Background in large-scale GPU fleet operations or hyperscale data center environments.

Benefits:

  • Industry competitive pay

  • Restricted Stock Units in a fast growing, well-funded technology company

  • Health insurance package options that include HDHP and PPO, vision, and dental for you and your dependents

  • Employer contributions to HSA accounts

  • Paid Parental Leave

  • Paid life insurance, short-term and long-term disability

  • Teladoc

  • 401(k) with a 100% match up to 4% of salary

  • Generous paid time off and holiday schedule

  • Cell phone reimbursement

  • Tuition reimbursement

  • Subscription to the Calm app

  • MetLife Legal

  • Company paid commuter benefit; $300 per month

Compensation Range

Compensation will be paid in the range of up to $208,000 - $253,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicants knowledge, education, and abilities, as well as internal equity and alignment with market data.

Crusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.

Posted 2026-04-04

Recommended Jobs

Construction Contract Administrator III

The Greenridge Group
Los Angeles, CA

The Greenridge Group is a prime contractor and consulting firm specializing in Project and Construction Management for major public-sector agencies. We are seeking an experienced Contract Admi…

View Details
Posted 2026-02-24

Analyst - Visa Crypto Sales & Operations

VISA
San Francisco, CA

Job Description As an Analyst within Visa’s Crypto organization, you will support both sales and sales operations activities across our portfolio of crypto‑native wallets, exchanges, ramps, orchest…

View Details
Posted 2026-01-30

Senior Footwear Product Testing Analyst - Energy (Los Angeles)

Nike
Los Angeles, CA

Become a Part of the NIKE, Inc. Team NIKE, Inc. does more than outfit the world’s best athletes. It’s a place where passionate individuals come together to create the futur…

View Details
Posted 2026-03-27

Occupational Therapist, Temp & Perm Available (Job ID: 128)

Blue United Sourcing
Los Angeles, CA

&##128680; IMMEDIATE NEED – Travel Occupational Therapist (OT) &##128680; Skilled Nursing Facility | Roseville, CA 💼 13-Week Contract ⏰ 36 Hours per Week 💰 $60–$65/hr 📅 Start ASAP We’…

View Details
Posted 2026-01-15

Cleaner/Limpiador(a) Part Time Huntington Beach, CA

Slate
Huntington Beach, CA

Slate is a professional and trusted commercial cleaning company dedicated to maintaining clean, safe, and inviting spaces for our clients. Known for reliability, attention to detail, and seamless dig…

View Details
Posted 2026-03-22

Project Manager - Heavy Civil Engineering

DeSilva Gates Construction
Chico, CA

We’re seeking a driven Project Manager to join our Chico team and take the lead on civil construction projects from start to finish. If you’re a hands-on Project Manager who enjoys solving problems, …

View Details
Posted 2026-02-25

Product Manager(Marketing and R&D)

Pluslife
San Diego, CA

1. Own the end-to-end product lifecycle—from concept definition through regulatory submission, launch, and post-launch release management. 2. Partner with R&D, Quality/Regulatory (QA/RA), and Manufa…

View Details
Posted 2026-01-21

Personal Care Aide (PCA)

Comfort Keepers
Stockton, CA

Become a Personal Care Aide (PCA) and Caregiver with Comfort Keepers and join a compassionate team of people, like you, who are dedicated to providing companionship and personal care for seniors and …

View Details
Posted 2026-01-30

Senior Software Engineer, Motion Controls

Waymo
Mountain View, CA

Waymo is an autonomous driving technology company with the mission to be the world's most trusted driver. Since its start as the Google Self-Driving Car Project in 2009, Waymo has focused on buildi…

View Details
Posted 2026-02-16

Case Manager - Personal Injury

Wilshire Law Firm
Los Angeles, CA

Case Manager - Personal Injury Wilshire Law Firm is a distinguished, award-winning legal practice with over 18 years of experience, specializing in Personal Injury, Employee Rights, and Consumer Cl…

View Details
Posted 2025-10-22