Software Engineer, Facilities
Software Engineering, Operations · Full-time
San Francisco, CA, USA · Austin, TX, USA · Seattle, WA, USA
About Fluidstack
We exist to make humanity more free. For most of human history, you farmed or you starved. Technology gave people more time for the things they wanted to do, instead of things they had to do. Powerful AI will be the biggest lever for human choice we've ever built - but only if models are aligned with what humanity actually wants. There are groups building AI who don't share these goals. Whoever deploys frontier compute infrastructure fastest will decide whether AI expands human freedom or shrinks it.
We're singularly focused on delivering 10 to 100s of GWs of compute faster than anyone else, rethinking every layer of the stack. We acquire power, design and build data centers, and operate them - with teams spanning hardware and software. Speed and scale are our key differentiators. Come be a part of building civilization-scale infrastructure for AI.
We hire people who care deeply about this problem space. If that is you, please apply!
How We Operate
Be a barrel. Full autonomy. Own things end to end, take on scope without being asked, no permission required to operate outside your core role.
Insane urgency. We drive everything forward as fast as possible.
Reason from first principles. Challenge every assumption. Zero analogy thinking, no egos, the best idea wins.
Love of the game. The frontier of AI is the most interesting problem of our time. We put in long hours at high intensity to push the frontier forward.
Build something that actually matters. If you're going to spend your time, spend it on something that matters to the world.
The Facility Controls Team
Examples of key problems the team is working on:
Controlling datacenter power demand in real time while keeping equipment within its operating limits.
Integrating power and cooling equipment and on-site generation so they can be monitored and controlled together.
Automating commissioning tests that verify equipment behavior and record the results.
Automating software deployment and configuration as new sites and equipment come online.
Processing telemetry from facility equipment and serving current and historical data through reliable APIs.
You'll build the production services behind this work: telemetry pipelines, equipment models, control APIs, and test orchestration. Most of our services use Go, NATS, ClickHouse, Redis, and Protobuf/gRPC, and run on Kubernetes.
Role Scope
Build and own the services other teams depend on. Device integrations, telemetry pipelines, equipment models, and APIs need to work together across sites.
Design the messaging and state layer. Preserve ordering where it matters, handle consumers falling behind, and recover after a restart without losing data or applying an action twice.
Own the data contracts and engineering standards others build against. Make device identity, units, timestamps, and data quality explicit, and evolve schemas and APIs without breaking consumers.
Build the path from a control request to a confirmed result. Enforce authorization, deadlines, and equipment constraints, track state changes, and make retries safe.
Build observability that explains failures. Measure ingestion lag, dropped data, stale readings, query latency, and command completion so an operator can tell where a problem started.
Prove behavior across the whole system. Use unit tests, integration tests, and simulated devices to exercise failure and recovery before changes reach live equipment.
What We're Looking For
Experience building large-scale software systems, particularly systems that produce and process large volumes of data, is relevant.
You've shipped production code in Go, Python, or TypeScript, and you pick up whatever language the problem demands.
You have worked with high-throughput messaging systems such as NATS or Kafka. You understand ordering, delivery guarantees, backpressure, and the tradeoffs between concurrency and correctness.
You have designed data models for large volumes of time-stamped data. You can reason about slow queries and explain the tradeoffs between write throughput, query performance, retention, and schema evolution.
You have instrumented production services with metrics, logs, and profiling. You can show how that instrumentation helped you find a bottleneck or explain an incident.
You design before you build. You can explain the alternatives you considered, why you chose a particular design, and when you changed it because it stopped fitting the problem.
You define interfaces and failure boundaries deliberately. You know what happens when a dependency is unavailable, and your tests exercise those failures and the recovery path.
You have followed data from its source through services and storage. You account for missing, duplicate, late, and invalid data, and handle a stale value differently from a wrong one.
Bonus: Prior datacenter or controls experience.
Useful Experience
Our stack: Go, NATS, Redis, ClickHouse, Protobuf/gRPC, Kubernetes, and ArgoCD. Prometheus and Grafana for production observability.
Industrial protocols such as Modbus or OPC UA.
We are committed to pay equity and transparency.
Fluidstack is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Fluidstack will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.
You will receive a confirmation email once your application has successfully been accepted. If there is an error with your submission and you did not receive a confirmation email, please email careers@fluidstack.io with your resume/CV, the role you've applied for, and the date you submitted your application-- someone from our recruiting team will be in touch.