hirly

Nscale

Design Reliability Engineer — Power and Energy

Houston · US

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Nscale first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.6M live jobs from 190,000+ employers in 200+ countries.

Tailor my resume for this job →

Apply from your AI assistant

Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.

Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.

hirly's read of this role

Seniority
Mid level
Stated salary
$130,000 – $200,000 per year
Country
US
Work mode
Remote-friendly
First seen by hirly
3 Oct 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

About Nscale

At Nscale, we are building the infrastructure for the AI revolution. Nscale is developing cutting-edge, sovereign generative AI solutions powered by a new generation of high-performance, sustainable data centers and GPUs built specifically for AI workloads.

The rapid growth of artificial intelligence is driving unprecedented global demand for compute and GPU infrastructure. Positioned at the heart of this transformation, Nscale is delivering the platforms that will enable the next decade of innovation.

Working closely with the world's most advanced AI technology providers — including Microsoft and NVIDIA — Nscale integrates next-generation compute hardware and GPU clusters across both partner and self-delivered facilities. This is a unique opportunity to join Nscale's journey, play a pivotal role in delivering transformative AI infrastructure, and help shape a company designed to scale rapidly across North America, Europe, and beyond.

About the Role

We are hiring a Design Reliability Engineer to own the analysis that proves our systems will hit their availability targets, and the discipline that makes those numbers real.

Nscale's Power and Energy group develops behind-the-meter generation colocated with our data center campuses: fleets of hundreds of reciprocating engines, battery energy storage, and medium- and high-voltage distribution operating as islanded microgrids, serving loads with contractual availability commitments measured at the GPU. At this scale and redundancy, availability is not a brochure number. It is an engineered quantity: allocated from committed SLAs down to systems and equipment, modeled with rigor, defended against common-mode risks, and eventually measured live against the models that predicted it. This role owns that quantity, end to end, across power and mechanical scopes together rather than in silos.

This is a hands-on technical leadership role built for someone who runs with minimal oversight. You will own the availability model of record for each campus, the FMEA program across major equipment, and the reliability foundations of our digital twin: the living, data-fed model that will carry RAM analysis from design studies into day-to-day operations. You will direct the RAM consultants and reliability data work supporting the portfolio, challenge methodology on its merits, and put analysis in front of executives that they can bet commercial commitments on.

You will sit within the Power and Energy technology organization, working alongside the electrical, controls, mechanical, and pipeline engineering managers whose designs your models test, and with the commissioning and operations teams whose data will eventually close the loop. You will be the owner's technical authority on reliability, directing the work of consultants rather than simply accepting it, and resolving the questions where textbook RAM practice does not account for a system this large, this redundant, or this consequential.

Location: Houston, TX (remote considered within the US, with regular site and vendor travel)

What You'll Be Doing

RAM modeling and availability assurance

Own the availability model of record for each generation campus: system-level RAM models spanning generation, electrical distribution, fuel supply, and cooling, built and maintained as living engineering assets.

Allocate committed SLA targets down through the system: availability budgets by subsystem, redundancy requirements, MTTR assumptions, and sparing and maintenance strategies that make the top-level number achievable.

Apply the right method for each question — Monte Carlo simulation (BlockSim or equivalent) where it earns its complexity, closed-form analytical models where they are faster and more auditable — and defend the choice on its merits.

Hunt common-mode and dependent failures relentlessly: shared fuel supply, shared cooling, control system dependencies, and site-wide events that redundancy counts conceal.

Report availability the way commitments are written: both single-path and contracted-capacity views, with sensitivities that show which assumptions carry the number.

FMEA and reliability engineering

Own the FMEA/FMECA program for major equipment and systems: engines and generators, BESS, switchgear and transformers, fuel gas systems, and cooling infrastructure, establishing baselines and keeping them current as designs mature.

Curate the failure rate and repair data underpinning every model: industry sources such as IEEE 493 and OREDA, OEM data challenged rather than transcribed, and field data as it accumulates.

Turn analysis into design influence: redundancy configuration, single-point-of-failure treatment, equipment selection input, and testability and maintainability requirements fed into the engineering teams while designs can still change.

Bring reliability analysis into design reviews, HAZOPs, and vendor evaluations as a routine discipline, not a report delivered after decisions are made.

Digital twin and reliability data systems

Own the reliability core of Nscale's digital twin: the model architecture, data structures, and analytical methods that let availability be recomputed on the fly as system state, configuration, and failure data change.

Define the operational data requirements — event capture, failure coding, downtime attribution, and run-hour tracking — so the twin is fed by trustworthy data from day one of operations.

Work with our software and data teams to move RAM analysis from static studies into instrumented, queryable tooling, and champion analytical approaches that are automatable and auditable rather than locked in desktop tools.

Establish the feedback loop: measured availability versus model prediction, model recalibration, and reliability growth tracking from commissioning onward.

Delivery, commissioning, and operations support

Support commissioning and startup with reliability input: burn-in and reliability run design, failure tracking during startup, and acceptance criteria grounded in the availability model.

Lead and support root cause analyses of significant failures and availability events, and drive corrective actions back into designs, models, and standards.

Develop Nscale's reliability engineering standards, methods, and reference models, so each campus builds on the last instead of starting over.

Team and vendor leadership

Direct RAM and reliability consultants: set the basis they work to, review and challenge their methodology and deliverables, and integrate their output into one coherent picture across power and mechanical scopes.

Coordinate with data center reliability and operations teams so that availability is engineered and measured across the full path to the GPU, not just to the fence line.

Grow the internal reliability engineering capability over time, building a small team as the portfolio scales.

Communicate complex reliability analysis clearly to project leadership and executives, with honest assessments of confidence, sensitivity, and risk.

About You

5+ years of reliability engineering experience on power generation, process, or mission-critical facilities, with significant time owning RAM analysis for large, redundant systems.

Deep RAM modeling capability across both discrete-event simulation (BlockSim, Raptor, or equivalent) and closed-form analytical methods, with the judgment to know which the question deserves.

Proven FMEA/FMECA leadership on major rotating, electrical, or process equipment, and fluency with the standard failure data sources (IEEE 493, OREDA, IEEE 3006 series or similar) and their limitations.

Experience allocating availability targets from commercial commitments down to systems, and defending the resulting analysis to executives, customers, or insurers.

A sharp eye for common-mode and dependent failure mechanisms, and a track record of finding the risks that redundancy arithmetic hides.

Work

Original posting on Nscale's site ↗

Listed on hirly, a job board. hirly is not the employer: Nscale is hiring for this role.

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job