hirly

Heidi

Senior AI Infrastructure Engineer

Melbourne

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Heidi first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.3M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

hirly's read of this role

Seniority
Senior
Country
AU
Work mode
On-site / unstated
First seen by hirly
28 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

We’re Heidi.

We're building the future of healthcare by giving every clinician the earth's finest AI Care Partner. In just 18 months, our clinical AI products have absorbed the administrative chaos of 73 million patient visits. Today, we support over 2.5 million patient sessions a week across 190+ countries.

Healthcare systems are failing us; clinicians spend more time on documentation than on patients, and the human connection that makes medicine worth practicing is eroding. Our mission is simple: double the world’s healthcare capacity and strengthen the human connection at its heart.

We found product-market fit with a freemium medical scribe that clinicians love. Now, we're expanding. Every task a clinician hands to Heidi is a patient who feels more attended to, a health system unclogged, and a clinician who gets to be a clinician again.

If you don’t choose easy and you want to build something way bigger than yourself then, choose the challenge, choose Heidi.

The role

This role sits in the model team, the researchers and engineers who train, deploy, and own the AI models behind every Heidi product. You’ll build and operate the infrastructure that makes those models fast, reliable, and cost-effective at scale.

Your work will span production model serving, GPU cluster management, and the infrastructure supporting training and evaluation. You’ll decide how workloads share compute, diagnose performance bottlenecks, and build the deployment and observability tools that help the team move quickly with confidence.

We’re looking for a hands-on engineer who has deployed and operated models in production, understands the demands of GPU workloads, and can take a system from initial design through rollout, incidents, and ongoing improvement. You’ll partner closely with researchers and our platform engineers, with ownership of the systems you build.

What you’ll do

Build and own model-serving infrastructure. Take models from checkpoint to production, with repeatable deployment pipelines, request routing, autoscaling, fallback paths, and controlled rollouts and rollbacks across regions.

Manage GPU clusters and workload scheduling. Improve resource allocation across inference, training, and evaluation. Build scheduling policies around workload priority, quotas, hardware topology, and recovery requirements so online services stay responsive while other workloads make productive use of capacity.

Improve inference performance. Profile real workloads and improve latency, throughput, and memory efficiency. Evaluate batching, KV-cache management, quantization, speculative decoding, and parallelism strategies against production traffic and quality requirements.

Support distributed training and model iteration. Give researchers reliable ways to launch fine-tuning and training jobs, manage model artifacts, save and restore checkpoints, and move validated models into serving. Reduce time lost to failed jobs, slow data loading, and manual setup.

Make deployments observable and incidents traceable. Connect application requests and sessions to the exact model, deployment configuration, and worker that served them. Build dashboards and alerts covering model latency, queueing, errors, GPU health, memory pressure, and workload performance.

Own production reliability. Define service objectives, investigate incidents across the application, inference engine, GPU, and network layers, and build recovery procedures that work. Turn recurring failures into fixes, automated checks, and useful runbooks.

Make compute costs actionable. Track GPU usage, idle capacity, and inference cost by model and workload. Use capacity forecasts and measured performance to guide deployment choices and improve cost per successful request without sacrificing quality or reliability.

Build a platform the model team can use independently. Automate provisioning, configuration, benchmarking, and releases. Partner with the engineers behind ASR, note generation, Evidence, and Dictate so new models can be deployed and evaluated through consistent, well-supported workflows.

What you'll need

A strong engineering foundation. Hands-on AI infrastructure experience. At least 1 year building and operating infrastructure for large language models, including model deployment, inference serving, or distributed training. You can design the system, write the code, and own it in production.

Production model deployment experience. You’ve deployed and maintained LLMs or other demanding ML workloads using engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or comparable systems. You understand the work between loading a model and running a reliable service.

GPU cluster and orchestration experience. You’ve managed GPU workloads using Kubernetes, Slurm, or an equivalent platform, with practical experience in scheduling, resource allocation, capacity planning, and failure recovery.

Performance debugging skills. You can use traces, metrics, and profiling tools to distinguish compute, memory, communication, and scheduling bottlenecks. You understand how batch size, context length, precision, and multi-GPU execution affect performance and cost.

Strong software and systems skills. You’re proficient in Python and comfortable with backend or systems development in Go, C++, Rust, or a comparable language. You have practical experience with Linux, containers, deployment automation, and distributed services.

Operational ownership. You’ve owned production incidents, built useful monitoring, and made releases recoverable. You can explain the trade-offs behind a design and work effectively with researchers, product engineers, and infrastructure partners.

Nice to have

Experience with distributed training frameworks such as PyTorch FSDP or Megatron, or infrastructure for reinforcement learning and rollout generation.

Experience tuning inference engines, serving MoE models, or implementing quantization, speculative decoding, and prefill/decode disaggregation.

Familiarity with GPU interconnects, NCCL, RDMA, topology-aware scheduling, or diagnosing multi-node communication problems.

CUDA or Triton kernel development, contributions to AI infrastructure projects, or experience building cluster operators and scheduling integrations.

Experience operating infrastructure across multiple regions or providers, particularly for healthcare or other sensitive production workloads.

How we show up

Build for the next decade, not next quarter . Our targets are outrageous on purpose. The world's health doesn't have the luxury of incrementalism.

Lead, don't wait . We treat tomorrow's problems today. Sometimes we build what's needed before it's wanted, and we're fine with that.

Follow the evidence . Trust the patient. We pursue truth relentlessly. But when the subjective and objective disagree, we treat the patient, not the numbers. Ego is a comorbidity we can't afford.

Own the outcome . Everyone here carries the company. Raise problems with solutions, solve them end-to-end, and never be a bystander.

Ship, measure, go again . A button today, a workflow tomorrow. More iterations beat better planning. We're precise at pace, not reckless.

Live in clinicians' reality . Not the ideal workflow, the twenty-patients-before-lunch actual one. We build for exhausted humans, and we'd better be decent ones while we do it.

Why Heidi?

You’ll join a team focused on real-world impact over imaginary valuations and glossy PR. We live and breathe the challenges of modern health systems, and are laser-focused on exacting the change we’d like to see. We’re medicos, engineers, builders, and designers who’ve felt the moral and practical toll of what non-care feels like. True A-players progress extremely fast here.

The nature of the scale-up game is demanding, but we value sustainable performance and mental health. You're trusted to perform, and you set your schedule. We operate on out

Original posting on Heidi's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job