hirly

Cantina

Machine Learning Engineer, Ops

Europe

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Cantina first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.5M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

Apply from your AI assistant

Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.

Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.

hirly's read of this role

Role family
Data & ML
Seniority
Mid level
Stated salary
$125,000 – $165,000 per year
Work mode
On-site / unstated
First seen by hirly
2 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

About Cantina:

Cantina Labs is a social AI company, developing a suite of advanced real-time models that push the boundaries of expression, personality, and realism. We bring characters to life, transforming how people tell stories, connect, and create. We build and power ecosystems. Cantina, our flagship social AI platform, is just the beginning.

If you're excited about the potential AI has to shape human creativity and social interactions, join us in building the future!

About the Role:

We are looking for an ML Engineer to own the production inference stack for our text-to-speech (TTS) and speech recognition (ASR) models, end to end. Our research team delivers trained model checkpoints; you will productionize them as low-latency, real-time streaming services that run cost-efficiently on GPUs. The role covers the inference engines, the Kubernetes infrastructure they run on, and the performance benchmarking and accuracy validation that let us ship changes safely. You will bridge research and production, and own latency, throughput, cost and accuracy as we scale.

What You’ll Do:

Build and optimize streaming inference engines for TTS and ASR models, including multi-stage model pipelines.

Optimize latency (time to first chunk, p50/p99) and throughput using continuous batching, CUDA Graphs, torch.compile and mixed precision.

Maintain numerical parity with the reference implementation through automated accuracy regression tests, per stage and end to end.

Deploy and operate GPU inference services on Kubernetes: autoscaling, containerization and infrastructure as code.

Build CI/CD for model artifacts, with model versioning and reproducible releases.

Own observability and performance benchmarking: p50/p99 latency, real-time factor (RTF), GPU utilization and cost per request.

Partner with research to productionize new models.

What You’ll Bring:

Hands-on experience deploying, scaling and optimizing ML models in production, beyond consuming model APIs.

Working knowledge of the full inference stack, from CUDA kernels to serving frameworks, with depth in some layers; expertise across all of them is not expected.

GPU performance optimization: profiling (Nsight Systems, PyTorch Profiler), memory bandwidth, CPU-GPU synchronization.

Solid understanding of model serving: continuous batching, KV cache management, request scheduling.

Production experience with Kubernetes, CI/CD and cloud infrastructure.

Strong Python and PyTorch, and a rigorous approach to performance benchmarking.

Experience with speech or audio models (TTS, ASR, voice conversion).

Experience with inference frameworks such as vLLM (vLLM-Omni), SGLang, TensorRT-LLM, TensorRT or Triton Inference Server.

Compensation:

The anticipated annual base salary range for this role is between $125,000-$165,000 (€110,000-€145,000). When determining compensation, a number of factors will be considered, including skills, experience, job scope, location, and competitive compensation market data.

Benefits for U.S.-based roles:

Competitive salary and generous company equity

Medical, dental, and vision insurance – 99.99% of premiums covered by Cantina

42 days of paid time off, including:

15 PTO days

10 sick days

15 company holidays

2 floating holidays

Generous parental leave & fertility support

401(k) retirement savings plan

Lifestyle spending account – $500/month to use however you’d like

Complimentary lunch and snacks for in-office employees

One Medical membership, and more!

Original posting on Cantina's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job
Machine Learning Engineer, Ops – Cantina · Europe | hirly.me