hirly

Sarvam

Performance Engineer, On-Device Inference

Bengaluru

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Sarvam first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.3M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

hirly's read of this role

Seniority
Mid level
Country
IN
Work mode
On-site / unstated
First seen by hirly
28 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

About Sarvam

Sarvam is building the bedrock of Sovereign AI for India. The company is developing India’s full-stack sovereign AI platform, building across research, models, infrastructure and applications with a singular focus on making AI genuinely work for India. Sarvam works with leading enterprises and public institutions and is backed by Lightspeed, Peak XV, and Khosla Ventures. Sarvam partners with India’s leading brands, including Tata Capital, SBI Life, CRED, IDFC, and LIC.

About the Role

Take Sarvam models from research-handoff state to production-ready artifacts on at least two of our target chipsets (Intel xPU, ARM xPU, Apple xPU, Nvidia / AMD GPUs). You'll own 1 - 2 (model, chipset) pairs end-to-end and partner with the consuming app team during integration.

What You’ll Do

Quantize, validate accuracy, benchmark, and document. Author the deployment workbook for each pair you own.

Embed part-time with consuming teams during integration; debug perf and accuracy issues alongside them.

Maintain and extend the team's benchmark harness.

What We're Looking For

3+ years on ML systems.

Solid PyTorch + ONNX export experience including the gotchas (dynamic shapes, control flow, custom ops).

Quantization in production on at least one real model.

Comfort with at least two of: ONNX Runtime, TensorRT, CoreML, OpenVINO, QNN, LiteRT.

Profiling fluency on at least one platform.

Bonus Points

Custom op authoring in any runtime.

Why Sarvam?

Sarvam is a fast-moving, high talent-density team building full-stack AI for India, working on problems that push the frontiers of AI with real population-scale impact.

Work alongside researchers, engineers, builders, and business leaders who move fast and hold each other to a very high bar

High ownership and high impact, from day one

Everything we do is AI-first, from the way we build and ship to the way we think about problems

You can work on problems that could change how an entire country learns, works, and communicates

If you want to work on problems at the frontier of AI in India, Sarvam is the place to be.

Original posting on Sarvam's site ↗

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job