Eka.care
Senior Data Scientist ( Kernel Optimisation & Inference Engineer)
Bengaluru, India
Apply through hirly
hirly scores this role against your resume, shows its reasoning, then writes a resume and cover letter for it and fills the application with you. Free to start — no card required.
hirly's read of this role
- Role family
- Data & ML
- Seniority
- Senior
- Country
- IN
- Work mode
- On-site / unstated
- First seen by hirly
- 23 Sept 2026
Derived automatically from the posting. Sign up to see how the role scores against your own resume.
the posting
Kernel Optimisation & Inference Engineer
Bengaluru · Full-time · 2–5 yrs
Somewhere between the model and the silicon, 10–20% of a training budget goes missing. Your job is to go get it back, and then make the same model fast enough to run in a clinic, and small enough to run on a phone.
About EkaCare and the mission
EkaCare is India's connected healthcare platform: an EMR that doctors run their practices on, a personal health record used by millions of Indians, and one of the deepest integrations with India's ABDM digital-health rails. Our Parrotlet family of medical models already serves Indian doctors in production, and we open-source our work where it counts.
The role:
You'll work with our performance lead on making everything fast: training-side fused kernels and MFU on the 30B MoE, inference-side latency and throughput, and the quantised 2B/4B on-device tier. Hardware-up: profiler first, roofline reasoning always, custom kernels when the math says so.
What you'll do
- Profile training and inference workloads and hunt utilisation gaps across kernels, memory and comms.
- Write and tune CUDA kernels where existing ops leave real performance on the table, and know when they don't.
- Optimise MoE-specific paths: grouped GEMMs, all-to-all communication, expert load imbalance.
- Build the fast inference path: vLLM-class serving, continuous batching, prompt/prefix caching for clinical-context workloads, speculative decoding.
- Own quantisation for the 2B/4B variants (AWQ/GPTQ-class, fp8) — with eval-parity verification, not just perplexity.
- Make on-device inference real for the hardware Indian clinics actually have.
What we look for
- 2–5 years in GPU performance work; you've profiled real workloads and shipped optimisations with before/after numbers you can defend.
- Working fluency in CUDA, and memory-hierarchy reasoning (coalescing, occupancy, SRAM tiling; you can explain *why* FlashAttention is fast).
- Hands-on with a modern serving stack (vLLM, TensorRT-LLM, SGLang or similar) beyond just running it.
- Measurement discipline: you profile before optimising and verify correctness after.
Bonus
- fp8 experience on H100/H200-class hardware; torch.compile/inductor internals.
- Quantisation research or on-device/mobile inference experience.
- Open-source kernels or serving contributions.
Why this is a rare gig
- Open source, with your name on it : weights and technical reports ship publicly.
- India-scale mission : models for a billion people in their own languages.
- Compute that’s rare to fine : dedicated multi-node H200 training under a national grant.
- Small senior team: you work with the people who own the recipe.
- A live deployment path : Government institutes, EkaCare's doctors and patients use what you ship.
Full-Time Employee Benefits
- Medical Insurance & Accidental Insurance
- Maternity & Paternity Benefits
- PF, Gratuity, & Leave Encashment
- Salary Advance Policy
Browse similar roles
Is this role actually a fit for you?
hirly answers with a score and its reasoning, then writes the resume and cover letter if you decide to go for it.
Score it against my resume