This role has closed. ixigo has taken the posting down.
hirly last saw it live on 29 September 2026. See similar open roles below, or browse the live board.
ixigo
Research Engineer - Agent Intelligence & Evaluation
New Delhi, DL, India
Similar open jobs
- AI Research Engineer (Kernel & Inference Optimization) - 100% Remote WorldwideTether Operations Limited · IndiaFirst seen todayremote
- ML Research Engineer, SpeechBlue Machines AI · IndiaFirst seen yesterday
- Research Engineer, Applied MLTriomics · India OfficeFirst seen yesterday
- Research EngineerAtlys · Delhi HQFirst seen yesterday
- Security Research EngineerCowbell · Pune, Maharashtra, IndiaFirst seen 2d ago
- Research Engineer - Inspection TechnologiesGevernova · BengaluruFirst seen 2d ago
- Patch Research EngineerQualys · PuneFirst seen 2d ago
- Security Research EngineerQualys · PuneFirst seen 2d ago
- Scientific Computing & Research Engineering Expert – Earth ScienceseDataBae · IndiaFirst seen 3d ago
- Scientific Computing & Research Engineering Expert - AI Evaluation ProjecteDataBae · IndiaFirst seen 3d ago
- Security Research EngineerCowbellcyber · Pune (India)First seen 4d ago
- Robotics Research Engineer - L2Perceptyne Technologies Private Limited · Hyderabad, IndiaFirst seen 6d ago
- ML Research Engineer - Pre-training (LLMs)Eka.care · Bengaluru, IndiaFirst seen 6d ago
- ML Research Engineer - Post-training (LLMs)Eka.care · Bengaluru, IndiaFirst seen 6d ago
- Research EngineerMonotype · NoidaFirst seen 10d ago
hirly's read of this role
- Seniority
- Mid level
- Country
- IN
- Work mode
- On-site / unstated
- First seen by hirly
- 23 Sept 2026
Derived automatically from the posting.
the posting
Voice agents fail in ways traditional software doesn't. ASR confidence drops on an accent and a tool call misfires. Latency breaks turn-taking and the LLM hallucinates a policy. A model swap silently regresses production and nobody catches it for a week. 
We're building self-healing voice agents for enterprise customer support. This role owns the intelligence layer: the evals that catch failures before shipping, the observability that traces them across the pipeline, and the feedback loops that let agents fix themselves 
  What you'll own 
Evaluation infrastructure. Audio-native metrics for barge-in, prosody, and turn-taking. Adversarial datasets across accents and edge cases. LLM-as-judge rubrics for task success, tool-use correctness, and recovery. 
Observability across the pipeline. Tracing that correlates audio, STT, LLM reasoning, tool calls, and TTS to a single conversation. Analysis and alerting that surfaces cascade failures instead of hiding them. 
Self-improvement systems. Mine production traces for failure patterns, generate targeted training or prompt data, validate fixes with adversarial replay, and guardrail against regressions.
 
Who we're looking for 
3 to 5 years in ML engineering, research engineering, or applied research. Strong Python and modern ML tooling. Depth in at least two of: speech and audio models, LLM agent systems, and eval or observability infrastructure. 
You've shipped something non-trivial where research met production. You read papers, spot when a benchmark measures the wrong thing, and translate ideas from Interspeech, ACL, or NeurIPS into systems that run on real traffic. Publications welcome, not required. 
 
Nice to have 
Real-time systems or telephony experience. Work on RLHF, DPO, or synthetic data pipelines. Familiarity with enterprise deployment (SOC 2, PII, data residency). 
 
What you'll get 
Senior seat on a small team where research and production aren't separate orgs. Real enterprise conversation data under proper governance. Meaningful equity, autonomy over tooling, and support to publish. 
Candidates are responsible for safeguarding sensitive company data against unauthorized access, use, or disclosure, and for reporting any suspected security incidents in line with the organization's ISMS (Information Security Management System) policies and procedures.
We’re building self-healing voice agents for enterprise customer support within ixigo. The system has to know when it’s failing, why it’s failing, and how to fix itself before a human notices. This fellowship sits at the intelligence layer behind that work.