Haparz Pvt Ltd
Senior/Staff AI Evaluation Engineer
India
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.3M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →hirly's read of this role
- Seniority
- Lead / management
- Country
- IN
- Work mode
- On-site / unstated
- First seen by hirly
- 25 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
Job Description
Role: Senior/Staff AI Evaluation Engineer
Experience: 6+ Years
Location: Remote (Pan India)
Mode: Hyparz Payroll
About the Role
We are looking for a Senior/Staff AI Evaluation Engineer to lead the benchmarking, validation, reliability, and safety evaluation of next-generation Agentic AI platforms .
This role sits at the intersection of Quality Engineering, Software Engineering, and AI/ML . You will build automated evaluation frameworks, benchmark AI behavior, validate LLM/RAG/agentic workflows, and establish measurable quality standards for production AI systems.
What You'll Do
- Design and implement automated evaluation frameworks for LLM, RAG, and agent-based applications .
- Build and maintain golden datasets, benchmark datasets, test datasets, and regression suites for AI evaluation.
- Develop measurable evaluation criteria for accuracy, relevance, consistency, safety, reliability, and agent behavior.
- Perform adversarial testing to identify hallucinations, prompt vulnerabilities, unsafe behavior, and edge cases.
- Evaluate LangGraph-based and multi-agent systems at the node, state-transition, routing, and workflow levels.
- Use LangSmith for tracing, debugging, experiments, datasets, and evaluation workflows.
- Work extensively with LangChain and LangGraph , including sub-graphs and conditional routing.
- Build Python-based automation for functional, regression, integration, and end-to-end AI testing.
- Evaluate prompts, embeddings, vector search, RAG pipelines, tool calling, and agentic workflows.
- Validate AI systems against safety, security, privacy, and Responsible AI expectations.
- Integrate evaluation and regression testing into GitHub-based CI/CD workflows .
- Work with AWS and Amazon Bedrock to validate AI-powered application architectures.
- Perform API and backend validation using REST APIs, JSON, SQL, and modern application architectures.
- Collaborate with Engineering, Product, Data, and Quality teams to investigate failures and drive improvements.
- Take ownership of ambiguous AI quality problems and convert them into repeatable, measurable evaluation approaches.
What We're Looking For
- 6+ years of experience in software engineering, quality engineering, test automation, AI engineering, ML engineering, or a related technical discipline.
- Demonstrable hands-on experience testing or evaluating LLM, NLP, ML, RAG, or agentic AI applications .
- Strong Python development and automation experience.
- Deep hands-on knowledge of LangChain, LangGraph, and LangSmith .
- Experience evaluating LangGraph node execution, state transitions, conditional routing, sub-graphs, or multi-agent workflows .
- Strong understanding of LLMs, prompt engineering, embeddings, vector databases/search, RAG, and AI agents.
- Experience building golden datasets, benchmark datasets, or structured AI test datasets.
- Experience with adversarial testing and AI safety evaluation.
- Practical understanding of security, privacy, and Responsible AI evaluation .
- Strong GitHub experience covering repositories, branching, pull requests, code reviews, and CI/CD.
- Experience with Claude Code or comparable AI-assisted development tools .
- Experience with AWS and Amazon Bedrock .
- Strong REST API, JSON, and SQL knowledge.
- Strong experience in automated, regression, integration, and end-to-end testing.
- Excellent analytical, troubleshooting, communication, and problem-solving abilities
Skills:- Large Language Models (LLM), Retrieval Augmented Generation (RAG), AI Evaluation, Agentic AI and AWS Bedrock
Similar jobs
- Lead Design Evaluation EngineerAnalogdevices · India, HyderabadFirst seen 7d ago
- Lead Design Evaluation EngineerAnalogdevices · India, HyderabadFirst seen 7d ago
- ML Model Evaluation EngineerTriomics · India OfficeFirst seen 2d ago
- Senior Test & Evaluation EngineerEmerson · PUNE, MAHARASHTRA, IndiaFirst seen 3d ago
- Agentic AI Evaluation EngineerEY · Kolkata, WB, INFirst seen 3d ago
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job