hirly

Haparz Pvt Ltd

Senior/Staff AI Evaluation Engineer

India

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Haparz Pvt Ltd first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.3M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

hirly's read of this role

Seniority
Lead / management
Country
IN
Work mode
On-site / unstated
First seen by hirly
25 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

Job Description

Role: Senior/Staff AI Evaluation Engineer

Experience: 6+ Years

Location: Remote (Pan India)

Mode: Hyparz Payroll

About the Role

We are looking for a Senior/Staff AI Evaluation Engineer to lead the benchmarking, validation, reliability, and safety evaluation of next-generation Agentic AI platforms .

This role sits at the intersection of Quality Engineering, Software Engineering, and AI/ML . You will build automated evaluation frameworks, benchmark AI behavior, validate LLM/RAG/agentic workflows, and establish measurable quality standards for production AI systems.

What You'll Do

  • Design and implement automated evaluation frameworks for LLM, RAG, and agent-based applications .
  • Build and maintain golden datasets, benchmark datasets, test datasets, and regression suites for AI evaluation.
  • Develop measurable evaluation criteria for accuracy, relevance, consistency, safety, reliability, and agent behavior.
  • Perform adversarial testing to identify hallucinations, prompt vulnerabilities, unsafe behavior, and edge cases.
  • Evaluate LangGraph-based and multi-agent systems at the node, state-transition, routing, and workflow levels.
  • Use LangSmith for tracing, debugging, experiments, datasets, and evaluation workflows.
  • Work extensively with LangChain and LangGraph , including sub-graphs and conditional routing.
  • Build Python-based automation for functional, regression, integration, and end-to-end AI testing.
  • Evaluate prompts, embeddings, vector search, RAG pipelines, tool calling, and agentic workflows.
  • Validate AI systems against safety, security, privacy, and Responsible AI expectations.
  • Integrate evaluation and regression testing into GitHub-based CI/CD workflows .
  • Work with AWS and Amazon Bedrock to validate AI-powered application architectures.
  • Perform API and backend validation using REST APIs, JSON, SQL, and modern application architectures.
  • Collaborate with Engineering, Product, Data, and Quality teams to investigate failures and drive improvements.
  • Take ownership of ambiguous AI quality problems and convert them into repeatable, measurable evaluation approaches.

What We're Looking For

  • 6+ years of experience in software engineering, quality engineering, test automation, AI engineering, ML engineering, or a related technical discipline.
  • Demonstrable hands-on experience testing or evaluating LLM, NLP, ML, RAG, or agentic AI applications .
  • Strong Python development and automation experience.
  • Deep hands-on knowledge of LangChain, LangGraph, and LangSmith .
  • Experience evaluating LangGraph node execution, state transitions, conditional routing, sub-graphs, or multi-agent workflows .
  • Strong understanding of LLMs, prompt engineering, embeddings, vector databases/search, RAG, and AI agents.
  • Experience building golden datasets, benchmark datasets, or structured AI test datasets.
  • Experience with adversarial testing and AI safety evaluation.
  • Practical understanding of security, privacy, and Responsible AI evaluation .
  • Strong GitHub experience covering repositories, branching, pull requests, code reviews, and CI/CD.
  • Experience with Claude Code or comparable AI-assisted development tools .
  • Experience with AWS and Amazon Bedrock .
  • Strong REST API, JSON, and SQL knowledge.
  • Strong experience in automated, regression, integration, and end-to-end testing.
  • Excellent analytical, troubleshooting, communication, and problem-solving abilities

Skills:- Large Language Models (LLM), Retrieval Augmented Generation (RAG), AI Evaluation, Agentic AI and AWS Bedrock

Original posting on Haparz Pvt Ltd's site ↗

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job