hirly

Metaforms

Senior AI Engineer

Bengaluru

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Metaforms first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.4M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

Apply from your AI assistant

Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.

Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.

hirly's read of this role

Seniority
Senior
Country
IN
Work mode
On-site / unstated
First seen by hirly
28 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

About Metaforms

Market research runs on 30-year-old survey platforms and armies of specialists hand-coding questionnaires in proprietary languages. Metaforms is the agent layer that does that work. Every survey is a program — full of skip logic, piping, quotas, and loops — and a single wrong number in a client report is unrecoverable. Our AI agents write production survey code, QA live deployments, process and clean large structured datasets, configure analysis, and generate client-ready reports, so agencies like Dynata, Savanta, and Borderless Access ship more projects with far less friction.

1,000+ surveys processed monthly

Serving Fortune 500 companies across the globe

Rapid month-over-month growth

We’re Series A funded and scaling fast, aggressively growing our AI engineering team to build the next generation of production-grade AI agent systems.

The Role

We’re hiring a Senior AI Engineer to own the design, development, and continuous improvement of the AI agent systems that power modern research operations.

This is a high-ownership, high-impact role at the intersection of applied AI and systems engineering. You’ll work on genuinely hard problems: agent reliability at scale, long-context handling, cascading error mitigation, and evaluation infrastructure — like codegen agents that write in proprietary DSLs, computer-use agents that QA live deployments, data agents that clean tabular exports and configure multi-step analysis, and evals for outputs where “correct” is genuinely ambiguous. And you’ll do it on a team that ships fast and treats quality as non-negotiable.

What You’ll Own

Agent Harness and Architecture

Own the agent harness our production agents run on — the loop where agents plan, use tools, check their work, and recover from failures

Lead research and implementation for long-context handling and cascading-error challenges in multi-step agent pipelines

Drive context engineering strategy and experimentation frameworks across the team

Evaluation and Production Monitoring

Define structured rubrics for evaluating AI outputs on nuanced, ambiguous research tasks

Build continuous monitoring, tracing, and failure-mode analysis for agents in production — including the loop that turns production failures into test cases

Create tooling that lets domain experts refine and evolve the skill files, eval sets, and knowledge bases our agents consume

Reliability for High-Stakes Outputs

Build eval suites — regression sets, golden datasets, LLM-as-judge pipelines — that catch regressions before deploy

Develop evaluation datasets for DSLs, structured data transforms, and computed outputs to systematically find and close model weaknesses

Design human-in-the-loop and review workflows for outputs where a single wrong number in a client report is unrecoverable

What We’re Looking For

Must-Have

Built and operated agentic systems in production — multi-step pipelines, tool use, codegen, computer-use, or data and reporting agents — not just prototypes

4+ years of engineering experience, with at least 1 year focused on LLM/agent systems in production

Deep hands-on experience with frontier model APIs (Anthropic, OpenAI, Gemini), evaluation frameworks, and AI system optimization

Strong Python skills; Go or TypeScript a plus

Solid grasp of context engineering and evaluation methodology

Strong instincts for debugging complex, non-deterministic system failures

High ownership: you drive problems to resolution independently and pull others in when it matters

Nice to Have

Experience with LLM observability and eval tooling (Braintrust, Langfuse, LangSmith, Weave, promptfoo, or in-house equivalents)

Background in semantic parsing, DSLs, or structured-output generation

Prior work on computer-use or browser agents

Experience with human-in-the-loop agent workflows where proposals are reviewed before apply, or agents over large structured datasets

Why Metaforms

Work at the frontier of production AI: systems handling 1,000+ research projects a month, with the reliability bar that implies

A small, senior team where your decisions carry real architectural weight

Zero-bureaucracy culture: high autonomy, fast feedback loops, direct access to leadership

Well-funded and financially stable, with a clear roadmap and the runway to execute on it

Benefits

Full family health insurance

$1,000 USD annual learning and development budget

Dedicated mentor and coaching support

Free snacks and dinner at the office

Original posting on Metaforms's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job