hirly

Geico

Sr. Staff AI Engineer, AI Agent Platform (Agent & Evaluation Harness)

New York City, NY · Palo Alto, CA · Bethesda, MD · Seattle, WA

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Geico first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.3M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

hirly's read of this role

Seniority
Lead / management
Country
US
Work mode
On-site / unstated
First seen by hirly
1 Oct 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

Why Join GEICO?

At GEICO, we offer a rewarding career where your ambitions are met with endless possibilities.

Every day we honor our iconic brand by offering quality coverage to millions of customers and being there when they need us most. We thrive on relentless innovation to exceed our customers' expectations while making a real impact on local communities nationwide.

Founded in 1936, GEICO is a member of the Berkshire Hathaway family of companies and one of the largest auto insurers in the United States. When you join our company, we want you to feel valued, supported, and proud to work here. That's why we offer the GEICO Pledge: Great Company, Great Culture, Great Rewards, and Great Careers.

Sr. Staff AI Engineer, AI Agent Platform (Agent & Evaluation Harness)

Why Join GEICO?

GEICO is transforming how AI is built and deployed across the enterprise. As one of the largest insurers in the United States, we are investing heavily in next-generation AI platforms that empower more than 30,000 associates and enhance experiences for millions of customers.

We are looking for a Sr. Staff AI Engineer to push the frontier of what GEICO's AI Agent Platform can do, with a focus on two core systems: the agent harness that determines how agents plan, reason, use tools, and manage context, and the eval harness that tells us, with evidence, whether they are getting better. The field is moving faster than any job description can capture. Techniques that define best practice today, such as context engineering, MCP-based tool integration, skills, sub-agent orchestration, and LLM-as-judge evaluation, barely existed a few years ago and will look different a year from now. We are looking for someone who thrives in that environment and can move quickly from a new idea to a working prototype to a data-backed conclusion about whether it delivers value.

The Opportunity

The AI Agent Platform team builds use-case-agnostic infrastructure that many business workflows across GEICO build on, from claims and underwriting to internal associate tools. In this role, you'll own the intelligence layer of that platform: the agent design patterns, context strategies, and evaluation methodology that make agents on the platform reliably good at their jobs.

You'll work alongside backend platform engineers who own durable execution, scaling, and operations. Your focus is agent behavior and quality, and the science of measuring and improving it, grounded in a solid understanding of distributed systems so that what works in an experiment also works at enterprise scale. Success here comes from experimentation: forming hypotheses, designing experiments, reading traces, and iterating until performance meets the bar for real business workflows.

The model is only the starting point. Your job is to find out what actually works, prove it, and make it reusable.

What You Will Do

Agent Harness

Design the core agent loop, including planning, reasoning, tool selection and calling, memory, and error recovery, as reusable patterns that generalize across many business workflows.

Develop context engineering strategies: prompt architecture, context window management, summarization and compaction, memory, and retrieval (RAG) integration.

Define how agents connect to tools and enterprise systems through the Model Context Protocol (MCP), including tool design, server patterns, and the conventions that make tools reliable for agents to use.

Design how agents discover, load, and apply skills (packaged instructions, scripts, and resources for specific tasks) so capabilities built once can be reused across many workflows.

Explore sub-agent and multi-agent patterns, and decide with evidence what belongs in the platform.

Evaluate new frontier and open-weight models as they are released, and understand their tradeoffs in capability, latency, and cost for platform use.

Evaluation Harness

Design evaluation methodology: task-level metrics, rubrics, golden datasets, LLM-as-judge graders, and human review, with automated graders calibrated against human judgment.

Build the eval harness that lets any team define, run, and compare evaluations for their agents, spanning offline benchmarks, simulated users and environments, and online quality monitoring.

Make eval results trustworthy by accounting for variance, statistical significance, dataset contamination, and grader bias.

Diagnose failure modes by analyzing traces and eval results, and turn qualitative observations into measurable hypotheses.

Experiment and Iterate

Run fast, well-designed experiments on prompts, architectures, models, and tools, owning the loop from hypothesis to result to decision.

Iterate until agent quality meets the bar for production business workflows, and help define what that bar should be.

Track the research frontier and industry practice, bring the best ideas into GEICO quickly, and filter out the hype.

Share what you learn through write-ups, internal talks, reusable patterns, and playbooks that raise the level of GenAI practice across the company.

Lead

Set technical direction for agent design and evaluation across the AI Agent Platform, in partnership with senior AI and engineering leaders.

Mentor engineers and data scientists on GenAI engineering practice and experimental rigor.

Partner with platform engineers, product managers, and business teams to turn workflow needs into measurable quality targets and generalized platform capabilities.

How You Work

Intellectual curiosity. You want to understand why something works, not just whether it does.

Speed to learn. A new model, paper, protocol, or framework drops, and you're productive with it within days.

Adaptability. You change direction when the evidence says so, and hold strong opinions loosely.

Empirical rigor. You trust data over intuition, and your intuition is well calibrated because of it.

Bias to action. You'd rather build a prototype and measure it than debate it in a meeting.

Minimum Qualifications

10+ years of experience in software engineering, ML engineering, or applied science, with strong Python proficiency.

Tech-lead experience: owned the design and delivery of significant systems involving multiple engineers or teams.

Significant hands-on experience building GenAI applications and agentic systems with LLMs such as GPT, Claude, Llama, or Qwen, taken beyond prototype into real use.

Demonstrated experience designing and running experiments to improve LLM or agent performance, including defining metrics, building evaluation sets, analyzing results, and iterating.

Deep practical understanding of prompt and context engineering, tool calling, RAG, and agent architectures, including hands-on experience with MCP.

Solid distributed systems fundamentals, including concurrency, fault tolerance, latency and throughput tradeoffs, and API design, with experience building or working closely with production services at scale.

A track record of quickly learning and applying new techniques in a fast-moving field.

Demonstrated technical leadership influencing direction across teams.

Preferred Qualifications

Experience building evaluation frameworks or benchmarks for LLMs or agents, including LLM-as-judge methods and human annotation pipelines, using tools such as Inspect AI, Braintrust, DeepEval, or RAG.

Experience building MCP servers and tools, or authoring skills using the Agent Skills standard (or similar packaged agent capabilities) used across teams.

Experience with agent frameworks such as LangGraph, Microsoft Agent Framework, OpenAI Agents SDK, Claude Agent SDK, or Google ADK, and eval and observability tools such as Langfuse, LangSmith, Braintrust, or Arize Phoenix.

Background in ML or applied science, such as statistics and experimental design, fine-tuning, reinforcement fine-tuning (e.g., RLHF, RLVR), or model evaluation.

Experience with cloud AI platforms such a

Original posting on Geico's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job