Geico
Senior Software Engineer, AI Agent Platform
Palo Alto, CA · New York City, NY · Bethesda, MD · Seattle, WA
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.3M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →hirly's read of this role
- Role family
- Engineering
- Seniority
- Senior
- Country
- US
- Work mode
- On-site / unstated
- First seen by hirly
- 1 Oct 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
Why Join GEICO?
At GEICO, we offer a rewarding career where your ambitions are met with endless possibilities.
Every day we honor our iconic brand by offering quality coverage to millions of customers and being there when they need us most. We thrive on relentless innovation to exceed our customers' expectations while making a real impact on local communities nationwide.
Founded in 1936, GEICO is a member of the Berkshire Hathaway family of companies and one of the largest auto insurers in the United States. When you join our company, we want you to feel valued, supported, and proud to work here. That's why we offer the GEICO Pledge: Great Company, Great Culture, Great Rewards, and Great Careers.
Senior Software Engineer, AI Agent Platform
GEICO
Why Join GEICO?
GEICO is transforming how AI is built and deployed across the enterprise. As one of the largest insurers in the United States, we are investing heavily in next-generation AI platforms that empower more than 30,000 associates and enhance experiences for millions of customers.
We are looking for a Senior Software Engineer with strong backend engineering experience to help build GEICO's enterprise AI Agent Platform. The team is building several core systems: an agent harness that runs agentic workflows durably in production, a retrieval-augmented generation (RAG) platform that grounds agents in enterprise knowledge, an insight engine that turns agent execution data and signals from related systems into live domain knowledge and context for both people and agents, an eval harness that measures agent quality and catches regressions, and a skills marketplace where teams publish and reuse agent capabilities. Each of them is a distributed systems problem at its core. We care about your experience building durable, scalable systems as well as your experience in AI/ML.
The Opportunity
As a Senior Software Engineer on the AI Agent Platform team, you will design, build, and own significant components of the core systems that power use-case-agnostic agentic workflows across the enterprise, from claims and underwriting to internal associate tools. You'll work closely with Sr. Staff and Staff engineers who set technical direction. This is platform work: generalized services and capabilities that many teams build on, not a single agent application.
We welcome candidates from application, platform, or infrastructure backgrounds. What matters is that you have designed and built systems that stay correct, available, and operable under real production load, that you've helped run them (on call, incidents, and all), and that you're comfortable working across the full backend stack. Hands-on experience building GenAI applications or agentic harnesses is a strong plus, and the foundation is solid engineering judgment.
What You Will Do
Own and Collaborate
Own the design and delivery of significant platform components and features end to end, from technical design and API definition through implementation, rollout, and iteration.
Write clear design docs, contribute actively to design and code reviews, and make sound tradeoffs within your area of ownership.
Mentor junior and mid-level engineers, and help raise the engineering bar on the team.
Collaborate with product managers, data scientists, and consuming teams to understand workflow needs and translate them into platform features.
Build
Depending on your strengths, you'll build and own work in areas such as:
Agent harness: the durable execution engine for agent workflows, with state management, checkpointing, retries, idempotency, and recovery for long-running, multi-step, tool-calling workflows.
Agent harness integrations: services, APIs, and integration contracts (including protocols like MCP) that let many teams connect agents to internal services, data sources, and external systems consistently.
RAG platform: shared ingestion, chunking, embedding, indexing, and retrieval services that ground agents in enterprise documents and data, with access controls, freshness guarantees, and low-latency retrieval at scale.
Insight engine: pipelines and services that collect data from agent executions and related enterprise systems and turn it into continuously updated domain knowledge and agent context, served to people through dashboards and reports and to agents through retrieval and tool interfaces. Where the RAG platform grounds agents in relatively static documents, the insight engine captures what is changing: emerging patterns, outcomes, and operational signals.
Eval harness: high-throughput pipelines for running offline and online evaluations, simulating scenarios, capturing execution traces, detecting regressions, and storing large volumes of execution data, as a shared framework every team can use.
Skills marketplace: a multi-tenant service where teams publish, version, discover, and reuse agent skills and tools, backed by permissions, security review, and governance.
Reliability, isolation, and cost controls: sandboxing, rate limiting, quotas, SLOs, and capacity planning at enterprise scale.
Work across the stack as needed, from storage and data modeling to services, deployment infrastructure, and the developer-facing surfaces (SDKs, CLIs, portals) teams use to build on the platform.
Operate
Own the production health of the components you build: help define SLIs and SLOs, and build the dashboards and alerts that tell the team when something is wrong before users do.
Build observability in from the start, with metrics, structured logging, and distributed tracing that make failures in the agent harness and in complex, multi-step agent workflows quick to diagnose.
Design for scale and resilience through load testing, capacity planning, graceful degradation, and failure-mode analysis, so the platform holds up as adoption grows across the enterprise.
Participate in the on-call rotation, respond to incidents, and contribute to blameless postmortems that result in durable fixes rather than repeat incidents.
Contribute to operational excellence through runbooks, alert hygiene, and safe deployment practices (canaries, feature flags, rollbacks).
Minimum Qualifications
5+ years of professional software engineering experience building and operating production backend systems at scale.
Experience owning the design and delivery of services or major features end to end.
Strong proficiency in one or more of Python, Java, Go, or comparable languages.
Strong distributed systems fundamentals, including concurrency, consistency, fault tolerance, data modeling, API design, and scalability.
Experience operating production services with CI/CD, automated testing, and Kubernetes or other container platforms.
Hands-on operational experience supporting production systems, including observability, monitoring and alerting, on-call participation, incident response, and postmortems.
Experience scaling systems or improving reliability, such as meeting SLOs, performance tuning, capacity planning, or eliminating recurring failure modes.
Background in application, platform, or infrastructure engineering; all are relevant.
Preferred Qualifications
Experience with durable execution or workflow orchestration systems (e.g., Temporal, Restate, DBOS, Azure Durable Task, AWS Step Functions) or event-driven architectures (e.g., Kafka).
Experience building internal developer platforms, multi-tenant services, or SDKs used by many teams.
Experience building GenAI applications and agentic harnesses, using frontier and open-weight models (e.g., GPT, Claude, Gemini, Llama, Qwen) and frameworks such as LangGraph, Microsoft Agent Framework, OpenAI Agents SDK, or Claude Agent SDK.
Experience building RAG systems at scale, including ingestion pipelines, hybrid retrieval (dense plus keyword search), reranking, and permission-aware retrieval, using engines such as Azure AI Search, OpenSearch, or Vespa.
Exp
Similar jobs
- Senior Frontend Software Engineer - Global Commercial Services TechnologyAmerican Express · Seattle, WA, United States; Phoenix, AZ, United States; New York, NY, United States; WA, United States; Sunrise, FL, United StatesFirst seen today
- Sr Software Engineer - AndroidUber · San Francisco, CA, United States; New York City, NY, United States; Sunnyvale, CA, United StatesFirst seen today
- Senior Software EngineerClean Harbors · Norwell, MA, United StatesFirst seen today
- Senior Software Engineer, Core InfrastructureOracle · Seattle, WA, United States; Austin, TX, United States; Nashville, TN, United StatesFirst seen today
- Senior Platform Software EngineerOracle · Nashville, TN, United StatesFirst seen today
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job