hirly

Geico

Senior Staff Engineer - SRE - Incident Prevention / Post Incident Correction of Errors

Bethesda, MD · Palo Alto, CA · Richardson, TX · Seattle, WA

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Geico first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.3M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

hirly's read of this role

Role family
Engineering
Seniority
Lead / management
Country
US
Work mode
On-site / unstated
First seen by hirly
30 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

Why Join GEICO?

At GEICO, we offer a rewarding career where your ambitions are met with endless possibilities.

Every day we honor our iconic brand by offering quality coverage to millions of customers and being there when they need us most. We thrive on relentless innovation to exceed our customers' expectations while making a real impact on local communities nationwide.

Founded in 1936, GEICO is a member of the Berkshire Hathaway family of companies and one of the largest auto insurers in the United States. When you join our company, we want you to feel valued, supported, and proud to work here. That's why we offer the GEICO Pledge: Great Company, Great Culture, Great Rewards, and Great Careers.

Senior Staff Engineer - SRE - Incident Prevention / Post Incident Corre cti on of Errors

Position Summary

Correction or Error’s / Post Incident Review is at the core of GEICO’s culture for improving the reliability and availability of our applications. We are rethinking how we do CoEs at scale and the tooling we need to enable GEICO’s engineering teams to learn from the incidents and identify and eliminate patterns leading to them.

GEICO is seeking an experienced SRE Software Engineer with a passion for building high-performance, low-maintenance, zero-downtime complex distributed platforms and applications. You will help drive our transformation to a tech organization with engineering excellence and site reliability as its mission, while co-creating a culture of psychological safety and continuous improvement.

This role focuses on improving the Correction of Error (COE) tooling and process across GEICO. It is a hands-on technical leadership role focused on better COEs, stronger root cause analysis, fewer repeat incidents, and building the tools and automation that help all GEICO engineering teams learn from incidents and act on that learning.

Position Description

Success in this role requires strong technical depth and equally strong process leadership. The right candidate can go deep on incident analysis, system behavior, and architecture, while also improving how teams run COEs, learn from incidents, and turn those lessons into engineering improvements.

Why This Role Is Different

You are shaping how GEICO learns from incidents and turns that learning into better engineering practices.

Your work directly impacts root cause quality, repeat incidents, availability, MTTR, and operational confidence.

This role blends deep technical understanding, hands on execution and ability to build SW with real time incident leadership, platform and process

Position Responsibilities

As a Senior Staff Engineer, you will:

Design, develop and operate automation, self-service tools, dashboards, and data pipelines that automate and scale COE workflows and reduce manual tracking and follow-ups.

Run and moderate weekly GEICO-wide COE presentation sessions for qualified high-severity incidents, ensuring the right owners are prepared to present.

Lead and improve the COE process across the engineering organization.

Provide technical leadership in system design, architecture, and hands-on engineering for COE improvements, automation, and incident tooling.

Coach application engineering teams so they can identify the true root cause, explain incidents clearly, and produce clear, complete, high-quality COEs.

Provide technical leadership for root cause analysis across distributed systems using logs, metrics, traces, and observability data.

Identify gaps exposed by incidents, such as missing alerts, weak monitoring, incomplete runbooks, poor testing, or bypassed deployment controls.

Look across incidents to find repeat patterns, connect lessons across COEs, and share insights that improve prevention and reliability.

Partner with application engineering teams, platform teams, SRE, and operational stakeholders to drive accountability and continuous improvement, by driving improvements in their technology strategy and roadmaps.

We have adopted “You Build it You Run it strategy.

All our senior technologists take an active role in leading and managing high-severity incidents requiring strong technical judgment, clear communication, and calm execution under pressure.

All our engineers have on-call responsibilities as part of a 24x7 rotation supporting incident response and production support for mission-critical platforms and processes they build and operate.

Qualifications

Hands on proficiency in multiple languages, including Go, Java, Python, C# for building production-grade full stack applications on Kubernetes and serverless technologies (KNative) in Azure and AWS.

Experience with SQL and NoSQL technologies

Experience with building and using data pipelines analytics, and dashboards for operational metrics, trends, and KPIs using technologies such as Spark, Trino, Grafana, Superset, PowerBI

Experience with OpenTelemetry and observability platforms such as Grafana, Datadog, Splunk, Azure Monitor.

Experience with incident management platforms such as PagerDuty.

Proficiency with AI assisted development processes and tools such as Claude Code, Cursor and GitHub Copilot.

Experience improving incident, post incident review, or reliability processes at scale through automation, data, and cross-team influence.

Deep incident forensics and root cause analysis skills, with the ability to raise COE quality through clear action items and follow-through across teams.

Strong understanding of observability, reliability engineering, incident management, and post-incident improvement practices.

Experience supporting incident response and high-severity production incidents in complex environments.

Strong software engineering fundamentals and system design skills, with experience building reliable production systems at scale.

Ability to lead technical design and architecture decisions in complex distributed systems.

Strong communication skills and the ability to coach engineering teams and present findings clearly to leadership.

Experience

10+ years of professional software engineering experience, preferably in platform engineering, reliability engineering, backend engineering, distributed systems, or operational tooling.

8+ years of experience with architecture, design, system reliability, scalability, and technical leadership for production systems.

6+ years of experience with open-source frameworks, modern engineering practices, or platform technologies.

4+ years of experience with Azure, AWS, GCP, or another cloud service provider, or equivalent experience in complex hybrid environments.

Demonstrated ownership of mission-critical systems operating in 24x7 production environments.

Education

Bachelor's degree in Computer Science, Information Systems, or equivalent education or work experience.

Additional Job Requirements

Ability to influence engineering outcomes across teams in complex organizations.

Must be able to communicate in a clear, concise, professional oral and written manner with customers, clients, co-workers, leadership, and other employees of the organization.

Must be able to perform effectively under pressure and in stressful situations, including during production support and high-severity incident response.

Must be able to participate in a 24x7 on-call rotation for incident response and production support of mission-critical platforms.

Annual Salary

$110,000.00 - $260,000.00 The above annual salary range is a general guideline. Multiple factors are taken into consideration to arrive at the final hourly rate/ annual salary to be offered to the selected candidate. Factors include, but are not limited to, the scope and responsibilities of the role, the selected candidate’s work experience, education and training, the work location as well as market and business considerations.

At this time, GEICO will not sponsor a new applicant for employment author

Original posting on Geico's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job