hirly

Tihinsurance

Site Reliability Engineer

Charlotte NC - 600 S Tryon St.

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Tihinsurance first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.6M live jobs from 190,000+ employers in 200+ countries.

Tailor my resume for this job →

Apply from your AI assistant

Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.

Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.

hirly's read of this role

Role family
Engineering
Seniority
Mid level
Country
US
Work mode
On-site / unstated
First seen by hirly
16 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

The position is described below. If you want to apply, click the Apply button at the top or bottom of this page. You'll be required to create an account or sign in to an existing one.

If you have a disability and need assistance with the application, you can request a reasonable accommodation. Send an email to Accessibility (accommodation requests only; other inquiries won't receive a response).

Regular or Temporary:

Regular Language Fluency: English (Required)

Work Shift:

1st Shift (United States of America) Please review the following job description:

Lead Site Reliability & Environment Monitoring Engineer (Azure / Dynatrace / ServiceNow)

We are seeking a Lead Site Reliability & Environment Monitoring Engineer to establish and evolve our enterprise observability and monitoring strategy across cloud and application platforms. This is a full-time leadership role responsible for owning monitoring design, driving platform decisions, and guiding engineering teams toward modern SRE practices.

This individual will act as the technical authority for monitoring and alerting, shaping how signals from Dynatrace flow into ServiceNow and enterprise messaging/paging platforms, and enabling a shift toward automated, intelligent, and self-healing operations.

Location: This role is hybrid based in Charlotte, NC.

Key Responsibilities

Strategic Leadership & Decision-Making

Define and own the enterprise monitoring and SRE observability strategy

Serve as the subject matter expert for Dynatrace, ServiceNow integration, and alerting architecture

Evaluate and recommend tooling, integration patterns, and platform direction

Drive decisions on alerting philosophy, noise reduction, and signal quality improvement

Platform Ownership & Architecture

Architect and standardize end-to-end monitoring and SRE pipelines:

Dynatrace → ServiceNow incident lifecycle

Alert correlation, deduplication, and prioritization

Integration with paging systems (PagerDuty, SMS, voice, Teams)

Establish best practices for:

Event ingestion and enrichment

Incident routing and automated assignment

Integration with CMDB and service mapping

Site Reliability Engineering (SRE) Leadership

Lead adoption of SRE principles, including:

SLIs, SLOs, and error budgets

Reliability engineering practices across services

Proactive monitoring and resilience design

Champion a shift from reactive operations to proactive reliability engineering

Influence application and platform teams to build observable, resilient systems by design

Automation & Self-Healing Enablement

Drive development of automated remediation and self-healing capabilities

Leverage Dynatrace workflows, Azure services, and automation frameworks to:

Reduce manual incident handling

Eliminate repeatable operational tasks

Minimize unnecessary paging

ServiceNow & Observability Integration Leadership

Own integration between Dynatrace and ServiceNow ITSM/ITOM, including:

Incident, Event Management, and CMDB alignment

Service mapping and dependency visibility

Governance for application/service tagging

Define standards for:

Automated incident creation and resolution

Priority assignment and routing logic

Monitoring-to-ITSM data synchronization

Team Leadership & Cross-Functional Influence

Provide technical leadership and mentorship across SRE, platform, and application teams

Act as a central point of coordination between engineering, cloud, and ITSM teams

Lead workshops and working sessions to:

Drive monitoring standardization

Align teams on reliability practices

Influence upstream architectural decisions

Operational Excellence

Establish KPIs and drive improvement in:

Incident response and resolution times

Alert quality and paging effectiveness

Monitoring coverage across critical services

Provide leadership with clear visibility into service health and reliability trends

Required Qualifications

7+ years in Site Reliability Engineering, monitoring, or production engineering

Proven experience in a technical leadership or lead engineer role

Deep hands-on experience with:

Dynatrace (or equivalent observability platforms)

Microsoft Azure (IaaS, PaaS, networking, identity)

ServiceNow ITSM / ITOM (incident, event management, CMDB)

Demonstrated ability to:

Design and lead enterprise monitoring/SRE architectures

Drive platform and tooling decisions

Integrate observability, ITSM, and paging solutions

Preferred Qualifications

Experience leading SRE or observability transformation initiatives

Strong expertise with Dynatrace–ServiceNow integrations

Experience modernizing or consolidating paging/on-call tooling

Familiarity with:

Azure-based SRE tooling or AI-assisted operations

Automation frameworks (GitHub Actions, Runbooks, etc.)

Infrastructure as Code (Terraform, ARM, Bicep)

Success Metrics

Reduction in alert noise and unnecessary paging

Improved incident routing accuracy and MTTR

Increased adoption of self-healing and automated workflows

Strong alignment between monitoring, CMDB, and service ownership

Enterprise-wide adoption of SRE and monitoring standards

General Description of Available Benefits for Eligible Employees of CRC Group: At CRC Group, we're committed to supporting every aspect of teammates' well-being – physical, emotional, financial, social, and professional. Our best-in-class benefits program is designed to care for the whole you, offering a wide range of coverage and support. Eligible full-time teammates enjoy access to medical, dental, vision, life, disability, and AD&D insurance; tax-advantaged savings accounts; and a 401(k) plan with company match. CRC Group also offers generous paid time off programs, including company holidays, vacation and sick days, new parent leave, and more. Eligible positions may also qualify for restricted stock units and/or a deferred compensation plan.

CRC Group supports a diverse workforce and is an Equal Opportunity Employer that does not discriminate against individuals on the basis of race, gender, color, religion, citizenship or national origin, age, sexual orientation, gender identity, disability, veteran status or other classification protected by law. CRC Group is a Drug Free Workplace.

EEO is the Law Pay Transparency Nondiscrimination Provision E-Verify

Original posting on Tihinsurance's site ↗

Listed on hirly, a job board. hirly is not the employer: Tihinsurance is hiring for this role.

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job