Tihinsurance
Site Reliability Engineer
Charlotte NC - 600 S Tryon St.
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.6M live jobs from 190,000+ employers in 200+ countries.
Tailor my resume for this job →Apply from your AI assistant
Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.
Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.
hirly's read of this role
- Role family
- Engineering
- Seniority
- Mid level
- Country
- US
- Work mode
- On-site / unstated
- First seen by hirly
- 16 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
The position is described below. If you want to apply, click the Apply button at the top or bottom of this page. You'll be required to create an account or sign in to an existing one.
If you have a disability and need assistance with the application, you can request a reasonable accommodation. Send an email to Accessibility (accommodation requests only; other inquiries won't receive a response).
Regular or Temporary:
Regular Language Fluency: English (Required)
Work Shift:
1st Shift (United States of America) Please review the following job description:
Lead Site Reliability & Environment Monitoring Engineer (Azure / Dynatrace / ServiceNow)
We are seeking a Lead Site Reliability & Environment Monitoring Engineer to establish and evolve our enterprise observability and monitoring strategy across cloud and application platforms. This is a full-time leadership role responsible for owning monitoring design, driving platform decisions, and guiding engineering teams toward modern SRE practices.
This individual will act as the technical authority for monitoring and alerting, shaping how signals from Dynatrace flow into ServiceNow and enterprise messaging/paging platforms, and enabling a shift toward automated, intelligent, and self-healing operations.
Location: This role is hybrid based in Charlotte, NC.
Key Responsibilities
Strategic Leadership & Decision-Making
Define and own the enterprise monitoring and SRE observability strategy
Serve as the subject matter expert for Dynatrace, ServiceNow integration, and alerting architecture
Evaluate and recommend tooling, integration patterns, and platform direction
Drive decisions on alerting philosophy, noise reduction, and signal quality improvement
Platform Ownership & Architecture
Architect and standardize end-to-end monitoring and SRE pipelines:
Dynatrace → ServiceNow incident lifecycle
Alert correlation, deduplication, and prioritization
Integration with paging systems (PagerDuty, SMS, voice, Teams)
Establish best practices for:
Event ingestion and enrichment
Incident routing and automated assignment
Integration with CMDB and service mapping
Site Reliability Engineering (SRE) Leadership
Lead adoption of SRE principles, including:
SLIs, SLOs, and error budgets
Reliability engineering practices across services
Proactive monitoring and resilience design
Champion a shift from reactive operations to proactive reliability engineering
Influence application and platform teams to build observable, resilient systems by design
Automation & Self-Healing Enablement
Drive development of automated remediation and self-healing capabilities
Leverage Dynatrace workflows, Azure services, and automation frameworks to:
Reduce manual incident handling
Eliminate repeatable operational tasks
Minimize unnecessary paging
ServiceNow & Observability Integration Leadership
Own integration between Dynatrace and ServiceNow ITSM/ITOM, including:
Incident, Event Management, and CMDB alignment
Service mapping and dependency visibility
Governance for application/service tagging
Define standards for:
Automated incident creation and resolution
Priority assignment and routing logic
Monitoring-to-ITSM data synchronization
Team Leadership & Cross-Functional Influence
Provide technical leadership and mentorship across SRE, platform, and application teams
Act as a central point of coordination between engineering, cloud, and ITSM teams
Lead workshops and working sessions to:
Drive monitoring standardization
Align teams on reliability practices
Influence upstream architectural decisions
Operational Excellence
Establish KPIs and drive improvement in:
Incident response and resolution times
Alert quality and paging effectiveness
Monitoring coverage across critical services
Provide leadership with clear visibility into service health and reliability trends
Required Qualifications
7+ years in Site Reliability Engineering, monitoring, or production engineering
Proven experience in a technical leadership or lead engineer role
Deep hands-on experience with:
Dynatrace (or equivalent observability platforms)
Microsoft Azure (IaaS, PaaS, networking, identity)
ServiceNow ITSM / ITOM (incident, event management, CMDB)
Demonstrated ability to:
Design and lead enterprise monitoring/SRE architectures
Drive platform and tooling decisions
Integrate observability, ITSM, and paging solutions
Preferred Qualifications
Experience leading SRE or observability transformation initiatives
Strong expertise with Dynatrace–ServiceNow integrations
Experience modernizing or consolidating paging/on-call tooling
Familiarity with:
Azure-based SRE tooling or AI-assisted operations
Automation frameworks (GitHub Actions, Runbooks, etc.)
Infrastructure as Code (Terraform, ARM, Bicep)
Success Metrics
Reduction in alert noise and unnecessary paging
Improved incident routing accuracy and MTTR
Increased adoption of self-healing and automated workflows
Strong alignment between monitoring, CMDB, and service ownership
Enterprise-wide adoption of SRE and monitoring standards
General Description of Available Benefits for Eligible Employees of CRC Group: At CRC Group, we're committed to supporting every aspect of teammates' well-being – physical, emotional, financial, social, and professional. Our best-in-class benefits program is designed to care for the whole you, offering a wide range of coverage and support. Eligible full-time teammates enjoy access to medical, dental, vision, life, disability, and AD&D insurance; tax-advantaged savings accounts; and a 401(k) plan with company match. CRC Group also offers generous paid time off programs, including company holidays, vacation and sick days, new parent leave, and more. Eligible positions may also qualify for restricted stock units and/or a deferred compensation plan.
CRC Group supports a diverse workforce and is an Equal Opportunity Employer that does not discriminate against individuals on the basis of race, gender, color, religion, citizenship or national origin, age, sexual orientation, gender identity, disability, veteran status or other classification protected by law. CRC Group is a Drug Free Workplace.
EEO is the Law Pay Transparency Nondiscrimination Provision E-Verify
Listed on hirly, a job board. hirly is not the employer: Tihinsurance is hiring for this role.
Similar jobs
- SW Engineer- Developer Systems Reliability EngineeringVisa · US - Austin, TXFirst seen today
- Site Reliability EngineerTruist · 3 LocationsFirst seen today
- Site Reliability EngineerTruist · 3 LocationsFirst seen today
- Site Reliability EngineerThales · Austin - Arboretum PlazaFirst seen today
- Site Reliability EngineerThales · AustinFirst seen today
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job