hirly

DigitalBridge

Principal, Staff Site Reliability

Boca Raton, Florida · New York, New York

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at DigitalBridge first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.4M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

Apply from your AI assistant

Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.

Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.

hirly's read of this role

Role family
Engineering
Seniority
Lead / management
Stated salary
$210,000 – $230,000 per year
Country
US
Work mode
On-site / unstated
First seen by hirly
2 Oct 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

We are hiring a Staff SRE / Infrastructure Engineer to own the reliability, scalability, and developer experience of the platforms that power our investment, portfolio, and enterprise operations. You will set the technical direction for our hybrid cloud footprint (AWS + Azure +

on-prem), harden production for a regulated buy-side environment, and lead an emerging body of work applying AI to incident response and root-cause analysis. This is a hands-on senior IC role with material influence over architecture, tooling standards, and how engineers ship software.

What you 'll do

Own end-to-end reliability for business-critical services — SLOs, error budgets, capacity planning, DR, and incident command — with a five-nines mindset appropriate to financial services workloads.

Design and evolve our multi-cloud and on-prem infrastructure across AWS, Azure, and colocated environments; drive workload placement, cost, and resilience trade-offs.

Build and maintain the Terraform, Ansible, and CI/CD backbone that lets product and data teams ship safely and quickly; codify golden paths and paved roads.

Advance observability (metrics, logs, traces, profiling) so on-call engineers can localize failures in minutes, not hours; instrument reliability as a first-class product surface.

Lead the buildout of AI-assisted SRE capabilities — LLM-driven triage, incident summarization, runbook synthesis, and automated RCA — with human-in-the-loop guardrails.

Partner with security, data platform, and application teams on hardening, patching, secrets, network segmentation, and change management appropriate to a regulated environment.

Participate in a leader-level on-call rotation; run blameless postmortems and drive systemic fixes to closure.

Mentor senior engineers; set the bar for infrastructure code review, production readiness reviews, and reliability practice across the org.

Required experience

10+ years building and operating production infrastructure at scale, including hybrid cloud

+ on-prem.

Deep expertise in AWS and Azure compute, networking, IAM, and managed data services; comfortable with account/project topology, landing zones, and org-level guardrails.

Expert-level Linux , Terraform , Ansible , and modern CI / CD (GitHub Actions, GitLab CI, Argo, or equivalent).

Track record of measurable improvements in platform reliability, MTTR, deployment velocity, and developer productivity.

Strong scripting/software skills in Python and/or Go; can read and refactor application code well enough to debug across the stack.

Production experience with Kubernetes, service meshes, and container security posture.

Fluency with observability stacks (Datadog, Prometheus/Grafana, OpenTelemetry, ELK/Splunk).

Experience running incident command and driving durable postmortem outcomes.

Nice to have

Prior experience in financial services, buy-side, or another regulated environment (SOC 2, SOX, GLBA).

Hands-on work with AI / LLM tooling for SRE — agentic incident response, RAG over runbooks/telemetry, or automated RCA.

Experience with data-center automation, colocation, or modular infrastructure.

FinOps depth: unit economics, cost attribution, and budget governance across AWS/Azure.

The compensation range for this position is for a full-time employee in New York. The base salary offered will depend on qualifications, market data and internal equity.

Base Salary Range

$210,000 — $230,000 USD

At DigitalBridge, we strive to create an inclusive environment where diverse employees want to work and where they can flourish professionally. In furtherance of our culture, all qualified applicants will receive consideration for employment without regard to race, national origin, gender, age, religion, disability, sexual orientation, veteran status, marital status or any other characteristics protected by law.

Original posting on DigitalBridge's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job