Backblaze
Site Reliability Engineer III (DBA)
Remote - US
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.3M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →hirly's read of this role
- Role family
- Engineering
- Seniority
- Senior
- Country
- US
- Work mode
- Remote-friendly
- First seen by hirly
- 30 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
About Backblaze
Backblaze is the object storage leader in the open cloud movement, fueling customer success with cloud storage built purposefully to unlock budgets, unburden administrators, and unleash innovators. Together with our partners, we’re helping customers break free from the restrictive, overpriced legacy solutions that hold them back, and blaze forward with the full power of the open cloud in their hands.
Founded in 2007, we scaled the business with less than $3 million in outside funding until 2021, when we did a traditional IPO on the Nasdaq stock exchange. Today, Backblaze generates over $100m in revenue and is the leading specialized storage cloud - managing over three billion gigabytes of data storage for 500K+ customers in 175+ countries, including businesses, developers, IT professionals, and individuals.
But while there is a lot to celebrate in our past, there is almost as much opportunity ahead of us. We’re seeking a Site Reliability Engineer III (DBA) to join our team!
About the Role:
Individuals fulfilling this role will be responsible for ensuring the stability, scalability, and reliability of our production database systems, primarily Vitess (distributed MySQL) and Cassandra, alongside the rest of our production services and infrastructure. This role carries the same on-call, incident response, and service ownership expectations as other SRE IIIs, with database systems serving as the area of deepest technical ownership.
Because our SRE Database Engineering function is new, this role will also help establish its operational foundation by designing database architecture and developing the runbooks, escalation guidance, procedures, and training materials that our Level 1 and Level 2 SRE Database Engineers will use as they onboard. The ideal candidate will have strong experience with production database systems, Linux, automation, distributed systems, Kubernetes, observability, and incident response, with a proactive approach to reliability and operational excellence.
What You'll Do:
Database Architecture & Administration:
Design, deploy, and own highly available database architecture for Vitess (distributed MySQL) and Cassandra
Establish and document operational procedures, runbooks, and escalation guidance for Level 1 and Level 2 SRE Database Engineers
Optimize database performance through query tuning, indexing strategies, schema design, and capacity planning
Own database backup, recovery, replication, and disaster recovery strategies
Perform and validate disaster recovery testing and database recovery procedures
Drive database security, access control, patching, hardening, and compliance practices
Partner with the DBA and Data Infrastructure teams on resharding, capacity planning, replication, and architecture decisions for sharded MySQL environments
Service Reliability & Operations:
Support the availability and durability of critical services across production environments
Monitor service health using SLIs, SLOs, error budgets, monitoring, logging, and alerting platforms
Partner with service owners to define and improve SLIs, SLOs, error budget policies, and alerting
Participate in on-call rotations, incident response, root cause analysis, and post-incident reviews
Serve as an escalation point for complex database production incidents
Follow established ITIL/OSS processes including incident, change, problem, and capacity management
Take ownership of operational issues and drive projects from problem discovery through resolution
Automation & Tooling:
Develop automation for common operational and database administration tasks to reduce manual intervention and operational toil
Contribute to monitoring, logging, and alerting frameworks including Prometheus, Grafana, Catchpoint, and ELK
Help integrate operational runbooks and incident response workflows with FireHydrant
Work with CI/CD pipelines, configuration management, and infrastructure as code tools including Terraform, Ansible, and Jenkins
Develop scripts using Bash, Python, Go, or similar technologies to improve reliability and operational efficiency
Operate and troubleshoot containerized production environments using Kubernetes and Docker
Work within Kubernetes and Vitess environments using technologies such as kubectl, mysqlsh, and Vitess keyspaces
Project Management:
Lead Production Readiness Reviews (PRRs) for functionality being handed off from engineering partner teams
Support the operational readiness of new database-backed services before they enter production
Build training plans, onboarding materials, and technical documentation for new Level 1 and Level 2 SRE Database Engineers
Partner with Engineering, Product, Operations, and DBA/Data Infrastructure teams on reliability initiatives
Assist with capacity planning, disaster recovery exercises, database migrations, and infrastructure projects
Work with vendors and service providers to troubleshoot service issues and track SLA performance
Identify opportunities for automation and process efficiency
Incident response
Respond to and resolve production database, infrastructure, and service incidents
Troubleshoot and escalate database, Linux, networking, application, and infrastructure issues as needed
Participate in the on-call rotation and serve as an escalation point for database-related incidents
Lead or contribute to root cause analysis and post-incident reviews
Identify recurring issues and develop long-term corrective actions to improve reliability
What we value:
A proactive mindset with a can-do attitude
Someone who can work independently, take ownership, and drive complex technical problems through resolution
Someone who steps up, supports teammates, mentors others, and shares knowledge freely
Strong problem-solving skills and a willingness to learn new technologies
Curiosity, reliability, and a desire to improve the reliability and scalability of production systems
Required Qualifications:
6–8 years of experience in site reliability engineering, systems engineering, infrastructure operations, database engineering, or similar roles, with meaningful experience supporting production database systems.
Deep hands-on experience with MySQL and distributed or sharded database systems.
Experience with Vitess in a production environment strongly preferred.
Experience administering and supporting NoSQL databases such as Cassandra.
Experience designing high-availability database architecture, replication topology, backup strategies, and disaster recovery processes.
Strong SQL skills, including query performance analysis, indexing, schema design, and troubleshooting.
Solid Linux systems administration and troubleshooting skills.
Experience with security-focused operations including patching, system hardening, access controls, and vulnerability remediation.
Strong understanding of service reliability concepts including monitoring, alerting, incident response, root cause analysis, SLIs, SLOs, and error budgets.
Experience working with containers and orchestration platforms including Kubernetes and Docker.
Comfortable operating in Kubernetes and Vitess environments using tools such as kubectl, mysqlsh, and Vitess keyspaces.
Experience with infrastructure and configuration management technologies including Terraform, Ansible, Jenkins, and HashiCorp products such as Vault and Nomad.
Proficiency in at least one scripting language such as Python, Bash, or Go.
Experience establishing operational procedures, runbooks, documentation, and escalation processes.
Experience mentoring, training, or helping onboard engineers into complex technical environments.
Experience in SaaS, cloud services, service provider, or large-scale distributed systems environments preferred.
Experience with AWS, GCP, Azure, or similar cloud platforms prefer
Similar jobs
- Senior Site Reliability EngineerWesco · Atlanta, GA, United StatesFirst seen today
- Senior Site Reliability EngineerWesco · Dallas, TX, United StatesFirst seen today
- Site Reliability Engineer III - Performance Engineer- Service nowJPMorgan Chase · San Francisco, CA, United StatesFirst seen today
- Senior Site Reliability Engineer (Digital Banking)Mtb · Wilmington, DEFirst seen today
- Senior Site Reliability EngineerMastercard · O'Fallon, MissouriFirst seen today
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job