hirly

IBM

Principal Site Reliability Engineer

Brno, Czech Republic

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at IBM first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.3M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

hirly's read of this role

Role family
Engineering
Seniority
Lead / management
Country
CZ
Work mode
On-site / unstated
First seen by hirly
30 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

At IBM Finance & Operations, we are the backbone of IBM’s transformation driving efficiency, transparency, and smart decision-making across the business. Our teams provide the insight and discipline that guide strategy, ensure financial strength, and enable IBM to invest in innovation and growth. Working in Finance & Operations means combining analytical skills with collaboration and curiosity. You’ll partner with colleagues across functions and geographies, using data, technology, and process excellence to create solutions that improve performance and deliver measurable impact. IBM offers continuous learning, career development, and a culture that values diverse perspectives. Join us and be part of a global team that keeps IBM moving forward, while building your own future in a dynamic and evolving environment. The IT Site Reliability Engineering team is seeking a Principal/Senior Site Reliability Engineer (SRE) to design, develop, scale, and operate our AI & OpenShift Platforms based on Red Hat technologies, including OpenShift AI (RHOAI) and Red Hat Enterprise Linux AI (RHEL AI). As a Principal SRE you will contribute to running core AI services at scale by enabling customer self-service, making our monitoring system more sustainable, and eliminating toil through automation. In our team, you will have the opportunity to lead and influence the complex challenges of scale which are unique to Red Hat IT managed platform services, while using your skills in coding, operations, and large-scale distributed system design. We develop, deploy, and maintain Red Hat’s next-generation AI application deployment environment for custom applications and services across a range of hybrid cloud infrastructures. We are a global team operating on-premise and in the public cloud, using the latest technologies from Red Hat and beyond. What will you do? * Working with live systems and coding automation * Design, build, and manage our large-scale infrastructure and platform services, including public cloud, private cloud, and datacenter-based * Automate cloud infrastructure through use of technologies (e.g. auto scaling, load balancing, etc.), scripting (bash, python and golang), monitoring and alerting solutions (e.g. Splunk, Splunk IM, Prometheus, Grafana, Catchpoint etc) * Design, develop, and become expert in AI capabilities leveraging emerging industry standards * Breakdown complex engineering efforts into consumable chunks while working with teams to understand deliverables * Design and development of software like Kubernetes operators, webhooks, cli-tools * Implement and maintain intelligent infrastructure and application monitoring designed to enable application engineering teams * Ensure the production environment is operating in accordance with established procedures and best practices * Lead escalation support for high severity and critical platform-impacting events * Provide feedback around bugs and feature improvements to the various Red Hat Product Engineering teams * Design software tests and lead peer reviews to increase the quality of our codebase * Help and develop peers’ capabilities through knowledge sharing, mentoring, and collaboration * Participate in a regular on-call schedule, supporting the operation needs of our tenants * Drive sustainable incident response and lead blameless postmortems * Work within a small agile team to develop and improve SRE methodologies, support your peers, plan and self-improve Agile, Amazon Web Services, Azure, Business Networking, C++, Collaboration, Communication, Computer Programming, Computer Science, Configuration Management, Distributed Systems, Domain Name System, Engineering, Golang, Google Cloud Platform, Hypertext Transfer Protocol, Internet Protocol Suite, Java, Kubernetes, Linux, OpenStack, Operating Systems, Passion, Platform as a Service, Public Clouds, Puppet, Python, Red Hat, Red Hat Ansible, Site Reliability, Software, Software as a Service, Systems Engineering, Troubleshooting, Unix n/a Czech Republic Infrastructure & Technology Remote Professional Brno, CZ (2150) IBM Middleware Czechia s.r.o

Original posting on IBM's site ↗

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job