Aift
Senior Site Reliability Engineer, Vulcan (AI Security product)
UAE
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.3M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →hirly's read of this role
- Role family
- Engineering
- Seniority
- Senior
- Country
- AE
- Work mode
- Remote-friendly
- First seen by hirly
- 10 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
Job Overview
We are looking for a hands-on infrastructure engineer to own the deployment, migration and troubleshooting of on-premise Kubernetes environments for enterprise and government clients , including airgapped, high-security data center environments where remote access is not possible.
This is a client-facing, on-site role: you will be the technical authority in the room, responsible for executing complex infrastructure changes correctly the first time, diagnosing failures independently under pressure and communicating clearly with client stakeholders throughout.
This role carries real ownership; y ou will be expected to understand the systems deeply enough to make sound judgment calls when things don't go to plan, without waiting on remote support.
In this role, you will play a vital part in supporting our Cybersecurity b usiness , Vulcan. Vulcan is a cybersecurity solution for GenAI, providing red and blue team services to ensure compliance and security.
Learn more about us 👉
Vulcan product: https://vulcanlab.ai/
Vulcan LinkedIn: https://www.linkedin.com/company/vulcanlab-ai/
AIFT group: https://aift.io/
Responsibilities
Plan and execute on-prem Kubernetes cluster deployments, upgrades and infrastructure migrations (including IP re-addressing, certificate rotation and cluster reconfiguration) in production and airgapped environments
Diagnose and resolve failures independently on-site
Own the full infrastructure stack end-to-end: Kubernetes control plane and data plane, PostgreSQL (primary/replica replication), distributed storage (e.g. SeaweedFS /Ceph/similar), private container registries and centralized logging (ELK or equivalent)
Validate deployment tooling (scripts, installers, automation) thoroughly in lab/staging environments before any client-facing execution
Represent the technical work directly to client stakeholders on-site: explain status, failures and remediation plans clearly
Travel to client data centers (including airgapped /restricted-access sites) as required , sometimes on short notice, for deployment and go-live support
Write clear, structured runbooks, decision trees and incident reports that others (including less experienced engineers) can follow under pressure
Escalate risks proactively to internal leadership, not just after something has gone wrong
Requirements
Technical:
5-6 years of hands-on experience with Kubernetes in production, including at least one on-premise (not purely cloud-managed) deployment
Solid understanding of etcd internals. Quorum, peer membership, failure recovery, not just kubectl-level familiarity
Experience with kubeadm-based cluster bootstrapping and certificate management (SANs, CA rotation, renewal)
Working knowledge of PostgreSQL replication, Linux networking fundamentals (DNS, NTP, firewalls) and container registries (Docker Distribution or similar)
Comfortable working entirely from the Linux command line, writing and debugging bash scripts and reading unfamiliar automation tooling under time pressure
Experience with at least one distributed storage system (SeaweedFS, Ceph, MinIO or similar) is a strong plus
GPU-enabled Kubernetes nodes (NVIDIA device plugin, container toolkit) experience is a plus, not required
Working style:
Demonstrated ability to work independently in high-pressure, high-stakes environments without live support
Strong incident communication. Can explain technical failures to non-technical stakeholders factually and calmly, without over-promising or minimizing
A track record of validating changes in test environments before touching production and the judgment to insist on this even under deadline pressure
Comfortable with travel, including to secure/restricted facilities where personal devices, internet access or remote assistance may not be available
Nice to have:
Prior consulting, systems integration or professional services experience, ideally on enterprise or government accounts
Experience specifically in the GCC/Middle East region, or with government-sector clients
Security background (the ability to reason about access controls, credential handling and airgapped operational discipline is valuable given the environments involved)
Interview Process
HR phone interview: 1 hour
Online interview: 1.5 - 2 hours, meet with hiring manager
Online interview: 1 hour, meet with hiring team
Why Join Us?
Innovative Environment: Be part of a company at the forefront of technology to provide security in GenAI, with opportunities to work on groundbreaking projects.
Growth Opportunities: Take your career to new heights with our career development programs and growth-focused culture.
Dynamic Team: Join a multi-cultural and dynamic team of dedicated professionals who inspire and support each other .
Compensation : Competitive salary and benefits package, commensurate with experience and performance.
Similar jobs
- Site Reliability EngineerLoftorbital · Abu DhabiFirst seen 10d ago
- Senior Site Reliability EngineerHive · BerlinFirst seen today
- Senior Site Reliability Engineer (m/f/d)Da Vinci Engineering GmbH · Raunheim, Hessen, GermanyFirst seen today
- Senior Site Reliability EngineerOracle · Bangalore, Karnataka, IndiaFirst seen today
- Senior Site Reliability EngineerUnitedHealth Group · Noida, Ghaziabad, Uttar Pradesh, IndiaFirst seen today
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job