EggAI
IT Operations Lead - Banking
Paris, Paris, France · München, Bavaria, Germany · Rome, Italy · London, England, United Kingdom
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.4M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →Apply from your AI assistant
Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.
Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.
hirly's read of this role
- Seniority
- Lead / management
- Countries
- FR, DE, IT, GB
- Work mode
- Remote-friendly
- First seen by hirly
- 29 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
EggAI · Banking transformation programme (UK) · Remote, UK-facing — London for gate weeks
About the Programme
EggAI is the transformation partner on a core banking system replacement for a client to lay the foundations of operating Agentic Generative AI in production.
EggAI's role is to specify, set the standard and programme-manage. You will define what operational readiness means and prove it at each gate, working with the client's own teams and its managed service providers, as well as EggAI's platform engineers.
The platform is a mix of classical banking infrastructure and agentic AI workloads, so the reliability problem runs across both deterministic and probabilistic systems.
The Role
Own the definition of operational readiness for go-live: what "ready to run in production" means, how the bank proves it at each regulatory gate, and what has to change in the organisation for that proof to hold.
The bank's operations teams, processes and provider contracts were built to run classical banking infrastructure. Agentic AI workloads need a different operating model, and you will shape it: redraw responsibilities between the client's teams and its managed service providers, set the standards each party works to, and lead the people through the change — so that those who run the platform today are ready and willing to run what comes next.
This is a lead/principal position reporting to the Programme Plan and Execution Lead. It is remote, with regular London presence required for gate weeks and workshops.
Responsibilities
Reliability targets — set the SLO and error-budget framework alongside the client's existing structure, have the service owners adopt it, and commission the behavioural quality measures that do the same job for agentic workloads
Observability and detection — define what Tier-1 services and agentic workloads must emit, including trajectory logging; require an automated path from a firing alert to the affected services and their owners; and hold the operating teams to closing the logging gaps in what runs today
The incident and change path — set the runbook standard and its required coverage, extend incident classification and intervention mechanisms to agentic failure modes, and have rollback and change-failure rates measured and reported by those who own them
Capacity and performance — direct the headroom position through the parallel-run peak at cutover, sponsor the chaos and game-day programme for payment rails and the core ledger, and mandate that capacity modelling covers inference workloads
Gate evidence — own the technical content of the Operational Acceptance Report and Operational Readiness Certificate, the reliability section of each gate pack, and the reliability NFRs defended as ADRs through the client's Architecture Design Authority — written by the teams, reviewed and signed off by you
The operating organisation — decide how responsibilities sit between the client's teams and its managed service providers for the new platform, and lead the people concerned through that change
Expected experience
10+ years in leading IT operations teams — You have operated L1, L2, and L3 support as SRE, platform reliability or infrastructure organisations, with the people and the providers reporting to you
Aligns operations with development, engineering and the business — has set operating goals that engineering teams and business stakeholders signed up to, and kept them focused on those goals through delivery
Has introduced SLOs and error budgets into an organisation that did not have them, and can build on the experienced success factors
Infrastructure in depth — has deep infrastructure knowledge, and what is required for smooth business operations: telemetry, redundancies, backups, etc.
Incident management , with post-incident reviews that led to real changes in the system and in the organisation
Regulated financial services , or another environment where an external body audits operational evidence.
Has worked through a managed service provider and can set standards they do not personally execute
Comfortable operating in a complex, multi-stakeholder programme where scope and ownership are still being established across parallel workstreams
Valuable, not required
LLM or agentic system operations: eval harnesses, drift detection, trajectory analysis, prompt and model versioning treated as change types
Core banking or payments exposure, particularly scheme gateway integration
What This Role Is Not
Not a platform architect. A separate role owns agentic platform design and build
Not an on-call engineer. You will not carry a client pager; out-of-hours operations are run by a third party
Not a service delivery manager. This is engineering judgement applied to readiness, not ITSM process administration
About EggAI
Agentic workforces are inevitable. Making them work is our mission.
EggAI is building the engineering methods, operating models, and reusable technology to deploy agentic workforces with the quality and control required for sustained business impact.
Working closely with enterprise clients, we design, implement and operate agentic systems that evolve from augmenting individual tasks to autonomously executing end-to-end processes alongside people—resilient, controlled, and scalable. We begin with the client problem, select the technical approach that best serves it, and remain accountable through deployment, adoption, and operations. Each deployment strengthens the next by turning production learning into reusable capability.
We are an experienced, international team working across Europe. We set high standards, take ownership of outcomes, and value clear thinking, candid feedback, and the ability to turn ideas into tangible results.
Similar jobs
- IT Operations & ERP Manager m/w/d in HamburgTopteq Tankstellentechnik GmbH · Hamburg, GermanyFirst seen today
- IT Operations Manager - IT Software Engineering (m/f/d) for AIRBUSOrizon GmbH, Unit Aviation · Taufkirchen, Kreis München, Bayern, GermanyFirst seen today
- IT Operations Manager (m/w/d)FERCHAU GmbH Niederlassung Bielefeld · Bielefeld, Nordrhein-Westfalen, GermanyFirst seen today
- IT Operations Engineer (m/w/d) Virtual Desktop Infrastructure AIRBUSOrizon GmbH, Unit Aviation · Taufkirchen, Kreis München, Bayern, GermanyFirst seen today
- IT Operations Specialist (m/f/d)Statista · HamburgFirst seen todayremote
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job