hirly

Absa

Senior AI Platform Engineer (Cloud)

Sandton

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Absa first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included
Tailor my resume for this job →

Apply from your AI assistant

Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.

Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.

hirly's read of this role

Role family
Engineering
Seniority
Senior
Country
ZA
Work mode
On-site / unstated
First seen by hirly
9 Oct 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

Empowering Africa’s tomorrow, together…one story at a time.

With over 100 years of rich history and strongly positioned as a local bank with regional and international expertise, a career with our family offers the opportunity to be part of this exciting growth journey, to reset our future and shape our destiny as a proudly African group.

Belonging at Absa

Absa is committed to creating an inclusive workplace where everyone can thrive. We are an equal opportunity employer and welcome applications from suitably qualified individuals from diverse backgrounds. In support of our Diversity, Equity, Inclusion and Belonging (DEIB) commitments and Employment Equity objectives, preference may be given to candidates from underrepresented designated groups, including persons with disabilities. We encourage applicants who may require reasonable accommodation during the recruitment process to let us know so that appropriate support can be provided.

Job Summary

Absa Group’s Chief Data Analytics and Applied AI Office (CDAIO) requires a technically exceptional and commercially grounded AI Platform Engineer (Cloud) to design, build, operate, and continuously optimise the multi-cloud AI infrastructure that powers the bank's enterprise AI capability. The AI capability must enable the CDAIO to fulfil its mandate as steward of the bank’s AI capabilities through the end-to-end delivery of the AI platform enablement, governance and acceptable use in service of the bank’s strategic and commercial objectives. This role is the engineering backbone of a platform that supports various live AI projects across four business units (CIB, PPB, BB, and AR) and ten countries. This role demands deep technical mastery in cloud AI infrastructure, AI FinOps, zero-trust security architecture, agentic AI infrastructure, and platform observability, combined with the commercial fluency to govern AI compute costs at enterprise scale and communicate trade-offs to senior business and finance stakeholders. The role includes but not limited to applying critical thinking, design thinking, and problem-solving skills in an agile team environment to solve complex platform engineering challenges, delivering high-quality solutions at optimal cost to serve, in full compliance with Absa's Enterprise-Wide Risk Management Framework, Group Architecture standards, and AI Responsible Use Policy. The successful candidate carries full accountability for building high-performing, scalable, enterprise-grade Platform services. As well as build capability in others to do the same.

Job Description

KEY FOCUS AREAS

AI Platform Engineering and Architecture : Design and operation of enterprise-grade, multi-cloud AI platform infrastructure supporting bank-wide AI delivery at scale across the AI platform stack (AWS Bedrock, Databricks AI, Microsoft Azure AI Foundry, Hugging Face, and GPU clusters).

AI FinOps and Compute Cost Governance : Full accountability for AI compute cost models, chargeback and showback frameworks, provisioned throughput optimisation, and monthly cost-per-use-case reporting to Group Finance across all four business units.

Platform Observability and SLA Engineering : AI-specific service reliability standards, observability tooling, and incident management for production AI workloads serving 43 live projects across ten countries.

AI Security Architecture and Zero Trust : Zero-trust security design, OAuth / OIDC integration, prompt injection controls, and data residency compliance protecting Absa's AI platform across the different country jurisdictions.

Agentic AI Infrastructure: Design and operation of the infrastructure layer enabling multi-agent AI systems, autonomous workflows, tool-calling architectures, and agent orchestration at enterprise scale.

ACCOUNTABILITIES

Platform Engineering and Architecture

Lead the design, deployment, and continuous optimisation of Absa's multi-cloud AI platform stack: AWS Bedrock, Databricks AI, Microsoft Azure AI Foundry, Hugging Face Model Hub, and on-demand GPU clusters.

Architect scalable, resilient, and reusable platform components including AI Gateway configuration, model serving infrastructure, vector database deployments, and data pipeline integration to support bank-wide AI delivery.

Define and maintain infrastructure-as-code (IaC) standards (e.g. using Terraform or Pulumi), enabling repeatable, auditable multi-cloud AI deployments across Absa's operating territories (10 countries).

Lead the design and operation of agentic AI infrastructure: orchestration runtime environments (e.g. Microsoft Foundry Agent Service, AWS Bedrock Agents), tool-calling schemas, agent memory and state management patterns, and multi-agent communication protocols.

Develop and enforce cloud-agnostic model serving patterns to reduce platform lock-in and ensure workload portability across the CDAIO's multi-vendor stack.

Identify and select appropriate internal and external technologies to deliver AI platform services; apply excellent judgement in continuously improving platform engineering practices.

Take full accountability for end-to-end platform quality, completeness, and user experience across the development, deployment, and operational lifecycle.

Positively contribute to the design and evolution of Group Architecture, infrastructure standards, and AI platform governance frameworks

AI FinOps and Compute Cost Governance

Own the AI compute cost model for the CDAIO, including chargeback and showback frameworks for Databricks DBU consumption, AWS Bedrock token-based pricing, Azure AI Foundry provisioned throughput units, and GPU cluster utilisation across all four business units.

Design and maintain FinOps dashboards and cost attribution reports using AWS Cost Explorer, Databricks System Tables cost analytics, and Azure OpenAI utilisation tooling — providing monthly cost-per-use-case reporting to Group Finance and the CDAIO COO.

Evaluate and manage provisioned throughput versus on-demand consumption trade-offs for production AI workloads, presenting optimisation recommendations to the CDAIO and BU technology leads.

Identify and execute AI compute cost optimisation opportunities: workload scheduling, spot instance strategies for training workloads, model distillation to reduce inference cost, and right-sizing of GPU clusters.

Create business cases and solution specifications for AI platform investments and governance processes, including CTO and architecture approvals.

Collaborate with the FinOps capability within the CDAIO COO to align AI platform costs to agreed budget envelopes and ensure spend anomalies are detected and escalated proactively

Platform Observability and SLA Engineering

Define, implement, and own AI-specific SLAs and OLAs covering inference latency, platform availability, token throughput, API gateway response times, and model serving reliability, with explicit targets agreed with each business unit technology lead.

Implement and maintain AI platform observability tooling (e.g. Prometheus, Grafana, Datadog, Databricks Lakehouse Monitoring, or equivalent) providing real-time visibility of platform health, model drift alerts, and capacity utilisation.

Design and operate incident management processes for AI platform failures: on-call runbooks, escalation paths, post-incident reviews, and root-cause remediation, ensuring minimal disruption to live AI projects across Absa's footprint.

Lead service improvement initiatives, translating performance data into platform enhancement programmes and continuously reducing mean time to recovery (MTTR) across the platform estate.

Own the release and change management process for AI platform components, including change governance, cutover management, and operational readiness sign-off in alignment with Absa's Group Technology change framework.

Use production performance monitoring and customer data to inform technical design and implementation decisio

Original posting on Absa's site ↗

Listed on hirly, a job board. hirly is not the employer: Absa is hiring for this role.

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job