Allegion
Senior AI Operations Engineer
Bangalore, India
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.3M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →hirly's read of this role
- Seniority
- Senior
- Country
- IN
- Work mode
- On-site / unstated
- First seen by hirly
- 1 Oct 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
Creating Peace of Mind by Pioneering Safety and Security
At Allegion, we help keep the people you know and love safe and secure where they live, work and visit. With more than 40 brands, 14,000+ employees globally and products sold in 130 countries, we specialize in security around the doorway and beyond.
Additionally, Allegion is proud to be recognized with the 2026 Gallup Exceptional Workplace Award (GEWA) for the third consecutive year, earning distinction in both the employee engagement and strengths categories. This year, Allegion also received Gallup’s With Distinction honor — a designation reserved for a select group of organizations that go above and beyond in building exceptional workplace cultures.
KEY RESPONSIBILITIES
AI Platform Operations
AI Platform Administration
Operate and maintain enterprise AI platforms and services across development, test, and production environments.
Support Azure OpenAI, Azure AI Foundry, Azure Machine Learning, and related Azure AI platform services.
Operate the application hosting platform for AI services across Azure App Service, Azure Functions, and Azure Container Apps, including networking, scaling, and runtime configuration.
Ensure platform availability, performance, reliability, scalability, and operational stability.
Perform platform configuration, lifecycle management, upgrades, and operational maintenance.
AI Service Deployment & Operations
Deploy and operationalize AI and Generative AI workloads in controlled enterprise environments.
Support production AI solutions including copilots, AI assistants, chatbots, agentic applications, retrieval-augmented generation (RAG) services, and machine learning applications.
Manage release activities, prompt and model version control, runtime configurations, deployment validations, and operational handovers.
Deploy and operate machine learning models and endpoints developed by Data Science teams, providing the pipelines, environments, and runtime platform they promote through.
Troubleshoot deployment, integration, performance, and runtime issues across AI services.
Container & Kubernetes Platform
Support containerised AI workloads using Docker and Kubernetes operational best practices, including image build, hardening, registries, and runtime configuration.
Help design, build, and operate Azure Kubernetes Service environments as the container platform is introduced for AI and platform workloads.
Establish and monitor cluster health, capacity, autoscaling, networking, ingress, security policies, and operational readiness.
Contribute to platform resiliency, backup, recovery, and disaster recovery readiness.
Model Access & Gateway Operations
Operate the governed model-access layer serving AI applications, agents, and developer tooling — provider routing, failover, model catalogue, and version management across Azure OpenAI, AWS Bedrock, and Google Vertex AI.
Manage model access control, user and application entitlements, quota allocation, rate limiting, and credential provisioning for the gateway.
Monitor and attribute token consumption, throttling and 429 events, latency, and cost by model, application, and consumer.
Support gateway upgrades, provider onboarding, configuration changes, and controlled rollout of new model versions.
Monitoring, Reliability & Support
Monitoring & Observability
Implement monitoring, logging, alerting, dashboards, and operational metrics for AI platforms and services.
Monitor AI system utilization, availability, latency, cost consumption, token usage, and performance trends.
Monitor AI-specific signals including model and agent quality, quota and throttling events, data or model drift, and grounding index freshness.
Use Azure Monitor, Application Insights, Log Analytics, and OpenTelemetry instrumentation to identify incidents and improvement areas.
Provide operational insights to engineering, security, governance, and leadership stakeholders.
Incident & Problem Management
Investigate and resolve platform incidents, production issues, service disruptions, and operational risks.
Perform root cause analysis and document corrective and preventive actions.
Support production support processes, escalation management, and incident communications.
Participate in post-incident reviews and implement service reliability improvements.
Operational Excellence
Create and maintain runbooks, support procedures, platform documentation, and operational knowledge articles.
Automate recurring support, validation, monitoring, and deployment activities.
Continuously improve platform efficiency, stability, security posture, and cost governance.
Support service-level objectives and operational performance targets for AI platforms.
Cloud Engineering & Automation
Infrastructure as Code
Develop and maintain Infrastructure as Code using Terraform, including reusable modules, remote state management, validation, and peer review.
Automate Azure AI platform provisioning, configuration, policy enforcement, and environment standardization across development, test, and production.
Manage controlled promotion of infrastructure changes across environments, with drift detection, change tracking, and infrastructure compliance.
Support configuration consistency, version control, and secure state management for all platform infrastructure.
CI/CD & Release Engineering
Build and maintain CI/CD pipelines for AI platform services, application workloads, containers, and supporting infrastructure.
Support deployment automation using Azure DevOps, GitHub Actions, or equivalent tools, including approvals, artifact management, and tested rollback.
Implement validation checks, release controls, and operational readiness gates, including automated model and prompt evaluations, content safety, latency, and cost checks before production release.
Bring unmanaged or manually deployed services under governed, automated delivery.
Maintain traceability and reproducibility across code, configuration, models, prompts, evaluation results, and deployed versions.
Developer Enablement & Self-Service
Build and maintain reusable Terraform modules, pipeline templates, and standardized provisioning patterns consumed by other engineering teams.
Establish golden paths and self-service capabilities that allow AI and application teams to onboard workloads with reduced manual effort and lower operational risk.
Promote reusable platform patterns and controlled deployment models across environments.
Maintain platform documentation, module usage guidance, and onboarding references for consuming teams.
Security & Identity Management
Implement secure cloud configurations and platform controls for AI services, including private networking, private endpoints, and network security controls.
Support identity and access management using managed identities, workload identities, RBAC, and least-privilege practices.
Manage secrets, certificates, keys, rotation, and secure configuration using Azure Key Vault or approved tools.
Implement and operate privileged access controls, including just-in-time elevation and periodic entitlement review.
Remediate identified platform security and hardening findings, and prevent recurrence through platform defaults, policy, and pipeline checks.
Partner with Information Security teams to address vulnerabilities, access reviews, and security findings.
REQUIRED QUALIFICATIONS
Education
Bachelor's degree in Computer Science, Information Technology, Engineering, Data Science, or a related STEM discipline, or equivalent practical experience.
Experience
8+ years of experience in Cloud Platform Engineering, Infrastructure Engineering, DevOps, Platform Operations, SRE, MLOps, or AI Operations.
Hands-on experience supporting Microsoft Azure-based enterprise environments and production workloads.
Experience op
Similar jobs
- Industrial Operations EngineerCapgemini Engineering · Bangalore, INFirst seen today
- Industrial Operations EngineerCapgemini Engineering · Bangalore, INFirst seen today
- Senior Network Operations EngineerSovos Compliance · Mulund East, Mumbai, Maharashtra, IndiaFirst seen 2d ago
- Senior AWS Operations EngineerHapag Lloyd · Chennai, IndiaFirst seen 3d ago
- Shift Operations Engineer (O&M- Electrical)Adani · Anantapur, Andhra Pradesh, IndiaFirst seen today
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job