This role has closed. OpenTeams has taken the posting down.
hirly last saw it live on 5 September 2026. See similar open roles below, or browse the live board.
OpenTeams
Senior Infrastructure Engineer - AI/ML Platform
United States - Remote
Similar open jobs
- Infrastructure Engineering Senior Advisor- HybridCigna · Raleigh, NCFirst seen today
- Infrastructure Engineer Sr (ETL Administrator)Pnc · 6 LocationsFirst seen today
- CICS Systems Programmer (Infrastructure Engineering Senior Advisor ) - Evernorth Health Services - RemoteCigna · 7 LocationsFirst seen todayremote
- Senior Infrastructure EngineerAdonis · New York CityFirst seen today
- Infrastructure Engineer Senior (ETL Administrator)Pnc · 5 LocationsFirst seen yesterday
- Senior Site Reliability Infrastructure Engineer- FedrampCisco · 2 LocationsFirst seen yesterday
- Infrastructure Engineer IIICherokee Federal · Tulsa, OK, United StatesFirst seen yesterday
- Senior Core Infrastructure Engineer - OCI Product Catalog (Nashville-TN)Oracle · Nashville, TN, United StatesFirst seen yesterday
- Senior Infrastructure Engineer, Systems – CollaborationSalesforce · 2 LocationsFirst seen yesterday
- Cloud Infrastructure EngineerLeidos · Gaithersburg, MDFirst seen today
- Cloud Infrastructure Engineer / OCI EngineerLeidos · 2 LocationsFirst seen today
- Infrastructure EngineerGhr · 2 LocationsFirst seen today
- Infrastructure EngineerGhr · 2 LocationsFirst seen today
- Automated Driving Performance HIL Infrastructure EngineerGeneralmotors · Milford, Michigan, United States of AmericaFirst seen today
- Staff Data Infrastructure EngineerFaire · New York City, NY; San Francisco, CAFirst seen todayremote
hirly's read of this role
- Seniority
- Senior
- Stated salary
- $145,000 – $250,000 per year
- Country
- US
- Work mode
- Remote-friendly
- First seen by hirly
- 5 Sept 2026
Derived automatically from the posting.
the posting
Who We Are
Every organization runs on intelligence: years of accumulated knowledge, decisions, and context. As AI takes on more of that work, companies face a choice: rent that intelligence from vendors who keep the data, the context, and the results, or own it.
OpenTeams exists to make ownership possible.
Founded by Travis Oliphant, creator of NumPy and SciPy, and built by people with deep roots across the open-source ecosystem, including NumPy, SciPy, PyTorch, and Jupyter, we help enterprises and governments build AI they control, govern, and evolve themselves.
If that sounds like your kind of work, we'd like to meet you.
Senior Infrastructure Engineer - AI/ML Platform
Location: U.S - Remote
Work Authorization: U.S. citizenship required
Clearance: An active clearance is not required at the time of hire. Candidates must be able to obtain and maintain a Secret security clearance, which includes a federal background investigation.
Salary Range: $145,000–$250,000 USD, dependent on experience level and location
About the Role
We’re seeking a Senior Infrastructure Engineer to build and operate the platform underlying secure AI test and evaluation capabilities for Government teams assessing AI systems.
You’ll own a Kubernetes-based platform supporting demanding AI/ML workloads, including GPU scheduling, large-scale data movement, reproducible test execution, and multi-tenant isolation. The platform must operate reliably within Government environments that may have restricted networks, accreditation boundaries, limited connectivity, and no assumption of outbound internet access or managed cloud services.
You’ll use open-source technologies such as OpenTofu, Terraform, Helm, Argo CD, Kubernetes operators, and Nebari to build reusable and composable infrastructure. You’ll also own platform reliability, observability, capacity planning, upgrade paths, hardened configurations, and the documentation needed to deploy and maintain the platform.
This is a fully remote, U.S.-based role working with a distributed team that relies heavily on asynchronous communication.
Key Responsibilities
Build and operate the Kubernetes platform supporting AI test and evaluation frameworks
Implement GPU scheduling, workload orchestration, resource management, and multi-tenant isolation across evaluation teams
Design infrastructure-as-code, GitOps workflows, and automated deployment pipelines that make the platform reproducible from source
Develop reusable and modular infrastructure components that can be composed into independently owned and operated platforms
Contribute to Nebari and other open-source Kubernetes, infrastructure, and MLOps projects used by the platform
Own platform reliability, including capacity planning, upgrade strategies, failure-mode analysis, backup and recovery considerations, and operational readiness
Design and implement observability, monitoring, logging, tracing, and alerting for large-scale AI/ML workloads
Develop operational runbooks and documentation that enable other engineers to deploy, operate, and troubleshoot the platform
Deploy, configure, and harden infrastructure within secure, restricted, disconnected, or limited-connectivity Government environments
Support security authorization and compliance activities through infrastructure documentation, hardened configurations, control evidence, and repeatable deployment processes
Integrate automated security tooling for container scanning, static and dynamic analysis, artifact signing, and policy enforcement
Collaborate with Government stakeholders, security personnel, software engineers, and ML engineers to translate platform requirements into reliable infrastructure
Provide technical leadership, contribute to engineering standards, and mentor less-experienced team members
Collaborate effectively within a remote and distributed team using asynchronous communication practices
Required Skills & Experience
U.S. citizenship and ability to obtain and maintain a Secret security clearance
6+ years of hands-on infrastructure, platform, DevOps, or site reliability engineering experience supporting production systems
Strong understanding of infrastructure engineering principles, including scalability, reliability, observability, security, and automation
Production experience with Kubernetes, including workload scheduling, resource management, and multi-tenant environments
Experience with automated security tooling, such as container scanning, SAST/DAST, artifact signing, and policy enforcement
Proficiency with infrastructure-as-code tools such as Terraform, OpenTofu, Pulumi, or equivalent technologies
Experience with at least one major cloud platform—AWS, Azure, or Google Cloud—including networking, security, storage, and compute services
Experience implementing monitoring and observability using tools such as OpenTelemetry, Prometheus, Grafana, or equivalent technologies
Strong programming or automation skills using Python, Go, or a comparable language
Experience with CI/CD practices, GitOps workflows, and infrastructure automation
Experience creating maintainable operational documentation, deployment procedures, and runbooks
Experience leading technical initiatives, establishing engineering practices, or mentoring other engineers
Ability to work independently and collaborate effectively within a remote, distributed team
Ability to give and receive constructive technical feedback
Nice to Have
Experience deploying or operating infrastructure in air-gapped, disconnected, or highly restricted environments
Experience engineering systems subject to the Risk Management Framework, NIST 800-53, NIST 800-171, or comparable security requirements
Experience supporting security personnel in achieving an Authorization to Operate for complex systems in Government or Intelligence Community environments
Experience supporting systems operating at Impact Level 5 or higher
Current DoD 8140 qualifying certification, such as Security+, CISSP, CISM, or an equivalent IAT/IAM Level II or III credential
Experience building MLOps pipelines or infrastructure supporting AI/ML workloads
Experience with GPU scheduling, large-scale data pipelines, or reproducible ML evaluation workloads
Experience with model-serving frameworks such as KServe, vLLM, LLM-D, or equivalent technologies
Familiarity with data sovereignty, privacy, and security requirements for enterprise or Government AI systems
Contributions to open-source Kubernetes, infrastructure, MLOps, or observability projects
Experience with Nebari
What We Offer
Medical, Dental & Vision – 100% paid for employees, 75% for dependents
401(k) Match – Up to 5% with full vesting after 2 years
Unlimited PTO – With a required minimum of 15 days off annually
Fully Remote Setup – Includes up to $3,000 equipment reimbursement
Continuous Education – Includes up to $500 reimbursement
Disability & Life Insurance – 100% employer-paid
HSA & FSA Options – With monthly HSA contributions from OpenTeams
Grow With Us
At OpenTeams, growth isn’t just about the company—it’s about you.
We believe the best careers are built at the edge of your potential. That is where new tools, ideas, and technologies change the world. Here, you’ll work alongside pioneers of AI, solving problems that matter: making AI more transparent, more ethical, and more empowering. As your skills grow, our career framework provides a pathway and recognition of that increased impact.
Opportunities aren’t limited by geography. You’ll collaborate with global experts, contribute to open source projects that power the world’s technology, and stretch your skills daily. That global perspective and diversity makes our solution more universal and robust. We are committed to continuing to celebrate diversity on our team.
Supported people are successful people. We offer 100% e