Openteams
Platform Engineer, AI/ML Infrastructure
United States - Remote OR Hybrid
Apply through hirly
Upload your resume and get a version tailored to this job, plus a cover letter, in about thirty seconds — before you create an account.
Apply with hirlyhirly's read of this role
- Role family
- Engineering
- Seniority
- Mid level
- Stated salary
- $120,000 – $250,000 per year
- Country
- US
- Work mode
- Remote-friendly
- First seen by hirly
- 24 Sept 2026
Derived automatically from the posting. Sign up to see how the role scores against your own resume.
the posting
Who We Are
Every organization runs on intelligence: years of accumulated knowledge, decisions, and context. As AI takes on more of that work, companies face a choice: rent that intelligence from vendors who keep the data, the context, and the results, or own it.
OpenTeams exists to make ownership possible.
Founded by Travis Oliphant, creator of NumPy and SciPy, and built by people with deep roots across the open-source ecosystem, including NumPy, SciPy, PyTorch, and Jupyter, we help enterprises and governments build AI they control, govern, and evolve themselves.
If that sounds like your kind of work, we'd like to meet you.
Platform Engineer, AI/ML Infrastructure
Location: U.S - Remote OR Hybrid - Washington, DC, Denver, CO or Colorado Springs, CO.
Work Authorization: U.S. citizenship required
Clearance: U.S.-Remote Opening: An active clearance is not required. Candidates must be eligible and willing to obtain and maintain a U.S. security clearance. Hybrid Opening: An active TS/SCI clearance is preferred. Candidates may also be considered if they previously held a TS/SCI with CI polygraph or currently hold an active TS or Secret clearance.
Salary Range: $120,000–$250,000 USD, dependent on experience level and location
Openings: Two positions are available:
One hybrid position: Candidates must be located in or willing to work hybrid from Washington, DC; Denver, CO; or Colorado Springs, CO. An active TS/SCI clearance is required. Up to 15% travel is required.
One U.S.-remote position : Candidates may work remotely from anywhere in the United States. An active clearance is not required, but candidates must be willing and able to undergo the process required to obtain and maintain a U.S. security clearance.
Candidates will be considered for the opening that best aligns with their location, clearance status, experience, and work preferences. Candidates who meet the requirements for multiple openings may be considered for more than one.
Candidates will be considered for the opening that best aligns with their location, clearance status, experience, and work preferences. Candidates who meet the requirements for multiple openings may be considered for more than one.
About the Role
OpenTeams builds AI platforms that governments and enterprises own outright: the infrastructure, the data, the models, and the evidence that the whole thing does what it claims. We're hiring several engineers to build and run that infrastructure.
The work spans the full depth of an AI platform. The Kubernetes clusters that schedule GPU workloads, move large datasets, and keep tenants isolated from one another. The services that make it a platform rather than a cluster: workflow orchestration, data ingest, model serving, policy enforcement, audit logging. The delivery path that gets released into production reliably and can prove what it shipped. The cloud infrastructure that ties it all together, and the operational practices that keep a distributed system resilient.
Much of this has to run where you can't assume normal cloud resources, or even an internet connection. That constraint is the interesting part of the job. Portability, reproducibility, and operability are design inputs from the first commit rather than problems handed downstream.
We build on open source and contribute back. Kubernetes, Terraform and OpenTofu, Argo, Prometheus, Nebari, and others. Upstream work is part of the job, not something you do on weekends.
This posting covers multiple roles, spanning mid-level through senior. We understand nobody spans every area above, so tell us where you fit. We assign level-based roles based on what you've actually done rather than a year count.
Key Responsibilities
Build and operate Kubernetes-based infrastructure for demanding AI/ML workloads, including GPU scheduling, resource management, and multi-tenant isolation
Design and implement platform services for orchestration, data ingest, model serving, and results management behind documented APIs
Write infrastructure as code and build GitOps pipelines so environments are reproducible from source
Build and operate CI/CD pipelines that produce versioned, signed, scanned release artifacts along with the documentation needed to deploy them
Own reliability: capacity planning, upgrade paths, failure-mode analysis, backup and recovery, incident response, and postmortems
Implement monitoring, logging, tracing, and alerting, and define the service level objectives they're measured against
Deploy and validate the platform in restricted, disconnected, or limited-connectivity environments, and verify parity after each release
Keep the platform portable by constraining dependencies to what's confirmed available in target environments
Write runbooks and operational documentation that other engineers can execute without you in the room
Contribute to Nebari and other open-source infrastructure, Kubernetes, and MLOps projects
Work with security engineers, government stakeholders, and other engineers to turn requirements into systems that hold up
Collaborate asynchronously across a distributed team
Required Skills & Experience
U.S. citizenship, and the ability to obtain and maintain a U.S. security clearance
Four or more years of hands-on experience building or operating production infrastructure, platforms, or distributed systems
Production experience with Kubernetes and containerized workloads
Experience with at least one major cloud platform: AWS, Azure, or Google Cloud
Experience with infrastructure as code and CI/CD, using tools such as Terraform, OpenTofu, Pulumi, or Helm
Working proficiency in Python, Go, Bash, or a comparable language
Experience implementing or operating production monitoring and observability
Ability to write documentation, runbooks, and deployment procedures that other people can actually follow
Ability to work independently and collaborate well in a remote, distributed team
Nice to Have
You will not have all of these, and very few people will. They're the things that would help, not a checklist. If the required list above describes you, apply.
An active U.S. security clearance, particularly TS/SCI with CI polygraph
Experience deploying or operating software in air-gapped, disconnected, or otherwise restricted environments
Experience with Department of Defense, Intelligence Community, or comparably regulated programsFamiliarity with the Risk Management Framework, NIST 800-53 or 800-171, or similar frameworks, and with producing the evidence they require
Experience supporting an Authorization to Operate, or with continuous ATO modelsSupply chain security work: hardened images, artifact signing, SBOM generation, dependency and container scanning, policy enforcement
Familiarity with cross-domain solutions, guards, data diodes, or similar transfer mechanisms
Experience with classified cloud environments, including AWS Secret or Top Secret regionsA DoD 8140/8570 qualifying certification such as Security+, CISSP, CASP+, or CISM, or willingness to obtain one after joining
Experience building MLOps pipelines or infrastructure for AI/ML workloads
Experience with GPU scheduling, distributed inference, or large-scale data and evaluation pipelines
Experience with model-serving or gateway frameworks such as KServe, vLLM, or LLM-DExperience designing API-first services and vendor-agnostic platforms that run across multiple environments
Experience with agentic workflow frameworks or multi-step AI pipeline orchestration
Contributions to open-source Kubernetes, infrastructure, MLOps, or observability projects, and experience with Nebari specifically
Familiarity with data sovereignty and privacy requirements for enterprise or government AI systems
Experience leading technical initiatives, setting engineering standards, or mentoring other engineers
Experience supporting rapid prototyping programs
Similar jobs
- Cloud Platform EngineerVaridesk LLC · HQ - Coppell, TXFirst seen today
- AI Platform EngineerRSC2 · Hanover, MDFirst seen today
- Cloud Platform Engineer - Clearance RequiredPhoenix Operations Group · Bethesda, MDFirst seen today
- Cloud Platform Engineer - Clearance RequiredPhoenix Operations Group · Falls Church, VAFirst seen today
- Intelligent Systems Platform EngineerHyper Careers Page · Richmond, VAFirst seen today
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job