Snorkel AI
Senior / Staff AI Engineer
New York, New York, United States · San Francisco
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.4M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →Apply from your AI assistant
Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.
Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.
hirly's read of this role
- Seniority
- Lead / management
- Country
- US
- Work mode
- Remote-friendly
- First seen by hirly
- 16 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
About Snorkel
Snorkel AI is the frontier AI data lab, helping teams build the data and environments behind high-performing frontier and agentic AI. We combine technology with research-driven AI data development to create datasets, benchmarks, evals, and custom solutions for real-world AI systems. Founded out of the Stanford AI Lab in 2019, Snorkel works with leading AI labs and enterprises to move from better data to better outcomes.
Excited to help us redefine how AI is built? Apply to be the newest Snorkeler!
About The Team
Snorkel's AI Platform organization builds the infrastructure and systems that power AI development at scale - synthetic data generation, evaluation, agentic workflows, simulation environments, LLM infrastructure, and distributed compute. Our platform enables engineering and research teams to rapidly experiment with models and agents, measure their behavior, and turn successful experiments into reliable production systems.
We're a small team operating at the intersection of distributed systems and applied AI, and we're in the middle of a foundational shift toward agent-first workflows where models interact with tools, environments, data, and other agents over long-running trajectories. The systems we build need to make these inherently non-deterministic workloads observable, reproducible, measurable, and scalable. You will help define how we do that.
About The Role
We're looking for AI Engineers who combine strong software and distributed systems fundamentals with experience operating AI systems in production. You'll build the infrastructure that lets teams create, experiment with, evaluate, and operate LLM and agentic workloads at significant scale - from synthetic data and evaluation pipelines to simulation environments, orchestration systems, and LLM infrastructure.
You'll work on systems where correctness is not defined by a single deterministic output. Instead, you'll build the infrastructure needed to understand behavior across models, prompts, tools, environments, and multi-step trajectories, and to continuously improve those systems through experimentation and evaluation.
We are looking to grow our team of AI Engineers, and are hiring at multiple levels.
What You'll Do
Design and build infrastructure for running large-scale agentic workloads, including multi-step agents interacting with tools, external services, sandboxes, and simulated environments
Build scalable synthetic data generation and automated labeling systems that allow teams to create, refine, and evaluate high-quality training and evaluation datasets
Design evaluation infrastructure for measuring AI system behavior across models, prompts, tools, environments, and multi-step trajectories - including reproducible experiments, benchmark execution, regression detection, and continuous evaluation
Build orchestration and distributed compute systems for running thousands to millions of AI experiments and simulations reliably across heterogeneous compute environments
Develop infrastructure for agent simulation environments, including environment provisioning, isolation, lifecycle management, and scalable execution
Build and operate LLM infrastructure for routing, rate limiting, retries, caching, provider failover, cost attribution, and efficient execution across multiple model providers
Instrument agent and model workloads so failures are observable and debuggable - capturing traces, model interactions, tool calls, environment state, evaluation results, latency, reliability, and cost
Design systems that make non-deterministic workloads reproducible and measurable, allowing engineers to compare experiments, diagnose behavioral regressions, and understand why an agent succeeded or failed
Improve the developer experience for AI experimentation by building APIs, SDKs, workflow abstractions, and tooling that make it easy to move workloads from local development to large-scale production execution
Collaborate with research, product, and engineering teams to turn experimental AI workflows into reliable, reusable platform capabilities
What You'll Bring
5+ years building production software systems, with experience in AI/ML infrastructure, ML platforms, distributed systems, data platforms, or backend infrastructure
Experience operating non-deterministic AI or ML workloads in production or at significant scale — you are comfortable reasoning about behavior across models, tools, environments, and multi-step execution
Experience building infrastructure for experimentation, evaluation, model development, synthetic data, agentic workflows, training, inference, or production ML systems
Strong proficiency in Python and experience building production-quality APIs, services, and developer tooling
Strong background in distributed systems and cloud platforms (AWS preferred), including compute orchestration, storage, networking, isolation, and failure handling
Experience with workflow or distributed execution frameworks such as Prefect, Airflow, Dagster, Ray, Kubernetes, or similar systems
Strong understanding of production system fundamentals - observability, telemetry, reliability, performance, debugging, incident response, and cost management
Ability to reason about AI system quality beyond traditional service metrics, including evaluation design, experiment reproducibility, behavioral regressions, and model or agent variability
Track record of leading complex engineering initiatives, influencing stakeholders, and delivering measurable impact
Ability to work in a fast-paced environment with strong technical communication skills
Fluency with modern AI and developer tooling and a willingness to rapidly evaluate and adopt new models, frameworks, infrastructure, and techniques as the ecosystem evolves
Nice to Have
Experience building or operating LLM or agent infrastructure, including model gateways, agent runtimes, tool execution, tracing, or multi-agent systems
Experience building evaluation or experimentation platforms for LLMs, agents, or other probabilistic systems
Experience with synthetic data generation, automated labeling, data refinement, or dataset quality systems
Experience building reinforcement learning environments, agent simulations, benchmarks, or other environment-based evaluation systems
Experience running large-scale distributed AI workloads across containers, Kubernetes, serverless compute, sandboxes, or heterogeneous compute environments
Experience with LLM observability, tracing, prompt/version management, token and cost attribution, rate limiting, caching, or multi-provider routing
Experience designing isolation and sandboxing infrastructure for executing model-generated code or tool calls safely
Experience building shared AI platform libraries or SDKs consumed by multiple teams, including versioning, backwards compatibility, and migration support
Experience in hyper-growth startup environments or scaling engineering organizations
Prior experience as a Tech Lead, Team Lead, or hands-on Engineering Manager
Why This Role
You'll have meaningful ownership over the infrastructure that determines how quickly Snorkel can experiment with, evaluate, and productionize new AI systems. This isn't a role focused on training a single model or maintaining traditional ML pipelines - you'll build the platform that makes large-scale AI experimentation possible.
You'll work on some of the hardest emerging infrastructure problems in AI: operating non-deterministic systems reliably, reproducing agent behavior across environments, evaluating long-running trajectories, scaling simulations and experiments, and turning rapidly evolving research workflows into robust production systems.
The architecture decisions made by this team will define how Snorkel builds and operates AI systems for years to come.
Snorkel is proud to be an Equal Employment Opportunity employer and is committ
Similar jobs
- Lead AI Engineer - AI PlatformLowes · Lowe's Charlotte Technology Hub 3505First seen today
- Principal Generative AI EngineerMaxar · Remote (United States)First seen todayremote
- Lead AI Engineer (Hybrid)Blattner · Remote/Traveling (For Corporate Use Only)First seen todayremote
- AI Engineer, LeadBah · McLean, VAFirst seen today
- Principal Applied AI Engineer - Entity AgentsZoominfo · RemoteFirst seen todayremote
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job