Snorkel AI
Senior/Staff FDE - Synthetic Data Generation
New York, New York, United States · San Francisco
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.7M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →Apply from your AI assistant
Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.
Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.
hirly's read of this role
- Seniority
- Lead / management
- Stated salary
- $180,000 – $320,000 per year
- Country
- US
- Work mode
- On-site / unstated
- First seen by hirly
- 2 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
About Snorkel
Snorkel AI is the frontier AI data lab, helping teams build the data and environments behind high-performing frontier and agentic AI. We combine technology with research-driven AI data development to create datasets, benchmarks, evals, and custom solutions for real-world AI systems. Founded out of the Stanford AI Lab in 2019, Snorkel works with leading AI labs and enterprises to move from better data to better outcomes.
Excited to help us redefine how AI is built? Apply to be the newest Snorkeler!
About the Role
Snorkel AI is hiring a Forward Deployed Engineer focused on Synthetic Data Generation to partner with leading AI labs and enterprises on their most critical AI initiatives.
In this role, you will lead the technical execution of complex customer engagements where synthetic data is used to improve model training, evaluation, and performance. You will translate ambiguous model and data challenges into effective data strategies, build scalable generation and evaluation pipelines, and use experimentation to continuously improve data quality and downstream model outcomes.
You will work across the full delivery lifecycle—from technical discovery and solution design through implementation, evaluation, and production delivery. You will also identify patterns across engagements and turn successful approaches into reusable capabilities, technical standards, and product improvements.
Main Responsibilities
Synthetic Data Generation & Evaluation
Design and build scalable synthetic data generation, transformation, filtering, and evaluation pipelines for complex AI use cases
Translate model objectives, failure modes, and data gaps into synthetic data strategies, experiments, and technical specifications
Develop LLM- and ML-assisted workflows to generate high-quality training and evaluation datasets across targeted behaviors, domains, and edge cases
Build automated evaluators, quality checks, and measurement frameworks to assess correctness, relevance, diversity, coverage, and adherence to customer requirements
Design and run experiments to measure the impact of synthetic data on downstream model performance and iteratively improve generation approaches
Package and deliver production-grade datasets with standardized formats, quality assurance, and clear documentation
Forward Deployed Engineering & Customer Partnership
Lead technical workstreams from initial solution design through production delivery, navigating ambiguity and making sound technical decisions
Build, refine, and iterate on solutions that address customer needs, incorporating feedback to ensure the delivered work provides tangible value
Rapidly prototype and productionize solutions across models, data pipelines, APIs, and custom applications
Communicate technical tradeoffs, experimental results, and recommendations clearly to technical and cross-functional stakeholders
Serve as a trusted technical partner to customers and internal delivery teams, resolving complex blockers and driving alignment
Technical Leadership & Scale
Identify recurring patterns across customer engagements and turn successful solutions into reusable pipelines, evaluators, tooling, and best practices
Define and improve technical standards for synthetic data generation, experimentation, evaluation, and delivery
Partner with DaaS Engineering and Product teams to influence platform and product capabilities based on real-world customer needs
Lead technical design reviews, share expertise, and provide guidance to other engineers
Stay current with emerging synthetic data, LLM evaluation, and data curation techniques and assess their applicability to customer problems
What We're Looking For
5+ years of experience in machine learning engineering, data science, applied AI, forward deployed engineering, or a similar technical role
Strong Python skills and experience building reliable production data or ML systems, including containerizing with Docker and deploying on cloud platforms (e.g., AWS, GCP, or Azure)
Hands-on experience with LLMs—building model-based applications and data workflows with the modern GenAI/LLM stack, and integrating systems, models, and data sources through APIs
Strong understanding of ML experimentation and evaluation, including defining metrics and using empirical results to guide technical decisions
Experience building synthetic data, data augmentation, or model-generated training and evaluation datasets
Experience with LLM evaluation techniques, including LLM-as-a-judge, model-based evaluation, rubric-based evaluation, or custom evaluators
Demonstrated ability to take ambiguous technical problems from problem definition through delivery, with strong technical communication and experience working directly with customers and cross-functional stakeholders
Experience serving as a technical lead—setting technical direction, driving architecture and key decisions, mentoring engineers, and creating reusable approaches that influence broader engineering or product outcomes
Preferred Qualifications
Experience developing datasets for fine-tuning, preference optimization, or benchmarking, including human-in-the-loop generation and review workflows
Experience building agentic environments and tasks—repo-scale coding tasks, tool-agent-user interaction design, and agent tool protocols and interoperability
Experience with reinforcement learning for LLMs, including reward and verifier design and RL with verifiable rewards (RLVR)
Experience working in fast-paced, customer-facing environments where requirements and technical approaches evolve quickly
Compensation
The base salary range for this position is $180,000–$320,000 , with an additional variable compensation opportunity. The exact mix of base salary and variable compensation will depend on the role level and work location. Final compensation will be determined based on job-related skills, experience, relevant education or training, interview performance, and other business considerations.
Most offers include equity and benefits.
Actual compensation will be determined based on factors including skills, qualifications, experience, and geographic location.
Actual compensation will be determined based on factors including skills, qualifications, experience, and geographic location.
Salary range(s) for this role
$180,000 — $320,000 USD
Be Your Best at Snorkel
Joining Snorkel AI means becoming part of a company that has market proven solutions, robust funding, and is scaling rapidly—offering a unique combination of stability and the excitement of high growth. As a member of our team, you’ll have meaningful opportunities to shape priorities and initiatives, influence key strategic decisions, and directly impact our ongoing success. Whether you’re looking to deepen your technical expertise, explore leadership opportunities, or learn new skills across multiple functions, you’re fully supported in building your career in an environment designed for growth, learning, and shared success.
Snorkel AI is proud to be an Equal Employment Opportunity employer and is committed to building a team that represents a variety of backgrounds, perspectives, and skills. Snorkel AI embraces diversity and provides equal employment opportunities to all employees and applicants for employment. Snorkel AI prohibits discrimination and harassment of any type on the basis of race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state, or local law. All employment is decided on the basis of qualifications, performance, merit, and business need.
We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, to perform essential job fu
Listed on hirly, a job board. hirly is not the employer: Snorkel AI is hiring for this role.
Similar jobs
- Principal Research Scientist, Synthetic Data GenerationNvidia · 5 LocationsFirst seen today
- Member of Technical Staff - Simulation (Synthetic Data Generation), Frontier AI & Robotics (FAR)Amazon · San Francisco, California, USAFirst seen 6d ago
- Product Manager, Real World Evidence (Data Generation and QA Systems)Flatironhealth · NY officeFirst seen 29d agoremote
- Senior Scientist, Synthetic Data GenerationNvidia · 5 LocationsFirst seen today
- Procedural Data Generation EngineerVinci · Palo Alto HQFirst seen 12d ago
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job