Cosine
ML Systems Engineer - Model Training and Infrastructure (SWE-focused LLMs)
London Office
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Apply from your AI assistant
Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.
Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.
hirly's read of this role
- Role family
- Engineering
- Seniority
- Mid level
- Stated salary
- £80,000 – £110,000 per year
- Country
- GB
- Work mode
- On-site / unstated
- First seen by hirly
- 9 Oct 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
Job title: ML Systems Engineer - Model Training and Infrastructure (SWE-focused LLMs)
Location: London; full in-office working as default
Start date: ASAP
Compensation: £80,000 - £110,000 Base Salary & £80,000 - £110,000 Share options.
___________________________________________________________________________
Help build the software engineers of the future
Cosine is building autonomous AI engineers that plan, write and ship code inside real development workflows.
Our agents work across complex software systems, and our Lumen models are trained to do more than produce code that looks correct. They are built to understand existing architectures, follow established patterns and produce software that engineers can actually maintain.
We develop our agent tooling entirely in-house and post-train open-source models for reliable, enterprise-grade coding performance. Our products are designed for on-premise, VPC and fully air-gapped environments, including security-critical settings where control, privacy and robustness are non-negotiable.
In 2024, Cosine achieved a 72% score on OpenAI’s SWE-Lancer benchmark, placing us among the strongest real-world software-engineering AI systems evaluated.
We’re now looking for an ML Systems Engineer to help train the next generation of Lumen models.
This is a highly hands-on role at the intersection of machine learning, software engineering, data and infrastructure. You’ll build the environments in which models learn to write software, develop the pipelines that generate and curate training data, and run the fine-tuning and reinforcement-learning workloads that shape model behaviour.
If you’re excited by the idea that the future quality of coding agents will be determined not just by model architecture, but by the quality of their data, environments and reward functions, this is an opportunity to work directly on that problem.
___________________________________________________________________________
The role
You’ll work closely with ML researchers, infrastructure engineers and product teams to decide:
What Lumen should learn next.
How to generate the right training data.
How to design environments that reflect real software-engineering work.
How to reward models for producing useful, maintainable code.
How to measure whether a new training run genuinely improves the experience for engineers.
The systems you build will sit directly inside our model-training loop. Models will write code, use tools, run tests and interact with real repositories. Your work will determine how those interactions are generated, evaluated and fed back into future training.
This is not a narrow research role and it is not traditional MLOps. You’ll move between custom PyTorch code, distributed data pipelines, Dockerised services, RL environments, evaluation infrastructure and production-quality software.
___________________________________________________________________________
What you’ll do
Build the training systems behind Lumen
Contribute to the end-to-end training of software-engineering models.
Implement supervised fine-tuning pipelines using curated code and conversation datasets.
Build reinforcement-learning loops in which models write code, run tests and use development tools.
Develop custom PyTorch dataloaders, training objectives and evaluation workflows.
Run and analyse fine-tuning and RL experiments across large, modern open-source models.
Create the data that teaches models to engineer
Develop synthetic data-generation pipelines for future RL and fine-tuning runs.
Design systems for producing, filtering, transforming and sampling large-scale datasets.
Work with object storage, dataset sharding and data-quality checks.
Investigate which examples, tasks and sampling strategies lead to better model behaviour.
Turn model failures into concrete improvements to training data and future experiments.
Build reliable RL infrastructure
Design, build and deploy containerised services that support model training and evaluation.
Use Docker and orchestration platforms such as Kubernetes to operate RL infrastructure.
Build environments where agents can modify repositories, run tests, use tools and receive meaningful feedback.
Improve the reliability, reproducibility and observability of large-scale training runs.
Work across Python, Go and the surrounding infrastructure required to run these systems.
Improve how we evaluate SWE models
Help maintain and extend evaluation suites for code models.
Build evaluations around unit tests, benchmark suites, repository-level tasks and real engineering workflows.
Analyse model failure modes, reward-hacking risks and regressions.
Develop better ways to measure code quality, maintainability, correctness and architectural fit.
Feed evaluation results back into model, data and infrastructure decisions.
Shape the next training direction
Work with research, infrastructure and product teams to identify the highest-value training problems.
Turn broad goals such as “make Lumen better at X” into clear experiments and measurable outcomes.
Develop opinionated reward functions for software-engineering agents.
Document decisions, communicate trade-offs and help the wider team understand what the results mean.
Take ownership of projects from initial idea through to deployment and iteration.
___________________________________________________________________________
What we’re looking for
You may come from software engineering, ML infrastructure, data engineering, applied machine learning or a closely related field.
You should be comfortable with:
Strong software engineering fundamentals
Typically 3–5 years of experience, or equivalent evidence of strong technical ability.
Reading, debugging and writing non-trivial production code.
Working primarily in Python and Go.
Caring about correctness, maintainability and code quality as much as model metrics.
Taking ownership of ambiguous technical problems and turning them into working systems.
Training frameworks and ML systems
Practical experience with at least one of PyTorch, TensorFlow or JAX.
Implementing custom training loops, losses or dataloaders.
Understanding the relationship between data, objectives, evaluation and model behaviour.
A willingness to work close to the code rather than treating training infrastructure as a black box.
Containers and cloud infrastructure
Experience with Docker and container-management or orchestration platforms such as Kubernetes.
Experience with at least one major cloud platform, such as GCP, AWS or Azure.
An understanding of how to build services that are reliable, observable and straightforward to operate.
Comfort working across application code, infrastructure and deployment systems.
Data engineering instincts
Experience working with large-scale datasets and object storage.
Understanding of sharding, filtering, sampling and dataset versioning.
The judgement to recognise that data quality can matter as much as model architecture.
An interest in building data pipelines that are reproducible and useful for experimentation.
Clear communication and ownership
The ability to explain technical decisions and experimental results clearly.
Comfort documenting trade-offs and walking others through your reasoning.
A practical approach to deciding what to build, what to measure and what to leave out.
The curiosity to investigate failures rather than dismissing them as noise.
You do not need to have worked on every part of the stack. We care more about strong engineering fundamentals, technical judgement and the ability to learn quickly than about matching every keyword.
Nice to have
You don’t need all of these, but experience in the following areas would help you get up to speed quickly:
Synthetic data generation for code or language models.
Training LLMs in
Listed on hirly, a job board. hirly is not the employer: Cosine is hiring for this role.
Similar jobs
- Principal Systems Engineer / Solutions Architect - Naval CommunicationsThales · 3 LocationsFirst seen 8d ago
- Lead Systems Engineer - Data DevOps/MLOpsEPAM Systems · Coimbatore, Tamil Nadu, IndiaFirst seen today
- Systems Engineer - Modelling & Simulation (Seeker)Mbda · 3 LocationsFirst seen today
- Systems Engineer - DefenceALTEN · Glasgow, Scotland, United KingdomFirst seen today
- Advanced RF Systems EngineerAerospace · Edinburgh, Midlothian, United KingdomFirst seen yesterday
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job