hirly

Turing

Research Engineer

Colombia, Huila, Colombia

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Turing first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.5M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

Apply from your AI assistant

Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.

Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.

hirly's read of this role

Seniority
Mid level
Country
CO
Work mode
Remote-friendly
First seen by hirly
2 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

About Turing

Turing’s mission is to accelerate superintelligence to drive real economic progress. Headquartered in San Francisco, Turing works with frontier AI labs to generate high-quality datasets, reinforcement learning environments, and frontier research benchmarks that improve model capabilities in software engineering, enterprise knowledge work, and advanced STEM reasoning. In software engineering, Turing is the largest and longest-running data provider in the category. Turing also works with Fortune 500 enterprises across financial services, life sciences, healthcare, retail, automotive, and CPG to build and deploy end-to-end agentic AI systems inside mission-critical workflows. By operating on both sides, Turing closes the loop between frontier research and enterprise deployment, turning real-world deployment signals into better data, evaluations, and more capable models. Learn more at www.turing.com .

*This is a remote role and can be performed anywhere in Colombia.*

The Role

We are looking for a Research Engineer to help deliver frontier-quality datasets, RL environments, and evaluations that improve state-of-the-art models for leading AI labs and enterprise clients.

This is a hands-on, research-facing technical leadership role. You will work directly with customer researchers & engineers to translate their model and post-training goals into concrete data and environment specifications, and drive the production of data that meets extremely high standards for correctness, realism, diversity, difficulty, and measurable model lift.

This role is designed for candidates with roughly 4 to 5 years of experience building and improving deep learning systems, especially where strong results depend on data quality, data curation, denoising, synthetic data generation, and rigorous evaluation. You’ll operate in one or more of the following capability areas:

Coding and software engineering agents (repositories, unit tests, debugging, tool use, code reviews, long-horizon workflows)

RL environments and verifier-based training (tasks, rewards/verifiers, trajectories, evaluation harnesses)

Multimodal data and reasoning (text + images + documents + tables/charts; optional audio/video)

STEM reasoning (math, physics, chemistry, bio, engineering – solution verification and error analysis)

Modern embodied AI / VLM-driven agents (vision-language(-action) models, embodied task suites, tool/sensor/action abstractions, long-horizon interaction data)

What You’ll Do

1) Own data and environment quality from an AI researcher perspective

Translate ambiguous research goals into clear data requirements: target skills, failure modes, difficulty calibration, coverage, and success metrics.

Define what “good” looks like by creating detailed rubrics, counterexamples, and boundary cases (what to include vs. exclude).

Perform deep, detail-oriented audits of produced data: spot subtle errors, reward hacking opportunities, leakage, ambiguity, inconsistent assumptions, and distribution shifts.

Drive iterative improvements using evidence: error taxonomies, slice-based quality metrics, and model-behavior-informed refinements.

2) Design and build datasets and RL environments for your capability area(s)

Contribute to or lead the design of:

Task suites (single-step and long-horizon workflows)

Ground-truth signals (verifiers, unit tests, structured checks, reward functions, automatic validators)

Environment interfaces (APIs, tool schemas, state abstractions, database schemas, simulator-like dynamics)

Depending on your mapped capability area(s), you may focus on:

Coding / SWE agents: data reflecting real development work (codebase navigation, bug localization, patching, tests, code reviews, CI-like constraints, refactors, security fixes).

Multimodality: tasks that test true multimodal reasoning (chart reading, document QA, UI understanding, diagram-based STEM reasoning, OCR-aware tasks).

STEM: tasks with verifiable solutions (symbolic checks, reference solvers, numerical validation, step consistency, unit sanity).

Modern embodied AI / VLM-driven agents: interaction data and environments for vision-language(-action) models (long-horizon tasks, instruction following grounded in visual context, robust action selection, safety/constraint adherence, adversarial state coverage).

3) Build robust validation, denoising, and synthetic data systems

Implement automated validation and filtering to achieve frontier-grade signal-to-noise:

Deduplication, decontamination, leakage checks

Consistency checks (format, schema, invariants)

Difficulty and diversity controls (coverage, novelty, long-tail)

Develop synthetic data generation and augmentation pipelines where appropriate:

Programmatic task generators

Controlled perturbations to create hard negatives

Scenario templating with diversity constraints

Simulator-/tool-driven rollouts for trajectory data

Create documentation and data cards: dataset intent, known limitations, recommended use, and evaluation linkage.

4) Use evaluations and training runs to prove impact

Design and run evals that reflect the customer’s intended usage.

Produce analysis that connects data to outcomes:

Pre/post comparisons on targeted capability slices

Error breakdowns and “why the model failed” narratives

Ablations to identify which data attributes drive lift

When needed, run in-house fine-tuning or RL-style experiments (or partner with research) to demonstrate that the data/environment improves model behavior in measurable ways.

5) Collaborate effectively with large production teams without being ops-heavy

Work with cross-functional teams (engineers, researchers, QAs, domain SMEs, and large-scale data production groups) by providing:

Clear specs, examples, and edge cases

Fast feedback loops based on audits and quantitative signals

Structured review processes focused on quality, not throughput alone

You are expected to be highly engaged in reviewing and improving outputs from large annotation/creation efforts, but not primarily responsible for hiring, staffing, or people operations.

Who We’re Looking For

4–5 years of experience building or improving deep learning systems where data quality mattered materially (training, post-training, evals, or agentic systems).

Strong intuition for the “data ingredients” that drive model improvements: what to collect, what to filter, what to synthesize, and how to measure.

Ability to communicate clearly with researchers and engineers: turning research objectives into concrete specs, and turning messy outputs into actionable insights.

Demonstrated ability to be extremely detail-oriented in diagnosing subtle data quality issues and failure modes.

Solid programming ability with a bias for shipping:

Python proficiency required

Comfort with SQL/structured data workflows strongly preferred

For coding-focused work: proficiency in one or more major languages (e.g., C++, Java, Go, Rust, JS/TS) is a plus

Comfort designing quality systems:

Rubrics, validation scripts, gold sets, sampling strategies

Statistical checks and slice-based evaluation

Human-in-the-loop review loops grounded in measurable criteria

Strong pluses

RL or post-training experience (any of: RLHF/RLAIF, verifier training, reward modeling, RL fine-tuning, environment design).

Experience with agentic evaluation (tool use, multi-step workflows, long-horizon tasks, trajectory analysis).

Multimodal expertise (document understanding, charts, diagrams, OCR, UI/vision grounding; audio/video optional).

STEM depth (math/physics/engineering) with an eye for verifiability and rigorous correctness.

Modern embodied AI / VLM-driven agent experience (vision-language(-action) models, interaction datasets, embodied evals, long-horizon grounding, tool/sensor/action interfaces).

Systems thinking: ability to “simulate” an appl

Original posting on Turing's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job