hirly

Anyone AI

Machine Learning Engineer – ML Evaluation & Experiment Design

Argentina - Fully Remote · Uruguay · Chile · Ecuador - Fully Remote · Portugal · Mexico - Fully Remote · Colombia - Fully Remote · Spain · Brazil

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Anyone AI first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.3M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

hirly's read of this role

Role family
Data & ML
Seniority
Mid level
Countries
AR, UY, CL, EC, PT, MX, CO, ES, BR
Work mode
Remote-friendly
First seen by hirly
28 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

Anyone AI is recruiting experienced Machine Learning Engineers for a specialized project focused on reviewing and evaluating machine learning challenges used in AI model training and evaluation.

The work involves analyzing ML experiments, datasets, metrics, and pipelines to determine whether challenges are technically sound, reproducible, appropriately difficult, and genuinely require strong machine learning reasoning.

What You’ll Work On

You’ll review ML challenges involving:

Experiment design and model selection

Small and synthetic datasets

Data quality and preprocessing

Distribution shift and data contamination

Label noise and feature leakage

Model evaluation and metric selection

Hyperparameter tuning

Train / validation / test methodology

Reproducibility and deterministic pipelines

Statistical significance of model improvements

A key part of the role is determining whether a challenge actually rewards good ML reasoning , rather than simply being solvable through brute-force model selection or large hyperparameter searches.

What We’re Looking For

3+ years of hands-on applied machine learning experience

Strong experience with:

ML experiment design

Model selection

Hyperparameter tuning

Model evaluation

Data preprocessing and validation

Strong understanding of train, validation, and test splits

Ability to identify:

Data leakage

Label noise

Distribution shift

Spurious correlations

Feature leakage

Data contamination

Experience evaluating whether performance improvements are statistically meaningful rather than random fluctuations

Strong understanding of ML evaluation metrics and when different metrics are appropriate

Experience debugging ML workloads across CPU and GPU environments

Ability to analyze technical problems and provide clear written feedback

Nice to Have

Experience creating or participating in Kaggle, DrivenData, or similar ML competitions

Experience designing benchmark datasets or ML challenges

Background in data-centric AI or dataset quality

Experience with synthetic data generation and validation

Familiarity with statistical testing, confidence intervals, and effect sizes

Experience with ML evaluation pipelines, RLHF, or AI model evaluation

Experience developing ML curricula or technical assessments

Understanding of common ML failure modes such as:

Shortcut learning

Spurious correlations

Goodhart’s Law

Simpson’s paradox

Metric gaming

What You’ll Be Responsible For

Reviewing ML challenges and determining whether they are well designed and technically solvable

Evaluating whether datasets contain meaningful and learnable signals

Identifying unintended shortcuts or artifacts in synthetic datasets

Determining whether tasks require genuine diagnosis of the underlying ML problem

Reviewing evaluation metrics and improvement thresholds

Detecting metric gaming, data leakage, and evaluation flaws

Verifying reproducibility across the complete data → model → evaluation pipeline

Assessing whether challenge difficulty is appropriately calibrated

Providing clear recommendations for improving, recalibrating, or excluding problematic tasks

Engagement

  • Work Type: Remote
  • Engagement: Part-time, project-based consulting
  • Focus: Applied machine learning, experiment design, data quality, and model evaluation

This role is a strong fit for ML engineers who enjoy debugging experiments, understanding why models succeed or fail, identifying problems in datasets and evaluation pipelines, and designing rigorous machine learning experiments.

Original posting on Anyone AI's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job