hirly

Clera

Research Engineer, Benchmarks

Singapore

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Clera first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.6M live jobs from 190,000+ employers in 200+ countries.

Tailor my resume for this job →

Apply from your AI assistant

Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.

Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.

hirly's read of this role

Seniority
Mid level
Stated salary
$150,000 – $250,000 per year
Country
SG
Work mode
On-site / unstated
First seen by hirly
3 Oct 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

About the Role

Join a small technical research and engineering team building benchmarks for evaluating AI agents on realistic, domain-specific workflows. You will own the design and implementation of rigorous evaluations that help technical teams understand how agents perform in real-world tasks.

What You'll Do

Design, implement, and maintain benchmarks for evaluating AI agents on domain-specific tasks.

Work with subject-matter experts to translate real workflows into benchmark tasks and evaluation criteria.

Build and operate reliable infrastructure for running models and agents against tasks at scale.

Develop metrics and analyses to assess benchmark difficulty, reliability, and failure modes.

Validate whether benchmark results align with real-world performance and technical user needs.

Write clear documentation and reports for research and engineering audiences.

What We're Looking For

Two to four years of experience in research engineering, machine learning engineering, or a related technical role, including at least two years building AI benchmarks, evaluation infrastructure, or agent environments.

Strong Python skills and practical experience with Docker and Linux.

Experience designing, implementing, and running benchmarks or evaluation environments for AI agents or large language models.

Experience collaborating with subject-matter experts and analyzing workflows across technical or business domains.

Experience developing metrics, statistical analyses, or validation studies for evaluations.

Strong technical writing, attention to detail, and ability to work independently in an early-stage environment.

A technical educational background; experience with reinforcement learning pipelines, published evaluation work, or widely used benchmarks is a plus.

Compensation & Benefits

Salary range: USD 150,000 to 250,000 annually. Visa sponsorship is available.

Location

On-site in Singapore, Singapore.

Original posting on Clera's site ↗

Listed on hirly, a job board. hirly is not the employer: Clera is hiring for this role.

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job