Firecrawl
Research Engineer - Benchmarks
San Francisco HQ · Toronto Hub
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.3M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →hirly's read of this role
- Seniority
- Mid level
- Stated salary
- $250,000 – $290,000 per year
- Countries
- US, CA
- Work mode
- On-site / unstated
- First seen by hirly
- 30 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
Research Engineer - Benchmarks
Every week, someone asks which data provider is actually best: for company records, for finding the right person, for fresh job listings. Right now the answers come from the vendors themselves. You'll build the independent version. You'll run rigorous, automated, public benchmarks of the providers in Alexandria and beyond, and publish them weekly. The same results will feed straight back into Alexandria so it learns which provider to call for which job.
This isn't our internal evals role. You're measuring the market, in public, where every number will be challenged by the vendors it ranks. You'll own it end to end: the datasets, the ground truth, the scoring, the harness, the weekly release, and the loop into Alexandria's provider selection. No one hands you a methodology. You write it, defend it, and ship it every week.
Salary Range: $250,000–$290,000 USD/year (SF) / $210,000–$224,000 CAD/year (Toronto)
Equity Range: Competitive equity. Details shared during the process.
Location: San Francisco, CA (SF HQ) or Toronto, ON (Toronto Hub). Hybrid, onsite 3+ days a week.
Equity Range: Competitive equity. Details shared during the process.
Location: San Francisco, CA (SF HQ). On-site, five days a week.
Job Type: Full-Time
Experience: 4+ years in ML, research engineering, or data engineering, with evaluation or benchmark work you've shipped
Work Authorization: Must be authorized to work in the United States or Canada. We're not able to sponsor US visas right now. For Canada, we'll consider sponsorship on a case-by-case basis through our Toronto Hub.
About Firecrawl
Firecrawl is the easiest way to turn the web into data AI agents can use. One API call converts any URL into clean, LLM-ready markdown or structured data. It's the boring-hard problem everyone building with LLMs eventually hits, solved.
In September 2026, we raised a $75M Series B led by Smash Capital, and we're spending it building the largest repository of knowledge in the world. We hit 8 figures in ARR in year one and more than doubled it in year two. We have 187k+ GitHub stars, putting us in the top 40 repositories of all time, and developers, agents, and category-defining AI companies build on us every day. Growth like this is rare, and we're just getting started.
We're a small team punching far above our weight. Everyone here owns a real piece of the product and company, end to end, and runs it themselves. No hiding behind process or headcount.
This is a place for people who want to work at the frontier: an AI company building the infrastructure other AI companies run on, not one bolting AI onto an existing product. We move fast, go deep, and are building the tools superintelligence will rely on to gather data from the web. That library is called Alexandria, and it starts now.
What You'll Do
Design and run head-to-head benchmarks of data providers against verified ground truth. For example:
Apollo vs. FullEnrich vs. DataLegion on company details.
FullEnrich vs. DataLegion on finding the right person and their current employer and title.
Built In vs. ZipRecruiter on relevant, fresh, non-duplicate job listings.
Build and maintain the test datasets and ground truth, and keep them from going stale or leaking.
Own the automated harness and the weekly benchmark release. Every run has to be reproducible, versioned, and defensible.
Measure what buyers actually care about: accuracy, coverage, freshness, speed, and cost.
Work with marketing to ship public leaderboard pages that are useful and hold up under scrutiny.
Close the loop into Alexandria so benchmark results change which provider gets called for what.
Keep expanding the categories we benchmark, inside Alexandria and beyond it.
What We're Looking For
You've shipped evals or benchmarks, and you can explain exactly why your results were trustworthy.
You've done this somewhere that matters: a frontier lab, a data company like Scale, Surge, micro1 or Mercor, a third-party benchmark org, or a public benchmark project. OSS contributors very welcome.
You're strong in Python and API integrations, and comfortable with messy vendor APIs, rate limits, and inconsistent schemas.
You know how to build test sets and scoring methods: sampling, labeling, inter-rater agreement, and when to trust an LLM judge and when not to.
You use AI heavily and keep upgrading your own workflow. You ship without waiting for instructions.
You write findings clearly for engineers, marketers, and the vendors on the other side of the leaderboard.
What We're Not Looking For
Someone who wants to run benchmarks someone else designed.
Someone who picks the metric that makes the story look good. Our numbers have to survive the vendors who lose.
A pure researcher who won't build the harness, or a pure engineer who won't think hard about methodology.
Someone who needs a fully specced ticket, or a quarter, to ship the first result.
A Note On Pace
We operate at an absurd level of urgency because the window for what we're building won't stay open forever. If that excites you, keep reading. If it doesn't, no hard feelings, but this role probably isn't for you.
Benefits & Perks
Available to all employees
Salary that makes sense: $250,000–$290,000 USD/year, / $210,000–$224,000 CAD/year (Toronto), based on impact, not tenure
Own a piece: Gain competitive equity in what you're helping build
Generous PTO: 15 days mandatory, anything after 24 days, just ask (holidays excluded). Take the time you need to recharge
Parental leave: 12 weeks fully paid, for all parents
Wellness stipend: $100 USD/month for the gym, therapy, massages, or whatever keeps you human
Learning & Development: Expense up to $1,000 USD/year toward anything that helps you grow professionally
Team offsites: A change of scenery, minus the trust falls
Sabbatical: 3 paid months off after 4 years, do something fun and new
Available to US-based full-time employees
Full coverage, no red tape: Medical, dental, and vision (100% for employees, 50% for partner and kids). No weird loopholes, just care that works
Life & Disability insurance: Employer-paid basic life and AD&D, short-term disability, and long-term disability. Coverage for life's curveballs
Virtual care and a health guide: Teladoc for the couch doctor visit, plus Rightway to answer coverage questions and fight billing errors for you
Mental health: Talkspace, therapy and psychiatry on your schedule
Fertility and family building: Carrot, covering you and your partner
EAP: Free confidential counseling, legal and financial consults, and online will prep through Guardian
401(k) plan: Retirement might be a ways off, but future-you will thank you
Pre-tax benefits: HSA, FSA, and commuter benefits to help your wallet out a bit
Supplemental options: Extra life and AD&D, accident, critical illness, hospital indemnity, plus pet, legal, and identity protection through MetLife
Available to Canada-based full-time employees
Full coverage, no red tape: Extended health, dental, and vision through Manulife (Diamond, the top tier), 100% employer-paid for you, your partner, and your kids
Life & Disability insurance: Employer-paid life, AD&D, short-term disability, and long-term disability. Coverage for life's curveballs
Virtual care: Dialogue Premium, so you can see a doctor or nurse from your couch, any hour
Mental health: Talkspace Elite, therapy and psychiatry on your schedule
Fertility and family building: Carrot, covering you and your partner
Retirement: Group RRSP through Wealthsimple, so future-you can thank you
Available to SF-based employees
SF HQ perks: Snacks, drinks, team lunches, intense ping pong, and peak startup energy
E-Bike transportation: A loaner electric bike to get you around the city, on us
Available to Toronto-based employees
Toronto Hub perks: Snacks, drinks, team lunches, glass-
Similar jobs
- Co-Op, Research Engineering and AI IntegrationModernatxFirst seen today
- Applied Research EngineerHigharc · Remote (New York City, NY)First seen todayremote
- Research Engineer / Research Scientist, RL FrontiersAnthropic · San Francisco, CA | New York City, NY | Seattle, WAFirst seen todayremote
- Research Engineer / Performance Engineer, RL Distributed SystemsAnthropic · San Francisco, CA | New York City, NY | Seattle, WAFirst seen todayremote
- Research Engineer - DistributionSouthern Company · Birmingham, AL, United States; Atlanta, GA, United StatesFirst seen today
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job