hirly

Thinking Machines Lab

Research Engineer, Web Crawling

San Francisco

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Thinking Machines Lab first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.4M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

Apply from your AI assistant

Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.

Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.

hirly's read of this role

Role family
Engineering
Seniority
Mid level
Stated salary
$350,000 – $475,000 per year
Country
US
Work mode
On-site / unstated
First seen by hirly
28 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

About Thinking Machines

The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.

About the Role

We're hiring a Software Engineer to build and own our web-crawling systems, from distributed collection at internet scale through filtering, deduplication, and deciding what data we keep.

The ideal candidate has built and scaled a web crawler or large-scale data-acquisition systems. In this role, you'll write and own production systems: the crawler itself, the infrastructure that runs it at scale, and the pipelines that turn raw crawls into usable pretraining data. You'll work closely with our pretraining and data teams to understand what's actually moving model quality, but this is fundamentally an engineering role, not a research one.

What You'll Do

Design and scale the web crawler and ingestion infrastructure that sources Inkling's pretraining data

Build pipelines for large-scale extraction, deduplication, and data quality filtering

Build specialized crawlers for high-value or hard-to-reach data sources

Work with the pretraining team to understand how changes in crawled data affect model performance

Improve the reliability and efficiency of crawling and ingestion infrastructure at petabyte scale

Help set technical direction for this area as it grows, and bring other engineers up to speed on what you've learned

Skills & Qualifications

Minimum Qualifications

8+ years designing, building, and scaling web crawlers, scrapers, or large-scale distributed data-acquisition systems

A track record of owning crawler or data-acquisition infrastructure at internet scale

Strong software engineering skills in a language such as Python, Go, or Rust, with real experience in distributed systems

Working knowledge of the practical and legal considerations of large-scale web data collection (robots.txt, rate limiting, licensing)

Preferred Qualifications

Experience applying machine learning to crawl selection, extraction, or data quality classification at internet scale

Experience setting technical direction for a crawling, data acquisition, or search infrastructure team, whether or not that was your formal title

Experience designing systems for petabyte-scale storage and processing

Track record of open-source contributions to crawling, scraping, or data infrastructure tools

Background at a search engine (crawling, indexing) or a frontier AI lab's data acquisition team

Logistics

Location: This role is based in San Francisco, CA.

Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $350,000-$475,000 USD.

Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.

Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

As set forth in Thinking Machines' Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law.

Thinking Machines Lab will consider for employment qualified applicants with criminal histories in a manner consistent with the requirements of the California Fair Chance Act, the San Francisco Fair Chance Ordinance, and any other applicable state or local fair chance ordinance or law.

Original posting on Thinking Machines Lab's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job
Research Engineer – Thinking Machines Lab | hirly.me