hirly

Fundamental

Senior Applied Research Engineer

Barcelona

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Fundamental first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.3M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

hirly's read of this role

Seniority
Senior
Country
ES
Work mode
On-site / unstated
First seen by hirly
28 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

About Fundamental

Fundamental is an AI research lab pioneering the future of enterprise decision-making. Our flagship model, NEXUS is the world's most powerful Large Tabular Model (LTM) - purpose-built for the structured records that contain trillions of dollars in business value. With $275m in funding from leading investors and trusted by Fortune 100 companies, Fundamental is giving businesses the Power to Predict.

At Fundamental, you'll work on unprecedented technical challenges in foundation model development and build technology that transforms how the world's largest companies make decisions. This is your opportunity to be part of a category-defining company from the ground-up. Join the team defining the future of enterprise AI.

Key responsibilities

Profile end-to-end distributed training runs to identify bottlenecks across compute, GPU memory, and inter-GPU communication

Contribute to architectural decisions that improve the efficiency and reliability of large-scale training jobs, including developing Triton/CUDA kernels when needed

Design and implement model scaling, parallelization, and memory optimization techniques for training workloads with very large context sizes

Collaborate closely with ML Researchers to diagnose architectural inefficiencies, ensure new research ideas scale efficiently in practice, and spread internal knowledge about model efficiency and optimization

Drive the productionization and serving of our models from the research side, including improving inference efficiency through techniques such as quantization

Must have

Strong understanding of modern ML architectures and large-scale training pipelines

Experience running distributed training jobs on multi-GPU systems

Advanced profiling and debugging skills across CPU, GPU, memory usage, latency, and inter-GPU communication

Strong programming skills in Python

Experience with model scaling and parallelization strategies, including tensor and pipeline parallelism

Nice to have

Familiarity with NCCL, MPI, and distributed communication primitives

Knowledge of PyTorch and Triton internals

Programming experience with C++ and CUDA

Benefits

Competitive compensation with salary and equity

Comprehensive health coverage for you and your dependents

Paid parental leave for all new parents, inclusive of adoptive and surrogate journeys

Relocation support for employees moving to join the team in one of our office locations

A mission-driven, low-ego culture that values diversity of thought, ownership, and bias toward action

Original posting on Fundamental's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job