hirly

Hcompany

Research Engineer, Model Inference & Serving - Paris

Hybrid Paris

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Hcompany first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.3M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

hirly's read of this role

Seniority
Mid level
Country
FR
Work mode
On-site / unstated
First seen by hirly
10 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

Research Engineer, Model Inference & Serving

About H: H exists to push the boundaries of superintelligence with agentic AI. By automating complex, multi-step tasks typically performed by humans, AI agents will help unlock full human potential. H is hiring the world's best AI talent, seeking those who are dedicated as much to building safely and responsibly as to advancing disruptive agentic capabilities. We promote a mindset of openness, learning, and collaboration, where everyone has something to contribute.

About the Team: The Inference team builds and operates the systems that serve H's foundational models in production. We focus on multimodal inference and serving for Computer Use Agents, optimizing across both the inference engine layer (e.g., vLLM, SGLang) and the model serving layer (e.g., disaggregated inference, intelligent routing). Agentic inference brings constraints around context length, multimodality, and tool calls, which we address by co-designing with the Models team on training-time choices and with the agent teams on how models are deployed. We operate at the intersection of research and production, translating cutting-edge inference techniques into the systems that power H's next generation of agents. We are looking for strong engineers excited about inference to join the team and help shape the systems behind superintelligent AI.

Key Responsibilities:

Build and operate the inference stack that serves H's multimodal agentic models

Improve latency, throughput, and cost of model serving across the stack

Research and implement inference techniques tailored to agent workloads

Co-design with the Models team on training-time decisions that affect inference

Collaborate with cross-functional teams to integrate inference into agentic AI products

Evaluate inference, serving, and hardware platforms, and communicate findings to stakeholders

Stay current with advancements in inference, model serving, and accelerator technology

Requirements:

Technical skills:

Strong software engineering track record

Proficient in Python and at least one systems language (Rust, C++, or Go)

Hands-on experience with deep learning frameworks (PyTorch, JAX), preferably in an industry setting

Solid distributed systems fundamentals

Experience working in a modern cloud environment and with production ML infrastructure (Kubernetes, etc.)

Working knowledge of modern ML, including transformers and multimodal architectures

Research skills:

Research engagement: an advanced degree with research output, or publications at top-tier AI or systems venues (e.g., NeurIPS, ICML, MLSys, OSDI), research internships, or substantive open-source contributions

Soft skills:

Excellent communication and presentation skills

Strong collaboration and teamwork skills

Passion for inference and AI

Preferred qualifications:

Startup experience

Hands-on experience with inference frameworks (vLLM, SGLang, TensorRT-LLM)

Writing or modifying GPU kernels (CUDA, Triton, etc.)

Edge or on-device inference experience (llama.cpp, MLX, ONNX Runtime, etc.)

Experience with quantization, speculative decoding, disaggregated inference or KV-cache compression

Experience with multimodal models and/or agentic systems

Location:

Paris or London.

This role is hybrid, and you are expected to be in the office 3 days a week on average.

Please expect some travel between offices on a reasonable cadence (e.g., every 4-6 weeks).

What We Offer:

Join the exciting journey of shaping the future of AI

Collaborate with a fun, dynamic and multicultural team, working alongside world-class AI talent in a highly collaborative environment

Enjoy a competitive salary

Unlock opportunities for professional growth, continuous learning, and career development

If you want to change the status quo in AI, join us.

Original posting on Hcompany's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job