River AI Inc.
Software Engineer, GPU Kernels
Palo Alto, CA
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.6M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →Apply from your AI assistant
Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.
Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.
hirly's read of this role
- Role family
- Engineering
- Seniority
- Mid level
- Stated salary
- $200,000 – $420,000 per year
- Country
- US
- Work mode
- On-site / unstated
- First seen by hirly
- 3 Oct 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
At River AI, our mission is to create personal AI owned and shaped by each individual. To achieve this, we are rewriting the entire stack from scratch: personal hardware for local inference, bespoke training infrastructure, next-generation UIs, and frontier deep learning research.
Who we are
We are scientists, engineers, and builders from the industry's top tech companies and AI labs. We bring a proven track record of scaling consumer systems for hundreds of millions of users and architecting the pre-training infrastructure behind today's frontier models.
About the Role
We are looking for exceptional GPU kernel engineers to build the compute primitives behind River’s training and inference infrastructure. Your goal is to make large models faster to train and more efficient to serve.
You will own performance-critical operations, including attention, matrix multiplication, mixture-of-experts execution, and low-precision computation. Working closely with researchers and systems engineers, you will identify bottlenecks, implement kernels, validate correctness, and bring improvements into production.
What You’ll Do
Build fast GPU kernels for attention, matrix multiplication, expert routing, and related operations.
Optimize memory access, tiling, and synchronization to make efficient use of GPU hardware.
Develop FP8, FP4, and mixed-precision kernels while preserving numerical correctness.
Accelerate fine-tuning and RL through optimized adapters, backward passes, and fused operations.
Profile real workloads and integrate improvements into training and inference runtimes.
Build reproducible benchmarks that verify correctness, gradients, and performance.
Skills & Qualifications
Minimum Qualifications:
Bachelor’s degree in Computer Science, Computer Engineering, or equivalent practical experience.
Experience optimizing GPU kernels with CUDA, Triton, CUTLASS, CuTe, or comparable tools.
Strong understanding of GPU architecture, memory hierarchies, and parallel execution.
Proficiency in C++ and Python.
Strong foundations in linear algebra, floating-point arithmetic, and numerical computing.
Strong debugging and profiling skills, with a collaborative approach to engineering.
Preferred Qualifications: (We encourage you to apply even if you don't meet all of these)
Experience optimizing for NVIDIA Blackwell or Hopper GPUs.
Work on attention, mixture-of-experts kernels, grouped GEMMs, or low-rank adapters.
Experience implementing backward passes and validating gradients.
Familiarity with FP8, FP4, and quantized weight layouts.
Experience integrating custom operators into PyTorch, SGLang, vLLM, or similar frameworks.
Open-source contributions or a track record of shipping substantial kernel optimizations.
Logistics & Benefits
Location: Palo Alto, California.
Compensation: Depending on experience and skills the expected base pay is $200,000 - $420,000 USD per year, plus equity.
Benefits: Comprehensive health, dental, and vision insurance; unlimited PTO; and relocation assistance as needed.
Visa Sponsorship: We sponsor visas and are committed to supporting the process for the right candidate.
Similar jobs
- Systems Software Engineer IACCO Engineered Systems · Pleasanton, CA, United StatesFirst seen today
- Software EngineerCapgemini · New York, USFirst seen today
- Software EngineerSpatial Front, Inc · College Park, MDFirst seen today
- Quantitative Software EngineerPmsi · Henderson, NevadaFirst seen today
- Software Engineer II – Java/Spring boot Engineer (Microservice Systems)WorldpayFirst seen today
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job