hirly

Anyone AI

GPU Kernel Engineer – CUDA, Triton & Accelerator Performance

Argentina - Fully Remote · Uruguay · Chile · Ecuador - Fully Remote · Portugal · Mexico - Fully Remote · Colombia - Fully Remote · Spain · Brazil

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Anyone AI first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.3M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

hirly's read of this role

Seniority
Mid level
Countries
AR, UY, CL, EC, PT, MX, CO, ES, BR
Work mode
Remote-friendly
First seen by hirly
28 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

Anyone AI is recruiting experienced GPU Kernel Engineers for a specialized project focused on reviewing, debugging, and evaluating high-performance compute kernels used in AI workloads.

We’re looking for engineers with hands-on experience writing and optimizing kernels across frameworks such as CUDA, Triton, NKI, or Pallas , with a strong understanding of numerical correctness, GPU performance, memory optimization, and benchmarking.

What You’ll Work On

You’ll work with GPU and accelerator kernel tasks involving:

Kernel implementation and debugging

CUDA and Triton optimization

Translation between kernel frameworks

Hardware migration

Operator fusion

Performance profiling and benchmarking

Numerical correctness verification

Compilation and runtime debugging

Memory hierarchy optimization

Kernel-level AI workload performance

You’ll assess whether implementations are technically correct, efficiently designed, reproducible, and appropriately optimized for the target hardware.

What We’re Looking For

3+ years of hands-on experience developing, optimizing, or debugging GPU or accelerator kernels

Strong experience with at least two of the following:

CUDA

Triton

NKI / AWS Neuron

Pallas / JAX

Strong understanding of GPU performance optimization

Experience with kernel profiling tools such as Nsight, NCU, roofline analysis, or framework-native profilers

Understanding of:

Memory bandwidth

Compute throughput

GPU occupancy

Shared memory

Register pressure

Memory coalescing

Bank conflicts

Strong understanding of floating-point numerical correctness and tolerance thresholds

Experience debugging kernel compilation and runtime issues

Ability to distinguish software defects, environment problems, and genuine optimization challenges

Relevant Experience

Candidates should have experience with several of the following types of work:

Writing kernels from technical specifications

Translating kernels between CUDA, Triton, or other frameworks

Migrating kernels across hardware platforms

Debugging incorrect kernel implementations

Optimizing kernel performance

Fusing multiple operations into optimized kernels

Nice to Have

Experience across both NVIDIA GPU and custom accelerator ecosystems

Experience with AWS Trainium, TPU, JAX, or other accelerators

Compiler engineering experience

Familiarity with MLIR, XLA, or intermediate representation lowering

Contributions to GPU or ML kernel libraries

Experience with cuBLAS, cuDNN, Triton community kernels, or JAX/XLA custom calls

Experience with AI model evaluation, RLHF, or technical benchmark development

What You’ll Be Responsible For

Reviewing GPU and accelerator kernel implementations for correctness

Comparing outputs against reference implementations

Evaluating numerical tolerance thresholds

Reviewing kernel benchmarks and determining whether comparisons are fair

Identifying performance bottlenecks and optimization opportunities

Assessing whether performance targets are realistic given hardware limits

Reviewing kernel translations and hardware migrations

Identifying compilation, driver, memory, shape, and runtime issues

Determining whether technical tasks are genuinely difficult or incorrectly configured

Providing clear, actionable technical feedback

Engagement

  • Work Type: Remote
  • Engagement: Part-time, project-based consulting
  • Focus: GPU kernels, performance engineering, debugging, and technical evaluation

This role is ideal for engineers who enjoy working close to the hardware, optimizing GPU workloads, debugging low-level performance issues, and pushing AI compute systems toward their performance limits.

Original posting on Anyone AI's site ↗

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job