Nvidia
Software Engineer Intern, AI and DL Kernel Libraries - 2027
China, Shanghai
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.3M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →hirly's read of this role
- Role family
- Engineering
- Seniority
- Internship
- Country
- CN
- Work mode
- On-site / unstated
- First seen by hirly
- 30 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
NVIDIA is looking for outstanding Software Engineer Interns to help develop groundbreaking technologies for AI and deep learning kernel libraries. Our team builds core software that accelerates high-impact AI workloads on NVIDIA GPUs, with a strong focus on deep learning primitives, kernel libraries, and performance-critical GPU software. As an intern on the team, you will contribute to the design, development, optimization, and delivery of software that powers NVIDIA's AI platform.
This internship is centered on foundational library engineering, with opportunities to work on low-level kernels, performance primitives, and efficient implementations for modern AI and deep learning workloads. You may contribute to GPU-accelerated deep learning primitives, attention kernel implementations, runtime components, code generation systems, and other performance-critical infrastructure for large language models and advanced AI applications. You will collaborate with world-class engineers across deep learning software, compilers, GPU architecture, and open-source inference ecosystems, and your work can directly impact the performance of real-world workloads at scale.
What you'll be doing
- Contribute to production-quality software that ships as part of NVIDIA's AI software stack, including cuDNN, FlashInfer, and optimized support for large language model inference workloads.
- Help develop new AI systems technologies for efficient inference, with a focus on performance, scalability, maintainability, and usability.
- Support the design, implementation, and optimization of kernels for high-impact AI workloads across LLM inference, generative AI, computer vision, autonomous driving, and recommender systems.
- Assist in building extensible software abstractions for deep learning libraries, LLM serving engines, and runtime systems.
- Contribute to just-in-time compilation, code generation, and runtime technologies for performance-critical GPU workloads.
- Analyze workload performance, tune current software, and help propose improvements to future software and hardware-software interfaces.
- Collaborate closely with engineers across deep learning frameworks, libraries, kernels, compilers, and GPU architecture teams at NVIDIA.
- Contribute to open-source communities and ecosystem integrations where relevant, including projects such as FlashInfer, vLLM, and SGLang.
What we need to see
- Currently pursuing a Bachelor's, Master's, or PhD degree in Computer Science, Electrical Engineering, or a related field.
- Coursework, research, or hands-on project experience in machine learning, deep learning systems, compilers, systems software, or GPU programming.
- Strong programming skills in C/C++ and Python.
- Familiarity with CUDA development and GPU programming fundamentals.
- Experience developing with or using deep learning frameworks such as PyTorch, JAX, TensorFlow, or ONNX.
- Understanding of linear algebra, performance analysis, profiling, and code optimization.
- Interest in software abstractions, APIs, and higher-level system architecture for performance-sensitive systems.
- Interest in modern machine learning and inference system trends, especially around LLMs and generative AI.
- Strong problem-solving skills, curiosity, and the ability to work effectively in a collaborative environment.
Ways to stand out from the crowd
- Hands-on experience with inference engines and runtimes such as vLLM, SGLang, MLC, TensorRT-LLM, or similar systems.
- Background in domain-specific compilers, code generation, or library solutions for LLM inference and training.
- Exposure to machine learning compilers or IR systems such as MLIR, Apache TVM, TensorIR, or related technologies.
- Practical experience with GPU performance modeling, computer architecture, or accelerator-oriented software design.
- Open-source project ownership or meaningful contributions in deep learning systems, compilers, kernels, or inference infrastructure.
Similar jobs
- Software Engineer Intern 磁共振软件开发实习生Philips · SuzhouFirst seen 7d ago
- MCU Software Test Engineer InternNxp · SuzhouFirst seen 18d ago
- AI/ML Software Engineer InternAgilent · China-ShanghaiFirst seen 29d ago
- 2027 Embedded Software Engineer Intern – Rolling Meadows ILNgc · United States-Illinois-Rolling MeadowsFirst seen today
- Quality Assurance Engineer InternSertis · Bangkok, Bangkok Metropolis, ThailandFirst seen todayremote
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job