Samsungsemiconductor
Technical Director, Large-Scale AI Model Inferencing
San Jose, California, United States
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.6M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →Apply from your AI assistant
Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.
Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.
hirly's read of this role
- Seniority
- Director
- Stated salary
- $219,000 – $351,000 per year
- Country
- US
- Work mode
- Remote-friendly
- First seen by hirly
- 21 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
Please Note:
To provide the best candidate experience amidst our high application volumes, each candidate is limited to 10 applications across all open jobs within a 6-month period.
Advancing the World’s Technology Together
Our technology solutions power the tools you use every day--including smartphones, electric vehicles, hyperscale data centers, IoT devices, and so much more. Here, you’ll have an opportunity to be part of a global leader whose innovative designs are pushing the boundaries of what’s possible and powering the future.
We believe innovation and growth are driven by an inclusive culture and a diverse workforce. We’re dedicated to empowering people to be their true selves. Together, we’re building a better tomorrow for our employees, customers, partners, and communities.
What You’ll Do
Inference is becoming a memory-bandwidth business. As models scale past what any single GPU can hold — KV caches grow with context, MoE expert weights spill beyond HBM, and new architectures change the rules of what "model state" even means — the winners will be the companies that treat memory as the core product of AI inference , not an afterthought.
We are looking for a Hands-on Principal Engineer who combines deep, first-principles knowledge of AI model architectures (dense Transformers, Mixture-of-Experts, State Space Models, and hybrids) with production-scale inference expertise , to own the requirement for full-stack AI memory solutions at scale — spanning GPU HBM, host DRAM, CXL-attached memory pools, and NVMe/SSD tiers and Samsung Cognos, AI memory software that moves model state intelligently across them.
This person will be the technical authority who connects model behavior to memory-system design: someone who can explain why an MoE router's activation pattern dictates an LRU expert cache policy, why a Mamba state cache breaks the assumptions of PagedAttention, and why disaggregated prefill/decode changes the required memory bandwidth per token by an order of magnitude — and then build the products that exploit those facts.
- Level: Principal Engineer
- Team: Memory Solutions Lab / Data Fabric Solutions
- Reports to: Chief Technologist, Memory Solutions Lab
Location: Daily onsite presence preferred at our San Jose office/headquarters in alignment with our Flexible Work policy; remote/hybrid option available.
Job ID : 43027
Model Architecture Expertise — The Foundation
Serve as expert on how different model families consume and move memory, and translate that into memory-product requirements:
Dense Transformers : MHA/MQA/GQA/MLA attention, KV-cache growth characteristics, long-context behaviors, attention sinks and prefix locality.
Mixture-of-Experts : routed vs. shared experts, expert-parallel execution, routing skew and hot-expert locality, expert-weight offloading and cache-admission policies, per-token weight-read economics.
State Space Models (Mamba/Mamba-2) and hybrid SSM-attention architectures : recurrent state vs. KV cache semantics, state size per sequence and per layer, cache-swapping behavior for context switching and batching, and what "cache-aware scheduling" means when the state is a fixed-size tensor instead of a token-indexed table.
Emerging architectures : linear attention, sliding-window/hybrid layers, diffusion and multimodal transformers — and how each changes the memory hierarchy math.
Model the memory footprint, bandwidth demand, and access patterns of frontier open-weight models (e.g., Llama/Qwen-class dense, DeepSeek/Kimi-class MoE, Jamba-class hybrids) and publish internal reference architectures for each.
Track the model landscape as a roadmap input: anticipate what coming architectures (longer contexts, agentic multi-session reuse, reasoning-loop workloads, speculative decoding drafts) will demand from memory systems 12–24 months out.
Large-Scale Inference Expertise
Own deep expertise in production inference stacks — SGLang (HiCache), vLLM (PagedAttention, LMCache integration), NVIDIA Dynamo, TensorRT-LLM, llama.cpp-class engines — including their memory-management internals, not just their flags.
Drive inference performance engineering: continuous batching, chunked prefill, disaggregated prefill/decode, prefix and radix caching, speculative decoding, CUDA Graphs, and their interactions with memory tiering.
Own the latency/throughput/cost envelope: TTFT and TBT/TPOT SLOs, tokens-per-second per dollar, GPU memory utilization as the binding constraint, and the tradeoff curves between cache hit rate, memory capacity, and bandwidth.
Define benchmarking and characterization methodology: realistic agentic and long-context workloads (multi-turn reuse, session persistence, RAG prefixes), KV-cache reuse-rate measurement, and bandwidth-latency profiling across the full hierarchy (Nsight, PyTorch Profiler, vendor memory tools).
Full-Stack AI Memory Solutions — The Core Mandate
Define engineering requirements, with proof, for tiered memory systems for inference at fleet scale : HBM as L1, host DRAM (pinned, NUMA-aware pools) as L2, CXL-attached memory pools as an elastic tier, and NVMe/SSD as capacity tier — with the policies (admission, eviction, prefetch, placement) that make the hierarchy behave like one memory.
Design expert-weight offloading solutions for MoE serving: host-resident expert pools, GPU-resident expert caches with bandwidth-adaptive fill/evict policies, and CPU/CXL-execution hybrid paths — informed by the routing statistics of real models.
Translate model knowledge into product: write the requirements, reference architectures, and performance models that guide memory hardware and firmware roadmaps (HBM capacity/bandwidth, CXL device behavior, SSD QoS for cache tiers), and validate with end-to-end prototypes on real inference workloads.
Develop and Deliver POCs: demos and published benchmarks showing inference TCO improvement from the memory stack — e.g., context capacity multiplied at constant GPU count, or cost-per-token reduced through cache-hit-rate gains — credible to both CTOs and PhD researchers.
Technical Leadership
Set multi-year technical strategy for AI memory solutions; own build-vs-adopt-vs-contribute decisions across the open-source inference and caching ecosystem (vLLM, SGLang, LMCache, Cognos-style KV stores) and drive upstream contributions where strategic.
Lead architecture reviews and deep-dive design sessions; write the documents that become the company's standard for how we talk about memory for AI.
Represent the company with customers and partners at the deepest technical level: serve as the expert voice in CTO-to-CTO conversations, design wins, and standards discussions.
Mentor senior engineers and grow a bench of architecture talent across the model-to-memory boundary.
What You Bring
BS in Computer/Electrical/Electronic Engineering or Computer Science, and 10+ years of relevant experience MS in Computer/Electrical/Electronic Engineering or Computer Science with 8 years of relevant experience preferred.
12+ years in systems engineering, with 4+ years hands-on in large-scale LLM inference or GPU systems performance — you have personally profiled, diagnosed, and fixed memory bottlenecks in production serving, not just read about them.
First-principles understanding of transformer-class model internals : you can derive KV-cache size formulas from attention math, explain MQA/GQA/MLA tradeoffs, and reason about activation-memory peaks during prefill.
Working expertise with MoE model behavior : routing, expert parallelism, load skew, and the weight-memory economics of serving models larger than GPU capacity.
Direct experience with at least one major inference stack's memory-management internals (vLLM PagedAttention/block manager, SGLang HiCache/token pools, TensorRT-LLM KV manager, or llama.cpp compute buffers) — code-level, not configuration-level.
Strong performance-engineering skills: bandwidth-bound vs. com
Similar jobs
- Assistant Technical Director, Texas Performing ArtsUtaustin · UT MAIN CAMPUSFirst seen today
- Technical Director, CNNWarnerbros · DC Washington 820 1st Street NEFirst seen today
- Engineering Technical DirectorPowerdesigninc · FL St PetersburgFirst seen today
- Engineering Technical DirectorPowerdesigninc · FL St PetersburgFirst seen today
- PMTDP - PM Technical Director/ProducerNexstar · CA-San Diego; 4575 Viewridge Ave (Nexstar - KUSI & KSWB)First seen today
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job