Trendmicro
Staff/Sr. ML Infrastructure / Platform Engineer
Taipei
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.3M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →hirly's read of this role
- Role family
- Engineering
- Seniority
- Lead / management
- Country
- TW
- Work mode
- On-site / unstated
- First seen by hirly
- 29 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
Join Trend ‧ Join New Generation
- 趨勢科技 - 全球雲端資安領航者 / 全亞洲最大軟體公司 / 企業版圖橫跨五大洲 / 趨勢全球研發基地在台灣
- ===============================================================
About the Role
We are building a production-grade, GPU-accelerated LLM serving platform that powers multiple AI products at enterprise scale. You will be responsible for designing, building, and operating the infrastructure that serves large language models — from raw Kubernetes cluster management to multi-GPU inference optimization and autoscaling.
Required Qualifications
Model Serving & Inference
Operate multi-model LLM serving infrastructure
Tune autoscaling policies to balance GPU cost and latency SLAs
Kubernetes & GPU Infrastructure
Operate production K8s clusters with NVIDIA GPU nodes
Handle GPU node lifecycle: NVIDIA driver setup
Infrastructure as Code
Write and maintain Terraform/Terragrunt modules for AWS/GCP cloud
Package platform components and model deployments as Helm charts
Manage multi-environment configurations
Observability & Performance
Maintain monitoring stack: Prometheus , Grafana ,
Build dashboards for GPU utilization, KV cache occupancy, TTFT/ITL latency, and cost per token
Set up alerting for SLA violations and OOM events
Bonus Skills
These are not required, but candidates with these skills will stand out.
LoRA / PEFT fine-tuning workflows
MLflow for experiment tracking, model registry, and automated adapter deployment
Experience building LoRA adapter CI/CD pipelines (training → registry → serving)
Experience with alternative inference frameworks such as SGLang or NVIDIA NIM , including deep Parameter Tuning for Continuous Batching , KV Cache management, and Speculative Decoding .
- ===============================================================
- 連結智慧 守護世界 --- Connected Intelligence for Securing a Connected World
Similar jobs
- [Agentic AI] Systems Software & Platform Engineering (OS & Firmware)Hp · Taipei, Taipei City, Taiwan RegionFirst seen 5d ago
- [Agentic AI] Systems Software & Platform Engineering (OS & Firmware)Hp · Taipei, Taipei City, Taiwan RegionFirst seen 5d ago
- Senior Platform EngineerRaccoon AI (JTCG) · Taipei OfficeFirst seen yesterday
- RDA & Metrology Data Platform EngineerMicron · Taoyuan - Fab 11, TaiwanFirst seen 2d ago
- Platform Engineer (Principal / Lead)Careers at Aura Cloud · Bengaluru, IndiaFirst seen today
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job