Coupanginternal
Tech Infra Engineer
z-Test & Templates Only
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.5M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →Apply from your AI assistant
Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.
Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.
hirly's read of this role
- Seniority
- Mid level
- Country
- CN
- Work mode
- Remote-friendly
- First seen by hirly
- 14 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
Please complete the attached Internal Transfer Request Form and submit.
Please make sure to apply with your Coupang e-mail address .
As a Staff Systems Engineer in Developer Platform, you will partner with leaders of multiple platform teams. You will work closely with product to define and implement simple solutions to complex orchestration problems, building a highly scalable, reliable, and efficient platform for our customers. You will engineer and develop Kubernetes controllers, operators, and node-level daemons for the application runtime; drive performance tuning and scaling; and design multi-cluster control-plane capabilities that scale to millions of pods across thousands of clusters.
What You Will Do
Engineer and develop a unified application platform for hybrid (multi-cluster, multi-region, multi-cloud) application management using Kubernetes controllers and feedback-driven control systems to meet SLOs.
Deliver end-to-end automation for application lifecycle (deployments, rollouts, failovers, policy enforcement) to minimize manual work for users.
Drive fleet-wide optimization for cost, performance, and latency through data-informed controls and capacity management, improving $/RPS and tail latency.
Build resilient, multi-tenant control planes and workflows that safely scale to millions of pods across thousands of clusters.
Ensure reliability, security, and governance with clear guardrails, safe defaults, and automated remediation.
Partner with product and customers to turn complex orchestration problems into simple, reusable platform primitives and great developer experiences.
Champion observability and continuous improvement with measurable, outcome-focused metrics.
Basic Qualifications
Bachelor’s degree in Computer Science, Electrical Engineering, Math, or a closely related field (or equivalent experience)
10+ years in backend software development and operations
Recent experience designing and operating large-scale distributed systems (last 3 years) • Fluency in one or more among Go, C/C++, Python, or Java
Proven track record of delivering mission-critical systems
Experience with cloud computing using AWS or Azure or GCP
Preferred Qualifications
Kubernetes API machinery and semantics: SSA, SMP, server-side dry-run, watches/informers/listers, rate-limited workqueues, finalizers, owner references, leader election, API Priority and Fairness
Controllers/operators and node daemons in Go: client-go/controller-runtime, reconciliation patterns, backoff and retry, idempotency, partitioned/sharded controllers, HA and failover
CRDs and webhooks: versioning, conversion functions/webhooks, validating/mutating admission webhooks, policy frameworks and best practices
Pod/runtime semantics: sidecars, init/ephemeral containers, probes (readiness/liveness/startup), lifecycle hooks, termination behavior, PDBs, QoS classes, ResourceQuota/LimitRange, topology spread, affinity/anti-affinity
Scaling systems: HPA (resource/custom/external metrics), VPA, cluster autoscaler; multi-dimensional scaling, health-aware/autopilot-style policies; external metrics adapters and SLO-driven scaling
Federated and multi-cluster: placement/propagation, failover, drift detection, reconciliation strategies; consistent hashing and partitioning for scale
Distributed systems: CRDTs and eventual consistency paradigms; Raft/memberlist/gossip; deep familiarity with etcd, Kafka, Redis and their operational characteristics (compaction, backpressure, retention, failover)
Observability and data: Prometheus (cardinality control, recording rules), tracing; experience with vector databases for search and diagnostics; strong time-series forecasting (classical + ML) and statistical modeling for proactive optimization
Languages and interfaces: Go (primary), Java/Python as needed; gRPC/protobuf; JSON/YAML/Jsonnet
Leadership: ability to handle multiple competing priorities in a fast-paced environment and lead the delivery of large-scale services for complex business offerings
Recruitment Process
Application Review - Phone Interview - Onsite (or Virtual Onsite) Interview – Offer
The exact nature of the recruitment process may vary according to the specific job and may be changed due to scheduling or other circumstances.
Interview schedules and the results will be informed to the applicant via the e-mail address submitted at the application stage.
Details to Consider
This job posting may be closed prior to the stated end date for application if all openings are filled.
Coupang has the right to rescind an offer of employment if a candidate is found to have submitted false information as part of the application process.
Those eligible for employment protection (recipients of veteran’s benefits, the disabled, etc.) may receive preferential treatment for employment in accordance with applicable laws.
Privacy Notice
Your personal information will be collected and managed by Coupang as stated in the Application Privacy Notice located below: https://www.coupang.jobs/privacy-policy/
Please complete the attached Internal Transfer Request Form and submit.
Please make sure to apply with your Coupang e-mail address .
Similar jobs
- Infra EngineerClaimSorted · LondonFirst seen todayremote
- Infra EngineerScaleOps · United States - RemoteFirst seen todayremote
- Infra Engineer ScaleOps · Tel AvivFirst seen todayremote
- Infra Engineer ClaimSorted · LondonFirst seen today
- LLM Infra EngineerRicursive Intelligence · Palo AltoFirst seen 5d ago
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job