hirly

Amazon

Software Development Engineer II, AWS SageMaker AI

Bellevue, Washington, USA

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Amazon first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.4M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

Apply from your AI assistant

Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.

Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.

hirly's read of this role

Seniority
Mid level
Country
US
Work mode
On-site / unstated
First seen by hirly
27 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

At AWS SageMaker AI, we're making it easy to build state-of-the-art foundation models on the cloud. Model Factory is our platform for building, training, customizing, and evaluating foundation models at scale. Instead of hand-chaining data prep, distributed training, evaluation, and deployment across thousands of GPU and AWS Trainium devices, Model Factory lets teams express the whole lifecycle as a single, contract-validated workflow — orchestrated, reproducible, and fully managed. As LLMs and Generative AI scale, Model Factory is the platform that turns frontier training research into a reliable, repeatable pipeline for our internal teams and customers.

We're looking for a Software Development Engineer to help design, build, and operate the distributed systems at the core of this platform — the orchestration engine, compute integrations, SDK and contract layer, and the infrastructure that runs large-scale training and customization jobs. You'll own components end-to-end, from design through delivery and on-call operations, and work closely with the ML scientists and platform teams who depend on Model Factory every day. You'll turn requirements into robust, scalable, supportable services that fit cleanly into the overall architecture, uphold a high engineering bar, and grow your scope and technical leadership as you go.

A successful candidate has a strong software-engineering foundation, writes high-quality distributed-systems and services code, communicates clearly, and is motivated to deliver results in a fast-paced, ambiguous environment.

  • Key job responsibilities
  • As a Software Development Engineer on the SageMaker AI team, you will:
  • - Design, build, test, and operate services that orchestrate foundation-model data preparation, training, evaluation, and deployment as reliable, contract-validated workflows.
  • - Own delivery of individual components end-to-end — from design and implementation through deployment, monitoring, and on-call operations.
  • - Build and extend compute-backend integrations and job launchers — submitting, monitoring, and recovering large-scale training jobs across SageMaker (Training/Processing/HyperPod), EMR, AWS Batch, and Kubernetes/EKS.
  • - Improve the platform's resiliency and operability for long-running distributed jobs — checkpoint/resume, fault detection and recovery, retries, and observability (metrics, logging, experiment tracking).
  • - Contribute to the SDK, workflow orchestration, and schema/contract layer that teams use to declare and run jobs, and to the CDK infrastructure that deploys the platform.
  • - Integrate containerized training and evaluation frameworks (e.g., PyTorch/FSDP, verl, NeMo/Megatron) into the platform's task and recipe model.
  • - Contribute to design and architecture discussions, write clear technical designs, and uphold engineering best practices (code review, testing, operational readiness).
  • - Collaborate with ML scientists and internal customers to translate training requirements into reliable, self-service platform capabilities, and help onboard and mentor interns and new engineers as you grow.

Basic qualifications

  • - 3+ years of non-internship professional software development experience
  • - 2+ years of non-internship design or architecture (design patterns, reliability and scaling) of new and existing systems experience
  • - 1+ years of software development engineer or related occupational experience
  • - 1+ years of designing and developing large-scale, multi-tiered, multi-threaded, embedded or distributed software applications, tools, systems, and services using: C#, C++, Java, or Perl experience
  • - 1+ years of Object Oriented Design experience
  • - Bachelor's degree or foreign equivalent in Computer Science, Engineering, Mathematics, or a related field
  • - Experience programming with at least one software programming language

Preferred qualifications

  • - 3+ years of full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations experience
  • - Bachelor's degree in computer science or equivalent
  • - - Experience with workflow/pipeline orchestration (Airflow, Step Functions, or similar) and event-driven or service-oriented architectures.
  • - - Experience with container and cluster compute (Kubernetes/EKS, Ray, Slurm, AWS Batch) and cloud infrastructure-as-code (AWS CDK/CloudFormation).
  • - - Experience building and operating fully-managed cloud services at scale, including resiliency, checkpointing, and fault tolerance for long-running jobs.
  • - - Familiarity with machine-learning / deep-learning training workflows, GPU/accelerator compute (SageMaker HyperPod, AWS Trainium, P5-class GPUs), or distributed-training frameworks (PyTorch FSDP, Megatron-LM, DeepSpeed, verl).

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits .

USA, WA, Bellevue - 143,700.00 - 194,400.00 USD annually

Original posting on Amazon's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job
Software Development Engineer II – Amazon | hirly.me