hirly

Amazon

Sr. Cloud Technical Account Manager, ES - NAMER - US-Frontier AI

San Francisco, California, USA

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Amazon first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.4M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

Apply from your AI assistant

Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.

Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.

hirly's read of this role

Role family
Sales
Seniority
Lead / management
Country
US
Work mode
On-site / unstated
First seen by hirly
27 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

AWS Applied AI Solutions (AAIS) is building toward a future where every business innovates with Amazon AI teammates. To get there, we build AI solutions that improve human capabilities and transform entire business functions. We create end-to-end products that surprise and delight out-of-the-box, making complex things easy and hard things possible, with no cloud experience required. We start with customers who embrace the future and build bridges to meet the rest where they are. We pursue ambitious opportunities with conviction, and we are looking for builders who share that mindset.

Are you ready to transform how businesses leverage artificial intelligence and machine learning at scale? Join the Amazon Web Services (AWS) Support team and become a strategic partner in delivering Amazon AI/ML solutions that empower our Frontier AI customers to innovate optimize and achieve unprecedented operational excellence.

Amazon Web Services (AWS) is seeking an experienced Sr. TAM with expertise in AI/ML HPC and/or other technologies to join our Frontier AI Technical Account Management (TAM) team.

You'll be at the forefront of solving complex AI/ML model training and inference implementation challenges guiding Frontier research customers through their most ambitious machine learning transformation journeys. By combining deep technical expertise with collaborative problem-solving you'll help organizations unlock the full potential of artificial intelligence and machine learning technologies - from distributed model training on GPU clusters to production-grade inference at scale.

The TAM role is not directly hands on keyboard within the customer's environment for troubleshooting customer support issues rather you will work with appropriate engineers and service teams to see issues through to resolution. You will help our customers design build operate and secure their cloud environments.

More importantly you will work proactively to help craft and execute strategies to drive our customers' adoption and use of AWS services including EC2 S3 DDB RDS and many more.

Your technical acumen and customer-facing skills will enable you to effectively represent AWS within a customer environment and drive discussions with senior leadership regarding incidents trade-offs support and risk management. You will provide advocacy and strategic technical guidance to help plan and build solutions using best practices and proactively keep your customers' AWS environments operationally healthy and resilient.

The close relationships developed with your customers will allow you to understand their business/operational needs and technical challenges to help them achieve the greatest value from AWS. This position will require the ability to travel 10% or more as needed.

The TAM is the centerpiece of value to our Enterprise Support customers. If you wish to be at the forefront of innovation come join us!

  • Key job responsibilities
  • Deliver Strategic Technical Engagements - Lead comprehensive technical deep-dives and performance optimization for enterprise AI/ML workloads including distributed training cluster architecture using AWS Parallel Computing Service (PCS) and AWS ParallelCluster the latest GPU-accelerated computing (i.e. P6/P6e G7/G7e instances) AWS Trainium-based training (Trn3 UltraServers) and multi-node NCCL communication tuning over EFA's SRD protocol.
  • Architect and Validate Innovative Solutions - Supporting customers who design and implement production-grade AI/ML training and inference solutions leveraging Slurm-based job scheduling distributed training frameworks (PyTorch FSDP DDP DeepSpeed Megatron-LM) SageMaker HyperPod for managed GPU clusters with automated health checks and node replacement high-performance parallel storage (Amazon FSx for Lustre) and container runtimes on Deep Learning AMIs (DLAMIs) against reference architectures and HPC lens to ensure performance reliability and cost governance at scale. Architect solutions using P6e UltraServers for multi-trillion parameter frontier models and Trn3 with the AWS Neuron SDK for cost-optimized training and inference.
  • Enable Customer Success - Support customers in implementing business-critical HPC capabilities including the development of large language model (LLM) (Llama GPT-class models) physics-informed neural networks (PINNs) and surrogate models MLOps pipelines simulation-ML hybrid architectures orchestrated by AWS Step Functions and AWS Batch distributed data processing cluster observability and governance controls for GPU/Trainium-intensive workloads.
  • Enable Business Critical Outcomes - Partner with service teams to enhance model training throughput optimize NCCL collective communications improve GPU/Trainium utilization across multi-node UltraClusters and drive operational efficiency through proactive monitoring automated failure recovery (HyperPod health checks) and capacity planning (EC2 Capacity Blocks for ML). Contribute to product roadmap PFR share reference architecture performance and benchmarks with broader Technical communities Serve as Trusted Advisor and Advocate - Develop and nurture technical partnerships with enterprise stakeholders serving as the trusted advisor for AI/ML infrastructure decisions spanning compute networking (Elastic Fabric Adapter with SRD) storage orchestration and the HPC-to-AI convergence journey.
  • A day in the life
  • Your day will be dynamic and impactful involving deep technical consultations on distributed training architectures strategic solution design for GPU and Training cluster deployments and collaborative problem-solving across multi-node ML environments. You'll engage with technical leaders architect innovative AI/ML implementations - from Slurm-managed training clusters and SageMaker HyperPod to PyTorch FSDP/DeepSpeed training jobs and Neuron SDK compilation workflows - and provide expert guidance that bridges machine learning infrastructure with business objectives.

You will partner with Solution Architects Frontline Engineers and Service teams to provide customers with AWS AI/ML best practice guidance diving deep into machine learning infrastructure services (PCS ParallelCluster HyperPod Batch) promoting customers' AI/ML workloads to production developing regional AI/ML strategies advising on HPC-to-AI convergence patterns (simulation-surrogate loops physics-informed neural networks) and training field teams on distributed training patterns GPU/Trainium cluster operations and the use cases and benefits of artificial intelligence and machine learning at scale. You establish trust with your customers to understand their business/operational needs and technical challenges and help them achieve the greatest value from AWS. This position will require the ability to travel 10% or more as needed.

  • About the team
  • Amazon values diverse experiences. Even if you do not meet all of the preferred qualifications and skills listed in the job description, we encourage candidates to apply. If your career is just starting, hasn't followed a traditional path, or includes alternative experiences, don't let it stop you from applying.

Amazon Web Services (AWS) is the world's most comprehensive and broadly adopted cloud platform. We pioneered cloud computing and never stopped innovating — that's why customers from the most successful startups to Global 500 companies trust our robust suite of products and services to power their businesses.

We value work-life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why flexible work hours and arrangements are part of our culture. When we feel supported in the workplace and at home, there's nothing we can't achieve in the cloud.

Here at AWS, it's in our nature to learn and be curious. Our employee-led affinity groups foster a culture of inclusion that empower us to be proud of our differences. Ongoing events and learning experiences, including our Conversations on Race and Ethni

Original posting on Amazon's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job