FriendliAI
Forward Deployed Engineer - AI Inference
San Francisco
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Apply from your AI assistant
Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.
Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.
hirly's read of this role
- Seniority
- Mid level
- Country
- US
- Work mode
- On-site / unstated
- First seen by hirly
- 8 Oct 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
About the job
FriendliAI is seeking a Forward Deployed Engineer to assist enterprises in deploying, scaling, and operating generative and agentic AI workloads on FriendliAI infrastructure. You will work directly with customers to solve and implement production-grade applications using our products, such as Serverless Endpoints, Dedicated Endpoints, or Container.
Friendli Container is our service that allows customers to download our inference engine as Docker images and deploy it in their chosen environment, such as private clouds or on-premises. Our Friendli Container can be adopted directly to AWS EKS clusters using our EKS add-on product.
You will work directly on our customers’ projects, collaborating with their engineering teams to solve AI inference challenges like scaling, orchestration, and monitoring. This is a hands-on, customer-embedded role. If you have worked in DevOps, platform engineering, or SRE for AI applications, this is your ideal position.
Key Responsibilities
Design and implement large-scale deployment architectures for LLM and multimodal inference
Deploy and manage containerized workloads across Kubernetes clusters
Diagnose production issues, such as performance bottlenecks, and implement temporary fixes as needed
Collaborate with customers’ DevOps teams to integrate FriendliAI’s infrastructure into their CI/CD workflows
Develop scripts, Helm charts, and Terraform modules that simplify repeated deployments
Contribute field insights to shape our platform reliability, observability, and scaling strategies
Lead workshops, technical sessions, or webinars to help customers master infrastructure best practices.
Qualifications
3+ years of experience in cloud infrastructure, DevOps, or reliability engineering
Bachelor’s or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent
Proficiency with Kubernetes, Docker, Terraform, and Helm
Strong foundation in distributed systems, networking, and performance tuning
Experience with GPU-based computing and generative AI model serving workloads
Strong technical background in backend systems or AI tooling
Experience operating workloads on AWS, GCP, or OCI
Excellent problem-solving and debugging skills in real-world environments
Preferred Experience
Experience deploying large models (LLMs, diffusion models) on GPUs or clusters
Familiarity with inference frameworks (Triton, vLLM, TensorRT, DeepSpeed-Inference)
Familiarity with observability stacks (Prometheus, Grafana, Loki, ELK, OTEL)
Understanding of networking security and compliance frameworks (e.g., SOC 2)
Experience supporting on-prem or hybrid-cloud deployments
Benefits
A front-row seat to the generative AI infrastructure revolution
Competitive compensation and benefits package
Daily lunch and dinner provided; unlimited snacks and beverages
Health check-up and top-tier hardware support
Flexible working hours and a highly collaborative environment
About us
FriendliAI is the fastest inference cloud for agents, built to run frontier open-weight models in production at scale. It delivers up to 7x faster output token speed, up to 90% lower inference costs, and 99.99% uptime across the most demanding agent workloads — long-context inference, real-time streaming, and accurate tool calling.
We are a small, fast-moving team doing work that matters at one of the most exciting moments in the history of technology. With our world-class inference stack, we are building the platform teams can actually rely on.
Listed on hirly, a job board. hirly is not the employer: FriendliAI is hiring for this role.
Similar jobs
- Founding Forward Deployed EngineerClera · San FranciscoFirst seen 2d ago
- Forward Deployed Engineer - Platform Ditto · San Francisco, CAFirst seen 2d ago
- Forward Deployed Engineer (New Grad) Roadrunner · San FranciscoFirst seen 3d ago
- Forward Deployed Engineer - USDryft · San FranciscoFirst seen 3d ago
- Forward Deployed Engineer, Enterprise AI Automation (SF/NYC)Lavendo · San FranciscoFirst seen 3d ago
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job