hirly

Ollama

Software Engineer, Cloud

Palo Alto

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Ollama first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.4M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

Apply from your AI assistant

Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.

Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.

hirly's read of this role

Role family
Engineering
Seniority
Mid level
Country
US
Work mode
On-site / unstated
First seen by hirly
28 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

Ollama is the most popular way for developers to access open models. What started as an open-source, local-first runtime is now the largest developer network in the open-model ecosystem: 8.9 million monthly active developers and over 67,000+ community-built integrations. We're backed by Y Combinator, Benchmark, 8VC, and Theory Ventures.

Our team is small and talent dense. We're flat, low-ego, and fast-moving. We like people who are truth-seeking, passionate, design-driven, and who enjoy shipping code.

About the role

You'll build Ollama’s cloud, a scalable inference platform that lets developers run large, capable open models in their workflow. You'll work on high-throughput, low-latency distributed systems — inference serving, GPU fleet management, routing, metering, and the platform that Pro, Max, Team, and Enterprise customers rely on to process trillions of tokens.

What you'll do

Build and scale the inference platform that serves every request from ollama.com .

Design the routing and capacity layer that places workloads across GPUs and regions for cost, latency, and availability.

Own multi-tenant infrastructure: isolation, quotas, usage metering, billing, and Pro/Max/team/enterprise tiering.

Build the reliability, observability, and cost controls for our team and customers

You may be a fit if

You have deep experience with high-throughput, low-latency distributed systems — inference serving, traffic routing, real-time data pipelines, or large-scale APIs.

You're comfortable with cost/performance tradeoffs at scale and have owned a production service end-to-end.

You've worked with Kubernetes, GPU scheduling, or inference infrastructure.

You think in terms of reliability, SLOs, and honest capacity planning.

Bonus: experience building an inference platform, GPU fleet management, or billing/metering for an AI service.

Original posting on Ollama's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job
Software Engineer, Cloud – Ollama · Palo Alto | hirly.me