Spellbrush
HPC/ML Infrastructure Engineer
San Francisco or Tokyo
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.6M live jobs from 190,000+ employers in 200+ countries.
Tailor my resume for this job →Apply from your AI assistant
Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.
Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.
hirly's read of this role
- Seniority
- Mid level
- Countries
- US, JP
- Work mode
- On-site / unstated
- First seen by hirly
- 2 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
We’re looking for an experienced HPC infrastructure engineer to lead bringup, administration, and operations on is probably the largest anime AI training cluster in the world . You’ll serve as the bridge between our researchers and the bare GPU machines, helping to make sure that SLURM jobs are running, parallel filesystems are serving, network is transmitting, and that the anime models are training.
You may be a good fit if:
You love anime and the anime aesthetic.
This probably one of the only jobs in the world where you will get to combine your love of anime and large-scale GPU systems.
You’re familiar with the modern HPC software landscape
Once upon a time, our team could install SLURM on a few bare metal nodes and get away with it. Now the landscape has become unbelievable complex, with SLURM deploys through Slinky on K8s, provisioning through warewulf/MAAS/ansible, filesystems through WEKA/VAST/Ceph, VPN and access through tailscale, and monitoring via the Grafana/Prometheus stack. We’re looking for someone with relevant experience up and down the stack (and maybe a papercut or two to show for it!)
As well as the traditional sysadmin landscape
Bringing up and managing cluster still requires good old linux sysadmin skills, including wrangling ldap, triaging dmesg, and setting sticky bits on directories for misbehaving users and tools.
You're not afraid of physical computers
We’re building out edge datacenters and our CEO is still personally racking, stacking, and provisioning HGX-based nodes in our living room. Also his VLAN design sucks and he’s bad at fiber routing. Please send help.
And you're comfortable working on small, fast-paced teams.
We currently have a very tiny research team, and you’ll be directly helping some of the AI researchers in the world train the best anime image model in the world.
We also believe in the unmatched speed of in-person teams, and prefer on-site collaboration in either our primary research office in Tokyo (downtown Akihabara), or San Francisco (dogpatch!). Bay area is strongly preferred as we have physical hardware in the Bay Area. Visa sponsorships are available.
Listed on hirly, a job board. hirly is not the employer: Spellbrush is hiring for this role.
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job