This posting is no longer listed by Orcristtechnologies.
hirly last saw it live on 1 September 2026. Similar roles are on the live board.
Orcristtechnologies
Infrastructure/Systems Engineer
Hybrid / Berlin
Apply through hirly
hirly scores this role against your resume, shows its reasoning, then writes a resume and cover letter for it and fills the application with you. Free to start — no card required.
hirly's read of this role
- Seniority
- Mid level
- Country
- DE
- Work mode
- Remote-friendly
- First seen by hirly
- 1 Sept 2026
Derived automatically from the posting. Sign up to see how the role scores against your own resume.
the posting
Infrastructure Engineer – Platform
Company
Orcrist is building a next generation data intelligence platform using cutting-edge technologies. We’re handling petabyte-scale data with sub-second queries. Our product is a Kubernetes-based platform delivered as B2B SaaS or as a self-hosted on-prem solution, including air-gapped deployments. We enable customers across defense, law enforcement, and enterprise to turn mission-critical data into actionable intelligence. Our Platform team owns the infrastructure that powers every deployment, from the metal up.
Role
You'll own the layer everything else runs on: bare-metal servers, operating systems, data-center networking, and storage across on-prem and fully air-gapped sites — the physical and infrastructure foundation our platform and GPU fleets are built on. You design, build, and operate server fleets with a strong automation and DevOps mindset, then partner with our SRE, MLOps, and ML teams to ensure everything running above the metal — including GPU inference — performs reliably at scale. Some of this work is hands-on at customer sites, where you size, rack, and commission self-contained server environments with no internet uplink.
We weight depth in modern data-center infrastructure, networking, and automation more heavily than GPU-specific experience. A strong infrastructure and network engineer with a genuine automation mindset — even without prior GPU exposure — is a better fit for this role than a candidate with GPU expertise whose networking background is rooted in legacy, corporate-style L2 designs.
What you'll do
Design, size, provision, and operate bare-metal server fleets across on-prem and air-gapped environments (firmware/BIOS/UEFI, BMC via Redfish/IPMI, OS, RAID, kernel and storage tuning) using zero-touch provisioning (PXE/iPXE, MAAS/Metal3/Tinkerbell/Ironic) and automation (Ansible, Terraform, or equivalent — the tooling matters less than the automation mindset).
Build and run modern data-center networking: L2/L3 design, IP Fabric, BGP and switching, and RDMA fabrics (RoCE/InfiniBand) sized to scale without ripping out the core.
Engineer resilient, highly available storage (Ceph/Rook, NVMe) with capacity planning and encryption at rest.
Operate confidently in air-gapped and on-prem environments: offline mirrors and registries, signed artifacts, firmware/driver lifecycle without internet access, and system hardening.
Support MLOps and inference-serving fundamentals — GPU model serving (Triton/KServe/vLLM), GPU scheduling and sharing, and throughput/latency optimization — in partnership with our SRE and ML teams.
Plan and run on-site build-outs: rack integration, power budgets, thermal/cooling and UPS sizing, commissioning, capacity planning, runbooks, and operator handover, with SWaP awareness for field sites.
About You
5+ years in bare-metal, data-center, or systems infrastructure engineering, with hands-on ownership of physical and compute infrastructure at scale.
Strong bare-metal Linux (Ubuntu, RKE2, Talos, or similar): provisioning, firmware/BMC management, PXE/iPXE, kernel and storage tuning, systemd, RAID.
Real experience with infrastructure automation (Ansible, Terraform, or equivalent), Git and CI/CD, and scripting in Python, Bash, or Go.
Solid, current data-center networking fundamentals: L2/L3, IP Fabric, BGP, and switching — this is a hard requirement, not a nice-to-have. RDMA (RoCE/InfiniBand) experience is a strong plus.
Comfortable operating in air-gapped or on-prem environments and traveling to customer sites for builds and deployments.
Practical hardware sizing literacy: power budgets, thermal/cooling, UPS sizing, and rack integration.
Documentation-focused, methodical, and calm during hardware incidents. Eligible to work in Germany.
Nice‑to‑haves
German language (B1+); exposure to regulated or security-critical environments (e.g. BSI C5, ISO 27001, or defense-sector delivery) is a plus.
NVIDIA GPU stack knowledge (drivers, CUDA, GPU Operator, MIG, DCGM) and cross-node GPU interconnect experience (NVLink, InfiniBand, NCCL).
Kubernetes bare-metal fundamentals — how cluster bring-up and GPU device plugins interact with the underlying hardware, not day-to-day cluster operation.
Inference optimization (vLLM, TensorRT-LLM, quantization) and familiarity with switch NOS (SONiC/Cumulus).
Relevant certifications (NVIDIA, Red Hat, CKA/CKS) or field/forward-deployed engineering experience.
What We Offer
Modern architecture & stack.
Remote/Hybrid setup in Berlin with occasional team events in Berlin.
Home office budget and great equipment.
30 days vacation.
Direct impact on critical missions across private and public-sector customers.
Is this role actually a fit for you?
hirly answers with a score and its reasoning, then writes the resume and cover letter if you decide to go for it.
Score it against my resume