hirly

UltaHost

AI Systems Architect

Bosnia and Herzegovina, Bosnia and Herzegovina

Apply through hirly

Upload your resume and get a version tailored to this job, plus a cover letter, in about thirty seconds — before you create an account.

Apply with hirly

hirly's read of this role

Seniority
Mid level
Work mode
On-site / unstated
First seen by hirly
23 Sept 2026

Derived automatically from the posting. Sign up to see how the role scores against your own resume.

the posting

About UltaHost

UltaHost is a fast-growing global web hosting and cloud infrastructure company delivering high-performance, reliable, and scalable technology solutions to customers worldwide. Since our founding in 2018, we have grown rapidly across global markets, supporting entrepreneurs, developers, agencies, startups, SaaS companies, and businesses with modern hosting and cloud infrastructure.

Our services span VPS hosting, dedicated servers, shared hosting, domains, game servers, and cloud-based infrastructure across a growing international footprint.

As UltaHost continues to scale, Artificial Intelligence is becoming an increasingly important part of our infrastructure, products, customer experience, and internal operations.

We are investing in: GPU-powered infrastructure, private and self-hosted LLM environments, commercial and open-source AI models, AI-powered applications, intelligent infrastructure automation, RAG and agentic systems, customer-facing AI services, and reusable AI platform capabilities.

We are now looking for an AI Systems Architect to help design and build the technical foundation of this next stage.

Job Overview

We are looking for an AI Systems Architect to help the company adopt AI across departments, reduce manual work, improve sales performance, and build internal AI-powered tools, scripts, platforms, and agents. This is not a traditional AI engineering role. We are looking for a visionary technology leader who can identify opportunities where Artificial Intelligence can fundamentally improve how the company operates, sells, supports customers, develops products, and scales globally.

As AI Systems Architect, you will define and execute the organization’s AI roadmap, lead AI innovation initiatives, build internal AI platforms and intelligent automation systems, and establish our position as an AI-first hosting, cloud, and SaaS company. This person should be both technical and product-minded: able to understand business problems, design practical AI solutions, build prototypes, work with tools such as Lovable and similar AI development platforms, and help promote the company as an AI-first technology brand.

The role is ideal for someone who has strong technical knowledge, good taste in design and user experience, and a clear vision for how AI can improve operations, customer support, sales, marketing, development, and management workflows. The AI Systems Architect will design, build, integrate, deploy, and improve AI systems across UltaHost’s infrastructure and application ecosystem.

This is a highly hands-on technical role operating at the intersection of: GPU infrastructure, Linux systems, Proxmox and KVM virtualization, containers, LLM deployment and inference, AI applications, RAG and vector search, agentic systems, cloud and hosting platforms, backend engineering, infrastructure automation, and production reliability.

You will work directly with our CTO and collaborate closely with Product, Engineering, Infrastructure, Operations, Support, and other business departments. This position does not currently include direct team-management responsibilities. It is an individual-contributor role with significant technical ownership and architectural influence.

We are looking for someone who understands the full production stack:

GPU infrastructure → virtualization and containers → model serving → AI platform services → APIs and integrations → AI applications → monitoring, security, and optimization.

Main Goal of the Role

Your main goal will be to help UltaHost build a practical, scalable, and production-ready AI ecosystem.

You will:

Design and build infrastructure for running AI and LLM workloads on GPU-enabled environments.

Deploy, benchmark, optimize, and operate self-hosted and third-party LLMs.

Build production AI applications, agents, copilots, APIs, RAG systems, and automation.

Connect AI systems with UltaHost infrastructure, products, customer portals, support systems, databases, and internal tools.

Explore how technologies such as GPUs, Proxmox, containers, open-source LLMs, vector databases, and modern AI frameworks can become part of UltaHost's technology stack.

Identify internal processes where AI and deterministic automation can reduce repetitive work and improve efficiency.

Build reusable AI infrastructure and services that can support multiple future applications instead of isolated one-off experiments.

Help UltaHost evolve toward an AI-enabled hosting and cloud platform.

The objective is not experimentation for its own sake.

We want working systems that create measurable technical or business value.

Key Responsibilities

GPU & AI Infrastructure

Design, deploy, configure, and operate infrastructure for LLM and AI workloads using GPU-enabled servers.

Work with GPU environments and understand practical considerations related to:GPU compute, VRAM capacity, model size, precision, quantization, concurrency, batching, utilization, throughput, latency, thermal and power constraints,and workload allocation.

Evaluate the hardware and infrastructure requirements of different AI models and use cases.

Design virtualization strategies for AI workloads using Proxmox VE, KVM, virtual machines, Linux containers, and Docker.

Configure or support GPU and PCIe passthrough in virtualized environments where appropriate.

Design secure resource-isolation models for internal and potentially customer-facing AI workloads.

Contribute to GPU infrastructure capacity planning, availability, monitoring, backup, and disaster-recovery strategies.

Develop repeatable deployment processes rather than relying on manual server configuration.

Work with Infrastructure and Engineering teams to create stable environments for development, testing, staging, and production AI workloads.

Proxmox, Virtualization & Platform Engineering

Design and maintain Proxmox-based infrastructure for AI and general-purpose workloads.

Work with Proxmox clusters, virtual machines, LXC containers, networking, storage, resource allocation, templates, and backups.

Design infrastructure with appropriate isolation, high availability, performance, and operational simplicity.

Automate VM, container, and infrastructure provisioning using APIs, infrastructure-as-code, scripts, or orchestration tools.

Help standardize deployment patterns across GPU servers and AI application environments.

Troubleshoot performance, networking, storage, virtualization, and hardware-resource issues.

Evaluate when workloads should run on bare metal, virtual machines, containers, or orchestration platforms.

LLM Deployment & Inference

Deploy and operate open-source LLMs on private/self-hosted infrastructure.

Evaluate and work with inference technologies such as vLLM, Ollama, LiteLLM, Hugging Face, TGI, or comparable platforms.

Benchmark models based on quality, latency, throughput, VRAM consumption, concurrency, and operating cost.

Understand and apply inference optimization techniques such as quantization, batching, caching, context management, and model routing where appropriate.

Compare self-hosted models with commercial APIs such as OpenAI, Anthropic Claude, and Google Gemini and select the appropriate architecture for each use case.

Design model gateways and reusable inference APIs that can serve multiple UltaHost applications.

Build resilient model integrations with appropriate fallbacks, retries, rate limits, timeouts, and error handling.

AI Applications, RAG & Agentic Systems

Design and build production-ready AI applications rather than demonstration-only chat interfaces.

Develop internal AI assistants connected to authorized company knowledge, documentation, support content, operational runbooks, product information, and other approved data sources.

Build RAG systems using embeddings, vector databases, metadata filtering, reranking, access controls, and retrieval evaluation.

Original posting on UltaHost's site ↗

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job
AI Systems Architect at UltaHost · Apply with a tailored resume