hirly

This posting is no longer listed by Quberesearchandtechnologies.

hirly last saw it live on 1 September 2026. Similar roles are on the live board.

Quberesearchandtechnologies

Production Support Engineer — LLM Platform

Hong Kong

Apply through hirly

hirly scores this role against your resume, shows its reasoning, then writes a resume and cover letter for it and fills the application with you. Free to start — no card required.

hirly's read of this role

Role family
Customer support
Seniority
Mid level
Country
HK
Work mode
Remote-friendly
First seen by hirly
1 Sept 2026

Derived automatically from the posting. Sign up to see how the role scores against your own resume.

the posting

Qube Research & Technologies (QRT) is a global quantitative and systematic investment manager, operating in all liquid asset classes across the world. We are a technology and data driven group implementing a scientific approach to investing. Combining data, research, technology, and trading expertise has shaped our collaborative mindset, which enables us to solve the most complex challenges. QRT’s culture of innovation continuously drives our ambition to deliver high quality returns for our investors.

Your future role within QRT

Provide first- and second-line support for LLM gateway platform, investigating and resolving issues raised by engineering and business users across the firm

Monitor and maintain the platform's underlying infrastructure to ensure availability, stability and predictable performance under rapidly growing load

Support and troubleshoot model-serving backends and provider integrations, model providers, covering latency, throughput, error-rate and capacity issues

Triage incidents affecting model availability — provider instability, connection resets, timeouts, regional slowness — determine whether the cause is platform-side or upstream, and drive to resolution with vendors where required

Support the tooling layer built on top of LLM gateway: integrations, developer workspaces (e.g. Coder), coding assistants and API clients, including diagnosing issues introduced by upstream vendor releases running against a gateway-fronted API

Coordinate with platform engineering, cloud infrastructure and end-user teams to resolve incidents and minimise disruption

Support release management and change processes to keep production stable, including staged rollouts, non-prod validation and rollback

Build tooling and automation to improve monitoring, diagnostics and operational visibility, and to reduce repetitive manual work

Contribute to the design and implementation of monitoring, dashboards and alerting — for example extending Grafana dashboards covering TTFT, TPOT, percentile latency and failure-rate reporting

Own and improve operational documentation, runbooks and user-facing status communication

Your present skillset

Experience in a production support, SRE or platform operations role within a fast-paced environment, with strong ownership of issue resolution end to end

Strong Linux and Windows system administration skills

Proficiency scripting and automating in Python, Bash and/or PowerShell

Solid experience with relational databases such as PostgreSQL or SQL Server, including writing queries for investigation and supporting routine operational processes

Practical understanding of monitoring and observability: metrics, logs, traces, dashboards and alerting, and the ability to analyse system data to distinguish a platform-wide problem from a localised one

Comfortable debugging distributed, API-driven services: HTTP status and error semantics, timeouts, retries, connection resets, rate limiting, caching and latency percentiles

Familiarity with large language model concepts and hosting environments — inference APIs, model gateways/proxies, prompt and context handling, token accounting, streaming responses, prompt caching

Exposure to public cloud, ideally AWS (Bedrock, networking, IAM, logging/metrics), and to containerised or Kubernetes-based workloads

Ability to communicate clearly with both engineers and non-technical users, and to manage expectations of senior stakeholders during live incidents

Awareness of data-sensitivity and access-control considerations when routing workloads to third-party model providers

Beneficial

Experience supporting developer tooling and AI coding assistants (e.g. Claude Code, OpenCode) or IDE/workspace platforms

Experience with Grafana, Prometheus or equivalent observability stacks, including building dashboards and alert rules

Experience with CI/CD and infrastructure-as-code (Terraform, Ansible, or similar)

Experience operating multi-region services and troubleshooting region-specific performance issues (e.g. APAC latency)

Experience acting as the operational interface to third-party vendors and cloud providers during degradations

QRT is an equal opportunity employer. We welcome diversity as essential to our success. QRT empowers employees to work openly and respectfully to achieve collective success. In addition to professional achievement, we are offering initiatives and programs to enable employees achieve a healthy work-life balance.

Is this role actually a fit for you?

hirly answers with a score and its reasoning, then writes the resume and cover letter if you decide to go for it.

Score it against my resume
Production Support Engineer — LLM Platform at Quberesearchandtechnologies — hirly