hirly

Towerresearchcapital

Machine Learning Performance Engineer (Inference)

New York

Apply through hirly

hirly scores this role against your resume, shows its reasoning, then writes a resume and cover letter for it and fills the application with you. Free to start — no card required.

hirly's read of this role

Seniority
Mid level
Stated salary
$200,000 – $300,000 per year
Country
US
Work mode
On-site / unstated
First seen by hirly
2 Sept 2026

Derived automatically from the posting. Sign up to see how the role scores against your own resume.

the posting

Tower Research Capital is a leading quantitative trading firm founded in 1998. Tower has built its business on a high-performance platform and independent trading teams. We have a 25+ year track record of innovation and a reputation for discovering unique market opportunities.

Tower is home to some of the world’s best systematic trading and engineering talent. We empower portfolio managers to build their teams and strategies independently while providing the economies of scale that come from a large, global organization.

Engineers thrive at Tower while developing electronic trading infrastructure at a world class level. Our engineers solve challenging problems in the realms of low-latency programming, FPGA technology, hardware acceleration and machine learning. Our ongoing investment in top engineering talent and technology ensures our platform remains unmatched in terms of functionality, scalability and performance.

At Tower, every employee plays a role in our success. Our Business Support teams are essential to building and maintaining the platform that powers everything we do — combining market access, data, compute, and research infrastructure with risk management, compliance, and a full suite of business services. Our Business Support teams enable our trading and engineering teams to perform at their best.

At Tower, employees will find a stimulating, results-oriented environment where highly intelligent and motivated colleagues inspire each other to reach their greatest potential.

Summary:

As part of Tower Research's Core Engineering team, you will bridge the gap between quantitative research and high-performance production systems, architecting inference pipelines that operate at the physical limits of hardware. Your objective will be to drive the speed, efficiency, and reliability of our ML inference pipelines to their absolute limits, ensuring our predictive models consistently achieve microsecond-level latency.

Responsibilities:

Benchmarking & Strategy:

Lead the technical evaluation of diverse inference platforms - ranging across CPUs, GPUs, and FPGAs - to guide Tower's infrastructure deployment decisions.

System Architecture Optimization:

Analyze and enhance execution across deep memory hierarchies to maximize resource utilization and parallel processing. You will assess and resolve memory subsystem and interconnect bottlenecks across the end-to-end inference lifecycle.

Infrastructure & Deployment Feasibility:

Collaborate with Infrastructure teams to understand thermal, power, and operational constraints of hardware platforms to design inference strategies for our latency-critical trading strategies that fit within those envelopes.

GPU Kernel Development:

Develop highly optimized kernels and integrate specialized performance libraries to extract maximum computational throughput from the underlying silicon.

Model Optimization & Deployment:

Implement advanced model reduction techniques (quantization, pruning, distillation) to ensure compact memory footprints and numerical stability. Prioritize optimization for low-latency, event-level inference workloads to meet real-time trading requirements.

Cross-Functional Collaboration:

Collaborate closely with ML Researchers, HPC Engineers, FPGA Engineers, and Datacenter Engineers to bring to fruition target deployments.

Qualifications:

2+ years of experience optimizing deep learning inference in latency-sensitive or high-throughput production environments, in any domain.

ML Frameworks: Deep expertise in lower-level ML framework development (PyTorch/JAX), paired with strong Python/C++ skills and a thorough understanding of mixed-precision computation.

Kernel Development & Optimization Tooling: Proven experience in custom GPU kernel development. Deep familiarity with advanced optimization libraries and compilers (e.g., Triton, TensorRT, ONNX, IREE, HLS4ML, cuBLAS, CUTLASS) as well as profiling tools (e.g., Nsight Systems, Nsight Compute).

GPU Architecture Mastery: Deep expertise in GPU microarchitecture, encompassing SM execution, warp scheduling, and full memory hierarchy optimization (registers to HBM).

Cross-Architecture Benchmarking: Proven record of rigorous, data-driven approach to evaluating inference performance across heterogeneous compute architectures.

Bonus: Practical experience targeting and optimizing inference workloads on specialized hardware ecosystems, including FPGAs and ASICs.

Prior experience in financial trading is not required.

Anticipated annual base salary range $200,000-$300,000, plus eligible for discretionary bonus.

Tower’s headquarters are in the historic Equitable Building, right in the heart of NYC’s Financial District and our impact is global, with over a dozen offices around the world.

At Tower, we believe work should be both challenging and enjoyable. That is why we foster a culture where smart, driven people thrive – without the egos. Our open concept workplace, casual dress code, and well-stocked kitchens reflect the value we place on a friendly, collaborative environment where everyone is respected, and great ideas win.

Our benefits include:

Generous paid time off policies

Savings plans and other financial wellness tools available in each region

Hybrid working opportunities

Free breakfast, lunch, and snacks daily

In-office wellness experiences and reimbursement for select wellness expenses (e.g., gym, personal training and more)

Company-sponsored sports teams and fitness events (JPM Corporate Challenge, Cycle for Survival, Wall Street Rides FAR and more)

Volunteer opportunities and charitable giving

Social events, happy hours, treats, and celebrations throughout the year

Workshops and continuous learning opportunities

At Tower, you’ll find a collaborative and welcoming culture, a diverse team and a workplace that values both performance and enjoyment. No unnecessary hierarchy. No ego. Just great people doing great work – together.

Tower Research Capital is an equal opportunity employer.

Is this role actually a fit for you?

hirly answers with a score and its reasoning, then writes the resume and cover letter if you decide to go for it.

Score it against my resume
Machine Learning Performance Engineer (Inference) at Towerresearchcapital — hirly