Tower Research Capital is a leading quantitative trading firm founded in 1998. Tower has built its business on a high-performance platform and independent trading teams. We have a 25+ year track record of innovation and a reputation for discovering unique market opportunities.
Tower is home to some of the world’s best systematic trading and engineering talent. We empower portfolio managers to build their teams and strategies independently while providing the economies of scale that come from a large, global organization.
Engineers thrive at Tower while developing electronic trading infrastructure at a world class level. Our engineers solve challenging problems in the realms of low-latency programming, FPGA technology, hardware acceleration and machine learning. Our ongoing investment in top engineering talent and technology ensures our platform remains unmatched in terms of functionality, scalability and performance.
At Tower, every employee plays a role in our success. Our Business Support teams are essential to building and maintaining the platform that powers everything we do — combining market access, data, compute, and research infrastructure with risk management, compliance, and a full suite of business services. Our Business Support teams enable our trading and engineering teams to perform at their best.
At Tower, employees will find a stimulating, results-oriented environment where highly intelligent and motivated colleagues inspire each other to reach their greatest potential.
Summary
You will bridge the gap between quantitative research and high-performance computing, building and optimizing the systems used to train machine learning models at scale. You will focus on accelerating the end-to-end training lifecycle—from data ingestion and distributed execution to kernel performance and hardware utilization—enabling researchers to iterate more quickly across increasingly complex models and datasets.
Responsibilities
Training Performance and Benchmarking
Distributed Training Optimization
End-to-End Training Efficiency
GPU Kernel and Framework Development
Model and Numerical Optimization
Training Infrastructure
Cross-Functional Collaboration
Qualifications
3+ years of experience optimizing machine learning training workloads in high-performance, distributed, or large-scale computing environments.
Deep knowledge of machine learning frameworks such as PyTorch or JAX, including their execution models, compilation paths, autograd systems, and distributed-training capabilities.
Strong programming skills in Python and C++, with experience developing or optimizing performance-critical systems.
Proven experience with GPU kernel development and optimization using technologies such as CUDA, Triton, CUTLASS, cuBLAS, cuDNN, or related libraries.
Strong understanding of GPU architecture, including streaming multiprocessor execution, warp scheduling, tensor cores, and the memory hierarchy from registers through HBM.
Experience with distributed-training technologies and communication libraries such as NCCL, FSDP, DeepSpeed, Megatron-LM, XLA, or equivalent systems.
Proficiency with performance-analysis tools such as Nsight Systems, Nsight Compute, PyTorch Profiler, or comparable tracing and profiling platforms.
Understanding of high-performance networking, storage, and accelerator interconnects, including technologies such as InfiniBand, RDMA, NVLink, or NVSwitch.
Demonstrated ability to benchmark heterogeneous compute platforms and make rigorous, data-driven recommendations about performance, scalability, and cost.
Preferred Qualifications
Experience optimizing training workloads for transformer-based, time-series, reinforcement-learning, or other computationally intensive models.
Experience with cluster orchestration and scheduling technologies such as Kubernetes, Slurm, Ray, or similar platforms.
Familiarity with fault-tolerant distributed training, large-scale checkpointing, experiment reproducibility, and GPU-cluster observability.
Practical experience with specialized accelerators, custom hardware, or compiler technologies for machine learning.
Prior experience in financial trading is not required.
Anticipated New York annual base salary of $200,000, plus eligible for discretionary bonus.
Benefits
Tower’s headquarters are in the historic Equitable Building, right in the heart of NYC’s Financial District and our impact is global, with over a dozen offices around the world.
At Tower, we believe work should be both challenging and enjoyable. That is why we foster a culture where smart, driven people thrive – without the egos. Our open concept workplace, casual dress code, and well-stocked kitchens reflect the value we place on a friendly, collaborative environment where everyone is respected, and great ideas win.
Our benefits include:
Generous paid time off policies
Savings plans and other financial wellness tools available in each region
Hybrid working opportunities
Free breakfast, lunch and snacks daily
In-office wellness experiences and reimbursement for select wellness expenses (e.g., gym, personal training and more)
Volunteer opportunities and charitable giving
Social events, happy hours, treats and celebrations throughout the year
Workshops and continuous learning opportunities
At Tower, you’ll find a collaborative and welcoming culture, a diverse team and a workplace that values both performance and enjoyment. No unnecessary hierarchy. No ego. Just great people doing great work – together.
Tower Research Capital is an equal opportunity employer.
Summary
Build and optimize large-scale systems for training machine learning models, covering data ingestion, distributed execution, GPU kernels and hardware utilization. Requires 3+ years optimizing ML training workloads plus deep Python, C++, CUDA/Triton and PyTorch or JAX expertise.
Responsibilities
Benchmark model-training workloads across CPUs, GPUs and accelerators; Design and optimize distributed training with data, tensor, pipeline and model parallelism; Improve end-to-end training pipeline including data loading, memory, checkpointing; Develop and optimize GPU kernels and framework components; Apply mixed-precision, operator fusion and memory-efficient techniques; Partner with HPC teams on scheduling and infrastructure
Qualifications
Deep knowledge of PyTorch or JAX including execution, compilation, autograd and distributed training; Strong Python and C++ for performance-critical systems; GPU kernel development with CUDA, Triton, CUTLASS, cuBLAS, cuDNN; GPU architecture; Distributed training with NCCL, FSDP, DeepSpeed, Megatron-LM, XLA; Profiling with Nsight, PyTorch Profiler; Networking with InfiniBand, RDMA, NVLink, NVSwitch
Experience requirements
3+ years of experience optimizing machine learning training workloads in high-performance, distributed, or large-scale computing environments
Benefits
Generous paid time off policies; Savings plans and financial wellness tools; Free breakfast, lunch and snacks daily; In-office wellness experiences and reimbursement for wellness expenses; Volunteer opportunities