← All jobs
DR

AI Inference Platform Engineer

DRWChicago
Type
Full-time
Work setup
On-site
Experience
Mid
Posted
Today
$200k - $250k
midpoint above market median

What engineering roles in crypto pay

231 salaries · our own data
this role$145kmedian$264k

This role pays $200k-$250k, above the $202k median for engineering roles in crypto on this board.

As an AI Inference Platform Engineer at DRW, you build, operate, and optimize the systems that serve large language, vision, multimodal, and embedding models across the firm. You own the inference serving platform end-to-end, from early model evaluation through reliable production deployment, working across inference runtimes, distributed systems, and production platform engineering with deep GPU literacy.

What you'll do

  • Optimize LLM inference performance across modern NVIDIA GPU architectures and inference runtimes
  • Build end-to-end performance profiling and observability to identify bottlenecks from individual GPU kernels through multi-node inference systems
  • Design and optimize KV cache and distributed inference architectures, including caching, routing, memory tiering, and prefill/decode strategies
  • Own day-0 model onboarding, determining the appropriate runtime, precision, sharding, memory, batching, cache policy, and serving configuration for new models
  • Maintain validated performance profiles for important model and hardware combinations, including performance and quality regression testing
  • Measure and monitor quality equivalence across serving configurations, including KV cache quantization, speculative decoding acceptance thresholds, precision choices, and model routing
  • Manage the production serving lifecycle of models, including versioning, compatibility, staging, canarying, promotion, rollback, and retirement
  • Partner with SRE and platform teams to automate model deployment, distribution, production readiness, observability, and reliable operation across environments
  • Optimize model placement, scaling, and resource allocation across the inference fleet to improve utilization and cost efficiency while meeting performance and reliability requirements
  • Design and operate multi-tenant scheduling and isolation across shared GPU capacity, balancing latency SLOs, throughput, and priority across concurrent workloads

What you bring

  • Hands-on experience serving LLMs on NVIDIA GPUs, with familiarity across current and emerging architectures (Hopper, Blackwell, and successors), HBM, Tensor Cores, NVLink/NVSwitch, and the compute and memory bottlenecks that shape serving decisions
  • Deep expertise in at least one modern inference runtime such as TensorRT-LLM, vLLM, or SGLang
  • Practical knowledge of inference optimization techniques including continuous batching, scheduling, chunked prefill, speculative decoding, quantization, CUDA Graphs, and paged attention
  • Understanding of KV cache architecture, including prefix caching, block management, sizing, eviction, quantization, cache-aware routing, and multi-tier caching
  • Experience measuring model quality equivalence across serving configurations, including evaluation harnesses, task-specific benchmarks, and regression detection for quantization, KV cache, and speculative decoding changes
  • Experience designing and tuning distributed inference systems, including tensor parallelism, multi-node deployments, and disaggregated prefill and decode
  • Experience with multi-tenant GPU scheduling, workload isolation, and QoS across concurrent inference workloads
  • Proficiency with GPU performance and observability tooling such as Nsight, DCGM, OpenTelemetry, Prometheus, and Grafana
  • Strong Linux and systems performance fundamentals, with the ability to diagnose bottlenecks across hardware, drivers, runtimes, networking, and application layers
  • Production experience with model serving infrastructure, including CI/CD, automated testing, observability, and production readiness
  • A measurement-driven approach to performance optimization
  • The ability to take ownership of performance problems across hardware, runtime, model, and infrastructure boundaries
  • The ability to move quickly and reprioritize as trading needs change, while maintaining a high bar for production systems
  • Understanding of the importance of reliability, predictability, and performance when AI systems are integrated into trading workflows and decision-making processes
  • The ability to evaluate unfamiliar models, runtimes, and hardware quickly and make sound engineering decisions with limited prior guidance
  • Strong communication skills to explain complex performance tradeoffs across engineering teams

What we offer

  • Annual base salary of $200,000 to $250,000 depending on experience, qualifications, and relevant skill set
  • Annual discretionary bonus eligibility
  • Comprehensive employee benefits including group medical, pharmacy, dental and vision insurance
  • 401k with discretionary employer match
  • Short and long-term disability insurance
  • Life and AD&D insurance
  • Health savings accounts and flexible spending accounts

About DRW

DRW is a diversified trading firm with over 30 years of experience, operating globally with offices throughout the U.S., Canada, Europe, and Asia. The firm trades a variety of asset classes including Fixed Income, ETFs, Equities, FX, Commodities and Energy across all major global markets, and has expanded into real estate, venture capital, and cryptoassets.

AI Inference Platform Engineer | CryptoJobsHQ