Technical Lead - GPU Infrastructure (100% Remote - Worldwide)
What engineering roles in crypto pay
205 salaries · our own dataMost engineering roles in crypto pay between $144k and $255k, with a median of $202k.
As Technical Lead for Cosmic AC at Tether Data, you own the architecture and delivery of a GPU compute and managed inference platform expanding from orchestrated workloads on managed clusters to a full stack on bare-metal GPU infrastructure. You lead an engineering team of about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India, reporting to the Senior Technical Product Manager while owning architecture, implementation, delivery plans, and serving as the primary technical interface to infrastructure partners. This is a hands-on infrastructure leadership role with a fixed delivery window in its first six months.
What you'll do
- Own the platform architecture end to end: architecture proposals, high-level and low-level designs, driven through review and kept current as the baseline
- Lead and line-manage the distributed engineering team across backend (Node.js), frontend (React), DevOps, QA and documentation: set engineering standards, conduct code and design review, manage release gates, run one-to-ones, and provide growth and performance input
- Design, build and operate a managed Slurm service for research users: controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline and upgrades, stalled-job and node-health detection, drain and autohealing, storage visibility, identity and isolation
- Own cluster bootstrap and lifecycle on partner-provided bare metal, NVIDIA GPU Operator and Network Operator, VM-based GPU isolation via KubeVirt and VFIO, and day-2 operations including upgrades, backup and recovery, and node replacement
- Define serving architecture for managed inference at scale: multi-GPU and multi-node parallelism, autoscaling, request routing and endpoint reliability, and confidential-compute-capable capacity for sensitive workloads
- Implement observability and operations across control plane, GPU fleet and application tiers: metrics, logging, alerting and SLOs, incident response and post-incident review, and an on-call model a small team can sustain
- Serve as primary technical interface to infrastructure partners and vendors: turn requirements into written specifications and acceptance tests, run escalations to closure, and provide technical input to capacity planning and hardware sourcing
- Work directly with research, model-training and product teams to translate their workloads into platform requirements and broker capacity when it is short
- Complete the platform team and set the technical bar for the engineers who join it
What you bring
- Proven hands-on infrastructure leadership: you own architecture and delivery while leading and line-managing an engineering team
- Deep production experience with Kubernetes and container orchestration at scale
- Experience designing and operating GPU infrastructure and scheduling systems (Slurm or equivalent)
- Strong proficiency with bare-metal infrastructure, Linux kernel and driver management
- Expertise in observability: metrics, logging, alerting and SLO-driven operations
- Excellent English communication skills and ability to work remotely in a distributed, global team
What we offer
- 100% remote work, worldwide
About Tether
Tether operates a suite of reserve-backed digital asset solutions including USDT, the world's most trusted stablecoin. Through divisions like Tether Power (sustainable Bitcoin mining with eco-friendly practices), Tether Data (AI and peer-to-peer infrastructure including KEET for secure data sharing), Tether Education (digital learning), and Tether Finance, the company serves hundreds of millions globally. Tether is a remote-first, distributed team operating as a lean leader in the fintech and blockchain space.
