
As a Senior Principal Site Reliability Engineer at Bybit, you will design and build an enterprise-grade chaos engineering platform for one of the world's leading cryptocurrency exchanges. You'll architect resilience validation frameworks, lead a small team of engineers, and drive production safety practices across multi-cluster, multi-region, and multi-environment deployments serving over 80 million users.
What you'll do
- Design and build an enterprise-grade chaos engineering platform supporting multi-cluster (Kubernetes and EC2 hybrid), multi-region, and multi-environment (testnet and mainnet) deployments
- Develop core fault injection capabilities including pod-level, node-level, and AZ-level fault simulation; network latency, packet loss, and partition injection; and dependency timeout and error injection
- Engineer production safety assurance features: blast radius control, one-click Kill Switch, automatic rollback, and real-time impact monitoring
- Build traffic isolation and fault isolation systems to ensure experiments do not impact real users while achieving precise impact scoping at service, cluster, and AZ granularity
- Design experiment orchestration capabilities supporting complex fault scenario composition (for example, simultaneous network latency plus downstream timeout plus cache invalidation)
- Integrate deeply with existing monitoring, alerting, and SLO systems to automate the closed loop: inject fault, observe impact, determine pass or fail
- Define safety standards and approval workflows for mainnet fault injection
- Design and drive routine chaos experiments: daily patrol-level experiments, monthly or quarterly resilience verification of critical paths, and large-scale cross-AZ or cross-region disaster recovery failover validation
- Establish a resilience scoring system to quantify system health based on experiment results
- Deliver improvement recommendations and drive business teams to remediate identified weaknesses
- Evaluate and select the technology foundation (Chaos Mesh, Litmus, or custom components using a hybrid strategy)
- Develop chaos engineering best practices and playbooks to enable SRE teams and application developers
- Mentor and grow a team of 2 to 3 engineers in chaos engineering capabilities
- Stay current with industry developments and introduce cutting-edge practices such as AI-driven fault scenario discovery
What you bring
- 8+ years of backend or infrastructure engineering experience, with 3+ years dedicated to chaos engineering or stability engineering
- Hands-on experience with large-scale fault injection in production environments (not just test environments), with deep understanding of production safety constraints
- Expert-level proficiency in Kubernetes fault injection using Chaos Mesh, Litmus, or custom solutions, with familiarity in CRD and Operator development
- Proficiency in at least one backend language (Go preferred) and platform-level system architecture design capability
- Deep understanding of distributed system failure modes including network partitions, split-brain, cascading failures, and data inconsistency
- Familiarity with observability tech stacks such as Prometheus, Grafana, Thanos, and OpenTelemetry
- Excellent technical documentation and solution design skills
- Ability to balance safety and validation depth, comfortable with production injection while maintaining strict risk control
- Strong cross-team collaboration and influence skills; chaos engineering requires buy-in from business teams and demands persuasion capability
- Self-driven and capable of independently planning and executing in ambiguous situations
Nice to have
- Experience in financial or trading system stability, including understanding of transaction consistency and fund safety constraints
- Experience building SLO or Error Budget frameworks
- Experience building automated fault recovery (self-healing) systems
- Familiarity with AWS infrastructure (EC2, EKS, Multi-AZ, Multi-Region)
- Knowledge of Netflix Chaos Engineering, AWS Fault Injection Simulator, or Gremlin
- Open-source community contributions to Chaos Mesh, Litmus, or similar projects
What we offer
- Study Growth Fund to support your professional development and continuous learning
- Internal events including team-building activities, workshops, and innovation sessions
- Global collaboration with a diverse, international team
- Career advancement opportunities within a rapidly expanding global company
- Internal mobility and opportunities to grow your long-term development path
About Bybit
Established in 2018, Bybit is one of the world's leading cryptocurrency exchanges and digital financial platforms, serving over 80 million users across more than 200 countries and regions. The company delivers a seamless ecosystem across trading, payments, wealth management, custody, institutional services, and Web3, backed by world-class technology and recognized as one of the most trusted and transparent platforms in the digital asset industry.
What engineering roles in crypto pay
383 salaries · our own dataMost engineering roles in crypto pay between $157k and $250k, with a median of $205k.