
As a Senior Principal Site Reliability Engineer at Bybit, you design and build an enterprise-grade chaos engineering platform that validates production resilience across a cryptocurrency exchange serving over 80 million users. You lead the architecture, development, and rollout of fault injection capabilities spanning multi-cluster Kubernetes and EC2 hybrid environments, multiple regions, and testnet/mainnet deployments. This role combines platform engineering, resilience validation, and team leadership to ensure Bybit's infrastructure can withstand real-world failure scenarios.
What you'll do
- Design and build an enterprise-grade chaos engineering platform supporting multi-cluster (Kubernetes and EC2 hybrid), multi-region, and multi-environment (testnet/mainnet) deployments
- Develop a Fault Injection Engine capable of pod-level, node-level, and AZ-level fault simulation, including network latency, packet loss, partition, dependency timeout, and error injection
- Engineer Production Safety Assurance features: blast radius control, one-click Kill Switch, automatic rollback, and real-time impact monitoring
- Implement Traffic Isolation to tag and segregate experiment traffic, ensuring fault injection does not impact real users
- Build Fault Isolation capabilities to scope precise impact at service, cluster, and AZ granularity
- Design experiment orchestration supporting complex fault scenario composition (simultaneous network latency plus downstream timeout plus cache invalidation)
- Integrate deeply with existing monitoring, alerting, and SLO systems to achieve automated closed-loop fault injection, impact observation, and pass/fail determination
- Define safety standards and approval workflows for mainnet fault injection
- Design and drive routine chaos experiments: daily patrol-level experiments, monthly/quarterly resilience verification, and cross-AZ/cross-region disaster recovery drills
- Establish a resilience scoring system to quantify system health based on experiment results
- Deliver improvement recommendations and drive business teams to remediate identified weaknesses
- Evaluate and select technology foundations (Chaos Mesh, Litmus, custom components, or hybrid strategy)
- Develop chaos engineering best practices and playbooks for SRE teams and application developers
- Mentor and grow a team of 2–3 engineers in chaos engineering capabilities
- Stay current with industry developments and introduce cutting-edge practices, such as AI-driven fault scenario discovery
What you bring
- 8+ years of backend and infrastructure engineering experience, including 3+ years dedicated to chaos engineering or stability engineering
- Hands-on experience running large-scale fault injection in production environments, with deep understanding of production safety constraints
- Expert-level proficiency in Kubernetes fault injection (Chaos Mesh, Litmus, or custom solutions), with familiarity in CRD and Operator development
- Proficiency in at least one backend language (Go preferred) and capability in platform-level system architecture design
- Deep understanding of distributed system failure modes: network partitions, split-brain, cascading failures, and data inconsistency
- Familiarity with observability tech stacks such as Prometheus, Grafana, Thanos, and OpenTelemetry
- Excellent technical documentation and solution design skills
- Ability to balance safety and validation depth, comfortable with production injection while maintaining strict risk control
- Strong cross-team collaboration and influence skills to drive buy-in from business teams
- Self-driven capability to independently plan and execute in ambiguous situations
Nice to have
- Experience in financial or trading system stability, including understanding of transaction consistency and fund safety constraints
- Experience building SLO and Error Budget frameworks
- Experience building automated fault recovery and self-healing systems
- Familiarity with AWS infrastructure: EC2, EKS, Multi-AZ, and Multi-Region
- Knowledge of Netflix Chaos Engineering, AWS Fault Injection Simulator, or Gremlin
- Open-source community contributions to Chaos Mesh, Litmus, or similar projects
What we offer
- Study Growth Fund to support your professional development and continuous learning
- Internal events including team-building activities, workshops, and events promoting collaboration and innovation
- Global collaboration within a diverse, international team
- Career advancement opportunities within a rapidly expanding global company
- Internal mobility and long-term development opportunities
About Bybit
Established in 2018, Bybit is one of the world's leading cryptocurrency exchanges and digital financial platforms, serving over 80 million users across more than 200 countries and regions. The platform delivers a seamless ecosystem spanning trading, payments, wealth management, custody, institutional services, and Web3, supported by 24/7 multilingual customer service and recognized as one of the most trusted and transparent platforms in the digital asset industry.
What engineering roles in crypto pay
383 salaries · our own dataMost engineering roles in crypto pay between $157k and $250k, with a median of $205k.