Senior Principal Site Reliability Engineer
AI Summary
Senior Principal Site Reliability Engineer designing and operating a large-scale chaos engineering platform, running fault injection experiments, and improving production resilience.
About this role
About Us
- Design and build an enterprise-grade chaos engineering platform supporting multi-cluster (K8s + EC2 hybrid), multi-region, and multi-environment (testnet/mainnet) deployments
- Core capability development:
- Fault Injection Engine: Pod-level / Node-level / AZ-level fault simulation, network latency / packet loss / partition, dependency timeout / error injection
- Production Safety Assurance: Blast radius control, one-click Kill Switch, automatic rollback, real-time impact monitoring
- Traffic Isolation: Experiment traffic tagging and isolation to ensure fault injection does not impact real users
- Fault Isolation: Precise impact scoping at service / cluster / AZ granularity
- Design experiment orchestration capabilities supporting complex fault scenario composition (e.g., simultaneous network latency + downstream timeout + cache invalidation)
- Deep integration with existing monitoring, alerting, and SLO systems to achieve an automated closed loop: inject fault → observe impact → determine pass/fail
- Define safety standards and approval workflows for mainnet fault injection
- Design and drive routine chaos experiments:
- Daily patrol-level experiments: Low-risk experiments executed automatically on a daily/weekly basis
- Periodic validation experiments: Monthly/quarterly resilience verification of critical paths
- Large-scale drills: Cross-AZ / cross-region disaster recovery failover validation
- Establish a resilience scoring system to quantify system health based on experiment results
- Deliver improvement recommendations and drive business teams to remediate identified weaknesses
- Evaluate and select the technology foundation (Chaos Mesh / Litmus / custom components — hybrid strategy)
- Develop chaos engineering best practices and playbooks to enable SRE teams and application developers
- Mentor and grow the team (2–3 engineers) in chaos engineering capabilities
- Stay current with industry developments and introduce cutting-edge practices (e.g., AI-driven fault scenario discovery)
- 8+ years of backend / infrastructure engineering experience, with 3+ years dedicated to chaos engineering or stability engineering
- Hands-on experience with large-scale fault injection in production environments (not just test environments), with deep understanding of production safety constraints
- Expert-level proficiency in Kubernetes fault injection (Chaos Mesh / Litmus / custom solutions), familiar with CRD / Operator development
- Proficient in at least one backend language (Go preferred), with platform-level system architecture design capability
- Deep understanding of distributed system failure modes (network partitions, split-brain, cascading failures, data inconsistency, etc.)
- Familiarity with observability tech stack (Prometheus / Grafana / Thanos / OpenTelemetry)
- Excellent technical documentation and solution design skills
- Experience in financial / trading system stability (understanding of transaction consistency and fund safety constraints)
- Experience building SLO / Error Budget frameworks
- Experience building automated fault recovery (self-healing) systems
- Familiarity with AWS infrastructure (EC2 / EKS / Multi-AZ / Multi-Region)
- Knowledge of Netflix Chaos Engineering / AWS Fault Injection Simulator / Gremlin
- Open-source community contributions (Chaos Mesh / Litmus or similar projects)
- Ability to balance "safety" and "validation depth" — not afraid of production injection, while maintaining strict risk control
- Strong cross-team collaboration and influence — chaos engineering requires buy-in from business teams; this role demands persuasion skills
- Self-driven, capable of independently planning and executing in ambiguous situations
Why Join Us
At Bybit, we are committed to fostering a supportive and enriching work environment.
Our benefits include:
- Study Growth Fund: We support your professional development and continuous learning.
- Internal Events: Participate in regular team-building activities, workshops, and events designed to promote collaboration and innovation.
- Global Collaboration: Be part of a diverse, international team, working alongside colleagues from around the world.
- Career Advancement: Access opportunities for growth and advancement within a rapidly expanding global company.
- Internal Mobility: Grow with us- Your long-term development is important to us. We offer internal job opportunities to help build your career path.
Skills
Explore related jobs
More jobs at bybit
Backend Development EngineerKuala Lumpur, Malaysia
Senior AML Specialist, Name Screening & Travel RuleKuala Lumpur, Malaysia
Senior AML Specialist, InvestigationsKuala Lumpur, Malaysia
Lead AML Product SpecialistAbu Dhabi, UAE
AML Expert, Process ExcellenceVilnius, Lithuania
Senior AML Specialist, Transaction MonitoringKuala Lumpur, Malaysia
Similar AWS jobs
Jobs in Kuala Lumpur
- SAP Consultant/ Senior Consultant/ Principal/ Manager/ - Finance (FI/CO)cbs APAC · Kuala Lumpur, Wilayah Persekutuan Kuala Lumpur
Lead, Strategy and Operations (Kuala Lumpur based)Agoda · Kuala Lumpur- Senior IT Support SpecialistMoneyHero Group · Kuala Lumpur, Malaysia
- Internship Program 2026 (Marketing)cbs APAC · Kuala Lumpur, Wilayah Persekutuan Kuala Lumpur
- Talent Acquisition CoordinatorLinesight · Kuala Lumpur, Malaysia
- Sr Associate Technical Consultant - Japanese/ Chinese speakingSAS · Kuala Lumpur, Malaysia
Browse these categories
Market data for devops / sre / platform engineer roles
All reports →- SeriesRole reportsOne role family at a time: how many openings, what changed this week, who is hiring, what it pays.
- SeriesSalary reportsWhat employers publish in job postings, by level and workplace. Not self-reported pay.
- Market overviewState of tech hiring, September 2026: up 4.8%Tech hiring rose 4.8% month over month in September 2026, with 411,122 new listings. Customer support and account executive roles led the growth.