Posted 24 days ago
Senior Site Reliability Engineer (SRE) – Observability & Platform Systems
AI Summary
Senior Site Reliability Engineer owning the observability strategy and platform services for a digital banking environment, leading incident management and reliability engineering.
About this role
Location - Gurugram (On-site)
We are seeking a highly autonomous and technically profound Senior Site Reliability Engineer to own our observability strategy and critical platform services. In this role, you will not just monitor systems; you will engineer their reliability, performance, and security. You will serve as the primary authority for our monitoring stack and the operational lead for our core middleware layers.
This position demands a proactive problem-solver who thrives in complex, regulated environments. You must be comfortable working independently, making high-stakes technical decisions, and leading incident resolution efforts from detection to post-mortem.
Responsibilities
Observability Leadership: Architect, maintain, and optimize our end-to-end observability platform using Grafana, Prometheus and Elasticsearch. You will define the standards for logging, metrics, and tracing.
Platform Engineering: Operate and engineer critical surrounding systems including MongoDB, Kafka, Redis, HashiCorp Vault, and WSO2 API Manager. You are responsible for their stability, scaling, and security within our OpenShift environment.
Incident Command: Lead L3 support and Major Incident Management. You will diagnose complex cross-layer issues (application to infrastructure) and drive permanent resolutions.
Automation: Build self-healing systems and automated recovery workflows using Python and Bash to reduce manual toil and improve system resilience.
Proactive Reliability: Identify potential failure modes before they occur and implement architectural improvements to prevent outages.
Reliability Engineering: Define and track service reliability objectives (SLIs/SLOs) and help engineering teams improve platform reliability through measurable outcomes and error budget management.
Your Profile
7+ years of experience in SRE, DevOps, or Platform Engineering.
Proven expertise in Grafana and the Prometheus ecosystem.
Experience defining and operating SLOs, SLIs, Error Budgets, and availability targets for production services.
Deep operational experience with Statefulset applications in production Kubernetes/OpenShift environments.
Proven track record of owning the full Incident Management lifecycle, including Root Cause Analysis and implementation of preventive measures.
Strong proficiency in Python and Bash for automation and tooling development.
Exceptional analytical skills with the ability to work independently and solve ambiguous technical challenges without constant supervision.
Experience in the banking or highly regulated sectors is a strong plus.
What We Offer
The opportunity to work on a critical, large-scale digital banking platform.
A culture that values engineering excellence, ownership, and proactive innovation.
A collaborative international team environment with a focus on high-quality delivery.
We get curious people invested in the world
When you work at Saxo, you become a Saxonian and part of a purpose-driven organisation, where good ideas are always taken seriously, and where you can make a true impact. We are invested in your development, and you can expect a robust career from day one when you join Saxo – no matter which role you take on.
You will join 2,500 other ambitious colleagues across 11 countries and become part of an international organisation. Working in Saxo, you will get to meet colleagues from many different cultures and backgrounds, and you should know that we value diversity and inclusion and see it as a genuine source of strength to drive growth, foster innovation and position us for long-term success.
We encourage an open feedback culture and supportive team environments enabling employees to grow and fulfil their career aspirations.
When you bring passion, curiosity, drive and team spirit, your learning journey will be dynamic and your career opportunities in Saxo will be immense.
At Saxo we don’t just offer a job – we offer an opportunity to invest in your future!
How to apply :
Click here to create an account and upload your resume and a short motivation. We look forward to getting to know you better!
Skills
Explore related jobs
More jobs at Saxo Group India Private Limited (SGIPL)
Similar Bash jobs
Jobs in Gurugram
- WoW Software EngineerNatWest Markets N.V. · Gurugram
- Senior Associate, Client Solutions - Global Corporate (Flexible with shift timings)GLG · Gurugram
- Associate, Client Solutions - Global Corporate (Flexible with shift timings)GLG · Gurugram
- Associate, Project Support (Flexible with shift timings)GLG · Gurugram
- Ad Operations SpecialistTaboola · Gurugram, India
- Senior Software Developer - BackendIxigo · Gurugram, India
Browse these categories
Market data for devops / sre / platform engineer roles
All reports →- SeriesRole reportsOne role family at a time: how many openings, what changed this week, who is hiring, what it pays.
- SeriesSalary reportsWhat employers publish in job postings, by level and workplace. Not self-reported pay.
- Market overviewState of tech hiring, September 2026: up 4.8%Tech hiring rose 4.8% month over month in September 2026, with 411,122 new listings. Customer support and account executive roles led the growth.