Posted 2 months ago
Site Reliability Engineer | Hybrid - Centris/Makati
Makati CityHybrid
AI Summary
Site Reliability Engineer responsible for designing and improving monitoring, observability, and reliability standards for critical systems, reducing alert noise and MTTR while embedding postmortem learnings into patterns and automation.
About this role
Non-negotiable skills we are looking for:
- Excellent communication skills to drive continuous improvement by reducing alert noise, shorten MTTR, and improve change success by embedding postmortem learnings into patterns, rules, and pipeline:
- Cloud Observability: Azure Monitor/App Insights/Log Analytics (KQL)
- Knowledge of Grafana, Prometheus, App Dynamics, ThousandEyes
- Uses SLI/SLOs, postmortems, and CMDB and other context to reduce noise, drive self-healing, and measurably improve MTTR and KPIs.
High Level Summary:
Ensure the reliability of critical Chevron’s Monitoring and Event Management systems and services, provide standards and governance around monitoring and observability which enable Chevron business-critical processes to operate safely, efficiently, and reliably.
Key Responsibilities:
- You will design and define standards, patterns, and automations opportunities that elevate monitoring and reliability across platforms and applications, with a strong focus on Azure Monitor, ServiceNow ITOM Event Management, Grafana, and APM/Synthetics tooling
- You’ll partner with product teams to implement SLO/SLI-driven operations, reduce alert noise, accelerate incident response, and embed self-healing where it matters most.
- Engineer enterprise monitoring & event patterns by authoring and maintaining reference architectures, runbooks, and event management models (alert → event → incident) with actionable alerts and incidents routing.
- Contribute to Monitoring and Observability & Event Management Strategy and tooling intake/governance checkpoints and coach product teamsInternal - General Use
- Excellent communication skills to drive continuous improvement by reducing alert noise, shorten MTTR, and improve change success by embedding postmortem learnings into patterns, rules, and pipelines.
Technologies and Tools:
- SRE Practices: Observability and Monitoring
- Cloud Observability: Azure Monitor/App Insights/Log Analytics (KQL)
- Grafana/Prometheus for metrics visualization where applicable
- ServiceNow ITOM Event Management
- Azure Fundamentals, Azure Monitor
- DevOps and Automation Tools
- Grafana, Prometheus, App Dynamics, ThousandEyes
- Application Performance Monitoring and Digital User Experience tools
Required Qualifications:
- Bachelor’s degree in IT /Computer science/Engineering, or related field.
- 3+ years in monitoring/observability/SRE roles with hands-on experience in Azure Monitor/App Insights (KQL) and ServiceNow Event Management.
- Strong knowledge in Azure Log Analytics, KQL, Telemetry, APM implementations
- Demonstrated ability to collaborate across IT Operations team, platform, cyber, network, and product teams, strong written verbal communication for standards and enablement.
Preferred Qualifications:
- 5+ years of experience with SRE role and deep understanding of monitoring and application performance management
- Knowledge of SLO platforms (e.g., Nobl9) and experience contributing to standards/governance artifacts.
- Knowledge of proactive monitoring using Azure monitor services, telemetry, and synthetic transactions.
- Understanding of network architecture and security: WAN/LAN, TCP/IP, PKI.Internal - General Use
- Familiarity with ITSM processes and tools (e.g., ServiceNow), and compliance processes
- Have AIOps vision and awareness
Critical Selection Criteria:
- Communication & Teaming – Able to translate complex reliability patterns into consumable standards and coach IT operations team via office hours/CoP sessions.
- Technical Depth in Monitoring and Observability Stack – Hands-on in ServiceNow Event Management, Azure Monitor/KQL, and automation.
- Analytical & Systems Thinking – Uses SLI/SLOs, postmortems, and CMDB context to reduce noise, drive self-healing, and measurably improve MTTR and KPIs.
Additional Details:
- Work set-up: Hybrid 3x / RTO 2x per week | Eton, Centris
- Work shift: Nightshift
Skills
App DynamicsApp InsightsAzure MonitorGrafanaKQLLog AnalyticsPrometheusServiceNow ITOM Event ManagementSLI/SLOThousandEyes
Explore related jobs
More jobs at Tasq Staffing Solutions, Inc.
Similar App Dynamics jobs
- Microsoft Biz Apps Technical Architect (Power Platform and Dynamics CE)Hitachi Solutions · London, England
- Microsoft Biz Apps Technical Senior (Power Platform and Dynamics CE)Hitachi Solutions · London, England
- Microsoft Biz Apps Technical Lead (Power Platform and Dynamics CE)Hitachi Solutions · London, England
Jobs in Makati City
- ITechnical Support TechnicianInternationalkvh · Makati City, Philippines
- GUS Frontline Talent Acquisition Sourcing SpecialistGlobalpepsico · Makati City, Philippines
- EHS Specialist 1 (Environmental Sampling)SGS · Makati City, Metro Manila
- Specialist, Renewal SalesConcentrix Services Mexico, S.A. de C.V. · PHL Makati City - SLC
- Policy Admin AssistantCompany 19 - John Hancock Life Insurance Company (U.S.A.) · Makati City
- ISales Operations AssociateISSINDIA · Makati City, Philippines
Browse these categories
Market data for devops / sre / platform engineer roles
All reports →- SeriesRole reportsOne role family at a time: how many openings, what changed this week, who is hiring, what it pays.
- SeriesSalary reportsWhat employers publish in job postings, by level and workplace. Not self-reported pay.
- Market overviewState of tech hiring, September 2026: up 4.8%Tech hiring rose 4.8% month over month in September 2026, with 411,122 new listings. Customer support and account executive roles led the growth.