Member of Technical Staff | Observability & Reliability
AI Summary
Joins the Platform team as an expert on observability and reliability, evolving telemetry stacks, monitoring dataplanes and agents, defining SLOs, leading incident response, and reducing MTTR across cloud and on-premise environments.
About this role
About the role
At Avra, every technical IC is a Member of Technical Staff (MTS). The title doesn't put anyone in a silo: you own systems and outcomes, not steps in a function, and you keep building depth in your area. Seniority shows up in your scope, level, and compensation, not in titles.
In this role, you'll join the Platform team as our go-to expert on observability and reliability. Our customers make real-time decisions based on our responses, so when we're down, their operations stop. Avra's cloud is just one more dataplane, alongside the dataplanes we operate inside customer environments — so observability and reliability have to work the same way everywhere.
What you'll do
Evolve our observability stack for logs, metrics, traces, and alerting.
Make sure every dataplane, in our cloud and on-premise, reports its active release, health, heartbeat, logs, metrics, and usage to the control plane.
Bring telemetry into customer clusters within a model where agents only make outbound connections.
Detect drift between the desired state and what's actually running in each environment.
Monitor the health of our deployment and runtime agents.
Provide visibility into ephemeral workloads, such as the Ray clusters that run our batch inference.
Define SLOs, lead incident response and postmortems, and reduce MTTR — including when a fix requires coordinating with the customer.
Reduce telemetry cost: less redundant data, more useful signal.
How we measure success
99.9% serving availability, with incidents trending down.
MTTR, including on-premise incidents.
Near-zero drift between desired and actual state.
All agents active and reporting, across every dataplane.
What we're looking for
Deep experience with OpenTelemetry and observability backends.
Hands-on practice with SLOs, error budgets, actionable alerting, and incident management.
Strong experience with Kubernetes and infrastructure as code (Terraform / Helm ).
Experience operating software in environments you don't fully control.
Production-quality code and reviews, and a willingness to operate what you build.
Nice to have
Shipping software to customer-hosted Kubernetes (e.g., Helm, outbound-only connectivity).
GCP or GKE, AWS or EKS.
ML multi-node/multi-cluster workloads in production.
Financial services or regulated environments.
Skills
Explore related jobs
More jobs at Avra
Similar Alerting jobs
Jobs in São Paulo
Brazil - Operations Representative (Boarding Agent/Maritime Ship Agent)ISS Group Holdings Limited · Santos, São Paulo- Analista Júnior em Infraestrutura | SREXP Inc. · São Paulo, SP
Analista de Asset Servicing- Credit & Illiquid AssetsBTG Pactual · São Paulo, Brazil- Analista Pleno em Infraestrutura | SREXP Inc. · São Paulo, SP
Desenvolvedor .NET C# - Electronic TradingBTG Pactual · São Paulo, Brazil
Consultor de Vendas - Tucuruvi/SPAgibank · São Paulo, São Paulo
Browse these categories
Market data for this role
All reports →- SeriesRole reportsOne role family at a time: how many openings, what changed this week, who is hiring, what it pays.
- SeriesSalary reportsWhat employers publish in job postings, by level and workplace. Not self-reported pay.
- Market overviewState of tech hiring, September 2026: up 4.8%Tech hiring rose 4.8% month over month in September 2026, with 411,122 new listings. Customer support and account executive roles led the growth.
