Jobless Developer
Jobgether logo

Posted 4 days ago

Open

Staff DevOps Engineer, Developer Experience (DevEx)

United StatesRemoteFull-time

AI Summary

Staff DevOps Engineer, Developer Experience (DevEx) at a partner company. Designs and leads the internal developer platform strategy, architecture, and tooling across AWS and GCP, focusing on Kubernetes, CI/CD, GitOps, observability, security, and cost optimization to enable reliable, scalable software delivery.

About this role

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff DevOps Engineer, Developer Experience (DevEx) based in the United States.

This is a senior technical leadership role focused on building the platforms, tooling, and workflows that enable engineering teams to deliver software reliably at scale.
You will define multi-quarter platform strategy across development, testing, deployment, and operations for cloud-based services and data platforms.
The role combines hands-on systems engineering with architecture, platform product thinking, and organization-wide technical influence.
You will build self-service developer experiences that reduce infrastructure complexity, cognitive load, and operational friction.
A major focus will be Kubernetes, cloud infrastructure, CI/CD, GitOps, observability, security, and cost optimization across AWS and GCP environments.
You will also help establish effective and secure practices for AI-assisted and agentic engineering workflows across the broader organization.
This is an opportunity to shape the engineering “paved road,” mentor senior engineers, and make reliable, secure delivery the default across a growing technology organization.

Accountabilities:

  • Set and drive the technical roadmap for the Internal Developer Platform, establishing long-term platform strategy rather than simply executing an existing plan.
  • Author and review RFCs, architecture documents, and technical proposals adopted across multiple engineering teams, serving as an architecture reviewer for infrastructure-impacting changes.
  • Act as a senior escalation point for high-severity platform incidents, leading retrospectives and translating findings into structural improvements.
  • Mentor senior and staff-level engineers while raising engineering standards around platform architecture, reliability, security, and software delivery.
  • Own and evolve the Internal Developer Platform as a product, including Backstage service/API/group catalogs, golden-path scaffolder templates, and CI/CD components.
  • Partner with software, data, security, and product engineering leaders to identify systemic development friction and address it through automation, standardization, and platform abstraction.
  • Define standards for AI-assisted and agentic development workflows, evaluating and rolling out appropriate coding-agent tools with effective security and governance controls.
  • Own the architecture of CI/CD orchestration, including GitLab CI/CD and reusable components, supporting rapid releases, strong security controls, and local-to-production consistency.
  • Standardize build, testing, and deployment practices across application and data workloads and lead the evolution of GitOps-based deployment strategies.
  • Own the architecture and lifecycle of multi-cluster Kubernetes platforms across AWS and GCP environments, including upgrades, node lifecycle management, networking, and security.
  • Establish cloud IAM, networking, security, and Infrastructure-as-Code standards across multiple accounts and environments using technologies such as Terraform and Crossplane.
  • Lead service mesh architecture and policies across multi-cluster environments, including Istio configuration, sidecar strategies, and traffic controls.
  • Own critical platform systems such as GitOps tooling, secrets management, and service mesh as production-grade services.
  • Define observability strategy across platforms, balancing actionable telemetry and signal quality with infrastructure and data costs.
  • Drive FinOps practices for platform infrastructure, including cost visibility, unit economics, and waste reduction.
  • Establish and evolve organization-wide SLIs, SLOs, alerting strategies, and incident-response practices while promoting a blameless postmortem culture.
  • Requirements:

    • 10+ years of overall engineering experience, including meaningful hands-on software development experience in Go, Python, Java, or similar languages.
    • 7+ years of experience designing, building, and operating cloud infrastructure at significant scale.
    • Demonstrated ability to provide technical leadership across multiple teams through architecture decisions, RFCs, platform strategy, and cross-functional influence.
    • Deep hands-on experience designing and operating multi-account, multi-cluster Kubernetes platforms in production, particularly EKS and/or GKE environments supporting dozens or more services or clusters.
    • Strong proficiency with Infrastructure as Code using tools such as Terraform and Crossplane, combined with organizational-scale GitOps practices using Flux, Argo CD, or comparable technologies.
    • Extensive experience designing and operating CI/CD platforms that support high release velocity while maintaining strong security and governance controls, with GitLab CI/CD experience preferred.
    • Hands-on experience with service mesh technologies such as Istio in multi-cluster environments, including security and traffic policies.
    • Strong experience building and operating observability platforms using technologies such as Prometheus, Datadog, Grafana, Loki, Mimir, or Thanos, with a strong focus on cost efficiency.
    • Working knowledge of container and software supply-chain security, including image scanning, hardened base images, vulnerability remediation, and related practices.
    • Experience owning or substantially contributing to cloud cost governance and FinOps initiatives.
    • Familiarity with AI-assisted and agentic development tools such as Cursor, Claude Code, Codex, or equivalent, along with the ability to establish practical approaches for safe organizational adoption.
    • Strong understanding of SLIs, SLOs, alerting, and incident response, including experience leading blameless postmortems rather than simply participating in them.
    • Excellent written and verbal communication skills, with the ability to influence through technical documentation, architecture reviews, presentations, and collaborative decision-making.
    • Ability to work effectively across engineering, data, security, and product organizations and translate complex infrastructure challenges into scalable solutions.
    • Experience supporting data pipeline and batch-processing platforms such as Airflow, EMR, or Dataproc is a plus.
    • Benefits:

      • Fully remote position based in the United States.
      • Opportunity to shape developer experience and platform strategy across an entire engineering organization.
      • High-impact technical leadership role with significant influence over cloud infrastructure, developer tooling, security, reliability, and engineering standards.
      • Opportunity to work with a modern technology stack spanning AWS, GCP, Kubernetes, Terraform, Crossplane, GitOps, service mesh, CI/CD, observability, and AI-assisted development tooling.
      • Strong opportunities for mentorship, technical leadership, and cross-functional collaboration.
      • Inclusive working environment focused on diversity, belonging, professional growth, and equal opportunity.
      • Encouragement to apply even when candidates do not meet every listed qualification, provided they have the experience and capabilities to succeed in the role.
      • Opportunity to contribute to systems and standards designed to help engineering teams move quickly while maintaining reliability, security, and operational excellence.

Skills

AirflowArchitecture ReviewArgo CDAWSBackStageBlameless PostmortemsCloud IAMCost OptimizationCrossplaneDataDogDataprocEMRFinOpsFluxGCPGitLab CI/CDGitOpsGOGrafanaIncident ResponseInfrastructure As CodeIstioJavaKubernetesLokiMimirNetworkingObservabilityPrometheusPythonRFCsSecuritySLIsSLOsTerraformThanos

Explore related jobs

Browse these categories

Market data for devops / sre / platform engineer roles

All reports →