A live role

Site Reliability Engineer - Dedicated Hosted Runners

from GitLab · read the original and apply on their site

Bangalore, India

Our take: The real work here is owning AWS runner fleets end to end, from Terraform and Go tooling to the SLO dashboards and on-call runbooks that keep customer CI/CD pipelines running.

If you love owning production infrastructure end to end, from Terraform modules and Go tooling to the dashboards and runbooks you actually operate by, this is your kind of work. You will build and run single-tenant runner fleets that big customers depend on, tuning autoscaling for cost and reliability and automating away the toil you find. It is remote, async, and calm about the important things, with real ownership over a platform that keeps other people's software shipping.

Needs

  • operating aws infra
  • terraform as code
  • writing go tooling
  • incident response
  • automating toil
  • building dashboards

Rewards

  • tuning for cost
  • fleet at scale
  • impact on fortune 100
  • clear reliability numbers
  • proper craftsmanship

Demands

  • carrying the pager
  • async across time zones
  • fully remote
  • keeping ci reliable

Grows

  • daily go work
  • slo monitoring
  • scale testing
  • zero-downtime deploys

Values

  • learning culture
  • operational rigour
  • cost efficiency
  • end-to-end ownership
  • customer focus

Drains

  • incident firefighting
  • qa validation
  • writing runbooks

Would you love this work?

See how what you love doing lines up with this role - it takes about ten minutes to find out.

Find what I'd love to do next