A live role

Staff Site Reliability Engineer, Environment Automation

from GitLab · read the original and apply on their site

Bangalore, India

Our take: The real work here is automating the entire lifecycle of hundreds of GitLab environments so manual operations disappear and everything stays reliable at scale.

If you love making manual operations disappear, this role hands you hundreds of GitLab environments to provision, upgrade, and heal entirely through code. You will live in Terraform, Kubernetes, and Prometheus, chasing down why deployments fail at scale and building the automation that stops them failing again. It is remote, high-stakes when things break, and deeply satisfying for someone who would rather write a pipeline than repeat a task.

Needs

  • infrastructure tending
  • automating toil
  • debugging kubernetes
  • monitoring and tuning
  • reading go and ruby
  • systems design

Rewards

  • leading incidents
  • visible platform impact
  • predicting capacity
  • finding root causes

Demands

  • on-call rotation
  • incident pressure
  • managing at scale
  • cross-team partnering

Grows

  • scaling systems
  • operational standards
  • resilience design
  • technical influence

Values

  • pragmatic operations
  • automation craft
  • remote working
  • customer reliability

Drains

  • firefighting outages
  • compliance rollouts
  • large-org coordination

Would you love this work?

See how what you love doing lines up with this role - it takes about ten minutes to find out.

Find what I'd love to do next