A live role

Site Reliability Engineer, Infrastructure Platforms — UK (Intermediate to Senior Staff)

from GitLab · read the original and apply on their site

Remote

Our take: The real work here is keeping GitLab's user-facing systems reliable at scale by building automation, operating Kubernetes, and turning every incident into a lasting improvement.

If you love keeping big systems humming and would rather build the automation that kills toil than click through the same manual fix twice, this is your kind of work. You will operate GitLab.com at real scale on Kubernetes, write infrastructure as code, and turn every incident into a lasting improvement, all in a calm, async, remote-first team that trusts you to own your outcomes. Come for the reliability puzzles, stay because early detection and clean systems here are treated as craft, not chores.

Needs

  • tending infrastructure
  • automating toil
  • debugging code
  • infra as code
  • monitoring and slos
  • writing runbooks

Rewards

  • incident response
  • root cause work
  • scaling systems
  • systems design

Demands

  • being on call
  • troubleshooting under pressure
  • working async and remote
  • clear written comms

Grows

  • learning fast
  • setting direction
  • improving how teams work

Values

  • pragmatism over toil
  • clear ownership
  • working in the open
  • early detection rigour
  • long-term reliability

Drains

  • firefighting alerts
  • runbook grind

Would you love this work?

See how what you love doing lines up with this role - it takes about ten minutes to find out.

Find what I'd love to do next