A live role

Site Reliability Engineer, Intermediate to Senior Staff — Infrastructure Platforms

from GitLab · read the original and apply on their site

Remote, Canada; Remote, United Kingdom; Remote, US

Our take: The real work here is keeping GitLab.com reliable at scale by writing automation and infrastructure as code, operating Kubernetes in production, and turning incidents into lasting improvements as a self-directed manager-of-one.

This is for the engineer who finds deep satisfaction in making systems that stay up under real load, and who would rather write the automation that kills a repetitive task than do it a second time. You will operate GitLab.com on Kubernetes, ship infrastructure as code, and turn every incident into a lasting improvement, all as a self-directed manager-of-one in an all-remote, async team. If you love following symptoms to their root cause and closing the loop with metrics, you will feel at home here.

Needs

  • tending infrastructure
  • automating toil
  • troubleshooting production
  • writing infra code
  • monitoring and slos
  • incident response
  • documenting findings
  • root cause work

Rewards

  • automation over toil
  • owned outcomes
  • measurable results
  • real autonomy
  • a learning culture

Demands

  • carrying the pager
  • calm under fire
  • fully remote async
  • self-directed solo work
  • working across teams

Grows

  • learning fast
  • explaining clearly
  • mentoring engineers
  • reliability strategy
  • better patterns

Values

  • automating toil
  • ownership
  • measurable results
  • autonomy
  • learning culture

Drains

  • alert triage
  • runbook upkeep
  • escalation paperwork

Would you love this work?

See how what you love doing lines up with this role - it takes about ten minutes to find out.

Find what I'd love to do next