A live role
Site Reliability Engineer, Intermediate to Senior Staff — Infrastructure Platforms
from GitLab · read the original and apply on their site
Remote, Canada; Remote, United Kingdom; Remote, US
Our take: The real work here is keeping GitLab.com reliable at scale by writing automation and infrastructure as code, operating Kubernetes in production, and turning incidents into lasting improvements as a self-directed manager-of-one.
This is for the engineer who finds deep satisfaction in making systems that stay up under real load, and who would rather write the automation that kills a repetitive task than do it a second time. You will operate GitLab.com on Kubernetes, ship infrastructure as code, and turn every incident into a lasting improvement, all as a self-directed manager-of-one in an all-remote, async team. If you love following symptoms to their root cause and closing the loop with metrics, you will feel at home here.
Needs
- tending infrastructure
- automating toil
- troubleshooting production
- writing infra code
- monitoring and slos
- incident response
- documenting findings
- root cause work
Rewards
- automation over toil
- owned outcomes
- measurable results
- real autonomy
- a learning culture
Demands
- carrying the pager
- calm under fire
- fully remote async
- self-directed solo work
- working across teams
Grows
- learning fast
- explaining clearly
- mentoring engineers
- reliability strategy
- better patterns
Values
- automating toil
- ownership
- measurable results
- autonomy
- learning culture
Drains
- alert triage
- runbook upkeep
- escalation paperwork
Would you love this work?
See how what you love doing lines up with this role - it takes about ten minutes to find out.
Find what I'd love to do next