Staff Site Reliability Engineer, Environment Automation
GitLab
Bengaluru, India · Posted 2 days ago · 18 Sept 2026
City
Bengaluru
Type
Full-time
Field
IT / Software
Pay
On apply page
GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation.
Overview
- Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab.
GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster.
The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software.
As a Staff Site Reliability Engineer (SRE) at GitLab, you’ll help keep all user-facing services and production systems reliable, scalable, and efficient. Our SREs combine a pragmatic operations mindset with strong software engineering practices to drive automation, reduce toil, and improve resilience across our platform.
In the Environment Automation specialization, your focus is on operating and automating hundreds of GitLab environments—from initial provisioning to day-to-day maintenance tasks.
Unlike other SRE roles, this position centers on automating the lifecycle of many tenant environments, ensuring they remain secure, consistent, and reliable at scale.
Some examples of the projects you could work on:
Designing infrastructure automation that provisions and operates GitLab environments using Terraform, Ansible, and Kubernetes
Creating and maintaining deployment packages for GitLab, such as Helm Charts and omnibus-gitlab
Building and operating Dedicated GitLab instances integrated with cloud-native services (e.g., GCP, AWS)
Developing tools to orchestrate infrastructure-as-code workflows across multiple tenants
Deploying and managing microservices on Kubernetes clusters at scale
Enhancing GitLab’s observability stack (e.g., Prometheus, ELK) to support proactive monitoring and incident response
Integrating with and operating infrastructure in cloud provider ecosystems (e.g., IAM, networking, storage)
Championing and implementing cloud security best practices across automated infrastructure
What You’ll Do
Build & Scale Multi-Tenant Infrastructure: Design and implement automation that provisions and manages hundreds of isolated GitLab environments using Terraform, Ansible, and Kubernetes. Manage complex state strategies and workspace configurations to support scale and maintainability.
Debug & Resolve Production Issues: Troubleshoot issues across Kubernetes clusters, cloud services, and GitLab apps—identifying root causes of failed deployments, crash loops, and scheduling conflicts to ensure service continuity.
Automate Operations at Scale: Replace manual workflows with infrastructure-as-code solutions, including automated version upgrades, configuration rollouts, and provisioning pipelines that operate reliably across all tenants.
Monitor & Predict Capacity: Build observability systems that detect bottlenecks, predict usage trends, and optimize resource consumption using tools like Prometheus, ELK, and Grafana.
Respond & Lead During Incidents: Lead incident response and postmortem efforts, applying technical depth to resolve issues and establish operational standards that reduce future risk.
Architect & Collaborate: Influence architectural decisions around automation, scalability, and operational excellence. Partner with engineering teams to improve automation, platform resilience, and production-readiness.
What You’ll Bring
Production-Scale Experience: Proven ability to operate and troubleshoot production workloads across multiple tenants or environments. Deep understanding of how distributed systems fail at scale and how to build in resilience.
Terraform & IaC Mastery: Strong hands-on experience with Terraform, including workspace strategies, state management, and automation patterns that scale. Comfortable solving state isolation issues and building reliable, reusable infrastructure code. Experience with Ansible and templating tools like Jsonnet is a plus.
Kubernetes in Production: Skilled at diagnosing deployment failures, interpreting pod logs, and debugging scheduling issues and rollback scenarios in live environments.
Application helper
Tailor my application
Add only what you want. Tadabbur creates a job-specific resume and email in this browser. Nothing is saved or submitted automatically.
Interested?
Sign in to send a short interest note to the poster (in addition to calling or WhatsApp).
Sign inSee something wrong? Sign in to report
