Description:
We’re looking for a Site Reliability Engineer to join our Production Sandbox Reliability team – the first line of defense ensuring enterprise customer sandboxes remain stable, up-to-date, and well-monitored. You'll work in a globally distributed team and act as the operational guardian for these environments.
This role is ideal for someone who thrives in high-ownership environments, enjoys hands-on infrastructure work, and wants to grow their SRE skills while partnering closely with engineering and customer-facing teams.
Key Responsibilities
- Maintain reliability: Monitor the health of enterprise customer sandbox environments and ensure high availability, uptime, and stability across all services
- Stay up-to-date: Regularly roll out updates to microservices inside each sandbox to ensure alignment with the latest versions
- Alert response & escalation: Triage infrastructure and application alerts, perform initial investigation and escalate incidents to the appropriate engineering teams with clear context
- Improve observability: Enhance metrics, logs and tracing coverage using Datadog and the Grafana stack (Mimir, Loki, Tempo), identifying gaps and driving better alerting practices
- Support incident workflows: Collaborate in post-incident reviews and ensure root cause analysis is followed up with actionable items and improvements by relevant teams
- Communicate proactively: Act as the bridge between internal engineering teams and customer-facing teams, providing timely updates during incidents, maintenance and version upgrades
- Participate in on-call rotation: Provide continuous coverage across APAC, EMEA and LATAM time zones as part of a rotating on-call schedule (follow the sun)
What We're Looking For
- 4+ years of experience in SRE, DevOps or Infrastructure Engineering roles
- Experience with Node.js or Go
- Familiarity with AWS cloud services (EKS, S3, RDS)
- Hand-on experience with Kubernetes, including Helm and ArgoCD
- Experience with observability stacks: Datadog, Grafana, Mimir, Loki, Tempo, Zabbix
- Strong verbal and written communication skills – able to interface effectively with both technical and non-technical stakeholders
- Self-starter mindset with an eye for operational excellence and continuous improvement