Remote
Site Reliability Engineer
About this role
To enhance platform reliability and performance for millions of readers, the remote Site Reliability Engineer will manage production services on Heroku and AWS, optimize application performance, and collaborate with engineering teams to implement automated solutions. Key responsibilities Share responsibility for the health and performance of production services, proactively identifying and resolving issues Define and track Service Level Objectives (SLOs) and error budgets, using data to guide reliability decisions Optimize application and database performance by identifying bottlenecks and implementing improvements Required qualifications 3-5+ years of experience in Site Reliability, DevOps, or related roles managing infrastructure Deep experience with applications on Heroku and AWS Proven expertise in Infrastructure as Code (IaC) using Terraform Hands-on experience with CI/CD pipelines, preferably using GitHub Actions Strong understanding of web application performance, particularly in a Ruby on Rails environment
Source listing: virtualvocations_main