Remote
Site Reliability Engineer
About this role
To enhance platform reliability and performance for millions of readers, the remote Site Reliability Engineer will apply a software engineering mindset to solve infrastructure challenges, manage production services on Heroku and AWS, and optimize application performance. Key responsibilities Share responsibility for the health and performance of production services, proactively identifying and resolving issues using monitoring tools Define and track Service Level Objectives (SLOs) and error budgets, using data to inform reliability decisions Collaborate with the backend team to refactor inefficient code and improve overall system performance Required qualifications 3-5+ years of experience in Site Reliability Engineering, DevOps, or a related role managing infrastructure Deep experience managing applications on Heroku and AWS Proven experience with Infrastructure as Code (IaC), specifically Terraform Hands-on experience building and maintaining CI/CD pipelines, preferably with GitHub Actions Strong understanding of web application performance in a Ruby on Rails environment
Source listing: virtualvocations_main