Remote
Site Reliability Engineer
About this role
To support a growing product team, the full-time remote Site Reliability Engineer will apply a software engineering mindset to enhance platform reliability and performance, collaborating with engineers to build automated solutions and optimize application performance. Key responsibilities Share responsibility for the health and performance of production services on Heroku and AWS, using monitoring tools to identify and resolve issues Define and track Service Level Objectives (SLOs) and error budgets for key services, using data to inform reliability decisions Collaborate with the backend team to refactor inefficient code and improve overall system performance Required qualifications 3-5+ years of experience in Site Reliability, DevOps, or related roles managing infrastructure Deep experience managing applications on Heroku and AWS Proven experience with Infrastructure as Code (IaC), specifically Terraform Hands-on experience building and maintaining CI/CD pipelines, preferably with GitHub Actions Strong understanding of web application performance, particularly in a Ruby on Rails environment
Source listing: virtualvocations_main