← All jobs

Remote

Site Reliability Engineer

About this role

Working remotely, the full-time Site Reliability Engineer will apply a software engineering mindset to enhance the reliability and performance of production services on Heroku and AWS, collaborating closely with the engineering team to build automated solutions for scaling the platform. Key responsibilities Share responsibility for the health and performance of production services, proactively identifying and resolving issues using monitoring tools Define and track Service Level Objectives (SLOs) and error budgets, using data to inform decisions about reliability work Optimize application and database performance by identifying bottlenecks and implementing improvements in collaboration with the backend team Required qualifications 3-5+ years of experience in Site Reliability Engineering, DevOps, or related roles managing infrastructure Deep experience managing applications on Heroku and AWS Proven experience with Infrastructure as Code (IaC), specifically Terraform Hands-on experience building and maintaining CI/CD pipelines, preferably with GitHub Actions Strong understanding of web application performance, particularly in a Ruby on Rails environment

Source listing: virtualvocations_main