Remote
Site Reliability Engineer
About this role
To support a growing product team, the full-time remote Site Reliability Engineer will manage the health and performance of production services on Heroku and AWS, optimize application performance, and collaborate on infrastructure improvements. Key responsibilities Share responsibility for the health and performance of production services, proactively identifying and resolving issues using monitoring tools Define and track Service Level Objectives (SLOs) and error budgets, using data to inform reliability decisions Optimize application and database performance by identifying bottlenecks and implementing improvements Required qualifications 3-5+ years of experience in Site Reliability Engineering, DevOps, or a similar role managing infrastructure Deep experience managing applications on Heroku and AWS Proven experience with Infrastructure as Code (IaC), specifically Terraform Hands-on experience building and maintaining CI/CD pipelines, preferably with GitHub Actions Strong understanding of web application performance, particularly in a Ruby on Rails environment
Source listing: virtualvocations_main