Remote
Site Reliability Engineer
About this role
Working remotely, the full-time Site Reliability Engineer will apply a software engineering mindset to enhance the reliability and performance of production services on Heroku and AWS, while collaborating with the engineering team to build automated solutions for scaling operations. Key responsibilities Share responsibility for the health and performance of production services, using monitoring tools to proactively identify and resolve issues Define and track Service Level Objectives (SLOs) and error budgets, utilizing data to inform reliability decisions Optimize application and database performance by identifying bottlenecks and implementing necessary improvements Required qualifications 3-5+ years of experience in a Site Reliability, DevOps, or related role managing infrastructure Deep experience managing applications on Heroku and AWS Proven experience with Infrastructure as Code (IaC), specifically Terraform Hands-on experience building and maintaining CI/CD pipelines, preferably with GitHub Actions Strong understanding of web application performance, particularly in a Ruby on Rails environment
Source listing: virtualvocations_main