As a Site Reliability Engineer, you will focus on improving the reliability and performance of our applications. You will collaborate with cross-functional teams to automate processes and ensure our systems are robust and scalable.
Responsibilities
Develop scripts for automating routine tasks and system monitoring.
Engage with software development teams to influence architectural decisions.
Utilise metrics to identify trends in system performance and proactively address issues.
Conduct disaster recovery testing and refine processes.
Act as a mentor for junior SREs and share knowledge across the team.
Foster a culture of continuous improvement and operational excellence.
Build tools to improve the reliability of the systems and improve team processes.
Requirements
Education
Bachelor's degree in Information Technology or related field
Experience
5+ years of experience in Site Reliability Engineering or Operations
Technical Skills
Automation Tools (Terraform, Ansible)
Database Management (MySQL, NoSQL)
Soft Skills
Analytical thinking
Adaptability
Certifications
Certified Kubernetes Administrator (CKA)
CompTIA Security+
Languages
English: Fluent
Advantageous
Experience with monitoring solutions: Familiarity with creating dashboards and alerts in various monitoring tools.
Experience in Agile/Scrum teams: Participation in Agile methodologies and daily stand-ups.
Benefits
Life insurance and wellness programs
Paid time off and holidays
Access to training and certification programs
Open and collaborative work environment
Company Culture
Diversity and Inclusion: We are committed to creating an inclusive and diverse workplace for everyone.
Continuous Improvement: We focus on constant growth and improvement in our practices and technology.
Transparency and Communication: We value open communication and ensure transparency in all our processes.
Status: Closed
Other Jobs in Information Technology (IT) and Software Development