As a Site Reliability Engineer, you will play a crucial role in the operational success of our company. Your responsibilities will include enhancing system reliability, troubleshooting and resolving incidents, and monitoring system health.
Responsibilities
Enhance system observability with monitoring and alerting systems.
Work closely with development teams to design scalable solutions.
Observe and troubleshoot system performance in production.
Manage cloud infrastructure and help with DevOps practices.
Assist in the development of disaster recovery procedures.
Engage in capacity planning and forecasting.
Implement security best practices in operational processes.
Drive initiatives to improve operational efficiency.
Requirements
Education
Bachelor's degree in Computer Science, IT, or equivalent experience
Master's degree is advantageous
Experience
5+ years of experience in Site Reliability Engineering or a related field
Technical Skills
Linux Administration
Python
Containerization (Docker/Kubernetes)
Cloud Computing (AWS/Azure)
Soft Skills
Analytical Skills
Team Collaboration
Certifications
AWS Certified Solutions Architect
Certified Kubernetes Administrator (CKA)
Languages
English: Fluent
Advantageous
Familiarity with Infrastructure as Code (IaC): Experience with configurations, preferably Terraform or Ansible.
Experience with monitoring tools: Proficient in using monitoring and logging tools such as Prometheus and Grafana.
Benefits
Competitive salary with performance bonuses.
Flexible working hours.
Health and wellness benefits.
Training and professional development opportunities.
Company Culture
Innovative Environment: We foster a culture of innovation and creativity, encouraging new ideas and strategies.
Supportive Team: Our team collaborates closely and supports each other’s growth and development.
Status: Open
Other Jobs in Information Technology (IT) and Software Development