As a Site Reliability Engineer in Sandton, you'll be responsible for enhancing the reliability of our services. You will engage in incident management, service monitoring, and infrastructure as code to deliver high-quality systems.
Responsibilities
Integrate and streamline tooling to enhance operational processes.
Analyze system metrics and provide insights for improvement.
Engage in root cause analysis of service outages and incidents.
Work with cross-functional teams to identify reliability gaps.
Oversee change management processes to mitigate risk.
Requirements
Education
Bachelor's degree in Information Technology or a related field
Master's degree preferred
Experience
5+ years in operations or SRE roles
Technical Skills
Kubernetes
Terraform
Soft Skills
Leadership
Adaptability
Certifications
Certified Kubernetes Administrator (CKA)
Google Associate Cloud Engineer
Languages
English: Fluent
Advantageous
Familiarity with service mesh technologies: Understanding of service mesh concepts and implementation.
Experience in performance benchmarking: Experience in performance tuning and benchmarking systems.
Benefits
Health insurance and wellness programs
Annual performance bonuses
Remote work flexibility
Training and conference attendance opportunities
Company Culture
Growth and Learning: We invest in our employees' growth with continuous learning opportunities.
Work-Life Balance: We promote a healthy work-life balance to ensure employee well-being.
Community Engagement: We engage with the community through various outreach programs and initiatives.
Status: Closed
Other Jobs in Information Technology (IT) and Software Development