We are looking for a passionate Site Reliability Engineer to enhance our infrastructure's reliability. You will be instrumental in implementing monitoring systems and ensuring continuous delivery pipelines are operating smoothly.
Responsibilities
Design and implement scalable infrastructure solutions.
Monitor system performance and troubleshoot issues proactively.
Collaborate with development teams to enhance system integration.
Automate deployment processes and management tasks.
Develop and maintain system and application monitoring frameworks.
Conduct root cause analysis of critical incidents and drive improvements.
Document operational procedures and best practices.
Participate in 24/7 on-call rotation to support infrastructure reliability.
Drive reliability improvements through engineering initiatives.
Advise teams on capacity planning and performance optimization.
Requirements
Education
Bachelor's degree in Computer Science or related field
Master's degree preferred
Experience
3+ years of experience in site reliability or DevOps roles
Technical Skills
AWS
Docker
Kubernetes
Terraform
Soft Skills
Communication
Problem-solving
Certifications
Google Cloud Professional
Certified Kubernetes Administrator (CKA)
Languages
English: Fluent
Advantageous
Experience with CI/CD tools: Familiarity with CI/CD tools such as Jenkins or GitLab CI for continuous integration.
Knowledge of monitoring tools: Experience with monitoring and logging solutions like Prometheus and Grafana.
Benefits
Competitive salary package with performance-based bonuses
Health insurance benefits
Flexible working arrangements
Career development funding
Company Culture
Innovation: We encourage innovative thinking and foster a culture of experimentation.
Teamwork: Collaboration is at the heart of our success.
Sustainability: We are committed to sustainable practices in our operations.
Status: Closed
Other Jobs in Information Technology (IT) and Software Development