DEVOPS Site Reliability Engineer (fully Remote)

Job ID: 36813776

Budget: $8 – $15 USD

Requirements:
• 10+ years of Software Development Life Cycle experience preferably Agile Scrum.
• 5+ years of public cloud experience (GCP or AWS).
• Experience with Incident Management Process, SRE best practices and continuous improvements.
• Excellent problem solving and analytical skills.
• Experience with Data Engineering frameworks such as Spark, Flink or similar
• Orchestration and management of containers using Kubernetes, Helm, GKE and Docker.
• Experience with Infrastructure as Code (Terraform), Python or other scripting languages.
• Experience with config management tools such as Ansible, Puppet, or Chef.
• Strong English verbal and written communications skills.

Key Responsibilities:
• Design, implementation, and automation of large-scale distributed systems.
• Build tools and automation that help company achieve higher availability, scalability, latency, and efficiency.
• Work with Engineering teams to deliver high quality software in a fast-paced environment.
• Monitor production and development environments to build preventive measures and provide a seamless customer experience.
• Work with delivery teams on software improvements to achieve higher availability and lower MTTD.
Related categories: Python Amazon Web Services Docker Kubernetes GCP AI