$2,053.00 Fixed
NexGen Media
Contract · Flexible hours
About the role
NexGen Media seeks a Senior Site Reliability Engineer to design, implement, and maintain high‑availability services for a fast‑growing streaming platform. The project focuses on automating infrastructure, improving observability, and ensuring resilient deployments across cloud environments.
Key responsibilities
- Design and implement automated CI/CD pipelines for micro‑service deployments.
- Develop and maintain Terraform/IaC modules for AWS resources.
- Configure monitoring, alerting, and logging using Prometheus, Grafana, and ELK stack.
- Optimize Kubernetes clusters for scalability and fault tolerance.
- Conduct incident response, root‑cause analysis, and post‑mortems.
- Collaborate with development teams to embed reliability best practices.
Must-have skills
- Extensive experience with AWS services and cloud networking.
- Strong proficiency in Docker and Kubernetes orchestration.
- Deep knowledge of Terraform or similar IaC tools.
- Solid Linux system administration and scripting (Bash/Python).
- Proven track record in SRE practices, monitoring, and alerting.
Nice to have
- Experience with service mesh technologies (e.g., Istio).
- Familiarity with chaos engineering tools.
- Proposal: 0
- Less than 3 month
James Powell
,
Member since
Oct 28, 2025
Total Job