$2,653.00 Fixed
NexGen Media
Contract · Flexible hours
About the role
NexGen Media is scaling its digital platform and needs a Senior Site Reliability Engineer to ensure high availability, performance, and reliability of critical services. You will work on a focused, short‑term project to design and implement robust monitoring, automation, and incident response processes.
Key responsibilities
- Design and deploy automated scaling and self‑healing mechanisms on AWS.
- Implement end‑to‑end monitoring, alerting, and logging pipelines.
- Optimize CI/CD pipelines for rapid, reliable releases.
- Conduct root‑cause analysis of incidents and create post‑mortem reports.
- Collaborate with developers to embed reliability best practices into code.
- Document runbooks and standard operating procedures.
Must-have skills
- Extensive experience with AWS services (EC2, RDS, S3, CloudWatch).
- Strong proficiency in Docker and Kubernetes orchestration.
- Deep knowledge of Linux system administration and networking.
- Expertise in Infrastructure as Code tools (Terraform, CloudFormation).
- Proven track record in incident management and automation.
Nice to have
- Experience with Prometheus/Grafana monitoring stack.
- Familiarity with chaos engineering practices.
- Proposal: 0
- Less than 3 month
Arthur Gray
,
Member since
Oct 28, 2025
Total Job