$2,393.00 Fixed
BrightMart Solutions
Contract · Flexible hours
About the role
We are seeking a Senior Site Reliability Engineer to strengthen the reliability and performance of our e‑commerce platform. The contractor will design, implement, and automate monitoring, incident response, and scaling solutions for a high‑traffic environment.
Key responsibilities
- Develop and maintain automated CI/CD pipelines for microservices.
- Implement robust monitoring, alerting, and logging using Prometheus, Grafana, and ELK.
- Design auto‑scaling strategies and capacity planning on AWS.
- Troubleshoot production incidents and lead post‑mortems.
- Collaborate with development teams to embed reliability best practices.
- Document runbooks and operational procedures.
Must-have skills
- Extensive experience with Linux system administration.
- Deep knowledge of Docker and Kubernetes orchestration.
- Proficiency in AWS services (EC2, RDS, S3, CloudWatch).
- Strong scripting skills (Python, Bash) for automation.
- Solid understanding of SRE principles and incident management.
Nice to have
- Experience with Terraform or other IaC tools.
- Familiarity with service mesh technologies.
- Proposal: 0
- Less than 3 month
Michael Houck
,
Member since
Oct 28, 2025
Total Job