$2,717.00 Fixed
Vertex Analytics
Contract · Remote · Flexible hours
About the role
Vertex Analytics is modernizing the reliability of its flagship customer‑facing platform, AuroraStream. As a Site Reliability Engineer you will drive a performance and reliability overhaul, ensuring seamless user experiences under peak load.
Key responsibilities
- Instrument AuroraStream micro‑services with OpenTelemetry and create real‑time dashboards in Grafana.
- Design and implement auto‑scaling policies in AWS EC2 Auto Scaling and EKS for traffic spikes.
- Refactor existing CI/CD pipelines in GitHub Actions to include canary deployments and automated rollback.
- Conduct chaos engineering experiments using Gremlin to validate fault tolerance.
- Collaborate with product owners to define SLOs/SLA thresholds and embed them in Service Level Objectives.
- Document runbooks and hand‑off procedures for incident response within Confluence.
Must-have skills
- Deep experience with Linux system administration and performance tuning.
- Proficiency in Docker, Kubernetes (EKS), and Helm chart management.
- Strong knowledge of AWS services (EC2, RDS, CloudWatch, S3).
- Expertise in monitoring, alerting, and observability tools (Prometheus, Grafana, Loki).
- Solid scripting skills in Bash/Python for automation.
Nice to have
- Experience with Terraform or CloudFormation for IaC.
- Background in chaos engineering or fault injection frameworks.
- Proposal: 0
- Less than 3 month
Sam Vasquez
,
Member since
Oct 27, 2025
Total Job