$1,823.00 Fixed
NexGen Media
Contract · Remote · Flexible hours
About the role
NexGen Media is building the StreamPulse analytics enablement layer to give operations teams real‑time visibility into data pipelines. As a Site Reliability Engineer you will own the reliability, scaling, and observability of this platform, ensuring data flows uninterrupted from ingestion to dashboard.
Key responsibilities
- Design and implement Terraform modules for AWS EKS clusters supporting StreamPulse workloads.
- Set up Prometheus‑Grafana monitoring stacks with custom alerts for pipeline latency and failure rates.
- Automate CI/CD pipelines using GitHub Actions to deploy micro‑services written in Go and Python.
- Optimize Kafka consumer groups and Spark job configurations for cost‑effective throughput.
- Conduct chaos engineering experiments with Gremlin to validate resilience under load spikes.
- Document runbooks and incident response procedures for data‑pipeline incidents.
Must-have skills
- Strong experience with Kubernetes orchestration and Helm charts.
- Proficiency in AWS services (EKS, S3, CloudWatch, IAM).
- Infrastructure‑as‑Code expertise using Terraform or CloudFormation.
- Solid Linux systems administration and scripting (Bash, Python).
- Deep understanding of observability tools (Prometheus, Grafana, Loki).
Nice to have
- Experience with data‑streaming frameworks such as Kafka or Flink.
- Familiarity with SRE practices like SLO/SLI definition and error budgeting.
- Proposal: 0
- Less than 2 month
George Russel
,
Member since
Oct 28, 2025
Total Job