$1,713.00 Fixed
Northbay Systems
Contract · Remote · Flexible hours
About the role
Northbay Systems is scaling its flagship B2B platform to support multiple tenants on a single codebase. As a Site Reliability Engineer you will own the reliability and performance of the new multi‑tenant delivery pipeline, ensuring seamless rollout for enterprise customers.
Key responsibilities
- Design and implement Terraform modules for tenant isolation across AWS accounts.
- Automate canary deployments of the tenant onboarding service using Argo CD and Helm.
- Build observability dashboards in Grafana for per‑tenant latency and error rates.
- Set up chaos‑engineering experiments with Gremlin to validate fault tolerance.
- Collaborate with product owners to define SLOs for each tenant segment.
- Document runbooks and incident response procedures for multi‑tenant failures.
Must-have skills
- Strong experience with Kubernetes (EKS) and container orchestration.
- Proficiency in Terraform or CloudFormation for IaC.
- Deep knowledge of Linux system internals and networking.
- Expertise in monitoring tools such as Prometheus, Grafana, and CloudWatch.
- Solid background in CI/CD pipelines using GitHub Actions or Jenkins.
Nice to have
- Familiarity with service mesh technologies like Istio.
- Experience with multi‑tenant SaaS architectures.
- Proposal: 0
- Less than 2 month
Gregory Brockett
,
Member since
Oct 28, 2025
Total Job