archer56

Sr Staff Site Reliability Engineer

Apply Now

At a Glance

Location
San Jose, California, United States
Experience
3+ years
Compensation
targeting a base pay between $133,400 - $200,000. Actual compensation offered
Posted
2026-05-28T13:48:56-04:00

Key Requirements

Required Skills

AWSBashCI/CDDevOpsDockerJenkinsKafkaKubernetesPython

Domain Knowledge

  • Engineering
  • Regulatory

Requirements

3+ years of experience in Site Reliability Engineering, DevOps, or a similar role with a strong focus on operational excellence.

Deep expertise in Amazon EKS, including cluster provisioning, management, and troubleshooting.

Extensive experience with observability tools and practices, including Prometheus, Grafana, ELK stack, or similar.

Proven track record in designing and implementing robust data pipelines (e.g., Kafka, Airflow, Spark).

Strong background in CI/CD methodologies and tools (e.g., Jenkins, GitLab CI, ArgoCD).

Expert-level knowledge of cloud platforms (AWS preferred), including infrastructure-as-code principles.

Responsibilities

Implement and maintain the infrastructure and pipeline required for an internal LLM-powered chat service, potentially leveraging platforms like OpenRouter or similar alternatives.

implement and maintain highly available, scalable, and secure cloud-native infrastructure on Amazon Elastic Kubernetes Service (EKS).

Develop and implement comprehensive observability strategies, including monitoring, logging, and alerting, to ensure the health and performance of our systems.

Architect and optimize data pipelines to ensure efficient and reliable data flow across various platforms.

Drive the continuous improvement of our CI/CD pipelines, promoting best practices for automated testing, deployment, and release management.

Champion cloud-first strategies, leveraging the full capabilities of cloud platforms for infrastructure, services, and operations.