ripple

Site Reliability Engineer, Observability

Apply Now

At a Glance

Location
New York, United States
Experience
5+ years
Posted
2026-05-26T13:10:52-04:00

Key Requirements

Required Skills

AzureCI/CDDevOpsTerraform

Domain Knowledge

  • Automation
  • Engineering

Requirements

5+ years of experience in Site Reliability Engineering, DevOps, or Platform Engineering with hands-on exposure across multiple disciplines

Strong expertise in application observability tools (New Relic, Datadog, or similar), including structured logging, distributed tracing, and dashboard/alert creation

Proven experience designing and implementing CI/CD pipelines using Azure DevOps, GitHub Actions, and Octopus Deploy, with deep knowledge of deployment automation (required)

Solid understanding of application security practices including SAST, DAST, SCA tools, OWASP Top 10, and DevSecOps integration into CI/CD pipelines

Strong experience with Infrastructure as Code (IaC) using Terraform or similar tools, and knowledge of trunk-based development and version control best practices (required)

Understanding of SLO/SLI definitions, error budgets, reliability engineering principles, and familiarity with secrets management solutions (Azure Key Vault, HashiCorp Vault)

Compensation & Benefits

$160,000

$200,000 USD

Employee giving match

Mobile phone stipend

Responsibilities

Embed with stream-aligned teams in time-boxed engagements (6–12 weeks) to build capabilities across observability, CI/CD, and security domains

Coach teams on instrumenting code with logs, metrics, and traces; creating dashboards, alerts, and defining SLOs/SLIs and error budgets to improve production troubleshooting

Guide teams in designing and optimizing CI/CD pipelines using Azure DevOps, GitHub Actions, and Octopus Deploy, including trunk-based development and progressive delivery strategies

Teach teams to integrate security scanning (SAST, DAST, SCA) into pipelines, perform threat modeling, implement secrets management, and remediate vulnerabilities

Identify and remove operational bottlenecks across deployment, monitoring, and security practices to reduce lead time, MTTR, and vulnerability remediation time

Develop reusable templates, documentation, and golden path examples for pipelines, dashboards, and security controls that teams can adopt independently

About the Company

Do Your Best Work