clockworksystems

Senior Software Engineer – Network Observability

Apply Now

At a Glance

Location
Onsite Palo Alto, California, United States
Compensation
expected base salary range is $180,000 - $260,000. The offered compensation pac
Posted
2026-05-01T20:56:12-04:00

Key Requirements

Required Skills

AWSAzureGCPKubernetesLinuxPythonRust

Domain Knowledge

  • Automation
  • Cloud
  • Engineering
  • Regulatory

Benefits & Perks

Health Insurance

mpensation. A great benefits package. Catered lunch. Compensation for this p

Requirements

Strong hands-on programming experience in C++, Go, Python, Rust, or similar systems programming languages.

Proven experience leading engineering teams, major technical initiatives, or complex infrastructure projects.

Experience building distributed systems, backend services, telemetry pipelines, or observability platforms.

Hands-on experience with RDMA, RoCE, InfiniBand, or other high-performance network fabrics.

Strong knowledge of Linux networking, TCP/IP, DNS, HTTP, routing, MTU, congestion control, packet loss, latency, and performance tuning.

Experience with traceroute-style diagnostics, path discovery, network reachability checks, synthetic probes, or active network measurements.

Responsibilities

We are seeking an experienced Tech Lead to lead the architecture, development, and scaling of a high-performance network monitoring and observability platform.

This role will focus on building systems that provide deep visibility into RDMA, RoCE, InfiniBand, and TCP/IP networks.

The ideal candidate has strong experience in distributed systems, Linux networking, and modern observability stacks (e.g., Grafana/Prometheus).

Lead architecture, design, and development of scalable network monitoring platforms for high-performance RDMA, RoCE, InfiniBand, and TCP/IP infrastructure.

Build backend telemetry services, observability dashboards, alerts, diagnostics, anomaly detection, SLA monitoring, and traffic analysis workflows.

Troubleshoot complex production issues across application, OS, server, RDMA, and network layers while optimizing low-latency collection, aggregation, and alerting.