penninteractive

Incident Commander

Apply Now

At a Glance

Location
Canada
Work Regime
remote
Posted
2026-07-24T16:51:12-04:00

Key Requirements

Required Skills

AWSDockerGCPKafkaKubernetesLinuxMySQLPostgreSQLPythonRedisTerraform

Domain Knowledge

  • Automation
  • Engineering

Requirements

Experience and understanding of Containerization (Docker & Kubernetes preferred)

Automation: Understanding of configuration management and infrastructure as code tools.

Terraform, Ansible, Helm, etc.

Comfortable within Linux environments and needs.

Experience working with AWS, GCP, and on-premises environments.

Nice to have: Postgres, MySQL, Elastic Search, Kafka, Redis, Terragrunt,  Prometheus, Python, Talos Linux

Compensation & Benefits

Competitive compensation package.

Fun, relaxed work environment.

Education and conference reimbursements.

Parental leave top up.

Opportunities for career progression and mentoring others.

#LI-REMOTE

Responsibilities

As part of the team, you will be working with a team of smart, friendly, and dedicated Engineers, Product Managers and Designers determined to deliver some of the best apps the market has to offer.

We are looking for an Incident Commander to join our site reliability team, to work cross-functionally across engineering, and be the front line for incidents and working with Release Engineering to help prevent new events.

Classifying and documenting all incidents and carrying out support, assisting and driving all incidents, regarding investigation, hierarchical and technical escalation, diagnosis and recovery and root cause analysis.

Additionally driving improvements to our service delivery and release processes based on disruption reports.

About the Company

• Drive and enhance collaboration with other Incident Commanders, Customer Support, Application and Engineering teams - cross-functional teams to lead real-time incident management.

• Provides Leadership for developing Practices, Frameworks, Process Flows, Templates  and Process Guides

• Continuously improve and enhance the internal framework, methodology, processes, and tools

• Developing and maintaining key practical capabilities

• Collaborating with SRE Teams and Infrastructure teams to identify requirements and gaps resulting in downtime or blindspots.

• Recommends innovative solutions that enable the organization to deliver on its objectives and goals.