penninteractive
Incident Commander
At a Glance
- Location
- Canada
- Work Regime
- remote
- Posted
- 2026-07-24T16:51:12-04:00
Key Requirements
Required Skills
Domain Knowledge
- Automation
- Engineering
Requirements
Experience and understanding of Containerization (Docker & Kubernetes preferred)
Automation: Understanding of configuration management and infrastructure as code tools.
Terraform, Ansible, Helm, etc.
Comfortable within Linux environments and needs.
Experience working with AWS, GCP, and on-premises environments.
Nice to have: Postgres, MySQL, Elastic Search, Kafka, Redis, Terragrunt, Prometheus, Python, Talos Linux
Compensation & Benefits
Competitive compensation package.
Fun, relaxed work environment.
Education and conference reimbursements.
Parental leave top up.
Opportunities for career progression and mentoring others.
#LI-REMOTE
Responsibilities
As part of the team, you will be working with a team of smart, friendly, and dedicated Engineers, Product Managers and Designers determined to deliver some of the best apps the market has to offer.
We are looking for an Incident Commander to join our site reliability team, to work cross-functionally across engineering, and be the front line for incidents and working with Release Engineering to help prevent new events.
Classifying and documenting all incidents and carrying out support, assisting and driving all incidents, regarding investigation, hierarchical and technical escalation, diagnosis and recovery and root cause analysis.
Additionally driving improvements to our service delivery and release processes based on disruption reports.
About the Company
• Drive and enhance collaboration with other Incident Commanders, Customer Support, Application and Engineering teams - cross-functional teams to lead real-time incident management.
• Provides Leadership for developing Practices, Frameworks, Process Flows, Templates and Process Guides
• Continuously improve and enhance the internal framework, methodology, processes, and tools
• Developing and maintaining key practical capabilities
• Collaborating with SRE Teams and Infrastructure teams to identify requirements and gaps resulting in downtime or blindspots.
• Recommends innovative solutions that enable the organization to deliver on its objectives and goals.