storyteller
Site Reliability / Production Engineer
At a Glance
- Employment
- full_time
- Compensation
- 🇹🇷 Up to EUR 32,000 per year, on a full-time, cont
Key Requirements
Required Skills
Certifications
- SAFe
Domain Knowledge
- Advertising
- Automation
- Media
- Regulatory
- SaaS
Requirements
Agency and ownership
You take responsibility for ambiguous live problems, gather evidence, choose a path and follow through after the immediate pressure has passed.
You can separate customer impact, symptoms and likely causes, make practical decisions under uncertainty and recognise when an intervention is no longer safe or bounded.
You are comfortable exploring unfamiliar systems through code, logs, APIs, data, infrastructure and command-line tools, and can make hands-on changes with a clear validation plan.
You use AI for substantive technical work - investigation, hypothesis generation, code, automation, incident analysis and workflow improvement - while supervising the agent and challenging its conclusions.
You actively look for false confidence and verify outcomes through appropriate technical and customer signals.
Responsibilities
Working closely with Support, you will assess customer impact, investigate the system, take proportionate action and bring in product developers only when their specific knowledge or judgement is genuinely needed.
You will own the technical response, solve what you reasonably can yourself and make escalations specific and useful.
Depending on the incident, you may restart or scale services, roll back deployments, change configuration, repair data, deploy a bounded fix or make a small code change.
When incidents are quiet, you will improve the reliability system: reduce alert noise, strengthen customer-outcome monitoring, improve diagnostics, create runbooks and AI Skills, automate repeated work and make our products easier to operate.
Receive automated alerts and technical escalations from Support, then establish customer impact, severity, blast radius and the current system state.
Investigate using logs, metrics, traces, dashboards, deployment history, infrastructure, databases, queues, background jobs, APIs and application code.
About the Company
Storyteller is a high-growth B2B SaaS platform that lets companies integrate Stories into their own apps and websites. Popularised by Instagram and Snapchat, Stories help our clients increase engagement, retention and revenue.
Our platform includes SDKs for Web, iOS and Android, alongside publishing tools, analytics and advertising support—giving enterprises a complete Stories solution in days. We work with globally recognised sports and media brands, and your work will be used live by millions of people.
Our production environment spans Storyteller, Storypilot and the services that support our customers’ live workflows. Reliability is therefore about more than infrastructure: we need to understand when customers are affected, respond quickly, coordinate the right people and improve our systems after every material incident.
About the Role
We are hiring two Site Reliability / Production Engineers. You will be the first technical response for live incidents during your coverage window.
Working closely with Support, you will assess customer impact, investigate the system, take proportionate action and bring in product developers only when their specific knowledge or judgement is genuinely needed.