In 8 months I replaced a Finnish retailer's manual Excel process with an engineering-grade SLO program: error budgets, burn rate alerting, and a status page the business can read. All of it built on the observability tools they already owned.
When I joined, the Loyalty platform team was already trying to track availability. Every month, team leads got the same question: "How much downtime did you have?" The answers went into a shared Excel sheet. They had the right instinct: measure reliability, report on it, hold teams accountable.
In practice, though, most teams had no idea how much downtime they'd actually had. Nothing measured it automatically, nobody shared a definition of what counted as a failure, and no one could compare services or spot a trend over time. A few engineers had experimented with click-ops SLO tooling, but none of it reached production scale.
When something went wrong, answering "is this service healthy?" meant tracking down the right engineer and asking them to read the dashboards. Nobody else could read the signal.
Technical fragmentation, organisational breadth, and a knowledge gap all needed to move together.
Teams used AWS CloudWatch, Dynatrace, and Splunk, each with its own SLO concept. Splunk users had built SLIs from log searches, which worked in principle but were fragile and inconsistent. No view across all of them existed.
Definitions weren't versioned, reviewable, or team-owned, so no one could answer "what changed, when, and why?" Click-ops also capped how precise an SLI could get: anything conditional, like a different latency threshold per endpoint, was impractical to maintain by hand.
When shipping to production, teams added alerts on every metric they could think of. Noise was high. Burn rate alerting required a different mental model entirely.
Stakeholders and product managers had to ask an engineer to find out if a service was healthy. There was no self-serve signal, and no urgency to create one.
The Loyalty platform spanned 10+ services across several product teams with different stacks, tooling, and working styles. Any rollout had to work for all of them.
Teams had stress-tested their services, but none of them had worked out how those services actually fail, what the observable symptoms look like, or whether their alerting would catch it.
Excel is the default tech stack for anyone who needs to do some math. It's what you reach for first, and for a lot of jobs it's the right answer. The sheet the team had built asked a sound question and put every service's availability in one place.
It stopped fitting at everything around the math. A spreadsheet can't collect its own data, can't hold teams to one definition of a failed request, and can't fill itself in when the month gets busy. Those numbers were only ever as good as someone's recollection of a bad week.
So I kept the report and rebuilt the plumbing underneath it. In each phase below I swapped a manual step for real measurement, while keeping the questions the business already asked.
One constraint shaped all of it: don't grow the tooling footprint. The teams already had CloudWatch, Dynatrace, Sentry, Splunk and synthetic monitors, and none of it was going away. Rather than introduce another platform to own, I worked around the limits of what was already there and built only the components needed to make those tools serve SLOs. Nothing here asks the organisation to run one more thing.
I started by getting SLO definitions out of tool UIs and into version control. I wrote a Terraform module that defines SLOs as code for AWS CloudWatch, the primary observability platform for most Loyalty services, and moved every definition into a central SLO monorepo: versioned, reviewable, team-owned.
That gave teams their first automated reliability data without asking them to change how they thought about their services yet.
Writing them in code also let me make the SLIs sharper than a tool UI allows. I
made latency thresholds conditional within a single SLI definition,
so what counts as a "good" event depends on the combination of HTTP method and
API route. A
GET on a read endpoint is held to a tight threshold; a
POST on the same API that legitimately does more work gets its own.
That gives one SLI and one error budget, with thresholds that match how each
endpoint actually behaves, instead of a single loose number that hides read
regressions or a tight one that cries wolf on every write.
This applies an Observability 2.0 instinct to SLOs. I keep the high-cardinality dimensions, method and route, inside the definition of a good event instead of flattening them into one pre-aggregated latency number and setting a single threshold on top. The SLI gets to be as specific as the system it measures.
With CloudWatch SLOs flowing, fragmentation became the next problem. CloudWatch, Dynatrace, and Splunk each showed only their own services, so I built a normalisation layer.
Sliisloo is the Splunk application I wrote for it, and it runs inside the Splunk the organisation was already paying for. It pulls 10-minute samples of SLO state into a shared index from CloudWatch, Dynatrace, Sentry, and synthetic monitors. For Splunk-native log-search SLOs I added a report that pushes records into the same index. One schema. One dashboard. Regardless of source.
Most of the work was in the seams. CloudWatch will hold an SLO object but won't show it next to anything else; Splunk will visualise anything you can get into an index but has no native concept of an SLO. Neither limitation was going to change, so Sliisloo is the set of small components that make one tool's output legible to the other. No new platform, no migration, no second place to look.
I applied the same Observability 2.0 principle at the platform level: one event store as the source of truth, with the status page, trends, and historical reporting all derived from it, instead of four tools each maintaining a partial picture nobody can reconcile.
Beyond status and trends, I gave Sliisloo an annotations layer: engineers log what happened at any point on the SLO timeline, whether that is an incident, a deployment, or a campaign launch. The context lives alongside the data, so the chart tells the full story, not just the numbers.
The name came from the team, not from me. One of the Loyalty engineers coined Sliisloo (pronounced SlìSlò), a portmanteau that solved a genuine problem. I'm Czech, the team is Finnish, and saying "S-L-I, S-L-O" aloud in every meeting was a mouthful in either language. It stuck, and ended up an internal brand, stickers included.
I took that as a good sign. People don't nickname something they're merely tolerating, and by that point the team knew the concepts well enough to play with them.
With SLOs measured and visible, I made them actionable. I configured burn rate alerting for every service, so alerts fire when the error budget burns too fast, before a breach becomes inevitable. That replaced the alert-on-everything default with a signal that tells on-call engineers what actually matters. I wired PagerDuty in as the alerting destination to close the end-to-end on-call loop.
Ahead of the live session I recorded a 17-minute video covering SLO fundamentals, error budget mechanics, and burn rate calculations, using examples from the Sliisloo dashboard. Engineers watched it at their own pace as preparation.
For the live workshop I wrote a fictional scenario: teams took a made-up service to production together, defining its SLIs and SLOs as a group exercise, and I facilitated the discussion afterwards. I built the format to stand on its own, so future engineers can work through the same video and workshop to onboard themselves, without me in the room.
With SLOs defined and alerting running, one question remained: does any of this actually work under failure? I set up fault injection testing on core Loyalty services, deliberately inducing failures to confirm the SLIs measured them accurately and the burn rate alerts fired as expected. Along the way we documented the failure symptoms the teams had never mapped.
I built a service status page powered entirely by live SLO data. No manual updates. No engineer needed to interpret it. Leadership reads system health at a glance, and it feeds straight into wider business reporting.
The shift that made the program stick wasn't a technical one. Putting an SLO on a new service still takes deliberate work. Someone writes the definition, picks the thresholds, opens the pull request. Nothing lands automatically. What changed is that nobody argues about whether to do it.
Before I started, measuring a new service meant building the measurement first: no module to call, no repository to add to, no agreed definition of a good event to borrow. That made an SLO easy to defer, and deferring it was the normal outcome. With the module and the monorepo in place, the work is small enough that skipping it is harder to justify than doing it.
So teams now treat an SLO as part of shipping a service rather than something to retrofit once it's already in production. That expectation is the durable outcome. The tooling only made it cheap enough to hold.
A service's status is its worst-performing SLO. That's the entire rule I built the status page on. It removes engineering judgement from reading system health, gives leadership an honest signal they can act on, and creates a forcing function: a weak SLO drags down the whole service state, which keeps SLO quality high over time.
These are the Sliisloo dashboards I built. In production they run on live SLO data, with no manual updates and no spreadsheets.
About these screenshots: the layouts, views, and functionality are exactly what I built, but the data is mocked. I replaced the service names, attainment figures, and statuses with representative sample values, so no customer metrics appear here.
I designed every deliverable to outlast me. The team runs all of it independently now that the engagement has ended.
CloudWatch SLOs defined as code, with latency thresholds that vary by HTTP method and route inside a single SLI. Reusable, versioned, reviewed like any other infrastructure change.
FoundationCentral GitOps repository as the single source of truth for all SLO definitions across the Loyalty platform.
Team-ownedPutting an SLO on a new service is now part of shipping it, not a retrofit. It still takes deliberate work. The module and monorepo just made it cheap enough that teams stopped deferring it.
Cultural shiftUnified SLO dashboard pulling from CloudWatch, Dynatrace, Sentry, synthetic monitors, and Splunk log-search SLOs. One view, regardless of source.
Built & deployedRunning on all 10+ services. On-call engineers get a meaningful signal before a breach is inevitable, not after.
All servicesWired in as the alerting destination for burn rate alerts, closing the end-to-end on-call loop.
Integrated17-minute fundamentals video plus a hands-on fictional scenario workshop. Reusable for onboarding, with no external facilitation needed.
Self-serveI broke core services on purpose to verify SLI accuracy, confirm alert behaviour, and document failure symptoms.
Core servicesLive SLO-powered status page. Each service's health = its worst SLO. No manual updates, no engineer needed to read it.
Real-timeSLOs alongside business KPIs, making reliability a business metric rather than just an engineering one. In progress.
NextVítek was the missing link in the success of rolling out SLOs for the organisation.
The progress of observability improvements this year has been awesome.
This case study covers the SLO program. Across the same 12 months I also worked on alert fatigue reduction, post-mortem facilitation, incident management coaching, engineering leadership, and FinOps. Each of those will get its own case study.
Whether you're starting from scratch or scaling up existing SLO work, let's talk.
Book a Discovery Call Send an Email