Fully Booked Through 2026 • Now Considering 2027 Consulting Engagements • Open to Full-Time Consulting Engagements

Engineering operations.
Minus the chaos.

I'm Vítek Urbanec, and I help engineering teams turn operational chaos into coordinated response. From incident management to operational readiness, I bring 13+ years of experience building resilient systems at scale.

13+
Years of experience
4
Industries served
Too Many
Incidents managed

What I Bring

Thirteen-plus years of operational depth, applied to long-term engineering engagements

☁

Platform Engineering at Scale

Led cloud infrastructure for 100,000+ AWS Lambda functions at Zapier. AWS published the work on its Architecture Blog.

🔌

Get the Best From Your Existing Stack

Most teams already own capable observability tools that aren't working together. I build the pieces that connect it: Terraform modules, Splunk apps, collectors. You get SLOs and error budgets without another platform to run, and the recommendation arrives already built.

🌱

Built to Outlast the Engagement

Modules, documentation, and self-serve workshops designed so your team owns the work once I leave, rather than depending on me.

📈

13+ Years in Production

Production readiness and operational resilience across enterprise, cloud, and high-scale consumer systems.

🏦

Regulated & High-Availability

Financial services, betting and gaming platforms where uptime is contractual, and AI systems.

🔥

Tested Under Real Pressure

Hands-on preparing teams for production and running incidents at companies where downtime carries serious business impact.

Case Studies

Consulting engagements and published engineering work

#SLI #SLO #Observability2.0 #SRE

SLI/SLO: From Spreadsheets to Error Budgets

In 8 months I took a Finnish retailer's Loyalty platform from monthly Excel self-reports to automated SLO measurement, burn rate alerting, and a business status page.

Read the Case Study
#AWSLambda #Terraform #PlatformEngineering #Serverless

Published · AWS Architecture Blog

Upgrading 100,000+ Lambda Functions at Scale

As Cloud Infrastructure Team Lead at Zapier, I led the team behind this work. We standardised the Terraform modules and built a canary-based upgrade tool with automatic rollback, which cut functions on deprecated runtimes by 95%. Co-authored with AWS.

Read on AWS ↗

About Me

Vítek Urbanec

I started in enterprise incident response in 2011 and have worked on production readiness and operational resilience ever since. Over time my focus moved from surviving incidents to preventing them, which mostly means building measurement teams actually trust and defaults they don't have to think about. Each role below taught me something I still use.

📊 Keep the instinct, replace the implementation (Finnish retailer, Loyalty platform)

A 12-month consulting engagement. The team was already tracking availability in a monthly Excel sheet. The shape of that report was right. The plumbing underneath it wasn't. So I kept the report and rebuilt what fed it: SLOs defined as code, a Splunk application normalising five observability sources into one view, burn rate alerting wired to on-call, and a status page leadership reads without asking an engineer. Read the case study.

⚡ Make the safe path the default path (Zapier)

As Cloud Infrastructure Team Lead I led the team that brought 100,000+ AWS Lambda functions under managed runtime upgrades: standardised Terraform modules, plus a canary tool that rolled back automatically when an upgrade misbehaved. Functions on deprecated runtimes fell by 95% and nobody had to remember to be careful. Published on the AWS Architecture Blog. I also wrote the on-call and operational readiness standards that let dozens of teams own their own services.

🎮 Structure beats heroics (Unity)

Designed Unity Finland's incident management process from the ground up, taking them from ad-hoc firefighting to something structured enough to scale as the organisation grew.

🎯 What you cannot observe, you cannot defend (Paddypower Betfair)

Ran incident response for one of Europe's largest online betting platforms, where peak load arrives on a fixture list rather than as a surprise. Built observability and resiliency for the betting exchange and the private OpenStack cloud underneath it.

🖥 The cheapest incident is the one that never happens (Rackspace)

Managed incident response from London for Rackspace's largest accounts across Europe, where every minute of downtime landed on someone's business. Owned capacity planning and operational resiliency for the hosted estate, which is the unglamorous work that stops the call coming at all.

🏦 I learned reliability where there was no margin (Dell/EMC)

Customer Engineer handling SAN/NAS storage incidents for major banks and financial institutions. Downtime there wasn't a degraded experience, it was a financial and regulatory event. It set the bar I have worked to ever since.

🤝 Ideas get better when you defend them in public (Community)

Co-founder of the SRE Finland meetup group, where I regularly present on incident management and post-mortem practice.

Consulting Expertise

Specialized support to transform your engineering operations

🔥

Production Readiness Assessment

With Focus on Incident Response & Operational Resilience

Before your service goes live, ensure your team is ready when (not if) things go wrong. Comprehensive assessment of your production readiness—covering traditional services and AI systems—with deep focus on incident response capabilities.

What's Included:

  • Review of 5-10 recent incidents
  • Interviews with 6-8 team members
  • Operational maturity scorecard
  • Gap analysis with root causes
  • Prioritized improvement roadmap
  • Smart automation & AI recommendations
  • Quick wins you can implement immediately
  • 90-min findings workshop with leadership

Typical engagement: ~2 months

Learn More
☁️

Cloud Governance & FinOps Assessment

Optimize cloud spend, improve governance, and build sustainable infrastructure practices

What's Included:

  • Cloud cost optimization & FinOps strategy
  • Multi-cloud governance frameworks
  • Infrastructure efficiency assessments
  • Cost allocation & chargeback models
  • Cloud policy & compliance design

Typical engagement: ~2 months

Learn More

Free Resources

Practice incident response and improve operational readiness

🛡️

CDN Resilience Worksheet

Practical steps to improve your CDN resilience without expensive multi-CDN setups. Lessons from the Cloudflare outage.

Open Worksheet
📋

Black Friday Operational Playbook

Your guide to managing extreme load and operational stress during high-traffic events.

Read the Playbook
🎮

Incident Commander Game

Navigate realistic incident scenarios and make decisions under pressure without real-world consequences.

Play the Game

Ready to Transform Your Operations?

Let's discuss how I can help your team build better systems and processes