I'm Vítek Urbanec, and I help engineering teams turn operational chaos into coordinated response. From incident management to operational readiness, I bring 13+ years of experience building resilient systems at scale.
Thirteen-plus years of operational depth, applied to long-term engineering engagements
Led cloud infrastructure for 100,000+ AWS Lambda functions at Zapier. AWS published the work on its Architecture Blog.
Most teams already own capable observability tools that aren't working together. I build the pieces that connect it: Terraform modules, Splunk apps, collectors. You get SLOs and error budgets without another platform to run, and the recommendation arrives already built.
Modules, documentation, and self-serve workshops designed so your team owns the work once I leave, rather than depending on me.
Production readiness and operational resilience across enterprise, cloud, and high-scale consumer systems.
Financial services, betting and gaming platforms where uptime is contractual, and AI systems.
Hands-on preparing teams for production and running incidents at companies where downtime carries serious business impact.
Consulting engagements and published engineering work
In 8 months I took a Finnish retailer's Loyalty platform from monthly Excel self-reports to automated SLO measurement, burn rate alerting, and a business status page.
Read the Case StudyPublished · AWS Architecture Blog
As Cloud Infrastructure Team Lead at Zapier, I led the team behind this work. We standardised the Terraform modules and built a canary-based upgrade tool with automatic rollback, which cut functions on deprecated runtimes by 95%. Co-authored with AWS.
Read on AWS ↗
I started in enterprise incident response in 2011 and have worked on production readiness and operational resilience ever since. Over time my focus moved from surviving incidents to preventing them, which mostly means building measurement teams actually trust and defaults they don't have to think about. Each role below taught me something I still use.
A 12-month consulting engagement. The team was already tracking availability in a monthly Excel sheet. The shape of that report was right. The plumbing underneath it wasn't. So I kept the report and rebuilt what fed it: SLOs defined as code, a Splunk application normalising five observability sources into one view, burn rate alerting wired to on-call, and a status page leadership reads without asking an engineer. Read the case study.
As Cloud Infrastructure Team Lead I led the team that brought 100,000+ AWS Lambda functions under managed runtime upgrades: standardised Terraform modules, plus a canary tool that rolled back automatically when an upgrade misbehaved. Functions on deprecated runtimes fell by 95% and nobody had to remember to be careful. Published on the AWS Architecture Blog. I also wrote the on-call and operational readiness standards that let dozens of teams own their own services.
Designed Unity Finland's incident management process from the ground up, taking them from ad-hoc firefighting to something structured enough to scale as the organisation grew.
Ran incident response for one of Europe's largest online betting platforms, where peak load arrives on a fixture list rather than as a surprise. Built observability and resiliency for the betting exchange and the private OpenStack cloud underneath it.
Managed incident response from London for Rackspace's largest accounts across Europe, where every minute of downtime landed on someone's business. Owned capacity planning and operational resiliency for the hosted estate, which is the unglamorous work that stops the call coming at all.
Customer Engineer handling SAN/NAS storage incidents for major banks and financial institutions. Downtime there wasn't a degraded experience, it was a financial and regulatory event. It set the bar I have worked to ever since.
Co-founder of the SRE Finland meetup group, where I regularly present on incident management and post-mortem practice.
Specialized support to transform your engineering operations
With Focus on Incident Response & Operational Resilience
Before your service goes live, ensure your team is ready when (not if) things go wrong. Comprehensive assessment of your production readiness—covering traditional services and AI systems—with deep focus on incident response capabilities.
Typical engagement: ~2 months
Optimize cloud spend, improve governance, and build sustainable infrastructure practices
Typical engagement: ~2 months
Practice incident response and improve operational readiness
Practical steps to improve your CDN resilience without expensive multi-CDN setups. Lessons from the Cloudflare outage.
Open WorksheetYour guide to managing extreme load and operational stress during high-traffic events.
Read the PlaybookNavigate realistic incident scenarios and make decisions under pressure without real-world consequences.
Play the GameLet's discuss how I can help your team build better systems and processes