Available for high-signal work
Systems that refuse to fail.
Pruthviraj — SRE at Microsoft. I build Kubernetes, Go automation, and observability that keep production calm at scale.
Not a badge.
A live readout.
Anyone can put a “99.9% uptime” badge on a homepage. This panel instead pulls real numbers, live, every time you load this page: the round-trip latency your own browser just measured, and the exact commit that built the response you’re looking at right now — sourced from the deployment platform itself, not typed in by hand. Nothing here is scripted or simulated.

I design outages
out of existence.
I build the platforms other engineers ship on — multi-region Kubernetes, Go tooling that deletes toil, and observability that catches failure before users feel it. My rule is simple: an outage should be a design flaw I already fixed, not a pager at 3am. I engineer systems that stay boring under pressure.
Reliability is not an accident; it's the result of high-intention engineering.
SLO-backed across multi-region AKS
self-healing controllers + smart alerting
20+ runbooks automated in Go
hardened CI/CD for 100+ microservices
how I operate
Design the outage away
Every incident review ends with a system change, not just a postmortem doc.
Automate the 3am page
If a human keeps fixing it the same way twice, it becomes a controller.
Measure, don't guess
SLOs and error budgets decide what ships — not gut feel or deadline pressure.
One discipline.
Zero shortcuts.
A running log of the systems I've operated — from single clusters to multi-region platforms at Microsoft. Expand any role to read the deployment notes.
Architecting and deploying zero-downtime certificate lifecycle platforms and managing high-availability Kubernetes environments.
Architected a zero-downtime certificate platform using Go and Azure Key Vault, reducing incidents by 80%.Developed custom Kubernetes controllers in Go to automate certificate rotation and lifecycle management.Managed reliability and capacity for large-scale production AKS environments, ensuring 99.9% availability.Built centralized observability platforms using Prometheus, Grafana, and Go exporters, reducing detection time by 40%.Azure Key VaultKubernetesPrometheusGrafanaAzure MonitorPythonGoPowerShellBash
The full stack of
staying online.
Seven domains that, together, keep distributed systems observable, recoverable, and calm under load.
Cloud & Infrastructure
Containers & Platform Engineering
Programming & Automation
Reliability Engineering
Observability & Monitoring
Linux & Networking
Distributed Systems
Credentials,
independently verified.
Not self-assessed. Every credential is an active, proctored certification from issuers including The Linux Foundation, Microsoft — each with a public verification ID.
- ActiveCKAKubernetes AdministratorThe Linux Foundation · LF-z5b8ayu6tqissued Feb 2024
- ActiveCKADKubernetes App DeveloperThe Linux Foundation · LF-1dlpfviz1wissued Apr 2024
- ActiveKCSACloud Native Security AssociateThe Linux Foundation · LF-dhfy9eb1mzissued Mar 2025
- ActiveKCNACloud Native AssociateThe Linux Foundation · LF-j5mlp40yuissued Jan 2025
- No expiryAZ-900Azure FundamentalsMicrosoft · 822D96D257A12E44issued Feb 2025
- In progressCKSKubernetes Security SpecialistThe Linux Foundationstudying now
From first commit to engineer who ships anything.
22 courses across 7 progressive stages. Each stage builds on the last — no skipping, no shortcuts. The way production skills actually compound.
24 tracks · 1254 lessons · 160+ hours of content
Browse all coursesReal incidents, real numbers, no listicles.
Scenario-driven write-ups on SRE, DevOps, backend, and platform engineering — root causes, trade-offs, and the fix that actually worked. Written for engineers who've been paged at 3am, not for search engines.
Let's build systems
that don't page at 3am.
Reliability engineering, platform architecture, or a hard scaling problem — if it needs to stay up, I want to hear about it. I read every message.