
Engineering systemsthat stay calmunder pressure.
I'm a Site Reliability Engineer at Microsoft, based in Pune, India. I specialize in Kubernetes at scale, Go automation, and observability that catches failure first.

I design outages
out of existence.
I build the platforms other engineers ship on — multi-region Kubernetes, Go tooling that deletes toil, and observability that catches failure before users feel it. My rule is simple: an outage should be a design flaw I already fixed, not a pager at 3am. I engineer systems that stay boring under pressure.
Reliability is not an accident; it's the result of high-intention engineering.
SLO-backed across multi-region AKS
self-healing controllers + smart alerting
20+ runbooks automated in Go
hardened CI/CD for 100+ microservices
how I operate
Design the outage away
Every incident review ends with a system change, not just a postmortem doc.
Automate the 3am page
If a human keeps fixing it the same way twice, it becomes a controller.
Measure, don't guess
SLOs and error budgets decide what ships — not gut feel or deadline pressure.
One discipline.
Zero shortcuts.
A running log of the systems I've operated — from single clusters to multi-region platforms at Microsoft. Expand any role to read the deployment notes.
Architecting and deploying zero-downtime certificate lifecycle platforms and managing high-availability Kubernetes environments.
Architected a zero-downtime certificate platform using Go and Azure Key Vault, reducing incidents by 80%.Developed custom Kubernetes controllers in Go to automate certificate rotation and lifecycle management.Managed reliability and capacity for large-scale production AKS environments, ensuring 99.9% availability.Built centralized observability platforms using Prometheus, Grafana, and Go exporters, reducing detection time by 40%.Azure Key VaultKubernetesPrometheusGrafanaAzure MonitorPythonGoPowerShellBash
The full stack of
staying online.
Seven domains that, together, keep distributed systems observable, recoverable, and calm under load.
Cloud & Infrastructure
Containers & Platform Engineering
Programming & Automation
Reliability Engineering
Observability & Monitoring
Linux & Networking
Distributed Systems
Credentials,
independently verified.
Not self-assessed. Every credential is an active, proctored certification from issuers including The Linux Foundation, Microsoft — each with a public verification ID.
- ActiveCKAKubernetes AdministratorThe Linux Foundation · LF-z5b8ayu6tqissued Feb 2024
- ActiveCKADKubernetes App DeveloperThe Linux Foundation · LF-1dlpfviz1wissued Apr 2024
- ActiveKCSACloud Native Security AssociateThe Linux Foundation · LF-dhfy9eb1mzissued Mar 2025
- ActiveKCNACloud Native AssociateThe Linux Foundation · LF-j5mlp40yuissued Jan 2025
- No expiryAZ-900Azure FundamentalsMicrosoft · 822D96D257A12E44issued Feb 2025
- In progressCKSKubernetes Security SpecialistThe Linux Foundationstudying now
Real incidents, real numbers, no listicles.
Scenario-driven write-ups on SRE, DevOps, backend, and platform engineering — root causes, trade-offs, and the fix that actually worked. Written for engineers who've been paged at 3am, not for search engines.
Let's build systems
that don't page at 3am.
Reliability engineering, platform architecture, or a hard scaling problem — if it needs to stay up, I want to hear about it. I read every message.