SOURAV KUNDU sys-eng / sre --:--:-- operational
Sourav Kundu
svc/identity · tier-0 · owner: self
devops · sre · platform · infrastructure

"You didn't notice anything.That was the job."

Five years keeping production boring — three and a half of them on 24/7 streaming infrastructure at Warner Bros. Discovery, and the last stretch as the infrastructure owner for a six-brand iGaming operator where a bad deploy is measured in revenue per minute. I build the layer nobody thinks about until it's gone: pipelines, clusters, edge, observability, and the pager that wakes someone up.

iGamingstreaming at scale multi-region24/7 on-call Go · Terraform · Kubernetes
where it runs local time, right now
currentlyopen to work

Skills

Not a list of logos — the path a request takes through everything I run.
a user types https://brand.example/ 
https://brand.example
↵ enter
someone, somewhere, presses enter
brand.example.com ?
root13 servers
TLD.com
authoritativens · cloudflare
A → 104.21.42.7
the name becomes an address
WAF · rate limit ✕ bot blocked
TLS 1.3 handshake🔒 secure
Worker · route to healthy regionedge fn
200+ cities. it never reached my origin yet
vpc · private subnets
load balancer
target✓ healthy
target✓ healthy
target✕ drained
routed inside the perimeter, around the sick one
node · eks4 pods
▲ hpa · queue depth rising → scheduling 3 more
capacity arrives before the user notices
func Handle(w http.ResponseWriter, r *http.Request) {
ctx := otel.Start(r.Context(), "render")
page, err := svc.Build(ctx, brandFrom(r))
if err != nil { // fail loudly, not silently }
w.WriteHeader(200)
}
the process I write, not only operate
redis · session + cachemiss
state, and whether it survives
trace · spans
metrics · p95
logs · correlated by trace id
info render.complete brand=spinbet 187ms
info cache.miss key=page:/ ttl=60
info hpa.scaled from=4 to=7
now I can answer questions at 3am
200 OK
page rendered in 187 ms · the user noticed nothing
that was the job
running underneath every step of that journey
daily — in my hands most weeks shipped — built and run in production working — used it for real, not daily exploring — learning it now

Runbook

problem classes · ranked by blast radius
Not a log of anyone's outages — the recurring problems I get called for, and the way I work them.

Deploy pipeline

Pick one. Each opens with its architecture, why it exists, and how it works.
index

Changelog

git log --author=sourav
Where the five years went.

Shell access

tty1 · it actually works
Type help. Or sudo, if you're feeling brave.
sourav@prod — zshtty1
keyboardshortcuts

Page me

escalation policy
Escalates in 15 minutes if unacknowledged. I would know — I wrote the bot.
availabilitycurrent

Open to senior DevOps, SRE and platform engineering conversations — remote, or relocation for the right team. I care most about owning a system end to end rather than a slice of one.

remote-friendlyUTC+5:30 iGamingstreaminghigh-traffic
LOG