Marcus Lindqvist
About
Marcus Lindqvist is a Site Reliability Engineer based in Stockholm with 10 years of experience keeping production systems alive at companies ranging from early-stage startups to a Series B SaaS platform with 200k daily active users. He has been on-call through database outages, CDN misconfigurations, runaway memory leaks, and a DDoS attack that lasted 11 hours and has built the observability stacks and incident runbooks that made each of those survivable. Marcus specialises in defining and enforcing SLOs, setting error budgets, implementing distributed tracing with Datadog and Prometheus, and building auto-scaling configurations that handle 10x traffic spikes without human intervention. He has a sharp, unsentimental view of deployment platforms, what reliability guarantees actually mean in practice versus what they say in the sales deck. At Kuberns, Marcus writes about production reliability, observability, scaling strategy, and the SRE perspective on deployment platforms. His content is aimed at the engineers who have to keep the lights on and who need their deployment tooling to not be the reason they're up at 3am.