What Running Kafka on VMs Taught Us About Systems Thinking
The article details the engineering challenges of managing Kafka clusters on virtual machines and the transition to Strimzi for better scalability. It emphasizes the importance of proactive architectural changes over reactive crisis management in infrastructure engineering.
Why it matters
It highlights the operational shift required for high-scale data pipelines, moving from manual, error-prone VM management to automated, containerized orchestration.
Most infrastructure stories get written after something breaks. This one is different. Nobody paged Celestina Amadi's team at 3 AM about Kafka. The VMs were working, yet she made the call to rebuild anyway because she could see where things were headed before they became a problem.
Celestina Amadi leads the Cloud Engineering team behind the infrastructure that powers Moniepoint's payment and savings products, building systems that move transaction and savings data reliably for millions of customers every day. She's also a Grafana Champion, a HashiCorp Ambassador, and an IBM Champion 2026.
We moved to Strimzi because we looked honestly at how we were running Kafka and made a deliberate call: this does not scale, and we can do better.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in