I got into streaming through factory IT: nine years at an automotive supplier, connecting production lines to enterprise systems. A stalled pipeline there did not page anyone - machines just stopped. Since 2022 I design central Kafka platforms for enterprise clients, and on the side I teach Kafka to engineering teams at Ultra Tendency Academy. I have the certificates, the training slides, and a healthy respect for what clusters get up to at 3 a.m.
Every cluster taught me the same lesson. Dashboards are great at telling you a broker is overloaded and useless at telling you what to do about it. Which partitions? Where should they go? What can move without hurting producers? Stock Kafka hands you kafka-reassign-partitions.sh and a JSON file of replica assignments you edit by hand. The generated plan balances partition counts, not load, the replication throttle needs babysitting until the last replica catches up, and nothing writes down what moved.
It is not that nothing exists. Cruise Control proved automated rebalancing works, and it deserves the respect it gets. But it brings its own metrics reporter onto every broker, demands real tuning and understanding before you can trust it, and I could rarely explain to a client why it moved what it moved. Enterprise basics like SSO and an audit trail are not part of the package. The rest of the Kafka ecosystem is full of good tools that each solve a slice. None covered my list.
That gap bugged me for years, because Kafka already knows the answers - it just never shows them in a form you can act on. So in 2025 I started Calinora and built the tool I kept telling clients should exist: point it at your bootstrap servers, see what is really going on, and when something has to move, get a plan you can defend - throttled, audited, reversible. That is Pilot. I hope it makes your on-call quieter. It made mine.
Julian Bergner, Founder of Calinora