ArchitectureCustom Software Development
Multi-team Kafka architecture: turning one product's pipeline into a shared platform
A Kafka pipeline designed for one product line bakes that team's assumptions into topic names, event formats, partition keys and retention. When other product lines start publishing and consuming, those assumptions become shared risk. This page sets out the architecture decisions that turn one team's pipeline into a platform several teams can depend on, and a migration sequence that avoids downtime.
On this page
- Signs a pipeline has outgrown its first owner
- The layers of a shared event platform
- Three ways to lay out topics across product lines
- Schema compatibility modes and who upgrades first
- Partition keys decide what stays in order
- Isolating teams: ACLs, quotas or separate clusters
- Delivery guarantees, retries and dead-letter topics
- Retention, compaction and the ClickHouse sink
- Migrating to the shared platform without downtime
- Questions and answers
- Sources
Signs a pipeline has outgrown its first owner
The first product line usually owns everything: it named topics after its own services, defined event formats in its own code and picked partition keys to suit its own consumers. That holds until other teams arrive. Then a field rename in one service breaks a consumer nobody knew existed. One team's batch job saturates the brokers and delays everyone's events. Nobody can say who owns a topic, so nobody dares delete it. Analytics double-counts after a consumer restart.
Each symptom traces back to a missing platform decision about ownership, contracts, isolation or delivery semantics. The sections below take those decisions in turn. They assume Kafka stays, which is usually right once several consumers depend on replay and independent consumption.
Three ways to lay out topics across product lines
| Criterion | Topic per event type | Topic per product line | One shared topic with a type field |
|---|---|---|---|
| Ordering across event types | None between topics | Possible for one product's events on a key | Possible for all events on a key |
| Access control | Fine-grained, per event type | Clean per team, coarse inside it | All or nothing |
| Schema governance | One schema per topic; simplest | Several types per topic need per-type subjects | Many types in one topic; strict discipline needed |
| Consumer cost | Consumers read only what they need | Consumers filter out other types | Every consumer reads and filters everything |
| Topic count | Grows with event types | Grows with product lines | Stays small |
Most platforms settle on a topic per event type within a domain, with names such as billing.invoice-issued, and combine types in one topic only where strict ordering across them is required.
Schema compatibility modes and who upgrades first
A schema registry rejects a new version that breaks the topic's compatibility mode. Confluent Schema Registry, a common choice, defaults to BACKWARD.2
- BACKWARD
- Consumers on the new schema can read data written with the previous version, so consumers upgrade first, then producers.2
- FORWARD
- Consumers still on the previous version can read data written with the new schema, so producers upgrade first.2
- FULL
- Both directions hold between the new and previous versions, so producers and consumers can upgrade in either order.2
- Transitive variants
- BACKWARD_TRANSITIVE, FORWARD_TRANSITIVE and FULL_TRANSITIVE check against every earlier version, not only the latest, which matters for topics replayed from the start.2
- NONE
- Checks disabled. Reserve it for topics still in development that nobody else consumes.
Partition keys decide what stays in order
Kafka orders events only within a partition, so the key determines which events stay in sequence. Key by the entity whose history must be read in order: an account, an order, a device. Keying by tenant or product line looks tidy but concentrates traffic, and one large tenant becomes a hot partition that caps throughput for everything sharing it.
Choose partition counts with growth in mind. Adding partitions later changes which partition a key maps to, so one entity's events can briefly straddle two partitions and order-sensitive consumers must cope. Treat a partition increase on such topics as a planned migration with a cut-over point, not a configuration tweak.
Isolating teams: ACLs, quotas or separate clusters
Isolation costs operations effort and makes cross-team data access harder. Choose the lightest option that contains the risk.
- If
Teams share a risk profile and mainly need to stop accidental writes to each other's topics.
ThenOne cluster with prefixed ACLs per team and service principal.
Access control gives clear ownership without duplicating infrastructure.
- If
One team's batch loads or replays slow everyone else.
ThenClient quotas on produce and fetch bandwidth and on request rate.
Brokers delay responses to clients that exceed their quota, protecting the rest of the cluster.1
- If
A product line has different residency, retention or regulatory requirements.
ThenA separate cluster, replicating only the events others are allowed to see.
Policy boundaries are easier to prove when they are physical.
- If
A shared upgrade or incident is unacceptable for one critical product.
ThenA separate cluster for that product, with the same conventions applied to both.
Blast radius, not traffic volume, justifies the extra operating work.
Delivery guarantees, retries and dead-letter topics
Kafka's default is at-least-once delivery: events are not lost but may arrive more than once, so consumers must tolerate duplicates.1 For most consumers the answer is idempotence: store processed event identifiers, or make each write an upsert keyed by the event, so a redelivery changes nothing. The idempotent producer separately removes duplicates caused by producer retries.1
Exactly-once processing is available for read-process-write flows that stay inside Kafka, using transactions and consumers that read only committed records.1 When output lands in an external system, exactness depends on that system, typically by storing the consumer's offset in the same place and transaction as the output.1 Reserve this for flows where a duplicate has real cost, such as ledger postings.
Separate transient errors from bad data. Retry transient failures with backoff, through retry topics if waiting would block the partition. Send events that can never succeed to a dead-letter topic, with headers recording the error, source topic and offset, and build the replay tool before you need it.
Retention, compaction and the ClickHouse sink
Questions and answers
Should every product line get its own Kafka cluster?
Usually not. Separate clusters multiply operations work and make cross-team data access harder, and most of the isolation teams need comes from ACLs and quotas on a shared cluster. Separate clusters are justified by policy or blast radius: different residency or retention rules, regulatory separation, or a product so critical that a shared upgrade or incident is unacceptable.
When does exactly-once processing in Kafka justify its complexity?
When a duplicate has a real cost and the flow stays inside Kafka, such as a stream application updating balances from one topic into another. For most consumers, at-least-once delivery with idempotent processing is simpler and just as correct. Where output goes to an external database or service, exactness depends on writing the output and the consumer position together, which that system must support.
Which schema format should we use: Avro, Protobuf or JSON Schema?
All three work with common schema registries. Avro is compact and has well-understood evolution rules, which suits high-volume topics. Protobuf fits naturally where services already use it for APIs. JSON Schema is easiest for teams new to schemas but produces larger messages. Consistency matters more than the choice: enforce compatibility in CI and record the decision in an architecture decision record.
How do we stop a dead-letter topic becoming a graveyard?
Give each dead-letter topic an owning team, an alert on new arrivals and a slot in that team's regular review. Put enough in the headers, such as error, source topic, partition and offset, to diagnose an event without trawling logs. Provide a replay tool that republishes fixed events to the original topic, and track how long items wait. A dead-letter topic nobody watches is slow data loss.
Sources
- Apache Kafka documentation: Design (message delivery semantics, log compaction, quotas) — The Apache Software Foundation · checked 10 October 2026
- Schema Evolution and Compatibility for Schema Registry — Confluent · checked 10 October 2026