ArchitectureCustom Software Development

Multi-team Kafka architecture: turning one product's pipeline into a shared platform

A Kafka pipeline designed for one product line bakes that team's assumptions into topic names, event formats, partition keys and retention. When other product lines start publishing and consuming, those assumptions become shared risk. This page sets out the architecture decisions that turn one team's pipeline into a platform several teams can depend on, and a migration sequence that avoids downtime.

Reviewed 7 min read

On this page
  1. Signs a pipeline has outgrown its first owner
  2. The layers of a shared event platform
  3. Three ways to lay out topics across product lines
  4. Schema compatibility modes and who upgrades first
  5. Partition keys decide what stays in order
  6. Isolating teams: ACLs, quotas or separate clusters
  7. Delivery guarantees, retries and dead-letter topics
  8. Retention, compaction and the ClickHouse sink
  9. Migrating to the shared platform without downtime
  10. Questions and answers
  11. Sources

Signs a pipeline has outgrown its first owner

The first product line usually owns everything: it named topics after its own services, defined event formats in its own code and picked partition keys to suit its own consumers. That holds until other teams arrive. Then a field rename in one service breaks a consumer nobody knew existed. One team's batch job saturates the brokers and delays everyone's events. Nobody can say who owns a topic, so nobody dares delete it. Analytics double-counts after a consumer restart.

Each symptom traces back to a missing platform decision about ownership, contracts, isolation or delivery semantics. The sections below take those decisions in turn. They assume Kafka stays, which is usually right once several consumers depend on replay and independent consumption.

The layers of a shared event platform

Product teams sit at the top, producing and consuming; the layers beneath them are what a platform team owns.

Producing teams01Schema registry02Topics and ownership03Access and quotas04Consumers and sinks05Operations06
  1. Producing teams

    Each product team publishes the events it owns, using a shared client configuration with idempotence enabled.

  2. Schema registry

    Each topic's schema is registered with a compatibility mode and checked in CI before deployment.

  3. Topics and ownership

    Domain-based names and a catalogue entry per topic recording owner, retention and data classification.

  4. Access and quotas

    Prefixed ACLs per team and client quotas on bandwidth and request rate.

  5. Consumers and sinks

    Independent consumer groups, retry and dead-letter topics, and sinks into analytics stores such as ClickHouse.

  6. Operations

    Lag, error and quota monitoring per team, with replay tooling and runbooks.

Conceptual layering of a multi-team Kafka platform. Real deployments often add stream processing and connectors between layers.

Three ways to lay out topics across product lines

CriterionTopic per event typeTopic per product lineOne shared topic with a type field
Ordering across event typesNone between topicsPossible for one product's events on a keyPossible for all events on a key
Access controlFine-grained, per event typeClean per team, coarse inside itAll or nothing
Schema governanceOne schema per topic; simplestSeveral types per topic need per-type subjectsMany types in one topic; strict discipline needed
Consumer costConsumers read only what they needConsumers filter out other typesEvery consumer reads and filters everything
Topic countGrows with event typesGrows with product linesStays small

Most platforms settle on a topic per event type within a domain, with names such as billing.invoice-issued, and combine types in one topic only where strict ordering across them is required.

Schema compatibility modes and who upgrades first

A schema registry rejects a new version that breaks the topic's compatibility mode. Confluent Schema Registry, a common choice, defaults to BACKWARD.2

BACKWARD
Consumers on the new schema can read data written with the previous version, so consumers upgrade first, then producers.2
FORWARD
Consumers still on the previous version can read data written with the new schema, so producers upgrade first.2
FULL
Both directions hold between the new and previous versions, so producers and consumers can upgrade in either order.2
Transitive variants
BACKWARD_TRANSITIVE, FORWARD_TRANSITIVE and FULL_TRANSITIVE check against every earlier version, not only the latest, which matters for topics replayed from the start.2
NONE
Checks disabled. Reserve it for topics still in development that nobody else consumes.

Partition keys decide what stays in order

Kafka orders events only within a partition, so the key determines which events stay in sequence. Key by the entity whose history must be read in order: an account, an order, a device. Keying by tenant or product line looks tidy but concentrates traffic, and one large tenant becomes a hot partition that caps throughput for everything sharing it.

Choose partition counts with growth in mind. Adding partitions later changes which partition a key maps to, so one entity's events can briefly straddle two partitions and order-sensitive consumers must cope. Treat a partition increase on such topics as a planned migration with a cut-over point, not a configuration tweak.

Isolating teams: ACLs, quotas or separate clusters

Isolation costs operations effort and makes cross-team data access harder. Choose the lightest option that contains the risk.

  • If

    Teams share a risk profile and mainly need to stop accidental writes to each other's topics.

    Then

    One cluster with prefixed ACLs per team and service principal.

    Access control gives clear ownership without duplicating infrastructure.

  • If

    One team's batch loads or replays slow everyone else.

    Then

    Client quotas on produce and fetch bandwidth and on request rate.

    Brokers delay responses to clients that exceed their quota, protecting the rest of the cluster.1

  • If

    A product line has different residency, retention or regulatory requirements.

    Then

    A separate cluster, replicating only the events others are allowed to see.

    Policy boundaries are easier to prove when they are physical.

  • If

    A shared upgrade or incident is unacceptable for one critical product.

    Then

    A separate cluster for that product, with the same conventions applied to both.

    Blast radius, not traffic volume, justifies the extra operating work.

Delivery guarantees, retries and dead-letter topics

Kafka's default is at-least-once delivery: events are not lost but may arrive more than once, so consumers must tolerate duplicates.1 For most consumers the answer is idempotence: store processed event identifiers, or make each write an upsert keyed by the event, so a redelivery changes nothing. The idempotent producer separately removes duplicates caused by producer retries.1

Exactly-once processing is available for read-process-write flows that stay inside Kafka, using transactions and consumers that read only committed records.1 When output lands in an external system, exactness depends on that system, typically by storing the consumer's offset in the same place and transaction as the output.1 Reserve this for flows where a duplicate has real cost, such as ledger postings.

Separate transient errors from bad data. Retry transient failures with backoff, through retry topics if waiting would block the partition. Send events that can never succeed to a dead-letter topic, with headers recording the error, source topic and offset, and build the replay tool before you need it.

Retention, compaction and the ClickHouse sink

Migrating to the shared platform without downtime

  1. Inventory what exists

    List every topic, producer, consumer group, schema and retention setting, and find an owner for each. Topics with no owner and no consumers are retirement candidates.

  2. Register today's formats as they are

    Register current event formats in the schema registry unchanged, with a compatibility mode. Accidental breaking changes stop before anything is improved.

  3. Create new topics alongside the old

    Create topics under the new naming convention, sized and retained for growth, and mirror or dual-publish events into them.

  4. Move consumers group by group

    Point each consumer group at the new topic from a known position, so nothing is skipped and idempotent processing absorbs any overlap.

  5. Switch producers

    Once every consumer reads the new topic, producers publish there directly and mirroring stops.

  6. Retire the old topics

    Keep old topics read-only for an agreed period, then delete them and update the catalogue.

Questions and answers

Should every product line get its own Kafka cluster?

Usually not. Separate clusters multiply operations work and make cross-team data access harder, and most of the isolation teams need comes from ACLs and quotas on a shared cluster. Separate clusters are justified by policy or blast radius: different residency or retention rules, regulatory separation, or a product so critical that a shared upgrade or incident is unacceptable.

When does exactly-once processing in Kafka justify its complexity?

When a duplicate has a real cost and the flow stays inside Kafka, such as a stream application updating balances from one topic into another. For most consumers, at-least-once delivery with idempotent processing is simpler and just as correct. Where output goes to an external database or service, exactness depends on writing the output and the consumer position together, which that system must support.

Which schema format should we use: Avro, Protobuf or JSON Schema?

All three work with common schema registries. Avro is compact and has well-understood evolution rules, which suits high-volume topics. Protobuf fits naturally where services already use it for APIs. JSON Schema is easiest for teams new to schemas but produces larger messages. Consistency matters more than the choice: enforce compatibility in CI and record the decision in an architecture decision record.

How do we stop a dead-letter topic becoming a graveyard?

Give each dead-letter topic an owning team, an alert on new arrivals and a slot in that team's regular review. Put enough in the headers, such as error, source topic, partition and offset, to diagnose an event without trawling logs. Provide a replay tool that republishes fixed events to the original topic, and track how long items wait. A dead-letter topic nobody watches is slow data loss.

Sources

  1. Apache Kafka documentation: Design (message delivery semantics, log compaction, quotas) — The Apache Software Foundation · checked 10 October 2026
  2. Schema Evolution and Compatibility for Schema Registry — Confluent · checked 10 October 2026

More in Custom Software Development

Back to Custom Software Development

Next step

Send us the topic list your new product lines will share

Outline your topics, consumer groups and the product lines joining the pipeline. We will reply with the platform decisions that look most urgent and how a migration could be sequenced around your releases.

Discuss your event platform