ArchitectureTechnology

Internal developer platform architecture: planes, golden paths and team model

An internal developer platform architecture is a set of planes, not a single tool: a developer portal and catalog on top, then delivery, resource provisioning, observability, and security and policy underneath, all reached through golden paths. This page sets out each plane, the choices inside it, how to decide what to build or buy, and how to tell whether the platform is actually helping the teams that use it.

Reviewed 8 min read

On this page
  1. A product for developers, not a pile of tools
  2. Five planes of a reference internal developer platform
  3. Building blocks inside the portal and delivery planes
  4. Guardrails in code: infrastructure, GitOps and policy
  5. Build, buy or adopt the developer portal
  6. Running the platform team as a product team
  7. Rolling an IDP out in phases
  8. Where internal developer platforms go wrong
  9. Questions and answers
  10. Sources

A product for developers, not a pile of tools

An internal developer platform (IDP) is the curated set of self-service capabilities that lets a product team go from an idea to a running, observable, compliant service without filing tickets with five other teams. The word that matters is curated. Most organizations already own a CI server, a cloud account structure, a secrets store and a monitoring tool; they do not have a platform until those pieces are wired together behind an interface a developer can use in a single sitting.

Two failure patterns are common. One is the portal-first project that installs a developer portal, populates a catalog once and then watches it decay because nothing depends on it. The other is the infrastructure-first project that builds a powerful Kubernetes estate nobody outside the platform team can operate. Both skip the question a product manager would ask first: which developer task is slow or risky today, and what would make it fast and safe?

ColdAI's technology practice describes platform engineering as building internal platforms with self-service infrastructure, golden paths and automated governance1. The architecture below is how those three ideas fit together.

Five planes of a reference internal developer platform

Each plane can be swapped independently. The portal plane is the only one most developers see directly.

Developer portal and catalog01Integration and delivery02Resource provisioning03Observability04Security and policy05
  1. Developer portal and catalog

    Service catalog, ownership metadata, software templates, documentation and scorecards in one interface.

  2. Integration and delivery

    CI pipelines, artifact registries and GitOps controllers that promote changes between environments.

  3. Resource provisioning

    Infrastructure as code modules and APIs for databases, queues, clusters and cloud accounts.

  4. Observability

    Telemetry collection, dashboards, alerting and service-level objectives wired in by default.

  5. Security and policy

    Identity, secrets, supply-chain checks and policy-as-code applied across every other plane.

Conceptual layering of an internal developer platform, adapted from common platform-engineering reference models. It is not a product diagram or a measured deployment.

Building blocks inside the portal and delivery planes

These terms come up in every platform design review. Agreeing what each means early prevents a lot of argument later.

Golden path
An opinionated, supported route for a common job, such as creating a new HTTP service, that bundles a template, pipeline, infrastructure and default telemetry. Teams may leave it, but then they own what they build.
Software template
A scaffold that creates a repository, pipeline and catalog entry in one action. In Backstage, an open-source portal framework hosted by the CNCF, this is the Scaffolder feature23.
Service catalog
The register of every component, API and resource, each with an owning team, lifecycle stage, dependencies and links to its runbooks and dashboards.
Scorecard
A set of automated checks per service, such as an on-call rotation, a defined SLO, no critical vulnerabilities and a recent dependency update, shown as production-readiness status.
Thinnest viable platform
The Team Topologies idea that a platform should be the smallest set of capabilities that speeds up stream-aligned teams; sometimes a wiki page and a few scripts are enough4.
Platform API
The contract developers use to request resources, often a declarative file in the repository that the platform reconciles, rather than a ticket or a console click.

Guardrails in code: infrastructure, GitOps and policy

The resource plane works best when every environment is described in version-controlled code. Infrastructure as code modules encode the organization's defaults, such as encryption, tagging and network boundaries, so a developer asking for a database gets a compliant one without reading a standard. Keep modules small and versioned; a single module that provisions an entire environment becomes impossible to change safely.

GitOps extends this to deployment. The OpenGitOps principles describe a system whose desired state is declarative, versioned and immutable, pulled automatically by agents and continuously reconciled5. In practice that means a controller in each cluster or account applies what is in the repository, and drift is corrected or flagged rather than discovered during an incident.

Policy-as-code is the security plane's main tool. Engines such as Open Policy Agent evaluate rules written as code against deployment manifests, infrastructure plans and API requests6. Run the same policies in the pipeline, where they give developers fast feedback, and at admission time, where they stop anything that bypassed the pipeline. Observability belongs on the golden path too: emitting traces, metrics and logs through a vendor-neutral standard such as OpenTelemetry keeps the backend replaceable7.

Build, buy or adopt the developer portal

The portal is the plane most often bought or adopted. The planes beneath it are usually assembled from existing cloud and open-source tooling.

FactorAdopt open source (for example Backstage)Buy a commercial portalBuild in-house
Time to a usable catalogModerate: needs hosting, plugins and front-end skillsFastest: hosted, with integrations readySlowest: every feature is new code
Fit to unusual workflowsHigh, through custom pluginsLimited to the vendor's extension modelComplete, at the cost of maintenance
Ongoing engineering loadReal: upgrades and plugin maintenance need an ownerLow for the portal itselfHighest, and it never ends
Lock-in and exitLow; catalog descriptors stay in your repositoriesDepends on data export and descriptor formatsNone external, but knowledge sits with a few people
Best suited toLarger engineering groups with a funded platform teamGroups that want results without a front-end teamRarely justified unless needs are truly unusual

The highlighted column is a sensible default for a funded platform team, not a universal answer. A small organization may get more from a commercial portal or no portal at all.

Running the platform team as a product team

Team Topologies describes a platform team whose purpose is to reduce the cognitive load on stream-aligned teams, delivering internal services they consume on a self-service basis4. That requires a product owner, a published roadmap, user research with developers and a willingness to retire features nobody uses.

Measure outcomes the way a product team would. The DORA research programme (DevOps Research and Assessment, not the EU Digital Operational Resilience Act) publishes software-delivery metrics covering throughput and instability, such as change lead time, deployment frequency, change failure rate and failed deployment recovery time8. Pair them with a short, regular developer-experience survey and adoption data from the catalog: how many services are on a golden path, and how many teams left it and why.

Rolling an IDP out in phases

  1. Interview developers and map one journey

    Pick the journey that hurts most, often creating a new service or getting a change to production, and time every hand-off and ticket in it.

    Output
    Journey map with waiting points
    Owner
    Platform product owner
  2. Ship one golden path end to end

    Cover template, pipeline, infrastructure, telemetry and catalog registration for that single service type, and move two willing teams onto it.

    Output
    Working golden path and two pilot services
    Owner
    Platform engineers
  3. Populate the catalog from code

    Generate catalog entries from repository descriptors so ownership stays current, then add the first scorecard checks that teams agree are fair.

    Output
    Catalog with owners and baseline scorecards
    Owner
    Platform team with service owners
  4. Add policy and widen the path

    Move recurring review findings into policy-as-code, then add a second golden path for the next most common workload, such as event consumers or data jobs.

    Output
    Policy library and a second path
    Owner
    Security and platform teams
  5. Review adoption and prune

    Each quarter, compare delivery metrics, survey results and path adoption, and remove capabilities that add support load without users.

    Output
    Revised platform roadmap
    Owner
    Platform product owner

Where internal developer platforms go wrong

Mandated adoption of an immature path

Early signalTeams file exceptions faster than the platform team can answer them.

MitigationMake paths attractive rather than compulsory until they handle the common cases well, and track why teams leave them.

A catalog nobody maintains

Early signalOwners listed in the portal have left, and on-call pages go to the wrong team.

MitigationGenerate entries from code, fail builds when descriptors are missing and show stale ownership on scorecards.

Abstractions that hide too much

Early signalDevelopers cannot debug production because the platform conceals logs, configuration or network paths.

MitigationExpose the underlying resources read-only and document the escape hatch for advanced users.

Platform team as a ticket queue

Early signalRequests arrive as tickets and the backlog grows each sprint.

MitigationTurn the most frequent requests into self-service APIs and stop accepting tickets for anything already automated.

Questions and answers

How large does an engineering organization need to be before an internal developer platform pays off?

There is no fixed threshold. The trigger is repeated, avoidable work: several teams solving the same pipeline, infrastructure and compliance problems separately. A handful of teams can often manage with shared templates and good documentation, which is the thinnest viable platform. A dedicated platform team and portal make sense once coordination costs and inconsistent production readiness are visibly slowing delivery.

When should a company not build an internal developer platform?

Skip a full platform when one or two teams own everything, when the product is a single application on a managed platform-as-a-service, or when there is nobody to own the platform as a product after launch. In those cases, invest in a good pipeline template, infrastructure as code and clear runbooks. A platform without an owner becomes another legacy system.

How do you move existing teams onto golden paths without a rewrite?

Start with new services, which adopt the path from day one. For existing services, offer incremental steps: register in the catalog, adopt the shared pipeline, then the telemetry defaults, then the infrastructure modules. Scorecards show each team where it stands. Retire old patterns only when the path supports their real needs, which may require extending it first.

Is Backstage required for an internal developer platform?

No. Backstage is a widely used open-source framework for the portal plane, but the portal is only one of five planes, and many effective platforms start with a command-line tool, templates in a repository and documentation. Choose a portal when catalog, ownership and discoverability problems justify the hosting and plugin maintenance it brings.

Sources

  1. Technology capability — ColdAI
  2. What is Backstage? — Backstage · checked 10 October 2026
  3. Backstage project page — Cloud Native Computing Foundation · checked 10 October 2026
  4. Key concepts — Team Topologies · checked 10 October 2026
  5. OpenGitOps principles — OpenGitOps (CNCF) · checked 10 October 2026
  6. Open Policy Agent documentation — Open Policy Agent · checked 10 October 2026
  7. What is OpenTelemetry? — OpenTelemetry · checked 10 October 2026
  8. DORA's software delivery metrics — DORA · checked 10 October 2026

More in Technology

Back to Technology

Next step

Have one developer journey reviewed against this architecture

Send a short description of how a new service reaches production today and which tools are involved. We will reply with the plane that most needs attention and a suggested first golden path.

Share your current journey