GuideOperations

Building a spend cube: taxonomy, supplier normalization and classification

A spend cube organizes every purchase by what was bought, from which supplier and by which part of the business. Building one is mostly data work: consolidating sources, normalizing supplier names, choosing a taxonomy and classifying transactions accurately enough to act on. This guide covers each step, the taxonomy decision, where machine-learning and LLM classifiers fit, and how to turn the cube into sourcing waves.

Reviewed 7 min read

On this page
  1. Three questions a spend cube answers
  2. Building the cube from raw transactions
  3. UNSPSC, ECLASS or a custom taxonomy
  4. Rules, machine learning or an LLM classifier
  5. Quality checks before anyone trusts the numbers
  6. A hypothetical first cube at a multi-site manufacturer
  7. Turning the cube into sourcing waves
  8. Questions and answers
  9. Sources

Three questions a spend cube answers

A spend cube is a dataset with three main axes: the category of what was bought, the supplier it came from and the business unit or cost center that bought it. A fourth view, the channel, separates spend through contracts and catalogs from off-contract and card spend. Together they answer the questions a category strategy starts from: where the money goes, how concentrated it is and how much of it follows agreed terms.

Spend analytics is one element of ColdAI's procurement transformation offering, alongside AI-powered sourcing, supplier risk management and contract optimization1. The steps below describe how a cube is built so that its numbers survive scrutiny from finance and from category managers.

Building the cube from raw transactions

  1. Consolidate the sources

    Pull accounts payable invoices, purchase orders, purchasing-card and travel and expense data, plus contract metadata. AP is the backbone because it shows what was actually paid; orders add item descriptions and requesters; contracts show which spend should be on terms. Agree the period, currency conversion and treatment of tax and intercompany lines up front.

    Output
    Consolidated transaction set
    Owner
    Procurement analytics with finance
  2. Normalize supplier names

    The same supplier appears under many spellings, legal entities and remit-to addresses. Cleanse names, match on tax IDs, addresses and bank details, and roll entities up into a parent-child hierarchy using a business registry or data provider. Skip this and concentration, and with it your negotiating position, is understated.

    Output
    Supplier master with parent mapping
  3. Choose the taxonomy and its owner

    Select a standard or custom taxonomy, decide how many levels you need and name an owner who approves changes. Go deep enough to separate categories that are sourced differently, and no deeper than you can classify reliably.

    Output
    Approved taxonomy with definitions
    Owner
    Chief procurement officer or delegate
  4. Classify transactions

    Use rules for the obvious cases, such as single-category suppliers or clean GL mappings, and a classifier for the rest that reads description, supplier and GL account together. Each prediction carries a confidence score; low-confidence lines go to category experts, and their corrections become training data or new rules.

    Output
    Classified lines with confidence scores
  5. Test quality

    Run the checks listed below, sampling by spend value. Fix systematic errors at the rule or model level rather than line by line.

    Output
    Quality report
  6. Build the views

    Publish category, business unit and supplier views, supplier concentration per category and off-contract spend by requester and channel. Keep drill-down to invoice level so category managers can test surprising numbers themselves.

    Output
    Spend dashboards and extracts
    Owner
    Analytics team
  7. Set refresh and governance

    Refresh on a fixed cycle, classify new lines each time, work the low-confidence queue and log taxonomy change requests. A cube rebuilt from scratch every year loses the corrections that made it accurate.

    Output
    Refresh calendar and change log

UNSPSC, ECLASS or a custom taxonomy

CriterionUNSPSCECLASSCustom taxonomy
StructureFour levels (segment, family, class, commodity) in an eight-digit code, with an optional fifth level2Four levels in an eight-digit code, two digits per level, with the lowest level described by properties3Whatever the organization defines, usually three or four levels
Strength for spend analysisBroad cross-industry coverage of goods and services; common in e-procurement catalogsDetailed product properties, strong in industrial and technical goodsMirrors how category teams are organized and how sourcing decisions are made
WeaknessProduct-oriented levels do not always match sourcing categories; some services are thinMore product detail than spend analysis needs; releases to manageNo external comparability; drifts without disciplined ownership
Best fitOrganizations exchanging catalogs with suppliers or public buyersManufacturers and distributors that also need product master dataMost spend programs, often with a mapping to a standard
MaintenanceTrack periodic releases and remapAdopt new releases and map between versionsInternal change control by the taxonomy owner

Many programs classify to a custom taxonomy for decisions and keep a mapping to UNSPSC or ECLASS for catalog and supplier data exchange.

Rules, machine learning or an LLM classifier

Most cubes use all three, each where it is strongest.

  • If

    A supplier sells only one kind of thing, or the GL account maps cleanly to a category.

    Then

    Classify with rules.

    Rules are transparent, instant and easy for auditors to follow.

  • If

    There is a history of classified lines and descriptions follow familiar patterns.

    Then

    Train a supervised classifier on description, supplier and GL fields.

    It handles the large middle cheaply and improves with every correction.

  • If

    Descriptions are free text, multilingual or new, with little labelled history.

    Then

    Use an LLM classifier prompted with the taxonomy definitions, logging its rationale for each line.

    It works from definitions rather than examples, which helps on the long tail.

  • If

    A prediction falls below the confidence threshold or lands in a sensitive category.

    Then

    Route it to a category expert and feed the decision back.

    Thresholds keep human effort on the lines where it changes the answer.

Quality checks before anyone trusts the numbers

0 of 6 checked

A hypothetical first cube at a multi-site manufacturer

Turning the cube into sourcing waves

A cube earns its keep when it changes what procurement does next. Rank categories by addressable spend, supplier fragmentation and contract coverage, then group them into waves the team can run in sequence. Fragmented categories with clear specifications are usually early candidates for competitive sourcing; categories driven by internal demand, such as travel or IT peripherals, often respond better to policy and demand management than to tendering.

For engineered parts where few suppliers can compete, the cube shows where the money is but not what it should cost; that is the point to build bottom-up should-cost models. Supplier concentration surfaced by the cube also feeds risk work such as supplier risk monitoring.

Questions and answers

How accurate does spend classification need to be?

Accurate enough for the decision at hand. For prioritizing sourcing waves, high accuracy by value in the largest categories matters far more than perfection in the tail. Measure accuracy on a value-weighted sample at the taxonomy level you act on, and raise the bar for categories about to go to tender, where one misclassified supplier can change the scope of an event.

Can an LLM classify our spend without training data?

It can produce a usable first pass from taxonomy definitions and transaction descriptions, which helps when labelled history is thin. It still needs confidence thresholds, expert review of uncertain lines and consistency checks, because a language model can classify the same description differently across runs. Logging the rationale for each line makes review faster and errors easier to correct.

Which data source should a spend cube start from?

Accounts payable, because it records what was actually paid to every supplier and ties to the general ledger. Purchase orders, card and expense data and contracts add detail but each covers only part of spend. Starting from AP and enriching it avoids a cube that looks complete for purchase-order spend while missing everything bought outside the procure-to-pay system.

How often should the spend cube be refreshed?

Monthly or quarterly for most organizations, classifying new transactions each time and working the low-confidence queue in between. Match the cycle to how the cube is used: monthly if category managers track contract compliance, quarterly if it mainly feeds the sourcing plan. An annual rebuild throws away corrections and lets supplier hierarchies drift.

Sources

  1. Operations capability: procurement transformation offering — ColdAI
  2. UNSPSC: United Nations Standard Products and Services Code — Wikipedia · checked 10 October 2026
  3. ECLASS technical specification: classification class — ECLASS e.V. · checked 10 October 2026

More in Operations

Back to Operations

Next step

Tell us what your spend data looks like today

Describe your data sources, the taxonomy you use, if any, and the sourcing decision the cube should support. We will reply with a view on where to start and what a first cube would need.

Discuss a spend analysis build