Deep diveMetals & Mining
How machine-learning mineral prospectivity models are built and validated
A prospectivity model ranks ground by how closely its evidence resembles the footprint of a target deposit type, so an exploration team can spend a limited drilling budget where the geology argues hardest. Building one worth trusting means translating a mineral systems model into mappable evidence, choosing a method that suits the labels you actually have, and validating it in a way that spatial data cannot flatter. Its output guides targeting; it is not a resource.
On this page
- Target ranking and drill budget allocation
- From a mineral systems model to ranked drill targets
- Terms that come up in prospectivity work
- Preparing evidence layers before any modelling starts
- Knowledge-driven and data-driven methods side by side
- Scarce labels and the negative-sampling problem
- Validating a prospectivity model so it cannot flatter itself
- Where a prospectivity map stops and public reporting begins
- Ranking porphyry copper targets across a hypothetical district
- Questions and answers
- Sources
Target ranking and drill budget allocation
Exploration is a sequence of spending decisions made under uncertainty: which licences to keep, which areas to map in detail, which anomalies deserve drill holes. A prospectivity model supports those decisions by scoring every cell of a grid on how strongly the available evidence matches what a deposit of the target style would leave behind. The useful output is a ranking, ideally with uncertainty attached, that geologists can interrogate and argue with.
The model does not find deposits; it concentrates attention. It earns its place if, in ground it has not seen, it would have led the team to the known occurrences while flagging only a small share of the area, and if the geologists can say in plain terms why a given cell scores high.
From a mineral systems model to ranked drill targets
- Deposit model
The ore-forming process for the commodity and deposit style, such as porphyry copper.
- Critical processes
Source, fluid pathway, trap and preservation: each a necessary ingredient of the system.
- Mappable proxies
Observable features standing in for each process, such as intrusive contacts, faults or alteration.
- Evidence layers
Gridded maps of each proxy built from geology, geophysics, geochemistry and remote sensing.
- Model and validate
Combine the layers with a knowledge-driven or data-driven method and test it by area.
- Ranked targets
A shortlist with scores, uncertainty and the evidence behind each, ready for field checks.
Terms that come up in prospectivity work
- Mineral systems approach
- Treating a deposit as the outcome of geological processes acting across scales, then targeting the evidence each process leaves rather than the deposit alone1.
- Targeting criterion (mappable proxy)
- A feature that can be mapped across the whole study area and represents one critical process, for example proximity to a fault set that could have channelled fluids.
- Evidence layer
- One proxy rendered as a grid at a shared resolution and projection, carrying its own uncertainty and source scale.
- Positive and negative labels
- Cells known to host an occurrence of the target type, and cells assumed barren. The second set is usually inferred rather than observed.
- Spatial autocorrelation
- The tendency of neighbouring cells to resemble one another, which lets a model score well by memorising neighbourhoods instead of learning geology.
- Prediction-area plot
- A chart comparing the share of known occurrences captured with the share of ground flagged, used to judge how efficiently a map concentrates targets2.
Preparing evidence layers before any modelling starts
Knowledge-driven and data-driven methods side by side
| Criterion | Fuzzy logic or index overlay | Weights of evidence | Random forest or gradient boosting | Neural networks |
|---|---|---|---|---|
| What sets the weights | Expert judgement on each proxy | Statistics from known occurrences, layer by layer | Learned from labels, including interactions between layers | Learned from labels, including spatial patterns |
| Labels needed | None, which suits frontier ground | A modest set of occurrences | More occurrences plus credible negatives | The most; risky in sparse districts |
| Correlated layers | Handled only as well as the expert handles them | Poorly, because it assumes conditional independence | Reasonably well | Well, given enough data |
| Explainability | Fully transparent | Transparent per layer | Good with feature-attribution methods | Hardest to explain to a geologist |
| Typical failure | Encodes the expert's blind spots | Inflated scores from dependent layers | Memorises clustered training areas | Overfits small datasets |
| Where it fits | Greenfield belts with few known deposits | Mature districts with a clear deposit model | Districts with enough occurrences and even coverage | Large, dense datasets such as continental grids |
Hybrid workflows are common: expert-built proxies feed a data-driven model, and an expert overlay acts as a sense check on its ranking.
Scarce labels and the negative-sampling problem
A district may contain only a handful of known occurrences of the target style, often clustered where outcrop is good or where earlier explorers happened to look. Data-driven methods also need negatives, and genuinely barren ground is rarely known. Teams typically sample negatives at random from areas far from known occurrences, sample them from areas an expert rates as unfavourable, or use positive-unlabelled learning, which treats everything without a label as uncertain.
Every choice biases the result. Random negatives can include undiscovered deposits; expert-filtered negatives import the expert's assumptions; clustered positives teach the model what well-explored ground looks like rather than what mineralised ground looks like. Test how much the ranking moves under different negative strategies and report that sensitivity, instead of presenting one run as the answer.
Validating a prospectivity model so it cannot flatter itself
Split by area, not at random
Divide the study area into spatial blocks or hold out whole sub-districts, so test cells never sit next to training cells. Random splits on autocorrelated data report accuracy the model will not achieve on new ground3.
Hold some occurrences out of everything
Keep a few known deposits out of every stage, including the choice of layers, and check where they land in the final ranking.
Plot success-rate and prediction-area curves
Show what share of held-out occurrences fall inside the top-ranked share of the area. A model that needs most of the map to capture them adds little2.
Beat a simple baseline
Benchmark against distance to known mineralisation or an expert overlay. A complex model has to outperform the obvious answer to justify itself.
Map uncertainty, not only the score
Rerun the model across resampled labels and parameter choices, then show where the ranking is stable and where it flips between runs.
Where a prospectivity map stops and public reporting begins
Ranking porphyry copper targets across a hypothetical district
Questions and answers
How much data is enough for a machine-learning prospectivity model?
There is no fixed threshold. What matters is whether the known occurrences are numerous and spread out enough to survive a spatial hold-out, and whether every evidence layer covers the whole area evenly. With only a few clustered occurrences, a knowledge-driven or weights-of-evidence approach is usually more honest than a complex learner, and data-driven methods can be added as drilling produces more labels.
Do prospectivity models transfer from one region to another?
Only with care. A model trained in one belt learns that belt's data coverage, mapping conventions and erosion level as much as its geology. The mineral systems logic and the choice of proxies usually transfer well; the trained weights often do not. Retrain or recalibrate on local data, and validate on held-out local occurrences before using a transferred model to rank ground.
How should drilling data be managed for prospectivity work?
Keep collars, surveys, lithology logs and assays in one versioned database with original document references, quality-control results and the date each record was added. Record holes that found nothing as carefully as those that hit mineralisation. Freeze the exact extract used for each model run, so a ranking can be reproduced and audited when it is challenged by management or a partner.
Should barren drill holes be used as negative labels?
They are among the best negatives available, because they are observed rather than assumed. Use them with their limits in mind: a hole tests a small volume, may have stopped short of the target, or may have been drilled for a different deposit style. Treat them as strong but local evidence, and combine them with a sampling strategy for the large areas no one has drilled.
Sources
- Translating the mineral systems approach into an effective exploration targeting system (McCuaig, Beresford and Hronsky), Ore Geology Reviews — Elsevier · checked 10 October 2026
- Prediction-area (P-A) plot and C-A fractal analysis to classify and evaluate evidential maps for mineral prospectivity modeling (Yousefi and Carranza), Computers & Geosciences — Elsevier · checked 10 October 2026
- Cross-validation strategies for data with temporal, spatial, hierarchical, or phylogenetic structure (Roberts et al.), Ecography — Wiley · checked 10 October 2026
- JORC Code review: update and timeline — Joint Ore Reserves Committee · checked 10 October 2026
- NI 43-101 Standards of Disclosure for Mineral Projects — Ontario Securities Commission · checked 10 October 2026
- CRIRSCO International Reporting Template and member reporting standards — Committee for Mineral Reserves International Reporting Standards · checked 10 October 2026