What belongs in a supplier risk score
Most risk scores are one-dimensional: delivery performance. Real risk comes from three places, and all three need measuring at once.
The supplier risk score is one of the most promised and least used features in purchasing software. The reason is simple: most scores are one-dimensional, and that dimension is usually delivery performance. “94% on-time delivery” may be accurate, but it does not tell you what to do about the next order.
A usable score needs three distinct risks separated out.
First risk: delivery
The best-known dimension, and the easiest to measure. Even here there is a common mistake: looking at average lateness.
Two suppliers averaging three days late can be very different. One is three days late every time — that is predictable and can be absorbed into planning. The other is usually on time but once a month is fifteen days late — that is the one that stops production.
What needs measuring is not the average but the tail of the distribution: what happens in the worst decile?
Second risk: quality
The second dimension is acceptance rate and non-conformance history for incoming goods. This data usually sits in the quality module and never reaches purchasing.
Joining the two is a gain on its own. A supplier who delivers on time but has eight percent of its lots returned looks good on a delivery score and stops looking good on a combined one.
Third risk: commercial behaviour
The third dimension is the least measured and often the earliest signal: price and quotation behaviour. A price departing from its own trend, a quote range narrowing abruptly, a lead time stretching quietly — none is a problem alone, but together they form a pattern.
Anomaly detection is useful here. The aim is not to penalise the supplier; it is to get the buyer on the phone sooner.
The score is not the decision
Once the three dimensions are combined, collapsing them into one number is tempting. It has a cost: a single number hides which dimension is bad. A supplier weak on delivery does not warrant the same action as one weak on quality.
We prefer showing the three separately and giving the recommendation with its reasoning. “Consider an alternative supplier for this line — the delivery tail deteriorated over the last six months” is far more useful than “risk score 62.”
Where to start
Building this score requires three modules to be talking: purchasing holds quote and order history, quality holds acceptance data, and inventory holds the actual delivery dates. As long as the three live in separate systems the score does not get built — and if it does, nobody trusts it.
Which means the work here is integration work, not modelling work. The model is the easy part, and it comes last.
Topics
- procurement
- artificial intelligence