Fit is an inverse problem.

What you feel is the gap between your foot and the inside of the shoe. In e-commerce, neither surface is known.

Ask someone why a shoe was uncomfortable and they will not describe a measurement. They will describe a place. It rubbed at the heel. It pinched across the ball. My toes hit the end. The pain has an address.

That address is the whole problem, and it is worth being precise about what is actually happening there.

What a person experiences as fit is the relationship between two surfaces: the outside of their foot, and the inside of the shoe. Where those surfaces are separated, there is room. Where they touch, there is contact. Where the shoe would have to occupy space the foot is already using, there is pressure.

Comfort is the shape of that gap, mapped over the whole foot.

The quantity nobody measures

Now consider what is available at the moment of purchase online.

The shopper's foot: not measured. A number they believe about themselves, inherited from a shop fitting years ago or from whatever they last managed to keep.

The inside of the shoe: not measured either. It is a cavity. It is sealed. It has never been recorded in any database the retailer holds. What exists instead is a size label, a set of photographs, and a description of the outside.

So the one quantity that determines the outcome — the gap between two surfaces — is a quantity where neither surface is known.

Fit is not a hard prediction problem. It is an inverse problem: the thing that causes the outcome is never observed, and has to be reconstructed from evidence that only points at it indirectly.

This is a familiar shape of problem in other fields. A geophysicist does not see the rock layer; they see how a seismic wave came back, and reconstruct the layer from it. A radiologist does not see the tumour; they see attenuation, and reconstruct. In each case the interesting work is not the measurement — it is the reconstruction, and knowing how much to trust it.

Footwear has the same structure and has almost never treated it that way.

Four kinds of indirect evidence

If the inner surface cannot be observed, it has to be estimated. There are broadly four sources of evidence available, and they are not equal.

Each of these is weak on its own and each is weak in a different way. That is the useful part: their failures are not correlated. Evidence that fails independently can be combined into something stronger than any of its parts — which is an old idea in measurement science and a surprisingly new one in footwear.

Why correlation runs out of road

Most fit recommendation available today rests on the fourth source alone. What did people who resemble you keep? It is a reasonable approach and it works within its limits.

The limits are structural. A correlation model knows what happened; it does not know why. So it degrades exactly where you most need it: on a product that has not sold yet, on a brand new to the catalogue, on a shopper without much history, and on any foot far enough from the population average that there are few similar people to average over.

That last case deserves saying plainly. The people most poorly served by fit recommendation are the people whose feet are least typical — which is to say, the people who have the most trouble buying shoes. A method built on averaging is weakest precisely where the need is greatest.

Causal evidence does not have that failure mode. A geometric comparison between a foot and a form does not care how many people bought the shoe last month.

The number you should not trust

There is a last requirement, and it is the one most easily skipped.

If fit is an estimate reconstructed from indirect evidence, then any honest fit system has to report how uncertain that estimate is. A recommendation with no uncertainty attached is not a measurement. It is an assertion.

And the uncertainty is not constant. It depends on how good the evidence is for that particular shoe. A model where the last has been physically measured is a very different epistemic situation from one estimated by family resemblance to something similar — and a system that presents both with the same confident tone is misrepresenting one of them.

The industry has trained people to expect a single number and no error bar. That is a habit worth breaking, because the error bar is where the honesty lives.

Fit was never a lookup. It is a reconstruction — and a reconstruction is only as trustworthy as its willingness to say how wrong it might be.

← All insights