The Missing Slice

Derek · part 8 of 8

Unreasonable Future Claims

On a Sentence Experian Cannot Possibly Mean, and an Industry That Repeated It Without Blinking

Every so often an industry says the quiet part out loud, and it turns out the quiet part was not a secret but a fiction. Derek has been paying for this one since before the meerkat.

· Insurance & the Consumer · 3,174 words, about 15 minutes

I have been staring at one sentence of marketing copy for a fortnight, in the manner of a man who suspects he has been sold a horse.

Experian markets its Delphi Generation 11 scorecard to insurers on a claimed fourteen percent improvement over its predecessor at identifying good customers. It defines a good customer, in its own words, as one of “those who don’t make unreasonable future claims or fail to make direct-debit payments.”

I do not believe it.

I want to be precise about which part I disbelieve, because I am not one of those people who thinks credit scoring in insurance is a racket, and I will come to why. I disbelieve the first half of the conjunction. I do not believe Experian is in any position to know whether you have made a motor insurance claim. I am fairly confident it cannot know this at population scale, by any lawful route, and I am entirely confident that it cannot know whether such a claim was unreasonable, for the excellent reason that no such quantity exists anywhere in the recorded world.

What Would Have To Be True

Consider what the sentence requires. To train a model that identifies people who will not make unreasonable future claims, you need a training set: a large population of individuals, each labelled with whether they subsequently made a claim, and each of those claims further labelled as reasonable or otherwise. Then you need credit attributes matched to the same individuals. Then you need enough of both to fit something.

Where, in the United Kingdom, does the first of those live?

It lives in CUE, the Claims and Underwriting Exchange, which holds something like thirty-four million records of reported motor, home and injury incidents over a rolling six years — including, notably, incidents where no claim was ever made. CUE is run by the Motor Insurers’ Bureau. And the terms of access are not obscure: MIB states plainly that organisations must contribute claims data in order to receive it. Insurers, brokers, solicitors, delegated authorities. Every insurer that draws from it is required to feed it.

Experian is not an insurer. Experian does not settle claims, does not appoint engineers, does not repudiate, does not contribute a single record to CUE, and therefore does not draw a single record from it.

I find this genuinely funny. In the earlier instalment I spent several thousand words complaining about the Principles of Reciprocity — the credit industry’s private compact governing who may take what out of the shared pot, administered by a committee with no powers, honoured mainly in the observance of its own convenience. It turns out the insurance industry operates a reciprocity regime of its own at the claims end, and unlike the credit one it is actually enforced. You contribute or you are not in the room. Experian is not in the room.

And then there is the adjective, which is where the sentence stops being merely unsupported and starts being silly.

Unreasonable.

A claim is not reasonable or unreasonable. A claim is valid or invalid, settled or repudiated, withdrawn, disputed, referred, or flagged for investigation. Those are the states in which a claim exists, because those are the states a claims system can record. Nobody in a claims department has ever opened a file and typed unreasonable into a field, because there is no such field, because reasonableness is not a property of a contractual entitlement. It is a property of a mood.

You cannot model an adjective that no system has ever recorded. What you can do is write it in a brochure, because a brochure has no schema.

Three Readings, All Of Them Worse

I have spent a fortnight trying to construct a reading under which that sentence is true. I have found three. I invite anyone at Experian, or anyone who buys from them, to supply a fourth, and I mean that sincerely rather than rhetorically, because I would genuinely like to be wrong.

The first reading is that Delphi is what its name and its whole commercial history say it is — a credit-default scorecard, trained on credit outcomes, which is what Experian has and what Experian is good at — and that the insurance-claims half of the sentence is decoration. Aspiration bolted on by someone in a content team who was asked to describe a good insurance customer and reached for the two things insurers are known to dislike. On this reading the product is honest and the copy is not, and every insurer in Britain has been buying a default model while telling itself it bought a risk model.

The second reading is that the claim rests on bespoke work: an insurer hands Experian its own loss data under a sharing agreement, Experian matches its attributes to that book and fits a model. This happens; it is legitimate; it is how a great deal of joint bureau modelling is done. But then the resulting model is that carrier’s, built on that carrier’s claims, on that carrier’s mix, and the fourteen percent is a number from somebody’s book that has been promoted to a general property of a product. It is also, on this reading, the insurer’s own claims data doing the explanatory work while the bureau takes the billing.

The third reading is the one I offered in an earlier draft of this piece, and I now think it was too generous by half: that the sentence means literally what it says, and the label is a compound target — the union of two events, one a collision and one a bounced payment.

Take that third reading at face value for a moment, because it is instructive even though I no longer believe it.

A model trained against A-or-B distributes its explanatory effort across A and B in whatever proportion the data makes easiest, and you cannot recover the split by inspecting the output, because the output is one number. So ask which leg the bureau is better at. Predicting a missed direct debit is the thing credit files were built for: sixty years of refinement, outcome observed monthly, at enormous volume and negligible cost. Predicting whether someone reverses into a bollard is a borrowed use of the same records, resting on a correlation that two decades of literature has failed to explain.

Take ten thousand policies, worst credit decile against best, a thousand in each. Assume — for illustration, not as measurement — claim frequency of fourteen percent against nine, a spread of about 1.55 times, roughly what survives controls in the American work. That is 140 events against 90. Now assume direct-debit failure at fifteen percent against two, a spread of seven and a half times, which is unremarkable for an instrument doing its home job. The compound target separates 270 from 108: a handsome two and a half times, and a score built to hit it will validate beautifully.

But of the 180 excess bad events driving that separation, 130 are missed payments and 50 are claims. Seven-tenths of the discriminating power has nothing whatever to do with a car meeting a bollard. Move the assumptions where you like; the direction will not move, because it is structural.

So: reading one says the claims component does not exist. Reading two says it exists but belongs to the insurer who supplied it. Reading three says it exists and is drowned out. There is no version of this in which an insurer is buying what the sentence says it is buying, and I would be glad of a fourth.

Why This Is Gratuitous

Now let me say the thing that will disappoint anyone hoping for a straightforward denunciation of credit scoring in insurance, because I am not going to provide one.

The signal is real.

Credit-type data predicts insurance loss, and not feebly, and not merely as a smuggled reissue of the postcode. Monaghan’s paper for the Casualty Actuarial Society in 2000 found loss ratios about half again as high in the worst credit decile. The EPIC study ran 2.69 million earned car years across fifty states and placed credit second or third by variable importance in a full multivariate plan. Texas replicated it across nine carriers. The Federal Trade Commission’s 2007 report to Congress found roughly a twofold spread best to worst, compressing to about 1.7 times after controlling for age, driving history and geography. Golden and colleagues, peer-reviewed and non-industry, examined 175,647 Texas policies in 2016 and found average losses of $918 in the worst decile against $558 in the best, surviving controls for age and sex.

The figure that matters is the survival rate: something like sixty percent of the raw univariate spread persists after controls. The signal is incremental. It is not a redundant encoding of where you live and how old you are, and anyone who has put the variable into a model has watched the Gini move and knows it. I have. I buy this data. I would defend the practice over the second drink and have done so, at length, to the visible boredom of others.

Which is exactly what makes the sentence so contemptible.

They did not have to say it. Nobody compelled Experian to assert a claims-prediction capability it cannot demonstrate and probably does not possess. The honest pitch was available and was perfectly good: our data carries incremental signal against your loss experience, as twenty years of overseas evidence suggests and as your own validation will confirm. That pitch is true, sellable, and defensible in front of a regulator. Instead somebody reached for a sentence about unreasonable claims, and an entire industry of quantitatively literate people read it, nodded, and raised a purchase order.

There is a note in the Honesty Tax piece about what can be asserted without evidence being priced without evidence. I want to go a step further here, because this is worse. This is not an assertion made without evidence. It is an assertion made without access. The party making it has no route to the data that would be required to know whether it were true.

The Two Sentences I Want Back

I have twice written things in this series that this discovery obliges me to retract.

The first was in the account of the afternoon I spent wearing a premium-finance lender’s coat in order to buy credit scores I intended to use for something else. I said I had no burning wish to know whether your direct debit would bounce; that was the lender’s anxiety and the lender was welcome to it; I wanted the number to price you. I now think that on the first and most likely reading above, the lender’s anxiety is very largely what I bought, and that my fastidiousness about the distinction was a good deal more elegant than it was accurate.

The second was worse, because it was an argument. In dismantling the payment-preference loading I conceded — grandly, in passing, as one does when one is enjoying oneself — that credit risk was not the issue, because credit risk is already priced, quite openly, in the APR. And I listed the credit score among the channels through which the insurer already knew everything the tickbox could tell it: already pulled through a soft search, already rated.

I still think the loading is indefensible and I withdraw none of that. But I described the score as though it were clean claims signal, when it is quite possibly a default model with an insurance sticker on it. If so, the credit risk was not neatly housed in the APR at all. It was in both places, and I helped wave it through.

I notice, setting the two retractions side by side, that they share a shape with the finding they retract. In the preference business the industry was pricing one quantity on evidence drawn from another — trained on who actually paid, levied on who said they would. Here a vendor describes one product and sells another. I had taken the first for an embarrassing blunder that an undergraduate would have caught. I am now inclined to think it is simply how the stack is built: the label is never quite the thing, the two are close enough that validation never complains, and the gap between them is where the money sits.

Derek, Who Wins One At Last

Derek, by now, has read the literature, and intends to stop paying the honesty tax, and is quite right to. He will tick annual. The loading will fall away. He will be so pleased with himself that he will raise it at Christmas between the starter and the main, and for once in fifteen years he will be entirely correct about something, and I hope somebody buys him a drink.

Then April will come, and he will take the instalments anyway, because he still does not have seven hundred pounds lying about and never has, and no tickbox ever changed that.

Here is what he cannot restate. Insurance is not credit and does not go on his file. Premium finance is credit and does. The moment he accepts the instalments there is a hard search, a credit agreement, and a payment record reported month upon month into the shared reservoir — which feeds the score, which prices next year’s premium. He can revise a declared preference. He cannot revise April.

And the corollary is the ugliest thing in this whole business. The only customers whose insurance behaviour becomes credit data are those who could not find the lump sum. They are rendered legible; the comfortable are not. A missed instalment is a sharp, well-modelled, expensive signal, while three years of paying on time is a faint one. So the apparatus learns most about the people it can least afford to be wrong about, and it learns about them through the one channel that opens only when they are short of money.

By Whom, And Under What Authority

On the word prohibited I have written elsewhere and will not repeat myself, beyond noting the one point that bears on this. What that trade compact restricts is a data format, not a use: the insurer may not read your accounts, but the bureau may distil those accounts into a score and sell the score. And a scorecard is a compression — it discards what fails to predict the target and preserves what succeeds, approaching in the limit a sufficient statistic. The better the modelling, the less of the restriction survives it. A thoroughly incompetent score would honour the Principles of Reciprocity rather well. One assumes Experian’s does not.

And nobody decided any of this. A committee in the 1990s decided something about lending, in a document that never contemplated insurance because insurance is not credit. A bureau built a product, and someone in marketing described it in terms nobody checked. A pricing actuary, doing precisely the job the profession asks of him, evaluated a candidate factor, observed genuine incremental lift, documented it properly, and put it in the plan — and he was not wrong to, because the lift was there. A privacy notice was updated by counsel and published, accurately. A regulator studied renewal pricing, then premium finance, then area ethnicity, and never once asked what was in the inputs. Every step defensible. No signature anywhere on the whole.

The Objection, Fairly Put

The strongest objection is that every rating factor is a proxy and none is causal. Being nineteen does not cause a crash; a postcode does not cause a theft, it correlates with the proximity of people who steal cars, almost none of whom are you. We have run insurance pricing on unexplained correlation for a century, because the alternative — pricing solely on demonstrated individual causation — is not insurance at all. Why should credit meet a burden of proof that age has never been asked to meet?

Mostly right, and I would not have it otherwise. But three things are true of this factor and of nothing else in the plan. You cannot see it, whereas you know your own age. Its composition is proprietary even to the firm applying it, so the Equality Act’s actuarial exemption — which asks for reliable statistical grounds — would require an insurer to defend a variable whose innards it does not hold. And no vendor has ever described the age factor in terms that the vendor could not possibly substantiate.

The practical objection I make against myself. Banning the factor does not abolish the risk, it relocates it. Four American states have tried, and the cost surfaced on the young and the urban, who are not obviously more deserving. Unlike the preference loading, this is not a straightforward extraction with competitors proving it unnecessary. There is real signal here, which is what makes it hard, and what makes the sales copy such a betrayal of it.

Where I Might Be Wrong

Perhaps Experian has an arrangement I know nothing about. Perhaps there is a claims-linked panel, a consented dataset, a research agreement with a carrier that permits the general claim. Perhaps unreasonable is a translation of some internal flag, clumsily rendered. All possible, and none of it public.

And that is the finding underneath the finding. If the copy is unreliable, then the public domain contains no reliable account of what this score is optimised against at all: composition proprietary, weight in the rating plan confidential, UK validation unpublished, no peer-reviewed British study in existence, and no regulator having asked in a decade of market studies. The footnote is not good evidence. It is merely the only evidence.

A variable of unknown weight and unknown composition, trained against an undisclosed target on unpublished data, sets the price of a legally compulsory product for tens of millions of people. The sole public description of what it measures is a line of sales copy that its author cannot have believed.

I still think the signal is real. I still buy the number. What I can no longer tell you is what I am buying.

Quo Vadis, Derek

The title of this piece is Experian’s phrase and I have handed it back to them, since on the evidence the unreasonable future claims in this market are not being made by policyholders.

Derek, for his part, will go on paying in April, and being written down for it, and telling anyone who will listen that he has the system beaten. He is nearer to right than he has ever been. That is not the same as being right, but after seven of these it will have to do.

The meerkat, on the mantelpiece, is not required to have an opinion.

Themes: Prices That Lie The Corruption of Language Bureaucracy & the Vanishing Decision-Maker Evidence & the Wish to Believe

← All essays