Reading in standalone mode. Open this treatise in the complete 2-Column Sovereign Research Wiki Engine:Open Wiki Dashboard (117 Treatises) →
Fooled by Best Practice VDigital Twin Architecture

Fooled by Best Practice, Part V: Now, Next, Never

100% Complete & Untruncated 44 min read
Return to Research Tracks

J. McKenney

Paper 5 of five in Fooled by Best Practice, and the last. Part I made the argument in ordinary language. Part II gave the distinction it rests on. Part III gave the method and its limits. Part IV gave the engineering response. This part is about the decision, and it is written for whoever signs.

Licence: CC BY 4.0. 16 September 2026.

FieldValue
Document IDWG-02-DT-FBP-5
Slugfooled-by-best-practice-5-now-next-never
Working groupWG-02-DT, Digital Twin
SeriesFooled by Best Practice, part 5 of 5 published
AuthorJ. McKenney
Published16 September 2026
Revision
StatusPublished
Preceded byWG-02-DT-FBP-1, WG-02-DT-FBP-2, WG-02-DT-FBP-3, WG-02-DT-FBP-4
Followed by

Executive Abstract#

Four papers argued that the loss which matters in a critical facility sits in the tail, and that the industry's instruments cannot see it. This part is about spending a budget on Monday morning.

The usual method is wrong. It multiplies a bad year's cost by the chance of one and spends a fraction, averaging over futures the operation may not survive. If a single event can end you, you are managing whether you keep playing, not an average.

There is a real published limit, widely misquoted. Gordon and Loeb showed in 2002 that the best defence spend never exceeds about 37 percent of expected loss. Read against potential rather than expected loss, that ceiling can inflate a budget tenfold.

The allocation is three buckets, NOW, NEXT and NEVER, sorted by five tests, after the budget splits into licence-to-operate spending and what changes exposure. Confusing the two is how organizations think they bought risk reduction when they bought permission.

Then insurance: since 31 March 2023, standalone cyber policies at Lloyd's must exclude state backed cyber-attacks, under market bulletin Y5381, so the events most capable of ruining a critical operator are, in part, the ones the policy now excludes. The last section weighs conformance against simulation; I build one and leave the choice open.

Abstract#

This paper concludes a five-part series applying Taleb's epistemology to operational-technology security. If exposure lives in the unobserved branches and the severe ones dominate the mean, expected-value budgeting targets the wrong statistic; the operator's constraint is survival. It restates the Gordon and Loeb bound, whose 2002 result caps optimal investment at 1/e of the expected loss under risk neutrality and two families of breach probability function, and corrects three common misstatements. The budget is first split into licence-to-operate and risk spending. Risk spending is sorted into NOW, NEXT and NEVER by five sequential tests: whether it can cause the loss it addresses, whether it acts on tail or body, its cost-benefit asymmetry, whether failure would be observable, and its lead time and reversal cost. Worked placements cover five budget lines. Lloyd's market bulletin Y5381 of 16 August 2022 required state backed cyber-attack exclusions in standalone cyber policies at Lloyd's from 31 March 2023, with an agreed basis for attribution. It states the consequence for an operator who treated cyber loss as transferred, poses four broker questions, disclaims advice, and withdraws claims about deductibles and legal defence, including the author's own. A final section weighs conformance against simulation, including the case against its own method; the author declares a commercial interest and declines to recommend.

1. Introduction#

1.1 The Monday morning question#

I have had the same conversation in a control room in Queensland, a substation office outside Hamburg and a rail operations centre in the American Midwest, and it always arrives at the same place. Someone has followed the argument. They accept that the incident-free record is one sample path. They accept that the audit measures something narrower than they thought. And then they ask the only question that matters to them, which is what they should do differently with the money they are about to commit.

That question deserves a straight answer, and four papers of epistemology do not constitute one.

So this part is the answer, as far as I can give it. It is shorter on philosophy than the ones before it and longer on arithmetic and procedure. It assumes you have a budget, a plant, a regulator and a finite number of outage windows a year, and that you are going to have to defend your choices to somebody who has not read any of this.

1.2 What I am claiming and what I am not#

I am claiming that the standard method for sizing a security budget is built on a statistic that does not describe the operator's constraint, and that a different allocation follows from taking the tail seriously.

I am not claiming that the allocation I give is optimal, or that it is novel. Three-bucket prioritization is old. What I think is worth the pages is the sorting rule, because the buckets are useless without one and most published versions of this scheme supply the buckets and leave the sorting to taste.

I am also not claiming that any of this removes uncertainty. It moves the uncertainty into places where you can see it and argue about it, which is a smaller achievement than the industry usually advertises and the largest one I think is available.

One caution before the arithmetic. This paper uses illustrative numbers in several places. They are illustrative. They are not measurements of any facility I have worked on, and I have not put a client's figures into a public document.

2. The wrong objective#

2.1 What is actually in the spreadsheet#

Open almost any security business case in this sector and you will find a version of the following. There is a figure for what a serious event would cost, usually assembled from lost production, replacement plant, regulatory penalty and some allowance for the rest. There is a figure for how often such an event is expected, often expressed as a return period. The two are multiplied. The product is an annualized expected loss, and the proposed spend is justified as a defensible fraction of it.

This is a respectable piece of work and everyone involved has done their job. It is also the wrong instrument for the thing it is being used on, and the reason is not that the inputs are uncertain, although they are.

2.2 Averaging over the futures you do not have#

An expected value is an average across the whole set of possible futures, weighted by their probabilities. It answers a question of the form: if this situation were repeated many times, what would the typical cost be.

That question is well posed for a portfolio of many small independent exposures. An insurer writing a hundred thousand motor policies genuinely does face something close to the average, because the law of large numbers does the work.

A single facility does not face the average. It faces one future. And if one of the branches in that future involves the operator losing its licence, its plant, or a person, then the branches on the other side of that event are not available to be averaged with the good ones. You cannot recover a bad year with a good one when the bad year removes you from the game.

Taleb makes this point in several places across the Incerto, most directly in Skin in the Game, where the distinction is between an average taken across a population at one moment and an average taken across time for one participant [4]. They coincide only when nothing absorbing can happen. Ruin is absorbing. Where a ruin branch exists, the two averages come apart, and the one that describes the operator is not the one in the spreadsheet.

This is not an argument that expected values are useless. It is an argument about which decisions they are the right tool for. For deciding how many spare protection relays to hold, expected value is exactly right, because you will face that decision hundreds of times and the errors will wash out. For deciding whether the plant can survive a determined adversary, it is not, because you face that once.

2.3 What the tail does to the number#

There is a second problem, and it is mechanical rather than philosophical.

The expected loss figure in the spreadsheet is almost always built from a single scenario, or a small number of them, chosen because they are the ones people can picture. A ransomware event that takes the corporate estate out for two weeks. A control system outage of a certain duration. These are body-of-the-distribution events. They are frequent enough to imagine, and their costs are estimable because similar things have happened to comparable operators and some of it has been published.

The branches that dominate a fat-tailed loss distribution are not those. They are the low-probability, high-consequence branches in which several things go wrong in sequence: the intrusion, plus the human response failing under load, plus the physical consequence not being contained. Part III explains why sampling is the only honest way to find those and what the sampling cannot tell you. The point here is narrower. If your expected loss figure omits them, it is not a slightly conservative estimate. It is an estimate of a different quantity, because in a fat-tailed distribution the rare branches contribute a large share of the mean, and omitting them does not shave the answer, it relocates it.

So the spreadsheet is wrong twice. It uses an average where the operator faces a single path, and the average it uses is computed over a sample that excludes the branches that make the average what it is.

2.4 What replaces it#

Not a better point estimate. A different question.

The question the tail argument produces is: which of my candidate spends changes the shape of the severe branches, and which of them only reduces the cost of the branches I was going to survive anyway.

That is a question about structure rather than about magnitude, and it is answerable without agreeing on a loss figure at all, which turns out to be its main practical virtue. Two engineers who cannot agree within a factor of five on what an event would cost can usually agree, quickly and without rancour, on whether a proposed control touches the branch where the process goes out of control.

The rest of this paper is built on that question.

3. The Gordon-Loeb ceiling, stated properly#

3.1 What they proved#

In 2002, Lawrence Gordon and Martin Loeb published a paper in ACM Transactions on Information and System Security which has been cited widely and quoted carelessly ever since [5].

Their setup is a single information set with three parameters. There is a potential loss if the set is breached. There is a threat probability, meaning the chance an attack is attempted. There is a vulnerability, meaning the conditional probability that an attack succeeds given that it is attempted. The product of the three is the expected loss in the absence of any security investment, and in the paper's compact notation that expected loss is written as v times L, where v is the vulnerability and L absorbs the potential loss and the threat probability together.

They then introduce a security breach probability function, which gives the probability of a breach as a function of the vulnerability and the amount invested. They require it to behave sensibly: an invulnerable asset stays invulnerable at zero cost, zero investment leaves the vulnerability where it was, more investment never increases the breach probability, and investment has diminishing returns, so each additional pound buys less reduction than the one before. They construct two families of function with these properties.

Their result, for both families, is that the investment which maximizes the expected net benefit is bounded above by 1/e of the expected loss. The number 1/e is approximately 0.368, so the optimum is never more than about 37 percent of the expected loss.

That is a genuinely interesting theorem. It says that diminishing returns alone, without any appeal to budget limits or to management preference, caps what it is rational to spend. You do not need to argue about whether security is worth it. Past a point, the arithmetic decides.

3.2 The assumptions, which are load-bearing#

The bound is a consequence of a model, and the model makes commitments. These are the ones that matter when somebody tries to apply the result to a plant.

The decision maker is risk neutral. The objective being maximized is expected net benefit, so the entire argument of section 2 applies to it. Gordon and Loeb are explicit that this is their framework, and it is the correct framework for the problem they set themselves, which is about information assets rather than about physical ruin.

The decision is static and single-period. Investment happens once, the world does not respond, and the adversary does not adapt to what you did. In a plant where the adversary is a funded organization reading your public filings, that assumption is doing real work.

Investment is perfectly divisible and continuously variable. You can spend any amount. In practice, security investments in operational technology are lumpy in the extreme. An interlock either exists or it does not. A segmentation project either happens in the outage window or waits a year. The continuous optimum can sit in a region that contains no purchasable option.

The vulnerability and the loss are known. The bound is a function of parameters that the operator is being asked to supply, and section 3.4 says what that does to the answer.

There is a single information set and no interdependence. Real facilities have controls whose value depends on other controls, and adversary paths that traverse several sets. The paper's later literature takes this up and it is not a small correction.

And the bound is proved for the two families of breach probability functions the authors construct. It is not a statement about every function that could describe a real defence. This is a point the paper itself is careful about and its popularizers usually are not. If your defence has a different response to money than either family assumes, the bound derived from those families does not automatically carry over.

None of this makes the result worthless. It makes it what it is: a clean theorem under stated conditions, which tells you something true about diminishing returns and does not tell you what to spend.

3.3 The misquotation, corrected#

I have heard "Gordon-Loeb says spend 37 per cent" in three separate procurement meetings, and I have seen it in vendor material. It is wrong in at least four distinguishable ways and they compound.

It is a ceiling, not a target. The theorem says the optimum is at most 1/e of the expected loss. For many parameter values the optimum is far below that, and for a low-vulnerability asset it approaches zero, because there is little left to buy. Reading an upper bound as a recommendation converts a limit on waste into a licence for it.

It is a fraction of the expected loss, not of the potential loss. This is the expensive error. Expected loss is the potential loss multiplied by the probability that it occurs. If a facility's serious-event loss is 40 million and the annual probability is 5 percent, the expected loss is 2 million and the ceiling is about 736,000. Applying 37 percent to the 40 million instead gives 14.8 million, which is twenty times the ceiling and more than seven times the entire expected loss. A budget built that way is not conservative. It is spending seven pounds to avoid an expected one.

It is per information set and per period. The bound applies to the asset being modeled, over the period the parameters describe. Summing a 37 percent allowance across every asset in a plant is not a use of the theorem, it is a coincidence of arithmetic wearing the theorem's name.

And it assumes risk neutrality, which is exactly the assumption section 2 said does not hold for an operator facing a ruin branch. Under ruin the objective is not expected net benefit, and a bound derived by maximizing expected net benefit does not bind the decision you are actually making. This is the most interesting of the four, because it cuts in the opposite direction to the others: the first three say people spend too much on the strength of the number, and this one says the number cannot tell you that you are spending too little on survival.

3.4 What the ceiling is good for#

Two things, and I use it for both.

The first is stopping a runaway subscription. If an operator is committing an annual recurring spend that exceeds the expected loss of the assets it protects, that is not a judgment call, it is an arithmetic error, and the ceiling makes it visible in one line. I have used the calculation in exactly that way and it ends the discussion quickly, because there is no counterargument available that does not require restating the loss figure upward, which then has to be defended.

The second is forcing the loss estimate into the open. The ceiling is linear in the expected loss. Double the loss figure and you double the permitted spend. That means the entire budget rests on a number that was probably produced in an afternoon by whoever was free, and the ceiling makes the dependence explicit rather than leaving it buried. In my experience the most useful output of a Gordon-Loeb calculation is not the ceiling at all. It is the argument the calculation provokes about the input.

There is a sharper version of that argument. Suppose the operator's 2 million expected loss is built from a 5 percent chance of a 40 million event. Now add one branch that nobody costed: a half percent chance of a 400 million event, where the process is destroyed rather than interrupted. The expected loss becomes 4 million and the ceiling doubles to about 1.47 million. The permitted budget has doubled on the strength of a branch estimated at one chance in two hundred, which is precisely the branch nobody can evidence.

That sensitivity is not a defect of the model. It is the tail problem arriving inside the finance. The quantity that most determines what you are allowed to spend is the quantity you know least about, and no amount of care in the multiplication fixes it.

3.5 So what does the ceiling actually decide#

It decides the size of the risk budget and nothing about its contents.

This matters because the contents are where the entire argument of this series lives. Two operators can spend the same permitted amount, both within the ceiling, both defensible on paper, and end up with completely different exposures, because one of them bought controls that reduce the probability of the branches they were going to survive and the other bought structure that changes what happens on the branches they were not.

The next section is about that choice.

4. Now, next, never#

4.1 First, split the budget in two#

Before any sorting rule can work, the money has to be separated into two accounts, and most organizations I have worked with keep them in one.

The first account is licence to operate. It holds everything you buy because a regulator, a standard, a customer contract or an insurer requires it. Conformance work under IEC 62443 [8], the NIST framework [12], NERC CIP [13] and the national implementations of NIS2 [11] mostly lands here, along with the audit itself, the evidence collection and the annual training refresh. This is not a criticism of any of those documents. Part I said I use 62443 constantly and I meant it. It is a statement about which account the spend belongs in.

The second account is risk. It holds spending whose justification is that it changes your exposure.

The two overlap, and where an item genuinely does both, it should be booked in both with the amount split, which is uncomfortable and correct. What must not happen is the thing that happens by default, where conformance spend is reported to the board as risk reduction. When that happens, the board sees a large and growing security investment and concludes, reasonably, that exposure must be falling. The organization has bought permission to operate and recorded it as protection.

The test for which account an item belongs in is short. If the regulator withdrew the requirement tomorrow, would you still buy it. If the honest answer is no, it is licence to operate. Buy it anyway, because operating without a licence is a worse outcome than any of this, but do not count it twice.

Everything from here applies only to the risk account.

4.2 The five gates#

A candidate control goes through these in order. The order matters, because the first gate is disqualifying and the later ones only sort.

Gate one: can this control cause the loss it is meant to prevent. Part IV set out the mechanisms in detail and I will not repeat them: scan cycle preemption, connection buffer exhaustion on legacy stacks, kernel privilege on machines governing thermodynamics, and syntactic validation standing in for physical constraint. The gate is the question those mechanisms produce. If the control, by failing in its own normal failure mode, can produce an uncommanded shutdown, a loss of view, or a loss of control, then it is NEVER, for the zone in question, regardless of how well it scores on everything else.

This gate is absolute and it is the only one that is. It is absolute because the failure it screens for is correlated with the event you are defending against: the conditions that stress your control system are the conditions under which the agent misbehaves, so the control fails exactly when it is needed. A defence whose failures cluster with the incident is not a defence, it is an additional branch in the tail.

Note the qualifier, for the zone in question. The same product can fail gate one on a controller network and pass it comfortably on a business network two zones up. The gate is asked about a deployment, not about a product, and the commonest error I see is applying the answer from one zone to another.

Gate two: does this act on the tail or on the body. Ask which branches change if the control is present. If the answer is that incidents you would have contained become incidents you contain slightly faster, it is a body control. If the answer is that a branch where the process goes out of control becomes a branch where it does not, it is a tail control.

Body controls are not worthless. They reduce real cost and they are often cheap. They are simply not what the risk account is for, and they compete for it on an unequal footing because their benefits are easier to evidence. That asymmetry is worth naming: the controls whose value you can prove are systematically the ones whose value is smallest, because provability and frequency travel together. If a control's benefit is easy to demonstrate, it is because the events it addresses happen often enough to count, which means they are body events.

Gate three: which way does the asymmetry run. Taleb's barbell, which Part IV applied to architecture, is an allocation rule before it is an architecture [3]. What it asks of a position is whether the downside is bounded and known while the upside is open, or the reverse.

For a control, the question becomes: is the cost bounded and the benefit open-ended, or is the cost open-ended and the benefit bounded. A hardwired trip has a known one-time cost, a twenty-year life and a benefit that is as large as the worst event it prevents. A per-node annual subscription has a cost that grows with your estate, with the vendor's pricing changes and with the length of the relationship, and a benefit capped by what the product detects. Those are opposite positions and the accounting treats them as the same kind of thing because both appear as a number in a budget line.

Gate four: what observation would show this had failed. This is Part I's argument turned into a procurement question, and it is the one that makes people uncomfortable, which is why I keep asking it.

If a control is installed and no incident occurs, is that evidence it worked. For most detection products the honest answer is no, because no incident is also what you would see if it were doing nothing, and you cannot distinguish the two from the outcome. That is the survivorship structure exactly, and installing more of the thing does not resolve it.

Some controls survive this gate cleanly. A mechanical interlock can be proof-tested, and the functional safety standards prescribe how and how often [9][10]. A segmentation change can be verified by attempting the traversal and failing to complete it. A removed access path can be confirmed removed. These produce evidence that is independent of whether an attack happened, which is the property the gate is looking for.

Controls that cannot pass this gate are not automatically NEVER. They are barred from NOW, because NOW is where you spend on the strength of an argument and the argument has to be checkable.

Gate five: what is the lead time and the reversal cost. This gate does no filtering at all. It sorts what remains into NOW and NEXT. A control that can be executed inside the current period, without an outage, and undone cheaply if it turns out to be wrong, is a NOW candidate. A control that needs a shutdown, a capital cycle, a vendor contract renegotiation or a design change is a NEXT candidate.

The distinction is scheduling, not importance. Some of the most valuable items in this scheme are NEXT items, and treating NEXT as a lower tier is the single most common way the framework gets corrupted in practice.

4.3 What NOW means#

NOW is the set of items that pass gate one, act on the tail, carry bounded cost against open benefit, can be shown to have worked, and can be done in this period.

There are usually fewer of these than people expect, and they are usually smaller. In the assessment work I have done, the NOW list is dominated by removals rather than purchases: a permanent vendor access path reduced to a scheduled one, a bidirectional link replaced with a one-way path where the downstream consumer was only ever reading, a dual-homed engineering workstation given up, a protocol gateway configuration that was permitting a function code nobody uses.

None of that is interesting to present. All of it passes five gates, and most of it costs engineering time rather than capital.

There is one category in NOW that does not reduce any risk by itself, and it needs stating separately because otherwise the gates exclude it wrongly. Spending that buys information rather than protection belongs in NOW when the information is the binding constraint on everything else. A survey that establishes what is actually connected in the zone that carries the consequence is the clearest example. It prevents nothing. It also determines whether every other number in this paper is a measurement or a guess, because Part I's reference architecture illusion means the drawing you would otherwise use as input is wrong in ways nobody has enumerated. If you cannot see the paths, you cannot allocate against them, and no amount of care downstream repairs that.

The test for an information purchase is whether a specific downstream decision is currently blocked on it. If yes, it is NOW. If it is a survey commissioned because surveys are good, it is not.

4.4 What NEXT means#

NEXT is the set of items that pass the same gates and cannot be done in this period.

These are the structural changes: the segmentation that changes the topology rather than filtering traffic across it, the hardwired interlock on the process whose failure carries the consequence, the replacement of a legacy gateway that cannot be defended, the contractual change that ends a standing remote access arrangement.

Two rules make NEXT work, and without them it fails reliably.

The first is that a NEXT item is funded now and delivered later. The money is committed in the current budget against a named future window. If NEXT means the item goes on a list and the funding decision is deferred, then NEXT is a polite refusal and everyone in the room knows it.

The second is that a NEXT item has a named owner and a named window, and if it has neither it is not a NEXT item. It is a NEVER item that nobody wanted to argue about. I have seen registers with items that have been NEXT for six years, and every one of them was a decision that had been made and not admitted. Writing it in the right column costs nothing and is honest. Leaving it in NEXT consumes the attention of everyone who reviews the register, every quarter, for years.

4.5 What NEVER means#

NEVER means not from this account, on this evidence. It does not mean the control is bad, and it does not mean the vendor is dishonest.

Four things land here.

Anything that fails gate one for the zone in question. This is the largest category in operational technology and it is the one that generates the most resistance, because the products in it are competently built and widely deployed and the person proposing them has usually been told by their peers that it is standard practice. The engineering case is in Part IV and it is specific.

Anything whose only justification is that a framework lists it, where no sampled path touches it. That is a licence-to-operate item that has been filed in the wrong account. Move it, buy it, and stop counting it as risk reduction.

Anything whose cost scales with your estate while its benefit is capped. This is gate three failing. It is not that the product does nothing. It is that you are taking an open-ended obligation to buy a bounded benefit, and the position gets worse every year as the estate grows.

And anything you cannot falsify, where the spend is large. Gate four bars these from NOW. If the item is also expensive, it is NEVER, because you are being asked to commit substantial capital to a proposition that no future observation could disconfirm, and the correct response to that request is no.

A NEVER decision should be written down with its reason. That sounds bureaucratic and it is the part that saves you, because these decisions get revisited. When a new CISO arrives and asks why there is no endpoint agent on the controller network, the answer needs to be a recorded engineering judgment with a named author, not an absence.

4.6 Placing five budget lines#

The gates are only worth anything if they sort real items, so here are five.

An endpoint detection agent, licensed per node, proposed for all engineering workstations, operator stations and the controller network. Gate one is asked separately for each zone. On the controller network the answer is that a user-space agent contending for CPU and memory on a machine running a deterministic scan cycle can produce the failure the plant is defended against, so that portion is NEVER. On operator stations the answer depends on whether losing the station loses the view, and in most facilities it does, so that portion is NEVER as well. On engineering workstations that sit above the boundary, that can be rebooted without consequence, and that are not in any control path, gate one passes, and the item then meets gate three, where a per-node annual subscription against a bounded detection benefit puts it into NEXT at best. The recommendation that arrives as one line item leaves as three different decisions, which is the normal outcome and the reason the gates are worth running.

A hardwired overpressure trip on the single process whose failure carries the largest physical consequence. Gate one passes trivially, since a mechanical device with no firmware has no software failure mode to correlate with the incident. Gate two is a tail control by construction, because it acts only on the branch where containment fails. Gate three is a one-time capital cost against an open-ended benefit. Gate four passes, because it can be proof-tested on a schedule under the functional safety standards [9][10]. Gate five requires an outage. So it is NEXT, funded in this budget, scheduled against a named window, with an owner. It is also probably the most valuable item on the list, which is the clearest illustration of why NEXT is not a lower tier.

One qualification I want to make honestly. A safety integrity level is a claim about random hardware failure rates and about systematic capability in the design process. It is not a claim about an adversary, and nothing in those standards was written with one in mind. What a hardwired trip gives you against an adversary is not a rating, it is the absence of an attack surface, which is a different and in this context better thing.

A permanent vendor remote access path into the process network, currently always available. Gate one passes, since removing a path cannot cause the loss. Gate two is a tail control, because that path appears in a large share of the sampled sequences in every engagement I have run, and it appears early in them. Gate three is bounded cost, since it is engineering time plus a contractual conversation, against an open benefit. Gate four passes cleanly: you can verify the path is gone by trying to use it. Gate five is days, not outages. NOW.

The reason this one is not already done is never technical. It is that the maintenance contract assumes the path, and the person who can change the contract is not in the security function. That is a real obstacle and it is a different kind of obstacle, and knowing which kind you are facing is most of the work.

The annual security awareness training refresh, required by the framework. Ask the licence-to-operate test first. If the requirement vanished, would you still buy it at this scale. In most organizations the answer is no, so it moves to the other account, gets bought, and stops appearing in the risk conversation. That is not a judgment on whether training has value. It is a statement about what the item is for and which budget it is defending.

An asset and connectivity survey of the zone that carries the consequence. It prevents nothing, so gate two would exclude it. The information exception in section 4.3 applies, because the segmentation decision, the interlock siting and the access path inventory are all currently blocked on not knowing what is connected. NOW, on the condition that the downstream decisions are named in advance. Commission it against those three decisions, not against a desire for completeness, or you will get a document rather than an input.

4.7 How this framework fails#

I would rather say this than have it discovered.

NEXT becomes a holding pen. Covered above and it is the most common failure by a wide margin.

The gates get gamed. Once people know that tail controls get funded, every proposal acquires a paragraph about the tail. The defence is that gate two has to be answered with a specific branch and a specific mechanism, not with the word catastrophic. If the proposer cannot say which branch changes and how, the gate has not been passed.

Gate one gets negotiated. There is always pressure to make an exception for a particular product on a particular network, usually because a peer operator has done it. The correct answer to the peer argument is Part I: you are being shown a site that deployed it and did not have an incident, which is a sample conditioned on the outcome.

The whole scheme gets used to justify spending nothing. The gates are restrictive, and a determined reader can use them to reject everything and bank the money. That is a misuse, and the tell is an empty NEXT column. An organization with genuine exposure and no structural work scheduled has not applied the framework, it has applied the framework's vocabulary to a decision it had already made.

And the sorting depends on model outputs that carry their own error. If your view of which branches are tail branches comes from a simulation, then the limits Part III sets out apply to every placement in this section. I return to that in the final part of this paper, because it is the strongest objection to the method and it deserves to be stated by me rather than by someone else.

5. The part where you find out what is insured#

5.1 Why this section is here#

An allocation decision depends on which part of the loss you are carrying. If a tail event is insured, the operator's exposure to it is the deductible plus the uninsured consequences, and the budget should reflect that. Every operator I have discussed this with had a view on what their cyber cover did. A number of them turned out to be working from the position before 2023.

I am not a lawyer and I am not an insurance adviser. What follows is a description of a published market instruction and its plain consequences. Your policy wording governs, not this paper, and the wordings differ from one another in ways that matter.

5.2 What Y5381 says#

On 16 August 2022, Lloyd's issued market bulletin Y5381, addressed to managing agents, on state backed cyber-attack exclusions [6].

It required that from 31 March 2023, at inception or on renewal, all standalone cyber-attack policies written at Lloyd's exclude liability for losses arising from any state backed cyber-attack. It set out requirements that the exclusion must meet. The clause must exclude losses arising from a war, whether or not declared, where the policy does not already carry a separate war exclusion. It must exclude losses arising from state backed cyber-attacks that significantly impair the ability of a state to function or that significantly impair a state's security capabilities. It must be clear about whether cover is excluded for computer systems located outside a state affected in that way. It must set out an agreed basis on which a state backed cyber-attack will be attributed to one or more states. And it must be clear about whether the definition of computer system extends to third parties' systems as well as the insured's.

Lloyd's did not prescribe a wording. The market has largely used the model clauses published by the Lloyd's Market Association in November 2021, numbered LMA5564 to LMA5567, which differ from one another in how much is carved back into cover [7].

Three limits on scope are worth stating plainly, because the bulletin is routinely described more broadly than it is. It applies to standalone cyber-attack policies. It applies to policies written at Lloyd's. And it is an instruction to managing agents about what their policies must contain, not a change in law.

5.3 The attribution problem#

The requirement that the clause set out how attribution will be decided is the one an operator should read closely, because it determines who decides whether you are covered and on what evidence.

The model wordings commonly provide that the primary basis is attribution by the government of the state in which the affected computer system is located. They then address what happens when no such attribution is made, or when it is slow, and the usual provision allows the insurer to rely on an objectively reasonable inference from the nature or the apparent objective of the attack. Wordings differ on this and yours may not say what I have described, which is why this paragraph is a prompt to read it rather than a summary you can rely on.

Two consequences follow for planning, whatever your specific wording turns out to say.

The first is timing. Public attribution of a significant cyber-attack to a state, where it happens at all, has historically taken months. Your recovery does not wait for it. So the realistic case is not a denied claim, it is a claim that is unresolved while you are funding restoration from your own balance sheet, and the liquidity question is separate from the coverage question and arrives much earlier.

The second is that attribution is a judgment made about an adversary who has an interest in how it comes out. Sophisticated intrusions are routinely conducted so as to be ambiguous in exactly this respect, and that was true before anyone wrote an exclusion that turned on it.

5.4 What this means for an operator who thought cyber was covered#

Here is the uncomfortable shape of it.

The exclusions apply to state backed attacks, and particularly those that significantly impair a state's ability to function or its security capabilities. Operators of critical national infrastructure are, by definition, among the assets whose impairment does that. The adversaries with the capability and the motive to cause the branches this series has been about are disproportionately the adversaries the exclusions describe.

So the correlation runs the wrong way. The events most capable of ending the operation are, in significant part, the events the policy has been rewritten to exclude. Insurance is transferring the body of your loss distribution and leaving the tail where it was.

That is not a scandal and I do not think the market did anything improper. An insurer cannot write systemic correlated risk at a scale it cannot diversify across, and a war exclusion in some form is as old as the industry. It is simply a fact about your position that changes the arithmetic in sections 3 and 4, and the operators I have met who were most exposed to it were the ones who had not read the renewal.

There is a second gap that is easy to miss. Physical damage and business interruption arising from a cyber cause can fall between two policies: excluded from the property programme by a cyber exclusion, and excluded from the cyber programme by a state backed exclusion. Whether that gap exists in your programme is a question about your two wordings read together, and it is not answerable from either one alone.

Four questions to put to your broker, in this order. Is our cyber cover standalone or an endorsement, and where does it sit relative to our property and business interruption programme. Which exclusion wording is in the policy, by number, and what is carved back. Who determines attribution under that wording, on what evidence, and what is the process if we dispute it. And if a cyber-caused physical loss occurred and attribution were contested, which policy responds in the interim and what do we fund ourselves while it is contested.

5.5 Claims I will not repeat#

There is material in circulation, including material my own company produced before I looked at it properly, which states that a particular architecture secures deductibles compressed by up to ninety percent, that it obtains an affirmative legal defence under the so far as is reasonably practicable standard, and that it insulates directors from personal liability under European instruments.

I cannot support any of those and I am withdrawing them here.

What I can say is narrower and I think it is still worth something. Underwriters price what they can verify. A physical protective function that does not depend on software integrity is verifiable in a way that a software control posture is not, because it can be inspected, proof-tested and evidenced on a schedule under standards the engineering world already accepts [9][10]. Whether any given underwriter prices that difference, and by how much, is a commercial matter between you and them, and it will depend on your programme, your loss history and the market at the time you renew. Anybody quoting you a percentage in advance is quoting you a number they cannot know.

The same applies to the liability point. Duty holders in several jurisdictions are required to reduce risk so far as is reasonably practicable, and NIS2 places responsibilities on management bodies for approving and overseeing cybersecurity risk-management measures [11]. Whether a particular engineering decision discharges a particular duty is a question for a lawyer who has read your facts, and I am not one.

6. What the five parts established#

The argument, stripped of its examples, is four moves and I want them on one page before the last section.

An incident-free record is a single realized path and not a property of the distribution it was drawn from, and the industry's evidence base is depleted twice over, once by outcomes that never occurred and once by outcomes that occurred and were not disclosed. That was Part I.

The instruments the sector uses are built on the observable side of that distinction and are then read as though they described the unobservable side. Conformance measures the presence of controls. The quantity operators care about is the exposure of a specific plant at a specific moment, and no general document can carry it. That was Part II.

If the exposure lives on the unobserved side, the only instrument that addresses it directly is one that samples it, which means generating a facility's counterfactual histories and examining the distribution rather than the realized path. That method has hard limits, and the limits are where the vendors go quiet. That was Part III.

And the engineering response is to stop putting fragile software in the path between an adversary and physical consequence. Weight the extremes: deterministic physical protection that contains no software, and analytical complexity confined to a plane that cannot write to the process. Evacuate the middle, which is where the tooling sits that can fail in the same conditions that produce the incident. That was Part IV.

What the four together produce is not a conclusion about products. It is a conclusion about what counts as evidence, and the reason I have spent five papers on it is that every allocation decision in this sector is downstream of that question and almost none of them acknowledge it.

7. The choice#

7.1 Declaring my interest#

I run a company that builds simulation tooling for exactly the purpose this series has argued for. I have a direct commercial interest in one of the two approaches described below, and the reader should apply whatever discount that deserves.

I am setting the choice out and I am not closing it. That is a deliberate decision and not a rhetorical device. An argument that ends by telling you which supplier to call has converted itself into a sales document retrospectively, and everything before the last page gets reread in that light. I would rather have the argument stand.

7.2 The conformance approach, on its merits#

The case for the conformance approach is stronger than its critics allow and I want to make it properly.

It finds real defects cheaply. An audit against a serious framework will surface unpatched systems, undocumented connections, expired accounts and missing procedures, at a cost per defect that no other method approaches. These are genuine findings and fixing them genuinely helps.

It is verifiable by a third party. Someone who does not work for you can check the claim and reach the same answer. That property is what makes it usable in contracts, in regulatory filings, in insurance submissions and in board reporting, and no other instrument in this sector currently has it. The value of that is easy to underrate from inside an engineering argument and impossible to replace.

It encodes accumulated knowledge. The frameworks were built from real incidents by people who investigated them. An operator who conforms is starting from a position that took the industry thirty years and a great deal of damage to reach.

It is stable enough to build an organization on. You can hire against it, train against it, write job descriptions and contracts against it, and know that the person you hire from another operator will recognize what you are doing. A method that changes every time the model is retrained does not have that property.

And its failure mode is bounded. When conformance is wrong, you have bought controls you did not need. That is waste. It is not ruin, and the distinction matters more than anything else in this paper.

What it cannot do is the thing Part II established. It measures the presence of specified controls, which is a real quantity and a different one from the exposure of your plant this week. It produces an ordinal result that is sensitive to the effort that produced it. And it is systematically blind to the branches that dominate the tail, because it was constructed from the branches that have already happened somewhere and been written up.

7.3 The simulation approach, on its merits#

The case for it is the four papers preceding this one, so I will state it briefly and spend the space on the objections, which is the part of this comparison that does not usually get written.

It produces a distribution rather than a score, so it can be asked about the tail directly. It ranks candidate actions by their contribution to consequence rather than by their position in a framework. It generates sequences that no analyst wrote down, which matters precisely because the branches that ruin facilities are the ones nobody imagined. And its output is denominated in a way that a budget decision can consume: a probability, a consequence and an interval, rather than a percentage of clauses satisfied.

Now the objections.

The model is built from the same records that Part I said were wrong. A twin is constructed from asset registers, network drawings and configuration exports, and the whole argument of the reference architecture illusion is that those documents diverged from the plant years ago. A simulation over an inaccurate topology produces a confident distribution over a facility that does not exist. This is the deepest problem with the method and no amount of sampling addresses it, because it is an error in the object being sampled.

The outputs look more precise than they are. When a run reports a probability with a confidence interval, the interval describes sampling variation in the model. It does not describe the probability that the model is wrong about the plant, about the adversary or about the people. Model error dominates sampling error in this domain, usually by a margin nobody can quantify, and a number carrying an interval invites a reader to believe the interval is the uncertainty. It is a lower bound on the uncertainty.

It cannot be backtested against the thing it is for. Tail events are rare by construction, so the realized record contains almost no instances of the branches the method exists to estimate. You can validate the mechanics, the traversals, the protocol behavior and the physical consequence models against known behavior. You cannot validate the tail probability against outcomes, because the outcomes are not there. This is precisely the objection Taleb raised against the risk models used in finance [2], and intellectual honesty requires noticing that it applies to a model with a Monte Carlo engine in it just as much as it applied to the ones he was writing about. A simulation is not exempt from the epistemology it was built to serve.

It decays. A twin that is not maintained against the plant drifts away from it at the same rate the drawings did, and it drifts silently. An out-of-date twin is worse than no twin, because it produces numbers with the same confidence as a current one and nothing in the output indicates which you have. The maintenance is not an operating overhead attached to the method. It is the method.

Nobody can audit it. There is no third-party assurance regime for this class of tool, no accepted body of practice for reviewing one, and no way for your regulator, your insurer or your board to check a vendor's claim independently. Compare that with the conformance approach's central strength and the asymmetry is stark. You are being asked to accept a number on the strength of the supplier's competence.

And the survivorship argument applies to it in full. A simulation vendor's reference customers deployed the product and did not have a serious incident. That list is conditioned on the outcome exactly as Part I described, the failures are handled under the same non-disclosure arrangements, and the fact that the vendor's product is philosophically committed to noticing this bias does not exempt the vendor from it. The question Part I proposed for security products is the right question to ask a simulation supplier, and it should be asked of mine.

7.4 What the choice actually is#

It is not a choice between a rigorous method and a lax one. Both approaches are rigorous about something and blind to something, and the blindnesses are not the same shape.

The conformance approach gives you a verifiable statement about a quantity that is not your exposure. The simulation approach gives you an unverifiable statement about a quantity that is. You are choosing which of those two properties you can less afford to be without, and that is a judgment about your position rather than about the methods.

There is a real answer for some operators. If you have a regulatory obligation you must discharge and a limited budget, the conformance approach is not optional and the question of whether to add the other is a question about money you may not have. If you operate a facility where a physical consequence branch would end the business, then an instrument that cannot see that branch is not sufficient no matter how well verified it is, and you have to decide what you will accept in place of verification.

The evidence that should move you, in either direction, is specific. For the conformance approach: how much of your exposure is in scope of the framework you are conforming to, and how would you know. For the simulation approach: how was the topology it runs on established, how often is it reconciled against the plant, what does it claim about model error as opposed to sampling error, and what would falsify its output.

Ask those of anyone selling you either, including me.

If the answers are unsatisfactory in both directions, that is information too, and it is the position most operators in this sector are actually in. It is not resolved by buying something. It is resolved by finding out which of your questions has no answer yet and deciding what you will do while it does not.

8. References#

[1] N. N. Taleb, Fooled by Randomness: The Hidden Role of Chance in Life and in the Markets. New York: Texere, 2001. Cited for the realised path as one draw from a distribution, and for alternative histories as the corrective.

[2] N. N. Taleb, The Black Swan: The Impact of the Highly Improbable. New York: Random House, 2007. Cited for the contribution of rare branches to the mean of a fat-tailed loss distribution, and for the objection that a risk model cannot be validated against events its sample does not contain.

[3] N. N. Taleb, Antifragile: Things That Gain from Disorder. New York: Random House, 2012. Cited for the barbell as an allocation rule, specifically the preference for bounded cost against open-ended benefit over the reverse.

[4] N. N. Taleb, Skin in the Game: Hidden Asymmetries in Daily Life. New York: Random House, 2018. Cited for ruin as an absorbing state, and for the divergence between an average taken across a population and an average taken across time for a single participant.

[5] L. A. Gordon and M. P. Loeb, "The economics of information security investment," ACM Transactions on Information and System Security, vol. 5, no. 4, pp. 438-457, November 2002. Cited for the bound on optimal security investment at 1/e of expected loss, for the two families of security breach probability function under which it is derived, and for the assumptions of risk neutrality, single-period decision, divisible investment and known parameters.

[6] Lloyd's, Market Bulletin Y5381: Cyber-attack exclusions, Corporation of Lloyd's, 16 August 2022. Cited for the requirement that standalone cyber-attack policies written at Lloyd's carry a state backed cyber-attack exclusion at inception or renewal from 31 March 2023, and for the requirements those clauses must meet, including an agreed basis for attribution.

[7] Lloyd's Market Association, model clauses LMA5564 to LMA5567, cyber war and cyber operation exclusions, November 2021. Cited as the model wordings in common market use to meet the requirement in [6], and for the differences between them in what is carved back into cover.

[8] International Electrotechnical Commission, IEC 62443, Industrial communication networks, Security for industrial automation and control systems. Cited as a conformance framework in general use, including by the author, and as a source of obligations that belong in the licence-to-operate account.

[9] International Electrotechnical Commission, IEC 61508, Functional safety of electrical, electronic and programmable electronic safety-related systems. Cited for proof testing and for the meaning of a safety integrity level as a claim about random hardware failure and systematic capability rather than about an adversary.

[10] International Electrotechnical Commission, IEC 61511, Functional safety, Safety instrumented systems for the process industry sector. Cited in the same sense as [9], for the proof-test regime that makes a physical protective function verifiable.

[11] Directive (EU) 2022/2555 on measures for a high common level of cybersecurity across the Union, NIS2. Cited for the responsibilities it places on management bodies in relation to cybersecurity risk-management measures, and as a source of licence-to-operate obligations.

[12] National Institute of Standards and Technology, Framework for Improving Critical Infrastructure Cybersecurity. Cited in the same sense as [8].

[13] North American Electric Reliability Corporation, Critical Infrastructure Protection reliability standards. Cited in the same sense as [8].

Note on this bibliography#

The numbers used in sections 3.3, 3.4 and 3.5 are illustrative and are constructed to show the arithmetic. They are not measurements of any facility and no client figures appear in this paper. Where an arithmetic result is stated, it follows from the stated inputs and requires no source beyond the model in [5].

The field observations throughout sections 4 and 5, including the composition of NOW lists in assessment work, the recurrence of standing vendor access paths in sampled sequences, and the behaviour of NEXT registers over time, are the author's own, drawn from assessment and engineering work in energy, rail, maritime and manufacturing in North America, Australasia and Europe. They carry no citation because no published source records them, and they are offered as testimony rather than as evidence of a general rate.

Section 5 describes the content of [6] and the common structure of the wordings in [7]. Wordings vary between insurers and between policy years, and the description here is not a substitute for reading the policy. The author is neither a lawyer nor an insurance adviser and the section is not advice.

The claims withdrawn in section 5.5 appeared in earlier material issued under the author's own company's name. They are withdrawn because no source supports them, and they are recorded here rather than quietly removed.

Eigenia Labs Open Scientific Publishing Standard
Licensed CC BY 4.0
Exact Verification Audit: 59,230 chars