Fooled by Best Practice, Part III: Running the Paths That Did Not Happen
J. McKenney
Paper 3 of five in Fooled by Best Practice. Part I made the argument in ordinary language. Part II set out the distinction it rests on, between the outcomes a facility can observe and the distribution those outcomes were drawn from. This part gives the method, with its mathematics and its limits. Part IV gives the engineering response. Part V gives the allocation decision.
Licence: CC BY 4.0. 16 September 2026.
| Field | Value |
|---|---|
| Document ID | WG-02-DT-FBP-3 |
| Slug | fooled-by-best-practice-3-running-the-paths |
| Working group | WG-02-DT, Digital Twin |
| Series | Fooled by Best Practice, part 3 of 5 published |
| Author | J. McKenney |
| Published | 16 September 2026 |
| Revision | |
| Status | Published |
| Preceded by | WG-02-DT-FBP-1, WG-02-DT-FBP-2 |
| Followed by | WG-02-DT-FBP-4, WG-02-DT-FBP-5 |
Executive Abstract#
If the risk lives in the histories that did not happen, no care applied to the history that did happen will find it. Incident records read what occurred, audits what is installed, dashboards what is alerting: honest instruments pointed at the wrong quantity.
The only way to look at the unobserved side is to generate it: model the facility, draw many randomized histories, run each to outcome, and read the distribution not a single one. The same plant that rejects a probabilistic statement about its cyber exposure will specify a safety loop whose whole specification is a probability of failing on demand. Choosing this Monte Carlo method declares the quantity a distribution whose uncertainty you show rather than compress into a score.
The method is worthless unless the model is of the actual plant: the firmware as loaded, the configuration as drifted, and the human response as a variable, which decides more outcomes than any other and is left out most. Adversaries build long sequences from individually permitted transitions, each compliant and none the failure, so a model that cannot walk twenty steps across network, software, configuration, people and process cannot see the paths that get used. What survives is a ranking of changes by how much each collapses the paths reaching something that matters.
Abstract#
If defensive adequacy is a property of the distribution of possible histories rather than the realized one, the instrument must sample that distribution, and counterfactual histories must be generated because they cannot be observed. Monte Carlo over a model of the facility is not an implementation choice but an epistemic commitment: treating the object as a distribution, representing uncertainty, and accepting the falsifiability that follows. Across the five steps, adversary, entry, timing, plant state and human response are randomized while the facility is held fixed, making the distribution a property of that plant and permitting comparative use. A twin representing only the designed architecture yields a precise distribution over a facility that does not exist, so the model must carry deployed state, actual reachability and human response as a distributed variable. The combinatorics of length-twenty walks over the facility graph require guided sampling with likelihood-ratio correction, not enumeration. The limits are stated plainly: the prior is an assumption, importance sampling needs prior knowledge of where the tail lies, unrepresented mechanisms receive probability zero, calibration is largely unavailable at these event rates, and a distribution is not a forecast. The claim that simulation finds paths no analyst conceived is corrected to compositions of represented elements no analyst enumerated.
1. Introduction#
1.1 The instrument has to point at the quantity#
I have spent a large part of my working life producing assessments of facilities, and for most of that time the instrument I reached for was a control assessment against IEC 62443 [4]. It is a good instrument. It finds real defects, it finds them at reasonable cost, and I would still run one tomorrow on a plant I had never seen.
What changed for me was not a loss of faith in the instrument. It was noticing what the instrument was measuring. A control assessment reads the configured state of a facility against a catalog of controls known to matter across a class of installations. It answers the question of whether specified things are present. It is a careful reading of the observed side of the table.
So is everything else on the shelf. An incident record reads what occurred. A detection platform reads what is alerting now. A penetration test reads what one team found in two weeks with a defined scope. Each of these is an honest instrument. Every one of them takes its reading from the side of the table where the risk does not live.
Part II set out why that matters, and I am not going to re-argue it here. The claim I am starting from is that the thing an operator wants to know is a property of the distribution of possible histories of their plant, and that the history they actually got is a single draw from it.
If that is right, then the instrument has to sample the distribution. And because the counterfactual histories of a facility do not exist anywhere to be read, the only way to sample them is to make them.
1.2 What this part does#
This part gives the method. It says what it means to generate a facility's counterfactual histories, what is randomized and what is held fixed in each one, and what a single run actually represents. It states, as strongly as I can, what the underlying model has to contain for any of the sampling to be meaningful, because this is the point at which most of the work in this field goes wrong. It explains why a long walk across a graph that crosses the boundaries between network, software, configuration, people and physics is the substance of the approach rather than an implementation detail.
And then it spends a long section on what the method cannot tell you. That section is not a disclaimer. It is the part I would read first if someone handed me this paper, because a method that produces numbers and does not state its own limits is doing exactly what Part I accused the compliance score of doing, at higher resolution and with more authority.
Two things I am not doing. I am not making the engineering argument about what should be deployed in the plant, which is Part IV's subject. I am not making the budget argument, which is Part V's. This part is about measurement.
One admission before I start. I build this kind of system. Where I describe something as a requirement of the approach I mean it is a requirement of the approach, and where I describe something as a particular construction of mine I will say so and say what its evidential standing is. I would rather a reader take the requirements and reject my implementation than accept the implementation on my say-so.
2. Monte Carlo as a position, not a feature#
2.1 Why you cannot integrate and have to sample#
The formal reason for Monte Carlo is not philosophical. It is that the integral cannot be done.
If a facility's exposure could be written as a tractable function of a handful of parameters, you would write the function and evaluate it. Reliability engineering does exactly that for simple series and parallel systems, and the answers are exact and cheap. The reason security exposure in a real plant does not yield to that treatment is structural. The state space is high dimensional and discrete. Outcomes depend on the order in which things happen, not only on which things happen. The variables are strongly and conditionally dependent, so the failure of one control changes the effective strength of several others. And some of the most consequential variables, the behavior of people under load in particular, are not stationary over the period you care about.
For a process like that there is no closed form to evaluate. There is, however, a process you can run forward. That is the exact condition under which Monte Carlo is the correct tool rather than a compromise: you can simulate the process but you cannot integrate it, so you estimate the integral by sampling the process.
That is the ordinary situation in applied probability over complex systems. It is not an exotic move.
2.2 The sector already accepts probability where it matters most#
There is an objection I meet often enough that it deserves an answer early. It goes: probability is not how we run plants, we run plants on determinism and engineering margin.
That is not true, and the people who say it usually know it is not true.
Functional safety in this sector is specified probabilistically and has been for a long time. A safety integrity level under IEC 61508 and its process-sector application IEC 61511 is a band of average probability of failure on demand [5], [6]. Specifying a loop at SIL 2 in low-demand mode is a statement that the protective function is expected to fail when called on with probability between one in a thousand and one in a hundred. Nobody regards that as unrigorous. It is the most rigorous thing in the plant.
The same engineer who signs that specification will then be told that his facility is eighty-seven percent compliant and accept it as a risk statement, when it is not a probability of anything and carries no interval.
So the sector does not object to probability. It objects, reasonably, to probability it does not trust. That is a question about the quality of the model and the honesty of the reporting, which is what sections 4 and 6 of this paper are about. It is not a question about whether probabilistic statements belong in a control room.
2.3 What choosing the method commits you to#
Treating Monte Carlo as a product feature to be compared on a datasheet misses what it is. Adopting it commits you to four things, and they are uncomfortable in ways a score is not.
The first is that the object of interest is a distribution. You are giving up the single number that fits in a board pack. What replaces it is a shape, with a body and a tail, and the tail is the part you convened the meeting about.
The second is that your uncertainty gets shown rather than hidden. A compliance score of eighty-seven has no error bar, which reads as precision and is in fact the absence of any statement about error at all. A simulated probability with an interval around it is a less comfortable object precisely because it is telling you something true about how much you know.
The third is that your assumptions become inspectable. Every input distribution is a choice somebody made and can be argued with. A score hides its assumptions inside a scoring rubric that nobody reads as an assumption. Taleb's objection to the risk systems of the financial industry was not that they used mathematics, it was that the mathematics was arranged to confirm the comfortable reading rather than to expose the exposure [1]. A method whose priors are stated on the same page as its output is harder to use that way.
The fourth is cost. Building a model with enough fidelity to sample from is more work than running a scan, by a wide margin, and most of the work is not in the mathematics. It is in the tedious business of finding out what is actually installed.
I regard all four as arguments in favour. A method that is expensive, that produces an uncomfortable output, that exposes its own assumptions and that can be shown to be wrong is behaving like an instrument. A method that is cheap, produces a reassuring number, buries its assumptions and cannot be falsified is behaving like a ritual.
3. Generating a facility's counterfactual histories#
3.1 The five steps, on a plant#
The structure of a Monte Carlo simulation is standard and has five steps: define the domain of inputs, define the probability distribution over that domain, generate samples from it, run a deterministic process on each sample, and aggregate the outputs. The steps are not the interesting part. What each one means when the object is a facility is the interesting part.
Define the domain. The inputs are not "attacks". An attack is an output. The input space is the set of conditions under which an attempt occurs, and it has at least five groups of coordinates. Adversary properties: capability tier, objective, tooling available, tolerance for dwell time, willingness to be noisy. Entry opportunity: which of the available initial access routes is the one that opens, whether that is a contractor connection, an exposed remote access service, a supply chain component, or physical presence. Timing: which shift, which season, whether the plant is in a maintenance window with interlocks bypassed and vendors on site. Plant state at that moment: the patch state as of that date, which loops are in manual, which temporary configuration is currently in place. And the human draw: who is on call, what the alert load has been that week, how many of the last hundred alerts resolved to nothing.
Define the distribution. This is where the honesty of the whole exercise is decided, and I am deferring it to section 6 rather than glossing it here. For now the important structural point is that these coordinates are not independent. Maintenance windows correlate with vendor presence, which correlates with the initial access route being open, which correlates with the alerting baseline being noisy because half the plant is being worked on. A joint distribution that treats these as independent will systematically understate the probability of the conjunctions that actually produce incidents, because the conjunctions are exactly what correlation produces.
Generate samples. Draw N joint samples from that distribution. What N needs to be is an arithmetic question with an unpleasant answer, and I will give it in 3.3.
Run the process. For each sampled set of conditions, propagate through the model of the facility. Given this adversary, entering here, at this time, against this plant in this state, with these people responding, how far does it get and what does it reach. This step is where model fidelity is everything, and it is section 4.
Aggregate. The output is a distribution over outcomes, and the outcomes have to be stated in terms the plant recognizes. Not "risk score improved". Reached the supervisory network. Achieved persistence in the control zone. Reached a function with a safety role. Drove the process outside its licensed operating envelope. Reached a state from which there is no stable recovery without an outage.
3.2 What is randomized, what is held fixed, and what one run is#
The distinction between what varies and what does not is the thing people get wrong most often when they first look at one of these runs, so I want to be exact.
What is randomized is the circumstance. The adversary, the timing, the entry, the human draw, and the stochastic elements of propagation.
What is held fixed is the facility. The topology, the asset inventory with its actual states, the dependency graph, the placement of controls, the physics of the process and the declared failure thresholds. These are fixed across all N runs because they are the object under study. The distribution you get out is a property of this plant, conditional on the priors you chose.
That is what makes the method useful for a decision. Hold the priors fixed, change one thing about the facility, re-run, and read the difference in the distribution. The difference is the value of the change. I will argue in 6.7 that this comparative reading is considerably more defensible than the absolute number, and that the absolute number is the part most likely to be quoted.
A single run is one counterfactual history. It is an internally consistent story in which a particular adversary arrived at a particular moment against the plant in a particular state and the people responded in a particular way, and the story was generated by the model's rules rather than written by anybody.
That last clause is the whole difference from a tabletop exercise, and I want to be fair to tabletops because I have run many and they are worth running. A tabletop samples the imagination of the people in the room. The scenarios are authored, which means they are drawn from the set of things those people already thought were worth worrying about, which is very close to the set of things that are already controlled. A generated run is drawn from the model's joint distribution, which includes combinations nobody proposed because nobody was thinking about that conjunction. Ten thousand generated runs will contain conjunctions that no workshop would have produced. That is the gain, and it is a real one, and section 6.3 states precisely how far it extends and where it stops.
3.3 How many runs, and the arithmetic nobody quotes#
Suppose you want to estimate the probability that a campaign reaches a function with a safety role. You run independent simulations and count the fraction that do. That estimator is unbiased, and its standard error is
Put numbers in it. If the true is 0.02 and you run 1,000 simulations, the standard error is about 0.0044, so a rough ninety-five percent interval is 1.1 percent to 2.9 percent. That is a usable answer. You would not use it to the second decimal place, and anyone who reports 2.1 percent from a thousand runs without the interval is reporting three significant figures of which one is real.
Now make the event rarer, which is the case you actually care about. If the true is 0.0001 and you run 1,000 simulations, the expected number of runs in which the event occurs is 0.1. Most executions of that experiment return zero occurrences. Zero occurrences in a thousand draws is consistent with a true probability anywhere from zero up to about three in a thousand. You have learned essentially nothing about the quantity you convened the project to estimate.
The naive response is more runs. To get a ten percent relative standard error on by plain sampling you need on the order of runs, and each run is a full propagation across a large graph, not a coin flip. For rarer events the requirement grows as and becomes infeasible before it becomes interesting.
The correct response is to stop sampling uniformly. Instead of drawing from the true input distribution , you draw from a proposal distribution that deliberately over-represents the conditions under which the rare outcome occurs, and then correct each sample by its likelihood ratio so the estimate remains unbiased:
This is importance sampling. Done well it reduces the variance of a rare-event estimate by orders of magnitude and makes the problem tractable. It also introduces a dependency that I regard as the most serious limitation of the entire method, and I have given it its own subsection in section 6 rather than burying it here.
The point for this section is narrower. The number of runs is not a specification you pick because it sounds substantial. It follows from the rarity of the outcome you are trying to estimate and the precision you need, and a vendor who quotes a run count without quoting the resulting interval has told you about their compute budget and nothing about their answer.
4. What the model has to represent#
4.1 The failure mode that makes all of this worthless#
Here is the sentence that the rest of this section exists to support.
A twin that models the designed architecture samples the counterfactual histories of a facility that does not exist.
This is not a fidelity quibble. It is a category error with a specific and nasty consequence. The sampling machinery does not know that the graph it has been given is wrong. It will run ten thousand histories against the drawing, aggregate them correctly, produce a properly formed distribution with a properly computed interval, and hand you a precise answer about an imaginary plant. The precision is real. The precision is about the model.
And that output is worse than no output, because it carries authority. A compliance score at least announces itself as a compliance score. A probability distribution with a confidence interval announces itself as a measurement, and the people receiving it will treat it as one. If the underlying graph is the reference architecture, the whole apparatus has done nothing except reformat the reference architecture illusion of Part I into a form that is harder to argue with.
So the requirements below are not a specification of a nice implementation. They are the conditions under which the arithmetic in section 3 means anything at all.
4.2 Deployed state, not catalog state#
The model needs two layers where most tools have one.
The first is the reference layer: what a given product is, what its designed configuration looks like, what its published vulnerabilities are, and how a newly published vulnerability inherits down to every instance of that product. This layer is genuinely useful and it is what most asset databases actually hold.
The second is the deployed layer: what is physically installed, where, with what serial number, running which firmware revision, with which configuration currently loaded, patched to what date, in what operational state. Part II treated the gap between these two as the concrete form of the map and territory problem, and I will not re-derive it. What I will insist on is the modeling consequence.
Firmware state in particular decides outcomes. In operational technology, firmware is frequently years behind the vendor's current release, and the reason is not negligence. It is that the controller governs a process that cannot be stopped, that the update has not been validated against the safety case, or that the vendor who would perform it went out of business. I have walked plants where the installed revision on a critical controller predated the engineer showing it to me. If the model takes its firmware version from the catalog, every vulnerability calculation downstream of it is a calculation about a plant that was commissioned in a design office and never built.
The same applies to configuration. The state of a facility is the accumulated product of every operational decision since commissioning, and those decisions are individually defensible and collectively enormous. The model has to carry the drift as data, not as an error term.
4.3 Software reality, several levels down#
A vulnerability scan reports on components it can identify. The code actually executing on a device includes everything those components depend on, and everything those dependencies depend on, and so on down. In my experience the ratio between the components a tool names and the components actually present is not close to one, and the difference is concentrated exactly where nobody is looking.
A software bill of materials in a structured format, SPDX or CycloneDX, is the mechanism for carrying this, and the requirement is that the analysis follows the transitive closure rather than the direct dependencies. A vulnerability five levels down in a dependency tree is reachable by the same execution that reaches the top-level component. It is simply not visible to anything that reads the manifest and stops.
Two honest qualifications. Depth of dependency does not by itself imply exploitability. Whether a given vulnerable function is reachable in the particular configuration the plant runs is a separate question, and treating every transitive vulnerability as live will produce a model that overstates exposure so badly that nobody believes its output. Exploit probability estimates, of the kind published as EPSS scores, are useful for weighting this, and they are themselves model outputs with their own error, not observations. If you feed a model output into a model as though it were data, you have to carry its uncertainty with it, and most pipelines do not.
4.4 Connectivity as it is#
Every facility I have assessed has a network diagram and the diagram is a subset of the reachability. The model needs the reachability, derived from configuration rather than from the drawing: routing and firewall rules as loaded, not as designed, dual-homed machines, engineering workstations that sit on both sides of a boundary because that is the only way the engineer can do the job, historians that reach into the control zone to collect and into the corporate zone to publish, remote access paths opened for a commissioning activity and never revoked, wireless and cellular links installed by a subcontractor, and shared credentials that make two nominally separated zones into one zone from the adversary's point of view.
This is not an exotic list. It is the ordinary condition of a plant that has been running for fifteen years and doing useful work the whole time.
4.5 Adversary capability, with its own known bias#
The model needs a representation of who might attempt this and what they can do: capability tiers, typical technique repertoires, sector targeting, campaign activity, objectives that make a given facility interesting. Frameworks that catalog adversary techniques are the right substrate for this.
I want to flag the problem with it here and take it apart properly in section 6.6. Every catalog of adversary behavior is built from campaigns that were detected and then described. Campaigns that succeeded quietly and campaigns that were detected under a non-disclosure agreement contribute nothing. Part I called this the second mechanism by which evidence goes missing, and it means the threat prior in a simulation intended to sample the unobserved side of the table is itself sampled from the observed side. That is a real contamination and it does not fully wash out.
The partial mitigation is to model capability rather than only history. Rather than asking only which techniques have been seen against this sector, ask what a resourced adversary could do given this specific graph, and let the propagation find it. That widens the support of the distribution beyond the catalog. It does not make the prior correct.
4.6 The physical process and what counts as a bad outcome#
A simulation that terminates at "the adversary obtained code execution on a controller" has not produced an outcome anybody can act on. The consequence is a physical event, and the model needs enough of the process to reach it: the control loops and their setpoints, the interlocks and what they are keyed to, the protective functions and their demand rates, the thermodynamics or the electro-mechanics, and the failure thresholds that the facility itself has declared in its own safety case.
That last point matters for the honesty of the exercise. The failure thresholds are the plant's own numbers. The topology is the plant's own topology. The component list is the plant's own inventory. The model takes what the facility asserts about itself and derives consequences from it by stated relations. None of it is a field measurement made by the model, and every conclusion stays open to challenge by challenging the inputs or the relations. That is the correct standing for this kind of work and it should be said out loud rather than allowed to blur into a claim of observation.
The outcome class that matters most is also the hardest to state. It is not "damage occurred". It is the transition to a regime with no stable operating point, where the process cannot be returned to a safe condition by any control action available to the operators. In dynamical terms this is a saddle-node bifurcation, whose normal form is
with two fixed points while , one stable and one unstable, which approach each other as rises and annihilate at . Beyond that there is no attractor. The feature of this that should worry anyone running a plant is that the distance to the transition scales as , so the warning time compresses non-linearly on approach. A system can drift towards that point for months and show observable signs only in the final fraction of the approach. A consequence model that can only express degrees of damage cannot express this outcome at all, and this is the outcome the facility exists to prevent.
4.7 The human response system, which is the one that gets left out#
An attack path is realized only if nobody interrupts it. That sentence is trivially true and almost universally ignored in modeling, because the probability of interruption is treated as a constant folded into a detection coverage figure.
It is not a constant. It is a random variable with a wide distribution and strong dependence on exactly the conditions that correlate with attacks succeeding.
Consider what actually determines whether an alert becomes a response. Who is on shift. How many alerts they have handled that week and how many resolved to nothing, because an analyst whose last ninety-seven alerts were noise is running a prior that makes the ninety-eighth look like noise too, and that prior is rational. Whether the escalation path terminates in someone who is awake. Whether the plant is in a maintenance window where anomalous behavior is expected and therefore invisible. Whether the organization is in a period of high turnover, budget contraction or leadership distraction, which changes the coherence of collective response in a way that is not reflected in any headcount figure. Part II treats the phase behavior of that last point and I defer to it.
The requirement is therefore: represent the response system as a distributed variable, sample it jointly with the technical conditions rather than independently, and let the correlations stand.
The specific apparatus I use for this is my own and I will name it as such. I represent each person in the response system as a vector built by combining two personality instruments, the Big Five and DISC, as a tensor product, so that each pairing of a behavioral tendency with a trait becomes its own coordinate and the person becomes an object with a distance, a direction and a trajectory under stress. I then model the responding team as a system with a pairwise interaction term, so that the friction between two incompatible profiles is a quantity that grows as the load grows rather than a fixed property of the org chart. When the simulation puts that team under a high-tempo incident with incomplete information, the interaction terms determine how much of the available response capacity is converted into coordination friction rather than into action.
I find it useful. I would not defend it as a measurement. The Big Five has a substantial validation literature behind it and DISC has considerably less. The mapping from either instrument to incident response behavior under fog of war is an assumption I have made, not a relationship anyone has established. A reader who rejects the entire construction should still accept the requirement underneath it, which does not depend on my formalism at all: the probability that a human interrupts an attack path is a variable with a distribution, and a model that replaces it with a constant has assumed away the factor that decides most real outcomes.
4.8 Fidelity should not be uniform#
One practical note, because the list above reads as a counsel of perfection and no facility will produce all of it.
Errors in the inputs do not matter equally. Some inputs the output is highly sensitive to and some it is nearly indifferent to, and which is which is not obvious in advance. It is, however, computable: vary each input across its plausible range, re-run, and see how far the output moves. Spend the data collection effort where the sensitivity is high and accept crude representations where it is low.
This is ordinary engineering practice and it changes the character of a deployment. The question stops being "do we have complete data" and becomes "do we have adequate data for the specific inputs this conclusion depends on", which is answerable, and which also tells you honestly when the answer is no.
5. Why the walk is long, and why it crosses domains#
5.1 Single-hop questions find single-hop failures#
A control assessment asks whether a given control prevents a given technique. A vulnerability scan asks whether a given host has a given defect. A penetration test asks whether a team can get in within a scope and a fortnight. All of these are one-step or few-step questions, and they find one-step and few-step failures, which is worth doing because those failures exist and are cheap to fix.
Adversaries do not operate in one step. The path from a compromised contractor laptop to a function with a safety role is a long sequence of transitions, and the important property of that sequence is that no individual transition is a violation. The contractor is permitted to connect. The remote access service is permitted to accept the connection. The jump host is permitted to reach the historian. The historian is permitted to reach into the control zone because that is how it collects. The engineering workstation is permitted to write to the controller because that is what an engineering workstation is for. The controller is permitted to actuate the valve, since that is the entire point of the controller.
Every hop passes its own assessment. The composition is the failure, and no instrument that evaluates hops one at a time will see it.
5.2 The facility as a graph and the attack as a walk#
State it formally and the requirement becomes obvious.
Represent the facility as a directed graph whose nodes are not assets but triples of asset, privilege level and state, and whose edges are permitted transitions, each carrying a conditional probability of success given that the adversary has reached the source node with the relevant capability. An attack is a walk on this graph from an entry node to a consequence node. The question the operator needs answered is not whether any individual edge is dangerous. It is whether a walk exists from some entry to some consequence within the adversary's budget of time, capability and tolerance for detection, and with what probability.
That reformulation is the method. Everything in section 4 is in service of getting the graph right, and everything in section 3 is in service of estimating probabilities over walks on it.
5.3 Why you cannot enumerate, and therefore must sample#
The reason this is a sampling problem rather than a search problem is combinatorial and it is easy to state.
If the graph has average out-degree , the number of distinct walks of length from a given node is on the order of . Take a modest of ten and a walk length of twenty and you have on the order of walks. There is no enumeration of that set, on any hardware, ever.
Set that against the number of paths a group of experienced people will construct in a day-long workshop, which in my experience is somewhere between ten and fifty. The workshop is not doing badly. It is doing what a human process can do. The gap between and is the entire reason this method exists.
There is a second piece of arithmetic here that changes how people think about long paths, and it is worth doing explicitly. Suppose a twenty-hop path where each hop succeeds with probability one half. The path completes with probability , about one in a million. That sounds like something you can ignore. But the number of such paths is large, and if there are on the order of distinct paths of that shape leading to consequence nodes, the expected number that complete is on the order of a thousand.
I want to be careful with that number rather than let it do more work than it can. Those paths are not independent, they share edges heavily, and the expected count of successes is not the probability that at least one succeeds. The calculation is an order-of-magnitude argument and nothing more. But the order of magnitude is the whole point: per-path improbability does not imply aggregate improbability, and an instrument that dismisses a path because that path is unlikely has made an error about a sum.
5.4 Guided walking, and the honest cost of it#
Random walking on a graph of this size is close to useless. A uniformly sampled walk spends its length wandering through low-consequence regions and almost never terminates anywhere interesting, which is another way of saying that the estimator has enormous variance.
So the walk is guided. The proposal distribution weights transitions by plausibility given the adversary's capability and objective, and biases towards regions of the graph that lead to consequence nodes. Then each sampled walk is re-weighted by its likelihood ratio, as in section 3.3, so the estimate is corrected back to the true distribution.
This works. It is standard, it is what makes rare-event estimation on large graphs tractable, and it is the difference between a method and a compute bill. It also means the answer depends on the proposal distribution in a way that is not visible in the output, and section 6.2 is about exactly that dependency.
5.5 One graph, because the important edges cross the boundaries#
The reason all of this has to be one graph rather than several coupled tools is the single most practical argument in this paper.
Take a path and label the domains it passes through. Hop one is a person clicking something, which is psychology. Hop four is a defect in a library four levels down a dependency tree, which is software composition. Hop nine is a credential shared between two zones because a maintenance procedure required it, which is configuration. Hop fourteen is a trust relationship between an engineering workstation and a controller, which is network and identity. Hop eighteen is a write to a setpoint that the process cannot tolerate, which is physics.
Now partition the model by domain, which is how the tooling market is actually organized. The vulnerability tool owns hop four. The identity tool owns hop nine. The network tool owns hop fourteen. The awareness programme owns hop one. Nothing owns the process, usually.
The edges that cross between those partitions have nowhere to live. They are not in any of the tools because each tool's model stops at its own boundary, and they are not in the integration layer because an integration layer correlates alerts rather than representing transitions. Those crossing edges are precisely the ones that make the long path possible. A model that is partitioned by domain is structurally incapable of representing the mechanism that produces the outcomes it was bought to prevent.
5.6 Different layers move at different speeds#
There is a technical consequence of putting everything into one graph that I have to address, because it is the part that is easiest to get wrong.
The layers of that graph have radically different time constants. New vulnerability disclosures arrive hourly. Adversary campaign activity shifts over days. Organisational stress, staffing and alert fatigue move over weeks and months. Network topology and installed equipment move over years. If you propagate state across the whole graph at a single rate, one of two things happens. Either the fast, noisy layers overwrite the slow ones, so a spike in geopolitical reporting contaminates your estimate of the response team's coherence, which it has no business touching on that timescale. Or you damp everything to protect the slow layers, and the model stops responding to a live disclosure that genuinely changes the picture today.
The requirement is that propagation across layers be gated by timescale, so that each signal moves at the rate its underlying process moves and information crossing between layers is filtered rather than merged.
The mechanism I use for this is a gated graph neural network with per-layer gates controlling what is allowed to propagate between layers at each step, which is an architectural choice I have made and built, not a published benchmark and not a claim about the state of the art. A reader should take the requirement, which is that the timescales be separated, and treat my particular network as one way of meeting it. There is a further caveat that belongs with it: learned propagation needs training data, and in this domain the labeled data is thin. Most of what such a model does well is structural, and where it is doing something learned the amount it has learned from is small enough that I would not build a conclusion on that component alone.
6. What the method cannot tell you#
This is the section that decides whether the rest of the paper was worth writing. I have read a great deal of vendor material on simulated risk and the limits section is almost always absent, occasionally present as a sentence about how no model is perfect, which is a way of raising the subject in order to close it.
These limits are specific, they are severe, and knowing them changes what you should do with the output.
6.1 The distribution is an assumption, not a measurement#
The output of the method is not the probability of compromise. It is the probability of compromise conditional on this model and these priors. The conditioning bar is carrying the entire weight of the exercise, and it is the first thing to fall off when a result travels from the analysis to a slide.
The priors are the weak point. The technical inputs, what firmware is installed and what a dependency tree contains, are at least in principle observable, and the effort to observe them is the effort described in section 4. The behavioral inputs are not observable in the same way. There is no census of adversaries. Nobody can tell you the rate at which capable actors initiate campaigns against mid-sized water utilities in a given region, because the numerator is unknown for the reasons Part I gave and the denominator is definitional. Every figure of that kind in any model, including mine, is an estimate constructed from partial reporting and judgment.
Garbage in. The phrase is old because the failure is old. A simulation does not purify its inputs. It propagates them, and it propagates them with enough machinery in between that the relationship between an assumed input and a reported output is no longer legible to the person reading the output. A compliance score is wrong in ways you can see. A simulated probability is wrong in ways that are laundered.
The defence is not better priors, since better priors are not available. The defence is to report the priors next to the result, to show how far the result moves across the plausible range of each one, and to treat any conclusion that flips inside that range as an artifact of assumption rather than a finding. I would put it as a rule: report the interval and the priors together, or do not report the number.
6.2 Rare events are rare inside the simulation, and importance sampling has a catch#
Section 3.3 gave the arithmetic. An event of true probability appears about times in draws, and when is small the estimate is noise. That is the ordinary version of this problem and importance sampling is the ordinary answer.
The catch is deeper than the compute cost, and it is the thing I would most want a reader to take from this section.
Importance sampling works by biasing the proposal distribution towards the region where the rare event occurs. To do that, you need to know approximately where that region is. You are constructing a proposal distribution out of your existing beliefs about which conditions are dangerous.
Now consider the tail you most need to find: the one nobody has thought about. By construction, your proposal distribution puts little mass near it, because you built the proposal out of the things you already suspected. You will sample intensively and efficiently in the neighbourhood of the dangers you already knew about, produce a tight interval, and report a confident number about a region that was never the problem.
This is not an argument against importance sampling, which remains necessary. It is a statement that the method's efficiency at finding rare events is conditional on its prior about where rare events are, and that this conditionality is invisible in the output. The variance reduction is real and the coverage guarantee is not. Anyone who tells you their simulation searches the tail should be asked which tail, and how the proposal was constructed, and what happens to the answer if it was constructed differently. Taleb's work on estimating tail quantities from finite samples is the general statement of the underlying difficulty, and it does not become easier because the samples are synthetic [3].
6.3 The model cannot contain what it does not represent#
Every mechanism absent from the graph has probability zero in the output.
A technique the model does not encode. A component the bill of materials does not list because it arrived in firmware rather than through a package manager. A physical access route through a door that a contractor props open. An insider with legitimate authority whose actions are indistinguishable from work. A supplier compromise upstream of anything the facility can see. A failure mode of the process that the plant's own safety case did not anticipate, which is a real category and not a rare one.
A probability of zero produced by a model is not a probability of zero in the world. It is the model's silence rendered as a number, and a number is more persuasive than a silence. This is the objection Taleb makes to induction over an unrepresentative sample [2], turned back against the method that was meant to answer it, and I do not think there is a clean escape from it. Widening the model helps at the margin and the margin recedes as you approach it.
There is one specific claim in this field that I want to correct, including in my own earlier writing, because it is the point at which overreach usually enters.
It is often said that a Monte Carlo attack simulation finds paths that no analyst has conceived. As stated, that is false. What it finds is compositions of represented elements that no analyst enumerated. The elements were all in the model, put there by people. The novelty is combinatorial, arising from the gap between and in section 5.3.
That capability is real, it is valuable, and it is the honest case for the method. It is also categorically different from discovering a mechanism that is not in the model, which the method cannot do, and which is where black swans actually come from. Combinatorial novelty is not genuine novelty. Conflating the two is how a useful instrument gets sold as an oracle.
6.4 A distribution over futures is not a forecast of the one you get#
If the output says eight percent, the facility does not experience eight percent of a compromise. It experiences one history. The distribution is a statement about a population of possible futures, and you are issued exactly one draw from it.
This cuts in both directions and the second direction is the one people miss. A low number is not an assurance, because a low-probability outcome that occurs is not evidence the model was wrong; that is what low probability means. And a high number followed by an uneventful year is not evidence that the model was wrong either. The single realized outcome is very nearly uninformative about the distribution it came from, which is the argument of Part II applied to the method's own results, and it applies with full force.
The result also decays. It is conditional on a snapshot of the facility, and the drift described in section 4.2 resumes the moment the snapshot is taken. A maintenance window, a vendor visit, a disclosure against a component deep in the dependency tree, a resignation in the control room: each of these moves the graph and therefore the distribution. Any report of this kind should carry its snapshot date and an honest statement about the rate at which it goes stale, and for most plants that rate is faster than the annual cycle on which such work is typically commissioned.
6.5 Calibration is mostly unavailable, and that should be said plainly#
A probabilistic weather forecast can be scored. You issue thousands of forecasts, you observe thousands of outcomes, and you can check whether the events you called thirty percent occurred about thirty percent of the time. That is calibration, and it is what makes a probabilistic forecast trustworthy rather than merely quantitative.
You cannot do this for catastrophic facility outcomes. The event rate is far too low to accumulate a scoreable sample within any period over which the facility, the threat environment and the model all stay constant. By the time you had enough realizations, none of the conditioning would hold.
So the validation available is indirect, and it is weaker. You can show generated paths to the people who run the plant and ask whether each hop is physically possible, which catches gross errors in the graph and catches nothing subtle. You can back-test against incidents that did occur where the facts are known, which is a small sample and a selected one. You can run sensitivity analysis, which tests the model's internal structure rather than its correspondence to the world. You can check internal consistency.
All of that is worth doing and none of it is calibration. I think the honest description is that these models are structurally argued rather than empirically validated, and that their outputs should be read as the consequences of stated assumptions rather than as measurements. Anyone offering you a probabilistic risk figure for an industrial facility and describing it as validated should be asked the obvious question: validated against what, over what sample, across what period. I have asked it. As with the vendor question in Part I, the refusals are informative.
6.6 The threat prior comes from the side of the table the method was built to escape#
This is the recursive problem and I flagged it in 4.5.
The method exists to sample the unobserved side. Its adversary prior is assembled from catalogd campaigns, which are campaigns that were detected, attributed and published. The campaigns that succeeded without detection are not in it. The campaigns detected under commercial confidentiality are not in it. Part I described both mechanisms and neither is fixable by anyone operating outside a national intelligence apparatus.
So the instrument is partially contaminated by exactly the bias it was constructed to correct. Its technical layers genuinely sample the unobserved side, because the graph contains compositions nobody enumerated. Its threat layer is drawn from a survivorship-selected record.
Capability-based adversary modeling mitigates this, as section 4.5 described, by asking what a resourced actor could do against this graph rather than only what has been observed being done. The mitigation is partial. A structural bias in a prior does not disappear because you have widened the support around it. It should be stated as a standing limitation of the method rather than a defect of a particular implementation.
6.7 What survives: the comparative result#
Everything above attacks the absolute number, and I think the attacks land. What survives is the comparative use, and it survives for a specific reason.
Take two runs against the same facility with the same priors, differing only in one change: a segmentation cut, a controller removed from a reachable path, a firmware campaign on one class of device, a change to the escalation procedure. The errors in the priors are present in both runs. In the difference between the two distributions, the shared errors largely cancel.
This is the same reason a differential measurement is more trustworthy than an absolute one in any instrument with systematic error. The absolute reading inherits the full bias. The difference between two readings taken under the same bias does not.
The practical conclusion is that the defensible output of this method is a ranking of candidate changes by how much each one collapses the set of paths reaching a consequence, together with the cost of each change. It is not a certificate that the facility's probability of catastrophic compromise is some specific figure. I have watched the second of these be quoted from a report that only supported the first, and the fault was in the reporting rather than in the reader.
6.8 Precise numbers carry authority they have not earned#
The last limit is about how the output behaves once it leaves the room.
A statement of the form "in 8.4 per cent of simulated campaigns the adversary reaches a function with a safety role, with a ninety-five per cent interval of 7.1 to 9.8 per cent" reads as a measurement. The formatting is the formatting of a measurement, the interval is the furniture of a measurement, and none of that is an accident of typography. It is the output of a model whose behavioral inputs are estimates, whose threat prior is selected, and whose proposal distribution encodes the analyst's prior belief about where danger lies. The numbers here are illustrative of the form of a result and are not the output of any particular facility, and even written that way I am aware they read as more solid than the sentence around them.
Taleb's charge against the risk systems of the financial industry was not that they produced numbers. It was that producing numbers of exactly this shape allowed an institution to feel it had measured something [1]. The apparatus for computing a value at risk was elaborate, internally consistent, and blind in precisely the region that eventually mattered, and its elaborateness was part of why nobody looked.
I do not think that charge exempts this method. I think it applies to it directly, and that the only protection is procedural: state the priors on the page with the number, state what the model does not represent on the page with the number, and refuse to report a figure with more significant digits than the run count supports. None of that is difficult. It is simply unattractive, because it makes the deliverable look less like an answer.
7. What an honest result contains#
I will end with the form of the deliverable, because the form is where most of the failures in this section become visible or get hidden.
The outcome variables should be stated in the facility's own terms, not in the security industry's. Reached the supervisory network. Achieved persistence in the control zone. Reached a function with a safety role. Drove the process outside its licensed envelope. Reached a state with no stable recovery. An operator can act on those and can dispute them. Nobody can act on a risk score.
Every figure carries its interval, and the interval is computed from the run count rather than asserted. Where the run count does not support an estimate, the correct output is a statement that the estimate is not supported, which is a legitimate finding and a rare one to see in print.
The priors are printed alongside the results, with the sensitivity of each conclusion to each prior. A conclusion that survives the plausible range of its inputs is a finding. A conclusion that flips inside that range is an assumption, and should be labeled as one in the same document rather than in a methodology appendix.
The snapshot date is stated, with a statement of what would invalidate the result, since drift resumes immediately and some of the drift is scheduled.
The results are presented comparatively wherever a decision is at stake: this change collapses this fraction of the paths to this outcome at this cost, ranked against the alternatives, for the reason given in 6.7.
And the report contains an explicit list of what was not represented. Not a disclaimer about the limits of modeling in general, which is a way of saying nothing at length, but a specific enumeration: these asset classes have no software bill of materials, this process area is modeled only to the interlock and not beyond, the physical access routes were not in scope, these two subsidiaries were not surveyed, the adversary prior was constructed this way and here is what it excludes.
That list is where every zero in the output comes from. It is the most important page in the document, it is the page that tells a reader which of the confident numbers to distrust, and in the reports I have seen from this industry, my own earlier ones included, it is the page that does not get written.
8. References#
[1] N. N. Taleb, Fooled by Randomness: The Hidden Role of Chance in Life and in the Markets. New York: Texere, 2001. Cited for the use of simulated alternative histories as a corrective to reasoning from a single realised path, and for the argument that elaborate risk apparatus can function to conceal exposure rather than to reveal it.
[2] N. N. Taleb, The Black Swan: The Impact of the Highly Improbable. New York: Random House, 2007. Cited for silent evidence, for the limits of induction over an unrepresentative sample, and for the objection in section 6.3 that a model cannot assign probability to a mechanism it does not contain.
[3] N. N. Taleb, Statistical Consequences of Fat Tails: Real World Preasymptotics, Epistemology, and Applications. STEM Academic Press, 2020. Cited for the difficulty of estimating tail quantities from finite samples, which section 6.2 applies to synthetic samples as well as observed ones.
[4] International Electrotechnical Commission, IEC 62443, Industrial communication networks, Security for industrial automation and control systems. Cited as the control catalogue and assessment method in general use, including by the author, and specifically for the observation in section 5.1 that control-based assessment asks a single-hop question.
[5] International Electrotechnical Commission, IEC 61508, Functional safety of electrical, electronic and programmable electronic safety-related systems. Cited for the precedent in section 2.2 that safety integrity in this sector is already specified as an average probability of failure on demand.
[6] International Electrotechnical Commission, IEC 61511, Functional safety, Safety instrumented systems for the process industry sector. Cited in the same sense as [5], as the process sector application of the same probabilistic specification.
Note on this bibliography#
The statistical results in sections 3.3 and 5.3, the standard error of a proportion, the importance sampling estimator with its likelihood ratio correction, the growth of the number of walks with path length, and the normal form of the saddle-node bifurcation in section 4.6, are standard and are not attributed to any author. They are stated here because the argument turns on the specific magnitudes, not because they are novel.
The description of what a facility actually contains in section 4, including firmware age, configuration drift, the gap between diagrammed and actual reachability, and the behaviour of alert handling under load, is the author's own observation from assessment and engineering work in the energy, rail, water and manufacturing sectors in North America, Australasia and Europe. It carries no citation because no published source records it, and it is offered as testimony.
The psychometric construction in section 4.7 and the gated propagation architecture in section 5.6 are the author's own implementations. They are described so that a reader can judge and reject them. Neither is presented as a validated result, and the requirements they are built to satisfy stand independently of whether these particular constructions are the right way to satisfy them.
No source is cited here for any claim about the frequency of adversary activity against industrial facilities, because section 6.1 argues that no trustworthy source for such a claim exists.