The Gun Did Not Win

Building in Steel, part two of three. Why the countries outside the frontier-model race are being told to compete on the wrong axis.

15 August 2026

Listen to this essay (24 min)

Narrated by Charlie via ElevenLabs

There is a story about gunpowder that almost everyone knows. The knight spends fifteen years learning to fight and inherits the land that pays for the horse and the armour. Then a peasant with three weeks of drill and a firearm shoots him off it, and a thousand years of social order falls over. Skill dethroned by technology. Power moving from those who trained to those who bought.

It is a wonderful story, it is largely wrong, and the way it is wrong is the most useful thing in it.

What the historians actually say

The story survives in popular history. In the specialist literature it has been under sustained attack for thirty-five years, and the attack is more interesting than the story.

Michael Roberts proposed a military revolution in the 1950s and Geoffrey Parker expanded it in the 1980s. Bert Hall and Kelly DeVries reviewed Parker in 1990 and made the charge that stuck: the thesis was overly broad and leaned too hard on technological determinism, the assumption that the hardware drove the social change. Clifford Rogers answered the whole framing by arguing there was no single revolution but a sequence of them, infantry, artillery, artillery-fortress, and administrative, spread across centuries.

Note the last item on Rogers's list. It is not a weapon.

Jeremy Black pushed furthest, and his conclusion is the one this essay is built on. He argued the decisive changes came after 1660 rather than in Roberts's 1560 to 1660 window, that warfare in that supposedly revolutionary century was inconclusive for reasons of manpower, logistics, finance and fortification rather than infantry weapons, and, most pointedly, that state development enabled military growth rather than military change producing the state.

That is the causal arrow reversed. Not: new weapon, therefore new kind of state. Rather: a state that had already built certain capacities could use the new weapon, and one that had not, could not.

The powder was on sale to everyone

Gunpowder was not a secret. The Chinese had it first, the Ottomans had it, every European power could buy it. If possession conferred advantage, the advantage would have been evenly spread. It was not.

What varied was the capacity to absorb it, and the clearest case has a name and a date. In the 1590s Maurice of Nassau rebuilt the Dutch army around the problem of making ordinary men useful with firearms. He standardised drill, equipment and pike lengths. He standardised firearm length and calibre, so a weapon and its ammunition were not a bespoke pairing. He drilled ranks to load, step and fire in sequence until fire became continuous rather than one ragged volley followed by a long helpless pause. He founded a military academy, because officers now had to be trained rather than born. And he insisted on regular pay, because a formation that holds under fire is made of men not wondering whether they will be paid. Gustavus Adolphus took that system into the Thirty Years' War.

Drill. Standardisation. A training institution. Payroll. Not one is a weapon, every one is something a state either had the administrative capacity to do or did not, and they took generations to build, which is exactly why they could not be bought in a hurry by whoever had just noticed they were behind.

The powder diffused instantly. The institutions did not diffuse at all. Which gives the analogy in its uncomfortable form: the interesting question was never who had the better gun, but who had built the kind of state that could make an average recruit dangerous with an average one.

Standardisation is the absorption mechanism

Of everything on Maurice's list, standardisation is the one that keeps recurring across four centuries, and the one happening right now in front of us.

Maurice standardised calibre in the 1590s so a barrel from one workshop could be served by shot from another. Two and a half centuries later, in 1841, Joseph Whitworth proposed a national standard for screw threads, fixing the angle at 55 degrees and the pitch by diameter. Before that, workshops cut their own threads and a bolt was married to the hole it was made for. After it, a part made in one factory fitted a machine built in another. Same move, three hundred years apart: stop treating the pairing of component with component as bespoke, and a capability held by specialists becomes a capability held by an economy.

Standards are how capability reaches everyone, and they are reliably the least interesting thing in the room at the moment they are set. Two are being set right now.

The API shape settled by cloning. OpenAI's Chat Completions format became the de facto standard for talking to a language model, not by decree but because everyone copied it. Inference servers, gateways and rival providers all expose an OpenAI-compatible endpoint, and a client written against that shape can point at a different model by changing a base URL. The detail I like most: OpenAI now steers new projects towards its newer Responses API, and the cloned standard carries on regardless. Once a standard is load-bearing it stops belonging to whoever wrote it.

The tool protocol just dropped its handshake. In the 2026-07-28 revision, the Model Context Protocol moved from a bidirectional stateful protocol to a request-and-response stateless one, retiring the initialize exchange and the session-ID header so a request now carries what it needs to be understood alone. That is statelessness at the protocol layer; application state can still live in payloads or storage, so it is not a claim that state has gone away. But the consequence is the one that matters for diffusion: a server holding no session can sit behind an ordinary load balancer and be scaled by people who do not understand it, which is what it means for a capability to become infrastructure.

Both are the Whitworth thread again. Interchangeability is not a feature anyone gets excited about. It decides whether a technology stays with specialists or reaches everyone with a workshop.

The two blocs, and everyone else

The map is easy to draw. One bloc runs the frontier closed models: OpenAI, Anthropic, xAI, on American capital, chips and energy. A second runs the frontier open-weights models: DeepSeek, Qwen, Kimi, on Chinese industrial policy and a deliberate strategy of giving the weights away.

Everyone else is invited to conclude they have lost. The United Kingdom, France, Germany, the Netherlands, Australia, Canada, Japan, South Korea, Israel, Singapore, India. Each has some combination of world-class research, real engineering and genuine mathematical depth, and none will out-capitalise the two blocs on pre-training. The sums do not work and are not going to start.

So the honest question is not how to catch up. It is which axis decides outcomes, given the powder is already on sale to everyone.

Diffusion, not invention

The best empirical work here is Jeffrey Ding's, in Technology and the Rise of the Great Powers. He sets GPT diffusion theory against leading sectors theory, and the contrast is the whole point. Leading sectors theory says a country wins by building institutions that let it monopolise innovation in a new, fast-growing industry. Ding's argument is that for a general purpose technology this is the wrong model: the productivity effects emerge only after a long diffusion across many sectors, so the country that first develops or commercialises the technology is not necessarily the one that gains most from it. Britain did not invent every element of the industrial revolution; the United States did not invent the internal combustion engine or the transistor. Both won on spread.

That distinction matters here because building a national frontier lab is a leading-sectors move, and it is being made almost everywhere by governments who have not asked whether AI is a leading sector or a general purpose technology.

Two further things get flattened when people repeat the argument.

The first is the mechanism. Diffusion capacity is not a vague willingness to adopt. In Ding's account it requires institutional adaptations that widen the base of engineering skills associated with the technology, with education systems and technical associations as the complementarities that do the widening. A country diffuses a general purpose technology when it can produce, at scale, people who are merely competent with it. Which returns us to Maurice, whose response to firearms was to found an academy. The academy is the diffusion mechanism. The gun is not.

The second is timescale. These technologies act on national power over decades, through economy-wide productivity, not through one decisive application. Nothing here predicts who wins the next three years, and any essay claiming otherwise is borrowing the framework's authority rather than using it.

One assumption should be visible rather than smuggled: this presupposes AI is a general purpose technology in Ding's sense, pervasive across sectors, improving over time, spawning complementary innovation. That looks right to me and is not proven. If AI turns out powerful but narrow, the historical analogies stop applying and so does most of what follows.

All of which is bad news for the "we need a national frontier lab" instinct, the reflex of nearly every government currently writing an AI strategy. A sovereign model nobody deploys is a trophy: the equivalent of buying cannon and never reforming the tax system that would let you field them.

The gate on high-consequence diffusion

Diffusion into consumer software is easy. You ship it, and when the model is wrong a summary is bad and someone is mildly annoyed.

Diffusion into the parts of an economy that carry the actual value is blocked by something specific. You cannot deploy a system you cannot check into law, medicine, energy, aviation, banking, defence or public administration. Not from squeamishness, but because those domains have liability, regulators, and failure modes ending in a courtroom or a coroner's report. The gate is not "is the model good enough". It is can anyone demonstrate what it will and will not do.

That reframes the competition. Where the money and strategic weight sit, the binding constraint is not model quality but checkability, and checkability is a different endowment entirely. It is not bought with capital and energy. It is built from mathematical depth, institutional patience, and a legal system capable of stating its own rules precisely enough to be checked against.

That endowment is distributed very differently from GPU capacity, and it favours exactly the countries currently being told they have lost.

Who actually holds this

Australia produced seL4, an operating system microkernel with a machine-checked proof of functional correctness from specification down to its C implementation, completed in 2009 in the L4.verified project at NICTA, with Gerwin Klein leading the verification. It is the reference achievement of the field, and it went into real cyber-physical systems rather than staying in a paper.

France has the deepest formal-methods culture in the world. Coq, now Rocq, came out of INRIA, as did CompCert, Xavier Leroy's C compiler carrying a proof that its output means what its input said. The mathematical tradition behind that is not acquired by announcing a programme.

Cambridge and Munich are the two poles of Isabelle/HOL, through Larry Paulson and Tobias Nipkow among many others, and Isabelle is the prover seL4 was verified in. Note what that means: the Australian achievement ran on European tooling. This capability is already international and no single state owns it. Lean and Mathlib, where much of the current energy sits, are likewise distributed and belong to nobody.

None of that is a consolation prize. It is accumulated capital in exactly the discipline the high-consequence gate requires, built over decades, which is another way of saying it cannot be bought quickly by anyone who lacks it.

What the UK is actually doing

The clearest institutional bet on this axis is ARIA's Safeguarded AI programme, whose stated purpose is to combine formal-methods approaches with AI to raise confidence in AI-based systems. Its opportunity space rests on an explicit claim that traditional empirical testing will not be reliable enough for certification, and that mathematical proof is the underexplored foundation. Its programme director's stated focus includes formal verification and the coordination needed to make safety claims between parties that do not trust each other, which is precisely the problem regulated industries have.

ARIA's chief executive, Kathleen Fisher, arrived from DARPA, where she ran HACMS, the programme that used formal methods and seL4 to build cyber-physical systems resisting a professional red team. She has publicly described formal methods as mathematically rigorous techniques for proving software behaves correctly, tied Safeguarded AI explicitly to combining them with AI, and emphasised translating research strength into real-world impact.

I want to be careful here, because I got this wrong myself and it is exactly the sort of error that propagates. As far as I can establish, Fisher has not publicly framed this as "the UK cannot lead on frontier models, so it should lead on formal verification". That framing is mine and I will own it rather than borrow authority for it. What is on the record is the substance: an agency of a mid-sized economy has put real money behind proof-based assurance instead of a national frontier lab, and hired someone whose entire career is formal methods for security to run it.

That is a strategic choice, and on the argument above it is the correct one.

Why this position compounds

The frontier-model race has miserable economics for anyone not already in it: enormous capital, brutal depreciation, and no durable position at the end, because the frontier moves and last year's model is worth approximately nothing.

The verification layer behaves in the opposite way, and the Whitworth thread is why. It is a standards position, and standards positions compound. Whoever's proof format, assurance toolkit and certification pathway become the ones regulators and insurers accept holds something that does not depreciate when a better model ships next quarter. It is durable precisely because it sits at the boundary rather than inside the technology, and the boundary is where liability lives. Recall that Chat Completions outlived its own author's preference. That is what a standards position looks like from the inside.

It is also why this is not work a company can do alone. Standards are set by institutions with the patience for a decade-long process, which a mid-sized state with a good legal tradition is well equipped for and a lab racing a competitor is not.

Four objections, and they are all serious

Scope. Nobody can currently prove anything meaningful about the internal behaviour of a frontier model, and anyone implying otherwise is selling something. What is verifiable is not the model. It is the specification it was asked to satisfy, the capability envelope around it, the properties of the artefacts it emits, and the behaviour of the containing system. That is far narrower than "verified AI" and the only version I would defend. An essay blurring the two is worse than useless, because it discredits the real thing.

Diffusion, turned around. Verification capacity is itself subject to Ding's argument. France has had Coq for four decades without a commanding industrial position, and academic strength has repeatedly failed to become deployed capability. Having the mathematicians is necessary and nowhere near sufficient. The countries that win this get proof out of the seminar and into procurement rules, insurance underwriting and court-admissible evidence, and almost none are currently organised to do that. The tax-and-drill problem again, in modern dress.

Theatre. A verified component does not make a secure system. seL4 is proved and the drivers, firmware, network stack and people around it are not, and attackers go where the proof is not. Worse for a strategy built on certification: certificates decouple from what they certify with depressing regularity, and we already have industries where a compliance badge is a purchase rather than an achievement, bought precisely because it is cheaper than the security. A country specialising in verification could quite easily produce a large, profitable, internationally respected apparatus for issuing assurances that do not correlate with whether systems fail.

I have no clean answer, only a structural one: this is why the predicate must be executable and the evidence machine-checkable rather than attested, because a certificate anyone can independently re-run is much harder to launder than a signature on a PDF. That raises the cost of theatre without abolishing it, and any honest version of this concedes the failure mode is not hypothetical, it is the default outcome of most certification regimes ever built.

The vassal objection, which worries me most. A country could do all of this and end up as the compliance department for other people's models, doing the expensive unglamorous checking while value accrues to whoever owns the weights. Not hypothetical: roughly what happened to several European countries in the last platform cycle.

My instinctive answer was that in regulated industries liability sits with whoever signs off, and signing off is where pricing power ends up. I went looking for evidence and it does not support that, at least not in the form I wanted. In audit markets, which are the closest analogue we have, fee premiums track market power and differentiation rather than the bearing of liability as such: large firms with few close substitutes charge more, some work finds the market broadly competitive with no abnormal profits, and at least one major pricing analysis found no relationship between concentration and fees at all.

So the honest position is worse than my instinct and worth stating plainly. Being the party that signs off does not, on its own, buy you anything. What buys you something is being hard to replace. That makes the vassal risk real and the mitigation specific: a country pursuing this has to end up holding something differentiated, the format or the tooling or the expertise that others cannot readily substitute, rather than merely holding the obligation. Do the checking without holding the format and you have volunteered for the low-margin half of somebody else's industry.

What exists today, plainly

Every essay in this series carries this section, because strategy writing is even worse than technical writing at describing the destination in the present tense.

Exists. seL4 and its 2009 proof. Rocq, CompCert, Isabelle, Lean and Mathlib, all in active use. ARIA's Safeguarded AI programme, funded and running. The EU AI Act's presumption-of-conformity machinery, in force. Verification of bounded components in aviation, rail signalling and chip design, routine for decades.

Does not exist. Any national strategy organised around verification capacity as the primary bet, rather than a programme inside a broader AI strategy. Any machine-checkable conformity standard published by a regulator. Any demonstration that a country capturing this layer captures value rather than cost. Any method for verifying meaningful properties of a frontier model itself.

Is an argument, not a finding. Everything connecting the first list to national advantage. The historical claims are sourced and the institutional ones checkable, but the inference to strategy is mine, it rests on an analogy, and analogies are where reasoning goes to die comfortably. Treat it as a hypothesis with a stated mechanism, which is the most a strategic essay can honestly offer.

The formation, not the powder

The knight did not fall to a chemical. He fell to a bundle of causes historians still argue about, fiscal, tactical, social and demographic, in which the weapon is one strand and not the thickest. What is not disputed is the other half: the states that got value out of gunpowder were the ones that could pay men, drill them, supply them and standardise their ammunition.

The countries outside the two blocs are being urged to buy cannon: fund a national lab, announce a sovereign model, lose slowly and expensively on an axis where the arithmetic will never favour them. The alternative is not resignation. It is to build the thing that decides whether any of this can be used where it matters, the capacity to check it, and then to do the boring work of making that capacity interchangeable.

The gun did not win. The drill, the supply line, the standardised calibre, and the state that could pay for all three won.

That is still true, and it is still available.


Part one, How High Can You Build in Mud?, sets out the mechanism. Part three, Make the Safe Harbour Compile, asks what regulators should do with it.


About the author: Eduardo Aguilar Pelaez is CTO and co-founder at Legal Engine Ltd. He writes on formal methods, AI agents, and the discipline of building systems that survive being walked away from.