A number and a margin of error tell you nothing about the model that produced them.
A twin drifts from the business it models, and nobody finds out until a decision goes badly. In Polytrion the rules are written down, checked before a run and watched on every step, so a breach arrives as an address — which step, which part, which rule — and every result replays exactly from four values, on anyone's hardware, years later. The engine underneath is our own, built and maintained in-house — internal technology, not a product — and it was qualified against nine published answers before it was allowed to answer yours.
A model is validated against the world it represents. The engine underneath one represents nothing.
So we did not validate it.
We qualified it.
Asking whether a simulation engine is “accurate” is a category error, and answering it with a plausible-looking curve is how simulation software acquires unearned credibility.
You describe the operation. The rules become declared artefacts, checked before a run and on every step of it, and every answer comes back with the evidence attached. Drift becomes something you are shown on a schedule instead of something you discover.
Test the policy you actually deploy. The same policy artefact runs in the twin and in the live system, so the twin cannot quietly diverge from the business.
Fewer runs to pick a winner. Stock levels, routing, staffing — the comparison is the question you care about, and it is the part this does cheaply.
Disruptions with the assumptions on the page. Outages and recovery times are named, versioned inputs, so any result traces back to the assumption behind it.
Publish results a referee can replay. Four values beside a figure let a reader re-derive the number instead of taking your word for it.
Check the mechanism first. Establish that a reaction scheme or move set is consistent before spending months of compute on it.
Sharper comparisons on the same budget. Sensitivity studies are where the savings land, and they grow as the effect you are chasing gets smaller.
Most of what a twin is asked is a sampling question. Sweep the parameters, rank the options, get the distribution.
Before a catalogue of mechanisms is handed to the fast sampler, the engine establishes that it is complete and consistent — exhaustively, before a single trajectory exists. A defect is named by state and channel, not discovered in a wrong average.
There is very little to say about the internals here on purpose — they are ours to maintain, not yours to learn. What matters to somebody deciding whether to use the platform is what holds, not how it is built.
Engine version, scenario, seed, replication index. Quote those beside a figure and anyone re-derives it digit for digit on their own hardware. Nothing has to be shipped, stored or trusted.
Because the mechanism is data rather than buried in the code that scans the state, an independent checker establishes structural properties of it exhaustively — before a single trajectory exists.
Not sampled, not checked at the end. Between eight and eleven invariants per world, evaluated over 105 to 108 journal rows per study. A breach names the node, the tick and the rule, and stops the run.
This is the rule that costs something. Two studies in the suite report underpowered with the exact value they could not estimate and the replication count that would have been needed — one after four thousand replications returned exactly zero.
Every entry was selected on one rule: the right answer must already be known to more digits than the engine can deliver, and it must not be a simulation. A closed form, an exactly solvable finite system, or a linear solve over a finite chain — in that order. A heavily replicated numerical constant is second best; another code's output is not accepted at all. That rule excludes most of the simulation literature, and it is what makes agreement informative — if the reference is somebody's program, reproducing it means reproducing their conventions, and disagreement localises nothing.
| Field | The known answer it was run against | What happened |
|---|---|---|
| Wave one — is the suite viable at all? · complete | ||
| Epidemiology | Reed–Frost final-size distribution, exact rationals | Every point within 0.0030, against tolerances of 0.0044–0.0085; the mean covers at three transmission rates |
| Queueing | Jackson product form and Gordon–Newell, solved exactly | Occupancy 1.2976 against 857090/660203; joint deviation 0.0020 against a 0.0124 budget |
| Traffic | Nagel–Schreckenberg at vmax=1, exact | Five randomisation rates, each against the finite ring's own exact current; all covered |
| Wave two — the substantive claim · complete | ||
| Non-equilibrium physics | Open TASEP, J50 = 26/101 | 0.257276 ± 0.000130 — −1.15σ. The value a defective tie-break would produce is excluded by 57σ |
| Stochastic chemistry | Schlögl network: birth–death sums and an absorbing chain's Green's function | First passage 0.16σ; occupation histogram 0.42σ; the driven cycle behaved as predicted |
| Operations research | Zheng–Federgruen (s, S) optimum, three published triples | All three to the digit — −0.92σ, +0.52σ, −0.41σ |
| Computer networks | Bianchi's 802.11 DCF fixed point | Within 0.7% at every station count; RTS/CTS flat at 1.98% against Basic Access's 34.3% |
| Wave three — the suite's last entry is running | ||
| Rare events | M/M/1 first passage, gambler's ruin | γ20 = 9.670×10−7 ± 1.5×10−8 against 9.537×10−7 |
| Statistical mechanics | 2D Ising on an L×L torus, against Kaufman's exact finite-lattice partition function — differentiated analytically, and cross-derived from the exact density of states | Fifteen points, thirty scored quantities, zero findings. Energy and heat capacity inside the declared rule at every temperature; two starts two units of energy per site apart met at −0.02σ |
| Market microstructure | zero-intelligence continuous double auction — the deep book as an exact infinite-server law, Exponential(δ) share lifetimes, and the dimensional collapse that makes the spread scorable without anyone's fitted prefactor | Built — 34 tests, zero findings across runs. The deep-book law is CLEAN at both order sizes: variance-to-mean 2.30 against an exact 2.5 at σ = 4, the statistic that separates per-share from per-order cancellation. The three collapse arms read 0.4228 / 0.4360 / 0.4325 against Smith et al.’s ≈0.45. Production runs in flight: the deep study at 12 replications, then the ε-collapse on a 120-level ladder. |
Every row's acceptance rule is a rule, not a number — “three standard errors of the ensemble mean at this replication count, plus the declared tick-quantisation allowance,” written into the experiment document before any run exists. “Within one per cent” is a number that can be chosen after seeing the answer. Every completed study reports zero invariant findings.
These are where the construction's other claim — that a mechanism can be checked before it is sampled — was exercised.
| Field | What it was run against | What happened |
|---|---|---|
| Surface chemistry | a published diamond CVD reaction catalogue | Five structural properties of the mechanism checked exhaustively before any trajectory existed. Eight deliberately introduced defects, eight rejected — each with the offending state or channel named. |
| Mathematics | Dhar's abelian sandpile theorem | Four different orderings produced four different executions and one identical answer, exactly as the theorem requires. |
| Quantum error correction | surface-code decoding, against Stim | Agrees with the reference tool in that field. |
| Quantum optics | a continuously measured two-level system, against QuTiP | Agrees. |
The quantum-error-correction row above is a reference-tool agreement. This is what came out of the same world once it was asked something nobody had the answer to.
Polytrion is open to a small number of teams while wave three finishes. Most engagements begin the same way: one model whose result your team can already check independently. It is the fastest way to find out whether this is the right tool, and the fastest way for you to find something we missed.