Fun With Science  /  Globe Deconstruction  /  Sections p. 203 + p. 415  /  Draft

The Standard of Proof Cuts Both Ways

Miller's charge of asymmetric scrutiny is correct and it lands on this review. What it does not reach is the reason the globe stopped being an extraordinary claim — which was never trust, and is checkable this afternoon.

Unreviewed first draft. This page is published for review only. It has not been through the editorial pass the rest of this site has had, it is not linked as an answer in the catalogue, and it is marked noindex. Known outstanding issues on this page: Numbers here have been cross-checked against their research record, but treat any figure as provisional until this notice is removed.
Globe Deconstruction? — Extraordinary Evidence or Fallacy, p. 205 · his words, quoted “Remember, we all agree that extraordinary claims require extraordinary evidence. Which side of the debate is being honest with the scientific method? The globe side is making a very long list of positive claims that should be verified in multiple ways (ideally by a handful of independent 3rd parties). The goal of the skeptical side is simply to demonstrate fallacy in the globe model.”

This page grants Levi Miller his p.203 argument in full: the flat and globe models are not held to the same standard of proof, the "extraordinary claims" maxim cannot repair that, and the charge applies to this review specifically. Granting it costs less than it looks, because the maxim was never what made the difference. What did was three thousand years of the same result being measured again by instruments and methods that did not exist the last time — refined every time, refuted never — until it ended up wired into bridges, radar, flight decks, telescope drives and eclipse tables that would fail this week if it were wrong. That is the reply to the standard-of-proof charge, and it is the first thing below. We then ask what, if anything, does license treating two hypotheses differently without simply announcing one's own starting point — and find one answer that survives, Lakatos's, which is symmetric by construction and imposes obligations on us. To test whether we can actually meet those obligations, we audit ourselves on the narrowest item we have: his p.415 remark about magnetic north in Antarctica. That item is low-stakes and we treat it as such. His observations there are right. The field geometry underneath them turns out to be measured, published, error-budgeted and re-issued on a schedule by people with no interest in this argument, and that — not the strangeness of anyone's claim — is the whole of the asymmetry we think we are entitled to.

p.203 asymmetry charge: correctSagan maxim: indefensible as deployedThis review has applied a double standardp.415 polarity naming: correctAntarctic compass problems: realDip pole is not on the continentCompass works at the geographic South PoleAsymmetry licensed by deliverables, not priorsThree thousand years: refined, never refutedThe book argues the edges, not the centre
Where this lands

Miller is right on p.203 and we concede it before anything else. "Extraordinary claims require extraordinary evidence" is not a standard of evidence; in the Bayesian reconstruction it says only that the speaker's prior is low, and Bayes' theorem supplies no rule for setting priors. The charge of asymmetric scrutiny lands on this review, and we name a concrete instance of our own from the research behind this very page. What replaces the maxim is not a softer version of it but a different criterion entirely — Lakatos's, which never mentions how surprising anyone finds anything, and which binds the incumbent model exactly as hard as the challenger. On p.415 his observations are also right: the Antarctic magnetic pole is, in physics polarity terms, a north pole; compass needles really do need rebalancing; parts of Antarctica really are compass-hostile. What the folklore adds to that is wrong in ways the published record settles quickly. The dip pole left the continent for the Southern Ocean around the early 1960s. The region where a compass genuinely fails is a few hundred kilometres across and centred on the dip pole, not on the continent or the geographic pole. At the geographic South Pole the horizontal field is 16,816 nT, about 86 percent of London's, and the compass works fine; what fails there is the reference direction. This is a small item and we say so.

The charge is correct, and it lands here first

Miller's argument on p.203 is that the flat and globe models are not held to the same standard of proof — that an observation which would count as suggestive for the incumbent is required to be conclusive for the challenger, and that an anomaly which the incumbent may absorb as an open problem is treated as fatal to the challenger. Before we say anything else about it, we should say that we think he is right.

We do not mean right in the narrow sense that some skeptics are rude about it. We mean that the standard usually invoked to justify the asymmetry does not justify it, that we have invoked that standard, and that the asymmetry we have practised has not so far been defended. That is three separate concessions and we grant all three.

The sharpest version of the point is not ours and is not Miller's. It belongs to Marcello Truzzi, who coined the modern form of the maxim in 1978 and spent much of the rest of his career objecting to what was done with it. In "On Pseudo-Skepticism" he wrote that if a critic asserts there is evidence for disproof — that he has a negative hypothesis, that a result was actually due to an artifact — then "he is making a claim and therefore also has to bear a burden of proof."

We have not seen the Zetetic Scholar originals. Both Truzzi quotations reached us through the secondary literature and we attribute them that way rather than dressing them as primary. Given what this page is about, pretending otherwise would be absurd.

This review is precisely the kind of critic Truzzi described. We do not take an agnostic posture and wait to be convinced. We assert, page after page, that specific claims are wrong for specific reasons. Every one of those assertions is a positive claim carrying its own burden, and Truzzi's rule puts that burden on us.

Here is a concrete instance, from the research behind this very page rather than a hypothetical. While preparing the Antarctic section below, we accepted a set of Geoscience Australia observatory annual means from an automated document summary rather than from the source document. One figure looked wrong, so we checked it against the definitional identity H = F·cos I — an identity the tabulated elements must satisfy, because that is what the symbols mean. It failed. The summary had returned a horizontal intensity of 3,264 nT for Mawson at epoch 2007.5; the primary table gives 18,512 nT. The values were fabricated. We discarded them and went to the PDF.

We caught that by habit, not by policy. Miller's arithmetic gets checked every time, as a matter of course; our own sources had not been getting the same treatment, and that is the double standard he is describing, occurring in our own workflow, on the page where we were preparing to describe it.

We have adopted the identity check as standing practice for every geomagnetic number this site publishes. That fixes one instance. It does not answer the general charge, and the rest of this page is an attempt to answer it honestly rather than to change the subject.

What turned the extraordinary into the ordinary, and it was not trust

Before the epistemics, the thing the maxim is groping at and getting wrong.

The claim was extraordinary once. A man in Alexandria measuring the size of a planet from the length of two shadows was an extraordinary claim in 240 BC. So was putting the Sun eighteen lunar distances away from the angle of a half-lit Moon. Nobody had seen the thing being described, the instruments were a stick and a protractor, and the people making the claims were doing something nobody had done. Scepticism was the correct response. It was novel work by the best minds available, and it was entitled to be doubted.

What happened next is the part the maxim cannot describe. The result did not sit there accumulating trust. It got measured again, with a better instrument. Then again, by a method that had not existed when it was first made. Eratosthenes had two sticks. Picard had a quadrant and a triangulation chain. Maupertuis had a survey of Lapland run to settle an argument between Newton and Cassini that had nothing to do with the shape being flat. Struve had 2,821 km of arc and 265 stations. Then satellite geodesy, which needs no ground at all. Then a sealed box in an aircraft bay that finds its own latitude by sensing the planet turn underneath it, and will refuse the alignment if the crew types in a position that disagrees. Different instruments, different physics, different centuries, different motives, no shared apparatus that could fail the same way twice.

In three thousand years of documented measurement the core result has been refined and never once refuted. The number moved — 252,000 stadia became 40,007.9 km — and each refinement moved it less than the one before. That is what a converging measurement looks like. A wrong one does not do this; a wrong one gets worse when you improve the instrument, because the error was in the thing, not in the ruler.

And that convergence is not sitting in a journal being admired. It is load-bearing. It is inside things that would fail — visibly, expensively, this week — if it were wrong. Which is the whole of what “ordinary” means here: not that people stopped asking, but that the answer got wired into enough independent machinery that it is checked continuously by people who are not thinking about it at all.

The edges, and what sits underneath them

Which brings us to a pattern in this book worth naming plainly, because it is why the standard-of-proof argument feels stronger than it is. The questions cluster at the edges of measurement: where an effect is small, an observation ambiguous, an instrument pushed past its comfortable range, or a photograph hard to read. That is a legitimate place to look. Edges are where models break, and a model that has never been probed at its edges has not really been tested.

But an edge case only threatens a model whose centre is untested. Here the centre is not merely tested. It is in daily industrial use, in trades that would notice within a week.

The edge the book examinesWhat runs on the same physics, every day
Whether an object vanishes bottom-first across a few hundred metres of pond — Claim #1, p. 2Radio and radar path planning. Every microwave link, broadcast transmitter and air-traffic radar is sited with an explicit earth-bulge term and a four-thirds-radius refraction factor. Get it wrong and the link does not close.
Whether water in a controlled channel curves — Q2, pp. 88–116Long-baseline construction. The Verrazzano-Narrows towers stand 1⅝ inches farther apart at the top than at the base across a 4,260-foot span, because they are built vertical and vertical is not parallel. LIGO’s 4 km arms had to be laid straight through 1.25 m of departure from level.
Whether the Sun shrinks rather than sets — Q8, p. 51Solar arrays and daylight studies. Every photovoltaic installation and every planning-permission shadow study runs a solar-position algorithm built on a spherical Earth and a Sun at one astronomical unit.
Whether refraction invalidates shadow angles — Q12, p. 55Land surveying — the trade the refraction constant he cites belongs to, where it is applied as a seventh of a curvature correction nobody in that trade disputes.
Whether a flight recorder would show a turn on the equator — Q3, pp. 117–123Inertial reference alignment, on the ramp, before every departure, tens of thousands of times a day, in a procedure that computes latitude and rejects the crew’s entry if it disagrees.
Whether the atmosphere co-rotates — Q10, pp. 179–202Weather forecasting and gunnery. Every numerical weather model carries a Coriolis term going as sin φ, and artillery firing tables carried the correction long before computers did.
Whether the star field moves at one rate — Q5, p. 48Telescope drives, from a hobbyist’s mount to a professional observatory, all tracking at the sidereal rate of 360° in 23h 56m 04s.
Whether Jupiter’s moon shadows align — Q6, p. 49Amateur astronomy. Shadow-transit times are published years ahead and checked from back gardens; the same moons were used to fix longitude before chronometers existed.
Whether the Moon’s silhouette appears before an eclipse — Q9, p. 52Eclipse prediction. Paths of totality are published years in advance, to the kilometre and the second, and then walked into by millions of people who would notice.
Whether the rotational south pole shows landscape movement — Q11, p. 54The marker at Amundsen–Scott is dug up and repositioned every 1 January, because the ice under it flows about ten metres a year. The pole stays put; the ice does not.
Whether magnetic north in Antarctica is what the folklore says — pp. 415–473Aviation charts, reissued on the World Magnetic Model’s five-year cycle, and compass swings performed on aircraft against surveyed ground marks.
Whether rockets need air to push against — Claim #2 and Q1, pp. 56–87The GPS fix on the device reading this page.
The obvious objection to that table, which we would make ourselves. An engineering practice can encode a working approximation for centuries without the model underneath it being right — Ptolemaic epicycles predicted planetary positions well enough to navigate by. So “it is used every day” is not on its own a proof of anything. Two things separate this case. The entries do not share a model that could be jointly wrong: radio propagation, Newtonian statics, inertial mechanisation, celestial mechanics and geomagnetism are different physics with different failure modes, and one wrong planetary shape would have to produce exactly compensating errors in all of them independently. And several are predictive and dated in advance — an eclipse path, a shadow transit, a sidereal rate — which is the thing an encoded fudge cannot do, because a fudge is fitted to what already happened.

Nor is the evidence concentrated where a conspiracy argument needs it to be. The shadow-transit table is checked by amateurs with telescopes they paid for. The offsets from the tangent to the parallel are run by survey crews. The tower is built by ironworkers who plumb it vertical and find the top wider than the drawings would suggest on a plane. The sidereal drive is in a back garden. And the shortest version of the whole thing is a stick, a level lake and an afternoon. None of that arrives from a government, and most of it is done by people with no stake in this argument who would be delighted to find an anomaly.

So questioning is not the problem, and we should say that clearly. Everything in that left-hand column started as a question somebody asked when the answer genuinely was not known, and the people who asked them are the reason the right-hand column exists. The division is not between people who question and people who accept. It is between a question that gets taken to a measurement and a question that stays a question — and the measurements are, almost without exception, ones a determined person can make or check themselves.

What the maxim actually says, once you write it in odds

"Extraordinary claims require extraordinary evidence" sounds like a standard of evidence. Written out, it is not one.

The Bayesian reconstruction is short enough to check in a line. Posterior odds equal prior odds multiplied by the likelihood ratio:

> O(H | E) = O(H) × P(E | H) / P(E | ¬H)

In that expression, "extraordinary claim" means exactly one thing: low prior odds. "Extraordinary evidence" means exactly one thing: a likelihood ratio large enough to overcome them. If your prior odds against a hypothesis are 1 in 10^k, you need a likelihood ratio of at least 10^k to bring it to even money. That much is arithmetic and nobody disputes it.

Now the problem. Bayes' theorem is a constraint on how beliefs move. It contains no rule whatever for where they start. There is no term in it that fixes O(H), no procedure inside the calculus that derives a prior from the world. So a demand for extraordinary evidence, made without any defence of the prior that generated the demand, is a demand backed by nothing except the demander's own starting point. It reports a fact about the speaker in the grammar of a fact about the claim.

The maxim, deployed bare, is not a standard. It is a starting point wearing the costume of a standard.

This is not a fringe reading and it is not our invention. It is how the standard Bayesian formalisation renders the maxim, and it is why the maxim's own coiner objected to its use.

The history is worth having in view, partly because it shows the idea is old and partly because one of its most famous applications was on the wrong side. David Hume in 1748 held that the evidence for a report receives "a diminution, greater or less, in proportion as the fact is more or less unusual". Laplace is generally rendered as holding that the weight of evidence for an extraordinary claim must be proportioned to its strangeness. Thomas Jefferson, in 1808, wrote that where facts are suggested "bearing no analogy with the laws of nature as yet known to us, their verity needs proofs proportioned to their difficulty" — and the fact in question was that stones fall from the sky. Meteorites are real. The extraordinary claim was true, the maxim was applied correctly by an intelligent man, and it produced the wrong answer.

The Hume, Laplace and Jefferson formulations above reached us through a secondary compilation of the maxim's history. We have not read them in the originals and do not present them as primary quotation. Anyone building an argument on the precise wording should go to the sources; we are using them only to establish that the idea long predates its modern slogan.

Truzzi gave the maxim its modern form in 1978. Sagan popularised it across the late 1970s. Neither addition supplied the missing piece, because the missing piece cannot be supplied from inside the theorem.

Three places the maxim breaks, all of them in Miller's favour

It is observer-relative, and never says whose observer. "Extraordinary" is not a property a claim carries around. It is a relation between a claim and some background body of belief. The moment anyone uses the word, they have selected a background — and they have almost never said which one, or defended it. Two people with different backgrounds will assign the same claim different degrees of extraordinariness and both will be using the word correctly. Nothing in the maxim adjudicates between them. The reference class is not chosen by probability theory. The usual repair is to appeal to a base rate: most fringe claims turn out false, therefore the prior is low. But a base rate requires a reference class, and Alan Hájek's argument — that reference-class selection is a problem for every interpretation of probability, not merely for frequentism, and is not itself settled by the calculus — applies with full force here. We paraphrase him rather than quote him; we located the paper but did not extract a verbatim thesis statement in this pass.

The practical consequence is easy to display. Put Miller's claim in three classes, each of them defensible:

Reference classRough implied base rate
Fringe geometrical claims about the EarthVery low
Claims that a widely-taught model contains an unexamined assumptionNot low at all
Claims by a careful autodidact whose arithmetic has checked out repeatedlyOn this review's own record, better than even

We are not asserting a numerical prior from that third row, and no reader should read one into it. The point is narrower and it is fatal to the base-rate repair: the base rate swings enormously with the choice of class, and probability theory does not choose the class for you. Whoever picks the class has already picked the answer.

The ratchet. If any anomaly for the incumbent is absorbable as a research problem, and any anomaly for the challenger is disqualifying, then no possible observation can move the verdict. A procedure with that property is not a test. It has the form of scrutiny and the function of a conclusion.
This review is exposed to the ratchet, and it has no immunity that we can point to. We have on occasion treated the fact that a number came from an institution as though the institution were evidence. Provenance is not a likelihood ratio. Where we have done that, we were not weighing, we were deferring, and Miller is entitled to say so.

Two results that constrain us more than they constrain him

Two pieces of formal work bear directly on this, and both of them tighten the screws on the incumbent rather than the challenger.

The first is Blackwell and Dubins, "Merging of Opinions with Increasing Information", published in the Annals of Mathematical Statistics in 1962. Their main theorem: suppose P is a predictive probability and Q is absolutely continuous with respect to P; then for each conditional distribution of the future given the past under P, there is a corresponding one under Q such that, except on a set of histories of Q-probability zero, the distance between them converges to zero as the evidence accumulates.

In plain terms: two agents who start with different priors, and who update on the same growing stream of evidence, end up agreeing — provided neither has assigned probability zero where the other assigns something positive.

Two consequences follow, and both cut our way rather than his.

First, if a disagreement really is about priors, then evidence dissolves it. Refusing to look at evidence because the prior is low is exactly the behaviour the theorem forbids, since it is the one move that prevents the merging from happening. "That claim is too extraordinary to examine" is not a shortcut to the conclusion the theorem would eventually deliver. It is a refusal to run the process at all.

Second, the absolute-continuity hypothesis is a condition on us. This is Cromwell's rule, in Lindley's formulation: "Leave a little probability for the moon being made of green cheese; it can be as small as 1 in a million, but have it there since otherwise an army of astronauts returning with samples of the said cheese will leave you unmoved." The mathematics is immediate — multiply zero prior odds by any finite likelihood ratio and you still have zero.

If our prior on the flat model were literally zero, then nothing on this site would be epistemology; it would be a very long restatement of where we began.

So we state it on the record. Our prior on the flat model is not zero. We are not able to give it a number, and we would distrust anyone who offered one to four significant figures, but it is a quantity that evidence can move, and we have set out below the specific evidence that would move it.

The replacement: Lakatos, not Sagan

If the maxim cannot license asymmetric appraisal, is there anything that can? We think there is exactly one candidate that does not beg the question, and it works by never mentioning surprise at all.

Imre Lakatos appraised rival research programmes on their track records rather than their plausibility. A programme is theoretically progressive when each successive version has excess empirical content over its predecessor — when it "must predict novel and hitherto unexpected facts", where novel is comparative: predicted neither by any rival programme in the offing nor by conventional wisdom. It is empirically progressive when some of that novel content is subsequently corroborated. It is degenerating when its modifications merely accommodate anomalies after the fact and generate nothing new to test.

We take Lakatos's criterion from the Stanford Encyclopedia entry, which quotes the Methodology of Scientific Research Programmes; we have not worked from the primary text in this pass and attribute it accordingly.

Notice what this criterion does not contain. It contains no term for how strange a claim feels, no reference class, no base rate, no background body of belief. It is symmetric by construction — the same test, applied the same way, to both programmes. And it is settled by inspecting publication records, which either exist or do not, rather than by consulting intuitions, which always exist and are always available in whatever quantity is required.

The obligation this creates for us is immediate and we accept it in advance. Incumbency earns nothing. If the globe programme were reduced to absorbing anomalies without staking risky new predictions, it would be degenerating under this criterion, and this review would be obliged to report that. We invite the test rather than merely permitting it.

There is a companion argument we have leaned on elsewhere and which needs its weakness stated plainly here. It runs: the flat model cannot be adopted alone, because adopting it requires a long list of independently established results to be false, each established by a different method, by a different community, for a different purpose. Formally, if pieces of evidence E₁…Eₙ are conditionally independent given the hypotheses, the total likelihood ratio is the product ∏LRᵢ, so n individually modest tests compound very fast.

That argument has a load-bearing assumption and it is the one most often false. Conditional independence is not free. Two "independent" confirmations that share an instrument, a calibration chain, a datum, or a modelling assumption are one confirmation wearing two coats. Every item we enumerate has to be defended on independence specifically, and this review has not always done that work.

We think this is the single most productive line of attack available to Miller, and we would rather he took it than argued about priors: show that two things we have counted as independent share a dependency, and the exponent in our product falls, item by item, without any need to dispute a single measurement.
Three sections removed after publication. This page carried a worked example on Antarctic compass behaviour — where the south dip pole sits, and whether a compass works at the geographic South Pole. Checking the catalogue against the book showed that Miller does not make that claim. His actual argument, at pp. 452–462, is that magnetic declination is itself an invention: “What if magnetic declination was a concept created to make our current maps appear accurate? Was it used to remove our trust in the compass? What if the compass was right all along?” — from which he builds a “Magnetic Declination Removed” flat map. We answered a question he did not ask and left the one he did ask untouched, so the material has been withdrawn rather than left standing. The epistemology above is unaffected.

Where this page could be wrong

The circularity limit, stated without softening. IGRF and WMM are spherical-harmonic expansions written in a geocentric spherical frame, with a geomagnetic reference radius of 6,371,200 m baked into the formalism. Citing the internal coherence of such a model as proof of sphericity would be circular, and we say so plainly rather than waiting to be caught. A determined alternative programme could in principle fit a different basis on a different surface. The honest answer is not that this is impossible; it is that nobody has done it — which is a claim about deliverables, not a proof about geometry.

What is not circular is out-of-sample predictive success against instruments that were not in the fit, and certification of safety-of-life navigation against a published error bar that the model is then held to. Those two things are the whole of our argument in the previous section, and we must not let them blur into the circular version. If a reader finds us doing so anywhere on this site, we want to know where.

Almost every Antarctic number above is ours, not published. The station table, the GVS figure, the pole positions, the drift rate, the displacement since 1900 — all are our computation from IGRF-14, validated three ways (against NCEI's, BGS's and AAD's published pole positions, and against the Geoscience Australia observatory means) but derived nonetheless. We have labelled them so. If we ever present them as published values, that is an error and it is a serious one. Our own intuitions are worthless here, and we tested that. We predicted Antarctic declination from a centred-dipole picture — horizontal field directed away from the south geomagnetic pole at 80.85°S, 107.24°E — and compared it with IGRF-14 at nine stations. The residuals run up to about 125°, with an RMS of about 70°. The reason is physical: near a dip pole the horizontal component is a small residual of much larger vertical terms, so its direction is dominated by non-dipole structure, and the dip pole itself lies some 2,066 km from the geomagnetic pole. This is not a point against Miller. It is a caution against everybody's mental bar magnet, including any version of it we might be tempted to use. Several of our epistemology citations are second-hand. The Truzzi quotations, the Hume and Laplace formulations, the Jefferson line and the Hájek thesis all reached us through secondary sources. We have flagged each in place. On a page whose entire argument is about holding oneself to the standard one demands of others, quoting these as primary would be a self-inflicted wound of exactly the kind the page diagnoses. Two things we dropped rather than publish unverified. We had a derivation of the maximum elevation angle of a GPS satellite as seen from the geographic pole, and a comparison of Mawson, David and Mackay's 1909 South Magnetic Pole determination against IGRF-14. We are not publishing either. The first rests on constellation parameters we did not confirm from the current performance standard. The second is a consistency check between a model and data of the kind the model was fitted to, not an independent test, and we did not verify the 1909 coordinate from the expedition's own record. And the tone risk. The concessions in the first half of this page are extensive. A reader in a hurry might take us to be conceding the substantive question. We are not. We concede the epistemological point in full, and then argue that a different, prior-independent criterion still separates the two programmes. If that transition reads as a bait-and-switch, the fault is ours in the writing, and the remedy is that the obligation Lakatos's criterion places on us is exactly as binding as the concession we made to him.

What would change our mind

Stated in advance, specifically, so that nobody has to take our word for it later.

One audited novel prediction. A single quantitative prediction derived from Miller's model, stated in advance with an error bar, that the globe model does not make, and then confirmed by a measurement taken by someone with no stake in the question. One such result would carry a likelihood ratio that no prior of ours could survive, and it would satisfy Lakatos's own criterion for a progressive programme on the criterion's own terms. We will say so on this page if it happens. A broken independence claim. A demonstration that two or more of the lines of evidence we have enumerated as independent are not — that they share an instrument, a calibration chain, a datum, or a modelling assumption. Every such demonstration reduces the exponent in our compounding product and correspondingly reduces the entanglement cost we have claimed. We regard this as the most productive line available to him and we would rather he took it than argued about priors. A model-dependent "independent" measurement. A demonstration that something we have described as made for reasons unrelated to this debate was in fact model-dependent in the relevant way — that its reduction from raw instrument output to published value assumed the geometry in dispute. The Geoscience Australia observatory annual means are the specific target. If the absolute-observation reduction at Mawson or Casey presupposes a spherical Earth in a way that would change the published D, I, H or F, we want to know, and this page will be corrected rather than defended. A failing observatory. On the Antarctic case specifically: a published magnetic observatory record from a station south of 60°S whose measured declination departs from IGRF-14 by more than a few degrees under geomagnetically quiet conditions, and which cannot be attributed to a known local crustal anomaly. That would falsify the audit the previous section rests on. The eight observatories in the WMM2025 fit are the obvious places to look, and the data are public. A rival field model. A quantitative field model — of any geometry, on any surface — fitted to the same 140 observatory series and the same Swarm data, with a stated coefficient count and a stated RMS error budget, that meets or beats WMM2025's 0.36° global declination RMS. Not an argument that one could exist. A model, with coefficients, that can be run. If the flat programme produced one, the Lakatosian asymmetry we have claimed would evaporate and we would have to withdraw the argument in the last two sections entirely. Another double standard of ours. Evidence that this review has applied asymmetric scrutiny on some specific page — a place where we accepted a mainstream number without the check we demanded of his. We have already found one such instance in our own research and disclosed it above. We expect there are others.

We will publish those corrections rather than defend them, and we will name the pages. That is the only version of this argument we are entitled to make.

Sources & further reading