Fun With Science / Globe Deconstruction / Sections p. 203 + p. 415 / Draft
Miller's charge of asymmetric scrutiny is correct and it lands on this review. What it does not reach is the reason the globe stopped being an extraordinary claim — which was never trust, and is checkable this afternoon.
This page grants Levi Miller his p.203 argument in full: the flat and globe models are not held to the same standard of proof, the "extraordinary claims" maxim cannot repair that, and the charge applies to this review specifically. Granting it costs less than it looks, because the maxim was never what made the difference. What did was three thousand years of the same result being measured again by instruments and methods that did not exist the last time — refined every time, refuted never — until it ended up wired into bridges, radar, flight decks, telescope drives and eclipse tables that would fail this week if it were wrong. That is the reply to the standard-of-proof charge, and it is the first thing below. We then ask what, if anything, does license treating two hypotheses differently without simply announcing one's own starting point — and find one answer that survives, Lakatos's, which is symmetric by construction and imposes obligations on us. To test whether we can actually meet those obligations, we audit ourselves on the narrowest item we have: his p.415 remark about magnetic north in Antarctica. That item is low-stakes and we treat it as such. His observations there are right. The field geometry underneath them turns out to be measured, published, error-budgeted and re-issued on a schedule by people with no interest in this argument, and that — not the strangeness of anyone's claim — is the whole of the asymmetry we think we are entitled to.
p.203 asymmetry charge: correctSagan maxim: indefensible as deployedThis review has applied a double standardp.415 polarity naming: correctAntarctic compass problems: realDip pole is not on the continentCompass works at the geographic South PoleAsymmetry licensed by deliverables, not priorsThree thousand years: refined, never refutedThe book argues the edges, not the centre
Where this lands
Miller is right on p.203 and we concede it before anything else. "Extraordinary claims require extraordinary evidence" is not a standard of evidence; in the Bayesian reconstruction it says only that the speaker's prior is low, and Bayes' theorem supplies no rule for setting priors. The charge of asymmetric scrutiny lands on this review, and we name a concrete instance of our own from the research behind this very page. What replaces the maxim is not a softer version of it but a different criterion entirely — Lakatos's, which never mentions how surprising anyone finds anything, and which binds the incumbent model exactly as hard as the challenger. On p.415 his observations are also right: the Antarctic magnetic pole is, in physics polarity terms, a north pole; compass needles really do need rebalancing; parts of Antarctica really are compass-hostile. What the folklore adds to that is wrong in ways the published record settles quickly. The dip pole left the continent for the Southern Ocean around the early 1960s. The region where a compass genuinely fails is a few hundred kilometres across and centred on the dip pole, not on the continent or the geographic pole. At the geographic South Pole the horizontal field is 16,816 nT, about 86 percent of London's, and the compass works fine; what fails there is the reference direction. This is a small item and we say so.
Miller's argument on p.203 is that the flat and globe models are not held to the same standard of proof — that an observation which would count as suggestive for the incumbent is required to be conclusive for the challenger, and that an anomaly which the incumbent may absorb as an open problem is treated as fatal to the challenger. Before we say anything else about it, we should say that we think he is right.
We do not mean right in the narrow sense that some skeptics are rude about it. We mean that the standard usually invoked to justify the asymmetry does not justify it, that we have invoked that standard, and that the asymmetry we have practised has not so far been defended. That is three separate concessions and we grant all three.
The sharpest version of the point is not ours and is not Miller's. It belongs to Marcello Truzzi, who coined the modern form of the maxim in 1978 and spent much of the rest of his career objecting to what was done with it. In "On Pseudo-Skepticism" he wrote that if a critic asserts there is evidence for disproof — that he has a negative hypothesis, that a result was actually due to an artifact — then "he is making a claim and therefore also has to bear a burden of proof."
This review is precisely the kind of critic Truzzi described. We do not take an agnostic posture and wait to be convinced. We assert, page after page, that specific claims are wrong for specific reasons. Every one of those assertions is a positive claim carrying its own burden, and Truzzi's rule puts that burden on us.
Here is a concrete instance, from the research behind this very page rather than a hypothetical. While preparing the Antarctic section below, we accepted a set of Geoscience Australia observatory annual means from an automated document summary rather than from the source document. One figure looked wrong, so we checked it against the definitional identity H = F·cos I — an identity the tabulated elements must satisfy, because that is what the symbols mean. It failed. The summary had returned a horizontal intensity of 3,264 nT for Mawson at epoch 2007.5; the primary table gives 18,512 nT. The values were fabricated. We discarded them and went to the PDF.
We have adopted the identity check as standing practice for every geomagnetic number this site publishes. That fixes one instance. It does not answer the general charge, and the rest of this page is an attempt to answer it honestly rather than to change the subject.
Before the epistemics, the thing the maxim is groping at and getting wrong.
The claim was extraordinary once. A man in Alexandria measuring the size of a planet from the length of two shadows was an extraordinary claim in 240 BC. So was putting the Sun eighteen lunar distances away from the angle of a half-lit Moon. Nobody had seen the thing being described, the instruments were a stick and a protractor, and the people making the claims were doing something nobody had done. Scepticism was the correct response. It was novel work by the best minds available, and it was entitled to be doubted.
What happened next is the part the maxim cannot describe. The result did not sit there accumulating trust. It got measured again, with a better instrument. Then again, by a method that had not existed when it was first made. Eratosthenes had two sticks. Picard had a quadrant and a triangulation chain. Maupertuis had a survey of Lapland run to settle an argument between Newton and Cassini that had nothing to do with the shape being flat. Struve had 2,821 km of arc and 265 stations. Then satellite geodesy, which needs no ground at all. Then a sealed box in an aircraft bay that finds its own latitude by sensing the planet turn underneath it, and will refuse the alignment if the crew types in a position that disagrees. Different instruments, different physics, different centuries, different motives, no shared apparatus that could fail the same way twice.
And that convergence is not sitting in a journal being admired. It is load-bearing. It is inside things that would fail — visibly, expensively, this week — if it were wrong. Which is the whole of what “ordinary” means here: not that people stopped asking, but that the answer got wired into enough independent machinery that it is checked continuously by people who are not thinking about it at all.
Which brings us to a pattern in this book worth naming plainly, because it is why the standard-of-proof argument feels stronger than it is. The questions cluster at the edges of measurement: where an effect is small, an observation ambiguous, an instrument pushed past its comfortable range, or a photograph hard to read. That is a legitimate place to look. Edges are where models break, and a model that has never been probed at its edges has not really been tested.
But an edge case only threatens a model whose centre is untested. Here the centre is not merely tested. It is in daily industrial use, in trades that would notice within a week.
| The edge the book examines | What runs on the same physics, every day |
|---|---|
| Whether an object vanishes bottom-first across a few hundred metres of pond — Claim #1, p. 2 | Radio and radar path planning. Every microwave link, broadcast transmitter and air-traffic radar is sited with an explicit earth-bulge term and a four-thirds-radius refraction factor. Get it wrong and the link does not close. |
| Whether water in a controlled channel curves — Q2, pp. 88–116 | Long-baseline construction. The Verrazzano-Narrows towers stand 1⅝ inches farther apart at the top than at the base across a 4,260-foot span, because they are built vertical and vertical is not parallel. LIGO’s 4 km arms had to be laid straight through 1.25 m of departure from level. |
| Whether the Sun shrinks rather than sets — Q8, p. 51 | Solar arrays and daylight studies. Every photovoltaic installation and every planning-permission shadow study runs a solar-position algorithm built on a spherical Earth and a Sun at one astronomical unit. |
| Whether refraction invalidates shadow angles — Q12, p. 55 | Land surveying — the trade the refraction constant he cites belongs to, where it is applied as a seventh of a curvature correction nobody in that trade disputes. |
| Whether a flight recorder would show a turn on the equator — Q3, pp. 117–123 | Inertial reference alignment, on the ramp, before every departure, tens of thousands of times a day, in a procedure that computes latitude and rejects the crew’s entry if it disagrees. |
| Whether the atmosphere co-rotates — Q10, pp. 179–202 | Weather forecasting and gunnery. Every numerical weather model carries a Coriolis term going as sin φ, and artillery firing tables carried the correction long before computers did. |
| Whether the star field moves at one rate — Q5, p. 48 | Telescope drives, from a hobbyist’s mount to a professional observatory, all tracking at the sidereal rate of 360° in 23h 56m 04s. |
| Whether Jupiter’s moon shadows align — Q6, p. 49 | Amateur astronomy. Shadow-transit times are published years ahead and checked from back gardens; the same moons were used to fix longitude before chronometers existed. |
| Whether the Moon’s silhouette appears before an eclipse — Q9, p. 52 | Eclipse prediction. Paths of totality are published years in advance, to the kilometre and the second, and then walked into by millions of people who would notice. |
| Whether the rotational south pole shows landscape movement — Q11, p. 54 | The marker at Amundsen–Scott is dug up and repositioned every 1 January, because the ice under it flows about ten metres a year. The pole stays put; the ice does not. |
| Whether magnetic north in Antarctica is what the folklore says — pp. 415–473 | Aviation charts, reissued on the World Magnetic Model’s five-year cycle, and compass swings performed on aircraft against surveyed ground marks. |
| Whether rockets need air to push against — Claim #2 and Q1, pp. 56–87 | The GPS fix on the device reading this page. |
Nor is the evidence concentrated where a conspiracy argument needs it to be. The shadow-transit table is checked by amateurs with telescopes they paid for. The offsets from the tangent to the parallel are run by survey crews. The tower is built by ironworkers who plumb it vertical and find the top wider than the drawings would suggest on a plane. The sidereal drive is in a back garden. And the shortest version of the whole thing is a stick, a level lake and an afternoon. None of that arrives from a government, and most of it is done by people with no stake in this argument who would be delighted to find an anomaly.
"Extraordinary claims require extraordinary evidence" sounds like a standard of evidence. Written out, it is not one.
The Bayesian reconstruction is short enough to check in a line. Posterior odds equal prior odds multiplied by the likelihood ratio:
> O(H | E) = O(H) × P(E | H) / P(E | ¬H)
In that expression, "extraordinary claim" means exactly one thing: low prior odds. "Extraordinary evidence" means exactly one thing: a likelihood ratio large enough to overcome them. If your prior odds against a hypothesis are 1 in 10^k, you need a likelihood ratio of at least 10^k to bring it to even money. That much is arithmetic and nobody disputes it.
Now the problem. Bayes' theorem is a constraint on how beliefs move. It contains no rule whatever for where they start. There is no term in it that fixes O(H), no procedure inside the calculus that derives a prior from the world. So a demand for extraordinary evidence, made without any defence of the prior that generated the demand, is a demand backed by nothing except the demander's own starting point. It reports a fact about the speaker in the grammar of a fact about the claim.
The maxim, deployed bare, is not a standard. It is a starting point wearing the costume of a standard.
This is not a fringe reading and it is not our invention. It is how the standard Bayesian formalisation renders the maxim, and it is why the maxim's own coiner objected to its use.
The history is worth having in view, partly because it shows the idea is old and partly because one of its most famous applications was on the wrong side. David Hume in 1748 held that the evidence for a report receives "a diminution, greater or less, in proportion as the fact is more or less unusual". Laplace is generally rendered as holding that the weight of evidence for an extraordinary claim must be proportioned to its strangeness. Thomas Jefferson, in 1808, wrote that where facts are suggested "bearing no analogy with the laws of nature as yet known to us, their verity needs proofs proportioned to their difficulty" — and the fact in question was that stones fall from the sky. Meteorites are real. The extraordinary claim was true, the maxim was applied correctly by an intelligent man, and it produced the wrong answer.
Truzzi gave the maxim its modern form in 1978. Sagan popularised it across the late 1970s. Neither addition supplied the missing piece, because the missing piece cannot be supplied from inside the theorem.
The practical consequence is easy to display. Put Miller's claim in three classes, each of them defensible:
| Reference class | Rough implied base rate |
|---|---|
| Fringe geometrical claims about the Earth | Very low |
| Claims that a widely-taught model contains an unexamined assumption | Not low at all |
| Claims by a careful autodidact whose arithmetic has checked out repeatedly | On this review's own record, better than even |
We are not asserting a numerical prior from that third row, and no reader should read one into it. The point is narrower and it is fatal to the base-rate repair: the base rate swings enormously with the choice of class, and probability theory does not choose the class for you. Whoever picks the class has already picked the answer.
The ratchet. If any anomaly for the incumbent is absorbable as a research problem, and any anomaly for the challenger is disqualifying, then no possible observation can move the verdict. A procedure with that property is not a test. It has the form of scrutiny and the function of a conclusion.Two pieces of formal work bear directly on this, and both of them tighten the screws on the incumbent rather than the challenger.
The first is Blackwell and Dubins, "Merging of Opinions with Increasing Information", published in the Annals of Mathematical Statistics in 1962. Their main theorem: suppose P is a predictive probability and Q is absolutely continuous with respect to P; then for each conditional distribution of the future given the past under P, there is a corresponding one under Q such that, except on a set of histories of Q-probability zero, the distance between them converges to zero as the evidence accumulates.
In plain terms: two agents who start with different priors, and who update on the same growing stream of evidence, end up agreeing — provided neither has assigned probability zero where the other assigns something positive.
Two consequences follow, and both cut our way rather than his.
First, if a disagreement really is about priors, then evidence dissolves it. Refusing to look at evidence because the prior is low is exactly the behaviour the theorem forbids, since it is the one move that prevents the merging from happening. "That claim is too extraordinary to examine" is not a shortcut to the conclusion the theorem would eventually deliver. It is a refusal to run the process at all.
Second, the absolute-continuity hypothesis is a condition on us. This is Cromwell's rule, in Lindley's formulation: "Leave a little probability for the moon being made of green cheese; it can be as small as 1 in a million, but have it there since otherwise an army of astronauts returning with samples of the said cheese will leave you unmoved." The mathematics is immediate — multiply zero prior odds by any finite likelihood ratio and you still have zero.
If our prior on the flat model were literally zero, then nothing on this site would be epistemology; it would be a very long restatement of where we began.
So we state it on the record. Our prior on the flat model is not zero. We are not able to give it a number, and we would distrust anyone who offered one to four significant figures, but it is a quantity that evidence can move, and we have set out below the specific evidence that would move it.
If the maxim cannot license asymmetric appraisal, is there anything that can? We think there is exactly one candidate that does not beg the question, and it works by never mentioning surprise at all.
Imre Lakatos appraised rival research programmes on their track records rather than their plausibility. A programme is theoretically progressive when each successive version has excess empirical content over its predecessor — when it "must predict novel and hitherto unexpected facts", where novel is comparative: predicted neither by any rival programme in the offing nor by conventional wisdom. It is empirically progressive when some of that novel content is subsequently corroborated. It is degenerating when its modifications merely accommodate anomalies after the fact and generate nothing new to test.
Notice what this criterion does not contain. It contains no term for how strange a claim feels, no reference class, no base rate, no background body of belief. It is symmetric by construction — the same test, applied the same way, to both programmes. And it is settled by inspecting publication records, which either exist or do not, rather than by consulting intuitions, which always exist and are always available in whatever quantity is required.
The obligation this creates for us is immediate and we accept it in advance. Incumbency earns nothing. If the globe programme were reduced to absorbing anomalies without staking risky new predictions, it would be degenerating under this criterion, and this review would be obliged to report that. We invite the test rather than merely permitting it.
There is a companion argument we have leaned on elsewhere and which needs its weakness stated plainly here. It runs: the flat model cannot be adopted alone, because adopting it requires a long list of independently established results to be false, each established by a different method, by a different community, for a different purpose. Formally, if pieces of evidence E₁…Eₙ are conditionally independent given the hypotheses, the total likelihood ratio is the product ∏LRᵢ, so n individually modest tests compound very fast.
That argument has a load-bearing assumption and it is the one most often false. Conditional independence is not free. Two "independent" confirmations that share an instrument, a calibration chain, a datum, or a modelling assumption are one confirmation wearing two coats. Every item we enumerate has to be defended on independence specifically, and this review has not always done that work.
What is not circular is out-of-sample predictive success against instruments that were not in the fit, and certification of safety-of-life navigation against a published error bar that the model is then held to. Those two things are the whole of our argument in the previous section, and we must not let them blur into the circular version. If a reader finds us doing so anywhere on this site, we want to know where.
Almost every Antarctic number above is ours, not published. The station table, the GVS figure, the pole positions, the drift rate, the displacement since 1900 — all are our computation from IGRF-14, validated three ways (against NCEI's, BGS's and AAD's published pole positions, and against the Geoscience Australia observatory means) but derived nonetheless. We have labelled them so. If we ever present them as published values, that is an error and it is a serious one. Our own intuitions are worthless here, and we tested that. We predicted Antarctic declination from a centred-dipole picture — horizontal field directed away from the south geomagnetic pole at 80.85°S, 107.24°E — and compared it with IGRF-14 at nine stations. The residuals run up to about 125°, with an RMS of about 70°. The reason is physical: near a dip pole the horizontal component is a small residual of much larger vertical terms, so its direction is dominated by non-dipole structure, and the dip pole itself lies some 2,066 km from the geomagnetic pole. This is not a point against Miller. It is a caution against everybody's mental bar magnet, including any version of it we might be tempted to use. Several of our epistemology citations are second-hand. The Truzzi quotations, the Hume and Laplace formulations, the Jefferson line and the Hájek thesis all reached us through secondary sources. We have flagged each in place. On a page whose entire argument is about holding oneself to the standard one demands of others, quoting these as primary would be a self-inflicted wound of exactly the kind the page diagnoses. Two things we dropped rather than publish unverified. We had a derivation of the maximum elevation angle of a GPS satellite as seen from the geographic pole, and a comparison of Mawson, David and Mackay's 1909 South Magnetic Pole determination against IGRF-14. We are not publishing either. The first rests on constellation parameters we did not confirm from the current performance standard. The second is a consistency check between a model and data of the kind the model was fitted to, not an independent test, and we did not verify the 1909 coordinate from the expedition's own record. And the tone risk. The concessions in the first half of this page are extensive. A reader in a hurry might take us to be conceding the substantive question. We are not. We concede the epistemological point in full, and then argue that a different, prior-independent criterion still separates the two programmes. If that transition reads as a bait-and-switch, the fault is ours in the writing, and the remedy is that the obligation Lakatos's criterion places on us is exactly as binding as the concession we made to him.Stated in advance, specifically, so that nobody has to take our word for it later.
One audited novel prediction. A single quantitative prediction derived from Miller's model, stated in advance with an error bar, that the globe model does not make, and then confirmed by a measurement taken by someone with no stake in the question. One such result would carry a likelihood ratio that no prior of ours could survive, and it would satisfy Lakatos's own criterion for a progressive programme on the criterion's own terms. We will say so on this page if it happens. A broken independence claim. A demonstration that two or more of the lines of evidence we have enumerated as independent are not — that they share an instrument, a calibration chain, a datum, or a modelling assumption. Every such demonstration reduces the exponent in our compounding product and correspondingly reduces the entanglement cost we have claimed. We regard this as the most productive line available to him and we would rather he took it than argued about priors. A model-dependent "independent" measurement. A demonstration that something we have described as made for reasons unrelated to this debate was in fact model-dependent in the relevant way — that its reduction from raw instrument output to published value assumed the geometry in dispute. The Geoscience Australia observatory annual means are the specific target. If the absolute-observation reduction at Mawson or Casey presupposes a spherical Earth in a way that would change the published D, I, H or F, we want to know, and this page will be corrected rather than defended. A failing observatory. On the Antarctic case specifically: a published magnetic observatory record from a station south of 60°S whose measured declination departs from IGRF-14 by more than a few degrees under geomagnetically quiet conditions, and which cannot be attributed to a known local crustal anomaly. That would falsify the audit the previous section rests on. The eight observatories in the WMM2025 fit are the obvious places to look, and the data are public. A rival field model. A quantitative field model — of any geometry, on any surface — fitted to the same 140 observatory series and the same Swarm data, with a stated coefficient count and a stated RMS error budget, that meets or beats WMM2025's 0.36° global declination RMS. Not an argument that one could exist. A model, with coefficients, that can be run. If the flat programme produced one, the Lakatosian asymmetry we have claimed would evaporate and we would have to withdraw the argument in the last two sections entirely. Another double standard of ours. Evidence that this review has applied asymmetric scrutiny on some specific page — a place where we accepted a mainstream number without the check we demanded of his. We have already found one such instance in our own research and disclosed it above. We expect there are others.We will publish those corrections rather than defend them, and we will name the pages. That is the only version of this argument we are entitled to make.