by John Pabst
In the late 19th century, weather observation was a decentralized, non-standardized task performed by a vast network of disparate nodes. Data were collected by postmasters, lighthouse keepers, and railroad agents—individuals whose primary professional responsibilities were administrative or logistical, not meteorological. While these observers were often diligent, they operated within a system that lacked the metrological infrastructure we require today. Observations were recorded without the benefit of universally standardized shielding, precise, synchronized time-stamping, or consistent siting requirements. Consequently, the temperature was often logged based on the equipment and location that was practically available at each specific site, rather than through a uniform, instrument-grade protocol.
The observations were real, but the measurement system was highly heterogeneous. Instruments, exposure conditions, and recording procedures evolved and varied over decades. If you look at the raw data from that era, it is a patchwork of gaps, “round-number” biases, and inconsistent reading times. The challenge is not whether temperatures were measured, but whether those measurements can be combined into a century-scale record with the degree of accuracy often implied in public discussion.¹
The “Fix”
To turn that chaotic journal into a clean, global trend line, modern climate agencies have to perform a massive act of translation. The original thermometer reading goes in one side of the process; the final temperature record comes out the other. Between the two sits what most people never see: a black box of adjustments, reconciliations, and statistical corrections designed to account for station moves, changing instruments, urban growth, and missing data.

Here is the problem: while known biases can be estimated and adjusted, information that was never measured cannot be fully recovered. If a thermometer in 1890 was sitting on a sun-drenched porch in a growing town, no algorithm in 2026 can know with certainty how much that reading differed from the true ambient air temperature.² The black box may produce a cleaner record, but confidence in the final trend depends not only on the original measurements—it also depends on the assumptions embedded in the reconstruction process.³
The Hidden Uncertainty
We are told these records are accurate to within fractions of a degree—typically cited as an uncertainty interval of approximately ±0.15°C for the late 19th-century global mean. However, we must be precise about what this figure actually represents. This is not a measurement of physical reality; it is a calculated margin of error derived from the homogenization model itself.
In other words, this uncertainty value is a report card on the algorithm’s internal consistency, not an audit of the Earth’s temperature. It measures how well the model believes it has “corrected” the data, not the degree of truth in the original observations. This budget assumes that the errors are random and can be smoothed away, but it cannot account for systemic biases—such as the widespread conversion of land use—that the model itself may be blind to. Statistical methods can improve the model’s precision, but they cannot confer physical accuracy onto a century of heterogeneous, unverifiable data.
In 1979, everything changed. We launched the first global satellite sensing systems. For the first time, we had a continuous, globally consistent, and physically traceable way to measure the temperature of the atmosphere. We moved from an era of anecdotal, human-mediated estimates to an era of systematic, physical observation.⁴
The Question of Confidence
This brings us to the central tension in climate policy. Should we treat the “human-journal” record of the 19th century as if it were the same quality as the “satellite-sensor” record of the 21st?
This is where the public narrative falters. We are routinely presented with a single, seamless temperature graph that ignores the fundamental shift in measurement architecture that occurred in 1979. Treating the pre-1979 “adjusted” surface record and the post-1979 “satellite-sensor” record as a single, consistent data stream is a category error. They are two fundamentally different forms of evidence: one is reconstructed from heterogeneous historical observations whose original conditions often cannot be independently verified, while the other is produced by a globally consistent sensing system whose measurements, telemetry, calibration records, and processing steps can be audited and reproduced. When we conflate the two, we conceal the difference between a record that must be reconstructed and a record that can be systematically verified.
If this were any other high-consequence system—like the design of a bridge, the safety of a drug, or the management of a pension fund—we would demand a clear separation between “anecdotal estimates” and “verified facts.” We would insist on knowing exactly how much of our warming trend is observation and how much is statistical correction.
My argument is not that the historical record is “fake.” My argument is that it is imprecise. By pretending that a messy, century-old volunteer log has the same authority as modern satellite data, we are masking the true level of uncertainty in our climate models. We are making multi-trillion-dollar decisions based on a ledger that is far “fuzzier” than the public is led to believe.
The thermometer is an observation. The black box is an interpretation.
If this were any other high-consequence system—like the design of a bridge or the management of a pension fund—we would demand a clear separation between ‘anecdotal estimates’ and ‘verified facts’. We would insist on knowing exactly how much of our warming trend is observation and how much is statistical correction. We don’t need to throw the history away, but we must stop treating an interpreted ledger as if it were a direct physical observation. Until we differentiate between what we can verify and what we have ‘fixed,’ we are not managing a climate—we are ignoring the limitations of our own instruments.
Notes
¹ For a discussion on the limitations of 19th-century observational networks, see James R. Fleming, Meteorology in America, 1800–1870 (Baltimore: Johns Hopkins University Press, 1990).
² Matthew J. Menne et al., “On the Reliability of the U.S. Surface Temperature Record,” Journal of Geophysical Research: Atmospheres 115, no. D11 (2010).
³ Ross R. McKitrick, “On the Adjustment of Surface Temperature Records,” Energy & Environment 21, no. 8 (2010): 945–966.
⁴ John R. Christy and Roy T. McNider, “Satellite Bulk Tropospheric Temperatures as a Metric for Climate Sensitivity,” Asia-Pacific Journal of Atmospheric Sciences 53, no. 4 (2017): 511–518.


