Defining Temperature

I never thought I would participate in a serious debate on the definition of “temperature.” But it happened on twitter, and with people who have degrees in physics and other hard sciences! Since I worked with kinetic and effective temperature estimates for 42 years as a petrophysicist, I knew these topics intimately and never gave them a second thought. So, to hear people say (paraphrasing) that they were not really temperatures came as a shock. Their proposed sole definition of “temperature:” is the thermodynamic statistical definition:

Equation 1:

Where T is temperature, U is the internal energy, S is entropy, ∂U/∂S is the partial derivative of internal energy with respect to entropy, N is the number of particles, and V is the volume. The derivative is taken while keeping the volume and number of particles fixed, thus it is an equilibrium temperature and has no meaning outside the laboratory. Besides, entropy is a well-defined state function only in equilibrium thermodynamics; outside equilibrium, entropy can be defined statistically (see Boltzmann), but it is not unique and does not serve the same role.

Equation 1 can be called a definition of the “equilibrium temperature,” and when a system is in equilibrium it is an exact definition. It follows directly from the first law of thermodynamics. The first law states that energy cannot be created or destroyed, only converted from one form to another. That said, the first law is not the sole definition of temperature, as the word is used today. There are many other temperature definitions in common usage as well as in science. Let’s look at some of them:

Kinetic temperature

You can measure the instantaneous temperature of molecules in a shock wave, a plasma, a turbulent flow, or a laser-heated gas. It is a function of the mean kinetic energy of the molecules in the gas (or other fluid) (Reif, 2009). This temperature is widely used in shock physics, plasma physics, petrophysics (my former profession), and in molecular dynamics simulations. It is a valid temperature that does not require thermodynamic equilibrium.

Effective temperature

There are some systems that have no equilibrium temperature at all. Yet, we routinely define an effective temperature for these systems (Herzberg, 1950). These include turbulence, granular materials, active matter, moving dissipative systems, etc. These are legitimate temperatures and useful.

Non-equilibrium systems have a temperature, sometimes more than one depending upon the parameters used. For example, a molecular beam can have a vibrational or rotational temperature and they can be different. Plasmas may have an electron temperature or an ion temperature. NMRs have a spin temperature. Electronic circuits have a noise temperature. These temperatures are meaningful and important especially when the system is far from equilibrium.

Local temperature

This is the temperature used in climate and meteorological studies. It is the foundation of heat conduction, and hydrodynamics. One assumes local thermodynamic equilibrium over some volume and measures or assumes a local temperature for the volume. This is the logic used in the famous Navier-Stokes equations. This assumed “local equilibrium” is valid over small volumes (for example an “air parcel”) for short time periods. The problem with some climate models is that they assume “local equilibrium” for volumes and time periods that are too large and too long.

Discussion

Temperature is not a primitive mechanical property like mass; it is an emergent statistical property that characterizes the distribution of energy among degrees of freedom (Landau & Lifshitz, 1980). Mass can be measured independently of the rest of the universe or system and does not change as the system around it changes. Temperature follows from the 0th law of thermodynamics, which says:

If system A is in thermal equilibrium with system B,

and system B is in thermal equilibrium with system C,

then A is in thermal equilibrium with C.

The 0th law establishes that “temperature” is a meaningful physical quantity because thermal equilibrium is transitive. It does not define temperature, but it guarantees that temperature is meaningful and measurable. We often hear that temperature is transitive, but this is only true in laboratory settings or at equilibrium. Outside equilibrium, temperature can lose transitivity and different degrees of freedom can produce different temperatures. Transitivity is the basis of equilibrium temperature, but not all temperature measurements.

In summary, the thermodynamic definition of temperature applies only to equilibrium states, but physics uses many other temperature concepts like kinetic, effective, local, and generalized that are essential for describing real systems far from equilibrium. Restricting “temperature” to its equilibrium definition ignores the vast range of physical systems where temperature is well-defined and indispensable.

Some additional non-equilibrium temperature references are cited and discussed here.

Bibliography

Herzberg, G. (1950). Molecular Spectra and Molecular Structure. Volume I: Spectra of Diatomic Molecules. Second Edition. D. Van Nostrand.

Landau, L. D., & Lifshitz, E. M. (1980). Statistical Physics, Third Edition. Butterworth-Heinemann.

Reif, F. (2009). Fundamentals of Statistical and Thermal Physics. Waveland Press, Inc.

The climate data they don't want you to find — free, to your inbox.
Join readers who get 5–8 new articles daily — no algorithms, no shadow bans.
5 22 votes
Article Rating
737 Comments
Sweet Old Bob
August 9, 2026 6:29 am

“a serious debate on the definition of “temperature.”

Knowing more and more about less and less until they know everything about nothing ?

😉

Kevin Kilty
Reply to  Sweet Old Bob
August 9, 2026 7:15 am

Very funny. Now for the serious part. Pedagogy.

I received a B.Sc. in Physics in 1975. Physics was taught out of Zemansky the last quarter of Junior year. The treatment of thermo in Zemansky any engineer would still recognize in modern engineering thermo texts like Cengel and Boles. In postgraduate physics thermo was being taught as statistical physics, my textbook was Thermal Physics by Morse. The reason for concentrating on statistical physics by the time one reaches post graduate education is that matter is composed of molecules and, by gawd, we need all of physics to be consistent with the molecular picture of matter. So far, I agree.

Trouble is that dependence on statistical physics has crept into the undergraduate curriculum. Thermo is not taught the way it was. The pedagogy has changed. Assembling a reasonably useful statistical description of some ensemble of particles is way more difficult than anything else. So, student who once knew some thermo by the time they got a B.Sc. are a bit lost now in my view that they are expected to find state functions from statistics.

In 2016 or so I had trouble finding a T.A. for the thermal-fluids laboratory class (it’s an MechE junior level course) I was teaching. I thought to myself that I could just go get a physics graduate student for my T.A. and that would work as well as a graduate student in ME. What resulted was that he knew very little fluid mechanics and not much of thermo either. I mean he’d never seen Bernoulli’s equation and was pretty baffled about calculating heat content or its changes from temperature and what else he could reasonably surmise about a system (I.e the difference between enthalpy and internal energy).

Knowing way too much about very little is not helpful. Pedagogy is important!

2hotel9
Reply to  Kevin Kilty
August 9, 2026 12:26 pm

Almost sounds as if education has been sabotaged from within, not by outsiders.

Reply to  Kevin Kilty
August 9, 2026 5:29 pm

…he’d never seen Bernoulli’s equation and was pretty baffled about calculating heat content or its changes from temperature …

We learned that in high school physics and regurgitated it in Physics 1.

Reply to  Sweet Old Bob
August 9, 2026 1:33 pm

When your story isn’t selling, change the meanings of words.

Sparta Nova 4
Reply to  Retired_Engineer_Jim
August 10, 2026 6:06 am

Example: The Trans-Reality Alarmist lexicon.

Kevin Kilty
August 9, 2026 6:41 am

U is the internal temperature

Typo, Andy. U is internal energy.

What is measurement of themperature useful for? Answer that question as it will lead you to the definition of themperature appropriate for any particular use. Local temperature seems fine. It’s defined over large enough ensemble to behave like true equilibrium temperature, it provides a reasonable parameter for calculation of internal energy, enthalpy, and so forth. Its gradient provides a potential for calculating heat transport sufficiently accurate for all but extreme engineering.

Reply to  Andy May
August 9, 2026 9:03 am

When one can travel sometimes less than a mile on a clear day and see the car thermometer change 1 to 4 degrees, one must question the thermodynamic equilibrium assumption. Insolation is never in equilibrium. A simple look at a temperature graph of either soil or air temperature one can see that equilibrium is seldom happening unless one declares a small change of temperature as equilibrium.

I have graphed enough insolation and 2 meter surface temperatures over land to see that maximum insolation occurs about two hours prior to Tmax. Something is not in equilibrium for that to occur. Gradients rule.

Reply to  Jim Gorman
August 9, 2026 3:22 pm

I had to go back 1000 years to be certain I had identified the period of minimum tropical sunlight for the present precession cycle. The minimum occurred in the 13th century. By all recorded accounts, it was about 300 years before Europe reached its lowest winter temperatures.

The increase in tropical sunlight is now around 3E18 joules per year, which is not much when the current tropical sunlight is up around 3.1E27 joules per year. But give it a few centuries and that small change is observed to have a measurable impact on mostly atmospheric moisture over the tropics, which increases poleward heat transport thereby increasing most of the surface temperature.

So thermal lag is a big factor in understanding what is happening in Earth’s climate.

If you make the wild assumption that the only factor causing temperature to increase is carbon combustion then you are bound to make the wrong assumption that Earth was in an equilibrium state before carbon combustion started ramping up.

Kevin Kilty
Reply to  Andy May
August 9, 2026 9:11 am

Andy, you may have heard this but engineers say that averages are the basis of bad designs. Physical scientists could pay heed to the same.

Temperature is more difficult to state accurately than many people realize. I often argued with geophysicists because they were seeing significance in borehole analyses (yes, those seeking to measure past climate) depending on temperatures allegedly resolved to 0.05K.

bdgwx
Reply to  Andy May
August 9, 2026 11:26 am

As I’ve said before…just because average temperatures have little value to you doesn’t mean they aren’t valuable to other people.

Reply to  bdgwx
August 9, 2026 12:56 pm

Yes, the value is often found in the motive.

Reply to  bdgwx
August 9, 2026 1:19 pm

The question is, are temperature averages scientifically useful. Averages may appeal to you as a mathematician and statistician, but to a physical scientist temperature averages are useless in a constantly changing multi-variable world. Temperature averages, especially those without a statement of quality, called uncertainty, have no meaning. How well a mean has been calculated as evidenced by an SDOM where the SD has been divided by √n simply does not provide the necessary information that describes the variance of the measurements used to calculate the mean.

The GUM says this in its very first paragraph.

0.1 When reporting the result of a measurement of a physical quantity, it is obligatory that some quantitative indication of the quality of the result be given so that those who use it can assess its reliability. Without such an indication, measurement results cannot be compared, either among themselves or with reference values given in a specification or standard. It is therefore necessary that there be a readily implemented, easily understood, and generally accepted procedure for characterizing the quality of a result of a measurement, that is, for evaluating and expressing its uncertainty.

The word obligatory means something to those who deal with physical measurements. The very fact that averaging measurements of observations of the same thing leaves doubt as to what the actual true value is carries a great weight. If you don’t know the amount of doubt there is no way to know the value of the measurement.

A “global average temperature” is a construct that has no physical meaning. It is not an accurate depiction of the earth’s energy because it ignores latent heat energy. From a birds eye view of the earth, if temperature is a good proxy for energy, then latent heat would have to be falling in order to maintain an energy balance. Can you show this?

bdgwx
Reply to  Andy May
August 9, 2026 2:44 pm

Temperature cannot be averaged in a valid way and if it is, the average has little meaning.

That is an absurd statement. Temperature is averaged in valid ways used for meaningful purposes all of the time.

Reply to  bdgwx
August 9, 2026 6:44 pm

That is an absurd statement. Temperature is averaged in valid ways used for meaningful purposes all of the time.”

This has been explained to you ad infinitum.

  1. Multiple measurements of the same thing using the same instrument under repeatable conditions can be averaged to find a “BEST ESTIMATE” for the value of the measurand. In order for this to work the distribution of the measurements should be random and Gaussian. Otherwise the average is *NOT* the “best estimate” of the value of the measurand. This is *NOT* averaging temperatures, it is averaging a set of measurements of a measurand.
  2. Single measurements of multiple things using different instruments under non-repeatable conditions is attempting to average the temperatures and not the measurements. And you certainly cannot make the assumption that the distribution of those single measurements will be random and Gaussian. The average is almost certainty *NOT* an accurate statistical descriptor of the distribution of the measurements.

In order to average measurements and get a value that has physical meaning the measurements must be able to be combined into a single larger system, e.g. masses – an extensive value.

Averaging measurements that can *NOT* be combined into a larger system does not provide a physically meaningful value. Two rocks, one at 270K and the other at 250K, held in your hand does *NOT* give a a larger system at 520K. The “average” value of 260K is physically meaningless.

And this doesn’t even begin to address the thermodynamics of moist air with sensible heat that can be measured and latent heat that cannot be measured. The enthalpy (dry air + wet air) of a parcel of air over a sandy desert will be vastly different than the enthalpy of that same sized parcel of air over the ocean. Trying to “average” the air temperatures of each together is a fool’s errand. It’s why throwing the air temperature in Phoenix and the air temperature in Miami into the same data set and expecting to get a physically meaningful average is impossible.

All you are doing here is confirming for everyone that you simply don’t care about obtaining physically meaningful results. It’s all “random, Gaussian, and cancels” coupled with “number is just numbers”.

Reply to  bdgwx
August 9, 2026 9:29 pm

Is the average of mid-range values as useful as the average of a time-series with consistent intervals of sampling? Can one tell from mid-range averages weather the population is skewed, or what the daily variance is?

Reply to  Andy May
August 9, 2026 6:55 pm

Temperature is not a true physical property and only has meaning in comparison to other temperatures.”

Temperature can’t even be compared to other temperatures unless they are in similar environments. Comparing the temperature in Phoenix with the temperature in Miami is meaningless since the environments (humidity, pressure, geography, terrain, etc) are vastly different.

Reply to  Tim Gorman
August 9, 2026 7:52 pm

Have you never heard of the 0th law? Two things with the same temperature are the same temperature. All the other things you mention are not temperature.

Reply to  Bellman
August 10, 2026 3:20 am

Have you never heard of the 0th law? Two things with the same temperature are the same temperature. All the other things you mention are not temperature.”

WHY do you insist in coming on here and lecturing people about things you have absolutely no understanding of?

The 0th law has to do with energy TRANSFER, not with energy content.

All the other things you mention are not temperature.”

It is energy CONTENT that drives the heat engine known as Earth, not temperature. Trying to use temperature as a proxy for energy content is just one more idiotic, garbage assumption that climate science makes that you just gobble up.

You simply cannot ignore latent heat and its effect.

Moist air has latent heat as well as sensible heat. It is the latent heat that is a large driver of convection, i.e. moist air rises faster than dry air. It is latent heat and its release that is a large driver of storms and precipitation.

Energy transfer, in fact, has more factors than just temperature. Cooling, i.e. the transfer of energy out, consists of conduction, convection, radiation, and evaporation (latent heat).

Two objects can be at the same temperature yet the energy flux can be different, meaning they are not actually in thermal equilibrium. Take evaporation as an example. Evaporation from an object occurs more rapidly in dry air than in moist air even if both objects are at the same temperature.

Bottom line: the 0th law assumes DRY AIR surrounding all three systems. No latent heat. No evaporation. Equal conduction, convection, and radiation. The 0th law only applies when the involved systems are restricted to exchanging heat, not mass. I.e. closed systems.

This is hardly the situation in the biosphere.

Reply to  Tim Gorman
August 10, 2026 4:18 am

“The 0th law has to do with energy TRANSFER, not with energy content. ”

And temperature is not energy content.

“It is energy CONTENT that drives the heat engine known as Earth, not temperature. ”

But you were the one talking about temperature, remember?

Temperature can’t even be compared to other temperatures unless they are in similar environments.

Reply to  Bellman
August 10, 2026 4:46 am

But you were the one talking about temperature, remember?”

So what? My statement is true!

tpg: “Temperature can’t even be compared to other temperatures unless they are in similar environments.”

It’s because the systems of Earth are open systems and not closed systems. The systems are characterized by MORE than just their temperature – they can exchange mass as well as energy. The 0th law has only to do with heat exchange between CLOSED systems.

As usual, you come on here making assertions that only highlight how little you actually know on the subject. It stems from your continued cherry picking instead of actually studying for meaning and context.

Reply to  Tim Gorman
August 10, 2026 5:34 am

“My statement is true!”

Well if you say it’s true I suppose it must be.

“The systems are characterized by MORE than just their temperature…”

Again, you were talking about temperature.

Of course the effects of temperature will depend on many factors. That’s why you have weather reports stating “feels like temperature”. But that doesn’t mean you can’t compare the actual temperatures, however you define it.

Reply to  Bellman
August 10, 2026 10:29 am

Well if you say it’s true I suppose it must be.”

It *is*true and you have absolutely no refutation to show otherwise.

“Again, you were talking about temperature.”

No, YOU were talking about temperature when you brought up the 0th law of thermodynamics.

I am talking about INTENSIVE PROPERTIES. Temperature is an intensive property and I used it as an example. I could have used density instead. Nothing I said would be any different.

An intensive property is independent of the size of the system. Therefore you cannot add intensive properties of multiple systems to create a larger system. If you can’t add systems to create a larger system then you can’t calculate a physically meaningful average from the larger system by equally subdividing the larger system.

Some day you REALLY should migrate out of statistical world and join the rest of us in the real world.

Reply to  Andy May
August 10, 2026 1:18 pm

“Outside equilibrium (eg. the whole world) it loses it intensivity.”

Really? That doesn’t seem right.

“Thermodynamically, the Earth has no temperature, average or otherwise.”

I think too many people have a problem with the concept of an average. They want it to be a property of the things being averaged, but that isn’t usually the case. An average height is not a height it’s the average of heights. An average is a statistic. An average temperature is not a temperature, it’s an average of temperatures.

The average surface temperature is not the temperature of the earth, it’s the average of the surface temperature across the earth. The main use of an average is to compare populations. For surface temperature this can mean seeing if some parts of the earth are hotter than others, or seeing if the temperature, on average, is changing over time.

Reply to  Bellman
August 10, 2026 2:19 pm

The average surface temperature is not the temperature of the earth, it’s the average of the surface temperature across the earth.

You admit that the GST is an average. An average only exists when you have a distribution of information. That distribution has both a mean AND A STANDARD DEVIATION. Why do you not post the SD ever. Don’t use the excuse that none of the standard sources don’t list it. If that is all you have, you shouldn’t use those sources.

Reply to  Jim Gorman
August 10, 2026 2:40 pm

You admit that the GST is an average.

There’s a clue when it’s described as the average global surface temperature.

That distribution has both a mean AND A STANDARD DEVIATION.

It can have as many properties as you like. But the mean is the most useful in comparing populations.

Why do you not post the SD ever.

I’ve posted lots of SDs. But their usefulness depends on which distribution you are talking about.

Reply to  Bellman
August 10, 2026 2:45 pm

The average surface temperature is not the temperature of the earth, it’s the average of the surface temperature across the earth. 

It just dawned on me what this really means. You have just destroyed the use of average temperatures as inputs to the Stefan-Boltzmann and Planck equations. Radiative calculations are meaningless if measured temperatures are not used.

Here is what Andy said.

 Outside equilibrium (eg. the whole world) it loses it intensivity. Then “temperature” becomes ambiguous, can differ between subsystems, and loses the strict meaning it has in equilibrium thermodynamics. Thermodynamically, the Earth has no temperature, average or otherwise.

You just made his case. When temperature becomes ambiguous it loses its meaning.

Why are spending billions on reference stations and supercomputers for modeling. Let’s just use averaged temperatures from Weather Underground stations. I mean if accurate measurements with low uncertainty are not necessary, why bother? You can get as much resolution as you need from a pocket calculator in order to find a signal. Numbers is just numbers, right?

bdgwx
Reply to  Andy May
August 11, 2026 5:46 pm

The CMB isn’t in equilibrium either. That doesn’t mean we can’t measure its average temperature of ~2.73 K and find meaning in it.

Reply to  bdgwx
August 11, 2026 6:18 pm

The CMB (Cosmic Microwave Backround) is not a measured temperature. The “temperature” is derived similar to UAH from microwave EM waves. It was found to be isotropic, the same wherever the antenna was pointed in the sky. That led to the conclusion that it is a “background” phenomena.

It was discovered by Arno Penzias and his partner when they could not get the noise factor of their microwave receiver below a certain value yet their test showed it was working properly. I got to meet Dr. Penzias once at Bell Labs when I was there for another unassociated project.

bdgwx
Reply to  Jim Gorman
August 12, 2026 12:21 pm

First…the CMB average temperature is measured. That’s how we know it is ~2.73 K. We measured it.

Second…the CMB is not in equilibrium; at least not global equilibrium. We know this because of the mapping COBE and WMAP did.

Third…the average temperature is actually exploited along side the dipole anisotropy to measure Earth’s velocity relative to the CMB of about 370 km/s. The dipole’s deviation from average is about 0.1% or 0.003 K.

Fourth…regarding Penzias and Wilson…duh. Everybody knows the story.

Fifth…I bring up the CMB for a reason. Andy said that climate models assume local equilibrium that is too large and too long. But if the local equilibrium assumptions made while analyzing the CMB aren’t considered too large and too long then how could the relatively miniscule scales of climate models be?

Reply to  bdgwx
August 12, 2026 7:03 pm

First…the CMB average temperature is measured. That’s how we know it is ~2.73 K. We measured it.

You can bluff but you must show your cards. Tell us the exact device that directly measured the CMB temperature.

Give us a reference that describes how that device “measures”, not calculates, the temperature directly.

I suppose you think an infrared thermometer actually directly measures temperature, right? How about UAH? Do the satellites directly measure temperature like a thermometer?

Fourth…regarding Penzias and Wilson…duh. Everybody knows the story.

If you know the story, then relate how they estimated the temperature of the background EM waves. That should give you some thought about direct measurement of temperature.

But if the local equilibrium assumptions made while analyzing the CMB aren’t considered too large and too long then how could the relatively miniscule scales of climate models be?

You are showing your lack of knowledge here.

First, local thermodynamic equilibrium of the atmosphere has absolutely nothing to do with using a parabolic antenna to intercept microwave EM waves from outer space. The antenna is NOT measuring “temperature”, it is measuring the intensity of the EM signal.

Second, I don’t think you have a clue about what isotropic EM radiation actually is. In this scenario it means the EM signal is ubiquitous and has the same characteristics. You can point that parabolic antenna at any point in the sky, and it will receive the same signal.

Lastly, you have given the impression that you understood the operation of the satellites used by UAH. Does the term microwave sounder (MSU) ring a bell with you? You don’t appear to understand what they actually intercept. They intercept microwave EM radiation from O2 molecules. Got that? Microwave EM radiation. Not temperature, EM intensity.

bdgwx
Reply to  Jim Gorman
August 14, 2026 5:59 pm

Tell us the exact device that directly measured the CMB temperature.

I already did. COBE and WMAP are 2 examples. There are others.

Give us a reference that describes how that device “measures”, not calculates, the temperature directly.

I suppose you think an infrared thermometer actually directly measures temperature, right?

Just so we’re clear all temperature measurements require a measurement model that involves some kind of calculation.

And your insinuation that a measurement cannot involve a calculation is contrary to what the GUM says.

You are showing your lack of knowledge here.

I lack a lot of knowledge. And the more I learn the more I realize how much knowledge I lack.

Reply to  bdgwx
August 15, 2026 1:05 pm

Just so we’re clear all temperature measurements require a measurement model that involves some kind of calculation.

Just so we are clear, measurement devices like LIG and RTD directly change in response to temperature. Their accuracy of indication relies on a calibration procedure. Thus, after proper calibration, they read temperature directly. No calculation is used other than perhaps a correction whose value is determined during calibrating occurs.

Measurement devices that change in response to radiation intensity do not directly read temperature. A temperature must be calculated based upon assumptions of the values of the various constants that may or may not meet reality. Thus, temperature is an estimate and is not read directly.

Tom Johnson
Reply to  Bellman
August 11, 2026 4:03 am

seeing if the temperature, on average, is changing over time.”
The illusion of an “average earth temperature is meaningless. You’re averaging a system where most of the population lives in only one hemisphere while the other hemisphere is mostly ocean, with a tropical zone that has thunderstorm temperature regulation, with polar regions that are out of annual phase where one goes up while the other goes down, and one is is mostly floating ice and the other is mostly high elevation glaciers,…and more. An average of that is nothing more than a farce. How can you possibly consider a climate “tipping point” to such a meaningless number. Even the UAH satellite data has an implied surface area weighting factor which hardly makes the number any better considering the above caveats.

Reply to  Tom Johnson
August 11, 2026 5:07 am

Temperature is *not* climate. It’s probably the worst metric you can use for climate.

bdgwx
Reply to  Tom Johnson
August 11, 2026 5:38 pm

The illusion of an “average earth temperature is meaningless.

As I keep saying…just because it is meaningless to you does not make it meaningless to everyone else.

Anyway, it might surprise you to know that many of the temperature values you see in everyday use are actually averages. This includes your daily weather reports. Yep, those airport temperature reports are actually averages. Surely you can find at least some meaning in those.

Reply to  bdgwx
August 11, 2026 7:07 pm

Yep, those airport temperature reports are actually averages. Surely you can find at least some meaning in those.

You refuse to learn and just keep on plugging your unsupported positions.

ASOS (at airports) take measurements at seconds intervals and average them to obtain a five minute value.

Guess what that is measuring, THE SAME volume multiple times in a short period of time. Look up repeatable conditions in the GUM. Just like measuring a stirred water bath at different locations with the same device. Those are all special conditions for allowable temperature averages.

How about averaging ASOS stations in Chicago O’Hare and Reagan International and publishing the average to pilots. You reckon that is a meaningful average?

Your refusal to acknowledge the difference between statistical operations and physical science requirements in thermodynamics buys you no points.

Reply to  Bellman
August 11, 2026 4:26 am

“I think too many people have a problem with the concept of an average. “

The mean of a distribution (i.e. the average) is a STATISTICAL DESCRIPTOR. It describes the distribution and does *NOT* describe anything to do with the values or properties of the individual data items making up the distribution.

An average is a statistic.”

Meaning it is *NOT* a measurement. How precisely you can locate the population mean of a set of measurements has *nothing* to do with the uncertainty associated with those measurements.

The average surface temperature is not the temperature of the earth, it’s the average of the surface temperature across the earth.”

You got *SO* close with your statement of “An average is a statistic”. And then you turn around and fail again. Cognitive dissonance at its finest.

You cannot average intensive properties of systems that can’t be added to form a larger system. That means the average of a set of temperatures HAS NO PHYSICAL MEANING, the average simply doesn’t exist. It is *NOT* the average of the surface temperature across the earth. It’s an application of the meme in Statistical World of “numbers is just numbers”.

The main use of an average is to compare populations. For surface temperature this can mean seeing if some parts of the earth are hotter than others, or seeing if the temperature, on average, is changing over time.

You can *NOT* compare surface populations of temperatures using an average value. You can’t use the average temperature to even determine if some parts of the earth are hotter than others. Temperatures in Phoenix and Miami can be exactly the same yet you simply cannot compare the two in order to claim that they are both equally “hot”. They are *NOT* equally “hot” since the latent heat component in Miami is almost always greater in Miami than in Phoenix.

Nor can the average of a set of numbers be depended upon to identify if changes in temperature have occurred. The average of the numbers 60 and 70 is 65. The average of the numbers 55 and 75 is 65. The numbers can change significantly yet the average will stay the same. Looking at the average simply won’t tell you if anything has changed.

Couple this with the fact that temperature is a minor factor in climate compared to precipitation, and temperature is a terrible metric for climate. The average temperature of the central US savannah is significantly different than the average temperature of the savannahs of central Africa – yet they are both savannah type climates, very similar in both flora and fauna in their natural states with large populations of grazing herbivores. Primarily because the precipitation is similar in both geographies.

The “average” of intensive properties tells you literally NOTHING. It is PHYSICALLY MEANINGLESS.

Reply to  Tim Gorman
August 11, 2026 5:20 am

“The mean of a distribution (i.e. the average) is a STATISTICAL DESCRIPTOR”

Yes, that’s what I said.

“It describes the distribution and does *NOT* describe anything to do with the values or properties of the individual data items making up the distribution.”

That’s just silly. Of course it has something to do with values of the individual data.

“Meaning it is *NOT* a measurement. ”

One day you are going to hVe to actually define what you mean by measurement, then explain why you are so obsessed with the measurement uncertainty of something you don’t believe us a measurement.

“How precisely you can locate the population mean of a set of measurements has *nothing* to do with the uncertainty associated with those measurements.”

Again, just silly. If the measurements are uncertain, that uncertainty will be reflected in the uncertainty of the average.

“You got *SO* close with your statement of “An average is a statistic”. And then you turn around and fail again. ”

Huh?

“You cannot average intensive properties of systems that can’t be added to form a larger system.”

You can repeat that nonsense as many times as you like. There’s only so many times I can explain to you why it’s nonsense.

“That means the average of a set of temperatures HAS NO PHYSICAL MEANING, the average simply doesn’t exist. ”

I’ll ask again, what do you mean by “physical meaning”, and for that matter by “exist”. Do you think the average of people’s heights exists, and has physical meaning?

“It is *NOT* the average of the surface temperature across the earth.”

You are saying the average if surface temperatures is not the average if surface temperatures? In that case what is it?

“It’s an application of the meme in Statistical World of “numbers is just numbers”. ”

I’ll say again that statistics world is the same as science world. Statistics describe the real world, and most of science for the past century or so understands the importance of statistical analysis.

Reply to  Bellman
August 11, 2026 5:44 am

“You can *NOT* compare surface populations of temperatures using an average value. ”

It might be beyond your limited mind set, but most people can understand that an average of 20°C is bigger than average if 18°C. Hence you can compare two averages.

“Temperatures in Phoenix and Miami can be exactly the same yet you simply cannot compare the two in order to claim that they are both equally “hot”. ”

Since the 17th century the word temperature has been used as a measure of “hotness”. Maybe you prefer the original meaning as a balance if the humours. Hot, cold, wet and dry.

One place may feel hotter because of the humidity or wind, but it does not mean it’s actually hotter. That’s why we invented thermometers so we didn’t have to use our subjective experience as a proxy for hotness.

“They are *NOT* equally “hot” since the latent heat component in Miami is almost always greater in Miami than in Phoenix. ”

You are basically trying to refine “hot” to include absorbed latent heat. What you really mean is specific enthalpy.

“Nor can the average of a set of numbers be depended upon to identify if changes in temperature have occurred. ”

But it’s a start. If the average changes it can only be because the there has been a change in some or all of the temperature set. The converse is not true.

“Couple this with the fact that temperature is a minor factor in climate compared to precipitation, and temperature is a terrible metric for climate. ”

Who said anything about climate? My claim is that an average surface temperature can tell you about changes in average surface temperature. If you want to look at changes in precipitation you need to look at average precipitation.

“The “average” of intensive properties tells you literally NOTHING.”

It might tell you literally nothing, but don’t confuse your own stupity with the rest of the population.

I look at the average UK temperatures for June and compare them with January. I see that June is on average a lot hotter than January. That tells me that I can expect June to be hotter than January. It’s literally told me something. I compare average UK temperatures with average Spanish temperatures. That literally tells me that Spain is generally a hotter country than the UK.

Reply to  Bellman
August 11, 2026 6:29 am

It might be beyond your limited mind set

As predicted yesterday…condescending superior bellman has arrived.

Reply to  karlomonte
August 11, 2026 7:36 am

ROFL! YEP!

Reply to  Tim Gorman
August 11, 2026 9:53 am

You can dish it out, but can’t take it.

Reply to  Bellman
August 11, 2026 6:59 am

 most people can understand that an average of 20°C is bigger than average if 18°C. Hence you can compare two averages”

Most people understand that 18C in Miami means something different than 20C in Phoenix!

Comparing the two averages tells you NOTHING that has any physical meaning.

This is just one more application of the meme: “numbers is just numbers” that you are so fond of.

Reply to  Bellman
August 11, 2026 7:01 am

Since the 17th century the word temperature has been used as a measure of “hotness”. Maybe you prefer the original meaning as a balance if the humours. Hot, cold, wet and dry.”

Not by physical scientists. Or engineers.

Reply to  Bellman
August 11, 2026 7:15 am

You are basically trying to refine “hot” to include absorbed latent heat. What you really mean is specific enthalpy.”

How many times have I told you that before? Enthalpy is what climate science should be using, not temperature!

Are you just now beginning to understand why I say that?

Latent heat EXISTS. Enthalpy (H) = h_d + h_w. It is TOTAL enthalpy that determines the internal energy of a parcel of atmosphere. h_w contains latent energy.

It is total energy that is the contributor to climate, not just h_d.

Temperature is only PART of the system. It’s why you can’t compare the air temperature in Phoenix with the air temperature in Miami and get a physically meaningful comparison. How “hot” the air is includes both sensible and latent heat, not just sensible heat.

I’m not “refining” or “re-defining” anything. Just like with metrology, several of us have been trying to educate you on basic thermodynamics for LITERALLY years. And just like with metrology, all you do is stubbornly cling to the same misconceptions you started out with!

Reply to  Bellman
August 11, 2026 7:19 am

But it’s a start. If the average changes it can only be because the there has been a change in some or all of the temperature set. The converse is not true.”

It’s *not* a start! If the converse is not true that means the average can remain stable even though the data values change. What the hell kind of a metric is that? It’s not fit-for-purpose!

Because of measurement uncertainty it’s not even possible to tell if the GAT has actually changed or not! What the hell kind of a metric is that? It’s not fit-for-purpose!

Reply to  Bellman
August 11, 2026 7:28 am

Who said anything about climate?”

“Climate change” is not using temperature as a metric?

My claim is that an average surface temperature can tell you about changes in average surface temperature.”

If the average surface temperature can *NOT* tell you whether a change has occurred, then how can it be a metric for identifying temperature change?

In fact, climate science doesn’t even use AVERAGES. It uses mid-range values. The daily mid-range value is determined from the range of the distribution of temperatures, it is *NOT* the statistical descriptor known as the “mean” or “average”.

Reply to  Bellman
August 11, 2026 7:32 am

That’s why we invented thermometers so we didn’t have to use our subjective experience as a proxy for hotness.

“They are *NOT* equally “hot” since the latent heat component in Miami is almost always greater in Miami than in Phoenix. ”

You are basically trying to refine “hot” to include absorbed latent heat. What you really mean is specific enthalpy.

It is pretty simple really. I know its hard for you to understand, but if you are going to deal with energy, especially energy by radiation, you must include latent heat in the equations. The law of conservation of energy is incomplete without using latent heat in the calculations.

If you believe that “hotness” is temperature only, then you also dismiss the use of “feels like” or heat index temperatures. Funny how much of the world wants to know the heat index, i.e. “hotness”, that you dismiss so blithely.

You have fallen back to using layman terms for scientific concepts. Tsk, tsk!

Reply to  Jim Gorman
August 11, 2026 9:41 am

“It is pretty simple really.”

It would be if you didn’t keep changing the subject. This was a discussion about temperature, but you keep wanting to turn it into one about energy.

“If you believe that “hotness” is temperature only”

It was a vague term for the purpose of sarcasm. The point however is that however you define temperature it is not taking into account potential energy. If you want to talk about specific enthalpy do so, but don’t call it temperature.

“then you also dismiss the use of “feels like” or heat index temperatures”

I don’t dismiss them. They are very useful, but again they are not actual temperature. I can just imagine the howels here if a met office started using feels like in place of temperature.

Reply to  Bellman
August 11, 2026 7:35 am

I look at the average UK temperatures for June and compare them with January. I see that June is on average a lot hotter than January.”

You can’t even get this one correct! The averages you are talking about are average MEASUREMENTS of the same thing!

It’s the same thing as TN1900!

It’s no different than taking 10 measurements of the same water bath to get a best estimate for the temperature of the water bath.

We can argue about whether the measurements are averages of intensive properties of physically different locations in the UK and are physically meaningful BUT averaging the average measurements is averaging the measurements and not the intensive properties themselves.

Reply to  Tim Gorman
August 11, 2026 9:51 am

“The averages you are talking about are average MEASUREMENTS of the same thing!”

Just zero internal consistency in your arguments. Everything’s special pleading. All averages are meaningless, except when you can see a meaning in which case they aren’t really averages.

So what are your rules for determine when an average of multiple places and times are measuring the same thing and when they are different? Is it just when it’s convenient for you and when it’s not?

And have you ever squared the circle of claiming that you can’t average intensive properties with the claim that you can if it’s the same thing? If an average of two different tickets is impossible because the sum isn’t physically meaningful, how can you measure the same rock twice and get an average? What is the physical meaning if the sum of two measurements of the same rock?

Reply to  Bellman
August 11, 2026 6:05 am

That’s just silly. Of course it has something to do with values of the individual data.”

A statistical descriptor has NOTHING to do with the individual data. The average does *NOT*, in any way, shape, or form determine the value of any individual piece of data. A statistical descriptor only describes the distribution.

One day you are going to hVe to actually define what you mean by measurement, then explain why you are so obsessed with the measurement uncertainty of something you don’t believe us a measurement.”

The measurement uncertainty is associated with the DATA, not with the average. The average is only used as a best estimate of the value of the measurand – and it is only such if certain, restrictive requirements are met.

If you would just once, JUST ONCE, stop and look at how a measurement is given, i.e.

“best estimate” +/- “measurement uncertainty”

it would be obvious to most people that the average value is only associated with the left-most component of the measurement, *NOT* with the right-most component.

*YOU* continue to try and pass off SAMPLING UNCERTAINTY as measurement uncertainty. But the sampling uncertainty only has to do with the value calculated for the best estimate, not with the propagation of measurement uncertainty.

Reply to  Tim Gorman
August 11, 2026 6:34 am

“A statistical descriptor has NOTHING to do with the individual data. ”

You are beyond help. Where do you think the statistics come from? Will the average of 1, 2, 3 be different to the average of 4, 5, 6? Do you think any difference be due to the difference in the Individual data?

“The average does *NOT*, in any way, shape, or form determine the value of any individual piece of data.”

Why would you expect it to?

“The measurement uncertainty is associated with the DATA, not with the average. ”

So? You either accept that an average can be considered a measurement of a measurand, and use the rules for figuring out the combined uncertainty for that measurement, or you can say it isn’t a measurement and there is no measurand, in which case, by your preferred definition, it’s not possible for the average to have a measurement uncertainty.

“it would be obvious to most people that the average value is only associated with the left-most component of the measurement, *NOT* with the right-most component.”

Which is the same for any measuremet taking from a function. The best estimate is based on the best estimates of the inputs to the function, and the uncertainty is propagated using the general equation for propagating uncertainties.

“YOU* continue to try and pass off SAMPLING UNCERTAINTY as measurement uncertainty.”

No I do not. What I keep trying to explain to you is that if you are treating an average as a random sample from a population, the uncertainty caused by the randomness of the sample will usually be much greater than the uncertainty caused by the uncertainty in the individual measurements. This follows from the fact that you would generally want to be measuring with an instrument capable of distinguishing the main variation in the population.

Exceptions may occur, in particular if the uncertainty of the measurements includes an unknown systematic error

Reply to  Bellman
August 11, 2026 8:12 am

You are beyond help. Where do you think the statistics come from? Will the average of 1, 2, 3 be different to the average of 4, 5, 6? Do you think any difference be due to the difference in the Individual data?”

Here we go again. The garbage meme that “numbers is just numbers”.

If I tell you that 1,2,3 have the unit of grams and 4,5,6 have the units of inches then can you compare the statistical descriptors known as the mean of the two distributions?

I know what you are going to say: “But you have to compare like things”.

And I’m going to answer that it is the individual elements that determine whether they are “like things”, not the average value of the distribution.

The statistical descriptor known as the mean does *NOT* have anything to do with the individual data elements, it only describes the distribution of the data elements.

You have two memes so baked into your brain that you can’t get around them, over them, under them or through them. They color literally *everything* you assert.

  1. measurement uncertainty is always random, Gaussian, and cancels and,
  2. numbers is just numbers.
Reply to  Tim Gorman
August 11, 2026 10:03 am

“Here we go again. The garbage meme that “numbers is just numbers”.”

They ate just numbers. Numbers produced from head to illustrate an example. If it helps imagine they are boards of a specific length.

“If I tell you that 1,2,3 have the unit of grams and 4,5,6 have the units of inches then can you compare the statistical descriptors known as the mean of the two distributions?”

Pathetic twisting. No, you can not compare a unit if length with a unit of mass.

“I know what you are going to say: “But you have to compare like things”. ”

Then why even try to dodge the issue. As always you are just trying to drag the conversation in any direction you can in order to avoid the obvious point.

Your assertion was that an average had nothing to do with the individual data. My example demonstrates why that is trivially wrong. This has nothing to do with comparing length and mass.

“The statistical descriptor known as the mean does *NOT* have anything to do with the individual data elements, it only describes the distribution of the data elements.”

What do you think, describes the distribution of the data elements, means? How is that not saying the mean depends on the data elements?

“You have two memes so baked into your brain that you can’t get around them, over them, under them or through them.”

The irony is too painfull. You lie about me having these memes, yet you repeat the same lies every time irrespective of the relevance to the subject. They are backed into your brain, not mine.

Reply to  Bellman
August 11, 2026 8:23 am

So? You either accept that an average can be considered a measurement of a measurand, and use the rules for figuring out the combined uncertainty for that measurement, or you can say it isn’t a measurement and there is no measurand, in which case, by your preferred definition, it’s not possible for the average to have a measurement uncertainty.”

ONE MORE TIME.

SAMPLING UNCERTAINTY IS *NOT* MEASUREMENT UNCERTAINTY!

A measurement is given as”

“best estimate” +/- “measurement uncertainty”

The average is associated with the best estimate. How precisely it estimates the population average is SAMPLING UNCERTAINTY.

The MEASUREMENT UNCERTAINTY has to do with the reasonable values that can be attributed to the value of the measurand. It is related to the variance of the distribution and It has NOTHING to do with the “best estimate” (i.e. the average) other than as a boundary condition. Your “best estimate” should be within the interval of reasonable values for the measurand or something is wrong somewhere.

The average does *NOT* have a measurement uncertainty. It has a SAMPLING uncertainty. They are *NOT* the same thing.

This is why some authorities on metrology have suggested just foregoing the “best estimate” and just giving the uncertainty interval in absolute values instead of as a standard deviation. I.e. instead of 10 +/- 2 grams just give 8-12 grams and the uncertainty interval. Then “sampling uncertainty” will no longer be confused with “measurement uncertainty” as you do all the time.

Reply to  Bellman
August 11, 2026 8:32 am

So? You either accept that an average can be considered a measurement of a measurand, and use the rules for figuring out the combined uncertainty for that measurement, or you can say it isn’t a measurement and there is no measurand, in which case, by your preferred definition, it’s not possible for the average to have a measurement uncertainty.

Going off the deep end here.

As Tim has told you, the mean is determined from a distribution of multiple measurements of the same thing. However, the “mean” is only a BEST ESTIMATE and then only when the distribution is Gaussian or mostly symmetric.

Read Chapter 2 of Bevington. There are many different types of distributions and that chapter covers 4 of them, binomial, Poisson, Gaussian, and Lorentzian. There are many more such as uniform, and triangular. In a lot of them the mean is not the most likely estimate of the measurand.

You only ever argue about the Type A evaluation using statistics. There are other influence quantities that are also part of the uncertainty evaluation. the WMO classifies station into classes. The lower classes can have extra uncertainty added up to 5°C. Anthony’s studies have determined there are a consequential number of stations like that in the U.S and other have done so in Great Britian.

Here are questions for you to answer. How are those uncertainties weighted into station, local, regional, and global data series? How are they weighted into baseline and anomaly averages?

Reply to  Bellman
August 11, 2026 9:26 am

Which is the same for any measuremet taking from a function. The best estimate is based on the best estimates of the inputs to the function, and the uncertainty is propagated using the general equation for propagating uncertainties.”

The average is *NOT* a function, it is a statistical descriptor of a distribution. The average is *NOT* a measurement, it is STATISTICAL DESCRIPTOR of a distribution.

A statistical descriptor equation is not a FUNCTION, it is a mathematical operator, little different than >.

It is called a “functional” relationship, a mapping of a distribution to a value. The uncertainty associated with that mapping is SAMPLING uncertainty, not measurement uncertainty.

“No I do not. What I keep trying to explain to you is that if you are treating an average as a random sample from a population, the uncertainty caused by the randomness of the sample will usually be much greater than the uncertainty caused by the uncertainty in the individual measurements”

Did you read this before you hit post?

Randomness of the sample is the SEM. SD/sqrt(n) where SD is the SD of the distribution.

Measurement uncertainty is the SD of the distribution..

You are basically saying that SD/sqrt(n) is always greater than SD.

Uncertainty is a variance. You add variances just like you add uncertainties. Var_total = (VAR1 + VAR2 + ….) So SD = sqrt(Var_total). Just like u_total^2 = RSS( u1,u2,….)

In other words, UNCERTAINTIES ADD.

So SD ≥ SD/sqrt(n) ALWAYS!

Reply to  Tim Gorman
August 11, 2026 3:08 pm

“The average is *NOT* a function, it is a statistical descriptor of a distribution.”

Yes we know you don’t understand maths or what a function is, yet you still insist you can propagate the uncertainty for an average. You just want to make up any old nonsense as long as you get a large enough uncertainty.

And yet you still believe you are trying to educate me.

“A statistical descriptor equation is not a FUNCTION, it is a mathematical operator, little different than >. ”

What in your mind does FUNCTION mean?

“The uncertainty associated with that mapping is SAMPLING uncertainty, not measurement uncertainty.”

Make your mind up. Do you want to treat the data as a sample and use the SEM or do you want to propagate measurement uncertainty? Or are you still just spouting anything in the hope you can make the uncertainty impossibly large?

“Measurement uncertainty is the SD of the distribution.”

Nope, that’s just another of your hallucinations.

“You are basically saying that SD/sqrt(n) is always greater than SD. ”

Guess again.

“You add variances just like you add uncertainties. ”

Your record is stuck. I suppose it’s pointless to explain again that what you do with variances depends on what you are doing to the random variables. You obviously learnt at school that adding independent random variables adds the variances, but dropped out before they explained what happens when you scale random variables.

“So SD ≥ SD/sqrt(n) ALWAYS!”

Keep thinking, maybe you will understand that this means the larger n is the smaller the uncertainty. But that would require you understanding why the measurement uncertainty of an average is not the standard deviation of all the measurements.

Reply to  Bellman
August 11, 2026 5:06 pm

Keep thinking, maybe you will understand that this means the larger n is the smaller the uncertainty. But that would require you understanding why the measurement uncertainty of an average is not the standard deviation of all the measurements.

Total trendology nonsense, you still refuse to grasp the basics of metrology, obviously because of your vested interest in claiming impossibly tiny “error bar” numbers.

Reply to  Bellman
August 11, 2026 6:31 am

One day you are going to hVe to actually define what you mean by measurement,

Projection alert.

Reply to  Bellman
August 11, 2026 6:36 am

Again, just silly. If the measurements are uncertain, that uncertainty will be reflected in the uncertainty of the average.”

Here we go again. You are conflating measurement uncertainty with sampling uncertainty!

You just absolutely refuse to specifically identify the differences. You just continue to use the term “uncertainty of the average” without actually defining what you are speaking of.

If the stated values of the individual data make up the entire population then the population mean can be calculated very precisely, there is no “uncertainty of the average”. Yet those stated values CAN have measurement uncertainty that gets propagated into a sum for the distribution.

If the entire population consists of:

(10+/-2, 11+/-2, 12+/-2)

the average can be very precisely found, it is 11. NO UNCERTAINTY OF THE AVERAGE. But the measurement uncertainty is at most +/- 6 (direct addition) or if partial cancellation is assumed +/- 4.

The measurement uncertainty value describes the interval of reasonable values that can be attributed to the value of the measurand. It does *NOT* tell you how accurately you have located the mean of the data distribution. Nor does how precisely you have located the average determine the measurement uncertainty of the measurement.

If those three values are actually the means of three samples then the SEM, how precisely you have located the mean of the population, is approximately 0.5. This is the SAMPLING uncertainty and not the measurement uncertainty. And it is really only applicable to judging the best estimate if the distribution is not skewed in any way, i.e. that the mean = mode.

After LITERALLY years of being educated on metrology basics you *STILL* have absolutely no understanding of measurement uncertainty. You stubbornly cling to the same misconceptions you started with.

Reply to  Tim Gorman
August 11, 2026 8:20 am

“Here we go again. You are conflating measurement uncertainty with sampling uncertainty!”

Nope. You just refuse to read what I write.

“You just continue to use the term “uncertainty of the average” without actually defining what you are speaking of. ”

By uncertainty of the average I mean how certain you are that your calculated average reflects the actual average. There may be many components to this uncertainty, including how you sampled the data, what measurement uncertainty there is in the individual measurements, any systematic errors in either, how you corrected for those errors, etc.

“If the stated values of the individual data make up the entire population then the population mean can be calculated very precisely…”

That’s why we keep coming back to equation 10. It’s the uncertainty (of measurement) when the measurement is an exact average of all your individual measurements. If that’s all you want then you can say that is the uncertainty of your average. It’s just that thus is rarely what you want. The population and the sample are usually different and the uncertainty is about how well your sample average reflects the population average. This uncertainty will include the uncertainty if the individual measurements, but usually that will not be so important.

“there is no “uncertainty of the average””

Unless there is any uncertainty in your individual measurements.

Reply to  Bellman
August 11, 2026 8:46 am

“If the entire population consists of:

(10+/-2, 11+/-2, 12+/-2) ”

And we are back to the usual nonsense. First claiming there is no uncertai Ty, then claiming there is measurement uncertainty, but then, as always completely failing to understand how to propagate those uncertainties onto the average.

It’s just not worth arguing anymore. You believe in nonsense and no amount of logic will convince you you are wrong.

“If those three values are actually the means of three samples then the SEM, how precisely you have located the mean of the population, is approximately 0.5. ”

Ease try to learn what these words mean. The three values are not taken from the SEM. The SEM describes the sampling distribution.

And in this case you are saying there is no sampling distribution as these values are the entire population.

I think you are trying to create an absurd example to change my claim that usually the uncertainty from sampling is bigger than from measurement. But in so doing you are not describing a meaningful exercise. The standard deviation of your three values is smaller than your measurement uncertainty. Why would you use an instrument that has an uncertainty of ±2 when the things you are measuring only varied by 1 or 2 units? Moreover if you were doing this you would expect your measurements to vary by more than the population standard deviation.

If you want to have a fantasy that you are educating me, you need to try and behave like a teacher and day things that indicate you understand the subject. Just yelling and saying you’re the expert and cannot be questioned, is not good teaching.

Reply to  Bellman
August 11, 2026 9:54 am

By uncertainty of the average I mean how certain you are that your calculated average reflects the actual average.”

That is SAMPLING UNCERTAINTY, not measurement uncertainty. It is associated with the best estimate and not with the measurement uncertainty.

“what measurement uncertainty there is in the individual measurements”

Measurement uncertainty is *NOT* used to determine SAMPLING uncertainty.

” If that’s all you want then you can say that is the uncertainty of your average.”

NO, you can’t say that. It is the measurement uncertainty of the measurements, not of the average. The average is *NOT* a measurement.

Reply to  Tim Gorman
August 12, 2026 6:11 am

“That is SAMPLING UNCERTAINTY, not measurement uncertainty.”

Whatever. If you want to know how uncertain your estimate of the average is you need to know the reasonable range that the actual average could have. That’s what I mean by the uncertainty if an average. This uncertainty calculation can take on as many factors as seem reasonable, including measurement uncertainty.

What you claim is measurement uncertainty has no relation to an actual uncertainty. You are now claiming the standard deviation of all values that go into the average is the measurement uncertainty. This is just absurd and is just a other example of you misunderstanding everything you read in order to claim the largest uncertainty possible.

Taking your current definition, how is that on any way going to tell you about how uncertain your actual average is. What practical deduction could you make by being told the measurement uncertainty of a UAH anomaly of 0.5°C has a measurement uncertainty of say 10°C?

Reply to  Bellman
August 12, 2026 6:30 am

This uncertainty calculation can take on as many factors as seem reasonable, including measurement uncertainty.

The why do you ignore real measurement uncertainty?

Reply to  Bellman
August 12, 2026 3:17 pm

What you claim is measurement uncertainty has no relation to an actual uncertainty.

You have no idea what you are talking about. If you measure the same thing multiple times under repeatable conditions, along with a decent device there will be small random variations of each observation. If the observations have a symmetric distribution, there will be an interval surrounding the mean where a single standard deviation should cover 68% of the observations.

Actual uncertainty arises when you are measuring different but similar things or other changes in the measurement conditions. This is known as reproducible conditions. When that occurs, you must determine the variance of the observations and add it to the component determined by analyzing the same thing.

Reply to  Jim Gorman
August 12, 2026 6:04 pm

If you measure the same thing multiple times under repeatable conditions

But as you keep pointing out we are not measuring the same thing under repeatable conditions.

You can’t have it both ways. If you want to treat the global average as a measurand and each reading from a station as a single measurement of the global mean, you can say the standard deviation is the uncertainty of any one measurement. That is treat it as a type A uncertainty, as in the GUM 4.2.2. But then you also have to accept that you can use the “experimental standard deviation of the mean”, as in GUM 4.2.3 as the uncertainty of the mean of all your individual measurements. This is what TN1900 does in example 2.

But you keep insisting that you cannot do that for the global mean because they are not observations obtained under repeatability, or measuring the same thing. In which case it makes no sense to use 4.2.2 which is assuming you are measuring the same thing.

Of course, if you don’t think the mean is a measurand, or the individual temperature are measuring the global mean under conditions of repeatability – you can just treat it as a statistic, and calculate the SEM in the usual manor as SD/√N, which is exactly the same equation as 4.2.3, just applied to a random sample rather than observations made under repeatability conditions.

Reply to  Bellman
August 12, 2026 6:18 pm

you can just treat it as a statistic, and calculate the SEM in the usual manor as SD/√N

…and get the impossibly tiny numbers you need! Problem solved!

Reply to  Bellman
August 13, 2026 7:26 am

If you want to treat the global average as a measurand and each reading from a station as a single measurement of the global mean, you can say the standard deviation is the uncertainty of any one measurement”

The global mean is *NOT* a single measurand. It is not even a measurand. There is no place you can go to actually measure it.

from a station as a single measurement of the global mean”

You STILL haven’t internalized TN1900, Ex 2. Tmax is a measurable measurand. In TN1900 the measurements of Tmax are taken in the same micro-environment using the same instrument. The rest of the assumptions in the example are meant to make the measurements into “repeatable” measurements, no different than taking 10 measurements of a water bath to estimate its temperature.

Nor is there a functional relationship of measurable components that can be used to calculate the global mean from measurable components. A STATISTICAL DESCRIPTOR operation is *NOT* a functional relationship. It is a “functional operation” providing a mapping of a distribution to a set of descriptive values. The statistical descriptors tell you about the distribution, and not about the physical meaning of measurements.

Even worse, the temperature data components are not themselves averages. They are mid-point values that are determined by the range of the data and not by the distribution of the data.

But you keep insisting that you cannot do that for the global mean because they are not observations obtained under repeatability, or measuring the same thing.”

The are *NOT* the same thing. Different micro-climates, different measuring devices, different variances, etc. THEY AREN’T EVEN MEASUREMENTS – they are mid-range values of a diurnal profile!

you can just treat it as a statistic, and calculate the SEM in the usual manor as SD/√N, which is exactly the same equation as 4.2.3, just applied to a random sample rather than observations made under repeatability conditions.”

CAN YOU STOP CHERRY PICKING? Just *STOP* it!

The SEM is *NOT* a measurement uncertainty. It is a SAMPLING UNCERTAINTY!

4.2.3 defines s^2(q_bar).

Read 4.2.2
—————–
The individual observations qk differ in value because of random variations in the influence quantities, or random effects (see 3.2.2). The experimental variance of the observations, which estimates the variance σ 2 of the probability distribution of q, is given by

s^(q_k) = [1/(n-1)] Σ (q_j – q_bar)^2

This estimate of variance and its positive square root s(q_k), termed the experimental standard deviation (B.2.17), characterize the variability of the observed values q_k , or more specifically, their dispersion about their mean q_bar.

———————

———————–
2.2.3

uncertainty (of measurement)
parameter, associated with the result of a measurement, that characterizes the dispersion of the values that could reasonably be attributed to the measurand
—————————–

Reply to  Tim Gorman
August 13, 2026 9:26 am

The global mean is *NOT* a single measurand. It is not even a measurand.

Round and round we go. If it isn’t a measurand it can’t have a measurement uncertainty, by definition. So can we just stop this endless nonsense? No, I thought not.

4.2.3 defines s^2(q_bar)

What do you think that means? It’s the variance of the average. Take the square root and you have the standard deviation of the mean. Which is taken to be the measurement uncertainty of the mean.

specifically, their dispersion about their mean q_bar.

Yes, that’s the standard deviation. A measure of the dispersion of values about the mean.

parameter, associated with the result of a measurement, that characterizes the dispersion of the values that could reasonably be attributed to the measurand

And this is where you keep falling over your own feet. What do you think the measurand is in this case and what is the measurement? You deny the mean is a measurand, so what is the measurand?

Assuming you allow that the mean is the measurand, then your measurement is your estimate of the mean – that is your sample mean.

Now 4.2.2 describes the standard deviation of all measurements about the measurand (the mean). But that’s only telling you the measurement uncertainty of any one measurement. But you are not using one individual measurement as your measurement, but averaging all of them together, which is exactly what 4.2.3 describes.

Reply to  Bellman
August 13, 2026 3:26 pm

 as in the GUM 4.2.2. But then you also have to accept that you can use the “experimental standard deviation of the mean”, as in GUM 4.2.3 as the uncertainty of the mean of all your individual measurements. This is what TN1900 does in example 2.

This is where I disagree with TN 1900. It says:

In these circumstances, the {ti} will be like a sample from a Gaussian distribution with mean r and standard deviation ( (both unknown).

Each {ti} is A sample from a Gaussian distribution. And each {ti} has how many “n”? They are single temperatures so “n = 1” for the sample size. To find the SDOM you divide by “n”, NOT the number of samples. Consequently, the standard uncertainty of the mean should be “σ/1″.

Go back and study the GUM more carefully. Here is what it says about the uncertainty of the mean.

4.2.3 Thus, for an input quantity Xᵢ determined from n independent repeated observations Xᵢ,ₖ, the standard uncertainty u(xᵢ) of its estimate xᵢ = Xᵢ is u(xᵢ) = s(Xᵢ), with s²(X) calculated according to Equation (5). For convenience, u² (x) = s²(X) and u(x) = s(X) are sometimes called a Type A variance and a Type A standard uncertainty, respectively.

See that Xᵢ,ₖ, it means “k” multiple observations of each input quantity. In the GUM “k = n”, that is the size of the sample for each input quantity. Consequently, NIST should have divided by the √1 and not the √22. Remember, NIST offered the assumption each input quantity {ti} was a sample, not me.

Reply to  Jim Gorman
August 13, 2026 4:09 pm

Consequently, NIST should have divided by the √1 and not the √22. Remember, NIST offered the assumption each input quantity {ti} was a sample, not me.

He (and they) will never acknowledge this, the SEM is the raison d’etre sits upon which their entire house of cards sits.

Reply to  Jim Gorman
August 13, 2026 5:53 pm

Each {ti} is A sample from a Gaussian distribution. And each {ti} has how many “n”?

Good grief – it’s not each {ti}. {ti} is a set of values, each daily maximum. It’s a single sample of size 22.

You’ve been making this mistake for years.

Reply to  Bellman
August 14, 2026 1:16 pm

Good grief – it’s not each {ti}. 

Just what do you think “Each {ti} is a sample” tells you?

Read this very carefully.

If Ei denotes the combined result of such effects, then t = τ +εi where ε denotes a random variable with mean 0, for i =1,…,m, where m = 22 denotes the number of days in which the thermometer was read. This so-called measurement error model (Freedman et al., 2007)may be specialized further by assuming that ε₁, …, εₘ are modeled independent random variables with the same Gaussian distribution with mean 0 and standard deviation σ. In these circumstances, the {t} will be like a sample from a Gaussian distribution with mean τ and standard deviation σ (both unknown).

t, = τ +ε,
t = τ +ε
.
.
.
t₂₂ = τ +ε₂₂

assuming that ε₁, …, εₘ are modeled independent random variables with the same Gaussian distribution with mean 0 and standard deviation σ.

ε, is a random variable with mean 0 and SD of σ. This makes t, a sample from a Gaussian distribution..
ε₂ is a random variable with mean 0 and SD of σ. This makes t a sample.
ε₂₂ is a random variable with mean 0 and SD of σ. This makes t₂₂ a sample.

So we have 22 samples, each with a mean of 0, and an SD of σ. Geez, perfect IID samples.

Lastly, we only have 22 observations. Exactly how many of those fit into each of the 22 samples? Maybe one per sample?

Reply to  Jim Gorman
August 14, 2026 3:30 pm

Read this very carefully.

Stop with this nonsense. Stop expecting me to read exactly the same cut and paste for the hundredth time just becasue you don’t understand it.

So we have 22 samples

That’s your misunderstanding – not mine. You need to reread it carefully, or better still try to understand how the maths works. Your the one claiming that NIST miscalculated it – maybe you need to consider the possibility that they know what they are doing and you are the one not understanding how sampling works.

There’s only so many times I can explain that a sample of size N is made up of adding N iid random variables and dividing by N. Each random variable is one element of the sample, and in this case each daily reading is considered to be value taken from one of the random variables.

Reply to  Bellman
August 15, 2026 3:28 pm

 Each random variable is one element of the sample, and in this case each daily reading is considered to be value taken from one of the random variables.

That is simply untrue. Read this again.

assuming that ε₁, …, εₘ are modeled independent random variables with the same Gaussian distribution with mean 0 and standard deviation σ.

τ is constant. The expected value. Therefore, each tᵢ has an independent value.

Then you have:
t, = τ +ε,
t = τ +ε

  • ε₁, …, εₘ are INDEPENDENT random variables.
  • ε₁, …, εₘ each have the same Gaussian distribution, i.e. the same mean.
  • ε₁, …, εₘ each have the same SD of σ.

The example also says this.

In these circumstances, the {t} will be like a sample from a Gaussian distribution with mean τ and standard deviation σ (both unknown).

Maybe you have a different definition of sample than I do but I am pretty sure I can follow the math here. If each error term is independent, then each {t} will also be independent, and not part of single sample.

Why don’t you post something from the example that shows NIST intended for the {t}’s to be a member of a larger random variable.

Please read GUM F.1.1.2. This treats how independent measurements of different samples should be treated.

if sampling is part of the measurement procedure because the measurand is the property of a material (as opposed to the property of a given specimen of the material), then the observations have not been independently repeated; an evaluation of a component of variance arising from possible differences among samples must be added to the observed variance of the repeated observations made on the single sample.

Variance arising from differences among samples is the operative phrase here.

Reply to  Jim Gorman
August 15, 2026 3:54 pm

ε₁, …, εₘ are INDEPENDENT random variables.
ε₁, …, εₘ each have the same Gaussian distribution, i.e. the same mean.
ε₁, …, εₘ each have the same SD of σ.”

Yes, because they are modeled as IID random variables. I means independent. ID means identically distributed.

Maybe you have a different definition of sample than I do but I am pretty sure I can follow the math here.

A sample is a subset of individuals selected from a larger population. The goal is to use the sample to make inferences or draw conclusions about the whole population.

https://www.geeksforgeeks.org/maths/sample-definition-types-formula-examples/

However, you have to understand that in this case the population is all possible daily values, and each daily value is one random individual drawn from that population.

If each error term is independent, then each {tᵢ} will also be independent, and not part of single sample.

I think you are getting confused by the symbols here. There is no “each” {tᵢ}. The curly braces indicate a set, the subscript i is an index. {tᵢ} means the set of all tᵢ for i from 1 to m.

Why don’t you post something from the example that shows NIST intended for the {tᵢ}’s to be a member of a larger random variable.

I keep telling you you don’t understand what a random variable is. As I said each tis modeled by a single random variable – 22 in all, each one independent and identically distributed.

Please read GUM F.1.1.2.

Why. It will say the same as it did every other time you insisted I read it – and will still be irrelevant to this example. What do you think

if sampling is part of the measurement procedure because the measurand is the property of a material (as opposed to the property of a given specimen of the material)

means?

The only way this is relevant to this example is if you want to take the average temperature for this specific month, and claim it’s a property of all months. That would be a silly thing to do, and you certainly couldn’t claim you have 22 independent observations for all months of May. But that is obviously not what NIST are doing here.

Reply to  Bellman
August 11, 2026 5:07 pm

Unless there is any uncertainty in your individual measurements.

Which you (and climatology) just ignore.

Reply to  Bellman
August 11, 2026 6:40 am

I’ll ask again, what do you mean by “physical meaning””

Just how dense can you be?

Once again, I’ll ask the same question that you absolutely refuse to answer:

If I put two rocks in your right hand, one at 60F and one at 70F, are you holding 130F in your right hand?

The answer will tell us if you understand what “physical meaning” means in any way, shape, or form.

Reply to  Tim Gorman
August 11, 2026 8:06 am

“Just how dense can you be? ”

Dense enough to keep asking you a question I know you won’t answer.

I’ll ask again, what do you mean by the term “physically meaningful”? How would you distinguish a “physically meaningful” quantity from one that isn’t “physically meaningful”?

“Once again, I’ll ask the same question that you absolutely refuse to answer:”

In other words, whenever you are asked a pertinent question you try to distract from it with an impertinent question. Classic deflection.

“If I put two rocks in your right hand, one at 60F and one at 70F, are you holding 130F in your right hand?”

The answer is exactly the same as the other 200 hundred times you’ve asked.

No.

The fact you think I keep refusing to answer it is why you should get your memory checked.

Reply to  Bellman
August 11, 2026 9:50 am

I gave you the answer. I can’t help that you refuse to understand it.

If you are *NOT* holding 130F in your hand then the average has no physical meaning since the sum of the temperatures has no physical meaning!

*YOU* want to attribute some physical meaning to the average of the two temperatures even though you can’t add them together into a larger system.

Willful ignorance – the worst kind of ignorance.

Reply to  Tim Gorman
August 11, 2026 11:40 am

“If you are *NOT* holding 130F in your hand then the average has no physical meaning since the sum of the temperatures has no physical meaning!”

I think this is why you keep talking about “math world” as if it’s an alien concept. You simply don’t understand abstraction so keep deriding it as numbers is numbers. If you could understand it, it you might understand that not everything that is useful has to exist as a physical object, and not every step of an equation has to be physically meaningful.

You seem to be working on the assumption that an average is sharing out some quantity equally, and in some cases you could look at it in those terms. It is what the original word meant – sharing out loses from a shipwreck were shared out evenly.

But when it started being used in statistics it takes on a whole new meaning. Rather than sharing things out it is used to determine the central tendency, a mid point of all values. You can still calculate it in the same way, but you are not literally combing all the values into a big pool and then sharing them out evenly. That’s just a way of finding the mid point. If you take the heights of 20 people you are not literally placing them end to end, and then cutting bits of the taller people off and sticking them to the smaller ones until everyone is the same height. The point is to find the central tendency if all the heights, adding the heights and dividing by the number of people is just a convenient way of doing it.

The same with your two stones. I can add the two temperatures to get a sum of temperatures, but doesn’t have any meaning. Even if I used Kelvin. But as soon as I divide by 2, I get the value exactly mid way between the two. And that’s the value that may be useful. The fact I had to sum their temperatures is irrelevant to the end result. I could just as easily have obtained it by taking the difference in temperatures, dividing by 2 and adding to the smaller value. Or by any other method that will tell me what the mid point between the two is.

The fact that you say there is no problem with averaging multiple temperature measurements of the same object should be a clue as to why it doesn’t matter if the sum is not physically meaningful. The fact that you also don’t have a problem averaging temperatures across the UK, should make you question why you think it’s impossible to average the temperature of two stones.

Reply to  Bellman
August 11, 2026 11:57 am

“*YOU* want to attribute some physical meaning to the average of the two temperatures even though you can’t add them together into a larger system.”

No I do not I keep telling you that the average is a statistic not a physical object, and that this applies to extensive or intensive properties. The average mass of two stones is not the mass of a stone. The average temperature of two stones is not the temperature of a stone

Reply to  Bellman
August 11, 2026 6:41 am

Do you think the average of people’s heights exists, and has physical meaning?”

Height is *NOT* an intensive value. I can lay all those people down head-to-foot and get a sum of all the heights. That sum has a physical meaning.

YOU *STILL* REFUSE TO ADMIT THAT INTENSIVE VALUES CAN’T BE SUMMED INTO A LARGER SYSTEM.

Reply to  Tim Gorman
August 11, 2026 8:52 am

“Height is *NOT* an intensive value.”

Not answering the question. Do you think it is physically meaningful?

“That sum has a physical meaning.”

Not answering the question. I asked if the average had a physical meaning.

“YOU *STILL* REFUSE TO ADMIT THAT INTENSIVE VALUES CAN’T BE SUMMED INTO A LARGER SYSTEM.”

Do I have to write in all caps before you will accept that I keep answering?

THE SUM OF INTENSIVE VALUES IS NOT MEANINGFUL. THAT IS PRETTY MUCH THE DEFINITION OF AN INTENSIVE PROPERTY. IT DOES NOT GROW WITH THE SYSTEM SIZE.

WAS THAT LOUD ENOUGH FOR YOU?

Reply to  Bellman
August 11, 2026 6:47 am

You are saying the average if surface temperatures is not the average if surface temperatures? In that case what is it?”

IT IS NOTHING! You can’t add temperatures to create a larger system.

For the umpteenth time:

—————-
If I put two rocks in your right hand, one at 60F and one at 70F, are you holding 130F in your right hand?
—————-

If I give you this relationship, 60 < 70, does that tell you anything physically meaningful?

It’s a mathematical operation that is true if you assume “numbers is just numbers”. Just like adding temperatures of different things.

What’s the meaning of 60 < 70 in the real, physical world?

Reply to  Bellman
August 11, 2026 6:54 am

I’ll say again that statistics world is the same as science world.”

No! In science world you can’t average intensive properties.

Your insistence that you *can* average intensive properties means you are in Statistical World, not in science world.

 most of science for the past century or so understands the importance of statistical analysis.”

Non sequitur. The assertion has no relationship to the issue at hand. Science understands that the statistical descriptor known as the mean only has meaning if the data being analyzed can be added to create a larger system.

Physical science simply doesn’t subscribe to the meme of “numbers is just numbers”.

Reply to  Tim Gorman
August 11, 2026 9:08 am

“No! In science world you can’t average intensive properties. ”

Yet many scientists seem to manage it. Even Dr Spencer is capable of producing a monthly average temperature.

“Science understands that the statistical descriptor known as the mean only has meaning if the data being analyzed can be added to create a larger system.”

Does it? Where did “science” say this?

You really can’t get your head around the fact that an average is not a physical thing. The average height of a person is not the height of a person. The average body temperature if a person is not the body temperature of a person. Both are statistical tools for asking questions about populations. Both can be used for that purpose. It really makes no difference if one is an average of an extensive property and the other is the average of an intensive property.

Reply to  Bellman
August 11, 2026 9:51 am

You really can’t get your head around the fact that an average is not a physical thing.

That is funny. We’ve been telling you that averages of intensive properties is meaningless. Here is something I just ran across at climate.gov. Pre-industrial global temperature is 13.7°C. Guess what the 20th century average global temperature is? 13.9°C.

What CAGW do you get when you compare those two averages?

Reply to  Jim Gorman
August 11, 2026 10:53 am

” We’ve been telling you that averages of intensive properties is meaningless”

And I keep telling you why that’s wrong. Unless you are prepared to accept the concept you might be wrong on this we will just keep going round in circles.

” Pre-industrial global temperature is 13.7°C. Guess what the 20th century average global temperature is? 13.9°C. ”

Could you provide a link?

Edit. OK, found it

https://www.climate.gov/news-features/understanding-climate/climate-change-global-temperature

You slightly missed the headline that 2024 was 1.2°C above the 20th century average.

The point about the 20th century is that is that most of the warming only really started on the last quarter of it. The worry about global warming isn’t about how warm the 20th century was. It’s about how warm it is now and how much warmer it will get.

Reply to  Bellman
August 12, 2026 8:32 am

Yet many scientists seem to manage it. Even Dr Spencer is capable of producing a monthly average temperature.”

Which doesn’t make it into anything physically meaningful.

“You really can’t get your head around the fact that an average is not a physical thing.”

If it isn’t physical then how can it describe the physical world?

 Both can be used for that purpose. It really makes no difference if one is an average of an extensive property and the other is the average of an intensive property.”

The mathematical operation of calculating the mean of a distribution in no way depends on physical characteristics of the data in the distribution nor does the mean of a distribution affect the data in the distribution in any way.

All the mathemenatical operation requires is a set of numbers.

 Both can be used for that purpose.”

Actually the average temperature CAN NOT be used to characterize any indivduals temperature as normal or not.

Just like climate science, you totally ignore the variance of the data. Ask your primary care person at your next visit what they consider a normal temperature to be. Mine says 97F to 100F. AN INTERVAL of temperatures, not an AVERAGE. You are considered to have a low grade fever if your current temperature is 2F above YOUR normal temperature, not above the human average. E.g. if YOUR normal temperature is 97.5F then you are considered to have a fever if your temperature is above 99.5F, not the “clinical” fever point of 100.4F.

Since climate science NEVER provides the variance of the “global” temperature data, it is impossible to judge when the “globe” has a fever!



Reply to  Tim Gorman
August 12, 2026 8:50 am

“Which doesn’t make it into anything physically meaningful”

Are you ever going to explain what you mean by “physically meaningful”?

Regardless, your claim was that in “science world” it wasn’t possible. Not whether the result is physically meaningful.

“If it isn’t physical then how can it describe the physical world?”

Numbers are not physical things but they can describe the physical world. An average family size is not a physical thing, but it describes something about families who live in the physical world.

“The mathematical operation of calculating the mean of a distribution in no way depends on physical characteristics of the data in the distribution …”

I worry that this makes sense to you.

“… nor does the mean of a distribution affect the data in the distribution in any way. ”

Of course it doesn’t. You are mixing up cause and effect. It’s the distribution that affects the mean, not t’other way round”

“Actually the average temperature CAN NOT be used to characterize any indivduals temperature as normal or not.”

Why don’t you ever address what I’m saying. I said nothing about “normal”. I said it can tell you about populations. E.g. you test a drug by giving it to volunteers and compare their average temperature against a control group. If there is a significant difference in body temperature you can conclude that the drug had an effect.

“Just like climate science, you totally ignore the variance of the data.”

The variance in the data is essential for determining significance, and is also essential for determine the range of effects. E.g for determining maximum safe dose. Nobody should be ignoring the variance.

Reply to  Bellman
August 13, 2026 5:44 am

“Are you ever going to explain what you mean by “physically meaningful”?”

Physically meaningful: relating to the physical, real world. Describing the REAL world and not just a “blackboard” world.

“An average family size is not a physical thing, but it describes something about families who live in the physical world.”

Once again, you are trying to conflate EXTENSIVE measurements with INTENSIVE measurements. And you don’t even realize that you are doing so.

Because you have no real understanding of physical science.

I worry that this makes sense to you.”

Do *YOU* need to know the units that go with the measurements in the data set in order to calculate the mean of the distribution?

Or do you just need to know the stated value NUMBERS?

You are mixing up cause and effect. It’s the distribution that affects the mean, not t’other way round””

And yet you think you have to know the units that go with the individual data components in order to calculate the average?

The dimensions of the individual data is *NOT* needed to calculate the statistical descriptors of the distribution. Nor do the statistical descriptors *change* the distribution or the individual components in any way, shape, or form.

And yet you think that my statement doesn’t make any sense to you?

————————-
“The mathematical operation of calculating the mean of a distribution in no way depends on physical characteristics of the data in the distribution …”
————————

” I said it can tell you about populations.”

bellman: ” The average body temperature if a person is not the body temperature of a person. Both are statistical tools for asking questions about populations.”

YOU are the one that brought up body temperature. And the average body temperature of the human population is *NOT* used for asking questions about the individual – again, the average does not apply to the individual data.

In the REAL world, it is the average temperature of the individual that matters, not the average temperature of the population.

You just can’t get out of Statistical World, can you?

“E.g. you test a drug by giving it to volunteers and compare their average temperature against a control group.”

The blind leading the blind again. My son will tell you that each individual has to be weighted based on their individual characteristic temperature.

You simply cannot ASSume that the average temperature of the control group is descriptive of anything. Typically the control group will be given a placebo and their weighted average temperatures will be compared to the weighted average temperature of the test group.

You are the PERFECT example of what I pointed out to my son when he started his studies. Math majors can’t judge the reasonableness of the data and the biologists can’t judge the reasonableness of the statistical analysis. THE BLIND LEADING THE BLIND.

You really don’t know anything about physical science and yet here you are again, trying to lecture on it!

Reply to  Tim Gorman
August 13, 2026 6:27 am

“Physically meaningful: relating to the physical, real world. Describing the REAL world and not just a “blackboard” world. ”

Then I would say an average can be physically meaningful. E.g. observing differences in average heigh betwee people exposed to a particular substance in childhood and people not exposed, can tell us something about how that substance affected growth in the real world. Changes in average temperature of products coming of a production line can relate to real world changes in the production process.

In general any average based on real world measurements is going to relate to the physical real world, and will have meaning related to that real physical world.

“Once again, you are trying to conflate EXTENSIVE measurements with INTENSIVE measurements. And you don’t even realize that you are doing so.”

Stop lying. I said nothing about extensive or intensive. I used an example of an average not being a physical thing. It makes no difference if it’s in or ex, averages are not normally physical things. Average family size is not the size of a family, average height is not the height if a person, average temperature is not normally a temperature.

Reply to  Bellman
August 13, 2026 7:03 am

And yet you think that my statement doesn’t make any sense to you?

Yes, because what you said was

The mathematical operation of calculating the mean of a distribution in no way depends on physical characteristics of the data in the distribution nor does the mean of a distribution affect the data in the distribution in any way.

That seems to imply you think the mean does not depend on the values of the physical characteristics. But I take it now that what you mean is that it doesn’t depend on the type of the characteristics.

You are still wrong though. Of course the calculations just uses values, but you still need to know if the values make sense. You cannot add different units, you need to interpret the meaning of your values, you have to know when your result won’t make sense in the real world.

This is just another example of your misunderstanding of what maths actually is and how it’s used. It deals with abstractions, but when dealing with applied maths those abstractions will be put back into concrete realities.

My son will tell you that each individual has to be weighted based on their individual characteristic temperature.

What has that got to do with anything? I gave a hypothetical of how an average temperature can be useful. I said nothing about the specifics of how you conduct the experiment of calculate the mean.

You simply cannot ASSume that the average temperature of the control group is descriptive of anything.

Stop shouting. It’s especially weird when you do it for half the word, it just makes you look like you’ve got a donkey fetish.

And yes, you assume that the control groups average is descriptive of something. That’s the foundation of medical testing.

Typically the control group will be given a placebo and their weighted average temperatures will be compared to the weighted average temperature of the test group.

Yes. That’s what a control group means. I’m not sure what weighting would be relevant in this hypothetical exercise. More likely I would expect you look at temperature change.

Math majors can’t judge the reasonableness of the data and the biologists can’t judge the reasonableness of the statistical analysis. THE BLIND LEADING THE BLIND.

Remind me about not being allowed to lecture you in your claimed area of expertise. Oh you did –

“You really don’t know anything about physical science and yet here you are again, trying to lecture on it!”

You can lecture mathematicians, biologists, climate scientists, etc, on how they don’t understand their subjects. But nobody is allowed to point out any of your mistakes.

Reply to  Bellman
August 14, 2026 6:15 am

That seems to imply you think the mean does not depend on the values of the physical characteristics”

Your lack of reading skills is showing again.

 nor does the mean of a distribution affect the data in the distribution in any way.”

There is no mention of “values” in there. But if the mean doesn’t affect the data in any way then how can you read phrase to say it changes the values of the data?

But I take it now that what you mean is that it doesn’t depend on the type of the characteristics.”

No, I mean that the average doesn’t affect the data IN ANY WAY.

but you still need to know if the values make sense.”

What in Pete’s name do you think I was trying to tell you with my post about my son and knowing statistics as a biologist?

The issue is that if a statistician is given the values then the functional operation known as finding the mean can be done. And you keep trying to say that the value thus determined has meaning!

 I gave a hypothetical of how an average temperature can be useful.”

Judas H. Priest! You just said you have to know if the values make sense! And your example just ignored this!

The average of a control group cannot be compared to the average of a test group unless you can confirm that both populations are similar! It’s part of the problem in biological science with duplicating experiment results. Creating mice populations with the same genetic makeup is an entire industry!

You have just provided a PERFECT example of my BLIND LEADING THE BLIND example!

Reply to  Tim Gorman
August 14, 2026 9:08 am

“Your lack of reading skills is showing again. ”

Does it ever occur to you that it might be your writing skills that are the problem?

Reply to  Tim Gorman
August 14, 2026 11:16 am

“No, I mean that the average doesn’t affect the data IN ANY WAY.”

What you said was,

The mathematical operation of calculating the mean of a distribution in no way depends on physical characteristics of the data in the distribution

“The issue is that if a statistician is given the values then the functional operation known as finding the mean can be done. And you keep trying to say that the value thus determined has meaning!”

Yes, I think the value has meaning. What would be the point if it didn’t?

“Judas H. Priest! You just said you have to know if the values make sense! And your example just ignored this!”

The values make sense. I am not ignoring it.

“The average of a control group cannot be compared to the average of a test group unless you can confirm that both populations are similar! ‘

Of course they are similar. What’s the point of you trying to pick fault in a hypothetical example?

Reply to  Bellman
August 14, 2026 6:22 am

 I’m not sure what weighting would be relevant in this hypothetical exercise. More likely I would expect you look at temperature change.”

You simply have no grasp of physical science at all.

You can lecture mathematicians, biologists, climate scientists, etc, on how they don’t understand their subjects. “

Because I do know something about these subjects! You do not know physical science at all nor do you know anything about metrology. All you ever do is make cherry picked assertions or assertions that are obviously garbage – e.g. averaging can increase the resolution of measurments.

Why do you think engineers are required to have at least 6 hours and usually 9 hours in statistical analysis alone? Statisticians aren’t required to be trained in physical science at all – you are a perfect example!

The final argument on this is that you and bdgwx are *NOT* mathematicians, biologists, or climate scientists at all. It’s obvious from the garbage assertions you all make on here like “you can average intensive properties and get something physically meaningful”.

Reply to  Bellman
August 14, 2026 5:43 am

Then I would say an average can be physically meaningful.”

ROFL!!! Not if it is an average of intensive properties of different things.

“E.g. observing differences in average heigh”

For at least the third time, length (be it vertical or horizontal) IS NOT AN INTENSIVE VALUE.

In general any average based on real world measurements is going to relate to the physical real world”

Not if the average is of intensive properties.

Stop lying. I said nothing about extensive or intensive”

No lying. You keep trying to say that since you can average length (an extensive property) you can average temperature (an intensive property).

That’s trying to conflate extensiveproperites with intensive properties.

averages are not normally physical things”

FINALLY! You are admitting that the average of measurements is not itself a measurement. That does not keep the average from being used to compare the EXTENSIVE properties of different populations. It DOES keep the average from being used to compare the INTENSIVE properties of different populations!

Reply to  Tim Gorman
August 14, 2026 8:55 am

“ROFL!!! ”

A persuasive argument.

bdgwx
Reply to  Andy May
August 10, 2026 5:50 pm

Temperature of a system in equilibrium is an intensive property. Outside equilibrium (eg. the whole world) it loses it intensivity.

Perhaps you should think on that statement for awhile.

Thermodynamically, the Earth has no temperature, average or otherwise.

And yet people are able to measure it; at least certain parts of it like the TLT layer. sea surface, 2m surface, etc.

Reply to  bdgwx
August 11, 2026 9:35 am

And yet people are able to measure it; at least certain parts of it like the TLT layer. sea surface, 2m surface, etc.

It is a heterogenous body. It has conduction and convection and latent heat. It has topography.

Yes, you can measure temperature at small pieces of it but you cannot measure the “whole thing” as one simple property.

I liken it to a cube of steel. You can test the hardness at various points and average them to get a number. Does that average standalone as to the value of the hardness of the entire cube when you know full well that it varies throughout (heterogenous)? Or, do you have to supply a quantitative value of quality of the average, i.e., how much does it vary? That is, the measurement uncertainty.

Can you use that global temperature in SB to calculate how much radiation should be leaving? How do you include the hidden energy contained by latent heat?



Reply to  bdgwx
August 13, 2026 5:46 am

Perhaps you should think on that statement for awhile.”

Perhaps *YOU* need to think on it some more. If either of the systems are open where mass can be exchanged you can have equal temperatures in both systems but *NOT* thermal equilibrium.

You seem to be stuck in radiative equilibrium between two closed systems.

Reply to  Tim Gorman
August 10, 2026 12:07 pm

“No, YOU were talking about temperature when you brought up the 0th law of thermodynamics.”

I think your memory is failing. I was responding to your comment

Temperature can’t even be compared to other temperatures unless they are in similar environments. Comparing the temperature in Phoenix with the temperature in Miami is meaningless since the environments (humidity, pressure, geography, terrain, etc) are vastly different.

If that’s not talking about temperature what are you talking about?

“I am talking about INTENSIVE PROPERTIES. Temperature is an intensive property and I used it as an example. I could have used density instead. Nothing I said would be any different.”

So now you are saying you can’t compare densities?

Reply to  Bellman
August 10, 2026 8:31 am

Two things with the same temperature are the same temperature.

That is not the 0th law. The zeroth law is a transitive statement involving three bodies.

Symbolically it is:

If A = B and B = C, then A = C.

From: Zeroth Law – Thermal Equilibrium | Glenn Research Center | NASA

The zeroth law of thermodynamics is an observation. When two objects are separately in thermodynamic equilibrium with a third object, they are in equilibrium with each other. 

Mathematically, this is known as the Transitive Property of Equality.

Reply to  Jim Gorman
August 10, 2026 8:44 am

“The zeroth law is a transitive statement involving three bodies.”

And that’s the basis of saying that temperature is a meaningful value. It means temperature is an equivalence relationship. The logic is that you can partition the set of all things by equal temperatures. If one thing has a temperature of 20°C, then it has an equivalent temperature to all other 20°C things.

Of course all this is using the thermodynamic definition of temperature, but other definitions should still have an equivalence relationship. If you define temperature by average kinetic energy, then two objects with the same average kinetic energy will have the same average kinetic energy.

Reply to  Bellman
August 10, 2026 11:17 am

The logic is that you can partition the set of all things by equal temperatures.”

You can only do this for closed systems. Open systems which can exchange mass cannot be partioned by equal temperatures.

” If one thing has a temperature of 20°C, then it has an equivalent temperature to all other 20°C things.”

Again, ONLY if you are speaking of closed systems.

Where in the Earth’s biosphere do you think there are closed thermodynamic systems? Be specific!

Of course all this is using the thermodynamic definition of temperature,”

It is using the thermodynamic definition of CLOSED SYSTEMS that can’t exchange mass.

Give us an example in the earth’s biosphere that is a closed system. Be specfic.

Reply to  Tim Gorman
August 10, 2026 2:43 pm

Open systems which can exchange mass cannot be partioned by equal temperatures.

Time to throw away all your thermometers.

Reply to  Bellman
August 10, 2026 11:25 am

It means temperature is an equivalence relationship. 

The REQUIREMENT of the law is that thermodynamic equilibrium exists between A and B, that is, their temperatures are the same and no heat is being transferred. It is also a REQUIREMENT that thermodynamic equilibrium exists between B and C, that is their temperatures are the same and no heat is being transferred.

The logic is that you can partition the set of all things by equal temperatures. If one thing has a temperature of 20°C, then it has an equivalent temperature to all other 20°C things.

You are partially correct. If I have a block at 20C and cut it in half, then I have two blocks of 20C. However, this is also the definition of an intensive property. An extensive property such as mass of a 20g block, would provide you with two block of 10g.

If I then add the two blocks of 20C back together, I still have a larger block at 20C, not 40C. After adding the two masses, I WOULD have a block of 20g.

This is also where you go wrong. The zeroth law only applies to sensible heat. Bodies that are at the same temperature and for all intents and purpose identical.

The other laws deal with things like enthalpy which includes both sensible heat and latent heat.

Read this page, What is Thermodynamics? | Glenn Research Center | NASA.

You appear to not have a detailed education in thermodynamics. I had 9 hours at university about thermodynamics of heat in high pressure steam power plants, the design of heat sinks for electronics, and fluid flow in heat exchangers. Keeping track of temperature, latent heat, sensible heat, entropy, steam tables, material characteristics, etc. is intense. Delve into this at your peril.

Reply to  Jim Gorman
August 10, 2026 2:29 pm

You are partially correct. If I have a block at 20C and cut it in half, then I have two blocks of 20C

I think you are misunderstanding the word partition here. I’m talking about the set of all things in the universe – not splitting one thing.

Mass is an extensive property, but it still forms an equivalence relationship. Anything that has a mass of 1kg has an equivalent mass to anything else that has a mass of 1kg. That does not mean you can cut a 1kg weight in half and have two objects with a mass of 1kg each.

The zeroth law only applies to sensible heat.

Once again, heat and temperature are not the same thing.

Bodies that are at the same temperature and for all intents and purpose identical.

There is absolutely nothing in the zeroth law that requires bodies to be almost identical. How do you think thermometers work. They can measure the temperature of any object, not just things that are almost identical to thermometers.

Reply to  Tim Gorman
August 9, 2026 9:33 pm

The same temperatures in the two cities may be comfortable in Phoenix and oppressive in Miami.

Reply to  Clyde Spencer
August 10, 2026 3:28 am

bellman simply doesn’t understand that the 0th law only applies to closed systems and the exchange of heat energy between them, e.g. (Th – Tc).

That is *not* the Earth’s biosphere. The exchange of mass, as in convection and evaporation, is involved as well as conduction and radiation. That means that Earth is one big set of OPEN systems that can exchange energy AND mass. Thus the 0th law doesn’t apply.

Climate science makes the assumption that U = T where U is total energy and T is temperature. And bellman just gobbles it up. The fact is that U ≠ T when open systems and latent heat are involved.

It’s just one more garbage assumption in climate science.

Reply to  Clyde Spencer
August 10, 2026 4:14 am

“The same temperatures in the two cities may be comfortable in Phoenix and oppressive in Miami.”

Yes, but that doesn’t mean the temperatures are not the same.

It’s like arguing you can’t compare people’s heights because two people with the same height can have different weights. Temperature and humidity are different things.

Reply to  Bellman
August 10, 2026 4:41 am

Both temperature and humidity are RELATED to energy content. It is energy content that drive the system we know as Earth. You cannot just look at temperature while ignoring humidity.

Earth is a collection of OPEN systems, not a collection of closed systems. They can exchange mass as well as heat.

It is *NOT* like people of the same height having different weights. These are descriptions of not just closed systems but of ISOLATED systems. Person A doesn’t exchange either height or mass with Person B.

Reply to  Tim Gorman
August 10, 2026 5:24 am

“Both temperature and humidity are RELATED to energy content.”

So? You are just doing your usual distraction routine. The question was about comparing temperatures. Not energy content.

Reply to  Bellman
August 10, 2026 7:14 am

YOU were the one who dredged up the 0th Law, and then did your sidestep shuffle when you ignorance was exposed.

Reply to  karlomonte
August 10, 2026 10:15 am

As usual, bellman is here trying to lecture people on things he doesn’t understand.

Reply to  Tim Gorman
August 10, 2026 10:23 am

Yep. Pretty soon he’ll shift into his superiority mode.

Reply to  Bellman
August 10, 2026 10:14 am

You cannot compare temperatures unless they have the same environment – which includes humidity.

Did you actually READ what I posted?

Temperature defines heat transport between CLOSED SYSTEMS. That is the ONLY way it makes any sense.

The Earth’s biosphere is *NOT* a set of closed systems. They are a set of OPEN SYSTEMS, where mass as well as heat can be transported. Therefore it makes NO sense to compare temperatures between different locations in the Earth’s biosphere expecting them to define the heat transport involved.

You can’t even get the simplest things right when it comes to thermodynamics. The 0th Law only applies between closed systems. That is a “specialized” case and not the general case. The general case involves open systems.

Reply to  Bellman
August 10, 2026 9:52 am

It makes a difference when you include a variable that is associated. A 6’8″ at 125 lbs sitting on your head is vastly different than a 6’8″ at 365 libs sitting on your head.

Reply to  Jim Gorman
August 10, 2026 10:09 am

I’m sure having 365 “libs” sitting on my head would be a different experience.

But the point is that the heights are the same, not different just because they are associated with different weights.

Reply to  Bellman
August 10, 2026 9:26 pm

Do you realize that you are acknowledging the shortcomings of focusing exclusively on temperature? It is like someone needing a complex number (as in the sq rt of -1) to do a calculation and ignoring the imaginary part of the complex number because it is a little more awkward to work with.

Reply to  Clyde Spencer
August 11, 2026 6:59 am

“Do you realize that you are acknowledging the shortcomings of focusing exclusively on temperature?”

Focus on whatever you like. This article is very much about temperature, Tim’s comment was about comparing temperature, and as this web site is mostly concerned with global warming, temperature seems an appropriate thing to focus on.

If I’m interested in what tomorrow’s weather will be like, my primary focus is more on whether it will rain or be sunny, rather than the specific temperature.

From the perspective of climate change, then global warming is going to be the primary effect as that’s what is affected directly by greenhouse gases. Other changes are going to be feedbacks from warming. But you certainly don’t want to ignore those effects.

Reply to  Bellman
August 11, 2026 9:45 am

 temperature seems an appropriate thing to focus on.”

Somehow you never seem to want to focus on it. You just want to be able to say you can average the intensive property values of different things.

“If I’m interested in what tomorrow’s weather will be like, my primary focus is more on whether it will rain or be sunny, rather than the specific temperature.”

So you wear your wool coat year round? Regardless of the temperature?

Reply to  Bellman
August 11, 2026 12:11 pm

temperature seems an appropriate thing to focus on.

Except you are focused on AVERAGING temperatures. It is impossible to average intensive properties. That is a fact. A simple question to an AI will give you answer.

Q: How to average two intensive properties like temperature

CoPilot:

Correct approach

To find a meaningful average, you must weight the values by the amount of substance (or other relevant extensive variable) in each system. This is because the total energy or other extensive quantity depends on the amount of matter.

Grok:

Intensive properties (like temperature, pressure, or density) are not averaged by a simple arithmetic mean in the same way extensive properties are, because they are independent of system size and are not additive. The physically meaningful average depends on the context and is typically a weighted average based on the relevant conserved extensive quantity.

This means one should be converting to enthalpy, an extensive value, if comparisons or sums are done. Climate science has ignored enthalpy for going on 50 years just so they can treat water vapor as a feedback quantity rather than an important radiant factor in cooling the earth. If they used enthalpy, sensible and latent heat get properly combined but water vapor disappears into the combined quantity.

You just keep hammering that numbers are just numbers. You can certainty do that, but then you become unable to relate the average of those numbers to intensive physical attributes.

Reply to  Jim Gorman
August 11, 2026 3:30 pm

Except you are focused on AVERAGING temperatures.

I was talking about the difference between temperature and enthalpy.

“That is a fact. A simple question to an AI will give you answer. ”

It is impossible to average intensive properties. That is a fact. A simple question to an AI will give you answer.

Not AI again. I take it you managed to get an answer you liked, so AI is OK again.

To find a meaningful average, you must weight the values by the amount of substance (or other relevant extensive variable) in each system. This is because the total energy or other extensive quantity depends on the amount of matter.

So copilot says it’s not impossible to average intensive properties.

The physically meaningful average depends on the context and is typically a weighted average based on the relevant conserved extensive quantity.

So again, even Grok admits it’s possible to average intensive properties.

Of course you don;t quote what prompt you used, or how much arguing you had before you got those two answers. But again, I can’t stress this enough, AI’s do not know anything.

But if you insist, here’s what I just asked copilot

“Is it possible to take a sample mean of a set of temperature measurements.”

Response

Yes — you absolutely can take a sample mean of a set of temperature measurements. In fact, in metrology and climate science, the sample mean is one of the most fundamental statistics used to summarize repeated temperature observations.

The key nuance is what the mean represents and what assumptions you are making about the underlying physical quantity.

Reply to  Bellman
August 11, 2026 5:47 pm

So copilot says it’s not impossible to average intensive properties.

That is not what CoPilot said. It said intensive properties must be weighted by extensive properties in order to obtain a meaningful quantity.

I’ll repeat it:

To find a meaningful average, you must weight the values by the amount of substance (or other relevant extensive variable) in each system. This is because the total energy or other extensive quantity depends on the amount of matter.

That is not what Grok said. I’ll repeat it:

Intensive properties (like temperature, pressure, or density) are not averaged by a simple arithmetic mean. … The physically meaningful average depends on the context and is typically a weighted average based on the relevant conserved extensive quantity.

They both say that a simple average of temperature is not correct. You must weight by extensive property of the parcels.

But if you insist, here’s what I just asked copilot

You didn’t ask it the right question. You needed to also ask if the mean of a distribution of temperatures is meaningful when values in the distribution are from different parcels of air that are not identical.

Q: Is it possible to take a sample mean of a set of temperature measurements when the distribution contains data values from different parcels of air. Is the mean of the different values meaningful.

CoPilot

A simple arithmetic mean of temperatures from different air parcels is not physically meaningful unless the parcels represent a single mixed mass of air or you have a justified weighting scheme (mass, volume, or flow). Otherwise, the mean is just a statistical descriptor of the dataset — not a thermodynamic state variable.

Reply to  Jim Gorman
August 11, 2026 6:42 pm

” It said intensive properties must be weighted by extensive properties in order to obtain a meaningful quantity.”

Hence not impossible. You seem to be assuming that a weighted average is not an average. And as I’ve been trying to explain to you many times, if you weight an average temperature you are converting the intensive property temperature into an extensive property of temperature times some extensive property.

“You didn’t ask it the right question.”

What question to did you ask? My question was directed at exactly the point I’ve been arguing. That an average is usualy a statistic estimating a population.

” You needed to also ask if the mean of a distribution of temperatures is meaningful when values in the distribution are from different parcels of air that are not identical.”

You are framing the question to get the answer you want, and I expect training it to give answers that will please you.

Reply to  Bellman
August 11, 2026 6:55 pm

I fed copilot the identical question you asked and got this response:

Q: Is it possible to take a sample mean of a set of temperature measurements when the distribution contains data values from different parcels of air. Is the mean of the different values meaningful.

CoPilot

Yes, you can compute a sample mean of temperatures from different parcels of air, but whether that mean is meaningful depends entirely on what physical question you’re trying to answer. In many cases, the mean is mathematically valid but physically meaningless.

Now let me change your question to get a clearer answer, relevant to global temperatures.

Q: Is it possible to take a sample mean of a set of temperature measurements when the distribution contains data values from different parcels of air. Is the mean of the different values useful. For example is a global mean temperature a meaningful value for assessing global changes?

CoPilot:

Yes, you can compute a mean temperature from measurements taken in different air parcels, and the global mean surface temperature (GMST) is indeed a meaningful and scientifically essential quantity for assessing global climate change. But its meaning is statistical, not thermodynamic.

The key is understanding what the mean represents and why climate scientists use it.

Core Takeaway

The global mean temperature is not the temperature of a single physical air parcel. It is a planetary-scale statistical indicator that tracks long‑term changes in Earth’s climate system. It is meaningful because it is constructed carefully, consistently, and with physical justification — but it is not a thermodynamic average of mixed air masses.

Final synthesis

Final synthesis

  • Is it possible to take a mean of temperatures from different air parcels?
  • Yes, mathematically — but it is not a thermodynamic average.
  • Is the mean useful?
  • Yes, if interpreted as a statistical indicator rather than a physical temperature.
  • Is global mean temperature meaningful for assessing global change?
  • Absolutely. It is one of the most important and scientifically validated metrics for detecting and quantifying global warming.
Reply to  Bellman
August 11, 2026 8:45 pm

But not particularly useful because water vapor will cause two apparently otherwise similar air parcels to have different temperatures, if they have different absolute humidities, because the the water vapor carries latent heat and also affects the lapse rate. It’s like trying to bake a cake and leaving out the eggs.

Reply to  Clyde Spencer
August 12, 2026 5:53 am

“useful” depends on context. What questions you are trying to answer

Reply to  Bellman
August 12, 2026 7:32 am

“useful” depends on context. What questions you are trying to answer

Useful for the purpose of scientific thermodynamic evaluation of what is occurring with the ocean, land, and atmosphere, for both heat and temperature.

Any other purpose is simply playing with numbers for the sake of playing with numbers.

Reply to  Jim Gorman
August 12, 2026 8:16 am

“Useful for the purpose of scientific thermodynamic evaluation of what is occurring with the ocean, land, and atmosphere, for both heat and temperature. ”

Then I’d say it’s useful but you will need a lot of additional information. This is true whatever single property you look at. Could you do that knowing only the enthalpy?

Reply to  Bellman
August 13, 2026 4:05 am

Then I’d say it’s useful but you will need a lot of additional information. This is true whatever single property you look at. Could you do that knowing only the enthalpy?”

While still not perfect enthalpy is MUCH BETTER at defining climate than a mid-range temperature. Enthalpy will at least give a clue as to why the climate in Miami is different than in Phoenix!

Reply to  Tim Gorman
August 13, 2026 5:52 am

How? Enthalpy is just a number. If one volume if air has a higher enthalpy than another you can’t tell if that’s because it has a higher temperature or more latent heat. Or for that matter if it’s because you have a larger volume of air.

Reply to  Bellman
August 13, 2026 7:39 am

How? Enthalpy is just a number. If one volume if air has a higher enthalpy than another you can’t tell if that’s because it has a higher temperature or more latent heat. Or for that matter if it’s because you have a larger volume of air.”

You don’t even know what you just said!

Miami and Phoenix can have the EXACT SAME TEMPERATURE while the enthalpy is different.



Reply to  Tim Gorman
August 13, 2026 9:34 am

Miami and Phoenix can have the EXACT SAME TEMPERATURE while the enthalpy is different.

What has that got top do with my comment? You’re just mindlessly shouting your usual cliches rather than trying to engage in conversation.

My question was how using enthalpy alone would allow you to tell that the climate in two cities are different? To places could have identical specific enthalpys but different temperatures and climates.

Reply to  Bellman
August 13, 2026 10:59 am

My question was how using enthalpy alone would allow you to tell that the climate in two cities are different? 

Read this site. Humidity Control & AC Sizing Guide | Dehumidification BTUs | ACCalculator

Does this statement at the site answer your question?

Your AC has two jobs: remove heat (sensible load) and remove moisture (latent load). Most homeowners—and many contractors—only calculate sensible load, leading to cold, clammy homes that feel uncomfortable even when thermostats show 72°F. In humid climates, latent load can be 30-50% of total cooling needs.

“many contractors—only calculate sensible load, leading to cold, clammy homes”

This epitomizes you to a T. No idea about what you are arguing about!

Don’t be confused, “latent load” is actually latent thermodynamic heat.

Reply to  Jim Gorman
August 13, 2026 11:08 am

“Does this statement at the site answer your question?”

Not at all. If all you have is the enthalpy, you do not know how hot or humid it is.

Reply to  Tim Gorman
August 13, 2026 2:37 pm

It is hilarious!

Reply to  Bellman
August 13, 2026 8:59 am

How? Enthalpy is just a number. If one volume if air has a higher enthalpy than another you can’t tell if that’s because it has a higher temperature or more latent heat. Or for that matter if it’s because you have a larger volume of air.

You know how I know you’ve never taken advanced physics and thermodynamics? Because you have no idea about heat and energy and how they relate.

Of course a unique value of enthalpy cannot be broken into its constituent parts without knowing what the parts are.

Take the ideal gas law, PV = nRT. If I give you a temperature, can you decipher what the other parts are that give that temperature. You need to know two of the other three variables to do that. This is simple algebra.

The general form for enthalpy is:
h=(cₚ,ₐT)+(mₕ₂ₒ/mₐᵢᵣ)[(cₕ₂ₒT+ hᵥᵥₑ]
Where:

  • h = specific enthalpy of moist air (kJ/kg or Btu/lb)
  •  cₚ,ₐ = specific heat of dry air at constant pressure
  • x = humidity ratio (kg water/kg dry air)
  • hᵥᵥₑᵥ = specific enthalpy of water vapor 
  • T = temperature

There are 7 variables here. You need to know 6 of them to calculate the 7th. That has nothing to do with summing enthalpy values. Enthalpy is extensive, it can be summed and averaged. Will an average value be useful, only in terms of energy. The piece parts of T, mass, and specific heats still can’t be averaged individually.

I’m not going to show it here, but do the work and add parcels together and then solve for T. I’ll guarantee T is not a simple averaged value.

Reply to  Jim Gorman
August 13, 2026 10:03 am

“You know how I know you’ve never taken advanced physics and thermodynamics?”

Because I’ve told you? But if you and Tim are examples of people who have, maybe I don’t miss much.

Nothing you say addresses my point. You keep saying enthalpy tells you more about the climate than temperature? I disagree. The specific claim was that enthalpy would tell you why the climate of one city was different than another. But that only works if you also know the temperature of the two cities, or the humidity.

But I would also say that knowing both temperature and humidity will make it easier to assess the different climates than knowing the enthalpy and one of the others. There’s probably a reason why weather forcasts are given as temperature, pressure, humidty etc, but not enthalpy.

Reply to  Bellman
August 13, 2026 10:48 am

The specific claim was that enthalpy would tell you why the climate of one city was different than another. But that only works if you also know the temperature of the two cities, or the humidity.

Do you have a problem reading a formula? Do you know what “cₚ,ₐ”, the specific heat at a constant pressure of air is? How about the specific heat of water, hᵥᵥₑᵥ?

You can’t just know temperatures OR humidity. You need to know them both. You also need to know the proper specific heat values for dry air and humid air. BTW, they are not constant.

I will drag you back to the original subject. Thermodynamic heat, that is, energy in a parcel of air. If you think when you walk off a plane in either city, Las Vegas or Miami, that they will feel equally hot because they have the same temperature, think again. If you think you can install equally sized air conditioning units, think again. If you think rainfall is the same, think again.

Reply to  Jim Gorman
August 13, 2026 11:04 am

“Thermodynamic heat, that is, energy in a parcel of air.”

You keep claiming to understand thermodynamics, yet keep making the same mistake. Heat is not energy in a parcel of air. It’s energy being transferred from one thing to another.

” If you think when you walk off a plane in either city, Las Vegas or Miami, that they will feel equally hot because they have the same temperature, think again.”

Feel free to correct me, but I think you are misunderstanding something here. You seem to think that the reason you feel hotter in a humid climate is because there is more enthalpy in the air. I don’t think that’s correct. The reason you feel hotter is because your ability to lose heat by sweating is reduced. If enthalpy controlled heat transfer as you are suggesting it would have the same effect on inanimate objects such as a thermometer.

Reply to  Bellman
August 15, 2026 7:58 am

“Feel free to correct me, but I think you are misunderstanding something here. You seem to think that the reason you feel hotter in a humid climate is because there is more enthalpy in the air.”

Go talk to an HVAC engineer. What do you think

H_total = h_d + h_w

is for?

Reply to  Tim Gorman
August 15, 2026 10:14 am

“Go talk to an HVAC engineer. ”

Just answer the question. Is it the enthalpy that makes you feel hotter or is it the humidities effect on your sweat?

https://www.hvacbase.org/how-does-humidity-affect-temperature

Humidity makes the air feel hotter than the actual temperature because high moisture levels prevent sweat from evaporating, which is your body’s primary cooling mechanism. At 90°F with 30% humidity, the air feels like 90°F. But at 90°F with 80% humidity, it feels like 113°F — a 23°F perceived temperature increase from humidity alone.

Reply to  Bellman
August 16, 2026 10:59 am

Is it the enthalpy that makes you feel hotter or is it the humidities effect on your sweat?

Two different things entirely. Enthalpy is a value of total energy in a parcel of air. Inability to evaporate sweat when the air already contains lots of water vapor is a different subject entirely.

I gave you a good article on HVAC. It explained that humidity, i.e., latent heat must also be removed through the system.

A good HVAC should be able to achieve a 20 to 30F drop. Maybe 35F for really good systems. If the humidity is not removed that drops significantly. A good HVAC removing humidity allow your skin to cool better through evaporation.

Reply to  Jim Gorman
August 16, 2026 5:13 pm

Through convection as well.

Reply to  Bellman
August 16, 2026 5:13 pm

I answered the question. h_total = h_d + h_w.

This is used to determine heating loads, cooling loads, humidification energy, dehumidification energy, coil performance, etc.

It is *not* due just to sweating, i.e. evaporation. Since the enthalpy of humid air is higher the gradient between the enthalpy of the air and the enthalpy of your body get smaller – so less heat transfer. This is mostly convective cooling. Evaporation is typically a large component of the thermodynamics but it is not the only one.

Again, why are you here lecturing us on how things work when you obviously don’t know how things work?

Reply to  Tim Gorman
August 16, 2026 6:38 pm

It is *not* due just to sweating, i.e. evaporation. Since the enthalpy of humid air is higher the gradient between the enthalpy of the air and the enthalpy of your body get smaller – so less heat transfer.

But doesn’t heat transfer depend in the difference in temperature not enthalpy?

When you keep going on about air conditioning etc, are you just saying you need more energy to change the temperature of humid air? That’s true, but I’m not interested in changing the temperature of the air, I’m interested in the effect of the air on my body.

Reply to  Bellman
August 17, 2026 8:56 am

But doesn’t heat transfer depend in the difference in temperature not enthalpy?”

Heat transfer has several different possible components, typically radiative, conductive, evaporative, and convective among them.

Evaporation is controlled by vapor-pressure gradient, not temperature gradient.

Conductive is controlled by the temperature gradient.

Convective is controlled by a combination of temperature difference, air buoyancy, air density, air viscosity, and air velocity.

Radiative is controlled by the absolute temperature, not by a temperature difference.

Enthalpy is a direct control on evaporative heat transfer.

Enthalpy indirectly affects conduction because it has an impact on skin temperature.

Enthalpy of air affects the heat removed via convection because of the ability of the air to absorb energy. Enthalpy indirectly affects convection because it has an impact on skin temperature.

Enthalpy doesn’t affect radiative heat loss.

You *could* learn all this if you were willing to actually study the subject instead of coming on here and lecturing others on things you really don’t know anything about.

Reply to  Clyde Spencer
August 12, 2026 8:09 am

Two different objects can be at the same temperature yet the total heat flux from each can be totally different. That means they are *NOT* in thermal equilibrium even though the temperatures are the same.

That means you cannot say the two objects are in thermal equilibrium even though they have the same temperature.

The global mean temperature is *NOT* able to assess global change. The global mean temperature can remain the same while massive temperature changes occur in the temperatures involved. That results in the “mean” temperature is not fit-for-purpose.

100K and 200K have the average of 150K. So does 50K and 250K. Big difference in temperatures but no difference in the mean.

So would the “global mean temperature” be able to identify the changes in temperature? Ans: NO

Would the changes in temperature have an impact on the global “climate”? Ans: YES

It’s even worse when you consider that the “global mean temperature” is not a mean. It is based on mid-range temperatures which are defined by the RANGE of temperatures and not on an AVERAGE temperature.

Not-fit-for-purpose. Period.

Reply to  Clyde Spencer
August 12, 2026 11:03 am

Worse! It isn’t scientific to ignore the energy involved with latent heat.

Reply to  Bellman
August 12, 2026 8:17 am

“And as I’ve been trying to explain to you many times, if you weight an average temperature you are converting the intensive property temperature into an extensive property of temperature times some extensive property.”

Weighting something is *NOT* the same thing as converting something.

You can “weight” densities based on their values, e.g. multiply the densities by a value. But unless that value CONVERTS the density into an extensive value, the weighting is meaningless as far as intensive properties are concerned.

Mass/m^3 (density) can be weighted by multiplying by a factor of 2 in one case and a factor of 3 in another case. You still wind up with mass/m^3, an intensive property.

You have to multiply mass/m^3 by m^3 (i.e. by volume) in order to CONVERT the value to an extensive value of mass.

What do you multiply temperature by in order to convert it from an intensive property to an extensive property?

Reply to  Tim Gorman
August 12, 2026 8:35 am

“Weighting something is *NOT* the same thing as converting something.”

Huh? If you are weighting by an extensive value you are multiplying each value by an extensive value. That gives you units of intensive times intensive units. How is this not converting to an extensive property?

“You can “weight” densities based on their values, e.g. multiply the densities by a value.”

And if that value is the volume, you get density times volume. I wonder what that is?

“Mass/m^3 (density) can be weighted by multiplying by a factor of 2 in one case and a factor of 3 in another case. ”

And why are you doing that? The weighting is meant to reflect a real world factor that is relevant to your average. You don’t just multiply by random numbers.

“You have to multiply mass/m^3 by m^3 (i.e. by volume) in order to CONVERT the value to an extensive value of mass.”

Yes that’s the point. You really need to try to understand what I’m saying rather than going on rants about what you think I’m saying. It would save so much time.

“What do you multiply temperature by in order to convert it from an intensive property to an extensive property?”

Depends on what you want an average of. If you are talking about surface area, then you multiply by area. In the case of your two stones you would multiply by mass.

Reply to  Bellman
August 13, 2026 3:08 am

Huh? If you are weighting by an extensive value you are multiplying each value by an extensive value. That gives you units of intensive times intensive units. How is this not converting to an extensive property?”

That is not weighting. That is CONVERTING – from an intensive property value to an extensive property value.

Weighting is multiplication by a coefficient reflecting importance or proportion, such as weighting temperature measurements by their associated variance.

Converting measurements of intensive property values to an extensive property value is transforming from one scale to another using a physical factor.

“You can “weight” densities based on their values, e.g. multiply the densities by a value.””

This is changing the measurement value based on importance, proportion, etc. You are changing the relationship between the measurements. E.g. multiplying a measurement with a large variance by a factor of 1 while multiplying a different measurement with a smaller variance by a factor of 5 in order to reflect the influence of each on the accuracy of the result.

This does not change the units associated with the measurement in any way, shape, or form.

And if that value is the volume, you get density times volume. I wonder what that is?”

This is a CONVERSION, not a weighting. Each measurement is multiplied by the same conversion factor. The relationship between the measurands is not changed in any manner. When converting temperatures from Fahrenheit to Celsius you don’t multiply one measurement by one factor and a different measurement by a different factor.

Weighting factors are UNITLESS. Conversion factors have units. Weighting factors *typically* sum to 1 although this is not required. Conversion factors are constants, they do not sum.

And why are you doing that? “

Perhaps to reflect a different impact, such as the radius of the orbit of two different objects around a third one. Or maybe to reflect the different influence the shape of the orbits will have. Or perhaps to reflect the different impacts a single-point ignition will have at high rpm versus a dual-point ignition. Or perhaps to equalize the impact of different skill levels by participants in an athletic endeavor, e.g. bowling golf, or drag racing.

Yes that’s the point. You really need to try to understand what I’m saying”

Really? Like when you speak of the “uncertainty of the average” we are supposed to understand what you are saying? When the context of the phrase changes from one assertion to another?

If you are going to lecture us on things scientific then you should use language appropriate to what you are speaking of. Weighting is not conversion. Sampling uncertainty is not measurement uncertainty.

Reply to  Tim Gorman
August 13, 2026 5:09 am

What do you think “area weighted average” means? You are weighting each value by the area it represents, hence you are multiplying by area.

Reply to  Bellman
August 13, 2026 6:02 am

What do you think “area weighted average” means? You are weighting each value by the area it represents, hence you are multiplying by area.”

You are *NOT* weighting anything!

You are converting from an intensive value to a extensive value. In essence you are SCALING, not weighting.

W/m^2 is an intensive property. It is an INTENSITY. Every m^2 being radiated gets the same flux. It is *NOT* based on the system size. It doesn’t matter if the total area is 10 m^2 or 100 m^2.

If you want TOTAL rate (another intensive property) that energy is being transmitted over the area, you have to multiply by the area. That converts the INTENSITY to a rate: joules/sec. You then have to multiply by the time interval to get the EXTENSIVE property value of joules.

Each step is a scaling from one measure to another. It does *NOT* change the relationship of the properties in any way, shape, or form.

Weighting *does* change the relationship of the properties.

Weighting does not change the physical meaning of the property, converting does.

Reply to  Tim Gorman
August 13, 2026 6:42 am

If you want TOTAL rate (another intensive property) that energy is being transmitted over the area, you have to multiply by the area. That converts the INTENSITY to a rate: joules/sec. You then have to multiply by the time interval to get the EXTENSIVE property value of joules.

If averaging is the only tool in your toolbox, integration is a really tough hill to climb.

Reply to  Tim Gorman
August 13, 2026 8:01 am

“You are *NOT* weighting anything!!”

You are multiplying each value by the area and dividing by total area. That’s what weighting means.

Weighting *does* change the relationship of the properties.

Yes. In this case by giving more weight to the larger areas.

Reply to  Bellman
August 13, 2026 6:13 am

What do you think “area weighted average” means? You are weighting each value by the area it represents, hence you are multiplying by area.

Your whole argument has gone beyond the reason for using extensive values in averages.

It is as simple as the fact that you can directly add extensive values to find a sum. You cannot add intensive values to obtain a sum except under very specific circumstances.

Averaging physical properties is basically a mixing function. If you mix two masses, each with different values, the final mass will be the sum.

If I mix two intensive substances such as gases, each with different temperatures, the resulting temperature will not be the sum of the individual temperatures. Characteristics of the gases such as latent heat, specific heat values, pressure, and mass of each gas will determine the final temperature of the mixture. Consequently, an average of the two individual temperatures gives the wrong answer.

If every parcel of air, at all altitudes, over the entire globe was identical in its characteristics, temperature could be averaged. That is not the case. Therefore, to compare, and add, temperatures, they must be converted to an extensive value which is generally done to a property called enthalpy.

As much as you would like, you will never justify averaging temperatures of air parcels that have varying characteristics. THE SCIENCE doesn’t support it. You have yet to support your argument with any scientific reference that indicates it is proper to do so. Quit whining and go find a serious physical science reference that sums the temperature of non-identical parcels of air.

Reply to  Jim Gorman
August 13, 2026 8:45 am

You cannot add intensive values to obtain a sum except under very specific circumstances.

Yes, it’s your delusion. You’ve made it abundantly clear you won;t accept any arguments that try to help you see why you are wrong. It doesn’t bother you that people have been using average temperatures for centuries without the fabric of the universe collapsing.

Averaging physical properties is basically a mixing function. If you mix two masses, each with different values, the final mass will be the sum.

No it is not. The main purpose of a mean is to find the central tendency. Yes in some cases you can think of this as what happens when you put all the extensive values together and share them out equally, but that’s just a means to an end. The point of averaging two masses is to find the point mid-way between them.

If I mix two intensive substances such as gases, each with different temperatures, the resulting temperature will not be the sum of the individual temperatures.

Which is why you don’t do that. You don’t have to mix them, just add up all the values. The sum itself has no meaning, but the average does. The sum is only a method to help you get the average.

Characteristics of the gases such as latent heat, specific heat values, pressure, and mass of each gas will determine the final temperature of the mixture.

Again the purpose of averaging is not to find the final temperature of a mixture (although a pro[properly weighted average will do that.).

“Consequently, an average of the two individual temperatures gives the wrong answer.”

“wrong” depends on what you want the average for. An average is a statistic, not a physical thing. The mean of the temperatures of two objects is not their temperature combined, any more than the average of the masses of two objects is the mass of those objects combined.

Therefore, to compare, and add, temperatures, they must be converted to an extensive value which is generally done to a property called enthalpy.

And as I keep saying, if you are averaging by surface area of any other extensive property, that is what you are doing. The average surface area is implicitly the sum (or integral) of temperature multiplied by area, and then divided by total area. This gives you the average surface temperature.

Converting to enthalpy is a different thing. You are measuring something different to temperature and you will have a value that is not the average temperature.

You also need to be clear about what you want to do with enthalpy. Are you measuring enthalpy or specific enthalpy?

As much as you would like, you will never justify averaging temperatures of air parcels that have varying characteristics.

I’ve asked before – but you seem to have a very flexible approach to what you think is possible or impossible. Your initial claim is you can’t average temperatures at all because it’s impossible to add the value up. Then you make an exception when you are averaging the same thing, without explain the difference. Then you extend that to averaging different daily temperatures at the same location, because you can assume they are measuring the same thing. Tim says that averaging temperatures across the UK is fine, becasue again it’s the same thing. You seem to imply it’s alright to average different things provided the air parcels have the same characteristics.

It’s almost as if this has nothing to do with not being able to add intensive properties, and just an excuse to ignore a global average when it’s going up. (You didn’t;t have a problem with global averages when you believed they were pausing.)

Reply to  Bellman
August 14, 2026 4:34 am

 It doesn’t bother you that people have been using average temperatures for centuries without the fabric of the universe collapsing.”

Argumentative fallacy known as Argument to Tradition. Just because it was done in the past does *NOT* mean it was done correctly.

The difference between intensive and extensive properties wasn’t developed till around 1920. The concept was developed to distinguish properties that depend on system size (extensive) and those that do not depend on system size (intensive).

As usual, you demonstrate absolutely *NO* knowledge of physical science at all.

Reply to  Tim Gorman
August 14, 2026 7:31 am

“Argumentative fallacy known as Argument to Tradition. ”

From people who insist the significant figures rules must be obeyed because that’s what they learnt at school.

“The difference between intensive and extensive properties wasn’t developed till around 1920. ”

And people have been averaging intensive properties since then, with no I’ll effects.

Reply to  Bellman
August 14, 2026 7:34 am

From people who insist the significant figures rules must be obeyed because that’s what they learnt at school.

No, because there are sound technical reasons for using them.

Next diversion please…

Reply to  Bellman
August 14, 2026 4:37 am

The point of averaging two masses is to find the point mid-way between them.”

Here comes your assumption that everything is Gaussian. What garbage.

You aren’t even good at statistics let alone physical science.

The general case has no limitation on the number of objects being averaged. The general case does *NOT* assume a symmetrical distribution such as Gaussian.

Reply to  Tim Gorman
August 14, 2026 7:55 am

“Here comes your assumption that everything is Gaussian. What garbage.”

You are so obsessed with Gaussian distributions it’s driving you mad. No. I said nothing about the distribution of the two masses. Two things can it really have a Gaussian distribution. Maybe you mean they are random values from a Gaussian distribution, but that’s irrelevant to what I said.

“The general case has no limitation on the number of objects being averaged.”

Which is why I used the simple example of averaging two values to illustrate the point. For more than two objects you can still see the mean as the mid-point of all the objects, for a suitable definition of middle.

Reply to  Bellman
August 14, 2026 4:40 am

The sum itself has no meaning, but the average does. “

Since the average will be INCORRECT for a final result then it has no physical meaning. It becomes nothing but mathematical masturbation. *YOU* may like doing the calculation but it has no purpose.

Reply to  Tim Gorman
August 14, 2026 8:01 am

“Since the average will be INCORRECT for a final result ”

Why do you think the average will be wrong? How are defining “correct”?

Reply to  Bellman
August 14, 2026 4:42 am

“An average is a statistic, not a physical thing.”

Statistics that have no physical meaning has what purpose? Be exact in your answer. What can you use it for?

A statistical descriptor have one purpose. To summarize the shape, location, and spread of a dataset without assuming any physical relationship between the values. The are the result of a functional operation, not a functional relationship.

Statistics are typically used to compare different distributions. But if those distributions are of intensive properties then the different distributions cannot be compared physically.

It is tempting to say you can compare the distributions by using the statistical descriptors, but you can’t really do that either. It would be similar to comparing the density of ore samples from different mines. All you can do is compare the SHAPE of the distributions.

Exactly what does comparing the shape of the distributions tell you? Be specific.

Reply to  Tim Gorman
August 14, 2026 8:25 am

“Statistics that have no physical meaning has what purpose?”

This is just going round in circles. I didn’t say it had no physical meaning, I said it wasn’t a physical thing.

Reply to  Bellman
August 14, 2026 8:31 am

“But if those distributions are of intensive properties then the different distributions cannot be compared physically. ”

Why? What property does a distribution of intensive properties have that makes it impossible to compare one with another? And how would it be different for an extensive property?

“It would be similar to comparing the density of ore samples from different mines. ”

Why on earth do you think you cannot do that?

Reply to  Bellman
August 14, 2026 5:11 am

The average surface area is implicitly the sum (or integral) of temperature multiplied by area, and then divided by total area.”

Huh? ∫ (T * A) dA? Exactly what do you think this gives you?

The integral becomes something like the integral of °C-m^2 d(m^2) –> (°C * m^3)/3

The derivative of (°C * m^3)/3 dA = (3/3) (°C * m^2) = °C-m^2.

Divide (°C-m^3)/3 by m^2 and you get °C-m/3, that’s a temperature per length divided by 3!

I think your inability to do calculus is showing again.

°C-m^2 has no physical meaning. Temperature does not scale with area – temperature is an intensive property that is not dependent on system size.

Reply to  Tim Gorman
August 14, 2026 8:46 am

“Huh? ∫ (T * A) dA? Exactly what do you think this gives you?”

That’s not what you are doing. It’s ∫ ∫ (T) dA.

Reply to  Bellman
August 14, 2026 5:17 am

“Your initial claim is you can’t average temperatures at all because it’s impossible to add the value up”

Temperatures DO NOT ADD. The values in a distribution may add up. You can do the functional operation known as the “mean” on the VALUES, but the mean that is calculated has NO PHYSICAL MEANING. It’s nothing but mathematical masturbation. You can’t even use the mean you calculate to compare to other distributions of temperatures and obtain a physical comparison. All you can do is to use the shape descriptors, primarily the mean and standard deviation, to see if the shape of the distributions are the same. But that shape comparison does not provide for a useful correlation comparison, causation comparison, or determination of any thermodynamic quantity.

Reply to  Bellman
August 14, 2026 5:23 am

Then you extend that to averaging different daily temperatures at the same location, because you can assume they are measuring the same thing.”

You remain willfully ignorant. Why?

Averaging measurements to obtain a best estimate of the value of the measurand is NOT averaging multiple components of an intensive physical property obtained from multiple samples.

You are *NOT* averaging intensive properties by determining the best estimate of AN INTENSIVE PROPERTY OF A SINGLE SAMPLE by averaging the measurement indications from that single sample.

How many times in how many ways does this have to be explained to you and bdgwx before it finally sinks in?

Reply to  Tim Gorman
August 14, 2026 8:51 am

“You remain willfully ignorant.”

Then what is your excuse for TN1900?

“Averaging measurements to obtain a best estimate of the value of the measurand is NOT averaging multiple components of an intensive physical property obtained from multiple samples.”

That’s what I said. That’s why you claim it’s ok to average daily values to get the monthly average.

“You are *NOT* averaging intensive properties by determining the best estimate of AN INTENSIVE PROPERTY OF A SINGLE SAMPLE by averaging the measurement indications from that single sample. ‘

Could you reword that so it makes sense?

Reply to  Bellman
August 14, 2026 5:28 am

Tim says that averaging temperatures across the UK is fine, becasue again it’s the same thing”

I did *NOT* say that. Your reading comprehension skills are showing again.

I said: We can argue about whether the measurements are averages of intensive properties of physically different locations in the UK and are physically meaningful

And then I went on to say that you can average measurements of a single measurand to obtain a “best estimate” of the value of the measurand.

I did *NOT* say that averaging temperatures across the UK is fine.

Reply to  Tim Gorman
August 14, 2026 5:37 am

You can’t keep your own lies straight.

I used the example of UK average temperature being useful. You replied

You can’t even get this one correct! The averages you are talking about are average MEASUREMENTS of the same thing!

It’s the same thing as TN1900!

It’s no different than taking 10 measurements of the same water bath to get a best estimate for the temperature of the water bath.

https://wattsupwiththat.com/2026/08/09/defining-temperature/#comment-4228175

If you accept it’s legitimate to average 10 measurements of the same thing, you are also saying it’s legitimate to average UK temperature

Reply to  Bellman
August 14, 2026 6:44 am

If you accept it’s legitimate to average 10 measurements of the same thing, you are also saying it’s legitimate to average UK temperature

Well this is nonsense — you are reduced to whining that all temperatures in the English Isles are identical!

Reply to  karlomonte
August 14, 2026 7:18 am

“…whining that all temperatures in the English Isles are identical! ”

That’s not what Tim is saying, at least not if he understands what he:s saying. He’s saying you can treat the mean temperature as the mean of a probability distribution and every measurement as a measurement of thatean plus a random error. That’s the logic of invoking TN1900 Ex 2.

Reply to  Bellman
August 14, 2026 7:37 am

That’s not what Tim is saying, at least not if he understands what he:s saying. He’s saying you can treat the mean temperature as the mean of a probability distribution and every measurement as a measurement of thatean plus a random error. That’s the logic of invoking TN1900 Ex 2.

So via your reading comprehension problems, you are able to see inside someone else’s thoughts and mind?

Impressive.

Reply to  Bellman
August 14, 2026 5:51 am

Which is why you don’t do that. You don’t have to mix them, just add up all the values. The sum itself has no meaning, but the average does. The sum is only a method to help you get the average.

You have also said that, paraphrasing, “if you know the average enthalpy, you don’t know the temperature of each”.

That is the whole point. Even if you know the temperature average, you have no idea what the associated energies are in the different parcels. One parcel can have 70% humidity and another 5%. It makes a large difference in the total heat in a given parcel.

Did you not read the HVAC reference I mentioned earlier? Humidity, that is latent heat, is a significant factor in sizing an HVAC system. You can’t adequately do that by just looking at temperature. Why do you think in a steam power plant that enthalpy is used account for the energy available to turn a turbine? Simple temperature would be a whole lot simpler to use.

Reply to  Jim Gorman
August 14, 2026 7:27 am

“That is the whole point.”

No that’s you shifting the argument. I was talking about how you can look at averaging temperatures, and you’ve shifted to talking about the difference between temperature and enthalpy.

Temperature won’t tell you the enthalpy and enthalpy won’t tell you the temperature. This has nothing to do with averaging.

” It makes a large difference in the total heat in a given parcel.”

If you want people to believe you are an expert in thermodynamics you really need to learn what heat means in the thermodynamic sense.

“Humidity, that is latent heat, is a significant factor in sizing an HVAC system. You can’t adequately do that by just looking at temperature.”

And you can’t do that just knowing the enthalpy. The site you provided doesn’t even mention enthalpy.

Reply to  Bellman
August 14, 2026 7:38 am

I was talking about how you can look at averaging temperatures

Why do you care so deeply about averaging air temperatures?

Is this some kind of religious experience for you?

Reply to  Bellman
August 13, 2026 3:09 am

“Depends on what you want an average of.”

Your lack of reading comprehension skill is showing again. I said:

““What do you multiply temperature by in order to convert it from an intensive property to an extensive property?”

I specified exactly what measurement type I want to convert from intensive to extensive.

I’m still waiting for what conversion factor you use to convert temperature from an intensive property to an extensive property.

This question is very important to understand what Andy is covering in his essay.

My guess is that you don’t get it at all.

Reply to  Tim Gorman
August 13, 2026 5:12 am

“I specified exactly what measurement type I want to convert from intensive to extensive.”

No you didn’t. You need to know if you are averaging by area, volume, mass etc.

Reply to  Bellman
August 13, 2026 6:09 am

tpg: “I specified exactly what measurement type I want to convert from intensive to extensive.”

“No you didn’t. You need to know if you are averaging by area, volume, mass etc.”

tpg:”““What do you multiply temperature by in order to convert it from an intensive property to an extensive property?””

Are you unable to read at all? Does the word temperature somehow not register on your retina’s?

Converting from intensive to extensive doesn’t involve AVERAGING at all! It involves SCALING using physical multipliers and changes the physical meaning of the property being considered.

Reply to  Tim Gorman
August 13, 2026 8:03 am

Are you unable to read at all? Does the word temperature somehow not register on your retina’s?

What temperature? This is an average temperature. That average is in relationship to something. E..g Surface area, mass, volume or whatever.

Reply to  Bellman
August 12, 2026 11:57 am

You are framing the question to get the answer you want, and I expect training it to give answers that will please you.

I asked two AI’s the same question and got the same answer. I have used Grok maybe 5 times in the last two years so little training has been given.

How many other AI’s must I ask for you to believe the answer?

Here is a response from ChatGPT.

Temperature is an intensive property, so averaging temperatures does not automatically give the equilibrium temperature after mixing. That calculation depends on mass, heat capacity, moisture, pressure, and possible phase changes.

So the mean is always mathematically defined, but it is meaningful only when accompanied by a clear statement of the population, sampling method, and weighting. For mixed air masses, reporting subgroup means, the distribution, and variability is often more informative than reporting one overall mean.

ChatGPT approached it from a slightly different standpoint. A mean from different air parcels, that is in widely separated in time or even different stations essentially provides a value for assuming they have been mixed and obtaining a unique value for the combination. Therefore, yes, you can find a mean, but it is only useful when all the air parcels have common characteristics.

Note the need for weighting. When dealing with thermodynamics the weighting should be the mass and even the specific heat of each air parcel. And here we are back to converting temperature into an enthalpy, an extensive value.

Reply to  Jim Gorman
August 11, 2026 8:35 pm

You didn’t ask it the right question.

That is what Deep Thought said to the philosophers after waiting for 7 1/2 million years of compute-bound hibernation for the answer to the question of the meaning of life and everything. 42 was a computed number, but was not a useful number. 🙂

Sparta Nova 4
Reply to  Tim Gorman
August 10, 2026 6:11 am

I typically use Sahara and antarctica for that point.

Reply to  Sparta Nova 4
August 10, 2026 10:16 am

You could also use the savannahs of central Africa and of central US.

Reply to  Jim Gorman
August 9, 2026 3:53 pm

For those of us unfamiliar with the acronym, what is the GUM?

bdgwx
Reply to  Retired_Engineer_Jim
August 9, 2026 4:16 pm

Guide to the Expression of Uncertainty in Measurement.

https://www.bipm.org/en/committees/jc/jcgm/publications

Reply to  bdgwx
August 9, 2026 7:19 pm

Jeez, tough crowd tonight…

Mario Barbafiera
Reply to  Jim Gorman
August 12, 2026 7:05 pm

Bit like the case of a man with one foot in ice and one in boiling water. Supposedly he will be close to a comfortable average. The recent European Heat wave was in the news everyday, however the area above 80deg N was experiencing record cold for that time of the year. In climatology, averages are not always useful.

High Arctic Sets 36 Cold Records In 41 Days; The High Arctic has endured an extraordinary run of summer cold, with 36 of the 41 days from July 1 through August 10 setting record-low daily means.

https://electroverse.uk/high-arctic-sets-36-cold-records-in-41-days-southern-south-america-freezes-reflect-orbitals-mirrors-ipcc-opens-ar7-review/

Reply to  bdgwx
August 10, 2026 1:26 am

Average temperatures are certainly useful to climate alarmists, but not to people who understand Physics and Metrology.

bdgwx
Reply to  Graemethecat
August 10, 2026 3:48 am

You don’t think NIST and BIPM understand physics and metrology?

Reply to  bdgwx
August 10, 2026 4:17 am

They do *NOT* understand metrology at all. If they do understand metrology they do not exhibit that they understand it. All they seem to know is sampling uncertainty and not measurement uncertainty. No sqrt(n) in measurement uncertainty.

Since NIST AND BIPM obviously believe you can average intensive properties of different things and get a physically meaningful value they obviously do not understand physics.

Sparta Nova 4
Reply to  bdgwx
August 10, 2026 6:15 am

Not the way they employ them, no, I do not think they understand.

Robert Cutler
Reply to  Andy May
August 10, 2026 6:02 am

Andy: “I agree, averages of temperatures have little value, especially areal averages”

comment image

ENSO indices: Areal and temporal averages of SST which are clearly not in equilibrium. Do they have any value?

comment image

Average temperature is an index, and indices should be defined based on the application. For example, hemispheric averages do a reasonable job of capturing seasonal trends and there’s no reason to believe they couldn’t capture longer-term trends. No one is being stopped from defining and promoting a different index.

comment image

I suspect Cohler’s goal for his preferred definition of temperature is to discredit those who would use average temperature (e.g. IPCC). Yet, he uses ENSO without acknowledging that it’s an average and refers to it a local temperature.

Reply to  Robert Cutler
August 10, 2026 9:18 am

Here is my problem with ENSO and global temperatures. First, temperature is NOT a good proxy of heat especially with H2O that has a high specific heat.

Second, and more complicated. Let’s describe ENSO.

  • Trade winds push warmed water westward in the Pacific.
  • Warm water is pooled at the west side of the Pacific.
  • Trade winds cease or even reverse letting the warm water slosh eastward.
  • The heat contained in the warm water spreads over a large area.

I will admit that the spreading warmth will change weather patterns over a large part of the earth. However, has the heat content of the earth really changed?

The current treatment by warmists is that somehow the heat content of the earth has increased drastically and is caused by GHG’s. When what is really happened is that the heat stored is simply spread over a large area.

I have pretty much reached the conclusion that temperature change in the Pacific due to ENSO should be subtracted from the global temperature as a bias that upsets the accuracy of the energy being stored at the earth’s surface. Temperature is just not the appropriate proxy with a natural event such as this.

Reply to  Andy May
August 10, 2026 2:55 pm

Thanks for the link. My point is that when temperatures are spread over a larger area from heat already absorbed, that is not a good proxy for increased warming of the earth.

bdgwx
Reply to  Robert Cutler
August 10, 2026 4:53 pm

Yet, he uses ENSO without acknowledging that it’s an average and refers to it a local temperature.

People use a lot of temperature measurements without realizing they are averages. Even the plain old ASOS METAR reports are 5-minute averages.

Reply to  bdgwx
August 12, 2026 8:10 am

Yet ASOS and USCRN meet the requirement of multiple measurement of the same thing over a short period of time (repeatability). Please note that is a component of measurement uncertainty in an uncertainty budget.

JCGM 100:2008

F.1.1.2 It must first be asked, “To what extent are the repeated observations completely independent repetitions of the measurement procedure?” If all of the observations are on a single sample, and if sampling is part of the measurement procedure because the measurand is the property of a material (as opposed to the property of a given specimen of the material), then the observations have not been independently repeated; an evaluation of a component of variance arising from possible differences among samples must be added to the observed variance of the repeated observations made on the single sample.

Please understand what this is telling you. A Type A evaluation on a single sample provides a component of uncertainty but does not determine the difference among various samples. There is also a component determined from the variance between samples. The two are added to find the combined uncertainty.

Averaging temperatures over time is only valid when several physical, statistical, and measurement conditions are satisfied. You may only average temperature values that represent the same physical quantity, measured consistently, over intervals that do not distort the underlying thermodynamics. 

Some necessary conditions. Same sensor properly calibrated, same siting/exposure such as height, shielding, ventilation, and no environmental changes. If these conditions are violated, the average is not physically meaningful because the values do not represent the same underlying quantity.

This is why (Tmax + Tmin)/2 is not a satisfactory average. They are different points on a curve of changing environment conditions. They do not represent the same underlying quantity due to different values of dew point, humidity, wind, heating of the enclosure, etc.

Reply to  Jim Gorman
August 13, 2026 4:02 am

This is why (Tmax + Tmin)/2 is not a satisfactory average.”

It’s not even an “average” statistically. The statistical descriptor known as the “mean” or “average’ is a measure of the center of gravity of a distribution. It is derived from all components in the distribution.

(Tmax + Tmin)/2 is a mid-point value. It is *NOT* derived from all components in the distribution, it is derived from the two end points of the distribution, i.e. the range. The “mean” or “average” is the first moment of a distribution. The mid-point value has no such attribute, it is not a moment of any kind.

Think of a teeter-totter. If you put equal weights on each end then the balance point is exactly in the middle, the same place the mid-range value is at. Put a heavier weight on one end and the balance point changes, this is a skewed “distribution” with more weight at one end than the other. But the mid-point between the end stays the same.

The center-of-gravity describes the *system*. The mid-point value does not describe the system, just the point mid-way between the ends.

The mid-point is pretty much useless for describing the system unless you assume the system has the same weight at each end, i.e. similar to a Gaussian distribution.

And this is what climate science is famous for. Everything is Gaussian. Measurement uncertainty is Gaussian. The daily temperature profile is Gaussian. The daily outgoing radiative flux is Gaussian.

Garbage after garbage after garbage ……

Robert Cutler
Reply to  Jim Gorman
August 13, 2026 6:10 am

“Averaging temperatures over time is only valid when several physical, statistical, and measurement conditions are satisfied”

This is not entirely true. You could average max temperature and dew point if you thought the index would be useful. This was my original point. ENSO is an average of temperature over an area that has a dynamic thermal gradient. That doesn’t mean it can’t be useful for seasonal weather predictions.

Here I compare ENSO metrics to average lower troposphere temperature. The temporal relationships tell us that the 2023 temperature spike is different; it’s coincident with ENSO rather than lagging. I find that useful.

comment image

“This is why (Tmax + Tmin)/2 is not a satisfactory average.”

With modern instruments we can, and should develop better indices. However, if we want to compare contemporary data to historical data then historical indices must also be computed, that is unless you want to fabricate historical data, or you want to limit the comparison, e.g. Tmax-to-Tmax.

“Perfect is the enemy of the good.” — Voltaire

As flawed as they are, I find that most temperature metrics do a good job of capturing temperature dynamics, e.g. is it warmer, or colder than earlier? When did it get warmer?

Reply to  Robert Cutler
August 13, 2026 7:43 am

We’ve had the ability to actually calculate the actual average value of the daily temperature for 40 years. Climate science says you only need 30 years to classify climate.

There is absolutely NO reason why climate science can’t do both in parallel. Use mid-range temperatures AND average temperatures.

I suspect that if they did it would soon become apparent just how piss poor mid-range temperature is as a metric for climate.

Robert Cutler
Reply to  Tim Gorman
August 13, 2026 8:28 am

Why don’t you do that? I’d be interested in your results.

For the UAH data in my plot above, a location is measured approximately twice per day (per sun-synchronous satellite), so not max/min. Their data tracks other datasets, but is not biased by UHI.

Reply to  Robert Cutler
August 13, 2026 9:25 am

or you want to limit the comparison, e.g. Tmax-to-Tmax.

Climate science should move beyond (Tmax+Tmin)/2. If nothing else it hides the changes in each component. Its only use is to scare people into thinking that higher daytime temperatures are what is happening. At least over land, in general, Tmin has risen faster than Tmax creating a higher average. People will never know that using Tavg.

Even individual Tmax and Tmin should be weighted by a UHI component. There are a number brightness maps that use satellites to show the light put off in an area (house and business, street lights, etc.)

Temperatures over large times and large areas really shouldn’t be averaged as they are intensive values. They should be converted to enthalpy instead. We are going on 50 years of automated stations, most of which report humidity values that can be used.

It is far passed the time that climate science moves on to better and more comprehensive thermodynamic evaluations of what is occurring on this planet. Keeping Tavg just because it is useful to compare with 19th and 20th century recorded history is not a good view on science. Time marches on as the song says.

Robert Cutler
Reply to  Jim Gorman
August 13, 2026 10:36 am

I’ve already suggested that we should develop better indices, but for my solar-forcing research I need as much historical data as possible.

Reply to  Jim Gorman
August 13, 2026 11:52 am

Have you looked at ERA5? But it gives hourly values at a high resolution for a huge number of attributes.

I’m guessing you’ll say it doesn’t count as it’s a reanalysis project, but it does demonstrate that climate science is not limited to just min and max temperatures.

Reply to  Bellman
August 13, 2026 12:48 pm

I have examined it but, as you say, it is highly adjusted. My research right now is concentrated on USCRN and other soil temperature sites. Insolation warms the soil which then warms the air. Trying to jump from insolation to air temperatures will never provide a scientific basis for what is actually happening thermodynamically. Focusing on temperature alone is a waste of good time.

Jeff Alberts
Reply to  Kevin Kilty
August 9, 2026 8:33 am

There’s another “typo”: “twitter”.

Reply to  Kevin Kilty
August 9, 2026 6:51 pm

Local temperature seems fine.”

It may be fine for local *use*. It is not fine to couple it with the local temperature from different locations that can have vastly different enthalpy in order to obtain an “average” temperature value. The temperature gradient between the “Arch” in St. Louis and the airport west of St Louis will *not* be sufficiently accurate for any kind of engineering because of different enthalpy for each.

August 9, 2026 6:46 am

On X, I’ve had AGW-promoters argue that LiG temperature measurements are proxies for temperature – equivalent in meaning to tree-ring metrics.

Reply to  Pat Frank
August 9, 2026 6:59 am

After a short search:

The term LiG most commonly refers to:

The Last Interglacial (LIG) period in paleoclimatology (~129,000 to 116,000 years ago),

OR

Laser-Induced Graphene (LIG) in material science and flexible electronics

Reply to  Steve Case
August 9, 2026 7:20 am

Well, try Liquid in Glass thermometry?

“U.S.DEPARTMENT OF COMMERCE/National Bureau of Standards
Liquid-in-Glass Thermometry” by Jacquelyn A. Wise

https://nvlpubs.nist.gov/nistpubs/Legacy/MONO/nbsmonograph150.pdf

Reply to  _Jim
August 9, 2026 7:52 am

“Liquid-in-Glass Thermometry” by Jacquelyn A. Wise also available here:

https://archive.org/details/liquidinglassthe150wise_0/mode/2up

Reply to  _Jim
August 9, 2026 10:23 am

That’s what I had in mind, Jim – Liquid in Glass thermometers. Apologies for being unclear.

Reply to  Pat Frank
August 9, 2026 10:30 am

The ‘regulars’ here knew what you meant, Pat. Not everyone’s shorthand acronyms/nomenclature has the same meaning across all STEM/science/technical pursuits however …

Reply to  Pat Frank
August 9, 2026 1:16 pm

Thanks for clearing that up.
Here I was wondering “LiG” was a Chinese knockoff of a Korean “LG” TV! 😎

Reply to  Pat Frank
August 9, 2026 9:13 am

And they turn around and tell you they know the correct temperature to the one-hundredth (or one-thousandth) of a degree.

They are mathematicians that see numbers first. They are not physical scientists making exasperating measurements on an experiment or prototype design.

Reply to  Jim Gorman
August 9, 2026 9:39 am

“tell you they know the correct temperature to the one-hundredth (or one-thousandth) of a degree. ”

Who says that?

Reply to  Bellman
August 9, 2026 10:37 am

Virtually every climate “scientist” in the World. See Mann’s Hockey Stick graph for an example.

bdgwx
Reply to  Graemethecat
August 9, 2026 12:24 pm

[Mann et al. 1999] report an uncertainty of about 0.5 C.

Reply to  bdgwx
August 10, 2026 1:33 am

Strangely enough, Mann’s famous Hockey Stick graph has the temperature axis in 0.2°C increments.

bdgwx
Reply to  Graemethecat
August 10, 2026 5:31 pm

It is 0.25 C increments.

Reply to  Bellman
August 9, 2026 10:38 am

Trendologists.

Reply to  Bellman
August 9, 2026 11:11 am

Who says that?

comment image

From NOAA. 1895 contiguous U.S. temperatures. Do you really think that temperatures were recorded to the one-hundredths digit with an uncertainty of ±0.005°F in 1895?

bdgwx
Reply to  Jim Gorman
August 9, 2026 12:48 pm

NOAA reports an uncertainty for nClimDiv/USHCN for the temperatures you’ve shown (pre-1900 maximum) to approach ±1.0 F. So no, NOAA does NOT say the uncertainty is one-thousandth or even one-hundredth of degree as you claimed. [Vose et al. 2014]

Reply to  bdgwx
August 9, 2026 2:39 pm

The divisional dataset also has four major weaknesses that render it suboptimal for certain applications, including, to some extent, the estimation of spatial means and temporal trends. First, each divisional value from 1931 to the present is just the arithmetic average of the station data within it, a computational practice that results in a bias when a division is spatially undersampled in a month (e.g., because some stations did not report) or is climatologically inhomogeneous in general (e.g., due to large variations in topography). Second, all divisional values before 1931 stem from state averages published by the U.S. Department of Agriculture (USDA) rather than from actual station observations, producing an artificial discontinuity in both the mean and variance for 1895 to 1930 relative to 1931 to the present (Guttman and Quayle 1996). Third, many divisions experienced a systematic change in average station location and elevation during the twentieth century, resulting in spurious historical trends in some regions (Keim et al. 2003; Keim et al. 2005; Allard et al. 2009). Finally, none of the station-based temperature records contain adjustments for historical changes in observation time, station location, or temperature instrumentation—inhomogeneities that further bias temporal trends (Peterson et al. 1998).

bias adjustments were computed specifically for version 2 to account for changes in observation time, station location, temperature instrumentation, and siting conditions. The first step in this process entailed using the method of Karl et al. (1986) to address documented changes in observation time at COOP stations and to adjust the records to a midnight local standard time (LST) observation schedule (matching ASOS, RAWS, and SNOTEL). COOP station histories were obtained from the NCDC Historical Observing Metadata Repository (HOMR) and the U.S. Historical Climatology Network (HCN; Menne et al. 2009). The second step in the adjustment process involved using the “pairwise” method of Menne and Williams (2009) to address all other documented and undocumented changes at any station in any network.

Figure 7 depicts the differences in average temperature at both the divisional and the national levels. From a divisional perspective, most differences in the annual average are less than 0.5°C, particularly in the East, and most differences in excess of 1.0°C are in the West.

There are a multitude of adjustments I haven’t even put here. Not once, and I emphasize, not once was measurement uncertainty ever addressed. All records were treated as 100% accurate and only varied due to moves and device changes. Even those had no assessment of quality, just a roundabout method of seeing if other stations were similar in order to adjust prior temperature reading.

It is basically a hodgepodge of various modifications made to the actual data in order to create long records that can be used to show the accuracy of a trend. Constructs created from this modified data has no scientific measurement uncertainty ever evaluated or propagated.

Reply to  bdgwx
August 9, 2026 2:50 pm

NOAA reports an uncertainty for nClimDiv/USHCN for the temperatures you’ve shown (pre-1900 maximum) to approach ±1.0 F.

±1.0 F? ASOS stations have an accuracy of ±1.8° F. If you assume that is the repeatability uncertainty and that it has a uniform distribution, when you divide by √3 end up with a standard uncertainty of ±1.8° F. And, that is just one component of an overall uncertainty budget. LIG thermometers are much higher than that. And the study you mention never addresses that at all, not one word.

You might want to show a reference where a measurement that starts with an integer in the measurement uncertainty can support stated values with multiple decimal places.

Reply to  Jim Gorman
August 9, 2026 6:44 pm

re: “LIG thermometers are much higher [inaccuracy] than that.”

Seems to be the common misconception (applicable for common, low-priced budget LiG thermometers however); You would be surprised how accurate LiG thermometers used in lab work can be (once run thru the cal /certification lab), see this earlier reference I cited: https://archive.org/details/liquidinglassthe150wise_0/mode/2up

Reply to  Jim Gorman
August 9, 2026 1:01 pm

Where does NOAA say they know the correct temperature to 0.01°F let alone 0.001°F?

They could be clearer about the uncertainty, but at no point do they claim they are correct to the hundredth if a degree.

Reply to  Bellman
August 9, 2026 2:39 pm

Sorry, but stating a temperature to 2 decimal places implies an accuracy of 0.01ºF..

They should not be stating it to 2 dps.

At most they should be stating whole numbers only…

Reply to  Bellman
August 9, 2026 7:24 pm

Try the North American Dataset and nClimGrid-Daily,

Reply to  Bellman
August 9, 2026 11:17 pm

Where does NOAA say they know the correct temperature to 0.01°F let alone 0.001°F?

Read the very first entry from the table I posted. It is 42.08 F!

The logic is this, If the number is not correct, then it is incorrect. Why is NOAA posting incorrect values? That is far beyond scientific for an agency that supposedly “knows” science.

Reply to  Jim Gorman
August 10, 2026 4:33 am

And my question is, where do they claim that is a “correct” value? I know you believe that every number written down is claiming to be an exact known value. I disagree.

“If the number is not correct, then it is incorrect”

Or it’s unknown. Most things in the real world are unknown. That’s why you have uncertainty.

It would be better if they explicitly stated what the uncertainty is, but this is no different to UAH which also posts monthly anomalies to 2 decimal places with no mention of the uncertainty.

Reply to  Bellman
August 10, 2026 4:49 am

It would be better if they explicitly stated what the uncertainty”

But then they couldn’t justify taking the last decimal place out to the hundredths digit!

“but this is no different to UAH which also posts monthly anomalies to 2 decimal places with no mention of the uncertainty.”

And UAH does *NOT* state their temperature measurements correctly. They show more decimal places than the uncertainty provides for.

If they were to show the actual measurement uncertainty, then they couldn’t be identifying differences out to the hundredths digit! Which means their data sets are actually not fit-for-purpose!

Reply to  Bellman
August 10, 2026 7:19 am

I know you believe that every number written down is claiming to be an exact known value. 

No, he DOESN’T. You are exposing your ignorance again.

Reply to  Bellman
August 10, 2026 1:23 pm

And my question is, where do they claim that is a “correct” value? I know you believe that every number written down is claiming to be an exact known value. I disagree.

What I claim is that every number is an estimate created from a measuring device with a given resolution. Do you know where the term “good enough for government work” originated and how it morphed into the current meaning of being barely adequate?

From the GUM.

1.2 This Guide is primarily concerned with the expression of uncertainty in the measurement of a well-defined physical quantity — the measurand — that can be characterized by an essentially unique value. If the phenomenon of interest can be represented only as a distribution of values or is dependent on one or more parameters, such as time, then the measurands required for its description are the set of quantities describing that distribution or that dependence.

3.1.2 In general, the result of a measurement (B.2.11) is only an approximation or estimate (C.2.26) of the value of the measurand and thus is complete only when accompanied by a statement of the uncertainty (B.2.18) of that estimate.

The fact that a stated value is only an estimate and is representative of an interval within which the true value may lay defines a measurement. Please note, the result is only complete when accompanied by an uncertainty of the estimated value.

Reply to  Jim Gorman
August 10, 2026 2:13 pm

What I claim is that every number is an estimate created from a measuring device with a given resolution.

Stop changing the subject. Your claim was that as NOAA gave their monthly average to 2 decimal places it meant they were claiming it was correct to 2 decimal places.

Reply to  Jim Gorman
August 9, 2026 1:28 pm

What the stat guys do is calculate a mean value but they do not round it the back to the calibration accuracy of thermometer. Back in those old days temperatures were only measured to +/- 1° F. All the temperature data in the table should be round to the nearest whole degree F. Also back in those days special max-min thermometers were used in the weather stations.

Presently, automated weather stations use platinum resistance thermometers which measure temperatures in short intervals.

For a recent US temperature check, I went to:
https://www.extremeweatherwatch/countries/united-states/avergage-temperature-by-year. The Thi and Tlo data from 1901 to 2024 are displayed in table. Here is the data gor these two dates:
Year——-Thi——-Tlo——-Tav Temperatures are ° C
2024——16.8——4.3——-10.5
1901——-14.9——1.6——–8.2
Change—+1.9—-+2.7—–+2.3

Range Thi: 14.7-16.8
Range Tlo: 0.7—4.3

After 123 years the US has warmed by 2.3° C. Note that most of the warming is in Tlo which occurs just before sunrise. Although the Tav has exceeded to the 2015 Paris Accord limit of 1.5° C, I have not read reports of any climate catastrophes occurring in the US except for the mega drought in the US southwest. Lake Mead is now at 23% of capacity.

Reply to  Bellman
August 9, 2026 1:16 pm

NASA formerly had tabular data published on the internet that was expressed to one-thousandth of a degree. They have since more appropriately rounded those to one-hundredth of a degree. However, it is still difficult to defend that when they are averages of mid-range averages.

Reply to  Clyde Spencer
August 9, 2026 1:51 pm

We keep going over the same ground, but expressing something to 3 decimal places is not the same as claiming it is correct to 3 decimal places.

Most global data sets, including UAH, are reported to at least 2 decimal places, but nobody would suggest they are correct to the hundredth if a degree. It’s difficult to even know what “correct” would even mean for the hundredth of a degree.

Reply to  Bellman
August 9, 2026 2:42 pm

something to 3 decimal places is not the same as claiming it is correct to 3 decimal places.”

Well, yes.. it is. ! If you don’t know the correct value for the second or third decimal place, you should not write it.

Reply to  Bellman
August 9, 2026 3:22 pm

We keep going over the same ground, but expressing something to 3 decimal places is not the same as claiming it is correct to 3 decimal places.

Not correct? What happened to your dividing an anomaly of 0.5 by the √1000 stations to obtain a mean correct to ±0.02?

From the GUM:

7.2.6 … Output and input estimates should be rounded to be consistent with their uncertainties; for example, if y = 10,057 62 Ω with uc(y) = 27 mΩ, y should be rounded to 10,058 Ω.

That means a stated value of 79.89 ±0.3 should be rounded to 79.9. If the uncertainty is ±1.0, it should be rounded to 80.



Reply to  Jim Gorman
August 10, 2026 4:52 am

” What happened to your dividing an anomaly of 0.5 by the √1000 stations to obtain a mean correct to ±0.02? ”

That would just be the uncertainty caused by random errors. Any actual estimate of a national average temperature has to take into account all the procedures used in extrapolation the data.

At the simplest if you consider the 1000 stations as a random sample the uncertainty of the average is SEM, which would be the SD of the temperatures divided by √1000.

“That means a stated value of 79.89 ±0.3 should be rounded to 79.9. If the uncertainty is ±1.0, it should be rounded to 80.”

Firstly, I doubt the uncertainty of a monthly US temperature would be as large as ±0.3. Secondly you would normally report uncertainty to 2 significant figures, so it might be 79.89 ±0.32. And thirdly, you can include an extra digit when the figures will be used in further computation, to avoid rounding errors.

Reply to  Bellman
August 10, 2026 6:51 am

Any actual estimate of a national average temperature has to take into account all the procedures used in extrapolation the data.

Nice word salad, shooting from the hip again?

 I doubt the uncertainty of a monthly US temperature would be as large as ±0.3.

What you “doubt” has been shown over and over to be nonsense.

“The error bars can’t be this big!!” — you

Reply to  Bellman
August 10, 2026 7:56 am

That would just be the uncertainty caused by random errors. Any actual estimate of a national average temperature has to take into account all the procedures used in extrapolation the data.

Look up the USCRN manual. It quotes the accuracy (calibration) as ±0.3°C. See it here.
comment image
That is only one part of an uncertainty budget that includes many components including systematic errors such as calibration accuracy. ASOS is no different. It has an accuracy uncertainty of ±1.0°C

Firstly, I doubt the uncertainty of a monthly US temperature would be as large as ±0.3. Secondly you would normally report uncertainty to 2 significant figures, so it might be 79.89 ±0.32. And thirdly, you can include an extra digit when the figures will be used in further computation, to avoid rounding errors.

Let’s address these one by one.

Monthly measurement uncertainty.

Let’s examine NIST TN 1900 Ex. 2.

For example, proceeding as in the GUM (4.2.3, 4.4.3, G.3.2), the average of the m = 22 daily readings is t̄ = 25.6 ◦C, and the standard deviation is s =4.1 ◦C. Therefore, the standard uncertainty associated with the average is u(τ)= s∕m =0.872 ◦C. The coverage factor for 95% coverage probability is k =2.08, which is the 97.5th percentile of Student’s t distribution with 21 degrees of freedom. In this conformity, the shortest 95% coverage interval is t̄± ks∕√n = (23.8 ◦C, 27.4 ◦C).

This example has u(τ)= ±0.872 ◦C and the expanded uncertainty U(τ)= ±1.8 ◦C. They even include a statement that another uncertainty calculation would have a ±2.0◦C which doesn’t rely on an assumption of a particular probability distribution.

It should be noted that this example relies on the assumption that all other measurement uncertainty is negligible and particularly mentions calibration error. This basically says that the uncertainty budget has only one component, reproducibility.

Significant figures.

Let’s see what Bevington says.

When quoting an experimental result, the number of significant figures should be approximately one more that that dictated by the experimental precision. The reason for including the extra digit is to avoid errors that might be caused by rounding errors in later calculations. If the result of the measurement of Example 1.1 is L = 1.979 m with an uncertainty of 0.012 m, this result could be quoted a L = (1.979 ± 0.012) m. However, if the first digit of the uncertainty is large, such as 0.082 m then we should probably quote L = (1.98 ±0.08) m. In other words, we let the uncertainty define the precision to which we quote our result.

Let’s see what Taylor says.

Rule 2.5 Experimental uncertainties should almost always be rounded to one significant figure.

The rule (2.5) has only one significant exception. If the leading digit in the uncertainty dx is a 1, then keeping two significant figures in dx may be better.

Rule 2.9 The last significant figure in any stated answer should usually be of the same order of magnitude (in the same decimal position) as the uncertainty.

Include one extra digit for future calculations.

This is fine if you are going to use the stated value in a multi-variable calculation such as P=nRT/V. However, “P”, after calculation, should not contain extra digits beyond the uncertainty and Taylor recommends that the uncertainty should be a single significant digit at this point.

Reply to  Jim Gorman
August 10, 2026 8:33 am

“Look up the USCRN manual.”

Did you read what I said. Nothing you say at this point has any relevance to my point, which was about how the uncertainty if a regional average is not just the uncertainty if the individual measurements.

“Let’s examine NIST TN 1900 Ex. 2. ”

That’s for the average for a single station. Not a regional average. And as you keep pointing out they are not looking at the uncertainty of the actual average. They treat each daily value as if it were a measurement of a hypothetical average. This is not the uncertainty you want when you are talking about the actual average for that month.

“Let’s see what Bevington / Taylor says.”

I know what they both say. I’ve quoted them to you enough times. They both say you should quote the uncertainty to 1 or 2 significant places, and report the best estimate to the same level of decimal places. This is also the same in the GUM except they are less proscriptive about the number of digits.

Note your TN1900 example. The uncertainty is initially written to 3sf, and the final result it’s rounded to 2sf.

“This is fine if you are going to use the stated value in a multi-variable calculation such as P=nRT/V. ”

Or, say, calculating a linear trend, or an annual average.

“Taylor recommends that the uncertainty should be a single significant digit at this point.”

Funny how you now want to go back to Taylor and ignore the GUM and NIST.

But none of this has anything to do with your original claim that if you right something to 2 decimal places you are claiming that the hundredth figure is correct. Even using Taylor’s 1 SF rule you are invritabky saying you do not know the actual value of the last digit.

Reply to  Bellman
August 10, 2026 10:51 am

which was about how the uncertainty if a regional average is not just the uncertainty if the individual measurements.”

It is the SUM of the uncertainties of the individual measurements.

As usual, you want to ASSume that the measurement uncertainties of the individual measurements are either 0 (zero) or that they all cancel, i.e. a random/Gaussian distribution.

Of course you are correct, in fact. The MEASUREMENT uncertainty of a regional average is not just the MEASUREMENT uncertainty of the individual data points. It also includes the sampling uncertainty – and the sampling uncertainty itself has multiple sampling sub-uncertainties.

But I’m sure that is *NOT* what you mean. Again, you just want to assume that all measurement uncertainty is random, Gaussian, and cancels. You can’t help yourself. That meme is so ingrained in your brain you can’t even recognize when you employ it.

Reply to  Bellman
August 10, 2026 12:33 pm

how the uncertainty if a regional average is not just the uncertainty if the individual measurements.

Not a regional average. And as you keep pointing out they are not looking at the uncertainty of the actual average. They treat each daily value as if it were a measurement of a hypothetical average. This is not the uncertainty you want when you are talking about the actual average for that month.

You are so far out in left field you can’t be seen.

Reply to  Bellman
August 10, 2026 12:51 pm

Note your TN1900 example. The uncertainty is initially written to 3sf, and the final result it’s rounded to 2sf.

Look at what Bevington said again.

The reason for including the extra digit is to avoid errors that might be caused by rounding errors in later calculations.

Was the 0.872 used in a later calculation? It sure was. No problem there. Was the mean rounded to one decimal place just like the most significant digit in uncertainty. Was the uncertainty rounded to one decimal place? Sure was.

I’m unclear where you think NIST didn’t follow any significant digit rules.

But none of this has anything to do with your original claim that if you right something to 2 decimal places you are claiming that the hundredth figure is correct. Even using Taylor’s 1 SF rule you are invritabky saying you do not know the actual value of the last digit.

You need to tell us what junior or senior level university lab courses you have taken. I can assure you that a professor would not let you state a value to two decimals without showing how you measured observations to that resolution.

I guess if you are happy with the government issuing official data that is not correct, that is your prerogative. Some of us were brought up differently.

Reply to  Jim Gorman
August 10, 2026 1:26 pm

“Was the 0.872 used in a later calculation? It sure was”

Yes that was the point I was making. They used 3sf for calculation and 2sf for the final result. I was contrasting it to your claim that you should only use 1sf for the final result.

” Was the uncertainty rounded to one decimal place? ”

1 decimal place because it has 2 significant figures.

Reply to  Bellman
August 10, 2026 1:54 pm

They used 3sf for calculation and 2sf for the final result. I was contrasting it to your claim that you should only use 1sf for the final result.

You have never bothered to work through the example have you? Here are some of the values.

Mean – 25.5909
Var – 16.74729
SD – 4.09234
u(τ) – 0.87249
k – 2.08
U – 1.81477

Now explain why did the numbers end up as:

Mean – 25.6
SD – 4.1
u(τ) – 0.872
U – 1.8

Nothing here stands out as to why there aren’t 4 significant digits using your logic.

Reply to  Jim Gorman
August 10, 2026 2:10 pm

Now explain why did the numbers end up as

Because they round the final uncertainty to 2 significant figures. 1.8 is two figures.

Nothing here stands out as to why there aren’t 4 significant digits using your logic.

4 figures of what? 4 figures of uncertainty is rarely much use, even to avoid rounding errors. If you mean the actual result, the number of significant figures is irrelevant. Write the answer in K and you will have 4 significant figures. What matters is that the final decimal place of the estimate matches the final decimal place of the uncertainty. Quote the uncertainty to the tenth of a degree and the result also has to be to the tenth of a degree.

Reply to  Bellman
August 10, 2026 3:38 pm

Because they round the final uncertainty to 2 significant figures. 1.8 is two figures.

That is not an answer to why. It is regurgitating something that you see. Why did they end up with 1.8? Every calculation would support at least 2 decimal digits.

25.6 has three significant figures. Why did they not follow through and make everything 3 significant figures?

The actual data has 4 significant figures. Why not use 4 SF’s throughout?

Look deeper into why. If you want to argue, you have to answer the tough questions.

Reply to  Jim Gorman
August 10, 2026 4:25 pm

That is not an answer to why. It is regurgitating something that you see. Why did they end up with 1.8?

I’ve no idea what question you want me to answer. They use 2 significant figures because that’s what their convention says. 2 figures seems fair to me. The expanded uncertainty is 1.81477, which rounded to 2 sf is 1.8.

25.6 has three significant figures. Why did they not follow through and make everything 3 significant figures?

Because the number of significant figures is not relevant to the uncertainty. The rule is that you round to the same number of decimal places as the uncertainty. It would make no sense to say the value was 26 when the uncertainty is given as 1.8.

And this makes even less sense for temperatures. The number of significant figures depends on where the zero is in your units. Convert to Kelvin and you would have to write 300 ± 1.8K, with the 300 possibly being anywhere from 250 – 350K.

The actual data has 4 significant figures. Why not use 4 SF’s throughout?

You illustrate one of the problems with using number of figures as a convention. It’s dependent on the base 10 system, it doesn’t allow for fractional values. The temperatures recorded to 4 figures, but that does not mean they represent values to the hundredths of a degree. It’s obvious that the actual values are only recorded to the nearest quarter of a degree.

Here’s a NIST document I found which gives a convention for significant figures in calibration.

https://www.nist.gov/system/files/documents/2019/05/14/glp-9-rounding-20190506.pdf

The summary is that you should do no rounding whilst calculating, then for reporting.

Identify the first two significant digits in the expanded uncertainty. Moving from left to right, the first non-zero number is considered the first significant digit.

Then round using any of three different methods, and finally

Round the reported measurement result, correction, or error to the same number of decimal places as the least significant digit of the uncertainty with both value and uncertainty being in the same units. Both the measurement result and uncertainty will be rounded to the same level of significance.

There are a number of examples given. Here’s one

The volume of a given flask is computed to be 2000.714 431 mL and the uncertainty is 0.084 024 mL. First, round the uncertainty to two significant figures, that is, 0.084 mL. (Do not count the first zero after the decimal point.) Round the calculated volume to the same number of decimal places as the uncertainty statement, that is, 2000.714 mL. Report the volume as 2000.714 mL ± 0.084 mL. Options A and B follow this example. Option C will round the uncertainty to 0.085 mL since it rounds up for evaluation. The result for Option C will be 2000.714 mL ± 0.085 mL.

Does that make it clearer?

Reply to  Bellman
August 11, 2026 6:27 pm

There are a number of examples given. Here’s one

“The volume of a given flask is computed to be 2000.714 431 mL and the uncertainty is 0.084 024 mL. First, round the uncertainty to two significant figures, that is, 0.084 mL. (Do not count the first zero after the decimal point.) Round the calculated volume to the same number of decimal places as the uncertainty statement, that is, 2000.714 mL. Report the volume as 2000.714 mL ± 0.084 mL. Options A and B follow this example. Option C will round the uncertainty to 0.085 mL since it rounds up for evaluation. The result for Option C will be 2000.714 mL ± 0.085 mL.”

Does that make it clearer?

This is an absurd example, whoever came up with it. Resolving 1 microliter in a 2 liter measurement? Good luck.

Then the uncertainty — this is a relative uncertainty of 0.084 / 2000 * 100 = 0.004 %.

Again, good luck with this, you’ll need it.

But numbers is numbers, right?

Reply to  Bellman
August 10, 2026 9:27 am

That would just be the uncertainty caused by random errors.”

No, that would be the SAMPLING UNCERTAINTY, sampling uncertainty is *not* caused by random errors. It is caused by only sampling the population. It is the uncertainty that is generated from not being able to calculate the population mean but only estimating it.

Both random fluctuations and systematic effects affect the measurement value read from the measuring instrument. That generates MEASUREMENT uncertainty, not sampling uncertainty.

You have returned to using the Equivocation fallacy by trying to have the term “uncertainty” mean what you need it to mean in the moment instead of using the more definitive terms “measurement uncertainty” and “sampling uncertainty”.

Measurement uncertainty defines the interval of measurement values that can be reasonably assigned to the value of the measurand. Sampling uncertainty does *NOT* define that interval of measurement values that can be reasonably assigned to the value of the measurand.

Sampling uncertainty only defines how precisely you can locate the mean from sampling. That is *NOT* measurement uncertainty.



Reply to  Tim Gorman
August 10, 2026 9:30 am

Any actual estimate of a national average temperature has to take into account all the procedures used in extrapolation the data.”

You do *NOT* extrapolate the data.

Merriam-Webster:
————————–
extrapolate
1a
: to predict by projecting past experience or known data
—————————-

You do *NOT* predict measurements by projecting past experience or known data. You MEASURE a value of a measurand, you do *NOT* predict what the measurement value will be.

Reply to  Tim Gorman
August 10, 2026 10:06 am

“You do *NOT* extrapolate the data. ”

This nit-picking is just a distraction.

I wasn’t using the word in a strict mathematical sense. It just means you draw a conclusion from the available data.

Simple Merriam-Webster definition

to form an opinion or to make an estimate about something from known facts

Reply to  Bellman
August 10, 2026 10:59 am

This nit-picking is just a distraction.”

It’s not nit picking. It’s pointing out, once again ad infinitum, that you simply throw shite against the wall hoping some of it will stick.

“It just means you draw a conclusion from the available data.”

You mean you calculate statistical descriptors – even when they don’t have physical meaning?

I don’t know what Merriam-Webster reference you are using but that is *NOT* the definition of “extrapolate” in the on-line version.

————————————–
Merriam-Webster
extrapolate
verb transitivs

1
a
: to predict by projecting past experience or known data
Party officials extrapolated public sentiment on one issue from known public reaction on others.
b
: to project, extend, or expand (known data or experience) into an area not known or experienced so as to arrive at a usually conjectural knowledge of the unknown area
Researchers extrapolate present trends to construct an image of the future.

2

: to infer (values of a variable in an unobserved interval) from values within an already observed interval
—————————

Reply to  Tim Gorman
August 10, 2026 9:35 am

“No, that would be the SAMPLING UNCERTAINTY,”

Then Jim needed to be clearer. What he said was

What happened to your dividing an anomaly of 0.5 by the √1000 stations to obtain a mean correct to ±0.02?

I assumed by “anomal” he meant measurement uncertainty. If he meant there was a standard deviation of 0.5, then yes that would be the standard error of the mean, but it’s pretty unlikely that temperatures across a country were all that close together.

The rest of your monotonous word salad if ored. This has nothing to do with the theme of this article.

Reply to  Bellman
August 10, 2026 10:50 am

it’s pretty unlikely that temperatures across a country were all that close together.

Really? So they warmists who tout the fact that anomalies are reliable across 1500 to 2000 meters are wrong?

Reply to  Jim Gorman
August 10, 2026 12:09 pm

1. The NOAA data you were talking about is temperature not anomaly.

2. The USA is a little bigger than a couple of kilometers.

Reply to  Bellman
August 10, 2026 11:06 am

I assumed by “anomal” he meant measurement uncertainty.”

Bullshite! An anomaly is the difference between stated best estimates and NOT the measurement uncertainty of the best estimates.

You just got caught once again and how you are just making excuses – POOR excuses.

 If he meant there was a standard deviation of 0.5, then yes that would be the standard error of the mean, but it’s pretty unlikely that temperatures across a country were all that close together.”

As usual, your reading comprehension skills are atrocious.

“obtain a mean correct to ±0.02?”

Measurement uncertainty doesn’t tell you how correct the mean is. It tells you what interval the value of the measurand is expected to be in.

You are *still* just throwing shite against the wall hoping something will stick!

The rest of your monotonous word salad if ored. This has nothing to do with the theme of this article”

In other words: “Stop pointing out where I am wrong.”

Reply to  Tim Gorman
August 10, 2026 10:08 am

At the simplest if you consider the 1000 stations as a random sample the uncertainty of the average is SEM, which would be the SD of the temperatures divided by √1000.”

There you go again with your Equivocation!

The uncertainty of the average, the SEM, only describes the “best estimate” of the value of the measurand. It does *NOT* describe the interval of values that can reasonably be assigned to the value of the measurand. The accuracy of the “best estimate” is SAMPLING UNCERTAINTY. The interval of values that can be reasonably assigned to the value of the measurand is the MEASUREMENT UNCERTAINTY.

They are *NOT* the same thing.

A measurement is given as:

“best estimate +/- measurement uncertainty

↑ ↑

sampling measurement data

————————-

As usual, you are caught in a catch-22. Are these temperature databases

1.a single sample of large size, or
2.multiple samples of size 1

Pick one and stick with it.

Reply to  Bellman
August 10, 2026 9:34 pm

There is no justification for “normally” reporting uncertainty as 2 digits.

Yes, guard digits are sometimes used in the reporting of something like physical constants. However, the protocol is to use something like brackets or an over-bar to indicate clearly that it is not truly a significant figure.

Reply to  Clyde Spencer
August 11, 2026 6:53 am

“There is no justification for “normally” reporting uncertainty as 2 digits.”

Take it up with NIST.

This is the problem I have with all these discussions about significant figures. I see the “laws” as just being rules of thumb, conventions or stylistic guidelines, but many here seem to treat them as immutable laws, and breaking them is considered scientifc fraud.

The fact that nobody agrees in exactly what those rules are is why I prefer to use common sense and context rather than blindly following rules.

Reply to  Bellman
August 11, 2026 9:41 am

 I see the “laws” as just being rules of thumb, conventions or stylistic guidelines”

Bullshite. You have been given the rules and the reasons for them multiple times. You have ALWAYS ignored them.

The basic concept is that you cannot increase resolution by averaging. The base uncertainty is determined by the resolution of the measurement instrument. That’s the first entry in the Uncertainty Budget. The uncertainty increases from there!

You have never, NOT ONCE, read any of the reference documents for meaning and context. Not even the ISO ones that define how to produce an uncertainty budget. NOT ONCE.

You just continue to cherry pick pieces and parts that you think confirm your misconceptions. And when shown how your assertions based on the cherry picking are wrong, you just ignore the reasons why and continue with the same misconceptions.

There is a REASON for these rules. It’s so that others performing the same experiments or measurements can reliably judge whether or not the results are reasonable. Otherwise metrology just turns out to be a jumble of meaningless garbage, kind of like climate science and temperatures are.

It’s right there in the very first paragraph of the GUM:
——————
0.1 When reporting the result of a measurement of a physical quantity, it is obligatory that some quantitative indication of the quality of the result be given so that those who use it can assess its reliability. Without such an indication, measurement results cannot be compared, either among themselves or with reference values given in a specification or standard. It is therefore necessary that there be a readily implemented, easily understood, and generally accepted procedure for characterizing the quality of a result of a measurement, that is, for evaluating and expressing its uncertainty.
———————-

Reply to  Tim Gorman
August 11, 2026 2:42 pm

“You have been given the rules and the reasons for them multiple times. You have ALWAYS ignored them.”

I quoted what NIST say and was told by Spencer said it wasn’t justified.

Reply to  Bellman
August 11, 2026 6:28 pm

Appeals to Authority, as usual.

Reply to  karlomonte
August 11, 2026 7:04 pm

If you can’t appeal to an authority, how can you claim the rules for significant figures must be obeyed?

Reply to  Bellman
August 12, 2026 7:51 am

The rules for significant figures are based on the logic of being able to compare values for reasonableness. They are not set by any single “authority”.

If you are a mathematician, statistician, or climate scientist to whom numbers is just numbers then being able to compare two different sets of numbers for reasonableness doesn’t mean anything.

If you are a physical scientist performing an experiment and you want to compare your results with the result of the same experiment performed at a different location using different instruments then it is IMPORTANT to know the resolution involved in the measurements made for each experiment. Significant figures are how resolution is transmitted for comparison, along with the other components of measurement uncertainty.

It’s the difference between living in statistical world and the real world.

Reply to  Tim Gorman
August 12, 2026 8:19 am

“If you are a mathematician, statistician, or climate scientist to whom numbers is just numbers then being able to compare two different sets of numbers for reasonableness doesn’t mean anything.”

Because statisticians never want to compare numbers for “reasonableness”? What exactly do you think statistics is?

Reply to  Bellman
August 13, 2026 4:16 am

Because statisticians never want to compare numbers for “reasonableness”? What exactly do you think statistics is?”

I’ll give you the example my son encountered one more time. When he entered university to study microbiology his advisor told him that he didn’t need to take any math course. If he needed statistical analysis of a set of data, just find a math major, especially one majoring in statistics, and have them analyze the data.

What a LOAD OF GARBAGE. I convinced my son to take at least 6 hours of statistics plus any requirements to get them. And it’s a good thing he did.

Some of his fellow students would get math majors to analyze their data and they would get garbage back because the math majors couldn’t identify what data was reasonable and what numbers weren’t. Numbers is just numbers. And his fellows couldn’t judge if the statistics were meaningful or not.

THE BLIND LEADING THE BLIND.

And that is *exactly* what you represent. Numbers is just numbers. They don’t have to be meaningful in the real world. They don’t even have to be reasonable in the real world.

Reply to  Tim Gorman
August 13, 2026 5:55 am

Just answer the question.

Reply to  Bellman
August 13, 2026 6:27 am

Just answer the question.”

I did. The issue is your inability to read and comprehend what you’ve read.

The mean is *NOT* a reasonableness test. Comparing two means is *NOT* a reasonableness test.

You’ve been told over and over and over …, ad infinitum, that systematic uncertainty is not amenable to statistical analysis.

You’ve been given the references from recognized experts in metrology that this is the case.

And yet here you are, trying to lecture us that you *can* identify systematic uncertainty by comparing the statistical descriptor known as the mean. Thus the mean is a valid test of reasonableness.

Unfreakingbelievable.

Reply to  Bellman
August 12, 2026 7:55 am

If you can’t compare measurements then the world becomes a jumble of chaos.

Significant figures are a primary way of communicating the resolution of measurements which allows for comparison of results.

It’s pretty damn obvious that you don’t care about being able to compare results of measurements in order to determine if the difference between two different sets of measurements is reasonable or not. That’s because you are *NOT* a physical scientist. You live in Statistical World and not in the Real World.

Reply to  Tim Gorman
August 12, 2026 8:22 am

“Significant figures are a primary way of communicating the resolution of measurements which allows for comparison of results. ”

They are a simplistic method used before you get into proper uncertainty analysis.

“It’s pretty damn obvious that you don’t care about being able to compare results of measurements in order to determine if the difference between two different sets of measurements is reasonable or not. ”

What do you think statistics is? The whole point is to compare different values to see if they are reasonably similar.

Reply to  Bellman
August 12, 2026 10:41 am

They are a simplistic method used before you get into proper uncertainty analysis.

Hnarf. As if you actually know anything about the subject.

bdgwx
Reply to  karlomonte
August 12, 2026 12:08 pm

Hnarf. As if you actually know anything about the subject.

Well…the IUPAC, IUPAP, BIPM, ISO, and other notable organizations all agreed on the GUM including its handling of significant figures.

Note that BIPM itself has 103 member states including their signatories like NIST, NPL, NMIA, etc.

Don’t get me wrong. These older “rules of thumb” regarding significant figures had their time and place. But now with the worldwide standardization and adoption of a more rigorous approach to handling and reporting of uncertainty (and metrology in general) those old rules are out of date.

Reply to  bdgwx
August 12, 2026 12:46 pm

But now with the worldwide standardization and adoption of a more rigorous approach to handling and reporting of uncertainty (and metrology in general) those old rules are out of date.

The “expert” pontificates … you have zero interest in using real measurement metrology.

Numbers is numbers!

Reply to  bdgwx
August 12, 2026 6:23 pm

Well…the IUPAC, IUPAP, BIPM, ISO, and other notable organizations all agreed on the GUM including its handling of significant figures.

I hate to tell you but significant figures rules have been around for a lot longer than any of the agencies and bodies you name., We were required to use them in high school physics and chemistry in 1968.

Isaac Newton proposed the very first one saying that the number of significant digits couldn’t exceed those in a product’s factors.

In the 1800’s the concept was expanded to include the concept of meaningful measurement precision/resolution. More formal rules came after that so that scientists could be sure that experimental results were reported with information that was actually measured.

All of my university lab classes taught this in the early 1970’s and when taking junior and senior labs taught by professors you would get zero on lab reports that didn’t include a section reviewing measured values and how significant digits were manipulated.

You fluff significant digit rules off as something only fuddy duddy’s do. That is not the case. I have shown you and Bellman a number of current lab notes on the internet for actual classes being taught. These have nothing to do with uncertainty. They have to do with the resolution of recorded data regardless of its value and uncertainty.

Reply to  Bellman
August 13, 2026 4:29 am

They are a simplistic method used before you get into proper uncertainty analysis.”

Proper uncertainty analysis REQUIRES the uncertainties to be representative of the real world. That is, in part enforced by the significant digit rules. If they measurement uncertainties are not given based on significant digit rules then they can’t be compared any better than best estimates can.

“What do you think statistics is? The whole point is to compare different values to see if they are reasonably similar.”

The statistical DESCRIPTOR known as the “mean” describes the DISTRIBUTION, not the individual components. The mean can be the same for two different sets of data. That will not let you determine if either or both are reasonable.

In fact, two different distributions can have the exact same mean AND standard deviation. One distribution can by symmetric and another can be skewed while both have the same mean and standard deviation.

The 5-number statistical descriptors help overcome this but how often do you see ANYONE in climate science offering up 5-number statistical descriptors for the data sets they use?

Reply to  Tim Gorman
August 13, 2026 6:04 am

“Proper uncertainty analysis REQUIRES the uncertainties to be representative of the real world.”

And you think rounding numbers makes them more representative of the real world?

“If they measurement uncertainties are not given based on significant digit rules then they can’t be compared any better than best estimates can.”

You can if you know the actual uncertainty.

“The statistical DESCRIPTOR known as the “mean” describes the DISTRIBUTION, not the individual components.”

Just answer the question. How do you think statisticians compare two means?

“The mean can be the same for two different sets of data.”

A rid of iron and a rod of wood can both have the same length. Therefore measuring things is a waste of time.

“That will not let you determine if either or both are reasonable. ”

What determines the “reasonableness” is how much confidence you have in the mean. That is where the SEM comes in.

“In fact, two different distributions can have the exact same mean AND standard deviation”

Have you given your lecture to statisticians on this matter. I’m sure it will come as great surprise to them.

Reply to  Bellman
August 13, 2026 6:38 am

And you think rounding numbers makes them more representative of the real world?

Yes. Why? Because the rounded number is more representative of the actual measured resolution. You cannot artificially create resolution out of thin air by arithmetic operations.

You conveniently overlook the fact that measurements are ESTIMATES. As the GUM specifies, a measurement is not complete without a statement of the QUANTITATIVE value of the quality of the measurement, i.e., the measurement uncertainty.

You just can’t lose the numbers is numbers attitude, can you?

Reply to  Jim Gorman
August 13, 2026 6:50 am

You cannot artificially create resolution out of thin air by arithmetic operations.

bellman and bg-whatever believe it is possible!

You just can’t lose the numbers is numbers attitude, can you?

Nope.

Reply to  Jim Gorman
August 13, 2026 9:05 am

You cannot artificially create resolution out of thin air by arithmetic operations.

You still don’t get it. The resolution of a mean can be higher than the individual measurements because a mean is not a single measurement. There is nothing “artificial” about it. You are not plucking anything out of thin air, the resolution comes from the data.

You conveniently overlook the fact that measurements are ESTIMATES.

When have I ever claimed otherwise. That’s why you need to quote the uncertainty. But removing digits doesn’t make it a better estimate – it will make it a worse estimate. You are adding uncertainty by removing digits.

The fact that you accept retaining additional digits before calculating in order to avoid rounding errors should be a clue why that is. If the rounded number better reflects reality, why would you care about rounding errors?

You just can’t lose the numbers is numbers attitude, can you?

You seem to think that’s a bad thing. Numbers are numbers. The main thrust of mathematics over the past 3 millennia has been to understand that numbers are numbers and how that allows you to figure things about about the real world. Children need time to learn the importance of abstracting numbers. You start of thinking in concrete terms – “if I have 2 apples and you give me 3 apples how many apples will I have.” Before making the leap to understanding that 2 + 3 = 5. That’s numbers are numbers. It can be a difficult concept that 2, 3,m and 5 have no physical meaning – they are just numbers, but that the beauty is that if the statement is true is true for all (almost all) real world sums.

Reply to  Bellman
August 13, 2026 11:59 am

The resolution of a mean can be higher than the individual measurements because a mean is not a single measurement.

Numbers is numbers. You can divide 1 by 3 and get 0.3333333333 if you wish.

You have no idea what measurement resolution is because you’ve never been accountable for making published measurements. The rest of your scree is meaningless because of that.

You keep making arguments that are not supported with resources. Find one or two textbooks, papers, etc. that say significant digits are not necessary when making and reporting measurements.

Here is a reference for you to read.

https://phys.libretexts.org/

When combining measurements with different degrees of precision, the number of significant digits in the final answer can be no greater than the number of significant digits in the least-precise measured value. There are two different rules, one for multiplication and division and the other for addition and subtraction.

The bolding is not mine, it is in the article. It shows how important this is in physics.

Show some references supporting your assertions.

Reply to  Jim Gorman
August 13, 2026 12:25 pm

” You can divide 1 by 3 and get 0.3333333333 if you wish. ”

Or 1/3 as us mathematins would call it. Any finite number of 3s will just be an approximation. But what I would never call it is 0.

Reply to  Jim Gorman
August 13, 2026 1:28 pm

” Find one or two textbooks, papers, etc. that say significant digits are not necessary when making and reporting measurements. ”

I didn’t say they are not useful, just that they shouldn’t been seen as absolute rules, should be based on calculated uncertainty, and that the “rule” about averaging is just wrong.

I’ve give you sources. But you either ignore them, or claim they are special cases. I quoted a NIST document in these comments and was accused of making an argument from authority. I’ve multiple times pointed you to the two exercises in Taylor where he specifically says that the mean can be quoted to more figures than the individual measurements. I’ve pointed out that the TN1900 example is quoting the average to a higher resolution than the individual measurements

Reply to  Bellman
August 13, 2026 4:13 pm

should be based on calculated uncertainty, and that the “rule” about averaging is just wrong.

Ah yes, SEM Uber Alles.

Reply to  Bellman
August 13, 2026 4:57 pm

“”I’ve pointed out that the TN1900 example is quoting the average to a higher resolution than the individual measurements””

You might want to reread it. The measurements are given to the one-hundredths digit. (18.75 …)

The uncertainty starts at the tenths digit and the average is rounded to the tenths digit to match.

Reply to  Jim Gorman
August 13, 2026 5:56 pm

Oops!

Reply to  Jim Gorman
August 13, 2026 5:59 pm

And as I keep having to explain, this is one of the reasons why just counting significant figures is simplistic. It’s the base 10 system. The measurements are clearly made to the nearest quarter of a degree. But in base 10 a quarter is 0.25, so blindly counting digits makes you think they have a resolution of 0.01, rather than 0.25.

Reply to  Bellman
August 13, 2026 6:46 pm

For what it’s worth, this is what ASTM E29-13 has to say on the subject of significant figures and averages.

7.6 Averages and Standard Deviations—When reporting the average and standard deviation of replicated measurements or repeated samplings of a material, a suggested rule for most cases is to round the standard deviation to two significant digits and round the average to the same last place of significant digits. When the number of observations is large (more than 15 when the lead digit of the standard deviation is 1, more than 50 with lead digit 2, more than 100 in other cases), an additional digit may be advisable.

This uses the standard deviation (not the SEM) as the basis, using 2 sf, but also says you may include an extra digit if the sample is large.

But it then gives an alternative using the SEM

7.6.1 Alternative approaches for averages include reporting x¯ to within 0.05 to 0.5 times the standard deviation of the average σ/√n , or applying rules for retaining significant digits to the calculation of x¯.

Reply to  Bellman
August 14, 2026 6:58 am

Another oops for you:

When reporting the average and standard deviation of replicated measurements or repeated samplings of a material,

Air temperature measurements don’t qualify!

N is always equal to 1!

When will you ever acknowledge this?

Reply to  Bellman
August 14, 2026 12:08 pm

For what it’s worth, this is what ASTM E29-13 has to say on the subject of significant figures and averages.

Good! I am glad that you have changed your mind about significant figures rules being appropriate. In the past, you have denigrated their use, so something has changed your mind.

I don’t have access to the document so I can’t tell the exact context. However, I don’t see a large difference from the GUM. Two sig figs for the SD, where the GUM says one and sometimes two. The value should “round the average to the same last place of significant digits.” So 1.2345 with an SD of 0.056 would be rounded to 1.235 ±0.056 for an interval of 1.179 to 1.291. Using 1.23 ±0.06 give an interval of 1.17 to 1.29. Not much difference. If the rule used is quoted with the document used, there should be no problem.

Reply to  Jim Gorman
August 14, 2026 3:14 pm

Good! I am glad that you have changed your mind about significant figures rules being appropriate.

Stop putting words in my mouth. I said nothing about the appropriateness of the rules, that’s why I said “for what it’s worth”. I just offered them as yet another standard, which gives different results to other standards.

I don’t have access to the document so I can’t tell the exact context.

This was the version I found online – I can;t say how legitimate it is.

https://www.galvanizeit.com/uploads/resources/ASTM-E-29-yr-13.pdf

However, I don’t see a large difference from the GUM. Two sig figs for the SD, where the GUM says one and sometimes two.

The context was the uncertainty of the average. You keep claiming that this can be written to no more decimal places than the individual measurements. Using this if the SD is small enough you could easily be reporting to more places than the individual measurements.

And there a couple of other points you may have missed. If the number of measurements is large it’s advisable to use an additional digit.

And they then go on to give an alternative method based on the size of σ/√N. The result would be equivalent to using the SEM as the uncertainty to 1 significant figure.

Either way, they seem to allow quoting the average to more digits than the resolution of the measurements.

Reply to  Bellman
August 14, 2026 5:30 pm

You are cherry picking (again) — read the title of the standard:

“Standard Practice for Using Significant Digits in Test Data to Determine Conformance with Specifications

It is a very narrow application. Look at the table in 5.3.3.

“Volumetric Tolerance** ± mL”

“**Tolerance limits specified are absolute limits as defined in Practice E29, for Using Significant Digits in Test Data to Determine Conformance with Specifications.”

Also, while the ASTM statistical standards are quite rigorous, ASTM as a whole does not use uncertainty, i.e. the GUM.

Reply to  karlomonte
August 14, 2026 6:28 pm

You are cherry picking (again) — read the title of the standard:

OK, so these all import significant figure rules do not apply when having to conform to specifications.

Given that you rule out ASTM and NIST which international specifications for reporting significant figures are we supposed to follow? The claim is always that you are only allowed to report averages to the same decimal place as the individual measurements, yet no standard ever states this.

Reply to  Bellman
August 15, 2026 6:19 am

Given that you rule out ASTM and NIST which international specifications for reporting significant figures are we supposed to follow?

I did NOT rule them out! You read what you want to read.

YOU are taking the standard out its context and scope!

I dare you — find the term “measurement uncertainty” in ASTM E29.

Reply to  Bellman
August 17, 2026 5:49 am

You need to read up on how “tolerance” and “measurement uncertainty” relate.

If the measurement uncertainty associated with a “part” is significant then the acceptance zone for the pass/fail of a produced part is:

acceptance zone = Tolerance interval – Uncertainty.

In other words the tolerance interval is REDUCED in size to become the acceptance zone. It is part of ISO 17025.

The measurement uncertainty is *NOT* used as the pass/fail interval. Tolerance is defined during design, it is a specification and is not a statistical descriptor.

Uncertainty doesn’t tell you how much variation in a “part” is allowable. Tolerance, i.e. variation, is design, not metrology. Metrology is used in determining if a “part” meets tolerance requirements.

While you can specify design tolerance limits out to several decimal places, measurement uncertainty limits the ability to determine if a “part” meets tolerance limits.

If you look at the acceptance limit equation:

Acceptance zone = Tolerance – Uncertainty

if Uncertainty > tolerance then the acceptance zone goes negative! If the tolerance is REQUIRED by the engineering design then there are basically only two (maybe three) options available. 1.Create a better measurement protocol, e.g. get measurement instruments, or 2. redo the design. The third “maybe” option is to outsource the measurement operation to a source with better capabilities.

Bottom line? If an engineer working for me CONSISTENTLY defines tolerance limits beyond what can be measured, and does *not* provide for upgrading the measurement equipment (costly) he is not going to get a good performance evaluation.

So stop comparing *tolerance* specifications to “measurement uncertainty”. They are *not* the same and the use of significant digits are different for both. One is a DESIGN specification and the other is a measurement capability.

Reply to  Tim Gorman
August 17, 2026 6:26 am

So stop comparing *tolerance* specifications to “measurement uncertainty”. They are *not* the same and the use of significant digits are different for both. One is a DESIGN specification and the other is a measurement capability.

As always, he’s shooting from the hip (again), hoping no one will notice.

Reply to  karlomonte
August 17, 2026 9:00 am

My walls are stained all over from his cherry picked piles of shite he has tried to stick to them.

Reply to  Bellman
August 14, 2026 6:20 pm

This was the version I found online – I can;t say how legitimate it is.

Because you have no idea what you are looking at — cherry picking anything to justify your SEM pseudoscience.

ASTM E29 has this footnote on the first page:

“Current edition approved Aug. 1, 2013. Published August 2013. Originally approved in 1940.”

The ASTM material standards predate measurement uncertainty by many years — they use standard reference materials.

And here you are trying to pound a square peg into a round hole.

And of course you ignored this provision:

7.6 Averages and Standard Deviations—When reporting the average and standard deviation of replicated measurements or repeated samplings of a material,

Once again, your holy air temperature measurements DON’T QUALIFY.

Reply to  Bellman
August 15, 2026 12:32 pm

The context was the uncertainty of the average. You keep claiming that this can be written to no more decimal places than the individual measurements. Using this if the SD is small enough you could easily be reporting to more places than the individual measurements.

You still don’t understand the dance that resolution and uncertainty has do you?

Tell us how the SD can have a smaller value starting value than what the measurement resolution is. Show your calculations.

Here are mine.

25.1, 24.9, 25.2, 25.0, 25.3, 24.8, 25.2, 25.0
Mean = 25.06
SD = = 0.168

Value -> 25.1 ±0.2

Reply to  Jim Gorman
August 15, 2026 1:28 pm

“Tell us how the SD can have a smaller value starting value than what the measurement resolution is.”

Say you have values with a resolution of 1.

1 × 24, 1 × 26, 8 × 25.

SD = 0.222

0.22 is less than 1.

“Here are mine.

25.1, 24.9, 25.2, 25.0, 25.3, 24.8, 25.2, 25.0
Mean = 25.06
SD = = 0.168

Value -> 25.1 ±0.2”

If you are allowing the old ASTM rules as applicable, then they say to quote the SD to 2 significant figures, so your result should say

25.06 ± 0.17.

And if you have a larger sample size you would use an extra digit.

Reply to  Bellman
August 15, 2026 1:50 pm

How about a really really large size, do you get to tack more random digits on the end?

Reply to  Bellman
August 15, 2026 6:30 pm

If you are allowing the old ASTM rules as applicable, then they say to quote the SD to 2 significant figures, so your result should say

They don’t say to always use two digits for the uncertainty. It is a suggested rule.

And if you have a larger sample size you would use an extra digit.

Only if the standard deviation begins with a 1 for 15 samples or if starts with a 2 for 50 samples. This is standard practice.

Reply to  Jim Gorman
August 15, 2026 7:28 pm

It is a suggested rule.

Correct.

This is standard practice.

Funny, you haven’t mentioned it before. You just keep saying the number of decimal places has to agree with the measurements.

Reply to  Bellman
August 15, 2026 6:38 pm

1 × 24, 1 × 26, 8 × 25.

This is far from a Gaussian, Poisson, uniform distribution. It is unimodal. It has an absurd variance, basically a fractional uncertainty of 0.0088%. Do you know how hard that is to obtain?

In the real world I would ask an employee how the outliers occurred. Contamination, blunder, not following procedure properly, etc.

Reply to  Jim Gorman
August 15, 2026 7:56 pm

This is far from a Gaussian, Poisson, uniform distribution.

Moving the goalposts again. You asked how it was possible that the standard deviation could be less than the resolution of the individual measurements.

basically a fractional uncertainty of 0.0088%

I think you might want to check your maths there. 0.22 / 25 = 0.88%.

In any event it’s irrelevant as the uncertainty is not relative. I could have just as easily made the mean 5 and had a SD of 0.22.

In the real world I would ask an employee how the outliers occurred.

If there were no outliers the sd would be 0.

Let’s take an random example. I’ll generate 100 values from a Gaussian distribution with a sd of 0.25 and mean of 25, with each value rounded to the nearest integer.

First try I get 3 ✕ 24, 96 ✕ 25, 1 ✕ 26.

Sample SD = 0.2.

Reply to  Bellman
August 16, 2026 10:34 am

I think you might want to check your maths there. 0.22 / 25 = 0.88%.

Yeah I missed that one.

In any event it’s irrelevant as the uncertainty is not relative. I could have just as easily made the mean 5 and had a SD of 0.22.

Even if the uncertainty is 1%, that is indicative of a measurement device with the resolution to the one-hundredths decimal place. To state a value correctly, you would need to show it as 25.00 ±0.22. Since the recorded data only had integer values, how can you say you know with any certainty, that the decimal uncertainty is correct?

The interval with this uncertainty becomes 24.78 to 25.22. How do you get those values with a device that only measures to integer values? What does someone wanting to duplicate your measurements do? Do i go purchase an expensive device that measure the one-hundredths? You are still doing the numbers is numbers game.

Reply to  Bellman
August 14, 2026 5:56 am

I didn’t say they are not useful, just that they shouldn’t been seen as absolute rules, should be based on calculated uncertainty, and that the “rule” about averaging is just wrong.”

In other words the international convention on significant figures and uncertainty are NOT necessary in order to compare measurements. It’s ok to just make up your own convention.

“I’ve give you sources”

No, you haven’t. Not a single quote. I can give you the title of something and say I’ve given you a source but unless I can quote from that source it’s an illegitimate reference! It’s an argumentative fallacy known as a False Appeal to Authority.

I quoted a NIST document in these comments and was accused of making an argument from authority.”

Here are your quotes on the subject:

“Take it up with NIST”
I quoted what NIST say and was told by Spencer said it wasn’t justified.”

Jeesh, I could say “Take it up with the New Testament” and say I have given you a source!

The appeal to authority accusation was in response to your quote of “Take it up with NIST”. You may as well have said “Take it up with the New Testament”.

Reply to  Tim Gorman
August 14, 2026 9:05 am

“In other words the international convention on significant figures and uncertainty are NOT necessary …”

The logical fallacy if an appeal to authority.

Yest every international convention seems to have different rules, so the argument to authority is not helpful.

“Not a single quote”

If given you numerous quotes, you just have no memory. There’s a lengthy quote just a few comments above.

https://wattsupwiththat.com/2026/08/09/defining-temperature/#comment-4229403

“Here are your quotes on the subject:”

Here’s the actual quote:

Identify the first two significant digits in the expanded uncertainty. Moving from left to right, the first non-zero number is considered the first significant digit.

And the example

The volume of a given flask is computed to be 2000.714 431 mL and the uncertainty is 0.084 024 mL. First, round the uncertainty to two significant figures, that is, 0.084 mL. (Do not count the first zero after the decimal point.) Round the calculated volume to the same number of decimal places as the uncertainty statement, that is, 2000.714 mL. Report the volume as 2000.714 mL ± 0.084 mL. Options A and B follow this example. Option C will round the uncertainty to 0.085 mL since it rounds up for evaluation. The result for Option C will be 2000.714 mL ± 0.085 mL.

https://wattsupwiththat.com/2026/08/09/defining-temperature/#comment-4227963

“The appeal to authority accusation was in response to your quote of “Take it up with NIST””

You do realise I was sending up your and Jim’s common refrain?

Reply to  Bellman
August 14, 2026 12:43 pm

The volume of a given flask is computed to be 2000.714 431 mL and the uncertainty is 0.084 024 mL. First, round the uncertainty to two significant figures, that is, 0.084 mL.

Pulling out this absurd ridiculous example again?

Please find a way into the real world.

Reply to  karlomonte
August 14, 2026 3:33 pm

Repeating a quote for the benefit of Tim who insisted I hadn’t quoted anything.

Feel free to ignore their recommendations, as I say there should be no strict rules about significant figures.

Reply to  Bellman
August 14, 2026 6:12 pm

That you would dig deep and pull a stupid example (twice) about an impossible measurement with TEN DIGITS of resolution shows how desperate you are, and how little you know about the real world.

Reply to  Bellman
August 13, 2026 6:44 am

And you think rounding numbers makes them more representative of the real world?”

When the rounding is done as part of providing proper information about the measurement resolution and uncertainty then YES – IT IS REPRESENTATIVE OF THE REAL WORLD!

Using your logic an infinitely repeating decimal, as in a result from calculating an average, would be 100% accurate to quote out to as many digits as you want.

That is just more STATISTICAL WORLD garbage.

You can if you know the actual uncertainty.”

Unfreakingbelievable! You don’t even *KNOW* the actual uncertainty. A standard deviation of a Gaussian distribution only provides 68% of the possible values. 2 SD’s covers about 95% and 3 SD’s covers about 99.7% of values.

Do you have any idea at all what the term “6 sigma” means? Even “6 sigma” certainty doesn’t mean you know the actual uncertainty – it’s just pretty damn close!

Just answer the question. How do you think statisticians compare two means?”

*YOU* are the one asserting that you can compare means to tell the difference between two distributions. *YOU* tell us how you can identify systematic uncertainty using the mean!

A rid of iron and a rod of wood can both have the same length. Therefore measuring things is a waste of time.”

Garbage! Pure shite thrown against the wall hoping something will stick.

If I go to two different lumberyards and measure all the lumber in each, the distributions from yard1 and from yard2 can have the exact same mean and standard deviation.

Yet the actual data values in each distribution can be different!

So how do you compare the two using the means and the standard deviations?

What determines the “reasonableness” is how much confidence you have in the mean.”

You haven’t learned ONE SOLITARY THING about metrology over the years we’ve tried to educate you. The *MEAN* is *not* the factor you look at to determine reasonableness, the MEASUREMENT UNCERTAINTY IS!

It is the measurement uncertainty that used to determine reasonableness. Does your “best estimate” fall within the uncertainty interval of similar measurments.

tpg: ““In fact, two different distributions can have the exact same mean AND standard deviation””

Have you given your lecture to statisticians on this matter. I’m sure it will come as great surprise to them.”

What is the mean and standard deviation of these two distributions?

-2, 0, +2
-4, +1, +3

(hint: mean = 0)

bdgwx
Reply to  Clyde Spencer
August 11, 2026 5:29 pm

There is no justification for “normally” reporting uncertainty as 2 digits.

Most of the examples in the GUM use 2 significant figures. Several of the examples use 3.

Reply to  bdgwx
August 11, 2026 8:50 pm

I don’t consider the GUM to be sacrosanct. It seems to me that it is sometimes logically inconsistent. I prefer reason over authority.

bdgwx
Reply to  Clyde Spencer
August 12, 2026 8:52 am

I mean the GUM is considered to be THE book on the subject of uncertainty in measurement and is authored by a consortium of the worlds leading metrology institutions like NIST, BIPM, ISO, IUPAC, IUPAP etc. via JCGM.

It used to be sacrosanct here until some of us started posting what it actually says. And with JCGM now working along side the WMO and IPCC on climate related topics my guess is that the WUWT community is going to become increasingly antagonistic toward the GUM.

Reply to  bdgwx
August 12, 2026 12:48 pm

I mean the GUM is considered to be THE book on the subject of uncertainty in measurement

WRONG! You don’t even understand the title of the document — it is a standard method for expressing uncertainty.

Reply to  karlomonte
August 12, 2026 5:54 pm

Makes you wonder how many of these guys telling everyone how and why and measurement uncertainty occurs have actually made measurements that were subject to legal and regulatory review.

Reply to  Jim Gorman
August 12, 2026 6:20 pm

Oh yeah, this number is vanishingly small.

Reply to  bdgwx
August 13, 2026 5:48 am

my guess is that the WUWT community is going to become increasingly antagonistic toward the GUM.”

Is that why most of us QUOTE THE GUM to you when we are trying to show your assertions to be garbage?

bdgwx
Reply to  Tim Gorman
August 14, 2026 5:30 pm

Is that why most of us QUOTE THE GUM to you when we are trying to show your assertions to be garbage?

I have no idea why you quote the GUM. Here you told me that the authors do not understand metrology or physics.

Reply to  bdgwx
August 15, 2026 6:24 am

WHY did he write that?

It is you who refuses to understand the limitations of intensive properties, because they make your holy air temperature trends look stoopid.

bdgwx
Reply to  karlomonte
August 15, 2026 8:21 am

WHY did he write that?

He tells us why. For the physics part It’s because NIST and BIPM average intensive properties of different things. For the metrology part it is because they scale the uncertainty of the average by 1/sqrt(N).

And it’s not just NIST and BIPM that do these things. The IUPAC and IUPAP are also authors of the GUM.

Reply to  bdgwx
August 15, 2026 11:01 am

For the metrology part it is because they scale the uncertainty of the average by 1/sqrt(N).

Trendology pseudoscience — air temperature measurements DO NOT QUALIFY.

Reply to  bdgwx
August 12, 2026 2:22 pm

Most of the examples in the GUM use 2 significant figures. Several of the examples use 3.

7.2.6 The numerical values of the estimate y and its standard uncertainty uc(y) or expanded uncertainty U should not be given with an excessive number of digits. It usually suffices to quote uc(y) and U [as well as the standard uncertainties u(xi) of the input estimates xi] to at most two significant digits, although in some cases it may be necessary to retain additional digits to avoid round-off errors in subsequent calculations.

What does “at most” mean to you? It tells me one digit and in rare cases two digits.

Three digits are only needed when following calculations require it. Even then, when quoting the final values, one digit is the most likely and two digits sometimes.

bdgwx
Reply to  Jim Gorman
August 15, 2026 5:48 pm

What does “at most” mean to you?

Less than or equal. And when combined with “usually suffices” I take it to mean most of the time less than or equal, but sometimes greater than. How much greater than 2? As long as it isn’t “excessive” is the rule given in the GUM. The GUMs guideline is thus distinguished from the A2LA guideline in this regard which makes any argument that there is a standard more difficult to depend.

Reply to  bdgwx
August 15, 2026 6:00 pm

from the A2LA guideline

You have not the faintest idea about which you type—A2LA documents are not “guidelines”.

Significant digit rules are just another roadblock to be driven through in your holy quest to get impossibly tiny “error bars” on your trendology models, so you dredge up anything you think might cast them in a bad light.

No one is buying what you are selling.

bdgwx
Reply to  karlomonte
August 16, 2026 7:38 am

No one is buying what you are selling.

I’m not the one making a big deal about significant figures. That was you, the Gorman’s, and Clyde.

And the way I’ve seen you handle significant figures is arbitrary. You have a set of rules for scientists and a different relaxed set of rules for contrarians.

Anyway, if there is anything that these endless discussions has taught us is that there is no universal standard regarding significant figures. Each institution has their own rules and guidelines.

Reply to  bdgwx
August 16, 2026 8:12 am

“”Anyway, if there is anything that these endless discussions has taught us is that there is no universal standard regarding significant figures.””

Here are some links that refute your statement that there is no universal standard. The issue is not that there is some variation in how to accomplish the proper number of digits.

The goal is that stated values must accurately reflect what resolution was measured without having calculations add resolution and modify detection limits.

https://eng.libretexts.org/Bookshelves/Electrical_Engineering/Electronics/DC_Electrical_Circuit_Analysis_-_A_Practical_Approach_(Fiore)/01%3A_Fundamentals/1.2%3A_Significant_Digits_and_Resolution

https://sites.middlebury.edu/chem103lab/2018/01/05/significant-figures-lab/

https://www2.chem21labs.com/labfiles/jhu_significant_figures.pdf

Every instrument has a smallest resolvable increment. Reporting digits beyond that increment is epistemically invalid.

Significant digits are not about arithmetic—they are about honesty in reporting what is known.

Reply to  Jim Gorman
August 16, 2026 8:36 am

Every instrument has a smallest resolvable increment. Reporting digits beyond that increment is epistemically invalid.

Significant digits are not about arithmetic—they are about honesty in reporting what is known.

Exactly right, and they will never acknowledge this because they are not honest persons.

bdgwx
Reply to  karlomonte
August 16, 2026 5:57 pm

Exactly right, and they will never acknowledge this because they are not honest persons.

Then you should be scolding Jim for presenting a source that says otherwise instead of praising him.

Reply to  Jim Gorman
August 16, 2026 8:41 am

Here are some links that refute your statement that there is no universal standard.

Logic’s not your strong point. You can’t prove a universal truth by looking at specific results.

In any event your three papers don’t agree about averages.

The first says

When performing calculations, the results will generally be no more accurate than the accuracy of the initial measurements.

Which is I think what you claim. But note, it says “in general” not “always”.

The third says

The mean cannot be more accurate than the original measurements. For example, when averaging measurements with 3 digits after the decimal point the mean should have a maximum of 3 digits after the decimal point.

Which is what you usually claim.

But the second says

We have special rules for averaging multiple measurements. Ideally, if you measure the same thing 3 times, you should get exactly the same result three times, but you usually don’t. The spread of your answers affects the number of significant digits in your average; a bigger spread leads to a less precise average. The last significant digit of the average is the first decimal place in the standard deviation. For example, if your average is 3.025622 and your standard deviation is 0.01845, then this is the correct number of significant figures for the average: 3.03, because the first digit of the standard deviation is in the hundredths place, so the last significant digit of the average is in the hundredths place.

So again is using the SD to determine number of significant figures, and it’s entirely possible that this will give you more digits than your individual measurements.

Reply to  Bellman
August 16, 2026 8:45 am

if your average is 3.025622 and your standard deviation is 0.01845, 

You quote this stuff while climatology (and you) claim milli-Kelvins from 1-degree (at best) data.

more digits than your individual measurements

Pseudoscientific nontechnical nonsense.

Reply to  karlomonte
August 16, 2026 9:37 am

I’m quoting the document Jim provided.

Reply to  Bellman
August 17, 2026 8:23 am

No, you aren’t. You can’t read well enough to accurately quote.

Reply to  Tim Gorman
August 17, 2026 9:05 am

Not worth arguing if you are just going to blatantly lie. You can follow the link above, you can see exactly the section I quoted verbatim. If you think I didn’t you need to say exactly where you think I misquoted.

Reply to  karlomonte
August 17, 2026 8:22 am

The measurements were out to three digits. The average is given out to two decimal places. Bellman thinks Two is greater than Three.

It’s not pseudoscientific or nontechnical, it’s just plain magical thinking!

Reply to  Tim Gorman
August 17, 2026 9:18 am

“Bellman thinks Two is greater than Three.”

This constant lying says more about the strength of your argument than it does mine.

The standard deviation has the most significant digit in the SD in the hundredths, so the average is quoted to the second decimal place. That’s the rule that particular document states. The document makes no mention of how many digits the individual measurements were recorded. They may have been made to 100 places or to the nearest integer. If you use this method the only thing that matters is the size of the SD.

And in case you still don’t get it I am not agreeing with the idea of using the SD as the uncertainty. I’m just pointing out that if the 3 documents that were supposed to demonstrate there is a universal rule, disagree as to what that rule is.

Reply to  Bellman
August 17, 2026 9:41 am

The document makes no mention of how many digits the individual measurements were recorded.”

The quote is: “ For example, when averaging measurements with 3 digits after the decimal point the mean should have a maximum of 3 digits after the decimal point.”

Your reading comprehension skills are showing again.

Reply to  Tim Gorman
August 17, 2026 10:05 am

You are mixing up two different quotes from two different documents.

Reply to  Bellman
August 17, 2026 10:44 am

This constant lying says more about the strength of your argument than it does mine.

Inability to understand sarcasm noted.

And in case you still don’t get it I am not agreeing with the idea of using the SD as the uncertainty.

Who cares what you don’t agree with? You have zero authority on the subject.

Standard deviation is how uncertainty is quantified, regardless of your wild fantasies to the contrary.

Deal.

Reply to  karlomonte
August 17, 2026 10:54 am

“Inability to understand sarcasm noted.”

The full quote was

The measurements were out to three digits. The average is given out to two decimal places. Bellman thinks Two is greater than Three.

It’s not pseudoscientific or nontechnical, it’s just plain magical thinking!

None of that is true. The sarcasm is irrelevant.

“Who cares what you don’t agree with? ”

I assume the three or four idiots who who spend so much of their time telling me I’m wrong.

Reply to  Bellman
August 17, 2026 11:41 am

I assume the three or four idiots who who spend so much of their time telling me I’m wrong.

More irony—here you are, endlessly spewing your magic about averaging.

And the main point you ran away from—standard deviation is how uncertainty is quantified, regardless of your wild fantasies to the contrary.

Reply to  karlomonte
August 18, 2026 10:25 am

And the main point you ran away from—standard deviation is how uncertainty is quantified, regardless of your wild fantasies to the contrary.”

He *still* doesn’t get this. It’s the SD of the DATA, not of the sample means.

It’s the SD of the DATA that determines the uncertainty, i.e. the values that can reasonably be assigned to the measurand.

SEM = SD/sqrt(n) IS A SHORTCUT estimate of the sample uncertainty associated with finding the population mean. It is supposed to be the SD of the *population” that is used, not the SD of a single sample. The SEM is technically defined as the standard deviation of the sample meanS determined by the sample distribution. One sample doesn’t give a sample distribution, so the SEM doesn’t even exist when you have only one sample! If you have only two samples the sampling uncertainty, i.e. the SEM, is over 20%!!!!! The shortcut calculation isn’t even minimally accurate unless you have at least 10 samples.

Reply to  Tim Gorman
August 18, 2026 11:18 am

“He *still* doesn’t get this.”

Because it’s wrong. You *still* don’t get that.

Reply to  Bellman
August 18, 2026 12:16 pm

In other words all you have is the Argument by Dismissal fallacy.

Exactly what is wrong with what I said? You continue to use the “uncertainty of the mean” phrase without ever once actually defining if you are talking about the sampling uncertainty of the mean or the uncertainty interval surrounding best estimate.

You try to feed us the bullshite that you can get an accurate SEM with just one sample when the fact is that the SEM doesn’t exist if you have just one sample because you don’t have a sample distribution.

And then you try to feed us the bullshite that measurement uncertainty always cancels so the SEM is the measurement uncertainty.

As KM keeps saying, it’s all so you can make the uncertainty of the temperature data so small it can be used to estimate differences in the milii-kelvin range.

Reply to  Bellman
August 17, 2026 8:20 am

Logic’s not your strong point. You can’t prove a universal truth by looking at specific results.”

And reading comprehension is certainly *NOT* your strong point.

JG: “epistemically”

Epistemic has to do with the limits of what can be known. It’s about the ability to know.

JG: “The goal is that stated values must accurately reflect what resolution was measured”

Resolution determines the limit of what you can know.

So again is using the SD to determine number of significant figures, and it’s entirely possible that this will give you more digits than your individual measurements.”

“For example, when averaging measurements with 3 digits after the decimal point ”

ROFL! The measurement were given out to THREE decimal places. The average is given to TWO decimal places. And you think that TWO is greater than THREE – i.e. more decimal places in the average then in the measurements?

Reply to  Jim Gorman
August 16, 2026 9:36 am

they are about honesty in reporting what is known.

You know what the average is. You know in many cases it will be a better measure than a single measurement, and the more measurements you take the better the average will be. Rounding the average will likely give you a less accurate value.

Let me generate some random variables. I start with taking just 5 measurements of something that has an actual weight of 5.4g. The device has a resolution of 1g, and the uncertainty (standard deviation) is 0.6g.

First try: 5 5 6 6 6
Mean = 5.6g

Rounding to resolution of device 6g.

The mean is closer to actual weight than your rounded figure. Using SD would allow us to write 5.6g.

Second try 5 5 5 5 6
Mean = 5.2g
Rounded mean = 5g

5.2 is closer to actual value than 5.

I repeat the experiment 1000 times and in 80% of cases the figure to 1 decimal place is more accurate than the rounded value is. The rounded value is better in only 7% of cases.

And that’s just with a small sample size. Repeat this with averages based on 30 measurements, and over 1000 runs I had 999 where the raw average value was closer to the actual value compared than the rounded average was. With one tie.

The extra digits in a mean are not fictitious, they are the data telling you something about where the value actually lies, with usually more precision that a rounded value will.

Reply to  Bellman
August 16, 2026 12:10 pm

Nice try. Your numbers are in integers. Remember we are talking about measurements not numbers is numbers.

You should not quote your measurement value with a decimal because you did not measure it. The value you have is created by calculation and not by physical measurement.

The numbers you quote could actually be from (5.0 to 5.9) or (6.0 to 6.9) and you have no way to know.

There are online examples of measurement uncertainty using real data, try to find some and show their results here.

Other sources are Taylor and Bevington books. Show some of their examples using real data.

Reply to  Jim Gorman
August 16, 2026 2:23 pm

Your numbers are in integers. Remember we are talking about measurements not numbers is numbers.

What do you think a measurement is?

The value you have is created by calculation and not by physical measurement.

Yes, becasue I’m simulating making a measurement, to illustrate how they work.

The numbers you quote could actually be from (5.0 to 5.9) or (6.0 to 6.9) and you have no way to know.

I think you mean 4.5 to 5.5 and 5.5 to 6.5, and I do know what they are because I generated them before rounding to the nearest integer.

Other sources are Taylor and Bevington books.

Surely you remember all the times I’ve pointed out what Taylor has to say on averaging and significant figures. Exercise 4.15 and 4.17 for instance.

And a better example 4.17

(a) Based on the 30 measurements in Problem 4.13, what would be your best estimate for the time involved and its uncertainty, assuming all uncertainties are random?

(b) Comment on the number of significant digits in your best estimate, as compared with the number of significant digits in the data.

The data from 4.13 is (time in seconds)

8.16, 8.14, 8.12, 8.16, 8.18, 8.10, 8.18, 8.18, 8.18, 8.24,
8.16, 8.14, 8,17, 8.18, 8.21, 8.12, 8.12, 8.17, 8.06, 8.10,
8.12, 8.10, 8.14, 8.09, 8.16, 8.16, 8.21, 8.14, 8.16, 8.13.

Here’s the answer given in the book

(a) (Final answer for time) = mean ± SDOM = 8.149 ± 0.007 s.

(b) The data have three significant figures, whereas the final answer has four; this result is what we should expect with a large number of measurements because the SDOM is then much smaller than the SD.

Reply to  Bellman
August 16, 2026 3:50 pm

Problem 4.13 only asks for the mean and SD. It does not ask for how the measurement should be stated.

Mean = 8.149
SD = 0.039

Both have three digits which is one more than what was measured that allows one to minimize rounding errors.

It would be reported as 8.15 ±0.04.

Reply to  Jim Gorman
August 16, 2026 3:59 pm

That’s why I was talking about 4.17. I quoted the answer. It specifically says the final answer should have 4 digits.

Reply to  Bellman
August 16, 2026 2:43 pm

You know in many cases it will be a better measure than a single measurement, and the more measurements you take the better the average will be. 

Only for repetitions of the same quantity, air temperatures don’t qualify.

The extra digits in a mean are not fictitious, they are the data telling you something about where the value actually lies, with usually more precision that a rounded value will.

Keep telling yourself this fiction…

Reply to  Bellman
August 17, 2026 8:33 am

You know what the average is. You know in many cases it will be a better measure than a single measurement, and the more measurements you take the better the average will be”

The average is a BEST ESTIMATE. It only applies in certain restrictive situations. Primarily that the measurement uncertainty be a Gaussian distribution.

For instance, if there is any systematic effects in the measurements, more measurements will *NOT* make the average more accurate.

For instance, if the measurement uncertainty is not Gaussian, perhaps because of hysteresis in the instrument, the distribution will be skewed. More measurements will *NOT* make the average any more accurate as a “best estimate” for the value of the measurand.

All you are doing here is highlighting your ingrained meme of “all measurement uncertainty is random, Gaussian, and cancels”.

Reply to  Tim Gorman
August 17, 2026 9:21 am

“For instance, if there is any systematic effects in the measurements, more measurements will *NOT* make the average more accurate.”

The familiar Gorman deflection. Sig fig rules will not help you with systematic errors.

Reply to  Bellman
August 17, 2026 9:44 am

The familiar Gorman deflection. Sig fig rules will not help you with systematic errors.”

*YOU* are the one that said: “You know in many cases it will be a better measure than a single measurement”

That is *NOT* the usual case – because of systematic uncertainty.

It all goes back to your meme of “all measurement uncertainty is random, Gaussian, and cancels” – so the SEM becomes the measurement uncertainty and it can be made arbitrarily small.



Reply to  Tim Gorman
August 17, 2026 10:10 am

“That is *NOT* the usual case – because of systematic uncertainty.”

Then why even bother taking multiple measurements of the same thing with the same instrument?

“so the SEM becomes the measurement uncertainty and it can be made arbitrarily small. ”

I’ll ignore your usual pathetic lies, but will point out that you are describing Taylor, Bevington, the GUM here. They all tell you the SEM, usually under a different name, is the uncertainty of the mean. And yes, they all assume in this that the errors / uncertainties are independent and random. You cannot detect or handle systematic errors by making multiple measurements. That does not mean you believe they do not exist, just that they have to be handled differently.

Reply to  Bellman
August 17, 2026 10:47 am

Then why even bother taking multiple measurements of the same thing with the same instrument?

Because this is the definition of a Type A evaluation of uncertainty.

And again, for the 13,666th time, air temperature measurements don’t qualify for Type A evaluations.

Reply to  karlomonte
August 17, 2026 10:57 am

“Because this is the definition of a Type A evaluation of uncertainty.”

And it would be a useless definition if what Tim said was correct.

Reply to  Bellman
August 17, 2026 11:43 am

A real uncertainty analysis needs BOTH.

But go ahead and keep claiming averaging increases resolution, after all, it is What You Do.

Reply to  Bellman
August 18, 2026 11:03 am

And it would be a useless definition if what Tim said was correct.”

That is *NOT* what I said at all. You simply refuse to read *anything* for meaning and context.

I said: “For instance, if there is any systematic effects in the measurements, more measurements will *NOT* make the average more accurate.”

As Taylor points out, u_total is the sum of u(ran) and u(sys).

u_total = u(ran) + u(sys) for direct addition
u_total = sqrt[ u^2(rand) + u^2(sys) ] for quadrature

Thus u(total) can NEVER be less than u(sys), even if u(ran) is driven to zero. Which, as Bevington points out, is impossible because the larger you make the sample the more likely you are to generate outliers.

Once you have made u(ran) < u(sys) there is no purpose in going further by using larger sampling sizes. Your accuracy limit becomes the systematic effect.

For field temperature measurements it is probably a very good assumption that u(ran) is smaller than u(sys). It therefore becomes idiotic to assume that u(sys) cancels among a collection of measurement devices in different environments – yet that is what climate science does so they can assume that u(total) = u(ran) ==> 0 as n ↑.

Reply to  Tim Gorman
August 18, 2026 11:15 am

What you said I relation to the average being better than individual measurements was

“That is *NOT* the usual case – because of systematic uncertainty.”

Reply to  Bellman
August 19, 2026 4:43 pm

Systematic uncertainty can only be eliminated through calibration corrections.

This is the reason uncertainty budgets are necessary. It itemizes the influence quantities that determine total uncertainty. Funny how you never discuss the elements involved in this. Most of them are not evaluated statistically (Type A uncertainty).

Resolution is one Type B that is normally considered to be systematic. It is evaluated based on knowledge of the instrument.  Resolution uncertainty is introduced when an instrument can only display values in discrete increments, such as the smallest graduation on a scale like an LIG thermometer or the last digit for a digital thermometer like an RTD.

Reply to  Jim Gorman
August 20, 2026 6:02 am

“Resolution is one Type B that is normally considered to be systematic.”

I keep explaining the obvious logic if this, and you just ignore what I say.

Uncertainty will be a systematic error if the variation in your measurements are too small to get past the resolution. This would mean that every measurement you take of the same thing gave you an identical result, so a type A analysis would be pointless.

This isn’t the case if there’s enough random variety on you measurements to keep getting different results. That’s the point of a Type A uncertainty. The assumption has to be that the resolution is high enough that you can get a different result each time.

If you are measuring different thingd, e.g different temperatures across the globe, the variation should always be much greater than the resolution of you instruments, and any errors from the resolution of your instruments become effectively random.

Measure something that is exactly 1.23 with an accurate instrument with a resolution of 0.1, and you always get a result of 1.2 and a systematic error of -0.3.

But measure lots of different things with different sizes, then whilst any individual measurment will have an error between ±0.05, it will be a different error each time. Hence the errors will tend to cancel, just like any random error.

Reply to  Bellman
August 20, 2026 7:04 am

Uncertainty will be a systematic error if the variation in your measurements are too small to get past the resolution. This would mean that every measurement you take of the same thing gave you an identical result, so a type A analysis would be pointless.

This is nonsense, you don’t know what you are talking about.

Type A and Type B are not either-or.

Measure something that is exactly 1.23 with an accurate instrument with a resolution of 0.1, and you always get a result of 1.2 and a systematic error of -0.3.

Uncertainty is not error!

You still can’t grasp the basics, yet you try to lecture others.

Reply to  Bellman
August 20, 2026 12:19 pm

Uncertainty will be a systematic error if the variation in your measurements are too small to get past the resolution. This would mean that every measurement you take of the same thing gave you an identical result, so a type A analysis would be pointless.

Resolution is a floor on uncertainty. You need to use accurate descriptions of measurement. If you obtain multiple measurements of the same value, the first cause is tht resolution is insufficient. The term for a device giving the same readings is precision. That is, it returns to the same value in multiple measurements. You can have low resolution yet high precision. It is up to the person making the measurement to determine what is happening.

If I have a one decimal digit voltmeter (I had one a long time ago), and I measure a voltage reference (calibrated by NIST) of 1.02 volts, I will always obtain 1.0 volts for the reading. The minimum uncertainty interval is 0.95 to 1.05 or ±0.05 volts. My readings are highly precise as I always get the same value. However, my resolution is simply too small to recognize smaller changes.

bdgwx
Reply to  karlomonte
August 17, 2026 4:41 pm

And again, for the 13,666th time, air temperature measurements don’t qualify for Type A evaluations.

You should contact Possolo and tell him that. Post your email and his response in this forum that we can all see what he says.

Reply to  bdgwx
August 17, 2026 5:13 pm

YOU DO IT.

And for the 1,355th time, you forgot to list the assumptions used in that NIST paper.

Reply to  bdgwx
August 20, 2026 8:01 am

You should contact Possolo and tell him that. Post your email and his response in this forum that we can all see what he says.

The real question is why have climate scientists and trendologists not asked NIST to review their work and publish information about how measurement uncertainty should be handled in temperature trends.

Why has NOAA never published any collaboration with NIST about the assessment of uncertainty in climate data for the globe. Lots of NOAA web pages about climate, but none about the temperature data validity whatsoever. No numbers, no equations, no uncertainty budgets, no mention of the JCGM documents and how they are used, nada, nothing. Try and obtain the information about how accuracy values for ASOS and USCRN stations was calculated. The results from that proved to be a worthless endeavor.

You would think with all the money in modeling, some could have been spent working with NIST or other standards bodies to evaluate the proper treatment of uncertainty in global temperature predictions.

Can you show any reference whatsoever that NIST approves averaging the intensive property of temperatures from a variety of different climates provide a relevant value of anything associated with thermodynamic heat?

So no, your statement that anyone on WUWT should contact NIST to register their concerns is simply abdicating your responsibility for doing the appropriate design and theoretical underpinning work yourself.

Reply to  Bellman
August 18, 2026 10:40 am

Then why even bother taking multiple measurements of the same thing with the same instrument?”

Because it can be done in a controlled environment over a short period of time using a calibrated instrument. A Type A measurement uncertainty.

As KM points out, the temperature data base sets will *NOT* meet the requirements for doing a Type A evaluation.

“They all tell you the SEM, usually under a different name, is the uncertainty of the mean.”

Here we go again with your Equivocation. They tell you that it is the SAMPLING UNCERTAINTY associated with trying to find the population mean from a set of samples. IT IS NOT THE MEASUREMENT UNCERTAINTY.

“That does not mean you believe they do not exist, just that they have to be handled differently.”

All of the references tell you that systematic effects can only be handled by making them small enough to be insignificant. Otherwise systematic effects have to be included in the measurement uncertainty budget and ADD to the measurement uncertainty.

Taylor has a whole section, Sect 4.6, on handling random and systematic uncertainties. First, you have to identify what the systematic uncertainties *are*. That’s the purpose of the measurement uncertainty budget.

bdgwx
Reply to  Jim Gorman
August 16, 2026 11:21 am

Here are some links that refute your statement that there is no universal standard.

On the contrary it proves my point. Those rules aren’t consistent with each other and they don’t agree with either the A2LA or the GUM.

Every instrument has a smallest resolvable increment. Reporting digits beyond that increment is epistemically invalid.

Your 2nd source does not agree with this statement.

Reply to  bdgwx
August 16, 2026 3:24 pm

They are consistent. They are based upon measurements and the discipline it takes to post measurements that are reliable and replicable. You are not going to find a universal algorithm used by everyone that has defined step by step instructions.

You have yet to show any resource from a university level course that permits reporting averages with more significant digits than was measured. If you can not locate even one, then you should question what you asserting.

bdgwx
Reply to  Jim Gorman
August 16, 2026 5:55 pm

They are consistent.

They are not. One source says for an average use the number of significant digits from the individual measurements. Another source says to use the most significant digit from the standard deviation.

You have yet to show any resource from a university level course that permits reporting averages with more significant digits than was measured.

Your own source from Middlebury College. Note that I make no endorsement of the college. I have no idea what its accreditation status is in the STEM fields.

Reply to  bdgwx
August 16, 2026 9:40 pm

They are not. One source says for an average use the number of significant digits from the individual measurements. Another source says to use the most significant digit from the standard deviation.

Only because you don’t understand (or refuse to accept) metrology concepts.

bdgwx
Reply to  karlomonte
August 17, 2026 9:52 am

Only because you don’t understand (or refuse to accept) metrology concepts.

I’m not the one who wrote either of the sources Jim posted. If you have an issue regarding the understanding expressed within those sources take it up with him or the authors of those sources.

Reply to  bdgwx
August 17, 2026 10:48 am

 take it up with him

Request DENIED.

bdgwx
Reply to  karlomonte
August 17, 2026 4:40 pm

Request DENIED.

Rules for thee, but not me.

Reply to  bdgwx
August 17, 2026 5:15 pm

Stop whining, not falling for your typical sophistry.

Reply to  bdgwx
August 16, 2026 8:34 am

WTF is a “contrarian”?

Someone who doesn’t buy into the IPCC global warming line you are pushing?

Reply to  bdgwx
August 17, 2026 8:06 am

 Each institution has their own rules and guidelines.”

Because each of them is involved in different things. The number of significant digits used in a tolerance specification has different requirements than the number of significant digits used with measurements that are made with instruments having limited resolution.

YOU are trying to imply that everything should have a common set of rules regardless of the situation. It just shows your total lack of knowledge concerning physical science.

bdgwx
Reply to  Tim Gorman
August 17, 2026 10:11 am

YOU are trying to imply that everything should have a common set of rules regardless of the situation. It just shows your total lack of knowledge concerning physical science.

That is gaslighting at its finest. I have never said that. In fact, I’ve gone to great lengths to minimize your, karlomonte, and Clyde’s arbitrary mandate that there is only one single right way to handle significant digits.

Let me make my position clear.

Stating a measurement to a specific number of digits does not always imply the measurand is known to within those digits.

Stating a measurement to a specific number of digits never implies the measurand is known to within those digits when the uncertainty is explicitly stated.

Having rules or guidelines for the expression of the measurement and/or uncertainty is a good thing.

Having multiple differing rules or guidelines for the expression of the measurement and/or uncertainty is not necessarily a bad thing.

Reply to  bdgwx
August 17, 2026 10:51 am

Stating a measurement to a specific number of digits does not always imply the measurand is known to within those digits.

Translation: “I retain the right to add as many digits as I please to make my results look GOOD.”

Stating a measurement to a specific number of digits never implies the measurand is known to within those digits when the uncertainty is explicitly stated.

Only if you need to put the rules in the rubbish bin and quote impossibly tiny “error bars”.

bdgwx
Reply to  karlomonte
August 17, 2026 4:39 pm

Translation: “I retain the right to add as many digits as I please to make my results look GOOD.”

I have never said that, insinuated that, or advocated for that.

The most obvious example of this is IEEE 754. Just because most computation tools output 7 or 16 digits does not mean that the value is known to 7 or 16 digits respectively.

Reply to  bdgwx
August 17, 2026 5:16 pm

Inability to grok the word “translation” noted.

Reply to  bdgwx
August 17, 2026 6:05 pm

So you lot are crying: “Buh, buh, buh GUM”, “Buh, buh, NIST” and claiming you deserve to manufacture an extra digit via the magic of averaging. This is a loooooong way from finding milli-Kelvin resolution inside Kelvin data.

Looks to me like y’all need to get busy and find something else to cherrypick.

bdgwx
Reply to  karlomonte
August 19, 2026 2:13 pm

So you lot are crying: “Buh, buh, buh GUM”,

You were the one that told me a had to use the GUM. After taken your advice and giving it careful consideration over the years I find nothing offensive enough about it to disqualify it as a legitimate source.

Reply to  bdgwx
August 19, 2026 2:35 pm

So you have no qualms about cherrypicking and misinterpreting it.

Reply to  bdgwx
August 18, 2026 11:27 am

I have never said that, insinuated that, or advocated for that.”

Then why are you saying that using decimal places BEYOND WHAT YOU ACTUALLY KNOW is acceptable?

You are either trying to make yourself look good or you are trying to con people into thinking you know more than you can possibly know.

They are *BOTH* bad. Either a narcissist or a carnival fortune teller. Take your pick.

bdgwx
Reply to  Tim Gorman
August 19, 2026 2:10 pm

Then why are you saying that using decimal places BEYOND WHAT YOU ACTUALLY KNOW is acceptable?

I didn’t say that. If you want to engage in a genuine discussion about something I said then great. But as I keep telling you I’m not going to defend your absurd arguments.

You are either trying to make yourself look good or you are trying to con people into thinking you know more than you can possibly know.

Hardly. I’ve repeated said that I’m not an expert. I’m not even sure I qualify as an amateur. I’ve also said that my knowledge is very weak relative to the information available. And every time I learn something new it reminds me of how little I actually know.

Reply to  bdgwx
August 19, 2026 3:28 pm

The only thing you care about is division by root-N.

Reply to  bdgwx
August 20, 2026 7:17 am

tpg: “Then why are you saying that using decimal places BEYOND WHAT YOU ACTUALLY KNOW is acceptable?”

I didn’t say that.”

You said exactly that.

bdgwx:”Stating a measurement to a specific number of digits never implies the measurand is known to within those digits when the uncertainty is explicitly stated.”

If that stated value, given with digits you don’t know, is used in a subsequent calculation then you are introducing information YOU DON’T KNOW into the sequence.

And you seem to be fine with that based on the statement above.

Such as quoting temperature data to the hundredths digit when the measurement uncertainty is in at least the tenths or units digit (and maybe more) and then using that in the “global average” calculation.

Reply to  Tim Gorman
August 20, 2026 7:36 am

If that stated value, given with digits you don’t know, is used in a subsequent calculation then you are introducing information YOU DON’T KNOW into the sequence.

And you seem to be fine with that based on the statement above.

He is not an honest person.

Reply to  karlomonte
August 20, 2026 9:44 am

Numbers is just numbers.

Reply to  Tim Gorman
August 20, 2026 7:52 am

If that stated value, given with digits you don’t know, is used in a subsequent calculation then you are introducing information YOU DON’T KNOW into the sequence.

That’s why there is uncertainty. Rounding uncertain digits away doesn’t remove that uncertainty, it just adds more uncertainty.

Reply to  Bellman
August 20, 2026 7:55 am

That’s why there is uncertainty. Rounding uncertain digits away doesn’t remove that uncertainty, it just adds more uncertainty.

HAHAHAHHAHAHAHHAHA

The expert speeeeeeeeks, glad I wasn’t sipping coffee when I read this one.

Reply to  Bellman
August 20, 2026 9:49 am

That’s why there is uncertainty. Rounding uncertain digits away doesn’t remove that uncertainty, it just adds more uncertainty.”

That is *NOT* why there is uncertianty.

If this stated value plus uncertainty were put into a database, like a temperature data base from NOAA, with too many digits in the stated value then it would get USED by anyone downloading that database and calculating an average from it. They would be using inaccurate data REGARDLESS of what the total measurement uncertainty is.

Your stated value should state WHAT YOU KNOW – even if it is only a best estimate. Adding digits conveys the idea that you measured with higher resolution than you did – even if you are unsure about their accuracy.

Reply to  Bellman
August 20, 2026 12:00 pm

That’s why there is uncertainty. Rounding uncertain digits away doesn’t remove that uncertainty, it just adds more uncertainty.

You are so full of it. In metrology, known digits in a measurement refer to the digits in a recorded value that are certain and can be trusted to reflect the actual precision of the measurement device, based on its resolution and other uncertainty components.

Reply to  bdgwx
August 16, 2026 10:47 am

As long as it isn’t “excessive” is the rule given in the GUM. 

I showed you this rule in the GUM. Did you not read it or understand it?

7.2.6 The numerical values of the estimate y and its standard uncertainty uc(y) or expanded uncertainty U should not be given with an excessive number of digits. It usually suffices to quote uc(y) and U [as well as the standard uncertainties u(xi) of the input estimates xi] to at most two significant digits, although in some cases it may be necessary to retain additional digits to avoid round-off errors in subsequent calculations.

The GUM says not to use excessive digits and then goes on to define how to keep from excessive digits.

bdgwx
Reply to  Jim Gorman
August 16, 2026 11:06 am

The GUM says not to use excessive digits and then goes on to define how to keep from excessive digits.

Right. And since the GUM gives examples where 3 significant digits are included in the final reported uncertainties then surely 2 significant digits would not be excessive. So I don’t see what the big hang up is here. Does your position align with that of your brother’s that the authors of the GUM do not understand physics and metrology? If so then it would at least be a reason why you continue to challenge the GUM in this regard.

Reply to  bdgwx
August 16, 2026 11:48 am

And since the GUM gives examples where 3 significant digits are included in the final reported uncertainties

What examples did you find with 3-digit uncertainties as a final value. I searched quite a few and didn’t see any.

4.35, 4.37, 4.38, 5.15, 5.22, 7.24, 7.26, F.2.2.1, F.2.2, F2.2.3, H.1.3.2, H.1.6, H.1.7, H.2.3, H.3.3, H.3.4, H.6.5

All these end up with uncertainty of two digits. Did I miss one?

bdgwx
Reply to  Jim Gorman
August 16, 2026 2:26 pm

All these end up with uncertainty of two digits. Did I miss one?

Tables H.3 and H.5.

Sections 4.3.4, 5.2.2

Reply to  bdgwx
August 16, 2026 4:18 pm

H.3.3 Calculation of results

The data to be fitted are given in the second and third columns of Table H.6. Taking t0 = 20 °C as the reference temperature, application of Equations (H.13a) to (H.13g) yields

y1 = 0,171 2° C s(y1) = 0,002 9° C

y2 = 0,002 18 s(y2) = 0,000 67

r(y1,y2) = -0,930 s = 0,003 5° C

I don’t see any 3-digit sig fig 0.0029, 0.0067, and 0.0035 all seem to be two sig figs.

H.3.4 Uncertainty of a predicted value

uc [b(30 °C)] = 0,004 1 °C

Again, I only see 2 sig figs here.

4.3.4

EXAMPLE

u(RS) = (129 μΩ)/2,58 = 50 μΩ

Do you understand what an expanded uncertainty U is and how to determine it? Again, I don’t see an uncertainty with more than 2 sig figs.

Example 2), that equation yields for the combined standard uncertainty of Rref,

10

uc (Rref ) =Σi=1u(Rs ) = 10×(100 mΩ) =. The result

uc (Rref ) = 0,32 Ω obtained from Equation (10) is incorrect because it does not take into account that all of the calibrated values of the ten resistors are correlated.

Again, please point out where the uncertainty has more that 2 sig figs.

bdgwx
Reply to  Jim Gorman
August 16, 2026 5:49 pm

I don’t see any 3-digit sig fig 0.0029, 0.0067, and 0.0035 all seem to be two sig figs.

I don’t either. But I do see 3 significant digits in uc(X) = 0.295 Ω, uc(Z) = 0.236 Ω in the location I said (table H.3).

Again, I only see 2 sig figs here.

Yep. But I do see 3 significant digits in uc(R) = 0.195 Ω, uc(X) = 0.201 Ω, and uc(Z) = 0.204 Ω in the location I said (table H.5).

Do you understand what an expanded uncertainty U is and how to determine it?

Yes.

Again, I don’t see an uncertainty with more than 2 sig figs.

Does your copy of JCGM 100:2008 not have the example of a calibration certificate showing 10.000742 Ω ± 129 µΩ in the location I said (section 4.3.4)?

Again, please point out where the uncertainty has more that 2 sig figs.

It’s where I said it was (section 5.2.2). The example given is of a calibration certificate showing  u(Rs) = 100 mΩ.

Reply to  bdgwx
August 12, 2026 3:29 pm

What do ISO 17025 and A2LA have to say about the matter?

Come on, you’re the expert…you should know.

Reply to  karlomonte
August 13, 2026 6:27 am

Well this is no surprise, none of the metrology/uncertainty experts could answer this basic question…

Reply to  karlomonte
August 14, 2026 5:32 am

You didn’t really expect an answer, did you?

ILAC P14, required under ISO 17025, says the numerical value of an expanded uncertainty should be given to, at most, two significant figures.

The GUM says the same thing, no more than two significant figures. It also says the measurement result should be rounded to the same decimal digit as the rounded uncertainty.

Reply to  Tim Gorman
August 14, 2026 7:03 am

I did not expect any answer — A2LA requires a calibration lab to give relative expanded measurement uncertainties to only one-and-a-half digits!

bdgwx
Reply to  karlomonte
August 14, 2026 5:25 pm

ISO 17025 doesn’t have much to say on the matter. A2LA says it is okay to report the uncertainty to 2 significant digits.

Reply to  bdgwx
August 15, 2026 6:25 am

So you have experience being audited as an ISO 17025 calibration lab by A2LA?

You’ve read all the documents?

Reply to  bdgwx
August 15, 2026 7:22 am

A2LA says it is okay to report the uncertainty to 2 significant digits.

You mischaracterize what it says.

Section 5.3 of ILAC P14

The numerical value of the expanded uncertainty shall be given to, at most, two significant digits. Where the measurement result has been rounded, that rounding shall be applied when all calculations have been completed; resultant values may then be rounded for presentation. For the process of rounding, the usual rules for rounding of numbers shall be used, subject to the guidance on rounding provided i.e in Section 7 of the GUM.

Does that mean two significant digits is ok, yes. Does it mean all, or even most have two digits, no.

Let me repeat what the GUM says.

7.2.6 The numerical values of the estimate y and its standard uncertainty uc(y) or expanded uncertainty U should not be given with an excessive number of digits. It usually suffices to quote uc(y) and U [as well as the standard uncertainties u(xi) of the input estimates xi] to at most two significant digits, although in some cases it may be necessary to retain additional digits to avoid round-off errors in subsequent calculations.

I don’t see much room between the two interpretations.

bdgwx
Reply to  Jim Gorman
August 15, 2026 8:17 am

I don’t see much room between the two interpretations.

I don’t either except for maybe the fact that the GUM always more than 2 significant digits for the uncertainty part, but that isn’t that relevant in this specific context. So I don’t know why you are opposed to the way NASA reports temperature. NASA is well within the guidelines of both A2LA and GUM to report the temperature to 3 decimal places if the uncertainty is given as 0.0xx C despite only actually reporting to 2 decimal places in the tabular data file.

Clyde is opposed to 2 decimal places because he advocates for a different set of rules in which the expression measurement value itself also embodies the uncertainty portion based on the digits presented which is different than the A2LA or GUM guidelines.

Reply to  bdgwx
August 15, 2026 11:03 am

 if the uncertainty is given as 0.0xx C

Here is your problem.

Reply to  bdgwx
August 15, 2026 2:32 pm

NASA is well within the guidelines of both A2LA and GUM to report the temperature to 3 decimal places if the uncertainty is given as 0.0xx C despite only actually reporting to 2 decimal places in the tabular data file.

As I told Bellman, you are making up hypotheticals that are not reality.

Resolution is floor for uncertainty. Read junior and senior level college lab notes that are all over the internet. No amount of averaging, interpolation, extrapolation, or statistical thuggery will allow the divination of values beyond what was measured.

The significant digit rule for addition/subtraction is that the value with the smallest number of significant digits controls the total amount of sig figs allowed. When you find the standard deviation, what value controls the number of sig figs? It isn’t the mean, it is the value itself.

Let’s examine a value like 25.2. Exactly what can the next digit in the one-hundredths be? It could be 25.20 to 25.25. Do you know which one in that interval is correct? If it wasn’t measured, there is no way to know. It is part of the great unknown. And that is just on the positive side. It also depends on the instrument. It could be anywhere from 25.11 to 25.29.

We are looking at temperatures but remember measurements can be anything, length, width, depth, hardness, volts, amps, resistance, force, velocity, etc. The measuring devices are numerous. In order to make a system that one can rely on for accuracy, precision, resolution, etc. there needs to be common rules.

Reply to  Jim Gorman
August 15, 2026 3:07 pm

” Read junior and senior level college lab notes that are all over the internet. No amount of averaging, interpolation, extrapolation, or statistical thuggery will allow the divination of values beyond what was measured. ”

Yet nobody has come up with a grown up standards document that says that. Every timevthetvarevsaying to use the uncertainty to determine the number of digits.

“The significant digit rule for addition/subtraction is that the value with the smallest number of significant digits controls the total amount of sig figs allowed.”

No. All these SF rules tell you that when adding or subtracting it’s the value with the fewest decimal places that determine the result.

“Let’s examine a value like 25.2. Exactly what can the next digit in the one-hundredths be? It could be 25.20 to 25.25. Do you know which one in that interval is correct?”

You do if you accept the sum is correct. If you have 10 values that sum to 252.3, you know the average of 25.23 is as correct as your sum is.

Reply to  Bellman
August 15, 2026 4:51 pm

No. All these SF rules tell you that when adding or subtracting it’s the value with the fewest decimal places that determine the result.

If I have a measurements of 3251, 3240, 3255, and 3248, exactly where do the decimal digits come from?

Mean, x̄: 3248.5
SD = 6

Tell us how the decimal in the mean is measured or important.

The value should be quoted as 3249 ±6. The 0.5 is subsumed into the uncertainty. The interval will be 3243 to 3255.

This is indicative your obsession with adding resolution to measurements. Think of the consequences of reporting to the first decimal. The next person will need to have a device with 10 times the resolution to reach accurate readings at the one-tenth decimal place. Anything beyond the resolution that was measured is simply fictional.

Here is the lab notes for a physics 2 class at Purdue Univ. that will help. I can list more if you like.
https://web.ics.purdue.edu/~lewicki/physics218/significant

Reply to  Jim Gorman
August 15, 2026 5:48 pm

1) Data are Kelvin, at best.

2) Trendologists need and must have milli-Kelvin.

3) Significant digit rules hinder and prevent #2.

4) Therefore significant digit rules must be circumvented.

5) Trendologists spend hours and hours scouring internet for any manner of loopholes.

This is “science.”

Reply to  Jim Gorman
August 15, 2026 7:25 pm

If I have a measurements of 3251, 3240, 3255, and 3248, exactly where do the decimal digits come from?

Do you every just accept you made a mistake. You said

The significant digit rule for addition/subtraction is that the value with the smallest number of significant digits controls the total amount of sig figs allowed.

Which is the rule for multiplication and division, not for addition and subtraction. Coming up with an example where all your values have both the same number of significant figures and decimal places is just dodging the point.

If the calculation is an addition or a subtraction, the rule is as follows: limit the reported answer to the rightmost column that all numbers have significant figures in common. For example, if you were to add 1.2 and 4.71, we note that the first number stops its significant figures in the tenths column, while the second number stops its significant figures in the hundredths column. We therefore limit our answer to the tenths column.

https://chem.libretexts.org/Bookshelves/Introductory_Chemistry/Introductory_Chemistry_(LibreTexts)/02%3A_Measurement_and_Problem_Solving/2.04%3A_Significant_Figures_in_Calculations

Reply to  Bellman
August 16, 2026 5:47 am

Do you every just accept you made a mistake.

Irony alert.

Reply to  karlomonte
August 16, 2026 6:07 am

You are correct. I mistyped “ever”. Stupid mistake on my part.

Reply to  Bellman
August 16, 2026 6:20 am

Why do you care so much about significant digits?

Are they a threat to your global warming worldview?

Reply to  karlomonte
August 16, 2026 6:57 am

Pay attention. I’m the one who doesn’t care so much about significant figures. I’m arguing against those who insist it’s fraudulent to use the wrong number of digits.

Reply to  Bellman
August 16, 2026 7:01 am

Oh yeah, sure, I believe you…

Reply to  karlomonte
August 16, 2026 7:09 am

You wouldn’t have to “believe’ me if you actually read any of these discussions.

Reply to  Bellman
August 16, 2026 8:37 am

Anything to prop up the warmunist line…

Reply to  Bellman
August 16, 2026 9:26 am

I’m arguing against those who insist it’s fraudulent to use the wrong number of digits.

Go do a Google search on something like “epistemological underpinning of significant digits in physical science”.

You will find that significant digits are the epistemic “footprint” of a measurement. They show how much of the value we can trust, how much is estimated, and how much is uncertain. Please note, this has nothing to do with the rules. Significant digits have been around for centuries, so they are not a newfangled way to tell others the resolution of what was measured.

Reply to  Jim Gorman
August 16, 2026 4:29 pm

Go do a Google search on something like “epistemological underpinning of significant digits in physical science”.

Didn’t come up with anything specific. Did you have an actual philosophical paper in mind?

However it did come up with a few amusing articles you might like:

The important thing to remember about sig figs is that they are an imprecise but typically “good enough” way to deal with errors in basic arithmetic. They’re not an exact science, and are more at home in the “rules of punctuation” schema than they are in the toolbox of a rigorous scientist.

“Rounding off” does a terrible violence to math. Now the error, rather than being a respectable standard deviation that was painstakingly and precisely derived from multiple trials and tabulations, is instead an order-of-magnitude stab in the dark.

That said, sig figs are kind of a train wreck. They are not a good way to accurately keep track of errors. What they do is save people a little effort, manage errors and fudges in a could-be-worse kind of way, and instill a deep sense of fatalism. Significant figures underscore at every turn the limits either of human expertise or concern.

By far the most common use of sig figs is in grading. When a student returns an exam with something like “I have calculated the mass of the Earth to be 5.97366729297353452283 x 1024 kg”, the grader knows immediately that the student doesn’t grok significant figures (the correct answer is “the Earth’s mass is 6 x 1024 kg, why all the worry?”). With that in mind, the grader is now a step closer to making up a grade. The student, for their part, could have saved some paper.

https://www.askamathematician.com/2014/06/q-where-do-the-rules-for-significant-figures-come-from/

Some Simple Rules That Apply Whenever You Write Down a Number

1. Use many enough digits to avoid unintended loss of information.

2. Use few enough digits to be reasonably convenient.

Important note: The previous two sentences tell you everything you need to know for most purposes, including real-life situations as well as academic situations at every level from primary school up to and including introductory college level. You can probably skip the rest of this document.

Seriously: The primary rule is to use plenty of digits. You hardly even need to think about it. Too many is vastly better than too few. To say the same thing the other way: If you ever have more digits than you need and they are causing major inconvenience, then you can think about reducing the number of digits. If you want more-detailed guidance, some ultra-simple procedures are outlined below.

No matter what you are trying to do, significant figures are the wrong way to do it.

People who care about their data don’t use sig figs.

https://openbooks.library.umass.edu/toggerson-131/front-matter/a-note-about-significant-figures/

I also liked this comment in a Reddit thread.

Not going to reproduce all of that here, but essentially sigfigs are a way of getting a set of rules that sort of approximates statistical error analysis without actually doing statistical error analysis. They’re useful only really in the context of a particular application where those particular sigfig rules are applicable. They aren’t generally or mathematically applicable in a rigorous way.

For high school and general education (where no particular application is intended) they are primarily an exercise in rule-following. The education system is trying to make sure that given 10 or so byzantine rules you can in-fact follow those rules in calculations. If you enter an occupation where you don’t ever get exposed to more math and statistics, this means you will do the right thing by applying the rules. You’re going to be graded on the rules at this point, not whether anything makes sense because that is the point of this stage of education. Don’t bother trying to make sense of it at this stage, it’s a waste of everyone’s time. Sometimes you even have an instructor who doesn’t know anything beyond sigfigs and will insist the sigfig rules are some sort of gospel to justify their combination of authority and ignorance.

https://www.reddit.com/r/AskPhysics/comments/1d8sqwc/why_are_the_rules_for_significant_figures_the_way/

Reply to  Bellman
August 16, 2026 9:43 pm

respectable standard deviation that was painstakingly and precisely derived from multiple trials and tabulations

…which you cannot get from air temperature measurements.

N=1

Reply to  Bellman
August 17, 2026 4:49 am

““Rounding off” does a terrible violence to math. Now the error, rather than being a respectable standard deviation that was painstakingly and precisely derived from multiple trials and tabulations, is instead an order-of-magnitude stab in the dark.” (bolding mine, tpg)

This was written by a mathematician or statistician. The very words “precisely derived” is a dead giveaway. Measurement uncertainty affects the ability to “precisely derive” a standard deviation in the real world. The standard deviation is “precisely derived” from imprecise stated values “best estimates”, the imprecision is inherent in the use of imperfect devices.

Only someone with absolutely no understanding of metrology in the real world would make such a statement as “precisely derived”.

“The important thing to remember about sig figs is that they are an imprecise but typically “good enough” way to deal with errors in basic arithmetic. They’re not an exact science, and are more at home in the “rules of punctuation” schema than they are in the toolbox of a rigorous scientist.” (bolding mine, tpg)

Again, this statement was made by a mathematician or statistician. Mathematicians and statisticians living in YOUR blackboard world. Significant digits are meant to convey information about the level of actual, real world knowledge that is available – such as the resolution limit of actual, real world measurements. It’s been that way for literally centuries. Significant digits are *NOT* meant to correct for “errors in basic arithmetic“.

“2. Use few enough digits to be reasonably convenient.”

“real-life situations”

Again, statements made by a mathematician or statistician living in a blackboard world. Convenience is *NOT* the purpose behind the use of significant figures except on a blackboard exercise that consider all measurements to be 100% accurate. Measurement uncertainty *does* exist in real-life situations and must be accounted for whether it is based on resolution limits, random fluctuations, systematic effects, or a combination of all three.

YOU ARE CHERRY PICKING AGAIN. AND ARE DOING IT FROM THE WRITINGS OF MATHEMATICIANS AND STATISTICIANS WHO ARE NOT SUBJECT TO ANY KIND OF SANCTIONS FOR MISREPRESENTING THE ACCURACY OF MEASUREMENTS – be it in bridge span design, capital project cost estimates, safety margins in equipment, and any other situations that those living in the real world are involved in.

Reply to  Tim Gorman
August 17, 2026 5:17 pm

This was written by a mathematician or statistician.
Again, this statement was made by a mathematician or statistician.

Firstly, you have no evidence for those claims, and secondly, it’s such a weird ad hominem argument, as if mathematicians or statisticians are incapable of understanding maths or statistics.

Have you tried clicking on the links to see who might have written them? The first one is from a site called “Ask a Mathematician / Ask a Physicist.” run by two anonymous people calling themselves “the mathematician” and “the physicist”. The article on significant figures was written by “the physicist”.

The second was borrowed from an online book called “Uncertainty as Applied to Measurements and Calculations by John Denker.

https://www.av8n.com/physics/uncertainty.htm

According to the site:

John Denker was an undergrad at Caltech. During his junior year, he founded a successful small software and electronics company which did pioneering work in many fields including security systems, Hollywood special effects, hand-held electronic games, and video games. Also while still an undergrad, he created and taught a course at Caltech: “Designing with Microprocessors”.

His doctoral research at Cornell examined the properties of a gas of hydrogen atoms at temperatures only a few thousandths of a degree above absolute zero, and showed that quantum spin transport and long-lived “spin wave” resonances occur in this dilute Bose gas. Other research concerned the design of ultra-low-noise measuring devices, in which the fundamental quantum-mechanical limitations play an important role.

In 1986-87 he was Visiting Professor at the Institute for Theoretical Physics (University of California, Santa Barbara). He has served on the organizing committee of several major scientific conferences.

There’s quite a bit more – but that seems to be enough. Note, I make no claims about how accurate any of those statements are.

Reply to  Bellman
August 18, 2026 11:50 am

Firstly, you have no evidence for those claims, and secondly, it’s such a weird ad hominem argument, as if mathematicians or statisticians are incapable of understanding maths or statistics.”

No, they are usually incapable of understanding the real world. They usually have the same problem *YOU* do – they never work with actual measurements under conditions of liability. It all just happens on a blackboard where you can ignore all real world impacts such as uncertainty!

The clue is that all of the quotes you give talk about *arithmetic*.

errors in basic arithmetic.”
“a terrible violence to math”
” painstakingly and precisely derived”
“They are not a good way to accurately keep track of errors.”
” Use few enough digits to be reasonably convenient.”
” sigfigs are a way of getting a set of rules that sort of approximates statistical error analysis without actually doing statistical error analysis.”

How does uncertainty translate into errors in basic arithmentic? How does uncertainty do terrible violence to math? How can uncertain quantities be precisely derived? How do you keep track of things you don’t know like true value and errors? How does not stating what you don’t know like you *do* know it become “reasonably convenient”? How is finding the standard deviation of a distribution *not* doing statistical analysis? How do you do statstical analysis on something you don’t know – i.e. error?

Have you tried clicking on the links to see who might have written them?”

it doesn’t matter who wrote them. What they wrote is what matters. And what they wrote indicate they are blackboard mathematicians and statisticians that have no understanding of metrology. I’m not surprised – you have yet to provide a reference to a university bachelor level mathematics or statistics textbook that addresses measurement uncertainty let alone a syllabus that gives a metrology tome as a course textbook. Most of the training I got was in engineering courses and even they didn’t treat metrology as a course subject. Much of it was in lab work where “close enough” was acceptable – but it was hardly ever explained why it was acceptable. It was usually because the university couldn’t afford to keep undergraduate lab equipment calibrated to the needed accuracy even if it could provide adequate resolution.

You are a prime example. You’ve obviously had some statistics training. But NONE of it has allowed you to step outside your blackboard bubble to understand uncertainty. You still believe in the outdated concept of “true value” and “error” and even then it is surprising you’ve heard the terms “true value” and “error”.

You consistently return to the concept that “best estimates” are true values and therefore measurement uncertainty can just be ignored or assumed to cancel. A true blackboard meme.

Reply to  Tim Gorman
August 18, 2026 1:03 pm

You are a prime example. You’ve obviously had some statistics training. But NONE of it has allowed you to step outside your blackboard bubble to understand uncertainty. You still believe in the outdated concept of “true value” and “error” and even then it is surprising you’ve heard the terms “true value” and “error”.

You consistently return to the concept that “best estimates” are true values and therefore measurement uncertainty can just be ignored or assumed to cancel. A true blackboard meme.

Don’t forget his refusal to acknowledge that uncertainty is quantified by standard deviation, even though the GUM says exactly this.

Reply to  karlomonte
August 18, 2026 3:19 pm

Don’t forget his refusal to acknowledge that uncertainty is quantified by standard deviation

That’s not what I said.

Standard uncertainty is the standard deviation of measurements. What I said is that the uncertainty of the mean is not the standard deviation of the values that went into calculating the mean.

The uncertainty of the mean is the standard deviation of the mean, or the standard error of the mean, or whatever you want to call it.

Reply to  Bellman
August 19, 2026 4:54 am

“What I said is that the uncertainty of the mean is not the standard deviation of the values that went into calculating the mean.

definition of the SEM is [ Population-SD]/sqrt(n)

The population SD is the standard deviation of the values that went into calculating the mean.

The SEM is a property of a SAMPLING DISTRIBUTION. A single sample is not a sampling distribution.

SEM = σ / √n is the definition

SEM = s / √n is an ESTIMATOR that doesn’t even exist if all you have is 1 sample because there is no “s”.

In addition, “s” must be assumed to be iid with the population in order for the estimator to be valid.

So, once again, is the UAH data base

  1. one sample of size n, or
  2. multiple samples of size 1?

If 1. then there is no SEM. Thus the measurement uncertainty propagated from the values used to calculate the mean becomes the measurement uncertainty of the data set.

If 2. then the measurement uncertainty of the data set is the propagated measurement uncertainties of the values used to calculate the mean.

Tell us again, which you choose.

Reply to  Tim Gorman
August 19, 2026 7:30 am

Sigh, we’ve gone over this so many times. You never want to learn, so why keep asking questions?

“The SEM is a property of a SAMPLING DISTRIBUTION. A single sample is not a sampling distribution. ”

You just refuse to understand what a sampling distribution is. It’s why your dismissal of maths is your problem.

A sampling distribution is a probability distribution. Any sample will be taken from that distribution. It existsif you take a single sample. It exists if you have zero samples. You can say that if you take a large number of samples their means will tend towards that distribution. But you do not need a large number of samples for the distribution to exist.

“SEM = σ / √n is the definition”

It is not. It’s an equation you can use to determine the SEM under certain assumptions.

“SEM = s / √n is an ESTIMATOR that doesn’t even exist if all you have is 1 sample because there is no “s”. ”

s is the sample standard deviation. Any sample with more than one elements will have a sampling standard deviation.

“In addition, “s” must be assumed to be iid with the population in order for the estimator to be valid. ”

You don’t understand what iid means, or when it is relevant.

Reply to  Bellman
August 19, 2026 7:36 am

“So, once again, is the UAH data base

one sample of size n, or
multiple samples of size 1?

Once again any UAH average will be based on a single sample. The global monthly average is based on the sample of all readings for that month. The average of a smaller region is based on a smaller sample of the relevant sub population.

But note that UAH is not a random sample. It’s based on systematic values taken around the globe.

Reply to  Bellman
August 19, 2026 7:45 am

But note that UAH is not a random sample. It’s based on systematic values taken around the globe.

Of different quantities!

N = 1!

With all your claimed statistical expertise it is difficult/impossible to understand how you refuse to acknowledge this.

Reply to  Bellman
August 20, 2026 6:31 am

Once again any UAH average will be based on a single sample.”

Meaning there is no SEM telling you how close their average is to the population average.

SEM_sample = s/sqrt(n)

THIS IS THE SAMPLING UNCERTANTY FOR THE MEAN OF THE SAMPLE! It is *NOT* the sampling uncertainty for the mean of the population!

I am not even a professional statistician, just a humble engineer that has worked with measurements my entire life.

And yet I keep finding all kinds of garbage statistics associated with climate science.

  1. measurement uncertainty is always random, Gaussian and cancels,
  2. numbers is just numbers
  3. mid-range values = average
  4. Single samples are always iid with the population (sample SEM = pop SEM)
  5. the average of intensive property values of different things is physically meaningful
  6. equal temperatures mean thermal equilibrium
  7. heat loss/gain is a state value and not a time function
  8. measurement uncertainty in an iterative process does not accumulate (variances don’t add)
  9. Significant figures rules are not important in conveying measurement information
  10. averaging can increase measurement resolution (see 2.)

I am still amazed that any self-respecting physical scientist is *not* a critic of climate science results and conclusions.

Reply to  Tim Gorman
August 20, 2026 7:44 am

Meaning there is no SEM telling you how close their average is to the population average.

SEM_sample = s/sqrt(n)

Again, and again – the uncertainty of a global monthly anomaly is not arrived at by dividing the SD by √N. You keep trying to jump from your misunderstanding of elementary statistics to actual uncertainty analysis of complex models.

THIS IS THE SAMPLING UNCERTANTY FOR THE MEAN OF THE SAMPLE! It is *NOT* the sampling uncertainty for the mean of the population!

Writing nonsense in capital letters doesn’t make it less nonsensical, it just looks like nonsense spouted by a mad man.

measurement uncertainty is always random, Gaussian and cancels,“!

You keep “finding” that because it’s your delusion. Nobody apart from you says that.

The same with your other points – they are just you making up silly catch phrases and assuming it’s what everyone believes.

Reply to  Bellman
August 20, 2026 7:59 am

Again, and again – the uncertainty of a global monthly anomaly is not arrived at by dividing the SD by √N. You keep trying to jump from your misunderstanding of elementary statistics to actual uncertainty analysis of complex models.

There is/are no uncertainty of “global monthly anomal[ies]”!

It is exactly zero!

Just read any UAH post to figure out this latest con job of yours.

Reply to  Bellman
August 20, 2026 9:09 am

Again, and again – the uncertainty of a global monthly anomaly is not arrived at by dividing the SD by √N”

Bullshite! There is no other way to get from an individual data measurement uncertainty in the units digit to one in the thouandths digit other than by making the SEM into the measurement uncertainty by using sqrt(n).

Reply to  Tim Gorman
August 20, 2026 11:56 am

“Bullshite! ”

Then could you provide a reference to a single global temperature set whose uncertainty analysis is simply SD/√N.

Reply to  Bellman
August 21, 2026 4:01 am

UAH does not even measure path loss in its measurements. It estimates a value using a model that only uses the absorption coefficient of oxygen, totally ignoring water vapor absorption. The humidity here over the past 24 hours has ranged from 60% to 95%, a range of about 50%. Meaning path loss seen a large variation resulting in a large measurement uncertainty for any one specific measurement including this location.

There simply isn’t any way this amount of path loss can result in a measurement uncertainty as small as the thousandths digit for lower troposphere temperature.

The only way to make the measurement uncertainty this small is to ignore measurement uncertainty and use the sampling uncertainty, SD/√N as the measurement uncertainty.

The same logic applies to *all* temperature data sets used to calculate an “average” temperature. In order to develop a measurement uncertainty that allows differentiating values as small as the hundredths digit the measurement uncertainty must be in the thousandths digit. Since the measurement uncertainty of the temperature values is at least in the units digit (e.g. +/- 1C), the only way to come up with a measurement uncertainty in the thousandths digit is to ignore the actual measurement uncertainty and use the sampling uncertainty as the measurement uncertainty.

In essence, they *all* use your meme of “all measurement uncertainty is random, Gaussian, and cancels” in order to justify using the sampling uncertianty instead.

Reply to  Bellman
August 20, 2026 9:44 am

The same with your other points – they are just you making up silly catch phrases and assuming it’s what everyone believes.

Let’s get down to brass tacks. Since you believe you are knowledgeable enough to criticize folks we need to see what you do know.

  1. Pick an ASOS station of your own choosing.
  2. Create a representative uncertainty budget for that station.
  3. Choose a month of your own choosing.
  4. Compute an average daily temperature For Tavg.
  5. Compute a combined uncertainty for a daily Tavg.
  6. Compute an average monthly Tmax.
  7. Compute a combined uncertainty for Tmax for the month.
  8. Compute an average monthly Tmin.
  9. Compute a combined uncertainty for Tmin for the month.
  10. Compute an average monthly Tavg.
  11. Compute a combined uncertainty for Tavg for the month.
  12. Compute an average baseline temperature for that month.
  13. Compute a combined uncertainty for the baseline average.
  14. Compute an anomaly for that month.
  15. Compute a combined uncertainty for the anomaly average.

Show your math at each step.

Each average will be computed from a random variable containing the appropriate observations and will consist of a mean and standard deviation.
Example: q = {T₁, T₂, …, Tₖ} where,
q̅ = (1/n)ΣTₖ |(1,k) and,
s²(Tₖ) = [1/(n-1)]Σ(Tₖ – q̅)² |(1,k), and,
s²() = s²(Tₖ)/ n.

I should warn you that an uncertainty budget uses standard uncertainties, predominantly from Type B evaluations which we haven’t even discussed.

PS: you won’t find an uncertainty budget for weather stations on the internet. If you need help, just ask.

Reply to  Bellman
August 20, 2026 4:19 am

You just refuse to understand what a sampling distribution is. It’s why your dismissal of maths is your problem.”

1.wikipedia:
——————
The sampling distribution of a statistic is the distribution of that statistic, considered as a random variable, when derived from a random sample of size n. It may be considered as the distribution of the statistic for all possible samples from the same population of a given sample size. The sampling distribution depends on the underlying distribution of the population, the statistic being considered, the sampling procedure employed, and the sample size used. There is often considerable interest in whether the sampling distribution can be approximated by an asymptotic distribution, which corresponds to the limiting case either as the number of random samples of finite size, taken from an infinite population and used to produce the distribution, tends to infinity, or when just one equally-infinite-size “sample” is taken of that same population.
For example, consider a normal population with mean μ and variance σ. Assume we repeatedly take samples of a given size from this population and calculate the arithmetic mean x¯ for each sample – this statistic is called the sample mean. The distribution of these means, or averages, is called the “sampling distribution of the sample mean”.

———————(bolding mine, tpg)

2.Sampling Distribution: Definition, Formula & Examples – Statistics By Jim
—————————-
A sampling distribution of a statistic is a type of probability distribution created by drawing many random samples of a given size from the same population. These distributions help you understand how a sample statistic varies from sample to sample.

………….

Let’s start with a simple example and move on from there!
Sampling Distribution of the Mean ExampleFor starters, I want you to fully understand the concept of a sampling distribution. So, here’s a simple example!
Imagine you draw a random sample of 10 apples. Then you calculate the mean of that sample as 103 grams. That’s one sample mean from one sample. However, you realize that if you were to draw another sample, you’d obtain a different mean. A third sample would produce yet another mean. And so on.

With this in mind, suppose you decide to collect 50 random samples of the same apple population. Each sample contains 10 apples, and you calculate the mean for each sample.

At this point, you have 50 sample means for apple weights. You plot these sample means in the histogram below to display your sampling distribution of the mean.

…………………

Statisticians refer to the standard deviation for a sampling distribution as the standard error. Because we’re assessing the mean, the variability of that distribution is the standard error of the mean.
—————-(bolding mine, tpg)
3.Sampling Distribution – GeeksforGeeks
——————
Sampling distribution is essential in various aspects of real life, essential in inferential statistics. A sampling distribution represents the probability distribution of a statistic (such as the mean or standard deviation) that is calculated from multiple samples of a population.
——————–(bolding mine, tpg)

I’m pretty sure I know what a sampling distribution is and how it is used to determine the standard error, i.e. the standard deviation of the sample means.

It works when you have MULTIPLE samples because the CLT pushes the sample mean distribution toward being Gaussian, even if the parent distribution is skewed.

Reply to  Tim Gorman
August 20, 2026 6:19 am

“wikipedia”

You do see that your quote agrees with what I’m saying, don’t you?

Of course you don’t. That would require reading for understanding rather than cherry picking something you don’t understand and having to admit you are wrong about something.

What does

considered as a random variable

mean to you, if not that the sampling distribution is considered as a random variable?

What does

for all possible samples from the same population

mean if you think a sampling distribuion means an actual distribution of a set if samples?

Assume we repeatedly take samples of a given size from this population and calculate the arithmetic mean x¯ for each sample

What dies the word “assume” mean to you?

“I’m pretty sure I know what a sampling distribution is and how it is used to determine the standard error, i.e. the standard deviation of the sample means. ”

And I’m pretty sure you are wrong. Not least because it’s a futile exercise to take multiple samples if a fixed size just to figure out how uncertain a single sample is. Usually if you can do that you would just pool the individual samples into one bigger sample with less uncertainty.

You can cut and paste your misunderstood definitions all you like, but I doubt you can find a single source that describes practically taking multiple samples to determine the SEM. The pnlybtines you do that is in a simulation designed to illustrate how the SEM works, or when using Monte Carlo methods.

Any introductory text will explain the usual method is to take 1 sample of a given size. Use the sample SD as an estimate of the population SD, and divide by √N.

Reply to  Bellman
August 20, 2026 6:32 am

From the “Statistics by Jim” article (my emphasis)

As you saw earlier, it’s possible to accurately produce sampling distributions using equations rather than drawing many samples. Hypothesis tests take your sample data and do just that for the test statistic.

Reply to  Bellman
August 20, 2026 8:48 am

“As you saw earlier, it’s possible to accurately produce sampling distributions using equations rather than drawing many samples. Hypothesis tests take your sample data and do just that for the test statistic.”

And do *YOU* and climate science DO THAT?

I can’t find where its done anywher!

Reply to  Bellman
August 20, 2026 1:03 pm

Nothing like making more mud! Hypothesis testing has nothing to do with measurement uncertainty!

Reply to  Bellman
August 20, 2026 7:11 am

 That would require reading for understanding rather than cherry picking something you don’t understand and having to admit you are wrong about something.

Irony alert from the master cherrypicker.

Uncertainty is quantified as standard deviation, not sd over root-N.

But you will never admit this.

Reply to  Bellman
August 20, 2026 8:47 am

“You do see that your quote agrees with what I’m saying, don’t you?”

As usual, you either won’t or can’t read.

wikipedia: “Assume we repeatedly take samples of a given size from this population and calculate the arithmetic mean x¯ for each sample – this statistic is called the sample mean. The distribution of these means, or averages, is called the “sampling distribution of the sample mean”.”

Do the words I bolded above mean NOTHING to you? Or did you just not bother to read them?

“What does

considered as a random variable

mean to you, if not that the sampling distribution is considered as a random variable?”

A random variable NEEDS MORE THAN ONE ENTRY TO QUALIFY AS A VARIABLE.

Meaning you need more than ONE sample mean in order to get a distribution of sample means.

“What does

for all possible samples from the same population

mean if you think a sampling distribuion means an actual distribution of a set if samples?”

Stop CHERRY PICKING. Provide th whole statement!

” It may be considered as the distribution of the statistic for all possible samples from the same population”

It means that if you have enough samples to form a good distribution that *all* possible samples will conform to that distribution!

IT’S THE CLT IN OPERATION! Enough samples will create a distribution of sample means to tends to be Gaussian. Adding more samples will just add more entries into that Gaussian distribution.

You keep confusing the SEM of the population mean with the SEM of a sample mean.

The SEM of a sample mean can only be the same as the SEM of the population mean if the sample is iid. Otherwise the SEM of the sample is just an ESTIMATOR of the population SEM and it may or may not be accurate. It introduces MORE uncertainty to the population measurement uncertainty.

The *TRUE* SEM absolutely requires multiple samples to be taken to get a distribution of sample means that defines the SD (i.e. the uncertainty) of the population mean.

Using the SD of a single sample is only an estimator and is only valid under IID assumption.

And I’m pretty sure you are wrong. Not least because it’s a futile exercise to take multiple samples if a fixed size just to figure out how uncertain a single sample is.”

You aren’t figuring out how uncertain a single sample is. You are figuring out how uncertain the “best estimate” is.

That’s why we keep trying to pin you down. Either the temperature data sets are a single sample or they are multiple samples.

If they are a single sample then the “true” SEM doesn’t exist, only an estimator exists. Unless you assume that the single temperature sample is IID with the total population of all temperatures everywhere on the earth that estimator becomes uncertain.

And there is a second consideration for the SEM from one sample. Is that one sample representative of the population? In reality, temperatures vary widely on the surface over distances as small as 1 km. If we up that to assuming that the temperature on the surface is constant over a 10km x 10km square and we sample every square we would get about 5 x 10^6 samples. That would be our population size. If we only sample 10,000 locations that’s only about 0.2% of the total population. That’s hardly big enough to claim the sample is IID with the population. In fact it is simply not large enough to even claim the sample SD is an accurate representation of the population SEM at all.

If they are multiple samples then the SD of the multiple samples is the “true” SEM.

So what do you and climate science do? Just assume that one small sample is IID with the total population of temperatures on the surface of the earth. And you and they do that by homogenizing and infilling of the temperatures to even get the sample size they have!

Reply to  Tim Gorman
August 20, 2026 9:32 am

“Do the words I bolded above mean NOTHING to you?”

Strangely you didn’t bold the word “assume”. Di you only notice it if it’s written ASSume?

“A random variable NEEDS MORE THAN ONE ENTRY TO QUALIFY AS A VARIABLE. ”

Learn what random variable means. The random variable in this case is continuous. It has an infinite number of values. The mean of you sample is just one of those values.

“It means that if you have enough samples to form a good distribution that *all* possible samples will conform to that distribution!”

Do you enjoy displaying your ignorance?

At best your “enough” samples can be used to estimate the distribution. But they can never be considered as all possible samples from the distribution.

“IT’S THE CLT IN OPERATION! Enough samples will create a distribution of sample means to tends to be Gaussian.”

No it isn’t. In future consider why you feel the need to write in all caps. It’s probably a good clue that you are wrong.

The CLT is about sample size, not the number of samples. If you roll two dice 1000000 times, the distribution of their mean will still not be normal. If you roll 1000000 dice once you know that the average will have come from something very close to a normal distribution.

“You keep confusing the SEM of the population mean with the SEM of a sample mean.”

There is no such thing as the SEM of a population mean. You are the one confusing yourself.

“The SEM of a sample mean can only be the same as the SEM of the population mean if the sample is iid. ”

There is no such thing as the SEM of the population, and you don’t understand what iid means.

“Otherwise the SEM of the sample is just an ESTIMATOR of the population SEM and it may or may not be accurate.”

Learn what these words mean. Don’t keep making up your own private language and expect others to understand.

I’ll say this just once in the sligybhope you might try to learn something.

The mean of a sample is an estimate of the population mean.

The SEM is a measure of how much variation there is in the sampling distribution, and hence an indication of how good your sample mean is likely to be as an estimate of the population mean.

It’s meaningless to say that the sample distribution has to be iid with the population distribution.

What matters is if the sample is a reasonable representation of the population. This is usually assumed if he sample is randomly drawn from the population, but will not be correct if there was a bias in your sampling. E.g. if you were more likely to pick big values than small ones.

This is not the same as saying that the distribution of the sample has to be identical to the population. With small sample sizes that will be unlikely. That is covered by the fact that smaller sample sizes will have greater SEMs.

Reply to  Bellman
August 20, 2026 10:34 am

It’s probably a good clue that you are wrong.

Please, just stop.

Reply to  karlomonte
August 20, 2026 10:40 am

Request DENIED.

Reply to  Bellman
August 20, 2026 10:56 am

You can’t even come up with original snark.

Where is your comprehensive uncertainty analysis of complex global air temperature anomalies?

Reply to  Bellman
August 20, 2026 10:51 am

tpg: ““A random variable NEEDS MORE THAN ONE ENTRY TO QUALIFY AS A VARIABLE. ”

Learn what random variable means. The random variable in this case is continuous. It has an infinite number of values.”

ROFL!! Infinite is not greater than 1!!!!!

tpg: ““IT’S THE CLT IN OPERATION! Enough samples will create a distribution of sample means to tends to be Gaussian.”

No it isn’t.”

ROFL!!

wikipedia: “In probability theory, the central limit theorem (CLT) states that, under appropriate conditions, the distribution of a normalized version of the sample mean converges to a standard normal distribution. This holds even if the original variables themselves are not normally distributed.”

The CLT is about sample size, not the number of samples.”

ROFL!

Multiple samples generate multiple sample means. When those multiple sample means are combined into a distribution THEY DEFINE THE SAMPLE SIZE OF THE SAMPLE MEANS!

If you roll two dice 1000000 times”

If I sample the results enough times and form a distribution of the sample means the CLT says that distribution of those sample means will tend toward Gaussian.

Look at wikepedia’s definition again: “This holds even if the original variables themselves are not normally distributed.”

There is no such thing as the SEM of a population mean.”

There is when that population mean is estimated FROM SAMPLING!

Again:
———————–
Standard Error of the Mean (SEM) for a Population MeanThe standard error of the mean (SEM) is a measure of how much the sample mean is expected to vary from the true population mean when you take repeated random samples of the same size from the population Statistics by Jim+1.
Definition and Purpose

  • SEM quantifies the precision of the sample mean as an estimate of the population mean.
  • It is the standard deviation of the sampling distribution of the mean — that is, the distribution of all possible sample means you could get from repeated sampling Wikipedia+1.
  • A smaller SEM means the sample mean is likely to be close to the true population mean, indicating higher precision Biology for Life+1.

——————

The mean of a sample is an estimate of the population mean.”

The mean of a single sample is the mean of that sample. It is only a good estimator of the population mean if it is IID with the population – which is *NOT* the case for most real world measurements.

The SD of that single sample can *NOT* tell you how precisely you have located the population mean if it isn’t a good estimator of the population mean!

The standard deviation of the distribution formed from means of multiple samples *can* tell you how precisely you have located the population mean because the mean of those multiple sample means is formed by the CLT.

and hence an indication of how good your sample mean is likely to be as an estimate of the population mean.”

The SEM using the SD of a single sample is *NOT* a good estimator. It is a crutch used when more samples can’t be done.

Take monthly data where your sample size is only 30 measurements. The 95% confidence intervals for the ratio of s/σ is SD_error= +/-20% and SEM_error = +/- 20%.

Think about what that means for estimating monthly temperature average. It could be off by 20% or more.

Reply to  Bellman
August 20, 2026 10:55 am

It’s meaningless to say that the sample distribution has to be iid with the population distribution.”

If your single sample is *NOT* IID, then just how good are the estimators for the population mean and standard deviation?

Reply to  Tim Gorman
August 20, 2026 5:10 pm

ROFL!! Infinite is not greater than 1!!!!!

I’m worried I’ve broken Tim.

“wikipedia”

Why don’t you read the second paragraph. It’s not easy to cut and paste but it says that the distribution tends to normal as n tends to infinity, where n is the sample size.

Multiple samples generate multiple sample means. When those multiple sample means are combined into a distribution THEY DEFINE THE SAMPLE SIZE OF THE SAMPLE MEANS!

Completely clueless.

If I sample the results enough times and form a distribution of the sample means the CLT says that distribution of those sample means will tend toward Gaussian.

No it does not. But feel free to try it out. In fact the wiki page demonstrates just that under the section “Applications and examples”.

https://en.wikipedia.org/wiki/Central_limit_theorem#/media/File:Dice_sum_central_limit_theorem.svg

“Look at wikepedia’s definition again: “This holds even if the original variables themselves are not normally distributed.””

Yes, that’s the whole point. But obviously you don’t understand what it means.

Again

You just refuse to accept that this is only describing what the sampling distribution means, it is not saying you literally have to take an infinite number of samples in order for the sampling distribution to exist. Here’s the bit you keep missing from Jim on Statistics

Fortunately, you don’t need to repeat your study an insane number of times to obtain the standard error of the mean. Statisticians know how to estimate the properties of sampling distributions mathematically, as you’ll see later in this post. Consequently, you can assess the precision of your sample estimates without performing the repeated sampling.

The mean of a single sample is the mean of that sample.

You don’t say. And that mean of the sample is also an estimate of the population mean.

It is only a good estimator of the population mean if it is IID with the population

Did I mention you don;t understand what iid means? You’ve had plenty of time to explain what you think iid means, and what you actually mean the sample being iid with the population. I’ve even tried to help you out by suggesting that what you really mean is the sample has to be representative of the population – i.e. not biased. But it just goes in one ear and out the other.

The standard deviation of the distribution formed from means of multiple samples *can* tell you how precisely you have located the population mean

Of course it can, but what’s the point when you could just combine all those samples into one big one, and get a much more precise estimate?

because the mean of those multiple sample means is formed by the CLT.”

Still wrong. Take the thousands of a samples of a pair of dice. The distribution will not be normal.

Reply to  Bellman
August 21, 2026 6:19 am

Why don’t you read the second paragraph. It’s not easy to cut and paste but it says that the distribution tends to normal as n tends to infinity, where n is the sample size.”

Which means that as you take more samples the distribution of the sample means tends to normal!

tpg: “Multiple samples generate multiple sample means. When those multiple sample means are combined into a distribution THEY DEFINE THE SAMPLE SIZE OF THE SAMPLE MEANS!

Completely clueless.

As usual you have nothing but the argumentative fallacy of Argument by Dismissal.

More sample means push the distribution of sample means toward Gaussian. A truth you can’t refute so you just dismiss it out of hand. WILLFUL IGNORANCE IS THE WORST KIND OF IGNORANCE.

No it does not. But feel free to try it out. In fact the wiki page demonstrates just that under the section “Applications and examples”.”

The link you gave says this: “Comparison of probability density functions p(k) for the sum of n fair 6-sided dice to show their convergence to a normal distribution with increasing n, in accordance to the central limit theorem”

This is EXACTLY what having more sample means does!

Once again, you are cherry picking something without understanding the meaning and context. Probably because you can’t read!

You just refuse to accept that this is only describing what the sampling distribution means, it is not saying you literally have to take an infinite number of samples in order for the sampling distribution to exist.”

NO ONE HAS SAID YOU NEED INFINITE SAMPLES IN ORDEER TO HAVE A SAMPLING DISTRIBUTION! Least of all me.

 Here’s the bit you keep missing from Jim on Statistics”

Here is the bit *YOU* keep missing! Because you can’t read!

“Statisticians know how to estimate the properties of sampling distributions mathematically” (bolding mine, tpg)

Estimates are UNCERTAIN. They ADD to the uncertainty, they don’t minimize it! Something you refuse to admit when you assert that one sample of a population is all you need!

You don’t say. And that mean of the sample is also an estimate of the population mean.”

And the uncertainty of that single sample is 20% or more with only 30 entries in the distribution. HUGE!

tpg: “It is only a good estimator of the population mean if it is IID with the population””

Did I mention you don;t understand what iid means?”

You’ve been given the definition of IID MULTIPLE times by me. You have yet to refute the definition.

IID:

  1. data components are independent
  2. same expected mean
  3. same variance
  4. same range
  5. same distribution shape

If you want to take one sample and expect it to properly represent the population then that sample has to be IID with the population: Same expected mean, same variance, same range, same distribution shape.

what you really mean is the sample has to be representative of the population”

No, it *has* to be IID. Period. One sample from an unknown parent does *NOT* contain enough information to even judge if it is IID with the unknown parent. Multiple samples do a much better job of identifying the statistical descriptors of an unknown population because of the CLT and because they make it easier to identify non-IID.

The *real* issue here is that you simply cannot confirm that any specific sample is IID. You *can* use some tests to exclude the sample from being IID but you can never confirm IID. Larger sample sizes make it easier to detect non-IID but still don’t confirm it. More samples make it easier to identify non-IID.

I know it must be confusing for you but IID does *not* mean that each sample in a group of samples has to have the same mean. But they have to have the same *expected* mean. The SEM, i.e. the standard deviation of the multiple sample means is a metric for how much the sample means deviate from the parent expected mean IF the multiple samples are IID. With a single sample you have no metric for determining how much the sample mean deviates from the expected mean of the parent. Nor is the ability to exclude the sample for being non-IID as easy to accomplish.

Think about this. The daily temperature average is estimated from TWO entries, i.e. a sample of size 2. That’s equivalent of an SEM of 70% or larger meaning the daily average could be from 40F to 80F for a sample average of 60F. And then that gets put into a sample of size 30 for monthly average temperature. Do you understand just how uncertain that monthly average is with data that has that kind of uncertainty?

Reply to  Tim Gorman
August 21, 2026 8:25 am

“Which means that as you take more samples the distribution of the sample means tends to normal!”

I give up. I clearly can’t compete with level of wilful ignorance.

N is the sample size, not the number of samples. The fact that you still don’t get it at this point suggests you never will.

Reply to  Bellman
August 22, 2026 5:15 am

tpg: “Which means that as you take more samples the distribution of the sample means tends to normal!”

bellman: “I give up. I clearly can’t compete with level of wilful ignorance.”

wikipedia:
—————–
In probability theory, the central limit theorem (CLT) states that, under appropriate conditions, the distribution of a normalized version of the sample mean converges to a standard normal distribution. This holds even if the original variables themselves are not normally distributed.”
——————–

wikipedia:
————————
In other words, suppose that a large sample of observations is obtained, each observation being randomly produced in a way that does not depend on the values of the other observations, and the average (arithmetic mean) of the observed values is computed. If this procedure is performed many times, resulting in a collection of observed averages, the central limit theorem says that if the sample size is large enough, the probability distribution of these averages will closely approximate a normal distribution.
——————(bolding mine, tpg)

statisticisfundamentals.com
————————
Take a population of any shape, draw a random sample, and record its average. Do it again, and again, a few thousand times. Plot all those averages. They form a bell curve, even when the original data is lopsided, flat, or full of spikes. That single fact is the Central Limit Theorem, and it is the reason a poll of 1,000 people can describe a country of millions.
__________________________(bolding mine, tpg)

investopedia.com
——————
Definition
The Central Limit Theorem (CLT) describes how sample means from a population, regardless of the population’s distribution, tend to form a normal distribution as the sample size increases.

  • Sample sizes equal to or greater than 30 are often considered sufficient under the central limit theorem.

——————–

stats.libretexts.org
————————

  • Understand the Central Limit Theorem, stating that the distribution of sample means approaches a normal distribution as sample size increases, regardless of the. population’s shape

———————

cuemath.com
———————
Central Limit Theorem says that the probability distribution of arithmetic means of different samples taken from the same population will closely resemble a normal distribution. In general, for the central limit theorem to hold, the sample size should be equal to or greater than 30.
——————–

How many definitions do you require for me to show that my statement is correct?

I can get you more if needed. But it probably won’t help – you either won’t read the definitions or you won’t comprehend what you’ve red.

N is the sample size, not the number of samples. The fact that you still don’t get it at this point suggests you never will.”

“N” is meaningless if you only have one sample! You require MULTIPLE samples with MULTIPLE means in order to create a SAMPLE MEAN DISTRIBUTION. All “N” does is increase how well the SAMPLE MEAN DISTRIBUTION adheres to being a Gaussian distribution.

With just one sample the uncertainty of the estimate of the SEM is MORE THAN 20% regardless of “N”.

Leave it to you and climate science to believe that an average of values in the units digit and having a minimum uncertainty of 20% is accurate to the hundredths digit!

Reply to  Tim Gorman
August 22, 2026 7:18 pm

Not a single one of your quotes says that the CLT is about the number of samples. You still can’t get it into your head that describing a sampling distribution in terms of taking a large number of samples does not mean that the sampling distribution only exists after you have take a large number of samples.

From some of your sources:

and it is the reason a poll of 1,000 people can describe a country of millions

Note, “a poll of 1,000 people” – not hundreds of polls.

tend to form a normal distribution as the sample size increases.

As sample size increases – not as the number of samples increases.

Understand the Central Limit Theorem, stating that the distribution of sample means approaches a normal distribution as sample size increases, regardless of the. population’s shape

As sample size increases – not as the number of samples increases.

In general, for the central limit theorem to hold, the sample size should be equal to or greater than 30.

The sample size should be equal to or greater than 30, nothing about the number of samples. (And not value you should take too seriously)

How many definitions do you require for me to show that my statement is correct?

One would be a start.

All “N” does is increase how well the SAMPLE MEAN DISTRIBUTION adheres to being a Gaussian distribution.

What do you think the CLT says?

With just one sample the uncertainty of the estimate of the SEM is MORE THAN 20% regardless of “N”.

Could you show your workings on that.

Reply to  Tim Gorman
August 20, 2026 5:16 pm

For the last time, what do you mean by IID? A single sample cannot be IID with the population. Each element in the sample is a random variable which should be IID with each other, and the distribution will be the same as the population. That’s what you get if you can truly take values from at random from the population.

Reply to  Bellman
August 20, 2026 6:07 pm

Each element in the sample is a random variable which should be IID with each other, and the distribution will be the same as the population.

Each element in a sample IS NOT a random variable. A sample is a random variable and is a container for multiple elements such as temperature.

A single sample cannot be IID with the population.

Of course it can. That is one reason the type of sampling is important. Sampling can be any number of types, random, stratified, grouped, etc. The closer each sample is to IID shrinks the sample means standard deviation, I.e., the SEM.

Reply to  Bellman
August 21, 2026 6:26 am

A single sample cannot be IID with the population.”

Of course it can be IID with the population. Same expected mean, same variance, etc. The issue is how do you know?

“Each element in the sample is a random variable which should be IID with each other,”

No, they should be IID with the parent distribution, not with each other.

Reply to  Tim Gorman
August 21, 2026 8:20 am

“Same expected mean, same variance, etc.”

Finally we ate getting somewhere. The problem is still that what you are describing is not iid. Like so many other times, you seem to have gotten hold of a term, guessed at it’s meaning, and then refuse to accept you are misusing it.

“The issue is how do you know?”

You can’t know for certain, but you try to do the best you can to minimise having a biased sample.

“No, they should be IID with the parent distribution, not with each other. ”

Do you want me to go through that and explain all the reasons why it’s wrong?

Reply to  Bellman
August 22, 2026 4:57 am

The problem is still that what you are describing is not iid”

It is the very defintion of IID! How many sources must I post to show you that?

1.statisticshowto.com:
——————–
Each distribution has its own characteristics. Let’s say we are looking at a sample of n random variables,
X1, X2,…, Xn. Since they are IID, each variable Xi has the same mean (μ), and variance(σ)2. In equation form, that’s:
E(Xi) = μ ; Var(Xi) = σ2
for all i = 1, 2,…, n.
————————-

2.https://medium.com/data-science-at-microsoft/understanding-identical-and-independent-distribution-iid-in-machine-learning-dc2e98b6609d
————————
When we say that data is identically and independently distributed, we are essentially saying two things:

  1. Identically distributed: Every data point comes from the same probability distribution. This means that the properties of the data points, like the mean, variance, and so on, are the same across all samples.
  2. Independence: Data points do not influence each other. The outcome of one data point provides no information about the outcome of another.

—————–

These are actually more restrictive than the definition I used. The actual mean of each sample does *not* have to be equal to the parent mean, if they were all equal to the parent mean then there would never be a standard deviation of the sample means. All that is required is that the samples have the same EXPECTED value – the mean of the population. This implies that we EXPECT the sample to mimic the parent’s mean but that doesn’t force each sample to have the same mean as the parent.

Your restriction that the samples be pulled from the same parent distribution is not a sufficient requirement. If the parent distribution has any of a number of issues, e.g. autocorrelation, differing variances for each variable in the parent, then the assumption that the samples are IID may be impossible to meet.

This is *exactly* the problem with temperature data. Climate science claims that temperatures in the same region are correlated – meaning they are not independent. Certainly the temperatures from the NH and the SH have different variances. So do coastal vs inland temperatures. Homogenizing and infilling among stations DECREASES independence.

But climate science IGNORES ALL OF THIS. They just treat the temperature data sets as IID with the parent global temperature field when IID is an impossible restriction to meet!

When it comes to climate science it’s just one garbage assumption after another. And here you are trying to defend it!

Reply to  Tim Gorman
August 22, 2026 6:56 am

“It is the very defintion of IID! ”

It might be the Gorman definition. It certainly isn’t the definition.

In probability theory and statistics, a collection of random variables is independent and identically distributed (i.i.d., iid, or IID) if each random variable has the same probability distribution as the others and all are mutually independent.[1] IID was first defined in statistics and finds application in many fields, such as data mining and signal processing.

https://en.wikipedia.org/wiki/Independent_and_identically_distributed_random_variables

To claim a sample is iid with the population you have to explain how a population is a random variable, explain how a single sample is a random variable, explain how the two are independent, and then explain how the two random variable have the same distribution.

Reply to  Bellman
August 22, 2026 9:39 am

To claim a sample is iid with the population you have to explain how a population is a random variable, explain how a single sample is a random variable, explain how the two are independent, and then explain how the two random variable have the same distribution.”

You have to do FAR more than that. I listed out some (not all, just some) of the restrictions on the parent distribution that prevents samples from being IID.

If the values in the parent are random and independent and the sampling method provides random selection of values from the parent then it can be assumed that the values in the sample are independent.

The issue is that the temperature values in the temperature data sets are NOT independent if, as climate science says, temperatures in one location are related to surrounding temperatures so that you can homogenize and infill temperatures that are missing by using temperatures from surrounding measurements. Doing so means they are correlated and a data set using those violate IID rules.

Not only are they spatially correlated but they are temporally correlated, both daily and seasonally.

If temperatures at two locations are correlated then:

  1. the effective sample size is smaller than n
  2. the SEM is underestimated by that smaller n thus the mean is less accurate than σ/ √n indicates
  3. The CLT converges much more slowly thus a large number of samples will be required to obtain a Gaussian sampling distribution of the mean
  4. the student-t distribution is not valid

The more one looks at the global average temperature generated by climate science the larger the number of garbage assumptions that will be found.

Reply to  Tim Gorman
August 22, 2026 9:44 am

“You have to do FAR more than that.”

Could you start with what I said, or better just accept you mis-used a term you didn’t understand.

Reply to  Bellman
August 22, 2026 10:18 am

I already told you what some of them are. Like usual, you didn’t bother reading for comprehension.

Reply to  Tim Gorman
August 22, 2026 10:27 am

Then spell them out here, rather than expect me to travel through your 2000 hysterical rants on the off chance that you hid your explaination as to how a population was a random variable in one.

Reply to  Bellman
August 22, 2026 12:57 pm

I am *NOT* your puppet to dance at the end of your strings. They are upthread. *YOU* go read them.

Reply to  Tim Gorman
August 20, 2026 3:44 pm

A random variable NEEDS MORE THAN ONE ENTRY TO QUALIFY AS A VARIABLE.

Actually a random variable can have just one entry. The variable part comes from the fact that a random variable is a “container” that holds something. It can have discreet values from −∞ -> or it can have a continuous function within an interval, even −∞ -> ∞. The function is usually a probability distribution function.

Reply to  Jim Gorman
August 21, 2026 4:55 am

Actually a random variable can have just one entry.”

A variable with only one entry is a constant! It contributes nothing to the slope of a function.

if “x” is a random variable and x = 1, then a graph of y = x is a horizontal line with a slope of 0. I.e. dy/dx = 0. Same for x = 100.

Reply to  Tim Gorman
August 22, 2026 8:02 am

A random variable’s range can be a single value, but only if that value is the only possible outcome in the experiment. And you are correct, it is a constant over the entire domain. The probability mass function (PMF) assigns probability 1 to that value and 0 to all others. In essence, a uniform distribution. It is why that single value is divided by the √3 to obtain a normalized standard uncertainty value expressed as a standard deviation.

That agrees with your slope of zero. That is, no actual calculatable standard deviation.

Reply to  Tim Gorman
August 20, 2026 9:52 am

One little change.

calculated from multiple samples, of a given size,of a population.

The given size is important for the calculation of the SEM from the standard deviation of the population.

Reply to  Bellman
August 20, 2026 4:55 am

Any sample will be taken from that distribution.”

No, the samples are drawn from the parent distribution.

The distribution of the sample means doesn’t exist if you have only one sample. There is no standard deviation of the sample means if you only have one sample.

s is the sample standard deviation.”

It is the standard deviation of the sample means distribution. It is *NOT* the standard deviation of one sample.

You don’t understand what iid means, or when it is relevant.”

How many different definitions from how many different sources will you need for you to understand how wrong your assertion here is?

I. Statistics by Jim:
The standard error of the mean is the variability of sample means in a sampling distribution of means.”

2.completeera.com
The **sampling distribution of the mean** is a **hypothetical distribution** of all possible **sample means** you could obtain from samples of size **n** taken from a population. The SEM is the **standard deviation of this distribution**.”

What you keep wanting to call the standard error of the mean is actually THE STANDARD ERROR OF THE SAMPLE.

The standard error of the sample is *NOT* the standard error of the mean which is the standard deviation of the means accumulated from multiple samples.

The standard error of a single sample will *NOT* tell you how close you are to the population mean unless you can show the sample distribution is iid with the parent distribution.

Using your logic there would NEVER be a reason to EVER take more than one sample.

Reply to  Tim Gorman
August 20, 2026 7:32 am

No, the samples are drawn from the parent distribution.

You are just determined to misunderstand every point aren’t you.

Let me rephrase the point. The mean of an individual sample is taken from the sampling distribution of the mean. The standard deviation of the sample is taken form the sampling distribution of the standard debilitation. The same with any property of the sample.

The distribution of the sample means doesn’t exist if you have only one sample.

What’s the point? I explain multiple times why that’s wrong. You just ignore all my explanations and just repeat your mistake.

How many different definitions from how many different sources will you need for you to understand how wrong your assertion here is?

Then tell me what your definition of iid is. Here’s mine

In probability theory and statistics, a collection of random variables is independent and identically distributed (i.i.d., iid, or IID) if each random variable has the same probability distribution as the others and all are mutually independent. IID was first defined in statistics and finds application in many fields, such as data mining and signal processing.

https://en.wikipedia.org/wiki/Independent_and_identically_distributed_random_variables

“The **sampling distribution of the mean** is a **hypothetical distribution** of all possible **sample means**

Do you understand what the word “hypothetical” means?

“What you keep wanting to call the standard error of the mean is actually THE STANDARD ERROR OF THE SAMPLE.

No it is not. Standard error just means the standard deviation of a property of a sample in it’s sampling distribution. The Standard Error of the Mean refers to the specific sampling distribution for the mean of the sample. Saying the “Standard Error of the Sample” is meaningless.

The standard error of a single sample will *NOT* tell you how close you are to the population mean …

Nothing will tell you how close you are tot he population mean, unless you actually know what the population mean is. The whole point of the SEM is to yell you how close your sample mean is likely is to be to the population mean.

unless you can show the sample distribution is iid with the parent distribution.

I’m sure you think you are making a point, and maybe it’s a reasonable one – but it gets lost if you keep misusing the same term and refusing to learn what it means.

If you are trying to say that your sample will be no good if it’s biased, then you’re correct. If that’s not what you are trying to say then you are probably misunderstanding something.

Using your logic there would NEVER be a reason to EVER take more than one sample.

By my logic, you mean the same as used in the GUM, Taylor, NIST and every other source you keep pointing me to.

Of course there isn’t normally a need to more than one sample, that’s why you generally don’t take more than one sample. The only reason I can see for taking more than one sample is to pool multiple sources in to one meta sample. There’s also methods like bootstrapping, where you simulate taking multiple samples by resampling your single sample.

But there’s no point in taking a large sample and splitting it into smaller sub-samples just to determine the SEM for your individual sub-sample size.

Reply to  Bellman
August 20, 2026 7:38 am

Let me rephrase the point. The mean of an individual sample is taken from the sampling distribution of the mean. The standard deviation of the sample is taken form the sampling distribution of the standard debilitation. The same with any property of the sample.

And none of this is uncertainty.

I’m sure you think you are making a point, and maybe it’s a reasonable one – but it gets lost if you keep misusing the same term and refusing to learn what it means.

From someone who still doesn’t understand that uncertainty is not error.

Reply to  karlomonte
August 20, 2026 9:51 am

I have asked bellman twice if he would revise his “best estimate” of the value of the measurand based on his calculation of the SEM – I’ve not seen an answer yet!

Reply to  Tim Gorman
August 20, 2026 12:24 pm

“I’ve not seen an answer yet!”

This pathetic crap again. You throw up 50 rambling all caps comments a day, and then claim that if I failed to spot one of your idiotic questions, you claim I couldn’t answer and do a little victory dance.

I’m not even sure what your question is meant to mean? Why would the best estimate change based on a different SEM? If you mean you have two different experiments with different SEMs then yes, you would give more weight to the more certain result. A result based on a sample size of 1000 would be stronger than one based on a sample of size 10 – everything else being equal.

Reply to  Bellman
August 21, 2026 4:21 am

I’m not even sure what your question is meant to mean? Why would the best estimate change based on a different SEM? If you mean you have two different experiments with different SEMs then yes, you would give more weight to the more certain result. A result based on a sample size of 1000 would be stronger than one based on a sample of size 10 – everything else being equal.”

Would you change your best estimate of the value of a measurand based on the SEM of the values in your data set?

A simple question. It has nothing to do comparing distributions from different objects with differently sized sample entries.

you would give more weight to the more certain result. A result based on a sample size of 1000 would be stronger than one based on a sample of size 10 – everything else being equal.”

Since variance is a direct metric for the uncertainty of the average value why aren’t the entries used to calculate the global temperature average weighted based on their variances?

A result from a distribution with a smaller variance would be “stronger” than one from a distribution having a larger variance.

In the past you’ve asserted that weighting temperature values based on the variance of the data used to determine the value, i.e. “more certain result”, IS NOT NECESSARY.

So, yes or no, should the average values of the various temperature data sets be calculated from weighted values or is weighting unnecessary?

Reply to  Tim Gorman
August 21, 2026 5:20 am

“Would you change your best estimate of the value of a measurand based on the SEM of the values in your data set?”

When I point out that it isn’t clear what your question is asking, just repeating it isn’t very helpful.

The obvious answer is no. Why on earth would your best estimate change? The SEM is about the uncertainty of your best estimate, not it’s value.

But it’s such a pointless question I assume you have some different question in mind.

“Since variance is a direct metric for the uncertainty of the average value why aren’t the entries used to calculate the global temperature average weighted based on their variances?”

Because, as I keep having to explain to you, the variance in temperatures is not generally done to measurement uncertainty. It’s caused by temperatures actually varying. And because you are not averaging the same temperature.

Variance weighting works because you are measuring the same thing with different instruments and want to give more weight to the more precise instruments.

“In the past you’ve asserted that weighting temperature values based on the variance of the data used to determine the value, i.e. “more certain result”, IS NOT NECESSARY.”

It isn’t that it’s not necessary, it’s that it would be wrong. If you measure the height fva child with a precise instrument and the height of a tall adult with an imprecise instrument, it would make no sense to say their average was closer to that of the child.

It makes no sense to say the average temperature of the earth is hotter, just because the warmest parts of the earth also have the least variance.

If you do want to factor in variance, then you could do what I did a little while ago, and work out the average standard deviation of temperatures. In that way somewhere in the tropics that was 1°C above average would have a larger SD score than somewhere in the far north that was 1°C above average. It’s a measure of how unusual the temperature is across the globe, bit doesn’t claim to actually change the average temperature.

“So, yes or no, should the average values of the various temperature data sets be calculated from weighted values or is weighting unnecessary?”

No. But that wasn’t the question you asked.

Reply to  Bellman
August 21, 2026 7:32 am

When I point out that it isn’t clear what your question is asking, just repeating it isn’t very helpful.”

The question is exact and clear.

The obvious answer is no. Why on earth would your best estimate change? The SEM is about the uncertainty of your best estimate, not it’s value.”

Isn’t the uncertainty of the best estimate specified in the measurement uncertainty interval – the interval of reasonable values that can be attributed to the measurand?

What does knowing the SEM provide? Does it expand the uncertainty interval? Does it cause a change in the best estimate?

If the SEM is significant compared to the measurement uncertainty, then what does that tell you about your measurement?

Because, as I keep having to explain to you, the variance in temperatures is not generally done to measurement uncertainty. It’s caused by temperatures actually varying. And because you are not averaging the same temperature.”

What does this mean? The measurement uncertainty of a temperature is equivalent to a variance. Why then shouldn’t temperatures used to calculate an average be weighted based on their variance/measurement uncertainty. Taylor says that measurements with lower measurement uncertainty/variance should be weighted so the more accurate measurement is given more weight. See his Chapter 7.

Are you saying the global average temperature is *NOT* an average of temperatures?

Are you saying that measurement uncertainty is not related to variance?

Variance weighting works because you are measuring the same thing with different instruments and want to give more weight to the more precise instruments.”

So you shouldn’t give more accurate measurements of different things more weight in their average?

If I am measuring three 2″x4″ boards and get

8′ +/- 1″
7.95′ +/- 2″
8.1′ +/- 2″

that I shouldn’t give the 8′ +/- 1″ measurement more weight when calculating the average length?

This logic is nothing more than a cousin to your ingrained meme of “all measurement uncertainty is random, Gaussian, and cancels”.

It makes no sense to say the average temperature of the earth is hotter, just because the warmest parts of the earth also have the least variance.”

The issue isn’t making the earth “hotter”. The issue is making the average temperature MORE ACCURATE!

 It’s a measure of how unusual the temperature is across the globe”

Unfreakingbelievable. Standard deviation is the square root of variance. Variance is a metric for uncertainty. The GUM *names* it as a variance.

All you are doing here is justifying why measurement uncertainty/variance should just be ignored and use the average of the stated values as 100% accurate!

“No”

Accuracy be damned!

Reply to  Tim Gorman
August 21, 2026 7:44 am

“Because, as I keep having to explain to you, the variance in temperatures is not generally done to measurement uncertainty. It’s caused by temperatures actually varying. And because you are not averaging the same temperature.”

What does this mean? The measurement uncertainty of a temperature is equivalent to a variance.

Just what I was trying to figure out, then gave up.

Reply to  Tim Gorman
August 21, 2026 9:27 am

The question was

“if he would revise his “best estimate” of the value of the measurand based on his calculation of the SEM”

I’ve given you the clear answer, no. But pointed out it depends on exactly what’s going on in your head, which is something I try to avoid thinking about. Every question you asked is based on your own definitions and misunderstandings.

“Isn’t the uncertainty of the best estimate specified in the measurement uncertainty interval – the interval of reasonable values that can be attributed to the measurand?”

You didn’t ask about the uncertainty, you asked about the best estimate.

“What does knowing the SEM provide? Does it expand the uncertainty interval? Does it cause a change in the best estimate?”

The uncertainty of the mean. Yes. No.

“If the SEM is significant compared to the measurement uncertainty, then what does that tell you about your measurement?”

You’ve now changed the framing of the question from a random sample to a measurement. And as you have odd ideas about what measurement uncertainty means, I’m not sure it’s worth answering knowing you will twist the meaning.

Could you provide a specific example?

“What does this mean? ”

It means my phone auto corrected “down” to “done”.

The rest seems clear enough. Some places are hot one month and cold the next, not because your measurements have a large uncertainty, but simply because it was actually hot one month and cold the next.

“The measurement uncertainty of a temperature is equivalent to a variance.”

No, it’s equivalent to the standard deviation. But as you say, that’s a temperature, not multiple temperatures. You should know this. You keep going on about repeatability and how the rules of a Type A uncertainty only applies to multiple measurements if the same thing.

“If I am measuring three 2″x4″ boards and get

8′ +/- 1″
7.95′ +/- 2″
8.1′ +/- 2″”

The clue is in the fact that the boards are all supposed to be the same length. Now suppose you wanted the average of three different length of board.

2.02 ± 0.01m
4.05 ± 0.02m
4.56 ± 0.02m

Weighing by the uncertainty will bias the average towards the shorter plank, just because it has slightly more certainty in it’s length.

“The issue is making the average temperature MORE ACCURATE! ”

But it won’t be more accurate. That’s my point.

“Unfreakingbelievable.”

You’ve clearly lost all sense if the plot at this point, and I can’t be bothered to correct you.

Reply to  Bellman
August 22, 2026 4:13 am

I’ve given you the clear answer, no.”

Then what good is the SEM? If it is encompassed in the measurement uncertainty interval what else do you need to know concerning the measurement?

“You didn’t ask about the uncertainty, you asked about the best estimate.”

You simply can’t read English. I’ll ask again more simply:

If the uncertainty of the Best Estimate is included in the interval of reasonable values that can be attributed to the value of the measurand then what use is the SEM?

The uncertainty of the mean. Yes. No.”

“Yes.” -> The SEM increases the measurement uncertainty interval? How does it do that? Is the SEM typically larger than the SD of a distribution? How can that be if the SEM = SD/sqrt(n)? Can “n” be a fraction?

“No.” -> Then what use is the SEM? Stop with saying it is the uncertainty of the mean. That isn’t defining the USE to which it can be put! If it doesn’t change the value of the Best Estimate and if it is less than the measurement uncertainty interval then what can it be used for?

You’ve now changed the framing of the question from a random sample to a measurement.”

Bullshite! Ten measurements of the same thing under repeatability conditions *IS* supposed to generate a random distribution of values! It is a random sample of the value of the measurand AND IT IS A SET OF MEASUREMENTS!

And as you have odd ideas about what measurement uncertainty means, I’m not sure it’s worth answering knowing you will twist the meaning.”

What a load of crappola! The only one in this sub-thread that has odd ideas about what measurement uncertainty is, IS YOU!

“Could you provide a specific example?”

You’ve already asserted this:

bellman: “the uncertainty caused by the randomness of the sample will usually be much greater than the uncertainty caused by the uncertainty in the individual measurements””

When called on it you said that wasn’t what you meant but you never clarified what you meant.

The issue is that how do you create an example where SD/sqrt(n) > SD? Which is what you are claiming!

Keep dancing!

The rest seems clear enough. Some places are hot one month and cold the next, not because your measurements have a large uncertainty, but simply because it was actually hot one month and cold the next.”

So what? The monthly averages get put into a distribution of 12 entries for the purpose of determining an annual average. The span of the monthly averages stated values creates the variance of the distribution stated values! BUT the measurement uncertainty associated with that annual average is the sum of the uncertainties associated with each monthly average. It is *NOT* the square root of the variance of the stated values divided by 12.

Unless, that is, it is you and climate science doing the annual average and are applying the meme of “all measurement uncertainty is random, Gaussian, and cancels”. Then the measurement uncertainty of the annual average becomes [sqrt(variance of the stated values)] / sqrt(12).

No, it’s equivalent to the standard deviation.”

It is calculated using VARIANCES!

From the GUM:
-“The experimental variance of the observations, which estimates the variance σ^2 of the probability distribution of q, is given by”
-“The experimental variance of the mean s^2(q_bar) and the experimental standard deviation of the mean s(q_bar) (B.2.17, Note 2), equal to the positive square root of s^2(q_bar), quantify how well q estimates the expectation μ_q of q, and either may be used as a measure of the uncertainty of q .” (bolding mine, tpg)

You calculate the VARIANCE first, then find the standard deviation!

READ THE DAMN GUM SOMETIME INSTEAD OF JUST CHERRY PICKING FROM IT!

Reply to  Bellman
August 22, 2026 4:26 am

Weighing by the uncertainty will bias the average towards the shorter plank, just because it has slightly more certainty in it’s length.”

So what? If you are going to try and find the sum of the lengths of the board then the most accurate measurement *should* be given extra weight! And you have to do the sum in order to find teh average!

But it won’t be more accurate. That’s my point.”

If you weight the measurements based on accuracy, then the mean WILL be more accurate!

Exactly what do you think you are doing when you add the measurement uncertainties to find the measurement uncertainty of the average value of different things? The most accurate measurement contributes less to the total measurement uncertainty. That’s called WEIGHTING!

All you are doing here is throwing more crap against the wall hoping some of it will stick. Do you *ever* wonder why everything you assert is *always* crap that doesn’t stick to the wall?

Reply to  Tim Gorman
August 22, 2026 7:37 pm

If you weight the measurements based on accuracy, then the mean WILL be more accurate!

I have to lumps of wood. One is exactly 1m long, the other 2m long. The exact average is 1.5m.

Now I measure them. The first one with an instrument that has an uncertainty of 1mm the other an uncertainty of 1cm.

Lets say my measurements are 1.001 ± 0.001 and the other is 1.99 ± 0.01.

I use inverse variance weighting to average them, and the average is 1.011 ± 0.001m.

If I average them without variance weighting the result is 1.50 ± 0.01m.

Which of those methods is most accurate?

Reply to  Bellman
August 21, 2026 10:09 am

Because, as I keep having to explain to you, the variance in temperatures is not generally done to measurement uncertainty. It’s caused by temperatures actually varying. And because you are not averaging the same temperature.

Here are two things from the GUM you need to read until you understand their implications.

B.2.16

reproducibility (of results of measurements)

closeness of the agreement between the results of measurements of the same measurand carried out under changed conditions of measurement

NOTE 3 Reproducibility may be expressed quantitatively in terms of the dispersion characteristics of the results.

F.1.1.2 It must first be asked, “To what extent are the repeated observations completely independent repetitions of the measurement procedure?” If all of the observations are on a single sample, and if sampling is part of the measurement procedure because the measurand is the property of a material (as opposed to the property of a given specimen of the material), then the observations have not been independently repeated; an evaluation of a component of variance arising from possible differences among samples must be added to the observed variance of the repeated observations made on the single sample.

Let’s break this down.

Reproducibility is a quantitative measure of the dispersion of the results. If a monthly average is a measurand, the dispersion of the results is a quantitative measure of the dispersion.

Repeated observations of a single sample of a monthly temperature are not available. That means a Type B uncertainty must be used and an uncertainty budget is necessary to determine its value. In addition, to the uncertainty of that single sample, one must add the variance arising from possible differences among other samples.

You need to explain why these are not used in determining a monthly average temperature.

Reply to  Jim Gorman
August 21, 2026 1:05 pm

You need to explain why these are not used in determining a monthly average temperature.

Good luck ever seeing this from the trendologists.

Reply to  Bellman
August 20, 2026 9:22 am

The mean of an individual sample is taken from the sampling distribution of the mean”

The mean of an individual sample is taken from the SAMPLE DATA in that individual sample. There is no sampling distribution defined for ONE SAMPLE.

The standard deviation of the sample is taken form the sampling distribution of the standard debilitation”

The standard deviation of the individual sample is taken from the SAMPLE DATA in that individual sample. There is no sampling distribution defined for one sample.

“What’s the point? I explain multiple times why that’s wrong. You just ignore all my explanations and just repeat your mistake”

You haven’t explained ANYTHING, not once. You just keep making the same assertions over and over. You’ve not provided one single reference or math derivation showing that you can get a distribution of sample means from one sample!

You keep tyring to play with words. Trying to equate “distribution of the data in a sample” to being a “sample distribution”.

Again, from wikepedia:
——————————-
For example, consider a normal population with mean μ and variance σ. Assume we repeatedly take samples of a given size from this population and calculate the arithmetic mean x¯ for each sample – this statistic is called the sample mean. The distribution of these means, or averages, is called the “sampling distribution of the sample mean”. This distribution is normal N(μ,σ2/n) (n is the sample size) since the underlying population is normal, although sampling distributions may be close to normal even when the population distribution is not (see central limit theorem).
—————(bolding mine, tpg)

You must have multiple samples in order to calculate multiple means in order to form a “sample distribution of the sample mean”.

The distribution of data in a single sample is *NOT* the “sampling distribution”. It is just a “sample distribution”.

Reply to  Bellman
August 20, 2026 9:33 am

“same probability distribution”

There is *NO* guarantee that a single sample will have the same probability distribution as the parent, including the mean, the SD, etc.

No it is not. Standard error just means the standard deviation of a property of a sample in it’s sampling distribution” (bolding mine, tpg)

Your lack of reading comprehension skills is showing again.

Read it again: “ The distribution of these means, or averages, is called the “sampling distribution of the sample mean”.””

I bolded the word “means” for you. “MeanS” is plural, not singular. “a sample” is SINGULAR, not plural.

You do *NOT* get a sampling distribution from a singular value.

Of course there isn’t normally a need to more than one sample”

Really? So there is no need for multiple experiments to confirm experimental results?

If you just do one experiment to find the speed of light and take multiple measurements during the experiment, i.e. one sample, that’s all you will ever need?

Unfreakingbelievable. And you wonder why I keep telling you that you live in Statistical world and not the real world.

Reply to  Tim Gorman
August 20, 2026 6:26 pm

There is *NO* guarantee that a single sample will have the same probability distribution as the parent

It’s pretty much guaranteed it won’t have the same distribution. That’s why you need to estimate the SEM.

Read it again: “ The distribution of these means, or averages, is called the “sampling distribution of the sample mean”.””

I’ve no idea what you think your point is here. Just your usual ranting without reading or understanding. I was correcting your claim that

What you keep wanting to call the standard error of the mean is actually THE STANDARD ERROR OF THE SAMPLE.

Sample mean is not the same thing as sample.

I bolded the word “means” for you. “MeanS” is plural, not singular. “a sample” is SINGULAR, not plural.

So? The sampling distribution of the mean is the probability of the means of all possible samples. Still no idea what you think “THE STANDARD ERROR OF THE SAMPLE” is.

You do *NOT* get a sampling distribution from a singular value.

You don’t “get” it from any value. It’s something that has to be calculated from your knowledge of probability theory.

Really?

Yes, really.

So there is no need for multiple experiments to confirm experimental results?

Not what I said.

If you just do one experiment to find the speed of light and take multiple measurements during the experiment, i.e. one sample, that’s all you will ever need?

Not what I said.

Unfreakingbelievable.

Not really – you do it all the time. Just make up things you think I might have been saying rather than trying to understand what I’m actually saying.

And you wonder why I keep telling you that you live in Statistical world and not the real world.

Is it becasue I understand the statistics, and you don’t?

Reply to  Bellman
August 20, 2026 10:43 am

The mean of an individual sample is taken from the sampling distribution of the mean. 

This makes no sense at all. Yes, the mean of a single individual sample is a point (a sample if you will), on the sample means distribution. However, without additional multiple samples, you have no sample means distribution to draw an inference from.

You have two choices, you can call each station a sample or you can call all the stations together a single sample. You can’t do both.

If each station is a sample, then **n** must equal “1”. If each station is a member of a single sample, then you do not have multiple samples and the definition of σ/√n being the SEM falls apart.

Which is it?

You might want to read Taylor, Section 5.7. He says:

To answer this, we naturally imagine repeating our N

measurements many times; that is, we imagine performing a sequence of experiments, in each of which we make N measurements and compute the average.

A sequence of experiment, each of which have N measurements. In other words, multiple samples each having the size **n**. The complete mathematical derivation also requires each sample to be IID or the derivation fails.

Here is another similar definition from: Sample Mean Distribution: Formula, Standard Error & CLT (2026)

The sampling distribution of the sample mean is the probability distribution formed by the means of all possible random samples of a fixed size n drawn from a population. Its mean equals the population mean (μ = μ) and its standard deviation — called the standard error — equals σ/√n. When n is large enough, this distribution is approximately normal, regardless of the shape of the population distribution. (bold by me)

Reply to  Jim Gorman
August 20, 2026 6:04 pm

This makes no sense at all.

Obviously not to you.

However, without additional multiple samples, you have no sample means distribution to draw an inference from.

I did not say “sample means distribution” whatever that would mean, I said the sampling distribution of the mean. And whilst it’s pointless to keep explaining this, a sampling distribution is a probability distribution.

You have two choices

Hopefully more than that.

you can call each station a sample

A station is not a sample. A set of measurements from a station would be a sample.

“or you can call all the stations together a single sample”

Indeed you can. Although context matters. If you are talking about weather stations you will not have a random sample.

You can’t do both.

You can’t stop me.

If each station is a sample, then **n** must equal “1”.

Only if you take a single measurement. And that’s not going to tell you much, about that station let alone about the globe.

If each station is a member of a single sample, then you do not have multiple samples and the definition of σ/√n being the SEM falls apart.

This dumb act is getting pretty stale. If you really want to pretend that you think the only way to determine the SME of a sample is to take many different samples and look at the standard deviation of their means, then why do you keep talking about σ/√n?

“Which is it?

Which is what? If you want to know the temperature of the earth and you can get a genuine random sample of measurements for a specific period, you can do the standard sampling technique of averaging all the measurements, taking that mean as the estimate of the population mean, and then using the sample standard deviation divided by √n as your estimate of the SEM.

The actual global anomaly calculations are rather more involved, and require much more detailed uncertainty analysis.

You might want to read Taylor, Section 5.7.

Why? Nobody else seems to.

To answer this, we naturally imagine repeating our N

measurements many times; that is, we imagine performing a sequence of experiments, in each of which we make N measurements and compute the average.

What part of the word “imagine” don’t you understand?

Note, he’s using that to prove the result. If you want a practical example of using the SEM (SDOM) look at the example in 4.5. There he takes a sample of 10 measurements of two sides of a rectangle and calculates the SDOM for each. No having to repeat the experiment multiple times in order to find out the SDOM, just a single sample of size 10.

Reply to  Bellman
August 21, 2026 6:56 am

I did not say “sample means distribution” whatever that would mean, I said the sampling distribution of the mean. And whilst it’s pointless to keep explaining this, a sampling distribution is a probability distribution.”

One point does not define a probability distribution unless its probability is 1!

Only if you take a single measurement. And that’s not going to tell you much, about that station let alone about the globe.”

That *IS* how each entry is treated. As a single measurement with a measurement uncertainty!

If you really want to pretend that you think the only way to determine the SME of a sample is to take many different samples and look at the standard deviation of their means, then why do you keep talking about σ/√n?”

σ/√n IS AN ESTIMATE OF THE SEM! Since you don’t know σ, “s” is used instead. The uncertainty of the SEM becomes the uncertainty in “s”.

An estimate of the uncertainty is “s” is given by

RelativeStandardError of s = sqrt[ 1/2(n-1) ].

With n = 5, the relative uncertainty of “s” is approx 30%. Even at n=100 the relative uncertainty of “s” is 7%.

If you have only one sample, there is no σ and there is no “s”. A distribution of sample means with only one entry is not a distribution.

There is simply no way to estimate how far a single sample mean is from the population mean.

The standard deviation of a sample is *NOT* the same thing as the standard deviation of sample means. The standard deviation of a single sample can be very small but the sample mean can be very far from the mean of the population. Without multiple samples to create a distribution of sample means you have no way to estimate how far the sample mean is from the population mean.



Reply to  Tim Gorman
August 21, 2026 8:45 am

“One point does not define a probability distribution unless its probability is 1!”

Still just as clueless. Once again a sample is one point taken from a probability distribution
It does not define the probability distribution. The only thing that defines the probability distribution is the existence of the probability distribution.

If I throw a 6-sided die and it comes up 5, that 5 was a random value taken from the actual probability distribution of the die. It does not define the probability distribution. The probability doesn’t suddenly become a constant 5 with probability 1 as a result of a single roll.

“With n = 5, the relative uncertainty of “s” is approx 30%. Even at n=100 the relative uncertainty of “s” is 7%. ”

Now apply that logic to your sample of sample methods. If you have 100 samples, the SEM estimate will still have an uncertainty of 7%. You’ve just had to spend a lot more in order to get the same amount of uncertainty.

“If you have only one sample, there is no σ and there is no “s”. A distribution of sample means with only one entry is not a distribution. ”

And this is where you are just confusing yourself. Understandable giving how many different standard deviations there are floating about. s is the sample standard deviation. That it it’s the standard deviation if your one sample. It exists and can be calculated in the usual way.

It is not, as you seem to think, the standard deviation of the means of multiple samples. That would be the SEM.

“There is simply no way to estimate how far a single sample mean is from the population mean.”

That’s why you have additional uncertainty . It’s why you use a student distribution and not a Gaussian. It’s also another reason why a larger sample size is better than a small one.

“The standard deviation of a sample is *NOT* the same thing as the standard deviation of sample means. ”

Please try to remember that.

“Without multiple samples to create a distribution of sample means you have no way to estimate how far the sample mean is from the population mean.”

Apart from the method every text book will tell you. Standard deviation of the sample divided by √N.

Reply to  Bellman
August 21, 2026 9:47 am

Once again a sample is one point taken from a probability distribution

A sample of a single point on a PDF has an **n** =1. A single point cannot define a PDF, only multiple samples of size **n** can define a PDF.

If you already know the values of the PDF, why are you sampling it to begin with? Do you think the value of the mean of the PDF is going to change due to sampling? Do you think the SEM is going to better define the mean?

Apart from the method every text book will tell you. Standard deviation of the sample divided by √N.

But, for that to be true, the sample must be IID with the population. Otherwise, you have no guarantee that the statistics from a sample have any meaning at all.

Let’s look at Tmax for a month at a single station. Are the 30 days of Tmax a sample or a population. If it is an entire population, that is, all the possible values, why are we calling it a sample? If it is not a population, what is the population?

Remember, Tmax is the maximum value of a periodic function similar to a sinusoid. It is generally considered to be the mean of Gaussian function rather than a trig function. Is that the use of a Gaussian PDF a correct assumption?

Reply to  Jim Gorman
August 21, 2026 2:44 pm

A sample of a single point on a PDF has an **n** =1.

What I said was a single sample is one point taken from a probability distribution. To be clear, that’s one value of the sample (e.g. the mean).

You keep thinking that a set of data defines a PDF. The PDF in this case is defined by the set of all possible samples of a specific size. It is not defined by any one, or any number of specific samples. It’s the other way round.

If you already know the values of the PDF, why are you sampling it to begin with?

There are a number of possible ways a sampling distribution is used. What we are talking here is the case where you do not know what the distribution is, but are taking a sample in order to infer what the distribution may be. That is the best estimate of the population mean is the sample mean, and the SEM is estimated from the sample standard deviation. This provides the uncertainty for your sample mean.

Another possibility is you already know the sampling distribution, and you use it to figure out the probability of getting a sample with a specific mean.

And a third option is that you assume a sampling distribution, and then see if your sample is likely to have come from the assumed distribution. This is hypothesis testing.

The examples in the document you linked to all involve a known or assumed distribution. E.g.

A snack manufacturer fills bags of potato chips to a target weight of μ = 283 grams. The population standard deviation of bag weights is σ = 9 grams. A quality inspector randomly selects a sample of n = 36 bags from the production line. What is the probability that the sample mean weight is less than 280 grams? (A sample mean below 280g would trigger a compliance audit.)

Do you think the value of the mean of the PDF is going to change due to sampling?

The mean of the population and the expected value of the sample mean should not change with sample size.

But, for that to be true, the sample must be IID with the population.

Please look up what iid means.

What you really mean is that the sample should be representative of the population. This should be the case if the sample is genuinely random, or if proper sampling methods are applied. A biased sample will not be accurate – that should be obvious.

Are the 30 days of Tmax a sample or a population.

Yes.

It really depends on what you want, what question you are trying to answer. I’ve said this many times in regard to the TN1900 example. If you want the average maximum temperature at that station for that month, then the 30 maximums are a population. The average of those 30 days is the average of those 30 days, and the only uncertainty is that from the measurements.

If you treat them as a sample, as in TN1900 you are assuming a hypothetical average temperature and distribution from which all your daily readings are randomly taken. The base uncertainty then is then SEM.

Of course in the case of TN1900 they only have 22 out of the 31 daily values so it is always going to be a sample, which adds to the confusion as to exactly what they are trying to do in the example.

Reply to  Bellman
August 21, 2026 5:03 pm

The base uncertainty then is then SEM.

Uncertainty quantified by variance (standard deviation), not sigma over root-n.

Reply to  karlomonte
August 22, 2026 6:33 am

Bellman is *NEVER* going to figure this one out.

s/sqrt(n), where s is the standard deviation of the sample, tells you the precision with which you have located the mean of the sample.

It has *nothing* to do with the precision with which you have located the mean of the parent distribution.

s/sqrt(n) can approach zero while the mean of the sample can be nowhere near the population mean.

You *have* to have a sampling distribution, i.e. multiple samples with multiple means all with sufficient sample size, in order to determine how precisely you have located the population mean. The standard deviation of the sample mean distribution, i.e. the sampling distribution, tells you how precisely you have located the population mean.

Bellman’s assumption is that the single sample duplicates the parent distribution either exactly or, at least, very closely for both mean and variance.

But there is no guarantee that this is the case, especially if the ratio of (sample size-n)/(population size-N) doesn’t approach 1. It’s why the SEM is undefined for one single sample.

With just one sample you cannot tell if the sample is IID, whether the sample is biased, whether the sample is from a multi-modal distribution, whether there is any drift in the process being sampled, or “whether the sample is representative of the parent distribution“.

To fix all of these the ratio of n/N must be very large – an impossibility with global temperature measurement.

It’s just one more garbage assumption that bellman and climate science makes in order to “justify” their global average temperature.

Reply to  Bellman
August 22, 2026 6:42 am

That is the best estimate of the population mean is the sample mean, and the SEM is estimated from the sample standard deviation.”

The SEM of the sample is how precisely you have located the mean of the sample. It is *NOT* how precisely you have located the mean of the population.

In order for the SEM to tell you how precisely the sample mean approaches the population mean you must assume that the sample perfectly mimics the parent distribution, same mean and same standard deviation.

But being IID does *NOT* guarantee the sample mean and variance perfectly mimic the parent mean and variance.

In order to make this assumption the ratio of the sample size (n) to the population size (N) must approach 1. n/N -> 1.

Otherwise you cannot detect bias in the sample, drifting in the parent process being sampled, multi-modal existence, non-stationarity, and a host of other issues.

You are basically ASSuming that the sample distribution perfectly mimics the parent distribution for mean, variance, kurtosis, etc.

n/N for global temperatures simply doesn’t approach 1, not even close.

Reply to  Tim Gorman
August 22, 2026 6:46 pm

The SEM of the sample is how precisely you have located the mean of the sample.

You keep using phrases like this and it makes no sense. You know how precisely you have located the mean of the sample. It’s just the mean of the sample. There is no uncertainty about what the mean of the sample is.

It is *NOT* how precisely you have located the mean of the population.

You cannot locate the mean of the sample. It’s unknown. What the SEM does is give you a range of values about your sample mean in which the population mean is likely to lie – or if you prefer the range of values that it is reasonable to attribute to the population mean.

In order for the SEM to tell you how precisely the sample mean approaches the population mean you must assume that the sample perfectly mimics the parent distribution, same mean and same standard deviation.

I’m sure this means something to you, and I could guess what you are trying to say. But as it stands it’s nonsense. If the sample had the same mean as. The whole reason for talking about a sampling distribution is becasue you know the sample is not going to be identical to the population.

But being IID

You still haven’t learned why this is nonsense. You refuse to read what iid means. You re either just saying this to annoy me or you are genuinely incapable of learning or admitting your misunderstandings.

In order to make this assumption the ratio of the sample size (n) to the population size (N) must approach 1. n/N -> 1.

Thus demonstrating you have no idea what a sample is.

Reply to  Jim Gorman
August 22, 2026 5:51 am

But, for that to be true, the sample must be IID with the population”

This is a hard to grasp point. Being IID does not mean that the sample has the same mean and variance as the parent. It only means that the EXPECTED mean is the same as the parent, the realized mean from the sample may not be such.

It’s why you get a sample means distribution. The samples can be IID, i.e. you EXPECT to get the parent mean from each sample, but the actual values don’t match it exactly.

But all kinds of issues with the parent can prevent IID, e.g. autocorrelation in the parent data, or different variances in the random variables making up the parent distribution.



Reply to  Tim Gorman
August 22, 2026 6:59 am

But all kinds of issues with the parent can prevent IID, e.g. autocorrelation in the parent data, or different variances in the random variables making up the parent distribution.

Read Taylor Section 5.7 and list all of the assumptions necessary for mathematically deriving the formula for SDOM. I’ll start.

To answer this, we naturally imagine repeating our N measurements many times; that is, we imagine performing a sequence of experiments, in each of which we make N measurements and compute the average.

  1. Multiple samples of the same thing, each of which has N measurements.
  2. The only unusual feature of the function (5.64) [the mean of each sample] is that all the measurements x,,…, xₙ happen to be measurements of the same quantity, with the same true value X and the same width σₓ.


There are several other assumptions necessary for the derivation, please list them.

Reply to  Tim Gorman
August 22, 2026 7:43 am

If you examine Taylor, equation 5.66 requires that ALL the partial derivatives of each sample uncertainty be the same in order to obtain the quantity of (1/N). Otherwise, each individual term will have its own unique value of its partial derivative. This means you can no longer estimate the SDOM by simply dividing by the √N.

This also requires ALL σ’s be the same as the population σ. In other words, IID.

If there is not IID, then Equation 5.66 falls apart because N is no longer a common term, and the SDOM must be calculated using Equation 5.65.

Guess what Equation 5.65 is? It is the RSS value of the individual sample uncertainty, i.e., the SD’s of each sample added in quadrature!

Reply to  Jim Gorman
August 22, 2026 6:30 pm

ALL the partial derivatives of each sample uncertainty be the same in order to obtain the quantity of (1/N).

You are confusing the word “sample” again. It’s the uncertainty of each measurement that are assumed to be the same. That’s the “identically distributed” part of iid.

Guess what Equation 5.65 is? It is the RSS value of the individual sample uncertainty, i.e., the SD’s of each sample added in quadrature!

No it isn’t. 5.65 is just the general equation for propagating errors. It still involves multiplying each term by the partial derivative, and that will still be 1/n for the average. The only difference if the uncertainties for each value are different is that you end up with the sum of those uncertainties in quadrature divided by n.

Reply to  Bellman
August 22, 2026 5:33 am

Still just as clueless.”

I have given you MULTIPLE definition of the CLT and MULTIPLE definitions of the SAMPLING DISTRIBUTION.

And you have learned NOTHING. You are stubbornly clinging to your misconception that one sample creates an SEM.

“If I throw a 6-sided die and it comes up 5, that 5 was a random value taken from the actual probability distribution of the die.”

A SAMPLE requires at least 30 entries from a parent distribution to be a part of a valid means distribution from which an estimate of the population mean can be made.

You are floundering around. STOP IT. FOCUS!

“Now apply that logic to your sample of sample methods. If you have 100 samples, the SEM estimate will still have an uncertainty of 7%. You’ve just had to spend a lot more in order to get the same amount of uncertainty.”

With your n = 1 sample, there isn’t even an SEM! There isn’t any uncertainty because there is an SEM!

And this is where you are just confusing yourself. Understandable giving how many different standard deviations there are floating about. s is the sample standard deviation.”

No, the sample standard deviation is *NOT* the SEM. Nor is it the meaasurment uncertainty of the average.

The SEM is the standard deviation of the sample MEANS. One mean is not a distribution with a mean!

How many more definitions of the SEM is required before it sinks into your head? Why do you think all of the definitions speak of the CLT and SEM together? With enough samples the CLT will push the means of those multiple samples into a Gaussian distribution whose mean approximates the population mean! How closely it approximates the population mean is the standard deviation of the SAMPLE MEANS. One sample cannot form a distribution of sample means.

If you only have one sample then you MUST ASSume that its mean and standard deviation EXACTLLY MATCHES THAT OF THE POPULATION. But being IID does *not* define that the mean and standard deviation of ONE sample exactly match that of the population distribution.

All larger sample sizes do is push the distribution of sample means closer to being Gaussian with a mean matching that of the population.

It’s one more garbage assumption that you and climate science makes!

Reply to  Tim Gorman
August 22, 2026 1:35 pm

“I have given you MULTIPLE definition of the CLT and MULTIPLE definitions of the SAMPLING DISTRIBUTION.”

And you are still clueless about what hey mean. You still think that the CLT depends on the number of samples and not on sample size.

“A SAMPLE requires at least 30 entries from a parent distribution to be a part of a valid means distribution from which an estimate of the population mean can be made. ”

Clueless.

“If you only have one sample then you MUST ASSume that its mean and standard deviation EXACTLLY MATCHES THAT OF THE POPULATION.”

Clueless. Have you noticed how you keep typing ASS in capitals. Maybe your subconscious is telling you something.

“All larger sample sizes do is push the distribution of sample means closer to being Gaussian with a mean matching that of the population. ”

Keep thinking? Maybe you will realise that you are describing the CLT. The larger the sample size, the closer to a Gaussian and the smaller the SEM.

Reply to  Bellman
August 22, 2026 4:22 pm

And you are still clueless about what hey mean. You still think that the CLT depends on the number of samples and not on sample size.”

Sample size reduces the variance of each sample mean and smooths the distribution. It makes the sampling distribution more Gaussian.

The number of samples affects how well the sampling distribution fits a Gaussian distribution. If you take 5 samples (of whatever size) you won’t be able to even see a Gaussian shape of the sampling distribution. If you take 100 samples the histogram of the sample means will look visibly Gaussian.

If you don’t form a visible Gaussian distribution by having enough samples how do you know its Gaussian? The uncertainty of the mean determined by the means of the samples, i.e. the sampling distribution, goes up if the shape of the curve is not Gaussian.

A non-Gaussian shape is a clue that some kind of dependence is at play, autocorrelation, drift, spatial/temporal correlation, etc.

Not being able to see a Gaussian shape doesn’t change the mean, the mean doesn’t depend on the shape. But not being able to see the Gaussian shape harms the ability to trust the standard deviation of the sample means, i.e. the SEM.

Even with large n, the sampling distribution of the means can fail to be Gaussian if the data are not independent. The CLT won’t fix this. You need to be able to *see* that the sampling distribution is Gaussian. You can’t just assume a Gaussian shape from large “n”.

The effective “n” (n_eff) is < n if there is dependence. The CLT, SEM, etc depends on n_eff, not n.

For instance, if you have spatial correlation, which climate science says you have between close measuring stations, then

n_eff ∝ f(n,p) where n is the number of locations and p is the correlation factor. If I remember its something like

n_eff = (n)/ [(n-1)p + 1

If you have two temperature sensors where the correlation factor is 0.9 (i.e. high correlation) then

n_eff = (2)/ 1.9 = 1.05.

In other words you don’t have n = 2. you have n=1.

Autocorrelation is even worse but its been so long I don’t remember the formula any more.

Bottom line? You simply don’t know enough about sampling to even discuss it intelligently. Just like metrology.

Apparently climate science doesn’t either. You don’t divide anything by “n” if there is any dependence involved in the data, you divide by n_eff < n. And climate science apparently doesn’t do this even though they claim significant dependence with the data.

  1. If the dependence isn’t high then homogenization and infilling doesn’t work!
  2. If the dependence is high then the SEM, etc depends on n_eff.

You *really* need to stop trying to lecture on subjects you simply have no knowledge of.

Reply to  Tim Gorman
August 22, 2026 4:37 pm

If you take 5 samples (of whatever size) you won’t be able to even see a Gaussian shape of the sampling distribution. If you take 100 samples the histogram of the sample means will look visibly Gaussian.

You are confusing an illustration of the CLT, with the actuality. The distribution exists whether you can see it or not.

If you don’t form a visible Gaussian distribution by having enough samples how do you know its Gaussian?

Maths!

Reply to  Bellman
August 22, 2026 5:44 am

That’s why you have additional uncertainty . It’s why you use a student distribution and not a Gaussian. It’s also another reason why a larger sample size is better than a small one.”

That is *NOT* why you use a student-T! You use a student-T because the variance of the parent distribution is not known! The student-T describes the sampling distribution of the sample means when the parent variance is unknown. It prevents heavy tails in the sampling distribution of the sample means!

Standard deviation of the sample divided by √N.”

Again, this is *NOT* the SEM. It is not *anything* when you only have one sample.

It is a shortcut estimation of the SEM and it doesn’t work for one sample unless you make the garbage ASSumption that the single sample exactly matches the mean and standard deviation of the parent distribution.

Think about it. “s” is the sample standard deviation. If it is *NOT* the same as the standard deviation of the parent then how can it estimate the standard deviation of the parent?

Your logic implies that one sample IS ALL THAT IS EVER REQUIRED to estimate the mean and standard deviation of the parent!

Using that logic why would you ever take more than one sample of anything?

Reply to  Tim Gorman
August 22, 2026 9:40 am

“You use a student-T because the variance of the parent distribution is not known! ”

Yes that’s the point. You are estimating the population SD from the sample and that increases the uncertainty interval.

“Again, this is *NOT* the SEM. It is not *anything* when you only have one sample. ”

Could you actually say how you think this works? In the equation s/√N, what do you think s is?

“It is a shortcut estimation of the SEM and it doesn’t work for one sample unless you make the garbage assumption that the single sample exactly matches the mean and standard deviation of the parent distribution.”

Yes it’s an estimate. You don’t know what the population is so you have to estimate it it’s the nature of working with data in the real world.

“Think about it. “s” is the sample standard deviation. If it is *NOT* the same as the standard deviation of the parent …”

That’s why it’s called the sample standard deviation rather than the parent standard deviation. Thanks for enlightening me, oh master.

“… Then how can it estimate the standard deviation of the parent?”

By being the best estimate you usually have. The larger the sample the better the estimate.

“Your logic …”

I can’t claim credit for it. It’s been known for over a century. And elementary text book explains it.

“Your logic implies that one sample IS ALL THAT IS EVER REQUIRED to estimate the mean and standard deviation of the parent!”

Yes, you can estimate the mean and standard deviation from a single sample, the bigger the better. What do you think we’ve been trying to tell you? That doesn’t mean you can’t improve that the estimate by adding more observations, either by increasing the sample size or by pooling multiple samples.

Reply to  Bellman
August 22, 2026 10:17 am

Yes that’s the point. You are estimating the population SD from the sample and that increases the uncertainty interval.”

NO! The student-t lessens the impact of the tails! How in Pete’s name does that INCREASE the standard deviation?

Could you actually say how you think this works? In the equation s/√N, what do you think s is?”

With a single sample “s” is the standard deviation of the sample. s/√n IS THE SEM OF THE SAMPLE! It is a metric for how precisely you have located the mean of the sample!

It is *NOT* a metric for how precisely you have located the mean of the population unless you ASSume that the one single sample has a distribution that perfectly mimics the population distribution, i.e. both have exactly the same standard deviation.

The use of multiple samples is *exactly* the same logic as that behind a Type A measurement uncertainty. The multiple means of the multiple samples will BRACKET the mean of the population. With enough samples to generate a Gaussian distribution from the sample means, the mean of the sample means is the BEST ESTIMATE of the population mean. The uncertainty interval associated with that Best Estimate is the standard deviation of the sample mean distribution. EXACTLY the same logic as a Type A measurement uncertainty.

Taking one sample is exactly like making one measurement. You cannot develop a sample means distribution from one sample just like you can’t develop a measurement distribution from one measurement.

Think of a production line where you want to know an average property value of the output units. Would you pull one sample of 30 consecutive units or would you pull 30 samples of 30 consecutive units where each of the 30 samples is pulled at a different time?

Would that single sample of 30 units give you a good value for the average of some property over the entire production?

Yes it’s an estimate. You don’t know what the population is so you have to estimate it it’s the nature of working with data in the real world.”

It’s not even a valid estimate since you have no way to determine its validity. And that is *NOT* the nature of working with data in the real world. See the production line example above.

By being the best estimate you usually have. The larger the sample the better the estimate.”

If that is the best you can do then you need to change your entire sampling process. NO quality engineer will take one sample as a valid estimate unless n/N is very high. And that requires more testing than pulling multiple smaller samples that are independent and are not correlated in time.

“Yes, you can estimate the mean and standard deviation from a single sample, the bigger the better. What do you think we’ve been trying to tell you?”

What elementary textbook recommends this? Demming recommended pulling multiple random samples across the entire production run – i.e. over time. This minimizes auto-correlation and the samples become closer to IID. And it minimizes both testing time AND interference with the production output.

I don’t think you know any more about sampling than you know about metrology! NOTHING for both!

Reply to  Tim Gorman
August 22, 2026 3:30 pm

NO! The student-t lessens the impact of the tails! How in Pete’s name does that INCREASE the standard deviation?

It doesn’t increase the SEM, what it does is widen the tails for small sample size.

With a single sample “s” is the standard deviation of the sample.

s is always the standard deviation of the sample. What else do you think it is?

It is *NOT* a metric for how precisely you have located the mean of the population unless you ASSume that the one single sample has a distribution that perfectly mimics the population distribution, i.e. both have exactly the same standard deviation.

Grow up. Then think about what you are saying. The sample SD is almost never going to be exactly the same as the population SD, and even if it were how would you know? That’s why you have to estimate the population SD from the sample.

The use of multiple samples is *exactly* the same logic as that behind a Type A measurement uncertainty.

A Type A measurement uncertainty is derived from a single sample. The Experimental standard deviation of the mean is just the standard deviation of that sample divided by √N. Do you think the standard deviation of that sample will be exactly the same as the population? Everything is uncertain. Everything is an estimate. And it would be just as pointless to take multiple samples of samples just to estimate the experimental standard deviation of the mean.

Taking one sample is exactly like making one measurement.

You are very confused. Each element in the sample is one measurement.

Would you pull one sample of 30 consecutive units or would you pull 30 samples of 30 consecutive units where each of the 30 samples is pulled at a different time?

You can do either. But 30 consecutive units is not a random sample. They won’t necessarily be representative of the units over the entire time frame.

Pooling 30 samples of 30 units each will obviously be better, as you know have a sample of 900 rather than 30, and the values will be less dependent on the specific time.

But what you would not do, is take the mean of each of your 30 samples, and use the SD of them to estimate SEM for a single sample.

NO quality engineer will take one sample as a valid estimate unless n/N is very high.

Huh? What are n and N in your scenario? If you mean N is the population size, that’s irrelevant to the SEM. The assumption is that N is infinite, or that the sample is made with replacement. If you are sampling nearly all the population you are not really sampling. Why not just test everything?

And that requires more testing than pulling multiple smaller samples that are independent and are not correlated in time.

I think you are losing your own plot. You were the one saying you had to take multiple samples. I’m saying that in general you only want one representative sample.

Demming recommended pulling multiple random samples across the entire production run – i.e. over time.

You are changing the subject again. Your claim was that you couldn’t estimate the SEM from a single sample and you had to take the standard deviation of the means of many samples in order to estimate the SEM. Now you’ve switched to the idea of pooling multiple samples in order to ensure you have a representative sample, that’s a completely different issue.

Reply to  Bellman
August 22, 2026 3:53 pm

What elementary textbook recommends this?

A few random examples from the internet.

Our goal in sampling is to determine the value of a statistic for an entire population of interest, using just a small subset of the population. We do this primarily to save time and effort – why go to the trouble of measuring every individual in the population when just a small sample is sufficient to accurately estimate the statistic of interest?

In the election example, the population is all registered voters in the region being polled, and the sample is the set of 1000 individuals selected by the polling organisation. The way in which we select the sample is critical to ensuring that the sample is representative of the entire population, which is the main goal of statistical sampling. It’s easy to imagine a non-representative sample; if a pollster only called individuals whose names they had received from the local Greens Party, then it would be unlikely that the results of the poll would be representative of the population as a whole. In general, we would define a representative poll as being one in which every member of the population has an equal chance of being selected. When this fails, then we have to worry about whether the statistic that we computed using the sample is biased – that is, whether its value is systematically different from the population value (which we will refer to as a parameter). Keep in mind that we generally don’t know this population parameter, because if we did then we wouldn’t need to sample!

https://stats.libretexts.org/Bookshelves/Applied_Statistics/A_Contemporary_Approach_to_Research

What Can We Conclude from 1 Sample?

§More than you might think

§Thanks to the Central Limit Theorem

To Estimate Mean from a Single Sample

§1) Choose sample size based on estimate of skew in

population

§2) Chose a random sample from the population

§3) Compute the mean and standard deviation of that

sample

§4) Use the standard deviation of that sample to

estimate the SE

§5) Use the estimated SE to generate confidence

intervals around the sample mean

https://ocw.mit.edu/courses/6-0002-introduction-to-computational-thinking-and-data-science-fall-2016/02f2ed5e44be5415b2cba8175c924092_MIT6_0002F16_lec8.pdf

Consider the following scenario:

A researcher ‘X’ is collecting data from a large population of voters. For practical reasons he can’t reach out to each and every voter. So, only a small randomized sample (of voters) is selected for data collection.

Once the data for the sample is collected, you calculate the mean (or any statistic) of that sample. But then, this mean you just computed is only the sample mean. It cannot be considered the entire population’s mean. You can however expect it to be somewhere close to population’s mean.

So how can you know the actual population mean?

While its not possible to compute the exactly value, you can use standard error to estimate how far the sample mean may spread from the actual population mean.

https://machinelearningplus.com/statistics/standard-error/

Reply to  Bellman
August 22, 2026 9:11 pm

This is a nice example of sampling a population. It is totally unrelated to measurement uncertainty.

I asked you a question before but never saw an answer.

Q: Is the 30 observations of a monthly Tmax a population or a sample of a population?

This is a very important first decision made in determining the statistical impacts. If it is a population, why are we treating it as a sample? If it is a sample of a larger population of Tmax in that time period, one must explain why.

Reply to  Jim Gorman
August 23, 2026 5:41 am

It is totally unrelated to measurement uncertainty.

We were talking about sampling, not measurement uncertainty. In particular the claim that you would never take just one sample.

I asked you a question before but never saw an answer.

I answered it here.

https://wattsupwiththat.com/2026/08/09/defining-temperature/#comment-4232397

This is a very important first decision made in determining the statistical impacts.

Yes, that’s what I said. Whether you treat them as a sample or a population depends on what question you are asking.

Reply to  Bellman
August 22, 2026 5:09 pm

A Type A measurement uncertainty is derived from a single sample. 

You are very confused. Each element in the sample is one measurement.

It is you who has everything upside-down, or have a highly esoteric/fluid definition of the word sample.

Reply to  karlomonte
August 22, 2026 5:37 pm

GUM B.2.17 experimental standard deviation
for a series of n measurements of the same measurand, the quantity s(q_k) characterizing the dispersion of the
results…
Note 1: Considering the series of n values as a sample of a distribution, q^bar is an unbiased estimate of the mean μ_q…

Reply to  Bellman
August 23, 2026 4:54 am

“It doesn’t increase the SEM, what it does is widen the tails for small sample size.”

You are correct about widening the tails. I misspoke on that. BUT, widening the tails means the variance goes up — which INCREASES the SEM, it doesn’t decrease it.

“s is always the standard deviation of the sample. What else do you think it is?”

The standard deviation of the distribution formed by the multiple sample means.

“Grow up. Then think about what you are saying. The sample SD is almost never going to be exactly the same as the population SD, and even if it were how would you know? That’s why you have to estimate the population SD from the sample.”

How would you know? By taking MULTIPLE samples! Multiple samples whose means bracket the population mean. And whose standard deviation is a metric for how precisely you have bracketed the population mean.

With ONE sample you have no metric to use in determining how precisely you have bracketed the population mean. The standard deviation of the sample is NOT an indicator for how close you are to the population mean – it is an indicator for how accurately you have located the sample mean.

You are *still* tied to the ASSumption that a single sample will *ALWAYS* be a perfect imitator of the population distribution.

“A Type A measurement uncertainty is derived from a single sample”

You totally MISSED THE ENTIRE POINT. A single measurement would be like having a single sample!

With the Type A evaluation you are involved with ONE thing and are generating a distribution of measurements of the same thing to be used in developing a Best Estimate whose uncertainty metric is the standard deviation.

With the temperature data you are involved with MULTIPLE things and are generating multiple samples to create a distribution of sample means used in developing a Best Estimate of the average of the multiple things.

You want to use ONE measurement of that single thing to be an acceptable estimate of the value being measured. You want ONE sample of those multiple things to be an acceptable estimate of the average value of the entire group of the multiple things.

You are making a fool of yourself trying to justify that ONE sample from a distribution of different things will *always* give acceptable results for the mean of the group.

This is especially garbage when considering the auto-correlation of the values from the different things, the spatial correlation of the different things, and the temporal correlation of the different things. Because of all this n_eff is far smaller than n, yet all you want to use is sqrt(n) and not sqrt(n_eff).

It also means that the ratio of n/N becomes n_eff/N and the ratio of your sample size vs the population get small – meaning the uncertainty goes way up!

Reply to  Bellman
August 23, 2026 5:28 am

Huh? What are n and N in your scenario? If you mean N is the population size, that’s irrelevant to the SEM.”

It is TOTALLY relevant to the SEM! It doesn’t change the *value* of the calculated SEM but the n/N ratio changes what the SEM *means*!

The SEM describes the uncertainty of the population IF the sample is representative of the population, i.e. same mean and same standard deviation.

But you can *NOT* test whether the sample is representative of the population FROM ONE SAMPLE unless n/N -> 1.

You wind up with an SEM whose applicability is unknown. ASSuming that the single sample is representative of the population is a GARBAGE ASSumption unless n/N -> 1.

You can do either. But 30 consecutive units is not a random sample. They won’t necessarily be representative of the units over the entire time frame.”

And what is a mid-range temperature using Tmax and Tmin? Those are representative of the daily temperature profile, i.e. over the entire time frame?

You are digging yourself a hole. Maybe you should stop digging?

Pooling 30 samples of 30 units each will obviously be better, as you know have a sample of 900 rather than 30, and the values will be less dependent on the specific time.”

30 samples of 30 units taken at different times *IS NOT* creating a single sample of 900!

It is creating 30 sample distributions whose means form a sampling distribution. They do *NOT* have temporal correlation.

“But what you would not do, is take the mean of each of your 30 samples, and use the SD of them to estimate SEM for a single sample.”

Why would you use the sampling distribution standard deviation as the standard deviation of any specific sample that is part of a set of samples? You are trying to estimate the population mean and the standard deviation of the sample means tells you how precisely the estimate of the population mean has been located.

You have 30 separate samples from the population created by the production run. The means of those 30 samples create the sample mean distribution whose mean is the Best Estimate of the population mean and whose SD is the uncertainty associated with that mean of the sample means.

Hell, using your logic YOU WOULD NEVER, EVER GET MULTIPLE SAMPLES. You would always wind up with 1 sample! And 1 sample does not create a sample distribution!

wikipedia: “The sampling distribution of a mean is generated by repeated sampling from the same population and recording the sample mean per sample. This forms a distribution of different sample means, and this distribution has its own mean and variance. Mathematically, the variance of the sampling mean distribution obtained is equal to the variance of the population divided by the sample size. This is because as the sample size increases, sample means cluster more closely around the population mean.” (bolding mine, tpg)

This appears to be a result of your poor reading comprehension skill. No one, including me, has ever said the SD of the sample means is the SD of each individual sample!

Reply to  Bellman
August 23, 2026 5:38 am

I think you are losing your own plot. You were the one saying you had to take multiple samples. I’m saying that in general you only want one representative sample.”

As usual you simply can’t read!

I said that creating ONE SAMPLE whose n/N ratio is sufficient to make the assumption that the single sample is representative of the population requires MORE testing than taking multiple, smaller samples!

How can you so misinterpret that?

You are changing the subject again. Your claim was that you couldn’t estimate the SEM from a single sample and you had to take the standard deviation of the means of many samples in order to estimate the SEM. Now you’ve switched to the idea of pooling multiple samples in order to ensure you have a representative sample, that’s a completely different issue.”

You STILL CAN’T READ!

I never said *anything* about combining multiple samples into ONE BIG SAMPLE!

You use the multiple samples to create a sampling distribution!

All you did was CHERRY PICK from what I wrote in order to create a strawman to argue with.

Just for reference here is what I said:

tpg:”What elementary textbook recommends this? Demming recommended pulling multiple random samples across the entire production run – i.e. over time. This minimizes auto-correlation and the samples become closer to IID. And it minimizes both testing time AND interference with the production output.” (bolding mine, tpg)



Reply to  Tim Gorman
August 22, 2026 5:05 pm

Demming recommended pulling multiple random samples across the entire production run – i.e. over time. This minimizes auto-correlation and the samples become closer to IID. And it minimizes both testing time AND interference with the production output.

ASTM has some very useful standards for SPC.

Reply to  Tim Gorman
August 22, 2026 6:58 pm

It prevents heavy tails in the sampling distribution of the sample means!

I think you’ll find it does the opposite. It makes the tails heavier.

Reply to  Bellman
August 19, 2026 6:49 am

Don’t forget his refusal to acknowledge that uncertainty is quantified by standard deviation

That’s not what I said.

Of course it is what you wrote, backpedaling now?

Standard uncertainty is the standard deviation of measurements.

Wrong — uncertainty is quantified by standard deviation, regardless if Type A or B.

What I said is that the uncertainty of the mean is not the standard deviation of the values that went into calculating the mean.

A totally meaningless number, but it is what you hang you entire existence upon.

Reply to  karlomonte
August 19, 2026 6:28 am

Don’t forget his refusal to acknowledge that uncertainty is quantified by standard deviation, even though the GUM says exactly this.

The GUM doesn’t really say this. It says the uncertainty is EXPRESSED as a standard deviation where a Type A uncertainty is evaluated statistically. Type B uncertainty is not necessarily determined by a statistical analysis although its value is expressed as a standard uncertainty equivalent to a standard deviation. The GUM goes to great lengths to create an equivalency between Type A and Type B even though they are obtained differently. That should be a clue to you that statistics is not the end all be all of uncertainty. It is also appropriate to know that not all uncertainty uses the GUM. One such is range uncertainty that uses intervals. It is a reason NIST and other bodies recommend not using ±values. Intervals with a confidence level are more comparable.

Reply to  Jim Gorman
August 19, 2026 7:00 am

The GUM doesn’t really say this. It says the uncertainty is EXPRESSED as a standard deviation where a Type A uncertainty is evaluated statistically. Type B uncertainty is not necessarily determined by a statistical analysis although its value is expressed as a standard uncertainty equivalent to a standard deviation. The GUM goes to great lengths to create an equivalency between Type A and Type B even though they are obtained differently.

3.3.5 and 3.3.6 use variance rather than s.d.; combining them needs both on the same basis (units).

3.3.4 says:

“Both types of evaluation are based on probability distributions (C.2.3), and the uncertainty components resulting from either type are quantified by variances or standard deviations.”

That should be a clue to you that statistics is not the end all be all of uncertainty. It is also appropriate to know that not all uncertainty uses the GUM. One such is range uncertainty that uses intervals. It is a reason NIST and other bodies recommend not using ±values. Intervals with a confidence level are more comparable.

This is absolutely correct, something that b & b will probably never figure out, all they want is loopholes.

Reply to  Tim Gorman
August 18, 2026 3:40 pm

No, they are usually incapable of understanding the real world.

You are not a mathematician or statistician, and you constantly demonstrate you understand little of either subject. To use one of your usual put-downs, what give you the right to lecture all mathematicians and statisticians on your superior understanding of subjects you know nothing about?

Statistics in particular is about the real world. It’s describing a world of uncertainty and randomness. Without statistics there would be no uncertainty analysis. Everything in Taylor, Bevington, the GUM etc, is statistics.

it doesn’t matter who wrote them

It mattered to you when you dismissed them as being written by mathematicians.

What they wrote is what matters.

Then argue about what they wrote, rather than your usual ad hominems.

You consistently return to the concept that “best estimates” are true values

Stop lying. It really should be possible for you to argue with someone without deliberately misunderstanding everything they say. By definition a “best estimate” is not a “true value”. There’s a clue in the word “estimate”. That should be obvious to you, whether or not you believe a “true value” exists or not.,

…measurement uncertainty can just be ignored or assumed to cancel.

Again, stop this nonsense. It’s not surprising you never understand these basic concepts if you just keep arguing with bogymen of your own invention. When have I ever claimed that measurement uncertainty can be ignored? I’ve been discussing measurement uncertainty with you for over 5 years – you might have figured out by now that means I don;t think it should just be ignored.

Reply to  Bellman
August 19, 2026 4:57 am

You are not a mathematician or statistician, and you constantly demonstrate you understand little of either subject.

You are not a trained physical science professional and you constantly demonstrate you understand little of how measurements are used in physical science.

Quoting the uncertainty of the mean is only important to a statistician that wants to know how accurately the sample mean predicts the real mean. For physical science use, the uncertainty of the mean can only be used for one specific measurand with multiple observations under repeatable conditions in a short time. That uncertainty DOES NOT translate to the next measurand. There is no guarantee the next measurand will have either the same value or the same uncertainty.

A physical scientist or engineer wants to know the range of possible values that were measured. Replicating experiments to validate measurement estimates is how theories are proven to predict dependent results. Knowing the range of values measured allows someone working in the physical sciences to determine safety margins. It allows one to ensure that product specifications are legally defendable.

Ranges of expected values allow control of calibration costs. Rather than having equipment recalibrated the second it slips out of the range of the uncertainty of the mean is unnecessary primarily because that range will be outside the ability of the device to even measure.

Concentrating on achieving the smallest “error” quotation won’t protect you when someone says “why didn’t you tell me that I couldn’t depend on the exact value”. Examples: The Challenger Shuttle, the Titanic, numerous auto recalls for safety.

Have you ever studied Deming’s quality control practices? He was a famous statistician. He recognized that every product did not need to meet exacting measurements. His methods meld measurements into a process that results in high-quality output, without every piece reaching one, very specific measurement in the range of the uncertainty of the mean. His theories allow one to use standard deviations to define acceptability, not just exacting measurements.

Here is a little story whose moral should tell you statistics are not the be all, end all and that the uncertainty of the mean is not what measurements are about.

The boss walks into the quality manager’s office and says “What is going on in the manufacture of our rods. We are getting half of them back because they are too long or too short. The quality manager says “That can’t be,” and pulls out a graph. He tells the boss we examine every tenth rod, measure it, and put it into a sample of ten. When we have 100 samples of ten each, we compute the average of each sample, plot it and find the mean of all the samples. We then calculate the error of the mean, and “You can see the error of the mean is well within our tolerance criteria”. The boss says, “You dumba**, what is the standard deviation and/or variance? I want to know the range of products we are sending out the door. The manager responds, “I don’t know we don’t look at that”. Do you think this fellow will keep his position?

Reply to  Jim Gorman
August 19, 2026 5:49 am

 For physical science use, the uncertainty of the mean can only be used for one specific measurand with multiple observations under repeatable conditions in a short time.”

It is typically glossed over, but there is also the restriction that the observations must be a Gaussian distribution. Otherwise the CLT doesn’t push the sample means distribution to Gaussian very quickly, i.e. for a small number of samples.

Reply to  Jim Gorman
August 19, 2026 7:47 am

“Quoting the uncertainty of the mean is only important to a statistician that wants to know how accurately the sample mean predicts the real mean.”

Why do you think that’s only of interest to statisticians?

” For physical science use, the uncertainty of the mean can only be used for one specific measurand with multiple observations under repeatable conditions in a short time. ”

You have a very low opinion of physical scientists. Even physists should understand the benefit of being able to compare one figure with another.

” There is no guarantee the next measurand will have either the same value or the same uncertainty. ”

Again, you have an extremely narrow understanding of what measurements are used for. You obviously only use them in one context, to avoid litigation. And refuse to understand that they can also be used to investigate the real world. If you are looking at how the world is changing you don’t expect every measurement to be the same. You are not manufacturing global temperatures to a fixed standard, you are measuring them to find out what the temperature is and how it’s changing over time.

Reply to  Bellman
August 19, 2026 8:11 am

You are not manufacturing global temperatures to a fixed standard, you are measuring them to find out what the temperature is and how it’s changing over time.

By extracting milli-Kelvins from 1 degree data?

Reply to  karlomonte
August 20, 2026 5:42 am

It comes from assuming that all measurement uncertainty cancels so the stated values can be considered 100% accurate. Then milli-Kelvin calculations are possible.

Reply to  Bellman
August 20, 2026 5:35 am

Again, you have an extremely narrow understanding of what measurements are used for. You obviously only use them in one context, to avoid litigation. And refuse to understand that they can also be used to investigate the real world.”

This from the man that has absolutely *NO* understanding of physical science?

You understand that even the speed of light has wound up being DEFINED instead of measured because of the measurement uncertainty associated with all the different experimental results?

You are not manufacturing global temperatures to a fixed standard, you are measuring them to find out what the temperature is and how it’s changing over time.”

If your measurement uncertainty is larger than changes then how do you know what the actual change is? You *still* can’t figure out that measurement uncertainty *HAS* to be considered when examining measurements and their use. You can’t just assume that the stated values are 100% accurate and can give you milli-Kelvin differences when the measurement uncertainty is in the tenths digit at best and probably in at least the units digit!

Reply to  Tim Gorman
August 20, 2026 7:15 am

“Again, you have an extremely narrow understanding of what measurements are used for. You obviously only use them in one context, to avoid litigation. And refuse to understand that they can also be used to investigate the real world.”

This from the man that has absolutely *NO* understanding of physical science?

Yeah, this one was an absolute hooter.

You understand that even the speed of light has wound up being DEFINED instead of measured because of the measurement uncertainty associated with all the different experimental results?

Think he knows what epsilon-zero and mu-zero are? Not a chance.

Reply to  Bellman
August 19, 2026 5:34 am

You are not a mathematician or statistician, and you constantly demonstrate you understand little of either subject.”

*I* am not the one that claims the (Single-Sample-SD)/sqrt(n) is the SEM. YOU are. CLIMATE SCIENCE is.

It’s just one more piece of garbage math you and climate science attempt to propagate.

“To use one of your usual put-downs, what give you the right to lecture all mathematicians and statisticians on your superior understanding of subjects you know nothing about?”

I don’t claim a right to just say “you are wrong”. I SHOW HOW THEY ARE WRONG. Just like above with the SEM being calculated from one sample!

You have yet to actually disprove any of my math or statistics – especially when it comes to metrology. All you ever offer are garbage assertions like “you can calculate an SEM from one sample”, or that “all measurement uncertainty is random, Gaussian, and cancels”, or that you *can* average different intensive property values to get a physically meaningful average value, or that mid-range values are “average values”.

Statistics in particular is about the real world”

No, it isn’t. It’s about “numbers is just numbers”. It’s about the shape of a distribution of a set of numbers, not about what the numbers represent.

If statistics were about the real world then statisticians would refuse to calculate the average of different intensive properties because the result is meaningless in the real world!

Without statistics there would be no uncertainty analysis.”

So what? The issue isn’t about the math itself, it’s about how the result of the math is USED! Such as saying that the SEM is MEASUREMENT uncertainty instead of SAMPLING uncertainty. Or saying that you can calculate the SEM from one sample. Or saying that the average of different intensive properties is meaningful.

Then argue about what they wrote, rather than your usual ad hominems.”

I gave you a complete explanation of why what they wrote was garbage. Go back and reread it. Pointing out that saying that significant figures are used to correct “arithmetic errors” is *NOT* an ad hominem.

You just *never* bother to actually read anything, do you?

 By definition a “best estimate” is not a “true value”.”

Then why do you treat it as one? If it is just an estimate and not a true value then exactly what does the SEM tell you? Are you going to change your best estimate value based on what the SEM is?

Even your claim that the SEM is usually larger than the measurement uncertainty implies that you believe the “best estimate” to be a “true value” estimate.

Your misconceptions are so ingrained in your brain, and you are so adamant about remaining willfully ignorant on metrology that you can’t even recognize when your misconceptions come to the front of what you say.

When have I ever claimed that measurement uncertainty can be ignored?”

When you claim the SEM is usually greater than the measurement uncertainty for one! When you claim the average is a measurement and not just a statistical descriptor for a distribution shape. When you assume that all measurement distributions are random and Gaussian – which is a requirement for the SEM = s/√n to be useful, especially when the number of samples is small.

Reply to  Tim Gorman
August 19, 2026 7:49 am

“*I* am not the one that claims the (Single-Sample-SD)/sqrt(n) is the SEM. ”

Yes, because you don’t understand how any of this works that’s your problem.

Reply to  Bellman
August 19, 2026 8:13 am

Intellectually superior condescending bellman returns…

Reply to  Bellman
August 20, 2026 6:34 am

Single_Sample_SD / sqrt(n)

IS THE SEM OF THE SAMPLE MEAN!

IT IS NOT THE SEM OF THE POPULATION MEAN.

(unless you assume the sample is Iid with the population. Is this one more garbage climate science assumption?)

Reply to  Tim Gorman
August 20, 2026 7:03 am

This level of anger can’t be good for your health.

Yes, the SEM is the standard error of the sample mean. There is no such thing as the error of the population mean – the population mean has no uncertainty, it cannot be treated as a random variable, it has no probability distribution. (Unless we are getting in to Bayesian statistics).

unless you assume the sample is Iid with the population”

You still don’t know what iid means.

Reply to  Bellman
August 20, 2026 7:33 am

Yes, the SEM is the standard error of the sample mean. 

And no, it is NOT uncertainty.

Reply to  Bellman
August 20, 2026 9:05 am

There is no such thing as the error of the population mean”

You simply can’t or won’t read.

What do you think SEM = σ/sqrt(n)

means?

You are confusing the fact that you may not know what “σ” is with the fact that it is part of the definition of how accurate a sample of the population can locate the mean of the population.

*IF* you know “σ”, then the precision with which the population mean can be located is determined by your sample size.

If you don’t know “σ” then the precision is best estimated by taking MULTIPLE samples and finding the standard deviation of the sample means. The CLT pushes the multiple sample means toward a Gaussian distribution with an average value approaching that of the population average with a level of uncertainty defined by the standard deviation of the sample means.

In fact, for the equation SEM = σ/sqrt(n) to be totally valid *ALL* of the samples have to be IID as well. That actually means that if the measurement values used have different measurement uncertainties, IID probably isn’t satisfied either.

It’s *all* a mess. And somehow climate science makes all the ASSumptions they need to in order to somehow get uncertainty down to the thousandths digit for temperature so they can identify differences in the hundredths digit accurately!

Reply to  Tim Gorman
August 20, 2026 11:52 am

“What do you think SEM = σ/sqrt(n)”

It’s the standard error of the mean. Hence SEM. It is not the standard error if the population mean, because that is a meaningless term you have just invented. The mean of the population is fixed, it is not a random variable, the only thing that can have an error related to the population mean is a sample.

“If you don’t know “σ” then the precision is best estimated by taking MULTIPLE samples and finding the standard deviation of the sample means.”

Define “best”? Is your definition of best taking into account how much time and cost it will take you? Say you can afford to take 100 measurements. Is it better to treat those 100 measurements as a single sample, and estimebthe SEM from its standard deviation. Or is it better to treat it as 10 samples each of size 10, and estimate the SEM from those 10 samples?

Reply to  Bellman
August 21, 2026 3:43 am

It’s the standard error of the mean. Hence SEM. It is not the standard error if the population mean, because that is a meaningless term you have just invented. The mean of the population is fixed, it is not a random variable, the only thing that can have an error related to the population mean is a sample.”

What in Pete’s name do you think I’ve been trying to teach you?

A sample can only ESTIMATE the population mean. How well it estimates the population mean due to sampling is SAMPLING UNCERTAINTY, commonly referred to as the standard error.

Why do you think everyone keeps asking you if temperature data sets like UAH is a single sample of multiple data components or is a collection of individual sample means?

As usual, you simply cannot read. I said “a sample of the population” and the rest of the post was about “a sample of the population”.

In essence your whole post here is nothing but a non sequitur because its an argument you are making in order to support a red herring you have created – likely because you can’t read simple English.

“Define “best”? Is your definition of best taking into account how much time and cost it will take you? Say you can afford to take 100 measurements. Is it better to treat those 100 measurements as a single sample, and estimebthe SEM from its standard deviation. Or is it better to treat it as 10 samples each of size 10, and estimate the SEM from those 10 samples?”

If those 100 measurements are of the same thing under repeatable conditions then they are a single population and there is *NO* SEM. If the distribution of the measurements is random and Gausian then the mean is the best estimate of the value of the measurand and the measurement uncertainty of the measurements is the standard deviation. If the distribution of the measurements is not random and Gaussian then a different method for determining a best estimate should be determined.

If they are 10 measurements each of a property associated with ten similar but different things then each set of 10 will have a best estimate and a measurement uncertainty and those ten measurement uncertainties adds to a total. Add them directly or by root-sum-square depending on the situation. The SEM of this situation is irrelevant. It gets subsumed into the measurement uncertainty. The mean of the ten measurements is the BEST ESTIMATE of the value of the measurand in question if the distribution is random and Gaussian. If it isn’t random and Gaussian a different method of determining the Best Estimate will be required, e.g. the mode. The measurement uncertainty is the interval of reasonable values that could be assigned to the value of the measurand.

I could explain this to a third grader and expect them to understand it. And you “can’t”, after three years of having measurement uncertainty explained to you?

Reply to  Tim Gorman
August 20, 2026 7:16 am

And it is not uncertainty of anything.

Reply to  Bellman
August 17, 2026 3:29 pm

Didn’t come up with anything specific. Did you have an actual philosophical paper in mind?

Here is one,

Evolution of the Significant Figure Rules 
340_1_1.4818368.pdf

A phrase form the document is enlightening.

Faculty would like students to realize that measurements have a precision, and that the precision of a measurement ultimately sets the precision for any calculation.

The key here is that you are not going to ever get an “equation” or formal rule for sig figs and measurement calculations. No one should expect that from estimates to begin with. The key is to be honest on what information was gathered in the measurement and how the uncertainty of that measurement should inform people the quality of the equipment you used to make the measurements. Creating resolution WILL mislead as to the precision of your instrument.

Reply to  Jim Gorman
August 17, 2026 4:54 pm

Thanks, but I’m not sure where the epistemological underpinning of significant digits appears in that review.

It does seem to agree with me that the sig fig rules are just rules of thumb and approximations for proper uncertainty analysis. It’s conclusion is

In addition, we need to emphasize that these are general rules of thumb, not rules. We should not ask our students “how many significant figures” a particular calculation has based on the rules. Instead we should ask our students “approximately what is the precision” of a particular calculation. The rules are not exact. Quizzing our students as if they are seems like a waste of time.

bdgwx
Reply to  Jim Gorman
August 16, 2026 11:37 am

The value should be quoted as 3249 ±6. The 0.5 is subsumed into the uncertainty. The interval will be 3243 to 3255.

Per JCGM 100:2008 section 4.2.3 and 7.2 it should be reported as 3248.5 ± 2.8 k=1.

Reply to  bdgwx
August 16, 2026 3:14 pm

Per JCGM 100:2008 section 4.2.3 and 7.2 it should be reported as 3248.5 ± 2.8 k=1.

Let’s look at what the GUM sections you reference actually say.

4.2.3 The best estimate of σ 2 (q ) = σ 2 n , the variance of the mean, is given by

Equation 5

The experimental variance of the mean s2(q) and the experimental standard deviation of the mean s(q) (B.2.17, Note 2), equal to the positive square root of s2(q), quantify how well q estimates the expectation μq of q, and either may be used as a measure of the uncertainty of q .

Thus, for an input quantity Xi determined from n independent repeated observations Xi,k, the standard uncertainty u(xi) of its estimate xi = Xi is u(xi) = s(Xi ), with s2 (Xi ) calculated according to Equation (5). For convenience, u2 (xi ) = s 2(Xi ) and u(xi ) = s(Xi ) are sometimes called a Type A variance and a Type A standard uncertainty, respectively.

NOTE 1 The number of observations n should be large enough to ensure that q provides a reliable estimate of the expectation μq of the random variable q and that s2(q ) provides a reliable estimate of the variance σ 2(q ) = σ 2 n (see 4.3.2, note). The difference between s2(q ) and σ 2(q ) must be considered when one constructs confidence intervals (see 6.2.2). In this case, if the probability distribution of q is a normal distribution (see 4.3.4), the difference is taken into account

through the t-distribution (see G.3.2).

NOTE 2 Although the variance s2(q ) is the more fundamental quantity, the standard deviation s(q ) is more convenient in practice because it has the same dimension as q and a more easily comprehended value than that of the variance.

I don’t see any indication here of how the number should be quoted. Let’s examine 7.2.

7.2.2

1) “mS = 100,021 47 g with (a combined standard uncertainty) uc = 0,35 mg.”

2) “mS = 100,021 47(35) g, where the number in parentheses is the numerical value of (the combined standard uncertainty) uc referred to the corresponding last digits of the quoted result.”

3) “mS = 100,021 47(0,000 35) g, where the number in parentheses is the numerical value of (the combined standard uncertainty) uc expressed in the unit of the quoted result.”

4) “mS = (100,021 47 ± 0,000 35) g, where the number following the symbol ± is the numerical value of (the combined standard uncertainty) uc and not a confidence interval.”

Nowhere else in 7.2 does it mention actual numbers. If you will notice item 3, it should be plain that the 2 significant digits in the uncertainty, (0.00035) have the same place as the value (100.02147), the 4th and 5th decimal places.

Why don’t you cut and paste the exact passage in these paragraphs that leads you to believe that your statement of 3248.5 ± 2.8 is correct.

Let me also add that

you are missing the part of calculating the variance that has (xᵢ – μ). Addition/subtraction rules state that the difference cannot exceed the value with the least number of significant digits. That means that 4 significant digits should be used.You did not measure the tenths digit. To state a measurement to the tenths digit means you MEASURED it to that resolution. Quoting the tenths digit is misleading especially when the uncertainty is in the units digit.I’ll ask one more time. Show some 3rd or 4th level lab notes in a physical science lab that lets you quote a value beyond what was measured.

Check out this site that has instructions for a university level chemistry course. Pay attention to Item 2.

The MAXIMUM possible number of sig figs in an average is the number of sig figs in the data. The sig figs may be limited further by the decimal position of the least significant digit as determined from the st dev.

bdgwx
Reply to  Jim Gorman
August 16, 2026 5:37 pm

Section 4.2.3 says:

“Thus, for an input quantity Xi determined from n independent repeated observations Xik, the standard uncertainty u(xi) of its estimate xi = Xi is u(xi) = s(Xi) with s^2(Xi) calculated according to Equation 5.”

Given q = {3251, 3240, 3255, 3248} then q_bar = 3248.5 and s(q_bar) = 5.5 / sqrt(4) = 2.75. Therefore u(q) = s(q) = 2.75.

Section 7.2.6 says:

“The numerical values of the estimate y and its standard uncertainty uc(y) or expanded uncertainty U should not be given with an excessive number of digits. It usually suffices to quote uc(y) and U [as well as the standard uncertainties u(xi) of the input estimates xi] to at most two significant digits, although in some cases it may be necessary to retain additional digits to avoid round-off errors in subsequent calculations.”

and

“Output and input estimates should be rounded to be consistent with their uncertainties; for example, if y = 10,057 62 Ω with uc(y) = 27 mΩ, y should be rounded to 10,058 Ω.”

Therefore we state the whole thing as 3248.5 ± 2.8 k=1.

To be pedantic the GUM actually discourages ±, but doesn’t strictly prohibit it.

Check out this site that has instructions for a university level chemistry course. Pay attention to Item 2.

I already did. I pointed out that it’s rule is not consistent with the other sources your cited. Furthermore, it’s rule bases the number of digits on the sd whereas the GUM uses the number of digits in the standard uncertainty calculated as sd/sqrt(N). Hopefully this drives home my point that there is no universal standardized handling of significant figures. Despite that the GUM guidelines are the most logical since there is less ambiguity.

Reply to  bdgwx
August 16, 2026 9:44 pm

To be pedantic

Yes, you have a special talent here.

Reply to  bdgwx
August 17, 2026 6:36 am

Given q = {3251, 3240, 3255, 3248} then q_bar = 3248.5 and s(q_bar) = 5.5 / sqrt(4) = 2.75. Therefore u(q) = s(q) = 2.75.

Therefore we state the whole thing as 3248.5 ± 2.8 k=1.

Only in climate science, where extra resolution is manufactured out of the ether or wishful thinking.

The measurements are to the singles, so the real answer is

3249 ± 6 with k = 2

bdgwx
Reply to  karlomonte
August 17, 2026 9:44 am

Only in climate science, where extra resolution is manufactured out of the ether or wishful thinking.

It’s the GUM. It’s authored by the BIPM, International Union of Pure and Applied Chemistry, and International Union of Pure and Applied Physics among others.

Reply to  bdgwx
August 17, 2026 10:53 am

Four digits, and you believe you can get five via the magic of averaging.

I think you are not an honest person.

bdgwx
Reply to  karlomonte
August 17, 2026 4:34 pm

Four digits, and you believe you can get five via the magic of averaging.

That’s what the GUM says.

I think you are not an honest person.

It’s not my rule. If you think there is dishonesty in it then tell BIPM, IUPAC, IUPAP, etc.

Reply to  bdgwx
August 17, 2026 5:18 pm

Doubling-down on your magic averaging claptrap.

Reply to  bdgwx
August 17, 2026 6:07 pm

“Buh, buh, buh, GUM!”
“Buh, buh, buh, <anything else I can misquote>!”

Reply to  bdgwx
August 17, 2026 8:40 am

Therefore we state the whole thing as 3248.5 ± 2.8 k=1.

You quote it the way you want. Since you obviously have never had experience in a machine shop or making measurements out in the wild, you can play with the numbers all you want. Foks who have actually done the work of making measurements know the correct methods and you do not convince them otherwise.

Here is an example from the real world for you to criticize.

Let’s examine a reading from a typical micrometer. The image below is a typical reading that one might obtain. How is the value obtained? 1st, one must read the value on the sleeve. This has a value of 6.5 mm.  2nd,One reads the thimble and obtains a value of 0.11. The total is 6.5 + 0.11 = 6.61 mm. As you can see, the reference line is between 11 and 12. There are several options available for deciding what the recorded value should be. 1) Estimate to the nearest 0.005. 2) Use the next line BELOW the reference line. 3) Use a micrometer with a vernier scale. The most often recommendation is option 2 which gives a reading of 6.61 ±0.005 mm. Next is option 3. Last is option 1. One must recognize that using option 1, the uncertainty will still lie 0.005 decimal position because the last digit is estimated. It is where uncertainty begins. So one could estimate 6.615 ±0.005 mm. That would give an interval of 6.610 to 6.620 mm.  Whereas option 2 gives an interval of 6.605 to 6.615 mm.
comment image

The image below is an image of a micrometer with a vernier reading as described in option 3. It has a third scale to estimate the third decimal value. In this image one has a spindle reading of 5.5, a thimble reading of 0.28, and a vernier reading of 0.003. That gives a total of 5.5 + 0.28 + 0.003 = 5.783 ±0.0005 mm for an interval of 5.7825 to 5.7835 mm. As you can see the vernier allows a much higher resolution but only at an added cost.  
comment image

This how measurements are made. It is why LIG temperatures were estimated to the nearest graduation. It is why resolution is so important when reporting measurements. These two examples show where uncertainty starts and how it must be handled in calculations. Ultimately, it results in an interval that reported.

These are analog instruments where one can read between the resolution of graduations. Digital instruments do not allow any interpretation. If you have a two decimal place display, you cannot know where the signal lies between the “graduations”. This is where error bars on graphs originate.

Tell us how you would report these measurements.

bdgwx
Reply to  Jim Gorman
August 17, 2026 9:46 am

You quote it the way you want.

I have no preference either way. I’m just tell you It’s the way the GUM says to do it.

Reply to  bdgwx
August 17, 2026 10:55 am

Bull.

Averaging does NOT get you extra digits.

bdgwx
Reply to  karlomonte
August 17, 2026 4:32 pm

Averaging does NOT get you extra digits.

Tell BIPM, IUPAC, IUPAP, etc. They’re the ones who wrote the GUM.

Reply to  bdgwx
August 17, 2026 5:19 pm

And you wonder why you climatology trendology nuts are laughed at…

Reply to  Jim Gorman
August 16, 2026 6:10 pm

I don’t see any indication here of how the number should be quoted. 

Oops #1.

Nowhere else in 7.2 does it mention actual numbers. If you will notice item 3, it should be plain that the 2 significant digits in the uncertainty, (0.00035) have the same place as the value (100.02147), the 4th and 5th decimal places.

Oops #2.

Check out this site that has instructions for a university level chemistry course. Pay attention to Item 2.

The MAXIMUM possible number of sig figs in an average is the number of sig figs in the data. The sig figs may be limited further by the decimal position of the least significant digit as determined from the st dev.

Oops #3.

Reply to  Bellman
August 17, 2026 6:32 am

You do if you accept the sum is correct. If you have 10 values that sum to 252.3, you know the average of 25.23 is as correct as your sum is.”

The issue isn’t with the average value you calculate. The issue is how you use that value. If you are going to use it as the BEST ESTIMATE of the measurand in a statement of the value that has been measured, then that statement of the value HAS TO FOLLOW RESOLUTION AND SIGNIFICANT FIGURE rules in order for it to be used as a reference in subsequent measurements.

You would *NOT* state the measurement as

25.23 +/- 0.5 units

It would be 25.3 +/- 0.5 units at best. It might even be better stated as 25 +/- 0.5 units.

You keep getting stuck in blackboard world.

Reply to  Tim Gorman
August 17, 2026 6:56 am

“You do if you accept the sum is correct. If you have 10 values that sum to 252.3, you know the average of 25.23 is as correct as your sum is.”

The issue isn’t with the average value you calculate. The issue is how you use that value. If you are going to use it as the BEST ESTIMATE of the measurand in a statement of the value that has been measured, then that statement of the value HAS TO FOLLOW RESOLUTION AND SIGNIFICANT FIGURE rules in order for it to be used as a reference in subsequent measurements.

His little blackboard thought experiment has the advantage that all he gave was the total sum, no individual values, no uncertainties. A perfect instance of numbers is numbers.

Problem ignored!

Reply to  Bellman
August 9, 2026 7:05 pm

We keep going over the same ground, but expressing something to 3 decimal places is not the same as claiming it is correct to 3 decimal places.”

If it isn’t considered correct then it is FRAUD on the reader to express it as such! It implies that your measurements *are* accurate to 3 decimal places.

The international metrology standard is to ONLY REPORT DIGITS IN THE STATED VALUE THAT ARE JUSTIFIED BY THE MEASUREMENT’S UNCERTAINTY.

You do *NOT* give a measurement as 10.01C +/- 0.1C.

Reply to  Tim Gorman
August 10, 2026 4:41 am

“The international metrology standard is to ONLY REPORT DIGITS IN THE STATED VALUE THAT ARE JUSTIFIED BY THE MEASUREMENT’S UNCERTAINTY.”

I keep asking if you think writing things on capital letters makes it true. We’ve been over what the international standards say many times. All the GUM says, along with all other metrology sources I can find, is that the number of decimal places should match the reported uncertainty, and that the reported uncertainty should not have too many significant figures. If the uncertainty if a measurement is 0.034 then the stated value would be to 3 decimal places. It is not saying that result is correct to the hundredth of a degree.

“You do *NOT* give a measurement as 10.01C +/- 0.1C.”

But you would say 10.02 ± 0.12°C.

Reply to  Bellman
August 10, 2026 5:01 am

 If the uncertainty if a measurement is 0.034 then the stated value would be to 3 decimal places. It is not saying that result is correct to the hundredth of a degree.”

Did you ACTUALLY read this before you hit the POST button?

“But you would say 10.02 ± 0.12°C.”

Really? How do you state a measurement to the hundredths digit when the uncertainty is in the tenths digit?

I would say 10.0 +/- 0.1C

10.02 + 0.12 = 10.14
10.02 – 0.12 = 9.9

Meaning you don’t actually know if the measurement is known in the UNITS digit, let alone the tenths or hundredths digit.

Reply to  Tim Gorman
August 10, 2026 5:19 am

“Did you ACTUALLY read this before you hit the POST button?”

YES!

bdgwx
Reply to  Bellman
August 10, 2026 5:28 pm

All the GUM says, along with all other metrology sources I can find, is that the number of decimal places should match the reported uncertainty, and that the reported uncertainty should not have too many significant figures.

Exactly. And since the GUM provides at least one example reporting the uncertainty to 3 significant figures we know that a value stated as say 12.345 ± 0.678 C complies with the guidelines contained within the GUM.

Reply to  Bellman
August 9, 2026 9:38 pm

Why bother to report it if it isn’t reliable and doesn’t provide information? The number of significant figures implies the resolution or precision. It is therefore misleading if there are more significant figures than are justified.

Reply to  Clyde Spencer
August 10, 2026 4:34 am

“Why bother to report it if it isn’t reliable and doesn’t provide information?”

Because it is reliable and does provide information.

Reply to  Bellman
August 10, 2026 4:51 am

Because it is reliable and does provide information.”

It is *NOT* reliable because it tries to show resolution the uncertainty doesn’t support.

And if it is not reliable then it doesn’t provide information either! It becomes opinion and not fact.

Reply to  Bellman
August 10, 2026 6:54 am

Because it is reliable and does provide information.

Only to climate trendologists, the rest of the world is laughing at you.

Reply to  Bellman
August 10, 2026 9:37 pm

Extra digits do NOT provide information if the digits are not significant! Do you not understand the meaning of “significant?”

Reply to  Bellman
August 10, 2026 1:38 am

Expressing a temperature to 3 dp IS EXACTLY THE SAME as claiming it is correct to 3 dp.

Sparta Nova 4
Reply to  Bellman
August 10, 2026 6:20 am

But it is done that way so people will assume it is.

It violates scientific notation.
The precision (number of decimal places or digits) of a calculation can be no more that the least precise value in the calculation.

1 x 2.1 = 2, not 2.1.

Reply to  Bellman
August 10, 2026 6:42 am

Most global data sets, including UAH, are reported to at least 2 decimal places,

The UAH absolute temperature data is reported to FIVE significant digits, which is a relative uncertainty of ±0.005%. If you knew anything about real metrology, you might known this level of uncertainty needs a calibration lab with carefully controlled environment and premium instrumentation.

But deep down you don’t care, plowing ahead and ignoring reality just like the rest all the other climate trendologists.

but nobody would suggest they are correct to the hundredth if a degree.

Don’t lie, of course they (and you) do, whenever the impossibly tiny “error bars” are used or requoted.

“Buh, buh, buh whatabout rounding??” — you.

bdgwx
Reply to  Clyde Spencer
August 9, 2026 3:04 pm

That doesn’t mean they are stating that the uncertainty is 0.01 C (or 0.005 C). In fact, they explicitly state their uncertainty and it is definitely higher than 0.01 C sometimes significantly so. [Lenssen et al. 2024] [GISTEMP Uncertainty Analysis]

And as we’ve already discussed your rules for significant figures are not only arbitrarily applied, but are themselves arbitrary and not consistent with the guidelines and examples published in [JCGM 100:2008]. NASA’s publication of temperatures are consistent with the guidelines and examples contained within the GUM.

Reply to  bdgwx
August 9, 2026 7:22 pm

And as we’ve already discussed your rules for significant figures are not only arbitrarily applied”

They are *NOT* arbitrarily applied. THEY ARE INTERNATIONAL METROLOGY STANDARDS.

NOAA *does* report temperatures in the North American Dataset to the hundredths of a degree. Even when the uncertainty is in at least the UNITS DIGIT.

NOAA reports temperatures in the GHN-Daily to the tenths of a degree – even when the measurement uncertainty is in at least the UNITS DIGIT.

Even the Local Climate Data is stored in the tenths digit Celsius when the measurement uncertainty of most of the instruments is +/- 1.0C.

The international metrology standard is that the stated value should have no more decimal places than is supported by the measurement uncertainty. If the uncertainty (i.e. +/- 1.0C) is in the units digit, the temperatures should be rounded to the appropriate units digit.

That doesn’t mean they are stating that the uncertainty is 0.01 C (or 0.005 C). In fact, they explicitly state their uncertainty and it is definitely higher than 0.01 C sometimes significantly so”

Reporting to decimal places not supported by the measurement uncertainty is perpetrating a fraud upon science. The whole purpose of the measurement uncertainty paradigm is to allow judgement as to how accurate a measurement is so subsequent measurements can be judged for accuracy.

Since almost no one in climate science ever reports measurement uncertainty (what is the measurement uncertainty of the GAT? Not the sampling uncertainty, the MEASUREMENT uncertainty) only the stated value is available for use in judging accuracy. And reporting the GAT to the hundredths digit with no accompanying measurement uncertainty interval leads to implying an accuracy in the hundredths digit!

Reply to  Tim Gorman
August 10, 2026 1:43 am

I believe the main reason these meteorological temperatures are reported to 2 or even 3 dp is to impress the rubes in the general public with a superficial “sciencey” veneer.

Reply to  bdgwx
August 9, 2026 9:26 pm

Who are “we”?

oeman50
Reply to  Jim Gorman
August 10, 2026 5:30 am

Good one. Just because you can mathematically produce an average temperature or a calculated temperature less than the instrumental uncertainty doesn’t mean is has value.

Reply to  Pat Frank
August 9, 2026 1:09 pm

Depending on one’s working definition of temperature, LiG may be considered to be a proxy. However, the correlation between the volume expansion coefficient of the liquid and the temperature is far higher than the correlation between the width of tree rings and temperature, largely because there are no spurious, confounding variables like insolation, water availability, and temperature. That is, anomalous heat or cold can shut down tree growth temporarily. One does not have that issue with mercury.

Reply to  Clyde Spencer
August 10, 2026 1:20 pm

Expansion of the liquid is a glass thermometer, due to a change in density, is a direct measure of the KE of the Hg atoms or the alcohol molecules. As an energetic measure, expansion is a direct indication of temperature; not a proxy.

Reply to  Pat Frank
August 10, 2026 9:57 pm

Because mass is conserved, a change in density is accomplished by a change in volume. The change in volume of a column of mercury, constrained by the bore of the thermometer, is in turn expressed as a change in length. Thus, the length of the mercury column is a function of temperature as calculated with the expansion coefficient. The length of the column can be referenced to a standard. However, there isn’t a way to compare the average kinetic energy of the air mass, whose temperature we would like to measure, by comparing it to any standard in Paris.

We can do direct comparisons between lengths and masses with standards, which define the units. But we cannot do similar comparisons for temperatures, except at special temperature and pressure points where phase changes occur; intermediate temperatures have to be obtained by interpolation.

Reply to  Clyde Spencer
August 11, 2026 7:01 am

That’s all fine, Clyde. But LiG measurements are not a temperature proxy.

Reply to  Pat Frank
August 11, 2026 8:53 pm

Apparently we agree to disagree.

Reply to  Clyde Spencer
August 14, 2026 3:32 pm

An increase in the KE of the fluid in the glass capillary is responsible for the expansion. The KE of the fluid comes into equilibrium with the KE of the measurand.

KE = 1/2 mv² = 3/2kT. The conversion is direct. The LiG thermometer is a direct measure of 3/2kT. It is not a proxy for T.

paul courtney
Reply to  Pat Frank
August 9, 2026 2:57 pm

Mr. Frank: Well, that’s just silly! Every one knows (oops, gotta use science lingo) the consensus is tree ring metric is unsurpassed, validated by every other consensus metric including LiG and data from models.
Sorry, I enjoy performing imitations, that was final nail.

Reply to  paul courtney
August 10, 2026 1:21 pm

You had me going. 🙂

August 9, 2026 6:56 am

Temperature is the kinetic energy of stuff.
We define the scales.
Boiling water at sea level is 100 C or 212 F or 373 Kelvin
Ice water slurry is 0 C or 32 F or 273 Kelvin.
A Celsius unit is (boiling – ice)/100

There are Celsius units on the Celsius scale.
There are Celsius units on the Kelvin scale.
There is no such thang as a Kelvin unit.

Reply to  Nicholas Schroeder
August 9, 2026 8:34 am

288 K – 255 K = 33C not 33 K.
33 C is delta.
33 K is 33 C above abs 0.

PV = nRT
T is Celsius degrees on the Kelvin scale.

Q = sigma epsilon A T^4.
T is Celsius degrees on the Kelvin scale.

You better know which to use when or your project will find itself in a canyon instead of on a plain.

Reply to  Nicholas Schroeder
August 9, 2026 4:00 pm

Pure water, at exactly 1 atmosphere pressure.

Reply to  Retired_Engineer_Jim
August 9, 2026 6:48 pm

With regard to this parameter (atmospheric pressure), NBS now NIST has a procedure they use during formal verification/certification of thermometers submitted for cal/verification/certification, noted in the reference I cited earlier: https://archive.org/details/liquidinglassthe150wise_0/mode/2up

LT3
August 9, 2026 7:15 am

People arguing about the validity of using temperature as evidence for whatever, ignore the foundational use of all of the temperature measurements throughout history. It’s disrespectful to dismiss the work that teams of people that dealt with the differences of measuring temperature from buckets in shipping during sailing era and all the complexities of dealing with ship engines raw water intake temperature readings to modern buoy measurements. These datasets have allowed the discovery of ENSO patterns, multi-centennial ocean current structures as well as weather forecasting 2 weeks out that is fairly accurate.

Reply to  LT3
August 9, 2026 8:11 am

It’s disrespectful to dismiss the work that teams of people that dealt with the differences of measuring temperature from buckets in shipping during sailing era and all the complexities of dealing with ship engines raw water intake temperature readings to modern buoy measurements

The problem with the work of these people is that many, and even most, do not adequately address the uncertainty involved. One cannot discover “bias” between measuring devices and procedures without an adequate common basis between them to make judgements from.

Different devices and procedures will result in varying results of measurement. The readings from different devices and procedures can be 100% accurate but be different. Adjusting a series becomes nothing more than a choice to make differing data series match a preconceived conclusion. That isn’t science, it is pseudoscience.

Any science that takes measurements read and recorded at integer resolution and adds more resolution that was not observed originally is ignoring the information that is known in favor of using calculator derived decimal digits. It is creating resolution out of thin air. That isn’t science, it is pseudoscience.

Jeff Alberts
Reply to  LT3
August 9, 2026 12:21 pm

as well as weather forecasting 2 weeks out that is fairly accurate.”

Now that right there is funny.

Richard Mott
August 9, 2026 7:24 am

Andy – so if local temperature assumes a small volume, what is it that the UAH microwave sounding instruments measure? I give them more credibility precisely because they sample huge volumes of air rather than point thermometer measurements subject to all kinds of local influences. If it is somehow the “average” of an enormous number of local temperatures what does than mean in physics terms? (Forgive my ignorance – I’m an EE who just pushes electrons around with voltages without worrying too much about the teleological implications.)

John Hultquist
Reply to  Richard Mott
August 9, 2026 7:58 am

When you push electrons around, where do they go and how fast?

Reply to  John Hultquist
August 9, 2026 1:38 pm

John, I haven’t found a citation, but I knew before looking that it took several minutes (8-12?) for telegraph messages sent from England to India to arrive. I suspect that the travel time depends on the voltage used. While use of copper is ubiquitous, different metals with different dielectric constants probably have different transmission speeds. (dielectric constant is related to the complex refractive index). I’m pretty sure that the electrons that come out at the India telegraph station are not the same ones pushed in at the station in England (Wales?). They act like ‘bumper cars’ and it is the electromagnetic field that propagates, not individual electrons.

Reply to  Clyde Spencer
August 9, 2026 7:11 pm

re: “but I knew before looking that it took several minutes (8-12?) for telegraph messages sent from England to India to arrive.”

The problem they had, was, very slow rise and fall times … they were driving the undersea line with high voltages to get the usual voltage levels at the other end that they had always used on land telegraphs … a nearly matched line, or a line with equalization inductance (effectively raising the line’s impedance) still results in a velocity of propagation up near the speed of light … coaxial cable itself today ranges from 66% (solid polyethylene) to a little over 88% (for low density foamed polyethylene.)

But they didn’t understand line impedance yet.

Recall back in time something called “The Telegrapher’s Equation” developed by Oliver Heaviside starting in 1876.

https://duckduckgo.com/?q=%22The+Telegrapher%27s+Equation%22&t=ffcm&ia=web

Reply to  Clyde Spencer
August 9, 2026 7:35 pm

If cable pairs are being used then the transmission delays are also related to the cable geometry and physical makeup. If you are using a cable pair then the inductance and capacitance of the cable pair determines the delay. The propagation delay is a function like t_d ∝ sqrt(LC). Since L and C are frequency dependent and a square wave is made up of multiple frequencies, a transition from no-voltage to voltage (i.e. the leading edge of a square wave) the “pulse” gets smeared, rounded, however you want to describe it. EE’s call it “rise time” of the leading edge. It takes a certain amount of time for the transition to be recognized at the far end. That determines the speed at which pulses can be sent down the line. (the same thing happens at the end of the pulse, it’s just called the “fall time”)

They act like ‘bumper cars’ and it is the electromagnetic field that propagates, not individual electrons.”

Yep!

Reply to  Tim Gorman
August 10, 2026 8:55 am

I read an article about this once. There were two theories of what happens.

One, is that electrons are knocked loose and immediately start moving toward the lower potential. No collisions, just electron flowing all the way from the positive potential to the other end.
Two, the bumper car effect where an electron knocks another one loose and the collisions happen all the down the line.

There are arguments for and against each theory. Like how would moving electrons not collide with molecules. However, collisions would require time resulting in a delay of electron flow so there goes the blinding electronic speed of light.

I think the final theory is that the electric field propagates at the speed of light from one end to the other causing electron drift to occur simultaneously all along the path.

Reply to  Jim Gorman
August 10, 2026 11:35 am

In semiconductors both happen (and more, like tunneling through a barrier).

Reply to  karlomonte
August 10, 2026 2:59 pm

I remember at KU as a senior, the school had hired two professors who did special research at Bell Labs in tunneling. Trying to put all the concepts and equations together in the mind stretched it somewhat. Useful later on, but wasn’t circuit design which is what I was interested in.

Reply to  Jim Gorman
August 10, 2026 3:58 pm

Still remember growing a diode as a senior…wasn’t the greatest, but it rectified.

Reply to  Jim Gorman
August 10, 2026 10:20 pm

However, collisions would require time resulting in a delay of electron flow so there goes the blinding electronic speed of light.

The thing is, light that propagates through a diamond only does so at about half the speed it does in a vacuum. Thus, it has a real index of refraction of about 2.

Cerenkov radiation (blue glow in nuclear reactors) is, as I understand it, produced from shock waves created when particles (probably mostly neutrons) exceed the speed of light in water.

Electromagnetic pulses propagate even more slowly through metals, with some metals having a real component (n) of about 5, suggesting a photon speed that is about 20% of the speed of light at optical frequencies. I’m not sure what happens at much lower frequencies, like 60 Hz.

While I’m probably the world’s living expert on the optical constants of opaque minerals, I can’t tell you how the extinction coefficient changes the speed of propagation, or how it changes with wavelength. Nor do I know if or how the speed of electrons varies with the energy they carry. That is why I speculated that the speed may vary with the EMF. I may have to look into that.

Reply to  Clyde Spencer
August 11, 2026 5:04 am

Electromagnetic pulses propagate even more slowly through metals”

The propagation of an EM wave is not dependent on “electron” movements. An EM wave can propagate through a vacuum where no electrons exist.

It is the permittivity/permeability of a medium that determines propagation speed. These properties are a measure of the electric/magnetic field ability to store energy in the medium. Similar to the inductance/capacitance of a cable pair propagating an EM field.

Even a vacuum has a permittivity, it’s a basic property of “space”. Think of it as somewhat like inertia. It’s why the speed of propagation of an EM wave isn’t infinite.

The propagation is not based on electron movement, drift, or energy. It’s a “field” that propagates, not electrons.

Reply to  Andy May
August 9, 2026 12:00 pm

[The UAH] team validated their estimates against weather balloon radiosonde data.”

Weather balloon radiosonde data are known to be no better than about ±0.4 C.

Given the ±0.3 C resolution of the microwave sounders, the UAH calibration uncertainty is about ±0.5 C.

But their published global TLT record never includes uncertainty bounds.

Reply to  Pat Frank
August 9, 2026 7:14 pm

Would one expect the ‘bias’ to be always in one direction, IOW an offset, rather than a +-x value and random within that range? WHAT factor determines the error(s)?

Reply to  _Jim
August 10, 2026 4:10 am

In the real world the uncertainty is many times asymmetric. Most materials expand when heated. The expansion grows over time if the heating remains. That expansion changes the material. Sending current through a material causes heat, even in a PRT. The big question is what change is caused in the material.

In measurement devices that expansion in the materials result in calibration drift over time. The drift is *usually* in one direction since the characteristic known as “resistance” changes in the same direction for most materials.

Reply to  Tim Gorman
August 10, 2026 6:59 am

Manufacturers of precision resistors routinely quote drift specifications as something like ±X% per year, meaning they don’t state in which direction the drift will move.

Reply to  karlomonte
August 10, 2026 11:11 am

Most manufacturers of precision parts will give a temperature drift specification caused by continuous heating.

The plus and minus covers drifting caused by environmental conditions, e.g. using the part at the South Pole where it is cold vs in Rio where it is warm.

If a part gets colder it shrinks. If it gets warmer it expands. Both cause drift. But continuous heating, be it from current through the part or from some other means, very seldom causes a random drift with some parts drifting down and some drifting up.

Reply to  Tim Gorman
August 10, 2026 4:25 pm

re: “In the real world the uncertainty is many times asymmetric. Most materials expand when heated. The expansion …”

I’m talking MICROWAVE SOUNDERS as the nature and accuracy of radiosonde temperature sensors (platinum sensors) are known.

The QUESTION was in regards to MICROWAVE spectrometer / thermometry accuracy, as a stated “±0.3 C resolution” of the microwave sounders does not imply their accuracy.

So, sensing the airmass may have other factors that make ‘reading’, sensing it such that ambiguity gives an error range a bit larger than platinum sensors used in radiosondes?

Reply to  _Jim
August 10, 2026 1:35 pm

Instrumental resolution is the base uncertainty. Instruments have detection limits determined by their construction.

Typically, resolution defines a rectangular uncertainty – a pixel within which there is no information. The pixel center, plus or minus the pixel half-width, determines the minimal uncertainty in a measurement. It cannot be reduced.

In a laboratory-grade 1C/division LiG thermometer, 2σ resolution is about ±0.2 C.

Biases can be detected only if one has a high-accuracy reference standard to which the measurement can be compared.

Typically, in a field measurement, the physically true magnitude of the measurand cannot be known. Which means whatever bias there may be also cannot be known.

Systematic errors can be caused by uncontrolled external variables. Such errors are virtually never random. Instrumental field calibration provides an estimate of the plus/minus measurement uncertainty.

Reply to  Pat Frank
August 10, 2026 4:05 pm

The cited NBS/NIST doc goes into this, but thanks.

(Of course, no on actually read that doc I cited either, so I understand the need that some have to make a post?)

Reply to  Pat Frank
August 9, 2026 7:46 pm

The measurement uncertainty is more than +/- 0.5C. The “hot” calibration source never gets calibrated and it *does* drift. That adds measurement uncertainty to the measurements. The path loss the “emissions” encounter cannot be determined on a “per measurement” basis. It’s just like clouds, climate science parameterizes a value to use. This adds more measurement uncertainty. The satellites are cross-calibrated meaning any measurement uncertainty in one gets propagated throughout the entire system. (of course climate science uses the old “random, Gaussian, and cancels” meme to ignore this propagation of measurement uncertainty throughout the system.

My guess is that the satellite data is just like the Argo data. A measurement uncertainty of +/- 1.0C. It all adds up. It doesn’t cancel. And it is *not* sampling uncertainty, no dividing by sqrt(n).

It’s at least an interval wide enough that the actual average temperature delta’s are all part of the Great Unknown. I sure as hell wouldn’t design a bridge span using beam measurement protocols this imprecise.

Reply to  Tim Gorman
August 10, 2026 1:50 pm

You’re right, Tim. I was shooting only for a lower limit.
I’ve brought the problem up to Roy Spencer, but he just won’t hear it.

Reply to  Richard Mott
August 9, 2026 8:41 am

Mott — what’s in a legendary name?

[Then] what is it that the UAH microwave sounding instruments measure?

It strikes us as something of a miracle, to achieve (absolute) thermometry by microwave spectroscopy, in a remote-sensing mode of operation, using nature’s own molecules in their natural (high, ~ 1/5th) abundance.
These are molecules of ‘ordinary’ di-oxygen, its magnetic ground-state is ‘triplet-sigma’, and has an unusually large intrinsic splitting of ~ 60 GHz [ ~ 0.5 millimeter optical wavelength], itself weakly coupled (hence displaced spectrally) to each molecular rotational states — these being ‘populated’ according to the thermal (‘Boltzmann’) distribution, right up to the relevant ~ 300 K (~ 200 cm-1) energy levels.

bdgwx
Reply to  Richard Mott
August 9, 2026 11:12 am

so if local temperature assumes a small volume, what is it that the UAH microwave sounding instruments measure?

In a nutshell…brightness temperature of O2 microwave emissions. In the case of the MSUs the source code computes the brightness temperature from the raw radiometer counts using a proprietary model that maps the counts to a brightness temperature. In the case of the AMSUs the source just ingests it directly since the AMSUs provide it in that form already.

Once the brightness temperature has been computed or acquired for each channel and for each viewing angle it is assigned to one of the 10368 grid cells (of which only 9504 are actually used). Each grid cell can have up to 12/30 brightness temperatures (6/15 view angles * 2x per day for MSU/AMSU) for each satellite and channel assigned to it. These values are accumulated through the month. At the end of the month a quadratic model is fit to the brightness temperatures. The final monthly layer temperature for that grid cell is then estimated via interpolation using the quadratic model determined in the previous step.

Once the grid has been built the global average temperature is computed as the weighted area average of the 9504 cells. Since UAH publishes the average temperature as global spanning 90S-90N the remaining 10368 – 9504 = 864 unused cells are effectively assumed to behave like the average of the 9504. Even though the 864 represent 8% of the cells because they are located at the poles they only represent about 1% of the area so this shouldn’t be a significant concern.

There is a special case for the TLT layer. It is computed via the simple equation LT = 1.538 * MT – 0.548 * LT + 0.01 * LS so there is one additional model in the whole model chain that computes the temperature we most frequently discuss.

I have the source code downloaded so if you want more details on how this is done I can do my best to extract the core ideas from the code.

Reply to  bdgwx
August 9, 2026 7:51 pm

 proprietary model”
“quadratic model”

And not a single measurement knows what the path loss is for that “brightness”. So each model uses a GUESS at an average that is used as a parameterization. Hint: water vapor affects Ghz frequencies SIGNIFICANTLY. And the water vapor in any specific measurement path is highly variable. Using a parameterization is just like assuming CO2 is well-mixed in the atmosphere so a “constant” value can be used for it. It introduces a LARGE measurement uncertainty.

Philip Mulholland
August 9, 2026 7:24 am

Andy,
Thank you for clarifying the distinction between the strict equilibrium thermodynamic definition of temperature and the operational concepts actually used in the physical sciences. Your emphasis on local temperature as the working foundation for meteorology and hydrodynamics is particularly useful.
In the Dew-Point Anchor Hypothesis, the surface dew-point temperature (via the LCL) is treated as the primary observable thermodynamic boundary condition for the convective column. Modelling proceeds from local temperatures, moist-adiabatic processes and hydrostatic balance rather than any assumption of global thermodynamic equilibrium. The stochastic Markovian implementations explore stationary distributions of those local states across ascent and descent regimes, and the resulting profiles remain physically meaningful.
A central practical distinction arises when these local, dynamic temperatures are compared with the static radiative-equilibrium frameworks that dominate the standard climate narrative. Radiosonde data record a continuously evolving, convective atmosphere in which temperature is measured under real non-equilibrium conditions. By contrast, many climate-model constructions still rely on idealized radiative equilibrium as the organising principle, with surface temperature treated as a dependent response to top-of-atmosphere energy balance. These two approaches are not interchangeable: one is grounded in the measured behaviour of a dynamic system; the other is an equilibrium idealisation. Your discussion of temperature definitions helps make that difference explicit.
Philip Mulholland

Reply to  Philip Mulholland
August 9, 2026 7:53 pm

 rather than any assumption of global thermodynamic equilibrium.”

Kudo’s. Finally – a realistic, real world approach not based on averages and parameterizations.

Curious George
August 9, 2026 8:28 am

A nice trick – to define temperature, you have to define internal energy, entropy, and then somehow measure a partial derivative. Not my way.

Jeff Alberts
August 9, 2026 8:31 am

But it happened on twitter”

So it was three years ago?

August 9, 2026 8:33 am

Nice article, Andy, but I’ll offer this additional comment that I believe was not addressed:

Your Equation 1 mathematical-scientific definition of temperature does not apply across changes of state that involve changes of phase of a substance, such as vaporization/condensation or melting/freezing which occur at constant sensible temperature.

As a simple example, water vapor at 75 deg-F and 14.7 psia has an internal energy (U) of 1036 Btu/lbm (2409 kJ/kg), compared to liquid water at 75 deg-F and 14.7 pisa having an internal energy of 43 Btu/lbm (100 kJ/kg), about a 24:1 difference at the same temperature. And please note that each of these phases can exist in an equilibrium state in the absence of any additional heat transfer.

bdgwx
August 9, 2026 8:37 am

Defining Temperature

This is a great article idea.

The problem with some climate models is that they assume “local equilibrium” for volumes and time periods that are too large and too long.

When does it get too large or too long?

There are many other temperature definitions in common usage as well as in science. Let’s look at some of them:

Yeah, there are many for sure. I’m not sure how deep in the weeds you want to go with this. If you want to include any equation where T appears and then solve for T we could probably list dozens.

Anyway, one notable definition that arises frequently in climate discussions is the blackbody temperature T = (F/σ)^(1/4).

Reply to  bdgwx
August 9, 2026 8:04 pm

When does it get too large or too long?”

When you try to define the equilibrium for the entire globe. And you try to do it over a 24 hour period.

If climate science actually wanted to attempt an equilibrium measurement they would have all measurements taken at 0000GMT everywhere on the globe.

 blackbody temperature”

The earth is not a black body. It is not isothermal. Its radiation is not the Plank curve.

The emissivity of the earth on an operational basis is a ratio (actual emission spectrum)/(black body spectrum). The emissivity can be constant while the Temperature varies widely at different points on the surface. In fact, the emissivity of the surface of the earth varies widely as well. Climate science does it’s typical “assume an average emissivity” and calculate an average Temperature for the globe from that.

It’s a joke. The uncertainty is so large from such averaging and parameterization that it is impossible to find the signal within the noise.

strativarius
August 9, 2026 9:41 am

Climate science – aka the science – incorporates a new form of [Orwellian] newspeak: an official ‘language’ of climate scientists which has been devised to meet the ideological needs of the UN (IPCC), and Global Socialism (one world government).
One early trail blazer of the new art was a certain Dr Stephen Schneider:

So we have to offer up scary scenarios, make simplified, dramatic statements, and make little mention of any doubts we might have. This ‘double ethical bind’ we frequently find ourselves in cannot be solved by any formula. Each of us has to decide what is the right balance between being effective and being honest.” —Dr. Stephen Schneider, former IPCC Coordinating Lead Author, APS Online, Aug./Sep. 1996

Expect some “scary scenarios“, my friends….

Reply to  strativarius
August 9, 2026 1:46 pm

What does your comment got to do with temperature? Go post it in the Open Thread.

Tom Shula
August 9, 2026 9:42 am

Andy, upon reading this post I was having a hard time understanding the point you were trying to make. Without context (you did not provide a link to the X exchange) it seemed like a lot of rambling, some that I would take issue with.

I found Johnathan Cohler’s response to your article here:

https://x.com/cohler/status/2086164299999756394?s=46&t=P4MClTGmb1Wkp5Dsb3NRdw

At a foundational physics level, which is what matters in the end, Johnathan is correct here.

It is not what you know that gets you into trouble. It is what you “think you know” and what you “don’t know that you don’t know.”

Tom Shula
Reply to  Andy May
August 9, 2026 11:47 am

Andy, I find it interesting that you assert “he is clearly wrong.” Apparently you have not understood his response to this article which is what I linked in my post. There he provides quite succinct definitions of the various quantities in physics that are properly considered temperatures, as well as a description of the various faux “temperatures” used in some fields.

I don’t agree with everything Cohler has to say, but he has done a lot of work regarding the concept of temperature and its proper use. He is clearly more expert in the area than you at this time.

It appears that you are also missing the distinction between defining temperature, which is done by relating it to fundamental physical constants and quantities that may be difficult to measure, and the measurement of temperature which is a phenomenological process using an instrument that is designed to respond to the macroscopic properties of what is being measured.

You are a talented writer. Your article is “warm and fuzzy” and no doubt appeals to much of your base. When Cohler is asking for a definition he expects scientific rigor. Whether you failed to recognize this or you are not equipped to provide the rigor he did does not change the circumstance. Science needs to be rigorous and clear. Your article was neither.

Reply to  Andy May
August 9, 2026 2:11 pm

Excellent response! This bears repeating, especially what is in bold text:

“Both of you forget temperature is not a primitive property like mass; it has no meaning except relative to other temperatures. There are multiple ways to define it and measure it.”

To emphasize the point: there is no means to measure “temperature” other than by the indirect effects it has on other properties of matter, such a volumetric expansion, bimetallic flexing, electrical resistance, thermoelectric voltage, and EM radiation emissions.

The fundamental reason this is true is that “temperature” is a statistical average of atomic/molecular movements, meaning it has no independent physical dimension for an individual particle in isolation. By the theory of relativity, the “velocity” of an isolated atom or molecule in an inertial reference frame cannot be established, thus the equation 1/2*m*V^2 = 3/2*k*T (governing “temperature” due to kinetic, translational energy only) has no meaning for that isolated particle.

Reply to  ToldYouSo
August 9, 2026 8:12 pm

The fundamental reason this is true is that “temperature” is a statistical average of atomic/molecular movements”

How does this definition, in any way, allow for the latent heat in moist air?

We know that latent heat exists and we know it has an effect in the transport of heat from one place to another.

Yet we can’t measure it, not even statistically by the movement of quantum particles.

“By the theory of relativity, the “velocity” of an isolated atom or molecule in an inertial reference frame cannot be established,”

So you do what Planck did. Do your analysis using a definition for volume that contains sufficient particles (Planck called them oscillators) to do the analysis. Planck did this so he could assume spherical radiation per volume instead of directional radiation per particle.



Reply to  Andy May
August 10, 2026 8:02 am

You can determine the heat involved by calculation, but you cannot measure it directly.

Reply to  Andy May
August 10, 2026 10:34 am

You can calculate it but you can’t “measure” it.

Total heat energy – sensible heat energy = latent heat energy.

And even then you have to measure the Δheat. Calorimetry can’t tell you the latent heat in an stable sample because you can’t actually measure total heat energy.

You can measure humidity but that is not measuring latent heat. You still have to do a calculation (or table lookup) to convert humidity into latent heat.

Reply to  Tim Gorman
August 10, 2026 7:19 am

“Yet we can’t measure it, not even statistically by the movement of quantum particles.”

Not true. Scientists can—and do, reference NIST for the theory, techniques and instruments—precisely measure the difference in enthalpy (which includes internal energy by definition) between a vapor phase and a liquid phase at a given material’s phase change temperature and pressure (i.e., latent heat).

Tabulations of the latent heat of vaporization/condensation for a wide range of substances can be found across the Web and in many textbooks for Physical Chemistry and for Properties of Materials.

Reply to  ToldYouSo
August 10, 2026 8:13 am

reference NIST for the theory, techniques and instruments

You cannot measure it directly with an enthalpy meter. It must be determined by other measurements and functions.

Here is a formula, “h = 1.006 × T + w × (2501 + 1.86 × T)” for determining it.
Even “enthalpy meters” use two sensors, one for temp and one for humidity, to internally calculate the enthalpy using a formula. There are no direct measurements.

Reply to  Jim Gorman
August 10, 2026 9:52 am

“You cannot measure it directly with an enthalpy meter.”

Ahhhh . . . you mean just as is the case with so many other aspects of nature.

For example: one cannot measure the mass of the Earth or any other celestial body directly with a “mass meter”.

Another: one cannot measure the energy content of radioactive materials directly with an “energy meter”.

Another: one cannot measure distance to the nearest star directly using a physical ruler/tape measure.

Reductio ad absurdum.

It was you—not me—that chose to insert the word “directly” regarding my reference to NIST “theory, techniques and instruments” for quantifying latent heats of various substances.

Reply to  ToldYouSo
August 10, 2026 9:59 am

For example: one cannot measure the mass of the Earth or any other celestial body directly with a “mass meter”.

Another: one cannot measure distance to the nearest star directly using a physical ruler/tape measure.

Not good analogies. If you had a big enough scale or ruler, you COULD measure these directly.

Another: one cannot measure the energy content of radioactive materials directly with an “energy meter”.

This is appropriate because measuring the potential energy would be a calculated quantity derived from other measurements such as mass and half-life.

Reply to  Jim Gorman
August 11, 2026 8:36 am

“Not good analogies. If you had a big enough scale or ruler, you COULD measure these directly.”

Really? How does one directly measure mass—not weight, a force—but mass?

This interchange has reached my limit for the onset of absurdity.

Reply to  ToldYouSo
August 10, 2026 10:39 am

 difference in enthalpy”

The operative word here is “difference”. I have yet to see an “enthalpy meter” that can tell you the sensible and latent heat total in a stable sample.

Tabulations of the latent heat of vaporization/condensation”

That’s like saying you have a steam table. You can calculate the enthalpy state of a substance but you can’t *MEASURE* it.



Reply to  Tim Gorman
August 11, 2026 8:51 am

You simply fail to recognized that the parameter “latent heat” is a value representing a difference in states of matter . . . not a difference in measurements of total enthalpy for a “stable sample” (your own words).

Reply to  ToldYouSo
August 12, 2026 7:34 am

I’m not sure how this makes a difference operationally.

Enthalpy is itself not an absolute, it is defined as the difference of the sample to a reference, e.g. 0C water.

h_total = h_d + h_w
where h_d is enthalpy of dry air (sensible heat) and h_w is the enthalpy of water vapor (latent heat)

All three are just “differences” from a reference.

Total enthalpy is not actually total internal energy. Enthalpy is a measure of heat energy while total internal energy has multiple components including potential energy.

You still can’t directly measure latent heat. There is no such thing as an “enthalpy meter”.

Reply to  Tim Gorman
August 12, 2026 7:48 am

“Enthalpy is itself not an absolute, it is defined as the difference of the sample to a reference, e.g. 0C water.”

Are you sure? I thought it was total internal energy plus volume times pressure.

Reply to  Bellman
August 13, 2026 3:16 am

Are you sure? I thought it was total internal energy plus volume times pressure.”

A parcel of air has MULTIPLE energy components. Among them are

-internal energy
-work (part of pressure and volume)
-kinetic energy
-gravitational energy
-latent energy

Why are you still trying to lecture people on things scientific when you simply don’t know enough to do so?

Reply to  Tim Gorman
August 13, 2026 5:46 am

I did not lecture you, I asked you a question. And I’m not taking any advise on what subjects I can lecture on from someone who routinely lectures on statistics, climate science and many other subjects he has no understanding of.

You claimed enthalpy was not defined as an absolute value, I think you are wrong. You only have to look at the definition to see that it is absolute.

Where the confusion lies, is that absolute enthalpy is rarely known it used, and you normally use changes in enthalpy, i.e. relative enthalpy. But that isn’t the definition of enthalpy, it’s the definition of a change in enthalpy.

Reply to  Tim Gorman
August 13, 2026 9:23 am

Gravity exists as a field (and Einstein would say that field is nothing more than the curvature of spacetime caused by the presence of mass). There is really no such thing as “gravitational energy”.

Scientists refer to vertical movements within a gravitational field as involving the concepts of “potential” energy and “kinetic” energy of any arbitrary mass particle within that field . . . but NOT that the field itself has energy that is conveyed to matter.

I could insert a comment here about “knowing enough”, but won’t.

Tom Shula
Reply to  Andy May
August 9, 2026 4:22 pm

Ok, Andy. I’ll put the ad hominem attacks in the last paragraph on hold. They are uninformed.

Please clarify with some specificity what in my prior post you are referring to. That is, what that I stated makes no sense and why. Properly, you probably should state that it makes no sense to you.

You made similar statements in our correspondence when we tried to collaborate early in 2024. You must accept the possibility that some of these things may just be beyond your comprehension in your present frame of reference.

Do you still believe that “all matter with a temperature above zero kelvin emits radiation”?

Thank you for the link to the original X exchange. Many tried responding to Cohlers’ “challenge” to provide a DEFINITION, and most failed. In your case, you started with “It depends on the context…”. And then you provided multiple contexts without definitions.

I believe Jonathan’s point is that temperature is a precise concept, not to be taken casually. Those who failed tried to define kinetic/thermodynamic temperature, that which is most commonly used. Other “temperatures” are usually accompanied by a descriptor, such as “radiation temperature.”

In that respect, you missed the point of his post. All others with incorrect responses accepted it and moved on. You, on the other hand, did not.

From my perspective you challenged his premise, he “spanked” you, and you retaliated by digging an even deeper hole for yourself.

A more constructive approach rather than playing the victim card would be to try and understand the significance of his point, and if you don’t understand it ask questions so that you can. This is foundational stuff whether you can appreciate it or not.

Reply to  Andy May
August 10, 2026 7:25 am

+42 intergalactic credits for that comment. Spend them as you wish!

Tom Shula
Reply to  Andy May
August 11, 2026 8:33 am

Andy,

First, let’s stick with thermodynamic temperature, as that is our topic of concern.

You said:

This is how a professional physicist, with any sense, sees it:
“Temperature is simply a measurement, relative to other temperature measurements, and has little meaning without context.”

You are confusing the OPERATIONAL USE of temperature with the PHYSICAL DEFINITION of temperature.

You also said:

This is not controversial. It’s standard

It is not “standard”, it is “conventional”.

By analogy, the standard meter is defined relative to the speed of light in a vacuum. We do not measure distances by measuring how long it takes light to travel that distance, but a meter stick does not define what a meter is.

I understand that you would not have explored what to you probably consider “abstract” concepts in your education. Nonetheless there is a difference between a physical definition and a measurement.

Here is a link to a more proper discussion of what temperature is and what it tells us about the system its measurement is assigned to.

https://pmc.ncbi.nlm.nih.gov/articles/PMC7761766/

Reply to  Andy May
August 12, 2026 7:42 am

temperature is not a physical standard based on physical constants. A meter is based on physical constants – the speed of light and the second.

pV/nR = T

p is not a physical constant
V is not a physical constant

Therefore T is not a physical constant. There are multiple ways to measure T besides pV/nR. They don’t all come out to the same value.

Tom Shula
Reply to  Andy May
August 12, 2026 8:07 am

Apologies for the delayed response. I was traveling yesterday. Based on the short time for your response yesterday and the lack of comment on it, I’m guessing you didn’t read the paper I linked for you.

I went back to the original X thread once again. Everyone who responded to Cohler’s challenge (except you) interpreted his question as referring to thermodynamic temperature. That is conventional and normative. When one hears “temperature” without any qualifiers, that is what is meant. A few got the right answer, others came close.

You, on the other hand, decided to take the “gotcha” tact. You lost.

Temperature (meaning thermodynamic in this context without any qualifiers as intended by Cohler) is not a measurement.

“Temperature” is a state variable of a thermodynamic system in equilibrium.

Temperature itself is not a measurement, the measurement is what we do with temperature. It does not define it.

It has a definition. The definition of (thermodynamic) temperature does not change. Part of that definition includes the requirement that the system be in equilibrium.

All you are doing is taking a fundamental physical concept that is clear within its limits of its applicability and muddying the conversation by throwing in a lot of unrelated chaff in the mix.

Like it or not, Johnathan Cohler is making an important point. Temperature is at the center of the climate debate and it is treated by most as a proxy for energy which it is not in all cases. It is also defined only in thermal equilibrium and the models that manifest the “greenhouse effect” and “radiative forcing”, leading to concepts such as “equilibrium climate sensitivity” and “earth energy imbalance” are all based on models that are in both thermal and radiative equilibrium.

Our planet is a system that constantly is driven toward equilibrium but can never achieve it. Its energy flows are never in balance. That is why weather, and climate, vary.

It is unclear to me what your motivation is to try and discredit Cohler’s claim. He is correct, and no obfuscation (which you are doing to try and make his point less clear) will change that.

BTW, I asked a question in a previous post. Do you still insist all matter at a temperature above 0 K emits radiation? It’s a simple yes or no question.

Tom Shula
Reply to  Andy May
August 12, 2026 8:56 am

So this is a language police issue for you?

You appear to be conflating the dictionary definition of the word “temperature” with the scientific definition of “temperature” as an intensive state variable of a system.

These are two different things. The discussion was intended to be a scientific one. If you cannot discern the difference, I don’t know what to say.

Are you intentionally avoiding answering the question I have presented now two times, or is it simply an oversight?

Reply to  Tom Shula
August 13, 2026 3:34 am

Everything above 0K radiates THERMAL energy. But two different objects at the same temperature may radiate the same flux (an intensive property btw) while *not* being in thermal equilibrium unless they are separate closed systems.

If either or both of the objects are open systems they can exchange mass with their surroundings, meaning each can have different conduction, evaporation, and convection rates. Thus they may not be in thermal equilibrium even at the same temperature.

Temperature is an intensive property. What conversion factor do you use to convert that intensive property value to an extensive property value?

August 9, 2026 12:20 pm

The assumption of equilibrium is only valid for very small volumes ( in in a highly controlled laboratory).

August 9, 2026 12:50 pm

…, and with people who have degrees in physics and other hard sciences!

At least they claim to have that background. How can one verify that when all that you know is their ‘handle,’ such as “Spud” or “Ridiculous Rich is a Lying Loser”? Those are actual ‘handles’ for commenters on Yahoo whom I have interacted with. Despite complaining that the latter is inherently a violation of Community Guidelines, being an intended insult, Yahoo lets him spew inanities and other insults. The news media are not our friends.

Quondam
August 9, 2026 1:53 pm

Curiously, I haven’t seen the expression “integrating factor” in this discussion. In classical thermodynamics (dS = q/T) , 1/T is an integrating factor rendering dS a complete differential (Guggenheim, 1950). The significance of complete is that thermodynamic states are path-independent. Their properties depend only on boundary conditions and their history reduced to entropy.

Statistical thermodynamics is based on the Boltzmann distribution and isothermal systems or infinitesimal (linear) perturbations thereof. Classical thermodynamics may encompass steady-state nonlinear dissipative systems, e.g. a tungsten light bulb. It’s SOP to assume local velocity distributions may obey Maxwell-Boltzmann statistics, but I’m unaware of a proof that local CO2 vibrational distributions share this temperature.

michael fellion
August 9, 2026 3:04 pm

On and on. Temperature is pretty important when it is 110F and one is standing in the sun for 4 hours without a drop to drink. It is even more important when it is -119F and one is freezing. What the temperature is at the north pole is irrelevant or at the bottom of the ocean or at 35,000 ft unless one is there. The relevant temperature of the world as a whole can be calculated, well or bad is simply how one does the calculation, how to apply the on and on to the data. It is important to note if that temperature is increasing over time as well or decreasing. The on and on is simply nonsense otherwise to the subject of world temperature. Most if not all the comments refer to transitory temperatures having nothing to do with climate long term mean temperature changes where most people actually live on the planet.

Reply to  michael fellion
August 10, 2026 3:51 am

The relevant temperature of the world as a whole can be calculated, well or bad is simply how one does the calculation, how to apply the on and on to the data.”

This actually isn’t true. The relevant temperature of the earth can *NOT* be calculated because we don’t actually know with enough granularity how it varies either spatially or temporally. The temperature of the earth is *not* isothermal. The temperature of the atmosphere, the surface, and the sub-surface all vary in a complex manner, T = f(x,y,z,heat capacity, humidity, ….)

It’s like saying the average of T^4 is (T_start – Tend)/2.

 It is important to note if that temperature is increasing over time as well or decreasing”

The measurement accuracy has to be of a magnitude that allows identifying the change. That simply isn’t the case for climate today. Climate science makes several garbage assumptions that allow ignoring this simple truth – e.g. “measurement uncertainty is random, Gaussian, and cancels”.

Phillip Chalmers
August 9, 2026 4:20 pm

From my point of view and understanding, this topic has not been resolved in the negative or the affirmative.
Is anyone else similarly perplexed by the words entropy and enthalpy, indeed by the entire field of thermodynamics.
Quite clearly the field is of immense practical importance and seems to be dealing with reality in a useful and practical way.
This comment section under this topic reminds me of a postgraduate tutorial which turns into a master class in the philosophy of the science of the concept of energy. Could we agree that not only the science is not settled but also the phenomenology is not settled?

Reply to  Phillip Chalmers
August 10, 2026 3:59 am

This comment section under this topic reminds me of a postgraduate tutorial which turns into a master class in the philosophy of the science of the concept of energy”

It’s not a class on “philosophy” of a “concept”. It’s a class on how the real world works and how to measure it and understand the measurement data, what it tells you and what it doesn’t.

Could we agree that not only the science is not settled but also the phenomenology is not settled?”

Phenomenology is description without theory. If you don’t understand the underlying mechanisms, then it is impossible to actually predict future states. All you wind up doing is data matching and extending the matching algorithm – which may or may not give reliable estimates.

Phenomenology is not actually settled, as you so aptly state. If your observations are not adequately determined and stated, then the data has limited capability for describing reality.

ntesdorf
August 9, 2026 4:21 pm

Temperature, like the definition of a woman, is a hot topic for some people.

ferdberple
August 9, 2026 5:15 pm

Question. Water temperature varies linearly (near enough) with energy. In a perfectly insulated container it take the same number of calories to raise a mass of water from 10 to 20 C as it does to raise from 40 to 50 C. This let’s you add and subtract temperatures for equal masses of water. However radiation is not linear nor is it mass dependent, it is a 4th power of temperature over area. So how is it we add and subtract radiation to get temperature. The radiation difference between 40 and 50 C is not equal to the radiation difference between 10 and 20 C. Thus the effects of radiation are temperature dependent. It takes less radiation to warm cold objects.

ferdberple
Reply to  Andy May
August 10, 2026 12:23 pm

Heat capacity for liquid water is near constant. It is why i chose water.

Reply to  Andy May
August 11, 2026 5:05 am

Perfect!

Erik Magnuson
August 9, 2026 5:46 pm

When working with NMR, it is possible for the spins to have negative spin temperatures.

Bob
August 9, 2026 6:27 pm

Very nice Andy but this topic needs some serious discussion. The definition that seems to pop up all the time says something like temperature is a representation of the motion in the molecule or something like that. I saw that definition decades ago and think it is as useless today as when I first heard it to the average guy. We absolutely need something the average guy can wrap his head around. If you’re getting the crap burned out of your hand the last thing you’re thinking about is the energy in the molecule.

August 10, 2026 2:35 am

Ooooh! Making my head hurt from remembering Chemical Engineering 201, Thermodynamics I.

Reply to  Andy May
August 10, 2026 3:54 pm

Scatchard-Hamer anyone?

Sparta Nova 4
August 10, 2026 6:30 am

Temperature is not a physical reality.
Temperature is a mathematical model of reality.
A model is an imperfect symbolic representation of reality.

All of math is based on symbols.
1 is a unit quantity. You cannot show me 1. You can show me 1 of something.
I will leave out the rest of the Theory of Mathematics for now.

As discussed in the article, temperature is an expression of underlying energy content.

For atmospheric discussions, I prefer the kinetic energy model.

ferdberple
Reply to  Sparta Nova 4
August 10, 2026 12:35 pm

As discussed in the article, temperature is an expression of underlying energy content.
≈========
Yes, while radiation itself is not energy, it is power. Yet climate science adds and subtracts power and calls it temperature.