Forecasting accuracy of Global Climate Models is something that has been at the very heart of the global warming debate for some time. Leif Svalgaard turned me on to this paper in GRL today:
Reifen, C., and R. Toumi (2009), Climate projections: Past performance no guarantee of future skill?, Geophys. Res. Lett., 36, L13704, doi:10.1029/2009GL038082.
PDF available here
It makes a very interesting point about the “stationarity” of climate feedback strengths. In a nutshell, it says that climate models break down after a time because both forcings and feedbacks don’t remain static, and the program can’t predict such changes.
Gavin Schmidt of NASA GISS says something similar in a recent interview:
The problem with climate prediction and projections going out to 2030 and 2050 is that we don’t anticipate that they can be tested in the way you can test a weather forecast. It takes about 20 years to evaluate because there is so much unforced variability in the system which we can’t predict — the chaotic component of the climate system — which is not predictable beyond two weeks, even theoretically. That is something that we can’t really get a handle on.
From Edge: THE PHYSICS THAT WE KNOW: A Conversation With Gavin Schmidt [with video]
Some excerpts from the paper:
The principle of selecting climate models based on their agreement with observations has been tested for surface temperature using 17 of the IPCC AR4 models.
…
There is no evidence that any subset of models delivers significant improvement in prediction accuracy compared to the total ensemble.
With the ever increasing number of models, the question arises of how to make a best estimate prediction of future temperature change. The Intergovernmental Panel on Climate Change (IPCC) Fourth Assessment Report (AR4) combines the results of the available models to form an ensemble average, giving all models equal weight. Other studies argue in favor of treating some models as more reliable than others [Shukla et al., 2006; Giorgi and Mearns,
2002]. However, determining which models, if any, are superior is not straightforward. The IPCC comments:
‘‘What does the accuracy of a climate model’s simulation of past or contemporary climate say about the accuracy of its projections of climate change? This question is just beginning to be addressed. . .’’[Intergovernmental Panel on Climate Change, 2007, p. 594].
One key assumption, on which the principle of performance-based selection rests, is that a model which performs better in one time period will continue to perform better in the future. This has been studied in terms of pattern-scaling using the ‘‘perfect model assumption’’ [Whetton et al., 2007]. We examine the question in an observational context for temperature here
for the first time. We will also quantify the effect of ensemble size on the global mean, Siberian and European temperature error. [3] The principle of averaging results from different
models to form a multi-model ensemble prediction also has potential problems, since models share biases and there is no guarantee that their errors will neatly cancel out. For this reason groups of models thus combined have been termed ‘‘ensembles of opportunity’’ [Piani et al., 2005]. Various studies have showed that multi-model ensembles produce more accurate results than single models [Kiktev et al., 2007; Mullen and Buizza, 2002]. Our examination of ensemble performance aims to address the question in the context of the current generation of climate models.
…
In our analysis there is no evidence of future prediction skill delivered by past performance-based model selection. There seems to be little persistence in relative model skill, as illustrated by the percentage turnover in Figure 3. We speculate that the cause of this behavior is the non-stationarity of climate feedback strengths. Models that respond accurately in one period are likely to have the correct feedback strength at that time. However, the feedback strength and forcing is not stationary, favoring no particular model or groups of models consistently. For example, one could imagine that in certain time periods the sea-ice albedo feedback is more important favoring those models that simulate sea-ice well. In another period,
El Nino may be the dominant mode, favoring those models that capture tropical climate better. On average all models have a significant signal to contribute.
…
While the authors of this paper still profess faith in model ensembles, the issues they point out with non-staionarity call into question the ability for any model to remain on-track for an extended forecast period.

Maybe it’s time to do what stock-market observers sometimes do: Pick future performance with a dart board and compare with the accuracy of “official” predictions. Or, better, write some “simulation code” which is driven purely by a suitably weighted pseudo-random number generator.
Let me get this straight. One of the major proponents of AGW claims that the current climate models will invariably fail the farther into the future their predictions are extrapolated because they cannot predict accurately? (Who knew?) Yet they expect us to commit trillions to their cause and destroy our economies and ways of life? And they call us delusional?
Are some of the warmists finally getting aroud to admitting that AGW is reductio ad absurdum?
Ah, well. They use models to show something that is inherent from first principles in the way models are constructed.
Ever since two years, when I started reading into this mess of models, I have been saying that there are inherent problems in the construction of the models.
Ignoring for the moment the large problems of data sampling and approximations, which are another large factor, and the dubious concept of “forcings” that contorts physics, I will stress, and have been stressing, the nonlinearity of the solutions of the differential equations used for the modeling.
If one looks at the structure of the models it is obvious that linear approximations to the solutions of the fluid equations are used all over, explicitly and implicitly ( for example average values are the first order term in the expansion of a well behaved solution in a power series) . Linear approximations of solutions can be used if the solutions are expandable in a perturbative expansion, where the next highest term is always much smaller than the previous. This is not true in the case of climate. The solutions are notoriously nonlinear. Also the implied solutions of equations not even considered but where the average value has been used, will also diverge from reality after a similar interval. Hence the whole construct will fail after a number of time steppings.
In the case of GCM used for weather, we see that a week or two weeks at most make the projections irrelevant. For the climate models the time interval is larger but they still fail, as we have seen, in about a few years.
It is sad that such theoretically obvious conclusions, and I am an experimentalist but these are elementary concepts for users of models, need to go through the rigmarole of model testing to be acceptable to what is “the climate community”.
I disagree that relevant models cannot be constructed. Tsonis et al ( there is a thread here and in CA) have made a start at creating a nonlinear model that takes into account the chaotic nature of weather/climate. That is the way modeling should go, IMO.
Are you telling me that these models can’t tell the future ? Who would have thought that was a possibilty ? Surely no one committing stellar quantities of money to ‘tackling climate change’.
Are they finally admitting that the models may not be that great a prediction tool? Too bad back-pedaling isn’t an olympic event. Models certainly have their place, but there has to be awareness of their short comings and the humility to admit that this is so.
Gavin is checking that the unbolting mechanism on the fire escape is working ok.
He knew all this years ago, but only now publicly admits it?
What happened to ‘Robust’?
Any model defines the relationship between the parameters responsible for change in weather and climate. Its a set of inputs on the one side and output on the other. Lets ignore interactions and feedbacks.
Let us assume that equatorial stratospheric ozone is influential in determining cloud cover and sea surface temperature in the tropics. (Some ENSO models are reported to include elements of this dynamic).
Let us further assume that stratospheric ozone in the tropics depends upon the episodic influx of mesospheric nitrogen oxides into the stratophere via the polar vortexes.
Let us further assume that mesospheric nitrogen oxide concentration varies with solar activity.
If we are not in a position to predict the level of solar activity we can not build a model to show how weather and climate respond to change in ozone levels in the stratosphere.
In this situation we are unable to define the relationship between a critical input and the output we want to predict. Model building is possible, outputs may be predictable so long as stratospheric ozone levels are static, but that’s as far as it can go.
That such fundamental problems with models are being accepted by the likes of Gavin Schmidt is a step in the right direction. However, surely the fact that all 40 models all assume that increased water vapour is a positive feedback, when clearly it is not, makes their linear limitations fairly irrelevant?
“While the authors of this paper still profess faith in model ensembles, the issues they point out with non-staionarity call into question the ability for any model to remain on-track for an extended forecast period.”
I don’t know whether to be awed or disgusted by the sheer persistence and bloody mindedness here.
In the face of even their own evidence, there still seems to be a fundamental belief that aggregating any number of clearly flawed models with no skill will magically produce an average with real predictive power.
You’d hope that at some stage someone might realise that the whole concept of “400 wrongs make a right” is not a valid avenue for public funding, let alone public policy. Sadly, in these times of over regulation and institutionally bloated government, somebody predicting something/anything is considered a vital grease for decision making.
The realism of uncertainty is policy making taboo I’m afraid.
However, surely the fact that all 40 models all assume that increased water vapour is a positive feedback, when clearly it is not, makes their linear limitations fairly irrelevant?
Why do you say “clearly not”? I’d agree that observations suggest a lower feedback than that which is evident in the models, but I’m not sure it’s totally “clear” yet.
Who would have predicted a frost in eastern Newfoundland on July 8 … now my veggies are dead 🙁
Andrew P (01:03:36) :
However, surely the fact that all 40 models all assume that increased water vapour is a positive feedback, when clearly it is not, makes their linear limitations fairly irrelevant?
Not really. If linearity were correct it would mean that there was a chance to get a model to fit the past data and project correctly in the future. I am saying there is no such chance anyway by construction.
I see that they glibly are talking of ensemble errors, whereas it is demonstrated that there are no true errors calculated for the model samples: errors from varying by 1 sigma ( true error) of the parameters entering the fits. If you do that for albedo, for example, the fits go all over the place. Error is this virtual reality construct of climate modelers :
Error in the ensemble
mean decreases systematically with ensemble size, N, and
for a random selection as approximately 1/Na, where a lies
between 0.6 and 1.N
from the abstract of the paper.
Why bother? Just throw them chicken bones.
James Griffiths (02:20:19) :
I agree. The use of statistics to models and their ensembles is dangerous. Let’s dig this with a very simple climate model that just contains one equation
Temperature change=Sensitivity*ln(CO2 increase factor)
Set the CO2 increase factor to 1.01 according to IPCC:n common exponential 1% yearly growth scenario. Sensitivity value 5.33 fits well with the historical temperature record.
Now let’s create an ensemble. To show how confident we are let’s have runs with sensitivities 5.321..5.5339. If you want to be sure that our predictions fit to the future measurements, we could use 4.8, 5,0, 5.2, 5.4, 5.6. You see the point. Inventing new runs is based on the decisions of the researchers. I admit that you could see ensembles as research groups’ opinion polls but is that a right way to do climate research.
In ensembles you have entirely different models, not just runs. So, we could replace the growth factor with Michal Hammer̈́’s second order polynomial (see Jennifer Marohasy’s blog). Now we have generated more models. Satisfied?
Mother Nature uses just a single very exact model that is based on physics and other sciences. Our goal is to find it.
The stock market allusion is not far off. Technical trading statisticians and climate modelers do basically the same thing: empirically test their prediction models on past scenarios and continually tweak them until their results match the observation in hindsight.
It doesn’t work. Economic data from, say a 4-month period, 50 years ago may match data seen today. But there are other macroeconomic and geopolitical factors that exist today which could not possibly have existed then. How do you predict the impact of those factors? You can’t. Or, you guess and get lucky. It’s all very scientific.
The models AGW uses to sell its policies are engineering models, not physics models.
They will never be accurate until all of the interactions between the variables are known.
AGW promoters have not even defined all of the variables, yet.
“…all of our models have errors which mean that they will inevitably fail to track reality within a few days irrespective of how well they are initialized.” – James Annan, William Connolley, RealClimate.org
“These codes are what they are – the result of 30 years and more effort by dozens of different scientists (note, not professional software engineers), around a dozen different software platforms and a transition from punch-cards of Fortran 66, to fortran 95 on massively parallel systems. […] No complex code can ever be proven ‘true’ (let alone demonstrated to be bug free). Thus publications reporting GCM results can only be suggestive.” – Gavin Schmidt, RealClimate.org
“No complex code can ever be proven ‘true’ (let alone demonstrated to be bug free). Thus publications reporting GCM results can only be suggestive.” – Gavin Schmidt, RealClimate.org
Maybe Dr Schmidt is starting to realise that all politicians are ultimately answerable to their voters and that the tide may be turning (deity of your choice willing).
You don’t then need the crystal-ball gazing and rune-stone reading powers of climate models to understand that, should such a turnaround come to pass, the aforementioned public servants will hang the scientists out to dry; especially if the same academics had not uttered any words of caution or restraint during the hyperbole of earlier times…
Cheers
Mark
There is a fundamental psychology at work with this. People are naturally averse to uncertainty, because it threatens our survival. Going back to the stone age caveman days, if you could not predict with some accuracy where your next meal was coming from, you would have less chance of finding a mate (she wanted a male who could provide for offspring ). The same psychology has been with us ever since, in the form of oracles, fortune tellers, etc. This is no different simply because it is dressed up with fancy mathematics and pretty charts. It’s still the same subconscious striving for certainty about the future. Sorry to say, that we really aren’t any better at prediction today than we were 10,000 years ago.
You have to crawl before you can walk. A skyscraper starts with foundation first. After years and years of study, the 5 day forecast is still filled with much uncertainty. If the short term forecast still has much work to do, how much more so long term forecasts. When a tropical system forms, look at all the uncertainty in the forecast of it. You have to perfect the more simple short term forecast right before you can put faith in the much more complex long term forecast.
I suggest that Demetris Koutsoyiannis has provided a comprehensive analysis of these problems. The relevant papers are on his university homepage here, http://www.itia.ntua.gr/dk
Readers of WUWT could profitably spend months studying Demetris’ published papers. The ones most relevant to the general questions being discussed around the theme of “climate projections: past performance no guarantee of future skill” namely, ‘are future climate projections reliable and do they provide grounds to assess impacts in hydrological processes? And how well does current climate research represent the intrinsic climate uncertainty?’ can be found opposite his statement of this under the heading “climate stochastics”.
I wonder if Gavin Schmidt mentions Demetris’ penetrating analysis and devastating findings and responds to them? If he doesn’t he has not addressed the relevant science.
I wonder if Reifen and Toumi cite Demtris’ work (and that of Cohen and Linns) and builds on the results of both?
I suggest further that any useful discussion of these questions has to have regard to Demetris’ analysis and findings.
Richard Mackey
An “ensemble” is when I use Langrangian and Eulerian Finite Element Models, each with a different formulation of the equation of state, to model the behavior of a system. I compare the results, and after analysis (a process occurring in a human brain), make a prediction about how the real world will behave.
Then I run a test or experiment.
After I’ve repeated the simulate-analyze-test cycle a few times, and the code results are doing a reasonable job of predicting test outcomes, I can claim that the codes are useful for making predictions of real world behavior of the subject system.
This is not what our climate modeling friends are doing. Their activity is political agenda driven numerology. I’ll also add, that millions of scientists and engineers working in the pharmaceutical and defense industries would wind up fired (if we’re lucky) or in jail (if we’re not) for pulling the fraudulent crap these clowns have. Which is why, among my colleagues, it is rare to find anyone who gives AGW more than a snicker.
Hi Anthony
Off-topic, so please forgive the intrusion. Not sure how to contact you hence this, but one of the readers emailed me a week or so back to point out a reply where you said you had not seen a review copy of my new book Air Con yet.
This was my fault, as we’ve been securing an on-demand print facility in the US and UK and wanted to ensure we could reliably supply Amazon and all orders within a few days. We didn’t want to rattle anyone’s cages too much until we knew people could access Air Con swiftly.
I’m happy to announce that as of the start of this week Air Con is now officially available on tap from distributors in the US, Canada and Europe, which means I can finally send you a review copy, if you wish to flick me a postal address via email to editorial@investigatemagazine.com I’ll get one sent immediately.
Regards
Ian Wishart
PS, to you and all those who post and comment here, this website is truly a credit to collective wisdom and the ability to maintain a healthy skepticism in the face of AGW spin
REPLY: Ian, thank you for the kind words, I’ll be happy to take a look. Email is on the way to you – Anthony
On a slightly different tack, there is an interesting piece in the July 2nd edition of Nature magazine, which describes statements of James Hansen of about a decade ago, in an article about soot/particulate carbon’s role in climate change (which apparently might now be more than thought years ago).
10 years ago, Hansen said that ‘we can address soot emissions now, whereas we can’t address carbon dioxide emissions’.
Now the implication of that for any sane politician should be this: well, right now, we’ll address the soot issue as a near term research goal leading to implementable policy, whereas we’ll need to fund carbon capture/storage technology research until we find a way to bring that within range.
What was the article this week about? Scientists ‘GETTING EXCITED’ about research into soot control.
So here we are: a decade after something addressible was identified, scientists are talking about RESEARCH?
And we wonder why such skepticism exists around this entire field?????
Any chance of one of your experts checking firstly that I got that interpretation of the nature article right and, if so, what their take on policy direction should be in that arena??
The models seem reasonably accurate when they are hindcasting – running the models against the known temperature record.
But they have not shown any skill in predicting the future climate.
They have only been producing actual predictions for about 20 years now, going back to Hansen’s ABC predictions from 1988.
Hansen’s Scenario B, which used input assumptions that are very close to what actually happened, and this prediction is way off.
http://img4.imageshack.us/img4/5277/hansenscenariobandc.png
The predictions from the IPCC’s First Assessment Report in 1990 are way off (it had temps at about +0.8C by now). The predictions from the IPCC’s Second Report from 1996 are closer since they dropped the warming prediction to +2.0C by 2100 (which is now considered too low even though the trendline to date is close).
Here are the predictions made by the IPCC’s Third Assessment Report in 2000 (at least the climate model predictions that are available from the Climate Explorer) – off by 0.25C in just 9 years.
http://img213.imageshack.us/img213/2509/ippctaraverage.png
Here is the Spaghetti graph of the individual IPCC TAR models. (Hard to put much faith in Spaghetti).
http://img18.imageshack.us/img18/8950/ippctarmodels.png
Here is GISS’s predictions submitted to the IPCC’s Fourth Report. The cut-off date for using actual temp records to date was the beginning of 2006 so GISS is off by 0.25C in just 3 years.
http://img189.imageshack.us/img189/7442/gissar4forecasts.png
So, yeah, they can hindcast the climate after tweaking and plugging and knowing what the actual climate has done – Hindcasting models of any type are known for being accurate – the newly reconstructed financial market models are accurately predicting the market meltdown in September for example.