Climate projections: Past performance no guarantee of future skill?

crystal_ball2

Forecasting accuracy of Global Climate Models is something that has been at the very heart of the global warming debate for some time. Leif Svalgaard turned me on to this paper in GRL today:

Reifen, C., and R. Toumi (2009), Climate projections: Past performance no guarantee of future skill?, Geophys. Res. Lett., 36, L13704, doi:10.1029/2009GL038082.

PDF available here

It makes a very interesting point about the “stationarity” of climate feedback strengths. In a nutshell, it says that climate models break down after a time because both forcings and feedbacks don’t remain static, and the program can’t predict such changes.

Gavin Schmidt of NASA GISS says something similar in a recent interview:

The problem with climate prediction and projections going out to 2030 and 2050 is that we don’t anticipate that they can be tested in the way you can test a weather forecast. It takes about 20 years to evaluate because there is so much unforced variability in the system which we can’t predict — the chaotic component of the climate system — which is not predictable beyond two weeks, even theoretically. That is something that we can’t really get a handle on.

From Edge: THE PHYSICS THAT WE KNOW: A Conversation With Gavin Schmidt [with video]

Some excerpts from the paper:

The principle of selecting climate models based on their agreement with observations has been tested for surface temperature using 17 of the IPCC AR4 models.

There is no evidence that any subset of models delivers significant improvement in prediction accuracy compared to the total ensemble.

With the ever increasing number of models, the question arises of how to make a best estimate prediction of future temperature change. The Intergovernmental Panel on Climate Change (IPCC) Fourth Assessment Report (AR4) combines the results of the available models to form an ensemble average, giving all models equal weight. Other studies argue in favor of treating some models as more reliable than others [Shukla et al., 2006; Giorgi and Mearns,

2002]. However, determining which models, if any, are superior is not straightforward. The IPCC comments:

‘‘What does the accuracy of a climate model’s simulation of past or contemporary climate say about the accuracy of its projections of climate change? This question is just beginning to be addressed. . .’’[Intergovernmental Panel on Climate Change, 2007, p. 594].

One key assumption, on which the principle of performance-based selection rests, is that a model which performs better in one time period will continue to perform better in the future. This has been studied in terms of pattern-scaling using the ‘‘perfect model assumption’’ [Whetton et al., 2007]. We examine the question in an observational context for temperature here

for the first time. We will also quantify the effect of ensemble size on the global mean, Siberian and European temperature error. [3] The principle of averaging results from different

models to form a multi-model ensemble prediction also has potential problems, since models share biases and there is no guarantee that their errors will neatly cancel out. For this reason groups of models thus combined have been termed ‘‘ensembles of opportunity’’ [Piani et al., 2005]. Various studies have showed that multi-model ensembles produce more accurate results than single models [Kiktev et al., 2007; Mullen and Buizza, 2002]. Our examination of ensemble performance aims to address the question in the context of the current generation of climate models.

In our analysis there is no evidence of future prediction skill delivered by past performance-based model selection. There seems to be little persistence in relative model skill, as illustrated by the percentage turnover in Figure 3. We speculate that the cause of this behavior is the non-stationarity of climate feedback strengths. Models that respond accurately in one period are likely to have the correct feedback strength at that time. However, the feedback strength and forcing is not stationary, favoring no particular model or groups of models consistently. For example, one could imagine that in certain time periods the sea-ice albedo feedback is more important favoring those models that simulate sea-ice well. In another period,

El Nino may be the dominant mode, favoring those models that capture tropical climate better. On average all models have a significant signal to contribute.

While the authors of this paper still profess faith in model ensembles, the issues they point out with non-staionarity call into question the ability for any model to remain on-track for an extended forecast period.

The climate data they don't want you to find — free, to your inbox.
Join readers who get 5–8 new articles daily — no algorithms, no shadow bans.
0 0 votes
Article Rating
78 Comments
Charlie
July 8, 2009 6:12 am

Richard Mackey says “I suggest that Demetris Koutsoyiannis has provided a comprehensive analysis of these problems. The relevant papers are on his university homepage here, http://www.itia.ntua.gr/dk
Thanks for the pointer. In addition to climate stochastics he also has an interesting editorial in the Hydrological Sciences–Journal–des Sciences Hydrologiques titled “The peer-review system: prospects and
challenges”. August 2005.
http://www.atypon-link.com/IAHS/doi/pdf/10.1623/hysj.2005.50.4.577

July 8, 2009 6:37 am

The public is convinced that climate models are reliable, and that we are all doomed. The news and science media slams them in the head with this message of certain climate doom every day.
Politicians believe it too, and they fund the climate centers like NASA GISS.
It is a circle of doom: Government funded climate scientists –> News and Science media –> Gullible Public –> Politicians –> Government funding for climate

Lazlo
July 8, 2009 6:39 am

“No complex code can ever be proven ‘true’ (let alone demonstrated to be bug free). Thus publications reporting GCM results can only be suggestive.” – Gavin Schmidt, RealClimate.org
The dumbing down of scientific journals, for left wing political motives of course. Happened in the social sciences years ago. Now in so-called climate science.

Hank
July 8, 2009 6:39 am

Here are some modelers from University of Texas. They seem to be focused on abrupt climate change.
http://www.tacc.utexas.edu/research/users/features/climatechange.php

Edward
July 8, 2009 6:41 am

Richard
Here is the team’s and Gavin’s response to Koutsoyiannis:
[Response:Your comments suggest a misunderstanding of a fundamental issue here, namely the distinction between stochastic and deterministic behavior of the climate. There is a fundamental difference between the underlying statistical behavior of climate forcings and the underlying statistical behavior of the climate response to a specified forcing. The stochastic model of AR(1) noise is only ever invoked to explain the unforced component of surface temperature variability. It is nonsensical to attempt to fit a stochastic model to the sum of both unforced and deterministic forced variability. The changes in mean temperature in simulations of e.g. the past 1000 years show that the low-frequency changes in hemispheric mean temperature can be explained quite well in terms of an approximately linear response to changes in natural changes in radiative forcing (see our discussion here, and the additional reviews cited). This is analogous to the fact that the annual cycle in surface temperature at most locations can be described well in terms of an essentially linear response to seasonal changes in insolation. Obviously, the underlying statistical behavior of the forcings themselves on these two timescales is quite different. But it doesn’t matter, from the point of view of understanding the physics of the climate system, what the underlying statistical nature of the variations in forcing is. The response of the system to those changes in forcing is deterministic—any two realizations with small differences in initial conditions will converge, not diverge, in their trajectories with regard to e.g. the global or hemispheric mean temperature, and those trajectories are essentially linearly related to the changes in forcing themselves, at least over the time interval and range of changes in forcing over this timescale. Finally, none of this has any bearing on the statistical description of “noise” present in proxy climate records. It is extremely difficult to reject the null hypothesis of weakly autocorrelated AR(1) red noise in this case. Why one would entertain highly elaborate models of “long-range dependence” and “random walk behavior” when such a simple null hypothesis cannot be rejected, is beyond me. –mike ]
[Response:If I may interject, David’s point is that if the non-climatic ‘noise’ in the proxy series can be modelled as AR(1), what is the likely magnitude of the correlation? He points out that it is much smaller than was recently assumed. Your statements are related to whether the whole series can be modelled as AR(1). These are obviously different issues. However, your statements above seem to imply that climatic series should be thought of as purely stochastic with no deterministic component. This is not likely to be well accepted by most climatologists – because of course would imply that there is no predictability of climate response to any change in external conditions. The fact that many climate changes can be understood in terms of changing solar forcing, volcanic eruptions, greenhouse gas changes, orbital forcing etc. are obvious counter examples to this idea. Therefore the more accepted description is that climate time series consist of a deterministic component together with intrinsic variability and some ‘noise’. In your stochastic descripitions, I am unaware of how you can distinguish these different components, and thus make a claim about the intrinsic variability characteristics. As Mike indicates, I don’t think you can reject the simplest AR(1) hypothesis. – gavin]

Paul Linsay
July 8, 2009 6:46 am

The resort to ensembles of models shows how unscientific this entire bunch is. Each model embodies a different set of physical assumptions about how the climate works, otherwise why have all those models in the first place? At best one is right and the rest are wrong, though it’s certainly possible for them all to be wrong. They clearly do the averaging because none of them can decide which, if any model, is correct. How does any kind of averaging improve the predictions of a bunch of incorrect models, and if one of them by chance is correct, why isn’t its predictions swamped by the bad models?

henrychance
July 8, 2009 6:47 am

tallbloke (00:18:07) :
Gavin is checking that the unbolting mechanism on the fire escape is working ok.
He knew all this years ago, but only now publicly admits it?
What happened to ‘Robust’?
<<<<<<<<<<
I appreciate your clarity in summarizing Schmidt. Also your vividness.

JIm Clarke
July 8, 2009 6:57 am

This information was known 20 years ago! The GCMs only ‘suggest’ future climates based entirely on the assumption that CO2 is a primary driver of global climate. In other words, the models do nothing more than what they were programed to do and have no predictive skills whatsoever! The model output is almost TOTALLY determined by the input assumptions!
The only pertinent question in the debate is the sensitivity of global climate to increasing greenhouse gases. This has always been the ‘only question’ and has almost always been avoided by AGW supporters! All real world evidence shows a low sensitivity and no climate crisis! All observed climate change fits a natural variability pattern and not an anthroprogenic one.
I am beginning to think that Jim Jones had a more compelling argument for drinking his koolaide than AGW supporters have for carbon mitigation. The result of both actions, however, is remarkably similar!

July 8, 2009 7:04 am

Slightly on topic, here’s a link to an article on the Nature.com website illustrating that models are only as good as the data entered. It’s reporting on a paper that examines how AGW could impact the geographic range of Sasquatch. A key point is that “…even if all the data are all highly dubious, a model based on them can still give a plausible-looking result…”.
http://www.nature.com/news/2009/090707/full/news.2009.641.html?s=news_rss
Mike.

henrychance
July 8, 2009 7:09 am

In the course of pharmaceutical research we use testing and apply “double blind” methods to the subjects. We can test the testers. We can eliminate infuence from the testers.
I see a way to test the “model writers” I have conducted experiments in several fields that are not related.
If we took gavin for example, let’s say 10 data sets. Many of the data sets would be not real but look like the real data sets. Give Gavin the sets from assorted time periods and ask him to predict using his model what the next 10 years of mean temp results would look like. we could give him 20 years like from 1928 to 1948, 1957 to 1977 and some other randomly selected 20 year periods. We could include some periods that were totally made up readings but readings within a sensible range.
If we told Mr math wizard what we were doing, he would refuse to “run” the data. There would be too much fear in being wrong. Again to keep the experiment clean, we would not disclose the actual years for the range.
Let me give you a medical example of a non drug nature. People say some are born gay. It is genetic. If that is true, then they would be very comfortible if I brought them some DNA samples and asked them which were gay.

Antonio San
July 8, 2009 7:15 am

“The problem with climate prediction and projections going out to 2030 and 2050 is that we don’t anticipate that they can be tested in the way you can test a weather forecast. It takes about 20 years to evaluate because there is so much unforced variability in the system which we can’t predict — the chaotic component of the climate system — which is not predictable beyond two weeks, even theoretically. That is something that we can’t really get a handle on.” says Gavin Schmidt
Ah the chaotic element of weather and thus climate… Of course these people want anyone to believe that 1) weather is not climate 2) weather is chaotic therefore you cannot use weather to predict climate. Yet when climate models run, they all predict weather…
This is the classic argument that was debunked by the late Marcel Leroux: “Observation of concrete reality suppresses the so-called border between meteorology and climatology, between weather and climate”. And he proves it. Weather is highly regulated and logical and thus its evolution can be used to predict climatic trends. This is of course weather in a slightly different sense than “is it going to rain 5mm in this county?” type of predictive value. But meteorology offers observation based rebuke to the AGW theory and that is why in Schmidt’s viewpoint weather has to be and stay chaotic, unpredictable thus unusable.

Chris
July 8, 2009 7:19 am

Joe Weizenbaum on models… he knew just a tad about computer models:
What is important in the present context is that models embody only the essential features of whatever it is they are intended to represent. … What aspects of reality are and what are not embodied in a model is entirely a function of the model builder’s purpose. But no matter what the purpose, a model, and here I am concerned with computer models of aspects of reality, must necessarily leave out almost everything that is actually present in the real thing. Whoever knows and appreciates this fact, and keeps it in mind while teaching students about the use of computers, has a chance to immunize his or her students against believing or making excessive claims for much of their computer work.
(Weizenbaum 1984, xvii)
Weizenbaum, J. (1984). Computer Power and Human reason. From Judgement to Calculation. Harmondsworth, Middlesex: Penguin.

July 8, 2009 7:35 am

Someone said it before – that anyone who thinks he/she can model something as complex as the earth’s climate simply does not understand the complexity of the earth’s climate.
The earth’s climate is not a game of chess – chess programmes can now beat the best human players – because there are a finite number of chess moves, albeit a very large finite number, but finite nonetheless.
Others have also said recently – and I agree – that GS is looking for a way out. I think he wants to be the first into the lifeboats.

Demesure
July 8, 2009 7:38 am

The notion of the average of untested models is junk science.
It’s like saying because the mean of 21 meteo models is 20 °C for next week, so the “ensemble prediction” would be better than any invidual model. It’s not only theorical nonsense, it’s proven false.

J. Bob
July 8, 2009 8:08 am

Hunter – These are not even engineering models. Engineering models MUST reflect reality, or standard practices, or you could end up in court.

Jim
July 8, 2009 8:17 am

@anna v (23:56:16) : Modeling the climate with the goal of predicting the temperature some years X in the future is a futile endeavor. If we had a model that took into account all the physics of the Sun and Earth perfectly, it would tell us how the climate behaves. It would give us limits on temperature, precipitation, etc.; and show us generally how it works. But climate is chaotic. That renders any model, no matter how perfect, incapable of predicting the future climate.

James Griffiths
July 8, 2009 8:27 am

Bill Illis (05:56:38) :
“The models seem reasonably accurate when they are hindcasting – running the models against the known temperature record.”
In the case of the climate, I would suggest that hindcasting is a gigantic waste of time.
The averages of temperatures over time and areas that are used in measuring the climate give such little information that at the scales the models work on there must be a practically infinite way of hitting the target, and all but one of those ways are wrong.
If the models could hindcast a temperature and explain accurately all the fluxes and processes involved at the resolution necessary to make an accurate prediction going forwards, then that would be something special! Then again, that wouldn’t really be a model, it would be more of an actual backwards running earth!
I’m waffling a little, but my point is I imagine there’s an awful lot of ways to tune a model to the past that work perfectly. Even if you get it right, chaos theory suggests even the tiniest difference in resolution of any input means you won’t recreate the actual sequence of events anyway!

July 8, 2009 8:28 am

“People say some are born gay. It is genetic. If that is true, then they would be very comfortible if I brought them some DNA samples and asked them which were gay.”
/rasp. Why did you even bring this up? It’s not relevant to your point.
It’s quite possible to deduce that a environmental influence is a factor in some end effect without knowing what the direct mechanism is.
Eye color is undoubtedly genetic, but we can’t (yet) examine a set of DNA samples and determine the eye color of each.
Mike.

July 8, 2009 8:35 am

Climate Modelers are basically dowsers. When they know where the target is (past climate), they’re always “accurate”. But when presented with a blind test (future climate), they do no better than chance would dictate.

hunter
July 8, 2009 8:38 am

Hindcasting is simply knowing the answer to the test before you take it.

theduke
July 8, 2009 8:42 am

Next step ahead for Warmists: introduction of models that purportedly account for non-stationarity?
One step ahead of the posse . . .

Poptech
July 8, 2009 8:43 am

Testing a model against past climate (hindcasting) is an advanced exercise in curve fitting, nothing more and proves absolutely nothing. What this means is you are attempting to have your model’s output match the existing historical output that has been recorded. For example matching the global mean temperature curve over 100 years. Even if you match this temperature curve with your model it is meaningless. Your model could be using some irrelevant calculation that simply matches the curve but does not relate to the real world. With a computer model there are an infinite number of ways to match the temperature curve but only one way that represents the real world. It is impossible for computer models to prove which combination of climate physics correctly matches the real world.
Virtual reality can be whatever you want it to be and computer climate models are just that, they are the code based on the subjective opinions of the scientists creating them. The real world has no such bias.

Russ R.
July 8, 2009 9:25 am

A few years back, when a good portion of the arctic ice was blown into warmer waters, the folks over at RC, were risking shoulder dislocation, to pat themselves on the back, at their amazing predictive abilities.
I posted that it was a short time period, based on a short timescale, and that they no real idea if it was an unusual event, or a natural event, or temperature driven, or due to wind, or warmer water temps. I compared it to having a few good holes in golf, and then extrapolating those results to a round.
Gavin explained to me the error of my ways. He said it was more like rolling balls down an incline, and using the results to predict future rolls on different inclines.
I am not hearing the same Gavin, I heard back in those glory days. I am sure a few of these “climate gurus”, have decided it is time to engineer a soft landing.

July 8, 2009 9:35 am

One issue vis-a-vis hindcasting: Curve fitting to WHAT?
Curve fitting to GISS will cause you to tune the model for much greater sensitivity than curve fitting to UAH or RSS.
Since we now have 30 years of satellite data (climate!), it would be interesting to see how the various models fit that data, even in hindcast…

July 8, 2009 9:52 am

Ian Wishart (05:19:43) :
Good to see you here Sir!
I’ve been looking for your book here in Thailand where I live but I haven’t found it yet…
The Thais don’t have much concept of climate change: they don’t like the temperature above 30C (86F) or below 27C (79F)!