Saturday, December 12, 2009

Stealthy UAVs and Drone Wars

USAF Stealth UAV Has Ties To Previous Designs.

Aerospace Daily and Defense Report (12/10, Fulghum) reported the recently-revealed USAF stealth RQ-170 has "linkages to earlier designs from Lockheed Martin's Advanced Development Programs, including the stealthy DarkStar and Polecat UAVs." The RQ-170 has a "tailless flying wing" featuring communication sensors and is currently serving in Afghanistan.

Drones Being Used Successfully In Pakistan.

On its "Danger Room" blog, Wired (12/10, Schactman) reported on the US military's use of drones in the Afghanistan and Pakistan. According to some estimates, drones have killed "as many as a thousand people" around Pakistan. While America is currently not invading Pakistan, the drones are allowed to pursue militants as long as "the government in Islamabad [is] notified first." These drone strikes are "widely credited for taking out senior leaders of both the Pakistani Taliban and Al Qaeda," but have been criticized as an "extension of the war in Central Asia fought under uncertain authority and with questionable morality."



That "questionable morality" comment is just a smear, but it is interesting stuff anyway.

Friday, December 11, 2009

Lord Monckton on Climategate

Pointed at by theAirVent:

Lord Monckton on Climategate at the 2nd International Climate Conference from CFACT EUROPE on Vimeo.



Gore isn't the only one who can make propaganda...just for entertainment purposes, don't go re-jiggering your economy based on this video, but enjoy.

Ha ha, traffic light tendency:


Calling those guys from East Anglia crooks is a bit over the top (they are just stealth advocates), but it's still entertaining.

Do you validate?


To sum up:

Indeed, much of what is presented as hard scientific evidence for the theory of global warming is false. "Second-rate myth" may be a better term, as the philosopher Paul Feyerabend called science in his 1975 polemic, Against Method.

"This myth is a complex explanatory system that contains numerous auxiliary hypotheses designed to cover special cases, as it easily achieves a high degree of confirmation on the basis of observation," Feyerabend writes. "It has been taught for a long time; its content is enforced by fear, prejudice and ignorance, as well as by a jealous and cruel priesthood. Its ideas penetrate the most common idiom, infect all modes of thinking and many decisions which mean a great deal in human life ... ".
Times Higher Education -- Beyond Debate?

Dueling Bayesians

Percontations: The Nature of Probability

Interest points:

  • Fun with coin-flipping (13:43)

  • Can probabilistic thinking be completely automated? (04:31)

  • The limits of probability theory (11:07)

  • How Andrew got shot down by Daily Kos (06:55)

  • Is the academic world addicted to easy answers? (11:11)

    This part of the discussion was very brief, but probably the most interesting. What Gelman is referring to is maximum entropy sampling, or optimal sequential design of experiments. This has some cool implications for model validation I think (see below).

  • The difference between Eliezer and Nassim Taleb (06:20)


Links mentioned:


Some things Dr Gelman said that I think are interesting:

I was in some ways thinking like a classical statistician, which was, well,I'll be wrong 5% of the time, you know, that's life. We can be wrong a lot, but you're never supposed to knowingly be wrong in Bayesian statistics. If you make a mistake, you shouldn't know that you made a mistake, that's a complete no-no.


With great power comes great responsibility. [...] A Bayesian inference can create predictions of everything, and as a result you can be much more wrong as a Bayesian than as a classical statistician.


When you have 30 cases your analysis is usually more about ruling things out than proving things.

Towards the end of the discussion Yudowski really sounds like he's parroting Jaynes (maybe they are just right in the same way).

Bayesian design of validation experiments

As I mentioned in the comments about the Gelman/Yudowski discussion, the most interesting thing to me was the ’adaptive testing’ that Gelman mentioned. This is a form of sequential design of experiments , and the Bayesian versions are the most flexible. That is because Bayes theorem provides a consistent and coherent (if not always conveneint) way of updating our knowledge state as each new test result arrives. Then, and Gelman’s comment about ’making predictions about everything’ is germane here, we assess our predictive distributions and find the areas of our parameter space that have the most uncertainty (highest entropy of the predictive distribution). This place in our parameter space with the highest predictive distribution entropy is where we should test next to get the most information. The example of academic testing that Gelman gives does exactly that, the question chosen is the one that the test-taker has equal chance of getting right or wrong.

The same idea applies to testing to validate models. Here’s a little passage from a relevant paper that provides some background and motivation [1]:

Under the constraints of time, money, and other resources, validation experiments often need to be optimally designed for a clearly defined purpose, namely computational model assessment. This is inherently a decision theoretic problem where a utility function needs to be first defined so that the data collected from the experiment provides the greatest opportunity for performing conclusive comparisons in model validation.

The method suggested to achieve this is based on choosing a test point from the area of the parameter space with the highest predictive entropy and also one from the area with the lowest predicitive entropy [2]. This addresses the little comment Gelman made about not being able to asses the goodness of the model very well if you only choose points in the high entropy area. Each round of two test points gives you an opportunity to make the most conclusive comparison of the model prediction to reality.

If you were just trying to calibrate a model, then you would only want to choose test points in the high-entropy-areas because these would do the most to reduce your uncertainty about the reality of interest (and hence give you better parameter estimates). Since we are trying to validate the model though, we want to evaluate its performance where we expect it to give the best predictions and where we expect it to give the worst predictions. Here the idea explained in a bit more technical language [1]:

Consider the likelihood ratio Λ(y) in Eq. (9) [or Bayes factor in [6]] as a validation metric. Suppose an experiment is conducted with the minimization result, and the experimental output is compared with model prediction. We expect a high value Λ(y)min , where the subscript min indicates that the likelihood is obtained from the experimental output in the minimization case. If Λ(y)min < η, then clearly this experiment rejects the model, since the validation metric Λ(y), even under the most favorable conditions, does not meet the threshold value η. On the other hand, suppose an experiment is conducted with the maximization result, and the experimental output is compared with the model prediction. We expect a low value Λ(y)max < η in this case. If Λ(y)max > η, then clearly this experiment accepts the model, since it is performed under the worst condition and still produces the validation metric to be higher than η. Thus, the cross entropy method provides conclusive comparison as opposed to an experiment at any arbitrary point.


Here's a graphical depiction of the placement of the optimal Bayesian decision boundary (image taken from [3]):


It would be nice to see these sorts of decision theory concepts applied to the public policy decisions that are being driven by the output of computational physics codes.

References

[1] Jiang, X., Mahadevan, S., “Bayesian risk-based decision method for model validation under uncertainty,” Reliability Engineering & System Safety, No. 92, pp 707-718, 2007.

[2] Jiang, X., Mahadevan, S., “Bayesian cross entropy methodology for optimal design of validation experiments,” Measurement Science & Technology, 2006.

[3] Jiang, X., Mahadevan, S., “Bayesian validation assessment of multivariate computational models ,” Journal of Applied Statistics, Vol. 35, No. 1, Jan 2008.

Thursday, December 10, 2009

Successful Vortex Hybrid Test

Orbital Technologies announced on 8 Dec 2009 that it successfully static tested its big vortex hybrid rocket motor. What's a 'vortex hybrid'? It's a hybrid that controls the fuel regression rate and combustion stability by swirling or recirculating the oxygen that is injected into the cavity of the fuel grain.

The effect of swirling or recirculating the oxygen injection is seen in the photos below taken from this paper:


And here's a little CFD marketing snapshot taken from
Orbital's vortex hybrid data sheet:


And finally here's a nice photo of the test from the ORBITEC press release:

Wednesday, December 9, 2009

Verification, Validation, and Uncertainty Quantification

Notes on Chapter 8: Verification, Validation, and Uncertainty Quantification by George Em Karniadakis [1]. Karniadakis provides the motivation for the topic right off:

In time-dependent systems, uncertainty increases with time, hence rendering simulation results based on deterministic models erroneous. In engineering systems, uncertainties are present at the component, subsystem, and complete system levels; therefore, they are coupled and are governed by disparate spatial and temporal scales or correlations.

His definitions are based on those published by DSMO and subsequently adopted by AIAA and others.

Verification is the process of determining that a model implementation accurately represents the developers conceptual description of the model and the solution to the model. Hence, by verification we ensure that the algorithms have been implemented correctly and that the numerical solution approaches the exact solution of the particular mathematical model typically a partial differential equation (PDE). The exact solution is rarely known for real systems, so “fabricated” solutions for simpler systems are typically employed in the verification process. Validation, on the other hand, is the process of determining the degree to which a model is an accurate representation of the real world from the perspective of the intended uses of the model. Hence, validation determines how accurate are the results of a mathematical model when compared to the physical phenomenon simulated, so it involves comparison of simulation results with experimental data. In other words, verification asks “Are the equations solved correctly?” whereas validation asks “Are the right equations solved?” Or as stated in Roache (1998) [2], “verification deals with mathematics; validation deals with physics.”

He addresses the constant problem of validation succinctly:

Validation is not always feasible (e.g., in astronomy or in certain nanotechnology applications), and it is, in general, very costly because it requires data from many carefully conducted experiments.

Getting decision makers to pay for this experimentation or testing is especially problematic when they were initially sold on using modeling and simulation as a way to avoid testing.

After this the chapter goes into an unnecessary digresion on inductive reasoning. An unfortunate common thread that I’ve noticed in many of the V&V reports I’ve read is they seem to think Karl Popper had the last word on scientific induction! I think the V&V community would profit greatly by studying Jayne’s theoretically sound pragmatism. They would quickly recognize that the ’problems’ they perceive in scientific induction are little more than misunderstandings of probability theory as logic.

The chapter gets back on track with the discussion of types of error in simulations:

Uncertainty quantification in simulating physical systems is a much more complex subject; it includes the aforementioned numerical uncertainty, but often its main component is due to physical uncertainty. Numerical uncertainty includes in addition to spatial and temporal discretization errors, errors in solvers (e.g., incomplete iterations, loss of orthogonality), geometric discretization (e.g., linear segments), artificial boundary conditions (e.g., infinite domains), and others. Physical uncertainty includes errors due to imprecise or unknown material properties (e.g., viscosity, permeability, modulus of elasticity, etc.), boundary and initial conditions, random geometric roughness, equations of state, constitutive laws, statistical potentials, and others. Numerical uncertainty is very important and many scientific journals have established standard guidelines for how to document this type of uncertainty, especially in computational engineering (AIAA 1998 [3]).

The examples given for effects of ’uncertainty propagation’ are interesting. The first is a direct numerical simulation (DNS) of turbulent flow over a circular cylinder. In this resolved simulation, the high-wave numbers (smallest scales) are accurately captured, but there is disagreement at the low wave numbers (largest scales). This somewhat counter-intuitive result occurs because the small scales are insensitive to experimental uncertainties about boundary and initial conditions, but the large scales of motion are not.

The section on methods for dealing with modelling uncertain inputs is sparse on details. Passing mention is made of Monte Carlo and Quasi-Monte Carlo methods, sensitivity-based methods and Bayesian methods.

The section on ’Certification / Accreditation’ is interesting. Karniadakis recomends designing experiments for validation based on the specific use or application rather than based on a particular code. This point deserves some emphasis. It is an often voiced desire from decision makers to have a repository of validated codes that they can access to support their various and sundry efforts. This is an unrealistic desire. A code can not be validated as such, only a particular use of a code can be validated. In most decisions that engineering simulation supports, the use is novel (research and new product development), therefore the validated model will be developed concurrently with (in the case of product development) or as a result of (in the case of research) the broader effort in question.

The suggested hierarchical validation framework is similar to the ’test driven development’ methodologies in software engineering and the ’knowledge driven product development’ championed in the GAO’s reports on government acquisition efforts. Small component (unit) tests followed by system integration tests and then full complex system tests. When the details of ’model validation’ are understood, it is clear that rather than replacing testing, simulation truly serves to organize test designs and optimize test efforts.

The conclusions are explicit (emphasis mine):

The NSF SBES report (Oden et al. 2006 [4]) stresses the need for new developments in V&V and UQ in order to increase the reliability and utility of the simulation methods at a profound level in the future. A report on European computational science (ESF 2007 [5]) concludes that “without validation, computational data are not credible, and hence, are useless.” The aforementioned National Research Council report (2008) on integrated computational materials engineering (ICME) states that, “Sensitivity studies, understanding of real world uncertainties and experimental validation are key to gaining acceptance for and value from ICME tools that are less than 100 percent accurate.” A clear recommendation was reached by a recent study on Applied Mathematics by the U.S. Department of Energy (Brown 2008 [6]) to “significantly advance the theory and tools for quantifying the effects of uncertainty and numerical simulation error on predictions using complex models and when fitting complex models to observations.”

References

[1] WTEC Panel Report on International Assessment of Research and Development in Simulation-Based Engineering and Science, 2009, http://www.wtec.org/sbes/SBES-GlobalFinalReport.pdf

[2] Roache, P.J. 1998. Verification and validation in computational science and engineering. Albuquerque,: Hermosa Publishers.

[3] AIAA Guide for the Verification and Validation of Computational Fluid Dynamics Simulations, Reston, VA, AIAA. AIAA-G-077-1998.

[4] Oden, J.T., T. Belytschko, T.J.R. Hughes, C. Johnson, D. Keyes, A. Laub, L. Petzold, D. Srolovitz, and S. Yip. 2006. Revolutionizing engineering science through simulation: A report of the National Science blue ribbon panel on simulation-based engineering science. Arlington: National Science Foundation. Available online

[5] European Computational Science Forum of the European Science Foundation (ESF). 2007. The Forward Look Initiative. European computational science: The Lincei Initiative: From computers to scientific excellence. Information available online.

[6] Brown, D.L. (chair). 2008. Applied mathematics at the U.S. Department of Energy: Past, present and a view to the future. May, 2008.


Concepts of Model Verification and Validation has a glossary that defines most of the relevant terms.

Tuesday, December 8, 2009

Simulation-based Engineering Science

gmcrews has a couple interesting posts on model verification and validation (V&V) and his commment on scientific software has links to a couple ’state of the practice’ reports [1] [2]. The reports are about something called Simulation-based Engineering Science (SBES), which is the (common?) jargon they use to describe doing research and development with computational modelling and simulation.

1 Notes and Excerpts from [1]

Below are some excerpts from the executive summary along with a little commentary.

Simulation has today reached a level of predictive capability that it now firmly complements the traditional pillars of theory and experimentation/observation. Many critical technologies are on the horizon that cannot be understood, developed, or utilized without simulation. At the same time, computers are now affordable and accessible to researchers in every country around the world. The near-zero entry-level cost to perform a computer simulation means that anyone can practice SBE&S, and from anywhere.

1.1 Major Trends Identified

  1. Data-intensive applications, including integration of (real-time) experimental and observational data with modelling and simulation to expedite discovery and engineering solutions, were evident in many countries, particularly Switzerland and Japan.
  2. Achieving millisecond time-scales with molecular resolution for proteins and other complex matter is now within reach using graphics processors, multicore CPUs, and new algorithms.
  3. The panel noted a new and robust trend towards increasing the fidelity of engineering simulations through inclusion of physics and chemistry.
  4. The panel sensed excitement about the opportunities that petascale speeds and data capabilities would afford.

1.2 Threats to U.S. Leadership

  1. The world of computing is flat, and anyone can do it. What will distinguish us from the rest of the world is our ability to do it better and to exploit new architectures we develop before those architectures become ubiquitous.

    Furthermore, already there are more than 100 million NVIDIA graphics processing units with CUDA compilers distributed worldwide in desktops and laptops, with potential code speedups of up to a thousand-fold in virtually every sector to whomever rewrites their codes to take advantage of these new general programmable GPUs.

  2. Inadequate education and training of the next generation of computational scientists threatens global as well as U.S. growth of SBE&S. This is particularly urgent for the United States; unless we prepare researchers to develop and use the next generation of algorithms and computer architectures, we will not be able to exploit their game-changing capabilities.

    Students receive no real training in software engineering for sustainable codes, and little training if any in uncertainty quantification, validation and verification, risk assessment or decision making, which is critical for multiscale simulations that bridge the gap from atoms to enterprise.

  3. A persistent pattern of subcritical funding overall for SBE&S threatens U.S. leadership and continued needed advances amidst a recent surge of strategic investments in SBE&S abroad that reflects recognition by those countries of the role of simulations in advancing national competitiveness and its effectiveness as a mechanism for economic stimulus.

I don’t know of any engineering curriculums that have a good program for training people in all of the areas (the physics, numerical methods, design of experiments, statistics and software carpentry) to be competent high-performance simulation developers (in scientific computing the users and developers tend to be the same people). It’s requires multi-disciplinary, and deeply technical knowledge at the same time. The groups that try to go broad with the curriculum tend to treat the simulations as a black box. Those sorts of programs tend to produce people who can turn the crank on a code, but don’t have the deeper technical understanding needed to add the next increment of physics, or apply the newer more efficient solver, or adapt the current code to take advantage of new hardware. Right now that sort of expertise is achieved in an ad-hoc or apprenticeship kind of manner (see for example MIT’s program). That works for producing a few experts at a time (after a lot of time), but it doesn’t scale well.

1.3 Opportunities for Investment

  1. There are clear and urgent opportunities for industry-driven partnerships with universities and national laboratories to hardwire scientific discovery and engineering innovation through SBE&S.
  2. There is a clear and urgent need for new mechanisms for supporting R&D in SBE&S.

    investment in algorithm, middleware, and software development lags behind investment in hardware, preventing us from fully exploiting and leveraging new and even current architectures. This disparity threatens critical growth in SBE&S capabilities needed to solve important worldwide problems as well as many problems of particular importance to the U.S. economy and national security.

  3. There is a clear and urgent need for a new, modern approach to educating and training the next generation of researchers in high performance computing specifically, and in modeling and simulation generally, for scientific discovery and engineering innovation.

    Particular attention must be paid to teaching fundamentals, tools, programming for performance, verification and validation, uncertainty quantification, risk analysis and decision making, and programming the next generation of massively multicore architectures. At the same time, students must gain deep knowledge of their core discipline.

The third finding is interesting, but it is a tall order. So we need to train subject matter experts who are also experts in V&V, decision theory, software development and exploiting unique (and rapidly evolving) hardware. Show me the curiculum that accomplishes that, and I’d be quite impressed (really, post a link in the comments if you know of one).

More on validation:

Experimental validation of models remains difficult and costly, and uncertainty quantification is not being addressed adequately in many of the applications. Models are often constructed with insufficient data or physical measurements, leading to large uncertainty in the input parameters. The economics of parameter estimation and model refinement are rarely considered, and most engineering analyses are conducted under deterministic settings. Current modeling and simulation methods work well for existing products and are mostly used to understand/explain experimental observations. However, they are not ideally suited for developing new products that are not derivatives of current ones.

One of the mistakes that the scientific computing community made early on was in letting the capabilities of our simulations be over-sold without stressing the importance of concurrent, supporting efforts in theory and experiment. There are a significant number of consultants who make outrageous claims about replacing testing with modeling and simulation. It is far to easy for our decision makers to be impressed by the really awesome movies we can make from our simulations, and the claims from the consultants begin to get traction. It is our job to make sure the decision makers understand that the simulation is only as real as our empirical validation of it.

2 Notes and Exerpts from [2]

Below are some exerpts from the executive summary along with a little commentary.

Major Findings:

  1. SBES is a discipline indispensable to the nations continued leadership in science and engineering. […] There is ample evidence that developments in these new disciplines could significantly impact virtually every aspect of human experience.
  2. Formidable challenges stand in the way of progress in SBES research. These challenges involve resolving open problems associated with multiscale and multi-physics modeling, real-time integration of simulation methods with measurement systems, model validation and verification, handling large data, and visualization. Significantly, one of those challenges is education of the next generation of engineers and scientists in the theory and practices of SBES.
  3. There is strong evidence that our nations leadership in computational engineering and science, particularly in areas key to Simulation-Based Engineering Science, is rapidly eroding. Because competing nations worldwide have increased their investments in research, the U.S. has seen a steady reduction in its proportion of scientific advances relative to that of Europe and Asia. Any reversal of those trends will require changes in our educational system as well as changes in how basic research is funded in the U.S.

The ’Principle Recommendations’ in the report amount to ’Give the NSF more money’, which is not surprising if you consider the source. It is interesting that their finding about education is largely the same as the other report.

References

[1] A Report of the National Science Foundation Blue Ribbon Panel on Simulation-Based Engineering Science: Revolutionizing Engineering Science through Simulation, National Science Foundation, May 2006, http://www.nsf.gov/pubs/reports/sbes_final_report.pdf

[2] WTEC Panel Report on International Assessment of Research and Development in Simulation-Based Engineering and Science, 2009, http://www.wtec.org/sbes/SBES-GlobalFinalReport.pdf

Monday, December 7, 2009

Final Causes

This is the conclusion of the comments section in the model comparison chapter (Chapter 20) from Jayne’s book (emphasis original).


It seems that every discussion of scientific inference must deal, sooner or later, with the issue of belief or disbelief in final causes. Expressed views range all the way from Jaques Monod (1970) forbidding us even to mention purpose in the Universe, to the religious fundamentalist who insists that it is evil not to believe in such a purpose. We are astonished by the dogmatic, emotional intensity with which opposite views are proclaimed, by persons who do not have a shred of supporting factual evidence for their positions.

But almost everyone who has discussed this has supposed that by a ’final cause’ one means some supernatural force that suspends natural law and takes over control of events (that is, alters positions and velocities of molecules in a way inconsistent with the equations of motions) in order to ensure that some desired final condition is attained. In our view, almost all past discussions have been flawed by failure to recognize that operation of a final cause does not imply controlling molecular details.

When the author of a textbook says: ’My purpose in writing this book was to…’, he is disclosing that there was a true ’final cause’ governing many activities of writer, pen, secretary, word processor, extending usually over several years. When a chemist imposes conditions on his system which forces it to have a certain volume and temperature, he is just as truly the wielder of a final cause dictating the final thermodynamic state that he wished it to have. A bricklayer and a cook are likewise engaged in the art of invoking final causes for definite purposes. But – and this is the point almost always missed – these final causes are macroscopic; they do not determine any particular ’molecular’ details. In all cases, had those fine details been different in any one of billions of ways, the final cause would have been satisfied just as well.

The final cause may then be said to possess an entropy, indicating the number of microscopic ways in which its purpose can be realized; and the larger that entropy, the greater is the probability that it will be realized. Thus the principle of maximum entropy applies also here.

In other words, while the idea of a microscopic final cause runs counter to all the instincts of a scientists, a macroscopic final cause is a perfectly familiar and real phenomenon, which we all invoke daily. We can hardly deny the existence of purpose in the Universe when virtually everything we do is done with some definite purpose in mind. Indeed, anybody who fails to pursue some definite long-term purpose in the conduct of his life is dismissed as an idler by his colleagues. Obviously, this is just a familiar fact with no religious connotations – and no anti-religious ones. Every scientist believes in macroscopic final causes without thereby believing in supernatural contravention of the laws of physics. The wielder of the final cause is not suspending physical law; he is merely choosing the Hamiltonian with which some system evolves according to physical law. To fail to see this is to generate the most fantastic, mystical nonsense.


So we have the wager from Pascal, and God as the ultimate ’Hamiltonian chooser’ from Jaynes?

Sunday, December 6, 2009

Sociology of Science

This is the first part of the comments section in the model comparison chapter (Chapter 20) from Jayne’s book (emphasis mine).

Actual scientific practice does not really obey Ockham’s razor, either in its previous ’simplicity’ form or in our revised ’plausibility’ form. As so many of us have deplored, the attractive new hypothesis or model, which accounts for the facts in such a neat, plausible way that you want to believe it at once, is usually pooh-poohed by the official Establishment in favor of some drab, complicated, uninteresting one; or, if necessary, in favor of no alternative at all. The progress of science is carried forward mostly by the few fundamental dissenting innovators, such as Copernicus, Galileo, Newton, Laplace, Darwin, Mendel, Pasteur, Boltzmann, Einstein, Wegener, Jeffreys – all of whom had to undergo this initial rejection and attack. In the cases of Galileo, Laplace, and Darwin, these attacks continued for more than a century after their deaths. This is not because their new hypothesis were faulty – quite the contrary – but because this is the part of the sociology of science (and, indeed of all scholarship). In any field, the Establishment is seldom in pursuit of the truth, because it is composed of those who sincerely believe that they are already in possession of it.

The sociology of science is an interesting topic that’s been brought forcefully into the public perception by the recent kerfuffle over the leaked UEA CRU emails. Hans von Storch has an interesting guest post over on Roger Pielke’s site discussing some of the concerns along with suggestions for improving the sustainability of science.

I think ’sustainability of science’ is his way of saying maintaining long-term credibility. Being honest about the uncertainties and not using science to support a ’preconceived political agenda of something good’. This is an unarguably good thing. A hard thing for sure, but something no one would argue against out loud. The term Pielke gives for the behaviour exhibited by the CRU scientists is ’stealth advocacy’. When you wrap the mantle of Science (relevant Anchorman audio clip, it really is relevant, the relevant part is at the very end) around your advocacy and misrepresent the actual state of knowledge to decision makers and laypeople, then you aren’t living up to that particular sort of honesty that Feynman exhorted scientists to uphold.

Friday, December 4, 2009

Bayesian Climate Model Averaging

Another instalment in the 'lack-of-climate-model-validation-bothers-me' series (see Lindzen's talk for a good intro). I've been reading Jaynes' book lately so naturally the Bayesian approach to the issue seems most germane. The whole climate science / public policy intersection can be viewed as one big decision theory problem (acting to maximize utility under uncertainty). To come out of that game well you generally need to have good models (hopefully with nice, tight predictive distributions) and smooth, gently sloping loss functions. I'll leave the loss functions for now and focus on the modelling aspect (since that's what I'm familiar with).

Validation (comparing the model predictions to experimental observations) is generally what allows you to find out if you've made good choices of what model structure to use, what physics to include and what physics to neglect. In most applications of computational physics this is a straight-forward (if sometimes expensive) process. The problem is harder with climate models. We can't do designed experiments on the Earth.

Here are a couple of choice quotes from Reichler and Kim 2007 about the difficulties of climate model validation.
Several important issues complicate the model validation process. First, identifying model errors is difficult because of the complex and sometimes poorly understood nature of climate itself, making it difficult to decide which of the many aspects of climate are important for a good simulation. Second, climate models must be compared against present (e.g., 1979-1999) or past climate, since verifying observations for future climate are unavailable. Present climate, however, is not an independent data set since it has already been used for the model development (Williamson 1995). On the other hand, information about past climate carries large inherit uncertainties, complicating the validation process of past climate simulations (e.g., Schmidt et al. 2004). Third, there is a lack of reliable and consistent observations for present climate, and some climate processes occur at temporal or spatial scales that are either unobservable or unresolvable. Finally, good model performance evaluated from the present climate does not necessarily guarantee reliable predictions of future climate (Murphy et al. 2004).

The above quoted paper is a comparison of three generations of IPCC-family models. The study shows improvement in prediction of modern climate as the models improve from 1990 to 2001 to 2007. It also shows that the ensemble mean is more skilled than any individaul model (more on this later). The reasons given to explain the improvement make intuitive sense:
Two developments, more realistic parameterizations and finer resolutions, are likely to be most responsible for the good performance seen in the latest model generation. For example, there has been a constant refinement over the years in how sub-grid scale processes are parameterized in models. Current models also tend to have higher vertical and horizontal resolution than their predecessors. Higher resolution reduces the dependency of models on parameterizations, eliminating problems since parameterizations are not always entirely physical. That increased resolution improves model performance has been shown in various previous studies (e.g., Mullen and Buizza 2002, Mo et al. 2005, Roeckner et al. 2006).

A problem faced by climate modelers is that it is unlikely that we'll be able to run grid-resolved solutions of the climate within the lifetime of anyone now living (us CFD guys have the same problem with grid resolution scaling for high Reynolds number flows). There will always be a need for 'sub-grid' parameterizations, the hope is that eventually they will become "entirely physical" and well calibrated (if you think they already are, then you have been taken by someone's propaganda).

Bayesian model averaging (BMA) is one way to account for our uncertainty in model structure / physics choices. Instead of choosing a 'right' model, we get predictive distributions for things we care about by marginalizing over the uncertain model structures (and the uncertain parameters too).This paper shows that it is a useful procedure for short-term forecasting. The benefit with short-term forecasts is that we can evaluate the accuracy by closing the loop between predictions and observations. Min and Hense apply this idea to the IPCC AR4 coupled-climate models. Here's a short snippet from that paper providing some motivation for the use of BMA:
However, more than 50% of the models with anthropogenic-only forcing cannot reproduce the observed warming reasonably. This indicates the important role of natural forcing although other factors like different climate sensitivity, forcing uncertainty, and a climate drift might be responsible for the discrepancy in anthropogenic-only models. Besides, Bayesian and conventional skill comparisons demonstrate that a skill-weighted average with the Bayes factors (Bayesian model averaging, BMA) overwhelms the arithmetic ensemble mean and three other weighted averages based on conventional statistics, illuminating future applicability of BMA to climate predictions.

The ensemble means or Bayesian averages tend to outperform individual models, but why is this? Here's what R&K2007 has to say:
Our results indicate that multi-model ensembles are a legitimate and effective means to improve the outcome of climate simulations. As yet, it is not exactly clear why the multi-model mean is better than any individual model. One possible explanation is that the model solutions scatter more or less evenly about the truth (unless the errors are systematic), and the errors behave like random noise that can be efficiently removed by averaging. Such noise arises from internal climate variability (Barnett et al. 1994), and probably to a much larger extent from uncertainties in the formulation of models (Murphy et al. 2004; Stainforth et al. 2005).

Another interesting paper that explores this finds that models which have good scores on the calibration data do not tend to outperform other models over a subsequent validation period.
Error in the ensemble mean decreases systematically with ensemble size, N, and for a random selection as approximately 1∕Na, where a lies between 0.6 and 1. This is larger than the exponent of a random sample (a = 0.5) and appears to be an indicator of systematic bias in the model simulations.

This should not be surprising, it is very difficult to get all of the physics right (and remove the systematic bias) when you aren't able to do no-kidding validation experiments. They begin their conclusion with
In our analysis there is no evidence of future prediction skill delivered by past performance-based model selection. There seems to be little persistence in relative model skill, as illustrated by the percentage turnover in Figure 3. We speculate that the cause of this behavior is the non-stationarity of climate feedback strengths. Models that respond accurately in one period are likely to have the correct feedback strength at that time. However, the feedback strength and forcing is not stationary, favoring no particular model or groups of models consistently.

This means it is very difficult to protect ourselves from 'over-fitting' the models to our available historical record, and it certainly indicates that we should be cautious in basing policy decision on climate model forecasts. The 'science is settled' crowd, while busy banging the consensus drum and clamouring for urgent action (NOW!), never seem to offer this sort of nuanced approach to policy though.

If you have read any good, recent climate model validation papers please post them in the comments. Please don't post polemics about polar bears and arctic sea ice, my skepticism is honest, your activism should be too.

For some further Bayes Model Averaging / Model Selection check out:

(isn't Google Books cool?)