Showing posts with label risk management. Show all posts
Showing posts with label risk management. Show all posts

Saturday, March 25, 2017

Innovation, Entropy and Exoplanets

I enjoy Shipulski on Design for the short articles on innovation. They are generally not technical at all. I like to think of most of the posts as innovation poetry to put your thoughts along the right lines of effort. This recent post has a huge, interesting technical iceberg riding under the surface though.
If you run an experiment where you are 100% sure of the outcome, your learning is zero. You already knew how it would go, so there was no need to run the experiment. The least costly experiment is the one you didn’t have to run, so don’t run experiments when you know how they’ll turn out. If you run an experiment where you are 0% sure of the outcome, your learning is zero. These experiments are like buying a lottery ticket – you learn the number you chose didn’t win, but you learned nothing about how to choose next week’s number. You’re down a dollar, but no smarter.

The learning ratio is maximized when energy is minimized (the simplest experiment is run) and probability the experimental results match your hypothesis (expectation) is 50%. In that way, half of the experiments confirm your hypothesis and the other half tell you why your hypothesis was off track.
Maximize The Learning Ratio

Sunday, March 22, 2015

Reliability Growth: Enhancing Defense System Reliability


This report (pdf) from the National academies on reliability growth is interesting. There's a lot of good stuff on design for reliability, physics of failure, highly accelerated life testing, accelerated life testing and reliability growth modeling. Especially useful is the discussion about the suitability of assumptions underlying some of the different reliability growth models.

The authors provide a thorough critique of MIL-HDBK-217, Reliability Prediction of Electronic Equipment, in Appendix D, which is probably worth the price of admission by itself. If you're concerned with product reliability you should read this report (lots of good pointers to the lit).

Tuesday, January 13, 2015

Guidelines for Planning and Evidence for Assessing a Well-Designed Experiment


This paper is full of great guidance for planning a campaign of experimentation, or assessing the sufficiency of a plan that already exists. The authors break up the effort into four phases:
  1. Plan a Series of Experiments to Accelerate Discovery
    1. Design Alternatives to Span the Factor Space
    2. Decide on a Design Strategy to Control the Risk of Wrong Conclusions
  2. Execute the Test
  3. Analyze the Experimental Design
They give a handy checklist for each phase (reproduced below). The checklists are comprehensive (significantly more than my little list of questions) and I think they stand-alone, but the whole paper is well worth a read. Design of experiments is more than just math, as this paper stresses it is a strategy for discovery.

Wednesday, August 21, 2013

3-D Printing in DoD: Who's Dragging Their Feet?

I found this article, Why is the Pentagon Dragging Its Feet on 3D Printing, by way of Small Wars Journal. It has some interesting information. The Army is deploying mobile Fab Labs, which seems like a mini MIT FabLab in a shipping container. I think this is a really neat idea. How this can be characterized as feet dragging, I'm not sure. The feet dragging accusation is based on some hand-waving from an article on Disruptive Thinkers, and another article that seems to be worried that there is no Pentagon overlord in charge of an additive manufacturing strategy:
With possible dwindling budgets on the horizon, a clear strategy and cohesive approach is essential to create efficiencies in the area of research and development as well as eliminating duplicative efforts. In order for DoD to take advantage of what is anticipated to be an explosion in the commercial sector within the next ten years, the Department must take an active approach, partnering with the private sector to keep up with this relatively nascent technology and shaping/guiding it towards the desired end state the department has in mind.

One step towards a clear strategy and cohesive approach is for DoD to designate an AM Czar within the Department. They could serve as a single point for all things AM and not the myriad of technical advisory boards that currently exist. This office could then work with policy makers to execute and monitor a strategy which will allow DoD to take full advantage of this technology. Logically, this office would interface directly with the National Additive Manufacturing and Innovation Institute (NAMII) as DoD's representative
3-D Printing Revolution in Military Logistics
I think an "additive manufacturing Czar" sounds like a terrible idea (so I'm sure it will secure funding for some beltway bandits to do a study). I know my recent success with qualifying a particular additive manufacturing process and supplier for use in 3D printing wind tunnel models did not need a Pentagon king-pin to tell me about DoD's strategy for additive manufacturing. Using this technology just made sense as a way to solve my problem: get a complex wind-tunnel model rapidly, and at an affordable cost. I did not receive top-down direction or guidance to use AM, I simply took the initiative to solve my problem. After reading that article I'm left wondering, just how exactly is waiting on direction from the very heights of the bureaucracy supposed to lead to innovation?

Sunday, May 8, 2011

Storms of Our Grandfathers

Are we "rolling 13s" and getting thousand year storms every year?

NOAA April 2011 Precipitation Anomaly

The contour plots below are taken from Theory of the hydraulic jump and backwater curves. These studies of historical storm records were used to inform design decisions for the Miami Valley Conservancy District's retarding basins and channel improvements following the 1913 floods. My previous post has pictures of the hydraulic jump below Huffman Dam in operation.

My question to Dr Curry about what value high-fidelity (read: relatively expensive to run and analyze) climate simulations have for decision makers was motivated by reading up on infrastructure projects like the retarding basins and channel improvements in the Miami Valley. I think it would be interesting to take a look at a historical project like this that included rudimentary analysis of climate (weather event frequency and magnitude) in its design, and say, "here's how it would be informed differently using modern tools."

The design philosophy taken by the engineers working for the Miami Valley Conservancy District was to design for the worst possible case (historical records from Europe were also considered since they went back further and more reliably) plus roughly twenty percent margin due to the inherent uncertainty in estimating the worst possible case.

If it were necessary to depend wholly on the records of storms which have occurred in the United States, it might be thought possible for moderately great storms to occur over a period of a few hundred years, and then to find, as an exception, a storm three or four times as great. Theoretically that is very improbable, simply because water vapor in sufficient quantities cannot be transported from the ocean or gulf fast and long enough to cause such exceptional storms. As stated in chapter XI, however, records were collected of the stages of rivers in Europe for long periods of time, and these furnish fairly conclusive proof that such great exceptional storms actually do not occur. On the Danube at Vienna, for instance, we have records since about the year 1000 A.D.; fairly accurate records are available for stages of floods in the Tiber at Rome for more than 2,000 years; and records have been made of floods on the Seine at Paris for a long period of years.
Relation of Great Storms to Maximum Possible
After making the extensive investigation of storms in the eastern United States, it is believed that the March, 1913, flood is one of the great floods of centuries in the Miami Valley. In the course of three or four hundred years, however, a flood 15 or 20 per cent greater may occur. We do not believe a flood will ever occur which is more than 20 or 25 per cent in excess of that of March 1913. There is a factor of ignorance, however, against which we must provide, and the only way to do this is arbitrarily to increase the size of the maximum flood to be provided for. If longer records were available a closer estimate could be made, but in planning works on which the protection of the Miami Valley depends, it is necessary to go beyond human judgment. This has been done on all the other phases of the design, and we believe it would not be good engineering practice to stop at our judgment on this phase. We must be able to say that the engineering works are absolutely safe in every respect. For this reason provision is made for a flood nearly 40 per cent greater than that of March 1913. This is 15 or 20 per cent in excess of what is believed to be the greatest possible flood that will ever occur.
Reasons for Choosing as a Basis for Design a Flood 40% Greater than that of March 1913
Would modern tools cut the design margin due to reduced uncertainty or would they indicate that the project is now under-designed due to projected climate change? The latter seems unlikely considering that the magnitude of the purported effects has been repeatably shown to be smaller than we can reliably detect given the length of our data record. Would there be any practically significant changes to the decisions and designs? If your system already has sufficient margin for projected changes in weather-event magnitude do projected changes in frequency matter?

Friday, May 6, 2011

Technocrats and Philosopher Kings can Save our Impotent Polity

Wow, really awesome article on Climate Resistance, Trust Me, I Speak for Science. I liked these parts from the concluding paragraphs especially. I think you'll notice the parallels to my posts, The Social Ethic and Appeals for Technocracy and No Fluid Dynamicist Kings in Flight-Test.
This metaphysical confusion runs throughout Mooney’s argument. For Mooney, ‘ideology’ is some insidious, toxic force, the antithesis to ‘truth’ itself. The thrust of his argument is that we need particular scientific institutions to ameliorate this intrinsic weakness of human nature. And as such, these institutions deserve elevated status above the reach of those prone to ideology. Otherwise, we would tend towards creationism, to MMR-scares, to climate-change denial. In other words, our flawed minds would create a catastrophe, and it is this possibility of catastrophe that seemingly legitimises the elevated position of scientific institutions. Mooney reinvents Plato’s city state administrated by Philosopher Kings, the main differences being that Mooney conceives of a global polity, and the wisdom of the Guardians only produces the possibility of mere survival, not even a better way of life. To bring this back the matter of trust, Mooney doesn’t trust humans. Their minds are flawed. Their ambitions and ideas are mere fictions. The institutions they create are accordingly founded on false premises, which, instituted and acted upon, will cause disaster. Even when humans are exposed to ‘the truth’, it is, on Mooney’s view, absorbed into the poisonous, ideological programmes of partisans: liars and cheats who distort it. But without a disaster looming, this instance of a politics of fear would collapse.
He simply can’t make a popular argument for his political idea, and so turns to ‘science’ to identify the necessity of such a programme — i.e. the crisis — and to identify reasons why conventional democratic processes cannot realise it...
It's always a good day when you can throw a little Plato into the mix ; - )

Wednesday, May 4, 2011

Dayton Flood Control Infrastructure at Work

Photos of the flood control contrivances in and around Dayton. All of these were taken on 3 May 2011. But first, a little history and engineering detail so you'll have a better appreciation of the pictures.

Arthur P. Morgan came to Dayton after the 1913 flood to design a flood control system to protect the entire Miami Valley. One element of this system was a dry dam—a dam that held water only during a flood and released the water at a rate that the downstream riverbed could carry. The problem was that the speed of the water through the dam made it powerful and destructive. To solve that problem, Morgan went with Col. Edward Deeds to his farm in Moraine where they built models in his swimming pool. They developed the hydraulic jump, which sends water through a series of baffles and steps, and then finally into a low wall that forces the water back onto itself, dissipating its own energy. This process of turning water onto itself is the hydraulic jump. From there, the water flows downstream calmly. This technology is still used in hydrological engineering throughout the world.

From Dayton Inventors River Walk

Here's a graphical depiction of a hydraulic jump.

A basic 1-D analysis follows the figure (clicking the image should take you to the free Google e-book).

Something that's kind of neat is that this system of flood control was designed in the days when "computers" were people (often women) not machines:

These engineers were certainly sure of themselves:
The bottom line is the important part. The hydraulic jump works by increasing the rate of turbulent kinetic energy production, this leads rather quickly (immediately if you assume equilibrium turbulence) to an increased rate of turbulent kinetic energy dissipation at the bottom of the energy cascade. The destructive capability (momentum) of the water is greatly reduced in exchange for raising its temperature ever so slightly.

The following figure shows the cuts that had to be made for the outlet channels and hydraulic jump pools (note Huffman Dam in the center).

And this one shows an aerial shot of the Huffman Dam just after completion.

These views from the top of the Huffman Dam show the turbulence at the end of the outlet channels due to the hydraulic jump.

Huffman Dam Outlet Channels
Huffman Dam Hydraulic Jump Pool Turbulence
View From the Top of Huffman Dam

These types of momentum dissipation mechanisms are also used throughout the city. The submerged dams take momentum out of the four streams that come together in the Dayton city limits: Miami, Mad and Stillwater Rivers and Wolf Creek.

Low Dam North-West of Downtown Dayton
Low Dam downstream of Dayton Canoe Club
There's talk of replacing these with something more water-sport (canoe / kayak) friendly.

Friday, February 4, 2011

Validation and Calibration: more flowcharts

In a previous post we developed a flow-chart for model verification and validation (V&V) activities. One thing I noted in the update on that post was that calibration activities were absent. My google alerts just turned up a new paper (they reference the Thacker et al. paper the previous post was based on, I think you’ll notice the resemblance of flow-charts) which adds the calibration activity in much the way we discussed.


PIC

Figure 1: Model Calibration Flow Chart of Youn et al. [1]

The distinction between calibration and validation is clearly highlighted, “In many engineering problems, especially if unknown model variables exist in a computational model, model improvement is a necessary step during the validation process to bring the model into better agreement with experimental data. We can improve the model using two strategies: Strategy 1 updates the model through calibration and Strategy 2 refines the model to change the model form.”


PIC

Figure 2: Flow chart from previous post

The well-founded criticism of calibration-based arguments for simulation credibility is that calibration provides no indication of the predictive capability of a model so-tuned. The statistician might use the term generalization risk to talk about the same idea. There is no magic here. Applying techniques such as cross-validation merely add a (hyper)parameter to the model (this becomes readily apparent in a Bayesian framework). Such techniques, while certainly useful, are no silver bullet against over-confidence. This is a fundamental truth that will not change with improving technique or technology, and that is because all probability statements are conditional on (among other things) the choice of model space (particular choices of which must by necessity be finite, though the space of all possible models is countably infinite).
One of the other interesting things in that paper is their argument for a hierarchical framework for model calibration / validation. A long time ago, in a previous life, I made a similar argument [2]. Looking back on that article is a little embarrassing. I wrote that before I had read Jaynes (or much else of the Bayesian analysis and design of experiments literature), so it seems very technically naive to me now. The basic heuristics for product development discussed in it are sound though. They’re based mostly on GAO reports [3456], a report by NAS [7], lessons learned from Live Fire Test and Evaluation [8] and personal experience in flight test. Now I understand better why some of those heuristics have sound theoretical underpinnings.
There are really two hierarchies though. There is the physical hierarchy of system, sub-system and component that Youn et al. emphasize, but there is also a modeling hierarchy. This modeling hierarchy is delineated by the level of aggregation, or the amount of reductive-ness, in the model. All models are reductive (that’s the whole point of modeling: massage the inordinately complex and ill-posed into tractability), some are just more reductive than others.


PIC

Figure 3: Modeling Hierarchy (from [2])

Figure 3 illustrates why I care about Bayesian inference. It’s really the only way to coherently combine information from the bottom of the pyramid (computational physics simulations), with information higher up the pyramid which rely on component and subsystem testing.
A few things I don’t like about the approach in [1]
  • The partitioning of parameters into “known” and “unknown” based on what level of the hierarchy (component, subsystem, system) you are at in the “bottom-up” calibration process. Our (properly formulated) models should tell us how much information different types of test data give us about the different parameters. Parameters should always be described by a distribution rather than discrete switches like known or unknown.
  • The approach is based entirely on the likelihood (but they do mention something that sounds like expert priors in passing).
  • They claim that the proposed calibration method enhances “predictive capability” (section 3), however this is misleading abuse of terminology. Certainly the in-sample performance is improved by calibration, but the whole point of making a distinction between calibration and validation is based on recognizing that this says little about the out-of-sample performance (in fairness, they do equivocate a bit on this point, “The authors acknowledge that it is difficult to assure the predictive capability of an improved model without the assumption that the randomness in the true response primarily comes from the the randomness in random model variables.”).
Otherwise, I find this a valuable paper that strikes a pragmatic chord, and that’s why I wanted to share my thoughts on it.
[Update: This thesis that I linked at Climate Etc. has a flow-chart too.
]

References

[1]   Youn, B. D., Jung, B. C., Xi, Z., Kim, S. B., and Lee, W., “A hierarchical framework for statistical model calibration in engineering product development,” Computer Methods in Applied Mechanics and Engineering, Vol. 200, No. 13-16, 2011, pp. 1421 – 1431.
[2]   Stults, J. A., “Best Practices for Developmental Testing of Modern, Complex Munitions,” ITEA Journal, Vol. 29, No. 1, March 2008, pp. 67–74.
[3]   Defense Acquisitions: Assesment of Major Weapon Programs,” Tech. Rep. GAO-03-476, U.S. General Accounting Office, May 2003.
[4]   Best Practices: Better Support of Weapon System Program Managers Needed to Improve Outcomes,” Tech. Rep. GAO-06-110, U.S. General Accounting Office, 2006.
[5]   Precision-Guided Munitions: Acquisition Plans for the Joint Air-to-Surface Standoff Missile,” Tech. Rep. GAO/NSIAD-96-144, U.S. General Accounting Office, 1996.
[6]   Best Practices: A More Constructive Test Approach is Key to Better Weapon System Outcomes,” Tech. Rep. GAO/NSIAD-00-199, U.S. General Accounting Office, July 2000.
[7]   Michael L. Cohen, John E. Rolph, D. L. S., editor, Statistics, Testing and Defense Acquisition: New Approaches and Methodological Improvements, National Academy Press, Washington D.C., 1998.
[8]   O’Bryon, J. F., editor, Lessons Learned from Live Fire Testing: Insights Into Designing, Testing, and Operating U.S. Air, Land, and Sea Combat Systems for Improved Survivability and Lethality, Office of the Director, Operational Test and Evaluation, Live Fire Test and Evaluation, Office of the Secretary of Defense, January 2007.

Tuesday, August 3, 2010

No Fluid Dynamicist Kings in Flight-Test

This was a guest post over on Pielke's site.

Dr Pielke's Honest Broker concepts resonate with me because of practical decision support experiences I've had, and this post is an attempt to share some of those from a realm pretty far removed from the geosciences. All the views and opinions expressed are my own and in no way represent the position or policy of the US Air Force, Department of Defense or US Government. I am writing as a simple student of good decision making. My background is not climate science. I am an Aeronautical Engineer with a background in computational fluid dynamics, flight test and weapons development. I got interested in the discussions of climate policy because the intersection of computational physics and decision making under uncertainty is an interesting one no matter what the subject area. The discussion in this area is much more public than the ones I'm accustomed to, so it makes a great target of opportunity. The decision support concepts Dr Pielke discusses make so much sense to me now, but I can see how hard they are for technical folks to grasp because I used to be a very linear thinker when I was a young engineer right out of school.

My journeyman's education in decision support came when I got the chance to lead a small team doing Live Fire Test and Evaluation for the Air Force (you may not be familiar with LFT&E, it is a requirement that grew out of the Army gaming testing of the Bradley fighting vehicle in the 1980s, a situation that was fairly accurately lampooned in the movie "Pentagon Wars"). The competing values of the different stakeholders (folks appointed by congress to ensure sufficient realistic testing compared to folks at the service level doing product development) was really an eye-opening education for a technical nerd like me. I initially thought, "if only everyone can agree on the facts, the proper course of action will be clear". How naive I was! Thankfully, the very experienced fellows working for me didn't mind training up a rash, newly-minted, young Captain.

It's tough for some technical specialists (engineers/scientists) to recognize worthy objectives their field of study doesn't encompass. The reaction I see from the more technically oriented folks like Tobis (see how he struggles) reminds me a lot of the reaction that engineers in product development offices would have to the role of my little Live Fire office. A difficulty we often encountered was the LFT&E oversight folks wanted to accomplish testing that didn't have direct payoff to narrower product development goals that concerned the engineers. "What those people want to do is wasteful and stupid!" This parallels the recent sand berm example. The preferred explanation from the technician's perspective is that the other guy is bat-shit crazy, and his views should be ridiculed and de-legitimized. The truth is usually closer to the other guy having different objectives that aren't contained within the realm of the technician's expertise. In fact, the other person is probably being quite rational, given their priors, utility function and state of knowledge.

In my little Live Fire Office we had lots of discussion about what to call the role we did, and how to best explain it to the program managers. I wish I had heard of Dr Pielke's book back then, because "Honest Broker" would have been an apt description for much of the role. We acted as a broker between the folks in the Pentagon with the mandate from congress for sufficient, realistic testing, and the Air Force level program office with the mandate for product development. The value we brought (as we saw it), was that we were separate from the direct program office chain of command (so we weren't advocates for their position), but we understood the technical details of the particular system, and we also understood the differing values of the folks in the Pentagon (which the folks in the program office loved to refuse to acknowledge as legitimate, sound familiar?). That position turns out to be a tough sell (program managers get offended if you seem to imply they are dishonest), so I can empathize with the virulent reaction Dr Pielke gets on applying the Honest Broker concepts to climate policy decision support. People love to take offense over their honor. That's a difficult snare to avoid while you try to make clear that, while there's nothing dishonest about advocacy, there remains significant value in honest brokering. Maybe Honest Broker wouldn't be the best title to assume though. The first reaction out of a tight-fisted program manager would likely be "I'm honest, why do I need you?"

One of the reason my little office existed was because of some "lessons learned" from the Tri-Service Standoff Missile debacle (all good things in defense acquisition must grow out of historical buffoonery). The broader Air Force leadership realized that it was counterproductive to have product development engineers and program managers constantly trying to de-legitimize the different values that the oversight stake-holders brought (the differences springing largely from different appetites for risk and priors for deception) by wrangling over largely inconsequential, technical nits (like tree rings in the Climate Wars). The wiser approach was to maintain an expertise whose sole job was to recognize and understand the legitimate concerns of the oversight folks and incorporate those into a decision that meets the service's constraints as quickly and efficiently as possible. Rather than wasting time arguing, product development folks could focus on product development.

The other area where I've seen this dynamic play out is in making flight test decisions. In that case though, the values of all the stake-holders tend to align more closely, so the separation between technical expertise and decision making is less contentious (Dr Pielke's Tornado analogy). In contrast to the climate realm where it's argued that science compels because we're in the Tornado mode, the flight-test engineers understand that the boss is taking personal responsibility for putting lives at risk based on their analysis. They tend to be respectful of their crucial, but limited, role in the broader risk management process. Computational fluid dynamics can't tell us if it's worth risking the life of an air crew to collect that flight test data. In that case there is no confusion about who is king, and over what questions the technical expert must "pass over in silence."

Wednesday, March 17, 2010

Zen Uncertainty

Zen Uncertainty: Attempts to understand uncertainty are mere illusions; there is only suffering.
-- WARNING: Physics Envy May Be Hazardous To Your Wealth!
Should we give up? No, there's plenty we can do to make the suffering more bearable. Lo and Mueller give an uncertainty taxonomy of five levels in their 'Physics Envy' paper:
  1. Complete Certainty: the idealized deterministic world
  2. Risk without Uncertainty: an honest casino
  3. Fully Reducible Uncertainty: the odds in the honest casino are not posted, we have to learn them from limited experience
  4. Partially Reducible Uncertainty: we're not quite sure which game at the casino we're playing so we have to learn that as well as the odds based on limited experience
  5. Irreducible Uncertainty: we're not even sure if we're in the casino, we might be outside splashing around in the fountain...
At the bottom of the decent we find level infinity, Zen Uncertainty.

Section 2 of the paper provides a nice historical overview of the early work of Paul A. Samuelson, who single-handedly brought statistical mechanics to the economists, and they have never been the same since. Samuelson acknowledged the deep connection between his work and physics:
Perhaps most relevant of all for the genesis of Foundations, Edwin Bidwell Wil- son (1879–1964) was at Harvard. Wilson was the great Willard Gibbs’s last (and, essentially only) protege at Yale. He was a mathematician, a mathematical physicist, a mathematical statistician, a mathematical economist, a polymath who had done first-class work in many fields of the natural and social sciences. I was perhaps his only disciple . . . I was vaccinated early to understand that economics and physics could share the same formal mathematical theorems (Euler’s theorem on homogeneous functions, Weierstrass’s theorems on constrained maxima, Jacobi determinant identities underlying Le Chatelier reactions, etc.), while still not resting on the same empirical foundations and certainties.
Related to this theme, there's an interesting recent article over on Mobjectivist site about using ideas from physics to model income distributions.

Lo and Mueller propose to operationalize their uncertainty taxonomy with a 2-D checklist (table). The levels provide the columns across the top, and there is a row for each business component of the activity being evaluated, here's their description:
The idea of an uncertainty checklist is straightforward: it is organized as a table whose columns correspond to the five levels of uncertainty of Section 3, and whose rows correspond to all the business components of the activity under consideration. Each entry consists of all aspects of that business component falling into the particular level of uncertainty, and ideally, the individuals and policies responsible for addressing their proper execution and potential failings.
This seems like an idea that could be adapted and combined with best practices for model validation (and checklist sorts of approaches) in helping to define what sorts of uncertainties we are operating under when we make decisions using science-based decision support products.

Their final paragraph echos Lindzen's sentiments about climate science:
While physicists have historically been inspired by mathematical elegance and driven by pure logic, they also rely on the ongoing dialogue between theoretical ideals and experimental evidence. This rational, incremental, and sometimes painstaking debate between idealized quantitative models and harsh empirical realities has led to many breakthroughs in physics, and provides a clear guide for the role and limitations of quantitative methods in financial markets, and the future of finance.
-- WARNING: Physics Envy May Be Hazardous To Your Wealth!

Tuesday, February 9, 2010

A Few Nits About Ensembles and Decision Support

First off, thanks to Steve Easterbrook for pointing at a new report on ensemble climate predictions. This post is some of my thoughts on that report.

Here's an interesting snippet from the first section of the report:
Within the last decade the causal link between increasing concentrations of anthropogenic greenhouse gases in the atmosphere and the observed changes in temperature has been scientifically established.
From a nice little summary of how to establish causation:
C. Establishing causation: The best method for establishing causation is an experiment, but many times that is not ethically or practically possible (e.g., smoking and cancer, education and earnings). The main strategy for learning about causation when we can’t do an experiment is to consider all lurking variables you can think of and look at how Y is associated with X when the lurking variables are held “fixed.”

D. Criteria for establishing causation without an experiment: The following criteria make causation more credible when we cannot do an experiment.
(i) The association is strong.
(ii) The association is consistent.
(iii) Higher doses are associated with stronger responses.
(iv) The alleged cause precedes the effect in time.
(v) The alleged cause is plausible.
The point being, if you can't do experiments, the causal link you establish will always be a rather contingent one (and if your population of 'lurkers' is only 4 or 5 then perhaps we aren't to the point of exhausting our imaginations yet).  I say that not to be disingenuous and sow doubt unnecessarily, but merely to show that I have an honest place to stand in my skepticism (I'd like to head off the "you're a willfully ignorant pseudo-scientific jerk" sorts of flames that seem to be popular in discussions on this topic).

That little digression aside, the specific aim of the work is to
develop an ensemble prediction system for climate change based on the principal state-of-the-art, high-resolution, global and regional Earth system models developed in Europe, validated against quality-controlled, high-resolution gridded datasets for Europe, to produce for the first time an objective probabilistic estimate of uncertainty in future climate at the seasonal to decadal and longer time-scales;
A side-note on climate alarmism: If the quality of the body of knowledge was such that it demanded action NOW! Then this would not have been the first such study. Rational decision support requires these sorts of uncertainty quantification efforts, it is totally irresponsible to demand political action without them.
The improvements for example, add skill to seasonal forecasting while multi-decadal models, for the first time, have produced probabilistic climate change projections for Europe.
Again, now that these projections have been made for the first time, we could actually attempt to validate them. I use that term in the technical sense of comparing a model's predictions to innovative experimental results. Since we can't do experiments on the earth (or can we?) we have to settle for either not validating, or validating by comparing the predictions to what actually happens. I realize that would take a couple decades to make useful quantitative comparisons. But think about this, if we've already bought a millenia of warming then can't we spend a decade or two to build the credibility in our tool-set which we'll be using to 'fly' the climate into the future for centuries to come? The fact that this set of ensemble results claims to be skillfull at the decadal time-scales would actually make the validation task a quicker one than it would be with less accurate models because you've taken some of the 'noise' and explained it with your model.

I got really excited about this, it sounds promising:
The multi-model ensemble builds on the experience of previous projects where it has been shown to be a successful method to improve the skill of seasonal forecasts from individual models. The perturbed parameter approach reflects uncertainty in physical model parameters, while the newly developed stochastic physics methodology represents uncertainty due to inherent errors in model parameterisations and to the unavoidably finite resolution of the models.
Then I got to this:
These results illustrate that initialised decadal forecasts have the potential to provide improved information compared with traditional climate change projections, but the optimal strategy for building improved decadal prediction systems in the presence of model biases remains an open question for future work.
Which reflects my impression of the state of the art from my little mini-lit review on Bayes Model Averaging. That's the fundamental difficulty, isn't it?

This is the sort of thing that worries folks who are used to being able to draw a bright line between calibration and validation:
The ENSEMBLES gridded observation data set was used along with other datasets to verify and calibrate both global and regional models, and also to assess the uncertainties in model response to anthropogenic forcing.

Recall Kelvin,
Conditional PDFs, which encompass the sampled uncertainty, were constructed from the statistically and dynamically downscaled output (and from GCM output) for temperature and/or precipitation for a number of areas and points. These are, however, qualitative constructions.
There's still some work to do here to make the product more useful for decision makers (quantitative rather than qualitative). I think the polynomial chaos expansion approaches being explored in the uncertainty quantification community have a lot of promise here (3 or 4 orders of magnitude speed-up over standard Monte Carlo approaches). The other slight difficulty with this approach is that these qualitative PDFs were then used as inputs into the impact assessments.

In light of the previous ensembles post and linked discussion thread, I found this snippet interesting:
The non-linear nature of the climate system makes dynamical climate forecasts sensitive to uncertainty in both the initial state and the model used for their formulation. Uncertainties in the initial conditions are accounted for by generating an ensemble from slightly different atmospheric and ocean analyses. Uncertainty in model formulation arises due to the inability of dynamical models of climate to simulate every single aspect of the climate system with arbitrary detail. Climate models have limited spatial and temporal resolution, so that physical processes that are active at smaller scales (e.g., convection, orographic wave drag, cloud physics, mixing) must be parameterised using semi-empirical relationships.

In that Adventures Among Alarmists post I made a sort of hand-wavy claim of everything about forecasting (from data assimilation on through to predictive distributions) being one big ill-posed problem with noise. I think reading section 3 of this report will give you a flavor of what I mean. One minor quibble: hindcasts are an ok sort of 'sanity' check, but we should take care to remember their dangers, and not mistake them for true validation. Taking too much confidence from hind-casts is a recipe for fooling ourselves.

The ensemble's skill changes with lead-time:
The skill increases for longer lead times, being larger for 6–10 years ahead than for 3–14 months or 2–5 years ahead. This is because the forced climate change signal, the sign of which is highly predictable, is greater at longer lead times.
This squares with the results discussed over here about BMA in climate and weather forecasting. Depending on the time-frame at which you are looking to forecast, different model weightings, and spin-up times are optimal. This also goes to the problem about the past-performance / future-skill connection. Since the feedbacks are not stationary, models which performed well in the past won't tend to perform well in the future (eg. the BMA weighting changes through time).
Encouragingly, the multi-model ensemble mean, which consists of the average of twelve individual projections, gives somewhat higher scores than any of the individual models, whose projections are derived from three members with perturbed initial conditions.
So, truth-centered or not?

The results also show that the skill increases for more recent hindcasts. In order to diagnose sources of skill, the blue curve of Figure 3.6 shows ensemble mean results from a parallel ensemble of ‘NoAssim’ hindcasts containing the same external forcing from greenhouse gases, sulphate aerosols, volcanoes and solar variations, but initialised from randomly selected model states rather than analyses of observations.
This seems like a reasonable use of hindcasts. Look for insight into the reasons particular models / realizations might have performed well on certain historical periods. They also found that initialization matters even with climate predictions (though their randomly selected initializations were pretty darn skill-full).

The product, decision support:
For many grid boxes there are significant probabilities of both drier and wetter future climates, and this may be important for impacts studies.
I think as regional projections start becoming more and more available, the extreme, alarmist policy prescriptions will be less and less well supported by a rational cost-benefit analysis.

The sensitivity study discussed in the report (done by climateprediction.net) is worth noting simply for the fact that they found an interesting interaction (that's always a fun part of experimentation). Some of the criticism of this effort has focused on the plausibility of some of the parameter combinations in the thousands of runs of this computer experiment. The distinction to keep in mind is that this is a sensitivity study rather than a complete uncertainty quantification study.

Well it said in the executive summary that they generated 'qualitative PDFs', but it seems like section 3.3.2 is describing a quantitative Bayesian approach. They sample their parameter space with a variety of models with varying levels of complexity and then fit a simple surrogate (some people might call it a response surface) so that they can get approximate 'results' for the whole space. Then they get posterior probabilities by weighting expert obtained priors by likelihoods, a straight-forward application of Bayes Theorem. That seems as quantitative as anyone could ask for, maybe I'm missing something?

They give a nice summary of the different types of uncertainties:
Also, the three techniques for sampling modelling uncertainty are essentially complementary to one another, so should not be seen as competing alternatives: the multi-model approach samples structural variations in model formulation, but does not systematically explore parameter uncertainties for a given set of structural choices, whereas the perturbed parameter approach does the reverse. The stochastic physics approach recognises the uncertainty inherent in inferring the effects of parameterised processes from grid box average variables which cannot account for unresolved sub-grid-scale organisations in the modelled flow, whereas the other methods do not. There is likely to be scope to develop better prediction systems in future by combining aspects of the separate systems considered in ENSEMBLES.

I think section 3 was the meat of the report (at least for what I'm interested in), so I'll stop with the commentary on that report there. I want to end with an answer to the question often suggested for dealing with us ornery skeptics, "What evidence would it take to convince you?" This is usually meant to be a jab, because obviously skeptics are really deniers in disguise and we couldn't possibly be reasonable or consider evidence (and it also displays one of the common fallacies of regarding disagreement about policy with ignorance of science, science demands nothing but a method). My answer is a simple one: Rational policy tied to skillful prediction. By rational policy I mean it is foolish to focus on the cost-benefit of extreme events far out into the future weighed against mitigation today or tomorrow (because the tails of those future cost distributions are so uncertain, and the immediate costs of mitigation are relatively well known). It is far more rational to look at the near-term cost-benefit of adaption to climate changes (no matter their cause) and evolutionary improvements to our irrigation, flood management, public health and energy diversity problems. The skillfulness of near-term predictions can be reasonably validated, and then used to guide policy. This should be a natural extension of the way we already make agriculture and infrastructure decisions based on weather predictions and an understanding of our current climate. There's no need for the slashing and burning of evil western capitalism (or whatever the rallying cry at Copenhagen was). The ability to skillfully predict changes further and further into the future can be gradually validated (by making a prediction and then waiting, sorry that's what seems reasonable to me) and then the results of those tools can be incorporated into policy decisions. Markets and people don't respond well to shocks. Gradualism may not be sexy, but it's smart.

Wednesday, February 3, 2010

Joint Targeting Zen

Sometimes the most important part of the targeting cycle is deciding what targets not to engage to achieve the effects we want.
"It's possible (a strike) could be used to play to nationalist tendencies," Petraeus, head of the U.S. Central Command region, which includes Iran, said in an interview this week. "There is certainly a history, in other countries, of fairly autocratic regimes almost creating incidents that inflame nationalist sentiment. So that could be among the many different, second, third, or even fourth order effects (of a strike)." --Patraeus Says Strike on Iran Could Provoke Nationalism
The response of the British populace to The Blitz provides a good historical example of the kind of thing Gen Patraeus is describing.
Thirty spokes
Round one hub.
Employ the nothing inside
And you can use a cart.
Knead the clay to make a pot.
Employ the nothing inside
And you can use a pot.
Cut out doors and windows.
Employ the nothing inside
And you can use a room.
What is achieved is something,
By employing nothing it can be used.
--Tao Te Ching, 11

Tuesday, December 22, 2009

Chaos: A Very Short Introduction (Book Review)

I got an early Christmas present from my favourite experimentalist. It's a book called "Chaos: A Very Short Introduction," by Leonard Smith (2007, Oxford University Press, 180pp, paperback, ISBN: 978-0-19-285-378-3), and it is a good, quick read. There is a short review in the Journal of Physics A which says, in part,
Anyone who ever tried to give a popular science account of research knows that this is a more challenging task than writing an ordinary research article. Lenny Smith brilliantly succeeds to explain in words, in pictures and by using intuitive models the essence of mathematical dynamical systems theory and time series analysis as it applies to the modern world.

[...]

However, the book will be of interest to anyone who is looking for a very short account on fundamental problems and principles in modern nonlinear science.
The only criticism offered in that review was of the low-resolution of some of the figures (which is hard to fix since it is a pocket-size format).

I'm reproducing most of one of the reviews from Amazon here because the criticisms which it claims make the book unsuitable as an 'intro to chaos' are the things I enjoyed about it:
This book starts out promising but, as one goes along, it drifts farther and farther from what an introduction to chaos should be.

In particular, the book turns out to be largely a discussion of modeling and forecasting, with some emphasis on the relevant implications of chaos. Moreover, most of the examples and applications relate to weather and climate, which becomes boring after a while (especially considering the abundance of other options). Smith's bio reveals that this is exactly his specialty, so the book appears to be heavily shaped by his background and interests, rather than what's best for a general audience. As a result, many standard and important topics in chaos theory recieve little or no mention, and I think the book fails as a proper introduction to chaos.

[...]

Considering all of this, I can recommend the book only to people who are particularly interested in modeling, forecasting, and the relevant implications of chaos, especially as this relates to weather and climate. In this context, Smith's discussion of the differences between mathematical, physical, statistical, and philosophical perspectives is particularly insightful and useful.

Well, since I think the intersection between public policy and computational physics is an interesting one, this book turned out to be right up my alley. It was an entertaining read, and I did not have to work too hard to translate the simple language Smith used to appeal to a wide audience back into familiar technical concepts. That's no mean feat.

I do have a somewhat significant bit of criticism about his treatment of the tractability of getting probabilistic forecasts in the case of chaotic physical systems for which we don't know the correct model. If you've read some of my posts on Jaynes' book you can probably guess what I'm going to say. But first, here's what Smith says:
With her perfect model, our 21st-century demon can compute probabilities that are useful as such. Why can't we? There are statisticians who argue we can, including perhaps a reviewer of this book, who form one component of a wider group of statisticians who call themselves Bayesians. Most Bayesians quite reasonably insist on using the concepts of probability correctly [this was Jaynes' main pedagogical point], but there is a small but vocal cult among them that confuse the diversity seen in our models for uncertainty in the real world. Just as it is a mistake to use the concept of probability incorrectly, it is an error to apply them where they do not belong.
There is then some illustration of this 'model inadequacy' problem which is correct as far as noting that the model is not reality, only a possibly useful shadow, but fails to support the assertion that the 'vocal group' is misapplying probability theory. Smith continues,
Would it not be a double-sense to proffer probability forecasts one knew were conditioned on an imperfect model as if they reflected the likelihood of future events, regardless of what small print appeared under the forecast?
This is an oblique criticism of the Bayesian approach, which would, of course, give predictive distributions conditional on the model or models used in the analysis. Smith's criticism is that the ensemble of models may not contain the 'correct' model, so the posterior predictive distribution is not a probability in the frequentist sense. Of course, no Bayesian would claim that it is, only that it best captures our current state of knowledge about the future and is the only procedure that enables coherent inference in general. Everything else is ad hockery, as Jaynes would say. Any prediction of the future is conditioned on our present state of knowledge (which includes, among other things, the choice of models) and the data we have. The only question then is, do we explicitly acknowledge that fact or not?

Another thing that bothered me was the sort of dismissive way he commented on the current state of model adequacy in the physical sciences:
... is the belief in the existence of mathematically precise Laws of Nature, whether deterministic or stochastic, any less wishful thinking than the hope that we will come across any of our various demons offering forecasts in the woods?

In any event, it seems we do not currently know the relevant equations for simple physical systems, or for complicated ones.
There are plenty of practising engineers and applied physicists using non-linear models to make successful predictions who I think would be quite surprised to hear that their models, or the conservation laws on which they are based, do not exist.

Other than those two minor quibbles, it was a very good book and an enjoyable read.

A nice feature of the book (quite suitable for a 'very short intro') is the 'further reading' list at the end, here's a couple that looked interesting:

Friday, December 4, 2009

Bayesian Climate Model Averaging

Another instalment in the 'lack-of-climate-model-validation-bothers-me' series (see Lindzen's talk for a good intro). I've been reading Jaynes' book lately so naturally the Bayesian approach to the issue seems most germane. The whole climate science / public policy intersection can be viewed as one big decision theory problem (acting to maximize utility under uncertainty). To come out of that game well you generally need to have good models (hopefully with nice, tight predictive distributions) and smooth, gently sloping loss functions. I'll leave the loss functions for now and focus on the modelling aspect (since that's what I'm familiar with).

Validation (comparing the model predictions to experimental observations) is generally what allows you to find out if you've made good choices of what model structure to use, what physics to include and what physics to neglect. In most applications of computational physics this is a straight-forward (if sometimes expensive) process. The problem is harder with climate models. We can't do designed experiments on the Earth.

Here are a couple of choice quotes from Reichler and Kim 2007 about the difficulties of climate model validation.
Several important issues complicate the model validation process. First, identifying model errors is difficult because of the complex and sometimes poorly understood nature of climate itself, making it difficult to decide which of the many aspects of climate are important for a good simulation. Second, climate models must be compared against present (e.g., 1979-1999) or past climate, since verifying observations for future climate are unavailable. Present climate, however, is not an independent data set since it has already been used for the model development (Williamson 1995). On the other hand, information about past climate carries large inherit uncertainties, complicating the validation process of past climate simulations (e.g., Schmidt et al. 2004). Third, there is a lack of reliable and consistent observations for present climate, and some climate processes occur at temporal or spatial scales that are either unobservable or unresolvable. Finally, good model performance evaluated from the present climate does not necessarily guarantee reliable predictions of future climate (Murphy et al. 2004).

The above quoted paper is a comparison of three generations of IPCC-family models. The study shows improvement in prediction of modern climate as the models improve from 1990 to 2001 to 2007. It also shows that the ensemble mean is more skilled than any individaul model (more on this later). The reasons given to explain the improvement make intuitive sense:
Two developments, more realistic parameterizations and finer resolutions, are likely to be most responsible for the good performance seen in the latest model generation. For example, there has been a constant refinement over the years in how sub-grid scale processes are parameterized in models. Current models also tend to have higher vertical and horizontal resolution than their predecessors. Higher resolution reduces the dependency of models on parameterizations, eliminating problems since parameterizations are not always entirely physical. That increased resolution improves model performance has been shown in various previous studies (e.g., Mullen and Buizza 2002, Mo et al. 2005, Roeckner et al. 2006).

A problem faced by climate modelers is that it is unlikely that we'll be able to run grid-resolved solutions of the climate within the lifetime of anyone now living (us CFD guys have the same problem with grid resolution scaling for high Reynolds number flows). There will always be a need for 'sub-grid' parameterizations, the hope is that eventually they will become "entirely physical" and well calibrated (if you think they already are, then you have been taken by someone's propaganda).

Bayesian model averaging (BMA) is one way to account for our uncertainty in model structure / physics choices. Instead of choosing a 'right' model, we get predictive distributions for things we care about by marginalizing over the uncertain model structures (and the uncertain parameters too).This paper shows that it is a useful procedure for short-term forecasting. The benefit with short-term forecasts is that we can evaluate the accuracy by closing the loop between predictions and observations. Min and Hense apply this idea to the IPCC AR4 coupled-climate models. Here's a short snippet from that paper providing some motivation for the use of BMA:
However, more than 50% of the models with anthropogenic-only forcing cannot reproduce the observed warming reasonably. This indicates the important role of natural forcing although other factors like different climate sensitivity, forcing uncertainty, and a climate drift might be responsible for the discrepancy in anthropogenic-only models. Besides, Bayesian and conventional skill comparisons demonstrate that a skill-weighted average with the Bayes factors (Bayesian model averaging, BMA) overwhelms the arithmetic ensemble mean and three other weighted averages based on conventional statistics, illuminating future applicability of BMA to climate predictions.

The ensemble means or Bayesian averages tend to outperform individual models, but why is this? Here's what R&K2007 has to say:
Our results indicate that multi-model ensembles are a legitimate and effective means to improve the outcome of climate simulations. As yet, it is not exactly clear why the multi-model mean is better than any individual model. One possible explanation is that the model solutions scatter more or less evenly about the truth (unless the errors are systematic), and the errors behave like random noise that can be efficiently removed by averaging. Such noise arises from internal climate variability (Barnett et al. 1994), and probably to a much larger extent from uncertainties in the formulation of models (Murphy et al. 2004; Stainforth et al. 2005).

Another interesting paper that explores this finds that models which have good scores on the calibration data do not tend to outperform other models over a subsequent validation period.
Error in the ensemble mean decreases systematically with ensemble size, N, and for a random selection as approximately 1∕Na, where a lies between 0.6 and 1. This is larger than the exponent of a random sample (a = 0.5) and appears to be an indicator of systematic bias in the model simulations.

This should not be surprising, it is very difficult to get all of the physics right (and remove the systematic bias) when you aren't able to do no-kidding validation experiments. They begin their conclusion with
In our analysis there is no evidence of future prediction skill delivered by past performance-based model selection. There seems to be little persistence in relative model skill, as illustrated by the percentage turnover in Figure 3. We speculate that the cause of this behavior is the non-stationarity of climate feedback strengths. Models that respond accurately in one period are likely to have the correct feedback strength at that time. However, the feedback strength and forcing is not stationary, favoring no particular model or groups of models consistently.

This means it is very difficult to protect ourselves from 'over-fitting' the models to our available historical record, and it certainly indicates that we should be cautious in basing policy decision on climate model forecasts. The 'science is settled' crowd, while busy banging the consensus drum and clamouring for urgent action (NOW!), never seem to offer this sort of nuanced approach to policy though.

If you have read any good, recent climate model validation papers please post them in the comments. Please don't post polemics about polar bears and arctic sea ice, my skepticism is honest, your activism should be too.

For some further Bayes Model Averaging / Model Selection check out:

(isn't Google Books cool?)

Wednesday, November 25, 2009

Converging and Diverging Views

I was brushing up on my maximum entropy and probability theory the other day and came across a great passage in Jaynes' book about convergence and divergence of views. He applies basic Bayesian probability theory to the concept of changing public opinion in the face of new data, especially the effect prior states of knowledge (prior probabilities) can have on the dynamics. The initial portion of section 5.3 is reproduced below.

5.3 Converging and diverging views (pp. 126 – 129)

Suppose that two people, Mr A and Mr B have differing views (due to their differing prior information) about some issue, say the truth or falsity of some controversial proposition S. Now we give them both a number of new pieces of information or ’data’, D1,D2,,Dn, some favorable to S, some unfavorable. As n increases, the totality of their information comes to be more nearly the same, therefore we might expect that their opinions about S will converge toward a common agreement. Indeed, some authors consider this so obvious that they see no need to demonstrate it explicitly, while Howson and Urbach (1989, p. 290) claim to have demonstrated it.

Nevertheless, let us see for ourselves whether probability theory can reproduce such phenomena. Denote the prior information by IA, IB, respectively, and let Mr A be initially a believer, Mr B be a doubter:



P (S|IA) ≃ 1, P(S |IB ) ≃ 0
(5.16)

after receiving data D, their posterior probabilities are changed to



 P (D |SIA) P(S |D IA) = P (S|IA)---------- P (D |IA )
(5.17)




 P-(D-|SIB-) P(S |D IB) = P (S|IB) P (D |IB )
(5.17)

If D supports S, then since Mr A already considers S almost certainly true, we have P(D|S IA), and so



P (S |D IA) ≃ P (S |IA)
(5.18)

Data D have no appreciable effect on Mr A’s opinion. But now one would think that if Mr B reasons soundly, he must recognize that P(D|S IB) > P(D|IB), and thus



P (S |D I ) > P (S |I ) B B
(5.19)

Mr B’s opinion should be changed in the direction of Mr A’s. Likewise, if D had tended to refute
S, one would expect that Mr B’s opinions are little changed by it, whereas Mr A’s will move in the direction of Mr B’s. From this we might conjecture that, whatever the new information D, it should tend to bring different people into closer agreement with each other, in the sense that



|P (S|D I ) - P (S |D I )| < |P (S|I ) - P (S|I )| A B A B
(5.20)

Although this can be verified in special cases, it is not true in general.

Is there some other measure of ‘closeness of agreement’ such as log[P(S|D Ia)∕P(S|D IB], for which this converging of opinions can be proved as a general theorem? Not even this is possible; the failure of probability theory to give this expected result tells us that convergence of views is not a general phenomenon. For robots and humans who reason according to the consistency desiderata of Chapter 1, something more subtle and sophisticated is at work.

Indeed, in practice we find that this convergence of opinions usually happens for small children; for adults it happens sometimes but not always. For example, new experimental evidence does cause scientists to come into closer agreement with each other about the explanation of a phenomenon.

Then it might be thought (and for some it is an article of faith in democracy) that open discussion of public issues would tend to bring about a general consensus on them. On the contrary, we observe repeatedly that when some controversial issue has been discussed vigorously for a few years, society becomes polarized into opposite extreme camps; it is almost impossible to find anyone who retains a moderate view. The Dreyfus affair in France which tore the nation apart for 20 years, is one of the most thoroughly documented examples of this (Bredin, 1986). Today, such issues as nuclear power, abortion, criminal justice, etc., are following the same course. New information given simultaneously to different people may cause a convergence of views; but it may equally well cause a divergence.

This divergence phenomenon is observed also in relatively well-controlled psychological experiments. Some have concluded that people reason in a basically irrational way; prejudices seem to be strengthened by new information which ought to have the opposite effect. Kahneman and Tversky (1972) draw the opposite conclusion from such psychological tests, and consider them an argument against Bayesian methods.

But now in view of the above ESP example, we wonder whether probability theory might also account for this divergence and indicate that people may be, after all, thinking in a reasonably rational, Bayesian way (i.e. in a way consistent with their prior information and prior beliefs). The key to the ESP example is that our new information was not

S fully adequate precautions against error or deception were taken, and Mrs Stewart did in fact deliver that phenomenal performance.

It was that some ESP researcher has claimed that S is true. But if our prior probability for S is lower than our prior probability that we are being deceived, hearing this claim has the opposite effect on our state of belief from what the claimant intended.

The same is true in science and politics; the new information a scientist gets is not that an experiment did in fact yield this result, with adequate protection against error. It is that some colleague has claimed that it did. The information we get from TV evening news is not that a certain event actually happened in a certain way; it is that some news reporter claimed that it did.

Scientists can reach agreement quickly because we trust our experimental colleagues to have high standards of intellectual honesty and sharp perception to detect possible sources of error. And this belief is justified because, after all, hundreds of new experiments are reported every month, but only about once in a decade is an experiment reported that turns out later to have been wrong. So our prior probability for deception is very low; like trusting children, we believe what experimentalists tell us.

In politics, we have a very different situation. Not only do we doubt a politician’s promises, few people believe that news reporters deal truthfully and objectively with economic, social, or political topics. We are convinced that virtually all news reporting is selective and distorted, designed not to report the facts, but to indoctrinate us in the reporter’s socio-political views. And this belief is justified abundantly by the internal evidence in the reporter’s own product – every choice of words and inflection of voice shifting the bias invariably in the same direction.

Not only in political speeches and news reporting, but wherever we seek for information on political matters, we run up against this same obstacle; we cannot trust anyone to tell us the truth, because we perceive that everyone who wants to talk about it is motivated either by self-interest or by ideology. In political matters, whatever the source of information, our prior probability for deception is always very high. However, it is not obvious whether this alone can prevent us from coming to agreement.

Jaynes, E.T., Probability Theory: The Logic of Science (Vol 1), Cambridge University Press, 2003.