Wednesday, November 11, 2020
Missiles and Rockets
Sunday, March 22, 2015
Reliability Growth: Enhancing Defense System Reliability
This report (pdf) from the National academies on reliability growth is interesting. There's a lot of good stuff on design for reliability, physics of failure, highly accelerated life testing, accelerated life testing and reliability growth modeling. Especially useful is the discussion about the suitability of assumptions underlying some of the different reliability growth models.
The authors provide a thorough critique of MIL-HDBK-217, Reliability Prediction of Electronic Equipment, in Appendix D, which is probably worth the price of admission by itself. If you're concerned with product reliability you should read this report (lots of good pointers to the lit).
Tuesday, January 13, 2015
Guidelines for Planning and Evidence for Assessing a Well-Designed Experiment
This paper is full of great guidance for planning a campaign of experimentation, or assessing the sufficiency of a plan that already exists. The authors break up the effort into four phases:
- Plan a Series of Experiments to Accelerate Discovery
- Design Alternatives to Span the Factor Space
- Decide on a Design Strategy to Control the Risk of Wrong Conclusions
- Execute the Test
- Analyze the Experimental Design
Monday, January 13, 2014
Flight Demo Program Lessons Learned
It must be remembered that there is nothing more difficult to plan, more doubtful of success, nor more dangerous to manage than creation of a new system. For the initiator has the enmity of all who would profit by the preservation of the old institutions, and merely lukewarm defenders in those who would gain by the new ones.
The Prince, Machiavelli, 1513
Here are the rules compiled based on previous flight demonstration program experience:
- Agree to clearly defined program objectives in advance
- Single manager under one agency
- Small government and contractor program offices
- Build competitive hardware, not paper
- Focus on key demonstrations, not everything
- Streamlined documentation and reviews
- Contractor integrates and tests prototype
- Develop minimum realistic funding profiles
- Track cost/schedule in near real time
- Mutual trust essential
Friday, January 10, 2014
RAND: no life-cycle cost savings from joint aircraft
- Joint aircraft programs have not historically saved overall life cycle cost. On average, such programs experienced substantially higher cost growth in acquisition (research, development, test, evaluation, and procurement) than single-service programs. The potential savings in joint aircraft acquisition and operations and support compared with equivalent single-service programs is too small to offset the additional average cost growth that joint aircraft programs experience in the acquisition phase.
- The difficulty of reconciling diverse service requirements in a common design is a major factor in joint cost outcomes. Diverse service requirements and operating environments work against commonality which is the source of potential cost savings, and are a major contributor to the joint acquisition cost-growth premium identified in the cost analysis.
- Historical analysis suggests joint programs are associated with contraction of the industrial base and a decline in potential future industry competition, as well as increased strategic and operational risk due to dependency across the services on a single type of weapon system which may experience unanticipated safety, maintenance, or performance issues with no alternative readily available.
In the past 50 years, the U.S. Department of Defense has pursued numerous joint aircraft programs, the largest and most recent of which is the F-35 Joint Strike Fighter (JSF). Joint aircraft programs are thought to reduce Life Cycle Cost (LCC) by eliminating duplicate research, development, test, and evaluation efforts and by realizing economies of scale in procurement, operations, and support. But the need to accommodate different service requirements in a single design or common design family can lead to greater program complexity, increased technical risk, and common functionality or increased weight in excess of that needed for some variants, potentially leading to higher overall cost, despite these efficiencies. To help Air Force leaders (and acquisition decisionmakers in general) select an appropriate acquisition strategy for future combat aircraft, this report analyzes the costs and savings of joint aircraft acquisition programs. The project team examined whether historical joint aircraft programs have saved LCC compared with single-service programs. In addition, the project team assessed whether JSF is on track to achieving the joint savings originally anticipated at the beginning of full-scale development. Also examined were the implications of joint fighter programs for the health of the industrial base and for operational and strategic risk.JSF is now expected to be more expensive than 3 F-22-like single-service programs:
Saturday, November 2, 2013
LockMart SR-72
![]() |
| NASA SP-2001-4525 |
![]() |
| YF-12 Inlet Schematic |
Tuesday, October 29, 2013
Defense Acquisitions: Where Should Reform Aim Next?
Let's just skip the acquisition reform charade... Another blue-ribbon study, more legislation and a new slogan will not make it happen at last...
Prof. H. Sapolsky, MIT
Luckily, this latest GAO report on the problem has no sweeping "blue ribbon"-style reform suggestions. There seems to be a pretty pragmatic recognition that perverse incentives are largely the reason that many of the grandiose reform efforts of the past were doomed to failure:
Reforms that focus on the methodological procedures of the acquisition process are only partial remedies because they do not address incentives to deviate from sound practices. Weapons acquisition is a complicated enterprise, complete with unintended incentives that encourage moving programs forward by delaying testing and employing other problematic practices. These incentives stem from several factors. For example, the different participants in the acquisition process impose conflicting demands on weapon programs so that their purpose transcends just filling voids in military capability. Also, the budget process forces funding decisions to be made well in advance of program decisions, which encourages undue optimism about program risks and costs. Finally, DOD program managers' short tenures and limitations in experience and training can foster a short-term focus and put them at a disadvantage with their industry counterparts.
Drawing on its extensive body of work in weapon systems acquisition, GAO sees several areas of focus regarding where to go from here:
These areas are not intended to be all-encompassing, but rather, practical places to start the hard work of realigning incentives with desired results.
- at the start of new programs, using funding decisions to reinforce desirable principles such as well-informed acquisition strategies;
- identifying significant risks up front and resourcing them;
- exploring ways to align budget decisions and program decisions more closely; and
- attracting, training, and retaining acquisition staff and managers so that they are both empowered and accountable for program outcomes.
Wednesday, August 21, 2013
3-D Printing in DoD: Who's Dragging Their Feet?
With possible dwindling budgets on the horizon, a clear strategy and cohesive approach is essential to create efficiencies in the area of research and development as well as eliminating duplicative efforts. In order for DoD to take advantage of what is anticipated to be an explosion in the commercial sector within the next ten years, the Department must take an active approach, partnering with the private sector to keep up with this relatively nascent technology and shaping/guiding it towards the desired end state the department has in mind.I think an "additive manufacturing Czar" sounds like a terrible idea (so I'm sure it will secure funding for some beltway bandits to do a study). I know my recent success with qualifying a particular additive manufacturing process and supplier for use in 3D printing wind tunnel models did not need a Pentagon king-pin to tell me about DoD's strategy for additive manufacturing. Using this technology just made sense as a way to solve my problem: get a complex wind-tunnel model rapidly, and at an affordable cost. I did not receive top-down direction or guidance to use AM, I simply took the initiative to solve my problem. After reading that article I'm left wondering, just how exactly is waiting on direction from the very heights of the bureaucracy supposed to lead to innovation?
One step towards a clear strategy and cohesive approach is for DoD to designate an AM Czar within the Department. They could serve as a single point for all things AM and not the myriad of technical advisory boards that currently exist. This office could then work with policy makers to execute and monitor a strategy which will allow DoD to take full advantage of this technology. Logically, this office would interface directly with the National Additive Manufacturing and Innovation Institute (NAMII) as DoD's representative
3-D Printing Revolution in Military Logistics
Monday, March 4, 2013
Drone Marketing, Same Song Second Verse
Update: A new one from LockMart.
Friday, December 16, 2011
McCain's Hangar Queen
Another example of how flawed the Pentagon’s weapons procurement process is can be found in the F-22 RAPTOR program. When the Pentagon and the defense industry originally conceived of the F-22 in the mid-1980s, they intended it to serve as a revolutionary solution to the Air Force’s need to maintain air superiority in the face of the Soviet threat during the Cold War. The F-22 obtained ‘full operational capability’ twenty years later -- well after the Soviet Union dissolved. When it finally emerged from its extended testing and development phase, the F-22 was recognized as a very capable tactical fighter, probably the best in the world for some time to come. But, plagued with developmental and technical issues that caused the cost of buying to go through the roof, not only was the F-22 twenty years in the making, but the process has proved so costly that the Pentagon could ultimately afford only 187 of the planes -- rather than the 750 it originally planned to buy.
“Unfortunately, the F-22 also ended up being effectively too expensive to operate compared to the legacy aircraft it was designed to replace. It also ended up largely irrelevant to the most predominant current threats to national security -- terrorists, insurgencies, and other non-state actors. In fact, if one were to set aside the F-22’s occasional appearances in recent big-budget Hollywood movies where it has been featured fighting aliens and giant robots, the F-22 has to this day not flown a single combat sortie -- despite that we have been at war for 10 years as of this September and recently supported a no-fly zone in Libya.
“Politically engineered to draw in over 1,000 suppliers from 44 states represented by key Members of Congress and, by the estimates of prime contractor Lockheed Martin, directly or indirectly supporting 95,000 jobs, there can be little doubt that the program kept being extended far longer than it should have been -- ultimately to the detriment to the taxpayer and the warfighter. As such, it remains an excellent example of how much our defense procurement process has been in need in reform. We may fight a near-peer military competitor with a fifth-generation fighter capability someday, but we have been at war for 10 years and until a few months ago had been helping NATO with a no-fly zone in Libya. And, this enormously expensive aircraft sat out both campaigns.
“Moreover, as a result of problems with its OBOG (On-Board Oxygen Generating) system, which has caused pilots to get dizzy or, in some cases, lose consciousness from lack of oxygen, on May 3, 2011, the Air Force grounded its entire fleet of F-22s. While this grounding was lifted earlier this year, exactly why F-22 pilots have been experiencing hypoxia remains unknown -- but similar unexplained incidents continue.
“And then, there is the issue of the sky-rocketing maintenance costs to the Air Force in trying to sustain a barely adequate ‘mission capable rate’ for the F-22. Its seems that the ‘plug and play’ component maintenance features that were supposed to reduce costs for the Air Force over the life cycle of the aircraft doesn’t really play well. And, each time a panel is opened for maintenance, the costs to repair the ‘low-observable’ surface in order to maintain its stealthiness have made this critical feature of the aircraft cost-prohibitive to sustain over the long-run. Finally, it seems that the engineers and technicians designing the F-22 forgot a basic law of physics during some point of the development phase -- that dissimilar metals in contact with each other have a tendency to corrode. The Air Force is now faced with a huge maintenance headache costing over hundreds of millions of dollars-and-growing to keep all 168 F-22s sitting on the ramp from corroding from the inside out.
One thing is clear: because of a problem directly attributable to how aggressively the F-22 was acquired -- procuring significant quantities of aircraft without having conducted careful developmental testing and reliably estimating how much they will cost to own and operate -- the 168 F-22s, costing over $200 million each, may very well become the most expensive corroding hanger queens ever in the history of modern military aviation.
REMARKS BY SENATOR JOHN McCAIN ON THE “MILITARY-INDUSTRIAL-CONGRESSIONAL” COMPLEX
What bothers me most about the F-22 is not the maintenance problems, operational difficulties, or danger to air-crews. These require mere engineering fixes. What really bothers me is the damage these huge programs cause to rational risk assessment and national security.
"The reality is we are fighting two wars, in Iraq and Afghanistan, and the F-22 has not performed a single mission in either theater," Gates told a Senate committee last week.
Carlson ["We think that [187 planes] is the wrong number"], however, told a group of reporters earlier in the week that the Air Force was "committed to funding 380" of the fighters, regardless of the Bush administration's decision.
According to an Air Force official briefed on the Thursday rebuke, Gates telephoned Air Force Secretary Michael W. Wynne, who was on vacation at the time, to express his displeasure with Carlson.
The senior defense official said Carlson's remarks, reported Thursday by the trade publication Aerospace Daily, angered the Pentagon's top leadership, adding that they were "completely unacceptable and out of line."
Fighter Dispute Hits Stratosphere
That's not something that more money or more technology can fix.
Friday, September 2, 2011
Airframer for Dayton Aerospace Cluster
Here's some of the defense contractors on the Fortune 100:
- 36. Boeing (already makes UAVs: Insitu)
- 44. United Technologies (claims it doesn't want to enter the UAV market, but UT subsidiary Sikorsky has made significant investments in a small UAV company)
- 52. Lockheed Martin (already makes UAVs: Sentinel)
- 72. Northrup Grumman (already makes UAVs: Global Hawk)
- 86. General Dynamics (sort of makes UAVs: joint venture with Elbit, joint venture with Aurora, joint venture with Aeronautics Defense Systems)
Sunday, April 17, 2011
Acquisition Death Spirals
This isn't about the normal death spiral of increasing unit costs driving production cuts, which increases unit costs, which drives production cuts, which.... It's about another sort of price spiral caused by the US government's infatuation with sole-sourcing critical capabilities. I think this is largely due to technocrats trusting simple, static industrial-age cost models which support decisions dominated by returns from economies of scale. The basic logic of the decisions these models support (an equilibrium solution) is: "things will be cheaper with one supplier because the overhead will be amortized over bigger quantities."
The basic mistake these models make is neglecting the dynamics. Price is a dynamic thing. If the capability is very critical, and there is only one supplier, then there is almost no ceiling on how high the price can rise. The price level reached under those dynamics is just below the point where you'd stop paying for the capability in favor of a more important one (I'll call this the buyer's "level of pain"). The alternative dynamics occurs when there are multiple competing suppliers. The price reached under these dynamics asymptotes towards the economic costs (this takes into account barriers to entry / opportunity costs). Here's a simple graph illustrating price behavior under these two situations.
![]() |
| Price Dynamics |
We see these dynamics play out in a variety of defense acquisitions. The F-35 engine program and the Evolved Expendable Launch Vehicle program are two exemplars that are currently making the news.
The F-35 engine procurement was initially structured to support two suppliers during development, much like the engine programs for F-15 and F-16. There are operational advantages to having two engines. If a problem is found in one model, only half of the Air Force's tactical aircraft would have to be grounded while the solution is found. The other advantage comes from the suppliers competing with each other on price for various lots of engines. A disadvantage is overhead and development cost for the two suppliers and possibly increased logistics footprint for supporting two different engine models.
The recent news is that the budget does not include funds for the second engine. Not a week after this budget is passed which makes the engine buy a sole-source deal, we have an Undersecretary of Defense for Acquisition complaining about the price from the remaining supplier.
"I'm not happy, as I am with so many parts of all our programs, with (the P&W engine's) cost performance so far," Carter told the House of Representatives Appropriations subcommittee on defense on Wednesday. "We need to drive the costs down." [...] "Our analysis does not show the payback," Carter told the subcommittee. He added that "people of good will come to different conclusions on this issue." U.S. "not happy" with F-35 engine cost overrunsDid the analysis include the second supplier offering to assume the risks and go fixed price on the development? Is complaining about the price, being really unhappy about paying it, but paying it anyway because there is no alternative anything but empty political theater? Makes for great content in the trade rags: Acquisition Official gives contractor a stern talking to! Contractor hangs head in a suitably chastened way, "yes, our prices are very high for these unique capabilities, we are working hard to contain the costs for our customer." Much harrumphing is heard from various congresscritters, meanwhile the price continues to spiral higher...
In the case of the EELV, despite the fact that the Air Force paid for two parallel rocket development programs we now have just a single supplier. The two launch service providers were so expensive, they could not compete in the commercial market. The business case for the two rockets hinged on them being able to make money in the commercial market and get their launch rates up. When no other customers but the US government could afford their high prices they had to combine into the single consortium: ULA. So it's sole-source with two vehicles. Recall the various advantages and disadvantages of developing two products discussed above in the case of the F-35 engine. Now EELV has the worst of both worlds: high overhead and logistics costs to support two vehicles, and no competition or customer diversification to get flight rates up and bring prices down.
Launch-service providers agree that their viability, as well as their ability to keep costs down, is based on launch rhythm. The more often a vehicle launches, the more reliable it becomes. Scale economies are introduced as well in a virtuous cycle. One U.S. government official agreed that if SpaceX is now allowed to break ULA’s monopoly on U.S. government satellite launches as indicated by the memorandum of agreement, it could force ULA’s already high prices even higher as it eats into ULA’s current market. “In the longer term we may be faced with questions about whether one of them [ULA or SpaceX] can remain viable without direct subsidies — the same questions we faced with ULA,” this official said. “Then what do we do? We have a policy of assured access to space, which means at least two vehicles. The demand for launches has not increased since ULA was formed, so we could be heading toward a nearly identical situation in a few years. But we are spending taxpayers’ money and if we can find reliable launches that are less expensive, we are not going to ignore that.” SpaceX Receives Boost in Bid To Loft National Security SatellitesIt is interesting to note the thinking of the unnamed government official. He doesn't recognize the dynamics of the situation. He's living in a sole-source mindset, it just happens that he's going to change to this new, lower cost source.
NASA has a pricing model that shows savings from "outsourcing development", but not because of any interesting dynamics that they've included. The justification is the same as those underlying the broken decisions about aircraft engines: "returns from economies of scale". The dynamics of competition remains ignored.
NASA Deputy Administrator Lori Garver, in a separate presentation here April 12, said the agency’s policy of pushing rocket-development work onto the private sector will only reach maximum benefit if other customers also purchase the vehicles developed initially with NASA funding. Referring specifically to SpaceX, Garver said a conventional NASA procurement of a Falcon 9-class rocket would cost nearly $4.5 billion according to a NASA-U.S. Air Force cost model that includes the vehicle’s first flight. Outsourcing development to SpaceX, she said, would cut that figure by 60 percent, but only if other customers purchase the vehicle, thus permitting scale economies to reach maximum effect. After Servicing Space Station SpaceXs Priority is Taking on EELV
The way out of the death spiral is program dependent. In the EELV case it took a new entrant who has signed commercial contracts in addition to chasing the government launches (from a couple different agencies). SpaceX has built in significant customer diversification that ULA never developed (though this was hoped for in the early justifications of the program structure). Why did Lockheed-Martin and Boeing, and subsequently ULA never develop this customer diversification? Because they didn't have to. The government guaranteed their continued existence. SpaceX, on the other hand, has no such guarantee. The only option for their continued existence is to make a profit from more than one customer. In the F-35 engine case, DoD has decided that even the assumption of development risk by the second supplier under a fixed price contract is not enough to close the case. This puzzles me. Maybe the second engine has become a symbol of "duplication and waste" rather than "competition and efficiency". If so, then DoD has given the primary engine supplier a way to "frame" their competition out of existence with political argument (removing the need for them to earn market share honestly).
The outlook for lower cost space access looks good. However, as long as the DoD is stuck in its current frame, there is great opportunity for political theater that will serve mainly as a content generator for defense trade publications and a distraction from the root cause of steadily rising tactical aircraft engine costs into the future.
Friday, February 4, 2011
Validation and Calibration: more flowcharts
|
|
- The partitioning of parameters into “known” and “unknown” based on what level of the hierarchy (component, subsystem, system) you are at in the “bottom-up” calibration process. Our (properly formulated) models should tell us how much information different types of test data give us about the different parameters. Parameters should always be described by a distribution rather than discrete switches like known or unknown.
- The approach is based entirely on the likelihood (but they do mention something that sounds like expert priors in passing).
- They claim that the proposed calibration method enhances “predictive capability” (section 3), however this is misleading abuse of terminology. Certainly the in-sample performance is improved by calibration, but the whole point of making a distinction between calibration and validation is based on recognizing that this says little about the out-of-sample performance (in fairness, they do equivocate a bit on this point, “The authors acknowledge that it is difficult to assure the predictive capability of an improved model without the assumption that the randomness in the true response primarily comes from the the randomness in random model variables.”).
]
References
Monday, January 31, 2011
Tuesday, August 3, 2010
No Fluid Dynamicist Kings in Flight-Test
This was a guest post over on Pielke's site.
Dr Pielke's Honest Broker concepts resonate with me because of practical decision support experiences I've had, and this post is an attempt to share some of those from a realm pretty far removed from the geosciences. All the views and opinions expressed are my own and in no way represent the position or policy of the US Air Force, Department of Defense or US Government. I am writing as a simple student of good decision making. My background is not climate science. I am an Aeronautical Engineer with a background in computational fluid dynamics, flight test and weapons development. I got interested in the discussions of climate policy because the intersection of computational physics and decision making under uncertainty is an interesting one no matter what the subject area. The discussion in this area is much more public than the ones I'm accustomed to, so it makes a great target of opportunity. The decision support concepts Dr Pielke discusses make so much sense to me now, but I can see how hard they are for technical folks to grasp because I used to be a very linear thinker when I was a young engineer right out of school.
My journeyman's education in decision support came when I got the chance to lead a small team doing Live Fire Test and Evaluation for the Air Force (you may not be familiar with LFT&E, it is a requirement that grew out of the Army gaming testing of the Bradley fighting vehicle in the 1980s, a situation that was fairly accurately lampooned in the movie "Pentagon Wars"). The competing values of the different stakeholders (folks appointed by congress to ensure sufficient realistic testing compared to folks at the service level doing product development) was really an eye-opening education for a technical nerd like me. I initially thought, "if only everyone can agree on the facts, the proper course of action will be clear". How naive I was! Thankfully, the very experienced fellows working for me didn't mind training up a rash, newly-minted, young Captain.
It's tough for some technical specialists (engineers/scientists) to recognize worthy objectives their field of study doesn't encompass. The reaction I see from the more technically oriented folks like Tobis (see how he struggles) reminds me a lot of the reaction that engineers in product development offices would have to the role of my little Live Fire office. A difficulty we often encountered was the LFT&E oversight folks wanted to accomplish testing that didn't have direct payoff to narrower product development goals that concerned the engineers. "What those people want to do is wasteful and stupid!" This parallels the recent sand berm example. The preferred explanation from the technician's perspective is that the other guy is bat-shit crazy, and his views should be ridiculed and de-legitimized. The truth is usually closer to the other guy having different objectives that aren't contained within the realm of the technician's expertise. In fact, the other person is probably being quite rational, given their priors, utility function and state of knowledge.
In my little Live Fire Office we had lots of discussion about what to call the role we did, and how to best explain it to the program managers. I wish I had heard of Dr Pielke's book back then, because "Honest Broker" would have been an apt description for much of the role. We acted as a broker between the folks in the Pentagon with the mandate from congress for sufficient, realistic testing, and the Air Force level program office with the mandate for product development. The value we brought (as we saw it), was that we were separate from the direct program office chain of command (so we weren't advocates for their position), but we understood the technical details of the particular system, and we also understood the differing values of the folks in the Pentagon (which the folks in the program office loved to refuse to acknowledge as legitimate, sound familiar?). That position turns out to be a tough sell (program managers get offended if you seem to imply they are dishonest), so I can empathize with the virulent reaction Dr Pielke gets on applying the Honest Broker concepts to climate policy decision support. People love to take offense over their honor. That's a difficult snare to avoid while you try to make clear that, while there's nothing dishonest about advocacy, there remains significant value in honest brokering. Maybe Honest Broker wouldn't be the best title to assume though. The first reaction out of a tight-fisted program manager would likely be "I'm honest, why do I need you?"
One of the reason my little office existed was because of some "lessons learned" from the Tri-Service Standoff Missile debacle (all good things in defense acquisition must grow out of historical buffoonery). The broader Air Force leadership realized that it was counterproductive to have product development engineers and program managers constantly trying to de-legitimize the different values that the oversight stake-holders brought (the differences springing largely from different appetites for risk and priors for deception) by wrangling over largely inconsequential, technical nits (like tree rings in the Climate Wars). The wiser approach was to maintain an expertise whose sole job was to recognize and understand the legitimate concerns of the oversight folks and incorporate those into a decision that meets the service's constraints as quickly and efficiently as possible. Rather than wasting time arguing, product development folks could focus on product development.
The other area where I've seen this dynamic play out is in making flight test decisions. In that case though, the values of all the stake-holders tend to align more closely, so the separation between technical expertise and decision making is less contentious (Dr Pielke's Tornado analogy). In contrast to the climate realm where it's argued that science compels because we're in the Tornado mode, the flight-test engineers understand that the boss is taking personal responsibility for putting lives at risk based on their analysis. They tend to be respectful of their crucial, but limited, role in the broader risk management process. Computational fluid dynamics can't tell us if it's worth risking the life of an air crew to collect that flight test data. In that case there is no confusion about who is king, and over what questions the technical expert must "pass over in silence."
Monday, May 17, 2010
Wednesday, December 9, 2009
Verification, Validation, and Uncertainty Quantification
Notes on Chapter 8: Verification, Validation, and Uncertainty Quantification by George Em Karniadakis [1]. Karniadakis provides the motivation for the topic right off:
In time-dependent systems, uncertainty increases with time, hence rendering simulation results based on deterministic models erroneous. In engineering systems, uncertainties are present at the component, subsystem, and complete system levels; therefore, they are coupled and are governed by disparate spatial and temporal scales or correlations.
His definitions are based on those published by DSMO and subsequently adopted by AIAA and others.
Verification is the process of determining that a model implementation accurately represents the developers conceptual description of the model and the solution to the model. Hence, by verification we ensure that the algorithms have been implemented correctly and that the numerical solution approaches the exact solution of the particular mathematical model typically a partial differential equation (PDE). The exact solution is rarely known for real systems, so “fabricated” solutions for simpler systems are typically employed in the verification process. Validation, on the other hand, is the process of determining the degree to which a model is an accurate representation of the real world from the perspective of the intended uses of the model. Hence, validation determines how accurate are the results of a mathematical model when compared to the physical phenomenon simulated, so it involves comparison of simulation results with experimental data. In other words, verification asks “Are the equations solved correctly?” whereas validation asks “Are the right equations solved?” Or as stated in Roache (1998) [2], “verification deals with mathematics; validation deals with physics.”
He addresses the constant problem of validation succinctly:
Validation is not always feasible (e.g., in astronomy or in certain nanotechnology applications), and it is, in general, very costly because it requires data from many carefully conducted experiments.
Getting decision makers to pay for this experimentation or testing is especially problematic when they were initially sold on using modeling and simulation as a way to avoid testing.
After this the chapter goes into an unnecessary digresion on inductive reasoning. An unfortunate common thread that I’ve noticed in many of the V&V reports I’ve read is they seem to think Karl Popper had the last word on scientific induction! I think the V&V community would profit greatly by studying Jayne’s theoretically sound pragmatism. They would quickly recognize that the ’problems’ they perceive in scientific induction are little more than misunderstandings of probability theory as logic.
The chapter gets back on track with the discussion of types of error in simulations:
Uncertainty quantification in simulating physical systems is a much more complex subject; it includes the aforementioned numerical uncertainty, but often its main component is due to physical uncertainty. Numerical uncertainty includes in addition to spatial and temporal discretization errors, errors in solvers (e.g., incomplete iterations, loss of orthogonality), geometric discretization (e.g., linear segments), artificial boundary conditions (e.g., infinite domains), and others. Physical uncertainty includes errors due to imprecise or unknown material properties (e.g., viscosity, permeability, modulus of elasticity, etc.), boundary and initial conditions, random geometric roughness, equations of state, constitutive laws, statistical potentials, and others. Numerical uncertainty is very important and many scientific journals have established standard guidelines for how to document this type of uncertainty, especially in computational engineering (AIAA 1998 [3]).
The examples given for effects of ’uncertainty propagation’ are interesting. The first is a direct numerical simulation (DNS) of turbulent flow over a circular cylinder. In this resolved simulation, the high-wave numbers (smallest scales) are accurately captured, but there is disagreement at the low wave numbers (largest scales). This somewhat counter-intuitive result occurs because the small scales are insensitive to experimental uncertainties about boundary and initial conditions, but the large scales of motion are not.
The section on methods for dealing with modelling uncertain inputs is sparse on details. Passing mention is made of Monte Carlo and Quasi-Monte Carlo methods, sensitivity-based methods and Bayesian methods.
The section on ’Certification / Accreditation’ is interesting. Karniadakis recomends designing experiments for validation based on the specific use or application rather than based on a particular code. This point deserves some emphasis. It is an often voiced desire from decision makers to have a repository of validated codes that they can access to support their various and sundry efforts. This is an unrealistic desire. A code can not be validated as such, only a particular use of a code can be validated. In most decisions that engineering simulation supports, the use is novel (research and new product development), therefore the validated model will be developed concurrently with (in the case of product development) or as a result of (in the case of research) the broader effort in question.
The suggested hierarchical validation framework is similar to the ’test driven development’ methodologies in software engineering and the ’knowledge driven product development’ championed in the GAO’s reports on government acquisition efforts. Small component (unit) tests followed by system integration tests and then full complex system tests. When the details of ’model validation’ are understood, it is clear that rather than replacing testing, simulation truly serves to organize test designs and optimize test efforts.
The conclusions are explicit (emphasis mine):
The NSF SBES report (Oden et al. 2006 [4]) stresses the need for new developments in V&V and UQ in order to increase the reliability and utility of the simulation methods at a profound level in the future. A report on European computational science (ESF 2007 [5]) concludes that “without validation, computational data are not credible, and hence, are useless.” The aforementioned National Research Council report (2008) on integrated computational materials engineering (ICME) states that, “Sensitivity studies, understanding of real world uncertainties and experimental validation are key to gaining acceptance for and value from ICME tools that are less than 100 percent accurate.” A clear recommendation was reached by a recent study on Applied Mathematics by the U.S. Department of Energy (Brown 2008 [6]) to “significantly advance the theory and tools for quantifying the effects of uncertainty and numerical simulation error on predictions using complex models and when fitting complex models to observations.”
References
[2] Roache, P.J. 1998. Verification and validation in computational science and engineering. Albuquerque,: Hermosa Publishers.
[3] AIAA Guide for the Verification and Validation of Computational Fluid Dynamics Simulations, Reston, VA, AIAA. AIAA-G-077-1998.
[4] Oden, J.T., T. Belytschko, T.J.R. Hughes, C. Johnson, D. Keyes, A. Laub, L. Petzold, D. Srolovitz, and S. Yip. 2006. Revolutionizing engineering science through simulation: A report of the National Science blue ribbon panel on simulation-based engineering science. Arlington: National Science Foundation. Available online
[5] European Computational Science Forum of the European Science Foundation (ESF). 2007. The Forward Look Initiative. European computational science: The Lincei Initiative: From computers to scientific excellence. Information available online.
[6] Brown, D.L. (chair). 2008. Applied mathematics at the U.S. Department of Energy: Past, present and a view to the future. May, 2008.
Concepts of Model Verification and Validation has a glossary that defines most of the relevant terms.
Tuesday, November 17, 2009
Cut F-35 Flight Testing?
I would have just dismissed this article as more of the "same-ole-same-ole", but at the bottom they have a quote I just love:
But scrimping on flight testing isn’t a good idea, said Bill Lawrence of Aledo, a former Marine test pilot turned aviation legal consultant.
"They build the aircraft . . . and go fly it," he said. "Then you come back and fix the things you don’t like."
No amount of computer processing power, Lawrence said, can tell Lockheed and the military what they don’t know about the F-35.
Thank God for testers. This goes to the heart of the problems I have with the model validation (or lack there of) in climate change science.
From reading some of these posts, you might think I'm some sort of Luddite who doesn't like CFD or other high-fidelity simulations, but you'd be wrong. I love CFD, I've spent a big chunk of my (admittedly short) professional career developing or running or consuming CFD (and similar) analyses. I just didn't fall in love with the colourful fluid dynamics. I understand that bending metal (or buying expensive stuff) based on the results of modeling unconstrained by reality is the height of folly.
That folly is exacerbated by an 'old problem'. In another article on the F-35, I found this gem:
No one should be shocked by the continuing delays and cost increases, said Hans Weber, a prominent aerospace engineering executive and consultant. "The airplane is so ambitious, it was bound to have problems."
The real problem, Weber said, is an old one. Neither defense contractors nor military or civilian leaders in government will admit how difficult a program will be and, when problems do arise, how challenging and costly they will be to fix. "It's really hard," Weber said, "for anyone to be really honest."
Defense acquisition in a nutshell, Pentagon Wars anyone?
Defense procurement veterans said they fear that the Pentagon will be tempted to cut the flight-testing plan yet again to save money.
"You need to do the testing the engineers originally said needed to be done," Weber said. By cutting tests now, "you kick the can down the road," and someone else has to deal with problems that will inevitably arise later.
Classic.
Monday, October 19, 2009
JASSM Bootstrap Reliability II
empirical_rand() and public statements by various folks concerning the flight tests. As reported by Reuters, JASSM has had success in recent tests. To be quantitative, 15 out of 16 test flights were a success. Certainly good news for a program that's been bedevilled by flight test failures and Nunn-McCurdy breaches.
The previous analysis can be extended to include this new information. This time we'll use Python rather than Octave. There is not (that I know of) a built-in function like
empirical_rand() in any of Python's libraries, but it's relatively straightforward to use random integers to index into arrays and accomplish the same thing. def bootstrap(data, nboot):
"""Draw nboot samples from the data array randomly with replacement, ie
a bootstrap sample."""
bootsample = sp.random.random_integers
return(data[bootsample(0, len(data)-1, (len(data), nboot))])
Applying this function to our new data vectors to get bootstrap samples is easy.
# reliability data for July 2009, based on Reuters report
d09_jul_tests = 19
d09_jul = sp.ones(d09_jul_tests, dtype=float)
d09_jul[0:3] = 0.0
# generate the bootstrap samples:
d09_jul_boot = sp.sum(bootstrap(d09_jul, nboot), axis=0) / float(d09_jul_tests)
# find the number of unique reliabilities to set the number of bins for the histogram:
d09_jul_nbins = sp.unique(d09_jul_boot).size

So, is it a 0.9 missile or a 0.8 missile? We should probably be a little more modest in the inferences we wish to draw from such small samples. As reported by Bloomberg the Air Force publicly contemplated cancellation for a missile with reliability distributions like the blue histogram shown below (six of ten failed), while stating that if 13 of 16 were successful (the green histogram) that would be acceptable. The figure below shows these two reliability distributions along with the actual recent performance.

This seems to be a reasonably supported inference from these sample sizes, the fixes that resulted in the recent 15 out of 16 successes had a measurable effect when compared to the 4 out of 10 performance.
Not that binary outcomes for reliability are a great measure, they just happen to be easily scrape-able from press releases. The figure below illustrates the problem, there just isn't much info in each 1 or 0, so really cranking up the number of samples only slowly improves the power (reduces the area of overlap between the two distributions).
Wednesday, September 23, 2009
Five Points for Acquisition Reform
Concluding Observations on Achieving Lasting Reform:
I would like to offer a few thoughts about other factors that should be considered so that we make the most out of today's opportunity for meaningful change. First, I think it is useful to think of the processes that affect weapon system outcomes (requirements, funding, and acquisition) as being in a state of equilibrium. Poor outcomes--delays, cost growth, and reduced quantities--have been persistent for decades. If we think of these processes as merely "broken", then some targeted repairs should fix them. I think the challenge is greater than that. If we think of these processes as being in equilibrium, where their inefficiencies are implicitly accepted as the cost of doing business, then the challenge for getting better outcomes is greater. Seen in this light, it will take considerable and sustained effort to change the incentives and inertia that reinforce the status quo.
Second, while actions taken and proposed by DOD and Congress are constructive and will serve to improve acquisition outcomes, one has to ask the question why extraordinary actions are needed to force practices that should occur normally. The answer to this question will shed light on the cultural or environmental forces that operate against sound management practices. For reforms to work, they will have to address these forces as well. For example, there are a number of proposals to make cost estimates more rigorous and realistic, but do these address all of the reasons why estimates are not already realistic? Clearly, more independence, methodological rigor, and better information about risk areas like technology will make estimates more realistic. On the other hand, realism is compromised as the competition for funding encourages programs to appear affordable. Also, when program sponsors present a program as more than a weapon system, but rather as essential to new fighting concepts, pressures exist to accept
less than rigorous cost estimates. Reform must recognize and counteract these pressures as well.
Third, decisions on individual systems must reinforce good practices. Programs that have pursued risky and unexecutable acquisition strategies have succeeded in winning approval and funding. If reform is to succeed, then programs that present realistic strategies and resource estimates must succeed in winning approval and funding. Those programs that continue past practices of pushing unexecutable strategies must be denied funding before they begin. This will take the cooperative efforts of DOD and Congress.
Fourth, consideration should be given to setting some limits on what is a reasonable length of time for developing a system. For example, if a program has to complete development within 5 or 6 years, this could serve as a basis to constrain requirements and exotic programs. It would also serve to get capability in the hands of the warfighter sooner.
Fifth, the institutional resources we have must match the outcomes we desire. For example, if more work must be done to reduce technical risk before development start--milestone B--DOD needs to have the organizational, people, and financial resources to do so. Once a program is approved for development, program offices and testing organizations must have the workforce with the requisite skills to manage and oversee the effort. Contracting instruments must be used
that match the needs of the acquisition and protect the government's interests. Finally, DOD must be judicious and consistent in how it relies on contractors.
--Statement of Paul Francis, Managing Director:
Acquisition and Sourcing Management
I especially like the focus on the system and incentives that tend to lead to certain outcomes rather than a focus on particular techniques or methodologies. Mandating techniques for cost estimation or risk management is a waste of time, any mandated technique will be gamed because of the pressures and incentives inherent in the system.















