← All posts Next · Part 4 →

The compounding cost of deferred systems engineering

Four laps around the V, and what each lap pays for the reasoning the last one didn't record.

There's a reason "we'll add systems engineering later" sounds like a schedule decision rather than a compounding one. Picture the V as a single trip (down the left arm, along the bottom and up the right) and deferring the paperwork looks like a fixed debt. Some quarter in the future, someone spends a few weeks reconstructing links, and the books balance.

The picture is wrong, and the reason is the thing engineers know perfectly well from experience but rarely state in the diagram: the V is one lap, not the whole program.

Development is laps, not a pass

A B-sample. A C-sample. A regulation amended. A variant scoped for a new market. A field issue that needs a root cause. Each of those sends the team back to the top of the V and down again.

This is not a new observation. Boehm's spiral model put iteration at the centre in 1988, and Forsberg and Mooz themselves later extended the Vee to handle concurrent and evolving architectures. INCOSE's current definition of systems engineering spans a system's realization, use and retirement, not just its first build. Nobody in the field believes complex systems are built in one pass.

What's worth stating precisely is a narrower point about what each lap needs from the last one.

Artifacts survive a lap. Reasoning survives only if someone captured it.

The requirements document is still there next year. The CAD is still there. The test reports are still there. What isn't there, unless somebody deliberately wrote it down, is why a limit sits where it does, what assumption sits underneath it, and which alternative was considered and rejected when that limit was set.

Part 2 covered why that can't be recovered afterward: rationale is a co-product of design, built alongside the artifact, not extractable from it later. Put those two facts together and you get the compounding mechanism. Every lap asks a question that only the previous lap could have answered, and the answer expires.

Four V-shapes in sequence, each with a return arrow from the top of its right arm back to the top of the next one's left arm. Solid document icons carry forward between laps; small diamond decision markers fade out with each successive lap.

Four kinds of lap, four different bills

The laps a complex systems program runs are not all the same. They come in a handful of recurring kinds, and each kind asks the previous laps for something different. Four of them cover most of what I've seen.

The spec change

An input moves. A supplier changes a part, a test result comes back marginal, a customer tightens a number. Someone updates a requirement.

That requirement was the basis for decisions elsewhere: a firmware parameter, a placement choice, a line in the test plan. If none of those are linked to it, the change lands in one document and nowhere else.

What it costs: a meeting to figure out what the change touches, and a set of downstream assumptions that nobody revisits because nobody remembers they derived from that number. The team has introduced a silent inconsistency, and the clock on finding it starts running.

The regulation

A standard is amended, or a new market brings a new one, with a compliance date attached.

The compliance engineer needs to show which existing requirements the change touches, which components implement them, and which tests provide evidence. As Part 2 set out, this is the first role that needs the whole chain at once.

What it costs: weeks of reconstruction under someone else's deadline. The work is not only tedious but uncertain: the team ends up attesting to links it inferred rather than links it recorded. This is the audit scramble from Part 1, and it's the lap where the cost is external and non-negotiable.

The variant

Someone wants a version with more range, more power, a different target audience, a different price point. A newer engineer, reasonably, proposes an option the original team looked at and turned down.

Nobody can say why it was turned down. The analysis, the margins that made the chosen option viable, and the trade-off that came with it are gone, often along with some of the people who were in the room when the original decision was made.

What it costs: the trade study gets run again from scratch. Worse, it may be run on different assumptions and reach a different answer, and the team now has two architectures whose relationship nobody can articulate.

The field issue

Something in service behaves worse than expected, and the root cause crosses subsystem boundaries. It usually traces back, at least in part, to an assumption set on an earlier lap that was never revisited when something around it changed.

What it costs: the investigation has to reconstruct several laps of reasoning simultaneously. Each reconstruction is harder than it would have been at the time, because the engineers involved have moved on, and the artifacts show what was decided without showing why.

Why the curve bends upward

Three effects stack, and it's worth separating them because they have different fixes.

Changes propagate. Clarkson, Simons and Eckert modelled how a change to one component spreads through the connections between components, and showed that propagation can be predicted when those connections are known [1]. Eckert, Clarkson and Zanker's case study of change in an aerospace product describes the underlying mechanism: change becomes problematic when it propagates to other systems because individual parameters exceed their tolerance margins [2]. Knowing the connections is precisely the capability SE produces. Without it, every change is a search rather than a traversal.

The decision context decays. People leave, memory fades, and the assumptions behind a number become invisible once the number is in a document. Fricke and colleagues, surveying how firms cope with change, treat understanding the causes and handling of change as a systems engineering problem rather than an administrative one [3].

Each lap builds on the last lap's gaps. A field issue on the fourth lap can depend on an assumption set on the first. If the first lap didn't record it, the fourth doesn't just pay its own cost, it pays the first lap's cost too, at fourth-lap prices.

That third effect is the compounding one, and it's the reason the total is not the sum of four retrofits. Reconstruction gets more expensive the further you are from the event, and the laps queue up.

About the multipliers

Part 1 cited the NASA study putting the cost of fixing a requirements error at 1 unit in the requirements phase, rising to 3–8 in design, 7–16 in build, 21–78 in integration and test, and 29 to 1,615 in operations [4]. Those numbers are the most-quoted in the field, and they deserve the caveats that usually get dropped.

The study modelled a hardware/software system with characteristics like a large complex spacecraft, a military aircraft, or a small communications satellite [4]. The wide ranges come from three different estimation methods that disagree with each other, which is honest of the authors and should make anyone quoting a single dramatic figure uncomfortable.

Line chart on a log scale of the relative cost to fix a requirements error by the phase in which it is found: requirements, design, build, test, operations. Three estimation methods start together at 1× and diverge. Method 1, based on spacecraft modifications, reaches 29× in operations; Method 2, a major aircraft program, 157–186×; Method 3, a communications satellite, 1,615×. A shaded band shows the composite range across all methods. Source: Stecklein et al. (2004), NASA JSC.

The older software-era version of this curve, usually attributed to Boehm, has been challenged on its evidence base. Laurent Bossavit's The Leprechauns of Software Engineering traces that curve and several other industry "ground truths" back to their original sources and finds the chain weaker than its reputation [5].

So don't take the multipliers as a forecast for your program. The direction is what survives scrutiny across very different sources, and the direction is all the argument needs: errors in early decisions get more expensive the further they travel, and unrecorded reasoning turns every later lap into a search.

What would have changed this

None of the four laps above requires a heavyweight process to make cheap. A rejected alternative needs one record: the question, the options, the criteria, and what won. Perhaps a paragraph. That record is what the variant lap needs and can't otherwise get.

A requirement needs an ID and a link to the assumptions derived from it. That's a minute of work at the time it was set, and the difference between a traversal and an investigation when the field issue arrives.

The asymmetry is the whole argument. Capture costs minutes at the moment of decision and is impossible at any later point. That's an unusual cost profile, and it's why the usual intuitions about deferring work don't apply here.

Which raises the question Part 4 takes up. If retrofitting is so much more expensive than capturing, why can't a better process fix it after the fact? The answer has less to do with process discipline than with the shape of the thing you're trying to rebuild.


Srivardhan Chandrapati is the founder of SysenAI. He spent about twelve years in systems engineering and functional safety across advanced mobility and energy, at Cummins, Rivian, EnerDel, Fisker and BorgWarner, and now builds from Bengaluru.

References

  1. Clarkson, P. J., Simons, C. & Eckert, C. (2004). "Predicting Change Propagation in Complex Design." Journal of Mechanical Design, 126(5), pp. 788–797. DOI: 10.1115/1.1765117
  2. Eckert, C., Clarkson, P. J. & Zanker, W. (2004). "Change and customisation in complex engineering domains." Research in Engineering Design, 15(1), pp. 1–21. DOI: 10.1007/s00163-003-0031-7
  3. Fricke, E., Gebhard, B., Negele, H. & Igenbergs, E. (2000). "Coping with changes: Causes, findings, and strategies." Systems Engineering, 3(4), pp. 169–179.
  4. Stecklein, J. M., Dabney, J., Dick, B., Haskins, B., Lovell, R. & Moroney, G. (2004). Error Cost Escalation Through the Project Life Cycle. NASA Johnson Space Center; presented at the INCOSE International Symposium 2004. NTRS 20100036670 (PDF)
  5. Bossavit, L. The Leprechauns of Software Engineering: How folklore turns into fact and what to do about it. Leanpub (first published 2012). leanpub.com/leprechauns

See the memory working, not just described.

Walk through one project end to end, from the decision to the graph that holds it.