JUNE 2026 I Volume 47, Issue 2
JUNE 2026
Volume 47 I Issue 2
IN THIS JOURNAL:
- Issue at a Glance
- Chairman’s Message
Technical Articles - 2026 AI in T&E Forum
- Decision Assurance for AI-Enabled Mission Systems: From Test Evidence to Operational Authority
- Accelerating Test & Evaluation with AI Across the Systems Engineering Lifecycle
- SECC: An AI-Powered Assurance Agent for Complex Government System Integration
- Human Oversight for AI-Generated Test Artifacts
- Toward an Integrated T&E Framework for AI-enabled Systems: A Conceptual Model
Technical Articles
- Retrieval-Augmented Generation for Departmental Test & Evaluation
- Avoiding Vendor Lock-In in AI Procurement
- Developing Winning Proposals through the Lens of Test and Evaluation
- Defining T&E as a Discipline
News
- Association News
- Chapter News
- Corporate Member News
![]()
Toward an Integrated T&E Framework for AI-enabled Systems: A Conceptual Model

Karen O’Brien
Technical Fellow at Modern Technology Solutions, Inc., M.S. in Predictive Analytics from Northwestern University
Abstract
The classic DoW T&E paradigm built around Operational Effectiveness, Suitability, Survivability, and Safety, benefits from decades of refinement, specialized disciplines, and rigorous analysis. T&E of AI-enabled systems is far newer, and the rapid growth of testing concepts has produced an overemphasis on performance, risking a loss of focus on broader effectiveness measures. Stepping back from confusion-matrix metrics reveals a much wider landscape of evaluation issues rooted in policy and operational need, including safety, transparency, and robustness, that are at risk of being overlooked. Borrowing from the classic “integrated survivability onion” that was used to make T&E within Systems of Systems tractable, we propose a nested set of evaluation questions for AI-enabled systems that spans traditional T&E concerns and those unique to military AI as implied by DoW Responsible AI Guidance. The model preserves analytical rigor while highlighting opportunities to apply test science across the full hierarchy.
Keywords: Artificial Intelligence, AI-Enabled Systems, System of Systems, Nested Hierarchical Model, Ethical AI
Introduction
As the test and evaluation (T&E) community stands on the cusp of our latest challenge, the T&E of AI-enabled systems, and as we consider how to grapple with this new surge of complexity and non-determinism, it is worth remembering that we have been here before. We have faced transformative shifts in system behavior, scale, and interdependence, and through analytical discipline and test science, we have found ways to reason about them. With that history in mind, we asked what lessons could be drawn from those earlier moments in T&E and how they might help us approach today’s challenge.
The result is a conceptual model for T&E intended to support conversations between stakeholders and evaluators, elicit meaningful test objectives, and help focus the scope of test on the issues that matter most, issues that are far beyond system performance. Rather than prescribing a rigid methodology, the model is meant to provide a shared structure for thinking through the nuances of AI-enabled systems. Through this paper, we aim to spark a conversation about how best to integrate diverse stakeholder perspectives into a unified T&E program that embraces the inherent complexity of AI/ML systems and supports their responsible and effective deployment.
Motivation and Historical Context
In the late 1990s, the U.S. Army began a major transformation effort informed by lessons from AirLand Battle doctrine, the post–Cold War security environment, and operational outcomes from the Gulf War (Benson 2012). Within this emerging system of systems construct, the survivability of a given ground asset, previously determined primarily by physical armor, became increasingly dependent on network enabled situational awareness and the ability to avoid engagement through interconnected sensors and weapons platforms (Crane, Lynch, and Reilly 2018).
This increase in complexity created significant challenges across the acquisition community. While engineers and program managers struggled to account for interdependencies between systems, the T&E community faced a genuine methodological crisis. Evaluators were now responsible for assessing systems whose effectiveness, suitability, and survivability were now partially derived from the broader system of interoperating systems. Traditional test designs, which centered on the performance of the individual platform, no longer captured the full picture. This shift also raised an essential operator focused question: how could a tank crew trust in their equipment when their survivability depended on factors other than their armor?
In 2004, Gary L. Guzie of the Army Research Laboratory introduced a conceptual model to address part of the methodology crisis. The Integrated System of Systems “Survivability Onion” (Figure 1) provided both a structure and a methodology for addressing measurement challenges within a system of systems. It offered a way to assess the survivability of a single ground asset even when a significant portion of that survivability was attributed to other interconnected systems. These systems contributed by preventing an adversary from detecting, acquiring, or engaging the blue asset.

Figure 1: The Survivability Onion
The model also helped clarify how crews could develop trust in these distributed survivability mechanisms. Reading the model from the center outward clarified how armor, countermeasures, and networked sensors collectively reduced vulnerability and gave crews a clearer understanding of how the system of systems’ operations protected them. By understanding that part of the survivability burden was carried by networked sensors and interoperating weapons platforms, operators gained a clearer basis for confidence in their equipment because they prevented the system from being detected, acquired, and fired upon. This improved understanding, in turn, supported more effective employment of the system.
The Current Challenge
Today, AI and ML-enabled systems present challenges of even greater complexity and raise parallel questions about how to measure performance and how to establish trust. Although these technologies introduce unprecedented capabilities, the underlying T&E difficulties are not entirely new. Just as Guzie’s Survivability Onion helped the community reason about distributed survivability within a system of systems, it also provides an actionable precedent for approaching the multifaceted evaluation needs of AI-enabled systems. Building on this foundation, we adapt the logic of that nested hierarchical model to create a Conceptual Model for T&E of AI-enabled systems (Figure 2) that addresses contemporary complexity in two essential ways.

Figure 2: A nested hierarchical model for T&E of AI-enabled systems, inspired by Guzi’s Survivability Onion.
First, it provides a structured framework for understanding the interrelationships among these diverse evaluation perspectives, enabling productive collaboration and debate about test requirements. Second, it highlights potential gaps in test planning, revealing deficiencies in system design, and uncovering opportunities for reusing data to address higher-order evaluation needs. Thus, this model not only informs T&E efforts but also influences system design itself by facilitating a conversation among stakeholders. While Guzie sought a mathematical integration for survivability, our vision for our model is to be primarily used in a workshop setting to elicit stakeholder needs and concerns as a basis for a test program.
Survivability in a system of systems was more than calculating a probability of kill given a hit, it was an understanding about whether the total system could be trusted to provide survivability. Similarly, AI-enabled systems carry an inherent trustworthiness burden.
Calibrated Trust
In AI-enabled systems, calibrated trust is the top-level metric of interest, superseding system performance. It can be understood as “Goldilocks trust,” meaning trust that is neither excessive nor insufficient, but appropriately matched to the system’s trustworthiness in the operational context at a given moment. Trust is also inherently dynamic (Hoffman 2017). In military environments where conditions evolve rapidly, both over trust and under trust can have significant operational consequences. An AI/ML system must therefore provide users with timely and sufficient information to judge whether its outputs should be accepted, questioned, or discounted in real time.
With this focus on trust, we begin at the center of the model and work outward, identifying the potential contributors to calibrated trust that should be incorporated into the test program. These contributors become evaluation questions that guide test design. The model also helps maintain a broader perspective on testing needs, enabling more efficient allocation of test resources and ensuring that the resulting program supports an accurate assessment of calibrated trust.
System Performance
Starting at the center of the model, we ask a foundational question: “Does the system successfully do what we want it to do?” System performance occupies this central position because it addresses whether the AI/ML system accomplishes its intended function. Traditional AI T&E evaluates models in controlled development environments using metrics such as accuracy, precision, recall, F1 score, and error measures like mean average error (MAE). These measures are essential, but reliance on them alone can obscure higher order considerations, including operational effectiveness, suitability, survivability, and the development of calibrated trust. They are necessary, but not sufficient. Such limitation is well documented in the machine learning systems literature, where system-level failures and “hidden technical debt” often arise even when model-level performance metrics appear strong (Sculley et al. 2015).
There are two important nuances. First is that the list of desired functions and level of performance should be treated as tentative. As the discussion moves beyond the inner layer of the model, additional perspectives will emerge, requiring updates to the initial assumptions. Second is that the metrics that are applied to the performance measurements can be used as response variables for tests associated with other regions of the model. System performance forms the core of the hierarchy, but a complete evaluation requires situating these metrics within a broader framework that reflects real world complexity, mission objectives, and the need to support calibrated trust.
AI Safety
Next, we ask, “Does the system successfully not do the things we don’t want it to do?” Framing the question in this somewhat awkward way draws attention to a persistent challenge: negative requirements are seldom treated with the same deliberation as positive ones in test design. Engineering best practice typically expresses requirements in positive terms, specifying what a system must do. In hardware systems, we proactively deal with hardware failures using engineering principles learned over millennia. For instance, it goes without saying that we do not want a bridge to collapse, yet every step of bridge design has methodology to prevent that negative outcome. AI-enabled systems have no such legacy of lessons-learned and engineering procedures, so negative requirements should be treated explicitly in test and evaluation. Such challenges align with widely recognized AI safety concerns, in which unintended behaviors, specification errors, and reward misalignment can produce harmful outcomes even when systems appear to be functioning as designed (Amodei et al. 2016). System safety is co-located with system performance at the center of the model because it is as foundational as system performance. Evaluating performance against negative requirements often calls for distinct test designs. These may emphasize stressing conditions, edge cases, or larger sample sizes to support robust analysis. They may also require different metrics, particularly error-based measures. Such attention ensures that the evaluation addresses not only desired system behaviors but also the prevention of unintended or unsafe outcomes.
Transparency and Explainable AI
A recurring topic in discussions of the trustworthiness of AI enabled systems is the need to understand how the system is making decisions and whether the underlying rationale is acceptable from a user perspective. Accordingly, the next question extends from the previous two: “Does the system do what it should do and not do what it shouldn’t do for the right reasons?” An early and essential step in evaluating whether a system’s decision-making process is satisfactory is achieving sufficient transparency into the system itself. Recent work has emphasized structured documentation approaches, such as model cards, to systematically communicate model behavior, intended use, and limitations to stakeholders (Mitchell et al. 2019). The information gained at this stage is not an end in itself; it provides essential diagnostic inputs to higher layers of the hierarchy, informing assessments of robustness, trustworthiness, and ultimately calibrated trust. Gaining satisfactory transparency can be challenging if the system was designed with transparency as an afterthought. Many options exist, such as whitebox testing, and the right choices are driven by user needs. However, transparency methods have strengths and limitations. For instance, saliency maps are a popular choice (Skliarov et al. 2025) for understanding computer vision models but may not be robust to mission-critical shifts in model function (Zhang et al. 2022). Transparency testing capabilities are critical inputs to Human Systems Integration studies (Vössing et al. 2022). Importantly, the insights gained from transparency form the basis for evaluating how the system performs in an operational setting, which may be beyond ideal or designed-for conditions.
Operational Robustness
Moving one layer outward in the model brings us to the operational environment, which encompasses the full range of conditions under which the AI-enabled system must function. The next question asks, “If the system can do what we desire it to do, at a desired level of performance and for the right reasons, can it do so across a range of operational conditions?” Addressing this question introduces a new dimension of test design. It may require substantial additional data to ensure adequate coverage, ideally across the full operational environment and at a minimum across the defined requirements envelope. This problem is closely related to the widely studied issue of distribution shift, where models that perform well under development conditions can exhibit degraded or poorly calibrated behavior when exposed to new data environments (Ovadia et al. 2019). Often thought of as “Natural Robustness,” testing toolkits (Hu et al 2024), and the associated hardening methods, focus on variables that vary in nature (atmospheric conditions, temperature) or variations from equipment (focal length, lens flare, signal noise) but not necessarily the variation that comes from changes in the operational context: operational tempo, rules of engagement, civil constraints, and mission objectives. The resulting test requirements should capture all relevant factors and conditions found in the operational environment.
Adversarial Robustness
Moving out a step further, we extend our testing perspective to include the adversarial lens and ask, “Can the system sustain its performance across a range operational conditions, against a motivated and capable adversary?” Adversarial robustness encompasses a wide attack surface (Shayea et al. 2025), much of which requires evaluators to look beyond the AI-enabled system itself and examine interoperating systems and the full lifecycle of data.
This echoes the earlier system of systems challenge of evaluating survivability in the context of a system of interoperating systems. Here, too, we must consider whether the system can alert the user to issues emerging from an interoperating system, detect possible data poisoning, withstand non-local adversarial events, or rely on an application that provides anomaly detection across the broader system of systems.
The term “adversarial” is often interpreted as referring solely to offensive actions taken to defeat a blue system. However, a motivated and rational adversary may also employ defensive, or “counter-AI”, measures such as evasion, deception, decoys, camouflage, or obscurants. Systems must be evaluated against these possibilities, even though relevant data may be limited.
It is also important to recognize that not every adversarial influence arises from a hostile actor. A well-intentioned but undertrained user may inadvertently perform actions that compromise system performance or safety. The system must therefore be robust to such user-generated disruptions as well.
Ethical AI
Moving further outward in the hierarchy brings us into the socio-technical domain, where the focus shifts from performance to potential harms (Ravi 2025). This shift reflects a broader understanding that AI systems operate within complex sociotechnical contexts, where failures often arise from mismatches between abstract system models and real-world social dynamics (Selbst et al. 2019). At this layer we ask, “Does the system accomplish its intended functions in a manner that does not unintentionally harm or mislead?”
The test data generated in the inner layers of the model are input to legal and ethical analysis. It is conceivable that future requirements documents will explicitly reference human or civil rights, but for now these considerations operate as implied requirements. As AI-enabled systems are increasingly scrutinized in legal contexts, the resulting case law will shape expectations for both system design and testing. A central question for evaluators is therefore: What legal tests should be incorporated into the evaluation process?
Systemic biases in AI/ML algorithms or in the underlying data can lead to mission-impacting failure modes, particularly in operational environments where fairness, equity, or proportionality carry legal or strategic significance. In addition, evaluators must consider not only whether harms occur, but how the system behaves when failures arise. When the system does fail, does it do so gracefully, and does it limit the potential for unintended harm? These considerations form a critical part of assessing the overall ethical suitability of the system, far beyond its core performance.
AI and Law
As we move to the outermost layer of the model, the focus shifts to a hypothetical yet plausible future in which AI-enabled systems support missions of global significance and a mishap occurs. In such a scenario, we must ask, “Does the system accomplish its intended function in a manner consistent with law, treaty obligations, and applicable rules of engagement?” It is reasonable to expect that, at some point, an incident will require a forensic technical and legal investigation. The question is whether we will have the necessary capabilities, documentation, and data artifacts in place to support that investigation.
Current Department of War ethical AI guidance emphasizes “traceability” (U.S. Department of Defense 2022), but this alone does not ensure the technical or legal records required for forensic reconstruction, auditability, or demonstrations of compliance. Designers and evaluators must therefore determine what data should be retained, under which conditions, and in what form. Identifying the essential elements of the AI equivalent of an airline black box recorder is a design decision with direct impacts to T&E. Any proposed logging or traceability mechanism should be exercised whenever possible to confirm that it can support future investigative and legal needs.
Implementing the Conceptual Model
The conceptual model functions primarily as an organizing framework for eliciting and refining test and evaluation requirements. The structure of the model also aligns with emerging risk-based AI governance approaches, which advocate continuous evaluation across the lifecycle rather than isolated test events (NIST 2023). Its value lies not in prescribing specific metrics or test methods but in helping stakeholders reason systematically across the full set of considerations that contribute to calibrated trust. As such, its implementation begins with facilitated engagements, such as interviews, workshops, or tabletop exercises, through which stakeholders articulate their expectations and concerns layer by layer.
This structured elicitation is intended to reveal gaps in formal requirements, inconsistencies in assumptions about system behavior, or unrecognized dependencies on operational context, all with the aim of ensuring test coverage, data, and artifacts inform calibrated trust. Because these observations frequently reshape initial test concepts, implementation should be approached as an iterative process rather than a single elicitation event (Figure 3). Returning to stakeholders to validate the synthesized requirements and emerging test designs is essential, both to maintain alignment and to ensure that the evaluation remains grounded in operational need.

The model works within the expectations of the Adaptive Acquisition Framework (AAF), particularly the Software Acquisition and Business Systems Pathways (DoDI 5000.02). Both emphasize rapid and continuous delivery of capability, frequent engagement with users, and the integration of developmental testing, operational demonstration, evaluation, and design updates before incremental fielding. Small, cross functional teams that include users, testers, developers, system engineers, cybersecurity specialists, and acquisition professionals are expected to work together in short cycles. The conceptual model reinforces these practices by providing a shared structure that enables diverse testing perspectives to be translated into testable questions and by supporting continuous refinement of requirements as new evidence is generated. To obtain best value from the model, these workshops should be conducted early in the pathway, regardless of the chosen pathway.
The DoD Responsible AI Toolkit (CDAO 2026) integrates into the model. The toolkit supplies checklists, methods, and governance artifacts that help teams express responsible AI concerns in testable terms and ensure that evaluation activities reflect Department of War expectations for traceability, reliability, mitigated bias, and lawful operation.
Legal expertise has an important role in the conceptual model. The outer layers of the model involve issues of compliance, authority, and potential liability which are important inputs to calibrated trust. Early involvement of legal advisors helps evaluators identify the legal tests that should be incorporated into the T&E program and ensures that design decisions, data retention strategies, and operational demonstrations remain consistent with applicable law, policy, and rules of engagement. Their participation helps teams anticipate documentation needs and identifies requirements for the creation of audit trails and traceability mechanisms. It is important to note that such traceability requirements may become a system design requirement.
The model can also be applied selectively when the evaluation problem is narrower, such as assessing cybersecurity. Even in narrow use cases, the hierarchical structure prevents critical issues from being overlooked by explicitly tying specialized concerns back to underlying performance assumptions and forward to missionlevel consequences. In this way, the model supports a disciplined expansion of focus without diluting analytical rigor.
Limitations of the Conceptual Model
As with any conceptual framework, the model trades specificity for generalizability. Its purpose is to establish (and sustain) a broader context for the test and evaluation of AI-enabled systems, not to prescribe detailed performance metrics or domain-specific test procedures. Users should adapt the model, placing additional emphasis on certain layers or omitting others when appropriate. Regardless of such modifications, we recommend developing a clear definition of calibrated trust for the specific context and ensuring that each layer of the model is examined in relation to that definition.
Feedback Sought
The earlier challenge of finding a path to evaluate survivability of a system in the context of a system of interoperating systems was ultimately addressed through robust debate and collaboration across testing communities. We invite the same spirit of collaboration today and hope to refine the model through feedback and lessons-learned from applying the model in different contexts. We welcome the opportunity to discuss and incorporate your findings. Please reach out to the corresponding author.
References
Amodei, Dario, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané. “Concrete Problems in AI Safety.” arXiv preprint arXiv:1606.06565, 2016. https://doi.org/10.48550/arXiv.1606.06565.
Benson, Bill. “Unified Land Operations: The Evolution of Army Doctrine for Success in the 21st Century.” Military Review (March–April 2012).
Crane, Conrad C., Michael E. Lynch, and Shane P. Reilly. A History of the Army’s Future: 1990–2018, v. 2.0. U.S. Army Heritage and Education Center, 2018. https://ahec.armywarcollege.edu/documents/History-of-the-Future.pdf.
Guzie, G. L. Integrated Survivability Assessment. ARL-TR-3186. Army Research Laboratory, April 2004.
Hu, Brian, Brandon Richard Webster, Paul Tunison, Emily Veenhuis, Bharadwaj Ravichandran, Alexander Lynch, Stephen Crowell, Alessandro Genova, Vicente Bolea, Sebastien Jourdain, and Austin Whitesell. “NRTK: An Open-Source Natural Robustness Toolkit for the Evaluation of Computer Vision Models.” In Assurance and Security for AI-enabled Systems, Proc. SPIE 13054, 130540C (7 June 2024). https://doi.org/10.1117/12.3012769.
Mitchell, Margaret, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. “Model Cards for Model Reporting.” In Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT* ’19), 220–229. New York: ACM, 2019. https://doi.org/10.1145/3287560.3287596.
NIST. Artificial Intelligence Risk Management Framework (AI RMF 1.0). National Institute of Standards and Technology, January 2023. https://doi.org/10.6028/NIST.AI.100-1.
Ovadia, Yaniv, Emily Fertig, Jie Ren, Zachary Nado, David Sculley, Dan Zhang, and others. “Can You Trust Your Model’s Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift.” In Advances in Neural Information Processing Systems (NeurIPS 2019), vol. 32. 2019.Ravi, Kalluri. Socio Technical System Challenges in the Era of Artificial Intelligence: A Comprehensive Analysis. International Journal of Business & Management Studies 6, no. 9 (September 2025). https://doi.org/10.56734/ijbms.v6n9a8.
Sculley, D., Gary Holt, Daniel Golovin, Eugene Davydov, Todd Phillips, Dietmar Ebner, Vinay Chaudhary, and others. “Hidden Technical Debt in Machine Learning Systems.” In Advances in Neural Information Processing Systems (NeurIPS 2015), vol. 28. 2015.
Selbst, Andrew D., Danah Boyd, Sorelle A. Friedler, Suresh Venkatasubramanian, and Janet Vertesi. “Fairness and Abstraction in Sociotechnical Systems.” In Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT* ’19), 59–68. New York: ACM, 2019. https://doi.org/10.1145/3287560.3287598.
Shayea, Ghadeer Ghazi, Mohd Hazli Mohammed Zabil, Mustafa Abdulfattah Habeeb, Yahya Layth Khaleel, and A. S. Albahri. “Strategies for Protection Against Adversarial Attacks in AI Models: An In-Depth Review.” Journal of Intelligent Systems 34, no. 1 (2025): 20240277. https://doi.org/10.1515/jisys-2024-0277.
Skliarov, M., R. E. Shawi, C. Dhaoui, et al. “A Comparative Evaluation of Explainability Techniques for Image Data.” Scientific Reports 15 (2025): 41898. https://doi.org/10.1038/s41598-025-25839-y.
U.S. Department of Defense. Responsible Artificial Intelligence Strategy and Implementation Pathway. June 2022. https://media.defense.gov/2024/Oct/26/2003571790/-1/-1/0/2024-06-RAI-STRATEGY-IMPLEMENTATION-PATHWAY.PDF.
Department of War. DOW Instruction 5000.02: Operation of the Adaptive Acquisition Framework. Office of the Under Secretary of War for Acquisition and Sustainment. January 23, 2020. Change 2 effective April 8, 2026.
U.S. Department of War. “RAI Toolkit. Reliable AI Toolkit, Accessed May 29, 2026. https://rai.acqbot.com/.
Vössing, M., N. Kühl, M. Lind, et al. “Designing Transparency for Effective Human-AI Collaboration.” Information Systems Frontiers 24 (2022): 877–895. https://doi.org/10.1007/s10796-022-10284-3.
Zhang, J., H. Chao, G. Dasegowda, G. Wang, M. K. Kalra, and P. Yan. “Overlooked Trustworthiness of Saliency Maps.” In Medical Image Computing and Computer Assisted Intervention – MICCAI 2022, edited by L. Wang, Q. Dou, P. T. Fletcher, S. Speidel, and S. Li, Lecture Notes in Computer Science, vol. 13433. Cham: Springer, 2022. https://doi.org/10.1007/978-3-031-16437-8_43.
Author Biographies
Karen O’Brien is a Technical Fellow at Modern Technology Solutions, Inc. In this capacity, she leverages her 20-year Army civilian career as a scientist, evaluator, ORSA, and analytics leader to aid DoW agencies in implementing AI/ML and advanced analytics solutions, including the T&E and V&V methods to make these technologies a success. Her analytics career ranged “from ballistics to logistics,” primarily within Army Test and Evaluation Command and through supporting roles at the Army Research Laboratory. Significant roles include former Deputy Chief Analytics Officer at Army Materiel Command and former Chief Evaluator at ATEC for reliability growth, data science, and AI. She publishes innovative research on T&E for AI/ML solutions and recently completed a study for the National Academy of Sciences, Engineering, and Medicine. She uses her M.S. in Predictive Analytics from Northwestern University to help her DoW clients tackle the toughest analytics challenges in support of the nation’s Warfighters.
Dewey Classification: L 681 12

