JUNE 2026 I Volume 47, Issue 2
JUNE 2026
Volume 47 I Issue 2
IN THIS JOURNAL:
- Issue at a Glance
- Chairman’s Message
Technical Articles - 2026 AI in T&E Forum
- Decision Assurance for AI-Enabled Mission Systems: From Test Evidence to Operational Authority
- Accelerating Test & Evaluation with AI Across the Systems Engineering Lifecycle
- SECC: An AI-Powered Assurance Agent for Complex Government System Integration
- Human Oversight for AI-Generated Test Artifacts
- Toward an Integrated T&E Framework for AI-enabled Systems: A Conceptual Model
Technical Articles
- Retrieval-Augmented Generation for Departmental Test & Evaluation
- Avoiding Vendor Lock-In in AI Procurement
- Developing Winning Proposals through the Lens of Test and Evaluation
- Defining T&E as a Discipline
News
- Association News
- Chapter News
- Corporate Member News
![]()
Accelerating Test and Evaluation with AI Across the Systems Engineering Lifecycle

Muhammad F. Islam
MITRE Corporation, McLean, VA

Tomi Esho
MITRE Corporation, McLean, VA

Jyotirmay Gadewadikar
MITRE Corporation, McLean, VA
Abstract
Modern systems engineering efforts often have to deliver detailed test and validation outcomes under strict timelines with limited resources. This can result in insufficient testing efforts, release delays, production issues or even critical failures after launch. This research proposes an AI-enabled, human-in-the-loop framework that automates requirements-based test case generation while keeping test engineers in the decision loop. At first, a set of requirements is provided to a large language model (LLM)-based AI framework and to test engineering subject matter experts for draft test case generation, each of whom generates test cases independently as defined by a shared evaluation rubric. The resulting test cases are then evaluated blindly using a common set of criteria by expert evaluators under consistent assessment conditions. The test cases are compared for traceability to requirements, reproducibility, objectivity of expected results, precision of test data, and robustness under boundary and non-standard conditions. This framework enables a comprehensive assessment of how AI-generated tests perform relative to human-developed tests. The proposed framework aims to enable test engineers performing rapid test and validation activities, while also providing empirical evidence on where it is effective and where it can be improved in this exploratory study. This research seeks to reduce manual testing efforts, rework in test development while enabling rapid test delivery for mission-critical systems in defense, aerospace, and other high-assurance domains.
Introduction
Modern systems engineering efforts frequently operate under tight timelines and limited resources. This increases the risk of inadequate testing leading to potential system failures (Li & Fang, 2025; Zhang et al., 2026). This research investigates whether AI-assisted test generation can improve testing efficiency while maintaining the evaluation quality across the systems engineering lifecycle (Esho et al., 2025).
The proposed approach includes the following characteristics and evaluation goals:
- AI systems and human engineers generate test cases independently
- Utilizes blind evaluation with common quality criteria
- Compares AI-generated and human-written test cases
- Evaluates traceability, reproducibility, objectivity, precision and robustness
- Explores opportunities to reduce engineering effort and accelerate test delivery
- Supports mission assurance for high-assurance systems and operational environments
Methodology:
AI-generated (AI4SE) and handwritten test cases are developed using same prompts. 10 AI-generated and 10 handwritten test cases are provided to expert evaluators randomly without identifying which test cases are AI-generated or written by human.Handwritten test case development requires approximately two hours (~120 minutes) for the 10 test cases. In contrast, AI-generated test cases require approximately five seconds for initial generation and approximately 10 minutes for quality validation. This equals roughly 10 minutes of total effort for the 10 test cases. This came out to be a 92% reduction in effort, corresponding to approximately twelve-times faster test development.

Figure 1: AI-generated (AI4SE) Test Case Workflow
After the test cases are generated, expert evaluators assessed each test case using a 1 to 5 Likert scale across five quality dimensions (Joshi et al., 2015; Tran et al., 2025; White & Krinke, 2022):
- Traceability:Does this test case map directly to a single, unique requirement ID in the Traceability Matrix? (White & Krinke, 2022)
- Reproducibility:Are the pre-conditions and environment setup described well enough for an independent third party to replicate the test?
- Objectivity:Are the expected results based on measurable, quantitative data rather than subjective observations?
- Precision:Does the test procedure use specific test data (exact inputs) rather than generalized instructions?
- Robustness:Does the test challenge boundary values or off-nominal conditions instead of only the “happy path”?
Likert scale responses included:
- Strongly Disagree
- Disagree
- Neutral
- Agree
- Strongly Agree
The expert evaluation process focuses on determining whether AI-generated test cases can achieve comparable quality as the handwritten test cases while substantially reducing time and manual engineering efforts (Isley et al., 2026).
Results
The evaluation results demonstrate comparable performance between AI-generated and handwritten test cases across all assessment categories. AI4SE-generated test cases achieved slightly higher overall evaluation scores while substantially minimizing the manual development efforts.
Table 1: Comparative Assessment Results by Evaluation Criterion
| Category (Mean ± Standard Deviation) | AI4SE-Generated | Handwritten |
| Traceability | 5.00 ± 0.00 | 5.00 ± 0.00 |
| Reproducibility | 4.00 ± 0.78 | 3.93 ± 0.69 |
| Objectivity | 4.40 ± 0.67 | 4.28 ± 0.75 |
| Precision | 3.38 ± 0.90 | 3.18 ± 0.93 |
| Robustness | 2.45 ± 0.82 | 2.45 ± 0.85 |
| Overall | 3.85 ± 0.38 | 3.77 ± 0.47 |
Additional statistical analyses are conducted to compare AI-generated and handwritten test cases under consistent evaluation conditions.
Table 2: Statistical Analysis Results
| Statistical Measure | Result | Interpretation |
| Wilcoxon p-value (Fagerland & Sandvik, 2009) | 0.25 | No statistically reliable evaluator-consistent differences found |
| Mann-Whitney p-value (Divine et al., 2013; Fagerland & Sandvik, 2009) | 0.3093 | No reliable evidence that score distributions are statistically different |
| Cliff’s Delta (Macbeth et al., 2011) | 0.131 | Small effect favoring that AI-generated test cases score higher |
Key findings from the results analysis are summarized below:
- AI-generated and handwritten test cases achieved comparable evaluation quality
- No statistically significant differences are detected between the two approaches
- AI4SE-generated test cases show slightly higher average evaluation scores
- Effect size differences between AI-generated and handwritten outputs are negligible
- AI-assisted test case generation substantially reduces test development effort and turnaround time
Example Use Case
An example application scenario includes a smart disaster rescue system-of-systems for a metropolitan area given in Figure 2. The operational environment includes flood detection, response coordination, rescue operations, medical treatment, interoperable data interfaces and real-time operational updates for the key stakeholders. This scenario is used for generating and evaluating requirements-based test cases for this integrated operational scenario.

Figure 2: Smart Disaster Rescue System-of-systems Scenario
One example of AI4SE-generated and handwritten test case each are provided in Figure 3.

Figure 3: AI4SE-generated and Handwritten Test Case Examples
Conclusion and Future Work:
This research presents an AI-enabled with human-in-the-loop framework for requirements-based test case generation across the systems engineering lifecycle. Results demonstrate that AI-generated test cases can achieve overall quality comparable to human expert test cases while significantly reducing time and development efforts. These findings show promise that AI-assisted test engineering can augment human experts in developing rapid and scalable test and validation activities for mission-critical systems. Future work includes expanded expert evaluations, larger-scale beta testing, statistical power analysis and assessment across broader operational and mission environments.
References:
Divine, G., Norton, H. J., Hunt, R., & Dienemann, J. (2013). A review of analysis and sample size calculation considerations for Wilcoxon tests. Anesthesia & Analgesia, 117(3), 699–710.
Esho, T., Hoyt, C., Marshall, J., & Gadewadikar, J. (2025). Artificial Intelligence Enabled Systems Engineering Modeling With Retrieval Augmented Generation. Systems Engineering, e70032.
Fagerland, M. W., & Sandvik, L. (2009). The wilcoxon–mann–whitney test under scrutiny. Statistics in Medicine, 28(10), 1487–1497.
Isley, C., Gilbert, J., Kassos, E., Kocher, M., Nie, A., Brunskill, E., Domingue, B., Hofman, J., Legewie, J., & Svoronos, T. (2026). Assessing the quality of AI-Generated exams: A large-scale field study. 40(45), 38626–38634.
Joshi, A., Kale, S., Chandel, S., & Pal, D. K. (2015). Likert scale: Explored and explained. British Journal of Applied Science & Technology, 7(4), 396–403.
Li, W., & Fang, C.-C. (2025). Applying a system dynamics approach for decision-making in software testing projects. PLoS One, 20(5), e0323765. https://doi.org/https://doi.org/10.1371/journal.pone.0323765
Macbeth, G., Razumiejczyk, E., & Ledesma, R. D. (2011). Cliff’s Delta Calculator: A non-parametric effect size program for two groups of observations. Universitas Psychologica, 10(2), 545–555.
Tran, H. K. V., Ali, N. bin, Unterkalmsteiner, M., Börstler, J., & Chatzipetrou, P. (2025). Quality attributes of test cases and test suites–importance & challenges from practitioners’ perspectives. Software Quality Journal, 33(1), 9.
White, R., & Krinke, J. (2022). TCTracer: Establishing test-to-code traceability links using dynamic and static techniques. Empirical Software Engineering, 27(3), 67.
Zhang, W., Cockburn, C., Henshaw, M., Douglas, P., Palmer, P., Olivier‐Myall, J., & Ji, S. (2026). MBSE Co‐Pilot: A research roadmap. Systems Engineering, 29(1), 20–33.
Author Biographies
Muhammad F. Islam, Ph.D., PMP, CISSP, CSEP is an AI and systems engineering leader providing technical guidance and advisory support to the U.S. Federal Government. He has experience in Systems Engineering, Enterprise Architecture, Systems Security, Project Management, and Analytics. He earned a Ph.D. in Systems Engineering from the George Washington University, an M.S. in Electrical Engineering from the University of South Alabama, and a B.S. in Electrical Engineering from West Virginia University. His professional honors include membership in Tau Beta Pi and Eta Kappa Nu. His industry credentials include PMP, CSEP, CISSP, and TOGAF. Dr. Islam is a recipient of the INCOSE Outstanding Leadership Award.
Tomi Esho, Ph.D. is a Senior Systems Engineer at The MITRE Corporation, where he works at the intersection of AI and systems engineering. While at MITRE, he has developed various technical tools and authored several journal and conference papers on AI-enabled systems engineering. Prior to joining MITRE, he received his Ph.D. in Materials Science from the California Institute of Technology, where he was a National Science Foundation Research Fellow and developed computational tools to understand electron transport in semiconductor devices. Dr. Esho received his B.S. degree in Mechanical Engineering from the University of Texas at Arlington with highest honors.
Jyotirmay Gadewadikar, Ph.D. Dr. Jyotirmay Gadewadikar is an AI strategist and technology leader with deep expertise in AI integration and enterprise systems. Currently serving as Chief of AI and Systems Engineering, He has previously held leadership roles, including Chief Product Officer for Conversational AI at Deloitte and AI Strategy Lead at MIT. His expertise spans AI system design, risk management, and innovation. He holds a PhD in Electrical Engineering from the University of Texas System and a System Design and Management Degree from MIT Sloan School of Management. Dr. Gadewadikar is a recipient of the U.S. Department of Homeland Security’s Scientific Leadership Award.
©2026 The MITRE Corporation. ALL RIGHTS RESERVED. Approved for public release. Distribution unlimited 26-00538-2.
Dewey Classification: L 681 12

