JUNE 2026 I Volume 47, Issue 2

Accelerating Test and Evaluation with AI Across the Systems Engineering Lifecycle

Accelerating Test and Evaluation with AI Across the Systems Engineering Lifecycle

Muhammad F. Islam

Muhammad F. Islam

MITRE Corporation, McLean, VA

Tomi Esho

Tomi Esho

MITRE Corporation, McLean, VA

Jyotirmay Gadewadikar

Jyotirmay Gadewadikar

MITRE Corporation, McLean, VA

DOI: 10.61278/itea.47.2.1006

Abstract

Modern systems engineering efforts often have to deliver detailed test and validation outcomes under strict timelines with limited resources. This can result in insufficient testing efforts, release delays, production issues or even critical failures after launch. This research proposes an AI-enabled, human-in-the-loop framework that automates requirements-based test case generation while keeping test engineers in the decision loop. At first, a set of requirements is provided to a large language model (LLM)-based AI framework and to test engineering subject matter experts for draft test case generation, each of whom generates test cases independently as defined by a shared evaluation rubric. The resulting test cases are then evaluated blindly using a common set of criteria by expert evaluators under consistent assessment conditions. The test cases are compared for traceability to requirements, reproducibility, objectivity of expected results, precision of test data, and robustness under boundary and non-standard conditions. This framework enables a comprehensive assessment of how AI-generated tests perform relative to human-developed tests. The proposed framework aims to enable test engineers performing rapid test and validation activities, while also providing empirical evidence on where it is effective and where it can be improved in this exploratory study. This research seeks to reduce manual testing efforts, rework in test development while enabling rapid test delivery for mission-critical systems in defense, aerospace, and other high-assurance domains.

Introduction

Modern systems engineering efforts frequently operate under tight timelines and limited resources. This increases the risk of inadequate testing leading to potential system failures (Li & Fang, 2025; Zhang et al., 2026). This research investigates whether AI-assisted test generation can improve testing efficiency while maintaining the evaluation quality across the systems engineering lifecycle (Esho et al., 2025).

The proposed approach includes the following characteristics and evaluation goals:

  • AI systems and human engineers generate test cases independently
  • Utilizes blind evaluation with common quality criteria
  • Compares AI-generated and human-written test cases
  • Evaluates traceability, reproducibility, objectivity, precision and robustness
  • Explores opportunities to reduce engineering effort and accelerate test delivery
  • Supports mission assurance for high-assurance systems and operational environments

Methodology:

AI-generated (AI4SE) and handwritten test cases are developed using same prompts. 10 AI-generated and 10 handwritten test cases are provided to expert evaluators randomly without identifying which test cases are AI-generated or written by human.Handwritten test case development requires approximately two hours (~120 minutes) for the 10 test cases. In contrast, AI-generated test cases require approximately five seconds for initial generation and approximately 10 minutes for quality validation. This equals roughly 10 minutes of total effort for the 10 test cases. This came out to be a 92% reduction in effort, corresponding to approximately twelve-times faster test development.

Figure 1: AI-generated (AI4SE) Test Case Workflow

After the test cases are generated, expert evaluators assessed each test case using a 1 to 5 Likert scale across five quality dimensions (Joshi et al., 2015; Tran et al., 2025; White & Krinke, 2022):

  • Traceability:Does this test case map directly to a single, unique requirement ID in the Traceability Matrix? (White & Krinke, 2022)
  • Reproducibility:Are the pre-conditions and environment setup described well enough for an independent third party to replicate the test?
  • Objectivity:Are the expected results based on measurable, quantitative data rather than subjective observations?
  • Precision:Does the test procedure use specific test data (exact inputs) rather than generalized instructions?
  • Robustness:Does the test challenge boundary values or off-nominal conditions instead of only the “happy path”?

Likert scale responses included:

  1. Strongly Disagree
  2. Disagree
  3. Neutral
  4. Agree
  5. Strongly Agree

The expert evaluation process focuses on determining whether AI-generated test cases can achieve comparable quality as the handwritten test cases while substantially reducing time and manual engineering efforts (Isley et al., 2026).

Results

The evaluation results demonstrate comparable performance between AI-generated and handwritten test cases across all assessment categories. AI4SE-generated test cases achieved slightly higher overall evaluation scores while substantially minimizing the manual development efforts.

Table 1: Comparative Assessment Results by Evaluation Criterion

Category (Mean ± Standard Deviation) AI4SE-Generated Handwritten
Traceability 5.00 ± 0.00 5.00 ± 0.00
Reproducibility 4.00 ± 0.78 3.93 ± 0.69
Objectivity 4.40 ± 0.67 4.28 ± 0.75
Precision 3.38 ± 0.90 3.18 ± 0.93
Robustness 2.45 ± 0.82 2.45 ± 0.85
Overall 3.85 ± 0.38 3.77 ± 0.47

Additional statistical analyses are conducted to compare AI-generated and handwritten test cases under consistent evaluation conditions.

Table 2: Statistical Analysis Results

Statistical Measure Result Interpretation
Wilcoxon p-value (Fagerland & Sandvik, 2009) 0.25 No statistically reliable evaluator-consistent differences found
Mann-Whitney p-value (Divine et al., 2013; Fagerland & Sandvik, 2009) 0.3093 No reliable evidence that score distributions are statistically different
Cliff’s Delta (Macbeth et al., 2011) 0.131 Small effect favoring that AI-generated test cases score higher

Key findings from the results analysis are summarized below:

  • AI-generated and handwritten test cases achieved comparable evaluation quality
  • No statistically significant differences are detected between the two approaches
  • AI4SE-generated test cases show slightly higher average evaluation scores
  • Effect size differences between AI-generated and handwritten outputs are negligible
  • AI-assisted test case generation substantially reduces test development effort and turnaround time

Example Use Case

An example application scenario includes a smart disaster rescue system-of-systems for a metropolitan area given in Figure 2. The operational environment includes flood detection, response coordination, rescue operations, medical treatment, interoperable data interfaces and real-time operational updates for the key stakeholders. This scenario is used for generating and evaluating requirements-based test cases for this integrated operational scenario.

Figure 2: Smart Disaster Rescue System-of-systems Scenario

One example of AI4SE-generated and handwritten test case each are provided in Figure 3.

Figure 3: AI4SE-generated and Handwritten Test Case Examples

Conclusion and Future Work:

This research presents an AI-enabled with human-in-the-loop framework for requirements-based test case generation across the systems engineering lifecycle. Results demonstrate that AI-generated test cases can achieve overall quality comparable to human expert test cases while significantly reducing time and development efforts. These findings show promise that AI-assisted test engineering can augment human experts in developing rapid and scalable test and validation activities for mission-critical systems. Future work includes expanded expert evaluations, larger-scale beta testing, statistical power analysis and assessment across broader operational and mission environments.

References:

Divine, G., Norton, H. J., Hunt, R., & Dienemann, J. (2013). A review of analysis and sample size calculation considerations for Wilcoxon tests. Anesthesia & Analgesia, 117(3), 699–710.

Esho, T., Hoyt, C., Marshall, J., & Gadewadikar, J. (2025). Artificial Intelligence Enabled Systems Engineering Modeling With Retrieval Augmented Generation. Systems Engineering, e70032.

Fagerland, M. W., & Sandvik, L. (2009). The wilcoxon–mann–whitney test under scrutiny. Statistics in Medicine, 28(10), 1487–1497.

Isley, C., Gilbert, J., Kassos, E., Kocher, M., Nie, A., Brunskill, E., Domingue, B., Hofman, J., Legewie, J., & Svoronos, T. (2026). Assessing the quality of AI-Generated exams: A large-scale field study. 40(45), 38626–38634.

Joshi, A., Kale, S., Chandel, S., & Pal, D. K. (2015). Likert scale: Explored and explained. British Journal of Applied Science & Technology, 7(4), 396–403.

Li, W., & Fang, C.-C. (2025). Applying a system dynamics approach for decision-making in software testing projects. PLoS One, 20(5), e0323765. https://doi.org/https://doi.org/10.1371/journal.pone.0323765

Macbeth, G., Razumiejczyk, E., & Ledesma, R. D. (2011). Cliff’s Delta Calculator: A non-parametric effect size program for two groups of observations. Universitas Psychologica, 10(2), 545–555.

Tran, H. K. V., Ali, N. bin, Unterkalmsteiner, M., Börstler, J., & Chatzipetrou, P. (2025). Quality attributes of test cases and test suites–importance & challenges from practitioners’ perspectives. Software Quality Journal, 33(1), 9.

White, R., & Krinke, J. (2022). TCTracer: Establishing test-to-code traceability links using dynamic and static techniques. Empirical Software Engineering, 27(3), 67.

Zhang, W., Cockburn, C., Henshaw, M., Douglas, P., Palmer, P., Olivier‐Myall, J., & Ji, S. (2026). MBSE Co‐Pilot: A research roadmap. Systems Engineering, 29(1), 20–33.

Author Biographies

Muhammad F. Islam, Ph.D., PMP, CISSP, CSEP is an AI and systems engineering leader providing technical guidance and advisory support to the U.S. Federal Government. He has experience in Systems Engineering, Enterprise Architecture, Systems Security, Project Management, and Analytics. He earned a Ph.D. in Systems Engineering from the George Washington University, an M.S. in Electrical Engineering from the University of South Alabama, and a B.S. in Electrical Engineering from West Virginia University. His professional honors include membership in Tau Beta Pi and Eta Kappa Nu. His industry credentials include PMP, CSEP, CISSP, and TOGAF. Dr. Islam is a recipient of the INCOSE Outstanding Leadership Award.

Tomi Esho, Ph.D. is a Senior Systems Engineer at The MITRE Corporation, where he works at the intersection of AI and systems engineering. While at MITRE, he has developed various technical tools and authored several journal and conference papers on AI-enabled systems engineering. Prior to joining MITRE, he received his Ph.D. in Materials Science from the California Institute of Technology, where he was a National Science Foundation Research Fellow and developed computational tools to understand electron transport in semiconductor devices. Dr. Esho received his B.S. degree in Mechanical Engineering from the University of Texas at Arlington with highest honors.

Jyotirmay Gadewadikar, Ph.D. Dr. Jyotirmay Gadewadikar is an AI strategist and technology leader with deep expertise in AI integration and enterprise systems. Currently serving as Chief of AI and Systems Engineering, He has previously held leadership roles, including Chief Product Officer for Conversational AI at Deloitte and AI Strategy Lead at MIT. His expertise spans AI system design, risk management, and innovation. He holds a PhD in Electrical Engineering from the University of Texas System and a System Design and Management Degree from MIT Sloan School of Management. Dr. Gadewadikar is a recipient of the U.S. Department of Homeland Security’s Scientific Leadership Award.

©2026 The MITRE Corporation. ALL RIGHTS RESERVED. Approved for public release. Distribution unlimited 26-00538-2.

ITEA_Logo2021
ISSN: 1054-0229, ISSN-L: 1054-0229
Dewey Classification: L 681 12

  • Join us on LinkedIn to stay updated with the latest industry insights, valuable content, and professional networking!