Creating Decision Dominance with Agentic AI Planning & Simulations

Presented by Scale AI Scale AI's logo

In August 2026, Chinese state TV aired footage of a PLA Air Force strike-planning system after live-fire testing. It coordinated over 100 tactical units to compress hours of campaign planning into minutes. Today, the Joint Warfighting Concept (JWC 3.0) makes AI-enabled planning doctrine, not an option. But commanders are right to distrust black box systems, and speed is no advantage when the plan behind it was never evaluated.

Agentic AI accelerates the staff work that consumes a planning cycle, and it shortens the path to the evidence a commander needs before committing. Previously, building a scenario, adjusting force posture, debugging the model used in wargames, and interpreting the output all required people who could code against the simulation platform, and there have never been enough of them.  

The effect of agentic AI is measurable. In independent testing by a leading nonprofit defense research organization, professional planning teams using Scale's agentic planning capability completed joint planning mission analysis 36% faster, with output quality equal to or better than the control group. The capability is deployed today at U.S. Pacific Command (PACOM) and U.S. European Command (EUCOM) through the Defense Innovation Unit’s (DIU’s) Thunderforge program.1

In independent testing, teams using Scale's capability completed joint mission analysis 36% faster, with quality equal or better. It is deployed today at PACOM and EUCOM through DIU's Thunderforge Project.

What Changes When Scenarios Take Hours

Agents can now write and modify simulation code, so constructing a scenario, altering force posture or terrain assumptions, and debugging the result become tasks a planner can direct without coding. And modeling specialists can move onto the scenarios that genuinely require them, when setup only takes hours, instead of weeks.

The immediate applications of agentic AI sit at the slowest stages of the Joint Planning Process: estimates, data synthesis, and red team review compressed from days to hours, with existing simulation platforms configured and rerun without dedicated coding support. That speed changes what a staff can attempt — courses of action (COAs) built, tested, and compared as intelligence arrives rather than developed once and shelved, and second and third order effects carried into the plan while there is still time to act. The same approach extends to logistics modeling and coalition planning.

"Military planning processes are decades old and mismatched to the speed of modern warfare. Reliable, domain-specific agentic AI can help commands accelerate tempo, anticipate threats, and create decision advantage."

—Dan Tadross, VP, GM, Public Sector, Scale

What Separates a Deployment from a Demonstration

Two risks separate a deployment from a demonstration: automation bias, which demands that human-on-the-loop design mean something operationally. For example, it should include review points at every major decision, override authority, an audit trail, and outputs with citable sources, because a recommendation a commander cannot explain is one they cannot act on. 

Next, most AI evaluations happen once, but testing must continue long after delivery, when agents drift, data goes stale, and multi-agent workflows fail in ways single-model evaluation misses. Commands need embedded test and evaluation against doctrine-specific criteria and their own mission context, continuous red-teaming, and private benchmarks rather than public leaderboards an adversary can study.

Evaluation criteria can include configuration of your environment, model infrastructure independence, data readiness, orchestration across tasks, capability transfer and more. 

90% of the world's leading frontier models and AI labs, and the DoW, rely on Scale for AI testing and evaluation.

Scale's frontier-trained evaluations are built by vetted, cleared domain experts and run against private benchmarks instead of public leaderboards. Agents are scored against doctrine-specific criteria and mission context, and checked continuously against DoW Ethical AI Principles, including safeguards against exposure of classified or sensitive data.

Where to Start

Start with the planning steps most constrained by specialist availability or speed. Simulation-dependent steps usually top the list. Then, pilot with instrumentation, capturing usability, performance, and mission impact from day one rather than reconstructing it later. Require vendors to benchmark against your mission and your rules, know your data sources and classification levels, and route findings to the J7, CDAO, and your Service, so pilot work carries into doctrine instead of ending as a demonstration. 

Commands that start experimenting with agentic planning now accumulate the institutional knowledge that only comes from running this inside real planning cycles.

Read the eBook, "Creating Decision Dominance at Machine Speed: How Agentic AI Is Helping to Evolve Strategic and Operational Planning written for senior leaders at geographic and functional commands and the J2, J3, J5, and J7 staff who support strategic and operational planning.

Visit www.scale.com/defense to learn more.

Sources

1. Defense Innovation Unit, "DIU's Thunderforge Project to Integrate Commercial AI-Powered Decision Making," 2025. diu.mil/latest/dius-thunderforge-project-to-integrate-commercial-ai-powered-decision-making

This content is made possible by our sponsor Scale AI, Inc., it is not written by and does not necessarily reflect the views of Defense One's editorial staff.

NEXT STORY: GDIT unveils revolutionary modular approach to zero trust