Operation Epic Fury demonstrated AI’s central role in contemporary U.S. military operations. Central Command utilized Claude via Palantir’s Maven Smart System to identify and prioritize approximately 1,000 targets within the first 24 hours of the campaign—a pace exceeding twice the rate of the 2003 Iraq invasion’s initial phase. Over 38 days, the campaign executed 13,000 strikes, per Pentagon data. Project Maven reports this efficiency has since scaled to 5,000 targets daily. The critical issue remains whether the Department of Defense can adequately assess systems before deploying them.
In the past 18 months, the Pentagon dismantled non-statutory barriers to AI adoption—a necessary but incomplete step. The next challenge lies in developing policy, workforce, and infrastructure to enable widespread integration. Adoption success hinges on how broadly an capability is embedded across the military, not just initial deployment. Without systemic capacity, AI use will remain limited to early adopters.
The Pentagon’s AI Acceleration Strategy, launched in January, established a Barrier Removal Board authorized to expedite non-statutory requirements across testing, contracting, and hiring. It mandated adoption of cutting-edge AI models within 30 days of release, deferring detailed guidance. This strategy aligns with an April 2025 executive order streamlining acquisition and the July 2025 AI Action Plan, creating political impetus for rapid implementation. Congress reinforced this through the FY26 National Defense Authorization Act.
The Department set July as the deadline for initial demonstrations of its priority AI projects. These were intended to validate that the strategy enabled real capability gains rather than merely accelerating systems the department cannot scale. Reports indicate the deadline was missed, and no updates have been provided on project status.
Generative AI Demands Custom Systems, Policies, and Resources
Unlike legacy narrow AI systems designed for specific tasks, generative models handle diverse functions while constantly evolving. This necessitates new training frameworks to prevent operator over- or under-reliance on AI systems.
Frontier models exhibit significant reliability gaps. On the HalluLens benchmark, they produce incorrect responses in 27–85% of factual queries and fabricate answers about non-existent entities up to 94% of the time. Pentagon policy has not adequately addressed this reality.
Automation bias risks are well-documented in human-machine teams. Generative models deliver detailed recommendations with precise coordinates and target prioritization, potentially leading operators to unquestioningly accept their outputs. Accelerated deployment timelines undermine their purpose: unvetted models with unmapped failure modes reduce operator trust, leading to bypassing in the field.
On May 1, following a dispute with Anthropic over usage restrictions, the Department expanded its list of authorized frontier AI vendors for classified networks from one to eight. However, this expansion lacks corresponding evaluations, workforce, or legal resources.
National Security Presidential Memo 11 Prioritizes Speed Over Safeguards
Signed June 5, NSPM 11 reinforced the “adopt now, build safeguards later” approach. It called for removing “unnecessary barriers to rapid deployment,” emphasizing assurance—ensuring AI technologies are “reliable, robust, steerable, and controllable”—as one pillar. However, it lacked assigned responsibility or funding for this pillar, pairing operational pressure with a 30-day deployment mandate. This effectively placed deployment ahead of independent testing.
Two offices could theoretically handle assurance: the Chief Digital and AI Office’s Responsible AI Office, and the Director of Operational Test and Evaluation. Neither has the capacity. The Responsible AI Office, tasked with governance and assurance processes, began testing frameworks this spring but lost staff to the deferred-resignation program and return-to-office policies. The Test and Evaluation office suffered a 50% staff cut in May 2025, with only seven days to adapt, and has since removed 100 programs from its oversight list without adjusting methods for rapidly updating models. Independent evaluation capacity for generative AI simply does not exist at the required scale.
The memo ignored a critical truth: operators adopt AI they trust and reject systems they don’t, increasing risks of authorizing unvetted systems. While it directed standardized testing and a compute roadmap—priorities that could accelerate adoption if funded—it currently lacks resources.
Five Core Capacity Challenges
Frontier models update frequently, but the Department lacks sufficient ML engineers, acquisition experts for consumption-based contracts, or operators trained to detect model inaccuracies. Security clearances alone take nearly a year. On compute, Pentagon facilities require 6–18 months for accreditation, or over two years for complex systems. Data access is fragmented: the Department often cannot use operational-generated data for training due to vendor contract limitations.
Evaluation relies on vendor-provided benchmarks via a “rent-a-bench” model rather than independent methodologies. As of 2024, Maven achieved 60% accuracy in object identification versus 84% for human analysts in the 18th Airborne Corps. The Department has no way to track if this gap has narrowed, despite intelligence officials aiming for 1,000 high-quality targeting decisions hourly. With two-year-old data, no external party can assess current performance.
Authority-to-operate processes also hinder progress. Certification for government use can take 12–18 months, designed for static systems, not models updated monthly. Congress has mandated continuous authorization and cross-service reciprocity, but reciprocity often remains a full redo of reviews.
Consequences of Inadequate Oversight
The Department has witnessed failures when deploying systems without mapped failure modes. In 2003, Patriot air defense batteries mistakenly shot down a British aircraft and a Navy jet, killing three aircrew. The system operated autonomously, conditioning operators to trust its outputs unconditionally. A Defense Science Board later found operators had no way to question sensor data when assumptions failed. The Patriot was mature and narrowly scoped; today’s generative models are neither.
When machines outpace human review, scrutiny deteriorates into a formality. +972 Magazine reported Israel’s Lavender system flagged 37,000 Gazans as militants with ~10% error rates. Human reviewers spent ~20 seconds per target. The pattern aligns with automation bias research: fluent AI recommendations reduce human input to a ritual. (While the U.S. strike on an Iranian school killing 156 people remains under investigation, current evidence suggests broader targeting system failures rather than a specific AI error.)
Departmental targeting doctrine requires positive identification, collateral damage assessment, and legal review—processes requiring time. During Operation Epic Fury, Maven’s tempo left no room for such rigor. No one can confirm how many of its identifications were correct, and silence on this gap highlights the urgency of capacity building.
Immediate Actions Required
None of this halts pace-setting projects. They will demonstrate sustainability once overdue demonstrations materialize. The Secretary should consolidate workforce, compute, data, and evaluation infrastructure under a single accountable leader from the Responsible AI Office. While this presents immediate capacity challenges, formal pace-setting status ensures senior leadership visibility and dedicated resourcing. Success metrics should track outputs: engineers cleared, compute delivered, and data rights clauses executed—rather than generic milestones.
The Department must adopt provisional security accreditations as the default for rapidly evolving models. Provisional accreditation focuses on cybersecurity risks rather than static configurations. While this may seem like a speed-over-safety concession, continuous monitoring addresses dynamic threats better than one-time audits. Capability evaluation—assessing whether systems perform as claimed—must remain a separate, rigorous check.
Congress should fund independent testing centers at universities and research institutes. The FY26 National Defense Authorization Act directed the creation of a National Security and Defense AI Institute at a university, but this focuses on workforce development rather than evaluation. Independent testing should measure target-recognition accuracy against human baselines using classified data, conduct adversarial hallucination tests, and evaluate operator scrutiny at operational speeds. Institutions like MIT Lincoln Laboratory could anchor this effort.
Operation Epic Fury proved AI can enable large-scale warfare. Pace-setting projects will show if this momentum can be sustained. However, neither initiative guarantees the military can verify system accuracy in combat or prevent critical failures. Addressing this gap is essential for sustainable adoption.
Jordan Kane advised Congress and the Pentagon for nine years and now works with the Horizon Institute for Public Service. Hamza Chaudhry focuses on AI safety at the Future of Life Institute. This article is adapted from a Modern War Institute publication.
Image: Amn Alenne Mojica via DVIDS


