Ivanti AI Security Guidance Cut Hours, but Its Claude Skill Invented Details
Ivanti disclosed that its Claude-based patch analysis skill invented details, despite reducing a recurring workflow from hours to minutes. The Ivanti AI security guidance system confused Microsoft Office editions and repeatedly misread Adobe’s release cadence during development.
That admission matters because the output helps security teams prioritize vulnerabilities after Microsoft’s monthly Patch Tuesday. A plausible fabrication can distort which systems receive attention first, even when every cited vulnerability is real.
Ivanti responded with a human approval gate and publicly identified Claude’s role in a June chart. CrowdStrike and Palo Alto Networks also keep people involved in consequential security automation, but neither has publicly described catching a comparable fabrication.
The story is therefore larger than one model making mistakes. It is a test of whether vendors will disclose how AI shapes operational guidance, what their reviewers actually verify, and which errors remain invisible.
Ivanti AI Security Guidance Entered Production After Months of Corrections
Ivanti moved its Claude skill into production only after its developers repeatedly caught fabricated or conflated details during training.
Chris Goettl, Ivanti’s vice president of product management for endpoint security, built the skill around a process he had performed manually for roughly a decade. It ingests published vendor advisories and Ivanti’s own spreadsheets, then applies the company’s threat-risk prioritization method.
A skill in this context is a reusable set of instructions, sources, and workflow rules provided to an AI model. It does not mean the model independently understands every vendor’s disclosure system.
Goettl trained the system one vendor at a time. Microsoft came first because its downloadable spreadsheets offered relatively structured input. Adobe required a different approach because the model had to interpret individual pages and their changing layouts.
That difference exposed an important weakness. The skill repeatedly produced the wrong Adobe release cadence. It also conflated Microsoft Office editions, which Microsoft separates into several product families.
These were not stylistic flaws. Release cadence and product identity affect how analysts interpret unusual activity, affected software, and patch urgency. A polished summary with the wrong pattern can still send a team in the wrong direction.
Goettl described confronting the system when an answer appeared obviously unsupported. According to his account in the AI-assisted pipeline, Claude acknowledged that the invented material lacked a referenceable foundation.
Ivanti says the skill only uses public vendor advisories and company-controlled spreadsheets. It does not process customer data, and every generated output remains a draft until a person approves it.
Senior product manager Todd Schell replayed the system against two months of briefings previously assembled by hand. Ivanti reported 98 percent alignment, while acknowledging remaining edge cases.
That figure needs careful interpretation. VentureBeat did not independently audit the source data, scoring method, error taxonomy, or replay results. Alignment also does not establish that every relevant vulnerability appeared in the generated output.
The first production run occurred in August 2026 while Goettl was on vacation. A spreadsheet that previously required about four hours from each of two employees reportedly ran in under 30 minutes.
The broader monthly workflow had consumed about 48 combined hours around each Patch Tuesday. Ivanti estimated that effort represented at least 5 percent of Goettl and Schell’s combined working time.
The operational gain is easy to understand. A machine can collect, normalize, and sort a large release faster than two specialists can manually copy fields across spreadsheets.
However, Ivanti did not remove the specialist from the process. The company changed the specialist’s job from assembling every row to reviewing a machine-produced draft.
That distinction defines the central conflict. Automation saves time by reducing manual construction, while reliable security guidance still depends on human judgment and source-level verification.
Ivanti also left a visible disclosure before the broader story emerged. Its June Patch Tuesday post identified a chart as generated with Claude using author-designed prompts and Goettl’s dataset.
No regulation required that caption. No common industry standard dictated its wording. The disclosure remained publicly available for about three months before receiving wider scrutiny.
That quiet caption became significant once the development failures were known. It connected an AI-generated artifact to a named model, a human author, a date, and a defined dataset.
That is more transparency than a generic “AI assisted” label. It still does not reveal which rows were checked, how completeness was measured, or whether the published risk tiers received independent validation.
The Patch Tuesday Data Problem Made Automation Attractive
The immediate pressure came from a Patch Tuesday process that became larger, less standardized, and more dependent on each vendor’s interpretation.
On September 8, 2026, Microsoft released what several security researchers described as its largest Patch Tuesday. Yet the published totals differed substantially across organizations analyzing the same release.
Tenable counted 964 vulnerabilities, including 104 rated critical and 860 rated important. Its September analysis also identified two vulnerabilities exploited in the wild before patches became available.
Ivanti counted 973. Senserva counted 1,169 because it mapped vulnerabilities against knowledge-base articles instead of using the same advisory-centered method as other trackers.
The difference between Tenable’s and Senserva’s totals was 205 vulnerabilities. That gap exceeded the previous Patch Tuesday record cited by Ivanti, which was 175 vulnerabilities in October 2025.
These differences do not prove that an AI model made an error. They can result from legitimate choices about scope, duplicated entries, product variants, advisory boundaries, and knowledge-base mappings.
They do show why a single monthly total cannot be treated as a neutral fact. Every published count now reflects a parsing method and a definition of what belongs in the set.
Microsoft intensified that problem when it stopped presenting one consolidated monthly CVE list in its Security Update Guide. Rapid7 documented the change while preparing its July patch review.
Each security vendor must now reconstruct the release from Microsoft’s available data. That reconstruction can involve deterministic code, manual analysis, an AI model, or a combination of all three.
A deterministic parser follows explicit rules and produces the same result when given identical input. A large language model can interpret inconsistent pages more flexibly, but it can also generate unsupported relationships.
That tradeoff becomes especially important when the output is not merely a count. Security teams need prioritization based on active exploitation, public disclosure, severity, asset exposure, and business importance.
An incorrect count can confuse reporting. An incorrect priority can redirect scarce patching capacity away from a threat that attackers are already exploiting.
Ivanti’s webinar briefing reportedly reaches between 500 and 700 attendees each month. Those participants are not necessarily copying every recommendation directly into production, but the audience gives each prioritization decision meaningful reach.
The time pressure is also real. Security teams cannot calmly inspect nearly one thousand entries before attackers act.
CrowdStrike reported that 88 percent of observed exploitation involving a public proof of concept began within 48 hours. A proof of concept is publicly available code or technical documentation that demonstrates how a vulnerability can be exploited.
Federal requirements can be even tighter. CISA’s Binding Operational Directive 26-04 established a three-day remediation window for the highest-risk vulnerabilities at federal civilian agencies.
Under those conditions, vendors have a strong incentive to automate collection and preliminary classification. Waiting for a perfect manual review can itself create risk.
The problem is not whether organizations should use AI. It is whether they can prove that acceleration has not removed critical vulnerabilities, corrupted product mappings, or elevated weak signals.
Ivanti’s experience shows why this proof cannot come from fluency. The model’s incorrect output looked convincing enough to require expert recognition and direct confrontation.
Security buyers are therefore pressured from both sides. They need faster guidance, yet they cannot assume that a faster, well-written briefing contains a complete or accurately prioritized set.
The Real Reversal Is Disclosure, Not Fabrication
The surprising part is not that Claude fabricated details; it is that Ivanti described the failures and preserved a visible human gate.
Large language models generate text by predicting likely sequences from their training and supplied context. They do not inherently guarantee that every statement maps to a provided source.
Fabrication, often called hallucination, occurs when a model presents unsupported information as if it were grounded. In this workflow, the failure appeared as incorrect vendor patterns and conflated software families.
Those weaknesses are already familiar in general-purpose AI systems. What changes the stakes is their placement inside a pipeline that produces security recommendations.
A consumer chatbot error might waste a reader’s time. A missing vulnerability row in operational guidance can delay remediation on an exposed system.
Ivanti’s disclosure did not eliminate that risk. It made the risk inspectable.
The company attached a model name and human ownership to at least one published graphic. Goettl also explained which problems emerged during training and why the review gate remained necessary.
That level of specificity helps customers ask better questions. They can distinguish model-assisted collection from autonomous prioritization, and they can ask which stage receives human verification.
The admission also complicates conventional trust messaging. Vendors usually emphasize accuracy gains, processing speed, or analyst productivity when announcing AI features.
Development failures rarely receive equal attention. That imbalance encourages buyers to evaluate a workflow using its best benchmark rather than its known failure modes.
Ivanti’s reported 98 percent alignment illustrates the danger. The number sounds reassuring, but its practical meaning depends on what fell within the remaining difference.
An unimportant formatting discrepancy and an omitted actively exploited vulnerability should not receive equal weight. A useful evaluation must classify errors by operational consequence.
Completeness deserves separate treatment from correctness. A reviewer can detect a wrong severity or malformed description in a visible row, but cannot easily notice an absent row.
This is the sharpest challenge to the human-review story. Reading generated output is not the same as reconciling it against an authoritative source inventory.
Kayne McGladrey, an independent virtual chief information security officer and senior IEEE member, argued that customers need a written account of the pipeline. He said vendors should explain where the model operates, how CVEs are parsed, who reviews output, and how much receives source verification.
His criticism went further than requesting a human signature. A replay against previous human-produced months can measure similarity without proving present-month correctness.
A human reviewer can also share the model’s blind spot. If both focus on visible entries, neither will discover an item that disappeared before the review began.
The relevant standard is therefore traceability. Each published recommendation should connect to an authoritative input, while each authoritative input should have a recorded disposition.
That second direction matters most. It converts review from “Does this draft look reasonable?” into “Can every source item be accounted for?”
This is where disciplined knowledge blending becomes relevant beyond patching. Combining sources is useful only when the workflow preserves provenance and exposes conflicts instead of smoothing them away.
Ivanti’s disclosure provides a starting point, not a completed standard. It tells buyers that AI participated and that known fabrications shaped the review process.
It does not publicly establish row-level lineage, independent risk-tier testing, false-negative rates, or the percentage of source material reconciled automatically.
Still, the admission creates pressure on competitors. A vendor that markets AI-assisted security guidance without describing its validation chain now offers customers less information than Ivanti.
The reversal is uncomfortable but constructive. Publicly acknowledging a model’s failures can become evidence of process maturity, while silence can conceal either excellent controls or no controls at all.
Human Review Works Only When It Can Find Missing Evidence
A reviewer adds value only when the review design targets omissions, unsupported claims, and high-impact classification errors.
Ivanti’s workflow keeps a person between Claude’s draft and the published briefing. That is safer than allowing a model to release prioritization automatically.
However, “human in the loop” is a design description, not a quality measurement. Its effectiveness depends on what the person sees, checks, and can stop.
A reviewer scanning prose can identify strange wording, a familiar product assigned to the wrong family, or an implausible release pattern. Goettl’s expertise appears to have caught those visible anomalies during training.
The harder failure is silent exclusion. If a parser or model never creates a row for a vulnerability, a reviewer examining only the final spreadsheet receives no obvious warning.
A stronger system needs reconciliation controls outside the language model. Deterministic checks can compare source identifiers, flag unmatched records, detect duplicate CVEs, and count expected items by vendor.
The model can then handle tasks that benefit from interpretation. It might summarize advisories, normalize inconsistent descriptions, or propose risk categories for expert approval.
This division assigns machines different responsibilities. Code protects completeness and repeatability, while the model assists with ambiguous language and prioritization context.
It also creates clearer failure signals. A reconciliation mismatch can halt publication even when the generated narrative looks polished.
Risk-tier validation needs similar rigor. The system should record why an item received its ranking, which source facts supported that decision, and what changed after human review.
Without that record, a final approval establishes accountability but offers limited evidence about quality. It cannot show whether the reviewer examined every recommendation or merely sampled the highest-risk rows.
Ivanti says its known-exploit, public-disclosure, and abnormal-volume signals trigger escalation. That is a sensible rule set, but the public reporting does not independently verify its application.
The Adobe and Office failures also show why vendor-specific testing matters. A workflow that performs well on Microsoft spreadsheets can fail when another publisher uses web pages, different naming conventions, or irregular release schedules.
Evaluation must therefore cover each data source and each transformation stage. A blended average can hide a weak vendor connector beneath strong Microsoft performance.
CrowdStrike described a different review pattern at Fal.Con 2026. Its “human on the loop” approach reportedly has an analyst work the same detection in parallel with the agent, then compare verdicts.
Parallel work can expose disagreements that sequential review misses. It also retains more human effort than a workflow where one person checks an already completed draft.
Palo Alto Networks introduced Cortex XSIAM AgentiX in February 2026 with prebuilt agents and approval gates for high-impact actions. Its agentic SOC design similarly preserves human control where automated decisions carry greater consequences.
Neither design automatically solves the missing-input problem. A human and an agent can both reason from an incomplete feed, while an approval gate can authorize an action based on faulty evidence.
The comparison nevertheless reveals an emerging consensus. Major vendors are not treating unrestricted autonomy as an acceptable default for high-impact security operations.
That consensus matters because industry language often blurs assistance and autonomy. A system that drafts a briefing differs materially from one that deploys a patch, isolates a device, or closes an incident.
Buyers should map each AI action to its reversibility and potential harm. Low-impact summaries can tolerate lighter controls than decisions that alter production systems.
They should also demand evidence from actual failure cases. A benchmark that reports only agreement conceals whether errors involved wording, scope, severity, or omissions.
Ivanti’s known fabrication catches provide more useful information than a bare accuracy claim. They identify concrete conditions under which the system became unreliable.
The unresolved question is whether the production controls detect new failure modes, not just the ones discovered during training. Vendor sites change, taxonomies evolve, and unusual releases can invalidate yesterday’s parsing assumptions.
Human review remains necessary, but it should sit inside a measurable control system. Otherwise, the phrase can become reassurance without demonstrating that the most dangerous errors are discoverable.
Ivanti’s Admission Raises the Standard for Every Security Vendor
Security vendors now face pressure to disclose not merely that they use AI, but exactly how AI influences customer-facing priorities.
Ivanti is not alone in automating security analysis. CrowdStrike, Palo Alto Networks, and other vendors are embedding models and agents into detection, investigation, triage, and response workflows.
The competitive difference exposed here is transparency. Ivanti named its model, described its inputs, identified known fabrication patterns, and explained that publication requires human approval.
VentureBeat reported that it found no comparable public fabrication disclosure from CrowdStrike or Palo Alto Networks. That absence does not establish that their systems failed or that their controls are weaker.
It means customers lack comparable information. One vendor has exposed part of its failure history, while others mainly describe architecture and safeguards.
Public disclosure creates a difficult incentive problem. A company that reports mistakes can appear less reliable than a competitor that publishes only successful evaluations.
Security procurement can reverse that incentive by rewarding evidence. Buyers can ask every vendor the same questions and treat missing answers as an unresolved control gap.
First, customers should ask which artifacts are AI-generated. The answer should distinguish collection, parsing, summarization, scoring, recommendation, and automated action.
Second, they should ask how completeness is verified. A valid response should address missing rows, duplicate records, source changes, and failed ingestion.
Third, they should ask who reviews the result and what that review covers. A named approval role is useful, but a documented checklist and audit trail offer stronger assurance.
Fourth, they should request error categories rather than one accuracy percentage. Buyers need to know whether failures affect grammar, product attribution, exploit status, severity, or inclusion.
These requests are proportionate to the decision at stake. Patch teams use prioritization because they cannot remediate every issue simultaneously.
The September release demonstrates the scale problem. Tenable identified 964 CVEs, Ivanti identified 973, and Senserva identified 1,169 under a different counting approach.
A buyer does not necessarily need every vendor to produce an identical number. It does need each vendor to explain its scope and reconcile its recommendations to that scope.
The same logic applies to risk tiers. Vendors can reasonably weigh exploitability, exposure, and business context differently, but those judgments should remain traceable.
Transparency also protects vendors from unfair comparisons. A documented methodology can show that two totals differ because of defined scope, not because one system silently lost data.
This issue reaches beyond cybersecurity. Any AI-generated research, compliance brief, financial summary, or operational report can contain plausible omissions that surface review will miss.
Security makes the problem unusually visible because CVEs have identifiers and authoritative advisories. That structure gives vendors a practical way to test completeness.
Organizations working with less structured evidence face a harder task. They still need provenance, conflict detection, and explicit treatment of missing material.
Ivanti’s approach is therefore notable without being sufficient. Its disclosure tells customers more than silence would, while its reported controls still leave important verification questions unanswered.
The company’s past security context makes scrutiny especially important. Customers evaluating patch guidance will judge not only the model’s efficiency, but also Ivanti’s ability to communicate risk accurately.
That scrutiny should remain evidence-based. The Claude skill discussed here analyzed public patch data and did not reportedly access customer environments.
It was not an autonomous remediation agent. It produced drafts for a recurring briefing, with a person responsible for release.
Conflating those categories would overstate the event. Minimizing the fabrication because “a human checked it” would understate the control challenge.
The measured conclusion sits between those extremes. Ivanti created a materially faster process, caught real model failures, disclosed AI involvement, and retained expert review.
It has not publicly demonstrated that its process catches every missing vulnerability or independently validates every priority decision. That is the standard customers should now ask it, and its competitors, to meet.
Three Signals Will Show Whether AI Patch Guidance Earns Trust
The next test is whether vendors turn broad human oversight into visible, repeatable, and source-complete validation.
The first signal will arrive with the next Patch Tuesday on October 13, 2026. Analysts should compare published totals, scope definitions, known-exploit labels, and the treatment of unusual vendor data.
If Ivanti’s output remains fast while clearly accounting for source discrepancies, its case for supervised automation strengthens. An unexplained omission or mapping error would weaken it.
The second signal is more detailed disclosure from Ivanti or its competitors. Useful documentation would identify where models operate, which deterministic controls protect completeness, and which decisions require human approval.
That information would strengthen the argument that transparency can become a competitive standard. Continued reliance on broad “human in the loop” language would leave the central verification gap unresolved.
The third signal is evidence from production evaluations. Vendors should report error categories, source coverage, reviewer overrides, and high-impact false negatives without exposing exploitable customer details.
Such evidence would show whether the systems improve after encountering new formats and edge cases. Aggregate agreement alone would not answer that question.
CrowdStrike’s report about rapid exploitation explains why this work cannot return entirely to manual processing. Its exploitation findings show that defenders often operate inside a shrinking response window.
The right outcome is therefore not less automation. It is automation whose evidence can be inspected before its recommendations move people or machines.
For security leaders, the practical response is to inventory every external guidance source that uses AI. Ask whether the model counts, interprets, ranks, or acts, because each role creates a different failure path.
Then test the vendor’s review claim against one difficult question: How would the process detect a vulnerability that never appeared in the generated draft?
If the answer depends on a person noticing something absent, the control is incomplete. If it includes source reconciliation, exception handling, and recorded human decisions, the workflow is more credible.
Ivanti AI security guidance now provides a public case study for that conversation. Its Claude skill saved substantial analyst time, fabricated details during development, and forced the company to design around those failures.
The disclosure should not earn automatic trust, but it deserves attention. It gives customers concrete weaknesses to examine and gives competitors a transparency benchmark they can exceed.
Before relying on the next AI-assisted security briefing, ask for the model’s role, the source-completeness check, and the reviewer’s actual task. Those answers will reveal more than any headline accuracy score.



