Ludicity Says AI Mania Is Eviscerating Global Decision-Making, and the Evidence Supports the Warning
- Aisha Washington

- 3 days ago
- 14 min read
Updated: 3 days ago
Ludicity published a blunt warning on July 18, 2026: AI mania is eviscerating global decision-making, despite mounting evidence that many enterprise projects never deliver measurable value.
The essay is not another prediction about an approaching AI bubble. It describes institutions already changing how they approve projects, hire employees, evaluate performance, and report failures. Its central allegation is that leaders increasingly reward visible AI adoption over useful outcomes.
That claim comes from one practitioner’s experience, not a representative global study. The author reports roughly 300 professional conversations and says every observed AI project failed during an 18-month period. That reported zero percent success rate has not been independently verified.
Yet the broader warning does not stand alone. RAND found that more than 80 percent of AI projects fail by some estimates. IBM found that only 25 percent of surveyed CEOs said their AI initiatives had delivered expected returns.
The conflict is no longer AI believers against AI skeptics. It is the institutional promise of rational investment against a system that can punish anyone who asks whether an AI project works.
A Viral Essay Turns AI Failure Into a Governance Story
The important change is not another disappointing chatbot. It is the growing inability of organizations to recognize and stop disappointing projects.
The AI mania warning begins with an unusually broad observation. Its author says interactions with professionals, executives, and public-sector institutions reveal organizations trapped between executive enthusiasm and private doubt.
According to the essay, leaders often lack a practical AI plan. Other decision-makers recognize the problem but avoid challenging it because their careers depend on supporting the approved narrative.
That distinction moves the story beyond ordinary technology failure. Software projects have always exceeded budgets, missed deadlines, and disappointed users. Those failures become more dangerous when employees cannot discuss them without creating professional risk.
The essay’s most provocative number is its least verifiable. The author says their team observed a zero percent success rate across AI projects during the previous year and a half.
That sample is neither defined nor independently audited. We do not know how many projects it includes, how success was measured, or whether the author disproportionately encounters distressed organizations. Consulting work often exposes practitioners to difficult projects rather than representative ones.
The number should therefore be read as a reported experience, not a global failure rate. However, the examples surrounding it show how distorted incentives can hide weak performance.
One example concerns internal chatbots. The essay argues that employees rarely adopt them because the underlying company documentation is incomplete, outdated, or inaccessible.
A large language model, or LLM, generates responses from patterns in its accessible data. It cannot reliably recover institutional knowledge that was never recorded or connected to the system.
This creates an uncomfortable dependency. A chatbot sold as a solution to fragmented knowledge actually depends on an organization having already solved much of that fragmentation.
Companies sometimes reverse that relationship in their planning. They buy the interface first, then discover that the information beneath it is missing, contradictory, or governed by incompatible permissions.
The essay also describes a Mitsubishi customer-service interaction. A natural-sounding voice bot collected the author’s automotive problem and promised a return call, but the author says no call came.
The original article says this happened six months before publication. The AIHOT summary describes a six-week delay, but the source essay itself states six months. The longer figure is the one supported by the published account.
Mitsubishi has not independently confirmed the incident. Still, the example captures a measurement problem that applies far beyond one company.
A dashboard might record that the bot completed the call, recognized the request, or avoided an immediate transfer. The customer experienced something different: an unresolved problem and a broken promise.
If management tracks containment instead of resolution, the failure can appear as success. Containment measures whether automation prevented contact with an employee. Resolution measures whether the customer’s actual need was met.
That gap explains why AI mania is eviscerating global decision-making. Leaders can receive positive operational metrics while customers, employees, and unfinished work absorb the real costs.
The essay offers another striking example involving a natural-language analytics demonstration. Prospective buyers reportedly became eager to purchase the tool even after being warned that its answers were not reliable enough for production.
A demonstration compresses uncertainty into a persuasive moment. It shows the correct response to a prepared question, but it rarely displays permission failures, ambiguous terminology, missing records, or unusual user behavior.
Those problems emerge after deployment. By then, the executive sponsor has often attached personal credibility to the project.
The news, then, is not that generative AI sometimes fails. The news is that AI demonstrations, mandates, and career incentives can weaken the feedback systems that normally expose failure.
AI Mania Is Eviscerating Global Decision-Making Through Bad Incentives
When AI usage becomes a target, employees optimize for visible usage instead of better work.
The essay describes companies requiring employees to demonstrate that AI cannot solve a problem before approving additional headcount. Similar policies present automation as the default and human labor as an exception requiring justification.
That rule sounds financially disciplined. It asks managers to test a cheaper option before expanding payroll.
The problem lies in what the rule actually measures. It does not ask whether automation can complete a task reliably, preserve customer trust, or survive unusual conditions. It asks whether someone can claim they tried AI.
The employee also faces an asymmetric risk. Reporting success supports leadership’s strategy. Reporting failure can suggest resistance, poor technical ability, or an inability to adapt.
Once that incentive becomes clear, honest reporting stops being the rational career choice. Employees learn to describe ordinary work as AI-assisted, avoid documenting weak outputs, or select metrics that make adoption look healthy.
The essay claims some organizations use token-consumption leaderboards and usage quotas. In those environments, higher model usage can become evidence of commitment.
Consumption is not productivity. A system that rewards token volume encourages longer prompts, redundant experiments, and automated activity that produces no usable output.
This is a familiar management failure disguised as technical measurement. Organizations choose a proxy because the desired outcome is difficult to measure, then employees optimize the proxy.
The result resembles Goodhart’s law: once a measure becomes a target, it often stops being a useful measure. AI makes that problem worse because it can generate enormous volumes of visible activity at low marginal effort.
Executives may see rising usage, more generated code, or shorter first-draft times. Those numbers omit review costs, correction work, security testing, customer escalation, and the opportunity cost of abandoned priorities.
AI-assisted coding illustrates the issue. Counting generated lines rewards volume, although maintainable software often requires deleting code. Measuring accepted suggestions ignores the time required to inspect them.
The same distortion appears in customer service. Counting automated conversations favors containment. It says nothing about whether customers returned, escalated through another channel, or quietly left.
Knowledge work presents an even harder problem. An AI-generated summary can reduce drafting time while introducing a subtle factual error. The saved minutes are visible, but the later decision based on that error is difficult to attribute.
Organizations need trustworthy information before they can automate decisions around it. A searchable knowledge base can help people retrieve evidence, but it does not replace ownership, access controls, or source verification.
The pressure extends beyond individual metrics. Executives also face incentives from boards, investors, competitors, and vendors.
A leader who delays an AI investment risks appearing passive if a rival announces one. A leader who approves a weak project can blame immature technology, integration difficulty, or changing market conditions later.
This imbalance rewards action over judgment. Announcing an AI initiative produces immediate reputational value, while its operational cost appears gradually across several departments.
IBM’s 2025 CEO study captured that pressure clearly. The research surveyed 2,000 CEOs across 33 countries and 24 industries.
Sixty-four percent said fear of falling behind drives some technology investment before leaders clearly understand its value. Only 37 percent preferred being fast and wrong to being right and slow.
Those answers reveal a contradiction. Most CEOs reject careless speed in principle, yet nearly two-thirds acknowledge investing before value becomes clear.
The same study found that 50 percent said rapid investment had created disconnected technology. Only 25 percent reported expected returns from AI initiatives, and 16 percent reported enterprise-wide scaling.
These findings do not prove that most AI spending is irrational. IBM sells AI and consulting services, and its survey records executive responses rather than audited project outcomes.
However, the data confirms the organizational tension described by Ludicity. Leaders feel compelled to invest while their data, systems, and management structures remain unprepared.
The strongest pressure falls on middle managers and technical specialists. Executives set an AI-first direction, while those below them must translate an abstract demand into production software.
They inherit conflicting responsibilities. They must support the strategy, protect operations, manage compliance, and report progress without appearing obstructive.
This is where governance starts to fracture. The people closest to implementation possess the best evidence, but they can have the least freedom to challenge the project.
The Real Failure Often Starts Before the Model
Enterprise AI fails most predictably when organizations select technology before defining the problem, user, and acceptable error rate.
Ludicity’s argument becomes more credible when compared with independent research into AI project failure.
RAND interviewed 65 experienced data scientists and engineers for its 2024 AI failure research. Participants came from academia, different industries, and companies of varying sizes.
The researchers identified five recurring causes. Organizations misunderstood the problem, lacked suitable data, chased new technology, lacked deployment infrastructure, or attempted tasks beyond current AI capabilities.
These causes matter because only one centers on model performance. The others concern management, information, infrastructure, and problem selection.
RAND also cited estimates that more than 80 percent of AI projects fail. It noted that this would be twice the failure rate of corporate IT projects without AI.
The scope requires care. RAND explicitly excluded projects that only used pretrained LLMs through prompt engineering. Its findings therefore should not be treated as a direct measurement of every chatbot or generative AI deployment.
The mechanisms still transfer. A customer-service bot cannot succeed without accurate policies, clear escalation rules, reliable integrations, and someone accountable for unresolved requests.
An internal assistant cannot answer questions reliably if departments use inconsistent definitions or restrict access to necessary documents. A coding assistant cannot compensate for missing tests and unclear system ownership.
AI adds uncertainty to every weakness already present in software delivery. Outputs can vary between runs, model behavior can change after updates, and confident language can conceal an incorrect answer.
That makes evaluation more important, not less. Yet evaluation often receives less attention than the demonstration because it creates friction during procurement.
A reliable evaluation starts with the job. Teams must define who uses the system, what decision it supports, what evidence it needs, and what errors remain acceptable.
They must also identify the fallback. If an AI system lacks confidence, encounters missing data, or produces conflicting answers, responsibility must move somewhere explicit.
Many projects never settle these questions. They begin with a directive to “use AI,” then search for a workflow that can justify the chosen technology.
That is the exact reversal RAND warns against. Its researchers recommend focusing on enduring problems rather than chasing the latest tool.
RAND advises leaders to choose problems deserving at least a year of sustained team commitment. Constantly shifting priorities prevents teams from understanding data and reaching production.
This creates a direct conflict with mania-driven procurement. The executive wants visible movement this quarter, while a dependable system requires patient integration, evaluation, and workflow redesign.
Generative AI pilots reveal the same divide. A 2025 MIT NANDA report, summarized in the GenAI Divide findings, examined 300 public deployments and included interviews and employee surveys.
The report said about 5 percent of pilots achieved rapid revenue acceleration. The remaining deployments stalled or produced little measurable profit-and-loss impact.
The widely repeated 95 percent figure needs precise framing. It concerns measurable business impact among the studied enterprise implementations, not a universal technical failure rate for all generative AI usage.
A pilot can function technically while failing financially. Employees might use it, yet the organization might see no measurable improvement after licensing, integration, review, training, and support costs.
The report attributed much of the divide to integration and organizational learning. Generic assistants served individuals well because users could adapt them to varied tasks.
Enterprise systems faced a different requirement. They had to learn from particular workflows, preserve context, connect to systems of record, and fit existing responsibilities.
The report also found a purchasing contrast. Specialized external solutions reportedly succeeded more often than internal builds, although “success” depended on the study’s definitions and sample.
That does not mean companies should outsource every AI project. It suggests that building a model interface is easier than maintaining an application around a durable operational problem.
Successful vendors often narrow the job. They build around one workflow, integrate deeply, and absorb the continuing burden of evaluation.
Internal teams can do the same, but only when they possess stable ownership and domain knowledge. A temporary AI task force rarely has either.
This is why model comparisons can distract executives. A modest model embedded in a well-designed process can create more value than a leading model connected to unreliable data.
The hard work remains familiar: map the workflow, clean the records, define authority, test edge cases, train users, monitor outcomes, and fund maintenance.
AI does not remove those obligations. It makes organizations pay for ignoring them in less predictable ways.
The Zero Percent Claim Is Unverified, but the Warning Survives
The essay overreaches when it turns a consulting sample into a global verdict, yet its incentive analysis remains difficult to dismiss.
A responsible reading must separate three claims.
First, Ludicity reports that its team observed zero successful AI projects during 18 months. That is a firsthand claim, but readers cannot independently assess the project list or success criteria.
Second, the essay suggests most public claims about major AI productivity gains are untrue. Available evidence does not establish that sweeping conclusion.
Third, it argues that organizations increasingly distort decisions to signal AI enthusiasm. Independent survey and project research provides meaningful support for this narrower claim.
Keeping those claims separate prevents skepticism from becoming another form of group loyalty. Readers do not need to accept the essay’s entire worldview to recognize the governance problem.
The zero percent figure has obvious selection bias. A consultancy may encounter organizations precisely because their systems, processes, or internal capabilities are weak.
Successful companies might also disclose less information. A narrow internal automation can deliver value without producing a public case study or requiring outside intervention.
Definitions create another complication. One team defines success as deployment. Another requires adoption, financial return, customer satisfaction, or a sustained operational improvement.
A chatbot can meet its launch date and still damage customer trust. A coding tool can accelerate drafting while increasing review time.
Conversely, a small productivity gain can matter even when it never appears as a distinct profit line. Employees may complete research faster or explore more alternatives without reducing headcount.
The essay sometimes treats ineffective enterprise programs as evidence against LLM utility generally. That inference goes too far.
Individual users have found real value in coding assistance, transcription, translation, drafting, and information retrieval. The presence of failed institutional deployments does not erase those uses.
The MIT findings make this distinction important. Generic AI tools can succeed for flexible individual tasks while custom enterprise systems stall during integration.
A balanced conclusion is therefore not that AI does nothing. It is that useful local capability does not automatically become an effective organization-wide program.
Scale changes the problem. A personal assistant can tolerate occasional errors because its user reviews the output and understands the context.
An enterprise service crosses departments, permissions, regulations, and customer relationships. The organization must define who is responsible when the model is wrong.
AI supporters can also point to successful applications outside conversational assistants. RAND cited uses in drug development, supply-chain forecasting, defense sensing, and autonomous flight.
These cases usually involve narrow objectives, specialist teams, strong data pipelines, and extensive testing. They support the problem-first model rather than the adoption-first model criticized by Ludicity.
The essay’s strongest contribution is therefore diagnostic, not statistical. It explains why weak projects can persist after evidence turns against them.
An executive sponsor announces the initiative. Procurement commits resources. A team builds the demonstration. Managers report usage. Employees learn that skepticism carries risk.
Each step increases the social cost of reversal. Canceling the project then requires leaders to admit that earlier confidence exceeded the available evidence.
Organizations often respond by changing the measure. A revenue goal becomes an engagement goal. An adoption goal becomes a license-distribution milestone.
The project survives, but its original purpose disappears. That is a decision failure even if the software remains online.
AI can intensify this pattern because its outputs look intelligent. Fluent text and natural speech encourage people to infer reasoning, understanding, and reliability beyond what the system has demonstrated.
A polished interface also makes technical uncertainty feel smaller. The user sees a confident answer rather than the inaccessible documents, probabilistic generation, and brittle integrations behind it.
This is why governance cannot rely on demonstrations or executive confidence. It needs independent evaluation and metrics tied to the underlying job.
The AI risk framework from the National Institute of Standards and Technology offers one useful reference. It organizes AI risk work around governing, mapping, measuring, and managing systems.
The value of that structure is procedural. It forces organizations to connect model behavior with context, affected users, measurement, and ongoing responsibility.
A framework alone cannot overcome a culture that punishes bad news. Leaders must also protect the people who report failed tests, low adoption, and customer harm.
Without that protection, formal governance becomes theater. Committees approve documents while operational teams continue optimizing whatever number leadership wants to see.
The key uncertainty is cultural. Companies know how to run controlled trials, stage deployments, and measure customer outcomes.
What remains unclear is whether leaders will accept those results when the evidence conflicts with an AI-first strategy they have already promoted.
Three Signals Will Show Whether Institutions Recover
The next phase will be decided by cancellation discipline, outcome-based reporting, and whether employees can challenge AI mandates safely.
The first signal is the treatment of failed pilots. Over the next several months, companies should begin disclosing which projects moved into production and which were stopped.
Cancellation is not automatically evidence of failure. Experiments exist to test uncertainty, and a healthy portfolio should close projects that cannot meet their targets.
The revealing question is whether executives treat cancellation as learning or quietly redefine the project. Transparent closure would strengthen the case that institutional decision-making is recovering.
A project should have a named user, a baseline, an acceptable error rate, an owner, and a stop condition before launch. Those fields make retrospective storytelling harder.
Watch whether companies publish these operational details. Announcements focused only on model partnerships, licenses, or the number of enabled employees provide little evidence of value.
The second signal is a move from adoption metrics to outcome metrics. Token usage, generated documents, and chatbot conversations measure activity.
More useful measures include completed cases, verified accuracy, repeat usage, escalation rates, cycle time, customer retention, and total review effort.
These metrics must be paired. Faster output means little if correction time rises. Automated containment means little if unresolved customers disappear.
IBM reported that 68 percent of surveyed CEOs believed their organizations had clear innovation-return metrics. That confidence sits awkwardly beside the same survey’s limited scaling and return figures.
The discrepancy deserves attention. If metrics are clear, companies should be able to explain which workflows improved and how they separated AI effects from other changes.
More detailed reporting would weaken Ludicity’s broad suspicion that organizations are hiding poor performance. Continued reliance on vague productivity claims would strengthen it.
The third signal is whether companies preserve dissent. Employees should be able to question an AI use case without being labeled resistant or technically deficient.
This does not require allowing endless obstruction. It requires a review process where concerns about data, reliability, security, and customer impact receive documented answers.
Leaders can test their culture with a simple question: can the project’s technical owner recommend cancellation without creating career risk?
If the answer is no, the organization has no effective stop mechanism. It has an AI mandate supported by ceremonial evaluation.
Knowledge workers should also examine their own habits. The danger is not limited to executive strategy.
An employee can become dependent on generated summaries without checking sources. A manager can confuse a polished plan with a feasible plan. A buyer can mistake a smooth demonstration for production readiness.
Tools that blend personal notes and source material can support judgment when they preserve evidence. A personal knowledge workflow is most useful when users can trace conclusions back to original records.
The objective is not to remove AI from decision-making. It is to prevent AI enthusiasm from deciding what evidence the organization is willing to see.
That distinction matters for developers, enterprise buyers, and ordinary AI users. Developers inherit the systems created by vague mandates. Buyers absorb integration risk. Users experience failures that dashboards can hide.
The Ludicity essay will not settle the enterprise AI debate. Its anecdotes are too selective, and its largest claims remain unverified.
It has nevertheless identified a critical failure mode. AI mania is eviscerating global decision-making when institutions reward technological allegiance, suppress negative evidence, and measure activity instead of resolved problems.
The next move belongs to the people approving these systems. Before funding another pilot, ask what result would justify stopping it, who can report that result, and whether leadership is prepared to listen.


