ISD Chatbot Election Study Finds a 29% Failure Rate, and Spanish Answers Fall Further
The ISD chatbot election study found flawed answers to 29% of basic voting questions, despite every tested system having access to current web information. The problems included inaccurate, incomplete, ambiguous, or outdated guidance. Responses became markedly less reliable when researchers asked the same questions in Spanish.
That language gap turns an ordinary chatbot accuracy problem into an access problem. A mistaken product recommendation can waste money. A wrong registration deadline, identification rule, or mail-ballot instruction can prevent someone from voting.
The Institute for Strategic Dialogue, or ISD, tested models associated with OpenAI, Google, Anthropic, xAI, Meta, and DeepSeek. Its researchers evaluated 2,400 prompts and responses covering election procedures across 10 states.
The results do not show that every chatbot answer is wrong. OpenAI’s tested model performed substantially better than several competitors. Most systems also rejected false claims presented through adversarial prompts.
The reversal is harder to dismiss. AI companies increasingly position chatbots as fast, accessible gateways to current information. Yet the systems struggled with factual questions that have official answers, especially when those answers required Spanish-language precision.
The ISD Chatbot Election Study Tested Real Voting Decisions
The study examined whether chatbots could provide enough correct, current detail for a voter to act safely.
ISD published its chatbot election study on September 3, 2026. Researchers Valeria de la Fuente, Max Read, and Peter Benzoni conducted the testing during June.
The team examined six models available at that time. Those systems were OpenAI’s GPT-5.5, Google’s Gemini 3.5 Flash, Anthropic’s Sonnet 4.6, xAI’s Grok 4.3, Meta’s Muse Spark, and DeepSeek’s V4 Pro.
Researchers submitted 400 prompts to each model, producing 2,400 prompt-and-response pairs. The dataset covered English and Spanish responses across Arizona, Utah, North Carolina, Ohio, Texas, Pennsylvania, Michigan, Georgia, Colorado, and Minnesota.
Those states were not chosen randomly. ISD selected jurisdictions with changing election procedures, pending legislation, active litigation, or histories of disputes over election administration.
The prompts included seven broadly applicable questions about registration, mail voting, identification, and other essential procedures. Eight additional questions were adapted to the rules and circumstances of each state.
Researchers also submitted five adversarial prompts per state. These questions deliberately introduced disputed or false claims, allowing the team to test whether models corrected or amplified them.
The evaluation considered four dimensions. Answers needed to be accurate and timely, handle disputed claims responsibly, cite authoritative material, and maintain quality across English and Spanish.
This approach matters because election guidance cannot be judged only by whether its central sentence sounds plausible. An answer can state a general rule correctly while omitting an exception that determines whether a particular ballot counts.
Pennsylvania supplied a useful example. Researchers asked how residents could register, whether identification was required, and what happened when a mail-ballot envelope lacked a date.
Each question appears straightforward. However, a useful answer must account for state procedures, legal developments, voter circumstances, and relevant deadlines. Generic language is not enough.
The study classified answers that were incomplete, ambiguous, outdated, or inaccurate as problematic. That broader standard explains why the headline figure covers more than obvious hallucinations.
A hallucination is invented information presented as fact. In this study, the more common failure was often less dramatic. Models supplied part of the answer but left out details required for action.
That distinction should shape how readers interpret the 29% figure. The result does not mean chatbots fabricated nearly one-third of everything they said. It means nearly one-third of basic answers failed a practical standard for dependable voting guidance.
Spanish Chatbot Accuracy Fell Across Every Tested Model
Spanish was not an isolated weakness affecting one model. Every tested system delivered less accurate responses in Spanish than in English.
ISD found that accuracy on generic prompts fell by an average of 16% when researchers switched from English to Spanish. Spanish answers were also 6% more likely to be outdated or inaccurate.
The largest problem involved completeness. A response might identify the basic rule while omitting exceptions, procedural steps, or local details needed to use that rule correctly.
That failure becomes serious when the subject is voter registration or ballot handling. A voter cannot recover from a missed deadline by asking for a revised answer after Election Day.
The gap also appeared in the quality and quantity of citations. Researchers found that Spanish answers were much more likely to provide no outside reference, limiting a user’s ability to verify the guidance.
This weak sourcing suggests that Spanish chatbot accuracy depends on more than direct translation. A model must retrieve current Spanish-language material, identify authoritative sources, and preserve procedural distinctions while synthesizing them.
Translation errors added another layer. De la Fuente described models confusing terms for polling places and electoral districts. One response used a word associated with racial categories when referring to an electoral race.
Another model reportedly converted the English word “registration” into the awkward Spanish term “registración.” The appropriate election term in that context was “registro.”
These examples may sound minor to an English speaker. They are not minor when a user needs to identify the correct office, form, jurisdiction, or registration process.
Meta’s tested model recorded the widest reported language gap. Muse Spark’s rate of accurate and complete answers fell from 61.3% in English to 38% in Spanish.
Google’s Gemini 3.5 Flash declined from 84% in English to 64% in Spanish. That 20-point difference appeared even though Google says Gemini aims to provide accurate, timely information across languages.
GPT-5.5 delivered the strongest overall results. It produced accurate and complete responses 89.3% of the time in English and 82% in Spanish.
That relative lead is important, but it does not erase the underlying concern. An 82% success rate still leaves consequential room for error when users request instructions about exercising a legal right.
The findings also continue a documented pattern. A 2024 evaluation found wrong information in 52% of Spanish answers, compared with 43% in English, across five then-current models.
That earlier language accuracy test involved different models, prompts, and election circumstances. Its percentages should not be treated as a direct time series.
Still, both evaluations found the same directional problem. Switching to Spanish reduced reliability, even as the tested models and their web-search capabilities changed.
The Real Conflict Is Accessibility Versus Dependability
Chatbots make complicated election information easier to ask about, but that accessibility becomes dangerous when confidence exceeds accuracy.
Conversational interfaces remove several barriers. Users can describe unusual circumstances, ask follow-up questions, and request simpler language without navigating a complex government website.
Those benefits matter in the United States, where election administration is decentralized. Rules differ across states and can vary among counties, cities, and townships.
The U.S. Election Assistance Commission says the country has more than 10,000 election jurisdictions. Each can involve different officials, procedures, voting methods, and local resources.
A chatbot appears well suited to that complexity. It can search several pages, summarize the rules, and tailor an answer to a voter’s location.
The ISD results show why this apparent fit is also a trap. The more specific the question becomes, the more damaging a missing qualification can be.
Language models generate responses by predicting text from patterns and retrieved information. They do not inherently understand which omitted detail will invalidate a registration or ballot.
Web access should reduce outdated answers, but it does not guarantee dependable retrieval. A model can find an old page, misunderstand a current one, or merge rules from different jurisdictions.
DeepSeek’s V4 Pro reportedly referred to 2024 election dates in 15 responses to questions about the 2026 election. Anthropic’s Sonnet 4.6 described the current election as the 2025 cycle three times.
Meta’s Muse Spark gave the wrong election date twice, identifying November 4 rather than November 3. This was not a hidden legal nuance. It was a basic calendar fact.
These failures pressure every company offering search-connected assistants. The product promise increasingly extends beyond creative drafting into factual synthesis, research, and personal decision support.
Election information exposes the weakness in that promise. A polished answer may compress several sources while concealing which sentence came from which authority.
The user then faces an asymmetric risk. Accepting a correct answer saves several minutes. Accepting a confident but incomplete answer can affect registration, polling access, or ballot validity.
The risk is particularly relevant because chatbot use is already mainstream. A 2026 AI usage survey found that 49% of Hispanic adults used chatbots at least occasionally.
The survey found daily use among 26% of Hispanic adults, compared with 24% of all U.S. adults. These figures cover general chatbot use, not election questions specifically.
ISD also noted that relatively few voters knowingly seek political information through chatbots. However, users increasingly encounter AI summaries inside search products, browsers, and productivity tools.
That indirect exposure changes the stakes. Someone does not need to open a dedicated chatbot and request election advice to receive an AI-generated account of voting rules.
Better Debunking Did Not Fix Routine Election Errors
The systems generally resisted planted falsehoods, yet they still mishandled ordinary questions that voters were more likely to ask.
ISD’s adversarial prompts included false or disputed claims. Only 1.3% of English responses and 2.3% of Spanish responses showed a tendency to affirm or amplify those claims.
That is the study’s most encouraging result. It suggests election safeguards can help models reject familiar narratives about fraud or previously debunked controversies.
The stronger performance also indicates that companies have invested in recognizing contentious election prompts. Several developers have introduced policies, safety tuning, or referrals to trusted voting resources.
Yet those protections focus on a visible category of harm. A prompt alleging mass fraud clearly signals that the model should slow down, challenge the premise, or consult vetted sources.
Routine questions may not activate the same safeguards. Asking where to register or whether identification is required sounds harmless, even when an incomplete answer carries direct consequences.
This produces a revealing reversal. The chatbots often handled deliberately misleading premises better than mundane procedural requests.
The result challenges a narrow definition of election misinformation. Harm does not require a coordinated influence campaign or a fabricated conspiracy.
An old deadline can mislead. A missing eligibility exception can mislead. An answer that names a general state rule without explaining local administration can also mislead.
The official AI guidance from the Election Assistance Commission makes this distinction clear. It warns that plausible AI-generated information is often inaccurate and can be particularly harmful to voters.
The study’s sourcing results help explain the failures. Official election information sits across secretary-of-state sites, county pages, court decisions, legislative updates, and voter assistance documents.
Some of those resources receive frequent updates. Others remain online after the relevant election cycle, creating a retrieval environment where old and current guidance can appear equally authoritative.
The problem becomes worse in Spanish. Authoritative translations may appear later, cover fewer questions, or use language that differs across jurisdictions and communities.
A chatbot can therefore retrieve a correct English page without finding an equivalent Spanish resource. It may translate the English content itself, summarize a less authoritative page, or answer without citations.
Max Read, ISD’s director of civic innovation, described chatbots as tools that sometimes provide more than users requested. Their confident elaboration can distort information while appearing helpful.
That tendency is familiar across generative AI products. Models are optimized to answer, not simply to declare that an official source should make the decision.
For election questions, restraint can be the safer capability. A responsible system should recognize when local rules, litigation, or missing context prevent a definitive response.
The Benchmark Has Limits, but the Warning Survives Them
The study is a significant snapshot, not a permanent ranking of AI companies or a complete simulation of consumer chatbot use.
Most models were tested through OpenRouter and application programming interfaces, or APIs. An API lets software send prompts directly to a model without using its normal consumer interface.
Meta’s Muse Spark was tested through its public portal because researchers lacked API access. That difference complicates direct comparisons between Meta and the other developers.
Consumer applications can add system instructions, search tools, election-specific warnings, and safety filters around the underlying model. An API response may not include the same protections.
Anthropic emphasized this limitation in its response to reporting about the study. The company said the third-party API setup did not reflect how most people use Claude or include its consumer voting-information feature.
The objection is reasonable. A model benchmark should not automatically become a product benchmark when the product applies additional retrieval and safety layers.
However, the limitation does not invalidate the broader finding. APIs are used by developers to build services that people encounter through other websites and applications.
The study also tested each tailored question only once per model. Language models can produce different answers when a prompt is repeated, even without an obvious product change.
One response therefore cannot establish a model’s stable failure rate for every user. A stronger longitudinal test would repeat prompts over time, across interfaces, and under controlled account conditions.
The models have also changed since June. Several companies released updates before the report appeared, so the exact rankings may already differ.
These limitations matter most when comparing individual developers. They matter less to the central conclusion that all six tested systems produced election-information failures.
The study should therefore be read as a stress test of an information channel, not a final league table. GPT-5.5’s leading score does not certify every ChatGPT answer.
Likewise, Muse Spark’s low score does not prove every Meta AI interface will perform identically. It shows that the evaluated system failed frequently under the recorded conditions.
The historical evidence strengthens that interpretation. In early 2024, election officials and researchers tested five earlier models and rated more than half their answers inaccurate.
Participants in that earlier election chatbot test classified 40% of responses as harmful. Some systems invented polling places or relied on outdated information.
The newer ISD chatbot election study indicates improvement in several areas, especially resistance to planted falsehoods. It also shows that procedural reliability remains unresolved.
That is the appropriate balance. The evidence does not support saying chatbots always misinform voters. It also does not support treating current assistants as authoritative election officials.
What Voters and AI Companies Should Watch Next
Three signals will show whether the language gap is shrinking: repeated public testing, stronger official sourcing, and measurable parity across consumer products.
The first signal is repeated evaluation before the November 3 election. Researchers should rerun identical prompts across several dates because election rules and model behavior both change.
Repeated tests would reveal whether incorrect answers persist or disappear after developers update their systems. They would also separate random variation from systematic weaknesses.
The second signal is citation quality. A useful response should direct voters to the correct state or local authority, not merely attach several links.
Official sourcing must also be current and language appropriate. A Spanish answer should connect users with Spanish guidance when that resource exists.
The Election Assistance Commission maintains multilingual voter FAQs and directs voters toward state and territorial election websites. Those official destinations should remain the final authority for deadlines, identification rules, and polling locations.
The third signal is parity inside consumer interfaces. Companies should test the complete products that people actually use, including system prompts, web retrieval, safety notices, and location features.
Publishing English-only accuracy scores will not reveal whether Spanish users receive the same detail. Developers need language-specific measures for correctness, completeness, freshness, and authoritative citations.
The companies have offered different responses. Meta says its assistant is designed to provide localized information or direct users toward trusted government sources.
Google says it continually improves Gemini’s ability to provide timely, accurate answers across languages. Anthropic says Claude is designed to answer voting questions accurately and challenge false claims.
OpenAI and Anthropic have also announced work with Democracy Works, a nonprofit provider of voting information. Such partnerships can improve retrieval, but their effectiveness requires independent testing.
The ISD chatbot election study establishes a practical baseline for that scrutiny. The issue is not whether one model wins a benchmark by several percentage points.
The real test is whether a voter can act on an answer without discovering that a missing date, local exception, or translation mistake changed its meaning.
For now, users should treat chatbots as starting points for forming questions, not as final authorities on voting procedures. Verify every actionable answer with a state or local election office. If an answer lacks a current official source, ask the system for one and inspect it directly. AI companies, meanwhile, should publish repeated bilingual evaluations instead of relying on broad claims about multilingual quality. Election officials can help by maintaining current, searchable guidance in both English and Spanish. Before the midterms, the decisive question is simple: will independently repeated tests show that Spanish-speaking voters finally receive the same complete, sourced guidance as English-speaking voters?



