Stephen Aarons ChatGPT Brief Invented Witnesses, and the Court Imposed a $5,000 Fine
Stephen Aarons filed a ChatGPT-assisted murder appeal containing invented witnesses, and New Mexico’s Supreme Court responded with contempt and a $5,000 sanction. The court also removed Aarons from the appeal, referred him for disciplinary review, and struck every brief filed in the case.
The Stephen Aarons ChatGPT brief represents more than another embarrassing citation failure. According to the court, the filing attributed testimony to people who never appeared at trial. It also changed statements from real witnesses and misrepresented existing legal authorities.
That distinction raises the stakes. Earlier legal AI failures often involved nonexistent cases that lawyers could have found through routine citation checks. This filing reportedly altered the factual record underlying a murder conviction, where the client was serving a life sentence.
The case now places a clear opponent at the center of legal AI adoption: automated drafting speed versus a lawyer’s personal duty to verify the record. Courts are not treating software output as an independent participant. They are holding the attorney who signs the filing responsible for every sentence.
The Court Found Fabricated Witnesses Inside a Murder Appeal
The sanction followed admitted factual inventions, not a disagreement over how a lawyer interpreted the evidence.
The New Mexico Supreme Court issued its dispositional order on September 9, 2026, in State v. Sandoval. The case concerns Oscar Renee Sandoval’s appeal from a murder conviction and carries docket number S-1-SC-40845.
Aarons acknowledged using ChatGPT while preparing the opening appellate brief. He also admitted that he had not verified its factual claims or legal authorities before signing and filing it.
The court identified four wholly fabricated witnesses: Officer Michelle Amarillo, Officer Sanchez, Manal Al-Jibury, and Teresa Marquez. Those were not minor spelling mistakes or uncertain identifications. The order says the witnesses themselves were false.
The filing also attributed invented testimony to actual witnesses. It claimed Danny Stanton received threats from Sandoval and took them seriously. It included false statements about threats received by Linda Stanton’s husband.
Other passages reportedly misrepresented testimony about the shooter’s clothing and appearance. The order connected those statements to Mariah Chavez and the fabricated Teresa Marquez.
The errors extended into legal research. The court said the filing misrepresented authority from State v. Lopez and State v. Manus. Those cases exist, which makes the problem different from simply citing an imaginary decision.
A fabricated case can sometimes be exposed by searching a citation. A distorted real decision demands a closer comparison between the proposition and the source. Altered testimony requires another verification process against transcripts and exhibits.
The contempt order records Aarons’ admission that he performed neither review. It also says he did not tell his client about the filing’s factual and legal misrepresentations.
The court had ordered Aarons to appear on August 21 and explain why he should not face contempt and disciplinary referral. He submitted a written response and presented an oral argument before all five justices.
After that hearing, the justices concluded that Aarons showed a lack of remorse and concern for his client. The final order found him in direct contempt and referred the matter to New Mexico’s Disciplinary Board.
Aarons must pay $5,000 to the State Bar of New Mexico Client Protection Fund within 30 days. He must also notify the court after making the payment.
The practical consequences extend beyond the fine. Aarons cannot appear before the state Supreme Court while the disciplinary process remains pending. The court reserved the right to make further decisions after that process ends.
It appointed the Law Office of the Public Defender to provide new counsel for Sandoval. Every existing brief was struck, and a new briefing schedule will follow.
The court intends to hear the restarted appeal during its 2026–2027 term. Therefore, the client faces another round of briefing because his previous lawyer filed unverified AI output.
That procedural reset creates the article’s central tension. ChatGPT promised faster review of a long record, but unchecked output increased the workload for everyone involved.
The Stephen Aarons ChatGPT Brief Put the Client at Risk
A hallucinated answer becomes materially different when it enters a court record carrying an attorney’s signature.
Generative AI produces text by predicting plausible sequences of words. It does not independently guarantee that a statement matches a trial transcript, cited decision, or evidentiary record.
That limitation is widely described as hallucination. In this context, an AI hallucination is a confident statement that lacks support in the available legal sources.
The term can sound harmless because it originated as technical shorthand. Inside a criminal appeal, however, the result can falsely describe who testified and what those people told a jury.
A murder appeal depends on a closed record. Appellate judges review what happened in the trial court rather than conducting a new trial or hearing replacement witnesses.
A lawyer can challenge evidentiary rulings, jury instructions, constitutional issues, or the sufficiency of evidence. That lawyer cannot populate the record with additional people or rewrite testimony to strengthen an argument.
The Stephen Aarons ChatGPT brief allegedly crossed that boundary repeatedly. Its errors affected the raw materials that appellate judges use to evaluate legal claims.
The client consequences matter more than the spectacle surrounding the chatbot. Sandoval did not receive an adverse ruling on his appeal through the contempt order. Still, the filing delayed and complicated his path to review.
The court said Aarons failed to inform Sandoval about the inaccuracies. It also found that he did not tell his client about the show-cause proceedings or provide the related filings.
Those omissions deprived the client of information about a serious disruption in his representation. They also left the court to arrange replacement counsel after identifying the problem.
The public defender must now review the trial record, assess viable issues, and prepare new briefing. Prosecutors and court staff must revisit an appeal they had already begun processing.
This is how an apparent efficiency tool can transfer work instead of eliminating it. One user saves time initially, while several other participants absorb the cost of detecting and repairing unsupported output.
According to Reuters coverage, the filing appears to go beyond earlier incidents centered on fictional precedent. It inserted fabricated testimony into a criminal appeal.
That escalation explains the force of the court’s response. A false precedent can mislead a judge about the law. A false witness can mislead the judge about what happened in the case itself.
Both failures violate the expectation that counsel will check a filing. Yet factual inventions can be harder to detect because trial records are lengthy, case-specific, and often unavailable through ordinary web searches.
The opposing party might catch an incorrect quotation from a published decision quickly. Detecting invented testimony can require searching thousands of transcript pages and comparing several witness accounts.
That burden makes traceability essential. A drafting system should preserve the source behind each factual proposition and show where the supporting language appears.
Even then, source links cannot replace legal judgment. A human reviewer must confirm that the quoted passage belongs to the correct witness and supports the surrounding claim.
The same discipline applies outside law. Teams using AI to summarize technical documents need a searchable knowledge base that keeps conclusions connected to original materials.
Court submissions make the duty unusually visible because judges can impose sanctions. The underlying lesson also applies to compliance reports, medical summaries, audits, and other evidence-sensitive work.
Faster Drafting Collided With Personal Accountability
The court treated ChatGPT as a tool under supervision, while responsibility remained entirely with the lawyer who filed its output.
Aarons reportedly used ChatGPT to help summarize trial transcripts and prepare the appellate brief. That use case sounds attractive because criminal records can contain extensive testimony, exhibits, motions, and procedural history.
Summarization can help a lawyer locate themes and organize notes. It can also compress language so aggressively that qualifications, contradictions, and speaker identities disappear.
A general-purpose chatbot adds another danger. If it cannot retrieve a requested fact, it can generate a plausible substitute instead of clearly identifying the gap.
The danger rises when the output matches legal writing conventions. Invented names, quotations, and citations can arrive in polished prose that looks ready for filing.
Professional formatting is not evidence of factual reliability. Fluency only describes how natural the output sounds, not whether each proposition has record support.
Aarons’ explanation, reported before the written order appeared, centered on excessive reliance. He said he did not expect the tool to create fictitious witnesses, testimony, quotations, and authorities in coherent form.
That explanation exposes the exact failure mode. A user who expects extractive summarization may not recognize that a generative model can produce material absent from the source.
Yet professional rules place the verification duty on counsel. The model has no license, client relationship, duty of candor, or exposure to professional discipline.
The American Bar Association addressed this division in Formal Opinion 512. Its AI ethics guidance says existing duties govern lawyers who use generative systems.
Those duties include competence, confidentiality, communication, supervision, candor, and reasonable billing. Lawyers must understand a tool’s capabilities and limitations well enough to use it responsibly.
The State Bar of New Mexico had also issued specific guidance before this filing. Its ethics opinion permits lawyers to use generative AI, but only while meeting existing professional obligations.
The opinion highlights candor toward tribunals and truthfulness. It also covers confidentiality, conflicts, supervision, and the treatment of fees and costs.
Therefore, the problem was not a missing statewide rule that prohibited ChatGPT by name. The established duties already required accurate factual assertions and honest legal authorities.
The court grounded its order in those familiar responsibilities. It cited briefing rules that authorize sanctions, including contempt, when lawyers fail to comply.
It also invoked the judiciary’s authority to regulate proceedings and discipline attorneys. The stated purpose includes protecting the public, the profession’s reputation, and the orderly administration of justice.
This approach avoids making courts dependent on a particular model or software version. A lawyer remains accountable whether an error came from ChatGPT, another assistant, a search platform, or careless manual drafting.
That principle pressures both law firms and legal technology vendors. Firms need review procedures that match the risk of each task. Vendors need interfaces that discourage unsupported copying into final work.
A low-risk brainstorming session does not require the same safeguards as a factual statement in a murder appeal. The workflow should become stricter as the consequence of an error increases.
For transcript work, that means every factual claim needs a page-level source. Every quotation needs comparison against the original record. Every named person needs confirmation through the witness list or transcript.
For legal authority, reviewers must open the decision and read the relevant passage. Confirming that a case exists is insufficient when the AI has misstated its holding.
These checks reduce some of the time savings that generative AI promises. That is not evidence that the technology lacks value. It means reliable adoption requires accounting for verification work.
The core tradeoff is therefore measurable. Automation can accelerate the first draft, but the professional still owns the cost of proving that draft accurate.
Legal AI Tools Have Not Eliminated Hallucinations
Purpose-built legal systems can reduce unsupported answers, but no interface removes the reviewer’s duty to test the result.
The Stephen Aarons ChatGPT brief involved a general-purpose AI service. It did not establish that every legal AI product behaves identically or carries the same risk.
Some legal research platforms use retrieval-augmented generation. RAG is a method that connects generated answers to a selected collection of source documents.
In theory, retrieval narrows the model’s evidence and provides citations for review. It can make verification easier because the user can open the cited authority beside the generated response.
However, retrieval does not guarantee that an answer faithfully represents its sources. A model can cite a real document that does not support the proposition beside it.
Researchers at Stanford evaluated AI-assisted legal research products from LexisNexis and Thomson Reuters. Their study also compared those systems with a general-purpose GPT model.
The researchers reported hallucination rates between 17 percent and 33 percent for the specialized products they tested. The systems performed better than general-purpose models, but unsupported or misgrounded answers remained.
The legal AI study argued that hallucinations had not been solved. It called for public benchmarking and more rigorous evaluations.
Those findings do not predict the accuracy of every current product or every legal task. Models, retrieval systems, prompts, and underlying databases continue to change.
The benchmark also focused on research questions, not transcript summarization in a particular criminal case. A product’s performance on published decisions does not establish reliability on uploaded evidence.
That distinction matters when assessing ChatGPT legal errors. Research databases generally contain structured authorities with recognizable citations. Trial records often contain imperfect transcripts, abbreviations, overlapping names, and fragmented testimony.
A system can accurately locate a case while confusing two witnesses in a record. It can also quote a transcript correctly but omit nearby language that changes the meaning.
The correct skepticism runs in both directions. One court failure does not show that lawyers must abandon AI. Vendor assurances also do not justify treating generated text as verified.
The relevant question is whether a workflow makes unsupported claims easy to detect before filing. A useful system should expose uncertainty, preserve citations, and make source comparison simple.
It should also prevent a model from silently filling gaps. When the record lacks an answer, “not found” is safer than a polished reconstruction.
Firms can reinforce those controls through task design. They can separate brainstorming from factual extraction and restrict sensitive documents to approved systems.
They can require attorneys to maintain a claim ledger. Each factual statement would point to a transcript page, exhibit, docket entry, or published authority.
Supervisors can then audit high-risk claims instead of rereading every generated sentence without context. That process creates evidence that verification occurred, although documentation alone cannot prove careful review.
Training must focus on behavior, not general warnings. Telling lawyers that AI sometimes hallucinates offers little guidance during an actual filing deadline.
A practical standard asks whether another reviewer can reproduce each claim from the cited source. If the answer is no, the passage is not ready for court.
There is also a confidentiality issue separate from accuracy. Uploading a trial record into a third-party service can expose protected client information, depending on the product’s terms and controls.
The New Mexico order did not resolve that question. It focused on factual and legal misrepresentations, client communication, and Aarons’ conduct before the court.
Readers should not infer a separate confidentiality finding that the justices did not make. Nor does the order establish that ChatGPT created every false passage without user modification.
What it establishes is narrower and still serious. Aarons admitted using the tool, filing unverified output, and failing to inform his client about the resulting problems.
AI Hallucinations in Court Now Carry Predictable Consequences
The New Mexico case belongs to an established sanctions pattern, but its fabricated trial facts push that pattern into more dangerous territory.
In 2023, lawyers in Mata v. Avianca submitted nonexistent judicial decisions generated through ChatGPT. A federal judge imposed a $5,000 sanction after examining their conduct and later responses.
That case became an early warning about AI hallucinations in court. It showed that a plausible citation can survive drafting, internal review, and filing when nobody opens the underlying decision.
The New Mexico incident arrived after that warning had become widely known. It also followed national and state-level ethics guidance addressing lawyers’ use of generative AI.
Courts can therefore expect lawyers to understand the basic risk. A claim of surprise becomes less persuasive when public sanctions, bar opinions, and professional training already describe the problem.
The factual scope also changed. Mata involved fictional precedents in a civil aviation dispute. The Stephen Aarons ChatGPT brief reportedly invented participants and testimony within a murder record.
An Australian murder proceeding produced another warning in 2025. Defense submissions contained fabricated quotations and nonexistent cases, causing a delay while the court investigated.
The lawyer there apologized and accepted responsibility. The judge said independent and thorough verification was necessary before lawyers used AI-generated material.
The Australian filing demonstrates that the risk is not limited to one American court or one practice area. Similar failures appear wherever fluent output outruns source checking.
Still, sanctions are not automatic whenever a filing contains an error. Courts consider the lawyer’s conduct, response, candor, corrective steps, and applicable procedural rules.
A typographical mistake differs from presenting invented testimony as part of the record. Prompt correction differs from concealing a problem or leaving a client uninformed.
The New Mexico Supreme Court emphasized Aarons’ admissions and his response to the proceedings. Its finding concerning remorse and concern for his client helped distinguish the matter from an accidental error quickly repaired.
The order is nonprecedential, meaning the court did not issue it as a formal opinion governing later cases. That status limits how lawyers should cite it as authority.
Its operational effect remains real. Aarons faces contempt, a monetary sanction, removal from the appeal, a bar on appearing before the court, and disciplinary review.
The public record also creates a reputational consequence that can outlast the payment. Searches for the attorney’s name will connect him with fabricated witnesses in a murder appeal.
For law firms, that consequence changes the risk calculation surrounding adoption. The cost of a deficient review can include replacement counsel, duplicated work, disciplinary exposure, and damage to client trust.
For courts, repeated incidents create pressure to require certifications or disclosures. Some judges already demand statements confirming that AI-generated text and citations received human review.
Such rules can remind lawyers of their obligations, but they cannot test whether verification was meaningful. A checked box does not reveal whether anyone opened the transcript.
Technical controls face the same limitation. A system might attach a citation to every sentence while still connecting claims to irrelevant or contradictory passages.
The strongest protection combines product constraints with professional review. Software should make evidence visible, while lawyers must compare the evidence with the assertion being filed.
That combination preserves useful automation without assigning judgment to a probabilistic model. It also gives supervisors a concrete process to inspect.
The case therefore should not be reduced to “ChatGPT made something up.” Generative systems are known to produce unsupported text, especially when asked to complete missing details.
The consequential decision occurred when unverified output entered a signed court filing. The court sanctioned the human professional responsible for that transition.
What Courts, Firms, and AI Vendors Should Watch Next
Three developments will show whether this case produces durable safeguards or becomes another warning that legal workflows fail to absorb.
The first signal is the New Mexico disciplinary process. The Supreme Court referred Aarons to the Disciplinary Board, but the September order did not announce a final professional sanction.
Investigators can examine issues beyond the contempt proceeding and recommend an appropriate response under disciplinary rules. The court said it would make further determinations after that process, if proceedings occur.
A substantial additional sanction would strengthen the message that unverified AI content represents a professional-conduct failure. A limited response would suggest that contempt and removal were considered sufficient for the filing itself.
The second signal is the restarted Sandoval appeal. New counsel must enter an appearance, rebuild the briefing, and identify arguments supported by the actual trial record.
That process will reveal the practical damage caused by the false filing. The court intends to place the appeal on its 2026–2027 calendar, but a specific hearing date was not provided.
A prompt, orderly restart would reduce the harm to the client. Further delays or disputes about the record would show how long one defective AI-assisted filing can affect a criminal case.
The third signal is the response from courts and legal technology providers. Courts can adopt verification certifications, while vendors can add source-level controls for transcript and citation work.
The most meaningful product change would not be another general accuracy claim. It would be a workflow that refuses unsupported factual completion and exposes the evidence behind each answer.
Independent testing will remain important. Any vendor can demonstrate a favorable example, but users need evaluations that measure performance across unfamiliar records and adversarial questions.
Firms should also monitor whether their own review procedures work under deadline pressure. A policy that requires verification means little if lawyers lack time, training, or source access.
The immediate lesson is straightforward. AI can assist with organizing a record, outlining an argument, or locating possible issues. It cannot assume the lawyer’s duty of candor.
Every factual statement in a court filing needs a path back to admissible material. Every legal proposition needs support from an authority that says what the filing claims.
That standard is demanding because appellate work has always been demanding. Generative AI changes the drafting interface, but it does not lower the burden of accuracy.
The Stephen Aarons ChatGPT brief turned a promised shortcut into contempt, replacement counsel, and an entirely new briefing cycle. Legal professionals now have another documented warning, this time involving invented witnesses rather than only invented cases.
The next question is whether organizations will redesign their workflows before another client absorbs the consequences. Lawyers using AI should test one thing before filing: can every important sentence be proven from the original record?



