New Mexico Supreme Court Holds AI-Using Lawyer in Contempt
On a Wednesday in September 2026, the New Mexico Supreme Court issued an order that every attorney in the United States should read twice. The state's highest court found that attorney Stephen Aarons — a criminal defense lawyer with more than four decades of experience — had submitted a brief to the court containing fabricated testimony attributed to wholly invented witnesses, including fake police testimony. The source of these inventions was not a dishonest client or a rogue paralegal. It was ChatGPT.
The court held Aarons in direct contempt and referred him to the state's attorney disciplinary board for further proceedings. In its order, the court found that Aarons "demonstrated a lack of remorse and a lack of concern for his client" — a phrase that carries enormous weight when it comes from a state's highest judicial authority evaluating the fitness of a licensed attorney. Aarons admitted that he did not verify the factual claims and legal authority in his AI-generated brief before signing it and filing it, and that he never disclosed this failure to his client.
The AI hallucination legal consequences in this case are not abstract. A man sits in prison for life. A lawyer of forty years faces potential discipline. And a state supreme court has now established, in writing, exactly what happens when a lawyer treats an AI chatbot as a reliable co-counsel.
The Case Behind the Ruling: A Murder Conviction Appeal
The stakes in the underlying matter were as serious as criminal law gets. Oscar Renee Sandoval was convicted of killing Shiereen Al-Jibury, his partner and the mother of his children. In February 2025, he was sentenced to life in prison. His family subsequently retained Aarons to appeal the conviction — a final, desperate attempt to overturn a verdict that would effectively end Sandoval's life as a free person.
Appeals in murder cases demand precise factual accuracy. Courts review trial records line by line. Every evidentiary claim, every citation to testimony, every procedural argument must be grounded in what actually happened in the courtroom below. Fabricated witnesses are not merely unhelpful — they are a category of fraud upon the court, an affront to the integrity of the entire appellate process.
Aarons apparently used ChatGPT to generate or assist with the brief and then, by his own admission, submitted it without verification. The brief contained false testimony attributed to witnesses who did not exist. It is not yet publicly detailed how many fabricated witnesses appeared in the filing, but the court's characterization — "wholly fabricated" — leaves no ambiguity about their nature. This was not a citation error or a misremembered page number. These were invented people saying invented things in a document filed in the highest court of New Mexico on behalf of a man serving a life sentence.
Why This Ruling Sets a Legal Precedent for AI Use in Law
The legal significance of this order extends well beyond Stephen Aarons or New Mexico. To understand why, it helps to distinguish between the two consequences the court imposed: direct contempt of court and referral to a disciplinary board.
Direct contempt — specifically, contempt committed in the presence of or directly affecting the court — is a judicial finding that a party's conduct obstructed or degraded the administration of justice. It carries its own immediate consequences, including potential fines or sanctions imposed by the court itself. This is a judicial remedy, swift and self-executing.
A disciplinary board referral is categorically different in its long-term implications. Bar disciplinary proceedings evaluate whether an attorney has violated the professional rules of conduct governing licensure. A finding of professional misconduct can result in public censure, suspension, or disbarment — the permanent loss of the right to practice law. The New Mexico Supreme Court did not simply punish Aarons for one bad brief; it initiated a process that could end his career entirely.
Together, these two actions signal to the legal profession that AI-related misconduct will be treated as a compound offense: a harm to the court and a breach of professional duty. No prior case had paired both responses so explicitly in a single order.
The Growing Problem of AI Hallucinations in Legal Filings
The Aarons ruling did not arrive in a vacuum. It is the most consequential entry in a pattern of AI hallucination legal consequences that courts have been documenting since large language models became accessible to the general public.
The most prominent earlier case was Mata v. Avianca, decided in the Southern District of New York in 2023. In that matter, attorney Steven Schwartz used ChatGPT to research case law for a personal injury filing and submitted citations to cases that did not exist — complete with fabricated quotations from fictional judicial opinions. Judge P. Kevin Castel sanctioned Schwartz and the supervising attorneys at his firm, finding that their conduct warranted a $5,000 penalty and letters to the judges whose names had been falsely attached to invented opinions.
What made the Mata case a landmark was the transparency it forced: ChatGPT was shown, in open court, confidently inventing legal precedent from nothing. But even Mata involved fabricated citations to nonexistent cases. The New Mexico case escalates that harm considerably by introducing fabricated human witnesses — testimony attributed to people who never existed, in a matter where the actual human record was the entire contested terrain.
Research from institutions including Stanford's CodeX center and various law review surveys conducted after 2023 has consistently demonstrated that AI language models hallucinate at meaningful rates when generating legal content. Studies have found fabricated citations appearing in AI-generated legal research at rates that, even conservatively estimated, would be professionally catastrophic if a law firm relied on AI output without independent verification. Large language models are designed to produce fluent, confident text — not accurate text — and in legal drafting, that distinction is the difference between competent representation and fraud.
Bar Associations and Courts Respond: Emerging Rules on AI Verification
The legal profession has not been passive in the face of these failures. In 2024, the American Bar Association issued Formal Opinion 512, which addressed a lawyer's obligations when using generative AI tools. The opinion made clear that existing professional responsibility rules — particularly those governing competence under Model Rule 1.1 and candor toward the tribunal under Model Rule 3.3 — fully apply to AI-assisted work product. The ABA did not create a special AI exception to professional duty. It affirmed that no such exception exists.
Rule 1.1's competence requirement, the ABA's opinion emphasized, encompasses the obligation to understand the tools being used. A lawyer who submits AI-generated content without understanding that such content can and does fabricate authority is not acting competently, full stop. Rule 3.3's candor requirement means that a lawyer who knowingly files false statements of fact or law — regardless of how those false statements were generated — violates their core duty to the tribunal.
Multiple state bars have since issued their own guidance in the same vein. The California State Bar's 2023 guidance, and subsequent opinions from state bars in New York, Florida, and elsewhere, have converged on a consistent principle: AI tools may assist legal work, but the supervising attorney bears full professional responsibility for everything filed under their name.
The New Mexico Supreme Court's contempt finding is now the most severe judicial enforcement of exactly this principle.
What Lawyers Must Do Before Submitting AI-Generated Legal Work
The Aarons case distills to a simple professional catastrophe: a lawyer with forty years of experience signed a document he did not read carefully enough to catch invented witnesses. That failure — at that experience level, in that high-stakes context — represents a cautionary case study that bar examiners and legal ethics professors will be citing for years.
Several concrete practices separate competent AI use from the conduct that landed Aarons in contempt.
First, every factual claim in an AI-assisted brief must be traced to a primary source. This means checking court transcripts, police reports, and trial exhibits directly — not accepting AI summaries of those documents. In the Aarons case, fake police testimony would have collapsed immediately against the actual trial record.
Second, every legal citation must be verified through a recognized legal research database such as Westlaw or LexisNexis. ChatGPT and similar tools cannot reliably distinguish real cases from plausible-sounding fictional ones. A two-minute search in a proper database would catch a fabricated citation every time.
Third, disclosure to the client is not optional when AI tools are used in a manner that creates material risk to the representation. Aarons admitted he did not inform Sandoval of his reliance on unverified AI output. The New Mexico Supreme Court treated that silence as evidence of a failure of loyalty, compounding the original filing error into a broader breach of fiduciary duty.
Fourth, signing a document filed with a court is a certification — under existing rules governing attorney conduct — that the contents are accurate to the best of the signer's knowledge. The standard is not "generated by a competent AI"; it is "verified by a competent lawyer." These are not the same thing, and no court has suggested they ever will be.
The New Mexico ruling does not prohibit AI in legal practice. It enforces what the rules have always required: that a lawyer who signs a filing owns everything in it. The tool used to draft it is irrelevant. The responsibility is not.
Source: [Ars Technica - All content](https://arstechnica.com/tech-policy/2026/09/chatgpt-using-lawyer-punished-for-citing-fake-testimony-from-made-up-witnesses/)