Pakistan's JudgeGPT Trial Raised Case Resolution 6.3 Percent
Researchers found judges using JudgeGPT resolved 6.3 percent more cases with no clear drop in quality — the first large-scale evidence on generative AI in courts.

Updated
Why it matters
- JudgeGPT increased resolved cases by 6.3 percent across 1,559 Pakistani trial judges with no measured drop in judgment quality.
- The tool combines GPT-4 with retrieval-augmented generation over 128,292 judicial opinions and 943 statutes.
- Trained judges resolved 38.5 more cases per month, yielding roughly $38.50 in savings per dollar spent on the tool.
Judges in Pakistan who used a custom GPT-4-based research assistant resolved 6.3 percent more cases, with no obvious drop in the quality of their judgments, according to a nationwide field experiment covering 1,559 trial judges — roughly half the country's judiciary. The study, by economist Sultan Mehmood of the New Economic School in Moscow and collaborators including ETH Zurich's Elliott Ash, is the first major independent assessment of ongoing judicial use of generative AI.
The stakes for Pakistan are concrete. The country's courts carry a backlog of 2.26 million cases, with fewer than two judges per 100,000 people — compared with 22 in the EU and eight in Brazil. The researchers built their tool, JudgeGPT, in consultation with the judiciary and began offering it to judges in 2024.
"We do find an increase in cases resolved, and we don't find any corresponding decrease in decision quality," Mehmood says.
AI tools for judges are already being deployed in Brazil and India, and University of Chicago law professor Eric Posner has compared LLM judgments to human judgments in a single case study. But no prior study has evaluated live, at-scale use of generative AI by sitting judges. That gap matters well beyond Pakistan: judges elsewhere have made headlines for illicitly using chatbots in drafting rulings, including error-ridden decisions that drew scrutiny from U.S. Senator Chuck Grassley.
A grounded tool, not a raw chatbot
Mehmood says Pakistani judges were enthusiastic from the start. "They were more techno-optimist than we were," he says. "The delays are so huge, this is something which they thought was worth trying anyway to reduce people's suffering."
Some judges were already using commercial chatbots, which performed poorly on Pakistani legal queries and frequently hallucinated case law. The team's answer was retrieval-augmented generation. JudgeGPT queries a knowledge base of 128,292 Pakistani judicial opinions and 943 statutes, and every response includes footnotes linking to the underlying cases and laws.
"It turns out that actually the way to fix [hallucinations] isn't just more intelligent models," says Ash, an associate professor of law, economics, and data science at ETH Zurich. "It's to attach the models to a tool that can do a search and verify the sources." The researchers do not report hallucination rates.
Training drove adoption
The tool alone was not enough. The team put 1,197 judges through six 90-minute Zoom training sessions, developed with Pakistan's Federal Judicial Academy, covering how LLMs work, their limitations, bias and hallucination risks, and the need to verify outputs. Another 180 judges received only generic technology training; a final group got none.
The difference showed in usage. JudgeGPT-trained judges logged in a median 56 times and sent 212 prompts over the study period, versus 10 logins and 25 prompts after generic training. Untrained judges typically used the tool for about a month and then stopped. "Just giving people the technology does not necessarily make them use it persistently," Mehmood says.
By the time 487 judges had completed the program, the median district saw the 6.3 percent rise in resolved cases, and districts with more trained judges saw larger effects. Appeal rates fell slightly, suggesting faster resolution was not producing sloppier decisions.
MIT economics professor David Autor calls the experiment's scale remarkable. "It's pretty amazing that he's able to pull this off," Autor says. "It's not easy to do large-scale field experiments in civil service, but especially where the stakes are so high." He describes the 6.3 percent gain as credible, if not overwhelming, and likely to improve with wider use.
Measuring quality — with an LLM
Evaluating thousands of judgments with legal experts was infeasible, so the team asked OpenAI's GPT-5-mini to compare pairs of judgments from the same judge before and after training. The model preferred post-training judgments 59 percent of the time. Two experienced Pakistani lawyers reviewed the model's analysis of 90 judgment pairs and agreed with GPT-5-mini 70.6 percent of the time — nearly matching their 73 percent agreement with each other.
The economics are notable. A trained judge resolved 38.5 more cases per month than the baseline, which the researchers calculate translates to roughly US $38.50 saved in judicial costs for every dollar spent running the tool. Ash cautions that these figures cover the first nine months of the trial, and the team has since updated both the underlying model and the database.
One participating trial judge, speaking on condition of anonymity, said their caseload has not dropped below 1,000 cases in more than a decade. "For research, it's just one prompt away, whereas before I had to search for the precedents and laws for hours," the judge says. "If I have to read 10 pages of a precedent, now I ask JudgeGPT to just summarize it for me and give me the crux, and it does that work in seconds."
The open question: quality of justice
Efficiency is not the only measure of a justice system, says John Zeleznikow, professor of law and technology at La Trobe University in Australia. "What they've tried to do is be effective, [to] deal with more cases more quickly, and they're able to do that," he says. "What's not that clear is whether what you call the quality of justice is better."
The data bears out his caution. Roughly a fifth of participants' prompts involved what the authors call "substantial AI delegation" — asking the tool for the best decision, legal reasoning, or drafted opinions with little judicial input. Training reduced the proportion of such inappropriate delegation but did not eliminate it.
Ash argues the alternative — unmanaged AI use — is worse. "There are risks for using these AIs, for sure, even with all these safeguards. But at some point you have to just put the judges in as strong a position as you can," he says. "Have technological safeguards, but then try to encourage the judges not to rely on it too much."
With judicial AI rollouts already underway in Brazil and India and no comparable evaluations published, the Pakistan experiment now serves as the reference point for what grounded tools plus training can — and cannot — deliver in courts.
Original: bbc.com
More from Rebecca Stone
Show full bio
Correspondent covering consumer brands and retail at AI In Context.
135 articles