In a country where millions of cases sit in court backlogs for years, an AI tool is quietly making waves in Pakistan's judiciary — and the results might change how we think about AI in the public sector.
A new field experiment by researchers from ETH Zurich, the New Economic School, and Imperial College London covered roughly half of all Pakistani trial court judges: 1,559 across 118 courts. They deployed JudgeGPT, an AI assistant built on OpenAI's GPT-4, designed specifically for Pakistani legal proceedings.
The results? Districts with more trained judges resolved 1,848 extra cases per year, a 6.3% bump. The estimated return on investment: $38.50 saved for every dollar invested — based on what it would cost to hire enough additional judges to match the same output.
AI Alone Doesn't Work — Training Does
Here's the critical finding: giving judges AI access did almost nothing by itself.
The researchers split participants into three groups. One got JudgeGPT plus targeted training — six 90-minute lectures over three weeks taught by ETH Professor Elliott Ash after court hours. A second group got AI access but only a general seminar on technology and law. A control group got the seminar but no AI access.
The difference was stark. Judges with targeted training used JudgeGPT four times as much as those in the general seminar group. After 40 weeks, trained judges averaged nearly 60 logins and over 200 prompts. The comparison group averaged about 20 logins and fewer than 50 prompts.
This is a message for every organization deploying AI: technology without training is just expensive wallpaper.
What Judges Actually Used It For
Looking at anonymized chat logs from about 1,500 judges, the most common tasks were legal research, text editing, and text generation. Around 60% of queries sought information about laws, procedures, or legal concepts.
Trained judges shifted their usage patterns significantly — they asked fewer broad legal questions (where hallucination risk is higher) and more editing and summarizing tasks (where language models are more reliable). Only about 20% of requests involved what researchers called "substantive AI delegation," where judges asked JudgeGPT to evaluate decisions or draft reasoning on its own.
Training actually made judges more likely to decide cases themselves and use AI only to write up their reasoning. This is responsible AI use — the human stays in the loop, AI handles the drudgery.
Judgment Quality Held Steady (or Improved)
Perhaps the biggest concern with any AI in the courtroom: does it make rulings worse?
The answer: no. Appeal rates per 1,000 resolved cases fell slightly. An LLM-based quality check, validated by two Pakistani lawyers, showed a slight improvement. Rulings from trained judges were rated better in 59% of pairwise comparisons, up from 42% in the control group.
Readability, length, and number of legal arguments held steady. Critically, no evidence of increased gender or religious bias in judicial language.
Work-life balance didn't change either — judges worked the same hours. The AI didn't make them work faster; it made them work smarter within existing constraints.
The Real Story: It Was Built on GPT-4
Here's something that should give everyone pause: JudgeGPT ran on GPT-4, a pre-reasoning model that was the best available at the time. Today's reasoning models write better, hallucinate less, and handle complex tasks more reliably.
The researchers stress their findings represent a floor, not a ceiling. If GPT-4 can deliver these results, what happens when you deploy it on Claude Opus 5 or GPT-5.6 Sol?
Why This Matters for the Global South
Pakistan has one of the world's largest court backlogs. According to the World Justice Project, the country ranks near the bottom on access to justice, with cases taking years to resolve. The AI tool doesn't replace judges — it gives them superpowers within the existing system.
The $38.50-per-dollar return figure is conservative. Even the most conservative estimates put ROI at "at least" $10 per dollar. Compare this to the billions being spent on frontier model development in Silicon Valley, and you start to see a different picture of AI's real-world impact.
This isn't about AGI or autonomous agents or the singularity. This is about a judge in Lahore using a tool to process one more case before going home to their family. Small, incremental, life-changing.
🔥 Hot Takes
1. The best AI deployment strategy isn't technical — it's pedagogical. Every company spending millions on AI rollout should read this paper. The technology was the easy part; getting judges to use it correctly required 540 minutes of targeted training. That's the lesson nobody in Silicon Valley wants to hear.
2. GPT-4 in Pakistan courts > frontier models in Silicon Valley boardrooms. While American tech companies debate whether AI will destroy civilization, a GPT-4 wrapper is helping Pakistani judges clear a backlog that's kept citizens waiting decades for justice. Sometimes the most impactful AI applications are the boring ones.
3. This is what "AI for good" actually looks like. Not viral demos or CEO keynotes — but a randomized controlled trial with 1,559 judges, transparent methodology, and results that hold up under scrutiny. If we want to measure AI's real impact, look at what it does in places where the stakes are highest.
Sources: The Decoder, Mehmood, Goessmann & Ash (ETH Zurich), International Journal