Second Chair
§ Families · the corpus-level reading

What this literature argues

The 106 topic clusters were grouped bottom-up into 20 themes and 5 families. Nobody chose those divisions in advance — they are what the 600 sources fell into. This page reads each family for its argument. Cluster names open the cluster; slugs like open the source. When you want to go further down, the explorer holds all 106.

Family 01 · 53 topics · 554 sources

The gains are real, bounded, and the constraint is absorption

The largest family. Generative AI reliably compresses the production of legal text while judgment, accountability and client context stay with lawyers — so value depends on redesigning how work is organised rather than on buying a better model.

puts the mechanism most compactly in : the bottleneck has moved, and it is no longer intelligence. What becomes scarce is the person who can direct the machine, test the answer, understand what the client is actually trying to achieve, and make the call when the answer is hard.

A proprietary platform is worth exactly as much as the changed practice it is wired into.

The experimental record is unusually good, and it does not say one simple thing. found GPT-4 “only slightly and inconsistently improved the quality of participants' legal analysis but induced large and consistent increases in speed”. The successor trial, , found retrieval and reasoning tools produced gains “of anywhere from 50% to 130%” across five of six tasks and, unlike the earlier work, real quality gains — but concentrated in litigation-shaped work: they “do not appear to extend to the one transactionally oriented task we evaluate, which involved drafting a short contract”.

Two results complicate any headline. Access alone does little — found that “access to an LLM did not noticeably improve performance on the examination absent user training”. And the benefit is not evenly distributed within a person's own work: found weak first drafts improved after AI revision, while strong ones got worse.

Participants whose initial memos were relatively weak produced stronger memos after revising with AI. Surprisingly, however, participants whose initial memos were relatively strong produced weaker memos.

The error mode is subtle, not spectacular

makes the family's strongest claim: the danger is not the invented case but the plausible mis-statement. measured hallucination “between 58% (ChatGPT 4) and 88% (Llama 2)” on direct, verifiable questions about randomly selected federal cases, with models “systematically overestimat[ing] their confidence relative to their actual rate of hallucination”. Grounding helps and does not cure: found the retrieval-based commercial tools “each hallucinate between 17% and 33% of the time”. Even citation formatting is unreliable — reports fully compliant citations only 69–74% of the time.

These errors are potentially more dangerous than fabricating a case outright, because they are subtler and more difficult to spot.

From which follows the family's most interesting open question, put by : in a text-dominated profession, the more AI is used the more there is to check, so the net value of the tool is the efficiency gain minus the verification cost. concedes that “expecting manual verification of every AI-generated assertion is unrealistic at scale”, and supplies the sting: lawyers could not reliably tell when AI had helped them. Nobody in this corpus has measured the verification side of that equation.

The courts did not wait for new rules

shows accountability arriving through ordinary doctrine. : “there is nothing inherently improper about using a reliable artificial intelligence tool for assistance. But existing rules impose a gatekeeping role on attorneys.” imposed $15,000 per attorney plus fees and double costs, holding that citing even a single fake case can be sanctionable; extended it to supervision, faulting a supervising attorney for failing to keep abreast of the technology. makes the standard use-specific rather than categorical, and doubts more is needed at all: “I'm not sure that AI-specific bans or certifications add much (if anything) to Rule 11.”

Where it genuinely divides

On billing (), frames the tension — “By every measure of productivity, AI has been a success in legal. By the logic of the billable hour, it is a slow-motion disaster” — and reports “nearly two-thirds (64%) expect to rely less on outside counsel”. The counter-argument runs the causation the other way: holds that “the billable hour persists because it solves problems no other pricing model can”, and that efficiency raises expectations of insight and responsiveness rather than lowering bills. splits it: the hour survives for work whose scope genuinely cannot be fixed in advance.

On access (), grants that “there are simply not enough attorneys to deliver legal services at the scale required” and then inverts the conclusion — as the cost of starting a claim approaches zero, the likely result is “more disputes and more assembly-line litigation against low- and middle-income people”, with the wealthy still holding superior access. offers the counter-standard: the test is not whether the tool matches the best available lawyer, but whether it beats the alternative, which for most people is nothing.

Open , and first — the second and third are where the corpus argues with itself rather than agreeing with itself.

Family 02 · 18 topics · 261 sources

The task is the unit, and verification is the frontier

The labour-economics branch. Its argument is that “lawyer” is the wrong thing to measure, and that what makes work automatable is not how routine it is but how cheaply its output can be checked.

states the premise: a role is a bundle of tasks, some automatable and some not, so “the scope of what computers can do better than humans does not map neatly onto existing human roles”. then advances the sharper claim, from , which displaces the older routine/non-routine frontier that worked with.

The critical delineation is no longer routine versus non-routine, but measurable versus non-measurable… tasks become automatable not when they are simple, but when they can be measured — creating a risk zone where execution is cheap but verification remains tacit and expensive.

Legal work sits squarely in that zone. Drafting is cheap; establishing that a draft is accurate, well-grounded and professionally suitable is not. This is a model rather than a measurement, and worth reading as one.

Levelling, or a new divide?

organises the corpus's liveliest empirical dispute. found productivity up “14% on average, including a 34% improvement for novice and low-skilled workers but with minimal impact on experienced and highly skilled workers”, and legal work replicated the pattern. Against that, reports that gains “were not predicted by GPA or prior knowledge, but by AI Interaction Competence — the ability to elicit, filter, and verify model outputs”, with low-competence users seeing “limited or even negative marginal returns”. That is a new stratification, not a levelled field. rejects the framing entirely: what matters is whether removing tasks raises or lowers the expertise required for the tasks that remain.

Realised labour markets split by setting. finds freelancers in exposed occupations losing both employment and earnings. uses Danish administrative data to “estimate precise null effects… ruling out effects larger than 2% two years after ChatGPT's launch” — while insisting that is not nothing.

The absence of measurable labor market effects is not evidence that nothing is happening—but rather that transformation is occurring on margins that conventional economic statistics do not capture.

The nearest professional analogue is auditing: associates AI investment with “a 5% reduction in the likelihood of restatements” alongside an 11% decade decline in accounting employees — better work, fewer juniors — concluding that partners benefit while junior employees bear the displacement.

Why supervision is not the answer

supplies the oldest argument here, and the one most often skipped. observed that skills decay when unused, so the person monitoring an automated process becomes progressively less able to take over — “yet manual takeover is needed precisely when something is wrong”. names the accountability trap: reviewers become “moral crumple zones”, unable to exercise control but bearing the blame for failures.

The most uncomfortable result comes from radiology. found that AI assistance “does not improve radiologists' average diagnostic quality, even though its predictions are more accurate than 78% of them”, because clinicians underweight the model and treat its output as independent of their own reasoning.

The optimal human-AI collaboration design delegates cases either to humans or to AI, but rarely to AI assisted humans.

If that transfers to legal work — and it is medicine, so it may not — then “human-in-the-loop” is not automatically the safe design, and choosing which matters go to which is the real decision.

Aggregate stability, individual harm

offers the demand-side counterweight: projects “a 33% reduction in hours worked by radiologists in 5 years” yet concludes productivity gains will be offset by growth in imaging volume. And supplies the distributional warning that should govern how any headline figure is read: when telephone switching was mechanised, later cohorts of young women were unaffected, but incumbents were roughly 40% more likely to change careers and older operators measurably less likely to be working at all.

Family 03 · 11 topics · 37 sources

General signals cannot certify particular legal work

Every cheap proxy for quality — a benchmark score, a price, a disclosure, a star rating, a vendor's contractual label — measures something adjacent to legal quality rather than legal quality itself. Followed through, this is the most policy-relevant branch of the corpus.

The intellectual core is credence-goods economics, and its premise predates AI by fifty years. defines the category as qualities “consumers cannot evaluate through normal use”, and identifies the structural hazard — when the seller diagnoses the need and also supplies the remedy, costly discovery “can permit a positive amount of fraud even under competition”. That is a description of a law firm.

separates two things the market conflates. In the experimental market buyers can see every seller's posted prices, yet still cannot learn which service they actually need. Price competition can be complete while diagnostic competence stays invisible — so publishing fixed fees for AI-assisted matters disciplines price without telling a client whether the diagnosis was right.

then attacks the assumption most legal-AI policy rests on: that more information protects clients. reports that “disclosure can backfire, by leading to more biased advice and at the same time reducing the likelihood that such advice is turned down by consumers”. And turns it into a design result that should worry anyone building intake or triage tools.

The expert will optimally choose imperfect diagnostic information even when information acquisition is costless.

Reputation is no substitute either: finds reviews of credence providers track “experience attributes, such as promptness, which consumers can typically evaluate, rather than credence attributes, such as knowledge”. A five-star legal AI product is evidence about its interface.

What does work: liability, not verifiability

carries the family's most consequential result, and it is experimental rather than argued. crossed liability against verifiability and found the two are not equivalent: efficiency runs 0.13–0.27 across the no-liability conditions and 0.72–0.96 with liability, and undertreatment — 0.53–0.73 without — disappears.

While theory predicts that either liability or verifiability yields efficiency, we find that liability has a crucial, but verifiability only a minor effect.

Almost every instrument currently proposed for legal AI — model cards, audit trails, explainability, disclosure that AI was used — is a verifiability instrument. This evidence says the binding problem is inadequate service, and that enforceable responsibility for outcomes is what fixes it. Competition does not substitute: seller competition “drives down prices and yields maximal trade, but does not lead to higher efficiency as long as liability is violated”, which contradicts the hope that cheap AI-delivered legal help will discipline itself through market entry. Two honest limits: liability does not touch overtreatment or overcharging, and these are laboratory markets in abstract goods, so the application to legal services is structural analogy.

Scores that do not decompose

supplies the measurement half. In , on identical contract-clause tasks one model reaches 97.6 and 91.6 while others sit at coin-flip — 50.0, 46.8, 49.3. On merger tasks the ordering scrambles: GPT-4 at 40.1 against Claude-1 at 0.0 on one question, then GPT-4 at 15.4 against GPT-3.5 at 31.9 on the next. “Capable at legal work” is therefore not a property a model has; only model–task pairs have it, and a firm procuring on an aggregate score is buying a number that does not break down into the work it actually does. The scores are not even commensurable with each other — some questions carry roughly 700 words of source text, others 122, others 25, and each counts once.

Where the risk actually lands

holds the family's most practical finding, and it comes from reading two documents against each other. tells buyers plainly that the vendor “is not an insurance company” and that the customer “should look solely to its insurance or self-insurance programs”, with liability capped at twelve months of fees and consequential loss excluded. Turn to the insurance side, and — catalogued as carrying a generative-AI exclusion — contains no such operative wording. What it does exclude is disclosure of confidential or personal information, loss or corruption of electronic data, failures of authentication and IT security, and “fraudulent instructions transmitted by electronic means, including through ‘social engineering fraud’”, regardless of fault or intent. Those are precisely the AI-era incidents a firm should fear. The vendor points at insurance; the policy points away; the residual sits with the firm.

extends the warning to ownership. The concentration mechanism is documented — private-equity-acquired physician sites rose from 816 across 119 metros in 2012 to 5,779 across 307 by 2021, and, as puts it, “because each individual acquisition is small in scope, consolidation does not trigger scrutiny from antitrust regulators”. Prices move reliably (colonoscopy +4.5%, dental charges +3.3%); quality does not move reliably in either direction, with finding no effect across six measures. All of it is healthcare, so its bearing on AI legal services is predictive analogy rather than evidence.

Family 04 · 21 topics · 19 sources

Exposure is not displacement, and the instruments disagree by construction

This family holds one genuine methodological dispute about how automation risk is measured, wrapped in the claims of a single large employer survey. Read the first for the argument; read the second as expectation rather than observation.

is the substantive cluster, and its point is that the instruments people cite when arguing about whether legal work is automatable are not measuring the same thing, so their apparent conflict is largely definitional. scores “Title Examiners, Abstractors, and Searchers” at 0.99 — a probability of technical computerisation for an occupation as classified, not a forecast of employment, and extrapolated from a hand-labelled set of just 70 occupations. measures exposure instead, and splits legal work across its whole range: legal secretaries high, legal associate professionals low, judges “Not Exposed”. Exposure there tracks clerical and text-processing content, not the sector label. measures something else again — observed use — and records lawyers at 0.62%, below web developers.

The cluster's conclusion is the useful part: an occupation can be highly computerisable under one framework, moderately exposed under another, and already using AI at a modest rate, without any of those findings contradicting each other. Comparisons have to align the occupational definition, period, technology and measured outcome first. Its practical instruction is to assess research, document review, drafting, title examination, secretarial work, intake, negotiation and advocacy separately, and to track actual use, error rates and supervision needs rather than treating an occupational score as a forecast.

What the employer survey claims

The rest of the family reads , a 2024 survey of employers published in 2025 — and it is employer expectation, not measurement. Its headline is structural churn of 22% of today's jobs by 2030: 170 million created, 92 million displaced, net 78 million. The task split moves from “48% of tasks performed predominantly by people, 30% by a combination and 22% predominantly by technology” toward an even three-way division. Analytical thinking remains the most sought-after core skill even as AI and big data grow fastest.

The distributional finding is the sharpest thing here, and and build on it: of a representative 100 workers, 41 need no training, 29 are upskilled in role, 19 are upskilled and redeployed — and 11 are unlikely to receive the training they need. Unequal access to training, rather than exposure, may decide who is augmented and who is displaced. notes the unresolved bill: employers expect to fund their own programmes while also ranking public reskilling funding as their top policy ask.

draws the conclusion the family is best used for — that outcomes turn on employer implementation choices rather than on the technology, with 41% of employers anticipating headcount reductions where AI can replicate human work and 85% prioritising upskilling. Two of its own clusters state the limit plainly: this is “economy- and industry-wide rather than legal-sector-specific”, and records “employer expectations, not observed outcomes”.

Family 05 · 3 topics · 6 sources

AI acts on the one rung where the profession's composition changed

The smallest family, and a methodological caveat with real force behind it: workforce effects have to be disaggregated by seniority, firm size and place, because the national average describes almost nobody.

shows where the profession's gains actually sit, and they sit at entry. Associates of colour reached a record 31.46% and women became a majority of associates at 51.62% — against 12.73% partners of colour and 28.83% women partners, a gap NALP estimates would take nearly 27 years to close at current rates. Summer associates of colour hit a record 43.07%, while Black summer associates fell 1.5 points to 10.24%.

Set that beside the apprenticeship argument in family 02 and the implication is direct: if AI compresses junior document, research and drafting work, it operates on precisely the junior and summer associate roles where representation has been won. is where that reading lives.

The aggregate also conceals enormous variation (, ). Firms over 1,000 lawyers report 30.54% women partners and 15.31% partners of colour; firms of 250 or fewer report under 10% partners of colour. San Francisco leads on women partners at 36.66%, while Silicon Valley and Miami both exceed 25% partners of colour and both have majority-of-colour associate populations. A national figure of 12.73% describes almost none of these places — so a claim that AI will raise or erode opportunity “in the profession” is a claim about a population that does not exist in aggregate.

Two older sources supply the mechanism. argues that measurement itself destabilised the promotion tournament — billing systems and performance metrics made individual contributions and rival firms' results legible, driving comparison-shopping, lateral mobility and internal scrutiny. And found that computerisation changed the task content of jobs whose titles never changed, with a large share of its effect on demand for educated labour occurring inside nominally unchanged occupations. Measuring AI's effect on law by counting lawyers will therefore miss most of it — which is worth holding against any employment projection cited as reassurance.

Note on this family: its three clusters rest on table data extracted from PDF appendices, and the quoted figures above come from those sources' own reporting rather than from the extracted rows. Treated as a caveat on the workforce reading it is sound; treated as a standalone branch of the literature it is the weakest of the five.

Going further

These five readings name perhaps thirty of the corpus's topic clusters. The rest are one click away.

106
topic clusters, each with its own proposition
20
themes between the clusters and the families
1,654
source-level stances, each with its own framing and quotes

Positions in this corpus are recorded against each cluster's own proposition, so they measure stance toward a framing rather than agreement in the abstract — and comparisons between voices or eras carry more than any single number does.

Open the explorer → · How the reading was produced →