How the record was read
This app reads a single cloud database built by the insight-bridge pipeline from a curated corpus on AI and the legal profession, assembled by David B. Wilkins and Anthea Roberts for the AI and the Legal Profession research program of the Center on the Legal Profession at Harvard Law School: 1,501 sources, 47,381 passages, 61,333 source–topic memberships, 9,642 source-level stances and 1,947 lens perspectives. Everything on every page is a live query against Cloudflare D1 (the relational residual) and Vectorize (the embeddings), through fully-typed Drizzle. The same data is exposed to agents over an MCP server at /mcp.
The pipeline
- Curation & ingest
The sources — scholarship, industry evidence, primary legal materials and practitioner interview transcripts, 1949–2026 — are collected by hand, normalised into one document each, chunked into 400-token passages and embedded. Cross-domain, historical and theoretical sources (other professions, medicine, economics, sociology) are included deliberately as analogy and framework evidence. This is a hand-built evidence base, not a sampling frame: the selection is the first and largest analytical choice in the whole pipeline.
- Facets — the two lenses and the curator’s vocabularies
Each source carries the curator’s coding from the corpus manifest: a voice (who is speaking, ten values), a source class and type, jurisdiction, evidence type and topic dimension (the last three multi-valued), year, publication, author and url. The era — five ordinal buckets over publication year — is derived from the year with the run profile’s cut-points. Voice answers who is speaking; era answers when. Every comparative view in this app is parameterised by one or the other. The curator’s own working facets (which thesis paper a source feeds, reading priority) are not carried.
- Per-source extraction
Each document is read end-to-end by a language model against this run’s profile: key points with verbatim supporting quotes, every quote checked against the source text; and a claims analysis under four fixed subtopics — the concrete claim or prediction, the evidence basis offered for it, the stated or implied timeframe and conditions, and who is affected and how. Views are attributed to the author or interviewed speaker, not to hosts or people they merely quote.
- Embedding & topic clustering
Topics are found from the passages themselves, bottom-up, by grouping the passages across the corpus that say the same thing; nothing supplies a list of themes beforehand. A topic has to be raised by more than one source to form: one source's volume cannot make a topic on its own, and the rare topic that only one source carries is flagged as such. Each passage, and each source, is then placed against every topic at a graded strength (exemplar, high-value, member), so a source that touches a topic in passing is recorded as well as the sources that define it, and a passage can belong to more than one topic. Each document is assigned its dominant cluster. An advisory junk review flagged clusters whose syntheses reported no substantive shared content; a reviewer confirmed them, and they are excluded from every aggregate here.
- Perspective synthesis
Each topic is written up as a proposition with key points, and every exemplar and high-value source's position is recorded against that proposition, from Supports to Opposes. Positions are relative to the topic's own framing, never absolute agreement, and every source is measured against the same proposition. Each position carries its framing, analysis, key points and quotes. The same is done per voice and per era, giving the two lenses. Every quotation is verbatim from the source document: each is verified against the document's text before it is stored, and one that cannot be found there is never presented as verified.
- The grouping tree
The substantive clusters are grouped upward, generation by generation, into named groups — themes, then families in this run. The app reads the tree as the pipeline left it, however many generations it has; nothing above the topics is chosen by hand.
- Preparation for this app
Outside the pipeline, before loading: the facet table was flattened to columns and the era derived; every quote the pipeline could not verify was re-checked with a rule that reads across page furniture; the chunk text and vectors of restricted-distribution sources were withheld; the tree was flattened per generation; and the entity review’s suggested duplicate groups were recorded, not collapsed.
How to read the output
These rules apply to every Insight Bridge corpus. They are properties of the method, not caveats about a particular run.
- Positions are relative to a proposition. A source’s position records how it stands against that cluster’s particular framing — not whether it agrees with some absolute claim. The same source can support one cluster and redirect a neighbouring one that covers similar ground differently.
- Propositions are written from the cluster’s own members. Because the argument is built from the sources that were grouped together, and positions are read relative to it, a degree of agreement is built into the method. Comparisons between groups carry weight; a corpus-wide agreement rate does not.
- Counts describe the corpus, not the world. Every corpus here is curated. “N sources say X” measures what was collected and is never a measure of how common X is in the field.
- Clusters differ in how many distinct sources back them. A long document can fragment across many clusters, so weight a theme by the distinct sources beneath it rather than by how many clusters it contains.
- Every extraction and position is a model judgement. Key points, stances, propositions and syntheses are produced by a language model reading the source. They inherit its calibration and are not determinations of fact.
Reading positions in this corpus
Where the disagreement is. The propositions here are things this literature actually argues, and 83.7% of source stances read supportive against them. Dissent therefore rarely surfaces as an Opposes count. It surfaces in the Supports vs Builds on split (plain endorsement versus endorsement-with-conditions), in voice divergence inside a single cluster, and in who is absent from a cluster altogether.
The tree is the pipeline's. Groups above the topics were produced by clustering the clusters, generation by generation, and named from their members. A group is a neighbourhood in embedding space, not a category anyone drew.