The citable data context for legal AI
Nearly twice the recall of keyword search, on a 97,913-query holdout. The harness is published.
What does your Legal AI base its trust on?
One pipeline, a widening corpus. Patents ship today; the rest are on the roadmap.
0.0M+
US patents in the live index
Currently processing
EP 4.7MJP 144kKR 40kCN 1.1M translating48.5M NPL references
Held today. The full EP, JP, KR and CN corpora are next.
0%
of examiner citations retrieved in our top-100 (97,913-query holdout)
4 / 4
domain slices won vs the research state of the art, same harness
0%
of claims: a top-20 hit covers more claim elements than the examiner’s citation
Examiner citations retrieved in top-100 · identical harness
Measured on a pinned 97,913-query examiner-citation holdout, a head-to-head with the research state of the art, and the public DAPFAM benchmark. The negative results are published next to the wins, and everything reruns on one GPU in an afternoon.
Read the full benchmarks →The layer nobody sells
Legal AI has to cite.
Legal AI is moving from storing documents to acting on them. Anything that drafts or argues has to cite, and its output is only as good as the citation under it.
Nobody sells that retrieval layer. Aggregators ship raw feeds; the platforms keep their pipelines internal. So every team rebuilds it before it can build anything. We sell it, as an API and an MCP server.
The pipeline every team rebuilds
raw documents → text
parties, citations, dates
the decision that sets accuracy
two neural stages
against citations of record
Chunking alone drives about 35% of RAG answer accuracy, ahead of re-ranking at 28%. This step decides answer quality, and it was nobody’s product.
What we do
Data layer
Transformation layer
Every legal corpus, parsed for retrieval.
Patents now. Case law and dockets next.
Tuned for conflict, not matching words.
Ranked by what an examiner or litigator would cite.
What we are
The retrieval layer, as an API.
The two layers, rejoined. One retrieval layer your platform searches, cites, and acts on: parsed from every legal corpus, ranked by what an examiner or litigator would cite, and traceable back to the record.
MCP · REST
Antecedint API
One call runs the whole search.
An MCP server and a REST endpoint over 10.5M+ patents. One call; a cited answer comes back.
The wedge
Why patents first
Patents are the hardest corpus in law to retrieve against. That is the reason to start there, not a reason to avoid it.
- 01
The ground truth exists
Examiner citations label millions of claim-to-art pairs. We structured them into training and evaluation data. They are why our numbers can be checked rather than asserted.
- 02
The hardest retrieval task in law
A claim is a single-sentence legal hierarchy, and the decisive art is often filed in another field and another language. Solve claim-to-art retrieval and case law, statutes, and dockets are downhill.
- 03
The spend is mandatory
Every filing, office action, and invalidity challenge needs a defensible citation, in every jurisdiction, every year. About 3.5 million patent applications are filed globally each year.
Beyond patents
One pipeline, widening corpus.
Each corpus runs on the pipeline the last one paid for. The corpus widens; the retrieval discipline does not change. Phase 1 is the product; phases 2 to 4 are strategy, not a roadmap.
- 1
In production
Patents and prior art
Claim hierarchies, cross-domain art, six filing languages. USPTO, EPO, WIPO, JPO, KIPO, CNIPA.
- 2
Strategy
Case law and statutes
Nine million precedential decisions from over 2,000 courts, already bulk-available. The archive runs through the same pipeline; we do not rebuild it.
- 3
Strategy
Litigation and dockets
Where a claim’s language was litigated, before which judge, to what outcome. Bridges prosecution to enforcement.
- 4
Strategy
Corporate filings
Patents are financial assets. Resolving assignees to SEC registrants closes the loop from claim to cap table.
- Honest numbers
- Benchmarks published with the caveats and the negative results. Reproducible on one GPU in an afternoon.
- No black box
- Our own retrieval models and our own citation warehouse. Every ranking explainable, every result auditable.
- Yours stays yours
- Query text is processed inside our cloud account and never used to train models. No customer data by design.
Find the clearing.
Bring an idea. Leave with the closest art, and the receipts.






