The citable data context for legal AI

Nearly twice the recall of keyword search, on a 97,913-query holdout. The harness is published.

What does your Legal AI base its trust on?

One pipeline, a widening corpus. Patents ship today; the rest are on the roadmap.

0.0M+

US patents in the live index

Currently processing

EP 4.7MJP 144kKR 40kCN 1.1M translating48.5M NPL references

Held today. The full EP, JP, KR and CN corpora are next.

0%

of examiner citations retrieved in our top-100 (97,913-query holdout)

4 / 4

domain slices won vs the research state of the art, same harness

0%

of claims: a top-20 hit covers more claim elements than the examiner’s citation

Examiner citations retrieved in top-100 · identical harness

Antecedint (production)45.3%
Qwen3-0.6B (untrained base)37.0%
GTE-ModernBERT29.8%
BM2524.3%
PatentSBERTa21.4%
How this was measured: the holdout, the split, and the metrics →

Measured on a pinned 97,913-query examiner-citation holdout, a head-to-head with the research state of the art, and the public DAPFAM benchmark. The negative results are published next to the wins, and everything reruns on one GPU in an afternoon.

Read the full benchmarks →

The layer nobody sells

Legal AI has to cite.

Legal AI is moving from storing documents to acting on them. Anything that drafts or argues has to cite, and its output is only as good as the citation under it.

Nobody sells that retrieval layer. Aggregators ship raw feeds; the platforms keep their pipelines internal. So every team rebuilds it before it can build anything. We sell it, as an API and an MCP server.

The pipeline every team rebuilds

Parse

raw documents → text

Resolve

parties, citations, dates

Chunk35% of accuracy

the decision that sets accuracy

Embed

two neural stages

Evaluate

against citations of record

Chunking alone drives about 35% of RAG answer accuracy, ahead of re-ranking at 28%. This step decides answer quality, and it was nobody’s product.

What we do

Data layer

Transformation layer

Every legal corpus, parsed for retrieval.

Patents now. Case law and dockets next.

Tuned for conflict, not matching words.

Ranked by what an examiner or litigator would cite.

What we are

antecedint

The retrieval layer, as an API.

The two layers, rejoined. One retrieval layer your platform searches, cites, and acts on: parsed from every legal corpus, ranked by what an examiner or litigator would cite, and traceable back to the record.

MCP · REST

Antecedint API

One call runs the whole search.

An MCP server and a REST endpoint over 10.5M+ patents. One call; a cited answer comes back.

Toolslocatecomparetrace

The wedge

Why patents first

Patents are the hardest corpus in law to retrieve against. That is the reason to start there, not a reason to avoid it.

  1. 01

    The ground truth exists

    Examiner citations label millions of claim-to-art pairs. We structured them into training and evaluation data. They are why our numbers can be checked rather than asserted.

  2. 02

    The hardest retrieval task in law

    A claim is a single-sentence legal hierarchy, and the decisive art is often filed in another field and another language. Solve claim-to-art retrieval and case law, statutes, and dockets are downhill.

  3. 03

    The spend is mandatory

    Every filing, office action, and invalidity challenge needs a defensible citation, in every jurisdiction, every year. About 3.5 million patent applications are filed globally each year.

Beyond patents

One pipeline, widening corpus.

Each corpus runs on the pipeline the last one paid for. The corpus widens; the retrieval discipline does not change. Phase 1 is the product; phases 2 to 4 are strategy, not a roadmap.

  1. 1

    In production

    Patents and prior art

    Claim hierarchies, cross-domain art, six filing languages. USPTO, EPO, WIPO, JPO, KIPO, CNIPA.

  2. 2

    Strategy

    Case law and statutes

    Nine million precedential decisions from over 2,000 courts, already bulk-available. The archive runs through the same pipeline; we do not rebuild it.

  3. 3

    Strategy

    Litigation and dockets

    Where a claim’s language was litigated, before which judge, to what outcome. Bridges prosecution to enforcement.

  4. 4

    Strategy

    Corporate filings

    Patents are financial assets. Resolving assignees to SEC registrants closes the loop from claim to cap table.

Honest numbers
Benchmarks published with the caveats and the negative results. Reproducible on one GPU in an afternoon.
No black box
Our own retrieval models and our own citation warehouse. Every ranking explainable, every result auditable.
Yours stays yours
Query text is processed inside our cloud account and never used to train models. No customer data by design.

Find the clearing.

Bring an idea. Leave with the closest art, and the receipts.