> ## Documentation Index
> Fetch the complete documentation index at: https://docs.vainona.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Match one record to another

> Find which open invoice a payment settles, or which role a candidate fits: candidates that share a key, and one judgment that chooses among them.

export const productName = "Vainona";

"Which open invoice does this payment settle?" and "which of these roles does this candidate fit?" are matching questions. The usual answer is a pipeline: blocking rules, a pair document for every candidate, a yes/no judgment on each pair, and a rerun whenever either side changes. That costs one judgment per pair and leaves you to pick the winner.

In {productName} a matching judgment has two halves. Its **candidates** are the documents that share a key with the judged document, read by a blocking relation, one of the [relations](/concepts/relations). And the judgment **chooses among them** in one call, so the engine sees the alternatives side by side and names the one that fits, or says none does. On standard entity-matching benchmarks, choosing among a record's candidates this way matched judging each pair on its own.

This is one-to-one matching: each judged document gets at most one match. A payment settles one invoice; a candidate fits one role. If you need every duplicate of a record, many-to-many, keep one pair document per candidate pair with a yes/no judgment on each, as in [judge a document with the document it points at](/guides/referenced-document).

## Write the key

A block is every document that holds the same value in one attribute you write: its **key**. Payments and invoices in block `acme-gbp` are candidates for each other; nothing in block `acme-usd` is.

The key should encode what a match requires. A payment and an invoice in different currencies can never match, so put the currency in the key. A key that is too coarse, such as the week, puts hundreds of unrelated invoices in front of every payment, and each of them costs a re-judge when it changes. Write the key as a short string of letters, digits, `-`, `_` and `.`, such as a normalised counterparty name and a currency. A document with no key, or a key that is not such a string, is in no block: it reads no candidates and is nobody's candidate.

## Define the judgment

This judgment asks, for every unmatched payment, which of its block's open invoices it settles:

```json theme={"theme":{"light":"css-variables","dark":"css-variables"}}
{
  "name": "settles",
  "type": "choice",
  "applies_to": {"attributes.kind": "payment", "attributes.status": "unmatched"},
  "question": "Which of these open invoices does this payment settle? Answer none_of_the_above if it settles none of them.",
  "options": {"from": "candidates", "label": ["state.number", "state.amount", "state.counterparty"]},
  "context": {
    "fields": ["state.amount", "state.reference", "state.received_at", "state.payer"],
    "related": {
      "candidates": {
        "match": {"attributes.kind": "invoice", "attributes.status": "open"},
        "join": {"theirs": "attributes.block", "mine": "attributes.block"},
        "window": "180d",
        "last_n": 10,
        "fields": ["state.number", "state.amount", "state.issued_at", "state.counterparty"]
      }
    }
  },
  "engine": {"name": "jev", "version": "current"},
  "freshness": {
    "policy": "on_change",
    "fanout": {"scope": {"created_within": null}}
  },
  "outcomes": {"rules": [{"on": "state.settles", "from": "state.settles"}]},
  "confirm": true
}
```

* **The blocking relation.** `join` names the same attribute on both sides: the payment's `attributes.block` and the invoice's. It reads the block's newest open invoices, created within `window` of the block's newest one, at most `last_n`. `window` is required. A choice that chooses among a relation's documents reads at most the [candidate limit](/limits) of them. A blocking relation reads its candidates' attributes and state, not their answers: a path under `answers.` is refused.
* **The options.** `options.from` names the relation. Each candidate the payment reads becomes one option, in the relation's order, newest first: its `value` is the invoice's `id` and its description its `label` fields. After them comes `none_of_the_above`, which the engine reads as "none of these: no match".
* **The scope.** `created_within: null` keeps old unmatched payments in scope: they are the ones your operations team cares about.

## Read the answer

The answer is a choice answer. `value` is the matched invoice's `id`, or `none_of_the_above`. `escape_p` is the probability that no candidate matches. `dist` is a shortlist, not a ranking: the probabilities are rounded to two decimals, so past the first few candidates they tie at zero, and only the candidates with any probability are listed.

A payment whose block has no open invoice is answered `none_of_the_above` with `escape_p` 1, with no engine call and nothing billed.

**The unmatched queue is a threshold on `escape_p`.** Out of the box the answers err toward "none": the pick among the candidates is right nearly every time it is taken, and most misses are payments that did have a match where "none" won. So the threshold on `escape_p` is the setting that decides how much is matched for you:

```json theme={"theme":{"light":"css-variables","dark":"css-variables"}}
{"thresholds": {"unmatched": {"value": "none_of_the_above", "gte": 0.6}}}
```

A payment whose `escape_p` is below it has its pick taken; at or above it, it waits for a person. Query the queue with `["answers.settles.thresholds.unmatched", "Eq", true]`. Once outcomes arrive, `GET .../thresholds/recommend?target=precision:0.98` recommends the threshold at which the picks it lets through are right at least that often (see [pick thresholds with the recommender](/guides/measure-improve-tune#pick-thresholds-with-the-recommender)).

**The order of the candidates matters a little.** Shown the same candidates in another order, the engine changes its pick for a few in a hundred payments. The relation's order is fixed, newest first, so the same candidates always give the same answer; a new candidate arriving can reorder them, and with it the pick.

## Record the match: the reconciliation shape

When your team or your own system settles a payment, write the match on the payment, `state.settles` with the invoice's `id`, and the payment's new status, **in one write**:

```json theme={"theme":{"light":"css-variables","dark":"css-variables"}}
{"patch": [{"id": "pay_881", "attributes": {"status": "matched"}, "state": {"settles": "inv_2043"}}]}
```

That write does three things:

* The payment leaves `applies_to`, so it is no longer judged or answered. The evaluation that found the match stays in the evaluation log, and the invoice is in the payment's own state. Re-judging it against a block its invoice has just left would flip the answer to "no match" right after you acted on it.
* The outcome rule `{"on": "state.settles", "from": "state.settles"}` fires, because `state.settles` went from absent to set on a payment that matched `applies_to` before the write. The outcome joins the revision just before the write: the evaluation whose candidates still held the invoice.
* The invoice, when you mark it paid, leaves the relation's `match`, which changes its block, so the other payments in the block stop reading it as open.

Write the settlement and the status change together. If they are two writes, a block change between them can re-judge the payment and flip its answer first.

A payment that never matches stays in scope, and is re-judged whenever a candidate in its block changes. The key is what bounds that churn. Write payments you know will never match, such as refunds, out of `applies_to`.

## Freshness

A write to a candidate makes every payment in its block `pending`. When the block has been quiet for the fan-out's `debounce_ms`, or its `max_wait_ms` has passed, {productName} reads the block once and compares what it shows with what it showed before:

* If nothing the payments read changed, such as an edit to a field the relation does not show, the payments are `fresh` again, with no engine call.
* If it changed, every in-scope payment in the block is re-judged. A new open invoice, an invoice that settled, and an invoice replaced by another with the same fields are all changes. A payment whose own candidates did not change is not asked again: its context is the same, so it keeps its answer.

## Cost and limits

**Size classes.** A matching judgment's question includes every candidate with its label fields, so it counts toward the [size class](/pricing#judgments) with the context. With the relation above, a payment with ten candidates of three short label fields each has a question and context of about a thousand tokens: standard, counted as 1. The same payment with pair documents is one yes/no judgment per candidate: ten. Long label fields or more candidates can make the question large, counted as 4, so label with the fields that tell candidates apart, not everything you have.

**What it costs a month.** Re-judgments are block changes a month times the payments in scope per block. The create response's replay estimate counts the payments' own writes and every invoice created in the last 30 days, the common case: a new invoice changes every payment's candidates in its block. It cannot see invoices edited or deleted in that time, since storage keeps only each document's newest version, so it says `"excludes": ["candidate_edits", "candidate_deletes"]` and shows its figures as "at least". In receivables, where settling an invoice is an edit, expect about twice the estimate.

**The limits it runs under**, set in `freshness.fanout` and changed with a `PATCH`:

* `block_cap`: the most payments one block may hold. A create whose data already has a larger block is refused, naming the key value in `details.key`, so a coarse key is caught before its first fan-out. A block that grows past it later waits: its payments read `stale` with `stale_reason` `limit_reached`, the events feed carries `judgment.limit_reached` with `limit` `block_cap` and the key value in `key`, and raising the cap runs it.
* `rolling_limit`: the most re-judgments the judgment's block changes may cause in any 30 days. Past it, changes wait and their payments read `stale` (`limit_reached`) until the window has room. When it is short, the blocks that have spent the most of it wait first, so a few busy counterparties cannot leave everyone else's payments stale.

Each blocking key takes one of the namespace's [reference indexes](/limits).

## The oldest open invoice

The relation reads the newest candidates first, so in a large block the oldest open invoices are the ones left out, and in receivables the invoice a payment settles is often the oldest one. The answer is a finer key: add the currency, the amount band or the month the invoice is due, so each block holds only the invoices a payment could settle. The calibration report's `match_misses.beyond_last_n` counts the outcomes that named an invoice the payment's block held but its evaluation did not read; if it grows, your key is too coarse.

## Measure it

The [calibration report](/concepts/calibration) of a matching judgment reports two things from the same outcomes:

* **Whether a match exists**, which `1 − escape_p` estimates. Calibration fits it, and the calibrated answer's `escape_p` uses the fit.
* **Whether the pick was right**, in `selection`: the share of outcomes naming an invoice where the answer picked that invoice.

`match_misses` counts the outcomes that named an invoice the evaluation did not read, by cause: `candidate_newer_than_evaluation` (created after the evaluation), `beyond_last_n` (in the block, but cut by `last_n`), `not_in_block` (another key: a key that does not encode the match), and `evaluation_failed`.

## Duplicates within one kind of record

The judged documents can be their own candidates: listings with `"match": {"attributes.kind": "listing"}` and the same key on both sides. Each listing reads the others in its block, never itself, and names at most one as its duplicate. Every write to a listing changes its block, so a block of n listings that change often costs up to n judgments per change. When you need every duplicate pair rather than the best one per listing, use pair documents.


## Related topics

- [Relations](/concepts/relations.md)
- [Judgments](/concepts/judgments.md)
- [Writing a context recipe](/guides/context-recipes.md)
- [Import existing data](/guides/import-existing-data.md)
- [Roll up answers from related documents](/guides/roll-ups.md)
