> ## Documentation Index
> Fetch the complete documentation index at: https://docs.vainona.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Roll up answers from related documents

> Judge an account from its conversations' answers: read another judgment's answer at a path in a relation, banded or cut, kept current as those answers change.

export const productName = "Vainona";

"How is this account doing?" is often a question about what its conversations were judged to be, not only their text. If you already run a judgment on each conversation, such as `frustrated` or `legal_threat`, a judgment on the account can read those answers the way it reads any field of a related document, beside the account's own facts. The account's answer stays current as the conversations' answers change, and each evaluation says which answer of which conversation it read.

Without it, you copy each conversation's answer into the account and rewrite the account whenever one changes: a pipeline. With it, {productName} re-judges the account only when a conversation's answer moves in a way the account can see.

Roll-ups are about keeping that current without a pipeline, not about accuracy. In our test, a judgment reading a roll-up of child answers did no better than one reading the children's text, and neither beat plain counts ([what we measured](#what-we-measured)).

This guide builds on [judging a document with its related documents](/guides/related-documents). Read that first: roll-ups add answers to what a relation can show, and everything else about [relations](/concepts/relations) stays the same.

## Define one

A relation reads a judgment's current answer at `answers.<judgment>.<field>`, the same paths a [query](/concepts/answers) filters on: `.p`, `.value`, `.score`, `.dist.<option>` and `.thresholds.<name>`.

```json theme={"theme":{"light":"css-variables","dark":"css-variables"}}
{
  "name": "account_health",
  "type": "bool",
  "applies_to": {"attributes.kind": "account"},
  "question": "Is this account at risk of cancelling in the next 60 days?",
  "context": {
    "fields": ["state.name", "attributes.plan"],
    "related": {
      "conversations": {
        "match": {"attributes.kind": "conversation"},
        "join": {"theirs": "attributes.account_id", "mine": "id"},
        "window": "30d",
        "last_n": 10,
        "fields": [
          "state.subject",
          {"path": "answers.frustrated.p", "bands": [0.5, 0.8], "labels": ["calm", "uneasy", "frustrated"]}
        ],
        "aggregate": {
          "count": true,
          "count_where": {"answers.frustrated.p": {"gte": 0.8}, "answers.legal_threat.thresholds.any": true},
          "max": [{"path": "answers.legal_threat.p", "bands": [0.5], "labels": ["none", "some"]}]
        }
      }
    }
  },
  "engine": {"name": "jev", "version": "current"},
  "freshness": {"policy": "on_change", "debounce_ms": 600000},
  "horizon": "60d",
  "confirm": true
}
```

The context shows each of the last ten conversations' `frustrated` answer as a word, how many of them are both frustrated and a legal threat, and whether any of them is a legal threat at all:

```json theme={"theme":{"light":"css-variables","dark":"css-variables"}}
"related.conversations": {
  "count": 3,
  "count_where": 1,
  "max(answers.legal_threat.p)": "none",
  "records": [
    {"answers.frustrated.p": "frustrated", "state.subject": "Export keeps timing out"},
    {"answers.frustrated.p": "uneasy", "state.subject": "Invoice looks wrong"},
    {"answers.frustrated.p": "calm", "state.subject": "How do I add a seat?"}
  ]
}
```

A conversation with no answer yet shows no answer field and counts for nothing in `count_where`.

What the context shows is not necessarily what the judgment weighs. In our test the engine's answers barely moved with the counts rendered into its context. When a number should decide the answer, such as how many conversations were frustrated, give it to a [composite judgment](/guides/composite-judgments#features), which fits a relation's aggregates as features against your outcomes, rather than relying on the engine to read it.

## Every rendered answer is stable

A probability moves a little every time a conversation is judged again. If the account saw the raw number, it would be judged again every time too, and nothing would ever be a [dedup](/concepts/judgments) hit. So what a relation shows of an answer only changes when the answer crosses a line you chose:

| Where | What it can read |
| - | - |
| `fields` | `.value` and `.thresholds.<name>` as they are; `.p`, `.score` and `.dist.<option>` only with `bands` |
| `aggregate.count_where` | documents whose answer matches: `.value` or `.thresholds.<name>` by value, or `.p` or `.score` by a cut such as `{"gte": 0.8}` (one of `gte`, `gt`, `lte` and `lt`) |
| `aggregate.min`, `aggregate.max` | `.p`, `.score` or `.dist.<option>` with `bands`: the band of the lowest or highest one |
| `aggregate.latest` | `.value` and `.thresholds.<name>` of the newest document that has one |
| `aggregate.sum` | never an answer |
| `match` | `.value` or `.thresholds.<name>` by value, or a cut on `.p` or `.score`, beside attributes |

`sum` over an answer is refused because a sum changes on every evaluation, and even a banded sum can cross a band when no single conversation's answer did. A [named threshold](/concepts/calibration) means what you decided it means from your outcomes; an inline cut is for when you have not set one, and is just as stable.

A banded `max` says whether any of the documents reaches a band, so over many documents it almost always does. In our test it read the top band for 99% of users once a relation held 50 or more documents, and told the judgment nothing. Keep `max` and `min` to a short relation, such as the last ten, and use `count_where` for many.

## What a judgment can read

The judgment whose answers a relation reads must be:

* **plain:** its own context reads no related documents and it is not a composite, so one hop stays one hop and nothing reads an answer that moves without being judged;
* **active and `on_change`**, so its answers stay current as they are read;
* **in the same namespace or template;**
* **answering every document the relation selects:** its `applies_to` covers the relation's `match`.

A relation reads at most 4 judgments' answers, and answer paths count toward its 8 aggregate paths. A create that breaks a rule is refused with `invalid_request`, and `details.path` names the path.

Once another judgment reads it, the judgment has **readers**, and its `GET` lists them. Deleting or deactivating it, switching it to a policy other than `on_change`, or activating a version that is not plain or whose `applies_to` no longer covers a reader's `match` is refused with `conflict`, naming the readers in `details.readers`.

## When the account is judged again

A conversation's new answer re-judges the account only when it changes what one of the account's relations shows of that conversation: its band, whether it matches a cut or threshold in `count_where` or `match`, the band it gives a `min` or `max`, or its `latest` value. A conversation judged a thousand times with `p` moving inside a band costs the account nothing.

From there it is an ordinary change to a related document: the account's `debounce_ms` and its ceiling apply, so a burst of conversation answers costs one evaluation of the account.

A conversation deleted and written again with the same id is a new conversation: its first answer re-judges the account even when the number is the same as before.

**Selecting by answer.** A `match` on an answer, such as only the frustrated conversations, keeps the documents whose answer holds before `last_n` keeps the newest. Membership changes only when a conversation crosses the cut. To find them, the relation reads at most 4 × `last_n` of the newest documents that match on attributes; if it reads that many before `last_n` qualify, the relation shows `capped: true`. Ordering by an answer is not available.

**The document it points at.** A relation that reads its [referenced document](/guides/referenced-document) can read that document's answers the same way. A change to what it shows is a referenced change, and fans out to the documents that point at it under the same limits.

## Pending conversations

The account reads each conversation's newest answer. When a conversation has been written and not judged again yet, the account reads its last answer, and the account's evaluation says `child_answers_stale: true`. The conversation's new answer re-judges the account anyway if it changes what the account shows, so this is a flag to wait out, not to act on.

## The audit

An account's evaluation lists every conversation it read in `related_documents`, and for each one whose answers it read, `evaluations` names the evaluation behind each answer: `{"frustrated": "ev_…", "legal_threat": "ev_…"}`. Evaluations are kept forever, so you can always follow an account's answer back to the conversation answers it read. `answers_generation` says which answers the context read, beside `watermark` for the documents.

## What it costs

A conversation's answer moving is a cost to every judgment that reads it. So the three requests that move conversation answers, or change what an account shows of them, say what they cost the readers before you confirm, as `downstream`, one line per reader:

* **a backfill estimate** of `frustrated`: the accounts it re-judges, in the month it runs;
* **the shadow report** of a new `frustrated` version: the accounts its answers would re-judge, from the sample;
* **a threshold change** on `frustrated` that a reader's relation names (below).

```json theme={"theme":{"light":"css-variables","dark":"css-variables"}}
{"downstream": [{"judgment": "account_health", "judgments_per_month": 18400, "cost_usd": 4.6}]}
```

Each line counts, from the last 30 days of the conversations' answers, how often their band or cut flipped, and prices those evaluations of the reader at your tiers. It is exact except for a conversation that moved between accounts in that time, which the history does not record.

When you already run `frustrated` for its own sake, reading it adds nothing on the conversation side: its answers exist anyway, and the account's re-judgments are ones it would have had from the conversations' writes. So read judgments you use yourself. Running a judgment only to feed another means paying for every child evaluation too, and in our test that bought no accuracy.

## Changing a threshold a reader reads

A threshold is a setting: changing one applies to every answer at once, with no evaluation. When a reader's relation names that threshold, such as `answers.legal_threat.thresholds.any` above, the change moves what the reader shows with no answer to re-judge it by. So the `PATCH` returns the `downstream` estimate and changes nothing, the rest of the request included:

<CodeGroup>
  ```ts TypeScript theme={"theme":{"light":"css-variables","dark":"css-variables"}}
  const result = await ns.judgments.update("legal_threat", { thresholds: { any: 0.7 } });
  if ("downstream" in result && !("judgment" in result)) {
    // Nothing changed. Show each reader's cost, then confirm.
    console.log(result.downstream);
    await ns.judgments.update("legal_threat", { thresholds: { any: 0.7 }, confirm: true });
  }
  ```

  ```python Python theme={"theme":{"light":"css-variables","dark":"css-variables"}}
  result = ns.judgments.update("legal_threat", thresholds={"any": 0.7})
  if "downstream" in result and "judgment" not in result:
      # Nothing changed. Show each reader's cost, then confirm.
      print(result["downstream"])
      ns.judgments.update("legal_threat", thresholds={"any": 0.7}, confirm=True)
  ```
</CodeGroup>

With `confirm: true` the thresholds apply to every read at once, as ever, and a `resync` job starts. It goes through `legal_threat`'s answers once and re-judges every account whose conversation now reads differently. Until the job reaches an account, that account's answer stays `fresh` under the old threshold. The response has `job_id` and the same `downstream` lines: if you are applying a [threshold recommendation](/guides/measure-improve-tune), this is often where you learn that another judgment reads it.

A threshold no reader names changes at once, as it always has. An inline cut lives in the reader's own definition, so changing it is a new version of the reader, with its own backfill estimate.

## Two levels, not three

"Message, conversation, account" is not a chain of roll-ups: a judgment can read a plain judgment's answers, and a judgment that reads related documents is not plain. It works when the leaf carries both keys, which most schemas do. A message with `attributes.conversation_id` and `attributes.account_id` can be read by a judgment on its conversation and by a judgment on its account, each one hop:

```mermaid theme={"theme":{"light":"css-variables","dark":"css-variables"}}
flowchart LR
  m["message"] -- "attributes.conversation_id" --> c["conversation"]
  m -- "attributes.account_id" --> a["account"]
```

So judge each message (`angry`), read `answers.angry` from the conversation, and read `answers.angry` again from the account, not the conversation's answer. If your messages carry only their conversation's id, write the account's id on them too.

## What we measured

We tested roll-ups on a public Q\&A dataset: will a user who posted at least 50 times in their first six months still be active six months later? That is the case where a user's posts no longer fit in one context, so a roll-up could have helped. It did not:

| What the judgment read | AUROC |
| - | -: |
| A roll-up of three judgments on each post: `count_where` and banded maxima, no text | 0.53 |
| The newest posts that fit in one standard judgment, as text | 0.61 |
| Counts of the user's posts, with no model | 0.76 |

The roll-up scored at chance, below the text, and neither added anything to the counts. The numbers it showed did carry some signal, but the engine's answers barely moved with them, and each `max` read the top band for almost every user. So this guide makes no claim that roll-ups are more accurate than text or than counts. Use them to judge an account from what its conversations were judged to be, kept current without a pipeline, and put the account's own facts, such as counts, beside them. Measure your own judgment against your outcomes (see [measure, improve and tune](/guides/measure-improve-tune)).


## Related topics

- [Judge a document with its related documents](/guides/related-documents.md)
- [Relations](/concepts/relations.md)
- [Writing a context recipe](/guides/context-recipes.md)
- [Judge a document with the document it points at](/guides/referenced-document.md)
- [Composite judgments](/guides/composite-judgments.md)
