Skip to main content
When many documents are judged against one referenced document, one of the relations, such as posts against a community’s rulebook or order lines against a product policy, changing it re-judges all of them. Before you write the change, ask what it would flip: simulate evaluates a sample of those documents with the rulebook as it is and as you propose it, and tells you how many answers would cross each of your thresholds.

Run one

Send the proposed document as you would upsert it, and the relation that reads it:
It returns 202 with a simulation job. relation must be one of the judgment’s relations that reads a referenced document, the document must be one the write API would accept, and sample is 1 to 1,000, 200 by default. Nothing is written: the rulebook stays as it is until you write it. The job:
  1. counts the judged documents that point at the rulebook and are inside the judgment’s re-judge scope (freshness.fanout.scope): the ones writing the change would re-judge;
  2. draws sample of them at random, so each sampled document stands for the same share of them;
  3. evaluates each sampled document twice, at the same moment and on the same model, once with the relation showing the rulebook as it is and once as you propose it.

Read the result

GET /v1/jobs/{id} shows simulation once the job is done:
  • thresholds: per named threshold, the sampled documents it would newly hold for (sampled_on) and stop holding for (sampled_off), and the flips estimated across the whole population, either way, with a 95% interval.
  • examples: up to 20 sampled documents that flipped, to read for yourself.
  • excluded: sampled documents left out: failed when either evaluation failed, out_of_scope when the document was gone by the time it was evaluated. The estimate counts only the documents evaluated both ways.
  • fanout: what writing the change would re-judge, priced like a backfill of the documents in scope, at the tiers your organization is in. Writing it runs that fan-out under the judgment’s rolling limit, 300,000 re-judgments in any 30 days by default: past it, the fan-out is deferred and catches up on its own, never refused.
  • larger_sample: when an interval is too wide to decide, wider than 5% of the population either way, the smallest sample size up to 1,000 that would narrow it enough; run again with that sample. null when every interval decides, as here.
Zero flips in the sample shows as an interval above zero, never as “nothing changes”: 0 of 200 is consistent with up to about 900 of 48,210. When the sample is the whole population, the count is exact and the interval is a single number.

Cost and limits

A simulation is free: its evaluations are not billed and write nothing. Because it can be run again and again, each one takes one of the judgment’s suggest_parts calls for the day, 20 per judgment and 50 per organization; past either, the route answers rate_limited with details.limit and details.resets_at, the next midnight UTC. See limits.