Skip to main content
A choice judgment answers with one of its options, or with none_of_the_above when none of them fits. That escape option is where a new kind of complaint, a new abuse pattern or a new product question shows up first. Discovery reads what landed there and proposes the options your judgment is missing. It never changes the judgment: you review the proposal, edit it, and create the new version yourself. Discovery is for choice judgments with fixed options. A bool or score judgment has no escape option; watch its calibration report instead.

Know when to look: the escape alert

escape_alert is a setting of a choice judgment: the share of its answers over the last 7 days that were none_of_the_above, above which you want to know. It is off until you set it, and changing it creates no version.
When the share rises above it, three things happen, and nothing else:
  • The judgment gets a warning. GET on the judgment lists taxonomy_drift in warnings, with the share and when it rose. The dashboard shows it on the judgment’s page. It clears when the share falls back.
  • One event. The events feed, and every webhook endpoint that asks for it, gets one judgment.taxonomy_drift per rise, not one per answer. Staying above sends nothing more.
  • Nothing changes on its own. Your answers, options and bill stay as they are. You run discovery when you want to.
The share needs at least 20 answers in the 7 days before it can raise the warning, so a quiet judgment’s first escape isn’t a drift. The alert takes a number above 0 and at most 1; null turns it off. A namespace that inherits the judgment from a template follows the template’s alert, and gets its own warning and event from its own answers.

Turn on suggestions

Discovery sends a sample of your documents to a general-purpose LLM provider, the same one that suggests parts, listed as a subprocessor in the data processing agreement. So it is off by default: an org admin turns on Suggestions in the organization’s settings in the dashboard. Until then, discover is refused with forbidden.

Run discovery

POST /namespaces/{ns}/judgments/{name}/discover with how far back to look (window, 7 days by default) and the most new options you want (count, 1 to 10, 5 by default):
It answers 202 with a discover job, which:
  1. Samples the escape. It takes the documents whose answer in the window is none_of_the_above, or whose escape_p is at or above the judgment’s threshold on none_of_the_above if you set one, and samples up to 200 of them.
  2. Asks for options. Each sampled document is compiled with the judgment’s context recipe, exactly as it is judged, and the sample is cut to 100,000 tokens together. The LLM gets your question, your current options and the sample, and proposes up to count new options, each with a description and the sampled documents it covers.
  3. Tries them. The current options plus the proposed ones are judged over the sample, as a shadow report judges a new version. These shadow judgments are free and change no answer.
Discovery is free. Each call takes one of the judgment’s suggest_parts calls for the day (20 per judgment and 50 per organization, together with suggested parts); past them it is rate_limited, with details.limit and when the allowance renews in details.resets_at. The call is taken when the job starts, and a busy model is waited out rather than failing the job.

Read the proposal

When the job is done, its proposal is a new version’s options, the current ones first:
  • escape_documents is how many documents in the window escaped; sampled is how many of them the job read.
  • absorbs is how many sampled documents the proposed version answered with that option; sample_ids are the ones the LLM said the option covers. Open a few of them to see what the option really means.
  • unlabelled is how many sampled documents the proposed version still answered none_of_the_above. A large number means the escape holds more than one new thing, or things no option should cover.

Review before you create the version

The proposal is a draft. The review is where most of its value is, so don’t accept it unedited. We tested discovery on a public dataset of banking support questions: 20 kinds of question, 80 each, with three kinds removed from the options so that their questions escaped. Two of the removed kinds sat close to options that stayed, and one was unlike any of them. All three came back as proposed options, each covering almost exactly the escaped questions of its kind. The same run showed what the review is for:
  • Some proposals are a narrower case of an option you have. Two of the five proposals were sub-topics of existing options: top-ups made with a phone wallet, under top-ups, and accounts for children, under accounts. Added as they were, they would have taken 73 and 39 of the 80 questions in their parent options. Fold a proposal like that into the parent option’s description instead of adding it, or delete it.
  • A new option next to an old one won’t take all of its documents. About a quarter of the questions of each removed kind that had a close neighbour never escaped: they were answered with the neighbour. After the new option was active, one of them still lost about a quarter of its questions to its neighbour. Sharpen both descriptions so the line between them is clear, and read the shadow report when you activate.
  • The conditions were easy. The removed kinds made up most of the escaped questions, and the dataset is public, so the model may have seen its categories before. A new kind that is a small part of your escape may not come back as its own option.
Then create the version with the options you settled on, as you create any version. Activating it runs the ordinary shadow report over up to 1,000 of your documents, which shows what would change before anything does. In the dashboard, the judgment’s Discover tab shows the proposal as an editable list of options, with the change to the definition beside it, and creates the version from what you edited.

Limits