plan_required (HTTP 402).
Export evaluations
An export is a job. Start it for a namespace, then poll the job until it isdone:
POST /namespaces/{ns}/evaluations/exports takes:
since, the earliestcreated_atexported, inclusive. It defaults to 30 days beforeuntil.until, the end of the range, exclusive. It defaults to now, and a later time is taken as now.judgments, to export only these. Leave it out for every judgment the namespace has evaluations of, deleted judgments included.
{} exports the last 30 days of every judgment. One export covers at most 366 days: export a longer history in several. Starting one needs a read_write key whose scope covers the namespace, and is in your audit log as job.create.
The job is GET /jobs/{id} with type: "evaluation_export". While it runs, progress.documents_done and export.evaluations count the evaluations written so far. You can pause, resume and cancel it like any job. When it is done, export.files lists its files:
- Each link works for an hour and needs no API key, so hand it to whatever downloads the file. Read the job again for new links.
- If links can’t be made right now, the job still reads as
donewith its files listed, but theirurlandurl_expires_atarenull, andlinks_unavailablesays so, with a request id to quote to support. The files are unaffected: read the job again later. - Files are kept 7 days after they are written.
expires_atis when the first one is deleted; after it the job lists no files, and you start a new export. - Files hold up to 100,000 evaluations, and an export writes at most 1,000 files. One that needs more fails with an
errorthat says so: narrow the range, or export fewer judgments at a time. - An export runs in the background and resumes where it stopped if it is interrupted, so a large one can take a while. Deleting the namespace deletes its exports too.
What a file holds
Each file is gzipped JSON Lines: one evaluation per line, asGET /namespaces/{ns}/evaluations/{id} returns it, plus the outcomes joined to it. Most tools read it as it is: BigQuery, Snowflake and Athena load .jsonl.gz directly, and gunzip -c 00001.jsonl.gz | jq . shows it.
contextis the exact compiled context the engine read, andrelated_documents, for a judgment with relations, what it read of other documents.rawis the engine’s response.outputis the answer’s raw numbers:p,value,dist,escape_p,scoreorparts. Thresholds and calibration are applied when answers are read, so the export holds what the engine said.engine_versionis the epoch the evaluation was made in: calibration fits each epoch on its own (see when the engine changes).outcomesare the outcomes calibration joins to this evaluation: posted outcomes, those your outcome rules derived (source: "rule") and labelling-queue labels ("queue"). Implicit negatives are calibration’s inference rather than outcomes, so they aren’t listed. It is empty when none joins.- Every evaluation is included: failed ones with their
error, shadow evaluations from shadow reports markedshadow: true, and replays after an engine change withreplay_ofset andcontext: null(the context is the evaluation it names). Evaluations of documents deleted since are included too. - Files aren’t in any order. Sort by
created_atif you need to. Rarely, an evaluation appears twice, when the history was reorganized while the export ran:idis unique, so keep one line perid.
Export the audit log
The dashboard’s Audit page lists your organization’s audit events on every plan. On Scale, Export CSV downloads the events that match its filters: each event’s time, actor, action, target, and the values before and after.After a downgrade
Below Scale, new exports are refused withplan_required, and the audit log’s CSV export goes. An export’s files stay downloadable until they are deleted, and the audit log stays in the dashboard. A downgrade takes effect on the 1st of a month at least 30 days after you ask (changing plan).