Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions docs/evaluations/jev.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -81,8 +81,8 @@ Very long sessions are read in excerpts and combined. When a session is too long
- **A classifier always produces a score**, never a metric or an assertion.
- **No reasoning**, as above. If a number will make someone ask "why?", write a judge instead.

## Testing and backfill
## Backfill

Unlike a judge, a classifier evaluation **can** be tested before you deploy it — [test it](/evaluations/test) against real sessions the same way you would a code evaluation, and read the scores before anything goes live.
Like a judge, a classifier evaluation **cannot** be tested before you deploy it. Both spend your organization's model budget, and that spend is only authorized once the evaluation is deployed. Deploy it to a narrow scope first and read the early scores instead.

It can also be [backfilled](/evaluations/deploy#score-sessions-you-already-have) over sessions you already have. It costs a model call per session, so scope the window deliberately rather than replaying everything.
4 changes: 3 additions & 1 deletion docs/evaluations/test.mdx
Original file line number Diff line number Diff line change
@@ -1,11 +1,13 @@
---
title: "Test an evaluation"
description: "Run an evaluation against your real sessions before deploying it. Nothing is stored."
description: "Test a code evaluation against real sessions before deployment. Nothing is stored."
icon: "flask-conical"
---

**test this evaluation**, on the authoring page, runs the code against real sessions of yours on the evaluator fleet without deploying it. Nothing is stored: a failure here is a preview, and deploying is always allowed.

**Code evaluations only.** Judge and classifier evaluations cannot be tested here — both spend your organization's model budget, which is only authorized once the evaluation is deployed. The panel tells you so if you try.

<Steps>
<Step title="Check that it compiles">
Select **check** to compile the code and condition against the sandbox's rules without running them on any session.
Expand Down