From f1fc85f42ecdded93461aed2634bb8191a06ee9d Mon Sep 17 00:00:00 2001 From: Deepanshu Pal <40927968+DeepanshuPal@users.noreply.github.com> Date: Sun, 27 Sep 2026 19:17:09 +0530 Subject: [PATCH 1/3] docs: correct classifier testing guidance --- docs/evaluations/jev.mdx | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/evaluations/jev.mdx b/docs/evaluations/jev.mdx index 2dce90532..0176906a3 100644 --- a/docs/evaluations/jev.mdx +++ b/docs/evaluations/jev.mdx @@ -81,8 +81,8 @@ Very long sessions are read in excerpts and combined. When a session is too long - **A classifier always produces a score**, never a metric or an assertion. - **No reasoning**, as above. If a number will make someone ask "why?", write a judge instead. -## Testing and backfill +## Backfill -Unlike a judge, a classifier evaluation **can** be tested before you deploy it — [test it](/evaluations/test) against real sessions the same way you would a code evaluation, and read the scores before anything goes live. +Like a judge, a classifier evaluation **cannot** be tested before you deploy it. Both spend your organization's model budget, and that spend is only authorized once the evaluation is deployed. Deploy it to a narrow scope first and read the early scores instead. It can also be [backfilled](/evaluations/deploy#score-sessions-you-already-have) over sessions you already have. It costs a model call per session, so scope the window deliberately rather than replaying everything. From ede75a36e6bdc1847eec2810cebd96e35b016847 Mon Sep 17 00:00:00 2001 From: Deepanshu Pal <40927968+DeepanshuPal@users.noreply.github.com> Date: Sun, 27 Sep 2026 19:17:51 +0530 Subject: [PATCH 2/3] docs: clarify code-only evaluation test panel --- docs/evaluations/test.mdx | 2 ++ 1 file changed, 2 insertions(+) diff --git a/docs/evaluations/test.mdx b/docs/evaluations/test.mdx index 3649f587e..cc84bbe2b 100644 --- a/docs/evaluations/test.mdx +++ b/docs/evaluations/test.mdx @@ -6,6 +6,8 @@ icon: "flask-conical" **test this evaluation**, on the authoring page, runs the code against real sessions of yours on the evaluator fleet without deploying it. Nothing is stored: a failure here is a preview, and deploying is always allowed. +**Code evaluations only.** Judge and classifier evaluations cannot be tested here — both spend your organization's model budget, which is only authorized once the evaluation is deployed. The panel tells you so if you try. + Select **check** to compile the code and condition against the sandbox's rules without running them on any session. From b081a0c28f5bc903a6dcefd01a68c0b4bf1b78dd Mon Sep 17 00:00:00 2001 From: Deepanshu Pal <40927968+DeepanshuPal@users.noreply.github.com> Date: Sun, 27 Sep 2026 19:52:49 +0530 Subject: [PATCH 3/3] docs: limit evaluation test page description to code --- docs/evaluations/test.mdx | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/evaluations/test.mdx b/docs/evaluations/test.mdx index cc84bbe2b..8a484ad81 100644 --- a/docs/evaluations/test.mdx +++ b/docs/evaluations/test.mdx @@ -1,6 +1,6 @@ --- title: "Test an evaluation" -description: "Run an evaluation against your real sessions before deploying it. Nothing is stored." +description: "Test a code evaluation against real sessions before deployment. Nothing is stored." icon: "flask-conical" ---