Skip to content

Held-out F1 reports 0.000 for a correct round; gateway backends have no retry #3

Description

@MrJev

Hi — reviewing jev-align for a write-up. Two things I want to note before the bugs, because they're the parts I'd point other projects at: the holdout is frozen before the run starts, excluded from sampling, and never passed to GEPA — with test_holdout_labels_are_added_separately_and_never_sent_to_gepa asserting it — and no proposal is ever accepted without a keystroke. 127 tests pass offline at 49753df, and ruff is clean.

I drove two full rounds in Docker against local stand-ins for both the decision model and the reflection LLM (no real endpoints were contacted). Findings, most consequential first.

1. A perfectly correct holdout round reports F1 0.000.

src/jev_align/metrics.py:23-25 treats undefined as zero:

precision = tp / (tp + fp) if tp + fp else 0.0
recall    = tp / (tp + fn) if tp + fn else 0.0
f1 = 2 * precision * recall / (precision + recall) if precision + recall else 0.0

This is reached constantly because the holdout batch is one row: holdout_batch_size = max(1, round(batch_size * holdout_fraction)) (models.py:499-503) is 1 at the default batch_size=5. My single holdout row was a true negative — model said no, human said no — and the report was:

Held-out F1   0.000 · P 0.000 · R 0.000

with {"f1":0.0,"precision":0.0,"recall":0.0,"tn":1,"tp":0} in the round report, and holdout_score=0.0 written into state.json's permanent history (session.py:572) and shown on the jeva functions card.

A batch with no positives has an undefined F1, not a zero one. Returning None and rendering it as — (or falling back to accuracy when tp+fp+fn == 0) would stop a correct round from looking like a total failure.

2. Gateway backends have no retry, and a 429 says the wrong thing.

jev_gateways.py:96-103:

with urlopen(request, timeout=60) as response:
    payload = json.load(response)
except HTTPError as error:
    raise RuntimeError(f"Jev provider returned HTTP {error.code}; check credentials, access, and model availability")

The TypeSafe SDK path gets the SDK's RetryPolicy (429/5xx, backoff). The Cloudflare and Vercel paths get neither, so one 429 aborts the round — and with asyncio.Semaphore(16) and one request per row, a 1,000-row pool scan is a good way to earn one. The labels already written survive (labels.jsonl is append+fsync, nice), but the round's GEPA work is lost.

The message is also misleading for a 429 specifically: credentials and model availability are fine, you're being rate limited.

3. A round can spend its whole budget and propose nothing.

ClassAwareBatchSampler.next_minibatch_ids (optimizer.py:59-66) memoises self._selected, so the reflection minibatch is chosen once per run and never changes. F1BatchEvaluator returns the batch-wide F1 for every row (optimizer.py:216), so skip_perfect_score=True, perfect_score=1.0 (optimizer.py:396-398) effectively asks "is this fixed 5-row batch perfect?"

When it is, every iteration skips reflection but still pays for the evaluation. My second round: 40 of 40 metric calls spent, 0 reflection calls, 0 candidates, diff "No textual change." ScoreThresholdStopper(1.0) didn't fire because it reads the valset aggregate (0.933), not the minibatch.

Resampling the minibatch per iteration, or stopping the round when the parent is perfect on the sampled batch, would turn a wasted round into an early exit.

4. Reflection candidates aren't validated for binary/multiclass/multilabel.

BinaryTaskSpec.from_gepa (models.py:34-41) checks the key set and non-emptiness, nothing else. ScoreTaskSpec.from_gepa (models.py:134-147) has real guards — length caps and a cross-component marker check.

Your own comment at optimizer.py:385-387 names the hazard ("prevents a reflection response for the whole rubric from being pasted into every component") and applies module_selector="round_robin" to multilabel and score, leaving binary on "all". In my run a reflection model returned a fenced JSON object and GEPA pasted that whole object into all three components, which then went out in paid requests. A talkative reflection model can inflate every subsequent request without anything noticing.

5. Documentation, mostly about where data goes.

The reflection prompt I captured was 5,383 characters and contained, per selected example: the row's full contents, the human label, the prediction, and the human rationale verbatim. All 24 reflection calls carried them. That's inherent to GEPA and not a bug — but the README describes the reflection model only as "GEPA's reflection model", and someone choosing --reflection-model gpt-5.6 may not realise the data they're labelling, and the reasons they type, go to that provider. One sentence in the Reflection models section would cover it.

Smaller ones in the same spirit:

  • "An optional 20% held-out evaluation set" reads as 20% of rows being evaluated. 20% are reserved, but only max(1, round(batch_size * 0.2)) are labelled per round — ten rounds gives a ten-example evaluation set. The wizard's "add 20% extra held-out annotations each round" (cli.py:516) is the accurate phrasing.
  • "Selects ambiguous rows plus a random audit sample": the audit sample is hardcoded to exactly 1 (session.py:360) regardless of batch_size, so at 20 annotations a round it's still one.
  • "Uses your accumulated labels and optional rationales to run GEPA": accumulated labels score candidates, but the reflection LLM only ever sees min(5, len(examples)) of them (optimizer.py:381), fixed for the run by (3).
  • The default pool is the first 1,000 rows (data.py:190-194), not a random sample. For a dataset sorted by date or class that's a real sampling bias, and it's worth a line.
  • --max-metric-calls is soft: I asked for 40 and got 55. The UI calls it "one in-flight step past the budget"; that was +37.5%.
  • The CLI checks PyPI on every interactive start and offers an upgrade with "Upgrade now" as the default (cli.py:246-370). Worth documenting, alongside JEVA_DISABLE_UPDATE_CHECK=1.

One design note, not a bug. The labelling card pre-selects the model's own answer (cli.py:1569-1575 for binary, :1552 multiclass, :1522 score) and prints its probability above the row. Enter accepts it. That's an efficient default for the training batch; on holdout rows it means the ground truth used to grade the model defaults to the model's answer. Leaving holdout cards unselected — or hiding the probability on them — would cost one keystroke and remove the anchor.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions