Skip to content

docs: update scoring methodology - #241

Merged
sanchitmehtagit merged 2 commits into
mainfrom
docs/scoring-methodology-applicability
Aug 26, 2026
Merged

docs: update scoring methodology#241
sanchitmehtagit merged 2 commits into
mainfrom
docs/scoring-methodology-applicability

Conversation

@brth31

@brth31 brth31 commented Aug 26, 2026

Copy link
Copy Markdown
Contributor
  • Hallucination/Security no longer default to 100 with zero graders on CLI-only evals (and Correctness no longer defaults to 0) — both become N/A and the remaining weights renormalize.
  • Efficiency gets a command-precision variant for cli-only evals: measures probe-and-retry against the declared target operation instead of file-based waste that can't fire without files.
  • Weight changes now require a before/after sensitivity check against stored results before merge.
  • Baseline (zero-tool) demoted from the headline comparison — it conflates the generic lift of having any tool access with the specific lift of Auth0's tools, and likely overstates the published delta.
  • Hallucination, Security, and Correctness given concrete, bounded criteria instead of open-ended labels.

…ine comparison

- Hallucination/Security no longer default to 100 with zero graders on
  CLI-only evals (and Correctness no longer defaults to 0) — both become
  N/A and the remaining weights renormalize.
- Efficiency gets a command-precision variant for cli-only evals: measures
  probe-and-retry against the declared target operation instead of
  file-based waste that can't fire without files.
- Weight changes now require a before/after sensitivity check against
  stored results before merge.
- Baseline (zero-tool) demoted from the headline comparison — it conflates
  the generic lift of having any tool access with the specific lift of
  Auth0's tools, and likely overstates the published delta.
- Hallucination, Security, and Correctness given concrete, bounded
  criteria instead of open-ended labels.
@brth31
brth31 requested a review from sanchitmehtagit August 26, 2026 04:45
@coderabbitai

coderabbitai Bot commented Aug 26, 2026

Copy link
Copy Markdown

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro Plus

Run ID: 4abe2ffa-7b8a-441c-9b88-bf2e428b98bf


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@brth31
brth31 requested review from frederikprijck, sanchitmehtagit and subhankarmaiti and removed request for sanchitmehtagit August 26, 2026 04:45
@sanchitmehtagit
sanchitmehtagit merged commit a4469f4 into main Aug 26, 2026
6 checks passed
@sanchitmehtagit
sanchitmehtagit deleted the docs/scoring-methodology-applicability branch August 26, 2026 04:52
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants