Skip to content

Feat (llm/eval): Standard deviation to perplexity computation - #1520

Open
pablomlago wants to merge 4 commits into
Xilinx:masterfrom
pablomlago:feat/std-ppl
Open

pablomlago wants to merge 4 commits into
Xilinx:masterfrom
pablomlago:feat/std-ppl

Conversation

@pablomlago

Copy link
Copy Markdown
Collaborator

Reason for this PR

compute_perplexity reports a single point estimate, making it hard to tell whether perplexity differences between configurations (float vs. quantized, recipe A vs. B) are meaningful. Other ecosystem tools (lm-eval-harness, lighteval) report a bootstrap stderr alongside perplexity for this reason.

Changes Made in this PR

  • src/brevitas_examples/llm/llm_quant/eval.py
    • compute_perplexity now returns Tuple[float, float]: (ppl, ppl_std).
    • Stderr computed via scipy.stats.bootstrap over the per-sequence NLLs with statistic exp(mean(x)), n_resamples=1000, method="BCa".
  • src/brevitas_examples/llm/main.py
    • Both call sites unpack the tuple and print "<X> perplexity (<dataset>): <ppl:.3f> ± <std:.3f>".
  • src/brevitas_examples/llm/benchmark/llm_benchmark.py
    • eval_metrics extended with float_ppl_std and quant_ppl_std.
    • parse_log regex updated to capture the new ± <std> suffix as an optional group, keeping backward compatibility with older logs.

Bootstrap was chosen over a closed-form stderr because perplexity (exp(mean(nll))) is non-linear in the NLLs; this matches the lighteval/lm-eval convention.

Testing Summary

  • Sanity import of the new compute_perplexity signature.
  • Manual float + quantized eval run confirming the new print format and that parse_log extracts both ppl and std.
  • Verified parse_log backward compatibility on logs without the ± <std> suffix (returns None for _std fields).

Risk Highlight

  • This PR includes code from another work (please detail).
    • Bootstrap design follows lighteval's bootstrap_stderr_scipy (MIT, EleutherAI / HuggingFace 2024).
  • This PR contains API-breaking changes.
    • compute_perplexity return type changes from float to Tuple[float, float]. In-tree call sites are updated.
  • This PR depends on work in another PR (please provide links/details).
  • This PR introduces new dependencies (please detail).
    • First direct use of scipy.stats.bootstrap in this module (scipy is already a transitive dep).
  • There are coverage gaps not covered by tests.
    • No dedicated unit test for the bootstrap stderr; covered indirectly via existing eval/benchmark paths.
  • Documentation updates required in subsequent PR.

Checklist

  • Code comments added to any hard-to-understand areas, if applicable.
  • Changes generate no new warnings.
  • Updated any relevant tests, if applicable.
  • No conflicts with destination dev branch.
  • I reviewed my own code changes.
  • Initial CI/CD passing.
  • 1+ reviews given, and any review issues addressed and approved.
  • Post-review full CI/CD passing.

@pablomlago pablomlago changed the title Add standard deviation to perplexity computation Feat (llm/eval): Standard deviation to perplexity computation May 14, 2026
@pablomlago
pablomlago changed the base branch from dev to master August 6, 2026 09:25
@pablomlago
pablomlago marked this pull request as ready for review August 6, 2026 09:25
@pablomlago
pablomlago requested a review from Giuseppe5 August 24, 2026 09:23

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant