Feat (eval): implementing EAR and KL-div evaluation metrics - #1598
Merged
Merged
Conversation
Giuseppe5
reviewed
Aug 31, 2026
Giuseppe5
reviewed
Aug 31, 2026
i-colbert
force-pushed
the
feat/more_metrics
branch
from
September 4, 2026 01:08
e5bfbcb to
3324c4a
Compare
i-colbert
marked this pull request as ready for review
September 4, 2026 18:02
Giuseppe5
reviewed
Sep 8, 2026
i-colbert
requested review from
Giuseppe5
and removed request for
Giuseppe5
September 10, 2026 00:22
Giuseppe5
reviewed
Sep 10, 2026
pablomlago
approved these changes
Sep 16, 2026
| self.dtype = dtype | ||
|
|
||
| @abstractmethod | ||
| def update(self, output: torch.Tensor, target: torch.Tensor) -> None: |
Collaborator
There was a problem hiding this comment.
I would maybe rename these methods, e.g. update to accumulate and finalize to aggregate, or something along those lines.
|
|
||
|
|
||
| @torch.no_grad() | ||
| def compute_float_evaluation_metrics( |
Collaborator
There was a problem hiding this comment.
Should the method name reflect that it is doing some caching apart on top of computing metrics?
| @@ -31,6 +31,11 @@ def parse_log(job_log: str) -> Dict[str, Any]: | |||
| # Find the line containing Quant PPL number | |||
Collaborator
There was a problem hiding this comment.
Maybe it would be worth extracting the common functionality, e.g.:
metric_logs = {"quant_ppl": r"Quantized perplexity \((.*?)\): (\d+\.\d+)", ...}
metrics = {metric_key: re.search(metric_value) for metric_key, metric_value in metric_logs.items()}
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Reason for this PR
Extend LLM evaluation with distribution-based metrics that quantify the effect of quantization beyond perplexity.
Perplexity (PPL) measures the quality of the quantized model on the observed target tokens. EAR and KLD provide additional information about how closely the quantized model preserves the float model's output distribution.
Changes Made in this PR
--evalas the single switch for PPL, EAR, and KLD evaluation.*Both EAR and KLD are normalized by the reference top-K probability mass by default.
Testing Summary
The following tests were added or updated:
tests/brevitas_examples/llm/test_eval.pytests/brevitas_examples/llm/test_llm.pytests/brevitas_examples/llm/test_benchmark.pyLocal evaluation was run with the following config:
Environment versions used for the evaluation:
Baseline
Results from master at
6df6a69:This PR
Using the above configuration with
--weight-bit-width=8we get:Below is a minimal evaluation of the overhead for collecting the additional metrics: