Skip to content

Feat (eval): implementing EAR and KL-div evaluation metrics - #1598

Merged
Giuseppe5 merged 17 commits into
Xilinx:masterfrom
i-colbert:feat/more_metrics
Sep 24, 2026
Merged

Giuseppe5 merged 17 commits into
Xilinx:masterfrom
i-colbert:feat/more_metrics

Conversation

@i-colbert

@i-colbert i-colbert commented Aug 27, 2026 •

Copy link
Copy Markdown
Collaborator

Reason for this PR

Extend LLM evaluation with distribution-based metrics that quantify the effect of quantization beyond perplexity.

Perplexity (PPL) measures the quality of the quantized model on the observed target tokens. EAR and KLD provide additional information about how closely the quantized model preserves the float model's output distribution.

Changes Made in this PR

  • Add Expected Acceptance Rate (EAR) to quantized LLM evaluation*.
  • Add KL divergence (KLD) to quantized LLM evaluation*.
  • Compute both metrics over the float model’s top-K tokens, either normalized or unnormalized.
  • Cache the float model’s top-K token IDs and probabilities for reuse during quantized evaluation.
  • Keep --eval as the single switch for PPL, EAR, and KLD evaluation.
  • Add EAR and KLD to the returned evaluation results and benchmark output.
  • Separate float and quantized evaluation paths while sharing logits and evaluation-window handling.

*Both EAR and KLD are normalized by the reference top-K probability mass by default.

Testing Summary

The following tests were added or updated:

  • tests/brevitas_examples/llm/test_eval.py
    • Tests EAR and KLD calculations.
    • Tests normalized and unnormalized metric behavior.
    • Tests reference-cache validation and one-forward-per-chunk evaluation.
  • tests/brevitas_examples/llm/test_llm.py
    • Adds small-model integration coverage for EAR and KLD.
  • tests/brevitas_examples/llm/test_benchmark.py
    • Tests metric parsing and benchmark result propagation.

Local evaluation was run with the following config:

bos_preprocessing: sequence
dtype: bfloat16
eval: true
model: meta-llama/Llama-3.2-1B-Instruct
weight_bit_width: 4
weight_quant_granularity: per_channel
weight_quant_type: asym
weight_scale_precision: float_scale

Environment versions used for the evaluation:

  • Python: 3.12.13
  • PyTorch: 2.6.0+rocm6.1
  • Transformers: 5.16.1

Baseline

Results from master at 6df6a69:

Float perplexity (wikitext2): 12.154
Quantized perplexity (wikitext2): 21.806

This PR

Float perplexity (wikitext2): 12.154
Quantized perplexity (wikitext2): 21.806
Quantized expected acceptance rate (wikitext2): 0.635022
Quantized KL divergence (wikitext2): 0.687447

Using the above configuration with --weight-bit-width=8 we get:

Float perplexity (wikitext2): 12.154
Quantized perplexity (wikitext2): 12.161
Quantized expected acceptance rate (wikitext2): 0.978353
Quantized KL divergence (wikitext2): 0.002100

Below is a minimal evaluation of the overhead for collecting the additional metrics:

Float eval: current 30.026s vs master 29.050s — +0.976s (+3.4%)
Quant eval: current 14.941s vs master 14.561s — +0.380s (+2.6%)
Combined eval: current 44.966s vs master 43.611s — +1.355s (+3.1%)

Comment thread src/brevitas_examples/llm/llm_quant/eval.py Outdated
Comment thread src/brevitas_examples/llm/llm_quant/eval.py Outdated
@i-colbert
i-colbert marked this pull request as ready for review September 4, 2026 18:02
Comment thread src/brevitas_examples/llm/llm_quant/eval.py Outdated
@i-colbert
i-colbert requested a review from Giuseppe5 September 9, 2026 01:05
@i-colbert
i-colbert requested review from Giuseppe5 and removed request for Giuseppe5 September 10, 2026 00:22
Comment thread src/brevitas_examples/llm/llm_quant/eval.py

@pablomlago pablomlago left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Minor suggestions:

Comment thread src/brevitas_examples/llm/llm_quant/eval.py Outdated
Comment thread src/brevitas_examples/llm/llm_quant/eval.py Outdated
Comment thread src/brevitas_examples/llm/llm_quant/eval.py Outdated
self.dtype = dtype

@abstractmethod
def update(self, output: torch.Tensor, target: torch.Tensor) -> None:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I would maybe rename these methods, e.g. update to accumulate and finalize to aggregate, or something along those lines.



@torch.no_grad()
def compute_float_evaluation_metrics(

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should the method name reflect that it is doing some caching apart on top of computing metrics?

@@ -31,6 +31,11 @@ def parse_log(job_log: str) -> Dict[str, Any]:
# Find the line containing Quant PPL number

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe it would be worth extracting the common functionality, e.g.:

metric_logs = {"quant_ppl": r"Quantized perplexity \((.*?)\): (\d+\.\d+)", ...}
metrics = {metric_key: re.search(metric_value) for metric_key, metric_value in metric_logs.items()}

@Giuseppe5
Giuseppe5 merged commit d525232 into Xilinx:master Sep 24, 2026
25 of 27 checks passed
@i-colbert
i-colbert deleted the feat/more_metrics branch September 24, 2026 16:19
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants