Skip to content

Feat (gguf): broaden format support and reproduce canonical recipes - #1537

Open
Giuseppe5 wants to merge 49 commits into
Xilinx:masterfrom
Giuseppe5:gguf_rework
Open

Giuseppe5 wants to merge 49 commits into
Xilinx:masterfrom
Giuseppe5:gguf_rework

Conversation

@Giuseppe5

@Giuseppe5 Giuseppe5 commented Jun 29, 2026 •

Copy link
Copy Markdown
Collaborator

Reason for this PR

This PR extends Brevitas GGUF support to add: (1) K-quants like Q2_K, Q3_K, etc, (2) nested scale
quantization, and (3) recipe-specific mixed precision. Users can now evaluate a GGUF recipe in
Brevitas and export the same quantized model.

Changes Made in this PR

  • Added GGUF quantizers from Q2_K through Q8_0. The K-quant implementations model nested scales
    and zero points. MSE scale search improves alignment with llama.cpp.
  • Added registered quantizer plugins and file-based recipe plugins. Llama 3.2 1B and 3B recipes
    apply the required mixed-precision tensor layouts.
  • Added gguf:<type> export targets and target validation. The exporter reads cached quantizer
    metadata, preserves the selected GGUF tensor types, and accepts an exact output path.
  • Added GGUF usage, configuration, and validation documentation in src/brevitas_examples/papers/gguf.
  • Removed Q6_K quantizer from src/brevitas_examples/llm/gguf_export/quant.py so that q6_k_quant_block only handles packing, just like the other implementations.

Testing Summary

The table compares Brevitas exports with fresh llama-quantize outputs for
Llama-3.2-1B-Instruct. Size values use decimal MB.

Recipe Brevitas Size (MB) Ref. Size (MB)
Q8_0 1321.1 1321.1
Q6_K 1021.8 1021.8
Q5_K_M 911.5 911.5
Q5_K 892.6 892.6
Q4_1 831.7 831.7
Q4_K_M 807.7 807.7
Q4_K_S 775.6 775.6
Q4_0 773.0 773.0
Q4_K 770.9 770.9
Q3_K_L 732.5 732.5
Q3_K_M 690.8 690.8
Q3_K_S 641.7 641.7
Q2_K 580.9 580.9

All 13 pairs have matching tensor names, quantization types, and shapes. File sizes differ by at
most 352 bytes because of GGUF metadata and header padding.

The reference flow uses an importance matrix when the llama.cpp recipe requires one. The Q5_K row
uses the llama.cpp Q5_K_S recipe. The Q4_K row uses Q4_K_S with
--pure --token-embedding-type q6_k. These options reproduce the Brevitas tensor layouts.

Interoperability with Qronos was also verified. The table below reports llama-perplexity results for GGUF models exported by Brevitas. Round-to-nearest (RTN) provides a baseline where each weight is rounded to the nearest supported value on a fixed grid.

Recipe Algorithm 1B 3B
Q2_K RTN 30.60 15.04
Q2_K Qronos 18.17 12.60

The Qronos results above use the provided configurations for Llama-3.2-1B-Instruct and Llama-3.2-3B-Instruct.

Risk Highlight

  • This PR includes code from another work. The GGUF layouts and mixed-precision rules adapt
    behavior from llama.cpp. The recipe files include the applicable MIT license.
  • This PR contains API-breaking changes. --quantize-first-last-layer replaces
    --quantize-last-layer. The save_quantized_as_gguf function signature also changes.
  • This PR depended on work in PR Feat (core): adding scale_shift_zero_point_impl to ParameterFromStatsFromParameterZeroPoint #1585 to support injecting a custom scale_shift_zero_point_impl into ParameterFromStatsFromParameterZeroPoint

Follow-up TODOs

  1. Generalize the nested quantizer cache in Brevitas core, either by adding a new nested cache or a new tensor subclass. This would render the out-of-source caching infrastructure redundant.
  2. Add a registry for out-of-source InferenceManager handlers. This registry can provide a supported extension path for the GGUF handler.
  3. Add activation quantization emulation. The common configuration uses dynamic 8-bit quantization, but activation quantization formats can vary by backend.

Comment thread src/brevitas_examples/llm/gguf_export/convert.py Outdated
Comment thread src/brevitas_examples/llm/gguf_export/base_quantizers.py Outdated
Comment thread src/brevitas_examples/llm/gguf_export/base_quantizers.py Outdated
Comment thread src/brevitas_examples/llm/gguf_export/base_quantizers.py Outdated
Comment thread src/brevitas_examples/llm/gguf_export/base_quantizers.py
Comment thread src/brevitas_examples/llm/gguf_export/base_quantizers.py
for wq in _gguf_proxies(model):
prior_settings[wq] = (
wq.cache_inference_quant_weight, wq.cache_inference_quant_weight_metadata_only)
wq.cache_inference_quant_weight = False

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This should not be necessary, same below

@i-colbert i-colbert Aug 17, 2026 •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It may be necessary actually. The cache needs to be cleared to be updated (unless we also make a change to Brevitas core), and quant_inference_mode enables metadata-only weight caching by default without clearing the cache on exit. Clearing the weight cache avoid the possibility of having stale metadata.

@i-colbert i-colbert changed the title GGUF rework Feat (gguf): broaden format support and reproduce canonical recipes Aug 14, 2026
@i-colbert
i-colbert marked this pull request as ready for review August 18, 2026 00:35
@i-colbert
i-colbert changed the base branch from dev to master August 18, 2026 00:45
Comment thread src/brevitas_examples/llm/main.py Outdated
@i-colbert
i-colbert self-requested a review August 24, 2026 17:26
@i-colbert
i-colbert requested review from i-colbert and removed request for i-colbert August 25, 2026 01:52
@Giuseppe5 Giuseppe5 self-assigned this Aug 26, 2026

@pablomlago pablomlago left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Minor comments from my first pass:

Comment thread src/brevitas_examples/llm/gguf_export/base_quantizers.py Outdated
Comment thread src/brevitas_examples/llm/gguf_export/base_quantizers.py Outdated
Comment thread src/brevitas_examples/llm/gguf_export/base_quantizers.py
Comment thread src/brevitas_examples/llm/gguf_export/base_quantizers.py Outdated
Comment thread src/brevitas_examples/llm/gguf_export/base_quantizers.py Outdated
# TODO: Generalize this to have a map between GGUF quant type
# and our preprocessing for quant_modules
if data_qtype == gguf.GGMLQuantizationType.Q4_K:
def _modify_if_tensor(value):

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I would change the name of this function (and that of modify_tensors) to something more self-explanatory.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the suggestion. modify_tensors is a general callback inherited from the llama.cpp conversion scripts. Each model can use it to rename, reshape, reorder, split, combine, or skip tensors.

A rename is therefore not a local change. It would also affect the inheritance and tests. I suggest keeping the established name, but welcome alternative function names.

hparams = ModelBase.load_hparams(tmp_work_dir)
model_architecture = hparams["architectures"][0]
try:
# TODO: every tensor now carries its own qtype via

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

When are you planning to address this? Before/after we merge this?

d_wmin_m=scale_zp)
weight_quant = module.weight_quant

if isinstance(weight_quant, GGUFGroupwiseWeightQuantProxyFromInjector):

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What happens when data_qtype is not defined? Does that happen for Q4_0?

tmp_work_dir = Path(tempfile.mkdtemp(prefix='brevitas_gguf_export_'))
is_training = model.training
prior_gguf_cache_settings = dict()
try:

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why is there a try with no except? In general I prefer explicit if conditions if possible

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants