Conversation
| for wq in _gguf_proxies(model): | ||
| prior_settings[wq] = ( | ||
| wq.cache_inference_quant_weight, wq.cache_inference_quant_weight_metadata_only) | ||
| wq.cache_inference_quant_weight = False |
There was a problem hiding this comment.
This should not be necessary, same below
There was a problem hiding this comment.
It may be necessary actually. The cache needs to be cleared to be updated (unless we also make a change to Brevitas core), and quant_inference_mode enables metadata-only weight caching by default without clearing the cache on exit. Clearing the weight cache avoid the possibility of having stale metadata.
pablomlago
left a comment
There was a problem hiding this comment.
Minor comments from my first pass:
| # TODO: Generalize this to have a map between GGUF quant type | ||
| # and our preprocessing for quant_modules | ||
| if data_qtype == gguf.GGMLQuantizationType.Q4_K: | ||
| def _modify_if_tensor(value): |
There was a problem hiding this comment.
I would change the name of this function (and that of modify_tensors) to something more self-explanatory.
There was a problem hiding this comment.
Thanks for the suggestion. modify_tensors is a general callback inherited from the llama.cpp conversion scripts. Each model can use it to rename, reshape, reorder, split, combine, or skip tensors.
A rename is therefore not a local change. It would also affect the inheritance and tests. I suggest keeping the established name, but welcome alternative function names.
…filename with export-path
da79dec to
e72b816
Compare
| hparams = ModelBase.load_hparams(tmp_work_dir) | ||
| model_architecture = hparams["architectures"][0] | ||
| try: | ||
| # TODO: every tensor now carries its own qtype via |
There was a problem hiding this comment.
When are you planning to address this? Before/after we merge this?
| d_wmin_m=scale_zp) | ||
| weight_quant = module.weight_quant | ||
|
|
||
| if isinstance(weight_quant, GGUFGroupwiseWeightQuantProxyFromInjector): |
There was a problem hiding this comment.
What happens when data_qtype is not defined? Does that happen for Q4_0?
| tmp_work_dir = Path(tempfile.mkdtemp(prefix='brevitas_gguf_export_')) | ||
| is_training = model.training | ||
| prior_gguf_cache_settings = dict() | ||
| try: |
There was a problem hiding this comment.
Why is there a try with no except? In general I prefer explicit if conditions if possible
Reason for this PR
This PR extends Brevitas GGUF support to add: (1) K-quants like Q2_K, Q3_K, etc, (2) nested scale
quantization, and (3) recipe-specific mixed precision. Users can now evaluate a GGUF recipe in
Brevitas and export the same quantized model.
Changes Made in this PR
and zero points. MSE scale search improves alignment with llama.cpp.
apply the required mixed-precision tensor layouts.
gguf:<type>export targets and target validation. The exporter reads cached quantizermetadata, preserves the selected GGUF tensor types, and accepts an exact output path.
src/brevitas_examples/papers/gguf.src/brevitas_examples/llm/gguf_export/quant.pyso thatq6_k_quant_blockonly handles packing, just like the other implementations.Testing Summary
The table compares Brevitas exports with fresh
llama-quantizeoutputs forLlama-3.2-1B-Instruct. Size values use decimal MB.
All 13 pairs have matching tensor names, quantization types, and shapes. File sizes differ by at
most 352 bytes because of GGUF metadata and header padding.
The reference flow uses an importance matrix when the llama.cpp recipe requires one. The Q5_K row
uses the llama.cpp Q5_K_S recipe. The Q4_K row uses Q4_K_S with
--pure --token-embedding-type q6_k. These options reproduce the Brevitas tensor layouts.Interoperability with Qronos was also verified. The table below reports
llama-perplexityresults for GGUF models exported by Brevitas. Round-to-nearest (RTN) provides a baseline where each weight is rounded to the nearest supported value on a fixed grid.The Qronos results above use the provided configurations for Llama-3.2-1B-Instruct and Llama-3.2-3B-Instruct.
Risk Highlight
behavior from llama.cpp. The recipe files include the applicable MIT license.
--quantize-first-last-layerreplaces--quantize-last-layer. Thesave_quantized_as_gguffunction signature also changes.scale_shift_zero_point_impltoParameterFromStatsFromParameterZeroPoint#1585 to support injecting a customscale_shift_zero_point_implintoParameterFromStatsFromParameterZeroPointFollow-up TODOs
InferenceManagerhandlers. This registry can provide a supported extension path for the GGUF handler.