-
Notifications
You must be signed in to change notification settings - Fork 603
[OMNIML-5899] Add IQ1_S and IQ2_XS quantization and unified checkpoint export #2381
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Closed
Closed
Changes from all commits
Commits
Show all changes
8 commits
Select commit
Hold shift + click to select a range
5a075fb
Add IQ1_S and IQ2_XS unified checkpoint export
ChenhanYu 1e3a537
Organize IQ codecs under GGML package
ChenhanYu 524356e
Support IQ export from Megatron
ChenhanYu 3563cdd
Document GGML FP8 compatibility limitation
ChenhanYu 7019f2c
Format IQ1_S CUDA kernel
ChenhanYu f0850ee
Fix IQ2_XS API documentation markup
ChenhanYu cd3e661
Use canonical weight keys for IQ payloads
ChenhanYu 73baffe
Preserve packed fused expert weights during export
ChenhanYu File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Oops, something went wrong.
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
[SUGGESTION] This branch hardcodes
"group_size": 256and silently ignores thegroup_sizeargument (gs) that every other branch honours. That's correct for GGML — the 256-value block is baked into the 50/74-byte layout — but a caller who passesgroup_size=128gets a config claiming 256 with no warning, and the mismatch surfaces only when a consumer tries to unpack. Worth asserting the contract instead of dropping the argument:Also
"type": "int"with"num_bits": 1is a lossy description of the format for anything reading this generically — the values are ternary{-1,0,1}(IQ1_S) / magnitude-grid codes (IQ2_XS) with a shared fp16 block scale, not 1-bit integers.effective_bitsandpacking: "ggml"carry the real information; consider a comment notingnum_bitsis nominal here so a future reader doesn't try to derive the size from it.