Fix benchmark defects and flatten the benchmark package - #72
Merged
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This fixes ten defects in the benchmark package and removes about 2000 lines from it: exported source labels that carry paths no longer split into nested directories, the tiling sources report their capacity as the known optimum instead of rebuilding the ground truth (and no longer return None without a seed), the subset sync pattern merges at least two threads, ONNX models take their id from the graph name and no longer inflate zero-size tensors, a campaign id of 0 is kept and the default save path no longer doubles its prefix, the Hugging Face variant hook no longer mutates the source mid-campaign, concatenated Hugging Face models get prefixed ids through a fixed-source base shared with Minimalloc, progress bars show the resolved allocator name, and PowerOf2Source checks its duration bound. run_benchmark now loops source, variant, allocator and records every skip in one list, the results package drops the double-built metadata and nested writers, Timer is a plain context manager, the Hugging Face tests build tiny local models instead of downloading 91 MB per test, and the examples are trimmed to minimal showcases of the API.