[Presets] Benchmark dataset support - #4134
Open
peterschmidt85 wants to merge 1 commit into
Open
Conversation
`dataset` selects what every benchmark in a preset session measures: the synthetic `random` prompts shaped by `input_tokens` and `output_tokens`, a dataset the benchmark tool supports, or a Hugging Face dataset ID. A custom dataset provides the requests, so the request-shape properties can't be set with it, and the preset records the measured means instead. The dataset is part of the contract: it is written to the session constraints, the agent reports it with the benchmark, the preset records it, and `dstack preset` shows it. A session that doesn't set it renders exactly as before. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Dataset
By default, benchmarks use synthetic prompts shaped by
input_tokensandoutput_tokens. Setdatasetto benchmark on real text instead: a dataset the benchmark tool supports, or a Hugging Face dataset ID.The dataset provides the requests, so
input_tokens,output_tokens, andshared_prefix_tokenscan't be set with it, and the preset records the measured means. A gated dataset requiresHF_TOKENinenv.dstack presetshows the dataset in place of the request shape:🤖 Generated with Claude Code