Skip to content

Cache (4/5): Adopt inference caching in task runners - #124

Merged
ErlisLushtaku merged 13 commits into
mainfrom
cache-on-118/04-runners
Sep 22, 2026
Merged

ErlisLushtaku merged 13 commits into
mainfrom
cache-on-118/04-runners

Conversation

@ErlisLushtaku

@ErlisLushtaku ErlisLushtaku commented Sep 9, 2026 •

Copy link
Copy Markdown
Collaborator

Description

Adopts the inference cache in the task runners through --run.store_root.

  • Pairwise, meta-eval, MT-Bench, and Elo generation and judging share the do_inference cache.
  • Elo temperature calibration uses the same judgement cache.
  • Cache rows get instruction, model pair, and orientation metadata at the direct inference call.
  • Dataset-provided completions stay as direct inputs when they do not run inference.
  • Removes ignore_cache, cache_function_dataframe, and the old cache tokens.

This is stacked on #123.

@ErlisLushtaku
ErlisLushtaku force-pushed the cache-on-118/03-providers branch from df52a6d to e28ee43 Compare September 9, 2026 13:01
@ErlisLushtaku
ErlisLushtaku force-pushed the cache-on-118/04-runners branch from 269ccdf to 4f5d127 Compare September 9, 2026 13:01
@ErlisLushtaku ErlisLushtaku changed the title Cache: WIP (4/6) Adopt inference caching in task runners Cache: WIP (4/5) Adopt inference caching in task runners Sep 9, 2026
kargibora added a commit that referenced this pull request Sep 14, 2026
@ErlisLushtaku ErlisLushtaku changed the title Cache: WIP (4/5) Adopt inference caching in task runners Cache (4/5): Adopt inference caching in task runners Sep 15, 2026
@ErlisLushtaku
ErlisLushtaku marked this pull request as ready for review September 15, 2026 07:40

@kargibora kargibora left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall this PR looks quite clear. It wires all the things that have been implemented and removes the legacy code. It is good that we got rid of many lines that were making the code harder to parse!

use_tqdm=use_tqdm,
inference_cache=build_completion_cache(cfg),
**extra_kwargs,
)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nice that we get rid of this!

@kargibora kargibora left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I also did not understand the reason of cache_metadata. Do we need it in , do_inference(), annotate_battles() and judge_and_parse_prefs()?

@ErlisLushtaku

ErlisLushtaku commented Sep 19, 2026 •

Copy link
Copy Markdown
Collaborator Author

@kargibora Yes, it needs to reach all three in the current flow. It is not part of the cache key, but rather the info that doesn't necessarily affect determine the output e.g. instruction_id, model_a, orientation etc. do_inference uses it only when storing a miss, annotate_battles keeps it aligned with the rendered prompts, and judge_and_parse_prefs swaps model_a/model_b and orientation for the reversed prompt. I renamed it to cache_row_metadata to make this distinction clearer.

We could avoid passing a separate argument by creating an object such as:
InferenceInput(payload=prompt, cache_row_metadata=metadata)

What do you think?

@kargibora

Copy link
Copy Markdown
Member

Thanks for making it clear! I think that can work

@ErlisLushtaku
ErlisLushtaku changed the base branch from cache-on-118/03-providers to main September 22, 2026 22:32
@ErlisLushtaku
ErlisLushtaku merged commit df2c187 into main Sep 22, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants