Skip to content

Update BrowserCode model comparison plot - #23

Merged
Alezander9 merged 1 commit into
mainfrom
refresh-bcode-harness-plot
Aug 4, 2026
Merged

Update BrowserCode model comparison plot#23
Alezander9 merged 1 commit into
mainfrom
refresh-bcode-harness-plot

Conversation

@Alezander9

@Alezander9 Alezander9 commented Aug 4, 2026

Copy link
Copy Markdown
Member

What

Resyncs official_plots/browser_harness_by_model_{light,dark}.png (the
"Comparing Models for BrowserCode" plot) from
benchmark-x-laminar@45b563b.

The committed copies were generated 2026-06-17 and are the stalest published
copy of this plot anywhere. They still show glm-5.1 (79.7%) as the top
open-weight model and predate every model added since:

  • Added: claude-opus-5 (87.0%), claude-fable-5 (87.0%), grok-4.5 (86.3%), kimi-k3 (86.0%), gpt-5.6-sol (84.0%), gpt-5.6-luna xhigh (82.0%) / default (72.4%), gpt-5.6-terra (72.0%), qwen3.8-max (79.0%), minimax-m3 (74.0%), deepseek-v4-flash-0731 (76.0%)
  • Dropped: mimo-v2.5/-pro, glm-4.7, glm-5.1, kimi-k2.6, qwen3.6-plus, grok-4.3, gpt-5.4-mini, claude-sonnet-4-6, base deepseek-v4-flash
  • Open-weight bars changed from a solid blue fill to a gray fill with a blue dashed outline

Image-only change. generate_plots.py in this repo produces accuracy_by_model
and accuracy_vs_throughput; the three BU Bench plots in the README are
imported copies from benchmark-x-laminar, so nothing here regenerates them.

Companion PR for the same plot in the BrowserCode README: browser-use/browsercode#143

Not in this PR

browser_use_framework_by_model_{light,dark}.png has also drifted from
upstream, which now has two extra bars (qwen3.627b 29.0%, glm-5.1 25.0%);
the top eight are unchanged. Left out because it is a browser-use-framework
plot and unaffected by the BrowserCode results this PR propagates -- say the
word and I will fold it in.

best_of_frameworks_public_dark.png differs from upstream only in PNG
metadata; the rendered content is byte-identical to the light version's
content. Not worth a churn commit.


Summary by cubic

Refreshes the BrowserCode model comparison plots official_plots/browser_harness_by_model_{light,dark}.png from benchmark-x-laminar@45b563b to replace the stale 2026-06-17 images. The plots now include new models (e.g., claude-opus-5, claude-fable-5, grok-4.5, kimi-k3, gpt-5.6-*, qwen3.8-max, deepseek-v4-flash-0731), remove dropped ones, and change open-weight bars to a gray fill with a blue dashed outline.

Written for commit 13bac02. Summary will update on new commits.

Review in cubic

Resyncs official_plots/browser_harness_by_model_{light,dark}.png from
benchmark-x-laminar@45b563b. The committed copies were generated 2026-06-17
and are the stalest of any published copy -- they still show glm-5.1 (79.7%)
as the top open-weight model and predate claude-opus-5, claude-fable-5,
grok-4.5, kimi-k3, the gpt-5.6 family, qwen3.8-max and
deepseek-v4-flash-0731.

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No issues found across 2 files

Re-trigger cubic

@Alezander9
Alezander9 merged commit 6c4efd8 into main Aug 4, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant