Skip to content
#

goodput

Here are 4 public repositories matching this topic...

Measurement harness for the sliding window attention premium in the vLLM TPU Ragged Paged Attention v3 kernel: per layer decode cost, block size control, throughput, and goodput for Gemma 4 31B on TPU v6e.

  • Updated Jul 29, 2026
  • Python

Simulates LLM serving admission control strategies (None, MaxConcurrency, TokenBudget, SLOAware, Predictive) across 135 configurations. Key findings: tight token budget (4096) achieves 10x better goodput than loose budget; predictive control is the only strategy maintaining SLO compliance at 100ms; SLOAware is fundamentally unstable under sustained

  • Updated Jul 10, 2026
  • Python

Improve this page

Add a description, image, and links to the goodput topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the goodput topic, visit your repo's landing page and select "manage topics."

Learn more