FreeCite: A Judge-Free Benchmark for Granular Citation Evaluation in Large Language Models
-
Updated
Feb 22, 2026 - Python
FreeCite: A Judge-Free Benchmark for Granular Citation Evaluation in Large Language Models
Research benchmark for evidence-grounded OS-agent collaboration, continuous state diagnosis, scoped memory reuse, and stale-state rejection.
Canonical specification and sealed execution traces for VERITAS Ω-CODE v2.0 — deterministic software verification protocol.
As of 07/09/2026 the only open-source PT-PT benchmark that avoids LLM judges and targets email-style structured generation. Deterministic CPU-friendly benchmark for European Portuguese (PT-PT) email LLMs. No LLM judges.
A realistic RL environment for training LLM agents on enterprise email triage—featuring multi-step decision making, ambiguity handling, tool usage, and deterministic evaluation.
AI-assisted policy formalization and governance platform for source-traceable rule extraction, human review, immutable versioning, deterministic evaluation, quality analysis, and regression testing.
Auditable address-level rental housing law screening with verified property facts, deterministic three-valued evaluation, and source-backed change tracking.
HCAC is a completion contract for agents and tools. It defines when something is truly “done” — not just structurally, but behaviorally.
Deterministic multi-lane evaluation domain for SnorkelAI Terminus3. Cockpit, lane orchestration, agent benchmarking, reasoning-drift detection.
Public, fully local PoCs for counterfactually auditable lifecycle certification: exact paired replay, drift monitoring, post-drift replanning, and bridge-aware ledger control on synthetic tasks.
Ultra pipeline framework for Gravity Binary, including deterministic evaluation logic, capsule workflows, and EC/Ultra integration artifacts.
SnorkelAI Terminus2 task suite with deterministic evaluators, Dockerized environments, and reproducible agent-benchmarking primitives.
Core deterministic architecture for multi-lane ML evaluation systems. Foundational logic, capsule interfaces, reproducible execution patterns.
Deterministic offline ComtradeBench judge for evaluating agent robustness under pagination, retries, duplicates, page drift, and totals traps.
To associate your repository with the deterministic-evaluation topic, visit your repo's landing page and select "manage topics."