Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -48,6 +48,12 @@ SEMANTIC_MODEL_PATH=sample_data/anime_streaming/semantic_model.yml
# always active; set this for table/column allowlists and LIMIT enforcement.
SQL_SECURITY_POLICY_PATH=

# Server-side registry of published data domains. When set, callers may pass a
# `domain_id` (REST/MCP/Gateway) and QueryForge resolves the database, semantic
# model, and SQL policy from this file instead of trusting caller-supplied paths.
# A missing file means "no domains published", which is not an error.
DOMAIN_REGISTRY_PATH=.queryforge/domains/registry.json

# Transport hardening for network deployments (REST / SSE / Gateway / MCP).
# When QUERYFORGE_API_KEY is set, every endpoint except /health requires
# `Authorization: Bearer <key>` or `X-API-Key: <key>`.
Expand Down
115 changes: 115 additions & 0 deletions .github/workflows/model-eval.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,115 @@
name: model-evaluation

# Tier 3 of the step-16 evaluation: a real model, real API spend, real latency.
#
# It is deliberately NOT part of push/PR CI:
# * the numbers are not comparable with the deterministic tier-1/2 gates, and
# mixing them would let a model flake look like a product regression;
# * it needs provider credentials and costs money per run.
# Scheduled and manual only, and it refuses to run without credentials so a
# missing secret is an explicit failure instead of a silent pass.

"on":
schedule:
# Sunday 10:00 Asia/Shanghai (02:00 UTC).
- cron: "0 2 * * 0"
workflow_dispatch:
inputs:
limit:
description: "Maximum number of gold cases to evaluate"
required: false
default: "40"
provider:
description: "Model provider (must match the configured secret)"
required: false
default: "openai"
model:
description: "Model name"
required: false
default: "gpt-4o-mini"

permissions:
contents: read

concurrency:
group: model-evaluation
cancel-in-progress: false

jobs:
model-evaluation:
runs-on: ubuntu-latest
timeout-minutes: 60
env:
LLM_PROVIDER: ${{ github.event.inputs.provider || 'openai' }}
LLM_MODEL: ${{ github.event.inputs.model || 'gpt-4o-mini' }}
LLM_API_KEY: ${{ secrets.LLM_API_KEY }}
LLM_BASE_URL: ${{ secrets.LLM_BASE_URL }}
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.12"
cache: pip
cache-dependency-path: pyproject.toml
- run: python -m pip install --upgrade pip
- run: python -m pip install -e .
- run: python -m queryforge --prepare-sample-data
- run: python sample/generate_aux_datasets.py
- name: Require provider credentials
run: |
if [ -z "${LLM_API_KEY}" ]; then
echo "LLM_API_KEY is not configured; tier 3 cannot run." >&2
echo "Add the repository secret or dispatch this workflow manually with credentials." >&2
exit 1
fi
- name: Evaluate the NL2SQL gold set with a real model
run: |
python scripts/evaluate_sql.py \
--cases evaluation/gold/nl2sql_multidomain.jsonl \
--limit "${{ github.event.inputs.limit || 40 }}" \
--output evaluation/reports/nl2sql_model_eval.json
- name: Tier-3 agent task report (recorded, not gating)
run: |
python scripts/benchmark_agent.py \
--tier 3 \
--provider "${LLM_PROVIDER}" \
--model "${LLM_MODEL}" \
--split regression \
--report evaluation/reports/agent_benchmark_tier3.json
- name: Publish summary
if: always()
run: |
python - <<'PY'
import json
import os
from pathlib import Path

summary = Path(os.environ["GITHUB_STEP_SUMMARY"])
lines = [
"# Tier-3 model evaluation",
"",
"Model numbers are reported separately from the deterministic tier-1/2",
"gates: a model flake must never read as a product regression.",
"",
]
for name in (
"evaluation/reports/nl2sql_model_eval.json",
"evaluation/reports/agent_benchmark_tier3.json",
):
path = Path(name)
if not path.is_file():
lines.append(f"- `{name}`: not produced")
continue
payload = json.loads(path.read_text(encoding="utf-8"))
lines.append(f"- `{name}`: produced")
metrics = payload.get("metrics") or {}
for key in sorted(metrics)[:12]:
lines.append(f" - {key}: `{metrics[key]}`")
summary.write_text("\n".join(lines) + "\n", encoding="utf-8")
PY
- uses: actions/upload-artifact@v4
if: always()
with:
name: model-evaluation-report
path: evaluation/reports/
if-no-files-found: warn
31 changes: 31 additions & 0 deletions .github/workflows/quality.yml
Original file line number Diff line number Diff line change
Expand Up @@ -28,6 +28,10 @@ jobs:
working-directory: web

offline-acceptance:
# Tier 1 of the step-16 evaluation: deterministic, offline, no model call.
# `scripts/run_acceptance.py --full` includes the agent task benchmark gate
# (`scripts/benchmark_agent.py --tier 1 --gate`), so a regression in task
# success, evidence coverage or tool legality fails this job.
runs-on: ubuntu-latest
timeout-minutes: 20
strategy:
Expand All @@ -45,6 +49,33 @@ jobs:
- run: python -m pip install -e .
- run: python scripts/run_acceptance.py --full

offline-acceptance-integration:
# Tier 2 of the step-16 evaluation: installs the optional transport
# integrations (REST API, MCP) and then *requires* them. A missing
# dependency fails the tier instead of silently skipping it (16-R1), which
# is why `benchmark_agent.py --tier 2 --gate` is part of this job.
runs-on: ubuntu-latest
timeout-minutes: 25
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.12"
cache: pip
cache-dependency-path: pyproject.toml
- run: python -m pip install --upgrade pip
- run: python -m pip install -e ".[all,duckdb]"
- run: python -m queryforge --prepare-sample-data
- run: python sample/generate_aux_datasets.py
- run: python -m unittest discover -s tests -q
- run: python scripts/benchmark_agent.py --tier 2 --gate
- run: python docs/demo/run_all.py
- uses: actions/upload-artifact@v4
if: always()
with:
name: integration-reports
path: evaluation/reports/

package:
runs-on: ubuntu-latest
timeout-minutes: 10
Expand Down
3 changes: 3 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -24,3 +24,6 @@ Thumbs.db
*.sqlite-wal
*.build.json
semantic-weekly-report.json
# Generated benchmark reports (the summarized evidence lives in
# docs/optimization/step-16-acceptance.md; CI uploads its own artifacts).
evaluation/reports/
6 changes: 5 additions & 1 deletion Makefile
Original file line number Diff line number Diff line change
@@ -1,11 +1,12 @@
.PHONY: help install install-all test acceptance check sample semantic-check package web-install web-dev web-check web-build
.PHONY: help install demo install-all test acceptance check sample semantic-check package web-install web-dev web-check web-build

PYTHON ?= python

help:
@echo "install Install QueryForge in editable mode"
@echo "install-all Install all optional integrations"
@echo "test Run the complete unittest suite"
@echo "demo Run the step-17 end-to-end acceptance demos"
@echo "acceptance Run offline acceptance checks"
@echo "check Run repository hygiene and full acceptance"
@echo "sample Validate the bundled anime dataset"
Expand All @@ -25,6 +26,9 @@ install-all:
test:
LOG_LEVEL=CRITICAL $(PYTHON) -m unittest discover -s tests -q

demo:
$(PYTHON) docs/demo/run_all.py

acceptance:
$(PYTHON) scripts/run_acceptance.py --full

Expand Down
31 changes: 29 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,7 @@ policy enforcement, bounded recovery, and production-friendly delivery interface
![Python](https://img.shields.io/badge/Python-3.11%20%7C%203.12-3776AB?logo=python&logoColor=white)
![SQLite](https://img.shields.io/badge/SQLite-read--only-003B57?logo=sqlite&logoColor=white)
![SQLGlot](https://img.shields.io/badge/SQL%20policy-SQLGlot-6B4FBB)
![Tests](https://img.shields.io/badge/tests-270%20passing-2EA44F)
![Tests](https://img.shields.io/badge/tests-691%20passing-2EA44F)
![Semantic contracts](https://img.shields.io/badge/semantic%20checks-82%20passing-7C3AED)

</div>
Expand Down Expand Up @@ -429,7 +429,34 @@ python scripts/evaluate_sql.py \
--output .queryforge/evaluations/openai.json
```

CI runs the offline acceptance gate on Python 3.11 and 3.12.
CI runs the offline acceptance gate (including the deterministic agent benchmark) on Python 3.11 and 3.12, plus an integration job that requires the optional transport dependencies.

## What is verified (and what is not)

Every claim in this section is reproducible from the repository; the linked
acceptance record contains the gaps as well as the passes.

| Capability | How you can check it | Status |
| --- | --- | --- |
| Full offline test suite | `make test` — **806 tests, 0 skipped** | verified |
| Repository + integration gate | `make check` (`scripts/run_acceptance.py --full`, 13/13 checks) | verified |
| End-to-end demos (upload → publish → query; semantic catch; multi-step analysis; transports/refusal/recovery) | `make demo` — four narrated, asserting scripts under `docs/demo/` | verified |
| Deterministic agent benchmark (32 gold tasks, 3 independent schemas, ablation, effect gate) | `python scripts/benchmark_agent.py --tier 1 --gate` | verified (32/32) |
| Optional-dependency integration tier | `python scripts/benchmark_agent.py --tier 2 --gate` — a missing dependency **fails** the tier | verified with `.[api,mcp]` installed |
| Real-model NL2SQL evaluation | `python scripts/evaluate_sql.py --cases evaluation/gold/nl2sql_multidomain.jsonl --model-provider <p> --model <m>` | **not run here** — no numbers, no accuracy claim |

Demo output is offline and deterministic (no model call, no network, no API key).
The agent benchmark's tier 1 gives the SQL as a fixture, so its 32/32 measures the
*engineering* chain (governance, execution, evidence, budget, failure
classification) — **not model accuracy**. Real-model numbers must come from a
tier-3 run with credentials and are reported separately
(`.github/workflows/model-eval.yml`).

Deployment level: **controlled environment, single tenant, read-only data access**.
SQLite is the default backend; a DuckDB adapter exists behind an optional extra
(see [Database adapters](docs/database_adapters.md)). The system is not hardened
for arbitrary untrusted multi-tenant input, and the known gaps are listed per
per capability in the docs listed above; the two honest blank spots are real-model evaluation (no accuracy numbers) and the PostgreSQL backend (implemented, not yet verified against a live server).

## Documentation

Expand Down
23 changes: 22 additions & 1 deletion README.zh-CN.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,7 @@
![Python](https://img.shields.io/badge/Python-3.11%20%7C%203.12-3776AB?logo=python&logoColor=white)
![SQLite](https://img.shields.io/badge/SQLite-只读执行-003B57?logo=sqlite&logoColor=white)
![SQLGlot](https://img.shields.io/badge/SQL%20治理-SQLGlot-6B4FBB)
![Tests](https://img.shields.io/badge/tests-270%20passing-2EA44F)
![Tests](https://img.shields.io/badge/tests-691%20passing-2EA44F)
![Semantic contracts](https://img.shields.io/badge/semantic%20checks-82%20passing-7C3AED)

</div>
Expand Down Expand Up @@ -409,6 +409,27 @@ python scripts/evaluate_sql.py \

CI 会在 Python 3.11 和 3.12 上执行离线验收。

## 已验证的能力(以及未验证的部分)

本节每条声明都能从仓库复现;对应验收记录里同时写着通过与缺口。

| 能力 | 怎么验证 | 状态 |
| --- | --- | --- |
| 完整离线测试套件 | `make test` —— **806 个测试,0 skip** | 已验证 |
| 仓库 + 集成门禁 | `make check`(`scripts/run_acceptance.py --full`,13/13 项通过) | 已验证 |
| 端到端 Demo(上传→发布→查询;语义校验抓错;多步分析;跨传输/拒绝/恢复) | `make demo` —— `docs/demo/` 下四个带断言的叙事脚本 | 已验证 |
| 确定性 Agent Benchmark(32 个金标任务、3 个独立 schema、消融、效果门禁) | `python scripts/benchmark_agent.py --tier 1 --gate` | 已验证(32/32) |
| 可选依赖集成层 | `python scripts/benchmark_agent.py --tier 2 --gate` —— 依赖缺失**判定失败**而非跳过 | 已装 `.[api,mcp]` 后通过 |
| 真实模型 NL2SQL 评测 | `python scripts/evaluate_sql.py --cases evaluation/gold/nl2sql_multidomain.jsonl --model-provider <p> --model <m>` | **本机未跑**——没有数字,因此不声称准确率 |

Demo 全部离线、确定性(无模型调用、无网络、无需 API key)。Agent Benchmark 的 tier 1 由金标提供 SQL,
所以 32/32 衡量的是**工程链路**(治理、执行、证据、预算、失败分类),**不是模型准确率**。
真实模型数字必须来自带凭证的 tier 3 运行,并单独报告(`.github/workflows/model-eval.yml`)。

部署等级:**受控环境、单租户、只读数据访问**。默认后端为 SQLite,另有可选的 DuckDB 适配器
(见 [数据库适配器](docs/database_adapters.md))。系统未针对任意不可信的多租户输入做加固,
逐条记在上方对应能力的文档中;两处诚实的空白是:真实模型评测(无准确率数字)与 PostgreSQL 后端(已实现、未在真实服务器上验证)。

## 项目文档

| 主题 | 文档 |
Expand Down
3 changes: 3 additions & 0 deletions docs/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,3 +20,6 @@
- [GitHub release checklist](github_release.md)
- [Report artifacts](report_artifact.md)
- [NL2SQL evaluation](nl2sql_evaluation.md)
- [Agent benchmark, ablation and effect gates](database_adapters.md)
- [End-to-end acceptance demos](demo/README.md)
- [Database adapters (SQLite default, DuckDB optional)](database_adapters.md)
Loading
Loading