Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 3 additions & 1 deletion .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -55,7 +55,9 @@ SQL_SECURITY_POLICY_PATH=
DOMAIN_REGISTRY_PATH=.queryforge/domains/registry.json

# Transport hardening for network deployments (REST / SSE / Gateway / MCP).
# When QUERYFORGE_API_KEY is set, every endpoint except /health requires
# When QUERYFORGE_API_KEY is set, every endpoint except the four public
# paths requires Authorization: Bearer <key> or X-API-Key:
# /health, /openapi.json, /docs, /redoc
# `Authorization: Bearer <key>` or `X-API-Key: <key>`.
QUERYFORGE_API_KEY=
# Comma-separated files/directories that remote callers may open as `database`
Expand Down
48 changes: 45 additions & 3 deletions .github/workflows/model-eval.yml
Original file line number Diff line number Diff line change
Expand Up @@ -38,10 +38,32 @@ jobs:
runs-on: ubuntu-latest
timeout-minutes: 60
env:
# ``load_config`` only reads the *provider-specific* variables declared in
# ``models.yml`` (openai -> OPENAI_API_KEY / OPENAI_MODEL /
# OPENAI_BASE_URL, and so on). A generic ``LLM_API_KEY`` is never consulted
# by the loader, so subscribing only that name let the credential gate below
# pass while every real model call failed with "No API key is configured
# for provider 'openai'". Every provider therefore maps its own secret here.
#
# All of them are populated from the same optional secret so a single
# repository secret can drive whichever provider is dispatched; a provider
# whose variable stays empty simply fails at the gate below.
#
# Only the api-key variables are fanned out. The model/base-url variables
# are deliberately NOT set here: ``--model-provider``/``--model`` are passed
# explicitly to both evaluation steps, so a fanned-out ``*_MODEL`` would be
# a second, contradicting source of truth.
LLM_PROVIDER: ${{ github.event.inputs.provider || 'openai' }}
LLM_MODEL: ${{ github.event.inputs.model || 'gpt-4o-mini' }}
LLM_API_KEY: ${{ secrets.LLM_API_KEY }}
LLM_BASE_URL: ${{ secrets.LLM_BASE_URL }}
OPENAI_API_KEY: ${{ secrets.LLM_API_KEY }}
ANTHROPIC_API_KEY: ${{ secrets.LLM_API_KEY }}
GEMINI_API_KEY: ${{ secrets.LLM_API_KEY }}
DEEPSEEK_API_KEY: ${{ secrets.LLM_API_KEY }}
QWEN_API_KEY: ${{ secrets.LLM_API_KEY }}
GLM_API_KEY: ${{ secrets.LLM_API_KEY }}
OPENAI_BASE_URL: ${{ secrets.LLM_BASE_URL }}
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
Expand All @@ -55,16 +77,36 @@ jobs:
- run: python sample/generate_aux_datasets.py
- name: Require provider credentials
run: |
if [ -z "${LLM_API_KEY}" ]; then
echo "LLM_API_KEY is not configured; tier 3 cannot run." >&2
echo "Add the repository secret or dispatch this workflow manually with credentials." >&2
# Check the variable ``load_config`` actually reads for the dispatched
# provider — not the generic LLM_API_KEY, which the loader ignores.
case "${LLM_PROVIDER}" in
openai) key="${OPENAI_API_KEY}" ;;
claude) key="${ANTHROPIC_API_KEY}" ;;
gemini) key="${GEMINI_API_KEY}" ;;
deepseek) key="${DEEPSEEK_API_KEY}" ;;
qwen) key="${QWEN_API_KEY:-${DASHSCOPE_API_KEY}}" ;;
glm) key="${GLM_API_KEY:-${ZAI_API_KEY}}" ;;
*)
echo "Unsupported provider '${LLM_PROVIDER}' for this workflow." >&2
exit 1
;;
esac
if [ -z "${key}" ]; then
echo "No API key is configured for provider '${LLM_PROVIDER}'; tier 3 cannot run." >&2
echo "Add the repository secret LLM_API_KEY (mapped to the provider-specific variable)" >&2
echo "or dispatch this workflow manually with credentials." >&2
exit 1
fi
- name: Evaluate the NL2SQL gold set with a real model
run: |
# ``--model-provider`` / ``--model`` are passed explicitly: without them
# this step silently ignored the dispatched ``model`` input and used the
# provider default from models.yml instead.
python scripts/evaluate_sql.py \
--cases evaluation/gold/nl2sql_multidomain.jsonl \
--limit "${{ github.event.inputs.limit || 40 }}" \
--model-provider "${LLM_PROVIDER}" \
--model "${LLM_MODEL}" \
--output evaluation/reports/nl2sql_model_eval.json
- name: Tier-3 agent task report (recorded, not gating)
run: |
Expand Down
6 changes: 4 additions & 2 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -24,6 +24,8 @@ Thumbs.db
*.sqlite-wal
*.build.json
semantic-weekly-report.json
# Generated benchmark reports (the summarized evidence lives in
# docs/optimization/step-16-acceptance.md; CI uploads its own artifacts).
# Generated benchmark reports: raw JSON is regenerable, so it is not tracked.
# The summarized, versioned evidence lives in docs/optimization/baselines.md
# (the comment here previously pointed at docs/optimization/step-16-acceptance.md,
# which does not exist in the repository).
evaluation/reports/
136 changes: 136 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,136 @@
# AGENTS.md

Project harness for agent-assisted development on **QueryForge** — a governed
NL2SQL / data-analysis platform (Python 3.11/3.12 + Next.js Studio).

Keep this file short. Project facts live in `docs/`; this file is routing and invariants.

## Startup Workflow

Before writing code:

1. **Confirm the working directory is the repository root** with `pwd` — the directory
containing `pyproject.toml` and this file
2. **Read this file** completely
3. **Run `./init.sh`** — checks the environment, runs the full offline test suite, reports
repository state. Exits non-zero on failure.
4. **Read the routed docs** for the area you are touching (see "Where Facts Live")

If baseline verification fails, **repair that first**. Do not add scope on top of a red baseline.

## Invariants — Do Not Violate

These are load-bearing. Any change that weakens them is a regression, not a refactor.

1. **Single SQL execution boundary.** `DatabaseTool` is the only sanctioned SQL execution
path. `DatabaseAdapter.execute_sql` is a *trusted primitive* that performs no policy
check — never call it directly from application code.
2. **Deterministic gates outrank models.** Final truth for these is decided by code, never by
an LLM verdict:
- `SQLPolicyEngine` (`queryforge/domain/security/sql_policy.py`) — AST policy
- `SemanticSQLValidator` (`queryforge/domain/semantic/sql_validator.py`) — business semantics
- `EvidenceStore` (`queryforge/domain/analysis/evidence.py`) — evidence ids
- `QualityGateEvaluator` (`queryforge/orchestration/gates.py`) — phase gates
- Terminal run status (`queryforge/core/outcomes.py`)
A model may *propose*. It may not declare success.
3. **Fail closed.** Unknown table/column, unsupported shape, or unverifiable claim must
surface as `unsupported` / `blocked` / `unverified` — never silently as `success`.
4. **Read-only by default.** Analysis opens the database read-only; write/admin SQL is
rejected at the AST layer *and* at the engine layer.

## Working Rules

- **No completion claim without evidence.** Run the relevant verification command and report
its actual output. "Should work" is not evidence.
- **Stay in scope.** Do not opportunistically refactor unrelated modules.
- **Prefer existing mechanisms over new ones.** Tool permissions, `PlanValidator`, the
ablation switches, the evidence layer and journal/resume already exist — check before
adding a mechanism.
- **Leave the repo verifiable.** `./init.sh` must pass when you stop.

## Verification Commands

```bash
./init.sh # full: environment + test suite + repo state

# Individual checks (from the venv)
LOG_LEVEL=CRITICAL .venv/bin/python -m unittest discover -s tests -q
make check # repository hygiene / required files + offline acceptance
make acceptance # offline acceptance only

# Evaluation (tier-1 is offline; tier-3 costs real API spend)
LOG_LEVEL=CRITICAL .venv/bin/python -B scripts/benchmark_agent.py --tier 1 --report /tmp/tier1.json
```

**Static and build checks** — `./init.sh` does not run these; run the relevant one when you
touch that area:

```bash
make check # repository hygiene (check_repository.py) + offline acceptance
make web-check # Studio: eslint + tsc --noEmit + build + node --test (TypeScript)
make web-build # Studio production build
make semantic-check # semantic-model drift against the sample database
```

> **Tier-3 requires credentials and a funded account — ask the human before running it.**
> It spends real money. See "Known Traps" below.

## Known Traps

Each of these has already cost time or produced a wrong conclusion. Read before touching
the relevant area.

1. **`evaluation/reports/` is gitignored.** Raw evaluator JSON there is regenerable and
untracked. Any number you want to cite must be written into a tracked document, with the
command that produced it and the versions it depends on.
2. **Tier-1 "all green" does not measure model capability.** `scripts/benchmark_runners.py`
swaps in `ScriptedModel`, which returns `reference_sql` verbatim. Tier-1 measures the
governance pipeline and control flow only.
3. **Tier-3 needs a funded account.** A depleted balance returns HTTP 402. Such cases are
classified as `environment_error` and excluded from the accuracy denominators, but the run
still cannot complete. Confirm balance before a baseline run.
4. **A candidate/no-candidate comparison is confounded unless you force the count.** The gold
set's per-case `candidate_selection` flag correlates perfectly with category
(`multi_table`/`metric` set it; `single_table`/`time` never do), so grouping by it measures
task difficulty, not the mechanism. Use `--parallel-candidates N` to hold inputs fixed.
5. **Provider credentials are provider-specific.** `load_config` reads `DEEPSEEK_API_KEY` /
`OPENAI_API_KEY` / … as declared in `models.yml`. A generic `LLM_API_KEY` is silently
ignored. `.env` is gitignored — never commit it.
6. **`ScriptedModel` vs real model, in the same runner.** `runner: "scripted_workflow"` tasks
call the real model when a provider is configured (tier-3) and the fake one otherwise.
Do not conclude from the runner name alone.
7. **Do not trust a measurement without reproducing its verdict against the data.** Several
"semantic errors" in the first baseline were the evaluator's fault, not the model's, and a
"parallel candidates are 3× slower" finding was pure difficulty confounding. Verify the
per-case evidence before quoting an aggregate.
8. **Check what the evaluator is *suppressing*, not just what it measures.** It once passed
`skills=[]` for every case, which takes the manual skill path and silently disables
automatic skill selection — so the catalogue was inert in every measured run and the
headline accuracy described a configuration that does not match production. Use
`--skill-mode auto` (the default) for a production-aligned number.
9. **A run that fails in ~60 ms with zero measured tokens never called the model.** That is
the signature of a provider-contract mismatch (for example a wrapper whose
`generate_with_messages` lags the adapter signature), not of bad model output.
10. **A single run of ~30 cases cannot resolve small accuracy differences.** Deltas of ±0.03
have been observed across *identical* code. Do not present such a difference as an
improvement; increase `--repeat` instead.

## Escalation

- **Architecture decisions** → read `docs/agent_team_architecture.md` and `docs/README.md`, then ask.
- **Anything that changes an invariant above** → ask before implementing.
- **Tier-3 / any real API spend** → ask before running.
- **Repeated test failures** → report them rather than weakening the assertions.

## Where Facts Live

| Need | Doc |
|---|---|
| Documentation index | `docs/README.md` |
| Architecture overview | `docs/agent_team_architecture.md` |
| Configuration / env vars | `docs/configuration.md` |
| Semantic model authoring | `docs/semantic_authoring.md`, `docs/semantic_contracts.md` |
| REST / MCP surfaces | `docs/api_reference.md`, `docs/mcp_server.md` |
| Evaluation method + frozen numbers | `docs/nl2sql_evaluation.md` |
| Database backends | `docs/database_adapters.md` |
| Release process | `docs/github_release.md` |
58 changes: 58 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,8 +17,66 @@ and does not yet claim semantic-versioning stability.
trust inspection, and persistent run history.
- D1/R2-backed atomic uploads that require a reviewed semantic contract.
- GitHub Actions quality gates and repository contribution/security metadata.
- Unified terminal outcomes (`succeeded`, `needs_clarification`, `blocked`,
`partial`, `failed`, `cancelled`) derived in one place, so a run's persisted
status, its delivery report and its event stream cannot disagree.
- A deterministic 32-task agent benchmark over three independent SQLite schemas,
with split separation, repeats, an ablation switch and a frozen effect gate.
- A 21-case multi-turn and compound-intent gold set that separates capability
failures from infrastructure failures.
- Run-budget coverage of every model call, including a persisted usage record and a
`RunContext` carrying data, semantic and policy versions.
- `docs/nl2sql_evaluation.md` now records the frozen real-model baselines and the
`--skill-mode auto|off` and candidate-count ablations, with their caveats.
- `./init.sh` — a single offline verification entry point (environment, full test
suite, repository state).

### Changed

- Canonical CLI implementation now lives in `queryforge.cli`; root `main.py` remains
a compatibility launcher.
- Workspace-relative paths (`.queryforge/` state, sample data, evaluation sets) are
resolved by `queryforge.core.paths` from `QUERYFORGE_ROOT`, an enclosing source
checkout, or the working directory — instead of `Path(__file__).parents[N]`, which
pointed into `site-packages` after an install. Packaged resources such as
`bundled_skills/` deliberately still resolve relative to the package.
- `SQLiteConnector.capabilities` is taken from the single frozen capability matrix
rather than rebuilt from dataclass defaults, so a reader of that attribute now
sees what the adapter actually enforces.
- `DatabaseTool.last_policy_decision` is read-only: it records decisions for calls
made through that tool and can no longer be overwritten by a caller, which had let
an audit record describe a call that never happened.
- Automatic skill selection is now measured rather than assumed: it costs p50
latency +139% and output tokens +44% with no demonstrated accuracy benefit, so it
is switchable and the question is recorded as open.
- For complex requests the candidate count is no longer raised implicitly; parallel
candidates are an explicit opt-in.
- The evaluator compares result projections tolerantly and classifies account-level
failures as environment errors instead of model errors, and reports measured token
usage rather than a character-count estimate.
- `make` targets prefer the repository virtualenv, so `make check` runs the same
interpreter as `./init.sh` instead of whatever `python` is first on `PATH`.

### Fixed

- Correct answers were scored wrong when the model returned extra columns or an
equivalent shape; correctness is now decided by semantic result equivalence.
- A reflection step that asked for clarification discarded the answer and the run's
artifacts; it now terminates as a structured `needs_clarification`.
- Chinese follow-up questions were not recognised, so multi-turn context was lost.
- A depleted provider balance was recorded as a model failure, corrupting accuracy
denominators; it is now an environment error, excluded from those denominators.
- The `reasoning` audit payload was silently dropped when the model returned
type-compatible-by-intent but schema-incompatible shapes (string lists, the string
`"None"`, word confidences); it is now normalised, and a discarded payload is
reported rather than ignored.
- Preview execution evaluated a rewritten statement, so a `require_limit` policy was
satisfied by the injected `LIMIT` rather than by the caller's SQL.
- Truncated result sets overwrote the row count, hiding that the bound had been hit;
`truncated` and `fetched_row_count` now preserve both numbers.
- Time-filter validation ignored `BETWEEN` and per-call time boundaries, so an
unbounded or partially bounded time filter could validate as complete.
- Studio uploads accepted `.sqlite`/`.db` files that publication could never accept;
the accepted set is now csv/parquet, and rejected database files explain why.
- The evaluator's isolation guard (independently graded code must not import the
runtime it grades) is now recursive and covers plain `import` statements.
17 changes: 17 additions & 0 deletions Makefile
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,23 @@

PYTHON ?= python

# Every target below needs the project's dependencies (sqlglot, pydantic, ...).
# ``PYTHON ?= python`` alone takes whatever ``python`` is first on PATH, which on a
# machine with a system or conda interpreter is *not* the project venv, so
# ``make check`` failed with ``ModuleNotFoundError: No module named 'sqlglot'``
# while ``init.sh`` (which hardcodes ``.venv/bin/python``) stayed green — the
# verification command documented in AGENTS.md could not be run as written.
# ``?=`` reports origin as "file", so we cannot tell a caller-supplied value from
# our own fallback by origin alone. Test the value instead: if it is still the bare
# ``python`` we defaulted to, upgrade it to the repository venv. An explicit
# ``make PYTHON=/usr/bin/python3.12 check`` (or an environment ``PYTHON``) names a
# different interpreter and is left untouched.
ifeq ($(PYTHON),python)
ifneq ($(wildcard .venv/bin/python),)
PYTHON := .venv/bin/python
endif
endif

help:
@echo "install Install QueryForge in editable mode"
@echo "install-all Install all optional integrations"
Expand Down
Loading
Loading