Greenfield horizontally scaled AI platform with two verticals on shared infra:
- Autonomous coding agents — plan → isolated workspace → patch → pytest verifier → merge status (multi-agent retry vs single-agent baseline).
- Behaviour analytics — synthetic trading/research product journeys (funnels, anomalies, evidence-grounded NL insights) with consent + deletion.
Placement pitch: I can ship distributed agent systems with evals and a real analytics product surface—not a chat demo.
No third-party API keys required. Planners and workers use offline stubs; CI and local demos never call OpenAI/Anthropic/etc. Local auth uses the documented demo token dev-token (see .env.example).
Repo: https://github.com/Wojtek-06/AgentGrid · License: MIT
Non-goals: Fitness-App code/domain; live extractive agents against private remotes; LLM network calls in CI.
Full write-up: docs/ARCHITECTURE.md.
┌─────────────────────────────────────┐
│ Dashboard (static UI) │
│ board · metrics · funnel · SSE │
└──────────────────┬──────────────────┘
│ Bearer token
▼
┌──────────────┐ enqueue ┌────────────────┐ dequeue ┌─────────────────┐
│ FastAPI API │─────────────►│ Queue (Redis │────────────►│ Coding workers │
│ jobs · eval │ │ or in-proc) │ │ plan→patch→test │
│ analytics │◄─────────────│ │◄────────────│ + merge path │
└──────┬───────┘ status └────────────────┘ results └────────┬────────┘
│ │
│ ┌───────────────────────────────┘
▼ ▼
┌──────────────┐ ┌─────────────────┐
│ SQLite / PG │ │ .artifacts/ │
│ jobs·events │ │ patch + checklist│
└──────────────┘ └─────────────────┘
Dogfood issues (local sandboxes, QuantForge/ChainVenue-shaped):
| Issue ID | Story | Eval contrast |
|---|---|---|
qf-leakage-guard |
Look-ahead mid bug | multi retries, single fails |
qf-ewma-alpha |
EWMA weights swapped | multi retries, single fails |
cv-basis-bps |
Wrong basis sign | both modes can fix |
Source of truth: data/eval_results.json (regenerate with python scripts/run_eval.py).
| Mode | Success rate | Avg tokens |
|---|---|---|
| Single-agent | 33% (1/3) | ~200 |
| Multi-agent | 100% (3/3) | ~287 |
Per-issue breakdown and interview script: docs/EVIDENCE_PACK.md.
A separate worker process needs a shared Redis queue (in-process queue is for pytest only).
Redis (pick one):
- WSL:
sudo service redis-server start(orredis-server) onlocalhost:6379 - Docker:
docker compose up -d redis
cd C:\Projekty\Quant\AgentGrid
python -m pip install -r requirements.txt
# Terminal A — API (sets USE_REDIS=1 + demo token)
.\scripts\run_api.ps1
# Terminal B — coding worker (heartbeats on /api/health)
.\scripts\run_worker.ps1Tests (no Redis required): $env:PYTHONPATH="backend"; python -m pytest -q
- Open http://127.0.0.1:8000 — token
dev-token; health strip should showredis on · workers ≥ 1. - Select issue
qf-leakage-guard, mode single → Enqueue → status goesfailed. - Same issue, mode multi → Enqueue → status goes
succeeded(retry path). - Click Load published (instant) or Run eval → multi 100% vs single 33%.
Optional: python scripts\seed_analytics.py then Refresh for the research funnel.
Protected routes expect:
Authorization: Bearer <AGENTGRID_API_TOKEN>Default token is dev-token (see .env.example) — change it outside local demos.
SSE (EventSource) cannot set headers, so only /api/jobs/stream accepts ?token= (Bearer everywhere else).
Health (GET /api/health) is open and reports queue depth, Redis reachability, and live worker heartbeats.
# API + Redis + one worker
docker compose up --build
# Horizontal scale story — extra workers share the Redis queue
docker compose up --build --scale worker=2Optional Postgres profile (SQLite remains CI/default):
docker compose --profile postgres up --build api-pg worker-pg postgres redisExtras:
python scripts\run_eval.py # refresh data/eval_results.json
python scripts\seed_analytics.py # demo funnel + retentionAll rows except health require Authorization: Bearer <token> (or ?token= on SSE).
| Method | Path | Auth | Purpose |
|---|---|---|---|
| GET | /api/health |
— | Liveness, queue, Redis, worker heartbeats |
| GET | /api/jobs/issues |
Bearer | Dogfood catalog |
| POST | /api/jobs |
Bearer | {issue_id, mode, idempotency_key?} |
| GET | /api/jobs |
Bearer | Board snapshot |
| GET | /api/jobs/stream |
Bearer or ?token= |
SSE live board |
| GET | /api/jobs/{id} |
Bearer | Job detail + patch/log |
| POST | /api/jobs/{id}/cancel |
Bearer | Cancel queued/running |
| POST | /api/jobs/{id}/retry |
Bearer | Re-enqueue failed/cancelled |
| GET | /api/eval/latest |
Bearer | Published data/eval_results.json |
| POST | /api/eval/run |
Bearer | Re-run multi vs single harness |
| GET | /api/metrics/overview |
Bearer | Tokens / $ / latency / queue |
| POST | /api/analytics/events |
Bearer | Batch ingest |
| GET | /api/analytics/funnel |
Bearer | Funnel + anomalies + insight |
| GET | /api/analytics/retention |
Bearer | Day-1 retention cohorts |
| GET | /api/analytics/operator-funnel |
Bearer | Operator telemetry funnel |
| POST | /api/analytics/consent |
Bearer | Consent flag |
| DELETE | /api/analytics/users/{id} |
Bearer | Erase + block |
Responses include X-Request-ID (echo client header or generate). API + worker logs use structured request_id=… fields.
| Deliverable | Status |
|---|---|
| Coordinator + workers + queue + verifier | Done (local/Redis) |
| Multi vs single eval | Done (scripts/run_eval.py → data/eval_results.json) |
| Merge artifacts + human review checklist | Done |
| Merge conflict risk surfacing + tests | Done |
| Cancel / retry / queue backpressure | Done |
| Observability metrics + request IDs / structured logs | Done |
| SSE live job board | Done |
| Analytics + privacy | Done (research journeys) |
| Operator telemetry + retention cohorts | Done |
| Optional Postgres compose profile | Done (SQLite remains CI default) |
| Horizontal scale story | Documented + compose --scale worker=N |
| Dogfood on QF/CV-shaped issues | 3 local sandboxes |
| Evidence pack / demo video | Docs ready; video user-owned |
Sibling status: docs/PORTFOLIO_STATUS.md
Docs: docs/EVIDENCE_PACK.md · docs/THREAT_PRIVACY.md · Screenshots: docs/images/
- QuantForge — C++ LOB MM lab (dogfood-shaped issues)
- ChainVenue — Foundry CLOB–AMM lab (dogfood-shaped issue)