Skip to content

Run on Pixeltable 0.7.6, declare the schema, and add CI - #24

Merged
pierrebrunelle merged 5 commits into
mainfrom
fix/pixeltable-0.7.6
Sep 10, 2026
Merged

pierrebrunelle merged 5 commits into
mainfrom
fix/pixeltable-0.7.6

Conversation

@pierrebrunelle

Copy link
Copy Markdown
Member

Why

pixelagent does not run on any Pixeltable >= 0.7.3. Reproduced against 0.7.6 on all four providers:

provider before this PR
openai chat() raises RequestError: Error in column image: expected non-None value
anthropic cannot be constructed: missing a required argument: 'max_tokens'
bedrock same image failure
gemini same image failure

core/base.py declared image: pxt.Image and chat() inserts None into it; 0.7.3 made schema types non-nullable by default. Because memory was written before the failing insert, every failed turn also orphaned a user message permanently.

This went unnoticed because there was no CI and poetry.lock pinned pixeltable 0.3.8. Local installs resolved to a version where the code worked, while pip install pixelagent got 0.7.6 and a package that could not run its own README.

What is here

1. Green baseline + CI (622e34b) — nullable image, anthropic max_tokens + model_kwargs, chat_kwargs/tool_kwargs routed through model_kwargs=/inference_config= instead of **splat, pixeltable>=0.7.6,<0.8, poetry → uv.

Three jobs. test runs a keyless suite on 3.11/3.12. canary-unpinned installs the newest pixeltable regardless of our own bound — it is what would have caught 0.7.3 the day it shipped. live runs provider tests on a schedule with secrets.

The keyless suite asserts more than "it constructed": with no credentials, upstream columns compute with errortype is None while response.errortype is AuthorizationError. That distinguishes a broken schema from a missing key. Against the pre-fix source it reports 13 failures.

2. Delete blueprints/multi_provider/ (6440506, −2006 lines) — a near-verbatim copy of the package that had already drifted, and that still carried every bug fixed in the commit before it.

3. Declare the schema instead of assembling it (8364074) — tables are now TableModel classes reconciled by update_all() rather than a sequence of add_computed_column(..., if_exists="ignore") calls, which had no notion of drift.

pxt.model_base() returns a fresh registry per call, so each agent gets its own and model/tools stay closure literals while everything else becomes row data. They have to be literals: the three providers with a resource pool raise Could not determine resource pool on a non-literal model, and invoke_tools compiles the toolset into the column.

system_prompt is now genuinely read. The column already existed and chat() already populated it; three of four providers ignored it in favour of a baked-in value. chat() also stops inserting and then selecting the row back, and writes memory only after a turn succeeds.

Providers go from four copies of one pipeline to a ProviderSpec each: 164/182/178/166 → 65/71/71/70 lines; core 272 → 195.

Public API

Unchanged. Agent(name, system_prompt, model, tools, ...), .chat(), .tool_call(), and pxt.get_table(f"{name}.memory")/.agent all still work.

One deliberate break: reconstructing a name with a different model or tools now raises SchemaConflict naming reset=True. Previously the new argument was silently discarded and the agent kept using the old model while reporting the new one.

Still to do

This is partial; I would rather land it reviewable than sit on it. Not yet done: folding the tool handshake into the agent table, modernizing examples/, app.py + deleting blueprints/single_provider/, and the README/0.2.0 pass.

Notes for review

  • blueprints/single_provider/ is still broken and its READMEs embed the broken code inline — the README tells people to paste that into Cursor. It needs app.py as a replacement first.
  • Add MiniMax provider #23 (MiniMax provider) will need a rebase onto the new ProviderSpec layer.
  • Two things need a keyed run to settle, and keyless CI cannot: whether anthropic's model_kwargs={"system": ...} reaches the SDK as a top-level system, and whether bedrock's inference_config= maps temperature as the old splat did. Both compile; the live job needs secrets configured.
  • Model ids are stale. A keyless probe reached Bedrock via ambient AWS creds and got ValidationException: Invocation of model ID amazon.nova-pro-v1:0 with on-demand throughput isn't supported. The default model ids across all four providers should be refreshed.
  • schema.py carries a scoped F821 ignore: ruff cannot parse TableModel class bodies, where an annotated name becomes an attribute later lines reference. Scoped to that file so F821 still catches real typos elsewhere.

🤖 Generated with Claude Code

pierrebrunelle and others added 5 commits September 9, 2026 20:31
pixelagent is non-functional on every provider against Pixeltable >= 0.7.3.

- core/base.py declared `image: pxt.Image` and chat() inserts None into it.
  0.7.3 made schema types non-nullable by default, so every text-only chat()
  raised RequestError. Now `pxt.Image | None`.
- anthropic could not even be constructed: `messages()` no longer accepts a
  top-level `system=` and now requires `max_tokens`. The system prompt moves
  into `model_kwargs`, and `max_tokens` is a constructor arg defaulting to 4096.
- chat_kwargs/tool_kwargs were `**splat` into the provider call, which predates
  the model_kwargs refactor: `temperature=0.0` was an unexpected-kwarg error.
  They now route through `model_kwargs=` (openai, anthropic) and
  `inference_config=` (bedrock). Gemini already used `config=` and is unchanged.
- `Optional[pxt.tools]` annotated a function object; `pxt.Tools` is the type.

None of this was caught because there was no CI and poetry.lock pinned
pixeltable 0.3.8, so local installs resolved to a version where the code
worked while `pip install pixelagent` got 0.7.6 and a package that could not
run its own README.

Adds three jobs. `test` runs a keyless suite on 3.11/3.12. `canary-unpinned`
installs the newest pixeltable regardless of our own bound, so the next
breaking release turns our build red instead of a user's install; it is what
would have caught 0.7.3 on the day it shipped. `live` runs the provider tests
on a schedule with secrets, so a retired model id is a red weekly build.

The keyless suite asserts more than "it constructed": with no credentials the
upstream prompt column computes with errortype None while response.errortype
is AuthorizationError, which distinguishes a broken schema from a missing key.
Against the pre-fix source it reports 13 failures; against this commit, 25 pass.

Also: provider tests no longer share the `financial_assistant` catalog
directory, poetry is replaced by uv, and requires-python is >=3.11 (pixeltable
dropped 3.10 in 0.7.2).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
blueprints/multi_provider/ was a near-verbatim duplicate of pixelagent/:
core/base.py was byte-identical, and the provider agent.py files differed
only by `from ..core.base import` vs `from pixelagent.core.base import`.

It had already drifted. gemini/agent.py disagreed with the package, and the
bedrock copy carried copy-pasted comments describing Anthropic.

As of the previous commit it is actively harmful: it still declares
`image: pxt.Image`, still passes `system=` to anthropic's messages(), and
still splats `**chat_kwargs`. Every bug just fixed in pixelagent/ lives on
here, in a tree the README tells people to copy into their editor.

A "blueprint" that is one import line away from `pip install pixelagent`
teaches nothing the package does not. The README's multi-provider link now
points at the Quick Start.

blueprints/single_provider/ has the same problems and is a separate matter:
it is the paste-into-Cursor reference, so it is replaced rather than simply
deleted once app.py exists.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
An agent was built by a sequence of catalog mutations: create the tables,
then add_computed_column(..., if_exists="ignore") for each pipeline stage.
That has no notion of drift. Re-running it against an existing agent kept
whatever was already there, so `Agent(name="x", model="gpt-4.1")` against an
agent built with gpt-4o-mini silently discarded the new model and kept
answering with the old one.

The tables are now TableModel classes reconciled by update_all(), which diffs
against the catalog and fails loudly instead.

What makes this possible in a library, as opposed to an app.py where the
schema is fixed at authoring time, is that pxt.model_base() returns a fresh
model registry per call. Each agent gets its own, so `model` and `tools` can
be closed over as literals while everything else becomes row data. They have
to be literals: chat_completions, messages and generate_content each declare a
resource pool keyed on the model, and a non-literal model raises "Could not
determine resource pool"; invoke_tools compiles the toolset into the column.

Everything else is now a column: system_prompt, model_kwargs, max_tokens,
n_latest, and a new conversation_id. The agent table already had a
system_prompt column that chat() populated and three of four providers then
ignored in favour of the baked-in value. It is now actually read, so one agent
can serve turns with different system prompts, and chat() takes per-call
overrides for all four.

Two bugs fall out of the restructuring:

- A failed turn no longer orphans a message. Memory was written before the
  provider call, so a failure left an unanswered user message in history
  permanently. It is now written after the turn produces a reply.
- chat() no longer inserts and then selects the row back to read what it just
  computed. insert(return_rows=True) carries every computed column, which also
  drops a full table scan per turn and the dependence on message_id being
  unique.

The providers stop being four copies of one pipeline: each is now a
ProviderSpec plus a constructor. core/base.py 272 -> 195 lines, and the four
agent.py files 164/182/178/166 -> 65/71/71/70.

Public API is unchanged. The one deliberate break is the conflict above: a
name reconstructed with a different model or toolset now raises SchemaConflict
naming reset=True, rather than silently ignoring the argument.

Keyless suite goes 25 -> 41 tests, adding: schema re-declaration is a no-op
(the guard for pixeltable PR #1599), the conflict raises, system prompts vary
per row, and a failed turn writes no memory.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every tool-calling example failed before reaching pixelagent at all:

  RequestError: Defining the UDF 'stock_price' directly in the global
  namespace of a Python script is not allowed.

A UDF has to be importable by name so a stored computed column can find it
again on a later run, and a script's globals are not. Nine example scripts
defined their tools inline, including the getting-started tutorial. Each now
has a sibling *_tools.py, which is the fix pixeltable's own error message
prescribes.

video-summarizer/table.py was broken a second way: AudioSplitter.create was
passed chunk_duration_sec / overlap_sec / min_chunk_duration_sec, and the
iterator shims forward **kwargs verbatim with no name mapping, so those are
hard errors rather than warnings. They are duration / overlap /
min_segment_duration now, and the output column it reads is audio_segment,
not audio_chunk.

Deprecations across agentic-rag: the pixeltable.iterators shims give way to
document_splitter / audio_splitter / string_splitter / frame_iterator, and
openai.vision to chat_completions with an image_url content block. Eight
files still called the positional .similarity(query), deprecated since 0.5.7.

tests/test_examples.py gates all of it: every example must compile, and none
may carry a script-level @pxt.udf, a deprecated iterator import, openai.vision,
a stale iterator kwarg or output column, or a positional .similarity(). Greps
and compiles rather than executions, because the examples need media files,
provider keys and network access to actually run.

Verified with PixeltableDeprecationWarning promoted to an error: frame_iterator
plus the chat_completions vision form, document_splitter and audio_splitter all
build clean on 0.7.6.

Suite 41 -> 81.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Fallout from swapping the deprecated iterator imports in the previous commit;
ruff's I001 gate catches it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@pierrebrunelle
pierrebrunelle merged commit 0c4bb4f into main Sep 10, 2026
4 checks passed
pierrebrunelle added a commit that referenced this pull request Sep 10, 2026
Follow-up to #24. Two things it left behind, plus the version bump.

## The blueprints were the last copy of the bugs

`blueprints/single_provider/` still carried everything #24 fixed. All
four `agent.py` files declared `"image": pxt.Image`, passed `system=` to
anthropic's `messages()`, and splatted `**chat_kwargs` — and the four
`README.md` files embedded that same broken code inline, under a heading
telling people to connect them to Cursor, Windsurf and Cline. Anyone
following that instruction pasted code that cannot run.

Deleted: 15 files, ~3,200 lines. Same reasoning as `multi_provider` in
#24 — all four providers ship in the package, so a blueprint is one
import away from `pip install pixelagent`. That README section is now
the install line and a five-line Quick Start.

## The ReAct example had never executed

It defined a `@pxt.udf` in a script's global namespace (a hard error
since the UDF must be importable by name), and twice called
`Agent(name="financial_planner")` without the required `system_prompt`.
After #24 it would also have hit `SchemaConflict`, since it rebuilt the
same agent name with a different toolset.

Rewritten around what #24 made possible. The loop existed to vary the
system prompt per step, which used to mean reconstructing the agent; the
prompt is row data now, so it is one agent and a per-call override:

```python
response = agent.chat(question, system_prompt=REACT_PROMPT.format(step=step, ...))
```

The loop gets shorter and the example demonstrates the change instead of
working around it. All five README python blocks now compile.

## Version

`0.1.6` → `0.2.0`, plus a **Migrating from 0.1.x** section covering the
four caller-visible changes: Python 3.11+, `SchemaConflict` on a changed
model/toolset, per-turn config overrides, and the new `conversation_id`.

**This still needs publishing to reach anyone.** PyPI serves **0.1.5**
with an unbounded `pixeltable>=0.3.15`, so a fresh `pip install
pixelagent` today resolves Pixeltable 0.7.6 and gets the broken code.
The merge fixed the repo, not the install. I have not published — that
is yours to run.

## Unchanged from #24's open items

- `live` job still needs `OPENAI_API_KEY` / `ANTHROPIC_API_KEY` /
`GEMINI_API_KEY` secrets. Two things stay unverified without it: whether
anthropic's `model_kwargs={"system": ...}` reaches the SDK as a
top-level `system`, and whether bedrock's `inference_config=` maps
`temperature` as the old splat did.
- Default model ids are stale — bedrock's `amazon.nova-pro-v1:0` returns
`ValidationException` on on-demand throughput.
- Separately: GitHub reports 40 dependabot vulnerabilities on `main` (1
critical, 23 high). Pre-existing, unrelated to this work.

81 keyless tests, ruff clean.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
pierrebrunelle added a commit that referenced this pull request Sep 10, 2026
0.2.0 has been on `main` since #25, but PyPI still serves **0.1.5**, so
nobody has actually received the fix. This makes releasing a tag push.

## The old script was broken

`scripts/release.sh` could not have worked as-is:

- required a git remote named `home`; a normal clone has `origin`
- read a long-lived `PYPI_API_KEY` out of `~/.pixeltable/config.toml`
- called `poetry build` after #24 moved the project to uv
- **published without ever installing what it built**

That last one is how 0.1.x shipped a package that could not import
against a current Pixeltable. Nothing in the release path ever tried it.

## Releasing now

```bash
git tag v0.2.0 && git push origin v0.2.0
```

**There is no token.** PyPI Trusted Publishing mints a short-lived
credential from the workflow's OIDC identity, so there is nothing to
leak, rotate, or paste into a shell.

The build job refuses to publish an artifact it has not exercised:

1. tag must match `pyproject.toml` version
2. `uv build` + `twine check`
3. install the built wheel into a clean venv and **construct a real
agent with it**

Verified locally on both projects. Worth noting: the wheel resolves
**pixeltable 0.7.7** through the `<0.8` bound and works — newer than the
0.7.6 we targeted, so the bound is doing its job.

## What you need to do once

On PyPI, **Manage project → Publishing → Add a pending publisher**:

| field | value |
|---|---|
| owner | `pixeltable` |
| repository | `pixelagent` |
| workflow | `release.yml` |
| environment | `pypi` |

Then tag. `RELEASING.md` has the full process.

Two optional extras: add a required reviewer to the `pypi` environment
(Settings → Environments) if you want a human gate before each publish,
and run the workflow from the Actions tab with `dry_run` checked to
exercise the whole path without publishing.

## Still not done by this PR

Adding `OPENAI_API_KEY` / `ANTHROPIC_API_KEY` / `GEMINI_API_KEY` as repo
secrets, so the `live` job can run. Without it, two things stay
unverified: whether anthropic's `model_kwargs={"system": ...}` reaches
the SDK as a top-level `system`, and whether bedrock's
`inference_config=` maps `temperature` as the old splat did. I did not
add those — I don't handle credentials.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant