Repository navigation
Run on Pixeltable 0.7.6, declare the schema, and add CI - #24
Merged
Merged
Conversation
pixelagent is non-functional on every provider against Pixeltable >= 0.7.3. - core/base.py declared `image: pxt.Image` and chat() inserts None into it. 0.7.3 made schema types non-nullable by default, so every text-only chat() raised RequestError. Now `pxt.Image | None`. - anthropic could not even be constructed: `messages()` no longer accepts a top-level `system=` and now requires `max_tokens`. The system prompt moves into `model_kwargs`, and `max_tokens` is a constructor arg defaulting to 4096. - chat_kwargs/tool_kwargs were `**splat` into the provider call, which predates the model_kwargs refactor: `temperature=0.0` was an unexpected-kwarg error. They now route through `model_kwargs=` (openai, anthropic) and `inference_config=` (bedrock). Gemini already used `config=` and is unchanged. - `Optional[pxt.tools]` annotated a function object; `pxt.Tools` is the type. None of this was caught because there was no CI and poetry.lock pinned pixeltable 0.3.8, so local installs resolved to a version where the code worked while `pip install pixelagent` got 0.7.6 and a package that could not run its own README. Adds three jobs. `test` runs a keyless suite on 3.11/3.12. `canary-unpinned` installs the newest pixeltable regardless of our own bound, so the next breaking release turns our build red instead of a user's install; it is what would have caught 0.7.3 on the day it shipped. `live` runs the provider tests on a schedule with secrets, so a retired model id is a red weekly build. The keyless suite asserts more than "it constructed": with no credentials the upstream prompt column computes with errortype None while response.errortype is AuthorizationError, which distinguishes a broken schema from a missing key. Against the pre-fix source it reports 13 failures; against this commit, 25 pass. Also: provider tests no longer share the `financial_assistant` catalog directory, poetry is replaced by uv, and requires-python is >=3.11 (pixeltable dropped 3.10 in 0.7.2). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
blueprints/multi_provider/ was a near-verbatim duplicate of pixelagent/: core/base.py was byte-identical, and the provider agent.py files differed only by `from ..core.base import` vs `from pixelagent.core.base import`. It had already drifted. gemini/agent.py disagreed with the package, and the bedrock copy carried copy-pasted comments describing Anthropic. As of the previous commit it is actively harmful: it still declares `image: pxt.Image`, still passes `system=` to anthropic's messages(), and still splats `**chat_kwargs`. Every bug just fixed in pixelagent/ lives on here, in a tree the README tells people to copy into their editor. A "blueprint" that is one import line away from `pip install pixelagent` teaches nothing the package does not. The README's multi-provider link now points at the Quick Start. blueprints/single_provider/ has the same problems and is a separate matter: it is the paste-into-Cursor reference, so it is replaced rather than simply deleted once app.py exists. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
An agent was built by a sequence of catalog mutations: create the tables, then add_computed_column(..., if_exists="ignore") for each pipeline stage. That has no notion of drift. Re-running it against an existing agent kept whatever was already there, so `Agent(name="x", model="gpt-4.1")` against an agent built with gpt-4o-mini silently discarded the new model and kept answering with the old one. The tables are now TableModel classes reconciled by update_all(), which diffs against the catalog and fails loudly instead. What makes this possible in a library, as opposed to an app.py where the schema is fixed at authoring time, is that pxt.model_base() returns a fresh model registry per call. Each agent gets its own, so `model` and `tools` can be closed over as literals while everything else becomes row data. They have to be literals: chat_completions, messages and generate_content each declare a resource pool keyed on the model, and a non-literal model raises "Could not determine resource pool"; invoke_tools compiles the toolset into the column. Everything else is now a column: system_prompt, model_kwargs, max_tokens, n_latest, and a new conversation_id. The agent table already had a system_prompt column that chat() populated and three of four providers then ignored in favour of the baked-in value. It is now actually read, so one agent can serve turns with different system prompts, and chat() takes per-call overrides for all four. Two bugs fall out of the restructuring: - A failed turn no longer orphans a message. Memory was written before the provider call, so a failure left an unanswered user message in history permanently. It is now written after the turn produces a reply. - chat() no longer inserts and then selects the row back to read what it just computed. insert(return_rows=True) carries every computed column, which also drops a full table scan per turn and the dependence on message_id being unique. The providers stop being four copies of one pipeline: each is now a ProviderSpec plus a constructor. core/base.py 272 -> 195 lines, and the four agent.py files 164/182/178/166 -> 65/71/71/70. Public API is unchanged. The one deliberate break is the conflict above: a name reconstructed with a different model or toolset now raises SchemaConflict naming reset=True, rather than silently ignoring the argument. Keyless suite goes 25 -> 41 tests, adding: schema re-declaration is a no-op (the guard for pixeltable PR #1599), the conflict raises, system prompts vary per row, and a failed turn writes no memory. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Every tool-calling example failed before reaching pixelagent at all: RequestError: Defining the UDF 'stock_price' directly in the global namespace of a Python script is not allowed. A UDF has to be importable by name so a stored computed column can find it again on a later run, and a script's globals are not. Nine example scripts defined their tools inline, including the getting-started tutorial. Each now has a sibling *_tools.py, which is the fix pixeltable's own error message prescribes. video-summarizer/table.py was broken a second way: AudioSplitter.create was passed chunk_duration_sec / overlap_sec / min_chunk_duration_sec, and the iterator shims forward **kwargs verbatim with no name mapping, so those are hard errors rather than warnings. They are duration / overlap / min_segment_duration now, and the output column it reads is audio_segment, not audio_chunk. Deprecations across agentic-rag: the pixeltable.iterators shims give way to document_splitter / audio_splitter / string_splitter / frame_iterator, and openai.vision to chat_completions with an image_url content block. Eight files still called the positional .similarity(query), deprecated since 0.5.7. tests/test_examples.py gates all of it: every example must compile, and none may carry a script-level @pxt.udf, a deprecated iterator import, openai.vision, a stale iterator kwarg or output column, or a positional .similarity(). Greps and compiles rather than executions, because the examples need media files, provider keys and network access to actually run. Verified with PixeltableDeprecationWarning promoted to an error: frame_iterator plus the chat_completions vision form, document_splitter and audio_splitter all build clean on 0.7.6. Suite 41 -> 81. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Fallout from swapping the deprecated iterator imports in the previous commit; ruff's I001 gate catches it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
pierrebrunelle
added a commit
that referenced
this pull request
Sep 10, 2026
Follow-up to #24. Two things it left behind, plus the version bump. ## The blueprints were the last copy of the bugs `blueprints/single_provider/` still carried everything #24 fixed. All four `agent.py` files declared `"image": pxt.Image`, passed `system=` to anthropic's `messages()`, and splatted `**chat_kwargs` — and the four `README.md` files embedded that same broken code inline, under a heading telling people to connect them to Cursor, Windsurf and Cline. Anyone following that instruction pasted code that cannot run. Deleted: 15 files, ~3,200 lines. Same reasoning as `multi_provider` in #24 — all four providers ship in the package, so a blueprint is one import away from `pip install pixelagent`. That README section is now the install line and a five-line Quick Start. ## The ReAct example had never executed It defined a `@pxt.udf` in a script's global namespace (a hard error since the UDF must be importable by name), and twice called `Agent(name="financial_planner")` without the required `system_prompt`. After #24 it would also have hit `SchemaConflict`, since it rebuilt the same agent name with a different toolset. Rewritten around what #24 made possible. The loop existed to vary the system prompt per step, which used to mean reconstructing the agent; the prompt is row data now, so it is one agent and a per-call override: ```python response = agent.chat(question, system_prompt=REACT_PROMPT.format(step=step, ...)) ``` The loop gets shorter and the example demonstrates the change instead of working around it. All five README python blocks now compile. ## Version `0.1.6` → `0.2.0`, plus a **Migrating from 0.1.x** section covering the four caller-visible changes: Python 3.11+, `SchemaConflict` on a changed model/toolset, per-turn config overrides, and the new `conversation_id`. **This still needs publishing to reach anyone.** PyPI serves **0.1.5** with an unbounded `pixeltable>=0.3.15`, so a fresh `pip install pixelagent` today resolves Pixeltable 0.7.6 and gets the broken code. The merge fixed the repo, not the install. I have not published — that is yours to run. ## Unchanged from #24's open items - `live` job still needs `OPENAI_API_KEY` / `ANTHROPIC_API_KEY` / `GEMINI_API_KEY` secrets. Two things stay unverified without it: whether anthropic's `model_kwargs={"system": ...}` reaches the SDK as a top-level `system`, and whether bedrock's `inference_config=` maps `temperature` as the old splat did. - Default model ids are stale — bedrock's `amazon.nova-pro-v1:0` returns `ValidationException` on on-demand throughput. - Separately: GitHub reports 40 dependabot vulnerabilities on `main` (1 critical, 23 high). Pre-existing, unrelated to this work. 81 keyless tests, ruff clean. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
pierrebrunelle
added a commit
that referenced
this pull request
Sep 10, 2026
0.2.0 has been on `main` since #25, but PyPI still serves **0.1.5**, so nobody has actually received the fix. This makes releasing a tag push. ## The old script was broken `scripts/release.sh` could not have worked as-is: - required a git remote named `home`; a normal clone has `origin` - read a long-lived `PYPI_API_KEY` out of `~/.pixeltable/config.toml` - called `poetry build` after #24 moved the project to uv - **published without ever installing what it built** That last one is how 0.1.x shipped a package that could not import against a current Pixeltable. Nothing in the release path ever tried it. ## Releasing now ```bash git tag v0.2.0 && git push origin v0.2.0 ``` **There is no token.** PyPI Trusted Publishing mints a short-lived credential from the workflow's OIDC identity, so there is nothing to leak, rotate, or paste into a shell. The build job refuses to publish an artifact it has not exercised: 1. tag must match `pyproject.toml` version 2. `uv build` + `twine check` 3. install the built wheel into a clean venv and **construct a real agent with it** Verified locally on both projects. Worth noting: the wheel resolves **pixeltable 0.7.7** through the `<0.8` bound and works — newer than the 0.7.6 we targeted, so the bound is doing its job. ## What you need to do once On PyPI, **Manage project → Publishing → Add a pending publisher**: | field | value | |---|---| | owner | `pixeltable` | | repository | `pixelagent` | | workflow | `release.yml` | | environment | `pypi` | Then tag. `RELEASING.md` has the full process. Two optional extras: add a required reviewer to the `pypi` environment (Settings → Environments) if you want a human gate before each publish, and run the workflow from the Actions tab with `dry_run` checked to exercise the whole path without publishing. ## Still not done by this PR Adding `OPENAI_API_KEY` / `ANTHROPIC_API_KEY` / `GEMINI_API_KEY` as repo secrets, so the `live` job can run. Without it, two things stay unverified: whether anthropic's `model_kwargs={"system": ...}` reaches the SDK as a top-level `system`, and whether bedrock's `inference_config=` maps `temperature` as the old splat did. I did not add those — I don't handle credentials. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
pixelagent does not run on any Pixeltable >= 0.7.3. Reproduced against 0.7.6 on all four providers:
chat()raisesRequestError: Error in column image: expected non-None valuemissing a required argument: 'max_tokens'core/base.pydeclaredimage: pxt.Imageandchat()insertsNoneinto it; 0.7.3 made schema types non-nullable by default. Because memory was written before the failing insert, every failed turn also orphaned a user message permanently.This went unnoticed because there was no CI and
poetry.lockpinned pixeltable 0.3.8. Local installs resolved to a version where the code worked, whilepip install pixelagentgot 0.7.6 and a package that could not run its own README.What is here
1. Green baseline + CI (
622e34b) — nullable image, anthropicmax_tokens+model_kwargs,chat_kwargs/tool_kwargsrouted throughmodel_kwargs=/inference_config=instead of**splat,pixeltable>=0.7.6,<0.8, poetry → uv.Three jobs.
testruns a keyless suite on 3.11/3.12.canary-unpinnedinstalls the newest pixeltable regardless of our own bound — it is what would have caught 0.7.3 the day it shipped.liveruns provider tests on a schedule with secrets.The keyless suite asserts more than "it constructed": with no credentials, upstream columns compute with
errortype is Nonewhileresponse.errortypeisAuthorizationError. That distinguishes a broken schema from a missing key. Against the pre-fix source it reports 13 failures.2. Delete
blueprints/multi_provider/(6440506, −2006 lines) — a near-verbatim copy of the package that had already drifted, and that still carried every bug fixed in the commit before it.3. Declare the schema instead of assembling it (
8364074) — tables are nowTableModelclasses reconciled byupdate_all()rather than a sequence ofadd_computed_column(..., if_exists="ignore")calls, which had no notion of drift.pxt.model_base()returns a fresh registry per call, so each agent gets its own andmodel/toolsstay closure literals while everything else becomes row data. They have to be literals: the three providers with a resource pool raiseCould not determine resource poolon a non-literal model, andinvoke_toolscompiles the toolset into the column.system_promptis now genuinely read. The column already existed andchat()already populated it; three of four providers ignored it in favour of a baked-in value.chat()also stops inserting and then selecting the row back, and writes memory only after a turn succeeds.Providers go from four copies of one pipeline to a
ProviderSpeceach: 164/182/178/166 → 65/71/71/70 lines; core 272 → 195.Public API
Unchanged.
Agent(name, system_prompt, model, tools, ...),.chat(),.tool_call(), andpxt.get_table(f"{name}.memory")/.agentall still work.One deliberate break: reconstructing a name with a different
modelortoolsnow raisesSchemaConflictnamingreset=True. Previously the new argument was silently discarded and the agent kept using the old model while reporting the new one.Still to do
This is partial; I would rather land it reviewable than sit on it. Not yet done: folding the tool handshake into the agent table, modernizing
examples/,app.py+ deletingblueprints/single_provider/, and the README/0.2.0 pass.Notes for review
blueprints/single_provider/is still broken and its READMEs embed the broken code inline — the README tells people to paste that into Cursor. It needsapp.pyas a replacement first.ProviderSpeclayer.model_kwargs={"system": ...}reaches the SDK as a top-levelsystem, and whether bedrock'sinference_config=mapstemperatureas the old splat did. Both compile; thelivejob needs secrets configured.ValidationException: Invocation of model ID amazon.nova-pro-v1:0 with on-demand throughput isn't supported. The default model ids across all four providers should be refreshed.schema.pycarries a scopedF821ignore: ruff cannot parseTableModelclass bodies, where an annotated name becomes an attribute later lines reference. Scoped to that file so F821 still catches real typos elsewhere.🤖 Generated with Claude Code