Skip to content

feat(examples): add a self-correcting coding agent harness - #951

Open
haigou-web wants to merge 6 commits into
apache:mainfrom
haigou-web:feat/coding-agent-harness
Open

haigou-web wants to merge 6 commits into
apache:mainfrom
haigou-web:feat/coding-agent-harness

Conversation

@haigou-web

Copy link
Copy Markdown

Adds an end-to-end example under examples/coding-agent/: a small coding agent built as a Burr state machine, plus the tools it drives.

The interesting part is that the harness is self-correcting — a failed tool call or a failing test feeds back into the next action instead of ending the run. That feedback loop is what makes the state machine do real work rather than just sequence steps, so it's the thing the example is meant to demonstrate.

Changes

  • examples/coding-agent/ — the state machine, the application wiring, and the tool surface, plus a README and a notebook walkthrough.
  • tools.py carries two hardening fixes found in review:
    • Tool arguments that arrive as a JSON string are now parsed instead of being handed to code expecting a dict (this previously raised).
    • Path containment uses normcase plus a separator-stripped prefix comparison, so the same path with a differently-cased Windows drive letter no longer reads as an escape.

How I tested this

py -3 -m pytest examples/coding-agent/test_coding_agent.py -q
# 16 passed

The two tools.py fixes each have a test that fails without the fix.

Notes

  • Rebased onto current main (81611591); the diff is confined to examples/coding-agent/.
  • Happy to split the docs/notebook into a separate PR if you'd rather keep this one code-only.

Agentic-JJ-Web3 and others added 5 commits October 6, 2026 12:08
Address the review findings on the coding-agent example:

- Drop the unused imports in tools.py; `requests` was not in
  requirements.txt, so a clean install failed to import at all.
- Route tool actions back through create_prompt, so the prompt-shaping
  seam runs before every model call rather than only the first.
- Report an unknown tool name back to the model as a tool error instead
  of leaving call_llm with no matching transition.
- Catch malformed tool-call JSON in OpenAIClient and hand it to the model
  as an error rather than raising.
- Wrap tool execution in try/except so a raising tool becomes an error
  result the model can react to.
- Create the workspace before running a command, so a model that starts
  with run_bash does not hit FileNotFoundError.
- Resolve paths with realpath so a symlink inside the workspace cannot be
  followed out of it.
- Skip the graphviz render when the binary is unavailable, instead of
  crashing before the example runs.
- Add the ASF license header to README.md.
Cover the scripted happy path, step-budget termination, unknown tools,
malformed tool arguments, tool failures, and workspace confinement for
reads and writes. Driven by ScriptedClient, so no API key is needed, and
tools.WORKSPACE is patched to tmp_path so tests never touch a real
./workspace.
…ive paths

Two defects in the new example, both found in review:

- `json.loads(None)` raises TypeError, which `except JSONDecodeError` did not
  catch. A backend that sends `arguments=None` (some OpenAI-compatible
  providers do when streaming fails to assemble the call) would therefore
  abort the run -- exactly the failure this block was added to prevent.
  Non-string arguments are now rejected with a clear message.

- `_resolve` compared a realpath-derived path against WORKSPACE with a plain
  string prefix test. On case-insensitive filesystems the two can differ in
  case, so a legitimate path could be rejected as an escape. The comparison is
  now normcased, and `rstrip` handles the drive-root case where
  `root + os.sep` would double the separator.

Also corrects the transition-order comment: the `next_tool=None` branch must
come *after* the `done=True` branch (a finished run has no tool calls, so both
conditions hold and the first match in insertion order wins), not merely
"before the tool transitions".

Tests: the existing malformed-arguments test hand-wrote the `error` field, so
the parsing path had no coverage at all. Added tests driving OpenAIClient
through a stand-in `openai` module (None + invalid JSON), plus workspace
confinement edges (root accepted, sibling sharing the prefix rejected, case
differences tolerated). 16 passed.
@github-actions github-actions Bot added the area/examples Relates to /examples label Oct 6, 2026
os.path.normcase is a no-op on POSIX, so the documented macOS coverage was not
actually provided by it; on macOS realpath already reports the on-disk case.
os.path.commonpath compares components and is case-insensitive exactly where
the platform is. The containment rule moves into _is_within so it can be
asserted directly instead of through the filesystem — the old test passed on
macOS without exercising it.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/examples Relates to /examples

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants