Make Tool failures actionable to the model - #953
Draft
Y1fe1Zh0u wants to merge 2 commits into
Draft
Conversation
Set the existing protocol repair, safe-read replay, and model-visible Tool episode limits to ten while preserving their current independent state and execution semantics. Update focused tests and planning artifacts to make the off-by-one behavior explicit. Constraint: Tool-related retry and repair limits must be ten without restructuring the existing counters Rejected: Unify protocol, Receipt, and model-visible repair state now | counter redesign is intentionally deferred Confidence: high Scope-risk: moderate Directive: Keep the independent counters until the planned repair-control refactor; do not infer identical attempt semantics from the shared numeric limit Tested: 911 Runtime and Tool pytest cases; scoped Ruff; fatal-level caller Ruff; py_compile; git diff --check Not-tested: Live Provider credentials
Runtime Tool results already retain sanitized failure summaries and optional remediation, but the model boundary discarded the failure signal and remediation text. Carry one provider-neutral error bit, render actionable failure text, and map it to native Anthropic and Gemini representations. Constraint: Keep Runtime control fields such as model_action and side_effect_state out of the model protocol. Rejected: Serialize the complete internal Tool outcome | couples provider prompts to Runtime control metadata. Confidence: high Scope-risk: narrow Directive: Expected Tool failures should provide sanitized result_summary and safe_remediation at their source. Tested: 151 scoped model-step, provider-boundary, and Tool-step tests; scoped Ruff; git diff --check. Not-tested: Live Anthropic, Gemini, and OpenAI provider calls.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changed
is_errorbit on model-facing Tool Resultstool_result.is_errorresponse.errorfor failures andresponse.outputfor successWhy
Runtime already retained the Tool failure summary and optional safe remediation, but
_prompt_messages()only forwarded plain content and the original Call ID. Anthropic and Gemini therefore received no native failure signal, and the remediation never reached the model. This made argument repair depend on guessing from an unlabelled string.The change intentionally does not serialize internal Runtime control fields such as
model_action,side_effect_state, Receipt metadata, or reconciliation state into the model protocol.Impact
A model now receives a call-linked failure such as:
Provider adapters additionally express the failure using their native representation where available.
Validation
git diff --checkpassedStack
This draft targets
002-tool-runtime-contractbecause the Runtime Tool outcome and StepToolContext implementation is currently in PR #945 and has not yet merged intomain.