Summary
run_factory can be exposed with an unconstrained name argument even when no matching factory is registered. The model can interpret tool availability as factory availability, invent plausible names, and retry them serially.
Originally reported from the VS Code Agents Window in microsoft/vscode#329551.
Observed behavior
With Copilot CLI 1.0.78 and GPT-5.6 Sol, one session produced 13 distinct model-authored run_factory calls. Every call failed immediately with:
RpcResponseError: No factory registered with name "<name>"
The model first tried eight implementation-oriented names:
apply_finding
code_implementation_factory
single_task
implementation
code-task
fix-factory
repo_task
review_then_implement
It later tried five more names for video analysis:
video_analysis
analyze_video
single_video_analysis
video-analysis
media_analysis
These were separate assistant-authored calls after each rejection, not host retries. No factory ran.
Steps to reproduce
- Expose
run_factory in a session where no matching factories are registered.
- Ask the agent to perform a routine single-session task, such as a small implementation or media-analysis task.
- Observe that the model may invent and repeatedly retry factory names.
Expected behavior
Invalid factory names should be impossible or strongly bounded. Ideally:
- Dynamically constrain
name to registered factory names.
- Do not expose
run_factory when there are no registered factories or resumable runs.
- If dynamic enum generation is not possible, require registry discovery before invocation and include available names in validation errors.
- Enforce exactly one of
name and resumeFromRunId at the schema level.
Actual behavior
name is effectively a free-form string, so failures occur only after a costly tool round trip. Prompt guidance to use factories selectively did not prevent repeated guesses.
Summary
run_factorycan be exposed with an unconstrainednameargument even when no matching factory is registered. The model can interpret tool availability as factory availability, invent plausible names, and retry them serially.Originally reported from the VS Code Agents Window in microsoft/vscode#329551.
Observed behavior
With Copilot CLI 1.0.78 and GPT-5.6 Sol, one session produced 13 distinct model-authored
run_factorycalls. Every call failed immediately with:The model first tried eight implementation-oriented names:
apply_findingcode_implementation_factorysingle_taskimplementationcode-taskfix-factoryrepo_taskreview_then_implementIt later tried five more names for video analysis:
video_analysisanalyze_videosingle_video_analysisvideo-analysismedia_analysisThese were separate assistant-authored calls after each rejection, not host retries. No factory ran.
Steps to reproduce
run_factoryin a session where no matching factories are registered.Expected behavior
Invalid factory names should be impossible or strongly bounded. Ideally:
nameto registered factory names.run_factorywhen there are no registered factories or resumable runs.nameandresumeFromRunIdat the schema level.Actual behavior
nameis effectively a free-form string, so failures occur only after a costly tool round trip. Prompt guidance to use factories selectively did not prevent repeated guesses.