Problem
LocalHFBackend extracts tool calls from raw model output with
parse_tools() (mellea/backends/tools.py:474), which only handles JSON
tool calls (its docstring says so). From 4.2, the Granite chat template
tells the model to emit tool calls in an XML format instead:
<tool_call>
<function=get_weather>
<parameter=city>
Boston
</parameter>
</function>
</tool_call>
Checked against the granite-4.2-3b chat_template.jinja. The 4.0 and 4.1
templates use JSON inside <tool_call></tool_call> blocks, which parses
fine (verified on 4.0-h-micro and 4.1-3b on the Hub).
When a Granite 4.2 model makes a tool call on the local HF path:
parse_tools() returns [], so to_tool_calls() returns None and
mot.tool_calls stays empty (mellea/backends/huggingface.py:2179-2180)
- the tool never runs
- the raw XML stays in
mot.value and reaches the caller as response text
- nothing is logged:
to_tool_calls() only warns about unknown function
names
Observed: parse_tools() returns [] for the format the template
prescribes (repro below, no model needed).
Expected: the XML format parses. Failing that, the backend should say
the combination is unsupported instead of dropping the call silently.
Reproduction (no model, no GPU)
from mellea.backends.tools import parse_tools
# The shape the granite-4.2 chat template tells the model to emit
granite_xml = (
"<tool_call>\n<function=get_weather>\n<parameter=city>\nBoston\n"
"</parameter>\n</function>\n</tool_call>"
)
print(parse_tools(granite_xml))
# []
# The Granite 4.1 shape (JSON inside a tool_call block) parses
print(parse_tools(
"<tool_call>\n{\"name\": \"get_weather\", \"arguments\": {\"city\": \"Boston\"}}\n</tool_call>"
))
# [('get_weather', {'city': 'Boston'})]
Scope
- Affected:
LocalHFBackend with Granite 4.2 (checked on 3b), plus any
other model whose chat template prescribes this format.
- Not affected: Ollama and OpenAI-compatible servers (vLLM with a
tool-call parser, for example). They parse tool calls server-side and
hand Mellea structured calls. test_mellea_tool.py runs on the library
default granite4.2:3b through Ollama, so 4.2 tool calling works there.
- Why the tests miss it:
test_huggingface_tools.py uses Mistral 7B,
which emits JSON. test_huggingface.py::test_intrinsic_tools_in_generate_input
uses Granite 4.1-3b and only checks that the tools appear in the prompt;
it never parses output. test_tool_calls.py runs on granite4:micro-h
(4.0) through Ollama, where the server parses.
- The tool-calling docs (
docs/docs/how-to/execute-tool-calls.md,
docs/docs/how-to/tools-and-agents.md) never mention the HF backend, so
this combination was never documented. Users will still hit it.
Options
- Parse the XML format in
parse_tools(), tests first. The format is
fixed and spelled out in the template. Parameter values arrive as raw
text, so non-string arguments need decoding;
validate_tool_arguments() already coerces simple cases like "30" to
30. Preferred.
- Declare it unsupported. Document it, and have
LocalHFBackend
raise when tools are passed to a model whose chat template uses this
format. Detect it from the template, not the model name, so fine-tunes
are caught too.
Either is better than the silent drop. Whichever we pick, logging a
warning when the output contains tool-call markup but nothing parses would
make the failure visible straight away.
Suggested priority: p2. Served backends are fine, but tool calling on
the local HF path is broken on Granite 4.2, which is also the library
default.
Problem
LocalHFBackendextracts tool calls from raw model output withparse_tools()(mellea/backends/tools.py:474), which only handles JSONtool calls (its docstring says so). From 4.2, the Granite chat template
tells the model to emit tool calls in an XML format instead:
Checked against the granite-4.2-3b
chat_template.jinja. The 4.0 and 4.1templates use JSON inside
<tool_call></tool_call>blocks, which parsesfine (verified on 4.0-h-micro and 4.1-3b on the Hub).
When a Granite 4.2 model makes a tool call on the local HF path:
parse_tools()returns[], soto_tool_calls()returnsNoneandmot.tool_callsstays empty (mellea/backends/huggingface.py:2179-2180)mot.valueand reaches the caller as response textto_tool_calls()only warns about unknown functionnames
Observed:
parse_tools()returns[]for the format the templateprescribes (repro below, no model needed).
Expected: the XML format parses. Failing that, the backend should say
the combination is unsupported instead of dropping the call silently.
Reproduction (no model, no GPU)
Scope
LocalHFBackendwith Granite 4.2 (checked on 3b), plus anyother model whose chat template prescribes this format.
tool-call parser, for example). They parse tool calls server-side and
hand Mellea structured calls.
test_mellea_tool.pyruns on the librarydefault
granite4.2:3bthrough Ollama, so 4.2 tool calling works there.test_huggingface_tools.pyuses Mistral 7B,which emits JSON.
test_huggingface.py::test_intrinsic_tools_in_generate_inputuses Granite 4.1-3b and only checks that the tools appear in the prompt;
it never parses output.
test_tool_calls.pyruns ongranite4:micro-h(4.0) through Ollama, where the server parses.
docs/docs/how-to/execute-tool-calls.md,docs/docs/how-to/tools-and-agents.md) never mention the HF backend, sothis combination was never documented. Users will still hit it.
Options
parse_tools(), tests first. The format isfixed and spelled out in the template. Parameter values arrive as raw
text, so non-string arguments need decoding;
validate_tool_arguments()already coerces simple cases like"30"to30. Preferred.LocalHFBackendraise when tools are passed to a model whose chat template uses this
format. Detect it from the template, not the model name, so fine-tunes
are caught too.
Either is better than the silent drop. Whichever we pick, logging a
warning when the output contains tool-call markup but nothing parses would
make the failure visible straight away.
Suggested priority: p2. Served backends are fine, but tool calling on
the local HF path is broken on Granite 4.2, which is also the library
default.