Description
Both OpenAI chat clients decide the input_audio.format for audio content by substring-matching content.media_type against "wav" and "mp3":
python/packages/openai/agent_framework_openai/_chat_client.py:2121-2127 (Responses API)
python/packages/openai/agent_framework_openai/_chat_completion_client.py:1316-1322 (Chat Completions API)
The registered media type for MP3 is audio/mpeg (RFC 3008), which contains neither "mp3" nor "wav". It therefore falls into the else branch, and the whole content item is dropped: the Responses client logs Unsupported audio media type: audio/mpeg and returns {}, the Chat Completions client logs the same at debug level and returns {}. Only the non-standard spelling audio/mp3 produces an input_audio part, so an MP3 input is silently discarded depending on which of two equivalent media type strings the caller happened to write.
This is reachable through paths the framework itself documents and ships:
detect_media_type_from_base64() returns "audio/mpeg" for MP3 magic bytes (python/packages/core/agent_framework/_types.py:171), and the Content.from_data() docstring tells callers to use exactly that helper to fill in media_type (_types.py:718-721).
agent-framework-hosting-telegram defaults inbound audio messages to "audio/mpeg" (python/packages/hosting-telegram/agent_framework_hosting_telegram/_parsing.py:30-35), so a Telegram audio message forwarded to an OpenAI-based agent disappears from the request with no error surfaced to the caller.
Expected behavior: audio/mpeg maps to format="mp3", like audio/mp3 already does. Truly unsupported types (audio/ogg, audio/flac, ...) should keep being skipped as they are today.
Code Sample
from agent_framework import Content
from agent_framework.openai import OpenAIChatClient
client = OpenAIChatClient(model="gpt-4o-mini", api_key="...")
# ID3-tagged MP3 bytes: core's detector labels these "audio/mpeg"
mp3 = Content.from_data(data=b"ID3\x03\x00\x00\x00\x00\x00\x00", media_type="audio/mpeg")
print(client._prepare_content_for_openai("user", mp3))
# actual -> {} (warning: Unsupported audio media type: audio/mpeg)
# expected -> {'type': 'input_audio', 'input_audio': {'data': ..., 'format': 'mp3'}}
same_file = Content.from_data(data=b"ID3\x03\x00\x00\x00\x00\x00\x00", media_type="audio/mp3")
print(client._prepare_content_for_openai("user", same_file)["input_audio"]["format"]) # 'mp3'
Error Messages / Stack Traces
No exception is raised; the content is dropped silently. Covered by the two audio tests, which fail on the added audio/mpeg case:
> assert result["type"] == "input_audio"
E KeyError: 'type'
packages\openai\tests\openai\test_openai_chat_completion_client.py:613
Package Versions
agent-framework-core: 1.19.0, agent-framework-openai: 1.14.4 (source main @ 273e635)
Additional Context
OpenAI's input_audio part only accepts wav and mp3, so this is not about adding formats — it is about accepting the standard name for the format that is already supported.
Description
Both OpenAI chat clients decide the
input_audio.formatfor audio content by substring-matchingcontent.media_typeagainst"wav"and"mp3":python/packages/openai/agent_framework_openai/_chat_client.py:2121-2127(Responses API)python/packages/openai/agent_framework_openai/_chat_completion_client.py:1316-1322(Chat Completions API)The registered media type for MP3 is
audio/mpeg(RFC 3008), which contains neither"mp3"nor"wav". It therefore falls into theelsebranch, and the whole content item is dropped: the Responses client logsUnsupported audio media type: audio/mpegand returns{}, the Chat Completions client logs the same at debug level and returns{}. Only the non-standard spellingaudio/mp3produces aninput_audiopart, so an MP3 input is silently discarded depending on which of two equivalent media type strings the caller happened to write.This is reachable through paths the framework itself documents and ships:
detect_media_type_from_base64()returns"audio/mpeg"for MP3 magic bytes (python/packages/core/agent_framework/_types.py:171), and theContent.from_data()docstring tells callers to use exactly that helper to fill inmedia_type(_types.py:718-721).agent-framework-hosting-telegramdefaults inboundaudiomessages to"audio/mpeg"(python/packages/hosting-telegram/agent_framework_hosting_telegram/_parsing.py:30-35), so a Telegram audio message forwarded to an OpenAI-based agent disappears from the request with no error surfaced to the caller.Expected behavior:
audio/mpegmaps toformat="mp3", likeaudio/mp3already does. Truly unsupported types (audio/ogg,audio/flac, ...) should keep being skipped as they are today.Code Sample
Error Messages / Stack Traces
No exception is raised; the content is dropped silently. Covered by the two audio tests, which fail on the added
audio/mpegcase:Package Versions
agent-framework-core: 1.19.0, agent-framework-openai: 1.14.4 (source
main@273e635)Additional Context
OpenAI's
input_audiopart only acceptswavandmp3, so this is not about adding formats — it is about accepting the standard name for the format that is already supported.