Repository navigation
How to use the VAD API? #5824
Replies: 1 comment
|
There's a documented Request/response // POST /v1/vad
{
"model": "silero-vad",
"audio": [0.001, -0.002, ...] // float32[], 16kHz mono PCM samples
}{
"segments": [
{ "start": 0.5, "end": 2.3 },
{ "start": 3.1, "end": 5.8 }
]
}If Model setup Add a model config: name: silero-vad
backend: silero-vadLocalAI will pull the gallery entry (backed by If you're already running the Full example (Python) import soundfile as sf
import numpy as np
import requests
audio, sr = sf.read("speech.wav") # must be 16kHz mono
if audio.ndim > 1:
audio = audio.mean(axis=1)
samples = audio.astype(np.float32).tolist()
vad = requests.post(
"http://localhost:8080/v1/vad",
json={"model": "silero-vad", "audio": samples},
).json()
if not vad["segments"]:
print("no speech detected, skipping whisper")
else:
# proceed to /v1/audio/transcriptions as usual
...Source and doc references:
|
Uh oh!
There was an error while loading. Please reload this page.
I cannot find any informations about using the VAD (voice activity detection) API #4204
The current situation is: when I pass a silent audio file to the whisper API, it returns "Thank you.".
I like to use the VAD API to prefent this.
Any ideas?
All reactions