Skip to content
Open
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
14 changes: 14 additions & 0 deletions index.bs
Original file line number Diff line number Diff line change
Expand Up @@ -120,6 +120,8 @@ This does not preclude adding support for this as a future API enhancement, and
<li>The user agent may also give the user a longer explanation the first time speech input is used, to let the user know what it is and how they can tune their privacy settings to disable speech recording if required.</li>

<li>To mitigate the risk of fingerprinting, user agents MUST NOT personalize speech recognition when performing speech recognition on a {{MediaStreamTrack}}.</li>

<li>To mitigate micro-architectural timing attacks and hardware fingerprinting, user agents may reduce the resolution of {{SpeechRecognitionResult/audioStartTime}} and {{SpeechRecognitionResult/audioEndTime}} or introduce jitter, in accordance with the user agent's security and privacy policies (similar to [[HR-TIME-3]] and [[HTML]]).</li>
</ol>

<h3 id="implementation-considerations">Implementation considerations</h3>
Expand Down Expand Up @@ -258,6 +260,8 @@ interface SpeechRecognitionResult {
readonly attribute unsigned long length;
getter SpeechRecognitionAlternative item(unsigned long index);
readonly attribute boolean isFinal;
readonly attribute double audioStartTime;
readonly attribute double audioEndTime;
};

// A collection of responses (used in continuous mode)
Expand Down Expand Up @@ -356,6 +360,16 @@ interface SpeechRecognitionPhrase {
</dd>
</dl>

<h4 id="speechrecoresult-attributes">SpeechRecognitionResult Attributes</h4>

<dl>
<dt><dfn attribute for=SpeechRecognitionResult>audioStartTime</dfn> attribute</dt>
<dd>A {{double}} representing the start time of the audio segment corresponding to this recognition result, in seconds relative to the start of the audio stream consumed by the speech recognizer.</dd>

<dt><dfn attribute for=SpeechRecognitionResult>audioEndTime</dfn> attribute</dt>
<dd>A {{double}} representing the end time of the audio segment corresponding to this recognition result, in seconds relative to the start of the audio stream consumed by the speech recognizer.</dd>
</dl>

<p class=issue>The group has discussed whether WebRTC might be used to specify selection of audio sources and remote recognizers.
See <a href="https://lists.w3.org/Archives/Public/public-speech-api/2012Sep/0072.html">Interacting with WebRTC, the Web Audio API and other external sources</a> thread on public-speech-api@w3.org.</p>

Expand Down
Loading