#1267: WG Revision: Web Speech API: SpeechRecognitionResult Timestamps

Visit on Github

Opened Aug 22, 2026

Specification

https://github.com/WebAudio/web-speech-api/pull/192

Explainer

https://github.com/WebAudio/web-speech-api/blob/main/explainers/speech-recognition-result-timestamps.md

Links

Feature 1:

The specification

Where and by whom is the work is being done?

Feedback so far

You should also know that...

W3C Audio WG/CG Meeting Discussion: This feature was presented and discussed during the W3C Audio WG/CG Teleconference on August 13, 2026 (see Meeting Minutes). The group aligned on developer needs for media timeline association (e.g. generating VTTCue captions, click-to-seek transcripts, and synchronizing captions with recorded video).

<!-- Content below this is maintained by @w3c-tag-bot -->

Track conversations at https://tag-github-bot.w3.org/gh/w3ctag/design-reviews/1267

Discussions

Comment by @alan33d Aug 24, 2026 (See Github)

There is some additional feedback we've received which requires updates to the Explainer. We will circle back here once those updates are made.

Comment by @alan33d Sep 10, 2026 (See Github)

The Explainer was updated, and TAG review can resume.

Comment by @lolaodelola Sep 23, 2026 (See Github)

Hi @alan33d,

Thank you for submitting this design review.

The only note I would make is that the Web IDL uses double for both the speech start and end times, we recommend DOMHighResTimeStamp as it allows for "comparison of timestamps, regardless of the user’s time settings" (Web Platform Design Principle 8.4), if there's a valid reason that the double is being used then we defer to the authors' judgment.

This issue is marked with "Resolution: decline" to reflect that we didn't do a thorough review as the architectural impact is minor. If you feel that there is a bigger architectural impact than what we've mentioned, please reopen the issue.

Comment by @alan33d Sep 23, 2026 (See Github)

Thanks for the feedback!

CC @padenot (Audio WG) since using double was their suggestion.

Using double aligns this API with attributes like BaseAudioContext.currentTime and HTMLMediaElement.currentTime which uses double and is measured in seconds. These are the timestamps we expect developers to compare these values with and avoid unit conversions when piping audio streams between Web Audio and VTTCue (https://github.com/WebAudio/web-speech-api/issues/191).