In the ML<>Apps sync, the Apps team mentioned that it would be useful to have a follow-along reading experience while audio is playing:
To support this, we are going to add word-level timestamps so that the spoken audio can be synchronized with the corresponding text in a way that is compatible with standard web caption formats such as W3C WebVTT.