What is live here?
Incremental wrapper, not a causal encoder
The current 606M Phase37 w2v-BERT checkpoint scores 21.84% WER on the fixed development benchmark. The heavier offline research stack reached 18.63% but is not used for interactive partials. This wrapper produces partial text every ~4 seconds and finalizes on speech pauses or a 12-second cap.
Why this matters
Collect real failure cases
Each final utterance is stored with audio, model fingerprint, latency and your QA correction. That gives us high-value supervised examples for later ASR improvement.
Later speech generation
ASR-approved ≠ TTS-approved
Podcast and live audio can help recognition, but TTS requires additional speaker consent, identity, quality and voice-use rights. Those fields remain separate.