The model transcribes audio, provides word timestamps, and generates speech embeddings, operating within a single 16.9 MB file on the CPU.

The model transcribes audio, provides word timestamps, and generates speech embeddings, operating within a single 16.9 MB file on the CPU.