Extremely Fast and Accurate Transcription with Granite Speech 5.0 Turbo CTC
IBM has released two new models in the Granite Speech family, offering strong accuracy and unprecedented speed. The models, Granite Speech 5.0 Turbo CTC, can transcribe more than 3.5 hours of speech in one second using batched inference.
Intelligence analysis by Llama

IBM has released two new models in the Granite Speech family, offering strong accuracy and unprecedented speed. The models, Granite Speech 5.0 Turbo CTC, can transcribe more than 3.5 hours of speech in one second using batched inference. The models are encoder-only, unlike prior Granite Speech models, and have a small memory footprint of only 470M parameters.
Imagine you're having a conversation with a friend, and you want to understand what they're saying. The Granite Speech 5.0 models are like super-fast and accurate listeners that can transcribe what you're saying in real-time, allowing you to focus on the conversation without worrying about taking notes.
Analysis
Performance Metrics
The Granite Speech 5.0 models have been evaluated on the public, English short-form test sets from the OpenASR Leaderboard. The results show that the models offer high accuracy, with the noncommercial model scoring an aggregate 4.85% WER and the Apache 2.0 model scoring 5.00% WER. The models also demonstrate unprecedented aggregate throughput in excess of 12,600 RTFx.
Model Architecture
The Granite Speech 5.0 models are encoder-only models, unlike prior Granite Speech models which comprise an acoustic encoder, projector, and Granite LM with LoRA adapters. The encoder-only design provides strong transcription performance, a small memory footprint of only 470M parameters, and over 20x faster throughput than previous Granite Speech models.
Training Data
The Granite Speech 5.0 models are trained with a combination of natural and synthetic data. The natural data includes a large corpus of transcribed speech, while the synthetic data includes artificially generated speech. The models are trained using a combination of supervised and unsupervised learning techniques.
Key points
- IBM has released two new models in the Granite Speech family, offering strong accuracy and unprecedented speed.
- The models, Granite Speech 5.0 Turbo CTC, can transcribe more than 3.5 hours of speech in one second using batched inference.
- The models are encoder-only, unlike prior Granite Speech models, and have a small memory footprint of only 470M parameters.
- The models have been evaluated on the public, English short-form test sets from the OpenASR Leaderboard, demonstrating high accuracy and unprecedented aggregate throughput.
- The models are trained with a combination of natural and synthetic data, using a combination of supervised and unsupervised learning techniques.
The release of these models could lead to significant advancements in the field of speech recognition, enabling faster and more accurate transcription capabilities. This could have a major impact on industries such as customer service, healthcare, and finance, leading to improved efficiency and productivity.
However, the development and deployment of these models could also raise concerns about data privacy and security. As with any new technology, there is a risk that sensitive information could be compromised, leading to potential security breaches.



