🤖 Cloudflare Unveils Multimodal Decision Model Clef-omni
Cloudflare's announcement expands its open weight decision model family by adding audio and video input alongside text and images, inside a unified sequence with video and audio...

🤖 Cloudflare Unveils Multimodal Decision Model Clef-omni
Cloudflare's announcement expands its open weight decision model family by adding audio and video input alongside text and images, inside a unified sequence with video and audio...
🤖 Falcon ASR Model Sets New Benchmark for Arabic Speech Recognition
The specification is the interesting part. Falcon ASR is trained on Emirati, Modern Standard Arabic, other Gulf dialects and English, on recordings with...
#BenchmarksEvaluation #ModelTraining #SpeechAudio #AI #AIPulse
🤖 Suno Expands Music Generator to Spoken Audio Amid Copyright Pressure
The feature announced is a spoken text generator with matching background music, described as a single audio track, and it is being released as a beta....
🤖 Language Discrimination Narrows Multilingual Speech Model Gap
The gap between multilingual and monolingual speech models is a familiar one, and the intervention proposed here attacks it from the beginning rather than...
🤖 ElevenLabs v4 Takes on Voice Cloning with New Architecture
The gap that Eleven v4 is closing is a familiar one. Mistral's Voxtral TTS addresses the same problem with a model that uses two completely different modelling paradigms,...
🤖 SageMaker AI Enables Real-Time Speech Streaming
A voice agent cannot pause waiting for the model to finish speaking. SageMaker's demonstration of Qwen3 TTS on a bidirectional WebSocket stream shows that voice can start...
#InferenceOptimization #SpeechAudio #EnterpriseAI #AI #AIPulse
🤖 Nvidia's 100M-Parameter Diarization Model Leads VoiceArena Benchmark
Nemotron 3 Diarisation is not a transcription model at all, though it is part of the same stack. Its task is to separate the voices in a...
🤖 Qwen3-TTS Model Enables 24 kHz Voice Cloning on SageMaker
The claim is that a model can reproduce a speaker's voice from a short recording and a transcript, without retraining, which is the condition under which cloning is interesting...
🤖 Apple Advances On-Device Speech Transcription with Compressed Tokenizer
The paper is about a tokenizer on a speech transcription system, because that tokenizer is the bottleneck when the model is sparsely...
#InferenceOptimization #SpeechAudio #Multimodal #AI #AIPulse
🤖 NVIDIA's Diarization Model Tracks More Speakers
Nemotron 3 Diarization is a 100M parameter model that tracks up to eight speakers, including when voices overlap, and runs on Linux through NVIDIA NeMo on Ampere, Ada Lovelace,...
🤖 Google Unveils Advanced TTS Models with Enhanced Voice Control
The pitch is that the new models let developers direct the voice line by line, with stage directions written in the script and controls for acting...
#SpeechAudio #EnterpriseAI #SoftwareDevelopment #AI #AIPulse
🤖 Alibaba Cuts AI Audio Prices Dramatically
Alibaba has released Qwen Audio 3.1, which adds a new ASR model that cleans filler words and repeats, multi speaker identification with timestamps, and emotions detected in the...
🤖 Brain Implant Translates Speech and Gestures Simultaneously
A brain implant that translates both verbal and non verbal communication at once is not a one step upgrade. It is a consequence of trying to restore one function at a time,...
🤖 OpenAI's GPT-Live-1 Sets New Benchmark for Full-Duplex Speech Interaction
The benchmark numbers are the claim. In full duplex interactivity, GPT Live 1 reaches 80.1 percent compared to 45.4 percent for the older model,...