Tren
dar
Dashboard
Papers
News
GitHub
Updates
KO
EN
Sign in
Updates
Papers
Dashboard
News
GitHub
“speech”
Papers, GitHub repos, and news related to this keyword, in one place.
Papers
12
All →
Semantic Scholar
Speech · Audio
355 citations
CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
OpenAlex
ML Methods
148 citations
A Recurrent Latent Variable Model for Sequential Data
arXiv
Speech · Audio
0 citations
Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs
arXiv
Speech · Audio
0 citations
WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS
arXiv
Safety
0 citations
Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity
arXiv
ML Methods
0 citations
Execution-State Capsules: Graph-Bound Execution-State Checkpoint and Restore for Low-Latency, Small-Batch, On-Device Physical-AI Serving
arXiv
Generative Models
0 citations
Reference-Driven Multi-Speaker Audio Scene Generation from In-the-Wild Priors
arXiv
Speech · Audio
0 citations
TRADE: Transducer-Augmented Decoder for Speech LLM
bioRxiv
NLP · LLMs
0 citations
Neural decoding of speech using deep neural ensembles
bioRxiv
NLP · LLMs
0 citations
Human-like sequential sound-to-meaning transfer drives artificial speech comprehension
Semantic Scholar
NLP · LLMs
985 citations
Intelligent Clinical Documentation: Harnessing Generative AI for Patient-Centric Clinical Note Generation
Semantic Scholar
NLP · LLMs
4 citations
The Structured Output Benchmark: A Multi-Source Benchmark for Evaluating Structured Output Quality in Large Language Models
GitHub
5
All →
Speech · Audio
Python
★ 30.5K
OpenBMB/VoxCPM
Product
Python
★ 11.9K
huggingface/speech-to-speech
Speech · Audio
Python
★ 18.8K
modelscope/FunASR
Speech · Audio
Python
★ 3.6K
OpenMOSS/MOSS-TTS
News
5
All →
Hacker News
Product
▲ 4
AI and brain-computer interface allow speechless ALS patient to work full-time
Google DeepMind
Product
▲ 0
Fluid, natural voice translation with Gemini 3.5 Live Translate
Hacker News
Product
Agents
Python
★ 4.5K
dograh-hq/dograh
▲ 4
Show HN: Imagent – agentic image/video/speech generation
Hacker News
▲ 13
If AI Outputs Aren't Speech, Who Has to Prove They're Human?
OpenAI
▲ 0
How we built a realtime system for responsive voice AI in six months