🪴 Anil's Garden

❯

CASPER: A Large Scale Spontaneous Speech Dataset

19 Dec 20251 min read

paper

Title: CASPER: A Large Scale Spontaneous Speech Dataset
Authors: Cihan Xiao, Ruixing Liang, Xiangyu Zhang, Mehmet Emre Tiryaki, Veronica Bae, Lavanya Shankar, Rong Yang, Ethan Poon, Emmanuel Dupoux, Sanjeev Khudanpur, Leibny Paola Garcia Perera
Published: 30th May 2025 (Friday) @ 22:03:59
Link: http://arxiv.org/abs/2506.00267v3

Abstract

The success of large language models has driven interest in developing similar speech processing capabilities. However, a key challenge is the scarcity of high-quality spontaneous speech data, as most existing datasets contain scripted dialogues. To address this, we present a novel pipeline for eliciting and recording natural dialogues and release our dataset with 100+ hours of spontaneous speech. Our approach fosters fluid, natural conversations while encouraging a diverse range of topics and interactive exchanges. Unlike traditional methods, it facilitates genuine interactions, providing a reproducible framework for future data collection. This paper introduces our dataset and methodology, laying the groundwork for addressing the shortage of spontaneous speech data. We plan to expand this dataset in future stages, offering a growing resource for the research community.

Graph View

Backlinks

Datasets
Speech and Audio - Rolodex - Papers, Models and Releases

Website
Bluesky
Twitter/X
GitHub
LinkedIn
Instagram
Goodreads
Letterboxd
🍋

🪴 Anil's Garden

Explorer

CASPER: A Large Scale Spontaneous Speech Dataset

Graph View

Backlinks