Dataset

MiPS Lab Datasets

Śabdavṛndam (शब्दवृन्दम्): A low resource Sanskrit dataset for ASR

Śabdavṛndam is a comprehensive low-resource Sanskrit dataset designed for Automatic Speech Recognition (ASR) research. This dataset addresses the challenges of working with low-resource languages and provides a valuable resource for researchers working on Sanskrit speech recognition.

The dataset includes carefully curated audio samples with corresponding transcriptions, making it suitable for training and evaluating ASR models for Sanskrit language processing.

Features:

  • Low-resource language dataset for Sanskrit
  • Comprehensive audio samples with transcriptions
  • Suitable for ASR model training and evaluation
  • Designed to address challenges in low-resource language processing

Request Dataset Access

To request access to the Śabdavṛndam dataset, please send an email with your research purpose and affiliation details.

Send Email Request