303 Hours English-Mandarin Bilingual Speech Dataset – Mobile Phone Recordings
This dataset contains 303 hours of Chinese-English mixed speech, collected from monologue based on given Chinese and English Mixed prompts, covering general and human-computer interaction domains. Transcribed with text content and other attributes. Our dataset was collected from extensive and diversify speakers(1,113 speakers), geographicly speaking, enhancing model performance in real and complex tasks like ASR, TTS, code-switching, and bilingual speech-related AI tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.
chinese-english speech dataset bilingual speech dataset mixed language speech dataset chinese-english audio dataset code-switching speech dataset