en

Please fill in your name

Mobile phone format error

Please enter the telephone

Please enter your company name

Please enter your company email

Please enter the data requirement

Successful submission! Thank you for your support.

Format error, Please fill in again

Confirm

The data requirement cannot be less than 5 words and cannot be pure numbers

m.nexdata.datatang.com

20 Hours - American English Male Voice TTS Dataset

TTS english dataset
speech synthesis dataset
TTS male voice dataset
male voice dataset for tts
American English speech synthesis dataset

This dataset contains 20 hours of American English male voice recordings. It is recorded by Americans (native English speakers) with authentic accent. The phoneme coverage is balanced. Professional phonetician participates in the annotation. It is suitable for text-to-speech (TTS) model training, phoneme recognition research, and AI voice development.

Paid Datasets
This is a paid datasets for commercial use, research purpose and more. Licensed ready made datasets help jump-start AI projects.
SpecificationsSpecifications
Format
48,000Hz, 24bit, uncompressed wav, mono channel;
Recording environment
professional recording studio;
Recording content
general narrative sentences, interrogative sentences, etc;
Speaker
male, 20-30 years old, young and positive voice;
Device
microphone;
Language
American English;
Annotation
word and phoneme transcription, four-level prosodic boundary annotation;
Application scenarios
speech synthesis.
Sample Sample
  • Audio

    Look- at- the way they hear operas- and- see oil paintings%.L UH1 K3 / AE1 T / DH AX0 / W EY1 / DH EY1 / HH IY1 R / AA1 . P R AX0 Z / AX0 N D / S IY1 / OY1 L / P EY1 N . T IH0 NG Z

  • Audio

    Was there some discussion- about- whether I should- speak%?W AX1 Z / DH EH1 R / S AH1 M / D IH0 . S K AH1 . SH AX0 N3 / AX0 . B AW1 T / W EH1 . DH ER0 / AY1 / SH UH1 D / S P IY1 K

  • Audio

    The focus- of this chapter is the American revolution%.DH AX0 / F OW1 . K AX0 S3 / AX1 V / DH IH1 S / CH AE1 P . T ER0 / IH1 Z / DH IY0 / AX0 . M EH1 . R IH0 . K AX0 N / R EH2 . V AX0 . L UW1 . SH AX0 N

  • Audio

    Can I go calling any time%?K AE1 N / AY13 / G OW1 / K AO1 . L IH0 NG / EH1 . N IY0 / T AY1 M

  • Audio

    Is- it really take- time%?IH1 Z / IH1 T / R IY1 . AX0 . L IY0 / T EY1 K / T AY1 M3

Recommended DatasetsRecommended Dataset
Tell Us Your Special Needs

Current Project Maturity

Early exploration (no concrete specs yet)
Defined goals, need professional guidance
Active development or optimization phase
Data & labeling experts with clear specifications

By submitting, I agree to the Privacy Protection

What languages and voice characteristics are covered by Nexdata’s speech synthesis datasets?

Nexdata offers speech synthesis datasets covering a broad range of languages, dialects, accents, and voice types, supported by extensive global language resources. Our datasets include diverse speakers, speaking styles, emotions, and recording scenarios to support natural and expressive Text-to-Speech (TTS) model development.

Can Nexdata customize speech synthesis datasets for specific languages or requirements?

Yes. If our off-the-shelf TTS datasets do not fully meet your requirements, Nexdata provides flexible custom data collection, transcription, annotation, and quality control services. We can customize datasets based on target languages or dialects, speaker profiles, voice characteristics, emotions, speaking styles, recording environments, and data volume.

How does Nexdata ensure the quality and scalability of speech synthesis datasets?

Nexdata applies multi-stage quality control throughout voice data collection, transcription, annotation, and validation. Combined with our extensive language resources and scalable collection capabilities, we can support both large-scale multilingual TTS projects and specialized datasets for specific voices, accents, emotions, or speech scenarios.

718c92f0-340c-4cad-8594-dc4df59b7495

a232306b-8076-48a5-ad53-22be237bf895