en

Please fill in your name

Mobile phone format error

Please enter the telephone

Please enter your company name

Please enter your company email

Please enter the data requirement

Successful submission! Thank you for your support.

Format error, Please fill in again

Confirm

The data requirement cannot be less than 5 words and cannot be pure numbers

High-Quality Training Datasets

Boost the performance of your AI models with our high-quality, ready-to-use training datasets.

Language

All

Data Type

All

6.9 million - Chinese Multi-disciplinary Questions Text Parsing And Processing Data

6.9 million - Chinese Multi-disciplinary Questions Text Parsing And Processing Data, including multiple disciplines in primary school, middle school, high school and university. Each questions contain title, answer, parse, type, subject, grade. The dataset can be used for large model subject knowledge enhancement tasks.
Chinese multi-disciplinary Questions LLM Text

1 million - Chinese Code Questions Text Parsing And Processing Data

1 million - Chinese Code Questions Text Parsing And Processing Data, including c, c++, python, java, javascript multiple language code questions. Each test contains title, answer, parse, and language fields. This data can help model building and solidify code programming skills for better performance in programming tasks.
Code Questions LLM Text

32 million - Science Subjects Questions Text Parsing And Processing Data

32 million - Science Subjects Questions Text Parsing And Processing Data, including mathematics, physics, chemistry and biology in primary, middle and high school and university. Each questions contain title, answer, parse, type, subject, grade. The dataset can be used for large model subject knowledge enhancement tasks.
Science Subjects Questions LLM Text

100,000 - Logical Reasoning Questions Text Parsing And Processing Data

100,000 - Logical Reasoning Questions Text Parsing And Processing Data, including logical thinking, mathematical logic, detective reasoning, etc. Each questions contain fields such as question, answer, parse, type, etc. This data can improve the complex cognitive ability of the model in multiple dimensions, making it better at handling logically complex and semantic rich tasks.
Logical Reasoning Questions LLM Text

140,000 - Contest Questions Text Parsing And Processing Data

140,000 - Contest Questions Text Parsing And Processing Data, including mathematics, physics, chemistry and biology in primary, middle school, high school and university. Each questions contain title, answer, parse, subject, grade, question type. The dataset can be used for large model subject knowledge enhancement tasks, while contributing to the overall intelligence development of the model.
Contest Questions LLM Text

5,000 Images of Turkish Natural Scene OCR Data

5,000 Turkish natural scenarios OCR data include a variety of natural scenarios and multiple shooting angles. For annotation, quadrilateral or polygon bounding box annotation and transcription for the texts were annotated in the data. This data can be used for tasks such as the Turkish language OCR.
OCR,Turkish,Natural scenes

534 Hours - Taiwanese Accent Mandarin Spontaneous Dialogue Smartphone speech dataset

534 Hours - Taiwanese Accent Mandarin Spontaneous Dialogue Smartphone speech dataset, collected from dialogues based on given topics. Transcribed with text content, timestamp, speaker's ID, gender and other attributes. Our dataset was collected from extensive and diversify speakers, geographicly speaking, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.
Accent Mandarin Taiwanese Spontaneous Dialogue

31 million Southeast Asian language news text dataset

This dataset is multilingual news data from Southeast Asia, covering four languages: Indonesian, Malay, Thai, and Vietnamese. The total amount of data exceeds 31 million, stored in JSONL format, with each record running independently in a row for efficient reading and processing. The data sources are extensive, covering various news topics, and can comprehensively reflect the social dynamics, cultural hotspots, and economic trends in Southeast Asia. This dataset can help large models improve their multilingual capabilities, enrich cultural knowledge, optimize performance, expand industry applications in Southeast Asia, and promote cross linguistic research.
Minor languages Southeast Asia NEWS Journalism

302 Hours - Tagalog(the Philippines) Scripted Monologue Smartphone speech dataset

Tagalog(the Philippines) Scripted Monologue Smartphone speech dataset, covers several domains, including chat, interactions, in-home, in-car, numbers and more, mirrors real-world interactions. Transcribed with text content and other attributes. Our dataset was collected from extensive and diversify speakers, geographicly speaking, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.
Tagalog The Philippines Smartphone Reading Scripted Monologue

822 Hours - Tagalog(the Philippines) Scripted Monologue Smartphone speech dataset

Tagalog(the Philippines) Scripted Monologue Smartphone speech dataset, collected from monologue based on given prompts, covering generic category; news category. Transcribed with text content. Our dataset was collected from extensive and diversify speakers(711 native speakers), geographicly speaking, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.
Tagalog Filipino reading

30,696 Pairs of Portrait Retouched Before and After Image Data

30,696 Pairs of Portrait Retouched Before and After Image Data. Data collection scenarios include indoor and outdoor scenes, the country distribution is Algeria, Egypt, Hungary, Poland, and Japan. Data types include portrait photos and wedding photos. In terms of data annotation, detailed retouching and annotationing are performed on the collected studio portrait data. The data can be used for tasks such as studio portrait retouching, PS segmentation, and portrait segmentation.
Portrait data Retouch Comparison Photos

9,000 Images of 180 People - Driver Gesture 21 Landmarks Annotation Data

9,000 Images of 180 People - Driver Gesture 21 Landmarks Annotation Data. This data diversity includes multiple age periods, multiple time periods, multiple gestures, multiple vehicle types, multiple time periods. For annotation, the vehicle type, gesture type, person nationality, gender, age and gesture 21 landmarks (each landmark includes the attribute of visible and invisible) were annotated. This data can be used for tasks such as driver gesture recognition, gesture landmarks detection and recognition.
DMS driver gesture gesture 21 landmarks static gesture dynamic gesture driver gesture recognition gesture landmarks detection gesture landmarks recognition
. . .
loading

loading

53f84e9c-a311-4356-903f-8bf14b925295