en

Please fill in your name

Mobile phone format error

Please enter the telephone

Please enter your company name

Please enter your company email

Please enter the data requirement

Successful submission! Thank you for your support.

Format error, Please fill in again

Confirm

The data requirement cannot be less than 5 words and cannot be pure numbers

LLM Datasets

Instantly enhance AI model performance with high quality off-the-shelf datasets.

Type

All
27
Image Caption
11
SFT Datasets
5
Pre-training Text
11

200000 text data in German, Spanish, French, and Italian

This dataset contains 200000 pieces of multilingual text content, covering four languages: French, German, Spanish, and Italian, with 50000 pieces of data in each language. These data cover over 200 categories, including architecture, animals, automobiles, catering, movies, zodiac signs, cybersecurity, and more. Provides abundant resources for diverse language learning and text analysis.
Network Data,German French Spanish Italian

1.5 million - Korean Test Questions Structured Analysis Processing Data

Korean Test Questions Structured Analysis Processing Data, around 1.5 million questions, contains question types, questions, answers, explanations, etc..For subjects, include [Primary School] Korean, Mathematics, English, Social Studies, Science; [Middle School] Korean, English, Mathematics, Science, Social Studies; [High School] Korean, English, Mathematics, Physics, Chemistry, Biology, History, Geography; question Types indlude single-choice question, fill-in question, true or false question, short answer question, etc. This dataset can be used for large-scale subject knowledge enhancement tasks.
K12 questions Text LLM Korean

250,000 English Animals Medical dataset

English Animals Medical dataset, Including hospital treatment details, hospital test results, hospital pet prescription, allergy test list, vaccination history for multiple animals, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.
Medical report Aimals Pet

130000 Chinese standard text parsing and processing data

This dataset contains all Chinese standard information data up to March 2024, including national standards 20000+, local standards 60000+, industry standards 20000+, and group standards 20000+, covering standard statuses such as "In force", "about to be implemented", and "abolished". Each type of standard has an independent. xlsx statistical table, which contains standard information fields such as standard name, standard number, release date, implementation date, status, and standard attributes.

140,000 - Contest Questions Text Parsing And Processing Data

140,000 - Contest Questions Text Parsing And Processing Data, including mathematics, physics, chemistry and biology in primary, middle school, high school and university. Each questions contain title, answer, parse, subject, grade, question type. The dataset can be used for large model subject knowledge enhancement tasks, while contributing to the overall intelligence development of the model.
Contest Questions LLM Text

100,000 Sets of ICONS Image Caption Data

100,000 Sets of ICONS Image Caption Data. The data includes two major categories of icons, namely 3D Style Icons and Vector Illustration Icons, totaling 17 subcategories. In terms of annotation, the icon descriptions are in Chinese, with a description length of about 30 characters. The data can be used for tasks such as graphic recognition and interface interaction.
ICONS Image caption

6.9 million - Chinese Multi-disciplinary Questions Text Parsing And Processing Data

6.9 million - Chinese Multi-disciplinary Questions Text Parsing And Processing Data, including multiple disciplines in primary school, middle school, high school and university. Each questions contain title, answer, parse, type, subject, grade. The dataset can be used for large model subject knowledge enhancement tasks.
Chinese multi-disciplinary Questions LLM Text

1 million - Chinese Code Questions Text Parsing And Processing Data

1 million - Chinese Code Questions Text Parsing And Processing Data, including c, c++, python, java, javascript multiple language code questions. Each test contains title, answer, parse, and language fields. This data can help model building and solidify code programming skills for better performance in programming tasks.
Code Questions LLM Text

32 million - Science Subjects Questions Text Parsing And Processing Data

32 million - Science Subjects Questions Text Parsing And Processing Data, including mathematics, physics, chemistry and biology in primary, middle and high school and university. Each questions contain title, answer, parse, type, subject, grade. The dataset can be used for large model subject knowledge enhancement tasks.
Science Subjects Questions LLM Text

loading

Tailor Your Data Now

Why off-the-shelf Datasets

  • Copyright

    Copyright

    Clear Coyright and Ready to Check
  • Security

    Security

    Properly Authorized Secure to Use
  • Professional

    Professional

    Designed and produced by AI data experts
  • Diversity

    Diversity

    Collected from a varity of real scenes
  • Cost Effective

    Cost Effective

    More Cost-Efficient Than Tailored Data
  • Efficiency

    Efficiency

    Ready-To-Go Deliver in Seconds
1985edd1-cb5b-42eb-a63c-185368e99672