en

Please fill in your name

Mobile phone format error

Please enter the telephone

Please enter your company name

Please enter your company email

Please enter the data requirement

Successful submission! Thank you for your support.

Format error, Please fill in again

Confirm

The data requirement cannot be less than 5 words and cannot be pure numbers

LLM Datasets

Instantly enhance AI model performance with high quality off-the-shelf datasets.

Type

All
32
Image Caption
13
SFT Datasets
6
Pre-training Text
13

114,000 - Chinese Contest Questions Text Parsing And Processing Data

114,000 - Chinese Contest Questions Text Parsing And Processing Data, including mathematics, physics, chemistry and biology in primary, middle school, high school and university. Each questions contain title, answer, parse, subject, grade, question type. The dataset can be used for large model subject knowledge enhancement tasks, while contributing to the overall intelligence development of the model.
Contest Questions LLM Text

20,846 Groups Image Caption Data of Cookbook

20,846 Groups Image Caption Data of Cookbook. Each set of recipes contains 4-18 images and a text description for each image. Cuisines include Chinese Cuisine, Western Cuisine, Korean Cuisine, Japanese Cuisine and so on. Description languages are Chinese and English. In terms of text length, the Chinese description should be no less than 15 words, and the English description should be no less than 30 words. The data can be used for recipe recommendations, culinary education and more.
Cookbook Image caption AIGC

480000 corrected texts in German, Spanish, French, Italian

This dataset focuses on the four major European languages (French, German, Spanish, Italian) and contains 480000 pairs of original and corrected text pairs. Each piece of data is presented in JSON format, including two fields: input (raw text) and output (corrected text), which can assist in natural language processing, machine translation, and language teaching research.
German French Spanish Italian proofreading

100,440 Sets - Chart Image Describe&QA Data

100,440 Sets - Chart Image Describe&QA Data,100,440 sets in total, including 79,044 in Chinese, 21,396 in English, for image content, provide structured descriptions and Q&A annotations in Chinese or English, with 1 structured description and 2 Q&A.
QA Chart Description

50,000 Sets - Image Editing Data

50,000 Sets - Image Editing Data. The types of editing include target removal, target addition, target modification, and target replacement. The editing targets cover scenes such as people, animals, products, plants, and landscapes. In terms of annotation, according to the editing instructions, the targets that need to be edited in the image are cropped and annotated for removal/addition/modification/replacement. The data can be used for tasks such as image synthesis, data augmentation, and virtual scene generation.
Image Editing

200000 text data in German, Spanish, French, and Italian

This dataset contains 200000 pieces of multilingual text content, covering four languages: French, German, Spanish, and Italian, with 50000 pieces of data in each language. These data cover over 200 categories, including architecture, animals, automobiles, catering, movies, zodiac signs, cybersecurity, and more. Provides abundant resources for diverse language learning and text analysis.
Network Data,German French Spanish Italian

250,000 English Animals Medical dataset

English Animals Medical dataset, Including hospital treatment details, hospital test results, hospital pet prescription, allergy test list, vaccination history for multiple animals, enhancing model performance in real and complex tasks. Quality tested by various AI companies. We strictly adhere to data protection regulations and privacy standards, ensuring the maintenance of user privacy and legal rights throughout the data collection, storage, and usage processes, our datasets are all GDPR, CCPA, PIPL complied.
Medical report Aimals Pet

140,000,000 - Chinese Judgment Documents Text Parsing And Processing Data

140 million judgment documents text data from 1998 to December 2023. Each piece of content contains title, abstract, body, date and author.A detailed data dictionary is provided.
Judgment documents Text LLM

1.5 million - Korean Test Questions Structured Analysis Processing Data

Korean Test Questions Structured Analysis Processing Data, around 1.5 million questions, contains question types, questions, answers, explanations, etc..For subjects, include [Primary School] Korean, Mathematics, English, Social Studies, Science; [Middle School] Korean, English, Mathematics, Science, Social Studies; [High School] Korean, English, Mathematics, Physics, Chemistry, Biology, History, Geography; question Types indlude single-choice question, fill-in question, true or false question, short answer question, etc. This dataset can be used for large-scale subject knowledge enhancement tasks.
K12 questions Text LLM Korean

loading

Tailor Your Data Now

Why off-the-shelf Datasets

  • Copyright

    Copyright

    Clear Coyright and Ready to Check
  • Security

    Security

    Properly Authorized Secure to Use
  • Professional

    Professional

    Designed and produced by AI data experts
  • Diversity

    Diversity

    Collected from a varity of real scenes
  • Cost Effective

    Cost Effective

    More Cost-Efficient Than Tailored Data
  • Efficiency

    Efficiency

    Ready-To-Go Deliver in Seconds
62a288fb-a1d0-4f09-beeb-04a5abae1c64