Home
Information

AI Dataset Collection

Large-scale datasets and benchmarks for training, evaluating, and testing models to measure

Tools

Intelligent Document Recognition

Comprehensive Text Extraction and Document Processing Solutions for Users

AI Tutorial

Stratified-LLM-Subsets-100K-1M-Scale

Public

Stratified LLM Subsets delivers diverse training data at 100K-1M scales across pre-training (FineWeb-Edu, Proof-Pile-2), instruction-following (Tulu-3, Orca AgentInstruct), and reasoning distillation (Llama-Nemotron). Embedding-based k-means clustering ensures maximum diversity across 5 high-quality open datasets.

Creat2025-09-30T16:35:20
Update2025-10-06T03:43:33
https://amanpriyanshu.github.io/Stratified-LLM-Subsets-100K-1M-Scale/
1
Stars
0
Stars Increase