HomeAI Tutorial

RepoCapsule

Public

RepoCapsule is a Python toolkit for turning GitHub, local, and other text/code sources into clean JSONL corpora for LLM pre-training, fine-tuning, or RAG. It provides structure-aware chunking, robust Unicode decoding, pluggable quality/safety screening, and optional dataset card + deduplication support.

Creat2025-11-02T09:35:49
Update2025-12-08T11:05:33
2
Stars
0
Stars Increase