ChainForge-R1-SuperCoT
PublicA multi-stage pipeline that enhances Qwen2.5 language models with DeepSeek Reasoner's chain-of-thought capabilities. Implements the DeepSeek-R1 methodology through cold-start SFT, reasoning-oriented RL, rejection sampling, and optional model distillation.
Creat:2025-01-25T03:13:53
Update:2025-02-24T17:02:19
10
Stars
0
Stars Increase