ChainForge-R1-SuperCoT
PublicA multi-stage pipeline that enhances Qwen2.5 language models with DeepSeek Reasoner's chain-of-thought capabilities. Implements the DeepSeek-R1 methodology through cold-start SFT, reasoning-oriented RL, rejection sampling, and optional model distillation.
作成時間:2025-01-25T03:13:53
更新時間:2025-02-24T17:02:19
10
Stars
0
Stars Increase