AIBase
Home
AI NEWS
AI Tools
GEO & AEO
MCP
AI Models
EN

AI News

View More

​谷歌 DeepMind 通过强化学习微调提升 AI 决策能力

近期,谷歌 DeepMind 团队与约翰・开普勒林茨大学 LIT AI 实验室合作,开展了一项关于人工智能语言模型的新研究。他们采用了强化学习微调(RLFT)技术,旨在提升语言模型的决策能力。这项研究的重点在于,通过思维链的强化训练,解决了模型在决策过程中存在的一些关键问题。随着大数据的应用,现有的语言模型已经展现出处理文本的超越能力,甚至能够在交互环境中做出基于知识的决策。然而,这些模型在实际决策时却常常出现 “纸上谈兵” 的问题,虽然能推导出正确的策略,却无

15.6k 08-08
​谷歌 DeepMind 通过强化学习微调提升 AI 决策能力

Models

View More

DeepSeek-R1

Deepseek

DeepSeek-R1

$4

Input tokens/M

$16

Output tokens/M

32

Context Length

Spark X1

Iflytek

Spark X1

$2

Input tokens/M

-

Output tokens/M

-

Context Length

qwq-plus

Alibaba

qwq-plus

$1.6

Input tokens/M

$4

Output tokens/M

128

Context Length

DeepSeek-R1-Distill-Qwen-7B

Deepseek

DeepSeek-R1-Distill-Qwen-7B

$1

Input tokens/M

-

Output tokens/M

8

Context Length

Baichuan-M2-32B

Baichuan

Baichuan-M2-32B

-

Input tokens/M

-

Output tokens/M

32

Context Length

ERNIE X1.1 Preview

Baidu

ERNIE X1.1 Preview

$1

Input tokens/M

$4

Output tokens/M

64

Context Length

o1

Openai

o1

$105

Input tokens/M

$420

Output tokens/M

200

Context Length

Qwen_v2.5_3b_Instruct

Alibaba

Qwen_v2.5_3b_Instruct

$1

Input tokens/M

-

Output tokens/M

32

Context Length

o1-mini

Openai

o1-mini

$21

Input tokens/M

$84

Output tokens/M

128

Context Length

o1-preview

Openai

o1-preview

$105

Input tokens/M

$420

Output tokens/M

128

Context Length

AIBase
Empowering the future, your artificial intelligence solution think tank
English简体中文繁體中文にほんご
FirendLinks:
AI Newsletters AI ToolsMCP ServersAI NewsAI MarketingLLM LeaderboardAI Ranking
© 2026AIBase
Business CooperationSite Map