--- title: 强化学习之父萨顿中国演讲:大模型无法成为超级智能,真正的AI还在后面(中英全文) source: 市场资讯 datetime: 2026-07-21T20:40:00+08:00 importance: 0 canonical_url: https://finance.sina.com.cn/wm/2026-07-21/doc-iniirfsp2196633.shtml --- (来源:财经会议圈) 编辑:米奇    来源:财经会议圈2026世界人工智能大会(WAIC)上强化学习之父、图灵奖得主理查德·萨顿的发表演讲,这是萨顿宣布创办AI公司Oak Lab、离开Keen Technologies后的首次公开演讲,延续了他近两年“AI应从依赖静态数据转向与世界互动学习”的核心主张,核心追问是: 当下如日中天的大语言模型(LLM)是否是AI演进的阶段性产物? 他的答案是肯定的: 大模型会成为AI生态的重要组成部分,但不会发展为“超级心智”,更现实的定位是“超级百科全书”。 核心理论支撑:《苦涩的教训》 萨顿从自己2019年提出的《苦涩的教训》切入,该理论总结了AI七十年的发展规律:研究者最初倾向于将人类知识、规则写入系统,短期见效快但长期会遇到瓶颈,最终会被更依赖算力规模、搜索与通用机器学习的新方法取代,国际象棋、NLP等领域都已验证这一规律。 他认为大模型对《苦涩的教训》“既符合又违背”: 符合之处在于其崛起本身就依赖算力扩张、海量数据与规模化训练,而非人工规则,契合该规律的发展方向;违背之处在于其核心训练数据是人类的既有知识,擅长调用、重组已有知识,却难以像人类一样在真实世界中探索、创造新知识,这会成为其长期发展的天花板——“知道人类说过什么”和“在世界中发现新事物”本质完全不同。 “大世界视角”下的AI发展路径 萨顿提出“大世界假说”: 世界的复杂度远高于任何单一心智,即便模型参数、训练数据规模再大,也不可能容纳整个世界,所谓“无所不知的超级心智”是不切实际的想象。 这一观点也与人类学家项飙的观点形成呼应: 信息覆盖不等于理解深度,过度追求全知反而会消解人与人的交流、公共空间的活力。 基于此,他勾勒了三类AI的未来走向: 专用系统:如自动驾驶,仅需在特定场景完成任务,无需完整心智; 完整心智AI:是他真正期待的方向,类似人类可在一生中持续体验、学习、创造、形成专属经验,不会无所不知,但拥有独立的经验体系; 基础模型(大模型):定位为“超级百科全书”,承担保存、组织、传播人类共识性知识的角色,不应强行回答超出知识范围的问题,核心目标是可靠性而非“显得更聪明”——要准确呈现已知边界,坦诚承认未知,减少“幻觉”问题,根据不同用户的需求灵活解释知识。 AI时代的人之价值 大模型建立在人类将经验转化为标准化符号的“复制时代”基础上,其构建的文本世界容易让人误以为是现实;即时生成答案的特性,还可能弱化人类倾听、分析、独立判断的能力。 因此基础模型可以帮人接近世界,却无法替代人经历世界 创造新知识、理解具体情境、积累个人经验的责任,始终属于人类自身。 理查德・萨顿 WAIC 2026 演讲完整中文译文(全文) 演讲主题:强化学习第一性:从经验培育超级智能时间:2026 年 7 月 19 日背景:连续第二年登上 WAIC 舞台,离职 Keen Technologies、创立 Oak Lab 后的首次公开演讲 各位下午好,感谢大会连续第二年邀请我来到世界人工智能大会。能再次来到上海,和大家分享人工智能未来的思考,我由衷感到荣幸。 就在一周前,我对外公布了职业生涯的重要变动。在约翰・卡马克创立的 Keen Technologies 深耕三年、收获颇丰后,我和长期科研搭档库拉姆・贾维德一同离职,创立了全新 AI 公司 Oak Lab。今天是我官宣创业后的首次公开露面,接下来我会完整阐述支撑 Oak Lab 全部研发方向的核心理念。 过去十年,大语言模型取得了表层层面极为亮眼的进展,但我们必须客观认清一个事实: 当下主流基座模型的底层逻辑存在根本性缺陷。 现在所有大语言模型都采用离线一次性训练模式,依托人工整理好的静态数据集完成学习,训练结束后模型参数就彻底固定。 它们可以生成流畅通顺的文字,却无法通过和现实世界实时互动实现持续、真正的自主学习。 人类智能的运行逻辑恰恰与之相反。 从婴儿阶段开始,人类依靠行动、感官感知、环境反馈与反复试错不间断成长学习。我们不会依赖一套一次性、海量的预制数据完成认知构建;每一次和外界交互都会产生全新体验,每一段新体验都会重塑我们对世界的认知框架。 如今的人工智能还停留在 “静态人类数据时代”,而通用人工智能发展的下一阶段,必然是 “自主生成实时经验时代”。二者的差异绝非局部优化层面的小区别,而是两套完全割裂的技术范式。这也正是 Oak Lab 选择的道路,和当下行业主流的大模型路线截然不同。 我们时常高估当前 AI 的真实能力。这类系统擅长复刻数据中存在的文本模式,却不具备第一人称主观体验、自主探索世界的能力,更无法主动衍生出新的子问题,推动自身认知迭代。它们只是记忆人类已经总结好的知识,不能依靠自身探索发现全新规律。 今天我想再次提起我早年撰写的《苦涩的教训》一文:数十年来,AI 研究者耗费大量精力,把人类总结的先验规则硬编码进系统里。历史已经证明,这套思路长期走不通。真正的通用智能,必须完全依靠自身与环境的交互去推导世界规则、逻辑与世界模型,不能背负大量人为预设的偏见约束。 我十分敬重约翰・卡马克,也感恩所有在 Keen 共事过的同事。在 Keen 的三年充满价值,我们在打造基于强化学习的具身智能这件事上,目标高度一致。 但我意识到,如果要全身心落地我们坚信的 “经验优先” 技术架构,我需要一家完全独立的机构,所有资源只为这一条技术使命服务。在 Oak Lab 内部,不存在互相冲突的研发优先级,全部人力算力都聚焦研发一类智能体:依靠运行时实时交互持续学习,而非依赖静态离线数据集。 Oak Lab 分为两条互补的业务板块:商业化实验室负责将强化学习基础理论落地为可商用的具身智能体;我个人名下的非营利开放心智研究院,专门开展开源共享的基础前沿研究。Oak Lab 商业化业务产生的收益,会全额反哺开放学术研究,形成可持续的自循环体系。 我们长期的硬件与能效目标十分清晰:打造万亿参数规模的通用智能体,支持实时学习与规划,功耗对标人脑,仅 20 瓦左右。当前主流大模型动辄数千瓦功耗,核心痛点是存储与计算单元之间无休止的数据搬运,运行效率极低。只有采用类脑并行架构,实现数据就地运算、减少数据迁移,才能完成能效层面的跨越式突破。 Oak Lab 自研的核心技术框架命名为 OaK,全称 Options and Knowledge(行为选项与认知知识)。它补齐了当下强化学习、大语言模型体系中最大的短板 —— 自主分层抽象能力。 所谓 “行为选项”,指一套可复用的高层行为模板,是带有明确触发条件与终止规则的完整动作序列。比如喝咖啡、开门、导航前往充电站,全部属于行为选项。和底层单一零散动作不同,行为选项是成套可复用技能,智能体可以在全新场景中灵活调用。 OaK 体系中的 “知识”,是智能体自主搭建的预测世界模型,会记录执行每一项行为选项后产生的对应结果。最关键的一点:整套知识全部来源于智能体自身的实时交互体验,而非从人类文本语料库导入。 OaK 的学习闭环完全在线运行,单次处理样本量固定为 1。每一条全新交互体验都会立刻触发参数更新,不存在延后分批离线重训练流程。智能体探索环境、产生新经验、创造全新行为选项、扩充自身认知模型、衍生新的待求解问题,构成一套无限自我优化的闭环。 这套架构从根源上解决了主流大模型的核心缺陷:训练阶段和推理阶段完全割裂。现有模型预训练完成后学习就彻底停止;而 OaK 智能体只要持续和环境交互,学习过程就不会中断。 目前行业 99% 的商业投资,都流向基于静态数据集的生成式模型。但这条路线很快会撞上无法突破的硬性天花板:泛化能力存在上限、算力成本持续高企、运行稳定性脆弱,除人工微调外不存在自主进化能力。 通用人工智能迎来拐点,需要三大条件同步成熟:第一,具备依托实时交互持续学习能力的智能体架构,也就是 Oak Lab 深耕的 OaK 路线;第二,接近人脑能效标准的硬件体系,实现 20 瓦功耗下运行万亿参数智能体;第三,大规模落地真实具身场景,产出可量化的产业实际价值。 当下主流 AI 厂商没有一家同时达成以上三点,这正是 Oak Lab 要抓住的技术与市场窗口期。 很多人问我,具身智能体会不会取代大语言模型?我的答案是否定的,语言模型只会成为整体架构里一个次要子模块。语言只是人类全部体验中非常狭窄的一小部分;真正的通用智能,必须先掌握完整的感知、运动、空间、因果交互能力,语言仅仅作为次级沟通工具存在。 数十年来,我们一直在训练 AI 模仿人类的文字、语音、书面知识。下一个 AI 时代,我们要训练 AI 像人类与所有生命体一样,亲身感知、体验真实世界。 我十分感谢约翰・卡马克与 Keen 全体团队,和大家一同走过这段研究旅程。如今依托 Oak Lab,我会将全部精力投入打造以实时交互经验为根基的人工智能。我们不满足于对现有模型做局部微调,而是要彻底重构底层学习范式。 感谢各位,期待和中国各地的科研人员、企业携手,共同推进这场以自主经验为核心的 AI 全新变革。 Richard Sutton WAIC 2026 Full English Speech Speech Title: First Principles of Reinforcement Learning: Growing Superintelligence from ExperienceDate: July 19, 2026Context: Second consecutive speaking appearance at WAIC; first public speech after leaving Keen Technologies and founding Oak Lab Good afternoon everyone, thank you for inviting me back to the World Artificial Intelligence Conference for the second year running. It is a genuine honor to return to Shanghai and share my thoughts on the future of artificial intelligence with all of you. Just one week ago, I announced a major shift in my career path. After three fulfilling years at Keen Technologies, the company founded by John Carmack, I departed alongside my long-term research collaborator Khurram Javed to launch our new AI venture, Oak Lab. Today marks my first public appearance since that announcement, so I will fully lay out the core vision guiding all of Oak Lab’s research and development work. Over the past decade, large language models have delivered striking surface-level breakthroughs, yet we must face an unvarnished truth: mainstream foundation models today suffer from fundamental flaws in their underlying logic. All prevailing LLMs adopt an offline one-time training paradigm built on static human-curated datasets; once training concludes, model parameters are locked permanently. These systems can generate fluent, coherent text, yet they lack the capacity for genuine, continuous self-learning through real-time interaction with the physical world. Human intelligence operates on the exact opposite principle. From infancy onward, humans learn nonstop through physical action, sensory perception, environmental feedback, and constant trial and error. We do not construct our cognition from a single batch of precompiled static data; every interaction with our surroundings yields new experience, and every new experience reshapes our internal framework for understanding the world. Artificial intelligence today remains trapped in the “era of static human datasets.” The next phase of AGI development will inevitably belong to the “era of self-generated real-time experience.” The divide between these two paradigms is far more than incremental optimization—it represents two entirely separate technical trajectories. This is the path Oak Lab has chosen, one distinct from the dominant large language model approach embraced by most of the industry. We consistently overestimate the true capabilities of current AI systems. These models excel at replicating textual patterns embedded in training data, but they possess no first-person subjective experience, no innate drive for autonomous world exploration, and no ability to spontaneously generate new subproblems to advance their own cognition. They merely memorize knowledge summarized by humans, rather than discovering new rules through independent exploration. Today I revisit my classic essay The Bitter Lesson. For decades, AI researchers have poured immense effort into hardcoding human-derived prior rules into computational systems. History has proven this approach unsustainable in the long run. True general intelligence must deduce physical rules, logical structures, and world models entirely through its own environmental interactions, unshackled by heavy artificially imposed biases. I hold deep respect for John Carmack, and I am grateful to every colleague I collaborated with at Keen. My three years there were immensely rewarding, and our teams shared perfectly aligned goals centered on building embodied intelligence powered by reinforcement learning. Nevertheless, I recognized that to fully implement the experience-first architecture we believe in, I required an independent organization dedicated exclusively to this singular technical mission. Within Oak Lab, there will be no competing R&D priorities; all human and compute resources converge on developing one category of agent: systems that learn persistently through runtime real-time interaction, rather than relying on static offline datasets. Oak Lab operates two complementary divisions. The commercial laboratory translates core reinforcement learning theories into commercially viable embodied agents, while my nonprofit Openmind Research Institute conducts open, shared foundational frontier research. All revenue generated by Oak Lab’s commercial operations will be fully reinvested into open academic research, creating a self-sustaining, circular ecosystem. Our long-term hardware and energy efficiency targets are unambiguous: to build trillion-parameter general agents supporting real-time learning and planning with a power draw comparable to the human brain, approximately 20 watts. Today’s mainstream large models consume thousands of watts of power, their core inefficiency stemming from endless data transfer between storage and compute units. Only a brain-inspired parallel architecture that enables local in-place data processing, minimizing data movement, can deliver the step-change improvement in energy efficiency we seek. Oak Lab’s proprietary core technical framework is named OaK, short for Options and Knowledge. It addresses the single largest gap within modern reinforcement learning and large language model ecosystems: the capacity for autonomous hierarchical abstraction. An Option refers to a reusable high-level behavioral template, a complete sequence of actions paired with clear trigger conditions and termination rules. Activities such as drinking a cup of coffee, opening a door, or navigating to a charging station all qualify as Options. Unlike fragmented low-level individual actions, Options constitute complete reusable skill sets that agents can flexibly deploy across entirely novel scenarios. The “Knowledge” component within the OaK system is a predictive world model autonomously constructed by the agent, which records the outcome generated following the execution of every Option. Critically, this entire body of knowledge originates solely from the agent’s own real-time interactive experience, not imported human text corpora. The OaK learning loop operates entirely online with a fixed batch size of one. Every single new interactive experience triggers an immediate parameter update, eliminating delayed batch offline retraining cycles. The agent explores its environment, generates novel experience, creates new behavioral Options, expands its internal cognitive model, and derives fresh unsolved subproblems—forming an infinitely self-improving closed loop. This architecture resolves the defining flaw of mainstream large language models: the complete separation between training and inference phases. Learning ceases entirely for existing models once pre-training finishes; for OaK agents, the learning process continues indefinitely as long as interaction with an environment persists. Ninety-nine percent of the industry’s commercial investment currently flows into generative models built on static datasets. Yet this technical path will soon hit unbreakable hard ceilings: bounded generalization capacity, perpetually soaring compute costs, fragile operational stability, and zero capacity for autonomous evolution outside manual fine-tuning. Three conditions must converge simultaneously for artificial general intelligence to reach its inflection point:First, an agent architecture capable of continuous learning via real-time interaction—the OaK roadmap Oak Lab is pioneering;Second, a hardware ecosystem matching human brain-level energy efficiency, enabling trillion-parameter agents to run on just 20 watts of power;Third, large-scale deployment within real embodied use cases that deliver quantifiable industrial value. No major mainstream AI player has fully satisfied all three requirements today, and this represents the technical and market window Oak Lab aims to occupy. Many people ask me whether embodied agents will replace large language models. My answer is no—language models will only exist as a minor submodule within the broader architecture. Language constitutes an extremely narrow subset of all human experience. True general intelligence must first master full-spectrum perception, motor control, spatial reasoning, and causal interaction; language serves merely as a secondary communication tool. For decades, we have trained AI to mimic human writing, human speech, and human recorded knowledge. The next era of artificial intelligence demands training AI to perceive and experience the world just as humans and all living organisms do. I remain deeply thankful to John Carmack and the entire Keen team for our shared research journey. Now, through Oak Lab, I will channel my full focus into building artificial intelligence rooted in real-time interactive experience. We do not aim to implement minor iterative tweaks to existing models; our goal is a complete reset of the fundamental learning paradigm. Thank you all. I look forward to collaborating with researchers and enterprises across China to advance this new experience-driven wave of artificial intelligence.