本期内容


AI 能力的跃升正在多个维度同时发生:模型按用途分档、底层内核自动优化、自由职业任务完成率量化可测。与此同时,两个更深层的变化也在浮现:聊天对话这种交互方式本身可能正在走向幕后,而 AI 写代码带来的可维护性危机正在逼问人类工程师的核心价值。本期五件事,帮你在噪音里找到真正值得关注的信号。


本期要点


- GPT-5.6 发布三档模型,Sol、Terra、Luna 各司其职,OpenAI 开始引导用户建立任务分级思维

- Claude Fable 提交了 KernelBench-Mega 史上第一个 megakernel,对特定任务实现接近 19 倍加速,AI 加速 AI 的循环已可在 benchmark 上验证

- Center for AI Safety 发布远程劳动指数 RLI,用真实 Upwork 项目衡量 AI 代理的完成质量,给"AI 影响就业"这个讨论提供了可量化的尺子

- Ethan Mollick 认为聊天机器人正走向命令行的命运,主动完成任务的代理式交互将取代单轮对话成为主界面

- Hacker News 热议 AI 写代码的可维护性问题:AI 擅长生成局部可用的代码,但架构决策和整体意图的责任仍然在人


参考资料


GPT-5.6 is now the preferred model in Microsoft 365 Copilot — https://openai.com/index/gpt-5-6-preferred-model-microsoft-365-copilot/

ChatGPT is now a partner for your most ambitious work — https://openai.com/index/chatgpt-for-your-most-ambitious-work/

Import AI #464: Fables writes GPU kernels; AI automation; and analog computation — https://importai.substack.com/

A Significant Increase in Digital Labor Automation (RLI) — https://safe.ai/blog/significant-increase-in-digital-labor-automation

The Twilight of the Chatbots — Ethan Mollick, One Useful Thing — https://www.oneusefulthing.org/

Write code like a human will maintain it — https://unstack.io/


---


BearTalk 狗熊有话说播客,始于 2012 年。

订阅地址:https://beartalking.com/page/podcast

Podden och tillhörande omslagsbild på den här sidan tillhör Bear Liu. Innehållet i podden är skapat av Bear Liu och inte av, eller tillsammans med, Poddtoppen.