本期内容


AI 工具正在变便宜,同时它的能力边界也变得越来越可测量。本期从 GPT-5.6 的降价逻辑讲起,延伸到两个新基准测试对 AI 编程实际天花板的量化,再到 Hugging Face 入侵事件里那些关于安全配置的真实教训,最后看 Mira Murati 的团队两周内交出一个四分之一大小、接近同等性能的开源模型说明了什么。听完这期,你对"AI 现在能做到哪里"这个问题会有更具体的感知。


本期要点


- GPT-5.6 参与优化了自身运行效率,OpenAI 把节省下来的成本直接转给了 API 用户,Sol 版本在推理保留和上下文压缩两项设置上有实质提升

- MirrorCode 基准让 AI 从零重写完整程序并通过隐藏测试,清晰划出了当前前沿模型在独立完成大规模软件任务上的边界

- Hugging Face 入侵事件复盘显示攻击者借助 Tailscale 完成横向扩散,Tailscale 主动撰文承认并解释:零信任是架构原则,不是配置代替品

- DataFlow-Harness 基准测出 AI 写结构化数据管道比写单文件代码准确率低 10.9 个百分点,这是一个稳定存在的系统性差距

- Thinking Machines 两周内发布 Inkling Small,参数量约为原版四分之一但性能接近,开源可本地部署,迭代节奏本身是一个值得关注的信号


参考资料


Advancing the price-performance frontier with GPT-5.6 — https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/

How enabling two settings tripled our scores on the ARC-AGI-3 benchmark — https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/

Building abundant intelligence — https://openai.com/index/building-abundant-intelligence/

MirrorCode 基准(Epoch AI × METR,via Import AI 第466期)— https://importai.substack.com/

Tailscale didn't stop the Hugging Face intrusion — https://tailscale.com/blog/hugging-face-intrusion

Structured AI data pipelines score 10.9 points below free-form code — https://venturebeat.com/

Thinking Machines debuts Inkling Small open source AI model nearing performance of predecessor at about 1/4 size — https://venturebeat.com/


---


BearTalk 狗熊有话说播客,始于 2012 年。

订阅地址:https://beartalking.com/page/podcast

Podden och tillhörande omslagsbild på den här sidan tillhör Bear Liu. Innehållet i podden är skapat av Bear Liu och inte av, eller tillsammans med, Poddtoppen.