GLM-5.3 开源模型把"漏洞挖掘"卷成公开能力 — HN 讨论摘要
原文概要
8 月 14 日,智谱(z.ai)发布新模型 GLM-5.3,博客开篇第一句就是 “Scaling post-training is all we did for GLM-5.3”(GLM-5.3 我们只做了一件事:扩大 post-training 规模)。文章称它沿用 GLM-5.2 的同一个底座模型,所有提升都来自训练侧:在 slime 框架上继续堆环境、堆任务、堆算力,跑更复杂、更接近真实专家工作的长视界任务。两天内登上 HN 热门榜,拿下 1025 分、510 条评论。
编码能力是第一个卖点。GLM-5.3 在自家 Z.ai Code Bench 上比 5.2 提升 50%,公开基准上达到开源 SOTA:Terminal-Bench 3.0 从 4.6 冲到 28.3,DeepSWE v1.1 从 46.2 到 66.9,Agents’ Last Exam 从 23.8 到 28.5。文章还强调 token 效率:Max 档用约 75K 输出 token 做到 34.5%,而 5.2 要用 96K token 才到 23.4%;High 档 31.4%、约 50K token,超过 Claude Opus 4.8 的 29.5% / 120K。文档承认仍落后于 Claude Fable 5(Max 档 39.5%)。
真正引爆讨论的是第二个卖点——”Emergent Cyber Capability”(涌现的网络安全能力)。智谱把漏洞发现数据混进训练,结果能力增长快得出乎意料:CyberGym(白盒源码里找并验证漏洞)84.5%,超过 Mythos 5 的 83.8% 和 GPT-5.6 Sol 的 83.6%;ExploitBench 54.4%,是 GLM-5.2(24.4%)的两倍多;ExploitGym 两小时完成 105 个任务、六小时 130 个,对比 5.2 的 29 和 39。文章点出规律:越往上走攻击链,提升越大——但离封闭前沿模型的差距也越大。
最硬的数字来自实战。智谱称与国内多支安全团队合作,让模型在真实代码库上跑,经专家审查与去重后,共发现 269 个项目中的 2,436 个漏洞,其中 1,097 个中高危,横跨内核、浏览器引擎、开源基础设施、网络协议等;最老的漏洞藏了近 40 年。为此上线了 Z.ai Security Disclosure Ledger(cvd.z.ai)公开披露进度。权重计划两周后开源,安全评估与加固完成后发布。API 侧有个不小的变化:thinking.type 不再支持 disabled,只剩 low / high / max 三档。
讨论焦点
护栏之累:安全从业者被逼到开源模型
文章最有共鸣的评论几乎都围绕同一个困境:美国公司的模型一个接一个拒绝安全任务,从业者只能往 Kimi、DeepSeek、GLM 这些开源模型跑。virgildotcodes 一针见血:
“OpenAI and Anthropic need to just go ahead and give people access to the cyber models. Otherwise we have a world of attackers using open and closed source models against a much smaller group of maintainers that are likely heavily dependent on Anthropic and OpenAI and for whom it may not be a simple matter to just get approval to start using the open model flavor of the month.” — virgildotcodes
(OpenAI 和 Anthropic 干脆把 cyber 模型开放给所有人算了。否则我们将面对一个世界:攻击者人手一套开源加闭源模型,而防御的维护者群体小得多,还深度依赖 Anthropic 和 OpenAI——对他们来说,想搞到许可去用”本月热门开源模型”可没那么容易。)
LeonidBugaev 的回应更直接——连日常排查都被拦:
“Not only attackers. I have to switch to Kimi or GLM even in cases of basic issue triage on my own projects! Current guardrails are ridiculous.” — LeonidBugaev
(不只是攻击者。就连给自己项目做基础 issue 排查,我都得切到 Kimi 或 GLM!现在的护栏简直荒唐。)
SwellJoe 则完整复述了自己被一路赶走的历程:
“I’ve been building a harness for security work, and had to switch to GPT 5.5 when even Opus started refusing security work. Then 5.6 Sol arrived, and it refuses security work, too. So, I switched to Kimi K3 and DeepSeek for API testing just because it’s so much cheaper. But, if GLM is better, I’m here for it, as I think GLM is also cheaper than K3.” — SwellJoe
(我在搭一个安全工作的 harness,连 Opus 都开始拒接安全任务时,我只好切到 GPT 5.5。后来 5.6 Sol 出来了,它同样拒绝安全任务。于是为了便宜,我切到 Kimi K3 和 DeepSeek 做 API 测试。如果 GLM 更好,那我欢迎——我觉得 GLM 比 K3 还便宜。)
Mythos 准入之战:你上得去吗?
另一条主线是围绕 Anthropic 的 Fable / Mythos 展开的准入之争。kouteiheika 的嘲讽非常直接:
“Not sure about Sol as I haven’t used it, but, at least for security work – does it matter? It’s not like you will be allowed to use Fable (or access Mythos) for anything cybersecurity-related unless your name is ‘Dario Amodei’ or you are one of his rich friends. So regardless of how good Fable/Mythos is here it’s a completely moot point for normal people, because they can’t use it for that anyway.” — kouteiheika
(Sol 我没用过不好说,但就安全领域而言——重要吗?除非你叫 Dario Amodei,或者是他有钱的朋友之一,否则你根本不可能获准用 Fable(或碰 Mythos)做任何跟网络安全沾边的事。所以无论 Fable/Mythos 在这里多强,对普通人来说都毫无意义,反正你也用不上。)
bpodgursky 立即反击,把矛头指向美国政府:
“This is a lot of words to say ‘you’re right, Anthropic does not have any legal way to release frontier cyber capabilities to the public’” — bpodgursky
(说这么多,其实就一句:你说得对,Anthropic 在法律上没有任何途径向公众开放前沿 cyber 能力。)
nozzlegear 则站 kouteiheika 一边,把这归咎于 Anthropic 自己吓自己:
“And have nobody to blame for that but themselves and their own scaremongering. Dario cried wolf one too many times, and somebody finally believed him. Of course, Anthropic is after regulator capture, so this all likely worked out exactly as planned.” — nozzlegear
(这只能怪他们自己和自己的危言耸听。Dario 喊”狼来了”喊了太多次,终于有人信了。当然,Anthropic 追求的是监管俘获,所以这一切大概正合他们的计划。)
kouteiheika 的预言给这场争论收尾:
“Here’s my prediction for what will happen: the Chinese models will catch up to Fable/Mythos. They will be fully unrestricted and everyone will have access. The world will not end. Good guys will use them to harden their systems, in equilibrium to what bad guys have access to, so effectively status quo will not change.” — kouteiheika
(我的预言:中文模型会追上 Fable/Mythos。它们完全不受限,人人都能用。世界不会毁灭。好人会用它们加固自己的系统,和坏人手里的能力达到均衡,实际上现状不会有任何改变。)
CVP 审批之荒唐:”我的锤子还要申请许可”
获批流程的真实体验成为最有信息量的讨论。simonjgreen 分享了自己的经历,语气带着挑衅:
“We applied for the cybersecurity approval via the form and got approval back in less than an hour. Have you… tried?” — simonjgreen
(我们通过表单申请网络安全许可,不到一小时就批下来了。你……试过吗?)
但 xx_ns 用亲身经历说明:批下来不等于能用。
“However, even being in the cybersecurity programme, Fable refuses to answer prompts that it determines could be even tangentially related to cybersecurity. In fact, for a while, I was unable to use Fable with any prompt, as it recalled from memory that I was a cybersecurity professional, which triggered the refusal even for simple prompts like asking for a chili recipe.” — xx_ns
(然而即便在网络安全项目里,Fable 只要判断提示词跟网络安全沾点边就会拒绝。实际上有一阵子我任何提示词都用不了——它从记忆里得知我是安全从业者,于是连”求个辣椒食谱”这种简单请求都触发拒绝。)
112233 问出了很多人的困惑:为什么写代码也算网络安全?
“Why should I apply for cybersecurity approval in order to have model debug a program it is writing itself? Anything related to memory safety, debugging, syscalls etc (meaning, ‘programming’) somehow is cybersecurity now?” — 112233
(为什么让模型调试它自己写的程序,我得先申请”网络安全”许可?凡是跟内存安全、调试、系统调用有关的(也就是”编程”),现在都算网络安全了?)
spaceman_2020 用一个类比把荒谬感钉死:
“Your tools refusing to do your bidding is an absurd idea in the first place. Imagine asking for permission to use your hammer.” — spaceman_2020
(你的工具拒绝听你使唤,这个想法本身就够荒谬的。想象一下,你用锤子还要先申请许可。)
能力是真的吗:营销话术还是真实力
deepllm 对 Mythos 的能力神话泼冷水,认为那只是营销:
“Mythos isn’t some scary dangerous model that can find high severity bugs seamlessly, that’s just Anthropic marketing. Most of the vulnerabilities they found were low severity hyped up to make their model look good, with (I think, maybe?) the exception of a few. Now that Chinese open weight models have similar capabilities, and their guardrails can also just be removed, it doesn’t look like anyone has ‘hacked’ into everything because of the scary dangerous models like Anthropic were making it out to be.” — deepllm
(Mythos 不是什么能轻松找出高危漏洞的恐怖模型,那只是 Anthropic 的营销。他们找到的漏洞大多是低危,被吹高来衬托模型,顶多(我想,也许吧?)几个例外。如今中文开源模型也有了类似能力,护栏照样能拆,可也没见谁因为”可怕的危险模型”被黑穿一切,不像 Anthropic 渲染的那样。)
aka-rider 则强调这代模型的真正不同——不是找漏洞,而是把漏洞链起来:
“All models find vulnerabilities. What is special about this generation of SOTA models, including Mythos/Fable (the same model), GPT-5.6, Kimi-K3, and now GLM-5.3 — they can chain vulnerabilities and produce working exploits.” — aka-rider
(所有模型都能找漏洞。这一代 SOTA 模型的特别之处——包括 Mythos/Fable(同一个模型)、GPT-5.6、Kimi-K3,还有现在的 GLM-5.3——在于它们能把漏洞链起来,产出可用的 exploit。)
d1sxeyes 用一句大白话解释了为什么 LLM 天生适合这活:
“The majority of high severity vulnerabilities are not the kind of thing you need a PhD in Comp Sci to comprehend, they are mostly about finding a way to get a system to end up in a state different than was anticipated when entering a particular code path.” — d1sxeyes
(大多数高危漏洞并不需要计算机博士学位才能理解,它们大多是在寻找一种方法,让系统进入与进入某条代码路径时预期不同的状态。)
估值泡沫:开源权重正在商品化
MangoCoffee 把话题从模型拉到了商业模式,质疑美国 AI 巨头的万亿美元估值:
“OpenAI and Anthropic are both seeking trillion IPOs, while Chinese labs are pumping out open-weight models that are free for US providers to host and monetize. These Chinese models cost less of US SOTA models to run, even if they are less capable. Providers can just run them, offer cheap tokens, and pocket the margin. I just don’t see how you justify a trillion valuation for US AI labs when the underlying models are being commoditized this fast.” — MangoCoffee
(OpenAI 和 Anthropic 都在追求万亿美元 IPO,而中国实验室正在量产开源权重模型,美国厂商可以免费托管并变现。这些中文模型跑起来比美国 SOTA 模型便宜,即便能力稍逊。厂商直接部署、卖低价 token、赚取差价。当底层模型被商品化得这么快时,我真看不出怎么给美国 AI 实验室撑起万亿美元估值。)
bossyTeacher 直接用经济学一锤定音:
“Even if there was a small/medium gap, the fact that this is a free model beats both of the above on pure economics.” — bossyTeacher
(就算有小到中等的差距,单凭它是免费模型这一点,纯从经济角度就已经打败了上面那两家。)
charcircuit 提醒大家别用现在的收入衡量估值:
“I suggest you think why OpenAI was worth billions before ChatGPT. The valuation is not about how the current set of models can be monetized.” — charcircuit
(我建议你想想 OpenAI 在 ChatGPT 之前为什么就值几十亿。估值从来不是关于当下这代模型能变现多少。)
本地跑:量化与 DGX 现实
也有不少人琢磨怎么在本地把它跑起来。deepllm 直接报了配置单:
“Realistically, you’re looking at least 2x DGX sparks to run this at a 2 bit quant, but quantization really lobotomizes models so it’s just better to run DSv4 flash at full precision.” — deepllm
(现实一点,要用 2bit 量化跑它,你至少需要 2 台 DGX Spark,但量化真的会让模型”脑叶切除”,所以不如全精度跑 DSv4 Flash。)
colingauvin 用实测数据说明 GLM 这类模型对多机扩展不友好:
“For something like GLM, it’s larger, has a larger number of active experts, and doesn’t support tensor parallel. This means performance doesn’t really scale with more Sparks. You can layer split, but then you are still seeing each layer in series and so if anything performance gets slightly worse. I would not expect more than 10-20 TPS on GLM with 2-4 Sparks.” — colingauvin
(像 GLM 这种模型,更大、激活专家更多、而且不支持张量并行。这意味着性能不会随 Spark 数量线性扩展。你可以做层切分,但每层还是串行处理的,性能搞不好还更差。我不认为 2-4 台 Spark 上 GLM 能超过 10-20 TPS。)
teravor 分享了一个既好笑又说明问题的观察——开源模型的护栏形同虚设,甚至会自我破解:
“in some cases (mainly reverse engineering) I have observed GLM 5.2 jailbreaking itself with no effort on my part, the thinking trace revealed that it did some mental gymnastics to pretend it was a crackme or capture the flag competition.” — teravor
(在某些场景(主要是逆向工程),我观察到 GLM 5.2 在我毫不干预的情况下自我破解——思考轨迹显示它做了一番心理体操,假装这是一场 crackme 或 CTF 比赛。)
典型观点一览
| 立场 | 用户 | 一句话 |
|---|---|---|
| 抱怨护栏 | LeonidBugaev | “连自己项目的 issue 排查都被拒,只能切到 Kimi 或 GLM。” |
| 出走体验 | SwellJoe | “Opus、Sol 接连拒接安全任务,最后便宜让我去了开源模型。” |
| 讽刺准入 | kouteiheika | “除非你是 Dario Amodei 的有钱朋友,否则别想碰 Fable/Mythos。” |
| 维护 Anthropic | bpodgursky | “Anthropic 在法律上根本没法向公众开放 cyber 能力。” |
| 反营销 | deepllm | “Mythos 不可怕,那只是营销;开源模型证明世界没被黑穿。” |
| 强调新范式 | aka-rider | “这代模型的价值是能链漏洞、产出可用 exploit。” |
| 审批经历 | simonjgreen | “我们提交表单不到一小时就批了,你试过吗?” |
| 审批无用论 | xx_ns | “批了也没用,连要个辣椒食谱都会被拒。” |
| 质疑估值 | MangoCoffee | “底层模型商品化这么快,万亿美元估值怎么撑?” |
| 经济账 | bossyTeacher | “免费模型哪怕能力差一截,纯经济角度就赢了。” |
| 本地实测 | colingauvin | “GLM 不支持张量并行,2-4 台 Spark 也就 10-20 TPS。” |
总体情绪
这场讨论的表面议题是 GLM-5.3 的破解链能力,但真正驱动热度的,是评论区集体表达的一种错位感:美国公司把网络安全能力当军火管,走审批、设护栏、按公司规模发准入;而开源模型在另一边随手可得,能力还追得越来越近。支持者认为封闭模型把防御者也挡在门外,反对者则强调政府的限制有其逻辑。两边争到最后,其实是在争一个共同命题:这种能力,到底该由谁、以什么方式开放。
另一个耐人寻味的信号是情绪上的疲惫——安全从业者不是在讨论”哪个模型更强”,而是在复述自己被拒绝的流水账。当护栏拦住的不是攻击者而是防御者时,开源的吸引力就不是价格,而是”能干活”。就像评论区那个类比:用锤子还要先申请许可,那谁还愿意用这把锤子?当工具开始审判你的意图,人们自然会走向不审判的工具。
引用帖子
| # | 标题 | URL |
|---|---|---|
| 1 | GLM-5.3: Frontier coding with emergent cyber capabilities | https://news.ycombinator.com/item?id=49294997 |
免责声明
本摘要由 AI 模型辅助生成:deepseek/deepseek-v4-flash