Kimi K3 — 开源 3T 级模型挑战前沿 — HN 讨论摘要
中国 AI 实验室 Moonshot AI 于 7 月 16 日发布 Kimi K3,一个 2.8 万亿参数(自称”3T 级”)的开源权重模型。根据其自报基准,K3 在多数测试中超过 Claude Opus 4.8 和 GPT-5.5,仅次于 Claude Fable 5 和 GPT-5.6 Sol。权重承诺在 7 月 27 日前开放。
定价方面,K3 输入 $3/百万 token、输出 $15/百万 token,与 Anthropic Sonnet 系列持平,远高于此前 K2.6 的 $0.95/$4。Artificial Analysis 报告显示,K3 的单任务成本约 $0.94,接近 Sol 的 $1.04,约为 Opus 4.8 的 $1.80 的一半。模型仅支持单一推理 effort 级别(”max”),并在 Arena.ai 前端代码排行榜上超过 Fable 5 位列第一。
此帖在 HN 引发 1159 条讨论,以 1992 分登上当日榜首。同时,Simon Willison 发布了一篇通过”鹈鹕骑自行车”测试评估 K3 的博客(262 分),同样引发大量讨论。
定价与性价比:对标 Sonnet 的豪赌
K3 最令人意外的不是参数规模,而是定价策略。来自 Moonshot 中国背景的模型通常以低价取胜,但 K3 直接对标 Anthropic。
“This is 1:1 pricing of Anthropic’s Sonnet series (except Sonnet 5 which is currently on discount), and very close to 5.6 Terra pricing.” — Tiberium (这是 Anthropic Sonnet 系列的 1:1 定价(除了目前在打折的 Sonnet 5),非常接近 5.6 Terra 的价格。)
但评论者随即指出,原始 token 价格不能代表实际成本:
“reasoning efficiency matters directly for how expensive a model actually is in real use. GPT’s models are extremely reasoning efficient, and some Claude models like Fable at lower effort are as well. So if Sol spends 10K reasoning tokens to do something (at $30/1M) vs Kimi K3 that spends 50K reasoning tokens, Sol would win on cost effectiveness.” — Tiberium (推理效率直接影响模型的实际使用成本。GPT 模型推理效率极高,低 effort 的 Fable 也是如此。如果 Sol 花 1 万推理 token 完成某事($30/百万),而 K3 花 5 万推理 token,最终 Sol 的成本反而更低。)
几位 GLM 用户也表达了类似观点——GLM 5.2 的 token 利用率同样不理想:
“GLM is actually quite expensive in actual practice because it’s not very token efficient.” — cmrdporcupine (GLM 在实际使用中相当昂贵,因为其 token 利用率不高。)
当前市场下,订阅制已经让 API 定价的竞争力变得复杂:
“Right now unless you’re paying by the token, there’s no cost based reason to use the open weight models for daily coding work because the monthly coding plans from Anthropic and OpenAI are a better deal.” — cmrdporcupine (目前除非按 token 付费,否则日常编码工作没有成本层面的理由去用开源权重模型,因为 Anthropic 和 OpenAI 的月度订阅方案更划算。)
DeepSeek V4 则被多次提及为当前最超值的选择:
“I know GLM is relatively expensive and so is Kimi, in comparison to those DeepSeek V4 pro and flash are a godsend and are absolutely good value.” — computerex (我知道 GLM 相对贵,Kimi 也是,相比之下 DeepSeek V4 Pro 和 Flash 简直是天赐之物,绝对超值。)
但也有用户尝鲜后给出正面评价:
“I’ve been using it for a few hours now… K3 is in the same ballpark as Fable or Sol.” — InsideOutSanta (我用了几个小时……K3 和 Fable 或 Sol 在同一水平。)
推理透明度:K3 开放,Fable 隐藏
K3 的一个差异化优势在于完全暴露推理过程,这与 Anthropic 和 OpenAI 的最新策略形成对比。
“I’ve been avidly using Fable since it was re-released and while it has been excellent at building the apps I want, the reasoning has been completely opaque. Kimi, however, has exposed the whole reasoning trace, or enough of it to matter.” — ImageXav (Fable 重新发布以来我一直在用,它在构建应用方面表现出色,但推理过程完全不透明。而 Kimi 展现了完整的推理轨迹。)
一位用户分析,Anthropic 隐藏推理链并非为了恶化用户体验,而是防御性策略:
“It’s a defensive tactic to reduce the effectiveness of distillation… It’s not because they want to wrest control from users. It’s because they don’t want Chinese companies to do exactly what Moonshot (Kimi creators) and others have done.” — qeternity (这是为了降低蒸馏有效性的防御性策略……不是因为想剥夺用户控制权,而是不想让中国公司做 Moonshot 和其他公司已经做过的事。)
对此,另一位用户尖锐回应:
“Anthropic’s position being that it is entitled to train models on the creative works of anyone at any time, but its own slop generators’ outputs are sacred jewels that must be protected from being learned from.” — anon373839 (Anthropic 的立场是:它有权随时用任何人的创意作品训练模型,但它自己的垃圾输出却是必须保护、不得学习的圣物。)
关于蒸馏的实际影响,有评论者援引业内观点指出:
“According to Pat Toulme, the thinking traces and outputs that the Chinese researchers distilled are useful for getting initial trajectories, to prevent a cold start during RL. Once you get those initial correct trajectories… the distillation is already done.” — desterothx (据 Pat Toulme,中国研究者蒸馏的思考轨迹和输出对于获取初始路径、避免 RL 冷启动有用。一旦获得正确轨迹……蒸馏就已经完成了。)
“第二(仅次于 Fable 5 和 Sol)”——文字游戏?
Moonshot 的措辞引发幽默讨论:
“Among the models tested, its overall intelligence ranks second only to Claude Fable 5 and GPT-5.6 Sol.” — tw1984 (在测试模型中,其整体智能仅次于 Claude Fable 5 和 GPT-5.6 Sol。)
“So… it ranks THIRD?” — nkmnz (所以……是第三?)
“USSR is proud to announce that they won 2nd place in an Olympic contest. The filthy USA regime? Next to last! (There were only two countries competing in said event)” — polski-g (苏联自豪地宣布在奥运会上获得第二名。肮脏的美国政权?倒数第二!(只有两个国家参赛。))
另有用户指出:
“The literal interpretation of that sentence is ‘when it is second or third, it is only behind Fable 5 or 5.6 Sol’. And indeed they give benchmarks where it is ahead of one but not both models.” — sudosysgen (这句话的字面意思是”排在第二或第三时,仅落后于 Fable 5 或 Sol”。实际基准中确实有超越其中一项但未超越两项的情况。)
鹈鹕测试:一个”搞笑基准”的严肃侧面
Simon Willison 用他的经典 prompt “Generate an SVG of a pelican riding a bicycle” 测试了 K3,结果花了 25 美分——13,241 个推理 token 和 3,417 个输出 token。这引出了话题:这个”鹈鹕基准”到底有没有意义?
“This entire ‘benchmark’ is a performative joke for attention that only works on HN.” — rvz (这个所谓的”基准”完全是为了博眼球的表演性玩笑,只在 HN 管用。)
Simon Willison 亲自下场回应:
“I take exception to that! It’s a performative joke for attention that works far more widely than just Hacker News.” — simonw (我反对!这是一个表演性玩笑,但它的影响力远不止 HN。)
关于鹈鹕是否已进入训练集、是否已被基准污染,讨论更为深入:
“I’m still not convinced that labs are training for the benchmark—if they were, I’d expect much better results.” — simonw (from blog) (我仍然不相信各实验室在针对这个基准训练——如果是,我期待更好的结果。)
但持怀疑态度的用户给出了具体证据:
“I have shared examples of certain models by certain labs doing far better on the pelican cycling vs other, similar prompts… Qwen3.6-35B-A3B… was noticeably better [at pelican]… a sloth riding a skateboard? It is barely recognisable as anything.” — Topfi (我分享过某些模型在鹈鹕骑自行车上远超其他类似提示的例子……Qwen3.6-35B-A3B 的鹈鹕明显更好,但树懒滑滑板?几乎认不出来。)
Simon 承认方向偏见的观察:
“There was a glorious moment when I thought that the Chinese models were more likely to produce right-to-left cycling pelicans, but sadly that trend didn’t seem to hold up.” — simonw (曾有一个辉煌时刻,我以为中国模型更可能产出从右向左骑的鹈鹕,但可惜趋势没有持续。)
关于 25 美分是否算贵,也有趣味交锋:
“Engineers get unbelievably silly about evaluating costs of things. ‘The tokens are so expensive!’ Oh my sweet child, how much would even the least capable human effort cost?” — BugsJustFindMe (工程师在评估成本时简直不可思议地幼稚。”token 太贵了!”我亲爱的孩子,哪怕最差的人类劳动力要花多少钱?)
“Would anyone pay a human to create an SVG of a pelican riding a bike?” — bakugo (会有人花钱请人画一个骑自行车的鹈鹕 SVG 吗?)
开源前景:权重真的会来吗?
Moonshot 最初表示权重将”在近日内”发布,但随后从博客删除了相关段落,引发猜测:
“They’ve removed the paragraph about releasing model weights.” — markasoftware (他们删除了关于发布模型权重的段落。)
“Does that mean this one won’t be open source?” — xur17 (这意味着这个模型不会开源了吗?)
实际上权重发布承诺仍然存在,只是改成了”7 月 27 日前”:
“Still there for me: ‘The full model weights will be released by July 27, 2026’” — nmfisher (我这边还在:”完整模型权重将于 2026 年 7 月 27 日前发布。”)
K3 的 2.8T 参数规模也让本地部署前景存疑:
“API prices are amazing, but hosting this on-premise will be real challenge.” — fmind-dev (API 价格很好,但本地部署将是个真正的挑战。)
“also its pretty big model inference costs are high even with margins running a 2.8T model costs a lot. if they release oss may be it goes down to $10-12 per million tokens.” — darkbatman (2.8T 模型的推理成本很高。如果开源,每百万 token 可能降到 $10-12。)
典型观点一览
| 立场 | 用户 | 一句话 |
|---|---|---|
| K3 定价合理 | easygenes | 有能力的模型按能力定价,订阅制依然健康,不存在补贴消失 |
| 定价过高 | cmrdporcupine | 开源权重模型的订阅不如 Anthropic/OpenAI 划算 |
| 推理效率是关键 | Tiberium | token 价格是假象,实际花多少推理 token 才是真相 |
| 开放推理链是优势 | ImageXav | 看到完整推理过程让调试和信任都变得可能 |
| 对蒸馏的担忧被夸大 | desterothx | 初始蒸馏完成后不再需要持续访问推理链 |
| 鹈鹕基准已被污染 | Topfi | 不同模型在鹈鹕和其他动物上的差距说明了一切 |
| 鹈鹕基准仍有价值 | simonw | 这是了解模型特性的”hello world”,不是为了排名 |
总体情绪
HN 社区对 Kimi K3 的评价喜忧参半。一方面,作为首个 3T 级开源权重模型,其技术成就不容忽视,部分用户实测表明它确实与 Fable 和 Sol 同属第一梯队。另一方面,其定价策略打破了”中国模型=便宜”的预期,引发关于推理效率和实际成本的深入讨论。围绕”仅次于 Fable 5 和 Sol”的措辞游戏、权重是否真会开放的疑虑,以及鹈鹕基准的学术性与娱乐性之争,让讨论既有干货又有幽默感。
社区整体态度可概括为:欢迎竞争,但需要看到更多独立第三方的客观评测,而非自报基准。模型的真正价值——尤其是在 agent 工具调用和长对话可靠性上——还需要时间验证。K3 的开源发布将是关键节点:届时社区可以获得无 API 限制的真实评估。
引用帖子
| # | 标题 | URL |
|---|---|---|
| 1 | Kimi K3: Open Frontier Intelligence | https://news.ycombinator.com/item?id=48935342 |
| 2 | Kimi K3, and what we can still learn from the pelican benchmark | https://news.ycombinator.com/item?id=48947717 |
本摘要由 AI 模型辅助生成:deepseek/deepseek-v4-flash