#xAI
#xAI
今日 1 条
Dongxi 东锡 NLP@dongxi_nlpAI 评分3131
Tomer Tunguz 博客(VC 分析)AI 评分5353 The Economist 用世界价值观调查测 25 个前沿 AI 模型的世界观
The Economist 用 1981 年以来覆盖 100 国的世界价值观问卷测试 25 个前沿 AI 模型,按世俗到传统、生存到自我表达两轴绘制分布。
Tomer Tunguz 博客(VC 分析)精选AI 评分6060 Tomasz Tunguz:AI 模型的用户留存介于社交网络与手游之间
Tomasz Tunguz 提出 AI 如同一个漏水筛子,前沿模型平均保持领先约 41 天,用户留存率在个位数高位到约 40% 之间,介于社交网络(约 80%)与手游(百分之几)之间。同等智能水平的价格每年约下降 10 倍,如 Grok 4.5 以每任务 $0.31 达到 Intelligence Index 54 分,买家每 41 天获得更多议价筹码,创业公司可随每月赢家切换模型。
推荐理由:作者用用户留存曲线和智能单价数据,把 AI 模型的用户流失放到软件、社交网络与手游之间定位。
Tomer Tunguz 博客(VC 分析)AI 评分5757 Tomer Tunguz:AI harness 正成为企业数据竞争的新战场
VC Tomer Tunguz 撰文分析 AI 时代企业数据外流风险,引用 Satya Nadella 的“反向信息悖论”与 Palantir CEO Alex Karp 关于前沿实验室“窃取业务权重与 alpha”的言论。
Every:最新文章(网页)AI 评分3232 为什么 evals 现在这么火
Every 的 Context Window 栏目解释了 evals 为何突然无处不在,并评测了 Grok 4.7,称其表现"参差不齐"、可能是"一次退步"。Grok 4.7 已于周一公开发布。
Every:最新文章(网页)AI 评分6060 Every 团队实测 Opus 5.5、GPT-6 Sol 与 Grok 4.7
Every 团队对本周发布的三款模型做了 Vibe Check。Opus 5.5 在产品和设计工作上持平或超过 Fable 5.1,每 token 成本低 60%,但会把要点埋进长文且不限时运行;GPT-6 Sol 成为 Dan Shipper 的日常主力,更快更便宜;Grok 4.7 只赢回一位用户,Katie 认为其写作缺乏节奏感。
Elon Musk@elonmuskAI 评分1818
Elon Musk@elonmuskAI 评分2222引用jimmah@jamesdoumaThis really brought home to me this new machine creativity. I don’t know the recipe, but it’s more than vibes I see. There’s substance in the lyrics and there’s nuance for the critics And a story being told in this AI homily. And a message being sold by the Ai homily.
Gary Marcus:The Road to AI We Can Trust(RSS)AI 评分5959 Gary Marcus 批评白宫“超级智能”协议缺乏实质约束
Gary Marcus 评析特朗普政府发布的白宫“超级智能”协议,称其承诺的四层控制与审计本质上是“不受监管、不给公众发声”的自我监管。他质疑协议中“独立”审计人的含义,并指出两周前业界谈论的 AI 放缓(Pacing)议题未体现在协议中,称 Dario、Sam 和 Elon 都退缩了。
Elon Musk@elonmuskAI 评分2727引用Margo Martin@MargoMartin47.@POTUS hosts a Super Intelligence meeting at the White House 🇺🇸
Ethan Mollick@emollickAI 评分2222
Yuchen Jin@Yuchenj_UWAI 评分22223 周前我试了 Grok Bot。 2 周前装了 Instint。 上周装了 Muse。 现在显然我还得试试 Dots。 个人 AI 助手之战,开始吧。
lauren@potetoAI 评分88
elsewhere:文章(RSS)AI 评分4141 个人主义才能救AI:从办公Agent到Personal agent的锐评
一篇锐评指出,办公Agent把toB业务向toC宣传在法理和情理上都站不住脚,因为对非程序员群体而言,提高生产力并不能换来早下班。作者转而推崇Personal agent,并实测了Today.ai、Grok bot和Muse:Today.ai能连接Gmail和Notion后主动找活干,但无法连接微信;Grok bot偏开放,需自建助手,作者用它搭建了欧洲旅行、视频素材整理、时尚搭配等助手。
lauren@potetoAI 评分3232和 @petergyang 聊了我们的 grok @bot 设置,非常开心!
引用Peter Yang@petergyang"Everything I touch with my keyboard and mouse, I try to delegate to my bots." Here's my new episode with @poteto and @pengzheng_, the eng and design leads for Grok @bot, where they showed me the 14 bots they use for work and life, including: → A design bot that turns one keyframe into a full user flow → An eng lead bot that manages a team of eng bots → How to trust your bots with more of your work Some quotes from both: "I like to call it the Michelin kitchen…when you say software factory, it has this connotation of mass manufactured slop." "Sometimes I actually don't even look at the PR until after it's landed and then I'm like, 'Oh, okay. Yeah, that looks good.'" "I think it ultimately comes back to trust. First, watch your bot work and correct it. Turn what worked into a skill. Once it nails the task in one shot, make it a routine." 📌 Watch now: https://youtu.be/xZ5TEaleUdg Thanks to our sponsors: @meetgranola: AI meeting notes that don’t suck https://granola.ai/peter @RiversidedotFM: All-in-one AI studio for podcasts and video https://creators.riverside.com/PeterYang
karminski-牙医@karminski3AI 评分2929
eric zakariasson@ericzakariassonAI 评分3030引用aditya@adxtyahqGROK 4.7 IS ACTUALLY COMPETING WITH GPT-6 ASTRA. I gave GPT-6 Astra, Grok 4.7, Kimi K3 and Fable 5.1 the same prompt to build a flight simulator Astra was still #1 overall, but Grok 4.7 was surprisingly close Kimi K3 and Fable 5.1 were basically a draw and both produced a much smoother result, while Grok 4.7 was right up there with Astra in terms of overall quality. overall: Astra > Grok ≈ Kimi ≈ Fable GPT-6 Astra finally has some serious competition.
eric zakariasson@ericzakariassonAI 评分1212
a16z:News(RSS)AI 评分5353 a16z 图表周报:AI 代码生成让应用数量激增,但用户增长停滞
a16z 图表周报引用近期论文和 SensorTower 数据指出,AI 代码生成工具让每月新应用数量在 iOS、Android 和 Chrome 上翻倍甚至翻两番,但下载量和评分基本停滞,达到 10+ 评分或 100+ 下载等规模的应用占比大幅下降。
Lee Robinson@leerobAI 评分5454引用Lee Robinson@leerobGrok @Bot has made a few simple yet powerful technical decisions that I believe make it easy and enjoyable to use. 1. The best UI is none at all. The product interface is dramatically simpler than alternatives without sacrificing functionality. How is this possible? It's one of the first products designed for current frontier model capabilities and has a UI restrained enough to remain easy to use as models improve exponentially. Everyone knows how to text. 2. A thin harness for the client, a thick harness for the server. You might have noticed the app feels very fluid to use, even for a beta product. This is primarily because of everything we didn't have to build. The app harness is essentially a single tool to send messages between the client and server. The complexity moves to the server, where you can still use the coding agent harness with specialized tools as needed. This helps make the UI fast and responsive on desktop and mobile. 3. An always-on computer. Most coding agents and assistants today start fresh with every question you ask. Some of these sessions are on your local machine and others happen in the cloud. We believe strongly that cloud is the future, which is why it's the only option. Further, rather than spinning up virtual machines for every conversation, your bots connect to their own computer. This means you can still run agents on the bot's persistent filesystem. It's closer to what programmers have been doing by using Tailscale from their phones to connect to a remote computer and run an agent TUI. You get those capabilities without the hassle. 4. Your bots can use the browser. Coding agents have shown that most work on a computer can be expressed and run as code. You can ask for a task in natural language and the agent will decide to write a script to complete it. This is amazing, but there's still many tasks which can't be completed without logging into a website and clicking around the browser. Models and harnesses are now good enough to reliably handle this. The combination of writing code and using browsers means you can automate almost any task on a computer. Further, you can ask Grok Bot to record you doing the task, and then turn it into something repeatable.