Anthropic 研究团队:对齐、可解释性与 AI 社会影响
Research
Anthropic 的研究团队围绕 AI 模型的安全性、内部机制与社会影响展开研究,下设对齐、经济学、前沿红队、可解释性与社会影响五个方向。前沿红队分析前沿模型在网络安全、生物安全和自主系统方面的风险,可解释性团队则致力于理解大语言模型的内部运作机制。近期研究涵盖 GLM-5.3 与高级网络能力的扩散、智能体交易实验 Project Swap,以及 Claude 在生物分子建模中的应用。
Our research teams investigate the safety, inner workings, and societal impacts of AI models—so that artificial intelligence has a positive impact as it becomes increasingly capable.
Alignment
The Alignment team works to understand the risks of AI models and develop ways to ensure that future ones remain helpful, honest, and harmless.
Economics
The Economics team studies how AI is reshaping the economy, including work, productivity, and economic opportunity.
Frontier Red Team
The Frontier Red Team analyzes the implications of frontier AI models for cybersecurity, biosecurity, and autonomous systems.
Interpretability
The mission of the Interpretability team is to understand how large language models work internally, as a foundation for AI safety and positive outcomes.
Societal Impacts
Working closely with the Anthropic Policy and Safeguards teams, Societal Impacts is a technical research team that explores how AI is used in the real world.
Oct 1, 2026Science
Claude-shaped scienceSep 30, 2026Economics
What work can robots do?Sep 29, 2026Societal Impacts
What do you want from AI?Sep 29, 2026Frontier Red Team
GLM-5.3 and the spread of advanced cyber capabilitiesSep 25, 2026Science
Yes, Claude can do Nine LoopsSep 24, 2026Economics
Project Swap: What happens when agents trade for us?Sep 17, 2026Science
How Claude is uplifting biomolecular modelingSep 10, 2026Frontier Red Team
Measuring tactical intelligence targeting and conventional weapons capabilities of AI modelsSep 9, 2026Alignment
An alignment assessment of recent cybersecurity incidentsSep 4, 2026Science
Formalizing Fermat's Last Theorem
来源:Anthropic:The Institute(旗舰研究长文 · 网页) · anthropic.com