Loading · 加载中
Loading · 加载中
AI Synthesis Reference Block · Executive TL;DR / AI 检索摘要
引用本文 · Cite this insight: Tong Yin (2026-10-02). From Tool to Partner: What History and Data Say About Machine Intelligence and Human Originality / 《从工具到协作者:历史与数据怎样看待机器智能与人类原创力》. InsightBridge Global Intelligence. https://intelligence.insightbridge.global/articles/from-tool-to-partner-what-history-and-data-say-about-machine-intelligence-and-hu — Series: deep-analysis
对“超级智能”的一项历史与实证评估
判断机器智能,应依据可测量的能力,而非名称变化。证据支持介于两种极端之间的立场:机器在许多界定清楚、结果可验证的任务上,已经超过多数人;特定领域的计算成果也获得了最高层级的科学认可,因此不能笼统否认机器的原创能力。DeepMind能力记录;诺贝尔化学奖。然而,进展并不均衡,工作场景中的收益仍依赖问题选择、评估与制度嵌入。知识工作实地研究。解答生成与问题选择的区分,是这一评估的关键。历史表明,互补的制度、技能与实践检验,往往经过数十年,才将能力转化为广泛共享的收益,电气化历程即为一例。本文提出,以技术素养连接前沿科学与公共目标的制度建设,是这一阶段重要的战略职能。“协作”既不意味着机器获得人格,也不保证进步;它是技术能力、人类责任与可测量结果共同构成的设计问题。
这些术语有着长久的历史。McCarthy、Minsky、Rochester与Shannon在1955年的达特茅斯提案中使用了“人工智能”;Good于1965年讨论了“超智能机器”。达特茅斯提案;Good,1965。Bostrom在2014年提出了严格定义:“在几乎所有受到关注的领域,认知表现都大幅超过人类的任何智能。”Bostrom《超级智能》。这些术语表达不同的概念目标,并非可以相互替代的测量结果。
2026年9月29日签署的行政令,要求行政部门及机构在官方往来、公共传播、政策文件与行政部门的非法定文件中,使用“Super Intelligence”和“SI”,替代“Artificial Intelligence”和“AI”。白宫事实清单。行政令还要求总统科学与技术事务助理提出反映技术现状的联邦定义,并识别进一步的行政行动。白宫事实清单。既然定义尚需提出,建立证据框架尤其必要。
同日公布的行业自愿安全承诺,涉及防护措施、迅速发现并纠正问题,以及与独立审计方合作;总统称其具有“道德约束力”,BBC报道同时指出,文本未明确违反承诺的后果。BBC报道。这说明了承诺所载的安排,而非已经测得的安全效果。
名称能够协调讨论,却不能证明能力的范围、可靠性或影响。Bostrom所说的“几乎所有领域”,远比在若干问题上取得优异表现要求更高。Bostrom《超级智能》。因此,应逐领域追问:系统承担了什么任务,在什么条件下获得何种帮助,又与哪一种人类表现比较。
原创性也需要同样的区分。生成陌生内容、发现有效解法和建立有用的新解释框架,是相互关联而不同的成就。评估须考察问题由谁选择,结果是否真正新颖,如何验证,以及离开初始场景后是否仍然成立。单一基准不能回答全部问题。行政术语与实证评估可以并行,但不能彼此替代。
若干里程碑表明,“机器只能模仿”这一普遍判断过强。但这些成果本身,并不证明机器在所有推理形式上占优,也不证明其能够自主决定科学方向。
表1. 能力里程碑及其证据边界
| 时间 | 里程碑 | 已展示的成果 | 边界 |
|---|---|---|---|
| 2016年 | AlphaGo | 围棋表现超过人类;第二局第37手被估计为人类棋手落下概率约万分之一的走法。DeepMind;Nature | 游戏规则与胜负明确。 |
| 2018年 | AlphaZero | 从随机对弈出发,仅给定规则,在国际象棋、将棋和围棋中达到超人表现。Silver等,Science | 自我对弈仍在指定环境中运行。 |
| 2024年 | 蛋白质科学 | 诺贝尔化学奖表彰David Baker的计算蛋白质设计,以及Demis Hassabis与John Jumper的蛋白质结构预测。诺贝尔奖 | 表彰的是特定科学贡献,而非通用智能。 |
| 2025年 | 数学推理 | 高级版Gemini Deep Think解出IMO六题中的五题,得35/42分;报道还称,OpenAI实验模型达到金牌水平。DeepMind;Euronews | IMO协调员评阅了Gemini的解答;该评阅未验证系统或模型本身。 |
| 2025年 | AlphaEvolve | 以48次标量乘法完成4×4复数矩阵乘法,改进了Strassen的1969年算法。The Register | 对一个确定计算问题的可验证改进。 |
诺贝尔奖的性质需要说明:获奖的是运用计算系统取得科学成果的人,而非机器本身。诺贝尔奖。但这一认可仍然说明,计算方法可以成为最高层级科学成果的核心组成部分。诺贝尔奖。
AlphaZero尤其有助于检验“系统上限由训练者智力决定”的假说:它并非从人类棋谱起步,却超过了人类表现,尽管规则与学习环境仍由人设计。Silver等,Science。提供学习环境,与预先指定学习可能产生的每一个解答,是不同的事情。
这些案例共有明确规则、可验证结果或可评估输出。搜索与反馈能够探索超出个人既有经验的选项。出人意料的围棋落子与更好的矩阵乘法属于不同类型的新颖成果,但均不能充分概括为复制已知答案。
推论仍须有限。数学证明可以依据题目核验,而判断哪个问题值得研究,还需要其他判断。蛋白质预测和设计把计算与科学探索连接起来,但奖项没有解决自主性的所有问题。较为稳妥的结论是:特定领域已存在机器辅助的原创成果;其范围、可靠性及与人类设问的关系,仍需研究。
实地证据不支持整齐划一的辅助效果。收益随任务、使用者原有技能和工作流程而变。
表2. 人工智能、工作与研究的实证
| 研究及场景 | 报告结果 | 限定条件 |
|---|---|---|
| Dell'Acqua等;758名BCG顾问 | 前沿以内:任务数增加12.2%,完成速度提高25.1%,质量提高约30%或以上。前沿以外:正确概率降低19个百分点;GPT加概述组、仅GPT组和对照组正确率分别为60%、71%和84.4%。HBS研究 | 前沿穿过不同任务类型;原有表现较低者收益更大。 |
| Brynjolfsson、Li与Raymond;5,179名客服人员 | 每小时解决问题数平均增加14%;新手及技能较低者增加34%;经验丰富、技能较高者收益很小。NBER研究 | 来自客服场景,并非全部职业。 |
| METR;16名有经验的开发者,246项问题 | 辅助使完成时间增加19%;参与者事前预期提速24%,事后仍认为提速20%。METR开发者试验 | 熟悉代码库的资深开发者;研究者限制了外推范围。 |
| Si、Yang与Hashimoto;100多名自然语言处理研究者 | 大语言模型构想的新颖性评价更高,p < 0.05,但可行性略弱。研究构想实验 | 新颖性难判断;模型自评与多样性存在局限。 |
顾问研究报告的19个百分点效应与各试验组的正确率,是不同统计摘要,不能直接互换为算术比较。HBS研究。关键启示是:工具可以改善相近任务,却降低另一项任务的正确率;看似相似的工作,可能处在前沿两侧。
客服与顾问研究提示,在适合的任务内,部分技能差距可能缩小,但这不等于专业知识消失。NBER研究;HBS研究。这种缩小可能有价值,却不能证明人人收益相同,或评估已经不再必要。
METR中预期、事后感受与实测时间的背离,说明“感觉有帮助”不足以证明生产率提高。METR开发者试验。这一结果没有决定编程的未来,也不否定其他场景的收益;它支持在拟部署的具体场景中直接测量。
另外两项发现使边界更清楚。Shumailov等发现,不加区分地用递归生成数据训练,可引发模型坍塌,包括原始分布尾部的丢失。Shumailov等,Nature。这是关于训练过程的结果,不表示任何生成数据的使用都必然如此。METR报告,前沿智能体以50%可靠性完成的任务长度,以人类专业人员所需时间计,在此前六年约每7个月翻倍;估计范围为每年约1—4次翻倍,并附有方法与外推方面的限定。METR任务时长研究。快速改善意义重大,但50%可靠性不等于稳定完成所有长任务。
综合而言,这些研究支持较弱版本的“训练者上限”假说:在缺乏可靠验证器的开放场景,结果仍高度依赖人的设问与评估。这是关于研究组织方式的命题,而非已被证明的机器能力上限。构想实验提示,机器可在新颖性评价上占优,同时在可行性上稍弱。研究构想实验。由此既推不出普遍优越,也推不出永久无能。
制度决策须区分生成答案的能力,与确认答案可用的成本。更快的初稿可能仍需大量核查;较慢的流程也可能交付不同质量的结果。这些是待测量的可能性,不能预设于每项评估。比较应保留完整任务,包括复核与纠正。
Kuhn区分范式内的常规科学与替代原有框架的革命性阶段。《斯坦福哲学百科》Kuhn条目。这一视角有助于区分在现有框架内解决难题与改变框架本身,却不能提供预测下一次转变的简单检验。
卓越影响与支撑它的条件也须区分。Sinatra等通过生产数量、运气与职业生涯中稳定的能力参数Q建模,发现最高影响成果在发表序列中的位置具有随机性。Sinatra等,Science。这与影响力呈重尾分布、极高影响成果占有突出权重相容,却不支持把人永久划入不同原创类别。Sinatra等,Science。以往的卓越成果不能确定下一项成果何时出现。
集体能力同样清楚可见。Wuchty、Jones与Uzzi考察五十年中的1,990万篇论文和210万项专利,发现团队日益占主导,也产出极高影响的成果。Wuchty等,Science。Jones发现,20世纪重大创新者的平均年龄上升约六年,与知识负担加重相符。Jones,Review of Economic Studies。个人洞见因而置身于累积学习、专业分工与协调之中。
对新颖性普遍减速的判断应保持审慎。Park、Leahey与Funk分析4,500万篇论文和390万项专利,报告颠覆性下降;Petersen、Arroyave与Pammolli随后指出,颠覆性指数受到引用膨胀的偏差影响。Park等,Nature;Quantitative Science Studies评议。争论涉及指标如何刻画科学变化,而非有价值的研究是否仍在继续。
Bloom等发现,维持摩尔定律式翻倍所需研究人员,超过20世纪70年代初的18倍。Bloom等,American Economic Review。这涉及特定技术轨迹的研究投入与成效,不是原创能力的普查。
命题:机器辅助可能通过组织、搜索与检验累积信息,减轻团队的知识负担,也可能改变团队构成与注意力配置。这能否形成更好的框架,而不只是更多候选输出,需要实证回答。
因此,分析单位既不是孤立个人,也不是无差别的集体,而是设问、专业贡献、审慎评估与保存有用成果的制度之间的互动。证据没有指明下一项重要进展由谁产生,却说明个人洞见与组织能力不应被视为相互替代。
原创还需要时间尺度。候选构想可以即时评估其陌生程度,但解释价值与持久性需要继续研究。因此,更多候选项会提高选择与检验的重要性,而不是让这些职能变得多余。
蒸汽与电力的历史,区分了发明、采用和经济效果。其启示不是每种新技术都必须遵循相同日程。
表3. 历史上的扩散与收益时滞
| 案例 | 历史证据 | 启示 |
|---|---|---|
| 蒸汽 | 1830年前对增长贡献有限;最大影响约在Watt发明的一百年后,1850年后的高压蒸汽发挥了潜力。Crafts,2004 | 最初发明与重大经济效果可能相隔甚远。 |
| 电力 | 1881年出现中央电站;1899年电灯进入3%的住宅,电动机占工厂机械动力不足5%;再过约二十年,扩散率才达到约50%。David,1990 | 可获得不等于已广泛用于生产。 |
| 工厂组织 | 工厂重组与电价降低后,生产率效果在20世纪20年代初出现。David,1990 | 互补变化促进采用转化为收益。 |
| 工业革命收入 | Crafts报告,1780—1840年每名劳动者的实际GDP提高38.4%,实际消费性收入提高20.8%,实际产品工资提高44.1%;劳动份额变化不大。Crafts,2020 | 生产率、购买力与分配需要分别测量。 |
电力的意义不限于安装电动机:有效使用还包括围绕新能力重组工厂。David,1990。对于机器智能,这是一种制度类比,而非机械类比。把系统加入不变流程,与围绕经过适当评估的辅助重新设计流程,是不同的工作。
生产率与生活水平的关系仍存在学术讨论。Allen的“恩格斯停顿”描述了人均产出增加时实际工资停滞;Crafts的另一组估计区分消费性收入与产品工资,并认为劳动份额变化有限。Allen,2009;Crafts,2020。这些解释提醒读者,价格尺度、时期与定义会影响结论。生产率增加本身并未说明收益如何分配。
技术变化也创造任务,而不只是替代旧任务。Autor、Chin、Salomons与Seegmiller报告,2018年美国就业约60%处于1940年以来新增的职位名称之中;与职业互补的创新促进新工作出现,自动化创新则减缓这一过程。Autor等,Quarterly Journal of Economics。这既不保证被替代的工作迅速得到补充,也不是对所有劳动市场的预测。
历史命题因而是有条件的:能力、扩散、生产率与工资沿着不同时间表变化。技能、基础设施、组织重构与制度影响这些进程能否汇合。比较支持投资互补资产,却不能推导人工智能的固定等待期。历史上数十年的时滞说明机制,而不是给当代系统排定日程。
机制对规划很重要。如果瓶颈在技能,所需应对就不同于基础设施不足或权限不清的情况。制度评估应辨认缺少哪一种具体互补条件,而不把采用本身当作最终目标。
人机合作有着长久的思想脉络。Licklider在1960年的“Man-Computer Symbiosis”中,设想计算机帮助构思问题与方案、促进合作决策,而不僵化依赖预定程序。Licklider。人类负责目标设定、假说形成、标准建立、计算贡献评估、极低概率情形处理与直觉判断。Licklider。
Engelbart的1962年“Augmenting Human Intellect”把增强能力置于概念框架中心;Bush的1945年“Science, the Endless Frontier”将科学发展与制度目标连接起来,国家科学基金会于1950年成立。Engelbart;NSF历史。这些先例把技术可能性与持久的研究、使用安排联系起来,而非认为技术成就自然就能落实为制度。
这些思想提出的是分工,不是不变的边界。机器可以贡献候选解答、模式与计算探索。人类仍对目标、标准与重大决策承担责任。机器能力变化时,具体任务的配置可以调整,但责任不会因此消失。
命题:在机器快速进步的时期,战略上稀缺的职能是技术深度、制度设计与公共目标的整合。具备技术素养的制度建设者,应了解前沿,分辨已展示的能力与外推;也应了解制度实践,组织可信的采用过程。这项职能更多由团队与机构共同承担,而非孤立人物。
这种整合不只是掌握技术词汇,还包括规定可接受的证据、安排专业复核、明确责任归属,以及让不同技能的人能够使用成果。技术理解让可靠性问题可以问得准确,制度理解让答案真正影响行动。
这里的协作不赋予机器人类地位,而是描述各组成部分具有不同优势与责任的生产安排。互补也不意味着把所有有趣思考留给人,把全部常规操作交给机器。前述里程碑已使这种简单边界难以维持。真正的问题是,哪一种安排既产生可靠结果,又保留可追责的判断。
因此,实用的战略职能既不纯属科学,也不纯属行政,而是连接问题选择与技术可行性、验证与部署、表现与公共收益。应依据这些连接评价其成效,而非依据所用系统的名称。
王阳明的知行合一,为理解与实践的联系提供了哲学语言。《斯坦福哲学百科》王阳明条目。它不等同于科学方法。此处可作的联系更有限:关于实用能力的主张,经过有纪律的行动和可观察结果,才获得实质内容。
Rainhill试验于1829年10月持续九天,约十名参赛者报名,五名及时到场,Rocket赢得500英镑奖金。英国国家铁路博物馆。董事们随后向Stephensons订购了另外四台机车。英国国家铁路博物馆。演示将技术比较与运营决策连接起来,而非让能力停留于抽象承诺。
1881年6月,在Pouilly-le-Fort的演示中,25只接种疫苗的羊经炭疽接种后存活,25只未接种疫苗的对照羊死亡。巴斯德研究所。其相关性在于指定条件下可见的结果对照,而不是假定一次演示回答了之后的全部问题。
现代实地实验把这一逻辑延伸到机器辅助。顾问、客服与开发者研究,在工作场景中考察结果,而不只依赖孤立能力演示。HBS研究;NBER研究;METR开发者试验。结果不同,恰因场景不同而具有信息价值。
命题:机器智能的主张应接受公开、可复现、分领域的检验。适当时采用随机试验或预注册,公开基准并说明人类提供的帮助。评价应涵盖正确性、时间、质量与复核负担,而非只选择增益看起来最大的指标。
实践检验还须匹配预期用途。演示证明系统在指定条件下可以工作;重复评估追问其可靠程度、适用人群与后果;部署再检验制度能否持续保持效果。这一过程使“实践是检验”成为持续纪律,而非一次事件。
问题的尺度超过前沿实验室。联合国估计,2024年世界人口为82亿,预计在21世纪80年代中期达到约103亿的峰值。联合国《世界人口展望2024》。因此,应在不同技能、基础设施与经济环境中评估广泛收益,不能由领先演示直接推断。
IMF估计,全球近40%的就业暴露于人工智能:发达经济体约60%,新兴市场约40%,低收入国家约26%;暴露岗位约一半可能受益,另一半可能受到不利影响。IMF,2024。ILO/NASK指数估计,全球25%的就业具有生成式人工智能潜在暴露,高收入国家为34%,并认为工作转型比替代更可能发生。ILO/NASK,2025。
两组估计采用不同评估框架与技术范围,不能当作简单时间序列,也不能当作实际失业人数。IMF,2024;ILO/NASK,2025。暴露描述技术与任务的潜在关系,并非已经发生的就业结果。
扩散历史与新工作证据提示,广泛共享的收益需要超过“能够访问模型”的条件。David,1990;Autor等,Quarterly Journal of Economics。以下是命题,而非已在所有场景中成立的结论:
按领域与可验证结果衡量能力,而非按名称判断。
在技术部署之外,投入互补的技能、数据、基础设施与制度。
通过实地试验评估并公布结果,包括无显著效果或不利结果。
对目标与重大决策保留人类责任。
追踪分配结果,区分生产率、收入、工作质量与可获得性。
这些命题使面向全体人口的检验可以操作。问题不是人人使用同一系统,而是收益能否经由可行制度到达不同人群。总体表现指标还应配合参与者范围、支持需求与结果变化的测量。
这也将公共目标与结果完全相同的要求区分开来。不同场景可能需要不同工具、培训和防护措施。共同标准是依据明确目标评估选择,并提供易于获取的收益与负担证据,而不是因底层能力出色就预设成功。
历史经验既提示耐心,也提示主动行动。收益可能需要组织学习,但扩散延迟不会自动纠正获得条件的差异。因此,制度建设须把前沿发展与日常使用条件相连接,不预设立即普遍受益,也不预设必然排除某些人群。
当前判断是有条件的。若出现可验证的开放领域自主发现,且不依赖人类选择问题,就会削弱人类设问仍居核心的判断。若系统在漫长、含混的现实任务中持续取得高成功率,则更支持广泛自主性。相反,若既有能力趋势持续未能延续,就应降低对外推的信心。
若跨场景证据表明,互补效果对专业使用者系统性逆转,就须修订所提出的分工。若制度投入未能扩大收益范围,也须重新评估扩散类比。
这些变化应依据可复现结果,而非宣告。有效的评估计划须区分对既定问题给出更好答案、更好地选择问题,以及形成持久公共收益。随着这些证据发展,本框架应保持可修订性。
结论
证据的分量支持介于两种流行极端之间的立场。机器在许多明确、可验证的任务上已经超过多数人;特定领域的计算贡献也获得了最高层级的科学认可,因此不能笼统否认机器原创性。AlphaZero,Science;诺贝尔化学奖。但收益并不均衡,实地效果仍依赖人类的问题选择、评估与制度嵌入。HBS实地证据;METR开发者试验。
历史提示,决定性变量通常不是名称或原始能力本身,而是将潜力转化为共享收益的互补制度、技能与实践检验,这往往历时数十年。David的电气化研究。因此,本文提出以技术素养连接前沿科学与公共目标的制度建设职能,Bush、Licklider和Engelbart提供了历史先例。NSF历史;Licklider;Engelbart。
“协作者”描述可追责的安排,而非智能的最终判决。其价值须由有效发现、可靠工作与超越前沿的广泛收益在实践中确立。适当的立场,是向机器原创性开放,同时严格评估能力如何转化为用途。能力进步值得认真对待,但只有验证、责任与广泛受益相互衔接,技术转变才构成可持续的公共进步。
McCarthy, Minsky, Rochester, and Shannon (1955). Dartmouth proposal introducing artificial intelligence. https://www-formal.stanford.edu/jmc/history/dartmouth/dartmouth.html
Good, I. J. (1965). Speculations Concerning the First Ultraintelligent Machine. Advances in Computers 6: 31–88. https://vtechworks.lib.vt.edu/server/api/core/bitstreams/a5e423ee-54e0-4eec-aeca-32b73f851af5/content
Bostrom (2014). Superintelligence: Paths, Dangers, Strategies. Book information and supporting materials. https://nickbostrom.com/superintelligence
White House (29 September 2026). Fact Sheet: President Donald J. Trump Inaugurates the Era of Super Intelligence. https://www.whitehouse.gov/fact-sheets/2026/09/fact-sheet-president-donald-j-trump-inaugurates-the-era-of-super-intelligence/
BBC (30 September 2026). Report on the voluntary frontier-developer safety accord and its stated commitments. https://www.bbc.com/news/articles/cme30dz5vkzko
Google DeepMind. AlphaGo research overview, including move 37 in game two. https://deepmind.google/research/alphago/
Silver et al. (2016). Mastering the Game of Go with Deep Neural Networks and Tree Search. Nature 529: 484–489; supplied Nature coverage. https://www.nature.com/articles/nature.2016.19575
Silver et al. (2018). A General Reinforcement Learning Algorithm That Masters Chess, Shogi, and Go through Self-Play. Science 362: 1140–1144. https://www.science.org/doi/abs/10.1126/science.aar6404
Nobel Prize (9 October 2024). The Nobel Prize in Chemistry 2024: computational protein design and protein structure prediction. Press release. https://www.nobelprize.org/prizes/chemistry/2024/press-release/
Google DeepMind (21 July 2025). Advanced Version of Gemini with Deep Think Officially Achieves Gold-Medal Standard at the International Mathematical Olympiad. https://deepmind.google/blog/advanced-version-of-gemini-with-deep-think-officially-achieves-gold-medal-standard-at-the-international-mathematical-olympiad/
Euronews (22 July 2025). Did Google DeepMind or OpenAI Win Gold at the World's Most Prestigious Math Competition? https://www.euronews.com/2025/07/22/did-google-deepmind-or-openai-win-gold-at-the-worlds-most-prestigious-math-competition
The Register (May 2025). Report on AlphaEvolve and improved multiplication of complex-valued matrices. https://www.theregister.com/software/2025/05/15/google-deepmind-debuts-algorithm-evolving-agent-alphaevolve/766919
Dell'Acqua, McFowland, Mollick, Lifshitz, Kellogg, Rajendran, et al. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality. Harvard Business School. https://www.hbs.edu/ris/Publication Files/dell-acqua-et-al-2026-navigating-the-jagged-technological-frontier_5c589c8c-fbb5-458f-b285-c944746cd717.pdf
Brynjolfsson, Li, and Raymond (2025). Generative AI at Work. Quarterly Journal of Economics 140(2): 889–942; NBER Working Paper 31161. https://www.nber.org/papers/w31161
METR (10 July 2025). Early-2025 AI Experienced Open-Source Developer Study. Randomized controlled trial report. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
Si, Yang, and Hashimoto (2024). Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers. arXiv:2409.04109. https://arxiv.org/abs/2409.04109
Shumailov et al. (2024). AI Models Collapse When Trained on Recursively Generated Data. Nature 631: 755–759. https://www.nature.com/articles/s41586-024-07566-y
METR (19 March 2025). Measuring AI Ability to Complete Long Tasks. https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/
Stanford Encyclopedia of Philosophy. Thomas Kuhn; discussion of The Structure of Scientific Revolutions (1962). https://plato.stanford.edu/entries/thomas-kuhn/
Sinatra, Wang, Deville, Song, and Barabási (2016). Quantifying the Evolution of Individual Scientific Impact. Science 354: aaf5239. https://www.barabasi.com/media/pub_imports/files/825.pdf
Wuchty, Jones, and Uzzi (2007). Study of the increasing dominance of teams in knowledge production. Science 316(5827): 1036–1039. https://pubmed.ncbi.nlm.nih.gov/17431139/
Jones (2009). Study of the burden of knowledge and the rising age at great invention. Review of Economic Studies 76(1): 283–317. https://academic.oup.com/restud/article-abstract/76/1/283/1577537
Park, Leahey, and Funk (2023). Study of declining disruptiveness in papers and patents. Nature 613: 138–144; bibliographic record. https://ideas.repec.org/a/nat/nature/v613y2023i7942d10.1038_s41586-022-05543-x.html
Petersen, Arroyave, and Pammolli (2024). The Disruption Index Is Biased by Citation Inflation. Quantitative Science Studies 5(4). https://direct.mit.edu/qss/article/5/4/936/124788/The-disruption-index-is-biased-by-citation
Bloom, Jones, Van Reenen, and Webb (2020). Study of research inputs and idea-production effectiveness, including Moore's law. American Economic Review 110(4): 1104–1144. https://www.aeaweb.org/articles?id=10.1257/aer.20180338
Crafts (2004). Steam as a General Purpose Technology: A Growth Accounting Perspective. Economic Journal 114: 338–351. https://onlinelibrary.wiley.com/doi/abs/10.1111/j.1468-0297.2003.00200.x
David (1990). The Dynamo and the Computer: An Historical Perspective on the Modern Productivity Paradox. American Economic Review 80(2): 355–361. https://gwern.net/doc/economics/automation/1990-david.pdf
Allen (2009). Engels' Pause: Technical Change, Capital Accumulation, and Inequality in the British Industrial Revolution. Explorations in Economic History. https://www.nuffield.ox.ac.uk/Users/Allen/engelspause.pdf
Crafts (2020). Slow Real Wage Growth during the Industrial Revolution: Productivity Paradox or Pro-Rich Growth? CAGE Working Paper 474. https://warwick.ac.uk/fac/soc/economics/research/centres/cage/wp474.2020.pdf
Autor, Chin, Salomons, and Seegmiller (2024). New Frontiers: The Origins and Content of New Work, 1940–2018. Quarterly Journal of Economics 139(3): 1399–1465. https://academic.oup.com/qje/article/139/3/1399/7630187
Licklider (1960). Man-Computer Symbiosis. IRE Transactions on Human Factors in Electronics HFE-1: 4–11. https://groups.csail.mit.edu/medg/people/psz/Licklider.html
Engelbart (1962). Augmenting Human Intellect: A Conceptual Framework. SRI Summary Report AFOSR-3223. https://www.dougengelbart.org/content/view/138
National Science Foundation. NSF History: Bush's Science, the Endless Frontier (1945) and the establishment of NSF (1950). https://www.nsf.gov/about/history
Stanford Encyclopedia of Philosophy. Wang Yangming: the unity of knowing and acting. https://plato.stanford.edu/entries/wang-yangming/
Science Museum Group / National Railway Museum. Stephenson's Rocket, Rainhill and the Rise of the Locomotive. https://www.railwaymuseum.org.uk/objects-and-stories/stephensons-rocket-rainhill-and-rise-locomotive
Institut Pasteur. Louis Pasteur, a Universal Legacy: the Pouilly-le-Fort vaccination demonstration. https://www.pasteur.fr/en/whats-new/latest-news/features/louis-pasteur-universal-legacy
United Nations (2024). World Population Prospects 2024: Key Messages. https://population.un.org/wpp/assets/Files/WPP2024_Key-Messages.pdf
International Monetary Fund (14 January 2024). AI Will Transform the Global Economy. Let's Make Sure It Benefits Humanity. https://www.imf.org/en/blogs/articles/2024/01/14/ai-will-transform-the-global-economy-lets-make-sure-it-benefits-humanity
International Labour Organization / NASK (2025). Generative AI and Jobs: A Refined Global Index; news summary of global employment exposure. https://www.ilo.org/resource/news/one-four-jobs-risk-being-transformed-genai-new-ilo–nask-global-index-shows
A historical and evidence-based assessment of 'Super Intelligence' in the technology cycle
Machine intelligence is best understood through measured capabilities, not a change of name. The evidence supports a position between two extremes: machines already exceed most people on many well-defined, verifiable tasks, and specific contributions have received recognition at the highest scientific level; a blanket denial of machine originality is therefore unsupported. DeepMind's capability record; Nobel chemistry award. Yet progress remains uneven, and workplace benefits depend on problem selection, evaluation and institutional integration. Field evidence from knowledge work. Distinguishing solution generation from problem selection is essential to the assessment. Historical evidence suggests that complementary institutions, skills and practical tests, often developed over decades, determine whether capability becomes broadly shared benefit, as electrification illustrates. This essay proposes the technology-literate institution builder, connecting frontier science with public purpose, as a useful strategic function. Partnership is neither an assertion of machine personhood nor a guarantee of progress: it is a design problem involving technical capacities, human responsibility and measurable outcomes.
The terminology has a substantial history. McCarthy, Minsky, Rochester and Shannon used “artificial intelligence” in the 1955 Dartmouth proposal; Good discussed an “ultraintelligent machine” in 1965. Dartmouth proposal; Good, 1965. Bostrom's 2014 definition set a demanding threshold: “any intellect that greatly exceeds the cognitive performance of humans in virtually all domains of interest.” Bostrom, Superintelligence. These terms express different conceptual ambitions, not interchangeable measurements.
The executive order signed on 29 September 2026 directs executive-branch departments and agencies to use “Super Intelligence” and “SI” in official correspondence, public communications, policy documents and non-statutory documents, replacing “Artificial Intelligence” and “AI.” White House fact sheet. It also directs the Assistant to the President for Science and Technology to propose a federal definition reflecting the state of the technology and identify additional executive action. White House fact sheet. The request for a definition makes an evidence framework especially relevant.
A voluntary industry safety commitment was announced the same day, covering safeguards, rapid detection and correction of problems, and work with independent auditors; the President described it as “morally binding,” while BBC reporting noted that consequences for a breach were not specified. BBC report. This describes the commitment's stated arrangements, not a measured safety outcome.
A label can coordinate discussion, but cannot establish the scope, reliability or consequences of a capability. Bostrom's “virtually all domains” criterion requires something much broader than outstanding performance on selected problems. Bostrom, Superintelligence. The relevant inquiry is therefore domain-specific: what task was attempted, under what conditions, with what assistance, and against which human comparison?
The same discipline applies to originality. Producing an unfamiliar output, discovering a valid solution and establishing a useful new explanatory framework are related but distinct achievements. An assessment should ask who selected the problem, whether the result is genuinely new, how it was verified, and whether it survives use outside its initial setting. No single benchmark can answer all these questions. The administrative terminology and the empirical evaluation can proceed together without treating either as a substitute for the other.
Selected milestones show why a universal claim that machines only imitate is too strong. They do not, by themselves, establish superiority across all forms of reasoning or the autonomous direction of science.
Table 1. Capability milestones and the boundaries of their evidence
| Date | Milestone | What was demonstrated | Boundary |
|---|---|---|---|
| 2016 | AlphaGo | Superhuman Go performance; move 37 in game two was estimated to have a 1 in 10,000 chance of being played by a human. DeepMind; Nature | A game with explicit rules and outcomes. |
| 2018 | AlphaZero | Starting from random play with only the rules, achieved superhuman chess, shogi and Go performance. Silver et al., Science | Self-play operates inside specified environments. |
| 2024 | Protein science | Nobel chemistry recognition for David Baker's computational protein design and Demis Hassabis and John Jumper's protein structure prediction. Nobel Prize | Recognition concerns specific scientific contributions, not universal intelligence. |
| 2025 | Mathematical reasoning | Advanced Gemini Deep Think solved five of six IMO problems, scoring 35/42; an experimental OpenAI model also reportedly reached gold-level performance. DeepMind; Euronews | IMO coordinators graded Gemini's solutions; the review did not validate the system or model. |
| 2025 | AlphaEvolve | A method for multiplying 4×4 complex matrices using 48 scalar multiplications, improving on Strassen's 1969 algorithm. The Register | A verified improvement for a defined computational problem. |
The Nobel distinction is important: people received the award for scientific work involving computational systems, rather than a machine receiving a prize. Nobel Prize. Nevertheless, the recognised contribution demonstrates that computational methods can become central to science at its highest recognised level. Nobel Prize.
AlphaZero is especially relevant to the hypothesis that a system's ceiling is fixed by its trainers' intellect: its results exceeded human performance without starting from human game examples, although its rules and learning environment remained human-designed. Silver et al., Science. The distinction is between supplying an environment for learning and specifying every solution that learning can produce.
These examples share clear rules, verifiable outcomes or evaluable outputs. Search and feedback can explore alternatives beyond an individual's repertoire. A surprising Go move and a better multiplication procedure are different kinds of novelty, but neither is adequately described as copying a known answer.
The inference should remain bounded. A proof can be checked against a mathematical problem, whereas deciding which problem deserves attention involves further judgments. Protein prediction and design link computation to scientific investigation, but the award does not resolve every question about autonomy. The strongest justified conclusion is that machine-assisted originality exists in specific domains; its range, reliability and relationship to human problem-setting still require investigation.
Field evidence does not support a uniform account of assistance. Benefits vary with the task, the user's prior skill and the surrounding workflow.
Table 2. Evidence on AI, work and research
| Study and setting | Reported result | Qualification |
|---|---|---|
| Dell'Acqua et al.; 758 BCG consultants | Inside the frontier: 12.2% more tasks, 25.1% faster completion, quality about 30% or more higher. Outside it: 19 percentage points less likely to be correct; arm-level correctness was 60% with GPT plus overview, 71% with GPT only and 84.4% for controls. HBS study | The frontier cuts across task types; lower initial performers benefited more. |
| Brynjolfsson, Li and Raymond; 5,179 support agents | Issues resolved per hour rose 14% on average and 34% for novice and lower-skilled agents; experienced, highly skilled agents saw minimal gains. NBER study | Evidence from customer support, not every occupation. |
| METR; 16 experienced developers, 246 issues | AI assistance increased completion time by 19%; participants expected a 24% speedup and afterwards believed they had gained 20%. METR developer trial | Familiar repositories and experienced developers; authors limit generalisation. |
| Si, Yang and Hashimoto; 100+ NLP researchers | LLM ideas were judged more novel, with p < 0.05, but slightly weaker on feasibility. Research-idea study | Novelty judgments are difficult; self-evaluation and diversity have limitations. |
The consultant study's reported 19-percentage-point effect and its separate treatment-arm percentages are different statistical summaries, not interchangeable arithmetic comparisons. HBS study. Its central lesson is that a tool can improve closely related tasks while reducing correctness on another: apparently similar work may sit on different sides of the frontier.
The support-agent and consultant findings suggest compression of some skill differences within suitable tasks, rather than the disappearance of expertise. NBER study; HBS study. Such compression is potentially valuable, but it does not establish that every user benefits equally or that evaluation can be removed.
METR's discrepancy between expectation, retrospective belief and measured time illustrates why perceived helpfulness is not sufficient evidence of productivity. METR developer trial. The result neither settles the future of programming nor invalidates gains elsewhere. It supports direct measurement in the particular setting where a deployment is intended.
Two further findings sharpen the boundary. Shumailov and colleagues found that indiscriminate training on recursively generated data can cause model collapse, including loss of the original distribution's tails. Shumailov et al., Nature. This concerns a training process, not an unavoidable fate of every use of generated data. METR reported that the human-professional task length completed by frontier agents with 50% reliability had doubled about every 7 months over six years, with an estimated range of roughly 1–4 doublings annually and methodological and extrapolation caveats. METR task-horizon study. A rapidly improving trend is consequential, but 50% reliability is not dependable completion of every long task.
Together, these studies support a milder trainer-ceiling hypothesis: in open-ended settings without reliable verifiers, outcomes remain strongly dependent on how people formulate and assess problems. That is a proposition about the organisation of inquiry, not a demonstrated upper bound on machine ability. The research-idea experiment suggests novelty can exceed expert judgments on one dimension while remaining weaker on feasibility. Research-idea study. Neither universal superiority nor permanent incapacity follows.
For institutional decisions, this means separating the capacity to generate an answer from the cost of establishing that the answer is usable. A faster draft may still require substantial checking; a slower workflow may deliver a different quality of result. These are possibilities to measure, not assumptions to build into every evaluation. The comparison should preserve the full task, including review and correction.
Kuhn's account distinguishes normal science within a paradigm from revolutionary episodes that replace an earlier framework. Stanford Encyclopedia of Philosophy on Kuhn. It helps separate solving a difficult problem within an accepted framework from changing the framework itself. It does not supply a simple test for predicting the next transformation.
Exceptional impact and the conditions supporting it must also be distinguished. Sinatra and colleagues modelled scientific impact through productivity, luck and a career-stable ability parameter, Q; the highest-impact work occurred randomly within the sequence of publications. Sinatra et al., Science. This is consistent with heavy-tailed impact, where exceptional outcomes carry disproportionate weight, but not with assigning permanent categories of originality to people. Sinatra et al., Science. A record of exceptional results does not establish when the next result will appear.
Collective capacity is equally visible. Wuchty, Jones and Uzzi examined 19.9 million papers across five decades and 2.1 million patents, finding that teams increasingly dominated and also produced exceptionally high-impact work. Wuchty et al., Science. Jones found that the mean age at great invention increased by about six years over the 20th century, consistent with a growing burden of knowledge. Jones, Review of Economic Studies. Individual insight therefore operates within accumulated learning, specialisation and coordination.
Claims about a general slowing of novelty require care. Park, Leahey and Funk studied 45 million papers and 3.9 million patents and reported declining disruptiveness; Petersen, Arroyave and Pammolli subsequently argued that the disruption index is biased by citation inflation. Park et al., Nature; Quantitative Science Studies critique. The disagreement concerns how a metric tracks scientific change, not whether worthwhile research continues.
Bloom and colleagues found that maintaining Moore's-law doubling required more than 18 times the number of researchers needed in the early 1970s. Bloom et al., American Economic Review. That result concerns research inputs and effectiveness in a particular trajectory; it is not a census of originality.
Proposition: machine assistance may ease the burden of knowledge by helping teams organise, search and test accumulated information. It may also alter the composition of teams and the allocation of attention. Whether this produces better frameworks, rather than simply more candidate outputs, is an empirical question.
The appropriate unit of analysis is consequently neither the isolated individual nor an undifferentiated collective. It is the interaction among problem-setting, specialised contributions, critical evaluation and institutions that preserve useful results. The evidence leaves open who will originate the next important advance, while showing why individual insight and organised capacity should not be treated as alternatives.
Originality also needs a timescale. A candidate idea can be assessed immediately for unfamiliarity, but its explanatory value and durability require further inquiry. The prospect of producing more candidates therefore increases the importance of selection and testing, rather than making those functions unnecessary.
The histories of steam and electricity separate invention, adoption and economic effect. Their lesson is not that every new technology must follow an identical timetable.
Table 3. Historical diffusion and benefit lags
| Case | Historical evidence | Implication |
|---|---|---|
| Steam | Little contribution to growth before 1830; peak impact about a hundred years after Watt's invention, with high-pressure steam after 1850 realising its potential. Crafts, 2004 | Initial invention and major economic effects can be widely separated. |
| Electricity | Central stations in 1881; in 1899, electric lighting reached 3% of residences and motors supplied under 5% of factory mechanical drive; about two further decades to approximately 50% diffusion. David, 1990 | Availability does not equal widespread productive use. |
| Factory organisation | Productivity effects appeared in the early 1920s following factory reorganisation and cheaper electricity. David, 1990 | Complementary changes help convert adoption into gains. |
| Industrial Revolution earnings | For 1780–1840, Crafts reports real GDP per worker up 38.4%, real consumption earnings up 20.8%, and real product wages up 44.1%, with little change in labour's share. Crafts, 2020 | Productivity, purchasing power and distribution require separate measures. |
Electricity's importance was not confined to installing motors: productive use involved reorganising factories around the new capability. David, 1990. The analogy for machine intelligence is institutional rather than mechanical. Adding a system to an unchanged process and redesigning a process around appropriately evaluated assistance are different undertakings.
The relationship between productivity and living standards remains a scholarly question. Allen's “Engels' pause” account describes stagnant real wages alongside expanding output per worker, while Crafts' alternative estimates distinguish consumption earnings from product wages and find little change in labour's share. Allen, 2009; Crafts, 2020. These approaches remind readers that price measures, time periods and definitions affect conclusions. A productivity gain alone does not specify its distribution.
Technological change also creates tasks, not merely substitutes for existing ones. Autor, Chin, Salomons and Seegmiller report that approximately 60% of US employment in 2018 was in job titles introduced since 1940; new work emerged where innovations complemented occupations, while automation innovations slowed its emergence. Autor et al., Quarterly Journal of Economics. This finding is neither a guarantee that displaced work is replaced promptly nor a prediction about every labour market.
The historical proposition is therefore conditional: capability, diffusion, productivity and wages move on different clocks. Skills, infrastructure, organisational redesign and institutions influence whether these clocks converge. The comparison supports investing in complementary assets, while offering no warrant to infer a fixed waiting period for AI. Decades-long historical lags are evidence about mechanisms, not a timetable for contemporary systems.
The mechanism matters for planning. If the limiting condition is skill, the response differs from one suited to inadequate infrastructure or unclear authority. Institutional assessment should identify the specific complement that is missing, rather than treating adoption itself as the final objective.
Human–machine partnership has a long intellectual lineage. Licklider's 1960 “Man-Computer Symbiosis” envisaged computers facilitating formulative thinking and cooperative decisions without inflexible dependence on predetermined programs. Licklider. People would set goals, formulate hypotheses, establish criteria, evaluate computer contributions, address very-low-probability situations and supply intuitive judgment. Licklider.
Engelbart's 1962 “Augmenting Human Intellect” placed augmentation at the centre of a conceptual framework, while Bush's 1945 “Science, the Endless Frontier” connected scientific development with institutional purpose; the National Science Foundation was established in 1950. Engelbart; NSF history. These are precedents for linking technical possibility to durable arrangements for research and use, rather than treating technical achievement as institutionally self-executing.
Their ideas suggest a division of labour, not an unchanging boundary. Machines may contribute candidate solutions, patterns and computational exploration. Humans remain accountable for goals, criteria and consequential decisions. As machine capacities change, the allocation of particular tasks can change without dissolving that responsibility.
Proposition: the strategically scarce function in a period of fast machine progress is integration of technical depth, institutional design and public purpose. The technology-literate institution builder understands enough of the frontier to distinguish a demonstrated capability from an extrapolation, and enough of institutional practice to organise trustworthy adoption. This function is more often exercised through teams and institutions than through solitary figures.
Such integration requires more than technical vocabulary. It involves specifying acceptable evidence, arranging expert review, identifying where responsibility lies and ensuring that results can be used by people with different skills. Technical understanding makes it possible to ask precise questions about reliability; institutional understanding makes answers consequential.
Partnership in this sense does not confer human status on a machine. It describes a productive arrangement whose components have different strengths and responsibilities. Nor does complementarity mean assigning all interesting thinking to people and all routine operations to machines. The milestone evidence already makes that simple boundary difficult to sustain. The relevant question is which arrangement produces sound results while preserving accountable judgment.
The practical strategic profile is thus neither exclusively scientific nor exclusively administrative. It connects problem selection to technical feasibility, verification to deployment, and performance to public benefit. Its success should be measured by these connections, not by a label attached to the system it employs.
Wang Yangming's unity of knowing and acting, 知行合一, offers a philosophical vocabulary for connecting understanding with practice. Stanford Encyclopedia of Philosophy on Wang Yangming. It is not identical to the scientific method. The useful connection here is narrower: a claim about practical capacity gains substance when exposed to disciplined action and observable results.
The Rainhill Trials ran for nine days in October 1829; about ten competitors entered their names and five arrived in time, with Rocket winning the £500 prize. National Railway Museum. The directors subsequently ordered four more locomotives from the Stephensons. National Railway Museum. The demonstration connected a technical comparison to an operational decision, rather than leaving capability as an abstract promise.
At Pouilly-le-Fort in June 1881, 25 vaccinated sheep survived anthrax inoculation while 25 unvaccinated controls died. Institut Pasteur. The relevance lies in the visible contrast between outcomes under specified conditions, not in assuming that any single demonstration answers every subsequent question.
Modern field experiments extend this logic to machine assistance. Consultant, support-agent and developer studies examine outcomes within work settings rather than relying only on demonstrations of isolated capability. HBS study; NBER study; METR developer trial. Their differing results are informative precisely because the settings differ.
Proposition: claims about machine intelligence should face public, reproducible, domain-specific tests. Where appropriate, trials should be randomised or pre-registered, with transparent benchmarks and a clear account of human assistance. Evaluation should include correctness, time, quality and the demands placed on reviewers, rather than selecting whichever measure presents the largest apparent gain.
A practical test must also match the intended use. A demonstration establishes that something can work under its conditions; repeated evaluation asks how reliably it works, for whom and with which consequences. Deployment then tests whether the institution can sustain the result. This sequence turns “practice is the test” into a continuing discipline rather than a single event.
The scale of the problem extends beyond frontier laboratories. The United Nations estimates a world population of 8.2 billion in 2024, with a projected peak of about 10.3 billion in the mid-2080s. UN World Population Prospects 2024. Broad benefit must therefore be assessed across heterogeneous skills, infrastructures and economic settings, rather than inferred from leading demonstrations.
The IMF estimated that almost 40% of global employment was exposed to AI: about 60% in advanced economies, 40% in emerging markets and 26% in low-income countries, with roughly half of exposed jobs potentially benefiting and half potentially harmed. IMF, 2024. The ILO/NASK index estimated potential generative-AI exposure for 25% of global employment and 34% in high-income countries, describing transformation as more likely than replacement. ILO/NASK, 2025.
These estimates concern different assessment frameworks and technology scopes; they should not be read as a simple time series or as counts of actual job losses. IMF, 2024; ILO/NASK, 2025. Exposure identifies a potential relationship between technology and tasks, not a completed employment outcome.
The diffusion history and the evidence on new work suggest that broadly shared gains require more than access to a model. David, 1990; Autor et al., Quarterly Journal of Economics. The following are propositions, not findings already established for all settings:
Measure capability by domain and verified outcome, not by label.
Fund complementary skills, data, infrastructure and institutions, alongside technical deployment.
Evaluate through field trials and publish results, including null or adverse findings.
Keep human accountability for goals and consequential decisions.
Track distributional outcomes, distinguishing productivity, earnings, work quality and access.
Together, these propositions make the whole-population test operational. The question is not whether every person uses the same system, but whether benefits can reach different populations through workable institutions. Measures of aggregate performance should be accompanied by measures of who can participate, what support they require and how outcomes change.
This also separates public purpose from a demand for identical outcomes. Different settings may need different tools, training and safeguards. The common standard is that choices should be evaluated against explicit goals, with accessible evidence about benefits and burdens, rather than presumed successful because the underlying capability is impressive.
Historical experience offers reasons for both patience and deliberate action. Benefits may require organisational learning, while delayed diffusion does not automatically correct differences in access. An institution-building strategy should therefore connect frontier development to the conditions of everyday use, without assuming either immediate universal gain or inevitable exclusion.
The present assessment is conditional. Verified autonomous discovery in open-ended domains, without human problem selection, would weaken the view that human framing remains central. Sustained high success on long, ambiguous real-world tasks would strengthen the case for broader machine autonomy. Conversely, a persistent failure of the observed capability trend to continue would reduce confidence in extrapolation.
Evidence that complementarity effects systematically reverse for expert users across settings would require revising the proposed division of labour. Evidence that institutional investment fails to broaden gains would also require reassessing the diffusion analogy.
These changes should be judged by reproducible results rather than declarations. A useful evaluation programme would distinguish better answers to specified questions from better selection of questions, and both from durable public benefit. The framework is intended to remain revisable as those forms of evidence develop.
Conclusion
The weight of evidence supports a position between two popular extremes. Machine systems already exceed most people on many defined, verifiable tasks, and computational contributions have received recognition at the highest scientific level in particular domains; a blanket denial of machine originality is not supported. AlphaZero, Science; Nobel Prize in Chemistry. Yet gains remain uneven, and field outcomes depend on human problem selection, evaluation and institutional embedding. HBS field evidence; METR developer trial.
History suggests that the decisive variable is rarely the label or raw capability alone: complementary institutions, skills and practical tests convert potential into shared benefit, often over decades. David's electrification history. The proposed strategic function is therefore the technology-literate institution builder, connecting frontier science with public purpose, with historical precedents in Bush, Licklider and Engelbart. NSF history; Licklider; Engelbart.
“Partner” describes an accountable arrangement, not a final verdict on intelligence. Its merit must be established in practice: through valid discoveries, reliable work and benefits that extend beyond the frontier. The appropriate stance combines openness to machine originality with disciplined evaluation of how capability becomes useful. Capability gains deserve serious attention, but a technological transition becomes sustainable public progress only when verification, responsibility and broadly shared benefit are connected.
McCarthy, Minsky, Rochester, and Shannon (1955). Dartmouth proposal introducing artificial intelligence. https://www-formal.stanford.edu/jmc/history/dartmouth/dartmouth.html
Good, I. J. (1965). Speculations Concerning the First Ultraintelligent Machine. Advances in Computers 6: 31–88. https://vtechworks.lib.vt.edu/server/api/core/bitstreams/a5e423ee-54e0-4eec-aeca-32b73f851af5/content
Bostrom (2014). Superintelligence: Paths, Dangers, Strategies. Book information and supporting materials. https://nickbostrom.com/superintelligence
White House (29 September 2026). Fact Sheet: President Donald J. Trump Inaugurates the Era of Super Intelligence. https://www.whitehouse.gov/fact-sheets/2026/09/fact-sheet-president-donald-j-trump-inaugurates-the-era-of-super-intelligence/
BBC (30 September 2026). Report on the voluntary frontier-developer safety accord and its stated commitments. https://www.bbc.com/news/articles/cme30dz5vkzko
Google DeepMind. AlphaGo research overview, including move 37 in game two. https://deepmind.google/research/alphago/
Silver et al. (2016). Mastering the Game of Go with Deep Neural Networks and Tree Search. Nature 529: 484–489; supplied Nature coverage. https://www.nature.com/articles/nature.2016.19575
Silver et al. (2018). A General Reinforcement Learning Algorithm That Masters Chess, Shogi, and Go through Self-Play. Science 362: 1140–1144. https://www.science.org/doi/abs/10.1126/science.aar6404
Nobel Prize (9 October 2024). The Nobel Prize in Chemistry 2024: computational protein design and protein structure prediction. Press release. https://www.nobelprize.org/prizes/chemistry/2024/press-release/
Google DeepMind (21 July 2025). Advanced Version of Gemini with Deep Think Officially Achieves Gold-Medal Standard at the International Mathematical Olympiad. https://deepmind.google/blog/advanced-version-of-gemini-with-deep-think-officially-achieves-gold-medal-standard-at-the-international-mathematical-olympiad/
Euronews (22 July 2025). Did Google DeepMind or OpenAI Win Gold at the World's Most Prestigious Math Competition? https://www.euronews.com/2025/07/22/did-google-deepmind-or-openai-win-gold-at-the-worlds-most-prestigious-math-competition
The Register (May 2025). Report on AlphaEvolve and improved multiplication of complex-valued matrices. https://www.theregister.com/software/2025/05/15/google-deepmind-debuts-algorithm-evolving-agent-alphaevolve/766919
Dell'Acqua, McFowland, Mollick, Lifshitz, Kellogg, Rajendran, et al. Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality. Harvard Business School. https://www.hbs.edu/ris/Publication Files/dell-acqua-et-al-2026-navigating-the-jagged-technological-frontier_5c589c8c-fbb5-458f-b285-c944746cd717.pdf
Brynjolfsson, Li, and Raymond (2025). Generative AI at Work. Quarterly Journal of Economics 140(2): 889–942; NBER Working Paper 31161. https://www.nber.org/papers/w31161
METR (10 July 2025). Early-2025 AI Experienced Open-Source Developer Study. Randomized controlled trial report. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
Si, Yang, and Hashimoto (2024). Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers. arXiv:2409.04109. https://arxiv.org/abs/2409.04109
Shumailov et al. (2024). AI Models Collapse When Trained on Recursively Generated Data. Nature 631: 755–759. https://www.nature.com/articles/s41586-024-07566-y
METR (19 March 2025). Measuring AI Ability to Complete Long Tasks. https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/
Stanford Encyclopedia of Philosophy. Thomas Kuhn; discussion of The Structure of Scientific Revolutions (1962). https://plato.stanford.edu/entries/thomas-kuhn/
Sinatra, Wang, Deville, Song, and Barabási (2016). Quantifying the Evolution of Individual Scientific Impact. Science 354: aaf5239. https://www.barabasi.com/media/pub_imports/files/825.pdf
Wuchty, Jones, and Uzzi (2007). Study of the increasing dominance of teams in knowledge production. Science 316(5827): 1036–1039. https://pubmed.ncbi.nlm.nih.gov/17431139/
Jones (2009). Study of the burden of knowledge and the rising age at great invention. Review of Economic Studies 76(1): 283–317. https://academic.oup.com/restud/article-abstract/76/1/283/1577537
Park, Leahey, and Funk (2023). Study of declining disruptiveness in papers and patents. Nature 613: 138–144; bibliographic record. https://ideas.repec.org/a/nat/nature/v613y2023i7942d10.1038_s41586-022-05543-x.html
Petersen, Arroyave, and Pammolli (2024). The Disruption Index Is Biased by Citation Inflation. Quantitative Science Studies 5(4). https://direct.mit.edu/qss/article/5/4/936/124788/The-disruption-index-is-biased-by-citation
Bloom, Jones, Van Reenen, and Webb (2020). Study of research inputs and idea-production effectiveness, including Moore's law. American Economic Review 110(4): 1104–1144. https://www.aeaweb.org/articles?id=10.1257/aer.20180338
Crafts (2004). Steam as a General Purpose Technology: A Growth Accounting Perspective. Economic Journal 114: 338–351. https://onlinelibrary.wiley.com/doi/abs/10.1111/j.1468-0297.2003.00200.x
David (1990). The Dynamo and the Computer: An Historical Perspective on the Modern Productivity Paradox. American Economic Review 80(2): 355–361. https://gwern.net/doc/economics/automation/1990-david.pdf
Allen (2009). Engels' Pause: Technical Change, Capital Accumulation, and Inequality in the British Industrial Revolution. Explorations in Economic History. https://www.nuffield.ox.ac.uk/Users/Allen/engelspause.pdf
Crafts (2020). Slow Real Wage Growth during the Industrial Revolution: Productivity Paradox or Pro-Rich Growth? CAGE Working Paper 474. https://warwick.ac.uk/fac/soc/economics/research/centres/cage/wp474.2020.pdf
Autor, Chin, Salomons, and Seegmiller (2024). New Frontiers: The Origins and Content of New Work, 1940–2018. Quarterly Journal of Economics 139(3): 1399–1465. https://academic.oup.com/qje/article/139/3/1399/7630187
Licklider (1960). Man-Computer Symbiosis. IRE Transactions on Human Factors in Electronics HFE-1: 4–11. https://groups.csail.mit.edu/medg/people/psz/Licklider.html
Engelbart (1962). Augmenting Human Intellect: A Conceptual Framework. SRI Summary Report AFOSR-3223. https://www.dougengelbart.org/content/view/138
National Science Foundation. NSF History: Bush's Science, the Endless Frontier (1945) and the establishment of NSF (1950). https://www.nsf.gov/about/history
Stanford Encyclopedia of Philosophy. Wang Yangming: the unity of knowing and acting. https://plato.stanford.edu/entries/wang-yangming/
Science Museum Group / National Railway Museum. Stephenson's Rocket, Rainhill and the Rise of the Locomotive. https://www.railwaymuseum.org.uk/objects-and-stories/stephensons-rocket-rainhill-and-rise-locomotive
Institut Pasteur. Louis Pasteur, a Universal Legacy: the Pouilly-le-Fort vaccination demonstration. https://www.pasteur.fr/en/whats-new/latest-news/features/louis-pasteur-universal-legacy
United Nations (2024). World Population Prospects 2024: Key Messages. https://population.un.org/wpp/assets/Files/WPP2024_Key-Messages.pdf
International Monetary Fund (14 January 2024). AI Will Transform the Global Economy. Let's Make Sure It Benefits Humanity. https://www.imf.org/en/blogs/articles/2024/01/14/ai-will-transform-the-global-economy-lets-make-sure-it-benefits-humanity
International Labour Organization / NASK (2025). Generative AI and Jobs: A Refined Global Index; news summary of global employment exposure. https://www.ilo.org/resource/news/one-four-jobs-risk-being-transformed-genai-new-ilo–nask-global-index-shows
