OpenAI 在科学领域的工作源于一个简单的信念:先进的 AI 可以成为科学家的强大合作伙伴,帮助他们探索更多想法、连接不同概念、设计更好的实验,并加速造福人类的发现。我们已经分享了模型在数学领域贡献新颖成果的早期案例,包括单位距离问题的研究、理论物理学中关于胶子振幅的新成果,以及生物学中 GPT‑5 在自动化实验室帮助降低无细胞蛋白质合成成本的工作。我们还推出了 GPT‑Rosalind,这是一个专为支持生命科学研究和药物发现流程而构建的模型。
本项目将这一轨迹延伸至药物化学领域,在该领域中,进展无法仅通过推理来衡量。一个假设必须在实验室中通过真实的分子、仪器和实验噪声来验证。我们与 Molecule.one(在新窗口中打开)合作,将 GPT‑5.4 连接到 Maria——一个与高通量实验室集成的自主研究化学 AI 代理——并赋予其一个开放式的目标:改进几个重要反应类别之一。该系统生成了研究提案,设计并运行了实验,分析了实验数据,并提出了后续实验。人类通过设计引导和评分提示、选择待测试的提案来保持参与。他们还对实验计划进行了有限的修正,协助了基本的实验室操作,并独立验证了最终结果。
最有前景的提案 OAI-M1-03 聚焦于 Chan–Lam 偶联反应的一个困难但有用的变体,这是化学家用于形成碳-氮键的反应。从改进 Chan–Lam 偶联用于工艺化学的开放式目标出发,GPT‑5.4 独立识别出伯磺酰胺是一个具有挑战性且高价值的底物类别,并建议使用包括 TEMPO 在内的温和氧化剂来改善反应。
在 Maria 实验室的两轮实验中,这一想法带来了显著改进。在优化条件下,测试的 88% 的硼酸和 83% 的磺酰胺的测量产率得到提高。平均产率从 16.6% 上升到 25.2%,产率超过 30% 的反应比例从 15.6% 增加到 37.5%。随后,人类化学家在实验室规模重复了代表性反应。这些实验证实了微升级别的结果,显示 14 个底物对中有 11 个产率更高,大多数情况下提高超过两倍。这一点很重要,因为药物化学家需要的反应不仅要在微升级筛选实验中有效,还要在药物发现过程中使用的实际实验室工作流程中可行。
药物化学这一领域的改进尤其令人兴奋,因为合成通常是药物发现中的主要瓶颈:科学家只能测试他们能够制造或以其他方式获得的分子。磺酰胺基团出现在涵盖广泛治疗领域的药物中,包括抗癌药、抗菌药和利尿剂,然而伯磺酰胺与硼酸的 Chan–Lam 偶联历来产率较低。使这种反应形式更可靠,可以为药物化学家提供更广泛、更实用的方法来生产和探索潜在有用的分子。
虽然这仍是一个早期结果,但它为我们正在努力的更广泛方向提供了另一个具体例证:AI 系统可以在研究循环的大部分环节中成为科学家的宝贵合作伙伴。该模型回顾了文献,提出了一个意想不到的想法,帮助设计和分析了实验,并得出了人类化学家可以评估的科学发现。
Maria 实验室:Molecule.one 的专业高通量实验室,在 OAI-M1-03 中运行了 10,080 个反应
为什么化学问题很重要
有机化学是所有小分子药物以及农业、电子和材料科学产品的基础。当一个反应能够在多种不同起始材料上可靠地形成同一种化学键时,它尤其有用。当反应产率低或产生过多不需要的副产物时,化学家可能不得不放弃原本有前景的分子,或花费大量时间开发不同的路线。这使得合成成为药物发现中的主要瓶颈:科学家通常只能测试他们能够制造或以其他方式获得的分子。
Chan–Lam 偶联在药物化学中很有用,因为它能形成药物中常见的碳-氮键。然而,该反应并非对所有分子类别都同样有效。特别是,伯磺酰胺与硼酸的偶联历来产率较低。磺酰胺是一类重要的分子,存在于用于肿瘤学和传染病的药物中。使这种反应更可靠,可以为药物化学家提供更广泛、更实用的方法来生产和探索潜在有用的分子。
将 GPT‑5.4 连接到 Maria AI 和实验室
组合系统配对互补能力。由与 Maria AI 合作的科学家编写的提示,在框架内与 GPT‑5.4 一起使用,以生成和排序数千个可能的研究提案。人类化学家审查了根据系统排名最高的一小部分提案,并选择了四个进行实验室测试。然后,Maria AI 将选定的高级计划转化为详细的实验室指令,运行了数千个高通量实验,分析了原始数据,并将结构化结果返回给 GPT‑5.4。
四个选定的提案之一 OAI-M1-03 建议使用温和氧化剂(如 TEMPO)来改善 Chan-Lam 反应在磺酰胺合成中的性能。化学家发现这个建议既令人惊讶又有趣。我们在本博客文章和论文(在新窗口中打开)中分享了 OAI-M1-03 的详细发现。我们还分享了 OAI-M1-03 的重写模型思维链(在新窗口中打开)。
最终的研究提案随后被 Maria 用于生成实验网格,人类进行了轻微修正。最大的修正是避免使用二甲基亚砜(DMSO)作为溶剂,因为化学家担心它可能与用作对照的更强氧化剂发生反应。
整个过程耗时三个月,从 3 月 4 日的首次提示到 6 月 4 日与独立专家分享 OAI-M1-03 的结果。
我们将这一工作流程描述为近乎自主,而非完全自主,因为人类化学家在整个过程中仍做出了重要决策。模型提出了关键研究思路,而人类化学家则提供高层指导与判断、修正实验细节、协助准备实验耗材与试剂,并手动重复关键实验。
我们的发现
OAI-M1-03 识别出 TEMPO 是本研究中所考察的伯磺酰胺 Chan-Lam 偶联反应的有效添加剂。在优化条件下,反应在两方面得到改善:平均产率提高,且更多底物组合达到了实际可用的产率。
在两个周期中,Maria 共运行了 10,080 个反应——这比一位化学家每天运行三个反应、持续十年所完成的反应还要多。这一规模之所以重要,是因为化学结果若仅在少数几个例子上测试,可能会产生误导。一个反应在一对起始原料上可能看起来很有前景,但在更广泛的分子集合中却可能失败。数千个反应使得在十种测试氧化剂中识别出 TEMPO、观察其效果在多种组合中的重复性,并发现其局限性成为可能。
在分析第一轮数据后,系统提出了更聚焦的第二轮实验,以检验后续假设。一个有用的后续发现是,TEMPO 可被一种更便宜的类似物 4-羟基-TEMPO 替代,且性能损失很小。
该结果在 Maria 实验室的微升级筛选规模之外同样成立。人类化学家手动在实验台规模上重复了代表性反应,观察到 14 对底物中有 11 对产率提高;其中 8 对的提高幅度超过两倍。这一重复验证之所以重要,是因为极小规模实验有时会引入在更大规模下消失的假象。在科研论文发表前,实验台规模验证也是惯例。

手动实验台规模验证中的反应瓶。
TEMPO 在实验台规模下改善产物生成
四位外部化学专家审阅了描述 OAI-M1-03 的预印本。他们的评估支持我们的观点,即该结果具有新颖性,值得与科学界分享。更严格的考验将来自下一步:独立实验室能否重复该结果,以及化学家是否发现其在更广泛的分子范围内有用。
在 GPT‑5.4 生成并由 Maria 在三个月期间测试的其他三个提案中,OAI-M1-02 和 OAI-M1-04 在 Maria 实验室得到实验验证,而 OAI-M1-01 被证伪。对这些结果的分析仍在进行中。
局限性
这项工作表明,模型可以在有机化学中做出有用贡献。它不仅仅是总结文献或建议一次性实验:它提出了一个具体的、令人惊讶的假设并提交给人类审查,设计了实验,解释了实验数据,并设计了后续实验。
但这并不表明 AI 可以独立地从头到尾运行一个化学研究项目。人类判断仍然至关重要,且该工作流程依赖于专业的高通量基础设施。这也不证明该方法将推广到其他偶联反应、其他底物类别或生产条件。
产率估算来自高通量平台,实验台验证覆盖了 14 对代表性底物。需要更多工作来表征反应机理、界定底物范围、测量不同实验室条件下的性能,并独立重复该结果。
准备情况
化学能力需要谨慎对待,因为支持医学和材料科学的同一工具也可能被滥用。我们特意将这项工作限定在一个合法的药物化学问题上:改进一种用于制造类药分子的已知偶联反应。实验不涉及毒素、化学武器或设计有害化合物的请求。这些结果不应被解读为该系统能够帮助实现那些有害应用的证据。该项目并未测试或证明这一点。
我们通过准备框架评估和缓解来自高级模型能力的新兴风险,包括与化学和生物领域相关的风险。本工作中使用的模型已经过英国 AI 安全研究所的相关评估,且该系统被设计为拒绝专注于有害应用的请求。实验工作流程增加了另一层控制:人类化学家选择哪些提案进入实验室,审查实验计划,并保留对物理基础设施的控制。
我们认为这是研究 AI 在实验化学中潜力的负责任方式:选择具有明确科学价值的问题空间,将模型级保障与专家监督相结合,并通过受约束的物理实验评估系统。随着这些能力的提升,我们将继续评估新兴风险,加强保障措施,并具体说明一个结果意味着什么以及不意味着什么。
下一步计划
接下来的直接步骤是科学性的:测试更广泛的起始原料,研究添加剂改善反应的原因,绘制效果有效和失效的范围,并支持独立重复。这些研究将共同决定该方法的应用广度及其在实用药物化学工作流程中的有用程度。
我们的长期目标是使 AI 系统成为可靠的科研伙伴,帮助研究人员生成假设、设计实验、解释结果并决定下一步测试什么,同时始终基于专家判断、可靠测量和强有力的保障。有机化学是一个特别高杠杆的领域,因为小分子发现和制造的进步依赖于可靠地制造分子的能力。科学家只能测试他们能制造的分子,而更好的合成可以扩展他们在医学、农业、电子、能源和材料科学领域探索的思路范围。这一结果是这一更广泛方向的一个早期示例:一个前沿模型、专门化智能体、自动化实验室和人类化学家协同工作,更快地推进研究循环,并产生科学界可以评估、重复和在此基础上发展的发现。
我们感谢 Molecule.one 团队以及审阅本工作的独立化学家。
OpenAI’s work in science is motivated by a simple belief: advanced AI can become a powerful partner for scientists, helping them explore more ideas, connect distant concepts, design better experiments, and accelerate discoveries that benefit humanity. We have already shared early examples of models contributing to novel results in mathematics, including work on the unit distance problem, in theoretical physics, through a new result on gluon amplitudes, and in biology, where GPT‑5 helped lower the cost of cell-free protein synthesis in an automated lab. We also introduced GPT‑Rosalind, a purpose-built model to support life sciences research and drug discovery workflows.
This project extends that trajectory into medicinal chemistry, where progress cannot be measured by reasoning alone. A hypothesis has to work in the lab with real molecules, instruments, and experimental noise. Working with Molecule.one(opens in a new window), we connected GPT‑5.4 to Maria—an agentic chemistry AI integrated with a high-throughput laboratory for autonomous research—and gave it an open-ended goal: to improve one of several important reaction classes. The system generated research proposals, designed and ran experiments, analyzed experimental data, and proposed follow-up experiments. Humans remained in the loop by designing steering and grading prompts and selecting proposals to test. They also made limited corrections to experimental plans, assisted with basic laboratory operations, and independently validated the final result.
The most promising proposal, OAI-M1-03, focused on a difficult but useful version of Chan–Lam coupling, a reaction chemists use to form carbon-nitrogen bonds. Starting from the open-ended goal of improving Chan–Lam coupling for process chemistry, GPT‑5.4 independently identified primary sulfonamides as a challenging, high-value substrate class and suggested that mild oxidants, including TEMPO, could improve the reaction.
Across two cycles of experimentation in Maria Lab that idea produced a significant improvement. Under the optimized conditions, measured yields improved for 88% of the boronic acids and 83% of the sulfonamides tested. The mean yield rose from 16.6% to 25.2%, and the share of reactions above 30% yield increased from 15.6% to 37.5%. Human chemists then repeated representative reactions at bench scale. Those experiments confirmed the microliter-scale results, showing higher yields for 11 of 14 substrate pairs, with a more than twofold increase in most cases. That matters because medicinal chemists need reactions that work not just in micro-liter screening experiments, but also in practical lab workflows used during drug discovery.
Improvements in this area of medicinal chemistry are particularly exciting because synthesis is often a major bottleneck in drug discovery: scientists can only test the molecules they can make or otherwise obtain. The sulfonamide group appears in medicines across a wide range of therapeutic areas, including anticancer drugs, antimicrobials, and diuretics, yet the Chan–Lam coupling of primary sulfonamides with boronic acids has historically given low yields. Making this form of the reaction more reliable could give medicinal chemists a broader and more practical way to produce and explore potentially useful molecules.
While this is still an early result, it provides another concrete example of the broader direction we are working toward: AI systems that can become valuable partners to scientists across much of the research loop. The model reviewed the literature, proposed an unexpected idea, helped design and analyze experiments, and arrived at a scientific finding that human chemists could evaluate.
Maria Lab: Molecule.one's specialized high-throughput laboratory that ran 10,080 reactions in OAI-M1-03
Why the chemistry problem matters
Organic chemistry underpins all small-molecule medicines, as well as products in agriculture, electronics, and materials science. A reaction is especially useful when it can make the same kind of chemical bond reliably across many different starting materials. When reactions produce low yields or too many unwanted byproducts, chemists may have to abandon otherwise promising molecules or spend significant time developing a different route. This makes synthesis a major bottleneck in drug discovery: scientists can generally only test the molecules they can make or otherwise obtain.
Chan–Lam coupling is useful in medicinal chemistry because it forms carbon-nitrogen bonds, which are common in medicines. However, the reaction does not work equally well for every class of molecule. In particular, coupling primary sulfonamides with boronic acids has historically produced low yields. Sulfonamides are an important family of molecules found in medicines used in oncology and infectious disease. Making this reaction more reliable could give medicinal chemists a broader and more practical way to produce and explore potentially useful molecules.
Connecting GPT‑5.4 to Maria AI and Lab
The combined system paired complementary capabilities. Prompts written by scientists working with Maria AI were used with GPT‑5.4 within a harness to generate and rank thousands of possible research proposals. Human chemists reviewed the small subset of proposals that ranked highest according to the system and selected four for laboratory testing. Maria AI then translated selected high-level plans into detailed lab instructions, ran thousands of high-throughput experiments, analyzed the raw data, and returned structured results to GPT‑5.4.
One of the four selected proposals, OAI-M1-03, suggested using mild oxidants such as TEMPO to improve the performance of the Chan-Lam reaction for sulfonamide synthesis. Chemists found the suggestion both surprising and interesting. We share the detailed findings from OAI-M1-03 in this blog post and in the paper(opens in a new window). We’re also sharing the rewritten model chain-of-thought(opens in a new window) for OAI-M1-03.
The final research proposal was then used by Maria to generate experimental grids, with slight corrections by humans. The largest human correction was to avoid dimethyl sulfoxide, or DMSO, as a solvent because chemists were concerned it could react with the stronger oxidants used as comparisons.
The full process took three months, from the first prompt on March 4th to sharing the OAI-M1-03 results with independent experts on June 4th.
We describe this workflow as near-autonomous, not fully autonomous, because human chemists still made important decisions throughout the process. The model proposed the key research ideas, while human chemists provided high-level steering and judgment, corrected experimental details, helped prepare lab consumables and reagents, and repeated key experiments by hand.
What we found
OAI-M1-03 identified TEMPO as a useful additive for the primary sulfonamide Chan-Lam coupling studied here. Under the optimized conditions, the reaction improved in two ways: average yield went up, and more substrate combinations reached practically useful yields.
Across two cycles, Maria ran a total of 10,080 reactions – more than a chemist running three reactions every day would run in a decade. That scale mattered because chemistry results can be misleading when they are tested on only a few examples. A reaction can look promising on one pair of starting materials, but fail across a broader set of molecules. Thousands of reactions made it possible to identify TEMPO among ten tested oxidants, see the effect repeat across diverse combinations, and find its limitations.
After analyzing the first round of data, the system proposed a more focused second round of experiments to test follow-up hypotheses. One useful follow-up finding was that TEMPO could be replaced by a much cheaper analog, 4-hydroxy-TEMPO, with little loss in performance.
The result also held up beyond Maria Lab’s microliter-scale screening format. Human chemists reproduced representative reactions manually at bench scale and observed an increase in yield for 11 of 14 substrate pairs; for eight pairs the increase was greater than twofold. That replication matters because very small-scale experiments can sometimes introduce artifacts that disappear at a larger scale. Bench-scale validation is also customary before research is published in a scientific journal.

Reaction vials from the manual bench-scale validation.
TEMPO improves product formation at bench scale
Four external chemistry experts reviewed the preprint describing OAI-M1-03. Their assessments supported our view that the result was novel and worth sharing with the scientific community. The stronger test will come next: whether independent labs can reproduce the result, and whether chemists find it useful across a broader range of molecules.
Of the other three proposals generated by GPT‑5.4 and tested by Maria during the three-month period, OAI-M1-02 and OAI-M1-04 were experimentally proven in the Maria Lab, while OAI-M1-01 was disproven. Analysis of these results is ongoing.
Limitations
This work shows that a model can make a useful contribution in organic chemistry. It did more than summarize the literature or suggest a one-off experiment: it proposed a specific surprising hypothesis and surfaced it for human review, designed experiments, interpreted experimental data, and designed follow-up experiments.
It does not show that AI can independently run a chemistry research program from end to end. Human judgment remained essential, and the workflow depended on specialized high-throughput infrastructure. It also does not establish that the method will generalize to other coupling reactions, other substrate classes, or manufacturing conditions.
The yield estimates came from a high-throughput platform, and bench validation covered 14 representative substrate pairs. More work is needed to characterize the reaction mechanism, define the substrate scope, measure performance under different laboratory conditions, and reproduce the result independently.
Preparedness
Chemistry capabilities require careful treatment because the same tools that can support medicine and materials science could also be misused. We deliberately scoped this work to a legitimate medicinal-chemistry problem: improving a known coupling reaction used to make drug-like molecules. The experiments did not involve toxins, chemical weapons, or requests to design harmful compounds. These results should not be read as evidence that the system can help with those harmful applications. The project did not test or demonstrate that.
We assess and mitigate emerging risks from advanced model capabilities through our Preparedness Framework, including risks related to chemical and biological domains. The model used in this work had already undergone relevant evaluations with the UK AI Security Institute, and the system was designed to refuse requests focused on harmful applications. The experimental workflow added another layer of control: human chemists selected which proposals entered the lab, reviewed experimental plans, and retained control of the physical infrastructure.
We think this is the responsible way to study AI's potential in experimental chemistry: choose a problem space with clear scientific value, pair model-level safeguards with expert oversight, and evaluate the system through constrained physical experiments. As these capabilities improve, we will continue to assess emerging risks, strengthen safeguards, and be specific about what a result does and does not imply.
What's next
The immediate next steps are scientific: test a broader range of starting materials, investigate why the additives improve the reaction, map where the effect works and fails, and support independent replication. Together, these studies will determine how broadly the method can be applied and how useful it is in practical medicinal chemistry workflows.
Our longer-term goal is to make AI systems reliable scientific partners that help researchers generate hypotheses, design experiments, interpret results, and decide what to test next, while remaining grounded in expert judgment, reliable measurement, and strong safeguards. Organic chemistry is a particularly high-leverage area because progress in small-molecule discovery and manufacturing depends on being able to make molecules reliably. Scientists can only test molecules they can make, and better synthesis can expand the range of ideas they can explore across medicine, agriculture, electronics, energy, and materials science. This result is one early example of that broader direction: a frontier model, specialized agents, an automated laboratory, and human chemists working together to move faster through the research loop and produce findings the scientific community can evaluate, reproduce, and build on.
We are grateful to the Molecule.one team and to the independent chemists who reviewed this work.
本文内容采集自官方网站,排版和翻译可能与原页面存在差异。
阅读官方全文