代理型AI系统在执行科学任务方面的能力日益增强。然而,它们对生命科学研究者的实用性,取决于其处理真实研究复杂性的能力。真实研究工作很少表现为单一的事实回忆问题或清晰的预测问题。研究者需要解读不完整的证据、调和相互矛盾的结果、设计困难的实验、排除检测故障、评估转化风险,并在不确定情况下决定下一步行动。
当前的基准测试未能完全捕捉这些能力。许多生命科学评估聚焦于狭窄领域或孤立技能,导致问题采用结构化格式并配有清晰的参考答案。虽然这些评估有价值,但它们往往无法真正衡量模型能否在更广泛的研究级工作中做出贡献。
我们设计了LifeSciBench来帮助弥合这一差距。每项任务都基于具有博士水平训练、且在生物技术和制药领域拥有推进药物发现项目直接经验的执业生命科学家的判断。
LifeSciBench包含750个由专家撰写的任务,涵盖七个工作流程和七个生物学领域。
1,062 任务工件
173 科学家贡献者
19,020 评分标准
453 专家评审员
LifeSciBench衡量什么
LifeSciBench衡量AI系统能否支持现实的生命科学研究任务,而不仅仅是回答生物学问题。为定义基准分类法,我们调查了执业生命科学家在应用研究环境中最常使用的工作流程。然后,我们将他们的回答归纳为七个反复出现的类别:证据处理、分析、设计与优化、科学推理、验证与操作、转化以及科学沟通。
每个任务的结构都类似于科学家可能向知识渊博的合作者提出的请求:科学提示、任何相关背景或工件,以及自由回答。专家撰写的评分标准评估模型能否针对特定问题给出正确答案,并具备科学家所期望的适当详细程度、论证、注意事项和格式。
数据集构建
LifeSciBench评估科学推理能力,同时兼顾现实科学应用所需的不那么明确但实用的技能。其任务要求模型处理真实的研究问题:解读证据、做出基于领域的判断,并传达对专家评审员有用的结论。许多任务还要求模型处理不确定性,并基于支持性数据文件进行推理,而不仅仅依赖提示文本。
该基准旨在反映生命科学工作的复杂性。总体而言,79%的任务需要多个推理或决策步骤,平均每个任务包含四个步骤。LifeSciBench包含1,062个附加工件,涵盖图表、PDF、表格、序列文件、结构或化学文件以及网络参考文献。超过一半的任务(53%)要求模型解读或综合至少一个工件中的信息。
任务由173位来自不同生命科学学科的专家科学家创建。每位科学家均具有博士水平训练以及生物技术或制药行业经验。任务在验收前可根据需要经历任意轮次的修订,无固定上限;已验收任务平均经历六轮自主自动化审查周期,并完成至少两轮专家评审。评审基于可验证的正确答案或强烈的专家共识,相关领域评审员之间的一致性至少达到90%。这一流程有助于确保已验收任务具有科学依据、清晰可评分,并能代表应用研究。
评分与评分标准分解
LifeSciBench任务使用详细的、针对具体任务的评分标准进行评分,该标准将预期回答分解为具体的科学主张、计算、决策、论证等。在整个基准中,专家制定的评分标准包含19,020个标准——平均每个任务25个——用于评估科学正确性以及对研究决策的实用性。
这一设计反映了科学工作在实际中的评估方式:许多生命科学任务无法仅通过检查最终答案来评分。一个回答可能得出正确的高层结论,但如果它忽略了关键检测限制或未能主动提出高度重要的生物学细微差别,仍可能被视为不完整。反之,部分回答即使未完全解决任务,也可能包含高质量的推理。
细粒度的评分标准捕捉了这种细微差别。LifeSciBench不仅评估最终答案的准确性,还评估模型是否以科学有效且操作实用的方式得出答案。
从论文、图表、表格和实验记录中提取、核对和审计科学证据。
候选回答
结论: 就目前呈现的情况而言,该方案不足以支持基于微型肌营养不良蛋白表达作为合理可能预测临床获益的替代终点而加速批准。持怀疑态度的FDA评审员可能会认为生物标志物、功能性、持久性、安全性和普适性证据在实质上缺乏充分支持。
| 项目内容 | 关键失效模式 | 所需改进 |
|---|---|---|
| 蛋白质印迹定量 | MANEX1A抗体结合内源性全长肌营养不良蛋白与微型肌营养不良蛋白转基因共有的N端表位,因此检测无法清晰区分转基因产物与残留/回复突变肌营养不良蛋白。使用健康全长肌营养不良蛋白标准品定量138 kDa微型肌营养不良蛋白同样无效。 | 使用重组微型肌营养不良蛋白标准品,并采用能区分转基因产物与内源性肌营养不良蛋白的正交方法,如靶向质谱法或转基因特异性/表位特异性检测。 |
| 免疫荧光 | C端多克隆抗体不适用,因为138 kDa构建体缺乏C端结构域。许多DMD患者存在回复突变肌纤维,且回复突变肌营养不良蛋白可能保留C端表位。回复突变肌纤维可能随年龄增长发生克隆性扩增,从而干扰免疫荧光信号,尤其在年龄较大的男孩中。 | 使用针对转基因产物中存在但回复突变肌营养不良蛋白中缺失的表位的抗体重复免疫荧光实验。将转基因阳性肌纤维与回复突变肌纤维分别定量。 |
| 替代终点有效性 | 该方案将蛋白量与临床功能混为一谈。“健康对照蛋白质量的38%”并不意味着正常肌营养不良蛋白功能的38%,因为微型肌营养不良蛋白在结构上被截短。 | 在将表达量视为替代终点前,需通过经验验证微型肌营养不良蛋白质量百分比、肌膜定位、下游功能恢复与临床获益之间的关系。 |
| 活检设计 | 治疗前后对侧股外侧肌活检引入了左右侧及肌肉内空间变异性。疾病进展和纤维脂肪替代也可能改变总蛋白归一化信号。 | 使用一致的解剖标志标准化活检部位,归一化至肌肉特异性蛋白,并同步测量纤维脂肪成分。 |
| NSAA比较/统计 | 外部自然病史队列并非随机同期对照。试验入组条件、支持性护理、参与效应、基线NSAA、类固醇方案、年龄和外显子类别均可能使比较产生偏倚。配对t检验不足以解决问题。此外,+1.4的NSAA变化在该年龄组的测试-重测变异范围内。 | 开展随机同期安慰剂对照研究,或至少采用校正基线NSAA、年龄、类固醇方案、外显子类别及其他混杂因素的调整分析。 |
| 年龄窗口混杂 | 4-7岁男孩处于发育窗口期,未治疗的可步行DMD患者可能在功能衰退前获得运动功能提升。48周NSAA变化混合了发育增益、疾病进展及可能的治疗效果。 | 使用按年龄分层的随机同期对照,以区分发育轨迹与治疗效果。 |
| 既往临床先例 | 开放标签微型肌营养不良蛋白的功能信号未能可靠预测确证性获益;已发表的先例包括微型肌营养不良蛋白基因治疗的确证性试验未能重现开放标签的NSAA改善。 | 不依赖开放标签NSAA变化作为决定性支持证据。需提供对照的功能性证据。 |
| 构建体的结构限制 | 138 kDa构建体删除了含有nNOS结合位点的血影蛋白重复序列R16/17。nNOS招募缺失可能损害运动中的功能性交感神经舒张和缺血保护,从而在表达水平之外形成机制性功能恢复上限。 | 增加机制研究,验证该特定构建体是否能恢复相关肌营养不良蛋白复合体功能、nNOS定位、运动生理学及肌肉保护。 |
| AAV持久性 | 12周的载体基因组无法证明表达的持久性。AAV9基因组主要为非整合型附加体,可能随时间推移而减少。载体基因组持久性不等同于蛋白表达的持久性。 | 测量超过12周的纵向转基因蛋白表达及功能性生物标志物的持久性。 |
| 免疫/安全性特征 | 8/12患者出现转氨酶升高,与对AAV转导细胞的免疫反应一致,但机制尚未明确。鉴于AAV9的心脏趋向性,一例心肌炎病例值得关注。 | 提供更深入的免疫监测、肝脏/心脏安全性表征,并加强心脏随访。 |
| 患者选择/普适性 | 排除抗AAV9中和抗体阳性患者限制了普适性。排除外显子44缺失患者则限制了对该DMD亚组的适用性。n=12样本量过小,无法在更广泛的DMD人群中表征安全性和有效性。 | 在可能的情况下放宽入组条件,或在利用该结果支持广泛批准前,预先按抗体状态、基因型/外显子类别、年龄和基线功能进行分层分析。 |
监管结论: 该方案可能显示生物学活性,但尚未证明所测得的微型肌营养不良蛋白表达是合理可能预测临床获益的可靠替代终点。主要缺陷包括检测特异性不足、无效的定量标准、可能的回复突变肌纤维干扰、缺乏随机对照、年龄相关的NSAA混杂、持久性不确定以及未解决的安全性和普适性问题。
为弥补这些缺陷,该计划需要采用受控的、按年龄分层的临床设计,包括转基因特异性表达检测、正交蛋白定量、组织成分对照、纵向持久性数据、针对截短构建体的机制性功能检测,以及更强的安全性监测(尤其是肝脏和心脏)。
评分标准与等级
| 标准 | 分数 |
|---|---|
| 识别微型肌营养不良蛋白定量中的检测/测量问题,包括MANEX1A表位共享、无效的全长肌营养不良蛋白标准品,以及需要重组或正交转基因特异性测量。 | +24 |
| 解释为何微型肌营养不良蛋白表达水平不能自动作为功能性临床获益的有效替代终点。 | +22 |
| 指出活检部位、组织成分和年龄窗口混杂因素削弱了表达量和NSAA的解释。 | +19 |
| 批评NSAA比较/统计方法,尤其是依赖外部自然病史对照。 | +12 |
| 涉及AAV持久性、免疫反应、转氨酶升高、心肌炎,以及需要更长期的表达/安全性随访。 | +15 |
| 指出患者选择/普适性缺陷,包括抗AAV9排除、外显子44排除及小样本量。 | +8 |
验证LifeSciBench
我们通过独立专家评审验证了LifeSciBench。反馈来自453位未参与任务撰写的评审员。其中,97%拥有博士学位或同等学历,平均具有12年领域经验和14篇同行评审论文;88%报告曾获得至少一项奖项或奖学金。
评审员根据每项任务是否具备强基准问题所需的特质进行评分:与现实研究工作的契合度、对科学推理和领域专业知识的恰当测试、基于证据或专家共识的可靠性,以及评估模型性能的整体实用性。每个类别的评分一致性均超过96%。
评审员评论进一步强化了量化评分:
1/3
结果
我们报告两个互补指标。通过率指模型达到任务级成功阈值70%的任务百分比。得分指平均评分奖励,即使未完全解决任务,也会对单个标准给予部分分数。两者都很重要,因为对科学任务的回答可能部分正确或有用,但未必满足完整答案的所有要求。
模型性能因任务类型、工作流程和响应格式而显著不同。
AI系统展现早期优势的领域
LifeSciBench显示,前沿模型在涉及科学综合、沟通和结构化解释的任务中表现相对最强。绝对通过率仍然不高,因此这些基准领域远未饱和,但GPT‑Rosalind相比GPT‑5.5取得了有意义的进步,整体精确通过率从25.7%提升至36.1%。
模型能力进步最显著的方向出现在科学沟通与转化领域。例如,科学沟通通过率从GPT‑5.5的56.3%提升至GPT‑Rosalind的71.1%;该类别样本量较小(n=9),需谨慎解读,但这表明前沿模型在组织证据和生成令人信服的专家级解释方面能力快速提升。转化(药物开发的"从实验室到临床"过程)呈现类似模式,从GPT‑5.5的36.8%提升至GPT‑Rosalind的57.7%,表明模型在连接临床前证据与临床意义方面的能力快速提升。
评分级结果指向相同方向。在需要专家有用或可操作输出的任务上,GPT‑Rosalind得分为44.7%,而GPT‑5.5为29.1%。在需要处理不确定性和注意事项的任务上,其得分为44.8%,而GPT‑5.5为29.3%。这一模式表明,当任务具有明确的证据边界并要求结构化科学判断时,模型最为有用。
GPT‑Rosalind在行业和学术专家认定的科学价值任务上领先性能。
GPT‑Rosalind在核心生命科学工作流程上相比GPT‑5.5性能提升,转化和科学沟通领域进步最大。
AI系统仍显不足的领域
在依赖人工制品、设计密集型以及操作受限的科学工作中,性能仍然较弱。具体而言,设计、优化与预测仍是最困难的工作流程之一,GPT‑Rosalind通过率为30.7%;分析同样困难,通过率为30.3%。
人工制品使用是一个特别明显的差距。虽然GPT‑Rosalind在依赖人工制品场景下表现优于GPT‑5.5,但其通过率从纯文本任务的45.1%下降至涉及人工制品或URL任务的28.1%。GPT‑5.5呈现相同模式,从29.9%下降至21.9%。更详细的分析证实,前沿模型在从复杂图表或大型序列文件中提取信息并将其整合到最终答案方面存在困难。
当任务需要基于来源的推理或处理人工制品时,通过率下降
答案格式也很重要。需要精确序列、结构或构建级输出的任务通过率较低:GPT‑Rosalind在数值任务上仅达14.8%,在序列或结构输出上为24.0%。构建生成任务同样脆弱,GPT‑Rosalind为27.3%,相比GPT‑5.5改进甚微。部分差距可能反映了精确答案任务更严格的评分标准,计算或格式上的微小差异可能导致响应低于通过阈值。然而,这些失败具有科学意义,因为许多生命科学工作流程需要足够精确以直接使用的输出,例如CRISPR/HDR供体设计或siRNA设计。
模型也常常部分完成任务但未完全解决。约14%的任务中,模型尽管未达到精确通过阈值,却获得了大量评分奖励。对于GPT‑Rosalind,109个任务的通过率低于20%,但仍获得至少50%的评分奖励。实践中,这意味着模型可能识别相关证据或生成看似合理的部分答案,但仍因遗漏关键约束、使用错误证据、计算不完整或未将推理与科学有用的最终决策联系起来而失败。
局限性与未来方向
LifeSciBench是衡量AI系统对生命科学研究有用性的一个步骤,但不能替代在真实研究环境中研究模型。该基准聚焦于反映重复性行业工作流程的独立任务,同时将许多科学专业和任务类型排除在当前范围之外。真实研究是迭代的:科学家收集新证据、修正假设、设计后续实验,并根据结果调整计划。
因此,LifeSciBench上的强表现应被解读为现实任务级能力的证据,而非下游研究影响的直接衡量。该基准基于行业工作流程,但未捕捉真实研究项目的全部多样性或动态性,其中进展取决于随时间展开的因素。
下一步是将基准性能与真实研究工作流程中的部署研究联系起来。虽然LifeSciBench是与实践科学家共同开发的,但衡量AI系统是否加速发现或改善研发成果,将需要在真实研究环境中、更长时间跨度内、以及多轮推理、反馈和实验跟进中研究模型的使用和性能。
Agentic AI systems are becoming increasingly capable of performing scientific tasks. However, their usefulness to life science researchers depends on how well they handle the complexity of real research. That work rarely looks like a single fact-recall question or a clean prediction problem. Researchers interpret incomplete evidence, reconcile conflicting results, design difficult experiments, troubleshoot assays, evaluate translational risk, and decide what to do next under uncertainty.
Current benchmarks do not fully capture these capabilities. Many life science evaluations focus on narrow domains or isolated skills, resulting in questions with structured question formats and clean reference answers. While valuable, they often fail to truly assess whether a model can contribute across the broader span of research-level work.
We designed LifeSciBench to help close this gap. Every task is grounded in the judgment of practicing life scientists with Ph.D.-level training and direct experience advancing drug discovery programs in biotech and pharmaceutical settings.
LifeSciBench includes 750 expert-authored tasks spanning seven workflows and seven biological domains.
1,062
Task artifacts
173
Scientist contributors
19,020
Rubric criteria
453
Expert reviewers
What LifeSciBench measures
LifeSciBench measures whether AI systems can support realistic life science research tasks, not just answer biology questions. To define the benchmark taxonomy, we surveyed practicing life scientists about the workflows they use most often in applied research settings. Then, we grouped their responses into seven recurring categories: evidence handling, analysis, design and optimization, scientific reasoning, validation and operations, translation, and scientific communication.
Each task is structured like a request a scientist might give to a knowledgeable collaborator: scientific prompt, any relevant context or artifacts, and a free-response answer. Expert-written rubrics evaluate whether a model can produce the right answer for a specific problem, with the right level of detail, justification, caveats, and formatting a scientist would expect.
Dataset construction
LifeSciBench evaluates scientific reasoning alongside the less well-defined, practical skills necessary for real-world scientific use. Its tasks ask models to work through realistic research problems: interpreting evidence, making domain-grounded judgments, and communicating conclusions that would be useful to expert reviewers. Many tasks also require models to handle uncertainty and reason over supporting data files rather than relying on prompt text alone.
The benchmark is designed to reflect the complexity of life science work. Overall, 79% of tasks require multiple reasoning or decision-making steps, with an average of four steps per task. LifeSciBench includes 1,062 attached artifacts spanning figures, PDFs, tables, sequence files, structure or chemical files, and web references. More than half of tasks (53%) require models to interpret or synthesize information from at least one artifact.
Tasks were created by 173 expert scientists across different life science disciplines. Each scientist had Ph.D.-level training and biotechnology or pharmaceutical industry experience. Tasks could undergo as many revision cycles as needed before acceptance, with no fixed cap on the number of rounds; accepted tasks averaged six self-directed automated review cycles and completed at least two rounds of expert reviews. Reviews were anchored in either a verifiable correct answer or strong expert consensus, with at least 90% agreement among reviewers in the relevant domain. This process helped ensure that accepted tasks were scientifically grounded, clear enough to grade, and representative of applied research.
Grading and rubric breakdown
LifeSciBench tasks are graded with a detailed, task-specific rubric that breaks down the expected response into specific scientific claims, calculations, decisions, justifications, and so on. Across the benchmark, expert-developed rubrics include 19,020 criteria—an average of 25 per task—to assess both scientific correctness and usefulness for research decisions.
This design reflects how scientific work is evaluated in practice: many life science tasks cannot be graded by checking the final answer alone. A response may reach the correct high-level conclusion but still be judged incomplete if, for example, it overlooks a key assay limitation or fails to proactively bring up a highly consequential biological nuance. Conversely, a partial response may contain high-quality reasoning even if it does not fully solve the task.
The granular rubrics capture this nuance. LifeSciBench evaluates not only final-answer accuracy, but whether a model reaches its answer in a scientifically valid and operationally useful way.
Extracting, reconciling, and auditing scientific evidence from papers, figures, tables, and experimental records.
Candidate Response
Bottom line: As presented, this package is not strong enough to support accelerated approval on micro-dystrophin expression as a surrogate endpoint reasonably likely to predict clinical benefit. A skeptical FDA reviewer would likely view the biomarker, functional, durability, safety, and generalizability evidence as materially under-supported.
| Package Item | Key Failure Mode | What Would Be Needed |
|---|---|---|
| Western blot quantification | MANEX1A binds an N-terminal epitope shared by endogenous full-length dystrophin and the micro-dystrophin transgene, so the assay does not cleanly distinguish transgene from residual/revertant dystrophin. Quantifying a 138 kDa micro-dystrophin against a healthy full-length dystrophin standard is also invalid. | Use a recombinant micro-dystrophin standard and an orthogonal method that distinguishes transgene from endogenous dystrophin, such as targeted mass spectrometry or a transgene-specific/epitope-specific assay. |
| Immunofluorescence | The C-terminal polyclonal antibody is poorly suited because the 138 kDa construct lacks the C-terminal domain. Many DMD patients have revertant fibers, and revertant dystrophin can retain C-terminal epitopes. Revertant fibers may expand clonally with age, biasing IF signal, especially in older boys. | Repeat IF with an antibody against an epitope present in the transgene but absent from revertant dystrophin. Quantify transgene-positive fibers separately from revertant fibers. |
| Surrogate endpoint validity | The package conflates protein amount with clinical function. “38% of healthy-control protein mass” does not mean 38% of normal dystrophin function because micro-dystrophin is structurally truncated. | Empirically validate the relationship between micro-dystrophin mass-percent, sarcolemmal localization, downstream functional restoration, and clinical benefit before treating expression as a surrogate endpoint. |
| Biopsy design | Pre- and post-treatment contralateral vastus lateralis biopsies introduce left-right and intramuscular spatial variability. Disease progression and fibro-fatty replacement can also change total-protein-normalized signal. | Standardize biopsy site using consistent anatomical landmarks, normalize to muscle-specific proteins, and measure fibro-fatty composition in parallel. |
| NSAA comparator/statistics | An external natural-history cohort is not a randomized concurrent control. Trial eligibility, supportive care, participation effects, baseline NSAA, steroid regimen, age, and exon class can all bias the comparison. An unpaired t-test is not sufficient. Also, a +1.4 NSAA change is within test-retest variability for this age group. | Run a randomized concurrent placebo-controlled study, or at minimum use adjusted analyses accounting for baseline NSAA, age, steroid regimen, exon class, and other confounders. |
| Age-window confounding | Boys age 4–7 are in a developmental window where untreated ambulatory DMD patients may gain motor function before decline dominates. A 48-week NSAA change mixes developmental gain, disease progression, and possible treatment effect. | Use a concurrent randomized control with age stratification to separate developmental trajectory from treatment effect. |
| Prior clinical precedent | Open-label micro-dystrophin functional signals have not reliably predicted confirmatory benefit; published precedent includes micro-dystrophin gene therapy confirmatory trials failing to reproduce open-label NSAA improvements. | Do not rely on open-label NSAA change as decisive support. Require controlled functional evidence. |
| Structural limits of the construct | The 138 kDa construct deletes spectrin repeats R16/17, which contain nNOS-binding sites. Loss of nNOS recruitment can impair functional sympatholysis and ischemia protection during exercise, creating a mechanistic ceiling on rescue independent of expression level. | Add mechanistic studies showing whether this specific construct restores relevant dystrophin-associated complex function, nNOS localization, exercise physiology, and muscle protection. |
| AAV durability | Vector genomes at 12 weeks do not establish durable expression. AAV9 genomes are largely non-integrating episomes and may decline over time. Vector-genome persistence is not the same as persistent protein expression. | Measure longitudinal transgene protein expression and functional biomarker durability beyond 12 weeks. |
| Immune/safety profile | Transaminitis in 8/12 patients is consistent with immune response to AAV-transduced cells, but the mechanism is not established. One myocarditis case is concerning given AAV9 cardiac tropism. | Provide deeper immune monitoring, liver/cardiac safety characterization, and intensified cardiac follow-up. |
| Patient selection/generalizability | Excluding anti-AAV9 neutralizing-antibody-positive patients limits generalizability. Excluding exon-44 deletions limits applicability to that DMD subgroup. n=12 is too small to characterize safety and efficacy across the broader DMD population. | Broaden eligibility where possible or pre-specify stratified analyses by antibody status, genotype/exon class, age, and baseline function before using the result to support broad approval. |
Regulatory conclusion: The package may show biological activity, but it does not yet establish that the measured micro-dystrophin expression is a reliable surrogate reasonably likely to predict clinical benefit. The main gaps are assay specificity, invalid quantification standards, possible revertant-fiber confounding, lack of a randomized control, age-related NSAA confounding, uncertain durability, and unresolved safety/generalizability issues.
To close the gap, the program would need a controlled, age-stratified clinical design with transgene-specific expression assays, orthogonal protein quantification, tissue-composition controls, longitudinal durability data, mechanistic functional assays for the truncated construct, and stronger safety monitoring, especially hepatic and cardiac.
Rubric Criteria & Grades
Criterion
Points
Identifies assay/measurement problems in micro-dystrophin quantification, including MANEX1A epitope sharing, invalid full-length dystrophin standards, and need for recombinant or orthogonal transgene-specific measurement.
+24
Explains why micro-dystrophin expression level is not automatically a valid surrogate for functional clinical benefit.
+22
Flags biopsy-site, tissue-composition, and age-window confounding that weaken expression and NSAA interpretation.
+19
Critiques the NSAA comparator/statistics, especially reliance on external natural-history controls.
+12
Addresses AAV durability, immune response, transaminitis, myocarditis, and need for longer-term expression/safety follow-up.
+15
Notes patient-selection/generalizability gaps, including anti-AAV9 exclusion, exon-44 exclusion, and small sample size.
+8
Validating LifeSciBench
We validated LifeSciBench through an independent expert review. Feedback came from 453 reviewers who were not involved in writing the tasks. Of those reviewers, 97% held a Ph.D. or equivalent doctorate, with an average of 12 years of field experience and 14 peer-reviewed publications; 88% reported receiving at least one award or fellowship.
Reviewers scored whether each task reflected the qualities needed for a strong benchmark question: alignment with real-world research work, appropriate testing of scientific reasoning and domain expertise, grounding in evidence or expert consensus, and overall usefulness for assessing model performance. Agreement exceeded 96% in every category.
Reviewer comments reinforced the quantitative ratings:
1 of 3
Results
We report two complementary metrics. Pass rate is the percentage of tasks on which a model meets the task-level success threshold of 70%. Score is the average rubric reward, giving partial credit for individual criteria even when the full task is not solved. Both matter because a response to a scientific task can be partially correct or useful without meeting every requirement for a complete answer.
Model performance varies substantially by task type, workflow, and response format.
Where AI systems show early strength
LifeSciBench shows that frontier models are relatively strongest on tasks involving scientific synthesis, communication, and structured interpretation. Absolute pass rates are still modest, so these benchmark domains are far from saturated, but GPT‑Rosalind shows meaningful progress over GPT‑5.5, improving overall exact pass rate from 25.7% to 36.1%.
The strongest directions of progression in model capabilities appear in Scientific Communication and Translation. For example, the Scientific Communication pass rate increases from 56.3% for GPT‑5.5 to 71.1% for GPT‑Rosalind; this category is small (n=9), so it should be interpreted cautiously, but it suggests frontier models are improving rapidly in their ability to organize evidence and produce convincing expert-facing explanations. Translation (the "bench-to-bedside" process of drug development) shows a similar pattern, rising from 36.8% for GPT‑5.5 to 57.7% for GPT‑Rosalind, suggesting models are quickly improving on their ability to connect preclinical evidence to clinical implications.
Rubric-level results point in the same direction. On tasks requiring expert-useful or actionable outputs, GPT‑Rosalind scores 44.7%, compared with 29.1% for GPT‑5.5. On tasks requiring uncertainty and caveat handling, it scores 44.8%, compared with 29.3%. This pattern suggests models are most useful when the task has a clear evidence boundary and calls for structured scientific judgment.
GPT‑Rosalind leads performance across scientifically-valuable tasks identified by industry and academic experts.
GPT‑Rosalind improves performance over GPT‑5.5 across core life-science workflows, with the strongest gains in translation and scientific communication.
Where AI systems still fall short
Performance remains much weaker on artifact-heavy, design-heavy, and operationally constrained scientific work. Namely, Design, Optimization, & Prediction remains one of the hardest workflows, with GPT‑Rosalind passrate at 30.7%; Analysis is similarly difficult at 30.3%.
Artifact use is a particularly clear gap. While GPT‑Rosalind performs better than GPT‑5.5 in artifact-heavy settings, its pass rate still drops from 45.1% on text-only tasks to 28.1% on tasks with artifacts or URLs. GPT‑5.5 shows the same pattern, dropping from 29.9% to 21.9%. A more detailed analysis confirms that frontier models struggle at extracting information from complex figures or large sequence files and integrating that information into the final answer.
Pass rates drop when tasks require source-grounded reasoning or working with artifacts
The answer format also matters. Tasks requiring exact sequence, structure, or construct-level outputs show lower pass rates: GPT‑Rosalind reaches only 14.8% on numeric tasks and 24.0% on sequence or structure outputs. Construct-generation tasks are also brittle, with GPT‑Rosalind at 27.3% and showing little improvement over GPT‑5.5. Some of this gap may reflect a stricter grading surface for exact-answer tasks, where small differences in calculation or formatting can cause a response to fall under pass threshold. Still, these failures are scientifically meaningful because many life science workflows require outputs that are exact enough to be used directly, such as in CRISPR/HDR donor design or siRNA design.
Models also often get part of the way there without fully solving the task. In roughly 14% of tasks, models earned substantial rubric credit despite failing the exact-pass threshold. For GPT‑Rosalind, 109 tasks had pass rates below 20% while still earning at least 50% rubric reward. In practice, this means models may identify relevant evidence or produce a plausible partial answer, but still fail because they miss a key constraint, use the wrong evidence, make an incomplete calculation, or do not connect their reasoning to a scientifically useful final decision.
Limitations & what’s next
LifeSciBench is a step toward measuring how useful AI systems can be for life science research, but it is not a substitute for studying models in live research environments. The benchmark focuses on self-contained tasks that reflect recurring industry workflows, while leaving many scientific specialties and task types outside its current scope. Real research is iterative: scientists gather new evidence, revise hypotheses, design follow-up experiments, and adapt their plans as results emerge.
Strong performance on LifeSciBench should therefore be interpreted as evidence of realistic task-level capability, not as a direct measure of downstream research impact. The benchmark is grounded in industry workflows, but it does not capture the full diversity or dynamics of live research programs, where progress depends on factors that unfold over time.
The next step is to connect benchmark performance to deployment studies in live research workflows. While LifeSciBench was developed with practicing scientists, measuring whether AI systems accelerate discovery or improve R&D outcomes will require studying model use and performance in real research settings, over longer horizons, and across multiple rounds of reasoning, feedback, and experimental follow-up.
本文内容采集自官方网站,排版和翻译可能与原页面存在差异。
阅读官方全文