<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Scientific Workflows | Zhang Handuo's Site</title><link>https://handuo.top/tags/scientific-workflows/</link><atom:link href="https://handuo.top/tags/scientific-workflows/index.xml" rel="self" type="application/rss+xml"/><description>Scientific Workflows</description><generator>HugoBlox Kit (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Tue, 25 Aug 2026 00:00:00 +0000</lastBuildDate><image><url>https://handuo.top/media/icon_hu_39ba2e8de423d957.png</url><title>Scientific Workflows</title><link>https://handuo.top/tags/scientific-workflows/</link></image><item><title>FrontierChallenge: Evaluating Scientific Workflow Completion</title><link>https://handuo.top/publication/frontierchallenge/</link><pubDate>Tue, 25 Aug 2026 00:00:00 +0000</pubDate><guid>https://handuo.top/publication/frontierchallenge/</guid><description>&lt;p&gt;Scientific agents can run analyses and write convincing reports while still failing to deliver the code, data, figures, and evidence a researcher needs. &lt;strong&gt;FrontierChallenge measures whether a specified scientific workflow is actually finished.&lt;/strong&gt;&lt;/p&gt;
&lt;h2 id="my-contribution"&gt;My contribution&lt;/h2&gt;
&lt;p&gt;My primary responsibility was &lt;strong&gt;data synthesis&lt;/strong&gt; for FrontierChallenge. The benchmark design, evaluation, and results described below are the work of the broader team.&lt;/p&gt;
&lt;h2 id="from-raw-inputs-to-a-checkable-handoff"&gt;From raw inputs to a checkable handoff&lt;/h2&gt;
&lt;p&gt;The benchmark contains a pool of 300 workflows; the paper releases and evaluates &lt;strong&gt;97 tasks across six scientific domains and 21 workflow families&lt;/strong&gt;, with the other 203 retained internally. The released tasks span quantum chemistry, molecular dynamics, materials characterization, analytical chemistry, life science, and electrochemistry/environment.&lt;/p&gt;
&lt;p&gt;Each task supplies fixed inputs and an explicit set of required outputs. Task-specific graders check the submitted artifacts, including scientific calculations, structured files, figures, and reports. &lt;strong&gt;Pass Rate&lt;/strong&gt; measures complete delivery; &lt;strong&gt;Avg. Score&lt;/strong&gt; records partial progress. This distinction prevents a mostly correct analysis with missing or inconsistent outputs from being counted as a completed task.&lt;/p&gt;
&lt;p&gt;
&lt;figure id="figure-the-released-evaluation-set-domain-coverage-workflow-families-and-hardmedium-task-composition-source-frontierchallenge-figure-1"&gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;
&lt;img alt="Composition of the 97 released FrontierChallenge tasks across six domains and 21 workflow families."
srcset="https://handuo.top/publication/frontierchallenge/task-landscape_hu_4a8de7a1fa7d7447.webp 320w, https://handuo.top/publication/frontierchallenge/task-landscape_hu_7c9bf1d89d643f50.webp 480w, https://handuo.top/publication/frontierchallenge/task-landscape_hu_8b190dacc3262ea8.webp 760w"
sizes="(max-width: 480px) 100vw, (max-width: 768px) 90vw, (max-width: 1024px) 80vw, 760px"
src="https://handuo.top/publication/frontierchallenge/task-landscape_hu_4a8de7a1fa7d7447.webp"
width="760"
height="702"
loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;figcaption&gt;
The released evaluation set: domain coverage, workflow families, and hard/medium task composition. Source: FrontierChallenge, Figure 1.
&lt;/figcaption&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p&gt;
.&lt;/p&gt;
&lt;h2 id="what-the-evaluation-revealed"&gt;What the evaluation revealed&lt;/h2&gt;
&lt;p&gt;The v2 paper evaluates twelve models using Codex, Claude Code, and FrontierAgent scaffolds. Its central result is a substantial gap between partial progress and complete delivery:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The highest Pass Rate is &lt;strong&gt;20 of 97 tasks (20.6%)&lt;/strong&gt;, shared by GPT-5.6 Sol with Codex and Grok 4.6 with Claude Code. The highest Avg. Score reaches &lt;strong&gt;87.9/100&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;In analytical chemistry and electrochemistry/environment, the best partial scores reach &lt;strong&gt;87.6&lt;/strong&gt; and &lt;strong&gt;94.9&lt;/strong&gt;, while the highest Pass Rates are only &lt;strong&gt;4%&lt;/strong&gt; and &lt;strong&gt;0%&lt;/strong&gt;, respectively.&lt;/li&gt;
&lt;li&gt;Among 849 non-passing Claude Code runs, &lt;strong&gt;641 (75.5%)&lt;/strong&gt; end with language claiming completion. An agent&amp;rsquo;s final status message is therefore an unreliable substitute for checking its deliverables.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;These are the paper&amp;rsquo;s reported results, not a live leaderboard. The task is to execute an already specified scientific workflow; it does not test whether an agent can independently choose a research agenda. The domain slices also have different sizes and should not be read as a ranking of disciplinary difficulty.&lt;/p&gt;
&lt;h2 id="how-the-benchmark-fits-into-agent-evaluation"&gt;How the benchmark fits into agent evaluation&lt;/h2&gt;
&lt;p&gt;FrontierChallenge brings fixed scientific inputs, multiple interdependent artifacts, executable grading, and cross-domain coverage into one evaluation. The comparison below summarizes the paper&amp;rsquo;s positioning alongside related benchmarks; it is a comparison of design choices, not a performance ranking.&lt;/p&gt;
&lt;p&gt;
&lt;figure id="figure-benchmark-design-comparison-from-frontierchallenge-table-1-dots-mark-characteristics-central-to-each-benchmarks-design"&gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;
&lt;img alt="Comparison of scientific workflow coverage, fixed inputs, end-to-end execution, multiple artifacts, executable evaluation, and cross-domain science across related benchmarks."
srcset="https://handuo.top/publication/frontierchallenge/benchmark-comparison_hu_8224dc198012d7e6.webp 320w, https://handuo.top/publication/frontierchallenge/benchmark-comparison_hu_1901e2b57d1907a5.webp 480w, https://handuo.top/publication/frontierchallenge/benchmark-comparison_hu_60be068b9ed75aa4.webp 760w"
sizes="(max-width: 480px) 100vw, (max-width: 768px) 90vw, (max-width: 1024px) 80vw, 760px"
src="https://handuo.top/publication/frontierchallenge/benchmark-comparison_hu_8224dc198012d7e6.webp"
width="760"
height="381"
loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;figcaption&gt;
Benchmark design comparison from FrontierChallenge, Table 1. Dots mark characteristics central to each benchmark&amp;rsquo;s design.
&lt;/figcaption&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p&gt;
.&lt;/p&gt;
&lt;h2 id="code-and-related-work"&gt;Code and related work&lt;/h2&gt;
&lt;p&gt;The
are released within
, the agent runtime and research workbench I initiated and lead. The benchmark also connects to the process-based evaluation in
: complete outputs and the quality of the investigation provide complementary evidence about an agent&amp;rsquo;s reliability.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Sources:&lt;/strong&gt;
, especially Sections 3–4 and Table 2;
.&lt;/p&gt;</description></item></channel></rss>