Scientific AI has a quiet amplifier problem. Train a model on days of catalyst data and ask it to predict months of deactivation, and the small noise between laboratories appears to compound across the timescale gap. Lab predictions look calibrated. The extrapolation window is where the variance may accumulate.
Penn State and SLAC's round-robin in Nature Catalysis made that amplifier visible. Four U.S. laboratories ran identical tests on the same rhodium-based CO2-to-CO catalyst. Replicating the result proved harder than the team expected, exposing hidden variance that ordinary literature reviews absorb without flagging.
As one illustration, drug screening might cover weeks while clinical outcomes unfold over years; materials fatigue tests might run for hours while structures must last decades. The interpolation window sits inside the lab. The extrapolation window is where the noise may compound.
Rioux's team frames the fix as protocol, not pessimism. Shared procedures, reported uncertainty, and round-robin benchmarks treat the variance as a measurement problem. The path forward runs through better numbers, not fewer models.
Reported by Sky for Type0, from Variations between labs can misinform scientific AI models, team reports. Read the original: psu.edu