By scoring AI physics predictions against the underlying equations, researchers turn the violation into spatially adaptive confidence bands that hold across six physics problems.
Engineers want AI to do the slow part of their job: solve a partial differential equation. Heat through a turbine blade, stress through a structural part, fluid through a porous medium. A neural operator is a neural network trained to approximate the answer to that kind of physics problem, useful when running the real simulation takes hours and the engineer needs an answer in milliseconds. The catch is that the AI gives one number and no honest answer to the obvious follow-up: how much should I trust it?
A new arXiv preprint reframes the question. Instead of asking the model for an answer, it asks how badly the model breaks the equations it was supposed to learn. The authors call the approach Physics-Informed Conformal Prediction, and the headline empirical claim is that across four conformal methods and six physics scenarios, heat conduction in 2D and 3D, structural mechanics in 2D and 3D, Darcy flow, and Navier-Stokes, their framework delivers empirical coverage between 89% and 91%, while the field's standard uncertainty tricks, MC Dropout and Deep Ensembles, swing from 82% to 100% on the same problems.
Conformal prediction is the older half of the framework: a statistical recipe for turning any model's errors into a confidence band with a coverage guarantee that does not depend on the data distribution. Split-conformal, the variant the paper uses, holds out a calibration set, scores each example by some nonconformity measure, and uses the empirical quantile of those scores to draw prediction intervals. The new contribution is the nonconformity score itself. Instead of asking how far a prediction sits from the training data, the authors ask how large the residual of the governing PDE is at the model's output. Where the residual is small, the band tightens. Where the equations are visibly violated, the band widens.
That adaptivity is what the 89-91% number actually buys. A fixed-width band could hit 90% on average and still leave a hot spot in a stress concentration or a wake entirely uncovered. Scoring by PDE residual puts the honesty where the physics is hardest.
The 63x figure is narrower still. The authors prove that a popular neural-operator architecture, the Fourier Neural Operator (FNO), has a built-in approximation barrier for PDEs with Dirichlet boundary conditions, the standard "the value is fixed on this edge" problem, because the network is translation-equivariant by construction. Adding coordinate channels to the input breaks that constraint, and on the study's setup the fix cuts error by up to 63 times. FNO also outperforms CNN and DeepONet by 10-12x in the same benchmarks, which the paper reads as a separate vote of confidence in the architecture once the boundary problem is handled.
The honest limit list is long. This is an arXiv preprint submitted in late June 2026, not a peer-reviewed deployment study. The six scenarios are benchmarks, not engineering workloads. Conformal coverage is a finite-sample guarantee under exchangeability, the assumption that the calibration data and the test data are drawn from the same distribution, and that assumption does not survive a real distribution shift, such as a geometry the model has never seen or a load case outside the training envelope. The framework is also architecture-aware in a specific way: the 63x result is conditional on FNO plus coordinate channels on Dirichlet problems, not a general AI improvement.
The paper establishes a way to attach a distribution-free, spatially adaptive uncertainty estimate to a class of models that have typically returned only point predictions. The next test is whether the PDE-residual score travels: whether a calibration set built on one problem geometry still produces honest bands on a neighboring one. If it does, neural operators get closer to being a drop-in surrogate with a real error bar attached.