A new arXiv preprint on multi objective workflow generation describes a method that returns different AI agent workflows for accuracy, cost, latency, robustness, or consistency trade offs at inference time, without retraining — a distinct mechanism
A new arXiv preprint describes a method that, at inference time, returns different AI agent workflows for different trade-off weightings: accuracy, cost, latency, robustness, or consistency. The same trained model serves all of them without retraining when priorities shift.
The system, called MoFlow by its authors, frames workflow generation as a multi-objective planning problem. A tree search stores the range of trade-offs already covered at each node. A graph neural network scores partial workflows. A preference-conditioning step lets the same trained model return a different workflow for, say, a low-latency budget than for a high-accuracy one.
That trades one cost for another. The authors explicitly note that they designed the evaluation so the six adapted baselines are rerun for each new preference, while MoFlow runs its search once. The headline metric, the highest average hypervolume (a standard multi-objective score) the paper claims, is reported under a setup the authors themselves say makes apples-to-apples comparison impossible.
The paper also points to a code repository that currently returns a 404. No code, no reproduction, and no peer review has been verified. What sits on arXiv is a preprint: a research proposal with one conceptual lever — moving trade-off selection from search time to inference time — and a caveat built into its own evaluation.