breaking papers · 83 analyzed
AI-powered analysis of breakthrough research from arXiv and beyond. We surface the work that matters before it hits the news cycle.
A recoverable geometric falsehood signal inside model activations suggests misinformation detection is becoming an interpretability problem, not just a document-lookup one.
Scaling the judge does not scale the verdict. A new research preprint, posted to arXiv, argues the fix is structural: split the checklist.
A public research preprint uses a simulated predator-prey testbed to show that systems that keep a running score of how much to trust their perceptions survive noisy inputs. The lesson travels past the model.
EntropyMoE drops the tokenizer — the step that chops text into model-digestible chunks — and routes raw-byte patches through a Mixture-of-Experts setup where only some expert sub-networks fire per input.
A single framework retunes AI agents across social objectives by changing only the evaluation metric. The proof of concept runs on one virtual fishery.
The PUSH algorithm trades exhaustive look-ahead for staggered planning over a rotating slice of agents, and authors report scaling lifelong path-finding to massive simulated fleets.
A new method, CoCo (Contribution-Contrast), aims to read the 'judge' models that score AI training responses, exposing how each expert weighs in rather than just which one was selected.
A new stress-test framework puts a robot's simulation-trained control software through thousands of randomized scenarios and only clears it for real-world use if the robot stays inside a safety boundary set in advance.
A robotics preprint proposes splitting a robot's reflexes from its planning with a switch that picks which AI to call, and reports 93 control decisions per second on a standard simulated test.
Across 512 ovarian ultrasounds, a 2026 study found the top single AI model beat expert radiologists. The strongest score, though, came from combining them.
A 2025 peer-reviewed AI-security paper (MINJA, a memory-injection attack) shows the trick can sit dormant in an assistant's memory for weeks — and the attack class is already expanding to email, with January 2026 research flagging its success rates