breaking papers · 80 analyzed
AI-powered analysis of breakthrough research from arXiv and beyond. We surface the work that matters before it hits the news cycle.
Stanford found the most popular AI detectors flag 61% of non-native English essays as machine-generated, yet universities still treat their scores as evidence of cheating.
A new Amazon AGI paper says the field's FLOP-based budgeting rule can pick a design that costs several-fold more GPU-hours to train than the math predicted, sharpest for sparse mixture-of-experts (MoE) models.
MIT's robotics AI system VLASH plans a robot's next move while the current one runs, cutting reaction delays up to 11.8× on the same hardware. Demos are research benchmarks, not deployed robots.
Motif Technologies re-released its largest open-weights language model under an MIT license, letting any builder fine-tune, embed, or sell products built on the weights.
Liquid AI's 3.1B-parameter vision-language model (an AI that reads images and answers questions about them) swaps the standard AI architecture's growing memory buffer for a fixed-size state, fitting in 3.
A public photonic-quantum company and University of Alberta chemists are pairing quantum algorithms designed for future error-correcting hardware with classical benchmarks on light-activated cancer drug molecules (photosensitizers).
Researchers show that a single 'stop' label per unsafe moment can teach the whole trajectory, by redistributing the safety signal backward through every earlier action.
Anthropic researcher Levent Alpöge, with two co-authors and Claude, closed the smallest open Hadamard matrix, a +1/−1 array stuck since 2005, and filled 11 more below 2000, with AI benchmarking group Epoch AI's mark provisional.
A black-box adversarial attack — one that never needs the robot's model internals — from the Hong Kong University of Science and Technology (Guangzhou), accepted at the major computer science conference ACM Multimedia 2026, finds the most modern
A new framework called SFS-DPO (Self-Fix Step-DPO) splits step-level reasoning from self-verification, with reported gains on math and code benchmarks over prior step-level training methods.
A new arXiv paper proves that cooperative AI systems, teams of agents sharing one goal, can collectively pick worse moves than any in their shared playbook.
A 100-code-change test across Python, Java, and C++ shows why state-of-the-art tools that try to pin down where a bug entered the code still need human auditors: the code changes that introduced the flaws run about six times larger than the fixes
AI inference and thermal ceilings have made the old hardware-to-software handoff uneconomic, spurring a second attempt 30 years on to design chips and the software that runs on them as one system, with virtual-twin simulations — software models of
A new preprint on multi-agent AI systems argues that adding a smarter model can leave a team reinforcing the same wrong number. The fix isn't always more compute.
An arxiv preprint turns the bank-chatbot pitch into a control-theory recipe — borrowed from control engineering, the math behind thermostats and cruise control — built around an orchestration layer that picks the next message and estimates visitor
A long-standing discipline in networks solved a basic problem: separate the policy decisions from the traffic they govern.
On Conway's 99-graph, a question John Horton Conway posed about whether a 99-vertex network with strict local rules can exist, an AI hit 69.43%. The wall is the result.
Following the rules is the riskier move under current AI detection: guideline-compliant edits get flagged at 64–80% while humanizer-assisted rewrites slip past at under 4%, a 16x–20x sanction gap that runs through every detector-led integrity regime.