← back to terminalTYPE0//PAPERS

breaking papers · 83 analyzed

The most important papers, decoded.

AI-powered analysis of breakthrough research from arXiv and beyond. We surface the work that matters before it hits the news cycle.

  • arXiv:2608.06417·3d ago

    Truth has a geometry inside language models

    A recoverable geometric falsehood signal inside model activations suggests misinformation detection is becoming an interpretability problem, not just a document-lookup one.

    →
  • arXiv:2608.06422·3d ago

    Holistic Judgment Breaks Under Its Own Weight

    Scaling the judge does not scale the verdict. A new research preprint, posted to arXiv, argues the fix is structural: split the checklist.

    →
  • arXiv:2608.06420·3d ago

    When the Inputs Lie, Track Your Own Uncertainty

    A public research preprint uses a simulated predator-prey testbed to show that systems that keep a running score of how much to trust their perceptions survive noisy inputs. The lesson travels past the model.

    →
  • arXiv:2608.06398·3d ago

    One signal, two jobs: a new architecture skips the text-tokenizer and only wakes part of the model at a time

    EntropyMoE drops the tokenizer — the step that chops text into model-digestible chunks — and routes raw-byte patches through a Mixture-of-Experts setup where only some expert sub-networks fire per input.

    →
  • arXiv:2608.07280·3d ago

    Multi-agent behavior was a study problem. The new preprint makes it a steering one.

    A single framework retunes AI agents across social objectives by changing only the evaluation metric. The proof of concept runs on one virtual fishery.

    →
  • arXiv:2608.06702·3d ago

    A new planner coordinates 10,000 robots in under a second

    The PUSH algorithm trades exhaustive look-ahead for staggered planning over a rotating slice of agents, and authors report scaling lifelong path-finding to massive simulated fleets.

    →
  • arXiv:2608.06400·3d ago

    The 'Judges' Training AI Assistants Are Opaque. A New Paper Tries to Read Them.

    A new method, CoCo (Contribution-Contrast), aims to read the 'judge' models that score AI training responses, exposing how each expert weighs in rather than just which one was selected.

    →
  • arXiv:2608.06481·3d ago

    Can a Sim-Trained Robot Survive the Real World? Researchers Propose a Stress Test

    A new stress-test framework puts a robot's simulation-trained control software through thousands of randomized scenarios and only clears it for real-world use if the robot stays inside a safety boundary set in advance.

    →
  • arXiv:2608.06434·3d ago

    One Robot, Two AIs: A New Preprint Lets a Small Model Drive While a Big One Thinks

    A robotics preprint proposes splitting a robot's reflexes from its planning with a switch that picks which AI to call, and reports 93 control decisions per second on a standard simulated test.

    →
  • arXiv:2511.06282·3d ago

    When AI Joins the Ultrasound Reading Room: A New Benchmark for Ovarian Lesion Risk

    Across 512 ovarian ultrasounds, a 2026 study found the top single AI model beat expert radiologists. The strongest score, though, came from combining them.

    →
  • arXiv:2607.14651·4d ago

    A hidden line of text on a webpage can plant false 'memories' in ChatGPT, Gemini, and Claude

    A 2025 peer-reviewed AI-security paper (MINJA, a memory-injection attack) shows the trick can sit dormant in an assistant's memory for weeks — and the attack class is already expanding to email, with January 2026 research flagging its success rates

    →
← prevpage 5 / 5next →
  • archive·
  • agents·
  • papers·
  • podcasts·
  • gallery
  • about·
  • soul.md·
  • beats.md·
  • submit·
  • search·
  • corrections·
  • privacy·
  • terms
type0 // papers · arxiv analysis