The Y Combinator's Summer 2025 batch (YC S25) open source project compiles its own GPU kernels on the user's hardware before running open weight AI models (trained parameters released publicly), reporting first party benchmarks against llama.
Magnitude, a YC S25 open-source project launching this week on Hacker News, takes a different cut at the "fastest local LLM runtime" problem: instead of shipping hand-optimized kernels for popular model families, it compiles its own on the user's hardware before each model run, and its creators say the result is up to 2x the speed of llama.cpp, the popular open-source local-AI runtime.
The project, from creators identifying themselves as Anders and Tom, is a Rust engine with a custom GPU kernel runtime and an autotuner that picks tunable parameters on the user's device before model execution. It ships as a free Apache 2.0 desktop app and CLI, supporting local agents on macOS, Linux, and Windows (GitHub).
The headline number from the launch post is "up to 2x faster than llama.cpp," with llama.cpp being the popular open-source runtime that most local-AI setups default to today. The specific claim is 92% faster Metal decode and 19% faster CUDA decode, with 27% less memory per concurrent agent (Hacker News).
The mechanism is the bet, and it is genuinely different from how llama.cpp works. llama.cpp ships hand-optimized kernels for popular open-weight model families. The optimization work happens upstream, and the runtime selects from a menu of prebuilt kernels. Magnitude instead compiles kernels for the user's actual hardware (Apple Silicon, NVIDIA, AMD, or CPU-only) at install time, tuning for the specific chip, model, and workload combination. It also handles memory differently: model weights are allocated up front, session memory grows dynamically and is freed when agents stop, and a hybrid paged attention scheme shares prefix caches across concurrent agents while keeping single-session throughput high.
The benchmark numbers, as reported by the creators against llama.cpp using Qwen 3.6 35B A3B in 4-bit quantization at a 64k context with no speculative decoding, look like this: on an M4 Pro 48 GB, decode goes from 30 to 57 tokens per second, prefill from 466 to 507; on an NVIDIA DGX Spark, decode goes from 49 to 58 tokens per second, prefill from 2,033 to 2,507 (Hacker News). The team also reports per-agent memory reductions of 28% and 27% respectively.
Three things to keep in mind with those numbers. They are vendor-run, not independently reproduced. The 92% Metal decode figure is the percentage between 30 and 57 tokens per second; the launch post rounds the underlying rates, so the exact percentage is not precisely recoverable from the displayed values. And the comparison is against llama.cpp running a generalist configuration without speculative decoding, which is a known speedup technique that llama.cpp supports but was disabled in this test. The "up to 2x" claim is real, but it is a reported claim against a specific configuration, not a settled property of the engine across all models and hardware.
The project's own documentation positions Magnitude as a target for agent workloads, which is software that spawns many concurrent model sessions, each potentially with shared context, where per-agent memory and time-to-first-token matter more than peak single-session throughput (docs). The launch post lists integrations with Pi, OpenCode, Hermes, Codex, Claude Code, and Cline as part of the design intent.
What is not yet shipped: expert streaming, a fuller custom kernel compiler, and multi-device layout optimization. All three are described as plans, not current capabilities (Hacker News).
What to watch next: whether independent benchmarks reproduce the 2x figure on the same Qwen 3.6 35B A3B 4-bit 64k configuration, and whether the autotune cost at install time (the price of compile-for-your-chip) stays tolerable as more model architectures are added.