UpTrajectory Review
Anders and Tom, two software engineers who previously built an open-source browser agent to over 4,000 GitHub stars and 100,000 downloads, are introducing Magnitude: a self-optimizing local inference engine purpose-built for AI agents. Their pitch is straightforward — existing inference engines force an awkward tradeoff. Datacenter tools like vLLM and SGLang are optimized for batched throughput, not the single-session latency that matters when one agent is reasoning through a task. Generalist tools like llama.cpp and Ollama prioritize compatibility over squeezing performance from your specific hardware. And specialized options like oMLX or ds4 lack the completeness to serve as a general engine. Magnitude claims to thread that needle, and the founders say it runs up to twice as fast as llama.cpp on identical hardware.
For a small-business operator running AI agents on local machines — whether for customer service automation, document processing, code assistance, or internal tooling — the economics here are worth paying attention to. Local inference means no per-token API costs, no data leaving your machines, and no dependency on a provider's uptime or pricing changes. But local inference has historically meant accepting sluggish performance or buying beefier hardware than your budget allows. If Magnitude's 2x speedup claim holds up in real-world workloads, that could meaningfully extend the useful life of existing machines, or make agent workloads viable on hardware you already own. For a business that has been weighing whether to invest in local AI infrastructure versus continuing to pay API bills, a credible performance jump changes that calculus.
The technical approach is where this gets genuinely interesting, and where healthy skepticism is warranted. Magnitude's key idea is on-device compilation and tuning: kernels are written with flexible parameters, then tuned against your actual hardware before the model runs. That is a real departure from shipping precompiled binaries optimized for the lowest common denominator. Combined with dynamic memory allocation — reserving only what model weights need upfront, then growing the heap as agent sessions expand — the design acknowledges something llama.cpp and Ollama largely ignore: agents are long-running, often concurrent, and the machine still needs to function for everything else you are doing. The hybrid paged attention borrowing from SGLang's approach to concurrent sessions suggests the founders have studied what works in production engines and adapted it thoughtfully rather than reinventing everything from scratch.
That said, the 2x faster claim deserves scrutiny. Faster than llama.cpp on which models, which hardware, which quantization levels, and under what concurrency? llama.cpp is a mature, battle-tested project with enormous community investment. A new engine claiming to double its performance on any hardware is a bold claim that independent benchmarking will need to verify. The founders' focus on popular open-weights model families is pragmatic — it means they are not spreading themselves thin — but it also means coverage gaps are likely early on. If your workflow depends on a less common architecture, Magnitude may not support it yet. And while dynamic memory allocation sounds elegant in theory, memory management under concurrent long-running sessions is exactly where subtle bugs and performance cliffs tend to hide.
The bigger picture is that local agent inference is becoming a real product category, not just a hobbyist pursuit. As open-weights models improve and businesses get more serious about data privacy and cost predictability, the tooling layer matters enormously. Magnitude is entering a space where llama.cpp and Ollama have enormous mindshare and community trust, but where genuine innovation in single-session performance and memory efficiency is still possible. The founders' track record with a well-received open-source project suggests they understand developer experience and community building, which will matter for adoption.
Worth watching: independent benchmarks comparing Magnitude against llama.cpp and Ollama across a range of hardware — especially Apple Silicon and consumer GPUs, where local agent workloads are most common. If you are currently running local agents and hitting performance or memory walls, this is a credible project to evaluate. If you are still on the fence about local versus API-based inference, hold off until the benchmark data is independently validated. Either way, the fact that serious engineers are building purpose-built infrastructure for local agents signals that the local AI stack is maturing faster than most businesses realize.
“We're building Magnitude, an inference engine for agents that optimizes itself to run as fast as possible on your hardware.” — Hacker News (front page)
Takeaway: If you run AI agents locally and have hit performance ceilings with llama.cpp or Ollama, Magnitude's on-device kernel tuning and dynamic memory management make it a project worth benchmarking on your own hardware.
Excerpt from the original — Hacker News (front page)
Hey HN, Anders and Tom here. We're building Magnitude, an inference engine for agents that optimizes itself to run as fast as possible on your hardware. It works on Mac, Linux, and Windows on any hardware and is up to 2x faster than llama.cpp.We're both software engineers and previously built an open source browser agent to 4k+ GH stars and 100k+ downloads. We increasingly wanted to run it on local models, but found that no inference engine worked for our use case.Inference engines today all make a performance tradeoff. They are either:- Built for batched inference on datacenter hardware at the cost of single-session performance (vLLM, SGLang)
– Designed for broad compatibility instead of optimizing for specific hardware (llama.cpp, Ollama)
– Specialized for specific hardware or models but lacking engine completeness (oMLX, ds4)Plus none of them are designed for running agents locally …