What actually works when you put machine intelligence into a real product or workflow — evaluation, cost, latency, failure modes, and the org design around it. Written by someone who builds with these systems, not just about them.
Anthropic's cheaper Opus 5 beats its pricier Fable 5 flagship on many tasks. So the builder's question isn't "which is better" — it's a short honest decision rule, ending in a small eval on your own work.
The signal in Anthropic's Opus 5 launch is not the benchmark table. It is that the model verifies and iterates on its own output — attacking the exact failure that makes long agent chains collapse.
AI capability isn't one number climbing toward "human level." It's a jagged frontier — superhuman at some tasks, worse than a child at others, with no smooth link between them — and that jaggedness, not the average, is what makes deployment hard and "AGI" a category error.
Recursive self-improvement is not a future event. It is a present, mundane, human-supervised loop that is real and compounding — and nothing like the intelligence explosion the phrase is meant to summon.
Recursive self-improvement isn't gated by a system's ability to rewrite itself but by its ability to tell a better version from a worse one — and self-verification hits a regress only external ground truth can break.
A world model isn't a model of text or images but of dynamics: given a state and an action, predict the next state. That definition explains why it enables planning and counterfactuals — and why learning a good one is structurally hard.
One-to-one tutoring is the most effective intervention education has ever measured, and always too expensive to give everyone. An AI tutor is the first plausible way to afford it — but only if it makes the learner struggle instead of doing the work.
Kimi K3 tops one third-party leaderboard and places second on another, and both are true. That apparent contradiction is the whole lesson in how to read a model release without being played.
Scaled video prediction absorbs an approximate physics as a byproduct, and action-conditioning turns it into a simulator agents can plan inside. But it optimizes for looking real, not being real — and that gap is the whole story.
Moonshot's headline demo — K3 autonomously designing a chip over 48 hours — is the most impressive and least checkable claim in the launch. Reading it right means separating what a real long-horizon run would prove from what a curated demo shows.
A robot that plans by imagining outcomes needs a world model that is actionable, not merely plausible — and physical reality falsifies wrong models on contact. That is what makes embodiment the hardest and most honest test of the world-model bet.
The closed-loop autonomous lab is real, and it accelerates discovery exactly where the measurement is fast, cheap, and clean. Everywhere else it inherits the noise in your ground truth and optimizes it at scale.
No human can hold a field anymore; the corpus outgrew comprehension decades ago. The highest-value near-term use of research agents is careful synthesis across the whole literature — but only if they read the primary source and keep the caveats.
Your feed is now a contest between the platform's recommender optimizing your engagement and a swarm of AI generators optimizing to be recommended. Neither has your interest in its objective. The fix is a third agent that does.
The binding constraint on autonomous agents isn't intelligence — it's that per-step success probabilities multiply. A 95%-reliable agent finishes a 20-step task 36% of the time. The fix is topology, not IQ.
MAMMAL's real contribution is not a benchmark win. It's a bet that molecules, proteins, and gene expression can share one sequence-to-sequence language — and a 458M-parameter generalist that proves the bet pays.
The end-to-end playbook for deploying AI in a real business: find the high-leverage use case, build the eval before the feature, design the human-in-the-loop, and measure ROI honestly enough to decide scale, iterate, or kill.
Few-shot prompting looks like learning, but the leading research says it's the model selecting a skill it already has. That reframes what you can and can't teach an LLM in the prompt.
Neural networks pack more concepts than they have neurons by storing them as overlapping directions, so individual neurons fire for many unrelated things. That is why we built these systems but can only partly read them.
Enterprise agent pilots stall at "impressive demo, never shipped" because teams score final answers while agents operate on trajectories — path-dependent decision sequences where one demo tells you almost nothing.
The limit on agent autonomy isn't capability, it's accountability. Every high-trust role is built around liability, and an AI bears no consequences for being wrong, so a human stays on the hook permanently.
A GUI is a translation layer between human intent and machine state. When an agent is the user, that translation is overhead — so for whole software categories the callable capability becomes the product and the screen goes vestigial.
LLM hallucination isn't a bug to patch. Truth is not a term in the training objective, so a fluent, confident falsehood is exactly what the loss rewards when the true continuation is uncertain.
The next platform war is over the personal agent: the layer that holds your context and acts for you. Whoever owns it becomes the aggregator and reduces every service beneath it to a price-competed backend.
The consequential shift isn't agents running your errands, it's agents transacting with other agents. That needs identity, binding commitment, and settlement primitives the web never built, and it opens an adversarial surface it has never faced.
Today's agents are amnesiacs that re-solve your problem from scratch every session. The next advance isn't a smarter model but persistent, structured memory, and the accumulated record of working with you is where the real moat forms.
As agents act on our behalf, the binding constraint stops being capability and becomes trust: whether an agent serves your interest, resists hijacking, and is who it claims to be. The winners will compete on verifiable trust primitives, not raw IQ.
A serious research line bets the path past current limits is not a larger language model but a world model — a system that learns an environment's dynamics so it can simulate, plan, and reason about interventions. A live bet, not a proven result.
Every extra agent buys you coordination overhead and a new error surface. A multi-agent design earns its keep only when the task has a structure one agent can't serve: parallel work, independent verification, real role separation, or a chain too long to run reliably in one pass.
The multi-agent setups that catch errors are adversarial, not cooperative: debate, generator-versus-critic, independent-then-vote. Agreement between correlated agents is worth almost nothing; the game is uncorrelated errors and a real judge.
Every model that ranks "what drives outcome Y" hands you a correlation, but you spend money on causes. The gap between the two is where data-driven companies quietly bleed, and more data makes it worse.
AI drug discovery keeps slipping because biology's labels are scarce, confounded, and often non-reproducible. You can't learn a reliable function from unreliable data; more compute just delivers the wrong answer faster.
From inside a working lab: agents compress every part of science where a check is fast and cheap, and stall wherever the answer is gated by a wet-lab experiment that takes weeks. Difficulty was never the dividing line.
AlphaFold worked because protein folding met four rare conditions most scientific problems don't. Score your problem on them before betting on "AlphaFold for X" — or get confident wrong answers faster.
A clinical AI that is right 95% of the time is more dangerous, in one specific way, than one right 70% of the time: high reliability switches off the human vigilance the whole safety case depends on, and deskilling means the backstop never forms.
Fast learned surrogates now screen millions of candidates in the time one physical run used to take. But a surrogate is valid only where it was validated, and discovery means looking outside the known — so the regime with the most value is the one you can trust least.
Clinical AI's real future isn't a diagnosis-in-a-box. It's an agent that generates the full hypothesis space and proposes the cheapest discriminating test, while the physician stays the control layer that owns the priors and the cost of being wrong.
LLMs are confident, fluent pattern-matchers that will always produce a plausible answer, right or wrong. Medicine built a discipline for reasoning safely around exactly that kind of mind: the differential diagnosis.
9 min read
Go deeper on ai.
Get new Applied AI essays — and the best of the other six pillars — delivered as they publish.