
I wrote most of this a month ago, the week Sakana AI shipped Fugu, and then I didn’t publish it. The draft ended on a promise: the benchmark numbers were the company’s own, nobody had re-tested them, and I said I’d watch for independent results.
Sitting on it turned out to be the useful thing. A month is long enough to find out whether anyone checked. Nobody did.
This is opinion, and the reporting below is dated to today.
One conductor, many musicians
Picture a symphony. There’s a conductor at the front and dozens of musicians behind them. The conductor doesn’t play a single note. Their entire job is to know the piece, decide who plays when, bring in the strings here and the brass there, and blend it into one performance that sounds like a single thing.
That’s Fugu.
Almost every AI you’ve used is a soloist, one model trying to do everything itself. Fugu is the conductor. Underneath it sits a pool of other AI models, each with different strengths, and Fugu’s job is to direct them: read the request, decide which model or combination fits, hand out the work, check the results, and combine everything into one answer that comes back as if it came from a single AI.
A normal AI is a soloist trying to play every instrument. Fugu is the conductor who never plays a note, and gets a better performance out of the orchestra than any one musician could give alone.
The genuinely clever part is that Fugu is itself an AI trained to do the conducting. Its coordination rules weren’t hand-written by engineers. Per Sakana, it learned to delegate, verify, and synthesize the way other models learn to write or reason, and it can call itself recursively when a problem needs breaking down.
The orchestra metaphor isn’t mine, by the way, and it isn’t a journalist’s flourish either. Sakana says Fugu is built on two of its own research papers, and one of them is named Conductor.
Why bother, instead of building a bigger soloist
Two reasons, and they’re the whole point.
You get more out of what already exists. There are several excellent models in the world, each better at some things than others. A soloist locks you into one set of strengths and weaknesses. A conductor reaches for the right one every time.
No single point of failure. The pool is swappable. If one underlying model goes down, gets more expensive, or is beaten by something new, the conductor reaches for a different musician without you noticing. Sakana says the lineup refreshes roughly every two weeks.
That second point got sharper in July. When Fugu launched, Sakana’s technical report described a pool of three closed American frontier models: Gemini 3.1 Pro, Claude Opus 4.8, and GPT-5.5. An orchestrator built entirely on models you don’t control is a dependency story wearing an independence costume. Then on July 16 Sakana announced it was adding NVIDIA’s open-weight Nemotron models to the pool. That changes the argument. A conductor that can reach for open weights is doing something the original lineup couldn’t, and it lines up with the third way I wrote about earlier, and with the own-it-don’t-rent-it theme I keep coming back to.
The month produced no scoreboard
Here’s the part I waited for.
On launch day Sakana published a benchmark table showing its top configuration, Fugu Ultra, beating each of the individual models in its own pool across most tests. Sakana also wrote, in prose, that Fugu Ultra “stands shoulder-to-shoulder with leading models like Fable 5 and Mythos Preview.”
Read that carefully, because the two claims are not the same kind of claim. The table is numbers. The Fable 5 comparison is a sentence. Sakana’s own table does not contain Fable 5 at all, and Sakana notes in the same post that Fable 5 and Mythos Preview aren’t in Fugu’s pool because they weren’t publicly accessible. So the strongest claim in the announcement is the one with no number attached to it.
You can partly check it anyway. On SWE-bench Pro, Sakana reports Fugu Ultra at 73.7. The public llm-stats leaderboard currently lists Fable 5 at 0.800 on the same benchmark. Those come from different harnesses and I wouldn’t hang a verdict on the gap, but it is not nothing, and it is the comparison Sakana chose to make in words and not in a table.
Now the part that matters more. As of today I can’t find a single independent reproduction of Sakana’s numbers. Not one. A third-party tracker does list seven Fugu scores under a “Verified” label, which looks reassuring until you notice the numbers are identical to Sakana’s own published column. That isn’t verification. That’s transcription.
And this is where I want to be fair, because it would be easy to make this a Sakana problem and it isn’t. That same llm-stats SWE-bench Pro leaderboard shows 43 entries, every one of them self-reported, and zero independently verified. Fugu isn’t even on it. The entire public scoreboard for this benchmark is vendors grading their own homework. Sakana is playing by the rules of a game where nobody checks anything.
So the honest status is: interesting architecture, real product, unverified claims, and no reason to expect that to change.
What the month did produce
Plenty, actually, which is its own kind of answer to the skeptics.
Fugu is no longer one model. It’s a line: Fugu, Fugu Ultra, and as of July 21 a security-focused Fugu-Cyber, which Sakana gates behind manual approval. Credit where it’s due, Sakana’s own release for that one says raw models “will inevitably generate false positives” and require human review, which is more candid than most vendor copy manages.
On July 24 Sakana shipped Fugu Ultra v1.1, claiming gains of up to 7.9 points and, this time, that it beats Fable 5 on complex coding and reasoning. That release publishes no table at all. It also shipped a Claude Code compatible endpoint, so you can point Claude Code at Fugu’s orchestration instead of a single model.
And people are using it. Per OpenRouter, Fugu Ultra has processed north of three billion tokens. That’s the best counterargument to the loudest criticism of the launch, which is that Fugu is an expensive black box in front of other black boxes. It’s a fair complaint: Fugu Ultra’s listed pricing sits right at frontier rates, so you’re paying top-tier prices plus orchestration latency for a routing layer you can’t inspect. Three billion tokens says a meaningful number of people ran that math and shrugged.
Worth knowing if you’re tempted: it isn’t available in the EU or EEA yet, and Fugu Ultra’s pool is fixed, so you can’t choose the musicians.
Where I’ve landed
The architecture is the interesting thing and I think it’s directionally right. The future of useful AI is much more likely to be a system that picks the right tool per task and checks its own work than one enormous brain you rent. Fugu is that idea pushed further than anyone else has pushed it, because the coordinator is itself a trained model rather than a routing table someone maintains.
What I can’t tell you is whether it’s as good as its launch post says, and after a month of looking, neither can anyone else. That’s not a knock on Sakana specifically. It’s the condition of the whole field right now: everyone publishes, nobody audits, and the scoreboard is a stack of press releases in a trench coat.
The habit that follows is simple enough. When a lab announces it has leapfrogged everyone, the first question isn’t whether the number is impressive. It’s whose scoreboard that number is on, and whether anyone else can get on it.