Sakana AI, the Tokyo-based lab last valued around $2.65 billion in a November 2025 Series B, shipped two new models on September 11, 2026: Fugu Max v1.0 and Fugu Ultra v2.0. Neither is a conventional trained frontier model in the sense that GPT-6 Astra or Claude Opus 5 are. Both are the second generation of Sakana's "Fugu" line, which the company markets under the tagline "a Multi-Agent System, Delivered as One Model."
That is not marketing fluff to skip past. Fugu is a model Sakana trained to act as a coordinator: it sits behind a single OpenAI-compatible API endpoint, decides whether a prompt needs one pass or a team, and — when it needs a team — breaks the task apart, routes the pieces to a pool of other open-weight and specialized models, checks their work, and stitches together one final answer. The design traces back to two Sakana research projects called TRINITY (a lightweight coordinator that assigns Thinker, Worker and Verifier roles across models over multiple turns) and Conductor (trained with reinforcement learning to invent its own natural-language coordination strategies). The original Fugu launched June 22, 2026; Max and Ultra v2 are its first major refresh.
The headline act is Fugu Ultra v2: Sakana says it takes best-or-joint-best results on 5 of its own 8 benchmarks and lands in the top two on 7 of 8, at a moment when the underlying model pool it draws from reportedly no longer includes Claude Fable 5, Claude Fable 5.1, or the just-released GPT-6 Astra. Beating models it never actually calls is the part worth being skeptical about, and that's what this piece digs into.
Quick summary: Fugu Max ($2/$6 per million input/output tokens) and Fugu Ultra v2 ($5/$30, rising to $10/$45 above 272K context) are not single trained frontier models — they're orchestration systems that route a prompt across a pool of other AI models behind one API. Fugu Ultra v2 claims a 48.3 on Sakana's own "Chartography" benchmark versus 27.3 for Claude Opus 5 and 29.5 for Claude Fable 5, plus a 74.3 on "DeepSWE," but its model pool explicitly excludes Fable 5, Fable 5.1 and GPT-6 Astra (training/eval cutoff August 28, 2026). Every number above is Sakana-reported; no independent lab had replicated the results as of publication, and at least one third-party review found Fugu Ultra v2 trailing Anthropic's shipped Claude Mythos 5 on several overlapping benchmarks.
Two models, one mission each
Sakana positions Max and Ultra v2 as answers to two different questions rather than a simple good/better tier. Fugu Max asks what the best output is at the lowest cost, leaning on a large pool of open-weight and specialized models — including Nvidia's Nemotron family — and Sakana says it's 40 to 60% cheaper than Claude Sonnet 5, GPT-5.6 Terra and Kimi K3, while topping Sakana's own suite on 6 of 10 benchmarks (Terminal-Bench 2.1, GPQA Diamond, AA-LCR, GDP.pdf, AutomationBench and SWEFish) and pushing the reported cost-performance frontier on 7 of 10. Fugu Ultra v2 asks what the highest capability looks like on complex, multi-step reasoning, research and software-engineering work, using a deeper expert pool at a steeper price. Base Fugu still lets users opt individual providers or models in or out through console settings; Max and Ultra v2 ship with fixed pools by design, with no user-level customization outside custom enterprise deals. Sources: [Introducing Fugu Max and Fugu Ultra v2](https://sakana.ai/fugu-max-release/), [Sakana Fugu — Multi-agent System as A Model](https://sakana.ai/fugu/)
| Spec | Fugu Max | Fugu Ultra v2 |
|---|---|---|
| Optimizes for | Lowest cost per quality output | Highest capability on complex tasks |
| Price per 1M tokens (in/out) | $2 / $6 | $5 / $30 (standard); $10 / $45 above 272K context |
| Reported context window | Not disclosed by Sakana | Not independently confirmed (Sakana has not published an official Ultra v2 context-window spec in the sources checked) |
| Own-suite benchmarks led | 6 of 10 | Best/joint-best on 5 of 8; top-2 on 7 of 8 |
| Claimed cost edge | 40–60% cheaper than Sonnet 5, GPT-5.6 Terra, Kimi K3 | Not claimed as a cost leader |
| Model pool customization | Fixed, by design | Fixed, by design (excludes Fable 5, Fable 5.1, GPT-6 Astra) |
The benchmark numbers
Sakana's two most-quoted Fugu Ultra v2 results are Chartography, a visual-reasoning benchmark where it reports 48.3 against 27.3 for Claude Opus 5 and 29.5 for Claude Fable 5, and DeepSWE, a software-engineering benchmark where it reports 74.3. Both figures show up identically across Sakana's own release page and every outlet that covered the launch, which at least confirms they're being reported consistently, if not independently. Worth flagging: Sakana's eight-benchmark suite mixes established public evals like GPQA Diamond and Terminal-Bench 2.1 with benchmarks it built itself — Chartography, DeepSWE, SWEFish, Toolathon, GDP.pdf, AA-LCR and AutomationBench. Coverage of the launch specifically noted that SWEFish is an internal Sakana benchmark and should be read as a vendor signal rather than a neutral third-party score. Sources: [Sakana AI Launches Fugu Max and Fugu Ultra v2 — MarkTechPost](https://www.marktechpost.com/2026/09/10/sakana-ai-launches-fugu-max-and-fugu-ultra-v2-for-cheaper-stronger-multi-agent-orchestration/), [Sakana AI Splits Fugu Into Max and Ultra v2 — AlphaSignal](https://alphasignal.ai/news/sakana-ai-splits-fugu-into-max-and-ultra-v2-to-cut-costs-60)
The pool paradox
Here's the detail that makes this launch more interesting than a routine benchmark chart: Fugu Ultra v2's model pool has a training and evaluation cutoff of August 28, 2026, and Sakana confirms it deliberately excludes Claude Fable 5, Claude Fable 5.1 and GPT-6 Astra — the model OpenAI had only shipped to approved users on September 3, 2026, just eight days before Fugu Ultra v2 went live. Sakana frames the omission as proof its system doesn't depend on any single vendor's flagship rather than as a limitation. Sources: [Sakana Fugu — Multi-agent System as A Model](https://sakana.ai/fugu/), [OpenAI announces rollout of GPT-6 Astra model — CNBC](https://www.cnbc.com/2026/09/03/open-ai-astra-gpt-6-cyber.html)
There's real history behind that framing, though it's easy to conflate two separate events. When the original Fugu launched on June 22, 2026, Claude Fable 5 and Claude Mythos 5 were, for a few weeks, literally unreachable: Forbes and CNBC reported that the US Commerce Department had ordered Anthropic on June 12–13, 2026 to suspend both models for all foreign nationals under export-control rules tied to a flagged jailbreak vulnerability, and that because Anthropic had no way to selectively enforce that by nationality, it pulled both models worldwide; those outlets reported the order was lifted and Fable 5 access restored globally by June 30–July 1, 2026. So by the time Fugu Ultra v2 shipped in September, Fable 5 had been fully available again for more than two months — and Sakana chose to keep it, Fable 5.1, and GPT-6 Astra out of the pool anyway. That's a design decision, not a workaround for an outage. Sources: [Anthropic Disabled Fable 5 And Mythos 5 After A U.S. Export-Control Order — Forbes](https://www.forbes.com/sites/anishasircar/2026/06/16/anthropic-disabled-fable-5-and-mythos-5-after-a-us-export-control-order-heres-what-happened/), [Anthropic says Trump admin has lifted export controls on Claude Fable 5 and Mythos 5 — CNBC](https://www.cnbc.com/2026/06/30/anthropic-says-trump-admin-has-lifted-export-controls-on-claude-fable-5-and-mythos-5.html)
The skeptical read
Sakana's numbers are, by the company's own account, self-reported, and no independent benchmarking organization had published a replication of the Chartography or DeepSWE scores as of this writing. The sharpest public pushback came from Prime Intellect research engineer Elie Bakouch, who wrote on X that Fugu is "a closed source orchestrator on top of closed source models," adding that "if before you didn't control the models, now you don't even control which ones are used or how much," and directly disputing Sakana's framing of the system as offering "AI sovereignty." Bakouch also pointed out that Sakana doesn't disclose what share of a given benchmark run used closed-source partner models versus open-weight ones, calling that omission misleading given the company's independence pitch. Sources: [Elie Bakouch on X](https://x.com/eliebakouch/status/2068939729811468503)
A separate review by Kingy AI lined Fugu Ultra v2's reported scores up against Anthropic's official Claude Mythos 5 system card, rather than the "Mythos Preview" Sakana had benchmarked against, and found a mixed picture: Fugu Ultra v2 leads on GPQA Diamond (95.5 vs. 94.1) but trails on SWE-Bench Pro (73.7 vs. 80.3), Terminal-Bench 2.1 (82.1 vs. 88.0), and Humanity's Last Exam (50.0 vs. 59.0). The broader structural critique, echoed across several write-ups of the launch, is that when an orchestrator scores well by calling other labs' models, the score belongs to the system rather than to any individual model and is hard to audit from the outside — one summary of the reaction put it as "Fugu did not beat the flagship, it hired the flagship." Reaction to the original Fugu concept on Reddit converged on a similar read, characterizing it as "a highly advanced router/wrapper, not a fundamental leap in intelligence." Sources: [Sakana Fugu Ultra v2: a smaller pool, a higher score — OrcaRouter](https://www.orcarouter.ai/blog/fugu-ultra-v2-explained), [Did Sakana Fugu Ultra Really Match Fable 5 and Mythos 5? — Kingy AI](https://kingy.ai/blog/sakana-fugu-ultra-fable-mythos-benchmarks/)
| Benchmark | Fugu Ultra v2 (Sakana-reported) | Claude Opus 5 | Fable 5 / Mythos 5 | Note |
|---|---|---|---|---|
| Chartography (visual reasoning) | 48.3 | 27.3 | 29.5 (Fable 5) | Sakana's own benchmark |
| DeepSWE (software engineering) | 74.3 | — | — | Sakana's own benchmark; no rival scores published |
| GPQA Diamond | 95.5 | — | 94.1 (Mythos 5) | Per independent Kingy AI comparison |
| SWE-Bench Pro | 73.7 | — | 80.3 (Mythos 5) | Fugu trails, per Kingy AI |
| Terminal-Bench 2.1 | 82.1 | — | 88.0 (Mythos 5) | Fugu trails, per Kingy AI |
| Humanity's Last Exam | 50.0 | — | 59.0 (Mythos 5) | Fugu trails, per Kingy AI |
Treat any single-source "X beats Y" benchmark headline about Fugu with caution until an independent evaluator like Artificial Analysis or LMArena publishes a replication — and check whether the comparison model is the one you'd actually be choosing between in production.
01What is Sakana Fugu Ultra v2, in one sentence?
It's Sakana AI's second-generation "Fugu" system — not a single trained model but a learned orchestrator that decomposes a prompt, routes pieces of it to a pool of other AI models, verifies the results, and synthesizes one answer through a single OpenAI-compatible API endpoint.
02Is Fugu Ultra v2 really an AI model, or just a wrapper around other models?
Both, depending on who you ask. Sakana trained Fugu itself to decide when to answer directly and when to assemble a team of specialist models, a design it traces to its TRINITY and Conductor research. Critics, including Prime Intellect's Elie Bakouch, have called it a closed-source orchestrator layered on other closed-source models rather than a genuine leap in underlying intelligence.
03Why doesn't Fugu Ultra v2's model pool include GPT-6 Astra or Claude Fable 5?
Sakana says the exclusion is deliberate: Fugu Ultra v2's pool has a training/eval cutoff of August 28, 2026 and simply doesn't route to Fable 5, Fable 5.1 or GPT-6 Astra, which Sakana frames as proof the system isn't dependent on any single vendor's flagship. This is separate from a brief US export-control order in June 2026 that temporarily pulled Fable 5 and Mythos 5 worldwide; that order was lifted by early July, well before Ultra v2 shipped.
04How much does Fugu Ultra v2 cost to use?
Standard pricing is reported at $5 per million input tokens and $30 per million output tokens, rising to $10 and $45 per million above 272,000 tokens of context. Sakana's pricing does reference a 272,000-token tier, but we could not independently confirm a specific maximum context-window figure for Ultra v2. Fugu Max, the cheaper sibling, is priced at $2 input and $6 output per million tokens.
05Are Sakana's "beats GPT-6 Astra" style benchmark claims independently verified?
Not as of this writing. Every benchmark number in Sakana's launch materials, including the widely cited Chartography score of 48.3 versus Claude Opus 5's 27.3, is self-reported, and Sakana's suite mixes standard public evals with benchmarks it designed itself. No third-party evaluator had published an independent replication, and at least one outside comparison found Fugu Ultra v2 trailing Anthropic's shipped Mythos 5 on several benchmarks.
06What's the practical difference between Fugu Max and Fugu Ultra v2?
Fugu Max optimizes for cost, routing across a broad pool of open-weight and specialized models (including Nvidia's Nemotron family) to reportedly undercut Claude Sonnet 5, GPT-5.6 Terra and Kimi K3 by 40–60% on price while leading Sakana's own suite on 6 of 10 benchmarks. Fugu Ultra v2 optimizes for peak capability on complex, multi-step reasoning and coding work, at roughly two to five times the price.
Whether an orchestration layer like Fugu Ultra v2 counts as "beating" GPT-6 Astra or Claude Opus 5 is exactly the kind of question that's hard to settle from a spec sheet alone. LumiChats lets you compare model cards side by side and chat with many of today's leading AI systems in one place, so you can run your own prompts and judge a routed answer against a single frontier model for yourself, rather than taking any one lab's benchmark chart at face value.
