Sakana Fugu ULTRA Review (Better than Fable 5?!)
Useful orchestration, punishing quota burn
- Fugu Ultra is an orchestrator, not a new base model. It coordinates models such as GPT and Claude, then combines their work. The Fable 5 comparison is therefore about the quality of the assembled result, not one new “brain.” (source video zaLAonePx38, 01:09; source video zaLAonePx38, 01:58)
- Sakana’s launch-day benchmarks put Ultra around Fable 5 and Mythos on hard coding and reasoning tasks, but Ron says those results were still mostly self-reported. This video does not independently reproduce the benchmark. (source video zaLAonePx38, 01:27; source video zaLAonePx38, 02:03)
- Cost dominated Ron’s June 22, 2026 test: one Fugu Ultra prompt used 29% of his weekly allowance on the $20 monthly standard plan, and the directory-review run consumed 84% of a five-hour usage limit. These are plan-meter observations from that test, not current pricing or cost-per-token data. (source video zaLAonePx38, 00:31; source video zaLAonePx38, 03:14)
- Ron judged
xhighexcessive for a simple directory review. His practical split washighfor reviews, system checks, questions, and feature drafts;xhighfor difficult refactors or long-horizon tasks. (source video zaLAonePx38, 05:20; source video zaLAonePx38, 05:40) - The tested value is convenience: automatic multi-model routing inside one task. If you can already route strong models manually, measure whether that convenience is worth Ultra’s higher latency and price. (source video zaLAonePx38, 02:40; source video zaLAonePx38, 04:19; source video zaLAonePx38, 04:29)
Fugu Ultra is smart, but “better than Fable 5” is the wrong buying question. Fable 5 is treated here as one premium model; Fugu Ultra is a learned multi-agent orchestrator, a control layer that routes parts of a task across other models and combines the result. Ron’s simple project review used xhigh, consumed most of the five-hour allowance shown in his test, and led him to call that effort choice a mistake. Start with a small, capped experiment, use high by default, and pay for xhigh only when the task is difficult enough that automatic routing can justify the cost. (source video zaLAonePx38, 00:47; source video zaLAonePx38, 03:14; source video zaLAonePx38, 05:40; source video zaLAonePx38, 08:40)
Watch the test
- 00:00 · Fugu Ultra versus Sakana Fugu: Ron separates the model label from the wider orchestration system.
- 01:09 · Not a standalone frontier model: Fugu is described as learned routing across GPT, Claude, and other frontier models.
- 02:10 · Fugu versus Fugu Ultra: the default tier targets balanced everyday work; Ultra uses deeper, more aggressive orchestration.
- 03:14 · The quota burn: a project-directory review consumes 84% of the five-hour usage limit shown in Ron’s test.
- 04:29 · Convenience is the product: automatic routing is weighed against switching models manually.
- 05:20 ·
highversusxhigh: Ron traces his fast quota burn to using maximum reasoning on a simple review. - 06:46 · The result on Ron’s CBR project: Fugu identifies strengths, limitations, and specification drift in an unfinished indicator project.
- 08:40 · The honest Fable comparison: the strength comes from routing several models into one task.
Ron in his own words
“Sakana Fugu is not a single frontier LLM.” — Ron, source video zaLAonePx38, 01:09
“So it’s really about convenience to be honest.” — Ron, source video zaLAonePx38, 04:29
“think of Fugu Ultra as the orchestrator helping route all of those other models.” — Ron, source video zaLAonePx38, 08:40
“Figure out a way to manage your cost first before building out your project.” — Ron, source video zaLAonePx38, 09:23
Fugu coordinates models
The default Fugu tier is described as balancing performance and latency for everyday coding, code review, and chat-style work. It also allows specific providers or models to be excluded from the pool for data privacy or compliance constraints. Ultra uses a deeper mixed-agent pool and more aggressive orchestration; Ron describes it as perhaps one to three agents per task, trading higher latency and a much higher per-token price for quality on complex work. (source video zaLAonePx38, 02:14; source video zaLAonePx38, 02:30)
That architecture explains the headline without endorsing it. Ron says the output shown in Codex was a compilation of GPT and Claude models examining the same task and reaching a conclusion. Later, he could see which models had been called for different parts of the job. Calling that “Fable 5 intelligence” is his shorthand for the combined result, not evidence that Sakana trained a standalone model equal to Fable 5. (source video zaLAonePx38, 02:55; source video zaLAonePx38, 08:29; source video zaLAonePx38, 08:52)
A routing decision table
The table organizes Ron’s launch-day observations by workload. It does not add evidence from a later Fugu test.
| Your job | Route suggested by the video | Why | Evidence boundary |
|---|---|---|---|
| Review a directory, check a system, ask questions, or draft a feature | Fugu Ultra on high | Ron says high still suits deep reasoning and would have been the better choice for his simple review. (source video zaLAonePx38, 05:31; source video zaLAonePx38, 05:55) | The video does not run the same task on high, so it does not show the savings or result quality. |
| Difficult refactor or genuinely long-horizon task | Consider xhigh | Ron reserves maximum reasoning effort for work difficult enough to need it. (source video zaLAonePx38, 05:36; source video zaLAonePx38, 05:49) | No matched high versus xhigh test is included. |
| Routine work where you already know which model should do each step | Route manually | Ron questions paying for full automation when strong models are already available and can be switched manually. (source video zaLAonePx38, 04:18) | The video does not calculate the cost or operator time of manual routing. |
| Experimenting without a known budget profile | Use a small capped trial | Ron recommends pay-as-you-go with $5 or less rather than immediately relying on the $20 subscription. (source video zaLAonePx38, 03:55; source video zaLAonePx38, 04:07) | This was Ron’s June 22 recommendation; current plans and limits were not checked for this companion. |
The project review found the tradeoff
Ron tested Fugu Ultra on his CBR TradingView-indicator repository. It flagged a strong indicator and blueprint layer alongside an unreliable backtesting and research layer. Ron explains why that diagnosis mattered: the project needed a real backtesting engine and paid access to order-flow history for his intended methodology, while he was trying to proceed with free inputs; he describes the resulting mismatch as specification drift. (source video zaLAonePx38, 06:51; source video zaLAonePx38, 07:34; source video zaLAonePx38, 08:14)
That is useful diagnostic evidence, but it is not a controlled Fugu-versus-Fable comparison. Fable 5 had built the indicator and blueprint earlier; several other models had continued the data and research work. The test shows Fugu finding known weaknesses in a mixed-history project. It does not show Fugu building the project, repairing the backtest, or beating Fable 5 on the same prompt. (source video zaLAonePx38, 07:06; source video zaLAonePx38, 07:23)
Freshness note
The video was published June 22, 2026. This companion was reviewed against the source video on July 18, 2026. No current Sakana documentation, pricing page, plan limits, benchmark, API behavior, or later creative-task test was added. The $5 pay-as-you-go suggestion, $20 standard plan, $100 Pro plan, 29% weekly usage, 84% five-hour usage, high and xhigh settings, API details, model pool, and benchmark position above are therefore a dated record of Ron’s June 22 test, not confirmation of Sakana’s state on July 18.
