Sakana Fugu Ultra Tests (high vs xhigh reasoning)
High buys more attempts, not automatic quality
- Fugu Ultra at
highreasoning completed four projects before approaching the same five-hour Codex usage limit where one earlierxhighprompt had consumed 84%. This is a window-level observation from Ron’s two runs, not a controlled token-cost benchmark. (source video HaWddVF6GZo, 00:12, 00:18) - The clear win was a single-file multi-agent orchestration visualizer. Ron was impressed by its animated nodes, routing log, and latency counters. (source video HaWddVF6GZo, 01:30, 01:53)
- The cyberpunk Hong Kong task, Astro landing-page recreation, and original product page all disappointed. Ron connected the misses to the test environment, prompt detail, and conflicting offline/CDN requirements, but the tests did not isolate those causes. (source video HaWddVF6GZo, 02:09, 03:56, 04:15, 05:03)
- Ron’s practical recommendation is
highoverxhighon Codex, plus a PRD (product requirements document), context files, and very specific guidelines. (source video HaWddVF6GZo, 05:57, 06:13) - The video reports a community suggestion that Pi or Droid can lower Fugu Ultra token costs. Ron did not demonstrate or quantify that route in this test, so treat it as a lead to verify rather than a measured result. (source video HaWddVF6GZo, 06:30)
Use high if you are running Fugu Ultra through Codex; xhigh burned too much of the five-hour allowance in Ron’s earlier run. But cheaper reasoning did not turn Fugu into a reliable one-shot website builder. Only one of four outputs impressed him, and the misses show that the model still needs a proper spec, context, compatible constraints, and an environment suited to the task. High buys more attempts per usage window. It does not remove the need to review the work. (source video HaWddVF6GZo, 00:12, 05:45, 05:57, 06:13)
Watch the test
- 00:12 · The high-versus-xhigh usage result: one earlier
xhighprompt used 84% of the five-hour limit; thishighrun reached four projects before nearing 100%. - 00:58 · The orchestration visualizer maps the agents: Fugu shows work moving through planner, research, builder, sandbox, critique, and supervisor roles.
- 02:09 · Ron diagnoses the cyberpunk miss: he connects the weak 3D result to testing on a 4 GB VPS rather than a local Windows machine with a graphics card.
- 03:35 · The Astro recreation loses to GLM 5.2: similar UI/UX is not enough; Ron prefers the GLM design and finds Fugu’s opening color theme lacking.
- 05:03 · Conflicting constraints hurt the original page: the prompt asks for CDN-loaded Tailwind and Google Fonts while also requiring entirely offline support.
- 05:45 · One win, three disappointments: Ron gives the four-project result plainly.
- 06:13 · Use
highon Codex: Ron recommendshighrather thanxhighto squeeze more work from the usage window. - 06:30 · The Pi or Droid cost tip: Ron relays a community suggestion for lowering token costs outside Codex.
Ron in his own words
“This time we’re able to do four tasks, four projects with high reasoning effort, and we are already approaching 100% of the 5-hour usage limit.” — Ron, source video HaWddVF6GZo, 00:18
“I think this is a fail for me, uh, in my book.” — Ron, source video HaWddVF6GZo, 04:26
“Structures, context files, PRDs are a must if you’re going to be using Fugu.” — Ron, source video HaWddVF6GZo, 06:01
“If you want to drive the cost down even lower, don’t use it on Codex.” — Ron, source video HaWddVF6GZo, 06:25
high stretched the usage window
The strongest evidence is throughput within Codex’s five-hour usage window. Ron says the previous xhigh test spent 84% on one prompt. With high, he completed four projects and was then approaching 100%. The video does not establish a matched-task comparison, so this does not establish a four-times cost reduction or equal output quality. It does establish why Ron prefers high for this workflow: it let him test more ideas before the window filled. (source video HaWddVF6GZo, 00:12, 00:18, 06:13)
Reasoning effort is only one routing decision. A lower setting helps when it increases completed attempts without making the outputs unusable. In Ron’s test, the usage-window result improved while the project results remained mixed.
Four-project scorecard
The observations below come from the video. The “Practical reading” column applies them without pretending Ron isolated every cause.
| Project | What Ron observed | Practical reading |
|---|---|---|
| Multi-agent orchestration visualizer | A single HTML file showed animated nodes, a live monospace routing log, and latency counters. Ron called the task straightforward but said the visuals impressed him. (source video HaWddVF6GZo, 01:30, 01:53) | Best result. The prompt had a tight artifact boundary and named the interface elements to include. |
| Cyberpunk Hong Kong in Three.js | The game existed, including a moon-triggered singularity effect, but Ron called the result poor and struggled to find the moon. He believed the 4 GB VPS was the wrong environment and recommended local Windows with a graphics card for this kind of task. (source video HaWddVF6GZo, 01:57, 02:09, 02:44, 03:02) | Environment and model quality are confounded here. The video does not prove that better hardware alone fixes the output. |
| Astro.build recreation | Ron found the UI/UX similar but the design and opening color theme weak. He preferred a GLM 5.2 version produced from a full spec generated with Kimi 2.7 Code. (source video HaWddVF6GZo, 03:27, 03:35, 04:15) | This was not a matched model test: both the model and the prompt-construction process differed. Web search and vision did not replace a detailed design specification in this run. |
| Original product landing page | Even with design language, typography, sections, and output constraints, the result was underwhelming. Ron suspected the request for CDN-loaded Tailwind and Google Fonts conflicted with entirely offline support, contributing to a default-template fallback instead of the supplied image. (source video HaWddVF6GZo, 04:34, 04:46, 05:03, 05:35) | Treat the cause as Ron’s diagnosis, not a confirmed postmortem. Resolve offline-versus-network requirements before the run. |
Use high for iterative Codex testing where several attempts matter. The video supplies no positive result for xhigh; it only reports higher usage-window consumption on a different prompt. It does not test whether xhigh would repair any of these four failures. (source video HaWddVF6GZo, 00:12, 06:13)
Before the next Fugu run
Ron did not demonstrate a separate Fugu recipe. These checks come from the four recorded tests.
- Is the deliverable bounded? Name the artifact and the interface details, as Ron did for the single-file visualizer. (source video HaWddVF6GZo, 01:30)
- Does the environment fit the job? Do not read the VPS 3D miss as clean model evidence; rerun graphics-heavy work in the local setup Ron recommends. (source video HaWddVF6GZo, 02:09)
- Are the constraints compatible? Offline support cannot depend on network-loaded assets. Ron identifies that conflict in his fourth prompt. (source video HaWddVF6GZo, 05:03)
- Do you have a PRD and context files? If not, write them before spending the window. Ron says one-shot work needs very detailed, specific guidelines. (source video HaWddVF6GZo, 05:57)
- Can you inspect the output against acceptance criteria? Three finished artifacts still disappointed Ron. Completion is not the same as passing review. (source video HaWddVF6GZo, 05:45)
Freshness note
The video was published June 23, 2026. This companion was source-checked on July 18, 2026 against the immutable transcript and all 375 timestamp segments. No current Sakana documentation, Codex limit policy, Pi or Droid integration guide, pricing page, benchmark, or later Fugu default test was added. Treat the usage percentages, model availability, benchmark comparison, and community cost tip as a dated record of Ron’s test, not confirmation of the current product state. Verify today’s integrations, limits, and economics before spending production budget.
