Sakana Fugu Ultra Tests (high vs xhigh reasoning)

Published
Jun 23, 2026
Duration
7:18
Click to load the YouTube player

High buys more attempts, not automatic quality

  • Fugu Ultra at high reasoning completed four projects before approaching the same five-hour Codex usage limit where one earlier xhigh prompt had consumed 84%. This is a window-level observation from Ron’s two runs, not a controlled token-cost benchmark. (source video HaWddVF6GZo, 00:12, 00:18)
  • The clear win was a single-file multi-agent orchestration visualizer. Ron was impressed by its animated nodes, routing log, and latency counters. (source video HaWddVF6GZo, 01:30, 01:53)
  • The cyberpunk Hong Kong task, Astro landing-page recreation, and original product page all disappointed. Ron connected the misses to the test environment, prompt detail, and conflicting offline/CDN requirements, but the tests did not isolate those causes. (source video HaWddVF6GZo, 02:09, 03:56, 04:15, 05:03)
  • Ron’s practical recommendation is high over xhigh on Codex, plus a PRD (product requirements document), context files, and very specific guidelines. (source video HaWddVF6GZo, 05:57, 06:13)
  • The video reports a community suggestion that Pi or Droid can lower Fugu Ultra token costs. Ron did not demonstrate or quantify that route in this test, so treat it as a lead to verify rather than a measured result. (source video HaWddVF6GZo, 06:30)

Use high if you are running Fugu Ultra through Codex; xhigh burned too much of the five-hour allowance in Ron’s earlier run. But cheaper reasoning did not turn Fugu into a reliable one-shot website builder. Only one of four outputs impressed him, and the misses show that the model still needs a proper spec, context, compatible constraints, and an environment suited to the task. High buys more attempts per usage window. It does not remove the need to review the work. (source video HaWddVF6GZo, 00:12, 05:45, 05:57, 06:13)

Watch the test

Ron in his own words

“This time we’re able to do four tasks, four projects with high reasoning effort, and we are already approaching 100% of the 5-hour usage limit.” — Ron, source video HaWddVF6GZo, 00:18

“I think this is a fail for me, uh, in my book.” — Ron, source video HaWddVF6GZo, 04:26

“Structures, context files, PRDs are a must if you’re going to be using Fugu.” — Ron, source video HaWddVF6GZo, 06:01

“If you want to drive the cost down even lower, don’t use it on Codex.” — Ron, source video HaWddVF6GZo, 06:25

high stretched the usage window

The strongest evidence is throughput within Codex’s five-hour usage window. Ron says the previous xhigh test spent 84% on one prompt. With high, he completed four projects and was then approaching 100%. The video does not establish a matched-task comparison, so this does not establish a four-times cost reduction or equal output quality. It does establish why Ron prefers high for this workflow: it let him test more ideas before the window filled. (source video HaWddVF6GZo, 00:12, 00:18, 06:13)

Reasoning effort is only one routing decision. A lower setting helps when it increases completed attempts without making the outputs unusable. In Ron’s test, the usage-window result improved while the project results remained mixed.

Four-project scorecard

The observations below come from the video. The “Practical reading” column applies them without pretending Ron isolated every cause.

ProjectWhat Ron observedPractical reading
Multi-agent orchestration visualizerA single HTML file showed animated nodes, a live monospace routing log, and latency counters. Ron called the task straightforward but said the visuals impressed him. (source video HaWddVF6GZo, 01:30, 01:53)Best result. The prompt had a tight artifact boundary and named the interface elements to include.
Cyberpunk Hong Kong in Three.jsThe game existed, including a moon-triggered singularity effect, but Ron called the result poor and struggled to find the moon. He believed the 4 GB VPS was the wrong environment and recommended local Windows with a graphics card for this kind of task. (source video HaWddVF6GZo, 01:57, 02:09, 02:44, 03:02)Environment and model quality are confounded here. The video does not prove that better hardware alone fixes the output.
Astro.build recreationRon found the UI/UX similar but the design and opening color theme weak. He preferred a GLM 5.2 version produced from a full spec generated with Kimi 2.7 Code. (source video HaWddVF6GZo, 03:27, 03:35, 04:15)This was not a matched model test: both the model and the prompt-construction process differed. Web search and vision did not replace a detailed design specification in this run.
Original product landing pageEven with design language, typography, sections, and output constraints, the result was underwhelming. Ron suspected the request for CDN-loaded Tailwind and Google Fonts conflicted with entirely offline support, contributing to a default-template fallback instead of the supplied image. (source video HaWddVF6GZo, 04:34, 04:46, 05:03, 05:35)Treat the cause as Ron’s diagnosis, not a confirmed postmortem. Resolve offline-versus-network requirements before the run.

Use high for iterative Codex testing where several attempts matter. The video supplies no positive result for xhigh; it only reports higher usage-window consumption on a different prompt. It does not test whether xhigh would repair any of these four failures. (source video HaWddVF6GZo, 00:12, 06:13)

Before the next Fugu run

Ron did not demonstrate a separate Fugu recipe. These checks come from the four recorded tests.

  • Is the deliverable bounded? Name the artifact and the interface details, as Ron did for the single-file visualizer. (source video HaWddVF6GZo, 01:30)
  • Does the environment fit the job? Do not read the VPS 3D miss as clean model evidence; rerun graphics-heavy work in the local setup Ron recommends. (source video HaWddVF6GZo, 02:09)
  • Are the constraints compatible? Offline support cannot depend on network-loaded assets. Ron identifies that conflict in his fourth prompt. (source video HaWddVF6GZo, 05:03)
  • Do you have a PRD and context files? If not, write them before spending the window. Ron says one-shot work needs very detailed, specific guidelines. (source video HaWddVF6GZo, 05:57)
  • Can you inspect the output against acceptance criteria? Three finished artifacts still disappointed Ron. Completion is not the same as passing review. (source video HaWddVF6GZo, 05:45)

Freshness note

The video was published June 23, 2026. This companion was source-checked on July 18, 2026 against the immutable transcript and all 375 timestamp segments. No current Sakana documentation, Codex limit policy, Pi or Droid integration guide, pricing page, benchmark, or later Fugu default test was added. Treat the usage percentages, model availability, benchmark comparison, and community cost tip as a dated record of Ron’s test, not confirmation of the current product state. Verify today’s integrations, limits, and economics before spending production budget.

Continue learning