Another major model released! Claims to rival Fable 5 and Mythos

Japan’s AI unicorn Sakana AI has unveiled the Sakana Fugu series of orchestration models, including Fugu Ultra and Fugu. Notably, the Fugu Ultra model matches or outperforms top contenders such as Fable 5 and Mythos Preview in engineering, scientific, and reasoning benchmarks.

!(https://nimg.ws.126.net/?url=http%3A%2F%2Fdingyue.ws.126.net%2F2026%2F0622%2Fb0d0d079j00th19qk003yd000ro00omg.jpg6thumbnail=660x21474836476quality=806type=jpg)

Unlike traditional large language models, Sakana Fugu doesn’t answer questions itself—it calls upon various models worldwide to complete tasks. Put simply, Sakana Fugu serves as a “commander”, selecting the best models for each task.

“Fugu” means pufferfish in Japanese. The official animation depicts Sakana Fugu as gathering many small “fish” into one big, delicious “pufferfish”.

!(https://nimg.ws.126.net/?url=http%3A%2F%2Fdingyue.ws.126.net%2F2026%2F0622%2F505ed1e6g00th19qk02agd001ei00s4g.gif6thumbnail=660x21474836476quality=806type=gif)

Sakana AI is a Japanese AI unicorn founded in 2023 by Llion Jones, the fifth author of the Transformer paper. The company has previously demonstrated “evolutionary” approaches—combining small models to achieve capabilities rivaling much larger ones. Now, with Sakana Fugu, they propose a new direction in model training: teach one model to orchestrate many, organizing various large models with different specialties into a “collective intelligence”.

On their blog, Sakana AI suggests that orchestration models will surpass traditional large models as the next frontier. While AI advances in recent years have relied on sheer compute and data, real-world complex tasks require expertise far beyond any single model’s grasp. Maximizing performance calls for collective intelligence—knowing when to use which model, how to delegate, and how to coordinate models with distinct strengths.

At the same time, this orchestration is not just a technical advance, but a geopolitical necessity. Sakana AI, learning from recent export controls placed on Anthropic models, believes that dependence on a single vendor risks abrupt access loss. In contrast, Fugu’s underlying model pool is fully swappable—if one provider is cut off, another can step in. Sakana AI calls this their “blueprint for AI sovereignty”.

Sakana AI’s blog explains that Fugu itself is a language model designed specifically to know when to delegate tasks, how agents communicate, and how their outputs can be aggregated into a reliable answer. This technical approach builds on their research into model orchestration, notably their ICLR 2026 papers Trinity and Conductor.

Technical report:

Demo:

1. Surpassing Mythos Preview and Fable 5—Orchestrating the Strongest Models for the Task

The technical report details Fugu’s performance on eight benchmarks spanning code, reasoning, science, and agent abilities. The report shows that the Fugu series models reached or nearly matched the level of cutting-edge models on all these benchmarks.

!(https://nimg.ws.126.net/?url=http%3A%2F%2Fdingyue.ws.126.net%2F2026%2F0622%2F64dc6608j00th19qk006qd001ns00ymg.jpg6thumbnail=660x21474836476quality=806type=jpg)

Notably, smart orchestration alone enabled Fugu to outperform Mythos Preview and Fable 5 on three benchmarks.

!(https://nimg.ws.126.net/?url=http%3A%2F%2Fdingyue.ws.126.net%2F2026%2F0622%2F67ad1437j00th19qk0018d000zs00cwg.jpg6thumbnail=660x21474836476quality=806type=jpg)

For cross-domain adaptability, Terminal Bench data shows that Fugu and Fugu Ultra both focused on GPT-5.5—the best performer for that task. In GPQADiamond, with Gemini-3.1-Pro as the top model, both Fugu models built their orchestration around it.

!(https://nimg.ws.126.net/?url=http%3A%2F%2Fdingyue.ws.126.net%2F2026%2F0622%2F311bc25bj00th19qk002nd000zg00iwg.jpg6thumbnail=660x21474836476quality=806type=jpg)

Fugu’s high scores come from a completely different approach: rather than training a stronger single model to solve problems, it determines which model should handle each problem, how to break down and validate sub-tasks, and combines the results for a higher-quality final answer than any individual model alone.

This is the technical report’s main point: Fugu’s value isn’t in replacing GPT, Claude, or Gemini, but in combining their respective strengths. Today, some LLMs excel at mathematical reasoning, some at code, others at security analysis. As different models develop different strengths, the ability to orchestrate is becoming a competitive edge in its own right.

2. Four Core Mechanisms: Fugu Directs a Legion of Models

The report breaks down Fugu’s core capabilities:

First, recognizing problem types. Fugu identifies whether a user query is about code, math, reasoning, information retrieval, scientific analysis, or multimodal tasks. This step sets the entire dispatch logic in motion.

Second, choosing suitable worker models. Different models excel at different tasks. Fugu is trained to pick which to call on for which problem. Even within a category—say, competitive programming—different models might be best at direct implementation, solution planning, or combining algorithmic ideas. Fugu must weigh these subtle differences.

Third, designing agent workflows. For complex tasks, Fugu Ultra generates a full agentic workflow: splitting tasks, allocating sub-tasks, sharing context, and synthesizing results. This is all coordinated in natural language within the model.

Fourth, feedback-based optimization. Fugu uses more than supervised fine-tuning; it incorporates evolutionary algorithms and reinforcement learning, using real task results to refine orchestration strategies. This teaches it how best to assign models for specific sub-tasks.

There are two model versions: Fugu and Fugu-Ultra. Fugu is meant for everyday use, balancing performance and latency—it responds quickly with high quality but doesn’t always involve complex multi-agent collaboration. Instead, it uses a lightweight selection mechanism to rapidly pick the best worker model per task.

!(https://nimg.ws.126.net/?url=http%3A%2F%2Fdingyue.ws.126.net%2F2026%2F0622%2Fef12d3b1j00th19qk0024d000ys00eyg.jpg6thumbnail=660x21474836476quality=806type=jpg)

Fugu-Ultra emphasizes quality, leveraging more sophisticated orchestration—dividing tasks into sub-tasks, delegating to different agents, and synthesizing results. This means longer response times but better results for challenging queries such as complex code, mathematical reasoning, scientific questions, and multi-step planning.

Both models are fully modular and model-agnostic—Sakana Fugu doesn’t require access to worker model weights, or even open-sourced models. New models can be immediately added to the worker pool, and users can customize the model list according to their needs for cost, privacy, or compliance.

3. Solving Rubik’s Cube, Blindfold Chess, and Even the Classic Car Wash Problem

The Sakana Fugu technical report appendix includes several experiments:

One is a “one-shot Rubik’s Cube solver”. The model must output a Rubik’s Cube solver using only Python standard libraries, then test it on 300 scrambled cubes. Both Fugu and Fugu-Ultra correctly solved all cubes, with Fugu-Ultra using fewer moves and Fugu running faster.

!(https://nimg.ws.126.net/?url=http%3A%2F%2Fdingyue.ws.126.net%2F2026%2F0622%2Fb2b1e367g00th19qk02isd000qo00f0g.gif6thumbnail=660x21474836476quality=806type=gif)

Another experiment is “blindfold chess”. The model plays chess without seeing the board, without move lists, and without FEN—just using move history. This tests whether it can maintain internal long-term state. The report shows Fugu beating several baseline models and a restricted Stockfish.

!(https://nimg.ws.126.net/?url=http%3A%2F%2Fdingyue.ws.126.net%2F2026%2F0622%2F28e81807g00th19qk01lfd000dw007tg.gif6thumbnail=660x21474836476quality=806type=gif)

There’s also an “online stock trading” experiment. The model only sees past and current anonymized market data, not future prices, and must decide each week whether to buy, hold, or sell. Fugu-Ultra posted higher average returns across five runs.

!(https://nimg.ws.126.net/?url=http%3A%2F%2Fdingyue.ws.126.net%2F2026%2F0623%2F899cb870g00th1ktk043ad001hc00u0c.gif6thumbnail=660x21474836476quality=806type=gif)

While these experiments don’t directly define model capabilities, they showcase the point: orchestration models can handle tasks requiring long-term operation, strategy, and multi-step execution.

Users have thrown typically tricky problems at Fugu-Ultra—like how many "r"s are in strawberry, is 5.11 greater than 5.1, or the classic car wash puzzle—and were impressed to see it answer all correctly, “bringing back the Fable experience”.

!(https://nimg.ws.126.net/?url=http%3A%2F%2Fdingyue.ws.126.net%2F2026%2F0622%2F0b9f6fe4j00th19qk009jd000wo00zeg.jpg6thumbnail=660x21474836476quality=806type=jpg)

The most notable innovation in the Sakana Fugu technical report is that it proposes a new research path for models.

We usually ask which model is strongest, but Sakana Fugu reframes the question: how can multiple cutting-edge models work together to achieve even more?

There are several implications: First, model capabilities become modular. New models can instantly join the worker pool and act as specialists. Second, users gain much more control. Enterprises or individuals can configure the pool according to privacy, compliance, cost, latency, or vendor preferences. Third, the competition may shift from “single-model superiority” to “system-level orchestration ability”. Whoever can best allocate models, leverage tools, design workflows, and integrate feedback will have the upper hand.

Of course, these results are vendor-supplied—real capability will depend on developer experience. Multi-model orchestration also brings higher costs and latency, especially for deeply collaborative setups like Fugu-Ultra. And, error attribution in multi-model systems is more complex—if an answer is wrong, is it the routing, a worker model, or the synthesis process?

Additionally, the orchestrator itself can introduce bias—misclassifying task type or over-relying on a certain model harms performance overall. So, while Sakana Fugu’s vision is promising, much engineering validation lies ahead before it fully arrives.

Conclusion: A New Approach to Entering Large Model Training

The launch of the Sakana Fugu series signals that the next stage of AI may not be just bigger, stronger monoliths—but also model systems that collaborate better.

If previous LLM battles were about “training super-intelligence”, Sakana Fugu is about training a “super coordinator”—models focused on learning how to delegate, coordinate, verify, and synthesize. In a field dominated by a handful of cutting-edge LLM providers, this “just orchestrate, don’t execute” paradigm might well be the new way forward for entering large-scale model training today.

4 Likes

前排沙发,第一

1 Like

还不支持欧洲部分地区,里面没有,不知道国内能不能注册登录

1 Like

真的有这么强吗

经典融合大模型,这两天在论坛出现的免费Atomesus也是

好像没有免费体验的

调度其他的模型吗。
我不太乐观,它这模型也是需要训练的,要是针对目前模型训练时间太长,那其他新的模型出的时候跟得上吗?

Just released, let’s see if there’s a chance to try it out.

果真吗?这么强、

怎么能用到这个模型

复古威武 :xhj41:

1 Like

这玩意本身是大模型,还是它只是一个调度其他大模型的 Agent 的模型?