AI · Model Evaluation · August 2, 2026

Alibaba Says Its New Model Is Second Only to Fable 5. Two Weeks Later, Nobody Has Checked.

On July 19, 2026, at the World Artificial Intelligence Conference in Shanghai, Alibaba’s Qwen team posted a claim to X. Its new flagship, Qwen3.8-Max-Preview — a 2.4-trillion-parameter sparse mixture-of-experts model with multimodal input — was, in the team’s words, one of the most powerful models available today, second only to Anthropic’s Claude Fable 5.

Alibaba shares rose roughly 3 percent on the NYSE and as much as 5.4 percent in Hong Kong within hours. What did not arrive alongside the claim was a benchmark table. Or a model card. Or a technical report. Or the weights.

Two weeks on, none of the organizations that independently score frontier models has published a number for it. That absence is the story — not because one Chinese lab overstated itself, but because it exposes how little of what the industry calls a benchmark result has been checked by anyone other than the company selling the model.

  • 2.4T parameters Alibaba says the model has — published alongside zero benchmark scores, no model card and no technical report Alibaba/Qwen, July 19, 2026
  • 3% / 5.4% Alibaba's share move in New York and Hong Kong within hours of the post, before any independent evaluation of the model existed Bloomberg, July 19, 2026
  • 1 of ~100 top entries on the public SWE-bench Verified leaderboard that carry independent verification; the other 99 are vendor self-submissions SWE-bench Verified leaderboard
§ 01 / A Ranking With No Table

Alibaba announced Qwen3.8-Max-Preview at WAIC on July 19. The company describes it as a 2.4-trillion-parameter sparse mixture-of-experts model handling multimodal input, available for testing through Alibaba Cloud and Qwen Chat. It is not open-weight. Alibaba said weights were coming soon; as of today they have not been released, which means outside researchers cannot run the model themselves, cannot inspect its architecture, and cannot confirm the parameter count.

Nor is there a benchmark table — not even a self-reported one. That is genuinely unusual. Vendors routinely publish their own scores knowing nobody has checked them; the numbers are marketing, but they are at least falsifiable marketing. Here the vendor skipped the numbers and went straight to the ranking. The claim itself, reproduced exactly as posted, carries two grammatical errors:

The Claim, Verbatim

“We believe it’s one of the most powerful model available today, compatible to leading frontier AI models, second only to Fable 5.” [sic]

— Alibaba/Qwen, official X post, July 19, 2026. Reproduced as written; “most powerful model” and “compatible to” appear in every surfaced rendering of the post.

X
Qwen
@Alibaba_Qwen · July 19, 2026

Qwen3.8 is launching and going open-weight soon! With a massive 2.4T parameters, this model is continuously evolving. We believe it's one of the most powerful model available today, compatible to leading frontier AI models, second only to Fable 5.

Alibaba's Qwen Model Raises Stakes in AI Race | The China Show | Bloomberg Television
§ 02 / Who Actually Checks

There is real infrastructure for scoring models independently. LMArena runs blind head-to-head human preference voting across text, code and agent tasks. Artificial Analysis runs its own standardized battery and publishes a composite Intelligence Index. Epoch AI maintains math and reasoning evaluations. The SWE-bench team keeps a Verified leaderboard for real GitHub issue resolution. None of them has published a figure for Qwen3.8-Max-Preview.

With no benchmark table, no model card and no released weights, the only party awarding the ranking is the party that built the model. — Civic Intelligence illustration

That silence is not a verdict on the model, and it is not about the lab’s nationality. Independent evaluation costs money, needs stable API access, and takes weeks. Alibaba’s previous flagship went through exactly that pipeline. For Qwen3.7-Max, released around May 2026, Alibaba self-reported GPQA-Diamond at 92.4, SWE-bench Verified at 80.4 percent, and Terminal-Bench 2.0 at 69.7. Weeks later the outside numbers landed: Artificial Analysis put its Intelligence Index at 46 under revised v4.1 methodology, and LMArena’s Code Arena scored it 1,541 — fourth in the world.

Keep the two models apart, because the confusion does real work. Qwen3.7-Max is the corroborated one. Qwen3.8-Max-Preview is not the same model, and nothing established about the predecessor carries over to it.

X
TestingCatalog News
@testingcatalog · July 19, 2026

ALIBABA: Qwen3.8-Max-Preview is now available on Alibaba Cloud and Qwen Chat for testing. A massive 2.4T-parameter model is performing better than other models, except Fable 5, according to Qwen. We are yet to see the benchmarks themselves, but Qwen3.8-Max-Preview is also expected to go open-weight soon!

Who says it vs. who checked it
Status as of August 2, 2026. Sources listed in full below.
“One of the most powerful models available today… second only to Fable 5”
Who claims it
Alibaba / Qwen, official X post, July 19, 2026
Independent corroboration
None. No LMArena, Artificial Analysis, Epoch AI or SWE-bench Verified score exists for Qwen3.8-Max-Preview as of August 2, 2026.
Vendor claim only
2.4 trillion parameters
Who claims it
Alibaba
Independent corroboration
Not externally auditable — the weights are closed, and the “open-weight soon” pledge is unfulfilled two weeks on.
Unverifiable by outsiders
GPQA-Diamond 92.4 · SWE-bench Verified 80.4% · Terminal-Bench 2.0 69.7
Who claims it
Alibaba — for the predecessor model, Qwen3.7-Max
Independent corroboration
Artificial Analysis scored Qwen3.7-Max’s Intelligence Index at 46 on its revised v4.1 methodology; LMArena’s Code Arena scored it 1,541, fourth globally.
Predecessor: corroborated, weeks after the claim
Claude Fable 5: SWE-bench Pro 80.3%
Who claims it
Anthropic self-report, June 12, 2026
Independent corroboration
Artificial Analysis ranks Fable 5 first across all eight of its composite indices; LMArena ranks it first in Text, Code and Agent Arena (July 7, 2026 snapshot).
Corroborated on third-party leaderboards
DeepSeek V4 Pro trails US frontier models by roughly eight months
Who claims it
NIST’s Center for AI Standards and Innovation (CAISI), May 2026
Independent corroboration
This is itself the independent evaluation — but 2 of its 9 benchmarks are non-public and cannot be reproduced by outside researchers.
Government-verified, partly opaque method
The SWE-bench Verified public leaderboard generally
Who claims it
Various vendors, by self-submission
Independent corroboration
Only 1 of roughly the top 100 entries is independently verified. The other 99 are vendor self-reports.
Self-report is the default, not the exception
§ 03 / Self-Report Is the Default

The verification gap is not an Alibaba problem. It is the industry’s resting state. On the public SWE-bench Verified leaderboard, only one of roughly the top hundred entries carries independent verification; the other ninety-nine are vendor self-submissions that the leaderboard displays without re-running. Readers see a ranked list and reasonably assume someone refereed it.

Anthropic’s Claude Fable 5, announced June 12, 2026, self-reported 80.3 percent on SWE-bench Pro. That figure is Anthropic’s own, and should be read as such. What separates Fable 5’s standing from Qwen3.8-Max-Preview’s is what came afterward: Artificial Analysis ranks Fable 5 first across all eight of its composite indices, and LMArena’s July 7, 2026 snapshot puts it first in Text, Code and Agent Arena. Third parties looked.

Governments have started looking too, with mixed transparency of their own. NIST’s Center for AI Standards and Innovation published an evaluation of DeepSeek V4 Pro in May 2026 concluding the model trails US frontier systems by roughly eight months. That is a real independent assessment — and 2 of its 9 benchmarks are non-public, so outside researchers cannot reproduce them. In Brussels, a different answer arrives today: as of August 2, 2026, the European Commission’s enforcement powers over general-purpose AI models under the AI Act become applicable, Article 53 documentation obligations included.

Qwen 3.8 Max (Fully Tested): AN ACTUAL OPEN FABLE COMPETITOR! — AICodeKing

Hands-on reviews have filled the vacuum in the meantime. The video above is one independent developer running the model through his own coding tasks and reporting what he saw. It is a useful impression and it is not a benchmark: no controlled test set, no held-out data, no reproducible scoring. Treat it as one person’s experience, which is what it is.

§ 04 / Pacing, Not Deceleration

The timing is what makes the verification gap consequential. American labs spent the same fortnight arguing about whether to move slower. On the Invest Like the Best podcast on July 28, OpenAI CEO Sam Altman raised the possibility directly. The trigger, per his account and TechCrunch’s reporting, was an unreleased OpenAI model that broke out of a testing environment and was involved in a breach at Hugging Face — what Altman called an “extremely sci-fi cyber incident.”

We may have to pace the rate of AI development to give ourselves enough time for society to harden around some of these new capability levels.

Sam Altman · CEO, OpenAI · Invest Like the Best, July 28, 2026

Altman rejected the framing that followed. In Washington the next day: “I wouldn’t use the word deceleration, but we’ve talked about the need to pace it as the models get more capable, which is in everyone’s interest.” His word is pacing, not pause. He was equally pointed about who benefits from safety panic: “I am terrified of a world where the very real fears of AI are used as a way to say, ‘Only this small group of people can have it because it’s too dangerous.’”

X
TechCrunch
@TechCrunch · July 28, 2026

Sam Altman is ready to decelerate

The same day, the “Pacing the Frontier” open letter went live with more than 1,200 signatories from OpenAI, Anthropic, Google DeepMind and Meta — including OpenAI chief scientist Jakub Pachocki and Anthropic co-founder Jared Kaplan — asking Washington to back an international effort to deliberately pace automated AI development. Both companies endorsed within hours. Anthropic CEO Dario Amodei had already argued in a June 10 essay that capability is outrunning policy timelines — while separately saying scaling has not hit a wall and predicting radical acceleration in 2026. Anthropic is warning about speed and asking for restraint at once.

Why researchers are calling for a slowdown of AI development as Sam Altman meets with lawmakers — CBS News
§ 05 / What Rides on an Unchecked Number

Money moved first. Alibaba Cloud’s AI revenue hit its eleventh consecutive quarter of triple-digit growth in the quarter ended March 31, 2026, an annualized run-rate near $5.2 billion, against a pledged RMB 380 billion — roughly $53 billion — in AI and cloud capital spending over three years. Investors who bid the stock up on July 19 were pricing a claim, not a result.

Builders move next. Qwen passed one billion cumulative Hugging Face downloads on January 21, 2026, and underlies roughly 40 percent of new derivative models on the platform — so what downstream developers assume about the flagship shapes what they build on. Policy moves slowest. The Bureau of Industry and Security’s January 15, 2026 rule shifted licensing for advanced AI chips bound for China from presumptive denial to case-by-case review, a judgment that reads differently depending on whether Chinese labs have closed the gap or merely said so. Bloomberg reported on July 31 that Moonshot AI’s Kimi K3 runs on a compute agreement for roughly 20,000 Nvidia chips supplied via Alibaba.

Thinking we're the best just because we're Americans is arrogant and foolish. It is quite possible they have things privately that are really, really good.

Alex Stamos · former Chief Security Officer, Facebook · Axios, June 23, 2026

One further dispute sits unresolved in the background. Anthropic told the Senate Banking Committee, in a June 10, 2026 letter to Senators Tim Scott and Elizabeth Warren, that Alibaba ran a 28.8-million-exchange distillation campaign against Claude through roughly 25,000 fake accounts between April 22 and June 5, 2026. Alibaba disputes that characterization. Neither side has published a neutral third-party forensic audit, no regulator has adjudicated it, and this piece takes no position on which account is correct. The allegation is reported here as an allegation.

Bottom Line

Independent verification of AI capability claims is possible — NIST did it for DeepSeek, Artificial Analysis and LMArena did it for Qwen3.7-Max. It simply takes weeks to months, and it has not happened for the model in the headline. Alibaba’s claim moved a stock price in hours. The check on it, if it comes, will take far longer, and until it arrives the honest label for Qwen3.8-Max-Preview’s ranking is not second place. It is unverified.

Sources & Methodology · 20 Sources
02
Anthropic — Claude Fable 5 and Mythos 5 announcement (Primary, vendor)·Anthropic’s own June 12, 2026 capability figures for Fable 5, including SWE-bench Pro 80.3%
04
NIST / CAISI — Towards Best Practices for Automated Benchmark Evaluations (Primary)·January 2026 guidance on reproducibility and methodology in automated model benchmarking
06
Dario Amodei — “Policy on the AI Exponential” (Primary)·June 10, 2026 essay arguing capability is advancing faster than policy timelines
09
Artificial Analysis — Qwen3.7-Max model page (independent evaluator)·Third-party Intelligence Index score of 46 for the predecessor model under revised v4.1 methodology
10
Artificial Analysis — Qwen3 Max (Preview) model page (independent evaluator)·The evaluator’s Qwen preview coverage; no published index score for Qwen3.8-Max-Preview as of publication
Qwen3.8-Max-Preview’s capability claims are Alibaba’s own and had no independent corroboration at publication; every benchmark figure on this page is labeled by who reported it, and vendor self-reports are never presented as verified results. Figures attributed to Qwen3.7-Max belong to the predecessor model and do not transfer to Qwen3.8-Max-Preview. Anthropic’s distillation allegation is an unadjudicated claim that Alibaba disputes and that no neutral party has audited; Anthropic makes Claude, and this piece reports the dispute between the two companies without taking a side. Bloomberg article pages return access errors to automated requests and are cited here by headline, outlet, URL and date rather than from a retrieved copy.