Meta's AI push has been given a boost when Muse Spark 1.3, the new model with the best performance on industry benchmark, was shown up by Meta Chief AI Officer Alexandr Wang, who took a swipe at Google’s Gemini lineup. Wang seemed to use the results to highlight Meta’s growing position in the increasingly competitive generative AI market.
In a post on X, Wang shared a chart of the recent rankings of the Artificial Analysis Intelligence Index and said: “I really hate to say it, but... Gemini Who?” He was widely seen as a jab at Google’s Gemini models, as the chart showed Muse Spark 1.3 ahead of other Gemini models.
According to the benchmark information cited in the post, Muse Spark 1.3 scored 62 on the Artificial Analysis Intelligence Index. This put Meta in second place behind Anthropic’s Claude Fable 5.1 and Claude Opus 5, making it one of the best artificial intelligence systems included in the comparison.
The results showed Muse Spark 1.3 outperforming Google's Gemini 1.5 Pro and Gemini 1.5 Flash. The two Google models had 61 and 53 scores, respectively. While benchmark results are only one way of evaluating AI systems, the figures provide Meta with a strong talking point as competition between big AI companies continues to escalate.
Wang’s comments came after Artificial Analysis announced the release of Muse Spark 1.3. The announcement highlighted the rapid pace with which Meta was developing the Muse Spark family. As mentioned in the thread, Muse Spark 1.3 is the fourth Muse Spark model release in approximately five months.
The pace of releases is important as AI companies are increasingly competing not only on raw model intelligence but also how quickly they can improve reasoning, coding, tool use and autonomous task completion. Meta’s latest results suggest that the company is attempting to close the gap with some of the most advanced models available from competitors.
Artificial Analysis data also reported improvements in scientific reasoning and agentic task performance. Agentic AI refers to systems that can independently perform a sequence of actions to accomplish a task, often involving external tools, software environments or multiple decision-making steps.
Muse Spark 1.3 had a 12-point improvement over Muse Spark 1.2 on the Tau3-Bench Banking evaluation. It also had a five-point improvement on Terminal Bench 2.1. These gains are particularly important because agentic capabilities are at the center of next generation AI products.
A limited-preview version of Muse Spark 1.3 achieved 52 percent on the Tau3-Bench Banking benchmark, which was described as the highest score of any model on this test. If the results are replicated in the wider testing, Meta’s new model is going to become more competitive in actual work situations.
Meta’s pitch is not based solely on benchmark performance. Cost is another factor that could make Muse Spark attractive to developers and businesses. The high variant was reported to cost around $0.55 per task, significantly below some competing AI systems. Based on the information mentioned in the benchmark discussion, models such as GPT 5.6 Sol and Grok 4.6 were priced almost twice the cost of the Meta model for comparable tasks.
The combination of performance and pricing is vitally important in an AI industry where companies and developers are looking beyond benchmark scores. A model that can provide comparable intelligence but has lower operating costs can be particularly appealing for businesses deploying AI at scale.
Muse Spark 1.3 is also considered “the Pareto frontier” for intelligence versus cost, analysts say. Put into simple terms, the model is being positioned as an attractive tradeoff between capability and expense rather than only for the highest possible benchmark score.
Wang’s “Gemini Who?” comment also reflects the increasingly aggressive tone of competition among leading AI companies. Google, Meta, Anthropic, OpenAI and xAI are all investing heavily in increasingly capable models, with each company attempting to differentiate its systems through reasoning, multimodal capabilities, agentic functions, speed and pricing.
For Meta, Muse Spark’s rapid development could help to shape its AI strategy. The company has been investing heavily in artificial intelligence in consumer products, research and infrastructure as well as in the business of building up a big presence along with the companies that have dominated much of the generative AI conversation.
At the same time, benchmark rankings should not be considered a definitive measure of which AI model is “best.” Different evaluations test different abilities, and real-world performance can vary depending on the task, prompting technique, context length, tools and deployment environment. Still, the latest results give Meta a useful opportunity to demonstrate the progress of its AI models.
With Muse Spark 1.3 appearing at the top of the list, and with a purported cost advantage, Meta is targeting intelligence and affordability at the same time. Wang’s sharp comment about Gemini may have been meant to be provocative, but the message is plain: Meta wants the AI industry to take its latest models seriously.