Meta Muse Spark 1.2 edges GPT-5.6 on Terminal-Bench
- Meta on August 5 released Muse Code in beta with Muse Spark 1.2, positioning the company against Anthropic's Claude Code and OpenAI's Codex. - VentureBeat reported Muse Spark 1.2 scored 82.9% on Terminal-Bench 2.1, ahead of GPT-5.6 Terra at 81.8% and Grok 4.5 at 81.6%. - Terminal-Bench 2.1 results are published on tbench.ai, where verified entries are run by a Terminal-Bench team member.
Meta on August 5 released Muse Code, a terminal-based coding agent in beta, alongside Muse Spark 1.2, a coding-focused update to its Muse Spark model family. VentureBeat reported the launch as Meta’s most direct push into AI coding tools, a market led by Anthropic’s Claude Code and OpenAI’s Codex. The headline number in that coverage was a Terminal-Bench 2.1 score of 82.9% for Muse Spark 1.2 running in Muse Code, ahead of OpenAI’s GPT-5.6 Terra at 81.8% and xAI’s Grok 4.5 at 81.6%, but still behind Anthropic’s Opus 5 at 86.7%. ### Why are people focused on one benchmark score? Terminal-Bench 2.1 is a benchmark for terminal agents, and its maintainers describe it as a verified refresh of Terminal-Bench 2.0 covering 89 curated tasks across software engineering, system administration, data processing, model training and security. The Terminal-Bench team said version 2.1 fixed issues in 28 of the 89 tasks and added continuous validation. (venturebeat.com) The 82.9% figure matters because Meta used it to place Muse Spark 1.2 in the middle of a closely watched contest among coding models. VentureBeat said that score narrowly exceeded GPT-5.6 Terra and Grok 4.5 on the same benchmark, while Anthropic’s Opus 5 remained ahead. (tbench.ai) ### Does the official Terminal-Bench leaderboard show those exact numbers? The official Terminal-Bench 2.1 leaderboard visible on August 6 does not show Muse Spark 1.2 among its 17 displayed verified entries. The page says a Terminal-Bench team member ran the evaluation and verified the results, and it lists GPT-5.6 Terra in Codex at 78.4% and Muse Spark 1.1 in mini-SWE-agent at 76.2%, but not the newer Meta model score cited by VentureBeat. (venturebeat.com) That gap suggests the 82.9% result was reported in launch materials or media coverage before appearing on the public verified leaderboard. The same official page says users can submit through the terminal-bench-2-1 repository, which indicates a review process rather than instant publication. ### What exactly did Meta launch with Muse Spark 1.2? (tbench.ai) Meta on August 5 released Muse Code in beta as its first terminal-based coding agent, according to VentureBeat. Mark Zuckerberg wrote on X that Muse Code “takes on complete software engineering tasks across large repos: planning changes, writing code, validating the results.” (tbench.ai) VentureBeat said Muse Code is installable on macOS or Linux with a single command, requires a Meta account and billing details, and is built around what Meta calls async background agents. Those agents remain active through a session instead of being spawned for each task, according to the report, and can split work across isolated git worktrees. (venturebeat.com) ### Why is the comparison to Codex and Claude Code getting attention? Anthropic and OpenAI have already turned coding agents into major products, and VentureBeat said Meta had largely stayed on the sidelines while focusing its developer strategy on Llama. The Muse Code launch puts Meta into direct competition with Claude Code, Codex and other agentic coding tools used by developers to ship software. (venturebeat.com) The benchmark comparison is part of that positioning. VentureBeat framed Muse Spark 1.2’s Terminal-Bench result as a narrow lead over OpenAI and xAI on one test, while still leaving Anthropic in front on the same measure. ### Where can readers watch for the next verified update? (venturebeat.com) The Terminal-Bench 2.1 leaderboard on tbench.ai is the public page to watch for verified entries, and the site says submissions run through the terminal-bench-2-1 repository. Meta’s launch details for Muse Code and Muse Spark 1.2 were published August 5, so the next concrete milestone is whether Muse Spark 1.2 appears on that verified leaderboard with the reported score. (venturebeat.com) (tbench.ai)