Anthropic has revealed that it is already using Anthropic Model 2, an unreleased artificial intelligence system that the company says is somewhat more capable than Claude Mythos 5. The model was disclosed in Anthropic’s August 2026 Risk Report, offering a rare look at an advanced system that is being used internally but is not currently available to the public.
Model 2 Outperforms Mythos 5
The clearest evidence of Anthropic Model 2’s capabilities comes from CoBench v2, an internal benchmark designed around real-world research and engineering challenges.
The benchmark contains 449 problems that had previously been solved by Anthropic employees. Model 2 achieved a score of 62.8%, significantly ahead of Claude Mythos 5, which scored 50.3%. Mythos Preview recorded 54.8%.
The gap becomes even more notable when compared with earlier Claude models. Claude Opus 4.7 scored 27.4%, while Opus 4.6 reached only 15.6% on the same evaluation.
These results suggest that Anthropic has made considerable progress in the ability of its systems to handle complex technical and research-oriented tasks.
However, the company is not presenting Model 2 as a system capable of replacing its technical workforce. Anthropic estimates that a model would need to reach at least 85% on CoBench to fully replace its technical research staff.
With a 62.8% score, Anthropic Model 2 remains well below that hypothetical threshold.
A Noticeable but Smaller Improvement
Anthropic describes Model 2 as a “noticeable improvement” over Mythos 5 across many internal tasks. At the same time, the company says the improvement is smaller than the leap it previously observed when moving from Claude Opus 4.6 to Mythos Preview.
That distinction is important because progress in frontier AI is becoming increasingly difficult to measure simply by comparing model generations.
As systems become more capable, improvements can involve specialized reasoning, coding, tool use and long-running autonomous tasks rather than dramatic changes in ordinary chatbot conversations.
The CoBench results indicate that Model 2’s strongest advantages may be particularly relevant to technical work rather than everyday consumer applications.
Model 2 Is Already Used Internally
One of the most interesting aspects of Anthropic Model 2 is that it is not simply a research prototype created for benchmarking.
Anthropic says both Model 2 and Mythos 5 are heavily used inside the company for coding, research, engineering and data generation. They are also being used in agentic workloads, including persistent AI-agent deployments.
This means Anthropic is gaining practical experience with the system before deciding whether it should eventually be offered to customers.
Internal use can also provide valuable information about how an advanced model behaves during long-running tasks. Instead of evaluating the system only through isolated prompts, employees can observe its performance across broader workflows involving software development, research and other technical activities.
Why Anthropic Is Not Releasing It Yet
Despite its stronger benchmark performance, Anthropic Model 2 is not currently scheduled for public release.
The company says it has not completed its standard full set of predeployment assessments. As a result, Anthropic has less confidence in its understanding of Model 2’s capabilities than it does with systems that have gone through the complete evaluation process.
This approach reflects the growing importance of safety testing as AI models become more capable.
Anthropic also reported that its internal deployment review did not identify new or more concerning forms of misalignment compared with issues already observed in Mythos 5.
That does not mean the model has been declared completely risk-free. Rather, the assessment suggests that the company has not identified a fundamentally new category of concerning behavior in its current internal testing.
No Commercial Name Yet
There is currently no indication that Anthropic Model 2 will eventually become Claude Mythos 6 or receive another specific commercial name.
The company may continue using the system internally, conduct additional safety evaluations or eventually introduce a related model to customers. Its eventual product strategy remains unclear.
The decision also highlights how frontier AI development is increasingly happening behind the scenes. Companies are building and testing systems internally before making them available publicly, allowing them to evaluate performance, safety and practical usefulness before wider deployment.
What Model 2 Means for AI Development
The emergence of Anthropic Model 2 provides an interesting snapshot of where frontier AI is heading. Its 62.8% CoBench v2 score shows meaningful progress over Mythos 5, particularly on complex research and engineering problems, but it also demonstrates that substantial limitations remain.
The fact that Anthropic is already using the model for coding, research and persistent AI agents suggests that advanced systems are becoming increasingly integrated into the development process itself.
Model 2 remains an internal system rather than a public Claude product. Its eventual release, if Anthropic decides to proceed, could provide a much clearer indication of how much the company’s next generation of AI systems has advanced beyond Mythos 5.



