Frontier model watch: OpenAI, Anthropic, Google, xAI and DeepSeek
The best model is now a procurement question, not a leaderboard answer. Context, tool use, governance and cost per completed job matter more than a single benchmark.
The market has split into operating systems
ChatGPT, Claude, Gemini, Grok and DeepSeek are becoming work surfaces with files, tools, memory and agent behaviour. A team choosing between them is choosing how research, drafting, coding and approval move through the business—not simply which chatbot writes the cleanest paragraph.
Where the leaders differ
OpenAI remains the broad generalist; Anthropic is particularly persuasive for long documents and careful writing; Google is strongest when Workspace is already the centre of work; xAI is differentiated by real-time public conversation; DeepSeek is a cost challenger that warrants additional governance review. Availability and packaging change, so HubAI scores the buying job rather than declaring a permanent universal winner.
A practical evaluation
Build a 20-task test set from real company work. Include easy tasks, ambiguous requests, sensitive data boundaries and one deliberate failure case. Record completion rate, human correction time, citations, refusals and total usage cost. Run the same set again after a major model update.
Verify the underlying guidance.
HubAI separates editorial analysis from official guidance. Product capabilities, availability and prices can change after publication.
UK ICO guidance on AI and data protection ↗UK Government AI regulation approach ↗