ForgeNEX

The end of loyalty to a single AI model: what Musk and SMBs share

Elon Musk announces that Grok Bot will use the best model for each task, even Claude. This 'model triage' strategy is already being applied by SMBs to control costs.

Elon Musk has published an important note about Grok Bot, the AI agent that SpaceX launched in beta in August. In it, he announces that SpaceX will use the best backend model for each task, including Claude Opus 5.5, MidJourney, Suno, and other leading APIs. Grok Bot is a joint product of SpaceXAI (Musk's AI lab, now integrated into SpaceX) and Cursor, whose acquisition for $60 billion in stock closed in August. This move reflects an unstoppable trend: loyalty to a single AI model is dead.

The news, originally published in The New Stack, is not an isolated case. At the ScaleUp:AI event by Insight Partners, held this week in New York, 22 leaders of portfolio companies were interviewed. At least 10 of them described the same practice: choosing the model based on the task and reserving the most expensive one for what really needs it. This strategy, known as 'model triage,' is becoming the standard for controlling AI costs.

Illustrative detail: Elon Musk’s Grok Bot will pick Claude over Grok when it’s better. One-model loyalty is dead.

Model triage: how companies control AI spending

The most common pattern that is taking hold is that of a funnel. A security company, for example, first runs a rules engine over all data, and only what survives goes to small models. Large models only see what remains. An executive at that company noted that processing a petabyte of data with any model, even a small one, would cost millions of dollars. Another company routes requests to its agent through a cheap model to find out what the user wants before the expensive model does anything.

Cost is the main reason. One executive commented that his company started with unlimited budgets for AI and now wonders what they have actually bought with all those tokens. Another said that his employees resorted to the most advanced model even when it was overkill, so now they teach staff which model fits each job. A third summed it up bluntly: labs earn more when you spend more tokens.

Data from OpenRouter, a service that allows developers to access hundreds of models through a single connection, confirms this trend. In the week ending October 7, four of the ten most used models were 'Flash' versions, the cheap and fast variants that labs release alongside their flagship ones. Anthropic has just made its own economy tier cheaper: Haiku 5.5 launched on Wednesday at $0.10 per million input tokens, a 90% cut for requests under 100,000 tokens. These small models usually handle high-volume jobs such as classification and routing.

Risks of switching models: the lesson of evals

But switching models is not without risks. An engineering leader recounted that his team switched to a newer model three days before a demo for an important client because benchmarks said it was cheaper and just as good. The workflow broke and they had to revert the change over a weekend. His conclusion: 'evals' are needed, sets of tests that would have detected the problem before they found it manually. The New Stack documented a similar case last month, when Anthropic made Opus 5.5 20% cheaper than Opus 5 and broke four things that agents depend on.

The moral is clear: test before switching and follow the ABS rule (Always Be Switching).

What it means for an SMB or an IT team

For an SMB, the lesson is that you should not marry a single AI provider. Musk's strategy of choosing the best model for each task is applicable at any scale. At ForgeNEX we recommend:

  • Analyze your AI tasks: identify which are high-volume and low-value (classification, routing) and which require a powerful model.
  • Implement a funnel: use small models or rules to filter first and reserve the large ones for what is essential.
  • Establish evals: before switching models, test with real data to avoid surprises.
  • Train your team: make sure they know which model to use in each case and why.
  • Monitor costs: periodically review token spending and adjust.

This model triage strategy aligns with other trends we have analyzed, such as the need to manage the risks of AI agents and the speed paradox in security. Flexibility and continuous evaluation are key in an ecosystem where models evolve every week.

In short, loyalty to a single model no longer makes sense. The companies that are best controlling their costs and results are those that adopt a pragmatic approach: the best model for each task, with constant testing and a mindset of continuous change. As one executive said, labs win when you spend more, but you win when you spend better.

Source: The New Stack. Analysis and adaptation: ForgeNEX.

Keep reading