n o t
o n l y
t e c h n o l o g y
blog image

The Real Total Cost of Ownership (TCO) of AI Solutions

Roko Roić, Co-Founder

AI

AI

July 16, 2026

July 16, 2026

Based on my recent talk on TCO in the world of tokens, GPUs and new AI architectures.

Every conversation about AI in business eventually lands on the same question: how much will it cost to own the solution? The good news is that a precise answer exists. AI costs can be calculated and managed like any other technology investment. We developed this approach at Notch, grounded in actual token costs from live production systems.

In medias res

Let’s set the stage with some real-world unit economics. Tokens are the new player in town, and I want to start by showing you how they play out in real life. The following are actual numbers from production systems. Each unit of work burns some tokens. This is how much:

  • Preparing a new CV: About 0.07 Euros with OpenAI
  • Building a full course or lab: About 10 Euros per course with Anthropic
  • Creating retail product data from multiple sources: 0.14 to 0.18 USD per product with OpenAI
  • Run a certification process on an employee: 0.94 USD per certificate with OpenAI and Anthropic

Ok, there you have it, actual real-world numbers. Token costs sound abstract until you attach them to real work. These numbers are actual unit economics. Now let’s put them in the big picture.

 

What is the Total Cost of Ownership (TCO) in AI?

The total cost of ownership for an AI solution is the sum of four things: implementation, maintenance and retraining, software and infrastructure, and token usage. Three of these behave like any other technology cost. One, tokens, is new, and it’s more predictable than it looks.

The ROI game

The ROI formula is an industry classic

ROI = (gain from investment − cost of investment) / cost of investment

In the case of AI solutions, both sides of that equation can be broken down concretely.

The gain is huge

The move to AI is a move from automated to autonomous systems. The shift here is tectonic, and the gains are huge. Each case is a story on its own, but here are some categories where you might find your gains.

Lower direct costs. Process management costs and the cost of maintaining legacy systems go down.

Lower opportunity costs. Faster time to market, lower churn, and operations that run 24/7/365.

Revenue growth. New channels, faster fulfillment of customer requests, acquisition of new customers.

Lower cost of inefficiency. Fewer delays, fewer errors, fewer corrections.

There is a fifth category that is harder to measure but real: better data quality, stronger brand reputation, higher employee satisfaction, and better compliance. This might actually be the most important one for you, even if not measured in currency.

The cost is predictable

Ok, I showed you token costs from real life at the very start, but where do they fit in? They are just one part of the cost calculation, and you might be surprised that the other cost elements are quite familiar to you. This is still digital transformation after all. We are still building digital products, and AI-based solutions draw heavily on what we already know. You still need to build, maintain, and run. A story you know so well.

Cost Group Recurrence and stability
Implementation One-time and stable
Maintenance, retraining and adaptation Recurring and stable
Software and cloud infrastructure Recurring and stable
Tokens and models Recurring and somewhat elastic

Let’s tackle each one and understand how it affects your cost table.

Implementation cost

This is digital product development. You know it so well, and you know how to invest in it. In the case of AI solutions, key items are:

  • data preparation
  • model tuning or training
  • development of the agents and orchestrators
  • development of the application layer
  • testing and delivery

Managing this cost requires nothing exotic. It is a classic effort-estimate that results in a budgeted number. Any organization that has delivered software projects already has this muscle. It should be easy for you to calculate with your partners.

Maintenance, retraining and adaptation

Once the solution is live, the recurring items are monitoring and operations, workflow adjustments, upgrades, model retraining, and observability. Model retraining is the new one here. As models progress or die out, your token economics change. At that point, you might want to switch models. Or, in another case, you have gotten more data and can make your models better in reasoning, so you want to retrain them. In any case, it is just a number of working hours. You budget it, and it stays firm in the cost structure. Think of it as adapting your software solution through time. This is the adaptive maintenance you already know.

Anyway, this is a classic estimate of the work-required. It behaves like the maintenance budget of any other business system. We are still on familiar ground.

Software and cloud infrastructure

The building blocks here are agentic platforms, integration platforms, MCP servers and platforms, database licenses, application platform licenses and security platform licenses. To cut it short, these are software products you might need to integrate and run your AI solution on. You will either buy them off-the-shelf or use open-source versions. You have done this many times before with database systems or middleware, I am sure.

To complete the stack, these platforms need to run on infrastructure in private or public clouds. Again, business as usual.

Procurement again follows the usual process your organization already runs. Nothing about AI changes how you buy licenses and cloud services.

The Token: New player on the team

Token is a new cost item. The closest thing we have had to Tokens so far is cloud cost. Your cloud invoice goes up when usage goes up. While the unit price is the same, the total number changes.

This is exactly what we do with tokens: calculate unit prices for our meaningful work items. Keep those unit costs stable and allow only the total number to change if usage is up. Which is ok, because you usually want your usage to go up. It increases the gains as well. The TCO still works.

Token economics are abstract until you connect them to business outcomes, as we did at the start of this blog. Let’s give this topic a little bit of attention.

How to frame it at the start?

Check your regulatory constraints early. You might need self-hosted models. You might need EU-based or US-based models. This shapes everything downstream. This is your first question.

Then, decide how much you want to spend, based on your gains side in the ROI calc. A more expensive model is often better at reasoning, but does your business case have a budget for it? If not, you will shape your system to use cheaper or self-hosted models. Always use the Minimum Viable Model.

Plan for a good Agent orchestration platform. Caching, async processing, logging token costs and compressing context all reduce spending directly. Sovera AI™ might be what you are looking for. Or any other solution, but do not start without planning it in. You will have much better unit economics.

Lastly, plan for a hybrid setup from day one. What does hybrid mean in the case of Agentic solutions? It means that Agents can switch models as needed, either in real time or with a change in system configuration and some training. What this will do is substantially lower your risks of token costs rising. If a token provider sends their prices flying, you will be able to switch and keep the unit cost similar to what you planned and what works for your ROI.

How to control the cost before even committing

Each implementation will start with a validation phase for Agents’ work items. You will prototype your key agents and measure real consumption. Build your own unit economics. Treat the prototype phase as a “performance test” that happens at the start of the project instead of the end.

Once again, you will choose the Minimum Viable Model per agent. Every agent gets the cheapest model that does its job well.

If you are using frontier models, check model lifespans. Frontier models typically live 12 to 24 months. After that, they are replaced by either a new version of the model family or a completely new model. You might end up with a more token-hungry model after the end of life (EOL), which changes unit economics. That is not ideal.

Hopefully, model providers will start providing long-term models for business, but in this fast-paced development stage, you will use your wits to avoid the risk. Stay ready to switch. A hybrid approach gives you agility when a model is retired or repriced. Plan for retraining.

So how elastic is this cost then?

Firstly, tokens are just one part of your cost, as I have just shown. Do not overestimate their influence on overall price stability. But use the tricks above to make sure you control the unit economics. So, if you play your cards well, this is what you are looking at in terms of elasticity:

2 years of near 100% stability. The development phase and the first two years of operation are quite predictable. You will have your token economics fully under control.

3 years of predictable elasticity. Models get retired; token prices change. You retrain for other models and keep the unit cost where you want it. Plan it in the adaptive maintenance.

After five years? It stops mattering if your ROI lands within five years, and that holds true for most cases.

The honest answer to “how big is the elasticity?” is: as big as you allow it to be. Price it into your ROI calculation, and it becomes a managed variable instead of a fear.

Choosing a model type by cost-elasticity

Different hosting choices come with different levels of predictability.

Self-hosted open-weight models. Nearly 100% predictable in the long run.

Cloud-hosted open-weight models. Highly predictable, with small elasticity, same as any other cloud service.

Frontier models hosted by the model provider. Somewhat less predictable after two years. Your maneuvering room is to swap an agent’s model for a cheaper or self-hosted one when the time comes.

Hybrid approach. You define the elasticity yourself through the ratio of self-hosted, cloud-hosted, and frontier models, and you leave that room in the ROI calculation.

One more thing worth noting: open-weight models currently trail frontier models by roughly 12 months. In two years, when your elasticity window opens, they will be at the level your solution needs or beyond. Also, new model providers are popping up every month. Microsoft and Meta just released their frontier models for public availability. We already have Google, Anthropic, OpenAI and Mistral. The market will be competitive, and that should mean prices will be controlled by demand.

Final thoughts

I hope I have shown that the costs of AI business solutions can be calculated quite concretely. Token prices are transparent, and unit economics are clear. Any elasticity can be managed through model changes in the future, especially with a hybrid setup.

The risk of missing the opportunity is surely greater than the risk of controlling the cost.