A small model as the first stop
With Laya and Jev, something happened to me that hadn't happened with a model in a long time. I thought: this is exactly what I was trying to achieve. In our Gateway, we had set up…
When a model charges just a few dollars per million input tokens, it can look cheap at first glance.
What is much less obvious is everything hiding underneath that number.
Input cost is only one part of the story. In many cases, output generation is far more expensive. And as soon as the model starts explaining, detailing, or producing longer responses, the bill scales quickly.
Classification is not the same as reasoning. And in the end, much of the real value is precisely in the output.
If you also operate in Europe, handle sensitive data, or work in environments with still-maturing regulation and compliance requirements, the real cost becomes even higher. Not only because of the model itself, but because of everything you need around it to make it viable in production.
And even then, current pricing still does not look like the real cost of the system. It looks more like subsidized access.
During this phase, many companies are building as if token pricing were stable and predictable. As if the current cost structure were a constant. It simply is not.
In production, a model does not answer once. It reasons, retries, corrects itself, expands context, and chains workflows. Every iteration creates more text, more consumption, and more cost. Not because you sent much more, but because the system is doing much more work.
That is one of the most dangerous points: a lot of products being released right now are little more than polished interfaces sitting on top of a model call. While the model remains cheap, that works. But once output starts to dominate the bill, margin disappears.
Adopting AI without real direction is a serious risk. And it will probably be discovered in the worst possible way: with systems already in production and very little room left to correct course.
There is also another layer that still receives too little attention in many forecasts: the political and social pressure around automation. If AI displaces work at scale, it is reasonable to expect taxes, limits, or compensation mechanisms. And that is before we even consider the energy and natural-resource cost of this computational demand.
In the end, all of this comes down to something simple: the current price is not the real price.
It is an entry phase.
The important question is not whether it will change. The important question is whether what you are building will still hold when it does.