The Real Math of On-Premise AI and Open Models: Tokens Got 60 Times Cheaper, So Why Can't Enterprises Scale?
Open models now carry a large share of real-world usage, inference prices have fallen roughly 60 times in four years, and more companies are weighing a move of AI back into their own data centers. Yet Mozilla's September report shows open models still reach production 12 points less often than closed ones, with the gap widening as companies grow. Surveys from Taiwan's MIC and of local SMEs point the same way: what blocks AI from scaling is not the model or the budget, but data integration and operational capability.