Podcast: Play in new window | Download
Subscribe: Apple Podcasts |
I realize I am dating myself to some extent, but my first cell phone plan came with a fixed number of minutes and a small allotment of text messages, and anything beyond that was billed one at a time. You got in the habit of watching the clock on calls and thinking twice before you replied to a text. That model didn’t last. Once enough people wanted to use their phones without doing arithmetic, the carriers moved to unlimited plans, and now nobody counts anything.
AI is still in the counting stage. A company signs up with OpenAI or Anthropic, people start using it for real work, and then the invoice arrives, and someone has to explain how many tokens the team went through last month. Richard Luna, CEO of Protected Harbor, joins me on this episode of the TechSpective Podcast to talk about where that pricing model is headed, and why he thinks the better option for a lot of businesses is to run AI on hardware they own.
Who Pays for the Build-Out
My assumption going in was that token metering will eventually give way to something closer to unlimited use, because people won’t lean on a tool they have to ration. Richard comes at it from the other side. He points out that “no AI company is making money,” while enormous amounts of borrowed money are going into new data centers. Sooner or later that debt shows up in what customers pay.
If he’s right, the prices businesses pay today are an introductory rate, and nobody has said yet what the regular rate will be. Richard compares the moment to the subprime lending run-up before 2008 and to the dot-com era. The dot-com comparison is a useful one, because the bubble burst and the web kept going, while plenty of the companies that looked permanent in 1999 did not. Richard doesn’t doubt the technology itself. As he puts it, “AI is real, and its benefits are real.” The open question is who will be selling it, and at what price, once the correction happens.
The Case for Running It Yourself
Richard’s answer is local models. His view is that a model running on your own machine is the best way for a business owner to use AI, and the best way for a developer to “own their own destiny.” The data stays in-house, and the cost is tied to hardware you already bought rather than to a vendor’s need to earn back its investment.
The hardware side of that argument keeps getting stronger. My first computer was a Commodore 64 with 64K of memory. The Surface Studio laptop I use now has 64GB, and I’ve already downloaded a local LLM to run on it. Hugging Face has a huge catalog of open models to choose from, and some of the models coming out of China run on far less hardware than the big U.S. platforms use.
My kids object to AI for a number of reasons, including the water and power that data centers consume and the IP theft and general lack of ethics demonstrated by the big AI companies. Those are exceptionally valid concerns, but a few months ago Bruce Schneier shared some insightful wisdom that I keep referring back to: most complaints about AI are really complaints about the companies behind it. Running an open model on your own laptop doesn’t answer every one of those objections, but it takes a lot of them off the table.
Planning Around the Bill
None of this means businesses should drop cloud AI tomorrow. The frontier models are still better at some tasks, and plenty of teams don’t have the hardware or the in-house skills to run models themselves. It does make sense to track what you’re actually spending on tokens and to figure out which workloads could run locally before a price change forces the question.
Richard and I cover a lot more ground in the full conversation, including how he went from punch cards and a bulletin board system in his grandparents’ house to running an engineering-heavy company, and where he expects today’s AI giants to be in 20 years. Watch or listen to the full episode of the TechSpective Podcast below, and let me know whether you’re running AI locally yet.
