I posted already about me selfhosting LLM models. Right now I have two devices:
- Framework Desktop
- Custom Ryzen Threadripper server with R9700 32GB GPU
Both are decent with llama.cpp. The biggest issue is prefill times.
I have from work subscription for Claude and Codex which I tested too for last couple of months. It is OK but lack of thinking traces really hurt to keep up with what models are doing.
I decided to test some bigger open source models and I have read few good things about Fireworks.AI and I decided to register and test few models.
I started with just few dollars and tested few models.
I decided that good test will be to build something that I wanted to do for a long time… but I never had a chance because of complexity and lack of time to dwell into the topic: simple manager for private Fdroid repository. Managing this by hand is time consuming. I did that inside the cli, but again it was buggy and debugging it was time consuming.
So armed with the couple of prompts and energy I burned through all of my initial 10$!
And solution was totally broken and nothing was working.
It was a bit skill issue because I just gave it a vague prompt and left it working for couple of hours.
Then I decide to rework it from ground up, step by step. Few plans after that. I had to add another $10.
And then another to continue.
And another ten $ to finish it up… which happened to be not enough. So I had top it up one more time with another ten dollars.
After few another plans and few bucks burned I finally had working solution. And I had 36$ less.

Is 36$ much for new applications that solves your problems. Certainly not if this for longer time.
Is 36$ much for 3 days of checking here and there what agent is doing. Also not.
Is 36$ much for just 3 days of usage? It combines to about 230-240$ dollars each month. And it was pretty cheap model:

If I would be using more costly model, i.e. Kimi 3 that is priced around 15$/M it would be probably closer to 100$.
So this would mean cost around 700-1k$ each month. This is a lot for a tool.
On the other hand I spend about 5k$ for new hardware to run models (R9700 and Framework Desktop + some disks and new fans). Token generation even if they run almost all the time is not great either. For Frameworks Desktop for about day and a half it generated:

Which amounts for (considering that I am using Qwen 3.8 right now, which is priced around:

Table napkin math says that i would amount for 5$ a day about 105$ a month.
Similar graph for R9700:

It is also for about 1 and a half days of usage. This totals to about 10$, which would be 6,9$ which would mean 145$ each month.
This means that I can use those devices equivalent amount of 250$ a month. Is it much? Does not seems so. Seems like I will brake even in about 20 months. If the devices will lives through 24/7 usage for that time. Many things may brake. Software stops working, kernel panics, out of memory evictions or just plain no electrical power. Which probably means more of about 2 years worth of subscriptions.
But!
200$ subscription would probably buy me more output tokens then lousy 1,5M!
So probably more. It is hard to tell how much more I can get out of Claude or Codex or some Z.ai subscription for 200$ dollars.
Anyway 36$ dollars for software that I will be using for few years seems like a good deal.
5k$ for equipment that is hard to use and noisy and requires a lot of upkeep (because I did no account for electrical power cost!) is not worth it. Unless you want to learn few things. Then it is really good idea.
I am just a bit sad that I did not buy NVIDIA Blackwell. At he time 10k$ dollars seemed like crazy idea. Now it is closer to 20k$.
But I could run bigger models faster. And support for NVIDIA is soooo much better.

