I’m a noob to local LLMs. I want to use an LLM to create Python and Bash functions from well-defined specs written by me. I’ve used Claude’s Sonnet for this mostly.
I’m a light user of LLMs and almost never hit the request limit on Claude.
It seems like cloud LLMs are in a price war right now and last-gen capabilities are bottoming out in cost? Is that correct?
I pay about $0.12 per kWh, probably going up as more AI data centers get built.
Hardware I have:
- <16GB VRAM - AMD BC-250 APU - “16 GB total with approximately 12 GiB assigned to GPU UMA and 4 GiB left for the OS” (hardware unlocking changes by the day)
- 8GB RAM - 11th gen Intel laptop with Xe graphics
- 8GB RAM - Pixel 7a
- 32 GB DDR3 - 2nd gen intel - doubt this does anything
Hardware I’m eventually selling:
- 8GB VRAM RTX 3070 + 32 GB DDR5 + 1TB NVMe SSD - AMD Ryzen 5 7600X CPU - putting this here in case it’s substantially better than the BC-250
Among this hardware, I should look for a model that fits on the BC-250?


It’s funny the per KWh going up prediction just made me think that buying more solar would be better now then buying RAM.
Sorry for the off topic. I wish I had an opinion but the cloud models seem so random in quality I get out of them it’s hard to say for me, and almost change which model I run on a per project basis.
I will say for through put I do like vLLM better when I needed it and so I try to stick to models that can work on both that and Ollama.