Low to midrange systems (8-32 GB) vs. free cloud tiers
I'm a noob to local LLMs. I want to use an LLM to create Python and Bash functions from well-defined specs written by me. I've used Claude's Sonnet for this mostly.
I'm a light user of LLMs and almost never hit the request limit on Claude.
It seems like cloud LLMs are in a price war right now and last-gen capabilities are bottoming out in cost? Is that correct?
I pay about $0.12 per kWh, probably going up as more AI data centers get built.
Hardware I have:
- <16GB VRAM - AMD BC-250 APU - "16 GB total with approximately 12 GiB assigned to GPU UMA and 4 GiB left for the OS" (hardware unlocking changes by the day)
- 8GB RAM - 11th gen Intel laptop with Xe graphics
- 8GB RAM - Pixel 7a
- 32 GB DDR3 - 2nd gen intel - doubt this does anything
Hardware I'm eventually selling:
- 8GB VRAM RTX 3070 + 32 GB DDR5 + 1TB NVMe SSD - AMD Ryzen 5 7600X CPU - putting this here in case it's substantially better than the BC-250
Among this hardware, I should look for a model that fits on the BC-250?
5 replies
I'll add my vote to keeping the 32+8 system as that has the most ram.
I also have an 8GB RTX 3070 Ti and 32GB RAM that I run Qwen3.6-35B-A3B perfectly on (30+ t/s) and looking forward to a possible Qwen3.8 35B update.
Recently bought a used 8GB RTX 3070 for a bargain, planning to combine it with the 3070 Ti and hopefully run a lower quant Qwen3.8-27B which is crazy good for its size.
Yes this can run Qwen 3.6 35b-a3b pretty nicely! And they might be releasing an updated version of that soon. Your BC-250 only has 16GB total which is not enough for 35b.
I also have 32GB RAM and 8GB VRAM, my computer is a little slower than yours, see my guide: https://lemmus.org/post/24235317
For the BC-250 you might try smaller models like Ling 3.0 Tiny, Ornith 1.5 9b, or Gemma 4 12b QAT
For your 8GB RAM devices, you can run Gemma 4 e4b QAT, Qwen 3.5 4b, or maybe Ling 3.0 Tiny
Use Unsloth Desktop or llama.cpp. Then you can connect Zoo Code to it, that's a VSCode extension which I like for programming with my local LLMs.
Qwen3.5-35B-A3B is the most powerful model you could run on your sell-eventually desktop, because it is MOE (mixture of experts) you can split its 22GB 4bit variant over both CPU and GPU and achieve workable performance. Use llama.cpp-cuda directly to save overhead, Claude can even help you set up a performance tested start script for your configuration. Hopefully Qwen3.8 will drop soon in a similar variant. On the other hand, deepseek 4 flash on openrouter.ai beats it easily and is energy cost level cheap. I would take that route over smaller models on programming duty on your weaker hardware.
Qwen 3.8 27B is about on par with Sonnet 5 for coding according to benchmarks and anecdotally.
I had a 3070 that I traded in for a new 7900 XTX at a final price of around $730. It’s about the most sanely priced 24GB GPU available in 2026 and runs Qwen 3.8 27B at around 30 t/s decode with MTP on llama.cpp for me.
Otherwise, 3.6 35B worked fine in my 3070, though it made quite a bit more mistakes than 3.6 27B. Guess it’s more like Haiku.
It's funny the per KWh going up prediction just made me think that buying more solar would be better now then buying RAM.
Sorry for the off topic. I wish I had an opinion but the cloud models seem so random in quality I get out of them it's hard to say for me, and almost change which model I run on a per project basis.
I will say for through put I do like vLLM better when I needed it and so I try to stick to models that can work on both that and Ollama.