Spyke

Syndicated from the fediverse. Read and engage on the original instance.

View original on lemmy.frozeninferno.xyz
localllama·LocalLLaMAbyBigHeadMode

Low to midrange systems (8-32 GB) vs. free cloud tiers

I'm a noob to local LLMs. I want to use an LLM to create Python and Bash functions from well-defined specs written by me. I've used Claude's Sonnet for this mostly.

I'm a light user of LLMs and almost never hit the request limit on Claude.

It seems like cloud LLMs are in a price war right now and last-gen capabilities are bottoming out in cost? Is that correct?

I pay about $0.12 per kWh, probably going up as more AI data centers get built.

Hardware I have:

  • <16GB VRAM - AMD BC-250 APU - "16 GB total with approximately 12 GiB assigned to GPU UMA and 4 GiB left for the OS" (hardware unlocking changes by the day)
  • 8GB RAM - 11th gen Intel laptop with Xe graphics
  • 8GB RAM - Pixel 7a
  • 32 GB DDR3 - 2nd gen intel - doubt this does anything

Hardware I'm eventually selling:

  • 8GB VRAM RTX 3070 + 32 GB DDR5 + 1TB NVMe SSD - AMD Ryzen 5 7600X CPU - putting this here in case it's substantially better than the BC-250

Among this hardware, I should look for a model that fits on the BC-250?

View original on lemmy.frozeninferno.xyz
19

5 replies

I'll add my vote to keeping the 32+8 system as that has the most ram.

I also have an 8GB RTX 3070 Ti and 32GB RAM that I run Qwen3.6-35B-A3B perfectly on (30+ t/s) and looking forward to a possible Qwen3.8 35B update.

Recently bought a used 8GB RTX 3070 for a bargain, planning to combine it with the 3070 Ti and hopefully run a lower quant Qwen3.8-27B which is crazy good for its size.

2

8GB VRAM RTX 3070 + 32 GB DDR5 + 1TB NVMe SSD - AMD Ryzen 5 7600X CPU - putting this here in case it’s substantially better than the BC-250

Yes this can run Qwen 3.6 35b-a3b pretty nicely! And they might be releasing an updated version of that soon. Your BC-250 only has 16GB total which is not enough for 35b.

I also have 32GB RAM and 8GB VRAM, my computer is a little slower than yours, see my guide: https://lemmus.org/post/24235317

For the BC-250 you might try smaller models like Ling 3.0 Tiny, Ornith 1.5 9b, or Gemma 4 12b QAT

For your 8GB RAM devices, you can run Gemma 4 e4b QAT, Qwen 3.5 4b, or maybe Ling 3.0 Tiny

I’m a noob to local LLMs.

Use Unsloth Desktop or llama.cpp. Then you can connect Zoo Code to it, that's a VSCode extension which I like for programming with my local LLMs.

8

Qwen3.5-35B-A3B is the most powerful model you could run on your sell-eventually desktop, because it is MOE (mixture of experts) you can split its 22GB 4bit variant over both CPU and GPU and achieve workable performance. Use llama.cpp-cuda directly to save overhead, Claude can even help you set up a performance tested start script for your configuration. Hopefully Qwen3.8 will drop soon in a similar variant. On the other hand, deepseek 4 flash on openrouter.ai beats it easily and is energy cost level cheap. I would take that route over smaller models on programming duty on your weaker hardware.

4

Qwen 3.8 27B is about on par with Sonnet 5 for coding according to benchmarks and anecdotally.

I had a 3070 that I traded in for a new 7900 XTX at a final price of around $730. It’s about the most sanely priced 24GB GPU available in 2026 and runs Qwen 3.8 27B at around 30 t/s decode with MTP on llama.cpp for me.

Otherwise, 3.6 35B worked fine in my 3070, though it made quite a bit more mistakes than 3.6 27B. Guess it’s more like Haiku.

3

It's funny the per KWh going up prediction just made me think that buying more solar would be better now then buying RAM.

Sorry for the off topic. I wish I had an opinion but the cloud models seem so random in quality I get out of them it's hard to say for me, and almost change which model I run on a per project basis.

I will say for through put I do like vLLM better when I needed it and so I try to stick to models that can work on both that and Ollama.

1

You reached the end