The DeepSeek dense models are pretty obsolete, MoE models are where it's at now. If you have a 6 GB GPU you can get around 7 tokens per second, and more like 30 with a RTX 3060 12 GB.
https://huggingface.co/mradermacher/Gemma-4-26B-A4B-Preserving-Abliteration-i1-G...
https://huggingface.co/llmfan46/Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Pr...
Get the Q6 variants.
Personally I love the Deepseek model on their web site, so hopefully we'll see something in the 25B-100B total, 3B-15B active range from them soon. Their current v4 models are too large to run on consumer hardware though.
