Qwen 3.5 do I go dense or go bigger MoE?

r/LocalLLaMA
Generative AI AI Hardware Open Source AI

I have a workstation with dual AMAd 7900XT, so 40gb VRAM at 800gb/s it runs the likes of qwen3.5 35b-a3b, a 3-bit version of qwen-coder-next and qwen3.5 27b, slowly. I love 27b it’s almost good enough to replace a subscription for day to day coding for me (the things I code are valuable to me but not extremely complex). The speed isn’t amazing though… I am of two minds here I could either go bigger, reach for the 122b qwen (and the nvidia and mistral models…) or I could try to speed up the 27b, my upgrade paths: Memory over bandwidth: dual AMD 9700 ai pro, 64gb vram and 640 GB/s bandwidth.