Running a 35B MoE at 45+ tok/s on a $170 GPU: My Experience with FreeToken and Qwen3.6
A chronicle of pushing an 8 GB NVIDIA Quadro RTX 4000 with FreeToken and Qwen3.6-35B-A3B: patching Turing kernels, quant benchmarks, harness traps, and the final pragmatic choice.