Qwen 3.6 27B: MTP Benchmark on a 24 GB GPUCan Multi-Token Prediction double inference speed on a 24 GB GPU?...