Qwen 3.8 27B: powerful model, but default mode leads to wild overthinking
Qwen 3.8 27B is a powerful open-source model from Alibaba suitable for local deployment, but its default setting automatically engages in excessively deep analysis (xhigh reasoning mode). The author recommends running the model at lower reasoning levels to avoid wasting time on unnecessary computation.
It's nice to see how quantization and model size fundamentally affect output quality.
What are typical deployments for the Qwen 3.8 27B model?
Given its 27 billion parameters and 17GB GGUF version, the model runs well on adequately-resourced laptops or small servers. It is commonly deployed via LM Studio or llama-server.
What is the xhigh reasoning mode in Qwen models and why is it problematic?
The xhigh (extra high) mode increases the depth of the model's reasoning for complex tasks. While useful for analytical problems, it leads to significant time waste for simple requests like drawing a circle without corresponding benefit.
What vision capabilities does the Qwen 3.8 27B model demonstrate?
The model performs well on detection tasks, such as returning precise bounding boxes for objects in photographs in a requested scale. It shows competency in multimodal inputs including images.
- Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone — github.com 75 % match
- LFM2.5-2.6B: small and capable local AI model — liquid.ai 75 % match
- DeepSeek makes its cheapest model more powerful: DeepSeek-V4-Flash — huggingface.co 74 % match
- Qwen 3.8 27B
- Alibaba3
- Qwen
- LM Studio
- NVIDIA DGX Spark
- Simon Willison
- Qwen 3.6 27B
- Qwen 3.8 2.4T-A95B
- OpenRouter5