LFM2.5-2.6B: small and capable local AI model
Liquid AI releases LFM2.5-2.6B, a 2.6-billion-parameter model designed for agentic tasks that runs locally on devices including phones. The model undergoes four-stage training: supervised fine-tuning, teacher specialization, multi-domain on-policy distillation, and agentic reinforcement learning. Benchmarks show it competes with models nearly four times larger and has day-one support across major inference frameworks including llama.cpp, MLX, and vLLM.
Large frontier AI models are already so intelligent that many users waste their capabilities on tasks they could easily handle. New small local models are arriving that run even on mobile phones. The era is approaching when smart AI routers will efficiently dispatch tasks between local models and those running in the cloud.
Why are small local AI models more efficient than large cloud-based models?
Local models eliminate per-token costs and enable massive parallelization on local hardware at no additional expense. They can run continuously in the background for routine tasks where large models waste computational capacity.
What are the main benefits of agentic models running offline?
They provide zero latency, independence from network connectivity, complete data privacy, and ability to run on resource-constrained devices. They can perform complex multi-step workflows without calling cloud services.
Where does LFM2.5-2.6B fall behind larger models?
Primarily in coding and complex math problems, where smaller models have less computational capacity. For these domains, larger models remain the better choice.
- GLM-5.3 from Z.ai leads in vulnerability detection and will soon be available for download — z.ai 81 % match
- AI agents develop roles and compete over code in coordination experiment — anthropic.com 78 % match
- Harness Engineering for Self-Improvement — lilianweng.github.io 77 % match
- LFM2.5-2.6B
- Liquid AI
- Hugging Face12
- llama.cpp
- MLX3
- vLLM3
- SGLang
- OpenClaw
- Hermes Agent