Friday, August 7, 2026

Liquid AI Releases LFM2.5-2.6B: An On-Machine Agentic Mannequin With 128K Context, Software Calling, And Open Weights


Liquid AI launched LFM2.5-2.6B, an agentic mannequin that runs totally on-device. It plans, calls instruments, and works via multi-step duties on telephones, laptops, PCs, and robots. The mannequin has 2.69B complete parameters, a 131,072-token context window, and a 128,000-token vocabulary. Pre-training used roughly 34 trillion tokens. Two checkpoints shipped: LFM2.5-2.6B-Base for fine-tuning, and LFM2.5-2.6B post-trained for agentic workloads. As a result of inference stays native, information by no means leaves the machine and the marginal price of every run is close to zero. Liquid AI reviews tool-use and instruction-following scores aggressive with fashions almost 4 instances its dimension.

Is it deployable

The reply is Sure. Each checkpoints are public on Hugging Face beneath the lfm1.0 license. Weights ship in native, GGUF, MLX, and ONNX codecs, with day-one assist in llama.cpp, vLLM, SGLang, and LM Studio.

  • Which corporations: Solo builders and startups can pilot on {hardware} they already personal. The mannequin decodes at 220 tokens/s on an M5 Max in beneath 2.5 GB. Mid-market groups can self-host on one GPU: a single NVIDIA H100 SXM5 serves roughly 1.3B tokens per day. Enterprises and OEMs can push the identical weights to machine fleets via GGUF and ONNX. High-quality-tuning is out there through LoRA with TRL and Unsloth.
  • Which industries: Liquid AI targets automotive, client electronics, industrial robotics, healthcare, monetary providers, e-commerce, and protection. Regulated and air-gapped settings profit most, since no immediate reaches a third-party API.
  • Functions: Liquid AI recommends agentic workloads, instrument use, information extraction, RAG, and long-context workflows. Sensible builds embody on-device assistants, offline doc triage over 128K inputs, type and bill extraction, robotics command parsing, and background brokers that run repeatedly with out per-token price. Liquid AI explicitly doesn’t suggest the mannequin for agentic coding or knowledge-heavy duties.

Structure and coaching funds

LFM2.5-2.6B has 2.69B complete parameters throughout 30 layers. The stack is 22 double-gated brief convolution blocks plus 8 grouped-query consideration blocks. Vocabulary dimension is 128,000 and context size is 131,072 tokens. Pre-training used roughly 34 trillion tokens.

Liquid AI doubled the vocabulary to 128K by extending the present tokenizer in place moderately than retraining from scratch. A devoted mid-training part extends context to 128K. The mannequin covers 16 languages and is text-only.

4-stage post-training

The bottom checkpoint turns into an agent via 4 phases.

  • First, two consecutive supervised fine-tuning rounds, with an SFT combine roughly seven instances the dimensions used for LFM2.5-8B-A1B.
  • Second, trainer specialization: one professional per area, skilled with reinforcement studying with verifiable rewards.
  • Third, multi-domain on-policy distillation, the place the scholar rolls out beneath its personal coverage and every immediate routes to its area trainer.
  • Fourth, agentic reinforcement studying with GRPO inside actual harnesses, together with Hermes Agent and OpenClaw.

Benchmarks

Liquid AI in contrast LFM2.5-2.6B in opposition to gemma-4-E2B-it (5.1B), gemma-4-E4B-it (8B), Qwen3.5-4B (4.7B) and Qwen3.5-9B (9.7B).

Benchmark LFM2.5-2.6B gemma-4-E4B-it Qwen3.5-9B
ToolSandbox 77.83 65.00 76.44
Multi-IF 80.07 77.35 62.55
IFStruct 85.49 76.65 78.50
IFBench 59.17 39.24 56.47
BFCLv4 56.88 46.39 60.13

It leads each instruction-following benchmark reported and almost each instrument use benchmark, trailing Qwen3.5-9B solely on BFCLv4. Coding is the place bigger fashions maintain an edge: LiveCodeBenchv6 is 59.41 versus 69.86 for Qwen3.5-9B.

Interactive explainer

Key Takeaways

  • 2.69B params, 30 layers (22 short-conv + 8 GQA), 128K context, ~34T coaching tokens.
  • Beats gemma-4-E4B-it and Qwen3.5-9B on ToolSandbox, Multi-IF and IFStruct.
  • 220 tok/s on M5 Max, 30 tok/s on telephone, beneath 2.5 GB reminiscence.
  • Open weights beneath lfm1.0, with GGUF, MLX and ONNX from day one.

Take a look at the Technical particulars, LFM2.5-2.6B, and LFM2.5-2.6B-Base. Additionally, be at liberty to observe us on Twitter and don’t neglect to affix our 150k+ML SubReddit and Subscribe to our E-newsletter. Wait! are you on telegram? now you possibly can be part of us on telegram as effectively.

Have to associate with us for selling your GitHub Repo OR Hugging Face Web page OR Product Launch OR Webinar and so forth.? Join with us


Asif Razzaq is the CEO of Marktechpost Media Inc.. As a visionary entrepreneur and engineer, Asif is dedicated to harnessing the potential of Synthetic Intelligence for social good. His most up-to-date endeavor is the launch of an Synthetic Intelligence Media Platform, Marktechpost, which stands out for its in-depth protection of machine studying and deep studying information that’s each technically sound and simply comprehensible by a large viewers. The platform boasts of over 2 million month-to-month views, illustrating its reputation amongst audiences.

Related Articles

Latest Articles