xLAM

Confidence 0.80 · 3 sources · last confirmed 2026-09-04

Salesforce’s open-weight “Large Action Model” family — language models post-trained specifically for tool calling and multi-turn agentic interaction, released by Salesforce AI Research with weights and training data public on Hugging Face.

The xLAM-2-fc-r generation ships at 1B, 3B, 8B, 32B and 70B parameters on Llama 3.1/3.2 and Qwen 2.5 backbones, trained by filtered behavioural cloning on synthetic multi-turn trajectories produced by the APIGen-MT pipeline — see 2025-04-04-prabhakar-salesforce-apigen-mt-xlam-2 for the method and the numbers.

Why the wiki tracks it

xLAM is the corpus’s worked example of the small-language-models argument: an open, small, locally-runnable model that is genuinely competitive with frontier models on the narrow capability agents actually need.

  • The NVIDIA position paper cites xLAM-2-8B by name as evidence for its central claim: “achieves state-of-the-art performance on tool calling despite its relatively modest size, surpassing frontier models like GPT-4o and Claude 3.5.”
  • On BFCL v3 as of April 2025, xLAM-2-70b and xLAM-2-32b held ranks 1 and 2, and even the 1B model beat o1 and gpt-4o on the multi-turn column.
  • Sokolenko runs xLAM-2-32B at 4-bit quantization on a two-year-old laptop — the model chosen precisely because it is “not only open-source but also smallish” while sitting in BFCL’s top 20 of ~110.

The qualifier that travels with it

The advantage is narrow and dated. It is concentrated in BFCL’s multi-turn column (75.12 vs 41–47 for GPT-4o); on single-turn AST it is indistinguishable from frontier models because that category is saturated. On τ-bench overall it is beaten by Claude 3.5 Sonnet (new), Claude 3.7 and o1. And its April-2025 top-of-board position had slipped to roughly 18th by December 2025 as the frontier moved. See the source page for the full reading.

Appears in this wiki via