Can Transformers Run Llama.cpp Quants?
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Can Transformers Run Llama.cpp Quants? on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Hugging Face has added native GGUF support to its transformers library, letting users run llama.cpp-style quantized checkpoints through the familiar from_pretrained API. The feature reuses llama.cpp’s ggml kernels and currently targets Apple Silicon and the Qwen3.5 architecture, available on the main branch ahead of a stable release.

Hugging Face has added native GGUF support to its transformers library, allowing users to run llama.cpp-style quantized checkpoints directly through the familiar from_pretrained API on their own machines. The feature reuses llama.cpp’s underlying ggml kernels to keep performance close to llama.cpp itself, with initial support focused on Apple Silicon and the Qwen3.5 architecture. The capability is currently available on the transformers main branch, ahead of the next stable release.

According to Hugging Face’s announcement, users can pick any GGUF checkpoint from the Hub, load it by passing a gguf_file argument to from_pretrained, and generate text with no extra configuration. When weights stay packed on Metal, transformers automatically loads compatible ggml/Metal layer kernels and uses ggml-org/ggml-attn as the attention implementation. If that kernel cannot be fetched, the model falls back to the standard sdpa attention with a warning, and users can force sdpa explicitly via attn_implementation=”sdpa”.

The requirements are specific: an Apple Silicon Mac, a PyTorch version supported by the published ggml-quantization kernel builds (usually the two latest releases), and the latest version of transformers plus a compatible version of the kernels library. Without a compatible quantization kernel, the loader falls back to dequantizing the model, which uses more memory.

Beyond direct model loading, the same GGUF checkpoints can be served through transformers serve, which exposes an OpenAI-compatible API on localhost. Clients such as Jan or Pi can connect by adding a custom OpenAI-compatible provider pointing at that endpoint. Hugging Face states that its reference for local inference performance is llama.cpp, and its benchmark comparison covers three GGUF checkpoints: a small dense model, a larger dense model, and a mixture-of-experts model.

At a glance
announcementWhen: announced ahead of the next stable tran…
The developmentHugging Face announced that its transformers library can now run GGUF quantized checkpoints directly through from_pretrained, reusing llama.cpp’s ggml kernels for performance.
At a glance
announcementWhen: announced April 2026; available via tra…
The developmentHugging Face announced that the transformers library can now run llama.cpp-style GGUF quantized models natively, using ggml kernels for near-llama.cpp performance on Apple Silicon.

Two Local AI Ecosystems Merge

This development matters because it collapses two previously separate ecosystems. GGUF, developed by the llama.cpp team, has become the dominant format for local inference — it powers tools like Ollama, LM Studio, and Jan, and GGUF models have been downloaded millions of times. Until now, running those checkpoints generally meant using llama.cpp-derived tools rather than the PyTorch-based transformers stack.

For developers already building on transformers, this means access to the full range of quantized checkpoints published by Unsloth, LM Studio Community, bartowski, and ggml-org without changing their code. For users with limited hardware, quantization lets a model like Qwen3.5-4B shrink from 8.42 GB in BF16 to 2.74 GB in Q4_K_M, making laptop-scale inference practical.

Hugging Face frames the move as making local AI “much easier” for everyday use, a claim echoed by growing interest in local coding agents running mid-sized models on consumer Macs.

Amazon

Top picks for "transformer llama quant"

As an affiliate, we earn on qualifying purchases.

How GGUF Quantization Fits Together

GGUF packages model weights and metadata — including tokenizer information and an optional chat template — in a single file. It supports multiple quantization levels, letting users trade precision for memory footprint. Variants such as Q4_K_M use mixed tensor precision: mostly 4-bit weights while keeping sensitive tensors at higher precision.

Hugging Face’s published file sizes for Unsloth’s Qwen3.5-4B illustrate the tradeoffs: BF16 at 8.42 GB as the unquantized reference, Q6_K at 3.53 GB, Q5_K_M at 3.14 GB, and Q4_K_M at 2.74 GB. The company recommends starting with Q4_K_M and moving to Q5_K_M or Q6_K if more memory is available, while cautioning that the quality cost of aggressive quantization depends on the model and task — users should evaluate on their actual workload.

The timing follows a period of rapid improvement in local inference. Hugging Face co-founder Julien Chaumond recently posted a demonstration of Qwen3.6 27B running inside the Pi coding agent via llama.cpp on a MacBook Pro, writing that for non-trivial tasks on Hugging Face codebases it felt “very, very close” to hitting the latest Claude Opus.

“We’re adding support for running GGUF models efficiently in transformers, so you can use checkpoints sized for your laptop’s memory through the familiar transformers APIs.”

— Hugging Face announcement

Limits of the Current Rollout

Several boundaries remain. Support currently targets Apple Silicon only — there is no stated timeline for CUDA, Linux, or Windows support. Architecture coverage starts with Qwen3.5; it is unclear which additional model families will be added or when.

The feature is also only on the transformers main branch for now, and Hugging Face has not announced a date for the next stable release that would include it. Full benchmark numbers comparing transformers-GGUF performance against llama.cpp across the three test checkpoints were referenced in the announcement, but details depend on the specific hardware and models used, so exact performance parity claims cannot be verified without those specifics.

Hugging Face itself notes that quality loss from more aggressive quantization depends on the model and the task, and advises users to evaluate on their own workloads.

Roadmap for Broader Device Support

The immediate next step is the feature shipping in a stable transformers release, removing the need to install from GitHub. Beyond that, Hugging Face’s stated initial focus on Apple Silicon and Qwen3.5 implies likely expansion along two axes: additional hardware backends (such as CUDA GPUs) and additional model architectures, including the mixture-of-experts models already used in its benchmarking.

Users can track the Hub’s GGUF documentation for newly supported quantization types and the kernels library for backend coverage as it develops.

Key Questions

Can transformers now run any GGUF quantized model?

No. Initial support targets the Qwen3.5 architecture on Apple Silicon Macs. Other architectures and hardware backends such as CUDA have no stated timeline yet.

How do I load a GGUF checkpoint in transformers?

Pass a gguf_file argument to from_pretrained with a GGUF checkpoint from the Hub. Generation then works with no extra configuration, provided you have a compatible PyTorch version, the latest transformers, and the ggml kernels library.

What happens if the ggml attention kernel is unavailable?

Transformers falls back to the standard sdpa attention with a warning. Users can also force sdpa explicitly via attn_implementation=”sdpa”. Without a compatible quantization kernel, the loader dequantizes the model, which uses more memory.

Which quantization level should I choose?

Hugging Face recommends starting with Q4_K_M and moving to Q5_K_M or Q6_K if memory allows. For Qwen3.5-4B, that corresponds to roughly 2.74 GB, 3.14 GB, and 3.53 GB respectively, versus 8.42 GB for BF16.

Can I serve GGUF models through an API with transformers?

Yes. The same checkpoints can be served via transformers serve, which exposes an OpenAI-compatible API on localhost that clients like Jan or Pi can connect to as a custom provider.

Primary source: Hugging Face · via ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Path To Self-Engineered Agent Harnesses In LLMs: ByteDance Seed’s Insights

ByteDance Seed’s HarnessDev project tests whether large language models can autonomously design their operational scaffolding, revealing significant generalization challenges.

Show HN: Make Cursed Fonts Like Times New Bastard

A new joke tool on Show HN uses OpenType ligatures to create intentionally cursed fonts like ‘Times New Bastard,’ gaining rapid attention among developers and designers.

Asustek Computer Surges In Global Coverage

Asustek Computer is experiencing a surge in international media coverage, with 26 mentions in recent reports, reflecting increased global interest.

Show HN: I Wrote A BASIC Interpreter That Boots On UEFI Machines

A developer has released Thoreau BASIC, a small BASIC interpreter that boots directly on UEFI systems, enabling vintage-style programming on modern hardware.