IA al Día
Back to archive
Tools news Sep 1, 2026 2 min read

Hugging Face releases 207 WebGPU kernels to speed up AI in the browser

Hugging Face published 207 WebGPU kernels as individually benchmarked Hub repositories under Apache-2.0, with Transformers.js integration still to come.

primary source · huggingface.co — Introducing @huggingface/kernels: 200+ WebGPU Kernels for Local AI

A model running locally in the browser is only as fast as the GPU operations it dispatches. Hugging Face’s WebAI team, Nico Martin and Joshua “Xenova”, turned that layer into a public, benchmarked release on September 1: @huggingface/kernels, an npm package and Hub collection of 207 WebGPU kernels — the matrix multiplications, normalizations, attention operations, quantizations and convolutions a model needs to run on a browser’s GPU.

One Hub repository per kernel

Every kernel ships as its own repository in the webgpu-kernels organization on the Hub, carrying a manifest that defines the operation’s contract, provenance metadata, correctness tests, benchmarks and parameterized .wgsl.jinja shader templates. That structure is the point: instead of one generic implementation baked into a runtime, each operation becomes a separate, inspectable unit that a runtime can pick per device. The whole collection is Apache-2.0, and the npm package installs as a preview, @huggingface/kernels@preview.

2.57x, on less than half the cases tested

Hugging Face benchmarked the kernels against ONNX Runtime Web’s WebGPU backend on an Apple M4 GPU. Of 1,756 test cases across the 207 operations, only 809 produced matching outputs on both sides with reliable timing, and the reported numbers cover only that subset. Across those 809 cases the kernels won 629, lost 176 and tied 4, for a 2.57x speedup by geometric mean and 1.90x at the median. Results varied by operation: Add reached 3.52x, LayerNormalization 2.22x, Softmax 2.11x, and MatMul a comparatively modest 1.14x. Hugging Face also notes that results shift with GPU, browser and driver, and that for small operations the round trip to the GPU can cost more than the computation itself.

Still a preview, with Transformers.js to follow

The kernels are not yet wired into Transformers.js, Hugging Face’s in-browser inference library; that integration is planned, and the team says it is also working to upstream improvements to ONNX Runtime Web itself. Only Apple M4 numbers are published so far; Hugging Face points to Fleet, a crowdsourced benchmarking tool, to collect results from other GPUs. The figures describe individual operations, not a full model, so how much faster an actual browser-based model runs is not yet measured. For local, private inference — a model that runs entirely in the browser sends nothing to a server — Hugging Face has not yet published that end-to-end figure.