Hugging Face has released @huggingface/kernels, a library containing over 200 optimized WebGPU kernels for running AI models directly in web browsers. The release represents a significant step toward making local AI inference accessible to anyone with a modern browser.
The Browser AI Stack
Running AI models in browsers has long been held back by the need for efficient low-level operations. While WebGPU provides a portable API for GPU operations across browsers, achieving good performance requires carefully optimized kernels—individual GPU operations like matrix multiplications, normalizations, and attention primitives.
The challenge is that portability doesn’t automatically mean performance. Two shaders implementing the same operation can behave completely differently across different GPUs, browsers, and input shapes. The optimal choice depends on workgroup sizes, memory access patterns, vectorization, and fusion strategies.
207 Kernels and Counting
The initial release includes 207 WebGPU kernels covering operations used across a wide variety of machine learning architectures. Each kernel is published as a complete, versioned package containing:
- Kernel card: Documentation of the operation’s semantics, inputs, outputs, and supported data types
- Manifest: The source of truth defining the operation contract
- Correctness tests: Cases to verify the implementation produces expected results
- Benchmark cases: Workloads used to evaluate kernel performance
- WGSL shaders: Parameterized implementations for different devices and input shapes
All kernels are Apache-2.0 licensed and available at huggingface.co/webgpu-kernels.
Enter Fleet: Crowdsourced Benchmarking
Alongside the kernels, Hugging Face launched Fleet, a browser-based GPU benchmarking and testing suite. Fleet runs kernels on users’ actual hardware and collects performance and correctness evidence across real-world devices—a scale of testing impossible in conventional labs.
With user consent, each run contributes evidence that helps identify failures, improve kernel variants, and make better optimization decisions. The tool essentially turns GPU benchmarking into a community effort, leveraging the diversity of real-world hardware to improve browser-based AI for everyone.
Why This Matters
The ability to run AI models efficiently in browsers opens up several possibilities:
- Privacy: Inference happens locally, keeping data on the user’s device
- No installation: Users don’t need to install anything beyond a browser
- Cross-platform: Works on any device with a modern browser and WebGPU support
- Cost: No API calls or cloud inference costs
For developers, the kernel-first approach means optimizations can be made independently while maintaining stable contracts for higher-level runtimes. The published versions and clear interfaces make it easier to understand, test, and improve individual operations.
Getting Started
Developers can install the JavaScript loader via npm:
npm install @huggingface/kernelsThe library downloads, prepares, and runs kernels directly from the Hugging Face Hub. Combined with the Fleet benchmarking tool, developers can now both run and optimize browser-based AI inference with community-backed performance data.
This release positions Hugging Face as a key player in the emerging WebAI ecosystem, where local, privacy-preserving AI inference runs directly in users’ browsers.