镜像站点 · 本页由第三方 GitHub 只读镜像提供,非 GitHub 官方站点,不接受任何登录或凭据输入。前往 github.com
Skip to content

Repository files navigation

1bit engine

Documentation: 1bit.gg · Measured results: wiki · Community: Discord

The 1bit engine runs inside Lemonade. Lemonade stays the server you talk to (its catalog, downloads, router and UI), and for the models it hands to 1bit, Lemonade runs the engine as one of its backends, the same way it runs llama-server. The engine serves each model behind an OpenAI-compatible API (1bit serve), whatever device runs it:

  • the XDNA 2 NPU engine
  • HRX on the Radeon iGPU (AMD's ggml-hrx with our Loom kernels), the default GPU route (docs/hrx.md)
  • Vulkan, while it is leaving the engine (RFC #213): --device vulkan still runs upstream llama.cpp's latest release, but --device auto now means HRX (docs/vulkan.md)
  • MoE experts streamed from the drive on Vulkan (1bit serve --moe-slots N), for MoE models larger than memory (docs/moe-streaming.md)
  • a lean option, ROCmFPX's ROCmFP4 and ROCmI4 formats: faster, less accurate (docs/lean.md)
  • ZINC, which also reaches NVIDIA GPUs (CUDA) and Apple GPUs (Metal)
  • DwarfStar, for DeepSeek V4 Flash, GLM 5.x and Qwen3.8-Flash-Next in its own GGUFs, with SSD expert streaming
  • MLX on Apple Silicon, through lemon-mlx-engine
  • ONNX Runtime GenAI models (Lemonade's ONNX format) on the CPU (docs/onnx.md)
  • Laya, which decides where each request runs
  • ComfyUI.cpp: ComfyUI workflows (Stable Diffusion 1.5 text-to-image and image-to-image) in C++, matching ComfyUI's output to 50 dB (docs/comfyui.md)
  • every Hugging Face model architecture, kept current by a daily census

Every tuned setting 1bit serve gives a backend is a recipe with its measurement attached (docs/recipes.md), and the serve numbers in these docs come from tools/bench.py, which A/B-measures configurations against a baseline in the same run (docs/bench.md).

Packages ship every Sunday, rebuilt at that week's upstream pins: Linux, Lemonade with the engine, Windows and 1bit OS (docs/releases.md). The first release ships on Sunday, 4 October 2026.

Direction: the engine is moving to HRX (AMD's ggml-hrx, kernels in Loom) plus the NPU.1 HRX met the decode gates in RFC #213 and is the default GPU route now; Vulkan leaves the engine in stages, as the features that still need it are ported to HRX or dropped.

Status: the engine runs inside Lemonade through 1bit serve (docs/lemonade.md, docs/serve.md); the NPU, Vulkan, HRX and ZINC each pass its end-to-end test on Strix Halo, and the Lemonade recipe that runs it (onebit, in our fork 1bit-MONSTER/lemonade) passes Lemonade's own LLM test suite on Vulkan and HRX. Following a review,1 the engine no longer vendors Lemonade: Lemonade is the host, 1bit is the engine inside it. Ported so far: HRX on AMD's live ggml-hrx (docs/hrx.md; its decode-split race, #123/#140, is fixed and the kernel is on by default; Q2_K, IQ2 and IQ3_XXS GGUFs run on HRX instead of the CPU, and --mtp on Qwen3.8-27B runs NaN-free since #257), Vulkan from upstream llama.cpp's latest release (docs/vulkan.md), the NPU engine on full ELFs with the upstream XDNA stack pinned (docs/npu.md; its layer kernel is not yet built from source), ZINC (docs/zinc.md) and MLX (docs/apple.md). ZAYA1-8B (Zyphra) runs from our llama.cpp on Vulkan, HRX and ROCm, matching transformers, at 93 tok/s decode in Q4_K_M on Vulkan, and ZAYA1-74B-preview at 35 tok/s (docs/vulkan.md); the rest of Zyphra's family (Zamba, Zamba2, BlackMamba) runs on Vulkan too, and so do its vision models, ZAYA1-VL-8B and Zamba2-VL, through 1bit serve --mmproj. Qwen3.8-27B runs in one ROCm server with Hadamard W4A4 prompt processing and DFlash2 decode: 518 tok/s on a 1,800-token prompt and 46 tok/s decode on code (docs/lean.md). On ROCm the 256-wide attention heads of Qwen3.5/3.8 run on a WMMA kernel of ours: a 32K-token prompt at 310 tok/s instead of 260 (docs/lean.md). W4A4 covers MoE experts too, and 1bit serve --long-model sends long prompts there and short ones to Vulkan: on Qwen3-Coder-30B-A3B a 16K-token prompt gets its answer at 13.1 effective tok/s against 8.6 on Vulkan alone, first token 11 s sooner (docs/serve.md). Experimental, and closed source: Qwen3.6-35B-A3B on the NPU through a private add-on, parity against fp64 passes, 16.3-16.5 tok/s decode (docs/npu.md). GGUFs of six architectures (Qwen2.5, Qwen3 MoE, Qwen3.6-35B-A3B, MiniCPM4/5, GLM-4.7-Flash) also answer on the NPU through 1bit serve --device npu: correct, not yet fast (0.006-0.17 tok/s; docs/npu.md). DwarfStar reads 1BP packages from our fork (docs/dwarfstar.md). Step 4, the Laya router, has landed as an opt-in: 1bit serve --laya classifies each conversation (code, prose, short, long document; 95.5% on 200 labelled requests) and a measured policy picks the device; the scorer runs on HRX at 15-16 ms a decision (docs/laya.md). Step 5, the model registry, has landed: of 332,726 HF text-generation models with an architecture, 94.88% are mapped to a backend and 64.33% have an architecture checked end to end on Strix Halo; a daily census keeps the counts current (docs/registry.md). The working engine is being ported from 1bit-MONSTER, our private development repository; see docs/PORTING.md. Measured results are on the wiki. This repository holds the verified code without the development history.

License

Apache-2.0. See LICENSE.

Thank you

Osmantic / ODS comes first. ODS introduced me to vibecoding, and that is where all of this started. Without it, this engine would not exist.

The Lemonade team and AMD's developers. Lemonade is where this engine runs. AMD's developers left breadcrumbs all over the place: the XDNA driver and XRT, IRON and Peano, HRX, their tested llama.cpp integration, their issues, their examples. This engine is what following those breadcrumbs built.

The repositories this engine is built on, in order of importance:

# Repository What it gives this engine License
1 amd/xdna-driver The XDNA 2 NPU driver and its XRT shim Apache-2.0 (shim)
2 Xilinx/XRT The runtime every NPU kernel runs through: full ELFs, hardware contexts, buffers Apache-2.0 (userspace)
3 Xilinx/mlir-aie IRON and aiecc: how our own NPU kernels are written and compiled Apache-2.0 WITH LLVM-exception
4 Xilinx/llvm-aie Peano, the C++ compiler for the NPU's AI Engine cores Apache-2.0 WITH LLVM-exception
5 ggml-org/llama.cpp GGUF inference on the Radeon iGPU: Vulkan (upstream release) and HRX (AMD's tested pair) MIT
6 ROCm/hrx-system HRX, AMD's HIP Runtime Extended, behind the HRX0 device Apache-2.0
7 torvalds/linux The kernel, with amdxdna and amdgpu in-tree GPL-2.0 WITH Linux-syscall-note
8 zolotukhin/zinc Its own GPU kernels, and the engine's route to NVIDIA through CUDA MIT
9 huggingface/tokenizers Every model's tokenizer.json, byte-exact, behind our C ABI Apache-2.0
10 NandhaKishorM/laya The router that decides where each request runs Apache-2.0
11 ROCm/FastFlowLM The Q4NX NPU model format and its models on Hugging Face (FastFlowLM/*-NPU2), which the engine's NPU route runs on its own kernels MIT
12 antirez/ds4 (DwarfStar) DeepSeek V4 Flash, GLM 5.x and Qwen3.8-Flash-Next on its own kernels (ROCm on Strix Halo, CUDA, Metal) MIT

Also built on nlohmann/json (MIT) and yhirose/cpp-httplib (MIT).

Every third-party copyright and license is listed in NOTICE. Each project keeps its own license; nothing here relicenses anyone's work.

Footnotes

  1. The HRX + Loom + NPU direction, and embedding the engine into Lemonade rather than Lemonade into the engine, were proposed by geramyL, a moderator on Lemonade's Discord. Thank you. ↩ ↩2

About

1bit engine: local LLM inference for AMD Ryzen AI (Strix Halo) inside Lemonade. HRX with our Loom kernels on the Radeon iGPU, plus the XDNA 2 NPU; OpenAI-compatible API. Runs Qwen3.8, Zyphra ZAYA1, ternary Bonsai.

Topics

Resources

Contributing

Security policy

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages