flash-moe
An inference engine written in C and Metal that runs a very large 397 billion parameter language model on a laptop. It shows that big models can run on small consumer hardware.
Share on XNo license declared
Overview
Flash-MoE is an inference engine written in C, Objective-C and hand-tuned Metal shaders. It runs Qwen3.5-397B-A17B, a 397 billion parameter Mixture-of-Experts model, on a MacBook Pro with 48GB of memory. The 209GB model is streamed from the SSD through a custom Metal compute pipeline, with no Python and no frameworks.
Key features
- Runs a 397B parameter model on a 48GB MacBook Pro
- Streams the model from SSD through Metal compute
- Reported 4.36 tokens per second with 4-bit experts
- Supports tool calling in the 4-bit setup
Best for
Readers interested in how very large models can run on consumer Apple hardware. The 2-bit setting is faster but breaks JSON output and tool calling, so 4-bit is the production choice.
- Upstream
- danveloper/flash-moe
- Fork on GitHub
- Guo-astro/flash-moe
- Upstream stars
- 4.2k
- Category
- AI agents and LLM tools
- License
- No license declaredWithout a license, the author keeps all rights. Ask the upstream owner before reusing the code.
- Forked
- 2026-03-22
- Sync status
- In syncLast synced 2026-10-10
More in AI agents and LLM tools
An open-source curriculum for learning AI engineering by implementing model internals, retrieval pipelines, and agent runtimes. Study through lessons, interactive labs, or staged coding projects, and inspect the code and evaluation results.
Forked 2026-10-09Last synced 2026-10-10License: MITAI agents and LLM toolsGitHub
A small, hackable character-level language model that learns patterns from line-based text files and generates similar items, such as name ideas.
Forked 2026-10-09Last synced 2026-10-10License: MITAI agents and LLM toolsGitHub
uzu is an inference engine for running AI models in apps, helping developers keep inference on-device instead of relying on remote services. Use it to download a model and build chat features.
Forked 2026-10-09Last synced 2026-10-10License: MITAI agents and LLM toolsGitHub
Anti Slop gives AI coding agents rules to avoid generic UI, filler text, and AI-shaped code. Use it as a filter alongside your own design direction.
Forked 2026-10-08Last synced 2026-10-10License: MITAI agents and LLM toolsGitHub