Cluster of 7 ESP32-S3 boards runs language model with 1.58-bit ternary quantization
Open source project distributes a 0.4-0.5 billion parameter model across seven ESP32-S3 boards, using the BitNet technique to compress weights to three possible values per parameter.
A repository published on GitHub by Low-Zi-Hong and featured on Hacker News shows a cluster of seven ESP32-S3 boards running, in a distributed way, a language model quantized with the BitNet ternary technique (1.58 bit per weight). The project, named ESP32s3-LLM-Cluster, already has 168 stars and 7 forks on GitHub and is licensed under MIT.
The ESP32-S3 is a microcontroller costing a few dollars, with a dual-core Xtensa CPU and, in the variants used for this type of project, a few megabytes of PSRAM and flash. Running an entire language model on it, without aggressive quantization, is unfeasible: the weights simply don't fit in memory. The project works around this on two fronts: it compresses each weight to one of three possible values (the BitNet technique) and distributes the model's layers across several physical boards, each processing a piece of the network.
How the model is sliced across the boards
The architecture described in the README separates the cluster into one master node and six compute nodes, linked in a chain (daisy-chain) over a two-channel SPI bus. The master node doesn't perform inference on the transformer layers: it runs the BPE tokenizer, looks up the embedding table (packed in INT4, taking up about 14 MB of flash) and, at the end, the LM Head plus the greedy sampling that decides the next token.
The heavy lifting of the transformer is handled by the six compute nodes, each responsible for a block of four layers (the README lists Node 1 processing layers 0 to 3 and Node 6 closing with layers 20 to 23, totaling 24 layers). Each node executes, in sequence:
RMSNorm(computed in FP16, scaled to FP32);- 1.58-bit attention (Q, K, V, and O projections) with
RoPE; - key-value cache (
KV Cache) kept in PSRAM; - 1.58-bit MLP (Gate, Up, and Down projections).
The hidden state travels in FP32 from one node to the next over SPI channel A, while channel B receives from the previous node. After the signal passes through the last node, it returns to the master, which applies the final normalization (stored in a 64 KB partition called fnorm) and produces the output token.
What 1.58-bit quantization is
BitNet is the ternary quantization technique described by Microsoft Research, which reduces each network weight to just three possible values: -1, 0, or +1. This replaces floating-point multiplication, expensive on limited hardware, with simple additions and subtractions, in addition to drastically cutting storage space. The project's code implements this layer in bitlinear.cpp, with an assembly-optimized version (bitlinear_forward.S) to speed up ternary multiply-accumulate (MAC) operations, backed by precomputed lookup tables (lut_table.cpp).
The model itself is derived from a Qwen architecture: the file that handles attention on the compute nodes is called qwen_attention.cpp, and the README describes the preparation process as fine-tuning via Quantization-Aware Training (QAT), in the qat_158.py script. In other words, the author didn't invent an architecture from scratch: they took an existing Qwen model, in the range of 0.4 to 0.5 billion parameters according to the repository's own varying descriptions, and retrained it to operate at ternary precision before slicing it across the boards.
The preparation pipeline, on the PC side
The repository's python_tools/ folder concentrates the work that happens before anything reaches the ESP32-S3:
crop_token.pyreduces the tokenizer's vocabulary to 32,000 tokens, trimming what's not needed to fit in the boards' flash;crop_model_weight.pyslices the embedding matrix according to this reduced vocabulary;qat_158.pyperforms fine-tuning with Quantization-Aware Training to adapt the weights to ternary precision;bit4_embedding.pypacks the embeddings into INT4;pack_tokenizer_bin.pyandpack_model_bin.pyserialize the tokenizer and layers into.binfiles physically aligned to each board's partitions.
The firmware on each side (master and nodes) is built on the ESP-IDF framework, with its own partition tables (partitions.csv) defining where the tokenizer, model, and final normalization sit in each board's flash. Batch flashing scripts (flash_*.bat) handle simultaneous writing across the seven units.
What the material doesn't show
The README doesn't provide performance numbers: there's no tokens per second, latency per token, or measured power consumption. There's also no explicit publication date in the available material, nor details on which specific Qwen checkpoint was used as the base before QAT. For anyone wanting to reproduce the project, the repository's workflow.md file promises a step-by-step guide to flashing, model preparation, and physical wiring between the boards, but it falls outside the material analyzed here.
The author cites two earlier projects as references, called sources of inspiration in the README: one about deploying a quantized LLM on a single ESP32-S3 node, and another about distributed AI architecture across multiple microcontrollers. This indicates that ESP32s3-LLM-Cluster is part of a line of experimentation that had already been running in the maker community before this specific project consolidated tokenizer, embeddings, and a seven-node pipeline into a single repository.
Why this matters for builders
For those working with IoT and edge computing, the interest isn't in replacing a GPU with a microcontroller in production, but in showing the technical floor of what's already possible without the cloud, without an accelerator board, and with hardware costing a few dollars per unit. The combination of ternary quantization with a parallel pipeline across multiple cheap boards is a replicable architecture: any team that already works with ESP32 in embedded products can read the code in bitlinear.cpp and spi_bus.cpp as a reference for how to structure synchronous chained communication between microcontrollers for a task that, in isolation, none of them would solve alone.
It remains an open question, one that whoever reproduces the project would need to measure, what the real end-to-end latency is in a chain of seven boards connected via SPI, and how viable this is for applications that require real-time response, versus use cases that can tolerate a few seconds per generated token.
Translated from the Brazilian Portuguese original · Read the original
Managers See Fiscal Vacuum in the Election and High Interest Rates as a Brake on Venture Capital in Tech
At the TAG Summit, managers from Bradesco Asset, TAG Investimentos and the Zaftra fund point to the absence of fiscal proposals in the campaigns of Lula and Flávio Bolsonaro and project higher interest rates for longer, a scenario that raises the cost of funding for startups and tech companies in Brazil.