Home Content News Microcontrollers Now Run a Diffusion Model and 289M LLM

Microcontrollers Now Run a Diffusion Model and 289M LLM

0
1
On-device AI without servers: image generation, local LLMs, and 1-bit smart glasses models.
On-device AI without servers: image generation, local LLMs, and 1-bit smart glasses models.

Two open-source AI projects have shown that image generation and a 289-million-parameter language model can now run directly on bare microcontrollers without Linux or external accelerators.

Two open-source AI demonstrations have pushed TinyML into new territory by running advanced generative AI models directly on bare-metal microcontrollers. The projects show that a diffusion model for image generation and a 289-million-parameter language model can operate without Linux, GPUs or dedicated AI accelerators, relying instead on highly optimised inference techniques for resource-constrained hardware.

The first demonstration runs a compact diffusion model on an STM32N6570-DK development board featuring an Arm Cortex-M55 processor with an Ethos-U55 neural processing unit. The system generates 64 × 64 grayscale images using a quantised model that occupies about 3.36 MB of memory, illustrating how diffusion-based image generation can be adapted for embedded hardware with limited resources.

The second project focuses on language models rather than image generation. It runs a 289-million-parameter LLM on an NXP FRDM-MCXN947 microcontroller using aggressive quantisation, compact weight storage and optimised inference to fit the model within the available memory and processing constraints. Instead of targeting cloud-scale AI, the demonstration explores practical local inference on deeply embedded devices.

Both projects reflect a broader shift in TinyML, where increasingly capable AI models are being redesigned for low-power edge hardware. Running AI directly on microcontrollers reduces dependence on cloud connectivity, lowers latency and enables privacy-sensitive applications such as industrial sensing, robotics, wearables and IoT devices to perform inference locally.

The demonstrations do not claim that microcontrollers can replace GPU-based AI systems. Instead, they highlight how model compression, quantisation and hardware-aware optimisation are making advanced AI techniques practical on embedded hardware. By keeping both projects open-source, the developers are providing a foundation for researchers and makers to explore how increasingly capable AI models can operate on inexpensive microcontroller platforms.

Loading form…

LEAVE A REPLY

Please enter your comment!
Please enter your name here