Home Content News Edge AI Chip Targets Local Large Models

Edge AI Chip Targets Local Large Models

0
4
Acrab’s Agent Box
Acrab’s Agent Box

An edge AI processor combines heterogeneous computing, high-bandwidth memory, and software optimisation to run large language models locally without relying on cloud infrastructure.

GΞLIX 1 is the name of the 5 nm edge AI SoC that Acrab has launched that can process large language models with 100-billion-parameters locally. The processor executes AI workloads locally instead of relying on cloud computing, reducing latency, improving data privacy, and enabling real-time AI applications.

It has a 20-core Arm CPU, a 3 TFLOPS GPU, a neural processing unit (NPU), and multimedia processing engines in its heterogeneous architecture. The chip delivers up to 650 TOPS of AI performance. Acrab complements this with 8MB of L1 cache, 768 GB/s shared L2 cache bandwidth, and 256-bit LPDDR5X memory to provide high-bandwidth data transfer.

Hardware acceleration of transformer attention along with optimisations of the model prefill are implemented in GΞLIX 1 in order to increase inference efficiency. According to Acrab, these changes substantially reduce the first token latency and increase throughput. The manufacturer states the prefill performance of 1,416.8 tokens per second for a Gemma 26B A4B setup with a 40k KV cache and 10k-token input, reportedly outperforming the performance of similar desktop hardware in this particular application.

In addition to the processor, Acrab offers users a complete software stack, development tools, and reference designs to simplify the deployment of edge AI applications. The manufacturer has developed a proprietary Agent Box based on the GΞLIX 1 platform allowing developers to evaluate edge AI inference workloads on the edge without using the cloud resources.

Designed for edge AI deployment, GΞLIX 1 demonstrates the approach combining a specialised hardware and memory architecture as well as software optimisation to support increasingly larger AI models run locally. As demand grows for lower latency and greater control over sensitive data, processors such as GΞLIX 1 could expand the range of AI applications that operate independently of cloud infrastructure.

 

LEAVE A REPLY

Please enter your comment!
Please enter your name here