Home Content News Tether Open Sources VisionPsy-Nano To Drive On-Device Multimodal AI

Tether Open Sources VisionPsy-Nano To Drive On-Device Multimodal AI

0
1
Tether
Tether

Tether’s AI research arm, QVAC, has released a 460M-parameter vision-language model that tops sub-0.5B benchmarks and delivers up to 36x faster smartphone inference.

On 29 July 2026, QVAC, the AI research division of Tether Data, open-sourced VisionPsy-Nano, a 460-million-parameter vision-language model (VLM) engineered specifically for local on-device and edge execution. Architecturally built on the nanoVLM framework pairing a SigLIP2 vision encoder with a 360M-parameter SmolLM2 language backbone, the model secured the top normalized score of 62.3 across 17 industry benchmarks for sub-0.5B models.

Outperforming its own base model (nanoVLM-460M-8k, which scored 54.9) on 16 of 17 tests, VisionPsy-Nano surpassed rival compact VLMs such as Liquid AI’s LFM2.5-VL-450M (59.6) and Hugging Face’s SmolVLM2-500M (52.5), establishing a relative lead in visual perception (+4.6%) and visual reasoning (+7.4%).

Tether Data released two distinct open-weight variants supporting an 8,192-token context window: VisionPsy-Nano-460M, optimized for reference accuracy, and VisionPsy-Nano-460M-Flash, tuned for minimal latency on consumer smartphones. While the standard variant generates 1,088 visual tokens per image, the Flash variant preserves native image scale to process just 64 visual tokens. This enables the Flash model to retain roughly 99% of full quality (scoring 61.4) while delivering first-token generation up to 36 times faster on an iPhone 15 and 19–23 times faster on Android devices compared to SmolVLM2-500M.

Despite its compact scale, VisionPsy-Nano demonstrated strong performance on instruction-following and hallucination benchmarks such as MM-IFEval (42.3) and POPE (87.9), outpacing models up to 2.3 times its size, including FastVLM-0.5B, Qwen3.5-0.8B, and InternVL3.5-1B. Distributed under the Apache 2.0 license, both variants feature day-one integration via Hugging Face Transformers, llama.cpp GGUF quantized checkpoints, and high-throughput vLLM server inference.

Primary language capabilities remain restricted to English, with official guidance advising against safety-critical deployments involving complex mathematics or dense counting. Strategically, Tether CEO Paolo Ardoino framed the release as part of a wider effort to bypass centralized data center gatekeepers in favour of private, local-first inference.

LEAVE A REPLY

Please enter your comment!
Please enter your name here