Figure AI has unveiled HELIX, a pioneering Vision-Language-Action (VLA) model that integrates vision, language comprehension, and action execution into a single neural network. This innovation allows ...
Tether Data's AI research initiative, QVAC, today announced the open-source release of VisionPsy-Nano, a compact 460-million-parameter vision-language model (VLM) purpose-built for on-device and edge ...
The rise in Deep Research features and other AI-powered analysis has given rise to more models and services looking to simplify that process and read more of the documents businesses actually use.
TurboVLA achieves 97.7% on the LIBERO robot manipulation benchmark at 32 Hz on a consumer NVIDIA RTX 4090 GPU, using 0.9 GB ...
Imagine a world where your devices not only see but truly understand what they’re looking at—whether it’s reading a document, tracking where someone’s gaze lands, or answering questions about a video.
Robot foundation model startup Dyna Robotics unveiled DYNA-2, a new architecture trained on one million hours of human video ...
Microsoft announced a new version of its small language model, Phi-3, which can look at images and tell you what’s in them. Phi-3-vision is a multimodal model — aka it can read both text and images — ...
Canadian AI startup Cohere launched in 2019 specifically targeting the enterprise, but independent research has shown it has so far struggled to gain much of a market share among third-party ...
The AI model type capturing the most attention across robotics and autonomous vehicles right now is the vision-language-action model, or VLA. At embedded AI conferences this year, particularly the ...