Three New Studies Examine Vision-Language-Action Models for Robotics and Benchmark Limitations

Researchers release new VLA model, expose visual grounding issues, and benchmark policies on low-cost robots.

Three separate research papers published on June 15, 2026, address different aspects of Vision-Language-Action (VLA) models for robotic applications.

According to arxiv.org, researchers including He Zhang and colleagues introduced “Hy-Embodied-0.5-VLA,” described as moving “From Vision-Language-Action Models to a Real-World Robot Learning Stack,” though specific technical details were not provided in the available excerpt.

A second study titled “Mirage Probes: How Vision Models Fake Visual Understanding” examined a critical limitation in vision-language models. According to arxiv.org, the research found that VLMs “can answer image-based questions confidently, and often correctly, even when no image is provided.” The authors argue this “mirage behavior inflates benchmark scores without reflecting visual grounding” and identified two distinct failure modes: “textual biases, where the model answers from language priors without engaging visual representations, and spurious images, where it constructs false visual content in latent space.”

A third paper benchmarked VLA models on the SO-101 robotic platform. According to arxiv.org, researchers evaluated models including π₀.₅, SmolVLA, Wall-X, and ACT on “four representative manipulation tasks” using a “low-cost SO-101 robotic platform.” The study found that “stronger pretrained VLA policies generally outperform the imitation learning baseline, although performance remains highly task-dependent,” with “execution instability” identified as “the dominant failure source.”