DiffAttn Framework Uses Diffusion and LLMs to Predict Driver Visual Attention

Researchers propose DiffAttn, a diffusion-based system using LLMs to predict driver visual attention and improve vehicle safety.

According to arxiv.org, researchers have developed DiffAttn, a diffusion-based framework designed to predict drivers’ visual attention patterns to enhance traffic safety in intelligent vehicles.

The system formulates visual attention prediction as a conditional diffusion-denoising process, according to the paper published June 18, 2026. DiffAttn uses a Swin Transformer as encoder and incorporates a decoder combining a Feature Fusion Pyramid with dense, multi-scale conditional diffusion to model both local and global scene contexts.

A key innovation is the integration of a large language model (LLM) layer to “enhance top-down semantic reasoning and improve sensitivity to safety-critical cues,” according to the abstract. The researchers state that extensive experiments on four public datasets demonstrate DiffAttn achieves “state-of-the-art (SoTA) performance, surpassing most video-based, top-down-feature-driven, and LLM-enhanced baselines.”

According to arxiv.org, the framework supports “interpretable driver-centric scene understanding” and has potential applications in improving in-cabin human-machine interaction, risk perception, and driver state measurement in intelligent vehicles. The research appears in the Computer Vision and Pattern Recognition category with relevance to Artificial Intelligence.