Researchers have introduced DiffAttn, a diffusion-based framework designed to predict drivers’ visual attention patterns to improve traffic safety in intelligent vehicles, according to a paper published on arxiv.org.
According to the research, DiffAttn formulates visual attention prediction as a “conditional diffusion-denoising process” and incorporates a large language model (LLM) layer “to enhance top-down semantic reasoning and improve sensitivity to safety-critical cues.” The system uses Swin Transformer as an encoder and employs a Feature Fusion Pyramid decoder for cross-layer interaction with “dense, multi-scale conditional diffusion.”
The researchers tested DiffAttn on four public datasets, where it “achieves state-of-the-art (SoTA) performance, surpassing most video-based, top-down-feature-driven, and LLM-enhanced baselines,” according to the paper. The framework is designed to model both local and global scene features to capture drivers’ perception patterns.
According to arxiv.org, the system “supports interpretable driver-centric scene understanding and has the potential to improve in-cabin human-machine interaction, risk perception, and drivers’ state measurement in intelligent vehicles.” The research emphasizes that drivers’ visual attention “provides critical cues for anticipating latent hazards and directly shapes decision-making and control maneuvers, where its absence can compromise traffic safety.”