Researchers have proposed an emotion recognition network that fuses facial expressions and body language to improve accuracy in complex human-robot interaction scenarios. The AADPS (asymmetric attention and dual pooling screening) network comprises two subnets: an intra-frame spatial fusion subnet and an inter-frame spatio-temporal fusion subnet.
Key takeaways
- Asymmetric attention mechanism hierarchically captures body-language geometry guided by gesture cues and intra-frame emotional correlation semantics guided by facial cues, addressing scale differences between facial and body regions.
- Dual pooling key-frame screening strategy uses average and max pooling to generate video-level emotional semantics, measuring similarity with frame-level semantics to suppress non-peak frames.
- The network achieves 94.95% accuracy on FABO and 88.69% on CAER datasets, improving 12.22% and 13.49% over baseline, outperforming single-modal and multi-modal methods in complex scenes.
This work offers a robust approach for visual emotion recognition in automation applications such as human-robot interaction, smart elderly care, and mental health monitoring, where facial occlusion or emotional masking can degrade performance.
Source: 《机器人》期刊 (robot.sia.cn) · Published 2026-05-12 · “基于非对称注意与双池化筛选的情感识别网络”
