Opinion

Emotion Recognition Network with Asymmetric Attention and Dual Pooling Screening

Daily briefingRyan OkaforMay 12, 2026· 4,068 views

AADPS network fuses facial and body cues with asymmetric attention and dual pooling key-frame screening, achieving 94.95% and 88.69% accuracy on FABO and CAER datasets.

Researchers have proposed an emotion recognition network that fuses facial expressions and body language to improve accuracy in complex human-robot interaction scenarios. The AADPS (asymmetric attention and dual pooling screening) network comprises two subnets: an intra-frame spatial fusion subnet and an inter-frame spatio-temporal fusion subnet.

Key takeaways

  • Asymmetric attention mechanism hierarchically captures body-language geometry guided by gesture cues and intra-frame emotional correlation semantics guided by facial cues, addressing scale differences between facial and body regions.
  • Dual pooling key-frame screening strategy uses average and max pooling to generate video-level emotional semantics, measuring similarity with frame-level semantics to suppress non-peak frames.
  • The network achieves 94.95% accuracy on FABO and 88.69% on CAER datasets, improving 12.22% and 13.49% over baseline, outperforming single-modal and multi-modal methods in complex scenes.

This work offers a robust approach for visual emotion recognition in automation applications such as human-robot interaction, smart elderly care, and mental health monitoring, where facial occlusion or emotional masking can degrade performance.

Source: 《机器人》期刊 (robot.sia.cn) · Published 2026-05-12 · “基于非对称注意与双池化筛选的情感识别网络”