Timezone: »
Attention networks in multimodal learning provide an efficient way to utilize given visual information selectively. However, the computational cost to learn attention distributions for every pair of multimodal input channels is prohibitively expensive. To solve this problem, co-attention builds two separate attention distributions for each modality neglecting the interaction between multimodal inputs. In this paper, we propose bilinear attention networks (BAN) that find bilinear attention distributions to utilize given vision-language information seamlessly. BAN considers bilinear interactions among two groups of input channels, while low-rank bilinear pooling extracts the joint representations for each pair of channels. Furthermore, we propose a variant of multimodal residual networks to exploit eight-attention maps of the BAN efficiently. We quantitatively and qualitatively evaluate our model on visual question answering (VQA 2.0) and Flickr30k Entities datasets, showing that BAN significantly outperforms previous methods and achieves new state-of-the-arts on both datasets.
Author Information
Jin-Hwa Kim (SK T-Brain)
Jaehyun Jun (Seoul National University)
Byoung-Tak Zhang (Seoul National University & Surromind Robotics)
More from the Same Authors
-
2021 : Partition-based Local Independence Discovery »
Inwoo Hwang · Byoung-Tak Zhang · Sanghack Lee -
2021 : C^3: Contrastive Learning for Cross-domain Correspondence in Few-shot Image Generation »
Hyukgi Lee · Gi-Cheon Kang · Chang-Hoon Jeong · Hanwool Sul · Byoung-Tak Zhang -
2022 Poster: Robust Imitation via Mirror Descent Inverse Reinforcement Learning »
Dong-Sig Han · Hyunseo Kim · Hyundo Lee · JeHwan Ryu · Byoung-Tak Zhang -
2022 Poster: SelecMix: Debiased Learning by Contradicting-pair Sampling »
Inwoo Hwang · Sangjun Lee · Yunhyeok Kwak · Seong Joon Oh · Damien Teney · Jin-Hwa Kim · Byoung-Tak Zhang -
2021 Poster: Goal-Aware Cross-Entropy for Multi-Target Reinforcement Learning »
Kibeom Kim · Min Whoo Lee · Yoonsung Kim · JeHwan Ryu · Minsu Lee · Byoung-Tak Zhang -
2020 Workshop: BabyMind: How Babies Learn and How Machines Can Imitate »
Byoung-Tak Zhang · Gary Marcus · Angelo Cangelosi · Pia Knoeferle · Klaus Obermayer · David Vernon · Chen Yu -
2020 : Opening Remarks: BabyMind, Byoung-Tak Zhang and Gary Marcus »
Byoung-Tak Zhang · Gary Marcus -
2018 Poster: Answerer in Questioner's Mind: Information Theoretic Approach to Goal-Oriented Visual Dialog »
Sang-Woo Lee · Yu-Jung Heo · Byoung-Tak Zhang -
2018 Spotlight: Answerer in Questioner's Mind: Information Theoretic Approach to Goal-Oriented Visual Dialog »
Sang-Woo Lee · Yu-Jung Heo · Byoung-Tak Zhang -
2017 Poster: Overcoming Catastrophic Forgetting by Incremental Moment Matching »
Sang-Woo Lee · Jin-Hwa Kim · Jaehyun Jun · Jung-Woo Ha · Byoung-Tak Zhang -
2017 Spotlight: Overcoming Catastrophic Forgetting by Incremental Moment Matching »
Sang-Woo Lee · Jin-Hwa Kim · Jaehyun Jun · Jung-Woo Ha · Byoung-Tak Zhang -
2016 : PororoQA: Cartoon Video Series Dataset for Story Understanding »
KyungMin Kim · Min-Oh Heo · Byoung-Tak Zhang -
2016 Poster: Multimodal Residual Learning for Visual QA »
Jin-Hwa Kim · Sang-Woo Lee · Donghyun Kwak · Min-Oh Heo · Jeonghee Kim · Jung-Woo Ha · Byoung-Tak Zhang -
2010 Poster: Generative Local Metric Learning for Nearest Neighbor Classification »
Yung-Kyun Noh · Byoung-Tak Zhang · Daniel Lee