VLDB 2026 Research / reviewers in the wild / expert
Zhengxiao Sun
dblp:304/1529
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Deep learning architectures and training · 77% Vision and language · 23% |
Topics — the 2 heaviest of 2, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training › feature fusion
dynamic fusion |
0.5 | 1 | 2021 | Learning What and When to Drop: Adaptive Multimodal and Contextual Dynamics for Emotion Recognition in Conversation · ACM Multimedia 2021 |
Computer vision › Vision and language
multimodal fusion |
0.1 | 1 | 2021 | Learning What and When to Drop: Adaptive Multimodal and Contextual Dynamics for Emotion Recognition in Conversation · ACM Multimedia 2021 |
Methods — techniques the papers use, named apart from their topics
differentiable module-wise dropping · 0.5adaptive multimodal fusion · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Synesthesia Transformer with Contrastive Multimodal Learning
Zhengxiao Sun, Feiyu Chen 0001, Jie Shao 0001 |
ICONIP (1) | 1 |
| 2021 | Learning What and When to Drop: Adaptive Multimodal and Contextual Dynamics for Emotion Recognition in ConversationabstractMulti-sensory data has exhibited a clear advantage in expressing richer and more complex feelings, on the Emotion Recognition in Conversation (ERC) task. Yet, current methods for multimodal dynamics that aggregate modalities or employ additional modality-specific and modality-shared networks are still inadequate in balancing between the sufficiency of multimodal processing and the scalability to incremental multi-sensory data type additions. This incurs a bottleneck of performance improvement of ERC. To this end, we present MetaDrop, a differentiable and end-to-end approach for the ERC task that learns module-wise decisions across modalities and conversation flows simultaneously, which supports adaptive information sharing pattern and dynamic fusion paths. Our framework mitigates the problem of modelling complex multimodal relations while ensuring it enjoys good scalability to the number of modalities. Experiments on two popular multimodal ERC datasets show that MetaDrop achieves new state-of-the-art results. Feiyu Chen 0001, Zhengxiao Sun, Deqiang Ouyang, Xueliang Liu, Jie Shao 0001 |
ACM Multimedia | 2 |