EDBT 2026 Demo / reviewers in the wild / expert
Jia Bei
dblp:04/3383
· DBLP profile ↗
6ranked-venue papers in the field
0as first author
6since 2021 · last 2024
0009-0008-3731-7294ORCID · corroborated
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 3Other / Interdisciplinary · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Semantic-guided RGB-Thermal Crowd Counting with Segment Anything ModelabstractRGB-Thermal (RGB-T) crowd counting leverages the complementary nature of visible light and thermal modalities for accurate counting. However, real-world scenarios often introduce challenges, such as misidentifying background elements like trees and lampposts as individuals, leading to inaccurate counts. Existing methods utilize segmentation as a preliminary procedure, which is constrained by segmentation accuracy. In this paper, we propose a novel method, utilizing the Segment Anything Model (SAM), to distinguish between the foreground and background of images. Specifically, we begin by utilizing SAM to obtain the semantic map of the original image. Subsequently, we extract the modality features and semantic features corresponding to the RGB and thermal modalities through multimodal feature extraction. These features are then fused using the Semantic-guide Feature Fusion module. Finally, the Multi-level Decoder is employed to generate the density map and the ultimate counting results. Our approach achieves state-of-the-art performance on the RGBT-CC dataset. Yaqun Fang, Jia Bei, Tongwei Ren |
ICMR | 3 |
| 2024 | Reproducibility Companion Paper of "MMSF: A Multimodal Sentiment-Fused Method to Recognize Video Speaking Style"abstractTo support the replication of "MMSF: A Multimodal Sentiment-Fused Method to Recognize Video Speaking Style", which was presented at ICMR'23, this companion paper provides the details of the artifacts. Speaking style recognition is aimed at recognizing the styles of conversations, which provides a fine-grained description about talking. In the original paper, we proposed a novel multimodal sentiment-fused method, MMSF, which extracts and integrates visual, audio and textual features of videos and introduced sentiment in MMSF with cross-attention mechanism to enhance the video feature to recognize speaking styles. In this paper, we explain the details of the implement code and the dataset used for experiments. Fan Yu 0003, Beibei Zhang 0005, Yaqun Fang, Jia Bei, Tongwei Ren, Jiyi Li, Luca Rossetto |
ICMR | 4 |
| 2023 | MMSF: A Multimodal Sentiment-Fused Method to Recognize Video Speaking StyleabstractAs talking takes a large proportion of human lives, it is necessary to perform deeper understanding of human conversations. Speaking style recognition is aimed at recognizing the styles of conversations, which provides a fine-grained description about talking. Current works focus on adopting only visual clues to recognize speaking styles, which cannot accurately distinguish different speaking styles when they are visually similar. To recognize speaking styles more effectively, we propose a novel multimodal sentiment-fused method, MMSF, which extracts and integrates visual, audio and textual features of videos. In addition, as sentiment is one of the motivations of human behavior, we first introduce sentiment into our multimodal method with cross-attention mechanism, which enhance the video feature to recognize speaking styles. The proposed MMSF is evaluated on long-form video understanding benchmark, and the experiment results show that it is superior to the state-of-the-arts. Beibei Zhang 0005, Yaqun Fang, Fan Yu 0003, Jia Bei, Tongwei Ren |
ICMR | 4 |
| 2023 | ADNet: An Asymmetric Dual-Stream Network for RGB-T Salient Object DetectionabstractRGB-Thermal salient object detection (RGB-T SOD) aims to locate salient objects in images that include both RGB and thermal information. Previous approaches often suggest designing a symmetric network structure to tackle the challenge of dealing with low-quality RGB or thermal images. However, we contend that RGB and thermal modalities possess different numbers of channels and disparities in information density. In this paper, we propose a novel asymmetric dual-stream network (ADNet). Specifically, we leverage an asymmetric backbone to extract four stages of RGB features and four stages of thermal features. To enable effective interaction among low-level features in the first two stages, we introduce the Channel-Spatial Interaction (CSI) module. In the last two stages, deep features are enhanced using the Self-Attention Enhancement (SAE) module. Experimental results on the VT5000, VT1000, and VT821 datasets attest to the superior performance of our proposed ADNet compared to state-of-the-art methods. Yaqun Fang, Ruichao Hou, Jia Bei, Tongwei Ren, Gangshan Wu |
MMAsia | 3 |
| 2023 | RGB-D Tracking via Hierarchical Modality Aggregation and Distribution NetworkabstractThe integration of dual-modal features has been pivotal in advancing RGB-Depth (RGB-D) tracking. However, current trackers are less efficient and focus solely on single-level features, resulting in weaker robustness in fusion and slower speeds that fail to meet the demands of real-world applications. In this paper, we introduce a novel network, denoted as HMAD (Hierarchical Modality Aggregation and Distribution), which addresses these challenges. HMAD leverages the distinct feature representation strengths of RGB and depth modalities, giving prominence to a hierarchical approach for feature distribution and fusion, thereby enhancing the robustness of RGB-D tracking. Experimental results on various RGB-D datasets demonstrate that HMAD achieves state-of-the-art performance. Moreover, real-world experiments further validate HMAD’s capacity to effectively handle a spectrum of tracking challenges in real-time scenarios. Boyue Xu, Ruichao Hou, Jia Bei, Tongwei Ren, Gangshan Wu |
MMAsia | 4 |
| 2023 | Easy Travelogue: A Travelogue Editor with Automatic Image Recommendation and InsertionabstractTravelogues are a common media form that incorporates both text and images. Typically, they are composed after the completion of a travel period. Creating a travelogue demands substantial time and effort, particularly in the curation of suitable images from the extensive collection of photos taken during the journey to complement the text. Consequently, we have developed and implemented Easy Travelogue, a travelogue editor that utilizes visual and language models. It offers real-time image suggestions while writing the text and can automatically insert fitting images into the finished content. The editor is versatile and can be readily utilized for personal travelogues, travel blogs, and various social media platforms, facilitating users in effortlessly sharing and showcasing their travel experiences. Fan Yu 0003, Huanyu Xing, Jia Bei, Tongwei Ren |
MMAsia | 3 |