VLDB 2026 Research / reviewers in the wild / expert
Jia Bei
dblp:04/3383
· DBLP profile ↗
16ranked-venue papers
0as first author
7since 2021 · last 2024
0009-0008-3731-7294ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 7 since 2021Databases, data management, data science and information retrieval · 6 · 6 since 2021Artificial intelligence and machine learning · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Semantic-guided RGB-Thermal Crowd Counting with Segment Anything ModelabstractRGB-Thermal (RGB-T) crowd counting leverages the complementary nature of visible light and thermal modalities for accurate counting. However, real-world scenarios often introduce challenges, such as misidentifying background elements like trees and lampposts as individuals, leading to inaccurate counts. Existing methods utilize segmentation as a preliminary procedure, which is constrained by segmentation accuracy. In this paper, we propose a novel method, utilizing the Segment Anything Model (SAM), to distinguish between the foreground and background of images. Specifically, we begin by utilizing SAM to obtain the semantic map of the original image. Subsequently, we extract the modality features and semantic features corresponding to the RGB and thermal modalities through multimodal feature extraction. These features are then fused using the Semantic-guide Feature Fusion module. Finally, the Multi-level Decoder is employed to generate the density map and the ultimate counting results. Our approach achieves state-of-the-art performance on the RGBT-CC dataset. Yaqun Fang, Jia Bei, Tongwei Ren |
ICMR | 3 |
| 2024 | Reproducibility Companion Paper of "MMSF: A Multimodal Sentiment-Fused Method to Recognize Video Speaking Style"abstractTo support the replication of "MMSF: A Multimodal Sentiment-Fused Method to Recognize Video Speaking Style", which was presented at ICMR'23, this companion paper provides the details of the artifacts. Speaking style recognition is aimed at recognizing the styles of conversations, which provides a fine-grained description about talking. In the original paper, we proposed a novel multimodal sentiment-fused method, MMSF, which extracts and integrates visual, audio and textual features of videos and introduced sentiment in MMSF with cross-attention mechanism to enhance the video feature to recognize speaking styles. In this paper, we explain the details of the implement code and the dataset used for experiments. Fan Yu 0003, Beibei Zhang 0005, Yaqun Fang, Jia Bei, Tongwei Ren, Jiyi Li, Luca Rossetto |
ICMR | 4 |
| 2024 | Jointly modeling association and motion cues for robust infrared UAV tracking
Boyue Xu, Ruichao Hou, Jia Bei, Tongwei Ren, Gangshan Wu |
Vis. Comput. | 3 |
| 2023 | MMSF: A Multimodal Sentiment-Fused Method to Recognize Video Speaking StyleabstractAs talking takes a large proportion of human lives, it is necessary to perform deeper understanding of human conversations. Speaking style recognition is aimed at recognizing the styles of conversations, which provides a fine-grained description about talking. Current works focus on adopting only visual clues to recognize speaking styles, which cannot accurately distinguish different speaking styles when they are visually similar. To recognize speaking styles more effectively, we propose a novel multimodal sentiment-fused method, MMSF, which extracts and integrates visual, audio and textual features of videos. In addition, as sentiment is one of the motivations of human behavior, we first introduce sentiment into our multimodal method with cross-attention mechanism, which enhance the video feature to recognize speaking styles. The proposed MMSF is evaluated on long-form video understanding benchmark, and the experiment results show that it is superior to the state-of-the-arts. Beibei Zhang 0005, Yaqun Fang, Fan Yu 0003, Jia Bei, Tongwei Ren |
ICMR | 4 |
| 2023 | ADNet: An Asymmetric Dual-Stream Network for RGB-T Salient Object DetectionabstractRGB-Thermal salient object detection (RGB-T SOD) aims to locate salient objects in images that include both RGB and thermal information. Previous approaches often suggest designing a symmetric network structure to tackle the challenge of dealing with low-quality RGB or thermal images. However, we contend that RGB and thermal modalities possess different numbers of channels and disparities in information density. In this paper, we propose a novel asymmetric dual-stream network (ADNet). Specifically, we leverage an asymmetric backbone to extract four stages of RGB features and four stages of thermal features. To enable effective interaction among low-level features in the first two stages, we introduce the Channel-Spatial Interaction (CSI) module. In the last two stages, deep features are enhanced using the Self-Attention Enhancement (SAE) module. Experimental results on the VT5000, VT1000, and VT821 datasets attest to the superior performance of our proposed ADNet compared to state-of-the-art methods. Yaqun Fang, Ruichao Hou, Jia Bei, Tongwei Ren, Gangshan Wu |
MMAsia | 3 |
| 2023 | RGB-D Tracking via Hierarchical Modality Aggregation and Distribution NetworkabstractThe integration of dual-modal features has been pivotal in advancing RGB-Depth (RGB-D) tracking. However, current trackers are less efficient and focus solely on single-level features, resulting in weaker robustness in fusion and slower speeds that fail to meet the demands of real-world applications. In this paper, we introduce a novel network, denoted as HMAD (Hierarchical Modality Aggregation and Distribution), which addresses these challenges. HMAD leverages the distinct feature representation strengths of RGB and depth modalities, giving prominence to a hierarchical approach for feature distribution and fusion, thereby enhancing the robustness of RGB-D tracking. Experimental results on various RGB-D datasets demonstrate that HMAD achieves state-of-the-art performance. Moreover, real-world experiments further validate HMAD’s capacity to effectively handle a spectrum of tracking challenges in real-time scenarios. Boyue Xu, Ruichao Hou, Jia Bei, Tongwei Ren, Gangshan Wu |
MMAsia | 4 |
| 2023 | Easy Travelogue: A Travelogue Editor with Automatic Image Recommendation and InsertionabstractTravelogues are a common media form that incorporates both text and images. Typically, they are composed after the completion of a travel period. Creating a travelogue demands substantial time and effort, particularly in the curation of suitable images from the extensive collection of photos taken during the journey to complement the text. Consequently, we have developed and implemented Easy Travelogue, a travelogue editor that utilizes visual and language models. It offers real-time image suggestions while writing the text and can automatically insert fitting images into the finished content. The editor is versatile and can be readily utilized for personal travelogues, travel blogs, and various social media platforms, facilitating users in effortlessly sharing and showcasing their travel experiences. Fan Yu 0003, Huanyu Xing, Jia Bei, Tongwei Ren |
MMAsia | 3 |
| 2019 | Saliency detection on sampled images for tag ranking
Jingfan Guo, Tongwei Ren, Lei Huang 0004, Jia Bei |
Multim. Syst. | 4 |
| 2017 | Object proposal on RGB-D images via elastic edge boxes
Tongwei Ren, Yuantian Wang, Shenghua Zhong, Jia Bei, Shengchao Chen |
Neurocomputing | 5 |
| 2016 | Depth-aware layered edge for object proposalabstractObject proposal, typically served as preprocessing of various multimedia applications, aims to detect the bounding boxes of possible objects in an image. In this paper, we propose a novel object proposal method for RGB-D images based on layered edges, which can effectively eliminate the influence of the mixture of edges from objects and background and improve the accuracy of proposals. Firstly, we detect the sparse edges and correct depth on super-pixel representation. Then, we use depth-adaptive sliding windows in sampling of depth distribution and measure the objectness of each candidate box in multiple depth layers. Finally, the candidate boxes are ranked according to the integrated scores of all the depth layers, and the final proposals are generated. The experimental results show that the proposed method can outperform the state-of-the-art methods on the largest RGB-D image dataset for object proposal. Tongwei Ren, Bing-Kun Bao, Jia Bei |
ICME | 4 |
| 2016 | Salient object detection for RGB-D image via saliency evolutionabstractSalient object detection aims to detect the most attractive objects in images, which has been widely used as a fundamental of various multimedia applications. In this paper, we propose a novel salient object detection method for RGB-D images based on evolution strategy. Firstly, we independently generate two saliency maps on color channel and depth channel of a given RGB-D image based on its super-pixels representation. Then, we fuse the two saliency maps with refinement to provide an initial saliency map with high precision. Finally, we utilize cellular automata to iteratively propagate saliency on the initial saliency map and generate the final detection result with complete salient objects. The proposed method is evaluated on two public RGB-D datasets, and the experimental results show that our method outperforms the state-of-the-art methods. Jingfan Quo, Tongwei Ren, Jia Bei |
ICME | 3 |
| 2016 | Automatic Scribble Simulation for Interactive Image Segmentation Evaluation
Bingjie Jiang, Tongwei Ren, Jia Bei |
MMM (1) | 3 |
| 2016 | Elastic Edge Boxes for Object Proposal on RGB-D Images
Tongwei Ren, Jia Bei |
MMM (1) | 3 |
| 2015 | Soft-assigned bag of features for object tracking
Tongwei Ren, Zhongyan Qiu, Yan Liu 0004, Tong Yu 0001, Jia Bei |
Multim. Syst. | 5 |
| 2013 | Multi-operator Image Retargeting Based on Automatic Quality AssessmentabstractImage retargeting aims to avoid visual distortion while retaining important image content in resizing. However, no single image retargeting method can handle all cases. In this paper, we propose a novel multi-operator image retargeting approach, which utilizes an efficient and human perception based automatic quality assessment in operator selection. First, we calculate the importance map and distortion map for quality assessment. Then, we construct the resizing space and assess the performance of each operator in iterative width and/or height reduction. Finally, we select the optimal operator sequence by dynamic programming and generate the target image. Experiments demonstrate the effectiveness of the proposed approach. Zhongyan Qiu, Tongwei Ren, Yan Liu 0004, Jia Bei, Yang Yang 0222 |
ICIG | 4 |
| 2007 | Research on XML-Based Active Interest Management in Distributed Virtual Environment
Jia Bei, Shiguang Ju, Jingui Pan |
ICCSA (1) | 3 |