VLDB 2026 Research / reviewers in the wild / expert
Bin Jiang 0011
dblp:18/4625-11
· DBLP profile ↗
16ranked-venue papers
0as first author
16since 2021 · last 2025
0000-0002-2897-5745ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Enhancing Long Video Understanding via Hierarchical Event-Based MemoryabstractRecently, integrating visual foundation models into large language models (LLMs) to form video understanding systems has attracted widespread attention. Most of the existing models compress diverse semantic information within the whole video and feed it into LLMs for content comprehension. While this method excels in short video understanding, it may result in a blend of multiple event information in long videos due to coarse compression, which causes information redundancy. Consequently, the semantics of key events might be obscured within the vast information that hinders the model’s understanding capabilities. To address this issue, we propose a Hierarchical Event-based Memory-enhanced LLM (HEM-LLM) for better understanding of long videos. Firstly, we design a novel adaptive sequence segmentation scheme to divide multiple events within long videos. In this way, we can perform individual memory modeling for each event to establish intra-event contextual connections, thereby reducing information redundancy. Secondly, while modeling current event, we compress and inject the information of the previous event to enhance the long-term inter-event dependencies in videos. Finally, we perform extensive experiments on various video understanding tasks and the results show that our model achieves state-of-the-art performances. Dingxin Cheng, Yongxin Guo 0001, Bin Jiang 0011, Qingbin Liu, Xi Chen 0003 |
ICME | 5 |
| 2025 | Mamba-based Layer-wise Progressive Fusion Network with Depthwise Enhancement for Low-resource Speech RecognitionabstractExisting speech recognition models struggle to capture deep implicit representations and preserve key information during propagation when handling low-resource speech data. Furthermore, low-resource dialect datasets are even rarer than low-resource language datasets. To address these issues, we propose a model primarily consisting of layer-wise progressive fusion mamba (LPFMamba) and the concatenation-depthwise enhancement module (CDEM), named LPFMamba-CDEM. At its core, the layer-wise progressive fusion module (LPFM) employs a hierarchical selective fusion mechanism to integrate local features, global features extracted by the bidirectional mamba module, and information propagated from the preceding LPFM layer. This mechanism progressively accumulates and propagates effective representations, enhancing the model’s capacity under limited data conditions. The CDEM further enhances high-level feature representations processed through multiple encoder layers, increasing adaptability to low-resource speech. We also introduce a self-built low-resource Jilu dialect dataset with approximately 34 hours of speech, aiming to promote the equitable dissemination of technology. Extensive experiments conducted on multiple low-resource speech datasets, including both publicly available datasets and the Jilu dialect dataset, demonstrate the effectiveness of our approach. Xuanda Chen, Dingxin Cheng, Bin Jiang 0011, Meixia Qu |
IJCNN | 4 |
| 2025 | Multi-scale Weight-residual Transformer for Uniand Multi-modal Representation LearningabstractRecently, Transformer has demonstrated its comparable performance in several vision tasks and multimedia domains. However, the key self-attention in Transformer computes global attention to the image primarily in the spatial dimension, which is a non-local operation. Such spatial attention lacks the ability to model the relationship among local regions of an image and the learned representations are biased with redundant channel information perturbation. To address this problem, we propose a new Multi-scale Weight-residual Transformer (MWT) for uni-and multi-modal. Specifically, we generate local and regional tokens by different convolutions and use them as query and key-value, respectively. In this way, the computed self-attention ensures both global information interaction and focuses attention on regional information, which is more relevant to the local information. It not only retains the global attention of transformers but also obtains the ability of local attention as CNNs. Moreover, we introduce a weight-residual network for channel dimension to alleviate the feature weakening in deeper layers of the network, which can improve the sensitivity of the model for key channel representations. These can build a more abstract high-level feature representation. Extensive experiments demonstrate the effectiveness of MWT on several uni- and multimodal benchmark tasks. Dingxin Cheng, Jiawei Gu, Kang Xie, Mengyue Zhang, Gang Wang 0060, Bin Jiang 0011 |
IJCNN | 6 |
| 2025 | FAformer: Exploring Frequency and Attention in Transformers for Long-Term Time Series ForecastingabstractTransformers have been successfully applied to long-term time series forecasting (LTSF) owing to their ability to model long-term dependencies of time series by the multi-head attention mechanism. However, most existing Transformer-based forecasting models capture temporal dependency patterns, while ignoring the frequency patterns. To balance both patterns, we explore a novel approach of applying both frequency filtering and attention mechanism within the Transformer for the LTSF task. We design the frequency filtering layer that represents time series in terms of their frequency components to capture frequency features with log-linear complexity, providing deeper insights into global dependencies of the data. Then, we propose FAformer, a simple yet effective architecture built upon frequency and attention in Transformers for LTSF. In FAformer, the frequency filtering layer captures global dependencies in time series, and then the deeper attention layer further models the features. Extensive experiments demonstrate the effectiveness of our proposed method, which outperforms state-of-the-art (SOTA) baselines on nine real-world datasets. The code will be publicly available. Jiawei Gu, Dingxin Cheng, Qiang Guo 0003, Bin Jiang 0011, Meixia Qu |
IJCNN | 4 |
| 2025 | FedPPD: Towards effective subgraph federated learning via pseudo prototype distillation
Jishuo Jia, Yinlin Zhu, Xunkai Li, Bin Jiang 0011, Meixia Qu |
Neural Networks | 5 |
| 2025 | Effective Global Context Integration for Lightweight 3D Medical Image SegmentationabstractAccurate and fast segmentation of 3D medical images is crucial in clinical analysis. CNNs struggle to capture long-range dependencies because of their inductive biases, whereas the Transformer can capture global features but faces a considerable computational burden. Thus, efficiently integrating global and detailed insights is key for precise segmentation. In this paper, we propose an effective and lightweight architecture named GCI-Net to address this issue. The key characteristic of GCI-Net is the global-guided feature enhancement strategy (GFES), which integrates the global context and facilitates the learning of local information; 3D convolutional attention, which captures long-range dependencies; and a progressive downsampling module, which perceives detailed information better. The GFES can capture the local range of information through global-guided feature fusion and global-local contrastive loss. All these designs collectively contribute to lower computational complexity and reliable performance improvements. The proposed model is trained and tested on four public datasets, namely MSD Brain Tumor, ACDC, BraTS2021, and MSD Lung. The experimental results show that, compared with several recent SOTA methods, our GCI-Net achieves superior computational efficiency with comparable or even better segmentation performance. The code is available athttps://github.com/qintianjian-lab/GCI-Net. Qiang Qiao, Meixia Qu, Bin Jiang 0011, Qiang Guo 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | BEI-DETR: A Multimodal Remote Sensing Object Detection via Integrating EEG and Eye MovementabstractRecently, the DEtection TRansformer (DETR) and its variants have achieved success in object detection, but they still require manual intervention in complex scenes. Object detection methods cannot entirely replace human capabilities. Leveraging human perception to assist remote sensing object detection offers a new technical approach. In the process of manual verification, the most prominent signals reflecting human identification of specific targets are electroencephalogram (EEG) and eye movement. To our best knowledge, there is currently no publicly available dataset for this task. Therefore, we collect two types of human perception information while observing remote sensing images, forming a trimodal dataset including brain signals, eye movements, and remote sensing images. Upon this, we propose a multimodal object detection method, termed as BEI-DETR (Brain-Eye-Image DETR). Specifically, we decouple the classification and localization tasks in the object detection, using EEG for classification and eye movement information for localization. Subsequently, we propose a feature matching fusion module (FMFM) to fuse multimodal features and then put them into the decoder and prediction heads to obtain the final detection results. Comparative and ablation experiments demonstrate the satisfactory performance of BEI-DETR, indicating that our model can effectively learn human perceptual information for the remote sensing object detection task. The trimodal dataset and source code are publicly available at https://github.com/Hickey-Curry/BEI-DETR and https://github.com/Hickey-Curry/BEIDataset. Junfeng Huang, Mengyue Zhang, Gang Wang 0060, Meixia Qu, Bin Jiang 0011, Qiang Guo 0003 |
BIBM | 6 |
| 2024 | MEEG and AT-DGNN: Improving EEG Emotion Recognition with Music Introducing and Graph-based LearningabstractWe present the MEEG dataset, a multi-modal collection of music-induced electroencephalogram (EEG) recordings designed to capture emotional responses to various musical stimuli across different valence and arousal levels. This public dataset facilitates an in-depth examination of brainwave patterns within musical contexts, providing a robust foundation for studying brain network topology during emotional processing. Leveraging the MEEG dataset, we introduce the Attention-based Temporal Learner with Dynamic Graph Neural Network (AT-DGNN), a novel framework for EEG-based emotion recognition. This model combines an attention mechanism with a dynamic graph neural network (DGNN) to capture intricate EEG dynamics. The AT-DGNN achieves state-of-the-art (SOTA) performance with an accuracy of 83.74% in arousal recognition and 86.01% in valence recognition, outperforming existing SOTA methods. This study advances graph-based learning methodology in brain-computer interfaces (BCI), significantly improving the accuracy of EEG-based emotion recognition. The MEEG dataset and source code are publicly available at https://github.com/xmh1011/AT-DGNN. Minghao Xiao, Zhengxi Zhu, Kang Xie, Bin Jiang 0011 |
BIBM | 4 |
| 2024 | Long Term Memory-Enhanced Via Causal Reasoning for Text-To-Video RetrievalabstractThe T2VR task aims to retrieve videos that are semantically relevant to the given query text in a large number of unlabeled videos. Most of the existing methods adopt a representation encoding strategy that can only focus on limited contextual information, and lack the ability to focus on the long memory of representation sequences. Besides, they also ignore the semantic causal impact of the predecessor on the successor in the sequence. To tackle this issue, we propose a new long term memory-enhanced via causal reasoning to better learn the feature sequence of video and text. Firstly, we design semantic causal reasoning to allow video and text to adaptively capture their respective feature sequence’s full-memory contextual causal relations and enhance the consistency of the semantic relations. Secondly, we perform key feature reweighting in the memory space of video and text respectively to make the key information focused. Finally, extensive experiments on three public datasets, i.e., MSR-VTT, VATEX, and TGIF, demonstrate the effectiveness of our proposed method. Dingxin Cheng, Shuhan Kong, Meixia Qu, Bin Jiang 0011 |
ICASSP | 5 |
| 2024 | Medical Image Segmentation via Single-Source Domain Generalization with Random Amplitude Spectrum Synthesis
Qiang Qiao, Meixia Qu, Bin Jiang 0011, Qiang Guo 0003 |
MICCAI (9) | 5 |
| 2024 | Transferable dual multi-granularity semantic excavating for partially relevant video retrievalabstractPartially Relevant Video Retrieval (PRVR) aims to retrieve partially relevant videos from many unlabeled and untrimmed videos according to the query, which is defined as the multiple instance learning problem. The challenge of PRVR is that it utilizes untrimmed videos, which are much closer to reality. The existing methods excavate video-text semantic consistency information insufficiently and lack the capacity to highlight the semantics of key representations. To tackle these issues, we propose a transferable dual multi-granularity semantic excavating network, called T-D3N, to focus on enhancing the learning of dual-modal representations. Specifically, we first introduce a novel transferable textual semantic learning strategy by designing Adaptive Multi-scale Semantic Mining (AMSM) component to excavate significant textual semantic from multiple perspectives. Second, T-D3N distinguishes the feature differences from the frame-wise perspective to better perform contrastive learning between positive and negative samples in the video feature domain, which can further distance the positive and negative samples and improve the probability of positive samples being retrieved by query. Finally, our model constructs multi-grained video temporal dependencies and conducts cross-grained core feature perception, which enables more sufficient multimodal interactions . Extensive experiments are performed on three benchmarks, i.e., ActivityNet Captions, Charades-STA, and TVR, our T-D3N achieves state-of-the-art results. Furthermore, we also confirm that our model is transferable on a broad range of multimodal tasks such as T2VR, VMR, and MMSum. Dingxin Cheng, Shuhan Kong, Bin Jiang 0011, Qiang Guo 0003 |
Image Vis. Comput. | 3 |
| 2023 | Learning to Dub Movies via Hierarchical Prosody ModelsabstractGiven a piece of text, a video clip and a reference audio, the movie dubbing (also known as visual voice clone, V2C) task aims to generate speeches that match the speaker's emotion presented in the video using the desired speaker voice as reference. V2C is more challenging than conventional text-to-speech tasks as it additionally requires the generated speech to exactly match the varying emotions and speaking speed presented in the video. Unlike previous works, we propose a novel movie dubbing architecture to tackle these problems via hierarchical prosody modeling, which bridges the visual information to corresponding speech prosody from three aspects: lip, face, and scene. Specifically, we align lip movement to the speech duration, and convey facial expression to speech energy and pitch via attention mechanism based on valence and arousal representations inspired by the psychology findings. Moreover, we design an emotion booster to capture the atmosphere from global video scenes. All these embeddings are used together to generate mel-spectrogram, which is then converted into speech waves by an existing vocoder. Extensive experimental results on the V2C and Chem benchmark datasets demonstrate the favourable performance of the proposed method. The code and trained models will be made available at https://github.com/GalaxyCong/HPMDubbing Gaoxiang Cong 0001, Liang Li 0003, Yuankai Qi, Zhengjun Zha, Qi Wu 0001, Bin Jiang 0011, Ming-Hsuan Yang 0001, Qingming Huang |
CVPR | 7 |
| 2023 | Dynamic Contrastive Learning with Pseudo-samples Intervention for Weakly Supervised Joint Video MR and HDabstractJoint video moment retrieval (MR) and highlight detection (HD) aims to find relevant video moments according to the query text. Existing methods are fully supervised based on manual annotation, and their coarse multi-modal information interactions easily lose details about video and text. In addition, some tasks introduce weakly supervised learning with random masks, while the single masking forces the model to focus on masked words and ignore multi-modal contextual information. In view of this, we attempt weakly supervised joint tasks (MR+HD) and propose Dynamic Contrastive Learning with Pseudo-Sample Intervention (CPI) for better multi-modal video comprehension. First, we design pseudo-samples over random masks for a more efficient contrastive learning manner. We introduce a proportional sampling strategy for pseudo-samples to ensure the semantic difference between the pseudo-samples and the query text. This balances the over-reliance from single random mask to global text semantics and makes the model learn multimodal context from each word fairly. Second, we design dynamic intervention contrastive loss to enhance the core feature-matching ability of the model dynamically. We add pseudo-sample intervention when negative proposals are close to positive proposals. This can help the model overcome the vision confusion phenomenon and achieve semantic similarity instead of word similarity. Extensive experiments demonstrate the effectiveness of CPI and the potential of weakly supervised joint tasks. Shuhan Kong, Liang Li 0003, Beichen Zhang 0006, Bin Jiang 0011, Chenggang Yan 0001, Changhao Xu |
ACM Multimedia | 5 |
| 2023 | Effective hybrid graph and hypergraph convolution network for collaborative filtering
Xunkai Li, Ronghui Guo, Youpeng Hu, Meixia Qu, Bin Jiang 0011 |
Neural Comput. Appl. | 6 |
| 2023 | LoyalDE: Improving the performance of Graph Neural Networks with loyal node discovery and emphasis
Haotong Wei, Yinlin Zhu, Xunkai Li, Bin Jiang 0011 |
Neural Networks | 4 |
| 2022 | LS-GAN: Iterative Language-based Image Manipulation via Long and Short Term Consistency ReasoningabstractIterative language-based image manipulation aims to edit images step by step according to user's linguistic instructions. The existing methods mostly focus on aligning the attributes and appearance of new-added visual elements with current instruction. However, they fail to maintain consistency between instructions and images as iterative rounds increase. To address this issue, we propose a novel Long and Short term consistency reasoning Generative Adversarial Network (LS-GAN), which enhances the awareness of previous objects with current instruction and better maintains the consistency with the user's intent under the continuous iterations. Specifically, we first design a Context-aware Phrase Encoder (CPE) to learn the user's intention by extracting different phrase-level information about the instruction. Further, we introduce a Long and Short term Consistency Reasoning (LSCR) mechanism. The long-term reasoning improves the model on semantic understanding and positional reasoning, while short-term reasoning ensures the ability to construct visual scenes based on linguistic instructions. Extensive results show that LS-GAN improves the generation quality in terms of both object identity and position, and achieves the state-of-the-art performance on two public datasets. Gaoxiang Cong 0001, Liang Li 0003, Zhenhuan Liu, Yunbin Tu, Weijun Qin, Shenyuan Zhang, Chengang Yan, Bin Jiang 0011 |
ACM Multimedia | 9 |