Bimei Wang

dblp:322/3459 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0002-6400-943XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Primary Visual Cortex Inspired Point Cloud Analysis Framework
abstract
Despite significant advancements in point cloud analysis, reducing energy consumption and improving robustness remain understudied, largely due to the inherent limitations of Convolutional Neural Networks (CNNs). To address this, we take the cue from the primary visual cortex and propose a Dendritic-Connected Continuous-Coupled Neural Network (DC-CCNN), a novel Brain-Inspired Neural Network (BINN) architecture tailored for point cloud analysis. By leveraging the unique characteristics of point clouds, our design combines discrete and continuous encoding, replacing traditional Multilayer Perceptrons (MLPs) with more efficient and robust BINNs. Our approach substantially improves the performance of Brain-Inspired Neural Networks on point analysis tasks and maintaining performance comparable to state-of-the-art methods. Furthermore, DC-CCNN exhibits enhanced robustness against various point cloud deformations and corruptions. Our experimental results demonstrate that DC-CCNN achieves competitive performance on benchmark datasets, making it a promising alternative to traditional deep learning methods for point cloud analysis. With its high efficiency and robustness, DC-CCNN has the potential for widespread adoption in 3D computer vision, robotics, and autonomous systems.
Jisheng Dang, Delin Deng, Bimei Wang, Jingze Wu, Haijiang Li, Jingmei Jiao, Dengyue Pan, Mangang Xie, Jizhao Liu
AAAI3
2026 Cross-modal Prompt Disentangled Graph Neural Networks for incomplete conversational emotion recognition
Shi Qiao 0006, Xiaowei Zhang 0001, Qinglin Zhao, Bimei Wang, Jisheng Dang, Bin Hu 0001, Hong Peng 0003
Knowl. Based Syst.4
2026 HM-RAG: Long video reasoning and anomaly detection via hierarchical multi-agent retrieval-augmented generation
Jisheng Dang, Dewei Liu, Bimei Wang, Hong Peng 0003, Bin Hu 0001, Tat-Seng Chua
Pattern Recognit.5
2025 Quality-Guided Dynamic Memory for LLMs-based Long-Term Video Understanding
abstract
Using the impressive learning representation capacity of large language models (LLMs), LLM-based video understanding methods have made significant strides recently. However, most existing methods overlook the crucial importance discrepancy of frames, which often include massive low-quality frames, leading to limited performance and inferior inference efficiency, particularly for long-term videos. To this end, this paper proposes a new video understanding method called quality- guided dynamic memory network (QDM-Net). First, we design a memory quality evolution module (MQEM), which dynamically assigns weights to each frame according to contextual relationships between adjacent frames. Second, we devise a high- level quality memory bank updating mechanism (HQMBU), which selectively maintains high-quality frames in the memory bank, avoiding the negative influences of redundant frames and ensuring that the model focuses on the most informative visual cues. Extensive experiments on long-term video understanding benchmarks demonstrate that our QDM-Net consistently outperforms state-of-the-art methods, showcasing its potential in real-world applications. Our code and model will be publicly available.
Bimei Wang, Jingmei Jiao, Jisheng Dang, Qingrun Jiang, Jiyuan Lin, Zhixuan Chen, Teng Wang 0007
ICME1
2025 Instruction-aware Memory Network for Video Recognition
abstract
The rapid development of multimodal large language models (MLLMs) has highlighted their potential in video understanding. However, challenges remain in long video tasks, particularly in integrating visual features with prompt texts. Existing methods naively store processed video frames in a long-term memory bank, but neglect simple yet effective cross-modal integration. To address this, we introduce the instruction-aware memory construction (IaMC) model for long-term video understanding. By integrating visual and textual information, our model can obtain cross-modal features with robust understanding capabilities. These features are stored in a text-visual memory bank, enabling efficient long-term aggregation without surpassing LLM context or GPU memory limits. Experiments on the LVU dataset demonstrate state-of-the-art performance in video understanding and question answering, showcasing the IaMC model’s effectiveness and setting a new benchmark for long-term video analysis. The source code and trained models will be released publicly.
Bimei Wang, Haijiang Li, Jisheng Dang, Yun Wang 0053, Zhixuan Chen, Jiyuan Lin, Teng Wang 0007
ICME1
2025 AS-Memory: Adaptive Sparse Memory Meeting Video-Language Models
abstract
Long-term video understanding in intelligent transportation systems (ITS) has advanced significantly with the integration of large language models (LLMs) and vision foundation models. However, existing LLM-based multimodal approaches are limited by context length and memory constraints, restricting their effectiveness to short video scenarios. To address these challenges, we propose AS-Memory, a novel framework that combines adaptive sparse memory with LLMs for efficient and scalable long-term video understanding. AS-Memory introduces a plug-and-play memory bank, a lightweight module designed to seamlessly integrate with existing multimodal LLMs. This memory bank stores and retrieves historical video content, enabling long-term analysis while mitigating context length and GPU memory limitations. To further enhance efficiency, we propose a sparse adaptive mechanism that dynamically compresses redundant features and retains critical information, ensuring effective management of streaming video data. Comprehensive evaluations on long-term video understanding benchmarks demonstrate that AS-Memory consistently outperforms state-of-the-art methods in terms of accuracy. The source code and trained models will be made available to the public.
Bimei Wang, Huilin Song, Jisheng Dang, Fei Shen 0004, Mangang Xie, Jizhao Liu, Jia-Si Weng 0001
ICME1
2025 Mitigating Hallucination in Large Video-Language Models with Injected Semantics
abstract
Vision-Language Models (VLMs) have demonstrated remarkable performance across various tasks by encoding visual frames into tokens analogous to textual tokens, which are then processed by a Large Language Model (LLM) for task execution. To manage computational demands, current methods often employ a token compressor, such as Q-former, for efficient inference. However, these methods are typically trained on video-to-text generation loss, lacking sufficient supervision to align intermediate visual representations with textual semantics, resulting in hallucinations when identifying essential objects. To address this issue, we propose a novel visual-textual alignment framework, Semantic Supervision LLM (SS-LLM), which aligns video and text representations within the intermediate feature space, thereby enhancing the LLM’s decoding process. Additionally, we introduce a CLIP Loss to facilitate visual-text alignment in the intermediate feature space, reducing hallucinations in VLMs. Extensive experiments demonstrate that our approach not only mitigates hallucinations more effectively than existing models but also achieves state-of-the-art performance across several benchmarks, providing more accurate and semantically consistent video-text representations. We will make our source code and trained models publicly available.
Bimei Wang, Fan Wen, Jisheng Dang, Huiguo He, Nannan Zhu, Jia-Si Weng 0001
ICME1
2025 Diff-LMM: Diffusion Teacher-Guided Spatio-Temporal Perception for Video Large Multimodal Models
abstract
Dynamic spatio-temporal understanding is essential for video-based multimodal tasks, yet existing methods often struggle to capture fine-grained temporal and spatial relationships in long videos. Current approaches primarily rely on pre-trained CLIP encoders, which excel in semantic understanding but lack spatially-aware visual context. This leads to hallucinated results when interpreting fine-grained objects or scenes. To address these limitations, we propose a novel framework that integrates diffusion models into multimodal video models. By employing diffusion encoders at intermediate layers, we enhance visual representations through feature alignment and knowledge distillation losses, significantly improving the model's ability to capture spatial patterns over time. Additionally, we introduce a multi-level alignment strategy to learn robust feature correspondence from pre-trained diffusion models. Extensive experiments on benchmark datasets demonstrate our approach's state-of-the-art performance across multiple video understanding tasks. These results establish diffusion models as a powerful tool for enhancing multimodal video models in complex, dynamic scenarios.
Jisheng Dang, Ligen Chen, Jingze Wu, Ronghao Lin, Bimei Wang, Yun Wang 0053, Nannan Zhu, Teng Wang 0007
IJCAI5
2025 Hallucination Reduction in Video-Language Models via Hierarchical Multimodal Consistency
abstract
The rapid advancement of large language models (LLMs) has led to the widespread adoption of video-language models (VLMs) across various domains. However, VLMs are often hindered by their limited semantic discrimination capability, exacerbated by the limited diversity and biased sample distribution of most video-language datasets. This limitation results in a biased understanding of the semantics between visual concepts, leading to hallucinations. To address this challenge, we propose a Multi-level Multimodal Alignment (MMA) framework that leverages a text encoder and semantic discriminative loss to achieve multi-level alignment. This enables the model to capture both low-level and high-level semantic relationships, thereby reducing hallucinations. By incorporating language-level alignment into the training process, our approach ensures stronger semantic consistency between video and textual modalities. Furthermore, we introduce a two-stage progressive training strategy that exploits larger and more diverse datasets to enhance semantic alignment and better capture general semantic relationships between visual and textual modalities. Our comprehensive experiments demonstrate that the proposed MMA method significantly mitigates hallucinations and achieves state-of-the-art performance across multiple video-language tasks, establishing a new benchmark in the field.
Jisheng Dang, Shengjun Deng, Haochen Chang, Teng Wang 0007, Bimei Wang, Shude Wang, Nannan Zhu, Guo Niu, Jizhao Liu
IJCAI5
2025 External Memory Matters: Generalizable Object-Action Memory for Retrieval-Augmented Long-Term Video Understanding
abstract
Long video understanding with Large Language Models (LLMs) enables the description of objects that are not explicitly present in the training data. However, continuous changes in known objects and the emergence of new ones require up-to-date knowledge of objects and their dynamics for effective understanding of the open world. To alleviate this, we propose an efficient Retrieval-Enhanced Video Understanding method, dubbed REVU, which leverages external knowledge to enhance the performance of open-world learning. First, REVU introduces an extensible external text-object memory with minimal text-visual mapping, involving static and dynamic multimodal information to help LLMs-based models align text and vision features. Second, REVU retrieves object information from external databases and dynamically integrates frame-specific data from videos, enabling effective knowledge aggregation to comprehend the open world. We conducted experiments on multiple benchmark datasets, and our model demonstrates strong adaptability to out-of-domain data without requiring additional fine-tuning or re-training. Experiments on benchmark video understanding datasets reveal that our model achieves state-of-the-art performance and robust generalization.
Jisheng Dang, Huicheng Zheng, Jingmei Jiao, Bimei Wang, Bin Hu 0001, Jian-Huang Lai, Tat-Seng Chua
IJCAI5
2024 Temporo-Spatial Parallel Sparse Memory Networks for Efficient Video Object Segmentation
abstract
Memory-based networks have achieved tremendous success in video object segmentation. However, these methods still suffer from unfaithful segmentation and inferior efficiency under complicated video scenarios. The reasons are mainly threefold: 1) Weak perception of fast-moving targets due to individual frame memory patterns without capturing inter-frame motion; 2) Lack of discrimination to visually similar appearances due to the limited receptive field; 3) Redundant computation caused by matching with all memorized frames. To address these issues, we propose a Temporo-Spatial Parallel Sparse Memory network (TSPSM) for efficient video object segmentation. Our TSPSM constructs a temporal memory bank and a spatial memory bank in parallel to memorize complementary discriminative object cues. The temporal bank exploits discriminative temporal motion cues, while the spatial bank mines spatial context cues between adjacent frames with large receptive fields, thereby alleviating the ambiguity caused by similar instances and fast movements. To reduce redundant computation without sacrificing performance during the matching step, we further design a parallel sparse memory reader based on the constructed informative memory banks, which efficiently retrieves relevant temporal and spatial information in a parallel way. Experiments demonstrate that our TSPSM achieves state-of-the-art performance with real-time speed on DAVIS, and YouTube-VOS benchmarks. Furthermore, extensive experiments show that the proposed TSPMC module can be applied to existing methods as a generic plugin to significantly improve performance.
Jisheng Dang, Huicheng Zheng, Bimei Wang, Longguang Wang, Yulan Guo
IEEE Trans. Intell. Transp. Syst.3