Jisheng Dang

dblp:272/7915 · DBLP profile ↗
← Back
30ranked-venue papers
16as first author
28since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 11 first-author · 18 since 2021Artificial intelligence and machine learning · 13 · 7 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Primary Visual Cortex Inspired Point Cloud Analysis Framework
abstract
Despite significant advancements in point cloud analysis, reducing energy consumption and improving robustness remain understudied, largely due to the inherent limitations of Convolutional Neural Networks (CNNs). To address this, we take the cue from the primary visual cortex and propose a Dendritic-Connected Continuous-Coupled Neural Network (DC-CCNN), a novel Brain-Inspired Neural Network (BINN) architecture tailored for point cloud analysis. By leveraging the unique characteristics of point clouds, our design combines discrete and continuous encoding, replacing traditional Multilayer Perceptrons (MLPs) with more efficient and robust BINNs. Our approach substantially improves the performance of Brain-Inspired Neural Networks on point analysis tasks and maintaining performance comparable to state-of-the-art methods. Furthermore, DC-CCNN exhibits enhanced robustness against various point cloud deformations and corruptions. Our experimental results demonstrate that DC-CCNN achieves competitive performance on benchmark datasets, making it a promising alternative to traditional deep learning methods for point cloud analysis. With its high efficiency and robustness, DC-CCNN has the potential for widespread adoption in 3D computer vision, robotics, and autonomous systems.
Jisheng Dang, Delin Deng, Bimei Wang, Jingze Wu, Haijiang Li, Jingmei Jiao, Dengyue Pan, Mangang Xie, Jizhao Liu
AAAI1
2026 Cross-modal Prompt Disentangled Graph Neural Networks for incomplete conversational emotion recognition
Shi Qiao 0006, Xiaowei Zhang 0001, Qinglin Zhao, Bimei Wang, Jisheng Dang, Bin Hu 0001, Hong Peng 0003
Knowl. Based Syst.5
2026 Graph-enhanced dual low-rank correlation embedding for spatio-temporal EEG fusion in depression recognition
Lu Zhang 0071, Jisheng Dang, Wencheng Gan, Bin Hu 0001, Hong Peng 0003
Neural Networks2
2026 HM-RAG: Long video reasoning and anomaly detection via hierarchical multi-agent retrieval-augmented generation
Jisheng Dang, Dewei Liu, Bimei Wang, Hong Peng 0003, Bin Hu 0001, Tat-Seng Chua
Pattern Recognit.1
2026 Fast Track Anything With Sparse Spatio-Temporal Propagation for Unified Video Segmentation
abstract
Recent advances in "track-anything" models have significantly improved fine-grained video understanding by simultaneously handling multiple video segmentation and tracking tasks. However, existing models often struggle with robust and efficient temporal propagation. To address these challenges, we propose the Sparse Spatio-Temporal Propagation (SSTP) method, which achieves robust and efficient unified video segmentation by selectively leveraging key spatio-temporal features in videos. Specifically, we design a dynamic 3D spatio-temporal convolution to aggregate global multi-frame spatio-temporal information into memory frames during memory construction. Additionally, we introduce a spatio-temporal aggregation reading strategy to efficiently aggregate the relevant spatio-temporal features from multiple memory frames during memory retrieval. By combining SSTP with an image segmentation foundation model, such as the segment anything model, our method effectively addresses multiple data-scarce video segmentation tasks. Our experimental results demonstrate state-of-the-art performance on five video segmentation tasks across eleven datasets, outperforming both task-specific and unified methods. Notably, SSTP exhibits strong robustness in handling sparse, low-frame-rate videos, making it well-suited for real-world applications.
Jisheng Dang, Huicheng Zheng, Zhixuan Chen, Yulan Guo, Tat-Seng Chua
IEEE Trans. Image Process.1
2026 Video Decoupling Networks for Accurate, Efficient, Generalizable, and Robust Video Object Segmentation
abstract
Video object segmentation (VOS) is a fundamental task in video analysis, aiming to accurately recognize and segment objects of interest within video sequences. Conventional methods, relying on memory networks to store single-frame appearance features, face challenges in computational efficiency and capturing dynamic visual information effectively. To address these limitations, we present a Video Decoupling Network (VDN) with a per-clip memory updating mechanism. Our approach is inspired by the dual-stream hypothesis of the human visual cortex and decomposes multiple previous video frames into fundamental elements: scene, motion, and instance. We propose the Unified Prior-based Spatio-temporal Decoupler (UPSD) algorithm, which parses multiple frames into basic elements in a unified manner. UPSD continuously stores elements over time, enabling adaptive integration of different cues based on task requirements. This decomposition mechanism facilitates comprehensive spatial-temporal information capture and rapid updating, leading to notable enhancements in overall VOS performance. Extensive experiments conducted on multiple VOS benchmarks validate the state-of-the-art accuracy, efficiency, generalizability, and robustness of our approach. Remarkably, VDN demonstrates a significant performance improvement and a substantial speed-up compared to previous state-of-the-art methods on multiple VOS benchmarks. It also exhibits excellent generalizability under domain shift and robustness against various noise types.
Jisheng Dang, Huicheng Zheng, Yulan Guo, Jian-Huang Lai, Bin Hu 0001, Tat-Seng Chua
IEEE Trans. Image Process.1
2026 SynPO: Synergizing Descriptiveness and Preference Optimization for Video Detailed Captioning
abstract
Fine-grained video captioning aims to generate detailed, temporally coherent descriptions of video content. However, existing methods struggle to capture subtle video dynamics and rich detailed information. In this paper, we leverage preference learning to enhance the performance of vision-language models (VLM) in fine-grained video captioning, while mitigating several limitations inherent to Direct Preference Optimization (DPO). First, we propose a pipeline for constructing preference pairs that leverages the intrinsic properties of VLMs along with partial assistance from large language models, achieving an balance between cost and data quality. Then, we propose Synergistic Preference Optimization (SynPO), a novel optimization method offering significant advantages over DPO and its variants. SynPO prevents negative pReferences from dominating the training, explicitly preserves the model's language capability to avoid deviation of the optimization objective, thus obtains high-quality captions and improves training efficiency by eliminating the need for the reference model. We extensively evaluate our proposed data construction pipeline across three models: AuroraCap, LLaVA1.6-7B-Video and InterVL2-8B. Results demonstrate that our method improve performance in fine-grained video captioning significantly and consistenly. Source code is available at https://github.com/longmalongma/SynPO.
Jisheng Dang, Teng Wang 0007, Yulan Guo, Bin Hu 0001
IEEE Trans. Image Process.1
2026 GeoStyler: A Generalizable Geometry-Aware Diffusion-Based Approach for Direct 3D Gaussian Style Transfer
abstract
Direct 3D scene stylization from sparse views remains a significant challenge, as existing optimization-based methods are prohibitively slow and require dense inputs to prevent geometric corruption. While recent direct methods accelerate this process, their rigid decoupling of a static geometry from appearance often leads to visual artifacts, where stylistic textures conflict with and distort the underlying scene structure. To address these limitations, we introduce GeoStyler, a direct framework that generates high-fidelity, multi-view consistent stylized 3D scenes in seconds. Our approach reformulates the conventional pipeline by first leveraging a diffusion model to generate a set of geometrically consistent stylized 2D images. The core of this stage is a novel hybrid query formulation for the self-attention mechanism. Specifically, cross-view geometric information is directly embedded into the query to enforce 3D consistency, while style information is independently injected via the key and value to preserve scene structure. This process is further stabilized by a geometrically-aware latent initialization that provides a coherent starting point for the denoising process. Subsequently, a decoupled reconstruction network lifts these 2D stylized images to 3D Gaussians. A geometry branch predicts a robust 3D scaffold from the original content images, while a parallel style branch predicts the final appearance from our generated stylized images, ensuring structural integrity is not compromised. Extensive experiments on large-scale benchmarks, including RealEstate10K and ACID, demonstrate that GeoStyler significantly outperforms prior arts in stylization quality and multi-view consistency, achieving state-of-the-art performance with a dramatic speedup. Our project page: https://huhuhuxiao.github.io/Geo-Styler/.
Qibin Hu, Ye Zhang 0037, Jisheng Dang, Minglin Chen, Longguang Wang, Yulan Guo
IEEE Trans. Image Process.3
2025 Quality-Guided Dynamic Memory for LLMs-based Long-Term Video Understanding
abstract
Using the impressive learning representation capacity of large language models (LLMs), LLM-based video understanding methods have made significant strides recently. However, most existing methods overlook the crucial importance discrepancy of frames, which often include massive low-quality frames, leading to limited performance and inferior inference efficiency, particularly for long-term videos. To this end, this paper proposes a new video understanding method called quality- guided dynamic memory network (QDM-Net). First, we design a memory quality evolution module (MQEM), which dynamically assigns weights to each frame according to contextual relationships between adjacent frames. Second, we devise a high- level quality memory bank updating mechanism (HQMBU), which selectively maintains high-quality frames in the memory bank, avoiding the negative influences of redundant frames and ensuring that the model focuses on the most informative visual cues. Extensive experiments on long-term video understanding benchmarks demonstrate that our QDM-Net consistently outperforms state-of-the-art methods, showcasing its potential in real-world applications. Our code and model will be publicly available.
Bimei Wang, Jingmei Jiao, Jisheng Dang, Qingrun Jiang, Jiyuan Lin, Zhixuan Chen, Teng Wang 0007
ICME3
2025 Instruction-aware Memory Network for Video Recognition
abstract
The rapid development of multimodal large language models (MLLMs) has highlighted their potential in video understanding. However, challenges remain in long video tasks, particularly in integrating visual features with prompt texts. Existing methods naively store processed video frames in a long-term memory bank, but neglect simple yet effective cross-modal integration. To address this, we introduce the instruction-aware memory construction (IaMC) model for long-term video understanding. By integrating visual and textual information, our model can obtain cross-modal features with robust understanding capabilities. These features are stored in a text-visual memory bank, enabling efficient long-term aggregation without surpassing LLM context or GPU memory limits. Experiments on the LVU dataset demonstrate state-of-the-art performance in video understanding and question answering, showcasing the IaMC model’s effectiveness and setting a new benchmark for long-term video analysis. The source code and trained models will be released publicly.
Bimei Wang, Haijiang Li, Jisheng Dang, Yun Wang 0053, Zhixuan Chen, Jiyuan Lin, Teng Wang 0007
ICME3
2025 AS-Memory: Adaptive Sparse Memory Meeting Video-Language Models
abstract
Long-term video understanding in intelligent transportation systems (ITS) has advanced significantly with the integration of large language models (LLMs) and vision foundation models. However, existing LLM-based multimodal approaches are limited by context length and memory constraints, restricting their effectiveness to short video scenarios. To address these challenges, we propose AS-Memory, a novel framework that combines adaptive sparse memory with LLMs for efficient and scalable long-term video understanding. AS-Memory introduces a plug-and-play memory bank, a lightweight module designed to seamlessly integrate with existing multimodal LLMs. This memory bank stores and retrieves historical video content, enabling long-term analysis while mitigating context length and GPU memory limitations. To further enhance efficiency, we propose a sparse adaptive mechanism that dynamically compresses redundant features and retains critical information, ensuring effective management of streaming video data. Comprehensive evaluations on long-term video understanding benchmarks demonstrate that AS-Memory consistently outperforms state-of-the-art methods in terms of accuracy. The source code and trained models will be made available to the public.
Bimei Wang, Huilin Song, Jisheng Dang, Fei Shen 0004, Mangang Xie, Jizhao Liu, Jia-Si Weng 0001
ICME3
2025 Mitigating Hallucination in Large Video-Language Models with Injected Semantics
abstract
Vision-Language Models (VLMs) have demonstrated remarkable performance across various tasks by encoding visual frames into tokens analogous to textual tokens, which are then processed by a Large Language Model (LLM) for task execution. To manage computational demands, current methods often employ a token compressor, such as Q-former, for efficient inference. However, these methods are typically trained on video-to-text generation loss, lacking sufficient supervision to align intermediate visual representations with textual semantics, resulting in hallucinations when identifying essential objects. To address this issue, we propose a novel visual-textual alignment framework, Semantic Supervision LLM (SS-LLM), which aligns video and text representations within the intermediate feature space, thereby enhancing the LLM’s decoding process. Additionally, we introduce a CLIP Loss to facilitate visual-text alignment in the intermediate feature space, reducing hallucinations in VLMs. Extensive experiments demonstrate that our approach not only mitigates hallucinations more effectively than existing models but also achieves state-of-the-art performance across several benchmarks, providing more accurate and semantically consistent video-text representations. We will make our source code and trained models publicly available.
Bimei Wang, Fan Wen, Jisheng Dang, Huiguo He, Nannan Zhu, Jia-Si Weng 0001
ICME3
2025 A Chaotic Dynamics Framework Inspired by Dorsal Stream for Event Signal Processing
abstract
Event cameras are bio-inspired vision sensors that encode visual information with high dynamic range, high temporal resolution, and low latency. Current state-of-the-art event stream processing methods rely on end-to-end deep learning techniques. However, these models are heavily dependent on data structures, limiting their stability and generalization capabilities across tasks, thereby hindering their deployment in real-world scenarios. To address this issue, we propose a chaotic dynamics event signal processing framework inspired by the dorsal visual pathway of the brain. Specifically, we utilize Continuous-coupled Neural Network (CCNN) to encode the event stream. CCNN encodes polarity-invariant event sequences as periodic signals and polarity-changing event sequences as chaotic signals. We then use continuous wavelet transforms to analyze the dynamical states of CCNN neurons and establish the high-order mappings of the event stream. The effectiveness of our method is validated through integration with conventional classification networks, achieving state-of-the-art classification accuracy on the N-Caltech101 and N-CARS datasets, with results of 84.3% and 99.9%, respectively. Our method improves the accuracy of event camera-based object classification while significantly enhancing the generalization and stability of event representation.
Jing Lian 0001, Zhaofei Yu, Jizhao Liu, Jisheng Dang, Gang Wang 0031
ICML5
2025 Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model
abstract
Recent advances in conditional diffusion models have shown promise for generating realistic TalkingFace videos, yet challenges persist in achieving consistent head movement, synchronized facial expressions, and accurate lip synchronization over extended generations. To address these, we introduce the Motion-priors Conditional Diffusion Model (MCDM), which utilizes both archived and current clip motion priors to enhance motion prediction and ensure temporal consistency. The model consists of three key elements: (1) an archived-clip motion-prior that incorporates historical frames and a reference frame to preserve identity and context; (2) a present-clip motion-prior diffusion model that captures multimodal causality for accurate predictions of head movements, lip sync, and expressions; and (3) a memory-efficient temporal attention mechanism that mitigates error accumulation by dynamically storing and updating motion features. We also introduce the TalkingFace-Wild dataset, a multilingual collection of over 200 hours of footage across 10 languages. Experimental results demonstrate the effectiveness of MCDM in maintaining identity and motion continuity for long-term TalkingFace generation.
Fei Shen 0004, Cong Wang 0018, Junyao Gao 0002, Jisheng Dang, Jinhui Tang 0001, Tat-Seng Chua
ICML5
2025 Diff-LMM: Diffusion Teacher-Guided Spatio-Temporal Perception for Video Large Multimodal Models
abstract
Dynamic spatio-temporal understanding is essential for video-based multimodal tasks, yet existing methods often struggle to capture fine-grained temporal and spatial relationships in long videos. Current approaches primarily rely on pre-trained CLIP encoders, which excel in semantic understanding but lack spatially-aware visual context. This leads to hallucinated results when interpreting fine-grained objects or scenes. To address these limitations, we propose a novel framework that integrates diffusion models into multimodal video models. By employing diffusion encoders at intermediate layers, we enhance visual representations through feature alignment and knowledge distillation losses, significantly improving the model's ability to capture spatial patterns over time. Additionally, we introduce a multi-level alignment strategy to learn robust feature correspondence from pre-trained diffusion models. Extensive experiments on benchmark datasets demonstrate our approach's state-of-the-art performance across multiple video understanding tasks. These results establish diffusion models as a powerful tool for enhancing multimodal video models in complex, dynamic scenarios.
Jisheng Dang, Ligen Chen, Jingze Wu, Ronghao Lin, Bimei Wang, Yun Wang 0053, Nannan Zhu, Teng Wang 0007
IJCAI1
2025 Hallucination Reduction in Video-Language Models via Hierarchical Multimodal Consistency
abstract
The rapid advancement of large language models (LLMs) has led to the widespread adoption of video-language models (VLMs) across various domains. However, VLMs are often hindered by their limited semantic discrimination capability, exacerbated by the limited diversity and biased sample distribution of most video-language datasets. This limitation results in a biased understanding of the semantics between visual concepts, leading to hallucinations. To address this challenge, we propose a Multi-level Multimodal Alignment (MMA) framework that leverages a text encoder and semantic discriminative loss to achieve multi-level alignment. This enables the model to capture both low-level and high-level semantic relationships, thereby reducing hallucinations. By incorporating language-level alignment into the training process, our approach ensures stronger semantic consistency between video and textual modalities. Furthermore, we introduce a two-stage progressive training strategy that exploits larger and more diverse datasets to enhance semantic alignment and better capture general semantic relationships between visual and textual modalities. Our comprehensive experiments demonstrate that the proposed MMA method significantly mitigates hallucinations and achieves state-of-the-art performance across multiple video-language tasks, establishing a new benchmark in the field.
Jisheng Dang, Shengjun Deng, Haochen Chang, Teng Wang 0007, Bimei Wang, Shude Wang, Nannan Zhu, Guo Niu, Jizhao Liu
IJCAI1
2025 External Memory Matters: Generalizable Object-Action Memory for Retrieval-Augmented Long-Term Video Understanding
abstract
Long video understanding with Large Language Models (LLMs) enables the description of objects that are not explicitly present in the training data. However, continuous changes in known objects and the emergence of new ones require up-to-date knowledge of objects and their dynamics for effective understanding of the open world. To alleviate this, we propose an efficient Retrieval-Enhanced Video Understanding method, dubbed REVU, which leverages external knowledge to enhance the performance of open-world learning. First, REVU introduces an extensible external text-object memory with minimal text-visual mapping, involving static and dynamic multimodal information to help LLMs-based models align text and vision features. Second, REVU retrieves object information from external databases and dynamically integrates frame-specific data from videos, enabling effective knowledge aggregation to comprehend the open world. We conducted experiments on multiple benchmark datasets, and our model demonstrates strong adaptability to out-of-domain data without requiring additional fine-tuning or re-training. Experiments on benchmark video understanding datasets reveal that our model achieves state-of-the-art performance and robust generalization.
Jisheng Dang, Huicheng Zheng, Jingmei Jiao, Bimei Wang, Bin Hu 0001, Jian-Huang Lai, Tat-Seng Chua
IJCAI1
2025 Boosting Temporal Sentence Grounding via Causal Inference
abstract
Temporal Sentence Grounding (TSG) aims to identify relevant moments in an untrimmed video that semantically correspond to a given textual query. Despite existing studies having made substantial progress, they often overlook the issue of spurious correlations between video and textual queries. These spurious correlations arise from two primary factors: (1) inherent biases in the textual data, such as frequent co-occurrences of specific verbs or phrases, and (2) the model's tendency to overfit to salient or repetitive patterns in video content. Such biases mislead the model into associating textual cues with incorrect visual moments, resulting in unreliable predictions and poor generalization to out-of-distribution examples. To overcome these limitations, we propose a novel TSG framework, causal intervention and counterfactual reasoning that utilizes causal inference to eliminate spurious correlations and enhance the model's robustness. Specifically, we first formulate the TSG task from a causal perspective with a structural causal model. Then, to address unobserved confounders reflecting textual biases toward specific verbs or phrases, a textual causal intervention is proposed, utilizing do-calculus to estimate the causal effects. Furthermore, visual counterfactual reasoning is performed by constructing a counterfactual scenario that focuses solely on video features, excluding the query and fused multi-modal features. This allows us to debias the model by isolating and removing the influence of the video from the overall effect. Experiments on public datasets demonstrate the superiority of the proposed method. The code is available at https://github.com/Tangkfan/CICR.
Kefan Tang, Lihuo He, Jisheng Dang, Xinbo Gao 0001
ACM Multimedia3
2025 IPAD: Inverse Prompt for AI Detection - A Robust and Interpretable LLM-Generated Text Detector
abstract
Large Language Models (LLMs) have attained human-level fluency in text generation, which complicates the distinguishing between human-written and LLM generated texts. This increases the risk of misuse and highlights the need for reliable detectors. Yet, existing detectors exhibit poor robustness on out-of-distribution (OOD) data and attacked data, which is critical for real-world scenarios. Also, they struggle to provide interpretable evidence to support their decisions, thus undermining reliability. In light of these challenges, we propose IPAD (Inverse Prompt for AI Detection), a novel framework consisting of a Prompt Inverter that identifies predicted prompts that could have generated the input text, and two Distinguishers that examine the probability that the input texts align with the predicted prompts. Empirical evaluations demonstrate that IPAD outperforms the strongest baselines by 9.05% (Average Recall) on in-distribution data, 12.93% (AUROC) on out-of-distribution (OOD) data, and 5.48% (AUROC) on attacked data. IPAD also performs robust on structured datasets. Furthermore, an interpretability assessment is conducted to illustrate that IPAD enhances the AI detection trustworthiness by allowing users to directly examine the decision-making evidence, which provides interpretable support for its state-of-the-art detection results.
Yushi Feng, Jisheng Dang, Changyang He, Yue Deng 0003, Hongxi Pu, Bo Li 0001
NeurIPS3
2025 Adaptive Sparse Memory Networks for Efficient and Robust Video Object Segmentation
abstract
Recently, memory-based networks have achieved promising performance for video object segmentation (VOS). However, existing methods still suffer from unsatisfactory segmentation accuracy and inferior efficiency. The reasons are mainly twofold: 1) during memory construction, the inflexible memory storage mechanism results in a weak discriminative ability for similar appearances in complex scenarios, leading to video-level temporal redundancy, and 2) during memory reading, matching robustness and memory retrieval accuracy decrease as the number of video frames increases. To address these challenges, we propose an adaptive sparse memory network (ASM) that efficiently and effectively performs VOS by sparsely leveraging previous guidance while attending to key information. Specifically, we design an adaptive sparse memory constructor (ASMC) to adaptively memorize informative past frames according to dynamic temporal changes in video frames. Furthermore, we introduce an attentive local memory reader (ALMR) to quickly retrieve relevant information using a subset of memory, thereby reducing frame-level redundant computation and noise in a simpler and more convenient manner. To prevent key features from being discarded by the subset of memory, we further propose a novel attentive local feature aggregation (ALFA) module, which preserves useful cues by selectively aggregating discriminative spatial dependence from adjacent frames, thereby effectively increasing the receptive field of each memory frame. Extensive experiments demonstrate that our model achieves state-of-the-art performance with real-time speed on six popular VOS benchmarks. Furthermore, our ASM can be applied to existing memory-based methods as generic plugins to achieve significant performance improvements. More importantly, our method exhibits robustness in handling sparse videos with low frame rates.
Jisheng Dang, Huicheng Zheng, Xiaohao Xu, Longguang Wang, Qingyong Hu, Yulan Guo
IEEE Trans. Neural Networks Learn. Syst.1
2024 Fact : Teaching MLLMs with Faithful, Concise and Transferable Rationales
abstract
The remarkable performance of Multimodal Large Language Models (MLLMs) has demonstrated their proficient understanding capabilities in handling various visual tasks. Nevertheless, the opaque nature of black-box reasoning processes persists as an enigma, rendering them uninterpretable and struggling with hallucination. Their ability to execute intricate reasoning tasks is also constrained, culminating in stagnation of progression. In this work, we introduce Fact, a novel paradigm designed to generate multimodal rationales that are faithful, concise, and transferable for teaching MLLMs. This paradigm utilizes verifiable visual programming to generate executable code guaranteeing faithfulness. Through a series of operations including pruning, merging, and bridging, the rationale enhances its conciseness. Furthermore, we filter rationales that can be transferred to end-to-end paradigms from programming paradigms to guarantee transferability. Empirical evidence from experiments demonstrates the superiority of Fact across models of varying parameter sizes, significantly enhancing their compositional reasoning and generalization ability and reducing hallucinations owing to its high correlation between images and text.
Minghe Gao, Liang Pang 0001, Yuan Yao 0013, Jisheng Dang, Wenqiao Zhang, Juncheng Li 0006, Siliang Tang, Yueting Zhuang, Tat-Seng Chua
ACM Multimedia5
2024 Beyond Appearance: Multi-Frame Spatio-Temporal Context Memory Networks for Efficient and Robust Video Object Segmentation
abstract
Current video object segmentation approaches primarily rely on frame-wise appearance information to perform matching. Despite significant progress, reliable matching becomes challenging due to rapid changes of the object's appearance over time. Moreover, previous matching mechanisms suffer from redundant computation and noise interference as the number of accumulated frames increases. In this paper, we introduce a multi-frame spatio-temporal context memory (STCM) network to exploit discriminative spatio-temporal cues in multiple adjacent frames by utilizing a multi-frame context interaction module (MCI) for memory construction. Based on the proposed MCI module, a sparse group memory reader is developed to enable efficient sparse matching during memory reading. Our proposed method is generic and achieves state-of-the-art performance with real-time speed on benchmark datasets such as DAVIS and YouTube-VOS. In addition, our model exhibits robustness to sparse videos with low frame rates.
Jisheng Dang, Huicheng Zheng, Xiaohao Xu, Longguang Wang, Yulan Guo
IEEE Trans. Image Process.1
2024 Temporo-Spatial Parallel Sparse Memory Networks for Efficient Video Object Segmentation
abstract
Memory-based networks have achieved tremendous success in video object segmentation. However, these methods still suffer from unfaithful segmentation and inferior efficiency under complicated video scenarios. The reasons are mainly threefold: 1) Weak perception of fast-moving targets due to individual frame memory patterns without capturing inter-frame motion; 2) Lack of discrimination to visually similar appearances due to the limited receptive field; 3) Redundant computation caused by matching with all memorized frames. To address these issues, we propose a Temporo-Spatial Parallel Sparse Memory network (TSPSM) for efficient video object segmentation. Our TSPSM constructs a temporal memory bank and a spatial memory bank in parallel to memorize complementary discriminative object cues. The temporal bank exploits discriminative temporal motion cues, while the spatial bank mines spatial context cues between adjacent frames with large receptive fields, thereby alleviating the ambiguity caused by similar instances and fast movements. To reduce redundant computation without sacrificing performance during the matching step, we further design a parallel sparse memory reader based on the constructed informative memory banks, which efficiently retrieves relevant temporal and spatial information in a parallel way. Experiments demonstrate that our TSPSM achieves state-of-the-art performance with real-time speed on DAVIS, and YouTube-VOS benchmarks. Furthermore, extensive experiments show that the proposed TSPMC module can be applied to existing methods as a generic plugin to significantly improve performance.
Jisheng Dang, Huicheng Zheng, Bimei Wang, Longguang Wang, Yulan Guo
IEEE Trans. Intell. Transp. Syst.1
2024 Unified Spatio-Temporal Dynamic Routing for Efficient Video Object Segmentation
abstract
Existing methods for video object segmentation (VOS) have achieved significant success by performing semantic guidance, spatial constraint, or temporal consistency. However, VOS still remains highly challenging because it is difficult to collaboratively leverage spatial constraint, temporal consistency, and semantic guidance while reducing redundant information. In this paper, we propose an efficient unified spatio-temporal dynamic routing (STDR) framework to address VOS by achieving a better spatio-temporal balance while avoiding redundancy. Specifically, our unified spatio-temporal modeling contains three paths: 1) short-term spatial path is employed to mine the spatial constraints from the previous frame; 2) long-term semantic path is used to capture semantic cues from the first reference frame with ground-truth labels; 3) memory queue path is designed to efficiently exploit the temporal consistency of middle frames with a compact memory bank of constant size. To enhance the input of each path, we introduce a progressive contextual memory enhancement module to exploit the contextualized memory with growing receptive fields by progressively aggregating spatial contextual information from adjacent frames for each memory frame. Furthermore, we design a dynamic memory-routed module to globally refine the outputs of our three paths for unified modeling. Enhanced by the proposed modules, our STDR achieves state-of-the-art performance with fast speed on the DAVIS 2016, DAVIS 2017 Val/Test, YouTube-VOS 2018/2019, and real-world long-video benchmarks.
Jisheng Dang, Huicheng Zheng, Xiaohao Xu, Yulan Guo
IEEE Trans. Intell. Transp. Syst.1
2023 Dual-Memory Feature Aggregation for Video Object Detection
Diwei Fan, Huicheng Zheng, Jisheng Dang
PRCV (6)3
2023 Efficient and Robust Video Object Segmentation Through Isogenous Memory Sampling and Frame Relation Mining
abstract
Recently, memory-based methods have achieved remarkable progress in video object segmentation. However, the segmentation performance is still limited by error accumulation and redundant memory, primarily because of 1) the semantic gap caused by similarity matching and memory reading via heterogeneous key-value encoding; 2) the continuously growing and inaccurate memory through directly storing unreliable predictions of all previous frames. To address these issues, we propose an efficient, effective, and robust segmentation method based on Isogenous Memory Sampling and Frame-Relation mining (IMSFR). Specifically, by utilizing an isogenous memory sampling module, IMSFR consistently conducts memory matching and reading between sampled historical frames and the current frame in an isogenous space, minimizing the semantic gap while speeding up the model through an efficient random sampling. Furthermore, to avoid key information loss during the sampling process, we further design a frame-relation temporal memory module to mine inter-frame relations, thereby effectively preserving contextual information from the video sequence and alleviating error accumulation. Extensive experiments demonstrate the effectiveness and efficiency of the proposed IMSFR method. In particular, our IMSFR achieves state-of-the-art performance on six commonly used benchmarks in terms of region similarity & contour accuracy and speed. Our model also exhibits strong robustness against frame sampling due to its large receptive field.
Jisheng Dang, Huicheng Zheng, Jinming Lai, Xu Yan 0005, Yulan Guo
IEEE Trans. Image Process.1
2022 LHPHGCNN: Lightweight Hierarchical Parallel Heterogeneous Group Convolutional Neural Networks for Point Cloud Scene Prediction
abstract
Many previous works have achieved tremendous success for point cloud processing. However, they still suffer from inefficiency in memory and computation. In this paper, we introduce Lightweight Hierarchical Parallel Heterogeneous Group Convolutional Neural Networks (LHPHGCNN), an efficient and lightweight neural architecture to achieve better performance but lower complexity than most existing methods for point cloud processing. By designing different local structure encodings, LHPHGCNN fully mines rich local geometric features. Additionally, we further propose Hierarchical Parallel Heterogeneous Group Convolution (HPHGConv) to simultaneously capture the discriminative nonlocal features and fine-grained local geometric features of point clouds in heterogeneous groups with fewer parameters and lower computing costs, which helps to recognize elusive shapes. To further capture the contextual features along with rich semantics, we introduce a novel multi-scale semantics (MSS) strategy to progressively increase the receptive field for each local area through the information communication between different scale areas. Extensive experiments show that our LHPHGCNN significantly outperforms state-of-the-art approaches for shape classification on ModelNet40 and semantic segmentation on three large scale benchmarks S3DIS, vKITTI, ScanNet, and SemanticKITTI in terms of accuracy and complexity.
Jisheng Dang
IEEE Trans. Intell. Transp. Syst.1
2021 HIGCNN: Hierarchical Interleaved Group Convolutional Neural Networks for Point Clouds Analysis
abstract
Although previous works for point clouds analysis have achieved remarkable performance, it is difficult for them to achieve a good trade-off between accuracy and complexity. In this paper, we present an efficient and lightweight neural network for point clouds analysis, named HIGCNN, which can achieve better performance but lower complexity compared to existing methods. The key component in our approach is the hierarchical interleaved group convolution (HIGConv) operation. We first present a neighborhood attention convolution (NAC) operation to fully mine fine-grained local geometric features inside each local area. With the proposed NAC, we further design a HIGConv to encode both fine-grained local geometric features and discriminative nonlocal point-wise features with fewer parameters and lower computational costs. To further capture fine-grained contextual features, we propose a multi-scale relation (MSR) module to fully explore the relationship among different scale areas. Extensive experiments show that our HIGCNN surpasses state-of-the-art approaches for classification and semantic segmentation on four benchmarks ModelNet40, S3DIS, vKITTI and SemanticKITTI in terms of accuracy and complexity.
Jisheng Dang
ICASSP1
2020 HPGCNN: Hierarchical Parallel Group Convolutional Neural Networks for Point Clouds Processing
Jisheng Dang
ACCV (1)1
2020 PVFNet: Point-View Fusion Network for 3D Shape Recognition
Jisheng Dang
KSEM (1)2