EDBT 2026 Demo / reviewers in the wild / expert
Mohan Chen 0001
dblp:254/0459-1
· DBLP profile ↗
12ranked-venue papers
3as first author
11since 2021 · last 2025
0000-0002-9824-5852ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Sketch-based Point Cloud Generation with Diffusion Model and Pre-training EnhancementabstractDiffusion models, known for their success in various generative tasks like image generation and super-resolution, are applied in this study for point cloud generation, a field that has not been extensively explored due to the complexity of point clouds. We propose a novel method using a diffusion model to generate high-quality 3D point clouds from 2D sketches. This method employs a self-supervised contrastive learning scheme to align sketch and point cloud modalities. Additionally, it incorporates a specific partition mixing strategy to integrate edge information during pre-training. Evaluated on two benchmark datasets, our method outperforms existing state-of-the-art approaches, showcasing the potential of diffusion models in point cloud generation and setting a new direction for future research. Yangdong Chen, Mohan Chen 0001, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003 |
ICASSP | 2 |
| 2025 | TGSAM-2: Text-Guided Medical Image Segmentation Using Segment Anything Model 2
Runtian Yuan, Ling Zhou 0002, Jilan Xu, Qingqiu Li, Mohan Chen 0001, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003 |
MICCAI (10) | 5 |
| 2025 | Text-Promptable Propagation for Referring Medical Image Sequence SegmentationabstractReferring Medical Image Sequence Segmentation (Ref-MISS) is a novel and challenging task that aims to segment anatomical structures in medical image sequences (e.g., endoscopy, ultrasound, CT, and MRI) based on natural language descriptions. Existing 2D and 3D segmentation models struggle to explicitly track objects of interest across medical image sequences, and lack support for interactive, text-driven guidance. To address these limitations, we propose Text-Promptable Propagation (TPP), which enables the recognition of referred objects through cross-modal referring interaction, and maintains continuous tracking across the sequence via Transformer-based triple propagation, using text embeddings as queries. To support this task, we curate a large-scale benchmark, Ref-MISS-Bench, which covers 4 imaging modalities and 20 different organs and lesions. Experimental results on this benchmark demonstrate that TPP consistently outperforms state-of-the-art methods in both medical segmentation and referring video object segmentation. Code and data are available at https://github.com/yuanruntian/TPP. Runtian Yuan, Mohan Chen 0001, Jilan Xu, Ling Zhou 0002, Qingqiu Li, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003 |
ACM Multimedia | 2 |
| 2024 | Temporal Feature Aggregation for Efficient 2D Video GroundingabstractVideo grounding aims to locate the target video moment in an untrimmed video based on a text query. Most existing methods employ 3D CNNs as the video feature extractor, incurring substantial computational costs. Only a few methods use 2D backbones for video feature extraction, and they suffer from diminished accuracy due to the inherent lack of temporal information within 2D features. To address this problem, we propose a novel 2D video grounding method called TFA that improves accuracy while minimizing computational costs. Our approach involves a query-guided temporal feature aggregation module designed to explicitly capture temporal information. We disentangle time intervals of input video frames and prediction spans to reduce computational overhead. Additionally, we introduce deformable attention into the multi-modal encoder for further enhancement. Extensive experiments on two public datasets demonstrate that our method outperforms previous 2D video grounding methods and achieves competitive results with most 3D methods at significantly reduced costs. Mohan Chen 0001, Yiren Zhang, Jueqi Wei, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003 |
ICME | 1 |
| 2024 | Memory-Augmented Transformer for Efficient End-to-End Video GroundingabstractVideo grounding aims to localize a specific segment corresponding to a text query in an untrimmed video. Due to the tremendous computational cost required to process the video frames, the de facto paradigm of video grounding is to extract video features using pretrained video encoders. The parameters of the video encoders are fixed during training, which limits the performance of the localization model. To solve this problem, we propose a Memory-Augmented Transformer (MAT) model. Specifically, each video is split into non-overlapping clips, and our MAT processes videos in a clip-by-clip manner while caching video features into FIFO cached memory queues. By enabling early return, our MAT outperforms previous methods with only less than 60% frames seen. Extensive experimental results on three public benchmark datasets demonstrate that our MAT can achieve competitive performance while being much more efficient than currently prevailing two-stage methods. Code is available at https://github.com/xuyw1997/MAT. Yuanwu Xu, Mohan Chen 0001, Yuejie Zhang, Rui Feng 0001, Tao Zhang 0022, Shang Gao 0003 |
ICME | 2 |
| 2023 | Enhanced Knowledge Injection for Radiology Report GenerationabstractAutomatic generation of radiology reports holds crucial clinical value, as it can alleviate substantial workload on radiologists and remind less experienced ones of potential anomalies. Despite the remarkable performance of various image captioning methods in the natural image field, generating accurate reports for medical images still faces challenges, i.e., disparities in visual and textual data, and lack of accurate domain knowledge. To address these issues, we propose an enhanced knowledge injection framework, which utilizes two branches to extract different types of knowledge. The Weighted Concept Knowledge (WCK) branch is responsible for introducing clinical medical concepts weighted by TF-IDF scores. The Multimodal Retrieval Knowledge (MRK) branch extracts triplets from similar reports, emphasizing crucial clinical information related to entity positions and existence. By integrating this finer-grained and well-structured knowledge with the current image, we are able to leverage the multi-source knowledge gain to ultimately facilitate more accurate report generation. Extensive experiments have been conducted on two public benchmarks, demonstrating that our method achieves superior performance over other state-of-the-art methods. Ablation studies further validate the effectiveness of two extracted knowledge sources. Qingqiu Li, Jilan Xu, Runtian Yuan, Mohan Chen 0001, Yuejie Zhang, Rui Feng 0001, Shang Gao 0003 |
BIBM | 4 |
| 2023 | Conditional Video-Text Reconstruction Network with Cauchy Mask for Weakly Supervised Temporal Sentence GroundingabstractTemporal sentence grounding aims to detect the target segment most related to a given query in an untrimmed video. To alleviate the expensive annotation cost for temporal labels, researchers paid more attention to weakly supervised setting. Prior studies neglected the utilization of video representation reconstruction, which led to an unbalanced alignment learning. Moreover, they used different strategies to generate proposals which ignored the temporal structure in a query. In this paper, we propose a novel Conditional Video-Text Reconstruction Network (CVTRN). It supports conditional reconstruction of video and text representation. Specifically, video and text features are fused to compute semantic alignment, which is the condition of reconstruction. A new mask strategy for mask conditioned sentence reconstruction is also devised. This strategy focuses more on boundary regions than the widely used Gaussian mask in previous methods. Experimental results on two public benchmark datasets show that our CVTRN outperforms the state-of-the-art methods. Jueqi Wei, Yuanwu Xu, Mohan Chen 0001, Yuejie Zhang, Rui Feng 0001, Shang Gao 0003 |
ICME | 3 |
| 2023 | SPTNET: Span-based Prompt Tuning for Video GroundingabstractWhen a Pre-trained Language Model (PLM) is adopted in video grounding task, it usually acts as a text encoder without having its knowledge fully utilized. Also, there exists an inconsistency problem between the pre-training and downstream objectives. To solve the issues, we propose a new paradigm, named Span-based Prompt Tuning (SPTNet). It can convert the video grounding task into a cloze form. Specifically, a query is first changed into a form with mask token by a template, then the video and the query embeddings are integrated through a cross-modal transformer. The start and end points of the query matching time span are predicted with the embedding of the mask token. Experimental results on two public benchmarks ActivityNet Captions and Charades-STA show that our SPTNet achieves surpassing performance compared with state-of-the-art methods. Yiren Zhang, Yuanwu Xu, Mohan Chen 0001, Yuejie Zhang, Rui Feng 0001, Shang Gao 0003 |
ICME | 3 |
| 2023 | Dynamic Graph Message Passing NetworksabstractModelling long-range dependencies is critical for scene understanding tasks in computer vision. Although convolution neural networks (CNNs) have excelled in many vision tasks, they are still limited in capturing long-range structured relationships as they typically consist of layers of local kernels. A fully-connected graph, such as the self-attention operation in Transformers, is beneficial for such modelling, however, its computational overhead is prohibitive. In this paper, we propose a dynamic graph message passing network, that significantly reduces the computational complexity compared to related works modelling a fully-connected graph. This is achieved by adaptively sampling nodes in the graph, conditioned on the input, for message passing. Based on the sampled nodes, we dynamically predict node-dependent filter weights and the affinity matrix for propagating information between them. This formulation allows us to design a self-attention module, and more importantly a new Transformer-based backbone network, that we use for both image classification pretraining, and for addressing various downstream tasks (e.g. object detection, instance and semantic segmentation). Using this model, we show significant improvements with respect to strong, state-of-the-art baselines on four different tasks. Our approach also outperforms fully-connected graphs while using substantially fewer floating-point operations and parameters. Code and models will be made publicly available at https://github.com/fudan-zvg/DGMN2. Li Zhang 0040, Mohan Chen 0001, Anurag Arnab, Xiangyang Xue 0001, Philip Torr 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Rethinking Local and Global Feature Representation for Dense Prediction
Mohan Chen 0001, Li Zhang 0040, Rui Feng 0001, Xiangyang Xue 0001, Jianfeng Feng |
Pattern Recognit. | 1 |
| 2021 | Rethinking local and global feature representation for semantic segmentation
Mohan Chen 0001, Xinxuan Zhao, Bingfei Fu, Li Zhang 0040, Xiangyang Xue 0001 |
BMVC | 1 |
| 2020 | Dynamic Depth Fusion and Transformation for Monocular 3D Object Detection
Erli Ouyang, Li Zhang 0040, Mohan Chen 0001, Anurag Arnab, Yanwei Fu 0001 |
ACCV (1) | 3 |