EDBT 2026 Demo / reviewers in the wild / expert
Fangcen Liu
dblp:256/6631
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Fusion-Enhanced Network for Infrared and Visible High-Level Vision TasksabstractInfrared and visible dual-modality vision tasks such as semantic segmentation, object detection, and salient object detection can achieve robust performance even in extreme scenes by leveraging complementary information. However, most existing image fusion-based methods and task-specific frameworks exhibit limited generalization across multiple tasks. Moreover, summing the general representations obtained from foundation models poses challenges, including insufficient semantic information mining and feature fusion. In this paper, we propose a fusion-enhanced network, which effectively enriches semantic information and integrates features based on the complementary characteristics of infrared and visible modalities. The proposed network can extend to high-level vision tasks, showing strong generalization capabilities. Firstly, we adopt the infrared and visible foundation models to extract the general representations. Then, to enrich the semantic information of these general representations for high-level vision tasks, we design the feature enhancement module and the token enhancement module for feature maps and tokens, respectively. Besides, the attention-guided fusion module is proposed for effective fusion by exploring the complementary information of two modalities. Moreover, we adopt the cutout&mix augmentation strategy to conduct the data augmentation, which further improves the ability of the model to mine the regional complementarity between the two modalities. Extensive experiments show that the proposed method outperforms state-of-the-art dual-modality methods in the semantic segmentation, object detection, and salient object detection tasks. Fangcen Liu, Chenqiang Gao, Pengcheng Li 0017, Junjie Guo, Deyu Meng |
IEEE Trans. Multim. | 1 |
| 2024 | DAMSDet: Dynamic Adaptive Multispectral Detection Transformer with Competitive Query Selection and Adaptive Feature Fusion
Junjie Guo, Chenqiang Gao, Fangcen Liu, Deyu Meng, Xinbo Gao 0001 |
ECCV (27) | 3 |
| 2024 | InfMAE: A Foundation Model in the Infrared Modality
Fangcen Liu, Chenqiang Gao, Yaming Zhang, Junjie Guo, Deyu Meng |
ECCV (18) | 1 |
| 2024 | TopologyFormer: structure transformer assisted topology reconstruction for point cloud completion
Zhenwei Jiang, Chenqiang Gao, Chuandong Liu, Fangcen Liu, Lijie Zhu |
Multim. Tools Appl. | 5 |
| 2024 | THISNet: Tooth Instance Segmentation on 3D Dental Models via Highlighting Tooth RegionsabstractAutomatic tooth instance segmentation on 3D dental models is crucial for digitizing dental treatments and enabling computer-assisted treatment planning. However, It is challenging since the tight arrangement of dental structures and the consequential impact of dental ailments on their morphological characteristics. To address these challenges, we propose a novel method called THISNet. Unlike existing methods, THISNet focuses on highlighting tooth regions rather than relying on bounding box detection, leading to improved accuracy in tooth segmentation and labeling. By incorporating the highlighted tooth regions with a tooth object affinity module, our method effectively integrates global contextual information, considering the relationships between neighboring teeth and their surrounding structures. THISNet adopts an end-to-end learning approach, reducing complexity and enhancing segmentation efficiency compared to multi-stage training methods. Experimental results demonstrate the superiority of THISNet over existing approaches, highlighting its potential in various dental clinical applications. Pengcheng Li 0017, Chenqiang Gao, Fangcen Liu, Deyu Meng, Yan Yan 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Hierarchical Supervision and Shuffle Data Augmentation for 3D Semi-Supervised Object DetectionabstractState-of-the-art 3D object detectors are usually trained on large-scale datasets with high-quality 3D annotations. However, such 3D annotations are often expensive and time-consuming, which may not be practical for real applications. A natural remedy is to adopt semi-supervised learning (SSL) by leveraging a limited amount of labeled samples and abundant unlabeled samples. Current pseudo-labeling-based SSL object detection methods mainly adopt a teacher-student framework, with a single fixed threshold strategy to generate supervision signals, which inevitably brings confused supervision when guiding the student network training. Besides, the data augmentation of the point cloud in the typical teacher-student framework is too weak, and only contains basic down sampling and flip-and-shift (i.e., rotate and scaling), which hinders the effective learning of feature information. Hence, we address these issues by introducing a novel approach of Hierarchical Supervision and Shuffle Data Augmentation (HSSDA), which is a simple yet effective teacher-student framework. The teacher network generates more reasonable supervision for the student network by designing a dynamic dual-threshold strategy. Besides, the shuffle data augmentation strategy is designed to strengthen the feature representation ability of the student network. Extensive experiments show that HSSDA consistently outperforms the recent state-of-the-art methods on different datasets. The code will be released at https://github.com/azhuantou/HSSDA. Chuandong Liu, Chenqiang Gao, Fangcen Liu, Pengcheng Li 0017, Deyu Meng, Xinbo Gao 0001 |
CVPR | 3 |
| 2023 | Infrared Small and Dim Target Detection With Transformer Under Complex BackgroundsabstractThe infrared small and dim (S&D) target detection is one of the key techniques in the infrared search and tracking system. Since the local regions similar to infrared S&D targets spread over the whole background, exploring the correlation amongst image features in large-range dependencies to mine the difference between the target and background is crucial for robust detection. However, existing deep learning-based methods are limited by the locality of convolutional neural networks, which impairs the ability to capture large-range dependencies. Additionally, the S&D appearance of the infrared target makes the detection model highly possible to miss detection. To this end, we propose a robust and general infrared S&D target detection method with the transformer. We adopt the self-attention mechanism of the transformer to learn the correlation of image features in a larger range. Moreover, we design a feature enhancement module to learn discriminative features of S&D targets to avoid miss-detections. After that, to avoid the loss of the target information, we adopt a decoder with the U-Net-like skip connection operation to contain more information of S&D targets. Finally, we get the detection result by a segmentation head. Extensive experiments on two public datasets show the obvious superiority of the proposed method over state-of-the-art methods, and the proposed method has a stronger generalization ability and better noise tolerance. Fangcen Liu, Chenqiang Gao, Deyu Meng, Wangmeng Zuo, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2022 | SS3D: Sparsely-Supervised 3D Object Detection from Point CloudabstractConventional deep learning based methods for 3D object detection require a large amount of 3D bounding box annotations for training, which is expensive to obtain in practice. Sparsely annotated object detection, which can largely reduce the annotations, is very challenging since the missing-annotated instances would be regarded as the background during training. In this paper, we propose a sparsely-supervised 3D object detection method, named SS3D. Aiming to eliminate the negative supervision caused by the missing annotations, we design a missing-annotated instance mining module with strict filtering strategies to mine positive instances. In the meantime, we design a reliable background mining module and a point cloud filling data augmentation strategy to generate the confident data for iteratively learning with reliable supervision. The proposed SS3D is a general framework that can be used to learn any modern 3D object detector. Extensive experiments on the KITTI dataset reveal that on different 3D detectors, the proposed SS3D framework with only 20% annotations required can achieve onpar performance comparing to fully-supervised methods. Comparing with the state-of-the-art semi-supervised 3D objection detection on KITTI, our SS3D improves the benchmarks by significant margins under the same annotation workload. Moreover, our SS3D also out-performs the state-of-the-art weakly-supervised method by remarkable margins, highlighting its effectiveness. Chuandong Liu, Chenqiang Gao, Fangcen Liu, Jiang Liu 0011, Deyu Meng, Xinbo Gao 0001 |
CVPR | 3 |
| 2021 | Infrared and Visible Cross-Modal Image Retrieval Through Shared FeaturesabstractImage retrieval is one of the key techniques of computer vision, and has been studied for a long time. Nevertheless, little attention is paid to infrared and visible cross-modal retrieval which can be widely used in various applications, e.g., infrared and visible surveillance systems. In this paper, we propose a shared features based infrared-visible cross-modal image retrieval method. The similar visual features are extracted from infrared and visible images as the shared features, and the Euclidean distance is used to measure the similarity between these features. The core of the proposed method comes from three aspects: 1) Feature separation network can separate image features into shared features and exclusive features; 2) Maximum Mean Discrepancy (MMD) loss is employed to constrain the distribution of shared features, which can reduce the retrieval error caused by different imaging angles and similarity of infrared images. 3) The cross-layer fusion encoder compensates for the context loss in the convolution of infrared images. Experimental results on the Infrared-Visible dataset demonstrate the proposed method is effective and outperforms the state-of-the-art approaches. Fangcen Liu, Chenqiang Gao, Yongqing Sun, Yue Zhao 0012, Feng Yang 0015, Anyong Qin, Deyu Meng |
IEEE Trans. Circuits Syst. Video Technol. | 1 |