EDBT 2026 Demo / reviewers in the wild / expert
Yifan Liu 0001
dblp:23/4955-1
· DBLP profile ↗
48ranked-venue papers
8as first author
38since 2021 · last 2026
0000-0002-2746-8186ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 6 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 32 · 6 first-author · 25 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HumanRecon: Neural reconstruction of dynamic human using geometric cues and physical priors
Junhui Yin, Wei Yin 0006, Hao Chen 0041, Xuqian Ren, Zhanyu Ma, Jun Guo 0002, Yifan Liu 0001 |
Pattern Recognit. | 7 |
| 2026 | Retrieval-Enhanced Visual Prompt Learning for Few-Shot ClassificationabstractThe Contrastive Language-Image Pretraining (CLIP) model has been widely used in various downstream vision tasks. The few-shot learning paradigm has been widely adopted to augment its capacity for these tasks. However, current paradigms may struggle with fine-grained classification, such as satellite image recognition, due to widening domain gaps. To address this limitation, we propose retrieval-enhanced visual prompt learning (RePrompt), which introduces retrieval mechanisms to cache and reuse the knowledge of downstream tasks. RePrompt constructs a retrieval database from either training examples or external data if available, and uses a retrieval mechanism to enhance multiple stages of a simple prompt learning baseline, thus narrowing the domain gap. During inference, our enhanced model can reference similar samples brought by retrieval to make more accurate predictions. A detailed analysis reveals that retrieval helps to improve the distribution of late features, thus, improving generalization for downstream tasks. RePrompt attains state-of-the-art performance on a wide range of vision datasets, including 11 image datasets, 3 video datasets, 1 multi-view dataset, and 4 domain generalization benchmarks. Jintao Rong 0001, Hao Chen 0041, Linlin Ou, Tianxiao Chen, Yifan Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Efficient Document Shadow Removal with Contrast-Aware Guidance
Yifan Liu 0001, Jiyu Wu, Jiancheng Huang, Mingfu Yan, Yi Huang 0035, Shifeng Chen |
CGI (3) | 1 |
| 2025 | Component Adaptive Clustering for Generalized Category DiscoveryabstractGeneralized Category Discovery (GCD) tackles the challenging problem of categorizing unlabeled images into both known and novel classes within a partially labeled dataset, without prior knowledge of the number of unknown categories. Traditional methods often rely on rigid assumptions, such as predefining the number of classes, which limits their ability to handle the inherent variability and complexity of real-world data. To address these shortcomings, we propose AdaGCD, a cluster-centric contrastive learning framework that incorporates Adaptive Slot Attention (AdaSlot) into the GCD framework. AdaSlot dynamically determines the optimal number of slots based on data complexity, removing the need for predefined slot counts. This adaptive mechanism facilitates the flexible clustering of unlabeled data into known and novel categories by dynamically allocating representational capacity. By integrating adaptive representation with dynamic slot allocation, our method captures both instance-specific and spatially clustered features, improving class discovery in open-world scenarios. Extensive experiments on public and fine-grained datasets validate the effectiveness of our framework, emphasizing the advantages of leveraging spatial local information for category discovery in unlabeled image datasets. Mingfu Yan, Jiancheng Huang, Yifan Liu 0001, Shifeng Chen |
ICME | 3 |
| 2025 | Dual-Schedule Inversion: Training- and Tuning-Free Inversion for Real Image EditingabstractText-conditional image editing is a practical AIGC task that has recently emerged with great commercial and academic value. For real image editing, most diffusion model-based methods use DDIM Inversion as the first stage before editing. However, DDIM Inversion often results in reconstruction failure, leading to unsatisfactory performance for downstream editing. To address this problem, we first analyze why the reconstruction via DDIM Inversion fails. We then propose a new inversion and sampling method named Dual-Schedule Inversion. We also design a classifier to adaptively combine Dual-Schedule Inversion with different editing methods for user-friendly image editing. Our work can achieve superior reconstruction and editing performance with the following advantages: 1) It can reconstruct real images perfectly without fine-tuning, and its reversibility is guaranteed mathematically. 2) The edited object/scene conforms to the semantics of the text prompt. 3) The unedited parts of the object/scene retain the original identity. Jiancheng Huang, Yi Huang 0035, Jianzhuang Liu, Yifan Liu 0001, Shifeng Chen |
WACV | 5 |
| 2025 | Diffusion Model-Based Image Editing: A SurveyabstractDenoising diffusion models have emerged as a powerful tool for various image generation and editing tasks, facilitating the synthesis of visual content in an unconditional or input-conditional manner. The core idea behind them is learning to reverse the process of gradually adding noise to images, allowing them to generate high-quality samples from a complex distribution. In this survey, we provide an exhaustive overview of existing methods using diffusion models for image editing, covering both theoretical and practical aspects in the field. We delve into a thorough analysis and categorization of these works from multiple perspectives, including learning strategies, user-input conditions, and the array of specific editing tasks that can be accomplished. In addition, we pay special attention to image inpainting and outpainting, and explore both earlier traditional context-driven and current multimodal conditional methods, offering a comprehensive analysis of their methodologies. To further evaluate the performance of text-guided image editing algorithms, we propose a systematic benchmark, EditEval, featuring an innovative metric, LMM Score. Finally, we address current limitations and envision some potential directions for future research. Yi Huang 0035, Jiancheng Huang, Yifan Liu 0001, Mingfu Yan, Jiaxi Lv, Jianzhuang Liu, Wei Xiong 0008, He Zhang 0004, Liangliang Cao, Shifeng Chen |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Amodal Scene Analysis via Holistic Occlusion Relation Inference and Generative Mask CompletionabstractAmodal scene analysis entails interpreting the occlusion relationship among scene elements and inferring the possible shapes of the invisible parts. Existing methods typically frame this task as an extended instance segmentation or a pair-wise object de-occlusion problem. In this work, we propose a new framework, which comprises a Holistic Occlusion Relation Inference (HORI) module followed by an instance-level Generative Mask Completion (GMC) module. Unlike previous approaches, which rely on mask completion results for occlusion reasoning, our HORI module directly predicts an occlusion relation matrix in a single pass. This approach is much more efficient than the pair-wise de-occlusion process and it naturally handles mutual occlusion, a common but often neglected situation. Moreover, we formulate the mask completion task as a generative process and use a diffusion-based GMC module for instance-level mask completion. This improves mask completion quality and provides multiple plausible solutions. We further introduce a large-scale amodal segmentation dataset with high-quality human annotations, including mutual occlusions. Experiments on our dataset and two public benchmarks demonstrate the advantages of our method. code public available at https://github.com/zbwxp/Amodal-AAAI. Bowen Zhang 0009, Qing Liu 0017, Jianming Zhang 0001, Yilin Wang 0002, Liyang Liu, Zhe Lin 0001, Yifan Liu 0001 |
AAAI | 7 |
| 2024 | MambaDW: Semantic-Aware Mamba for Document Watermark Removal
Yifan Liu 0001, Mingfu Yan, He Hua, Jiancheng Huang, Shifeng Chen |
CGI (1) | 1 |
| 2024 | Robust Lightweight Depth Estimation Model via Data-Free DistillationabstractExisting Monocular Depth Estimation (MDE) methods often use large and complex neural networks. Despite the advanced performance of these methods, we consider the efficiency and generalization for practical applications with limited resources. In our paper, we present an efficient transformer-based monocular relative depth estimation network and train it with a diverse depth dataset to obtain good generalization performance. Knowledge distillation (KD) is employed to transfer the general knowledge from a pre-trained teacher network to the compact student network, demonstrating that KD can improve the generalization ability as well as the accuracy. Moreover, we propose a geometric label-free distillation method to improve the lightweight model in specific domains utilizing 3D geometric cues with unlabeled data. We show that our method outperforms other KD methods with or without ground truth supervision. Finally, we propose an application of the lightweight network to a two-stage depth completion task. Our method shows on par or even superior cross-domain generalization ability compared to large networks. Wei Yin 0006, Yifan Liu 0001, Zengchang Qin |
ICASSP | 4 |
| 2024 | Entwined Inversion: Tune-Free Inversion For Real Image Faithful Reconstruction and EditingabstractText-conditional image editing is a very practical AIGC task that has recently emerged with great commercial and academic research value. For real image editing, most diffusion model-based methods use DDIM Inversion as the first stage before editing, but DDIM Inversion often results in reconstruction failure, leading to unsatisfactory performance for all downstream edits. In order to solve this problem, we first mathematically analyze the reason for the reconstruction failure of DDIM Inversion, and then propose a new inversion and sampling method named Entwined Inversion that can achieve satisfactory reconstruction and editing performance, which can solve two major problems: 1) the object can retain the main content of the original image; 2) the edited object can conform to the semantics of the text prompt. In addition, our method does not require training the diffusion model itself on a large dataset, nor does it require any fine-tuning for some particular images. Jiancheng Huang, Yifan Liu 0001, Jiaxi Lv, Shifeng Chen |
ICASSP | 2 |
| 2024 | ICGNet: A Unified Approach for Instance-Centric GraspingabstractAccurate grasping is the key to several robotic tasks including assembly and household robotics. Executing a successful grasp in a cluttered environment requires multiple levels of scene understanding: First, the robot needs to analyze the geometric properties of individual objects to find feasible grasps. These grasps need to be compliant with the local object geometry. Second, for each proposed grasp, the robot needs to reason about the interactions with other objects in the scene. Finally, the robot must compute a collision-free grasp trajectory while taking into account the geometry of the target object. Most grasp detection algorithms directly predict grasp poses in a monolithic fashion, which does not capture the composability of the environment. In this paper, we introduce an end-to-end architecture for object-centric grasping. The method uses pointcloud data from a single arbitrary viewing direction as an input and generates an instance-centric representation for each partially observed object in the scene. This representation is further used for object reconstruction and grasp detection in cluttered table-top scenes. We show the effectiveness of the proposed method by extensively evaluating it against state-of-the-art methods on synthetic datasets, indicating superior performance for grasping and reconstruction. Additionally, we demonstrate real-world applicability by decluttering scenes with varying numbers of objects. Videos and Code icgraspnet.github.io. René Zurbrügg, Yifan Liu 0001, Francis Engelmann, Suryansh Kumar 0001, Marco Hutter 0001, Vaishakh Patil, Fisher Yu 0001 |
ICRA | 2 |
| 2024 | BPKD: Boundary Privileged Knowledge Distillation For Semantic SegmentationabstractCurrent knowledge distillation approaches in semantic segmentation tend to adopt a holistic approach that treats all spatial locations equally. However, for dense prediction, students’ predictions on edge regions are highly uncertain due to contextual information leakage, requiring higher spatial sensitivity knowledge than the body regions. To address this challenge, this paper proposes a novel approach called boundary-privileged knowledge distillation (BPKD). BPKD distills the knowledge of the teacher model’s body and edges separately to the compact student model. Specifically, we employ two distinct loss functions: (i) edge loss, which aims to distinguish between ambiguous classes at the pixel level in edge regions; (ii) body loss, which utilizes shape constraints and selectively attends to the inner-semantic regions. Our experiments demonstrate that the proposed BPKD method provides extensive refinements and aggregation for edge and body regions. Additionally, the method achieves state-of-the-art distillation performance for semantic segmentation on three popular benchmark datasets, highlighting its effectiveness and generalization ability. BPKD shows consistent improvements across a diverse array of lightweight segmentation structures, including both CNNs and transformers, underscoring its architecture-agnostic adaptability. The code is available at https://github.com/AkideLiu/BPKD. Liyang Liu, Minh Hieu Phan, Bowen Zhang 0009, Jinchao Ge, Yifan Liu 0001 |
WACV | 6 |
| 2024 | Scaling Up Multi-domain Semantic Segmentation with Sentence Embeddings
Wei Yin 0006, Yifan Liu 0001, Chunhua Shen, Baichuan Sun, Anton van den Hengel |
Int. J. Comput. Vis. | 2 |
| 2024 | SegViT v2: Exploring Efficient and Continual Semantic Segmentation with Plain Vision TransformersabstractAbstract This paper investigates the capability of plain Vision Transformers (ViTs) for semantic segmentation using the encoder–decoder framework and introduce SegViTv2 . In this study, we introduce a novel Attention-to-Mask (ATM) module to design a lightweight decoder effective for plain ViT. The proposed ATM converts the global attention map into semantic masks for high-quality segmentation results. Our decoder outperforms popular decoder UPerNet using various ViT backbones while consuming only about $$5\%$$ 5 % of the computational cost. For the encoder, we address the concern of the relatively high computational cost in the ViT-based encoders and propose a Shrunk ++ structure that incorporates edge-aware query-based down-sampling (EQD) and query-based up-sampling (QU) modules. The Shrunk++ structure reduces the computational cost of the encoder by up to $$50\%$$ 50 % while maintaining competitive performance. Furthermore, we propose to adapt SegViT for continual semantic segmentation, demonstrating nearly zero forgetting of previously learned knowledge. Experiments show that our proposed SegViTv2 surpasses recent segmentation methods on three popular benchmarks including ADE20k, COCO-Stuff-10k and PASCAL-Context datasets. The code is available through the following link: https://github.com/zbwxp/SegVit . Bowen Zhang 0009, Liyang Liu, Minh Hieu Phan, Zhi Tian, Chunhua Shen, Yifan Liu 0001 |
Int. J. Comput. Vis. | 6 |
| 2023 | ZegCLIP: Towards Adapting CLIP for Zero-shot Semantic SegmentationabstractRecently, CLIP has been applied to pixel-level zero-shot learning tasks via a two-stage scheme. The general idea is to first generate class-agnostic region proposals and then feed the cropped proposal regions to CLIP to utilize its image-level zero-shot classification capability. While effective, such a scheme requires two image encoders, one for proposal generation and one for CLIP, leading to a complicated pipeline and high computational cost. In this work, we pursue a simpler-and-efficient one-stage solution that directly extends CLIP's zero-shot prediction capability from image to pixel level. Our investigation starts with a straightforward extension as our baseline that generates semantic masks by comparing the similarity between text and patch embeddings extracted from CLIP. However, such a paradigm could heavily overfit the seen classes and fail to generalize to unseen classes. To handle this issue, we propose three simple-but-effective designs and figure out that they can significantly retain the inherent zero-shot capacity of CLIP and improve pixel-level generalization ability. Incorporating those modifications leads to an efficient zero-shot semantic segmentation system called ZegCLIP. Through extensive experiments on three public benchmarks, ZegCLIP demonstrates superior performance, outperforming the state-of-the-art methods by a large margin under both “inductive” and “transductive” zero-shot settings. In addition, compared with the two-stage method, our one-stage ZegCLIP achieves a speedup of about 5 times faster during inference. We release the code at https://github.com/ZiqinZhou66/ZegCLIP.git. Ziqin Zhou, Yinjie Lei, Bowen Zhang 0009, Lingqiao Liu, Yifan Liu 0001 |
CVPR | 5 |
| 2023 | Dynamic Token Pruning in Plain Vision Transformers for Semantic SegmentationabstractVision transformers have achieved leading performance on various visual tasks yet still suffer from high computational complexity. The situation deteriorates in dense prediction tasks like semantic segmentation, as high-resolution inputs and outputs usually imply more tokens involved in computations. Directly removing the less attentive tokens has been discussed for the image classification task but can not be extended to semantic segmentation since a dense prediction is required for every patch. To this end, this work introduces a Dynamic Token Pruning (DToP) method based on the early exit of tokens for semantic segmentation. Motivated by the coarse-to-fine segmentation process by humans, we naturally split the widely adopted auxiliary-loss-based network architecture into several stages, where each auxiliary block grades every token’s difficulty level. We can finalize the prediction of easy tokens in advance without completing the entire forward pass. Moreover, we keep k highest confidence tokens for each semantic category to uphold the representative context information. Thus, computational complexity will change with the difficulty of the input, akin to the way humans do segmentation. Experiments suggest that the proposed DToP architecture reduces on average 20% ∼ 35% of computational cost for current semantic segmentation methods based on plain vision transformers without accuracy degradation. The code is available through the following link: https://github.com/zbwxp/Dynamic-Token-Pruning. Quan Tang 0001, Bowen Zhang 0009, Jiajun Liu 0004, Fagui Liu, Yifan Liu 0001 |
ICCV | 5 |
| 2023 | Video Task Decathlon: Unifying Image and Video Tasks in Autonomous DrivingabstractPerforming multiple heterogeneous visual tasks in dynamic scenes is a hallmark of human perception capability. Despite remarkable progress in image and video recognition via representation learning, current research still focuses on designing specialized networks for singular, homogeneous, or simple combination of tasks. We instead explore the construction of a unified model for major image and video recognition tasks in autonomous driving with diverse input and output structures. To enable such an investigation, we design a new challenge, Video Task Decathlon (VTD), which includes ten representative image and video tasks spanning classification, segmentation, localization, and association of objects and pixels. On VTD, we develop our unified network, VTDNet, that uses a single structure and a single set of weights for all ten tasks. VTDNet groups similar tasks and employs task interaction stages to exchange information within and between task groups. Given the impracticality of labeling all tasks on all frames and the performance degradation associated with joint training of many tasks, we design a Curriculum training, Pseudo-labeling, and Fine-tuning (CPF) scheme to successfully train VTDNet on all tasks and mitigate performance loss. Armed with CPF, VTDNet significantly outperforms its single-task counterparts on most tasks with only 20% overall computations. VTD is a promising new direction for exploring the unification of perception tasks in autonomous driving. Thomas E. Huang, Yifan Liu 0001, Luc Van Gool, Fisher Yu 0001 |
ICCV | 2 |
| 2023 | 3DPPE: 3D Point Positional Encoding for Transformer-based Multi-Camera 3D Object DetectionabstractTransformer-based methods have swept the benchmarks on 2D and 3D detection on images. Because tokenization before the attention mechanism drops the spatial information, positional encoding becomes critical for those methods. Recent works found that encodings based on samples of the 3D viewing rays can significantly improve the quality of multi-camera 3D object detection. We hypothesize that 3D point locations can provide more information than rays. Therefore, we introduce 3D point positional encoding, 3DPPE, to the 3D detection Transformer decoder. Although 3D measurements are not available at the inference time of monocular 3D object detection, 3DPPE uses predicted depth to approximate the real point positions. Our hybrid-depth module combines direct and categorical depth to estimate the refined depth of each pixel. Despite the approximation, 3DPPE achieves 46.0 mAP and 51.4 NDS on the competitive nuScenes dataset, significantly outperforming encodings based on ray samples. The code is available at https://github.com/drilistbox/3DPPE. Changyong Shu, Jiajun Deng, Fisher Yu 0001, Yifan Liu 0001 |
ICCV | 4 |
| 2023 | CTVIS: Consistent Training for Online Video Instance SegmentationabstractThe discrimination of instance embeddings plays a vital role in associating instances across time for online video instance segmentation (VIS). Instance embedding learning is directly supervised by the contrastive loss computed upon the contrastive items (CIs), which are sets of anchor/positive/negative embeddings. Recent online VIS methods leverage CIs sourced from one reference frame only, which we argue is insufficient for learning highly discriminative embeddings. Intuitively, a possible strategy to enhance CIs is replicating the inference phase during training. To this end, we propose a simple yet effective training strategy, called Consistent Training for Online VIS (CTVIS), which devotes to aligning the training and inference pipelines in terms of building CIs. Specifically, CTVIS constructs CIs by referring inference the momentum-averaged embedding and the memory bank storage mechanisms, and adding noise to the relevant embeddings. Such an extension allows a reliable comparison between embeddings of current instances and the stable representations of historical instances, thereby conferring an advantage in modeling VIS challenges such as occlusion, re-identification, and deformation. Empirically, CTVIS outstrips the SOTA VIS models by up to +5.0 points on three VIS benchmarks, including YTVIS19 (55.1% AP), YTVIS21 (50.1% AP) and OVIS (35.5% AP). Furthermore, we find that pseudo-videos transformed from images can train robust models surpassing fully-supervised ones. Kaining Ying, Weian Mao, Zhenhua Wang 0003, Hao Chen 0041, Lin Wu 0001, Yifan Liu 0001, Chengxiang Fan, Yunzhi Zhuge, Chunhua Shen |
ICCV | 7 |
| 2023 | SegPrompt: Boosting Open-world Segmentation via Category-level Prompt LearningabstractCurrent closed-set instance segmentation models rely on pre-defined class labels for each mask during training and evaluation, largely limiting their ability to detect novel objects. Open-world instance segmentation (OWIS) models address this challenge by detecting unknown objects in a class-agnostic manner. However, previous OWIS approaches completely erase category information during training to keep the model’s ability to generalize to unknown objects. In this work, we propose a novel training mechanism termed SegPrompt that uses category information to improve the model’s class-agnostic segmentation ability for both known and unknown categories. In addition, the previous OWIS training setting exposes the unknown classes to the training set and brings information leakage, which is unreasonable in the real world. Therefore, we provide a new open-world benchmark closer to a real-world scenario by dividing the dataset classes into known-seen-unseen parts. For the first time, we focus on the model’s ability to discover objects that never appear in the training set images.Experiments show that SegPrompt can improve the overall and unseen detection performance by 5.6% and 6.1% in AR on our new benchmark without affecting the inference efficiency. We further demonstrate the effectiveness of our method on existing cross-dataset transfer and strongly supervised settings, leading to 5.5% and 12.3% relative improvement. Code and data are released at: https://github.com/aim-uofa/SegPrompt Muzhi Zhu, Hengtao Li, Hao Chen 0041, Chengxiang Fan, Weian Mao, Chenchen Jing, Yifan Liu 0001, Chunhua Shen |
ICCV | 7 |
| 2023 | Semi-supervised Semantic Segmentation with Mutual Knowledge DistillationabstractConsistency regularization has been widely studied in recent semi- supervised semantic segmentation methods, and promising per- formance has been achieved. In this work, we propose a new con- sistency regularization framework, termed mutual knowledge dis- tillation (MKD), combined with data and feature augmentation. We introduce two auxiliary mean-teacher models based on consis- tency regularization. More specifically, we use the pseudo-labels generated by a mean teacher to supervise the student network to achieve a mutual knowledge distillation between the two branches. In addition to using image-level strong and weak augmentation, we also discuss feature augmentation. This involves considering various sources of knowledge to distill the student network. Thus, we can significantly increase the diversity of the training samples. Experiments on public benchmarks show that our framework out- performs previous state-of-the-art (SOTA) methods under various semi-supervised settings. Code is available at https://github.com/jianlong-yuan/semi-mmseg. Jianlong Yuan, Jinchao Ge, Zhibin Wang 0004, Yifan Liu 0001 |
ACM Multimedia | 4 |
| 2023 | Object Detection Difficulty: Suppressing Over-aggregation for Faster and Better Video Object DetectionabstractCurrent video object detection (VOD) models often encounter issues with over-aggregation due to redundant aggregation strategies, which perform feature aggregation on every frame. This results in suboptimal performance and increased computational complexity. In this work, we propose an image-level Object Detection Difficulty (ODD) metric to quantify the difficulty of detecting objects in a given image. The derived ODD scores can be used in the VOD process to mitigate over-aggregation. Specifically, we train an ODD predictor as an auxiliary head of a still-image object detector to compute the ODD score for each image based on the discrepancies between detection results and ground-truth bounding boxes. The ODD score enhances the VOD system in two ways: 1) it enables the VOD system to select superior global reference frames, thereby improving overall accuracy; and 2) it serves as an indicator in the newly designed ODD Scheduler to eliminate the aggregation of frames that are easy to detect, thus accelerating the VOD process. Comprehensive experiments demonstrate that, when utilized for selecting global reference frames, ODD-VOD consistently enhances the accuracy of Global-frame-based VOD models. When employed for acceleration, ODD-VOD consistently improves the frames per second (FPS) by an average of 73.3% across 8 different VOD models without sacrificing accuracy. When combined, ODD-VOD attains state-of-the-art performance when competing with many VOD methods in both accuracy and speed. Our work represents a significant advancement towards making VOD more practical for real-world applications. The code will be released at https://github.com/bingqingzhang/odd-vod. Bingqing Zhang, Sen Wang 0001, Yifan Liu 0001, Branislav Kusy, Xue Li 0001, Jiajun Liu 0004 |
ACM Multimedia | 3 |
| 2023 | Segment Anything in High QualityabstractThe recent Segment Anything Model (SAM) represents a big leap in scaling up segmentation models, allowing for powerful zero-shot capabilities and flexible prompting. Despite being trained with 1.1 billion masks, SAM's mask prediction quality falls short in many cases, particularly when dealing with objects that have intricate structures. We propose HQ-SAM, equipping SAM with the ability to accurately segment any object, while maintaining SAM's original promptable design, efficiency, and zero-shot generalizability. Our careful design reuses and preserves the pre-trained model weights of SAM, while only introducing minimal additional parameters and computation. We design a learnable High-Quality Output Token, which is injected into SAM's mask decoder and is responsible for predicting the high-quality mask. Instead of only applying it on mask-decoder features, we first fuse them with early and final ViT features for improved mask details. To train our introduced learnable parameters, we compose a dataset of 44K fine-grained masks from several sources. HQ-SAM is only trained on the introduced detaset of 44k masks, which takes only 4 hours on 8 GPUs. We show the efficacy of HQ-SAM in a suite of 10 diverse segmentation datasets across different downstream tasks, where 8 out of them are evaluated in a zero-shot transfer protocol. Our code and pretrained models are at https://github.com/SysCV/SAM-HQ. Lei Ke, Mingqiao Ye, Martin Danelljan, Yifan Liu 0001, Yu-Wing Tai, Chi-Keung Tang, Fisher Yu 0001 |
NeurIPS | 4 |
| 2023 | QuantSR: Accurate Low-bit Quantization for Efficient Image Super-ResolutionabstractLow-bit quantization in image super-resolution (SR) has attracted copious attention in recent research due to its ability to reduce parameters and operations significantly. However, many quantized SR models suffer from accuracy degradation compared to their full-precision counterparts, especially at ultra-low bit widths (2-4 bits), limiting their practical applications. To address this issue, we propose a novel quantized image SR network, called QuantSR, which achieves accurate and efficient SR processing under low-bit quantization. To overcome the representation homogeneity caused by quantization in the network, we introduce the Redistribution-driven Learnable Quantizer (RLQ). This is accomplished through an inference-agnostic efficient redistribution design, which adds additional information in both forward and backward passes to improve the representation ability of quantized networks. Furthermore, to achieve flexible inference and break the upper limit of accuracy, we propose the Depth-dynamic Quantized Architecture (DQA). Our DQA allows for the trade-off between efficiency and accuracy during inference through weight sharing. Our comprehensive experiments show that QuantSR outperforms existing state-of-the-art quantized SR networks in terms of accuracy while also providing more competitive computational efficiency. In addition, we demonstrate the scheme's satisfactory architecture generality by providing QuantSR-C and QuantSR-T for both convolution and Transformer versions, respectively. Our code and models are released at https://github.com/htqin/QuantSR . Haotong Qin, Yulun Zhang 0001, Yifu Ding 0001, Yifan Liu 0001, Xianglong Liu 0001, Martin Danelljan, Fisher Yu 0001 |
NeurIPS | 4 |
| 2023 | A Dynamic Feature Interaction Framework for Multi-task Visual Perception
Yuling Xi, Hao Chen 0041, Ning Wang 0020, Peng Wang 0015, Yanning Zhang 0001, Chunhua Shen, Yifan Liu 0001 |
Int. J. Comput. Vis. | 7 |
| 2023 | Structured Knowledge Distillation for Dense PredictionabstractIn this work, we consider transferring the structure information from large networks to compact ones for dense prediction tasks in computer vision. Previous knowledge distillation strategies used for dense prediction tasks often directly borrow the distillation scheme for image classification and perform knowledge distillation for each pixel separately, leading to sub-optimal performance. Here we propose to distill structured knowledge from large networks to compact networks, taking into account the fact that dense prediction is a structured prediction problem. Specifically, we study two structured distillation schemes: i) pair-wise distillation that distills the pair-wise similarities by building a static graph; and ii) holistic distillation that uses adversarial training to distill holistic knowledge. The effectiveness of our knowledge distillation approaches is demonstrated by experiments on three dense prediction tasks: semantic segmentation, depth estimation and object detection. Code is available at https://git.io/StructKD. Yifan Liu 0001, Changyong Shu, Jingdong Wang 0001, Chunhua Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Towards Accurate Reconstruction of 3D Scene Shape From A Single Monocular ImageabstractDespite significant progress made in the past few years, challenges remain for depth estimation using a single monocular image. First, it is nontrivial to train a metric-depth prediction model that can generalize well to diverse scenes mainly due to limited training data. Thus, researchers have built large-scale relative depth datasets that are much easier to collect. However, existing relative depth estimation models often fail to recover accurate 3D scene shapes due to the unknown depth shift caused by training with the relative depth data. We tackle this problem here and attempt to estimate accurate scene shapes by training on large-scale relative depth data, and estimating the depth shift. To do so, we propose a two-stage framework that first predicts depth up to an unknown scale and shift from a single monocular image, and then exploits 3D point cloud data to predict the depth shift and the camera's focal length that allow us to recover 3D scene shapes. As the two modules are trained separately, we do not need strictly paired training data. In addition, we propose an image-level normalized regression loss and a normal-based geometry loss to improve training with relative depth annotation. We test our depth model on nine unseen datasets and achieve state-of-the-art performance on zero-shot evaluation. Code is available at: https://github.com/aim-uofa/depth/. Wei Yin 0006, Jianming Zhang 0001, Oliver Wang, Simon Niklaus, Simon Chen, Yifan Liu 0001, Chunhua Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | A Real-Time Memory Updating Strategy for Unsupervised Person Re-IdentificationabstractRecently, clustering-based methods have been the dominant solution for unsupervised person re-identification (ReID). Memory-based contrastive learning is widely used for its effectiveness in unsupervised representation learning. However, we find that the inaccurate cluster proxies and the momentum updating strategy do harm to the contrastive learning system. In this paper, we propose a real-time memory updating strategy (RTMem) to update the cluster centroid with a randomly sampled instance feature in the current mini-batch without momentum. Compared to the method that calculates the mean feature vectors as the cluster centroid and updating it with momentum, RTMem enables the features to be up-to-date for each cluster. Based on RTMem, we propose two contrastive losses, i.e., sample-to-instance and sample-to-cluster, to align the relationships between samples to each cluster and to all outliers not belonging to any other clusters. On the one hand, sample-to-instance loss explores the sample relationships of the whole dataset to enhance the capability of density-based clustering algorithm, which relies on similarity measurement for the instance-level images. On the other hand, with pseudo-labels generated by the density-based clustering algorithm, sample-to-cluster loss enforces the sample to be close to its cluster proxy while being far from other proxies. With the simple RTMem contrastive learning strategy, the performance of the corresponding baseline is improved by 9.3% on Market-1501 dataset. Our method consistently outperforms state-of-the-art unsupervised learning person ReID methods on three benchmark datasets. Code is made available at:https://github.com/PRIS-CV/RTMem. Junhui Yin, Xinyu Zhang 0015, Zhanyu Ma, Jun Guo 0002, Yifan Liu 0001 |
IEEE Trans. Image Process. | 5 |
| 2022 | Semantic-Guided Multi-mask Image Harmonization
Xuqian Ren, Yifan Liu 0001 |
ECCV (37) | 2 |
| 2022 | Controllable Shadow Generation Using Pixel Height Maps
Yichen Sheng, Yifan Liu 0001, Jianming Zhang 0001, Wei Yin 0006, A. Cengiz Öztireli, He Zhang 0004, Zhe Lin 0001, Eli Shechtman, Bedrich Benes |
ECCV (23) | 2 |
| 2022 | SegViT: Semantic Segmentation with Plain Vision TransformersabstractWe explore the capability of plain Vision Transformers (ViTs) for semantic segmentation and propose the SegViT. Previous ViT-based segmentation networks usually learn a pixel-level representation from the output of the ViT. Differently, we make use of the fundamental component—attention mechanism, to generate masks for semantic segmentation. Specifically, we propose the Attention-to-Mask (ATM) module, in which the similarity maps between a set of learnable class tokens and the spatial feature maps are transferred to the segmentation masks. Experiments show that our proposed SegViT using the ATM module outperforms its counterparts using the plain ViT backbone on the ADE20K dataset and achieves new state-of-the-art performance on COCO-Stuff-10K and PASCAL-Context datasets. Furthermore, to reduce the computational cost of the ViT backbone, we propose query-based down-sampling (QD) and query-based up-sampling (QU) to build a Shrunk structure. With our Shrunk structure, the model can save up to 40% computations while maintaining competitive performance. Bowen Zhang 0009, Zhi Tian, Quan Tang 0001, Xiangxiang Chu, Xiaolin Wei, Chunhua Shen, Yifan Liu 0001 |
NeurIPS | 7 |
| 2022 | Virtual Normal: Enforcing Geometric Constraints for Accurate and Robust Depth PredictionabstractMonocular depth prediction plays a crucial role in understanding 3D scene geometry. Although recent methods have achieved impressive progress in the evaluation metrics such as the pixel-wise relative error, most methods neglect the geometric constraints in the 3D space. In this work, we show the importance of the high-order 3D geometric constraints for depth prediction. By designing a loss term that enforces a simple geometric constraint, namely, virtual normal directions determined by randomly sampled three points in the reconstructed 3D space, we significantly improve the accuracy and robustness of monocular depth estimation. Importantly, the virtual normal loss can not only improve the performance of learning metric depth, but also disentangle the scale information and enrich the model with better shape information. Therefore, when not having access to absolute metric depth training data, we can use virtual normal to learn a robust affine-invariant depth generated on diverse scenes. Our experiments demonstrate state-of-the-art results of learning metric depth on NYU Depth-V2 and KITTI. From the high-quality predicted depth, we are now able to recover good 3D structures of the scene such as the point cloud and surface normal directly, eliminating the necessity of relying on additional models as was previously done. To demonstrate the excellent generalization capability of learning affine-invariant depth on diverse data with the virtual normal loss, we construct a large-scale and diverse dataset for training affine-invariant depth, termed Diverse Scene Depth dataset (DiverseDepth), and test on five datasets with the zero-shot test setting. Code is available at: https://git.io/Depth. Wei Yin 0006, Yifan Liu 0001, Chunhua Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Generic Perceptual Loss for Modeling Structured Output DependenciesabstractThe perceptual loss has been widely used as an effective loss term in image synthesis tasks including image super-resolution [16], and style transfer [14]. It was believed that the success lies in the high-level perceptual feature representations extracted from CNNs pretrained with a large set of images. Here we reveal that, what matters is the network structure instead of the trained weights. Without any learning, the structure of a deep network is sufficient to capture the dependencies between multiple levels of variable statistics using multiple layers of CNNs. This insight removes the requirements of pre-training and a particular network structure (commonly, VGG) that are previously assumed for the perceptual loss, thus enabling a significantly wider range of applications. To this end, we demonstrate that a randomly-weighted deep CNN can be used to model the structured dependencies of outputs. On a few dense per-pixel prediction tasks such as semantic segmentation, depth estimation and instance segmentation, we show improved results of using the extended randomized perceptual loss, compared to the baselines using pixel-wise loss alone. We hope that this simple, extended perceptual loss may serve as a generic structured-output loss that is applicable to most structured output learning tasks. Yifan Liu 0001, Hao Chen 0041, Yu Chen 0037, Wei Yin 0006, Chunhua Shen |
CVPR | 1 |
| 2021 | Channel-wise Knowledge Distillation for Dense Prediction*abstractKnowledge distillation (KD) has been proven a simple and effective tool for training compact dense prediction models. Lightweight student networks are trained by extra supervision transferred from large teacher networks. Most previous KD variants for dense prediction tasks align the activation maps from the student and teacher network in the spatial domain, typically by normalizing the activation values on each spatial location and minimizing point-wise and/or pair-wise discrepancy. Different from the previous methods, here we propose to normalize the activation map of each channel to obtain a soft probability map. By simply minimizing the Kullback–Leibler (KL) divergence between the channel-wise probability map of the two networks, the distillation process pays more attention to the most salient regions of each channel, which are valuable for dense prediction tasks.We conduct experiments on a few dense prediction tasks, including semantic segmentation and object detection. Experiments demonstrate that our proposed method outperforms state-of-the-art distillation methods considerably, and can require less computational cost during training. In particular, we improve the RetinaNet detector (ResNet50 backbone) by 3.4% in mAP on the COCO dataset, and PSPNet (ResNet18 backbone) by 5.81% in mIoU on the Cityscapes dataset. Code is available at: https://git.io/Distiller Changyong Shu, Yifan Liu 0001, Jianfei Gao 0002, Chunhua Shen |
ICCV | 2 |
| 2021 | A Simple Baseline for Semi-supervised Semantic Segmentation with Strong Data Augmentation*abstractRecently, significant progress has been made on semantic segmentation. However, the success of supervised semantic segmentation typically relies on a large amount of labeled data, which is time-consuming and costly to obtain. Inspired by the success of semi-supervised learning methods for image classification, here we propose a simple yet effective semi-supervised learning framework for semantic segmentation. We demonstrate that the devil is in the details: a set of simple designs and training techniques can collectively improve the performance of semi-supervised semantic segmentation significantly. Previous works [3], [25] fail to effectively employ strong augmentation in pseudo-label learning, as the large distribution disparity caused by strong augmentation harms the batch nor-malization statistics. We design a new batch normalization, namely distribution-specific batch normalization (DSBN) to address this problem and show the importance of strong augmentation for semantic segmentation. Moreover, we design a self-correction loss, which is effective in terms of noise resistance. We conduct a series of ablation studies to show the effectiveness of each component. Our method achieves state-of-the-art results in the semi-supervised settings on the Cityscapes and Pascal VOC datasets. Jianlong Yuan, Yifan Liu 0001, Chunhua Shen, Zhibin Wang 0004, Hao Li 0030 |
ICCV | 2 |
| 2021 | A Generative Adversarial Framework For Optimizing Image Matting And Harmonization SimultaneouslyabstractImage matting and image harmonization are two important tasks in image composition. Image matting, aiming to achieve foreground boundary details, and image harmonization, aiming to make the background compatible with the foreground, are both promising yet challenging tasks. Previous works consider optimizing these two tasks separately, which may lead to a sub-optimal solution. We propose to optimize matting and harmonization simultaneously to get better performance on both the two tasks and achieve more natural results. We propose a new Generative Adversarial (GAN) framework which optimizing the matting network and the harmonization network based on a self-attention discriminator. The discriminator is required to distinguish the natural images from different types of fake synthesis images. Extensive experiments on our constructed dataset demonstrate the effectiveness of our proposed method. Our dataset and dataset generating pipeline can be found in https://git.io/HaMaGAN. Xuqian Ren, Yifan Liu 0001, Chunlei Song |
ICIP | 2 |
| 2021 | Learning Structure Affinity for Video Depth EstimationabstractDepth estimation is a structure learning problem. The affinity among neighbouring pixels plays an important role in inferring depth values. In this paper, we propose to learn structure affinity in both spatial and temporal domain for accurate depth estimation from monocular videos. Specifically, we first propose a convolutional spatial temporal propagation network (CSTPN) that learns affinity among neighbouring video frames. Secondly, we employ a structure knowledge distillation scheme that transfers the spatial temporal affinity learned by cumbersome network to compact network. By calculating pixel-wise similarities between neighboring frames and neighbouring sequences, our knowledge distillation scheme efficiently captures both short-term and long-term spatial temporal affinity. Finally, we apply a warping loss based on optical flow between video frames to further enforce the temporal affinity. Experiment results show that our proposed depth estimation approach outperform the state-of-the-art methods on both indoor and outdoor benchmark datasets. Yuanzhouhan Cao, Yidong Li, Haokui Zhang, Chao Ren 0002, Yifan Liu 0001 |
ACM Multimedia | 5 |
| 2021 | Dynamic Neural Representational Decoders for High-Resolution Semantic SegmentationabstractSemantic segmentation requires per-pixel prediction for a given image. Typically, the output resolution of a segmentation network is severely reduced due to the downsampling operations in the CNN backbone. Most previous methods employ upsampling decoders to recover the spatial resolution.Various decoders were designed in the literature. Here, we propose a novel decoder, termed dynamic neural representational decoder (NRD), which is simple yet significantly more efficient. As each location on the encoder's output corresponds to a local patch of the semantic labels, in this work, we represent these local patches of labels with compact neural networks. This neural representation enables our decoder to leverage the smoothness prior in the semantic label space, and thus makes our decoder more efficient. Furthermore, these neural representations are dynamically generated and conditioned on the outputs of the encoder networks. The desired semantic labels can be efficiently decoded from the neural representations, resulting in high-resolution semantic segmentation predictions.We empirically show that our proposed decoder can outperform the decoder in DeeplabV3+ with only $\sim$$30\%$ computational complexity, and achieve competitive performance with the methods using dilated encoders with only $\sim$$15\% $ computation. Experiments on Cityscapes, ADE20K, and Pascal Context demonstrate the effectiveness and efficiency of our proposed method. Bowen Zhang 0009, Yifan Liu 0001, Zhi Tian, Chunhua Shen |
NeurIPS | 2 |
| 2020 | Instance-Aware Embedding for Point Cloud Instance Segmentation
Tong He 0001, Yifan Liu 0001, Chunhua Shen, Changming Sun |
ECCV (30) | 2 |
| 2020 | Efficient Semantic Video Segmentation with Per-Frame Inference
Yifan Liu 0001, Chunhua Shen, Changqian Yu, Jingdong Wang 0001 |
ECCV (10) | 1 |
| 2020 | Representative Graph Neural Network
Changqian Yu, Yifan Liu 0001, Changxin Gao, Chunhua Shen, Nong Sang |
ECCV (7) | 2 |
| 2020 | MobileFAN: Transferring deep hidden representation for face alignment
Yang Zhao 0019, Yifan Liu 0001, Chunhua Shen, Yongsheng Gao 0001, Shengwu Xiong 0001 |
Pattern Recognit. | 2 |
| 2019 | Structured Knowledge Distillation for Semantic SegmentationabstractIn this paper, we investigate the issue of knowledge distillation for training compact semantic segmentation networks by making use of cumbersome networks. We start from the straightforward scheme, pixel-wise distillation, which applies the distillation scheme originally introduced for image classification and performs knowledge distillation for each pixel separately. We further propose to distill the structured knowledge from cumbersome networks into compact networks, which is motivated by the fact that semantic segmentation is a structured prediction problem. We study two such structured distillation schemes: (i) pair-wise distillation that distills the pairwise similarities, and (ii) holistic distillation that uses adversarial training to distill holistic knowledge. The effectiveness of our knowledge distillation approaches is demonstrated by extensive experiments on three scene parsing datasets: Cityscapes, Camvid and ADE20K. Yifan Liu 0001, Chris Liu, Zengchang Qin, Zhenbo Luo, Jingdong Wang 0001 |
CVPR | 1 |
| 2019 | Pixel Level Data Augmentation for Semantic Image Segmentation Using Generative Adversarial NetworksabstractSemantic segmentation is one of the basic topics in computer vision, it aims to assign semantic labels to every pixel of an image. Unbalanced semantic label distribution could have a negative influence on segmentation accuracy. In this paper, we investigate using data augmentation approach to balance the semantic label distribution in order to improve segmentation performance. We propose using generative adversarial networks (GANs) to generate realistic images for improving the performance of semantic segmentation networks. Experimental results show that the proposed method can not only improve segmentation performance on those classes with low accuracy, but also obtain 1.3% to 2.1% increase in average segmentation accuracy. It shows that this augmentation method can boost the accuracy and be easily applicable to any other segmentation models. Shuangting Liu, Yifan Liu 0001, Zengchang Qin, Tao Wan 0001 |
ICASSP | 4 |
| 2019 | Enforcing Geometric Constraints of Virtual Normal for Depth PredictionabstractMonocular depth prediction plays a crucial role in understanding 3D scene geometry. Although recent methods have achieved impressive progress in evaluation metrics such as the pixel-wise relative error, most methods neglect the geometric constraints in the 3D space. In this work, we show the importance of the high-order 3D geometric constraints for depth prediction. By designing a loss term that enforces one simple type of geometric constraints, namely, virtual normal directions determined by randomly sampled three points in the reconstructed 3D space, we can considerably improve the depth prediction accuracy. Furthermore, we can not only predict accurate depth but also achieve high-quality other 3D information from the depth without retraining new parameters, Significantly, the byproduct of this predicted depth being sufficiently accurate is that we are now able to recover good 3D structures of the scene such as the point cloud and surface normal directly from the depth, eliminating the necessity of training new sub-models as was previously done. Experiments on two challenging benchmarks: NYU Depth-V2 and KITTI demonstrate the effectiveness of our method and state-of-the-art performance. Wei Yin 0006, Yifan Liu 0001, Chunhua Shen, Youliang Yan |
ICCV | 2 |
| 2018 | Emotion Classification with Data Augmentation Using Generative Adversarial Networks
Xinyue Zhu, Yifan Liu 0001, Tao Wan 0001, Zengchang Qin |
PAKDD (3) | 2 |
| 2018 | Auto-painter: Cartoon image generation from sketch by using conditional Wasserstein generative adversarial networks
Yifan Liu 0001, Zengchang Qin, Tao Wan 0001, Zhenbo Luo |
Neurocomputing | 1 |
| 2017 | Stock Volatility Prediction Using Recurrent Neural Networks with Sentiment Analysis
Yifan Liu 0001, Zengchang Qin, Tao Wan 0001 |
IEA/AIE (1) | 1 |