EDBT 2026 Demo / reviewers in the wild / expert
Chengxu Liu 0001
dblp:294/5677-1
· DBLP profile ↗
27ranked-venue papers
13as first author
27since 2021 · last 2026
0000-0001-8023-9465ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 10 first-author · 21 since 2021Artificial intelligence and machine learning · 12 · 8 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Density-aware adaptive label assignment for end-to-end dense detection in drone images
Chengxu Liu 0001, Nengzhong Yin, Guoshuai Zhao 0001, Xueming Qian |
Expert Syst. Appl. | 1 |
| 2026 | Towards Better Distortion Feature Learning for Object Detection in Top-View Fisheye CamerasabstractWith the development of deep learning in recent years, the performance of object detection under conventional cameras has been significantly improved. Nevertheless, due to the distortion caused by the fisheye cameras, detecting objects in this scenario remains a significant challenge. The dominant approaches focus on modifying the shape of the bounding box to better align the boundaries of the distorted object. However, these methods neglect the learning of spatial distortion information, which prevents them from satisfactory results. In this paper, we propose a novel fisheye camera detection network to learn distortion features better, dubbed SDANet. SDANet is composed of a series of SDABlocks, which are designed to learn spatial distortion features. Each SDABlock consists of multiple convolution kernels of different sizes, and it can generate the most suitable kernel based on the current input's distortion characteristics. Moreover, to address the limitations of the scarcity and uneven spatial distribution of fisheye image datasets on performance improvement, we propose a dedicated data augmentation strategy called Prominent Fisheye Distortion Augmentation (PFDAug). PFDAug can further introduce distortions to fisheye images, effectively alleviating these problems. Experimental results on the CEPDOF, MW-R, HABBOF, LOAF, and FishEye8k fisheye image datasets demonstrate that our method achieves state-ofthe-art performance. Pengbo Guo, Chengxu Liu 0001, Xingsong Hou, Xueming Qian |
IEEE Trans. Multim. | 2 |
| 2026 | Human Pose Estimation in Low-Light Condition With Decomposition and ModulationabstractHuman pose estimation (HPE) is a fundamental problem in computer vision, aiming to locate anatomical keypoints of the human body in a given picture. Benefiting from recent progress in deep learning, dominant HPE methods can achieve more advanced performance. Unfortunately, these methods rely heavily on large-scale, high-quality datasets captured expensively, resulting in limited learning capabilities in data-constrained low-light situations. Existing methods enhance the model's ability for low-light scenarios by performing intermediate feature alignment between low-light image and its well-lit counterpart. However, these methods fall short in fully exploiting explicit semantic feature exploitation that is independent of lighting conditions, resulting in sub-optimal performance. In this paper, we propose a Progressive Decomposition-Modulation network (PDMNet) for human pose estimation in extremely low-light condition. In particular, PDMNet mainly consists of 1) a semantic-specific decomposition module (SDM) for decomposing reflectance component with rich semantic information, and 2) a semantic-specific modulation mechanism (SMM) that enables the reflectance component to modulate the representation learning of human body parts in a tailored manner. Two closely-related components cooperate with each other to achieve more effective content-specific feature learning in low-light conditions. We further equip them progressively into different scales to enhance the feature learning. Experimental results demonstrate the superiority of PDMNet over state-of-the-art models on publicly available datasets. Our code will be released soon. Chengxu Liu 0001, Yujie Dun, Xueming Qian |
IEEE Trans. Multim. | 2 |
| 2025 | Frequency Domain-Based Diffusion Model for Unpaired Image DehazingabstractUnpaired image dehazing has attracted increasing attention due to its flexible data requirements during model training. Dominant methods based on contrastive learning not only introduce haze-unrelated content information, but also ignore haze-specific properties in the frequency domain (\ie,~haze-related degradation is mainly manifested in the amplitude spectrum). To address these issues, we propose a novel frequency domain-based diffusion model, named \ours, for fully exploiting the beneficial knowledge in unpaired clear data. In particular, inspired by the strong generative ability shown by Diffusion Models (DMs), we tackle the dehazing task from the perspective of frequency domain reconstruction and perform the DMs to yield the amplitude spectrum consistent with the distribution of clear images. To implement it, we propose an Amplitude Residual Encoder (ARE) to extract the amplitude residuals, which effectively compensates for the amplitude gap from the hazy to clear domains, as well as provide supervision for the DMs training. In addition, we propose a Phase Correction Module (PCM) to eliminate artifacts by further refining the phase spectrum during dehazing with a simple attention mechanism. Experimental results demonstrate that our \ours outperforms other state-of-the-art methods on both synthetic and real-world datasets. Chengxu Liu 0001, Lu Qi 0001, Jinshan Pan, Xueming Qian, Ming-Hsuan Yang 0001 |
ICCV | 1 |
| 2025 | Learning Deblurring Texture Prior From Unpaired Data with Diffusion Model
Chengxu Liu 0001, Lu Qi 0001, Jinshan Pan, Xueming Qian, Ming-Hsuan Yang 0001 |
ICCV | 1 |
| 2025 | Multi-Task Learning with Adaptive Fusion for Point-of-Interest RecommendationabstractPoint-of-Interest (POI) recommender systems [4, 29] have shown strong potential in personalized cultural exploration. Existing multi-branch frameworks combine graph, semantic, and sequential representations to capture diverse user behaviors, but often adopt static, intent-agnostic fusion strategies and fail to adapt to users’ varying preferences at different times of day. Moreover, the reasoning ability of Pre-trained Language Models (PLMs) [12] in these frameworks is typically confined to a single representation task, leaving their capacity for high-level intent inference largely untapped. We propose MAPRec, which improves this architecture through three key designs. First, an enhanced trajectory prompting mechanism enriches PLM inputs with fine-grained temporal patterns (e.g., day of week, time of day), providing richer contextual cues. Second, a multi-task PLM module jointly predicts the next POI and infers the user’s latent intent, encouraging the encoder to learn more generalizable representations. Finally, an adaptive fusion layer uses the inferred intent to dynamically adjust the contributions from the GNN, PLM, and sequential branches, enabling more tailored recommendations. Extensive experiments on two benchmark POI datasets demonstrate that MAPRec consistently outperforms strong baselines, validating the effectiveness of our design. Bosong Yang, Chengxu Liu 0001, Boyang Yan, Zimo Zhu |
MMAsia | 2 |
| 2025 | Rethinking Label Assignment and Sampling Strategies for Two-Stage Oriented Object DetectionabstractOriented object detection seeks to determine both the position and orientation of objects, yet angle periodicity often limits performance. To solve this issue, we rethink label assignment and sampling strategies and propose a pair of orientation-aware assigner and sampler (OAS) for a two-stage detector. The orientation-aware assigner (OA) incorporates angle and location to improve positive and negative sample assignment, while the orientation-aware sampler (OS) ranks positive samples by their angular difference from ground truth, adjusting learning weights by soft sampling. Such a design significantly mitigates the angle periodicity problem and enables detector focusing on high-quality samples with a more consistent orientation for training. Experimental results on two challenging oriented object detection benchmarks demonstrate that OAS can consistently boost the detection accuracy based on many existing two-stage detectors (e.g., Oriented R-CNN and RPGAOD) without additional cost. Both code and pretrained models are available athttps://github.com/skyandkibo/OAS. Jinjin Qian, Chengxu Liu 0001, Yubin Ai, Guoshuai Zhao 0001, Xueming Qian |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2025 | InstanceSR: Efficient Reconstructing Small Object With Differential Instance-Level Super-ResolutionabstractSuper-resolution (SR) aims to restore a high-resolution (HR) image from its low-resolution (LR) counterpart. Existing works try to achieve an overall average recovery over all regions to provide better visual quality for human viewing. If we desire to explore the potential that performs super-resolution for machine recognition instead of human viewing, the solution should change accordingly. From this insight, we propose a new SR pipeline, called InstanceSR, which treats each region in the LR image differentially and consumes more resources to focus on the recovery of the foreground region where the instances exist. In particular, InstanceSR consists of an encoder that formulates the LR image into a set of various difficulty tokens according to the instances distribution in each sub-region, and a decoder based on a multi-exit network structure to recover the sub-regions corresponding to various difficulty tokens by consuming different computational resources. Experimental results demonstrate the superiority of the proposed InstanceSR over state-of-the-art models, especially the recovery of regions where instances exist, by extensive quantitative and qualitative evaluations on three widely used benchmarks containing small instances. Besides, the comparisons using SR results on three challenging small object detection benchmarks verify that our InstanceSR can consistently boost the detection accuracy and has great potential for subsequent machine recognition. Yuanting Fan, Chengxu Liu 0001, Ruhao Tian, Xueming Qian |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Enhancing Automated Vending Machine Product Recognition Through Depth-Guided Regression RefinementabstractDeep neural network advancements have led to significant progress in industrial applications, particularly in product recognition for smart automated vending machines (AVMs). This area has seen increased market demand as a fundamental part of automatedretail. However, the existing works possess critical gaps: 1) densely placed and obscured objects in AVMs lead to inaccurate product recognition results, necessitating auxiliary information for achieving precise detection; 2) the lack of datasets with auxiliary information hinders further development in this field. To address these gaps, we propose a depth-guided product recognition network, which consists of two novel components: a depth-aware feature pyramid network (DFPN) and a depth-aware regression head (DRH). Our DFPN can adaptively select features that are beneficial for regression from both red, green, blue (RGB) and depth data, whereas the DRH refines the regression branch via depth information without affecting the classification process. In addition, to overcome dataset limitations, we develop an extended and fully annotated depth information dataset namedSmartUVM-D, which includes depth information for each image based on the existing SmartUVM dataset. The experimental results obtained on our SmartUVM-D benchmark show that our method effectively solves the inaccurate product recognition problem and achieves substantial gains over the baseline approaches. Specifically, our method (based on the ATSS framework) achieves a mean average precision of 84.4, representing a 2.3-point improvement over the previously developed ATSS method and establishing a new state-of-the-art approach. Jinjin Qian, Chengxu Liu 0001, Yubin Ai, Guoshuai Zhao 0001, Xueming Qian |
IEEE Trans. Ind. Informatics | 2 |
| 2025 | Single-Domain Generalized Object Detection With Frequency Whitening and Contrastive LearningabstractSingle-Domain Generalization Object Detection (Single-DGOD) refers to training a model with only one source domain, enabling the model to generalize to any unseen domain. For instance, a detector trained on a sunny daytime dataset should also perform well in scenarios such as rainy nighttime. The main challenge is to improve the detector's ability to learn the domain-invariant representation (DIR) while removing domain-specific information. Recent progress in Single-DGOD has demonstrated the efficacy of removing domain-specific information by adjusting feature distributions. Nonetheless, simply adjusting the global feature distribution in Single-DGOD task is insufficient to learn the potential relationship from sunny to adverse weather, as these ignore the significant domain gaps between instances across different weathers. In this paper, we propose a novel object detection method for more robust single-domain generalization. In particular, it mainly consists of a frequency-aware selective whitening module (FSW) for removing redundant domain-specific information and a contrastive feature alignment module (CFA) for enhancing domain-invariant information among instances. Specially, FSW extracts the magnitude spectrum of the feature and uses a group whitening loss to selectively eliminate redundant domain-specific information in the magnitude. To further eliminate domain differences among instances, we apply the style transfer method for data augmentation and use the augmented data in the CFA module. CFA formulates both the original and the augmentd RoI features into a series of groups with different categories, and utilizes contrastive learning across them to facilitate the learning of DIR in various categories. Experiments show that our method achieves favorable performance on existing standard benchmarks. Chengxu Liu 0001, Xueming Qian, Xubin Feng |
IEEE Trans. Multim. | 2 |
| 2025 | SLE: Out-of-Distribution Detection With Shallow Layer-Driven EnhancementabstractOut-of-distribution detection aims to protect models against overconfidently categorizing samples from unknown categories,i.e., out-of-distribution data (OOD), into known categories,i.e., in-distribution data (ID). From the perspective of feature distribution, the difference between OOD samples and ID samples can be decomposed into semantic shifts and covariate shifts. Most DL-based methods only extract deeper features, which represent semantic shifts, to discern feature variances in the data, ignoring the exploration of covariance shifts. In this paper, we propose a Shallow Layer-driven Enhanced OOD detection method (SLE), which enhances the difference of OOD samples by exploiting covariate shifts in shallow features. Specifically, it contains three main components: Hierarchical Feature Extractor (HFE), Adaptive Dimensionality Reduction Strategy (ADR), Cross-layer Score Aggregator (CSA). HFE is responsible for extracting both deeper and shallow features from the deep network. ADR adaptively reduces all hierarchical feature dimensionality according to sample characteristics, avoiding feature redundancy. CSA defines a novel confidence score for OOD samples, that effectively prevents confusion in the feature representation space at each layer. In SLE, these three closely related components cooperate with each other to effectively enhance the representation ability of OOD samples and divide OOD data better. We conduct extensive experiments to examine the performance of SLE in four benchmarks and discuss its individual components. This method performs well on the OOD datasets. Zhenni Yang, Chengxu Liu 0001, Xueming Qian |
IEEE Trans. Multim. | 2 |
| 2024 | Decoupling Degradations with Recurrent Network for Video Restoration in Under-Display CameraabstractUnder-display camera (UDC) systems are the foundation of full-screen display devices in which the lens mounts under the display. The pixel array of light-emitting diodes used for display diffracts and attenuates incident light, causing various degradations as the light intensity changes. Unlike general video restoration which recovers video by treating different degradation factors equally, video restoration for UDC systems is more challenging that concerns removing diverse degradation over time while preserving temporal consistency. In this paper, we introduce a novel video restoration network, called D2RNet, specifically designed for UDC systems. It employs a set of Decoupling Attention Modules (DAM) that effectively separate the various video degradation factors. More specifically, a soft mask generation function is proposed to formulate each frame into flare and haze based on the diffraction arising from incident light of different intensities, followed by the proposed flare and haze removal components that leverage long- and short-term feature learning to handle the respective degradations. Such a design offers an targeted and effective solution to eliminating various types of degradation in UDC systems. We further extend our design into multi-scale to overcome the scale-changing of degradation that often occur in long-range videos. To demonstrate the superiority of D2RNet, we propose a large-scale UDC video benchmark by gathering HDR videos and generating realistically degraded videos using the point spread function measured by a commercial UDC system. Extensive quantitative and qualitative evaluations demonstrate the superiority of D2RNet compared to other state-of-the-art video restoration and UDC image restoration methods. Chengxu Liu 0001, Xuan Wang 0018, Yuanting Fan, Xueming Qian |
AAAI | 1 |
| 2024 | Motion-Adaptive Separable Collaborative Filters for Blind Motion DeblurringabstractEliminating image blur produced by various kinds ofmotion has been a challenging problem. Dominant approaches rely heavily on model capacity to remove blurring by reconstructing residual from blurry observation in feature space. These practices not only prevent the capture of spatially variable motion in the real world but also ignore the tai-lored handling of various motions in image space. In this paper, we propose a novel real-world deblurring filtering model called the Motion-adaptive Separable Collaborative (MISC) Filter. In particular, we use a motion estimation net-work to capture motion information from neighborhoods, thereby adaptively estimating spatially-variant motion flow, mask, kernels, weights, and offsets to obtain the MISC Fil-ter. The MISC Filter first aligns the motion-induced blur-ring patterns to the motion middle along the predicted flow direction, and then collaboratively filters the aligned image through the predicted kernels, weights, and offsets to generate the output. This design can handle more general-ized and complex motion in a spatially differentiated man-ner. Furthermore, we analyze the relationships between the motion estimation network and the residual reconstruction network. Extensive experiments on four widely used bench-marks demonstrate that our method provides an effective solution for real-world motion blur removal and achieves state-of-the-art performance. Code is available at https://github.com/ChengxuLiu/MISCFilter. Chengxu Liu 0001, Xuan Wang 0018, Xiangyu Xu 0002, Ruhao Tian, Xueming Qian, Ming-Hsuan Yang 0001 |
CVPR | 1 |
| 2024 | AdaDiffSR: Adaptive Region-Aware Dynamic Acceleration Diffusion Model for Real-World Image Super-Resolution
Yuanting Fan, Chengxu Liu 0001, Nengzhong Yin, Changlong Gao, Xueming Qian |
ECCV (12) | 2 |
| 2024 | QueryCDR: Query-Based Controllable Distortion Rectification Network for Fisheye Images
Pengbo Guo, Chengxu Liu 0001, Xingsong Hou, Xueming Qian |
ECCV (15) | 2 |
| 2024 | AJENet: Adaptive Joints Enhancement Network for Abnormal Behavior Detection in Office ScenarioabstractWith the increasing popularity of intelligent surveillance systems, abnormal behavior detection of human beings based on computer vision is attracting more attention. It aims to classify and locate the abnormal behaviors and coordinates of human beings, respectively, and is a fundamental technology for intelligent security. Existing approaches mainly focus on exploring abnormal behavior features through object detectors. However, in office scenarios, almost all abnormal behaviors are closely associated with the fine-grained feature around the nose, wrist, elbow, and other human joint points regions. Detectors for generic objects cannot adequately capture such differences between abnormal behaviors, resulting in sub-optimal performance. In this paper, we focus on human joints and take one step further to enable effective behavior characteristics learning in office scenarios. In particular, we propose a novel Adaptive Joints Enhancement Network (AJENet), which includes two closely-related components, Joints Predict block (JP) and Adaptive Key Joints Enhancement block (AKJE). JP block is used to predict the human joints and facilitates the feature learning around them implicitly. By inputting the features around joints, the AKJE block enhances the feature representations of key joints according to the abnormal behavior characteristics adaptively. Experimental results demonstrate that our method outperforms other state-of-the-art methods on the collected real office scenario Office Behavior Dataset. Besides, to verify the generalization capabilities and potential of AJENet, we construct comparisons on another generic dataset PASCAL VOC 2012 Action. Chengxu Liu 0001, Yaru Zhang, Xueming Qian |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | HF-HRNet: A Simple Hardware Friendly High-Resolution NetworkabstractHigh-resolution networks have made significant progress in dense prediction tasks such as human pose estimation and semantic segmentation. To better explore this high-resolution mechanism on mobile devices, Lite-HRNet incorporates shuffle operations to reduce computational complexity in the channel dimension, while Dite-HRNet employs dynamic convolution and pooling to capture long-range interactions with low computational complexity in the spatial dimension. The core idea behind both approaches is to efficiently capture information in either the channel or spatial dimension. However, shuffle operations and dynamic operations are not hardware-friendly. As a result, both Lite-HRNet and Dite-HRNet cannot achieve the desired inference speed on specialized devices, including Neural Processing Units (NPUs) and Graphics Processing Units (GPUs). To overcome these limitations, we present a simple Hardware-Friendly Lightweight High-resolution Network (HF-HRNet) based on our proposed Hardware-Friendly Uniform-sized Mug (HUM) block. HUM block mainly consists of the Cascaded Depthwise (CAD) block and Multi-Scale Context Embedding (MCE) block. The CAD block cascades depthwise convolutions to obtain a larger receptive field in the spatial dimension, while the MCE block aggregates multi-scale spatial feature information from different scales and adjusts channel features. Extensive experiments are conducted on human pose estimation (COCO, MPII) and semantic segmentation (Cityscapes), resulting in a better trade-off between inference speed and accuracy on both NPUs and GPUs. It is noteworthy that on the COCO test-dev set, HF-HRNet-30 outperforms Dite-HRNet-30 and Lite-HRNet-30 by 1.9 AP and 2.8 AP, respectively, while running about 13 times faster and 9 times faster on NPUs, respectively. Our code are publicly available for use: https://github.com/zhanghao5201/HF-HRNet. Hao Zhang 0117, Yujie Dun, Yixuan Pei, Shenqi Lai, Chengxu Liu 0001, Kaipeng Zhang, Xueming Qian |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Split-Check: Boosting Product Recognition via Instance-Level RetrievalabstractAI-based methods are shining across a variety of industries, especially unmanned retail. Product recognition is the problem of recognizing the category and quantity of products (e.g., beverages and mineral water) in intelligent unmanned vending machines (UVMs) to automatic checkout during purchase. However, for similar products in hundreds of categories, the existing method is not accurate enough. Besides, they cannot be extended for new products without retraining. In this article, we propose a product recognition approach based on intelligent UVMs, calledSplit-Check, which first splits the region of interest of products by detection and then check product by instance-level retrieval. Split-Check is the combination of two important components. The preliminary detection distinguishes items that contain the different coarse-grained features, then locates items, and classifies them into coarse-grained categories as a candidate. The retrieval further distinguishes the candidate items that contain the different fine-grained features. Besides, we reconstruct a large-scale categories product dataset GOODS-85 based on actual UVMs scenarios, in which the number of categories of items is larger than the existing dataset. Experimental results demonstrate the effectiveness of the proposed approach. Our method significantly improves the recognition performance of hundreds of products and increases the scalability of products. Chengxu Liu 0001, Zongyang Da, Yuanzhi Liang, Guoshuai Zhao 0001, Xueming Qian |
IEEE Trans. Ind. Informatics | 1 |
| 2024 | SDPDet: Learning Scale-Separated Dynamic Proposals for End-to-End Drone-View DetectionabstractDetecting objects in large-scale drone-view images is notoriously challenging due to their uneven distribution and scale variation caused by photoing angles. Common approaches promote drone-view object detection by two-step detection (i.e., detecting sub-regions first) and multi-scale input. However, all these methods suffer from onerous computational costs since the high model complexity and input resolution. In this paper, we propose a novel one-step detector, calledSDPDet, to enable effective object learning in drone-view images. In particular, a Scale-separated Activation Pyramid (SAP) serves to focus on the regions with objects aggregated at each scale, and a Scale-separated Learnable Proposals (SLP) mechanism learns proposal boxes and corresponding features on these regions. By such design, the quantity of learnable proposals allows dynamic adjustment at each scale separately, which facilitates the objects learning of various distributions and scales with less computational costs. Experiments demonstrate SDPDet can significantly outperform the state-of-the-art one-step detectors on three widely-used benchmarks. On the most challenging VisDrone dataset, SDPDet with ResNet50 gains5.4%AP and6.9%AP$_{s}$improvements while running1.9×faster than previous models. Nengzhong Yin, Chengxu Liu 0001, Ruhao Tian, Xueming Qian |
IEEE Trans. Multim. | 2 |
| 2024 | Product Recognition for Unmanned Vending MachinesabstractRecently, the emerging concept of "unmanned retail" has drawn more and more attention, and the unmanned retail based on the intelligent unmanned vending machines (UVMs) scene has great market demand. However, existing product recognition methods for intelligent UVMs cannot adapt to large-scale categories and have insufficient accuracy. In this article, we propose a method for large-scale categories product recognition based on intelligent UVMs. It can be divided into two parts: 1) first, we explore the similarities and differences between products through manifold learning, and then we build a hierarchical multigranularity label to constrain the learning of representation; and 2) second, we propose a hierarchical label object detection network, which mainly includes coarse-to-fine refine module (C2FRM) and multiple granularity hierarchical loss (MGHL), which are used to assist in capturing multigranularity features. The highlights of our method are mine potential similarity between large-scale category products and optimization through hierarchical multigranularity labels. Besides, we collected a large-scale product recognition dataset GOODS-85 based on the actual UVMs scenario. Experimental results and analysis demonstrate the effectiveness of the proposed product recognition methods. Chengxu Liu 0001, Zongyang Da, Yuanzhi Liang, Guoshuai Zhao 0001, Xueming Qian |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | CSDA: Learning Category-Scale Joint Feature for Domain Adaptive Object DetectionabstractDomain Adaptive Object Detection (DAOD) aims to improve the detection performance of target domains by minimizing the feature distribution between the source and target domain. Recent approaches usually align such distributions in terms of categories through adversarial learning and some progress has been made. However, when objects are non-uniformly distributed at different scales, such category-level alignment causes imbalanced object feature learning, refer as the inconsistency of category alignment at different scales. For better category-level feature alignment, we propose a novel DAOD framework of joint category and scale information, dubbed CSDA, such a design enables effective object learning for different scales. Specifically, our framework is implemented by two closely-related modules: 1) SGFF (Scale-Guided Feature Fusion) fuses the category representations of different domains to learn category-specific features, where the features are aligned by discriminators at three scales. 2) SAFE (Scale-Auxiliary Feature Enhancement) encodes scale coordinates into a group of tokens and enhances the representation of category-specific features at different scales by self-attention. Based on the anchor-based Faster-RCNN and anchor-free FCOS detectors, experiments show that our method achieves state-of-the-art results on three DAOD benchmarks. Changlong Gao, Chengxu Liu 0001, Yujie Dun, Xueming Qian |
ICCV | 2 |
| 2023 | FSI: Frequency and Spatial Interactive Learning for Image Restoration in Under-Display CamerasabstractUnder-display camera (UDC) systems remove the screen notch for bezel-free displays and provide a better interactive experience. The main challenge is that the pixel array of light-emitting diodes used for display diffracts and attenuates the incident light, leading to complex degradation. Existing models eliminate spatial diffraction by maximizing model capacity through complex design and ignore the periodic distribution of diffraction in the frequency domain, which prevents these approaches from satisfactory results. In this paper, we introduce a new perspective to handle various diffraction in UDC images by jointly exploring the feature restoration in the frequency and spatial domains, and present a Frequency and Spatial Interactive Learning Network (FSI). It consists of a series of well-designed Frequency-Spatial Joint (FSJ) modules for feature learning and a color transform module for color enhancement. In particular, in the FSJ module, a frequency learning block uses the Fourier transform to eliminate spectral bias, a spatial learning block uses a multi-distillation structure to supplement the absence of local details, and a dual transfer unit to facilitate the interactive learning between features of different domains. Experimental results demonstrate the superiority of the proposed FSI over state-of-the-art models, through extensive quantitative and qualitative evaluations in three widely-used UDC benchmarks. Chengxu Liu 0001, Xuan Wang 0018, Yuzhi Wang, Xueming Qian |
ICCV | 1 |
| 2023 | Anomaly detection framework for unmanned vending machines
Zongyang Da, Yujie Dun, Chengxu Liu 0001, Yuanzhi Liang, Xueming Qian |
Knowl. Based Syst. | 3 |
| 2023 | TTVFI: Learning Trajectory-Aware Transformer for Video Frame InterpolationabstractVideo frame interpolation (VFI) aims to synthesize an intermediate frame between two consecutive frames. State-of-the-art approaches usually adopt a two-step solution, which includes 1) generating locally-warped pixels by calculating the optical flow based on pre-defined motion patterns (e.g., uniform motion, symmetric motion), 2) blending the warped pixels to form a full frame through deep neural synthesis networks. However, for various complicated motions (e.g., non-uniform motion, turn around), such improper assumptions about pre-defined motion patterns introduce the inconsistent warping from the two consecutive frames. This leads to the warped features for new frames are usually not aligned, yielding distortion and blur, especially when large and complex motions occur. To solve this issue, in this paper we propose a novel Trajectory-aware Transformer for Video Frame Interpolation (TTVFI). In particular, we formulate the warped features with inconsistent motions as query tokens, and formulate relevant regions in a motion trajectory from two original consecutive frames into keys and values. Self-attention is learned on relevant tokens along the trajectory to blend the pristine features into intermediate frames through end-to-end training. Experimental results demonstrate that our method outperforms other state-of-the-art methods in four widely-used VFI benchmarks. Both code and pre-trained models will be released at https://github.com/ChengxuLiu/TTVFI. Chengxu Liu 0001, Huan Yang 0005, Jianlong Fu, Xueming Qian |
IEEE Trans. Image Process. | 1 |
| 2023 | 4D LUT: Learnable Context-Aware 4D Lookup Table for Image EnhancementabstractImage enhancement aims at improving the aesthetic visual quality of photos by retouching the color and tone, and is an essential technology for professional digital photography. Recent years deep learning-based image enhancement algorithms have achieved promising performance and attracted increasing popularity. However, typical efforts attempt to construct a uniform enhancer for all pixels' color transformation. It ignores the pixel differences between different content (e.g., sky, ocean, etc.) that are significant for photographs, causing unsatisfactory results. In this paper, we propose a novel learnable context-aware 4-dimensional lookup table (4D LUT), which achieves content-dependent enhancement of different contents in each image via adaptively learning of photo context. In particular, we first introduce a lightweight context encoder and a parameter encoder to learn a context map for the pixel-level category and a group of image-adaptive coefficients, respectively. Then, the context-aware 4D LUT is generated by integrating multiple basis 4D LUTs via the coefficients. Finally, the enhanced image can be obtained by feeding the source image and context map into fused context-aware 4D LUT via quadrilinear interpolation. Compared with traditional 3D LUT, i.e., RGB mapping to RGB, which is usually used in camera imaging pipeline systems or tools, 4D LUT, i.e., RGBC(RGB+Context) mapping to RGB, enables finer control of color transformations for pixels with different content in each image, even though they have the same RGB values. Experimental results demonstrate that our method outperforms other state-of-the-art methods in widely-used benchmarks. Chengxu Liu 0001, Huan Yang 0005, Jianlong Fu, Xueming Qian |
IEEE Trans. Image Process. | 1 |
| 2022 | Learning Trajectory-Aware Transformer for Video Super-ResolutionabstractVideo super-resolution (VSR) aims to restore a sequence of high-resolution (HR) frames from their low-resolution (LR) counterparts. Although some progress has been made, there are grand challenges to effectively utilize temporal dependency in entire video sequences. Existing approaches usually align and aggregate video frames from limited adjacent frames (e.g., 5 or 7 frames), which prevents these approaches from satisfactory results. In this paper, we take one step further to enable effective spatio-temporal learning in videos. We propose a novel Trajectory-aware Transformer for Video Super-Resolution (TTVSR). In particular, we formulate video frames into several pre-aligned trajectories which consist of continuous visual tokens. For a query token, self-attention is only learned on relevant visual tokens along spatio-temporal trajectories. Compared with vanilla vision Transformers, such a design significantly reduces the computational cost and enables Transformers to model long-range features. We further propose a cross-scale feature tokenization module to over-come scale-changing problems that often occur in long-range videos. Experimental results demonstrate the superiority of the proposed TTVSR over state-of-the-art models, by extensive quantitative and qualitative evaluations in four widely-used video super-resolution benchmarks. Both code and pre-trained models can be downloaded at https://github.com/researchmm/TTVSR. Chengxu Liu 0001, Huan Yang 0005, Jianlong Fu, Xueming Qian |
CVPR | 1 |
| 2021 | Food and Ingredient Joint Learning for Fine-Grained RecognitionabstractFine-grained food recognition is the detailed classification that provides more specialized and professional attribute information of food. It is the basic work to realize healthy diet recommendations and cooking instructions, nutrition intake management, and cafeteria self-checkout system. Chinese food lacks structured information, and ingredients composition is an important consideration. The current approaches mostly focus on global dish appearance without any analysis of ingredient composition and fully considering the attention of regional features. In this paper, we propose an Attention Fusion Network (AFN) and Food-Ingredient Joint Learning module for fine-grained food and ingredients recognition. The AFN first focuses on the food discrimination region against unstructured defeat and generates the feature embeddings jointly aware of the ingredients and food. The Food-Ingredient Joint Learning module aims at alleviating the issue of ingredients imbalance. Therefore, we propose a balance focal loss to optimize the feature expression ability of the network for ingredients. In experiments, the results of ingredients recognition show the state-of-the-art performances on fine-grained Chinese food dataset VIREO Food-172. Chengxu Liu 0001, Yuanzhi Liang, Xueming Qian, Jianlong Fu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |