Lizhe Qi

dblp:121/6207 · DBLP profile ↗
← Back
41ranked-venue papers
1as first author
30since 2021 · last 2025
0000-0002-4348-1559ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 1 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 12 since 2021Systems, architecture and hardware · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 RetinexUNet++: A Retinex-Based UNet++ Model for Low-Light Image Enhancement
abstract
Low-light image enhancement (LLIE) aims to recover details in images captured under low-light conditions, crucial for improving the quality of low-level perception and enhancing the performance of high-level visual tasks. Most transformer-based methods show great potential in this field, but they rarely utilize the inherent structural details and color information of images that are unaffected by lighting, leading to susceptibility to noise interference when learning features. To address this issue, we propose a novel low-light image enhancement U-shaped network based on Retinex, called RetinexUNet++. Firstly, we design a reflectance perception module (RPM), which uses image intrinsic structure information as key clues to guide the network to learn low-light image features. Secondly, we design a transformer-based reinforced U-shaped network (RUN) to narrow the semantic gap between the encoder and decoder feature maps, learning multi-scale semantic features crucial for image restoration. Finally, to address the lack of real high-definition datasets in this field, we build a dataset of 6K resolution image pairs captured in desert low-light (D-LOL6K) scenes. Extensive experiments conducted on public datasets and the new dataset demonstrate that RetinexUNet++ outperforms the current state-of-the-art methods in both quantitative and visual performance.
Lijun Dai, Chenhuan Liu, Yunquan Sun, Lizhe Qi
CSCWD5
2025 DI-Net: Dynamic Interconnected Network for Accurate Industrial Defect Detection
abstract
Surface defect detection is crucial for ensuring product quality and improving production efficiency in real-world industrial scenes. Although deep learning-based detectors show remarkable performance, their effectiveness is hindered by the subtle distinction between defective objects and the background, as well as the varying sizes, shapes, and dense distribution of defects in industrial environments. Focusing on the above problems, this paper proposes a notable and innovative one-stage detector named Dynamic Interconnected Network (DI-Net). Firstly, the DI-Net consists of a dedicated Cross-Field Learning (CFL) module that leverages both convolution and involution feature extraction capabilities, enabling it to aggregate key information while avoiding over-fitting caused by excessive stacking of convolution layers, thus resolving the challenge of distinguishing subtle differences between the object and background. Secondly, we design the Dynamic Feature Fusion (DFF) Block, which is integrated into the neck network of DI-Net. This block employs flexible and adaptive dynamic convolution, enhanced by the CFL attention mask to further enrich feature representations, which enhance the model's detection performance for multi-size and shaped defect objects. In this paper, we refer to the overall module, which includes CFL and DFF, as the Dynamic Interconnected Module (DIM). To verify DI-Net's superiority and generalization, the proposed approach is tested on two sets of benchmark datasets including PKU-PCB and NEU-DET for industrial quality inspection. The experimental results demonstrate that our DI-Net achieves 80.9% and 95.6% mAP on the NEU-DET and PKU-PCB test datasets, respectively. Overall, it is a significant improvement for the detection performance at complex industrial, which better satisfies the accuracy requirements for the practical industrial applications.
Xinzhi Lin, Yunquan Sun, Lizhe Qi
CSCWD6
2025 Integrating Failures in Robot Skill Acquisition with Offline Action-Sequence Diffusion RL
abstract
Recent advancements in robot learning leverage large language models (LLMs) and sampling-based task and motion planning (TAMP) modules for automatic and scalable robot data generation. This method yields both success trajectories and a large number of failure trajectories. Prior works typically filter out failure data and adopt behavior cloning (BC) to train policy. However, this significantly reduces the sample efficiency of the method and results in a policy limited by the data collection behavior policy. In this paper, we introduce a vision-language-conditioned action-sequence diffusion policy and an action-sequence diffusion policy learning with Q-guided refinement for its training. We first redefine the reverse process of the diffusion model as the distribution of action sequences conditioned on visual observations and language instructions. We then employ BC to pretrain the policy on the success sub-dataset. Next, we optimize the action-sequence Q-value function by minimizing the temporal difference error across the complete dataset. Finally, we integrate guidance from the Q-value function into the BC loss of the reverse diffusion chain. Our method significantly outperforms baseline methods in terms of success rate and sample efficiency. By effectively leveraging failure data to optimize the policy, our method can achieve results comparable to those trained with the complete success sub-dataset while requiring 20%-30% less success data.
Hecheng Wang, Lizhe Qi, Yunquan Sun
ICASSP2
2025 A Novel Tendon-Driven Articulated Continuum Robot with Stabilized Self-Locking Joints
abstract
Articulated continuum robots (ACRs) are characterized by flexibility, controllability, and adaptability and perform excellently in complex and constrained environments. However, the large number of motor drives limit the ACRs' portability and make them cumbersome to control. This paper presents a novel tendon-driven ACR composed of stabilized self-locking joints (SLJs) connected in series. After triggering the mechanical constraints with shape memory alloy coils, each joint can be maintained in either a self-locking or release state with zero power consumption. Consequently, even with a single set of drive units, the ACR can operate in multiple modes, enabling variable motion performance and workspace adaptability, effectively reducing the number of motors. The ACR's stiffness also varies with the locking state of its SLJs, and no motor drive is required to maintain its shape when all SLJs are self-locking. The performance and reliability of the SLJ prototype were validated. The workspace of the ACR prototype model was analyzed, and its partial motion performance, motion error, and variable stiffness were verified.
Jiankun Ren, Lizhe Qi, Hecheng Wang, Yunquan Sun
ICRA2
2025 Hierarchical Visual Policy Learning for Long-Horizon Robot Manipulation in Densely Cluttered Scenes
abstract
In this work, we focus on addressing the long-horizon packing tasks in densely cluttered scenes. Such tasks require policies to effectively manage severe occlusions among objects and continually produce precise actions based on visual observations. We propose a vision-based Hierarchical policy for Cluttered-scene Long-horizon Manipulation (HCLM). It employs a high-level policy and three options to select and instantiate three parameterized action primitives: push, pick, and place. We first train the two-stream pick and place options by behavior cloning (BC). Subsequently, we use hierarchical reinforcement learning (HRL) to train the high-level policy and push option. During HRL, we propose a Spatially Extended Q-update (SEQ) to augment the updates for the push option and a Two-Stage Update Scheme (TSUS) to alleviate the non-stationary transition problem in updating the high-level policy. We demonstrate that HCLM significantly outperforms baseline methods in terms of success rate and efficiency in diverse tasks both in simulation and real world. The ablation studies also validate the key roles of SEQ and TSUS in HRL.
Hecheng Wang, Lizhe Qi, Jiankun Ren, Yunquan Sun
ICRA2
2025 RSNet: Reflectance-Guided and Semantic-Aware Network for Low-light Image Enhancement
abstract
Low-light image enhancement (LLIE) aims to restore the degraded details of low-light images to obtain high-quality images with appropriate brightness. In recent years, deep learning methods have become the mainstream solution due to their excellent performance. However, these methods rarely explicitly use the real color, spatial structure, and texture detail of the image that are not affected by illumination to guide model learning, resulting in insufficient robustness of the network to restore image details under different illumination. Another problem is that such methods often improve lowlight images by global enhancement, ignoring the semantic information of different regions, resulting in the enhancement results easily deviating from the original color of the region. To solve these two problems, we propose a novel Reflectance-guided and Semantic-aware Network (RSNet). Specifically, we design a Reflectance Perception Module (RPM) and a Semantic Guidance Module (SGM). The RPM guides the model to learn the reflectance map information of the image that is not affected by illumination, thereby achieving clear and natural images under different lighting conditions. The SGM guides the model to perform differential enhancement on different regions during image restoration by cascading a semantic segmentation model after the image enhancement model, thereby obtaining a visual image that is more consistent with the real world. In addition, we construct a dataset of LOw-Light image pairs with fine SEMantic annotations (LOL-SEM) to evaluate algorithms for joint low-level and high-level tasks. A series of experiments on public image enhancement datasets and the LOL-SEM dataset show that our proposed network achieves excellent performance.
Lijun Dai, Yunquan Sun, Lizhe Qi
IJCNN3
2025 HSDNet: Hybrid Super-Resolution Assisted Detection Network for Surface Defect Detection
abstract
Surface defect detection plays a vital role in ensuring product quality in real industrial scene. Currently, deep learning-based detectors have achieved notable performance on different open-source high-quality image datasets. Nevertheless, constrained by the deployment costs of imaging equipment in real industrial scenes, the collected images frequently suffer from low resolution (LR). Defect detection on LR images is quite challenging due to the insufficient texture details and blurred boundaries. To address this problem, this paper proposes an advanced two-stage defect detection network, named HSDNet, which consists of an a well-designed Hybrid Super-Resolution (HSR) sub-network and a downstream defect detector RetinaNet. Firstly, the HSR performs effective pixel-level enhancement through the key module, Hybrid Scale Alignment (HSA), comprising the sub-attention modules Expanded Receptive Attention (ERA) and Local Context Attention (LCA), enabling the smooth restoration of multi-scale defect textures and improving both image resolution and quality. Secondly, to preserve potential detection-related features, we introduce an attention selective strategy within the HSR sub-network based on Adaptive Selection Attention (ASA) module. This strategy automatically assigns adaptive parameters to the ERA and LCA modules, thereby improving the representation capability for defect objects of varying scales and sizes. Meanwhile, we adopt a cascaded training strategy, where the downstream frozen detector’s detection results guide the updates of the HSR sub-network, ensuring alignment with detection objectives. Experimental results on two public defect datasets, NEU-DET and PKU-PCB, demonstrate that our proposed HSDNet outperforms several competitive baselines in terms of PSNR, while surpassing other SR methods’ downstream detectors in terms of mAP, proving its effectiveness in defect detection.
Xinzhi Lin, Yunquan Sun, Lizhe Qi
IJCNN4
2025 IF-DETR: Incremental Few-Shot Detection Transformer for Surface Defect Detection
abstract
Recognizing new categories that continually emerge during production is a significant challenge in defect detection, and deep learning methods have become the mainstream solution in recent years. However, in industrial scenarios, the scarcity of new category samples in the early phases leads to catastrophic forgetting and poor generalization in these methods. To address this, we propose a novel end-to-end Incremental Few-shot DEtection TRansformer (IF-DETR), which aims to resolve the knowledge ambiguity between old and new categories in such settings. Specifically, IF-DETR follows a two-stage paradigm involving pre-training and fine-tuning, leveraging knowledge distillation (KD) to improve the fine-tuning process while distinguishing label structures to eliminate ambiguous knowledge. We design two KD losses: Feature-level Instance Aware (FIA) loss and Logit-level Hierarchy Aligned (LHA) loss. For FIA loss, we exploit the global modeling ability of transformers to generate attention masks, thereby assigning reasonable weights to valuable distillation regions and mitigating catastrophic forgetting. For LHA loss, we decouple the teacher model’s output logits based on the supervision information, applying auxiliary losses at each layer of the decoder to refine the student model’s predictions and enhance the generalization ability for new categories. To the best of our knowledge, this is the first work to introduce transformer-based detector into incremental few-shot defect detection. We conduct extensive experiments on two benchmark defect detection datasets, and IF-DETR achieves the best performance across all 9 splits and few-shot settings, demonstrating its effectiveness.
Zhangxun Li, Xinzhi Lin, Jiamu Sheng, Nailei Hei, Lijun Dai, Lizhe Qi
IJCNN7
2025 FBI-Net: Foreground and Background Isolate Knowledge Distillation Network for Surface Defect Detection
abstract
In intelligent manufacturing, surface defect detection is essential for ensuring product quality and enhancing production efficiency. Although deep learning-based methods for defect detection show significant potential, optimization strategies focused on single models often address either speed or accuracy, while still facing three challenges in industrial settings: limited computational resources on deployed devices, minimal differences between foreground and background, and multi-scale defects. To tackle these issues simultaneously, we propose a novel surface defect detection method based on knowledge distillation, named FBI-Net. (i) We introduce foreground-background isolate knowledge distillation (FBI-KD), which transfers foreground and background knowledge separately during training, thereby enhancing the feature differences while maintaining detection speed. (ii) We propose a group attention mechanism (GAM) that extracts key channels and positions from feature representations, allowing the student model to focus on the most valuable knowledge. (iii) We design a convolution selection module (CSM) that assigns weights to different convolution layers to improve the detection capability for multi-scale defects. Rigorous experiments demonstrate our FBI-Net is competitive with existing methods in both detection accuracy and speed, providing a reliable solution for real-time processing demands in practical industrial applications.
Xinzhi Lin, Yunquan Sun, Lizhe Qi
IJCNN5
2025 SLU-DQN: A Model for Anticipatory Steam Detection for Steamer-Filling in Baijiu Intelligent Distillation Systems
abstract
The true implementation of the Anticipatory Steam Detection for Steamer-Filling(ASDSF) process in baijiu intelligent distillation systems, which involves predicting and precisely spreading distillers’ grains before steam emerges, remains a critical unresolved challenge. In this study, we introduce the SLU model, which utilizes SwinLSTM as the core feature extraction module and adopts a U-shaped structure. This model achieves spatiotemporal feature extraction and dynamic change prediction. It is further enhanced by integrating a U-Net module for multi-scale feature fusion and optimized through a Deep Q-Network (DQN)-based decision-making process. The SLU-DQN model, specifically designed for anticipatory material spreading planning in the baijiu Steamer-Filling(SF) distillation system, predicts future steam emission areas. Finally, both quantitative and qualitative experimental results demonstrate the excellent performance of the SLU-DQN model in solving the ASDSF problem. The model achieved 91.1% reward accuracy, an F1-Score of 91% for material spreading point prediction, an MSE of 19.02, and an SSIM of 95.8%. These results not only highlight the model’s superior accuracy in predicting future steam emission areas but also provide a significant technical breakthrough for intelligent baijiu distillation systems, filling a crucial gap in the field.
Jiankun Ren, Hanwen Liang, Lizhe Qi, Yunquan Sun
IROS5
2025 RWKV3D: An RWKV-Based Model with Multiple Training Strategies for Point Cloud Analysis
abstract
Transformer-based models have achieved dominance in point cloud analysis, yet their quadratic computational complexity remains a fundamental limitation for practical applications. Recently, RWKV has emerged as a promising alternative for sequence modeling due to its linear computational complexity. However, it has yet to be effectively adapted to handle the unordered and sparse nature of point cloud data. In this paper, we propose RWKV3D, an innovative and computational framework tailored for point cloud analysis, which is adaptable to three training strategies: training from scratch, single-modal pre-training, and cross-modal pre-training. First, we replace the MLP layer with an advanced Local Feature Mixer (LFM), which not only enhances fine-grained feature extraction but also reduces the number of parameters. Second, we introduce a Bidirectional Multi-head Shift (BMS) mechanism to expand the receptive field, effectively capturing richer contextual information. Additionally, to enhance high-level feature processing, we strategically incorporate a Multi-head Self-Attention (MSA) block before the first RWKV3D block. Experimental results demonstrate that RWKV3D outperforms Transformer-based and Mamba-based methods while maintaining lower parameter counts and computational costs. Notably, it achieves several state-of-the-art results, including overall accuracies of 95.3% (training from scratch) and 95.9% (cross-modal pre-training) on the ModelNet40 dataset, as well as 95.28% (single-modal pre-training) on the ScanObjectNN (PB_T50_RS) dataset. These results underscore the superior efficacy of the RWKV architecture in 3D vision tasks and highlight its potential for broader multimodal learning scenarios.
Chenglong Sun, Shijie Pang, Lizhe Qi
ACM Multimedia4
2024 Out of Thin Air: Exploring Data-Free Adversarial Robustness Distillation
abstract
Adversarial Robustness Distillation (ARD) is a promising task to solve the issue of limited adversarial robustness of small capacity models while optimizing the expensive computational costs of Adversarial Training (AT). Despite the good robust performance, the existing ARD methods are still impractical to deploy in natural high-security scenes due to these methods rely entirely on original or publicly available data with a similar distribution. In fact, these data are almost always private, specific, and distinctive for scenes that require high robustness. To tackle these issues, we propose a challenging but significant task called Data-Free Adversarial Robustness Distillation (DFARD), which aims to train small, easily deployable, robust models without relying on data. We demonstrate that the challenge lies in the lower upper bound of knowledge transfer information, making it crucial to mining and transferring knowledge more efficiently. Inspired by human education, we design a plug-and-play Interactive Temperature Adjustment (ITA) strategy to improve the efficiency of knowledge transfer and propose an Adaptive Generator Balance (AGB) module to retain more data information. Our method uses adaptive hyperparameters to avoid a large number of parameter tuning, which significantly outperforms the combination of existing techniques. Meanwhile, our method achieves stable and reliable performance on multiple benchmarks.
Zhaoyu Chen 0001, Dingkang Yang, Pinxue Guo, Kaixun Jiang, Lizhe Qi
AAAI7
2024 De-Confounded Data-Free Knowledge Distillation for Handling Distribution Shifts
abstract
Data-Free Knowledge Distillation (DFKD) is a promising task to train high-performance small models to enhance actual deployment without relying on the original training data. Existing methods commonly avoid relying on private data by utilizing synthetic or sampled data. However, a long-overlooked issue is that the severe distribution shifts between their substitution and original data, which mani-fests as huge differences in the quality of images and class proportions. The harmful shifts are essentially the con-founder that significantly causes performance bottlenecks. To tackle the issue, this paper proposes a novel perspective with causal inference to disentangle the student models from the impact of such shifts. By designing a customized causal graph, we first reveal the causalities among the variables in the DFKD task. Subsequently, we propose a Knowledge Distillation Causal Intervention (KDCI) framework based on the backdoor adjustment to de-confound the confounder. KDCI can be flexibly combined with most existing state-of-the-art baselines. Experiments in combination with six representative DFKD methods demonstrate the effectiveness of our KDCI, which can obviously help existing methods under almost all settings, e.g., improving the base-line by up to 15.54% accuracy on the CIFAR-100 dataset.
Dingkang Yang, Zhaoyu Chen 0001, Yang Liu 0246, Siao Liu, Lihua Zhang 0002, Lizhe Qi
CVPR8
2024 Self-cooperation Knowledge Distillation for Novel Class Discovery
Zhaoyu Chen 0001, Dingkang Yang, Yunquan Sun, Lizhe Qi
ECCV (70)5
2024 An Arc Light Elimination Network Using Polarization and Prior Information
Wenhao Yu 0009, Lizhe Qi, Hanwen Liang, Yunquan Sun
IJCNN2
2024 Sampling to Distill: Knowledge Transfer from Open-World Data
abstract
Data-Free Knowledge Distillation (DFKD) is a novel task that aims to train high-performance student models using only the pre-trained teacher network without original training data. Most of the existing DFKD methods rely heavily on additional generation modules to synthesize the substitution data resulting in high computational costs and ignoring the massive amounts of easily accessible, low-cost, unlabeled open-world data. Meanwhile, existing methods ignore the domain shift issue between the substitution data and the original data, resulting in knowledge from teachers not always trustworthy and structured knowledge from data becoming a crucial supplement. To tackle the issue, we propose a novel Open-world Data Sampling Distillation (ODSD) method for the DFKD task without the redundant generation process. First, we try to sample open-world data close to the original data's distribution by an adaptive sampling module and introduce a low-noise representation to alleviate the domain shift issue. Then, we build structured relationships of multiple data examples to exploit data knowledge through the student model itself and the teacher's structured representation. Extensive experiments on CIFAR-10, CIFAR-100, NYUv2, and ImageNet show that our ODSD method achieves state-of-the-art performance with lower FLOPs and parameters. Especially, we improve 1.50%-9.59% accuracy on the ImageNet dataset and avoid training the separate generator for each class.
Zhaoyu Chen 0001, Jie Zhang 0107, Dingkang Yang, Zuhao Ge, Yang Liu 0246, Siao Liu, Yunquan Sun, Lizhe Qi
ACM Multimedia10
2024 Mixed noise-guided mutual constraint framework for unsupervised anomaly detection in smart industries
Qing Zhao 0007, Yan Wang 0068, Yuxuan Lin 0001, Shaoqi Yan, Wei Song 0007, Boyang Wang 0003, Yang Chang, Lizhe Qi
Comput. Commun.9
2023 Adversarial Contrastive Distillation with Adaptive Denoising
abstract
Adversarial Robustness Distillation (ARD) is a novel method to boost the robustness of small models. Unlike general adversarial training, its robust knowledge transfer can be less easily restricted by the model capacity. However, the teacher model that provides the robustness of knowledge does not always make correct predictions, interfering with the student’s robust performance. Besides, in the previous ARD methods, the robustness comes entirely from one-to-one imitation, ignoring the relationship between examples. To this end, we propose a novel structured ARD method called Contrastive Relationship DeNoise Distillation (CRDND). We design an adaptive compensation module to model the instability of the teacher. Moreover, we utilize the contrastive relationship to explore implicit robustness knowledge among multiple examples. Experimental results on multiple attack benchmarks show CRDND can transfer robust knowledge efficiently and achieves state-of-the-art performance.
Zhaoyu Chen 0001, Dingkang Yang, Yang Liu 0246, Siao Liu, Lizhe Qi
ICASSP7
2023 Explicit and Implicit Knowledge Distillation via Unlabeled Data
abstract
Data-free knowledge distillation is a challenging model lightweight task for scenarios in which the original dataset is not available. Previous methods require a lot of extra computational costs to update one or more generators and their naive imitate-learning lead to lower distillation efficiency. Based on these observations, we first propose an efficient unlabeled sample selection method to replace high computational generators and focus on improving the training efficiency of the selected samples. Then, a class-dropping mechanism is designed to suppress the label noise caused by the data domain shifts. Finally, we propose a distillation method that incorporates explicit features and implicit structured relations to improve the effect of distillation. Experimental results show that our method can quickly converge and obtain higher accuracy than other state-of-the-art methods.
Zuhao Ge, Zhaoyu Chen 0001, Chuangjia Ma, Yunquan Sun, Lizhe Qi
ICASSP7
2023 On the Importance of Spatial Relations for Few-shot Action Recognition
abstract
Deep learning has achieved great success in video recognition, yet still struggles to recognize novel actions when faced with only a few examples. To tackle this challenge, few-shot action recognition methods have been proposed to transfer knowledge from a source dataset to a novel target dataset with only one or a few labeled videos. However, existing methods mainly focus on modeling the temporal relations between the query and support videos while ignoring the spatial relations. In this paper, we find that the spatial misalignment between objects also occurs in videos, notably more common than the temporal inconsistency. We are thus motivated to investigate the importance of spatial relations and propose a more accurate few-shot action recognition method that leverages both spatial and temporal information. Particularly, a novel Spatial Alignment Cross Transformer (SA-CT) which learns to re-adjust the spatial relations and incorporates the temporal information is contributed. Experiments reveal that, even without using any temporal information, the performance of SA-CT is comparable to temporal based methods on 3/4 benchmarks. To further incorporate the temporal information, we propose a simple yet effective Temporal Mixer module. The Temporal Mixer enhances the video representation and improves the performance of the full SA-CT model, achieving very competitive results. In this work, we also exploit large-scale pretrained models for few-shot action recognition, providing useful insights for this research direction.
Yuqian Fu, Xingjun Ma, Lizhe Qi, Jingjing Chen 0001, Zuxuan Wu, Yu-Gang Jiang 0001
ACM Multimedia4
2023 Cleaning of object surfaces based on deep learning: a method for generating manipulator trajectories using RGB-D semantic segmentation
Lizhe Qi, Zhongwei Hua, Daming Du, Wenxuan Jiang, Yunquan Sun
Neural Comput. Appl.1
2022 Dual Attention Based Multi-scale Feature Fusion Network for Indoor RGBD Semantic Segmentation
abstract
RGBD semantic segmentation combined with color image information and depth information can effectively alleviate the problems of low classification accuracy and difficulty in accurately dividing edges between different semantic regions in indoor scenes caused by complex backgrounds, uneven lighting, similar object textures, spatial overlap, and occlusion. To fully fuse the color features with the spatial position and hierarchical information of objects, this paper proposes a multi-scale network model based on the dual attention mechanism (channel attention and spatial attention)(DAMFNet), which effectively integrates color texture features and spatial structure features, and further improves the semantic segmentation performance of indoor objects. We evaluate the proposed network model on the common indoor dataset SUNRGBD and achieve state-of-the-art results. In addition, this paper also demonstrates the excellent segmentation accuracy and effect of the proposed network model on self-built indoor datasets and in real-world application scenarios.
Zhongwei Hua, Lizhe Qi, Daming Du, Wenxuan Jiang, Yunquan Sun
ICPR2
2022 C2F-CFN: Coarse-to-Fine ClothFlow Network for High-Fidelity Virtual Try-On
abstract
Image-based virtual try-on refers to the task of transferring a target clothing item onto the corresponding area of a person. While deep generative models have made considerable progress in virtual try-on, current methods are incapable of handling the rendering of try-on details, and they generally suffer from unrealistic try-on images with artifacts due to the improper warping of clothing. In this paper, we propose a novel virtual try-on network, called C2F-CFN, which models the per-pixel appearance flow via a coarse-to-fine manner. Specifically, the affine transformation is learned first to obtain the coarse warped clothing, and then the texture and geometric details of clothing are refined through appearance flow estimation. Furthermore, the oracle flow supervision is introduced to produce more natural and reasonable flow maps. We evaluate our model on the VITON dataset, and demonstrate that our approach outperforms the state-of-the-art techniques quantitatively and qualitatively.
Yanli Bi, Lizhe Qi, Yunquan Sun
IJCNN2
2022 Density and Context Aware Network with Hierarchical Head for Traffic Scene Detection
abstract
We investigate traffic scene detection from surveillance cameras and UAVs. This task is rather challenging, mainly due to the spatial nonuniform gathering, large-scale variance, and instance-level imbalanced distribution of vehicles. Most existing methods that employed FPN to enrich features are prone to failure in this scenario. To mitigate the influences above, we propose a novel detector called Density and Context Aware Network(DCANet) that can focus on dense regions and adaptively aggregate context features. Specifically, DCANet consists of three components: Density Map Supervision(DMP), Context Feature Aggregation(CFA), and Hierarchical Head Module(HHM). DMP is designed to capture the gathering information of objects supervised by density maps. CFA exploits adjacent feature layers' relationships to fulfill ROI-level contextual information enhancement. Finally, HHM is introduced to classify and locate imbalanced objects employed in hierarchical heads. Without bells and whistles, DCANet can be used in any two-stage detectors. Extensive experiments are carried out on the two widely used traffic detection datasets, CityCam and VisDrone, and DCANet reports new state-of-the-art scores on the CityCam.
Zuhao Ge, Wenhao Yu 0009, Lizhe Qi, Yunquan Sun
IJCNN4
2022 Colour balance and contrast stretching for sand-dust image enhancement
abstract
Abstract The increasingly frequent sand‐dust weather in the inland areas seriously affects outdoor vision applications, especially autonomous vehicles and security monitoring. To moderate the image's colour cast and poor contrast caused by sand‐dust weather, an effective approach is proposed in this study to enhance the sand‐dust images. First, the original degraded image's colour cast is corrected by a new colour balance and compensation formula, which compensates the blue and green channel information through numerous yellow channel information caused by sand‐dust scattering before white balance. Next, in order to avoid the new colour deviation, the corrected image is converted from the RGB colour space to the HSV colour space and use the CLAHE to enhance the V component to improve the contrast. Then, a nonlinear gain function is defined to further adaptively sharpen the V component to enhance image details. Finally, the S component is stretched to improve image saturation. The extensive qualitative and quantitative evaluation shows that this method can improve the image edge clarity and contrast, restore good colour fidelity for all sand‐dust images tested. The verification also proves that this method is of much significance in improving the feature point extraction and the target detection results in the sand‐dust weather.
Zhongwei Hua, Lizhe Qi, Min Guan, Yunquan Sun
IET Image Process.2
2022 Text Representation Model for Multiple Language Forms in Spoken Chinese Expression
abstract
Mixture of multiple language forms in spoken Chinese is a common but unfavorable issue.. It increases the difficulty of intent understanding and leads to inconvenience for information communication. Existing studies on intent recognition mainly focus on single language form or parallel multilingual language while paying little attention to spoken texts including multiple language forms. In considering that it is hard to capture the semantics of an expression with multiple language forms, it is important to study the problem. To solve this issue, a text representation model for the spoken Chinese expression mixed with English and Chinese Pinyin is proposed. And the feature matrix is built to mine the composition information of English and Pinyin. Besides, the model can efficiently distinguish English from Chinese Pinyin even though both fragments are composed of English letters. Meanwhile, it can effectively process the problem of hidden text information since the problem has been transformed into the Chinese translation task of English and Pinyin. In addition, to verify the performance of the model, the texts processed by this model are used as the input of the classifier. extensive experiments on a large online logistics manual customer service corpus show that this text representation model is correct and effective. It can not only eliminate the obstacles of the mixing of multiple language forms but also bring better results for intent understanding.
Miao Hu 0004, Jingxiang Hu, Lizhe Qi, Huanxiang Zhang
Int. J. Pattern Recognit. Artif. Intell.5
2022 Zoom-and-Reasoning: Joint Foreground Zoom and Visual-Semantic Reasoning Detection Network for Aerial Images
abstract
Aerial image object detection remains rather challenging, due to the small object gathering and confusion of inter-class similarities and intra-class diversity. Confronting such challenges, we propose a two-stage framework, ‘Zoom&Reasoning Det,’ which performs detection in a foreground highlight manner and leverages contextual relations to assist detection. In the coarse foreground zoom stage, different from earlier works that divide original images into patches and perform detection on each patch separately, we design Foreground Zoom Strategy (FZS), which zooms foreground dense regions from a coarse detector and packs them into one image. In the fine reasoning detect stage, motivated by a human visual mechanism that can achieve correct recognition by reasoning through context, we present Visual-Semantic Reasoning Network (VSRNet), consisting of Visual Reasoning Graph (VRG) and Semantic Reasoning Graph (SRG), simultaneously considering local visual and global semantic contextual relational information for each instance. Each instance feature representation is further refined by aggregating the outputs of VSRNet. Comprehensive experiments conducted on two challenging aerial image datasets, VisDrone and UAVDT, demonstrate the advantage of our method over the state-of-the-art.
Zuhao Ge, Lizhe Qi, Yunquan Sun
IEEE Signal Process. Lett.2
2022 A Soft Robot With Variable Stiffness Multidirectional Grasping Based on a Folded Plate Mechanism and Particle Jamming
abstract
Despite good performance in grasping irregular fragile objects, soft grippers exhibiting low stiffness, and carrying capacity lack multidirectional grasping ability in the case of inclination. To address this problem, we propose a novel rigid and soft coupling variable stiffness module that employs a folded plate mechanism (FPM) to provide rigid multidirectional loading and combines it with particle jamming to achieve local variable stiffness with the characteristics of a finger grasping structure. Hence, a variable stiffness multidirectional soft grasping robot is developed to realize soft grasping and multidirectional rigid loading. The bending and stiffness control of the variable stiffness soft gripper is realized by a double-layer pneumatic driving structure with the advantages of simple control and corresponding speed. Moreover, good self-recovery is achieved with the soft outer layer since the particles can quickly return to the initial state due to partitioning of the FPM. Finally, prototype experiments verify its strong adaptability and stable multidirectional grasping ability, and experimental results show that the maximum grasping weight in each direction can be increased by more than three times with the variable stiffness.
Hang Wei 0004, Yu Shan, Yanzhi Zhao, Lizhe Qi, Xilu Zhao
IEEE Trans. Robotics4
2021 SimDet: Cross Similarity Attention for One-shot Object Detection
abstract
Object detection based on the convolutional neural network requires a large of datasets for training to achieve good results. However, it is labor-intensive or unrealistic to prepare such high-quality training data in most industrial applications. Recently, one-shot object detection task was proposed aiming to tackle this challenging problem by using only one sample for reference. In this work, a new framework named SimDet based on Faster R-CNN has been proposed for one-shot object detection. Specifically, query and target image features are extracted through a Siamese network, and target features are enhanced in where has high similarity with query feature. In order to solve the problem of RPN learning difficulty with one sample, we propose a new module of Cross Similarity Module, which focuses more on the difference between query and support rather than the feature itself. Furthermore, we design a similarity loss for learning the cross similarity between target feature and query feature. Finally, the extensive experiments show that our model achieves state-of-the-art performance on PASCAL VOC under one-shot setting of detecting objects from both seen and novel classes and on MS COCO from seen classes.
Rujia Cai, Yingjie Qin, Lizhe Qi, Yunquan Sun
IJCNN3
2021 An intention multiple-representation model with expanded information
Jingxiang Hu, Lizhe Qi, Miao Hu 0004, Huanxiang Zhang
Comput. Speech Lang.4
2020 Adaptive Graph Reasoning Network for Fashion Landmark Detection
abstract
In this paper, we address the fashion landmark detection task by enforcing structural fashion layout relationships among landmarks based on Graph Convolutional Networks (GCNs). Unlike previous works that detect each fashion landmark separately and ignore the rich semantic layout relation among different landmarks, we propose an Adaptive Graph Reasoning Network (AGRNet) to integrate the convolutional features with the human commonsense knowledge and make detected fashion landmarks be coherent with clothes layouts from a global perspective. Specifically, we design the Adaptive Graph Reasoning (AGR) module and stack it on top of Fully Convolutional Networks (FCNs), which enforces fashion layout constraints and semantic relations of fashion landmarks on deep representations. AGR maps the convolutional features into structural graph node representations and performs adaptive reasoning according to the correlation matrix, which is adaptively generated from defined basic fashion layout and confidence maps of all landmarks. The graph-based reasoning evolves the cloth node representations to achieve global layout coherency and then the evolved graph nodes are mapped back to enhance convolutional feature representations. Furthermore, we design the Dual Attention Upsample (DAU) module on each decoder layer to emphasize the spatial detailed and task-related features by modelling the semantic interdependencies in spatial and channel dimensions respectively. We achieve new state-of-the-art detection performance on two challenging fashion landmark datasets, i.e., Deepfashion and FLD dataset. In particular, a Normalized Error (NE) score of 0.0297 on the Deepfashion test set is achieved without any additional annotations.
Hang Ying, Yingjie Qin, Lizhe Qi, Yunquan Sun
ECAI4
2020 Multi-Scale Deep Feature Fusion for Vehicle Re-Identification
abstract
Vehicle re-identification (re-id) is challenging due to the small inter-class distance. The differences between similar vehicles can be extremely subtle and only captured at particular scales and semantic levels. In this paper, we propose a novel Multi-Scale Deep Feature Fusion Network (MSDeep) to conduct both multi-scale and multi-level features for precise vehicle re-id. Based on the backbone deep CNN, MS-Deep mainly consists of two modules: 1) Multi-Scale Fusion (MSF) Block which encapsulates combination of multi-scale streams as MSF feature; 2) Multi-Level Fusion (MLF) Block which fuses MSF features of multiple levels to build the final descriptor. Importantly, in MSF, Multi-Scale Attention (MSA) is introduced to dynamically emphasize important channels of each scale, and Level-Wise Attention(LWA) is utilized in MLF to determine the different weightings for each MSF feature of different levels. As a result, experiments show that our MSDeep outperforms state-of-the-art algorithms on challenging VeRi and VehicleID benchmarks in terms of abundant and hierarchical hyper-descriptors.
Yiting Cheng 0001, Chuanfa Zhang, Kangzheng Gu, Lizhe Qi, Zhongxue Gan 0001
ICASSP4
2020 All In One Network for Driver Attention Monitoring
abstract
Nowadays, driver drowsiness and driver distraction is considered as a major risk for fatal road accidents around the world. As a result, driver monitoring identifying is emerging as an essential function of automotive safety systems. Its basic features include head pose, gaze direction, yawning and eye state analysis. However, existing work has investigated algorithms to detect these tasks separately and was usually conducted under laboratory environments. To address this problem, we propose a multi-task learning CNN framework which simultaneously solve these tasks. The network is implemented by sharing common features and parameters of highly related tasks. Moreover, we propose Dual-Loss Block to decompose the pose estimation task into pose classification and coarse-to-fine regression and Objectcentric Aware Block to reduce orientation estimation errors. Thus, with such novel designs, our model not only achieves SOA results but also reduces the complexity of integrating into automotive safety systems. It runs at 10 fps on vehicle embedded systems which marks a momentous step for this field. More importantly, to facilitate other researchers, we publish our dataset FDUDrivers which contains 20000 images of 100 different drivers and covers various real driving environments. FDUDrivers might be the first comprehensive dataset regarding driver attention monitoring.
Xiaotian Dai 0001, Lizhe Qi, Zhe Jiang 0004
ICASSP5
2020 Intention Multiple-Representation Model for Logistics Intelligent Customer Service
Jingxiang Hu, Lizhe Qi, Miao Hu 0004, Huanxiang Zhang
KSEM (1)4
2020 Multi-condition Place Generator for Robust Place Recognition
Yiting Cheng 0001, Lizhe Qi
MMM (1)3
2019 Automatic Tongue Image Segmentation For Real-Time Remote Diagnosis
abstract
Tongue diagnosis, one of the essential diagnostic methods of Traditional Chinese Medicine (TCM), is considered an ideal candidate for remote diagnosis methods because of its convenience and noninvasiveness. However, the trade-off between accuracy and efficiency and the variation of tongue images pose great challenges in real-time tongue image segmentation. To remedy these problems, in this paper, a light weight architecture based on the encoder-decoder structure is proposed. The tongue image feature extraction (TIFE) module is designed to generate features with larger receptive fields without sacrificing spatial resolution. The context module is used to increase the performance by aggregating multi-scale contextual information. The decoder is designed as a simple yet efficient feature upsampling module to fuse different depth features and refine the segmentation results along tongue boundaries. The loss module is proposed to deal with misclassifications causing by class imbalance. A new tongue image dataset (FDU/SHUTCM) is constructed for model training and testing, which contains 5,600 tongue images and their corresponding high quality masks. We demonstrate the effectiveness of the proposed model on BioHit, PolyU/HIT, and our datasets, achieving the performance of 99.15%, 95.69%, and 99.03% IoU accuracy, respectively. Segmentation of a 513×513 image takes 165 ms on CPU.
Yan Wang 0068, Lizhe Qi, Fufeng Li, Zhongxue Gan 0001
BIBM5
2019 Reading Face, Reading Health: Exploring Face Reading Technologies for Everyday Health
abstract
With the recent advancement in computer vision, Artificial Intelligence (AI), and mobile technologies, it has become technically feasible for computerized Face Reading Technologies (FRTs) to learn about one's health in everyday settings. However, how to design FRT-based applications for everyday health practices remains unexplored. This paper presents a design study with a technology probe called Faced, a mobile health checkup application based on the facial diagnosis method from Traditional Chinese Medicine (TCM). A field trial of Faced with 10 participants suggests potential usage modes and highlights a number of critical design issues in the use of FRTs for everyday health, including adaptability, practicality, sensitivity, and trustworthiness. We end by discussing design implications to address the unique challenges of fully integrating FRTs into everyday health practices.
Xianghua Ding, Yanqi Jiang, Xiankang Qin, Yunan Chen 0001, Lizhe Qi
CHI6
2019 Learning Camera-Invariant Representation for Person Re-identification
Shizheng Qin, Kangzheng Gu, Lecheng Wang, Lizhe Qi
ICANN (2)4
2019 Focus Generator with Score Classification on Fabric Defect Detection
abstract
Fabric defect detection plays an important role in the production of fabrics. Thanks to deep learning and large-scale datasets, popular object detection tasks have made great progress. RetinaNet has been widely used in object detection tasks in various fields, such as face detection, due to its flexibility and operability. It's high accuracy and detection results are derived from datasets with rich features. So limited by the small-scale fabric defect dataset, RetinaNet is difficult to apply to fabric defect detection task. In this paper, we propose an effective neural network approach to solve the problem of small-scale fabric dataset and apply RetinaNet to this task. To overcome the insufficient features because of the small-scale dataset, we first propose a generative model to add Gaussian noise on latent space, called focus generator, which can be controlled with defect instances to generate more data. Then we add a classification model to limit influence produced by the focus generator, called score classification. Finally, we merge the focus generator and score classification with an improved RetinaNet to achieve fabric defect detection, therefore, we name our model FSR. By the way, the operations of adding noise are different on the steps of training and testing. The experimental result shows that our proposed method can achieve better performance comparing to our baseline RetinaNet and finally achieve accuracy of 83.4% on our small-scale fabric defect dataset.
Yingjie Qin, Lizhe Qi, Yunquan Sun
ICTAI3
2019 Implicit Rating Methods Based on Interest Preferences of Categories for Micro-Video Recommendation
Lizhe Qi, Gan Chen
KSEM (1)3
2019 L0 Gradient Smoothing and Bimodal Histogram Analysis: A Robust Method for Sea-sky-line Detection
abstract
Sea-sky-line detection is an important research topic in the field of object detection and tracking on the sea. We propose an L0 gradient smoothing and bimodal histogram analysis based method to improve the robustness and accuracy of sea-sky-line detection. The proposed method mainly depends on the brightness difference between the sea region and the sky region in the image. First, we use L0 gradient smoothing to eliminate discrete noise in the image and achieve the modularity of brightness. Differing from previous methods, diagonal dividing is applied to obtain the brightness thresholds for the sky and sea regions. Then the thresholds are used for bimodal histogram analysis which helps to obtain the brightness near the sea-sky-line and narrow the detection region. After narrowing the detection region, the sea-sky-line in the image is extracted by a linear fitting method. To evaluate the performance of the proposed method, we manually construct an dataset which includes 40, 000 images taken in five scenes. Moreover, we also mark the corresponding ground-truth positions of sea-sky-line in each of the images. Extensive experiments on the dataset demonstrate that our method outperforms the state-of-the-art methods tremendously.
Hong Lu 0001, Lizhe Qi
MMAsia5