VLDB 2026 Research / reviewers in the wild / expert
Qingli Li
dblp:77/8026
· DBLP profile ↗
85ranked-venue papers
0as first author
71since 2021 · last 2027
0000-0001-5063-8801ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 55 · 45 since 2021Artificial intelligence and machine learning · 35 · 33 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 17 since 2021Computer networks · 3 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | DiffCL: Diffusion contrastive learning for strip steel surface defect classification
Xingcheng Zhu, Qingli Li |
Expert Syst. Appl. | 6 |
| 2026 | SES-Net: Semantic-edge synergistic network for industrial small defect detection
Zhe Liu 0004, Luhao Xia, Kai Han 0006, Jun Chen 0030, Jinyao Zhu, Shiyu Gan, Xiaocheng Hu, Qingli Li, Yi Liu 0114 |
Pattern Recognit. | 10 |
| 2025 | MDN: Mamba-Driven Dualstream Network For Medical Hyperspectral Image SegmentationabstractMedical Hyperspectral Imaging (MHSI) offers potential for computational pathology and precision medicine. However, existing CNN and Transformer struggle to balance segmentation accuracy and speed due to high spatial-spectral dimensionality. In this study, we leverage Mamba’s global context modeling to propose a dual-stream architecture for joint spatial-spectral feature extraction. To address the limitation of Mamba’s unidirectional aggregation, we introduce a recurrent spectral sequence representation to capture low-redundancy global spectral features. Experiments on a public Multi-Dimensional Choledoch dataset and a private Cervical Cancer dataset show that our method outperforms state-of-the-art approaches in segmentation accuracy while minimizing resource usage and achieving the fastest inference speed. Our code will be available at https://github.com/DeepMed-Lab-ECNU/MDN. Shijie Lin, Boxiang Yun, Wei Shen 0002, Qingli Li, Anqiang Yang, Yan Wang 0033 |
ICASSP | 4 |
| 2025 | Self-Prompting Driven SAM2 for 3D Medical Image SegmentationabstractThe latest advancement in large foundational model, SAM2, has demonstrated significant potential in 3D medical image segmentation due to their capability to effectively segment video streams. However, its application in medical image segmentation presents challenges, requiring extensive training on medical images or high-quality prompts provided by experts to achieve optimal performance. To address the aforementioned limitations, we propose SAM2-SP, which adopts Low-Rank Adaption for parameter-efficient fine-tuning and introduces a novel dynamic self-prompting strategy that generates most confident prompt templates from voxel features, enabling SAM2 to achieve domain adaptation in medical image segmentation without reliance on expert-level prompts. Extensive experiments show that SAM2-SP achieves state-of-the-art performance on the public Synapse dataset and the private EDC dataset, and even outperforms the compared task-specific segmentation approaches, the vanilla SAM and other SAM-based approaches. Sheng Wei 0004, Song Qiu, Mei Zhou, He Zhang 0023, Yan Wang 0033, Qingli Li |
ICASSP | 6 |
| 2025 | RobusTReID: Defending Vision Transformer for Robust Image ReIDabstractVision Transformer (ViT) achieves competitive results in person ReID, not only due to the powerful ability on feature representation, but also its resistance to attacks. However, they are vulnerable to specially designed attacks. To enhance their robustness, this paper proposes RobusTReID to defend the ViT-based model against perturbed images without obvious performance drop on clean data. The basic idea is to incorporate the adversarial co-training into ReID, which first disturbs pixels by minimizing adversarial loss in primary feature branch, then optimizes model by ReID task loss computed in all branches on clean and perturbed data. We separate the representation paths for clean and perturbed images. Particularly, a learnable [ADV] token and low-rank positional embeddings (PE) are incorporated to build the feature for perturbed image. Extensive experiments on several ReID datasets show that our method effectively increases the robustness of ReID model under different types of attacks. Tingting Xiao, Li Sun 0012, Qingli Li |
ICME | 4 |
| 2025 | DragLoRA: Online Optimization of LoRA Adapters for Drag-based Image Editing in Diffusion ModelabstractDrag-based editing within pretrained diffusion model provides a precise and flexible way to manipulate foreground objects. Traditional methods optimize the input feature obtained from DDIM inversion directly, adjusting them iteratively to guide handle points towards target locations. However, these approaches often suffer from limited accuracy due to the low representation ability of the feature in motion supervision, as well as inefficiencies caused by the large search space required for point tracking. To address these limitations, we present DragLoRA, a novel framework that integrates LoRA (Low-Rank Adaptation) adapters into the drag-based editing pipeline. To enhance the training of LoRA adapters, we introduce an additional denoising score distillation loss which regularizes the online model by aligning its output with that of the original model. Additionally, we improve the consistency of motion supervision by adapting the input features using the updated LoRA, giving a more stable and accurate input feature for subsequent operations. Building on this, we design an adaptive optimization scheme that dynamically toggles between two modes, prioritizing efficiency without compromising precision. Extensive experiments demonstrate that DragLoRA significantly enhances the control precision and computational efficiency for drag-based image editing. The Codes of DragLoRA are available at: https://github.com/Sylvie-X/DragLoRA. Siwei Xia, Li Sun 0012, Tiantian Sun, Qingli Li |
ICML | 4 |
| 2025 | Advancing Stain Transfer for Multi-Biomarkers: A Human Annotation-Free Method Based on Auxiliary Task SupervisionabstractHistopathological examination primarily relies on hematoxylin and eosin (H&E) and immunohistochemical (IHC) staining. Though IHC provides more crucial molecular information for diagnosis, it is more costly than H&E staining. Stain transfer technology seeks to efficiently generate virtual IHC images from H&E images. While current deep learning-based methods have made progress, they still struggle to maintain pathological and structural consistency across biomarkers without pixel-level aligned reference. To address the problem, we propose an Auxiliary Task supervision-based Stain Transfer method for multi-biomarkers (ATST-Net), which pioneeringly employs human annotation-free masks as ground truth (GT). ATST-Net ensures pathological consistency, structural preservation and style transfer. It automatically annotates H&E masks in a cost-effective manner by utilizing consecutive IHC sections. Multiple auxiliary tasks provide diverse supervisory information on the location and intensity of biomarker expression, ensuring model accuracy and interpretability. We design a pretrained model-based generator to extract deep feature in H&E images, improving generalization performance. Extensive experiments demonstrate the effectiveness of ATST-Net's components. Compared to existing methods, ATST-Net achieves state-of-the-art (SOTA) accuracy on datasets with multiple biomarkers and intensity levels, while also reflecting high practical value. Code is available at https://github.com/SikangSHU/ATST-Net. Haofei Song, Yingjiao Deng, Jiansheng Wang, Yan Wang 0033, Qingli Li |
IJCAI | 6 |
| 2025 | Historical Report Guided Bi-modal Concurrent Learning for Pathology Report Generation
Boxiang Yun, Qingli Li, Yan Wang 0033 |
MICCAI (6) | 3 |
| 2025 | Deep Association Multimodal Learning for Zero-Shot Spatial Transcriptomics Prediction
Yijing Zhou, Yadong Lu, Qingli Li, Xinxing Li, Yan Wang 0033 |
MICCAI (6) | 3 |
| 2025 | Efficient Multi-Slide Visual-Language Feature Fusion for Placental Disease ClassificationabstractAccurate prediction of placental diseases via whole slide images (WSIs) is critical for preventing severe maternal and fetal complications. However, WSI analysis presents significant computational challenges due to the massive data volume. Existing WSI classification methods encounter critical limitations: (1) inadequate patch selection strategies that either compromise performance or fail to sufficiently reduce computational demands, and (2) the loss of global histological context resulting from patch-level processing approaches. To address these challenges, we propose an Efficient multimodal framework for Patient-level placental disease Diagnosis, named EmmPD. Our approach introduces a two-stage patch selection module that combines parameter-free and learnable compression strategies, optimally balancing computational efficiency with critical feature preservation. Additionally, we develop a hybrid multimodal fusion module that leverages adaptive graph learning to enhance pathological feature representation and incorporates textual medical reports to enrich global contextual understanding. Extensive experiments conducted on both a self-constructed patient-level Placental dataset and two public datasets demonstrating that our method achieves state-of-the-art diagnostic performance. The code is available at https://github.com/ECNU-MultiDimLab/EmmPD. Zixuan Gao, Siyuan Yang 0001, Shulin Peng, Xiang Tao, Yan Wang 0033, Qingli Li |
ACM Multimedia | 9 |
| 2025 | Learning to Optimize Low-Latency Live Streaming from Expertise: An Offline Meta-Reinforcement Learning ApproachabstractIn low-latency live streaming (LLLS), the adaptive bitrate algorithm plays a critical role in optimizing encoding bitrates to meet the strict end-to-end latency requirement, which is mostly several seconds or less. Existing methods usually rely on a fixed preset latency target, limiting their generalization capabilities across the heterogeneous scenarios. To address this, we propose an offline meta-reinforcement learning-based bitrate decision framework for LLLS, which incorporates the expertise of multiple state-of-the-art LLLS algorithms from their collected experience and constructs a universal bitrate selection policy via the offline training. Specifically, we establish an RL-based policy that adaptively adjusts the network throughput measurements and selects the encoding bitrates for LLLS based on the network conditions. The policy’s performance is enhanced by improving the robustness of short-term throughput estimations. To leverage the expertise of current LLLS algorithms, we collect their decision-making trajectories, and directly train the policy on these offline trajectories with implicit Q-learning. In addition, a meta-RL paradigm is further adopted to learn a universal policy that performs uniformly well across the heterogeneous network conditions and varying target latencies. Experimental results demonstrate that, compared to the other baselines, the proposed method achieves at least a 5.8% performance gain in terms of the overall quality of experience (QoE), across the diverse target-latency scenarios in the real-world throughput traces. Yuhui Du, Nuowen Kan, Junni Zou, Wenrui Dai, Qingli Li, Hongkai Xiong |
VCIP | 8 |
| 2025 | WFANet-DDCL: Wavelet-Based Frequency Attention Network and Dual Domain Consistency Learning for 7T MRI Synthesis From 3T MRIabstractUltra-high field magnetic resonance imaging (MRI), such as 7-Tesla (7T) MRI, provides significantly enhanced tissue contrast and anatomical details compared to 3T MRI. However, 7T MRI scanners are more costly and less accessible in clinical settings than 3T scanners. In this paper, we propose a wavelet-based frequency attention network (WFANet) and a semi-supervised method named dual domain consistency learning (DDCL), and combine them to form a WFANet-DDCL framework for 7T MRI synthesis. WFANet leverages the frequency sensitivity of the proposed wavelet-based frequency attention encoder (WFAE) along with the large receptive field of dilated convolution. WFAE is proposed as an independent module to capture multi-scale frequency attention via the proposed wavelet-based frequency attention (WFA) mechanism. WFAE can be integrated into any backbone network as a plug-and-play component and improve network performance. To tackle the challenge of limited paired data for network training, DDCL is proposed to take advantage of both paired and unpaired data. Frequency domain perturbation is proposed and combined with Gaussian noise to regularize the supervised learning process in dual domains, better avoiding overfitting. Extensive experimental results demonstrate that WFANet-DDCL can achieve comparable performance to state-of-the-art supervised methods even using 66% of all paired data. Song Qiu, Mei Zhou, Weijie Le, Qingli Li, Yan Wang 0033 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | S4R: Separated Self-Supervised Spectral Regression for Hyperspectral Histopathology Image DiagnosisabstractHyperspectral images (HSIs) offer great potential for computational pathology. But, limited by the lack of adequate annotated data and the high spectral redundancy of HSIs, traditional supervised learning techniques are usually bottlenecked. To exploit the structural properties of HSIs and learn representations with good transferability, we propose Separated Self-Supervised Spectral Regression (S4R). Concretely, we find one spectral band can be represented by a linear combination of the remaining bands. Regressing the distribution of the linear coefficients learns the inherent properties of HSIs and pathological information about the tissue. Besides, reconstructing the missing band, especially the tissue boundaries makes the model learn pathology details that are critical to downstream tasks. Coupling these two pretext tasks makes the self-supervised model understand spectral structures of HSIs w.r.t. pathological semantics and spatial micro details. Furthermore, we design two brand-new architectures to avoid the interference of extraneous signal based on S4R: S4R-CLS and S4R-SEG for HSI classification and segmentation, respectively. Two downstream tasks are incorporated into a unified framework, which first encodes different bands from HSIs via a depthwise separable encoder, and then selectively aggregates band features to generate final predictions. In S4R-SEG, we propose to pick the best matching bands with the guidance of a classification paradigm. Extensive experiments show S4R performs much better than competitors on both tasks. Theoretical analysis and clinical discussion also indicate the great potential for further medical applications. The code and pre-trained checkpoints are available at https://github.com/DeepMed-Lab-ECNU/S4R. Yan Wang 0033, Xingran Xie, Benyan Zhang, Chunhua Zhou, Duowu Zou, Le Lu 0001, Qingli Li |
IEEE Trans. Image Process. | 8 |
| 2025 | Clinical Stage Prompt Induced Multi-Modal PrognosisabstractHistology analysis of the tumor micro-environment integrated with genomic assays is widely regarded as the cornerstone for cancer analysis and survival prediction. This paper jointly incorporates genomics and Whole Slide Images (WSIs), and focuses on addressing the primary challenges involved in multi-modality prognosis analysis: 1) the high-order relevance is difficult to be modeled from dimensional imbalanced gigapixel WSIs and tens of thousands of genetic sequences, and 2) the lack of medical expertise and clinical knowledge hampers the effectiveness of prognosis-oriented multi-modal fusion. Due to the nature of the prognosis task, statistical priors and clinical knowledge are essential factors to provide the likelihood of survival over time, which, however, has been under-studied. To this end, we propose a prognosis-oriented image-omics fusion framework, dubbed Clinical Stage Prompt induced Multimodal Prognosis (CiMP). Concretely, we leverage the capabilities of the advanced LLM to generate descriptions derived from structured clinical records and utilize the generated clinical staging prompts to inquire critical prognosis-related information from each modality intentionally. In addition, we propose a Group Multi-Head Self-Attention module to capture structured group-specific features within cohorts of genomic data. Experimental results on five TCGA datasets show the superiority of our proposed method, achieving state-of-the-art performance compared to previous multi-modal prognostic models. Furthermore, the clinical interpretability and discussion also highlight the immense potential for further medical applications. Our code will be released at https://github.com/DeepMed-Lab-ECNU/CiMP/. Xingran Xie, Qingli Li, Xinxing Li, Yan Wang 0033 |
IEEE Trans. Medical Imaging | 3 |
| 2025 | Debiasing Medical Knowledge for Prompting Universal Model in CT Image SegmentationabstractWith the assistance of large language models, which offer universal medical prior knowledge via text prompts, state-of-the-art Universal Models (UM) have demonstrated considerable potential in the field of medical image segmentation. Semantically detailed text prompts, on the one hand, indicate comprehensive knowledge; on the other hand, they bring biases that may not be applicable to specific cases involving heterogeneous organs or rare cancers. To this end, we propose a Debiased Universal Model (DUM) to consider instance-level context information and remove knowledge biases in text prompts from the causal perspective. We are the first to discover and mitigate the bias introduced by universal knowledge. Specifically, we propose to extract organ-level text prompts via language models and instance-level context prompts from the visual features of each image. We aim to highlight more on factual instance-level information and mitigate organ-level's knowledge bias. This process can be derived and theoretically supported by a causal graph, and instantiated by designing a standard UM (SUM) and a biased UM. The debiased output is finally obtained by subtracting the likelihood distribution output by biased UM from that of the SUM. Experiments on three large-scale multi-center external datasets and MSD internal tumor datasets show that our method enhances the model's generalization ability in handling diverse medical scenarios and reducing the potential biases, even with an improvement of 4.16% compared with popular universal model on the AbdomenAtlas dataset, showing the strong generalizability. The code is publicly available at https://github.com/DeepMed-Lab-ECNU/DUM. Boxiang Yun, Shitian Zhao, Qingli Li, Alex Chichung Kot, Yan Wang 0033 |
IEEE Trans. Medical Imaging | 3 |
| 2024 | A Semi-Supervised Approach with Error Reflection for Echocardiography SegmentationabstractSegmenting internal structure from echocardiography is essential for the diagnosis and treatment of various heart diseases. Semi-supervised learning shows its ability in alleviating annotations scarcity. While existing semi-supervised methods have been successful in image segmentation across various medical imaging modalities, few have attempted to design methods specifically addressing the challenges posed by the poor contrast, blurred edge details and noise of echocardiography. These characteristics pose challenges to the generation of high-quality pseudo-labels in semi-supervised segmentation based on Mean Teacher. Inspired by human reflection on erroneous practices, we devise an error reflection strategy for echocardiography semi-supervised segmentation architecture. The process triggers the model to reflect on inaccuracies in unlabeled image segmentation, thereby enhancing the robustness of pseudo-label generation. Specifically, the strategy is divided into two steps. The first step is called reconstruction reflection. The network is tasked with reconstructing authentic proxy images from the semantic masks of unlabeled images and their auxiliary sketches, while maximizing the structural similarity between the original inputs and the proxies. The second step is called guidance correction. Reconstruction error maps decouple unreliable segmentation regions. Then, reliable data that are more likely to occur near high-density areas are leveraged to guide the optimization of unreliable data potentially located around decision boundaries. Additionally, we introduce an effective data augmentation strategy, termed as multi-scale mixing up strategy, to minimize the empirical distribution gap between labeled and unlabeled images and perceive diverse scales of cardiac anatomical structures. Extensive experiments on a public echocardiography dataset CAMUS, and a private clinical echocardiography dataset demonstrate the competitiveness of the proposed method. Xiaoxiang Han 0001, Yiman Liu, Jiang Shang, Qingli Li, Menghan Hu, Qi Zhang 0003, Yan Wang 0033 |
BIBM | 4 |
| 2024 | AdaRevD: Adaptive Patch Exiting Reversible Decoder Pushes the Limit of Image DeblurringabstractDespite the recent progress in enhancing the efficacy of image deblurring, the limited decoding capability constrains the upper limit of State-Of- The-Art (SOTA) methods. This paper proposes a pioneering work, Adaptive Patch Ex-iting Reversible Decoder (AdaRevD), to explore their in-sufficient decoding capability. By inheriting the weights of the well-trained encoder, we refactor a reversible de-coder which scales up the single-decoder training to multi-decoder training while remaining GPU memory-friendly. Meanwhile, we show that our reversible structure gradually disentangles high-level degradation degree and low-level blur pattern (residual of the blur image and its sharp counterpart) from compact degradation representation. Besides, due to the spatially-variant motion blur kernels, different blur patches have various deblurring difficulties. We further introduce a classifier to learn the degradation degree of image patches, enabling them to exit at different sub-decoders for speedup. Experiments show that our AdaRevD pushes the limit of image deblurring, e.g., achieving 34.60 dB in PSNR on GoPro dataset. Xintian Mao, Qingli Li, Yan Wang 0033 |
CVPR | 2 |
| 2024 | Spatially-Variant Degradation Model for Dataset-Free Super-Resolution
Shaojie Guo, Haofei Song, Qingli Li, Yan Wang 0033 |
ECCV (25) | 3 |
| 2024 | Medical Image Classification Attack Based on Texture Manipulation
Yunrui Gu, Cong Kong, Zhao-Xia Yin, Yan Wang 0033, Qingli Li |
ICPR (12) | 5 |
| 2024 | OD-DETR: Online Distillation for Stabilizing Training of Detection Transformer
Shengjian Wu, Li Sun 0012, Qingli Li |
IJCAI | 3 |
| 2024 | Spatial-Temporal Traffic Prediction Model Based on Adaptive Graphs Fusion and Dual-Graph Collaborative ConvolutionabstractTraffic flow prediction is crucial for intelligent transportation systems (ITS). Traditional graph convolutional networks (GCNs) have limitations in handling road network data. These GCNs can only handle binary relationships between nodes and cannot effectively capture the dynamic and nonlinear spatial dependencies among multiple nodes in the road network. This paper proposes an innovative spatial-temporal traffic flow prediction model to address this challenge. Firstly, by introducing novel graphs fusion and hypergraph encoding module, a fused graph and its dual hypergraph are constructed to provide richer structural information for the model. Then, the model can learn complex relationships among multiple nodes through the collaborative convolution of GCN and hypergraph convolutional network (HGCN). To comprehensively capture the dynamic nature of traffic flow, we utilize a variant of Transformer combined with time encoding information to capture the periodicity in the data, thereby enhancing the model’s ability to recognize periodic traffic flow patterns. Comprehensive experimental results on two publicly accessible real-world traffic datasets demonstrate the superiority of our proposed model over state-of-the-art traffic prediction models. Song Qiu, Li Sun 0012, Dingding Han, Qingli Li, Mingsong Chen 0001 |
IJCNN | 5 |
| 2024 | DSA-SCGC: A Dual Self-Attention Mechanism based on Space-Channel Grouped Compression for Vehicle Re-IdentificationabstractVehicle re-identification (re-ID) has attracted significant attention within the computer vision community due to its wide-ranging applications in intelligent transportation systems and law enforcement. Nevertheless, this field faces considerable challenges owing to the high inter-class similarity and the large intra-class difference among vehicles. To address these challenges, this paper proposes a novel network incorporating a dual self-attention mechanism based on a space-channel grouped compression operation (DSA-SCGC). This innovative approach combines channel and spatial self-attention mechanisms to selectively enhance pivotal channel features and spatial local details while minimizing attention toward backgrounds and occlusions commonly encountered in real-world scenarios. Moreover, to address the issue of spatial information loss in channel attention, we propose a space-channel grouped compression (SCGC) operation that effectively compresses spatial information into channels, thereby significantly preserving spatial information. Comprehensive experiments conducted on the VeRi-776 and VehicleID datasets validate the superiority of our proposed DSA-SCGC model over the existing state-of-the-art vehicle re-identification methods. Yuejun Jiao, Song Qiu, Li Sun 0012, Dingding Han, Qingli Li, Mingsong Chen 0001 |
IJCNN | 5 |
| 2024 | Multi-stage Multi-granularity Focus-Tuned Learning Paradigm for Medical HSI Segmentation
Haichuan Dong, Runjie Zhou, Boxiang Yun, Benyan Zhang, Qingli Li, Yan Wang 0033 |
MICCAI (8) | 6 |
| 2024 | Prompting Whole Slide Image Based Genetic Biomarker Prediction
Boxiang Yun, Xingran Xie, Qingli Li, Xinxing Li, Yan Wang 0033 |
MICCAI (4) | 4 |
| 2024 | LoFormer: Local Frequency Transformer for Image Deblurring
Xintian Mao, Jiansheng Wang, Xingran Xie, Qingli Li, Yan Wang 0033 |
ACM Multimedia | 4 |
| 2024 | GeNSeg-Net: A General Segmentation Framework for Any Nucleus in Immunohistochemistry Images
Haofei Song, Jiansheng Wang, Yan Wang 0033, Qingli Li |
ACM Multimedia | 6 |
| 2024 | RefineStyle: Dynamic Convolution Refinement for StyleGAN
Siwei Xia, Xueqi Hu, Li Sun 0012, Qingli Li |
PRCV (9) | 4 |
| 2024 | KD loss: Enhancing discriminability of features with kernel trick for object detection in VHR remote sensing images
Xi Chen 0004, Liyue Li, Qingli Li, Honggang Qi, Ying Wen 0003, Guitao Cao, Philip L. H. Yu |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | Scale-pyramid dynamic atrous convolution for pixel-level labeling
Xi Chen 0004, Yong Wang 0032, Qingli Li, Honggang Qi, Robert Laganière |
Expert Syst. Appl. | 6 |
| 2024 | SpecTr: Spectral Transformer for Microscopic Hyperspectral Pathology Image SegmentationabstractHyperspectral imaging (HSI) unlocks the huge potential to a wide variety of applications relying on high-precision pathology image segmentation, such as computational pathology. It can acquire biochemical properties even invisible to naked eyes from histological specimens. Since 1) spectra contain discriminative and continuous patterns for differentiating tissues/cells, and 2) the discriminability of spectra relies on both fine-grained relations in the high-resolution spectrum and coarse relations in the low-resolution spectrum, the key to achieving high-precision hyperspectral pathology image segmentation is to felicitously model the intra- and inter-scale context especially for spectra. In this paper, we propose a spectral transformer (SpecTr) for hyperspectral pathology image segmentation, which first captures global context for intra-scale spectral features, and subsequently extract coarse and fine-grained discriminative spectral information from inter-scale features, respectively. To learn intra-scale spectral context, we propose a Spectral Attentive Module (SAM). Unlike the existing Transformer model that is designed for modalities such as natural images, our proposed SAM is efficient in capturing sparse and pivotal spectral context while avoiding the heterogeneous underlying distributions and noises of different bands. Besides, to reduce the computational complexity of the HSI segmentation model, we further propose a global-local attention module to effectively learn a condensed spectral feature. Experiments show that HSIs can become a more powerful image modality for understanding microscopic pathology images than RGB images, and the proposed SpecTr outperforms other competing methods for hyperspectral pathology image segmentation, with an improvement of 3% compared with the popular 3D-nnUNet and other transformer-based methods. Our code is available at https://github.com/DeepMed-Lab-ECNU/SpecTr. Boxiang Yun, Bai Ying Lei, Jieneng Chen, Song Qiu, Wei Shen 0002, Qingli Li, Yan Wang 0033 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | I³Net: Inter-Intra-Slice Interpolation Network for Medical Slice SynthesisabstractMedical imaging is limited by acquisition time and scanning equipment. CT and MR volumes, reconstructed with thicker slices, are anisotropic with high in-plane resolution and low through-plane resolution. We reveal an intriguing phenomenon that due to the mentioned nature of data, performing slice-wise interpolation from the axial view can yield greater benefits than performing super-resolution from other views. Based on this observation, we propose an Inter-Intra-slice Interpolation Network ( [Formula: see text]Net), which fully explores information from high in-plane resolution and compensates for low through-plane resolution. The through-plane branch supplements the limited information contained in low through-plane resolution from high in-plane resolution and enables continual and diverse feature learning. In-plane branch transforms features to the frequency domain and enforces an equal learning opportunity for all frequency bands in a global context learning paradigm. We further propose a cross-view block to take advantage of the information from all three views online. Extensive experiments on two public datasets demonstrate the effectiveness of [Formula: see text]Net, and noticeably outperforms state-of-the-art super-resolution, video frame interpolation and slice interpolation methods by a large margin. We achieve 43.90dB in PSNR, with at least 1.14dB improvement under the upscale factor of ×2 on MSD dataset with faster inference. Code is available at https://github.com/DeepMed-Lab-ECNU/Medical-Image-Reconstruction. Haofei Song, Xintian Mao, Qingli Li, Yan Wang 0033 |
IEEE Trans. Medical Imaging | 4 |
| 2024 | Intelligent diagnosis of atrial septal defect in children using echocardiography with deep learningabstractAtrial septal defect (ASD) is one of the most common congenital heart diseases. The diagnosis of ASD via transthoracic echocardiography is subjective and time-consuming. The objective of this study was to evaluate the feasibility and accuracy of automatic detection of ASD in children based on color Doppler echocardiographic static images using end-to-end convolutional neural networks. The proposed depthwise separable convolution model identifies ASDs with static color Doppler images in a standard view. Among the standard views, we selected two echocardiographic views, i.e., the subcostal sagittal view of the atrium septum and the low parasternal four-chamber view. The developed ASD detection system was validated using a training set consisting of 396 echocardiographic images corresponding to 198 cases. Additionally, an independent test dataset of 112 images corresponding to 56 cases was used, including 101 cases with ASDs and 153 cases with normal hearts. The average area under the receiver operating characteristic curve, recall, precision, specificity, F1-score, and accuracy of the proposed ASD detection model were 91.99, 80.00, 82.22, 87.50, 79.57, and 83.04, respectively. The proposed model can accurately and automatically identify ASD, providing a strong foundation for the intelligent diagnosis of congenital heart diseases. Yiman Liu, Size Hou, Xiaoxiang Han 0001, Tongtong Liang, Menghan Hu, Qingli Li |
Virtual Real. Intell. Hardw. | 9 |
| 2023 | CLIP-ReID: Exploiting Vision-Language Model for Image Re-identification without Concrete Text LabelsabstractPre-trained vision-language models like CLIP have recently shown superior performances on various downstream tasks, including image classification and segmentation. However, in fine-grained image re-identification (ReID), the labels are indexes, lacking concrete text descriptions. Therefore, it remains to be determined how such models could be applied to these tasks. This paper first finds out that simply fine-tuning the visual model initialized by the image encoder in CLIP, has already obtained competitive performances in various ReID tasks. Then we propose a two-stage strategy to facilitate a better visual representation. The key idea is to fully exploit the cross-modal description ability in CLIP through a set of learnable text tokens for each ID and give them to the text encoder to form ambiguous descriptions. In the first training stage, image and text encoders from CLIP keep fixed, and only the text tokens are optimized from scratch by the contrastive loss computed within a batch. In the second stage, the ID-specific text tokens and their encoder become static, providing constraints for fine-tuning the image encoder. With the help of the designed loss in the downstream task, the image encoder is able to represent data as vectors in the feature embedding accurately. The effectiveness of the proposed strategy is validated on several datasets for the person or vehicle ReID tasks. Code is available at https://github.com/Syliz517/CLIP-ReID. Li Sun 0012, Qingli Li |
AAAI | 3 |
| 2023 | Intriguing Findings of Frequency Selection for Image DeblurringabstractBlur was naturally analyzed in the frequency domain, by estimating the latent sharp image and the blur kernel given a blurry image. Recent progress on image deblurring always designs end-to-end architectures and aims at learning the difference between blurry and sharp image pairs from pixel-level, which inevitably overlooks the importance of blur kernels. This paper reveals an intriguing phenomenon that simply applying ReLU operation on the frequency domain of a blur image followed by inverse Fourier transform, i.e., frequency selection, provides faithful information about the blur pattern (e.g., the blur direction and blur level, implicitly shows the kernel pattern). Based on this observation, we attempt to leverage kernel-level information for image deblurring networks by inserting Fourier transform, ReLU operation, and inverse Fourier transform to the standard ResBlock. 1 × 1 convolution is further added to let the network modulate flexible thresholds for frequency selection. We term our newly built block as Res FFT-ReLU Block, which takes advantages of both kernel-level and pixel-level features via learning frequency-spatial dual-domain representations. Extensive experiments are conducted to acquire a thorough analysis on the insights of the method. Moreover, after plugging the proposed block into NAFNet, we can achieve 33.85 dB in PSNR on GoPro dataset. Our method noticeably improves backbone architectures without introducing many parameters, while maintaining low computational complexity. Code is available at https://github.com/DeepMed-Lab/DeepRFT-AAAI2023. Xintian Mao, Fengze Liu, Qingli Li, Wei Shen 0002, Yan Wang 0033 |
AAAI | 4 |
| 2023 | Deep Adversarial Network Based Stain Unmixing for Brightfield Multiplex Immunohistochemistry ImagesabstractMultiplex immunohistochemistry (IHC) makes it possible to simultaneously label multiple protein biomarkers with different colored stains in a tissue section. Unmixing the multiplex IHC image provides an efficient way to obtain the rich diagnostic information each biomarker contains. However, due to the limitation of three-channel RGB images taken by a CCD color camera, it is challenging to unmix brightfield multiplex IHC images with more than three stains. The main technical challenge is that the unmixing is inherent underdeterminate, leading to the possibility of multiple solutions. In this paper, we propose a novel unmixing method using generative adversarial networks (GANs) for brightfield multiplex IHC images containing more than three stains. Our method takes advantage of slides stained only with individual biomarker to address the intriguing task without any human annotation. We propose to employ adversarial training to automatically learn the optimal unmixing, without relying on any inadequately designed priori. To the best of our knowledge, the method achieves state-of-the-art (SOTA) results in terms of unmixing quality, speed, and practicality, as evidenced by both pathologists’ visual comparisons and quantitative experiments. Mingxue Gu, Qingli Li |
BIBM | 4 |
| 2023 | Bidirectional Copy-Paste for Semi-Supervised Medical Image SegmentationabstractIn semi-supervised medical image segmentation, there exist empirical mismatch problems between labeled and un-labeled data distribution. The knowledge learned from the labeled data may be largely discarded if treating labeled and unlabeled data separately or in an inconsistent manner. We propose a straightforward method for alleviating the problem-copy-pasting labeled and unlabeled data bidirectionally, in a simple Mean Teacher architecture. The method encourages unlabeled data to learn comprehensive common semantics from the labeled data in both inward and outward directions. More importantly, the consistent learning procedure for labeled and unlabeled data can largely reduce the empirical distribution gap. In detail, we copy-paste a random crop from a labeled image (foreground) onto an unlabeled image (background) and an unlabeled image (foreground) onto a labeled image (background), respectively. The two mixed images are fed into a Student network and supervised by the mixed supervisory signals of pseudo-labels and ground-truth. We reveal that the simple mechanism of copy-pasting bidirectionally between labeled and unlabeled data is good enough and the experiments show solid gains (e.g., over 21% Dice improvement on ACDC dataset with 5% labeled data) compared with other state-of-the-arts on various semi-supervised medical image segmentation datasets. Code is avaiable at https://github.com/DeepMed-Lab-ECNU/BCP. Yunhao Bai, Duowen Chen 0002, Qingli Li, Wei Shen 0002, Yan Wang 0033 |
CVPR | 3 |
| 2023 | MagicNet: Semi-Supervised Multi-Organ Segmentation via Magic-Cube Partition and RecoveryabstractWe propose a novel teacher-student model for semi-supervised multi-organ segmentation. In teacher-student model, data augmentation is usually adopted on unlabeled data to regularize the consistent training between teacher and student. We start from a key perspective that fixed relative locationsand variable sizes of different organs can provide distribution information where a multi-organ CT scan is drawn. Thus, we treat the prior anatomy as a strong tool to guide the data augmentation and reduce the mismatch between labeled and unlabeled images for semi-supervised learning. More specifically, we propose a data augmentation strategy based on partition-and-recovery N3cubes cross-and within-labeled and unlabeled images. Our strategy encourages unlabeled images to learn organ semantics in relative locations from the labeled images (cross-branch) and enhances the learning ability for small organs (within-branch). For within-branch, we further propose to refine the quality of pseudo labels by blending the learned representations from small cubes to incorporate local attributes. Our method is termed as MagicNet, since it treats the CT volume as a magic-cube and N3-cube partition-and-recovery process matches with the rule of playing a magic-cube. Extensive experiments on two public CT multi-organ datasets demonstrate the effectiveness of MagicNet, and noticeably outperforms state-of-the-art semi-supervised medical image segmentation approaches, with + 7% DSC improvement on MACT dataset with 10% labeled images. Code is avaiable at https://github.com/DeepMed-Lab-ECNU/MagicNet. Duowen Chen 0002, Yunhao Bai, Wei Shen 0002, Qingli Li, Lequan Yu, Yan Wang 0033 |
CVPR | 4 |
| 2023 | Class Balanced Adaptive Pseudo Labeling for Federated Semi-Supervised LearningabstractThis paper focuses on federated semi-supervised learning (FSSL), assuming that few clients have fully labeled data (labeled clients) and the training datasets in other clients are fully unlabeled (unlabeled clients). Existing methods attempt to deal with the challenges caused by not independent and identically distributed data (Non-IID) setting. Though methods such as sub-consensus models have been proposed, they usually adopt standard pseudo labeling or consistency regularization on unlabeled clients which can be easily influenced by imbalanced class distribution. Thus, problems in FSSL are still yet to be solved. To seek for a fundamental solution to this problem, we present Class Balanced Adaptive Pseudo Labeling (CBAFed), to study FSSL from the perspective of pseudo labeling. In CBAFed, the first key element is a fixed pseudo labeling strategy to handle the catastrophic forgetting problem, where we keep a fixed set by letting pass information of unlabeled data at the beginning of the unlabeled client training in each communication round. The second key element is that we design class balanced adaptive thresholds via considering the empirical distribution of all training data in local clients, to encourage a balanced training process. To make the model reach a better optimum, we further propose a residual weight connection in local supervised training and global model aggregation. Extensive experiments on five datasets demonstrate the superiority of CBAFed. Code will be available at https://github.com/minglllli/CBAFed. Qingli Li, Yan Wang 0033 |
CVPR | 2 |
| 2023 | RecursiveDet: End-to-End Region-based Recursive Object DetectionabstractEnd-to-end region-based object detectors like Sparse R-CNN usually have multiple cascade bounding box decoding stages, which refine the current predictions according to their previous results. Model parameters within each stage are independent, evolving a huge cost. In this paper, we find the general setting of decoding stages is actually redundant. By simply sharing parameters and making a recursive decoder, the detector already obtains a significant improvement. The recursive decoder can be further enhanced by positional encoding (PE) of the proposal box, which makes it aware of the exact locations and sizes of input bounding boxes, thus becoming adaptive to proposals from different stages during the recursion. Moreover, we also design centerness-based PE to distinguish the RoI feature element and dynamic convolution kernels at different positions within the bounding box. To validate the effectiveness of the proposed method, we conduct intensive ablations and build the full model on three recent mainstream region-based detectors. The RecusiveDet is able to achieve obvious performance boosts with even fewer model parameters and slightly increased computation cost. Codes are available at https://github.com/bravezzzzzz/RecursiveDet. Li Sun 0012, Qingli Li |
ICCV | 3 |
| 2023 | A Generative Data Augmentation Trained by Low-quality Annotations for Cholangiocarcinoma Hyperspectral Image SegmentationabstractMicroscopic hyperspectral imaging technology combined with deep learning method emerges medical field recently as a multiplexed imaging technology. With the semantic segmentation of hyperspectral histopathological image of pathological tissue, doctors can quickly locate suspicious areas, diagnose and arrange treatment accurately and rapidly, reducing the workload of them. Cholangiocarcinoma is a rare and devastating disease with few hyperspectral histopathological data. Moreover, achieving high-quality annotations of hyperspectral histopathological image is challenging and costs time for pathologists, so generally, rough labels are annotated, but directly using the low-quality labels will reduce the performance of segmentation networks. So how to fully utilize few high-quality annotations and dozens of low-quality labels to enhance the segmentation performance of cholangiocarcinoma hyperspectral image remains to be resolved. In this paper, we proposed a two-stage hyperspectral segmentation deep learning framework based on Labels-to-Photo translation and Swin-Spec Transformer(L2P-SST). In stage-I, the OASIS generative network and the Swin-Spec Transformer discriminative network are used for adversarial training, and a spectral perceptual loss function is proposed to generate highquality hyperspectral images; in stage-II, parameters of the generative network is fixed and the generated hyperspectral images are used as data augmentation in the training of Swin-Spec Transformer segmentation network. The proposed framework achieved 76.16% mIoU(mean Intersection over Union), 85.80% mDice(mean Dice), 90.96% Accuracy and 71.65% Kappa coefficient in the semantic segmentation task of the Multidimensional Choledoch Database. Compared with other methods, the results demonstrate our framework provides a competitive segmentation performance. Kaijie Dai, Zehao Zhou, Song Qiu, Yan Wang 0033, Mei Zhou, Mingshuai Li, Qingli Li |
IJCNN | 7 |
| 2023 | CCH-YOLOX: Improved YOLOX for Challenging Vehicle Detection from UAV ImagesabstractIn this paper, we focus on vehicle detection algorithms from UAV images in intelligent transportation applications. We propose an improved YOLOX model called CCHYOLOX for the problems of dense distribution and drastic scale changes of vehicle targets from the UAV viewpoint. Firstly, we construct a novel plug-and-play module, called Correlation Extraction and Feature Fusion (CEFF), to process the multi-layer adaptive fusion of pyramidal features. It enables the adaptive fusion of features in adjacent layers by adding spatial-awareness and extracting global channels correlation of adjacent layers features to mitigate the effect of target size diversity. Then, we design a special cascade strategy, including a feature alignment module named DConv, for single-stage and anchor-free detectors by considering the feature offset problem to produce more accurate detection bounding boxes. The cascade strategy allows the model to improve performance with little increase in computational cost. Finally, a high-resolution branch is designed for the small-target detection task, which greatly improves the detection accuracy of the model. Experiments on the challenging Visdrone-vehicle and Drone Vehicle datasets show that the proposed method effectively tackle the detection accuracy decrease caused by the above problems. CCH-YOLOX achieves 43.2% mAP and 57.3% mAP on the above two datasets, respectively, which is about 3%-4% higher than the strong baseline model (YOLOX) and exceeds many current popular models. The code is available at https://github.com/lz06787/CCH-YOLOX Song Qiu, Mingsong Chen 0001, Dingding Han, Tiantian Qi, Qingli Li, Yue Lu 0001 |
IJCNN | 6 |
| 2023 | Gene-Induced Multimodal Pre-training for Image-Omic Classification
Xingran Xie, Renjie Wan, Qingli Li, Yan Wang 0033 |
MICCAI (6) | 4 |
| 2023 | Deep Mutual Distillation for Semi-supervised Medical Image Segmentation
Yushan Xie, Yuejia Yin, Qingli Li, Yan Wang 0033 |
MICCAI (3) | 3 |
| 2023 | Factor Space and Spectrum for Medical Hyperspectral Image Segmentation
Boxiang Yun, Qingli Li, Lubov B. Mitrofanova, Chunhua Zhou, Yan Wang 0033 |
MICCAI (4) | 2 |
| 2023 | Exploring Hyperspectral Histopathology Image Segmentation from a Deformable PerspectiveabstractHyperspectral images (HSIs) offer great potential for computational pathology. However, limited by the spectral redundancy and the lack of spectral prior in popular 2D networks, previous HSI based techniques do not perform well. To address these problems, we propose to segment HSIs from a deformable perspective, which processes different spectral bands independently and fuses spatiospectral features of interest via deformable attention mechanisms. In addition, we propose Deformable Self-Supervised Spectral Regression (DF-S3R), which introduces two self-supervised pre-text tasks based on the low rank prior of HSIs enabling the network learning with spectrum-related features. During pre-training, DF-S3R learns both spectral structures and spatial morphology, and the jointly pre-trained architectures help alleviate the transfer risk to downstream fine-tuning. Compared to previous works, experiments show that our deformable architecture and pre-training method perform much better than other competitive methods on pathological semantic segmentation tasks, and the visualizations indicate that our method can trace the critical spectral characteristics from subtle spectral disparities. Code will be released at https://github.com/Ayakax/DFS3R. Xingran Xie, Boxiang Yun, Qingli Li, Yan Wang 0033 |
ACM Multimedia | 4 |
| 2023 | Uni-Dual: A Generic Unified Dual-Task Medical Self-Supervised Learning FrameworkabstractRGB images and medical hyperspectral images (MHSIs) are two widely-used modalities in computational pathology. The former is cheap, easy and fast to obtain while lacking pathological information such as physiochemical state. The latter is an emerging modality which captures electromagnetic radiation matter interaction but suffers from problems such as high time cost and low spatial resolution. In this paper, we bring forward a unified dual-task multi-modality self-supervised learning (SSL) framework, called Uni-Dual, which takes the most use of both paired and unpaired RGB-MHSIs. Concretely, we design a unified SSL paradigm for RGB images and MHSIs. Two tasks are proposed: (1) a discrimination learning task which learns high-level semantics via mining the cross-correlation across unpaired RGB-MHSIs, (2) a reconstruction learning task which models low-level stochastic variations via furthering the interaction across RGB-MHSI pairs. Our Uni-Dual enjoys the following benefits: (1) A unified model which can be easily transferred to different downstream tasks on various modality combinations. (2) We consider multi-constituent and structured information learning from MHSIs and RGB images for low-cost high-precision clinical purposes. Experiments conducted on various downstream tasks with different modalities show the proposed Uni-Dual substantially outperforms other competitive SSL methods. Boxiang Yun, Xingran Xie, Qingli Li, Yan Wang 0033 |
ACM Multimedia | 3 |
| 2023 | Dense-scale dynamic network with filter-varying atrous convolution for semantic segmentation
Xi Chen 0004, Robert Laganière, Qingli Li, Honggang Qi, Yong Wang 0032 |
Appl. Intell. | 5 |
| 2023 | An online continual object detector on VHR remote sensing images with class imbalance
Xi Chen 0004, Honggang Qi, Qingli Li, Laiwen Zheng, Yongqiang Deng |
Eng. Appl. Artif. Intell. | 5 |
| 2023 | MIANet: Multi-level temporal information aggregation in mixed-periodicity time series forecasting tasks
Xi Chen 0004, Chen Wang 0007, Yong Wang 0032, Honggang Qi, Gongjian Zhou, Qingli Li |
Eng. Appl. Artif. Intell. | 8 |
| 2023 | Coupled Global-Local object detection for large VHR aerial images
Xi Chen 0004, Chaojie Wang 0009, Qingli Li, Honggang Qi, Yong Wang 0032 |
Knowl. Based Syst. | 5 |
| 2023 | A multi-grained unsupervised domain adaptation approach for semantic segmentation
Tai Ma, Yue Lu 0001, Qingli Li, Lianghua He, Ying Wen 0003 |
Pattern Recognit. | 4 |
| 2023 | Angel's Girl for Blind Painters: An Efficient Painting Navigation System Validated by Multimodal Evaluation ApproachabstractFor people who ardently love painting but unfortunately have visual impairments, holding a paintbrush to create a work is a very difficult task. People in this special group are eager to pick up the paintbrush, like Leonardo da Vinci, to create and make full use of their own talents. Therefore, to maximally bridge this gap, we propose a painting navigation system called “Angle’s Eyes” to assist blind people in artistic creation. The proposed system is composed of cognitive system and guidance system. The system adopts drawing board positioning based on QR code, brush navigation based on target detection and bush real-time positioning. Meanwhile, we design a simple yet efficient position information coding rule to remind the user of the current brush tip position. In addition, we design a criterion to efficiently judge whether the brush reaches the target or not. The numerous experiments are conducted to optimize and test the performance of the system. The results of real-world scenario experiments demonstrate that the developed system has great potential to help blind people with painting. This work also demonstrates that it is practicable for the blind people to feel the world through the brush in their hands. In the future, we plan to deploy “Angle’s Eyes” on the phone to make it more portable. The demo video of the proposed painting navigation system is available athttps://doi.org/10.6084/m9.figshare.9760004.v1. Menghan Hu, Qingli Li, Guangtao Zhai, Simon X. Yang, Xiao-Ping Zhang 0002, Xiaokang Yang 0001 |
IEEE Trans. Multim. | 4 |
| 2022 | IoU-Enhanced Attention for End-to-End Task Specific Object Detection
Shengjian Wu, Li Sun 0012, Qingli Li |
ACCV (5) | 4 |
| 2022 | Style Transformer for Image Inversion and EditingabstractExisting GAN inversion methods fail to provide latent codes for reliable reconstruction and flexible editing simultaneously. This paper presents a transformer-based image inversion and editing model for pretrained StyleGAN which is not only with less distortions, but also of high quality and flexibility for editing. The proposed model employs a CNN encoder to provide multi-scale image features as keys and values. Meanwhile it regards the style code to be determined for different layers of the generator as queries. It first initializes query tokens as learnable parameters and maps them into W+ space. Then the multi-stage alternate self-and cross-attention are utilized, updating queries with the purpose of inverting the input by the generator. Moreover, based on the inverted code, we investigate the reference-and label-based attribute editing through a pretrained latent classifier, and achieve flexible image-to-image translation with high quality results. Extensive experiments are carried out, showing better performances on both inversion and editing tasks within StyleGAN. Codes are available at https://github.com/sapphire497/style-transformer. Xueqi Hu, Qiusheng Huang, Zhengyi Shi, Changxin Gao, Li Sun 0012, Qingli Li |
CVPR | 7 |
| 2022 | QS-Attn: Query-Selected Attention for Contrastive Learning in I2I TranslationabstractUnpaired image-to-image (I2I) translation often requires to maximize the mutual information between the source and the translated images across different domains, which is critical for the generator to keep the source content and prevent it from unnecessary modifications. The self-supervised contrastive learning has already been successfully applied in the I2I. By constraining features from the same location to be closer than those from different ones, it implicitly ensures the result to take content from the source. However, previous work uses the features from random locations to impose the constraint, which may not be appropriate since some locations contain less information of source domain. Moreover, the feature itself does not reflect the relation with others. This paper deals with these problems by intentionally selecting significant anchor points for contrastive learning. We design a query-selected attention (QS-Attn) module, which compares feature distances in the source domain, giving an attention matrix with a probability distribution in each row. Then we select queries according to their measurement of significance, computed from the distribution. The selected ones are regarded as anchors for contrastive loss. At the same time, the reduced attention matrix is employed to route features in both domains, so that source relations maintain in the synthesis. We validate our proposed method in three different I2I datasets, showing that it increases the image quality with-out adding learnable parameters. Codes are available at https://github.com/sapphire497/query-selected-attention. Xueqi Hu, Xinyue Zhou, Qiusheng Huang, Zhengyi Shi, Li Sun 0012, Qingli Li |
CVPR | 6 |
| 2022 | Cross Attention Based Style Distribution for Controllable Person Image Synthesis
Xinyue Zhou, Mingyu Yin, Li Sun 0012, Changxin Gao, Qingli Li |
ECCV (15) | 6 |
| 2022 | S3R: Self-supervised Spectral Regression for Hyperspectral Histopathology Image Classification
Xingran Xie, Yan Wang 0033, Qingli Li |
MICCAI (2) | 3 |
| 2022 | Cross-Stage Class-Specific Attention for Image Semantic Segmentation
Zhengyi Shi, Li Sun 0012, Qingli Li |
PRCV (4) | 3 |
| 2022 | DSAM-GN: Graph Network Based on Dynamic Similarity Adjacency Matrices for Vehicle Re-identification
Yuejun Jiao, Song Qiu, Mingsong Chen 0001, Dingding Han, Qingli Li, Yue Lu 0001 |
PRICAI (1) | 5 |
| 2022 | Superdense-scale network for semantic segmentation
Xi Chen 0004, Honggang Qi, Qingli Li, Laiwen Zheng |
Neurocomputing | 5 |
| 2022 | Cascaded Multiscale Structure With Self-Smoothing Atrous Convolution for Semantic SegmentationabstractConvolutional neural networks (CNNs) have attracted great attention in the semantic segmentation of very-high-resolution (VHR) images of urban areas. However, large-scale variation of objects in the urban areas often makes it difficult to achieve good segmentation accuracy. Atrous convolution and atrous spatial pyramid pooling composed of atrous convolution can alleviate this problem by exploring multiscale contextual information. Unfortunately, atrous convolution causes gridding artifacts, where actual receptive fields are separated unit sets and fail to cover all the receptive fields. To address this problem, in this article, we first propose a self-smoothing atrous convolution (SS-AConv) that intrinsically improves atrous convolution, unlike existing methods. SS-AConv enhances sampling rates with low computational costs by adding several key parameters in the spatial dimension of its filter. Then, in the backbone, the Xception network, all max pooling operations are replaced with SS-AConvs for a large and effective receptive field. Moreover, to extract diverse features at multiple scales, we propose the SS-AConv cascaded multiscale structure (SCMS) by integrating SS-AConvs with different rates and the residual correction scheme (RCS) into a cascaded spatial pyramid. Finally, to extract diverse features at dense multiple scales, the SS-AConv convolutional network (SS-ACNet) is constructed by integrating SCMS into both the encoder and decoder layers of the modified Xception network. Extensive experimental results show that SS-ACNet outperforms some state-of-the-art methods on four open challenge datasets: the ISPRS Vaihingen and Potsdam, Cityscapes, and PASCAL VOC 2012. Xi Chen 0004, Zhen Han 0007, Qingli Li |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2021 | ID-Unet: Iterative Soft and Hard Deformation for View SynthesisabstractView synthesis is usually done by an autoencoder, in which the encoder maps a source view image into a latent content code, and the decoder transforms it into a target view image according to the condition. However, the source contents are often not well kept in this setting, which leads to unnecessary changes during the view translation. Al-though adding skipped connections, like Unet, alleviates the problem, but it often causes the failure on the view conformity. This paper proposes a new architecture by performing the source-to-target deformation in an iterative way. Instead of simply incorporating the features from multiple layers of the encoder, we design soft and hard deformation modules, which warp the encoder features to the target view at different resolutions, and give results to the decoder to complement the details. Particularly, the current warping flow is not only used to align the feature of the same resolution, but also as an approximation to coarsely deform the high resolution feature. Then the residual flow is estimated and applied in the high resolution, so that the deformation is built up in the coarse-to-fine fashion. To better constrain the model, we synthesize a rough target view image based on the intermediate flows and their warped features. The extensive ablation studies and the final results on two different data sets show the effectiveness of the proposed model. https://github.com/MingyuY/Iterative-view-synthesis Mingyu Yin, Li Sun 0012, Qingli Li |
CVPR | 3 |
| 2021 | Identification of Deep Breath While Moving Forward Based on Multiple Body Regions and Graph Signal AnalysisabstractThis paper presents an unobtrusive solution that can automatically identify deep breath when a person is walking past the global depth camera. Existing non-contact breath assessments achieve satisfactory results under restricted conditions when human body stays relatively still. When someone moves forward, the breath signals detected by depth camera are hidden within signals of trunk displacement and deformation, and the signal length is short due to the short stay time, posing great challenges for us to establish models. To over-come these challenges, multiple region of interests (ROIs) based signal extraction and selection method is proposed to automatically obtain the signal informative to breath from depth video. Subsequently, graph signal analysis (GSA) is adopted as a spatial-temporal filter to wipe the components unrelated to breath. Finally, a classifier for identifying deep breath is established based on the selected breath-informative signal. In validation experiments, the proposed approach outperforms the comparative methods with the accuracy, precision, recall and F1 of 75.5%, 76.2%, 75.0% and 75.2%, respectively. This system can be extended to public places to provide timely and ubiquitous help for those who may have or are going through physical or mental trouble. Yunlu Wang, Cheng Yang 0003, Menghan Hu, Jian Zhang 0060, Qingli Li, Guangtao Zhai, Xiao-Ping Zhang 0002 |
ICASSP | 5 |
| 2021 | Progressive Multi-Stage Feature Mix for Person Re-IdentificationabstractImage features from a small local region often give strong evidence in person re-identification task. However, CNN suffers from paying too much attention on the most salient local areas, thus ignoring other discriminative clues, e.g., hair, shoes or logos on clothes. In this work, we propose a Progressive Multi-stage feature Mix network (PMM), which enables the model to find out the more precise and diverse features in a progressive manner. Specifically, (i) to enforce the model to look for different clues in the image, we adopt a multi-stage classifier and expect that the model is able to focus on a complementary region in each stage. (ii) we propose an Attentive feature Hard-Mix (A-Hard-Mix) to replace the salient feature blocks by the negative example in the current batch, whose label is different from the current sample. (iii) extensive experiments have been carried out on reID datasets such as the Market-1501, DukeMTMC-reID and CUHK03, showing that the proposed method can boost the re-identification performance significantly. Source code1has been released. Binyu He, Li Sun 0012, Qingli Li |
ICASSP | 4 |
| 2021 | Bridging the Gap between Label- and Reference-based Synthesis in Multi-attribute Image-to-Image TranslationabstractThe image-to-image translation (I2TT) model takes a target label or a reference image as the input, and changes a source into the specified target domain. The two types of synthesis, either label- or reference-based, have substantial differences. Particularly, the label-based synthesis reflects the common characteristics of the target domain, and the reference-based shows the specific style similar to the reference. This paper intends to bridge the gap between them in the task of multi-attribute I2TT. We design the label- and reference-based encoding modules (LEM and REM) to compare the domain differences. They first transfer the source image and target label (or reference) into a common embedding space, by providing the opposite directions through the attribute difference vector. Then the two embeddings are simply fused together to form the latent code Srand(or Sref), reflecting the domain style differences, which is injected into each layer of the generator by SPADE. To link LEM and REM, so that two types of results benefit each other, we encourage the two latent codes to be close, and set up the cycle consistency between the forward and backward translations on them. Moreover, the interpolation between the Srandand Srefis also used to synthesize an extra image. Experiments show that label- and reference-based synthesis are indeed mutually promoted, so that we can have the diverse results from LEM, and high quality results with the similar style of the reference. Code will be available at https://github.com/huangqiusheng/BridgeGAN. Qiusheng Huang, Zhilin Zheng, Xueqi Hu, Li Sun 0012, Qingli Li |
ICCV | 5 |
| 2021 | Low-Cost and Unobtrusive Respiratory Condition Monitoring Based on Raspberry Pi and Recurrent Neural NetworkabstractThis paper presents a low-cost and unobtrusive intelligent respiratory monitoring system. To achieve low-cost and remote measurement of respiratory signal, an RGB camera collaborated with marker tracking is used as data acquisition sensor, and a Raspberry Pi is used as data processing platform. To overcome challenges in actual applications, the signal processing algorithms are designed for removing sudden body movements and smoothing the raw signal. To discover more specific information in the respiratory signal, respiratory rate is estimated by a translational cross point algorithm, and respiratory pattern is identified by recurrent neural network. Finally, the obtained decision-making information and some original information are sent to user's smartphone via a cloud service platform. For estimating respiratory rate, the Bland-Altman plot demonstrates the satisfactory results with agreement ranges of -0.13 ± 5.85 bpm. With respect to the classification of breathing patterns, the results validate that the system has the good performance with the accuracy, precision, recall, and F1 of 92.5%, 92.5%, 93.3%, and 92.9%, respectively. This work may contribute to the development of low-cost and non-contact respiratory monitoring products specific to home or work health care. Yunlu Wang, Menghan Hu, Jian Zhang 0060, Qingli Li, Guangtao Zhai, Simon X. Yang |
ISCAS | 6 |
| 2021 | A Coherent Cooperative Learning Framework Based on Transfer Learning for Unsupervised Cross-Domain Classification
Xinxin Shan, Ying Wen 0003, Qingli Li, Yue Lu 0001, Haibin Cai |
MICCAI (5) | 3 |
| 2021 | Respiratory Consultant by Your Side: Affordable and Remote Intelligent Respiratory Rate and Respiratory Pattern Monitoring SystemabstractThe aim of this study is to develop an affordable and remote intelligent respiratory monitoring system. To achieve low-cost and remote measurement of respiratory signal, an RGB camera collaborated with marker tracking is used as a data acquisition sensor, and a Raspberry Pi is used as a data processing platform. To overcome challenges in actual applications, the signal processing algorithms are designed for removing sudden body movements and smoothing the raw signal. Subsequently, respiratory rate (RR) is estimated by a translational cross-point algorithm, and the respiratory pattern is identified by the recurrent neural network. For estimating RR, the translational cross-point algorithm performs better than other methods with root-mean-square error (RMSE) of 3.29 bpm. With respect to the classification of breathing patterns, the established neural network performs better than support vector machine-based classifiers with the accuracy, precision, recall, and F1 of 89.0%, 89.0%, 90.5%, and 89.0%, respectively. The obtained decision-making information and some original information are sent to the user’s smartphone via a cloud service platform. In a way, due to its low-price, noncontact, and portable merits, the established system can be seen as a “respiratory consultant” by your side. Yunlu Wang, Menghan Hu, Jian Zhang 0060, Qingli Li, Guangtao Zhai, Simon X. Yang, Xiao-Ping Zhang 0002, Xiaokang Yang 0001 |
IEEE Internet Things J. | 6 |
| 2021 | Model-Based Transfer Learning and Sparse Coding for Partial Face RecognitionabstractWith the growing needs of practical applications such as security monitoring, partial face recognition is a challenging but important issue, because the captured faces in real-world surveillance videos may be occluded or with variations. Though current face recognition methods perform well in relatively constrained scenes, they may suffer from degradation for partial faces. In this paper, we propose a framework of model-based transfer learning and sparse coding (MTLSC) for partial face recognition. First, due to less information in partial face image, we exploit the mirrored image of an original probe sample as sample augment to provide further information. Considering the inadequacy of training face samples, we obtain face features based on model-based transfer learning VGGNet that is pre-trained on VGGFace dataset. Then we reconstruct face features by sliding window in view of different sizes of partial face hard to extract the same feature dimension. Finally we carry out sparse coding with rectification and calculate the minimum score of the probe and mirrored samples among all classes to get the results. Thus, by model-based transfer learning, sliding window for feature reconstruction and sparse coding with rectification, the proposed framework improves partial face recognition performance. Experimental results on three face databases (LFW, AR and NIR), and two person re-identification databases (iLIDS-VID and PKU-Reid) demonstrate our method is effective for partial face recognition. Xinxin Shan, Yue Lu 0001, Qingli Li, Ying Wen 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Adaptive Effective Receptive Field Convolution for Semantic Segmentation of VHR Remote Sensing ImagesabstractConvolutional neural networks (CNNs) have facilitated impressive improvements in the semantic segmentation of very high-resolution (VHR) remote sensing images. The success of semantic segmentation depends on an effective receptive field (RF) large enough to cover the entire object. Popular methods to enlarge the effective RF include dilated filters, subsampling operations, and stacking layers. Unfortunately, the methods are inefficient or able to cause grid artifacts. Moreover, although the object sizes vary greatly in remote sensing images, the size of the RF cannot reach a compromise between small and large objects. To tackle these problems, we propose adaptive effective receptive convolution (AERFC) for VHR remote sensing images. AERFC adaptively controls the sampling location of convolution and automatically adjusts the effective RF without significantly increasing the parameter number and computational cost. Thus, AERFC reduces the training difficulty, decreases overfitting risk, and reserves details in VHR images. AERFC is also integrated with spatial pyramid pooling (SPP) to aggregate diverse multiscale features for exploring contextual information. Experimental results of the quantitative and qualitative evaluation over four benchmark data sets show that AERFC outperforms state-of-the-art methods. Xi Chen 0004, Zhen Han 0007, Shiyi Deng, Qingli Li |
IEEE Trans. Geosci. Remote. Sens. | 9 |
| 2021 | Identification of Melanoma From Hyperspectral Pathology Image Using 3D Convolutional NetworksabstractSkin biopsy histopathological analysis is one of the primary methods used for pathologists to assess the presence and deterioration of melanoma in clinical. A comprehensive and reliable pathological analysis is the result of correctly segmented melanoma and its interaction with benign tissues, and therefore providing accurate therapy. In this study, we applied the deep convolution network on the hyperspectral pathology images to perform the segmentation of melanoma. To make the best use of spectral properties of three dimensional hyperspectral data, we proposed a 3D fully convolutional network named Hyper-net to segment melanoma from hyperspectral pathology images. In order to enhance the sensitivity of the model, we made a specific modification to the loss function with caution of false negative in diagnosis. The performance of Hyper-net surpassed the 2D model with the accuracy over 92%. The false negative rate decreased by nearly 66% using Hyper-net with the modified loss function. These findings demonstrated the ability of the Hyper-net for assisting pathologists in diagnosis of melanoma based on hyperspectral pathology images. Qian Wang 0046, Li Sun 0012, Yan Wang 0033, Mei Zhou, Menghan Hu, Ying Wen 0003, Qingli Li |
IEEE Trans. Medical Imaging | 8 |
| 2020 | Novel View Synthesis on Unpaired Data by Conditional Deformable Variational Auto-Encoder
Mingyu Yin, Li Sun 0012, Qingli Li |
ECCV (28) | 3 |
| 2020 | Disentangling The Spatial Structure And Style In Conditional VAEabstractThis paper proposes a structure in conditional variation autoencoder (cVAE) to disentangle the latent vector into a spatial structure and a style code, complementary to each other, with the one $( z_{s})$ being label relevant and the other $( z_{u})$ irrelevant. Different from traditional cVAE, our network maps the condition label into its relevant code zsthrough a separated module. Depending on whether the label directly relates to the image spatial structure or not, zsoutput from the condition mapping module is used either as the style code with the two spatial dimension of $1 \times 1$, or as the spatial structure code with a single channel. Based on the input image and its corresponding zs, the encoder provides the posterior distribution close to a common prior regardless of its label, thus zusampled from it becomes label irrelevant. The decoder employs zsand zuby two typical adaptive normalization modules to reconstruct the input image. Results on two datasets with different types of labels show the effectiveness of our method. Li Sun 0012, Zhilin Zheng, Qingli Li |
ICIP | 4 |
| 2020 | Wearable Visually Assistive Device for Blind People to Appreciate Real-world Scene and Screen ImageabstractDue to the loss of vision, the appreciation of the realworld scene and the images displayed on the screen becomes almost impossible for blind people. In an effort to meet the needs of the blind community, we develop a wearable visually assistive device to help them perceive images. With the help of various multimedia information processing technologies, the proposed device can first acquire image information through a depth camera, then implement an image-to-text transformation using image caption technology, and finally the obtained text sequence is fed back to the user via voice. In this way, blind people are able to perceive the outside world, thus creating an unprecedented experience for them. The main technical specifications of the system are: distance perception range is 0.1m to 10m; RGB field of view is 69.4°×42.5°×77°; depth field of view is 91.2°×65.5°×100.6°; maximum weight is 3.05kg. Two demo videos of the proposed navigation system which are respectively recorded for real-world scene and screen image are available at: https://doi.org/10.6084/m9.figshare.12520499.v1. Jin Ai, Menghan Hu, Guangtao Zhai, Jian Zhang 0060, Qingli Li, Wendell Q. Sun |
VCIP | 6 |
| 2020 | Special Cane with Visual Odometry for Real-time Indoor Navigation of Blind PeopleabstractIndoor navigation is urgently needed by blind people in their everyday lives. In this paper, we design an assistive cane with visual odometry based on actual requirements of the blind to aid them in attaining safe indoor navigation. Compared to the state-of-the-art indoor navigation systems, the proposed device is portable, compact, and adaptable. The main specifications of the system are: the perception range is respectively from 0.10m to 2.10m, and 0.08m to 1.60m for width and length dimensions; the maximum weight is 2.1kg; the detection range is from 0.15m and 3.00m; the cruising ability is about 8h; and the objects whose heights are below 80cm can be detected. The demo video of the proposed navigation system is available at: https://doi.org/10.6084/m9.figshare.12399572.v1. Menghan Hu, Qingli Li, Jian Zhang 0060, Xiaofeng Zhou 0002, Guangtao Zhai |
VCIP | 4 |
| 2020 | Unobtrusive and Automatic Classification of Multiple People's Abnormal Respiratory Patterns in Real Time Using Deep Neural Network and Depth CameraabstractRespiratory pattern is a representation of human breathing activity, which can reflect people's physical and psychological condition. Capturing the unexpected abnormal respiratory pattern unobtrusively of the patient or the potential patient has great significance. In the current work, we attempt to capitalize on depth camera and deep learning architecture to achieve the accurate and unobtrusive measurement of abnormal respiratory patterns, and the whole system can classify multiple people's respiratory patterns in a real-time manner. The challenges in this task are threefold: 1) the real-time online system means that the Region of Interest (ROI) needs to be located and tracked automatically; 2) the amount of real-world data is not enough for training to obtain the robust deep neural network; and 3) the intraclass variation is large and the outer class variation is small. Consequently, human joints tracking is applied to determine the location of subjects shoulder and chest. Based on the characteristics of actual respiratory signals, a novel and efficient respiratory simulation model (RSM) is proposed to generate abundant and high-quality training data. Finally, we apply a gated recurrent unit (GRU) neural network with bidirectional and attentional mechanisms (BI-AT-GRU) to classify six clinically significant respiratory patterns (Eupnea, Tachypnea, Bradypnea, Biots, Cheyne-Stokes, and Central-Apnea). The performance of the obtained BI-AT-GRU is tested by the data that is actually measured by the depth camera. The experimental results demonstrate that the proposed model can classify six different respiratory patterns with the accuracy, precision, recall, and F1 of 94.5%, 94.4%, 95.1%, and 94.8%, respectively. In comparative experiments, the obtained BI-AT-GRU specific to respiratory pattern classification outperforms the existing state-of-the-art, viz., BI-AT-LSTM, GRU, long short-term memory (LSTM), and BI-AT-GRU. Moreover, other experimental results indicate that the proposed online measuring system, deep neural network, and the modeling ideas have the potential to be extended to the large-scale applications, such as public places, sleep scenario, and office environment. The demo videos of the proposed system are available at: https://doi.org/10.6084/m9.figshare.11493666.v1. Yunlu Wang, Menghan Hu, Yuwen Zhou, Qingli Li, Nan Yao, Guangtao Zhai, Xiao-Ping Zhang 0002, Xiaokang Yang 0001 |
IEEE Internet Things J. | 4 |
| 2020 | Gabor Feature-Based LogDemons With Inertial Constraint for Nonrigid Image RegistrationabstractNonrigid image registration plays an important role in the field of computer vision and medical application. The methods based on Demons algorithm for image registration usually use intensity difference as similarity criteria. However, intensity based methods can not preserve image texture details well and are limited by local minima. In order to solve these problems, we propose a Gabor feature based LogDemons registration method in this paper, called GFDemons. We extract Gabor features of the registered images to construct feature similarity metric since Gabor filters are suitable to extract image texture information. Furthermore, because of the weak gradients in some image regions, the update fields are too small to transform the moving image to the fixed image correctly. In order to compensate this deficiency, we propose an inertial constraint strategy based on GFDemons, named IGFDemons, using the previous update fields to provide guided information for the current update field. The inertial constraint strategy can further improve the performance of the proposed method in terms of accuracy and convergence. We conduct experiments on three different types of images and the results demonstrate that the proposed methods achieve better performance than some popular methods. Ying Wen 0003, Yue Lu 0001, Qingli Li, Haibin Cai, Lianghua He |
IEEE Trans. Image Process. | 4 |
| 2020 | Blood Cell Classification Based on Hyperspectral Imaging With Modulated Gabor and CNNabstractCell classification, especially that of white blood cells, plays a very important role in the field of diagnosis and control of major diseases. Compared to traditional optical microscopic imaging, hyperspectral imagery, combined with both spatial and spectral information, provides more wealthy information for recognizing cells. In this paper, a novel blood cell classification framework, which combines a modulated Gabor wavelet and deep convolutional neural network (CNN) kernels, named as MGCNN, is proposed based on medical hyperspectral imaging. For each convolutional layer, multi-scale and orientation Gabor operators are taken dot product with initial CNN kernels. The essence is to transform the convolutional kernels into the frequency domain to learn features. By combining characteristics of Gabor wavelets, the features learned by modulated kernels at different frequencies and orientations are more representative and discriminative. Experimental results demonstrate that the proposed model can achieve better classification performance than traditional CNNs and widely used support vector machine approaches, especially as training small-sample-size situations. Wei Li 0032, Baochang Zhang 0001, Qingli Li, Ran Tao 0003, Nigel H. Lovell |
IEEE J. Biomed. Health Informatics | 4 |
| 2019 | Robust Real-Time Object Detection Based on Deep Learning for Very High Resolution Remote Sensing ImagesabstractRecently, the development of deep learning boosts the object detection for remote sensing images. The existing deep learning methods can be divided into two types. The region-based methods represented by Faster R-CNN have progressive performance in accuracy. However, their computational cost is massive due to the deep Convolutional Neural Network (CNN) backbones, which limits the efficiency. The regression-based methods such as YOLO and Single Shot MultiBox Detector (SSD) are advantageous in speed while the accuracy is not satisfactory. To meet the increasing demand in both speed and accuracy for object detection of remote sensing images, we employ the Reception Field Block Net (RFBNet) detector. It embeds the Receptive Field Block (RFB) module into SSD to obtain better feature representation. The experimental results on NWPU VHR-10 dataset demonstrate that the mAP of RFBNet-512 reaches 91.56%, which outperforms other state-of-the-art networks. Meanwhile, the speed is also competitive. Jinzheng Zhao, Weiyu Xiong, Qingli Li, Junli Yang |
IGARSS | 5 |
| 2019 | Angel Girl of Visually Impaired Artists: Painting Navigation System for Blind or Visually Impaired PaintersabstractFor those who love painting but unfortunately have visual impairments, holding a paintbrush to create a work is really a difficult task. For the purpose of solving this problem, a painting navigation system for visually impaired painters is introduced through the live demonstration. When painting, the developed system can endow visually impaired persons with the ability to perceive the surrounding environment, thus helping them realize their dream of painting. To achieve this goal, we designed four main modules viz., QR code based drawing board positioning module, brush real-time positioning module, color recognition module and human-computer interaction module, and integrated them into the system. In the validation experiments, the blindfolded users can successfully create a painting with the help of the developed navigation system. Moreover, the users told us that this system provided them with good experience. In a way, this painting navigation system can be seen as "angel's eyes" of visually impaired painters. The demo video of the proposed painting navigation system is available at: https://doi.org/10.6084/m9.figshare.9760004.v1. Menghan Hu, Guangtao Zhai, Huijing Huang, Wa Zhang, Qingli Li, Yinghong Tian, Yanling Shi |
VCIP | 6 |
| 2018 | Position-Squeeze and Excitation Block for Facial Attribute Analysis
Wanxia Shen, Li Sun 0012, Qingli Li |
BMVC | 4 |
| 2018 | Investigation in Spatial-Temporal Domain for Face Spoof DetectionabstractThis paper focuses on face spoofing detection using video. The purpose is to find out the best scheme for this task in the end-to-end learning manner. We investigate 4 different types of structure to fully exploit the raw data in its spatial-temporal domain, which are the pure CNN, CNN with 3D convolution, CNN+LSTM and CNN+Conv-LSTM. Moreover, another stream built on optical flow is also used, and with a proper fusion method, it can improve the accuracy. In experiments, we compare schemes on the raw data in single stream and fusion methods with optical flow in two streams. The performance are not only given within each dataset, but also measured across different datsets, which is crucial to avoid the overfitting. Zhonglin Sun, Li Sun 0012, Qingli Li |
ICASSP | 3 |
| 2018 | Person Re-id by Incorporating PCA Loss in CNN
Li Sun 0012, Song Qiu, Qingli Li |
MMM (2) | 5 |
| 2017 | Facial age estimation through self-paced learningabstractThis paper proposes an age estimation algorithm in Self-Paced Learning (SPL) framework. Facial samples in the training set inherently include both easy and complex images, which is caused by both the characteristic of age and the variation of pose or expression. Furthermore, by randomly hiding patches in face region, data with different difficulty levels can be gradually used by SPL, in which Convolution Neural Network (CNN) is trained to give the estimation. Alternative Optimization Strategy (AVO), for the weight of CNN and the latent weight in SPL regularizer, is adopted in SPL framework. Experiments show that the proposed algorithm is able to give an accurate results especially under pose and expression variation. Li Sun 0012, Song Qiu, Mei Zhou, Qingli Li |
VCIP | 5 |
| 2010 | Finger vein recognition with manifold learning
Zhi Liu 0004, Yilong Yin, Hongjun Wang 0004, Shangling Song, Qingli Li |
J. Netw. Comput. Appl. | 5 |