Liangqiong Qu

dblp:149/2634 · DBLP profile ↗
← Back
37ranked-venue papers
6as first author
28since 2021 · last 2026
0000-0001-8235-7852ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 21 · 4 first-author · 16 since 2021Artificial intelligence and machine learning · 19 · 3 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 7 since 2021
YearPublicationVenuePosition
2026 StyleTailor: Towards Personalized Fashion Styling via Hierarchical Negative Feedback
abstract
The advancement of intelligent agents has revolutionized problem-solving across diverse domains, yet solutions for personalized fashion styling remain underexplored, which holds immense promise for promoting shopping experiences. In this work, we present StyleTailor, the first collaborative agent framework that seamlessly unifies personalized apparel design, shopping recommendation, virtual try-on, and systematic evaluation into a cohesive workflow. To this end, StyleTailor pioneers an iterative visual refinement paradigm driven by multi-level negative feedback, enabling adaptive and precise user alignment. Specifically, our framework features two core agents, i.e., Designer for personalized garment selection and Consultant for virtual try-on, whose outputs are progressively refined via hierarchical vision-language model feedback spanning individual items, complete outfits, and try-on efficacy. Counterexamples are aggregated into negative prompts, forming a closed-loop mechanism that enhances recommendation quality. To assess the performance, we introduce a comprehensive evaluation suite encompassing style consistency, visual quality, face similarity, and artistic appraisal. Extensive experiments demonstrate StyleTailor's superior performance in delivering personalized designs and recommendations, outperforming strong baselines without negative feedback and establishing a new benchmark for intelligent fashion systems.
Hongbo Ma, Fei Shen 0004, Xiaoce Wang, Jinkai Zheng, Liangqiong Qu, Ming Li 0073
AAAI7
2026 Unleashing the Potential of Large Language Models for Text-to-Image Generation Through Autoregressive Representation Alignment
abstract
We present Autoregressive Representation Alignment (ARRA), a new training framework that unlocks global-coherent text-to-image generation in autoregressive LLMs without architectural modifications. Different from prior works that require complex architectural redesigns, ARRA aligns LLM's hidden states with visual representations from external visual foundational models via a global visual alignment loss and a hybrid token, . This token enforces dual constraints: local next-token prediction and global semantic distillation, enabling LLMs to implicitly learn spatial and contextual coherence while retaining their original autoregressive paradigm. Extensive experiments validate ARRA's plug-and-play versatility. When training T2I LLMs from scratch, ARRA reduces FID by 16.6% (ImageNet), 12.0% (LAION-COCO) for autoregressive LLMs like LlamaGen, without modifying original architecture and inference mechanism. For training from text-generation-only LLMs, ARRA reduces FID by 25.5% (MIMIC-CXR), 8.8% (DeepEyeNet) for advanced LLMs like Chameleon. For domain adaptation, ARRA aligns general-purpose LLMs with specialized models (e.g., BioMedCLIP), achieving an 18.6% FID reduction over direct fine-tuning on medical imaging (MIMIC-CXR). These results demonstrate that training objective redesign, rather than architectural modifications, can resolve cross-modal global coherence challenges. ARRA offers a complementary paradigm for advancing autoregressive models.
Jiawei Liu 0003, Ziyue Lin, Huijie Fan, Zhi Han, Yandong Tang, Liangqiong Qu
AAAI7
2026 Exploring the Vulnerabilities of Federated Learning: A Deep Dive Into Gradient Inversion Attacks
abstract
Federated Learning (FL) has emerged as a promising privacy-preserving collaborative model training paradigm without sharing raw data. However, recent studies have revealed that private information can still be leaked through shared gradient information and attacked by Gradient Inversion Attacks (GIA). While many GIA methods have been proposed, a detailed analysis, evaluation, and summary of these methods are still lacking. Although various survey papers summarize existing privacy attacks in FL, few studies have conducted extensive experiments to unveil the effectiveness of GIA and their associated limiting factors in this context. To fill this gap, we first undertake a systematic review of GIA and categorize existing methods into three types, i.e., optimization-based GIA (OP-GIA), generation-based GIA (GEN-GIA), and analytics-based GIA (ANA-GIA). Then, we comprehensively analyze and evaluate the three types of GIA in FL, providing insights into the factors that influence their performance, practicality, and potential threats. Our findings indicate that OP-GIA is the most practical attack setting despite its unsatisfactory performance, while GEN-GIA has many dependencies and ANA-GIA is easily detectable, making them both impractical. Finally, we offer a three-stage defense pipeline to users when designing FL frameworks and protocols for better privacy protection and share some future research directions from the perspectives of attackers and defenders that we believe should be pursued. We hope that our study can help researchers design more robust FL frameworks to defend against these attacks.
Pengxin Guo 0001, Runxi Wang, Shuang Zeng, Jinjing Zhu, Haoning Jiang, Yuyin Zhou, Hui Xiong 0001, Liangqiong Qu
IEEE Trans. Pattern Anal. Mach. Intell.10
2026 Understanding gait recognition through silhouette sequence disentanglement and fine-grained visualization
Shaoxiong Zhang 0001, Yixiu Liu, Jinkai Zheng, Liangqiong Qu, Ming Li 0073, Chenggang Yan 0001
Pattern Recognit.4
2026 DVG-Diffusion: Dual-View-Guided Diffusion Model for CT Reconstruction From X-Rays
abstract
Directly reconstructing 3D CT volume from few-view 2D X-rays using an end-to-end deep learning network is a challenging task, as X-ray images are merely projection views of the 3D CT volume. In this work, we facilitate complex 2D X-ray image to 3D CT mapping by incorporating new view synthesis, and reduce the learning difficulty through view-guided feature alignment. Specifically, we propose a dual-view guided diffusion model (DVG-Diffusion), which couples a real input X-ray view and a synthesized new X-ray view to jointly guide CT reconstruction. First, a novel view parameter-guided encoder captures features from X-rays that are spatially aligned with CT. Next, we concatenate the extracted dual-view features as conditions for the latent diffusion model to learn and refine the CT latent representation. Finally, the CT latent representation is decoded into a CT volume in pixel space. By incorporating view parameter guided encoding and dual-view guided CT reconstruction, our DVG-Diffusion can achieve an effective balance between high fidelity and perceptual quality for CT reconstruction. Experimental results demonstrate our method outperforms state-of-the-art methods. Based on experiments, the comprehensive analysis and discussions for views and reconstruction are also presented. The model and code are available at https://github.com/xiexing0916/DVG-Diffusion.
Jiawei Liu 0003, Huijie Fan, Zhi Han, Yandong Tang, Liangqiong Qu
IEEE Trans. Image Process.6
2025 MEPNet: Medical Entity-Balanced Prompting Network for Brain CT Report Generation
abstract
The automatic generation of brain CT reports has gained widespread attention, given its potential to assist radiologists in diagnosing cranial diseases. However, brain CT scans involve extensive medical entities, such as diverse anatomy regions and lesions, exhibiting highly inconsistent spatial patterns in 3D volumetric space. This leads to biased learning of medical entities in existing methods, resulting in repetitiveness and inaccuracy in generated reports. To this end, we propose a Medical Entity-balanced Prompting Network (MEPNet), which harnesses the large language model (LLM) to fairly interpret various entities for accurate brain CT report generation. By introducing the visual embedding and the learning status of medical entities as enriched clues, our method prompts the LLM to balance the learning of diverse entities, thereby enhancing reports with comprehensive findings. First, to extract visual embedding of entities, we propose Knowledge-driven Joint Attention to explore and distill entity patterns using both explicit and implicit medical knowledge. Then, a Learning Status Scorer is designed to evaluate the learning of entity visual embeddings, resulting in unique learning status for individual entities. Finally, these entity visual embeddings and status are elaborately integrated into multi-modal prompts, to guide the text generation of LLM. This process allows LLM to self-adapt the learning process for biased-fitted entities, thereby covering detailed findings in generated reports. We conduct experiments on two brain CT report generation benchmarks, showing the effectiveness in clinical accuracy and text coherence.
Xiaodan Zhang 0003, Yanzhao Shi, Junzhong Ji, Chengxin Zheng, Liangqiong Qu
AAAI5
2025 A New Federated Learning Framework Against Gradient Inversion Attacks
abstract
Federated Learning (FL) aims to protect data privacy by enabling clients to collectively train machine learning models without sharing their raw data. However, recent studies demonstrate that information exchanged during FL is subject to Gradient Inversion Attacks (GIA) and, consequently, a variety of privacy-preserving methods have been integrated into FL to thwart such attacks, such as Secure Multi-party Computing (SMC), Homomorphic Encryption (HE), and Differential Privacy (DP). Despite their ability to protect data privacy, these approaches inherently involve substantial privacy-utility trade-offs. By revisiting the key to privacy exposure in FL under GIA, which lies in the frequent sharing of model gradients that contain private data, we take a new perspective by designing a novel privacy preserve FL framework that effectively ``breaks the direct connection'' between the shared parameters and the local private data to defend against GIA. Specifically, we propose a Hypernetwork Federated Learning (HyperFL) framework that utilizes hypernetworks to generate the parameters of the local model and only the hypernetwork parameters are uploaded to the server for aggregation. Theoretical analyses demonstrate the convergence rate of the proposed HyperFL, while extensive experimental results show the privacy-preserving capability and comparable performance of HyperFL.
Pengxin Guo 0001, Shuang Zeng, Xiaodan Zhang 0003, Weihong Ren, Yuyin Zhou, Liangqiong Qu
AAAI7
2025 Selective Aggregation for Low-Rank Adaptation in Federated Learning
abstract
We investigate LoRA in federated learning through the lens of the asymmetry analysis of the learned $A$ and $B$ matrices. In doing so, we uncover that $A$ matrices are responsible for learning general knowledge, while $B$ matrices focus on capturing client-specific knowledge. Based on this finding, we introduce Federated Share-A Low-Rank Adaptation (FedSA-LoRA), which employs two low-rank trainable matrices $A$ and $B$ to model the weight update, but only $A$ matrices are shared with the server for aggregation. Moreover, we delve into the relationship between the learned $A$ and $B$ matrices in other LoRA variants, such as rsLoRA and VeRA, revealing a consistent pattern. Consequently, we extend our FedSA-LoRA method to these LoRA variants, resulting in FedSA-rsLoRA and FedSA-VeRA. In this way, we establish a general paradigm for integrating LoRA with FL, offering guidance for future work on subsequent LoRA variants combined with FL. Extensive experimental results on natural language understanding and generation tasks demonstrate the effectiveness of the proposed method. Our code is available at https://github.com/Pengxin-Guo/FedSA-LoRA.
Pengxin Guo 0001, Shuang Zeng, Huijie Fan, Liangqiong Qu
ICLR6
2025 Region-Based Text-Consistent Augmentation for Multimodal Medical Segmentation
Kunyan Cai, Chenggang Yan 0001, Liangqiong Qu, Shuai Wang 0003, Tao Tan 0002
MICCAI (3)4
2025 TextSplat: Text-Guided Semantic Fusion for Generalizable Gaussian Splatting
abstract
Recent advancements in Generalizable Gaussian Splatting have enabled robust 3D reconstruction from sparse input views by utilizing feed-forward Gaussian Splatting models, achieving superior cross-scene generalization. However, while many methods focus on geometric consistency, they often neglect the potential of text-driven guidance to enhance semantic understanding, which is crucial for accurately reconstructing fine-grained details in complex scenes. To address this limitation, we propose TextSplat-the first text-driven Generalizable Gaussian Splatting framework. Specifically, our framework employs three parallel modules to obtain complementary representations: the Diffusion Prior Depth Estimator for accurate depth information, the Semantic Aware Segmentation Network for detailed semantic information, and the Multi-View Interaction Network for refined cross-view features. Then, in the Text-Guided Semantic Fusion Module, these representations are integrated via the text-guided and attention-based feature aggregation mechanism, resulting in enhanced 3D Gaussian parameters enriched with detailed semantic cues. Experimental results on various benchmark datasets demonstrate improved performance compared to existing methods across multiple evaluation metrics, validating the effectiveness of our framework. The code will be publicly available.
Zhicong Wu, Ping Nie, Zhixin Yan, Jinkai Zheng, Liangqiong Qu, Ming Li 0073, Liqiang Nie
ACM Multimedia7
2025 Intra- and Inter-Head Orthogonal Attention for Image Captioning
abstract
Multi-head attention (MA), which allows the model to jointly attend to crucial information from diverse representation subspaces through its heads, has yielded remarkable achievement in image captioning. However, there is no explicit mechanism to ensure MA attends to appropriate positions in diverse subspaces, resulting in overfocused attention for each head and redundancy between heads. In this paper, we propose a novel Intra- and Inter-Head Orthogonal Attention (I2OA) to efficiently improve MA in image captioning by introducing a concise orthogonal regularization to heads. Specifically, Intra-Head Orthogonal Attention enhances the attention learning of MA by introducing orthogonal constraint to each head, which decentralizes the object-centric attention to more comprehensive content-aware attention. Inter-Head Orthogonal Attention reduces the heads redundancy by applying orthogonal constraint between heads, which enlarges the diversity of representation subspaces and improves the representation ability for MA. Moreover, the proposed I2OA is flexible to combine with various multi-head attention based image captioning methods and improve the performances without increasing model complexity and parameters. Experiments on the MS COCO dataset demonstrate the effectiveness of the proposed model.
Xiaodan Zhang 0003, Aozhe Jia, Junzhong Ji, Liangqiong Qu, Qixiang Ye
IEEE Trans. Image Process.4
2025 Benchmarking Radiology Report Generation From Noisy Free-Texts
abstract
Automatic radiology report generation can enhance diagnostic efficiency and accuracy. However, clean open-source imaging scan-report pairs are limited in scale and variety. Moreover, the vast amount of radiological texts available online is often too noisy to be directly employed. To address this challenge, we introduce a novel task called Noisy Report Refinement (NRR), which generates radiology reports from noisy free-texts. To achieve this, we propose a report refinement pipeline that leverages large language models (LLMs) enhanced with guided self-critique and report selection strategies. To address the inability of existing radiology report generation metrics in measuring cleanliness, radiological usefulness, and factual correctness across various modalities of reports in NRR task, we introduce a new benchmark, NRRBench, for NRR evaluation. This benchmark includes two online-sourced datasets and four clinically explainable LLM-based metrics: two metrics evaluate the matching rate of radiology entities and modality-specific template attributes respectively, one metric assesses report cleanliness, and a combined metric evaluates overall NRR performance. Experiments demonstrate that guided self-critique and report selection strategies significantly improve the quality of refined reports. Additionally, our proposed metrics show a much higher correlation with noisy rate and error count of reports than radiology report generation metrics in evaluating NRR.
Yujian Yuan, Yanting Zheng, Liangqiong Qu
IEEE J. Biomed. Health Informatics3
2024 Exploring Self- and Cross-Triplet Correlations for Human-Object Interaction Detection
abstract
Human-Object Interaction (HOI) detection plays a vital role in scene understanding, which aims to predict the HOI triplet in the form of . Existing methods mainly extract multi-modal features (e.g., appearance, object semantics, human pose) and then fuse them together to directly predict HOI triplets. However, most of these methods focus on seeking for self-triplet aggregation, but ignore the potential cross-triplet dependencies, resulting in ambiguity of action prediction. In this work, we propose to explore Self- and Cross-Triplet Correlations (SCTC) for HOI detection. Specifically, we regard each triplet proposal as a graph where Human, Object represent nodes and Action indicates edge, to aggregate self-triplet correlation. Also, we try to explore cross-triplet dependencies by jointly considering instance-level, semantic-level, and layout-level relations. Besides, we leverage the CLIP model to assist our SCTC obtain interaction-aware feature by knowledge distillation, which provides useful action clues for HOI detection. Extensive experiments on HICO-DET and V-COCO datasets verify the effectiveness of our proposed SCTC.
Weibo Jiang, Weihong Ren, Jiandong Tian, Liangqiong Qu, Zhiyong Wang 0009, Honghai Liu 0001
AAAI4
2024 Unleashing the Potential of SAM for Medical Adaptation via Hierarchical Decoding
abstract
The Segment Anything Model (SAM) has garnered significant attention for its versatile segmentation abilities and intuitive prompt-based interface. However, its application in medical imaging presents challenges, requiring either substantial training costs and extensive medical datasets for full model fine-tuning or high-quality prompts for optimal performance. This paper introduces H-SAM: a prompt-free adaptation of SAM tailored for efficient fine-tuning of medical images via a two-stage hierarchical decoding procedure. In the initial stage, H-SAM employs SAM's original decoder to generate a prior probabilistic mask, guiding a more intricate decoding process in the second stage. Specifically, we propose two key designs: 1) A class-balanced, mask-guided self-attention mechanism addressing the unbalanced label distribution, enhancing image embedding; 2) A learnable mask cross-attention mechanism spatially modulating the interplay among different image regions based on the prior mask. Moreover, the inclusion of a hierarchical pixel decoder in H-SAM enhances its proficiency in capturing fine-grained and localized details. This approach enables SAM to effectively integrate learned medical priors, facilitating enhanced adaptation for medical image segmentation with limited samples. Our H-SAM demonstrates a 4.78% improvement in average Dice compared to existing prompt-free SAM variants for multi-organ segmentation using only 10% of 2D slices. Notably, without using any unlabeled data, H-SAM even outperforms state-of-the-art semisupervised models relying on extensive unlabeled training data across various medical datasets. Our code is available at https://github.com/Cccccczh404/H-SAM.
Zhiheng Cheng, Qingyue Wei, Hongru Zhu, Yan Wang 0033, Liangqiong Qu, Wei Shao 0008, Yuyin Zhou
CVPR5
2024 Residual Denoising Diffusion Models
abstract
We propose residual denoising diffusion models (RDDM), a novel dual diffusion process that decouples the traditional single denoising diffusion process into residual diffusion and noise diffusion. This dual diffusion framework expands the denoising-based diffusion models, initially uninterpretable for image restoration, into a unified and interpretable model for both image generation and restoration by introducing residuals. Specifically, our residual diffusion represents directional diffusion from the target image to the degraded input image and explicitly guides the reverse generation process for image restoration, while noise diffusion represents random perturbations in the diffusion process. The residual prioritizes certainty, while the noise emphasizes diversity, enabling RDDM to effectively unify tasks with varying certainty or diversity requirements, such as image generation and restoration. We demonstrate that our sampling process is consistent with that of DDPM and DDIM through coefficient transformation, and propose a partially path-independent generation process to better understand the reverse process. Notably, our RDDM enables a generic UNet, trained with only an L1 loss and a batch size of 1, to compete with state-of-the-art image restoration methods. We provide code and pre-trained models to encourage further exploration, application, and development of our innovative framework (https://github.com/nachifurlRDDM).
Jiawei Liu 0003, Qiang Wang 0015, Huijie Fan, Yandong Tang, Liangqiong Qu
CVPR6
2024 FLHetBench: Benchmarking Device and State Heterogeneity in Federated Learning
abstract
Federated learning (FL) is a powerful technology that enables collaborative training of machine learning models without sharing private data among clients. The fundamental challenge in FL lies in learning over extremely hetero-geneous data distributions, device capacities, and device state availabilities, all of which adversely impact performance and communication efficiency. While data hetero-geneity has been well-studied in the literature, this paper introduces FLHetBench, the first FL benchmark targeted toward understanding device and state heterogeneity. FL-HetBench comprises two new sampling methods to generate real-world device and state databases with varying het-erogeneity and new metrics for quantifying the success of FL methods under these real-world constraints. Using FL-HetBench, we conduct a comprehensive evaluation of existing methods and find that they struggle under these settings, which inspires us to propose BiasPrompt+, a new method employing staleness-aware aggregation and fast weights to tackle these new heterogeneity challenges. Experiments on various FL tasks and datasets validate the effectiveness of our BiasPrompt+ method and highlight the value of FLHet-Bench in fostering the development of more efficient and robust FL solutions under real-world device and state constraints.
Junyuan Zhang, Shuang Zeng, Miao Zhang 0030, Runxi Wang, Yuyin Zhou, Paul Pu Liang, Liangqiong Qu
CVPR8
2024 Tackling Data Heterogeneity in Federated Learning via Loss Decomposition
Shuang Zeng, Pengxin Guo 0001, Yuyin Zhou, Liangqiong Qu
MICCAI (10)6
2024 Learning Self- and Cross-Triplet Context Clues for Human-Object Interaction Detection
abstract
Human-Object Interaction (HOI) detection aims to infer interactions between humans and objects, and it is very important for scene analysis and understanding. The existing methods usually focus on exploring instance-level (e.g., object appearance) or interaction-level (e.g., action semantic) features to conduct interaction prediction. However, most of these methods only consider the self-triplet feature aggregation, which may lead to learning ambiguity without exploring the cross-triplet context exchange. In this paper, from both visual and textual perspectives, we propose a novel method to jointly explore self-and cross-triplet interaction context clues for HOI detection. First, we employ a graph neural network to perform self-triplet aggregation, where human and object features represent graph nodes and visual interaction feature and textual prior knowledge are acted as two different edges. Furthermore, we also attempt to explore cross-triplet context exchange by incorporating symbiotic and layout relationships among different HOI triplets. Extensive experiments on two benchmarks demonstrate that our proposed method outperforms the state-of-the-art ones and achieves the impressive performance of 40.32 mAP on HICO-DET and 69.1 mAP on V-COCO datasets, respectively.
Weihong Ren, Jinguo Luo, Weibo Jiang, Liangqiong Qu, Zhi Han, Jiandong Tian, Honghai Liu 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 MCDNet: An Infrared Small Target Detection Network Using Multi-Criteria Decision and Adaptive Labeling Strategy
abstract
The success of deep learning methods heavily relies on the availability of adequate samples. However, in the task of infrared small target detection (ISTD), the lack of high-quality training samples is a challenging problem due to the confidentiality of the application field and the difficulty of labeling. This limitation often leads to suboptimal detection performance of convolutional neural networks (CNNs). To address this challenge, we propose an adaptive labeling strategy and an ISTD network called the multi-criteria decision network (MCDNet) to achieve higher-quality sample labeling and more accurate detection results. In the adaptive labeling strategy, we propose a second-order differential autocorrelation method to determine the size of fuzzy edge targets accurately. In addition, we introduce local backgrounds to enhance the saliency information in the labels and improve the richness and contrast of training label content. To obtain accurate and robust detection results with limited target feature information, we design MCDNet. In particular, we propose a multi-criteria decision method that can combine the CNN decisions and the infrared small target prior saliency decisions through weighted fusion, and set decision weights based on the importance of different decision criteria in the decision-making process. This method can integrate the advantages of both the CNN decisions and the prior saliency decisions, avoid the one-sidedness of a single criterion, and improve the reliability and stability of the decision-making. The experimental results indicate that our method has a higher accuracy compared to other contrastive methods.
Tianlei Ma, Zhen Yang 0026, Jing J. Liang, Jun Fu 0008, Yu Dou, Yanan Ku, Liangqiong Qu
IEEE Trans. Geosci. Remote. Sens.9
2023 Granularity Matters: Pathological Graph-driven Cross-modal Alignment for Brain CT Report Generation
abstract
The automatic Brain CT reports generation can improve the efficiency and accuracy of diagnosing cranial diseases.However, current methods are limited by 1) coarse-grained supervision: the training data in image-text format lacks detailed supervision for recognizing subtle abnormalities, and 2) coupled cross-modal alignment: visual-textual alignment may be inevitably coupled in a coarse-grained manner, resulting in tangled feature representation for report generation.In this paper, we propose a novel Pathological Graph-driven Cross-modal Alignment (PGCA) model for accurate and robust Brain CT report generation.Our approach effectively decouples the cross-modal alignment by constructing a Pathological Graph to learn finegrained visual cues and align them with textual words.This graph comprises heterogeneous nodes representing essential pathological attributes (i.e., tissue and lesion) connected by intra-and inter-attribute edges with prior domain knowledge.Through carefully designed graph embedding and updating modules, our model refines the visual features of subtle tissues and lesions and aligns them with textual words using contrastive learning.Extensive experimental results confirm the viability of our method.We believe that our PGCA model holds the potential to greatly enhance the automatic generation of Brain CT reports and ultimately contribute to improved cranial disease diagnosis.
Yanzhao Shi, Junzhong Ji, Xiaodan Zhang 0003, Liangqiong Qu
EMNLP4
2023 Label-Efficient Self-Supervised Federated Learning for Tackling Data Heterogeneity in Medical Imaging
abstract
The collection and curation of large-scale medical datasets from multiple institutions is essential for training accurate deep learning models, but privacy concerns often hinder data sharing. Federated learning (FL) is a promising solution that enables privacy-preserving collaborative learning among different institutions, but it generally suffers from performance deterioration due to heterogeneous data distributions and a lack of quality labeled data. In this paper, we present a robust and label-efficient self-supervised FL framework for medical image analysis. Our method introduces a novel Transformer-based self-supervised pre-training paradigm that pre-trains models directly on decentralized target task datasets using masked image modeling, to facilitate more robust representation learning on heterogeneous data and effective knowledge transfer to downstream models. Extensive empirical results on simulated and real-world medical imaging non-IID federated datasets show that masked image modeling with Transformers significantly improves the robustness of models against various degrees of data heterogeneity. Notably, under severe data heterogeneity, our method, without relying on any additional pre-training data, achieves an improvement of 5.06%, 1.53% and 4.58% in test accuracy on retinal, dermatology and chest X-ray classification compared to the supervised baseline with ImageNet pre-training. In addition, we show that our federated self-supervised pre-training methods yield models that generalize better to out-of-distribution data and perform more effectively when fine-tuning with limited labeled data, compared to existing FL algorithms. The code is available at https://github.com/rui-yan/SSL-FL.
Liangqiong Qu, Qingyue Wei, Shih-Cheng Huang, Liyue Shen, Daniel L. Rubin, Lei Xing 0001, Yuyin Zhou
IEEE Trans. Medical Imaging2
2023 A Decoupled Multi-Task Network for Shadow Removal
abstract
Shadow removal, which aims to restore the illumination in shadow regions, is challenging due to the diversity of shadows in terms of location, intensity, shape, and size. Different from most multi-task methods, which design elaborate multi-branch or multi-stage structures for better shadow removal, we introduce feature decomposition to learn better feature representations. Specifically, we propose a single-stage and decoupled multi-task network (DMTN) to explicitly learn the decomposed features for shadow removal, shadow matte estimation, and shadow image reconstruction. First, we propose several coarse-to-fine semi-convolution (SMC) modules to capture features sufficient for joint learning of these three tasks. Second, we design a theoretically supported feature decoupling layer to explicitly decouple the learned features into shadow image features and shadow matte features via weight reassignment. Last, these features are converted to a target shadow-free image, affiliated shadow matte, and shadow image, supervised by multi-task joint loss functions. With multi-task collaboration, DMTN effectively recovers the illumination in shadow areas while ensuring the fidelity of non-shadow areas. Experimental results show that DMTN competes favorably with state-of-the-art multi-branch/multi-stage shadow removal methods, while maintaining the simplicity of single-stage methods. We have released our code to encourage future exploration in powerful feature representation for shadow removalhttps://github.com/nachifur/DMTN
Jiawei Liu 0003, Qiang Wang 0015, Huijie Fan, Liangqiong Qu, Yandong Tang
IEEE Trans. Multim.5
2023 Breast Tumor Segmentation in DCE-MRI With Tumor Sensitive Synthesis
abstract
Segmenting breast tumors from dynamic contrast-enhanced magnetic resonance (DCE-MR) images is a critical step for early detection and diagnosis of breast cancer. However, variable shapes and sizes of breast tumors, as well as inhomogeneous background, make it challenging to accurately segment tumors in DCE-MR images. Therefore, in this article, we propose a novel tumor-sensitive synthesis module and demonstrate its usage after being integrated with tumor segmentation. To suppress false-positive segmentation with similar contrast enhancement characteristics to true breast tumors, our tumor-sensitive synthesis module can feedback differential loss of the true and false breast tumors. Thus, by following the tumor-sensitive synthesis module after the segmentation predictions, the false breast tumors with similar contrast enhancement characteristics to the true ones will be effectively reduced in the learned segmentation model. Moreover, the synthesis module also helps improve the boundary accuracy while inaccurate predictions near the boundary will lead to higher loss. For the evaluation, we build a very large-scale breast DCE-MR image dataset with 422 subjects from different patients, and conduct comprehensive experiments and comparisons with other algorithms to justify the effectiveness, adaptability, and robustness of our proposed method.
Shuai Wang 0003, Li Wang 0026, Liangqiong Qu, Fuhua Yan, Qian Wang 0001, Dinggang Shen
IEEE Trans. Neural Networks Learn. Syst.4
2022 Rethinking Architecture Design for Tackling Data Heterogeneity in Federated Learning
abstract
Federated learning is an emerging research paradigm enabling collaborative training of machine learning models among different organizations while keeping data private at each institution. Despite recent progress, there remain fundamental challenges such as the lack of convergence and the potential for catastrophic forgetting across real-world heterogeneous devices. In this paper, we demonstrate that self-attention-based architectures (e.g., Transformers) are more robust to distribution shifts and hence improve federated learning over heterogeneous data. Concretely, we conduct the first rigorous empirical investigation of different neural architectures across a range of federated algorithms, real-world benchmarks, and heterogeneous data splits. Our experiments show that simply replacing convolutional networks with Transformers can greatly reduce catastrophic forgetting of previous devices, accelerate convergence, and reach a better global model, especially when dealing with heterogeneous data. We release our code and pretrained models to encourage future exploration in robust architectures as an alternative to current research efforts on the optimization front.
Liangqiong Qu, Yuyin Zhou, Paul Pu Liang, Yingda Xia, Ehsan Adeli-Mosabbeb, Li Fei-Fei 0001, Daniel L. Rubin
CVPR1
2022 Relation constraint self-attention for image captioning
Junzhong Ji, Mingzhan Wang, Xiaodan Zhang 0003, Minglong Lei, Liangqiong Qu
Neurocomputing5
2022 A cascaded nested network for 3T brain MR image segmentation guided by 7T labeling
Zhengwang Wu, Li Wang 0026, Toan Duc Bui, Liangqiong Qu, Pew-Thian Yap, Yong Xia 0001, Gang Li 0001, Dinggang Shen
Pattern Recognit.5
2022 SplitAVG: A Heterogeneity-Aware Federated Deep Learning Method for Medical Imaging
abstract
Federated learning is an emerging research paradigm for enabling collaboratively training deep learning models without sharing patient data. However, the data from different institutions are usually heterogeneous across institutions, which may reduce the performance of models trained using federated learning. In this study, we propose a novel heterogeneity-aware federated learning method, SplitAVG, to overcome the performance drops from data heterogeneity in federated learning. Unlike previous federated methods that require complex heuristic training or hyper parameter tuning, our SplitAVG leverages the simple network split and feature map concatenation strategies to encourage the federated model training an unbiased estimator of the target data distribution. We compare SplitAVG with seven state-of-the-art federated learning methods, using centrally hosted training data as the baseline on a suite of both synthetic and real-world federated datasets. We find that the performance of models trained using all the comparison federated learning methods degraded significantly with the increasing degrees of data heterogeneity. In contrast, SplitAVG method achieves comparable results to the baseline method under all heterogeneous settings, that it achieves 96.2% of the accuracy and 110.4% of the mean absolute error obtained by the baseline in a diabetic retinopathy binary classification dataset and a bone age prediction dataset, respectively, on highly heterogeneous data partitions. We conclude that SplitAVG method can effectively overcome the performance drops from variability in data distributions across institutions. Experimental results also show that SplitAVG can be adapted to different base convolutional neural networks (CNNs) and generalized to various types of medical imaging tasks. The code is publicly available at https://github.com/zm17943/SplitAVG.
Miao Zhang 0030, Liangqiong Qu, Praveer Singh, Jayashree Kalpathy-Cramer, Daniel L. Rubin
IEEE J. Biomed. Health Informatics2
2021 Multi-Scale Context-Guided Deep Network for Automated Lesion Segmentation With Endoscopy Images of Gastrointestinal Tract
abstract
Accurate lesion segmentation based on endoscopy images is a fundamental task for the automated diagnosis of gastrointestinal tract (GI Tract) diseases. Previous studies usually use hand-crafted features for representing endoscopy images, while feature definition and lesion segmentation are treated as two standalone tasks. Due to the possible heterogeneity between features and segmentation models, these methods often result in sub-optimal performance. Several fully convolutional networks have been recently developed to jointly perform feature learning and model training for GI Tract disease diagnosis. However, they generally ignore local spatial details of endoscopy images, as down-sampling operations (e.g., pooling and convolutional striding) may result in irreversible loss of image spatial information. To this end, we propose a multi-scale context-guided deep network (MCNet) for end-to-end lesion segmentation of endoscopy images in GI Tract, where both global and local contexts are captured as guidance for model training. Specifically, one global subnetwork is designed to extract the global structure and high-level semantic context of each input image. Then we further design two cascaded local subnetworks based on output feature maps of the global subnetwork, aiming to capture both local appearance information and relatively high-level semantic information in a multi-scale manner. Those feature maps learned by three subnetworks are further fused for the subsequent task of lesion segmentation. We have evaluated the proposed MCNet on 1,310 endoscopy images from the public EndoVis-Ab and CVC-ClinicDB datasets for abnormal segmentation and polyp segmentation, respectively. Experimental results demonstrate that MCNet achieves [Formula: see text] and [Formula: see text] mean intersection over union (mIoU) on two datasets, respectively, outperforming several state-of-the-art approaches in automated lesion segmentation with endoscopy images of GI Tract.
Shuai Wang 0003, Yang Cong, Hancan Zhu, Xianyi Chen, Liangqiong Qu, Huijie Fan, Qiang Zhang 0008, Mingxia Liu 0001
IEEE J. Biomed. Health Informatics5
2020 Synthesized 7T MRI from 3T MRI via deep learning in spatial and wavelet domains
Liangqiong Qu, Yongqin Zhang, Shuai Wang 0002, Pew-Thian Yap, Dinggang Shen
Medical Image Anal.1
2020 CT Male Pelvic Organ Segmentation via Hybrid Loss Network With Incomplete Annotation
abstract
Sufficient data with complete annotation is essential for training deep models to perform automatic and accurate segmentation of CT male pelvic organs, especially when such data is with great challenges such as low contrast and large shape variation. However, manual annotation is expensive in terms of both finance and human effort, which usually results in insufficient completely annotated data in real applications. To this end, we propose a novel deep framework to segment male pelvic organs in CT images with incomplete annotation delineated in a very user-friendly manner. Specifically, we design a hybrid loss network derived from both voxel classification and boundary regression, to jointly improve the organ segmentation performance in an iterative way. Moreover, we introduce a label completion strategy to complete the labels of the rich unannotated voxels and then embed them into the training data to enhance the model capability. To reduce the computation complexity and improve segmentation performance, we locate the pelvic region based on salient bone structures to focus on the candidate segmentation organs. Experimental results on a large planning CT pelvic organ dataset show that our proposed method with incomplete annotation achieves comparable segmentation performance to the state-of-the-art methods with complete annotation. Moreover, our proposed method requires much less effort of manual contouring from medical professionals such that an institutional specific model can be more easily established.
Shuai Wang 0003, Dong Nie, Liangqiong Qu, Yeqin Shao, Jun Lian, Qian Wang 0001, Dinggang Shen
IEEE Trans. Medical Imaging3
2019 Wavelet-based Semi-supervised Adversarial Learning for Synthesizing Realistic 7T from 3T MRI
Liangqiong Qu, Shuai Wang 0002, Pew-Thian Yap, Dinggang Shen
MICCAI (4)1
2018 Evaluation of shadow features
abstract
Shadow features such as colour ratio, texture, and chromaticity have proved to be quite effective in shadow detection. Many shadow detection methods have been proposed on the basis of different features. However, previous works for shadow detection mainly focus on designing an effective classifier for existing shadow features, but pay less attention on the analysis of shadow features themselves. The majority of studies simply report the final shadow detection results rather than make an evaluation on each feature. Readers often do not know which features are more effective or whether these shadow features are complementary. The following problems are still unsolved: the robustness of each feature, which feature plays the most important role in a detection method, and what is the best performance that current features can reach. The purpose of this study is to answer these questions, and the authors hope that this study can offer guidance for future shadow detection algorithms via the evaluation of frequently used shadow features. Several useful and interesting conclusions are obtained after conducting extensive comparison experiments on a large dataset.
Liangqiong Qu, Jiandong Tian, Huijie Fan, Yandong Tang
IET Comput. Vis.1
2017 DeshadowNet: A Multi-context Embedding Deep Network for Shadow Removal
abstract
Shadow removal is a challenging task as it requires the detection/annotation of shadows as well as semantic understanding of the scene. In this paper, we propose an automatic and end-to-end deep neural network (DeshadowNet) to tackle these problems in a unified manner. DeshadowNet is designed with a multi-context architecture, where the output shadow matte is predicted by embedding information from three different perspectives. The first global network extracts shadow features from a global view. Two levels of features are derived from the global network and transferred to two parallel networks. While one extracts the appearance of the input image, the other one involves semantic understanding for final prediction. These two complementary networks generate multi-context features to obtain the shadow matte with fine local details. To evaluate the performance of the proposed method, we construct the first large scale benchmark with 3088 image pairs. Extensive experiments on two publicly available benchmarks and our large-scale benchmark show that the proposed method performs favorably against several state-of-the-art methods.
Liangqiong Qu, Jiandong Tian, Shengfeng He, Yandong Tang, Rynson W. H. Lau
CVPR1
2017 A hand pose tracking benchmark from stereo matching
abstract
In this paper we establish a long-term 3D hand pose tracking benchmark1. It contains 18,000 stereo image pairs as well as the ground-truth 3D positions of palm and finger joints from different scenarios. Meanwhile, to accurately segment hand from stereo images, we propose a novel stereo-based hand segmentation and depth estimation algorithm specially tailored for hand tracking here. The experiments indicate the effectiveness of the proposed algorithm by demonstrating that its tracking performance is comparable to the use of an active depth sensor under various of challenging scenarios.
Jiawei Zhang 0002, Jianbo Jiao, Liangqiong Qu, Xiaobin Xu 0001, Qingxiong Yang
ICIP4
2017 A New Intrinsic-Lighting Color Space for Daytime Outdoor Images
abstract
Extracting or separating intrinsic information and illumination from natural images is crucial for better solving computer vision tasks. In this paper, we present a new illumination-based color space, the IL (intrinsic information and lighting level) space. Its first two channels represent 2D intrinsic information, and the third channel is for lighting levels. The IL color space has a one-to-one correspondence with the RGB color space. One valuable benefit of the IL color space is that illumination-related processing can be realized by directly operating on the lighting channel. As an example, based on the extracted lighting channel, we propose a new algorithm to estimate the intrinsic lighting level of an image such that the shadow-free color image and relighting series are obtained. In contrast to the existing color spaces for display or printing, the IL color space intuitively shows the information of reflectance and lighting levels for colors separately.
Zhi Han, Jiandong Tian, Liangqiong Qu, Yandong Tang
IEEE Trans. Image Process.3
2017 RGBD Salient Object Detection via Deep Fusion
abstract
Numerous efforts have been made to design various low-level saliency cues for RGBD saliency detection, such as color and depth contrast features as well as background and color compactness priors. However, how these low-level saliency cues interact with each other and how they can be effectively incorporated to generate a master saliency map remain challenging problems. In this paper, we design a new convolutional neural network (CNN) to automatically learn the interaction mechanism for RGBD salient object detection. In contrast to existing works, in which raw image pixels are fed directly to the CNN, the proposed method takes advantage of the knowledge obtained in traditional saliency detection by adopting various flexible and interpretable saliency feature vectors as inputs. This guides the CNN to learn a combination of existing features to predict saliency more effectively, which presents a less complex problem than operating on the pixels directly. We then integrate a superpixel-based Laplacian propagation framework with the trained CNN to extract a spatially consistent saliency map by exploiting the intrinsic structure of the input image. Extensive quantitative and qualitative experimental evaluations on three data sets demonstrate that the proposed method consistently outperforms the state-of-the-art methods.
Liangqiong Qu, Shengfeng He, Jiawei Zhang 0002, Jiandong Tian, Yandong Tang, Qingxiong Yang
IEEE Trans. Image Process.1
2016 New spectrum ratio properties and features for shadow detection
Jiandong Tian, Xiaojun Qi 0001, Liangqiong Qu, Yandong Tang
Pattern Recognit.3