Qiang Li 0048

dblp:72/872-48 · DBLP profile ↗
← Back
27ranked-venue papers
4as first author
27since 2021 · last 2026
0000-0001-7129-1456ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 12 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Computer networks · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 NeuroSketch: Bloom Filter-Based Sketch for Accurate Network Measurement via Neural Networks
abstract
In network measurement, learning-based sketch is a hot topic recently, which combines traditional sketches with machine learning techniques to improve the accuracy of sketches, while reducing the deployment overhead on switches. So far, most learning-based sketches estimate the sizes of either error-prone flows or all flows using machine learning models. These models take the sketch counter values of flows as features and their real sizes as labels for training. However, the flow size distribution is highly skewed, resulting in the effect that the flow sizes estimated by models are biased toward the sizes of mouse flows, severely underestimating elephant flows. To this end, a network measurement framework via back propagation neural network (BPNN) called NeuroSketch is proposed, which can directly estimate flow sizes and flow cardinality without identifying error-prone flows. Meanwhile, in order to provide effective features for BPNNs, a novel bloom filter-based sketch named BF-Sketch is proposed in this paper. BF-Sketch not only records the count values, but also the number of hash collisions in counters as a new feature, which can efficiently reduce the underestimation of elephant flows by machine learning models. The experimental results show that NeuroSketch reduces the average absolute error (AAE) of flow size estimation by 65%, and relative errors of flow cardinality estimation by 72.23%, compared with learning-based sketches. Moreover, BF-Sketch is implemented on OVS platform and P4-programmable switch to justify its feasible deployment in commodity software and hardware switches.
Jindian Liu, Zhuo Li 0009, Hao Xun, Yu Zhang 0036, Peng Luo 0004, Qiang Li 0048
IEEE Trans. Netw.6
2025 FreeInsert: Disentangled Text-Guided Object Insertion in 3D Gaussian Scene without Spatial Priors
Weijie Wang 0002, Qiang Li 0048, Nicu Sebe, Bruno Lepri, Weizhi Nie
ACM Multimedia3
2025 Causal inference model for accurate medical diagnosis in Coronary Artery Bypass Graft operation
Qiyi Zhang, Wei Zhang 0390, Qiang Li 0048, Yunpeng Bai, Weizhi Nie, Keliang Xie
Artif. Intell. Medicine3
2025 The interpretable deep learning framework and validation for seizure detection in pediatric electroencephalography: An improved accuracy and performance analysis
Qiang Li 0048, Ai-Ping Yang, Ming-Lang Tseng
Artif. Intell. Medicine3
2025 Temporal and Spatial Analysis in Early Sepsis Prediction via Causal Disentanglements
abstract
Sepsis is one of the main causes of death in ICU patients, and accurate and stable early prediction is essential for clinical intervention. Existing methods mostly rely on traditional time series models (e.g., LSTM, Transformer) or clinical scoring criteria (e.g., SOFA, qSOFA), but face two major challenges: 1) spurious correlations in the data affect the robustness of the model; 2) Lack of modeling the underlying causal relationships in the data space. We propose a Serialized Causal Disentanglement Model (SCDM) that decouples latent variables into sepsis-related factors ($u$), other disease-related factors ($v$), and irrelevant confounders ($s$). Based on the MIMIC-IV v2.2 dataset (3,511 positive samples and 17,538 negative samples), SCDM took patient clinical indicators, personal information, and clinical notes as input, and achieved an AUC of 0.765-0.928in the prediction task 48 to 0 hours before the onset of sepsis. The performance is significantly better than the baseline models (e.g., Transformer's 0.662-0.910, MGP-AttTCN's 0.692-0.913). Experiments show that optimizing the time window (5 hours of continuous observation) and variable selection (45 key indicators) can improve the performance of the model. The effectiveness of causal unwinding is verified by the visualization of Grad CAM and t-SNE, key clinical indicators such as platelet count, lactic acid, and respiratory rate are further identified to provide interpretable decision support for doctors. Our study provides a high-precision and interpretable causal disentanglement framework for early prediction of sepsis, which is expected to promote the development of intelligent diagnosis and treatment in the ICU.
Qiang Li 0048, Weizhi Nie, He Jiao, Anan Liu
IEEE Trans. Knowl. Data Eng.1
2025 Structure Causal Models and LLMs Integration in Medical Visual Question Answering
abstract
Medical Visual Question Answering (MedVQA) aims to answer medical questions according to medical images. However, the complexity of medical data leads to confounders that are difficult to observe, so bias between images and questions is inevitable. Such cross-modal bias makes it challenging to infer medically meaningful answers. In this work, we propose a causal inference framework for the MedVQA task, which effectively eliminates the relative confounding effect between the image and the question to ensure the precision of the question-answering (QA) session. We are the first to introduce a novel causal graph structure that represents the interaction between visual and textual elements, explicitly capturing how different questions influence visual features. During optimization, we apply the mutual information to discover spurious correlations and propose a multi-variable resampling front-door adjustment method to eliminate the relative confounding effect, which aims to align features based on their true causal relevance to the question-answering task. In addition, we also introduce a prompt strategy that combines multiple prompt forms to improve the model's ability to understand complex medical data and answer accurately. Extensive experiments on three MedVQA datasets demonstrate that 1) our method significantly improves the accuracy of MedVQA, and 2) our method achieves true causal correlations in the face of complex medical data.
Zibo Xu, Qiang Li 0048, Weizhi Nie, Weijie Wang 0002, Anan Liu
IEEE Trans. Medical Imaging2
2025 DSDP: Real-Time Asymmetric Dual-Stream Instance Segmentation Embedding Depth-Predictive Architecture for Enhanced Scene Understanding
abstract
Instance segmentation can help vehicles or robots enhance their understanding of a scene through the pixel-level segmentation of different objects. However, occlusion and boundary blur, especially in cases with similar colors or textures, are still challenges encountered in real-time robust segmentation tasks. To segment a complete instance boundary, the existing 2D approaches fuse local and abstract semantic features derived from the color domain, which leads to homogeneous semantic information, and efficiently separating different objects is difficult in some cases. To address these complicated scenes, inspired by a human prediction processing strategy, where “the brain fills in missing information in advance to help make better decisions”, this study proposes a real-time asymmetric dual-stream instance segmentation algorithm embedding a depth-predictive architecture that provides the covisible depth information of objects. Furthermore, a cross-domain data fusion method and an enhancement-decoupling loss are designed to complement RGB data by utilizing the rich foreground and boundary details of the predicted depth map. In addition, our model can be fine-tuned to integrate it with real depth domain data provided by different input devices. Extensive experiments conducted on the COCO, OCHuman and CityScapes datasets demonstrate the effectiveness of our method. We further deployed our DSDP method on a UAV platform for validation purposes and qualitatively confirmed its validity.
Qiang Li 0048, Weizhi Nie, Jing Liu 0002, Jingjing Geng, Yongtao Ma
IEEE Trans. Multim.2
2025 DCDL: Dual Causal Disentangled Learning for Zero-Shot Sketch-Based Image Retrieval
abstract
Zero-shot sketch-based image retrieval (ZS-SBIR) is a challenging task that hinges on overcoming the cross-domain differences between sketches and images. Previous methods primarily address cross-domain differences by creating a common embedding space, improving final retrieval results. However, most previous approaches have overlooked a critical aspect: sketch-based image retrieval task actually requires only the cross-domain invariant information relevant to the retrieval. Irrelevant information (such as posture, expression, background, and specificity) may detract from retrieval accuracy. In addition, most previous methods perform well on traditional SBIR datasets but lack corresponding research on generalization and extensibility in the face of more diverse and complex data. To address these issues, we propose a Dual Causal Disentangled Learning (DCDL) for ZS-SBIR. This approach can mitigate the negative impact of irrelevant features by separating retrieval-relevant features in the latent variable space. Specifically, we constructed a causal disentanglement model using two Variational Autoencoders (VAE), each applied to the sketch and image domains, to obtain disentangled variables with exchangeable attributes. Our framework effectively integrates causal intervention with disentangled representation learning, enabling a clearer separation of cross-domain retrieval-relevant and intra-class irrelevant features, which can be recombined into new reconstructed samples. Concurrently, we designed a Dual Alignment Module (DAM), leveraging the accurate and comprehensive semantic features provided by a text encoder pre-trained on large-scale datasets to supplement semantic associations and align disentangled retrieval-relevant features. The Dual Alignment Module enhances the model's ability to generalize across diverse datasets by effectively aligning retrieval-relevant information from different domains. Extensive experiments demonstrate that our method achieves state-of-the-art (SOTA) performance on the Sketchy and TU–Berlin datasets. Additionally, more experiments on larger scale dataset QuickDraw, fine-grained datasets, Shoe-V2 and Chair-V2, as well as an inter-dataset further validate the generalization and extensibility of DCDL.
Qiang Li 0048, Wei Zhang 0390, Shaojin Bai, Weizhi Nie, Anan Liu
IEEE Trans. Multim.1
2024 Causal Intervention for Brain Tumor Segmentation
Hengxin Liu, Qiang Li 0048, Weizhi Nie, Zibo Xu, Anan Liu
MICCAI (9)2
2024 A deep convolutional neural network for the automatic segmentation of glioblastoma brain tumor: Joint spatial pyramid module and attention mechanism network
Hengxin Liu, Jingteng Huang, Qiang Li 0048, Ming-Lang Tseng
Artif. Intell. Medicine3
2024 An effective and accurate flow size measurement using funnel-shaped sketch
Jindian Liu, Zhuo Li 0009, Huipeng Du, Haodong Zhou, Leyang Li, Yi An, Yu Zhang 0036, Qiang Li 0048
Comput. Networks9
2024 Multi-modal fusion network guided by prior knowledge for 3D CAD model recognition
Qiang Li 0048, Zibo Xu, Shaojin Bai, Weizhi Nie, Anan Liu
Neurocomputing1
2024 Early prediction of sepsis using chatGPT-generated summaries and structured data
Qiang Li 0048, Hanbo Ma, Dan Song 0006, Yunpeng Bai, Keliang Xie
Multim. Tools Appl.1
2024 Hybrid QUS Radiomics: A Multimodal-Integrated Quantitative Ultrasound Radiomics for Assessing Ambulatory Function in Duchenne Muscular Dystrophy
abstract
BACKGROUND: Duchenne muscular dystrophy (DMD) is a neuromuscular disorder that affects ambulatory function. Quantitative ultrasound (QUS) imaging, utilizing envelope statistics, has proven effective in diagnosing DMD. Radiomics enables the extraction of detailed features from QUS images. This study further proposes a hybrid QUS radiomics and explores its value in characterizing DMD. METHODS: Patients (n = 85) underwent ultrasound examinations of gastrocnemius through Nakagami, homodyned K (HK), and information entropy imaging. The hybrid QUS radiomics extracted, selected, and integrated the retained features derived from each QUS image for classification of ambulatory function using support vector machine. Nested five fold cross-validation of the data was conducted, with the rotational process repeated 50 times. The performance was assessed by averaging the areas under the receiver operating characteristic curve (AUROC). RESULTS: Radiomics enhanced the average AUROC of B-scan, Nakagami, HK, and entropy imaging to 0.790, 0.911, 0.869, and 0.890, respectively. By contrast, the hybrid QUS radiomics using HK and entropy images for diagnosing ambulatory function in DMD patients achieved a superior average AUROC of 0.971 (p < 0.001 compared with conventional radiomics analysis). CONCLUSIONS: The proposed hybrid QUS radiomics incorporates microstructure-related backscattering information from various envelope statistics models to effectively enhance the performance of DMD assessment.
Qiang Li 0048, Chia-Wei Lin, Jeng-Yi Shieh, Wen-Chin Weng, Po-Hsiang Tsui
IEEE J. Biomed. Health Informatics2
2024 Frequency-Domain Robust PCA for Real-Time Monitoring of HIFU Treatment
abstract
High intensity focused ultrasound (HIFU) is a thriving non-invasive technique for thermal ablation of tumors, but significant challenges remain in its real-time monitoring with medical imaging. Ultrasound imaging is one of the main imaging modalities for monitoring HIFU surgery in organs other than the brain, mainly due to its good temporal resolution. However, strong acoustic interference from HIFU irradiation severely obscures the B-mode images and compromises the monitoring. To address this problem, we proposed a frequency-domain robust principal component analysis (FRPCA) method to separate the HIFU interference from the contaminated B-mode images. Ex-vivo and in-vivo experiments were conducted to validate the proposed method based on a clinical HIFU therapy system combined with an ultrasound imaging platform. The performance of the FRPCA method was compared with the conventional notch filtering method. Results demonstrated that the FRPCA method can effectively remove HIFU interference from the B-mode images, which allowed HIFU-induced grayscale changes at the focal region to be recovered. Compared to notch-filtered images, the FRPCA-processed images showed an 8.9% improvement in terms of the structural similarity (SSIM) index to the uncontaminated B-mode images. These findings demonstrate that the FRPCA method presents an effective signal processing framework to remove the strong HIFU acoustic interference, obtains better dynamic visualization in monitoring the HIFU irradiation process, and offers great potential to improve the efficacy and safety of HIFU treatment and other focused ultrasound related applications.
Qiang Li 0048, Meng-Xing Tang, Zhibiao Wang, Po-Hsiang Tsui
IEEE Trans. Medical Imaging2
2024 Commonsense-Guided Semantic and Relational Consistencies for Image-Text Retrieval
abstract
Image-text retrieval, as a fundamental task in the cross-modal field, aims to explore the relationship between visual and textual modalities. Recent methods address this task only by learning the conceptual and syntactical correspondences between cross-modal fragments, but these correspondences inevitably contain noise without considering external knowledge. To solve this issue, we propose a novelCommonsense-GuidedSemantic andRelationalConsistencies (CSRC) for image-text retrieval that can simultaneously expand the semantics and relations to reduce the cross-modal differences under the assumption that the semantics and relations of the true image-text pair should be consistent between two modalities. Specifically, we first explore commonsense knowledge to expand the specific concepts for visual and textual graphs and optimize the semantic consistency by minimizing the differences in cross-modal semantic importance. Then, we extend the same relations for cross-modal concept pairs with semantic consistency, which serves to implement relational consistency. After that, we combine external commonsense knowledge with internal correlation to enhance concept representation and further optimize relational consistency by regularizing the importance differences between association-enhanced concepts. Extensive experimental results on two popular image-text retrieval datasets demonstrate the effectiveness of our proposed method.
Wenhui Li 0001, Qiang Li 0048, Xuanya Li, Anan Liu
IEEE Trans. Multim.3
2023 External Knowledge Dynamic Modeling for Image-text Retrieval
abstract
Image-text retrieval is a fundamental branch in cross-modal retrieval. The core is to explore the semantic correspondence to align relevant image-text pairs. Some existing methods rely on global semantics and co-occurrence frequency to design knowledge introduction patterns for consistent representations. However, they lack flexibility due to the limitations of fixed information and empirical feedback. To address these issues, we develop an External Knowledge Dynamic Modeling~(EKDM) architecture based on the filtering mechanism, which dynamically explores different knowledge towards varied image-text pairs. Specially, we first capture abundant concepts and relationships from external knowledge to construct visual and textual corpus sets. Then, we progressively explores concepts related to images and texts by dynamic global representations. To endow the model with the capability of relationship decision, we integrate the variable spatial locations between objects for association exploration. Since the filtering mechanism is conditioned on dynamic semantics and variable spatial locations, our model can dynamically model different knowledge for different image-text pairs. Extensive experimental results on two benchmark datasets demonstrate the effectiveness of our proposed method.
Qiang Li 0048, Wenhui Li 0001, Min Liu 0008, Xuanya Li, Anan Liu
ACM Multimedia2
2023 Principal views selection based on growing graph convolution network for multi-view 3D model recognition
Qi Liang 0004, Qiang Li 0048, Weizhi Nie, Yuting Su 0001
Appl. Intell.2
2023 Multiscale lightweight 3D segmentation algorithm with attention mechanism: Brain tumor image segmentation
Hengxin Liu, Guoqiang Huo, Qiang Li 0048, Ming-Lang Tseng
Expert Syst. Appl.3
2023 Unsupervised Cross-Media Graph Convolutional Network for 2D Image-Based 3D Model Retrieval
abstract
With the rapid development of 3D construction technology, 3D models have been implemented in many applications. In particular, the fields of virtual and augmented reality have created a considerable demand for rapid access to large sets of 3D models in recent years. An effective method for addressing the demand is to search 3D models based on 2D images because 2D images can be easily captured by smartphones or other lightweight vision sensors. In this paper, we propose a novel unsupervised cross-media graph convolutional network (UCM-GCN) for 3D model retrieval based on 2D images. Here, we render views from 3D models to construct a graph model based on 3D model structural information. Then, we utilize the 2D image's visual information to bridge the gap between cross-modality data. Then, the proposed UCM-GCN is utilized to update the feature vector of the 2D image and the 3D model. Here, we introduce correlation loss to mitigate the distribution discrepancy across different modalities, which can fully consider the structural and visual similarities between the 2D image and 3D model to embed the final different modalities into the same feature space. To demonstrate the performance of our approach, we conducted a series of experiments on the MI3DOR dataset, which is utilized in SHREC19. We also compared it with other similar methods on the 3D-FUTURE dataset. The experimental results demonstrate the superiority of our proposed method over state-of-the-art methods.
Qi Liang 0004, Qiang Li 0048, Weizhi Nie, Anan Liu
IEEE Trans. Multim.2
2023 Semantic Completion and Filtration for Image-Text Retrieval
abstract
Image–text retrieval is a vital task in computer vision and has received growing attention, since it connects cross-modality data. It comes with the critical challenges of learning unified representations and eliminating the large gap between visual and textual domains. Over the past few decades, although many works have made significant progress in image–text retrieval, they are still confronted with the challenge of incomplete text descriptions of images, i.e., how to fully learn the correlations between relevant region–word pairs with semantic diversity. In this article, we propose a novel semantic completion and filtration (SCAF) method to alleviate the above issue. Specifically, the text semantic completion module is presented to generate a complete semantic description of an image using multi-view text descriptions, guiding the model to explore the correlations of relevant region–word pairs fully. Meanwhile, the adaptive structural semantic matching module is presented to filter irrelevant region–word pairs by considering the relevance score of each region–word pair, which facilitates the model to focus on learning the relevance of matching pairs. Extensive experiments show that our SCAF outperforms the existing methods on Flickr30K and MSCOCO datasets, which demonstrates the superiority of our proposed method.
Qiang Li 0048, Wenhui Li 0001, Xuanya Li, Rui Wang 0110, Anan Liu
ACM Trans. Multim. Comput. Commun. Appl.2
2022 LP-GAN: Learning perturbations based on generative adversarial networks for point cloud adversarial attacks
Qi Liang 0004, Qiang Li 0048
Image Vis. Comput.2
2022 PAGN: perturbation adaption generation network for point cloud adversarial defense
Qi Liang 0004, Qiang Li 0048, Weizhi Nie, Anan Liu
Multim. Syst.2
2022 LD-GAN: Learning perturbations for adversarial defense based on GAN structure
Qi Liang 0004, Qiang Li 0048, Weizhi Nie
Signal Process. Image Commun.2
2022 Dual-Level Representation Enhancement on Characteristic and Context for Image-Text Retrieval
abstract
Image-text retrieval is a fundamental and vital task in multi-media retrieval and has received growing attention since it connects heterogeneous data. Previous methods that perform well on image-text retrieval mainly focus on the interaction between image regions and text words. But these approaches lack joint exploration of characteristics and contexts of regions and words, which will cause semantic confusion of similar objects and loss of contextual understanding. To address these issues, a dual-level representation enhancement network (DREN) is proposed to strength the characteristic and contextual representations by innovative block-level and instance-level representation enhancement modules, respectively. The block-level module focuses on mining the potential relations between multiple blocks within each instance representation, while the instance-level module concentrates on learning the contextual relations between different instances. To facilitate the accurate matching of image-text pairs, we propose the graph correlation inference and weighted adaptive filtering to conduct the local and global matching between image-text pairs. Extensive experiments on two challenging datasets (i.e., Flickr30K and MSCOCO) verify the superiority of our method for image-text retrieval.
Qiang Li 0048, Wenhui Li 0001, Xuanya Li, Anan Liu
IEEE Trans. Circuits Syst. Video Technol.2
2021 Disease Correlation Enhanced Attention Network for ICD Coding
abstract
The automatic ICD coding task is to assign the most relevant International Classification of Disease diagnosis codes (ICD codes) to the patient’s clinical text, which is defined as a multi-label text classification problem. In practical applications, the clinical text is lengthy and the distribution of ICD codes is skewed with a long tail, making it difficult to predict the labels for rare disease. To tackle the problem, we propose a new deep learning framework which combines diseases correlation and multi-attention mechanism into Bidirectional Recurrent Neural Network. In particular, we apply the Graph Convolutional Network to capture the disease relevance between ICD codes, which can be highly valuable for those diseases with limited data. In addition, we leverage multi-head attention and label-wise attention to encode clinical text, thus our model can learn label-specific representations and consider disease correlation in the same time. Extensive experiments demonstrate that our proposed model is effective. On the MIMIC-III-full dataset, our model outperformed the state-of-the-art model in 5 out of 7 evaluation metrics. On the MIMIC-III-50 dataset, our model outperformed the state-of-the-art model in 4 out of 7 evaluation metrics.
Ping Gu, Qiang Li 0048, Jiangxing Wang
BIBM3
2021 MHFP: Multi-view based hierarchical fusion pooling method for 3D shape recognition
Qi Liang 0004, Qiang Li 0048, Lihu Zhang, Haixiao Mi, Weizhi Nie, Xuanya Li
Pattern Recognit. Lett.2