VLDB 2026 Research / reviewers in the wild / expert
Lili Huang 0006
dblp:49/4253-6
· DBLP profile ↗
17ranked-venue papers
5as first author
17since 2021 · last 2026
0000-0002-8753-3170ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Morphology-aware hierarchical mixture of experts for Chest X-ray anatomy segmentation
Lili Huang 0006, Yuanjun He, Chenglong Li 0002, Jin Tang 0001 |
Eng. Appl. Artif. Intell. | 1 |
| 2026 | Medical report generation via knowledge distillation and medical keywords
Lili Huang 0006, Chenglong Li 0002, Jin Tang 0001 |
Neurocomputing | 1 |
| 2026 | SequencePAR: Understanding pedestrian attributes via a sequence generation paradigm
Jiandong Jin, Xiao Wang 0014, Yin Lin, Chenglong Li 0002, Lili Huang 0006, Aihua Zheng, Jin Tang 0001 |
Pattern Recognit. | 5 |
| 2026 | RCNet: Reliable Co-Training Network for Weakly Supervised Change DetectionabstractFully supervised change detection (CD) methods in remote sensing (RS) perform well but depend on costly and time-consuming pixel-level annotations, which are impractical to obtain at scale. Therefore, it is essential to develop annotation-efficient alternatives that can narrow the performance gap with fully supervised methods. To this end, we propose a novel weakly supervised CD framework, named RCNet, which employs dual networks to implement reliable co-training using image-level annotations. Our framework is grounded in multi-view learning of co-training and the localization ability of class activation mapping (CAM). In our approach, two sub-nets with the same architecture perform image-level change classification and pixel-level segmentation from different views. Although CAM roughly localizes changes, ambiguity and noise in its pseudo labels may cause confirmation bias, limiting performance. Our approach mitigates this bias by introducing a feature discrepancy loss to enable cross-supervision between two sub-nets. Meanwhile, CAM tends to highlight a single object, but RS images commonly contain many dense and small changed objects with complexity, resulting in decreased reliability of pseudo labels. Therefore, we present an IoU-based reliable pseudo label screening (RPLS) strategy, which minimizes the likelihood of changed areas being misidentified as unchanged, enhancing the reliability of changed information obtained. Besides, to further improve boundary fineness and internal integrity of changed areas, we incorporate an additional strong perturbation branch for each sub-net and develop a consistency regularization loss. Extensive experiments on three challenging RS image CD datasets demonstrate that our RCNet achieves competitive performance with image-level labels. The source code is available athttps://github.com/Youzhihui/RCNet. Zhi-Hui You, Sibao Chen 0001, Chris Ding, Lili Huang 0006, Jia-Xin Wang, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | CMCNet:Cross-directional morphology-aware convolution network for chest X-ray anatomy segmentation
Lili Huang 0006, Yuhan Feng, Chenglong Li 0002, Jin Tang 0001 |
Neurocomputing | 1 |
| 2025 | Instant pose extraction based on mask transformer for occluded person re-identification
Qing-Ling Shu, Sibao Chen 0001, Lili Huang 0006, Bin Luo 0001 |
Pattern Recognit. | 4 |
| 2025 | Multidimensional Remote Sensing Change Detection Based on Siamese Dual-Branch NetworksabstractDeep learning models, particularly convolutional neural networks (CNNs), have demonstrated outstanding feature learning capabilities, leading to remarkable performance in remote sensing change detection (RSCD) tasks. However, their most critical drawback lies in the lack of effective modeling of global information. This deficiency affects the model’s understanding of the overall context and structure of the entire image, making it difficult to distinguish between background and target areas, thereby leading to the erroneous identification of change regions. Second, features extracted by traditional backbone networks contain a significant amount of noise, resulting in blurred boundaries of changed objects. The challenge of effectively fusing detailed and semantic information to accurately differentiate pseudo changes remains significant. Furthermore, how to fully exploit multiscale information is another issue worth considering. We propose a full-scale multidimensional interaction network called SDSN, which enhances feature representation by leveraging both detail and semantic branches. Initially, bi-temporal images are processed by the encoder to extract coarse multiscale features. The semantic branch guides shallow-scale features, while the detail branch focuses on deep-scale features. Multikernel receptive module (MRM) aggregates global information. The detail branch utilizes a diversity variance module (DVM) and differential operations to generate refined change maps with noise reduction and background suppression. A multidimensional cross-perception module (MCM) guides the fusion of these change maps, establishing multidimensional dependencies to enrich feature representation. Compared with previous methods, SDSN demonstrates greater performance under complex environmental conditions, particularly noteworthy for its fewer parameters (4.03 M) and lower computational costs (7.94 G). The code is publicly available athttps://github.com/dpt000121/dpt. Li-Rong Shen, Sibao Chen 0001, Lili Huang 0006, Zhi-Hui You, Chris Ding, Jin Tang 0001, Bin Luo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Knowledge-Guided Cross-Modal Alignment and Progressive Fusion for Chest X-Ray Report GenerationabstractThe task of chest X-ray report generation, which aims to simulate the diagnosis process of doctors, has received widespread attention. Compared with the image caption task, chest X-ray report generation is more challenging since it needs to generate a longer and more accurate description of each diagnostic part in chest X-ray images. Most of existing works focus on how to extract better visual features or more accurate text expression based on existing reports. However, they ignore the interactions between visual and text modalities and are thus obviously not in line with human thinking. A small part of works explore the interactions of visual and text modalities, but data-driven learning of cross-modal information mapping can not break the semantic gap between different modalities. In this work, we propose a novel approach called Knowledge-guided Cross-modal Alignment and Progressive fusion (KCAP), which takes the knowledge words from a created medical knowledge dictionary as the bridge to guide the cross-modal feature alignment and fusion, for accurate chest X-ray report generation. In particular, we create the medical knowledge dictionary by extracting medical phrases from the training set and then selecting some phrases with substantive meanings as knowledge words based on their frequency of occurrence. Based on the knowledge words from the medical knowledge dictionary, the visual and text modalities are interacted by a mapping layer for the enhancement of the features of two modalities, and then the alignment fusion module is introduced to mitigate the semantic gap between visual and text modalities. To retain the important details of the original information, we design a progressive fusion scheme to integrate the advantages of both salient fused and original features to generate better medical reports. The experimental results on IU-Xray and MIMIC datasets demonstrate the effectiveness of the proposed KCAP. Lili Huang 0006, Pengcheng Jia, Chenglong Li 0002, Jin Tang 0001, Chuanfu Li |
IEEE Trans. Multim. | 1 |
| 2024 | Semantics Guided Disentangled GAN for Chest X-Ray Image Rib Segmentation
Lili Huang 0006, Dexin Ma, Chenglong Li 0002, Haifeng Zhao 0001, Jin Tang 0001, Chuanfu Li |
PRCV (14) | 1 |
| 2024 | CNN-Transformer with Stepped Distillation for Fine-Grained Visual Classification
Lili Huang 0006, Jin Tang 0001 |
PRCV (9) | 4 |
| 2024 | DEGANet: Road Extraction Using Dual-Branch Encoder With Gated Attention MechanismabstractAutomatic identification and extraction of roads from high-resolution remote sensing images (RSIs) are important in remote sensing and computer vision. Advancements in remote sensing technology have increased the information in images, making road extraction more challenging. Conventional convolutional methods have limitations, such as loss of spatial details and inadequate fusion of multiscale features. To address these challenges, the letter introduces a novel encoder-decoder architecture called dual-branch encoder with gated attention mechanism network (DEGANet), for extracting road networks in remote sensing image (RSI). First, we propose a multigated informative self-attention (MGSA) module that combines information from dual-branch encoders. By integrating the ResNet and the dynamic snake convolution (DSC) block, which conforms to road shapes, the module emphasizes slender structures similar to roads, thus enhancing the extraction of road features and focusing on capturing more road details. Second, we also introduce the cascade receptive field enhancement (CRFE) module, which optimizes both accuracy and computational complexity. This module combines various receptive field enhancement modules to improve capture long-range dependencies and spatial information perception. Comprehensive experiments conducted on various public remote sensing road datasets demonstrate that our network attains greater segmentation accuracy (intersection over union (IoU) and$F1$score) and connectivity [average path length similarity (APLS)], validating the effectiveness of our proposed method. Sibao Chen 0001, Lili Huang 0006, Chris Ding, Jin Tang 0001, Bin Luo 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2024 | Few-Shot Object Detection in Remote Sensing Images With Multiscale Spatial Selective AttentionabstractFew-shot object detection (FSOD) leverages limited labeled data and substantial unlabeled data for detection. However, these approaches mainly target natural images and ignore the spatial relationships and contextual information between objects in remote sensing images (RSIs). To overcome these challenges, this letter introduces a novel method for detecting few-shot objects in RSI. First, we propose a new attention, called multiscale spatial selective attention (MSSSA). This attention spatially selects feature maps from convolution kernels of different scales through spatial selection, focusing the network on the most relevant region of spatial context. Then, our proposed pixel-level feature extractor module (PLFEM) was used in the first stage of FSOD, providing pixel-level object position information to reduce false and missed detection. To evaluate the proposed method, we carry out comprehensive experiments on the DIOR dataset. The results show that the novel class mAP of our method reaches 38.2% in ten shots, an increase of 3.0% compared with the baseline, significantly improving the accuracy of FSOD in RSI. Yingnan Yu, Sibao Chen 0001, Lili Huang 0006, Jin Tang 0001, Bin Luo 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2023 | G2PL: Lexicon Enhanced Chinese Polyphone Disambiguation Using Bert Adapter with a New DatasetabstractPolyphone disambiguation is the core of grapheme-to-phoneme(G2P) module for the Chinese speech synthesis system. However, there is a lack of datasets and only one public for polyphone disambiguation. Moreover, due to the double long-tail distribution of polyphones, the ratio of pronunciation data for most polyphones is extremely unbalanced after sampling. To solve these problems, we propose a new dataset with 57,000 sentences from various domains by a new strategy for sampling. In addition, we propose the G2PL, which integrates word features into the bottom of BERT to assist in predicting the correct pronunciation of polyphone. In the experiment, we train the G2PL model to outperform other methods on our and public datasets. Our dataset, codes and user-friendly package are freely available. Haifeng Zhao 0001, Hongzhi Wan, Lili Huang 0006, Mingwei Cao |
ICASSP | 3 |
| 2023 | Road Extraction by Multiscale Deformable Transformer From Remote Sensing ImagesabstractRapid progress has been made in the research of high-resolution remote sensing road extraction tasks in the past years, but due to the diversity of road types and the complexity of road context, extracting the perfect road network is still fraught with difficulties and challenges. Many Convolutional Neural Networks (CNNs) based on encoder-decoder structures have demonstrated their effectiveness. Transformer’s self-attention mechanism shows more powerful performance than CNNs in modeling global feature dependencies. In this paper, we propose a Multi-scale Deformable Transformer Network (MDTNet) based on encoder-decoder structure to extract road networks from remote sensing images. The core of MDTNet is our proposed Multi-scale Deformable Self-Attention (MDSA) mechanism. MDSA can capture more comprehensive features than conventional self-attention. In addition, roads are not present in certain blocks of areas like other objects, but are interwoven throughout the image in such a long, linear fashion that information about certain road segments may be overlooked. To minimize residual errors in road segmentations, our MDSA incorporates a deformable design on feature maps, which effectively enhances the salience of road features relative to their surroundings. Extensive experiments on several public remote sensing road datasets show that our MDTNet achieves higher segmentation [F1 score and Intersection over Union (IoU)] and connectivity [Average Path Length Similarity (APLS)] accuracy, which verifies the effectiveness of our approach. Pengcheng Hu 0001, Sibao Chen 0001, Lili Huang 0006, Guizhou Wang, Jin Tang 0001, Bin Luo 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2023 | Dynamic Hypergraph Convolution and Recursive Gated Convolution Fusion Network for Hyperspectral Image ClassificationabstractRecently, convolutional neural network (CNN) and graph convolutional network (GCN) have been used widely for hyperspectral image (HSI) classification which, respectively, specialize in characterizing the local receptive feature and structure feature. However, the existing CNN-based methods cannot learn the higher-order interactions of different spectral bands. The GCN-based methods mostly used the fixed or simple graph model for feature learning. To solve the problems, we propose the dynamic hypergraph convolution and recursive gated convolution fusion network (DHCRGCFN) for HSI classification. To learn the hidden and important relations represented in the HSI data, the dynamic hypergraph convolution network (DHCN) is designed which dynamically updates the hypergraph model and captures the global spatial information of HSI. To efficiently model the high-order interactions among the high spectral dimension, the recursive gated convolution network (RGCN) is developed for progressively capturing the interactions of spectral feature. The features extracted by the two branches are adaptively fused to achieve the complementary advantages. Extensive experiments are conducted on two public HSI datasets to demonstrate the effectiveness of the proposed DHCRGCFN. Shumeng Xu, Jinpei Liu, Lili Huang 0006 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Semi-supervised Learning via Multiple Layer Graph Regularized PerceptionabstractRecently, Graph Neural Networks (GNNs) have made remarkable achievements in semi-supervised classification tasks. Nevertheless, GNNs usually rely on a specific graph convolution which has high computational complexity. To overcome this issue, recent works attempt to implicitly use adjacency matrix to guide message propagation in multi-layer perception (MLP) via neighboring contrastive loss. However, existing works accomplish implicit message passing only, without considering multi-order graph topology information. In this paper, we propose a novel method called Multiple Layer Graph Regularized Perception (MLGP). The main advantage of MLGP is to incorporate multi-order neighboring information into MLP. Further, inspired by gated mechanism, we design a linear gating to capture important features of nodes. More discriminant features can be obtained to alleviate over-smoothing. MLGP is more effective and more robust than existing works when dealing with large-scale graph data and noisy adjacency information. The comparative experiment results show that our model achieves better performance and strong robustness. Haiyun Xu, Lili Huang 0006, Bo Jiang 0002, Jin Tang 0001, Shaojie Zhang 0002 |
ICPR | 2 |
| 2022 | Multi-head collaborative learning for graph neural networks
Haiyun Xu, Bo Jiang 0002, Lili Huang 0006, Jin Tang 0001, Shaojie Zhang 0002 |
Neurocomputing | 3 |