VLDB 2026 Research / reviewers in the wild / expert
Lixia Xue
dblp:148/6988
· DBLP profile ↗
38ranked-venue papers
5as first author
23since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 5 first-author · 12 since 2021Artificial intelligence and machine learning · 11 · 10 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CCANet : A Cognition-Inspired Framework for Few-Shot Segmentation from Category-Agnostic to Category-AwareabstractDue to the limited annotated data, semantic segmentation faces significant challenges in few shot generalization. Current methods combining prototype learning and affinity learning suffer from static prototypes failing to capture intra-class diversity. Even if these issues are addressed, they still encounter the critical problem of ineffective integration between global information acquired through prototype learning and local information obtained via affinity learning, which ultimately hinders their performance in few-shot segmentation scenarios. Inspired by human cognitive mechanisms, we propose CCANet - a two-stage framework that progressively combines prototype learning with affinity learning through "category-agnostic to category-aware" knowledge evolution. This innovative architecture effectively resolves the aforementioned limitations while achieving systematic knowledge integration. The Global category-aware Module employs category-agnostic localization and dynamic prototype evolution, leveraging category-agnostic global context modeling and history-driven prototype adaptation to enhance cross-category alignment and noise robustness. The Local Feature Enrichment Module introduces a self-attention-gated fusion paradigm, coordinating multi-scale global-local correlations to effectively guide the affinity modeling process across multi-scale contexts. Extensive experiments demonstrate state-of-the-art performance on PASCAL-5i and COCO-20i, validating the effectiveness of our approach. Our work bridges cognitive principles with computational models, offering new insights for few-shot segmentation. Ronggui Wang, Lixia Xue, Juan Yang 0001 |
MMAsia | 3 |
| 2025 | Multi-Directional Transformer Image Super-Resolution Network Based on Information EnhancementabstractABSTRACT With the advancement of deep learning, single‐image super‐resolution (SISR) has achieved significant progress. Recently, vision transformer‐based super‐resolution models have demonstrated remarkable performance; however, their high computational cost hinders their practical application. In this paper, we introduce a lightweight transformer‐based super‐resolution model termed information‐enhanced efficient multi‐directional transformer(IEMT). The model employs a dual‐branch architecture that integrates the strengths of both convolutional neural network (CNN) and transformer networks. The proposed high‐frequency extraction block (HEB) effectively captures high‐frequency information from the enhanced image. Furthermore, a multi‐directional attention mechanism is incorporated into the transformer branch to comprehensively learn latent features and details, thereby enhancing reconstruction quality. For attention computation, we propose a dynamic parameter‐sharing mechanism that adaptively adjusts parameter sharing based on local image features, significantly reducing the model's parameter count. Experimental results demonstrate that the proposed IEMT achieves superior performance on five benchmark datasets, with a significantly reduced parameter count, computational complexity, and memory usage. Ronggui Wang, Lixia Xue |
IET Image Process. | 4 |
| 2025 | BMFNet: Bidirectional Multimodal Fusion Network for image captioning
Lixia Xue, ZiQian Jin, Ronggui Wang, Juan Yang 0001 |
Multim. Syst. | 1 |
| 2025 | Dual-stage pixel transformer with enhanced visual context for image captioning
Juan Yang 0001, Anbo Liu, Ronggui Wang, Lixia Xue |
Multim. Syst. | 4 |
| 2025 | VTIENet: visual-text information enhancement network for image captioning
Juan Yang 0001, Yuhang Wei, Ronggui Wang, Lixia Xue |
Multim. Syst. | 4 |
| 2025 | Boosting few-shot learning via selective patch embedding by comprehensive sample analysis
Juan Yang 0001, Ronggui Wang, Lixia Xue |
Mach. Vis. Appl. | 4 |
| 2025 | A lightweight dual-student mean teacher semi-supervised semantic segmentation method for skin lesions
Jindian Lu, Yongyong Chen, Yuanhaonan Deng, Binghui Zhao, Lixia Xue |
Neural Networks | 7 |
| 2025 | Adpl: attentive dual-modality prompt learning for vision-language understanding
Zhiyong Deng, Ronggui Wang, Lixia Xue, Juan Yang 0001 |
J. Supercomput. | 3 |
| 2025 | Adaptive sparse triple convolutional attention for enhanced visual question answering
Ronggui Wang, Juan Yang 0001, Lixia Xue |
Vis. Comput. | 4 |
| 2025 | Modular dual-stream visual fusion network for visual question answering
Lixia Xue, Ronggui Wang, Juan Yang 0001 |
Vis. Comput. | 1 |
| 2025 | Semantically Enhanced Dual Visual Fusion Transformer for accurate image captioning
Ronggui Wang, Lixia Xue, Jiaping Zhang |
Vis. Comput. | 4 |
| 2024 | Multisource hierarchical neural network for knowledge graph embedding
Ronggui Wang, Lixia Xue, Juan Yang 0001 |
Expert Syst. Appl. | 3 |
| 2024 | Sample-Adaptive Classification Inference NetworkabstractAbstract Existing pre-trained models have yielded promising results in terms of computational time reduction. However, these models only focus on pruning simple sentences or less salient words, while neglecting the treatment of relatively complex sentences. It is frequently these sentences that cause the loss of model accuracy. This shows that the adaptation of the existing models is one-sided. To address this issue, in this paper, we propose a sample-adaptive training and inference model. Specifically, complex samples are extracted from the training datasets and a dedicated data augmentation module is trained to extract global and local semantic information of complex samples. During inference, simple samples can exit the model via the Sample Adaptive Exit Mechanism, Normal samples pass through the whole backbone model before inference, while complex samples are processed by the Characteristic Enhancement Module after passing through the backbone model. In this way, all samples are processed adaptively. Our extensive experiments on classification tasks datasets in the field of Natural Language Processing demonstrate that our method enhances model accuracy and reduces model inference time for multiple datasets. Moreover, our method is transferable and can be applied to multiple pre-trained models. Juan Yang 0001, Guanghong Zhou, Ronggui Wang, Lixia Xue |
Neural Process. Lett. | 4 |
| 2023 | FPIseg: Iterative segmentation network based on feature pyramid for few-shot segmentationabstractAbstract Few‐shot segmentation (FSS) enables rapid adaptation to the segmentation task of unseen‐classes object based on a few labelled support samples. Currently, the focal point of research in the FSS field is to align features between support and query images, aiming to improve the segmentation performance. However, most existing FSS methods implement such support/query alignment by solely leveraging middle‐level feature for generalization, ignoring the category semantic information contained in high‐level feature, while pooling operation inevitably lose spatial information of the feature. To alleviate these issues, the authors propose the Iterative Segmentation Network Based on Feature Pyramid (FPIseg), which mainly consists of three modules: Feature Pyramid Fusion Module (FPFM), Region Feature Enhancement Module (RFEM), and Iterative Optimization Segmentation Module (IOSM). Firstly, FPFM fully utilizes the foreground information from the support image to implement support/query alignment under multi‐scale, multi‐level semantic backgrounds. Secondly, RFEM enhances the foreground detail information of aligned feature to improve generalization ability. Finally, ISOM iteratively segments the query image to optimize the prediction result and improve segmentation performance. Extensive experiments on the PASCAL‐5 i and COCO‐20 i datasets show that FPIseg achieves considerable segmentation performance under both 1‐shot and 5‐shot settings. Ronggui Wang, Juan Yang 0001, Lixia Xue |
IET Image Process. | 4 |
| 2023 | MFFN: Multi-path feedback fusion network for lightweight image super resolutionabstractAbstract Recently, convolutional neural network (CNN) has shown great power in single image super resolution (SISR) reconstruction, and achieving significant improvements over traditional methods. Despite the great success of these CNN‐based methods, direct application of these methods to some edge devices is impractical due to the large computational overhead required. To address this problem, a novel, lightweight SISR network focusing on speed and accuracy, called the multi‐path feedback fusion network (MFFN), has been designed in this paper. Specifically, in order to extract features more effectively, the authors propose a novel fusion attention feedback block (FAFB) as the main building block of MFFN. The FAFB consists of a backbone branch and several hierarchical branches. The backbone branch is composed of stacked enhanced pixel attentional blocks (EPAB), which are responsible for incremental deep feature learning on the feature map. And the hierarchical branches are responsible for extracting feature maps with different sizes of receptive fields and fusing these feature maps with the features extracted from the trunk branches to achieve multi‐scale feature learning, which the authors refer to this design as the multi‐scale fusion block (MSFB). Extra enhancement information (EIE) is added to each EPAB input, which enables the backbone branch to learn more effectively. On the other hand, the outputs of the cascade branches are further complemented by an additional feedback fusion enhancement block (FFEB) before being fused with the output of the trunk branches to achieve more comprehensive and accurate feature learning. Numerous experiments have shown that MFFN achieves higher accuracy than other state‐of‐the‐art methods on benchmark test sets. Lixia Xue, JunHui Shen, Ronggui Wang, Juan Yang 0001 |
IET Image Process. | 1 |
| 2022 | Structural context-based knowledge graph embedding for link prediction
Ronggui Wang, Juan Yang 0001, Lixia Xue |
Neurocomputing | 4 |
| 2022 | Multiview feature augmented neural network for knowledge graph embedding
Ronggui Wang, Lixia Xue, Juan Yang 0001 |
Knowl. Based Syst. | 3 |
| 2022 | Knowledge graph embedding by reflection transformation
Ronggui Wang, Juan Yang 0001, Lixia Xue |
Knowl. Based Syst. | 4 |
| 2021 | Multi-domain few-shot image recognition with knowledge transfer
Ronggui Wang, Juan Yang 0001, Lixia Xue, Min Hu 0010 |
Neurocomputing | 4 |
| 2021 | Person Re-identification with Global-Local Background_bias Net
Yuxiu Gong, Ronggui Wang, Juan Yang 0001, Lixia Xue, Min Hu 0010 |
J. Vis. Commun. Image Represent. | 4 |
| 2021 | Kernel multi-attention neural network for knowledge graph embedding
Ronggui Wang, Juan Yang 0001, Lixia Xue |
Knowl. Based Syst. | 4 |
| 2021 | Knowledge graph embedding by translating in time domain space for link prediction
Ronggui Wang, Juan Yang 0001, Lixia Xue |
Knowl. Based Syst. | 4 |
| 2021 | Multi-scale feature self-enhancement network for few-shot learning
Ronggui Wang, Juan Yang 0001, Lixia Xue |
Multim. Tools Appl. | 4 |
| 2020 | Multi-scale feature network for few-shot learning
Mengya Han, Ronggui Wang, Juan Yang 0001, Lixia Xue, Min Hu 0010 |
Multim. Tools Appl. | 4 |
| 2020 | Learning semantic dependencies with channel correlation for multi-label classification
Lixia Xue, Ronggui Wang, Juan Yang 0001, Min Hu 0010 |
Vis. Comput. | 1 |
| 2019 | Hierarchical deep transfer learning for fine-grained categorization on micro datasets
Ronggui Wang, Xuchen Yao, Juan Yang 0001, Lixia Xue, Min Hu 0010 |
J. Vis. Commun. Image Represent. | 4 |
| 2019 | Enhanced two-phase residual network for single image super-resolution
Juan Yang 0001, Ronggui Wang, Lixia Xue, Min Hu 0010 |
J. Vis. Commun. Image Represent. | 4 |
| 2019 | Principal characteristic networks for few-shot learning
Ronggui Wang, Juan Yang 0001, Lixia Xue, Min Hu 0010 |
J. Vis. Commun. Image Represent. | 4 |
| 2019 | Multipath feedforward network for single image super-resolution
Mingyu Shen, Ronggui Wang, Juan Yang 0001, Lixia Xue, Min Hu 0010 |
Multim. Tools Appl. | 5 |
| 2019 | Multi-label image classification with recurrently learning semantic dependencies
Ronggui Wang, Juan Yang 0001, Lixia Xue, Min Hu 0010 |
Vis. Comput. | 4 |
| 2018 | Robust object tracking via superpixels and keypoints
Mingyu Shen, Ronggui Wang, Juan Yang 0001, Lixia Xue, Min Hu 0010 |
Multim. Tools Appl. | 5 |
| 2018 | Super-resolution via supervised classification and independent dictionary training
Ronggui Wang, Juan Yang 0001, Lixia Xue, Min Hu 0010 |
Multim. Tools Appl. | 4 |
| 2018 | Low - resolution vehicle recognition based on deep feature fusion
Lixia Xue, Ronggui Wang, Juan Yang 0001, Min Hu 0010 |
Multim. Tools Appl. | 1 |
| 2017 | Large scale automatic image annotation based on convolutional neural network
Ronggui Wang, Yunfei Xie, Juan Yang 0001, Lixia Xue, Min Hu 0010, Qingyang Zhang 0002 |
J. Vis. Commun. Image Represent. | 4 |
| 2016 | A novel method for image classification based on bag of visual words
Ronggui Wang, Juan Yang 0001, Lixia Xue |
J. Vis. Commun. Image Represent. | 4 |
| 2015 | An audio-visual human attention analysis approach to abrupt change detection in videos
Minglong Song, Lixia Xue, Xiaoxue Chen |
Signal Process. | 3 |
| 2014 | The analytic hierarchy process-based optimal forwarder selection in multi-hop broadcasting scheme for vehicular safetyabstractVehicular technologies have been recently introduced to highway safety applications and caught much attention of governments and management authorities. The vehicular ad hoc network (VANET) uses cars as mobile nodes to form a mobile network. VANET offers potential and promising technical solutions to vehicular safety. Multi-hop broadcasting schemes are particularly preferred methods to transmit time-sensitive safety warning information to potentially influenced vehicles. However, broadcasting may lead to frequent contention and serious collision, thus causing broadcast storms. In order to alleviate broadcast storm and promptly disseminate safety warnings, the paper proposes an optimal forwarder selection scheme to minimize the number of rebroadcasting nodes and guarantee fast and efficient safety warning information dissemination. The proposed forwarder selection scheme is based on the analytic hierarchy process (AHP). The criteria include inter-vehicular lateral and longitudinal distances, vehicular communication ranges, and vehicles covered within the communication range of previous forwarder vehicles. The AHP-based forwarder selection model is established and model's validity is verified mathematically in this paper. Dongxiu Ou, Lixia Xue, Tuo Shen |
Intelligent Vehicles Symposium | 3 |
| 2014 | On Channel Estimation for Multi-User MIMO in LTE-A UplinkabstractIn 3GPP Long Term Evolution-Advanced (LTE-A), uplink multi-input-multi-output(MIMO) transmission have been introduced as a key technology to improve the spectrum efficiency. Both single-user MIMO (SU-MIMO) and multi-user MIMO (MU-MIMO) modes are supported. With UL-MIMO, multiple layers of data are mapped on the same set of recourses and transmitted simultaneously. Accordingly, the demodulation reference signal (DMRS) in LTE-A applies both frequency domain code division multiplexing and time domain code division multiplexing to support channel estimation for each multiplexed layer. Discrete Fourier transform (DFT)-based channel estimation schemes and the applications for LTE UL-MIMO have been widely studied. However, the orthogonality among DMRS of different layers could be reduced due to the channel conditions. The effect is significant when the arrival times of the paired user equipment (UE) at the receiver are not strictly aligned if MU-MIMO is applied. Consequently, the performance of DFT-based channel estimation schemes is affected. In this paper, we demonstrate a novel DFT based channel estimation scheme that can restore the inter-layer orthogonality. Simulation results show that the proposed method provides significant improvement over conventional DFT-based channel estimation methods. Yu-chun Wu, Shulan Feng, Philipp Zhang, Lixia Xue, Yongxing Zhou |
VTC Spring | 5 |