VLDB 2026 Research / reviewers in the wild / expert
Xinyao Liu
dblp:251/1052
· DBLP profile ↗
18ranked-venue papers
6as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 2 first-author · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Deployment of Lightweight LLMs on Edge Devices: Kernel-Level Profiling and Cross-Platform Insights
Xinyao Liu, Xingzhou Zhang, Weisong Shi |
ICDCS | 1 |
| 2026 | PT-Herb: A Prompt-Tuned Network with Adaptive Contrastive Learning for Long-Tailed Traditional Chinese Medicine Herb Recognition
Jiehan Zhou, Xinyao Liu, Yuwu Lu |
Pattern Recognit. | 3 |
| 2025 | Mg-Mcc: Metadata-Guided Skin Lesion Classification Via Multimodal Consistency ConstraintsabstractWith the continuous rise in the prevalence of skin diseases, accurate classification has become an urgent demand in clinical diagnosis and treatment. In recent years, deep learningbased methods have demonstrated excellent performance in skin lesion classification tasks. To further improve diagnostic accuracy, multimodal fusion approaches integrating dermoscopic images, clinical photographs, and metadata have been explored. However, many existing methods still fail to effectively capture the deep semantic associations among these modalities. This oversight leads to semantic fragmentation across modalities, loss of critical diagnostic cues, and ultimately hinders classification accuracy. To address this challenge, we propose a Metadata-Guided Multimodal Consistency Constraints (MG-MCC) framework, which leverages metadata as a semantic anchor and enforces collaborative constraints through three specially designed loss functions: Metadata-Aware Alignment (MA), Lesion Feature Purification (LFP), and Semantic Consistency Preservation (SCP). Specifically, the MA loss enforces precise cross-modal alignment between image features and metadata in a unified space; the LFP loss strengthens semantic associations between core lesion features and metadata to accurately separate background features; and the SCP loss avoids the loss of diagnostic ROI information caused by Instance Normalization. Extensive experiments conducted on two public skin lesion datasets demonstrate that the proposed MG-MCC framework significantly outperforms state-of-the-art methods in skin lesion classification tasks. Shudi Zhang, Junchang Xin, Qi Shen 0002, Xinyao Liu, Zhiqiong Wang |
BIBM | 5 |
| 2025 | Constructing Ophthalmic MLLM for Positioning-Diagnosis Collaboration Through Clinical Cognitive Chain ReasoningabstractMultimodal large language models (MLLMs) demonstrate significant potential in the field of medical diagnosis. However, they face critical challenges in specialized domains such as ophthalmology, particularly the fragmentation of annotation granularity and inconsistencies in clinical reasoning logic, which hinder precise cross-modal understanding. This paper introduces FundusExpert, an ophthalmology-specific MLLM with integrated positioning-diagnosis reasoning capabilities, along with FundusGen, a dataset constructed through the intelligent Fundus-Engine system. Fundus-Engine automates localization and leverages MLLM-based semantic expansion to integrate global disease classification, local object detection, and fine-grained feature analysis within a single fundus image. Additionally, by constructing a clinically aligned cognitive chain, it guides the model to generate interpretable reasoning paths. FundusExpert, fine-tuned with instruction data from FundusGen, achieves the best performance in ophthalmic question-answering tasks, surpassing the average accuracy of the 40B MedRegA by 26.6%. It also excels in zero-shot report generation tasks, achieving a clinical consistency of 77.0%, significantly outperforming GPT-4o's 47.6%. Furthermore, we reveal a scaling law between data quality and model capability ($L \propto N^{0.068}$), demonstrating that the cognitive alignment annotations in FundusGen enhance data utilization efficiency. By integrating region-level localization with diagnostic reasoning chains, our work develops a scalable, clinically-aligned MLLM and explores a pathway toward bridging the visual-language gap in specific MLLMs. Our project can be found at https://github.com/MeteorElf/FundusExpert. Xinyao Liu, Diping Song |
ICCV | 1 |
| 2025 | ElaSleepNet: Exploring an Elastic Multimodal Neural Network for Sleep Staging via Temporal and Contextual Consistency LearningabstractIntegrating multimodal learning (ML) with polysomnography (PSG) has emerged as a research hotspot for reliable sleep staging. However, the complexity of these signals and the discomfort associated with wearing multi-lead devices somewhat limit the feasibility of daily and ubiquitous sleep monitoring. Unfortunately, most existing ML paradigms are constrained by consistent and fixed input patterns. When the number of modalities is less than required by the ML framework, it is easy to cause inference bias, resulting in a significant performance degradation. To this end, we propose an elastic multimodal sleep staging network (ElaSleepNet), consisting of multimodal information completion (MIC) and adaptive cross-modal (ACM) interaction. Specifically, MIC maximizes the consistency of multimodal signals on intra-epoch temporal level and inter-epoch contextual level, thereby enhancing the reasoning and completion abilities of the available modalities for the unavailable modalities. Moreover, we introduce learnable parameters and design the ACM attention mechanism, which allows handling multimodal information interaction while maintaining robustness in the absence of certain modalities. Our ElaSleepNet demonstrates its state-of-the-art on three multimodal sleep datasets. Compared with previous methods, ElaSleepNet can achieve better performance with fewer testing modalities, making it flexible for daily monitoring. Qi Shen 0002, Junchang Xin, Bing Tian Dai, Shudi Zhang, Xinyao Liu, Zhiqiong Wang |
ACM Multimedia | 5 |
| 2025 | GPLQ: A General, Practical, and Lightning QAT Method for Vision TransformersabstractVision Transformers (ViTs) are essential in computer vision but are computationally intensive, too. Model quantization, particularly to low bit-widths like 4-bit, aims to alleviate this difficulty, yet existing Post-Training Quantization (PTQ) and Quantization-Aware Training (QAT) methods exhibit significant limitations. PTQ often incurs substantial accuracy drop, while QAT achieves high accuracy but suffers from prohibitive computational costs, limited generalization to downstream tasks, training instability, and lacking of open-source codebase. To address these challenges, this paper introduces General, Practical, and Lightning Quantization (GPLQ), a novel framework designed for efficient and effective ViT quantization. GPLQ is founded on two key empirical insights: the paramount importance of activation quantization and the necessity of preserving the model's original optimization basin to maintain generalization. Consequently, GPLQ employs a sequential activation-first, weights-later strategy. Stage 1 keeps weights in FP32 while quantizing activations with a feature mimicking loss in only 1 epoch to keep it stay in the same basin, thereby preserving generalization. Stage 2 quantizes weights using a PTQ method. As a result, GPLQ is 100x faster than existing QAT methods, lowers memory footprint to levels even below FP32 training, and achieves 4-bit model performance that is highly competitive with FP32 models in terms of both accuracy on ImageNet and generalization to diverse downstream tasks, including fine-grained visual classification and object detection. We release an easy-to-use open-source toolkit supporting multiple vision tasks at [GPLQ code](https://github.com/wujx2001/GPLQ). Guang Liang, Xinyao Liu, Jianxin Wu 0001 |
NeurIPS | 2 |
| 2025 | A pre-trained data deduplication model based on active learning
Xinyao Liu, Fengmao Lv, Hongtao Xue, Jie Hu 0007, Shengdong Du, Tianrui Li 0001 |
Expert Syst. Appl. | 2 |
| 2025 | Soft label-guided transformer for radiology report generation
Xinyao Liu, Junchang Xin, Qi Shen 0002, Zhiqiong Wang |
J. Biomed. Informatics | 1 |
| 2025 | Cross-Domain Few-Shot Learning Method Based on Fractional Domain Information for Hyperspectral Image Multi-Class Change DetectionabstractHyperspectral image multi-class change detection (HSI-MCD) based on deep learning (DL) rely significantly on the number of labeled data. Due to the high cost of manually labeling for hyperspectral images (HSIs), obtaining a large amount of labeled samples is difficult. Moreover, for multi-class change detection (MCD) tasks, there is the phenomenon of semantic cross-coupling of changes due to complex change scenarios. To solve the above problems, a cross-domain few-shot learning method based on fractional domain information for HSI-MCD (FrCFSL) is proposed. Firstly, a spectral-spatial-fractional information extraction module is proposed, which can extract spectral-spatial-fractional domain joint feature. Thus, the module can obtain more comprehensive and discriminative representations of land cover categories, alleviating the phenomenon of semantic cross-coupling between classes. Afterward, a cross-domain fewshot learning strategy is introduced, where it learns task-relevant category discrimination meta-knowledge from a pair of richly labeled very high-resolution optical images (VHRIs) dataset and transfers it to the bitemporal HSIs dataset. Thus, the model can achieve better MCD performance with a small number of labeled samples. Finally, to mitigate the domain distribution differences between VHRIs data and HSIs data, a topological structure alignment module is proposed to align the intrinsic topological relationships between land cover categories, thus narrowing the gap between the two domain distributions. Through experiments conducted on three HSI-MCD datasets and comparative analysis with six state-of-the-art methods, the validity and stability of the proposed method are indicated. Shou Feng, Jinghe Zhang, Yuanze Fan, Xinyao Liu, Chunhui Zhao 0003, Wei Li 0032, Ran Tao 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | MaskEditor: Instruct 3D Object Editing with Learned Masks
Xinyao Liu, Kai Xu 0004, Yuhang Huang 0006, Renjiao Yi, Chenyang Zhu 0002 |
PRCV (6) | 1 |
| 2024 | GSEditPro: 3D Gaussian Splatting Editing with Attention-based Progressive LocalizationabstractAbstract With the emergence of large‐scale Text‐to‐Image(T2I) models and implicit 3D representations like Neural Radiance Fields (NeRF), many text‐driven generative editing methods based on NeRF have appeared. However, the implicit encoding of geometric and textural information poses challenges in accurately locating and controlling objects during editing. Recently, significant advancements have been made in the editing methods of 3D Gaussian Splatting, a real‐time rendering technology that relies on explicit representation. However, these methods still suffer from issues including inaccurate localization and limited manipulation over editing. To tackle these challenges, we propose GSEditPro, a novel 3D scene editing framework which allows users to perform various creative and precise editing using text prompts only. Leveraging the explicit nature of the 3D Gaussian distribution, we introduce an attention‐based progressive localization module to add semantic labels to each Gaussian during rendering. This enables precise localization on editing areas by classifying Gaussians based on their relevance to the editing prompts derived from cross‐attention layers of the T2I model. Furthermore, we present an innovative editing optimization method based on 3D Gaussian Splatting, obtaining stable and refined editing results through the guidance of Score Distillation Sampling and pseudo ground truth. We prove the efficacy of our method through extensive experiments. Yanhao Sun, Runze Tian, Xinyao Liu, Yan Zhang 0057, Kai Xu 0004 |
Comput. Graph. Forum | 4 |
| 2024 | Cyclic edge and cyclic vertex connectivity of (4,5,6)-fullerene graphsabstractCyclic vertex connectivity cκ and cyclic edge connectivity cλ are two important kinds of conditional connectivity, which reflect the number of vertices or edges that can be removed before the graph is disconnected and at least two components contain a cycle, respectively. They have important applications in various networks such as computer networks or biochemical networks. In addition, a fullerene is a special kind of molecule in chemistry. A classic fullerene graph is a 3-connected cubic planar graph with only pentagonal and hexagonal faces. (4, 5, 6)-fullerene graphs are atypical fullerene graphs which also contain 4-faces. In this paper, we prove that cκ=cλ for (4,5,6)-fullerene graphs except for four exceptional graphs with order less than 16. We also give O(ν)-algorithms to determine the cyclic vertex connectivity and the cyclic edge connectivity of (4,5,6)-fullerene graphs. Jun Liang 0002, Xinyao Liu, Dingjun Lou, Zan-Bo Zhang, Zixin Qin |
Discret. Appl. Math. | 2 |
| 2024 | Scale-aware local difference attention on pyramidal features for crowd counting
Qian Zhang 0046, Shizhou Zhang, Xinyao Liu, Yanning Zhang 0001 |
Multim. Tools Appl. | 3 |
| 2023 | A missing value filling model based on feature fusion enhanced autoencoder
Xinyao Liu, Shengdong Du, Tianrui Li 0001, Fei Teng 0001, Yan Yang 0001 |
Appl. Intell. | 1 |
| 2023 | Deep learning in medical image super resolution: a review
Hujun Yang, Xinyao Liu, Chuangang Li, Junchang Xin, Zhiqiong Wang |
Appl. Intell. | 3 |
| 2022 | Eureka: Neural Insight Learning for Knowledge Graph ReasoningabstractThe human recognition system has presented the remarkable ability to effortlessly learn novel knowledge from only a few trigger events based on prior knowledge, which is called insight learning. Mimicking such behavior on Knowledge Graph Reasoning (KGR) is an interesting and challenging research problem with many practical applications. Simultaneously, existing works, such as knowledge embedding and few-shot learning models, have been limited to conducting KGR in either “seen-to-seen” or “unseen-to-unseen” scenarios. To this end, we propose a neural insight learning framework named Eureka to bridge the “seen” to “unseen” gap. Eureka is empowered to learn the seen relations with sufficient training triples while providing the flexibility of learning unseen relations given only one trigger without sacrificing its performance on seen relations. Eureka meets our expectation of the model to acquire seen and unseen relations at no extra cost, and eliminate the need to retrain when encountering emerging unseen relations. Experimental results on two real-world datasets demonstrate that the proposed framework also outperforms various state-of-the-art baselines on datasets of both seen and unseen relations. Xuan Zhang 0009, Xun Liang 0001, Bo Wu 0026, Xiangping Zheng 0002, Sensen Zhang, Yuhui Guo, Xinyao Liu |
COLING | 8 |
| 2020 | HFuzz: Towards automatic fuzzing testing of NB-IoT core network protocols implementations
Xinyao Liu, Baojiang Cui, Junsong Fu 0001, Jinxin Ma |
Future Gener. Comput. Syst. | 1 |
| 2019 | A Distributed Position-Based Routing Algorithm in 3-D Wireless Industrial Internet of ThingsabstractSmart factory is a typical application scene of Internet of Things and wireless terminal devices naturally compose a three-dimensional (3-D) industrial wireless network. A primary requirement in the network is delivering packets from source node to destination node. Most geographic routing algorithms are designed for planar networks and they do not suit 3-D networks. In this paper, we extend a greedy perimeter stateless routing algorithm (GPSR) into three dimensions named GPSR-3D. In GPSR-3D, each node decides next hop of a packet by cooperating with only local neighbors and hence this algorithm is totally distributed. GPSR-3D comprises two packet forwarding patterns named greedy forwarding pattern (GFP) and surface forwarding pattern (SFP). In GFP, a node always sends the packet to a neighbor closest to destination and when it fails, SFP is employed for recovery. In SFP, we first divide the whole network space into a set of subspaces based on a novel 3-D geometric structure. Then, a parallel polyhedron traverse algorithm is proposed to recover local minima. A flowchart of GPSR-3D is given to clearly present the process of delivering a packet based on GFP and SFP. Simulation results show that GPSR-3D is of great reliability, energy efficiency and storage efficiency. Specifically, data transmission amount in GPSR-3D is about 67% and 71% to that of multihop Delaunay triangulation (MDT) and GDSTR-3D on average. Moreover, GPSR-3D performs much better than MDT and GDSTR-3D in terms of average storage cost and the average storage space in GPSR-3D is about 48% and 26% to that of MDT and GDSTR-3D, respectively. Junsong Fu 0001, Baojiang Cui, Na Wang 0003, Xinyao Liu |
IEEE Trans. Ind. Informatics | 4 |