VLDB 2026 Research / reviewers in the wild / expert
Xinping Zhao
dblp:05/2524
· DBLP profile ↗
7ranked-venue papers
4as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Language models and text generation · 45% Vision and language · 35% Information extraction and text analysis · 15% | |
| Computer graphics and multimedia
1 paper |
Image and video coding · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 10 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › vision-language model
prompt learning |
0.9 | 1 | 2025 | Adaptive Prompt Learning for Blind Image Quality Assessment with Multi-modal Mixed-datasets Training · ACM Multimedia 2025 |
Computer vision › Vision and language
vision-language model |
0.9 | 1 | 2025 | Adaptive Prompt Learning for Blind Image Quality Assessment with Multi-modal Mixed-datasets Training · ACM Multimedia 2025 |
Image and video coding
image quality assessment |
0.9 | 1 | 2025 | Adaptive Prompt Learning for Blind Image Quality Assessment with Multi-modal Mixed-datasets Training · ACM Multimedia 2025 |
Image and video coding › image quality assessment
no-reference image quality assessment |
0.9 | 1 | 2025 | Adaptive Prompt Learning for Blind Image Quality Assessment with Multi-modal Mixed-datasets Training · ACM Multimedia 2025 |
Natural language and speech › Language models and text generation
alignment |
0.8 | 1 | 2024 | Take Off the Training Wheels! Progressive In-Context Learning for Effective Alignment · EMNLP 2024 |
Natural language and speech › Information extraction and text analysis
evidence extraction |
0.8 | 1 | 2024 | SEER: Self-Aligned Evidence Extraction for Retrieval-Augmented Generation · EMNLP 2024 |
Natural language and speech › Language models and text generation
in-context learning |
0.8 | 1 | 2024 | Take Off the Training Wheels! Progressive In-Context Learning for Effective Alignment · EMNLP 2024 |
Natural language and speech › Language models and text generation
retrieval-augmented generation |
0.8 | 1 | 2024 | SEER: Self-Aligned Evidence Extraction for Retrieval-Augmented Generation · EMNLP 2024 |
Information retrieval
retrieval-augmented generation |
0.8 | 1 | 2024 | SEER: Self-Aligned Evidence Extraction for Retrieval-Augmented Generation · EMNLP 2024 |
Machine learning › Transfer learning and domain adaptation › multi-source learning
multi-dataset training |
0.3 | 1 | 2025 | Adaptive Prompt Learning for Blind Image Quality Assessment with Multi-modal Mixed-datasets Training · ACM Multimedia 2025 |
Methods — techniques the papers use, named apart from their topics
learnable prompt vectors · 1.7dual weight adjustment · 1.7conditional fusion · 1.7self-aligned learning · 1.5in-context learning · 0.8demonstration · 0.8ICL vector extraction · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Semi-supervised multi-label feature selection with consistent sparse graph learning
Yan Zhong 0001, Xinping Zhao, Li Zhang 0104, Xinyuan Song 0002, Lei Shi 0030, Bingbing Jiang 0001 |
Neural Networks | 3 |
| 2025 | The Research on Intelligent Medical Triage System Based on Large Language ModelsabstractTo address critical challenges in the field of medical AI, including data fragmentation, large language model (LLM) hallucinations, and the lack of humanistic interaction, this paper proposes an intelligent medical triage system solution that integrates LLM technology with clinical needs. The solution employs Retrieval-Augmented Generation (RAG) technology to fuse medical knowledge bases, enabling real-time updates of medical knowledge and suppression of LLM hallucinations. Meanwhile, a multimodal input parsing framework is constructed to effectively integrate text, speech, and image information, significantly improving the accuracy of information collection in primary care scenarios. Finally, an intelligent decision engine is established, combining functions such as medical insurance policy factors and emotion recognition to achieve precise department matching and humanistic care responses. The core innovation of this study lies in the proposal of a “technology-scene-ethics” three-dimensional collaborative architecture, breaking through the limitations of single-modal systems and enabling dynamic knowledge governance. The system adopts a microservice architecture and the Deepseek-V3 base model to ensure technological autonomy and data security. Experimental verification shows that the system can efficiently handle emergency triage processes, generating comprehensive decisions including patient symptom summaries, predictive directions, department recommendations, expert allocation, estimated costs, risk grading, and humanistic suggestions, providing a replicable paradigm for the practical application of medical LLMs. Xinping Zhao |
BIBM | 1 |
| 2025 | Adaptive Prompt Learning for Blind Image Quality Assessment with Multi-modal Mixed-datasets TrainingabstractDue to the high cost and small scale of Image Quality Assessment (IQA) datasets, achieving robust generalization remains challenging for prevalent Blind IQA (BIQA) methods. Traditional deep learning-based methods emphasize visual information to capture quality features, while recent developments in Vision-Language Models (VLMs) demonstrate strong potential in learning generalizable representations through textual information. However, applying VLMs to BIQA poses three major Challenges: (1) How to make full use of the multi-modal information. (2) The prompt engineering for appropriate quality description is extremely time-consuming. (3) How to use mixed data for joint training to enhance the generalization of VLM-based BIQA model. To this end, we propose a Multi-modal BIQA method with prompt learning, named MMP-IQA. For (1), we propose a conditional fusion module to better utilize the cross-modality information. By jointly adjusting visual and textual features, our model can capture quality information with a stronger representation ability. For (2), we model the quality prompt's context words with learnable vectors during the training process, which can be adaptively updated for superior performances. For (3), we jointly train a linearity-induced quality evaluator, a relative quality evaluator, and a dataset-specific absolute quality evaluator. In addition, we propose a dual automatic weight adjustment strategy to adaptively balance the loss weights between different datasets and among various losses within the same dataset. Extensive experiments illustrate the superior effectiveness of MMP-IQA. Yan Zhong 0001, Xinping Zhao, Li Zhang 0104, Xinyuan Song 0002, Tingting Jiang 0001 |
ACM Multimedia | 2 |
| 2024 | Take Off the Training Wheels! Progressive In-Context Learning for Effective AlignmentabstractRecent studies have explored the working mechanisms of In-Context Learning (ICL).However, they mainly focus on classification and simple generation tasks, limiting their broader application to more complex generation tasks in practice.To address this gap, we investigate the impact of demonstrations on token representations within the practical alignment tasks.We find that the transformer embeds the task function learned from demonstrations into the separator token representation, which plays an important role in the generation of prior response tokens.Once the prior response tokens are determined, the demonstrations become redundant.Motivated by this finding, we propose an efficient Progressive In-Context Alignment (PICA) method consisting of two stages.In the first few-shot stage, the model generates several prior response tokens via standard ICL while concurrently extracting the ICL vector that stores the task function from the separator token representation.In the following zero-shot stage, this ICL vector guides the model to generate responses without further demonstrations.Extensive experiments demonstrate that our PICA not only surpasses vanilla ICL but also achieves comparable performance to other alignment tuning methods.The proposed training-free method reduces the time cost (e.g., 5.45×) with improved alignment performance (e.g., 6.57+).Consequently, our work highlights the application of ICL for alignment and calls for a deeper understanding of ICL for complex generations. Dongfang Li 0002, Xinshuo Hu, Xinping Zhao, Yibin Chen, Baotian Hu, Min Zhang 0005 |
EMNLP | 4 |
| 2024 | SEER: Self-Aligned Evidence Extraction for Retrieval-Augmented GenerationabstractRecent studies in Retrieval-Augmented Generation (RAG) have investigated extracting evidence from retrieved passages to reduce computational costs and enhance the final RAG performance, yet it remains challenging.Existing methods heavily rely on heuristic-based augmentation, encountering several issues: (1) Poor generalization due to hand-crafted context filtering; (2) Semantics deficiency due to rulebased context chunking; (3) Skewed length due to sentence-wise filter learning.To address these issues, we propose a model-based evidence extraction learning framework, SEER, optimizing a vanilla model as an evidence extractor with desired properties through selfaligned learning.Extensive experiments show that our method largely improves the final RAG performance, enhances the faithfulness, helpfulness, and conciseness of the extracted evidence, and reduces the evidence length by 9.25 times.The code will be available at https://github.com/HITsz-TMG/SEER. Xinping Zhao, Dongfang Li 0002, Boren Hu, Yibin Chen, Baotian Hu, Min Zhang 0005 |
EMNLP | 1 |
| 2024 | Enhancing Attributed Graph Networks with Alignment and Uniformity Constraints for Session-based RecommendationabstractSession-based Recommendation (SBR), seeking to predict a user’s next action based on an anonymous session, has drawn increasing attention for its practicability. Most SBR models only rely on the contextual transitions within a short session to learn item representations while neglecting additional valuable knowledge. As such, their model capacity is largely limited by the data sparsity issue caused by short sessions. A few studies have exploited the Modeling of Item Attributes (MIA) to enrich item representations. However, they usually involve specific model designs that can hardly transfer to existing attribute-agnostic SBR models and thus lack universality. In this paper, we propose a model-agnostic framework, named AttrGAU (Attributed Graph Networks with Alignment and Uniformity Constraints), to bring the MIA’s superiority into existing attribute-agnostic models, to improve their accuracy and robustness for recommendation. Specifically, we first build a bipartite attributed graph and design an attribute-aware graph convolution to exploit the rich attribute semantics hidden in the heterogeneous item-attribute relationship. We then decouple existing attribute-agnostic SBR models into the graph neural network and attention readout sub-modules to satisfy the non-intrusive requirement. Lastly, we design two representation constraints, i.e., alignment and uniformity, to optimize distribution discrepancy in representation between the attribute semantics and collaborative semantics. Extensive experiments on three public benchmark datasets demonstrate that the proposed AttrGAU framework can significantly enhance backbone models’ recommendation performance and robustness against data sparsity and data noise issues. Our implementation codes will be available at https://github.com/ItsukiFujii/AttrGAU. Xinping Zhao, Chaochao Chen 0001, Jiajie Su, Yizhao Zhang, Baotian Hu |
ICWS | 1 |
| 2022 | SGLCMR: Self-supervised Graph Learning of Generalized Representations for Cross-Market RecommendationabstractCross-Market Recommendation (CMR) has been proposed to improve item recommendation performance in the data-insufficient markets by leveraging knowledge learned from other data-sufficient markets. CMR can be regarded as a subtask of Cross-Domain Recommendation (CDR), which overlaps in items and differs in users. Usually, there are a lot of overlapped items in different markets that make the knowledge transfer across markets feasible and valuable. Most existing graph-based methods construct a cross-domain graph by collecting interactions from different domains and then devise some kind of graph neural network (GNN) to learn graph representations for each node. However, we think these methods suffer from two limitations: (1) graph representations are vulnerable to noisy interactions, especially interactions collected from different domains that make the problem more serious; (2) graph representations are biased because they have been more influenced by data-sufficient domains, making graph representations less generalized. In this paper, we propose a novel framework, Self-supervised Graph Learning of Generalized Representations for Cross-Market Recommendation (SGLCMR), to address the foregoing limitations, and we treat different markets as different domains. First, we construct the single-market and cross-market graph to model market-specific and market-generalized features respectively. Second, we adopt Light Graph Convolution (LGC) to aggregate neighbors without loss of generality. Then, we devise two graph-structured data augmentation operators for self-supervised graph learning, i.e., Market-unaware Dropout and Market-aware Dropout, aiming to solve the limitations that graph representations are vulnerable and biased. Finally, we employ a multi-task learning strategy to optimize our model. Extensive experiments on seven pairs of real-world datasets show that our proposed SGLCMR is highly superior to the state-of-the-art CMR methods. Xinping Zhao, Yingchun Yang |
IJCNN | 1 |