Qianchi Zhang

dblp:332/5857 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2026
0009-0000-8359-8906ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Artificial intelligence
1 paper
Language models and text generation · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
hallucination mitigation
1.012026
Stable-RAG: Mitigating Retrieval-Permutation-Induced Hallucinations in Retrieval-Augmented Generation · ACL (1) 2026
Natural language and speech › Language models and text generation
retrieval-augmented generation
1.012026
Stable-RAG: Mitigating Retrieval-Permutation-Induced Hallucinations in Retrieval-Augmented Generation · ACL (1) 2026
Information retrieval › retrieval-augmented generation
context compression
1.012026
Less is More: Compact Clue Selection for Efficient Retrieval-Augmented Generation Reasoning · WWW 2026
Information retrieval
reranking
1.012026
Less is More: Compact Clue Selection for Efficient Retrieval-Augmented Generation Reasoning · WWW 2026
Information retrieval
retrieval-augmented generation
1.012026
Less is More: Compact Clue Selection for Efficient Retrieval-Augmented Generation Reasoning · WWW 2026

Methods — techniques the papers use, named apart from their topics

min-max optimization · 1.0large language model · 1.0hidden state clustering · 1.0fine-tuning · 1.0contrastive decoding · 1.0
YearPublicationVenuePosition
2026 Stable-RAG: Mitigating Retrieval-Permutation-Induced Hallucinations in Retrieval-Augmented Generation
abstract
Retrieval-Augmented Generation (RAG) has become a key paradigm for reducing factual hallucinations in Large Language Models (LLMs), yet little is known about how the order of retrieved documents affects model behavior.We empirically show that under a Top-5 retrieval setting with the gold document included, LLM answers vary substantially across permutations of the retrieved set, even when the gold document is fixed in the first position.This reveals a previously underexplored sensitivity to retrieval permutations.Although existing robust RAG methods focus primarily on enhancing LLM robustness to low-quality retrieval and mitigating positional bias to distribute attention fairly over long contexts, neither approach directly addresses permutation sensitivity.In this paper, we propose Stable-RAG, which exploits permutation sensitivity estimation to mitigate permutation-induced hallucinations.Stable-RAG runs the generator under multiple retrieval orders, clusters hidden states, and decodes from a cluster-center representation that captures the dominant reasoning pattern.It then uses these reasoning results to align hallucinated outputs toward the correct answer, encouraging the model to produce consistent and accurate predictions across document permutations.Experiments on three QA datasets show that Stable-RAG improves answer accuracy, reasoning consistency, and generalization across datasets, retrievers, and input lengths compared with strong baselines 1 .
Qianchi Zhang, Hainan Zhang 0001, Liang Pang 0001, Hongwei Zheng 0003, Zhiming Zheng 0001
ACL (1)1
2026 Less is More: Compact Clue Selection for Efficient Retrieval-Augmented Generation Reasoning
abstract
Current RAG retrievers are designed primarily for human readers, emphasizing complete, readable, and coherent paragraphs. However, Large Language Models (LLMs) benefit more from precise, compact, and well-structured input, which enhances reasoning quality and efficiency. Existing methods rely on reranking or summarization to identify key sentences, but may introduce semantic breaks and unfaithfulness. Thus, efficiently extracting and organizing answer-relevant clues from large-scale documents while reducing LLM reasoning costs remains challenging in RAG systems. Inspired by Occam's razor, we frame LLM-centric retrieval as MinMax optimization: maximizing the extraction of potential clues and reranking them for well-organization, while minimizing reasoning costs by truncating to the smallest sufficient set of clues. In this paper, we propose CompSelect, a compact clue selection mechanism for LLM-centric RAG, consisting of a clue extractor, a reranker, and a truncator. (1) The clue extractor first uses answer-containing sentences as fine-tuning targets, aiming to extract sufficient potential clues; (2) The reranker is trained to prioritize effective clues based on real LLM feedback; (3) The truncator uses the truncated text containing the minimum sufficient clues for answering the question as fine-tuning targets, thereby enabling efficient RAG reasoning. Experiments on three QA datasets demonstrate that CompSelect improves performance while reducing both total and online latency compared to a range of baseline methods. Further analysis also confirms its robustness to unreliable retrieval and generalization across different scenarios.
Qianchi Zhang, Hainan Zhang 0001, Liang Pang 0001, Yongxin Tong, Hongwei Zheng 0003, Zhiming Zheng 0001
WWW1
2026 HEFKVis: A visual analysis approach for exploring students' online learning behavior
abstract
The analysis and evaluation of students’ online learning behaviors are crucial tasks in online education. Analyzing various behaviors during the learning process and obtaining the corresponding evaluation results can help teachers understand students’ learning conditions, adjust teaching strategies on time, and enhance the quality of teaching. However, the existing methods for evaluating learning behaviors are often based on one specific dimension and have difficulty analyzing student-learning behavior data simultaneously across multiple online platforms, leading to incomplete and inaccurate analysis results. In this paper, we propose a visual analysis pipeline based on a comprehensive scoring model of diverse features of online learning process data that is capable of detecting and analyzing learning behavior anomalies in various learning scenarios. We also developed a visual analysis system that demonstrates the effectiveness of our pipeline in multi-platform and multi-scenario learning behavior analysis. We illustrate the effectiveness and usability of the system through two usage scenarios and in-depth user interviews.
Zhang Qing, Deyu Guo, Qianchi Zhang, Yining Quan, Xiaoyang Han
Vis. Informatics3
2022 Construction and Application of the Knowledge Graph in Endangered Plants
abstract
The current network information of endangered plants is scattered, making acquiring and reusing relevant knowledge difficult. The endangered plants' knowledge graph ePlantKG constructed in this paper can form a visual semantic information network. The data sources in this paper include China Rare and Endangered Plant Information Network, Baidu Encyclopedia, Wikipedia, and Kuaiming Encyclopedia. Firstly, we use a Python crawler to obtain network data and preprocess them. Then we use the obtained structured data combined with previously constructed plant ontology to define and build new ontologies. Next, through the rule-based method, we extract triples from semi-structured and unstructured data, fuse knowledge between heterogeneous data, and store them in the Neo4j database to form ePlantKG. Finally, we design and build a knowledge service platform to illustrate knowledge graphs and intelligent question answering. The intelligent question answering algorithm extracts features from the user's input text with TF-IDF and classifies questions with Naive Bayes. After realizing the similarity matching between entities and relations, we retrieve answers with the returned Cypher statement. The ePlantKG records 1926 species of endangered plants and 37860 species of common plants, and the image filling rate of endangered plants is more than 99 %. The platform implements several functions, e.g., graph display, entity recognition, and intelligent question answering. This paper realizes the information sharing and reuses on endangered plants, providing a method reference for applying knowledge graph in forestry intelligent question answering system and forestry big data analysis.
Haochuan Wei, Qianchi Zhang, Weixuan Gao, Xianghao Meng
ICIS2