Kangheng Liang

dblp:398/3537 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2026
0009-0005-1367-0606ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Information retrieval · 100%
Artificial intelligence
2 papers
Information extraction and text analysis · 50% Question answering and dialogue systems · 25% Vision and language · 25%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › retrieval models › neural retrieval
diffusion-based retrieval
1.922026
ADaFuSE: Adaptive Diffusion-generated Image and Text Fusion for Interactive Text-to-Image Retrieval · SIGIR 2026
Diffusion Augmented Retrieval: A Training-Free Approach to Interactive Text-to-Image Retrieval · SIGIR 2025
Natural language and speech › Information extraction and text analysis › syntactic parsing
chunking
1.012026
How Should Multimodal Information Be Chunked for Complex Question Answering? · SIGIR 2026
Natural language and speech › Information extraction and text analysis › fact-checking
evidence retrieval
1.012026
How Should Multimodal Information Be Chunked for Complex Question Answering? · SIGIR 2026
Computer vision › Vision and language
multimodal fusion
1.012026
ADaFuSE: Adaptive Diffusion-generated Image and Text Fusion for Interactive Text-to-Image Retrieval · SIGIR 2026
Natural language and speech › Question answering and dialogue systems
multimodal question answering
1.012026
How Should Multimodal Information Be Chunked for Complex Question Answering? · SIGIR 2026
Information retrieval
multimodal retrieval
1.012026
ADaFuSE: Adaptive Diffusion-generated Image and Text Fusion for Interactive Text-to-Image Retrieval · SIGIR 2026
Information retrieval › cross-modal retrieval
text-to-image retrieval
0.912025
Diffusion Augmented Retrieval: A Training-Free Approach to Interactive Text-to-Image Retrieval · SIGIR 2025
Information retrieval › query reformulation
conversational query rewriting
0.312025
Diffusion Augmented Retrieval: A Training-Free Approach to Interactive Text-to-Image Retrieval · SIGIR 2025
Information retrieval › query reformulation
query refinement
0.312025
Diffusion Augmented Retrieval: A Training-Free Approach to Interactive Text-to-Image Retrieval · SIGIR 2025

Methods — techniques the papers use, named apart from their topics

diffusion model · 2.9mixture of experts · 2.0adaptive gating · 2.0large language model dialogue refinement · 0.9
YearPublicationVenuePosition
2026 How Should Multimodal Information Be Chunked for Complex Question Answering?
abstract
Effective question answering over heterogeneous documents requires not only powerful reasoning models but also careful decisions about how evidence is structured and retrieved. In many real-world settings, information is inherently multimodal: a researcher assessing a scientific finding must reason jointly over paragraphs, figures, and result tables; a financial analyst combines textual summaries with tabular data; and a clinician interprets diagnostic imaging alongside patient records. Yet the question of how to decompose such heterogeneous documents into meaningful, retrievable units of evidence remains largely unexplored.
Kangheng Liang
SIGIR1
2026 ADaFuSE: Adaptive Diffusion-generated Image and Text Fusion for Interactive Text-to-Image Retrieval
abstract
Recent advances in interactive text-to-image retrieval (I-TIR) use diffusion models to bridge the modality gap between the textual information need and the images to be searched. However, existing frameworks fuse multi-modal user feedback by simple embedding addition. In this work, we show that this basic fusion strategy indiscriminately incorporates generative noise produced by the diffusion model, leading to performance degradation for up to 55.62% of samples. We further propose ADaFuSE (Adaptive Diffusion-Text Fusion with Semantic-aware Experts), a lightweight fusion model designed to align and calibrate multi-modal views for diffusion-augmented I-TIR, which can be plugged into existing frameworks without modifying the backbone encoder. Specifically, we introduce a dual-branch fusion mechanism that employs an adaptive gating branch to dynamically balance modality reliability, alongside a semantic-aware mixture-of-experts branch to capture fine-grained cross-modal nuances. Via thorough evaluation over four standard I-TIR benchmarks, ADaFuSE achieves state-of-the-art performance, surpassing the DAR model by up to 3.49% in Hits@10 with only a 5.29% parameter increase, while exhibiting stronger robustness to noisy and longer interactive queries. These results show that generative augmentation coupled with principled fusion provides a simple, generalizable alternative to fine-tuning for tasks like product search.
Xingwu Zhang, Kangheng Liang, Guanxuan Li, Richard McCreadie, Zijun Long
SIGIR3
2025 Diffusion Augmented Retrieval: A Training-Free Approach to Interactive Text-to-Image Retrieval
abstract
Interactive Text-to-image retrieval (I-TIR) is an important enabler for a wide range of state-of-the-art services in domains such as e-commerce and education.However, current methods rely on finetuned Multimodal Large Language Models (MLLMs), which are costly to train and update, and exhibit poor generalizability.This latter issue is of particular concern, as: 1) finetuning narrows the pretrained distribution of MLLMs, thereby reducing generalizability; and 2) I-TIR introduces increasing query diversity and complexity.As a result, I-TIR solutions are highly likely to encounter queries and images not well represented in any training dataset.To address this, we propose leveraging Diffusion Models (DMs) for text-to-image mapping, to avoid finetuning MLLMs while preserving robust performance on complex queries.Specifically, we introduce Diffusion Augmented Retrieval (DAR), a framework that generates multiple intermediate representations via LLM-based dialogue refinements and DMs, producing a richer depiction of the user's information needs.This augmented representation facilitates more accurate identification of semantically and visually related images.Extensive experiments on four benchmarks show that for simple queries, DAR achieves results on par with finetuned I-TIR models, yet without incurring their tuning overhead.Moreover, as queries become more complex through additional conversational turns, DAR surpasses finetuned I-TIR models by up to 7.61% in Hits@10 after ten turns, illustrating its improved generalization for more intricate queries.
Zijun Long, Kangheng Liang, Gerardo Aragon-Camarasa, Richard McCreadie, Paul Henderson
SIGIR2