EDBT 2026 Demo / reviewers in the wild / expert
Qingyi Si
dblp:227/6822
· DBLP profile ↗
21ranked-venue papers
4as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 4 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 9 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sparse Attention Across Multiple-Context KV CacheabstractLarge language models face significant cost challenges in long-sequence inference. To address this, reusing historical Key-Value (KV) Cache for improved inference efficiency has become a mainstream approach. Recent advances further enhance throughput by sparse attention mechanisms to select the most relevant KV Cache, thereby reducing sequence length. However, such techniques are limited to single-context scenarios, where historical KV Cache is computed sequentially with causal-attention dependencies. In retrieval-augmented generation (RAG) scenarios, where retrieved documents as context are unknown beforehand, each document’s KV Cache is computed and stored independently (termed multiple-context KV Cache), lacking cross-attention between contexts. This renders existing methods ineffective. Although prior work partially recomputes multiple-context KV Cache to mitigate accuracy loss from missing cross-attention, it requires retaining all KV Cache throughout, failing to reduce memory overhead. This paper presents SamKV, the first exploration of attention sparsification for multiple-context KV Cache. Specifically, SamKV takes into account the complementary information of other contexts when sparsifying one context, and then locally recomputes the sparsified information. Experiments demonstrate that our method compresses sequence length to 15% without accuracy degradation compared with full-recomputation baselines, significantly boosting throughput in multi-context RAG scenarios. Ziyi Cao, Qingyi Si, Bingquan Liu |
AAAI | 2 |
| 2026 | Test-time Prompt InterventionabstractTest-time compute has led to remarkable success in the large language model (LLM) community, particularly for complex tasks, where longer chains of thought (CoTs) are generated to enhance reasoning capabilities. However, growing evidence reveals that such reasoning models often produce CoTs plagued by excessive redundancy, including repetitive verification steps and unnecessary reasoning shifts. The root cause lies in post-training of them that overly rely on outcome reward paradigms, as the data of process reward paradigms, which regulate intermediate reasoning steps, is difficult to construct at scale. To address this, we propose PI, a novel framework for Test-time Prompt Intervention. PI provides an interface to dynamically guide and regulate reasoning paths during inference through timely (When module) and proper (How module) interventions and post-intervention sampling (Which module). This allows human problem-solving expertise and cognitive science principles to be seamlessly integrated into LLMs’ reasoning processes, enhancing controllability and interpretability. Extensive experiments across multiple models and datasets demonstrate that PI significantly shortens CoTs while reducing hallucination, yielding more concise and reliable reasoning. Chenxu Yang, Qingyi Si, Mz Dai, Dingyu Yao, Mingyu Zheng, Zheng Lin 0001, Weiping Wang 0005 |
AAAI | 2 |
| 2026 | Breaking the Trade-Off Between Faithfulness and Expressiveness for Large Language ModelsabstractGrounding responses in external knowledge represents an effective strategy for mitigating hallucinations in Large Language Models (LLMs). However, current LLMs struggle to seamlessly integrate knowledge while simultaneously maintaining faithfulness (or fidelity) and expressiveness, capabilities that humans naturally possess. This limitation results in outputs that either lack support from external knowledge, thereby compromising faithfulness, or appear overly verbose and unnatural, thus sacrificing expressiveness. In this work, to break the trade-off between faithfulness and expressiveness, we propose Collaborative Decoding (CoDe), a novel approach that dynamically integrates output probabilities generated with and without external knowledge. This integration is guided by distribution divergence and model confidence, enabling the selective activation of relevant and reliable expressions from the model's internal parameters. Furthermore, we introduce a knowledge-aware reranking mechanism that prevents over-reliance on prior parametric knowledge while ensuring proper utilization of provided external information. Through comprehensive experiments, our plug-and-play CoDe framework demonstrates superior performance in enhancing faithfulness without compromising expressiveness across diverse LLMs and evaluation metrics, validating both its effectiveness and generalizability. Chenxu Yang, Qingyi Si, Lanrui Wang, Zheng Lin 0001 |
AAAI | 2 |
| 2026 | Outcome-Grounded Advantage Reshaping for Fine-Grained Credit Assignment in Mathematical ReasoningabstractZiheng Li, Liu Kang, Feng Xiao, Luxi Xing, Qingyi Si, Zhuoran Li, Weikang Gong, Deqing Yang, Yanghua Xiao, Hongcheng Guo. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Liu Kang, Luxi Xing, Qingyi Si, Weikang Gong, Deqing Yang, Yanghua Xiao, Hongcheng Guo |
ACL (1) | 5 |
| 2025 | Multimodal Hypothetical Summary for Retrieval-based Multi-image Question AnsweringabstractRetrieval-based multi-image question answering (QA) task involves retrieving multiple question-related images and synthesizing these images to generate an answer. Conventional "retrieve-then-answer" pipelines often suffer from cascading errors because the training objective of QA fails to optimize the retrieval stage. To address this issue, we propose a novel method to effectively introduce and reference retrieved information into the QA. Given the image set to be retrieved, we employ a multimodal large language model (visual perspective) and a large language model (textual perspective) to obtain multimodal hypothetical summary in question-form and description-form. By combining visual and textual perspectives, MHyS captures image content more specifically and replaces real images in retrieval, which eliminates the modality gap by transforming into text-to-text retrieval and helps improve retrieval. To more advantageously introduce retrieval with QA, we employ contrastive learning to align queries (questions) with MHyS. Moreover, we propose a coarse-to-fine strategy for calculating both sentence-level and word-level similarity scores, to further enhance retrieval and filter out irrelevant details. Our approach achieves a 3.7% absolute improvement over state-of-the-art methods on RETVQA and a 14.5% improvement over CLIP. Comprehensive experiments and detailed ablation studies demonstrate the superiority of our method. Peize Li, Qingyi Si, Peng Fu 0008, Zheng Lin 0001, Yan Wang 0028 |
AAAI | 2 |
| 2025 | LEAP: An LLM-Based Evidence Augmented Pipeline for Table-Based Fact Verification
Hanwen Zhang 0010, Qingyi Si, Peng Fu 0008, Zheng Lin 0001, Zhigang Lu 0001, Weiping Wang 0005 |
ADMA (1) | 2 |
| 2025 | Towards Realistic Generation: A Multi-Task Agent for Imitating Diverse Character Linguistic StylesabstractThe advent of large language models (LLMs) has significantly propelled the advancement of Role-Playing Agents (RPAs). However, current Role-Playing Agents predominantly focus on mimicking a character’s fundamental attributes while neglecting the replication of linguistic style, and they are incapable of effectively replicating characters when performing tasks beyond multi-turn dialogues, which results in generated responses that lack authenticity. The reason current RPAs lack this capability is due to the nature of existing character datasets, which lack collections of character quotations and are limited to multi-turn dialogue tasks, constraining the RPA’s performance across other task domains and failing to mimic a character’s linguistic style. To address this gap, we developed a multi-task role-playing dataset named MRstyle, which encompasses a substantial number of real individuals along with their quotations and covers seven different tasks. On this basis, we develop StyleRPA, a Multi-Task Role-Playing Agent (MRPA) that significantly outperforms recent open-source LLMs and RPAs baselines on 7 tasks including Dialogue, Dictionary, Composition, Story Generation, Product Description, Music Commentary, and Open Question Answering. Qingyi Si, Chenxu Yang, Zheng Lin 0001, Yunzhi Liang, Siyang Tao, Weiping Wang 0005 |
IJCNN | 2 |
| 2025 | S-GRPO: Early Exit via Reinforcement Learning in Reasoning ModelsabstractAs Test-Time Scaling emerges as an active research focus in the large language model community, advanced post-training methods increasingly emphasize extending chain-of-thought (CoT) generation length, thereby enhancing reasoning capabilities to approach Deepseek R1-like reasoning models.
However, recent studies reveal that reasoning models (even Qwen3) consistently exhibit excessive thought redundancy in CoT generation. This overthinking issue arises from the inherent limitations of conventional outcome-reward reinforcement learning, which systematically overlooks the regulation of intermediate reasoning processes. This paper introduces Serial-Group Decaying-Reward Policy Optimization (S-GRPO), a novel reinforcement learning paradigm that enables models to implicitly evaluate the sufficiency of intermediate reasoning steps, thereby facilitating early exit in CoT generation.
Unlike GRPO, which samples multiple possible reasoning paths in parallel (parallel group), S-GRPO only samples one reasoning path and serially selects multiple temporal positions from the path to exit thinking and directly generate answers (serial group). For correct answers within a serial group, rewards gradually decrease based on the exit positions along the reasoning path from front to back. This design encourages the model to produce more accurate and concise thoughts, while also incentivizing early thinking termination when appropriate. Empirical evaluations demonstrate that S-GRPO is compatible with state-of-the-art reasoning models, including Qwen3 and Deepseek-distill. Across diverse benchmarks such as GSM8K, AIME 2024, AMC 2023, MATH-500, and GPQA Diamond, S-GRPO achieves a substantial reduction in sequence length (40.4%~61.1%) while simultaneously improving accuracy (absolute 0.72%~3.92%). Muzhi Dai, Chenxu Yang, Qingyi Si |
NeurIPS | 3 |
| 2024 | Object Attribute Matters in Visual Question AnsweringabstractVisual question answering is a multimodal task that requires the joint comprehension of visual and textual information. However, integrating visual and textual semantics solely through attention layers is insufficient to comprehensively understand and align information from both modalities. Intuitively, object attributes can naturally serve as a bridge to unify them, which has been overlooked in previous research. In this paper, we propose a novel VQA approach from the perspective of utilizing object attribute, aiming to achieve better object-level visual-language alignment and multimodal scene understanding. Specifically, we design an attribute fusion module and a contrastive knowledge distillation module. The attribute fusion module constructs a multimodal graph neural network to fuse attributes and visual features through message passing. The enhanced object-level visual features contribute to solving fine-grained problem like counting-question. The better object-level visual-language alignment aids in understanding multimodal scenes, thereby improving the model's robustness. Furthermore, to augment scene understanding and the out-of-distribution performance, the contrastive knowledge distillation module introduces a series of implicit knowledge. We distill knowledge into attributes through contrastive loss, which further strengthens the representation learning of attribute features and facilitates visual-linguistic alignment. Intensive experiments on six datasets, COCO-QA, VQAv2, VQA-CPv2, VQA-CPv1, VQAvs and TDIUC, show the superiority of the proposed method. Peize Li, Qingyi Si, Peng Fu 0008, Zheng Lin 0001, Yan Wang 0028 |
AAAI | 2 |
| 2024 | Multimodal Table UnderstandingabstractMingyu Zheng, Xinwei Feng, Qingyi Si, Qiaoqiao She, Zheng Lin, Wenbin Jiang, Weiping Wang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Mingyu Zheng, Xinwei Feng, Qingyi Si, Qiaoqiao She, Zheng Lin 0001, Wenbin Jiang 0002, Weiping Wang 0005 |
ACL (1) | 3 |
| 2024 | Are Large Language Models Table-based Fact-Checkers?abstractTable-based Fact Verification (TFV) aims to extract the entailment relationship between statements and structured tables. Existing TFV methods based on small-scale models suffer from insufficient labeled data and weak zero-shot ability. Recently, the appearance of Large Language Models (LLMs) has gained lots of attraction in research fields. They have shown strong zero-shot and in-context learning capabilities on several NLP tasks, but their potential on TFV is still unknown. In this work, we implement a preliminary study on whether LLMs are table-based fact-checkers. In detail, we design various prompts to explore how in-context learning can help LLMs in TFV, i.e., zero-shot and few-shot TFV capability. Besides, we carefully design and construct TFV instructions to investigate the performance gain brought by the instruction tuning of LLMs. Experimental results demonstrate that LLMs can achieve acceptable results on zero-shot and few-shot TFV with prompt engineering, while instruction-tuning can stimulate the TFV capability significantly. We also make some valuable findings about the format of zero-shot prompts and the number of in-context examples. Finally, we analyze some possible directions to promote the accuracy of TFV via LLMs, which will benefit further research on table reasoning. Hanwen Zhang 0010, Qingyi Si, Peng Fu 0008, Zheng Lin 0001, Weiping Wang 0005 |
CSCWD | 2 |
| 2024 | Towards Unified Interactive Visual Grounding in The WildabstractInteractive visual grounding in Human-Robot Interaction (HRI) is challenging yet practical due to the inevitable ambiguity in natural languages. It requires robots to disambiguate the user’s input by active information gathering. Previous approaches often rely on predefined templates to ask disambiguation questions, resulting in performance reduction in realistic interactive scenarios. In this paper, we propose TiO, an end-to-end system for interactive visual grounding in human-robot interaction. Benefiting from a unified formulation of visual dialog and grounding, our method can be trained on a joint of extensive public data, and show superior generality to diversified and challenging open-world scenarios. In the experiments, we validate TiO on GuessWhat?! and InViG benchmarks, setting new state-of-the-art performance by a clear margin. Moreover, we conduct HRI experiments on the carefully selected 150 challenging scenes as well as real-robot platforms. Results show that our method demonstrates superior generality to diversified visual and language inputs with a high success rate. Codes and demos are available on https://jxu124.github.io/TiO/. Hanbo Zhang, Qingyi Si, Xuguang Lan, Tao Kong |
ICRA | 3 |
| 2024 | Towards Flexible Evaluation for Generative Visual Question AnsweringabstractThroughout rapid development of multimodal large language models, a crucial ingredient is a fair and accurate evaluation of their multimodal comprehension abilities. Although Visual Question Answering (VQA) could serve as a developed test field, limitations of VQA evaluation, like the inflexible pattern of Exact Match, have hindered MLLMs from demonstrating their real capability and discourage rich responses. Therefore, this paper proposes the use of semantics-based evaluators for assessing unconstrained open-ended responses on VQA datasets. As characteristics of VQA have made such evaluation significantly different from the traditional Semantic Textual Similarity (STS) task, to systematically analyze the behaviour and compare the performance of various evaluators including LLM-based ones, we propose three key properties, i.e., Alignment, Consistency and Generalization, and a corresponding dataset Assessing VQA Evaluators (AVE) to facilitate analysis. In addition, this paper proposes a Semantically Flexible VQA Evaluator (SFVE) with meticulous design based on the unique features of VQA evaluation.Experimental results verify the feasibility of model-based VQA evaluation and effectiveness of the proposed evaluator that surpasses existing semantic evaluators by a large margin. The proposed training scheme generalizes to both the BERT-like encoders and decoder-only LLM. Relaed codes and data available at https://github.com/jihuishan/flexible_evaluation_for_vqa_mm24. Huishan Ji, Qingyi Si, Zheng Lin 0001, Weiping Wang 0005 |
ACM Multimedia | 2 |
| 2024 | Cross-modality Multiple Relations Learning for Knowledge-based Visual Question AnsweringabstractKnowledge-based visual question answering not only needs to answer the questions based on images but also incorporates external knowledge to study reasoning in the joint space of vision and language. To bridge the gap between visual content and semantic cues, it is important to capture the question-related and semantics-rich vision-language connections. Most existing solutions model simple intra-modality relation or represent cross-modality relation using a single vector, which makes it difficult to effectively model complex connections between visual features and question features. Thus, we propose a cross-modality multiple relations learning model, aiming to better enrich cross-modality representations and construct advanced multi-modality knowledge triplets. First, we design a simple yet effective method to generate multiple relations that represent the rich cross-modality relations. The various cross-modality relations link the textual question to the related visual objects. These multi-modality triplets efficiently align the visual objects and corresponding textual answers. Second, to encourage multiple relations to better align with different semantic relations, we further formulate a novel global-local loss. The global loss enables the visual objects and corresponding textual answers close to each other through cross-modality relations in the vision-language space, and the local loss better preserves semantic diversity among multiple relations. Experimental results on the Outside Knowledge VQA and Knowledge-Routed Visual Question Reasoning datasets demonstrate that our model outperforms the state-of-the-art methods. Yan Wang 0028, Peize Li, Qingyi Si, Hanwen Zhang 0010, Wenyu Zang, Zheng Lin 0001, Peng Fu 0008 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | Combo of Thinking and Observing for Outside-Knowledge VQAabstractOutside-knowledge visual question answering is a challenging task that requires both the acquisition and the use of open-ended real-world knowledge.Some existing solutions draw external knowledge into the cross-modality space which overlooks the much vaster textual knowledge in natural-language space, while others transform the image into a text that further fuses with the textual knowledge into the natural-language space and completely abandons the use of visual features.In this paper, we are inspired to constrain the cross-modality space into the same space of natural-language space which makes the visual features preserved directly, and the model still benefits from the vast knowledge in natural-language space.To this end, we propose a novel framework consisting of a multimodal encoder, a textual encoder and an answer decoder.Such structure allows us to introduce more types of knowledge including explicit and implicit multimodal and textual knowledge.Extensive experiments validate the superiority of the proposed method which outperforms the state-ofthe-art by 6.17% accuracy.We also conduct comprehensive ablations of each component, and systematically study the roles of varying types of knowledge.Codes and knowledge data can be found at https://github.com/ PhoebusSi/Thinking-while-Observing. Qingyi Si, Yuchen Mo, Zheng Lin 0001, Huishan Ji, Weiping Wang 0005 |
ACL (1) | 1 |
| 2023 | Compressing and Debiasing Vision-Language Pre-Trained Models for Visual Question AnsweringabstractDespite the excellent performance of visionlanguage pre-trained models (VLPs) on conventional VQA task, they still suffer from two problems: First, VLPs tend to rely on language biases in datasets and fail to generalize to outof-distribution (OOD) data.Second, they are inefficient in terms of memory footprint and computation.Although promising progress has been made in both problems, most existing works tackle them independently.To facilitate the application of VLP to VQA tasks, it is imperative to jointly study VLP compression and OOD robustness, which, however, has not yet been explored.This paper investigates whether a VLP can be compressed and debiased simultaneously by searching sparse and robust subnetworks.To this end, we systematically study the design of a training and compression pipeline to search the subnetworks, as well as the assignment of sparsity to different modality-specific modules.Our experiments involve 3 VLPs, 2 compression methods, 4 training methods, 2 datasets and a range of sparsity levels.Our results show that there indeed exist sparse and robust subnetworks, which are competitive with the debiased full VLP and clearly outperform the debiasing SoTAs with fewer parameters on OOD datasets VQA-CP v2 and VQA-VS. 1 Qingyi Si, Yuanxin Liu, Zheng Lin 0001, Peng Fu 0008, Yanan Cao 0001, Weiping Wang 0005 |
EMNLP | 1 |
| 2023 | Schema Item Matters in Knowledge Base Question AnsweringabstractKnowledge base question answering is a challenging task that aims to answer questions by querying knowledge bases. Recently, state-of-the-art methods tend to include an enumerator module and a ranker module. They first enumerate by searching the knowledge base and then rank the candidates to select the target logical form. However, these methods sometimes fail to cover the candidates which involve more complex combinations. A recent solution to this issue is to add a generator module after the ranker to generate the uncovered target logical form. However, the enumerator and ranker always discard partial ground truth schema items. Consequently, the lack of them in the generator input results in the failure to generate the target logical form. To address this problem, we present a novel framework, SIMQA, to reuse the neglected schema items, i.e., classes and relations. Specifically, we adopt a matcher module to select the most related schema items for the given question, and feed them to the generator. On this basis, we propose a novel generation model based on contrastive learning to force the model to focus on the supplemental schema items. Experiment results on GRAILQA and WEBQSP datasets demonstrate the highly competitive performance of the proposed method, and verify that schema item matters in KBQA. Zhe Wen, Qingyi Si, Zheng Lin 0001, Peng Fu 0008, Weiping Wang 0005 |
IJCNN | 2 |
| 2022 | Visual Dialog for Spotting the Differences between Pairs of Similar ImagesabstractVisual dialog has witnessed great progress after introducing various vision-oriented goals into the conversation. Much of previous work focuses on tasks where only one image can be accessed by two interlocutors, such as VisDial and GuessWhat. The work on situations where two interlocutors access different images has received less attention. Those situations are common in real world and bring some different challenges compared with one-image tasks. The lack of such types of dialog tasks and corresponding large-scale datasets makes it impossible to carry out in-depth research. This paper therefore first proposes a new visual dialog task named Dial-the-Diff, where two interlocutors accessing two similar images respectively try to spot the difference between the images through conversing in natural language. The task raises new challenges to the dialog strategy and the ability of categorizing objects. We then build a large-scale multi-modal dataset for the task, named DialDiff, which contains 87k Virtual Reality images and 78k dialogs. Some details of the data are given and analyzed to highlight the challenges behind the task. Finally, we propose benchmark models for this task, and conduct extensive experiments to evaluate their performance as well as its problems remained. Duo Zheng, Fandong Meng, Qingyi Si, Hairun Fan, Zipeng Xu, Jie Zhou 0016, Fangxiang Feng, Xiaojie Wang 0006 |
ACM Multimedia | 3 |
| 2021 | Check It Again: Progressive Visual Question Answering via Visual EntailmentabstractQingyi Si, Zheng Lin, Ming yu Zheng, Peng Fu, Weiping Wang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Qingyi Si, Zheng Lin 0001, Mingyu Zheng, Peng Fu 0008, Weiping Wang 0005 |
ACL/IJCNLP (1) | 1 |
| 2021 | Learning Class-Transductive Intent Representations for Zero-shot Intent DetectionabstractZero-shot intent detection (ZSID) aims to deal with the continuously emerging intents without annotated training data. However, existing ZSID systems suffer from two limitations: 1) They are not good at modeling the relationship between seen and unseen intents. 2) They cannot effectively recognize unseen intents under the generalized intent detection (GZSID) setting. A critical problem behind these limitations is that the representations of unseen intents cannot be learned in the training stage. To address this problem, we propose a novel framework that utilizes unseen class labels to learn Class-Transductive Intent Representations (CTIR). Specifically, we allow the model to predict unseen intents during training, with the corresponding label names serving as input utterances. On this basis, we introduce a multi-task learning objective, which encourages the model to learn the distinctions among intents, and a similarity scorer, which estimates the connections among intents more accurately. CTIR is easy to implement and can be integrated with existing ZSID and GZSID methods. Experiments on two real-world datasets show that CTIR brings considerable improvement to the baseline systems. Qingyi Si, Yuanxin Liu, Peng Fu 0008, Zheng Lin 0001, Weiping Wang 0005 |
IJCAI | 1 |
| 2019 | A Multi-channel Neural Network for Imbalanced Emotion RecognitionabstractImbalanced issue becomes one of major bottleneck for further popularizing of emotion recognition in actual applications. Recently, some resampling methods have been proposed to improve performance by balancing the training samples. However, over-sampling methods may lead to overfitting, and undersampling methods would lose useful emotion information. In this paper, we propose a multi-channel deep architecture to improve performance in both samples and features imbalance. Specifically, we design a class correction loss function to overcome the gap between majority and minority emotions. Meanwhile, emotionspecific word embedding and a fine-tuning BERT are used to increase the differentiation of emotion words and sentences. Experimental results on two Chinese micro-blog emotion classification datasets show that our proposed architecture outperforms state-of-the-art in imbalanced emotion recognition. Qingyi Si, Peng Fu 0008, Zheng Lin 0001, Weiping Wang 0005 |
ICTAI | 2 |