Quanyu Long

dblp:259/0397 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0002-4839-012XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 10 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Causality Matters: How Temporal Information Emerges in Video Language Models
abstract
Video language models (VideoLMs) have made significant progress in multimodal understanding. However, temporal understanding, which involves identifying event order, duration, and relationships across time, still remains a core challenge. Prior works emphasize positional encodings (PEs) as a key mechanism for encoding temporal structure. Surprisingly, we find that removing or modifying PEs in video inputs yields minimal degradation in the performance of temporal understanding. In contrast, reversing the frame sequence while preserving the original PEs causes a substantial drop. To explain this behavior, we conduct substantial analysis experiments to trace how temporal information is integrated within the model. We uncover a causal information pathway: temporal cues are progressively synthesized through inter-frame attention, aggregated in the final frame, and subsequently integrated into the query tokens. This emergent mechanism shows that temporal reasoning emerges from inter-visual token interactions under the constraints of causal attention, which implicitly encodes temporal structure. Based on these insights, we propose two efficiency-oriented strategies: staged cross-modal attention and a temporal exit mechanism for early token truncation. Experiments on two benchmarks validate the effectiveness of both approaches.
Yumeng Shi, Quanyu Long, Yin Wu 0001, Wenya Wang 0001
AAAI2
2026 Programming over Thinking: Efficient and Robust Multi-Constraint Planning
abstract
Multi-constraint planning involves identifying, evaluating, and refining candidate plans while satisfying multiple, potentially conflicting constraints.Existing large language model (LLM) approaches face fundamental limitations in this domain.Pure reasoning paradigms, which rely on long natural language chains, are prone to inconsistency, error accumulation, and prohibitive cost as constraints compound.Conversely, LLMs combined with coding-or solver-based strategies lack flexibility: they often generate problem-specific code from scratch or depend on fixed solvers, failing to capture generalizable logic across diverse problems.To address these challenges, we introduce the Scalable COde Planning Engine (SCOPE), a framework that disentangles query-specific reasoning from generic code execution.By separating reasoning from execution, SCOPE produces solver functions that are consistent, deterministic, and reusable across queries while requiring only minimal changes to input parameters.SCOPE achieves state-of-the-art performance while lowering cost and latency.For example, with GPT-4o, it reaches 93.1% success on TravelPlanner, a 61.6% gain over the best baseline (CoT) while cutting inference cost by 1.4x and time by 4.67x.Code is available at https://github.com/DerrickGXD/SCOPE.
Derrick Goh Xin Deik, Quanyu Long, Zhengyuan Liu, Nancy F. Chen, Wenya Wang 0001
ACL (1)2
2026 Coordinating Search-Informed Reasoning and Reasoning-Guided Search in Claim Verification
abstract
Multi-hop claim verification is inherently challenging, requiring multi-step reasoning to construct verification chains while iteratively searching for information to uncover hidden bridging facts. This process is fundamentally interleaved, as effective reasoning relies on dynamically retrieved evidence, while effective search demands reasoning to refine queries based on partial information. To achieve this, we propose Hierarchical Agent Reasoning and Information Search (HARIS), explicitly modeling the coordinated process of reasoning-driven searching and search-informed reasoning. HARIS consists of a high-level reasoning agent that focuses on constructing the main verification chain, generating factual questions when more information is needed, and a low-level search agent that iteratively retrieves more information, refining its search based on intermediate findings. This design allows each agent to specialize in its respective task, enhancing verification accuracy and interpretability. HARIS is trained using reinforcement learning with outcome-based rewards. Experimental results on the EX-FEVER and HOVER benchmarks demonstrate that HARIS achieves strong performance, greatly advancing multi-hop claim verification.
Qisheng Hu, Quanyu Long, Wenya Wang 0001
ACL (1)2
2026 From Competition to Synergy: Unlocking Reinforcement Learning for Subject-Driven Image Generation
abstract
Ziwei Huang, Ying Shu, Fanghao, Quanyu Long, Wenya Wang, Qiushi Guo, Tiezheng Ge, Leilei Gan. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Ziwei Huang 0005, Yin Shu, Quanyu Long, Wenya Wang 0001, Qiushi Guo, Tiezheng Ge, Leilei Gan
ACL (1)4
2026 Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
abstract
Retrieval-augmented generation (RAG) augments large language models (LLMs) with external knowledge to tackle knowledge-intensive question answering. While several benchmarks evaluate multimodal LLMs (MLLMs) under multimodal RAG settings, they predominantly retrieve from textual corpora and do not explicitly assess how models exploit visual evidence during answer generation. Consequently, there still lacks benchmark that cleanly isolates and measures the contribution of retrieved images in a visual knowledge-intensive RAG pipeline. We introduce Visual-RAG, a question-answering benchmark that targets visually-grounded, knowledge-intensive questions in a visual evidence-centric manner. Unlike prior work, Visual-RAG requires text-to-image retrieval and the integration of retrieved clue images whose pixel content explicitly encodes the visual knowledge necessary for answer generation. With Visual-RAG, we evaluate five open-source and three proprietary MLLMs and find that current systems still substantially underutilize the visual information available in retrieved images. Despite clear opportunities for multimodal evidence integration, state-of-the-art models struggle to extract and exploit fine-grained visual knowledge, and text-to-image retrieval itself remains challenging even under constrained entity-level corpora. These results underscore the need for improved visual retrieval, grounding, and attribution in multimodal RAG. Visual-RAG is publicly available at: github.com/visual-rag/visual-rag
Yin Wu 0001, Quanyu Long, Jing Li 0034, Jianfei Yu, Wenya Wang 0001
SIGIR2
2026 A causal saliency enhancement and Mamba-based multi-scale feature fusion encoding framework for railway fastener fault diagnosis
Quanyu Long, Dechen Yao, Jia Dong, Yuanting Dai
Adv. Eng. Informatics1
2025 T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive Concepts
abstract
Ziwei Huang, Wanggui He, Quanyu Long, Yandi Wang, Haoyuan Li, Zhelun Yu, Fangxun Shu, Weilong Dai, Hao Jiang, Fei Wu, Leilei Gan. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Ziwei Huang 0005, Wanggui He, Quanyu Long, Yandi Wang, Haoyuan Li 0002, Zhelun Yu, Fangxun Shu, Weilong Dai, Hao Jiang 0014, Fei Wu 0001, Leilei Gan
ACL (1)3
2025 Static or Dynamic: Towards Query-Adaptive Token Selection for Video Question Answering
abstract
Video question answering benefits from the rich information in videos, enabling various applications.However, the large volume of tokens generated from long videos presents challenges to memory efficiency and model performance.To alleviate this, existing works propose to compress video inputs, but often overlook the varying importance of static and dynamic information across different queries, leading to inefficient token usage within limited budgets.We propose a novel token selection strategy, EXPLORE-THEN-SELECT, that adaptively adjusts static and dynamic information based on question requirements.Our framework first explores different token allocations between key frames, which preserve spatial details, and delta frames, which capture temporal changes.Then it employs a query-aware attention-based metric to select the optimal token combination without model updates.Our framework is plug-and-play and can be seamlessly integrated within diverse video language models.Extensive experiments show that our method achieves significant performance improvements (up to 5.8%) on multiple video question answering benchmarks.Our code is available at https://github.com/ANDgate99/Explore- Then-Select.
Yumeng Shi, Quanyu Long, Wenya Wang 0001
EMNLP2
2025 Decomposition Dilemmas: Does Claim Decomposition Boost or Burden Fact-Checking Performance?
abstract
Qisheng Hu, Quanyu Long, Wenya Wang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Qisheng Hu, Quanyu Long, Wenya Wang 0001
NAACL (Long Papers)2
2023 Adapt in Contexts: Retrieval-Augmented Domain Adaptation via In-Context Learning
abstract
Large language models (LLMs) have showcased their capability with few-shot inference known as in-context learning.However, indomain demonstrations are not always readily available in real scenarios, leading to crossdomain in-context learning.Besides, LLMs are still facing challenges in long-tail knowledge in unseen and unfamiliar domains.The above limitations demonstrate the necessity of Unsupervised Domain Adaptation (UDA).In this paper, we study the UDA problem under an in-context learning setting to adapt language models from the source domain to the target domain without any target labels.The core idea is to retrieve a subset of cross-domain elements that are the most similar to the query, and elicit language model to adapt in an in-context manner by learning both target domain distribution and the discriminative task signal simultaneously with the augmented cross-domain incontext examples.We devise different prompting and training strategies, accounting for different LM architectures to learn the target distribution via language modeling.With extensive experiments on Sentiment Analysis (SA) and Named Entity Recognition (NER) tasks, we thoroughly study the effectiveness of ICL for domain transfer and demonstrate significant improvements over baseline models.
Quanyu Long, Wenya Wang 0001, Sinno Jialin Pan
EMNLP1
2022 Domain Confused Contrastive Learning for Unsupervised Domain Adaptation
abstract
In this work, we study Unsupervised Domain Adaptation (UDA) in a challenging selfsupervised approach.One of the difficulties is how to learn task discrimination in the absence of target labels.Unlike previous literature which directly aligns cross-domain distributions or leverages reverse gradient, we propose Domain Confused Contrastive Learning (DCCL) to bridge the source and the target domains via domain puzzles, and retain discriminative representations after adaptation.Technically, DCCL searches for a most domainchallenging direction and exquisitely crafts domain confused augmentations as positive pairs, then it contrastively encourages the model to pull representations towards the other domain, thus learning more stable and effective domain invariances.We also investigate whether contrastive learning necessarily helps with UDA when performing other data augmentations.Extensive experiments demonstrate that DCCL significantly outperforms baselines.
Quanyu Long, Tianze Luo, Wenya Wang 0001, Sinno Jialin Pan
NAACL-HLT1
2021 Generative Imagination Elevates Machine Translation
abstract
There are common semantics shared across text and images.Given a sentence in a source language, whether depicting the visual scene helps translation into a target language?Existing multimodal neural machine translation methods (MNMT) require triplets of bilingual sentence -image for training and tuples of source sentence -image for inference.In this paper, we propose ImagiT, a novel machine translation method via visual imagination.ImagiT first learns to generate visual representation from the source sentence, and then utilizes both source sentence and the "imagined representation" to produce a target translation.Unlike previous methods, it only needs the source sentence at the inference time.Experiments demonstrate that ImagiT benefits from visual imagination and significantly outperforms the text-only neural machine translation baselines.Further analysis reveals that the imagination process in ImagiT helps fill in missing information when performing the degradation strategy.
Quanyu Long, Mingxuan Wang, Lei Li 0005
NAACL-HLT1
2020 On the Robustness of Language Encoders against Grammatical Errors
abstract
We conduct a thorough study to diagnose the behaviors of pre-trained language encoders (ELMo, BERT, and RoBERTa) when confronted with natural grammatical errors.Specifically, we collect real grammatical errors from non-native speakers and conduct adversarial attacks to simulate these errors on clean text data.We use this approach to facilitate debugging models on downstream applications.Results confirm that the performance of all tested models is affected but the degree of impact varies.To interpret model behaviors, we further design a linguistic acceptability task to reveal their abilities in identifying ungrammatical sentences and the position of errors.We find that fixed contextual encoders with a simple classifier trained on the prediction of sentence correctness are able to locate error positions.We also design a cloze test for BERT and discover that BERT captures the interaction between errors and specific tokens in context.Our results shed light on understanding the robustness and behaviors of language encoders against grammatical errors.
Fan Yin, Quanyu Long, Kai-Wei Chang 0001
ACL2