EDBT 2026 Demo / reviewers in the wild / expert
Tianqing Fang
dblp:283/4921
· DBLP profile ↗
25ranked-venue papers
4as first author
25since 2021 · last 2026
0000-0002-0186-8253ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 3 first-author · 23 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI GroundingabstractWenkai Wang, Xiyun Li, Hongcan Guo, Wenhao Yu, Tianqing Fang, Haitao Mi, Dong Yu, Shengyu Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xiyun Li, Hongcan Guo, Tianqing Fang, Haitao Mi, Shengyu Zhang 0001 |
ACL (1) | 5 |
| 2026 | WebAggregator: Enhancing Compositional Reasoning Capabilities of Deep Research Agent Foundation ModelsabstractRui Wang, Ce Zhang, Jun-Yu Ma, Jianshu Zhang, Hongru Wang, Yi Chen, Boyang Xue, Tianqing Fang, Zhisong Zhang, Hongming Zhang, Haitao Mi, Dong Yu, Kam-Fai Wong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Rui Wang 0015, Ce Zhang 0009, Jun-Yu Ma, Hongru Wang 0003, Yi Chen 0007, Boyang Xue, Tianqing Fang, Zhisong Zhang, Hongming Zhang 0009, Haitao Mi, Dong Yu 0001, Kam-Fai Wong |
ACL (1) | 8 |
| 2025 | OpenWebVoyager: Building Multimodal Web Agents via Iterative Real-World Exploration, Feedback and OptimizationabstractHongliang He, Wenlin Yao, Kaixin Ma, Wenhao Yu, Hongming Zhang, Tianqing Fang, Zhenzhong Lan, Dong Yu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Hongliang He 0002, Wenlin Yao, Kaixin Ma, Wenhao Yu 0002, Hongming Zhang 0009, Tianqing Fang, Zhen-Zhong Lan, Dong Yu 0001 |
ACL (1) | 6 |
| 2025 | Attention Entropy is a Key Factor: An Analysis of Parallel Context Encoding with Full-attention-based Pre-trained Language ModelsabstractLarge language models have shown remarkable performance across a wide range of language tasks, owing to their exceptional capabilities in context modeling. The most commonly used method of context modeling is full self-attention, as seen in standard decoder-only Transformers. Although powerful, this method can be inefficient for long sequences and may overlook inherent input structures. To address these problems, an alternative approach is parallel context encoding, which splits the context into sub-pieces and encodes them parallelly. Because parallel patterns are not encountered during training, naively applying parallel encoding leads to performance degradation. However, the underlying reasons and potential mitigations are unclear. In this work, we provide a detailed analysis of this issue and identify that unusually high attention entropy can be a key factor. Furthermore, we adopt two straightforward methods to reduce attention entropy by incorporating attention sinks and selective mechanisms. Experiments on various tasks reveal that these methods effectively lower irregular attention entropy and narrow performance gaps. We hope this study can illuminate ways to enhance context modeling mechanisms. Zhisong Zhang, Yan Wang 0060, Xinting Huang, Tianqing Fang, Hongming Zhang 0009, Chenlong Deng, Shuaiyi Li, Dong Yu 0001 |
ACL (1) | 4 |
| 2025 | WebEvolver: Enhancing Web Agent Self-Improvement with Co-evolving World ModelabstractAgent self-improvement, where agents autonomously train their underlying Large Language Model (LLM) on self-sampled trajectories, shows promising results but often stagnates in web environments due to limited exploration and under-utilization of pretrained web knowledge.To improve the performance of self-improvement, we propose a novel framework that introduces a co-evolving World Model LLM.This world model predicts the next observation based on the current observation and action within the web environment.The World Model serves dual roles: (1) as a virtual web server generating self-instructed training data to continuously refine the agent's policy, and (2) as an imagination engine during inference, enabling look-ahead simulation to guide action selection for the agent LLM.Experiments in real-world web environments (Mind2Web-Live, WebVoyager, and GAIAweb) show a 10% performance gain over existing self-evolving agents, demonstrating the efficacy and generalizability of our approach, without using any distillation from more powerful close-sourced models 1 . Tianqing Fang, Hongming Zhang 0009, Zhisong Zhang, Kaixin Ma, Wenhao Yu 0002, Haitao Mi, Dong Yu 0001 |
EMNLP | 1 |
| 2025 | Recall with Reasoning: Chain-of-Thought Distillation for Mamba's Long-Context Memory and ExtrapolationabstractMamba's theoretical infinite-context potential is limited in practice when sequences far exceed training lengths.This work explores unlocking Mamba's long-context memory ability by a simple-yet-effective method, Recall with Reasoning (RwR), by distilling chain-ofthought (CoT) summarization from a teacher model.Specifically, RwR prepends these summarization as CoT prompts during fine-tuning, teaching Mamba to actively recall and reason over long contexts.Experiments on LONG-MEMEVAL and HELMET show that RwR outperforms existing long-term memory methods on the Mamba model.Furthermore, under similar pre-training conditions, RwR improves the long-context performance of Mamba relative to comparable Transformer/hybrid baselines while preserving short-context capabilities, all without changing the architecture. Jun-Yu Ma, Tianqing Fang, Zhisong Zhang, Hongming Zhang 0009, Haitao Mi, Dong Yu 0001 |
EMNLP | 2 |
| 2025 | UniGist: Towards General and Hardware-aligned Sequence-level Long Context CompressionabstractLarge language models are increasingly capable of handling long-context inputs, but the memory overhead of KV cache remains a major bottleneck for general-purpose deployment. While many compression strategies have been explored, sequence-level compression is particularly challenging due to its tendency to lose important details. We present UniGist, a gist token-based long context compression framework that removes the need for chunk-wise training, enabling the model to learn how to compress and utilize long-range context during training. To fully exploit the sparsity, we introduce a gist shift trick that transforms the attention layout into a right-aligned block structure and develop a block-table-free sparse attention kernel based on it. UniGist further supports one-pass training and flexible chunk sizes during inference, allowing efficient and adaptive context processing. Experiments across multiple long-context tasks show that UniGist significantly improves compression quality, with especially strong performance in recalling details and long-range dependency modeling. Chenlong Deng, Zhisong Zhang, Kelong Mao, Shuaiyi Li, Tianqing Fang, Hongming Zhang 0009, Haitao Mi, Dong Yu 0001, Zhicheng Dou |
NeurIPS | 5 |
| 2024 | CANDLE: Iterative Conceptualization and Instantiation Distillation from Large Language Models for Commonsense ReasoningabstractWeiqi Wang, Tianqing Fang, Chunyang Li, Haochen Shi, Wenxuan Ding, Baixuan Xu, Zhaowei Wang, Jiaxin Bai, Xin Liu, Cheng Jiayang, Chunkit Chan, Yangqiu Song. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Weiqi Wang 0001, Tianqing Fang, Wenxuan Ding 0001, Baixuan Xu, Zhaowei Wang 0003, Jiaxin Bai, Xin Liu 0039, Cheng Jiayang, Chunkit Chan, Yangqiu Song |
ACL (1) | 2 |
| 2024 | Complex Reasoning over Logical Queries on Commonsense Knowledge GraphsabstractEvent commonsense reasoning requires the ability to reason about the relationship between events, as well as infer implicit context underlying that relationship.However, data scarcity makes it challenging for language models to learn to generate commonsense inferences for contexts and questions involving interactions between complex events.To address this demand, we present COM 2 (COMplex COMmonsense), a new dataset created by sampling multi-hop logical queries (e.g., the joint effect or cause of both event A and B, or the effect of the effect of event C) from an existing commonsense knowledge graph (CSKG), and verbalizing them using handcrafted rules and large language models into multiple-choice and text generation questions.Our experiments show that language models trained on COM 2 exhibit significant improvements in complex reasoning ability, resulting in enhanced zero-shot performance in both indomain and out-of-domain tasks for question answering and generative commonsense reasoning, without expensive human annotations. 1* Work done during internship at EPFL. 1 Code and data are available at https://github.com/ tqfang/complex-commonsense-reasoning Tianqing Fang, Zeming Chen 0001, Yangqiu Song, Antoine Bosselut |
ACL (1) | 1 |
| 2024 | AbsInstruct: Eliciting Abstraction Ability from LLMs through Explanation Tuning with Plausibility EstimationabstractZhaowei Wang, Wei Fan, Qing Zong, Hongming Zhang, Sehyun Choi, Tianqing Fang, Xin Liu, Yangqiu Song, Ginny Wong, Simon See. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Zhaowei Wang 0003, Wei Fan 0001, Qing Zong, Hongming Zhang 0009, Sehyun Choi, Tianqing Fang, Xin Liu 0039, Yangqiu Song, Ginny Y. Wong, Simon See |
ACL (1) | 6 |
| 2024 | ConstraintChecker: A Plugin for Large Language Models to Reason on Commonsense Knowledge BasesabstractQuyet V. Do, Tianqing Fang, Shizhe Diao, Zhaowei Wang, Yangqiu Song. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Quyet V. Do, Tianqing Fang, Shizhe Diao, Zhaowei Wang 0003, Yangqiu Song |
EACL (1) | 2 |
| 2024 | MIND: Multimodal Shopping Intention Distillation from Large Vision-language Models for E-commerce Purchase UnderstandingabstractBaixuan Xu, Weiqi Wang, Haochen Shi, Wenxuan Ding, Huihao Jing, Tianqing Fang, Jiaxin Bai, Xin Liu, Changlong Yu, Zheng Li, Chen Luo, Qingyu Yin, Bing Yin, Long Chen, Yangqiu Song. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Baixuan Xu, Weiqi Wang 0001, Wenxuan Ding 0001, Huihao Jing, Tianqing Fang, Jiaxin Bai, Xin Liu 0039, Changlong Yu, Zheng Li 0018, Chen Luo 0003, Qingyu Yin, Long Chen 0016, Yangqiu Song |
EMNLP | 6 |
| 2024 | Neuro-Inspired Information-Theoretic Hierarchical Perception for Multimodal LearningabstractIntegrating and processing information from various sources or modalities are critical for obtaining a comprehensive and accurate perception of the real world in autonomous systems and cyber-physical systems. Drawing inspiration from neuroscience, we develop the Information-Theoretic Hierarchical Perception (ITHP) model, which utilizes the concept of information bottleneck. Different from most traditional fusion models that incorporate all modalities identically in neural networks, our model designates a prime modality and regards the remaining modalities as detectors in the information pathway, serving to distill the flow of information. Our proposed perception model focuses on constructing an effective and compact information flow by achieving a balance between the minimization of mutual information between the latent state and the input modal state, and the maximization of mutual information between the latent states and the remaining modal states. This approach leads to compact latent state representations that retain relevant information while minimizing redundancy, thereby substantially enhancing the performance of multimodal representation learning. Experimental evaluations on the MUStARD, CMU-MOSI, and CMU-MOSEI datasets demonstrate that our model consistently distills crucial information in multimodal learning scenarios, outperforming state-of-the-art benchmarks. Remarkably, on the CMU-MOSI dataset, ITHP surpasses human-level performance in the multimodal sentiment binary classification task across all evaluation metrics (i.e., Binary Accuracy, F1 Score, Mean Absolute Error, and Pearson Correlation). Xiongye Xiao, Gengshuo Liu, Defu Cao, Yaxing Li, Tianqing Fang, Mingxi Cheng, Paul Bogdan |
ICLR | 7 |
| 2024 | Acquiring and modeling abstract commonsense knowledge via conceptualization
Mutian He 0001, Tianqing Fang, Weiqi Wang 0001, Yangqiu Song |
Artif. Intell. | 2 |
| 2024 | : Visualizing and Understanding Commonsense Reasoning Capabilities of Natural Language ModelsabstractRecently, large pretrained language models have achieved compelling performance on commonsense benchmarks. Nevertheless, it is unclear what commonsense knowledge the models learn and whether they solely exploit spurious patterns. Feature attributions are popular explainability techniques that identify important input concepts for model outputs. However, commonsense knowledge tends to be implicit and rarely explicitly presented in inputs. These methods cannot infer models' implicit reasoning over mentioned concepts. We present CommonsenseVIS, a visual explanatory system that utilizes external commonsense knowledge bases to contextualize model behavior for commonsense question-answering. Specifically, we extract relevant commonsense knowledge in inputs as references to align model behavior with human knowledge. Our system features multi-level visualization and interactive model probing and editing for different concepts and their underlying relations. Through a user study, we show that CommonsenseVIS helps NLP experts conduct a systematic and scalable visual analysis of models' relational reasoning over concepts in different situations. Xingbo Wang 0001, Renfei Huang, Zhihua Jin, Tianqing Fang, Huamin Qu |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | COLA: Contextualized Commonsense Causal Reasoning from the Causal Inference PerspectiveabstractZhaowei Wang, Quyet V. Do, Hongming Zhang, Jiayao Zhang, Weiqi Wang, Tianqing Fang, Yangqiu Song, Ginny Wong, Simon See. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Zhaowei Wang 0003, Quyet V. Do, Hongming Zhang 0009, Jiayao Zhang 0001, Weiqi Wang 0001, Tianqing Fang, Yangqiu Song, Ginny Y. Wong, Simon See |
ACL (1) | 6 |
| 2023 | CAT: A Contextualized Conceptualization and Instantiation Framework for Commonsense ReasoningabstractCommonsense reasoning, aiming at endowing machines with a human-like ability to make situational presumptions, is extremely challenging to generalize.For someone who barely knows about meditation, while is knowledgeable about singing, he can still infer that meditation makes people relaxed from the existing knowledge that singing makes people relaxed by first conceptualizing singing as a relaxing event and then instantiating that event to meditation.This process, known as conceptual induction and deduction, is fundamental to commonsense reasoning while lacking both labeled data and methodologies to enhance commonsense modeling.To fill such a research gap, we propose CAT (Contextualized ConceptuAlization and InsTantiation), a semi-supervised learning framework that integrates event conceptualization and instantiation to conceptualize commonsense knowledge bases at scale.Extensive experiments show that our framework achieves state-of-the-art performances on two conceptualization tasks, and the acquired abstract commonsense knowledge can significantly improve commonsense inference modeling.Our code, data, and fine-tuned models are publicly available at https://github.com/HKUST- KnowComp/CAT. * Equal ContributionPersonX watches football game, as a result, PersonX will: feel relaxed PersonX plays with his dog, as a result, PersonX will: be happy and relaxed PersonX [observe] as a result, PersonX will: feel relaxed PersonX [relaxing event] as a result, PersonX will: feel relaxed Conceptualization (watches football game → relaxing event) Instantiation (relaxing event → plays with his dog) Weiqi Wang 0001, Tianqing Fang, Baixuan Xu, Chun Yi Louis Bo, Yangqiu Song, Lei Chen 0002 |
ACL (1) | 2 |
| 2023 | StoryAnalogy: Deriving Story-level Analogies from Large Language Models to Unlock Analogical UnderstandingabstractCheng Jiayang, Lin Qiu, Tsz Chan, Tianqing Fang, Weiqi Wang, Chunkit Chan, Dongyu Ru, Qipeng Guo, Hongming Zhang, Yangqiu Song, Yue Zhang, Zheng Zhang. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Cheng Jiayang, Tsz Ho Chan, Tianqing Fang, Weiqi Wang 0001, Chunkit Chan, Dongyu Ru, Qipeng Guo, Hongming Zhang 0009, Yangqiu Song, Yue Zhang 0004, Zheng Zhang 0001 |
EMNLP | 4 |
| 2023 | KCTS: Knowledge-Constrained Tree Search Decoding with Token-Level Hallucination DetectionabstractLarge Language Models (LLMs) have demonstrated remarkable human-level natural language generation capabilities.However, their potential to generate misinformation, often called the hallucination problem, poses a significant risk to their deployment.A common approach to address this issue is to retrieve relevant knowledge and fine-tune the LLM with the knowledge in its input.Unfortunately, this method incurs high training costs and may cause catastrophic forgetting for multitasking models.To overcome these limitations, we propose a knowledge-constrained decoding method called KCTS (Knowledge-Constrained Tree Search), which guides a frozen LM to generate text aligned with the reference knowledge at each decoding step using a knowledge classifier score and MCTS (Monte-Carlo Tree Search).To adapt the sequence-level knowledge classifier to token-level guidance, we also propose a novel token-level hallucination detection method called RIPA (Reward Inflection Point Approximation).Our empirical results on knowledge-grounded dialogue and abstractive summarization demonstrate the strength of KCTS 1 as a plug-and-play, model-agnostic decoding method that can effectively reduce hallucinations in natural language generation. Sehyun Choi, Tianqing Fang, Zhaowei Wang 0003, Yangqiu Song |
EMNLP | 2 |
| 2023 | Doolittle: Benchmarks and Corpora for Academic Writing FormalizationabstractShizhe Diao, Yongyu Lei, Liangming Pan, Tianqing Fang, Wangchunshu Zhou, Sedrick Keh, Min-Yen Kan, Tong Zhang. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Shizhe Diao, Yongyu Lei, Liangming Pan, Tianqing Fang, Wangchunshu Zhou, Sedrick Keh, Min-Yen Kan, Tong Zhang 0001 |
EMNLP | 4 |
| 2022 | SubeventWriter: Iterative Sub-event Sequence Generation with Coherence ControllerabstractIn this paper, we propose a new task of subevent generation for an unseen process to evaluate the understanding of the coherence of subevent actions and objects.To solve the problem, we design SubeventWriter, a sub-event sequence generation framework with a coherence controller.Given an unseen process, the framework can iteratively construct the subevent sequence by generating one sub-event at each iteration.We also design a very effective coherence controller to decode more coherent sub-events.As our extensive experiments and analysis indicate, SubeventWriter 1 can generate more reliable and meaningful sub-event sequences for unseen processes. Zhaowei Wang 0003, Hongming Zhang 0009, Tianqing Fang, Yangqiu Song, Ginny Y. Wong, Simon See |
EMNLP | 3 |
| 2022 | ASER: Towards large-scale commonsense knowledge acquisition via higher-order selectional preference over eventualities
Hongming Zhang 0009, Xin Liu 0039, Haojie Pan, Haowen Ke, Jiefu Ou, Tianqing Fang, Yangqiu Song |
Artif. Intell. | 6 |
| 2021 | Probing Toxic Content in Large Pre-Trained Language ModelsabstractNedjma Ousidhoum, Xinran Zhao, Tianqing Fang, Yangqiu Song, Dit-Yan Yeung. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Nedjma Ousidhoum, Tianqing Fang, Yangqiu Song, Dit-Yan Yeung |
ACL/IJCNLP (1) | 3 |
| 2021 | Benchmarking Commonsense Knowledge Base Population with an Effective Evaluation DatasetabstractReasoning over commonsense knowledge bases (CSKBs) whose elements are in the form of free-text is an important yet hard task in NLP.While CSKB completion only fills the missing links within the domain of the CSKB, CSKB population is alternatively proposed with the goal of reasoning unseen assertions from external resources.In this task, CSKBs are grounded to a large-scale eventuality (activity, state, and event) graph to discriminate whether novel triples from the eventuality graph are plausible or not.However, existing evaluations on the population task are either not accurate (automatic evaluation with randomly sampled negative examples) or of small scale (human annotation).In this paper, we benchmark the CSKB population task with a new large-scale dataset by first aligning four popular CSKBs, and then presenting a highquality human-annotated evaluation set to probe neural models' commonsense reasoning ability.We also propose a novel inductive commonsense reasoning model that reasons over graphs.Experimental results show that generalizing commonsense reasoning on unseen assertions is inherently a hard task.Models achieving high accuracy during training perform poorly on the evaluation set, with a large gap between human performance. Tianqing Fang, Weiqi Wang 0001, Sehyun Choi, Shibo Hao, Hongming Zhang 0009, Yangqiu Song |
EMNLP (1) | 1 |
| 2021 | DISCOS: Bridging the Gap between Discourse Knowledge and Commonsense KnowledgeabstractCommonsense knowledge is crucial for artificial intelligence systems to understand natural language. Previous commonsense knowledge acquisition approaches typically rely on human annotations (for example, ATOMIC) or text generation models (for example, COMET.) Human annotation could provide high-quality commonsense knowledge, yet its high cost often results in relatively small scale and low coverage. On the other hand, generation models have the potential to automatically generate more knowledge. Nonetheless, machine learning models often fit the training data well and thus struggle to generate high-quality novel knowledge. To address the limitations of previous approaches, in this paper, we propose an alternative commonsense knowledge acquisition framework DISCOS (from DIScourse to COmmonSense), which automatically populates expensive complex commonsense knowledge to more affordable linguistic knowledge resources. Experiments demonstrate that we can successfully convert discourse knowledge about eventualities from ASER, a large-scale discourse knowledge graph, into if-then commonsense knowledge defined in ATOMIC without any additional annotation effort. Further study suggests that DISCOS significantly outperforms previous supervised approaches in terms of novelty and diversity with comparable quality. In total, we can acquire 3.4M ATOMIC-like inferential commonsense knowledge by populating ATOMIC on the core part of ASER. Codes and data are available at https://github.com/HKUST-KnowComp/DISCOS-commonsense. Tianqing Fang, Hongming Zhang 0009, Weiqi Wang 0001, Yangqiu Song |
WWW | 1 |