Weiqi Wang 0001

dblp:51/5775-1 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
19since 2021 · last 2026
0000-0002-1617-9805ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 6 first-author · 16 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 InferenceDynamics: Adaptive LLM Routing through Structured Capability and Knowledge Profiling
abstract
Haochen Shi, Tianshi Zheng, Weiqi Wang, Baixuan Xu, Chunyang Li, Chunkit Chan, Tao Fan, Yangqiu Song. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Tianshi Zheng, Weiqi Wang 0001, Baixuan Xu, Chunkit Chan, Tao Fan 0002, Yangqiu Song
ACL (1)3
2026 arXiv2Table: Toward Realistic Benchmarking and Evaluation for LLM-Based Literature-Review Table Generation
abstract
Literature review tables are essential for summarizing and comparing collections of scientific papers.In this paper, we study automatic generation of such tables from a pool of papers to satisfy a user's information need.Building on recent work (Newman et al., 2024), we move beyond oracle settings by (i) simulating wellspecified yet schema-agnostic user demands that avoid leaking gold column names or values, (ii) explicitly modeling retrieval noise via semantically related but out-of-scope distractor papers verified by human annotators, and (iii) introducing a lightweight, annotation-free, utilization-oriented evaluation that decomposes utility (schema coverage, unary cell fidelity, pairwise relational consistency) and measures paper selection via a two-way QA procedure (gold→system and system→gold) with recall, precision, and F1.To support reproducible evaluation, we introduce ARXIV2TABLE, a benchmark of 1,957 tables referencing 7,158 papers, with human-verified distractors and rewritten, schema-agnostic user demands.We also develop an iterative, batch-based generation method that co-refines paper filtering and schema over multiple rounds.We validate the evaluation protocol with human audits and cross-evaluator checks.Extensive experiments show that our method consistently improves over strong baselines, while absolute scores remain modest, underscoring the task's difficulty.Our data and code is available at https://github.com
Weiqi Wang 0001, Jiefu Ou, Yangqiu Song, Benjamin Van Durme, Daniel Khashabi
ACL (1)1
2025 Controllable Style Arithmetic with Language Models
abstract
Language models have shown remarkable capabilities in text generation, but precisely controlling their linguistic style remains challenging. Existing methods either lack fine-grained control, require extensive computation, or introduce significant latency. We propose Style Arithmetic (SA), a novel parameter-space approach that first extracts style-specific representations by analyzing parameter differences between models trained on contrasting styles, then incorporates these representations into a base model with precise control over style intensity. Our experiments show that SA achieves three key capabilities: controllability for precise adjustment of styles, transferability for effective style transfer across tasks, and composability for simultaneous control of multiple style dimensions. Compared to alternative methods, SA offers superior effectiveness while achieving optimal computational efficiency. Our approach opens new possibilities for flexible and efficient style control in language models.
Weiqi Wang 0001, Wengang Zhou 0001, Zongmeng Zhang, Houqiang Li
ACL (1)1
2025 MARS: Benchmarking the Metaphysical Reasoning Abilities of Language Models with a Multi-task Evaluation Dataset
abstract
To enable Large Language Models (LLMs) to function as conscious agents with generalizable reasoning capabilities, it is crucial that they possess the ability to comprehend situational changes (transitions) in distribution triggered by environmental factors or actions from other agents.Despite its fundamental significance, this ability remains underexplored due to the complexity of modeling infinite possible changes in an event and their associated distributions, coupled with the lack of benchmark data with situational transitions.Addressing these gaps, we propose a novel formulation of reasoning with distributional changes as a three-step discriminative process, termed as MetAphysical ReaSoning.We then introduce the first-ever benchmark, MARS, comprising three tasks corresponding to each step.These tasks systematically assess LLMs' capabilities in reasoning the plausibility of (i) changes in actions, (ii) states caused by changed actions, and (iii) situational transitions driven by changes in action.Extensive evaluations with 20 (L)LMs of varying sizes and methods indicate that all three tasks in this process pose significant challenges, even after fine-tuning.Further analyses reveal potential causes for the underperformance of LLMs and demonstrate that pre-training on largescale conceptualization taxonomies can potentially enhance LMs' metaphysical reasoning capabilities.Our data and models are publicly accessible at https://github.com/HKUST- KnowComp/MARS.
Weiqi Wang 0001, Yangqiu Song
ACL (1)1
2025 EcomScriptBench: A Multi-task Benchmark for E-commerce Script Planning via Step-wise Intention-Driven Product Association
abstract
Weiqi Wang, Limeng Cui, Xin Liu, Sreyashi Nag, Wenju Xu, Chen Luo, Sheikh Muhammad Sarwar, Yang Li, Hansu Gu, Hui Liu, Changlong Yu, Jiaxin Bai, Yifan Gao, Haiyang Zhang, Qi He, Shuiwang Ji, Yangqiu Song. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Weiqi Wang 0001, Limeng Cui, Xin Liu 0039, Sreyashi Nag, Wenju Xu, Chen Luo 0003, Sheikh Muhammad Sarwar, Yang Li 0055, Hansu Gu, Hui Liu 0033, Changlong Yu, Jiaxin Bai, Yifan Gao 0001, Qi He 0002, Shuiwang Ji, Yangqiu Song
ACL (1)1
2025 From Automation to Autonomy: A Survey on Large Language Models in Scientific Discovery
abstract
Large Language Models (LLMs) are catalyzing a paradigm shift in scientific discovery, evolving from task-specific automation tools into increasingly autonomous agents and fundamentally redefining research processes and human-AI collaboration.This survey systematically charts this burgeoning field, placing a central focus on the changing roles and escalating capabilities of LLMs in science.Through the lens of the scientific method, we introduce a foundational three-level taxonomy-Tool, Analyst, and Scientist-to delineate their escalating autonomy and evolving responsibilities within the research lifecycle.We further identify pivotal challenges and future research trajectories such as robotic automation, self-improvement, and ethical governance.Overall, this survey provides a conceptual architecture and strategic foresight to navigate and shape the future of AI-driven scientific discovery, fostering both rapid innovation and responsible advancement.
Tianshi Zheng, Zheye Deng, Hong Ting Tsang, Weiqi Wang 0001, Jiaxin Bai, Zihao Wang 0001, Yangqiu Song
EMNLP4
2024 CANDLE: Iterative Conceptualization and Instantiation Distillation from Large Language Models for Commonsense Reasoning
abstract
Weiqi Wang, Tianqing Fang, Chunyang Li, Haochen Shi, Wenxuan Ding, Baixuan Xu, Zhaowei Wang, Jiaxin Bai, Xin Liu, Cheng Jiayang, Chunkit Chan, Yangqiu Song. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Weiqi Wang 0001, Tianqing Fang, Wenxuan Ding 0001, Baixuan Xu, Zhaowei Wang 0003, Jiaxin Bai, Xin Liu 0039, Cheng Jiayang, Chunkit Chan, Yangqiu Song
ACL (1)1
2024 Chain-of-Choice Hierarchical Policy Learning for Conversational Recommendation
Wei Fan 0001, Weiqi Wang 0001, Yangqiu Song
DASFAA (5)3
2024 Text-Tuple-Table: Towards Information Integration in Text-to-Table Generation via Global Tuple Extraction
abstract
The task of condensing large chunks of textual information into concise and structured tables has gained attention recently due to the emergence of Large Language Models (LLMs) and their potential benefit for downstream tasks, such as text summarization and text mining.Previous approaches often generate tables that directly replicate information from the text, limiting their applicability in broader contexts, as text-to-table generation in real-life scenarios necessitates information extraction, reasoning, and integration.However, there is a lack of both datasets and methodologies towards this task.In this paper, we introduce LIVESUM, a new benchmark dataset created for generating summary tables of competitions based on real-time commentary texts.We evaluate the performances of state-of-the-art LLMs on this task in both fine-tuning and zero-shot settings, and additionally propose a novel pipeline called T3 (Text-Tuple-Table ) to improve their performances.Extensive experimental results demonstrate that LLMs still struggle with this task even after fine-tuning, while our approach can offer substantial performance gains without explicit training.Further analyses demonstrate that our method exhibits strong generalization abilities, surpassing previous approaches on several other text-to-table datasets.
Zheye Deng, Chunkit Chan, Weiqi Wang 0001, Yuxi Sun 0010, Wei Fan 0001, Tianshi Zheng, Yauwai Yim, Yangqiu Song
EMNLP3
2024 GoldCoin: Grounding Large Language Models in Privacy Laws via Contextual Integrity Theory
abstract
Privacy issues arise prominently during the inappropriate transmission of information between entities.Existing research primarily studies privacy by exploring various privacy attacks, defenses, and evaluations within narrowly predefined patterns, while neglecting that privacy is not an isolated, context-free concept limited to traditionally sensitive data (e.g., social security numbers), but intertwined with intricate social contexts that complicate the identification and analysis of potential privacy violations.The advent of Large Language Models (LLMs) offers unprecedented opportunities for incorporating the nuanced scenarios outlined in privacy laws to tackle these complex privacy issues.However, the scarcity of open-source relevant case studies restricts the efficiency of LLMs in aligning with specific legal statutes.To address this challenge, we introduce a novel framework, GOLDCOIN 1 , designed to efficiently ground LLMs in privacy laws for judicial assessing privacy violations.Our framework leverages the theory of contextual integrity as a bridge, creating numerous synthetic scenarios grounded in relevant privacy statutes (e.g., HIPAA), to assist LLMs in comprehending the complex contexts for identifying privacy risks in the real world.Extensive experimental results demonstrate that GOLD-COIN markedly enhances LLMs' capabilities in recognizing privacy risks across real court cases, surpassing the baselines on different judicial tasks.
Wei Fan 0001, Haoran Li 0003, Zheye Deng, Weiqi Wang 0001, Yangqiu Song
EMNLP4
2024 MIND: Multimodal Shopping Intention Distillation from Large Vision-language Models for E-commerce Purchase Understanding
abstract
Baixuan Xu, Weiqi Wang, Haochen Shi, Wenxuan Ding, Huihao Jing, Tianqing Fang, Jiaxin Bai, Xin Liu, Changlong Yu, Zheng Li, Chen Luo, Qingyu Yin, Bing Yin, Long Chen, Yangqiu Song. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Baixuan Xu, Weiqi Wang 0001, Wenxuan Ding 0001, Huihao Jing, Tianqing Fang, Jiaxin Bai, Xin Liu 0039, Changlong Yu, Zheng Li 0018, Chen Luo 0003, Qingyu Yin, Long Chen 0016, Yangqiu Song
EMNLP2
2024 Miko: Multimodal Intention Knowledge Distillation from Large Language Models for Social-Media Commonsense Discovery
abstract
Social media has become ubiquitous for connecting with others, staying updated with news, expressing opinions, and finding entertainment. However, understanding the intention behind social media posts remains challenging due to the implicit and commonsense nature of these intentions, the need for cross-modality understanding of both text and images, and the presence of noisy information such as hashtags, misspelled words, and complicated abbreviations. To address these challenges, we present MIKO, a Multimodal Intention Knowledge DistillatiOn framework that collaboratively leverages a Large Language Model (LLM) and a Multimodal Large Language Model (MLLM) to uncover users' intentions. Specifically, our approach uses an MLLM to interpret the image, an LLM to extract key information from the text, and another LLM to generate intentions. By applying MIKO to publicly available social media datasets, we construct an intention knowledge base featuring 1,372K intentions rooted in 137,287 posts. Moreover, We conduct a two-stage annotation to verify the quality of the generated knowledge and benchmark the performance of widely used LLMs for intention generation, and further apply MIKO to a sarcasm detection dataset and distill a student model to demonstrate the downstream benefits of applying intention knowledge.
Feihong Lu, Weiqi Wang 0001, Yangyifei Luo, Ziqin Zhu, Qingyun Sun, Baixuan Xu, Shiqi Gao, Qian Li 0033, Yangqiu Song, Jianxin Li 0002
ACM Multimedia2
2024 Acquiring and modeling abstract commonsense knowledge via conceptualization
Mutian He 0001, Tianqing Fang, Weiqi Wang 0001, Yangqiu Song
Artif. Intell.3
2023 COLA: Contextualized Commonsense Causal Reasoning from the Causal Inference Perspective
abstract
Zhaowei Wang, Quyet V. Do, Hongming Zhang, Jiayao Zhang, Weiqi Wang, Tianqing Fang, Yangqiu Song, Ginny Wong, Simon See. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Zhaowei Wang 0003, Quyet V. Do, Hongming Zhang 0009, Jiayao Zhang 0001, Weiqi Wang 0001, Tianqing Fang, Yangqiu Song, Ginny Y. Wong, Simon See
ACL (1)5
2023 CAT: A Contextualized Conceptualization and Instantiation Framework for Commonsense Reasoning
abstract
Commonsense reasoning, aiming at endowing machines with a human-like ability to make situational presumptions, is extremely challenging to generalize.For someone who barely knows about meditation, while is knowledgeable about singing, he can still infer that meditation makes people relaxed from the existing knowledge that singing makes people relaxed by first conceptualizing singing as a relaxing event and then instantiating that event to meditation.This process, known as conceptual induction and deduction, is fundamental to commonsense reasoning while lacking both labeled data and methodologies to enhance commonsense modeling.To fill such a research gap, we propose CAT (Contextualized ConceptuAlization and InsTantiation), a semi-supervised learning framework that integrates event conceptualization and instantiation to conceptualize commonsense knowledge bases at scale.Extensive experiments show that our framework achieves state-of-the-art performances on two conceptualization tasks, and the acquired abstract commonsense knowledge can significantly improve commonsense inference modeling.Our code, data, and fine-tuned models are publicly available at https://github.com/HKUST- KnowComp/CAT. * Equal ContributionPersonX watches football game, as a result, PersonX will: feel relaxed PersonX plays with his dog, as a result, PersonX will: be happy and relaxed PersonX [observe] as a result, PersonX will: feel relaxed PersonX [relaxing event] as a result, PersonX will: feel relaxed Conceptualization (watches football game → relaxing event) Instantiation (relaxing event → plays with his dog)
Weiqi Wang 0001, Tianqing Fang, Baixuan Xu, Chun Yi Louis Bo, Yangqiu Song, Lei Chen 0002
ACL (1)1
2023 StoryAnalogy: Deriving Story-level Analogies from Large Language Models to Unlock Analogical Understanding
abstract
Cheng Jiayang, Lin Qiu, Tsz Chan, Tianqing Fang, Weiqi Wang, Chunkit Chan, Dongyu Ru, Qipeng Guo, Hongming Zhang, Yangqiu Song, Yue Zhang, Zheng Zhang. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
Cheng Jiayang, Tsz Ho Chan, Tianqing Fang, Weiqi Wang 0001, Chunkit Chan, Dongyu Ru, Qipeng Guo, Hongming Zhang 0009, Yangqiu Song, Yue Zhang 0004, Zheng Zhang 0001
EMNLP5
2023 Complex Query Answering on Eventuality Knowledge Graph with Implicit Logical Constraints
abstract
Querying knowledge graphs (KGs) using deep learning approaches can naturally leverage the reasoning and generalization ability to learn to infer better answers. Traditional neural complex query answering (CQA) approaches mostly work on entity-centric KGs. However, in the real world, we also need to make logical inferences about events, states, and activities (i.e., eventualities or situations) to push learning systems from System I to System II, as proposed by Yoshua Bengio. Querying logically from an EVentuality-centric KG (EVKG) can naturally provide references to such kind of intuitive and logical inference. Thus, in this paper, we propose a new framework to leverage neural methods to answer complex logical queries based on an EVKG, which can satisfy not only traditional first-order logic constraints but also implicit logical constraints over eventualities concerning their occurrences and orders. For instance, if we know that *Food is bad* happens before *PersonX adds soy sauce*, then *PersonX adds soy sauce* is unlikely to be the cause of *Food is bad* due to implicit temporal constraint. To facilitate consistent reasoning on EVKGs, we propose Complex Eventuality Query Answering (CEQA), a more rigorous definition of CQA that considers the implicit logical constraints governing the temporal order and occurrence of eventualities. In this manner, we propose to leverage theorem provers for constructing benchmark datasets to ensure the answers satisfy implicit logical constraints. We also propose a Memory-Enhanced Query Encoding (MEQE) approach to significantly improve the performance of state-of-the-art neural query encoders on the CEQA task.
Jiaxin Bai, Xin Liu 0039, Weiqi Wang 0001, Chen Luo 0003, Yangqiu Song
NeurIPS3
2021 Benchmarking Commonsense Knowledge Base Population with an Effective Evaluation Dataset
abstract
Reasoning over commonsense knowledge bases (CSKBs) whose elements are in the form of free-text is an important yet hard task in NLP.While CSKB completion only fills the missing links within the domain of the CSKB, CSKB population is alternatively proposed with the goal of reasoning unseen assertions from external resources.In this task, CSKBs are grounded to a large-scale eventuality (activity, state, and event) graph to discriminate whether novel triples from the eventuality graph are plausible or not.However, existing evaluations on the population task are either not accurate (automatic evaluation with randomly sampled negative examples) or of small scale (human annotation).In this paper, we benchmark the CSKB population task with a new large-scale dataset by first aligning four popular CSKBs, and then presenting a highquality human-annotated evaluation set to probe neural models' commonsense reasoning ability.We also propose a novel inductive commonsense reasoning model that reasons over graphs.Experimental results show that generalizing commonsense reasoning on unseen assertions is inherently a hard task.Models achieving high accuracy during training perform poorly on the evaluation set, with a large gap between human performance.
Tianqing Fang, Weiqi Wang 0001, Sehyun Choi, Shibo Hao, Hongming Zhang 0009, Yangqiu Song
EMNLP (1)2
2021 DISCOS: Bridging the Gap between Discourse Knowledge and Commonsense Knowledge
abstract
Commonsense knowledge is crucial for artificial intelligence systems to understand natural language. Previous commonsense knowledge acquisition approaches typically rely on human annotations (for example, ATOMIC) or text generation models (for example, COMET.) Human annotation could provide high-quality commonsense knowledge, yet its high cost often results in relatively small scale and low coverage. On the other hand, generation models have the potential to automatically generate more knowledge. Nonetheless, machine learning models often fit the training data well and thus struggle to generate high-quality novel knowledge. To address the limitations of previous approaches, in this paper, we propose an alternative commonsense knowledge acquisition framework DISCOS (from DIScourse to COmmonSense), which automatically populates expensive complex commonsense knowledge to more affordable linguistic knowledge resources. Experiments demonstrate that we can successfully convert discourse knowledge about eventualities from ASER, a large-scale discourse knowledge graph, into if-then commonsense knowledge defined in ATOMIC without any additional annotation effort. Further study suggests that DISCOS significantly outperforms previous supervised approaches in terms of novelty and diversity with comparable quality. In total, we can acquire 3.4M ATOMIC-like inferential commonsense knowledge by populating ATOMIC on the core part of ASER. Codes and data are available at https://github.com/HKUST-KnowComp/DISCOS-commonsense.
Tianqing Fang, Hongming Zhang 0009, Weiqi Wang 0001, Yangqiu Song
WWW3