Danqing Wang

dblp:226/6524 · DBLP profile ↗
← Back
14ranked-venue papers
8as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 8 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 TypedThinker: Diversify Large Language Model Reasoning with Typed Thinking
abstract
Large Language Models (LLMs) have demonstrated strong reasoning capabilities in solving complex problems. However, current approaches primarily enhance reasoning through the elaboration of thoughts while neglecting the diversity of reasoning types. LLMs typically employ deductive reasoning, proceeding step-by-step from given conditions, which limits their exploration during problem-solving. Our analysis reveals that certain problems are exclusively solvable through specific reasoning strategies like inductive, abductive, or analogical reasoning. However, incorporating diverse reasoning approaches presents two key challenges: identifying the appropriate reasoning type for each problem and exploiting this approach during problem-solving. Therefore, we propose the TypedThinker that predicts suitable reasoning types based on the problem and their previous effectiveness and provides relevant demonstrations to guide LLMs in applying these strategies. Experimental results show significant improvements across multiple benchmarks, with performance gains of 3.4\% for Mistral 7B, 6.5\% for LLaMA3 8B, and 7\% for Qwen 2 7B on logical and mathematical reasoning tasks. TypedThinker enhances LLM reasoning without requiring knowledge distillation from larger models. It can be integrated into more advanced systems like GPT-4o or specialized models like MetaMath to diversify their reasoning approaches and improve their problem-solving capabilities.
Danqing Wang, Fei Fang 0001, Lei Li 0005
ICLR1
2025 Scaling LLM Inference Efficiently with Optimized Sample Compute Allocation
abstract
Kexun Zhang, Shang Zhou, Danqing Wang, William Yang Wang, Lei Li. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Kexun Zhang, Shang Zhou, Danqing Wang, William Yang Wang, Lei Li 0005
NAACL (Long Papers)3
2024 Learning Personalized Alignment for Evaluating Open-ended Text Generation
abstract
Recent research has increasingly focused on evaluating large language models' (LLMs) alignment with diverse human values and preferences, particularly for open-ended tasks like story generation.Traditional evaluation metrics rely heavily on lexical similarity with humanwritten references, often showing poor correlation with human judgments and failing to account for alignment with the diversity of human preferences.To address these challenges, we introduce PERSE, an interpretable evaluation framework designed to assess alignment with specific human preferences.It is tuned to infer specific preferences from an in-context personal profile and evaluate the alignment between the generated content and personal preferences.PERSE enhances interpretability by providing detailed comments and fine-grained scoring, facilitating more personalized content generation.Our 13B LLaMA-2-based PERSE shows a 15.8% increase in Kendall correlation and a 13.7% rise in accuracy with zero-shot reviewers compared to GPT-4.It also outperforms GPT-4 by 46.01% in Kendall correlation on new domains, indicating its transferability 1 .
Danqing Wang, Kevin Yang, Hanlin Zhu, Andrew Cohen, Lei Li 0005, Yuandong Tian
EMNLP1
2024 Global Human-guided Counterfactual Explanations for Molecular Properties via Reinforcement Learning
abstract
Counterfactual explanations of Graph Neural Networks (GNNs) offer a powerful way to understand data that can naturally be represented by a graph structure. Furthermore, in many domains, it is highly desirable to derive data-driven global explanations or rules that can better explain the high-level properties of the models and data in question. However, evaluating global counterfactual explanations is hard in real-world datasets due to a lack of human-annotated ground truth, which limits their use in areas like molecular sciences. Additionally, the increasing scale of these datasets provides a challenge for random search-based methods. In this paper, we develop a novel global explanation model RLHEX for molecular property prediction. It aligns the counterfactual explanations with human-defined principles, making the explanations more interpretable and easy for experts to evaluate. RLHEX includes a VAE-based graph generator to generate global explanations and an adapter to adjust the latent representation space to human-defined principles. Optimized by Proximal Policy Optimization (PPO), the global explanations produced by RLHEX cover 4.12% more input graphs and reduce the distance between the counterfactual explanation set and the input set by 0.47% on average across three molecular datasets. RLHEX provides a flexible framework to incorporate different human-designed principles into the counterfactual explanation generation process, aligning these explanations with domain expertise. The code and data are released at https://github.com/dqwang122/RLHEX.
Danqing Wang, Antonis Antoniades, Kha-Dinh Luong, Edwin Zhang, Mert Kosan, Ambuj K. Singh, William Yang Wang, Lei Li 0005
KDD1
2023 Learning from Mistakes via Cooperative Study Assistant for Large Language Models
abstract
Large language models (LLMs) have demonstrated their potential to refine their generation based on their own feedback.However, the feedback from LLM itself is often inaccurate, thereby limiting its benefits.In this paper, we propose Study Assistant for Large LAnguage Model (SALAM), a novel framework with an auxiliary agent to assist the main LLM in learning from mistakes through interactive cooperation.In the gathering phase, the student assistant agent probes the main LLM, analyzes its errors, and collects the interaction in a mistake memory.During the examination phase, the study assistant provides guidelines by retrieving relevant cases to help the main LLM anticipate and avoid similar errors.We first investigate the effectiveness of a general study assistant and then customize it to provide LLMspecific guidance through imitation learning from successful guidance experiences.Our experiments on three LLMs using two challenging frameworks demonstrate that SALAM can significantly boost LLMs by an accuracy margin of up to 6.6 on BBH and 12.6 on BBQ 1 .
Danqing Wang, Lei Li 0005
EMNLP1
2023 INSTRUCTSCORE: Towards Explainable Text Generation Evaluation with Automatic Feedback
abstract
Automatically evaluating the quality of language generation is critical.Although recent learned metrics show high correlation with human judgement, these metrics do not provide explicit explanation of their verdict, nor associate the scores with defects in the generated text.To address this limitation, we present IN-STRUCTSCORE, a fine-grained explainable evaluation metric for text generation.By harnessing both explicit human instruction and the implicit knowledge of GPT-4, we fine-tune a text evaluation metric based on LLaMA, producing both a score for generated text and a human readable diagnostic report.We evaluate INSTRUCTSCORE on a variety of generation tasks, including translation, captioning, data-to-text, and commonsense generation.Experiments show that our 7B model surpasses all other unsupervised metrics, including those based on 175B GPT-3 and GPT-4.Surprisingly, our INSTRUCTSCORE, even without direct supervision from human-rated data, achieves performance levels on par with state-of-the-art metrics like COMET22, which were fine-tuned on human ratings.Prompt: You are evaluating a model output based on a reference.Reference: Normally the administration office downstairs would call me when there's a delivery.Output: Usually when there is takeaway, the management office downstairs will call.
Wenda Xu, Danqing Wang, Liangming Pan, Zhenqiao Song, Markus Freitag, William Yang Wang, Lei Li 0005
EMNLP2
2023 On Pre-training Language Model for Antibody
Danqing Wang, Hao Zhou 0012
ICLR1
2023 Accelerating Antimicrobial Peptide Discovery with Latent Structure
abstract
Antimicrobial peptides (AMPs) are promising therapeutic approaches against drug-resistant pathogens. Recently, deep generative models are used to discover new AMPs. However, previous studies mainly focus on peptide sequence attributes and do not consider crucial structure information. In this paper, we propose a latent sequence-structure model for designing AMPs (LSSAMP). LSSAMP exploits multi-scale vector quantization in the latent space to represent secondary structures (e.g. alpha helix and beta sheet). By sampling in the latent space, LSSAMP can simultaneously generate peptides with ideal sequence attributes and secondary structures. Experimental results show that the peptides generated by LSSAMP have a high probability of antimicrobial activity. Our wet laboratory experiments verified that two of the 21 candidates exhibit strong antimicrobial activity. The code is released at https://github.com/dqwang122/LSSAMP.
Danqing Wang, Zeyu Wen, Lei Li 0005, Hao Zhou 0012
KDD1
2023 ALGO: Synthesizing Algorithmic Programs with Generated Oracle Verifiers
abstract
Large language models (LLMs) excel at implementing code from functionality descriptions but struggle with algorithmic problems that require not only implementation but also identification of the suitable algorithm. Moreover, LLM-generated programs lack guaranteed correctness and require human verification. To address these challenges, we propose ALGO, a framework that synthesizes Algorithmic programs with LLM-Generated Oracles to guide the generation and verify their correctness. ALGO first generates a reference oracle by prompting an LLM to exhaustively enumerate all the combinations of relevant variables. This oracle is then utilized to guide an arbitrary search strategy in exploring the algorithm space and to verify the synthesized algorithms. Our study shows that the LLM-generated oracles are correct for 88% of the cases. With the oracles as verifiers, ALGO can be integrated with any existing code generation model in a model-agnostic manner to enhance its performance. Experiments show that when equipped with ALGO, we achieve an 8× better one-submission pass rate over the Codex model and a 2.6× better one-submission pass rate over CodeT, the current state-of-the-art model on CodeContests. We can also get 1.3× better pass rate over the ChatGPT Code Interpreter on unseen problems. The problem set we used for testing, the prompts we used, the verifier and solution programs, and the test cases generated by ALGO are available at https://github.com/zkx06111/ALGO.
Kexun Zhang, Danqing Wang, Jingtao Xia, William Yang Wang, Lei Li 0005
NeurIPS2
2021 Enhancing Scientific Papers Summarization with Citation Graph
abstract
Previous work for text summarization in scientific domain mainly focused on the content of the input document, but seldom considering its citation network. However, scientific papers are full of uncommon domain-specific terms, making it almost impossible for the model to understand its true meaning without the help of the relevant research community. In this paper, we redefine the task of scientific papers summarization by utilizing their citation graph and propose a citation graph-based summarization model CGSum which can incorporate the information of both the source paper and its references. In addition, we construct a novel scientific papers summarization dataset Semantic Scholar Network (SSN) which contains 141K research papers in different domains and 661K citation relationships. The entire dataset constitutes a large connected citation graph. Extensive experiments show that our model can achieve competitive performance when compared with the pretrained models even with a simple architecture. The results also indicates the citation graph is crucial to better understand the content of papers and generate high-quality summaries.
Chenxin An, Ming Zhong 0005, Yiran Chen 0013, Danqing Wang, Xipeng Qiu, Xuanjing Huang 0001
AAAI4
2021 CNewSum: A Large-Scale Summarization Dataset with Human-Annotated Adequacy and Deducibility Level
Danqing Wang, Jiaze Chen, Xianze Wu, Hao Zhou 0012, Lei Li 0005
NLPCC (1)1
2020 Heterogeneous Graph Neural Networks for Extractive Document Summarization
abstract
As a crucial step in extractive document summarization, learning cross-sentence relations has been explored by a plethora of approaches.An intuitive way is to put them in the graphbased neural network, which has a more complex structure for capturing inter-sentence relationships.In this paper, we present a heterogeneous graph-based neural network for extractive summarization (HETERSUMGRAPH), which contains semantic nodes of different granularity levels apart from sentences.These additional nodes act as the intermediary between sentences and enrich the cross-sentence relations.Besides, our graph structure is flexible in natural extension from a singledocument setting to multi-document via introducing document nodes.To our knowledge, we are the first one to introduce different types of nodes into graph-based neural networks for extractive document summarization and perform a comprehensive qualitative analysis to investigate their benefits.The code will be released on Github 1 .
Danqing Wang, Pengfei Liu 0003, Yining Zheng, Xipeng Qiu, Xuanjing Huang 0001
ACL1
2020 Extractive Summarization as Text Matching
abstract
This paper creates a paradigm shift with regard to the way we build neural extractive summarization systems.Instead of following the commonly used framework of extracting sentences individually and modeling the relationship between sentences, we formulate the extractive summarization task as a semantic text matching problem, in which a source document and candidate summaries will be (extracted from the original text) matched in a semantic space.Notably, this paradigm shift to semantic matching framework is well-grounded in our comprehensive analysis of the inherent gap between sentence-level and summary-level extractors based on the property of the dataset.Besides, even instantiating the framework with a simple form of a matching model, we have driven the state-of-the-art extractive result on CNN/DailyMail to a new level (44.41 in ROUGE-1).Experiments on the other five datasets also show the effectiveness of the matching framework.We believe the power of this matching-based summarization framework has not been fully exploited.To encourage more instantiations in the future, we have released our codes, processed dataset, as well as generated summaries in https://github. com/maszhongming/MatchSum.
Ming Zhong 0005, Pengfei Liu 0003, Yiran Chen 0013, Danqing Wang, Xipeng Qiu, Xuanjing Huang 0001
ACL4
2019 Searching for Effective Neural Extractive Summarization: What Works and What's Next
abstract
The recent years have seen remarkable success in the use of deep neural networks on text summarization.However, there is no clear understanding of why they perform so well, or how they might be improved.In this paper, we seek to better understand how neural extractive summarization systems could benefit from different types of model architectures, transferable knowledge and learning schemas.Additionally, we find an effective way to improve current frameworks and achieve the state-ofthe-art result on CNN/DailyMail by a large margin based on our observations and analyses.Hopefully, our work could provide more clues for future research on extractive summarization.Source code will be available on Github 1 .
Ming Zhong 0005, Pengfei Liu 0003, Danqing Wang, Xipeng Qiu, Xuanjing Huang 0001
ACL (1)3