Han Wu 0004

dblp:13/1864-4 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0002-8008-064XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 DeepOR: A Deep Reasoning Foundation Model for Optimization Modeling
abstract
Optimization modeling plays a critical role in supporting optimal decision-making across various domains. Previous works have demonstrated that large language models (LLMs) tailored for optimization modeling have significantly automated and simplified this process. However, these models typically employ a straightforward input-output paradigm and struggle with challenging instances. In contrast, recent advances in general-purpose reasoning LLMs (RLLMs), such as DeepSeek-R1, have shown impressive capabilities in complex domains like mathematics and coding. In this paper, we introduce DeepOR, the first RLLM specifically designed for optimization modeling. Instead of directly outputting solutions, DeepOR explicitly performs multiple intermediate reasoning steps. To adapt a base LLM into an RLLM, we begin by synthesizing long chain-of-thought (CoT) data guided by a flowchart, which is automatically generated using a self-exploration algorithm. Once the training data are prepared, we employ supervised fine-tuning on the base LLM to endow it with reasoning capabilities tailored for optimization modeling. To fully leverage the model's reasoning potential, we further apply reinforcement learning with reward-shaping derived from solver feedback. Experimental results on benchmarks confirm that DeepOR consistently and significantly outperforms existing state-of-the-art approaches.
Ziyang Xiao, Yuan Jessica Wang, Xiongwei Han, Shisi Guan, Jingyan Zhu, Jingrong Xie 0001, Lilin Xu, Han Wu 0004, Wing Yin Yu, Zehua Liu, Xiaojin Fu, Gang Chen 0001, Dongxiang Zhang
AAAI8
2025 BPP-Search: Enhancing Tree of Thought Reasoning for Mathematical Modeling Problem Solving
abstract
Teng Wang, Wing Yin Yu, Zhenqi He, Zehua Liu, HaileiGong HaileiGong, Han Wu, Xiongwei Han, Wei Shi, Ruifeng She, Fangzhou Zhu, Tao Zhong. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Wing Yin Yu, Zhenqi He, Zehua Liu, HaileiGong HaileiGong, Han Wu 0004, Xiongwei Han, Ruifeng She, Fangzhou Zhu, Tao Zhong 0004
ACL (1)6
2025 Determine-Then-Ensemble: Necessity of Top-k Union for Large Language Model Ensembling
abstract
Large language models (LLMs) exhibit varying strengths and weaknesses across different tasks, prompting recent studies to explore the benefits of ensembling models to leverage their complementary advantages. However, existing LLM ensembling methods often overlook model compatibility and struggle with inefficient alignment of probabilities across the entire vocabulary. In this study, we empirically investigate the factors influencing ensemble performance, identifying model performance, vocabulary size, and response style as key determinants, revealing that compatibility among models is essential for effective ensembling. This analysis leads to the development of a simple yet effective model selection strategy that identifies compatible models. Additionally, we introduce the \textsc{Uni}on \textsc{T}op-$k$ \textsc{E}nsembling (\textsc{UniTE}), a novel approach that efficiently combines models by focusing on the union of the top-k tokens from each model, thereby avoiding the need for full vocabulary alignment and reducing computational overhead. Extensive evaluations across multiple benchmarks demonstrate that \textsc{UniTE} significantly enhances performance compared to existing methods, offering a more efficient framework for LLM ensembling.
Han Wu 0004, Sichun Luo, Xiongwei Han, Jie Liu 0022, Zhijiang Guo, Linqi Song
ICLR2
2025 A Survey of Optimization Modeling Meets LLMs: Progress and Future Directions
abstract
By virtue of its great utility in solving real-world problems, optimization modeling has been widely employed for optimal decision-making across various sectors, but it requires substantial expertise from operations research professionals. With the advent of large language models (LLMs), new opportunities have emerged to automate the procedure of mathematical modeling. This survey presents a comprehensive and timely review of recent advancements that cover the entire technical stack, including data synthesis and fine-tuning for the base model, inference frameworks, benchmark datasets, and performance evaluation. In addition, we conducted an in-depth analysis on the quality of benchmark datasets, which was found to have a surprisingly high error rate. We cleaned the datasets and constructed a new leaderboard with fair performance evaluation in terms of base LLM model and datasets. We also build an online portal that integrates resources of cleaned datasets, code and paper repository to benefit the community. Finally, we identify limitations in current methodologies and outline future research opportunities.
Ziyang Xiao, Jingrong Xie 0001, Lilin Xu, Shisi Guan, Jingyan Zhu, Xiongwei Han, Xiaojin Fu, WingYin Yu, Han Wu 0004, Qingcan Kang, Jiahui Duan, Tao Zhong 0004, Mingxuan Yuan, Yuan Wang 0003, Gang Chen 0001, Dongxiang Zhang
IJCAI9
2025 Preserving LLM Capabilities through Calibration Data Curation: From Analysis to Optimization
abstract
Post-training compression has been a widely employed approach to scale down large language model (LLM) and facilitate efficient inference. In various proposed compression methods, including pruning and quantization, calibration data plays a vital role by informing the weight importance and activation dynamic ranges. However, how calibration data impacts the LLM capability after compression is less explored. Few of the existing works, though recognizing the significance of this study, only investigate the language modeling or commonsense reasoning performance degradation from limited angles, like the data sources or sample amounts. More systematic research is still needed to examine the impacts on different LLM capabilities in terms of compositional properties and domain correspondence of calibration data. In this work, we aim at bridging this gap and further analyze underlying influencing mechanisms from the activation pattern perspective. Especially, we explore the calibration data's impacts on high-level complex reasoning capabilities, like math problem solving and code generation. Delving into the underlying mechanism, we find that the representativeness and diversity in activation space more fundamentally determine the quality of calibration data. Finally, we propose a calibration data curation framework based on such observations and analysis, enhancing the performance of existing post-training compression methods on preserving critical LLM capabilities. Our code is provided in [Link](https://github.com/BokwaiHo/COLA.git).
Bowei He, Lihao Yin, Hui-Ling Zhen, Shuqi Liu 0001, Han Wu 0004, Xiaokun Zhang 0001, Mingxuan Yuan, Chen Ma 0001
NeurIPS5
2025 Activation-Guided Consensus Merging for Large Language Models
abstract
Recent research has increasingly focused on reconciling the reasoning capabilities of System 2 with the efficiency of System 1. While existing training-based and prompt-based approaches face significant challenges in terms of efficiency and stability, model merging emerges as a promising strategy to integrate the diverse capabilities of different Large Language Models (LLMs) into a unified model. However, conventional model merging methods often assume uniform importance across layers, overlooking the functional heterogeneity inherent in neural components. To address this limitation, we propose \textbf{A}ctivation-Guided \textbf{C}onsensus \textbf{M}erging (\textbf{ACM}), a plug-and-play merging framework that determines layer-specific merging coefficients based on mutual information between activations of pre-trained and fine-tuned models. ACM effectively preserves task-specific capabilities without requiring gradient computations or additional training. Extensive experiments on Long-to-Short (L2S) and general merging tasks demonstrate that ACM consistently outperforms all baseline methods. For instance, in the case of Qwen-7B models, TIES-Merging equipped with ACM achieves a \textbf{55.3\%} reduction in response length while simultaneously improving reasoning accuracy by \textbf{1.3} points. We submit the code with the paper for reproducibility, and it will be publicly available.
Shuqi Liu 0001, Zehua Liu, Qintong Li, Xiongwei Han, Zhijiang Guo, Han Wu 0004, Linqi Song
NeurIPS8
2024 Structure-Aware Dialogue Modeling Methods for Conversational Semantic Role Labeling
abstract
Conversational semantic role labeling (CSRL) is believed to be a crucial step toward dialogue understanding. By incorporating the CSRL information into the conversational models, previous work (Xu et al., 2021) has confirmed the usefulness of CSRL to downstream conversation-based tasks, including multi-turn dialogue rewriting and multi-turn dialogue response generation. However, (Xu et al., 2021) found that the quality of the extracted CSRL structures would consequently affect the performance of downstream dialogue tasks while the performance of existing CSRL models is still unsatisfactory. There are two major problems in existing CSRL models to handle predicate-aware and conversational structural information. First, they ignore the fact that explicitly correlating the predicate and the context utterances could help the model better identify the arguments. Secondly, these models do not encode some vital conversational structural information, such as the speaker information which is necessary for modeling inter-speaker dependency. In this paper, we model the conversational structure-aware features based on three components: 1) the predicate-aware module which aims to capture rich correlations between the predicate and utterances; 2) a speaker-aware graph network which explicitly encodes the speaker-dependent information; 3) a novel structure-aware dialogue modeling method for the model warm-up. Experimental results on benchmark datasets show that our model significantly outperforms the baselines. We also examine the efficiency of our model and its effectiveness in low-resource scenarios. We find that our model can achieve better performance with less training time and training data than the existing models. In addition, further improvements are observed when applying the CSRL information extracted by our model into downstream dialogue tasks, which consistently indicates the superiority of our model.
Han Wu 0004, Kun Xu 0005, Linqi Song
IEEE ACM Trans. Audio Speech Lang. Process.1
2023 Reconstruct Before Summarize: An Efficient Two-Step Framework for Condensing and Summarizing Meeting Transcripts
abstract
Meetings typically involve multiple participants and lengthy conversations, resulting in redundant and trivial content.To overcome these challenges, we propose a two-step framework, Reconstruct before Summarize (RbS), for effective and efficient meeting summarization.RbS first leverages a self-supervised paradigm to annotate essential contents by reconstructing the meeting transcripts.Secondly, we propose a relative positional bucketing (RPB) algorithm to equip (conventional) summarization models to generate the summary.Despite the additional reconstruction process, our proposed RPB significantly compressed the input, leading to faster processing and reduced memory consumption compared to traditional summarization methods.We validate the effectiveness and efficiency of our method through extensive evaluations and analysis.On two meeting summarization datasets, AMI and ICSI, our approach outperforms previous state-of-the-art approaches without relying on large-scale pretraining or expert-grade annotating tools.
Haochen Tan, Han Wu 0004, Wei Shao 0009, Xinyun Zhang 0001, Mingjie Zhan, Zhaohui Hou, Ding Liang, Linqi Song
EMNLP2
2023 Fine-grained Conversational Decoding via Isotropic and Proximal Search
abstract
General-purpose text decoding approaches are usually adopted for dialogue response generation.Although the quality of the generated responses can be improved with dialogue-specific encoding methods, conversational decoding methods are still under-explored.Inspired by Wu et al. (2023) that a good dialogue feature space should follow the rules of locality and isotropy, we present a fine-grained conversational decoding method, termed isotropic and proximal search (IPS).Our method is designed to generate the semantic-concentrated response, while still maintaining informativeness and discrimination against the context.Experiments show that our approach outperforms existing decoding strategies in the dialogue field across both automatic and human evaluation metrics.More in-depth analyses further confirm the effectiveness of our approach.
Han Wu 0004, Qiling Xu, Linqi Song
EMNLP2
2023 Learning Locality and Isotropy in Dialogue Modeling
Han Wu 0004, Haochen Tan, Mingjie Zhan, Gangming Zhao, Shaoqing Lu, Ding Liang, Linqi Song
ICLR1
2021 CSAGN: Conversational Structure Aware Graph Network for Conversational Semantic Role Labeling
abstract
Conversational semantic role labeling (CSRL) is believed to be a crucial step towards dialogue understanding.However, it remains a major challenge for existing CSRL parser to handle conversational structural information.In this paper, we present a simple and effective architecture for CSRL which aims to address this problem.Our model is based on a conversational structure-aware graph network which explicitly encodes the speaker dependent information.We also propose a multi-task learning method to further improve the model.Experimental results on benchmark datasets show that our model with our proposed training objectives significantly outperforms previous baselines.
Han Wu 0004, Kun Xu 0005, Linqi Song
EMNLP (1)1
2021 Conversational Semantic Role Labeling
abstract
Semantic role labeling (SRL) aims to extract the arguments for each predicate in an input sentence. Traditional SRL can fail to analyze dialogues because it only works on every single sentence, while ellipsis and anaphora frequently occur in dialogues. To address this problem, we propose the conversational SRL task, where an argument can be the dialogue participants, a phrase in the dialogue history or the current sentence. As the existing SRL datasets are in the sentence level, we manually annotate semantic roles for 3000 chit-chat dialogues (27198 sentences) to boost the research in this direction. Experiments show that while traditional SRL systems (even with the help of coreference resolution or rewriting) perform poorly for analyzing dialogues, modeling dialogue histories and participants greatly helps the performance, indicating that adapting SRL to conversations is very promising for universal dialogue understanding. Our initial study by applying CSRL to two mainstream conversational tasks, dialogue response generation and dialogue context rewriting, also confirms the usefulness of CSRL.
Kun Xu 0005, Han Wu 0004, Linfeng Song, Haisong Zhang, Linqi Song, Dong Yu 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 Semantic Role Labeling Guided Multi-turn Dialogue ReWriter
abstract
For multi-turn dialogue rewriting, the capacity of effectively modeling the linguistic knowledge in dialog context and getting rid of the noises is essential to improve its performance.Existing attentive models attend to all words without prior focus, which results in inaccurate concentration on some dispensable words.In this paper, we propose to use semantic role labeling (SRL), which highlights the core semantic information of who did what to whom, to provide additional guidance for the rewriter model.Experiments show that this information significantly improves a RoBERTa-based model that already outperforms previous stateof-the-art systems.
Kun Xu 0005, Haochen Tan, Linfeng Song, Han Wu 0004, Haisong Zhang, Linqi Song, Dong Yu 0001
EMNLP (1)4