EDBT 2026 Demo / reviewers in the wild / expert
Hongqiu Wu
dblp:195/2044
· DBLP profile ↗
19ranked-venue papers
5as first author
18since 2021 · last 2025
0009-0008-1557-7829ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 5 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Game Development as Human-LLM InteractionabstractGame development is a highly specialized task that relies on a complex game engine powered by complex programming languages, preventing many gaming enthusiasts from handling it.This paper introduces the Chat Game Engine (ChatGE) powered by LLM, which allows everyone to develop a custom game using natural language through Human-LLM interaction.To enable an LLM to function as a ChatGE, we instruct it to perform the following processes in each turn: ( 1) P script : configure the game script segment based on the user's input; (2) P code : generate the corresponding code snippet based on the game script segment; (3) P utter : interact with the user, including guidance and feedback.We propose a data synthesis pipeline based on LLM to generate game script-code pairs and interactions from a few manually crafted seed data.We propose a three-stage training strategy following curriculum learning principles to transfer the dialogue-based LLM to ChatGE smoothly.We construct ChatGE for poker games as a case study and comprehensively evaluate it from two perspectives: interaction quality and code correctness.Pool def flopx(self, x): self.deck.pop()for i in range(x): self.community+=[self.deck.pop()] Jiale Hong, Hongqiu Wu, Hai Zhao 0001 |
ACL (1) | 2 |
| 2025 | X-TURING: Towards an Enhanced and Efficient Turing Test for Long-Term Dialogue AgentsabstractThe Turing test examines whether AIs exhibit human-like behaviour in natural language conversations. The traditional setting limits each participant to one message at a time and requires constant human participation. This fails to reflect a natural conversational style and hinders the evaluation of dialogue agents based on Large Language Models (LLMs) in complex and prolonged interactions. This paper proposes X-Turing, which enhances the original test with a burst dialogue pattern, allowing more dynamic exchanges using consecutive messages. It further reduces human workload by iteratively generating dialogues that simulate the long-term interaction between the agent and a human to compose the majority of the test process. With the pseudo-dialogue history, the agent then engages in a shorter dialogue with a real human, which is paired with a human-human conversation on the same topic to be judged using questionnaires. We introduce the X-Turn Pass-Rate metric to assess the human likeness of LLMs across varying durations. While LLMs like GPT-4 initially perform well, achieving pass rates of 51.9% and 38.9% during 3 turns and 10 turns of dialogues respectively, their performance drops as the dialogue progresses, which underscores the difficulty in maintaining consistency in the long term. Weiqi Wu, Hongqiu Wu, Hai Zhao 0001 |
ACL (1) | 2 |
| 2025 | Towards Enhanced Immersion and Agency for LLM-based Interactive DramaabstractLLM-based Interactive Drama is a novel AIbased dialogue scenario, where the user (i.e. the player) plays the role of a character in the story, has conversations with characters played by LLM agents, and experiences an unfolding story.This paper begins with understanding interactive drama from two aspects: Immersion-the player's feeling of being present in the story-and Agency-the player's ability to influence the story world.Both are crucial to creating an enjoyable interactive experience, while they have been underexplored in previous work.To enhance these two aspects, we first propose Playwriting-guided Generation, a novel method that helps LLMs craft dramatic stories with substantially improved structures and narrative quality.Additionally, we introduce Plot-based Reflection for LLM agents to refine their reactions to align with the player's intentions.Our evaluation relies on human judgment to assess the gains of our methods in terms of immersion and agency. Hongqiu Wu, Weiqi Wu, Jiameng Zhang, Hai Zhao 0001 |
ACL (1) | 1 |
| 2025 | Driving Chinese Spelling Correction from a Fine-Grained PerspectiveabstractThis paper explores the task: Chinese spelling correction (CSC), from a fine-grained perspec- tive by recognizing that existing evaluations lack nuanced typology for the spelling errors. This deficiency can create a misleading impres- sion of model performance, incurring an “in- visible” bottleneck hindering the advancement of CSC research. In this paper, we first cate- gorize spelling errors into six types and con- duct a fine-grained evaluation across a wide variety of models, including BERT-based mod- els and LLMs. Thus, we are able to pinpoint the underlying weaknesses of existing state-of- the-art models - utilizing contextual clues and handling co-existence of multiple typos, asso- ciated to contextual errors and multi-typo er- rors. However, these errors occur infrequently in conventional training corpus. Therefore, we introduce new error generation methods to aug- ment their occurrence, which can be leveraged to enhance the training of CSC models. We hope this work could provide fresh insight for future CSC research. Linfeng Liu 0003, Hongqiu Wu, Hai Zhao 0001 |
COLING | 2 |
| 2025 | Evolving Chinese Spelling Correction with Corrector-Verifier CollaborationabstractRecent methods address Chinese Spelling Correction (CSC) with either BERT-based models or large language models (LLMs) independently.However, both of them face challenges.BERT-based models are efficient for this task but struggle with limited generalizability to error patterns, thus failing in opendomain CSC.LLMs are advantageous in their extensive knowledge but fall into low efficiency in character-level editing.To address this dilemma, we propose Automatic Corrector Iteration (ACI), a novel model collaboration pipeline to iteratively optimize a BERT-based corrector.This pipeline is free of human annotation, by leveraging the knowledge and reasoning ability of an LLM verifier to provide useful signals for the corrector.Experimental results demonstrate that our pipeline consistently improves the model performance across iterations and significantly outperforms existing data augmentation methods, achieving comparable performance with human annotation. Linfeng Liu 0003, Hongqiu Wu, Hai Zhao 0001 |
EMNLP | 2 |
| 2025 | Cost-aware Best Arm Identification in Stochastic BanditsabstractThe best arm identification problem in multi-armed bandit model has been widely applied into many practical applications, such as spectrum sensing, online advertising, and cloud computing. Although lots of works have been devoted into this area, most of them do not consider the cost of pulling actions, i.e., a player has to pay some cost when she pulls an arm. Motivated by this, we study a ratio-based best arm identification problem, where each arm is associated with a random reward as well as a random cost. For any \(\delta\in(0,1)\) , with probability at least \(1-\delta\) , the player aims to find the arm with the largest ratio of expected reward to expected cost using as few samplings as possible. Specifically, we consider two settings: (1) the precise setting, i.e., identifying the precise optimal one; (2) the Probably Approximate Correct (PAC) setting, which identifies the \(\epsilon\) -optimal one. For the precise setting, we design the elimination-type algorithms and provide a fundamental lower bound which asymptotically matches the upper bound, while in the PAC setting, an UCB-type algorithm which amed \(\epsilon\) -RCB algorithm is proposed. We show that for all algorithms, the sample complexities, i.e., the pulling times for all arms, grow logarithmically as \(\frac{1}{\delta}\) increases. Moreover, compared to existing works, the running of our algorithms is independent of the arm-related parameters, which is more practical. Finally, we validate our theoretical results through numerical experiments. Zhida Qin, Wenhao Xue, Xiaoying Gan, Hongqiu Wu, Haiming Jin, Luoyi Fu |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2024 | Chinese Spelling Correction as Rephrasing Language ModelabstractThis paper studies Chinese Spelling Correction (CSC), which aims to detect and correct potential spelling errors in a given sentence. Current state-of-the-art methods regard CSC as a sequence tagging task and fine-tune BERT-based models on sentence pairs. However, we note a critical flaw in the process of tagging one character to another, that the correction is excessively conditioned on the error. This is opposite from human mindset, where individuals rephrase the complete sentence based on its semantics, rather than solely on the error patterns memorized before. Such a counter-intuitive learning process results in the bottleneck of generalizability and transferability of machine spelling correction. To address this, we propose Rephrasing Language Modeling (ReLM), where the model is trained to rephrase the entire sentence by infilling additional slots, instead of character-to-character tagging. This novel training paradigm achieves the new state-of-theart results across fine-tuned and zero-shot CSC benchmarks, outperforming previous counterparts by a large margin. Our method also learns transferable language representation when CSC is jointly trained with other tasks. Linfeng Liu 0003, Hongqiu Wu, Hai Zhao 0001 |
AAAI | 2 |
| 2024 | Episodic Return Decomposition by Difference of Implicitly Assigned Sub-trajectory RewardabstractReal-world decision-making problems are usually accompanied by delayed rewards, which affects the sample efficiency of Reinforcement Learning, especially in the extremely delayed case where the only feedback is the episodic reward obtained at the end of an episode. Episodic return decomposition is a promising way to deal with the episodic-reward setting. Several corresponding algorithms have shown remarkable effectiveness of the learned step-wise proxy rewards from return decomposition. However, these existing methods lack either attribution or representation capacity, leading to inefficient decomposition in the case of long-term episodes. In this paper, we propose a novel episodic return decomposition method called Diaster (Difference of implicitly assigned sub-trajectory reward). Diaster decomposes any episodic reward into credits of two divided sub-trajectories at any cut point, and the step-wise proxy rewards come from differences in expectation. We theoretically and empirically verify that the decomposed proxy reward function can guide the policy to be nearly optimal. Experimental results show that our method outperforms previous state-of-the-art methods in terms of both sample efficiency and performance. The code is available at https://github.com/HxLyn3/Diaster. Haoxin Lin, Hongqiu Wu, Junyin Ye, Yang Yu 0001 |
AAAI | 2 |
| 2024 | Unveiling Vulnerability of Self-AttentionabstractPre-trained language models (PLMs) are shown to be vulnerable to minor word changes, which poses a significant threat to real-world systems. While previous studies directly focus on manipulating word inputs, they are limited by their means of generating adversarial samples, lacking generalization to versatile real-world attacks. This paper studies the basic structure of transformer-based PLMs, the self-attention (SA) mechanism. (1) We propose a powerful perturbation technique named ‘HackAttend,’ which perturbs the attention scores within the SA matrices via meticulously crafted attention masks. We show that state-of-the-art PLMs fall into heavy vulnerability, with minor attention perturbations (1%) resulting in a very high attack success rate (98%). Our paper extends the conventional text attack of word perturbations to more general structural perturbations. (2) We introduce ‘S-Attend,’ a novel smoothing technique that effectively makes SA robust via structural perturbations. We empirically demonstrate that this simple yet effective technique achieves robust performance on par with adversarial training when facing various text attackers. Khai Jiet Liong, Hongqiu Wu, Hai Zhao 0001 |
LREC/COLING | 2 |
| 2024 | Attack Named Entity Recognition by Entity Boundary InterferenceabstractNamed Entity Recognition (NER) is a cornerstone natural language processing task while its robustness has been given little attention. This paper rethinks the principles of the conventional text attack, as they can easily violate the label consistency between the original and adversarial NER samples. This is due to the fine-grained nature of NER, as even minor word changes in the sentence can result in the emergence or mutation of any entity, producing invalid adversarial samples. To this end, we propose a novel one-word modification NER attack based on a key insight, NER models are always vulnerable to the boundary position of an entity to make their decision. We thus strategically insert a new boundary into the sentence and trigger the victim model to make a wrong recognition either on this boundary word or on other words in the sentence. We call this attack Virtual Boundary Attack (ViBA), which is shown to be remarkably effective when attacking both English and Chinese models with a 70%-90% attack success rate on state-of-the-art language models, and also significantly faster than previous methods. Hongqiu Wu, Hai Zhao 0001 |
LREC/COLING | 2 |
| 2024 | A Coin Has Two Sides: A Novel Detector-Corrector Framework for Chinese Spelling CorrectionabstractChinese Spelling Correction (CSC) stands as a foundational Natural Language Processing (NLP) task, which primarily focuses on the correction of erroneous characters in Chinese texts. Certain existing methodologies opt to disentangle the error correction process, employing an additional error detector to pinpoint error positions. However, owing to the inherent performance limitations of error detector, precision and recall are like two sides of the coin which can not be both facing up simultaneously. Furthermore, it is also worth investigating how the error position information can be judiciously applied to assist the error correction. In this paper, we introduce a novel approach based on error detector-corrector framework. Our detector is designed to yield two error detection results, each characterized by high precision and recall. Given that the occurrence of errors is context-dependent and detection outcomes may be less precise, we incorporate the error detection results into the CSC task using an innovative feature fusion strategy and a selective masking strategy. Empirical experiments conducted on mainstream CSC datasets substantiate the efficacy of our proposed method. Xiangke Zeng, Zuchao Li, Lefei Zhang, Ping Wang 0028, Hongqiu Wu, Hai Zhao 0001 |
ECAI | 5 |
| 2023 | Adversarial Self-Attention for Language UnderstandingabstractDeep neural models (e.g. Transformer) naturally learn spurious features, which create a ``shortcut'' between the labels and inputs, thus impairing the generalization and robustness. This paper advances self-attention mechanism to its robust variant for Transformer-based pre-trained language models (e.g. BERT). We propose Adversarial Self-Attention mechanism (ASA), which adversarially biases the attentions to effectively suppress the model reliance on features (e.g. specific keywords) and encourage its exploration of broader semantics. We conduct comprehensive evaluation across a wide range of tasks for both pre-training and fine-tuning stages. For pre-training, ASA unfolds remarkable performance gain compared to naive training for longer steps. For fine-tuning, ASA-empowered models outweigh naive models by a large margin considering both generalization and robustness. Hongqiu Wu, Ruixue Ding, Hai Zhao 0001, Pengjun Xie, Fei Huang 0002, Min Zhang 0005 |
AAAI | 1 |
| 2023 | Rethinking Masked Language Modeling for Chinese Spelling CorrectionabstractIn this paper, we study Chinese Spelling Correction (CSC) as a joint decision made by two separate models: a language model and an error model.Through empirical analysis, we find that fine-tuning BERT tends to over-fit the error model while under-fit the language model, resulting in poor generalization to outof-distribution error patterns.Given that BERT is the backbone of most CSC models, this phenomenon has a significant negative impact.To address this issue, we are releasing a multidomain benchmark LEMON, with higher quality and diversity than existing benchmarks, to allow a comprehensive assessment of the open domain generalization of CSC models.Then, we demonstrate that a very simple strategyrandomly masking 20% non-error tokens from the input sequence during fine-tuning -is sufficient for learning a much better language model without sacrificing the error model.This technique can be applied to any model architecture and achieves new state-of-the-art results on SIGHAN, ECSpell, and LEMON 1 . Hongqiu Wu, Hai Zhao 0001 |
ACL (1) | 1 |
| 2023 | Empower Nested Boolean Logic via Self-Supervised Curriculum LearningabstractBeyond the great cognitive powers showcased by language models, it is crucial to scrutinize whether their reasoning capabilities stem from strong generalization or merely exposure to relevant data.As opposed to constructing increasingly complex logic, this paper probes into the boolean logic, the root capability of a logical reasoner.We find that any pre-trained language models even including large language models only behave like a random selector in the face of multi-nested boolean logic, a task that humans can handle with ease.To empower language models with this fundamental capability, this paper proposes a new self-supervised learning method Curriculum Logical Reasoning (CLR), where we augment the training data with nested boolean logic chain step-by-step, and program the training from simpler logical patterns gradually to harder ones.This new training paradigm allows language models to effectively generalize to much harder and longer-hop logic, which can hardly be learned through naive training.Furthermore, we show that boolean logic is a great foundation for improving the subsequent general logical tasks 1 . Hongqiu Wu, Linfeng Liu 0003, Hai Zhao 0001, Min Zhang 0005 |
EMNLP | 1 |
| 2023 | Contrastive Learning of Functionality-Aware Code EmbeddingsabstractUsing pre-trained language models to obtain code embeddings is a common and effective practice in the field of source code comprehension. However, language models pre-trained on natural language text fail to capture some intrinsic characteristics of code snippets since programming languages are in a quite different form from natural language. In this paper, we present Functionality-aware Code Embeddings (FaCE) in terms of contrastive learning. The key idea of this work is that when comprehending a code snippet, it is the functionality that counts rather than its semantic meaning that mainly comes from its entities (e.g. names of functions, variables and classes). We construct positive samples and hard negative samples according to the functionality of code snippets, then pre-train our model by standard contrastive learning framework. Experimental results and massive analysis on two code-related benchmarks have justified the effectiveness of our proposed FaCE by outperforming the baseline models with large margins. Yiyang Li 0002, Hongqiu Wu, Hai Zhao 0001 |
ICASSP | 2 |
| 2023 | Toward Adversarial Training on Contextualized Language Representation
Hongqiu Wu, Yongxiang Liu, Hanwen Shi, Hai Zhao 0001, Min Zhang 0005 |
ICLR | 1 |
| 2023 | Adversarial Counterfactual Environment Model LearningabstractAn accurate environment dynamics model is crucial for various downstream tasks in sequential decision-making, such as counterfactual prediction, off-policy evaluation, and offline reinforcement learning.
Currently, these models were learned through empirical risk minimization (ERM) by step-wise fitting of historical transition data. This way was previously believed unreliable over long-horizon rollouts because of the compounding errors, which can lead to uncontrollable inaccuracies in predictions. In this paper, we find that the challenge extends beyond just long-term prediction errors: we reveal that even when planning with one step, learned dynamics models can also perform poorly due to the selection bias of behavior policies during data collection.
This issue will significantly mislead the policy optimization process even in identifying single-step optimal actions, further leading to a greater risk in sequential decision-making scenarios.
To tackle this problem, we introduce a novel model-learning objective called adversarial weighted empirical risk minimization (AWRM). AWRM incorporates an adversarial policy that exploits the model to generate a data distribution that weakens the model's prediction accuracy, and subsequently, the model is learned under this adversarial data distribution.
We implement a practical algorithm, GALILEO, for AWRM and evaluate it on two synthetic tasks, three continuous-control tasks, and \textit{a real-world application}. The experiments demonstrate that GALILEO can accurately predict counterfactual actions and improve various downstream tasks, including offline policy evaluation and improvement, as well as online decision-making. Xiong-Hui Chen, Yang Yu 0001, Zhengmao Zhu, Zhihua Yu, Zhenjun Chen, Chenghe Wang, Rong-Jun Qin, Hongqiu Wu, Ruijin Ding, Fangsheng Huang |
NeurIPS | 9 |
| 2022 | Semantic-Preserving Adversarial Code ComprehensionabstractBased on the tremendous success of pre-trained language models (PrLMs) for source code comprehension tasks, current literature studies either ways to further improve the performance (generalization) of PrLMs, or their robustness against adversarial attacks. However, they have to compromise on the trade-off between the two aspects and none of them consider improving both sides in an effective and practical way. To fill this gap, we propose Semantic-Preserving Adversarial Code Embeddings (SPACE) to find the worst-case semantic-preserving attacks while forcing the model to predict the correct labels under these worst cases. Experiments and analysis demonstrate that SPACE can stay robust against state-of-the-art attacks while boosting the performance of PrLMs for code. Yiyang Li 0002, Hongqiu Wu, Hai Zhao 0001 |
COLING | 2 |
| 2020 | Exploring Best Arm with Top Reward-Cost Ratio in Stochastic BanditsabstractThe best arm identification problem in multi-armed bandit model has been widely applied into many practical applications, such as spectrum sensing, online advertising, and cloud computing. Although lots of works have been devoted into this area, most of them do not consider the cost of pulling actions, i.e., a player has to pay some cost when she pulls an arm. Motivated by this, we study a ratio-based best arm identification problem, where each arm is associated with a random reward as well as a random cost. For any δ ∈ (0, 1), with probability at least 1-δ, the player aims to find the optimal arm with the largest ratio of expected reward to expected cost using as few samplings as possible. To solve this problem, we propose three algorithms: 1) a genie-aided algorithm GA; 2) the successive elimination algorithm with unknown gaps SEUG; 3) the successive elimination algorithm with unknown gaps and variance information SEUG-V, where gaps denote the differences between the optimal arm and the suboptimal arms. We show that for all three algorithms, the sample complexities, i.e., the pulling times for all arms, grow logarithmically as 1/δ increases. Moreover, compared to existing works, the running of our elimination-type algorithms is independent of the arm-related parameters, which is more practical. In addition, we also provide a fundamental lower bound for sample complexities of any algorithms under Bernoulli distributions, and show that the sample complexities of the proposed three algorithms match that of the lower bound in the sense of log 1/δ. Finally, we validate our theoretical results through numerical experiments. Zhida Qin, Xiaoying Gan, Hongqiu Wu, Haiming Jin, Luoyi Fu |
INFOCOM | 4 |