VLDB 2026 Research / reviewers in the wild / expert
Dayong Wu
dblp:39/10712
· DBLP profile ↗
15ranked-venue papers
0as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 10 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Graph Reasoning Paradigm: Structured and Symbolic Reasoning with Topology-Aware Reinforcement Learning for Large Language ModelsabstractRunxuan Liu, Xianhao Ou, Xinyan Ma, Jiyuan Wang, Jiafeng Liang, Jiaqi Li, Tao He, Zheng Chu, Rongchuan Mu, Zekun Wang, Baoxin Wang, Dayong Wu, Ming Liu, Shijin Wang, Guoping Hu, Bing Qin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Runxuan Liu, Xianhao Ou, Xinyan Ma, Jiafeng Liang, Jiaqi Li 0004, Tao He 0014, Rongchuan Mu, Zekun Wang 0001, Baoxin Wang, Dayong Wu, Ming Liu 0004, Shijin Wang 0001, Bing Qin 0001 |
ACL (1) | 12 |
| 2026 | Question Tells You Where the Answer Is: Intention-aware Long-Context KV Cache CompressionabstractLiang Zhao, Xiaocheng Feng, Weihong Zhong, Lei Huang, Kun Zhu, Baoxin Wang, Dayong Wu, Guoping Hu, Ting Liu, Bing Qin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Weihong Zhong, Lei Huang 0021, Kun Zhu 0025, Baoxin Wang, Dayong Wu, Ting Liu 0001, Bing Qin 0001 |
ACL (1) | 7 |
| 2025 | Alleviating Hallucinations from Knowledge Misalignment in Large Language Models via Selective Abstention LearningabstractLei Huang, Xiaocheng Feng, Weitao Ma, Yuchun Fan, Xiachong Feng, Yuxuan Gu, Yangfan Ye, Liang Zhao, Weihong Zhong, Baoxin Wang, Dayong Wu, Guoping Hu, Lingpeng Kong, Tong Xiao, Ting Liu, Bing Qin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Lei Huang 0021, Weitao Ma, Yuchun Fan, Xiachong Feng, Yuxuan Gu 0004, Yangfan Ye, Weihong Zhong, Baoxin Wang, Dayong Wu, Lingpeng Kong, Tong Xiao 0001, Ting Liu 0001, Bing Qin 0001 |
ACL (1) | 11 |
| 2025 | Improving Contextual Faithfulness of Large Language Models via Retrieval Heads-Induced OptimizationabstractLei Huang, Xiaocheng Feng, Weitao Ma, Yuchun Fan, Xiachong Feng, Yangfan Ye, Weihong Zhong, Yuxuan Gu, Baoxin Wang, Dayong Wu, Guoping Hu, Bing Qin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Lei Huang 0021, Weitao Ma, Yuchun Fan, Xiachong Feng, Yangfan Ye, Weihong Zhong, Yuxuan Gu 0004, Baoxin Wang, Dayong Wu, Bing Qin 0001 |
ACL (1) | 10 |
| 2025 | Ontology-Guided Reverse Thinking Makes Large Language Models Stronger on Knowledge Graph Question AnsweringabstractRunxuan Liu, Luobei Luobei, Jiaqi Li, Baoxin Wang, Ming Liu, Dayong Wu, Shijin Wang, Bing Qin. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Runxuan Liu, Luobei Luobei, Jiaqi Li 0004, Baoxin Wang, Ming Liu 0007, Dayong Wu, Shijin Wang 0001, Bing Qin 0001 |
ACL (1) | 6 |
| 2025 | Chart2Code53: A Large-Scale Diverse and Complex Dataset for Enhancing Chart-to-Code GenerationabstractChart2Code has recently received significant attention in the multimodal community due to its potential to reduce the burden of visualization and promote a more detailed understanding of charts.However, existing Chart2Coderelated training datasets suffer from at least one of the following issues: (1) limited scale, (2) limited type coverage, and ( 3) inadequate complexity.To address these challenges, we seek more diverse sources that better align with real-world user distributions and propose dual data synthesis pipelines: (1) Synthesize based on online plotting code.(2) Synthesize based on the chart images in the academic paper.We create a large-scale Chart2Code training dataset Chart2Code53, including 53 chart types, 130K Chart-code pairs based on the pipeline.Experimental results demonstrate that even with few parameters, the model finetuned on Chart2Code53 achieves state-ofthe-art performance on multiple Chart2Code benchmarks within open-source models 1 . Tianhao Niu, Yiming Cui 0001, Baoxin Wang, Xiao Xu 0005, Qingfu Zhu, Dayong Wu, Shijin Wang 0001, Wanxiang Che |
EMNLP | 7 |
| 2024 | LM-Combiner: A Contextual Rewriting Model for Chinese Grammatical Error CorrectionabstractOver-correction is a critical problem in Chinese grammatical error correction (CGEC) task. Recent work using model ensemble methods based on voting can effectively mitigate over-correction and improve the precision of the GEC system. However, these methods still require the output of several GEC systems and inevitably lead to reduced error recall. In this light, we propose the LM-Combiner, a rewriting model that can directly modify the over-correction of GEC system outputs without a model ensemble. Specifically, we train the model on an over-correction dataset constructed through the proposed K-fold cross inference method, which allows it to directly generate filtered sentences by combining the original and the over-corrected text. In the inference stage, we directly take the original sentences and the output results of other systems as input and then obtain the filtered sentences through LM-Combiner. Experiments on the FCGEC dataset show that our proposed method effectively alleviates the over-correction of the original system (+18.2 Precision) while ensuring the error recall remains unchanged. Besides, we find that LM-Combiner still has a good rewriting performance even with small parameters and few training data, and thus can cost-effectively mitigate the over-correction of black-box GEC systems (e.g., ChatGPT). Baoxin Wang, Dayong Wu, Wanxiang Che |
LREC/COLING | 4 |
| 2023 | TiBERT: A Non-autoregressive Pre-trained Model for Text Editing
Baoxin Wang, Ziyue Wang 0002, Wanxiang Che, Dayong Wu, Shijin Wang 0001 |
NLPCC (3) | 4 |
| 2022 | CINO: A Chinese Minority Pre-trained Language ModelabstractMultilingual pre-trained language models have shown impressive performance on cross-lingual tasks. It greatly facilitates the applications of natural language processing on low-resource languages. However, there are still some languages that the current multilingual models do not perform well on. In this paper, we propose CINO (Chinese Minority Pre-trained Language Model), a multilingual pre-trained language model for Chinese minority languages. It covers Standard Chinese, Yue Chinese, and six other ethnic minority languages. To evaluate the cross-lingual ability of the multilingual model on ethnic minority languages, we collect documents from Wikipedia and news websites, and construct two text classification datasets, WCM (Wiki-Chinese-Minority) and CMNews (Chinese-Minority-News). We show that CINO notably outperforms the baselines on various classification tasks. The CINO model and the datasets are publicly available at http://cino.hfl-rc.com. Ziqing Yang 0001, Zihang Xu, Yiming Cui 0001, Baoxin Wang, Dayong Wu, Zhigang Chen 0003 |
COLING | 6 |
| 2022 | CCTC: A Cross-Sentence Chinese Text Correction Dataset for Native SpeakersabstractThe Chinese text correction (CTC) focuses on detecting and correcting Chinese spelling errors and grammatical errors. Most existing datasets of Chinese spelling check (CSC) and Chinese grammatical error correction (GEC) are focused on a single sentence written by Chinese-as-a-second-language (CSL) learners. We find that errors caused by native speakers differ significantly from those produced by non-native speakers. These differences make it inappropriate to use the existing test sets directly to evaluate text correction systems for native speakers. Some errors also require the cross-sentence information to be identified and corrected. In this paper, we propose a cross-sentence Chinese text correction dataset for native speakers. Concretely, we manually annotated 1,500 texts written by native speakers. The dataset consists of 30,811 sentences and more than 1,000,000 Chinese characters. It contains four types of errors: spelling errors, redundant words, missing words, and word ordering errors. We also test some state-of-the-art models on the dataset. The experimental results show that even the model with the best performance is 20 points lower than humans, which indicates that there is still much room for improvement. We hope that the new dataset can fill the gap in cross-sentence text correction for native Chinese speakers. Baoxin Wang, Xingyi Duan, Dayong Wu, Wanxiang Che, Zhigang Chen 0003 |
COLING | 3 |
| 2022 | Extending the Detection Range for Low-Channel Roadside LiDAR by Static Background ConstructionabstractFurther detection will increase traffic safety by perceiving unexpected traffic incidents earlier, allowing more time to make decisions and implement these decisions. Therefore, this article attempts to extend the detection range by using a low-channel roadside light detection and range sensor (LiDAR) sensor due to its low price and widespread employment in the future. The proposed method contains two major parts: static background construction and traffic objects detection. For static background construction, successive point cloud frames data were used to cover the most LiDAR scanning horizontal–vertical angle and obtain background information ultimately. Furthermore, the constructed static background was optimized to reduce the missing of the far range points and the occurrence of noise points. For vehicle and road user detection, the density-based spatial clustering of applications with noise (DBSCAN) algorithm was used to identify the near-range traffic objects. For far-range traffic objects detection, the trajectories of objects were extracted based on their moving direction and distance, and a fast Fourier transform (FFT) algorithm was used to filter the noise point and identify the vehicle and road user points. The experimental results show that the static background construction with the proposed method is more effective and uses fewer point cloud data. The average maximum detection range can be up to 100, 100, and 85 m for vehicles, cyclists, and pedestrians, respectively. The average precision (AP) and recall are 95.71% and 90.62%, respectively, which are better than the previous studies. Hui Liu 0054, Ciyun Lin, Bowen Gong, Dayong Wu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | NeurJudge: A Circumstance-aware Neural Framework for Legal Judgment PredictionabstractLegal Judgment Prediction is a fundamental task in legal intelligence of the civil law system, which aims to automatically predict the judgment results of multiple subtasks, such as charge, law article, and term of penalty prediction. Existing studies mainly focus on the impact of the entire fact description on all subtasks. They ignore the practical judicial scenario, where judges adopt circumstances of crime (i.e., various parts of the fact) to decide judgment results. To this end, in this paper, we propose a circumstance-aware legal judgment prediction framework (i.e., NeurJudge) by exploring circumstances of crime. Specifically, NeurJudge utilizes the results of intermediate subtasks to separate the fact description into different circumstances and exploits them to make the predictions of other subtasks. In addition, considering the popularity of confusing verdicts (i.e., charges and law articles), we further extend NeurJudge to a more comprehensive framework which is denoted by NeurJudge+. Particularly, NeurJudge+ utilizes a label embedding method to incorporate the semantics of labels (i.e., charges and law articles) into facts to generate more expressive fact representations for confusing verdicts problems. Extensive experimental results on two real-world datasets clearly validate the effectiveness of our proposed frameworks. Linan Yue, Qi Liu 0003, Binbin Jin, Han Wu 0002, Kai Zhang 0038, Yanqing An, Mingyue Cheng 0004, Biao Yin, Dayong Wu |
SIGIR | 9 |
| 2021 | Circumstances enhanced Criminal Court View GenerationabstractCriminal Court View Generation is an essential task in legal intelligence, which aims to automatically generate sentences interpreting judgment results. The court view could be seen as the summary of crime circumstances in a case, including ADjudging Circumstance (ADC) and SEntencing Circumstance (SEC). However, different circumstances vary widely, and adopting them to generate court views directly may limit the generation performance. Therefore, it is necessary to identify the ADC and SEC related sentences in case facts and enhance them into the court view generation, respectively. To this end, in this paper, we propose a novel Circumstances enhanced Criminal Court View Generation (C3VG) method, consisting of the extraction and generation stage. Specifically, in the extraction stage, we design a Circumstances Selector to select ADC and SEC related sentences. After that, we apply them to two generators to generate the circumstances enhanced court views, respectively. After merging the two types of court views, we could obtain the final court views. We evaluate C3VG by conducting extensive experiments on a real-world dataset and experimental results clearly validate the effectiveness of our proposed model. Linan Yue, Qi Liu 0003, Han Wu 0002, Yanqing An, Li Wang 0014, Senchao Yuan, Dayong Wu |
SIGIR | 7 |
| 2018 | Learning sequential features for cascade outbreak prediction
Chengcheng Gou, Huawei Shen, Pan Du 0001, Dayong Wu, Xueqi Cheng 0001 |
Knowl. Inf. Syst. | 4 |
| 2013 | A Self-learning Template Approach for Recognizing Named Entities from Web Text
Bingyang Liu, Dayong Wu, Xueqi Cheng 0001 |
IJCNLP | 3 |