VLDB 2026 Research / reviewers in the wild / expert
Wen Dai
dblp:118/1983
· DBLP profile ↗
6ranked-venue papers
2as first author
4since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Information extraction and text analysis · 62% Language models and text generation · 38% |
Topics — the 5 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis › named entity recognition
fine-grained entity recognition |
0.9 | 1 | 2025 | OmniNER2025: Diverse and Comprehensive Fine-Grained NER Dataset and Benchmark for Chinese · SIGIR 2025 |
Natural language and speech › Information extraction and text analysis
named entity recognition |
0.9 | 1 | 2025 | OmniNER2025: Diverse and Comprehensive Fine-Grained NER Dataset and Benchmark for Chinese · SIGIR 2025 |
Natural language and speech › Language models and text generation › text generation › text rewriting
dialogue rewriting |
0.7 | 1 | 2023 | Dialogue Rewriting via Skeleton-Guided Generation · AAAI 2023 |
Natural language and speech › Language models and text generation › text generation › constrained text generation
skeleton-guided generation |
0.7 | 1 | 2023 | Dialogue Rewriting via Skeleton-Guided Generation · AAAI 2023 |
Natural language and speech › Information extraction and text analysis › named entity recognition
chinese named entity recognition |
0.3 | 1 | 2025 | OmniNER2025: Diverse and Comprehensive Fine-Grained NER Dataset and Benchmark for Chinese · SIGIR 2025 |
Methods — techniques the papers use, named apart from their topics
teacher model guidance · 0.9error analysis · 0.9skeleton-guided generation · 0.7sequence-to-sequence generation · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | OmniNER2025: Diverse and Comprehensive Fine-Grained NER Dataset and Benchmark for ChineseabstractAs Named Entity Recognition (NER) tasks have evolved, artificial intelligence has been widely applied in this field. However, most benchmarks are limited to English, making it challenging to replicate successful experiences in other languages. To expand NER to informal and diverse Chinese text scenarios, we have proposed a new large-scale Chinese NER dataset, OmniNER2025. This dataset, obtained from user posts on a popular Chinese social media platform Xiaohongshu, contains 195,568 samples and 89 categories, all manually annotated. To our knowledge, it is currently the largest Chinese open-source NER dataset in terms of sample size, category diversity, and domain coverage. This dataset is more challenging than existing Chinese NER datasets and better reflects real-world applications. The large sample size and diverse entity types provide valuable research resources. Additionally, we introduced the ERRTA tool for error analysis and teacher model guidance, significantly reducing model errors and improving performance. In the future, we will refine the ERRTA framework and explore optimization strategies to enhance the practical value of NER models. By releasing the OmniNER2025 dataset and introducing the ERRTA tool, we have advanced fine-grained NER research and improved model performance, promoting its application and development in real-world scenarios. Shuaipeng Liu, Mengting Hu 0002, Wen Dai, Xiaowei Zhao 0003, Xiujuan Xu |
SIGIR | 5 |
| 2023 | Dialogue Rewriting via Skeleton-Guided GenerationabstractDialogue rewriting aims to transform multi-turn, context-dependent dialogues into well-formed, context-independent text for most NLP systems. Previous dialogue rewriting benchmarks and systems assume a fluent and informative utterance to rewrite. Unfortunately, dialogue utterances from real-world systems are frequently noisy and with various kinds of errors that can make them almost uninformative. In this paper, we first present Real-world Dialogue Rewriting Corpus (RealDia), a new benchmark to evaluate how well current dialogue rewriting systems can deal with real-world noisy and uninformative dialogue utterances. RealDia contains annotated multi-turn dialogues from real scenes with ASR errors, spelling errors, redundancies and other noises that are ignored by previous dialogue rewriting benchmarks. We show that previous dialogue rewriting approaches are neither effective nor data-efficient to resolve RealDia. Then this paper presents Skeleton-Guided Rewriter (SGR), which can resolve the task of dialogue rewriting via a skeleton-guided generation paradigm. Experiments show that RealDia is a much more challenging benchmark for real-world dialogue rewriting, and SGR can effectively resolve the task and outperform previous approaches by a large margin. Chunlei Xin, Xianpei Han, Bo Chen 0020, Wen Dai, Le Sun 0001 |
AAAI | 6 |
| 2022 | Using vertices of a triangular irregular network to calculate slope and aspectabstractTerrain derivative calculations from triangulated irregular network (TIN)-based digital elevation models (DEMs) have been extensively explored in geomorphometry. However, most calculation methods focus on the triangulation facets of TIN-based DEMs and ignore the vertices. In fact, these vertices are the original sampling points from the terrain surface and serve as the basis for triangulation. In this study, we argue that terrain derivative calculations using TIN-based DEMs should focus on the vertices. Employing examples with slope and aspect, we applied the TIN vertex-based method to a mathematical surface and a real topography using TIN-based DEMs with a range of sampling point densities. We performed a comparative analysis of the TIN vertex-based, TIN facet-based, and grid-based methods. Assessments on the mathematical surface showed that the TIN vertex-based method achieved the highest accuracy among the three methods. Error analysis for the real landform case indicated that the TIN vertex-based method performed slightly better than the grid-based method for slope calculation and slightly worse than the grid-based method for aspect calculation. Among the three methods, the TIN facet-based method was most sensitive to error. The TIN vertex-based method can provide a reference for the slope and aspect calculation based on point clouds. Wen Dai, Liyang Xiong, Guoan Tang, Josef Strobl |
Int. J. Geogr. Inf. Sci. | 4 |
| 2021 | The Solution of Xiaomi AI Lab to the 2021 Language and Intelligence Challenge: Multi-format Information Extraction Task
Wen Dai, Xinyu Hua, Rongrong Lv, Ruipeng Bo |
NLPCC (2) | 1 |
| 2020 | Integrated edge detection and terrain analysis for agricultural terrace delineation from remote sensing imagesabstractAgricultural terraces are important for agricultural production and soil-and-water conservation. They comprise treads and risers that require manual construction and maintenance. If managed improperly, risers will collapse, causing soil loss, gully erosion, and cultivation threats. However, mapping terrace risers remains a challenge. This study presents a novel approach to automatically map terrace risers by combining remote sensing images and digital elevation models (DEMs). First, a terraced hillslope was extracted via a hill-shading method and edges in the image were detected using a Canny edge detector. Next, the DEM was used to generate the contour direction, and edges along this direction were searched and coded as candidate terrace risers via directional detection. Finally, the results of directional detection and the edge image obtained from the Canny detector were overlaid to backtrack complete terrace risers. The approach was validated using four study areas with different topographic characteristics in the Loess Plateau, China. The results verify that the approach achieves outstanding performance and robustness in mapping terrace risers. The precision, recall, and F-measure were 90.81%–97.57%, 88.53%–94.10%, and 90.13%–95.80%, respectively. This approach is flexible and applicable with freely available images and DEM sources. Wen Dai, Jiaming Na, Nan Huang 0003, Xin Yang 0005, Guoan Tang, Liyang Xiong, Fayuan Li |
Int. J. Geogr. Inf. Sci. | 1 |
| 2013 | Investigating the pattern of syndrome based on the difference of symptom network in depressionabstractIn TCM theory, the syndrome is crucial to diagnose diseases and treat patients. In syndrome identification, the relation of symptoms usually correlates with syndrome and represents the pattern of syndrome at symptomatic level. Hence, we learn models for classifying syndromes in depression using 4 different algorithms, which are naive Bayes, Bayes network, SVM and C4.5. From the results of classification, we find that the dependence of symptoms has something to do with the accuracies of syndrome classification. Then, 8 symptom networks corresponding to depression and 7 syndromes are constructed to explore the interaction profile of symptoms under syndrome. By comparing syndrome-specific symptom network to the base network of depression, we discover the enriched edges and different nodes to represent the pattern of each syndrome. Literature and symptom ranking by Fisher score demonstrate the correctness of the different nodes selected through network comparison. After all, the enriched edges and different nodes associated with a given syndrome reveal the pattern of that syndrome at symptomatic level. Jianglong Song, Wen Dai, Yibo Gao, Yunling Zhang, Zhichen Zhang, Peng Lu 0001, Rongjuan Guo |
BIBM | 3 |