EDBT 2026 Demo / reviewers in the wild / expert
Zhenghui Wang
dblp:06/7757
· DBLP profile ↗
11ranked-venue papers
3as first author
4since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Theory of computation · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Information extraction and text analysis · 29% Graph learning · 28% Deep learning architectures and training · 19% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Medical and health informatics · 100% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Kernel, tree and ensemble methods › gradient boosting
gradient boosting decision tree |
0.5 | 1 | 2021 | Task-wise Split Gradient Boosting Trees for Multi-center Diabetes Prediction · KDD 2021 |
Machine learning › Deep learning architectures and training
data augmentation |
0.4 | 1 | 2020 | Local Additivity Based Data Augmentation for Semi-supervised NER · EMNLP (1) 2020 |
Natural language and speech › Information extraction and text analysis › relation extraction
distant supervision |
0.4 | 1 | 2020 | QuAChIE: Question Answering based Chinese Information Extraction System · SIGIR 2020 |
Machine learning › Deep learning architectures and training › data augmentation
interpolation-based augmentation |
0.4 | 1 | 2020 | Local Additivity Based Data Augmentation for Semi-supervised NER · EMNLP (1) 2020 |
Natural language and speech › Information extraction and text analysis
named entity recognition |
0.4 | 1 | 2020 | Local Additivity Based Data Augmentation for Semi-supervised NER · EMNLP (1) 2020 |
Natural language and speech › Information extraction and text analysis
relation extraction |
0.4 | 1 | 2020 | QuAChIE: Question Answering based Chinese Information Extraction System · SIGIR 2020 |
Machine learning › Learning paradigms
semi-supervised learning |
0.4 | 1 | 2020 | Local Additivity Based Data Augmentation for Semi-supervised NER · EMNLP (1) 2020 |
Machine learning › Graph learning
network embedding |
0.4 | 1 | 2019 | Sampled in Pairs and Driven by Text: A New Graph Embedding Framework · WWW 2019 |
Machine learning › Graph learning › network embedding
random-walk-based embedding |
0.4 | 1 | 2019 | Sampled in Pairs and Driven by Text: A New Graph Embedding Framework · WWW 2019 |
Machine learning › Graph learning › graph representation learning
text-attributed graph embedding |
0.4 | 1 | 2019 | Sampled in Pairs and Driven by Text: A New Graph Embedding Framework · WWW 2019 |
Machine learning › Learning paradigms
multi-task learning |
0.1 | 1 | 2021 | Task-wise Split Gradient Boosting Trees for Multi-center Diabetes Prediction · KDD 2021 |
Machine learning › Graph learning
link prediction |
0.1 | 1 | 2019 | Sampled in Pairs and Driven by Text: A New Graph Embedding Framework · WWW 2019 |
Methods — techniques the papers use, named apart from their topics
multi-task learning · 1.0gradient boosting decision tree · 1.0sequence interpolation · 0.4pre-trained language model · 0.4distant supervision · 0.4consistency loss · 0.4text-driven embedding · 0.4random walk · 0.4pair sampling · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MNiST: A deep learning framework for multi-scale spatial feature modeling and cellular landscape decoding in spatial
Zhenghui Wang, Ruoyan Dai, Kaitai Han, Mengqiu Wang, Lixin Lei, Jirui Zhang, Qianjin Guo |
Knowl. Based Syst. | 1 |
| 2024 | Attention-guided variational graph autoencoders reveal heterogeneity in spatial transcriptomicsabstractThe latest breakthroughs in spatially resolved transcriptomics technology offer comprehensive opportunities to delve into gene expression patterns within the tissue microenvironment. However, the precise identification of spatial domains within tissues remains challenging. In this study, we introduce AttentionVGAE (AVGN), which integrates slice images, spatial information and raw gene expression while calibrating low-quality gene expression. By combining the variational graph autoencoder with multi-head attention blocks (MHA blocks), AVGN captures spatial relationships in tissue gene expression, adaptively focusing on key features and alleviating the need for prior knowledge of cluster numbers, thereby achieving superior clustering performance. Particularly, AVGN attempts to balance the model's attention focus on local and global structures by utilizing MHA blocks, an aspect that current graph neural networks have not extensively addressed. Benchmark testing demonstrates its significant efficacy in elucidating tissue anatomy and interpreting tumor heterogeneity, indicating its potential in advancing spatial transcriptomics research and understanding complex biological phenomena. Lixin Lei, Kaitai Han, Chaojing Shi, Zhenghui Wang, Ruoyan Dai, Mengqiu Wang, Qianjin Guo |
Briefings Bioinform. | 5 |
| 2022 | Gini index based initial coin offering mechanism
Mingyu Guo 0001, Zhenghui Wang, Yuko Sakurai |
Auton. Agents Multi Agent Syst. | 2 |
| 2021 | Task-wise Split Gradient Boosting Trees for Multi-center Diabetes PredictionabstractDiabetes prediction is an important data science application in the social healthcare domain. There exist two main challenges in the diabetes prediction task: data heterogeneity since demographic and metabolic data are of different types, data insufficiency since the number of diabetes cases in a single medical center is usually limited. To tackle the above challenges, we employ gradient boosting decision trees (GBDT) to handle data heterogeneity and introduce multi-task learning (MTL) to solve data insufficiency. To this end, Task-wise Split Gradient Boosting Trees (TSGB) is proposed for the multi-center diabetes prediction task. Specifically, we firstly introduce task gain to evaluate each task separately during tree construction, with a theoretical analysis of GBDT's learning objective. Secondly, we reveal a problem when directly applying GBDT in MTL, i.e., the negative task gain problem. Finally, we propose a novel split method for GBDT in MTL based on the task gain statistics, named task-wise split, as an alternative to standard feature-wise split to overcome the mentioned negative task gain problem. Extensive experiments on a large-scale real-world diabetes dataset and a commonly used benchmark dataset demonstrate TSGB achieves superior performance against several state-of-the-art methods. Detailed case studies further support our analysis of negative task gain problems and provide insightful findings. The proposed TSGB method has been deployed as an online diabetes risk assessment software for early diagnosis. Mingcheng Chen, Zhenghui Wang, Zhiyun Zhao, Weinan Zhang 0001, Xiawei Guo, Jian Shen 0003, Yanru Qu, Jieli Lu, Wei-Wei Tu, Yong Yu 0001, Yufang Bi, Guang Ning |
KDD | 2 |
| 2020 | Local Additivity Based Data Augmentation for Semi-supervised NERabstractNamed Entity Recognition (NER) is one of the first stages in deep language understanding yet current NER models heavily rely on humanannotated data.In this work, to alleviate the dependence on labeled data, we propose a Local Additivity based Data Augmentation (LADA) method for semi-supervised NER, in which we create virtual samples by interpolating sequences close to each other.Our approach has two variations: Intra-LADA and Inter-LADA, where Intra-LADA performs interpolations among tokens within one sentence, and Inter-LADA samples different sentences to interpolate.Through linear additions between sampled training data, LADA creates an infinite amount of labeled data and improves both entity and context learning.We further extend LADA to the semi-supervised setting by designing a novel consistency loss for unlabeled data.Experiments conducted on two NER benchmarks demonstrate the effectiveness of our methods over several strong baselines.We have publicly released our code at Jiaao Chen, Zhenghui Wang, Diyi Yang |
EMNLP (1) | 2 |
| 2020 | QuAChIE: Question Answering based Chinese Information Extraction SystemabstractIn this paper, we present the design of QuAChIE, a Question Answering based Chinese Information Extraction system. QuAChIE mainly depends on a well-trained question answering model to extract high-quality triples. The group of head entity and relation are regarded as a question given the input text as the context. For the training and evaluation of each model in the system, we build a large-scale information extraction dataset using Wikidata and Wikipedia pages by distant supervision. The advanced models implemented on top of the pre-trained language model and the enormous distant supervision data enable QuAChIE to extract relation triples from documents with cross-sentence correlations. The experimental results on the test set and the case study based on the interactive demonstration show its satisfactory Information Extraction quality on Chinese document-level texts. Dongyu Ru, Zhenghui Wang, Hao Zhou 0012, Lei Li 0005, Weinan Zhang 0001, Yong Yu 0001 |
SIGIR | 2 |
| 2019 | Sampled in Pairs and Driven by Text: A New Graph Embedding FrameworkabstractIn graphs with rich texts, incorporating textual information with structural information would benefit constructing expressive graph embeddings. Among various graph embedding models, random walk (RW)-based is one of the most popular and successful groups. However, it is challenged by two issues when applied on graphs with rich texts: (i) sampling efficiency: deriving from the training objective of RW-based models (e.g., DeepWalk and node2vec), we show that RW-based models are likely to generate large amounts of redundant training samples due to three main drawbacks. (ii) text utilization: these models have difficulty in dealing with zero-shot scenarios where graph embedding models have to infer graph structures directly from texts. To solve these problems, we propose a novel framework, namely Text-driven Graph Embedding with Pairs Sampling (TGE-PS). TGE-PS uses Pairs Sampling (PS) to improve the sampling strategy of RW, being able to reduce ~ 99% training samples while preserving competitive performance. TGE-PS uses Text-driven Graph Embedding (TGE), an inductive graph embedding approach, to generate node embeddings from texts. Since each node contains rich texts, TGE is able to generate high-quality embeddings and provide reasonable predictions on existence of links to unseen nodes. We evaluate TGE-PS on several real-world datasets, and experiment results demonstrate that TGE-PS produces state-of-the-art results on both traditional and zero-shot link prediction tasks. Yanru Qu, Zhenghui Wang, Weinan Zhang 0001, Shaodian Zhang, Yong Yu 0001 |
WWW | 3 |
| 2018 | Label-Aware Double Transfer Learning for Cross-Specialty Medical Named Entity RecognitionabstractZhenghui Wang, Yanru Qu, Liheng Chen, Jian Shen, Weinan Zhang, Shaodian Zhang, Yimei Gao, Gen Gu, Ken Chen, Yong Yu. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Zhenghui Wang, Yanru Qu, Jian Shen 0003, Weinan Zhang 0001, Shaodian Zhang, Yimei Gao, Gen Gu, Yong Yu 0001 |
NAACL-HLT | 1 |
| 2012 | Information Complexity versus Corruption and Applications to Orthogonality and Gap-Hamming
Amit Chakrabarti, Ranganath Kondapally, Zhenghui Wang |
APPROX-RANDOM | 3 |
| 2011 | Lower Bound for Envy-Free and Truthful Makespan Approximation on Related Machines
Lisa Fleischer, Zhenghui Wang |
SAGT | 2 |
| 2009 | A Parameter-Based Scheme for Service Composition in Pervasive Computing EnvironmentabstractPervasive computing, the new computing paradigm aiming at providing services anywhere at anytime, poses great challenges on dynamic service composition. Existing service composition methods can hardly meet the requirements of dynamic characteristic and heterogeneity in pervasive computing environment. In this paper, we propose a parameter-based service model to accurately describe pervasive services. Based on the model, pervasive services are aggregated in a two-layer graph according to both semantic and syntactic information of the input and output parameters. Moreover, we design a novel service composition scheme to accomplish the user task while satisfy the QoS requirements. Both theoretical analysis and simulation experiments show that this service composition mechanism is effective in pervasive environment. Zhenghui Wang, Tianyin Xu, Zhuzhong Qian, Sanglu Lu |
CISIS | 1 |