EDBT 2026 Demo / reviewers in the wild / expert
Wei Zhang 0027
dblp:10/4661-27
· DBLP profile ↗
13ranked-venue papers
1as first author
3since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4Security and privacy · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A survey of anomaly detection in HPC systems using machine learningabstractAbstract High-performance computing (HPC) systems must remain stable and reliable to consistently deliver robust computational power and ensure the proper execution of user jobs. Anomaly detection is a key means to ensure the stability and reliability of these systems. With the expansion of HPC systems and changes in their architecture, accurately identifying anomalies in dynamic environments has become increasingly challenging. Traditional detection methods rely on experience and rules, which could be inefficient and inaccurate. To address these issues, researchers have proposed machine learning-based methods to automatically process large amounts of complex data, improving the efficiency of anomaly identification and diagnosis. In this survey, we conduct a comprehensive and in-depth investigation of machine learning-based anomaly detection methods in HPC systems. Firstly, we summarize and introduce the background and challenges of anomaly detection in HPC systems. Secondly, we compare a series of machine learning-based anomaly detection works in detail and summarize their frameworks. We conclude their advantages and disadvantages and application scenarios. Finally, we discuss several promising development trends of machine learning-based HPC system anomaly detection. Wei Zhang 0027, Yiqin Dai, Huijun Wu 0001, Zhenwei Wu, Hongyun Tian, Juan Chen 0001, Chubo Liu, Yong Dong |
CCF Trans. High Perform. Comput. | 2 |
| 2026 | A Survey on Machine Learning-Based HPC I/O Analysis and OptimizationabstractThe soaring computing power of HPC systems supports numerous large-scale applications, which generate massive data volumes and diverse I/O patterns, leading to severe I/O bottlenecks. Analyzing and optimizing HPC I/O is therefore critical. However, traditional approaches are typically customized and lack the adaptability required to cope with dynamic changes in HPC environments. To address the challenge, Machine Learning (ML) has been increasingly adopted to automate and enhance I/O analysis and optimization. Given sufficient I/O traces from HPC systems, ML can learn underlying I/O behaviors, extract actionable insights, and dynamically adapt to evolving workloads to improve performance. In this survey, we propose a novel taxonomy that aligns HPC I/O problems with learning tasks to systematically review existing studies. Through this taxonomy, we synthesize key findings on research distribution, data preparation, and model selection. Finally, we discuss several directions to advance the effective integration of ML in HPC I/O systems. Jingxian Peng, Huijun Wu 0001, Zhenwei Wu, Wei Zhang 0027, Yiqin Dai, Yong Dong |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2022 | Towards Scalable Resource Management for SupercomputersabstractToday's supercomputers offer massive computation resources to execute a large number of user jobs. Effectively managing such large-scale hardware parallelism and workloads is essential for supercomputers. However, existing HPC resource management (RM) systems fail to capitalize on the hardware parallelism by following a centralized design used decades ago. They give poor scalability and inefficient performance on today's supercomputers, which will worsen in exascale computing. We present ESlurm, a better RM for supercomputers. As a departure from existing HPC RMs, ESlurm implements a distributed communication structure. It employs a new communication tree strategy and uses job runtime estimation to improve communications and job scheduling efficiency. ESlurm is deployed into production in a real supercomputer. We evaluate ESlurm on up to 20K nodes. Compared to state-of-the-art RM solutions, ESlurm exhibits better scalability, significantly reducing the resource usage of master nodes and improving data transfer and job scheduling efficiency by a large margin. Yiqin Dai, Yong Dong, Kai Lu 0001, Ruibo Wang, Wei Zhang 0027, Juan Chen 0001, Mingtian Shao, Zheng Wang 0001 |
SC | 5 |
| 2020 | Improving Neural Relation Extraction with Positive and Unlabeled LearningabstractWe present a novel approach to improve the performance of distant supervision relation extraction with Positive and Unlabeled (PU) Learning. This approach first applies reinforcement learning to decide whether a sentence is positive to a given relation, and then positive and unlabeled bags are constructed. In contrast to most previous studies, which mainly use selected positive instances only, we make full use of unlabeled instances and propose two new representations for positive and unlabeled bags. These two representations are then combined in an appropriate way to make bag-level prediction. Experimental results on a widely used real-world dataset demonstrate that this new approach indeed achieves significant and consistent improvements as compared to several competitive baselines. Zhengqiu He, Wenliang Chen, Yuyi Wang 0001, Wei Zhang 0027, Guanchun Wang, Min Zhang 0005 |
AAAI | 4 |
| 2020 | Improving Relation Extraction with Relational Paraphrase SentencesabstractSupervised models for Relation Extraction (RE) typically require human-annotated training data.Due to the limited size, the human-annotated data is usually incapable of covering diverse relation expressions, which could limit the performance of RE.To increase the coverage of relation expressions, we may enlarge the labeled data by hiring annotators or applying Distant Supervision (DS).However, the human-annotated data is costly and non-scalable while the distantly supervised data contains many noises.In this paper, we propose an alternative approach to improve RE systems via enriching diverse expressions by relational paraphrase sentences.Based on an existing labeled data, we first automatically build a task-specific paraphrase data.Then, we propose a novel model to learn the information of diverse relation expressions.In our model, we try to capture this information on the paraphrases via a joint learning framework.Finally, we conduct experiments on a widely used dataset and the experimental results show that our approach is effective to improve the performance on relation extraction, even compared with a strong baseline. Tong Zhu 0002, Wenliang Chen, Wei Zhang 0027, Min Zhang 0005 |
COLING | 4 |
| 2020 | Towards Accurate and Consistent Evaluation: A Dataset for Distantly-Supervised Relation ExtractionabstractIn recent years, distantly-supervised relation extraction has achieved a certain success by using deep neural networks.Distant Supervision (DS) can automatically generate large-scale annotated data by aligning entity pairs from Knowledge Bases (KB) to sentences.However, these DSgenerated datasets inevitably have wrong labels that result in incorrect evaluation scores during testing, which may mislead the researchers.To solve this problem, we build a new dataset NYT-H, where we use the DS-generated data as training data and hire annotators to label test data.Compared with the previous datasets, NYT-H has a much larger test set and then we can perform more accurate and consistent evaluation.Finally, we present the experimental results of several widely used systems on NYT-H.The experimental results show that the ranking lists of the comparison systems on the DS-labelled test data and human-annotated test data are different.This indicates that our human-annotated data is necessary for evaluation of distantly-supervised relation extraction. Tong Zhu 0002, Haitao Wang 0019, Xiabing Zhou, Wenliang Chen, Wei Zhang 0027, Min Zhang 0005 |
COLING | 6 |
| 2019 | Syntax-aware entity representations for neural relation extraction
Zhengqiu He, Wenliang Chen, Zhenghua Li, Wei Zhang 0027, Hao Shao, Min Zhang 0005 |
Artif. Intell. | 4 |
| 2018 | SEE: Syntax-Aware Entity Embedding for Neural Relation ExtractionabstractDistant supervised relation extraction is an efficient approach to scale relation extraction to very large corpora, and has been widely used to find novel relational facts from plain text. Recent studies on neural relation extraction have shown great progress on this task via modeling the sentences in low-dimensional spaces, but seldom considered syntax information to model the entities. In this paper, we propose to learn syntax-aware entity embedding for neural relation extraction. First, we encode the context of entities on a dependency tree as sentence-level entity embedding based on tree-GRU. Then, we utilize both intra-sentence and inter-sentence attentions to obtain sentence set-level entity embedding over all sentences containing the focus entity pair. Finally, we combine both sentence embedding and entity embedding for relation classification. We conduct experiments on a widely used real-world dataset and the experimental results show that our model can make full use of all informative instances and achieve state-of-the-art performance of relation extraction. Zhengqiu He, Wenliang Chen, Zhenghua Li, Meishan Zhang, Wei Zhang 0027, Min Zhang 0005 |
AAAI | 5 |
| 2018 | Adversarial Learning for Chinese NER From Crowd AnnotationsabstractTo quickly obtain new labeled data, we can choose crowdsourcing as an alternative way at lower cost in a short time. But as an exchange, crowd annotations from non-experts may be of lower quality than those from experts. In this paper, we propose an approach to performing crowd annotation learning for Chinese Named Entity Recognition (NER) to make full use of the noisy sequence labels from multiple annotators. Inspired by adversarial learning, our approach uses a common Bi-LSTM and a private Bi-LSTM for representing annotator-generic and -specific information. The annotator-generic information is the common knowledge for entities easily mastered by the crowd. Finally, we build our Chinese NE tagger based on the LSTM-CRF model. In our experiments, we create two data sets for Chinese NER tasks from two domains. The experimental results show that our system achieves better scores than strong baseline systems. YaoSheng Yang, Meishan Zhang, Wenliang Chen, Wei Zhang 0027, Haofen Wang, Min Zhang 0005 |
AAAI | 4 |
| 2018 | Constructing a database for the relations between CNV and human genetic diseases via systematic text miningabstractBACKGROUND: The detection and interpretation of CNVs are of clinical importance in genetic testing. Several databases and web services are already being used by clinical geneticists to interpret the medical relevance of identified CNVs in patients. However, geneticists or physicians would like to obtain the original literature context for more detailed information, especially for rare CNVs that were not included in databases. RESULTS: The resulting CNVdigest database includes 440,485 sentences for CNV-disease relationship. A total number of 1582 CNVs and 2425 diseases are involved. Sentences describing CNV-disease correlations are indexed in CNVdigest, with CNV mentions and disease mentions annotated. CONCLUSIONS: In this paper, we use a systematic text mining method to construct a database for the relationship between CNVs and diseases. Based on that, we also developed a concise front-end to facilitate the analysis of CNV/disease association, providing a user-friendly web interface for convenient queries. The resulting system is publically available at http://cnv.gtxlab.com /. Xi Yang 0020, Chengkun Wu, Wei Wang 0130, Gen Li 0007, Wei Zhang 0027, Lingqian Wu, Kai Lu 0001 |
BMC Bioinform. | 6 |
| 2010 | An affine invariant interest point and region detector based on Gabor filtersabstractThis paper presents a novel approach for interest point and region detection which is invariant to affine transformations. Such transformations introduce significant changes in the point location as well as in the scale and the shape of the neighborhood of an interest point. Our approach allows to solve for these problems simultaneously. The approach is based on three key ideas: 1) Interest points can be extracted based on local maxima of the normalized local energy maps. 2) Local extrema over scale of the normalized energy function indicate the presence of characteristic local structures. 3) The maximum response along all the orientations indicates the principle orientation of the local structure. We first extract interest points at multi-scales from the local energy map constructed by Gabor filter responses, and then select points at which a local measure is maximal over scales. This allows a selection of distinctive points for which the characteristic scale is known. We then estimate the principle orientation through the orientational responses of Gabor filters and extend the detector to affine invariance by estimating the affine shape of a point neighborhood. The characteristic scale and the affine shape of neighborhood determine an affine invariant region for each point. Experimental results with synthetic images and natural images show the affine invariance performance of our approach. Comparative evaluation using the repeatability criteria demonstrates the comparable performance in the presence of large viewpoint changes. Wanying Xu, Xinsheng Huang, Xingwei Li, Ying Zhang 0032, Wei Zhang 0027 |
ICARCV | 6 |
| 2009 | Architecture- and OS-Independent Binary-Level Dynamic Test Generation
Gen Li 0002, Kai Lu 0001, Ying Zhang 0032, Xicheng Lu, Wei Zhang 0027 |
ICICS | 5 |
| 2009 | Providing Responsiveness Requirement Based Consistency in DVEabstractConsistency and responsiveness are two important factors in providing the sense of reality in Distributed Virtual Environment (DVE). However, it is not easy to optimize both aspects because of the trade-off between these two factors. As a result, most existing consistency maintenance methods ignored the responsiveness requirements, or just assumed a simple responsiveness requirement model which cannot meet the real need of DVE systems. In this paper, we first present a new responsiveness requirement model. The model can describe requirement satisfaction situation of each node. Base on this model, we propose a responsiveness requirement based consistency method. The method can adjust the utilization of time resource according to the requirements of different nodes and improve the overall responsiveness performance by at least 20%. Therefore, it provides a good support to increase the applicability of DVE systems. Wei Zhang 0027, Hangjun Zhou, Yuxing Peng 0001, Sikun Li |
ICPADS | 1 |