EDBT 2026 Demo / reviewers in the wild / expert
Zhenwei Wu
dblp:82/5441
· DBLP profile ↗
17ranked-venue papers
4as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Computer networks · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DeloopSGNN: Revisiting Spectral GNNs Through the Lens of Spatial AggregationabstractGraph Neural Networks (GNNs) have been studied from two primary perspectives: spectral, which employs global graph signal filtering and is theoretically more expressive, and spatial, which builds on local neighborhood aggregation and generalizes well across diverse graph structures. While spectral GNNs are expected to perform better in theory, they often underperform in practice compared to spatial models. To better understand this gap, we introduce a novel theoretical framework for converting spectral GNNs into the spatial domain, allowing for more intuitive analysis. This transformation reveals that signal looping and repeated high-order aggregation are major causes of over-smoothing in spectral GNNs. By addressing these issues in the spatial domain and converting the model back to the spectral domain, we propose DeloopSGNN, a spectral GNN with improved expressive capacity. Experiments on benchmark datasets show that DeloopSGNN achieves consistently strong performance in terms of accuracy and adversarial robustness, demonstrating that spectral GNNs can benefit significantly from careful architectural design grounded in our proposed framework. Duanyu Li, Huijun Wu 0001, Kai Lu 0001, Zhenwei Wu, Yong Dong, Ruibo Wang |
AAAI | 6 |
| 2026 | PallasGNN: Curriculum-Based Pattern Mining for Robust GNNs
Kaiwen Xia, Huijun Wu 0001, Ruibo Wang, Zhenwei Wu, Yong Dong |
PAKDD (3) | 6 |
| 2026 | A survey of anomaly detection in HPC systems using machine learningabstractAbstract High-performance computing (HPC) systems must remain stable and reliable to consistently deliver robust computational power and ensure the proper execution of user jobs. Anomaly detection is a key means to ensure the stability and reliability of these systems. With the expansion of HPC systems and changes in their architecture, accurately identifying anomalies in dynamic environments has become increasingly challenging. Traditional detection methods rely on experience and rules, which could be inefficient and inaccurate. To address these issues, researchers have proposed machine learning-based methods to automatically process large amounts of complex data, improving the efficiency of anomaly identification and diagnosis. In this survey, we conduct a comprehensive and in-depth investigation of machine learning-based anomaly detection methods in HPC systems. Firstly, we summarize and introduce the background and challenges of anomaly detection in HPC systems. Secondly, we compare a series of machine learning-based anomaly detection works in detail and summarize their frameworks. We conclude their advantages and disadvantages and application scenarios. Finally, we discuss several promising development trends of machine learning-based HPC system anomaly detection. Wei Zhang 0027, Yiqin Dai, Huijun Wu 0001, Zhenwei Wu, Hongyun Tian, Juan Chen 0001, Chubo Liu, Yong Dong |
CCF Trans. High Perform. Comput. | 6 |
| 2026 | Scribble-Guided Hierarchical Prompt for SAM-Based Weakly Supervised Salient Object Detection
Fen Xiao, Ruozhuo Huang, Zhenwei Wu, Xieping Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | A Survey on Machine Learning-Based HPC I/O Analysis and OptimizationabstractThe soaring computing power of HPC systems supports numerous large-scale applications, which generate massive data volumes and diverse I/O patterns, leading to severe I/O bottlenecks. Analyzing and optimizing HPC I/O is therefore critical. However, traditional approaches are typically customized and lack the adaptability required to cope with dynamic changes in HPC environments. To address the challenge, Machine Learning (ML) has been increasingly adopted to automate and enhance I/O analysis and optimization. Given sufficient I/O traces from HPC systems, ML can learn underlying I/O behaviors, extract actionable insights, and dynamically adapt to evolving workloads to improve performance. In this survey, we propose a novel taxonomy that aligns HPC I/O problems with learning tasks to systematically review existing studies. Through this taxonomy, we synthesize key findings on research distribution, data preparation, and model selection. Finally, we discuss several directions to advance the effective integration of ML in HPC I/O systems. Jingxian Peng, Huijun Wu 0001, Zhenwei Wu, Wei Zhang 0027, Yiqin Dai, Yong Dong |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2025 | MergeFS: Optimizing Node-Local Burst Buffers for Complex HPC WorkflowsabstractHigh-performance computing (HPC) applications are increasingly transitioning from traditional numerical simulations to an intelligent fusion paradigm integrating AI algorithms and big data analytics, exemplified by initiatives such as AI4Science. This evolution, coupled with rising problem complexity, results in workflows composed of interdependent subtasks. Existing HPC storage solutions, particularly burst buffer systems, have yet to adequately address the unique challenges posed by such workflows, including efficient cross-task data sharing and namespace fusion, leading to suboptimal resource utilization and performance bottlenecks in complex dependency scenarios. In this paper, we present MergeFS, a lightweight, workflow-aware burst buffer file system that incorporates a treestructured workflow registry for precise dependency management alongside a dynamic multi-namespace mechanism enabling rapid and isolated data access. MergeFS effectively integrates workflow management, namespace control, and data view fusion. Experimental evaluations demonstrate that MergeFS significantly outperforms current workflow-centric burst buffer optimizations in runtime performance with low management overhead. Zhaohao Zhong, Huijun Wu 0001, Yong Dong, Zhenwei Wu, Ruibo Wang |
ICPADS | 6 |
| 2024 | Talos: A More Effective and Efficient Adversarial Defense for GNN Models Based on the Global Homophily of GraphsabstractGraph neural network (GNN) models play a pivotal role in numerous tasks involving graph-related data analysis. Despite their efficacy, similar to other deep learning models, GNNs are susceptible to adversarial attacks. Even minor perturbations in graph data can induce substantial alterations in model predictions. While existing research has explored various adversarial defense techniques for GNNs, the challenge of defending against adversarial attacks on real-world scale graph data remains largely unresolved. On one hand, methods reliant on graph purification and preprocessing tend to excessively emphasize local graph information, leading to sub-optimal defensive outcomes. On the other hand, approaches rooted in graph structure learning entail significant time overheads, rendering them impractical for large-scale graphs. In this paper, we propose a new defense method named Talos, which enhances the global, rather than local, homophily of graphs as a defense. Experiments show that the proposed approach notably outperforms state-of-the-art defense approaches, while imposing little computational overhead. Duanyu Li, Huijun Wu 0001, Xugang Wu, Zhenwei Wu |
ECAI | 5 |
| 2024 | Federated Multi-View Clustering via Tensor Factorization
Wei Feng 0010, Zhenwei Wu, Qianqian Wang 0001, Bo Dong 0001, Zhiqiang Tao, Quanxue Gao |
IJCAI | 2 |
| 2024 | Efficient Federated Multi-View Clustering with Integrated Matrix Factorization and K-Means
Wei Feng 0010, Zhenwei Wu, Qianqian Wang 0001, Bo Dong 0001, Zhiqiang Tao, Quanxue Gao |
IJCAI | 2 |
| 2024 | Federated Fuzzy C-means with Schatten-p Norm MinimizationabstractMulti-view clustering has emerged as an important unsupervised method to process unlabelled multi-view data that provides a comprehensive description of an object. Existing multi-view clustering methods focus on centralized settings but ignore the fact that real-world multi-view data may be distributed across different entities. The sensitive information embedded in multi-view data hinders the cooperative training of multi-view clustering, since data of different views cannot be directly shared, leading to a great challenge to cooperatively exploit the consistent and complementary information of different views. To validate the multi-view clustering in distributed scenarios, in this paper, we propose a novel federated multi-view method named Federated Multi-View Fuzzy C-means with Schatten-p Norm Minimization (FMVFCMSP) which is based on fuzzy C-means and tensor Schatten p-norm. Specifically, we utilize the membership degrees to replace conventional hard clustering assignment in K-means, enabling improved uncertainty handling and less information loss. Moreover, we introduce a tensor Schatten p-norm-based regularizer to fully explore the inter-view complementary information and global spatial structure. We also develop a federated optimization algorithm enabling clients to collaboratively learn the clustering results. Extensive experiments on several datasets demonstrate that our proposed method exhibits superior performance in federated multi-view clustering. Wei Feng 0010, Zhenwei Wu, Qianqian Wang 0001, Bo Dong 0001, Quanxue Gao |
ACM Multimedia | 2 |
| 2022 | PMemTrace: Lightweight and Efficient Memory Access Monitoring for Persistent Memory
Yushuqing Zhang, Kai Lu 0001, Zhenwei Wu |
ICA3PP | 3 |
| 2022 | Efficient Dynamic Binary Translation with Accumulative Persistent Code CachingabstractDynamic binary translation technology can directly translate the binary code of one instruction architecture into the one of another instruction architecture at runtime, and execute it on the target machine. This technology has been applied in many fields, and several dynamic binary translators developed on this basis have shown great translation efficiency. While efficiently translating instruction streams is required, the challenge of mitigating translation overhead is also the essential issue. At present, there are many optimization techniques that can effectively improve the performance of dynamic binary translator, however, most of the techniques are ineffective for the applications with short run times or massive cold code regions. This paper introduces a method of persistent code caching that can be shared and reused between different executions. This method enables the same program to reuse the translated code extracted from former executions to skip part of the translation work, and allows the persistent code accumulated to raise reuse hit ratios. We have implemented a prototype based on Box64, a open source and newly released efficient dynamic binary translator, with persistent code caching method. The experimental results demonstrate that the optimization method could efficiently improve the performance of Box64. Haoming Lin, Yong Dong, Wanqing Chi, Zhenwei Wu, Hongqing Zeng |
ICPADS | 4 |
| 2020 | PMThreads: persistent memory threads harnessing versioned shadow copiesabstractByte-addressable non-volatile memory (NVM) makes it possible to perform fast in-memory accesses to persistent data using standard load/store processor instructions. Some approaches for NVM are based on durable memory transactions and provide a persistent programming paradigm. However, they cannot be applied to existing multi-threaded applications without extensive source code modifications. Durable transactions typically rely on logging to enforce failure-atomic commits that include additional writes to NVM and considerable ordering overheads. Zhenwei Wu, Kai Lu 0001, Andy Nisbet, Mikel Luján |
PLDI | 1 |
| 2020 | A survey on optimizations towards best-effort hardware transactional memory
Zhenwei Wu, Kai Lu 0001, Ruibo Wang |
CCF Trans. High Perform. Comput. | 1 |
| 2019 | POSTER: Quiescent and Versioned Shadow Copies for NVMabstractQuiescentNVM is a user-space runtime providing transparent failure-consistency guarantees for lock-based parallel programs executing on hybrid combinations of traditional DRAM and byte-addressable non-volatile memory (NVM) technologies. A dual-versioning mechanism performs in-place persistent writes over one copy, and consistent fallback guarantees are provided by the other copy. Thus, the two writes to NVM present in logging-based solutions (such as for durable memory transactions) are reduced to a single write. Further, we avoid the need to rewrite legacy applications to exploit durable transactions. Our system relies on its dual-copy framework operation that safely persists data during global quiescent states, where no thread must hold a lock on persistent data. For applications with low lock-contention, global lock-free quiescent states will occur sufficiently frequently, and we deliver better performance and lower wear to NVM than current systems. We do not cover high-lock contention scenarios while enforcing quiescent states to occur. Zhenwei Wu, Kai Lu 0001, Andy Nisbet, Mikel Luján |
PACT | 1 |
| 2016 | Reducing and Balancing Flow Table Entries in Software-Defined NetworksabstractSoftware-Defined Networking (SDN) allows flexible and efficient management of networks. However, the limited capacity of flow tables in SDN switches hinders the deployment of SDN. In this paper, we propose a novel routing scheme to improve the efficiency of flow tables in SDNs. To efficiently use the routing scheme, we formulate an optimization problem with the objective to maximize the number of flows in the network, constrained by the limited flow table space in SDN switches. The problem is NP-hard, and we propose the K Similar Greedy Tree (KSGT) algorithm to solve it. We evaluate the performance of KSGT against "traditional" SDN solutions with real-world topologies and traffic. The results show that, compared to the existing solutions, KSGT can reduce about 60% of flow entries when processing the same amount of flows, and improve about 25% of the successful installation and forwarding flows under the same flow table space. Xuya Jia, Yong Jiang 0001, Zehua Guo 0001, Zhenwei Wu |
LCN | 4 |
| 2016 | An Efficiency Pipeline Processing Approach for OpenFlow SwitchabstractOne primary component of OpenFlow switches is a pipeline of flow tables. Flows are directed through the pipeline by looking up a matched rule in each table. Sequentially, Packet matching starts at table 0 and continues to additional tables of the pipeline if necessary. The lookup stops when one flow entry is matched or the end of pipeline is reached. Though effective, this manner of processing is not efficient when popular rules are installed in back tables. If frequently matched rules appear earlier in the pipeline, the procedure of lookup and comparison can be improved. Unfortunately, a simple sorting algorithm is not feasible. In this paper, we formalize the problem of reducing lookup times, which is proven to be NP-hard. A heuristic approach, MILE(migrating flow rules), is proposed to minimize the average number of lookups. Experimental results show that MILE is able to reduce table lookups by 50%. Zhenwei Wu |
LCN | 1 |