EDBT 2026 Demo / reviewers in the wild / expert
Huijun Wu 0001
dblp:95/1731-1
· DBLP profile ↗
25ranked-venue papers
5as first author
17since 2021 · last 2026
0000-0002-9513-5359ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DeloopSGNN: Revisiting Spectral GNNs Through the Lens of Spatial AggregationabstractGraph Neural Networks (GNNs) have been studied from two primary perspectives: spectral, which employs global graph signal filtering and is theoretically more expressive, and spatial, which builds on local neighborhood aggregation and generalizes well across diverse graph structures. While spectral GNNs are expected to perform better in theory, they often underperform in practice compared to spatial models. To better understand this gap, we introduce a novel theoretical framework for converting spectral GNNs into the spatial domain, allowing for more intuitive analysis. This transformation reveals that signal looping and repeated high-order aggregation are major causes of over-smoothing in spectral GNNs. By addressing these issues in the spatial domain and converting the model back to the spectral domain, we propose DeloopSGNN, a spectral GNN with improved expressive capacity. Experiments on benchmark datasets show that DeloopSGNN achieves consistently strong performance in terms of accuracy and adversarial robustness, demonstrating that spectral GNNs can benefit significantly from careful architectural design grounded in our proposed framework. Duanyu Li, Huijun Wu 0001, Kai Lu 0001, Zhenwei Wu, Yong Dong, Ruibo Wang |
AAAI | 2 |
| 2026 | PallasGNN: Curriculum-Based Pattern Mining for Robust GNNs
Kaiwen Xia, Huijun Wu 0001, Ruibo Wang, Zhenwei Wu, Yong Dong |
PAKDD (3) | 2 |
| 2026 | A survey of anomaly detection in HPC systems using machine learningabstractAbstract High-performance computing (HPC) systems must remain stable and reliable to consistently deliver robust computational power and ensure the proper execution of user jobs. Anomaly detection is a key means to ensure the stability and reliability of these systems. With the expansion of HPC systems and changes in their architecture, accurately identifying anomalies in dynamic environments has become increasingly challenging. Traditional detection methods rely on experience and rules, which could be inefficient and inaccurate. To address these issues, researchers have proposed machine learning-based methods to automatically process large amounts of complex data, improving the efficiency of anomaly identification and diagnosis. In this survey, we conduct a comprehensive and in-depth investigation of machine learning-based anomaly detection methods in HPC systems. Firstly, we summarize and introduce the background and challenges of anomaly detection in HPC systems. Secondly, we compare a series of machine learning-based anomaly detection works in detail and summarize their frameworks. We conclude their advantages and disadvantages and application scenarios. Finally, we discuss several promising development trends of machine learning-based HPC system anomaly detection. Wei Zhang 0027, Yiqin Dai, Huijun Wu 0001, Zhenwei Wu, Hongyun Tian, Juan Chen 0001, Chubo Liu, Yong Dong |
CCF Trans. High Perform. Comput. | 5 |
| 2026 | A Survey on Machine Learning-Based HPC I/O Analysis and OptimizationabstractThe soaring computing power of HPC systems supports numerous large-scale applications, which generate massive data volumes and diverse I/O patterns, leading to severe I/O bottlenecks. Analyzing and optimizing HPC I/O is therefore critical. However, traditional approaches are typically customized and lack the adaptability required to cope with dynamic changes in HPC environments. To address the challenge, Machine Learning (ML) has been increasingly adopted to automate and enhance I/O analysis and optimization. Given sufficient I/O traces from HPC systems, ML can learn underlying I/O behaviors, extract actionable insights, and dynamically adapt to evolving workloads to improve performance. In this survey, we propose a novel taxonomy that aligns HPC I/O problems with learning tasks to systematically review existing studies. Through this taxonomy, we synthesize key findings on research distribution, data preparation, and model selection. Finally, we discuss several directions to advance the effective integration of ML in HPC I/O systems. Jingxian Peng, Huijun Wu 0001, Zhenwei Wu, Wei Zhang 0027, Yiqin Dai, Yong Dong |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2026 | Fully Decentralized Data Distribution for Large-Scale HPC SystemsabstractFor many years, in the HPC data distribution scenario, as the scale of the HPC system continues to increase, manufacturers have to increase the number of data providers to improve the IO parallelism to match the data demanders. In large-scale, especially exascale HPC systems, this mode of decoupling the demander and provider presents significant scalability limitations and incurs substantial costs. In our view, only a distribution model in which the demander also acts as the provider can fundamentally cope with changes in scale and have the best scalability, which is called all-to-all data distribution mode in this paper. We design and implement the BitTorrent protocol on computing networks in HPC systems and propose FD3, a fully decentralized data distribution method. We design the Requested-to-Validated Table (RVT) and the Highest ranking and Longest consecutive piece segment First (HLF) policy based on the features of the HPC networking environment to improve the performance of FD3. In addition, we design a torrent-tree to accelerate the distribution of seed file data and the aggregation of distribution state, and release the tracker load with neighborhood local-generation algorithm. Experimental results show that FD3 can scale smoothly to 11k+ computing nodes, and its performance is much better than that of the parallel file system. Compared with the original BitTorrent, the performance is improved by 8-15 times. FD3 highlights the considerable potential of the all-to-all model in HPC data distribution scenarios. Furthermore, the work of this paper can further stimulate the exploration of future distributed parallel file systems and provide a foundation and inspiration for the design of data access patterns for Exscale HPC systems. Ruibo Wang, Mingtian Shao, Huijun Wu 0001, Yiqin Dai, Kai Lu 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2025 | MergeFS: Optimizing Node-Local Burst Buffers for Complex HPC WorkflowsabstractHigh-performance computing (HPC) applications are increasingly transitioning from traditional numerical simulations to an intelligent fusion paradigm integrating AI algorithms and big data analytics, exemplified by initiatives such as AI4Science. This evolution, coupled with rising problem complexity, results in workflows composed of interdependent subtasks. Existing HPC storage solutions, particularly burst buffer systems, have yet to adequately address the unique challenges posed by such workflows, including efficient cross-task data sharing and namespace fusion, leading to suboptimal resource utilization and performance bottlenecks in complex dependency scenarios. In this paper, we present MergeFS, a lightweight, workflow-aware burst buffer file system that incorporates a treestructured workflow registry for precise dependency management alongside a dynamic multi-namespace mechanism enabling rapid and isolated data access. MergeFS effectively integrates workflow management, namespace control, and data view fusion. Experimental evaluations demonstrate that MergeFS significantly outperforms current workflow-centric burst buffer optimizations in runtime performance with low management overhead. Zhaohao Zhong, Huijun Wu 0001, Yong Dong, Zhenwei Wu, Ruibo Wang |
ICPADS | 2 |
| 2025 | From Islands to Archipelago: Towards Collaborative and Adaptive Burst Buffer for HPC SystemsabstractModern supercomputers increasingly use node-local storage as burst buffers (BB) to address I/O bottlenecks.However, current BBs do not naturally support workflows, a common workload in HPC consisting of many interconnected subtasks.Workflow I/O can be divided into three types: intratask I/O, inter-task I/O, and stage-in/out I/O.While BBs can accelerate intra-task I/O, they often overlook the other two.Inter-task I/O relies on migrating data through the Parallel File System (PFS), which can slow down overall performance.Although allocating more resources to create larger BBs could help, it increases costs.Additionally, temporary BBs lack permanent storage, requiring data migration between the PFS and BB for stage-in and stage-out I/O.This process often involves multiple data copies and reduces I/O efficiency.Even for intra-task I/O, unbalanced data distribution can cause bottlenecks on heavily loaded nodes.To improve workflow acceleration in BB systems, it is important to address the needs of all the above-mentioned three Mingtian Shao, Ruibo Wang, Kai Lu 0001, Yiqin Dai, Huijun Wu 0001 |
ICS | 6 |
| 2025 | Incomplete Multi-view Deep Clustering with Data Imputation and AlignmentabstractIncomplete multi-view deep clustering is an emerging research hot-pot to incorporate data information of multiple sources or modalities when parts of them are missing. Most of existing approaches encode the available data observations into multiple view-specific latent representations and subsequently integrate them for the next clustering task. However, they ignore that the latent representations are unique to a fixed set of data samples in all views. Meanwhile, the pair-wise similarities of missing data observations are also failed to utilize in latent representation learning sufficiently, leading to unsatisfactory clustering performance. To address these issues, we propose an incomplete multi-view deep clustering method with data imputation and alignment. Assuming that each data sample corresponds to a same latent representation among all views, it projects the latent representations into feature spaces with neural networks. As a result, not only the available data observations are reconstructed, but also the missing ones can be imputed accordingly. Moreover, a linear alignment measurement of linear complexity is defined to compute the pair-wise similarities of all data observations, especially including those of the missing. By executing the above two procedures iteratively, the discriminative latent representations can be learned and used to group the data into categories with off-the-shelf clustering algorithms. In experiment, the proposed method is validated on a set of benchmark datasets and achieves state-of-the-art performances. Jiyuan Liu 0003, Xinwang Liu 0002, Xinhang Wan, Ke Liang 0006, Weixuan Liang, Sihang Zhou 0001, Huijun Wu 0001, Kehua Guo |
NeurIPS | 7 |
| 2025 | MIST: Towards MPI Instant Startup and Termination on Tianhe HPC SystemsabstractAs the size of MPI programs grows with expanding HPC resources and parallelism demands, the overhead of MPI startup and termination escalates due to the inclusion of less scalable global operations. Global operations involving extensive cross-machine communication and synchronization are crucial for ensuring semantic correctness. The current focus is on optimizing and accelerating these global operations rather than removing them, as the latter involves systematic changes to the system software stack and may impact program semantics. Given this background, we propose a systematic solution named MIST to safely eliminate global operations in MPI startup and termination. Through optimizing the generation of communication addresses, designing reliable communication protocols, and exploiting the resource release mechanism, MIST eliminates all global operations to achieve MPI instant startup and termination while ensuring correct program execution. Experiments on Tianhe-2A supercomputer demonstrate that MIST can reduce theMPI_Init()time by 32.5-77.6% and theMPI_Finalize()time by 28.9-85.0%. Yiqin Dai, Ruibo Wang, Yong Dong, Juan Chen 0001, Huijun Wu 0001, Mingtian Shao, Kai Lu 0001 |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2024 | Fully Decentralized Data Distribution for Exascale-HPC: End of the Provider-Demander Matching PuzzleabstractFor many years, in the HPC data distribution scenario, as the scale of the HPC system continues to increase, manufacturers have to increase the number of data providers to improve the IO parallelism to match the data demanders. In the era of Exascale Computing, this mode of decoupling the demander and provider has limited scalability and huge costs. In our view, only a distribution model in which the demander also acts as the provider can fundamentally cope with changes in scale and have the best scalability, which is called all-to-all data distribution mode in this paper. We design and implement the BitTorrent protocol on computing networks in HPC systems and propose FD3, a fully decentralized data distribution method. We design the Requested-to-Validated Table (RVT) and the Nearest and Longest consecutive piece Segment First (NLSF) policy based on the features of the HPC networking environment to improve the performance of FD3. Experimental results show that FD3 can scale smoothly to 11k+ computing nodes, and its performance is much better than that of the parallel file system. Compared with the original BitTorrent, the performance is improved by 7–11 times. FD3 shows the great potential of the all-to-all model in HPC data distribution scenarios. At the same time, the work of this paper can further stimulate the exploration of future distributed parallel file systems and provide a foundation and inspiration for the design of data access patterns for Exscale HPC systems. Mingtian Shao, Ruibo Wang, Huijun Wu 0001, Yiqin Dai, Kai Lu 0001 |
CLUSTER | 4 |
| 2024 | Talos: A More Effective and Efficient Adversarial Defense for GNN Models Based on the Global Homophily of GraphsabstractGraph neural network (GNN) models play a pivotal role in numerous tasks involving graph-related data analysis. Despite their efficacy, similar to other deep learning models, GNNs are susceptible to adversarial attacks. Even minor perturbations in graph data can induce substantial alterations in model predictions. While existing research has explored various adversarial defense techniques for GNNs, the challenge of defending against adversarial attacks on real-world scale graph data remains largely unresolved. On one hand, methods reliant on graph purification and preprocessing tend to excessively emphasize local graph information, leading to sub-optimal defensive outcomes. On the other hand, approaches rooted in graph structure learning entail significant time overheads, rendering them impractical for large-scale graphs. In this paper, we propose a new defense method named Talos, which enhances the global, rather than local, homophily of graphs as a defense. Experiments show that the proposed approach notably outperforms state-of-the-art defense approaches, while imposing little computational overhead. Duanyu Li, Huijun Wu 0001, Xugang Wu, Zhenwei Wu |
ECAI | 2 |
| 2024 | AIO: Automating I/O Optimization Pipeline for Data-Intensive Applications in HPCabstractHigh-performance computing (HPC) systems has entered the exascale era, but I/O performance has lagged behind due to storage hardware limitations, creating a "storage wall effect" that hinders HPC systems full potential. Modern HPC storage systems employ a layered storage architecture with local node storage, burst buffers, and global parallel file systems to improve I/O performance. However, application workloads vary, and fast storage layers may not always outperform global parallel file systems, while a static storage software stack configuration may not achieve optimal performance across different applications. Traditional I/O optimization is time-consuming, laborintensive, and prone to human error. To address this issue, this paper proposes AIO, a data-driven I/O optimization method. AIO automatically selects the optimal storage hierarchy and software stack based on the task's I/O patterns, and further enhances performance by identifying and applying the most efficient configuration settings. We evaluated AIO with eight applications in a cluster that features shared burst buffers, local burst buffers, and a global shared file system. The experiments demonstrate that AIO effectively optimizes I/O performance, achieving an overall speedup of 2.74. Wanxin Wang, Huijun Wu 0001, Zhangyu Liu |
ISPA | 2 |
| 2024 | Automatic and Aligned Anchor Learning Strategy for Multi-View ClusteringabstractMulti-View Clustering (MVC) commonly utilizes the anchor technique to mitigate the computational complexity. Existing methods generally assume a pre-selection of anchors to facilitate subsequent clustering tasks. However, the determination of the optimal number of anchors is often non-trivial and necessitates their treatment as a tunable parameter, incurring additional computational overhead. Moreover, it is not reasonable to assume an identical number of anchors across all views, as this assumption restricts the representational capacity of anchors in individual views. To address the above issues, we propose a view adaptive anchor multi-view clustering called Multi-view Clustering with Automatic and Aligned Anchor (3AMVC). We introduce a Hierarchical Bipartite Neighbor Clustering (HBNC) strategy to adaptively select a suitable number of representative anchors in each view. Specifically, when the representative difference of anchors lies in a acceptable and satisfactory range, the HBNC process is halted and picks out the final anchors. Moreover, we propose an innovative anchor alignment strategy in response to the varying quantities of anchors across different views. This approach initially evaluates the quality of anchors on each view based on the intra-cluster distance criterion and then proceeds to align based on the view with the highest-quality anchors. The carefully organized experiments well validate the effectiveness and strengthens of 3AMVC. Siwei Wang 0001, Shengju Yu, Suyuan Liu, Junjie Huang 0001, Huijun Wu 0001, Xinwang Liu 0002, En Zhu |
ACM Multimedia | 6 |
| 2024 | GraphLearner: Graph Node Clustering with Fully Learnable AugmentationabstractContrastive deep graph clustering (CDGC) leverages the power of contrastive learning to group nodes into different clusters. The quality of contrastive samples is crucial for achieving better performance, making augmentation techniques a key factor in the process. However, the augmentation samples in existing methods are always predefined by human experiences, and agnostic from the downstream task clustering, thus leading to high human resource costs and poor performance. To overcome these limitations, we propose a Graph Node Clustering with Fully Learnable Augmentation, termed GraphLearner. It introduces learnable augmentors to generate high-quality and task-specific augmented samples for CDGC. GraphLearner incorporates two learnable augmentors specifically designed for capturing attribute and structural information. Moreover, we introduce two refinement matrices, including the high-confidence pseudo-label matrix and the cross-view sample similarity matrix, to enhance the reliability of the learned affinity matrix. During the training procedure, we notice the distinct optimization goals for training learnable augmentors and contrastive learning networks. In other words, we should both guarantee the consistency of the embeddings as well as the diversity of the augmented samples. To address this challenge, we propose an adversarial learning mechanism within our method. Besides, we leverage a two-stage training strategy to refine the high-confidence matrices. Extensive experimental results on six benchmark datasets validate the effectiveness of GraphLearner.The code and appendix of GraphLearner are available at https://github.com/xihongyang1999/GraphLearner on Github. Xihong Yang, Erxue Min, Ke Liang 0006, Yue Liu 0008, Siwei Wang 0001, Sihang Zhou 0001, Huijun Wu 0001, Xinwang Liu 0002, En Zhu |
ACM Multimedia | 7 |
| 2024 | Towards adaptive graph neural networks via solving prior-data conflictsabstractGraph neural networks (GNNs) have achieved remarkable performance in a variety of graph-related tasks. Recent evidence in the GNN community shows that such good performance can be attributed to the homophily prior; i.e., connected nodes tend to have similar features and labels. However, in heterophilic settings where the features of connected nodes may vary significantly, GNN models exhibit notable performance deterioration. In this work, we formulate this problem as prior-data conflict and propose a model called the mixture-prior graph neural network (MPGNN). First, to address the mismatch of homophily prior on heterophilic graphs, we introduce the non-informative prior, which makes no assumptions about the relationship between connected nodes and learns such relationship from the data. Second, to avoid performance degradation on homophilic graphs, we implement a soft switch to balance the effects of homophily prior and non-informative prior by learnable weights. We evaluate the performance of MPGNN on both synthetic and real-world graphs. Results show that MPGNN can effectively capture the relationship between connected nodes, while the soft switch helps select a suitable prior according to the graph characteristics. With these two designs, MPGNN outperforms state-of-the-art methods on heterophilic graphs without sacrificing performance on homophilic graphs. Xugang Wu, Huijun Wu 0001, Ruibo Wang, Xu Zhou 0004, Kai Lu 0001 |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2023 | Leveraging Free Labels to Power up Heterophilic Graph Learning in Weakly-Supervised Settings: An Empirical Study
Xugang Wu, Huijun Wu 0001, Ruibo Wang, Duanyu Li, Xu Zhou 0004, Kai Lu 0001 |
ECML/PKDD (3) | 2 |
| 2022 | Towards Defense Against Adversarial Attacks on Graph Neural Networks via Calibrated Co-Training
Xugang Wu, Huijun Wu 0001, Xu Zhou 0004, Kai Lu 0001 |
J. Comput. Sci. Technol. | 2 |
| 2020 | SMINT: Toward Interpretable and Robust Model Sharing for Deep Neural NetworksabstractSharing a pre-trained machine learning model, particularly a deep neural network via prediction APIs, is becoming a common practice on machine learning as a service (MLaaS) platforms nowadays. Although deep neural networks (DNN) have shown remarkable successes in many tasks, they are also criticized for the lack of interpretability and transparency. Interpreting a shared DNN model faces two additional challenges compared with interpreting a general model. (1) Limited training data can be disclosed to users. (2) The internal structure of the models may not be available. These two challenges impede the application of most existing interpretability approaches, such as saliency maps or influence functions, for DNN models. Case-based reasoning methods have been used for interpreting decisions; however, how to select and organize the data points under the constraints of shared DNN models is not discussed. Moreover, simply providing cases as explanations may not be sufficient for supporting instance level interpretability. Meanwhile, existing interpretation methods for DNN models generally lack the means to evaluate the reliability of the interpretation. In this article, we propose a framework named Shared Model INTerpreter (SMINT) to address the above limitations. We propose a new data structure called a boundary graph to organize training points to mimic the predictions of DNN models. We integrate local features, such as saliency maps and interpretable input masks, into the data structure to help users to infer the model decision boundaries. We show that the boundary graph is able to address the reliability issues in many local interpretation methods. We further design an algorithm named hidden-layer aware p-test to measure the reliability of the interpretations. Our experiments show that SMINT is able to achieve above 99% fidelity to corresponding DNN models on both MNIST and ImageNet by sharing only a tiny fraction of training data to make these models interpretable. The human pilot study demonstrates that SMINT provides better interpretability compared with existing methods. Moreover, we demonstrate that SMINT is able to assist model tuning for better performance on different user data. Huijun Wu 0001, Chen Wang 0008, Richard Nock, Wei Wang 0011, Jie Yin 0001, Kai Lu 0001, Liming Zhu 0001 |
ACM Trans. Web | 1 |
| 2019 | Adversarial Examples for Graph Data: Deep Insights into Attack and DefenseabstractGraph deep learning models, such as graph convolutional networks (GCN) achieve state-of-the-art performance for tasks on graph data. However, similar to other deep learning models, graph deep learning models are susceptible to adversarial attacks. However, compared with non-graph data the discrete nature of the graph connections and features provide unique challenges and opportunities for adversarial attacks and defenses. In this paper, we propose techniques for both an adversarial attack and a defense against adversarial attacks. Firstly, we show that the problem of discrete graph connections and the discrete features of common datasets can be handled by using the integrated gradient technique that accurately determines the effect of changing selected features or edges while still benefiting from parallel computations. In addition, we show that an adversarially manipulated graph using a targeted attack statistically differs from un-manipulated graphs. Based on this observation, we propose a defense approach which can detect and recover a potential adversarial perturbation. Our experiments on a number of datasets show the effectiveness of the proposed techniques. Huijun Wu 0001, Chen Wang 0008, Yuriy Tyshetskiy, Andrew Docherty, Kai Lu 0001, Liming Zhu 0001 |
IJCAI | 1 |
| 2019 | A Case Based Deep Neural Network Interpretability Framework and Its User Study
Rimmal Nadeem, Huijun Wu 0001, Hye-Young Paik, Chen Wang 0008 |
WISE | 2 |
| 2018 | One Size Does Not Fit All: The Case for Chunking Configuration in Backup DeduplicationabstractData backup is regularly required by both enterprise and individual users to protect their data from unexpected loss. There are also various commercial data deduplication systems or software that help users to eliminate duplicates in their backup data to save storage space. In data deduplication systems, the data chunking process splits data into small chunks. Duplicate data is identified by comparing the fingerprints of the chunks. The chunk size setting has significant impact on deduplication performance. A variety of chunking algorithms have been proposed in recent studies. In practice, existing systems often set the chunking configuration in an empirical manner. A chunk size of 4KB or 8KB is regarded as the sweet spot for good deduplication performance. However, the data storage and access patterns of users vary and change along time, as a result, the empirical chunk size setting may not lead to a good deduplication ratio and sometimes results in difficulties of storage capacity planning. Moreover, it is difficult to make changes to the chunking settings once they are put into use as duplicates in data with different chunk size settings cannot be eliminated directly. In this paper, we propose a sampling-based chunking method and develop a tool named SmartChunker to estimate the optimal chunking configuration for deduplication systems. Our evaluations on real-world datasets demonstrate the efficacy and efficiency of SmartChunker. Huijun Wu 0001, Chen Wang 0008, Kai Lu 0001, Yinjin Fu, Liming Zhu 0001 |
CCGrid | 1 |
| 2018 | HDM-MC in-Action: A Framework for Big Data Analytics across Multiple ClustersabstractBig data are increasingly collected and stored in a highly distributed infrastructures due to the development of several emerging technologies including sensor network, cloud computing, IoT and mobile computing among many other emerging technologies. In practice, the majority of existing big data processing frameworks (e.g., Hadoop, Spark, Flink) are designed based on the single-cluster setup with the assumptions of centralized management and homogeneous connectivity which makes them sub-optimal and sometimes infeasible to be applied for scenarios that require implementing data analytics jobs on highly distributed data sets (across racks, data centers or multi organizations). We demonstrate HDM-MC, a big data processing framework that is designed to enable the capability of performing large scale data analytics across multi-clusters with minimum extra overhead due to additional scheduling requirements. We describe the architecture and realization of the system using a step-by-step example scenario. Dongyao Wu, Sherif Sakr, Liming Zhu 0001, Sung Une Lee, Huijun Wu 0001 |
ICDCS | 5 |
| 2018 | Sharing Deep Neural Network Models with InterpretationabstractDespite outperforming humans in many tasks, deep neural network models are also criticized for the lack of transparency and interpretability in decision making. The opaqueness results in uncertainty and low confidence when deploying such a model in model sharing scenarios, where the model is developed by a third party. For a supervised machine learning model, sharing training process including training data is a way to gain trust and to better understand model predictions. However, it is not always possible to share all training data due to privacy and policy constraints. In this paper, we propose a method to disclose a small set of training data that is just sufficient for users to get the insight into a complicated model. The method constructs a boundary tree using selected training data and the tree is able to approximate the complicated deep neural network models with high fidelity. We show that data point pairs in the tree give users significantly better understanding of the model decision boundaries and paves the way for trustworthy model sharing. Huijun Wu 0001, Chen Wang 0008, Jie Yin 0001, Kai Lu 0001, Liming Zhu 0001 |
WWW | 1 |
| 2018 | A Differentiated Caching Mechanism to Enable Primary Storage Deduplication in CloudsabstractExisting primary deduplication techniques either use inline caching to exploit locality in primary workloads or use post-processing deduplication to avoid the negative impact on I/O performance. However, neither of them works well in the cloud servers running multiple services for the following two reasons: First, the temporal locality of duplicate data writes varies among primary storage workloads, which makes it challenging to efficiently allocate the inline cache space and achieve a good deduplication ratio. Second, the post-processing deduplication does not eliminate duplicate I/O operations that write to the same logical block address as it is performed after duplicate blocks have been written. A hybrid deduplication mechanism is promising to deal with these problems. Inline fingerprint caching is essential to achieving efficient hybrid deduplication. In this paper, we present a detailed analysis of the limitations of using existing caching algorithms in primary deduplication in the cloud. We reveal that existing caching algorithms either perform poorly or incur significant memory overhead in fingerprint cache management. To address this, we propose a novel fingerprint caching mechanism that estimates the temporal locality of duplicates in different data streams and prioritizes the cache allocation based on the estimation. We integrate the caching mechanism and build a hybrid deduplication system. Our experimental results show that the proposed mechanism provides significant improvement for both deduplication ratio and overhead reduction. Huijun Wu 0001, Chen Wang 0008, Yinjin Fu, Sherif Sakr, Kai Lu 0001, Liming Zhu 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2017 | Towards Big Data Analytics across Multiple ClustersabstractBig data are increasingly collected and stored in a highly distributed infrastructures due to the development of sensor network, cloud computing, IoT and mobile computing among many other emerging technologies. In practice, the majority of existing big-data-processing frameworks (e.g., Hadoop and Spark) are designed based on the single-cluster setup with the assumptions of centralized management and homogeneous connectivity which makes them sub-optimal and sometimes infeasible to apply for scenarios that require implementing data analytics jobs on highly distributed data sets (across racks, data centers or multi-organizations). In order to tackle this challenge, we present HDM-MC, a multi-cluster big data processing framework which is designed to enable the capability of performing large scale data analytics across multi-clusters with minimum extra overhead due to additional scheduling requirements. In this paper, we present the architecture and realization of the system. In addition, we evaluate the performance of our framework in comparison to other state-of-art single cluster big data processing frameworks. Dongyao Wu, Sherif Sakr, Liming Zhu 0001, Huijun Wu 0001 |
CCGrid | 4 |