EDBT 2026 Demo / reviewers in the wild / expert
Qiang Guan
dblp:20/1255
· DBLP profile ↗
72ranked-venue papers
13as first author
42since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 27 · 5 first-author · 13 since 2021Artificial intelligence and machine learning · 16 · 12 since 2021Security and privacy · 10 · 5 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 9 since 2021Computer networks · 8 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 8 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Uncertainty Quantification Driven Benchmarking and Characterization of Noisy Quantum Backends: A VQE Case StudyabstractBenchmarking noisy quantum computers for variational workloads requires more than a single best energy number: two backends may reach similar energies while differing sharply in convergence speed, stability, and parameter sensitivity. We present a compact uncertainty quantification driven benchmarking study for the Variational Quantum Eigensolver (VQE), combining Bayesian optimization with Metropolis-Hastings refinement to search a backend conditioned VQE landscape and characterizing each backend by its best observed energy, evaluations to target, objective dispersion, and Sobol sensitivity fingerprint. On a 92 parameters LiH VQE across seven IBM fake backends, Brisbane wins overall (best energy 1.0207, top normalized quality), Kyoto matches its energy but converges five times more slowly, and Osaka trades a modest energy penalty for near optimal speed and low dispersion. The Sobol fingerprints further show that the dominant VQE parameters differ across backends, indicating backend specific robust regions rather than a universal calibration. Uncertainty quantification therefore provides a practical application level view of backend quality for near term variational workloads. Priyabrata Senapati, Waylon Luo, Bo Peng 0024, Qiang Guan |
ACM Great Lakes Symposium on VLSI | 4 |
| 2026 | End-to-end motion detection via multi-scale spatial-temporal feature fusion for dual-view 3D macaque behavior quantification
Zongli Jiang, Qiang Guan, Xibo Ma |
Expert Syst. Appl. | 3 |
| 2026 | Ethics of trustworthy AI in healthcare: Challenges, principles, and practical pathwaysabstractArtificial Intelligence (AI) is transforming healthcare by enhancing diagnostics, personalizing treatment planning, and streamlining patient care. Yet, its adoption is hindered by persistent ethical challenges, including algorithmic bias, lack of transparency, privacy risks, and unclear accountability. Existing international frameworks articulate high-level principles but seldom provide operational guidance for clinical deployment. We bridge this gap by synthesizing trust dimensions for healthcare, with measurable metrics for fairness, explainability, privacy, accountability, and robustness, and proposing the Healthcare AI Trustworthiness Index (HAITI), a composite, context-aware readiness score with explicit normalization, weighting, and uncertainty reporting. We outline a development–deployment–governance blueprint and present two case studies (diagnostic bias mitigation; privacy-preserving federated learning). Together, these contributions translate ethical principles into measurable practices that can foster trust, improve equity, and accelerate responsible AI integration in clinical settings. Pegah Ahadian, Wei Xu 0020, Dongfang Liu, Qiang Guan |
Neurocomputing | 4 |
| 2026 | Towards Explainable Quantum AI: Informing the Encoder Selection of Quantum Neural Networks via VisualizationabstractQuantum Neural Networks (QNNs) represent a promising fusion of quantum computing and neural network architectures, offering speed-ups and efficient processing of high-dimensional, entangled data. A crucial component of QNNs is the encoder, which maps classical input data into quantum states. However, choosing suitable encoders remains a significant challenge, largely due to the lack of systematic guidance and the trial-and-error nature of current approaches. This process is further impeded by two key challenges: (1) the difficulty in evaluating encoded quantum states prior to training, and (2) the lack of intuitive methods for analyzing an encoder's ability to effectively distinguish data features. To address these issues, we introduce a novel visualization tool, XQAI-Eyes, which enables QNN developers to compare classical data features with their corresponding encoded quantum states and to examine the mixed quantum states across different classes. By bridging classical and quantum perspectives, XQAI-Eyes facilitates a deeper understanding of how encoders influence QNN performance. Evaluations across diverse datasets and encoder designs demonstrate XQAI-Eyes's potential to support the exploration of the relationship between encoder design and QNN effectiveness, offering a holistic and transparent approach to optimizing quantum encoders. Moreover, domain experts used XQAI-Eyes to derive two key practices for quantum encoder selection, grounded in the principles of pattern preservation and feature mapping. Shaolun Ruan, Rohan Ramakrishna, Chao Ren 0006, Rudai Yan, Qiang Guan, Jiannan Li, Yong Wang 0021 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | ALYMPICS: LLM Agents Meet Game TheoryabstractGame theory is a branch of mathematics that studies strategic interactions among rational agents. We propose Alympics (Olympics for Agents), a systematic framework utilizing Large Language Model (LLM) agents for empirical game theory research. Alympics creates a versatile platform for studying complex game theory problems, bridging the gap between theoretical game theory and empirical investigations by providing a controlled environment for simulating human-like strategic interactions with LLM agents. In our pilot case study, the “Water Allocation Challenge”, we explore Alympics through a challenging strategic game focused on the multi-round auction of scarce survival resources. This study demonstrates the framework’s ability to qualitatively and quantitatively analyze game determinants, strategies, and outcomes. Additionally, we conduct a comprehensive human assessment and an in-depth evaluation of LLM agents in rational strategic decision-making scenarios. Our findings highlight LLM agents’ potential to advance game theory knowledge and expand the understanding of their proficiency in emulating human strategic behavior. Shaoguang Mao, Yuzhe Cai, Yan Xia 0005, Wenshan Wu, Xun Wang 0012, Qiang Guan, Tao Ge 0001, Furu Wei |
COLING | 7 |
| 2025 | UQ-VarQA: Benchmarking and Characterizing NISQ Computers Through Uncertainty Quantification of Variational Quantum AlgorithmsabstractQuantum computing offers speedups, but NISQ processors face hardware-induced errors that degrade fidelity and reproducibility. We introduce an uncertainty-aware benchmarking framework that combines uncertainty quantification with global sensitivity analysis to evaluate not only peak fidelity but also its reliability over time. Using Bayesian Optimization with SGLD refinement under fixed budgets, seeds, bounds, and trust-region rules, and repeating runs across days, we capture calibration drift and quantify efficiency, stability, landscape complexity, and maintenance cost. A noise-aware Gaussian process surrogate provides scalable sensitivity estimates without error mitigation. Applied to VQAMET and VQC on three IBMQ backends, the framework delivers actionable guidance for co-tuning and backend selection, complementing and often outperforming quantum volume style metrics. Priyabrata Senapati, Shengye Zhu, Bo Peng 0024, Bo Fang 0002, Qiang Guan |
ICCD | 5 |
| 2025 | Can Large Language Models Understand Intermediate Representations in Compilers?abstractIntermediate Representations (IRs) play a critical role in compiler design and program analysis, yet their comprehension by *Large Language Models* (LLMs) remains underexplored. In this paper, we present an explorative empirical study evaluating the capabilities of six state-of-the-art LLMs—GPT-4, GPT-3, DeepSeek, Gemma 2, Llama 3, and Code Llama—in understanding IRs. Specifically, we assess model performance across four core tasks: *control flow graph reconstruction*, *decompilation*, *code summarization*, and *execution reasoning*. While LLMs exhibit competence in parsing IR syntax and identifying high-level structures, they consistently struggle with instruction-level reasoning, especially in control flow reasoning, loop handling, and dynamic execution. Common failure modes include misinterpreting branching instructions, omitting critical operations, and relying on heuristic reasoning rather than on precise instruction-level logic. Our findings highlight the need for IR-specific enhancements in LLM design. We recommend fine-tuning on structured IR datasets and integrating control-flow-sensitive architectures to improve the models’ effectiveness on IR-related tasks. All the experimental data and source code are publicly available at [https://github.com/hjiang13/LLM4IR](https://github.com/hjiang13/LLM4IR). Hailong Jiang, Yao Wan 0001, Bo Fang 0002, Hongyu Zhang 0002, Ruoming Jin, Qiang Guan |
ICML | 7 |
| 2025 | Adaptive Job Scheduling in Quantum Clouds Using Reinforcement LearningabstractPresent-day quantum systems face critical bottlenecks, including limited qubit counts, brief coherence intervals, and high susceptibility to errors—all of which obstruct the execution of large and complex circuits. The advancement of quantum algorithms has outpaced the capabilities of existing quantum hardware, making it difficult to scale computations effectively. Additionally, inconsistencies in hardware performance and pervasive quantum noise undermine system stability and computational accuracy. To optimize quantum workloads under these constraints, strategic approaches to task scheduling and resource coordination are essential. One of the persistent challenges in this domain is how to efficiently divide and execute large circuits across multiple quantum processors (QPUs), especially in error-prone environments. In response, we introduce a simulation-based tool that supports distributed scheduling and concurrent execution of quantum jobs on networked QPUs connected via real-time classical channels. The tool models circuit decomposition for workloads that surpass individual QPU limits, allowing for parallel execution through inter-processor communication. Using this simulation environment, we compare four distinct scheduling techniques—among them, a model informed by reinforcement learning. These strategies are evaluated across multiple metrics, including runtime efficiency, fidelity preservation, and communication costs. Our analysis underscores the trade-offs inherent in each approach and highlights how parallelized, noise-aware scheduling can meaningfully improve computational throughput in distributed quantum infrastructures. Waylon Luo, Jiapeng Zhao, Tong Zhan, Qiang Guan |
ICPP | 4 |
| 2025 | KIS-S: A GPU-Aware Kubernetes Inference Simulator with RL-Based Auto-ScalingabstractAutoscaling GPU inference workloads in Kubernetes remains challenging due to the reactive and threshold-based nature of default mechanisms such as the Horizontal Pod Autoscaler (HPA), which struggle under dynamic and bursty traffic patterns and lack integration with GPU-level metrics. We present KIS-S, a unified framework that combines KISim, a GPU-aware Kubernetes Inference Simulator, with KIScaler, a Proximal Policy Optimization (PPO)-based autoscaler. KISim enables safe, high-fidelity scheduling emulation with real GPU hardware and Prometheus integration, while KIScaler learns latency-aware and resource-efficient scaling policies entirely in simulation. KIScaler observes system metrics via Prometheus and adjusts replica counts via the Kubernetes API. We evaluate KIS-S across four synthetic traffic patterns-ramp, periodic, random, and spike-and compare it against conventional baselines including HPA and fixed-resource deployments. Despite training with synthetic feedback due to single-GPU hardware constraints, KIScaler's moving average reward improves from 1.05 to 1.84 (a 75.2 % increase) over 100 training episodes, reduces P95 latency by up to$6.7 \times$over CPU-only baselines, and generalizes across all traffic patterns without retraining. These results highlight the value of combining simulation and learning, bridging the gap between reactive autoscaling and intelligent orchestration for scalable, GPU-accelerated Kubernetes environments. Guilin Zhang, Wulan Guo, Qiang Guan, Hailong Jiang |
IPCCC | 4 |
| 2025 | A Digital Twin of Scalable Quantum Clouds
Waylon Luo, Betis Baheri, Travis S. Humble, Jiapeng Zhao, Tong Zhan, Rajan Maharjan, Qiang Guan |
SIGSIM-PADS | 7 |
| 2025 | QDockBank: A dataset for Ligand Docking on Protein Fragments Predicted on Utility-Level Quantum ComputersabstractProtein structure prediction is a core challenge in computational biology, particularly for fragments within ligand-binding regions, where accurate modeling is still difficult. Quantum computing offers a novel first-principles modeling paradigm, but its application is currently limited by hardware constraints, high computational cost, and the lack of a standardized benchmarking dataset. In this work, we present QDockBank—the first large-scale protein fragment structure dataset generated entirely using utility-level quantum computers, specifically designed for protein–ligand docking tasks. QDockBank comprises 55 protein fragments extracted from ligand-binding pockets. The dataset was generated through tens of hours of execution on superconducting quantum processors, making it the first quantum-based protein structure dataset with a total computational cost exceeding one million USD. Experimental evaluations demonstrate that structures predicted by QDockBank outperform those predicted by AlphaFold2 and AlphaFold3 in terms of both RMSD and docking affinity scores. QDockBank serves as a new benchmark for evaluating quantum-based protein structure prediction. Yuxin Yang 0001, Cheng-Chang Lu, Weiwen Jiang, Feixiong Cheng, Bo Fang 0002, Qiang Guan |
SC | 7 |
| 2024 | Adopting Trustworthy AI for Sleep Disorder Prediction: Deep Time Series Analysis with Temporal Attention Mechanism and Counterfactual ExplanationsabstractSleep disorders have a major impact on both lifestyle and health. Effective sleep disorder prediction from lifestyle and physiological data can provide essential details for early intervention. This research utilizes three deep time series models and facilitates them with explainability approaches for sleep disorder prediction. Specifically, our approach adopts Temporal Convolutional Networks (TCN), Long Short-Term Memory (LSTM) for time series data analysis, and Temporal Fusion Transformer model (TFT). Meanwhile, the temporal attention mechanism and counterfactual explanation with SHapley Additive exPlanations (SHAP) approach are employed to ensure dependable, accurate, and interpretable predictions. Finally, using a large dataset of sleep health measures, our evaluation demonstrates the effect of our method in predicting sleep disorders. Pegah Ahadian, Wei Xu 0020, Sherry Wang, Qiang Guan |
IEEE Big Data | 4 |
| 2024 | AMD: Automatic Multi-step Distillation of Large-Scale Vision Models
Cheng Han 0001, Qifan Wang 0001, Sohail A. Dianat, Majid Rabbani, Raghuveer M. Rao, Yi Fang 0008, Qiang Guan, Lifu Huang, Dongfang Liu |
ECCV (65) | 7 |
| 2024 | Radiance Field Learners As UAV First-Person Viewers
Liqi Yan, Qifan Wang 0001, Junhan Zhao, Qiang Guan, Dongfang Liu |
ECCV (61) | 4 |
| 2024 | A Deep Multimodal Representation Learning Framework for Accurate Molecular Properties PredictionabstractDrug discovery is a challenging process, requiring the optimization of compounds to become safe and effective. Predicting molecular properties is an indispensable step in the drug discovery pipeline. Traditionally, this process is costly, involving multiple rounds of experiments, rendering it impractical for every candidate compound. Deep learning techniques have emerged as a promising approach to drug discovery to reduce the cost during the process. However, prevalent research in deep learning models focused on predicting molecular properties has primarily fixated on single-modal models, neglecting the potential benefits of combining different data modalities. To overcome this limitation, we introduce MRL-Mol: a deep Multimodal Representation Learning framework for accurate Molecular properties prediction. MRL-Mol harnesses three data modalities: sequence, graph, and image, augmenting the depth of comprehension. Leveraging a large-scale unlabeled dataset ( 1M unique molecules), we pretrain MRL-Mol to extract inter- and intra-modal information. Our study demonstrates the superior performance of MRL-Mol in predicting molecular properties across six benchmark datasets. Notably, MRL-Mol outperforms other state-of-the-art molecular properties prediction models. These findings suggest that by combining information from multiple data modalities, MRL-Mol can comprehend molecules better than single-modal deep learning models and identify molecular properties with better accuracy. Yuxin Yang 0001, Pegah Ahadian, Abby Jerger, Jeremy Zucker, Feixiong Cheng, Qiang Guan |
ACM Great Lakes Symposium on VLSI | 8 |
| 2024 | Prototypical Transformer As Unified Motion LearnersabstractIn this work, we introduce the Prototypical Transformer (ProtoFormer), a general and unified framework that approaches various motion tasks from a prototype perspective. ProtoFormer seamlessly integrates prototype learning with Transformer by thoughtfully considering motion dynamics, introducing two innovative designs. First, Cross-Attention Prototyping discovers prototypes based on signature motion patterns, providing transparency in understanding motion scenes. Second, Latent Synchronization guides feature representation learning via prototypes, effectively mitigating the problem of motion uncertainty. Empirical results demonstrate that our approach achieves competitive performance on popular motion tasks such as optical flow and scene depth. Furthermore, it exhibits generality across various downstream tasks, including object tracking and video stabilization. Cheng Han 0001, Yawen Lu, James Liang, Zhiwen Cao, Qifan Wang 0001, Qiang Guan, Sohail A. Dianat, Raghuveer M. Rao, Tong Geng, Zhiqiang Tao, Dongfang Liu |
ICML | 7 |
| 2024 | A Lightweight Convolutional Neural Network for Personalized Blood Pressure Estimation Based on PhotoplethysmographyabstractContinuous blood pressure (BP) monitoring holds potential in preventing and detecting cardiovascular disease (CVD). Photoplethysmography (PPG)-based BP measuring systems with noninvasive and continuous properties are of enormous research value in biomedical science. However, mainstream technologies such as Transformer are not viable for implementation in compact devices due to their significant processing complexity. Additionally, the unique individual variability of biosignals limits the generalization performance of the model. In this work, we propose a personalized modeling approach employing a lightweight convolutional neural network(CNN) that streamlines the structure while preserving the model’s ability for long-term predictions. We leverage continuous records from 413 and 30 subjects extracted from MIMIC-III for model pretraining and personalization. The final results demonstrate that, calibrated with only 30 labeled windows for the target subject, the personalized model achieves an estimation error of -0.218±7.657 mmHg for systolic BP (SBP) and -0.261±4.898 mmHg for diastolic BP (DBP). This performance meets the standard requirements of the Association for the Advancement of Medical Instrumentation (AAMI). The lightweight structure and personalized modeling approach improve the accuracy of BP estimation while reducing the difficulty of deployment. Caijie Qin, Zhaoyi Ning, Qiang Guan, Xibo Ma |
IJCNN | 6 |
| 2024 | Diffusion-Inspired Truncated Sampler for Text-Video RetrievalabstractPrevalent text-to-video retrieval methods represent multimodal text-video data in a joint embedding space, aiming at bridging the relevant text-video pairs and pulling away irrelevant ones. One main challenge in state-of-the-art retrieval methods lies in the modality gap, which stems from the substantial disparities between text and video and can persist in the joint space. In this work, we leverage the potential of Diffusion models to address the text-video modality gap by progressively aligning text and video embeddings in a unified space. However, we identify two key limitations of existing Diffusion models in retrieval tasks: The L2 loss does not fit the ranking problem inherent in text-video retrieval, and the generation quality heavily depends on the varied initial point drawn from the isotropic Gaussian, causing inaccurate retrieval. To this end, we introduce a new Diffusion-Inspired Truncated Sampler (DITS) that jointly performs progressive alignment and modality gap modeling in the joint embedding space. The key innovation of DITS is to leverage the inherent proximity of text and video embeddings, defining a truncated diffusion flow from the fixed text embedding to the video embedding, enhancing controllability compared to adopting the isotropic Gaussian. Moreover, DITS adopts the contrastive loss to jointly consider the relevant and irrelevant pairs, not only facilitating alignment but also yielding a discriminatively structured embedding. Experiments on five benchmark datasets suggest the state-of-the-art performance of DITS. We empirically find that DITS can also improve the structure of the CLIP embedding space. Code is available at https://github.com/Jiamian- Wang/DITS-text-video-retrieval Jiamian Wang, Pichao Wang, Dongfang Liu, Qiang Guan, Sohail A. Dianat, Majid Rabbani, Raghuveer M. Rao, Zhiqiang Tao |
NeurIPS | 4 |
| 2024 | HAppA: A Modular Platform for HPC Application Resilience Analysis with LLMs EmbeddedabstractHigh-performance computing (HPC) systems are increasingly vulnerable to soft errors, which pose significant challenges in maintaining computational accuracy and reliability. Predicting the resilience of HPC applications to these errors is crucial for robust code protection and detailed resilience analysis. In this study, we present HAppA, a modular platform designed for HPC Application Resilience Analysis. Embedding Large Language Models (LLMs), HAppA addresses understanding the context information of long code sequences typical in HPC applications. HAppA implements a novel code representation module that chunks the code into fixed-size segments and aggregates the embeddings of these segments. Three aggregation methods have been explored: MeanPooling, MaxPooling, and LSTM-based techniques. We built a DAtaset for REsilience analysis using Fault Injection (FI), named DARE. Using our DARE dataset, HAppA is trained for regression prediction tasks. Our evaluation results demonstrate the predictive accuracy of HAppA compared to other models, particularly noting that the LSTM-based aggregation method - HAppA-LSTM - achieves a mean squared error (MSE) of 0.078 for SDC prediction, surpassing the existing state-of-the-art PARIS model, which recorded an MSE of 0.1172. Additionally, HAppA with the KeyBERT model extracts a list of key words representing the source code. A comprehensive importance analysis of these key words further elucidates the code patterns contributing to the error rate. These findings highlight the effectiveness of HAppA in analyzing the resilience of HPC applications and establish a new benchmark for predictive accuracy in resilience. Hailong Jiang, Bo Fang 0002, Kevin J. Barker, Ruoming Jin, Qiang Guan |
SRDS | 7 |
| 2024 | A Systematic Methodology to Compute the Quantum Vulnerability Factors for Quantum CircuitsabstractQuantum computing is one of the most promising technology advances of the latest years. Qubits are highly sensitive to noise, which can make the output useless. Lately, it has been shown that superconducting qubits are extremely susceptible to external sources of faults, such as ionizing radiation. When adopted in large scale, radiation-induced errors are expected to become a serious challenge for qubits reliability. We propose an evaluation of the impact of transient faults in the execution of quantum circuits on superconducting chips. Inspired by the Architectural and Program Vulnerability Factors, widely used for classical computation, we propose the Quantum Vulnerability Factor (QVF) to measure the impact of qubit corruption on the circuit output. We model faults, and design a fault injector, based on the latest studies on real machines and radiation experiments. We report the finding of more than 388,000,000 fault injections, considering single and double faults, on three algorithms, identifying the faults and qubits that are more likely to impact the output. We give guidelines on how to map the qubits in real devices to reduce the output error and to reduce the probability of having a radiation-induced corruption modifying the output. Finally, we compare simulations with experiments on physical quantum computers. Daniel Oliveira 0002, Edoardo Giusto, Betis Baheri, Qiang Guan, Bartolomeo Montrucchio, Paolo Rech |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2024 | An HPC-Container Based Continuous Integration Tool for Detecting Scaling and Performance Issues in HPC ApplicationsabstractTesting is one of the most important steps in software development–it ensures the quality of software. Continuous Integration (CI) is a widely used testing standard that can report software quality to the developer in a timely manner during development progress. Performance, especially scalability, is another key factor for High Performance Computing (HPC) applications. There are many existing profiling and performance tools for HPC applications, but none of these are integrated into CI tools. In this work, we propose BeeSwarm, an HPC container based parallel scaling performance system that can be easily applied to the current CI test environments. BeeSwarm is mainly designed for HPC application developers who need to monitor how their applications can scale on different compute resources. We demonstrate BeeSwarm using three different HPC applications: CoMD, LULESH and NWChem. We utilize GitHub Actions and provision resources from Google Compute Engine. Our results show that BeeSwarm can be used for scalability and performance testing of a variety of HPC applications, allowing developers to monitor application performance over time. Jake Tronge, Jieyang Chen, Patricia Grubel, Tim Randles, Rusty Davis, Quincy Wofford, Steven Anaya, Qiang Guan |
IEEE Trans. Serv. Comput. | 8 |
| 2024 | QuantumEyes: Towards Better Interpretability of Quantum CircuitsabstractQuantum computing offers significant speedup compared to classical computing, which has led to a growing interest among users in learning and applying quantum computing across various applications. However, quantum circuits, which are fundamental for implementing quantum algorithms, can be challenging for users to understand due to their underlying logic, such as the temporal evolution of quantum states and the effect of quantum amplitudes on the probability of basis quantum states. To fill this research gap, we propose QuantumEyes, an interactive visual analytics system to enhance the interpretability of quantum circuits through both global and local levels. For the global-level analysis, we present three coupled visualizations to delineate the changes of quantum states and the underlying reasons: a Probability Summary View to overview the probability evolution of quantum states; a State Evolution View to enable an in-depth analysis of the influence of quantum gates on the quantum states; a Gate Explanation View to show the individual qubit states and facilitate a better understanding of the effect of quantum gates. For the local-level analysis, we design a novel geometrical visualization dandelion chart to explicitly reveal how the quantum amplitudes affect the probability of the quantum state. We thoroughly evaluated QuantumEyes as well as the novel dandelion chart integrated into it through two case studies on different types of quantum algorithms and in-depth expert interviews with 12 domain experts. The results demonstrate the effectiveness and usability of our approach in enhancing the interpretability of quantum circuits. Shaolun Ruan, Qiang Guan, Paul Griffin 0001, Ying Mao 0001, Yong Wang 0021 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | VIOLET: Visual Analytics for Explainable Quantum Neural NetworksabstractWith the rapid development of Quantum Machine Learning, quantum neural networks (QNN) have experienced great advancement in the past few years, harnessing the advantages of quantum computing to significantly speed up classical machine learning tasks. Despite their increasing popularity, the quantum neural network is quite counter-intuitive and difficult to understand, due to their unique quantum-specific layers (e.g., data encoding and measurement) in their architecture. It prevents QNN users and researchers from effectively understanding its inner workings and exploring the model training status. To fill the research gap, we propose VIOLET, a novel visual analytics approach to improve the explainability of quantum neural networks. Guided by the design requirements distilled from the interviews with domain experts and the literature survey, we developed three visualization views: the Encoder View unveils the process of converting classical input data into quantum states, the Ansatz View reveals the temporal evolution of quantum states in the training process, and the Feature View displays the features a QNN has learned after the training process. Two novel visual designs, i.e., satellite chart and augmented heatmap, are proposed to visually explain the variational parameters and quantum circuit measurements respectively. We evaluate VIOLET through two case studies and in-depth interviews with 12 domain experts. The results demonstrate the effectiveness and usability of VIOLET in helping QNN users and developers intuitively understand and explore quantum neural networks. Shaolun Ruan, Zhiding Liang, Qiang Guan, Paul Griffin 0001, Xiaolin Wen, Yanna Lin, Yong Wang 0021 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2023 | Battle Against Fluctuating Quantum Noise: Compression-Aided Framework to Enable Robust Quantum Neural NetworkabstractRecently, we have been witnessing the scale-up of superconducting quantum computers; however, the noise of quantum bits (qubits) is still an obstacle for real-world applications to leveraging the power of quantum computing. Although there exist error mitigation or error-aware designs for quantum applications, the inherent fluctuation of noise (a.k.a., instability) can easily collapse the performance of error-aware designs. What’s worse, users can even not be aware of the performance degradation caused by the change in noise. To address both issues, in this paper we use Quantum Neural Network (QNN) as a vehicle to present a novel compression-aided framework, namely QuCAD, which will adapt a trained QNN to fluctuating quantum noise. In addition, with the historical calibration (noise) data, our framework will build a model repository offline, which will significantly reduce the optimization time in the online adaption process. Emulation results on an earthquake detection dataset show that QuCAD can achieve 14.91% accuracy gain on average in 146 days over a noise-aware training approach. For the execution on a 7-qubit IBM quantum processor, ibm-jakarta, QuCAD can consistently achieve 12.52% accuracy gain on earthquake detection. Zhirui Hu, Youzuo Lin, Qiang Guan, Weiwen Jiang |
DAC | 3 |
| 2023 | Deep Learning-based Student Learning Behavior Understanding Framework in Real Classroom SceneabstractDeep learning techniques have emerged as valuable tools for video analysis and motion detection. Recent advancements in this field have shown promising results. Our objective is to leverage these video understanding techniques to aid teachers in evaluating their teaching quality and enhancing their effectiveness in the classroom. However, existing research on student behavior analysis primarily focuses on recognizing actions pertaining to classroom management, neglecting the identification of “learning behaviors” exhibited by students. To address this limitation, we introduce a novel video dataset specifically designed to capture the nuances of “learning behaviors” displayed by primary-grade students in the mathematics classroom, along with a dedicated student localization dataset focused on detecting the location of individuals. Our approach introduces a framework that utilizes deep learning-based object detection and action recognition techniques trained on our curated datasets to analyze and comprehend student learning behaviors in the classroom. To assess the performance of our approach, we conduct separate tests on our object detection and action recognition models. Sub-sequently, our framework is applied to a collection of recorded 360-degree classroom videos, enabling a thorough evaluation of its capabilities. Yuxin Yang 0001, Zhengyong Ren, Chris Lenart, Ashton Corsello, Karl W. Kosko, Simon Su, Qiang Guan |
ICMLA | 7 |
| 2023 | Gaze Analysis System for Immersive 360° Video for Preservice Teacher EducationabstractUnified systems for multi-sensor devices, particularly eye-tracking in Virtual Reality (VR), are intricate and often require the listening and streaming of multichannel data. In this project, we propose a visual analysis framework for replicating a participant's viewing involvement by interpreting head movements as rotations and point-of-gaze (POG) as on-screen indicators. Our solution suggests an additional layer of system for near-real-time for processing and analyzing this multi-device data to connect with the data and enable both near-real-time or subsequent offline viewing of the entire VR eye-tracking session. Moreover, our method provides a no-batteries-need solution to create traditional eye-tracking visualization techniques. Finally, we apply three prior education technology analysis metrics: higher density gaze for students, shorter fixation time, and less fixation duration variance for students to determine expertise levels in this system. We systematically establish a ubiquitous, multi-device, eye-tracking solution to incorporate this approach. We evaluate the effectiveness of our system through a user study, using both expertise and non- expertise levels, and selectively surveying to ascertain the quality of the replicated experience and we test the system by running a real-world user study with sixty four different participants. We demonstrate the application's significance and potential to integrate prior analysis metrics using the collected data which this data collection and analysis have been approved by IRB. Chris Lenart, Pegah Ahadian, Yuxin Yang 0001, Simon Suo, Ashton Corsello, Karl W. Kosko, Qiang Guan |
ACM Multimedia | 7 |
| 2023 | Enhancing Detailed Feedback to Chinese Writing Learners Using a Soft-Label Driven Approach and Tag-Aware Ranking Model
Yuzhe Cai, Shaoguang Mao, Chenshuo Wang, Tao Ge 0001, Wenshan Wu, Yan Xia 0005, Chanjin Zheng, Qiang Guan |
NLPCC (1) | 8 |
| 2023 | Continuous Exploration via Multiple Perspectives in Sparse Reward Environment
Zhongpeng Chen, Qiang Guan |
PRCV (3) | 2 |
| 2023 | Visilience: An Interactive Visualization Framework for Resilience Analysis using Control-Flow GraphabstractSoft errors have become one of the main concerns for the resilience of HPC applications, as these errors can cause HPC applications to generate serious outcomes such as silent data corruption (SDC). Many approaches have been proposed to analyze the resilience of HPC applications. However, existing studies rarely address the challenges of analysis result perception. Specifically, resilience analysis techniques often produce a massive volume of unstructured data, making it difficult for programmers to perform resilience analysis due to non-intuitive raw data. Furthermore, different analysis models produce diverse results with multiple levels of detail, which can create obstacles to compare and explore the resilience of the HPC program execution. To this end, we present Visilience, an interactive VISual resILIENCE analysis framework to allow programmers to facilitate the resilience analysis of HPC applications. In particular, Visilience leverages an effective visualization approach, Control Flow Graph (CFG) to present a function execution. Furthermore, three widely used models for resilience analysis (i.e., Y-Branch, IPAS, and TRIDENT) are seamlessly integrated into the framework for resilience analysis and result comparison. Multiple case studies have been conducted to demonstrate the effectiveness of our proposed framework Visilience. Hailong Jiang, Shaolun Ruan, Bo Fang 0002, Yong Wang 0021, Qiang Guan |
PRDC | 5 |
| 2023 | VENUS: A Geometrical Representation for Quantum State VisualizationabstractAbstract Visualizations have played a crucial role in helping quantum computing users explore quantum states in various quantum computing applications. Among them, Bloch Sphere is the widely‐used visualization for showing quantum states, which leverages angles to represent quantum amplitudes. However, it cannot support the visualization of quantum entanglement and superposition, the two essential properties of quantum computing. To address this issue, we propose VENUS, a novel visualization for quantum state representation. By explicitly correlating 2D geometric shapes based on the math foundation of quantum computing characteristics, VENUS effectively represents quantum amplitudes of both the single qubit and two qubits for quantum entanglement. Also, we use multiple coordinated semicircles to naturally encode probability distribution, making the quantum superposition intuitive to analyze. We conducted two well‐designed case studies and an in‐depth expert interview to evaluate the usefulness and effectiveness of VENUS. The result shows that VENUS can effectively facilitate the exploration of quantum states for the single qubit and two qubits. Shaolun Ruan, Ribo Yuan, Qiang Guan, Yanna Lin, Ying Mao 0001, Weiwen Jiang, Zhepeng Wang 0001, Wei Xu 0020, Yong Wang 0021 |
Comput. Graph. Forum | 3 |
| 2023 | Elastic Resource Management for Deep Learning Applications in a Container ClusterabstractThe increasing demand for learning from massive datasets is restructuring our economy. Effective learning, however, involves nontrivial computing resources. Most businesses utilize commercial infrastructure providers (e.g., AWS) to host their computing clusters in the cloud, where various jobs compete for available resources. While cloud resource management is a fruitful research field that has made many advances in production, such as Kubernetes and YARN, few efforts have been invested to further optimize the system performance, especially for Deep Learning (DL) training jobs in a container cluster. This work introduces FlowCon, a system that is able to monitor the individual evaluation functions of DL jobs at runtime, and thus to make placement decisions and resource allocations elastically. We present a detailed design and implementation of FlowCon and conduct intensive experiments over various DL models. The results demonstrate that FlowCon significantly improves DL job completion time and resource utilization efficiency when compared to default systems. According to the results, FlowCon can improve the completion time by up to 68.8% and meanwhile, reduce the makespan by 18.0%, in the presence of various DL job workloads. Ying Mao 0001, Vaishali Sharma, Wenjia Zheng, Long Cheng 0003, Qiang Guan, Ang Li 0006 |
IEEE Trans. Cloud Comput. | 5 |
| 2023 | : A isualization pproah for Noie Awarenss in Quatum ComputingabstractQuantum computing has attracted considerable public attention due to its exponential speedup over classical computing. Despite its advantages, today's quantum computers intrinsically suffer from noise and are error-prone. To guarantee the high fidelity of the execution result of a quantum algorithm, it is crucial to inform users of the noises of the used quantum computer and the compiled physical circuits. However, an intuitive and systematic way to make users aware of the quantum computing noise is still missing. In this paper, we fill the gap by proposing a novel visualization approach to achieve noise-aware quantum computing. It provides a holistic picture of the noise of quantum computing through multiple interactively coordinated views: a Computer Evolution View with a circuit-like design overviews the temporal evolution of the noises of different quantum computers, a Circuit Filtering View facilitates quick filtering of multiple compiled physical circuits for the same quantum algorithm, and a Circuit Comparison View with a coupled bar chart enables detailed comparison of the filtered compiled circuits. We extensively evaluate the performance of VACSEN through two case studies on quantum algorithms of different scales and in-depth interviews with 12 quantum computing users. The results demonstrate the effectiveness and usability of VACSEN in achieving noise-aware quantum computing. Shaolun Ruan, Yong Wang 0021, Weiwen Jiang, Ying Mao 0001, Qiang Guan |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2022 | BatchLens: A Visualization Approach for Analyzing Batch Jobs in Cloud SystemsabstractCloud systems are becoming increasingly powerful and complex. It is highly challenging to identify anomalous execution behaviors and pinpoint problems by examining the overwhelming intermediate results/states in complex application workflows. Domain scientists urgently need a friendly and functional interface to understand the quality of the computing services and the performance of their applications in real time. To meet these needs, we explore data generated by job schedulers and investigate general performance metrics (e.g., utilization of CPU, memory and disk I/O). Specifically, we propose an interactive visual analytics approach, BatchLens, to provide both providers and users of cloud service with an intuitive and effective way to explore the status of system batch jobs and help them conduct root-cause analysis of anomalous behaviors in batch jobs. We demonstrate the effectiveness of BatchLens through a case study on the public Alibaba bench workload trace datasets. Shaolun Ruan, Yong Wang 0021, Hailong Jiang, Weijia Xu, Qiang Guan |
DATE | 5 |
| 2022 | QuFI: a Quantum Fault Injector to Measure the Reliability of Qubits and Quantum CircuitsabstractQuantum computing is an up-and-coming technology that is expected to revolutionize the computation paradigm in the next few years. Qubits, the primary computing elements of quantum circuits, exploit the quantum physics proprieties to increase the parallelism and speed of computation drastically. Unfortunately, besides being intrinsically noisy, qubits have also been shown to be highly susceptible to external sources of faults, such as ionizing radiation. The latest discoveries highlight a much higher radiation sensitivity of qubits than traditional transistors and identify a much more complex fault model than bit-flip.We propose a framework to identify the quantum circuits sensitivity to radiation-induced faults and the probability for a fault in a qubit to propagate to the output. Based on the latest studies and radiation experiments performed on real quantum machines, we model the transient faults in a qubit as a phase shift with a parametrized magnitude. Additionally, our framework can inject multiple qubit faults, tuning the phase shift magnitude based on the proximity of the qubit to the particle strike location. As we show in the paper, the proposed fault injector is highly flexible, and it can be used on both quantum circuit simulators and real quantum machines. We report the finding of more than 285, 249, 536 injections on the Qiskit simulator and 53, 248 injections on real IBM machines. We consider three quantum algorithms and identify the faults and qubits that are more likely to impact the output. We also consider the fault propagation dependence on the circuit scale, showing that the reliability profile for some quantum algorithms is scale-dependent, with increased impact from radiation-induced faults as we increase the number of qubits. Finally, we also consider multi qubits faults, showing that they are much more critical than single faults. The fault injector and the data presented in this paper are available in a public repository to allow further analysis. Daniel Oliveira 0002, Edoardo Giusto, Emanuele Dri, Nadir Casciola, Betis Baheri, Qiang Guan, Bartolomeo Montrucchio, Paolo Rech |
DSN | 6 |
| 2022 | Demo: A Multi-Perspective Video Streaming System with Privacy Preservation in Trauma RoomabstractMore and more hospitals are now deploying mul-tiple cameras in trauma room for a multi-perspective remote observation. but video surveillance system can cause privacy breach by showing and storing sensitive information of patients and staff. We use OpenPose which is a state-of-the-art human body skeletons estimation framework to extract 18 human key skeleton points. For privacy preservation, we can apply the image obfuscation techniques to human heads, we also can use human skeleton to replace the human body in the truth background. we proposed a head detection method based on the 5 key points of each head output from OpenPose. We applied the st-gcn algorithm to recognize human actions, we propose a interactive algorithm for multiple cameras to recognize and trace the same person in different cameras, Based on multi-view action recognition for the same person, we can take action recognition accuracy to a high level. Our experiment results prove that our proposed technique has a high performance in privacy protection applications. Now we focus on the interactive algorithm for multiple cameras. Zhengyong Ren, Yuxin Yang 0001, Kambiz Ghazinour, Sara Bayramzadeh, Qiang Guan |
SEC | 5 |
| 2022 | Quantum Noise in the Flow of Time: A Temporal Study of the Noise in Quantum ComputersabstractOver the last couple of years, Quantum Computing (QC) has captured the interest of computer scientists due to the fact of quantum speedup, the possibility of solving NPhard problems, and achieving higher compute power. However, mitigating the impact of the noise inside each quantum device presents an immediate challenge. These changes open up new opportunities to investigate the effect of calibration parameters for individual characteristics of each qubit in a manner of time. In this paper, we investigate the temporal behavior of noisy intermediate-scale quantum (NISQ) computers based on calibration data and the characteristics of individual devices. In particular, we collect calibration data of IBM-Q machines over the last two years and compare the quantum error robustness against the processor types, quantum topology, and quantum volumes of the IBM-Q machines. Betis Baheri, Qiang Guan, Vipin Chaudhary, Ang Li 0006 |
IOLTS | 2 |
| 2022 | MARS: Malleable Actor-Critic Reinforcement Learning SchedulerabstractIn this paper, we introduce MARS, a new scheduling system for HPC-cloud infrastructures based on a cost-aware, flexible reinforcement learning approach, which serves as an intermediate layer for next generation HPC-cloud resource manager. MARSensembles the pre-trained models from heuristic workloads and decides on the most cost-effective strategy for optimization. A whole workflow application would be split into several optimizable dependent sub-tasks, then based on the predefined resource management plan, a reward will be generated after executing a scheduled task. Lastly, MARSupdates the Deep Neural Network (DNN) model based on the reward. MARSis designed to optimize the existing models through reinforcement mechanisms. MARSadapts to the dynamics of workflow applications, selects the most cost-effective scheduling solution among pre-built scheduling strategies (backfilling, SJF, etc.) and self-learning deep neural network model at run-time. We evaluate MARSwith different real-world workflow traces. MARS can achieve 5%–60% increased performance compared to the state-of-the-art approaches. Betis Baheri, Jake Tronge, Bo Fang 0002, Ang Li 0006, Vipin Chaudhary, Qiang Guan |
IPCCC | 6 |
| 2022 | Making Invisible Visible: Data-Driven Seismic Inversion With Spatio-Temporally Constrained Data AugmentationabstractDeep learning and data-driven approaches have shown great potential in scientific domains. The promise of data-driven techniques relies on the availability of a large volume of high-quality training datasets. Due to the high cost of obtaining data through expensive physical experiments, instruments, and simulations, data augmentation techniques for scientific applications have emerged as a new direction for obtaining scientific data recently. However, existing data augmentation techniques originating from computer vision yield physically unacceptable data samples that are not helpful for the domain problems that we are interested in. In this article, we develop new data augmentation techniques based on convolutional neural networks. Specifically, our generative models leverage different physics knowledge (such as governing equations, observable perception, and physics phenomena) to improve the quality of the synthetic data. To validate the effectiveness of our data augmentation techniques, we apply them to solve a subsurface seismic full-waveform inversion using simulated CO2leakage data. Our interest is to invert for subsurface velocity models associated with very small CO2leakage. We validate the performance of our methods using comprehensive numerical tests. Via comparison and analysis, we show that data-driven seismic imaging can be significantly enhanced by using our data augmentation techniques. Particularly, the imaging quality has been improved by 15% in test scenarios of general-sized leakage and 17% in small-sized leakage when using an augmented training set obtained with our techniques. Yuxin Yang 0001, Xitong Zhang, Qiang Guan, Youzuo Lin |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Characterizing Impacts of Storage Faults on HPC Applications: A Methodology and InsightsabstractIn recent years, the increasing complexity in scientific simulations and emerging demands for training heavy artificial intelligence models require massive and fast data accesses, which urges high-performance computing (HPC) platforms to equip with more advanced storage infrastructures such as solid-state disks (SSDs). While SSDs offer high-performance I/O, the reliability challenges faced by the HPC applications under the SSD-related failures remains unclear, in particular for failures resulting in data corruptions. The goal of this paper is to understand the impact of SSD-related faults on the behaviors of complex HPC applications. To this end, we propose FFIS, a FUSE-based fault injection framework that systematically introduces storage faults into the application layer to model the errors originated from SSDs. FFIS is able to plant different I/O related faults into the data returned from underlying file systems, which enables the investigation on the error resilience characteristics of the scientific file format. We demonstrate the use of FFIS with three representative real HPC applications, showing how each application reacts to the data corruptions, and provide insights on the error resilience of the widely adopted HDF5 file format for the HPC applications. Bo Fang 0002, Daoce Wang, Sian Jin, Quincey Koziol, Zhao Zhang 0007, Qiang Guan, Surendra Byna, Sriram Krishnamoorthy, Dingwen Tao |
CLUSTER | 6 |
| 2021 | BEE Orchestrator: Running Complex Scientific Workflows on Multiple SystemsabstractIn this paper, we propose a workflow orchestration system that is able to run workflows on both HPC systems and in the cloud using HPC containers. Most existing workflow orchestration systems are only able to run workflows on one system at a time, and thus may be unable to run workflows that require more resources than what some platforms provide, and may also be unable to handle validation and fault tolerance requirements. Users may have access to a number of different systems, perhaps a mix of HPC systems and private and public clouds, but currently are only able to utilize one system at a time for running complex workflows. Utilizing HPC containers, such as Charliecloud, and a subset of the Common Workflow Language (CWL) for representing workflows, we extend the BEE Orchestration System to allow for possible scheduling of complex workflows across resources. We design and implement a scheduling component and a component for interacting with OpenStack-based HPC clusters and Google Compute Engine clouds to allow for communication between any combination of Cloud and HPC components. We demonstrate how BEE orchestrates workflows across systems, making it possible to run complex scientific applications across all systems that are available to a user. These results also show that BEE will become a viable alternative to other workflow orchestration systems. Jake Tronge, Patricia Grubel, Tim Randles, Quincy Wofford, Rusty Davis, Steven Anaya, Qiang Guan |
HiPC | 7 |
| 2021 | A Hybrid System for Learning Classical Data in Quantum StatesabstractDeep neural network powered artificial intelligence has rapidly changed our daily life with various applications. However, as one of the essential steps of deep neural networks, training a heavily-weighted network requires a tremendous amount of computing resources. Especially in the post Moore’s Law era, the limit of semiconductor fabrication technology has restricted the development of learning algorithms to cope with the increasing high intensity training data. Meanwhile, quantum computing has demonstrated its significant potential in terms of speeding up the traditionally compute-intensive workloads. For example, Google illustrated quantum supremacy by completing a sampling calculation task in 200 seconds, which is otherwise impracticable on the world’s largest supercomputers. To this end, quantum-based learning has become an area of interest, with the potential of a quantum speedup. In this paper, we propose GenQu, a hybrid and general-purpose quantum framework for learning classical data through quantum states. We evaluate GenQu with real datasets and conduct experiments on both simulations and real quantum computer IBM-Q. Our evaluation demonstrates that, compared with classical solutions, the proposed models running on GenQu framework achieve similar accuracy with a much smaller number of qubits, while significantly reducing the parameter size by up to 95.86% and converging speedup by 33.33% faster. Samuel A. Stein, Ryan L'Abbate, Wenrui Mu, Betis Baheri, Ying Mao 0001, Qiang Guan, Ang Li 0006, Bo Fang 0002 |
IPCCC | 7 |
| 2021 | BeeSwarm: Enabling Parallel Scaling Performance Measurement in Continuous Integration for HPC ApplicationsabstractTesting is one of the most important steps in software development–it ensures the quality of software. Continuous Integration (CI) is a widely used testing standard that can report software quality to the developer in a timely manner during development progress. Performance, especially scalability, is another key factor for High Performance Computing (HPC) applications. There are many existing profiling and performance tools for HPC applications, but none of these are integrated into CI tools. In this work, we propose BeeSwarm, an HPC container based parallel scaling performance system that can be easily applied to the current CI test environments. BeeSwarm is mainly designed for HPC application developers who need to monitor how their applications can scale on different compute resources. We demonstrate BeeSwarm using a multi-physics HPC application with Travis CI, GitLab CI and GitHub Actions while using ChameleonCloud and Google Compute Engine as the compute backends. Our results show that BeeSwarm can be used for scalability and performance testing of HPC applications. Jake Tronge, Jieyang Chen, Patricia Grubel, Tim Randles, Rusty Davis, Quincy Wofford, Steven Anaya, Qiang Guan |
ASE | 8 |
| 2020 | Chaser: An Enhanced Fault Injection Tool for Tracing Soft Errors in MPI ApplicationsabstractResilient computation has been an emerging topic in the field of high-performance computing (HPC). In particular, studies show that tolerating faults on leadership-class supercomputers (such as exascale supercomputers) is expected to be one of the main challenges. In this paper, we utilize dynamic binary instrumentation and virtual machine based fault injection to emulate soft errors and study the soft errors' impact on the behavior of applications. We propose Chaser, a fine-grained, accountable, flexible, and efficient fault injection framework built on top of QEMU. Chaser offers just-in-time fault injection, the ability to trace fault propagation, and flexible and programable interfaces. In the case study, we demonstrate the usage of Chaser on Matvec and a real DOE mini MPI application Qiang Guan, Xunchao Hu, Terence Grove, Bo Fang 0002, Hailong Jiang, Heng Yin 0001, Nathan DeBardeleben |
DSN | 1 |
| 2019 | TSM2: optimizing tall-and-skinny matrix-matrix multiplication on GPUsabstractLinear algebra operations have been widely used in big data analytics and scientific computations. Many works have been done on optimizing linear algebra operations on GPUs with regular-shaped input. However, few works are focusing on fully utilizing GPU resources when the input is not regular-shaped. Current optimizations lack of considering fully utilizing the memory bandwidth and computing power, therefore they could only achieve sub-optimal performance. In this paper, we propose a performant tall-and-skinny matrix-matrix multiplication algorithm on GPUs - TSM2. It focuses on optimizing linear algebra operation with none regular-shaped input. We implement the proposed algorithm and test on three different Nvidia GPU micro-architectures: Kepler, Maxwell, and Pascal. Experiments show that our TSM2 speedups the computation by 1.1x - 3x, improves memory bandwidth utilization by 8% - 47.6%, and improves computing power utilization by 7% - 37.3% comparing to the current state-of-the-art works. We replace the original matrix operations in K-means and Algorithm-Bases Fault Tolerance (ABFT) with TSM2 and achieve up to 1.89x and 1.90x speed up. Jieyang Chen, Nan Xiong, Xin Liang 0001, Dingwen Tao, Sihuan Li, Kaiming Ouyang, Kai Zhao 0008, Nathan DeBardeleben, Qiang Guan, Zizhong Chen |
ICS | 9 |
| 2019 | CARE: compiler-assisted recovery from soft failuresabstractAs processors continue to boost the system performance with higher circuit density, shrinking process technology and near-threshold voltage (NTV) operations, they are projected to be more vulnerable to transient faults, which have become one of the major concerns for future extreme-scale HPC systems. Despite being relatively infrequent, crashes due to transient faults are incredibly disruptive, particularly for massively parallel jobs on supercomputers where they potentially kill the entire job, requiring an expensive rerun or restart from a checkpoint. Chao Chen 0024, Greg Eisenhauer, Santosh Pande, Qiang Guan |
SC | 4 |
| 2018 | Build and Execution Environment (BEE): an Encapsulated Environment Enabling HPC Applications Running EverywhereabstractVariations in High Performance Computing (HPC) system software configurations mean that applications are typically configured and built for specific HPC environments. Building applications can require a significant investment of time and effort for application users and requires application users to have additional technical knowledge. Linux container technologies such as Docker and Charliecloud bring great benefits to the application development, build and deployment processes. While cloud platforms already widely support containers, HPC systems still have non-uniform support of container technologies. In this work, we propose a unified runtime framework - Build and Execution Environment (BEE) across both HPC and cloud platforms that allows users to run their containerized HPC applications across all supported platforms without modification. We design four BEE backends for four different classes of HPC or cloud platform so that together they cover the majority of mainstream computing platforms for HPC users. Evaluations show that BEE provides an easy-to-use unified user interface, execution environment, and comparable performance. Jieyang Chen, Qiang Guan, Xin Liang 0001, Paul Bryant, Patricia Grubel, Allen McPherson, Li-Ta Lo, Tim Randles, Zizhong Chen, James P. Ahrens |
IEEE BigData | 2 |
| 2018 | In situ TensorView: In situ Visualization of Convolutional Neural NetworksabstractConvolutional Neural Networks(CNNs) are complex systems trained to recognize images, texts and more. However, once trained, they are regarded as black-boxes that are not easy to analyze and understand. Visualizing the dynamics within such deep artificial neural networks can provide a better understanding of how they are learning and making predictions. In the field of scientific simulations, visualization tools like Paraview have long been utilized to provide insights. We present in situ TensorView to visualize the training and functioning of CNNs as if they are systems of scientific simulations. In situ TensorView is a loosely coupled in situ visualization open framework that provides multiple viewers with the ability to visualize and understand their networks. It leverages the capability of co-processing from Paraview to provide real-time visualization during training and predicting phases, and avoids heavy I/O overhead. Tensorview is easily coupled with Tensorflow, as it only requires the insertion of a few lines of code into a TensorFlow framework. In this work, we showcase visualizing LeNet-5 and VGG16 using in situ TensorView. With the insight provided by Tensorview, users can adjust network architectures, or compress pre-trained networks guided by visualization results. Qiang Guan, Li-Ta Lo, Simon Su, Zhengyong Ren, James P. Ahrens, Trilce Estrada |
IEEE BigData | 2 |
| 2018 | BeeFlow: A Workflow Management System for In Situ Processing across HPC and Cloud SystemsabstractIn this paper, we propose BeeFlow - an in situ analysis enabled workflow management system across multiple platforms using Docker containers. BeeFlow can support both traditional workflows as well as workflows with in situ analysis. BeeFlow leverages Docker containers to provide a portable, flexible, and reproducible workflow management system across HPC and cloud platforms. We showcase how current in situ visualization workflows can apply BeeFlow with DOE production codes VPIC and Flecsale. Jieyang Chen, Qiang Guan, Zhao Zhang 0007, Xin Liang 0001, Louis James Vernon, Allen McPherson, Li-Ta Lo, Patricia Grubel, Tim Randles, Zizhong Chen, James P. Ahrens |
ICDCS | 2 |
| 2018 | Modeling Application Resilience in Large-scale Parallel ExecutionabstractUnderstanding how the application is resilient to hardware and software errors is critical to high-performance computing. To evaluate application resilience, the application level fault injection is the most common method. However, the application level fault injection can be very expensive when running the application in parallel in large scales due to the high requirement for hardware resource during fault injection. Kai Wu 0006, Wenqian Dong, Qiang Guan, Nathan DeBardeleben, Dong Li 0001 |
ICPP | 3 |
| 2018 | Fault tolerant one-sided matrix decompositions on heterogeneous systems with GPUs
Jieyang Chen, Hongbo Li 0006, Sihuan Li, Xin Liang 0001, Panruo Wu, Dingwen Tao, Kaiming Ouyang, Yuanlai Liu, Kai Zhao 0008, Qiang Guan, Zizhong Chen |
SC | 10 |
| 2018 | Using virtualization to quantify power conservation via near-threshold voltage reduction for inherently resilient applications
Nathan DeBardeleben, Qiang Guan, Sean Blanchard, Michael Lang 0003 |
Parallel Comput. | 3 |
| 2017 | LetGo: A Lightweight Continuous Framework for HPC Applications Under FailuresabstractRequirements for reliability, low power consumption, and performance place complex and conflicting demands on the design of high-performance computing (HPC) systems. Fault-tolerance techniques such as checkpoint/restart (C/R) protect HPC applications against hardware faults. These techniques, however, have non negligible overheads particularly when the fault rate exposed by the hardware is high: it is estimated that in future HPC systems, up to 60% of the computational cycles/power will be used for fault tolerance. Bo Fang 0002, Qiang Guan, Nathan DeBardeleben, Karthik Pattabiraman, Matei Ripeanu |
HPDC | 2 |
| 2017 | Silent Data Corruption Resilient Two-sided Matrix FactorizationsabstractThis paper presents an algorithm based fault tolerance method to harden three two-sided matrix factorizations against soft errors: reduction to Hessenberg form, tridiagonal form, and bidiagonal form. These two sided factorizations are usually the prerequisites to computing eigenvalues/eigenvectors and singular value decomposition. Algorithm based fault tolerance has been shown to work on three main one-sided matrix factorizations: LU, Cholesky, and QR, but extending it to cover two sided factorizations is non-trivial because there are no obvious \textit{offline, problem} specific maintenance of checksums. We thus develop an \textit{online, algorithm} specific checksum scheme and show how to systematically adapt the two sided factorization algorithms used in LAPACK and ScaLAPACK packages to introduce the algorithm based fault tolerance. Panruo Wu, Nathan DeBardeleben, Qiang Guan, Sean Blanchard, Jieyang Chen, Dingwen Tao, Xin Liang 0001, Kaiming Ouyang, Zizhong Chen |
PPoPP | 3 |
| 2016 | Towards Practical Algorithm Based Fault Tolerance in Dense Linear AlgebraabstractAlgorithm based fault tolerance (ABFT) attracts renewed interest for its extremely low overhead and good scalability. However the fault model used to design ABFT has been either abstract, simplistic, or both, leaving a gap between what occurs at the architecture level and what the algorithm expects. As the fault model is the deciding factor in choosing an effective checksum scheme, the resulting ABFT techniques have seen limited impact in practice. In this paper we seek to close the gap by directly using a comprehensive architectural fault model and devise a comprehensive ABFT scheme that can tolerate multiple architectural faults of various kinds. We implement the new ABFT scheme into high performance linpack (HPL) to demonstrate the feasibility in large scale high performance benchmark. We conduct architectural fault injection experiments and large scale experiments to empirically validate its fault tolerance and demonstrate the overhead of error handling, respectively. Panruo Wu, Qiang Guan, Nathan DeBardeleben, Sean Blanchard, Dingwen Tao, Xin Liang 0001, Jieyang Chen, Zizhong Chen |
HPDC | 2 |
| 2015 | Towards Building Resilient Scientific Applications: Resilience Analysis on the Impact of Soft Error and Transient Error Tolerance with the CLAMR Hydrodynamics Mini-AppabstractIn this paper, we present a resilience analysis of the impact of soft errors on CLAMR, a hydrodynamics miniapp for high performance computing (HPC). Leveraging the conservation of mass law, we design a fault detection mechanism and checkpoint/restart fault tolerance approach to enhance the resilience of CLAMR. Overall, our approach can detect up to 88.3% of faults that propagate into SDC or crashes with minimal (less than 1%) overhead for the optimal configuration. We show that CLAMR's fault-tolerance depends on when a fault is injected into the simulation and we also evaluate the frequency of detection and checkpointing on performance. Qiang Guan, Nathan DeBardeleben, Brian Atkinson, Robert W. Robey, William M. Jones |
CLUSTER | 1 |
| 2015 | Differentiated Failure Remediation with Action Selection for Resilient ComputingabstractAs the fault frequency is increasing with the component count in modern and future computer systems, resilience becomes increasingly critical. Existing work on anomaly detection and fault prediction enables failure avoidance techniques to circumvent fault effects proactively. In addition, traditional fault tolerance techniques can be applied to handle faults reactively. Different types of faults may affect different components of a system and have various manifestations. They need to be treated differently. However, the existing fault handling techniques uniformly treat all faults without considering their types and distinct properties. In this paper, we present a differentiated fault remediation framework with action selection (DFRAS) which integrates both preventive and reactive remediation actions differentiated for different types of faults with their urgency requirements. We investigate four major types of faults and identify candidate remediation actions. We apply the urgency requirements as constraints for action selection. We propose formal performance models to quantify the wasted time of the candidate actions, and develop a decision making method to select the best actions that minimize the overall remediation cost. We have implemented a prototype of DFRAS and evaluated its performance by simulations and experiments. Simulation and experimental results show that the integrated fault remediation strategies can significantly reduce the remediation overhead. The developed DFRAS system is lightweight, making it feasible for online fault management in large-scale systems. Song Fu, Nathan DeBardeleben, Qiang Guan, Cheng-Zhong Xu 0001 |
PRDC | 4 |
| 2014 | Optimal self-learning battery control in smart residential grids by iterative Q-learning algorithmabstractIn this paper, a novel dual iterative Q-learning algorithm is developed to solve the optimal battery management and control problems in smart residential environments. The main idea is to use adaptive dynamic programming (ADP) technique to obtain the optimal battery management and control scheme iteratively for residential energy systems. In the developed dual iterative Q-learning algorithm, two iterations, including external and internal iterations, are introduced, where internal iteration minimizes the total cost of power loads in each period and the external iteration makes the iterative Q function converge to the optimum. For the first time, the convergence property of iterative Q-learning method is proven to guarantee the convergence property of the iterative Q function. Finally, numerical results are given to illustrate the performance of the developed algorithm. Qinglai Wei, Derong Liu 0001, Yu Liu 0078, Qiang Guan |
ADPRL | 5 |
| 2014 | A hierarchical classification algorithm for evaluating energy consumption behaviorsabstractResearches on office building energy consumption have been hot in these years, but few researchers consider the classification of office energy consumption performance which can evaluate user behaviors in order to offer a clear analysis of energy consumption and improve their energy saving consciousness. In this paper, we propose a novel hierarchical classification algorithm for evaluating energy consumption behaviors at a real energy management system, which combines fuzzy c-means clustering with GA (genetic algorithm)-based SVM (support vector machine) to fully utilize collected samples. The experiment results with real energy consumption data show that the proposed algorithm works well to distinguish the abnormal behaviors and classify energy consumption behaviors accurately on normal offices. Li Bu, Dongbin Zhao, Yu Liu 0005, Qiang Guan |
IJCNN | 4 |
| 2014 | F-SEFI: A Fine-Grained Soft Error Fault Injection Tool for Profiling Application VulnerabilityabstractAs the high performance computing (HPC) community continues to push towards exascale computing, resilience remains a serious challenge. With the expected decrease of both feature size and operating voltage, we expect a significant increase in hardware soft errors. HPC applications of today are only affected by soft errors to a small degree but we expect that this will become a more serious issue as HPC systems grow. We propose F-SEFI, a Fine-grained Soft Error Fault Injector, as a tool for profiling software robustness against soft errors. In this paper we utilize soft error injection to mimic the impact of errors on logic circuit behavior. Leveraging the open source virtual machine hypervisor QEMU, F-SEFI enables users to modify emulated machine instructions to introduce soft errors. F-SEFI can control what application, which sub-function, when and how to inject soft errors with different granularities, without interference to other applications that share the same environment. F-SEFI does this without requiring revisions to the application source code, compilers or operating systems. We discuss the design constraints for F-SEFI and the specifics of our implementation. We demonstrate use cases of F-SEFI on several benchmark applications to show how data corruption can propagate to incorrect results. Qiang Guan, Nathan DeBardeleben, Sean Blanchard, Song Fu |
IPDPS | 1 |
| 2013 | Wavelet-based multi-scale anomaly identification in cloud computing systemsabstractModern cloud computing systems contain thousands of computing and storage servers. Such a scale combined with ever-growing system complexity of their components and interactions, introduces a key challenge to failure and resource management for highly dependable cloud computing. Automated anomaly detection is a crucial technique for understanding emergent, cloud-wide phenomena and self-managing cloud resources for system level dependability assurance. In this paper, we present a wavelet-based multi-scale anomaly identification mechanism, that can analyze profiled cloud performance metrics in both time and frequency domains and identify anomalous cloud behaviors. Learning technologies are exploited to adapt the selection of mother wavelets and a sliding detection window is employed handle cloud dynamicity and improve anomaly detection accuracy. We test a prototype implementation of our cloud anomaly detection mechanism on an institute-wide cloud system. Experimental results show our approach can identify cloud failures accurately. Qiang Guan, Song Fu |
GLOBECOM | 1 |
| 2013 | Deadline-Aware Event Scheduling for Complex Event Processing Systems
Qiang Guan |
IDEAL | 2 |
| 2013 | Exploring Time and Frequency Domains for Accurate and Automated Anomaly Detection in Cloud Computing SystemsabstractCloud computing has become increasingly popular by obviating the need for users to own and maintain complex computing infrastructures. However, due to their inherent complexity and large scale, production cloud computing systems are prone to various runtime problems caused by hardware and software faults and environmental factors. Autonomic anomaly detection is crucial for understanding emergent, cloud-wide phenomena and self-managing cloud resources for system-level dependability assurance. To detect anomalous cloud behaviors, we need to monitor the cloud execution and collect runtime cloud performance data. For different types of failures, the data display different correlations with the performance metrics. In this paper, we present a wavelet-based multi-scale anomaly identification mechanism, that can analyze profiled cloud performance metrics in both time and frequency domains and identify anomalous cloud behaviors. Learning technologies are exploited to adapt the selection of mother wavelets and a sliding detection window is employed to handle cloud dynamicity and improve anomaly detection accuracy. We have implemented a prototype of the anomaly identification system and conducted experiments on an on-campus cloud computing environment. Experimental results show the proposed mechanism can achieve 93.3% detection sensitivity while keeping the false positive rate as low as 6.1% while outperforming other tested anomaly detection schemes. Qiang Guan, Song Fu, Nathan DeBardeleben, Sean Blanchard |
PRDC | 1 |
| 2013 | Adaptive Anomaly Identification by Exploring Metric Subspace in Cloud Computing InfrastructuresabstractCloud computing has become increasingly popular by obviating the need for users to own and maintain complex computing infrastructures. However, due to their inherent complexity and large scale, production cloud computing systems are prone to various runtime problems caused by hardware and software faults and environmental factors. Autonomic anomaly detection is a crucial technique for understanding emergent, cloud-wide phenomena and self-managing cloud resources for system-level dependability assurance. To detect anomalous cloud behaviors, we need to monitor the cloud execution and collect runtime cloud performance data. These data consist of values of performance metrics for different types of failures, which display different correlations with the performance metrics. In this paper, we present an adaptive anomaly identification mechanism that explores the most relevant principal components of different failure types in cloud computing infrastructures. It integrates the cloud performance metric analysis with filtering techniques to achieve automated, efficient, and accurate anomaly identification. The proposed mechanism adapts itself by recursively learning from the newly verified detection results to refine future detections. We have implemented a prototype of the anomaly identification system and conducted experiments in an on-campus cloud computing environment and by using the Google data center traces. Our experimental results show that our mechanism can achieve more efficient and accurate anomaly detection than other existing schemes. Qiang Guan, Song Fu |
SRDS | 1 |
| 2012 | AFD: Adaptive failure detection system for cloud computing infrastructuresabstractCloud computing has become increasingly popular by obviating the need for users to own and maintain complex computing infrastructure. However, due to their inherent complexity and large scale, production cloud computing systems are prone to various runtime problems caused by hardware and software failures. Autonomic failure detection is a crucial technique for understanding emergent, cloud-wide phenomena and self-managing cloud resources for system-level dependability assurance. To detect failures, we need to monitor the cloud execution and collect runtime performance data. These data are usually unlabeled, and thus a prior failure history is not always available in production clouds, especially for newly managed or deployed systems. In this paper, we present an Adaptive Failure Detection (AFD) framework for cloud dependability assurance. AFD employs data description using hypersphere for adaptive failure detection. Based on the cloud performance data, AFD detects possible failures, which are verified by the cloud operators. They are confirmed as either true failures with failure types or normal states. AFD adapts itself by recursively learning from these newly verified detection results to refine future detections. Meanwhile, AFD exploits the observed but undetected failure records reported by the cloud operators to identify new types of failures. We have implemented a prototype of the AFD system and conducted experiments in an on-campus cloud computing environment. Our experimental results show that AFD can achieve more efficient and accurate failure detection than other existing schemes. Husanbir Singh Pannu, Jianguo Liu 0001, Qiang Guan, Song Fu |
IPCCC | 3 |
| 2012 | An adaptive power management framework for autonomic resource configuration in cloud computing infrastructuresabstractPower is becoming an increasingly important concern for large-scale cloud computing systems. Meanwhile, cloud service providers leverage virtualization technologies to facilitate service consolidation and enhance resource utilization. However, the introduction of virtualization makes the cloud infrastructure more complex, and thus challenges cloud power management. In a virtualized environment, resource needs to be configured at runtime at the cloud, server and virtual machine levels to achieve high power efficiency. In addition, cloud power management should guarantee high users' SLA (service level agreement) satisfaction. In this paper, we present an adaptive power management framework in the cloud to achieve autonomic resource configuration. We propose a software and lightweight approach to accurately estimate the power usage of virtual machines and cloud servers. It explores hypervisor-observable performance metrics to build the power usage model. To configure cloud resources, we consider both the system power usage and the SLA requirements, and leverage learning techniques to achieve autonomic resource allocation and optimal power efficiency. We implement a prototype of the proposed power management system and test it on a cloud testbed. Experimental results show the high accuracy (over 90%) of our power usage estimation mechanism and our resource configuration approach achieves the lowest energy usage among the compared four approaches. Qiang Guan, Song Fu |
IPCCC | 2 |
| 2012 | Efficient and Accurate Anomaly Identification Using Reduced Metric Space in Utility CloudsabstractThe online detection of anomalies is a vital element of operations in utility clouds. Detection should function for different levels of abstraction including hardware and software, and for the various metrics used in cloud computing systems. Given ever-increasing cloud sizes coupled with the complexity of system components, continuous monitoring leads to the overwhelming volume of data collected by health monitoring tools. High metric dimensionality and existence of interacting metrics compromise the detection accuracy and lead to high detection complexity. In this paper, we present a metric selection framework and propose systematic approaches to effectively identify and select the most essential metrics for online anomaly detection in utility clouds. Specifically, a mutual information based approach selects metrics with the maximized mutual relevance and the minimized redundancy. Then metric space combination and separation are explored to reduce the metric dimensionality further. Experimental results on utility cloud scenarios demonstrate the viability and efficiency of this framework. The selected metrics contribute to a high efficiency and accuracy in anomaly detection. Qiang Guan, Chi-Chen Chiu, Song Fu |
NAS | 1 |
| 2012 | CDA: A Cloud Dependability Analysis Framework for Characterizing System Dependability in Cloud Computing InfrastructuresabstractCloud computing has become increasingly popular by obviating the need for users to own and maintain complex computing infrastructure. However, due to their inherent complexity and large scale, production cloud computing systems are prone to various runtime problems caused by hardware and software failures. Dependability assurance is crucial for building sustainable cloud computing services. Although many techniques have been proposed to analyze and enhance reliability of distributed systems, there is little work on understanding the dependability of cloud computing environments. As virtualization has been an enabling technology for the cloud, it is imperative to investigate the impact of virtualization on the cloud dependability, which is the focus of this work. In this paper, we present a cloud dependability analysis (CDA) framework with mechanisms to characterize failure behavior in cloud computing infrastructures. We design the failure-metric DAGs (directed a cyclic graph) to analyze the correlation of various performance metrics with failure events in virtualized and non-virtualized systems. We study multiple types of failures. By comparing the generated DAGs in the two environments, we gain insight into the impact of virtualization on the cloud dependability. This paper is the first attempt to study this crucial issue. In addition, we exploit the identified metrics for failure detection. Experimental results from an on-campus cloud computing test bed show that our approach can achieve high detection accuracy while using a small number of performance metrics. Qiang Guan, Chi-Chen Chiu, Song Fu |
PRDC | 1 |
| 2011 | Proactive Failure Management by Integrated Unsupervised and Semi-Supervised Learning for Dependable Cloud SystemsabstractCloud computing systems continue to grow in their scale and complexity. They are changing dynamically as well due to the addition and removal of system components, changing execution environments, frequent updates and upgrades, online repairs and more. In such large-scale complex and dynamic systems, failures are common. In this paper, we present a failure prediction mechanism exploiting both unsupervised and semi-supervised learning techniques for building dependable cloud computing systems. The unsupervised failure detection method uses an ensemble of Bayesian models. It characterizes normal execution states of the system and detects anomalous behaviors. After the anomalies are verified by system administrators, labeled data are available. Then, we apply supervised learning based on decision tree classier to predict future failure occurrences in the cloud. Experimental results in an institute-wide cloud computing system show that our proposed method can forecast failure dynamics with high accuracy. Qiang Guan, Song Fu |
ARES | 1 |
| 2011 | Ensemble of Bayesian Predictors for Autonomic Failure Management in Cloud ComputingabstractIn modern cloud computing systems, hundreds and even thousands of cloud servers are interconnected by multi-layer networks. In such large-scale and complex systems, failures are common. Proactive failure management is a crucial technology to characterize system behaviors and forecast failure dynamics in the cloud. To make failure predictions, we need to monitor the system execution and collect health-related runtime performance data. However, in newly deployed or managed cloud systems, these data are usually unlabeled. Supervised learning based approaches are not suitable in this case. In this paper, we present an unsupervised failure detection method using an ensemble of Bayesian models. It estimates the probability distribution of runtime performance data collected by health monitoring tools when cloud servers perform normally. It characterizes normal execution states of the system and detects anomalous behaviors. Experimental results in an institute-wide cloud computing system show that our methods can achieve high true positive rate and low false positive rate for proactive failure management. Qiang Guan, Song Fu |
ICCCN | 1 |
| 2010 | Anomaly detection in large-scale coalition clusters for dependability assuranceabstractIn large-scale high-performance computing systems, component failures become norms instead of exceptions. Failure occurrence as well as its impact on system performance and operation costs are becoming an increasingly important concern to system designers and administrators. When a compute node fails to function properly, health-related data are valuable for troubleshooting. However, it is challenging to effectively identify anomalies from the voluminous amount of noisy, high-dimensional data. Manual detection is time-consuming and error-prone. It does not scale well. In this paper, we present an autonomic mechanism for anomaly detection in coalition clusters. It is composed of a set of techniques that facilitates automatic analysis of system health data. We apply data transformation to format health data in a uniform manner. Then principal variables are chosen by feature selection, which reduces the data size. Clustering and outlier detection are explored to identify nodes with anomalous behavior. We evaluate our prototype implementation on a production institution-wide computational grid. The results show that our mechanism can effectively detect faulty nodes with high accuracy and low computation overhead. Qiang Guan, Derek Smith, Song Fu |
HiPC | 1 |
| 2010 | auto-AID: A data mining framework for autonomic anomaly identification in networked computer systemsabstractNetworked computer systems continue to grow in scale and in the complexity of their components and interactions. Component failures become norms instead of exceptions in these environments. A failure will cause one or multiple computer(s) to be unavailable, which affects the resource utilization and system throughput. When a computer fails to function properly, health-related data are valuable for troubleshooting. However, it is challenging to effectively identify anomalies from the voluminous amount of noisy, high-dimensional data. In this paper, we present auto-AID, an autonomic mechanism for anomaly identification in networked computer systems. It is composed of a set of data mining techniques that facilitates automatic analysis of system health data. The identification results are very valuable for the system administrators to manage systems and schedule the available resources. We implement a prototype of auto-AID and evaluate it on a production institution-wide compute grid. The results show that auto-AID can effectively identify anomalies with little human intervention. Qiang Guan, Song Fu |
IPCCC | 1 |
| 2007 | Solution to the Generalized Champagne Problem on simultaneous stabilization of linear systems
Qiang Guan, Long Wang 0001, Bican Xia, Wensheng Yu, Zhenbing Zeng |
Sci. China Ser. F Inf. Sci. | 1 |