EDBT 2026 Demo / reviewers in the wild / expert
Bin Nie
dblp:68/6005
· DBLP profile ↗
13ranked-venue papers
7as first author
3since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 6 first-author · 1 since 2021Security and privacy · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Hardware reliability and fault tolerance · 69% GPUs and heterogeneous computing · 26% High-performance computing · 3% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware reliability and fault tolerance › soft errors
soft error resilience |
1.0 | 2 | 2021 | Practical Resilience Analysis of GPGPU Applications in the Presence of Single- and Multi-Bit Faults · IEEE Trans. Computers 2021 Enabling Software Resilience in GPGPU Applications via Partial Thread Protection · ICSE 2021 |
Hardware reliability and fault tolerance
soft errors |
0.6 | 2 | 2018 | Fault Site Pruning for Practical Reliability Analysis of GPGPU Applications · MICRO 2018 A large-scale study of soft-errors on GPUs in the field · HPCA 2016 |
GPUs and heterogeneous computing
GPU reliability |
0.6 | 2 | 2021 | Enabling Software Resilience in GPGPU Applications via Partial Thread Protection · ICSE 2021 A large-scale study of soft-errors on GPUs in the field · HPCA 2016 |
Hardware reliability and fault tolerance
error resilience |
0.3 | 1 | 2018 | Fault Site Pruning for Practical Reliability Analysis of GPGPU Applications · MICRO 2018 |
Hardware reliability and fault tolerance
fault injection |
0.3 | 1 | 2018 | Fault Site Pruning for Practical Reliability Analysis of GPGPU Applications · MICRO 2018 |
GPUs and heterogeneous computing
GPU computing |
0.2 | 2 | 2021 | Practical Resilience Analysis of GPGPU Applications in the Presence of Single- and Multi-Bit Faults · IEEE Trans. Computers 2021 Fault Site Pruning for Practical Reliability Analysis of GPGPU Applications · MICRO 2018 |
High-performance computing › system resilience
application resilience |
0.1 | 1 | 2018 | Fault Site Pruning for Practical Reliability Analysis of GPGPU Applications · MICRO 2018 |
Performance modeling and evaluation
field data analysis |
0.1 | 1 | 2016 | A large-scale study of soft-errors on GPUs in the field · HPCA 2016 |
Methods — techniques the papers use, named apart from their topics
fault injection · 0.8thread remapping · 0.5partial replication · 0.5fault site pruning · 0.5dynamic instruction analysis · 0.3large-scale field data analysis · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Research on drug-drug interaction prediction using capsule neural network based on self-attention mechanismabstractBACKGROUND: Multi-drug combinations represent an effective strategy for treating complex diseases. However, due to the vast number of unknown interactions among drugs, accurately predicting drug-drug interactions (DDIs) is essential for preventing adverse drug reactions that may cause serious harm to patients. Therefore, DDI prediction plays a critical role in pharmacology. RESULTS: In this paper, we propose a novel DDI prediction model that integrates a self-attention mechanism with a capsule neural network, termed ACaps-DDI. The model effectively combines chemical information from internal drug substructures with biological information from external drug targets and drug-metabolizing enzymes to predict potential drug-drug interactions. CONCLUSIONS: Experimental results on two benchmark datasets show that the ACaps-DDI model outperforms six other classification models across seven evaluation metrics, demonstrating its strong predictive performance and generalization ability. Ablation studies further confirm the effectiveness of individual components within the ACaps-DDI architecture. Finally, case studies involving three drugs (cannabidiol, torasemide, and cyclophosphamide) validate the model's ability to predict previously unknown drug interactions. In conclusion, the ACaps-DDI model exhibits high predictive accuracy for known drugs and demonstrates promising predictive capability for unseen drugs, highlighting its practical significance for clinical research on drug interactions. Xingxin Chen, Zhen Miao, Bin Nie |
BMC Bioinform. | 4 |
| 2021 | Enabling Software Resilience in GPGPU Applications via Partial Thread ProtectionabstractGraphics Processing Units (GPUs) are widely used by various applications in a broad variety of fields to accelerate their computation but remain susceptible to transient hardware faults (soft errors) that can easily compromise application output. By taking advantage of a general purpose GPU application hierarchical organization in threads, warps, and cooperative thread arrays, we propose a methodology that identifies the resilience of threads and aims to map threads with the same resilience characteristics to the same warp. This allows to engage partial replication mechanisms for error detection/correction at the warp level. By exploring 12 benchmarks (17 kernels) from 4 benchmark suites, we illustrate that threads can be remapped into reliable or unreliable warps with only 1.63% introduced overhead (on average), and then enable selective protection via replication to those groups of threads that truly need it. Furthermore, we show that thread remapping to different warps does not sacrifice application performance. We show how this remapping facilitates warp replication for error detection and/or correction and achieves average reduction of 20.61% and 27.15% execution cycles, respectively comparing to standard duplication/triplication. Lishan Yang 0001, Bin Nie, Adwait Jog, Evgenia Smirni |
ICSE | 2 |
| 2021 | Practical Resilience Analysis of GPGPU Applications in the Presence of Single- and Multi-Bit FaultsabstractGraphics Processing Units (GPUs) have rapidly evolved to enable energy-efficient data-parallel computing for a broad range of scientific areas. While GPUs achieve exascale performance at a stringent power budget, they are also susceptible to soft errors, often caused by high-energy particle strikes, that can significantly affect the application output quality. Understanding the resilience of general purpose GPU (GPGPU) applications is especially challenging because unlike CPU applications, which are mostly single-threaded, GPGPU applications can contain hundreds to thousands of threads, resulting in a tremendously large fault site space in the order of billions, even for some simple applications and even when considering the occurrence of just a single-bit fault. We present a systematic way to progressively prune the fault site space aiming to dramatically reduce the number of fault injections such that assessment for GPGPU application error resilience becomes practical. The key insight behind our proposed methodology stems from the fact that while GPGPU applications spawn a lot of threads, many of them execute the same set of instructions. Therefore, several fault sites are redundant and can be pruned by careful analysis. We identify important features across a set of 10 applications (16 kernels) from Rodinia and Polybench suites and conclude that threads can be primarily classified based on the number of the dynamic instructions they execute. We therefore achieve significant fault site reduction by analyzing only a small subset of threads that are representative of the dynamic instruction behavior (and therefore error resilience behavior) of the GPGPU applications. Further pruning is achieved by identifying the dynamic instruction commonalities (and differences) across code blocks within this representative set of threads, a subset of loop iterations within the representative threads, and a subset of destination register bit positions. The above steps result in a tremendous reduction of fault sites by up to seven orders of magnitude. Yet, this reduced fault site space accurately captures the error resilience profile of GPGPU applications. We show the effectiveness of the proposed progressive pruning technique for a single-bit model and illustrate its application to even more challenging cases with three distinct multi-bit fault models. Lishan Yang 0001, Bin Nie, Adwait Jog, Evgenia Smirni |
IEEE Trans. Computers | 2 |
| 2020 | Characterizing Accuracy-Aware Resilience of GPGPU ApplicationsabstractGraphics Processing Units (GPUs) have rapidly evolved to enable energy-efficient data-parallel computing. In addition to achieving exascale performance at a stringent power budget, it is imperative for GPUs to provide reliable computing guarantees to the end user. In current commodity systems, such guarantees are often achieved by incurring high protection cost in terms of performance, power, and hardware resources. However, we argue that these strict guarantees are often not required (and that the associated protected overheads can be significantly reduced) because several GPGPU applications are either fault-tolerant or can accept a quantifiable loss in output quality. To this end, this paper characterizes in a hierarchical manner the accuracy-aware resilience of GPGPU applications consisting of thousands of threads. This characterization study shows that accuracy-aware error resilience exhibits several interesting patterns across threads at different hierarchies (i.e., kernel/thread-block/warp). The insights from this characterization study can be used to reduce the overheads of expensive protection or recovery mechanisms that are typically used by GPUs to ensure application reliability. Bin Nie, Adwait Jog, Evgenia Smirni |
CCGRID | 1 |
| 2020 | Mining Multivariate Discrete Event Sequences for Knowledge Discovery and Anomaly DetectionabstractModern physical systems deploy large numbers of sensors to record at different time-stamps the status of different systems components via measurements such as temperature, pressure, speed, but also the component's categorical state. Depending on the measurement values, there are two kinds of sequences: continuous and discrete. For continuous sequences, there is a host of state-of-the-art algorithms for anomaly detection based on time-series analysis, but there is a lack of effective methodologies that are tailored specifically to discrete event sequences. This paper proposes an analytics framework for discrete event sequences for knowledge discovery and anomaly detection. During the training phase, the framework extracts pairwise relationships among discrete event sequences using a neural machine translation model by viewing each discrete event sequence as a "natural language". The relationship between sequences is quantified by how well one discrete event sequence is "translated" into another sequence. These pairwise relationships among sequences are aggregated into a multivariate relationship graph that clusters the structural knowledge of the underlying system and essentially discovers the hidden relationships among discrete sequences. This graph quantifies system behavior during normal operation. During testing, if one or more pairwise relationships are violated, an anomaly is detected. The proposed framework is evaluated on two real-world datasets: a proprietary dataset collected from a physical plant where it is shown to be effective in extracting sensor pairwise relationships for knowledge discovery and anomaly detection, and a public hard disk drive dataset where its ability to effectively predict upcoming disk failures is illustrated. Bin Nie, Jianwu Xu, Jacob Alter, Evgenia Smirni |
DSN | 1 |
| 2018 | Machine Learning Models for GPU Error Prediction in a Large Scale HPC SystemabstractGPUs are widely deployed on large-scale HPC systems to provide powerful computational capability for scientific applications from various domains. As those applications are normally long-running, investigating the characteristics of GPU errors becomes imperative for reliability. In this paper, we first study the system conditions that trigger GPU errors using six-month trace data collected from a large-scale, operational HPC system. Then, we use machine learning to predict the occurrence of GPU errors, by taking advantage of temporal and spatial dependencies of the trace data. The resulting machine learning prediction framework is robust and accurate under different workloads. Bin Nie, Ji Xue, Saurabh Gupta 0002, Tirthak Patel, Christian Engelmann, Evgenia Smirni, Devesh Tiwari |
DSN | 1 |
| 2018 | Fault Site Pruning for Practical Reliability Analysis of GPGPU ApplicationsabstractGraphics Processing Units (GPUs) have rapidly evolved to enable energy-efficient data-parallel computing for a broad range of scientific areas. While GPUs achieve exascale performance at a stringent power budget, they are also susceptible to soft errors, often caused by high-energy particle strikes, that can significantly affect the application output quality. Understanding the resilience of general purpose GPU applications is the purpose of this study. To this end, it is imperative to explore the range of application output by injecting faults at all the potential fault sites. This problem is especially challenging because unlike CPU applications, which are mostly single-threaded, GPGPU applications can contain hundreds to thousands of threads, resulting in a tremendously large fault site space – in the order of billions even for some simple applications. In this paper, we present a systematic way to progressively prune the fault site space aiming to dramatically reduce the number of fault injections such that assessment for GPGPU application error resilience can be practical. The key insight behind our proposed methodology stems from the fact that GPGPU applications spawn a lot of threads, however, many of them execute the same set of instructions. Therefore, several fault sites are redundant and can be pruned by a careful analysis of faults across threads and instructions. We identify important features across a set of 10 applications (16 kernels) from Rodinia and Polybench suites and conclude that threads can be first classified based on the number of the dynamic instructions they execute. We achieve significant fault site reduction by analyzing only a small subset of threads that are representative of the dynamic instruction behavior (and therefore error resilience behavior) of the GPGPU applications. Further pruning is achieved by identifying and analyzing: a) the dynamic instruction commonalities (and differences) across code blocks within this representative set of threads, b) a subset of loop iterations within the representative threads, and c) a subset of destination register bit positions. The above steps result in a tremendous reduction of fault sites by up to seven orders of magnitude. Yet, this reduced fault site space accurately captures the error resilience profile of GPGPU applications. Bin Nie, Lishan Yang 0001, Adwait Jog, Evgenia Smirni |
MICRO | 1 |
| 2017 | Improved algorithm of C4.5 decision tree on the arithmetic average optimal selection classification attributeabstractTo try to decrease the preference of the attribute values for information gain and information gain ratio, in the paper, the authors puts forward a improved algorithm of C4.5 decision tree on the selection classification attribute. The basic thought of the algorithm is as follows: Firstly, computing the information gain of selection classification attribute, and then get an attribute of the information gain which is higher than the average level; Secondly, computing separately the arithmetic average value of the information gain ratio and information gain of the attribute, and then select the biggest attribute of the average value and set up a branch decision; Finally, to use recursive method to build a decision tree. The experiment shows that this method is applicable and effective. Bin Nie, Jigen Luo, Jianqiang Du, Ai Chen |
BIBM | 1 |
| 2017 | Fill-in the gaps: Spatial-temporal models for missing dataabstractEffective workload characterization and prediction are instrumental for efficiently and proactively managing large systems. System management primarily relies on the workload information provided by underlying system tracing mechanisms that record system-related events in log files. However, such tracing mechanisms may temporarily fail due to various reasons, yielding “holes” in data traces. This missing data phenomenon significantly impedes the effectiveness of data analysis. In this paper, we study real-world data traces collected from over 80K virtual machines (VMs) hosted on 6K physical boxes in the data centers of a service provider. We discover that the usage series of VMs co-located on the same physical box exhibit strong correlation with one another, and that most VM usage series show temporal patterns. By taking advantage of the observed spatial and temporal dependencies, we propose a data-filling method to predict the missing data in the VM usage series. Detailed evaluation using trace data in the wild shows that the proposed method is sufficiently accurate as it achieves an average of 20% absolute percentage errors. We also illustrate its usefulness via a use case. Ji Xue, Bin Nie, Evgenia Smirni |
CNSM | 2 |
| 2017 | Characterizing Temperature, Power, and Soft-Error Behaviors in Data Center Systems: Insights, Challenges, and OpportunitiesabstractGPUs have become part of the mainstream high performance computing facilities that increasingly require more computational power to simulate physical phenomena quickly and accurately. However, GPU nodes also consume significantly more power than traditional CPU nodes, and high power consumption introduces new system operation challenges, including increased temperature, power/cooling cost, and lower system reliability. This paper explores how power consumption and temperature characteristics affect reliability, provides insights into what are the implications of such understanding, and how to exploit these insights toward predicting GPU errors using neural networks. Bin Nie, Ji Xue, Saurabh Gupta 0002, Christian Engelmann, Evgenia Smirni, Devesh Tiwari |
MASCOTS | 1 |
| 2016 | A large-scale study of soft-errors on GPUs in the fieldabstractParallelism provided by the GPU architecture has enabled domain scientists to simulate physical phenomena at a much faster rate and finer granularity than what was previously possible by CPU-based large-scale clusters. Architecture researchers have been investigating reliability characteristics of GPUs and innovating techniques to increase the reliability of these emerging computing devices. Such efforts are often guided by technology projections and simplistic scientific kernels, and performed using architectural simulators and modeling tools. Lack of large-scale field data impedes the effectiveness of such efforts. This study attempts to bridge this gap by presenting a large-scale field data analysis of GPU reliability. We characterize and quantify different kinds of soft-errors on the Titan supercomputer's GPU nodes. Our study uncovers several interesting and previously unknown insights about the characteristics and impact of soft-errors. Bin Nie, Devesh Tiwari, Saurabh Gupta 0002, Evgenia Smirni, James H. Rogers |
HPCA | 1 |
| 2016 | Transformations and Soliton Solutions for a Variable-coefficient Nonlinear Schrödinger Equation in the Dispersion Decreasing Fiber with Symbolic ComputationabstractDescribing the dispersion decreasing fiber, a variable-coefficient nonlinear Schrödinger equation is hereby under investigation. Three transformations have been obtained from such a equation to the known standard and cylindrical nonlinear Schrödinger equations with the relevant constraints on the variable coefficients presented, which turn out to be more general than those previously published in the literature. Meanwhile, several families of exact dark-soliton-like and bright-soliton-like solutions are constructed. Also, we obtain some similarity solutions, which can be illustrated in terms of the elliptic and the second Painlevé transcendent equations. Zhi-Fang Zeng, Jian-Guo Liu 0007, Bin Nie |
Fundam. Informaticae | 4 |
| 2014 | Leveraging online social friendship to improve data swarming performance
Honggang Zhang 0003, Benyuan Liu, Bin Nie, Xiayin Weng |
Comput. Networks | 3 |