EDBT 2026 Demo / reviewers in the wild / expert
Yunfei Du 0001
dblp:64/5123-1
· DBLP profile ↗
41ranked-venue papers
1as first author
13since 2021 · last 2026
0000-0002-6541-2511ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 24 · 3 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Computer networks · 3 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Practical and Scalable RDMA Connection Sharing for HPC WorkloadabstractRDMA is a fundamental communication infrastructure in high-performance computing (HPC). However, as the number of RDMA connections increases, system performance rapidly declines and memory consumption increases sharply. Previous research demonstrates that sharing RDMA connections among processes is necessary and effective to address the scalability problem. Unfortunately, previous work shares connections in software, thus incurring substantial overhead to each packet operation, and fails to comprehensively explore control policies to achieve superior sharing decisions. Yuejie Wang, Tuo Fang, Biyu Peng, Xin Sun 0027, Chengchao Xu, Yuxin Ren 0001, Ning Jia 0004, Xinwei Hu, Yunfei Du 0001, Guyue Liu |
EuroSys | 11 |
| 2025 | CoffeeBoost: Gradient Boosting Native Conformal Inference for Bayesian OptimizationabstractBayesian optimization (BO) is a key technique for solving black-box optimization problems. This study extends the scope of BO from conventional applications (e.g., AutoML and robotics learning) to automated tuning of software systems. Despite GP (Gaussian Process) implementing a foundation formalism for exploitation and exploration in BO, its limited predictive power and unrealistic assumptions (e.g., continuity and Gaussianity) can severely affect its effectiveness and efficiency in tuning complex software systems. To overcome these limitations, we propose a BO framework CoffeeBoost, which implements exploitation and exploration with a GBDT-native distribution-free probabilistic surrogate model. CoffeeBoost constructs surrogate models via stochastic gradient boosting ensembles (SGBE) and quantifies probabilistic distributions via distribution-free conformal predictive systems. Moreover, CoffeeBoost leverages the residual paths in SGBE to improve the local adaptiveness of the resulting predictive distributions in a GBDT-native manner. Across eight auto-tuning benchmarks for database management systems (DBMS), we evaluate CoffeeBoost and show its superior learnability and optimizability against existing GP-based and tree-ensemble-based BO schemes. Detailed analysis further shows CoffeeBoost's predictive distributions excel in both coverage and tightness. Yuanhao Lai, Chenpeng Ji, Tingkai Wang, Songhan Zhang, Zhengang Wang, Yunfei Du 0001 |
AAAI | 8 |
| 2025 | Centrum: Model-based Database Auto-tuning with Minimal Distributional AssumptionsabstractGaussian Process (GP)-based Bayesian optimization (BO), i.e., GP-BO, emerges as a prevailing model-based framework for DBMS (Database Management System) auto-tuning. However, recent work shows GP-BO-based DBMS auto-tuners are significantly outperformed by auto-tuners based on SMAC, which features random forest surrogate models; such results motivate us to rethink and investigate the limitations of GP-BO in auto-tuner design. We find that the fundamental assumptions of GP-BO are widely violated when modeling and optimizing DBMS performance, while tree-ensemble-BOs (e.g., SMAC) can avoid the assumption pitfalls and deliver improved tuning efficiency and effectiveness. Moreover, we argue that existing tree-ensemble-BOs restrict further advancement in DBMS auto-tuning. First, existing tree-ensemble-BOs can only achieve distribution-free point estimates, but still impose unrealistic distributional assumptions on uncertainty (interval) estimates, which can compromise surrogate modeling and distort the acquisition function. Second, recent advances in (ensemble) gradient boosting, which can further enhance surrogate modeling against vanilla GP and random forest counterparts, have rarely been applied in optimizing DBMS auto-tuners. To address these issues, we propose a novel model-based DBMS auto-tuner, Centrum . Centrum achieves and improves distribution-free point and interval estimation in surrogate modeling with a two-phase learning procedure of stochastic gradient boosting ensembles (SGBE). Moreover, Centrum adopts a generalized SGBE-estimated locally-adaptive conformal prediction to facilitate a distribution-free interval (uncertainty) estimation and acquisition function. To our knowledge, Centrum is the first auto-tuner that realizes distribution-freeness to stress and enhance BO's practicality in DBMS auto-tuning, and the first to seamlessly fuse gradient boosting ensembles and conformal inference in BO. Extensive physical and simulation experiments on two DBMSs and three workloads show that Centrum outperforms 21 state-of-the-art (SOTA) DBMS auto-tuners based on BO with GP, random forest, gradient boosting, OOB (Out-Of-Bag) conformal ensemble and other surrogates, as well as that based on reinforcement learning and genetic algorithms. Yuanhao Lai, Chenpeng Ji, Yan Li 0139, Songhan Zhang, Rutao Zhang, Zhengang Wang, Yunfei Du 0001 |
Proc. ACM Manag. Data | 8 |
| 2023 | Evolution Strategies Enhanced Complex Multiagent CoordinationabstractMulti-agent coordination involves both the individual reward and team reward, where the former guides the agent to learn basic skills and the latter measures how well such a team cooperatively completes final tasks. However, in many complex scenarios, these two aspects can be contradictory, due to that one agent excessively pursuing its own profits may suppress the performance of other teammates and lead to the reduction of overall profits. Besides, such dual rewards are generally entangled which make the learning swing between optimizing either the former or latter, which further leads to the sub-optimal and unstable solutions. Moreover, the sparse reward problem commonly encountered in the multi-agent system would further exacerbate this contradiction. In the present work, we address these challenges by proposing CEMARL, a novel framework combining cross-entropy method (CEM) and off-policy multi-agent reinforcement learning (MARL). CEM is gradient-free and learns from the whole episode, whereas MARL is gradient-based and learns from the experiences of agents. The core idea behind CEMARL is that it explicitly decomposes the individual reward and team reward, and deals with them through gradient-based learning and gradient-free evolution, respectively. By means of this, it can simultaneously maximize the individual and team reward, and reconciles the contradiction between individual and team as well as the sparse reward problem. CEMARL shows both conciseness in framework and stability in training, and achieves significantly better performances than state-of-the-art baselines on a range of complex tasks. Yunfei Du 0001, Ya Cong, Shiliang Pu |
IJCNN | 1 |
| 2023 | LPV: A Log Parsing Framework Based on VectorizationabstractLogs are pervasive in modern computing systems, and are valuable to service and system management. Nevertheless, with the rapidly growing size and complexity of computing systems, the log volume is exploding, which makes automatic log analysis imperative. Generally, in automatic log analysis, the first and fundamental step is log parsing, to which a lot of effort has been devoted. However, in most existing log parsing methods, log messages are merely treated as plain text. In natural language processing (NLP) area, it is a common practice to represent words and sentences with vectors, then the similarity between two words or sentences can be measured by the distance between their vectors. Inspired by these, we put forward a novel log parsing framework, named LPV (LogParser based onVectorization), which performs log parsing by converting log messages and log templates into vectors, with the help of a vectorization method in NLP. LPV incorporates offline and online log parsing. In the offline log parsing, the central idea is to first represent log messages with vectors, so that the similarity between two log messages can be measured by the distance between their vectors, then we cluster log messages via clustering the vectors, and finally we extract log templates from the resultant clusters. By the end of the offline log parsing, each log template is assigned with an average vector, so that in the online log parsing, the similarity between an incoming log message and each log template can also be measured by the distance between their vectors. Extensive experiments have been conducted based on several public log datasets to evaluate LPV with three different vectorization methods. The results demonstrate that, with a proper vectorization method, LPV performs competitive with state-of-the-art log parsing methods, in both effectiveness and efficiency. Tong Xiao 0002, Zhe Quan, Zhi-Jie Wang 0009, Kaiqi Zhao 0001, Xiangke Liao, Yunfei Du 0001, Kenli Li 0001 |
IEEE Trans. Netw. Serv. Manag. | 7 |
| 2023 | Loader: A Log Anomaly Detector Based on TransformerabstractDetecting anomalies in logs is crucial for service and system management, since logs are widely used to record the runtime status, and are often the only data available for postmortem analysis. Since anomalies are usually rare in real-world services and systems, a common and feasible practice is to mine or learn normal patterns from logs, and deem those violating the normal patterns as anomalies. As log sequences are a kind of time series data, RNN (Recurrent Neural Network) and its variants have been extensively employed to capture the normal patterns. Nevertheless, the sequential nature of RNN and its variants makes them hard to parallelize and capture long-term dependencies, which may hinder their performance. To address this issue, in this paper we propose Loader, a novel semi-supervisedloganomalydetector based on Transformer, because the Transformer architecture eschews recurrence and is able to draw global dependencies. Loader leverages the Transformer encoder to capture normal patterns from normal log sequences. When detecting, it gives a set of candidate log templates, that may appear after the input log substring under normal conditions. If the template of the actual next log message is not within the candidate set, this implies an anomaly. Previous similar methods select the most possible$k$log templates as candidates in any case, so the performance is sensitive to$k$, and it is nontrivial to pick a proper$k$. To alleviate this, we design a more flexible and robust ‘top-$p$’ algorithm, which determines the candidate set based on the cumulative probability of the most possible log templates. Extensive experiments are conducted based on three public log datasets, the experimental results validate the effectiveness and competitiveness of our approach. Tong Xiao 0002, Zhe Quan, Zhi-Jie Wang 0009, Yuquan Le, Yunfei Du 0001, Xiangke Liao, Kenli Li 0001, Keqin Li 0001 |
IEEE Trans. Serv. Comput. | 5 |
| 2022 | Enhancing Distributed In-Situ CNN Inference in the Internet of ThingsabstractConvolutional neural networks (CNNS) enable machines to view the world as humans and become increasing prevalent for Internet of Things (IoT) applications. Instead of streaming the raw data to the cloud and executing CNN inference remotely, it would be very attractive to use local IoT devices to process as it enables IoT applications with independent decision-making ability. Since a single IoT device can hardly match the requirements of the CNN inference, especially for time-sensitive and high-accuracy tasks, the distributedin-situCNN inference becomes a potential solution. However, because of the inherently tightly coupled structure of existing CNN models, it is difficult to distribute the inference efficiently. In this article, we enhance the distributedin-situCNN inference in the IoT. We fundamentally reduce the communication overhead of distributed CNN inference by designing new loosely coupled structure (LCS). Experimental results demonstrate that LCS achieves the leading performance compared with other popular structures. Next, based on the LCS, we customize the partitioning method to reduce the synchronization points and design the decentralized asynchronous method to optimize communication in each synchronization point. To evaluate the effectiveness, we build a prototype system. When the number of IoT devices increases from 1 to 4, our system accelerates by up to$3.85\times $and reduces the memory footprint in each device by 70% with achieving a competitive accuracy and significantly outperforming other approaches. Jiangsu Du, Yunfei Du 0001, Dan Huang 0001, Yutong Lu, Xiangke Liao |
IEEE Internet Things J. | 2 |
| 2022 | Iteration number-based hierarchical gradient aggregation for distributed deep learning
Danyang Xiao, Jieying Zhou, Yunfei Du 0001, Weigang Wu |
J. Supercomput. | 4 |
| 2021 | Multi-Layer Networks for Ensemble Precipitation Forecasts PostprocessingabstractThe postprocessing method of ensemble forecasts is usually used to find a more precise estimate of future precipitation, because dynamic meteorology models have limitations in fitting fine-grained atmospheric processes and precipitation is driven more often by smaller-scale processes, while ensemble forecasts can hit this precipitation at times. However, the pattern of these hits cannot be easily summarized. The existing objective postprocessing methods tend to extend the rain area or false alarm the precipitation intensity categories. In this work, we introduce a multi-layer structure to simultaneously reduce the bias in forecast ensembles output by meteorology models and merge them to a quality deterministic (single-valued) forecast using cross-grid information, which differs quite dramatically from the previous statistical postprocessing method. The multi-layer network is designed to model the spatial distribution of future precipitation of different intensity categories(IC-MLNet). We provide a comparison of IC-MLNet to simple average as well as another two state-of-the-art ensemble quantitative precipitation forecasts (QPFs) postprocessing approaches over both single-model and multi-model ensemble forecasts datasets from TIGGE. The experimental results indicate that our model achieves superior performance over the compared baselines in precipitation amount prediction as well as precipitation intensities categories prediction. Fengyang Xu, Guanbin Li, Yunfei Du 0001, Zhiguang Chen 0001, Yutong Lu |
AAAI | 3 |
| 2021 | DeepPE: Emulating Parameterization in Numerical Weather Forecast Model Through Bidirectional Network
Fengyang Xu, Wencheng Shi, Yunfei Du 0001, Zhiguang Chen 0001, Yutong Lu |
ECML/PKDD (5) | 3 |
| 2021 | Robust graph convolutional networks with directional graph adversarial training
Weibo Hu, Chuan Chen 0001, Yaomin Chang, Zibin Zheng, Yunfei Du 0001 |
Appl. Intell. | 5 |
| 2021 | Learning deep discriminative representations with pseudo supervision for image clustering
Weibo Hu, Chuan Chen 0001, Fanghua Ye 0001, Zibin Zheng, Yunfei Du 0001 |
Inf. Sci. | 5 |
| 2021 | Model Parallelism Optimization for Distributed Inference Via Decoupled CNN StructureabstractIt is promising to deploy CNN inference on local end-user devices for high-accuracy and time-sensitive applications. Model parallelism has the potential to provide high throughput and low latency in distributed CNN inference. However, it is non-trivial to use model parallelism as the original CNN model is inherently tightly-coupled structure. In this article, we propose DeCNN, a more effective inference approach that uses decoupled CNN structure to optimize model parallelism for distributed inference on end-user devices. DeCNN is novel consisting of three schemes. Scheme-1 is structure-level optimization. It exploits group convolution and channel shuffle to decouple the original CNN structure for model parallelism. Scheme-2 is partition-level optimization. It is based on channel group to partition the convolutional layers, and then leverages input-based method to partition the fully connected layers, further exposing high degree of parallelism. Scheme-3 is communication-level optimization. It uses inter-sample parallelism to hide communications for better performance and robustness, especially in the weak network connections. We use ImageNet classification task to evaluate the effectiveness of DeCNN on a distributed multi-ARM platform. Notably, when using the number of devices from 1 to 4, DeCNN can accelerate the inference of large-scale ResNet-50 by 3.21×, and reduce 65.3 percent memory footprint, with 1.29 percent accuracy improvement. Jiangsu Du, Xin Zhu 0003, Minghua Shen, Yunfei Du 0001, Yutong Lu, Nong Xiao 0001, Xiangke Liao |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2020 | Dynamic reverse proxy chain generation for networks in data centersabstractReverse proxy, as one of the important components required by the data center to provide application services, has functions of access control, load balancing, connecting to different types of networks, etc. In the future, as the application services requiring reverse proxy further increase, network types become more and more diverse and complex, and the network hierarchy becomes higher and higher, reverse proxy will change from a single layer to multiple layers to form a reverse proxy chain. The construction of the reverse proxy chain will become one of the bottlenecks of data center networking operation and maintenance. In this paper, we propose a method of automatically constructing reverse proxy chains to avoid the problem of manual static configuration of the reverse proxy chain which is time-consuming, laborious, and difficult to maintain. In our software-defined networking experiment, we simulated a full-binary-tree-like topology of 1534 nodes. We recorded the time to generate and remove proxy chains of various lengths. The average time to generate all reverse proxy chains in the topology consisting of 1534 nodes with 100ms delay is around 5100ms, much smaller than manual configuration, which usually needs several hours. Guixin Guo, Kangyou Zhong, Jiang Li 0008, Yunfei Du 0001 |
APNOMS | 6 |
| 2020 | A Distributed In-Situ CNN Inference System for IoT ApplicationsabstractCNN is a popular deep learning structure able to provide intelligent processing in IoT applications. Instead of deploying the resource-hungry CNN inference workloads on the cloud, it would be promising to utilize local IoT devices for the in-situ processing. Since a single IoT device has only limited resources available, distributing over multiple local devices becomes a potential solution, especially for high-accuracy and time-sensitive tasks. However, it is non-trivial to distribute the inference of existing CNN models efficiently as they are inherently tightly-coupled structure. In this paper, we propose a distributed in-situ CNN inference system with the loosely-coupled CNN structure (LCS), the synchronization-oriented partitioning (SOP), and the decentralized asynchronous communication (DAC) for IoT applications. LCS is based on two novel design ideas, the homogeneous group and the intermittent shuffle. Experiments on ImageNet classification illustrate that LCS has the leading accuracy compared with other structures, under a given computation budget. SOP and DAC target on converting the loosely-coupled feature of LCS into practical performance improvement. SOP tries to partition LCS with fewer synchronization points and DAC reduces the communication overhead by overlapping communications. When the number of IoT devices increases from 1 to 4, our system accelerates by up to 3.85 ×, and reduces the memory footprint in each device by 70%, outperforming other approaches. Jiangsu Du, Minghua Shen, Yunfei Du 0001 |
ICCD | 3 |
| 2020 | Re-evaluation of Atomic Operations and Graph Coloring for Unstructured Finite Volume GPU SimulationsabstractIn general, race condition can be resolved by introducing synchronisations or breaking data dependencies. Atomic operations and graph coloring are the two typical approaches to avoid race condition. Graph coloring algorithms have been generally considered winning algorithms in the literature due to their lock free implementations. In this paper, we present the GPU-accelerated algorithms of the unstructured cell-centered finite volume Computational Fluid Dynamics (CFD) software framework named PHengLEI which was originally developed for aerodynamics applications with arbitrary hybrid meshes. Overall, the newly developed GPU framework demonstrate up to 4.8 speedup comparing with 18 MPI tasks run on the latest Intel CPU node. Furthermore, the enormous efforts have been invested to optimize data dependencies which could lead to race condition due to unstructured mesh indirect addressing and related reduction math operations. With careful comparison between our optimised graph coloring and atomic operations using a series of numerical tests with different mesh sizes, the results show that atomic operations are more efficient than our optimised graph coloring in all of the test cases on Nvidia Tesla GPU V100. Specifically, for the summation operation, using atomicAdd is twice as fast as graph coloring. For the maximum operation, a speedup of 1.5 to 2 is found for atomicMax vs. graph coloring. Xu Sun 0001, Xiaohu Guo, Yunfei Du 0001, Yutong Lu, Yang Liu 0005 |
SBAC-PAD | 4 |
| 2020 | UniIndex: An index and query middleware for parallel file systemsabstractSummary As data analysis scenarios keep increasing on high‐performance computing systems, the ability to select a small fraction of data from a large volume of scientific data sets is vital to accelerate scientific discovery. However, parallel file systems lack the ability to provide efficient data locating services at the granularity of both a file and a record. Existing methods for identifying and indexing data are often domain‐specific and do not scale to large scientific data sets. In this paper, we describe the design and implementation of UniIndex framework, which combines our proposed techniques for user‐annotation extraction, in‐memory cache layer, in‐situ indexing, and parallel query processing. Acting as middleware on top of production file systems, UniIndex enables efficient data locating services with minimal user effort. Our evaluations show that UniIndex can locate target files from directories containing millions of files in microseconds. By applying in situ indexing and the lightweight range‐bitmap index, record‐level index building time can be dramatically reduced while maintaining up to two orders of magnitude query speedup than scanning the entire data set. Peng Cheng 0012, Yutong Lu, Yunfei Du 0001, Zhiguang Chen 0001 |
Concurr. Comput. Pract. Exp. | 4 |
| 2019 | Understanding the Resource Demand Differences of Deep Neural Network Training
Jiangsu Du, Xin Zhu 0003, Yunfei Du 0001 |
ICA3PP (2) | 4 |
| 2019 | Optimizing Data Placement on Hierarchical Storage Architecture via Machine Learning
Peng Cheng 0012, Yutong Lu, Yunfei Du 0001, Zhiguang Chen 0001, Yang Liu 0005 |
NPC | 3 |
| 2019 | Tiered data management system: Accelerating data processing on HPC systems
Peng Cheng 0012, Yutong Lu, Yunfei Du 0001, Zhiguang Chen 0001 |
Future Gener. Comput. Syst. | 3 |
| 2019 | A high performance implementation of Zolo-SVD algorithm on distributed memory systems
Shengguo Li, Jie Liu 0002, Yunfei Du 0001 |
Parallel Comput. | 3 |
| 2019 | Toward fault-tolerant hybrid programming over large-scale heterogeneous clusters via checkpointing/restart optimization
Cheng Chen 0005, Yunfei Du 0001, Ke Zuo, Jianbin Fang, Canqun Yang |
J. Supercomput. | 2 |
| 2018 | Comparative Study of Distributed Deep Learning Tools on Supercomputers
Di Kuang, Mengqiang Chen, Yunfei Du 0001, Weigang Wu |
ICA3PP (1) | 6 |
| 2018 | Mimir+: An Optimized Framework of MapReduce on Heterogeneous High-Performance Computing System
Zhiguang Chen 0001, Yunfei Du 0001, Yutong Lu |
NPC | 3 |
| 2016 | Accelerating the Simulation of Thermal Convection in the Earth's Outer Core on Tianhe-2abstractNumerical simulation of thermal convection in the Earth's outer core requires extreme-scale computing due to the large temporal and spatial disparity, extreme physical parameters, rapid rotation and spherical geometry. In this work, the numerical simulation of the thermal convection in the Earth's outer core for CPU-MIC heterogeneous many-core systems is studied. Firstly, starting from a legacy parallel code based on the PETSc software package, a framework of the numerical simulation built on CPU-MIC heterogeneous many-core systems has been developed. Secondly, a sparse linear solver for CPUMIC heterogeneous many-core systems, which focuses on solving the two linear systems of the simulation, is presented and optimized. Thirdly, some computational kernels of the simulation, including sparse matrix-vector multiplication (SpMV) and polynomial preconditioner on distributed memory Xeon Phiaccelerated systems are implemented and optimized. In addition, in order to reduce the cost of data movement, we use methods to minimize the memory access, the PCI-E data transfer, and the MPI communication. Finally, some optimized measures are taken to the extended code. Experiments on Tianhe-2 Supercomputer show that as compared to the original code, our Xeon Phiaccelerated design is able to deliver 6.93x and 6.00x speedups for single MIC device and 64 MIC devices, respectively. Changmao Wu, Fangfang Liu 0004, Chao Yang 0002, Ligang Li, Yutong Lu, Leisheng Li, Yunfei Du 0001 |
ICPADS | 8 |
| 2016 | Accelerator-Centered Programming on Heterogeneous SystemsabstractParallel many cores contribute to heterogeneous architectures and achieve high computation throughput. Working as coprocessors and connected to general-purpose CPUs via PCIe, those special-purpose cores usually work as float computing accelerators (ACC). The popular programming models typically offload the computing intensive parts to accelerator then aggregate results, which would result in a great amount of data transfer via PCIe. In this paper, we introduce an ACC-centered model to leverage the limited bandwidth of PCIe, increase performance, reduce idle time of ACC. In order to realize dada-near-computing, our ACC-centered model arms to program centered on ACC and the control intensive parts are offloaded to CPU. Both CPU and ACC are devoted to higher performance with their architect feature. Validation on the Tianhe-2 supercomputer shows that the implementation of ACC-centered LU competes with the highly optimized Intel MKL hybrid implementation and achieves about 5× speedup versus the CPU version. Cheng Chen 0005, Yunfei Du 0001, Canqun Yang |
PDCAT | 2 |
| 2016 | ASDB: a resource for probing protein functions with small moleculesabstractUNLABELLED: : Identifying chemical probes or seeking scaffolds for a specific biological target is important for protein function studies. Therefore, we create the Annotated Scaffold Database (ASDB), a computer-readable and systematic target-annotated scaffold database, to serve such needs. The scaffolds in ASDB were derived from public databases including ChEMBL, DrugBank and TCMSP, with a scaffold-based classification approach. Each scaffold was assigned with an InChIKey as its unique identifier, energy-minimized 3D conformations, and other calculated properties. A scaffold is also associated with drugs, natural products, drug targets and medical indications. The database can be retrieved through text or structure query tools. ASDB collects 333 601 scaffolds, which are associated with 4368 targets. The scaffolds consist of 3032 scaffolds derived from drugs and 5163 scaffolds derived from natural products. For given scaffolds, scaffold-target networks can be generated from the database to demonstrate the relations of scaffolds and targets. AVAILABILITY AND IMPLEMENTATION: ASDB is freely available at http://www.rcdd.org.cn/asdb/with the major web browsers. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Peng Ding 0004, Xin Yan 0006, Minghao Zheng, Huihao Zhou, Yuehua Xu, Yunfei Du 0001, Qiong Gu, Jun Xu 0017 |
Bioinform. | 7 |
| 2015 | Performance Evaluation of HPGMG on Tianhe-2: Early Experience
Yulong Ao, Yiqung Liu 0005, Chao Yang 0002, Fangfang Liu 0004, Yutong Lu, Yunfei Du 0001 |
ICA3PP (4) | 7 |
| 2015 | FT-Offload: A Scalable Fault-Tolerance Programing Model on MIC Cluster
Cheng Chen 0005, Yunfei Du 0001, Zhen Xu 0004, Canqun Yang |
ICA3PP (4) | 2 |
| 2014 | HPCG: Preliminary Evaluation and Optimization on Tianhe-2 CPU-only NodesabstractHPCG has become a new metric for the design and ranking of HPC. By incorporating a local symmetric Gauss-Seidel preconditioned, HPCG implements the Conjugate Gradient method to solve a sparse linear system. HPCG performs poorly with irregular memory access and may consume a great deal of MPI resources when it is executed on supercomputers. This paper focuses on optimizing SpMV and the Gauss-Seidel preconditioned, the two most important kernels in HPCG. By evaluating the performance impacts of several representative sparse matrix formats, ELLPACK is selected due to its suitability for SIMD, resulting in a speedup of 2.3x for the SpMV kernel. Multi-coloring is performed for Gauss-Seidel, resulting in a speedup of 7.3x over the reference implementation. The CG convergence rate may also be improved after multi-coloring. Our experimental results show that our optimization process works well on supercomputers, achieving 6.5 Gflops on a CPU-only node. This has boosted the total HPCG Gflops by about 7x, giving rise to 80,151 Gflops on 8192 CPU-only Tianhe-2 nodes. Cheng Chen 0005, Yunfei Du 0001, Hao Jiang 0001, Ke Zuo, Canqun Yang |
SBAC-PAD | 2 |
| 2011 | Optimizing Linpack Benchmark on GPU-Accelerated Petascale Supercomputer
Feng Wang 0050, Canqun Yang, Yunfei Du 0001, Juan Chen 0001, Huizhan Yi, Weixia Xu 0001 |
J. Comput. Sci. Technol. | 3 |
| 2010 | Adaptive Optimization for Petascale Heterogeneous CPU/GPU ComputingabstractIn this paper, we describe our experiment developing an implementation of the Linpack benchmark for TianHe-1, a petascale CPU/GPU supercomputer system, the largest GPU-accelerated system ever attempted before. An adaptive optimization framework is presented to balance the workload distribution across the GPUs and CPUs with the negligible runtime overhead, resulting in the better performance than the static or the training partitioning methods. The CPU-GPU communication overhead is effectively hidden by a software pipelining technique, which is particularly useful for large memory-bound applications. Combined with other traditional optimizations, the Linpack we optimized using the adaptive optimization framework achieved 196.7 GFLOPS on a single compute element of TianHe-1. This result is 70.1% of the peak compute capability and 3.3 times faster than the result using the vendor's library. On the full configuration of TianHe-1 our optimizations resulted in a Linpack performance of 0.563PFLOPS, which made TianHe-1 the 5th fastest supercomputer on the Top500 list released in November 2009. Canqun Yang, Feng Wang 0050, Yunfei Du 0001, Juan Chen 0001, Jie Liu 0002, Huizhan Yi, Kai Lu 0001 |
CLUSTER | 3 |
| 2009 | Solving 2D Nonlinear Unsteady Convection-Diffusion Equations on Heterogenous Platforms with Multiple GPUsabstractSolving complex convection-diffusion equations is very important to many practical mathematical and physical problems. After the finite difference discretization, most of the time for equations solution is spent on sparse linear equation solvers. In this paper, our goal is to solve 2D Nonlinear Unsteady Convection-Diffusion Equations by accelerating an iterative algorithm named Jacobi-preconditioned QMRCGSTAB on a heterogenous platform, which is composed of a multi-core processor and multiple GPUs. Firstly, a basic implementation and evaluation for adapting the problem to this kind of platform is given. Then, we propose two optimization methods to improve the performance: kernel merging method and matrix boundary data processing. Our experimental evaluation on an AMD Opteron(tm) quad-core processor 2380 linked to an NVIDIA Tesla S1070 platform with four GPUs delivers the peak performance of 33 GFLOPS (double precision), which is a speedup of close to a factor 32 compared to the same problem running on 4 cores of the same CPU. Canqun Yang, Zhen Ge, Juan Chen 0001, Feng Wang 0050, Yunfei Du 0001 |
ICPADS | 5 |
| 2009 | FTPA: Supporting Fault-Tolerant Parallel Computing through Parallel RecomputingabstractAs the size of large-scale computer systems increases, their mean-time-between-failures are becoming significantly shorter than the execution time of many current scientific applications. To complete the execution of scientific applications, they must tolerate hardware failures. Conventional rollback-recovery protocols redo the computation of the crashed process since the last checkpoint on a single processor. As a result, the recovery time of all protocols is no less than the time between the last checkpoint and the crash. In this paper, we propose a new application-level fault-tolerant approach for parallel applications called the fault-tolerant parallel algorithm (FTPA), which provides fast self-recovery. When fail-stop failures occur and are detected, all surviving processes recompute the workload of failed processes in parallel. FTPA, however, requires the user to be involved in fault tolerance. In order to ease the FTPA implementation, we developed get it fault-tolerant (GiFT), a source-to-source precompiler tool to automate the FTPA implementation. We evaluate the performance of FTPA with parallel matrix multiplication and five kernels of NAS Parallel Benchmarks on a cluster system with 1,024 CPUs. The experimental results show that the performance of FTPA is better than the performance of the traditional checkpointing approach. Xuejun Yang, Yunfei Du 0001, Panfeng Wang, Hongyi Fu, Jia Jia 0004 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2008 | Static Analysis for Application-Level Checkpointing of MPI ProgramsabstractApplication-level checkpointing is a promising technology in the domain of large-scale scientific computing. The consistency of global checkpoint must be carefully guaranteed in order to correctly restore the computation. Usually, some complex coordinated protocols are employed to ensure the consistency of global checkpoint, which require logging orphan or in-transit messages during checkpointing. These protocols complicate the recovery of the computation and increase the checkpoint overhead due to logging message. In this paper, a new method which ensures the consistency of global checkpoint by static analysis is proposed. The method identifies the safe checkpointing regions in MPI programs, where the global checkpoint is always strongly consistent. All checkpoints are located in those safe checkpoint regions. During checkpointing, the method will not log any messages and introduce no extra overhead. The method was implemented and integrated into ALEC, which is a source-to-source precompiler for automating application-level checkpointing. The experimental results show that our method is effective. Panfeng Wang, Yunfei Du 0001, Hongyi Fu, Xuejun Yang, Haifang Zhou |
HPCC | 2 |
| 2008 | Optimal Placement of Application-Level CheckpointsabstractOne of the basic problems related to the efficient application-level checkpointing is the placement of checkpoints in the source codes. In this paper we discuss two common questions with a source-to-source precompiler ALEC: 1) if there are N checkpoints in the application's source code, how to pick M checkpoints out of them minimizing the total amount of checkpoint data? 2) if there are no checkpoint in the application's source code, how to insert a set of checkpoints minimizing the amount of checkpoint data? We reveal that these two questions can both be abstracted as a mathematic model which is similar to the 0-1 integer programming model, and the model can be solved using implicit enumeration method. The solving methods proposed in the paper have been implemented and integrated into ALEC. Experimental results show that the method is efficient. Panfeng Wang, Yunfei Du 0001, Xuejun Yang, Haifang Zhou |
HPCC | 3 |
| 2008 | Compiler-Assisted Application-Level Checkpointing for MPI ProgramsabstractApplication-level checkpointing can decrease the overhead of fault tolerance by minimizing the amount of checkpoint data. However this technique requires the programmer to manually choose the critical data that should be saved. In this paper, we firstly propose a live-variable analysis method for MPI programs. Then, we provide an optimization method of data saving for application-level checkpointing based on the analysis method. Based on the theoretical foundation, we implement a source-to-source precompiler (ALEC) to automate application-level checkpointing. Finally, we evaluate the performance of five FORTRAN/MPI programs which are transformed and integrated checkpointing features by ALEC on a 512-CPU cluster system. The experimental results show that i) the application-level checkpointing based on live-variable analysis for MPI programs can efficiently reduce the amount of checkpoint data, thereby decrease the overhead of checkpoint and restart; ii) ALEC is capable of automating application-level checkpointing correctly and effectively. Xuejun Yang, Panfeng Wang, Hongyi Fu, Yunfei Du 0001, Jia Jia 0004 |
ICDCS | 4 |
| 2008 | GiFT: Automating FTPA Implementation for MPI ProgramsabstractFault tolerance is a critical issue in the arena of large-scale computing. The fault-tolerant parallel algorithm (FTPA) is an application-level technique for tolerating hardware failures. FTPA achieves fast failure recovery making use of parallel recomputing. However, it complicates the coding of the application program. This paper uses compiler technology to automate the design of FTPA, and introduces the implementation of a tool called GiFT (Get it Fault-Tolerant). GiFT utilizes the extended data-flow analysis to choose the state needed by failure recovery, exploits the parallel recomputing time model to compute the optimal number of recomputing processes, and uses parallelization technologies to generate parallel recomputing codes. The experimental results show that original MPI programs can be transformed into the FTPA counterparts by GiFT correctly, and the performance of GiFT-generated FTPA programs is comparable to the performance of hand-modified FTPA programs. Hongyi Fu, Yunfei Du 0001, Panfeng Wang, Jia Jia 0004, Xuejun Yang |
ICPADS | 2 |
| 2008 | Automated application-level checkpointing based on live-variable analysis in MPI programsabstractThis paper proposes an optimization method of data saving for application-level checkpointing based on the live-variable analysis method for MPI programs. We presents the implementation of a source-to-source precompiler (CAC) for automating applicationlevel checkpointing based on the optimization method. The experiment shows that CAC is capable of automating application-level checkpointing correctly and reducing checkpoint data effectively. Panfeng Wang, Xuejun Yang, Hongyi Fu, Yunfei Du 0001, Zhiyun Wang, Jia Jia 0004 |
PPoPP | 4 |
| 2007 | The Fault Tolerant Parallel Algorithm: the Parallel Recomputing Based Failure Recovery
Xuejun Yang, Yunfei Du 0001, Panfeng Wang, Hongyi Fu, Jia Jia 0004, Guang Suo |
PACT | 2 |
| 2007 | A data-distributed parallel algorithm for wavelet-based fusion of remote sensing images
Xuejun Yang, Panfeng Wang, Yunfei Du 0001, Haifang Zhou |
Frontiers Comput. Sci. China | 3 |