EDBT 2026 Demo / reviewers in the wild / expert
Yong Dong
dblp:42/3197
· DBLP profile ↗
39ranked-venue papers
5as first author
25since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 23 · 3 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 4 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Revisiting free space fragmentation: A new garbage collection scheme for F2FSabstractFlash Friendly File System (F2FS) is a log-structured file system (LFS) optimized for Flash memory characteristics and is widely deployed on mobile devices, embedded systems, and some Linux platforms that use NAND Flash. File fragmentation and free space fragmentation both affect the performance of F2FS. In this work, we investigate the performance impact of fragmentation through energy consumption characterization. Our measurements indicate that energy consumption increases with the number of file and free space fragments, based on experiments evaluating F2FS while serving I/O requests across diverse workload scenarios. While considerable efforts have been devoted to mitigating file fragmentation, comparatively less attention has been paid to understanding and optimizing free space fragmentation, which predominantly arises from the distribution of invalid blocks. We observe that reclaiming invalid blocks via background garbage collection (GC) incurs over 100 mJ of energy per invocation, yet yields only marginal reductions in free space fragmentation. This motivates us to improve GC effectiveness in reducing free space fragmentation. We propose the free space frag mentation-aware (FragGC) and file system perf ormance-aware GC (PerfGC) scheme. We seek to both reduce GCs and enhance the efficiency of each GC operation. We reassess the definition of a free space fragment through empirical analysis and introduce the free space fragmentation factor as a lightweight metric to quantify the degree of free space fragmentation at the segment level. FragGC optimizes victim segment selection and valid block migration based on this metric. PerfGC adjusts GC frequency according to the impact of free space fragmentation on file system performance. Experimental results on a real platform demonstrate that FragGC and PerfGC reduce GC count compared to traditional F2FS and its latest GC optimization, ATGC. FragGC reduces the time to replay traces by 24.4% to 40.9% for large-scale applications. Ruibo Wang, Yong Dong, Weizhao Lin |
Future Gener. Comput. Syst. | 3 |
| 2026 | DeloopSGNN: Revisiting Spectral GNNs Through the Lens of Spatial AggregationabstractGraph Neural Networks (GNNs) have been studied from two primary perspectives: spectral, which employs global graph signal filtering and is theoretically more expressive, and spatial, which builds on local neighborhood aggregation and generalizes well across diverse graph structures. While spectral GNNs are expected to perform better in theory, they often underperform in practice compared to spatial models. To better understand this gap, we introduce a novel theoretical framework for converting spectral GNNs into the spatial domain, allowing for more intuitive analysis. This transformation reveals that signal looping and repeated high-order aggregation are major causes of over-smoothing in spectral GNNs. By addressing these issues in the spatial domain and converting the model back to the spectral domain, we propose DeloopSGNN, a spectral GNN with improved expressive capacity. Experiments on benchmark datasets show that DeloopSGNN achieves consistently strong performance in terms of accuracy and adversarial robustness, demonstrating that spectral GNNs can benefit significantly from careful architectural design grounded in our proposed framework. Duanyu Li, Huijun Wu 0001, Kai Lu 0001, Zhenwei Wu, Yong Dong, Ruibo Wang |
AAAI | 7 |
| 2026 | PallasGNN: Curriculum-Based Pattern Mining for Robust GNNs
Kaiwen Xia, Huijun Wu 0001, Ruibo Wang, Zhenwei Wu, Yong Dong |
PAKDD (3) | 7 |
| 2026 | A survey of anomaly detection in HPC systems using machine learningabstractAbstract High-performance computing (HPC) systems must remain stable and reliable to consistently deliver robust computational power and ensure the proper execution of user jobs. Anomaly detection is a key means to ensure the stability and reliability of these systems. With the expansion of HPC systems and changes in their architecture, accurately identifying anomalies in dynamic environments has become increasingly challenging. Traditional detection methods rely on experience and rules, which could be inefficient and inaccurate. To address these issues, researchers have proposed machine learning-based methods to automatically process large amounts of complex data, improving the efficiency of anomaly identification and diagnosis. In this survey, we conduct a comprehensive and in-depth investigation of machine learning-based anomaly detection methods in HPC systems. Firstly, we summarize and introduce the background and challenges of anomaly detection in HPC systems. Secondly, we compare a series of machine learning-based anomaly detection works in detail and summarize their frameworks. We conclude their advantages and disadvantages and application scenarios. Finally, we discuss several promising development trends of machine learning-based HPC system anomaly detection. Wei Zhang 0027, Yiqin Dai, Huijun Wu 0001, Zhenwei Wu, Hongyun Tian, Juan Chen 0001, Chubo Liu, Yong Dong |
CCF Trans. High Perform. Comput. | 11 |
| 2026 | Survey of storage systems in high performance computingabstractAbstract As high performance computing (HPC) moves towards exascale, storage systems face core challenges such as data flooding, bandwidth bottlenecks, mixed load coordination, and performance cost balancing. This article systematically reviews the cutting-edge technologies of high performance storage systems, covering four aspects: storage architecture, hardware, software, and networking. At the architecture level, storage computing separation, distributed and hierarchical architectures decouple computing and storage resources, and optimize latency and scalability through high-speed networks. Typical cases include supercomputer systems such as Frontier and Fugaku. In terms of hardware, persistent memory, all flash array, and integrated storage and computing chips significantly improve throughput and reduce latency, while ZNS SSD and QLC technology optimize cost and lifespan. At the software level, distributed parallel file systems respond to massive small files and high concurrency access through burst buffering technology. In network communication, low latency protocols such as Slingshot, InfiniBand, and RoCE support TB level bandwidth, while CXL technology promotes storage resource pooling. In the future, photon interconnection, AI native architecture, and green energy-saving technologies will further promote the development of high performance storage towards efficiency and intelligence, to support ZB level storage requirements in scenarios such as Exascale computing and AI training. Gen Zhang, Zhenlong Song, Xinhai Chen 0001, Yong Dong |
CCF Trans. High Perform. Comput. | 6 |
| 2026 | CML-PowF: Data Clustering Matching Based Low-overhead Multiple CPU Real-time Power ForecastingabstractEfficient CPU power capping is essential for energy saving and fault tolerance in parallel computing clusters, but its effectiveness depends on accurate and timely processor power forecasting with minimal sampling overhead. Existing methods often struggle to balance these factors under scalability constraints, as hardware limitations tightly bound the available sampling resources. This article focuses on the issue of high-precision real-time processor power forecasting while maintaining (or minimally increasing) the total overhead of multiprocessor power forecasting, particularly when the parallelism scale ranges from P to 2P processors or when the problem size scales from M to 2M . We propose CML-PowF , a low-overhead multiprocessor real-time power forecasting approach based on data clustering. CML-PowF integrates two key algorithms: Alg-CEF , which conducts cluster matching on the runtime characteristics of the program at the P / M scale, and models the tradeoff among forecasting error, time span, and sampling overhead. Alg-MSF , which leverages execution patterns from smaller-scale runs to determine the optimal sampling overhead and forecasting time span at the 2P / 2M scale. We evaluate CML-PowF on x86 and ARM platforms with up to 32 computing nodes (2,048 cores). Results show that it achieves 3–6% forecasting error at large scales with only 0.2–0.5% degradation compared to the P / M scale, without increasing total sampling overhead. Integrated with the PowC control system, CML-PowF effectively maintains real-time processor power below target thresholds. Rongyu Deng, Juan Chen 0001, Yuan Yuan 0034, Yong Dong, Aolin Cao, Yida Gu, Dingwen Tao |
ACM Trans. Archit. Code Optim. | 5 |
| 2026 | On Theoretical Stability Proof and Stability Margin Analysis of Enhanced Droop-Free Control Schemes for Islanded MicrogridsabstractThis article studies enhanced droop-free control strategies with sparse neighboring communication for achieving effective active power sharing of distributed energy resources (DERs) while maintaining the frequency stability of islanded microgrids. The normalized active power consensus (NAPC) based droop-free control can share the load among controllable DERs in proportion to their available capacities. However, existing literature exclusively takes the asymptotic stability of the NAPC-based droop-free control for granted, lacking a comprehensive theoretical proof that is critical for ensuring its effective design and practical implementation. This article, for the first time, provides a thorough theoretical proof of the asymptotic stability of two NAPC-based droop-free control schemes: ordinary NAPC (O-NAPC) and amplifier-equipped NAPC (A-NAPC), by testifying that all effective eigenvalues have negative real parts. The effect of various system settings on the stability margins is further analyzed with respect to the average admittance of the electrical network, the sparseness of the communication network, and the average available capacity of controllable DERs. Based on the sensitivity of eigenvalues with respect to perturbations, a vulnerability analysis is conducted to identify the weaknesses in the microgrids. Case studies demonstrate that the available capacity of controllable DERs has the most decisive influence on the stability margin of NAPC-based droop-free control, while O-NAPC/A-NAPC control scheme is more suitable for microgrids with DERs of larger/ smaller available capacities. Weipeng Liu, Upendra Prasad, Yutian Liu 0002, Yong Dong, Lei Wu 0004 |
IEEE Trans. Ind. Informatics | 4 |
| 2026 | A Survey on Machine Learning-Based HPC I/O Analysis and OptimizationabstractThe soaring computing power of HPC systems supports numerous large-scale applications, which generate massive data volumes and diverse I/O patterns, leading to severe I/O bottlenecks. Analyzing and optimizing HPC I/O is therefore critical. However, traditional approaches are typically customized and lack the adaptability required to cope with dynamic changes in HPC environments. To address the challenge, Machine Learning (ML) has been increasingly adopted to automate and enhance I/O analysis and optimization. Given sufficient I/O traces from HPC systems, ML can learn underlying I/O behaviors, extract actionable insights, and dynamically adapt to evolving workloads to improve performance. In this survey, we propose a novel taxonomy that aligns HPC I/O problems with learning tasks to systematically review existing studies. Through this taxonomy, we synthesize key findings on research distribution, data preparation, and model selection. Finally, we discuss several directions to advance the effective integration of ML in HPC I/O systems. Jingxian Peng, Huijun Wu 0001, Zhenwei Wu, Wei Zhang 0027, Yiqin Dai, Yong Dong |
IEEE Trans. Parallel Distributed Syst. | 9 |
| 2025 | MergeFS: Optimizing Node-Local Burst Buffers for Complex HPC WorkflowsabstractHigh-performance computing (HPC) applications are increasingly transitioning from traditional numerical simulations to an intelligent fusion paradigm integrating AI algorithms and big data analytics, exemplified by initiatives such as AI4Science. This evolution, coupled with rising problem complexity, results in workflows composed of interdependent subtasks. Existing HPC storage solutions, particularly burst buffer systems, have yet to adequately address the unique challenges posed by such workflows, including efficient cross-task data sharing and namespace fusion, leading to suboptimal resource utilization and performance bottlenecks in complex dependency scenarios. In this paper, we present MergeFS, a lightweight, workflow-aware burst buffer file system that incorporates a treestructured workflow registry for precise dependency management alongside a dynamic multi-namespace mechanism enabling rapid and isolated data access. MergeFS effectively integrates workflow management, namespace control, and data view fusion. Experimental evaluations demonstrate that MergeFS significantly outperforms current workflow-centric burst buffer optimizations in runtime performance with low management overhead. Zhaohao Zhong, Huijun Wu 0001, Yong Dong, Zhenwei Wu, Ruibo Wang |
ICPADS | 4 |
| 2025 | MIST: Towards MPI Instant Startup and Termination on Tianhe HPC SystemsabstractAs the size of MPI programs grows with expanding HPC resources and parallelism demands, the overhead of MPI startup and termination escalates due to the inclusion of less scalable global operations. Global operations involving extensive cross-machine communication and synchronization are crucial for ensuring semantic correctness. The current focus is on optimizing and accelerating these global operations rather than removing them, as the latter involves systematic changes to the system software stack and may impact program semantics. Given this background, we propose a systematic solution named MIST to safely eliminate global operations in MPI startup and termination. Through optimizing the generation of communication addresses, designing reliable communication protocols, and exploiting the resource release mechanism, MIST eliminates all global operations to achieve MPI instant startup and termination while ensuring correct program execution. Experiments on Tianhe-2A supercomputer demonstrate that MIST can reduce theMPI_Init()time by 32.5-77.6% and theMPI_Finalize()time by 28.9-85.0%. Yiqin Dai, Ruibo Wang, Yong Dong, Juan Chen 0001, Huijun Wu 0001, Mingtian Shao, Kai Lu 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2024 | Energy-Aware Task Mapping for Multi-Core Processor: A Machine Learning Based ApproachabstractWhen parallel workloads do not scale with the core, it is challenging to utilize parallelism effectively, especially on distributed cache/memory multi-core platforms. The mapping of processes to cores affects both the cache/memory latencies and the power consumption, determining the energy efficiency of computing subsystems. Consequently, performance and energy metrics usually cannot be simultaneously optimal, which demands a trade-off optimization mapping strategy. To this end, this paper presents a machine learning-based Energy-aware task Mapping Predictor (EMP). EMP adjusts memory-intensive and compute-intensive programs mapping on a single computing node by capturing hardware resources and program characteristics. For parallel programs on multiple nodes, EMP tries to add nodes and adjusts the mapping strategy to improve program performance while controlling energy consumption. We evaluate EMP on both ARM and x86 platforms. In single-node experiments, EMP reduces the Energy Delay Product (EDP) by 68.35% and 60.45% on ARM (A) and ARM (B), respectively, and in multi-node experiments, EMP reduces the EDP by 29.52% as compared to clustered task mapping. Tao Xu 0052, Juan Chen 0001, Yong Dong |
HPCC | 4 |
| 2024 | Towards Highly Compatible I/O-Aware Workflow Scheduling on HPC SystemsabstractScientific workflows on High-Performance Computing (HPC) consist of multiple data processing and computing tasks with dependencies. Efficiently scheduling computing resources and multi-tier storage across workflow tasks is crucial for optimizing performance. Existing solutions often fall short in achieving the co-scheduling of computing and 1/O resources and lack compatibility with HPC system software. In this paper, we introduce a performance model for scheduling workflows on HPC systems to enhance the understanding of workflow scheduling and facilitate the testing of the scheduling algorithm. Additionally, we propose THman, an open-source scientific workflow scheduler featuring our heuristic scheduling algorithm, Highest Contribution First (HCF). THman achieves online co-scheduling of computing and I/O resources for workflows and is designed to work with traditional HPC batch schedulers for high compatibility. We evaluate THman using simulated workloads and real-world workflow applications. Experimental results show that THman reduces workflow makespan by up to 30.9% compared to alternative methods. Yiqin Dai, Ruibo Wang, Yong Dong, Kai Lu 0001 |
SC | 3 |
| 2024 | Optimality of the Proper Gaussian Signal in Complex MIMO Wiretap ChannelsabstractThe multiple-input multiple-output (MIMO) wiretap channel (WTC) serves as a fundamental model for exploring information-theoretic secrecy in wireless communication systems, involving a transmitter, a legitimate user, and an eavesdropper. This paper investigates the optimality of proper complex signals in complex WTCs. Our primary contribution lies in the derivation of a determinant inequality, which establishes that the secrecy rate of degraded complex MIMO WTCs is maximized when the signal is proper, meaning that its pseudo-covariance matrix is a zero matrix. Remarkably, we extend this result beyond the degraded scenario to the general complex WTC by leveraging a min-max reformulation of the secrecy capacity. Thus, we demonstrate that focusing on proper signals is sufficient when examining the secrecy capacity of the complex WTC. Overall, this work highlights the significance of the determinant inequality we derive and its implications for optimizing secrecy rates in the complex WTC. Yong Dong, Yinfei Xu, Tong Zhang 0026, Yili Xia |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2024 | Faster and Scalable MPI Applications LaunchingabstractDistributed parallel MPI applications are the dominant workload in many high-performance computing systems. While optimizing MPI application execution is a well-studied field, little work has considered optimizing the initial MPI application launching phase, which incurs extensive cross-machine communications and synchronization. The overhead of MPI application launching can be expensive, accounting for more than million core hours per 10K nodes annually on the production Tianhe-2A supercomputer, which will increase as the number of parallel machines used grows. Therefore, it is critical to optimize the MPI application launching process. This paper presents a novel approach to optimizing the MPI application launch. Our approach adopts a location-aware address generation rule to eliminate the need for address exchange and a topology-aware global communication scheme to optimize cross-machine synchronization. We then design a new application launch procedure to support the proposed optimizations to further reduce the pressure of the shared I/O system. Our techniques have been deployed to production in the Tianhe-2A supercomputer and the Next Generation Tianhe Supercomputer. Experimental results show that our approach scales well and outperforms alternative schemes, reducing the MPI application launching time by over 29% with 320K MPI processes. Yong Dong, Yiqin Dai, Kai Lu 0001, Ruibo Wang, Juan Chen 0001, Mingtian Shao, Zheng Wang 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2023 | HighRPM: Combining Integrated Measurement and Sofware Power Modeling for High-Resolution Power MonitoringabstractIn an era where power and energy are the first-class constraints of computing systems, accurate power information is crucial for energy efficiency optimization in parallel computing systems. Existing power monitoring techniques rely on either software-centric power models that suffer from poor accuracy or integrated hardware measurement schemes that have a low reading update frequency and coarse granularity. These result in a low spatiotemporal resolution for power monitoring. This paper introduces HighRPM, a new method for accurately measuring power consumption on parallel computing systems. HighRPM combines coarse-grained power sensor readings and software power modeling techniques to improve temporal and spatial resolutions. To provide high-frequent power readings in the temporal domain, HighRPM employs statistical modeling and machine learning techniques to predict the long-term power trend and the short-term fluctuations in power consumption. To improve spatial coverage, HighRPM takes low-time resolution node-level power consumption and uses a neural network to distribute the power readings to lower-level computing components like CPUs and memory components. We evaluate HighRPM by applying it to both ARM-based and X86-based platforms. Experimental results show that HighRPM improves time resolution by 10 times, provides accurate readings for CPUs and memory, and reduces error by 7-24% compared to other power modeling methods. Xinxin Qi, Juan Chen 0001, Yong Dong, Yuan Yuan 0034, Tao Xu 0052, Rongyu Deng, Kexing Zhou, Zheng Wang 0001 |
ICPP | 3 |
| 2023 | BCN: A Fast Notified Backpressure Congestion ManagementabstractApplications such as cloud computing, big data processing, and artificial intelligence, demand high bandwidth and low latency in datacenter networks. Existing congestion control and flow control schemes at switches have limitations in granularity, fairness, and signal transmission delay. This paper proposes a fast notified per-hop per-flow backpressure congestion management, called BCN. BCN dynamically allocates queues for each flow at each switch hop, precisely pauses upstream queues based on queuing conditions, and enables non-paused queues to continue transmission. The switch sends network status notifications, collaborating with sender-side speed up and deceleration strategies to achieve rapid traffic rate adjustment. To the best of our knowledge, this is the first work focusing on fast rate adjustment of per-hop per-flow traffic. We evaluate BCN in real traffic scenarios and observe significant reductions in flow completion time of 53% and 60% compared to BFC and DCQCN, respectively. Under high workload conditions, BCN also improves network throughput. In high incast scenarios, BCN outperforms BFC and DCQCN, reducing flow completion time by 67.7% and 74.2%, respectively, while maintaining throughput comparable to BFC. Additionally, BCN minimizes transmission delay, ensuring nearly lossless transmission and optimizing overall throughput. Dinghuang Hu, Dezun Dong, Yong Dong |
IPCCC | 6 |
| 2023 | Optimality of Proper Gaussian Signaling for SIMO Wiretap ChannelsabstractIn this paper, we study the optimality of proper signals for wiretap channels (WTC). We characterize the secrecy capacity of WTC in both separate and augmented forms, and show that proper Gaussian signals can achieve the secrecy capacity for both single-input multiple-output (SIMO) and single-input single-output (SISO) WTC. We also use an accelerated difference of convex functions algorithm (ADCA) to verify our results. Yong Dong, Yinfei Xu, Tong Zhang 0026, Yili Xia |
WCNC | 1 |
| 2023 | Processor power forecasting through model sample analysis and clustering
Kexing Zhou, Yong Dong, Juan Chen 0001, Rongyu Deng, Yifei Guo, Zhixin Ou |
CCF Trans. High Perform. Comput. | 2 |
| 2023 | UrsaX: Integrating Block I/O and Message Transfer for Ultrafast Block Storage on SupercomputersabstractIt is increasingly important for the next-generation exascale supercomputers to extend its applications beyond traditional high-performance computing (HPC) scenarios, so as to achieve high social and economic benefit. Similar to Amazon Web Services (AWS) and Alibaba Cloud, cloud-style virtual HPC service is a promising application scenario on supercomputers, for which remote block storage is the key to provide tenants with supercomputers’ extremely high storage performance. Unfortunately, the state-of-the-art block storage software systems (such as URSA and Ceph) cannot adapt to the advanced hardware features of supercomputers. This article presents UrsaX, an efficient block storage service for our next-generation Tianhe exascale supercomputer that is equipped with the high-performance global express (GLEX) network and nonvolatile memory express (NVMe) SSDs. UrsaX’s virtual disks, which can be mounted like normal physical ones, enable not only traditional HPC applications but also supercomputer-oblivious POSIX applications to enjoy the high performance of supercomputers. At the core of UrsaX is with a novel design of the efficient integration of on-disk block I/O and in-network message transfer on supercomputers. UrsaX utilizes the NVMe Fabrics kernel module to expand the NVMe standard on the supercomputer network, and separates metadata I/O and data I/O of blocks, respectively, being handled over the mini packet (MP) and remote direct memory access (RDMA) protocols. We thoroughly explore the design space for remote block storage on supercomputers, including parallelism, scalability, fault tolerance, and consistency. We conduct an extensive evaluation on a subset of our exascale supercomputer consisting of 44 storage machines (each with four NVMe SSDs). The result shows that UrsaX achieves local-storage-level I/O latency (tens of microseconds) while being able to linearly increase the aggregate performance (IOPS and throughput) as the system scale increases, an order of magnitude higher than the state-of-the-art block storage systems. Shun Gai, Yiming Zhang 0003, Xuchao Xie, Yong Dong, Zhenlong Song |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2022 | Efficient Dynamic Binary Translation with Accumulative Persistent Code CachingabstractDynamic binary translation technology can directly translate the binary code of one instruction architecture into the one of another instruction architecture at runtime, and execute it on the target machine. This technology has been applied in many fields, and several dynamic binary translators developed on this basis have shown great translation efficiency. While efficiently translating instruction streams is required, the challenge of mitigating translation overhead is also the essential issue. At present, there are many optimization techniques that can effectively improve the performance of dynamic binary translator, however, most of the techniques are ineffective for the applications with short run times or massive cold code regions. This paper introduces a method of persistent code caching that can be shared and reused between different executions. This method enables the same program to reuse the translated code extracted from former executions to skip part of the translation work, and allows the persistent code accumulated to raise reuse hit ratios. We have implemented a prototype based on Box64, a open source and newly released efficient dynamic binary translator, with persistent code caching method. The experimental results demonstrate that the optimization method could efficiently improve the performance of Box64. Haoming Lin, Yong Dong, Wanqing Chi, Zhenwei Wu, Hongqing Zeng |
ICPADS | 2 |
| 2022 | The Fast and Scalable MPI Application Launch of the Tianhe HPC systemabstractFast and scalable MPI application launch helps achieve exascale performance and is becoming a common goal in high-performance computing. However, the traditional launch technique suffers from scalability deficiencies in the global information exchange and the global barrier operation. This drawback makes it challenging to launch MPI applications quickly in large-scale systems. In this paper, we propose a fast and scalable application launch technique and details its associated hardware and software support. The optimized launch technique includes a locality-aware static address generation rule for eliminating the need for address exchange and a topology-aware global communication scheme for improving global communication efficiency. We also propose an optimized application launch sequence for supporting the above launch technique. We implement and evaluate the proposed launch technique on the Tianhe-2A supercomputer and the Tianhe Exascale Prototype Upgrade System. Experimental results show that our technique can reduce the launch time by 26.1% when launching an application with 256K processes. Yiqin Dai, Yong Dong, Kai Lu 0001, Ruibo Wang, Mingtian Shao, Juan Chen 0001 |
IPDPS | 2 |
| 2022 | MIMO Gaussain Cognitive Interference Channels With Confidential MessagesabstractThe secrecy capacity of the multiple-input multiple-output (MIMO) Gaussian cognitive interference channel with common, private and confidential messages is characterized in this paper. To show the achievability, we use jointly Gaussian auxiliary random variables to evaluate the existing single-letter description of the capacity region for the discrete memoryless channel. The converse part is established by invoking linear estimation theory and a new version of the extremal inequality. To resolve the Gaussian optimality problem in the extremal inequality, the perturbation framework is employed. Our results explore the connections between the MIMO Gaussian broadcast channel and its distributive antennas counterpart, i.e., the MIMO Gaussian interference channel. Yinfei Xu, Tong Zhang 0026, Yong Dong, Yili Xia |
ISIT | 3 |
| 2022 | Towards Scalable Resource Management for SupercomputersabstractToday's supercomputers offer massive computation resources to execute a large number of user jobs. Effectively managing such large-scale hardware parallelism and workloads is essential for supercomputers. However, existing HPC resource management (RM) systems fail to capitalize on the hardware parallelism by following a centralized design used decades ago. They give poor scalability and inefficient performance on today's supercomputers, which will worsen in exascale computing. We present ESlurm, a better RM for supercomputers. As a departure from existing HPC RMs, ESlurm implements a distributed communication structure. It employs a new communication tree strategy and uses job runtime estimation to improve communications and job scheduling efficiency. ESlurm is deployed into production in a real supercomputer. We evaluate ESlurm on up to 20K nodes. Compared to state-of-the-art RM solutions, ESlurm exhibits better scalability, significantly reducing the resource usage of master nodes and improving data transfer and job scheduling efficiency by a large margin. Yiqin Dai, Yong Dong, Kai Lu 0001, Ruibo Wang, Wei Zhang 0027, Juan Chen 0001, Mingtian Shao, Zheng Wang 0001 |
SC | 2 |
| 2021 | Software/Hardware Co-Design Optimization for Sparse Convolutional Neural NetworksabstractDeep convolutional neural network (DNN) has been widely used in image recognition, target detection, and natural language processing. Unfortunately, the previous state-of-the-art convolutional neural network (CNN) increase model accuracy by increasing the number of layers of the network, which leads to the problem of large models and complex calculations. Previous work has demonstrated that network compression can be achieved by weight pruning approaches, and that specialized hardware can be used to speed up reasoning. However, the format of compressed sparse rows (CSR) or compressed sparse columns (CSC) is generally used in model pruning. The approach leads to coding operations before calculation and decoding operations during calculation by processing units on specialized hardware. And the approach degrades the performance of the model.In this paper, we propose a combination of hardware and software to address above bottlenecks. Firstly, we propose a novel software-based structured pruning, called vector-pruning. Secondly, we design a novel dataflow for our pruning method, called Vector-Sparse, our method eliminates complex coding and decoding Then, we design the corresponding hardware architecture for vector-prune. Finally, we evaluate the performance of our pruning approach for AlexNet on Xilinx xcvu19p-fsva3824-2-e. Experiments show that compared with a state-of-the-art convolution network accelerator, we proposed a software/hardware co-design approach that is 1.01x and 1.3x better in performance, effectively improving the inference speed of the model on Field Programmable Gate Array (FPGA). Wei Hu 0001, Yong Dong, Fang Liu 0031, Qiang Jiao |
SMC | 2 |
| 2021 | RISC-VTF: RISC-V Based Extended Instruction Set for TransformerabstractDeep learning model Transformer has been widely used in natural language processing(NLP) filed, and its demand for computing resources is also growing. However, general-purpose processors(CPU and GPU) have invested excessive hardware resources in the design because they have to flexibly support a variety of tasks, they are not efficient for the implementation of Transformer. Consequently, various software optimizations that towards general-purpose processors have been proposed one after another. But, under the condition of ensuring sufficient accuracy, the degree of software optimization is limited. It is requisite to friendly-support Transformer at the hardware level.After analysing the computational characteristics of the Transformer model, based on RISC-V, we designed a hardware friendly instruction set architecture for the Transformer model. In addition to the basic instruction, for the intensive and general computing part of the model, according to the expansion rules of RISC-V instruction, we design the matrix load/store instruction calculation instruction, softmax instruction, activation instruction and other user-defined instructions. They support any matrix scale, and deploy it on FPGA to realize a flexible and efficient custom processor RISC-VTF for Transformer. The design is integrated on the Xilinx toolkit zynq-7000 FPGA, and the resource consumption and performance are analyzed. Compared with the traditional common ISA(Instruction Set Architecture) such as x86, arm or MIPs, RISC-VTF provides higher code density and performance efficiency. Qiang Jiao, Wei Hu 0001, Fang Liu 0031, Yong Dong |
SMC | 4 |
| 2020 | Optimizing Accelerator on FPGA for Deep Convolutional Neural Networks
Yong Dong, Wei Hu 0001, Yonghao Wang, Qiang Jiao |
ICA3PP (2) | 1 |
| 2020 | Design of a Convolutional Neural Network Instruction Set Based on RISC-V and Its Microarchitecture Implementation
Qiang Jiao, Wei Hu 0001, Yuan Wen, Yong Dong, Zhenhao Li 0004, Yu Gan 0004 |
ICA3PP (2) | 4 |
| 2020 | PMC-Based Dynamic Adaptive CPU and DRAM Power Modeling
Yunfang Zhang, Yong Dong, Juan Chen 0001, Zhixin Ou, Yuan Yuan 0034 |
ICA3PP (1) | 2 |
| 2019 | Analyzing time-dimension communication characterizations for representative scientific applications on supercomputer systems
Juan Chen 0001, Yong Dong, Feihao Wu, Enqiang Zhou, Yuhua Tang |
Frontiers Comput. Sci. | 3 |
| 2019 | Co-regularized kernel ensemble regression
Dickson Keddy Wornyo, Xiangjun Shen, Yong Dong, Liangjun Wang, Shu-Cheng Huang |
World Wide Web | 3 |
| 2018 | Least squares kernel ensemble regression in Reproducing Kernel Hilbert Space
Xiangjun Shen, Yong Dong, Jianping Gou, Yongzhao Zhan 0001, Jianping Fan 0001 |
Neurocomputing | 2 |
| 2016 | Facial descriptor for Kinect depth using inner-inter-normal components local binary patterns and tensor histograms
Guangling Sun, Yong Dong, Xiaofei Zhou 0003, Zhi Liu 0003 |
Mach. Vis. Appl. | 2 |
| 2016 | Reducing Static Energy in Supercomputer Interconnection Networks Using Topology-Aware PartitioningabstractThe key to reducing static energy in supercomputers is switching off their unused components. Routers are the major components of a supercomputer. Whether routers can be effectively switched off or not has become the key to static energy management for supercomputers. For many typical applications, the routers in a supercomputer exhibit low utilization. However, there is no effective method to switch the routers off when they are idle. By analyzing the router occupancy in time and space, for the first time, we present a routing-policy guided topology partitioning methodology to solve this problem. We propose topology partitioning methods for three kinds of commonly used topologies (mesh, torus and fat-tree) equipped with the three most popular routing policies (deterministic routing, directionally adaptive routing and fully adaptive routing). Based on the above methods, we propose the key techniques required in this topology partitioning based static energy management in supercomputer interconnection networks to switch off unused routers in both time and space dimensions. Three topology-aware resource allocation algorithms have been developed to handle effectively different job-mixes running on a supercomputer. We validate the effectiveness of our methodology by using Tianhe-2 and a simulator for the aforementioned topologies and routing policies. The energy savings achieved on a subsystem of Tianhe-2 range from 3.8 to 79.7 percent. This translates into a yearly energy cost reduction of up to half a million US dollars for Tianhe-2. Juan Chen 0001, Yuhua Tang, Yong Dong, Jingling Xue |
IEEE Trans. Computers | 3 |
| 2014 | Parallel Bayesian Network Structure Learning for Genome-Scale Gene NetworksabstractLearning Bayesian networks is NP-hard. Even with recent progress in heuristic and parallel algorithms, modeling capabilities still fall short of the scale of the problems encountered. In this paper, we present a massively parallel method for Bayesian network structure learning, and demonstrate its capability by constructing genome-scale gene networks of the model plant Arabidopsis thaliana from over 168.5 million gene expression values. We report strong scaling efficiency of 75% and demonstrate scaling to 1.57 million cores of the Tianhe-2 supercomputer. Our results constitute three and five orders of magnitude increase over previously published results in the scale of data analyzed and computations performed, respectively. We achieve this through algorithmic innovations, using efficient techniques to distribute work across all compute nodes, all available processors and coprocessors on each node, all available threads on each processor and coprocessor, and vectorization techniques to maximize single thread performance. Sanchit Misra, Md. Vasimuddin, Kiran Pamnany, Sriram P. Chockalingam, Yong Dong, Maneesha Aluru, Srinivas Aluru |
SC | 5 |
| 2014 | Hybrid hierarchy storage system in MilkyWay-2 supercomputer
Yutong Lu, Enqiang Zhou, Zhenlong Song, Yong Dong, Dengping Wei, Jianying Xing, Yuan Yuan 0034 |
Frontiers Comput. Sci. | 6 |
| 2008 | Energy-Constrained OpenMP Static Loop SchedulingabstractIn high performance parallel computing, energy optimization for parallel loops becomes one key because the time of loops often takes a significant part of the whole execution time. Energy-constrained problem is one of the important research focuses. This paper studies energy-constrained problem based on OpenMP static loop scheduling. Firstly, we propose energy-constrained static scheduling algorithm (ECSS), which utilizes DVS to scale down voltage/frequency of the light-loaded processors in terms of energy constraint. Secondly, we propose Energy Constraint based Performance-Optimal Static Scheduling algorithm (ECPOSS), which combines loop rescheduling and DVS for the better performance under the same energy constraint. We prove ECPOSS can obtain the best performance under the same energy constraint. Through testing NPB3.2-OMP programs on 20-160 multiprocessor simulation environment, we evaluate the effectiveness of our algorithms. Experimental results show the performance of ECPOSS is better than that of ECSS by 4.81% under the 50% energy constraint on 100 processors. Juan Chen 0001, Yong Dong, Xuejun Yang, Panfeng Wang |
HPCC | 2 |
| 2008 | Energy-Oriented OpenMP Parallel Loop SchedulingabstractIn HPC, power-related concern becomes dominant aspects of hardware and software design. Significant research effort has been devoted towards the energy optimization of parallel loop. This article is focused on energy-oriented OpenMP static and dynamic parallel loop scheduling problem. Only DVS cannot obtain the maximum energy savings. It is necessary to combine parallel loop rescheduling and DVS. First, we propose an energy-saving static scheduling (ESSS) algorithm, which exploits the scheduling slack to save energy by DVS. Second, we propose an energy-saving optimal static scheduling (EOSS) algorithm, which obtains the maximum energy saving through combining loop rescheduling and DVS. Last, in order to reduce the energy of OpenMP dynamic scheduling, we shut down the processor when it is idle, which is called shut-down based dynamic scheduling (SBDS) algorithm. Finally, we demonstrate the effectiveness by experiments. Yong Dong, Juan Chen 0001, Xuejun Yang, Xuemeng Zhang |
ISPA | 1 |
| 2005 | Energy-Constrained Prefetching Optimization in Embedded Applications
Juan Chen 0001, Yong Dong, Huizhan Yi, Xuejun Yang |
EUC | 2 |
| 2005 | A Compiler-Directed Energy Saving Strategy for Parallelizing Applications in On-Chip MultiprocessorsabstractAs energy consumption becoming one of the key optimization objects in on-chip multiprocessor, compiling a parallelizing application combined with energy saving strategy is more significant. In this paper, we focus on an on-chip multiprocessors architecture, where each processor in on-chip multiprocessor can independently adjust its frequency and voltage for energy savings. Given an arrayintensive application, we simulate parallelizing application and analyze probable load imbalance; then our energy saving strategy determines each processor’s clock frequency and voltage level fit for each parallel fragment in terms of load imbalance. Here, parallel fragments mainly denote parallel loop nests. Further, we consider the serial code fragments as a severe load-unbalanced parallel partitioning when the redundant processors can be shut down. Initial experiment proves our energy saving strategy is successful in reducing the energy consumption of the parallel programs. Juan Chen 0001, Yong Dong, Xuejun Yang |
ISPDC | 2 |