Hai Jiang 0003

dblp:15/5983-3 · DBLP profile ↗
← Back
35ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0002-8136-5428ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 8 · 1 since 2021Software engineering, systems software and programming languages · 7Applied, interdisciplinary, general and emerging computing · 5Computer networks · 2Security and privacy · 2Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2026 FEditor: Consecutive Task Placement With Adjustable Shapes Using FPGA State Frames
abstract
Field Programmable Gate Arrays (FPGAs) are widely adopted in datacenters, where each FPGA is exclusively assigned to a task. This strategy results in significant resource waste and increased task rejections. To address this issue, placement algorithms adjust the locations and shapes of tasks based on Dynamic Partial Reconfiguration, which partitions an FPGA into multiple rectangular areas for sharing. However, existing schemes are designed for static task sets without adjustable shapes, incapable of optimizing the placement problem in datacenters. In this paper, FEditor is proposed as the first consecutive task placement scheme with adjustable shapes. It expands the planar FPGA models into three-dimensional ones with timestamps to accommodate consecutive tasks. To reduce the complexity of three-dimensional resource management,State Frames(SFs) are designed to compress the models losslessly. Three metrics and a nested heuristic algorithm are used for task placement. Experimental results demonstrate that FEditor has improved resource utilization by at least$19.8\%$and acceptance rate by at least$10\%$compared to the referenced algorithms.SFsand the nested algorithm accelerate the task placement by up to$10.26\times$. The suitability of FEditor in datacenter environments is verified by its time efficiency trends.
Zhiqian Xu 0001, Hai Jiang 0003, Keqin Li 0001
IEEE Trans. Parallel Distributed Syst.5
2025 WIC: Hiding Producer-Consumer Synchronization Delays with Warp-Level Interrupt-based GPU Communications
Jiajian Zhang, Fangyu Wu 0001, Hai Jiang 0003, Qiufeng Wang 0001, Genlang Chen, Chaoyi Pang
USENIX ATC3
2025 AlignMalloc: Warp-Aware Memory Rearrangement Aligned With UVM Prefetching for Large-Scale GPU Dynamic Allocations
abstract
As parallel computing tasks rapidly expand in both complexity and scale, the need for efficient GPU dynamic memory allocation becomes increasingly important. While progress has been made in developing dynamic allocators for substantial applications, their real-world applicability is still limited due to inefficient memory access behaviors. This paper introduces AlignMalloc, a novel memory management system that aligns with the Unified Virtual Memory (UVM) prefetching strategy, significantly enhancing both memory allocation and access performance in large-scale dynamic allocation scenarios. We analyze the fundamental inefficiencies in UVM access and first reveal the mismatch between memory access and UVM prefetching methods. To resolve this issue, AlignMalloc implements a warp-aware memory rearrangement strategy that exploits the regularity of warps to align with the UVM's static prefetching setup. Additionally, AlignMalloc introduces an OR tree-based structure within a host-co-managed framework to further optimize dynamic allocation. Comprehensive experiments demonstrate that AlignMalloc substantially outperforms current state-of-the-art systems, achieving up to$2.7 \times$improvement in dynamic allocation and$2.3 \times$in memory access. Additionally, eight real-world applications with diverse memory access patterns exhibit consistent performance enhancements, with average speedups$1.5 \times$.
Jiajian Zhang, Fangyu Wu 0001, Hai Jiang 0003, Qiufeng Wang 0001, Genlang Chen, Eng Gee Lim, Keqin Li 0001
IEEE Trans. Parallel Distributed Syst.3
2024 SyncMalloc: A Synchronized Host-Device Co-Management System for GPU Dynamic Memory Allocation across All Scales
abstract
Dynamic memory allocation on GPUs, increasingly crucial for applications with dynamic computational patterns, encounters significant challenges due to the complex calculations with intricate branches and substantial memory resources consumed by metadata from massive thread allocations. Despite the current research, there is a lack of a scalable and flexible solution that effectively manages dynamic memory allocation while minimizing memory usage on GPUs. This paper introduces SyncMalloc, a synchronized Host-Device Co-Management system that is specifically designed to adeptly handle dynamic memory allocations of diverse magnitudes. Through the integration of pipelining and producer-consumer mechanisms, SyncMalloc effectively reduces communication overhead and resolves architectural mismatches, further enhancing its capability through synergistic integration with CUDA’s unified memory to facilitate oversubscription. Moreover, SyncMalloc advances slab-based memory management to enhance the efficiency of small allocations, reducing conflict probabilities and overhead in high-activity scenarios. Finally, we present a comprehensive performance evaluation, expanding benchmarks and measurement dimensions to reflect the performance of real-world applications more accurately. The experimental results demonstrate the effectiveness of SyncMalloc in supporting dynamic GPU allocations scaled from 4B to 200GB from multiple perspectives. Our source code is available at https://github.com/jjZhang94/SyncMalloc.
Jiajian Zhang, Fangyu Wu 0001, Hai Jiang 0003, Genlang Chen, Qiufeng Wang 0001
ICPP3
2022 CRAC: An automatic assistant compiler of checkpoint/restart for OpenCL program
abstract
Summary Nowadays, people use multiple devices to meet the growing requirement for computing. With the application of multicard computing, fault tolerance, load balance, and resource sharing have been the hot issues and the checkpoint/restart (CPR) mechanism is critical in a preemptive system. This article proposes a CPR framework including the automatic compiler (CRAC) to achieve a feasible CPR system, especially for graphics processing unit applications on heterogeneous devices in OpenCL programs. By offering the positions of the CPR in source code, CRAC inserts primitives into programs and invokes the runtime support modules for final results. A comprehensive example and experiments have demonstrated the feasibility and effectiveness of proposed framework.
Genlang Chen, Jiajian Zhang, Zufang Zhu, Hai Jiang 0003, Chaoyi Pang
Concurr. Comput. Pract. Exp.5
2021 CASH: correlation-aware scheduling to mitigate soft error impact on heterogeneous multicores
abstract
With the exponential increase in the number of transistors under fast-paced technology progress, the soft error induced reliability issue is becoming even more challenging in heterogeneous multicore processor design. As there are significant opportunities to mitigate the soft error impacts through heterogeneous multicore scheduling, we show in this paper that the correlation among multiple applications exhibits important reliability characteristics, by defining a new metric to measure the system-level vulnerability factor of multiple applications and an approximate estimator to evaluate the metric fast and accurately for effective scheduling decisions. To approach these issues, we propose CASH, a Correlation-Aware Scheduling strategy to optimise heterogeneous multicore system reliability. Comprehensive simulation results demonstrate that the proposed approach is promising, achieving up to 21.4% reliability improvement with only 3.6% performance degradation when compared with performance-oriented scheduling policy.
Jiajia Jiao, Libao Wang, Yanxiang Li, Dezhi Han, Kuanching Li, Hai Jiang 0003
Connect. Sci.7
2021 CRState: checkpoint/restart of OpenCL program for in-kernel applications
Genlang Chen, Jiajian Zhang, Zufang Zhu, Qiangqiang Jiang, Hai Jiang 0003, Chaoyi Pang
J. Supercomput.5
2020 A Malware Detection Approach Using Malware Images and Autoencoders
abstract
Most machine learning-based malware detection systems use various supervised learning methods to classify different instances of software as benign or malicious. This approach provides no information regarding the behavioral characteristics of malware. It also requires a large amount of training data and is prone to labeling difficulties and can reduce accuracy due to redundant training data. Therefore, we propose a malware detection method based on deep learning, which uses malware images and a set of autoencoders to detect malware. The method is to design an autoencoder to learn the functional characteristics of malware, and then to observe the reconstruction error of autoencoder to realize the classification and detection of malware and benign software. The proposed approach achieves 93% accuracy and comparatively better F1-score values while detecting malware and needs little training data when compared with traditional malware detection systems.
Xiang Jin, Xiaofei Xing, Haroon Elahi, Guojun Wang 0001, Hai Jiang 0003
MASS5
2019 X-RDMA: Effective RDMA Middleware in Large-scale Production Environments
abstract
X-RDMA is a communication middleware deployed and heavily used in Alibaba's large-scale cluster hosting cloud storage and database systems. Unlike recent research projects which purely focus on squeezing out the raw hardware performance, it puts emphasis on robustness, scalability and maintainability of large-scale production clusters. X-RDMA integrates necessary features, not available in current RDMA ecosystem, to release the developers from complex and imperfect details. X-RDMA simplifies the programming model, extends RDMA protocols for application awareness, and proposes mechanisms for resource management with thousands of connections per machine. It also reduces the work for administration and performance tuning with built-in tracing, tuning and monitoring tools. X-RDMA has been deployed in several large-scale clusters with over 4000 servers in Alibaba cloud since 2016. It can save at least 70% development and maintenance time over RDMA, effectively improve performance and reduce network jitter especially when production servers are under pressure. It also helped locate over 30 issues in different layers of productions with over 5000 connections for each server on average.
Teng Ma 0006, Tao Ma 0006, Huaixin Chang, Kang Chen 0001, Hai Jiang 0003, Yongwei Wu 0001
CLUSTER7
2019 CRState: In-Kernel Checkpoint/Restart of OpenCL Program Execution on GPU
abstract
Checkpoint/restart is an important mechanism to achieve fault tolerance, load balancing and resources sharing in a preemptive system. As Graphics Processing Unit (GPU) becomes quite popular in high performance computing as well as OpenCL programs are portable across various CPUs and GPUs, checkpoint/restart of OpenCL programs on GPUs is in demand. However, due to the intricacy of computation states inside GPUs, there is no effective checkpoint/restart scheme for heterogeneous devices now. This paper proposes a feasible system, CRState, to achieve checkpoint/restart in GPU kernels. With the assistant of a pre-compiler, the primitives are inserted into programs. In run-time, the computation state existing in the underlying hardware is concretized and reconstructed at application level and is ported to heterogeneous devices. Comprehensive experiments have been conducted to demonstrate CRState's feasibility and effectiveness. The experimental results also indicate that CRState has the potential to reschedule resources and balance workload across heterogeneous devices.
Genlang Chen, Jiajian Zhang, Qiuru Lin, Hai Jiang 0003, Chaoyi Pang
ICPADS4
2019 GPU Acceleration of Ciphertext-Policy Attribute-Based Encryption
abstract
With the development of cloud computing, data security became popular in recent decades. However, traditional cryptography has some major limitations. For example, public key cryptography is not scalable in cases with many clients. Since Ciphertext-Policy Attribute-based encryption (CP-ABE) was developed in 2007, it has become as one of the major candidates to implement secure cloud storage. However, CP-ABE still cannot play a solid role due to its several limitations such as complexity of computation, lack efficiency revocation function, etc. This paper will review the CP-ABE and analyze the current CP-ABE toolkit. Major performance bottleneck will be identified and parallelized in CUDA. CP-ABE toolkit will be partially ported to GPU platform for acceleration. Some experiments have been conducted to demonstrate the effectiveness of the proposed approach.
Chaoyu Zhang, Ruiwen Shan, Hexuan Yu, Hai Jiang 0003
SNPD5
2019 Generic attribute revocation systems for attribute-based encryption in cloud storage
abstract
Attribute-based encryption (ABE) has been a preferred encryption technology to solve the problems of data protection and access control, especially when the cloud storage is provided by third-party service providers. ABE can put data access under control at each data item level. However, ABE schemes have practical limitations on dynamic attribute revocation. We propose a generic attribute revocation system for ABE with user privacy protection. The attribute revocation ABE (AR-ABE) system can work with any type of ABE scheme to dynamically revoke any number of attributes.
Genlang Chen, Zhiqian Xu 0001, Jiajian Zhang, Guojun Wang 0001, Hai Jiang 0003, Miaoqing Huang
Frontiers Inf. Technol. Electron. Eng.5
2018 Generic user revocation systems for attribute-based encryption in cloud storage
abstract
Cloud-based storage is a service model for businesses and individual users that involves paid or free storage resources. This service model enables on-demand storage capacity and management to users anywhere via the Internet. Because most cloud storage is provided by third-party service providers, the trust required for the cloud storage providers and the shared multi-tenant environment present special challenges for data protection and access control. Attribute-based encryption (ABE) not only protects data secrecy, but also has ciphertexts or decryption keys associated with fine-grained access policies that are automatically enforced during the decryption process. This enforcement puts data access under control at each data item level. However, ABE schemes have practical limitations on dynamic user revocation. In this paper, we propose two generic user revocation systems for ABE with user privacy protection, user revocation via ciphertext re-encryption (UR-CRE) and user revocation via cloud storage providers (UR-CSP), which work with any type of ABE scheme to dynamically revoke users.
Genlang Chen, Zhiqian Xu 0001, Hai Jiang 0003, Kuanching Li
Frontiers Inf. Technol. Electron. Eng.3
2016 P-CP-ABE: Parallelizing Ciphertext-Policy Attribute-Based Encryption for clouds
abstract
Recently, cloud storage has become quite attractive due to its elasticity, availability and scalability. However, the security issue has started to prevent public clouds from getting even more popular. Traditional encryption algorithms (both symmetric and asymmetric ones) fail to help achieve effective secure cloud storage due to their severe issues such as complex key management and heavy redundancy. Ciphertext-Policy Attribute Based Encryption (CP-ABE) scheme overcomes the aforementioned issues and provides fine-grained access control as well as deduplication features. CP-ABE has becomes a possible solution to cloud storage. However, its high complexity has prevented it from being widely adopted. This paper proposes a P-CP-ABE scheme to parallelize CP-ABE and port it to multi-core architecture machines. Major performance bottlenecks such as key management and encryption/decryption process are identified and accelerated. New AES encryption operation mode is adopted for further performance gains. Experimental results have demonstrated its effectiveness.
Lifeng Li, Xiaowan Chen, Hai Jiang 0003, Kuanching Li
SNPD3
2015 MR-Graph: A Customizable GPU MapReduce
abstract
The MapReduce programming model has been widely used in Big Data and Cloud applications. Criticism on its inflexibility when being applied to complicated scientific applications recently emerges. Several techniques have been proposed to enhance its flexibility. However, some of them exert special requirements on applications, while others fail to support the increasingly popular coprocessors, such as Graphics Processing Unit (GPU). In this paper, we propose MR-Graph, a customizable and unified framework for GPU-based MapReduce, which aims to improve the flexibility, scalability and performance of MapReduce. MR-Graph addresses the limitations and restrictions of the traditional MapReduce execution paradigm. The three execution modes integrated in MR-Graph facilitates users to write their applications in a more flexible fashion by defining a Map and Reduce function call graph. MR-Graph efficiently explores the memory hierarchy in GPUs to reduce the data transfer overhead between execution stages and accommodate big data applications. We have implemented a prototype of MR-Graph and experimental results show the effectiveness of using MR-Graph for flexible and scalable GPU-based MapReduce computing.
Zhi Qiao 0001, Shuwen Liang, Hai Jiang 0003, Song Fu
CSCloud3
2015 A customizable MapReduce framework for complex data-intensive workflows on GPUs
abstract
The MapReduce programming model has been widely used in big data and cloud applications. Criticism on its inflexibility when being applied to complicated scientific applications recently emerges. Several techniques have been proposed to enhance its flexibility. However, some of them exert special requirements on applications, while others fail to support the increasingly popular coprocessors, such as Graphics Processing Unit (GPU). In this paper, we propose MR-Graph, a customizable and unified framework for GPU-based MapReduce, which aims to improve the flexibility and performance of MapReduce. MR-Graph addresses the limitations and restrictions of the traditional MapReduce execution paradigm. The three execution modes integrated in MR-Graph facilitates users to write their applications in a more flexible fashion by defining a Map and Reduce function call graph. MR-Graph efficiently explores the memory hierarchy in GPUs to reduce the data transfer overhead between execution stages and accommodate big data applications.We have implemented a prototype of MR-Graph and experimental results show the effectiveness of using MR-Graph for flexible and scalable GPU-based MapReduce computing.
Zhi Qiao 0001, Shuwen Liang, Hai Jiang 0003, Song Fu
IPCCC3
2015 Portable parallelized blowfish via RenderScript
abstract
The recent rise in the popularity of mobile computing has brought the attention of mobile security to the forefront. As users depend more on tablets and smartphones, sensitive data is left to be secured using devices with vastly weaker resources than a typical computer. As mobile technology matures, the industry is starting to provide devices with multiple CPU cores in addition to other coprocessors such as GPUs. By using RenderScript, a new language technology on the Android platform, we hope to utilize the power of parallelism to increase the efficiency of the Blowfish encryption algorithm, while at the same time leveraging the power of RenderScript's heterogenous execution to cope with the quickly changing mobile architectures in order to make the use of data encryption more feasible on a mobile platform. Experimental results demonstrate the effectiveness of RenderScript.
Spencer Davis, Brandon Jones, Hai Jiang 0003
SNPD3
2015 A secure and scalable storage system for aggregate data in IoT
Hai Jiang 0003, Kuanching Li, Young-Sik Jeong
Future Gener. Comput. Syst.1
2014 GPU-in-Hadoop: Enabling MapReduce across distributed heterogeneous platforms
abstract
As the size of high performance applications increases, four major challenges including heterogeneity, programmability, failure resilience, and energy efficiency have arisen in the underlying distributed systems. To tackle with all of them without sacrificing performance, traditional approaches in resource utilization, task scheduling and programming paradigm should be reconsidered. As Hadoop has handled data-intensive applications well in Clouds, GPU has demonstrated its acceleration effectiveness for computation-intensive ones. This paper intends to integrate Hadoop with CUDA to exploit both CPU and GPU resources. Hadoop will schedule MapReduce's Map and Reduce functions across multiple nodes, whereas CUDA code helps accelerate them further on local GPUs. All available heterogeneous computational power will be utilized. MapReduce in Hadoop will ease the programming task by hiding communication details. Hadoop Distributed File System will help achieve data-level fault resilience. GPU's energy efficiency characteristics help reduce the power consumption of the whole system. To achieve Hadoop and GPU integration, four approaches including Jcuda, JNI, Hadoop Streaming, and Hadoop Pipes, have been accomplished. Experimental results have demonstrated their effectiveness.
Juanjuan Li, Erikson Hardesty, Hai Jiang 0003, Kuanching Li
ICIS4
2014 Analysis and acceleration of NTRU lattice-based cryptographic system
abstract
Lattice based cryptography is attractive for its quantum computing resistance and efficient encryption/decryption process. However, the big data problem has perplexed lattice based cryptographic systems with the slow processing speed. This paper intends to analyze one of the major lattice-based cryptographic systems, Nth-degree truncated polynomial ring (NTRU), and accelerate its execution with Graphic Processing Unit (GPU) for acceptable processing performance. Three strategies, including single GPU with zero copy, single GPU with data transfer, and multi-GPU versions are proposed. GPU computing techniques such as stream and zero copy are applied to overlap the computation and communication for possible speedup. Experimental results have demonstrated the effectiveness of GPU acceleration of NTRU. As the number of involved devices increases, better NTRU performance will be achieved.
Tianyu Bai, Spencer Davis, Juanjuan Li, Hai Jiang 0003
SNPD4
2014 Survey of attribute based encryption
abstract
In Attribute Based Encryption, a set of descriptive attributes is used as an identity to generate a secret key, as well as serving as the access structure that performs access control. It successfully integrates Encryption and Access Control and is ideal for sharing secrets among groups, especially in a Cloud environment. Most developed ABE schemes support key-policy or ciphertext-policy access control in addition to other features such as decentralized authority, efficient revocation and key delegation. This paper surveys mainstream papers, analyzes main features for desired ABE systems, and classifies them into different categories. With this high-level guidance, future researchers can treat these features as individual modules and select related ones to build their ABE systems on demand.
Zhi Qiao 0001, Shuwen Liang, Spencer Davis, Hai Jiang 0003
SNPD4
2013 Towards constructing application-level GPU computation states
abstract
Computation state construction is an indispensable step to achieve fault tolerance and computation mobility for scientific applications by saving and restoring the state during program execution. However, there is no effective state construction scheme yet due to the GPU's batch-mode execution manner as the GPU takes on a larger role in high performance computing. The GPU's complex memory hierarchy means the states are scattered in different memory locations that are difficult to fetch. Programs that are running in parallel make the states difficult to construct for each thread. The paper proposes an application-level computation state construction scheme to support GPU programs. A precompiler and run-time support module are developed to construct and save states in the CPU system memory dynamically. Memory blocks are registered, and new data structures are proposed to save and restore the computation states represented by variables and pointers in the GPU. Secondary storage can be utilized for scalability and long-term fault tolerance.
Xinyuan Guo, Hai Jiang 0003, Kuanching Li
ICIS3
2013 A Checkpoint/Restart Scheme for CUDA Applications with Complex Memory Hierarchy
abstract
Checkpoint/restart has been an effective mechanism to achieve fault tolerance for many scientific applications. However, as GPU becomes a much bigger role in high performance computing, there is no effective checkpoint/restart scheme yet due to GPU's batch-mode execution manner. The paper proposes an application-level checkpoint/restart scheme to save and restore GPU computation states. A precompiler and run-time support module are developed to construct and save states in CPU system memory dynamically. Secondary storage can be utilized for scalability and long-term fault tolerance. CUDA applications with complicated memory use are support as well. Experimental results have demonstrated the effectiveness of the proposed scheme.
Xinyuan Guo, Hai Jiang 0003, Kuanching Li
SNPD3
2012 Preemption of a CUDA Kernel Function
abstract
As graphics processing units (GPUs) gain adoption as general purpose parallel compute devices, several key problems need to be addressed in order for their use to become more practical and more user friendly. One such problem is special functions designed to execute on GPUs called kernel functions are non-preempt able. Once the kernel is issued to the GPU it will remain there till either execution finishes or it is killed. If the kernel uses all the execution units of the GPU, then no other kernels are able to be executed. This paper proposes a way to apply preemption to the executing kernel function. The kernel at some point in its execution will be able to save its state, halt execution, and free up the GPU's execution units for other kernels to run. After a given amount of time the halted kernel will be able to regain control of the GPU and complete its execution as if it never was halted in the first place. Experimental results have demonstrated the effectiveness of the proposed scheme.
Jon Calhoun 0001, Hai Jiang 0003
SNPD2
2012 A Secure Distributed File System Based on Revised Blakley's Secret Sharing Scheme
abstract
To support cloud storage effectively, a Distributed File System (DFS) should be well-rounded with excellent features in multiple major aspects and without significant drawbacks. The main design goals of a DFS in Clouds include security, reliability and scalability. Traditionally, cryptography, data duplication and powerful machines are common approaches to support DFS. However, the success of such a DFS will depend on cumbersome key management, large storage and costly infrastructure, respectively. This paper intends to revise Blakley's secret sharing and apply it to a DFS for both security and reliability without sacrificing the scalability in performance too much. A DFS is deployed with GPU (Graphics Processing Unit) as an acceleration option to tackle with scalability issue further. Experimental results have demonstrated the effectiveness of the new DFS.
Yi Chen 0018, Hai Jiang 0003, Laurence T. Yang, Kuanching Li
TrustCom3
2010 Pitcher: Enabling Distributed Parallel Computing with Automatic Thread and Data Assignments
abstract
Parallel and distributed systems have intended to exploit local and remote multi-cores and multiprocessors for high performance. However, without popular single-image distributed operating systems, few systems can utilize all resources automatically and effectively. This paper proposed an MPI-like middleware, Pitcher, to distribute multithreaded parallel workloads across networked computers without user involvement. Fine-grained computational units, threads, are regrouped into bundles based on data locality and scheduling requests. Such thread bundles are treated as load distribution units. Unlike MPI, data and synchronization variables in Pitcher will be automatically partitioned and distributed. A preprocessor transforms the source code so that programmers stick to shared virtual address space programming paradigm whereas the runtime support module dispatches thread bundles and data based on locality. The experimental results demonstrate the effectiveness of Pitcher.
Chonglei Mei, Hai Jiang 0003, Jeff Jenness
AINA2
2009 MTS: Multiresolution Thread Selection for Parallel Workload Distribution
Chonglei Mei, Hai Jiang 0003, Jeff Jenness
GPC2
2009 MM-DSM: Multi-threaded Multi-home Distributed Shared Memory Systems
abstract
Most traditional Distributed Shared Memory (DSM) systems support data sharing in multi-process applications. This paper proposed a Multi-threaded Multi-home DSM system (MM-DSM) to support both data sharing and computation synchronization in multi-threaded applications whose threads are grouped into bundles and distributed across multiple computers for parallel execution. Globally shared data are rearranged and assigned to different thread bundles based on their access patterns. As thread bundles move around, their hosting nodes will act as the homes of the associated data blocks to reduce communication cost. Programmers can still stick to the shared memory programming paradigm whereas data consistency, distributed lock, false sharing and multiple writes are taken care of by MM-DSM. Experimental results demonstrate its effectiveness and correctness.
Chonglei Mei, Hai Jiang 0003, Jeff Jenness
ISPA2
2008 Parallel and Distributed Particle Collision Simulation with Decentralized Control
Hai Jiang 0003, Hung-Chi Su, Bin Zhang 0005, Jeff Jenness
GPC2
2007 Flexible Speculative Parallelization of Many-Particle Collision Simulation
Hung-Chi Su, Hai Jiang 0003, Bin Zhang 0005
CAINE2
2007 Speculative and distributed simulation of many-particle collision systems
abstract
As problem size increases dramatically and quickly, many scientific computing applications, such as molecular dynamics and many-particle collision simulations, exhibit high demand on data capacity and computability of the hosting computer systems. Scalable solutions are on demand. Neither well optimized sequential algorithms nor expensive shared memory parallel computers could be the ultimate choice for such large-scale problems. This paper intends to solve this scalability problem by taking advantage of the availability of commodity computers. Aggregated memory and networked computers satisfy the capacity and computability requests. To overcome the strong event dependency nature in particle collision simulations, a speculative execution scheme is proposed to exploit parallelism. Performance analysis and experiments are provided to show the balance between scalability and performance gains.
Hai Jiang 0003, Hung-Chi Su, Bin Zhang 0005, Jeff Jenness
ICPADS2
2004 Experiments with Parallelizing Tribology Simulations
Vipin Chaudhary, William L. Hase, Hai Jiang 0003, Darshan Thaker
J. Supercomput.3
2003 Thread Migration/Checkpointing for Type-Unsafe C Programs
Hai Jiang 0003, Vipin Chaudhary
HiPC1
2003 Data Conversion for Process/Thread Migration and Checkpointing
abstract
Process/thread migration and checkpointing schemes support load balancing, load sharing and fault tolerance to improve application performance and system resource usage on workstation clusters. To enable these schemes to work in heterogeneous environments, we have developed an application-level migration and checkpointing package, MigThread, to abstract computation states at the language level for portability. To save and restore such states across different platforms, we propose a novel "receiver makes right" (RMR) data conversion method, called coarse-grain tagged RMR (CGT-RMR), for efficient data marshalling and unmarshalling. Unlike common data representation standards, CGT-RMR does not require programmers to analyze data types, flatten aggregate types, and encode/decode scalar types explicitly within programs. With help from MigThread's type system, CGT-RMR assigns a tag to each data type and converts nonscalar types as a whole. This speeds up the data conversion process and eases the programming task dramatically, especially for the large data trunks common to migration and checkpointing. Armed with this "plug-and-play" style data conversion scheme, MigThread has been ported to work in heterogeneous environments. Some microbenchmarks and performance measurements within the SPLASH-2 suite are given to illustrate the efficiency of the data conversion process
Hai Jiang 0003, Vipin Chaudhary, John Paul Walters
ICPP1
2002 On Improving Thread Migration: Safety and Performance
Hai Jiang 0003, Vipin Chaudhary
HiPC1