Lei Chai

dblp:02/4099 · DBLP profile ↗
← Back
29ranked-venue papers
13as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 4 first-authorArtificial intelligence and machine learning · 9 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Attribution Analysis-based Concept Alignment: A Human-in-the-loop Data Debugging Framework
abstract
Ensuring consistently high-quality training data is essential for developing reliable machine learning systems. Recent research demonstrates that incorporating human supervision into training set debugging effectively improves model performance, especially for text classification tasks. However, such methods often prove inapplicable to image understanding tasks, where inherently unstructured pixel data presents challenges in understanding and correcting biases. Inspired by human-AI alignment, we introduce AACA (Attribution Analysis-based Concept Alignment), a human-in-the-loop framework that mitigates bias in the training set by aligning the concepts used by humans and AI during the decision-making process. Specifically, AACA comprises two primary stages: interpretable data bug discovery and targeted data augmentation. During the data bug discovery stage, AACA identifies confounded and valid concepts to explain why prediction failure occurs and what concept the model should focus, using interpretability methods and human annotation. In the stage of targeted data augmentation, AACA adopts these concept-level attributions as clues to synthesize debugging instances via text-to-image generative model. The initial model is then retrained on the augmented set to correct prediction failures. Comparative experiments conducted on crowdsourced annotations and real-world datasets demonstrate that AACA can accurately identifies data bugs and effectively repairs prediction failures, thereby significantly improving prediction performance.
Lei Chai, Hailong Sun 0001, Jingxuan Xu
AAAI1
2026 Misclassification-Aware Robust Learning from Multiple Human Labelers (Student Abstract)
abstract
Adversarial training is an effective technique for enhancing the robustness of deep neural networks (DNNs). Prior research shows that misclassified examples influence final adversarial robustness much more than correctly classified examples. Ignoring this difference during training can hurt model performance. In crowdsourcing, varying annotator expertise causes noisy, inconsistent labels. As a result, it is hard to distinguish misclassified and correctly classified examples using only provided annotations. Thus, how to use the reliability and discrepancy between these example types to improve robustness within adversarial learning remains a critical but underexplored issue. In this work, we first explore how misclassified and correctly classified examples affect learning from crowds (LFC) in adversarial environments. Then, we formulate the problem of misclassification-aware robust learning from multiple human labelers as a bilevel min-max problem. After that, we introduce MALC, a new approach to make classifiers more robust to adversarial examples via iterative adversarial example generation and parameter estimation. We conduct an extensive evaluation of the proposed MALC, showing that MALC can outperform the state-of-the-art LFC methods in both white-box and black-box settings.
Zuoyuehe Wang, Chicheng Ma, Lei Chai, Yongqiang Yang, Jingzheng Li
AAAI4
2026 Learning and fusing the rasterized and serialized point cloud for 3D semantic segmentation in railway scenes
Ning Sun 0005, Jixin Liu 0001, Lei Chai, Cong Wu 0007
Expert Syst. Appl.4
2026 Quality control in open-ended crowdsourcing: a survey
abstract
Abstract Crowdsourcing provides a flexible approach for leveraging human intelligence to solve large-scale problems, gaining widespread acceptance in domains like intelligent information processing, social decision-making, and crowd ideation. However, the uncertainty of participants significantly compromises the answer quality, sparking substantial research interest. Existing surveys predominantly concentrate on quality control in Boolean tasks, which are generally formulated as simple label classification, ranking, or numerical prediction. Ubiquitous open-ended tasks like question-answering, translation, and semantic segmentation have not been sufficiently discussed. These tasks usually have large to infinite answer spaces and non-unique acceptable answers, posing significant challenges for quality assurance. This survey focuses on quality control methods applicable to open-ended tasks in crowdsourcing. We propose a two-tiered framework to categorize related works. The first tier presents a comprehensive overview of the quality model, covering essential aspects including tasks, workers, answers, and the system. The second tier further refines this classification by breaking it down into more detailed categories: ‘quality dimensions’, ‘evaluation metrics’, and ‘design decisions’. This breakdown provides deeper insights into the internal structure of the quality control model for each aspect. We thoroughly investigate how these quality control methods are implemented in state-of-the-art works and discuss key challenges and potential future research directions.
Lei Chai, Hailong Sun 0001
Frontiers Comput. Sci.1
2026 3D semantic segmentation for railway scenes via heterogeneous multimodal alignment and distillation
Ning Sun 0005, Maomao Sun, Jixin Liu 0001, Lei Chai, Cong Wu 0007
Multim. Syst.5
2026 Robust Multimodal Sentiment Analysis Based on Adaptive Information Distillation and Adversarial Learning
Ning Sun 0005, Weiliang Zhang, Wenming Zheng, Jixin Liu 0001, Lei Chai, Cong Wu 0007
IEEE Trans. Affect. Comput.5
2025 Multimodal Sentimental Privileged Information Embedding for Improving Facial Expression Recognition
abstract
Facial expression recognition (FER) has always been one of the key task in affective computing. Over the years, researchers have worked to improve the performance of FER by designing models with more powerful feature extraction, embedding attention mechanism, and reconstructing missing information, etc. Different from the paradigms above, we attempt to improve FER performance by using multimodal sentiment data, such as audio and text, as privileged information (PI) for facial images. To this end, a multimodal privileged information embedded facial expression recognition network (MPI-FER) is proposed in this paper. During the training phase, this model achieves the PI embedding of multimodal data for FER by developing cross-modality translation between multimodal sentiment data. During the test phase, input images alone are sufficient for the model inference to accomplish the FER task input. The MPI-FER is a large-scale, heterogeneous deep neural network. To achieve effective training of this model with limited training samples, we design a multi-stage training strategy of module-wise pre-training followed by end-to-end fine-tuning. In addition, a strategy of filling the multimodal sentiment quaternion is proposed for implementing our method on a facial expression database consisting only of face images. We conducted extensive experiments to evaluate the proposed method on two databases of multimodal sentiment analysis (CH-SIMS and CMU-MOSI) and two databases of FER in the wild (RAF-DB and AffectNet). The results show that embedding multimodal sentiment data as privileged information into the FER task based on face images can significantly improve the accuracy of FER. Furthermore, by only using image in the test phase, the proposed method can achieve better results of multimodal sentiment analysis than those methods achieved by using multimodal sentimental data fusion.
Ning Sun 0005, Changwei You, Wenming Zheng, Jixin Liu 0001, Lei Chai, Haian Sun
IEEE Trans. Affect. Comput.5
2024 RA3: A Human-in-the-loop Framework for Interpreting and Improving Image Captioning with Relation-Aware Attribution Analysis
abstract
Interpreting model behavior is crucial for model evaluation and optimization. Recent research demonstrates that incorporating human intelligence into the learning process effectively improve the interpretability and performance of the machine learning models, especially for simple classification tasks. However, the image captioning task has not received much attention. Such complex sequential tasks generally contain semantic relationships between different concepts, which pose challenges for interpreting model behavior and developing optimization methods. In this paper, we present RA3(Relation-Aware Attribution Analysis), a human-in-the-loop framework, for improving the interpretability, and further boosting the performance of the image captioning model. Specifically, we first engage human participants in two types of annotation tasks to identify what the model actually focuses on (model attribution) and what it should focus on (human rationale) at the conceptual level, supported by machine learning interpretability methods. Then, we identify and filter hard instances based on relation-aware model attribution for both validating the quality of the explanation and eliminating low-quality captions (this process is also considered as a kind of data debugging). We subsequently designed an explanation loss that penalizes the difference between model attribution and human rationale to optimize the model's behavior for improving caption quality. Through extensive experiments on crowdsourced annotations and MSCOCO, the experiment results indicate that the explanations produced by RA3can accurately describe the model's behavior, effectively identify difficult instances, and significantly improve the caption quality.
Lei Chai, Hailong Sun 0001, Jingzheng Li
ICDE1
2024 Target Structure Learning Framework for Unsupervised Multi-Class Domain Adaptation
abstract
Unsupervised multi-class domain adaptation (multi-class UDA) has recently been proposed to fill the gap between empirically practical methods for multi-class classification and well-founded theory with the setting of binary classification. Nevertheless, the multi-class UDA methods use model predictions to characterize the disagreement of multi-class scoring hypotheses, which is used to optimize the divergence between domain distributions. Such self-training manner may bring inaccurate model predictions, which would damage the target structure due to the absence of labels of the target domain, leading to sub-optimal performance. On the other hand, this disagreement between multi-class scoring hypotheses does not involve the relationships among all of the multiple classes. It causes that multi-class UDA cannot properly connect the advanced practical UDA methods that consider class-conditional distribution alignment. Thus, we propose to exploit the target structure information and then incorporate it into multi-class UDA to achieve class-conditional distribution alignment. We theoretically and experimentally explain the importance of accurate target structure information to reduce the expected error on the target domain. Notably, our method achieves state-of-the-art results on three commonly-used benchmarks with different scales. In addition, using the target structure information, we propose a variant to cope with noisy open-world source domains such as noisy labels and out-of-distribution samples, enhancing the robustness of our method. The source code is available at https://github.com/jingzhengli/Multi_Class_UDA .
Jingzheng Li, Hailong Sun 0001, Lei Chai, Jiyi Li
ACM Trans. Multim. Comput. Commun. Appl.3
2023 Joint sensor registration and multi-object tracking with PHD filter in distributed multi-sensor networks
Lei Chai, Wei Yi 0002, Lingjiang Kong
Signal Process.1
2022 Pyramid Attention For Source Code Summarization
abstract
This paper presents a multi-granularity method for source code summarization, which generates a concise functional description for the given code snippet. We notice that skilled programmers write and read source codes hierarchically and pay close attention to conceptual entities like statements, tokens, sub-tokens, and the mapping relations between them. The entities have specific emphasis according to their granularities, e.g., statements in coarse-granularity reveal the global logical semantics of code, and the sub-tokens in fine-granularity are more related to the textual semantics. Driven by this observation, we demonstrate that a multi-granularity formulation incorporating these conceptual entities benefit the code summarization task. Concretely, the source code is transformed into a pyramidal representation, and then a pyramid attention mechanism is applied for efficient feature aggregation among different hierarchies in it. We instantiate our multi-granularity method using the proposed pyramid attention and name it PA-former (Pyramid Attention transformer). We evaluated it on two source code summarization benchmarks where it surpasses the prior works and achieves new state-of-the-art results. Our code and data are available at https://github.com/leichainju/pa-former.
Lei Chai
NeurIPS1
2022 Stock movement prediction via gated recurrent unit network based on reinforcement learning with incorporated attention mechanisms
Hongfeng Xu, Lei Chai, Zhiming Luo, Shaozi Li
Neurocomputing2
2022 An error consistency based approach to answer aggregation in open-ended crowdsourcing
abstract
Crowdsourcing plays a vital role in today’s AI industry. However, existing crowdsourcing research mainly focuses on those simple tasks that are often formulated as label classification, while complex open-ended tasks such as question answering and translation have not received much attention. Such tasks usually have open solution spaces and non-unique true answers, which pose great challenges for designing effective crowdsourcing algorithms. In this work, we are concerned specifically with complex text annotation crowdsourcing tasks, where each answer of a task is in the form of free text. We propose an error consistency-based approach to inferring a satisfying result from a set of open-ended answers. First, each answer is represented with two vectors that capture the local word collocation and the global sentence semantics respectively. Second, the true answer is approximated by the sum of the answer vectors weighted by the reciprocals of their respective errors. Third, an algorithm called AEC (Aggregation based on Error Consistency) is designed to infer the aggregated result by maximizing the consistency of the errors of an answer in two vector spaces. Experimental results on two datasets demonstrate the effectiveness of our approach.
Lei Chai, Hailong Sun 0001, Zizhe Wang
Inf. Sci.1
2020 A multi-source heterogeneous data analytic method for future price fluctuation prediction
Lei Chai, Hongfeng Xu, Zhiming Luo, Shaozi Li
Neurocomputing1
2020 Stock movement predictive network via incorporative attention mechanisms based on tweet and historical prices
Hongfeng Xu, Lei Chai, Zhiming Luo, Shaozi Li
Neurocomputing2
2020 The multiple model multi-Bernoulli filter based track-before-detect using a likelihood based adaptive birth distribution
Lei Chai, Lingjiang Kong, Suqi Li, Wei Yi 0002
Signal Process.1
2019 A Distributed PHD Filter for On-line Joint Sensor Registration and Multi-target Tracking
Lei Chai, Wei Yi 0002, Lingjiang Kong
FUSION1
2008 Designing an Efficient Kernel-Level and User-Level Hybrid Approach for MPI Intra-Node Communication on Multi-Core Systems
abstract
The emergence of multi-core processors has made MPI intra-node communication a critical component in high performance computing. In this paper, we use a three-stepmethodology to design an efficient MPI intra-node communication scheme from two popular approaches: shared memory and OS kernel-assisted direct copy. We use an Intel quad-core cluster for our study. We first run microbenchmarks to analyze the advantages and limitations of these two approaches, including the impacts of processor topology, communication buffer reuse, process skew effects, and L2 cache utilization. Based on the results and the analysis, we propose topology-aware and skew-aware thresholds to build an optimized hybrid approach. Finally, we evaluate the impact of the hybrid approach on MPI collective operations and applications using IMB, NAS, PSTSWM, and HPL benchmarks. We observe that the optimized hybrid approach can improve the performance of MPI collective operations by up to 60%, and applications by up to 17%.
Lei Chai, Ping Lai, Hyun-Wook Jin, Dhabaleswar K. Panda 0001
ICPP1
2007 Understanding the Impact of Multi-Core Architecture in Cluster Computing: A Case Study with Intel Dual-Core System
abstract
Multi-core processors are growing as a new industry trend as single core processors rapidly reach the physical limits of possible complexity and speed. In the new Top500 supercomputer list, more than 20% processors belong to the multi-core processor family. However, without an in-depth study on application behaviors and trends on multi-core clusters, we might not be able to understand the characteristics of multi-core cluster in a comprehensive manner and hence not be able to get optimal performance. In this paper, we take on these challenges and design a set of experiments to study the impact of multi-core architecture on cluster computing. We choose to use one of the most advanced multi-core servers, Intel Bensley system with Woodcrest processors, as our evaluation platform, and use benchmarks including HPL, NAMD, and NAS as the applications to study. From our message distribution experiments, we find that on an average about 50% messages are transferred through intra-node communication, which is much higher than intuition. This trend indicates that optimizing intra- node communication is as important as optimizing inter- node communication in a multi-core cluster. We also observe that cache and memory contention may be a potential bottleneck in multi-core clusters, and communication middleware and applications should be multi-core aware to alleviate this problem. We demonstrate that multi-core aware algorithm, e.g. data tiling, improves benchmark execution time by up to 70%. We also compare the scalability of a multi-core cluster with that of a single-core cluster and find that the scalability of the multi-core cluster is promising.
Lei Chai, Qi Gao 0004, Dhabaleswar K. Panda 0001
CCGRID1
2007 Lightweight kernel-level primitives for high-performance MPI intra-node communication over multi-core systems
abstract
Modern processors have multiple cores on a chip to overcome power consumption and heat dissipation issues. As more and more compute cores become available on a single node, it is expected that node-local communication will play an increasingly greater role in overall performance of parallel applications such as MPI applications. It is therefore crucial to optimize intra-node communication paths utilized by MPI libraries. In this paper, we propose a novel design of a kernel extension, called LiMIC2, for high-performance MPI intra-node communication over multi-core systems. LiMIC2 can minimize the communication overheads by implementing lightweight primitives and provide portability across different interconnects and flexibility for performance optimization. Our performance evaluation indicates that LiMIC2 can attain 80% lower latency and more than three times improvement in bandwidth. Also the experimental results show that LiMIC2 can deliver bidirectional bandwidth greater than 11GB/s.
Hyun-Wook Jin, Sayantan Sur, Lei Chai, Dhabaleswar K. Panda 0001
CLUSTER3
2007 Efficient asynchronous memory copy operations on multi-core systems and I/OAT
abstract
Bulk memory copies incur large overheads such as CPU stalling (i.e., no overlap of computation with memory copy operation), small register-size data movement, cache pollution, etc. Asynchronous copy engines introduced by Intelpsilas I/O Acceleration Technology help in alleviating these overheads by offloading the memory copy operations using several DMA channels. However, the startup overheads associated with these copy engines such as pinning the application buffers, posting the descriptors and checking for completion notifications, limit their overlap capability. In this paper, we propose two schemes to provide complete overlap of memory copy operation with computation by dedicating the critical tasks to a single core in a multi-core system. In the first scheme, MCI (Multi-Core with I/OAT), we offload the memory copy operation to the copy engine and onload the startup overheads to the dedicated core. For systems without any hardware copy engine support, we propose a second scheme, MCNI (Multi-Core with No I/OAT) that onloads the memory copy operation to the dedicated core. We further propose a mechanism for an application-transparent asynchronous memory copy operation using memory protection. We analyze our schemes based on overlap efficiency, performance and associated overheads using several micro-benchmarks and applications. Our microbenchmark results show that memory copy operations can be significantly overlapped (up to 100%) with computation using the MCI and MCNI schemes. Evaluation with MPI-based applications such as IS-B and PSTSWM-small using the MCNI scheme show up to 4% and 5% improvement, respectively, as compared to traditional implementations. Evaluations with data-centers using the MCI scheme show up to 37% improvement compared to the traditional implementation. Our evaluations with gzip SPEC benchmark using application-transparent asynchronous memory copy show a lot of potential to use such mechanisms in several application domains.
Karthikeyan Vaidyanathan, Lei Chai, Wei Huang 0003, Dhabaleswar K. Panda 0001
CLUSTER2
2007 Designing NFS with RDMA for Security, Performance and Scalability
abstract
NFS has traditionally used TCP or UDP as the underlying transport. However, the overhead of these stacks has limited both the performance and scalability of NFS. Recently, high-performance network such as InfiniBand have been deployed. These networks provide low latency of a few microseconds and high bandwidth for large messages up to 20 Gbps. Because of the unique characteristics of NFS protocols, previous designs of NFS with RDMA were unable to exploit the improved bandwidth of networks such as InfiniBand. Also, they leave the server open to attacks from malicious clients. In this paper, we discuss the design principles for implementing NFS/RDMA protocols. We propose, implement and evaluate an alternate design for NFS/RDMA on InfiniBand, which can significantly improve the security of the server, compared to the previous design. In addition, we evaluate the performance bottlenecks of using RDMA operations in NFS protocols and propose strategies and designs that tackle these overheads. With the best of these strategies and designs, we demonstrate throughput of 700 MB/s on the OpenSolaris NFS/RDMA design and 900 MB/s on the Linux design and an application level improvement in performance of up to 50%. We also evaluate the scalability of the RDMA transport in a multi-client setting, with a RAID array of disks. Our design has been integrated into the OpenSolaris kernel.
Ranjit Noronha, Lei Chai, Thomas Talpey, Dhabaleswar K. Panda 0001
ICPP2
2007 Designing Efficient Asynchronous Memory Operations Using Hardware Copy Engine: A Case Study with I/OAT
abstract
Memory copies for bulk data transport incur large overheads due to CPU stalling, small register-size data movement, etc. Intel's I/O Acceleration Technology offers an asynchronous memory copy engine in kernel space which alleviates such overheads. In this paper, we propose a set of designs for asynchronous memory operations in user space for both single process (as an offloaded memcpy()) and lPC using the copy engine. We analyze our design based on overlap efficiency, performance and cache utilization. Our microbenchmark results show that using the copy engine for performing memory copies can achieve close to 87% overlap with computation. Further, the copy engine improves the copy latency of bulk memory data transfers by 50% and avoids cache pollution effects. With the emergence of multi-core architectures, the support for asynchronous memory operations holds a lot of promise in reducing the gap between the memory and processor performance.
Karthikeyan Vaidyanathan, Wei Huang 0003, Lei Chai, Dhabaleswar K. Panda 0001
IPDPS3
2006 MPI over uDAPL: Can High Performance and Portability Exist Across Architectures?
abstract
Looking at the TOP 500 list of supercomputers we can see that different architectures and networking technologies appear on the scene from time to time. The networking technologies are also changing along with the advances of processor technologies. While the hardware has been constantly changing, parallel applications written in different paradigms have remained largely unchanged. With MPI being the most popular parallel computing standard, it is crucial to have an MPI implementation portable across different networks and architectures. It is also desirable to have such an MPI deliver high performance. In this paper we take on this challenge. We have designed an MPI with both portability and portable high performance using the emerging uDAPL interface. We present the design alternatives and a comprehensive performance evaluation of this new design. The results show that this design can improve the startup time and communication performance by 30% compared with our previous work. It also delivers the same good performance as MPI implemented over native APIs of the underlying interconnect. We also present a multistream MPI design which aims to achieve high bandwidth across networks and operating systems. Experimental results on Solaris show that the multi-stream design can improve bandwidth over InfiniBand by 30%, and improve the application performance by up to 11%.
Lei Chai, Ranjit Noronha, Dhabaleswar K. Panda 0001
CCGRID1
2006 Designing High Performance and Scalable MPI Intra-node Communication Support for Clusters
abstract
As new processor and memory architectures advance, clusters start to be built from larger SMP systems, which makes MPI intra-node communication a critical issue in high performance computing. This paper presents a new design for MPI intra-node communication that aims to achieve both high performance and good scalability in a cluster environment. The design distinguishes small and large messages and handles them differently to minimize the data transfer overhead for small messages and the memory space consumed by large messages. Moreover, the design utilizes the cache efficiently and requires no locking mechanisms to achieve optimal performance even with large system size. This paper also explores various optimization strategies to reduce polling overhead and maintain data locality. We have evaluated our design on NUMA and dual core NUMA (non-uniform memory access) systems. The experimental results on NUMA system show that the new design can improve MPI intra-node latency by up to 35% and bandwidth by up to 50% compared to MVAPICH. While running the bandwidth benchmark, the measured L2 cache miss rate is reduced by half. The new design also improves the performance of MPI collective calls by up to 25%. The results on dual core NUMA system show that the new design can achieve 0.48 musec in CMP latency
Lei Chai, Albert Hartono, Dhabaleswar K. Panda 0001
CLUSTER1
2006 Efficient SMP-aware MPI-level broadcast over InfiniBand's hardware multicast
abstract
Most of the high-end computing clusters found today feature multi-way SMP nodes interconnected by an ultra-low latency and high bandwidth network. InfiniBand is emerging as a high-speed network for such systems. InfiniBand provides a scalable and efficient hardware multicast primitive to efficiently implement many MPI collective operations. However, employing hardware multicast as the communication method may not perform well in all cases. This is true especially when more than one process is running per node. In this context, shared memory channel becomes the desired communication medium within the node as it delivers latencies which are of an order of magnitude lower than the inter-node message latencies. Thus, to deliver optimal collective performance, coupling hardware multicast with shared memory channel becomes necessary. In this paper we propose mechanisms to address this issue. On a 16-node 2-way SMP cluster, the Leader-based scheme proposed in this paper improves the performance of the MPI/spl I.bar/Bcast operation by a factor of as much as 2.3 and 1.8 when compared to the point-to-point and original solution employing only hardware multicast. We have also evaluated our designs on NUMA based system and obtained a performance improvement of 1.7 using our designs on 2-node 4-way system. We also propose a dynamic attach policy as an enhancement to this scheme to mitigate the impact of process skew on the performance of the collective operation.
Amith R. Mamidala, Lei Chai, Hyun-Wook Jin, Dhabaleswar K. Panda 0001
IPDPS2
2006 Shared receive queue based scalable MPI design for InfiniBand clusters
abstract
Clusters of several thousand nodes interconnected with InfiniBand, an emerging high-performance interconnect, have already appeared in the Top 500 list. The next-generation InfiniBand clusters are expected to be even larger with tens-of-thousands of nodes. A high-performance scalable MPI design is crucial for MPI applications in order to exploit the massive potential for parallelism in these very large clusters. MVAPICH is a popular implementation of MPI over InfiniBand based on its reliable connection oriented model. The requirement of this model to make communication buffers available for each connection imposes a memory scalability problem. In order to mitigate this issue, the latest InfiniBand standard includes a new feature called shared receive queue (SRQ) which allows sharing of communication buffers across multiple connections. In this paper, we propose a novel MPI design which efficiently utilizes SRQs and provides very good performance. Our analytical model reveals that our proposed designs take only 1/10/sup th/ the memory requirement as compared to the original design on a cluster sized at 16,000 nodes. Performance evaluation of our design on our 8-node cluster shows that our new design was able to provide the same performance as the existing design while requiring much lesser memory. In comparison to tuned existing designs our design showed a 20% and 5% improvement in execution time of NAS Benchmarks (Class A) LU and SP, respectively. The high performance Linpack was able to execute a much larger problem size using our new design, whereas the existing design ran out of memory.
Sayantan Sur, Lei Chai, Hyun-Wook Jin, Dhabaleswar K. Panda 0001
IPDPS2
2006 RDMA read based rendezvous protocol for MPI over InfiniBand: design alternatives and benefits
abstract
Message Passing Interface (MPI) is a popular parallel programming model for scientific applications. Most high-performance MPI implementations use Rendezvous Protocol for efficient transfer of large messages. This protocol can be designed using either RDMA Write or RDMA Read. Usually, this protocol is implemented using RDMA Write. The RDMA Write based protocol requires a two-way handshake between the sending and receiving processes. On the other hand, to achieve low latency, MPI implementations often provide a polling based progress engine. The two-way handshake requires the polling progress engine to discover multiple control messages. This in turn places a restriction on MPI applications that they should call into the MPI library to make progress. For compute or I/O intensive applications, it is not possible to do so. Thus, most communication progress is made only after the computation or I/O is over. This hampers the computation to communication overlap severely, which can have a detrimental impact on the overall application performance. In this paper, we propose several mechanisms to exploit RDMA Read and selective interrupt based asynchronous progress to provide better computation/communication overlap on InfiniBand clusters. Our evaluations reveal that it is possible to achieve nearly complete computation/communication overlap using our RDMA Read with Interrupt based Protocol. Additionally, our schemes yield around 50% better communication progress rate when computation is overlapped with communication. Further, our application evaluation with Linpack (HPL) and NAS-SP (Class C) reveals that MPI_Wait time is reduced by around 30% and 28%, respectively, for a 32 node InfiniBand cluster. We observe that the gains obtained in the MPI_Wait time increase as the system size increases. This indicates that our designs have a strong positive impact on scalability of parallel applications.
Sayantan Sur, Hyun-Wook Jin, Lei Chai, Dhabaleswar K. Panda 0001
PPoPP3
2005 LiMIC: Support for High-Performance MPI Intra-node Communication on Linux Cluster
abstract
High performance intra-node communication support for MPI applications is critical for achieving best performance from clusters of SMP workstations. Present day MPI stacks cannot make use of operating system kernel support for intra-node communication. This is primarily due to the lack of an efficient, portable, stable and MPI friendly interface to access the kernel functions. In this paper we attempt to address design challenges for implementing such a high performance and portable kernel module interface. We implement a kernel module interface called LiMIC and integrate it with MVAPICH, an open source MPI over InfiniBand. Our performance evaluation reveals that the point-to-point latency can be reduced by 71% and the bandwidth improved by 405% for 64 KB message size. In addition, LiMIC can improve HPCC effective bandwidth and NAS IS class B benchmarks by 12% and 8%, respectively, on an 8-node dual SMP InfiniBand cluster.
Hyun-Wook Jin, Sayantan Sur, Lei Chai, Dhabaleswar K. Panda 0001
ICPP3