Zhou Lei 0001

dblp:90/3667-1 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
5since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 4 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3Software engineering, systems software and programming languages · 2Security and privacy · 1 · 1 first-author
YearPublicationVenuePosition
2024 A Comfortable and Robust DRL-based Car-following Policy Incorporating Lateral Information under Cut-in Scenarios
abstract
The cut-in behavior of adjacent vehicles presents a challenge for the Adaptive Cruise Control (ACC) system. Inability to proactively discern adjacent vehicles’ cut-in actions could impact driving safety. In addition, abrupt changes in ego vehicle’s following target might provoke excessive reactions, undermining passenger comfort. To address this challenge, this paper integrates trajectory prediction model into a deep reinforcement learning(DRL)-based car-following policy. Utilizing Finite State Machine(FSM), we proactively identify cut-in vehicles based on the predicted trajectories to enhance safety. In designing the DRL-based car-following policy, we propose a novel reward function by analyzing human driving data distribution and considering lateral information of cut-in vehicles. This method enhances driving comfort by significantly reducing abrupt maneuvers in both car-following and cut-in scenarios. Additionally, we investigate the impact of state observation configurations on the performance of the DRL policy. Our experimental findings reveal that incorporating the ego vehicle’s acceleration into the observation state contributes to optimizing comfort and enhancing robustness in scenarios where the observation of other vehicles’ motion state is not precise.
Zhifei Yang 0005, Weijia Lu, Wenfeng Shen, Zhou Lei 0001
IV5
2023 Optimizing Transformer Training Based on Computation and Accessing Memory Features with Deep Learning Processor
abstract
The Transformer model, which has significantly advanced natural language processing and computer vision, overcomes the limitations of recurrent neural networks and convolutional neural networks. However, it faces challenges with computational efficiency and memory management due to complex computations and variable-length inputs. Despite research efforts, these issues persist. This paper presents a novel optimization of the Transformer model, following an indepth analysis of its computational graph structure. Firstly, we utilize Deep Computing Unit (DCU) as our hardware platform. Secondly, we optimize element-wise and reduction operators through operator fusion and rewriting. Thirdly, we develop a fine-grained memory management algorithm using a greedy strategy. As a result, the training speed of the Transformer model increases by 1.2x - 1.4x without compromising accuracy.
Zhou Lei 0001, Qin Fang, Jingfeng Qian, Shengbo Chen, Qingguo Xu, Ninghua Yang
ICPADS1
2023 TGAS-ReID: Efficient architecture search for person re-identification via greedy decisions with topological order
Shengbo Chen, Xianrui Liu, Kangkang Yang, Zhou Lei 0001
Appl. Intell.5
2023 KubeGPU: efficient sharing and isolation mechanisms for GPU resource management in container cloud
Wenfeng Shen, Zhengsen Liu, Yunjie Tan, Zhaokai Luo, Zhou Lei 0001
J. Supercomput.5
2021 Load Balancing Optimization for Transformer in Distributed Environment
abstract
In recent years, the demand for artificial intelligence applications has increased dramatically. Complex models can promote machine learning to achieve excellent results, but computing efficiency has gradually reached a bottleneck. Therefore, more researchers are exploring the improvement of the efficiency of intelligent computing systems. Distributed machine learning can improve the efficiency of model training and inference, but problems such as communication delay and load imbalance between computing nodes still exist. In the multi-GPU distributed computing environment, this paper takes the vision field algorithm VIT (vision transformer) as the optimization object, which has the advantage of convenient parallel training, and proposes several related solutions. Firstly, the parameter server is used as the system logic architecture and in order to reduce the idleness of the computing devices during the training process, the device working status query mechanism is designed to realize load balancing. Secondly, combined with the pre-trained small VIT algorithm model, semi-asynchronous communication method is proposed to reduce the communication overhead of computing devices and accelerate global convergence. The results of this experiment carried out in the existing distributed environment has demonstrated that compared with the existing synchronization method, the computational efficiency has been improved well under the premise of slightly reducing the accuracy.
Delu Ma, Zhou Lei 0001, Shengbo Chen
ICPADS2
2016 An approach to improving the performance of CUDA in virtual environment
abstract
The GPU pass-through technology, under the virtual machine(VM), is always used in CUDA programming. The data will transfer between the VM main memories and GPU memory during the processing of CUDA programs. It is known that the transmission speed of the pinned memory is faster than the pageable when a large amount of data appears. However, the cost of pinned memory will occupy a lot of physical memory and lead to a lack of memory for the VM. So it can also affect the memory usage of other applications in the VM. In this paper, we propose a method to improve the performance of CUDA in virtual environment. In our experiment, the virtual machine manager is chosen as Kernel-based Virtual Machine(KVM), and the GPU pass-through is used for distributing the GPU for the VMs. We also defined a new module as CUDA Memory Management(CMM) which is added into the virtual environment. The result shows that the approach has high efficiency in CUDA program. And the efficiency of data transmission for using the CMM module is more than 16% than the pageable memory only using mode.
Shenquan Han, Zhou Lei 0001, Wenfeng Shen, Shengbo Chen, Huiran Zhang, Tao Zhang 0046, Baoyu Xu
SNPD2
2014 Distributed video transcoding based on MapReduce
abstract
Video transcoding is an important job in video processing and network service. With the improvement of devices and the Internet, the size of the video increases rapidly so that it takes a lot of resources to transcode. Low efficiency, high cost of upgrading hardware and low capacity of processing failure are problems of the traditional method of serial transcoding. Distributed transcoding can resolve these problems. To reduce the time of serial processing and be able to deal with the fault, this paper means to model a distributed video transcoding system which is based on MapReduce, an open source distribute computing model, and FFmpeg.
Chenwei Song, Wenfeng Shen, Lianqiang Sun, Zhou Lei 0001, Weimin Xu
ICIS4
2014 An Improved Image File Storage Method Using Data Deduplication
abstract
Recent years have seen a rapid growth in the number of virtual machines and virtual machine images that are managed to support infrastructure as a service (IaaS). For example, Amazon Elastic Compute Cloud (EC2) has 6,521 public virtual machine images. This creates several challenges in management of image files in a cloud computing environment. In particular, a large amount of duplicate data that exists in image files consumes significant storage space. To address this problem, we propose an effective image file storage technique using data deduplication with a modified fixed-size block scheme. When a user requests to store an image file, this technique first calculates the fingerprint for the image file, and then compares the fingerprint with the fingerprints in a fingerprint library. If the fingerprint of the image is already in the library, a pointer to the existing fingerprint is used to store this image. Otherwise this image will be processed using the fixed-size block image segmentation method. We design a metadata format for image files to organize image file blocks and a new MD5 index table of image files to reduce their retrieval time. The experiments show that our technique can significantly reduce the transmission time of image files that have already existed in storage. Also the deletion rate for image groups which have the same version of operating systems but different versions of software applications is up about 58%.
Zhou Lei 0001, Zhaoxin Li, Yu Lei 0001, Yanling Bi, Luokai Hu, Wenfeng Shen
TrustCom1
2013 A probabilistic integrity checking approach for dynamic data in untrusted cloud storage
abstract
This paper proposes a simple approach for client to verify whether his/her data have been modified on the cloud storage server without downloading the data. In our approach, a number of bytes of each data block is collected to compose its metadata and stored in the cloud server along with small extra data. The approach effectively supports dynamic data operations with light overheads, in case of both computation and bandwidth. Our experiment shows that the computation time consumed is much less than the traditional MD5 checksum depending on different variable.
Thanh Cuong Nguyen, Wenfeng Shen, Zhou Lei 0001, Weimin Xu, Wencong Yuan, Chenwei Song
ICIS3
2009 An innovative application execution toolkit for multicluster grids
abstract
Multicluster grids provide one promising solution to satisfying growing computation demands of compute-intensive applications by collaborating various networked clusters. However, it is challenging to seamlessly integrate all participating clusters in different domains into a virtual computation platform. In order to take full advantages of multicluster grids capability, computer scientists need to deal with how to collaborate practically and efficiently participating autonomic systems to execute Grid-enabled applications. We make efforts on grid resource management and implement a toolkit called Pelecanus to improve the overall performance of application execution in multicluster grids environment. The Pelecanus takes advantages of the DA-TC (Dynamic Assignment with Task Containers) execution model to improve resource interoperability and enhance application execution and monitoring. Experiments show that it can significantly reduce turnaround time and increase resource utilization for certain applications with large number of sequential jobs.
Zhifeng Yun, Zhou Lei 0001, Gabrielle Allen, Daniel S. Katz, Tevfik Kosar, Shantenu Jha, J. Ramanujam
CLUSTER2
2008 A Grid-enabled problem-solving environment for advanced reservoir uncertainty analysis
abstract
Abstract Uncertainty analysis is critical for conducting reservoir performance prediction. However, it is challenging because it relies on (1) massive modeling‐related, geographically distributed, terabyte, or even petabyte scale data sets (geoscience and engineering data), (2) needs to rapidly perform hundreds or thousands of flow simulations, being identical runs with different models calculating the impacts of various uncertainty factors, (3) an integrated, secure, and easy‐to‐use problem‐solving toolkit to assist uncertainty analysis. We leverage Grid computing technologies to address these challenges. We design and implement an integrated problem‐solving environmentResGridto effectively improve reservoir uncertainty analysis. The ResGrid consists of data management, execution management, and a Grid portal. Data Grid tools, such as metadata, replica, and transfer services, are used to meet massive size and geographically distributed characteristics of data sets. Workflow, task farming, and resource allocation are used to support large‐scale computation. A Grid portal integrates the data management and the computation solution into a unified easy‐to‐use interface, enabling reservoir engineers to specify uncertainty factors of interest and perform large‐scale reservoir studies through a web browser. The ResGrid has been used in petroleum engineering. Copyright © 2008 John Wiley & Sons, Ltd.
Zhou Lei 0001, Gabrielle Allen, Promita Chakraborty, Dayong Huang, Christopher D. White
Concurr. Comput. Pract. Exp.1
2007 An Integrated Grid Portal for Managing Energy Resources
abstract
The discovery and management of energy resources, especially at locations in the Gulf of Mexico, requires an economic but technically enhanced infrastructure. Research teams from Louisiana State University, University of Louisiana at Lafayette, and Southern University Baton Rouge are engaged in a collaborative effort to create a ubiquitous computing and monitoring system (UCoMS) for the discovery and management of energy resources. The UCoMS team has sucessfully addressed two difficult issues in this research: (1) the computational challenges faced by compute-intensive simulations for reservoir uncertainty analysis that requires thousands of simulations and deals with terabytes, and even petabytes, of data, (2) the development of a prototype wireless sensor network (WSN) infrastructure to collect and process realtime data from production locations. While the former requires the intensive computational power of the UCoMS grid resources, the latter requires efficient interfacing between WSN & grid. A unified workflow analysis has been performed to ensure smooth operation of both efforts and a unified portal has been created. This paper integrates the above two workflows and portals into a single platform. It illustrates the need for such integration for users with similar (but not same) goals and describes how to partition users among different groups with different access rights to ensure security within subgroups. Such a system can easily integrate future UCoMS sub-projects into a unified whole. Hence, our portal prototype serves as a good example of the benefit that may accrue from integrated workflows.
Promita Chakraborty, Gabrielle Allen, Zhou Lei 0001, Adam Wade Lewis, Ian Chang-Yen, Itthichok Jangjaimon, Nian-Feng Tzeng
eScience3
2006 ResGrid: A Grid-aware Toolkit for Reservoir Uncertainty Analysis
abstract
Many efforts in Grid communities have focused on middleware research and development. However, Grid application-level tools are needed which can build higherlevel functionality on top of core middleware services. We work with specific classes of scientific applications and present a Grid-aware toolkit ResGrid for reservoir uncertainty analysis. With the help of the ResGrid, a reservoir engineer can transparently take advantage of Grid resources and services for compute-intensive and dataintensive uncertainty analysis as well as enforce the understanding of reservoir modeling. In this paper, the ResGrid is introduced in terms of overview, architecture, and implementation status.
Zhou Lei 0001, Dayong Huang, Archit Kulshrestha, Santiago Peña, Gabrielle Allen, Christopher D. White, Richard Duff, John R. Smith, Subhash Kalla
CCGRID1
2006 Poster reception - Utilizing grid computing technologies for advanced reservoir studies
abstract
Reservoir studies are crucial to obtain accurate assessments and predictions of reservoir performance. However, this is a challenging issue because 1) it relies on massive modeling-related, geographically distributed, terabyte or even petabyte sized datasets (seismic and well-logging data), 2) needs to rapidly perform hundreds or thousands of simulations, being identical runs with different reservoir models circulating the impacts of various uncertainty factors, 3) the lack of easy-to-use problem solving toolkits to assist the uncertainty analysis.The poster focuses on leveraging Grid computing technologies to address the challenging issue mentioned above. It describes a newly developed data archive tool based on metadata and replica services and high performance file transfer. Our task farming framework enables a large amount of parallel job runs across a Grid, a related Grid portal eases the management of advanced reservoir studies. Our solutions are being employed by other Grid applications.
Zhou Lei 0001, Gabrielle Allen, Dayong Huang, Hartmut Kaiser, Christopher D. White
SC1