Mingfa Zhu

dblp:44/3722 · DBLP profile ↗
← Back
18ranked-venue papers
1as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7Applied, interdisciplinary, general and emerging computing · 5 · 1 first-authorSecurity and privacy · 2Computer networks · 1Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 70% Storage systems · 23% Processor architecture and microarchitecture · 7%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › cache management
cache partitioning
0.212013
LvtPPP: Live-Time Protected Pseudopartitioning of Multicore Shared Caches · IEEE Trans. Parallel Distributed Syst. 2013
Memory systems › cache management
cache replacement
0.212013
LvtPPP: Live-Time Protected Pseudopartitioning of Multicore Shared Caches · IEEE Trans. Parallel Distributed Syst. 2013
Memory systems › cache management
dead block prediction
0.212013
LvtPPP: Live-Time Protected Pseudopartitioning of Multicore Shared Caches · IEEE Trans. Parallel Distributed Syst. 2013
Storage systems › flash and SSD
victim block selection
0.212013
LvtPPP: Live-Time Protected Pseudopartitioning of Multicore Shared Caches · IEEE Trans. Parallel Distributed Syst. 2013
Processor architecture and microarchitecture
multicore design
0.012013
LvtPPP: Live-Time Protected Pseudopartitioning of Multicore Shared Caches · IEEE Trans. Parallel Distributed Syst. 2013

Methods — techniques the papers use, named apart from their topics

two-cascade victim selection · 0.2live-time protected counter · 0.2LRU replacement · 0.2
YearPublicationVenuePosition
2019 EDOA: an efficient delay optimization approach for mixed-polarity Reed-Muller logic circuits under the unit delay model
Zhenxue He, Limin Xiao 0002, Fei Gu 0001, Zhisheng Huo, Mingfa Zhu, Longbing Zhang, Rui Liu 0007, Xiang Wang 0006
Frontiers Comput. Sci.7
2017 A Power and Area Optimization Approach of Mixed Polarity Reed-Muller Expression for Incompletely Specified Boolean Functions
Zhenxue He, Limin Xiao 0002, Fei Gu 0001, Zhisheng Huo, Guangjun Qin, Mingfa Zhu, Longbing Zhang, Rui Liu 0007, Xiang Wang 0006
J. Comput. Sci. Technol.7
2016 EMA-FPRMs: An efficient minimization algorithm for fixed polarity Reed-Muller expressions
abstract
Fixed polarity Reed-Muller expressions (FPRMs) are well-suited for many practical applications due to they have many excellent properties. In order to obtain an optimal FPRM with fewest product terms, we propose an efficient minimization algorithm (EMA-FPRMs) for FPRMs. The main idea behind the EMA-FPRMs is that, firstly, the incompletely specified Boolean function is transformed into the zero polarity incompletely specified fixed polarity RM expression (ISFPRM) by using the proposed ISFPRM acquisition algorithm; secondly, the polarity and allocation of don't care terms of ISFPRM is encoded as chromosome; lastly, the optimal FPRM with fewest product terms is obtained by using genetic algorithm (GA), in which the FPRM that corresponds to the given chromosome is obtained by using the proposed chromosome conversion algorithm. The experimental results on MCNC benchmark circuits show that compared with the traditional polarity optimization approach which neglects the don't care terms, the EMA-FPRMs is highly effective in minimizing the number of product terms of FPRMs. Moreover, the EMA-FPRMs is faster than the GA based minimization algorithm which also considers the don't care terms.
Zhenxue He, Limin Xiao 0002, Longbing Zhang, Fei Gu 0001, Zhisheng Huo, Mingfa Zhu, Rui Liu 0007, Xiang Wang 0006
FPT6
2015 Lessen Interflow Interference Using Virtual Channels Partitioning
abstract
Interconnection networks are a significant consideration for high-performance computing and the datacenter. However, interflow interference seriously impacts the communication performance and even causes disastrous congestion. The paper reports a virtual channel (VC)-sharing scheme that is aimed to separate VCs into many groups, and assign them to data flows based on the destination address. The technique can effectively isolate various traffics into separate VCs groups such that heavy loads have a lesser influence on the other normal traffics. As a consequence, we achieve a slimming congestion tree. In the proposal, the routing algorithm is a two-stage selection that includes the port selection and the VC group selection, respectively. Each of them has an independently selecting algorithm so that routing algorithm is a combined tactic by Cartesian product. The experiment represents that our scheme has excellent performance on adversarial traffics. Using our scheme, the growth curve is linear and slow after crossing the saturation point. For benign traffic patterns, our scheme does not effect any oblivious change on the communication performance when the system receives a lower injection rate.
Guangjun Qin, Mingfa Zhu
Comput. J.2
2015 GPU accelerated sparse matrix-vector multiplication and sparse matrix-transpose vector multiplication
abstract
Summary Many high performance computing applications require computing both sparse matrix‐vector product (SMVP) and sparse matrix‐transpose vector product (SMTVP) for better overall performance. Under such a circumstance, it is critical to maintain a similarly high throughput for these two computing patterns with the underlying sparse matrix encoded in a single storage format. The compressed sparse block (CSB) format proposed by Buluç et al. allows computing both problems on multi‐core CPUs with nearly identical throughputs. On the other hand, a direct porting of CSB to graphics processing units (GPUs), which have been recently recognized as a powerful general purpose computing platform, turns out to be inefficient. In this work, we propose a new data structure, designated as expanded CSB (eCSB), to minimize the throughput gap between SMVP and SMTVP computations on GPUs, while at the same time enable a high computing throughput. We also use a hybrid storage format to store elements in each block, which can be selected dynamically at runtime. Experimental results show that the proposed techniques implemented on a Kepler GPU delivers similar throughput on both SMVP and SMTVP and the throughput is up to 13 times faster than that of the CPU‐based CSB implementation. In addition, our eCSB procedure outperforms the previous GPU results by up to 188% and 914% in computing SMVP and SMTVP, and we validate the effectiveness of eCSB by means of wall‐clock time of bi‐conjugate gradient algorithm; our eCSB is 25% faster than Compressed Sparse Rows (CSR) and 6% faster than HYB, respectively. Copyright © 2014 John Wiley & Sons, Ltd.
Yangdong Deng, Shuai Mu 0002, Zhenzhong Zhang, Mingfa Zhu
Concurr. Comput. Pract. Exp.5
2014 MIMP: Deadline and Interference Aware Scheduling of Hadoop Virtual Machines
abstract
Virtualization promised to dramatically increase server utilization levels, yet many data centers are still only lightly loaded. In some ways, big data applications are an ideal fit for using this residual capacity to perform meaningful work, but the high level of interference between interactive and batch processing workloads currently prevents this from being a practical solution in virtualized environments. Further, the variable nature of spare capacity may make it difficult to meet big data application deadlines. In this work we propose two schedulers: one in the virtualization layer designed to minimize interference on high priority interactive services, and one in the Hadoop framework that helps batch processing jobs meet their own performance deadlines. Our approach uses performance models to match Hadoop tasks to the servers that will benefit them the most, and deadline-aware scheduling to effectively order incoming jobs. The combination of these schedulers allows data center administrators to safely mix resource intensive Hadoop jobs with latency sensitive web applications, and still achieve predictable performance for both. We have implemented our system using Xen and Hadoop, and our evaluation shows that our schedulers allow a mixed cluster to reduce web response times by more than ten fold, while meeting more Hadoop deadlines and lowering total task execution times by 6.5%.
Wei Zhang 0052, Sundaresan Rajasekaran, Timothy Wood 0001, Mingfa Zhu
CCGRID4
2014 Atomic reduction based sparse matrix-transpose vector multiplication on GPUs
abstract
Sparse Matrix-Transpose Vector Product (SMTVP) is a frequently used computation pattern in High Performance Computing applications. It is typically solved by transposition followed by a Sparse Matrix-Vector Product (SMVP) in current linear algebra packages. However, the transposition process can be a serious bottleneck on modern parallel computing platforms. A previous work proposed a relatively complex data structure for efficiently computing SMTVP with multi-core CPUs, but it proved to be inefficient on GPUs. In this work, we show that the Compressed Sparse Row (CSR) based SMVP algorithm can also be efficient for SMTVP computation on modern GPUs. The proposed method exploits atomic operations to perform the reduce operation in the computation of each inner product of a row in the transposed matrix and the vector. Experimental results show that the simple technique can outperform the SMTVP flow of transposition plus SMVP released in the CUSPARSE package by up to 405-fold.
Yangdong Deng, Shuai Mu 0002, Mingfa Zhu, Zhibin Huang
ICPADS4
2014 A Load Balancing Strategy of SDN Controller Based on Distributed Decision
abstract
Software-Defined Networking (SDN), enabled by Open Flow, represents a paradigm shift from traditional network to the future Internet. Replicate or distributed controllers have been proposed to address the issues of availability and scalability that a centralized controller suffers from. However, it lacks a flexible mechanism to balance load among distributed controllers. To address this problem, this paper presents DALB, a dynamic and adaptive algorithm for controller load balancing totally based on distributed architecture, without any centralized component. This algorithm is running as a module of SDN controller. On one hand, it adopts an adjustable load collection threshold so as to reduce the overhead of exchanging messages for load collection, and on the other hand it can make policy and election locality in order to reduce the decision delay caused by network transmission. In this paper, we build the prototype system on floodlight to demonstrate our design and test the performance of our algorithm.
Yuanhao Zhou, Mingfa Zhu, Wenbo Duan, Deguo Li, Mingming Zhu
TrustCom2
2014 High availability for Non-stop network controller
abstract
Network controller is the core of the OpenFlow-Based Software-Defined Networks (SDN). High availability of network controller is an urgent need objectively. The appearance of controller instance fault should not be perceived by data plane. During the fault recovery, OpenFlow messages, especially asynchronous message from data plane, should not be discarded. Otherwise the control platform would hold the outdated network status. This is an interesting problem which we call Asynchronous Message Discard (ASMD) trouble. This paper proposes a Non-stop network controller (NSNC). It can troubleshoot the ASMD problem based on highly available TCP (HA-TCP) technology. During the fault recovery on distributed control platform, we want to elect the optimal controller to take over the network devices managed by the failed controller instance. We analyses the minimum metrics set for controller election. The view of this paper is verified on Floodlight controller. Experiments show that when the controller instance fails, asynchronous messages from switch will not be discarded, and using HA-TCP only adds negligible overhead to throughput and latency. We test the performance of election algorithm based on Mininet. Currently, there is only a theoretical analysis of metrics related to the election. We will test the validity of the election strategy in a real scenario as part of our future work.
Deguo Li, Mingfa Zhu, Wenbo Duan, Yuanhao Zhou, Mengxi Chen, Yinben Xia, Mingming Zhu
WoWMoM4
2013 Secure Distribution of Big Data Based on BitTorrent
abstract
Recently, big data becomes more and more widespread on the Internet, as various P2P protocols, especially BitTorrent, make great contribution. Accompanied with BitTorrent spreading, however, malicious activities, divulging sensitive data and other security problems arise. Relative researches and analyses indicate that existing means of protecting P2P network are sophisticated but intricate when distributing big data, leading inefficiency of implementation. In this paper, a scheme to distribute big data securely and efficiently on BitTorrent network is proposed, which can be implemented in server, authorizing peers' admittances and actions, hence protecting sensitive data in the network. To achieve this goal, identity verification and cipher system are embedded into BitTorrent protocol, enabling the server to regulate and keep trace of peers' behaviors and sensitive data. The experimental results show the functional effectiveness of this scheme, as well as acceptable overhead on server.
Chunjie Xu, Jingchao Qin, Guangjun Qin, Mingfa Zhu, Zhiyao Wang, Mingquan Li, Dongyu Tan
DASC5
2013 LvtPPP: Live-Time Protected Pseudopartitioning of Multicore Shared Caches
abstract
Partition enforcement policy is essential in the cache partition, and its main function is to protect the lines and retain the cache quota of each core. This paper focuses online protection based on its generation time rather than the CPU core ID that it belongs to or the position of the replacement stack, where it is located. The basic idea is that when a line is live, it must be protected and retained in the cache; when the line is “dead,” it needs to be evicted as early as possible. Therefore, the live-time protected counter (LvtP, four bits) is augmented to trace the lines' live time. Moreover, dead blocks are predicted according to the access event sequence. This paper presents a pseudopartition approach-LvtPPP and proposes a two-cascade victim selection mechanism to alleviate dead blocks based on the LRU replacement policy and the LvtP counter. LvtPPP also supports flexible handling of allocation deviation by introducing a parameter λ to adjust the generation time of the line. There is significant improvement of the performance and fairness in LvtPPP over PIPP and UCP according to the evaluation results based on Simics.
Zhibin Huang, Mingfa Zhu
IEEE Trans. Parallel Distributed Syst.2
2012 Memory Virtualization for MIPS Processor Based Cloud Server
Huixiang Wang, Mingfa Zhu, Feibo Li
GPC4
2012 LVMCI: Efficient and Effective VM Live Migration Selection Scheme in Virtualized Data Centers
abstract
Virtualization can provide significant benefits in virtualized data centers by enabling efficient and effective live migration to ensure service level agreement(SLA). Most of existing studies make decision on which bad virtual machines (VMs) should be migrated to which appropriate physical machines (PMs) in terms of resource utilizations. However, migration actions may degrade migrated application performance due to extra CPU and bandwidth consumptions. Furthermore, negative performance interferences amongst applications scheduled to the same PM may arise given the poor performance isolations of VMs on a PM. We design and implement a VM migration selection system with less migration costs and application performance interferences, called LVMCI (Live Virtual machine Migration with less Costs and application Interference). We propose a migration cost evaluation model to analyze quantitatively the aspects (i.e. throughput and response latency) of application performance degradation. Dirty rate and frequent dirty rate are two key factors that affect iteration time and downtime. We implement a tool that measures these parameters before VMs are migrated. We distinguish the performance degradation of migrated applications caused by memory iteration phase and stop-and-copy phase, which helps to select VM migrated. Besides that, we propose a performance interference model which helps to select the destination PM. The experimental results show that our system can estimate memory iteration time and downtime with high accuracy, and ensures a high level of SLAs by minimizing performance degradation during migration process and performance interference among co-located VMs at the destination PM.
Wei Zhang 0052, Mingfa Zhu, Yiduo Mei, Yunwei Gao, Yuzhong Sun
ICPADS2
2012 An Efficient Data Dissemination Approach for Cloud Monitoring
Xingjian Lu, Jianwei Yin, Ying Li 0001, Shuiguang Deng, Mingfa Zhu
ICSOC5
2012 Autonomic Resource Allocation in Virtualized Data Centers
abstract
Virtualization has been widely adopted in data centers for improving efficiency and flexibility. Multiple applications are co-hosted in virtualized data centers. In order to meet the Service Level Agreements (SLA), how to allocate resources for multiple applications is an important and challenging task, especially when dealing with fluctuating workloads and complex server applications. Virtual Machine Monitor provides fine-grained resource allocation and live migration. In this paper, we develop RTCOIN-Qclouds, a response time-aware, cost-aware and interference-aware control framework that tunes resource allocation, which ensures a high level of meeting the SLAs. Every physical machine's resources are assigned to multiple virtual machines which run on it based on application's response time, rather than traditional methods based on resource utilization. Virtual machine migration allows data centers to rebalance workloads across physical machines. However, migration actions may lead to performance impact during the migration process. Current virtualization techniques do not provide effective performance isolation between virtual machines (VMs). Specially, hidden contention for physical resources impacts performance differently in different virtual machines. As to the problem of selecting which virtual machines to be migrated, we consider migration cost. As to the problem of selecting which physical machine to be placed, we consider performance interference. Furthermore, we experimentally validate the effectiveness of response time-aware resource allocation in our framework using microbenchmarks.
Wei Zhang 0052, Mingfa Zhu, Qimeng Wu, Yuzhong Sun
ISPA2
2012 Performance Degradation-Aware Virtual Machine Live Migration in Virtualized Servers
abstract
Live migration of virtual machines(VMs) is widely used for system management in virtualized servers. When the loads increase and SLAs of some applications are violated, dynamic migration of virtual machines across physical machines (PMs) has the potential to ensure a high level of meeting the SLAs. Because of consuming extra CPU and bandwidth, application performance may be degraded during the migration process. However, different applications have different performance degradation. We design and implement a VM migration selection method that decides which VMs should be migrated. It can not only eliminate resouce competition on the PM, but also have less performance degradation during the migration process. We propose a performance degration-aware model to analyze applications' performance degradation which is directly sensitive to users. We analyze migration source code and find that memory size, dirty rate and frequent dirty rate are key factors that affect iteration time and downtime. We implement a tool that measures dirty rate and frequent dirty rate before VMs are migrated. we make a distinction between memory iteration phrase and stop-and-copy phrase owing to different performance degradation. The experimental results show that our method is effective.
Wei Zhang 0052, Mingfa Zhu, Yiduo Mei, Yuzhong Sun
PDCAT2
2010 DeepComp: towards a balanced system design for high performance computer systems
Mingfa Zhu, Qinfen Hao
Frontiers Comput. Sci. China1
1999 Exploiting the capabilities of the interconnection network on Dawning-1000
Mingfa Zhu
J. Comput. Sci. Technol.2