Xunchao Chen

dblp:168/0906 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
0since 2021 · last 2018
0000-0002-3272-1369ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Memory systems · 63% Storage systems · 16% Parallel and multicore computing · 11%

Topics — the 16 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
data locality
0.312018
Achieving Load Balance for Parallel Data Access on Distributed File Systems · IEEE Trans. Computers 2018
Storage systems › file systems
distributed file system
0.312018
Achieving Load Balance for Parallel Data Access on Distributed File Systems · IEEE Trans. Computers 2018
Parallel and multicore computing
load balancing
0.312018
Achieving Load Balance for Parallel Data Access on Distributed File Systems · IEEE Trans. Computers 2018
Memory systems
cache design
0.312017
Energy-Aware Adaptive Restore Schemes for MLC STT-RAM Cache · IEEE Trans. Computers 2017
Memory systems › cache › STT-RAM cache
MLC STT-RAM cache
0.312017
Energy-Aware Adaptive Restore Schemes for MLC STT-RAM Cache · IEEE Trans. Computers 2017
Memory systems
non-volatile memory
0.312017
Energy-Aware Adaptive Restore Schemes for MLC STT-RAM Cache · IEEE Trans. Computers 2017
Storage systems › flash and SSD › flash memory reliability
read disturbance mitigation
0.312017
Energy-Aware Adaptive Restore Schemes for MLC STT-RAM Cache · IEEE Trans. Computers 2017
Memory systems › non-volatile memory › magnetic random access memory
STT-MRAM
0.312017
Energy-Aware Adaptive Restore Schemes for MLC STT-RAM Cache · IEEE Trans. Computers 2017
Memory systems › non-volatile memory
write disturbance mitigation
0.312017
Energy-Aware Adaptive Restore Schemes for MLC STT-RAM Cache · IEEE Trans. Computers 2017
Memory systems
cache
0.212016
AOS: adaptive overwrite scheme for energy-efficient MLC STT-RAM cache · DAC 2016
Energy-efficient computing › power management
memory power management
0.212016
AOS: adaptive overwrite scheme for energy-efficient MLC STT-RAM cache · DAC 2016
Memory systems › non-volatile memory › magnetic random access memory › STT-MRAM
MLC STT-RAM
0.212016
AOS: adaptive overwrite scheme for energy-efficient MLC STT-RAM cache · DAC 2016
Memory systems › cache
STT-RAM cache
0.212016
AOS: adaptive overwrite scheme for energy-efficient MLC STT-RAM cache · DAC 2016
Parallel and multicore computing › parallel computing
parallel data access
0.112018
Achieving Load Balance for Parallel Data Access on Distributed File Systems · IEEE Trans. Computers 2018
Energy-efficient computing › power management › memory power management
cache energy reduction
0.112017
Energy-Aware Adaptive Restore Schemes for MLC STT-RAM Cache · IEEE Trans. Computers 2017
Energy-efficient computing
power management
0.112017
Energy-Aware Adaptive Restore Schemes for MLC STT-RAM Cache · IEEE Trans. Computers 2017

Methods — techniques the papers use, named apart from their topics

matching algorithm · 0.3heatmap monitoring · 0.3HM-LRU · 0.3forecasting · 0.3read reuse distance · 0.2adaptive overwrite scheme · 0.2
YearPublicationVenuePosition
2018 Achieving Load Balance for Parallel Data Access on Distributed File Systems
abstract
The distributed file system, HDFS, is widely deployed as the bedrock for many parallel big data analysis. However, when running multiple parallel applications over the shared file system, the data requests from different processes/executors will unfortunately be served in a surprisingly imbalanced fashion on the distributed storage servers. These imbalanced access patterns among storage nodes are caused because a). unlike conventional parallel file system using striping policies to evenly distribute data among storage nodes, data-intensive file system such as HDFS store each data unit, referred to as chunk file, with several copies based on a relative random policy, which can result in an uneven data distribution among storage nodes; b). based on the data retrieval policy in HDFS, the more data a storage node contains, the higher probability the storage node could be selected to serve the data. Therefore, on the nodes serving multiple chunk files, the data requests from different processes/executors will compete for shared resources such as hard disk head and networkbandwidth, resulting in a degraded I/O performance. In this paper, we first conduct a complete analysis on how remote and imbalanced read/write patterns occur and how they are affected by the size of the cluster. We then propose novel methods, referred to as Opass, to optimize parallel data reads, as well as to reduce the imbalance of parallel writes on distributed file systems. Our proposed methods can benefit parallel data-intensive analysis with various parallel data access strategies. Opass adopts new matching-based algorithms to match processes to data so as to compute the maximum degree of data locality and balanced data access. Furthermore, to reduce the imbalance of parallel writes, Opass employs a heatmap for monitoring the I/O statuses of storage nodes and performs HM-LRU policy to select a local optimal storage node for serving write requests. Experiments are conducted on PRObE's Marmot 128-node cluster testbed and the results from both benchmark and well-known parallel applications show the performance benefits and scalability of Opass.
Dan Huang 0001, Dezhi Han, Jun Wang 0001, Jiangling Yin, Xunchao Chen, Xuhong Zhang 0002, Jian Zhou 0004, Mao Ye 0008
IEEE Trans. Computers5
2017 DFS-container: achieving containerized block I/O for distributed file systems
abstract
Today BigData systems commonly use resource management systems such as TORQUE, Mesos, and Google Borg to share the physical resources among users or applications. Enabled by virtualization, users can run their applications on the same node with low mutual interference. Container-based virtualizations (e.g., Docker and Linux Containers) offer a lightweight virtualization layer, which promises a near-native performance and is adopted by some Big-Data resource sharing platforms such as Mesos. Nevertheless, using containers to consolidate the I/O resources of shared storage systems is still at an early stage, especially in a distributed file system (DFS) such as Hadoop File System (HDFS). To overcome this issue, we propose a distributed middleware system, DFS-Container, by further containerizing DFS. We also evaluate and analyze the unfairness of using containers to proportionally allocate the I/O resource of DFS. Based on these analyses and evaluations, we propose and implement a new mechanism, IOPS-Regulator, which improve the fairness of proportional allocation by 74.4% on average.
Dan Huang 0001, Jun Wang 0001, Qing Liu 0001, Xuhong Zhang 0002, Xunchao Chen, Jian Zhou 0004
SoCC5
2017 SideIO: A Side I/O system framework for hybrid scientific workflow
Jun Wang 0001, Dan Huang 0001, Huafeng Wu, Jiangling Yin, Xuhong Zhang 0002, Xunchao Chen
J. Parallel Distributed Comput.6
2017 Energy-Aware Adaptive Restore Schemes for MLC STT-RAM Cache
abstract
For the sake of higher cell density while achieving near-zero standby power, recent research progress in Magnetic Tunneling Junction (MTJ) devices has leveraged Multi-Level Cell (MLC) configurations of Spin-Transfer Torque Random Access Memory (STT-RAM). However, in orderto mitigate the write disturbance in an MLC strategy, data stored in the soft bit must be restored back immediately after the hard bit switching is completed. Furthermore, as the result of MTJ feature size scaling, the soft bit can be expected to become disturbed by the read sensing current, thus requiring an immediate restore operation to ensure the data reliability. In this paper, we design and analyze a novel Adaptive Restore Scheme for Write Disturbance (ARS-WD) and Read Disturbance (ARS-RD), respectively. ARS-WD alleviates restoration overhead by intentionally overwriting soft bit lines which are less likely to be read. ARS-RD, on the other hand, aggregates the potential writes and restore the soft bit line at the time of its eviction from higher level cache. Both of these two schemes are based on a lightweight forecasting approach for the future read behavior of the cache block. Our experimental results show substantial reduction in soft bit line restore operations, delivering 17.9 percent decrease in overall energy consumption and 9.4 percent increase in IPC, while incurring negligible capacity overhead. Moreover, ARS promotes advantages of MLC to provide a preferable L2 design alternative in terms of energy, area and latency product compared to SLC STT-RAM alternatives.
Xunchao Chen, Navid Khoshavi, Ronald F. DeMara, Jun Wang 0001, Dan Huang 0001, Wujie Wen, Yiran Chen 0001
IEEE Trans. Computers1
2016 AOS: adaptive overwrite scheme for energy-efficient MLC STT-RAM cache
abstract
Spin-Transfer Torque Random Access Memory (STT-RAM) has been identified as an advantageous candidate for on-chip memory technology due to its high density and ultra low leakage power. Recent research progress in Magnetic Tunneling Junction (MTJ) devices has developed Multi-Level Cell (MLC) STT-RAM to further enhance cell density. To avoid the write disturbance in MLC strategy, data stored in the soft bit must be restored back immediately after the hard bit switching is completed. However, frequent restores are not only unnecessary, but also introduce a significant energy consumption overhead. In this paper, we propose an Adaptive Overwrite Scheme (AOS) which alleviates restoration overhead by intentionally overwriting selected soft bits based on RRD (Read Reuse Distance). Our experimental results show 54.6% reduction in soft bit restoration, delivering 10.8% decrease in overall energy consumption. Moreover, AOS promotes MLC to be a preferable L2 design alternative in terms of energy, area and latency product.
Xunchao Chen, Navid Khoshavi, Jian Zhou 0004, Dan Huang 0001, Ronald F. DeMara, Jun Wang 0001, Wujie Wen, Yiran Chen 0001
DAC1
2015 Achieving up to zero communication delay in BSP-based graph processing via vertex categorization
abstract
The Bulk Synchronous Parallel (BSP) model, which divides a graphing algorithm into multiple supersteps, has become extremely popular in distributed graph processing systems. However, the high number of network messages exchanged in each superstep of the graph algorithm will create a long period of time. We refer to this as a communication delay. Furthermore, the BSP's global synchronization barrier does not allow computation in the next superstrep to be scheduled during this communication delay. This communication delay makes up a large percentage of the overall processing time of a superstep. While most recent research has focused on reducing number of network messages, but communication delay is still a deterministic factor for overall performance. In this paper, we add a runtime communication and computation scheduler into current graph BSP implementations. This scheduler will move some computation from the next superstep to the communication phase in the current superstep to mitigate the communication delay. Finally, we prototyped our system, Zebra, on Apache Hama, which is an open source clone of the classic Google Pregel. By running a set of graph algorithms on an in-house cluster, our evaluation shows that our system could completely eliminate the communication delay in the best case and can achieve average 2X speedup over Hama.
Xuhong Zhang 0002, Xunchao Chen, Jun Wang 0001, Tyler Lukasiewicz, Dezhi Han
NAS3