Rishan Chen

dblp:56/8165 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 2Databases, data management, data science and information retrieval · 2 · 1 first-authorArtificial intelligence and machine learning · 1Computer networks · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Hardware reliability and fault tolerance · 37% Parallel and multicore computing · 27% Cloud and datacenter computing · 18%
Databases, data mining, and information retrieval
2 papers
Information retrieval · 27% Data mining · 27% Web and social media mining · 27%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 77% Program analysis · 23%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware reliability and fault tolerance › software fault tolerance
hypervisor-based fault tolerance
0.212015
FTXen: Making hypervisor resilient to hardware faults on relaxed cores · HPCA 2015
Hardware reliability and fault tolerance › soft errors
soft error resilience
0.212015
FTXen: Making hypervisor resilient to hardware faults on relaxed cores · HPCA 2015
Cloud and datacenter computing
virtualization
0.212015
FTXen: Making hypervisor resilient to hardware faults on relaxed cores · HPCA 2015
Web and social media mining › event detection
burst detection
0.112012
EventSearch: a system for event discovery and retrieval on multi-type historical data · KDD 2012
Information retrieval › document retrieval › temporal information retrieval
event retrieval
0.112012
EventSearch: a system for event discovery and retrieval on multi-type historical data · KDD 2012
Data mining › text mining
temporal text mining
0.112012
EventSearch: a system for event discovery and retrieval on multi-type historical data · KDD 2012
Compilers and program optimization
compiler optimization
0.112012
Spotting Code Optimizations in Data-Parallel Pipelines through PeriSCOPE · OSDI 2012
Parallel and multicore computing
data-parallel programming
0.112012
Optimizing Data Shuffling in Data-Parallel Computation by Understanding User-Defined Functions · NSDI 2012
Distributed systems › distributed data processing
data shuffling
0.112012
Optimizing Data Shuffling in Data-Parallel Computation by Understanding User-Defined Functions · NSDI 2012
Parallel and multicore computing
parallel programming models
0.112012
Spotting Code Optimizations in Data-Parallel Pipelines through PeriSCOPE · OSDI 2012
Graph data management
graph processing
0.112010
Large graph processing in the cloud · SIGMOD Conference 2010
Program analysis
static analysis
0.012012
Spotting Code Optimizations in Data-Parallel Pipelines through PeriSCOPE · OSDI 2012
Parallel and multicore computing › data-parallel programming
mapreduce
0.012010
Large graph processing in the cloud · SIGMOD Conference 2010

Methods — techniques the papers use, named apart from their topics

static analysis · 0.3code optimization · 0.3fault injection · 0.2burst model · 0.1
YearPublicationVenuePosition
2015 FTXen: Making hypervisor resilient to hardware faults on relaxed cores
abstract
As CMOS technology scales, the Increasingly smaller transistor components are susceptible to a variety of in-field hardware errors. Traditional redundancy techniques to deal with the increasing error rates are expensive and energy inefficient. To address this emerging challenge, many researchers have recently proposed the idea of relaxed hardware design and exposing errors to software. For such relaxed hardware to become a reality, it is crucially important for system software, such as the virtual machine hypervisor, to be resilient to hardware faults. To address the above fundamental software challenge in enabling relaxed hardware design, we are making a major effort in restructuring an important part of system software, namely the virtual machine hypervisor, to be resilient to faulty cores. A fault in a relaxed core can only affect those virtual machines (and applications) running on that core, but the hypervisor and other virtual machines remain intact and continue providing services. We have redesigned every component of Xen, a large, popular virtual machine hypervisor, to achieve such error resiliency. This paper presents our design and implementation of the restructured Xen (we refer to it as FTXen). Our experimental evaluation on real systems shows that FTXen adds minimum application overhead, and scales well to different ratios of reliable and relaxed cores. Our results with random fault injection show that FTXen can successfully survive all injected hardware faults.
Xinxin Jin, Tianwei Sheng, Rishan Chen, Zhiyong Shan, Yuanyuan Zhou 0001
HPCA4
2012 Improving large graph processing on partitioned graphs in the cloud
abstract
As the study of large graphs over hundreds of gigabytes becomes increasingly popular for various data-intensive applications in cloud computing, developing large graph processing systems has become a hot and fruitful research area. Many of those existing systems support a vertex-oriented execution model and allow users to develop custom logics on vertices. However, the inherently random access pattern on the vertex-oriented computation generates a significant amount of network traffic. While graph partitioning is known to be effective to reduce network traffic in graph processing, there is little attention given to how graph partitioning can be effectively integrated into large graph processing in the cloud environment. In this paper, we develop a novel graph partitioning framework to improve the network performance of graph partitioning itself, partitioned graph storage and vertex-oriented graph processing. All optimizations are specifically designed for the cloud network environment. In experiments, we develop a system prototype following Pregel (the latest vertex-oriented graph engine by Google), and extend it with our graph partitioning framework. We conduct the experiments with a real-world social network and synthetic graphs over 100GB each in a local cluster and on Amazon EC2. Our experimental results demonstrate the efficiency of our graph partitioning framework, and the effectiveness of network performance aware optimizations on the large graph processing engine.
Rishan Chen, Mao Yang 0004, Xuetian Weng, Byron Choi, Bingsheng He, Xiaoming Li 0001
SoCC1
2012 EventSearch: a system for event discovery and retrieval on multi-type historical data
abstract
We present EventSearch, a system for event extraction and retrieval on four types of news-related historical data, i.e., Web news articles, newspapers, TV news program, and micro-blog short messages. The system incorporates over 11 million web pages extracted from "Web InfoMall", the Chinese Web Archive since 2001. The newspaper and TV news video clips also span from 2001 to 2011. The system, upon a user query, returns a list of event snippets from multiple data sources. A novel burst model is used to discover events from time-stamped texts. In addition to offline event extraction, our system also provides online event extraction to further meet the user needs. EventSearch provides meaningful analytics that synthesize an accurate description of events. Users interact with the system by ranking the identified events using different criteria (scale, recency and relevance) and submitting their own information needs in different input fields.
Dongdong Shan, Wayne Xin Zhao, Rishan Chen, Baihan Shu, Ziqi Wang 0002, Hongfei Yan, Xiaoming Li 0001
KDD3
2012 Optimizing Data Shuffling in Data-Parallel Computation by Understanding User-Defined Functions
Hucheng Zhou, Rishan Chen, Xuepeng Fan, Haoxiang Lin, Jack Li 0001, Wei Lin 0016, Jingren Zhou 0001, Lidong Zhou
NSDI3
2012 Spotting Code Optimizations in Data-Parallel Pipelines through PeriSCOPE
Xuepeng Fan, Rishan Chen, Hucheng Zhou, Sean McDirmid, Chang Liu 0021, Wei Lin 0016, Jingren Zhou 0001, Lidong Zhou
OSDI3
2010 Comet: batched stream processing for data intensive distributed computing
abstract
Batched stream processing is a new distributed data processing paradigm that models recurring batch computations on incrementally bulk-appended data streams. The model is inspired by our empirical study on a trace from a large-scale production data-processing cluster; it allows a set of effective query optimizations that are not possible in a traditional batch processing model.We have developed a query processing system called Comet that embraces batched stream processing and integrates with DryadLINQ. We used two complementary methods to evaluate the effectiveness of optimizations that Comet enables. First, a prototype system deployed on a 40-node cluster shows an I/O reduction of over 40% using our benchmark. Second, when applied to a real production trace covering over 19 million machine-hours, our simulator shows an estimated I/O saving of over 50%.
Bingsheng He, Mao Yang 0004, Rishan Chen, Wei Lin 0016, Lidong Zhou
SoCC4
2010 Large graph processing in the cloud
abstract
As the study of graphs, such as web and social graphs, becomes increasingly popular, the requirements of efficiency and programming flexibility of large graph processing tasks challenge existing tools. We propose to demonstrate Surfer, a large graph processing engine designed to execute in the cloud. Surfer provides two basic primitives for programmers - MapReduce and propagation. MapReduce, originally developed by Google, processes different key-value pairs in parallel, and propagation is an iterative computational pattern that transfers information along the edges from a vertex to its neighbors in the graph. These two primitives are complementary in graph processing. MapReduce is suitable for processing flat data structures, such as vertex-oriented tasks, and propagation is optimized for edge-oriented tasks on partitioned graphs.
Rishan Chen, Xuetian Weng, Bingsheng He, Mao Yang 0004
SIGMOD Conference1
2009 Wave Computing in the Cloud
Bingsheng He, Mao Yang 0004, Rishan Chen, Wei Lin 0016, Lidong Zhou
HotOS4