Yunhong Gu

dblp:83/6524 · DBLP profile ↗
← Back
18ranked-venue papers
7as first author
0since 2021 · last 2011
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 5 first-authorComputer networks · 2 · 1 first-authorArtificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
High-performance computing · 31% Storage systems · 21% Cloud and datacenter computing · 17%
Computer networks
3 papers
Transport protocols and congestion control · 100%

Topics — the 15 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
data locality
0.112011
Toward Efficient and Simplified Distributed Data Intensive Computing · IEEE Trans. Parallel Distributed Syst. 2011
Parallel and multicore computing › data-parallel programming
data-parallel frameworks
0.112011
Toward Efficient and Simplified Distributed Data Intensive Computing · IEEE Trans. Parallel Distributed Syst. 2011
High-performance computing › data-intensive computing
distributed data-intensive computing
0.112011
Toward Efficient and Simplified Distributed Data Intensive Computing · IEEE Trans. Parallel Distributed Syst. 2011
Storage systems › file systems
distributed file system
0.112011
Toward Efficient and Simplified Distributed Data Intensive Computing · IEEE Trans. Parallel Distributed Syst. 2011
High-performance computing
data transfer
0.122006
Bandwidth challenge - Transporting sloan digital sky survey data using SECTOR · SC 2006
Experiences in Design and Implementation of a High Performance Transport Protocol · SC 2004
High-performance computing › data transfer
high-speed data transfer
0.122006
Bandwidth challenge - Transporting sloan digital sky survey data using SECTOR · SC 2006
Experiences in Design and Implementation of a High Performance Transport Protocol · SC 2004
Storage systems
data management
0.112010
An overview of the Open Science Data Cloud · HPDC 2010
Cloud and datacenter computing
cloud storage
0.112008
Data mining using high performance data clouds: experimental studies using sector and sphere · KDD 2008
Distributed systems › distributed data processing
distributed data mining
0.112008
Data mining using high performance data clouds: experimental studies using sector and sphere · KDD 2008
Distributed systems
distributed data processing
0.112006
Bandwidth challenge - Transporting sloan digital sky survey data using SECTOR · SC 2006
Transport protocols and congestion control › transport protocols
high-speed transport protocol
0.012004
Experiences in Design and Implementation of a High Performance Transport Protocol · SC 2004
Transport protocols and congestion control › transport protocols
UDP-based transport
0.012004
Experiences in Design and Implementation of a High Performance Transport Protocol · SC 2004
Storage systems
distributed storage
0.012011
Toward Efficient and Simplified Distributed Data Intensive Computing · IEEE Trans. Parallel Distributed Syst. 2011
Transport protocols and congestion control
transport protocols
0.022006
Bandwidth challenge - Transporting sloan digital sky survey data using SECTOR · SC 2006
Supporting Configurable Congestion Control in Data Transport Services · SC 2005
Cloud and datacenter computing › cloud deployment model
distributed cloud
0.012010
An overview of the Open Science Data Cloud · HPDC 2010

Methods — techniques the papers use, named apart from their topics

data locality optimization · 0.1parallel data transfer · 0.1object-oriented design · 0.1bandwidth estimation · 0.1distributed data mining · 0.1
YearPublicationVenuePosition
2011 Toward Efficient and Simplified Distributed Data Intensive Computing
abstract
While the capability of computing systems has been increasing at Moore's Law, the amount of digital data has been increasing even faster. There is a growing need for systems that can manage and analyze very large data sets, preferably on shared-nothing commodity systems due to their low expense. In this paper, we describe the design and implementation of a distributed file system called Sector and an associated programming framework called Sphere that processes the data managed by Sector in parallel. Sphere is designed so that the processing of data can be done in place over the data whenever possible. Sometimes, this is called data locality. We describe the directives Sphere supports to improve data locality. In our experimental studies, the Sector/Sphere system has consistently performed about 2-4 times faster than Hadoop, the most popular system for processing very large data sets.
Yunhong Gu, Robert L. Grossman
IEEE Trans. Parallel Distributed Syst.1
2010 An overview of the Open Science Data Cloud
abstract
The Open Science Data Cloud is a distributed cloud based infrastructure for managing, analyzing, archiving and sharing scientific datasets. We introduce the Open Science Data Cloud, give an overview of its architecture, provide an update on its current status, and briefly describe some research areas of relevance.
Robert L. Grossman, Yunhong Gu, Joe Mambretti, Michal Sabala, Alex Szalay, Kevin P. White
HPDC2
2010 Sector: A high performance wide area community data storage and sharing system
Yunhong Gu, Robert L. Grossman
Future Gener. Comput. Syst.1
2009 Compute and storage clouds using wide area high performance networks
Robert L. Grossman, Yunhong Gu, Michal Sabala, Wanzhi Zhang
Future Gener. Comput. Syst.2
2008 Data mining using high performance data clouds: experimental studies using sector and sphere
abstract
We describe the design and implementation of a high performance cloud that we have used to archive, analyze and mine large distributed data sets. By a cloud, we mean an infrastructure that provides resources and/or services over the Internet. A storage cloud provides storage services, while a compute cloud provides compute services. We describe the design of the Sector storage cloud and how it provides the storage services required by the Sphere compute cloud. We also describe the programming paradigm supported by the Sphere compute cloud. Sector and Sphere are designed for analyzing large data sets using computer clusters connected with wide area high performance networks (for example, 10+ Gb/s). We describe a distributed data mining application that we have developed using Sector and Sphere. Finally, we describe some experimental studies comparing Sector/Sphere to Hadoop.
Robert L. Grossman, Yunhong Gu
KDD2
2007 UDT: UDP-based data transfer for high-speed wide area networks
Yunhong Gu, Robert L. Grossman
Comput. Networks1
2006 SDCS: Simplified Data Communications in Parallel/Distributed Applications
abstract
This paper presents SDCS (Simple Data Communication and Sharing), a programming model for data communications in parallel/distributed applications. With SDCS, developers can define data communications in shared memory style and have the model translate the declarations into corresponding message passing code. The translation from data sharing declarations to message passing code is based on simple mapping rules to lower runtime overhead and increase understandability of the model. Some frequently seen data communication modes are well supported to enhance its usability. SDCS can effectively reduce the difficulty in programming process communications.
Yong Mao, Yunhong Gu, Robert L. Grossman
CCGRID2
2006 Distributing the Sloan Digital Sky Survey Using UDT and Sector
abstract
In this paper, we describe a peer-to-peer storage system called Sector that is designed to access and transport large data sets over wide area high performance networks. We also describe our recent experience using Sector to distribute the Sloan Digital Sky Survey BESTDR4 catalog data.
Yunhong Gu, Robert L. Grossman, Alex Szalay, Ani Thakar
e-Science1
2006 Bandwidth challenge - Transporting sloan digital sky survey data using SECTOR
abstract
National Center for Data Mining at UICIn our SC06 BWC entry, we will transfer SDSS (Sloan Digital Sky Survey) Data Release 5 (DR5) between the SC06 show floor in Tampa and one of the NCDM labs on the UIC campus. We will use SECTOR, our newly developed distributed data space management system, to transfer DR5 in parallel between two Linux clusters in Tampa and Chicago, respectively. SECTOR transparently manages the file locating and data moving, while it employs UDT for actual data transfer. The data transfer will be from disk to disk over a 10Gb/s shared, router link between SC06 and UIC, via StarLight. We expect to reach 5Gb/s disk-to-disk data transfer rate between the two sites.
Robert L. Grossman, Yunhong Gu, Michal Sabala, Shirley Connelly, David Hanley, Joe Mambretti, Alex Szalay, Ani Thakar, Jan vandenBerg, Alainna Wonders
SC2
2006 Data mining middleware for wide-area high-performance networks
Robert L. Grossman, Yunhong Gu, David Hanley, Michal Sabala, Joe Mambretti, Alex Szalay, Ani Thakar, Kazumi Kumazoe, Yuji Oie, Yoonjoo Kwon, Woojin Seok
Future Gener. Comput. Syst.2
2005 Supporting Configurable Congestion Control in Data Transport Services
abstract
As wide area high-speed networks rapidly increase, new applications emerge and require new control mechanisms in data transport services to support them. In this paper, we present UDT/CCC, a data transport library that allows users to make use of a new control algorithm through simple configurations. We aim to provide a tool for fast implementation and deployment, as well as easy evaluation, of new congestion control algorithms. UDT/CCC uses an objected-oriented design. We show that our UDT/CCC library can be used to easily implement a large variety of control algorithms and can simulate the behavior of their native implementations as well. The UDT/CCC library is at the application level and it does not need root privilege to be installed. Meanwhile, it was specially developed to require very few changes to the existing applications. This paper describes its design, implementation, and evaluation.
Yunhong Gu, Robert L. Grossman
SC1
2005 Teraflows over Gigabit WANs with UDT
Robert L. Grossman, Yunhong Gu, Xinwei Hong, Antony Antony, Johan Blom, Freek Dijkstra, Cees T. A. M. de Laat
Future Gener. Comput. Syst.2
2004 Experiences in Design and Implementation of a High Performance Transport Protocol
abstract
This paper describes our experiences in the development of the UDP-based Data Transport (UDT) protocol, an application level transport protocol used in distributed data intensive applications. The new protocol is motivated by the emergence of wide area high-speed optical networks, in which TCP is often found to fail to utilize the abundant bandwidth. UDT demonstrates good efficiency and fairness (including RTT fairness and TCP friendliness) characteristics in high performance computing applications where a small number of bulk sources share the abundant bandwidth. It combines both rate and window control and uses bandwidth estimation to determine the control parameters automatically. This paper presents the rationale behind UDT: how UDT integrates these schemes to support high performance data transfer, why these schemes are used, and what the main issues are in the design and implementation of this high performance transport protocol.
Yunhong Gu, Xinwei Hong, Robert L. Grossman
SC1
2004 Experimental studies of data transport and data access of earth-science data over networks with high bandwidth delay products
Robert L. Grossman, Yunhong Gu, David Hanley, Xinwei Hong, Babu Krishnaswamy
Comput. Networks2
2003 Experimental studies using photonic data services at IGrid 2002
Robert L. Grossman, Yunhong Gu, Don Hamelburg, David Hanley, Xinwei Hong, Jorge Levera, David J. Lillethun, Marco Mazzucco, Joe Mambretti, Jeremy Weinberger
Future Gener. Comput. Syst.2
2003 The Photonic TeraStream: enabling next generation applications through intelligent optical networking at iGRID2002
Joe Mambretti, Jeremy Weinberger, Jim Hao Chen, Elizabeth Bacon, Fei Yeh, David J. Lillethun, Robert L. Grossman, Yunhong Gu, Marco Mazzucco
Future Gener. Comput. Syst.8
2003 SABUL: A Transport Protocol for Grid Computing
Yunhong Gu, Robert L. Grossman
J. Grid Comput.1
2003 Data webs for earth science data
Asvin Ananthanarayan, Rajiv Balachandran, Robert L. Grossman, Yunhong Gu, Xinwei Hong, Jorge Levera, Marco Mazzucco
Parallel Comput.4