Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Lianghong Xu

dblp:68/2865 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
0since 2021 · last 2017
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-authorComputer networks · 2Databases, data management, data science and information retrieval · 2 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
Storage systems · 67% Cloud and datacenter computing · 26% Performance modeling and evaluation · 7%
Computer networks
2 papers
Network management and operations · 40% Internet architecture and protocols · 30% Routing and switching · 30%
Databases, data mining, and information retrieval
1 paper
Database system architecture and tuning · 100%

Topics — the 10 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems
distributed storage
0.422014
Agility and Performance in Elastic Distributed Storage · ACM Trans. Storage 2014
SpringFS: bridging agility and performance in elastic distributed storage · FAST 2014
Storage systems › data reduction
data deduplication
0.312017
Online Deduplication for Databases · SIGMOD Conference 2017
Storage systems › data compression
delta compression
0.312017
Online Deduplication for Databases · SIGMOD Conference 2017
Cloud and datacenter computing
cloud storage
0.212014
SpringFS: bridging agility and performance in elastic distributed storage · FAST 2014
Storage systems
data migration
0.212014
Agility and Performance in Elastic Distributed Storage · ACM Trans. Storage 2014
Cloud and datacenter computing › datacenter storage
elastic storage
0.212014
Agility and Performance in Elastic Distributed Storage · ACM Trans. Storage 2014
Network management and operations › performance management
performance diagnosis
0.112011
Diagnosing Performance Changes by Comparing Request Flows · NSDI 2011
Internet architecture and protocols › packet processing
packet classification
0.112009
Packet Classification Algorithms: From Theory to Practice · INFOCOM 2009
Routing and switching
packet forwarding
0.112009
Packet Classification Algorithms: From Theory to Practice · INFOCOM 2009
Authentication and access control
access control
0.012009
Packet Classification Algorithms: From Theory to Practice · INFOCOM 2009

Methods — techniques the papers use, named apart from their topics

hypersplit · 0.2hicuts · 0.2HSM · 0.2trace analysis · 0.2read offloading · 0.2passive migration · 0.2bounded write offloading · 0.2
YearPublicationVenuePosition
2017 Online Deduplication for Databases
abstract
dbDedup is a similarity-based deduplication scheme for on-line database management systems (DBMSs). Beyond block-level compression of individual database pages or operation log (oplog) messages, as used in today's DBMSs, dbDedup uses byte-level delta encoding of individual records within the database to achieve greater savings. dbDedup's single-pass encoding method can be integrated into the storage and logging components of a DBMS to provide two benefits: (1) reduced size of data stored on disk beyond what traditional compression schemes provide, and (2) reduced amount of data transmitted over the network for replication services. To evaluate our work, we implemented dbDedup in a distributed NoSQL DBMS and analyzed its properties using four real datasets. Our results show that dbDedup achieves up to 37x reduction in the storage size and replication traffic of the database on its own and up to 61x reduction when paired with the DBMS's block-level compression. dbDedup provides both benefits with negligible effect on DBMS throughput or client latency (average and tail).
Lianghong Xu, Andrew Pavlo, Sudipta Sengupta, Gregory R. Ganger
SIGMOD Conference1
2015 Reducing replication bandwidth for distributed document databases
abstract
With the rise of large-scale, Web-based applications, users are increasingly adopting a new class of document-oriented database management systems (DBMSs) that allow for rapid prototyping while also achieving scalable performance. Like for other distributed storage systems, replication is important for document DBMSs in order to guarantee availability. The network bandwidth required to keep replicas synchronized is expensive and is often a performance bottleneck. As such, there is a strong need to reduce the replication bandwidth, especially for geo-replication scenarios where wide-area network (WAN) bandwidth is limited.
Lianghong Xu, Andrew Pavlo, Sudipta Sengupta, Jin Li 0001, Gregory R. Ganger
SoCC1
2014 Exploiting iterative-ness for parallel ML computations
abstract
Many large-scale machine learning (ML) applications use iterative algorithms to converge on parameter values that make the chosen model fit the input data. Often, this approach results in the same sequence of accesses to parameters repeating each iteration. This paper shows that these repeating patterns can and should be exploited to improve the efficiency of the parallel and distributed ML applications that will be a mainstay in cloud computing environments. Focusing on the increasingly popular "parameter server" approach to sharing model parameters among worker threads, we describe and demonstrate how the repeating patterns can be exploited. Examples include replacing dynamic cache and server structures with static pre-serialized structures, informing prefetch and partitioning decisions, and determining which data should be cached at each thread to avoid both contention and slow accesses to memory banks attached to other sockets. Experiments show that such exploitation reduces per-iteration time by 33--98%, for three real ML workloads, and that these improvements are robust to variation in the patterns over time.
Henggang Cui, Alexey Tumanov, Jinliang Wei, Lianghong Xu, Wei Dai 0003, Jesse Haber-Kucharsky, Qirong Ho, Gregory R. Ganger, Phillip B. Gibbons, Garth A. Gibson, Eric P. Xing
SoCC4
2014 SpringFS: bridging agility and performance in elastic distributed storage
Lianghong Xu, James Cipar, Elie Krevat, Alexey Tumanov, Nitin Gupta 0001, Michael A. Kozuch, Gregory R. Ganger
FAST1
2014 Agility and Performance in Elastic Distributed Storage
abstract
Elastic storage systems can be expanded or contracted to meet current demand, allowing servers to be turned off or used for other tasks. However, the usefulness of an elastic distributed storage system is limited by its agility: how quickly it can increase or decrease its number of servers. Due to the large amount of data they must migrate during elastic resizing, state of the art designs usually have to make painful trade-offs among performance, elasticity, and agility. This article describes the state of the art in elastic storage and a new system, called SpringFS, that can quickly change its number of active servers, while retaining elasticity and performance goals. SpringFS uses a novel technique, termed bounded write offloading , that restricts the set of servers where writes to overloaded servers are redirected. This technique, combined with the read offloading and passive migration policies used in SpringFS, minimizes the work needed before deactivation or activation of servers. Analysis of real-world traces from Hadoop deployments at Facebook and various Cloudera customers and experiments with the SpringFS prototype confirm SpringFS’s agility, show that it reduces the amount of data migrated for elastic resizing by up to two orders of magnitude, and show that it cuts the percentage of active servers required by 67--82%, outdoing state-of-the-art designs by 6--120%.
Lianghong Xu, James Cipar, Elie Krevat, Alexey Tumanov, Nitin Gupta 0001, Michael A. Kozuch, Gregory R. Ganger
ACM Trans. Storage1
2011 Diagnosing Performance Changes by Comparing Request Flows
Raja R. Sambasivan, Alice X. Zheng, Michael De Rosa, Elie Krevat, Spencer Whitman, Michael Stroucken, Lianghong Xu, Gregory R. Ganger
NSDI8
2009 Packet Classification Algorithms: From Theory to Practice
abstract
During the past decade, the packet classification problem has been widely studied to accelerate network applications such as access control, traffic engineering and intrusion detection. In our research, we found that although a great number of packet classification algorithms have been proposed in recent years, unfortunately most of them stagnate in mathematical analysis or software simulation stages and few of them have been implemented in commercial products as a generic solution. To fill the gap between theory and practice, in this paper, we propose a novel packet classification algorithm named HyperSplit. Compared to the well-known HiCuts and HSM algorithms, HyperSplit achieves superior performance in terms of classification speed, memory usage and preprocessing time. The practicability of the proposed algorithm is manifested by two facts in our test: HyperSplit is the only algorithm that can successfully handle all the rule sets; HyperSplit is also the only algorithm that reaches more than 6Gbps throughput on the Octeon3860 multi-core platform when tested with 64-byte Ethernet packets against 10K ACL rules.
Yaxuan Qi, Lianghong Xu, Baohua Yang, Yibo Xue, Jun Li 0003
INFOCOM2
2001 Geometric Hermite interpolation for space curves
Lianghong Xu, Jianhong Shi
Comput. Aided Geom. Des.1