VLDB 2026 Research / reviewers in the wild / expert
Min Lv
dblp:13/6723
· DBLP profile ↗
9ranked-venue papers
0as first author
0since 2021 · last 2020
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Storage systems · 62% Distributed systems · 22% Hardware reliability and fault tolerance · 11% | |
| Databases, data mining, and information retrieval
2 papers |
Web and social media mining · 58% Data mining · 42% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Computational social science and digital humanities · 100% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems
data layout |
0.4 | 1 | 2020 | PDL: A Data Layout towards Fast Failure Recovery for Erasure-coded Distributed Storage Systems · INFOCOM 2020 |
Storage systems › storage reliability
erasure coding |
0.4 | 1 | 2020 | PDL: A Data Layout towards Fast Failure Recovery for Erasure-coded Distributed Storage Systems · INFOCOM 2020 |
Distributed systems › fault tolerance
failure recovery |
0.4 | 1 | 2020 | PDL: A Data Layout towards Fast Failure Recovery for Erasure-coded Distributed Storage Systems · INFOCOM 2020 |
Data mining › structured data mining › graph mining
graph sampling |
0.4 | 1 | 2019 | Walking with Perception: Efficient Random Walk Sampling via Common Neighbor Awareness · ICDE 2019 |
Web and social media mining › social network sampling
random walk sampling |
0.4 | 1 | 2019 | Walking with Perception: Efficient Random Walk Sampling via Common Neighbor Awareness · ICDE 2019 |
Storage systems › file systems
distributed file system |
0.4 | 1 | 2019 | Explicit Data Correlations-Directed Metadata Prefetching Method in Distributed File Systems · IEEE Trans. Parallel Distributed Syst. 2019 |
Storage systems › metadata management
metadata prefetching |
0.4 | 1 | 2019 | Explicit Data Correlations-Directed Metadata Prefetching Method in Distributed File Systems · IEEE Trans. Parallel Distributed Syst. 2019 |
Computational social science and digital humanities › social computing
user behavior analysis |
0.1 | 1 | 2012 | Users sleeping time analysis based on micro-blogging data · UbiComp 2012 |
Hardware reliability and fault tolerance › soft errors
single-event upset |
0.1 | 1 | 2012 | Bitstream decoding and SEU-induced failure analysis in SRAM-based FPGAs · Sci. China Inf. Sci. 2012 |
Hardware reliability and fault tolerance
soft errors |
0.1 | 1 | 2012 | Bitstream decoding and SEU-induced failure analysis in SRAM-based FPGAs · Sci. China Inf. Sci. 2012 |
Reconfigurable computing and FPGAs › FPGA architecture
SRAM-based FPGA |
0.1 | 1 | 2012 | Bitstream decoding and SEU-induced failure analysis in SRAM-based FPGAs · Sci. China Inf. Sci. 2012 |
Distributed systems
fault tolerance |
0.1 | 1 | 2020 | PDL: A Data Layout towards Fast Failure Recovery for Erasure-coded Distributed Storage Systems · INFOCOM 2020 |
Methods — techniques the papers use, named apart from their topics
combinatorial design · 0.4weighted random walk · 0.4unbiased sampling · 0.4pattern matching · 0.4extended attributes · 0.4adaptive feedback · 0.4time zone inference · 0.3activity data mining · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | PDL: A Data Layout towards Fast Failure Recovery for Erasure-coded Distributed Storage SystemsabstractErasure coding becomes increasingly popular in distributed storage systems (DSSes) for providing high reliability with low storage overhead. However, traditional random data placement causes massive cross-rack traffic and severely unbalanced load during failure recovery, degrading the recovery performance significantly. In addition, various erasure coding policies coexisting in a DSS exacerbates the above problem. In this paper, we propose PDL, a PBD-based Data Layout, to optimize failure recovery performance in DSSes. PDL is constructed based on Pairwise Balanced Design, a combinatorial design scheme with uniform mathematical properties, and thus presents a uniform data layout. Then we propose rPDL, a failure recovery scheme based on PDL. rPDL reduces cross-rack traffic effectively and provides nearly balanced cross-rack traffic distribution by uniformly choosing replacement nodes and retrieving determined available blocks to recover the lost blocks. We implemented PDL and rPDL in Hadoop 3.1.1. Compared with existing data layout of HDFS, experimental results show that rPDL reduces degraded read latency by an average of 62.83%, delivers 6.27× data recovery throughput, and provides evidently better support for front-end applications. Liangliang Xu, Min Lv, Zhipeng Li 0005, Cheng Li 0001, Yinlong Xu 0001 |
INFOCOM | 2 |
| 2019 | HP-Mapper: A High Performance Storage Driver for Docker ContainersabstractDocker containers are widely deployed to provide lightweight virtualization, and they have many desirable features such as ease of deployment and near bare-metal performance. However, both the performance and cache efficiency of containers are still limited by their storage drivers due to the coarse-grained copy-on-write operations, and the large amount of redundancy in both I/O requests and page cache. To improve I/O performance and cache efficiency of containers, we develop HP-Mapper, a high performance storage driver for Docker containers. HP-Mapper provides a two-level mapping strategy to support fine-grained copy-on-write with low overhead, and an efficient interception method to reduce redundant I/Os. Furthermore, it uses a novel cache management mechanism to reduce duplicate cached data. Experiment results with our prototype system show that HP-Mapper significantly reduces copy-on-write latency due to its finer-grained copy-on-write scheme. Moreover, HP-Mapper can also reduce 65.4% cache usage on average due to elimination of duplicated data. As a result, HP-Mapper improves the throughput of real-world workloads by up to 39.4%, and improves the startup speed of containers by 2.0x. Fan Guo 0003, Yongkun Li 0001, Min Lv, Yinlong Xu 0001, John C. S. Lui |
SoCC | 3 |
| 2019 | Walking with Perception: Efficient Random Walk Sampling via Common Neighbor AwarenessabstractRandom walk is widely applied to sample large-scale graphs due to its simplicity of implementation and solid theoretical foundations of bias analysis. However, its computational efficiency is heavily limited by the slow convergence rate (a.k.a. long burn-in period). To address this issue, we propose a common neighbor aware random walk framework called CNARW, which leverages weighted walking by differentiating the next-hop candidate nodes to speed up the convergence. Specifically, CNARW takes into consideration the common neighbors between previously visited nodes and next-hop candidate nodes in each walking step. Based on CNARW, we further develop two efficient "unbiased sampling" schemes. Experimental results on real-world network datasets show that our approach converges remarkably faster than the state-of-the-art random walk sampling algorithms. Furthermore, to achieve the same estimation accuracy, our approach reduces the query cost (a measure of sampling budget) significantly. Lastly, we also use two case studies to demonstrate the effectiveness of our sampling framework in solving large-scale graph analysis tasks. Yongkun Li 0001, Zhiyong Wu 0004, Hong Xie 0004, Min Lv, Yinlong Xu 0001, John C. S. Lui |
ICDE | 5 |
| 2019 | D3: Deterministic Data Distribution for Efficient Data Reconstruction in Erasure-Coded Distributed Storage SystemsabstractDue to individual unreliable commodity components, failures are common in large-scale distributed storage systems. Erasure codes are widely deployed in practical storage systems to provide fault tolerance with low storage overhead. However, the commonly used random data placement in storage systems based on erasure codes induces to heavy crossrack traffic, load imbalance, and random access, which slow down the recovery process upon failures. In this paper, with orthogonal arrays, we define a Deterministic Data Distribution (D3) of blocks to nodes and racks, and propose an efficient failure recovery approach based on D3. D3not only uniformly distributes data/parity blocks among storage servers, but also balances the repair traffic among racks and storage servers for failure recovery. Furthermore, D3also minimizes the cross-rack repair traffic for data layouts against a single rack failure and provides sequential access for failure recovery. We implement D3in Hadoop Distributed File System (HDFS) with a cluster of 28 machines. Our experiments show that D3significantly speeds up the failure recovery process compared with random data distribution, e.g., 2.21 times for (6, 3)-RS code in a system consisting of eight racks and three nodes in each rack. Zhipeng Li 0005, Min Lv, Yinlong Xu 0001, Yongkun Li 0001, Liangliang Xu |
IPDPS | 2 |
| 2019 | Explicit Data Correlations-Directed Metadata Prefetching Method in Distributed File SystemsabstractMetadata performance in distributed file systems (DFS) is critical, due to the following trends: (a) the growing size of modern storage systems is expected to exceed billions of files and most files are small; (b) over half of the file accesses are metadata operations. In this work, we present SMeta, a metadata prefetching method that is seamlessly integrated into DFS for easy-of-use and significantly scales the metadata performance. Previous prefetching proposals primarily focus on mining groups of files that tend to be accessed together from the access history. Nevertheless, our study discovered that these solutions likely miss a huge number of correlated files whose co-occurrence frequency is not high enough. Unlike access correlations, we take a novel and completely different approach to explore explicit data correlations by understanding the reference relationships between files encoded in some forms of hyperlinks, which naturally exist in many applications. To embrace this new concept, SMeta explores correlations upon files are written via a light-weight pattern matching algorithm, stores correlations in the reserved extended attributes of file metadata to avoid changes in DFS APIs, and collapses multiple I/O rounds for accessing metadata of the target file and its data-correlated files into one round. A cost-efficient adaptive feedback mechanism is introduced to improve prefetching accuracy. We implemented SMeta atop of Ceph and evaluated it using synthetic and real system workloads. Compared to baselines, SMeta provides better metadata performance in terms of latency, throughput and scalability. Youxu Chen, Cheng Li 0001, Min Lv, Xinyang Shao, Yongkun Li 0001, Yinlong Xu 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2018 | LCR: Load-Aware Cache Replacement Algorithm for Flash-Based SSDsabstractFlash-based SSDs are usually equipped with an onboard cache to further improve system performance by smoothing the gap between the upper-level applications and lower-level flash chips. Since modern SSDs are usually composed of multiple flash chips, and the load of flash chips are significantly different, it is very meaningful to be aware of the chip load condition when designing a cache replacement algorithm. Nevertheless, existing cache replacement algorithms only consider to reduce the cache miss ratio so as to reduce the I/O requests to the underlying flash memory as much as possible, none of them considers the load condition of flash chips. In this paper, we propose a Load- aware Cache Replacement algorithm, called LCR, to improve the performance of flash-based SSDs. The basic idea is to give a higher priority to cache the blocks on overloaded flash chips. We evaluate the performance of our scheme by using a trace- driven simulator with multiple real-world workloads, and results show that compared with the most common algorithm LRU and the state-of-the-art algorithm GCaR, LCR reduces the average response time by as much as 39.2% and 12.3%, respectively. Caiyin Liu, Min Lv, Yubiao Pan, Hao Chen 0080, Yongkun Li 0001, Cheng Li 0001, Yinlong Xu 0001 |
NAS | 2 |
| 2016 | Reliability Modeling on Consecutive-kr -out-of-nr: F Linear Zigzag Structure and Circular Polygon StructureabstractWe propose the consecutive-kr-out-of-nr:F linear zigzag structure and circular polygon structure, which extend the consecutive-k-out-of-n:F system model. By employing the finite Markov chain imbedding approach, we have proven the equations for calculating the reliability of consecutive-k-out-of-n:F system with different patterns given the state of the first and(or) the last component(s). Based on those reliability equations, we provide the rules for obtaining the reliability of the linear zigzag structure and circular polygon structure. Numerical examples are presented as an illustration of our new system model and demonstrate the use of the equations and rules proposed in this paper. Lirong Cui, David W. Coit, Min Lv |
IEEE Trans. Reliab. | 4 |
| 2012 | Users sleeping time analysis based on micro-blogging dataabstractThe emergence of new social network services, often labeled as Web 2.0, has permitted an amazingly increase of user generated content. In particular, Sina Weibo, a popular Chinese micro-blogging service is designed as platforms allowing users to generate contents that open to the public. From analyzing activates of submitting posts to Sina Weibo, some features of users can be estimated. This paper aims to contribute to this growing body of literature by studying how users' frequent activities reflect their sleeping time and living time zones. By mining a large set of users' activates data from Sina Weibo, we demonstrate its possible role to detect the sleeping time of users and find a new method for judging users' time zone. Guangzhong Sun, Min Lv |
UbiComp | 3 |
| 2012 | Bitstream decoding and SEU-induced failure analysis in SRAM-based FPGAs
Zhongming Wang, Zhibin Yao, Min Lv |
Sci. China Inf. Sci. | 4 |