Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jing Li 0021

dblp:l/JingLi21 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
0since 2021 · last 2017
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 2 first-authorSystems, architecture and hardware · 3 · 1 first-authorArtificial intelligence and machine learning · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Storage systems · 86% Energy-efficient computing · 9% Cloud and datacenter computing · 4%
Databases, data mining, and information retrieval
2 papers
Recommender systems · 54% Query processing and optimization · 46%

Topics — the 11 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems › flash and SSD
solid-state drive
0.422017
Summarizer: trading communication with computing near storage · MICRO 2017
HippogriffDB: Balancing I/O and GPU Bandwidth in Big Data Analytics · Proc. VLDB Endow. 2016
Storage systems › flash and SSD › flash memory management
flash translation layer
0.312017
Summarizer: trading communication with computing near storage · MICRO 2017
Storage systems › computational storage
in-storage computing
0.312017
Summarizer: trading communication with computing near storage · MICRO 2017
Storage systems › computational storage
near-storage computing
0.312017
Summarizer: trading communication with computing near storage · MICRO 2017
Query processing and optimization
OLAP
0.212016
HippogriffDB: Balancing I/O and GPU Bandwidth in Big Data Analytics · Proc. VLDB Endow. 2016
Storage systems › i/o architecture
SSD-GPU peer-to-peer DMA
0.212016
HippogriffDB: Balancing I/O and GPU Bandwidth in Big Data Analytics · Proc. VLDB Endow. 2016
Storage systems › storage devices › storage media
mobile storage
0.212014
On the energy overhead of mobile storage systems · FAST 2014
Recommender systems
graph-based recommendation
0.112012
Challenging the Long Tail Recommendation · Proc. VLDB Endow. 2012
Recommender systems › beyond-accuracy recommendation
long-tail recommendation
0.112012
Challenging the Long Tail Recommendation · Proc. VLDB Endow. 2012
Cloud and datacenter computing
datacenter storage
0.112017
Summarizer: trading communication with computing near storage · MICRO 2017
Storage systems › flash and SSD › solid-state drive
NVMe SSD
0.112016
HippogriffDB: Balancing I/O and GPU Bandwidth in Big Data Analytics · Proc. VLDB Endow. 2016

Methods — techniques the papers use, named apart from their topics

operator fusion · 0.5double buffering · 0.5compression · 0.5embedded core offloading · 0.3hitting time · 0.1entropy-biased absorbing cost · 0.1absorbing time · 0.1
YearPublicationVenuePosition
2017 Summarizer: trading communication with computing near storage
abstract
Modern data center solid state drives (SSDs) integrate multiple general-purpose embedded cores to manage flash translation layer, garbage collection, wear-leveling, and etc., to improve the performance and the reliability of SSDs. As the performance of these cores steadily improves there are opportunities to repurpose these cores to perform application driven computations on stored data, with the aim of reducing the communication between the host processor and the SSD. Reducing host-SSD bandwidth demand cuts down the I/O time which is a bottleneck for many applications operating on large data sets. However, the embedded core performance is still significantly lower than the host processor, as generally wimpy embedded cores are used within SSD for cost effective reasons. So there is a trade-off between the computation overhead associated with near SSD processing and the reduction in communication overhead to the host system.
Gunjae Koo, Kiran Kumar Matam, Te I, Krishna Narra, Jing Li 0021, Hung-Wei Tseng 0001, Steven Swanson, Murali Annavaram
MICRO5
2016 Hippogriff: Efficiently moving data in heterogeneous computing systems
abstract
Data movement between the compute and the storage (e.g., GPU and SSD) has been a long-neglected problem in heterogeneous systems, while the inefficiency in existing systems does cause significant loss in both performance and energy efficiency. This paper presents Hippogriff to provide a high-level programming model to simplify data movement between the compute and the storage, and to dynamically schedule data transfers based on system load. By eliminating unnecessary data movement, Hippogriff can speedup single program workloads by 1.17×, and save 17% energy. For multi-program workloads, Hippogriff shows 1.25× speedup. Hippogriff also improves the performance of a GPU-based MapReduce framework by 27%.
Yang Liu 0044, Hung-Wei Tseng 0001, Mark Gahagan, Jing Li 0021, Yanqin Jin, Steven Swanson
ICCD4
2016 HippogriffDB: Balancing I/O and GPU Bandwidth in Big Data Analytics
abstract
As data sets grow and conventional processor performance scaling slows, data analytics move towards heterogeneous architectures that incorporate hardware accelerators (notably GPUs) to continue scaling performance. However, existing GPU-based databases fail to deal with big data applications efficiently: their execution model suffers from scalability limitations on GPUs whose memory capacity is limited; existing systems fail to consider the discrepancy between fast GPUs and slow storage, which can counteract the benefit of GPU accelerators. In this paper, we propose HippogriffDB, an efficient, scalable GPU-accelerated OLAP system. It tackles the bandwidth discrepancy using compression and an optimized data transfer path. HippogriffDB stores tables in a compressed format and uses the GPU for decompression, trading GPU cycles for the improved I/O bandwidth. To improve the data transfer efficiency, HippogriffDB introduces a peer-to-peer, multi-threaded data transfer mechanism, directly transferring data from the SSD to the GPU. HippogriffDB adopts a query-over-block execution model that provides scalability using a stream-based approach. The model improves kernel efficiency with the operator fusion and double buffering mechanism. We have implemented HippogriffDB using an NVMe SSD, which talks directly to a commercial GPU. Results on two popular benchmarks demonstrate its scalability and efficiency. HippogriffDB outperforms existing GPU-based databases (YDB) and in-memory data analytics (MonetDB) by 1-2 orders of magnitude.
Jing Li 0021, Hung-Wei Tseng 0001, Chunbin Lin, Yannis Papakonstantinou, Steven Swanson
Proc. VLDB Endow.1
2014 On the energy overhead of mobile storage systems
Jing Li 0021, Anirudh Badam, Ranveer Chandra, Steven Swanson, Bruce L. Worthington, Qi Zhang 0012
FAST1
2014 HAT: an efficient buffer management method for flash-based hybrid storage systems
Yanfei Lv, Bin Cui 0001, Xuexuan Chen, Jing Li 0021
Frontiers Comput. Sci.4
2013 Hotness-aware buffer management for flash-based hybrid storage systems
abstract
Flash solid-state drives (SSDs) provide much faster access to data compared with traditional hard disk drives (HDDs). The current price and performance of SSD suggest it can be adopted as a data buffer between main memory and HDD, and buffer management policy in such hybrid systems has attracted more and more interest from research community recently. In this paper, we propose a novel approach to manage the buffer in flash-based hybrid storage systems, named Hotness Aware Hit (HAT). HAT exploits a page reference queue to record the access history as well as the status of accessed pages, i.e., hot, warm and cold. Additionally, the page reference queue is further split into hot and warm regions which correspond to the memory and flash in general. The HAT approach updates the page status and deals with the page migration in the memory hierarchy according to the current page status and hit position in the page reference queue. Our empirical evaluation on benchmark traces demonstrates the superiority of the proposed strategy against the state-of-the-art competitors.
Yanfei Lv, Bin Cui 0001, Xuexuan Chen, Jing Li 0021
CIKM4
2012 Challenging the Long Tail Recommendation
abstract
The success of "infinite-inventory" retailers such as Amazon.com and Netflix has been largely attributed to a "long tail" phenomenon. Although the majority of their inventory is not in high demand, these niche products, unavailable at limited-inventory competitors, generate a significant fraction of total revenue in aggregate. In addition, tail product availability can boost head sales by offering consumers the convenience of "one-stop shopping" for both their mainstream and niche tastes. However, most of existing recommender systems, especially collaborative filter based methods, can not recommend tail products due to the data sparsity issue. It has been widely acknowledged that to recommend popular products is easier yet more trivial while to recommend long tail products adds more novelty yet it is also a more challenging task. In this paper, we propose a novel suite of graph-based algorithms for the long tail recommendation. We first represent user-item information with undirected edge-weighted graph and investigate the theoretical foundation of applying Hitting Time algorithm for long tail item recommendation. To improve recommendation diversity and accuracy, we extend Hitting Time and propose efficient Absorbing Time algorithm to help users find their favorite long tail items. Finally, we refine the Absorbing Time algorithm and propose two entropy-biased Absorbing Cost algorithms to distinguish the variation on different user-item rating pairs, which further enhances the effectiveness of long tail recommendation. Empirical experiments on two real life datasets show that our proposed algorithms are effective to recommend long tail items and outperform state-of-the-art recommendation techniques.
Hongzhi Yin, Bin Cui 0001, Jing Li 0021, Chen Chen 0056
Proc. VLDB Endow.3