Yijing Luo

dblp:23/2850 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
3since 2021 · last 2024
0000-0002-8846-4875ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Storage systems · 53% Memory systems · 47%
Databases, data mining, and information retrieval
2 papers
Indexing and storage engines · 64% Data mining · 36%
Artificial intelligence
1 paper
Trustworthy machine learning · 77% Learning theory · 23%

Topics — the 12 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › non-volatile memory
persistent memory
0.922024
FlatLSM: Write-Optimized LSM-Tree for PM-Based KV Stores · ACM Trans. Storage 2023
When Learned Indexes Meet Persistent Memory: The Analysis and the Optimization · IEEE Trans. Knowl. Data Eng. 2024
Indexing and storage engines
learned index
0.812024
When Learned Indexes Meet Persistent Memory: The Analysis and the Optimization · IEEE Trans. Knowl. Data Eng. 2024
Indexing and storage engines › learned index
persistent memory learned index
0.812024
When Learned Indexes Meet Persistent Memory: The Analysis and the Optimization · IEEE Trans. Knowl. Data Eng. 2024
Storage systems
key-value storage
0.712023
FlatLSM: Write-Optimized LSM-Tree for PM-Based KV Stores · ACM Trans. Storage 2023
Storage systems › key-value storage
LSM-tree
0.712023
FlatLSM: Write-Optimized LSM-Tree for PM-Based KV Stores · ACM Trans. Storage 2023
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels
0.412020
A Bi-level Formulation for Label Noise Learning with Spectral Cluster Discovery · IJCAI 2020
Data mining
clustering
0.412020
A Bi-level Formulation for Label Noise Learning with Spectral Cluster Discovery · IJCAI 2020
Data mining › clustering
spectral clustering
0.412020
A Bi-level Formulation for Label Noise Learning with Spectral Cluster Discovery · IJCAI 2020
Memory systems
non-volatile memory
0.212024
When Learned Indexes Meet Persistent Memory: The Analysis and the Optimization · IEEE Trans. Knowl. Data Eng. 2024
Memory systems › non-volatile memory › persistent memory
persistent memory indexing
0.212024
When Learned Indexes Meet Persistent Memory: The Analysis and the Optimization · IEEE Trans. Knowl. Data Eng. 2024
Storage systems › i/o optimization
write optimization
0.212023
FlatLSM: Write-Optimized LSM-Tree for PM-Based KV Stores · ACM Trans. Storage 2023
Machine learning › Learning theory
classification
0.112020
A Bi-level Formulation for Label Noise Learning with Spectral Cluster Discovery · IJCAI 2020

Methods — techniques the papers use, named apart from their topics

selective persistence · 1.5cost-based insertion pattern selection · 1.5low-rank approximation · 0.9bi-level optimization · 0.9affinity graph · 0.9parallel flush/compaction · 0.7KV separation · 0.7
YearPublicationVenuePosition
2024 When Learned Indexes Meet Persistent Memory: The Analysis and the Optimization
abstract
The emerging persistent memory (PM) is increasingly being leveraged to construct high-performance and persistent indexes. By exploiting data distribution, recent learned indexes open up a new index design paradigm. Some prior studies try to refit the learned index according to the features of PM. However, they neglect to analyze the performance of existing learned index schemes on PM. In this paper, we provide a comprehensive analysis of learned indexes on PM and propose two optimization methods to improve the performance. In particular, we evaluate ALEX, PGM-index, and XIndex after converting them to persistent indexes. With appropriate modifications, some design choices of volatile learned index still show favorable performance on PM under workloads with simple data distribution. But they perform poorly when the data distribution becomes complex. According to the experiment results, we summarize some instructive insights and optimize persistent learned indexes for complex data distributions with two methods: 1) a cost-based insertion pattern selection to minimize PM writes and 2) recoverable internal nodes selective persistence to decrease the overhead of internal lookups. Our evaluations demonstrate the performance of optimized ALEX is 2.09x/1.53x of the original ALEX in insert/search. Meanwhile, it also outperforms the specific-designed persistent learned index.
Lixiao Cui, Yijing Luo, Yusen Li, Gang Wang 0001, Xiaoguang Liu 0001
IEEE Trans. Knowl. Data Eng.2
2023 FlatLSM: Write-Optimized LSM-Tree for PM-Based KV Stores
abstract
The Log-Structured Merge Tree (LSM-Tree) is widely used in key-value (KV) stores because of its excwrite performance. But LSM-Tree-based KV stores still have the overhead of write-ahead log and write stall caused by slow L 0 flush and L 0 - L 1 compaction. New byte-addressable, persistent memory (PM) devices bring an opportunity to improve the write performance of LSM-Tree. Previous studies on PM-based LSM-Tree have not fully exploited PM’s “dual role” of main memory and external storage. In this article, we analyze two strategies of memtables based on PM and the reasons write stall problems occur in the first place. Inspired by the analysis result, we propose FlatLSM, a specially designed flat LSM-Tree for non-volatile memory based KV stores. First, we propose PMTable with separated index and data. The PM Log utilizes the Buffer Log to store KVs of size less than 256B. Second, to solve the write stall problem, FlatLSM merges the volatile memtables and the persistent L 0 into large PMTables, which can reduce the depth of LSM-Tree and concentrate I/O bandwidth on L 0 - L 1 compaction. To mitigate write stall caused by flushing large PMTables to SSD, we propose a parallel flush/compaction algorithm based on KV separation. We implemented FlatLSM based on RocksDB and evaluated its performance on Intel’s latest PM device, the Intel Optane DC PMM with the state-of-the-art PM-based LSM-Tree KV stores, FlatLSM improves the throughput 5.2× on random write workload and 2.55× on YCSB-A.
Kewen He, Yujie An, Yijing Luo, Xiaoguang Liu 0001, Gang Wang 0001
ACM Trans. Storage3
2022 Multi-class Label Noise Learning via Loss Decomposition and Centroid Estimation
abstract
In real-world scenarios, many large-scale datasets often contain inaccurate labels, i.e., noisy labels, which may confuse model training and lead to performance degradation. To overcome this issue, Label Noise Learning (LNL) has recently attracted much attention, and various methods have been proposed to design an unbiased risk estimator to the noise-free dataset to combat such label noise. Among them, a trend of works based on Loss Decomposition and Centroid Estimation (LDCE) has shown very promising performance. However, existing LNL methods based on LDCE are only designed for binary classification, and they are not directly extendable to multi-class situations. In this paper, we propose a novel multi-class robust learning method for LDCE, which is termed “MC-LDCE”. Specifically, we decompose the commonly adopted loss (e.g., mean squared loss) function into a label-dependent part and a label-independent part, in which only the former is influenced by label noise. Further, by defining a new form of data centroid, we transform the recovery problem of a label-dependent part to a centroid estimation problem. Finally, by critically examining the mathematical expectation of clean data centroid given the observed noisy set, the centroid can be estimated which helps to build an unbiased risk estimator for multi-class learning. The proposed MC-LDCE method is general and applicable to different types (i.e., linear and nonlinear) of classification models. The experimental results on five public datasets demonstrate the superiority of the proposed MC-LDCE against other representative LNL methods in tackling multi-class label noise problem.
Yongliang Ding, Tao Zhou 0002, Yijing Luo, Chen Gong 0002
SDM4
2020 A Bi-level Formulation for Label Noise Learning with Spectral Cluster Discovery
abstract
Practically, we often face the dilemma that some of the examples for training a classifier are incorrectly labeled due to various subjective and objective factors. Although intensive efforts have been put to design classifiers that are robust to label noise, most of the previous methods have not fully utilized data distribution information. To address this issue, this paper introduces a bi-level learning paradigm termed “Spectral Cluster Discovery'' (SCD) for combating with noisy labels. Namely, we simultaneously learn a robust classifier (Learning stage) by discovering the low-rank approximation to the ground-truth label matrix and learn an ideal affinity graph (Clustering stage). Specifically, we use the learned classifier to assign the examples with similar label to a mutual cluster. Based on the cluster membership, we utilize the learned affinity graph to explore the noisy examples based on the cluster membership. Both stages will reinforce each other iteratively. Experimental results on typical benchmark and real-world datasets verify the superiority of SCD to other label noise learning methods.
Yijing Luo, Bo Han 0003, Chen Gong 0002
IJCAI1
2007 Packet Scheduling for Scalable Video Streaming Over Lossy Packet Access Networks
abstract
Video streaming applications have gained in popularity in recent years. The quality of service offered by such applications is limited by the available transmission rates as well as time-varying conditions, such as, channel fading and network congestion, which lead to packet losses. Scalable video coding techniques that allow for the flexible adaptation of temporal resolution as well as quality of an encoded bitstream can be immensely useful in developing video streaming applications that can adapt to time-varying network and channel conditions. Scalable coding techniques, however, are generally designed to offer progressive refinement, which introduces dependencies between encoded video packets. Therefore, when determining a packet scheduling technique for scalable coded video, the possibility of random packet losses, which might affect the decodability of subsequent packets, must be taken into account. In this paper, we take into account the available transmission rate, possibly time-varying channel conditions, and the possibility of random packet losses, to design a scheduling technique for video packets in a scalable bit-stream. Since the optimal solution to the scheduling problem requires an exhaustive, and therefore, intractable computation, we propose a greedy algorithm that will schedule the optimal packet for transmission at a given transmission opportunity based on the encoded content and the available channel state information. Simulation results show significant gains in performance when the proposed technique is compared to content and channel independent packet scheduling techniques.
Ehsan Maani, Yijing Luo, Peshala V. Pahalawatta, Aggelos K. Katsaggelos
ICCCN2