Yiyu Liu

dblp:263/7231 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Language models and text generation · 46% Efficient and distributed learning · 25% Information extraction and text analysis · 22%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 56% Memory systems · 44%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
instruction following
1.012026
ASKD: Reinforcement Learning-Style Knowledge Distillation with Quality-Adaptive Skewness · AAAI 2026
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
1.012026
ASKD: Reinforcement Learning-Style Knowledge Distillation with Quality-Adaptive Skewness · AAAI 2026
Memory systems › memory disaggregation
memory pooling
0.812024
FaaSMem: Improving Memory Efficiency of Serverless Computing with Memory Pool Architecture · ASPLOS (3) 2024
Cloud and datacenter computing
serverless computing
0.812024
FaaSMem: Improving Memory Efficiency of Serverless Computing with Memory Pool Architecture · ASPLOS (3) 2024
Natural language and speech › Information extraction and text analysis › text classification
multi-task classification
0.412020
MATINF: A Jointly Labeled Large-Scale Dataset for Classification, Question Answering and Summarization · ACL 2020
Natural language and speech › Language models and text generation › text summarization
multi-task summarization
0.412020
MATINF: A Jointly Labeled Large-Scale Dataset for Classification, Question Answering and Summarization · ACL 2020
Natural language and speech › Information extraction and text analysis
text classification
0.412020
MATINF: A Jointly Labeled Large-Scale Dataset for Classification, Question Answering and Summarization · ACL 2020
Natural language and speech › Language models and text generation
text summarization
0.412020
MATINF: A Jointly Labeled Large-Scale Dataset for Classification, Question Answering and Summarization · ACL 2020
Machine learning › Reinforcement learning
policy optimization
0.312026
ASKD: Reinforcement Learning-Style Knowledge Distillation with Quality-Adaptive Skewness · AAAI 2026
Cloud and datacenter computing › serverless computing
container cold start
0.212024
FaaSMem: Improving Memory Efficiency of Serverless Computing with Memory Pool Architecture · ASPLOS (3) 2024
Data mining
dataset construction
0.112020
MATINF: A Jointly Labeled Large-Scale Dataset for Classification, Question Answering and Summarization · ACL 2020

Methods — techniques the papers use, named apart from their topics

reinforcement learning-style optimization · 1.0kullback-leibler divergence · 1.0gradient clipping · 1.0multi-task learning · 0.9memory offloading policy design · 0.8
YearPublicationVenuePosition
2026 ASKD: Reinforcement Learning-Style Knowledge Distillation with Quality-Adaptive Skewness
abstract
Knowledge distillation (KD) is a widely adopted technique for transferring the capabilities of large teacher models to smaller student models, thereby significantly reducing inference costs and memory consumption. However, existing KD methods are all constrained by an inherent greedy optimization objective, rooted in the assumption of teacher superiority: "Trust all teacher-generated outputs (TGOs)" and "Distrust any student-generated outputs (SGOs) unsupported by the teacher". We propose ASKD, a novel KD method with adaptive skewness determined by sample quality, refining this objective to: "Learn TGOs proportionally to their quality, and distrust only low-quality unsupported SGOs". ASKD comprises three key components: (1) A reinforcement learning-style optimization formulation to mitigate the inherent approximation bias in sample-based Kullback-Leibler (KL) divergence approximations used by previous KD methods; (2) Well-designed quality supervision signals to map and achieve adaptive skewness in skewed KL loss, pioneering the usage of sample quality to adjust learning magnitudes; (3) A gradient-clip function on high-quality SGOs for findings that high-quality SGOs in KL loss fail to yield positive updates and even cause adverse effects on some samples. Extensive experiments indicate that ASKD builds high-performance student models across various tasks, including instruction following, mathematical reasoning, and code generation, outperforming state-of-the-art methods comprehensively and surpassing GRPO-like approaches that use advantages as multiplicative factors. We also provide detailed mathematical proofs demonstrating properties such as Lipschitz continuity of the update coefficient and uniform convergence of the loss function, ensuring theoretical rigor for key components of ASKD.
Xiaoling Zhou, Yiyu Liu, Shikun Zhang, Wei Ye 0004
AAAI4
2026 High-resolution underwater camouflaged object detection: GBU-UCOD dataset and topology-aware and frequency-decoupled networks
Wenji Wu, Shuo Ye, Yiyu Liu, Jiguang He, Zitong Yu
Pattern Recognit. Lett.3
2025 Spectral Co-Clustering Based Wireless Network Decomposition for Resource Scheduling
abstract
Large-scale wireless networks pose significant challenges in resource scheduling, where the solution space grows exponentially with network size. While network decomposition offers a promising solution by breaking networks into manageable subnetworks, existing approaches, including spectral clustering, fail to effectively capture the complex service relationships between base stations (BSs) and users, particularly in networks with massive user populations. This paper presents BSCCD (Bidirectional Spectral Co-Clustering Based Decomposition), a new decomposition scheme that addresses these challenges through two key innovations: (i) a two-round spectral co-clustering framework that captures bidirectional BS-user relationships, and (ii) a user node merging strategy that handles massive user populations. Extensive experiments on real-world datasets from multiple Chinese cities demonstrate that BSCCD reduces computation latency by up to 61.91 % compared to global optimization, while achieving more than 10 % improvement in solution quality over traditional clustering approaches. The advantage is particularly pronounced in medium-scale networks, where BSCCD outperforms traditional methods by$\mathbf{2 8. 7 6 \%}$. Our results demonstrate BSCCD's practical viability for resource scheduling in contemporary wireless networks, especially in scenarios with complex BS-user interactions and large user populations.
Yiyu Liu, Yilin Xiao 0001, Ming Tang 0006, Lin Gao 0001, Jianwei Huang 0001
WiOpt1
2024 FaaSMem: Improving Memory Efficiency of Serverless Computing with Memory Pool Architecture
abstract
In serverless computing, an idle container is not recycled directly, in order to mitigate time-consuming cold container startup. These idle containers still occupy the memory, exasperating the memory shortage of today's data centers. By offloading their cold memory to remote memory pool could potentially resolve this problem. However, existing offloading policies either hurt the Quality of Service (QoS) or are too coarse-grained in serverless computing scenarios.
Chuhao Xu, Yiyu Liu, Zijun Li 0001, Quan Chen 0002, Han Zhao 0005, Deze Zeng, Xueqi Wu, Senbo Fu, Minyi Guo
ASPLOS (3)2
2024 Multi-Level Spatial-Temporal Feature Aggregation and Alignment-Based Selective Residual Dense Propagation Module for HDR Video Reconstruction
abstract
To reconstruct high dynamic range (HDR) video from alternating exposed low dynamic range (LDR) frames, the key is to address the misalignment and imprecise fusion caused by information loss and noise in ill-exposed regions. Following a coarse-to-fine manner, a Multi-level Spatial-Temporal feature aggregation and alignment-based Selective Residual Dense Propagation Network (MSTSRDPNet) is proposed. The Multi-level Spatial-Temporal aggregation extracts spatial-temporal features and aggregates them to mitigate information loss for fusion. The alignment-based Selective Residual Dense Propagation module reconstructs the aligned feature by using channel attention to redistribute feature weights while leveraging residual dense connections for information propagation. Experiments show that the proposed MSTSRDPNet outperforms all conventional methods on the synthetic dataset with PSNR-T, HDR-VQM, and HDR-VDP-2 scores of 44.64 dB, 86.83, and 73.9.
Yiyu Liu, Fengshan Zhao, Qin Liu 0002, Takeshi Ikenaga
ICASSP1
2021 Concept-Aware Denoising Graph Neural Network for Micro-Video Recommendation
abstract
Recently, micro-video sharing platforms such as Kuaishou and Tiktok have become a major source of information for people's lives. Thanks to the large traffic volume, short video lifespan and streaming fashion of these services, it has become more and more pressing to improve the existing recommender systems to accommodate these challenges in a cost-effective way. In this paper, we propose a novel concept-aware denoising graph neural network (named Conde) for micro-video recommendation. Conde consists of a three-phase graph convolution process to derive user and micro-video representations: warm-up propagation, graph denoising and preference refinement. A heterogeneous tripartite graph is constructed by connecting user nodes with video nodes, and video nodes with associated concept nodes, extracted from captions and comments of the videos. To address the noisy information in the graph, we introduce a user-oriented graph denoising phase to extract a subgraph which can better reflect the user's preference. Despite the main focus of micro-video recommendation in this paper, we also show that our method can be generalized to other types of tasks. Therefore, we also conduct empirical studies on a well-known public E-commerce dataset. The experimental results suggest that the proposed Conde achieves significantly better recommendation performance than the existing state-of-the-art solutions.
Yiyu Liu, Yu Tian 0008, Changping Wang, Yanan Niu, Yang Song 0008, Chenliang Li 0005
CIKM1
2020 MATINF: A Jointly Labeled Large-Scale Dataset for Classification, Question Answering and Summarization
abstract
Recently, large-scale datasets have vastly facilitated the development in nearly all domains of Natural Language Processing.However, there is currently no cross-task dataset in NLP, which hinders the development of multi-task learning.We propose MATINF, the first jointly labeled large-scale dataset for classification, question answering and summarization.MAT-INF contains 1.07 million question-answer pairs with human-labeled categories and usergenerated question descriptions.Based on such rich information, MATINF is applicable for three major NLP tasks, including classification, question answering, and summarization.We benchmark existing methods and a novel multi-task baseline over MATINF to inspire further research.Our comprehensive comparison and experiments over MATINF and other datasets demonstrate the merits held by MAT-INF. 1
Canwen Xu, Jiaxin Pei, Yiyu Liu
ACL4