Yihui Feng

dblp:192/8535 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
10since 2021 · last 2026
0009-0003-7853-9508ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2Computer networks · 1
YearPublicationVenuePosition
2026 Heterogeneous multi-objective contrastive learning for stock prediction
Metoh Adler Loua, Xiaocao Ouyang, Xiangkun Wang, Yihui Feng, Xin Yang 0012
Expert Syst. Appl.5
2025 Multi-granularity Knowledge Transfer for Continual Reinforcement Learning
abstract
Continual reinforcement learning (CRL) empowers RL agents with the ability to learn a sequence of tasks, accumulating knowledge learned in the past and using the knowledge for problemsolving or future task learning. However, existing methods often focus on transferring fine-grained knowledge across similar tasks, which neglects the multi-granularity structure of human cognitive control, resulting in insufficient knowledge transfer across diverse tasks. To enhance coarse-grained knowledge transfer, we propose a novel framework called MT-Core (as shorthand for Multi-granularity knowledge Transfer for Continual reinforcement learning). MT-Core has a key characteristic of multi-granularity policy learning: 1) a coarsegrained policy formulation for utilizing the powerful reasoning ability of the large language model (LLM) to set goals, and 2) a fine-grained policy learning through RL which is oriented by the goals. We also construct a new policy library (knowledge base) to store policies that can be retrieved for multi-granularity knowledge transfer. Experimental results demonstrate the superiority of the proposed MT-Core in handling diverse CRL tasks versus popular baselines.
Chaofan Pan, Lingfei Ren, Yihui Feng, Linbo Xiong, Wei Wei 0018, Yonghao Li, Xin Yang 0012
IJCAI3
2025 Agamotto: Scheduling of Deadline-Oriented Incremental Query Execution under Uncertain Resource Price
abstract
Incremental query processing is widely used in data warehouses and streaming systems. While many optimization techniques are developed to generate incremental query plans, the scheduling support for incremental processing remains preliminary. Typically, execution is triggered with fixed frequencies specified by the user. In this paper, we propose a novel scheduling problem for incremental query execution under a deadline, assuming the resource has a fluctuating and unforeseen price. We propose two naive solutions as well as a prophet scheduler that foresees the future. We present an end-to-end system Agamotto that models future probabilities offline with a Markov Decision Process (MDP) and makes cost-based and dynamic scheduling decisions online. We show how Agamotto can be extended to handle a workflow of dependent queries, so that they can all incrementally execute in an asynchronous fashion. Experiments show that Agamotto consistently outperforms the naive solutions, and the achieved cost is on average 10x closer to the theoretical lower bound provided by the prophet scheduler.
Botong Huang, Lianggui Weng, Wei Chen 0133, Zuozhi Wang, Kai Zeng 0002, Chen Li 0001, Yihui Feng, Bolin Ding, Jingren Zhou 0001
Proc. VLDB Endow.7
2025 Ten Challenging Problems in Federated Foundation Models
abstract
Federated Foundation Models (FedFMs) represent a distributed learning paradigm that fuses general competences of foundation models as well as privacy-preserving capabilities of federated learning. This combination allows the large foundation models and the small local domain models at the remote clients to learn from each other in a teacher-student learning setting. This paper provides a comprehensive summary of the ten challenging problems inherent in FedFMs, encompassing foundational theory, utilization of private data, continual learning, unlearning, Non-IID and graph data, bidirectional knowledge transfer, incentive mechanism design, game mechanism design, model watermarking, and efficiency. The ten challenging problems manifest in five pivotal aspects: “Foundational Theory,” which aims to establish a coherent and unifying theoretical framework for FedFMs. “Data,” addressing the difficulties in leveraging domain-specific knowledge from private data while maintaining privacy; “Heterogeneity,” examining variations in data, model, and computational resources across clients; “Security and Privacy,” focusing on defenses against malicious attacks and model theft; and “Efficiency,” highlighting the need for improvements in training, communication, and parameter efficiency. For each problem, we offer a clear mathematical definition on the objective function, analyze existing methods, and discuss the key challenges and potential solutions. This in-depth exploration aims to advance the theoretical foundations of FedFMs, guide practical implementations, and inspire future research to overcome these obstacles, thereby enabling the robust, efficient, and privacy-preserving FedFMs in various real-world applications.
Tao Fan 0002, Hanlin Gu, Xuemei Cao 0001, Chee Seng Chan, Qian Chen 0023, Yiqiang Chen 0001, Yihui Feng, Yang Gu 0001, Jiaxiang Geng, Bing Luo 0002, Shuoling Liu, WinKent Ong, Chao Ren 0006, Jiaqi Shao, Xiaoli Tang 0001, Hong Xi Tae, Yongxin Tong, Shuyue Wei 0001, Fan Wu 0006, Wei Xi 0003, Mingcong Xu, Xin Yang 0012, Jiangpeng Yan, Hao Yu 0023, Han Yu 0001, Xiaojin Zhang 0002, Zhenzhe Zheng 0001, Lixin Fan, Qiang Yang 0001
IEEE Trans. Knowl. Data Eng.7
2024 Overcoming Spatial-Temporal Catastrophic Forgetting for Federated Class-Incremental Learning
abstract
This paper delves into federated class-incremental learning (FCiL), where new classes appear continually or even privately to local clients. However, existing FCiL methods suffer from the problem of spatial-temporal catastrophic forgetting, i.e., forgetting the previously learned knowledge over time and the client-specific information owned by different clients. Additionally, private class and knowledge heterogeneity amongst local clients further exacerbate spatial-temporal forgetting, making FCiL challenging to apply. To address these issues, we propose Federated Class-specific Binary Classifier (FedCBC), an innovative approach to transferring and fusing knowledge across both temporal and spatial perspectives. FedCBC consists of two novel components: (1) continual personalization that distills previous knowledge from a global model to multiple local models, and (2) selective knowledge fusion that enhances knowledge integration of the same class from divergent clients and shares private knowledge with other clients. Extensive experiments using three newly-formulated metrics (termed GA, KRS, and KRT) demonstrate the effectiveness of the proposed approach. Our code is now hosted at: https://github.com/SkyOfBeginning/FedCBC.
Hao Yu 0023, Xin Yang 0012, Xin Gao 0038, Yihui Feng, Hao Wang 0068, Yan Kang 0001, Tianrui Li 0001
ACM Multimedia4
2024 PPS: Fair and efficient black-box scheduling for multi-tenant GPU clusters
Kaihao Ma, Zhenkun Cai, Xiao Yan 0002, Zhi Liu 0002, Yihui Feng, Chao Li 0009, Wei Lin 0022, James Cheng
Parallel Comput.6
2022 Content-Oriented Learned Image Compression
Meng Li 0050, Shangyin Gao, Yihui Feng, Yibo Shi, Jing Wang 0194
ECCV (19)3
2022 Fine-Grained Modeling and Optimization for Intelligent Resource Management in Big Data Processing
abstract
Big data processing at the production scale presents a highly complex environment for resource optimization (RO), a problem crucial for meeting performance goals and budgetary constraints of analytical users. The RO problem is challenging because it involves a set of decisions (the partition count, placement of parallel instances on machines, and resource allocation to each instance), requires multi-objective optimization (MOO), and is compounded by the scale and complexity of big data systems while having to meet stringent time constraints for scheduling. This paper presents a MaxCompute based integrated system to support multi-objective resource optimization via fine-grained instance-level modeling and optimization. We propose a new architecture that breaks RO into a series of simpler problems, new fine-grained predictive models, and novel optimization methods that exploit these models to make effective instance-level RO decisions well under a second. Evaluation using production workloads shows that our new RO system could reduce 37--72% latency and 43--78% cost at the same time, compared to the current optimizer and scheduler, while running in 0.02-0.23s.
Chenghao Lyu, Yanlei Diao, Wei Chen 0133, Yihui Feng, Yaliang Li, Kai Zeng 0002, Jingren Zhou 0001
Proc. VLDB Endow.8
2021 Asymmetric Gained Deep Image Compression With Continuous Rate Adaptation
abstract
With the development of deep learning techniques, the combination of deep learning with image compression has drawn lots of attention. Recently, learned image compression methods had exceeded their classical counterparts in terms of rate-distortion performance. However, continuous rate adaptation remains an open question. Some learned image compression methods use multiple networks for multiple rates, while others use one single model at the expense of computational complexity increase and performance degradation. In this paper, we propose a continuously rate adjustable learned image compression framework, Asymmetric Gained Variational Autoencoder (AG-VAE). AG-VAE utilizes a pair of gain units to achieve discrete rate adaptation in one single model with a negligible additional computation. Then, by using exponential interpolation, continuous rate adaptation is achieved without compromising performance. Besides, we propose the asymmetric Gaussian entropy model for more accurate entropy estimation. Exhaustive experiments show that our method achieves comparable quantitative performance with SOTA learned image compression methods and better qualitative performance than classical image codecs. In the ablation study, we confirm the usefulness and superiority of gain units and the asymmetric Gaussian entropy model.
Ze Cui, Jing Wang 0194, Shangyin Gao, Tiansheng Guo, Yihui Feng
CVPR5
2021 Scaling Large Production Clusters with Partitioned Synchronization
Yihui Feng, Zhi Liu 0002, Yunjian Zhao, Tatiana Jin, Yidi Wu 0001, James Cheng, Chao Li 0009
USENIX ATC1
2020 AntMan: Dynamic Scaling on GPU Clusters for Deep Learning
Wencong Xiao, Shiru Ren, Yong Li 0045, Yang Zhang 0102, Pengyang Hou, Yihui Feng, Wei Lin 0016, Yangqing Jia
OSDI7
2020 Color-Guided Depth Image Recovery With Adaptive Data Fidelity and Transferred Graph Laplacian Regularization
abstract
Depth images play an important role and are prevalently used in many computer vision and computational imaging tasks. However, due to the limitation of active sensing technology, the captured depth images in practice usually suffer from low resolution and noise, which prevents its further applications. To remedy this problem, in this paper, we first propose an adaptive data fidelity formulation to optimally generate each depth pixel from a mixture probability distribution, characterizing the similarity both in the depth map and the corresponding high-resolution guided color image. The proposed method is able to fit the distribution of the input depth signal as an optimization problem by maximizing the mixture probability. Furthermore, to promote the piecewise property that depth images exhibit, we propose a transferred graph Laplacian model as a regularization term, which is general and able to handle various depth recovery tasks such as super-resolution and denoising well. Specifically, each pixel within the recovered depth image is represented as a vertex in a graph with weights in connected edges representing the similarity between vertices. By minimizing the squared variations of the image signal, the task of depth image recovery can be converted to the problem of graph-based image filtering. Since the proposed graph Laplacian regularization model is able to fully exploit a priori information about the depth image, a much more accurate and robust estimation of the underlying depth can be obtained. Extensive experiment evaluations verify that the proposed method obtains recovered depth with higher quality in terms of both objective and subjective criteria, compared with most of the state-of-the-art methods.
Yongbing Zhang 0002, Yihui Feng, Xianming Liu 0005, Deming Zhai, Xiangyang Ji, Haoqian Wang, Qionghai Dai
IEEE Trans. Circuits Syst. Video Technol.2
2019 Who limits the resource efficiency of my datacenter: an analysis of Alibaba datacenter traces
abstract
Cloud platform provides great flexibility and cost-efficiency for end-users and cloud operators. However, low resource utilization in modern datacenters brings huge wastes of hardware resources and infrastructure investment. To improve resource utilization, a straightforward way is co-locating different workloads on the same hardware. To figure out the resource efficiency and understand the key characteristics of workloads in co-located cluster, we analyze an 8-day trace from Alibaba's production trace. We reveal three key findings as follows. First, memory becomes the new bottleneck and limits the resource efficiency in Alibaba's datacenter. Second, in order to protect latency-critical applications, batch-processing applications are treated as second-class citizens and restricted to utilize limited resources. Third, more than 90% of latency-critical applications are written in Java applications. Massive self-contained JVMs further complicate resource management and limit the resource efficiency in datacenters.
Zihao Chang, Sa Wang, Haiyang Ding, Yihui Feng, Yungang Bao
IWQoS5
2019 Yugong: Geo-Distributed Data and Job Placement at Scale
abstract
Companies like Alibaba operate tens of data centers (DCs) across geographically distributed locations. These DCs collectively provide the storage space and computing power for the company, storing EBs of data and serving millions of batch analytics jobs every day. In Alibaba, as our businesses grow, there are more and more cross-DC dependencies caused by jobs reading data from remote DCs. Consequently, the precious wide area network bandwidth becomes a major bottleneck for operating geo-distributed DCs at scale. In this paper, we present Yugong --- a system that manages data placement and job placement in Alibaba's geo-distributed DCs, with the objective to minimize cross-DC bandwidth usage. Yugong uses three methods, namely project placement, table replication, and job outsourcing, to address the issues of high bandwidth consumption across the DCs. We give the details of Yugong's design and implementation for the three methods, and describe how it cooperates with other systems (e.g., Alibaba's big data analytics platform and cluster scheduler) to improve the productivity of the DCs. We also report comprehensive performance evaluation results, which validate the design of Yugong and show that significant reduction in cross-DC bandwidth usage has been achieved.
Yingjie Shi, Yihui Feng, James Cheng, Haochuan Fan, Chao Li 0009, Jingren Zhou 0001
Proc. VLDB Endow.4
2017 Single depth image super-resolution and denoising based on sparse graphs via structure tensor
abstract
The existing single depth image super-resolution (SR) methods suppose that the image to be interpolated is noise free. However, the supposition is invalid in practice because noise will be inevitably introduced in the depth image acquisition process. In this paper, we address the problem of image denoising and SR jointly based on designing sparse graphs that are useful for describing the geometric structures of data domains. In our method, we first cluster similar patches in a noisy depth image and compute an average patch. Different from the majority of the graph Fourier transform (GFT) that assumed an underlying 4-connected graph structure with vertical and horizontal edges only, we select more general sparse graph structures and edges weights based on the difference of the blocks' structure tensors. For the average patch, a graph template with edges orthogonal to the principal gradient is designed. Finally, the graph based transform (GBT) dictionary is learned from the derived correlation graph for signal representation. As shown in our experimental results, the proposed method obtains a lot of improvement in performance.
Yihui Feng, Xianming Liu 0005, Yongbing Zhang 0002, Qionghai Dai
ICIP1
2017 Reliable Computing Service in Massive-Scale Systems through Rapid Low-Cost Failover
abstract
Large-scale distributed systems deployed as Cloud datacenters are capable of provisioning service to consumers with diverse business requirements. Providers face pressure to provision uninterrupted reliable services while reducing operational costs due to significant software and hardware failures. A widely adopted means to achieve such a goal is using redundant system components to implement user-transparent failover, yet its effectiveness must be balanced carefully without incurring heavy overhead when deployed—an important practical consideration for complex large-scale systems. Failover techniques developed for Cloud systems often suffer serious limitations, including mandatory restart leading to poor cost-effectiveness, as well as solely focusing on crash failures, omitting other important types, such as timing failures and simultaneous failures. This paper addresses these limitations by presenting a new approach to user-transparent failover for massive-scale systems. The approach uses soft-state inference to achieve rapid failure recovery and avoid unnecessary restart, with minimal system resource overhead. It also copes with different failures, including correlated and simultaneous events. The proposed approach was implemented, deployed and evaluated within Fuxi system, the underlying resource management system used within Alibaba Cloud. Results demonstrate that our approach tolerates complex failure scenarios while incurring at worst 228.5 microsecond instance overhead with 1.71 percent additional CPU usage.
Renyu Yang, Peter Garraghan, Yihui Feng, Jin Ouyang, Jie Xu 0007, Zhuo Zhang 0015
IEEE Trans. Serv. Comput.4
2016 Single image super-resolution via projective dictionary learning with anchored neighborhood regression
abstract
We propose a novel single image super-resolution (SR) algorithm based on the projective dictionary pair learning with anchored neighborhood regression. Different from previous dictionary learning methods that aim to learn only a synthesis or an analysis dictionary, our method would learn both types of dictionaries jointly for regression to achieve image SR. We first cluster the training features into K clusters in order to learn synthesis and analysis dictionaries. Moreover, we learn the regressions with the training samples at training phase and use them on reconstruction stage. As shown in our experimental results, the proposed method obtains high-quality SR results quantitatively and visually against state-of-the-art methods.
Yihui Feng, Yongbing Zhang 0002, Yulun Zhang 0001, Qionghai Dai
VCIP1