VLDB 2026 Research / reviewers in the wild / expert
Chen Ying
dblp:25/4657
· DBLP profile ↗
12ranked-venue papers
8as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Pegasus: Efficient Asynchronous Three-Layer Federated LearningabstractCompared to conventional two-layer federated learning (FL), three-layer FL, which adds a layer of edge servers between the central server and clients, is less studied but could provide better training performance in terms of reducing elapsed wall-clock time to complete training. However, three-layer FL inherits and magnifies some inherent challenges in two-layer FL, such as client heterogeneity. With different computing capabilities and network connections, slow clients require exceedingly long training times. To alleviate the negative effect due to slow clients, this paper is the first to study asynchronous three-layer FL, where edge servers conduct local aggregation without waiting for slow clients and the central server aggregates updates from fast edge servers. Unfortunately, asynchronous three-layer FL could suffer from performance degradation as when aggregating updates from slow clients and edge servers, those updates are not computed based on the newest global model. Therefore, we propose a staleness-aware framework, Pegasus, with newly designed client selection, compression, and update aggregation mechanisms to improve every important aspect during training. Our extensive evaluation of different training tasks demonstrates that Pegasus can achieve a reduction in elapsed wall-clock training time by at least 46.8% with an increase in converged accuracy of the trained global model by up to 0.3% as compared to the state-of-the-art. Chen Ying, Baochun Li |
ICCCN | 1 |
| 2025 | Symphony: Collective Coordination in Multi-Tenant GPU ClustersabstractMulti-tenant GPU clusters are designed to concurrently run multiple distributed ML training workloads. However, frequent data transfers among GPUs via collective operations can slow down training, as collectives from different tenants compete for network bandwidth. Recent work (e.g., CASSINI) has considered collective coordination to prevent network contention, but primarily focused on static job-level optimizations at deployment time, oblivious to runtime network conditions and the specific traffic pattern of each workload. In this paper, we present Symphony, an application-layer solution that dynamically coordinates collective operations across tenants at runtime. Symphony integrates seamlessly with existing clusters with minimal modifications to the collective communication library and includes a lightweight online scheduling mechanism that requires no advance information about the workloads or their collectives. We evaluate Symphony using both a real GPU testbed implementation and trace-driven simulations. Specifically, using realistic ML workloads in our testbed, we observe improvements of up to 13.2% in average communication time and 9.6% in training time compared to state-of-the-art solutions. Manaf Bin-Yahya, Amir Shani, Hossein Shafieirad, S. Hossein Mortazavi, Chen Ying, Aaron Wang, Majid Ghaderi |
ICNP | 5 |
| 2025 | An approach to microalgae identification based on joint optimization model of slicing and detection
Chen Ying, Si Yu, Chen Ting |
Expert Syst. Appl. | 2 |
| 2024 | Blade: Pushing the Performance Envelope of Asynchronous Federated LearningabstractAsynchronous federated learning (FL) has been proposed to decrease the training time in conventional FL where the communication paradigm is synchronous. Instead of aggregating after receiving updates from all the selected clients, an asynchronous FL server conducts aggregation without waiting for slow clients. Though superior to synchronous FL, the performance of existing works in asynchronous FL — measured by the wall-clock time of global training — leaves much to be desired, as the staleness of client updates may degrade the performance substantially. In this paper, we propose Blade, a new stalenessaware framework that seeks to push the performance envelope of asynchronous FL by designing new mechanisms in all important design aspects of FL training, including client selection, adaptive pruning, quantization, and update aggregation. Blade selects clients based on their staleness and the quality of their previous updates. Before reporting to the server, every client prunes its update with a pruning amount related to its staleness and quantizes the pruned update. When aggregating updates, Blade tunes the aggregation weight of each update according to its staleness and divergence from the previous global model. In an extensive array of performance evaluations with six benchmark datasets, Blade consistently showed its substantial performance superiority over its state-of-the-art competitors. It decreased the wall-clock training time by up to 64.6%. Chen Ying, Baochun Li, Bo Li 0001 |
IWQoS | 1 |
| 2022 | AoTI Minimization for Multi-Type Data Sampling in Industrial Wireless Sensor NetworksabstractFor practical industrial wireless sensor networks (IWSNs), the system freshness of a specific task is usually related to multiple and multitype sensing data. However, most existing research on freshness metrics, such as Age of Information (AoI) or Age of Processing (AoP), only considers a single-package setting with a single type of data. To fill this gap, we propose the Age of Task-oriented Information (AoTI) for measuring the freshness of industrial tasks in IWSNs. It measures the time elapsed of the latest analyzed results before arriving at the receiver since the generation of any type of sampling data belonging to one certain task. Furthermore, we aim to minimize the long-term AoTI for IWSNs applications by jointly optimizing access modes and sampling frequencies for all sensors. By first formulating the problem as a Mixed Integer Nonlinear Program-ming problem, we then transform it to a constrained Markov Decision Process (CMDP) and relax it as an un-constrained MDP using Lagrangian method. Finally, we develop a Learning-based Access mode selection and Sampling frequency Control (LASC) algorithm and verify its superiority through simulations. Chen Ying, Zhen Zhao 0001, Changyan Yi, You Shi, Ran Wang 0004 |
EUC | 1 |
| 2022 | Tempo: Improving Training Performance in Cross-Silo Federated LearningabstractDifferent from its commonly studied scenario to centrally store clients’ data in institutions, which implicitly neglects clients’ data privacy, we study cross-silo federated learning in a preferable setting to keep private data on clients, and train the global model with a three-layer structure, where the institutions aggregate model updates from their clients for several rounds before sending their aggregated updates to the central server. In this context, we mathematically prove that the number of clients’ local training epochs affects the global model performance and thus propose a new approach, Tempo, to adaptively tune the epoch number of each client through training. The results of our evaluation conducted under real network environments show that Tempo can not only improve training performance in terms of global model accuracy and communication efficiency, but also the elapsed training time. Chen Ying, Baochun Li, Bo Li 0001 |
ICASSP | 1 |
| 2022 | Raven: Scheduling Virtual Machine Migration During Datacenter Upgrades with Reinforcement Learning
Chen Ying, Baochun Li, Xiaodi Ke |
Mob. Networks Appl. | 1 |
| 2019 | Scheduling Virtual Machine Migration During Datacenter Upgrades with Reinforcement Learning
Chen Ying, Baochun Li, Xiaodi Ke |
QSHINE | 1 |
| 2017 | A Prior-Free Spectrum Auction for Approximate Revenue MaximizationabstractDynamic spectrum allocation has been proven as a promising solution to the spectrum scarcity problem. Auctions represent a natural allocation mechanism that generates a monetary remuneration for primary users. We study approximate revenue-maximizing spectrum auctions in a prior-free setting, when information on user valuations on channels is unavailable. A two-phase auction framework is presented. In Phase 1, a strategyproof mechanism computes a subset of users with an interference-free spectrum allocation, such that the potential revenue to be gained in the second phase is maximized. A carefully tailored payment scheme ensures truthful bidding at this stage. The selected users advance into Phase 2, where eventual auction winners are computed through a recursive random partitioning and revenue extraction procedure. While no strategyproof auction can achieve absolute optimal revenue in the prior-free setting, our random partition auction is both truthful in expectation and achieves the best known ratio 13 of the optimal revenue. Chen Ying, Hao Huang 0001, Ajay Gopinathan, Zongpeng Li |
Comput. J. | 1 |
| 2016 | Modeling for Noisy Labels of Crowd Workers
Qian Yan 0001, Hao Huang 0001, Yunjun Gao, Chen Ying, Qingyang Hu, Tieyun Qian, Qinming He |
APWeb (2) | 4 |
| 2016 | Image quality assessment based on the visual perception of image contentsabstractThis paper describes an image quality assessment (IQA) metric based on the visual perception of image contents (VPIC)). In the metric, VPIC is firstly modelled by simulating the nonlinearity of luminance perception, masking properties and contrast sensitivity characteristics of human visual system (HVS). Then the source and distorted images are processed by this model respectively, and their intensity differences are calculated. Finally, based on the intensity differences, an IQA model is built. And 47 reference images and 1549 distorted images in the LIVE, TID2008 and CSIQ databases are tested with the IQA metric. The results show that it is helpful to improve the consistency between the objective IQA scores and the subjective mean opinion scores (MOSs) combining the visual perception and complexity of image contents. Juncai Yao, Guizhong Liu, Chen Ying |
VCIP | 3 |
| 2005 | Shear-resize factorizations for fast image registrationabstractOwing to its effectiveness and simplicity, intensity-based method works well for registration of images. However, it needs a large amount of computation for geometric transformation. In this paper, we present two shear-resize matrix factorizations to accelerate the transformation. A transform matrix can be factorized into two shears and a fixed non-uniform resize, or three shears and a customizable resize. A customizable resize can be uniform in all dimensions or scaling just in one dimension. Shears can be implemented very fast by memory-shift, and a resize can be done by simple axis-aligned interpolation. The factorizations can be applied to both rigid-body and affine transformations. Their efficiency is performed by experiments on some standard test images and fingerprint images. The methods are quite promising for hardware implementation, and can also be extended to 3D or higher dimensional fast geometric transformation. Chen Ying, Pengwei Hao, Chao Zhang 0001 |
ICIP (3) | 1 |