VLDB 2026 Research / reviewers in the wild / expert
Lingyun Yang
dblp:16/1292
· DBLP profile ↗
23ranked-venue papers
9as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorComputer networks · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FlashPS: Efficient Generative Image Editing with Mask-aware Caching and SchedulingabstractGenerative image editing using diffusion models has become a prevalent application in today's AI cloud services. In production environments, image editing typically involves a mask that specifies the regions of an image template to be edited. The use of mask provides direct control over the editing process and introduces sparsity in the model inference. In this paper, we present FlashPS, a system that efficiently serves image editing requests. The key insight behind FlashPS is that image editing only modifies the masked regions of image templates, while preserving the original content in the unmasked areas. Driven by this insight, FlashPS judiciously skips redundant computations associated with the unmask areas by reusing cached intermediate activations from previous inferences. To mitigate the high cache loading overhead, FlashPS employs a bubble-free pipeline scheme that overlaps computation with cache loading. Additionally, to reduce queuing latency in online serving while improving the GPU utilization, FlashPS proposes a novel continuous batching strategy for diffusion model serving, allowing newly arrived requests to join the running batch in just one step of denoising computation, without waiting for the entire batch to complete. As heterogenous masks induce imbalanced load, FlashPS also develops a load balancing strategy that takes into account the loads of both computation and cache loading. Collectively, FlashPS outperforms state-of-the-art diffusion serving systems for image editing, achieving up to 3× higher throughput and reducing average request latency by up to 14.7× while ensuring image quality. Xiaoxiao Jiang, Suyi Li 0002, Lingyun Yang, Tianyu Feng, Zhipeng Di, Weiyi Lu, Guoxuan Zhu, Xiu Lin, Yinghao Yu, Tao Lan, Lin Qu, Liping Zhang 0013, Wei Wang 0030 |
EuroSys | 3 |
| 2026 | FractalGPU: Fair and Elastic GPU Sharing for General-Purpose Computing
Kaicheng Guo, Lingyun Yang, Wenda Tang, Pengwei Du, Qian Da, Zhengwei Qi |
ICDCS | 3 |
| 2026 | Close-Loop Controlled RC Parameter Damping-Based Multiplier for Wind Farm Equivalent Modeling
Jinpeng Lei, Lingyun Yang |
ISCAS | 2 |
| 2026 | CL-WTAL: Weakly-Supervised Temporal Complex Action Localization Based on Multi-Scale Contrast LearningabstractTemporal action localization in long-term untrimmed videos remains a critical yet challenging task in video understanding, with existing methods often relying on anchor-based or fully-supervised frameworks that incur heavy computation and require labor-intensive frame-level annotations. This paper presents a novel weakly-supervised approach, CL-WTAL, which leverages multi-scale contrast learning and graph convolution for accurate action localization and recognition. The method comprises three key components: (i) A multi-scale sliding window mechanism (long/normal/short sequences) to segment sub-actions from complex videos, adapting to diverse action durations; (ii) A spatio-temporal graph convolution network (ST-RGCN) to extract skeletal feature vectors, integrating human motion dynamics and environmental context; (iii) A contrastive learning-based similarity evaluation framework that combines cosine similarity and Dynamic Time Warping (DTW) distance to measure feature vector relationships, enabling precise action boundary detection without extensive fine-tuning. Experiments on daily-life video datasets demonstrate that CL-WTAL effectively localizes action intervals and classifies actions with high accuracy, outperforming state-of-the-art weakly-supervised methods. Weili Ding, Lingyun Yang, Shuo Hu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | GPU-Disaggregated Serving for Deep Learning Recommendation Models at Scale
Lingyun Yang, Yongchen Wang, Yinghao Yu, Qizhen Weng 0001, Jianbo Dong, Chi Zhang 0005, Yanyi Zi, Zechao Zhang, Menglei Zheng, Lanlan Xi, Binzhang Fu, Tao Lan, Liping Zhang 0013, Lin Qu, Wei Wang 0030 |
NSDI | 1 |
| 2025 | Katz: Efficient Workflow Serving for Diffusion Models with Many Adapters
Suyi Li 0002, Lingyun Yang, Xiaoxiao Jiang, Hanfeng Lu, Dakai An, Zhipeng Di, Weiyi Lu, Yinghao Yu, Tao Lan, Lin Qu, Liping Zhang 0013, Wei Wang 0030 |
USENIX ATC | 2 |
| 2023 | Beware of Fragmentation: Scheduling GPU-Sharing Workloads with Fragmentation Gradient Descent
Qizhen Weng 0001, Lingyun Yang, Yinghao Yu, Wei Wang 0030, Xiaochuan Tang, Liping Zhang 0013 |
USENIX ATC | 2 |
| 2022 | Workload consolidation in alibaba clusters: the good, the bad, and the uglyabstractWeb companies typically run latency-critical long-running services and resource-intensive, throughput-hungry batch jobs in a shared cluster for improved utilization and reduced cost. Despite many recent studies on workload consolidation, the production practice remains largely unknown. This paper describes our efforts to efficiently consolidate the two types of workloads in Alibaba clusters to support the company's e-commerce businesses. Yongkang Zhang 0003, Yinghao Yu, Wei Wang 0030, Qiukai Chen, Tianchen Ding, Qizhen Weng 0001, Lingyun Yang, Jian He 0004, Liping Zhang 0013 |
SoCC | 10 |
| 2022 | Elastic Full-Waveform Inversion With Unconverted-Wave Adjoint PropagatorsabstractElastic full-waveform inversion (EFWI) can restore high-resolution model parameters by minimizing the misfit function between the modeled and observed data. However, the coupling propagation of$P$- and$S$-waves will cause the crosstalk among elastic parameters and increase the nonlinearity of EFWI. The decoupled elastic wave equation can help EFWI to weaken the crosstalk effect, but it increases the computational cost of EFWI. In addition, the decomposition of the$S$-wave stress will produce artifacts. Hence, we have developed an EFWI approach with unconverted-wave adjoint propagators to recover the high-resolution model parameters. In the new EFWI, we use the unconverted-wave equation to construct the adjoint propagators without$S$-wave stress decomposition, which can reduce the artifacts. Since the unconverted-wave equation omits the cross term in the elastic wave equation, the computational cost of EFWI is reduced. Numerical examples have demonstrated that our EFWI can efficiently produce high-resolution models and reduce the computational cost of EFWI by about 30%. Guochen Wu, Zhanyuan Liang, Lingyun Yang |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2021 | Morphling: Fast, Near-Optimal Auto-Configuration for Cloud-Native Model ServingabstractMachine learning models are widely deployed in production cloud to provide online inference services. Efficiently deploying inference services requires careful tuning of hardware and runtime configurations (e.g., GPU type, GPU memory, batch size), which can significantly improve the model serving performance and reduce cost. However, existing autoconfiguration approaches for general workloads, such as Bayesian optimization and white-box prediction, are inefficient in navigating the high-dimensional configuration space of model serving, incurring high sampling cost. Lingyun Yang, Yinghao Yu, Wei Wang 0030, Bo Li 0001, Xianchao Sun, Jian He 0004, Liping Zhang 0013 |
SoCC | 2 |
| 2019 | The study of new features for video traffic classification
Lingyun Yang, Zaijian Wang |
Multim. Tools Appl. | 1 |
| 2018 | New Cross-Domain QoE Guarantee Method Based on Isomorphism Flow
Zaijian Wang, Xinheng Wang 0001, Lingyun Yang, Pingping Tang |
CollaborateCom | 4 |
| 2017 | On the largest Cartesian closed category of stable domains
Xiaoyong Xi, Qingyu He, Lingyun Yang |
Theor. Comput. Sci. | 3 |
| 2013 | Roughness in quantales
Lingyun Yang, Luoshan Xu |
Inf. Sci. | 1 |
| 2011 | Topological properties of generalized approximation spaces
Lingyun Yang, Luoshan Xu |
Inf. Sci. | 1 |
| 2009 | Algebraic aspects of generalized approximation spaces
Lingyun Yang, Luoshan Xu |
Int. J. Approx. Reason. | 1 |
| 2007 | Anomaly detection and diagnosis in grid environmentsabstractIdentifying and diagnosing anomalies in application behavior is critical to delivering reliable application-level performance. In this paper we introduce a strategy to detect anomalies and diagnose the possible reasons behind them. Our approach extends the traditional window-based strategy by using signal-processing techniques to filter out recurring, background fluctuations in resource behavior. In addition, we have developed a diagnosis technique that uses standard monitoring data to determine which related changes in behavior may cause anomalies. We evaluate our anomaly detection and diagnosis technique by applying it in three contexts when we insert anomalies into the system at random intervals. The experimental results show that our strategy detects up to 96% of anomalies while reducing the false positive rate by up to 90% compared to the traditional window average strategy. In addition, our strategy can diagnose the reason for the anomaly approximately 75% of the time. Lingyun Yang, Chuang Liu 0006, Jennifer M. Schopf, Ian T. Foster |
SC | 1 |
| 2006 | Statistical Data Reduction for Efficient Application Performance MonitoringabstractThere is a growing need for systems that can monitor and analyze application performance data automatically in order to deliver reliable and sustained performance to applications. However, the continuously growing complexity of high performance computer systems and applications makes this process difficult. We introduce a statistical data reduction method that can be used to guide the selection of system metrics that are both necessary and sufficient to describe observed application behavior, thus reducing the instrumentation perturbation and data volume to be managed. To evaluate our strategy, we applied it to one CPU-bound grid application using cluster machines and GridFTP data transfer in a wide area testbed. A comparative study shows that our strategy produces better results than other techniques. It can reduce the number of system metrics to be managed by about 80%, while still capturing enough information for performance predictions. Lingyun Yang, Jennifer M. Schopf, Catalin Dumitrescu, Ian T. Foster |
CCGRID | 1 |
| 2005 | Online resource matching for heterogeneous grid environmentsabstractIn this paper, we first present a linear programming based approach for modeling and solving the resource matching problem in grid environments with heterogeneous resources. The resource matching problem described takes into account resource sharing, job priorities, dependencies on multiple resource types, and resource specific policies. We then propose Web service style architecture for online matching of independent jobs with resources in a grid environment and describe a prototype implementation. Our preliminary performance results indicate that the linear programming based approach for resource matching is efficient in speed and accuracy and can keep up with high job arrival rates of an important criterion for online resource matching systems. Also, the Web service style architecture makes the system scalable and extendable. It can also be integrated with other existing grid services in a straightforward manner. Vijay K. Naik, Chuang Liu 0006, Lingyun Yang, Jonathan Wagner |
CCGRID | 3 |
| 2005 | Improving parallel data transfer times using predicted variances in shared networksabstractIt is increasingly common to use multiple distributed storage systems as a single data store within which large datasets may be replicated. Thus, we face the problem of how to access replicated data efficiently. Multiple-source parallel transfers can reduce access times by transferring data from several replicas in parallel. However, we then face the problem of deciding which data to fetch from which replicas. We propose a Tuned Conservative scheduling technique that uses predicted means and variances for network performance to make data selection decisions. This stochastic scheduling technique adjusts the amount of data fetched on a link according to not only the link performance but the expected variance in that performance. We incorporate our technique into the striped GridFTP server from the Globus Toolkit, and demonstrate that the technique can produce data transfer times that are significantly faster and less variable than those of other techniques. Lingyun Yang, Jennifer M. Schopf, Ian T. Foster |
CCGRID | 1 |
| 2005 | Efficient Relational Joins with Arithmetic Constraints on Multiple AttributesabstractWe introduce and study a new class of queries that we refer to as ACMA (arithmetic constraints on multiple attributes) queries. Such combinatorial queries require the simultaneous satisfaction of arithmetic constraints on three or more attributes from different relations, and thus often involve expensive multi-join operations. Building on techniques from constraint programming, we develop preprocessing methods, algorithms, and a new constrained join operator that allow ACMA queries to be evaluated efficiently within a conventional relational database engine. We present the results of a careful performance evaluation of both our new approach and the conventional nested-loop join algorithm. Measurements of tuples read, intermediate tuples generated, and execution time shows that our approach achieves superior performance for ACMA joins. Chuang Liu 0006, Lingyun Yang, Ian T. Foster |
IDEAS | 2 |
| 2003 | Conservative Scheduling: Using Predicted Variance to Improve Scheduling Decisions in Dynamic EnvironmentsabstractIn heterogeneous and dynamic environments, efficient execution of parallel computations can require mappings of tasks to processors whose performance is both irregular (because of heterogeneity) and time-varying (because of dynamicity). While adaptive domain decomposition techniques have been used to address heterogeneous resource capabilities, temporal variations in those capabilities have seldom been considered. We propose a conservative scheduling policy that uses information about expected future variance in resource capabilities to produce more efficient data mapping decisions. We first present techniques, based on time series predictors that we developed in previous work, for predicting CPU load at some future time point, average CPU load for some future time interval, and variation of CPU load over some future time interval. We then present a family of stochastic scheduling algorithms that exploit such predictions of future availability and variability when making data mapping decisions. Finally, we describe experiments in which we apply our techniques to an astrophysics application. The results of these experiments demonstrate that conservative scheduling can produce execution times that are both significantly faster and less variable than other techniques. Lingyun Yang, Jennifer M. Schopf, Ian T. Foster |
SC | 1 |
| 2002 | Design and Evaluation of a Resource Selection Framework for Grid ApplicationsabstractWhile distributed, heterogeneous collections of computers ("Grids") can in principle be used as a computing platform, in practice the problems of first discovering and then organizing resources to meet application requirements are difficult. We present a general-purpose resource selection framework that addresses these problems by defining a resource selection service for locating Grid resources that match application requirements. At the heart of this framework is a simple, but powerful, declarative language based on a technique called set matching, which extends the Condor matchmaking framework to support both single-resource and multiple-resource selection. This framework also provides an open interface for loading application-specific mapping modules to personalize the resource selector. We present results obtained when this framework is applied in the context of a computational astrophysics application, Cactus. These results demonstrate the effectiveness of our technique. Chuang Liu 0006, Lingyun Yang, Ian T. Foster, Dave Angulo |
HPDC | 2 |