VLDB 2026 Research / reviewers in the wild / expert
Yi Ding 0006
dblp:89/5503-6
· DBLP profile ↗
17ranked-venue papers
7as first author
9since 2021 · last 2026
0000-0003-2757-9182ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 5 first-author · 2 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Personalization Matters: Benchmarking ML, DL and LLMs for Mental Health Forecasting from Smartphone Sensing DataabstractSmartphone sensing offers a scalable, unobtrusive window into daily behavioral patterns linked to mental health. While prior work focuses on detection — identifying current conditions — forecasting future mental health states enables proactive, Just-in-Time Adaptive Interventions before crises emerge. We present the first benchmarking study comparing traditional machine learning (ML), deep learning (DL), and large language models (LLMs) for mental health forecasting using the College Experience Sensing dataset — the most extensive longitudinal passive sensing dataset on college student mental health to date. Our evaluation spans modeling paradigm, LLM adaptation strategy, and personalization, yielding three key findings: DL models, particularly Transformer, achieve the strongest overall performance, while LLMs exhibit a reasoning-prediction gap; few-shot consistently outperforms zero-shot and supervised fine-tuning (SFT); and personalization substantially improves forecasting. These findings lay groundwork for next-generation, human-centered adaptive systems supporting early mental health intervention. Our code is available at https://github.com/KaiDF/Study_MentalHealth_Forecasting. Kaidong Feng, Zhu Sun 0001, Roy Ka-Wei Lee, Yin Leng Theng, Yi Ding 0006 |
UMAP | 6 |
| 2026 | Trustworthy and Adaptive LLMs for Mental and Physical Wellbeing in RecommendationsabstractLarge Language Models (LLMs) are rapidly transforming recommender systems (RSs) and user modeling by enabling richer representations of users, context, and intent. In wellbeing-oriented applications, such as activity, diet, stress, and mental health support, the use of LLM-based RSs introduces both new opportunities and critical challenges related to trustworthiness, adaptivity, explainability, privacy, and human-centered evaluation. This workshop aims to bring together researchers and industry practitioners from user modeling, personalization, RSs and healthcare to examine how LLM-based models can be responsibly designed, evaluated, and deployed for mental and physical wellbeing, fostering AI for social good. The workshop focuses on adaptive LLM-based user modeling, trustworthy recommendation mechanisms, LLM-powered user simulation, and evaluation methodologies that capture longitudinal and affective user outcomes. Strongly aligned with the ACM UMAP community, the workshop emphasizes adaptive and personalized systems that place human values at the center. Through paper presentations, invited talks, and interactive discussions, the workshop will surface open challenges, emerging best practices, and future research directions for LLM-driven wellbeing recommendation. Zhu Sun 0001, Yi Ding 0006, Yin Leng Theng, W. Quin Yow, Roy Ka-Wei Lee |
UMAP | 2 |
| 2025 | Unveiling Environmental Impacts of Large Language Model Serving: A Functional Unit ViewabstractLarge language models (LLMs) offer powerful capabilities but come with significant environmental impact, particularly in carbon emissions.Existing studies benchmark carbon emissions but lack a standardized basis for comparison across different model configurations.To address this, we introduce the concept of functional unit (FU) as a standardized basis and develop FUEL, the first FU-based framework for evaluating LLM serving's environmental impact.Through three case studies, we uncover key insights and trade-offs in reducing carbon emissions by optimizing model size, quantization strategy, and hardware choice, paving the way for more sustainable LLM serving.The code is available at https://github.com/jojacola/FUEL. Yanran Wu, Inez Hua, Yi Ding 0006 |
ACL (1) | 3 |
| 2023 | CAFQA: A Classical Simulation Bootstrap for Variational Quantum AlgorithmsabstractClassical computing plays a critical role in the advancement of quantum frontiers in the NISQ era. In this spirit, this work uses classical simulation to bootstrap Variational Quantum Algorithms (VQAs). VQAs rely upon the iterative optimization of a parameterized unitary circuit (ansatz) with respect to an objective function. Since quantum machines are noisy and expensive resources, it is imperative to classically choose the VQA ansatz initial parameters to be as close to optimal as possible to improve VQA accuracy and accelerate their convergence on today’s devices. Gokul Subramanian Ravi, Pranav Gokhale, Yi Ding 0006, William M. Kirby, Kaitlin N. Smith, Jonathan M. Baker, Peter J. Love, Henry Hoffmann, Kenneth R. Brown, Fred Chong |
ASPLOS (1) | 3 |
| 2023 | AntTune: An Efficient Distributed Hyperparameter Optimization System for Large-Scale Data
Jun Zhou 0011, Qitao Shi, Yi Ding 0006, Lin Wang 0098, Feng Zhu 0011 |
DASFAA (4) | 3 |
| 2023 | A Rule-based Decision System for Financial ApplicationsabstractDecision rules have been widely applied in industrial applications such as finance, medicine, and biology, due to the critical requirement of interpretability. In order to make decision rules easier and more widely used in financial scenarios, an automatic intelligent rule system with rule learning and rule management capabilities is needed. However, the rule system for financial applications has distinctive challenges both in algorithms and systems. From the algorithm perspective, due to the characteristics of the financial data and scenarios, the rule learning algorithm faces the class-imbalanced issue, the scalability issue, and the diversity of optimization objectives. From the system perspective, a flexible rule learning and management framework is needed to adapt to fast-changing financial applications with heterogenous data, and engineering optimization is required to ensure the time and space efficiency of rule learning. In this work, we focus on developing a Rule-based Decision System (RDS) to deal with the algorithmic and systematic challenges mentioned above. RDS covers the full life cycle of the decision rules, including the rule learning module, rule management module, and rule deployment module. Moreover, the rule system offers an interactive interface to allow users to integrate the expert experiences into the decision rules and realize the human-in-the-loop. The RDS has been deployed on one of the world’s largest trading and money transfer platforms, serving hundreds of millions of users and transactions. Meng Li 0068, Jun Zhou 0011, Lu Yu 0006, Xiaoguang Huang, Yongfeng Gu, Yi Ding 0006 |
ICDE | 6 |
| 2023 | DistriBayes: A Distributed Platform for Learning, Inference and Attribution on Large Scale Bayesian NetworkabstractTo improve the marketing performance in the financial scenario, it is necessary to develop a trustworthy model to analyze and select promotion-sensitive customers. Bayesian Network (BN) is suitable for this task because of its interpretability and flexibility, but it usually suffers the exponentially growing computation complexity as the number of nodes grows. To tackle this problem, we present a comprehensive distributed platform named DistriBayes, which can efficiently learn, infer and attribute on a large-scale BN all-in-one platform. It implements several score-based structure learning methods, loopy belief propagation with backdoor adjustment for inference, and a carefully optimized search procedure for attribution. Leveraging the distributed cluster, DistriBayes can finish the learning and attribution on Bayesian Network with hundreds of nodes and millions of samples in hours. Yi Ding 0006, Jun Zhou 0011, Qing Cui, Lin Wang 0098, Mengqi Zhang 0003 |
WSDM | 1 |
| 2023 | Turaco: Complexity-Guided Data Sampling for Training Neural Surrogates of ProgramsabstractProgrammers and researchers are increasingly developing surrogates of programs, models of a subset of the observable behavior of a given program, to solve a variety of software development challenges. Programmers train surrogates from measurements of the behavior of a program on a dataset of input examples. A key challenge of surrogate construction is determining what training data to use to train a surrogate of a given program. We present a methodology for sampling datasets to train neural-network-based surrogates of programs. We first characterize the proportion of data to sample from each region of a program's input space (corresponding to different execution paths of the program) based on the complexity of learning a surrogate of the corresponding execution path. We next provide a program analysis to determine the complexity of different paths in a program. We evaluate these results on a range of real-world programs, demonstrating that complexity-guided sampling results in empirical improvements in accuracy. Alex Renda, Yi Ding 0006, Michael Carbin |
Proc. ACM Program. Lang. | 2 |
| 2021 | Generalizable and interpretable learning for configuration extrapolationabstractModern software applications are increasingly configurable, which puts a burden on users to tune these configurations for their target hardware and workloads. To help users, machine learning techniques can model the complex relationships between software configuration parameters and performance. While powerful, these learners have two major drawbacks: (1) they rarely incorporate prior knowledge and (2) they produce outputs that are not interpretable by users. These limitations make it difficult to (1) leverage information a user has already collected (e.g., tuning for new hardware using the best configurations from old hardware) and (2) gain insights into the learner’s behavior (e.g., understanding why the learner chose different configurations on different hardware or for different workloads). To address these issues, this paper presents two configuration optimization tools, GIL and GIL+, using the proposed generalizable and interpretable learning approaches. To incorporate prior knowledge, the proposed tools (1) start from known configurations, (2) iteratively construct a new linear model, (3) extrapolate better performance configurations from that model, and (4) repeat. Since the base learners are linear models, these tools are inherently interpretable. We enhance this property with a graphical representation of how they arrived at the highest performance configuration. We evaluate GIL and GIL+ by using them to configure Apache Spark workloads on different hardware platforms and find that, compared to prior work, GIL and GIL+ produce comparable, and sometimes even better performance configurations, but with interpretable results. Yi Ding 0006, Ahsan Pervaiz, Michael Carbin, Henry Hoffmann |
ESEC/SIGSOFT FSE | 1 |
| 2020 | Dynamical Systems Theory for Causal Inference with Application to Synthetic Control MethodsabstractIn this paper, we adopt results in nonlinear time series analysis for causal inference in dynamical settings. Our motivation is policy analysis with panel data, particularly through the use of “synthetic control" methods. These methods regress pre-intervention outcomes of the treated unit to outcomes from a pool of control units, and then use the fitted regression model to estimate causal effects post-intervention. In this setting, we propose to screen out control units that have a weak dynamical relationship to the treated unit. In simulations, we show that this method can mitigate bias from “cherry-picking" of control units, which is usually an important concern. We illustrate on real-world applications, including the tobacco legislation example of \citet{Abadie2010}, and Brexit. Yi Ding 0006, Panos Toulis |
AISTATS | 1 |
| 2020 | A polynomial-time algorithm for learning nonparametric causal graphsabstractWe establish finite-sample guarantees for a polynomial-time algorithm for learning a nonlinear, nonparametric directed acyclic graphical (DAG) model from data. The analysis is model-free and does not assume linearity, additivity, independent noise, or faithfulness. Instead, we impose a condition on the residual variances that is closely related to previous work on linear models with equal variances. Compared to an optimal algorithm with oracle knowledge of the variable ordering, the additional cost of the algorithm is linear in the dimension $d$ and the number of samples $n$. Finally, we compare the proposed algorithm to existing approaches in a simulation study. Yi Ding 0006, Bryon Aragam |
NeurIPS | 2 |
| 2019 | Generative and multi-phase learning for computer systems optimizationabstractMachine learning and artificial intelligence are invaluable for computer systems optimization: as computer systems expose more resources for management, ML/AI is necessary for modeling these resources' complex interactions. The standard way to incorporate ML/AI into a computer system is to first train a learner to accurately predict the system's behavior as a function of resource usage---e.g., to predict energy efficiency as a function of core usage---and then deploy the learned model as part of a system---e.g., a scheduler. In this paper, we show that (1) continued improvement of learning accuracy may not improve the systems result, but (2) incorporating knowledge of the systems problem into the learning process improves the systems results even though it may not improve overall accuracy. Specifically, we learn application performance and power as a function of resource usage with the systems goal of meeting latency constraints with minimal energy. We propose a novel generative model which improves learning accuracy given scarce data, and we propose a multi-phase sampling technique, which incorporates knowledge of the systems problem. Our results are both positive and negative. The generative model improves accuracy, even for state-of-the-art learning systems, but negatively impacts energy. Multi-phase sampling reduces energy consumption compared to the state-of-the-art, but does not improve accuracy. These results imply that learning for systems optimization may have reached a point of diminishing returns where accuracy improvements have little effect on the systems outcome. Thus we advocate that future work on learning for systems should de-emphasize accuracy and instead incorporate the system problem's structure into the learner. Yi Ding 0006, Nikita Mishra, Henry Hoffmann |
ISCA | 1 |
| 2017 | Large Scale Kernel Methods for Online AUC MaximizationabstractLearning to optimize AUC performance for classifying label imbalanced data in online scenarios has been extensively studied in recent years. Most of the existing work has attempted to address the problem directly in the original feature space, which may not suitable for non-linearly separable datasets. To solve this issue, some kernel-based learning methods are proposed for non-linearly separable datasets. However, such kernel approaches have been shown to be inefficient and failed to scale well on large scale datasets in practice. Taking this cue, in this work, we explore the use of scalable kernel-based learning techniques as surrogates to existing approaches: random Fourier features and Nyström method, for tackling the problem and bring insights to the differences between the two methods based on their online performance. In contrast to the conventional kernel-based learning methods which suffer from high computational complexity of the kernel matrix, our proposed approaches elevate this issue with linear features that approximate the kernel function/matrix. Specifically, two different surrogate kernel-based learning models are presented for addressing the online AUC maximization task: (i) the Fourier Online AUC Maximization (FOAM) algorithm that samples the basis functions from a data-independent distribution to approximate the kernel functions; and (ii) the Nyström Online AUC Maximization (NOAM) algorithm that samples a subset of instances from the training data to approximate the kernel matrix by a low rank matrix. Another novelty of the present work is the proposed mini-batch Online Gradient Descent method for model updating to control the noise and reduce the variance of gradients. We provide theoretical analyses for the two proposed algorithms. Empirical studies on commonly used large scale datasets show that the proposed algorithms outperformed existing state-of-the-art methods in terms of both AUC performance and computational efficiency. Yi Ding 0006, Peilin Zhao, Steven C. H. Hoi |
ICDM | 1 |
| 2017 | KunPeng: Parameter Server based Distributed Learning Systems and Its Applications in Alibaba and Ant FinancialabstractIn recent years, due to the emergence of Big Data (terabytes or petabytes) and Big Model (tens of billions of parameters), there has been an ever-increasing need of parallelizing machine learning (ML) algorithms in both academia and industry. Although there are some existing distributed computing systems, such as Hadoop and Spark, for parallelizing ML algorithms, they only provide synchronous and coarse-grained operators (e.g., Map, Reduce, and Join, etc.), which may hinder developers from implementing more efficient algorithms. This motivated us to design a universal distributed platform termed KunPeng, that combines both distributed systems and parallel optimization algorithms to deal with the complexities that arise from large-scale ML. Specifically, KunPeng not only encapsulates the characteristics of data/model parallelism, load balancing, model sync-up, sparse representation, industrial fault-tolerance, etc., but also provides easy-to-use interface to empower users to focus on the core ML logics. Empirical results on terabytes of real datasets with billions of samples and features demonstrate that, such a design brings compelling performance improvements on ML programs ranging from Follow-the-Regularized-Leader Proximal algorithm to Sparse Logistic Regression and Multiple Additive Regression Trees. Furthermore, KunPeng's encouraging performance is also shown for several real-world applications including the Alibaba's Double 11 Online Shopping Festival and Ant Financial's transaction risk estimation. Jun Zhou 0011, Xiaolong Li 0005, Peilin Zhao, Chaochao Chen 0001, Xinxing Yang, Qing Cui, Xu Chen 0017, Yi Ding 0006, Yuan Qi 0001 |
KDD | 10 |
| 2017 | Multiresolution Kernel Approximation for Gaussian Process RegressionabstractGaussian process regression generally does not scale to beyond a few thousands data points without applying some sort of kernel approximation method. Most approximations focus on the high eigenvalue part of the spectrum of the kernel matrix, $K$, which leads to bad performance when the length scale of the kernel is small. In this paper we introduce Multiresolution Kernel Approximation (MKA), the first true broad bandwidth kernel approximation algorithm. Important points about MKA are that it is memory efficient, and it is a direct method, which means that it also makes it easy to approximate $K^{-1}$ and $\mathop{\textrm{det}}(K)$. Yi Ding 0006, Risi Kondor, Jonathan Eskreis-Winkler |
NIPS | 1 |
| 2015 | An Adaptive Gradient Method for Online AUC MaximizationabstractLearning for maximizing AUC performance is an important research problem in machine learning. Unlike traditional batch learning methods for maximizing AUC which often suffer from poor scalability, recent years have witnessed some emerging studies that attempt to maximize AUC by single-pass online learning approaches. Despite their encouraging results reported, the existing online AUC maximization algorithms often adopt simple stochastic gradient descent approaches, which fail to exploit the geometry knowledge of the data observed in the online learning process, and thus could suffer from relatively slow convergence. To overcome the limitation of the existing studies, in this paper, we propose a novel algorithm of Adaptive Online AUC Maximization (AdaOAM), by applying an adaptive gradient method for exploiting the knowledge of historical gradients to perform more informative online learning. The new adaptive updating strategy by AdaOAM is less sensitive to parameter settings due to its natural effect of tuning the learning rate. In addition, the time complexity of the new algorithm remains the same as the previous non-adaptive algorithms. To demonstrate the effectiveness of the proposed algorithm, we analyze its theoretical bound, and further evaluate its empirical performance on both public benchmark datasets and anomaly detection datasets. The encouraging empirical results clearly show the effectiveness and efficiency of the proposed algorithm. Yi Ding 0006, Peilin Zhao, Steven C. H. Hoi, Yew-Soon Ong |
AAAI | 1 |
| 2014 | Learning Relative Similarity by Stochastic Dual Coordinate AscentabstractLearning relative similarity from pairwise instances is an important problem in machine learning and has a wide range of applications. Despite being studied for years, some existing methods solved by Stochastic Gradient Descent (SGD) techniques generally suffer from slow convergence. In this paper, we investigate the application of Stochastic Dual Coordinate Ascent (SDCA) technique to tackle the optimization task of relative similarity learning by extending from vector to matrix parameters. Theoretically, we prove the optimal linear convergence rate for the proposed SDCA algorithm, beating the well-known sublinear convergence rate by the previous best metric learning algorithms. Empirically, we conduct extensive experiments on both standard and large-scale data sets to validate the effectiveness of the proposed algorithm for retrieval tasks. Yi Ding 0006, Peilin Zhao, Chunyan Miao, Steven C. H. Hoi |
AAAI | 2 |