EDBT 2026 Demo / reviewers in the wild / expert
Yankai Cao
dblp:155/9335
· DBLP profile ↗
18ranked-venue papers
2as first author
17since 2021 · last 2027
0000-0001-9014-2552ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 10 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | A moving-horizon differential evolution algorithm for training deep classification trees on large datasetsabstractDecision trees are widely used in interpretable expert systems because their if-then rules can be inspected by domain experts. However, training accurate decision trees on large continuous-feature datasets remains difficult. Greedy methods such as CART scale well but optimize each split locally, while exact optimal-tree methods can produce high-quality trees only for relatively small datasets or shallow depths. Existing evolutionary approaches are more flexible, but full-tree search becomes increasingly difficult as tree depth grows. This creates a need for an optimization method that improves decision-tree accuracy for deep trees and large datasets. This paper presents the Moving-Horizon Differential Evolution algorithm for Optimal Classification Trees ( MH-DEOCT ). The proposed method combines three components: a moving-horizon decomposition that replaces full-tree optimization with a sequence of tractable subtree problems, a discrete tree decoding strategy that removes redundant threshold searches, and a GPU-accelerated fitness evaluation exploiting Single Instruction Multiple Data parallelism. Extensive experiments on 65 benchmark datasets show that MH-DEOCT trains depth-8 trees on datasets with up to 60 million samples, improves average test accuracy over CART by 2.73%, and exceeds DL8.5 by 2.27% on medium and large datasets. On a 60-million-sample wearable stress detection dataset, it improves test accuracy over CART by 5.64% at depth 8. Overall, these results show that MH-DEOCT improves predictive accuracy for deep decision trees on large datasets while retaining the transparency of the final tree. Jiayang Ren, Valentín Osuna-Enciso, Yankai Cao |
Expert Syst. Appl. | 3 |
| 2026 | Class-Missing Semi-supervised document key information extraction via synergistic refinement estimationabstractCurrent methods for document key information extraction (DKIE) rely heavily on labeled data with high annotation costs. To mitigate this issue, the semi-supervised learning (SSL) paradigm, which utilizes unlabeled document samples, has gained broad attention in DKIE. However, existing SSL methods require labeled and unlabeled data to share an identical label space, which is impractical in many DKIE tasks (i.e., some unlabeled samples do not belong to any known classes in the labeled set). In this paper, we formulate this problem as Class-Missing Semi-supervised (CMSS) DKIE. In DKIE, unknown classes usually belong to minority and fine-grained categories, intensifying the misconnections between known and unknown classes and making CMSS more challenging. To address this issue, we propose Synergistic Refinement Estimation (SRE), a progressive prototype estimation scheme that alleviates the unknown classes bias to the majority known classes on long-tailed unlabeled data. Furthermore, dynamic threshold hash rectification and structural calibration mechanisms are proposed to correct connections between fine-grained classes. Extensive experimental results demonstrate that SRE surpasses existing state-of-the-art methods on several DKIE benchmarks. Code is available at https://github.com/anonymoulink/SRE_DKIE . Yonghong Song, Boyu Wang 0004, Yankai Cao, Jiayang Ren, Chaojie Ji, Qi Zhang 0096, Qiangqiang Mao |
Inf. Process. Manag. | 4 |
| 2026 | Learning directed acyclic graphs via noising and denoisingabstractLearning directed acyclic graphs (DAGs) from observational data that involve a set of variables carrying intrinsic noise is a crucial yet challenging task. Recent approaches frame the DAG learning task as minimizing a reconstruction-based objective function, i.e., reconstructing observed data by learning a DAG, while adhering to an acyclic constraint. However, optimizing this objective does not always guarantee the correctness of the learned graphs. One reason for this is that the intrinsic noise entangled with the variables is inadvertently absorbed in the reconstruction process at the expense of inferring incorrect DAG structures. To address this issue, we propose a novel DAG learner that first injects artificial noise into observational variables that are contaminated by fixed intrinsic noise. The next step involves reconstructing these perturbed variables using a weighted structure estimator and a weighted noise estimator, instead of reconstructing the observational variables solely with the fixed intrinsic noise. This strategy effectively reduces the sensitivity of the structure estimator to the fixed intrinsic noise. Additionally, we observe a strong similarity between the proposed DAG learner and diffusion models. This similarity motivates us to replicate the well-known denoising capabilities of diffusion models in our DAG learner. We reformulate and adapt the denoising process in denoising diffusion probabilistic models (DDPMs), which allows us to derive a specific weight schedule for the weighted structure and noise estimators of our DAG learner. Extensive experiments conducted on synthetic and real datasets with varying scales demonstrate the outstanding performance of our proposed method. Chaojie Ji, Jialin Nan, Ruxin Wang 0001, Yankai Cao |
Inf. Sci. | 6 |
| 2026 | Causality-inspired latent feature augmentation for single domain generalization
Chaojie Ji, Yankai Cao, Ye Li 0002, Wei Zhao 0001, Ruxin Wang 0001 |
Pattern Recognit. | 3 |
| 2025 | Change Detection for Wide-Field Video Images in Foggy Weather Based on Enhanced K-Means Clustering
Yankai Cao, Xiaoqian Qu, Sensen Song, Zhenhong Jia |
ICIC (1) | 1 |
| 2025 | RGPest-YOLO: A YOLOv8 Pest Detection Method Based on Image Preprocessing
Xiaoqian Qu, Yankai Cao, Zhenhong Jia |
ICIC (1) | 2 |
| 2025 | Differentiable Decision Tree via "ReLU+Argmin" ReformulationabstractDecision tree, despite its unmatched interpretability and lightweight structure, faces two key issues that limit its broader applicability: non-differentiability and low testing accuracy.
This study addresses these issues by developing a differentiable oblique tree that optimizes the entire tree using gradient-based optimization. We propose an exact reformulation of hard-split trees based on "ReLU+Argmin" mechanism, and then cast the reformulated tree training as an unconstrained optimization task.
The ReLU-based sample branching, expressed as exact-zero or non-zero values, preserve a unique decision path, in contrast to soft decision trees with probabilistic routing. The subsequent Argmin operation identifies the unique zero-violation path, enabling deterministic predictions.
For effective gradient flow, we approximate Argmin behaviors by scaling softmin function. To ameliorate numerical instability, we propose a warm-start annealing scheme that solves multiple optimization tasks with increasingly accurate approximations.
This reformulation alongside distributed GPU parallelism offers strong scalability, supporting 12-depth tree even on million-scale datasets where most baselines fail.
Extensive experiments demonstrate that our optimized tree achieves a superior testing accuracy against 14 baselines, including an average improvement of 7.54\% over CART. Qiangqiang Mao, Jiayang Ren, Yixiu Wang, Chenxuanyin Zou, Yankai Cao |
NeurIPS | 6 |
| 2025 | AdaMSS: Adaptive Multi-Subspace Approach for Parameter-Efficient Fine-TuningabstractIn this paper, we propose AdaMSS, an adaptive multi-subspace approach for parameter-efficient fine-tuning of large models. Unlike traditional parameter-efficient fine-tuning methods that operate within a large single subspace of the network weights, AdaMSS leverages subspace segmentation to obtain multiple smaller subspaces and adaptively reduces the number of trainable parameters during training, ultimately updating only those associated with a small subset of subspaces most relevant to the target downstream task. By using the lowest-rank representation, AdaMSS achieves more compact expressiveness and finer tuning of the model parameters. Theoretical analyses demonstrate that AdaMSS has better generalization guarantee than LoRA, PiSSA, and other single-subspace low-rank-based methods. Extensive experiments across image classification, natural language understanding, and natural language generation tasks show that AdaMSS achieves comparable performance to full fine-tuning and outperforms other parameter-efficient fine-tuning methods in most cases, all while requiring fewer trainable parameters. Notably, on the ViT-Large model, AdaMSS achieves 4.7\% higher average accuracy than LoRA across seven tasks, using just 15.4\% of the trainable parameters. On RoBERTa-Large, AdaMSS outperforms PiSSA by 7\% in average accuracy across six tasks while reducing the number of trainable parameters by approximately 94.4\%. These results demonstrate the effectiveness of AdaMSS in parameter-efficient fine-tuning. The code for AdaMSS is available at https://github.com/jzheng20/AdaMSS. Wanglong Lu, Yiming Dong, Chaojie Ji, Yankai Cao, Zhouchen Lin |
NeurIPS | 5 |
| 2025 | Deep Learning-Based Approximation of Model Predictive Control Laws Using Mixture NetworksabstractIn recent years, researchers have proposed the approximation of model predictive control (MPC) using deep neural networks (DNNs). However, a limitation arises as DNNs inherently offer one-to-one mappings, posing a challenge when multiple optimal control inputs correspond to each system state, thereby leading to one-to-many mappings. Therefore, we propose an alternative scheme using mixture networks (MNs) with components of probability (density) distributions in the output layer. This method uses conditional probabilities offered by combining several estimated probability distributions, enabling generating multiple control inputs with the highest probabilities. Notably, this approach is applicable to several problems by choosing a suitable probability distribution, such as using a Gaussian distribution for nonlinear problems and a Bernoulli distribution for mixed-integer linear programming (MILP) problems. We investigate two case studies illustrating that the mixture network-based approximation outperforms the DNN-based approximation.Note to Practitioners—This article proposes using mixture (density) networks, a type of machine learning technique, to enable the online implementation of model predictive control. The prohibitive computation cost associated with model predictive control poses a significant challenge when implementing it in complex nonlinear and MILP problems. While approximation methods of control laws using deep neural networks have been studied to address this issue, it is unsuitable for problems where each state has multiple optimal control inputs. In contrast, our proposed approach can accurately approximate the control laws for these problems. This notable feature can facilitate the online implementation of model predictive control in a wide range of nonlinear and MILP problems. Morimasa Okamoto, Jiayang Ren, Qiangqiang Mao, Yankai Cao |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2025 | An Edge-Based Adaptive Event-Triggered Network Transmission Scheme for Fully Distributed Power and Frequency Control of Islanded AC MicrogridsabstractIn islanded AC microgrids, achieving optimal power distribution across distributed generators (DGs) and restoring frequency post-primary control under limited communication resources is paramount for ensuring stability and flexibility. In this paper, we develop a distributed adaptive secondary control strategy tailored for islanded AC microgrids. This strategy facilitates effective active power sharing and frequency regulation using solely local network information. A unique adaptive edge-based event-triggered transmission mechanism is also proposed, obviating Zeno behavior, to reduce communication overhead. It requires only the degree information of a single node in the DG network topology. Contrasting most existing research that necessitates global network topology details, like the Laplacian matrix, this work establishes a fully distributed control performance criterion that operates devoid of global data. The effectiveness of the proposed control approach is validated using a modified IEEE 34-bus test system. Sheng Han 0002, Hong Zhu 0001, Qishui Zhong, Kaibo Shi, Yankai Cao |
IEEE Trans. Netw. Serv. Manag. | 5 |
| 2024 | Handling the Non-smooth Challenge in Tensor SVD: A Multi-objective Tensor Recovery Framework
Wanglong Lu, Wenzhe Wang, Yankai Cao, Xiaoqin Zhang 0002, Xianta Jiang |
ECCV (14) | 4 |
| 2024 | Stable predictive control of continuous stirred-tank reactors using deep learning
Shulei Zhang, Runda Jia, Yankai Cao, Dakuo He, Feng Yu 0014 |
Inf. Sci. | 3 |
| 2022 | Global Optimization of K-Center Clusteringabstract$k$-center problem is a well-known clustering method and can be formulated as a mixed-integer nonlinear programming problem. This work provides a practical global optimization algorithm for this task based on a reduced-space spatial branch and bound scheme. This algorithm can guarantee convergence to the global optimum by only branching on the centers of clusters, which is independent of the dataset’s cardinality. In addition, a set of feasibility-based bounds tightening techniques are proposed to narrow down the domain of centers and significantly accelerate the convergence. To demonstrate the capacity of this algorithm, we present computational results on 32 datasets. Notably, for the dataset with 14 million samples and 3 features, the serial implementation of the algorithm can converge to an optimality gap of 0.1% within 2 hours. Compared with a heuristic method, the global optimum obtained by our algorithm can reduce the objective function on average by 30.4%. Mingfei Shi, Kaixun Hua, Jiayang Ren, Yankai Cao |
ICML | 4 |
| 2022 | A Scalable Deterministic Global Optimization Algorithm for Training Optimal Decision TreeabstractThe training of optimal decision tree via mixed-integer programming (MIP) has attracted much attention in recent literature. However, for large datasets, state-of-the-art approaches struggle to solve the optimal decision tree training problems to a provable global optimal solution within a reasonable time. In this paper, we reformulate the optimal decision tree training problem as a two-stage optimization problem and propose a tailored reduced-space branch and bound algorithm to train optimal decision tree for the classification tasks with continuous features. We present several structure-exploiting lower and upper bounding methods. The computation of bounds can be decomposed into the solution of many small-scale subproblems and can be naturally parallelized. With these bounding methods, we prove that our algorithm can converge by branching only on variables representing the optimal decision tree structure, which is invariant to the size of datasets. Moreover, we propose a novel sample reduction method that can predetermine the cost of part of samples at each BB node. Combining the sample reduction method with the parallelized bounding strategies, our algorithm can be extremely scalable. Our algorithm can find global optimal solutions on dataset with over 245,000 samples (1000 cores, less than 1% optimality gap, within 2 hours). We test 21 real-world datasets from UCI Repository. The results reveal that for datasets with over 7,000 samples, our algorithm can, on average, improve the training accuracy by 3.6% and testing accuracy by 2.8%, compared to the current state-of-the-art. Kaixun Hua, Jiayang Ren, Yankai Cao |
NeurIPS | 3 |
| 2022 | Global Optimal K-Medoids Clustering of One Million SamplesabstractWe study the deterministic global optimization of the K-Medoids clustering problem. This work proposes a branch and bound (BB) scheme, in which a tailored Lagrangian relaxation method proposed in the 1970s is used to provide a lower bound at each BB node. The lower bounding method already guarantees the maximum gap at the root node. A closed-form solution to the lower bound can be derived analytically without explicitly solving any optimization problems, and its computation can be easily parallelized. Moreover, with this lower bounding method, finite convergence to the global optimal solution can be guaranteed by branching only on the regions of medoids. We also present several tailored bound tightening techniques to reduce the search space and computational cost. Extensive computational studies on 28 machine learning datasets demonstrate that our algorithm can provide a provable global optimal solution with an optimality gap of 0.1\% within 4 hours on datasets with up to one million samples. Besides, our algorithm can obtain better or equal objective values than the heuristic method. A theoretical proof of global convergence for our algorithm is also presented. Jiayang Ren, Kaixun Hua, Yankai Cao |
NeurIPS | 3 |
| 2022 | An efficient improved African vultures optimization algorithm with dimension learning hunting for traveling salesman and large-scale optimization applicationsabstractExploring the finest shortest-path traveling salesman optimization application is a typical NP-hard problem. Similarly the solution of the large-scale optimization applications is also a big challenging issue in front of scientists. First, African Vultures Optimization Algorithm (AVOA) was developed to resolve continuous applications where it performed fine. In the last few months, many enhanced strategies of AVOA have been offered in recent literature works and it has been extensively utilized to resolve large-scale engineering optimization applications. This study offers a newly modified dimension learning hunting (DLH)-based AVOA called DLHAV algorithm to resolve highly complex continuous and discrete applications. It helps improve the imbalance amid the hunting (or exploitation) and search (or exploration), the lack of crowd diversity, slow convergence speed, trapping in local optima, and early convergence of the AVOA variant. The proposed strategy benefits from a newly driven approach called the DLH search approach congenital from the separate exploitation behavior of vultures in the search domain. DLH exploration strategy utilizes a distinct method to make the best neighborhood for all vultures in which the nearest member information can be supplied amid vultures. DLH helps in improving the balance amid global and local and sustains diversity. To scrutinize the performance of DLHAV, the solutions of the DLHAV method are verified on 29-CEC'17 and 10-CEC'20 with familiar comparative methods and some other classical optimization approaches over many familiar traveling salesman problem/large-scale instances. With the intention of attaining unbiased and rigorous comparison, descriptive statistics such as standard deviation and mean have been applied, and the statistical Friedman test is also conducted. The experimental solution carried out in this study has revealed that the proposed algorithm outperforms significantly over the other alternative optimizers. Narinder Singh, Essam H. Houssein, Seyedali Mirjalili, Yankai Cao, Ganeshsree Selvachandran |
Int. J. Intell. Syst. | 4 |
| 2021 | A Scalable Deterministic Global Optimization Algorithm for Clustering ProblemsabstractThe minimum sum-of-squares clustering (MSSC) task, which can be treated as a Mixed Integer Second Order Cone Programming (MISOCP) problem, is rarely investigated in the literature through deterministic optimization to find its global optimal value. In this paper, we modelled the MSSC task as a two-stage optimization problem and proposed a tailed reduced-space branch and bound (BB) algorithm. We designed several approaches to construct lower and upper bounds at each node in the BB scheme, including a scenario grouping based Lagrangian decomposition approach. One key advantage of this reduced-space algorithm is that it only needs to perform branching on the centers of clusters to guarantee convergence, and the size of centers is independent of the number of data samples. Moreover, the lower bounds can be computed by solving small-scale sample subproblems, and upper bounds can be obtained trivially. These two properties enable our algorithm easy to be paralleled and can be scalable to the dataset with up to 200,000 samples for finding a global $\epsilon$-optimal solution of the MSSC task. We performed numerical experiments on both synthetic and real-world datasets and compared our proposed algorithms with the off-the-shelf global optimal solvers and classical local optimal algorithms. The results reveal a strong performance and scalability of our algorithm. Kaixun Hua, Mingfei Shi, Yankai Cao |
ICML | 3 |
| 2019 | A scalable global optimization algorithm for stochastic nonlinear programs
Yankai Cao, Victor M. Zavala |
J. Glob. Optim. | 1 |