VLDB 2026 Research / reviewers in the wild / expert
Zhi-Lin Zhao 0001
dblp:189/1602 · also Zhilin Zhao 0001
· DBLP profile ↗
25ranked-venue papers
15as first author
18since 2021 · last 2026
0000-0001-6391-8423ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 14 first-author · 17 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DCAC: Dynamic Class-Aware Cache Creates Stronger Out-of-Distribution DetectorsabstractOut-of-distribution (OOD) detection remains a fundamental challenge for deep neural networks, particularly due to overconfident predictions on unseen OOD samples during testing. We reveal a key insight: OOD samples predicted as the same class, or given high probabilities for it, are visually more similar to each other than to the true in-distribution (ID) samples. Motivated by this class-specific observation, we propose DCAC (Dynamic Class-Aware Cache), a training-free, test-time calibration module that maintains separate caches for each ID class to collect high-entropy samples and calibrate the raw predictions of input samples. DCAC leverages cached visual features and predicted probabilities through a lightweight two-layer module to mitigate overconfident predictions on OOD samples. This module can be seamlessly integrated with various existing OOD detection methods across both unimodal and vision-language models while introducing minimal computational overhead. Extensive experiments on multiple OOD benchmarks demonstrate that DCAC significantly enhances existing methods, achieving substantial improvements, i.e., reducing FPR95 by 6.55% when integrated with ASH-S on ImageNet OOD benchmark. Yanqi Wu, Qichao Chen, Runhe Lai, Xinhua Lu, Jiaxin Zhuang, Zhi-Lin Zhao 0001, Wei-Shi Zheng 0001 |
AAAI | 6 |
| 2026 | Federated neural nonparametric point processesabstractTemporal point processes (TPPs) are effective for modeling event occurrences over time but struggle with sparse and uncertain events in federated systems, where privacy is a major concern. To address this, we propose FedPP , a federated neural nonparametric point process model. FedPP integrates neural embeddings into sigmoidal Gaussian Cox processes (SGCPs) on the client side. SGCPs is a flexible and expressive class of TPPs, allowing FedPP to generate highly flexible intensity functions that capture client-specific event dynamics and uncertainties while efficiently summarizing historical records. For global aggregation, FedPP introduces a divergence-based mechanism to communicate the distributions of kernel hyperparameters in SGCPs between the server and clients, while keeping client-specific parameters local to ensure privacy and personalization. FedPP effectively captures event uncertainty and sparsity. Extensive experiments demonstrate its superior performance in federated settings, showing global aggregation with the KL divergence and the Wasserstein distance. Hui Chen 0026, Xuhui Fan 0001, Hengyu Liu 0001, Yaqiong Li, Zhi-Lin Zhao 0001, Feng Zhou 0011, Christopher J. Quinn, Longbing Cao |
Artif. Intell. | 5 |
| 2025 | Mixture of Online and Offline Experts for Non-Stationary Time SeriesabstractWe consider a general and realistic scenario involving non-stationary time series, consisting of several offline intervals with different distributions within a fixed offline time horizon, and an online interval that continuously receives new samples. For non-stationary time series, the data distribution in the current online interval may have appeared in previous offline intervals. We theoretically explore the feasibility of applying knowledge from offline intervals to the current online interval. To this end, we propose the Mixture of Online and Offline Experts (MOOE). MOOE learns static offline experts from offline intervals and maintains a dynamic online expert for the current online interval. It then adaptively combines the offline and online experts using a meta expert to make predictions for the samples received in the online interval. Specifically, we focus on theoretical analysis, deriving parameter convergence, regret bounds, and generalization error bounds to prove the effectiveness of the algorithm. Zhi-Lin Zhao 0001, Longbing Cao, Yuan-Yu Wan |
AAAI | 1 |
| 2025 | Out-of-Distribution Detection by Regaining Lost Clues (Abstract Reprint)abstractOut-of-distribution (OOD) detection identifies samples in the test phase that are drawn from distributions distinct from that of training in-distribution (ID) samples for a trained network. According to the information bottleneck, networks that classify tabular data tend to extract labeling information from features with strong associations to ground-truth labels, discarding less relevant labeling cues. This behavior leads to a predicament in which OOD samples with limited labeling information receive high-confidence predictions, rendering the network incapable of distinguishing between ID and OOD samples. Hence, exploring more labeling information from ID samples, which makes it harder for an OOD sample to obtain high-confidence predictions, can address this over-confidence issue on tabular data. Accordingly, we propose a novel transformer chain (TC), which comprises a sequence of dependent transformers that iteratively regain discarded labeling information and integrate all the labeling information to enhance OOD detection. The generalization bound theoretically reveals that TC can balance ID generalization and OOD detection capabilities. Experimental results demonstrate that TC significantly surpasses state-of-the-art methods for OOD detection in tabular data. Zhi-Lin Zhao 0001, Longbing Cao, Philip S. Yu |
IJCAI | 1 |
| 2025 | SepDiff: Self-Encoding Parameter Diffusion for Learning Latent SemanticsabstractThe recently proposed Bayesian Flow Networks (BFNs) show great potential in modeling parameter spaces via a diffusion process, offering a unified strategy for handling continuous, discrete data. However, these parameter diffusion models cannot learn high-level semantic representation from the parameter space since common encoders, which encode data into one static representation, can- not capture semantic changes in parameters. This motivates a new direction: learning semantic representations hidden in the param- eter spaces to characterize noisy data. Accordingly, we propose a representation learning framework named SepDiff which operates in the parameter space to obtain parameter-wise latent semantics that exhibit progressive structures. Specifically, SepDiff proposes a self-encoder to learn latent semantics directly from parameters, rather than from observations. The encoder is then integrated into parameter diffusion model, enabling representation learning with various formats of observations. Mutual information terms further promote the disentanglement of latent semantics and capture mean- ingful semantics simultaneously. We illustrate seven representation learning tasks in SepDiff via expanding this parameter diffusion model, and extensive quantitative experimental results demonstrate the superior effectiveness of SepDiff in learning parameter repre- sentation. Zhangkai Wu, Xuhui Fan 0001, Jin Li 0028, Zhi-Lin Zhao 0001, Hui Chen 0026, Longbing Cao |
KDD (2) | 4 |
| 2025 | Exploring the Limits of Vision-Language-Action Manipulation in Cross-task GeneralizationabstractThe generalization capabilities of vision-language-action (VLA) models to unseen tasks are crucial to achieving general-purpose robotic manipulation in open-world settings.
However, the cross-task generalization capabilities of existing VLA models remain significantly underexplored.
To address this gap, we introduce **AGNOSTOS**, a novel simulation benchmark designed to rigorously evaluate cross-task zero-shot generalization in manipulation.
AGNOSTOS comprises 23 unseen manipulation tasks for test—distinct from common training task distributions—and incorporates two levels of generalization difficulty to assess robustness.
Our systematic evaluation reveals that current VLA models, despite being trained on diverse datasets, struggle to generalize effectively to these unseen tasks.
To overcome this limitation, we propose **Cross-Task In-Context Manipulation (X-ICM)**,
a method that conditions large language models (LLMs) on in-context demonstrations from seen tasks to predict action sequences for unseen tasks.
Additionally, we introduce a **dynamics-guided sample selection** strategy that identifies relevant demonstrations by capturing cross-task dynamics.
On AGNOSTOS, X-ICM significantly improves cross-task zero-shot generalization performance over leading VLAs, achieving improvements of 6.0\% over $\pi_0$ and 7.9\% over VoxPoser.
We believe AGNOSTOS and X-ICM will serve as valuable tools for advancing general-purpose robotic manipulation. Ke Ye, Teli Ma, Ronghe Qiu, Kun-Yu Lin, Zhi-Lin Zhao 0001, Junwei Liang 0001 |
NeurIPS | 8 |
| 2025 | Out-of-distribution detection by regaining lost cluesabstractOut-of-distribution (OOD) detection identifies samples in the test phase that are drawn from distributions distinct from that of training in-distribution (ID) samples for a trained network. According to the information bottleneck, networks that classify tabular data tend to extract labeling information from features with strong associations to ground-truth labels, discarding less relevant labeling cues. This behavior leads to a predicament in which OOD samples with limited labeling information receive high-confidence predictions, rendering the network incapable of distinguishing between ID and OOD samples. Hence, exploring more labeling information from ID samples, which makes it harder for an OOD sample to obtain high-confidence predictions, can address this over-confidence issue on tabular data. Accordingly, we propose a novel transformer chain (TC), which comprises a sequence of dependent transformers that iteratively regain discarded labeling information and integrate all the labeling information to enhance OOD detection. The generalization bound theoretically reveals that TC can balance ID generalization and OOD detection capabilities. Experimental results demonstrate that TC significantly surpasses state-of-the-art methods for OOD detection in tabular data. Zhi-Lin Zhao 0001, Longbing Cao, Philip S. Yu |
Artif. Intell. | 1 |
| 2025 | Escaping posterior collapse: Enhancing variational autoencoders with vine copulas
Dianlong You, Xiaoyi Ge, Chuan Lu, Dongyan Wang, Shunfu Jin, Zhi-Lin Zhao 0001 |
Inf. Sci. | 6 |
| 2025 | Distilling the Unknown to Unveil CertaintyabstractOut-of-distribution (OOD) detection is critical for identifying test samples that deviate from in-distribution (ID) data, ensuring network robustness and reliability. This paper presents a flexible framework for OOD knowledge distillation that extracts OOD-sensitive information from a network to develop a binary classifier capable of distinguishing between ID and OOD samples in both scenarios, with and without access to training ID data. To accomplish this, we introduce Confidence Amendment (CA), an innovative methodology that transforms an OOD sample into an ID one while progressively amending prediction confidence derived from the network to enhance OOD sensitivity. This approach enables the simultaneous synthesis of both ID and OOD samples, each accompanied by an adjusted prediction confidence, thereby facilitating the training of a binary classifier sensitive to OOD. Theoretical analysis provides bounds on the generalization error of the binary classifier, demonstrating the pivotal role of confidence amendment in enhancing OOD sensitivity. Extensive experiments spanning various datasets and network architectures confirm the efficacy of the proposed method in detecting OOD samples. Zhi-Lin Zhao 0001, Longbing Cao, Yixuan Zhang 0006, Kun-Yu Lin, Wei-Shi Zheng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Gray Learning From Non-IID Data With Out-of-Distribution SamplesabstractThe integrity of training data, even when annotated by experts, is far from guaranteed, especially for non-independent and identically distributed (non-IID) datasets comprising both in- and out-of-distribution samples. In an ideal scenario, the majority of samples would be in-distribution, while samples that deviate semantically would be identified as out-of-distribution and excluded during the annotation process. However, experts may erroneously classify these out-of-distribution samples as in-distribution, assigning them labels that are inherently unreliable. This mixture of unreliable labels and varied data types makes the task of learning robust neural networks notably challenging. We observe that both in- and out-of-distribution samples can almost invariably be ruled out from belonging to certain classes, aside from those corresponding to unreliable ground-truth labels. This opens the possibility of utilizing reliable complementary labels that indicate the classes to which a sample does not belong. Guided by this insight, we introduce a novel approach, termed gray learning (GL), which leverages both ground-truth and complementary labels. Crucially, GL adaptively adjusts the loss weights for these two label types based on prediction confidence levels. By grounding our approach in statistical learning theory, we derive bounds for the generalization error, demonstrating that GL achieves tight constraints even in non-IID settings. Extensive experimental evaluations reveal that our method significantly outperforms alternative approaches grounded in robust statistics. Zhi-Lin Zhao 0001, Longbing Cao, Chang-Dong Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Revealing Distribution Discrepancy by Sampling Transfer in Unlabeled DataabstractThere are increasing cases where the class labels of test samples are unavailable, creating a significant need and challenge in measuring the discrepancy between training and test distributions. This distribution discrepancy complicates the assessment of whether the hypothesis selected by an algorithm on training samples remains applicable to test samples. We present a novel approach called Importance Divergence (I-Div) to address the challenge of test label unavailability, enabling distribution discrepancy evaluation using only training samples. I-Div transfers the sampling patterns from the test distribution to the training distribution by estimating density and likelihood ratios. Specifically, the density ratio, informed by the selected hypothesis, is obtained by minimizing the Kullback-Leibler divergence between the actual and estimated input distributions. Simultaneously, the likelihood ratio is adjusted according to the density ratio by reducing the generalization error of the distribution discrepancy as transformed through the two ratios. Experimentally, I-Div accurately quantifies the distribution discrepancy, as evidenced by a wide range of complex data scenarios and tasks. Zhi-Lin Zhao 0001, Longbing Cao, Xuhui Fan 0001, Wei-Shi Zheng 0001 |
NeurIPS | 1 |
| 2024 | Weighting non-IID batches for out-of-distribution detectionabstractAbstract A standard network pretrained on in-distribution (ID) samples could make high-confidence predictions on out-of-distribution (OOD) samples, leaving the possibility of failing to distinguish ID and OOD samples in the test phase. To address this over-confidence issue, the existing methods improve the OOD sensitivity from modeling perspectives, i.e., retraining it by modifying training processes or objective functions. In contrast, this paper proposes a simple but effective method, namely Weighted Non-IID Batching (WNB), by adjusting batch weights. WNB builds on a key observation: increasing the batch size can improve the OOD detection performance. This is because a smaller batch size may make its batch samples more likely to be treated as non-IID from the assumed ID, i.e., associated with an OOD. This causes a network to provide high-confidence predictions for all samples from the OOD. Accordingly, WNB applies a weight function to weight each batch according to the discrepancy between batch samples and the entire training ID dataset. Specifically, the weight function is derived by minimizing the generalization error bound. It ensures that the weight function assigns larger weights to batches with smaller discrepancies and makes a trade-off between ID classification and OOD detection performance. Experimental results show that incorporating WNB into state-of-the-art OOD detection methods can further improve their performance. Zhi-Lin Zhao 0001, Longbing Cao |
Mach. Learn. | 1 |
| 2024 | Out-of-Distribution Detection by Cross-Class Vicinity Distribution of In-Distribution DataabstractDeep neural networks for image classification only learn to map in-distribution inputs to their corresponding ground-truth labels in training without differentiating out-of-distribution samples from in-distribution ones. This results from the assumption that all samples are independent and identically distributed (IID) without distributional distinction. Therefore, a pretrained network learned from in-distribution samples treats out-of-distribution samples as in-distribution and makes high-confidence predictions on them in the test phase. To address this issue, we draw out-of-distribution samples from the vicinity distribution of training in-distribution samples for learning to reject the prediction on out-of-distribution inputs. A cross-class vicinity distribution is introduced by assuming that an out-of-distribution sample generated by mixing multiple in-distribution samples does not share the same classes of its constituents. We, thus, improve the discriminability of a pretrained network by finetuning it with out-of-distribution samples drawn from the cross-class vicinity distribution, where each out-of-distribution input corresponds to a complementary label. Experiments on various in-/out-of-distribution datasets show that the proposed method significantly outperforms the existing methods in improving the capacity of discriminating between in- and out-of-distribution samples. Zhi-Lin Zhao 0001, Longbing Cao, Kun-Yu Lin |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Deep Spectral Copula Mechanisms Modeling Coupled and Volatile Multivariate Time SeriesabstractExploring inter- and intra-time series relations and handling volatile covariates form various challenges in modeling Coupled and Volatile Multivariate Time Series (CVMTS). A typical CVMTS data is the COVID-19 case time series across multiple countries, whose covariates may involve high volatility caused by missing samples. The existing approaches merely focus on a single set of multivariate time series or multiple multivariate time series without considering their volatile temporal covariates. They do not sufficiently characterize CVMTS features by explicitly modeling intra- and inter-MTS couplings and effectively handling volatile covariates in multiple multivariate time series. Accordingly, we propose Deep Spectral Copula Mechanisms (DSCM) to adapt CVMTS. Specifically, DSCM (1) incorporates a Singular Spectral Analysis (SSA) module to reduce the volatility of multiple covariates; (2) applies an intra-MTS coupling module to explicitly model the temporal couplings within a single set of multivariate time series; and (3) transforms target variables into joint probability distributions by Gaussian copula transformation to establish inter-MTS couplings across multiple multivariate time series. Substantial experiments on COVID-19 time-series data from multiple countries indicate the superiority of DSCM over state-of-the-art approaches. Yang Yang 0034, Zhi-Lin Zhao 0001, Longbing Cao |
DSAA | 2 |
| 2023 | R-divergence for Estimating Model-oriented Distribution DiscrepancyabstractReal-life data are often non-IID due to complex distributions and interactions, and the sensitivity to the distribution of samples can differ among learning models. Accordingly, a key question for any supervised or unsupervised model is whether the probability distributions of two given datasets can be considered identical. To address this question, we introduce R-divergence, designed to assess model-oriented distribution discrepancies. The core insight is that two distributions are likely identical if their optimal hypothesis yields the same expected risk for each distribution. To estimate the distribution discrepancy between two datasets, R-divergence learns a minimum hypothesis on the mixed data and then gauges the empirical risk difference between them. We evaluate the test power across various unsupervised and supervised tasks and find that R-divergence achieves state-of-the-art performance. To demonstrate the practicality of R-divergence, we employ R-divergence to train robust neural networks on samples with noisy labels. Zhi-Lin Zhao 0001, Longbing Cao |
NeurIPS | 1 |
| 2023 | Revealing the Distributional Vulnerability of Discriminators by Implicit GeneratorsabstractIn deep neural learning, a discriminator trained on in-distribution (ID) samples may make high-confidence predictions on out-of-distribution (OOD) samples. This triggers a significant matter for robust, trustworthy and safe deep learning. The issue is primarily caused by the limited ID samples observable in training the discriminator when OOD samples are unavailable. We propose a general approach for fine-tuning discriminators by implicit generators (FIG). FIG is grounded on information theory and applicable to standard discriminators without retraining. It improves the ability of a standard discriminator in distinguishing ID and OOD samples by generating and penalizing its specific OOD samples. According to the Shannon entropy, an energy-based implicit generator is inferred from a discriminator without extra training costs. Then, a Langevin dynamic sampler draws specific OOD samples for the implicit generator. Lastly, we design a regularizer fitting the design principle of the implicit generator to induce high entropy on those generated OOD samples. The experiments on different networks and datasets demonstrate that FIG achieves the state-of-the-art OOD detection performance. Zhi-Lin Zhao 0001, Longbing Cao, Kun-Yu Lin |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Supervision Adaptation Balancing In-Distribution Generalization and Out-of-Distribution DetectionabstractThe discrepancy between in-distribution (ID) and out-of-distribution (OOD) samples can lead to distributional vulnerability in deep neural networks, which can subsequently lead to high-confidence predictions for OOD samples. This is mainly due to the absence of OOD samples during training, which fails to constrain the network properly. To tackle this issue, several state-of-the-art methods include adding extra OOD samples to training and assign them with manually-defined labels. However, this practice can introduce unreliable labeling, negatively affecting ID classification. The distributional vulnerability presents a critical challenge for non-IID deep learning, which aims for OOD-tolerant ID classification by balancing ID generalization and OOD detection. In this paper, we introduce a novel supervision adaptation approach to generate adaptive supervision information for OOD samples, making them more compatible with ID samples. First, we measure the dependency between ID samples and their labels using mutual information, revealing that the supervision information can be represented in terms of negative probabilities across all classes. Second, we investigate data correlations between ID and OOD samples by solving a series of binary regression problems, with the goal of refining the supervision information for more distinctly separable ID classes. Our extensive experiments on four advanced network architectures, two ID datasets, and eleven diversified OOD datasets demonstrate the efficacy of our supervision adaptation approach in improving both ID classification and OOD detection capabilities. Zhi-Lin Zhao 0001, Longbing Cao, Kun-Yu Lin |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Shallow and Deep Non-IID Learning on Complex DataabstractNon-IID (i.i.d.) data holds complex non-IIDness, e.g., couplings and interactions (non-independent) and heterogeneities (not IID drawn from a given distribution). Non-IID learning emerges as a major challenge to shallow and deep learning, including classic statistical learning, mathematical modeling, shallow machine learning, and deep neural learning. Here, we outline the problem, research map, main challenges and topics of shallow and deep non-IID learning. Longbing Cao, Philip S. Yu, Zhi-Lin Zhao 0001 |
KDD | 3 |
| 2019 | LSCD: Low-rank and sparse cross-domain recommendation
Ling Huang 0002, Zhi-Lin Zhao 0001, Chang-Dong Wang 0001, Dong Huang 0001, Hongyang Chao |
Neurocomputing | 2 |
| 2018 | Low-Rank and Sparse Cross-Domain Recommendation Algorithm
Zhi-Lin Zhao 0001, Ling Huang 0002, Chang-Dong Wang 0001, Dong Huang 0001 |
DASFAA (1) | 1 |
| 2017 | Missing Value LearningabstractMissing value is common in many machine learning problems and much effort has been made to handle missing data to improve the performance of the learned model. Sometimes, our task is not to train a model using those unlabeled/labeled data with missing value but process examples according to the values of some specified features. So, there is an urgent need of developing a method to predict those missing values. In this paper, we focus on learning from the known values to learn missing value as close as possible to the true one. It's difficult for us to predict missing value because we do not know the structure of the data matrix and some missing values may relate to some other missing values. We solve the problem by recovering the complete data matrix under the three reasonable constraints: feature relationship, upper recovery error bound and class relationship. The proposed algorithm can deal with both unlabeled and labeled data and generative adversarial idea will be used in labeled data to transfer knowledge. Extensive experiments have been conducted to show the effectiveness of the proposed algorithms. Zhi-Lin Zhao 0001, Chang-Dong Wang 0001, Kun-Yu Lin, Jian-Huang Lai |
CIKM | 1 |
| 2017 | Low-Rank and Sparse Matrix Completion for Recommendation
Zhi-Lin Zhao 0001, Ling Huang 0002, Chang-Dong Wang 0001, Jian-Huang Lai, Philip S. Yu |
ICONIP (5) | 1 |
| 2017 | Multi-view Unit Intact Space Learning
Kun-Yu Lin, Chang-Dong Wang 0001, Yu-Qin Meng, Zhi-Lin Zhao 0001 |
KSEM | 4 |
| 2017 | An item orientated recommendation algorithm from the multi-view perspective
Qi-Ying Hu, Zhi-Lin Zhao 0001, Chang-Dong Wang 0001, Jian-Huang Lai |
Neurocomputing | 2 |
| 2016 | FTMF: Recommendation in social network with Feature Transfer and Probabilistic Matrix FactorizationabstractIt is well-known that recommendation system which is widely used in many e-commerce platforms to recommend items to the right users suffers from data sparsity, imbalanced rating and cold start problems. Matrix factorization is a good way to deal with the sparsity and imbalance problems, which is however unable to make prediction for new users due to the lack of auxiliary information. With the advent of online social networks, the trust relation in the social network can be utilized as auxiliary data to solve the aforementioned problems since consumers' buying behavior is usually affected by the people around them. This paper reports a study of exploiting the trust relationship in social network for personalized recommendation. Although previous studies have paid attention to this topic, we improve the quality of recommendation further and solve the cold start problem better. To this end, we propose a recommendation system in social network with Feature Transfer and Probabilistic Matrix Factorization (FTMF). The auxiliary data and matrix factorization technique are integrated to learn a social latent feature vector of users which represents the features transferred from trusted people. And an adaptive firm factor is introduced to balance the impact from user's own factors and trusted people on buying behavior for each user. The experimental results show that our model can effectively use the auxiliary data and outperforms the existing state-of-the-art social network based recommendation algorithms. Zhi-Lin Zhao 0001, Chang-Dong Wang 0001, Yuan-Yu Wan, Jian-Huang Lai, Dong Huang 0001 |
IJCNN | 1 |