EDBT 2026 Demo / reviewers in the wild / expert
Nevin Lianwen Zhang
dblp:z/NevinLianwenZhang · also Nevin L. Zhang
· DBLP profile ↗
86ranked-venue papers
28as first author
20since 2021 · last 2025
0000-0002-4662-3217ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 78 · 27 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 6 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Test-Time Adaptation on Noisy Data via Model-Pruning-Based Filtering and Flatness-Aware Entropy MinimizationabstractTest-time adaptation (TTA) deals with domain shifts during inference by training models based on only unlabeled test samples. Test samples may include noisy samples, which degrade domain adaptation. Existing methods rely on the model's output prediction to detect and filter noisy samples, and further search for flat regions during optimization, which makes the optimization more robust on noisy samples. However, there are two issues: (1) the output prediction tends to be inaccurate due to domain shifts, weakening noisy-sample detection; (2) current approaches for searching flat regions focus on optimization to enhance the worst case, which ignores achieving flatness by avoiding the quick changing of losses. To address these challenges, we propose a model pruning-based test-time adaptation model for noisy data streams, named MoTTA, which leverages a new proposed filtering, output difference under pruning (ODP)-based filtering, and a flatness-aware entropy minimization (FlatEM). Specifically, to reduce the impact of inaccurate output predictions, ODP-based filtering measures the output difference of a sample before and after model pruning, which works even under inaccurate output. To improve the search for flat loss surfaces, FlatEM integrates zeroth-order flatness and first-order flatness (minimize the maximal gradient normalization with a weight perturbation constrained in a small Euclidean ball) on entropy minimization. To solve these hard maximum problems, we leverage Taylor expansion to obtain approximated results for optimization. FlatEM also adopts a parameter regularization to mitigate incorrect updates from noisy samples. The experiments show our advantages in dealing with noisy data streams at TTA comparable to existing baselines. Xingzhi Zhou 0002, Zhiliang Tian, Ka Chun Cheung, Simon See, Nevin Lianwen Zhang |
AAAI | 9 |
| 2025 | Resilient Test-Time Adaptation by Mitigating Batch-Normalization OverfittingabstractTest-time domain adaptation adjusts a source domain model to accommodate previously unseen domain shifts in a target domain during inference. In real-world scenarios, domain shifts continually evolve, and test data are often non-independent and identically distributed (non-i.i.d.). Existing methods update batch normalization (BN) statistics (mean and variance) based on test batch statistics to mitigate domain shifts and use a memory bank to provide approximate i.i.d. sampling by selectively storing samples. However, excessive updates to BN statistics lead to overfitting to specific domain shifts. To address this issue, we propose a resilient practical test-time adaptation method (ResiTTA), employing soft constraints on the BN statistics and a low-entropy sampling strategy, which reduces overfitting on domain shifts and enables rapid adaptation. Specifically, we develop a resilient batch normalization (BN) with estimated statistics and soft constraints between the source and the estimated statistics. The soft constraints regularize the estimated statistics to mitigate overfitting caused by the excessive updates. To avoid overfitting, we design a low-entropy memory bank that accounts for sample uncertainty and class balance. We adapt the source domain model via a teacher-student self-training adaptation on the samples from the memory, incorporating the soft constraints’ updates to BN. Our ResiTTA obtains state-of-the-art results on various benchmarks. We release our code1. Xingzhi Zhou 0002, Zhiliang Tian, Xin Niu 0002, Ka Chun Cheung, Simon See, Nevin Lianwen Zhang |
ICASSP | 8 |
| 2025 | COSDA: Counterfactual-based Susceptibility Risk Framework for Open-Set Domain AdaptationabstractOpen-Set Domain Adaptation (OSDA) aims to transfer knowledge from the labeled source domain to the unlabeled target domain that contains unknown categories, thus facing the challenges of domain shift and unknown category recognition. While recent works have demonstrated the potential of causality for domain alignment, little exploration has been conducted on causal-inspired theoretical frameworks for OSDA. To fill this gap, we introduce the concept of Susceptibility and propose a novel Counterfactual-based susceptibility risk framework for OSDA, termed COSDA. Specifically, COSDA consists of three novel components: (i) a Susceptibility Risk Estimator (SRE) for capturing causal information, along with comprehensive derivations of the computable theoretical upper bound, forming a risk minimization framework under the OSDA paradigm; (ii) a Contrastive Feature Alignment (CFA) module, which is theoretically proven based on mutual information to satisfy the Exogeneity assumption and facilitate cross-domain feature alignment; (iii) a Virtual Multi-unknown-categories Prototype (VMP) pseudo-labeling strategy, providing label information by measuring how similar samples are to known and multiple virtual unknown category prototypes, thereby assisting in open-set recognition and intra-class discriminative feature learning. Extensive experiments demonstrate that our approach achieves state-of-the-art performance. Ruichun Tang, Nevin Lianwen Zhang |
ICML | 8 |
| 2024 | TCM-FTP: Fine-Tuning Large Language Models for Herbal Prescription PredictionabstractTraditional Chinese medicine (TCM) has relied on specific combinations of herbs in prescriptions to treat various symptoms and signs for thousands of years. Predicting TCM prescriptions poses a fascinating technical challenge with significant practical implications. However, this task faces limitations due to the scarcity of high-quality clinical datasets and the complex relationship between symptoms and herbs. To address these issues, we introduce DigestDS, a novel dataset comprising practical medical records from experienced experts in digestive system diseases. We also propose a method, TCM-FTP (TCM Fine-Tuning Pre-trained), to leverage pre-trained large language models (LLMs) via supervised fine-tuning on DigestDS. Additionally, we enhance computational efficiency using a low-rank adaptation technique. Moreover, TCM-FTP incorporates data augmentation by permuting herbs within prescriptions, exploiting their order-agnostic nature. Impressively, TCM-FTP achieves an F1-score of 0.8031, significantly outperforming previous methods. Furthermore, it demonstrates remarkable accuracy in dosage prediction, achieving a normalized mean square error of 0.0604. In contrast, LLMs without fine-tuning exhibit poor performance. Although LLMs have demonstrated wide-ranging capabilities, our work underscores the necessity of fine-tuning for TCM prescription prediction and presents an effective way to accomplish this. Xingzhi Zhou 0002, Xin Dong 0017, Chunhao Li, Yuning Bai, Ka Chun Cheung, Simon See, Xinpeng Song, Runshun Zhang, Xuezhong Zhou, Nevin Lianwen Zhang |
BIBM | 11 |
| 2024 | Tree-Instruct: A Preliminary Study of the Intrinsic Relationship between Complexity and AlignmentabstractTraining large language models (LLMs) with open-domain instruction data has yielded remarkable success in aligning to end tasks and human preferences. Extensive research has highlighted the importance of the quality and diversity of instruction data. However, the impact of data complexity, as a crucial metric, remains relatively unexplored from three aspects: (1)where the sustainability of performance improvements with increasing complexity is uncertain; (2)whether the improvement brought by complexity merely comes from introducing more training tokens; and (3)where the potential benefits of incorporating instructions from easy to difficult are not yet fully understood. In this paper, we propose Tree-Instruct to systematically enhance the instruction complexity in a controllable manner. By adding a specified number of nodes to instructions’ semantic trees, this approach not only yields new instruction data from the modified tree but also allows us to control the difficulty level of modified instructions. Our preliminary experiments reveal the following insights: (1)Increasing complexity consistently leads to sustained performance improvements of LLMs. (2)Under the same token budget, a few complex instructions outperform diverse yet simple instructions. (3)Curriculum instruction tuning might not yield the anticipated results; focusing on increasing complexity appears to be the key. Yingxiu Zhao, Bowen Yu 0002, Binyuan Hui, Haiyang Yu 0003, Fei Huang 0002, Nevin Lianwen Zhang, Yongbin Li 0001 |
LREC/COLING | 7 |
| 2024 | Dual Risk Minimization: Towards Next-Level Robustness in Fine-tuning Zero-Shot ModelsabstractFine-tuning foundation models often compromises their robustness to distribution shifts. To remedy this, most robust fine-tuning methods aim to preserve the pre-trained features. However, not all pre-trained features are robust and those methods are largely indifferent to which ones to preserve. We propose dual risk minimization (DRM), which combines empirical risk minimization with worst-case risk minimization, to better preserve the core features of downstream tasks. In particular, we utilize core-feature descriptions generated by LLMs to induce core-based zero-shot predictions which then serve as proxies to estimate the worst-case risk. DRM balances two crucial aspects of model robustness: expected performance and worst-case performance, establishing a new state of the art on various real-world benchmarks. DRM significantly improves the out-of-distribution performance of CLIP ViT-L/14@336 on ImageNet (75.9$\to$77.1), WILDS-iWildCam (47.1$\to$51.8), and WILDS-FMoW (50.7$\to$53.1); opening up new avenues for robust fine-tuning. Our code is available at https://github.com/vaynexie/DRM. Kaican Li, Weiyan Xie, Yongxiang Huang, Didan Deng, Lanqing Hong, Zhenguo Li, Ricardo Silva 0001, Nevin Lianwen Zhang |
NeurIPS | 8 |
| 2024 | Uncovering, Explaining, and Mitigating the Superficial Safety of Backdoor DefenseabstractBackdoor attacks pose a significant threat to Deep Neural Networks (DNNs) as they allow attackers to manipulate model predictions with backdoor triggers. To address these security vulnerabilities, various backdoor purification methods have been proposed to purify compromised models. Typically, these purified models exhibit low Attack Success Rates (ASR), rendering them resistant to backdoored inputs. However, \textit{Does achieving a low ASR through current safety purification methods truly eliminate learned backdoor features from the pretraining phase?} In this paper, we provide an affirmative answer to this question by thoroughly investigating the \textit{Post-Purification Robustness} of current backdoor purification methods. We find that current safety purification methods are vulnerable to the rapid re-learning of backdoor behavior, even when further fine-tuning of purified models is performed using a very small number of poisoned samples. Based on this, we further propose the practical Query-based Reactivation Attack (QRA) which could effectively reactivate the backdoor by merely querying purified models. We find the failure to achieve satisfactory post-purification robustness stems from the insufficient deviation of purified models from the backdoored model along the backdoor-connected path. To improve the post-purification robustness, we propose a straightforward tuning defense, Path-Aware Minimization (PAM), which promotes deviation along backdoor-connected paths with extra model updates. Extensive experiments demonstrate that PAM significantly improves post-purification robustness while maintaining a good clean accuracy and low ASR. Our work provides a new perspective on understanding the effectiveness of backdoor safety tuning and highlights the importance of faithfully assessing the model's safety. Zeyu Qin, Nevin Lianwen Zhang, Li Shen 0008, Minhao Cheng |
NeurIPS | 3 |
| 2024 | Consistency Regularization for Domain Generalization with Logit Attribution MatchingabstractDomain generalization (DG) is about training models that generalize well under domain shift. Previous research on DG has been conducted mostly in single-source or multi-source settings. In this paper, we consider a third lesser-known setting where a training domain is endowed with a collection of pairs of examples that share the same semantic information. Such semantic sharing (SS) pairs can be created via data augmentation and then utilized for consistency regularization (CR). We present a theory showing CR is conducive to DG and propose a novel CR method called Logit Attribution Matching (LAM). We conduct experiments on five DG benchmarks and four pretrained models with SS pairs created by both generic and targeted data augmentation methods. LAM outperforms representative single/multi-source DG methods and various CR methods that leverage SS pairs. The code and data of this project are available at https://github.com/Gaohan123/LAM. Han Gao 0016, Kaican Li, Weiyan Xie, Yongxiang Huang, Luning Wang, Caleb Chen Cao, Nevin Lianwen Zhang |
UAI | 8 |
| 2023 | Causal Document-Grounded Dialogue Pre-trainingabstractThe goal of document-grounded dialogue (DocGD) is to generate a response by anchoring the evidence in a supporting document in accordance with the dialogue context.This entails four causally interconnected variables.While task-specific pre-training has significantly enhanced performances on numerous downstream tasks, existing DocGD methods still rely on general pre-trained language models without a specifically tailored pre-training approach that explicitly captures the causal relationships.To address this, we present the first causallycomplete dataset construction strategy for developing million-scale DocGD pre-training corpora.Additionally, we propose a causallyperturbed pre-training strategy to better capture causality by introducing perturbations on the variables and optimizing the overall causal effect.Experiments conducted on three benchmark datasets demonstrate that our causal pretraining yields substantial and consistent improvements in fully-supervised, low-resource, few-shot, and zero-shot settings 1 . Yingxiu Zhao, Bowen Yu 0002, Bowen Li 0002, Haiyang Yu 0003, Jinyang Li 0003, Fei Huang 0002, Yongbin Li 0001, Nevin Lianwen Zhang |
EMNLP | 9 |
| 2023 | ViT-CX: Causal Explanation of Vision TransformersabstractDespite the popularity of Vision Transformers (ViTs) and eXplainable AI (XAI), only a few explanation methods have been designed specially for ViTs thus far. They mostly use attention weights of the [CLS] token on patch embeddings and often produce unsatisfactory saliency maps. This paper proposes a novel method for explaining ViTs called ViT-CX. It is based on patch embeddings, rather than attentions paid to them, and their causal impacts on the model output. Other characteristics of ViTs such as causal overdetermination are considered in the design of ViT-CX. The empirical results show that ViT-CX produces more meaningful saliency maps and does a better job revealing all important evidence for the predictions than previous methods. The explanation generated by ViT-CX also shows significantly better faithfulness to the model. The codes and appendix are available at https://github.com/vaynexie/CausalX-ViT. Weiyan Xie, Xiao-Hui Li 0009, Caleb Chen Cao, Nevin Lianwen Zhang |
IJCAI | 4 |
| 2023 | Two-stage holistic and contrastive explanation of image classificationabstractThe need to explain the output of a deep neural network classifier is now widely recognized. While previous methods typically explain a single class in the output, we advocate explaining the whole output, which is a probability distribution over multiple classes. A whole-output explanation can help a human user gain an overall understanding of model behaviour instead of only one aspect of it. It can also provide a natural framework where one can examine the evidence used to discriminate between competing classes, and thereby obtain contrastive explanations. In this paper, we propose a contrastive whole-output explanation (CWOX) method for image classification, and evaluate it using quantitative metrics and through human subject studies. The source code of CWOX is available at https://github.com/vaynexie/CWOX. Weiyan Xie, Xiao-Hui Li 0009, Leonard K. M. Poon, Caleb Chen Cao, Nevin Lianwen Zhang |
UAI | 6 |
| 2022 | Improving Meta-learning for Low-resource Text Classification and Generation via Memory ImitationabstractYingxiu Zhao, Zhiliang Tian, Huaxiu Yao, Yinhe Zheng, Dongkyu Lee, Yiping Song, Jian Sun, Nevin Zhang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Yingxiu Zhao, Zhiliang Tian, Huaxiu Yao, Yinhe Zheng, Yiping Song, Jian Sun 0021, Nevin Lianwen Zhang |
ACL (1) | 8 |
| 2022 | Emotion-Aware Multimodal Pre-training for Image-Grounded Emotional Response Generation
Zhiliang Tian, Zhihua Wen, Yiping Song, Jintao Tang, Dongsheng Li 0001, Nevin Lianwen Zhang |
DASFAA (3) | 7 |
| 2022 | Adaptive Label Smoothing with Self-Knowledge in Natural Language GenerationabstractOverconfidence has been shown to impair generalization and calibration of a neural network.Previous studies remedy this issue by adding a regularization term to a loss function, preventing a model from making a peaked distribution.Label smoothing smoothes target labels with a pre-defined prior label distribution; as a result, a model is learned to maximize the likelihood of predicting the soft label.Nonetheless, the amount of smoothing is the same in all samples and remains fixed in training.In other words, label smoothing does not reflect the change in probability distribution mapped by a model over the course of training.To address this issue, we propose a regularization scheme that brings dynamic nature into the smoothing parameter by taking model probability distribution into account, thereby varying the parameter per instance.A model in training self-regulates the extent of smoothing on the fly during forward propagation.Furthermore, inspired by recent work in bridging label smoothing and knowledge distillation, our work utilizes self-knowledge as a prior label distribution in softening target labels, and presents theoretical support for the regularization effect by knowledge distillation and the dynamic smoothing parameter.Our regularizer is validated comprehensively, and the result illustrates marked improvements in model generalization and calibration, enhancing robustness and trustworthiness of a model. Ka Chun Cheung, Nevin Lianwen Zhang |
EMNLP | 3 |
| 2022 | Hard Gate Knowledge Distillation - Leverage Calibration for Robust and Reliable Language ModelabstractIn knowledge distillation, a student model is trained with supervisions from both knowledge from a teacher and observations drawn from a training data distribution.Knowledge of a teacher is considered a subject that holds interclass relations which send a meaningful supervision to a student; hence, much effort has been put to find such knowledge to be distilled.In this paper, we explore a question that has been given little attention: "when to distill such knowledge."The question is answered in our work with the concept of model calibration; we view a teacher model not only as a source of knowledge but also as a gauge to detect miscalibration of a student.This simple and yet novel view leads to a hard gate knowledge distillation scheme that switches between learning from a teacher model and training data.We verify the gating mechanism in the context of natural language generation at both the token-level and the sentence-level.Empirical comparisons with strong baselines show that hard gate knowledge distillation not only improves model generalization, but also significantly lowers model calibration error. Zhiliang Tian, Yingxiu Zhao, Ka Chun Cheung, Nevin Lianwen Zhang |
EMNLP | 5 |
| 2022 | Prompt Conditioned VAE: Enhancing Generative Replay for Lifelong Learning in Task-Oriented DialogueabstractLifelong learning (LL) is vital for advanced task-oriented dialogue (ToD) systems.To address the catastrophic forgetting issue of LL, generative replay methods are widely employed to consolidate past knowledge with generated pseudo samples.However, most existing generative replay methods use only a single taskspecific token to control their models.This scheme is usually not strong enough to constrain the generative model due to insufficient information involved.In this paper, we propose a novel method, prompt conditioned VAE for lifelong learning (PCLL), to enhance generative replay by incorporating tasks' statistics.PCLL captures task-specific distributions with a conditional variational autoencoder, conditioned on natural language prompts to guide the pseudo-sample generation.Moreover, it leverages a distillation process to further consolidate past knowledge by alleviating the noise in pseudo samples.Experiments on natural language understanding tasks of ToD systems demonstrate that PCLL significantly outperforms competitive baselines in building lifelong learning models.We release the code and data at GitHub. Yingxiu Zhao, Yinhe Zheng, Zhiliang Tian, Jian Sun 0021, Nevin Lianwen Zhang |
EMNLP | 6 |
| 2022 | SeqPATE: Differentially Private Text Generation via Knowledge DistillationabstractProtecting the privacy of user data is crucial for text generation models, which can leak sensitive information during generation. Differentially private (DP) learning methods provide guarantees against identifying the existence of a training sample from model outputs. PATE is a recent DP learning algorithm that achieves high utility with strong privacy protection on training samples. However, text generation models output tokens sequentially in a large output space; the classic PATE algorithm is not customized for this setting. Furthermore, PATE works well to protect sample-level privacy, but is not designed to protect phrases in samples. In this paper, we propose SeqPATE, an extension of PATE to text generation that protects the privacy of individual training samples and sensitive phrases in training data. To adapt PATE to text generation, we generate pseudo-contexts and reduce the sequence generation problem to a next-word prediction problem. To handle the large output space, we propose a candidate filtering strategy to dynamically reduce the output space, and refine the teacher aggregation of PATE to avoid low agreement due to voting for a large number of candidates. To further reduce privacy losses, we use knowledge distillation to reduce the number of teacher queries. The experiments verify the effectiveness of SeqPATE in protecting both training samples and sensitive phrases. Zhiliang Tian, Yingxiu Zhao, Yu-Xiang Wang 0003, Nevin Lianwen Zhang |
NeurIPS | 5 |
| 2021 | Learning from My Friends: Few-Shot Personalized Conversation Systems via Social NetworksabstractPersonalized conversation models (PCMs) generate responses according to speaker preferences. Existing personalized conversation tasks typically require models to extract speaker preferences from user descriptions or their conversation histories, which are scarce for newcomers and inactive users. In this paper, we propose a few-shot personalized conversation task with an auxiliary social network. The task requires models to generate personalized responses for a speaker given a few conversations from the speaker and a social network. Existing methods are mainly designed to incorporate descriptions or conversation histories. Those methods can hardly model speakers with so few conversations or connections between speakers. To better cater for newcomers with few resources, we propose a personalized conversation model (PCM) that learns to adapt to new speakers as well as enabling new speakers to learn from resource-rich speakers. Particularly, based on a meta-learning based PCM, we propose a task aggregator (TA) to collect other speakers' information from the social network. The TA provides prior knowledge of the new speaker in its meta-learning. Experimental results show our methods outperform all baselines in appropriateness, diversity, and consistency with speakers. Zhiliang Tian, Wei Bi, Yiping Song, Nevin Lianwen Zhang |
AAAI | 6 |
| 2021 | Enhancing Content Preservation in Text Style Transfer Using Reverse Attention and Conditional Layer NormalizationabstractDongkyu Lee, Zhiliang Tian, Lanqing Xue, Nevin L. Zhang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Zhiliang Tian, Lanqing Xue, Nevin Lianwen Zhang |
ACL/IJCNLP (1) | 4 |
| 2021 | DeepRapper: Neural Rap Generation with Rhyme and Rhythm ModelingabstractLanqing Xue, Kaitao Song, Duocai Wu, Xu Tan, Nevin L. Zhang, Tao Qin, Wei-Qiang Zhang, Tie-Yan Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Lanqing Xue, Kaitao Song, Duocai Wu, Xu Tan 0003, Nevin Lianwen Zhang, Tao Qin 0001, Tie-Yan Liu |
ACL/IJCNLP (1) | 5 |
| 2020 | Not All Attention Is Needed: Gated Attention Network for Sequence DataabstractAlthough deep neural networks generally have fixed network structures, the concept of dynamic mechanism has drawn more and more attention in recent years. Attention mechanisms compute input-dependent dynamic attention weights for aggregating a sequence of hidden states. Dynamic network configuration in convolutional neural networks (CNNs) selectively activates only part of the network at a time for different inputs. In this paper, we combine the two dynamic mechanisms for text classification tasks. Traditional attention mechanisms attend to the whole sequence of hidden states for an input sentence, while in most cases not all attention is needed especially for long sequences. We propose a novel method called Gated Attention Network (GA-Net) to dynamically select a subset of elements to attend to using an auxiliary network, and compute attention weights to aggregate the selected elements. It avoids a significant amount of unnecessary computation on unattended elements, and allows the model to pay attention to important parts of the sequence. Experiments in various datasets show that the proposed method achieves better performance compared with all baseline models with global or local attention while requiring less computation and achieving better interpretability. It is also promising to extend the idea to more complex attention-based models, such as transformers and seq-to-seq models. Lanqing Xue, Nevin Lianwen Zhang |
AAAI | 3 |
| 2020 | Response-Anticipated Memory for On-Demand Knowledge Integration in Response GenerationabstractNeural conversation models are known to generate appropriate but non-informative responses in general.A scenario where informativeness can be significantly enhanced is Conversing by Reading (CbR), where conversations take place with respect to a given external document.In previous work, the external document is utilized by (1) creating a contextaware document memory that integrates information from the document and the conversational context, and then (2) generating responses referring to the memory.In this paper, we propose to create the document memory with some anticipated responses in mind.This is achieved using a teacher-student framework.The teacher is given the external document, the context, and the ground-truth response, and learns how to build a response-aware document memory from three sources of information.The student learns to construct a response-anticipated document memory from the first two sources, and the teacher's insight on memory creation.Empirical results show that our model outperforms the previous stateof-the-art for the CbR task. Zhiliang Tian, Wei Bi, Lanqing Xue, Yiping Song, Xiaojiang Liu, Nevin Lianwen Zhang |
ACL | 7 |
| 2020 | Learning the Structure of Auto-Encoding RecommendersabstractAutoencoder recommenders have recently shown state-of-the-art performance in the recommendation task due to their ability to model non-linear item relationships effectively. However, existing autoencoder recommenders use fully-connected neural network layers and do not employ structure learning. This can lead to inefficient training, especially when the data is sparse as commonly found in collaborative filtering. The aforementioned results in lower generalization ability and reduced performance. In this paper, we introduce structure learning for autoencoder recommenders by taking advantage of the inherent item groups present in the collaborative filtering domain. Due to the nature of items in general, we know that certain items are more related to each other than to other items. Based on this, we propose a method that first learns groups of related items and then uses this information to determine the connectivity structure of an auto-encoding neural network. This results in a network that is sparsely connected. This sparse structure can be viewed as a prior that guides the network training. Empirically we demonstrate that the proposed structure learning enables the autoencoder to converge to a local optimum with a much smaller spectral norm and generalization error bound than the fully-connected network. The resultant sparse network considerably outperforms the state-of-the-art methods like Mult-vae/Mult-dae on multiple benchmarked datasets even when the same number of parameters and flops are used. It also has a better cold-start performance. Farhan Khawar, Leonard K. M. Poon, Nevin Lianwen Zhang |
WWW | 3 |
| 2019 | Learning to Abstract for Memory-augmented Conversational Response GenerationabstractNeural generative models for open-domain chit-chat conversations have become an active area of research in recent years.A critical issue with most existing generative models is that the generated responses lack informativeness and diversity.A few researchers attempt to leverage the results of retrieval models to strengthen the generative models, but these models are limited by the quality of the retrieval results.In this work, we propose a memory-augmented generative model, which learns to abstract from the training corpus and saves the useful information to the memory to assist the response generation.Our model clusters query-response samples, extracts characteristics of each cluster, and learns to utilize these characteristics for response generation.Experimental results show that our model outperforms other competitive baselines. Zhiliang Tian, Wei Bi, Nevin Lianwen Zhang |
ACL (1) | 4 |
| 2019 | Conformative Filtering for Implicit Feedback Data
Farhan Khawar, Nevin Lianwen Zhang |
ECIR (1) | 2 |
| 2019 | A Novel Document Generation Process for Topic Detection Based on Hierarchical Latent Tree Models
Peixian Chen, Zhourong Chen, Nevin Lianwen Zhang |
ECSQARU | 3 |
| 2019 | Fast Structure Learning for Deep Feedforward Networks via Tree Skeleton Expansion
Zhourong Chen, Zhiliang Tian, Nevin Lianwen Zhang |
ECSQARU | 4 |
| 2019 | Modeling Multidimensional User Preferences for Collaborative FilteringabstractA popular idea in collaborative filtering is to map users and items to latent vectors in the same Euclidean space and make recommendations based on their inner products. The idea of user/item clustering has also been exploited. However, the possibility of obtaining latent user and item feature vectors from user/item clusters has not been investigated. In this paper, we propose such a method for implicit feedback data. We cluster users along multiple latent dimensions, with each latent dimension being defined by a distinct subset of items. User clustering along a latent dimension results in two soft groups of users: those who have a tendency to consume the corresponding items and those who do not. The first group is called a taste group. As there are multiple latent dimensions, we get multiple taste groups. We map users and items to latent feature vectors based on the taste groups such that the vector for a user tells us what tastes she possesses, and the vector for an item tells us how popular it is for users with various tastes. We call the method Multidimensional User Clustering for Collaborative Filter (MUC-CF). In comparison with other methods, MUC-CF leads to more meaningful latent factors and hence its recommendations are easier to explain. MUC-CF is also scalable and in empirical evaluations, it outperforms state-of-the-art baselines. Farhan Khawar, Nevin Lianwen Zhang |
ICDE | 2 |
| 2019 | Learning Latent Superstructures in Variational Autoencoders for Deep Multidimensional Clustering
Zhourong Chen, Leonard K. M. Poon, Nevin Lianwen Zhang |
ICLR (Poster) | 4 |
| 2019 | Cleaned Similarity for Better Memory-Based RecommendersabstractMemory-based collaborative filtering methods like user or item k-nearest neighbors (kNN) are a simple yet effective solution to the recommendation problem. The backbone of these methods is the estimation of the empirical similarity between users/items. In this paper, we analyze the spectral properties of the Pearson and the cosine similarity estimators, and we use tools from random matrix theory to argue that they suffer from noise and eigenvalues spreading. We argue that, unlike the Pearson correlation, the cosine similarity naturally possesses the desirable property of eigenvalue shrinkage for large eigenvalues. However, due to its zero-mean assumption, it overestimates the largest eigenvalues. We quantify this overestimation and present a simple re-scaling and noise cleaning scheme. This results in better performance of the memory-based methods compared to their vanilla counterparts. Farhan Khawar, Nevin Lianwen Zhang |
SIGIR | 2 |
| 2018 | Building Sparse Deep Feedforward Networks using Tree Receptive FieldsabstractSparse connectivity is an important factor behind the success of convolutional neural networks and recurrent neural networks. In this paper, we consider the problem of learning sparse connectivity for feedforward neural networks (FNNs). The key idea is that a unit should be connected to a small number of units at the next level below that are strongly correlated. We use Chow-Liu's algorithm to learn a tree-structured probabilistic model for the units at the current level, use the tree to identify subsets of units that are strongly correlated, and introduce a new unit with receptive field over the subsets. The procedure is repeated on the new units to build multiple layers of hidden units. The resulting model is called a TRF-net. Empirical results show that, when compared to dense FNNs, TRF-net achieves better or comparable classification performance with much fewer parameters and sparser structures. They are also more interpretable. Zhourong Chen, Nevin Lianwen Zhang |
IJCAI | 3 |
| 2018 | UC-LTM: Unidimensional clustering using latent tree models for discrete data
Leonard K. M. Poon, April H. Liu, Nevin Lianwen Zhang |
Int. J. Approx. Reason. | 3 |
| 2017 | Sparse Boltzmann Machines with Structure Learning as Applied to Text AnalysisabstractWe are interested in exploring the possibility and benefits of structure learning for deep models. As the first step, this paper investigates the matter for Restricted Boltzmann Machines (RBMs). We conduct the study with Replicated Softmax, a variant of RBMs for unsupervised text analysis. We present a method for learning what we call Sparse Boltzmann Machines, where each hidden unit is connected to a subset of the visible units instead of all of them. Empirical results show that the method yields models with significantly improved model fit and interpretability as compared with RBMs where each hidden unit is connected to all visible units. Zhourong Chen, Nevin Lianwen Zhang, Dit-Yan Yeung, Peixian Chen |
AAAI | 2 |
| 2017 | Latent Tree AnalysisabstractLatent tree analysis seeks to model the correlations amonga set of random variables using a tree of latent variables. It was proposed as an improvement to latent class analysis—a method widely used in social sciences and medicine to identify homogeneous subgroups in a population. It provides new and fruitful perspectives on a number of machine learningareas, including cluster analysis, topic detection, and deep probabilistic modeling. This paper gives an overview of the research on latent tree analysis and various ways it is used inpractice. Nevin Lianwen Zhang, Leonard K. M. Poon |
AAAI | 1 |
| 2017 | Latent tree models for hierarchical topic detection
Peixian Chen, Nevin Lianwen Zhang, Leonard K. M. Poon, Zhourong Chen, Farhan Khawar |
Artif. Intell. | 2 |
| 2016 | Progressive EM for Latent Tree Models and Hierarchical Topic DetectionabstractHierarchical latent tree analysis (HLTA) is recently proposed as a new method for topic detection. It differs fundamentally from the LDA-based methods in terms of topic definition, topic-document relationship, and learning method. It has been shown to discover significantly more coherent topics and better topic hierarchies. However, HLTA relies on the Expectation-Maximization (EM) algorithm for parameter estimation and hence is not efficient enough to deal with large datasets. In this paper, we propose a method to drastically speed up HLTA using a technique inspired by the advances in the method of moments. Empirical experiments show that our method greatly improves the efficiency of HLTA. It is as efficient as the state-of-the-art LDA-based method for hierarchical topic detection and finds substantially better topics and topic hierarchies. Peixian Chen, Nevin Lianwen Zhang, Leonard K. M. Poon, Zhourong Chen |
AAAI | 2 |
| 2015 | Unidimensional Clustering of Discrete Data Using Latent Tree ModelsabstractThis paper is concerned with model-based clustering of discrete data. Latent class models (LCMs) are usually used for the task. An LCM consists of a latent variable and a number of attributes. It makes the overly restrictive assumption that the attributes are mutually independent given the latent variable. We propose a novel method to relax the assumption. The key idea is to partition the attributes into groups such that correlations among the attributes in each group can be properly modeled by using one single latent variable. The latent variables for the attribute groups are then used to build a number of models and one of them is chosen to produce the clustering results. Extensive empirical studies have been conducted to compare the new method with LCM and several other methods (K-means, kernel K-means and spectral clustering) that are not model-based. The new method outperforms the alternative methods in most cases and the differences are often large. April H. Liu, Leonard K. M. Poon, Nevin Lianwen Zhang |
AAAI | 3 |
| 2015 | Bayesian adaptive matrix factorization with automatic model selectionabstractLow-rank matrix factorization has long been recognized as a fundamental problem in many computer vision applications. Nevertheless, the reliability of existing matrix factorization methods is often hard to guarantee due to challenges brought by such model selection issues as selecting the noise model and determining the model capacity. We address these two issues simultaneously in this paper by proposing a robust non-parametric Bayesian adaptive matrix factorization (AMF) model. AMF proposes a new noise model built on the Dirichlet process Gaussian mixture model (DP-GMM) by taking advantage of its high flexibility on component number selection and capability of fitting a wide range of unknown noise. AMF also imposes an automatic relevance determination (ARD) prior on the low-rank factor matrices so that the rank can be determined automatically without the need for enforcing any hard constraint. An efficient variational method is then devised for model inference. We compare AMF with state-of-the-art matrix factorization methods based on data sets ranging from synthetic data to real-world application data. From the results, AMF consistently achieves better or comparable performance. Peixian Chen, Naiyan Wang, Nevin Lianwen Zhang, Dit-Yan Yeung |
CVPR | 3 |
| 2015 | Greedy learning of latent tree models for multidimensional clustering
Nevin Lianwen Zhang, Peixian Chen, April H. Liu, Leonard K. M. Poon, Yi Wang 0006 |
Mach. Learn. | 2 |
| 2014 | Hierarchical Latent Tree Analysis for Topic Detection
Nevin Lianwen Zhang, Peixian Chen |
ECML/PKDD (2) | 2 |
| 2014 | Latent tree models for rounding in spectral clustering
April H. Liu, Leonard K. M. Poon, Nevin Lianwen Zhang |
Neurocomputing | 4 |
| 2013 | Model-based clustering of high-dimensional data: Variable selection versus facet determination
Leonard K. M. Poon, Nevin Lianwen Zhang, April H. Liu |
Int. J. Approx. Reason. | 2 |
| 2013 | LTC: A latent tree approach to classification
Yi Wang 0006, Nevin Lianwen Zhang, Leonard K. M. Poon |
Int. J. Approx. Reason. | 2 |
| 2013 | A Survey on Latent Tree Models and ApplicationsabstractIn data analysis, latent variables play a central role because they help provide powerful insights into a wide variety of phenomena, ranging from biological to human sciences. The latent tree model, a particular type of probabilistic graphical models, deserves attention. Its simple structure - a tree - allows simple and efficient inference, while its latent variables capture complex relationships. In the past decade, the latent tree model has been subject to significant theoretical and methodological developments. In this review, we propose a comprehensive study of this model. First we summarize key ideas underlying the model. Second we explain how it can be efficiently learned from data. Third we illustrate its use within three types of applications: latent structure discovery, multidimensional clustering, and probabilistic inference. Finally, we conclude and give promising directions for future researches in this field. Raphaël Mourad, Christine Sinoquet, Nevin Lianwen Zhang, Philippe Leray 0001 |
J. Artif. Intell. Res. | 3 |
| 2012 | A Model-Based Approach to Rounding in Spectral Clustering
Leonard K. M. Poon, April H. Liu, Nevin Lianwen Zhang |
UAI | 4 |
| 2012 | Model-based multidimensional clustering of categorical data
Nevin Lianwen Zhang, Leonard K. M. Poon, Yi Wang 0006 |
Artif. Intell. | 2 |
| 2011 | Latent Tree Classifier
Yi Wang 0006, Nevin Lianwen Zhang, Leonard K. M. Poon |
ECSQARU | 2 |
| 2010 | Variable Selection in Model-Based Clustering: To Do or To Facilitate
Leonard K. M. Poon, Nevin Lianwen Zhang, Yi Wang 0006 |
ICML | 2 |
| 2008 | Latent Tree Models and Approximate Inference in Bayesian Networks
Yi Wang 0006, Nevin Lianwen Zhang |
AAAI | 2 |
| 2008 | Latent tree models and diagnosis in traditional Chinese medicine
Nevin Lianwen Zhang, Shihong Yuan, Yi Wang 0006 |
Artif. Intell. Medicine | 1 |
| 2008 | Latent Tree Models and Approximate Inference in Bayesian NetworksabstractWe propose a novel method for approximate inference in Bayesian networks (BNs). The idea is to sample data from a BN, learn a latent tree model (LTM) from the data offline, and when online, make inference with the LTM instead of the original BN. Because LTMs are tree-structured, inference takes linear time. In the meantime, they can represent complex relationship among leaf nodes and hence the approximation accuracy is often good. Empirical evidence shows that our method can achieve good approximation accuracy at low online computational cost. Yi Wang 0006, Nevin Lianwen Zhang |
J. Artif. Intell. Res. | 2 |
| 2007 | Hierarchical Latent Class Models and Statistical Foundation for Traditional Chinese Medicine
Nevin Lianwen Zhang, Shihong Yuan, Yi Wang 0006 |
AIME | 1 |
| 2005 | Effective dimensions of partially observed polytrees
Tomas Kocka, Nevin Lianwen Zhang |
Int. J. Approx. Reason. | 3 |
| 2005 | Special Issue on ECSQARU-2003: The Seventh European Conference on Symbolic and Quantitative Approaches to Reasoning with Uncertainty: Message from the Guest Editors
Thomas D. Nielsen, Nevin Lianwen Zhang |
Int. J. Approx. Reason. | 2 |
| 2005 | Restricted Value Iteration: Theory and AlgorithmsabstractValue iteration is a popular algorithm for finding near optimal policies for POMDPs. It is inefficient due to the need to account for the entire belief space, which necessitates the solution of large numbers of linear programs. In this paper, we study value iteration restricted to belief subsets. We show that, together with properly chosen belief subsets, restricted value iteration yields near-optimal policies and we give a condition for determining whether a given belief subset would bring about savings in space and time. We also apply restricted value iteration to two interesting classes of POMDPs, namely informative POMDPs and near-discernible POMDPs. Nevin Lianwen Zhang |
J. Artif. Intell. Res. | 2 |
| 2004 | Efficient Learning of Hierarchical Latent Class ModelsabstractHierarchical latent class (HLC) models are tree-structured Bayesian networks where leaf nodes are observed while internal nodes are hidden. In earlier work, we have demonstrated in principle the possibility of reconstructing HLC models from data. We address the scalability issue and develop a search-based algorithm that can efficiently learn high-quality HLC models for realistic domains. There are three technical contributions: (1) the identification of a set of search operators; (2) the use of improvement in BIC score per unit of increase in model complexity, rather than BIC score itself, for model selection; and (3) the adaptation of structural EM for situations where candidate models contain different variables than the current model. The algorithm was tested on the COIL Challenge 2000 data set and an interesting model was found. Nevin Lianwen Zhang, Tomas Kocka |
ICTAI | 1 |
| 2004 | Latent variable discovery in classification models
Nevin Lianwen Zhang, Thomas D. Nielsen, Finn V. Jensen |
Artif. Intell. Medicine | 1 |
| 2004 | Effective Dimensions of Hierarchical Latent Class ModelsabstractHierarchical latent class (HLC) models are tree-structured Bayesian networks where leaf nodes are observed while internal nodes are latent. There are no theoretically well justified model selection criteria for HLC models in particular and Bayesian networks with latent nodes in general. Nonetheless, empirical studies suggest that the BIC score is a reasonable criterion to use in practice for learning HLC models. Empirical studies also suggest that sometimes model selection can be improved if standard model dimension is replaced with effective model dimension in the penalty term of the BIC score. Effective dimensions are difficult to compute. In this paper, we prove a theorem that relates the effective dimension of an HLC model to the effective dimensions of a number of latent class models. The theorem makes it computationally feasible to compute the effective dimensions of large HLC models. The theorem can also be used to compute the effective dimensions of general tree models. Nevin Lianwen Zhang, Tomas Kocka |
J. Artif. Intell. Res. | 1 |
| 2004 | Hierarchical Latent Class Models for Cluster Analysis
Nevin Lianwen Zhang |
J. Mach. Learn. Res. | 1 |
| 2003 | Effective Dimensions of Partially Observed Polytrees
Tomas Kocka, Nevin Lianwen Zhang |
ECSQARU | 2 |
| 2003 | Exploiting Contextual Independence In Probabilistic InferenceabstractBayesian belief networks have grown to prominence because they provide compact representations for many problems for which probabilistic inference is appropriate, and there are algorithms to exploit this compactness. The next step is to allow compact representations of the conditional probabilities of a variable given its parents. In this paper we present such a representation that exploits contextual independence in terms of parent contexts; which variables act as parents may depend on the value of other variables. The internal representation is in terms of contextual factors (confactors) that is simply a pair of a context and a table. The algorithm, contextual variable elimination, is based on the standard variable elimination algorithm that eliminates the non-query variables in turn, but when eliminating a variable, the tables that need to be multiplied can depend on the context. This algorithm reduces to standard variable elimination when there is no contextual independence structure to exploit. We show how this can be much more efficient than variable elimination when there is structure to exploit. We explain why this new method can exploit more structure than previous methods for structured belief network inference and an analogous algorithm that uses trees. David Poole 0001, Nevin Lianwen Zhang |
J. Artif. Intell. Res. | 2 |
| 2002 | Dimension Correction for Hierarchical Latent Class Models
Tomas Kocka, Nevin Lianwen Zhang |
UAI | 2 |
| 2001 | Space-Progressive Value Iteration: An Anytime Algorithm for a Class of POMDPs
Nevin Lianwen Zhang |
ECSQARU | 1 |
| 2001 | Speeding Up the Convergence of Value Iteration in Partially Observable Markov Decision ProcessesabstractPartially observable Markov decision processes (POMDPs) have recently become popular among many AI researchers because they serve as a natural model for planning under uncertainty. Value iteration is a well-known algorithm for finding optimal policies for POMDPs. It typically takes a large number of iterations to converge. This paper proposes a method for accelerating the convergence of value iteration. The method has been evaluated on an array of benchmark problems and was found to be very effective: It enabled value iteration to converge after only a few iterations on all the test problems. Nevin Lianwen Zhang |
J. Artif. Intell. Res. | 1 |
| 1999 | On the Role of Context-Specific Independence in Probabilistic Inference
Nevin Lianwen Zhang, David Poole 0001 |
IJCAI | 1 |
| 1999 | An Environment Model for Nonstationary Reinforcement Learning
Samuel P. M. Choi, Dit-Yan Yeung, Nevin Lianwen Zhang |
NIPS | 3 |
| 1999 | A Method for Speeding Up Value Iteration in Partially Observable Markov Decision Processes
Nevin Lianwen Zhang, Stephen S. Lee |
UAI | 1 |
| 1998 | Context-Specific Independence, Decomposition of Conditional Probabilities, and Inference in Bayesian Networks
Nevin Lianwen Zhang |
PRICAI | 1 |
| 1998 | Probabilistic Inference in Influence Diagrams
Nevin Lianwen Zhang |
UAI | 1 |
| 1998 | Planning with Partially Observable Markov Decision Processes: Advances in Exact Solution Method
Nevin Lianwen Zhang, Stephen S. Lee |
UAI | 1 |
| 1998 | Computational Properties of Two Exact Algorithms for Bayesian Networks
Nevin Lianwen Zhang |
Appl. Intell. | 1 |
| 1998 | Probabilistic Inference in Influence DiagramsabstractThis paper is about reducing influence diagram (ID) evaluation into Bayesian network (BN) inference problems that are as easy to solve as possible. Such reduction is interesting because it enables one to readily use one's favorite BN inference algorithm to efficiently evaluate IDs. Two such reduction methods have been proposed previously (Cooper 1988; Shachter and Peot 1992). This paper proposes a new method. The BN inference problems induced by the new method are much easier to solve than those induced by the two previous methods. Nevin Lianwen Zhang |
Comput. Intell. | 1 |
| 1998 | Independence of causal influence and clique tree propagation
Nevin Lianwen Zhang |
Int. J. Approx. Reason. | 1 |
| 1997 | Incremental Pruning: A Simple, Fast, Exact Method for Partially Observable Markov Decision Processes
Anthony R. Cassandra, Michael L. Littman, Nevin Lianwen Zhang |
UAI | 3 |
| 1997 | Region-Based Approximations for Planning in Stochastic Domains
Nevin Lianwen Zhang |
UAI | 1 |
| 1997 | Independence of Causal Influence and Clique Tree Propagation
Nevin Lianwen Zhang |
UAI | 1 |
| 1997 | Fast Value Iteration for Goal-Directed Markov Decision Processes
Nevin Lianwen Zhang |
UAI | 1 |
| 1997 | A Model Approximation Scheme for Planning in Partially Observable Stochastic DomainsabstractPartially observable Markov decision processes (POMDPs) are a natural model for planning problems where effects of actions are nondeterministic and the state of the world is not completely observable. It is difficult to solve POMDPs exactly. This paper proposes a new approximation scheme. The basic idea is to transform a POMDP into another one where additional information is provided by an oracle. The oracle informs the planning agent that the current state of the world is in a certain region. The transformed POMDP is consequently said to be region observable. It is easier to solve than the original POMDP. We propose to solve the transformed POMDP and use its optimal policy to construct an approximate policy for the original POMDP. By controlling the amount of additional information that the oracle provides, it is possible to find a proper tradeoff between computational time and approximation quality. In terms of algorithmic contributions, we study in details how to exploit region observability in solving the transformed POMDP. To facilitate the study, we also propose a new exact algorithm for general POMDPs. The algorithm is conceptually simple and yet is significantly more efficient than all previous exact algorithms. Nevin Lianwen Zhang |
J. Artif. Intell. Res. | 1 |
| 1996 | Irrelevance and ParameterLearning in Bayesian Networks
Nevin Lianwen Zhang |
Artif. Intell. | 1 |
| 1996 | Exploiting Causal Independence in Bayesian Network InferenceabstractA new method is proposed for exploiting causal independencies in exact Bayesian network inference. A Bayesian network can be viewed as representing a factorization of a joint probability into the multiplication of a set of conditional probabilities. We present a notion of causal independence that enables one to further factorize the conditional probabilities into a combination of even smaller factors and consequently obtain a finer-grain factorization of the joint probability. The new formulation of causal independence lets us specify the conditional probability of a variable given its parents in terms of an associative and commutative operator, such as ``or'', ``sum'' or ``max'', on the contribution of each parent. We start with a simple algorithm VE for Bayesian network inference that, given evidence and a query variable, uses the factorization to find the posterior distribution of the query. We show how this algorithm can be extended to exploit causal independence. Empirical studies, based on the CPCS networks for medical diagnosis, show that this method is more efficient than previous methods and allows for inference in larger networks than previous algorithms. Nevin Lianwen Zhang, David Poole 0001 |
J. Artif. Intell. Res. | 1 |
| 1995 | Inference with Causal Independence in the CPSC Network
Nevin Lianwen Zhang |
UAI | 1 |
| 1994 | Solving Asymmetric Decision Problems with Influence Diagrams
Runping Qi, Nevin Lianwen Zhang, David Poole 0001 |
UAI | 2 |
| 1994 | Intercausal Independence and Heterogeneous Factorization
Nevin Lianwen Zhang, David Poole 0001 |
UAI | 1 |
| 1994 | A computational theory of decision networks
Nevin Lianwen Zhang, Runping Qi, David Poole 0001 |
Int. J. Approx. Reason. | 1 |
| 1993 | Incremental computation of the value of perfect information in stepwise-decomposable influence diagrams
Nevin Lianwen Zhang, Runping Qi, David Poole 0001 |
UAI | 1 |
| 1992 | Stepwise-Decomposable Influence Diagrams
Nevin Lianwen Zhang, David Poole 0001 |
KR | 1 |