VLDB 2026 Research / reviewers in the wild / expert
Zhengdong Lu
dblp:33/3562
· DBLP profile ↗
54ranked-venue papers
12as first author
3since 2021 · last 2025
0000-0002-6418-6030ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 46 · 11 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-authorDatabases, data management, data science and information retrieval · 10 · 2 first-authorSoftware engineering, systems software and programming languages · 2 · 1 since 2021Systems, architecture and hardware · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
31 papers |
Language models and text generation · 15% Information extraction and text analysis · 15% Representation and self-supervised learning · 15% | |
| Databases, data mining, and information retrieval
15 papers |
Information retrieval · 31% Data mining · 23% Web and social media mining · 22% | |
| Network and information security
1 paper |
Security and privacy of machine learning · 100% |
Topics — the 30 heaviest of 86, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Machine translation
neural machine translation |
1.3 | 5 | 2017 | Deep Neural Machine Translation with Linear Associative Unit · ACL (1) 2017 Neural Machine Translation Advised by Statistical Machine Translation · AAAI 2017 Memory-enhanced Decoder for Neural Machine Translation · EMNLP 2016 |
Natural language and speech › Language models and text generation › trustworthy language model
privacy-preserving inference |
0.9 | 1 | 2025 | PrivacyRestore: Privacy-Preserving Inference in Large Language Models via Privacy Removal and Restoration · ACL (1) 2025 |
Security and privacy of machine learning
privacy-preserving inference |
0.9 | 1 | 2025 | PrivacyRestore: Privacy-Preserving Inference in Large Language Models via Privacy Removal and Restoration · ACL (1) 2025 |
Natural language and speech › Information extraction and text analysis
semantic parsing |
0.6 | 2 | 2018 | Object-oriented Neural Programming (OONP) for Document Understanding · ACL (1) 2018 Coupling Distributed and Symbolic Execution for Natural Language Queries · ICML 2017 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.5 | 3 | 2016 | genCNN: A Convolutional Architecture for Word Sequence Prediction · ACL (1) 2015 Encoding Source Language with Convolutional Neural Network for Machine Translation · ACL (1) 2015 Learning to Answer Questions from Image Using Convolutional Neural Network · AAAI 2016 |
Web and social media mining
social network analysis |
0.5 | 4 | 2013 | Collaborative boosting for activity classification in microblogs · KDD 2013 Clustered embedding of massive social networks · SIGMETRICS 2012 Supervised Link Prediction Using Multiple Sources · ICDM 2010 |
Natural language and speech › Question answering and dialogue systems › dialogue
short-text conversation |
0.4 | 2 | 2015 | Neural Responding Machine for Short-Text Conversation · ACL (1) 2015 A Dataset for Research on Short-Text Conversations · EMNLP 2013 |
Natural language and speech › Information extraction and text analysis
named entity recognition |
0.4 | 1 | 2019 | A Prism Module for Semantic Disentanglement in Name Entity Recognition · ACL (1) 2019 |
Machine learning › Representation and self-supervised learning › representation learning › disentangled representation learning › disentanglement
semantic disentanglement |
0.4 | 1 | 2019 | A Prism Module for Semantic Disentanglement in Name Entity Recognition · ACL (1) 2019 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › ontology
ontology construction |
0.3 | 1 | 2018 | Object-oriented Neural Programming (OONP) for Document Understanding · ACL (1) 2018 |
Natural language and speech › Information extraction and text analysis
text classification |
0.3 | 1 | 2018 | Jumper: Learning When to Make Classification Decision in Reading · IJCAI 2018 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › query answering
knowledge base querying |
0.3 | 1 | 2017 | Coupling Distributed and Symbolic Execution for Natural Language Queries · ICML 2017 |
Natural language and speech › Question answering and dialogue systems
knowledge base question answering |
0.3 | 1 | 2017 | Coupling Distributed and Symbolic Execution for Natural Language Queries · ICML 2017 |
Machine learning › Deep learning architectures and training
recurrent neural network |
0.3 | 1 | 2017 | Deep Neural Machine Translation with Linear Associative Unit · ACL (1) 2017 |
Knowledge graphs
link prediction |
0.3 | 2 | 2012 | Clustered embedding of massive social networks · SIGMETRICS 2012 Supervised Link Prediction Using Multiple Sources · ICDM 2010 |
Natural language and speech › Language models and text generation › text generation › neural text generation
copy mechanism |
0.2 | 1 | 2016 | Incorporating Copying Mechanism in Sequence-to-Sequence Learning · ACL (1) 2016 |
Natural language and speech › Question answering and dialogue systems › answer generation
generative question answering |
0.2 | 1 | 2016 | Neural Generative Question Answering · IJCAI 2016 |
Machine learning › Representation and self-supervised learning
multimodal representation learning |
0.2 | 1 | 2016 | Learning to Answer Questions from Image Using Convolutional Neural Network · AAAI 2016 |
Machine learning › Deep learning architectures and training › sequence modeling
sequence-to-sequence learning |
0.2 | 1 | 2016 | Incorporating Copying Mechanism in Sequence-to-Sequence Learning · ACL (1) 2016 |
Computer vision › Vision and language
visual question answering |
0.2 | 1 | 2016 | Learning to Answer Questions from Image Using Convolutional Neural Network · AAAI 2016 |
Data models and query languages
natural language interface |
0.2 | 1 | 2016 | Neural Enquirer: Learning to Query Tables in Natural Language · IJCAI 2016 |
Information retrieval › retrieval models
neural retrieval |
0.2 | 1 | 2016 | Deep Learning for Information Retrieval · SIGIR 2016 |
Information retrieval › question answering
table question answering |
0.2 | 1 | 2016 | Neural Enquirer: Learning to Query Tables in Natural Language · IJCAI 2016 |
Machine learning › Representation and self-supervised learning › representation learning
dimensionality reduction |
0.2 | 4 | 2010 | Parametric dimensionality reduction by unsupervised regression · CVPR 2010 Dimensionality reduction by unsupervised regression · CVPR 2008 Geometry-aware metric learning · ICML 2009 |
Data mining
clustering |
0.2 | 3 | 2009 | Clustering with Multiple Graphs · ICDM 2009 Constrained spectral clustering through affinity propagation · CVPR 2008 Semi-supervised Learning with Penalized Probabilistic Clustering · NIPS 2004 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
manifold learning |
0.2 | 3 | 2010 | Parametric dimensionality reduction by unsupervised regression · CVPR 2010 Dimensionality reduction by unsupervised regression · CVPR 2008 Geometry-aware metric learning · ICML 2009 |
Computer vision › Vision and language › cross-modal matching
image-text matching |
0.2 | 1 | 2015 | Multimodal Convolutional Neural Networks for Matching Image and Sentence · ICCV 2015 |
Natural language and speech › Question answering and dialogue systems › dialogue generation › dialogue response generation
neural response generation |
0.2 | 1 | 2015 | Neural Responding Machine for Short-Text Conversation · ACL (1) 2015 |
Machine learning › Representation and self-supervised learning › text embedding
sentence embedding |
0.2 | 1 | 2015 | Self-Adaptive Hierarchical Sentence Model · IJCAI 2015 |
Natural language and speech › Information extraction and text analysis
text matching |
0.2 | 1 | 2015 | Syntax-Based Deep Matching of Short Texts · IJCAI 2015 |
Methods — techniques the papers use, named apart from their topics
privacy restoration · 1.7privacy removal · 1.7convolutional neural network · 1.3neural network · 1.1reinforcement learning · 0.7recurrent neural network · 0.5semantic disentanglement · 0.4parallelization · 0.4coordinate descent · 0.4supervised learning · 0.3policy network · 0.3auxiliary classifier · 0.3deep learning · 0.2multimodal representation learning · 0.2retrieval-based conversation model · 0.2gradient descent · 0.2collaborative learning · 0.2boosting · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PrivacyRestore: Privacy-Preserving Inference in Large Language Models via Privacy Removal and RestorationabstractZiqian Zeng, Jianwei Wang, Junyao Yang, Zhengdong Lu, Haoran Li, Huiping Zhuang, Cen Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Ziqian Zeng, Junyao Yang, Zhengdong Lu, Huiping Zhuang, Cen Chen 0002 |
ACL (1) | 4 |
| 2025 | Subkv: Quantizing Long Context KV Cache for Sub-Billion Parameter Language Models on Edge DevicesabstractABSTRACT Background Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks. However, their substantial computational and memory requirements present significant challenges for widespread deployment on edge devices. Motivation In long‐context scenarios, even sub‐billion parameter LLMs face unavoidable memory and performance bottlenecks due to inefficient KV Cache utilization. Existing quantization methods fail to address these challenges effectively. Method This paper addresses these challenges by introducing advanced quantization techniques tailored for sub‐billion parameter LLMs. It specifically targets reducing memory consumption through the conversion of the model's KV Cache to lower‐bit integers. We present SubKV, a quantization method specifically designed to optimize the KV Cache in sub‐billion parameter LLMs. Our analysis reveals distinct distributional differences in the magnitude of key and value caches. Leveraging this insight, we apply Per‐Channel Quantization to the key cache and Per‐Token Quantization to the value cache. Furthermore, we introduce the Dynamic Window Quantization method to enhance attention computations. To mitigate the extreme sensitivity of the first token, we also introduce Attention Sink‐Aware Quantization. Results Experimental results demonstrate that SubKV significantly reduces the KV Cache size during long context inference while maintaining model performance, offering superior results to existing KV Cache quantization methods. Ziqian Zeng, Tao Zhang 0019, Zhengdong Lu, Huiping Zhuang, Hongen Shao, Sin G. Teo, Xiaofeng Zou |
Softw. Pract. Exp. | 3 |
| 2024 | Zero-shot Event Detection Using a Textual Entailment Model as an Enhanced AnnotatorabstractZero-shot event detection is a challenging task. Recent research work proposed to use a pre-trained textual entailment (TE) model on this task. However, those methods treated the TE model as a frozen annotator. We treat the TE model as an annotator that can be enhanced. We propose to use TE models to annotate large-scale unlabeled text and use annotated data to finetune the TE model, yielding an improved TE model. Finally, the improved TE model is used for inference on the test set. To improve the efficiency, we propose to use keywords to filter out sentences with a low probability of expressing event(s). To improve the coverage of keywords, we expand limited number of seed keywords using WordNet, so that we can use the TE model to annotate unlabeled text efficiently. The experimental results show that our method can outperform other baselines by 15% on the ACE05 dataset. Ziqian Zeng, Runyu Wu, Yuxiang Xiao, Xiaoda Zhong, Zhengdong Lu, Huiping Zhuang |
LREC/COLING | 6 |
| 2020 | Finding decision jumps in text classification
Xianggen Liu, Lili Mou, Haotian Cui, Zhengdong Lu, Sen Song |
Neurocomputing | 4 |
| 2020 | Outline Extraction with Question-Specific Memory CellsabstractOutline extraction has been widely applied in online consultation to help experts quickly understand individual cases. Given a specific case described as unstructured plain text, outline extraction aims to make a summary for this case by answering a set of questions, which in fact is a new type of machine reading comprehension task. Inspired by a recently popular memory network, we propose a novel question-specific memory cell network (QSMCN) to extract information related to multiple questions on-the-fly as it reads texts. QSMCN constructs a specific memory cell for each question, which is sequentially expanded in recurrent neural network style. Each cell contains three specific vectors to first identify whether current input is related to corresponding question and then update question-specific case representation. We add a penalization term in the loss function to make extracted knowledge more reasonable and interpretable. To support this study, we construct a new outline extraction corpus, InjuryCase, 1 which is composed of 3,995 real Chinese occupational injury cases. Experimental results show that our method makes a significant improvement. We further apply the proposed framework on two multi-aspect extraction tasks and find that the proposed model also remarkably outperforms existing state-of-the-art methods of the aspect extraction task. Haotian Cui, Si Li 0001, Sheng Gao 0001, Jun Guo 0002, Zhengdong Lu |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 6 |
| 2019 | A Prism Module for Semantic Disentanglement in Name Entity RecognitionabstractNatural Language Processing has been perplexed for many years by the problem that multiple semantics are mixed inside a word, even with the help of context.To solve this problem, we propose a prism module to disentangle the semantic aspects of words and reduce noise at the input layer of a model.In the prism module, some words are selectively replaced with task-related semantic aspects, then these denoised word representations can be fed into downstream tasks to make them easier.Besides, we also introduce a structure to train this module jointly with the downstream model without additional data.This module can be easily integrated into the downstream model and significantly improve the performance of baselines on named entity recognition (NER) task.The ablation analysis demonstrates the rationality of the method.As a side effect, the proposed method also provides a way to visualize the contribution of each word.1 Daqi Zheng, Zhengdong Lu, Sheng Gao 0001, Si Li 0001 |
ACL (1) | 4 |
| 2018 | Object-oriented Neural Programming (OONP) for Document UnderstandingabstractWe propose Object-oriented Neural Programming (OONP), a framework for semantically parsing documents in specific domains.Basically, OONP reads a document and parses it into a predesigned object-oriented data structure that reflects the domain-specific semantics of the document.An OONP parser models semantic parsing as a decision process: a neural netbased Reader sequentially goes through the document, and builds and updates an intermediate ontology during the process to summarize its partial understanding of the text.OONP supports a big variety of forms (both symbolic and differentiable) for representing the state and the document, and a rich family of operations to compose the representation.An OONP parser can be trained with supervision of different forms and strength, including supervised learning (SL) , reinforcement learning (RL) and hybrid of the two.Our experiments on both synthetic and real-world document parsing tasks have shown that OONP can learn to handle fairly complicated ontology with training data of modest sizes.* The work was done when these authors worked as interns at DeeplyCurious.ai. Zhengdong Lu, Xianggen Liu, Haotian Cui, Yukun Yan, Daqi Zheng |
ACL (1) | 1 |
| 2018 | Jumper: Learning When to Make Classification Decision in ReadingabstractIn early years, text classification is typically accomplished by feature-based classifiers; recently, neural networks, as powerful classifiers, make it possible to work with raw input as the text stands. In this paper, we propose a novel framework, Jumper, inspired by the cognitive process of text reading, that models text classification as a sequential decision process. Basically, Jumper is a neural system that can scan a piece of text sequentially and make classification decision at the time it chooses. Both the classification and when to make the classification are part of the decision process which are controlled by the policy net and trained with reinforcement learning to maximize the overall classification accuracy. Experimental results show that a properly trained Jumper has the following properties: (1) It can make decisions whenever the evidence is enough, therefore reducing the total text reading by 30~40% and often finding the key rationale of prediction. (2) It can achieve classification accuracy better or comparable to state-of-the-art model in several benchmark and industrial datasets. Xianggen Liu, Lili Mou, Haotian Cui, Zhengdong Lu, Sen Song |
IJCAI | 4 |
| 2017 | Neural Machine Translation Advised by Statistical Machine TranslationabstractNeural Machine Translation (NMT) is a new approach to machine translation that has made great progress in recent years. However, recent studies show that NMT generally produces fluent but inadequate translations (Tu et al. 2016b; 2016a; He et al. 2016; Tu et al. 2017). This is in contrast to conventional Statistical Machine Translation (SMT), which usually yields adequate but non-fluent translations. It is natural, therefore, to leverage the advantages of both models for better translations, and in this work we propose to incorporate SMT model into NMT framework. More specifically, at each decoding step, SMT offers additional recommendations of generated words based on the decoding information from NMT (e.g., the generated partial translation and attention history). Then we employ an auxiliary classifier to score the SMT recommendations and a gating function to combine the SMT recommendations with NMT generations, both of which are jointly trained within the NMT architecture in an end-to-end manner. Experimental results on Chinese-English translation show that the proposed approach achieves significant and consistent improvements over state-of-the-art NMT and SMT systems on multiple NIST test sets. Xing Wang 0007, Zhengdong Lu, Zhaopeng Tu, Hang Li 0001, Deyi Xiong, Min Zhang 0005 |
AAAI | 2 |
| 2017 | Deep Neural Machine Translation with Linear Associative UnitabstractDeep Neural Networks (DNNs) have provably enhanced the state-of-the-art Neural Machine Translation (NMT) with their capability in modeling complex functions and capturing complex linguistic structures.However NMT systems with deep architecture in their encoder or decoder RNNs often suffer from severe gradient diffusion due to the non-linear recurrent activations, which often make the optimization much more difficult.To address this problem we propose novel linear associative units (LAU) to reduce the gradient propagation length inside the recurrent unit.Different from conventional approaches (LSTM unit and GRU), LAUs utilizes linear associative connections between input and output of the recurrent unit, which allows unimpeded information flow through both space and time direction.The model is quite simple, but it is surprisingly effective.Our empirical study on Chinese-English translation shows that our model with proper configuration can improve by 11.7 BLEU upon Groundhog and the best reported results in the same setting.On WMT14 English-German task and a larger WMT14 English-French task, our model achieves comparable results with the state-of-the-art. Mingxuan Wang, Zhengdong Lu, Jie Zhou 0016, Qun Liu 0001 |
ACL (1) | 2 |
| 2017 | Coupling Distributed and Symbolic Execution for Natural Language QueriesabstractBuilding neural networks to query a knowledge base (a table) with natural language is an emerging research topic in deep learning. An executor for table querying typically requires multiple steps of execution because queries may have complicated structures. In previous studies, researchers have developed either fully distributed executors or symbolic executors for table querying. A distributed executor can be trained in an end-to-end fashion, but is weak in terms of execution efficiency and explicit interpretability. A symbolic executor is efficient in execution, but is very difficult to train especially at initial stages. In this paper, we propose to couple distributed and symbolic execution for natural language queries, where the symbolic executor is pretrained with the distributed executor’s intermediate execution results in a step-by-step fashion. Experiments show that our approach significantly outperforms both distributed and symbolic executors, exhibiting high accuracy, high learning efficiency, high execution efficiency, and high interpretability. Lili Mou, Zhengdong Lu, Hang Li 0001, Zhi Jin 0001 |
ICML | 2 |
| 2017 | Context Gates for Neural Machine TranslationabstractIn neural machine translation (NMT), generation of a target word depends on both source and target contexts. We find that source contexts have a direct impact on the adequacy of a translation while target contexts affect the fluency. Intuitively, generation of a content word should rely more on the source context and generation of a functional word should rely more on the target context. Due to the lack of effective control over the influence from source and target contexts, conventional NMT tends to yield fluent but inadequate translations. To address this problem, we propose context gates which dynamically control the ratios at which source and target contexts contribute to the generation of target words. In this way, we can enhance both the adequacy and fluency of NMT with more careful control of the information flow from contexts. Experiments show that our approach significantly improves upon a standard attention-based NMT system by +2.3 BLEU points. Zhaopeng Tu, Yang Liu 0005, Zhengdong Lu, Hang Li 0001 |
Trans. Assoc. Comput. Linguistics | 3 |
| 2016 | Learning to Answer Questions from Image Using Convolutional Neural NetworkabstractIn this paper, we propose to employ the convolutional neural network (CNN) for the image question answering (QA) task. Our proposed CNN provides an end-to-end framework with convolutional architectures for learning not only the image and question representations, but also their inter-modal interactions to produce the answer. More specifically, our model consists of three CNNs: one image CNN to encode the image content, one sentence CNN to compose the words of the question, and one multimodal convolution layer to learn their joint representation for the classification in the space of candidate answer words. We demonstrate the efficacy of our proposed model on the DAQUAR and COCO-QA datasets, which are two benchmark datasets for image QA, with the performances significantly outperforming the state-of-the-art. Zhengdong Lu, Hang Li 0001 |
AAAI | 2 |
| 2016 | Incorporating Copying Mechanism in Sequence-to-Sequence LearningabstractWe address an important problem in sequence-to-sequence (Seq2Seq) learning referred to as copying, in which certain segments in the input sequence are selectively replicated in the output sequence. A similar phenomenon is observable in human language communication. For example, humans tend to repeat entity names or even long phrases in conversation. The challenge with regard to copying in Seq2Seq is that new machinery is needed to decide when to perform the operation. In this paper, we incorporate copying into neural network-based Seq2Seq learning and propose a new model called CopyNet with encoder-decoder structure. CopyNet can nicely integrate the regular way of word generation in the decoder with the new copying mechanism which can choose sub-sequences in the input sequence and put them at proper places in the output sequence. Our empirical study on both synthetic data sets and real world data sets demonstrates the efficacy of CopyNet. For example, CopyNet can outperform regular RNN-based model with remarkable margins on text summarization tasks. Jiatao Gu, Zhengdong Lu, Hang Li 0001, Victor O. K. Li |
ACL (1) | 2 |
| 2016 | Modeling Coverage for Neural Machine TranslationabstractAttention mechanism has enhanced stateof-the-art Neural Machine Translation (NMT) by jointly learning to align and translate.It tends to ignore past alignment information, however, which often leads to over-translation and under-translation.To address this problem, we propose coverage-based NMT in this paper.We maintain a coverage vector to keep track of the attention history.The coverage vector is fed to the attention model to help adjust future attention, which lets NMT system to consider more about untranslated source words.Experiments show that the proposed approach significantly improves both translation quality and alignment quality over standard attention-based NMT. 1 Zhaopeng Tu, Zhengdong Lu, Yang Liu 0005, Hang Li 0001 |
ACL (1) | 2 |
| 2016 | Interactive Attention for Neural Machine TranslationabstractConventional attention-based Neural Machine Translation (NMT) conducts dynamic alignment in generating the target sentence. By repeatedly reading the representation of source sentence, which keeps fixed after generated by the encoder (Bahdanau et al., 2015), the attention mechanism has greatly enhanced state-of-the-art NMT. In this paper, we propose a new attention mechanism, called INTERACTIVE ATTENTION, which models the interaction between the decoder and the representation of source sentence during translation by both reading and writing operations. INTERACTIVE ATTENTION can keep track of the interaction history and therefore improve the translation performance. Experiments on NIST Chinese-English translation task show that INTERACTIVE ATTENTION can achieve significant improvements over both the previous attention-based NMT baseline and some state-of-the-art variants of attention-based NMT (i.e., coverage models (Tu et al., 2016)). And neural machine translator with our INTERACTIVE ATTENTION can outperform the open source attention-based NMT system Groundhog by 4.22 BLEU points and the open source phrase-based system Moses by 3.94 BLEU points averagely on multiple test sets. Fandong Meng, Zhengdong Lu, Hang Li 0001, Qun Liu 0001 |
COLING | 2 |
| 2016 | Memory-enhanced Decoder for Neural Machine TranslationabstractWe propose to enhance the RNN decoder in a neural machine translator (NMT) with external memory, as a natural but powerful extension to the state in the decoding RNN.This memory-enhanced RNN decoder is called MEMDEC.At each time during decoding, MEMDEC will read from this memory and write to this memory once, both with content-based addressing.Unlike the unbounded memory in previous work (Bahdanau et al., 2014) to store the representation of source sentence, the memory in MEMDEC is a matrix with predetermined size designed to better capture the information important for the decoding process at each time step.Our empirical study on Chinese-English translation shows that it can improve by 4.8 BLEU upon Groundhog and 5.3 BLEU upon on Moses, yielding the best performance achieved with the same training set. Mingxuan Wang, Zhengdong Lu, Hang Li 0001, Qun Liu 0001 |
EMNLP | 2 |
| 2016 | Neural Generative Question Answering
Xin Jiang 0002, Zhengdong Lu, Lifeng Shang, Hang Li 0001 |
IJCAI | 3 |
| 2016 | Neural Enquirer: Learning to Query Tables in Natural Language
Zhengdong Lu, Hang Li 0001, Ben Kao |
IJCAI | 2 |
| 2016 | Deep Learning for Information RetrievalabstractRecent years have observed a significant progress in information retrieval and natural language processing with deep learning technologies being successfully applied into almost all of their major tasks. The key to the success of deep learning is its capability of accurately learning distributed representations (vector representations or structured arrangement of them) of natural language expressions such as sentences, and effectively utilizing the representations in the tasks. This tutorial aims at summarizing and introducing the results of recent research on deep learning for information retrieval, in order to stimulate and foster more significant research and development work on the topic in the future. Hang Li 0001, Zhengdong Lu |
SIGIR | 2 |
| 2015 | Encoding Source Language with Convolutional Neural Network for Machine TranslationabstractFandong Meng, Zhengdong Lu, Mingxuan Wang, Hang Li, Wenbin Jiang, Qun Liu. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Fandong Meng, Zhengdong Lu, Mingxuan Wang, Hang Li 0001, Wenbin Jiang 0002, Qun Liu 0001 |
ACL (1) | 2 |
| 2015 | Neural Responding Machine for Short-Text ConversationabstractLifeng Shang, Zhengdong Lu, Hang Li. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Lifeng Shang, Zhengdong Lu, Hang Li 0001 |
ACL (1) | 2 |
| 2015 | genCNN: A Convolutional Architecture for Word Sequence PredictionabstractMingxuan Wang, Zhengdong Lu, Hang Li, Wenbin Jiang, Qun Liu. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Mingxuan Wang, Zhengdong Lu, Hang Li 0001, Wenbin Jiang 0002, Qun Liu 0001 |
ACL (1) | 2 |
| 2015 | Multimodal Convolutional Neural Networks for Matching Image and SentenceabstractIn this paper, we propose multimodal convolutional neural networks (m-CNNs) for matching image and sentence. Our m-CNN provides an end-to-end framework with convolutional architectures to exploit image representation, word composition, and the matching relations between the two modalities. More specifically, it consists of one image CNN encoding the image content and one matching CNN modeling the joint representation of image and sentence. The matching CNN composes different semantic fragments from words and learns the inter-modal relations between image and the composed fragments at different levels, thus fully exploit the matching relations between image and sentence. Experimental results demonstrate that the proposed m-CNNs can effectively capture the information necessary for image and sentence matching. More specifically, our proposed m-CNNs significantly outperform the state-of-the-art approaches for bidirectional image and sentence retrieval on the Flickr8K and Flickr30K datasets. Zhengdong Lu, Lifeng Shang, Hang Li 0001 |
ICCV | 2 |
| 2015 | Syntax-Based Deep Matching of Short Texts
Mingxuan Wang, Zhengdong Lu, Hang Li 0001, Qun Liu 0001 |
IJCAI | 2 |
| 2015 | Self-Adaptive Hierarchical Sentence Model
Han Zhao 0002, Zhengdong Lu, Pascal Poupart |
IJCAI | 2 |
| 2014 | A Parallel and Efficient Algorithm for Learning to MatchabstractMany tasks in data mining and related fields can be formalized as matching between objects in two heterogeneous domains, including collaborative filtering, link prediction, image tagging, and web search. Machine learning techniques, referred to as learning-to-match in this paper, have been successfully applied to the problems. Among them, a class of state-of-the-art methods, named feature-based matrix factorization, formalize the task as an extension to matrix factorization by incorporating auxiliary features into the model. Unfortunately, making those algorithms scale to real world problems is challenging, and simple parallelization strategies fail due to the complex cross talking patterns between sub-tasks. In this paper, we tackle this challenge with a novel parallel and efficient algorithm. Our algorithm, based on coordinate descent, can easily handle hundreds of millions of instances and features on a single machine. The key recipe of this algorithm is an iterative relaxation of the objective to facilitate parallel updates of parameters, with guaranteed convergence on minimizing the original objective function. Experimental results demonstrate that the proposed method is effective on a wide range of matching problems, with efficiency significantly improved upon the baselines while accuracy retained unchanged. Jingbo Shang, Tianqi Chen 0001, Hang Li 0001, Zhengdong Lu, Yong Yu 0001 |
ICDM | 4 |
| 2014 | Convolutional Neural Network Architectures for Matching Natural Language Sentences
Baotian Hu, Zhengdong Lu, Hang Li 0001, Qingcai Chen |
NIPS | 2 |
| 2013 | A Dataset for Research on Short-Text ConversationsabstractNatural language conversation is widely regarded as a highly difficult problem, which is usually attacked with either rule-based or learning-based models.In this paper we propose a retrieval-based automatic response model for short-text conversation, to exploit the vast amount of short conversation instances available on social media.For this purpose we introduce a dataset of short-text conversation based on the real-world instances from Sina Weibo (a popular Chinese microblog service), which will be soon released to public.This dataset provides rich collection of instances for the research on finding natural and relevant short responses to a given short text, and useful for both training and testing of conversation models.This dataset consists of both naturally formed conversations, manually labeled data, and a large repository of candidate responses.Our preliminary experiments demonstrate that the simple retrieval-based conversation model performs reasonably well when combined with the rich instances in our dataset. Hao Wang 0076, Zhengdong Lu, Hang Li 0001, Enhong Chen |
EMNLP | 2 |
| 2013 | Collaborative boosting for activity classification in microblogsabstractUsers' daily activities, such as dining and shopping, inherently reflect their habits, intents and preferences, thus provide invaluable information for services such as personalized information recommendation and targeted advertising. Users' activity information, although ubiquitous on social media, has largely been unexploited. This paper addresses the task of user activity classification in microblogs, where users can publish short messages and maintain social networks online. We identify the importance of modeling a user's individuality, and that of exploiting opinions of the user's friends for accurate activity classification. In this light, we propose a novel collaborative boosting framework comprising a text-to-activity classifier for each user, and a mechanism for collaboration between classifiers of users having social connections. The collaboration between two classifiers includes exchanging their own training instances and their dynamically changing labeling decisions. We propose an iterative learning procedure that is formulated as gradient descent in learning function space, while opinion exchange between classifiers is implemented with a weighted voting in each learning iteration. We show through experiments that on real-world data from Sina Weibo, our method outperforms existing off-the-shelf algorithms that do not take users' individuality or social connections into account. Yangqiu Song, Zhengdong Lu, Cane Wing-ki Leung, Qiang Yang 0001 |
KDD | 2 |
| 2013 | A Deep Architecture for Matching Short TextsabstractMany machine learning problems can be interpreted as learning for matching two types of objects (e.g., images and captions, users and products, queries and documents). The matching level of two objects is usually measured as the inner product in a certain feature space, while the modeling effort focuses on mapping of objects from the original space to the feature space. This schema, although proven successful on a range of matching tasks, is insufficient for capturing the rich structure in the matching process of more complicated objects. In this paper, we propose a new deep architecture to more effectively model the complicated matching relations between two objects from heterogeneous domains. More specifically, we apply this model to matching tasks in natural language, e.g., finding sensible responses for a tweet, or relevant answers to a given question. This new architecture naturally combines the localness and hierarchy intrinsic to the natural language problems, and therefore greatly improves upon the state-of-the-art models. Zhengdong Lu, Hang Li 0001 |
NIPS | 1 |
| 2013 | Learning bilinear model for matching queries and documents
Wei Wu 0014, Zhengdong Lu, Hang Li 0001 |
J. Mach. Learn. Res. | 2 |
| 2012 | Finding web appearances of social network users via latent factor modelabstractWith the rapid growing of Web 2.0, people spend more time on social networks such as Facebook and Twitter. In order to know the people they are interacting with, finding the web appearances of them will help the social network users to a great extent. We propose a novel and effective latent factor model to find web appearances of target social network users. Our method solves the name ambiguity problem by simultaneously exploring the link structure of social networks and the web. Experiments on real-world data show the superiority of our method over several baselines. Kailong Chen, Zhengdong Lu, Xiaoshi Yin, Yong Yu 0001, Zaiqing Nie |
SIGIR | 2 |
| 2012 | Clustered embedding of massive social networksabstractThe explosive growth of social networks has created numerous exciting research opportunities. A central concept in the analysis of social networks is a proximity measure, which captures the closeness or similarity between nodes in the network. Despite much research on proximity measures, there is a lack of techniques to efficiently and accurately compute proximity measures for large-scale social networks. In this paper, we embed the original massive social graph into a much smaller graph, using a novel dimensionality reduction technique termed Clustered Spectral Graph Embedding. We show that the embedded graph captures the essential clustering and spectral structure of the original graph and allow a wide range of analysis to be performed on massive social graphs. Applying the clustered embedding to proximity measurement of social networks, we develop accurate, scalable, and flexible solutions to three important social network analysis tasks: proximity estimation, missing link inference, and link prediction. We demonstrate the effectiveness of our solutions to the tasks in the context of large real-world social network datasets: Flickr, LiveJournal, and MySpace with up to 2 million nodes and 90 million links. Han Hee Song, Berkant Savas, Tae Won Cho, Vacha Dave, Zhengdong Lu, Inderjit S. Dhillon, Yin Zhang 0001, Lili Qiu |
SIGMETRICS | 5 |
| 2011 | Manifold Learning and Missing Data Recovery through Unsupervised RegressionabstractWe propose an algorithm that, given a high-dimensional dataset with missing values, achieves the distinct goals of learning a nonlinear low-dimensional representation of the data (the dimensionality reduction problem) and reconstructing the missing high-dimensional data (the matrix completion, or imputation, problem). The algorithm follows the Dimensionality Reduction by Unsupervised Regression approach, where one alternately optimizes over the latent coordinates given the reconstruction and projection mappings, and vice versa, but here we also optimize over the missing data, using an efficient, globally convergent Gauss-Newton scheme. We also show how to project or reconstruct test data with missing values. We achieve impressive reconstructions while learning good latent representations in image restoration with 50% missing pixels. Miguel Á. Carreira-Perpiñán, Zhengdong Lu |
ICDM | 2 |
| 2011 | A Denoising View of Matrix CompletionabstractIn matrix completion, we are given a matrix where the values of only some of the entries are present, and we want to reconstruct the missing ones. Much work has focused on the assumption that the data matrix has low rank. We propose a more general assumption based on denoising, so that we expect that the value of a missing entry can be predicted from the values of neighboring points. We propose a nonparametric version of denoising based on local, iterated averaging with mean-shift, possibly constrained to preserve local low-rank manifold structure. The few user parameters required (the denoising scale, number of neighbors and local dimensionality) and the number of iterations can be estimated by cross-validating the reconstruction error. Using our algorithms as a postprocessing step on an initial reconstruction (provided by e.g. a low-rank method), we show consistent improvements with synthetic, image and motion-capture data. Miguel Á. Carreira-Perpiñán, Zhengdong Lu |
NIPS | 3 |
| 2011 | Kernels for Longitudinal Data with Variable Sequence Length and Sampling IntervalsabstractWe develop several kernel methods for classification of longitudinal data and apply them to detect cognitive decline in the elderly. We first develop mixed-effects models, a type of hierarchical empirical Bayes generative models, for the time series. After demonstrating their utility in likelihood ratio classifiers (and the improvement over standard regression models for such classifiers), we develop novel Fisher kernels based on mixture of mixed-effects models and use them in support vector machine classifiers. The hierarchical generative model allows us to handle variations in sequence length and sampling interval gracefully. We also give nonparametric kernels not based on generative models, but rather on the reproducing kernel Hilbert space. We apply the methods to detecting cognitive decline from longitudinal clinical data on motor and neuropsychological tests. The likelihood ratio classifiers based on the neuropsychological tests perform better than than classifiers based on the motor behavior. Discriminant classifiers performed better than likelihood ratio classifiers for the motor behavior tests. Zhengdong Lu, Todd K. Leen, Jeffrey A. Kaye |
Neural Comput. | 1 |
| 2011 | Scalable Affiliation Recommendation using Auxiliary NetworksabstractSocial network analysis has attracted increasing attention in recent years. In many social networks, besides friendship links among users, the phenomenon of users associating themselves with groups or communities is common. Thus, two networks exist simultaneously: the friendship network among users, and the affiliation network between users and groups. In this article, we tackle the affiliation recommendation problem, where the task is to predict or suggest new affiliations between users and communities, given the current state of the friendship and affiliation networks. More generally, affiliations need not be community affiliations---they can be a user’s taste, so affiliation recommendation algorithms have applications beyond community recommendation. In this article, we show that information from the friendship network can indeed be fruitfully exploited in making affiliation recommendations. Using a simple way of combining these networks, we suggest two models of user-community affinity for the purpose of making affiliation recommendations: one based on graph proximity, and another using latent factors to model users and communities. We explore the affiliation recommendation algorithms suggested by these models and evaluate these algorithms on two real-world networks, Orkut and Youtube. In doing so, we motivate and propose a way of evaluating recommenders, by measuring how good the top 50 recommendations are for the average user, and demonstrate the importance of choosing the right evaluation strategy. The algorithms suggested by the graph proximity model turn out to be the most effective. We also introduce scalable versions of these algorithms, and demonstrate their effectiveness. This use of link prediction techniques for the purpose of affiliation recommendation is, to our knowledge, novel. Vishvas Vasuki, Nagarajan Natarajan, Zhengdong Lu, Berkant Savas, Inderjit S. Dhillon |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2010 | Parametric dimensionality reduction by unsupervised regressionabstractWe introduce a parametric version (pDRUR) of the recently proposed Dimensionality Reduction by Unsupervised Regression algorithm. pDRUR alternately minimizes reconstruction error by fitting parametric functions given latent coordinates and data, and by updating latent coordinates given functions (with a Gauss-Newton method decoupled over coordinates). Both the fit and the update become much faster while attaining results of similar quality, and afford dealing with far larger datasets (105points). We show in a number of benchmarks how the algorithm efficiently learns good latent coordinates and bidirectional mappings between the data and latent space, even with very noisy or low-quality initializations, often drastically improving the result of spectral and other methods. Miguel Á. Carreira-Perpiñán, Zhengdong Lu |
CVPR | 2 |
| 2010 | Supervised Link Prediction Using Multiple SourcesabstractLink prediction is a fundamental problem in social network analysis and modern-day commercial applications such as Face book and My space. Most existing research approaches this problem by exploring the topological structure of a social network using only one source of information. However, in many application domains, in addition to the social network of interest, there are a number of auxiliary social networks and/or derived proximity networks available. The contribution of the paper is twofold: (1) a supervised learning framework that can effectively and efficiently learn the dynamics of social networks in the presence of auxiliary networks, (2) a feature design scheme for constructing a rich variety of path-based features using multiple sources, and an effective feature selection strategy based on structured sparsity. Extensive experiments on three real-world collaboration networks show that our model can effectively learn to predict new links using multiple sources, yielding higher prediction accuracy than unsupervised and single-source supervised models. Zhengdong Lu, Berkant Savas, Inderjit S. Dhillon |
ICDM | 1 |
| 2010 | Affiliation recommendation using auxiliary networksabstractSocial network analysis has attracted increasing attention in recent years. In many social networks, besides friendship links amongst users, the phenomenon of users associating themselves with groups or communities is common. Thus, two networks exist simultaneously: the friendship network among users, and the affiliation network between users and groups. In this paper, we tackle the affiliation recommendation problem, where the task is to predict or suggest new affiliations between users and communities, given the current state of the friendship and affiliation networks. More generally, affiliations need not be community affiliations - they can be a user's taste, so affiliation recommendation algorithms have applications beyond community recommendation. In this paper, we show that information from the friendship network can indeed be fruitfully exploited in making affiliation recommendations. Using a simple way of combining these networks, we suggest two models of user-community affinity for the purpose of making affiliation recommendations: one based on graph proximity, and another using latent factors to model users and communities. We explore the two classes of affiliation recommendation algorithms suggested by these models. We evaluate these algorithms on two real world networks - Orkut and Youtube. In doing so, we motivate and propose a way of evaluating recommenders, by measuring how good the top 50 recommendations are for the average user, and demonstrate the importance of choosing the right evaluation strategy. The algorithms suggested by the graph proximity model turn out to be the most effective and efficient. This use of link prediction techniques for the purpose of affiliation recommendation is, to our knowledge, novel. Vishvas Vasuki, Nagarajan Natarajan, Zhengdong Lu, Inderjit S. Dhillon |
RecSys | 3 |
| 2009 | Clustering with Multiple GraphsabstractIn graph-based learning models, entities are often represented as vertices in an undirected graph with weighted edges describing the relationships between entities. In many real-world applications, however, entities are often associated with relations of different types and/or from different sources, which can be well captured by multiple undirected graphs over the same set of vertices. How to exploit such multiple sources of information to make better inferences on entities remains an interesting open problem. In this paper, we focus on the problem of clustering the vertices based on multiple graphs in both unsupervised and semi-supervised settings. As one of our contributions, we propose Linked Matrix Factorization (LMF) as a novel way of fusing information from multiple graph sources. In LMF, each graph is approximated by matrix factorization with a graph-specific factor and a factor common to all graphs, where the common factor provides features for all vertices. Experiments on SIAM journal data show that (1) we can improve the clustering accuracy through fusing multiple sources of information with several models, and (2) LMF yields superior or competitive results compared to other graph-based clustering methods. Zhengdong Lu, Inderjit S. Dhillon |
ICDM | 2 |
| 2009 | Geometry-aware metric learningabstractIn this paper, we introduce a generic framework for semi-supervised kernel learning. Given pair-wise (dis-)similarity constraints, we learn a kernel matrix over the data that respects the provided side-information as well as the local geometry of the data. Our framework is based on metric learning methods, where we jointly model the metric/kernel over the data along with the underlying manifold. Furthermore, we show that for some important parameterized forms of the underlying manifold model, we can estimate the model parameters and the kernel matrix efficiently. Our resulting algorithm is able to incorporate local geometry into the metric learning task; at the same time it can handle a wide class of constraints. Finally, our algorithm is fast and scalable -- unlike most of the existing methods, it is able to exploit the low dimensional manifold structure and does not require semi-definite programming. We demonstrate wide applicability and effectiveness of our framework by applying to various machine learning tasks such as semi-supervised classification, colored dimensionality reduction, manifold alignment etc. On each of the tasks our method performs competitively or better than the respective state-of-the-art method. Zhengdong Lu, Prateek Jain 0002, Inderjit S. Dhillon |
ICML | 1 |
| 2009 | A spatio-temporal approach to collaborative filteringabstractIn this paper, we propose a novel spatio-temporal model for collaborative filtering applications. Our model is based on low-rank matrix factorization that uses a spatio-temporal filtering approach to estimate user and item factors. The spatial component regularizes the factors by exploiting correlation across users and/or items, modeled as a function of some implicit feedback (e.g., who rated what) and/or some side information (e.g., user demographics, browsing history). In particular, we incorporate correlation in factors through a Markov random field prior in a probabilistic framework, whereby the neighborhood weights are functions of user and item covariates. The temporal component ensures that the user/item factors adapt to process changes that occur through time and is implemented in a state space framework with fast estimation through Kalman filtering. Our spatio-temporal filtering (ST-KF hereafter) approach provides a single joint model to simultaneously incorporate both spatial and temporal structure in ratings and therefore provides an accurate method to predict future ratings. To ensure scalability of ST-KF, we employ a mean-field approximation for inference. Incorporating user/item covariates in estimating neighborhood weights also helps in dealing with both cold-start and warm-start problems seamlessly in a single unified modeling framework; covariates predict factors for new users and items through the neighborhood. We illustrate our method on simulated data, benchmark data and data obtained from a relatively new recommender system application arising in the context of Yahoo! Front Page. Zhengdong Lu, Deepak Agarwal, Inderjit S. Dhillon |
RecSys | 1 |
| 2008 | Dimensionality reduction by unsupervised regressionabstractWe consider the problem of dimensionality reduction, where given high-dimensional data we want to estimate two mappings: from high to low dimension (dimensionality reduction) and from low to high dimension (reconstruction). We adopt an unsupervised regression point of view by introducing the unknown low-dimensional coordinates of the data as parameters, and formulate a regularised objective functional of the mappings and low-dimensional coordinates. Alternating minimisation of this functional is straightforward: for fixed low-dimensional coordinates, the mappings have a unique solution; and for fixed mappings, the coordinates can be obtained by finite-dimensional non-linear minimisation. Besides, the coordinates can be initialised to the output of a spectral method such as Laplacian eigenmaps. The model generalises PCA and several recent methods that learn one of the two mappings but not both; and, unlike spectral methods, our model provides out-of-sample mappings by construction. Experiments with toy and real-world problems show that the model is able to learn mappings for convoluted manifolds, avoiding bad local optima that plague other methods. Miguel Á. Carreira-Perpiñán, Zhengdong Lu |
CVPR | 2 |
| 2008 | Constrained spectral clustering through affinity propagationabstractPairwise constraints specify whether or not two samples should be in one cluster. Although it has been successful to incorporate them into traditional clustering methods, such as K-means, little progress has been made in combining them with spectral clustering. The major challenge in designing an effective constrained spectral clustering is a sensible combination of the scarce pairwise constraints with the original affinity matrix. We propose to combine the two sources of affinity by propagating the pairwise constraints information over the original affinity matrix. Our method has a Gaussian process interpretation and results in a closed-form expression for the new affinity matrix. Experiments show it outperforms state-of-the-art constrained clustering methods in getting good clusterings with fewer constraints, and yields good image segmentation with user-specified pairwise constraints. Zhengdong Lu, Miguel Á. Carreira-Perpiñán |
CVPR | 1 |
| 2008 | Detecting mild cognitive loss with continuous monitoring of medication adherenceabstractThis paper describes an approach for detecting early cognitive loss using medication adherence behavior. We investigate the discriminative power of a comprehensive set of recurrent medication timing features extracted from time-of-day and inter-dose timing statistics. We adopt information theoretic measures for feature ranking for initial dimensionality reduction and conduct exhaustive leave-one-out cross validation for final feature selection and regularization. The selected feature set is subjected to a support vector machine for classification. The results demonstrate that patterns of adherence based on the data from relatively unobtrusive behavior monitoring can make reliable inference for mild cognitive loss individuals. Yonghong Huang, Deniz Erdogmus, Zhengdong Lu, Todd K. Leen |
ICASSP | 3 |
| 2008 | A reproducing kernel Hilbert space framework for pairwise time series distancesabstractA good distance measure for time series needs to properly incorporate the temporal structure, and should be applicable to sequences with unequal lengths. In this paper, we propose a distance measure as a principled solution to the two requirements. Unlike the conventional feature vector representation, our approach represents each time series with a summarizing smooth curve in a reproducing kernel Hilbert space (RKHS), and therefore translate the distance between time series into distances between curves. Moreover we propose to learn the kernel of this RKHS from a population of time series with discrete observations using Gaussian process-based non-parametric mixed-effect models. Experiments on two vastly different real-world problems show that the proposed distance measure leads to improved classification accuracy over the conventional distance measures. Zhengdong Lu, Todd K. Leen, Yonghong Huang, Deniz Erdogmus |
ICML | 1 |
| 2008 | Hierarchical Fisher Kernels for Longitudinal DataabstractWe develop new techniques for time series classification based on hierarchical Bayesian generative models (called mixed-effect models) and the Fisher kernel derived from them. A key advantage of the new formulation is that one can compute the Fisher information matrix despite varying sequence lengths and sampling times. We therefore can avoid the ad hoc replacement of Fisher information matrix with the identity matrix commonly used in literature, which destroys the geometrical grounding of the kernel construction. In contrast, our construction retains the proper geometric structure resulting in a kernel that is properly invariant under change of coordinates in the model parameter space. Experiments on detecting cognitive decline show that classifiers based on the proposed kernel out-perform those based on generative models and other feature extraction routines. Zhengdong Lu, Todd K. Leen, Jeffrey A. Kaye |
NIPS | 1 |
| 2007 | People Tracking with the Laplacian Eigenmaps Latent Variable ModelabstractReliably recovering 3D human pose from monocular video requires constraints that bias the estimates towards typical human poses and motions. We define priors for people tracking using a Laplacian Eigenmaps Latent Variable Model (LELVM). LELVM is a probabilistic dimensionality reduction model that naturally combines the advantages of latent variable models---definining a multimodal probability density for latent and observed variables, and globally differentiable nonlinear mappings for reconstruction and dimensionality reduction---with those of spectral manifold learning methods---no local optima, ability to unfold highly nonlinear manifolds, and good practical scaling to latent spaces of high dimension. LELVM is computationally efficient, simple to learn from sparse training data, and compatible with standard probabilistic trackers such as particle filters. We analyze the performance of a LELVM-based probabilistic sigma point mixture tracker in several real and synthetic human motion sequences and demonstrate that LELVM provides sufficient constraints for robust operation in the presence of missing, noisy and ambiguous image measurements. Zhengdong Lu, Miguel Á. Carreira-Perpiñán, Cristian Sminchisescu |
NIPS | 1 |
| 2007 | Penalized Probabilistic ClusteringabstractWhile clustering is usually an unsupervised operation, there are circumstances in which we believe (with varying degrees of certainty) that items A and B should be assigned to the same cluster, while items A and C should not. We would like such pairwise relations to influence cluster assignments of out-of-sample data in a manner consistent with the prior knowledge expressed in the training set. Our starting point is probabilistic clustering based on gaussian mixture models (GMM) of the data distribution. We express clustering preferences in a prior distribution over assignments of data points to clusters. This prior penalizes cluster assignments according to the degree with which they violate the preferences. The model parameters are fit with the expectation-maximization (EM) algorithm. Our model provides a flexible framework that encompasses several other semisupervised clustering models as its special cases. Experiments on artificial and real-world problems show that our model can consistently improve clustering results when pairwise relations are incorporated. The experiments also demonstrate the superiority of our model to other semisupervised clustering methods on handling noisy pairwise relations. Zhengdong Lu, Todd K. Leen |
Neural Comput. | 1 |
| 2007 | Fast neural network surrogates for very high dimensional physics-based models in computational oceanography
Rudolph van der Merwe, Todd K. Leen, Zhengdong Lu, Sergey Frolov, António M. Baptista |
Neural Networks | 3 |
| 2007 | Erratum to "Fast neural network surrogates for very high dimensional physics-based models in computational oceanography" [Neural Netw. 20(4) (2007) 462-478]
Rudolph van der Merwe, Todd K. Leen, Zhengdong Lu, Sergey Frolov, António M. Baptista |
Neural Networks | 3 |
| 2004 | Semi-supervised Learning with Penalized Probabilistic ClusteringabstractWhile clustering is usually an unsupervised operation, there are circum- stances in which we believe (with varying degrees of certainty) that items A and B should be assigned to the same cluster, while items A and C should not. We would like such pairwise relations to influence cluster assignments of out-of-sample data in a manner consistent with the prior knowledge expressed in the training set. Our starting point is proba- bilistic clustering based on Gaussian mixture models (GMM) of the data distribution. We express clustering preferences in the prior distribution over assignments of data points to clusters. This prior penalizes cluster assignments according to the degree with which they violate the prefer- ences. We fit the model parameters with EM. Experiments on a variety of data sets show that PPC can consistently improve clustering results. Zhengdong Lu, Todd K. Leen |
NIPS | 1 |