VLDB 2026 Research / reviewers in the wild / expert
Jianshu Chen
dblp:11/3124
· DBLP profile ↗
54ranked-venue papers
9as first author
12since 2021 · last 2025
0000-0001-8216-2756ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 39 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-authorComputer networks · 4Theory of computation · 3 · 2 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Exploring Reasoning-Infused Text Embedding with Large Language Models for Zero-Shot Dense RetrievalabstractTransformer-based models such as BERT and E5 have significantly advanced text embedding by capturing rich contextual representations. However, many complex real-world queries require sophisticated reasoning to retrieve relevant documents beyond surface-level lexical matching, where encoder-only retrievers often fall short. Decoder-only large language models (LLMs), known for their strong reasoning capabilities, offer a promising alternative. Despite this potential, existing LLM-based embedding methods primarily focus on contextual representation and do not fully exploit the reasoning strength of LLMs. To bridge this gap, we propose Reasoning-Infused Text Embedding (RITE), a simple but effective approach that integrates logical reasoning into the text embedding process using generative LLMs. RITE builds upon existing language model embedding techniques by generating intermediate reasoning texts in the token space before computing embeddings, thereby enriching representations with inferential depth. Experimental results on BRIGHT, a reasoning-intensive retrieval benchmark, demonstrate that RITE significantly enhances zero-shot retrieval performance across diverse domains, underscoring the effectiveness of incorporating reasoning into the embedding process. Gourab Kundu, Tianyu Cao 0001, Guang Cheng 0003, Zhen Ge, Jianshu Chen, Qingjun Cui, Trishul Chilimbi |
CIKM | 7 |
| 2024 | Rewards-in-Context: Multi-objective Alignment of Foundation Models with Dynamic Preference AdjustmentabstractWe consider the problem of multi-objective alignment of foundation models with human preferences, which is a critical step towards helpful and harmless AI systems. However, it is generally costly and unstable to fine-tune large foundation models using reinforcement learning (RL), and the multi-dimensionality, heterogeneity, and conflicting nature of human preferences further complicate the alignment process. In this paper, we introduce Rewards-in-Context (RiC), which conditions the response of a foundation model on multiple rewards in its prompt context and applies supervised fine-tuning for alignment. The salient features of RiC are simplicity and adaptivity, as it only requires supervised fine-tuning of a single foundation model and supports dynamic adjustment for user preferences during inference time. Inspired by the analytical solution of an abstracted convex optimization problem, our dynamic inference-time adjustment method approaches the Pareto-optimal solution for multiple objectives. Empirical evidence demonstrates the efficacy of our method in aligning both Large Language Models (LLMs) and diffusion models to accommodate diverse rewards with only around 10% GPU hours compared with multi-objective RL baseline. Rui Yang 0010, Xiaoman Pan, Feng Luo 0003, Han Zhong 0001, Dong Yu 0001, Jianshu Chen |
ICML | 7 |
| 2024 | MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction TuningabstractFuxiao Liu, Xiaoyang Wang, Wenlin Yao, Jianshu Chen, Kaiqiang Song, Sangwoo Cho, Yaser Yacoob, Dong Yu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Fuxiao Liu, Xiaoyang Wang 0001, Wenlin Yao, Jianshu Chen, Kaiqiang Song, Sangwoo Cho, Yaser Yacoob, Dong Yu 0001 |
NAACL-HLT | 4 |
| 2024 | From Language Modeling to Instruction Following: Understanding the Behavior Shift in LLMs after Instruction TuningabstractXuansheng Wu, Wenlin Yao, Jianshu Chen, Xiaoman Pan, Xiaoyang Wang, Ninghao Liu, Dong Yu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Xuansheng Wu, Wenlin Yao, Jianshu Chen, Xiaoman Pan, Xiaoyang Wang 0001, Ninghao Liu 0001, Dong Yu 0001 |
NAACL-HLT | 3 |
| 2023 | Blessing of Class Diversity in Pre-trainingabstractThis paper presents a new statistical analysis aiming to explain the recent superior achievements of the pre-training techniques in natural language processing (NLP). We prove that when the classes of the pre-training task (e.g., different words in the masked language model task) are sufficiently diverse, in the sense that the least singular value of the last linear layer in pre-training (denoted as $\tilde{\nu}$) is large, then pre-training can significantly improve the sample efficiency of downstream tasks. Specially, we show the transfer learning excess risk enjoys an $O\left(\frac{1}{\tilde{\nu} \sqrt{n}}\right)$ rate, in contrast to the $O\left(\frac{1}{\sqrt{m}}\right)$ rate in the standard supervised learning. Here, $n$ is the number of pre-training data and $m$ is the number of data in the downstream task, and typically $n \gg m$. Our proof relies on a vector-form Rademacher complexity chain rule for disassembling composite function classes and a modified self-concordance condition. These techniques can be of independent interest. Yulai Zhao 0002, Jianshu Chen, Simon S. Du |
AISTATS | 2 |
| 2023 | How do Words Contribute to Sentence Semantics? Revisiting Sentence Embeddings with a Perturbation MethodabstractWenlin Yao, Lifeng Jin, Hongming Zhang, Xiaoman Pan, Kaiqiang Song, Dian Yu, Dong Yu, Jianshu Chen. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Wenlin Yao, Lifeng Jin, Hongming Zhang 0009, Xiaoman Pan, Kaiqiang Song, Dian Yu 0001, Dong Yu 0001, Jianshu Chen |
EACL | 8 |
| 2023 | Learning Language Representations with Logical Inductive Bias
Jianshu Chen |
ICLR | 1 |
| 2023 | Knowledge-in-Context: Towards Knowledgeable Semi-Parametric Language Models
Xiaoman Pan, Wenlin Yao, Hongming Zhang 0009, Dian Yu 0001, Dong Yu 0001, Jianshu Chen |
ICLR | 6 |
| 2023 | Thrust: Adaptively Propels Large Language Models with External KnowledgeabstractAlthough large-scale pre-trained language models (PTLMs) are shown to encode rich knowledge in their model parameters, the inherent knowledge in PTLMs can be opaque or static, making external knowledge necessary. However, the existing information retrieval techniques could be costly and may even introduce noisy and sometimes misleading knowledge. To address these challenges, we propose the instance-level adaptive propulsion of external knowledge (IAPEK), where we only conduct the retrieval when necessary. To achieve this goal, we propose to model whether a PTLM contains enough knowledge to solve an instance with a novel metric, Thrust, which leverages the representation distribution of a small amount of seen instances. Extensive experiments demonstrate that Thrust is a good measurement of models' instance-level knowledgeability. Moreover, we can achieve higher cost-efficiency with the Thrust score as the retrieval indicator than the naive usage of external knowledge on 88% of the evaluated tasks with 26% average performance improvement. Such findings shed light on the real-world practice of knowledge-enhanced LMs with a limited budget for knowledge seeking due to computation latency or costs. Hongming Zhang 0009, Xiaoman Pan, Wenlin Yao, Dong Yu 0001, Jianshu Chen |
NeurIPS | 6 |
| 2022 | Improving Machine Reading Comprehension with Contextualized Commonsense KnowledgeabstractTo perform well on a machine reading comprehension (MRC) task, machine readers usually require commonsense knowledge that is not explicitly mentioned in the given documents.This paper aims to extract a new kind of structured knowledge from scripts and use it to improve MRC.We focus on scripts as they contain rich verbal and nonverbal messages, and two relevant messages originally conveyed by different modalities during a short time period may serve as arguments of a piece of commonsense knowledge as they function together in daily communications.To save human efforts to name relations, we propose to represent relations implicitly by situating such an argument pair in a context and call it contextualized knowledge.To use the extracted knowledge to improve MRC, we compare several fine-tuning strategies to use the weakly-labeled MRC data constructed based on contextualized knowledge and further design a teacher-student paradigm with multiple teachers to facilitate the transfer of knowledge in weakly-labeled MRC data.Experimental results show that our paradigm outperforms other methods that use weaklylabeled data and improves a state-of-the-art baseline by 4.3% in accuracy on a Chinese multiple-choice MRC dataset C 3 , wherein most of the questions require unstated prior knowledge.We also seek to transfer the knowledge to other tasks by simply adapting the resulting student reader, yielding a 2.9% improvement in F1 on a relation extraction dataset DialogRE, demonstrating the potential usefulness of the knowledge for non-MRC tasks that require document comprehension.Interior.Runaway office.Day. Kai Sun 0006, Dian Yu 0001, Jianshu Chen, Dong Yu 0001, Claire Cardie |
ACL (1) | 3 |
| 2022 | Z-LaVI: Zero-Shot Language Solver Fueled by Visual ImaginationabstractLarge-scale pretrained language models have made significant advances in solving downstream language understanding tasks.However, they generally suffer from reporting bias, the phenomenon describing the lack of explicit commonsense knowledge in written text, e.g., "an orange is orange".To overcome this limitation, we develop a novel approach, Z-LaVI, to endow language models with visual imagination capabilities.Specifically, we leverage two complementary types of "imaginations": (i) recalling existing images through retrieval and (ii) synthesizing nonexistent images via text-toimage generation.Jointly exploiting the language inputs and the imagination, a pretrained vision-language model (e.g., CLIP) eventually composes a zero-shot solution to the original language tasks.Notably, fueling language models with imagination can effectively leverage visual knowledge to solve plain language tasks.In consequence, Z-LaVI consistently improves the zero-shot performance of existing language models across a diverse set of language tasks. 1 * Work was done during the internship at Tencent AI Lab.(a) Word Sense Disambiguation (b) Science Question Answering (c) Topic Classification sense1: bank(institute) sense2: bank(geography) Input: The species prefers the {bank} of pond. Wenlin Yao, Hongming Zhang 0009, Xiaoyang Wang 0001, Dong Yu 0001, Jianshu Chen |
EMNLP | 6 |
| 2021 | Connect-the-Dots: Bridging Semantics between Words and Definitions via Aligning Word Sense InventoriesabstractWord Sense Disambiguation (WSD) aims to automatically identify the exact meaning of one word according to its context.Existing supervised models struggle to make correct predictions on rare word senses due to limited training data and can only select the best definition sentence from one predefined word sense inventory (e.g., WordNet).To address the data sparsity problem and generalize the model to be independent of one predefined inventory, we propose a gloss alignment algorithm that can align definition sentences (glosses) with the same meaning from different sense inventories to collect rich lexical knowledge.We then train a model to identify semantic equivalence between a target word in context and one of its glosses using these aligned inventories, which exhibits strong transfer capability to many WSD tasks 1 .Experiments on benchmark datasets show that the proposed method improves predictions on both frequent and rare word senses, outperforming prior work by 1.2% on the All-Words WSD Task and 4.3% on the Low-Shot WSD Task.Evaluation on WiC Task also indicates that our method can better capture word meanings in context. Wenlin Yao, Xiaoman Pan, Lifeng Jin, Jianshu Chen, Dian Yu 0001, Dong Yu 0001 |
EMNLP (1) | 4 |
| 2020 | Logical Natural Language Generation from Open-Domain TablesabstractNeural natural language generation (NLG) models have recently shown remarkable progress in fluency and coherence.However, existing studies on neural NLG are primarily focused on surface-level realizations with limited emphasis on logical inference, an important aspect of human thinking and language.In this paper, we suggest a new NLG task where a model is tasked with generating natural language statements that can be logically entailed by the facts in an open-domain semi-structured table.To facilitate the study of the proposed logical NLG problem, we use the existing Tab-Fact dataset (Chen et al., 2019) featured with a wide range of logical/symbolic inferences as our testbed, and propose new automatic metrics to evaluate the fidelity of generation models w.r.t.logical inference.The new task poses challenges to the existing monotonic generation frameworks due to the mismatch between sequence order and logical order.In our experiments, we comprehensively survey different generation architectures (LSTM, Transformer, Pre-Trained LM) trained with different algorithms (RL, Adversarial Training, Coarse-to-Fine) on the dataset and made following observations: 1) Pre-Trained LM can significantly boost both the fluency and logical fidelity metrics, 2) RL and Adversarial Training are trading fluency for fidelity, 3) Coarse-to-Fine generation can help partially alleviate the fidelity issue while maintaining high language fluency. Wenhu Chen, Jianshu Chen, Yu Su 0001, Zhiyu Chen 0002, William Yang Wang |
ACL | 2 |
| 2020 | Recurrent Chunking Mechanisms for Long-Text Machine Reading ComprehensionabstractIn this paper, we study machine reading comprehension (MRC) on long texts, where a model takes as inputs a lengthy document and a question and then extracts a text span from the document as an answer.State-of-the-art models tend to use a pretrained transformer model (e.g., BERT) to encode the joint contextual information of document and question.However, these transformer-based models can only take a fixed-length (e.g., 512) text as its input.To deal with even longer text inputs, previous approaches usually chunk them into equally-spaced segments and predict answers based on each segment independently without considering the information from other segments.As a result, they may form segments that fail to cover the correct answer span or retain insufficient contexts around it, which significantly degrades the performance.Moreover, they are less capable of answering questions that need cross-segment information.We propose to let a model learn to chunk in a more flexible way via reinforcement learning: a model can decide the next segment that it wants to process in either direction.We also employ recurrent mechanisms to enable information to flow across segments.Experiments on three MRC datasets -CoQA, QuAC, and TriviaQA -demonstrate the effectiveness of our proposed recurrent chunking mechanisms: we can obtain segments that are more likely to contain complete answers and at the same time provide sufficient contexts around the ground truth answers for better predictions. Hongyu Gong, Yelong Shen, Dian Yu 0001, Jianshu Chen, Dong Yu 0001 |
ACL | 4 |
| 2020 | ZPR2: Joint Zero Pronoun Recovery and Resolution using Multi-Task Learning and BERTabstractZero pronoun recovery and resolution aim at recovering the dropped pronoun and pointing out its anaphoric mentions, respectively.We propose to better explore their interaction by solving both tasks together, while the previous work treats them separately.For zero pronoun resolution, we study this task in a more realistic setting, where no parsing trees or only automatic trees are available, while most previous work assumes gold trees.Experiments on two benchmarks show that joint modeling significantly outperforms our baseline that already beats the previous state of the arts. Linfeng Song, Kun Xu 0005, Yue Zhang 0004, Jianshu Chen, Dong Yu 0001 |
ACL | 4 |
| 2020 | Comprehensive Image Captioning via Scene Graph Decomposition
Yiwu Zhong, Liwei Wang 0009, Jianshu Chen, Dong Yu 0001, Yin Li 0003 |
ECCV (14) | 3 |
| 2020 | TabFact: A Large-scale Dataset for Table-based Fact Verification
Wenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang 0002, Hong Wang 0023, Xiyou Zhou, William Yang Wang |
ICLR | 3 |
| 2020 | Watch the Unobserved: A Simple Approach to Parallelizing Monte Carlo Tree Search
Anji Liu, Jianshu Chen, Mingze Yu, Xuewen Zhou |
ICLR | 2 |
| 2019 | Incorporating Structured Commonsense Knowledge in Story CompletionabstractThe ability to select an appropriate story ending is the first step towards perfect narrative comprehension. Story ending prediction requires not only the explicit clues within the context, but also the implicit knowledge (such as commonsense) to construct a reasonable and consistent story. However, most previous approaches do not explicitly use background commonsense knowledge. We present a neural story ending selection model that integrates three types of information: narrative sequence, sentiment evolution and commonsense knowledge. Experiments show that our model outperforms state-ofthe-art approaches on a public dataset, ROCStory Cloze Task (Mostafazadeh et al. 2017), and the performance gain from adding the additional commonsense knowledge is significant. Jiaao Chen, Jianshu Chen |
AAAI | 2 |
| 2019 | Semantically Conditioned Dialog Response Generation via Hierarchical Disentangled Self-AttentionabstractSemantically controlled neural response generation on limited-domain has achieved great performance.However, moving towards multi-domain large-scale scenarios are shown to be difficult because the possible combinations of semantic inputs grow exponentially with the number of domains.To alleviate such scalability issue, we exploit the structure of dialog acts to build a multi-layer hierarchical graph, where each act is represented as a rootto-leaf route on the graph.Then, we incorporate such graph structure prior as an inductive bias to build a hierarchical disentangled self-attention network, where we disentangle attention heads to model designated nodes on the dialog act graph.By activating different (disentangled) heads at each layer, combinatorially many dialog act semantics can be modeled to control the neural response generation.On the large-scale Multi-Domain-WOZ dataset, our model can yield a significant improvement over the baselines on various automatic and human evaluation metrics. Wenhu Chen, Jianshu Chen, Pengda Qin, Xifeng Yan, William Yang Wang |
ACL (1) | 2 |
| 2019 | Improving Pre-Trained Multilingual Model with Vocabulary ExpansionabstractRecently, pre-trained language models have achieved remarkable success in a broad range of natural language processing tasks.However, in multilingual setting, it is extremely resource-consuming to pre-train a deep language model over large-scale corpora for each language.Instead of exhaustively pre-training monolingual language models independently, an alternative solution is to pre-train a powerful multilingual deep language model over large-scale corpora in hundreds of languages.However, the vocabulary size for each language in such a model is relatively small, especially for low-resource languages.This limitation inevitably hinders the performance of these multilingual models on tasks such as sequence labeling, wherein in-depth token-level or sentence-level understanding is essential.In this paper, inspired by previous methods designed for monolingual settings, we investigate two approaches (i.e., joint mapping and mixture mapping) based on a pre-trained multilingual model BERT for addressing the out-of-vocabulary (OOV) problem on a variety of tasks, including part-of-speech tagging, named entity recognition, machine translation quality estimation, and machine reading comprehension.Experimental results show that using mixture mapping is more promising.To the best of our knowledge, this is the first work that attempts to address and discuss the OOV issue in multilingual settings. Hai Wang 0013, Dian Yu 0001, Kai Sun 0006, Jianshu Chen, Dong Yu 0001 |
CoNLL | 4 |
| 2019 | Evidence Sentence Extraction for Machine Reading ComprehensionabstractRemarkable success has been achieved in the last few years on some limited machine reading comprehension (MRC) tasks.However, it is still difficult to interpret the predictions of existing MRC models.In this paper, we focus on extracting evidence sentences that can explain or support the answers of multiplechoice MRC tasks, where the majority of answer options cannot be directly extracted from reference documents.Due to the lack of ground truth evidence sentence labels in most cases, we apply distant supervision to generate imperfect labels and then use them to train an evidence sentence extractor.To denoise the noisy labels, we apply a recently proposed deep probabilistic logic learning framework to incorporate both sentence-level and cross-sentence linguistic indicators for indirect supervision.We feed the extracted evidence sentences into existing MRC models and evaluate the end-to-end performance on three challenging multiplechoice MRC datasets: MultiRC, RACE, and DREAM, achieving comparable or better performance than the same models that take as input the full reference document.To the best of our knowledge, this is the first work extracting evidence sentences for multiple-choice MRC. Hai Wang 0013, Dian Yu 0001, Kai Sun 0006, Jianshu Chen, Dong Yu 0001, David A. McAllester, Dan Roth 0001 |
CoNLL | 4 |
| 2019 | Unsupervised Speech Recognition via Segmental Empirical Output Distribution Matching
Chih-Kuan Yeh, Jianshu Chen, Chengzhu Yu, Dong Yu 0001 |
ICLR (Poster) | 2 |
| 2019 | Stochastic Variance Reduced Primal Dual Algorithms for Empirical Composition OptimizationabstractWe consider a generic empirical composition optimization problem, where there are empirical averages present both outside and inside nonlinear loss functions. Such a problem is of interest in various machine learning applications, and cannot be directly solved by standard methods such as stochastic gradient descent (SGD). We take a novel approach to solving this problem by reformulating the original minimization objective into an equivalent min-max objective, which brings out all the empirical averages that are originally inside the nonlinear loss functions. We exploit the rich structures of the reformulated problem and develop a stochastic primal-dual algorithms, SVRPDA-I, to solve the problem efficiently. We carry out extensive theoretical analysis of the proposed algorithm, obtaining the convergence rate, the total computation complexity and the storage complexity. In particular, the algorithm is shown to converge at a linear rate when the problem is strongly convex. Moreover, we also develop an approximate version of the algorithm, named SVRPDA-II, which further reduces the memory requirement. Finally, we evaluate the performance of our algorithms on several real-world benchmarks and experimental results show that they significantly outperform existing techniques. Adithya M. Devraj, Jianshu Chen |
NeurIPS | 2 |
| 2019 | DREAM: A Challenge Dataset and Models for Dialogue-Based Reading ComprehensionabstractWe present DREAM, the first dialogue-based multiple-choice reading comprehension data set. Collected from English as a Foreign Language examinations designed by human experts to evaluate the comprehension level of Chinese learners of English, our data set contains 10,197 multiple-choice questions for 6,444 dialogues. In contrast to existing reading comprehension data sets, DREAM is the first to focus on in-depth multi-turn multi-party dialogue understanding. DREAM is likely to present significant challenges for existing reading comprehension systems: 84% of answers are non-extractive, 85% of questions require reasoning beyond a single sentence, and 34% of questions also involve commonsense knowledge. We apply several popular neural reading comprehension models that primarily exploit surface information within the text and find them to, at best, just barely outperform a rule-based approach. We next investigate the effects of incorporating dialogue structure and different kinds of general world knowledge into both rule-based and (neural and non-neural) machine learning-based reading comprehension models. Experimental results on the DREAM data set show the effectiveness of dialogue structure and general world knowledge. DREAM is available at https://dataset.org/dream/ . Kai Sun 0006, Dian Yu 0001, Jianshu Chen, Dong Yu 0001, Yejin Choi 0001, Claire Cardie |
Trans. Assoc. Comput. Linguistics | 3 |
| 2018 | XL-NBT: A Cross-lingual Neural Belief Tracking FrameworkabstractTask-oriented dialog systems are becoming pervasive, and many companies heavily rely on them to complement human agents for customer service in call centers.With globalization, the need for providing cross-lingual customer support becomes more urgent than ever.However, cross-lingual support poses great challenges-it requires a large amount of additional annotated data from native speakers.In order to bypass the expensive human annotation and achieve the first step towards the ultimate goal of building a universal dialog system, we set out to build a cross-lingual state tracking framework.Specifically, we assume that there exists a source language with dialog belief tracking annotations while the target languages have no annotated dialog data of any form.Then, we pre-train a state tracker for the source language as a teacher, which is able to exploit easy-to-access parallel data.We then distill and transfer its own knowledge to the student state tracker in target languages.We specifically discuss two types of common parallel resources: bilingual corpus and bilingual dictionary, and design different transfer learning strategies accordingly.Experimentally, we successfully use English state tracker as the teacher to transfer its knowledge to both Italian and German trackers and achieve promising results. Wenhu Chen, Jianshu Chen, Yu Su 0001, Xin Wang 0061, Dong Yu 0001, Xifeng Yan, William Yang Wang |
EMNLP | 2 |
| 2018 | SBEED: Convergent Reinforcement Learning with Nonlinear Function ApproximationabstractWhen function approximation is used, solving the Bellman optimality equation with stability guarantees has remained a major open problem in reinforcement learning for decades. The fundamental difficulty is that the Bellman operator may become an expansion in general, resulting in oscillating and even divergent behavior of popular algorithms like Q-learning. In this paper, we revisit the Bellman equation, and reformulate it into a novel primal-dual optimization problem using Nesterov’s smoothing technique and the Legendre-Fenchel transformation. We then develop a new algorithm, called Smoothed Bellman Error Embedding, to solve this optimization problem where any differentiable function class may be used. We provide what we believe to be the first convergence guarantee for general nonlinear function approximation, and analyze the algorithm’s sample complexity. Empirically, our algorithm compares favorably to state-of-the-art baselines in several benchmark control problems. Bo Dai 0001, Albert Eaton Shaw, Lihong Li 0001, Niao He, Zhen Liu 0019, Jianshu Chen |
ICML | 7 |
| 2018 | Coupled Variational Bayes via Optimization EmbeddingabstractVariational inference plays a vital role in learning graphical models, especially on large-scale datasets. Much of its success depends on a proper choice of auxiliary distribution class for posterior approximation. However, how to pursue an auxiliary distribution class that achieves both good approximation ability and computation efficiency remains a core challenge. In this paper, we proposed coupled variational Bayes which exploits the primal-dual view of the ELBO with the variational distribution class generated by an optimization procedure, which is termed optimization embedding. This flexible function class couples the variational distribution with the original parameters in the graphical models, allowing end-to-end learning of the graphical models by back-propagation through the variational distribution. Theoretically, we establish an interesting connection to gradient flow and demonstrate the extreme flexibility of this implicit distribution family in the limit sense. Empirically, we demonstrate the effectiveness of the proposed method on multiple graphical models with either continuous or discrete latent variables comparing to state-of-the-art methods. Bo Dai 0001, Hanjun Dai, Niao He, Weiyang Liu, Zhen Liu 0019, Jianshu Chen |
NeurIPS | 6 |
| 2018 | M-Walk: Learning to Walk over Graphs using Monte Carlo Tree SearchabstractLearning to walk over a graph towards a target node for a given query and a source node is an important problem in applications such as knowledge base completion (KBC). It can be formulated as a reinforcement learning (RL) problem with a known state transition model. To overcome the challenge of sparse rewards, we develop a graph-walking agent called M-Walk, which consists of a deep recurrent neural network (RNN) and Monte Carlo Tree Search (MCTS). The RNN encodes the state (i.e., history of the walked path) and maps it separately to a policy and Q-values. In order to effectively train the agent from sparse rewards, we combine MCTS with the neural policy to generate trajectories yielding more positive rewards. From these trajectories, the network is improved in an off-policy manner using Q-learning, which modifies the RNN policy via parameter sharing. Our proposed RL algorithm repeatedly applies this policy-improvement step to learn the model. At test time, MCTS is combined with the neural policy to predict the target node. Experimental results on several graph-walking benchmarks show that M-Walk is able to learn better policies than other RL-based methods, which are mainly based on policy gradients. M-Walk also outperforms traditional KBC baselines. Yelong Shen, Jianshu Chen, Po-Sen Huang, Yuqing Guo 0003, Jianfeng Gao 0001 |
NeurIPS | 2 |
| 2017 | Character-level deep conflation for business data analyticsabstractConnecting different text attributes associated with the same entity (conflation) is important in business data analytics since it could help merge two different tables in a database to provide a more comprehensive profile of an entity. However, the conflation task is challenging because two text strings that describe the same entity could be quite different from each other for reasons such as misspelling. It is therefore critical to develop a conflation model that is able to truly understand the semantic meaning of the strings and match them at the semantic level. To this end, we develop a character-level deep conflation model that encodes the input text strings from character level into finite dimension feature vectors, which are then used to compute the cosine similarity between the text strings. The model is trained in an end-to-end manner using back propagation and stochastic gradient descent to maximize the likelihood of the correct association. Specifically, we propose two variants of the deep conflation model, based on long-short-term memory (LSTM) recurrent neural network (RNN) and convolutional neural network (CNN), respectively. Both models perform well on a real-world business analytics dataset and significantly outperform the baseline bag-of-character (BoC) model. Zhe Gan, P. D. Singh, Ameet Joshi, Xiaodong He 0001, Jianshu Chen, Jianfeng Gao 0001, Li Deng 0001 |
ICASSP | 5 |
| 2017 | Stochastic Variance Reduction Methods for Policy EvaluationabstractPolicy evaluation is concerned with estimating the value function that predicts long-term values of states under a given policy. It is a crucial step in many reinforcement-learning algorithms. In this paper, we focus on policy evaluation with linear function approximation over a fixed dataset. We first transform the empirical policy evaluation problem into a (quadratic) convex-concave saddle-point problem, and then present a primal-dual batch gradient method, as well as two stochastic variance reduction methods for solving the problem. These algorithms scale linearly in both sample size and feature dimension. Moreover, they achieve linear convergence even when the saddle-point problem has only strong concavity in the dual variables but no strong convexity in the primal variables. Numerical experiments on benchmark problems demonstrate the effectiveness of our methods. Simon S. Du, Jianshu Chen, Lihong Li 0001, Dengyong Zhou |
ICML | 2 |
| 2017 | Q-LDA: Uncovering Latent Patterns in Text-based Sequential Decision ProcessesabstractIn sequential decision making, it is often important and useful for end users to understand the underlying patterns or causes that lead to the corresponding decisions. However, typical deep reinforcement learning algorithms seldom provide such information due to their black-box nature. In this paper, we present a probabilistic model, Q-LDA, to uncover latent patterns in text-based sequential decision processes. The model can be understood as a variant of latent topic models that are tailored to maximize total rewards; we further draw an interesting connection between an approximate maximum-likelihood estimation of Q-LDA and the celebrated Q-learning algorithm. We demonstrate in the text-game domain that our proposed method not only provides a viable mechanism to uncover latent patterns in decision processes, but also obtains state-of-the-art rewards in these games. Jianshu Chen, Chong Wang 0002, Lihong Li 0001, Li Deng 0001 |
NIPS | 1 |
| 2017 | Unsupervised Sequence Classification using Sequential Output StatisticsabstractWe consider learning a sequence classifier without labeled data by using sequential output statistics. The problem is highly valuable since obtaining labels in training data is often costly, while the sequential output statistics (e.g., language models) could be obtained independently of input data and thus with low or no cost. To address the problem, we propose an unsupervised learning cost function and study its properties. We show that, compared to earlier works, it is less inclined to be stuck in trivial solutions and avoids the need for a strong generative model. Although it is harder to optimize in its functional form, a stochastic primal-dual gradient method is developed to effectively solve the problem. Experiment results on real-world datasets demonstrate that the new unsupervised learning method gives drastically lower errors than other baseline methods. Specifically, it reaches test errors about twice of those obtained by fully supervised learning. Yu Liu 0061, Jianshu Chen, Li Deng 0001 |
NIPS | 2 |
| 2017 | CherryPick: Adaptively Unearthing the Best Cloud Configurations for Big Data Analytics
Omid Alipourfard, Hongqiang Harry Liu, Jianshu Chen, Shivaram Venkataraman, Minlan Yu, Ming Zhang 0005 |
NSDI | 3 |
| 2016 | Deep Reinforcement Learning with a Natural Language Action SpaceabstractJi He, Jianshu Chen, Xiaodong He, Jianfeng Gao, Lihong Li, Li Deng, Mari Ostendorf. Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2016. Jianshu Chen, Xiaodong He 0001, Jianfeng Gao 0001, Lihong Li 0001, Li Deng 0001, Mari Ostendorf |
ACL (1) | 2 |
| 2016 | Deep Reinforcement Learning with a Combinatorial Action Space for Predicting Popular Reddit ThreadsabstractWe introduce an online popularity prediction and tracking task as a benchmark task for reinforcement learning with a combinatorial, natural language action space.A specified number of discussion threads predicted to be popular are recommended, chosen from a fixed window of recent comments to track.Novel deep reinforcement learning architectures are studied for effective modeling of the value function associated with actions comprised of interdependent sub-actions.The proposed model, which represents dependence between sub-actions through a bi-directional LSTM, gives the best performance across different experimental configurations and domains, and it also generalizes well with varying numbers of recommendation requests. Mari Ostendorf, Xiaodong He 0001, Jianshu Chen, Jianfeng Gao 0001, Lihong Li 0001, Li Deng 0001 |
EMNLP | 4 |
| 2016 | Interpreting the prediction process of a deep network constructed from supervised topic modelsabstractIn this paper, we propose an approach to interpret the prediction process of the BP-sLDA model, which is a supervised Latent Dirichlet Allocation model trained by Back Propagation over a deep architecture. The model is shown to achieve state-of-the-art prediction performance on several large-scale text analysis tasks. To interpret the prediction process of the model, often demanded by business data analytics applications, we perform evidence analysis on each pair-wise decision boundary over the topic distribution space, which is decomposed into a positive and a negative components. Then, for each element in the current document, a novel evidence score is defined by exploiting this topic decomposition and the generative nature of LDA. Then the score is used to rank the relative evidence of each element for the effectiveness of model prediction. We demonstrate the effectiveness of the method on a large-scale binary classification task on a corporate proprietary dataset with business-centric applications. Jianshu Chen, Xiaodong He 0001, Jianfeng Gao 0001, Li Deng 0001 |
ICASSP | 1 |
| 2016 | The 2015 NIST Language Recognition Evaluation: The Shared View of I2R, Fantastic4 and SingaMSabstractTechnical report for NIST LRE 2015 Workshop Kong-Aik Lee, Haizhou Li 0001, Li Deng 0001, Ville Hautamäki, Wei Rao 0002, Anthony Larcher, Hanwu Sun, Trung Hieu Nguyen 0001, Guangsen Wang, Aleksandr Sizov, Jianshu Chen, Ivan Kukanov, Amir Hossein Poorjam, Trung Ngo Trong, Chenglin Xu, Haihua Xu 0001, Bin Ma 0001, Chng Eng Siong, Sylvain Meignier |
INTERSPEECH | 12 |
| 2016 | Deep Sentence Embedding Using Long Short-Term Memory Networks: Analysis and Application to Information RetrievalabstractThis paper develops a model that addresses sentence embedding, a hot topic in current natural language processing research, using recurrent neural networks (RNN) with Long Short-Term Memory (LSTM) cells. The proposed LSTM-RNN model sequentially takes each word in a sentence, extracts its information, and embeds it into a semantic vector. Due to its ability to capture long term memory, the LSTM-RNN accumulates increasingly richer information as it goes through the sentence, and when it reaches the last word, the hidden layer of the network provides a semantic representation of the whole sentence. In this paper, the LSTM-RNN is trained in a weakly supervised manner on user click-through data logged by a commercial web search engine. Visualization and analysis are performed to understand how the embedding process works. The model is found to automatically attenuate the unimportant words and detect the salient keywords in the sentence. Furthermore, these detected keywords are found to automatically activate different cells of the LSTM-RNN, where words belonging to a similar topic activate the same cell. As a semantic representation of the sentence, the embedding vector can be used in many different applications. These automatic keyword detection and topic allocation abilities enabled by the LSTM-RNN allow the network to perform document retrieval, a difficult language processing task, where the similarity between the query and documents can be measured by the distance between their corresponding sentence embedding vectors computed by the LSTM-RNN. On a web search task, the LSTM-RNN embedding is shown to significantly outperform several existing state of the art methods. We emphasize that the proposed model generates sentence embedding vectors that are specially useful for web document retrieval tasks. A comparison with a well known general sentence embedding method, the Paragraph Vector, is performed. The results show that the proposed method in this paper significantly outperforms Paragraph Vector method for web document retrieval task. Hamid Palangi, Li Deng 0001, Yelong Shen, Jianfeng Gao 0001, Xiaodong He 0001, Jianshu Chen, Xinying Song, Rabab K. Ward |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2016 | Excess-Risk of Distributed Stochastic LearnersabstractThis paper studies the learning ability of consensus and diffusion distributed learners from continuous streams of data arising from different but related statistical distributions. Four distinctive features for diffusion learners are revealed in relation to other decentralized schemes even under left-stochastic combination policies. First, closed-form expressions for the evolution of their excess-risk are derived for strongly convex risk functions under a diminishing step-size rule. Second, using these results, it is shown that the diffusion strategy improves the asymptotic convergence rate of the excess-risk relative to non-cooperative schemes. Third, it is shown that when the in-network cooperation rules are designed optimally, the performance of the diffusion implementation can outperform that of naive centralized processing. Finally, the arguments further show that diffusion outperforms consensus strategies asymptotically and that the asymptotic excess-risk expression is invariant to the particular network topology. The framework adopted in this paper studies convergence in the stronger mean-square-error sense, rather than in distribution, and develops tools that enable a close examination of the differences between distributed strategies in terms of asymptotic behavior, as well as in terms of convergence rates. Zaid J. Towfic, Jianshu Chen, Ali H. Sayed |
IEEE Trans. Inf. Theory | 2 |
| 2015 | End-to-end Learning of LDA by Mirror-Descent Back Propagation over a Deep ArchitectureabstractWe develop a fully discriminative learning approach for supervised Latent Dirichlet Allocation (LDA) model using Back Propagation (i.e., BP-sLDA), which maximizes the posterior probability of the prediction variable given the input document. Different from traditional variational learning or Gibbs sampling approaches, the proposed learning method applies (i) the mirror descent algorithm for maximum a posterior inference and (ii) back propagation over a deep architecture together with stochastic gradient/mirror descent for model parameter estimation, leading to scalable and end-to-end discriminative learning of the model. As a byproduct, we also apply this technique to develop a new learning method for the traditional unsupervised LDA model (i.e., BP-LDA). Experimental results on three real-world regression and classification tasks show that the proposed methods significantly outperform the previous supervised topic models, neural networks, and is on par with deep neural networks. Jianshu Chen, Yelong Shen, Xiaodong He 0001, Jianfeng Gao 0001, Xinying Song, Li Deng 0001 |
NIPS | 1 |
| 2015 | On the Learning Behavior of Adaptive Networks - Part I: Transient AnalysisabstractThis paper carries out a detailed transient analysis of the learning behavior of multiagent networks, and reveals interesting results about the learning abilities of distributed strategies. Among other results, the analysis reveals how combination policies influence the learning process of networked agents, and how these policies can steer the convergence point toward any of many possible Pareto optimal solutions. The results also establish that the learning process of an adaptive network undergoes three (rather than two) well-defined stages of evolution with distinctive convergence rates during the first two stages, while attaining a finite mean-square-error level in the last stage. The analysis reveals what aspects of the network topology influence performance directly and suggests design procedures that can optimize performance by adjusting the relevant topology parameters. Interestingly, it is further shown that, in the adaptation regime, each agent in a sparsely connected network is able to achieve the same performance level as that of a centralized stochastic-gradient strategy even for left-stochastic combination strategies. These results lead to a deeper understanding and useful insights on the convergence behavior of coupled distributed learners. The results also lead to effective design mechanisms to help diffuse information more thoroughly over networks. Jianshu Chen, Ali H. Sayed |
IEEE Trans. Inf. Theory | 1 |
| 2015 | On the Learning Behavior of Adaptive Networks - Part II: Performance AnalysisabstractPart I of this paper examined the mean-square stability and convergence of the learning process of distributed strategies over graphs. The results identified conditions on the network topology, utilities, and data in order to ensure stability; the results also identified three distinct stages in the learning behavior of multiagent networks related to transient phases I and II and the steady-state phase. This Part II examines the steady-state phase of distributed learning by networked agents. Apart from characterizing the performance of the individual agents, it is shown that the network induces a useful equalization effect across all agents. In this way, the performance of noisier agents is enhanced to the same level as the performance of agents with less noisy data. It is further shown that in the small step-size regime, each agent in the network is able to achieve the same performance level as that of a centralized strategy corresponding to a fully connected network. The results in this part reveal explicitly which aspects of the network topology and operation influence performance and provide important insights into the design of effective mechanisms for the processing and diffusion of information over networks. Jianshu Chen, Ali H. Sayed |
IEEE Trans. Inf. Theory | 1 |
| 2014 | Online dictionary learning over distributed modelsabstractIn this paper, we consider learning dictionary models over a network of agents, where each agent is only in charge of a portion of the dictionary elements. This formulation is relevant in big data scenarios where multiple large dictionary models may be spread over different spatial locations and it is not feasible to aggregate all dictionaries in one location due to communication and privacy considerations. We first show that the dual function of the inference problem is an aggregation of individual cost functions associated with different agents, which can then be minimized efficiently by means of diffusion strategies. The collaborative inference step generates local error measures that are used by the agents to update their dictionaries without the need to share these dictionaries or even the coefficient models for the training data. This is a useful property that leads to an efficient distributed procedure for learning dictionaries over large networks. Jianshu Chen, Zaid J. Towfic, Ali H. Sayed |
ICASSP | 1 |
| 2014 | Sequence classification using the high-level features extracted from deep neural networksabstractThe recent success of deep neural networks (DNNs) in speech recognition can be attributed largely to their ability to extract a specific form of high-level features from raw acoustic data for subsequent sequence classification or recognition tasks. Among the many possible forms of DNN features, what forms are more useful than others and how effective these DNN features are in connection with the different types of downstream sequence recognizers remained unexplored and are the focus of this paper. We report our recent work on the construction of a diverse set of DNN features, including the vectors extracted from the output layer and from various hidden layers in the DNN. We then apply these features as the inputs to four types of classifiers to carry out the identical sequence classification task of phone recognition. The experimental results show that the features derived from the top hidden layer of the DNN perform the best for all four classifiers, especially for the autoregressive-moving-average (ARMA) version of a recurrent neural network. The feature vector derived from the DNN's output layer performs slightly worse but better than any of the hidden layers in the DNN except the top one. Li Deng 0001, Jianshu Chen |
ICASSP | 2 |
| 2013 | Cooperative off-policy prediction of Markov decision processes in adaptive networksabstractWe apply diffusion strategies to propose a cooperative reinforcement learning algorithm, in which agents in a network communicate with their neighbors to improve predictions about their environment. The algorithm is suitable to learn off-policy even in large state spaces. We provide a mean-square-error performance analysis under constant step-sizes. The gain of cooperation in the form of more stability and less bias and variance in the prediction error, is illustrated in the context of a classical model. We show that the improvement in performance is especially significant when the behavior policy of the agents is different from the target policy under evaluation. Sergio Valcarcel Macua, Jianshu Chen, Santiago Zazo, Ali H. Sayed |
ICASSP | 2 |
| 2013 | Distributed inference over regression and classification modelsabstractWe study the distributed inference task over regression and classification models where the likelihood function is strongly log-concave. We show that diffusion strategies allow the KL divergence between two likelihood functions to converge to zero at the rate 1/Ni on average and with high probability, where N is the number of nodes in the network and i is the number of iterations. We derive asymptotic expressions for the expected regularized KL divergence and show that the diffusion strategy can outperform both non-cooperative and conventional centralized strategies, since diffusion implementations can weigh a node's contribution in proportion to its noise level. Zaid J. Towfic, Jianshu Chen, Ali H. Sayed |
ICASSP | 2 |
| 2013 | On distributed online classification in the midst of concept drifts
Zaid J. Towfic, Jianshu Chen, Ali H. Sayed |
Neurocomputing | 2 |
| 2013 | Cramer-Rao Bounds for Joint RSS/DoA-Based Primary-User Localization in Cognitive Radio NetworksabstractKnowledge about the location of licensed primary-users (PU) could enable several key features in cognitive radio (CR) networks including improved spatio-temporal sensing, intelligent location-aware routing, as well as aiding spectrum policy enforcement. In this paper we consider the achievable accuracy of PU localization algorithms that jointly utilize received-signal-strength (RSS) and direction-of-arrival (DoA) measurements by evaluating the Cramer-Rao Bound (CRB). Previous works evaluate the CRB for RSS-only and DoA-only localization algorithms separately and assume DoA estimation error variance is a fixed constant or rather independent of RSS. We derive the CRB for joint RSS/DoA-based PU localization algorithms based on the mathematical model of DoA estimation error variance as a function of RSS, for a given CR placement. The bound is compared with practical localization algorithms and the impact of several key parameters, such as number of nodes, number of antennas and samples, channel shadowing variance and correlation distance, on the achievable accuracy are thoroughly analyzed and discussed. We also derive the closed-form asymptotic CRB for uniform random CR placement, and perform theoretical and numerical studies on the required number of CRs such that the asymptotic CRB tightly approximates the numerical integration of the CRB for a given placement. Jun Wang 0007, Jianshu Chen, Danijela Cabric |
IEEE Trans. Wirel. Commun. | 2 |
| 2012 | Distributed learning via Diffusion adaptation with application to ensemble learning
Zaid J. Towfic, Jianshu Chen, Ali H. Sayed |
ESANN | 2 |
| 2012 | Performance of diffusion adaptation for collaborative optimizationabstractWe derive an adaptive diffusion mechanism to optimize global cost functions in a distributed manner over a network of nodes. The cost function is assumed to consist of the sum of individual components, and diffusion adaptation is used to enable the nodes to cooperate locally through in-network processing in order to solve the desired optimization problem. We analyze the mean-square-error performance of the algorithm, including its transient and steady-state behavior. We illustrate one application in the context of least-mean-squares estimation for sparse vectors. Jianshu Chen, Ali H. Sayed |
ICASSP | 1 |
| 2012 | Distributed throughput optimization over P2P mesh networks using diffusion adaptationabstractThis work develops a decentralized adaptive strategy for throughput maximization over peer-to-peer (P2P) networks. The adaptive strategy can cope with changing network topologies, is robust to network disruptions, and does not rely on central processors. The algorithm is obtained as a special case of a more general diffusion strategy for the distributed solution of optimization problems with constraints. Simulation results illustrate how the proposed technique is competitive with other methods. Zaid J. Towfic, Jianshu Chen, Ali H. Sayed |
ICC | 2 |
| 2011 | Bio-inspired cooperative optimization with application to bacteria motilityabstractInspired by bacterial motility, we propose an algorithm for adaptation over networks with mobile nodes. The nodes have limited abilities and they are allowed to cooperate with their neighbors to optimize a common objective function. In contrast to traditional adaptation formulations, an important consideration in this work is the fact that the nodes do not know the form of the cost function beforehand. The nodes can only sense variations in the values of the objective function as they diffuse through the space, such as sensing the variation in the concentration of nutrients in the environment. We propose a technique for the nodes to pick the search vector as a linear combination of the neighbors' last steps, by attempting to maximize the nutritional gradient. The procedure enables information to flow from "information-rich" nodes to the other nodes. Jianshu Chen, Ali H. Sayed |
ICASSP | 1 |
| 2009 | Analyzing amplify-and-forward and decode-and-forward cooperative strategies in Wyner's channel modelabstractThe benefits of Amplify-and-Forward (AF) and Decode-and-Forward (DF) cooperative relay for secure communication are investigated within Wyner's wiretap channel. We characterize the secrecy rate when source, destination, relay and eavesdropper all use single antenna and the channel conditions are fix. Both AF and DF cooperative strategies are proved theoretically to be able to facilitate secure communication. Detailed analysis of AF and DF scheme reveals a trade off between secrecy area and request secrecy rate. In addition, secrecy constraints in cooperative secure communication are discussed and are used to explain the differences in AF and DF scheme. Overall, our work establishes the utility of cooperation and compares each advantage of AF and DF scheme in facilitating secure communication over wireless channel. Jianshu Chen, Jian Wang 0030 |
WCNC | 3 |