EDBT 2026 Demo / reviewers in the wild / expert
Kai Fan 0002
dblp:20/3825-2
· DBLP profile ↗
33ranked-venue papers
7as first author
17since 2021 · last 2026
0000-0002-8256-0807ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 7 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient and Adaptive Simultaneous Speech Translation with Fully Unidirectional ArchitectureabstractSimultaneous speech translation (SimulST) produces translations incrementally while processing partial speech input. Although large language models (LLMs) have shown strong capabilities in offline translation tasks, applying them to SimulST poses notable challenges. Existing LLM-based SimulST approaches either incur significant computational overhead due to repeated encoding of bidirectional speech encoder, or they depend on a fixed read/write policy, limiting the efficiency and performance. In this work, we introduce Efficient and Adaptive Simultaneous Speech Translation (EASiST) with fully unidirectional architecture, including both speech encoder and LLM. EASiST includes a multi-latency data curation strategy to generate semantically aligned SimulST training samples and redefines SimulST as an interleaved generation task with explicit read/write tokens. To facilitate adaptive inference, we incorporate a lightweight policy head that dynamically predicts read/write actions. Additionally, we employ a multi-stage training strategy to align speech-text modalities and optimize both translation and policy behavior. Experiments on both in-domain (MuST-C) and out-of-domain (Europarl-ST) En-De and En-Es datasets demonstrate that EASiST offers superior latency-quality trade-offs compared to several strong baselines. Biao Fu, Donglei Yu, Minpeng Liao, Chengxi Li 0014, Xinjie Chen, Yidong Chen 0001, Kai Fan 0002, Xiaodong Shi |
AAAI | 7 |
| 2026 | From Data-Centric to Sample-Centric: Enhancing LLM Reasoning via Progressive OptimizationabstractXinjie Chen, Minpeng Liao, Guoxin Chen, Chengxi Li, Biao Fu, Kai Fan, Xinggao Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xinjie Chen, Minpeng Liao, Guoxin Chen, Chengxi Li 0014, Biao Fu, Kai Fan 0002, Xinggao Liu |
ACL (1) | 6 |
| 2025 | Fixing Distribution Shifts of LLM Self-Critique via On-Policy Self-Play TrainingabstractSelf-critique mechanisms significantly improve the performance of language models in complex reasoning tasks by giving them the ability to correct errors, conduct induction and deduction, and switch thinking insights. However, synthetic data methods often require human-introduced errors or sampling of the model’s reasoning results from the previous moment, and the current output distribution of the model cannot be obtained, makes the data for critique and reasoning face the problem of distribution shifts. In this work, we propose an on-policy reinforcement learning framework to synchronize the reasoning and critique capabilities of language models. To alleviate reward hacking caused by outcome-based supervision, we design a deliberate reward framework for different purposes. The reward framework not only supervises the model reasoning process based on the results, but also uses Monte Carlo sampling to give appropriate rewards to the critique content according to the success rate of the model’s correction after critique. In addition, we introduce a rule-based reward function to impose penalties on the model when it generates hallucinatory critiques. When our approach is applied to the DeepSeek-Math-7B-Base and Qwen2.5-7B-Base models, model performance improves 5.40 and 3.66 points, respectively, compared to the best baseline approach. This validates the significant advantages of our method in improving model’s reasoning and self-critique capability. Code will be made available at https://github.com/rbao2018/SCOP Rong Bao, Donglei Yu, Kai Fan 0002, Minpeng Liao |
ACL (1) | 3 |
| 2025 | C-3PO: Compact Plug-and-Play Proxy Optimization to Achieve Human-like Retrieval-Augmented GenerationabstractRetrieval-augmented generation (RAG) systems face a fundamental challenge in aligning independently developed retrievers and large language models (LLMs). Existing approaches typically involve modifying either component or introducing simple intermediate modules, resulting in practical limitations and sub-optimal performance. Inspired by human search behavior—typically involving a back-and-forth process of proposing search queries and reviewing documents, we propose C-3PO, a proxy-centric framework that facilitates communication between retrievers and LLMs through a lightweight multi-agent system. Our framework implements three specialized agents that collaboratively optimize the entire RAG pipeline without altering the retriever and LLMs. These agents work together to assess the need for retrieval, generate effective queries, and select information suitable for the LLMs. To enable effective multi-agent coordination, we develop a tree-structured rollout approach for reward credit assignment in reinforcement learning. Extensive experiments in both in-domain and out-of-distribution scenarios demonstrate that C-3PO significantly enhances RAG performance while maintaining plug-and-play flexibility and superior generalization capabilities. Guoxin Chen, Minpeng Liao, Peiying Yu, Dingmin Wang, Zile Qiao, Wayne Xin Zhao, Kai Fan 0002 |
ICML | 8 |
| 2025 | Markov Chain of Thought for Efficient Mathematical ReasoningabstractWen Yang, Minpeng Liao, Kai Fan. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Minpeng Liao, Kai Fan 0002 |
NAACL (Long Papers) | 3 |
| 2024 | Divergence-Guided Simultaneous Speech TranslationabstractTo achieve high-quality translation with low latency, a Simultaneous Speech Translation (SimulST) system relies on a policy module to decide whether to translate immediately or wait for additional streaming input, along with a translation model capable of effectively handling partial speech input. Prior research has tackled these components separately, either using ``wait-k'' policies based on fixed-length segments or detected word boundaries, or dynamic policies based on different strategies (e.g., meaningful units), while employing offline models for prefix-to-prefix translation. In this paper, we propose Divergence-Guided Simultaneous Speech Translation (DiG-SST), a tightly integrated approach focusing on both translation quality and latency for streaming input. Specifically, we introduce a simple yet effective prefix-based strategy for training translation models with partial speech input, and develop an adaptive policy that makes read/write decisions for the translation model based on the expected divergence in translation distributions resulting from future input. Our experiments on multiple translation directions of the MuST-C benchmark demonstrate that our approach achieves a better trade-off between translation quality and latency compared to existing methods. Xinjie Chen, Kai Fan 0002, Xinggao Liu, Zhongqiang Huang |
AAAI | 2 |
| 2024 | AlphaMath Almost Zero: Process Supervision without ProcessabstractAlthough recent advancements in large language models (LLMs) have significantly improved their performance on various tasks, they still face challenges with complex and symbolic multi-step reasoning, particularly in mathematical reasoning. To bolster the mathematical reasoning capabilities of LLMs, most existing efforts concentrate on seeking assistance from either domain experts or GPT-4 for high-quality process-supervised data, which is not only expensive but also labor-intensive. In our study, we propose an innovative framework, AlphaMath, that bypasses the need for process annotations (from humans or GPTs) by leveraging Monte Carlo Tree Search (MCTS). This framework focuses on unleashing the potential of a well-pretrained LLM to autonomously enhance its mathematical reasoning. Specifically, we integrate a value model with the LLM, automatically generating both process supervision and step-level evaluation signals in MCTS. Furthermore, we propose an efficient inference strategy—step-level beam search, where the value model is crafted to assist the policy model (i.e., LLM) in navigating more effective reasoning paths, rather than solely relying on prior probabilities. The experimental results on both in-domain and out-of-domain datasets demonstrate that even without GPT-4 or human-annotated process supervision, our AlphaMath framework achieves comparable or superior results to previous state-of-the-art methods. Guoxin Chen, Minpeng Liao, Chengxi Li 0014, Kai Fan 0002 |
NeurIPS | 4 |
| 2023 | Better Simultaneous Translation with Monotonic Knowledge DistillationabstractSimultaneous machine translation (SiMT) presents a unique challenge as it requires generating target tokens before the source sentence is fully consumed.This can lead to the hallucination problem, where target tokens are generated without support from the source sentence.The prefix-to-prefix training data used to train SiMT models are not always parallel, due to divergent word order between the source and target languages, and can contribute to the problem.In this paper, we propose a novel approach that leverages traditional translation models as teachers and employs a two-stage beam search algorithm to generate monotonic yet accurate reference translations for sequence-level knowledge distillation.Experimental results demonstrate the significant improvements achieved by our approach over multiple strong SiMT baselines, leading to new state-of-the-art performance across various language pairs.Notably, when evaluated on a monotonic version of the WMT15 De→En test set, which includes references generated in a more monotonic style by professional translators, our approach achieves even more substantial improvement over the baselines.The source code and data are publicly available for further exploration 1 . Shushu Wang, Kai Fan 0002, Zhongqiang Huang |
ACL (1) | 3 |
| 2023 | Adapting Offline Speech Translation Models for Streaming with Future-Aware Distillation and InferenceabstractA popular approach to streaming speech translation is to employ a single offline model with a wait-k policy to support different latency requirements, which is simpler than training multiple online models with different latency constraints.However, there is a mismatch problem in using a model trained with complete utterances for streaming inference with partial input.We demonstrate that speech representations extracted at the end of a streaming input are significantly different from those extracted from a complete utterance.To address this issue, we propose a new approach called Future-Aware Streaming Translation (FAST) that adapts an offline ST model for streaming input.FAST includes a Future-Aware Inference (FAI) strategy that incorporates future context through a trainable masked embedding, and a Future-Aware Distillation (FAD) framework that transfers future context from an approximation of full speech to streaming input.Our experiments on the MuST-C EnDe, EnEs, and EnFr benchmarks show that FAST achieves better trade-offs between translation quality and latency than strong baselines.Extensive analyses suggest that our methods effectively alleviate the aforementioned mismatch problem between offline training and online inference.1 Biao Fu, Minpeng Liao, Kai Fan 0002, Zhongqiang Huang, Boxing Chen, Yidong Chen 0001, Xiaodong Shi |
EMNLP | 3 |
| 2023 | Training Simultaneous Speech Translation with Robust and Random Wait-k-Tokens StrategyabstractSimultaneous Speech Translation (SimulST) is a task focused on ensuring high-quality translation of speech in low-latency situations.Despite this, the modality gap (e.g., unknown word boundaries) between audio and text presents a challenge.This gap hinders the effective application of policies from simultaneous text translation (SimulMT) and compromises the performance of offline speech translation.To address this issue, we first leverage the Montreal Forced Aligner (MFA) and utilize audio transcription pairs in pre-training the acoustic encoder, and introduce a token-level crossmodal alignment that allows the wait-k policy from SimulMT to better adapt to SimulST.This token-level boundary alignment simplifies the decision-making process for predicting read/write actions, as if the decoder were directly processing text tokens.Subsequently, to optimize the SimulST task, we propose a robust and random wait-k-tokens strategy.This strategy allows a single model to meet various latency requirements and minimizes error accumulation of boundary alignment during inference.Our experiments on the MuST-C dataset show that our method achieves a better tradeoff between translation quality and latency. Kai Fan 0002, Jiajun Bu, Zhongqiang Huang |
EMNLP | 2 |
| 2023 | Adaptive Policy with Wait-k Model for Simultaneous TranslationabstractSimultaneous machine translation (SiMT) requires a robust read/write (R/W) policy in conjunction with a high-quality translation model.Traditional methods rely on either a fixed waitk policy coupled with a standalone wait-k translation model, or an adaptive policy jointly trained with the translation model.In this study, we propose a more flexible approach by decoupling the adaptive policy model from the translation model.Our motivation stems from the observation that a standalone multi-path waitk model performs competitively with adaptive policies utilized in state-of-the-art SiMT approaches.Specifically, we introduce DaP, a divergence-based adaptive policy, that makes read/write decisions for any translation model based on the potential divergence in translation distributions resulting from future information.DaP extends a frozen wait-k model with lightweight parameters, and is both memory and computation efficient.Experimental results across various benchmarks demonstrate that our approach offers an improved trade-off between translation accuracy and latency, outperforming strong baselines.1 Kai Fan 0002, Shushu Wang, Ziqian Zeng, Zhongqiang Huang |
EMNLP | 2 |
| 2023 | StrokeNet: Stroke Assisted and Hierarchical Graph Reasoning NetworksabstractScene text detection is still a challenging task, as there may be extremely small or low-resolution strokes and close or arbitrary-shaped texts. In this paper, StrokeNet proposes to effectively detect the texts by capturing the fine-grained strokes and inferring structural relations between the hierarchical representations of each text area in the graph-based network. Different from existing approaches that represent the text area by a series of points or rectangular boxes, we directly localize the strokes of each text instance. We introduce Stroke Assisted Prediction Network (SAPN), which performs hierarchical representation learning of text areas, effectively capturing extremely small or low-resolution texts. We extract a series of text- and stroke-level rectangular boxes on the predicted text areas, which are treated as graph nodes and grouped to form the corresponding local graphs. Hierarchical Relation Graph Network (HRGN) then performs relational reasoning and predicts the likelihood of linkages among graph nodes of different levels. It efficiently splits the close text instances and grouping node classification results into the arbitrary-shaped text area. We introduce a novel dataset with stroke-level annotations, namelySynthStroke, for offline pre-training of widespread text detectors. Experiments on benchmarks verify the State-of-the-Art performance of our method. Lei Li 0051, Kai Fan 0002, Chun Yuan 0003 |
IEEE Trans. Multim. | 2 |
| 2022 | Efficient Cluster-Based k-Nearest-Neighbor Machine Translationabstractk-Nearest-Neighbor Machine Translation (kNN-MT) has been recently proposed as a non-parametric solution for domain adaptation in neural machine translation (NMT).It aims to alleviate the performance degradation of advanced MT systems in translating out-ofdomain sentences by coordinating with an additional token-level feature-based retrieval module constructed from in-domain data.Previous studies (Khandelwal et al., 2021; Zheng et al., 2021a) have already demonstrated that non-parametric NMT is even superior to models fine-tuned on out-of-domain data.In spite of this success, kNN retrieval is at the expense of high latency, in particular for large datastores.To make it practical, in this paper, we explore a more efficient kNN-MT and propose to use clustering to improve the retrieval efficiency.Concretely, we first propose a cluster-based Compact Network for feature reduction in a contrastive learning manner to compress context features into 90+% lower dimensional vectors.We then suggest a cluster-based pruning solution to filter out 10%~40% redundant nodes in large datastores while retaining translation quality.Our proposed methods achieve better or comparable performance while reducing up to 57% inference latency against the advanced non-parametric MT model on several machine translation benchmarks.Experimental results indicate that the proposed methods maintain the most useful information of the original datastore and the Compact Network shows good generalization on unseen domains.Codes are available at https: //github.com/tjunlp-lab/PCKMT.Datastore distribution C-II("bank") C-III("bank") C-VI("bill") Compact Layer Compact Layer centroid ✖ Datastore distribution C-VI("bill") C-II("bank") C-I("sell") C-III("bank") Datastore distribution C-VI("bill") C-II("bank") C-I("sell") C-III("bank") Join Repartition Prune C-III C-II A pruning example w.r.t."bank" ✔ C-I("sell") Kai Fan 0002, Boxing Chen, Deyi Xiong |
ACL (1) | 2 |
| 2022 | Competency-Aware Neural Machine Translation: Can Machine Translation Know its Own Translation Quality?abstractNeural machine translation (NMT) is often criticized for failures that happen without awareness.The lack of competency awareness makes NMT untrustworthy.This is in sharp contrast to human translators who give feedback or conduct further investigations whenever they are in doubt about predictions.To fill this gap, we propose a novel competency-aware NMT by extending conventional NMT with a selfestimator, offering abilities to translate a source sentence and estimate its competency.The selfestimator encodes the information of the decoding procedure and then examines whether it can reconstruct the original semantics of the source sentence.Experimental results on four translation tasks demonstrate that the proposed method not only carries out translation tasks intact but also delivers outstanding performance on quality estimation.Without depending on any reference or annotated data typically required by state-of-the-art metric and quality estimation methods, our model yields an even higher correlation with human quality judgments than a variety of aforementioned methods, such as BLEURT, COMET, and BERTScore.Quantitative and qualitative analyses show better robustness of competency awareness in our model.1 Pei Zhang 0011, Baosong Yang, Dayiheng Liu, Kai Fan 0002, Luo Si |
EMNLP | 5 |
| 2022 | Cross-modal Representation Learning and Relation Reasoning for Bidirectional Adaptive ManipulationabstractSince single-modal controllable manipulation typically requires supervision of information from other modalities or cooperation with complex software and experts, this paper addresses the problem of cross-modal adaptive manipulation (CAM). The novel task performs cross-modal semantic alignment from mutual supervision and implements bidirectional exchange of attributes, relations, or objects in parallel, benefiting both modalities while significantly reducing manual effort. We introduce a robust solution for CAM, which includes two essential modules, namely Heterogeneous Representation Learning (HRL) and Cross-modal Relation Reasoning (CRR). The former is designed to perform representation learning for cross-modal semantic alignment on heterogeneous graph nodes. The latter is adopted to identify and exchange the focused attributes, relations, or objects in both modalities. Our method produces pleasing cross-modal outputs on CUB and Visual Genome. Lei Li 0051, Kai Fan 0002, Chun Yuan 0003 |
IJCAI | 2 |
| 2022 | Unifying Cross-lingual Summarization and Machine Translation with Compression RateabstractCross-Lingual Summarization (CLS) is a task that extracts important information from a source document and summarizes it into a summary in another language. It is a challenging task that requires a system to understand, summarize, and translate at the same time, making it highly related to Monolingual Summarization (MS) and Machine Translation (MT). In practice, the training resources for Machine Translation are far more than that for cross-lingual and monolingual summarization. Thus incorporating the Machine Translation corpus into CLS would be beneficial for its performance. However, the present work only leverages a simple multi-task framework to bring Machine Translation in, lacking deeper exploration. Yu Bai 0018, Heyan Huang, Kai Fan 0002, Yang Gao 0016, Jiaao Zhan, Zewen Chi, Boxing Chen |
SIGIR | 3 |
| 2021 | Explore Hierarchical Relations Reasoning and Global Information Aggregation
Lei Li 0051, Chun Yuan 0003, Kai Fan 0002 |
ICDAR (1) | 3 |
| 2020 | Long-Short Term Masking Transformer: A Simple but Effective Baseline for Document-level Neural Machine TranslationabstractMany document-level neural machine translation (NMT) systems have explored the utility of context-aware architecture, usually requiring an increasing number of parameters and computational complexity.However, few attention is paid to the baseline model.In this paper, we research extensively the pros and cons of the standard transformer in document-level translation, and find that the auto-regressive property can simultaneously bring both the advantage of the consistency and the disadvantage of error accumulation.Therefore, we propose a surprisingly simple long-short term masking self-attention on top of the standard transformer to both effectively capture the long-range dependence and reduce the propagation of errors.We examine our approach on the two publicly available document-level datasets.We can achieve a strong result in BLEU and capture discourse phenomena. Pei Zhang 0011, Boxing Chen, Niyu Ge, Kai Fan 0002 |
EMNLP (1) | 4 |
| 2020 | Neural Zero-Inflated Quality Estimation Model for Automatic Speech Recognition SystemabstractThe performances of automatic speech recognition (ASR) systems are usually evaluated by the metric word error rate (WER) when the manually transcribed data are provided, which are, however, expensively available in the real scenario.In addition, the empirical distribution of WER for most ASR systems usually tends to put a significant mass near zero, making it difficult to simulate with a single continuous distribution.In order to address the two issues of ASR quality estimation (QE), we propose a novel neural zero-inflated model to predict the WER of the ASR result without transcripts.We design a neural zeroinflated beta regression on top of a bidirectional transformer language model conditional on speech features (speech-BERT).We adopt the pre-training strategy of token level masked language modeling for speech-BERT as well, and further fine-tune with our zero-inflated layer for the mixture of discrete and continuous outputs.The experimental results show that our approach achieves better performance on WER prediction compared with strong baselines. Kai Fan 0002, Bo Li 0121, Jiayi Wang 0010, Shiliang Zhang, Boxing Chen, Niyu Ge, Zhijie Yan |
INTERSPEECH | 1 |
| 2019 | "Bilingual Expert" Can Find Translation ErrorsabstractThe performances of machine translation (MT) systems are usually evaluated by the metric BLEU when the golden references are provided. However, in the case of model inference or production deployment, golden references are usually expensively available, such as human annotation with bilingual expertise. In order to address the issue of translation quality estimation (QE) without reference, we propose a general framework for automatic evaluation of the translation output for the QE task in the Conference on Statistical Machine Translation (WMT). We first build a conditional target language model with a novel bidirectional transformer, named neural bilingual expert model, which is pre-trained on large parallel corpora for feature extraction. For QE inference, the bilingual expert model can simultaneously produce the joint latent representation between the source and the translation, and real-valued measurements of possible erroneous tokens based on the prior knowledge learned from parallel data. Subsequently, the features will further be fed into a simple Bi-LSTM predictive model for quality estimation. The experimental results show that our approach achieves the state-of-the-art performance in most public available datasets of WMT 2017/2018 QE task. Kai Fan 0002, Jiayi Wang 0010, Bo Li 0121, Fengming Zhou, Boxing Chen, Luo Si |
AAAI | 1 |
| 2019 | Improving Distantly Supervised Relation Extraction with Neural Noise Converter and Conditional Optimal SelectorabstractDistant supervised relation extraction has been successfully applied to large corpus with thousands of relations. However, the inevitable wrong labeling problem by distant supervision will hurt the performance of relation extraction. In this paper, we propose a method with neural noise converter to alleviate the impact of noisy data, and a conditional optimal selector to make proper prediction. Our noise converter learns the structured transition matrix on logit level and captures the property of distant supervised relation extraction dataset. The conditional optimal selector on the other hand helps to make proper prediction decision of an entity pair even if the group of sentences is overwhelmed by no-relation sentences. We conduct experiments on a widely used dataset and the results show significant improvement over competitive baseline methods. Shanchan Wu, Kai Fan 0002 |
AAAI | 2 |
| 2019 | Lattice Transformer for Speech TranslationabstractRecent advances in sequence modeling have highlighted the strengths of the transformer architecture, especially in achieving state-of-theart machine translation results.However, depending on the up-stream systems, e.g., speech recognition, or word segmentation, the input to translation system can vary greatly.The goal of this work is to extend the attention mechanism of the transformer to naturally consume the lattice in addition to the traditional sequential input.We first propose a general lattice transformer for speech translation where the input is the output of the automatic speech recognition (ASR) which contains multiple paths and posterior scores.To leverage the extra information from the lattice structure, we develop a novel controllable lattice attention mechanism to obtain latent representations.On the LDC Spanish-English speech translation corpus, our experiments show that lattice transformer generalizes significantly better and outperforms both a transformer baseline and a lattice LSTM.Additionally, we validate our approach on the WMT 2017 Chinese-English translation task with lattice inputs from different BPE segmentations.In this task, we also observe the improvements over strong baselines. Pei Zhang 0011, Niyu Ge, Boxing Chen, Kai Fan 0002 |
ACL (1) | 4 |
| 2019 | Unsupervised Multi-Modal Neural Machine TranslationabstractUnsupervised neural machine translation (UNMT) has recently achieved remarkable results \cite{lample2018phrase} with only large monolingual corpora in each language. However, the uncertainty of associating target with source sentences makes UNMT theoretically an ill-posed problem. This work investigates the possibility of utilizing images for disambiguation to improve the performance of UNMT. Our assumption is intuitively based on the invariant property of image, i.e., the description of the same visual content by different languages should be approximately similar. We propose an unsupervised multi-modal machine translation (UMNMT) framework based on the language translation cycle consistency loss conditional on the image, targeting to learn the bidirectional multi-modal translation simultaneously. Through an alternate training between multi-modal and uni-modal, our inference model can translate with or without the image. On the widely used Multi30K dataset, the experimental results of our approach are significantly better than those of the text-only UNMT on the 2016 test dataset. Yuanhang Su, Kai Fan 0002, Nguyen Bach, C.-C. Jay Kuo, Fei Huang 0002 |
CVPR | 2 |
| 2019 | InverseNet: Solving Inverse Problems of Multimedia Data with Splitting NetworksabstractWe propose a novel network architecture, namely InverseNet, to solve the inverse problems of multimedia data. The inverse problem is cast in the form of learning an end-to-end mapping from observed multimedia data to the ground-truth data. Inspired by the splitting strategy to tackle inverse problems, the mapping is learned by InverseNet, a composition of two networks, with one handling the inversion of the physical forward model and the other handling the denoising of the output from the former network. Training InverseNet is annealing as the intermediate variable between these two networks bridges the gap between the input and output and progressively approaches to the ground-truth. Extensive experiments on synthetic and real multimedia datasets on the tasks, e.g., motion deblurring, super-resolution, and colorization, demonstrate the efficiency and accuracy of the proposed method compared with other image processing algorithms. Kai Fan 0002, Wenlin Wang, Tianhang Zheng, Amit Chakraborty, Katherine A. Heller, Changyou Chen, Kui Ren 0001 |
ICME | 2 |
| 2018 | Zero-Shot Learning via Class-Conditioned Deep Generative ModelsabstractWe present a deep generative model for Zero-Shot Learning (ZSL). Unlike most existing methods for this problem, that represent each class as a point (via a semantic embedding), we represent each seen/unseen class using a class-specific latent-space distribution, conditioned on class attributes. We use these latent-space distributions as a prior for a supervised variational autoencoder (VAE), which also facilitates learning highly discriminative feature representations for the inputs. The entire framework is learned end-to-end using only the seen-class training data. At test time, the label for an unseen-class test input is the class that maximizes the VAE lower bound. We further extend the model to a (i) semi-supervised/transductive setting by leveraging unlabeled unseen-class data via an unsupervised learning module, and (ii) few-shot learning where we also have a small number of labeled inputs from the unseen classes. We compare our model with several state-of-the-art methods through a comprehensive set of experiments on a variety of benchmark data sets. Wenlin Wang, Yunchen Pu, Vinay Kumar Verma, Kai Fan 0002, Yizhe Zhang 0002, Changyou Chen, Piyush Rai, Lawrence Carin |
AAAI | 4 |
| 2017 | Adversarial Feature Matching for Text GenerationabstractThe Generative Adversarial Network (GAN) has achieved great success in generating realistic (real-valued) synthetic data. However, convergence issues and difficulties dealing with discrete data hinder the applicability of GAN to text. We propose a framework for generating realistic text via adversarial training. We employ a long short-term memory network as generator, and a convolutional network as discriminator. Instead of using the standard objective of GAN, we propose matching the high-dimensional latent feature distributions of real and synthetic sentences, via a kernelized discrepancy metric. This eases adversarial training by alleviating the mode-collapsing problem. Our experiments show superior performance in quantitative evaluation, and demonstrate that our model can generate realistic-looking sentences. Yizhe Zhang 0002, Zhe Gan, Kai Fan 0002, Zhi Chen 0009, Ricardo Henao, Dinghan Shen, Lawrence Carin |
ICML | 3 |
| 2017 | An inner-loop free solution to inverse problems using deep neural networksabstractWe propose a new method that uses deep learning techniques to accelerate the popular alternating direction method of multipliers (ADMM) solution for inverse problems. The ADMM updates consist of a proximity operator, a least squares regression that includes a big matrix inversion, and an explicit solution for updating the dual variables. Typically, inner loops are required to solve the first two sub-minimization problems due to the intractability of the prior and the matrix inversion. To avoid such drawbacks or limitations, we propose an inner-loop free update rule with two pre-trained deep convolutional architectures. More specifically, we learn a conditional denoising auto-encoder which imposes an implicit data-dependent prior/regularization on ground-truth in the first sub-minimization problem. This design follows an empirical Bayesian strategy, leading to so-called amortized inference. For matrix inversion in the second sub-problem, we learn a convolutional neural network to approximate the matrix inversion, i.e., the inverse mapping is learned by feeding the input through the learned forward network. Note that training this neural network does not require ground-truth or measurements, i.e., data-independent. Extensive experiments on both synthetic data and real datasets demonstrate the efficiency and accuracy of the proposed method compared with the conventional ADMM solution using inner loops for solving inverse problems. Kai Fan 0002, Lawrence Carin, Katherine A. Heller |
NIPS | 1 |
| 2016 | A Unifying Variational Inference Framework for Hierarchical Graph-Coupled HMM with an Application to Influenza InfectionabstractThe Hierarchical Graph-Coupled Hidden Markov Model (hGCHMM) is a useful tool for tracking and predicting the spread of contagious diseases, such as influenza, by leveraging social contact data collected from individual wearable devices. However, the existing inference algorithms depend on the assumption that the infection rates are small in probability, typically close to 0. The purpose of this paper is to build a unified learning framework for latent infection state estimation for the hGCHMM, regardless of the infection rate and transition function. We derive our algorithm based on a dynamic auto-encoding variational inference scheme, thus potentially generalizing the hGCHMM to models other than those that work on highly contagious diseases. We experimentally compare our approach with previous Gibbs EM algorithms and standard variational method mean-field inference, on both semi-synthetic data and app collected epidemiological and social records. Kai Fan 0002, Chunyuan Li, Katherine A. Heller |
AAAI | 1 |
| 2016 | High-Order Stochastic Gradient Thermostats for Bayesian Learning of Deep ModelsabstractLearning in deep models using Bayesian methods has generated significant attention recently. This is largely because of the feasibility of modern Bayesian methods to yield scalable learning and inference, while maintaining a measure of uncertainty in the model parameters. Stochastic gradient MCMC algorithms (SG-MCMC) are a family of diffusion-based sampling methods for large-scale Bayesian learning. In SG-MCMC, multivariate stochastic gradient thermostats (mSGNHT) augment each parameter of interest, with a momentum and a thermostat variable to maintain stationary distributions as target posterior distributions. As the number of variables in a continuous-time diffusion increases, its numerical approximation error becomes a practical bottleneck, so better use of a numerical integrator is desirable. To this end, we propose use of an efficient symmetric splitting integrator in mSGNHT, instead of the traditional Euler integrator. We demonstrate that the proposed scheme is more accurate, robust, and converges faster. These properties are demonstrated to be desirable in Bayesian deep learning. Extensive experiments on two canonical models and their deep extensions demonstrate that the proposed scheme improves general Bayesian posterior sampling, particularly for deep models. Chunyuan Li, Changyou Chen, Kai Fan 0002, Lawrence Carin |
AAAI | 3 |
| 2016 | Triply Stochastic Variational Inference for Non-linear Beta Process Factor AnalysisabstractWe propose a non-linear extension to factor analysis with beta process priors for improved data representation ability. This non-linear Beta Process Factor Analysis (nBPFA) allows data to be represented as a non-linear transformation of a standard sparse factor decomposition. We develop a scalable variational inference framework, which builds upon the ideas of the variational auto-encoder, by allowing latent variables of the model to be sparse. Our framework can be readily used for real-valued, binary and count data. We show theoretically and with experiments that our training scheme, with additive or multiplicative noise on observations, improves performance and prevents overfitting. We benchmark our algorithms on image, text and collaborative filtering datasets. We demonstrate faster convergence rates and competitive performance compared to standard gradient-based approaches. Kai Fan 0002, Yizhe Zhang 0002, Ricardo Henao, Katherine A. Heller |
ICDM | 1 |
| 2016 | Towards Unifying Hamiltonian Monte Carlo and Slice SamplingabstractWe unify slice sampling and Hamiltonian Monte Carlo (HMC) sampling, demonstrating their connection via the Hamiltonian-Jacobi equation from Hamiltonian mechanics. This insight enables extension of HMC and slice sampling to a broader family of samplers, called Monomial Gamma Samplers (MGS). We provide a theoretical analysis of the mixing performance of such samplers, proving that in the limit of a single parameter, the MGS draws decorrelated samples from the desired target distribution. We further show that as this parameter tends toward this limit, performance gains are achieved at a cost of increasing numerical difficulty and some practical convergence issues. Our theoretical results are validated with synthetic data and real-world applications. Yizhe Zhang 0002, Xiangyu Wang 0006, Changyou Chen, Ricardo Henao, Kai Fan 0002, Lawrence Carin |
NIPS | 5 |
| 2015 | Hierarchical Graph-Coupled HMMs for Heterogeneous Personalized Health DataabstractThe purpose of this study is to leverage modern technology (mobile or web apps) to enrich epidemiology data and infer the transmission of disease. We develop hierarchical Graph-Coupled Hidden Markov Models (hGCHMMs) to simultaneously track the spread of infection in a small cell phone community and capture person-specific infection parameters by leveraging a link prior that incorporates additional covariates. In this paper we investigate two link functions, the beta-exponential link and sigmoid link, both of which allow the development of a principled Bayesian hierarchical framework for disease transmission. The results of our model allow us to predict the probability of infection for each persons on each day, and also to infer personal physical vulnerability and the relevant association with covariates. We demonstrate our approach theoretically and experimentally on both simulation data and real epidemiological records. Kai Fan 0002, Marisa C. Eisenberg, Alison Walsh, Allison Aiello, Katherine A. Heller |
KDD | 1 |
| 2015 | Fast Second Order Stochastic Backpropagation for Variational InferenceabstractWe propose a second-order (Hessian or Hessian-free) based optimization method for variational inference inspired by Gaussian backpropagation, and argue that quasi-Newton optimization can be developed as well. This is accomplished by generalizing the gradient computation in stochastic backpropagation via a reparametrization trick with lower complexity. As an illustrative example, we apply this approach to the problems of Bayesian logistic regression and variational auto-encoder (VAE). Additionally, we compute bounds on the estimator variance of intractable expectations for the family of Lipschitz continuous function. Our method is practical, scalable and model free. We demonstrate our method on several real-world datasets and provide comparisons with other stochastic gradient methods to show substantial enhancement in convergence rates. Kai Fan 0002, Jeffrey M. Beck, James T. Kwok, Katherine A. Heller |
NIPS | 1 |