EDBT 2026 Demo / reviewers in the wild / expert
Tianxing He
dblp:149/0111
· DBLP profile ↗
29ranked-venue papers
10as first author
19since 2021 · last 2026
0009-0008-6383-0307ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 6 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 1 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Visualized Framework for Event Cooperation with Generative AgentsabstractLarge Language Models (LLMs) have revolutionized the simulation of agent societies, enabling autonomous planning, memory formation, and social interactions. However, existing frameworks often overlook systematic evaluations for event organization and lack visualized integration with physically grounded environments, limiting agents' ability to navigate spaces and interact with items realistically. We develop MiniAgentPro, a visualization platform featuring an intuitive map editor for customizing environments and a simulation player with smooth animations. Based on this tool, we introduce a comprehensive test set comprising eight diverse event scenarios with basic and hard variants to assess agents' ability. Evaluations using GPT-4o demonstrate strong performance in basic settings but highlight coordination challenges in hard variants. Shunqiang Mao, Wenchang Gao, Lanlan Qiu, Tianxing He |
AAAI | 5 |
| 2026 | ChatAnime: Towards User-Centered Emotional Support in LLM-based Virtual Character ChatabstractWith the growing popularity of virtual character platforms like Character.AI, users are increasingly turning to role-playing agents for emotional support in daily life.Yet existing research mainly focuses on character consistency in fictional or game-based scenarios, overlooking user-centered interactions such as companionship and psychological support.To bridge this gap, we propose Emotionally Supportive Role-Playing (ESRP), a framework designed to align role-playing with real-world user scenarios and emotional needs.We focus on typical users of these platforms, i.e., anime enthusiasts-including students, office workers, freelancers, and self-employed individuals-and design scenario-based questions that reflect their everyday struggles such as work stress and social loneliness.Through a two-round data collection involving 40 anime fans and 10 Large Language Models (LLMs), we build ChatAnime: the first ESRP dataset with 2,400 human-written and 24,000 LLM-generated responses, supported by over 132,000 finegrained human annotations.We also provide the ESRP evaluation framework featuring 9 fine-grained metrics across three dimensions: basic dialogue, role-playing and emotional support, along with an overall metric for diversity.Experimental results under our evaluation setting show that top-performing LLMs surpass anime fans in role-playing and emotional support, while humans still lead in diversity.The ChatAnime dataset is available at https://github.com/LanlanQiu/ChatAnime. Lanlan Qiu, Xiao Pu 0003, Yeqi Feng, Wenchang Gao, Tianxing He |
ACL (1) | 5 |
| 2026 | Pay for The Second-Best Service: A Game-Theoretic Approach against Dishonest LLM ProvidersabstractThe widespread adoption of Large Language Models (LLMs) through Application Programming Interfaces (APIs) induces a critical vulnerability: the potential for dishonest manipulation by service providers. This manipulation can manifest in various forms, such as secretly substituting a proclaimed high-performance model with a low-cost alternative, or inflating responses with meaningless tokens to increase billing. This work tackles the issue through the lens of algorithmic game theory and mechanism design. We are the first to propose a formal economic model for a realistic user-provider ecosystem, where a user can iteratively delegate T queries to multiple model providers, and providers can engage in a range of strategic behaviors. As our central contribution, we prove that for a continuous strategy space and any ε∈(0,1/2), there exists an approximate incentive-compatible mechanism with an additive approximation ratio of O(T1-ε log T), and a guaranteed quasi-linear second-best user utility. We also prove an impossibility result, stating that no mechanism can guarantee an expected user utility that is asymptotically better than our mechanism. Furthermore, we demonstrate the effectiveness of our mechanism in simulation experiments with real-world API settings. Yuhan Cao 0003, Yixin Tao, Tianxing He |
WWW | 6 |
| 2026 | Toward Modality- and Sampling-Universal Learning Strategies for Accelerating Cardiovascular Imaging: Summary of the CMRxRecon2024 ChallengeabstractCardiovascular health is vital to human well-being, and cardiac magnetic resonance (CMR) imaging is considered the clinical reference standard for diagnosing cardiovascular disease. However, its adoption is hindered by long scan times, complex contrasts, and inconsistent quality. While deep learning methods perform well on specific CMR imaging sequences, they often fail to generalize across modalities and sampling schemes. The lack of benchmarks for high-quality, fast CMR image reconstruction further limits technology comparison and adoption. The CMRxRecon2024 challenge, attracting over 200 teams from 18 countries, addressed these issues with two tasks: generalization to unseen modalities and robustness to diverse undersampling patterns. We introduced the largest public multi-modality CMR raw dataset, an open benchmarking platform, and shared code. Analysis of the best-performing solutions revealed that prompt-based adaptation and enhanced physics-driven consistency enabled strong cross-scenario performance. These findings establish principles for generalizable reconstruction models and advance clinically translatable AI in cardiovascular imaging. Fanwen Wang, Zi Wang 0005, Yan Li 0064, Chen Qin, Shuo Wang 0011, Kunyuan Guo, Mengting Sun, Mingkai Huang, Michael Tänzer, Qirong Li, Yinzhe Wu 0001, Haosen Zhang, Kian Anvari Hamedani, Yuntong Lyu, Longyu Sun, Tianxing He, Lizhen Lan, Qiong Yao, Bingyu Xin, Dimitris N. Metaxas, Narges Razizadeh, Shahabedin Nabavi, George Yiasemis, Jonas Teuwen, Daniel B. Ennis, Zhihao Xue, Ruru Xu, Ilkay Öksüz, Donghang Lyu, Yanxin Huang, Xinrui Guo, Ruqian Hao, Jaykumar H. Patel, Guanke Cai, Binghua Chen, Sha Hua, Zhensen Chen, Qi Dou 0001, Xiahai Zhuang, Wenjia Bai, Harry Qin, He Wang 0016, Claudia Prieto, Michael Markl 0001, Alistair A. Young, Hao Li 0082, Xihong Hu, Lianming Wu, Xiaobo Qu 0001, Guang Yang 0006, Chengyan Wang |
IEEE Trans. Medical Imaging | 21 |
| 2025 | Jailbreak Large Vision-Language Models Through Multi-Modal LinkageabstractWith the rapid advancement of Large Vision-Language Models (VLMs), concerns about their potential misuse and abuse have grown rapidly.Prior research has exposed VLMs' vulnerability to jailbreak attacks, where carefully crafted inputs can lead the model to produce content that violates ethical and legal standards.However, current jailbreak methods often fail against cutting-edge models such as GPT-4o.We attribute this to the overexposure of harmful content and the absence of stealthy malicious guidance.In this work, we introduce a novel jailbreak framework: Multi-Modal Linkage (MML) Attack.Drawing inspiration from cryptography, MML employs an encryption-decryption process across text and image modalities to mitigate the over-exposure of malicious information.To covertly align the model's output with harmful objectives, MML leverages a technique we term evil alignment, framing the attack within the narrative context of a video game development scenario.Extensive experiments validate the effectiveness of MML.Specifically, MML jailbreaks GPT-4o with attack success rates of 99.40% on SafeBench, 98.81% on MM-SafeBench, and 99. Geyuan Zhang, Tianxing He |
ACL (1) | 5 |
| 2025 | Towards Black-Box Membership Inference Attack for Diffusion ModelsabstractGiven the rising popularity of AI-generated art and the associated copyright concerns, identifying whether an artwork was used to train a diffusion model is an important research topic. The work approaches this problem from the membership inference attack (MIA) perspective. We first identify the limitation of applying existing MIA methods for proprietary diffusion models: the required access of internal U-nets. To address the above problem, we introduce a novel membership inference attack method that uses only the image-to-image variation API and operates without access to the model’s internal U-net. Our method is based on the intuition that the model can more easily obtain an unbiased noise prediction estimate for images from the training set. By applying the API multiple times to the target image, averaging the outputs, and comparing the result to the original image, our approach can classify whether a sample was part of the training set. We validate our method using DDIM and Stable Diffusion setups and further extend both our approach and existing algorithms to the Diffusion Transformer architecture. Our experimental results consistently outperform previous methods. Jing Dong 0008, Tianxing He, Jingzhao Zhang |
ICML | 3 |
| 2025 | Learning Syntax Without Planting Trees: Understanding Hierarchical Generalization in TransformersabstractAbstract Transformers trained on natural language data have been shown to exhibit hierarchical generalization without explicitly encoding any structural bias. In this work, we investigate sources of inductive bias in transformer models and their training that could cause such preference for hierarchical generalization. We extensively experiment with transformers trained on five synthetic, controlled datasets using several training objectives and show that, while objectives such as sequence-to-sequence modeling, classification, etc., often fail to lead to hierarchical generalization, the language modeling objective consistently leads to transformers generalizing hierarchically. We then study how different generalization behaviors emerge during the training by conducting pruning experiments that reveal the joint existence of subnetworks within the model implementing different generalizations. Finally, we take a Bayesian perspective to understand transformers’ preference for hierarchical generalization: We establish a correlation between whether transformers generalize hierarchically on a dataset and if the simplest explanation of that dataset is provided by a hierarchical grammar compared to regular grammars exhibiting linear generalization. Overall, our work presents new insights on the origins of hierarchical generalization in transformers and provides a theoretical framework for studying generalization in language models. Kabir Ahuja, Vidhisha Balachandran, Madhur Panwar, Tianxing He, Noah A. Smith, Navin Goyal, Yulia Tsvetkov |
Trans. Assoc. Comput. Linguistics | 4 |
| 2024 | Stumbling Blocks: Stress Testing the Robustness of Machine-Generated Text Detectors Under AttacksabstractYichen Wang, Shangbin Feng, Abe Hou, Xiao Pu, Chao Shen, Xiaoming Liu, Yulia Tsvetkov, Tianxing He. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yichen Wang 0002, Shangbin Feng, Abe Bohan Hou, Xiao Pu 0003, Chao Shen 0001, Xiaoming Liu 0001, Yulia Tsvetkov, Tianxing He |
ACL (1) | 8 |
| 2024 | Knowledge Card: Filling LLMs' Knowledge Gaps with Plug-in Specialized Language ModelsabstractBy design, large language models (LLMs) are static general-purpose models, expensive to retrain or update frequently. As they are increasingly adopted for knowledge-intensive tasks, it becomes evident that these design choices lead to failures to generate factual, relevant, and up-to-date knowledge. To this end, we propose Knowledge Card, a modular framework to plug in new factual and relevant knowledge into general-purpose LLMs. We first introduce knowledge cards---specialized language models trained on corpora from specific domains and sources. Knowledge cards serve as parametric repositories that are selected at inference time to generate background knowledge for the base LLM. We then propose three content selectors to dynamically select and retain information in documents generated by knowledge cards, specifically controlling for relevance, brevity, and factuality of outputs. Finally, we propose two complementary integration approaches to augment the base LLM with the (relevant, factual) knowledge curated from the specialized LMs. Through extensive experiments, we demonstrate that Knowledge Card achieves state-of-the-art performance on six benchmark datasets. Ultimately, Knowledge Card framework enables dynamic synthesis and updates of knowledge from diverse domains. Its modularity will ensure that relevant knowledge can be continuously updated through the collective efforts of the research community. Shangbin Feng, Yuyang Bai, Vidhisha Balachandran, Tianxing He, Yulia Tsvetkov |
ICLR | 5 |
| 2024 | SemStamp: A Semantic Watermark with Paraphrastic Robustness for Text GenerationabstractAbe Hou, Jingyu Zhang, Tianxing He, Yichen Wang, Yung-Sung Chuang, Hongwei Wang, Lingfeng Shen, Benjamin Van Durme, Daniel Khashabi, Yulia Tsvetkov. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Abe Bohan Hou, Tianxing He, Yichen Wang 0002, Yung-Sung Chuang, Lingfeng Shen, Benjamin Van Durme, Daniel Khashabi, Yulia Tsvetkov |
NAACL-HLT | 3 |
| 2024 | KGQuiz: Evaluating the Generalization of Encoded Knowledge in Large Language ModelsabstractLarge language models (LLMs) demonstrate remarkable performance on knowledge-intensive tasks, suggesting that real-world knowledge is encoded in their model parameters. However, besides explorations on a few probing tasks in limited knowledge domains, it is not well understood how to evaluate LLMs' knowledge systematically and how well their knowledge abilities generalize, across a spectrum of knowledge domains and progressively complex task formats. To this end, we propose KGQuiz, a knowledge-intensive benchmark to comprehensively investigate the knowledge generalization abilities of LLMs. KGQuiz is a scalable framework constructed from triplet-based knowledge, which covers three knowledge domains and consists of five tasks with increasing complexity: true-or-false, multiple-choice QA, blank filling, factual editing, and open-ended knowledge generation. To gain a better understanding of LLMs' knowledge abilities and their generalization, we evaluate 10 open-source and black-box LLMs on the KGQuiz benchmark across the five knowledge-intensive tasks and knowledge domains. Extensive experiments demonstrate that LLMs achieve impressive performance in straightforward knowledge QA tasks, while settings and contexts requiring more complex reasoning or employing domain-specific facts still present significant challenges. We envision KGQuiz as a testbed to analyze such nuanced variations in performance across domains and task formats, and ultimately to understand, evaluate, and improve LLMs' knowledge abilities across a wide spectrum of knowledge domains and tasks. Yuyang Bai, Shangbin Feng, Vidhisha Balachandran, Zhaoxuan Tan, Shiqi Lou, Tianxing He, Yulia Tsvetkov |
WWW | 6 |
| 2023 | On the Blind Spots of Model-Based Evaluation Metrics for Text GenerationabstractTianxing He, Jingyu Zhang, Tianle Wang, Sachin Kumar, Kyunghyun Cho, James Glass, Yulia Tsvetkov. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Tianxing He, Tianle Wang 0003, Sachin Kumar 0009, Kyunghyun Cho, James R. Glass, Yulia Tsvetkov |
ACL (1) | 1 |
| 2023 | Learning Time-Invariant Representations for Individual Neurons from Population DynamicsabstractNeurons can display highly variable dynamics. While such variability presumably supports the wide range of behaviors generated by the organism, their gene expressions are relatively stable in the adult brain. This suggests that neuronal activity is a combination of its time-invariant identity and the inputs the neuron receives from the rest of the circuit. Here, we propose a self-supervised learning based method to assign time-invariant representations to individual neurons based on permutation-, and population size-invariant summary of population recordings. We fit dynamical models to neuronal activity to learn a representation by considering the activity of both the individual and the neighboring population. Our self-supervised approach and use of implicit representations enable robust inference against imperfections such as partial overlap of neurons across sessions, trial-to-trial variability, and limited availability of molecular (transcriptomic) labels for downstream supervised tasks. We demonstrate our method on a public multimodal dataset of mouse cortical neuronal activity and transcriptomic labels. We report >35\% improvement in predicting the transcriptomic subclass identity and >20\% improvement in predicting class identity with respect to the state-of-the-art. Lu Mi, Trung Le 0002, Tianxing He, Eli Shlizerman, Uygar Sümbül |
NeurIPS | 3 |
| 2023 | Can Language Models Solve Graph Problems in Natural Language?abstractLarge language models (LLMs) are increasingly adopted for a variety of tasks with implicit graphical structures, such as planning in robotics, multi-hop question answering or knowledge probing, structured commonsense reasoning, and more. While LLMs have advanced the state-of-the-art on these tasks with structure implications, whether LLMs could explicitly process textual descriptions of graphs and structures, map them to grounded conceptual spaces, and perform structured operations remains underexplored. To this end, we propose NLGraph (Natural Language Graph), a comprehensive benchmark of graph-based problem solving designed in natural language. NLGraph contains 29,370 problems, covering eight graph reasoning tasks with varying complexity from simple tasks such as connectivity and shortest path up to complex problems such as maximum flow and simulating graph neural networks. We evaluate LLMs (GPT-3/4) with various prompting approaches on the NLGraph benchmark and find that 1) language models do demonstrate preliminary graph reasoning abilities, 2) the benefit of advanced prompting and in-context learning diminishes on more complex graph problems, while 3) LLMs are also (un)surprisingly brittle in the face of spurious correlations in graph and problem settings. We then propose Build-a-Graph Prompting and Algorithmic Prompting, two instruction-based approaches to enhance LLMs in solving natural language graph problems. Build-a-Graph and Algorithmic prompting improve the performance of LLMs on NLGraph by 3.07% to 16.85% across multiple tasks and settings, while how to solve the most complicated graph reasoning tasks in our setup with language models remains an open research question. Heng Wang 0008, Shangbin Feng, Tianxing He, Zhaoxuan Tan, Xiaochuang Han, Yulia Tsvetkov |
NeurIPS | 3 |
| 2022 | Mining Python fix patterns via analyzing fine-grained source code changes
Tianxing He, Yang Feng 0003, Shaoying Liu, Baowen Xu |
Empir. Softw. Eng. | 2 |
| 2022 | A comprehensive empirical study on bug characteristics of deep learning frameworks
Tianxing He, Zhilong Xia, Yang Feng 0003 |
Inf. Softw. Technol. | 2 |
| 2021 | Analyzing the Forgetting Problem in Pretrain-Finetuning of Open-domain Dialogue Response ModelsabstractTianxing He, Jun Liu, Kyunghyun Cho, Myle Ott, Bing Liu, James Glass, Fuchun Peng. Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2021. Tianxing He, Kyunghyun Cho, Myle Ott, Bing Liu 0024, James R. Glass, Fuchun Peng |
EACL | 1 |
| 2021 | Joint Energy-based Model Training for Better Calibrated Natural Language Understanding ModelsabstractIn this work, we explore joint energy-based model (EBM) training during the finetuning of pretrained text encoders (e.g., Roberta) for natural language understanding (NLU) tasks.Our experiments show that EBM training can help the model reach a better calibration that is competitive to strong baselines, with little or no loss in accuracy.We discuss three variants of energy functions (namely scalar, hidden, and sharp-hidden) that can be defined on top of a text encoder, and compare them in experiments.Due to the discreteness of text data, we adopt noise contrastive estimation (NCE) to train the energy-based model.To make NCE training more effective, we train an autoregressive noise model with the masked language model (MLM) objective. Tianxing He, Bryan McCann, Caiming Xiong, Ehsan Hosseini-Asl |
EACL | 1 |
| 2021 | Exposure Bias versus Self-Recovery: Are Distortions Really Incremental for Autoregressive Text Generation?abstractExposure bias has been regarded as a central problem for auto-regressive language models (LM).It claims that teacher forcing would cause the test-time generation to be incrementally distorted due to the training-generation discrepancy.Although a lot of algorithms have been proposed to avoid teacher forcing and therefore alleviate exposure bias, there is little work showing how serious the exposure bias problem actually is.In this work, we focus on the task of open-ended language generation, propose metrics to quantify the impact of exposure bias in the aspects of quality, diversity, and consistency.Our key intuition is that if we feed ground-truth data prefixes (instead of prefixes generated by the model itself) into the model and ask it to continue the generation, the performance should become much better because the training-generation discrepancy in the prefix is removed.Both automatic and human evaluations are conducted in our experiments.On the contrary to the popular belief in exposure bias, we find that the the distortion induced by the prefix discrepancy is limited, and does not seem to be incremental during the generation.Moreover, our analysis reveals an interesting self-recovery ability of the LM, which we hypothesize to be countering the harmful effects from exposure bias. Tianxing He, Jingzhao Zhang, Zhiming Zhou 0001, James R. Glass |
EMNLP (1) | 1 |
| 2020 | Negative Training for Neural Dialogue Response GenerationabstractAlthough deep learning models have brought tremendous advancements to the field of opendomain dialogue response generation, recent research results have revealed that the trained models have undesirable generation behaviors, such as malicious responses and generic (boring) responses.In this work, we propose a framework named "Negative Training" to minimize such behaviors.Given a trained model, the framework will first find generated samples that exhibit the undesirable behavior, and then use them to feed negative training signals for fine-tuning the model.Our experiments show that negative training can significantly reduce the hit rate of malicious responses, or discourage frequent responses and improve response diversity. Tianxing He, James R. Glass |
ACL | 1 |
| 2020 | An Empirical Study of Transformer-Based Neural Language Model AdaptationabstractWe explore two adaptation approaches of deep Transformer based neural language models (LMs) for automatic speech recognition. The first approach is a pretrain-finetune framework, where we first pretrain a Transformer LM on a large-scale text corpus from scratch and then adapt it to relatively small target domains via finetuning. The second approach is a mixer of dynamically weighted models that are separately trained on source and target domains, aiming to improve simple linear interpolation with dynamic weighting. We compare the two approaches with three baselines - without adaptation, merging data, and simple interpolation - on Switchboard (SWBD) and Wall Street Journal (WSJ). Experiments show that the mixer model generally performs better than baselines and finetuning. Compared with no adaptation, finetuning and the mixer approach obtain up to relative 11.5% and 14.1% WER reductions on SWBD, respectively. The mixer model also outperforms linear interpolation and merging data. On WSJ, the mixer approach achieves a new state-of-the-art WER result. Ke Li 0018, Zhe Liu 0011, Tianxing He, Hongzhao Huang, Fuchun Peng, Daniel Povey, Sanjeev Khudanpur |
ICASSP | 3 |
| 2020 | Why Gradient Clipping Accelerates Training: A Theoretical Justification for Adaptivity
Jingzhao Zhang, Tianxing He, Suvrit Sra, Ali Jadbabaie |
ICLR | 2 |
| 2019 | Detecting Egregious Responses in Neural Sequence-to-sequence Models
Tianxing He, James R. Glass |
ICLR (Poster) | 1 |
| 2019 | From Data Quality to Model Quality: An Exploratory Study on Deep LearningabstractIn the field of deep learning, people strive to construct high-quality deep neural networks (DNNs) to improve the accuracy of predicting. As well known, the quality of training data have great impacts on the quality of DNN models, since all the DNN models are obtained by training using these training data. However, there is not any reported systematic study on how the quality of training data affects the quality of DNN model. To study the relationships between data quality and model quality, we mainly consider four aspects of data quality including Skewed Classes, Sample Complexity, Label Quality, and Noisy Data in this paper. We design experiments on MNIST and Cifar-10, and attempt to find out the influences of four aspects on the quality of DNN models. Pearson correlation coefficient and Spearman correlation coefficient are utilized to evaluate such influences. Experimental results show that all the four aspects of data quality have significant impacts on the quality of DNN models. It means that the decrease of data quality in these four aspects will reduce the accuracy of the DNN models. Tianxing He, Shengcheng Yu, Ziyuan Wang 0001, Jieqiong Li, Zhenyu Chen 0001 |
Internetware | 1 |
| 2016 | Exploiting LSTM structure in deep neural networks for speech recognitionabstractThe CD-DNN-HMM system has became the state-of-art system for large vocabulary continuous speech recognition (LVCSR) tasks, in which deep neural networks (DNN) plays a key role. However, DNN training suffers from the vanishing gradient problem, limiting training of deep models. In this work, we address this problem by incorporating the successful long-short term memory (LSTM) structure, which has been proposed to help recurrent neural network (RNN) to remember long term dependencies, into DNN. Also, we propose a generalized formulation of the LSTM block, which we name general LSTM(GLSTM). In our experiments, it is shown that our proposed (G)LSTM-DNN scales well with more layers, and achieves 8.2% relative word error rate reduction on the 2000-hour Switchboard data set. Tianxing He, Jasha Droppo |
ICASSP | 1 |
| 2015 | Recurrent neural network language model with structured word embeddings for speech recognitionabstractDue to effective word context encoding and long-term context preserving, recurrent neural network language model (RNNLM) has attracted great interest by showing better performance over back-off n-gram models and feed-forward neural network language models (FNNLM). However, it still has the difficulty of modelling words of very low frequency in training data. To address this issue, a new framework of structured word embedding is introduced to RNNLM, where both input and target word embeddings are factorized into weighted sum of the corresponding sub-word embeddings. The framework is instantiated for Chinese, where characters can be naturally used as the sub-word units. Experiments on a Chinese twitter LVCSR task showed that the proposed approach effectively outperformed the standard RNNLM, yielding a relative PPL improvement of 8:8% and an absolute 0:59% CER improvement in N-Best re-scoring. Tianxing He, Xu Xiang, Yanmin Qian, Kai Yu 0004 |
ICASSP | 1 |
| 2015 | Automatic model redundancy reduction for fast back-propagation for deep neural networks in speech recognitionabstractAlthough deep neural networks (DNNs) have achieved great performance gain, the immense computational cost of DNN model training has become a major block to utilize massive speech data for DNN training. Previous research on DNN training acceleration mostly focussed on hardware-based parallelization. In this paper, node pruning and arc restructuring are proposed to explore model redundancy after a novel lightly discriminative pretraining process. With some measures of node/arc importance, model redundancies are automatically removed to form a much more compact DNN. This significantly accelerates the subsequent back-propagation (BP) training process. Model redundancy reduction can be combined with multiple GPU parallelization to achieve further acceleration. Experiments showed that the combined acceleration framework can achieve about 85% model size reduction and over 4.2 times speed-up factor for BP training on 2 GPUs, at no loss of recognition accuracy. Yanmin Qian, Tianxing He, Kai Yu 0004 |
IJCNN | 2 |
| 2015 | Paragraph vector based topic model for language model adaptation
Wengong Jin, Tianxing He, Yanmin Qian, Kai Yu 0004 |
INTERSPEECH | 2 |
| 2014 | Reshaping deep neural network for fast decoding by node-pruningabstractAlthough deep neural networks (DNN) has achieved significant accuracy improvements in speech recognition, it is computationally expensive to deploy large-scale DNN in decoding due to huge number of parameters. Weights truncation and decomposition methods have been proposed to speed up decoding by exploiting the sparseness of DNN. This paper summarizes different approaches of restructuring DNN and proposes a new node pruning approach to reshape DNN for fast decoding. In this approach, hidden nodes of a fully trained DNN are pruned with certain importance function and the reshaped DNN is retuned using back-propagation. The approach requires no modification on code and can directly save computational costs during decoding. Furthermore, it is complementary to weight decomposition methods. Experiments on a switchboard task shows that, by using the proposed node-pruning approach, DNN complexity can be reduced to 37.9%. The complexity can be further reduced to 12.3% without accuracy loss when node-pruning is combined with weight decomposition. Tianxing He, Yuchen Fan 0001, Yanmin Qian, Tian Tan 0002, Kai Yu 0004 |
ICASSP | 1 |