Bo Xu 0002

dblp:26/1194-2 · DBLP profile ↗
← Back
358ranked-venue papers
5as first author
102since 2021 · last 2026
0000-0002-1111-1529ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 251 · 4 first-author · 77 since 2021Graphics, computer vision, multimedia, augmented reality and games · 203 · 5 first-author · 45 since 2021Databases, data management, data science and information retrieval · 10 · 1 since 2021Human-computer interaction and ubiquitous computing · 10Applied, interdisciplinary, general and emerging computing · 10 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 Speech-Aware Long Context Pruning and Integration for Contextualized Automatic Speech Recognition
abstract
Automatic speech recognition (ASR) systems have achieved remarkable performance in common conditions but often struggle to leverage long-context information in contextualized scenarios that require domain-specific knowledge, such as conference presentations. This challenge arises primarily due to constrained model context windows and the sparsity of relevant information within extensive contextual noise. To solve this, we propose the SAP^2 method, a novel framework that dynamically prunes and integrates relevant contextual keywords in two stages. Specifically, each stage leverages our proposed Speech-Driven Attention-based Pooling mechanism, enabling efficient compression of context embeddings while preserving speech-salient information. Experimental results demonstrate state-of-the-art performance of SAP^2 on the SlideSpeech and LibriSpeech datasets, achieving word error rates (WER) of 7.71% and 1.12%, respectively. On SlideSpeech, our method notably reduces biased keyword error rates (B-WER) by 41.1% compared to non-contextual baselines. SAP^2 also exhibits robust scalability, consistently maintaining performance under extensive contextual input conditions on both datasets.
Yiming Rong, Deyang Jiang, Bo Xu 0002
AAAI8
2026 MrCoM: A Meta-Regularized World-Model Generalizing Across Multi-Scenarios
abstract
Model-based reinforcement learning (MBRL) is a crucial approach to enhance the generalization capabilities and improve the sample efficiency of RL algorithms. However, current MBRL methods focus primarily on building world models for single tasks and rarely address generalization across different scenarios. Building on the insight that dynamics within the same simulation engine share inherent properties, we attempt to construct a unified world model capable of generalizing across different scenarios, named Meta-Regularized Contextual World-Model (MrCoM). This method first decomposes the latent state space into various components based on the dynamic characteristics, thereby enhancing the accuracy of world-model prediction. Further, MrCoM adopts meta-state regularization to extract unified representation of scenario-relevant information, and meta-value regularization to align world-model optimization with policy learning across diverse scenario objectives. We theoretically analyze the generalization error upper bound of MrCoM in multi-scenario settings. We systematically evaluate our algorithm's generalization ability across diverse scenarios, demonstrating significantly better performance than previous state-of-the-art methods.
Xuantang Xiong, Ni Mu, Runpeng Xie, Senhao Yang, Lexiang Wang, Yao Luan 0001, Siyuan Li 0003, Yiqin Yang, Bo Xu 0002
AAAI11
2026 TinyChemVL: Advancing Chemical Vision-Language Models via Efficient Visual Token Reduction and Complex Reaction Tasks
abstract
While Vision Language Models (VLMs) have demonstrated remarkable capabilities in general visual understanding, their application in the chemical domain has been limited, with previous works predominantly focusing on text and thus overlooking critical visual information, such as molecular structures. Current approaches that directly adopt standard VLMs for chemical tasks suffer from two primary issues: (i) computational inefficiency of processing entire chemical images with non-informative backgrounds. (ii) a narrow scope on molecular-level tasks that restricts progress in chemical reasoning. In this work, we propose TinyChemVL, an efficient and powerful chemical VLM that leverages visual token reduction and reaction-level tasks to improve model efficiency and reasoning capacity. Also, we propose ChemRxn-V, a reaction-level benchmark for assessing vision-based reaction recognition and prediction tasks. Directly predicting reaction products from molecular images poses a non-trivial challenge, as it requires models to integrate both recognition and reasoning capacities. Our results demonstrate that, with only 4B parameters, TinyChemVL achieves superior performance on both molecular and reaction tasks, while also demonstrating faster inference and training speeds compared to existing models. Notably, TinyChemVL outperforms ChemVLM while utilizing only 1/16th of the visual tokens. This work builds efficient yet powerful VLMs for chemical domains by co-designing model architecture and task complexity.
Xuanle Zhao, Shuxin Zeng, Xinyuan Cai, Duzhen Zhang, Xiuyi Chen, Bo Xu 0002
AAAI7
2026 FastGaze: An efficient and flexible model for human scanpath prediction
Jiahong Zhang, Hongjuan Pei, Tianxiang Hu, Richard D. Shang, Bo Xu 0002, Guoqi Li 0002
Expert Syst. Appl.7
2026 Critical-state-accelerated RNN-based reinforcement learning
Wangzi Yao, Bo Xu 0002, Tielin Zhang
Neurocomputing3
2026 Enhancing robustness of spiking neural networks through retina-like coding and memory-based neurons
Jiahong Zhang, Man Yao, Peng Zhou 0017, Bo Xu 0002, Guoqi Li 0002
Neural Networks6
2025 Spike2Former: Efficient Spiking Transformer for High-performance Image Segmentation
abstract
Spiking Neural Networks (SNNs) have a low-power advantage but perform poorly in image segmentation tasks. The reason is that directly converting neural networks with complex architectural designs for segmentation tasks into spiking versions leads to performance degradation and non-convergence. To address this challenge, we first identify the modules in the architecture design that lead to the severe reduction in spike firing, make targeted improvements, and propose Spike2Former architecture. Second, we propose normalized integer spiking neurons to solve the training stability problem of SNNs with complex architectures. We set a new state-of-the-art for SNNs in various semantic segmentation datasets, with a significant improvement of +12.7% mIoU and 5.0x efficiency on ADE20K, +14.3% mIoU and 5.2x efficiency on VOC2012, and +9.1% mIoU and 6.6x efficiency on CityScapes.
Zhenxin Lei, Man Yao, Xinhao Luo, Yanye Lu, Bo Xu 0002, Guoqi Li 0002
AAAI6
2025 Efficient 3D Recognition with Event-driven Spike Sparse Convolution
abstract
Spiking Neural Networks (SNNs) provide an energy-efficient way to extract 3D spatio-temporal features. Point clouds are sparse 3D spatial data, which suggests that SNNs should be well-suited for processing them. However, when applying SNNs to point clouds, they often exhibit limited performance and fewer application scenarios. We attribute this to inappropriate preprocessing and feature extraction methods. To address this issue, we first introduce the Spike Voxel Coding (SVC) scheme, which encodes the 3D point clouds into a sparse spike train space, reducing the storage requirements and saving time on point cloud preprocessing. Then, we propose a Spike Sparse Convolution (SSC) model for efficiently extracting 3D sparse point cloud features. Combining SVC and SSC, we design an efficient 3D SNN backbone (E-3DSNN), which is friendly with neuromorphic hardware. For instance, SSC can be implemented on neuromorphic chips with only minor modifications to the addressing function of vanilla spike convolution. Experiments on ModelNet40, KITTI, and Semantic KITTI datasets demonstrate that E-3DSNN achieves state-of-the-art (SOTA) results with remarkable efficiency. Notably, our E-3DSNN (1.87M) obtained 91.7% top-1 accuracy on ModelNet40, surpassing the current best SNN baselines (14.3M) by 3.0%. To our best knowledge, it is the first direct training 3D SNN backbone that can simultaneously handle various 3D computer vision tasks (e.g., classification, detection, and segmentation) with an event-driven nature.
Xuerui Qiu, Man Yao, Jieyuan Zhang, Yuhong Chou, Shibo Zhou, Bo Xu 0002, Guoqi Li 0002
AAAI7
2025 MMDEND: Dendrite-Inspired Multi-Branch Multi-Compartment Parallel Spiking Neuron for Sequence Modeling
abstract
Kexin Wang, Yuhong Chou, Di Shang, Shijie Mei, Jiahong Zhang, Yanbin Huang, Man Yao, Bo Xu, Guoqi Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yuhong Chou, Richard D. Shang, Shijie Mei 0001, Jiahong Zhang, Yanbin Huang, Man Yao, Bo Xu 0002, Guoqi Li 0002
ACL (1)8
2025 DPMT: Dual Process Multi-scale Theory of Mind Framework for Real-time Human-AI Collaboration
Xiyun Li, Yining Ding, Yuhua Jiang, Runpeng Xie, Yuanhua Ni, Yiqin Yang, Bo Xu 0002
CogSci9
2025 Episodic Novelty Through Temporal Distance
abstract
Exploration in sparse reward environments remains a significant challenge in reinforcement learning, particularly in Contextual Markov Decision Processes (CMDPs), where environments differ across episodes. Existing episodic intrinsic motivation methods for CMDPs primarily rely on count-based approaches, which are ineffective in large state spaces, or on similarity-based methods that lack appropriate metrics for state comparison. To address these shortcomings, we propose Episodic Novelty Through Temporal Distance (ETD), a novel approach that introduces temporal distance as a robust metric for state similarity and intrinsic reward computation. By employing contrastive learning, ETD accurately estimates temporal distances and derives intrinsic rewards based on the novelty of states within the current episode. Extensive experiments on various benchmark tasks demonstrate that ETD significantly outperforms state-of-the-art methods, highlighting its effectiveness in enhancing exploration in sparse reward CMDPs.
Yuhua Jiang, Qihan Liu, Yiqin Yang, Xiaoteng Ma, Dianyu Zhong, Hao Hu 0006, Jun Yang 0028, Bin Liang 0001, Bo Xu 0002, Chongjie Zhang, Qianchuan Zhao
ICLR9
2025 Fewer May Be Better: Enhancing Offline Reinforcement Learning with Reduced Dataset
abstract
Research in offline reinforcement learning (RL) marks a paradigm shift in RL. However, a critical yet under-investigated aspect of offline RL is determining the subset of the offline dataset, which is used to improve algorithm performance while accelerating algorithm training. Moreover, the size of reduced datasets can uncover the requisite offline data volume essential for addressing analogous challenges. Based on the above considerations, we propose identifying Reduced Datasets for Offline RL (ReDOR) by formulating it as a gradient approximation optimization problem. We prove that the common actor-critic framework in reinforcement learning can be transformed into a submodular objective. This insight enables us to construct a subset by adopting the orthogonal matching pursuit (OMP). Specifically, we have made several critical modifications to OMP to enable successful adaptation with Offline RL algorithms. The experimental results indicate that the data subsets constructed by the ReDOR can significantly improve algorithm performance with low computational complexity.
Yiqin Yang, Quanwei Wang, Chenghao Li 0002, Hao Hu 0006, Chengjie Wu, Yuhua Jiang, Dianyu Zhong, Ziyou Zhang, Qianchuan Zhao, Chongjie Zhang, Bo Xu 0002
ICLR11
2025 Integrate-and-Fire Compressor: Learning to Compress Context for LLMs Adaptively
abstract
Large language models (LLMs) face significant challenges in long-context modeling due to increased inference costs, higher latency, and performance degradation caused by information redundancy. Context compression offers a promising solution, but existing methods often rely on fixed strategies that don’t adapt to variations in information density. To address this, we propose the Integrate-and-Fire Compressor (IFC), an adaptive method inspired by the neural integrate-and-fire mechanism. IFC assesses token importance more accurately and dynamically adjusts compression, producing compact and effective context representations for LLMs, even at high compression rates. Experiments on three open-domain QA datasets show that IFC consistently outperforms existing baselines in both compression rate and task performance, highlighting its potential to improve the efficiency and effectiveness of LLMs in handling long contexts.
Xiyun Li, Minglun Han, Bo Xu 0002
ICME6
2025 CLARIFY: Contrastive Preference Reinforcement Learning for Untangling Ambiguous Queries
abstract
Preference-based reinforcement learning (PbRL) bypasses explicit reward engineering by inferring reward functions from human preference comparisons, enabling better alignment with human intentions. However, humans often struggle to label a clear preference between similar segments, reducing label efficiency and limiting PbRL’s real-world applicability. To address this, we propose an offline PbRL method: Contrastive LeArning for ResolvIng Ambiguous Feedback (CLARIFY), which learns a trajectory embedding space that incorporates preference information, ensuring clearly distinguished segments are spaced apart, thus facilitating the selection of more unambiguous queries. Extensive experiments demonstrate that CLARIFY outperforms baselines in both non-ideal teachers and real human feedback settings. Our approach not only selects more distinguished queries but also learns meaningful trajectory embeddings.
Ni Mu, Hao Hu 0006, Yiqin Yang, Bo Xu 0002, Qing-Shan Jia
ICML5
2025 β-DQN: Improving Deep Q-Learning By Evolving the Behavior
Hongming Zhang 0003, Fengshuo Bai, Chenjun Xiao, Chao Gao 0012, Bo Xu 0002, Martin Müller 0003
AAMAS5
2025 Integrating Large Language Models with Reinforcement Learning for Generalization in Strategic Card Games
Wannian Xia, Yali Du 0001, Bo Xu 0002
AAMAS5
2025 S-EPOA: Overcoming the Indistinguishability of Segments with Skill-Driven Preference-Based Reinforcement Learning
abstract
Preference-based reinforcement learning (PbRL) stands out by utilizing human preferences as a direct reward signal, eliminating the need for intricate reward engineering. However, despite its potential, traditional PbRL methods are often constrained by the indistinguishability of segments, which impedes the learning process. In this paper, we introduce Skill-Enhanced Preference Optimization Algorithm (S-EPOA), which addresses the segment indistinguishability issue by integrating skill mechanisms into the preference learning framework. Specifically, we first conduct the unsupervised pretraining to learn useful skills. Then, we propose a novel query selection mechanism to balance the information gain and distinguishability over the learned skill space. Experimental results on a range of tasks, including robotic manipulation and locomotion, demonstrate that S-EPOA significantly outperforms conventional PbRL methods in terms of both robustness and learning efficiency. The results highlight the effectiveness of skill-driven learning in overcoming the challenges posed by segment indistinguishability.
Ni Mu, Yao Luan 0001, Yiqin Yang, Bo Xu 0002, Qing-Shan Jia
IJCAI4
2025 Empowering LLMs with Parameterized Skills for Adversarial Long-Horizon Planning
abstract
Recent advancements in Large Language Models(LLMs) have led to the development of LLM-based AI agents. A key challenge is the creation of agents that can effectively ground themselves in complex, adversarial long-horizon environments. Existing methods mainly focus on (1) using LLMs as policies to interact with the environment through generating low-level feasible actions, and (2) utilizing LLMs to generate high-level tasks or language guides to stimulate action generation. However, the former struggles to generate reliable actions, while the latter relies heavily on expert experience to translate high-level tasks into specific action sequences. To address these challenges, we introduce the Plan with Language, Act with Parameter (PLAP) planning framework that facilitates the grounding of LLM-based agents in long-horizon environments. The PLAP method comprises three key components: (1) a skill library containing environment-specific parameterized skills, (2) a skill planner powered by LLMs, and (3) a skill executor converting the parameterized skills into executable action sequences. We implement PLAP in MicroRTS, a long-horizon real-time strategy game that provides an unfamiliar and challenging environment for LLMs. The experimental results demonstrate the effectiveness of PLAP. In particular, GPT-4o-driven PLAP in a zero-shot setting outperforms 80 percent of baseline agents, and Qwen2-72B-driven PLAP, with carefully crafted few-shot examples, surpasses the top-tier scripted agent, CoacAI. Additionally, we design comprehensive evaluation metrics and test 6 closed-source and 2 open-source LLMs within the PLAP framework, ultimately releasing an LLM leaderboard ranking long-horizon skill planning ability.
Sijia Cui, Aiyao He, Yanna Wang, Bo Xu 0002
IJCNN5
2025 STAR: Efficient Preference-based Reinforcement Learning via Dual Regularization
abstract
Preference-based reinforcement learning (PbRL) bypasses complex reward engineering by learning from human feedback. However, due to the high cost of obtaining feedback, PbRL typically relies on a limited set of preference-labeled samples. This data scarcity introduces two key inefficiencies: (1) the reward model overfits to the limited feedback, leading to poor generalization to unseen samples, and (2) the agent exploits the learned reward model, exacerbating overestimation of action values in temporal difference (TD) learning. To address these issues, we propose STAR, an efficient PbRL method that integrates preference margin regularization and policy regularization. Preference margin regularization mitigates overfitting by introducing a bounded margin in reward optimization, preventing excessive bias toward specific feedback. Policy regularization bootstraps a conservative estimate $\widehat{Q}$ from well-supported state-action pairs in the replay memory, reducing overestimation during policy learning. Experimental results show that STAR improves feedback efficiency, achieving 34.8\% higher performance in online settings and 29.7\% in offline settings compared to state-of-the-art methods. Ablation studies confirm that STAR facilitates more robust reward and value function learning. The videos of this project are released at https://sites.google.com/view/pbrl-star.
Fengshuo Bai, Rui Zhao 0001, Hongming Zhang 0003, Sijia Cui, Shao Zhang, Bo Xu 0002, Ying Wen 0001, Yaodong Yang 0001
NeurIPS6
2025 STAIR: Addressing Stage Misalignment through Temporal-Aligned Preference Reinforcement Learning
abstract
Preference-based reinforcement learning (PbRL) bypasses complex reward engineering by learning rewards directly from human preferences, enabling better alignment with human intentions. However, its effectiveness in multi-stage tasks, where agents sequentially perform sub-tasks (e.g., navigation, grasping), is limited by **stage misalignment**: Comparing segments from mismatched stages, such as movement versus manipulation, results in uninformative feedback, thus hindering policy learning. In this paper, we validate the stage misalignment issue through theoretical analysis and empirical experiments. To address this issue, we propose **ST**age-**A**l**I**gned **R**eward learning (STAIR), which first learns a stage approximation based on temporal distance, then prioritizes comparisons within the same stage. Temporal distance is learned via contrastive learning, which groups temporally close states into coherent stages, without predefined task knowledge, and adapts dynamically to policy changes. Extensive experiments demonstrate STAIR's superiority in multi-stage tasks and competitive performance in single-stage tasks. Furthermore, human studies show that stages approximated by STAIR are consistent with human cognition, confirming its effectiveness in mitigating stage misalignment.
Yao Luan 0001, Ni Mu, Yiqin Yang, Bo Xu 0002, Qing-Shan Jia
NeurIPS4
2025 DAIL: Beyond Task Ambiguity for Language-Conditioned Reinforcement Learning
abstract
Comprehending natural language and following human instructions are critical capabilities for intelligent agents. However, the flexibility of linguistic instructions induces substantial ambiguity across language-conditioned tasks, severely degrading algorithmic performance. To address these limitations, we present a novel method named DAIL (Distributional Aligned Learning), featuring two key components: distributional policy and semantic alignment. Specifically, we provide theoretical results that the value distribution estimation mechanism enhances task differentiability. Meanwhile, the semantic alignment module captures the correspondence between trajectories and linguistic instructions. Extensive experimental results on both structured and visual observation benchmarks demonstrate that DAIL effectively resolves instruction ambiguities, achieving superior performance to baseline methods. Our implementation is available at https://github.com/RunpengXie/Distributional-Aligned-Learning.
Runpeng Xie, Quanwei Wang, Hao Hu 0006, Zherui Zhou, Ni Mu, Xiyun Li, Yiqin Yang, Qianchuan Zhao, Bo Xu 0002
NeurIPS10
2025 Self-Verifying Reflection Helps Transformers with CoT Reasoning
abstract
Advanced large language models (LLMs) frequently reflect in reasoning chain-of-thoughts (CoTs), where they self-verify the correctness of current solutions and explore alternatives. However, given recent findings that LLMs detect limited errors in CoTs, how reflection contributes to empirical improvements remains unclear. To analyze this issue, in this paper, we present a minimalistic reasoning framework to support basic self-verifying reflection for small transformers without natural language, which ensures analytic clarity and reduces the cost of comprehensive experiments. Theoretically, we prove that self-verifying reflection guarantees improvements if verification errors are properly bounded. Experimentally, we show that tiny transformers, with only a few million parameters, benefit from self-verification in both training and reflective execution, reaching remarkable LLM-level performance in integer multiplication and Sudoku. Similar to LLM results, we find that reinforcement learning (RL) improves in-distribution performance and incentivizes frequent reflection for tiny transformers, yet RL mainly optimizes shallow statistical patterns without faithfully reducing verification errors. In conclusion, integrating generative transformers with discriminative verification inherently facilitates CoT reasoning, regardless of scaling and natural language.
Zhongwei Yu, Wannian Xia, Bo Xu 0002, Haifeng Zhang 0002, Yali Du 0001, Jun Wang 0012
NeurIPS4
2025 Enabling scale and rotation invariance in convolutional neural networks with retina like transformation
Jiahong Zhang, Guoqi Li 0002, Qiaoyi Su, Lihong Cao, Yonghong Tian 0001, Bo Xu 0002
Neural Networks6
2025 Scaling Spike-Driven Transformer With Efficient Spike Firing Approximation Training
abstract
The ambition of brain-inspired Spiking Neural Networks (SNNs) is to become a low-power alternative to traditional Artificial Neural Networks (ANNs). This work addresses two major challenges in realizing this vision: the performance gap between SNNs and ANNs, and the high training costs of SNNs. We identify intrinsic flaws in spiking neurons caused by binary firing mechanisms and propose a Spike Firing Approximation (SFA) method using integer training and spike-driven inference. This optimizes the spike firing pattern of spiking neurons, enhancing efficient training, reducing power consumption, improving performance, enabling easier scaling, and better utilizing neuromorphic chips. We also develop an efficient spike-driven Transformer architecture and a spike-masked autoencoder to prevent performance degradation during SNN scaling. On ImageNet-1k, we achieve state-of-the-art top-1 accuracy of 78.5%, 79.8%, 84.0%, and 86.2% with models containing 10 M, 19 M, 83 M, and 173 M parameters, respectively. For instance, the 10 M model outperforms the best existing SNN by 7.2% on ImageNet, with training time acceleration and inference energy efficiency improved by 4.5× and 3.9×, respectively. We validate the effectiveness and efficiency of the proposed method across various tasks, including object detection, semantic segmentation, and neuromorphic vision tasks. This work enables SNNs to match ANN performance while maintaining the low-power advantage, marking a significant step towards SNNs as a general visual backbone.
Man Yao, Xuerui Qiu, Tianxiang Hu, Yuhong Chou, Keyu Tian, Jianxing Liao, Luziwei Leng, Bo Xu 0002, Guoqi Li 0002
IEEE Trans. Pattern Anal. Mach. Intell.9
2024 SpikeVoice: High-Quality Text-to-Speech Via Efficient Spiking Neural Network
abstract
Brain-inspired Spiking Neural Network (SNN) has demonstrated its effectiveness and efficiency in vision, natural language, and speech understanding tasks, indicating their capacity to “see”, “listen”, and “read”. In this paper, we design SpikeVoice, which performs high-quality Text-To-Speech (TTS) via SNN, to explore the potential of SNN to “speak”. A major obstacle to using SNN for such generative tasks lies in the demand for models to grasp long-term dependencies. The serial nature of spiking neurons, however, leads to the invisibility of information at future spiking time steps, limiting SNN models to capture sequence dependencies solely within the same time step. We term this phenomenon “partial-time dependency”. To address this issue, we introduce Spiking Temporal-Sequential Attention (STSA) in the SpikeVoice. To the best of our knowledge, SpikeVoice is the first TTS work in the SNN field. We perform experiments using four well-established datasets that cover both Chinese and English languages, encompassing scenarios with both single-speaker and multi-speaker configurations. The results demonstrate that SpikeVoice can achieve results comparable to Artificial Neural Networks (ANN) with only 10.5% energy consumption of ANN. Both our demo and code are available as supplementary material.
Jiahong Zhang, Yong Ren 0006, Man Yao, Richard D. Shang, Bo Xu 0002, Guoqi Li 0002
ACL (1)6
2024 EP-Net: Automatic Artery/Vein Classification With Evidential Probability Map
abstract
Abnormal retinal vascular morphology is commonly associated with cardiac, cerebrovascular, and systemic diseases. Hence, automated artery/vein(A/V) classification is crucial for the diagnosis of ophthalmic and systemic diseases. However, existing methods still face limitations in A/V classification and are prone to errors especially in microvessels and in noisy backgrounds. To alleviate these problems, this paper proposes an Evidence Probability Network (EP-Net) to achieve accurate A/V classification. Concretely, the multi-scale feature module in the EP-Net learns various vessel features, and the evidence probability module measures uncertainty and evidence for each pixel to overcome misclassification because of over-/under-confidence. Experiments on two public fundus image datasets demonstrate the superiority of the proposed EP-Net over state-of-the-art A/V classification methods.
Rongchang Zhao, Bo Xu 0002, Xiaoliang Jia, Jin Liu 0012
BIBM2
2024 Integer-Valued Training and Spike-Driven Inference Spiking Neural Network for High-Performance and Energy-Efficient Object Detection
Xinhao Luo, Man Yao, Yuhong Chou, Bo Xu 0002, Guoqi Li 0002
ECCV (32)4
2024 A New Pre-Training Paradigm for Offline Multi-Agent Reinforcement Learning with Suboptimal Data
abstract
Offline multi-agent reinforcement learning (MARL) with pre-training paradigm, which uses a large quantity of trajectories for offline pre-training and online deployment, has become fashionable lately. While performing well on various tasks, conventional pre-trained decision-making models based on imitation learning typically require many expert trajectories or demonstrations, which limits the development of pre-trained policies in multi-agent case. To address this problem, we propose a new setting, where a multi-agent policy is pre-trained offline using suboptimal (non-expert) data and then tested online with the expectation of high rewards. In this practical setting inspired by contrastive learning , we propose YANHUI, a simple yet effective framework utilizing a well-designed reward contrast function for multi-agent policy representation learning from a dataset including various reward-level data instead of just expert trajectories. Furthermore, we enrich the multi-agent policy pre-training with mixture-of-experts to dynamically represent it. With the same quantity of offline StarCraft Multi-Agent Challenge datasets, YANHUI achieves significant improvements over offline MARL baselines. In particular, our method surprisingly competes in performance with earlier state-of-the-art approaches, even with 10% of the expert data used by other baselines and the rest replaced by poor data.
Linghui Meng 0001, Dengpeng Xing, Bo Xu 0002
ICASSP4
2024 ViLaS: Exploring the Effects of Vision and Language Context in Automatic Speech Recognition
abstract
Enhancing automatic speech recognition (ASR) performance by leveraging additional multimodal information has shown promising results in previous studies. However, most of these works have primarily focused on utilizing visual cues derived from human lip motions. In fact, context-dependent visual and linguistic cues can also benefit in many scenarios. In this paper, we first propose ViLaS (Vision and Language into Automatic Speech Recognition), a novel multimodal ASR model based on the continuous integrate-and-fire (CIF) mechanism, which can integrate visual and textual context simultaneously or separately, to facilitate speech recognition. Next, we introduce an effective training strategy that improves performance in modal-incomplete test scenarios. Then, to explore the effects of integrating vision and language, we create VSDial, a multimodal ASR dataset with multimodal context cues in both Chinese and English versions. Finally, empirical results are reported on the public Flickr8K and self-constructed VSDial datasets. We explore various cross-modal fusion schemes, analyze fine-grained cross-modal alignment on VSDial, and provide insights into the effects of integrating multimodal information on speech recognition.
Ziyi Ni, Minglun Han, Linghui Meng 0001, Jing Shi 0003, Bo Xu 0002
ICASSP7
2024 MaDE: Multi-Scale Decision Enhancement for Multi-Agent Reinforcement Learning
abstract
In the domain of multi-agent reinforcement learning (MARL), the limited information availability, complex agent interactions, and individual capabilities among agents often pose a bottleneck for effective decision-making. Previous studies frequently fall short due to insufficient consideration of these multi-dimensional challenges. Thus, this paper introduces a novel methodology, termed Multi-scale Decision Enhancement (MaDE), anchored by a dual-wise bisimulation framework for pre-training agent encoders. The MaDE framework aims to facilitate decision-making across three pivotal dimensions: macroscale awareness, mesoscale coordination, and microscale insight. At the macro level, a pretrained global encoder captures a situational awareness map to guide overall strategies. At the meso level, specialized local encoders generate cluster-based representations to promote inter-agent cooperation. At the micro level, individual agents focus on the accurate decision-making process. Empirical evaluations validate that MaDE outperforms state-of-the-art methods in various multi-agent environments, which shows the potential to tackle the intricate challenges of MARL, enabling agents to make more informed, coordinated, and adaptive decisions. Code is available at https://github.com/paper2023/MaDE.
Jingqing Ruan, Runpeng Xie, Xuantang Xiong, Bo Xu 0002
ICASSP5
2024 UNeC: Unsupervised Exploring In Controllable Space
abstract
In unsupervised reinforcement learning, agents traverse a reward-free environment, aiming for rapid generalisation to subsequent tasks. This strategy offers a compelling resolution to the challenges of sample efficiency. Nevertheless, environments are frequently saturated with excessive information. The omnipresence of uncontrollable states, akin to noise, can impede the efficacy of unsupervised exploration. To address this issue, we introduce the UNeC framework, standing for UNsupervised Exploring in Controllable Space. This approach leverages the intricate dependencies between states and actions, establishing constraints on controllable state representation. Building on this foundation, we employ particle entropy to evaluate states, adeptly directing agent exploration. Empirical evaluations on the unsupervised RL benchmark affirm the superiority of our method, with a particular emphasis on its dominance in noise-affected scenarios.
Xuantang Xiong, Linghui Meng 0001, Jingqing Ruan, Bo Xu 0002
ICASSP5
2024 Double Reverse Regularization Network Based on Self-Knowledge Distillation for SAR Object Classification
abstract
In current synthetic aperture radar (SAR) object classification, one of the major challenges is the severe overfitting issue due to the limited dataset (few-shot) and noisy data. Considering the advantages of knowledge distillation as a learned label smoothing regularization, this paper proposes a novel Double Reverse Regularization Network based on Self-Knowledge Distillation (DRRNet-SKD). Specifically, through exploring the effect of distillation weight on the process of distillation, we are inspired to adopt the double reverse thought to implement an effective regularization network by combining offline and online distillation in a complementary way. Then, the Adaptive Weight Assignment (AWA) module is designed to adaptively assign two reverse-changing weights based on the network performance, allowing the student network to better benefit from both teachers. The experimental results on OpenSARShip and FUSAR-Ship demonstrate that DRRNet-SKD exhibits remarkable performance improvement on classical CNNs, outperforming state-of-the-art self-knowledge distillation methods.
Bo Xu 0002, Hao Zheng 0009, Zhigang Hu 0001, Liu Yang 0015, Meiguang Zheng, Xianting Feng
ICASSP1
2024 Spike-driven Transformer V2: Meta Spiking Neural Network Architecture Inspiring the Design of Next-generation Neuromorphic Chips
abstract
Neuromorphic computing, which exploits Spiking Neural Networks (SNNs) on neuromorphic chips, is a promising energy-efficient alternative to traditional AI. CNN-based SNNs are the current mainstream of neuromorphic computing. By contrast, no neuromorphic chips are designed especially for Transformer-based SNNs, which have just emerged, and their performance is only on par with CNN-based SNNs, offering no distinct advantage. In this work, we propose a general Transformer-based SNN architecture, termed as ``Meta-SpikeFormer", whose goals are: (1) *Lower-power*, supports the spike-driven paradigm that there is only sparse addition in the network; (2) *Versatility*, handles various vision tasks; (3) *High-performance*, shows overwhelming performance advantages over CNN-based SNNs; (4) *Meta-architecture*, provides inspiration for future next-generation Transformer-based neuromorphic chip designs. Specifically, we extend the Spike-driven Transformer in \citet{yao2023spike} into a meta architecture, and explore the impact of structure, spike-driven self-attention, and skip connection on its performance. On ImageNet-1K, Meta-SpikeFormer achieves 80.0\% top-1 accuracy (55M), surpassing the current state-of-the-art (SOTA) SNN baselines (66M) by 3.7\%. This is the first direct training SNN backbone that can simultaneously supports classification, detection, and segmentation, obtaining SOTA results in SNNs. Finally, we discuss the inspiration of the meta SNN architecture for neuromorphic chip design.
Man Yao, Tianxiang Hu, Zhaokun Zhou, Yonghong Tian 0001, Bo Xu 0002, Guoqi Li 0002
ICLR7
2024 High-Performance Temporal Reversible Spiking Neural Networks with O(L) Training Memory and O(1) Inference Cost
Man Yao, Xuerui Qiu, Yuhong Chou, Yonghong Tian 0001, Bo Xu 0002, Guoqi Li 0002
ICML8
2024 SA-MPF: A Status-Aware Mask Prediction Framework for Online Disease Diagnosis
abstract
An increasing number of individuals are turning to online self-diagnosis by matching their symptoms with potential medical conditions. This process involves two primary components: symptom inquiry and disease prediction. Existing works employ two separate modules to learn these tasks individually. Nevertheless, this intuitive approach encounters low data efficiency due to the separate learning of each module. In addition, previous research incorporates symptom statuses solely as part of the input without any additional modeling. However, this oversight neglects the importance of symptom status, which indicates whether the user has experienced the symptom. The status significantly influences both symptom inquiry strategies and disease prediction. To address these challenges, we propose a Status-Aware Mask Prediction Framework for online disease diagnosis, called SA-MPF. SA-MPF formalizes symptom inquiry and disease prediction as a single masked token prediction task, distinguishing them solely through the masked token type. Furthermore, we introduce a masked status prediction task, which unifies the prediction of symptom or disease statuses in a similar manner to masked token prediction, thereby enhancing the modeling of symptom and disease statuses. We evaluate SA-MPF on several datasets collected from various sources. The experimental results demonstrate substantial improvements achieved by SA-MPF. For example, on the GMD-12 dataset, SAMPF demonstrates a noteworthy 5% improvement in diagnostic accuracy, from 82% to 87%.1
Zefa Hu, Linghui Meng 0001, Bo Xu 0002
IJCNN6
2024 T-Agent: A Term-Aware Agent for Medical Dialogue Generation
abstract
Large language models (LLMs) excel at providing general and comprehensive health advice in single-turn dialogues. However, the limited information in single-turn conversations provided by users results in generated advice lacking personalization and specificity. In real-world medical consultations, doctors typically gain a comprehensive understanding of a patient’s condition through a series of iterative inquiries, enabling them to subsequently offer effective and personalized advice. To enhance capabilities similar to those of doctors, existing approaches often learn by increasing multi-turn medical dialogue corpora. In this study, we consider capturing the transitions of medical terms in each turn crucial, as they aid in understanding the flow of the conversation and enhance the accuracy of generating medical term information in the next turn. Therefore, we propose a Term-aware Agent (T-Agent) and develop a corresponding term extraction tool and term prediction model. T-Agent explicitly models the flow of term information in the dialogue by invoking the term extraction tool and the term prediction model. To better learn the term prediction task, we adopt a two-stage training approach. In the first stage, we conduct mixed training on a single large model, simultaneously learning term prediction and the ability of T-Agent to invoke term tools for dialogue. This mixed training in the first stage allows the large model to initially adapt to the term prediction task. In the second stage, we independently train the term prediction model and T-Agent on this basis, enhancing their expertise and performance in their respective tasks. We validated the effectiveness of the proposed method on two Chinese multi-turn medical dialogue datasets, demonstrating significant performance improvements, particularly in the accuracy of term information within dialogues.
Zefa Hu, Haozhi Zhao, Bo Xu 0002
IJCNN5
2024 Long Short-Term Reasoning Network with Theory of Mind for Efficient Multi-Agent Cooperation
abstract
Enhancing the theory of mind (ToM) ability of agents is becoming more and more critical in the research area of cooperative multi-agent reinforcement learning (MARL). ToM describes the ability of agents to understand their partners’ logic first, and then reason their intentions and behaviors accurately. In cognitive science, dual-reasoning pathway theory (DRPT) is a statistical ToM process, which claims that humans can achieve accurate and rapid reasoning by combining long-term and short-term reasoning (LTR and STR) pathways that cover different brain regions. However, most existing works focus on the quick decision-making ability of the STR, while overlooking the significance of the long-term ToM reasoning ability from the LTR. To emphasize such ability, we propose a long short-term reasoning (LSTR) algorithm which contains a large language model (LLM) for long-term reasoning and an additional augmenting module to decode the semantic space in LLM as the action space in MARL. Experimental results demonstrate that our LSTR algorithm has achieved significant improvement over competitive MARL methods (e.g., value-based QMIX and policy-based COMA) from the perspective of reward scores, convergence speed, and scalability.
Xiyun Li, Tielin Zhang, Linghui Meng 0001, Bo Xu 0002
IJCNN5
2024 TaCoD: Tasks-Commonality-Aware World in Meta Reinforcement Learning
abstract
Improving sample efficiency is a key objective in the field of reinforcement learning, and a learned dynamics model plays a crucial role in enhancing this efficiency. However, a significant challenge arises in aligning the objectives of model learning with policy performance. Existing solutions to this mismatch problem frequently yield models with diminished generalisation capabilities, thereby restricting their effectiveness to other tasks. We introduce the tasks-commonality-aware world model, which balances the specificity and generality of task-aware models. The concept of meta-value, developed through meta-learning, is proposed to capture task commonalities from a set of source tasks. This meta-value then serves as an intermediary to align the optimisation objectives of the model. In the context of meta reinforcement learning tasks, our approach exhibits generalisation capabilities that substantially surpass previous methods tackling the mismatch issue, as well as state-of-the-art model-based meta reinforcement learning methods.
Xuantang Xiong, Bo Xu 0002
IJCNN3
2024 Bridge the Query and Document: Contrastive Learning for Generative Document Retrieval
abstract
Generative retrieval has garnered significant attention for its end-to-end optimization and exceptional performance. Compared with the dense retrieval paradigm, the generative retrieval paradigm maps a query to a relevant document ID only relying on its model parameters, greatly simplifying the retrieval process. However, generative retrieval faces two challenges: it does not explicitly model the semantic relevance between query and document, and there exists a gap between the representation of query and document. To this end, we propose the Contrastive Search Index (ConSI), a simple but effective contrastive learning framework for generative document retrieval, to address the above challenges. Experiments show that the proposed ConSI consistently surpasses the previous generative retrieval baselines, NCI. Further analysis of different factors and indicators verifies the performance enhancement brought by our method. Besides, our ConSI also achieves excellent performance in the dense retrieval paradigm, demonstrating that the designed framework boosts representation learning ability and can be directly used as a dense retriever.
Ziyi Ni, Zefa Hu, Bo Xu 0002
IJCNN5
2024 CIEASR: Contextual Image-Enhanced Automatic Speech Recognition for Improved Homophone Discrimination
abstract
Automatic Speech Recognition (ASR) models pre-trained on large-scale speech datasets have achieved significant breakthroughs compared with traditional methods. However, mainstream pre-trained ASR models encounter challenges in distinguishing homophones, which have close or identical pronunciations. Previous studies have introduced visual auxiliary cues to address this challenge, yet the sophisticated use of lip movements falls short in correcting homophone errors. On the other hand, the fusion and utilization of scene images remain in an exploratory stage, with performance still inferior to the pre-trained speech model. In this paper, we introduce CIEASR (Contextual Image-Enhanced Automatic Speech Recognition), a novel multimodal speech recognition model that incorporates a new cue fusion method, using scene images as soft prompts to correct homophone errors. To mitigate data scarcity, we refine and expand the VSDial dataset for extensive experiments, illustrating that scene images contribute to the accurate recognition of entity nouns and personal pronouns. Our proposed CIEASR achieves state-of-the-art results on VSDial and Flickr8K, significantly reducing the Character Error Rate (CER) on VSDial from 3.61% to 0.92%.
Yiming Rong, Deyang Jiang, Bo Xu 0002
ACM Multimedia6
2024 RSC-SNN: Exploring the Trade-off Between Adversarial Robustness and Accuracy in Spiking Neural Networks via Randomized Smoothing Coding
Keming Wu, Man Yao, Yuhong Chou, Xuerui Qiu, Bo Xu 0002, Guoqi Li 0002
ACM Multimedia6
2024 Exploiting the Replay Memory Before Exploring the Environment: Enhancing Reinforcement Learning Through Empirical MDP Iteration
abstract
Reinforcement learning (RL) algorithms are typically based on optimizing a Markov Decision Process (MDP) using the optimal Bellman equation. Recent studies have revealed that focusing the optimization of Bellman equations solely on in-sample actions tends to result in more stable optimization, especially in the presence of function approximation. Upon on these findings, in this paper, we propose an Empirical MDP Iteration (EMIT) framework. EMIT constructs a sequence of empirical MDPs using data from the growing replay memory. For each of these empirical MDPs, it learns an estimated Q-function denoted as $\widehat{Q}$. The key strength is that by restricting the Bellman update to in-sample bootstrapping, each empirical MDP converges to a unique optimal $\widehat{Q}$ function. Furthermore, gradually expanding from the empirical MDPs to the original MDP induces a monotonic policy improvement. Instead of creating entirely new algorithms, we demonstrate that EMIT can be seamlessly integrated with existing online RL algorithms, effectively acting as a regularizer for contemporary Q-learning methods. We show this by implementing EMIT for two representative RL algorithms, DQN and TD3. Experimental results on Atari and MuJoCo benchmarks show that EMIT significantly reduces estimation errors and substantially improves the performance of both algorithms.
Hongming Zhang 0003, Chenjun Xiao, Chao Gao 0012, Bo Xu 0002, Martin Müller 0003
NeurIPS5
2024 MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map
abstract
Various linear complexity models, such as Linear Transformer (LinFormer), State Space Model (SSM), and Linear RNN (LinRNN), have been proposed to replace the conventional softmax attention in Transformer structures. However, the optimal design of these linear models is still an open question. In this work, we attempt to answer this question by finding the best linear approximation to softmax attention from a theoretical perspective. We start by unifying existing linear complexity models as the linear attention form and then identify three conditions for the optimal linear attention design: (1) Dynamic memory ability; (2) Static approximation ability; (3) Least parameter approximation. We find that none of the current linear models meet all three conditions, resulting in suboptimal performance. Instead, we propose Meta Linear Attention (MetaLA) as a solution that satisfies these conditions. Our experiments on Multi-Query Associative Recall (MQAR) task, language modeling, image classification, and Long-Range Arena (LRA) benchmark demonstrate that MetaLA is more effective than the existing linear models.
Yuhong Chou, Man Yao, Yuqi Pan, Rui-Jie Zhu 0003, Jibin Wu, Yiran Zhong, Bo Xu 0002, Guoqi Li 0002
NeurIPS9
2024 Multi-scale full spike pattern for semantic segmentation
Qiaoyi Su, Weihua He, Xiaobao Wei, Bo Xu 0002, Guoqi Li 0002
Neural Networks4
2024 SNN-BERT: Training-efficient Spiking Neural Networks for energy-efficient BERT
Qiaoyi Su, Shijie Mei 0001, Xingrun Xing, Man Yao, Bo Xu 0002, Guoqi Li 0002
Neural Networks6
2024 SSCFormer: Push the Limit of Chunk-Wise Conformer for Streaming ASR Using Sequentially Sampled Chunks and Chunked Causal Convolution
abstract
Currently, the chunk-wise schemes are often used to make Automatic Speech Recognition (ASR) models to support streaming deployment. However, existing approaches are unable to capture the global context, lack support for parallel training, or exhibit quadratic complexity for the computation of multi-head self-attention (MHSA). On the other side, the causal convolution, no future context used, has become thede factomodule in streaming Conformer. In this letter, we propose SSCFormer to push the limit of chunk-wise Conformer for streaming ASR using the following two techniques: 1) A novel cross-chunks context generation method, named Sequential Sampling Chunk (SSC) scheme, to re-partition chunks from regular partitioned chunks to facilitate efficient long-term contextual interaction within local chunks. 2)The Chunked Causal Convolution (C2Conv) is designed to concurrently capture the left context and chunk-wise future context. Evaluations on AISHELL-1 show that an End-to-End (E2E) CER 5.33% can achieve, which even outperforms a strong time-restricted baseline U2. Moreover, the chunk-wise MHSA computation in our model enables it to train with a large batch size and perform inference with linear complexity.
Fangyuan Wang 0003, Bo Xu 0011, Bo Xu 0002
IEEE Signal Process. Lett.3
2024 Multi-Cue Guided Semi-Supervised Learning Toward Target Speaker Separation in Real Environments
abstract
To solve the cocktail party problem in real multi-talker environments, this article proposed a multi-cue guided semi-supervised target speaker separation method (MuSS). Our MuSS integrates three target speaker-related cues, including spatial, visual, and voiceprint cues. Under the guidance of the cues, the target speaker is separated into a predefined output channel, and the interfering sources are separated into other output channels with the optimal permutation. Both synthetic mixtures and real mixtures are utilized for semi-supervised training. Specifically, for synthetic mixtures, the separated target source and other separated interfering sources are trained to reconstruct the ground-truth references, while for real mixtures, the mixture of two real mixtures is fed into our separation model, and the separated sources are remixed to reconstruct the two real mixtures. Besides, in order to facilitate finetuning and evaluating the estimated source on real mixtures, we introduce a real multi-modal speech separation dataset, RealMuSS, which is collected in real-world scenarios and is comprised of more than one hundred hours of multi-talker mixtures with high-quality pseudo references of the target speakers. Experimental results show that the pseudo references effectively improve the finetuning efficiency and enable the model to successfully learn and evaluate estimating speech on real mixtures, and various cue-driven separation models are greatly improved in signal-to-noise ratio and speech recognition accuracy under our semi-supervised learning framework.
Jiaming Xu 0001, Yunzhe Hao, Bo Xu 0002
IEEE ACM Trans. Audio Speech Lang. Process.4
2024 Self-Lateral Propagation Elevates Synaptic Modifications in Spiking Neural Networks for the Efficient Spatial and Temporal Classification
abstract
The brain’s mystery for efficient and intelligent computation hides in the neuronal encoding, functional circuits, and plasticity principles in natural neural networks. However, many plasticity principles have not been fully incorporated into artificial or spiking neural networks (SNNs). Here, we report that incorporating a novel feature of synaptic plasticity found in natural networks, whereby synaptic modifications self-propagate to nearby synapses, named self-lateral propagation (SLP), could further improve the accuracy of SNNs in three benchmark spatial and temporal classification tasks. The SLP contains lateral pre (${\rm SLP}_{\rm pre}$) and lateral post (${\rm SLP}_{\rm post}$) synaptic propagation, describing the spread of synaptic modifications among output synapses made by axon collaterals or among converging synapses on the postsynaptic neuron, respectively. The SLP is biologically plausible and can lead to a coordinated synaptic modification within layers that endow higher efficiency without losing much accuracy. Furthermore, the experimental results showed the impressive role of SLP in sharpening the normal distribution of synaptic weights and broadening the more uniform distribution of misclassified samples, which are both considered essential for understanding the learning convergence and network generalization of neural networks.
Tielin Zhang, Bo Xu 0002
IEEE Trans. Neural Networks Learn. Syst.3
2023 PiCor: Multi-Task Deep Reinforcement Learning with Policy Correction
abstract
Multi-task deep reinforcement learning (DRL) ambitiously aims to train a general agent that masters multiple tasks simultaneously. However, varying learning speeds of different tasks compounding with negative gradients interference makes policy learning inefficient. In this work, we propose PiCor, an efficient multi-task DRL framework that splits learning into policy optimization and policy correction phases. The policy optimization phase improves the policy by any DRL algothrim on the sampled single task without considering other tasks. The policy correction phase first constructs an adaptive adjusted performance constraint set. Then the intermediate policy learned by the first phase is constrained to the set, which controls the negative interference and balances the learning speeds across tasks. Empirically, we demonstrate that PiCor outperforms previous methods and significantly improves sample efficiency on simulated robotic manipulation and continuous control tasks. We additionally show that adaptive weight adjusting can further improve data efficiency and performance.
Fengshuo Bai, Hongming Zhang 0003, Tianyang Tao, Yanna Wang, Bo Xu 0002
AAAI6
2023 Complex Dynamic Neurons Improved Spiking Transformer Network for Efficient Automatic Speech Recognition
abstract
The spiking neural network (SNN) using leaky-integrated-and-fire (LIF) neurons has been commonly used in automatic speech recognition (ASR) tasks. However, the LIF neuron is still relatively simple compared to that in the biological brain. Further research on more types of neurons with different scales of neuronal dynamics is necessary. Here we introduce four types of neuronal dynamics to post-process the sequential patterns generated from the spiking transformer to get the complex dynamic neuron improved spiking transformer neural network (DyTr-SNN). We found that the DyTr-SNN could handle the non-toy automatic speech recognition task well, representing a lower phoneme error rate, lower computational cost, and higher robustness. These results indicate that the further cooperation of SNNs and neural dynamics at the neuron and network scales might have much in store for the future, especially on the ASR tasks.
Tielin Zhang, Minglun Han, Yi Wang 0077, Duzhen Zhang, Bo Xu 0002
AAAI6
2023 Cardsformer: Grounding Language to Learn a Generalizable Policy in Hearthstone
abstract
Hearthstone is a widely played collectible card game that challenges players to strategize using cards with various effects described in natural language. While human players can easily comprehend card descriptions and make informed decisions, artificial agents struggle to understand the game’s inherent rules and are unable to generalize their policies through natural language. To address this issue, we propose Cardsformer, a method capable of acquiring linguistic knowledge and learning a generalizable policy in Hearthstone. Cardsformer consists of a Prediction Model trained with offline trajectories to predict state transitions based on card descriptions and a Policy Model capable of generalizing its policy on unseen cards. To our knowledge, this is the first work to consider language knowledge in a card game. Experiments show that our approach significantly improves data efficiency and outperforms the state-of-the-art in Hearthstone even when there are untrained cards in the deck, inspiring a new perspective of tackling problems as such with knowledge representation from large language models. As the game constantly releases new cards along with new descriptions and new effects, the challenge in Hearthstone remains. To encourage further research, we make our code publicly available and publish PyStone, the code base of Hearthstone on which we conducted our experiments, as an open benchmark.
Wannian Xia, Jingqing Ruan, Dengpeng Xing, Bo Xu 0002
ECAI5
2023 Task-Prompt Generalised World Model in Multi-Environment Offline Reinforcement Learning
abstract
Offline reinforcement learning (RL) circumvents costly interactions with the environment by utilising historical trajectories. Incorporating a world model into this method could substantially enhance the transfer performance of various tasks without expensive calculations from scratch. However, due to the complexity arising from different types of generalisation, previous works have focused almost exclusively on single-environment tasks. In this study, we introduce a multi-environment offline RL setting to investigate whether a generalised world model can be learned from large, diverse datasets and serve as a good surrogate for policy learning in different tasks. Inspired by the success of multi-task prompt methods, we propose the Task-prompt Generalised World Model (TGW) framework, which demonstrates notable performance in this setting. TGW comprises three modules: a task-state prompter, a generalised dynamics module, and a reward module. We implement the generalised dynamics module as a transformer-based recurrent state-space model and employ prompts to provide task-specific instructions, enabling TGW to address the internal stochasticity of the generalised world model. On the MuJoCo control benchmarks, TGW significantly outperforms previous offline RL algorithms in multi-environment setting.
Xuantang Xiong, Linghui Meng 0001, Jingqing Ruan, Qingyang Zhang 0004, Guoqi Li 0002, Dengpeng Xing, Bo Xu 0002
ECAI7
2023 Matching-Based Term Semantics Pre-Training for Spoken Patient Query Understanding
abstract
Medical Slot Filling (MSF) task aims to convert medical queries into structured information, playing an essential role in diagnosis dialogue systems. However, the lack of sufficient term semantics learning makes existing approaches hard to capture semantically identical but colloquial expressions of terms in medical conversations. In this work, we formalize MSF into a matching problem and propose a Term Semantics Pre-trained Matching Network (TSPMN) that takes both terms and queries as input to model their semantic inter-action. To learn term semantics better, we further design two self-supervised objectives, including Contrastive Term Discrimination (CTD) and Matching-based Mask Term Modeling (MMTM). CTD determines whether it is the masked term in the dialogue for each given term, while MMTM directly predicts the masked ones. Experimental results on two Chinese benchmarks show that TSPMN outperforms strong baselines, especially in few-shot settings1.
Zefa Hu, Xiuyi Chen, Minglun Han, Ziyi Ni, Jing Shi 0003, Bo Xu 0002
ICASSP8
2023 Inherent Redundancy in Spiking Neural Networks
abstract
Spiking Neural Networks (SNNs) are well known as a promising energy-efficient alternative to conventional artificial neural networks. Subject to the preconceived impression that SNNs are sparse firing, the analysis and optimization of inherent redundancy in SNNs have been largely overlooked, thus the potential advantages of spike-based neuromorphic computing in accuracy and energy efficiency are interfered. In this work, we pose and focus on three key questions regarding the inherent redundancy in SNNs. We argue that the redundancy is induced by the spatio-temporal invariance of SNNs, which enhances the efficiency of parameter utilization but also invites lots of noise spikes. Further, we analyze the effect of spatio-temporal invariance on the spatio-temporal dynamics and spike firing of SNNs. Then, motivated by these analyses, we propose an Advance Spatial Attention (ASA) module to harness SNNs’ redundancy, which can adaptively optimize their membrane potential distribution by a pair of individual spatial attention sub-modules. In this way, noise spike features are accurately regulated. Experimental results demonstrate that the proposed method can significantly drop the spike firing with better performance than state-of-the-art SNN baselines. Our code is available in https://github.com/BICLab/ASA-SNN.
Man Yao, Guang-She Zhao, Yaoyuan Wang, Bo Xu 0002, Guoqi Li 0002
ICCV6
2023 Replay Memory as An Empirical MDP: Combining Conservative Estimation with Experience Replay
Hongming Zhang 0003, Chenjun Xiao, Jun Jin 0001, Bo Xu 0002, Martin Müller 0003
ICLR5
2023 Make Spoken Document Readable: Leveraging Graph Attention Networks for Chinese Document-Level Spoken-to-Written Simplification
Bo Xu 0002
ICONIP (12)4
2023 Balancing Exploration and Exploitation in Hierarchical Reinforcement Learning via Latent Landmark Graphs
abstract
Goal-Conditioned Hierarchical Reinforcement Learning (GCHRL) is a promising paradigm to address the exploration-exploitation dilemma in reinforcement learning. It decomposes the source task into sub goal conditional subtasks and conducts exploration and exploitation in the subgoal space. The effectiveness of GCHRL heavily relies on sub goal representation functions and sub goal selection strategy. However, existing works often overlook the temporal coherence in GCHRL when learning latent sub goal representations and lack an efficient sub goal selection strategy that balances exploration and exploitation. This paper proposes HIerarchical reinforcement learning via dynamically building Latent Landmark graphs (HILL) to overcome these limitations. HILL learns latent subgoal representations that satisfy temporal coherence using a contrastive representation learning objective. Based on these representations, HILL dynamically builds latent landmark graphs and employs a novelty measure on nodes and a utility measure on edges. Finally, HILL develops a subgoal selection strategy that balances exploration and exploitation by jointly considering both measures. Experimental results demonstrate that HILL outperforms state-of-the-art baselines on continuous control tasks with sparse rewards in sample efficiency and asymptotic performance. Our code is available at https://github.com/papercode2022/HILL.
Qingyang Zhang 0004, Jingqing Ruan, Xuantang Xiong, Dengpeng Xing, Bo Xu 0002
IJCNN6
2023 Enhancing Visual Question Answering via Deconstructing Questions and Explicating Answers
abstract
A compositional question refers to a question that involves multiple visual objects, as well as their attributes and relationships, which requires compositional reasoning to answer.Existing VQA models can well answer a compositional question, but few works can give the reasoning process and explain why this answer is given.In this paper, we propose a novel model (DEEX) to enhance visual question answering via DEconstructing questions and EXplicating answers when answering compositional questions.Specifically, DEEX aims to accomplish three sub-tasks: (1) Compositional Question Answering (CQA), (2) Question Deconstructing (QD), and (3) Answer Explicating (AE).We utilize prompt-based multi-task learning to train the proposed DEEX to be able to answer questions and give explanations simultaneously.Experimental results on the GQA dataset demonstrate our method's effectiveness, which can enhance visual question answering by giving corresponding reasoning processes and explanations.
Minglun Han, Jing Shi 0003, Bo Xu 0002
INTERSPEECH5
2023 Knowledge Transfer from Pre-trained Language Models to Cif-based Speech Recognizers via Hierarchical Distillation
Minglun Han, Jing Shi 0003, Bo Xu 0002
INTERSPEECH5
2023 Generalized Robot Dynamics Learning and Gen2Real Transfer
abstract
Acquiring dynamics is critical for robot learning and is fundamental to planning and control. This paper concerns two fundamental questions: How can we learn a model that covers massive, diverse robot dynamics? Can we construct a model that lifts the data-collection pain and domain expertise required for building specific robot models? We learn the dynamics involved in a dataset containing a large number of serial articulated robots and propose a new concept, “Gen2Real”, to transfer simulated, generalized models to physical, specific robots. We generate a large-scale dataset by randomizing dynamics parameters, topology configurations, and model dimensions, which, in sequence, correspond to different properties, connections, and numbers of robot links. A structure modified from the generative pre-trained transformer is applied to approximate the dynamics of massive heterogeneous robots. In Gen2Real, we transfer the pre-trained model to a target robot using distillation, for the sake of real-time computation. The results demonstrate the superiority of the proposed method in terms of its accuracy in learning a tremendous amount of robot dynamics and its generality to transfer to different robots.
Dengpeng Xing, Zechang Wang, Bo Xu 0002
IROS5
2023 Spike-driven Transformer
abstract
Spiking Neural Networks (SNNs) provide an energy-efficient deep learning option due to their unique spike-based event-driven (i.e., spike-driven) paradigm. In this paper, we incorporate the spike-driven paradigm into Transformer by the proposed Spike-driven Transformer with four unique properties: (1) Event-driven, no calculation is triggered when the input of Transformer is zero; (2) Binary spike communication, all matrix multiplications associated with the spike matrix can be transformed into sparse additions; (3) Self-attention with linear complexity at both token and channel dimensions; (4) The operations between spike-form Query, Key, and Value are mask and addition. Together, there are only sparse addition operations in the Spike-driven Transformer. To this end, we design a novel Spike-Driven Self-Attention (SDSA), which exploits only mask and addition operations without any multiplication, and thus having up to $87.2\times$ lower computation energy than vanilla self-attention. Especially in SDSA, the matrix multiplication between Query, Key, and Value is designed as the mask operation. In addition, we rearrange all residual connections in the vanilla Transformer before the activation functions to ensure that all neurons transmit binary spike signals. It is shown that the Spike-driven Transformer can achieve 77.1\% top-1 accuracy on ImageNet-1K, which is the state-of-the-art result in the SNN field.
Man Yao, Zhaokun Zhou, Li Yuan 0007, Yonghong Tian 0001, Bo Xu 0002, Guoqi Li 0002
NeurIPS6
2023 ODE-based Recurrent Model-free Reinforcement Learning for POMDPs
abstract
Neural ordinary differential equations (ODEs) are widely recognized as the standard for modeling physical mechanisms, which help to perform approximate inference in unknown physical or biological environments. In partially observable (PO) environments, how to infer unseen information from raw observations puzzled the agents. By using a recurrent policy with a compact context, context-based reinforcement learning provides a flexible way to extract unobservable information from historical transitions. To help the agent extract more dynamics-related information, we present a novel ODE-based recurrent model combines with model-free reinforcement learning (RL) framework to solve partially observable Markov decision processes (POMDPs). We experimentally demonstrate the efficacy of our methods across various PO continuous control and meta-RL tasks. Furthermore, our experiments illustrate that our method is robust against irregular observations, owing to the ability of ODEs to model irregularly-sampled time series.
Xuanle Zhao, Duzhen Zhang, Liyuan Han, Tielin Zhang, Bo Xu 0002
NeurIPS5
2023 Filtered Observations for Model-Based Multi-agent Reinforcement Learning
Linghui Meng 0001, Xuantang Xiong, Yifan Zang 0001, Guoqi Li 0002, Dengpeng Xing, Bo Xu 0002
ECML/PKDD (4)7
2023 Meta neurons improve spiking neural networks for efficient spatio-temporal learning
Tielin Zhang, Shuncheng Jia, Bo Xu 0002
Neurocomputing4
2023 Origin of the efficiency of spike timing-based neural computation for processing temporal information
Jiaming Xu 0001, Tielin Zhang, Mu-Ming Poo, Bo Xu 0002
Neural Networks5
2023 Attention Spiking Neural Networks
abstract
Brain-inspired spiking neural networks (SNNs) are becoming a promising energy-efficient alternative to traditional artificial neural networks (ANNs). However, the performance gap between SNNs and ANNs has been a significant hindrance to deploying SNNs ubiquitously. To leverage the full potential of SNNs, in this paper we study the attention mechanisms, which can help human focus on important information. We present our idea of attention in SNNs with a multi-dimensional attention module, which infers attention weights along the temporal, channel, as well as spatial dimension separately or simultaneously. Based on the existing neuroscience theories, we exploit the attention weights to optimize membrane potentials, which in turn regulate the spiking response. Extensive experimental results on event-based action recognition and image classification datasets demonstrate that attention facilitates vanilla SNNs to achieve sparser spiking firing, better performance, and energy efficiency concurrently. In particular, we achieve top-1 accuracy of 75.92% and 77.08% on ImageNet-1 K with single/4-step Res-SNN-104, which are state-of-the-art results in SNNs. Compared with counterpart Res-ANN-104, the performance gap becomes -0.95/+0.21 percent and the energy efficiency is 31.8×/7.4×. To analyze the effectiveness of attention SNNs, we theoretically prove that the spiking degradation or the gradient vanishing, which usually holds in general SNNs, can be resolved by introducing the block dynamical isometry theory. We also analyze the efficiency of attention SNNs based on our proposed spiking response visualization method. Our work lights up SNN's potential as a general backbone to support various applications in the field of SNN research, with a great balance between effectiveness and energy efficiency.
Man Yao, Guang-She Zhao, Hengyu Zhang 0001, Yifan Hu 0013, Lei Deng 0003, Yonghong Tian 0001, Bo Xu 0002, Guoqi Li 0002
IEEE Trans. Pattern Anal. Mach. Intell.7
2023 A Brain-Inspired Approach for Probabilistic Estimation and Efficient Planning in Precision Physical Interaction
abstract
This article presents a novel structure of spiking neural networks (SNNs) to simulate the joint function of multiple brain regions in handling precision physical interactions. This task desires efficient movement planning while considering contact prediction and fast radial compensation. Contact prediction demands the cognitive memory of the interaction model, and we novelly propose a double recurrent network to imitate the hippocampus, addressing the spatiotemporal property of the distribution. Radial contact response needs rich spatial information, and we use a cerebellum-inspired module to achieve temporally dynamic prediction. We also use a block-based feedforward network to plan movements, behaving like the prefrontal cortex. These modules are integrated to realize the joint cognitive function of multiple brain regions in prediction, controlling, and planning. We present an appropriate controller and planner to generate teaching signals and provide a feasible network initialization for reinforcement learning, which modifies synapses in accordance with reality. The experimental results demonstrate the validity of the proposed method.
Dengpeng Xing, Tielin Zhang, Bo Xu 0002
IEEE Trans. Cybern.4
2022 Multi-Sacle Dynamic Coding Improved Spiking Actor Network for Reinforcement Learning
abstract
With the help of deep neural networks (DNNs), deep reinforcement learning (DRL) has achieved great success on many complex tasks, from games to robotic control. Compared to DNNs with partial brain-inspired structures and functions, spiking neural networks (SNNs) consider more biological features, including spiking neurons with complex dynamics and learning paradigms with biologically plausible plasticity principles. Inspired by the efficient computation of cell assembly in the biological brain, whereby memory-based coding is much more complex than readout, we propose a multiscale dynamic coding improved spiking actor network (MDC-SAN) for reinforcement learning to achieve effective decision-making. The population coding at the network scale is integrated with the dynamic neurons coding (containing 2nd-order neuronal dynamics) at the neuron scale towards a powerful spatial-temporal state representation. Extensive experimental results show that our MDC-SAN performs better than its counterpart deep actor network (based on DNNs) on four continuous control tasks from OpenAI gym. We think this is a significant attempt to improve SNNs from the perspective of efficient coding towards effective decision-making, just like that in biological networks.
Duzhen Zhang, Tielin Zhang, Shuncheng Jia, Bo Xu 0002
AAAI4
2022 A Multi Domain Knowledge Enhanced Matching Network for Response Selection in Retrieval-Based Dialogue Systems
abstract
Building a human-machine conversational agent is a core problem in Artificial Intelligence, where knowledge has to be integrated into the model effectively. In this paper, we propose a Multi Domain Knowledge Enhanced Matching Network (MDKEMN) to build retrievalbased dialogue systems that could leverage both explicit knowledge graph and implicit domain knowledge for response selection. Specifically, our MDKEMN leverages the self-attention mechanism of a single-stream Transformer to make deep interactions among the dialogue context, response candidate and external knowledge graph, and finally returns the matching degree of each context-response pair under the external knowledge. Furthermore, to leverage the implicit domain knowledge from all domains to improve the performance of each domain, we combine the multi-domain datasets for training and then finetune the pretrained model on each domain. Experimental results show (1) the effectiveness of both explicit and implicit knowledge incorporating and (2) the superiority of our approach over previous baselines on a Chinese multi-domain knowledge-driven dialogue dataset.
Xiuyi Chen, Bo Xu 0002
ICASSP4
2022 Improving Cross-Modal Understanding in Visual Dialog Via Contrastive Learning
abstract
Visual Dialog is a challenging vision-language task since the visual dialog agent needs to answer a series of questions after reasoning over both the image content and dialog history. Though existing methods try to deal with the cross-modal understanding in visual dialog, they are still not enough in ranking candidate answers based on their understanding of visual and textual contexts. In this paper, we analyze the cross-modal understanding in visual dialog based on the vision-language pre-training model VD-BERT and propose a novel approach to improve the cross-modal understanding for visual dialog, named ICMU. ICMU enhances cross-modal understanding by distinguishing different pulled inputs (i.e. pulled images, questions or answers) based on four-way contrastive learning. In addition, ICMU exploits the single-turn visual question answering to enhance the visual dialog model’s cross-modal understanding to handle a multi-turn visually-grounded conversation. Experiments show that the proposed approach improves the visual dialog model’s cross-modal understanding and brings satisfactory gain to the Vis-Dial dataset.
Xiuyi Chen, Bo Xu 0002
ICASSP4
2022 Improving End-to-End Contextual Speech Recognition with Fine-Grained Contextual Knowledge Selection
abstract
Nowadays, most methods for end-to-end contextual speech recognition bias the recognition process towards contextual knowledge. Since all-neural contextual biasing methods rely on phrase-level contextual modeling and attention-based relevance modeling, they may suffer from the confusion between similar context-specific phrases, which hurts predictions at the token level. In this work, we focus on mitigating confusion problems with fine-grained contextual knowledge selection (FineCoS). In FineCoS, we introduce fine-grained knowledge to reduce the uncertainty of token predictions. Specifically, we first apply phrase selection to narrow the range of phrase candidates, and then conduct token attention on the tokens in the selected phrase candidates. Moreover, we re-normalize the attention weights of most relevant phrases in inference to obtain more focused phrase-level contextual representations, and inject position information to help model better discriminate phrases or tokens. On LibriSpeech and an in-house 160,000-hour dataset, we explore the proposed methods based on an all-neural biasing method, collaborative decoding (ColDec). The proposed methods further bring at most 6.1% relative word error rate reduction on LibriSpeech and 16.4% relative character error rate reduction on the in-house dataset.
Minglun Han, Linhao Dong, Zhenlin Liang, Zejun Ma 0001, Bo Xu 0002
ICASSP7
2022 Motif-Topology and Reward-Learning Improved Spiking Neural Network for Efficient Multi-Sensory Integration
abstract
Network architectures and learning principles are key in forming complex functions in artificial neural networks (ANNs) and spiking neural networks (SNNs). SNNs are considered the new-generation artificial networks by incorporating more biological features than ANNs, including dynamic spiking neurons, functionally specified architectures, and efficient learning paradigms. In this paper, we propose a Motiftopology and Reward-learning improved SNN (MR-SNN) for efficient multi-sensory integration. MR-SNN contains 13 types of 3-node Motif topologies which are first extracted from independent single-sensory learning paradigms and then integrated for multi-sensory classification. The experimental results showed higher accuracy and stronger robustness of the proposed MR-SNN than other conventional SNNs without using Motifs. Furthermore, the proposed reward learning paradigm was biologically plausible and can better explain the cognitive McGurk effect caused by incongruent visual and auditory sensory signals.
Shuncheng Jia, Ruichen Zuo, Tielin Zhang, Bo Xu 0002
ICASSP5
2022 Kinematics Learning of Massive Heterogeneous Serial Robots
abstract
Kinematics and instantaneous kinematics are fundamental in many robotic tasks, such as positioning and collision avoidance. Existing learning methods mainly concern a single robot, and small-scale networks are sufficient for considerable approximation accuracy. A question is: Can we learn a kinematics model that can generalize to various robots rather than a single robot? This paper studies the kinematics learning of massive heterogeneous serial robots and the transfer of these general models to reality. We generate a dataset by randomizing dimensions, configurations, and link lengths and employ a network based on the generative pre-trained transformer to learn general kinematics mappings. We directly transfer our models for accuracy and use distillation-based transfer for computational efficiency. The results validate that our method can accurately approximate the kinematics of thousands of robot models and demonstrates generality in transfer.
Dengpeng Xing, Wannian Xia, Bo Xu 0002
ICRA3
2022 Recent Advances and New Frontiers in Spiking Neural Networks
abstract
In recent years, spiking neural networks (SNNs) have received extensive attention in brain-inspired intelligence due to their rich spatially-temporal dynamics, various encoding methods, and event-driven characteristics that naturally fit the neuromorphic hardware. With the development of SNNs, brain-inspired intelligence, an emerging research field inspired by brain science achievements and aiming at artificial general intelligence, is becoming hot. This paper reviews recent advances and discusses new frontiers in SNNs from five major research topics, including essential elements (i.e., spiking neuron models, encoding methods, and topology structures), neuromorphic datasets, optimization algorithms, software, and hardware frameworks. We hope our survey can help researchers understand SNNs better and inspire new works to advance this field.
Duzhen Zhang, Tielin Zhang, Shuncheng Jia, Bo Xu 0002
IJCAI5
2022 Learning in Bi-level Markov Games
abstract
Although multi-agent reinforcement learning (MARL) has demonstrated remarkable progress in tackling sophisticated cooperative tasks, the assumption that agents take simultaneous actions still limits the applicability of MARL for many real-world problems. In this work, we relax the assumption by proposing the framework of the bi-level Markov game (BMG). BMG breaks the simultaneity by assigning two players with a leader-follower relationship in which the leader considers the policy of the follower who is taking the best response based on the leader's actions. We propose two provably convergent algorithms to solve BMG: BMG-1 and BMG-2. The former uses the standard Q-learning, while the latter relieves solving the local Stackelberg equilibrium in BMG-1 with the further two-step transition to estimate the state value. For both methods, we consider temporal difference learning techniques with both tabular and neural network representations. To verify the effectiveness of our BMG framework, we test on a series of games, including Seeker, Cooperative Navigation, and Football, that are challenging to existing MARL solvers find challenging to solve: Seeker, Cooperative Navigation, and Football. Experimental results show that our BMG methods achieve competitive advantages in terms of better performance and lower variance.
Linghui Meng 0001, Jingqing Ruan, Dengpeng Xing, Bo Xu 0002
IJCNN4
2022 Joint Modeling of Document and Label with Clause Interaction Hypergraph for ICD Medical Code Assignment
abstract
Automatic medical code assignment for clinical records is the fundamental problem of medical statistical research and informatization. Due to the high dimension and sparse distribution of label space, it is necessary to make full use of the description information of the labels. However, most of the current work is based on similarity matching at the level of token or n-gram, and ignores information fusion with richer semantic structure and representation. In this paper, we propose a Clause Interaction HyperGraph (CIHG) to jointly model documents and label descriptions, which construct a richer semantic interaction at the level of clause. The CIHG models the high-order co-occurrence relationship between document and labels based on hypergraph, and uses the semantic structure of the document to constrain the encoding of labels. Experiments on widely used medical code assignment datasets show that our method successfully constrains the embedding of labels and significantly improves predictive precision11The code is available at https://github.com/CKRE/CIHG.
Linghui Meng 0001, Bo Xu 0002
IJCNN4
2022 Token-level Speaker Change Detection Using Speaker Difference and Speech Content via Continuous Integrate-and-fire
Zhiyun Fan, Zhenlin Liang, Linhao Dong, Jun Zhang 0066, Zejun Ma 0001, Bo Xu 0002
INTERSPEECH9
2022 Unsupervised and Pseudo-Supervised Vision-Language Alignment in Visual Dialog
abstract
Visual dialog requires models to give reasonable answers according to a series of coherent questions and related visual concepts in images. However, most current work either focuses on attention-based fusion or pre-training on large-scale image-text pairs, ignoring the critical role of explicit vision-language alignment in visual dialog. To remedy this defect, we propose a novel unsupervised and pseudo-supervised vision-language alignment approach for visual dialog (AlignVD). Firstly, AlginVD utilizes the visual and dialog encoder to represent images and dialogs. Then, it explicitly aligns visual concepts with textual semantics via unsupervised and pseudo-supervised vision-language alignment (UVLA and PVLA). Specifically, UVLA utilizes a graph autoencoder, while PVLA uses dialog-guided visual grounding to conduct alignment. Finally, based on the aligned visual and textual representations, AlignVD gives a reasonable answer to the question via the cross-modal decoder. Extensive experiments on two large-scale visual dialog datasets have demonstrated the effectiveness of vision-language alignment, and our proposed AlignVD achieves new state-of-the-art results. In addition, our single model has won first place on the visual dialog challenge leaderboard with a NDCG metric of 78.70, surpassing the previous best ensemble model by about 1 point.
Duzhen Zhang, Xiuyi Chen, Jing Shi 0003, Bo Xu 0002
ACM Multimedia6
2022 LiMuSE: Lightweight Multi-Modal Speaker Extraction
abstract
Multi-modal cues, including spatial information, facial expression and voiceprint, are introduced to the speech separation and speaker extraction tasks to serve as complementary information to achieve better performance. However, the introduction of these cues brings about an increasing number of parameters and model complexity, which makes it harder to deploy these models on resource-constrained devices. In this paper, we alleviate the aforementioned problem by proposing a Lightweight Multi-modal framework for Speaker Extraction (LiMuSE). We propose to use GC-equipped TCN, which incorporates Group Communication (GC) and Temporal Convolutional Network (TCN) in the Context Codec module, the audio block and the fusion block. The experiments on the MC_GRID dataset demonstrate that LiMuSE achieves on par or better performance with a much smaller number of parameters and less model complexity. We further investigate the impacts of the quantization of LiMuSE. Our code and dataset are provided11https://github.com/aispeech-lab/LiMuSE.
Yunzhe Hao, Jiaming Xu 0001, Bo Xu 0002
SLT5
2022 MCascade R-CNN: A Modified Cascade R-CNN for Detection of Calcified on Coronary Artery Angiography Images
abstract
Among cardiovascular diseases, coronary artery calcification (CAC) is a high-risk factor for worsening protopathy and increased mortality. However, the coronary artery an-giogram, which is the main approach for CAC diagnosis, suffers from plenty of photographing noise. This brings difficulties to detect calcification from the background. In this paper, a modified Cascade R-CNN (MCascade R-CNN) network is proposed to deal with the problem of calcium detection in angiograms. In the proposed network, we propose an innovative balanced aggregation pyramid structure, integrating multi-level features of every depth in the feature map, based on enhanced propagation of strong semantic features. In addition, a new convolutional attention mechanism is designed to improve the performance of the detector. Experiments show that the proposed method enjoys better performance in detecting and marking CAC in angiograms,
Wei Wang 0335, Honggang Zhang 0002, Lihua Xie 0003, Bo Xu 0002
VCIP6
2022 Train from scratch: Single-stage joint training of speech separation and recognition
Jing Shi 0003, Xuankai Chang, Shinji Watanabe 0001, Bo Xu 0002
Comput. Speech Lang.4
2022 Compressing speaker extraction model with ultra-low precision quantization and knowledge distillation
Yunzhe Hao, Jiaming Xu 0001, Bo Xu 0002
Neural Networks4
2022 Sequence-Level Speaker Change Detection With Difference-Based Continuous Integrate-and-Fire
abstract
Speaker change detection is an important task in multi-party interactions such as meetings and conversations. In this paper, we address the speaker change detection task from the perspective of sequence transduction. Specifically, we propose a novel encoder-decoder framework that directly converts the input feature sequence to the speaker identity sequence. The difference-based continuous integrate-and-fire mechanism is designed to support this framework. It detects speaker changes by integrating the speaker difference between the encoder outputs frame-by-frame and transfers encoder outputs to segment-level speaker embeddings according to the detected speaker changes. The whole framework is supervised by the speaker identity sequence, a weaker label than the precise speaker change points. The experiments on the AMI and DIHARD-I corpora show that our sequence-level method consistently outperforms a strong frame-level baseline that uses the precise speaker change labels.
Zhiyun Fan, Linhao Dong, Zejun Ma 0001, Bo Xu 0002
IEEE Signal Process. Lett.5
2022 A Brain-Inspired Approach for Collision-Free Movement Planning in the Small Operational Space
abstract
In a small operational space, e.g., mesoscale or microscale, we need to control movements carefully because of fragile objects. This article proposes a novel structure based on spiking neural networks to imitate the joint function of multiple brain regions in visual guiding in the small operational space and offers two channels to achieve collision-free movements. For the state sensation, we simulate the primary visual cortex to directly extract features from multiple input images and the high-level visual cortex to obtain the object distance, which is indirectly measurable, in the Cartesian coordinates. Our approach emulates the prefrontal cortex from two aspects: multiple liquid state machines to predict distances of the next several steps based on the preceding trajectory and a block-based excitation-inhibition feedforward network to plan movements considering the target and prediction. Responding to "too close" states needs rich temporal information, and we leverage a cerebellar network for the subconscious reaction. From the viewpoint of the inner pathway, they also form two channels. One channel starts from state extraction to attraction movement planning, both in the camera coordinates, behaving visual-servo control. The other is the collision-avoidance channel, which calculates distances, predicts trajectories, and reacts to the repulsion, all in the Cartesian coordinates. We provide appropriate supervised signals for coarse training and apply reinforcement learning to modify synapses in accordance with reality. Simulation and experiment results validate the proposed method.
Dengpeng Xing, Tielin Zhang, Bo Xu 0002
IEEE Trans. Neural Networks Learn. Syst.4
2022 Tuning Convolutional Spiking Neural Network With Biologically Plausible Reward Propagation
abstract
Spiking neural networks (SNNs) contain more biologically realistic structures and biologically inspired learning principles than those in standard artificial neural networks (ANNs). SNNs are considered the third generation of ANNs, powerful on the robust computation with a low computational cost. The neurons in SNNs are nondifferential, containing decayed historical states and generating event-based spikes after their states reaching the firing threshold. These dynamic characteristics of SNNs make it difficult to be directly trained with the standard backpropagation (BP), which is also considered not biologically plausible. In this article, a biologically plausible reward propagation (BRP) algorithm is proposed and applied to the SNN architecture with both spiking-convolution (with both 1-D and 2-D convolutional kernels) and full-connection layers. Unlike the standard BP that propagates error signals from postsynaptic to presynaptic neurons layer by layer, the BRP propagates target labels instead of errors directly from the output layer to all prehidden layers. This effort is more consistent with the top-down reward-guiding learning in cortical columns of the neocortex. Synaptic modifications with only local gradient differences are induced with pseudo-BP that might also be replaced with the spike-timing-dependent plasticity (STDP). The performance of the proposed BRP-SNN is further verified on the spatial (including MNIST and Cifar-10) and temporal (including TIDigits and DvsGesture) tasks, where the SNN using BRP has reached a similar accuracy compared to other state-of-the-art (SOTA) BP-based SNNs and saved 50% more computational cost than ANNs. We think that the introduction of biologically plausible learning rules to the training procedure of biologically realistic SNNs will give us more hints and inspiration toward a better understanding of the biological system's intelligent nature.
Tielin Zhang, Shuncheng Jia, Bo Xu 0002
IEEE Trans. Neural Networks Learn. Syst.4
2021 Consecutive Decoding for Speech-to-text Translation
abstract
Speech-to-text translation (ST), which directly translates the source language speech to the target language text, has attracted intensive attention recently. However, the combination of speech recognition and machine translation in a single model poses a heavy burden on the direct cross-modal cross-lingual mapping. To reduce the learning difficulty, we propose COnSecutive Transcription and Translation (COSTT), an integral approach for speech-to-text translation. The key idea is to generate source transcript and target translation text with a single decoder. It benefits the model training so that additional large parallel text corpus can be fully exploited to enhance the speech translation training. Our method is verified on three mainstream datasets, including Augmented LibriSpeech English-French dataset, TED English-German dataset, and TED English-Chinese dataset. Experiments show that our proposed COSTT outperforms the previous state-of-the-art methods. The code is available at https://github.com/dqqcasia/st.
Qianqian Dong, Mingxuan Wang, Hao Zhou 0012, Bo Xu 0002, Lei Li 0005
AAAI5
2021 Listen, Understand and Translate: Triple Supervision Decouples End-to-end Speech-to-text Translation
abstract
An end-to-end speech-to-text translation (ST) takes audio in a source language and outputs the text in a target language. Existing methods are limited by the amount of parallel corpus. Can we build a system to fully utilize signals in a parallel ST corpus? We are inspired by human understanding system which is composed of auditory perception and cognitive processing. In this paper, we propose Listen-Understand-Translate, (LUT), a unified framework with triple supervision signals to decouple the end-to-end speech-to-text translation task. LUT is able to guide the acoustic encoder to extract as much information from the auditory input. In addition, LUT utilizes a pre-trained BERT model to enforce the upper encoder to produce as much semantic information as possible, without extra data. We perform experiments on a diverse set of speech translation benchmarks, including Librispeech English-French, IWSLT English-German and TED English-Chinese. Our results demonstrate LUT achieves the state-of-the-art performance, outperforming previous methods. The code is available at https://github.com/dqqcasia/st.
Qianqian Dong, Rong Ye, Mingxuan Wang, Hao Zhou 0012, Bo Xu 0002, Lei Li 0005
AAAI6
2021 MACCIF-TDNN: Multi Aspect Aggregation of Channel and Context Interdependence Features in TDNN-Based Speaker Verification
abstract
Most of the recent state-of-the-art results for speaker verification are achieved by X-vector and its subsequent variants. In this paper, we propose a new network architecture which aggregates the channel and context interdependence features from multi aspects based on Time Delay Neural Network (TDNN). Firstly, we use the SE-Res2Blocks as in ECAPA-TDNN to explicitly model the channel interdependence to realize adaptive calibration of channel features, and process local context features in a multi-scale way at a more granular level compared with conventional TDNN-based methods. Secondly, we explore to use the encoder structure of Transformer to model the global context interdependence features at an utterance level which can capture better long term temporal characteristics. Before the pooling layer, we aggregate the outputs of SE-Res2Blocks and Transformer Encoders to leverage the complementary channel and context interdependence features learned by themself respectively. Finally, instead of performing a single attentive statistics pooling, we also find it beneficial to extend the pooling method in a multi-head way which can discriminate features from multiple aspects. The proposed MACCIF-TDNN architecture can outperform most of the state-of-the-art TDNN based systems on VoxCeleb1 test sets.
Fangyuan Wang 0003, Zhigang Song, Hongchen Jiang, Bo Xu 0002
ASRU4
2021 Cif-Based Collaborative Decoding for End-to-End Contextual Speech Recognition
abstract
End-to-end (E2E) models have achieved promising results on multiple speech recognition benchmarks, and shown the potential to become the mainstream. However, the unified structure and the E2E training hamper injecting context information into them for contextual biasing. Though contextual LAS (CLAS) gives an excellent all-neural solution, the degree of biasing to given contextual information is not explicitly controllable. In this paper, we focus on incorporating contextual information into the continuous integrate-and-fire (CIF) based model that supports contextual biasing in a more controllable fashion. Specifically, an extra context processing network is introduced to extract contextual embeddings, integrate acoustically relevant contextual information and decode the contextual output distribution, thus forming a collaborative decoding with the decoder of the CIF-based model. Evaluated on the named entity rich evaluation sets of HKUST/AISHELL-2, our method brings relative character error rate (CER) reduction of 8.83%/21.13% and relative named entity character error rate (NE-CER) reduction of 40.14%/51.50% when compared with a strong baseline. Besides, it keeps the performance on original evaluation set without degradation.
Minglun Han, Linhao Dong, Bo Xu 0002
ICASSP4
2021 Wase: Learning When to Attend for Speaker Extraction in Cocktail Party Environments
abstract
In the speaker extraction problem, it is found that additional information from the target speaker contributes to the tracking and extraction of the target speaker, which includes voiceprint, lip movement, facial expression, and spatial information. However, no one cares for the cue of sound onset, which has been emphasized in the auditory scene analysis and psychology. Inspired by it, we explicitly modeled the onset cue and verified the effectiveness in the speaker extraction task. We further extended to the onset/offset cues and got performance improvement. From the perspective of tasks, our onset/offset-based model completes the composite task, a complementary combination of speaker extraction and speaker-dependent voice activity detection. We also combined voiceprint with onset/offset cues. Voiceprint models voice characteristics of the target while onset/offset models the start/end information of the speech. From the perspective of auditory scene analysis, the combination of two perception cues can promote the integrity of the auditory object. The experiment results are also close to state-of-the-art performance, using nearly half of the parameters. We hope that this work will inspire communities of speech processing and psychology, and contribute to communication between them. Our code will be available in https://github.com/aispeech-lab/wase/.
Yunzhe Hao, Jiaming Xu 0001, Bo Xu 0002
ICASSP4
2021 Speaker and Direction Inferred Dual-Channel Speech Separation
abstract
Most speech separation methods, trying to separate all channel sources simultaneously, are still far from having enough generalization capabilities for real scenarios where the number of input sounds is usually uncertain and even dynamic. In this work, we employ ideas from auditory attention with two ears and propose a speaker and direction inferred speech separation network (dubbed SDNet) to solve the cocktail party problem. Specifically, our SDNet first parses out the respective perceptual representations with their speaker and direction characteristics from the mixture of the scene in a sequential manner. Then, the perceptual representations are utilized to attend to each corresponding speech. Our model generates more precise perceptual representations with the help of spatial features and successfully deals with the problem of the unknown number of sources and the selection of outputs. The experiments on standard fully-overlapped speech separation benchmarks, WSJ0-2mix, WSJ0-3mix, and WSJ0-2&3mix, show the effectiveness, and our method achieves SDR improvements of 25.31 dB, 17.26 dB, and 21.56 dB under anechoic settings. Our codes will be released at https://github.com/aispeech-lab/SDNet.
Chenxing Li, Jiaming Xu 0001, Nima Mesgarani, Bo Xu 0002
ICASSP4
2021 MixSpeech: Data Augmentation for Low-Resource Automatic Speech Recognition
abstract
In this paper, we propose MixSpeech, a simple yet effective data augmentation method based on mixup for automatic speech recognition (ASR). MixSpeech trains an ASR model by taking a weighted combination of two different speech features (e.g., mel-spectrograms or MFCC) as the input, and recognizing both text sequences, where the two recognition losses use the same combination weight. We apply MixSpeech on two popular end-to-end speech recognition models including LAS (Listen, Attend and Spell) and Transformer, and conduct experiments on several low-resource datasets including TIMIT, WSJ, and HKUST. Experimental results show that MixSpeech achieves better accuracy than the baseline models without data augmentation, and outperforms a strong data augmentation method SpecAugment on these recognition tasks. Specifically, MixSpeech outperforms SpecAugment with a relative PER improvement of 10.6% on TIMIT dataset, and achieves a strong WER of 4.7% on WSJ dataset.
Linghui Meng 0001, Jin Xu 0010, Xu Tan 0003, Jindong Wang 0001, Tao Qin 0001, Bo Xu 0002
ICASSP6
2021 Two-Stage Pre-Training for Sequence to Sequence Speech Recognition
abstract
The attention-based encoder-decoder structure is popular in automatic speech recognition (ASR). However, it relies heavily on transcribed data. In this paper, we propose a novel pre-training strategy for the encoder-decoder sequence-to-sequence (seq2seq) model by utilizing unpaired speech and transcripts. The pre-training process consists of two stages, acoustic pre-training and linguistic pre-training. In the acoustic pre-training stage, we use a large amount of speech to pre-train the encoder by predicting masked speech feature chunks with their contexts. In the linguistic pre-training stage, we first generate synthesized speech from a large number of transcripts using a text-to-speech (TTS) system and then use the synthesized paired data to pre-train the decoder. The two-stage pre-training is conducted on the AISHELL-2 dataset, and we apply this pre-trained model to multiple subsets of AISHELL-1 and HKUST for post-training. As the size of the subset increases, we obtain relative character error rate reduction (CERR) from 38.24% to 7.88% on AISHELL-1 and from 12.00% to 1.20% on HKUST.
Zhiyun Fan, Bo Xu 0002
IJCNN3
2021 Towards Modeling Auditory Restoration in Noisy Environments
Yunzhe Hao, Jiaming Xu 0001, Bo Xu 0002
IJCNN4
2021 Transfer Ability of Monolingual Wav2vec2.0 for Low-resource Speech Recognition
abstract
Recently, there are several domains that have their own feature extractors, such as ResNet, BERT, and GPT-x, which are widely used for various down-stream tasks. These models are pre-trained on large amounts of unlabeled data by self-supervision. In the speech domain, wav2vec2.0 starts to show its powerful representation ability and feasibility for ultra-low resource speech recognition tasks. This speech feature extractor is pre-trained on the monolingual audiobook corpus, whereas it has not been thoroughly examined in real spoken scenarios and other languages. In this work, we endeavor to transfer the knowledge from the pre-trained monolingual wav2vec2.0 to cross-lingual spoken ASR tasks with less than 20 hours of labeled data. We achieve more than 20% relative improvements in all the six languages compared with previous methods, establishing a strong benchmark on CALLHOME datasets. Compared with supervised pre-training, self-supervision training used in wav2vec2.0 has a better transfer ability. We also find that using coarse-grained modeling units, such as subword or character, usually achieves better results than fine-grained modeling units, such as phone or letter.
Jianzong Wang, Ning Cheng 0001, Bo Xu 0002
IJCNN5
2021 A Language Model Based Pseudo-Sample Deliberation for Semi-supervised Speech Recognition
abstract
End-to-end modeling requires tremendous amounts of transcribed speech to achieve an automatic speech recognition (ASR) model with high performance. For low-resource ASR tasks, it is a promising approach to utilize the highly accessible unlabeled speech and text corpus. Previous works have shown that training with pseudo samples, which are the inferring results given the unlabeled speech, can substantially improve the accuracy of a baseline ASR model. Besides the common data filtering to improve pseudo-label quality, we propose an alternative pseudo-sample deliberation method that operates on the output of the ASR model through a pre-trained bidirectional language model (BERT). It fixes the unreasonable tokens in the inference by substitution, which can distill knowledge from the large text corpus. Experiments on Librispeech show that assisted with our fixing operation, self-training on additional unlabeled samples can bridge up to 82.3 % of the gap with the supervised training.
Jianzong Wang, Ning Cheng 0001, Bo Xu 0002
IJCNN5
2021 Audio-Visual Speech Separation with Visual Features Enhanced by Adversarial Training
abstract
Audio-visual speech separation (AVSS) refers to separating individual voice from an audio mixture of multiple simultaneous talkers by conditioning on visual features. For the AVSS task, visual features play an important role, based on which we manage to extract more effective visual features to improve the performance. In this paper, we propose a novel AVSS model that uses speech-related visual features for isolating the target speaker. Specifically, the method of extracting speech-related visual features has two steps. Firstly, we extract the visual features that contain speech-related information by learning joint audio-visual representation. Secondly, we use the adversarial training method to enhance speech-related information in visual features further. We adopt the time-domain approach and build audio-visual speech separation networks with temporal convolutional neural networks block. Experiments on four audio-visual datasets, including GRID, TCD-TIMIT, AVSpeech, and LRS2, show that our model significantly outperforms previous state-of-the-art AVSS models. We also demonstrate that our model can achieve excellent speech separation performance in noisy realworld scenarios. Moreover, in order to alleviate the performance degradation of AVSS models caused by the missing of some video frames, we propose a training strategy, which makes our model robust when video frames are partially missing. The demo, code, and supplementary materials can be available at https://github.com/aispeech-lab/advr-avss.
Jiaming Xu 0001, Jing Shi 0003, Yunzhe Hao, Bo Xu 0002
IJCNN6
2021 Exploring wav2vec 2.0 on Speaker Verification and Language Identification
abstract
Wav2vec 2.0 is a recently proposed self-supervised framework for speech representation learning.It follows a two-stage training process of pre-training and fine-tuning, and performs well in speech recognition tasks especially ultra-low resource cases.In this work, we attempt to extend the self-supervised framework to speaker verification and language identification.First, we use some preliminary experiments to indicate that wav2vec 2.0 can capture the information about the speaker and language.Then we demonstrate the effectiveness of wav2vec 2.0 on the two tasks respectively.For speaker verification, we obtain a new state-of-the-art result, Equal Error Rate (EER) of 3.61% on the VoxCeleb1 dataset.For language identification, we obtain an EER of 12.02% on the 1 second condition and an EER of 3.47% on the full-length condition of the AP17-OLR dataset.Finally, we utilize one model to achieve the unified modeling by the multi-task learning for the two tasks.
Zhiyun Fan, Bo Xu 0002
Interspeech4
2021 MIMO Self-Attentive RNN Beamformer for Multi-Speaker Speech Separation
abstract
Recently, our proposed recurrent neural network (RNN) based all deep learning minimum variance distortionless response (ADL-MVDR) beamformer method yielded superior performance over the conventional MVDR by replacing the matrix inversion and eigenvalue decomposition with two RNNs.In this work, we present a self-attentive RNN beamformer to further improve our previous RNN-based beamformer by leveraging on the powerful modeling capability of self-attention.Temporal-spatial self-attention module is proposed to better learn the beamforming weights from the speech and noise spatial covariance matrices.The temporal self-attention module could help RNN to learn global statistics of covariance matrices.The spatial self-attention module is designed to attend on the cross-channel correlation in the covariance matrices.Furthermore, a multi-channel input with multi-speaker directional features and multi-speaker speech separation outputs (MIMO) model is developed to improve the inference efficiency.The evaluations demonstrate that our proposed MIMO self-attentive RNN beamformer improves both the automatic speech recognition (ASR) accuracy and the perceptual estimation of speech quality (PESQ) against prior arts.
Xiyun Li, Yong Xu 0004, Meng Yu 0003, Shixiong Zhang 0001, Jiaming Xu 0001, Bo Xu 0002, Dong Yu 0001
Interspeech6
2021 Counterfactual Supporting Facts Extraction for Explainable Medical Record Based Diagnosis with Graph Network
abstract
Providing a reliable explanation for clinical diagnosis based on the Electronic Medical Record (EMR) is fundamental to the application of Artificial Intelligence in the medical field.Current methods mostly treat the EMR as a text sequence and provide explanations based on a precise medical knowledge base, which is disease-specific and difficult to obtain for experts in reality.Therefore, we propose a counterfactual multi-granularity graph supporting facts extraction (CMGE) method to extract supporting facts from the irregular EMR itself without external knowledge bases in this paper.Specifically, we first structure the sequence of the EMR into a hierarchical graph network and then obtain the causal relationship between multi-granularity features and diagnosis results through counterfactual intervention on the graph.Features having the strongest causal connection with the results provide interpretive support for the diagnosis.Experimental results on real Chinese EMRs of the lymphedema demonstrate that our method can diagnose four types of EMRs correctly, and can provide accurate supporting facts for the results.More importantly, the results on different diseases demonstrate the robustness of our approach, which represents the potential application in the medical field 1 .
Wei Chen 0048, Bo Xu 0002
NAACL-HLT4
2021 Efficiently Fusing Pretrained Acoustic and Linguistic Encoders for Low-Resource Speech Recognition
abstract
End-to-end models have achieved impressive results on the task of automatic speech recognition (ASR). For low-resource ASR tasks, however, labeled data can hardly satisfy the demand of end-to-end models. Self-supervised acoustic pre-training has already shown its impressive ASR performance, while the transcription is still inadequate for language modeling in end-to-end models. In this work, we fuse a pre-trained acoustic encoder (wav2vec2.0) and a pre-trained linguistic encoder (BERT) into an end-to-end ASR model. The fused model only needs to learn the transfer from speech to language during fine-tuning on limited labeled data. The length of the two modalities is matched by a monotonic attention mechanism without additional parameters. Besides, a fully connected layer is introduced for the hidden mapping between modalities. We further propose a scheduled fine-tuning strategy to preserve and utilize the text context modeling ability of the pre-trained linguistic encoder. Experiments show our effective utilizing of pre-trained modules. Our model achieves better recognition performance on CALLHOME corpus (15 hours) than other end-to-end models.
Bo Xu 0002
IEEE Signal Process. Lett.3
2021 Simultaneous Control in Belief Space for Circular Insertion in Precision Assembly
abstract
Simultaneously inserting multiple objects is an essential topic in precision assembly to compose complicated shapes. This task involves the acquisition difficulty of unobservable interaction states, which makes it hard to plan the insertion. To solve it, this article investigates the circular assembly of multiple objects and proposes a strategy to control the simultaneous insertion in belief space. We first present the insertion state transition and observation models, in which the stochastic parts are modeled as Gaussian noise, and then estimate the belief state using an extended Kalman filter. An optimization approach is discussed, for the compensational movement planning, to decrease the estimated radial interaction forces and the toward-center movement is thus determined, considering the optimized compensational movement and the belief state. Experiments are carried out to demonstrate the validation of the proposed method.
Dengpeng Xing, Fangfang Liu 0006, De Xu, Bo Xu 0002
IEEE Trans. Ind. Informatics4
2020 DMRM: A Dual-Channel Multi-Hop Reasoning Model for Visual Dialog
abstract
Visual Dialog is a vision-language task that requires an AI agent to engage in a conversation with humans grounded in an image. It remains a challenging task since it requires the agent to fully understand a given question before making an appropriate response not only from the textual dialog history, but also from the visually-grounded information. While previous models typically leverage single-hop reasoning or single-channel reasoning to deal with this complex multimodal reasoning task, which is intuitively insufficient. In this paper, we thus propose a novel and more powerful Dual-channel Multi-hop Reasoning Model for Visual Dialog, named DMRM. DMRM synchronously captures information from the dialog history and the image to enrich the semantic representation of the question by exploiting dual-channel reasoning. Specifically, DMRM maintains a dual channel to obtain the question- and history-aware image features and the question- and image-aware dialog history features by a mulit-hop reasoning process in each channel. Additionally, we also design an effective multimodal attention to further enhance the decoder to generate more accurate responses. Experimental results on the VisDial v0.9 and v1.0 datasets demonstrate that the proposed model is effective and outperforms compared models by a significant margin.
Fandong Meng, Jiaming Xu 0001, Peng Li 0030, Bo Xu 0002, Jie Zhou 0016
AAAI5
2020 Knowledge Aware Emotion Recognition in Textual Conversations via Multi-Task Incremental Transformer
abstract
Emotion recognition in textual conversations (ERTC) plays an important role in a wide range of applications, such as opinion mining, recommender systems, and so on.ERTC, however, is a challenging task.For one thing, speakers often rely on the context and commonsense knowledge to express emotions; for another, most utterances contain neutral emotion in conversations, as a result, the confusion between a few non-neutral utterances and much more neutral ones restrains the emotion recognition performance.In this paper, we propose a novel Knowledge Aware Incremental Transformer with Multi-task Learning (KAITML) to address these challenges.Firstly, we devise a dual-level graph attention mechanism to leverage commonsense knowledge, which augments the semantic information of the utterance.Then we apply the Incremental Transformer to encode multi-turn contextual utterances.Moreover, we are the first to introduce multi-task learning to alleviate the aforementioned confusion and thus further improve the emotion recognition performance.Extensive experimental results show that our KAITML model outperforms the state-of-the-art models across five benchmark datasets.
Duzhen Zhang, Xiuyi Chen, Bo Xu 0002
COLING4
2020 Bridging the Gap between Prior and Posterior Knowledge Selection for Knowledge-Grounded Dialogue Generation
abstract
Knowledge selection plays an important role in knowledge-grounded dialogue, which is a challenging task to generate more informative responses by leveraging external knowledge.Recently, latent variable models have been proposed to deal with the diversity of knowledge selection by using both prior and posterior distributions over knowledge and achieve promising performance.However, these models suffer from a huge gap between prior and posterior knowledge selection.Firstly, the prior selection module may not learn to select knowledge properly because of lacking the necessary posterior information.Secondly, latent variable models suffer from the exposure bias that dialogue generation is based on the knowledge selected from the posterior distribution at training but from the prior distribution at inference.Here, we deal with these issues on two aspects: (1) We enhance the prior selection module with the necessary posterior information obtained from the specially designed Posterior Information Prediction Module (PIPM); (2) We propose a Knowledge Distillation Based Training Strategy (KDBTS) to train the decoder with the knowledge selected from the prior distribution, removing the exposure bias of knowledge selection.Experimental results on two knowledge-grounded dialogue datasets show that both PIPM and KDBTS achieve performance improvement over the state-of-theart latent variable model and their combination shows further improvement.
Xiuyi Chen, Fandong Meng, Peng Li 0030, Bo Xu 0002, Jie Zhou 0016
EMNLP (1)6
2020 CIF: Continuous Integrate-And-Fire for End-To-End Speech Recognition
abstract
In this paper, we propose a novel soft and monotonic alignment mechanism used for sequence transduction. It is inspired by the integrate-and-fire model in spiking neural networks and employed in the encoder-decoder framework consists of continuous functions, thus being named as: Continuous Integrate-and-Fire (CIF). Applied to the ASR task, CIF not only shows a concise calculation, but also supports online recognition and acoustic boundary positioning, thus suitable for various ASR scenarios. Several support strategies are also proposed to alleviate the unique problems of CIF-based model. With the joint action of these methods, the CIF-based model shows competitive performance. Notably, it achieves a word error rate (WER) of 2.86% on the test-clean of Librispeech and creates new state-of-the-art result on Mandarin telephone ASR benchmark.
Linhao Dong, Bo Xu 0002
ICASSP2
2020 Low-Frequency Guided Self-Supervised Learning For High-Fidelity 3d Face Reconstruction In The Wild
abstract
In this paper, we propose a low-frequency guided self-supervised learning method for high-fidelity 3D face reconstruction from an in-the-wild image. Unlike other self-supervised methods only using the color difference between the original image and the estimated image, we add low-frequency albedo information to enhance the self-supervised learning for more realistic albedo while insensitive to the non-skin regions. Specifically, based on a PCA albedo model, we first train a Boosting Network (B-Net) to provide illumination and intact albedo distribution. Then with above information, we learn an image-to-image non-linear Facial Albedo Network (FAN) by self-supervision to produce a high-fidelity albedo. We further propose a Detail Recovering Network (DRN) to recover geometric details such as wrinkles. FAN and DRN permit to reconstruct 3D faces with high-fidelity albedo and geometry details. Finally, experimental results demonstrate the effectiveness of the proposed method.
Pengrui Wang, Chunze Lin, Bo Xu 0002, Wujun Che
ICME3
2020 Class-Balanced Loss for Scene Text Detection
Randong Huang, Bo Xu 0002
ICONIP (2)2
2020 LISNN: Improving Spiking Neural Networks with Lateral Interactions for Robust Object Recognition
abstract
Spiking Neural Network (SNN) is considered more biologically plausible and energy-efficient on emerging neuromorphic hardware. Recently backpropagation algorithm has been utilized for training SNN, which allows SNN to go deeper and achieve higher performance. However, most existing SNN models for object recognition are mainly convolutional structures or fully-connected structures, which only have inter-layer connections, but no intra-layer connections. Inspired by Lateral Interactions in neuroscience, we propose a high-performance and noise-robust Spiking Neural Network (dubbed LISNN). Based on the convolutional SNN, we model the lateral interactions between spatially adjacent neurons and integrate it into the spiking neuron membrane potential formula, then build a multi-layer SNN on a popular deep learning framework, i.\,e., PyTorch. We utilize the pseudo-derivative method to solve the non-differentiable problem when applying backpropagation to train LISNN and test LISNN on multiple standard datasets. Experimental results demonstrate that the proposed model can achieve competitive or better performance compared to current state-of-the-art spiking neural networks on MNIST, Fashion-MNIST, and N-MNIST datasets. Besides, thanks to lateral interactions, our model processes stronger noise-robustness than other SNN. Our work brings a biologically plausible mechanism into SNN, hoping that it can help us understand the visual information processing in the brain.
Yunzhe Hao, Jiaming Xu 0001, Bo Xu 0002
IJCAI4
2020 Speaker-Conditional Chain Model for Speech Separation and Extraction
abstract
Speech separation has been extensively explored to tackle the cocktail party problem. However, these studies are still far from having enough generalization capabilities for real scenarios. In this work, we raise a common strategy named Speaker-Conditional Chain Model to process complex speech recordings. In the proposed method, our model first infers the identities of variable numbers of speakers from the observation based on a sequence-to-sequence model. Then, it takes the information from the inferred speakers as conditions to extract their speech sources. With the predicted speaker information from whole observation, our model is helpful to solve the problem of conventional speech separation and speaker extraction for multi-round long recordings. The experiments from standard fully-overlapped speech separation benchmarks show comparable results with prior studies, while our proposed model gets better adaptability for multi-round long recordings.
Jing Shi 0003, Jiaming Xu 0001, Yusuke Fujita, Shinji Watanabe 0001, Bo Xu 0002
INTERSPEECH5
2020 A Unified Framework for Low-Latency Speaker Extraction in Cocktail Party Environments
Yunzhe Hao, Jiaming Xu 0001, Jing Shi 0003, Bo Xu 0002
INTERSPEECH6
2020 Sequence to Multi-Sequence Learning via Conditional Chain Mapping for Mixture Signals
abstract
Neural sequence-to-sequence models are well established for applications which can be cast as mapping a single input sequence into a single output sequence. In this work, we focus on one-to-many sequence transduction problems, such as extracting multiple sequential sources from a mixture sequence. We extend the standard sequence-to-sequence model to a conditional multi-sequence model, which explicitly models the relevance between multiple output sequences with the probabilistic chain rule. Based on this extension, our model can conditionally infer output sequences one-by-one by making use of both input and previously-estimated contextual output sequences. This model additionally has a simple and efficient stop criterion for the end of the transduction, making it able to infer the variable number of output sequences. We take speech data as a primary test field to evaluate our methods since the observed speech data is often composed of multiple sources due to the nature of the superposition principle of sound waves. Experiments on several different tasks including speech separation and multi-speaker speech recognition show that our conditional multi-sequence models lead to consistent improvements over the conventional non-conditional models.
Jing Shi 0003, Xuankai Chang, Shinji Watanabe 0001, Yusuke Fujita, Jiaming Xu 0001, Bo Xu 0002, Lei Xie 0001
NeurIPS7
2020 A biologically plausible supervised learning method for spiking neural networks using the symmetric STDP rule
Yunzhe Hao, Xuhui Huang, Meng Dong, Bo Xu 0002
Neural Networks4
2020 Chinese Short Text Classification with Mutual-Attention Convolutional Neural Networks
abstract
The methods based on the combination of word-level and character-level features can effectively boost performance on Chinese short text classification. A lot of works concatenate two-level features with little processing, which leads to losing feature information. In this work, we propose a novel framework called Mutual-Attention Convolutional Neural Networks, which integrates word and character-level features without losing too much feature information. We first generate two matrices with aligned information of two-level features by multiplying word and character features with a trainable matrix. Then, we stack them as a three-dimensional tensor. Finally, we generate the integrated features using a convolutional neural network. Extensive experiments on six public datasets demonstrate improved performance of our new framework over current methods.
Bo Xu 0002, Jing-Yi Liang, Bowen Zhang 0011, Xu-Cheng Yin
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2019 Adapting Translation Models for Transcript Disfluency Detection
abstract
Transcript disfluency detection (TDD) is an important component of the real-time speech translation system, which arouses more and more interests in recent years. This paper presents our study on adapting neural machine translation (NMT) models for TDD. We propose a general training framework for adapting NMT models to TDD task rapidly. In this framework, the main structure of the model is implemented similar to the NMT model. Additionally, several extended modules and training techniques which are independent of the NMT model are proposed to improve the performance, such as the constrained decoding, denoising autoencoder initialization and a TDD-specific training object. With the proposed training framework, we achieve significant improvement. However, it is too slow in decoding to be practical. To build a feasible and production-ready solution for TDD, we propose a fast non-autoregressive TDD model following the non-autoregressive NMT model emerged recently. Even we do not assume the specific architecture of the NMT model, we build our TDD model on the basis of Transformer, which is the state-of-the-art NMT model. We conduct extensive experiments on the publicly available set, Switchboard, and in-house Chinese set. Experimental results show that the proposed model significantly outperforms previous state-ofthe-art models.
Qianqian Dong, Feng Wang 0023, Zhen Yang 0007, Wei Chen 0048, Bo Xu 0002
AAAI6
2019 A Working Memory Model for Task-oriented Dialog Response Generation
abstract
Recently, to incorporate external Knowledge Base (KB) information, one form of world knowledge, several end-to-end task-oriented dialog systems have been proposed.These models, however, tend to confound the dialog history with KB tuples and simply store them into one memory.Inspired by the psychological studies on working memory, we propose a working memory model ( WMM2Seq) for dialog response generation.Our WMM2Seq adopts a working memory to interact with two separated long-term memories, which are the episodic memory for memorizing dialog history and the semantic memory for storing KB tuples.The working memory consists of a central executive to attend to the aforementioned memories, and a short-term storage system to store the "activated" contents from the longterm memories.Furthermore, we introduce a context-sensitive perceptual process for the token representations of the dialog history, and then feed them into the episodic memory.Extensive experiments on two task-oriented dialog datasets demonstrate that our WMM2Seq significantly outperforms the state-of-the-art results in several evaluation metrics.
Xiuyi Chen, Jiaming Xu 0001, Bo Xu 0002
ACL (1)3
2019 Speaker-Aware Speech-Transformer
abstract
Recently, end-to-end (E2E) models become a competitive alternative to the conventional hybrid automatic speech recognition (ASR) systems. However, they still suffer from speaker mismatch in training and testing condition. In this paper, we use Speech-Transformer (ST) as the study platform to investigate speaker aware training of E2E models. We propose a model called Speaker-Aware Speech-Transformer (SAST), which is a standard ST equipped with a speaker attention module (SAM). The SAM has a static speaker knowledge block (SKB) that is made of i-vectors. At each time step, the encoder output attends to the i-vectors in the block, and generates a weighted combined speaker embedding vector, which helps the model to normalize the speaker variations. The SAST model trained in this way becomes independent of specific training speakers and thus generalizes better to unseen testing speakers. We investigate different factors of SAM. Experimental results on the AISHELL-1 task show that SAST achieves a relative 6.5% CER reduction (CERR) over the speaker-independent (SI) baseline. Moreover, we demonstrate that SAST still works quite well even if the i-vectors in SKB all come from a different data source other than the acoustic training set.
Zhiyun Fan, Jie Li 0032, Bo Xu 0002
ASRU4
2019 Self-attention Aligner: A Latency-control End-to-end Model for ASR Using Self-attention Network and Chunk-hopping
abstract
Self-attention network, an attention-based feedforward neural network, has recently shown the potential to replace recurrent neural networks (RNNs) in a variety of NLP tasks. However, it is not clear if the self-attention network could be a good alternative of RNNs in automatic speech recognition (ASR), which processes the longer speech sequences and may have online recognition requirements. In this paper, we present a RNN-free end-to-end model: self-attention aligner (SAA), which applies the self-attention networks to a simplified recurrent neural aligner (RNA) framework. We also propose a chunk-hopping mechanism, which enables the SAA model to encode on segmented frame chunks one after another to support online recognition. Experiments on two Mandarin ASR datasets show the replacement of RNNs by the self-attention networks yields a 8.4%-10.2% relative character error rate (CER) reduction. In addition, the chunk-hopping mechanism allows the SAA to have only a 2.5% relative CER degradation with a 320ms latency. After jointly training with a self-attention network language model, our SAA model obtains further error rate reduction on multiple datasets. Especially, it achieves 24.12% CER on the Mandarin ASR benchmark (HKUST), exceeding the best end-to-end model by over 2% absolute CER.
Linhao Dong, Feng Wang 0023, Bo Xu 0002
ICASSP3
2019 NRTR: A No-Recurrence Sequence-to-Sequence Model for Scene Text Recognition
abstract
Scene text recognition has attracted a great many researches due to its importance to various applications. Existing methods mainly adopt recurrence or convolution based networks. Though have obtained good performance, these methods still suffer from two limitations: slow training speed due to the internal recurrence of RNNs, and high complexity due to stacked convolutional layers for long-term feature extraction. This paper, for the first time, proposes a no-recurrence sequence-to-sequence text recognizer, named NRTR, that dispenses with recurrences and convolutions entirely. NRTR follows the encoder-decoder paradigm, where the encoder uses stacked self-attention to extract image features, and the decoder applies stacked self-attention to recognize texts based on encoder output. NRTR relies solely on self-attention mechanism thus could be trained with more parallelization and less complexity. Considering scene image has large variation in text and background, we further design a modality-transform block to effectively transform 2D input images to 1D sequences, combined with the encoder to extract more discriminative features. NRTR achieves state-of-the-art or highly competitive performance on both regular and irregular benchmarks, while requires only a small fraction of training time compared to the best model from the literature (at least 8 times faster).
Fenfen Sheng, Zhineng Chen, Bo Xu 0002
ICDAR3
2019 Efficient and Accurate Face Shape Reconstruction by Fusion of Multiple Landmark Databases
abstract
We propose an efficient and accurate regression-based 3D face shape reconstruction method. We use an encoder based on MobileNet to estimate parameters including face pose and coefficients of a parametric face model from a single face image. The encoder is trained only by 2D landmarks. Faces can be reconstructed by these parameters. Three contributions of our method are: 1) we propose a databases fusion method to train our network which can easily utilize multiple 2D landmark databases which have different landmark numbers and positions; 2) with the fusion method, we propose a simple MobileNet based network which is efficient, accurate and robust for face reconstruction even without complex training strategies; 3) we add an additional deformation field for shape correction to further improve our network's performance. Experiments demonstrate our method can bring about great performance improvement on most test databases and also compare favorably to some state-of-the-art methods in performance and speed.
Pengrui Wang, Wujun Che, Bo Xu 0002
ICIP4
2019 A Single-Shot Oriented Scene Text Detector with Learnable Anchors
abstract
Current regression based text detectors mainly use fixed anchors, where scales and positions can not be changed during network training. As scene texts tend to have large variation in orientations, aspect ratios and sizes, fixed anchors are insufficient to cover all varieties. This paper proposes a novel text detector with learnable anchors, named LATD. LATD contains two prediction branches. One aims to refine scales and locations of anchors according to the characteristics of scene texts. The other one receives refined anchors as defaults and regresses their offsets to text regions. These two branches are optimized jointly without sacrifices much speed. Meanwhile, we explore the class-imbalance issue between texts and backgrounds, and replace softmax loss with focal loss. Extensive experiments on both oriented and horizontal benchmarks demonstrate the effectiveness of LATD with new state-of-the-art performance. By visualizing qualitative results, as expected, LATD provides more accurate locations and lower rate of missed detections.
Fenfen Sheng, Zhineng Chen, Tao Mei 0001, Bo Xu 0002
ICME4
2019 Strong-Background Restrained Cross Entropy Loss for Scene Text Detection
abstract
In this paper, we investigate the issue of class imbalance in scene text detection. Class Balanced Cross Entropy (CBCE) loss is often adopted for addressing this imbalance problem. We find that CBCE excessively restrains the backward gradients of background. Negative samples own extremely small weights which are offered by CBCE during training of text detectors. These tiny weight values lead to insufficient learning of background. As a result, the CBCE-based text detection methods only can achieve sub-optimal performance. We propose a novel loss function, Strong-Background Restrained Cross Entropy (SBRCE), to deal with the disadvantage in CBCE. Specifically, SBRCE effectively down-weights the loss assigned to the strong background which means well-classified negative samples. Our SBRCE can make training focused on all positive samples and weak background(i.e., hard-classified negative samples). Moreover, it can prevent the enormous amount of strong background from overwhelming text detectors during training. Experimental results show that the proposed SBRCE can improve the performance of the efficient and accurate scene text detector (EAST) by F-score of 3.3% on ICDAR2015 dataset and 1.12% on MSRA-TD500 dataset, without sacrificing the training and testing speed of EAST.
Randong Huang, Bo Xu 0002
IJCNN2
2019 Text Attention and Focal Negative Loss for Scene Text Detection
abstract
This paper proposes a novel attention mechanism and a fancy loss function for scene text detectors. Specifically, the attention mechanism can effectively identify the text regions by learning an attention mask automatically. The fine-grained attention mask is directly incorporated into the convolutional feature maps of a neural network to produce graininess-aware feature maps, which essentially obstruct the background inference and especially emphasize the text regions. Therefore, our graininess-aware feature maps concentrate on text regions, in especial those of exceedingly small size. Additionally, to address the extreme text-background class imbalance during training, we also propose a newfangled loss function, named Focal Negative Loss (FNL). The proposed loss function is able to down-weight the loss assigned to easy negative samples. Consequently, the proposed FNL can make training focused on hard negative samples. To evaluate the effectiveness of our text attention module and FNL, we integrate them into the efficient and accurate scene text detector (EAST). The comprehensive experimental results demonstrate that our text attention module and FNL can increase the performance of EAST by F-score of 3.98% on ICDAR2015 dataset and 1.87% on MSRA-TD500 dataset.
Randong Huang, Bo Xu 0002
IJCNN2
2019 A Unified Multi-output Semi-supervised Network for 3D Face Reconstruction
abstract
In this paper, we propose a method to reconstruct fine-grained 3D faces from single images base on a nearly unified multi-output regression network. The network estimates the facial shape, normal and appearance jointly in 2D UV map which preserves spatial adjacency relations among vertexes and provides semantic meaning of each vertex. Three contributions of the proposed method are: 1) we generate the UV map by as-rigid-as-possible parametrization to address the overlapping problem caused by cylindrical unwarp; 2) we directly estimate face normal rather than compute it from the estimated shape to let it catch geometric details from face texture; 3) we propose a post process strategy to generating more realistic faces and to employing the estimated normal. Experiments show that our network is able to learn a uniform appearance and predict more accurate shape from the proposed UV map. Additionally, the post process procedure can improve the quality of facial shapes and add geometric details from estimated normals.
Pengrui Wang, Wujun Che, Bo Xu 0002
IJCNN4
2019 RevCuT Tree Search Method in Complex Single-player Game with Continuous Search Space
abstract
Monte-Carlo Tree Search (MCTS) has achieved great success in combinatorial game, which has the characteristics of finite action state space, deterministic state transition and sparse reward. AlphaGo Zero combined MCTS and deep neural networks defeated the world champion Lee Sedol in the Go game, proving the advantages of tree search in combinatorial game with enormous search space. However, when the search space is continuous and even with chance factors, tree search methods like UCT failed. Because each state will be visited repeatedly with probability zero and the information in tree will never be used, that is to say UCT algorithm degrades to Monte Carlo rollouts. Meanwhile, the previous exploration experiences cannot be used to correct the next tree search process, and makes a huge increase in the demand for computing resources. To solve this kind of problem, this paper proposes a step-by-step Reverse Curriculum Learning with Truncated Tree Search method (RevCuT Tree Search). In order to retain the previous exploration experiences, we use the deep neural network to learn the state-action values at explored states and then guide the next tree search process. Besides, taking the computing resources into consideration, we establish a truncated search tree focusing on continuous state space rather than the whole trajectory. This method can effectively reduce the number of explorations and achieve the effect beyond the human level in our well designed single-player game with continuous state space and probabilistic state transition.
Hongming Zhang 0003, Fangjuan Cheng, Bo Xu 0002
IJCNN3
2019 Boosting Character-Based Chinese Speech Synthesis via Multi-Task Learning and Dictionary Tutoring
Yuxiang Zou, Linhao Dong, Bo Xu 0002
INTERSPEECH3
2019 How social media usage affects employees' job satisfaction and turnover intention: An empirical study in China
Xin Zhang 0072, Bo Xu 0002
Inf. Manag.3
2019 Hybrid Attention for Chinese Character-Level Neural Machine Translation
Feng Wang 0023, Wei Chen 0048, Zhen Yang 0007, Bo Xu 0002
Neurocomputing5
2019 Effectively training neural machine translation models with monolingual data
Zhen Yang 0007, Wei Chen 0048, Feng Wang 0023, Bo Xu 0002
Neurocomputing4
2019 Concept learning through deep reinforcement learning with memory-augmented neural networks
Jing Shi 0003, Jiaming Xu 0001, Yiqun Yao, Bo Xu 0002
Neural Networks4
2019 Pyrboxes: An efficient multi-scale scene text detector with feature pyramids
Fenfen Sheng, Zhineng Chen, Wei Zhang 0031, Bo Xu 0002
Pattern Recognit. Lett.4
2018 Modeling Attention and Memory for Auditory Selection in a Cocktail Party Environment
abstract
Developing a computational auditory model to solve the cocktail party problem has long bedeviled scientists, especially for a single microphone recording. Although recent deep learning based frameworks have made significant progress in multi-talker mixed speech separation, most existing deep learning based methods, focusing on separating all the speech channels rather than selectively attending the target speech and ignoring other sounds, may fail to offer a satisfactory solution in a complex auditory scene where the number of input sounds is usually uncertain and even dynamic. In this work, we employ ideas from auditory selective attention of behavioral and cognitive neurosciences and from recent advances of memory-augmented neural networks. Specifically, a unified Auditory Selection framework with Attention and Memory (dubbed ASAM) is proposed. Our ASAM first accumulates the prior knowledge (that is the acoustic feature to one specific speaker) into a life-long memory during the training phase, meanwhile a speech perceptor is trained to extract the temporal acoustic feature and update the memory online when a salient speech is given. Then, the learned memory is utilized to interact with the mixture input to attend and filter the target frequency out from the mixture stream. Finally, the network is trained to minimize the reconstruction error of the attended speech. We evaluate the proposed approach on WSJ0 and THCHS-30 datasets and the experimental results demonstrate that our approach successfully conducts two auditory selection tasks: the top-down task-specific attention (e.g. to follow a conversation with friend) and the bottom-up stimulus-driven attention (e.g. be attracted by a salient speech). Compared with deep clustering based methods, our method conducts competitive advantages especially in a real noise environment (e.g. street junction). Our code is available at https://github.com/jacoxu/ASAM.
Jiaming Xu 0001, Jing Shi 0003, Guangcan Liu, Xiuyi Chen, Bo Xu 0002
AAAI5
2018 Unsupervised Neural Machine Translation with Weight Sharing
abstract
Unsupervised neural machine translation (NMT) is a recently proposed approach for machine translation which aims to train the model without using any labeled data.The models proposed for unsupervised NMT often use only one shared encoder to map the pairs of sentences from different languages to a shared-latent space, which is weak in keeping the unique and internal characteristics of each language, such as the style, terminology, and sentence structure.To address this issue, we introduce an extension by utilizing two independent encoders but sharing some partial weights which are responsible for extracting high-level representations of the input sentences.Besides, two different generative adversarial networks (GANs), namely the local GAN and global GAN, are proposed to enhance the cross-language translation.With this new approach, we achieve significant improvements on English-German, English-French and Chinese-to-English translation tasks.
Zhen Yang 0007, Wei Chen 0048, Feng Wang 0023, Bo Xu 0002
ACL (1)4
2018 Which Mapping Rule in the Fireworks Algorithm is Better for Large Scale Optimization
abstract
Fireworks algorithm(FWA), which is proposed for global optimization of complex function, becomes a hot spot in optimization field recently, caused by its competitive performance. Boundary handling for FWA, which maps the out-of-bound sparks into feasible space, is critical for its convergence efficiency. However, random mapping rule, which is widely used for boundary handling, always caused computing resource waste, especially for high-dimensional optimization. In this paper, we propose three novel mapping rules to speed up large scale optimization of FWA. Meanwhile, to evaluate the effectiveness of the new rules, we compare them by representative nine benchmark functions on different dimensionality scale. Experimental results indicate that the mirror rule which we proposed, achieve superior performance for most optimization functions.
Xuemei Yet, Junzhi Li 0001, Bo Xu 0002, Ying Tan 0002
CEC3
2018 Semi-Supervised Disfluency Detection
abstract
While the disfluency detection has achieved notable success in the past years, it still severely suffers from the data scarcity. To tackle this problem, we propose a novel semi-supervised approach which can utilize large amounts of unlabelled data. In this work, a light-weight neural net is proposed to extract the hidden features based solely on self-attention without any Recurrent Neural Network (RNN) or Convolutional Neural Network (CNN). In addition, we use the unlabelled corpus to enhance the performance. Besides, the Generative Adversarial Network (GAN) training is applied to enforce the similar distribution between the labelled and unlabelled data. The experimental results show that our approach achieves significant improvements over strong baselines.
Feng Wang 0023, Wei Chen 0048, Zhen Yang 0007, Qianqian Dong, Bo Xu 0002
COLING6
2018 Cascaded Mutual Modulation for Visual Reasoning
abstract
Visual reasoning is a special visual question answering problem that is multi-step and compositional by nature, and also requires intensive text-vision interactions.We propose CMM: Cascaded Mutual Modulation as a novel end-to-end visual reasoning model.CMM includes a multi-step comprehension process for both question and image.In each step, we use a Feature-wise Linear Modulation (FiLM) technique to enable textual/visual pipeline to mutually control each other.Experiments show that CMM significantly outperforms most related models, and reach stateof-the-arts on two visual reasoning benchmarks: CLEVR and NLVR, collected from both synthetic and natural languages.Ablation studies confirm that both our multistep framework and our visual-guided language modulation are critical to the task.Our code is available at https://github. com/FlamingHorizon/CMM-VR.
Yiqun Yao, Jiaming Xu 0001, Feng Wang 0023, Bo Xu 0002
EMNLP4
2018 Speech-Transformer: A No-Recurrence Sequence-to-Sequence Model for Speech Recognition
abstract
Recurrent sequence-to-sequence models using encoder-decoder architecture have made great progress in speech recognition task. However, they suffer from the drawback of slow training speed because the internal recurrence limits the training parallelization. In this paper, we present the Speech-Transformer, a no-recurrence sequence-to-sequence model entirely relies on attention mechanisms to learn the positional dependencies, which can be trained faster with more efficiency. We also propose a 2D-Attention mechanism, which can jointly attend to the time and frequency axes of the 2-dimensional speech inputs, thus providing more expressive representations for the Speech-Transformer. Evaluated on the Wall Street Journal (WSJ) speech recognition dataset, our best model achieves competitive word error rate (WER) of 10.9%, while the whole training process only takes 1.2 days on 1 GPU, significantly faster than the published results of recurrent sequence-to-sequence models.
Linhao Dong, Bo Xu 0002
ICASSP3
2018 CBLDNN-Based Speaker-Independent Speech Separation Via Generative Adversarial Training
abstract
In this paper, we propose a speaker-independent multi-speaker monaural speech separation system (CBLDNN-GAT) based on convolutional, bidirectional long short-term memory, deep feedforward neural network (CBLDNN) with generative adversarial training (GAT). Our system aims at obtaining better speech quality instead of only minimizing a mean square error (MSE). In the initial phase, we utilize log-mel filterbank and pitch features to warm up our CBLDNN in a multi-task manner. Thus, the information that contributes to separating speech and improving speech quality is integrated into the model. We execute GAT throughout the training, which makes the separated speech indistinguishable from the real one. We evaluate CBLDNN-GAT on WSJ0-2mix dataset. The experimental results show that the proposed model achieves 11.0d-B signal-to-distortion ratio (SDR) improvement, which is the new state-of-the-art result.
Chenxing Li, Bo Xu 0002
ICASSP5
2018 A Cascaded Framework for Model-Based 3D Face Reconstruction
abstract
This paper presents a general framework for model-based 3D face reconstruction from a single image, which can incorporate mature face alignment methods and utilize their properties. In the proposed framework, the final model parameters, i.e., mostly including pose, identity and expression, are achieved by estimating updating the face landmarks and 3D face model parameter alternately. In addition, we propose the parameter augmented regression method (PARM) as an novel derivation of the framework. Compared with existing methods, PARM is able to utilize mature face alignment methods and use fairly simple features in addition to image appearances for the reconstruction task. Experiments on three derivation methods of the framework show that the proposed framework is feasible and PARM is quite an effective and fast method. With face alignment method LBF, PARM can run over 90 fps on a desktop.
Pengrui Wang, Wujun Che, Bo Xu 0002
ICASSP3
2018 A Comparison of Modeling Units in Sequence-to-Sequence Speech Recognition with the Transformer on Mandarin Chinese
Linhao Dong, Bo Xu 0002
ICONIP (5)4
2018 Compression of Acoustic Model via Knowledge Distillation and Pruning
abstract
Recently, the performance of speech recognition system based on neural network has been greatly improved. Arguably, this huge improvement can be mainly attributed to deeper and wider layers. These systems are more difficult to be deployed on the embedded devices due to their large size and high computational complexity. To address these issues, we propose a method to compress deep feed-forward neural network (DNN) based acoustic model. In detail, a state-of-the-art acoustic model is trained as the baseline model. In this step, layer normalization is applied to accelerating the model convergence and improving the generalization performance. Knowledge distillation and pruning are then conducted to compress the model. Our final model can achieve 14.59× parameters reduction, 5× storage size reduction and comparable performance compared with the baseline model.
Chenxing Li, Bo Xu 0002
ICPR5
2018 Recurrent Neural Network Based Small-footprint Wake-up-word Speech Recognition System with a Score Calibration Method
abstract
In this paper, we propose a small-footprint wake-up-word speech recognition (WUWSR) system based on long short-term memory (LSTM) recurrent neural network, and we design a novel back-end calibration scoring method named modified zero normalization (MZN). First, LSTM is trained to predict posterior probability of context-dependent state. Next, MZN is adopted to transfer posterior probability to normalized score, which is then converted to confidence score by dynamic programming. Finally, a certain wake-up-word is recognized according to the confidence score. This WUWSR system can recognize multiple wake-up words and change wake-up words flexibly. This system can guarantee low latency by omitting decoding network. Equal error rate (EER) is adopted as the evaluation metric. Experimental results show that the proposed LSTM-based system achieves 33.33% relative improvement compared with a baseline system based on deep feed-forward neural network. Combining the front-end LSTM acoustic model with back-end MZN method, our WUWSR system can achieve 51.92% relative improvement.
Chenxing Li, Bo Xu 0002
ICPR5
2018 Self-Attention Based Network for Punctuation Restoration
abstract
Inserting proper punctuation into Automatic Speech Recognizer(ASR) transcription is a challenging and promising task in real-time Spoken Language Translation(SLT). Traditional methods built on the sequence labelling framework are weak in handling the joint punctuation. To tackle this problem, we propose a novel self-attention based network, which can solve the aforementioned problem very well. In this work, a light-weight neural net is proposed to extract the hidden features based solely on self-attention without any Recurrent Neural Nets(RNN) and Convolutional Neural Nets(CNN). We conduct extensive experiments on complex punctuation tasks. The experimental results show that the proposed model achieves significant improvements on joint punctuation task while being superior to traditional methods on simple punctuation task as well.
Feng Wang 0023, Wei Chen 0048, Zhen Yang 0007, Bo Xu 0002
ICPR4
2018 Unsupervised Domain Adaptation for Neural Machine Translation
abstract
Impressive neural machine translation (NMT) results are achieved in domains with large-scale, high quality bilingual training corpora. However, transferring to a target domain with significant domain shifts but no bilingual training corpora remains largely unexplored. To address the aforementioned setting of unsupervised domain adaptation, we propose a novel adversarial training procedure for NMT to leverage the widespread monolingual data in target domain. Two discriminative networks, namely the domain discriminator and pair discriminator, are introduced to guide the translation model. The domain discriminator evaluates whether the sentences generated by the translation model are indistinguishable from the ones in target domain. The pair discriminator assesses whether the generated sentences are paired with the source-side sentences. The translation model acts as an adversary to the two discriminators, which aims to generate sentences uneasily discriminated by the discriminators. We tested our approach on Chinese-English and English-German translation tasks. Experimental results show that our approaches achieve great success in unsupervised domain adaptation for NMT.
Zhen Yang 0007, Wei Chen 0048, Feng Wang 0023, Bo Xu 0002
ICPR4
2018 Listen, Think and Listen Again: Capturing Top-down Auditory Attention for Speaker-independent Speech Separation
abstract
Recent deep learning methods have made significant progress in multi-talker mixed speech separation. However, most existing models adopt a driftless strategy to separate all the speech channels rather than selectively attend the target one. As a result, those frameworks may be failed to offer a satisfactory solution in complex auditory scene where the number of input sounds is usually uncertain and even dynamic. In this paper, we present a novel neural network based structure motivated by the top-down attention behavior of human when facing complicated acoustical scene. Different from previous works, our method constructs an inference-attention structure to predict interested candidates and extract each speech channel of them. Our work gets rid of the limitation that the number of channels must be given or the high computation complexity for label permutation problem. We evaluated our model on the WSJ0 mixed-speech tasks. In all the experiments, our model gets highly competitive to reach and even outperform the baselines.
Jing Shi 0003, Jiaming Xu 0001, Guangcan Liu, Bo Xu 0002
IJCAI4
2018 Brain-inspired Balanced Tuning for Spiking Neural Networks
abstract
Due to the nature of Spiking Neural Networks (SNNs), it is challenging to be trained by biologically plausible learning principles. The multi-layered SNNs are with non-differential neurons, temporary-centric synapses, which make them nearly impossible to be directly tuned by back propagation. Here we propose an alternative biological inspired balanced tuning approach to train SNNs. The approach contains three main inspirations from the brain: Firstly, the biological network will usually be trained towards the state where the temporal update of variables are equilibrium (e.g. membrane potential); Secondly, specific proportions of excitatory and inhibitory neurons usually contribute to stable representations; Thirdly, the short-term plasticity (STP) is a general principle to keep the input and output of synapses balanced towards a better learning convergence. With these inspirations, we train SNNs with three steps: Firstly, the SNN model is trained with three brain-inspired principles; then weakly supervised learning is used to tune the membrane potential in the final layer for network classification; finally the learned information is consolidated from membrane potential into the weights of synapses by Spike-Timing Dependent Plasticity (STDP). The proposed approach is verified on the MNIST hand-written digit recognition dataset and the performance (the accuracy of 98.64%) indicates that the ideas of balancing state could indeed improve the learning ability of SNNs, which shows the power of proposed brain-inspired approach on the tuning of biological plausible SNNs.
Tielin Zhang, Yi Zeng 0001, Dongcheng Zhao, Bo Xu 0002
IJCAI4
2018 Distilled Binary Neural Network for Monaural Speech Separation
abstract
Monaural speech separation, aiming at solving the cocktail party problem, has many important application scenarios, most of which ask for the real-time response, high energy efficiency and efficient storage. However, the state-of-the-art Deep Neural Network based separation models usually require huge memory and computation for the 32-bit floating point multiply accumulations, hence most of them cannot meet those requirements. Recently, there are many methods proposed to solve the problem, and binary neural networks have drawn many attentions for they compress and speed up its counterparts at the cost of some performance. Hence, in this paper, we binarize Deep Neural Network based separation models, aiming to deploy them on embedded devices for real-time applications. Furthermore, we improve the separation performance by integrating knowledge distillation into the training phase of binary neural network based models, which is referred as Distilled Binary Neural Network (DBNN). To the best of our knowledge, DBNN is the first attempt to integrate two types of model compression. In the experiments, we demonstrate the effectiveness of our proposed method, which successfully binarizes the Deep Neural Network based separation models with a comparable performance.
Xiuyi Chen, Guangcan Liu, Jing Shi 0003, Jiaming Xu 0001, Bo Xu 0002
IJCNN5
2018 Improving Speech Separation with Adversarial Network and Reinforcement Learning
abstract
In contrast to the conventional deep neural network for single-channel speech separation, we propose a separation framework based on adversarial network and reinforcement learning. The purpose of the adversarial network inspired by the generative adversarial network is to make the separated result and ground-truth with the same data distribution by evaluating the discrepancy between them. Meanwhile, in order to enable the model to bias the generation towards desirable metrics and reduce the discrepancy between training loss (such as mean squared error) and testing metric (such as SDR), we present the future success based on reinforcement learning. We directly optimize the performance metric to accomplish exactly that. With the combination of adversarial network and reinforcement learning, our model is able to improve the performance of single-channel speech separation.
Guangcan Liu, Jing Shi 0003, Xiuyi Chen, Jiaming Xu 0001, Bo Xu 0002
IJCNN5
2018 Hierarchical Tree Long Short-Term Memory for Sentence Representations
abstract
A fixed-length feature vector is required for many machine learning algorithms in NLP field. Word embeddings have been very successful at learning lexical information. However, they can't capture the compositional meaning of sentences, which prevents them from a deeper understanding of language. In this paper, we introduce a novel hierarchical tree long short-term memory (HTLSTM) model that learns vector representations for sentences of arbitrary syntactic type and length. We propose to split one sentence into three hierarchies: short phrase, long phrase and full sentence level. The HTLSTM model gives our algorithm the potential to fully consider the hierarchical information and longterm dependencies of language. We design the experiments on both English and Chinese corpus to evaluate our model on sentiment analysis task. And the results show that our model outperforms several existing state of the art approaches significantly.
Xiuying Wang 0002, Changliang Li, Bo Xu 0002
IJCNN3
2018 Paraphrase Recognition via Combination of Neural Classifier and Keywords
abstract
Paraphrases are sentences or phrases that convey the same meaning using different words. Paraphrase recognition is of interest for many current Natural Language Processing (NLP) tasks. As understood in linguistics, the phenomenon of paraphrases is difficult to characterize. In this article, we present a novel approach to the task of paraphrase identification. The proposed approach measures similarity between two sentences based on both the lexical and semantic levels, via combining neural networks and keywords jointly. In particular, we employ a vector offset, which implies the relation of given inputs in vector space, as the representation of a neural classifier. We conduct experiments on the Microsoft Research Paraphrase Corpus (MSRP)1and SICK dataset, which are both standard datasets for evaluating approaches to paraphrase identification. The experiments showed that our proposed approach makes much progress and achieves state-of-the-art results.
Xiuying Wang 0002, Changliang Li, Zhijun Zheng, Bo Xu 0002
IJCNN4
2018 Syllable-Based Acoustic Modeling with CTC for Multi-Scenarios Mandarin speech recognition
abstract
With the improvement of speech recognition, voice products are gradually applied to every scene of life. The existing approaches to handle various scenarios are often to build many different acoustic models using scenario-dependent data only, with each for a special scene. The obvious weakness of these approaches is that it seriously hampers the large-scale application and maintenance of voice products. To address this issue, acoustic modeling based on context-independent syllables optimized with CTC loss is presented for multiple scenarios of Mandarin speech recognition. On the one hand, context-independent modeling overcomes the shortcomings of context-dependent modeling overfitting a particular scene. Also, it sidesteps decision trees used in context-dependent modeling so that there is no need to consider the building of decision tree and whether to start training again in a real application. On the other hand, choosing longer-length syllable acoustic units can effectively preserve the co-articulation effect that context-dependent phone can model. Also, syllables in the Chinese language have its inherent advantages, as its number is fixed and it is trainable, effective generalization and better robustness. This paper also explores the differences between wideband and narrowband data caused by the front-end signal acquisition block, and proposes a unified training method based on the use of VGG in the bottom layer, and introduces layer normalization. The experimental results demonstrate that the proposed syllable-based CTC acoustic model for multiple scenarios can achieve more than 15% and 7% relatively improvement for mobile phone data and telephone data separately compare with scenarios-dependent modeling.
Linhao Dong, Bo Xu 0002
IJCNN4
2018 Extending Recurrent Neural Aligner for Streaming End-to-End Speech Recognition in Mandarin
abstract
End-to-end models have been showing superiority in Automatic Speech Recognition (ASR).At the same time, the capacity of streaming recognition has become a growing requirement for end-to-end models.Following these trends, an encoder-decoder recurrent neural network called Recurrent Neural Aligner (RNA) has been freshly proposed and shown its competitiveness on two English ASR tasks.However, it is not clear if RNA can be further improved and applied to other spoken language.In this work, we explore the applicability of RNA in Mandarin Chinese and present four effective extensions: In the encoder, we redesign the temporal downsampling and introduce a powerful convolutional structure.In the decoder, we utilize a regularizer to smooth the output distribution and conduct joint training with a language model.On two Mandarin Chinese conversational telephone speech recognition (MTS) datasets, our Extended-RNA obtains promising performance.Particularly, it achieves 27.7% character error rate (CER), which is superior to current state-of-the-art result on the popular HKUST task.
Linhao Dong, Wei Chen 0048, Bo Xu 0002
INTERSPEECH4
2018 An End-to-End Text-Independent Speaker Identification System on Short Utterances
Ruifang Ji, Xinyuan Cai, Bo Xu 0002
INTERSPEECH3
2018 Single-channel Speech Dereverberation via Generative Adversarial Training
abstract
In this paper, we propose a single-channel speech dereverberation system (DeReGAT) based on convolutional, bidirectional long short-term memory and deep feed-forward neural network (CBLDNN) with generative adversarial training (GAT).In order to obtain better speech quality instead of only minimizing a mean square error (MSE), GAT is employed to make the dereverberated speech indistinguishable form the clean samples.Besides, our system can deal with wide range reverberation and be well adapted to variant environments.The experimental results show that the proposed model outperforms weighted prediction error (WPE) and deep neural network-based systems.In addition, DeReGAT is extended to an online speech dereverberation scenario, which reports comparable performance with the offline case.
Chenxing Li, Tieqiang Wang, Bo Xu 0002
INTERSPEECH4
2018 Syllable-Based Sequence-to-Sequence Speech Recognition with the Transformer in Mandarin Chinese
abstract
Sequence-to-sequence attention-based models have recently shown very promising results on automatic speech recognition (ASR) tasks, which integrate an acoustic, pronunciation and language model into a single neural network.In these models, the Transformer, a new sequence-to-sequence attention-based model relying entirely on self-attention without using RNNs or convolutions, achieves a new single-model state-of-the-art BLEU on neural machine translation (NMT) tasks.Since the outstanding performance of the Transformer, we extend it to speech and concentrate on it as the basic architecture of sequence-to-sequence attention-based model on Mandarin Chinese ASR tasks.Furthermore, we investigate a comparison between syllable based model and context-independent phoneme (CI-phoneme) based model with the Transformer in Mandarin Chinese.Additionally, a greedy cascading decoder with the Transformer is proposed for mapping CI-phoneme sequences and syllable sequences into word sequences.Experiments on HKUST datasets demonstrate that syllable based model with the Transformer performs better than CI-phoneme based counterpart, and achieves a character error rate (CER) of 28.77%, which is competitive to the state-of-the-art CER of 28.0% by the joint CTC-attention based encoder-decoder network.
Linhao Dong, Bo Xu 0002
INTERSPEECH4
2018 Improving Neural Machine Translation with Conditional Sequence Generative Adversarial Nets
abstract
Zhen Yang, Wei Chen, Feng Wang, Bo Xu. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Zhen Yang 0007, Wei Chen 0048, Feng Wang 0023, Bo Xu 0002
NAACL-HLT4
2018 Generative adversarial training for neural machine translation
Zhen Yang 0007, Wei Chen 0048, Feng Wang 0023, Bo Xu 0002
Neurocomputing4
2018 Efficient coding matters in the organization of the early visual system
Qingqun Kong, Jiuqi Han, Yi Zeng 0001, Bo Xu 0002
Neural Networks4
2018 Learning to activate logic rules for textual reasoning
Yiqun Yao, Jiaming Xu 0001, Jing Shi 0003, Bo Xu 0002
Neural Networks4
2018 Distant supervision for relation extraction with hierarchical selective attention
Peng Zhou 0009, Jiaming Xu 0001, Zhenyu Qi 0003, Hongyun Bao, Zhineng Chen, Bo Xu 0002
Neural Networks6
2017 Joint Extraction of Entities and Relations Based on a Novel Tagging Scheme
abstract
Joint extraction of entities and relations is an important task in information extraction.To tackle this problem, we firstly propose a novel tagging scheme that can convert the joint extraction task to a tagging problem.Then, based on our tagging scheme, we study different end-toend models to extract entities and their relations directly, without identifying entities and relations separately.We conduct experiments on a public dataset produced by distant supervision method and the experimental results show that the tagging based methods are better than most of the existing pipelined and joint learning methods.What's more, the end-to-end model proposed in this paper, achieves the best results on the public dataset.
Suncong Zheng, Feng Wang 0023, Hongyun Bao, Yuexing Hao, Peng Zhou 0009, Bo Xu 0002
ACL (1)6
2017 Towards Compact and Fast Neural Machine Translation Using a Combined Method
abstract
Neural Machine Translation (NMT) lays intensive burden on computation and memory cost.It is a challenge to deploy NMT models on the devices with limited computation and memory budgets.This paper presents a four stage pipeline to compress model and speed up the decoding for NMT.Our method first introduces a compact architecture based on convolutional encoder and weight shared embeddings.Then weight pruning is applied to obtain a sparse model.Next, we propose a fast sequence interpolation approach which enables the greedy decoding to achieve performance on par with the beam search.Hence, the time-consuming beam search can be replaced by simple greedy decoding.Finally, vocabulary selection is used to reduce the computation of softmax layer.Our final model achieves 10× speedup, 17× parameters reduction, <35MB storage size and comparable performance compared to the baseline model.
Wei Chen 0048, Feng Wang 0023, Bo Xu 0002
EMNLP5
2017 Combining unidirectional long short-term memory with convolutional output layer for high-performance speech synthesis
abstract
In this paper, we target improving the accuracy of acoustic modelling for statistical parametric speech synthesis (SPSS) and introduce the convolutional neural network (CNN) due to its powerful capacity in locality modelling. A novel model architecture combining unidirectional long short-term memory (LSTM) and a time-domain convolutional output layer (COL) is proposed and employed to acoustic modelling. The two components complement each other and result in a high-performance synthesis system. Specifically, the unidirectional LSTM can learn expressive feature representations from history context and the COL ingeniously absorbs some of these representations within a look-ahead window to advance predictions. This complementary mechanism significantly improve the predictive accuracy and the quality of synthetic speech. In addition, the unique operation mechanism of convolution makes COL a fine parameter trajectory smoother between consecutive frames. Subjective preference tests show that the proposed architecture can synthesize natural sounding speech without dynamic features.
Wenfu Wang, Bo Xu 0002
ICASSP2
2017 Measuring Word Semantic Similarity Based on Transferred Vectors
Changliang Li, Yujun Zhou 0001, Jian Cheng 0001, Bo Xu 0002
ICONIP (4)5
2017 End-to-End Chinese Image Text Recognition with Attention Model
Fenfen Sheng, Chuanlei Zhai, Zhineng Chen, Bo Xu 0002
ICONIP (3)4
2017 Towards a Brain-Inspired Developmental Neural Network by Adaptive Synaptic Pruning
Tielin Zhang, Yi Zeng 0001, Bo Xu 0002
ICONIP (4)4
2017 Word-Level Permutation and Improved Lower Frame Rate for RNN-Based Acoustic Modeling
Bo Xu 0002
ICONIP (6)4
2017 Hierarchical Hybrid Attention Networks for Chinese Conversation Topic Classification
Yujun Zhou 0001, Changliang Li, Bo Xu 0002, Jiaming Xu 0001, Bo Xu 0011
ICONIP (2)3
2017 Convolutional Neural Network with Word Embeddings for Chinese Word Segmentation
abstract
Character-based sequence labeling framework is flexible and efficient for Chinese word segmentation (CWS). Recently, many character-based neural models have been applied to CWS. While they obtain good performance, they have two obvious weaknesses. The first is that they heavily rely on manually designed bigram feature, i.e. they are not good at capturing n-gram features automatically. The second is that they make no use of full word information. For the first weakness, we propose a convolutional neural model, which is able to capture rich n-gram features without any feature engineering. For the second one, we propose an effective approach to integrate the proposed model with word embeddings. We evaluate the model on two benchmark datasets: PKU and MSR. Without any feature engineering, the model obtains competitive performance — 95.7% on PKU and 97.3% on MSR. Armed with word embeddings, the model achieves state-of-the-art performance on both datasets — 96.5% on PKU and 98.0% on MSR, without using any external labeled resource.
Chunqi Wang, Bo Xu 0002
IJCNLP(1)2
2017 A class-specific copy network for handling the rare word problem in neural machine translation
abstract
Neural machine translation (NMT) has shown promising results and rapidly gained adoption in many large-scale settings. With the NMT model being widely used in empirical productions, its long-standing weakness in handling the rare and out of vocabulary words has been amplified a lot. In order to release the model from the stress of “understanding” the rare words, copy mechanism has been proposed to deal with the rare and unseen words for the neural network models using attention. However the negative side of the copy mechanism is that the model is only able to decide whether to copy or not. It is unable to detect which class should the rare word be copied to, such as person, location, and organization. This paper deeply investigates this limitation of the NMT model. As a result, we propose a new NMT model by novelly incorporating a class-specific copy network. With the network, the proposed NMT model is able to decide which class the words in the target belong to and which class in the source should be copied to. Experimental results on Chinese-English translation tasks show that the proposed model outperforms the traditional NMT model with a large margin especially for sentences containing the rare words.
Feng Wang 0023, Wei Chen 0048, Zhen Yang 0007, Bo Xu 0002
IJCNN6
2017 Multi-sense based neural machine translation
abstract
Attention mechanism advances the neural machine translation (NMT) by reducing the confusion introduced by irrelevant words in long sentences. However, the confusion caused by ambiguous words hasn't been handled yet and it may be a bottleneck for the NMT model. This paper validates the hypothesis and proposes a simple and flexible framework, which enables the NMT model to only focus on the relevant sense type of the input word in current context. Experiments show that the proposed model achieves substantial improvements on every test set over competitive baselines. Our contributions come from twofold. Firstly, to the best of our knowledge, this is the first effort to introduce the multi-sense representation, which represents each sense type of the word with a sense-specific embedding, into NMT. Secondly, We propose a sense search module which can detect the sense type of the word automatically. Flexibility and versatility are the most attractive characteristic of the proposed sense search module. It can be applied to any other semantic related NLP tasks with little modification.
Zhen Yang 0007, Wei Chen 0048, Feng Wang 0023, Bo Xu 0002
IJCNN4
2017 Multilingual Recurrent Neural Networks with Residual Learning for Low-Resource Speech Recognition
abstract
The shared-hidden-layer multilingual deep neural network (SHL-MDNN), in which the hidden layers of feed-forward deep neural network (DNN) are shared across multiple languages while the softmax layers are language dependent, has been shown to be effective on acoustic modeling of multilingual low-resource speech recognition. In this paper, we propose that the shared-hidden-layer with Long Short-Term Memory (LSTM) recurrent neural networks can achieve further performance improvement considering LSTM has outperformed DNN as the acoustic model of automatic speech recognition (ASR). Moreover, we reveal that shared-hidden-layer multilingual LSTM (SHL-MLSTM) with residual learning can yield additional moderate but consistent gain from multilingual tasks given the fact that residual learning can allievate the degradation problem of deep LSTMs. Experimental results demonstrate that SHL-MLSTM can relatively reduce word error rate (WER) by 2.1-6.8\% over SHL-MDNN trained using six languages and 2.6-7.3\% over monolingual LSTM trained using the language specific data on CALLHOME datasets. Additional WER reduction, about relatively 2\% over SHL-MLSTM, can be obtained through residual learning on CALLHOME datasets, which demonstrates residual learning is useful for SHL-MLSTM on multilingual low-resource ASR.
Bo Xu 0002
INTERSPEECH4
2017 Constructing a Chinese Conversation Corpus for Sentiment Analysis
Yujun Zhou 0001, Changliang Li, Bo Xu 0002, Jiaming Xu 0001, Bo Xu 0011
NLPCC3
2017 Improving multi-layer spiking neural networks by incorporating brain-inspired rules
Yi Zeng 0001, Tielin Zhang, Bo Xu 0002
Sci. China Inf. Sci.3
2017 Joint entity and relation extraction based on a hybrid neural network
Suncong Zheng, Yuexing Hao, Dongyuan Lu, Hongyun Bao, Jiaming Xu 0001, Hongwei Hao, Bo Xu 0002
Neurocomputing7
2017 Self-Taught convolutional neural networks for short text clustering
Jiaming Xu 0001, Bo Xu 0002, Peng Wang 0079, Suncong Zheng, Guanhua Tian, Jun Zhao 0001
Neural Networks2
2017 Encoder-decoder recurrent network model for interactive character animation generation
Wujun Che, Bo Xu 0002
Vis. Comput.3
2016 Hierarchical Memory Networks for Answer Selection on Unknown Words
abstract
Recently, end-to-end memory networks have shown promising results on Question Answering task, which encode the past facts into an explicit memory and perform reasoning ability by making multiple computational steps on the memory. However, memory networks conduct the reasoning on sentence-level memory to output coarse semantic vectors and do not further take any attention mechanism to focus on words, which may lead to the model lose some detail information, especially when the answers are rare or unknown words. In this paper, we propose a novel Hierarchical Memory Networks, dubbed HMN. First, we encode the past facts into sentence-level memory and word-level memory respectively. Then, k-max pooling is exploited following reasoning module on the sentence-level memory to sample the k most relevant sentences to a question and feed these sentences into attention mechanism on the word-level memory to focus the words in the selected sentences. Finally, the prediction is jointly learned over the outputs of the sentence-level reasoning module and the word-level attention mechanism. The experimental results demonstrate that our approach successfully conducts answer selection on unknown words and achieves a better performance than memory networks.
Jiaming Xu 0001, Jing Shi 0003, Yiqun Yao, Suncong Zheng, Bo Xu 0002, Bo Xu 0011
COLING5
2016 A Character-Aware Encoder for Neural Machine Translation
abstract
This article proposes a novel character-aware neural machine translation (NMT) model that views the input sequences as sequences of characters rather than words. On the use of row convolution (Amodei et al., 2015), the encoder of the proposed model composes word-level information from the input sequences of characters automatically. Since our model doesn’t rely on the boundaries between each word (as the whitespace boundaries in English), it is also applied to languages without explicit word segmentations (like Chinese). Experimental results on Chinese-English translation tasks show that the proposed character-aware NMT model can achieve comparable translation performance with the traditional word based NMT models. Despite the target side is still word based, the proposed model is able to generate much less unknown words.
Zhen Yang 0007, Wei Chen 0048, Feng Wang 0023, Bo Xu 0002
COLING4
2016 Text Classification Improved by Integrating Bidirectional LSTM with Two-dimensional Max Pooling
abstract
Recurrent Neural Network (RNN) is one of the most popular architectures used in Natural Language Processsing (NLP) tasks because its recurrent structure is very suitable to process variable-length text. RNN can utilize distributed representations of words by first converting the tokens comprising each text into vectors, which form a matrix. And this matrix includes two dimensions: the time-step dimension and the feature vector dimension. Then most existing models usually utilize one-dimensional (1D) max pooling operation or attention-based operation only on the time-step dimension to obtain a fixed-length vector. However, the features on the feature vector dimension are not mutually independent, and simply applying 1D pooling operation over the time-step dimension independently may destroy the structure of the feature representation. On the other hand, applying two-dimensional (2D) pooling operation over the two dimensions may sample more meaningful features for sequence modeling tasks. To integrate the features on both dimensions of the matrix, this paper explores applying 2D max pooling operation to obtain a fixed-length representation of the text. This paper also utilizes 2D convolution to sample more meaningful information of the matrix. Experiments are conducted on six text classification tasks, including sentiment analysis, question classification, subjectivity classification and newsgroup classification. Compared with the state-of-the-art models, the proposed models achieve excellent performance on 4 out of 6 tasks. Specifically, one of the proposed models achieves highest accuracy on Stanford Sentiment Treebank binary classification and fine-grained classification tasks.
Peng Zhou 0009, Zhenyu Qi 0003, Suncong Zheng, Jiaming Xu 0001, Hongyun Bao, Bo Xu 0002
COLING6
2016 Gating recurrent mixture density networks for acoustic modeling in statistical parametric speech synthesis
abstract
Though recurrent neural networks (RNNs) using long short-term memory (LSTM) units can address the issue of long-span dependencies across the linguistic inputs and have achieved the state-of-the-art performance for statistical parametric speech synthesis (SPSS), another limitation of the intrinsic uni-Gaussian nature of mean square error (MSE) objective function still remains. This paper proposes a gating recurrent mixture density network (GRMDN) architecture to jointly address these two problems in neural network based SPSS. What's more, the gated recurrent unit (GRU), which is much simpler and has more intelligible work mechanism than LSTM, is also investigated as an alternative gating unit in RNN based acoustic modeling. Experimental results show that the proposed GRMDN architecture can synthesize more natural speech than its MSE-trained counterpart and both the two gating units (LSTM and GRU) show comparable performance.
Wenfu Wang, Bo Xu 0002
ICASSP3
2016 End-to-End Language Identification Using Attention-Based Recurrent Neural Networks
Wang Geng, Wenfu Wang, Xinyuan Cai, Bo Xu 0002
INTERSPEECH5
2016 Gating Recurrent Enhanced Memory Neural Networks on Language Identification
Wang Geng, Wenfu Wang, Xinyuan Cai, Bo Xu 0002
INTERSPEECH5
2016 First Step Towards End-to-End Parametric TTS Synthesis: Generating Spectral Parameters with Neural Attention
Wenfu Wang, Bo Xu 0002
INTERSPEECH3
2016 Multidimensional Residual Learning Based on Recurrent Neural Networks for Acoustic Modeling
Bo Xu 0002
INTERSPEECH3
2016 Joint Learning of Entity Semantics and Relation Pattern for Relation Extraction
Suncong Zheng, Jiaming Xu 0001, Hongyun Bao, Zhenyu Qi 0003, Hongwei Hao, Bo Xu 0002
ECML/PKDD (1)7
2016 SHTM: A neocortex-inspired algorithm for one-shot text generation
abstract
Text generation is a typical nature language processing task, and is the basis of machine translation and question answering. Deep learning techniques can get good performance on this task under the condition that huge number of parameters and mass of data are available for training. However, human beings do not learn in this way. People combine knowledge learned before and something new with only few samples. This process is called one-shot learning. In this paper, we propose a neocortex based computational model, Semantic Hierarchical Temporal Memory model (SHTM), for one-shot text generation. The model is refined from Hierarchical Temporal Memory model. LSTM is used for comparative study. Results on three public datasets show that SHTM performs much better than LSTM on the measures of mean precision and BLEU score. In addition, we utilize SHTM model to do question answering in the fashion of text generation and verifying its superiority.
Yi Zeng 0001, Bo Xu 0002
SMC3
2016 HMSNN: Hippocampus inspired Memory Spiking Neural Network
abstract
Human beings receive stimulations in primary sensory cortex and transfer them to higher brain regions automatically. What happened in this procedure? In this paper, we will focus on one of these regions (hippocampus) and try to simulate its working procedure by building an HMSNN (Hippocampus inspired Memory Spiking Neural Network) model. Dentate Gyrus (DG) and Cornu Ammonis area 3 (CA3) are the main regions of hippocampus and will be simulated by feed forward Spiking Neural Network (SNN) and recurrent Hopfield-like network respectively. From the structural perspective, the computational unit and the connectivity between neurons in HMSNN are all consistent with the anatomical-experimental results in hippocampus. From the functional perspective, the multi-scale memory formation, memory abstraction and memory retention will be shown in HMSNN model. In addition, the HMSNN is tested on MNIST handwritten digit dataset (with static images) and robot walking dataset (with dynamical images). The experimental result shows that: biological neural circuit inspired HMSNN shows comparable classification performance on both datasets compared to the state-of-art convolutional neural networks (CNNs), and shows significantly better performance compared to CNN when noises are introduced to the original images.
Tielin Zhang, Yi Zeng 0001, Dongcheng Zhao, Liwei Wang 0001, Yuxuan Zhao 0002, Bo Xu 0002
SMC6
2016 Compositional Recurrent Neural Networks for Chinese Short Text Classification
abstract
Word segmentation is the first step in Chinese natural language processing, and the error caused by word segmentation can be transmitted to the whole system. In order to reduce the impact of word segmentation and improve the overall performance of Chinese short text classification system, we propose a hybrid model of character-level and word-level features based on recurrent neural network (RNN) with long short-term memory (LSTM). By integrating character-level feature into word-level feature, the missing semantic information by the error of word segmentation will be constructed, meanwhile the wrong semantic relevance will be reduced. The final feature representation is that it suppressed the error of word segmentation in the case of maintaining most of the semantic features of the sentence. The whole model is finally trained end-to-end with supervised Chinese short text classification task. Results demonstrate that the proposed model in this paper is able to represent Chinese short text effectively, and the performances of 32-class and 5-class categorization outperform some remarkable methods.
Yujun Zhou 0001, Bo Xu 0011, Jiaming Xu 0001, Changliang Li, Bo Xu 0002
WI6
2016 Semantic expansion using word embedding clustering and convolutional neural network for improving short text classification
Peng Wang 0079, Bo Xu 0002, Jiaming Xu 0001, Guanhua Tian, Hongwei Hao
Neurocomputing2
2016 HCNN: A Neural Network Model for Combining Local and Global Features Towards Human-Like Classification
abstract
Brain-inspired algorithms such as convolutional neural network (CNN) have helped machine vision systems to achieve state-of-the-art performance for various tasks (e.g. image classification). However, CNNs mainly rely on local features (e.g. hierarchical features of points and angles from images), while important global structured features such as contour features are lost. Global understanding of natural objects is considered to be essential characteristics that the human visual system follows, and for developing human-like visual systems, the lost of consideration from this perspective may lead to inevitable failure on certain tasks. Experimental results have proved that well-trained CNN classifier cannot correctly distinguish fooling images (in which some local features from the natural images are chaotically distributed) from natural images. For example, a picture that is composed of yellow–black bars will be recognized as school bus with very high confidence by CNN. On the contrary, human visual system focuses on both the texture and contour features to form representation of images and would not mis-take them. In order to solve the upper problem, we propose a neural network model, named as histogram of oriented gradient (HOG) improved CNN (HCNN), that combines local and global features towards human-like classification based on CNN and HOG. The experimental results on MNIST datasets and part of ImageNet datasets show that HCNN outperforms traditional CNN for object classification with fooling images, which indicates the feasibility, accuracy and potential effectiveness of HCNN for solving image classification problem.
Tielin Zhang, Yi Zeng 0001, Bo Xu 0002
Int. J. Pattern Recognit. Artif. Intell.3
2016 A neural network framework for relation extraction: Learning entity semantic and relation pattern
Suncong Zheng, Jiaming Xu 0001, Peng Zhou 0009, Hongyun Bao, Zhenyu Qi 0003, Bo Xu 0002
Knowl. Based Syst.6
2015 Short Text Hashing Improved by Integrating Multi-granularity Topics and Tags
Jiaming Xu 0001, Bo Xu 0002, Guanhua Tian, Jun Zhao 0001, Fangyuan Wang 0003, Hongwei Hao
CICLing (1)2
2015 Semi-supervised Chinese Word Segmentation based on Bilingual Information
abstract
This paper presents a bilingual semisupervised Chinese word segmentation (CWS) method that leverages the natural segmenting information of English sentences.The proposed method involves learning three levels of features, namely, character-level, phrase-level and sentence-level, provided by multiple submodels.We use a sub-model of conditional random fields (CRF) to learn monolingual grammars, a sub-model based on character-based alignment to obtain explicit segmenting knowledge, and another sub-model based on transliteration similarity to detect out-of-vocabulary (OOV) words.Moreover, we propose a sub-model leveraging neural network to ensure the proper treatment of the semantic gap and a phrase-based translation sub-model to score the translation probability of the Chinese segmentation and its corresponding English sentences.A cascaded log-linear model is employed to combine these features to segment bilingual unlabeled data, the results of which are used to justify the original supervised CWS model.The evaluation shows that our method results in superior results compared with those of the state-of-the-art monolingual and bilingual semi-supervised models that have been reported in the literature.
Wei Chen 0048, Bo Xu 0002
EMNLP2
2015 Convolutional Neural Networks for Text Hashing
Jiaming Xu 0001, Peng Wang 0079, Guanhua Tian, Bo Xu 0002, Jun Zhao 0001, Fangyuan Wang 0003, Hongwei Hao
IJCAI4
2015 Multilingual tandem bottleneck feature for language identification
abstract
The deep bottleneck (BN) feature based ivector solution has been recognized as a popular pipeline for language identification (LID) recently. However, issues such as how to extract more effective BN features and how to fully utilize features extracted from deep neural networks (DNN) are still not well investigated. In this paper, these issues are empirically tackled by means as follows: First, two novel types of deep features, phone-discriminant and triphone-discriminate are extracted. Then, DNNs are trained both separately and jointly on multilingual corpuses to produce different BN features. Finally, tandem fashion on deep BN features is applied to build enhanced deep features. Experiment results show that systems built on top of tandem deep features obtain 19% and 42% relative equal error rate reduction on average on NIST LRE 2007 over the counterpart built on traditional deep BN features and the cepstral feature based LID system, respectively
Wang Geng, Jie Li 0032, Xinyuan Cai, Bo Xu 0002
INTERSPEECH5
2015 Multi-task learning deep neural networks for speech feature denoising
abstract
Traditional automatic speech recognition (ASR) systems usually get a sharp performance drop when noise presents in speech. To make a robust ASR, we introduce a new model using the multi-task learning deep neural networks (MTL-DNN) to solve the speech denoising task in feature level. In this model, the networks are initialized by pre-training restricted Boltzmann machines (RBM) and fine-tuned by jointly learning multiple interactive tasks using a shared representation. In multi-task learning, we choose a noisy-clean speech pair fitting task as the primary task and separately explore two constraints as the secondary tasks: phone label and phone cluster. In experiments, the denoised speech is reconstructed by the MTL-DNN using the noisy speech as input and it is respectively evaluated by the DNN-hidden Markov model (HMM) based and the Gaussian Mixture Model (GMM)-HMM based ASR systems. Results show that, using the denoised speech, the word error rate (WER) is respectively reduced by 53.14% and 34.84% compared with baselines. The MTL-DNN model also outperforms the general single-task learning deep neural networks (STL-DNN) model with a performance improvement of 4.93% and 3.88% respectively.
Dengfeng Ke, Hao Zheng 0009, Bo Xu 0002, Yanyan Xu 0001, Kaile Su
INTERSPEECH4
2015 Towards end-to-end speech recognition for Chinese Mandarin using long short-term memory recurrent neural networks
abstract
End-to-end speech recognition systems have been successfully designed for English. Taking into account the distinctive characteristics between Chinese Mandarin and English, it is worthy to do some additional work to transfer these approaches to Chinese. In this paper, we attempt to build a Chinese speech recognition system using end-to-end learning method. The system is based on a combination of deep Long Short-Term Memory Projected (LSTMP) network architecture and the Connectionist Temporal Classification objective function (CTC). The Chinese characters (the number is about 6,000) are used as the output labels directly. To integrate language model information during decoding, the CTC Beam Search method is adopted and optimized to make it more effective and more efficient. We present the first-pass decoding results which are obtained by decoding from scratch using CTC-trained network and language model. Although these results are not as good as the performance of DNN-HMMs hybrid system, they indicate that it is feasible to choose Chinese characters as the output alphabet in the end-toend speech recognition system.
Jie Li 0032, Heng Zhang 0028, Xinyuan Cai, Bo Xu 0002
INTERSPEECH4
2015 Modeling emotion entrainment of online users in emergency events
abstract
Emotion entrainment accounts for the rhythmic convergence of human emotions through social interactions. This phenomenon abounds in various disciplines, i.e. effervescency in soccer games, anger proliferation in violence incidents, or anxiety diffusion in disasters. Although emotion entrainment is highly relevant to the quality of human daily life, the principles underpinning this phenomenon is still unclear. Previous dynamic models try to explain entrainment phenomenon by assuming symmetrical coupling among identical individuals. Yet this assumption clearly does not hold in real-world human interactions. As such, we propose an alternative model that captures asymmetric relationships. In depicting the coupling mechanism, the effect of social influence is also encoded. Experimental results on two emergent social events suggest that the proposed model characterizes emotion trends with high accuracy. Also, we explain the emotion dynamics by analyzing the reconstructed entrainment matrix. Our work may present practical implications for those who want to guide or regulate the emotion evolution in emergency events discussed online.
Saike He, Xiaolong Zheng 0001, Daniel Dajun Zeng, Bo Xu 0002, Changliang Li, Guanhua Tian, Lei Wang 0062, Hongwei Hao
ISI4
2015 Bilingually-Constrained Recursive Neural Networks with Syntactic Constraints for Hierarchical Translation Model
abstract
Hierarchical phrase-based translation models have advanced statistical machine translation (SMT). Because such models can improve leveraging of syntactic information, two types of methods (leveraging source parsing and leveraging shallow parsing) are applied to introduce syntactic constraints into translation models. In this paper, we propose a bilingually-constrained recursive neural network (BC-RNN) model to combine the merits of these two types of methods. First we perform supervised learning on a manually parsed corpus using the standard recursive neural network (RNN) model. Then we employ unsupervised bilingually-constrained tuning to improve the accuracy of the standard RNN model. Leveraging the BC-RNN model, we introduce both source parsing and shallow parsing information into a hierarchical phrase-based translation model. The evaluation demonstrates that our proposed method outperforms other state-of-the-art statistical machine translation methods for National Institute of Standards and Technology 2008 (NIST 2008) Chinese-English machine translation testing data.
Wei Chen 0048, Bo Xu 0002
NLPCC2
2015 Parallel Recursive Deep Model for Sentiment Analysis
Changliang Li, Bo Xu 0002, Saike He, Guanhua Tian, Yujun Zhou 0001
PAKDD (2)2
2014 Learning New Semi-Supervised Deep Auto-encoder Features for Statistical Machine Translation
abstract
In this paper, instead of designing new features based on intuition, linguistic knowledge and domain, we learn some new and effective features using the deep autoencoder (DAE) paradigm for phrase-based translation model.Using the unsupervised pre-trained deep belief net (DBN) to initialize DAE's parameters and using the input original phrase features as a teacher for semi-supervised fine-tuning, we learn new semi-supervised DAE features, which are more effective and stable than the unsupervised DBN features.Moreover, to learn high dimensional feature representation, we introduce a natural horizontal composition of more DAEs for large hidden layers feature learning.On two Chinese-English tasks, our semi-supervised DAE features obtain statistically significant improvements of 1.34/2.45(IWSLT) and 0.82/1.52(NIST) BLEU points over the unsupervised DBN features and the baseline features, respectively.
Shi-xiang Lu, Zhenbiao Chen, Bo Xu 0002
ACL (1)3
2014 Characterizing emotion entrainment in social media
abstract
The sociological theory of entrainment accounts for the synchronization of human rhythmic modalities through social interactions: they coordinate in a variety of dimensions including linguistic styles, facial expressions, music pace, applause, and so on. Though highly relevant, emotion entrainment has received little attention to date. In addition, most previous studies on entrainment are done through small scale or controlled laboratory studies. In this paper, we investigate emotion entrainment in the context of online social media. To the best of our knowledge, this is the first time that emotion entrainment has been examined on a large scale, real world setting. For this purpose, we propose a framework that can model entrainment phenomenon and measure its effect. Our framework differentiates from previous research by its model-free essential and discerning in entrainment directions. These traits enable us to model entrainment dynamics under few assumptions, and distinguish emotion flow of entrainment. In our studies, we investigate entrainment patterns under different emotion states, i.e. positive, neutral and negative. We discover that entrainments under different emotions all follow a power law distribution. Besides, people are willing to entrain to others under positive emotion, and users with positive emotion are more likely to be entrained. By inspecting the interactions between entrainment and emotion, we reveal that entrainment has an effect of negotiating different emotion types toward an even distribution.
Saike He, Xiaolong Zheng 0001, Xiuguo Bao, Hongyuan Ma, Daniel Dajun Zeng, Bo Xu 0002, Changliang Li, Hongwei Hao
ASONAM6
2014 Obtaining Better Word Representations via Language Transfer
Changliang Li, Bo Xu 0002, Xiuying Wang 0002, Wendong Ge
CICLing (1)2
2014 Chinese Image Character Recognition Using DNN and Machine Simulated Training Samples
Jinfeng Bai, Zhineng Chen, Bailan Feng, Bo Xu 0002
ICANN4
2014 Chinese Image Text Recognition on grayscale pixels
abstract
This paper presents a novel scheme for Chinese text recognition in images and videos. It's different from traditional paradigms that binarize text images, fed the binarized text to an OCR engine and get the recognized results. The proposed scheme, named grayscale based Chinese Image Text Recognition (gCITR), implements the recognition directly on grayscale pixels via the following steps: image text over-segmentation, building recognition graph, Chinese character recognition and beam search determination. The advantages of gCITR lie in: (1) it does not heavily rely on the performance of binarization, which is not robust in practical and thus severely affects the performance of OCR, (2) grayscale image retains more information of the text thus facilitates the recognition. Experimental results on text from 13 TV news videos demonstrate the effectiveness of the proposed gCITR, from which significant performance gains are observed.
Jinfeng Bai, Zhineng Chen, Bailan Feng, Bo Xu 0002
ICASSP4
2014 Recursive neural network based word topology model for hierarchical phrase-based speech translation
abstract
Recursive word topology structure is commonly found in natural language sentences, and discovering this structure can help us to not only identify the units that a sentence contains but also how they interact to form a whole. In this paper, we explore a novel recursive neural network (RNN) based word topology model (WordTM) for hierarchical phrase-based (HPB) speech translation, which captures the topological structure of the words on the source side in a syntactically and semantically meaningful order. Experiments show that our WordTM significantly outperforms the state-of-the-art soft syntactic constraints.
Shi-xiang Lu, Wei Wei 0036, Xiaoyin Fu, Bo Xu 0002
ICASSP4
2014 An investigation of summed-channel speaker recognition with multi-session enrollment
abstract
This paper describes a general framework of speaker recognition on summed-channel condition for both enrolling and test data. We present several methods for clustering the target speaker who is involved in multiple summed-channel enrolling excerpts. In our approach, each excerpt is segmented separately by a speaker diarization system as the first stage. Then segments belonging to the same speaker are clustered to train the target speaker model, and speaker verification is applied finally. We propose several effective objective functions to measure the purity of clustered segments in multi-session enrollment. Different confidence measures for summed-channel scoring are also presented. We report experimental results on female part in the NIST 2008 speaker recognition evaluation data, which show that our approach applied on summed-channel condition loses only 1% of the performance measured by equal error rates (EER) compared to the two-channel condition.
Rong Zheng 0005, Bo Xu 0002
ICASSP3
2014 Variational Bayes based I-vector for speaker diarization of telephone conversations
abstract
In this paper, we investigate the variational Bayes based I-vector method for speaker diarization of telephone conversations. The motivation of the proposed algorithm is to utilize variational Bayesian framework and exploit potential channel effect of total variability modeling for diarization of conversation side. Other three well-known techniques are compared as follows: K-means clustering for eigenvoices and I-vector speaker diarization, and variational Bayes applied to eigenvoices. Performance evaluations are conducted on the summed-channel telephone data from the 2008 NIST speaker recognition evaluation. The paper discusses how the performance is influenced by different modules, e.g., VAD, initial speaker clustering and Viterbi re-segmentation. Comparison experiments show the interest of variational Bayesian probabilistic framework for speaker diarization.
Rong Zheng 0005, Ce Zhang 0006, Bo Xu 0002
ICASSP4
2014 Image character recognition using deep convolutional neural network learned from different languages
abstract
This paper proposes a shared-hidden-layer deep convolutional neural network (SHL-CNN) for image character recognition. In SHL-CNN, the hidden layers are made common across characters from different languages, performing a universal feature extraction process that aims at learning common character traits existed in different languages such as strokes, while the final softmax layer is made language dependent, trained based on characters from the destination language only. This paper is the first attempt to introduce the SHL-CNN framework to image character recognition. Under the SHL-CNN framework, we discuss several issues including architecture of the network, training of the network, from which a suitable SHL-CNN model for image character recognition is empirically learned. The effectiveness of the learned SHL-CNN is verified on both English and Chinese image character recognition tasks, showing that the SHL-CNN can reduce recognition errors by 16–30% relatively compared with models trained by characters of only one language using conventional CNN, and by 35.7% relatively compared with state-of-the-art methods. In addition, the shared hidden layers learned are also useful for unseen image character recognition tasks.
Jinfeng Bai, Zhineng Chen, Bailan Feng, Bo Xu 0002
ICIP4
2014 Short Text Hashing Improved by Integrating Topic Features and Tags
Jiaming Xu 0001, Bo Xu 0002, Jun Zhao 0001, Guanhua Tian, Heng Zhang 0028, Hongwei Hao
ICONIP (2)2
2014 A robust framework for short text categorization based on topic model and integrated classifier
abstract
In this paper, we propose a method for short text categorization using topic model and integrated classifier. To enrich the representation of short text, the Latent Dirichlet Allocation (LDA) model is used to extract latent topic information. While for classification, we combine two classifiers for achieving high reliability. Particularly, we train LDA models with variable number of topics using the Wikipedia corpus as external knowledge base, and extend labeled Web snippets by potential topics extracted by LDA. Then, the enriched representation of snippets are used to learn Maximum Entropy (MaxEnt) and support vector machine (SVM) classifiers separately. Finally, viewing that the most possible predicted result will appear in the top two candidates selected by MaxEnt classifier, we develop a novel scheme that if the gap between these candidates is large enough, the predicted result is considered to be reliable; otherwise, the SVM classifier will be integrated with MaxEnt classifier to make a comprehensive prediction. Experimental results show that our framework is effective and can outperform the state-of-the-art techniques.
Peng Wang 0079, Heng Zhang 0028, Yu-Fang Wu, Bo Xu 0002, Hongwei Hao
IJCNN4
2014 An empirical study of multilingual and low-resource spoken term detection using deep neural networks
Jie Li 0032, Bo Xu 0002
INTERSPEECH3
2014 Investigation of cross-lingual bottleneck features in hybrid ASR systems
Jie Li 0032, Rong Zheng 0005, Bo Xu 0002
INTERSPEECH3
2014 Improving wideband acoustic models using mixed-bandwidth training data via DNN adaptation
Zhao You, Bo Xu 0002
INTERSPEECH2
2014 CeleLabel: an interactive system for annotating celebrities in web videos
abstract
Manual annotation of celebrities in Web videos is an essential task in many people-related Web services. The task, however, poses a significant challenge even to skillful annotators, mainly due to the large quantity of unfamiliar and greatly varied celebrities, and the lack of a customized system for it. This work develops CeleLabel, an interactive system for manually annotating celebrities in the Web video domain. The peculiarity of CeleLabel is to exploit and display multiple types of information that could assist the annotation, including video content, context surrounding and within a video, celebrity images on the Web, and human factors. Using the system, annotators can interactively switch between two views, i.e., merging similar faces and labeling faces with names, to approach the annotation. User studies show that the CeleLabel leads to a much better labeling efficiency and satisfaction.
Zhineng Chen, Jinfeng Bai, Chong-Wah Ngo, Bailan Feng, Bo Xu 0002
ACM Multimedia5
2014 Spatial Similarity Measure of Visual Phrases for Image Retrieval
Jiansong Chen, Bailan Feng, Bo Xu 0002
MMM (2)3
2014 Video to Article Hyperlinking by Multiple Tag Property Exploration
Zhineng Chen, Bailan Feng, Hongtao Xie 0001, Rong Zheng 0005, Bo Xu 0002
MMM (1)5
2014 A Hybrid Method for Chinese Entity Relation Extraction
Zhenyu Qi 0003, Hongwei Hao, Bo Xu 0002
NLPCC4
2014 Short Text Feature Enrichment Using Link Analysis on Topic-Keyword Graph
Peng Wang 0079, Heng Zhang 0028, Bo Xu 0002, Hongwei Hao
NLPCC3
2014 Multiple style exploration for story unit segmentation of broadcast news video
Bailan Feng, Zhineng Chen, Rong Zheng 0005, Bo Xu 0002
Multim. Syst.4
2013 Understanding the dropout strategy and analyzing its effectiveness on LVCSR
abstract
The work by Hinton et al shows that the dropout strategy can greatly improve the performance of neural networks as well as reducing the influence of over-fitting. Nevertheless, there is still not a more detailed study on this strategy. In addition, the effectiveness of dropout on the task of LVCSR has not been analyzed. In this paper, we attempt to make a further discussion on the dropout strategy. The impacts on performance of different dropout probabilities for phone recognition task are experimented on TIMIT. To get an in-depth understanding of dropout, experiments of dropout testing are designed from the perspective of model averaging. The effectiveness of dropout is analyzed on a LVCSR task. Results show that the method of dropout fine-tuning combined with standard back-propagation gives significant performance improvements.
Jie Li 0032, Bo Xu 0002
ICASSP3
2013 Multi-modal topic unit segmentation in videos using conditional random fields
abstract
In this paper a novel approach of video segmentation into topic units is presented. This approach is built upon the design in which topic unit segmentation is transformed into label identification problem by defining four types of shots that reveal semantic structure of it. To implement our algorithm, four middle-level features including shot difference signal, scene transition graph, shot theme and audio type are extracted to depict the label properties of each shot, and then CRFs model is employed to identify the labels sequence. CRFs model integrates context information, so it produces accurate results in topic unit segmentation. The proposed approach is verified by two types of data: documentary and news. Experiments on testing data set yield average 86% F-measure, which illustrates that the proposed method can accurately detect most topic units in different genres of programs.
Bailan Feng, Bo Xu 0002
ICASSP3
2013 Asynchronous stochastic gradient descent for DNN training
abstract
It is well known that state-of-the-art speech recognition systems using deep neural network (DNN) can greatly improve the system performance compared with conventional GMM-HMM. However, what we have to pay correspondingly is the immense training cost due to the enormous parameters of DNN. Unfortunately, it is difficult to achieve parallelization of the minibatch-based back-propagation (BP) algorithm used in DNN training because of the frequent model updates. In this paper we describe an effective approach to achieve an approximation of BP - asynchronous stochastic gradient descent (ASGD), which is used to parallelize computing on multi-GPU. This approach manages multiple GPUs to work asynchronously to calculate gradients and update the global model parameters. Experimental results show that it achieves a 3.2 times speed-up on 4 GPUs than the single one, without any recognition performance loss.
Ce Zhang 0006, Zhao You, Rong Zheng 0005, Bo Xu 0002
ICASSP5
2013 Joint and Coupled Bilingual Topic Model Based Sentence Representations for Language Model Adaptation
Shi-xiang Lu, Xiaoyin Fu, Wei Wei 0036, Xingyuan Peng, Bo Xu 0002
IJCAI5
2013 Phrase-based Parallel Fragments Extraction from Comparable Corpora
Xiaoyin Fu, Wei Wei 0036, Shi-xiang Lu, Zhenbiao Chen, Bo Xu 0002
IJCNLP5
2013 A general Framework of video segmentation to logical unit based on conditional random fields
abstract
Segmenting video into logical units like scenes in movies and topic units in News videos is an essential prerequisite for a wide range of video related applications. In this paper, a novel approach for logical unit segmentation based on conditional random fields (CRFs) is presented. In comparison with previous approaches that handle scenes and topic units separately, the proposed approach deals with them in a general framework. Specifically, four types of shots are defined and represented by four middle-level features, i.e., shot difference, scene transition, shot theme and audio type. Then, the problem of logical unit segmentation is novelly formulated as a problem of identifying the type of shot based on the extracted features, by leveraging the CRFs model. The proposed framework effectively integrate visual, audio and contextual features, and it is able to produce ideal result for both scene and topic unit segmentation. The effectiveness of the proposed approach is verified on seven mainstream types of videos, from which average F-measures of 88% and 86% on scenes and topic units are reported respectively, illustrating that the proposed method can accurately segment logical units in different genres of videos.
Bailan Feng, Zhineng Chen, Bo Xu 0002
ICMR4
2013 Temporal Video Segmentation to Scene Based on Conditional Random Fileds
Bailan Feng, Bo Xu 0002
MMM (2)3
2013 Fusion of Audio-Visual Features and Statistical Property for Commercial Segmentation
Bailan Feng, Bo Xu 0002
MMM (1)3
2013 Simulated Spoken Dialogue System Based on IOHMM with User History
Changliang Li, Bo Xu 0002, Xiuying Wang 0002, Wendong Ge, Hongwei Hao
NLPCC2
2013 Pseudo In-Domain Data Selection from Large-Scale Web Corpus for Spoken Language Translation
Shi-xiang Lu, Xingyuan Peng, Zhenbiao Chen, Bo Xu 0002
NLPCC4
2013 A Fast Matching Method Based on Semantic Similarity for Short Texts
Jiaming Xu 0001, Pengcheng Liu 0001, Zhengya Sun, Bo Xu 0002, Hongwei Hao
NLPCC5
2013 Entity Conceptualization and Understanding Based on Web-Scale Knowledge Bases
abstract
This paper investigates on entity conceptualization and its applications based on a Web-scale knowledge base. Firstly, inspired by the spreading activation theory from Cognitive Psychology, a model for entity conceptualization in a context-free setting is introduced. Secondly, an extended model with additional consideration on environmental factors is proposed for context-sensitive entity conceptualization. With entity conceptualization, we propose a method for generating concept-level knowledge based on conceptualization of large-scale factual triple form knowledge acquired from the Wiki-Encyclopedias. We provide several typicality measures for understanding concept-level knowledge in different perspectives. Finally, we discuss how the concept-level knowledge and its typicality measures can be used for finding incorrect knowledge contributed by Wiki-Encyclopedia contents contributors.
Yi Zeng 0001, Hongwei Hao, Bo Xu 0002
SMC3
2013 A-STAR: Toward translating Asian spoken languages
Sakriani Sakti, Michael Paul, Andrew M. Finch, Shinsuke Sakai, Thang Tat Vu, Noriyuki Kimura, Chiori Hori, Eiichiro Sumita, Satoshi Nakamura 0001, Jun Park, Chai Wutiwiwatchai, Bo Xu 0002, Hammam Riza, Karunesh Arora, Haizhou Li 0001
Comput. Speech Lang.12
2012 Automated Essay Scoring Based on Finite State Transducer: towards ASR Transcription of Oral English Speech
Xingyuan Peng, Dengfeng Ke, Bo Xu 0002
ACL (1)3
2012 Translation Model Based Cross-Lingual Language Model Adaptation: from Word Models to Phrase Models
Shi-xiang Lu, Wei Wei 0036, Xiaoyin Fu, Bo Xu 0002
EMNLP-CoNLL4
2012 Multi-modal information fusion for news story segmentation in broadcast video
abstract
With the fast development of high-speed network and digital video recording technologies, broadcast video has been playing a more and more important role in our daily life. In this paper, we propose a novel news story segmentation scheme which can segment broadcast video into story units with multi-modal information fusion (MMIF) strategy. Compared with traditional methods, the proposed scheme extracts a wealth of semantic-level features including anchor person, topic caption, face, silence, acoustic change, audio keywords and textual content. Parallel to this, we make use of a multi-modal information fusion strategy for news story boundary characterization by joining these visual, audio and textual cues. Encouraging experimental results on News Vision dataset demonstrate the effectiveness of the proposed scheme.
Bailan Feng, Peng Ding 0003, Jiansong Chen, Jinfeng Bai, Bo Xu 0002
ICASSP6
2012 Unsupervised training of subspace gaussian mixture models for conversational telephone speech recognition
abstract
This paper presents our preliminary works on exploring unsupervised training of subspace gaussian mixture models for under-resourced CTS recognition task. The subspace model yields better performance than conventional GMM model, particularly in small or middle-sized training set. As an effective way to save human efforts, unsupervised learning is often applied to automatically transcribe a large amount of speech archives. The additional auto-transcribed data may help to improve model accuracy. In this paper, experiments are carried out on two publicly available English conversational telephone speech corpora. Both GMM and SGMM model in combination with unsupervised learning are examined and compared in this paper.
Zejun Ma 0001, Bo Xu 0002
ICASSP3
2012 Graph-based multi-modal scene detection for movie and teleplay
abstract
Automatic scene detection is a fundamental step for efficient video searching and browsing. This paper presents our current work on scene detection that integrates three effective strategies into a single framework. For each video, firstly, a coherence signal is constructed by graph modal obtained from the similarity matrix in a temporal interval. Secondly, the signal is optimized by scene transition graph (STG) analysis and audio classification, in which scene clues hidden in multimedia are discovered from the video. Finally, the scene boundaries are identified by window function. In experiments, we compare the proposed scene detection method with three typical algorithms on teleplay and movies, and the results of our method, yielding an average 0.85 F-measure, is the best one.
Bailan Feng, Peng Ding 0003, Bo Xu 0002
ICASSP4
2012 Discriminative training of weighted polynomial vector for acoustic language recognition
abstract
In this paper, we propose a discriminative method for the acoustic feature based language recognizer, which is a modification of the polynomial expansion in generalized linear discriminant sequence (GLDS) kernel. It is inspired by the Gaussian mixture model-support vector machine (GMM-SVM) system which has been successfully used in both speaker and language recognition. Because of the restriction of calculations in our method, it is nearly impossible to stack component dependent polynomial expansion vectors as GM-MSVM system does. Thus we introduce a set of language dependent weights to fuse these expansion vectors and utilize maximum mutual information (MMI) criterion and logistic regression to estimate the model parameters. Finally, we evaluate our method on the close-set, 30 seconds test condition of NIST LRE 2007 and up to 30% relative improvement can be achieved comparing to the baseline GLDS system.
Ce Zhang 0006, Rong Zheng 0005, Bo Xu 0002
ICASSP3
2012 Effective near-duplicate image retrieval with image-specific visual phrase selection
abstract
Near-duplicate image retrieval (NDIR) is an important topic for many applications such as multimedia content management, copyright infringement identification et al. In this work we propose a novel NDIR framework based on visual phrase. Compared with previous researches, this paper first introduces a spatial visual phrase (SVP) model enabling to capture relative geometry information between visual words. Then, it proposes an image-specific strategy to select descriptive SVPs. The strategy can not only handle the phrase sparseness problem which occurs in traditional selection strategy but also allow to select visual phrases according to the characteristic of each image. Experiments are carried out over Ukbench dataset and TRECVID dataset respectively, and encouraging experimental results demonstrate that both the SVP model and the selection strategy significantly improve the overall performance.
Jiansong Chen, Bailan Feng, Peng Ding 0003, Bo Xu 0002
ICIP5
2012 Statistical and Structural Analysis of Web-Based Collaborative Knowledge Bases Generated from Wiki Encyclopedia
abstract
Web-based collaborative knowledge bases collect human knowledge through the Web. They can be used for answering questions or support different knowledge intensive applications on the Web. From a statistical point of view, they usually reveal some interesting characteristics, which can be acquired through statistical analysis to get deeper understanding of these kinds of knowledge bases. In this paper, we build a semantic knowledge base using the triples extracted from a Chinese wiki Web site called Baidu Baike. We make an investigation on the statistical results on the building process and structural characteristics of this knowledge base. We explain what we have observed and inferred and how the conclusion can help to understand the process of building large scale Web based collaborative knowledge bases and how to make them better.
Yi Zeng 0001, Hongwei Hao, Bo Xu 0002
Web Intelligence4
2012 From English pitch accent detection to Mandarin stress detection, where is the difference?
Chongjia Ni, Bo Xu 0002
Comput. Speech Lang.3
2012 Automatic Prosodic Break Detection and Feature Analysis
Chongjia Ni, Aiying Zhang, Bo Xu 0002
J. Comput. Sci. Technol.4
2011 Structured precision modelling with Cholesky Basis Superposition for speech recognition
abstract
Structured precision modelling is an important approach to improve the intra-frame correlation modelling of the standard HMM, where Gaussian mixture model with diagonal covariance are used. Previous work has all been focused on direct structured representation of the precision matrices. In this paper, a new framework is proposed, where the structure of the Cholesky square root of the precision matrix is investigated, referred to as Cholesky Basis Superposition (CBS). Each Cholesky matrix associated with a particular Gaussian distribution is represented as a linear combination of a set of Gaussian independent basis upper-triangular matrices. Efficient optimization methods are derived for both combination weights and basis matrices. Experiments on a Chinese dictation task showed that the proposed approach can significantly outperformed the direct structured precision modelling with similar number of parameters as well as full covariance modelling.
Bo Xu 0002
ICASSP3
2011 Exploring nuisance attribute projection and score normalization for GLDS-SVM based automatic mispronunciation detection method
abstract
In the task of mispronunciation detection, the cross-speaker degradation and some other confusing nuisances are the challenging problems demanding prompt solution. In this paper, we will attempt to remove the non-pronunciation variations in the GLDS-SVM expansion space by using nuisance attribute projection strategy, in order to increase the separating capacity between different phoneme instances. Moreover, different kinds of score normalization methods with softmax, posterior probability vector (PPV), Z-norm and T-norm are comparatively discussed. The experiments on three kinds of speech corpora demonstrate the effectiveness of the above methods, and the performance improvement is not very significant, but sustainable.
Hongyan Li 0010, Shen Huang, Shijin Wang 0001, Jiaen Liang, Bo Xu 0002
ICASSP5
2011 Exploring implicit score normalization techniques in speaker verification
abstract
In this paper we first introduce four kinds of modification of Symmetric Scoring which produce likelihood ratios that do not need to be explicitly normalized, i.e. T-norm, Z-norm. To solve the numerical problem caused by large covariance matrix calculation, we propose three solutions and present the result for each of them according to different modifications. Then we introduce a new kernel function that contains the effect of score normalization for SVM-based Speaker Verification system. We also show that these methods consistently improve the performance of the original system by means of implicit score normalization. In order to achieve more efficient computation, we evaluate an attempt to explore implicit score normalization in the much lower dimensional speaker factor space. We evaluate the performance of the proposed algorithms on the core condition of the NIST SRE 2006 dataset.
Ce Zhang 0006, Rong Zheng 0005, Bo Xu 0002
ICASSP3
2011 Commercial detection by mining maximal repeated sequence in audio stream
abstract
Efficient detection of commercial is an important topic for many applications such as commercial monitoring, market investigation. This paper reports an unsupervised technique of discovering commercial by mining repeated sequence in audio stream. Compared with previous work, we focus on solving practical problems by introducing three principles of commercial: repetition principle, independence principle and equivalence principle. Based on these principles, we detect the commercials by first mining maximal repeated sequences (MRS) and then post-processing the MRS pairs based on independence principle and equivalence principle for final result. In addition, a coarse-to-fine scheme is adopted in the acoustic matching stage to save computational cost. Extensive experiments both on simulated data and real broadcast data demonstrate the effectiveness of our method.
Jiansong Chen, Peng Ding 0003, Bo Xu 0002
ICME5
2011 Prosody dependent Mandarin speech recognition
abstract
In this paper, we discuss how to model and train Mandarin prosody dependent acoustic model based on automatic prosody annotation corpus. Based on prosody annotation corpus, we first utilize our proposed methods to train prosody dependent and prosody independent tonal syllable model, and then use these models to get the mixed acoustic models. In this paper, we also utilize tone model to improve the correct rate of tonal syllable through revising the tone of the tonal syllable at certain significant level. When compared with the baseline system, the performance of our proposed mixed speech recognition system improves the correct rate of tonal syllable significantly.
Chongjia Ni, Bo Xu 0002
IJCNN3
2011 A Robust Approach to Mining Repeated Sequence in Audio Stream
Jiansong Chen, Bailan Feng, Peng Ding 0003, Bo Xu 0002
INTERSPEECH5
2011 Context-Dependent Duration Modeling with Backoff Strategy and Look-Up Tables for Pronunciation Assessment and Mispronunciation Detection
Hongyan Li 0010, Shen Huang, Shijin Wang 0001, Bo Xu 0002
INTERSPEECH4
2011 An Empirical Study of Multilingual Spoken Term Detection
Zejun Ma 0001, Bo Xu 0002
INTERSPEECH3
2011 Fusing Multiple Confidence Measures for Chinese Spoken Term Detection
Zejun Ma 0001, Bo Xu 0002
INTERSPEECH3
2011 Automatic Prosodic Events Detection by Using Syllable-Based Acoustic, Lexical and Syntactic Features
abstract
Automatic prosodic events detection and annotation are important for both speech understanding and natural speech synthesis. In this paper, the complementary model method is proposed to detect prosodic events. This method discards the independent assumption between the acoustic features and the lexical and syntactic features, models not only the features of the current syllable but also the contextual features of the current syllable at the model level, and realizes the complementarities by taking the advantages of each model. The experiments on Boston University Radio News Corpus show that the complementary model can yield 91.40% pitch accent detection accuracy rate, 95.19% intonational phrase boundaries (IPB) detection accuracy rate and 93.96% break index detection accuracy rate. When compared with the previous work, the results for pitch accent, IPB and break index detection are significantly better. Index Terms: complementary model, boosting classification and regression tree (CART), conditional random fields (CRFs)
Chongjia Ni, Bo Xu 0002
INTERSPEECH3
2011 Restoring the Residual Speaker Information in Total Variability Modeling for Speaker Verification
Ce Zhang 0006, Rong Zheng 0005, Bo Xu 0002
INTERSPEECH3
2011 Data-Driven Gaussian Component Selection for Fast GMM-Based Speaker Verification
Ce Zhang 0006, Rong Zheng 0005, Bo Xu 0002
INTERSPEECH3
2011 Data-Driven UBM Generation via Tied Gaussians for GMM-Supervector Based Accent Identification
Rong Zheng 0005, Ce Zhang 0006, Bo Xu 0002
INTERSPEECH3
2011 Ridge extraction of a smooth 2-manifold surface based on vector field
Wujun Che, Xiaopeng Zhang 0001, Yi-Kuan Zhang, Jean-Claude Paul, Bo Xu 0002
Comput. Aided Geom. Des.5
2011 Direct quad-dominant meshing of point cloud via global parameterization
Er Li, Wujun Che, Xiaopeng Zhang 0001, Yi-Kuan Zhang, Bo Xu 0002
Comput. Graph.5
2011 Monaural voiced speech segregation based on elaborate harmonic grouping strategies
Xueliang Zhang 0001, Wei Jiang 0030, Peng Li 0030, Bo Xu 0002
Sci. China Inf. Sci.5
2010 Simplified Residual Factor Analysis for Text-Independent Speaker Verification
Rong Zheng 0005, Bo Xu 0002
ICASSP3
2010 Mandarin stress detection using hierarchical model based boosting classification and regression tree
abstract
Automatic stress detection is important for both speech understanding and natural speech synthesis. In this paper, we develop hierarchical model based boosting classification and regression tree (CART) to detect Mandarin stress by using acoustic evidence and text information. When comparing with previous proposed method at the same training and test sets, there are 2.52% and 1.09% absolute accuracy rate improvements respectively. We also analyze the differences between Mandarin stress detection and English pitch accent prediction, and prove some linguistic conclusions based on the large corpus in a different way.
Chongjia Ni, Bo Xu 0002
IJCNN3
2010 Automatic reference independent evaluation of prosody quality using multiple knowledge fusions
Shen Huang, Hongyan Li 0010, Shijin Wang 0001, Jiaen Liang, Bo Xu 0002
INTERSPEECH5
2010 Exploring goodness of prosody by diverse matching templates
abstract
In automatic speech grading systems, rare research is followed through addressing the issue of GOR (Goodness Of pRosody). In this paper we propose a novel method by taking the advantage of our QBH (Query By Humming) techniques in 2008 MIREX evaluation task. A set of standard samples related to the top-cream students are initially picked up as templates, a cascade QBH structure is then taken from two metrics: the MOMEL stylization followed by DTW distance; the Fujisaki model followed by EMD distance. Sentence GOR is obtained by the fused confidence between target and each template, and forms a weighted sum as the goodness in the passage level. Experiment results indicate that performance increases with the count of template, and Fujisaki-EMD metric outperforms MOMEL-DTW one in terms of correlation. Their combination can be treated as template based GOR score, compensated with our previous feature based GOR score, the approach can achieve 0.432 in correlation and 17.90% in EER in our corpus. Index Terms: speech prosody, query by humming
Shen Huang, Hongyan Li 0010, Shijin Wang 0001, Jiaen Liang, Bo Xu 0002
INTERSPEECH5
2010 Using prosody to improve Mandarin automatic speech recognition
Chongjia Ni, Bo Xu 0002
INTERSPEECH3
2010 An investigation into direct scoring methods without SVM training in speaker verification
Ce Zhang 0006, Rong Zheng 0005, Bo Xu 0002
INTERSPEECH3
2010 On the use of Gaussian component information in the generative likelihood ratio estimation for speaker verification
Rong Zheng 0005, Bo Xu 0002
INTERSPEECH2
2010 A new approach for automatic tone error detection in strong accented Mandarin based on dominant set
Taotao Zhu, Dengfeng Ke, Zhenbiao Chen, Bo Xu 0002
INTERSPEECH4
2010 Monaural speech separation based on MAXVQ and CASA for robust speech recognition
Peng Li 0030, Shijin Wang 0001, Bo Xu 0002
Comput. Speech Lang.4
2009 The Asian network-based speech-to-speech translation system
abstract
This paper outlines the first Asian network-based speech-to-speech translation system developed by the Asian Speech Translation Advanced Research (A-STAR) consortium. The system was designed to translate common spoken utterances of travel conversations from a certain source language into multiple target languages in order to facilitate multiparty travel conversations between people speaking different Asian languages. Each A-STAR member contributes one or more of the following spoken language technologies: automatic speech recognition, machine translation, and text-to-speech through Web servers. Currently, the system has successfully covered 9 languages-namely, 8 Asian languages (Hindi, Indonesian, Japanese, Korean, Malay, Thai, Vietnamese, Chinese) and additionally, the English language. The system's domain covers about 20,000 travel expressions, including proper nouns that are names of famous places or attractions in Asian countries. In this paper, we discuss the difficulties involved in connecting various different spoken language translation systems through Web servers. We also present speech-translation results on the first A-STAR demo experiments carried out in July 2009.
Sakriani Sakti, Noriyuki Kimura, Michael Paul, Chiori Hori, Eiichiro Sumita, Satoshi Nakamura 0001, Jun Park, Chai Wutiwiwatchai, Bo Xu 0002, Hammam Riza, Karunesh Arora, Haizhou Li 0001
ASRU9
2009 Context Dependent Feature Based Bottom-up Rescoring SVM Classifier in Children's English Stress Mis-pronunciation Detection
abstract
Automatic assessment of word stress error is an integral part for oral language grading system. However, problems that the property of vowels depends on its context information and the data sparseness of different vowel class are yet to be solved. This paper shall briefly introduce a hybrid method consisting of both traditional prosodic features and proposed context dependent strategies. In classification word stress is determined by weighting a bottom-up fashioned group tree with modified distributed probability score. In experiment, the overall equal error rate of our proposed system achieves 9.41%, which exhibits relative reduction and its competence of use in stress error detection system.
Shen Huang, Hongyan Li 0010, Shijin Wang 0001, Jiaen Liang, Bo Xu 0002
ICALT5
2009 Chinese intonation assessment using SEV features
abstract
Intonation assessment is an important part of Chinese CALL system. Nowadays, most systems use the correlation and RMSE features to assess the quality of the intonation of a given speech. As correlation and RMSE assign unoptimized weights to different degrees of mismatching errors, they may lead to performance degradation. In this paper, we propose a new feature called sorted error vector (SEV) for intonation assessment. The basic idea is to calculate mismatching quantities, sort them with ascending order, and then re-sample them to a K-points vector. This feature has four benefits: first, it is text-length independent; second, weights are let to train by classifiers; third, the relationship between the errors and the final results is not limited to any assumption; fourth, SEV is not sensitive to the performance of different pitch extracting algorithms. Experiments show that no matter in which case, SEV feature performs the best.
Dengfeng Ke, Bo Xu 0002
ICASSP2
2009 An efficient mispronounciation detction method using GLDS-SVM and formant enhanced features
abstract
Mispronunciation detection is an important component in computer assisted language learning (CALL) system. In this work, we introduce an efficient GLDS-SVM based detection method, which is successfully used in language and speaker identification systems, and combine it with traditional methods. The main ideas include: extended MFCC features with normalized formant trajectory information, and then propose a novel multi-model strategy for model training to make full use of samples and solve the problem of data unbalance, finally combine GLDS-SVM method with UBM-GMM system to further improve the performance. Experiments show that GLDS-SVM is highly efficient than traditional RBF-SVM, and the fused system can achieve a significant relative improvement of 17.5% in EER reduction, compared with the baseline UBM-GMM system.
Hongyan Li 0010, Jiaen Liang, Shijin Wang 0001, Bo Xu 0002
ICASSP4
2009 Automatic pronunciation error detection based on linguistic knowledge and pronunciation space
abstract
This paper presents a new approach that uses linguistic knowledge and pronunciation space for automatic detection of typical phone-level errors made by non-native speakers of mandarin. Firstly, linguistic knowledge of common learner mistakes is embedded in the calculation of log-posterior probability and the revised log-posterior probability (RLPP) is regarded as the measure of mispronunciation; secondly, a restricted pronunciation space is constructed by using RLPP vectors to describe the characteristics of pronunciation and Support Vector Machine (SVM) classifier is applied into the detection of typical pronunciation errors. Experiments based on a nonnative speaker database of mandarin confirm the promising effectiveness of our methods.
Zhenbiao Chen, Bo Xu 0002
ICASSP4
2009 Monaural voiced speech segregation based on elaborate harmonic grouping strategy
abstract
Monaural speech segregation is a very challenging problem which has been studied by many researchers. In this paper, we focus on voiced speech segregation. Different strategies are used to segregate resolved and unresolved harmonics respectively. For resolved harmonics, “harmonicity” principle and a novel mechanism based on “minimum amplitude” principle are employed. Amplitude modulation rate is extracted by “enhanced” autocorrelation function of envelope to segregate unresolved harmonics which is more robust than previous method. An elaborate rule is also introduced to determine the regions dominated by resolved and by unresolved harmonics. Proposed algorithm is evaluated on Cooke's 100 mixtures and compared with a state-of-the-art algorithm Hu and Wang model. Results show that proposed algorithm is more robust than the Hu and Wang model.
Xueliang Zhang 0001, Peng Li 0030, Bo Xu 0002
ICASSP4
2009 High performance automatic mispronunciation detection method based on neural network and TRAP features
Hongyan Li 0010, Shijin Wang 0001, Jiaen Liang, Shen Huang, Bo Xu 0002
INTERSPEECH5
2008 Improved phonotactic language identification using random forest language models
abstract
Recently a new language model, the random forest language model (RFLM), has been proposed and shown encouraging results in speech recognition tasks. In this paper we applied the RFLM to language identification tasks. We proposed a shared backoff smoothing to deal with data sparseness problem. Experiments were conducted on a subset of NIST 2003 language recognition evaluation data. The RFLM obtained 15.7% relative error rate reduction comparing with the standard trigram LM. The RFLM can be used as a counterpart to n-gram LM and BTLM for system fusion. We also empirically studied the relation between system performance and the tree numbers in a RFLM.
Shijin Wang 0001, Jiaen Liang, Bo Xu 0002
ICASSP4
2008 Query by humming via multiscale transportation distance in random query occurrence context
abstract
Query by humming (QBH) is an interactive tool for retrieving favored songs from a large database of known media via acoustic input. In this task, common method for measuring similarity between query and candidate is either by symbolic notation distance or by framed based dynamic programming. However, the former has disadvantage of error-prone to the noted symbolic feature extraction stage, while the latter is time-consuming. It has been proved that transportation distance has its remarkable merit in image query. However, we adopt a new structure for handling QBH, which is based on an improved version of this measure in combination with a string searching algorithm. More practically, we extend this method in random piece context, which means users can hum at any part of music piece. Experimental results are evaluated in MIREX 2007. Final 92.82% MRR has shown its significant advantages.
Shen Huang, Lei Wang 0062, Hongchen Jiang, Bo Xu 0002
ICME5
2008 An effective microphone array post-filter in arbitrary environments
abstract
The theoretic foundation of traditional microphone array post-filters is the signal model in which the noise between sensors is assumed to be uncorrelated. However, this model is inaccurate in real environments since the correlated noise exists. In this paper, a more generalized signal model which considers both the correlated and uncorrelated noise is introduced. A general expression of the microphone array post-filter is proposed for this model. For better residual noise shaping, the human auditory property is incorporated into the post-filter estimation process. In experiments with real noise microphone array recordings, the proposed technique has shown to produce impressive results in terms of quality measures of the enhanced speech. Index Terms: post-filter, generalized signal model, human auditory property, speech enhancement
Ning Cheng 0001, Peng Li 0030, Bo Xu 0002
INTERSPEECH4
2008 Improving searching speed and accuracy of query by humming system based on three methods: feature fusion, candidates set reduction and multiple similarity measurement rescoring
Lei Wang 0062, Shen Huang, Jiaen Liang, Bo Xu 0002
INTERSPEECH5
2007 Ungreedy Methods for Chinese Deterministic Dependency Parsing
Xiangyu Duan, Jun Zhao 0001, Bo Xu 0002
AAAI3
2007 Probabilistic Models for Action-Based Chinese Dependency Parsing
Xiangyu Duan, Jun Zhao 0001, Bo Xu 0002
ECML3
2007 Probabilistic Parsing Action Models for Multi-Lingual Dependency Parsing
Xiangyu Duan, Jun Zhao 0001, Bo Xu 0002
EMNLP-CoNLL3
2007 A Novel Phone-State Matrix Based Vocabulary-Indenendent Keyword Spotting Method for Spontaneous Speech
abstract
Keyword spotting (KWS) is an essential technique for speech information retrieval. When doing offline keyword query on large volume spontaneous speech data, fast and accurate KWS methods are required. In this paper, a novel phone-state matrix based vocabulary-independent KWS method is proposed, which has merits of both hidden Markov model (HMM) based and lattice-based methods. Four KWS systems are compared in our experiments on conversational telephone speech test set. Result shows that compared to the high precision HMM-based KWS system the proposed phone-state matrix system has better equal-error-rate (EER) and false-alarm (FA) performance than the other two lattice-based systems.
Jiaen Liang, Peng Ding 0003, Bo Xu 0002
ICASSP (4)4
2007 Word Sense Disambiguation through Sememe Labeling
Xiangyu Duan, Jun Zhao 0001, Bo Xu 0002
IJCAI3
2006 Applying Pitch Target Model to Convert F0 Contour for Expressive Mandarin Speech Synthesis
abstract
In the paper, pitch target model is employed to represent and convert F0 contour for synthesizing an emotional Mandarin speech from a neutral speech. Compared with conventional F0 transforming methods, the proposed method converts F0 patterns described by pitch target parameters rather than F0 contours themselves, and uses Gaussian Mixture Model(GMM) and Classification and Regression Trees (CART) methods to build mapping functions for well-chosen pitch target parameters. Other prosodic parameters such as duration and intensity are also converted. Listening tests prove that these converted speeches express corresponding emotional states.
Yongguo Kang, Jianhua Tao 0001, Bo Xu 0002
ICASSP (1)3
2006 An Improved Mandarin Keyword Spotting System Using MCE Training and Context-Enhanced Verification
abstract
The task of keyword spotting is to detect a set of keywords in the input continuous speech. The main goal of this work is to develop an improved Mandarin keyword spotting (KWS) system for conversational telephone speech (CTS). In this paper, we propose an efficient online-garbage model based KWS system, which integrated with a word-level minimum classification error (MCE) training method and a novel context-enhanced verification method. Experiment showed that the proposed methods can reduce the equal-error-rate (EER) of the system by 13.8% in relative
Jiaen Liang, Peng Ding 0003, Bo Xu 0002
ICASSP (1)5
2006 A Novel Noise Robust Front-End Using First Order VTS in Construction of Mel-Warped Wiener Filter
abstract
In this paper, we first review two approaches in the context of robust recognition, e.g. speech enhancement based two-stage mel-warp Wiener filtering (MWF) (A. Agarwal and Y.M. Cheng, 1999) and first-order vector Taylor series (VTS) (P.J. Moreno et al., 1996) compensation in log power spectrum, which are widely used. A new noise robust front-end is proposed, in which VTS compensation derived statistics are used to construct the mel-warped Wiener filter. We will show that this noise robust front end is superior. The experiments results prove that our proposed method does show significant improvement over VTS and MWF
Mu Su, Peng Li 0030, Peng Ding 0003, Bo Xu 0002
ICASSP (1)5
2006 One-Pass Coarse-to-Fine Segmental Speech Decoding Algorithm
abstract
In this paper, a novel one-pass coarse-to-fine decoding algorithm is proposed to accelerate the speed of segment model (SM). The algorithm is originated from the segmentation similarity observation described in the paper and is specific for the SM based speech recognition. At each step, a coarse search is first implemented to get coarse segmentations and then a fine search is performed based on the derived segmentation information. This fast algorithm is successfully integrated into an SM based Mandarin LVCSR system and saves more than 50% decoding time without obvious influence on the recognition accuracy
Yun Tang 0002, Hua Zhang 0009, Bo Xu 0002, Guo-Hong Ding
ICASSP (1)4
2006 Fast SVM training based on the choice of effective samples for audio classification
Shilei Zhang, Hongchen Jiang, Shuwu Zhang, Bo Xu 0002
INTERSPEECH4
2006 A quality measure method using Gaussian mixture models and divergence measure for speaker identification
Rong Zheng 0005, Shuwu Zhang, Bo Xu 0002
INTERSPEECH3
2006 An approach to automatic acquisition of translation templates based on phrase structure extraction and alignment
abstract
In this paper, we propose a new approach for automatically acquiring translation templates from unannotated bilingual spoken language corpora. Two basic algorithms are adopted: a grammar induction algorithm, and an alignment algorithm using bracketing transduction grammar. The approach is unsupervised, statistical, and data-driven, and employs no parsing procedure. The acquisition procedure consists of two steps. First, semantic groups and phrase structure groups are extracted from both the source language and the target language. Second, an alignment algorithm based on bracketing transduction grammar aligns the phrase structure groups. The aligned phrase structure groups are post-processed, yielding translation templates. Preliminary experimental results show that the algorithm is effective.
Rile Hu, Chengqing Zong, Bo Xu 0002
IEEE Trans. Speech Audio Process.3
2006 Monaural Speech Separation Based on Computational Auditory Scene Analysis and Objective Quality Assessment of Speech
abstract
Monaural speech separation is a very challenging problem in speech signal processing. It has been studied extensively, and many separation systems based on computational auditory scene analysis (CASA) have been proposed in the last two decades. Although the research on CASA has tended to introduce high-level knowledge into separation processes using primitive data-driven methods, the knowledge on speech quality still has not been combined with it. This makes the performance evaluation of CASA mainly focused on the signal-to-noise ratio (SNR) improvement. Actually, the quality of the separated speech is not directly related to its SNR. In order to solve this problem, we propose a new method which combines CASA with objective quality assessment of speech (OQAS). In the grouping process of CASA, we use OQAS as the guide to instruct the CASA system. With this combination, the performance of the speech separation can be improved not only in SNR, but also in mean opinion score (MOS). Our system is systematically evaluated and compared with previous systems, and it yields substantially better performance, especially for the subjective perceptual quality of separated speech.
Peng Li 0030, Bo Xu 0002
IEEE Trans. Speech Audio Process.3
2005 Investigation of Emotive Expressions of Spoken Sentences
Wenjie Cao, Chengqing Zong, Bo Xu 0002
ACII3
2005 A Hybrid GMM and Codebook Mapping Method for Spectral Conversion
Yongguo Kang, Zhiwei Shuang, Jianhua Tao 0001, Wei Zhang 0031, Bo Xu 0002
ACII5
2005 Chinese prosodic phrasing with a constraint-based approach
abstract
The linguistic constraints and phrase-length constraints are the most important factors for Chinese prosodic phrasing in the natural speech. This paper presents a linguistic constraint model and a phrase-length constraint model to describe these two processes independently. Therefore, each part can be described in detail. In the linguistic constraints model, Chunk (base phrase) is considered as an important basic unit. And an HMM is used to model the phrase-length constraints, which concerns the distribution of the prosodic phrase lengths and the relationship between the prosodic phrase and the prosodic words. Then a k-candidate method is introduced to combine these two models. This approach makes full use of the linguistic constraints and the phrase-length constraints. The experiments show that this approach achieved a perfect performance with the phrasing f-score 82.9%.
Honghui Dong, Jianhua Tao 0001, Bo Xu 0002
INTERSPEECH3
2005 Optimal model order selection based on regression tree in speaker identification
Shilei Zhang, Junmei Bai, Shuwu Zhang, Bo Xu 0002
INTERSPEECH4
2004 Chinese-English bilingual phone modeling for cross-language speech recognition
abstract
In this paper, three different approaches to Chinese-English bilingual phone modeling are investigated and compared. The first approach is to simply combine Chinese and English phone inventories together without phone sharing across the languages. The second one is to map language-dependent phones to the inventory of the International Phonetic Association (IPA) based on phonetic knowledge to construct the bilingual phone inventory. The third one is to merge the language-dependent phone models by an hierarchical phone clustering algorithm to get a compact bilingual inventory. In the third approach, two distance measures are used to perform the bottom-up clustering. One is the Bhattacharyya distance. The other is the acoustic likelihood distance. Experimental results show that the phone clustering approach outperforms the IPA-based phone mapping approach, and it can also achieve comparable performance to the simple combination of language-dependent phone inventories with fewer model parameters, especially when using acoustic likelihood distance measurement.
Shengmin Yu, Shuwu Zhang, Bo Xu 0002
ICASSP (1)3
2004 Window-Based Method for Information Retrieval
Qianli Jin, Jun Zhao 0001, Bo Xu 0002
IJCNLP3
2004 Bilingual Chunk Alignment Based on Interactional Matching and Probabilistic Latent Semantic Indexing
Qianli Jin, Jun Zhao 0001, Bo Xu 0002
IJCNLP4
2004 Approach to interchange-format based Chinese generation
abstract
Interlingua-based machine translation is an important approach to implement multi-lingual speech-to-speech (S2S) translation. The natural language generation (NLG) is one of the key components in the interlingua-based machine translation systems. This paper introduces our approach to Chinese generation based on the Interchange Format (IF) developed by the C-STAR organization. In our approach, the hybrid method of feature-based deep generation method and template-based method are employed. The deep generator ensures that the generation component possesses the merits of flexibility and domain portability. The template-based generator makes the system more efficient. We also introduce another simplified Chinese generator applied in specific domain. The experimented results show that our approach is effective and practical for the natural language generation in the Interchange-Format (IF) based S2S translation system. 1.
Wenjie Cao, Chengqing Zong, Bo Xu 0002
INTERSPEECH3
2004 Exploring high-performance speech recognition in noisy environments using high-order taylor series expansion
Guo-Hong Ding, Bo Xu 0002
INTERSPEECH2
2004 A novel target-driven generalized JMAP adaptation algorithm
abstract
Adapting the parameters of a statistical speaker independent continuous speech recognizer to the speaker can significantly improve the recognition performance and robustness of the system. In this paper, we propose a novel target-driven speaker adaptation method, Generalized Joint Maximum a Posteriori (GJMAP), which extends and improves the previous successful method JMAP. GJMAP partitions the HMM parameters with respect to the adaptation data, using the priori phonetic knowledge. The generation of regression class trees is dynamically constructed on the target-driven principle in order to obtain the maximum increase of the auxiliary function. An off-line adaptation experiment on large vocabulary continuous speech recognition is carried out. The experimental results show GJMAP has more advantages than the conventional methods.
Zhaobing Han, Shuwu Zhang, Bo Xu 0002
INTERSPEECH3
2004 Combining agglomerative and tree-based state clustering for high accuracy acoustic modeling
Zhaobing Han, Shuwu Zhang, Bo Xu 0002
INTERSPEECH3
2004 Multi-layer structure MLLR adaptation algorithm with subspace regression classes and tying
Xiangyu Mu, Shuwu Zhang, Bo Xu 0002
INTERSPEECH3
2004 A new multicomponent AM-FM demodulation with predicting frequency boundaries and its application to formant estimation
abstract
In this paper, a method using dynamic programming to predict frequency boundaries is proposed for the joint demodulation of amplitude modulation (AM) and frequency modulation (FM) for speech signals. Because of the existence of modulations in speech signal, an algorithm called energy separation algorithm (ESA) has been developed to track the energy needed by a source to produce the speech signal, and this algorithm provides an efficient solution to separate output energy product into amplitude modulation and frequency modulation components. For multicomponent AM-FM signals like speech signals, a bank of bandpass filters or a set of individual bandpass filters, whose center frequency and critical bandwidth commonly are selected through experiential selection, is necessary to get monocomponent signals. Our experimental results provide that the bandpass filter with predicted frequency boundaries instead of experiential selection is more effective in AM-FM demodulation. Formant estimation based on this demodulating method also proves it is efficient and formant tracking algorithm is not necessary at all in the estimating procedure.
Bo Xu 0002, Jianhua Tao 0001, Yongguo Kang
INTERSPEECH1
2004 Suppression of additive noise using a power spectral density MMSE estimator
abstract
In this letter, we propose a novel speech enhancement approach, called power spectral density minimum mean-square error (PSD-MMSE) estimation-based speech enhancement, which is implemented in the power spectral domain where stationary stochastic noise can be modeled as the exponential distribution. Speech magnitude-squared spectra are modeled as the mixed exponential distribution. And an MMSE estimator is constructed based on the parametric distributions. Besides, a fast algorithm is presented to implement the approach in real time. Experimental results of Itakura-Saito distortion measures show that the proposed approach is superior to alternative speech enhancement algorithms.
Guo-Hong Ding, Taiyi Huang 0001, Bo Xu 0002
IEEE Signal Process. Lett.3
2003 A Maximum Entropy Approach for Spoken Chinese Understanding
Guodong Xie, Chengqing Zong, Bo Xu 0002
CICLing3
2003 Fast speaker adaptation using triple diagonal and shared block diagonal transform matrices
abstract
This paper proposes two fast and effective adaptation algorithms, which are called SATD and SASBD respectively. The two algorithms are implemented in the MLLR frame and the transform matrices have constrained forms. SATD uses triple diagonal matrices to describe the mismatch between speakers and the acoustic model in the log-spectral domain and the matrices can be transformed into the cepstral domain to adjust the acoustic model. SASBD is different from the traditional block-diagonal MLLR and shares the three transformations of basic MFCC and dynamic features with one matrix. Moreover, both algorithms provide multiple choices for the biases. Experiments are extensively implemented and the results prove the advantages of SATD and SASBD over traditional MLLR.
Guo-Hong Ding, Bo Xu 0002, Juha Iso-Sipilä
ICASSP (1)2
2003 Comparison and study of some variants of partially tied covariance modeling
abstract
Some practical implementation issues on partially tied covariance (PTC) modeling are discussed. First, from the view of model complexity and computational load, a comparison is made for some variants of PTC. From the analysis, two representatives, STC and Ortho-STC are compared in detail. Second, based on these variants, two techniques are studied. One technique is joint optimization of both transformation and HMM parameters, which will exploit the potential of PTC. The other technique is model selection by hierarchical tree via Bayesian information criterion (BIC), which will decide the number and structure of transformation classes thus to assure the generalization capacity. Experiment results showed that STC always outperforms Ortho-STC due to the effect of parameter tying and by the application of above two techniques the system performance can be much improved.
Peng Ding 0003, Shuwu Zhang, Bo Xu 0002
ICASSP (1)3
2003 A vector statistical piecewise polynomial approximation algorithm for environment compensation in telephone LVCSR
abstract
A vector statistical piecewise polynomial (VPP) approximation algorithm is proposed for environment compensation in speech signals that are degraded by both additive and convolutive noise. By investigating the model of the telephone environment, we address a piecewise polynomial, namely two linear polynomials and a quadratic polynomial, to approximate the environment function precisely. The VPP is applied either to stationary noise, or to non-stationary noise. In the first case, batch EM is used in the log-spectral domain; in the second case, recursive EM with iterative stochastic approximation is developed in the cepstral domain. Both approaches are based on the minimum mean squared error (MMSE) sense. Experimental results are presented on the application of this approach in improving the performance of Mandarin large vocabulary continuous speech recognition (LVCSR) in background noise and different transmission channels (such as fixed telephone line and GSM). The method can reduce the average character error rate (CER) by about 18%.
Zhaobing Han, Shuwu Zhang, Huayun Zhang, Bo Xu 0002
ICASSP (2)4
2003 Sequential MAP estimation based speech feature enhancement for noise robust speech recognition
abstract
In this paper, the environment mismatch due to additive noise is assumed as an additive bias in power spectral domain. It is viable to introduce some constraints on the values that the bias can take due to the internal relation between bias and noise power spectrum. We propose to introduce the noise priori knowledge into bias estimation process by using maximum a posteriori (MAP) criterion. Moreover, the mismatch is usually nonstationary in real application and sequential algorithm can be used to track time varying environment within a test utterance. This paper proposes to use the sequential techniques to estimate the bias in the MAP framework and update the parameters of noise priori adaptively. Speech recognition experiments demonstrated that the proposed algorithm outperformed sequential ML estimation method and was obviously better than the batch mode under non-stationary noise environment.
Chuan Jia, Peng Ding 0003, Bo Xu 0002
ICASSP (1)3
2003 Discriminative optimization of large vocabulary Mandarin conversational speech recognition system
Peng Ding 0003, Zhenbiao Chen, Shuwu Zhang, Bo Xu 0002
INTERSPEECH5
2003 Joint model and feature based compensation for robust speech recognition under non-stationary noise environments
Chuan Jia, Peng Ding 0003, Bo Xu 0002
INTERSPEECH3
2003 Statistical speech-to-speech translation with multilingual speech recognition and bilingual-chunk parsing
Bo Xu 0002, Shuwu Zhang, Chengqing Zong
INTERSPEECH1
2003 Dynamic channel compensation based on maximum a posteriori estimation
Huayun Zhang, Zhaobing Han, Bo Xu 0002
INTERSPEECH3
2003 Geometric constrained maximum likelihood linear regression on Mandarin dialect adaptation
Huayun Zhang, Bo Xu 0002
INTERSPEECH2
2002 Chinese Syntactic Parsing Based on Extended GLR Parsing Algorithm with PCFG*
Bo Xu 0002, Chengqing Zong
COLING2
2002 Asymmetrical Support Vector Machines and applications in speech processing
abstract
Support Vector Machines have merged as a pattern classifier and have been shown to be successful in some tasks in the realm of speech processing. This paper explores the issues involved in applying SVMs to asymmetrical situations, namely. beavy sample ratio bias between different classes and different costs for different types of misclassification error. We also present our revisions on the SMO algorithm to make the asymmetrical SVM training procedure practical. Experiments on both recognition of isolated spoken digits in mandarin and the learning of the decision function for speaker authentication yielded performance improvements, which show the effectiveness of asymmetrical SVMs.
Peng Ding 0003, Zhenbiao Chen, Yang Liu 0085, Bo Xu 0002
ICASSP4
2002 Study on prosodic boundary location in Chinaese mandarin
abstract
In this paper, based on large speech corpus with prosodic structure label (ASCCD), we present some statistic result on acoustic parameter at prosodic boundary. We study the syllable duration, intensity and pitch at the boundary and select a serial acoustic parameter to train a CART. Then the CART was employed to classify the prosodic boundary type. The result shows that the parameter characterize acoustic feature of the prosodic boundary and the trained CART can classify different boundary efficiency. So it is possible to train statistical model for prosodic boundary location in Mandarin, this is very important both for speech recognition and synthesis.
Weixiang Hu, Taiyi Huang 0001, Bo Xu 0002
ICASSP3
2002 Including detailed information feature in MFCC for large vocabulary contious speech recornition
abstract
This paper focuses on the inclusion of more detailed linguistically relevant speech information in the Mel-Frequency Cepstral Coefficients(MFCC) feature extraction process in order to improve the recognition accuracy of LVCSR. Detailed linguistically relevant speech information feature is extracted to reflect the change of energy spectrum in each mel-frequency bank(MFB). A normalized positive weighting vector is used to combine the log channel energy feature of the standard MFCC with the new detailed information features to form one energy feature for each MFB. The optimal weighting vector can be obtained by the Heteroscedastic Discriminant Analysis (HDA) before feature extraction. Experiments on two test sets show that the new feature extraction method is superior in performance to the standard MFCC and 10% relative error reduction for LVCSR is witnessed in the test set with standard accent speakers.
Bo Xu 0002
ICASSP2
2002 Using nonstandard SVM for combination of Speaker Verification and Verbal Information Verification in speaker authentication system
abstract
How to extract more identity-related information from speech signal is a crucial problem in the research of speaker authentication. In this paper, we present a novel framework for combining Speaker Verification (SV) System with Verbal Information Verification (VIV) System by proposing the nonstandard-SVM method. The scheme of nonstandard-SVMs makes it possible to adjust the tradeoff between False Acceptance Error and False Rejection Error without the loss of the good generalization performance of traditional SVMs. Further more, experiments show that the proposed method improves the authentication performance relatively 80%, 60% and 50% compared with SV, VIV and the conventional combination system respectively.
Yang Liu 0085, Peng Ding 0003, Bo Xu 0002
ICASSP3
2002 Pitch and tone's modeling in parametric trajectory model
abstract
It is described in this paper for the application of pitch/tone information in the parametric trajectory model. Pitch as a dynamic feature and its contours—tone as a segmental-level feature are deserved their own particular characteristics, which match case of parametric trajectory model better compared with MFCC and energy. Here we give an improved pitch extraction algorithm and especially the “total” and “parallel” integration methods to combine these information with the base model. In the experiment of Mandarin connected digit recognition, we achieve 22.87% and 33.54% error reduction respectively for them, moreover when combined with these two methods, 38.72% error reduction is obtained.
Bo Xu 0002, Huayun Zhang
ICASSP3
2002 Covariance-Tied Clustering Method In Speaker Identification
abstract
Gaussian mixture models (GMMs) have been successfully applied to the classifier for speaker modeling in speaker identification. However, there are still problems to solve, such as the clustering methods. The conditional k-means algorithm utilizes Euclidean distance taking all data distribution as sphericity, which is not the distribution of the actual data. In this paper we present a new method making use of covariance information to direct the clustering of GMMs, namely covariance-tied clustering. This method consists of two parts: obtaining covariance matrices using the data sharing technique based on a binary tree, and making use of covariance matrices to direct clustering. The experimental results prove that this method leads to worthwhile reductions of error rates in speaker identification. Much remains to be done to explore fully the covariance information.
Yang Liu 0085, Peng Ding 0003, Bo Xu 0002
ICMI4
2002 Factor analyzed Gaussian mixture models for speaker identification
Peng Ding 0003, Yang Liu 0085, Bo Xu 0002
INTERSPEECH3
2002 Implementing vocal tract length normalization in the MLLR framework
Guo-Hong Ding, Chengrong Li, Bo Xu 0002
INTERSPEECH4
2002 Parametric trajectory segment model for LVCSR
Bo Xu 0002
INTERSPEECH2
2002 Chinese spoken language analyzing based on combination of statistical and rule methods
abstract
A combination of statistical and rule methods has been developed for Chinese spoken language analyzing. The analyzing result is a middle semantic frame, which can be converted to different language according people’s needs. We adopt the statistical method in the stage of extracting semantic meaning and the rule method in the stage of mapping the semantic units to middle semantic frame. Experiment shows this method has high robustness and can analyzing Chinese spoken language effectively. 1.
Guodong Xie, Chengqing Zong, Bo Xu 0002
INTERSPEECH3
2002 Codebook dependent dynamic channel estimation for Mandarin speech recognition over telephone
Huayun Zhang, Zhaobing Han, Bo Xu 0002
INTERSPEECH3
2002 Improving parametric trajectory modeling by integration of pitch and tone information
Bo Xu 0002, Huayun Zhang
INTERSPEECH3
2001 A novel target-driven MLLR adaptation algorithm with multi-layer structure
Bo Xu 0002
INTERSPEECH2
2001 Study and auto-detection of stress based on tonal pitch range in Mandarin
abstract
In Mandarin, there is a special acoustic feature—tonal pitch range, which is relative to stress. In this paper, we present a novel concept—tonal range ratio (TRR), which is based on tonal pitch range, and make a study on the correlation between TRR and stress in Mandarin. And we developed a system to automatically detect stresses in words and sentences based on TRR in Mandarin. We obtained high success rate (92.67 % in words and 82.0 % in sentences). The results show that TRR has strong correlation with stress and is powerful in detecting stresses in Mandarin. 1.
Xipeng Shen, Bo Xu 0002
INTERSPEECH2
2001 The study of the effect of training set on statistical language modeling
abstract
In this work, we make a study on the effect of training set on statistical language modeling (SLM). A corpus selection system based on perplexity is presented. It is tested in two experiments: one is to select optimal training corpus for generating a domain-specific SLM; the other one is for generating an optimal SLM for a LVCSR system. The results show that the training corpus is important for the capability of SLM and our corpus selection system is powerful for optimal corpus selection. With the help of this system, we generated a SLM for a LVCSR system, which contributed 14.5%--17.7 % relative character error reduction. 1.
Xipeng Shen, Bo Xu 0002
INTERSPEECH2
2000 Decision tree based Mandarin tone model and its application to speech recognition
abstract
Tone is an essential language phenomenon for Mandarin Chinese language. Until now, we still do not know exactly how context affects tone pattern variation in continuous Mandarin speech. In this paper, we proposed a decision tree based approach to obtain the quantitative result of tone pattern variation in continuous Mandarin speech. Many possible factors other than tone of neighboring syllables were taken into consideration when the decision tree was constructed, After the tree was established, 29 tone patterns were automatically obtained, and we found that syllable position in the word together with consonant/vowel type of the syllable made an important contribution to tone pattern variation in continuous utterance. We also presented a novel approach to integrate tone information into the search process at word level. Experimental results showed that the character error rate was reduced by 15.2%.
Yonggang Deng, Taiyi Huang 0001, Bo Xu 0002
ICASSP5
2000 Acoustic modeling for Chinese speech recognition: a comparative study of Mandarin and Cantonese
abstract
This paper presents a comparative study on automatic speech recognition for two different Chinese dialects, namely Mandarin and Cantonese. It focuses on decision-tree based context-dependent acoustic modeling for large-vocabulary continuous speech recognition. Extensive phonological and phonetic knowledge are incorporated to design questions concerning the left and right context of sub-syllable units, namely INITIALs and FINALs. This results in a set of class-triphone models for each dialect. Syllable recognition accuracy of 81.7% and 75.5% are attained for Mandarin and Cantonese respectively. Such a performance gap is accountable by various linguistic and practical reasons, including: 1) phonological and phonetic discrepancies between the two dialects; 2) design of training databases; and 3) design of phonetic questions in decision-tree clustering.
Tan Lee, Yiu Wing Wong, Bo Xu 0002, Pak-Chung Ching, Taiyi Huang 0001
ICASSP4
2000 Mandarin accent adaptation based on context-independent/context-dependent pronunciation modeling
abstract
An accent adaptation approach using pronunciation variation modeling technology for the Mandarin accent was proposed in this paper. As the Chinese language is monosyllabic, the syllable pronunciation variation dictionary (SPVD) was built to depict the characteristics of accent. Firstly, the pronunciation modeling technology was utilized to get the context-independent and context-dependent accent-specific syllable confusion matrix according to the acoustic recognition results (pin-yin stream). Then the accent-specific Chinese SPVD was constructed from this confusion matrix. Finally, N-Best acoustic recognition candidates were re-scored with the help of SPVD. The proposed method was experimented on INTEL Shanghai-Accent Mandarin Corpus and the 863 standard Mandarin acoustic recognizer. The syllable recognition error rate was reduced 15% by context-independent SPVD, and 20% by context-dependent SPVD.
Mingkuan Liu, Bo Xu 0002, Taiyi Huang 0001, Yonggang Deng, Chengrong Li
ICASSP2
2000 Statistical Analysis of Chinese Language and Language Modeling Based on Huge Text Corpora
Bo Xu 0002, Taiyi Huang 0001
ICMI2
2000 Approach to Recognition and Understanding of the Time Constituents in the Spoken Chinese Language Translation
Chengqing Zong, Taiyi Huang 0001, Bo Xu 0002
ICMI3
2000 A stochastic polynomial tone model for continuous Mandarin speech
Taiyi Huang 0001, Bo Xu 0002, Chengrong Li
INTERSPEECH3
2000 Neural network based integration of multiple confidence measures for OOV detection
Yonggang Deng, Bo Xu 0002
INTERSPEECH3
2000 Towards high performance continuous Mandarin digit string recognition
Yonggang Deng, Taiyi Huang 0001, Bo Xu 0002
INTERSPEECH3
2000 Chinese spoken language understanding across domain
Yunbin Deng, Bo Xu 0002, Taiyi Huang 0001
INTERSPEECH2
2000 Update progress of Sinohear: advanced Mandarin LVCSR system at NLPR
Bo Xu 0002, Chengrong Li, Taiyi Huang 0001
INTERSPEECH2
2000 Accent-specific Mandarin adaptation based on pronunciation modeling technology
Mingkuan Liu, Bo Xu 0002
INTERSPEECH2
2000 A generation system for Chinese texts
Taiyi Huang 0001, Bo Xu 0002
INTERSPEECH3
2000 How to choose training set for language modeling
Bo Xu 0002, Taiyi Huang 0001
INTERSPEECH2
2000 Incorporating HMM-state sequence confusion for rapid MLLR adaptation to new speakers
Bo Xu 0002
INTERSPEECH2
2000 An improved template-based approach to spoken language translation
abstract
In this paper, we describe an improved template-based approach to Chinese-to-English Spoken Language Translation (SLT) and present experimental results. The improved template-based translation approach uses flexible expression format to describe the template condition. The condition of a template may consist of keywords, parts-of-speech and also semantic features, so the input may be matched with a template from shallow level to deep level. In the condition of a template, the distance between two fixed keywords is stretchable, thus some needless words in the input utterances may be skipped in matching operation. And also the translation results of the same template are alterable. The proper results are finally generated according to the specific context. That is, the relation between a template and translated utterance is one-to-n (where, n is an integer and n≥1). The experiments were performed with input of both text transcription and results of speech recognition. The preliminary experimental results have proven the approach is practical. 1.
Chengqing Zong, Taiyi Huang 0001, Bo Xu 0002
INTERSPEECH3
2000 Japanese-to-Chinese spoken language translation based on the simple expression
abstract
This paper describes a Japanese-to-Chinese spoken language translation (SLT) method based on simple expression and presents the experimental results. The method is aimed at developing a compact speech translation system, which is robust for spontaneous spoken language phenomena, including the recognition errors and different expression from various speakers. The idea of translation method based on simple expression is that the mechanism interprets speech-act rather than the direct translation of the speaker’s words. The method is realized by mapping the simple expression instead of deep
Chengqing Zong, Yumi Wakita, Bo Xu 0002, Zhenbiao Chen, Kenji Matsui
INTERSPEECH3
1999 A novel model TD-PSPTP for speech synthesis
Bo Xu 0002
EUROSPEECH2
1999 LODESTAR: a Mandarin spoken dialogue system for travel information retrieval
abstract
This paper described the THISL spoken document retrieval system for British and North American Broadcast News. The system is based on the ABBOT large vocabulary speech recognizer and a probabilistic text retrieval system. We discuss the development of a realtime British English Broadcast News system, and its integration into a spoken document retrieval system. Detailed evaluation is performed using a similar North American Broadcast News system, to take advantage of the TREC SDR evaluation methodology. We report results on this evaluation, with particular reference to the effect of query expansion and of automatic segmentation algorithms.
Xin Zhang 0072, Shubin Zhao, Taiyi Huang 0001, Bo Xu 0002
EUROSPEECH6
1999 Regression class selection and speaker adaptation with MLLR in Mandarin continuous speech recognition
Chengrong Li, Jingdong Chen, Bo Xu 0002
EUROSPEECH3
1998 A novel robust feature of speech signal based on the Mellin transform for speaker-independent speech recognition
abstract
This paper presents a novel kind of speech feature which is the modified Mellin transform of the log-spectrum of the speech signal (short for MMTLS). Because of the scale invariance property of the modified Mellin transform, the new feature is insensitive to the variation of the vocal tract length among individual speakers, and thus it is more appropriate for speaker-independent speech recognition than the popular used cepstrum. The preliminary experiments show that the performance of the MMTLS-based method is much better in comparison with those of the LPC- and MFC-based methods. Moreover, the error rate of this method is very consistent for different outlier speakers.
Jingdong Chen, Bo Xu 0002, Taiyi Huang 0001
ICASSP2
1996 Context-dependent acoustic models for Chinese speech recognition
abstract
Selecting good speech units and building precise acoustic models on these units are basic problems for HMM speech recognition systems. In this paper, we review some distinguished phonetic features of the Chinese language and show how these features could be considered for getting a better solution of speech units selection and acoustic model building. The initials and finals are suggested to be the recognition units. Several experiments have been carried out to find those proper acoustic models which can represent the co-articulation among initials and finals more accurately and give a higher recognition rate. Results of recognition error rate under different acoustic models are given and some comparisons are made. Experiments show that not only the intra-syllable context-dependent acoustic models but also the inter-syllable acoustic models are necessary for reducing recognition error rate of words and continuous speech recognition. Experiments also show the tone-pattern information is important for accurate syllable recognition.
Bin Ma 0012, Taiyi Huang 0001, Bo Xu 0002, Xijun Zhang, Fei Qu
ICASSP3
1996 Speaker-independent dictation of Chinese speech with 32k vocabulary
abstract
While early machines adopted isolated syllable as input units and needed boring enrollment, our research focus on the speaker-independent, word-based dictation.A deliberately designed 120-speaker database was built for training ; inter-syllable context ,tonal and endpoint dependent acoustic model are applied with promising MFCC feature; Two-pass acoustic matching accelerates the recognition making fully advantage of the monosyllabic structure of Chinese speech; A complete word bigram and trigram serve as language processing module.With all efforts, the system reaches 90% character accuracy performing in almost real-time on Pentium PC without DSP help.
Bo Xu 0002, Shuwu Zhang, Fei Qu, Taiyi Huang 0001
ICSLP1
1994 Adaptation of neural network model: comparison of multilayer perceptron and LVQ
Dongxin Xu, Dao Wen Chen, Bo Xu 0002, Taiyi Huang 0001
ICSLP4
1992 A. 46 500 word Chinese speech recognition system
Bo Xu 0002, Taiyi Huang 0001, Dongxin Xu
ICSLP1
1991 A real-time Chinese speech recognition system with unlimited vocabulary
abstract
A Chinese speech recognizer with unlimited vocabulary is described. The system has two major components: the acoustic recognition component, which includes an HMM (hidden Markov model)-based phone recognizer, a NN (neural network)-based initial refiner, and a NN-based tone classifier; and the lexical and homonym processor, which is based on a knowledge database extracted from large amounts of texts. This real-time recognizer is implemented on a PC-386 enhanced by only one digital signal processing board on which a TMS-320c25 chip operates as the CPU. On average, it takes only 0.19 s to recognize a one-syllable word. The recognition accuracy for syllables, tones, and words is 92.5%, 99.6%, and 97.5%, respectively.>
Taiyi Huang 0001, Bo Xu 0002, Dongxin Xu
ICASSP4