Shen Huang

dblp:90/839 · DBLP profile ↗
← Back
51ranked-venue papers
15as first author
20since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 31 · 8 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 4 first-author · 11 since 2021Databases, data management, data science and information retrieval · 10 · 5 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Evidence-Augmented Policy Optimization with Reward Co-Evolution for Long-Context Reasoning
abstract
While Reinforcement Learning (RL) has advanced LLM reasoning, applying it to longcontext scenarios is hindered by sparsity of outcome rewards.This limitation fails to penalize ungrounded "lucky guesses," leaving the critical process of needle-in-a-haystack evidence retrieval largely unsupervised.To address this, we propose EAPO (Evidence-Augmented Policy Optimization).We first establish the Evidence-Augmented Reasoning paradigm, validating via Tree-Structured Evidence Sampling that precise evidence extraction is the decisive bottleneck for long-context reasoning.Guided by this insight, EAPO introduces a specialized RL algorithm where a reward model computes a Group-Relative Evidence Reward, providing dense process supervision to explicitly improve evidence quality.To sustain accurate supervision throughout training, we further incorporate an Adaptive Reward-Policy Co-Evolution mechanism.This mechanism iteratively refines the reward model using outcome-consistent rollouts, sharpening its discriminative capability to ensure precise process guidance.Comprehensive evaluations across eight benchmarks demonstrate that EAPO significantly enhances long-context reasoning performance compared to SOTA baselines.
Shen Huang, Pengjun Xie, Jingren Zhou 0001, Jiuxin Cao
ACL (1)3
2025 Full-text Error Correction for Chinese Speech Recognition with Large Language Model
abstract
Large Language Models (LLMs) have demonstrated substantial potential for error correction in Automatic Speech Recognition (ASR). However, most research focuses on utterances from short-duration speech recordings, which are the predominant form of speech data for supervised ASR training. This paper investigates the effectiveness of LLMs for error correction in full-text generated by ASR systems from longer speech recordings, such as transcripts from podcasts, news broadcasts, and meetings. First, we develop a Chinese dataset for full-text error correction, named ChFT, utilizing a pipeline that involves text-to-speech synthesis, ASR, and error-correction pair extractor. This dataset enables us to correct errors across contexts, including both full-text and segment, and to address a broader range of error types, such as punctuation restoration and inverse text normalization, thus making the correction process comprehensive. Second, we fine-tune a pre-trained LLM on the constructed dataset using a diverse set of prompts and target formats, and evaluate its performance on full-text error correction. Specifically, we design prompts based on full-text and segment, considering various output formats, such as directly corrected text and JSON-based error-correction pairs. Through various test settings, including homogeneous, up-to-date, and hard test sets, we find that the finetuned LLMs perform well in the full-text setting with different prompts, each presenting its own strengths and weaknesses. This establishes a promising baseline for further research. The dataset is available on the website1.
Zhiyuan Tang, Shen Huang, Shidong Shang
ICASSP3
2025 A Structural Model of Truncated Gaussia princeps Luciferase Elucidating the Crucial Catalytic Function of No.76 Arginine towards Coelenterazine Oxidation
abstract
Gaussia Luciferase (GLuc) is a renowned reporter protein that can catalyze the oxidation of coelenterazine (CTZ) and emit a bright light signal. GLuc comprises two consecutive repeats that form the enzyme body and a central putative catalytic cavity. However, deleting the C-terminal repeat only limited reduces the activity (over 30% residual luminescence intensity detectable), despite being a key part of the cavity. How does the remaining GLuc (tGLuc) catalyze CTZ? To address this question, we built a structural model of tGLuc by removing the C-terminal repeat from the resolved structure of intact GLuc, and verified that the cavity-forming component in GLuc remains stable and provides an open-mouth cavity in tGLuc during 500 ns MD simulations in water. Docking simulation and a followed umbrella sampling analysis further revealed that the cavity on tGLuc has a high affinity for CTZ, with a binding energy of up to -114 kJ/mol. Moreover, R76, a validated activity-critical amino acid residue, resides in the cavity and forms a stable hydrogen bond with CTZ. Then, we constructed a cluster model to examine the CTZ oxidation pathway in the cavity using Density Functional Theory (DFT) calculations. The result showed that the pathway consists of four elementary reactions, with the highest Gibbs energy barrier being 65.4 kJ/mol. Both intramolecular electron transfer and the convergence of S1/S0 potential energy surfaces occurred in the last elementary reaction, which was regarded as the reported Chemically-Initiated-Electron-Exchange-Luminescence (CIEEL) reaction. Geometry and wavefunction analysis on the pathway indicated that R76 plays a vital role in CTZ oxidation, which first anchors the environmental oxygen molecule and induces it to form a singlet biradical state, facilitating its attack on CTZ. Subsequently, R76 and the adjacent Q88, positioned near R76 through the tGLuc refolding process, stabilize the transition states and facilitate the emergence of radical electrons on CTZ at the onset of the CIEEL reaction, which contributes to the subsequent intramolecular electron transfer and the production of excited amide product. This study provides a comprehensive explanation of tGLuc's catalytic mechanism. However, it is important to note that these findings are specific to tGLuc and may not extend to other CTZ-based luciferases, particularly those lacking arginine in their catalytic cavities, which likely operate via distinct mechanisms.
Zhi-Chao Xu, Kai-Dong Du, Shen Huang, Naohiro Kobayashi, Yutaka Kuroda, Yan-Hong Bai
PLoS Comput. Biol.4
2024 EcomGPT: Instruction-Tuning Large Language Models with Chain-of-Task Tasks for E-commerce
abstract
Recently, instruction-following Large Language Models (LLMs) , represented by ChatGPT, have exhibited exceptional performance in general Natural Language Processing (NLP) tasks. However, the unique characteristics of E-commerce data pose significant challenges to general LLMs. An LLM tailored specifically for E-commerce scenarios, possessing robust cross-dataset/task generalization capabilities, is a pressing necessity. To solve this issue, in this work, we proposed the first E-commerce instruction dataset EcomInstruct, with a total of 2.5 million instruction data. EcomInstruct scales up the data size and task diversity by constructing atomic tasks with E-commerce basic data types, such as product information, user reviews. Atomic tasks are defined as intermediate tasks implicitly involved in solving a final task, which we also call Chain-of-Task tasks. We developed EcomGPT with different parameter scales by training the backbone model BLOOMZ with the EcomInstruct. Benefiting from the fundamental semantic understanding capabilities acquired from the Chain-of-Task tasks, EcomGPT exhibits excellent zero-shot generalization capabilities. Extensive experiments and human evaluations demonstrate that EcomGPT outperforms ChatGPT in term of cross-dataset/task generalization on E-commerce tasks. The EcomGPT will be public at https://github.com/Alibaba-NLP/EcomGPT.
Yangning Li, Shirong Ma, Xiaobin Wang, Shen Huang, Chengyue Jiang, Hai-Tao Zheng 0002, Pengjun Xie, Fei Huang 0002, Yong Jiang 0005
AAAI4
2024 SeqGPT: An Out-of-the-Box Large Language Model for Open Domain Sequence Understanding
abstract
Large language models (LLMs) have shown impressive abilities for open-domain NLP tasks. However, LLMs are sometimes too footloose for natural language understanding (NLU) tasks which always have restricted output and input format. Their performances on NLU tasks are highly related to prompts or demonstrations and are shown to be poor at performing several representative NLU tasks, such as event extraction and entity typing. To this end, we present SeqGPT, a bilingual (i.e., English and Chinese) open-source autoregressive model specially enhanced for open-domain natural language understanding. We express all NLU tasks with two atomic tasks, which define fixed instructions to restrict the input and output format but still ``open'' for arbitrarily varied label sets. The model is first instruction-tuned with extremely fine-grained labeled data synthesized by ChatGPT and then further fine-tuned by 233 different atomic tasks from 152 datasets across various domains. The experimental results show that SeqGPT has decent classification and extraction ability, and is capable of performing language understanding tasks on unseen domains. We also conduct empirical studies on the scaling of data and model size as well as on the transfer across tasks. Our models are accessible at https://github.com/Alibaba-NLP/SeqGPT.
Tianyu Yu 0002, Chengyue Jiang, Chao Lou, Shen Huang, Xiaobin Wang, Wei Liu 0131, Jiong Cai, Yangning Li, Kewei Tu, Hai-Tao Zheng 0002, Ningyu Zhang 0001, Pengjun Xie, Fei Huang 0002, Yong Jiang 0005
AAAI4
2024 Pinyin Regularization in Error Correction for Chinese Speech Recognition with Large Language Models
Zhiyuan Tang, Shen Huang, Shidong Shang
INTERSPEECH3
2024 Exploring Key Point Analysis with Pairwise Generation and Graph Partitioning
abstract
Xiao Li, Yong Jiang, Shen Huang, Pengjun Xie, Gong Cheng, Fei Huang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Xiao Li 0043, Yong Jiang 0005, Shen Huang, Pengjun Xie, Gong Cheng 0001, Fei Huang 0002
NAACL-HLT3
2024 End-to-End Beam Retrieval for Multi-Hop Question Answering
abstract
Jiahao Zhang, Haiyang Zhang, Dongmei Zhang, Liu Yong, Shen Huang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Dongmei Zhang 0007, Yong Liu 0027, Shen Huang
NAACL-HLT5
2024 Confidence-Aware Sentiment Quantification via Sentiment Perturbation Modeling
abstract
Sentiment Quantification aims to detect the overall sentiment polarity of users from a set of reviews corresponding to a target. Existing methods equally treat and aggregate individual reviews' sentiment to judge the overall sentiment polarity. However, the confidence of each review is not equal in sentiment quantification where sentiment perturbation arising from high- and low-confidence reviews may degrade the accuracy of Sentiment Quantification. Specifically, fake reviews with deceptive sentiments are low confidence, which perturbs the overall sentiment prediction. Whereas, some reviews generated by responsible users are high confidence. They contain authoritative suggestions so they should be emphasized in Sentiment Quantification. In this paper, we design and build COSE, a confidence-aware sentiment quantification framework, which can measure the confidence of individual reviews to eliminate sentiment perturbation and facilitate sentiment quantification. We design a Review Graph that achieves review confidence modeling in an unsupervised manner and obtains review confidence representations. Moreover, we develop a dynamic fusion attention mechanism, which produces sentiment “de-perturbation” vectors to eliminate the sentiment perturbation based on the confidence representations. Extensive experiments on large-scale review datasets validate the significant superiority of COSE over the state-of-the-art.
Xiangyun Tang, Dongliang Liao, Meng Shen 0001, Liehuang Zhu, Shen Huang, Gongfu Li, Hong Man, Jin Xu 0014
IEEE Trans. Affect. Comput.5
2023 Curriculum-Style Fine-Grained Adaption for Unsupervised Cross-Lingual Dependency Transfer
abstract
Unsupervised cross-lingual transfer has been shown great potentials for dependency parsing of the low-resource languages when there is no annotated treebank available. Recently, the self-training method has received increasing interests because of its state-of-the-art performance in this scenario. In this work, we advance the method further by coupling it with curriculum learning, which guides the self-training in an easy-to-hard manner. Concretely, we present a novel metric to measure the instance difficulty of a dependency parser which is trained mainly on a Treebank from a resource-rich source language. By using the metric, we divide a low-resource target language into several fine-grained sub-languages by their difficulties, and then apply iterative-self-training progressively on these sub-languages. To fully explore the auto-parsed training corpus from sub-languages, we exploit an improved parameter generation network to model the sub-languages for better representation learning. Experimental results show that our final curriculum-style self-training can outperform a range of strong baselines, leading to new state-of-the-art results on unsupervised cross-lingual dependency parsing. We also conduct detailed experimental analyses to examine the proposed approach in depth for comprehensive understandings.
Peiming Guo, Shen Huang, Peijie Jiang, Yueheng Sun, Meishan Zhang, Min Zhang 0005
IEEE ACM Trans. Audio Speech Lang. Process.2
2022 Adversarial Sample Detection for Speaker Verification by Neural Vocoders
abstract
Automatic speaker verification (ASV), one of the most important technology for biometric identification, has been widely adopted in security-critical applications. However, ASV is seriously vulnerable to recently emerged adversarial attacks, yet effective counter-measures against them are limited. In this paper, we adopt neural vocoders to spot adversarial samples for ASV. We use the neural vocoder to re-synthesize audio and find that the difference between the ASV scores for the original and re-synthesized audio is a good indicator for discrimination between genuine and adversarial samples. This effort is, to the best of our knowledge, among the first to pursue such a technical direction for detecting time-domain adversarial samples for ASV, and hence there is a lack of established baselines for comparison. Consequently, we implement the Griffin-Lim algorithm as the detection baseline. The proposed approach achieves effective detection performance that outperforms the baselines in all the settings. We also show that the neural vocoder adopted in the detection framework is dataset-independent. Our codes will be made open-source for future works to do fair comparison1.
Po-Chun Hsu, Ji Gao, Shen Huang, Jian Kang 0006, Zhiyong Wu 0001, Helen M. Meng, Hung-yi Lee
ICASSP5
2022 ICPR 2022 Challenge on Multi-Modal Subtitle Recognition
abstract
Video subtitle recognition, as one of the basic elements of video editing, has received increasing attention recently. However, the misaligned between audio and subtitles as well as the costly manual annotation remain a demanding issue toward subsequent intelligent processing. In this paper, we introduce a multi-modal subtitle recognition challenge for ICPR 2022, in which we present a large-scale video dataset (215 hours in total for visual and audio annotations), and 3 tracks including: 1) extracting subtitles in visual modality with audio annotation (ESV); 2) extracting subtitles in audio modality with visual annotations (ESA); and 3) extracting subtitles with both visual and audio annotation (ESVA). The challenge attracts 376 participants, among which the methods of top 3 teams on each track have been elaborated.
Shen Huang, Pengfei Hu 0004, Jian Kang 0006, Weida Liang, Yaqiang Wu, Yong Liu 0027
ICPR2
2022 PM-MMUT: Boosted Phone-mask Data Augmentation using Multi-Modeling Unit Training for Phonetic-Reduction-Robust E2E Speech Recognition
abstract
Consonant and vowel reduction are often encountered in speech, which might cause performance degradation in automatic speech recognition (ASR).Our recently proposed learning strategy based on masking, Phone Masking Training (PMT), alleviates the impact of such phenomenon in Uyghur ASR.Although PMT achieves remarkably improvements, there still exists room for further gains due to the granularity mismatch between the masking unit of PMT (phoneme) and the modeling unit (word-piece).To boost the performance of PMT, we propose multi-modeling unit training (MMUT) architecture fusion with PMT (PM-MMUT).The idea of MMUT framework is to split the Encoder into two parts including acoustic feature sequences to phoneme-level representation (AF-to-PLR) and phoneme-level representation to word-piece-level representation (PLR-to-WPLR).It allows AF-to-PLR to be optimized by an intermediate phoneme-based CTC loss to learn the rich phoneme-level context information brought by PMT.Experimental results on Uyghur ASR show that the proposed approaches outperform obviously the pure PMT.We also conduct experiments on the 960-hour Librispeech benchmark using ES-Pnet1, which achieves about 10% relative WER reduction on all the test set without LM fusion comparing with the latest official ESPnet1 pre-trained model.
Pengfei Hu 0004, Nurmemet Yolwas, Shen Huang
INTERSPEECH4
2022 DAMO-NLP at NLPCC-2022 Task 2: Knowledge Enhanced Robust NER for Speech Entity Linking
Shen Huang, Yuchen Zhai, Xinwei Long, Yong Jiang 0005, Xiaobin Wang, Yin Zhang 0006, Pengjun Xie
NLPCC (2)1
2022 How to Boost Anti-Spoofing with X-Vectors
abstract
With the development of speech synthesis or voice conversion, speech spoofing countermeasures are increasingly required for protecting automatic speaker verification system. In our daily life, if we are familiar with the speaker, we tend to seek her/his traits in our memory to distinguish between bona fide and spoofed speech of her/him. Speaker label can not be directly used to guide the training of anti-spoofing network because it is difficult to obtain in real scenes. Motivated by this, we use x-vectors to represent speaker information and propose two novel methods by introducing x-vectors on the acoustic feature and embedding level into the two mainstream anti-spoofing methods (LightCNN and SeNet). An attention module is also added on the embedding level for further improvement. Experimental results on the ASVspoof 2019 logical access (LA) database show that the best EER and mintDCF in our methods are 0.98% and 0.0294, outperforming state-of-the-art single systems as far as we know.
Shen Huang, Ji Gao, Ying Hu 0005, Liang He 0003
SLT3
2021 Stacked Acoustic-and-Textual Encoding: Integrating the Pre-trained Models into Speech Translation Encoders
abstract
Chen Xu, Bojie Hu, Yanyang Li, Yuhao Zhang, Shen Huang, Qi Ju, Tong Xiao, Jingbo Zhu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Chen Xu 0008, Bojie Hu, Yanyang Li, Shen Huang, Qi Ju 0002, Tong Xiao 0001
ACL/IJCNLP (1)5
2021 Improved Lightcnn with Attention Modules for Asv Spoofing Detection
abstract
With the advent of many state-of-the-art speech synthesis or conversion techniques, speaker recognition system is confronted with the imperceptible interference caused by spoofed speech. In this paper, we realize that filter bank distributions of cepstral features in the frequency domain cause considerable influence to the system performance through experiments. Therefore, given attention mechanism may help to focus on key information, this paper presents improved light convolutional neural network (LCNN) with attention modules, separately named Squeeze-and-Excitation (SE) block, Convolutional Block Attention Module (CBAM) and Dual Attention Network (DANet). To our knowledge, we are the first to do systematic study of attention mechanisms in the field of speech anti-spoofing. Experimented on ASVspoof 2019 dataset, our proposed single system can get 22% min-tDCF reduction as well as 42% EER reduction over LCNN baseline and its performance can even rank fourth among all participating teams using fusion systems in ASVspoof 2019 competition.
Tianyu Liang, Shen Huang, Liang He 0003
ICME4
2021 Leveraging Phone Mask Training for Phonetic-Reduction-Robust E2E Uyghur Speech Recognition
abstract
In Uyghur speech, consonant and vowel reduction are often encountered, especially in spontaneous speech with high speech rate, which will cause a degradation of speech recognition performance. To solve this problem, we propose an effective phone mask training method for Conformer-based Uyghur end-to-end (E2E) speech recognition. The idea is to randomly mask off a certain percentage features of phones during model training, which simulates the above verbal phenomena and facilitates E2E model to learn more contextual information. According to experiments, the above issues can be greatly alleviated. In addition, deep investigations are carried out into different units in masking, which shows the effectiveness of our proposed masking unit. We also further study the masking method and optimize filling strategy of phone mask. Finally, compared with Conformer-based E2E baseline without mask training, our model demonstrates about 5.51% relative Word Error Rate (WER) reduction on reading speech and 12.92% on spontaneous speech, respectively. The above approach has also been verified on test-set of open-source data THUYG-20, which shows 20% relative improvements.
Pengfei Hu 0004, Jian Kang 0006, Shen Huang
Interspeech4
2021 The TNT Team System Descriptions of Cantonese and Mongolian for IARPA OpenASR20
Zhiqiang Lv, Ambyer Han, Guan-Bo Wang, Gui-Xin Shi, Jian Kang 0006, Jinghao Yan, Pengfei Hu 0004, Shen Huang, Weiqiang Zhang 0001
Interspeech9
2021 Residual attention-based multi-scale script identification in scene text images
Mengkai Ma, Qiufeng Wang 0001, Shen Huang, John Yannis Goulermas, Kaizhu Huang
Neurocomputing4
2020 Dynamic Curriculum Learning for Low-Resource Neural Machine Translation
abstract
Large amounts of data has made neural machine translation (NMT) a big success in recent years.But it is still a challenge if we train these models on small-scale corpora.In this case, the way of using data appears to be more important.Here, we investigate the effective use of training data for low-resource NMT.In particular, we propose a dynamic curriculum learning (DCL) method to reorder training samples in training.Unlike previous work, we do not use a static scoring function for reordering.Instead, the order of training samples is dynamically determined in two ways -loss decline and model competence.This eases training by highlighting easy samples that the current model has enough competence to learn.We test our DCL method in a Transformerbased system.Experimental results show that DCL outperforms several strong baselines on three low-resource machine translation benchmarks and different sized data of WMT'16 En-De.
Chen Xu 0008, Bojie Hu, Yufan Jiang, Zeyang Wang, Shen Huang, Qi Ju 0002, Tong Xiao 0001
COLING6
2020 CSP: Code-Switching Pre-training for Neural Machine Translation
abstract
This paper proposes a new pre-training method, called Code-Switching Pre-training (CSP for short) for Neural Machine Translation (NMT).Unlike traditional pre-training method which randomly masks some fragments of the input sentence, the proposed CSP randomly replaces some words in the source sentence with their translation words in the target language.Specifically, we firstly perform lexicon induction with unsupervised word embedding mapping between the source and target languages, and then randomly replace some words in the input sentence with their translation words according to the extracted translation lexicons.CSP adopts the encoderdecoder framework: its encoder takes the codemixed sentence as input, and its decoder predicts the replaced fragment of the input sentence.In this way, CSP is able to pre-train the NMT model by explicitly making the most of the cross-lingual alignment information extracted from the source and target monolingual corpus.Additionally, we relieve the pretrainfinetune discrepancy caused by the artificial symbols like [mask].To verify the effectiveness of the proposed method, we conduct extensive experiments on unsupervised and supervised NMT.Experimental results show that CSP achieves significant improvements over baselines without pre-training or with other pre-training methods.
Bojie Hu, Ambyer Han, Shen Huang, Qi Ju 0002
EMNLP (1)4
2020 Dominate and Non-dominate Hand Prediction for Handheld Touchscreen Interaction
abstract
People have their individual preference of which hand they preferentially use to do certain things. It is not unusual to see them use mobile devices with a touchscreen one-handedly. Depending on where they are and what they do, people may use one hand over the other to hold and interact with mobile devices. Few studies have looked into the implication of using a preferred hand versus a non-preferred in touchscreen interaction on mobile devices. As the screen size increases, the difference between using a preferred hand and a non-preferred hand on the touchscreen becomes more significant. In this paper, we show how to extract features from 3 different interaction gestures on touchscreen, tap, swipe, and drag to learn if a user is using the dominant hand or the non-dominant hand. We compare the performance of using different sets of features in prediction by considering the constraints of handheld devices. A random forest-based prediction system is also created and enhanced to recognize if the user is using a preferred hand or a non-preferred hand. This technique enables the user interface of a touchscreen to adapt to which hand the user hold and interact with mobile devices.
Shen Huang
HSI2
2020 Cognitive Representation Learning of Self-Media Online Article Quality
abstract
The automatic quality assessment of self-media online articles is an urgent and new issue, which is of great value to the online recommendation and search. Different from traditional and well-formed articles, self-media online articles are mainly created by users, which have the appearance characteristics of different text levels and multi-modal hybrid editing, along with the potential characteristics of diverse content, different styles, large semantic spans and good interactive experience requirements. To solve these challenges, we establish a joint model CoQAN in combination with the layout organization, writing characteristics and text semantics, designing different representation learning subnetworks, especially for the feature learning process and interactive reading habits on mobile terminals. It is more consistent with the cognitive style of expressing an expert's evaluation of articles. We have also constructed a large scale real-world assessment dataset. Extensive experimental results show that the proposed framework significantly outperforms state-of-the-art methods, and effectively learns and integrates different factors of the online article quality assessment.
Shen Huang, Gongfu Li, Qiang Deng, Dongliang Liao, Pengda Si, Yujiu Yang 0001
ACM Multimedia2
2020 Detection of High-Risk Depression Groups Based on Eye-Tracking Data
Simeng Lu, Shen Huang, Xiujuan Zheng, Danmin Miao, Zheru Chi
PRCV (2)2
2019 Verifying Deep Keyword Spotting Detection with Acoustic Word Embeddings
abstract
In this paper, in order to improve keyword spotting (KWS) performance in a live broadcast scenario, we propose to use a template matching method based on acoustic word embeddings (AWE) as the second stage to verify the detection from the Deep KWS system. AWEs are obtained via a deep bidirectional long short-term memory (BLSTM) network trained using limited positive and negative keyword candidates, which aims to encode variable-length keyword candidates into fixed-dimensional vectors with reasonable discriminative ability. Learning AWEs takes a combination of three specifically-designed losses: the triplet and reversed triplet losses try to keep same keyword candidates closer and different keyword candidates farther, while the hinge loss is to set a fixed threshold to distinguish all positive and negative keyword candidates. During keyword verification, calibration scores are used to reduce the bias between different templates for different keyword candidates. Experiments show that adding AWE-based keyword verification to Deep KWS achieves 5.6% relative accuracy improvement; the hinge loss brings additional 5.5% relative gain and the final accuracy climbs to 0.775 by using calibration scores.
Yougen Yuan, Zhiqiang Lv, Shen Huang, Lei Xie 0001
ASRU3
2019 Utterance-level End-to-end Language Identification Using Attention-based CNN-BLSTM
abstract
In this paper, we present an end-to-end language identification framework, the attention-based Convolutional Neural Network-Bidirectional Long-short Term Memory (CNN-BLSTM). The model is performed on the utterance level, which means the utterance-level decision can be directly obtained from the output of the neural network. To handle speech utterances with entire arbitrary and potentially long duration, we combine CNN-BLSTM model with a self-attentive pooling layer together. The front-end CNN-BLSTM module plays a role as local pattern extractor for the variable-length inputs, and the following self-attentive pooling layer is built on top to get the fixed-dimensional utterance-level representation. We conducted experiments on NIST LRE07 closed-set task, and the results reveal that the proposed attention-based CNN-BLSTM model achieves comparable error reduction with other state-of-the-art utterance-level neural network approaches for all 3 seconds, 10 seconds, 30 seconds duration tasks.
Weicheng Cai, Danwei Cai, Shen Huang, Ming Li 0026
ICASSP3
2019 Text Siamese Network for Video Textual Keyframe Detection
abstract
In this paper, we propose a novel approach of video text keyframe detection, to achieve the goal of representing the video with textual keyframes and reducing the waste of resources in video review. Different from the works of video summarization which mainly focus on the variation of the scenes in videos, we pay attention to the variances of the text between sequential frames. For the above purpose, Text Siamese Network (TSN) is developed to automatically detect the keyframes which contain text in videos. Specifically, the TSN is composed of two branches, text similarity measurement and text identification. The first branch is utilized to evaluate the similarity between consecutive frames. Furthermore, an attention block is used to select the informative features to identify whether a frame contains text or not in the second branch. Additionally, a new dataset called VTKD2019 is proposed for video text keyframe detection. VTKD2019 contains 571 videos and is spit into three levels (easy, medium and hard) for evaluation. The experimental results on the VTKD2019 and ICDAR2015 demonstrate the effectiveness of our method.
Hongzhen Wang, Pei Xu 0006, Shen Huang, Qi Ju 0002
ICDAR5
2019 A Multi-oriented Chinese Keyword Spotter Guided by Text Line Detection
abstract
Chinese keyword spotting is a challenging task as there is no visual blank for Chinese words. Different from English words which are split naturally by visual blanks, Chinese words are generally split only by semantic information. In this paper, we propose a new Chinese keyword spotter for natural images, which is inspired by Mask R-CNN. We propose to predict the keyword masks guided by text line detection. Firstly, proposals of text lines are generated by Faster R-CNN; Then, text line masks and keyword masks are predicted by segmentation in the proposals. In this way, the text lines and keywords are predicted in parallel. We create two Chinese keyword datasets based on RCTW-17 and ICPR MTWI2018 to verify the effectiveness of our method.
Pei Xu 0006, Hongzhen Wang, Shen Huang, Qi Ju 0002
ICDAR5
2019 Multimedia Simultaneous Translation System for Minority Language Communication with Mandarin
Shen Huang, Bojie Hu, Pengfei Hu 0004, Jian Kang 0006, Zhiqiang Lv, Jinghao Yan, Qi Ju 0002, Shiyin Kang, Deyi Tuo, Guangzhi Li, Nurmemet Yolwas
INTERSPEECH1
2017 Addressing Domain Adaptation for Chinese Word Segmentation with Global Recurrent Structure
abstract
Boundary features are widely used in traditional Chinese Word Segmentation (CWS) methods as they can utilize unlabeled data to help improve the Out-of-Vocabulary (OOV) word recognition performance. Although various neural network methods for CWS have achieved performance competitive with state-of-the-art systems, these methods, constrained by the domain and size of the training corpus, do not work well in domain adaptation. In this paper, we propose a novel BLSTM-based neural network model which incorporates a global recurrent structure designed for modeling boundary features dynamically. Experiments show that the proposed structure can effectively boost the performance of Chinese Word Segmentation, especially OOV-Recall, which brings benefits to domain adaptation. We achieved state-of-the-art results on 6 domains of CNKI articles, and competitive results to the best reported on the 4 domains of SIGHAN Bakeoff 2010 data.
Shen Huang, Xu Sun 0001, Houfeng Wang
IJCNLP(1)1
2011 Exploring nuisance attribute projection and score normalization for GLDS-SVM based automatic mispronunciation detection method
abstract
In the task of mispronunciation detection, the cross-speaker degradation and some other confusing nuisances are the challenging problems demanding prompt solution. In this paper, we will attempt to remove the non-pronunciation variations in the GLDS-SVM expansion space by using nuisance attribute projection strategy, in order to increase the separating capacity between different phoneme instances. Moreover, different kinds of score normalization methods with softmax, posterior probability vector (PPV), Z-norm and T-norm are comparatively discussed. The experiments on three kinds of speech corpora demonstrate the effectiveness of the above methods, and the performance improvement is not very significant, but sustainable.
Hongyan Li 0010, Shen Huang, Shijin Wang 0001, Jiaen Liang, Bo Xu 0002
ICASSP2
2011 Context-Dependent Duration Modeling with Backoff Strategy and Look-Up Tables for Pronunciation Assessment and Mispronunciation Detection
Hongyan Li 0010, Shen Huang, Shijin Wang 0001, Bo Xu 0002
INTERSPEECH2
2010 Automatic reference independent evaluation of prosody quality using multiple knowledge fusions
Shen Huang, Hongyan Li 0010, Shijin Wang 0001, Jiaen Liang, Bo Xu 0002
INTERSPEECH1
2010 Exploring goodness of prosody by diverse matching templates
abstract
In automatic speech grading systems, rare research is followed through addressing the issue of GOR (Goodness Of pRosody). In this paper we propose a novel method by taking the advantage of our QBH (Query By Humming) techniques in 2008 MIREX evaluation task. A set of standard samples related to the top-cream students are initially picked up as templates, a cascade QBH structure is then taken from two metrics: the MOMEL stylization followed by DTW distance; the Fujisaki model followed by EMD distance. Sentence GOR is obtained by the fused confidence between target and each template, and forms a weighted sum as the goodness in the passage level. Experiment results indicate that performance increases with the count of template, and Fujisaki-EMD metric outperforms MOMEL-DTW one in terms of correlation. Their combination can be treated as template based GOR score, compensated with our previous feature based GOR score, the approach can achieve 0.432 in correlation and 17.90% in EER in our corpus. Index Terms: speech prosody, query by humming
Shen Huang, Hongyan Li 0010, Shijin Wang 0001, Jiaen Liang, Bo Xu 0002
INTERSPEECH1
2009 Improving product review search experiences on general search engines
abstract
In the Web 2.0 era, internet users contribute a large amount of online content. Product review is a good example. Since these phenomena are distributed all over shopping sites, weblogs, forums etc., most people have to rely on general search engines to discover and digest others' comments. While conventional search engines work well in many situations, it's not sufficient for users to gather such information. The reasons include but are not limited to: 1) the ranking strategy does not incorporate product reviews' inherent characteristics, e.g., sentiment orientation; 2) the snippets are neither indicative nor descriptive of user opinions. In this paper, we propose a feasible solution to enhance the experience of product review search. Based on this approach, a system named "Improved Product Review Search (IPRS)" is implemented on the ground of a general search engine. Given a query on a product, our system is capable of: 1) automatically identifying user opinion segments in a whole article; 2) ranking opinions by incorporating both the sentiment orientation and the topics expressed in reviews; 3) generating readable review snippets to indicate user sentiment orientations; 4) easily comparing products based on a visualization of opinions. Both results of a usability study and an automatic evaluation show that our system is able to assist users quickly understand the product reviews within limited time.
Shen Huang, Catherine Baudin
ICEC1
2009 Discovering clues for review quality from author's behaviors on e-commerce sites
abstract
With the number of online reviews growing rapidly, it is increasingly difficult to digest all the information within limited time. To help users efficiently get concise information about a product, researchers have studied algorithms for automated opinion summarization. However, users might expect to further read detailed high-quality reviews in addition to a review outline. This raises another interesting problem not well studied yet: how to discover high quality product reviews? Previous research examined various properties of a product review to predict its quality. In this paper, we further explore this topic by incorporating another information resource: the behavior of review authors in an e-commerce community. First, we perform a high-level analysis on two kinds of data: product reviews and deal transactions. According to the results of this analysis, three features, including personal reputation, seller degree and expertise degree, are studied to assess the quality of a review from a credibility and expertise perspective. Our analysis shows that these features are strongly related to review quality and that they can help uncover review spamming by sellers. Furthermore, we propose a simulation model based on the above findings. The model is able to generate the basic properties of the review community, especially when the above three features are taken into account.
Shen Huang, Catherine Baudin
ICEC1
2009 Context Dependent Feature Based Bottom-up Rescoring SVM Classifier in Children's English Stress Mis-pronunciation Detection
abstract
Automatic assessment of word stress error is an integral part for oral language grading system. However, problems that the property of vowels depends on its context information and the data sparseness of different vowel class are yet to be solved. This paper shall briefly introduce a hybrid method consisting of both traditional prosodic features and proposed context dependent strategies. In classification word stress is determined by weighting a bottom-up fashioned group tree with modified distributed probability score. In experiment, the overall equal error rate of our proposed system achieves 9.41%, which exhibits relative reduction and its competence of use in stress error detection system.
Shen Huang, Hongyan Li 0010, Shijin Wang 0001, Jiaen Liang, Bo Xu 0002
ICALT1
2009 High performance automatic mispronunciation detection method based on neural network and TRAP features
Hongyan Li 0010, Shijin Wang 0001, Jiaen Liang, Shen Huang, Bo Xu 0002
INTERSPEECH4
2008 Query by humming via multiscale transportation distance in random query occurrence context
abstract
Query by humming (QBH) is an interactive tool for retrieving favored songs from a large database of known media via acoustic input. In this task, common method for measuring similarity between query and candidate is either by symbolic notation distance or by framed based dynamic programming. However, the former has disadvantage of error-prone to the noted symbolic feature extraction stage, while the latter is time-consuming. It has been proved that transportation distance has its remarkable merit in image query. However, we adopt a new structure for handling QBH, which is based on an improved version of this measure in combination with a string searching algorithm. More practically, we extend this method in random piece context, which means users can hum at any part of music piece. Experimental results are evaluated in MIREX 2007. Final 92.82% MRR has shown its significant advantages.
Shen Huang, Lei Wang 0062, Hongchen Jiang, Bo Xu 0002
ICME1
2008 Improving searching speed and accuracy of query by humming system based on three methods: feature fusion, candidates set reduction and multiple similarity measurement rescoring
Lei Wang 0062, Shen Huang, Jiaen Liang, Bo Xu 0002
INTERSPEECH2
2006 Subjectivity Categorization of Weblog with Part-of-Speech Based Smoothing
abstract
Experts from different domains try to mine users' comments on Weblogs for different reasons such as politics or commerce. All these needs necessitate automatically distinguishing subjective Weblog contents from objective ones, namely subjectivity categorization. Since Weblogs contain various topics from different domains, limited training data can hardly cover all the topics and "unseen words" becomes a serious problem for categorization tasks. In this paper, part-of-speech (POS) based smoothing is proposed to alleviate the "unseen words" problem. In conjunction with a naive Bayes model constructed from limited training data, the probability of an unseen word in a new domain can be well smoothed by the probability of its POS result. Empirical studies on five datasets show that our approach consistently outperforms the basic naive Bayes with Laplace smoothing. In a cross-domain experiment, our approach achieves 22.0% improvement in Macro Fl and 24.4% in Micro Fl over basic naive Bayes. These verify that POS based smoothing can indeed benefit subjectivity categorization, especially in the cases with a large number of unseen words.
Shen Huang, Jian-Tao Sun, Xuanhui Wang, Hua-Jun Zeng, Zheng Chen 0001
ICDM1
2006 Salient Phrases-based Clustering and Ranking in Chinese Bulletin Board System
Xiaoyuan Wu, Shen Huang, Yong Yu 0001
SEKE2
2006 Multitype Features Coselection for Web Document Clustering
abstract
Feature selection has been widely applied in text categorization and clustering. Compared to unsupervised selection, supervised feature selection is more successful in filtering out noise in most cases. However, due to a lack of label information, clustering can hardly exploit supervised selection. Some studies have proposed to solve this problem by "pseudoclass." As empirical results show, this method is sensitive to selection criteria and data sets. In this paper, we propose a novel feature coselection for Web document clustering, which is called multitype features coselection for clustering (MFCC). MFCC uses intermediate clustering results in one type of feature space to help the selection in other types of feature spaces. Our experiments show that for most selection criteria, MFCC reduces effectively the noise introduced by "pseudoclass," and further improves clustering performance.
Shen Huang, Zheng Chen 0001, Yong Yu 0001, Wei-Ying Ma
IEEE Trans. Knowl. Data Eng.1
2006 TSSP: Multi-features based reinforcement algorithm to find related papers
Shen Huang, Yong Yu 0001, Gui-Rong Xue, Benyu Zhang, Zheng Chen 0001, Wei-Ying Ma
Web Intell. Agent Syst.1
2005 Block-Based Language Modeling Approach Towards Web Search
Shengping Li, Shen Huang, Gui-Rong Xue, Yong Yu 0001
APWeb2
2005 Interactive Chinese Search Results Clustering for Personalization
Gui-Rong Xue, Shen Huang, Yong Yu 0001
WAIM3
2004 DHT Based Searching Improved by Sliding Window
Shen Huang, Gui-Rong Xue, Yan-Feng Ge, Yong Yu 0001
WAIM1
2004 TSSP: A Reinforcement Algorithm to Find Related Papers
abstract
Content analysis and citation analysis are two common methods in recommending system. Compared with content analysis, citation analysis can discover more implicitly related papers. However, the citation-based methods may introduce more noise in citation graph and cause topic drift. Some work combine content with citation to improve similarity measurement. The problem is that the two features are not used to reinforce each other to get better result. To solve the problem, we propose a new algorithm, Topic Sensitive Similarity Propagation (TSSP), to effectively integrate content similarity into similarity propagation. TSSP has two parts: citation context based propagation and iterative reinforcement. First, citation contexts provide clues for which papers are topic related to and filter out less irrelevant citations. Second, iteratively integrating content and citation similarity enable them to reinforce each other during the propagation. The experimental results of a user study show TSSP outperforms other algorithms in almost all cases.
Shen Huang, Gui-Rong Xue, Benyu Zhang, Zheng Chen 0001, Yong Yu 0001, Wei-Ying Ma
Web Intelligence1
2004 Multi-type Features Based Web Document Clustering
Shen Huang, Gui-Rong Xue, Benyu Zhang, Zheng Chen 0001, Yong Yu 0001, Wei-Ying Ma
WISE1
2004 Optimizing Web Search Using Spreading Activation on the Clickthrough Data
Gui-Rong Xue, Shen Huang, Yong Yu 0001, Hua-Jun Zeng, Zheng Chen 0001, Wei-Ying Ma
WISE2