VLDB 2026 Research / reviewers in the wild / expert
Baoxing Huai
dblp:152/3689
· DBLP profile ↗
40ranked-venue papers
1as first author
32since 2021 · last 2025
0000-0001-9625-2314ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 13 since 2021Databases, data management, data science and information retrieval · 11 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Alignment-Augmented Speculative Decoding with Alignment Sampling and Conditional VerificationabstractRecent works have revealed the great potential of speculative decoding in accelerating the autoregressive generation process of large language models.The success of these methods relies on the alignment between draft candidates and the sampled outputs of the target model.Existing methods mainly achieve draft-target alignment with training-based methods, e.g., EAGLE, Medusa, involving considerable training costs.In this paper, we present a trainingfree alignment-augmented speculative decoding algorithm.We propose alignment sampling, which leverages output distribution obtained in the prefilling phase to provide more aligned draft candidates.To further benefit from highquality but non-aligned draft candidates, we also introduce a simple yet effective flexible verification strategy.Through an adaptive probability threshold, our approach can improve generation accuracy while further improving inference efficiency.Experiments on 8 datasets (including question answering, summarization and code completion tasks) show that our approach increases the average generation score by 3.3 points for the LLaMA3 model.Our method achieves a mean acceptance length up to 2.39 and speed up generation by 2.23×. Zhenxu Tian, Juntao Li 0005, Qingrong Xia, Xinyu Duan, Zhefeng Wang 0001, Baoxing Huai, Min Zhang 0005 |
EMNLP | 7 |
| 2025 | IMPRESS: An Importance-Informed Multi-Tier Prefix KV Storage System for Large Language Model Inference
Weijian Chen 0002, Shuibing He, Haoyang Qu, Siling Yang, Baoxing Huai, Gang Chen 0001 |
FAST | 8 |
| 2025 | Taming the Titans: A Survey of Efficient LLM Inference ServingabstractLarge Language Models (LLMs) for Generative AI have achieved remarkable progress, evolving into sophisticated and versatile tools widely adopted across various domains and applications. However, the substantial memory overhead caused by their vast number of parameters, combined with the high computational demands of the attention mechanism, poses significant challenges in achieving low latency and high throughput for LLM inference services. Recent advancements, driven by groundbreaking research, have significantly accelerated progress in this field. This paper provides a comprehensive survey of these methods, covering fundamental instance-level approaches, in-depth cluster-level strategies, and emerging scenarios. At the instance level, we review model placement, request scheduling, decoding length prediction, storage management, and the disaggregation paradigm. At the cluster level, we explore GPU cluster deployment, multi-instance load balancing, and cloud service solutions. Additionally, we discuss specific tasks, modules, and auxiliary methods in emerging scenarios. Finally, we outline potential research directions to further advance the field of LLM inference serving. Ranran Zhen, Juntao Li 0005, Yixin Ji, Zhenlin Yang, Qingrong Xia, Xinyu Duan, Zhefeng Wang 0001, Baoxing Huai, Min Zhang 0005 |
INLG | 9 |
| 2024 | StyleSinger: Style Transfer for Out-of-Domain Singing Voice SynthesisabstractStyle transfer for out-of-domain (OOD) singing voice synthesis (SVS) focuses on generating high-quality singing voices with unseen styles (such as timbre, emotion, pronunciation, and articulation skills) derived from reference singing voice samples. However, the endeavor to model the intricate nuances of singing voice styles is an arduous task, as singing voices possess a remarkable degree of expressiveness. Moreover, existing SVS methods encounter a decline in the quality of synthesized singing voices in OOD scenarios, as they rest upon the assumption that the target vocal attributes are discernible during the training phase. To overcome these challenges, we propose StyleSinger, the first singing voice synthesis model for zero-shot style transfer of out-of-domain reference singing voice samples. StyleSinger incorporates two critical approaches for enhanced effectiveness: 1) the Residual Style Adaptor (RSA) which employs a residual quantization module to capture diverse style characteristics in singing voices, and 2) the Uncertainty Modeling Layer Normalization (UMLN) to perturb the style attributes within the content representation during the training phase and thus improve the model generalization. Our extensive evaluations in zero-shot style transfer undeniably establish that StyleSinger outperforms baseline models in both audio quality and similarity to the reference singing voice samples. Access to singing voice samples can be found at https://stylesinger.github.io/. Yu Zhang 0126, Rongjie Huang 0001, Ruiqi Li 0002, Jinzheng He, Yan Xia 0006, Feiyang Chen 0001, Xinyu Duan, Baoxing Huai, Zhou Zhao 0001 |
AAAI | 8 |
| 2024 | CopyNE: Better Contextual ASR by Copying Named EntitiesabstractEnd-to-end automatic speech recognition (ASR) systems have made significant progress in general scenarios.However, it remains challenging to transcribe contextual named entities (NEs) in the contextual ASR scenario.Previous approaches have attempted to address this by utilizing the NE dictionary.These approaches treat entities as individual tokens and generate them token-by-token, which may result in incomplete transcriptions of entities.In this paper, we treat entities as indivisible wholes and introduce the idea of copying into ASR.We design a systematic mechanism called CopyNE, which can copy entities from the NE dictionary.By copying all tokens of an entity at once, we can reduce errors during entity transcription, ensuring the completeness of the entity.Experiments demonstrate that CopyNE consistently improves the accuracy of transcribing entities compared to previous approaches.Even when based on the strong Whisper, CopyNE still achieves notable improvements. Shilin Zhou 0002, Zhenghua Li, Yu Hong 0001, Min Zhang 0005, Zhefeng Wang 0001, Baoxing Huai |
ACL (1) | 6 |
| 2024 | When and How to Grow? On Efficient Pre-training via Model Growth
Juntao Li 0005, Min Zhang 0005, Zechang Li, Qingrong Xia, Xinyu Duan, Zhefeng Wang 0001, Baoxing Huai |
ACML | 8 |
| 2024 | Improving Chinese Named Entity Recognition with Multi-grained Words and Part-of-Speech Tags via Joint ModelingabstractNowadays, character-based sequence labeling becomes the mainstream Chinese named entity recognition (CNER) approach, instead of word-based methods, since the latter degrades performance due to propagation of word segmentation (WS) errors. To make use of WS information, previous studies usually learn CNER and WS simultaneously with multi-task learning (MTL) framework, or treat WS information as extra guide features for CNER model, in which the utilization of WS information is indirect and shallow. In light of the complementary information inside multi-grained words, and the close connection between named entities and part-of-speech (POS) tags, this work proposes a tree parsing approach for joint modeling CNER, multi-grained word segmentation (MWS) and POS tagging tasks simultaneously. Specifically, we first propose a unified tree representation for MWS, POS tagging, and CNER.Then, we automatically construct the MWS-POS-NER data based on the unified tree representation for model training. Finally, we present a two-stage joint tree parsing framework. Experimental results on OntoNotes4 and OntoNotes5 show that our proposed approach of jointly modeling CNER with MWS and POS tagging achieves better or comparable performance with latest methods. Chenhui Dou, Chen Gong 0004, Zhenghua Li, Zhefeng Wang 0001, Baoxing Huai, Min Zhang 0005 |
LREC/COLING | 5 |
| 2024 | TextrolSpeech: A Text Style Control Speech Corpus with Codec Language Text-to-Speech ModelsabstractRecently, there has been a growing interest in the field of controllable Text-to-Speech (TTS). While previous studies have relied on users providing specific style factor values based on acoustic knowledge or selecting reference speeches that meet certain requirements, generating speech solely from natural text prompts has emerged as a new challenge for researchers. This challenge arises due to the scarcity of high-quality speech datasets with natural text style prompt and the absence of advanced text-controllable TTS models. In light of this, 1) we propose TextrolSpeech, which is the first large-scale speech emotion dataset annotated with rich text attributes. The dataset comprises 236,203 pairs of style prompt in natural text descriptions with five style factors and corresponding speech samples. Through iterative experimentation, we introduce a multi-stage prompt programming approach that effectively utilizes the GPT model for generating natural style descriptions in large volumes. 2) Furthermore, to address the need for generating audio with greater style diversity, we propose an efficient architecture called Salle. This architecture treats text controllable TTS as a language model task, utilizing audio codec codes as an intermediate representation to replace the conventional mel-spectrogram. Finally, we successfully demonstrate the ability of the proposed model by showing a comparable performance in the controllable TTS task. Audio samples are available on the demo page https://sall-e.github.io/. Shengpeng Ji, Jialong Zuo, Minghui Fang 0002, Ziyue Jiang 0004, Feiyang Chen 0001, Xinyu Duan, Baoxing Huai, Zhou Zhao 0001 |
ICASSP | 7 |
| 2024 | MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis
Qian Yang 0006, Jialong Zuo, Ziyue Jiang 0001, Zhou Zhao 0001, Feiyang Chen 0001, Zhefeng Wang 0001, Baoxing Huai |
INTERSPEECH | 9 |
| 2024 | Poisoning for Debiasing: Fair Recognition via Eliminating Bias Uncovered in Data PoisoningabstractNeural networks often tend to rely on bias features that have strong but spurious correlations with the target labels for decision-making, leading to poor performance on data that does not adhere to these correlations. Early debiasing methods typically construct an unbiased optimization objective based on the labels of bias features. Recent work assumes that bias label is unavailable and usually trains two models: a biased model to deliberately learn bias features for exposing data bias, and a target model to eliminate bias captured by the bias model. In this paper, we first reveal that previous biased models fit target labels, which resulted in failing to expose data bias. To tackle this issue, we propose poisoner, which utilizes data poisoning to embed the biases learned by biased models into the poisoned training data, thereby encouraging the models to learn more biases. Specifically, we couple data poisoning and model training to continuously prompt the biased model to learn more bias. By utilizing the biased model, we can identify samples in the data that contradict these biased correlations. Subsequently, we amplify the influence of these samples in the training of the target model to prevent the model from learning such biased correlations. Experiments show the superior debiasing performance of our method. Yi Zhang 0101, Zhefeng Wang 0001, Rui Hu 0011, Xinyu Duan, Yi Zheng 0007, Baoxing Huai, Jiarun Han, Jitao Sang 0001 |
ACM Multimedia | 6 |
| 2024 | Light POI-Guided Conversational Recommender System based on Adaptive SpaceabstractConversational Recommender Systems (CRS) have recently attracted significant attention. Despite existing GNN-based CRS methods have been proven to be effective in exploiting knowledge graphs (KGs), we note that these methods are not suitable for modeling scenarios with geographic positional information, which encompass the two key issues that have not been adequately solved: 1) Data noise is ubiquitous in the real world due to a variety of factors, and existing methods are prone to amplifying data noise, which can lead to a deterioration in downstream tasks; 2) Existing CRS models are designed solely in Euclidean space without considering space curvature, which implies that they may suffer from significant distortion when representing real-world graph structures, leading to a decrease in the accuracy and reliability of geographical POIs. To this end, we propose a Light POI-Guided Conversational Recommender based on Adaptive Space, namely PCRA, aiming to address the above problems by enhancing both embedding spaces and graph structures. Specifically, PCRA introduces the unified space to obtain high-quality embeddings compatible with hyperbolic space, Euclidean space, and spherical space. On the other hand, to extract the most valuable neighbors, we adopt a graph denoising module to eliminate noisy entities and ensure light information propagation. Finally, we further fuse the embeddings of utterances and entities to bridge the semantic gap of recommendation and conversation. Extensive experiments on MultiWOZ 2.0 and MultiWOZ 2.1 datasets demonstrate that our proposed PCRA has a significant improvement over the state-of-the-art CRS methods. Yiqi Tong, Yuxin Ying, Fuzhen Zhuang, Baoxing Huai |
SDM | 7 |
| 2024 | Multimodal Dialogue Systems via Capturing Context-aware Dependencies and Ordinal Information of Semantic ElementsabstractThe topic of multimodal conversation systems has recently garnered significant attention across various industries, including travel and retail, among others. While pioneering works in this field have shown promising performance, they often focus solely on context information at the utterance level, overlooking the context-aware dependencies of multimodal semantic elements like words and images. Furthermore, the ordinal information of images, which indicates the relevance between visual context and users’ demands, remains underutilized during the integration of visual content. Additionally, the exploration of how to effectively utilize corresponding attributes provided by users when searching for desired products is still largely unexplored. To address these challenges, we propose PMATE, a P osition-aware M ultimodal di A logue system with seman T ic E lements. Specifically, to obtain semantic representations at the element level, we first unfold the multimodal historical utterances and devise a position-aware multimodal element-level encoder. This component considers all images that may be relevant to the current turn and introduces a novel position-aware image selector to choose related images before fusing the information from the two modalities. Finally, we present a knowledge-aware two-stage decoder and an attribute-enhanced image searcher for the tasks of generating textual responses and selecting image responses, respectively. We extensively evaluate our model on two large-scale multimodal dialogue datasets, and the results of our experiments demonstrate that our approach outperforms several baseline methods. Weidong He, Zhi Li 0057, Hao Wang 0076, Tong Xu 0001, Zhefeng Wang 0001, Baoxing Huai, Nicholas Jing Yuan, Enhong Chen |
ACM Trans. Intell. Syst. Technol. | 6 |
| 2024 | A Survey on Arabic Named Entity Recognition: Past, Recent Advances, and Future TrendsabstractAs more and more Arabic texts emerged on the Internet, extracting important information from these Arabic texts is especially useful. As a fundamental technology, Named entity recognition (NER) serves as the core component in information extraction technology, while also playing a critical role in many other Natural Language Processing (NLP) systems, such as question answering and knowledge graph building. In this paper, we provide a comprehensive review of the development of Arabic NER, especially the recent advances in deep learning and pre-trained language model. Specifically, we first introduce the background of Arabic NER, including the characteristics of Arabic and existing resources for Arabic NER. Then, we systematically review the development of Arabic NER methods. Traditional Arabic NER systems focus on feature engineering and designing domain-specific rules. In recent years, deep learning methods achieve significant progress by representing texts via continuous vector representations. With the growth of pre-trained language model, Arabic NER yields better performance. Finally, we conclude the method gap between Arabic NER and NER methods from other languages, which helps outline future directions for Arabic NER. Xiaoye Qu, Yingjie Gu, Qingrong Xia, Zechang Li, Zhefeng Wang 0001, Baoxing Huai |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2023 | Distantly-Supervised Named Entity Recognition with Adaptive Teacher Learning and Fine-Grained Student EnsembleabstractDistantly-Supervised Named Entity Recognition (DS-NER) effectively alleviates the data scarcity problem in NER by automatically generating training samples. Unfortunately, the distant supervision may induce noisy labels, thus undermining the robustness of the learned models and restricting the practical application. To relieve this problem, recent works adopt self-training teacher-student frameworks to gradually refine the training labels and improve the generalization ability of NER models. However, we argue that the performance of the current self-training frameworks for DS-NER is severely underestimated by their plain designs, including both inadequate student learning and coarse-grained teacher updating. Therefore, in this paper, we make the first attempt to alleviate these issues by proposing: (1) adaptive teacher learning comprised of joint training of two teacher-student networks and considering both consistent and inconsistent predictions between two teachers, thus promoting comprehensive student learning. (2) fine-grained student ensemble that updates each fragment of the teacher model with a temporal moving average of the corresponding fragment of the student, which enhances consistent predictions on each model fragment against noise. To verify the effectiveness of our proposed method, we conduct experiments on four DS-NER datasets. The experimental results demonstrate that our method significantly surpasses previous SOTA methods. The code is available at https://github.com/zenhjunpro/ATSEN. Xiaoye Qu, Daizong Liu, Zhefeng Wang 0001, Baoxing Huai, Pan Zhou 0001 |
AAAI | 5 |
| 2023 | Reference Matters: Benchmarking Factual Error Correction for Dialogue Summarization with Fine-grained Evaluation FrameworkabstractFactuality is important to dialogue summarization.Factual error correction (FEC) of modelgenerated summaries is one way to improve factuality.Current FEC evaluation that relies on factuality metrics is not reliable and detailed enough.To address this problem, we are the first to manually annotate a FEC dataset for dialogue summarization containing 4000 items and propose FERRANTI, a fine-grained evaluation framework based on reference correction that automatically evaluates the performance of FEC models on different error categories.Using this evaluation framework, we conduct sufficient experiments with FEC approaches under a variety of settings and find the best training modes and significant differences in the performance of the existing approaches on different factual error categories. 1 Corrected (ref) Original Summary Corrected (hypo) Align Align Edits (hypo) Edits (ref) Compare Miley needs some rest and wants to go to work tomorrow.Miley needs some rest and does not want to go to work tomorrow.Miley doesn't want to go to work tomorrow.U "needs some rest and" → "" R "wants" → "doesn't want" R "wants" → "does not want" NegE Pred:VerbE TP: 1 TP: 0 FP: 0 FP: 1 FN: 0 TN: 0 Classify Classify Edits (hypo) Edits (ref) U:Pred:VerbE "needs some rest and" → "" R:Pred:NegE "wants" → "doesn't want" Mingqi Gao 0002, Xiaojun Wan 0001, Zhefeng Wang 0001, Baoxing Huai |
ACL (1) | 5 |
| 2023 | Mirror: A Universal Framework for Various Information Extraction TasksabstractTong Zhu, Junfei Ren, Zijian Yu, Mengsong Wu, Guoliang Zhang, Xiaoye Qu, Wenliang Chen, Zhefeng Wang, Baoxing Huai, Min Zhang. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Tong Zhu 0002, Junfei Ren, Zijian Yu, Mengsong Wu, Xiaoye Qu, Wenliang Chen, Zhefeng Wang 0001, Baoxing Huai, Min Zhang 0005 |
EMNLP | 9 |
| 2023 | VarietySound: Timbre-Controllable Video to Sound Generation Via Unsupervised Information DisentanglementabstractVideo-to-sound generation aims to generate realistic and natural sound given a video input. However, previous video-to-sound generation methods can only generate a random or average timbre without any controls of the generated sound timbre, leading to the problem that people cannot obtain the desired timbre under these methods sometimes. In this paper, we propose the task of generating sound with a specific timbre given a silent video input and a reference audio sample. To solve this task, we first use three encoders to disentangle each target sound audio into temporal, acoustic, and background information respectively, then we use a decoder to reconstruct the audio given these disentangled representations. To make the generated result achieve better quality and temporal alignment, we also adopt a mel discriminator and a temporal discriminator for the adversarial training. Our experimental results on the VAS dataset demonstrate that our method can generate high-quality audio samples with good synchronization with events in video and high timbre similarity with the reference audio. Our demos have been published on https://conferencedemos.github.io/icassp23/. Chenye Cui, Zhou Zhao 0001, Yi Ren 0006, Jinglin Liu, Rongjie Huang 0001, Feiyang Chen 0001, Zhefeng Wang 0001, Baoxing Huai, Fei Wu 0001 |
ICASSP | 8 |
| 2023 | CED: Catalog Extraction from Documents
Tong Zhu 0002, Zechang Li, Zijian Yu, Junfei Ren, Mengsong Wu, Zhefeng Wang 0001, Baoxing Huai, Pingfu Chao, Wenliang Chen |
ICDAR (3) | 8 |
| 2023 | Recognizing Unseen Objects via Multimodal Intensive Knowledge Graph PropagationabstractZero-Shot Learning (ZSL), which aims at automatically recognizing unseen objects, is a promising learning paradigm to understand new real-world knowledge for machines continuously. Recently, the Knowledge Graph (KG) has been proven as an effective scheme for handling the zero-shot task with large-scale and non-attribute data. Prior studies always embed relationships of seen and unseen objects into visual information from existing knowledge graphs to promote the cognitive ability of the unseen data. Actually, real-world knowledge is naturally formed by multimodal facts. Compared with ordinary structural knowledge from a graph perspective, multimodal KG can provide cognitive systems with fine-grained knowledge. For example, the text description and visual content can depict more critical details of a fact than only depending on knowledge triplets. Unfortunately, this multimodal fine-grained knowledge is largely unexploited due to the bottleneck of feature alignment between different modalities. To that end, we propose a multimodal intensive ZSL framework that matches regions of images with corresponding semantic embeddings via a designed dense attention module and self-calibration loss. It makes the semantic transfer process of our ZSL framework learns more differentiated knowledge between entities. Our model also gets rid of the performance limitation of only using rough global features. We conduct extensive experiments and evaluate our model on large-scale real-world data. The experimental results clearly demonstrate the effectiveness of the proposed model in standard zero-shot classification tasks. Likang Wu, Zhi Li 0057, Hongke Zhao, Zhefeng Wang 0001, Qi Liu 0003, Baoxing Huai, Nicholas Jing Yuan, Enhong Chen |
KDD | 6 |
| 2023 | Interaction-aware Drug Package Recommendation via Policy GradientabstractRecent years have witnessed the rapid accumulation of massive electronic medical records, which highly support intelligent medical services such as drug recommendation. However, although there are multiple interaction types between drugs, e.g., synergism and antagonism, which can influence the effect of a drug package significantly, prior arts generally neglect the interaction between drugs or consider only a single type of interaction. Moreover, most existing studies generally formulate the problem of package recommendation as getting a personalized scoring function for users, despite the limits of discriminative models to achieve satisfactory performance in practical applications. To this end, in this article, we propose a novel end-to-end Drug Package Generation (DPG) framework, which develops a new generative model for drug package recommendation that considers the interaction effects between drugs that are affected by patient conditions. Specifically, we propose to formulate the drug package generation as a sequence generation process. Along this line, we first initialize the drug interaction graph based on medical records and domain knowledge. Then, we design a novel message-passing neural network to capture the drug interaction, as well as a drug package generator based on a recurrent neural network. In detail, a mask layer is utilized to capture the impact of patient condition, and the deep reinforcement learning technique is leveraged to reduce the dependence on the drug order. Finally, extensive experiments on a real-world dataset from a first-rate hospital demonstrate the effectiveness of our DPG framework compared with several competitive baseline methods. Zhi Zheng 0008, Chao Wang 0086, Tong Xu 0001, Dazhong Shen, Penggang Qin, Xiangyu Zhao 0001, Baoxing Huai, Xian Wu 0001, Enhong Chen |
ACM Trans. Inf. Syst. | 7 |
| 2022 | Flow-Based Unconstrained Lip to Speech GenerationabstractUnconstrained lip-to-speech aims to generate corresponding speeches based on silent facial videos with no restriction to head pose or vocabulary. It is desirable to generate intelligible and natural speech with a fast speed in unconstrained settings. Currently, to handle the more complicated scenarios, most existing methods adopt the autoregressive architecture, which is optimized with the MSE loss. Although these methods have achieved promising performance, they are prone to bring issues including high inference latency and mel-spectrogram over-smoothness. To tackle these problems, we propose a novel flow-based non-autoregressive lip-to-speech model (GlowLTS) to break autoregressive constraints and achieve faster inference. Concretely, we adopt a flow-based decoder which is optimized by maximizing the likelihood of the training data and is capable of more natural and fast speech generation. Moreover, we devise a condition module to improve the intelligibility of generated speech. We demonstrate the superiority of our proposed method through objective and subjective evaluation on Lip2Wav-Chemistry-Lectures and Lip2Wav-Chess-Analysis datasets. Our demo video can be found at https://glowlts.github.io/. Jinzheng He, Zhou Zhao 0001, Yi Ren 0006, Jinglin Liu, Baoxing Huai, Nicholas Jing Yuan |
AAAI | 5 |
| 2022 | Parallel and High-Fidelity Text-to-Lip GenerationabstractAs a key component of talking face generation, lip movements generation determines the naturalness and coherence of the generated talking face video. Prior literature mainly focuses on speech-to-lip generation while there is a paucity in text-to-lip (T2L) generation. T2L is a challenging task and existing end-to-end works depend on the attention mechanism and autoregressive (AR) decoding manner. However, the AR decoding manner generates current lip frame conditioned on frames generated previously, which inherently hinders the inference speed, and also has a detrimental effect on the quality of generated lip frames due to error propagation. This encourages the research of parallel T2L generation. In this work, we propose a parallel decoding model for fast and high-fidelity text-to-lip generation (ParaLip). Specifically, we predict the duration of the encoded linguistic features and model the target lip frames conditioned on the encoded linguistic features with their duration in a non-autoregressive manner. Furthermore, we incorporate the structural similarity index loss and adversarial learning to improve perceptual quality of generated lip frames and alleviate the blurry prediction problem. Extensive experiments conducted on GRID and TCD-TIMIT datasets demonstrate the superiority of proposed methods. Jinglin Liu, Yi Ren 0006, Wencan Huang, Baoxing Huai, Nicholas Jing Yuan, Zhou Zhao 0001 |
AAAI | 5 |
| 2022 | Revisiting Pre-trained Language Models and their Evaluation for Arabic Natural Language ProcessingabstractAbbas Ghaddar, Yimeng Wu, Sunyam Bagga, Ahmad Rashid, Khalil Bibi, Mehdi Rezagholizadeh, Chao Xing, Yasheng Wang, Xinyu Duan, Zhefeng Wang, Baoxing Huai, Xin Jiang, Qun Liu, Phillippe Langlais. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Abbas Ghaddar, Yimeng Wu, Sunyam Bagga, Ahmad Rashid, Khalil Bibi, Mehdi Rezagholizadeh, Yasheng Wang, Xinyu Duan, Zhefeng Wang 0001, Baoxing Huai, Xin Jiang 0002, Qun Liu 0001, Philippe Langlais |
EMNLP | 11 |
| 2022 | ReCo: Reliable Causal Chain Reasoning via Structural Causal Recurrent Neural NetworksabstractCausal chain reasoning (CCR) is an essential ability for many decision-making AI systems, which requires the model to build reliable causal chains by connecting causal pairs.However, CCR suffers from two main transitive problems: threshold effect and scene drift.In other words, the causal pairs to be spliced may have a conflicting threshold boundary or scenario.To address these issues, we propose a novel Reliable Causal chain reasoning framework (ReCo), which introduces exogenous variables to represent the threshold and scene factors of each causal pair within the causal chain, and estimates the threshold and scene contradictions across exogenous variables via structural causal recurrent neural networks (SRNN).Experiments show that ReCo outperforms a series of strong baselines on both Chinese and English CCR datasets.Moreover, by injecting reliable causal chain knowledge distilled by ReCo, BERT can achieve better performances on four downstream causal-related tasks than BERT models enhanced by other kinds of knowledge. Kai Xiong 0002, Ting Liu 0001, Bing Qin 0001, Baoxing Huai |
EMNLP | 8 |
| 2022 | Efficient Document-level Event Extraction via Pseudo-Trigger-aware Pruned Complete GraphabstractMost previous studies of document-level event extraction mainly focus on building argument chains in an autoregressive way, which achieves a certain success but is inefficient in both training and inference. In contrast to the previous studies, we propose a fast and lightweight model named as PTPCG. In our model, we design a novel strategy for event argument combination together with a non-autoregressive decoding algorithm via pruned complete graphs, which are constructed under the guidance of the automatically selected pseudo triggers. Compared to the previous systems, our system achieves competitive results with 19.8% of parameters and much lower resource consumption, taking only 3.8% GPU hours for training and up to 8.5 times faster for inference. Besides, our model shows superior compatibility for the datasets with (or without) triggers and the pseudo triggers can be the supplements for annotated triggers to make further improvements. Codes are available at https://github.com/Spico197/DocEE . Tong Zhu 0002, Xiaoye Qu, Wenliang Chen, Zhefeng Wang 0001, Baoxing Huai, Nicholas Jing Yuan, Min Zhang 0005 |
IJCAI | 5 |
| 2022 | SingGAN: Generative Adversarial Network For High-Fidelity Singing Voice GenerationabstractDeep generative models have achieved significant progress in speech synthesis to date, while high-fidelity singing voice synthesis is still an open problem for its long continuous pronunciation, rich high-frequency parts, and strong expressiveness. Existing neural vocoders designed for text-to-speech cannot directly be applied to singing voice synthesis because they result in glitches and poor high-frequency reconstruction. In this work, we propose SingGAN, a generative adversarial network designed for high-fidelity singing voice synthesis. Specifically, 1) to alleviate the glitch problem in the generated samples, we propose source excitation with the adaptive feature learning filters to expand the receptive field patterns and stabilize long continuous signal generation; and 2) SingGAN introduces global and local discriminators at different scales to enrich low-frequency details and promote high-frequency reconstruction; and 3) To improve the training efficiency, SingGAN includes auxiliary spectrogram losses and sub-band feature matching penalty loss. To the best of our knowledge, SingGAN is the first work designed toward high-fidelity singing voice vocoding. Our evaluation of SingGAN demonstrates the state-of-the-art results with higher-quality (MOS 4.05) samples. Also, SingGAN enables a sample speed of 50x faster than real-time on a single NVIDIA 2080Ti GPU. We further show that SingGAN generalizes well to the mel-spectrogram inversion of unseen singers, and the end-to-end singing voice synthesis system SingGAN-SVS enjoys a two-stage pipeline to transform the music scores into expressive singing voices. Rongjie Huang 0001, Chenye Cui, Feiyang Chen 0001, Yi Ren 0006, Jinglin Liu, Zhou Zhao 0001, Baoxing Huai, Zhefeng Wang 0001 |
ACM Multimedia | 7 |
| 2021 | Read, Retrospect, Select: An MRC Framework to Short Text Entity LinkingabstractEntity linking (EL) for the rapidly growing short text (e.g. search queries and news titles) is critical to industrial applications. Most existing approaches relying on adequate context for long text EL are not effective for the concise and sparse short text. In this paper, we propose a novel framework called Multi-turn Multiple-choice Machine reading comprehension (M3) to solve the short text EL from a new perspective: a query is generated for each ambiguous mention exploiting its surrounding context, and an option selection module is employed to identify the golden entity from candidates using the query. In this way, M3 framework sufficiently interacts limited context with candidate entities during the encoding process, as well as implicitly considers the dissimilarities inside the candidate bunch in the selection stage. In addition, we design a two-stage verifier incorporated into M3 to address the commonly existed unlinkable problem in short text. To further consider the topical coherence and interdependence among referred entities, M3 leverages a multi-turn fashion to deal with mentions in a sequence manner by retrospecting historical cues. Evaluation shows that our M3 framework achieves the state-of-the-art performance on five Chinese and English datasets for the real-world short text EL. Yingjie Gu, Xiaoye Qu, Zhefeng Wang 0001, Baoxing Huai, Nicholas Jing Yuan, Xiaolin Gui |
AAAI | 4 |
| 2021 | Cross-Oilfield Reservoir Classification via Multi-Scale Sensor Knowledge TransferabstractReservoir classification is an essential step for the exploration and production process in the oil and gas industry. An appropriate automatic reservoir classification will not only reduce the manual workloads of experts, but also help petroleum companies to make optimal decisions efficiently, which in turn will dramatically reduce the costs. Existing methods mainly focused on generating reservoir classification in a single geological block but failed to work well on a new oilfield block. Indeed, how to transfer the subsurface characteristics and make accurate reservoir classification across the geological oilfields is a very important but challenging problem. To that end, in this paper, we present a focused study on the cross-oilfield reservoir classification task. Specifically, we first propose a Multi-scale Sensor Extraction (MSE) to extract the multi-scale feature representations of geological characteristics from multivariate well logs. Furthermore, we design an encoder-decoder module, Specific Feature Learning (SFL), to take advantage of specific information of both oilfields. Then, we develop a Knowledge-Attentive Transfer (KAT) module to learn the feature-invariant representation and transfer the geological knowledge from a source oilfield to a target oilfield. Finally, we evaluate our approaches by conducting extensive experiments with real-world industrial datasets. The experimental results clearly demonstrate the effectiveness of our proposed approaches to transfer the geological knowledge and generate the cross-oilfield reservoir classifications. Zhi Li 0057, Zhefeng Wang 0001, Zhicheng Wei, Xiangguang Zhou, Yijun Wang 0002, Baoxing Huai, Qi Liu 0003, Nicholas Jing Yuan, Renbin Gong, Enhong Chen |
AAAI | 6 |
| 2021 | An In-depth Study on Internal Structure of Chinese WordsabstractChen Gong, Saihao Huang, Houquan Zhou, Zhenghua Li, Min Zhang, Zhefeng Wang, Baoxing Huai, Nicholas Jing Yuan. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Chen Gong 0004, Saihao Huang, Houquan Zhou 0001, Zhenghua Li, Min Zhang 0005, Zhefeng Wang 0001, Baoxing Huai, Nicholas Jing Yuan |
ACL/IJCNLP (1) | 7 |
| 2021 | A Coarse-to-Fine Labeling Framework for Joint Word Segmentation, POS Tagging, and Constituent ParsingabstractThe most straightforward approach to joint word segmentation (WS), part-of-speech (POS) tagging, and constituent parsing (PAR) is converting a word-level tree into a char-level tree, which, however, leads to two severe challenges.First, a larger label set (e.g., ≥ 600) and longer inputs both increase computational cost.Second, it is difficult to rule out illegal trees containing conflicting production rules, which is important for reliable model evaluation.If a POS tag (like VV) is above a phrase tag (like VP) in the output tree, it becomes quite complex to decide word boundaries.To deal with both challenges, this work proposes a two-stage coarse-to-fine labeling framework for joint WS-POS-PAR.In the coarse labeling stage, the joint model outputs a bracketed tree, in which each node corresponds to one of four labels (i.e., phrase, subphrase, word, subword).The tree is guaranteed to be legal via constrained CKY decoding.In the fine labeling stage, the model expands each coarse label into a final label (such as VP, VP * , VV, VV * ).Experiments on Chinese Penn Treebank 5.1 and 7.0 show that our joint model consistently outperforms the pipeline approach on both settings of without and with BERT, and achieves new state-of-the-art performance. Yang Hou 0001, Houquan Zhou 0001, Zhenghua Li, Yu Zhang 0092, Min Zhang 0005, Zhefeng Wang 0001, Baoxing Huai, Nicholas Jing Yuan |
CoNLL | 7 |
| 2021 | VLAD-VSA: Cross-Domain Face Presentation Attack Detection with Vocabulary Separation and AdaptationabstractFor face presentation attack detection (PAD), most of the spoofing cues are subtle, local image patterns (e.g., local image distortion, 3D mask edge and cut photo edges). The representations of existing PAD works with simple global pooling method, however, lose the local feature discriminability. In this paper, the VLAD aggregation method is adopted to quantize local features with visual vocabulary locally partitioning the feature space, and hence preserve the local discriminability. We further propose the vocabulary separation and adaptation method to modify VLAD for cross-domain PAD task. The proposed vocabulary separation method divides vocabulary into domain-shared and domain-specific visual words to cope with the diversity of live and attack faces under the cross-domain scenario.The proposed vocabulary adaptation method imitates the maximization step of the k-means algorithm in the end-to-end training, which guarantees the visual words be close to the center of assigned local features and thus brings robust similarity measurement. We give illustrations and extensive experiments to demonstrate the effectiveness of VLAD with the proposed vocabulary separation and adaptation method on standard cross-domain PAD benchmarks. The codes are available at https://github.com/Liubinggunzu/VLAD-VSA. Zhou Zhao 0001, Weike Jin, Xinyu Duan, Zhen Lei 0001, Baoxing Huai, Yiling Wu, Xiaofei He 0001 |
ACM Multimedia | 6 |
| 2021 | Drug Package Recommendation via Interaction-aware Graph InductionabstractRecent years have witnessed the rapid accumulation of massive electronic medical records (EMRs), which highly support the intelligent medical services such as drug recommendation. However, prior arts mainly follow the traditional recommendation strategies like collaborative filtering, which usually treat individual drugs as mutually independent, while the latent interactions among drugs, e.g., synergistic or antagonistic effect, have been largely ignored. To that end, in this paper, we target at developing a new paradigm for drug package recommendation with considering the interaction effect within drugs, in which the interaction effects could be affected by patient conditions. Specifically, we first design a pre-training method based on neural collaborative filtering to get the initial embedding of patients and drugs. Then, the drug interaction graph will be initialized based on medical records and domain knowledge. Along this line, we propose a new Drug Package Recommendation (DPR) framework with two variants, respectively DPR on Weighted Graph (DPR-WG) and DPR on Attributed Graph (DPR-AG) to solve the problem, in which each the interactions will be described as signed weights or attribute vectors. In detail, a mask layer is utilized to capture the impact of patient condition, and graph neural networks (GNNs) are leveraged for the final graph induction task to embed the package. Extensive experiments on a real-world data set from a first-rate hospital demonstrate the effectiveness of our DPR framework compared with several competitive baseline methods, and further support the heuristic study for the drug package generation task with adequate performance. Zhi Zheng 0008, Chao Wang 0086, Tong Xu 0001, Dazhong Shen, Penggang Qin, Baoxing Huai, Tongzhu Liu, Enhong Chen |
WWW | 6 |
| 2020 | A High Precision Pipeline for Financial Knowledge Graph ConstructionabstractMotivated by applications such as question answering, fact checking, and data integration, there is significant interest in constructing knowledge graphs by extracting information from unstructured information sources, particularly text documents.Knowledge graphs have emerged as a standard for structured knowledge representation, whereby entities and their inter-relations are represented and conveniently stored as (subject, predicate, object) triples in a graph that can be used to power various downstream applications.The proliferation of financial news sources reporting on companies, markets, currencies, and stocks presents an opportunity for extracting valuable knowledge about this crucial domain.In this paper, we focus on constructing a knowledge graph automatically by information extraction from a large corpus of financial news articles.For that purpose, we develop a high precision knowledge extraction pipeline tailored for the financial domain.This pipeline combines multiple information extraction techniques with a financial dictionary that we built, all working together to produce over 342,000 compact extractions from over 288,000 financial news articles, with a precision of 78% at the top-100 extractions.The extracted triples are stored in a knowledge graph making them readily available for use in downstream applications. Sarah Elhammadi, Laks V. S. Lakshmanan, Raymond T. Ng, Michael Simpson 0001, Baoxing Huai, Zhefeng Wang 0001, Lanjun Wang |
COLING | 5 |
| 2020 | Object-Aware Multi-Branch Relation Networks for Spatio-Temporal Video GroundingabstractSpatio-temporal video grounding aims to retrieve the spatio-temporal tube of a queried object according to the given sentence. Currently, most existing grounding methods are restricted to well-aligned segment-sentence pairs. In this paper, we explore spatio-temporal video grounding on unaligned data and multi-form sentences. This challenging task requires to capture critical object relations to identify the queried target. However, existing approaches cannot distinguish notable objects and remain in ineffective relation modeling between unnecessary objects. Thus, we propose a novel object-aware multi-branch relation network for object-aware relation discovery. Concretely, we first devise multiple branches to develop object-aware region modeling, where each branch focuses on a crucial object mentioned in the sentence. We then propose multi-branch relation reasoning to capture critical object relationships between the main branch and auxiliary branches. Moreover, we apply a diversity loss to make each branch only pay attention to its corresponding object and boost multi-branch learning. The extensive experiments show the effectiveness of our proposed method. Zhou Zhao 0001, Zhijie Lin 0001, Baoxing Huai, Nicholas Jing Yuan |
IJCAI | 4 |
| 2020 | Text-Guided Image InpaintingabstractGiven a partially masked image, image inpainting aims to complete the missing region and output a plausible image. Most existing image inpainting methods complete the missing region by expanding or borrowing information from the surrounding source region, which work well when the original content in the missing region is similar to the surrounding source region. Unsatisfactory results will be generated if there is no sufficient contextual information can be referenced from source region. Besides, the inpainting results should be diverse and this kind of diversity should be controllable. Based on these observations, we propose a new inpainting problem that introduces text as a kind of guidance to direct and control the inpainting process. The main difference between this problem and previous works is that we need ensure the result to be consistent with not only the source region but also the textual guidance during inpainting. By this way, we want to avoid the unreasonable completion and meanwhile make it controllable. We propose a progressively coarse-to-fine cross-modal generative network and adopt the text-image-text training schema to generate visually consistent and semantically coherent images. Extensive quantitative and qualitative experiments on two public datasets with captions demonstrate the effectiveness of our method. Zijian Zhang 0002, Zhou Zhao 0001, Baoxing Huai, Nicholas Jing Yuan |
ACM Multimedia | 4 |
| 2020 | Multimodal Dialogue Systems via Capturing Context-aware Dependencies of Semantic ElementsabstractRecently, multimodal dialogue systems have engaged increasing attention in several domains such as retail, travel, etc. In spite of the promising performance of pioneer works, existing studies usually focus on utterance-level semantic representations with hierarchical structures, which ignore the context-aware dependencies of multimodal semantic elements, i.e., words and images. Moreover, when integrating the visual content, they only consider images of the current turn, leaving out ones of previous turns as well as their ordinal information. To address these issues, we propose a Multimodal diAlogue systems with semanTic Elements, MATE for short. Specifically, we unfold the multimodal inputs and devise a Multimodal Element-level Encoder to obtain the semantic representation at element-level. Besides, we take into consideration all images that might be relevant to the current turn and inject the sequential characteristics of images through position encoding. Finally, we make comprehensive experiments on a public multimodal dialogue dataset in the retail domain, and improve the BLUE-4 score by 9.49, and NIST score by 1.8469 compared with state-of-the-art methods. Weidong He, Zhi Li 0057, Dongcai Lu, Enhong Chen, Tong Xu 0001, Baoxing Huai, Nicholas Jing Yuan |
ACM Multimedia | 6 |
| 2020 | FastLR: Non-Autoregressive Lipreading Model with Integrate-and-FireabstractLipreading is an impressive technique and there has been a definite improvement of accuracy in recent years. However, existing methods for lipreading mainly build on autoregressive (AR) model, which generate target tokens one by one and suffer from high inference latency. To breakthrough this constraint, we propose FastLR, a non-autoregressive (NAR) lipreading model which generates all target tokens simultaneously. NAR lipreading is a challenging task that has many difficulties: 1) the discrepancy of sequence lengths between source and target makes it difficult to estimate the length of the output sequence; 2) the conditionally independent behavior of NAR generation lacks the correlation across time which leads to a poor approximation of target distribution; 3) the feature representation ability of encoder can be weak due to lack of effective alignment mechanism; and 4) the removal of AR language model exacerbates the inherent ambiguity problem of lipreading. Thus, in this paper, we introduce three methods to reduce the gap between FastLR and AR model: 1) to address challenges 1 and 2, we leverage integrate-and-fire (I&F) module to model the correspondence between source video frames and output text sequence. 2) To tackle challenge 3, we add an auxiliary connectionist temporal classification (CTC) decoder to the top of the encoder and optimize it with extra CTC loss. We also add an auxiliary autoregressive decoder to help the feature extraction of encoder. 3) To overcome challenge 4, we propose a novel Noisy Parallel Decoding (NPD) for I&F and bring Byte-Pair Encoding (BPE) into lipreading. Our experiments exhibit that FastLR achieves the speedup up to 10.97× comparing with state-of-the-art lipreading model with slight WER absolute increase of 1.5% and 5.5% on GRID and LRS2 lipreading datasets respectively, which demonstrates the effectiveness of our proposed method. Jinglin Liu, Yi Ren 0006, Zhou Zhao 0001, Chen Zhang 0020, Baoxing Huai, Nicholas Jing Yuan |
ACM Multimedia | 5 |
| 2020 | Distant Supervision for Multi-Stage Fine-Tuning in Retrieval-Based Question AnsweringabstractWe tackle the problem of question answering directly on a large document collection, combining simple “bag of words” passage retrieval with a BERT-based reader for extracting answer spans. In the context of this architecture, we present a data augmentation technique using distant supervision to automatically annotate paragraphs as either positive or negative examples to supplement existing training data, which are then used together to fine-tune BERT. We explore a number of details that are critical to achieving high accuracy in this setup: the proper sequencing of different datasets during fine-tuning, the balance between “difficult” vs. “easy” examples, and different approaches to gathering negative examples. Experimental results show that, with the appropriate settings, we can achieve large gains in effectiveness on two English and two Chinese QA datasets. We are able to achieve results at or near the state of the art without any modeling advances, which once again affirms the cliché “there’s no data like more data”. Yuqing Xie 0001, Wei Yang 0017, Luchen Tan, Kun Xiong, Nicholas Jing Yuan, Baoxing Huai, Ming Li 0001, Jimmy Lin |
WWW | 6 |
| 2014 | Learning to annotate via social interaction analytics
Tong Xu 0001, Hengshu Zhu, Enhong Chen, Baoxing Huai, Hui Xiong 0001, Jilei Tian |
Knowl. Inf. Syst. | 4 |
| 2014 | Toward Personalized Context Recognition for Mobile Users: A Semisupervised Bayesian HMM ApproachabstractThe problem of mobile context recognition targets the identification of semantic meaning of context in a mobile environment. This plays an important role in understanding mobile user behaviors and thus provides the opportunity for the development of better intelligent context-aware services. A key step of context recognition is to model the personalized contextual information of mobile users. Although many studies have been devoted to mobile context modeling, limited efforts have been made on the exploitation of the sequential and dependency characteristics of mobile contextual information. Also, the latent semantics behind mobile context are often ambiguous and poorly understood. Indeed, a promising direction is to incorporate some domain knowledge of common contexts, such as “waiting for a bus” or “having dinner,” by modeling both labeled and unlabeled context data from mobile users because there are often few labeled contexts available in practice. To this end, in this article, we propose a sequence-based semisupervised approach to modeling personalized context for mobile users. Specifically, we first exploit the Bayesian Hidden Markov Model (B-HMM) for modeling context in the form of probabilistic distributions and transitions of raw context data. Also, we propose a sequential model by extending B-HMM with the prior knowledge of contextual features to model context more accurately. Then, to efficiently learn the parameters and initial values of the proposed models, we develop a novel approach for parameter estimation by integrating the Dirichlet Process Mixture (DPM) model and the Mixture Unigram (MU) model. Furthermore, by incorporating both user-labeled and unlabeled data, we propose a semisupervised learning-based algorithm to identify and model the latent semantics of context. Finally, experimental results on real-world data clearly validate both the efficiency and effectiveness of the proposed approaches for recognizing personalized context of mobile users. Baoxing Huai, Enhong Chen, Hengshu Zhu, Hui Xiong 0001, Tengfei Bao, Qi Liu 0003, Jilei Tian |
ACM Trans. Knowl. Discov. Data | 1 |