VLDB 2026 Research / reviewers in the wild / expert
Feilong Bao
dblp:136/5318
· DBLP profile ↗
78ranked-venue papers
2as first author
49since 2021 · last 2026
0000-0001-7312-1629ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 56 · 1 first-author · 32 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 1 first-author · 15 since 2021Databases, data management, data science and information retrieval · 9 · 7 since 2021Computer networks · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CEDAR: A Chinese Evaluation Dataset for Computational ArgumentationabstractTian Lan, Jiang Li, Rong Yan, Feilong Bao, Weihua Wang, Guanglai Gao, Xiangdong Su. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jiang Li 0013, Feilong Bao, Weihua Wang 0006, Guanglai Gao, Xiangdong Su |
ACL (1) | 4 |
| 2026 | Construction and Evaluation of Large Language Models for the Mongolian Medicine Diagnostic and Treatment System
Jixieqi Bai, Feilong Bao, Hui Zhang 0031, Aruukhan Bai |
KSEM (6) | 2 |
| 2026 | TM-Bench: Benchmarking Large Language Models on Low-Resource Traditional MongolianabstractLarge language models (LLMs) have achieved remarkable success in high-resource languages, yet their performance on Traditional Mongolian remains highly limited. A primary bottleneck is the absence of a systematic evaluation framework, which precludes quantitative comparison and obscures directions for model optimization. In this paper, we introduce TM-Bench, the first comprehensive benchmark for LLMs on Traditional Mongolian. TM-Bench adopts a hybrid construction strategy consisting of human-verified Translation-based Adaptation, Expert-Original Authoring, and Semi-automated Synthesis. It comprises 18,357 instances spanning five tasks across both natural language understanding and generation to evaluate models' reasoning, knowledge application, and linguistic proficiency. We conduct systematic evaluations across representative model families. The results show that on understanding tasks, model performance lags significantly behind high-resource languages, with only a few models performing slightly above the random baseline. For generation tasks, both automatic metrics and double-blind human evaluations reveal severe semantic collapse, failing to generate coherent text and often producing unreadable gibberish. These findings underscore the critical role of TM-Bench as a foundational infrastructure for evaluating LLMs in Traditional Mongolian and catalyzing future model optimization. Our benchmark and code are available at https://github.com/gao1948083886/TM-Bench. Zhenjie Gao, Feilong Bao, Aruukhan Bai, Ruichen Hou, Xieqi Ji, Dabalgan Wang, Hugjil Ming |
SIGIR | 2 |
| 2026 | Selective Distillation for Continual Named Entity Recognition with Memory Replay
Weihua Wang 0006, Feilong Bao |
SIGIR | 3 |
| 2026 | How to teach and forget: Towards cross-modal semantic consistency for entity alignment
Cunda Wang, Chenglong Miao, Po Hu 0001, Weihua Wang 0006, Feilong Bao |
Inf. Process. Manag. | 5 |
| 2026 | Wisteria: A unified multi-scale feature learning framework for DNA language model
Weihua Wang 0006, Haoji Li, Feilong Bao, Guanglai Gao |
Pattern Recognit. | 3 |
| 2026 | Hyperbolic-Based Cross-Modal Semantic Remodeling Network for Zero-Shot Sketch-Based Image RetrievalabstractThe Zero-Shot Sketch-Based Image Retrieval (ZS-SBIR) task aims to retrieve images associated with sketches from unseen classes, bringing great convenience to the engineering field. To address the modality gap, most existing works project images and sketches into a shared Euclidean space. However, the hierarchical structure of image data makes the Euclidean space not the optimal choice as an embedding space for representing complex structured image data. Meanwhile, existing text and hierarchical models are not effective enough for addressing the problem of knowledge transfer. To address these issues, this article proposes an original Hyperbolic-Based Cross-Modal Semantic Remodeling Network (called HCMSN) for ZS-SBIR. Specifically, this article proposes to extract category-level word embeddings based on BERT model, then align image features and sketches with the word embeddings using adversarial methods. Meanwhile, this article further proposes a cross-modal retrieval feature reconstruction network for improving the informativeness and robustness of retrieval features. Moreover, this article presents a feature projection network that maps the retrieval features to the hyperbolic space to generate the hyperbolic retrieval features, thus effectively representing the data with hierarchical structure. Extensive experiments demonstrate that the mAP@all of our HCMSN model surpasses CNN-based models by 20.9% on the Sketchy dataset, 1.2% on the more difficult TU-Berlin dataset, and 13.6% on the more challenging QuickDraw dataset. Xiangdong Su, Feilong Bao, Guanglai Gao |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | Distance-Adaptive Quaternion Knowledge Graph Embedding with Bidirectional RotationabstractQuaternion contains one real part and three imaginary parts, which provided a more expressive hypercomplex space for learning knowledge graph. Existing quaternion embedding models measure the plausibility of a triplet either through semantic matching or distance scoring functions. However, it appears that semantic matching diminishes the separability of entities, while the distance scoring function weakens the semantics of entities. To address this issue, we propose a novel quaternion knowledge graph embedding model. Our model combines semantic matching with entity’s geometric distance to better measure the plausibility of triplets. Specifically, in the quaternion space, we perform a right rotation on the head entity and a reverse rotation on the tail entity to learn the rich semantic features. Then, we utilize distance adaptive translations to learn the geometric distance between entities. Furthermore, we provide mathematical proofs to demonstrate our model can handle complex logical relationships. Extensive experimental results and analyses show our model significantly outperforms previous models on well-known knowledge graph completion benchmark datasets. Our code is available at https://anonymous.4open.science/r/l2730. Weihua Wang 0006, Qiuyu Liang, Feilong Bao, Guanglai Gao |
COLING | 3 |
| 2025 | Unifying Dual-Space Embedding for Entity Alignment via Contrastive LearningabstractEntity alignment (EA) aims to match identical entities across different knowledge graphs (KGs). Graph neural network-based entity alignment methods have achieved promising results in Euclidean space. However, KGs often contain complex local and hierarchical structures, which are hard to represent in a single space. In this paper, we propose a novel method named as UniEA, which unifies dual-space embedding to preserve the intrinsic structure of KGs. Specifically, we simultaneously learn graph structure embeddings in both Euclidean and hyperbolic spaces to maximize the consistency between embeddings in the two spaces. Moreover, we employ contrastive learning to mitigate the misalignment issues caused by similar entities, where embeddings of similar neighboring entities become too close. Extensive experiments on benchmark datasets demonstrate that our method achieves state-of-the-art performance in structure-based EA. Our code is available at https://github.com/wonderCS1213/UniEA. Cunda Wang, Weihua Wang 0006, Qiuyu Liang, Feilong Bao, Guanglai Gao |
COLING | 4 |
| 2025 | Stand on The Shoulders of Giants: Building JailExpert from Previous Attack ExperienceabstractXi Wang, Songlei Jian, Shasha Li, Xiaopeng Li, Bin Ji, Ma Jun, Xiaodong Liu, Jing Wang, Jianfeng Zhang, Jie Yu, Feilong Bao, Wangbaosheng. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Songlei Jian, Ma Jun, Feilong Bao, Wangbaosheng |
EMNLP | 11 |
| 2025 | Hyperbolic Multimodal Knowledge Graph EmbeddingabstractMultimodal knowledge graph embedding refers to learning multimodal entities and their relation representations in a low-dimensional space. However, existing multimodal embedding models tend to ignore the inherent structure of knowledge graphs. To address this issue, we propose a novel multimodal knowledge graph embedding model to simultaneously learn semantic relation and hierarchical structure of entities within a hyperbolic space. Specifically, we project all modalities features embedding into a hyperbolic space and unify these embeddings to form a multimodal embedding. Then, we model the knowledge graph triplets by treating the relation as a Lorentzian linear transformation from head entity to tail entity. The plausibility of triplets is measured by Lorentz distance. Extensive experiments on multimodal knowledge graph completion benchmarks validate that our model achieves the state-of-the-art results across most metrics. In terms of training speed, our model is one order of magnitude faster than the best one. The visualization results further reveal our model’s ability to capture hierarchical structures. Our code is available at https://github.com/llqy123/HyME. Qiuyu Liang, Weihua Wang 0006, Cunda Wang, Feilong Bao, Jie Yu 0008 |
ICASSP | 4 |
| 2025 | A High-Precision Character Cartoon Style Transfer Method Based on VToonify and Diffusion ModelsabstractIn recent years, the technology for converting character images into cartoons has gained widespread attention on short video platforms, with diffusion models particularly standing out in the field of image style transfer. However, existing methods still have shortcomings in detail handling and character feature control. To address these issues, this paper proposes a high-precision character style transfer framework that combines the VToonify model with diffusion models. By using Mask-based integration of VToonify and Diffusion models, we significantly enhance character detail representation while maintaining overall style consistency. Experimental results demonstrate that this model outperforms existing similar generative models in both detail control and overall effectiveness, offering a new solution for high-quality personalized style transfer. Weiting Wang, Feilong Bao |
ICASSP | 3 |
| 2025 | OTMEA : Multi-modal Entity Alignment via Optimal TransportabstractMulti-modal Entity Alignment (MMEA) aims to identify the same entities exhibited in different knowledge graphs (KGs), where the entities are enriched by structure and visual information. Existing MMEA methods learn multi-modal joint entity embeddings by encompassing both modality interaction and modality alignment. However, these approaches predominantly emphasize modality interaction and fail to adequately address the issue of modality heterogeneity. In this paper, we propose a novel approach OTMEA, which leverages optimal transport to mitigate modality heterogeneity from the perspective of modality distributions. Specifically, we view the modality alignment problem as a Wasserstein minimum distance problem involving multimodal distributions. Furthermore, our experiments indicate that employing entity-level attention weights significantly enhances modality alignment through optimal transport. The effectiveness of our method is validated through extensive experiments conducted on five public datasets. The source code is available at https://github.com/wonderCS1213/OTMEA. Cunda Wang, Weihua Wang 0006, Qiuyu Liang, Feilong Bao |
ICASSP | 5 |
| 2025 | Multilingual Parameter-Sharing Adapters: A Method for Optimizing Low-Resource Neural Machine TranslationabstractAdapter-based Multilingual Neural Machine Translation (MNMT) has become a significant approach in low-resource language translation by mitigating data imbalances between high-resource and low-resource language pairs and reducing training costs. However, existing adapter-based methods lack generalization in cross-lingual settings, particularly under low-resource conditions, where their scalability is limited. Additionally, current methods often introduce independent adapter modules for each language, leading to a linear increase in model parameters with the number of languages. To address these challenges, we propose a multilingual parameter-sharing adapter approach. Moreover, we introduce a neural architecture search (NAS)-based strategy to improve translation performance. Experimental results demonstrate that the multilingual parameter-sharing adapter exhibits competitive performance on both low-resource and high-resource datasets. The multilingual parameter-sharing adapter method has only 400K trainable parameters, which is 20× lower than the parameters of the traditional adapter method. Yonghe Wang, Xiangdong Su, Feilong Bao |
ICASSP | 5 |
| 2025 | Structural-Aware Disentangled Learning with CLIP for Hyperbolic Zero-Shot Sketch-Based Image RetrievalabstractThe zero-shot sketch-based image retrieval task faces two key challenges: domain gap and knowledge transfer. Our innovation is recognizing that directly aligning cross-domain features weakens the discriminative ability of the model, as it overlooks the asymmetry between sketches and images. Additionally, Euclidean space is inadequate for capturing the hierarchical structure, which limits the performance of the model on complex data. To address these issues, we propose a Structural-Aware Disentangled Learning network (termed SADLnet) that incorporates CLIP and hyperbolic geometry. Specifically, we use CLIP to extract visual features from each domain to enhance the domain generalization of the model. Furthermore, we design a structure-guided disentanglement strategy to decompose image representations into sketch-related and sketch-unrelated features, addressing the domain gap. Moreover, we project the retrieval features into hyperbolic space to capture hierarchical information, improving feature discrimination in retrieval tasks. Extensive experiments demonstrate that SADLnet establishes new state-of-the-art performance on three datasets. Feilong Bao, Xiangdong Su, Guanglai Gao |
ICASSP | 3 |
| 2025 | Enhancing Mandarin Lip Reading with a Multimodal Conformer and Structured State-Space DecoderabstractLip reading is a visual recognition technology that interprets spoken content by decoding lip movements. Since speech perception is inherently a multimodal task, incorporating audio information during training is crucial to assist lip reading. This paper proposes a novel architecture that combines the Conformer network with a structured state space decoder under multimodal input to enhance Mandarin lip reading capabilities. As a tonal language, Mandarin benefits from audio cues that guide visual information learning, improving the accuracy of speech content recognition. Our approach leverages the Conformer to extract shared semantics from both audio and video, and employs a bidirectional structured state space decoder to decode, effectively capturing the temporal dynamics and complex dependencies of long sequences. This method achieved CERs of 54.97% and 12.53% in the CN-CVS and CMLR datasets, respectively. The research code is open source at: https://anonymous.4open.science/r/Lip-reading-model-D7B8. Meng Miao, Feilong Bao, Guanglai Gao |
IJCNN | 2 |
| 2025 | Zero-Shot Speech Recognition from Text-Only Data through Synthesized Spectrogram Refinement Using Style Truncation and Contextual Alignment LossabstractUtilizing pseudo speech-label pairs synthesized via Text-to-Speech (TTS) systems as supplementary training data for automatic speech recognition (ASR) has shown significant benefits. However, the mismatch between synthesized and real speech make them unsuitable for direct use in zero-shot speech recognition tasks. In this paper, we propose a generative-adversarial model with a style truncation strategy that enhances mel-spectrograms by projecting style vectors into high-density regions. We also introduce a Contextual alignment loss function to align synthesized and real mel-spectrograms by computing global distances between high-dimensional features. Experimental results on zero-shot ASR tasks using the SLURP and LibriSpeech datasets show that our method achieves WER reductions of 4.3%, 11.1% (clean), and 6.9% (other) compared to the baseline mel-spectrograms enhanced model. Yonghe Wang, Zhenjie Gao, Feilong Bao |
MMAsia | 4 |
| 2025 | MAD-HD: Multi-agent Debate-Driven Ungrounded Hallucination Detection
Zhenjie Gao, Feilong Bao |
NLPCC (1) | 2 |
| 2025 | Mongolian Speech Recognition Based on Semi-supervised Learning and Syllable Subword Modeling Units
Yonghe Wang, Zhenjie Gao, Feilong Bao |
NLPCC (4) | 4 |
| 2025 | Exploring Representation-Efficient Transfer Learning Approaches for Speech Recognition and Translation Using Pre-trained Speech Models
Yonghe Wang, Feilong Bao |
NLPCC (1) | 4 |
| 2025 | RSCAC-NET: A Remote Sensing Image Change Description Network Based on Change-Aware and Multi-stage Global Fusion
Hongyi Dong, Xiuzhen He, Yan Wang 0037, Jing Liu 0003, Feilong Bao, Bing Jia |
NPC (1) | 5 |
| 2025 | Robust Self-Localization of Wireless Acoustic Sensor NetworksabstractWireless acoustic sensor networks (WASNs), or the so-called Internet of Audio Things (IoAuT), have attracted increasing attention in the Internet of Things community. As the geometric structure of WASNs is required in audio/speech processing tasks like source localization or acoustic beamforming, automatic self-localization of sensors is necessary. However, most of the existing approaches suffer from poor stability, as their constructed cost functions involve nonconvex programming. To address this issue, we investigate the robust self-localization (or geometry calibration) of WASNs in this article. Specifically, a rough self-localization (RSL) method is first presented based on measurements including Time-Difference-of-Arrivals (TDoAs), direction-of-arrivals (DoA), and energy-rates (ERs), and its closed-form solution is further derived. As ER estimates are sensitive to acoustic environments, the performance of the RSL method is somewhat limited. Therefore, a precise self-localization (PSL) method is then developed by building a weighted (and nonconvex) TDoA-DoA cost function, after regarding the RSL approach as an initialization step. As the RSL offers better initial values compared with existing initialization strategies, the combination of RSL and PSL methods (named as RSL-PSL method) shows stronger robustness and stability. In addition, computational complexity of both RSL and PSL methods is analyzed in detail. Finally, the Cramér-Rao Bound (CRB) of the PSL method is derived to show its theoretical lower bound. The proposed RSL-PSL method outperforms the state-of-the-arts in terms of stability and accuracy, which is confirmed by numerical real-world and simulation experiments. Xu Wang 0058, De Hu, Rui Liu 0008, Feilong Bao |
IEEE Internet Things J. | 4 |
| 2025 | Domain disentanglement and fusion based on hyperbolic neural networks for zero-shot sketch-based image retrieval
Xiangdong Su, Yonghe Wang, Feilong Bao, Guanglai Gao |
Inf. Process. Manag. | 5 |
| 2024 | Hyperbolic Representations for Prompt LearningabstractContinuous prompt tuning has gained significant attention for its ability to train only continuous prompts while freezing the language model. This approach greatly reduces the training time and storage for downstream tasks. In this work, we delve into the hierarchical relationship between the prompts and downstream text inputs. In prompt learning, the prefix prompt acts as a module to guide the downstream language model, establishing a hierarchical relationship between the prefix prompt and subsequent inputs. Furthermore, we explore the benefits of leveraging hyperbolic space for modeling hierarchical structures. We project representations of pre-trained models from Euclidean space into hyperbolic space using the Poincaré disk which effectively captures the hierarchical relationship between the prompt and input text. The experiments on natural language understanding (NLU) tasks illustrate that hyperbolic space can model the hierarchical relationship between prompt and text input. We release our code at https://github.com/myaxxxxx/Hyperbolic-Prompt-Learning. Xiangdong Su, Feilong Bao |
LREC/COLING | 3 |
| 2024 | L\²GC: Lorentzian Linear Graph Convolutional Networks for Node Classification
Qiuyu Liang, Weihua Wang 0006, Feilong Bao, Guanglai Gao |
LREC/COLING | 3 |
| 2024 | Fully Hyperbolic Rotation for Knowledge Graph EmbeddingabstractHyperbolic rotation is commonly used to effectively model knowledge graphs and their inherent hierarchies. However, existing hyperbolic rotation models rely on logarithmic and exponential mappings for feature transformation. These models only project data features into hyperbolic space for rotation, limiting their ability to fully exploit the hyperbolic space. To address this problem, we propose a novel fully hyperbolic model designed for knowledge graph embedding. Instead of feature mappings, we define the model directly in hyperbolic space with the Lorentz model. Our model considers each relation in knowledge graphs as a Lorentz rotation from the head entity to the tail entity. We adopt the Lorentzian version distance as the scoring function for measuring the plausibility of triplets. Extensive results on standard knowledge graph completion benchmarks demonstrated that our model achieves competitive results with fewer parameters. In addition, our model get the state-of-the-art performance on datasets of CoDEx-s and CoDEx-m, which are more diverse and challenging than before. Our code is available at https://github.com/llqy123/FHRE. Qiuyu Liang, Weihua Wang 0006, Feilong Bao, Guanglai Gao |
ECAI | 3 |
| 2024 | Efficient Speech-to-Text Translation: Progressive Pruning for Accelerated Speech Pre-trained ModelabstractRecently, speech pre-trained models based on the Transformer architecture have become very popular for speech-to-text translation tasks. However, computing representation outputs of speech pre-trained models is highly time-consuming, primarily due to the length of speech sequences far exceeding the corresponding texts, leading to quadratic computation costs for the self-attention module. To address this issue, we propose a novel pruning method that progressively reduces representation sequence length layer by layer. We leverage attention scores to calculate importance scores for all tokens. Additionally, we introduce both fixed and scheduled pruning rate strategies to determine which tokens should be retained. Experiments demonstrate that our approach reduces the output tokens of the speech pre-trained model by 55%, with only a 0.7% performance decrease, and improves in practice encoding speed up to 1.76 ×. Our method also is an orthogonal and complementary direction to efficient speech pre-trained models. We release our code at https://github.com/myaxxxxx/pruning. Yonghe Wang, Xiangdong Su, Feilong Bao |
ICME | 4 |
| 2024 | Improving End-to-End Speech Recognition Through Conditional Cross-Modal Knowledge Distillation with Language ModelabstractRecently, cross-modal knowledge distillation methods for end-to-end automatic speech recognition (E2E-ASR) model training pointed out the potential help of text data for improving recognition performance. However, conventional optimization strategies can mislead student models to produce suboptimal performance due to erroneous predictions generated by teacher models. This paper addresses the issue by proposing a conditional cross-modal knowledge distillation strategy, a novel technique for selectively incorporating contextual linguistic information from language model into the E2E-ASR model for improving the recognition performance. We introduce a conditional selector to dynamically adjust the knowledge source of the student model to avoid knowledge distillation from erroneous predictions generated by teacher model. In pre-trained language model fine-tuning, we perform an analysis of the impact of unsupervised text data of varying scales on the quality of soft labels and the recognition performance of the E2E-ASR model. Our proposed method simultaneously improve two different non-autoregressive decoding approaches. Experiments on the Chinese speech datasets AISHELL-1 and AISHELL-2 show competitive performance. Yonghe Wang, Feilong Bao, Zhenjie Gao, Guanglai Gao |
IJCNN | 3 |
| 2024 | Hierarchy-Aware Quaternion Embedding for Knowledge Graph CompletionabstractKnowledge graph completion is an essential task in the fields of graph mining and graph machine learning. Most contemporary approaches rely on geometric transformation to achieve knowledge graph completion, as geometry offers a well-defined mathematical foundation. For example, rotation transformations in rigid body transformation are frequently employed within quaternion spaces to model complex relation types in knowledge graphs. However, these models cannot effectively handle the hierarchical structure in the knowledge graph. As a result, the performance of knowledge graph completion suffers. To address this shortcoming of quaternion space, we propose a novel model that integrates hyperbolic space. Specifically, we perform a translation transformation in a hyperbolic space to obtain support vector embeddings that imply relation embedding. We then perform a rotation transformation with the Hamilton product in tangent space, treating the relation embedding as a rotation from the head entity embedding to the tail entity embedding. We verify the validity and generalization ability of our model on standard benchmark datasets including WN18RR, FB15k-237 and YAGO3-10. The experimental results show that our model achieves competitive results on MRR and H@K metrics. Our code is publicly available at https://github.com/llqy123/HAQE-master. Qiuyu Liang, Weihua Wang 0006, Jie Yu 0008, Feilong Bao |
IJCNN | 4 |
| 2024 | Pre-training Language Model for Mongolian with Agglutinative Linguistic Knowledge InjectionabstractBERT based Pre-training Language Model (PLM) has become a crucial step in achieving the best results in various natural language processing (NLP) tasks. However, the current progress, which mainly focuses on major languages such as English and Chinese, has not thoroughly investigated the low-resource languages, particularly agglutinative languages like Mongolian, due to the scarcity of large-scale data resources and the difficulty of understanding agglutinative knowledge. In this paper, we propose a novel PLM for the Mongolian language, that incorporates a novel three-stage agglutinative knowledge injection strategy. Specifically, early-stage injection aims to convert the Mongolia word sequence to the fine-grained sub-word token that comprises a stem and some suffixes; Middle-stage injection designed a morphological knowledge-based masking strategy to enhance the model's ability to learn agglutinative knowledge; Late-stage injection not only involves the model restoring the masked tokens but also predicting the order of suffixes. To address the issue of data scarcity, we create a large-scale Mongolian PLM dataset and three datasets for three downstream tasks, that are News Classification, Name Entity Recognition (NER), and Part-of-Speech (POS) prediction, etc. The experimental results on three downstream tasks demonstrate that our method surpasses the traditional BERT approach and successfully learns agglutinative language knowledge in Mongolian. Muhan Na, Rui Liu 0008, Feilong Bao, Guanglai Gao |
IJCNN | 3 |
| 2024 | Parameter-Efficient Adapter Based on Pre-trained Models for Speech Translation
Yonghe Wang, Feilong Bao |
INTERSPEECH | 3 |
| 2024 | Knowledge-Preserving Pluggable Modules for Multilingual Speech Translation Tasks
Yonghe Wang, Feilong Bao |
INTERSPEECH | 3 |
| 2024 | Sign Value Constraint Decomposition for Efficient 1-Bit Quantization of Speech Translation Tasks
Yonghe Wang, Feilong Bao |
INTERSPEECH | 3 |
| 2024 | Effective Knowledge Graph Embedding with Quaternion Convolutional Networks
Qiuyu Liang, Weihua Wang 0006, Jie Yu 0008, Feilong Bao |
NLPCC (3) | 4 |
| 2024 | Segmentation-Free Todo Mongolian OCR and its Public Dataset
Feilong Bao |
PRCV (7) | 2 |
| 2024 | The image and ground truth dataset of Mongolian movable-type newspapers for text recognition
Feilong Bao, Hui Zhang 0031, Guanglai Gao |
Int. J. Document Anal. Recognit. | 2 |
| 2023 | TableSF: A Structural Bias Framework for Table-To-Text Generation
Weihua Wang 0006, Feilong Bao, Guanglai Gao |
ICANN (9) | 3 |
| 2023 | Few-Shot Table-to-Text Generation with Structural Bias Attention
Weihua Wang 0006, Feilong Bao, Guanglai Gao |
PRICAI (2) | 3 |
| 2023 | A Comparative Study on Selecting Acoustic Modeling Units for WFST-based Mongolian Speech RecognitionabstractTraditional weighted finite-state transducer– (WFST) based Mongolian automatic speech recognition (ASR) systems use phonemes as pronunciation lexicon modeling units. However, Mongolian is an agglutinative, low-resource language, and building an ASR system based on the phoneme pronunciation lexicon remains a challenge for various reasons. First, the phoneme pronunciation lexicon manually constructed by Mongolian linguists is finite, which is usually used to build a grapheme-to-phoneme conversion (G2P) model to frequently expand new words. However, the data sparsity decreases the robustness of the G2P model and affects the performance of the final ASR system. Second, homophones and polysyllabic words are common in Mongolian, which has a certain impact on the construction of the Mongolian acoustic model. To address these problems, in this work, we first propose a grapheme-to-phoneme alignment model to obtain the mapping relationship between phonemes and subword units. Then, we construct an acoustic subword segmentation set to segment words directly instead of using the traditional G2P method to predict phoneme sequences to expand the pronunciation lexicon. Further, by analyzing the Mongolian encoding form, we also propose an acoustic subword modeling units construction method that removes control characters. Finally, we investigate various acoustic subword modeling units for pronunciation lexicon construction for the Mongolian ASR system. Experiments on a Mongolian dataset with 325 hours of training show that the pronunciation lexicon based on the acoustic subword modeling unit can effectively construct the WFST-based Mongolian ASR system. Further, removing the control characters when building the acoustic subword modeling unit can further improve the ASR system performance. Yonghe Wang, Feilong Bao, Guanglai Gao |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2023 | Distributed Sensor Selection for Speech Enhancement With Acoustic Sensor NetworksabstractIn distributed acoustic sensor networks, only a few nodes make a significant contribution to speech enhancement tasks. Using these most informative nodes instead of the entire network not only avoids unnecessary energy consumption but also prolongs the lifetime of sensors. To this end, a sensor selection method for distributed speech enhancement is proposed. The best subset of microphone nodes is determined by maximizing the signal-to-noise ratio (SNR), while keeping the activated nodes connected with each other. The above criterion involves an integer and non-linear programming, which is linearized with multiple base-3 sub-optimization problems, and each of them is solved by a state-of-the-art steepest descent (SD) algorithm. In addition, a greedy searching strategy is presented to select sensors rapidly. Finally, a distributed SD algorithm is further derived, which is more suitable for distributed sensor networks. The proposed method can obtain the optimal subnetwork in noisy and reverberant environments. Unlike the existing approaches, it can select nodes from a microphone network with arbitrary communication graphs. Moreover, it requires only local communications among nodes without an external central processor. Experimental results confirm the validity of the proposed method. De Hu, Qintuya Si, Rui Liu 0008, Feilong Bao |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | Distributed Sampling Rate Offset Estimation Over Acoustic Sensor Networks Based on Asynchronous Network Newton OptimizationabstractSampling rate synchronization is an inevitable issue in distributed acoustic sensor networks. In this paper, an analytical sampling rate offset (SRO) estimation approach is first proposed, and then, it is extended to a distributed method that suitable for acoustic sensor networks with arbitrary communication graphs. Specifically, a linear-phase drift model in the short-time Fourier transform domain is used to approximate the SRO between each pair of microphone nodes. Next, after unwrapping the temporally averaged phase information, SROs are recovered analytically via a new weighted-sum criterion. Based on this, a distributed cost function is established at each node to obtain the SROs of all nodes simultaneously in a distributed manner. Finally, a state-of-the-art distributed algorithm named asynchronous network Newton optimization is adopted to carry out the distributed SRO estimation. The proposed method can effectively estimate the SROs among acoustic sensor nodes in noisy and reverberant environments. Compared with the existing approaches, it does not require an external central processor, and only local communications among nodes are needed. Experimental results confirm the validity of the proposed method. De Hu, Huaiwen Zhang, Feilong Bao, Rui Wang 0046 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | End-to-End Large-Scale Image Retrieval Network with Convolution and Vision Transformers
Feilong Bao, Xiangdong Su, Weihua Wang 0006, Guanglai Gao |
ICANN (4) | 2 |
| 2022 | Alignment-Learning Based Single-Step Decoding for Accurate and Fast Non-Autoregressive Speech RecognitionabstractNon-autoregressive transformer (NAT) based speech recognition models have gained more and more attention since they perform faster inference speed compared with autoregressive counterparts, especially when the single-step decoding is applied. However, the single-step decoding process with length prediction will suffer from the decoding stability problem and limited improvement for inference speed. To address this, in this paper, we propose an alignment learning based NAT model, named AL-NAT. Our idea is inspired by the fact that the encoder CTC output and the target sequence are monotonically related. Specifically, we design an alignment cost matrix between the CTC output tokens and the target tokens and define a novel alignment loss to minimize the distance between the alignment cost matrix and the ground truth monotonic alignment path. By eliminating the length prediction mechanism, our AL-NAT model achieves remarkable improvements in recognition accuracy and decoding speed. To learn the contextual knowledge to improve the decoding accuracy, we further add lightweight language model on both the encoder and decoder side. Our proposed method achieves WERs of 2.8%/6.3% and RTF of 0.011 on Librispeech test clean/other sets with a lightweight 3-gram LM, and a CER of 5.3% and RTF of 0.005 on Aishell1 without LM, respectively. Yonghe Wang, Rui Liu 0008, Feilong Bao, Hui Zhang 0031, Guanglai Gao |
ICASSP | 3 |
| 2022 | A Deep Investigation of RNN and Self-attention for the Cyrillic-Traditional Mongolian Bidirectional Conversion
Muhan Na, Rui Liu 0008, Feilong Bao, Guanglai Gao |
ICONIP (6) | 3 |
| 2021 | Joint Alignment Learning-Attention Based Model for Grapheme-to-Phoneme ConversionabstractSequence-to-sequence attention-based models for grapheme-to-phoneme (G2P) conversion have gained significant interests. The attention-based encoder-decoder framework learns the mapping of input to output tokens by selectively focusing on relevant information, and has been shown well performance. However, the attention mechanism can result in non-monotonic alignments, resulting in poor G2P conversion performance. In this paper, we present a novel approach to optimize the G2P conversion model directly alignment grapheme-phoneme sequence by using alignment learning (AL) as the loss function. Besides, we propose a multi-task learning method that uses a joint alignment learning model and attention model to predict the proper alignments and thus improve the accuracy of G2P conversion. Evaluations on Mongolian and CMUDict tasks show that alignment learning as the loss function can effectively train G2P conversion model. Further, our multi-task method can significantly outperform both the alignment learning-based model and attention-based model. Yonghe Wang, Feilong Bao, Hui Zhang 0031, Guanglai Gao |
ICASSP | 2 |
| 2021 | MCDALNet: Multi-scale Contextual Dual Attention Learning Network for Medical Image SegmentationabstractMedical image segmentation has been widely studied, and many methods have been proposed. Among the existing methods, U-Net and its variants have achieved a promising performance. However, these methods miss certain areas because they only generate fixed-scale receptive fields in each layer of the encoder and cannot establish rich contextual dependencies on the fusion features in the decoder. To solve these problems, this paper proposes a multi-scale contextual dual attention learning network (named MCDALNet) to capture multi-scale information and the dependencies of spatial and channel features. MCDALNet contains two components: an encoder with three multi-scale contextual learning (MCL) modules and a decoder with three dual attention modules. The MCL module extracts multi-scale context information from low-level features through the split-transform-merge-residual architecture. The dual attention module consists of a position attention sub-module and a channel attention submodule, which improve the feature representation and help the medical image segmentation. The position attention submodule captures spatial dependencies by learning similar spatial features, and the channel attention sub-module captures channel dependencies by learning relevant features on the channel maps. Experiment results show that our approach achieves significant improvement in medical image segmentation and outperforms the representative deep learning models on public datasets. Xiangdong Su, Feilong Bao |
IJCNN | 4 |
| 2021 | Panoptic-DLA: Document Layout Analysis of Historical Newspapers Based on Proposal-Free Panoptic Segmentation Model
Feilong Bao, Guanglai Gao |
KSEM | 2 |
| 2021 | Soft-BAC: Soft Bidirectional Alignment Cost for End-to-End Automatic Speech Recognition
Yonghe Wang, Hui Zhang 0031, Feilong Bao, Guanglai Gao |
PRICAI (2) | 3 |
| 2021 | Exploiting Morphological and Phonological Features to Improve Prosodic Phrasing for Mongolian Speech SynthesisabstractProsodic phrasing is an important factor that affects naturalness and intelligibility in text-to-speech synthesis. Studies show that deep learning techniques improve prosodic phrasing when large text and speech corpus are available. However, for low-resource languages, such as Mongolian, prosodic phrasing remains a challenge for various reasons. First, the database suitable for system training is limited. Second, word composition knowledge that is prosody-informing has not been used in prosodic phrase modeling. To address these problems, in this article, we propose a feature augmentation method in conjunction with a self-attention neural classifier. We augment input text with morphological and phonological decompositions of words to enhance the text encoder. We study the use of self-attention classifier, that makes use of global context of a sentence, as a decoder for phrase break prediction. Both objective and subjective evaluations validate the effectiveness of the proposed phrase break prediction framework, that consistently improves voice quality in a Mongolian text-to-speech synthesis system. Rui Liu 0008, Berrak Sisman, Feilong Bao, Guanglai Gao, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Incorporating Inner-word and Out-word Features for Mongolian Morphological SegmentationabstractMongolian morphological segmentation is regarded as a crucial preprocessing step in many Mongolian related NLP applications and has received extensive attention.Recently, end-to-end segmentation approaches with long short-term memory networks (LSTM) have achieved excellent results.However, the inner-word features among characters in the word and the out-word features from context are not well utilized in the segmentation process.In this paper, we propose a neural network incorporating inner-word and out-word features for Mongolian morphological segmentation.The network consists of two encoders and one decoder.The inner-word encoder uses the self-attention mechanisms to capture the inner-word features of the target word.The out-word encoder employs a two layers BiLSTM network to extract out-word features in the sentence.Then, the decoder adopts a multi-head double attention layer to fuse the inner-word features and out-word features and produces the segmentation result.The evaluation experiment compares the proposed network with the baselines and explores the effectiveness of the sub-modules. Xiangdong Su, Guanglai Gao, Feilong Bao |
COLING | 5 |
| 2020 | Teacher-Student Training For Robust Tacotron-Based TTSabstractWhile neural end-to-end text-to-speech (TTS) is superior to conventional statistical methods in many ways, the exposure bias problem in the autoregressive models remains an issue to be resolved. The exposure bias problem arises from the mismatch between the training and inference process, that results in unpredictable performance for out-of-domain test data at run-time. To overcome this, we propose a teacher-student training scheme for Tacotron-based TTS by introducing a distillation loss function in addition to the feature loss function. We first train a Tacotron2-based TTS model by always providing natural speech frames to the decoder, that serves as a teacher model. We then train another Tacotron2-based model as a student model, of which the decoder takes the predicted speech frames as input, similar to how the decoder works during run-time inference. With the distillation loss, the student model learns the output probabilities from the teacher model, that is called knowledge distillation. Experiments show that our proposed training scheme consistently improves the voice quality for out-of-domain test data both in Chinese and English systems. Rui Liu 0008, Berrak Sisman, Jingdong Li, Feilong Bao, Guanglai Gao, Haizhou Li 0001 |
ICASSP | 4 |
| 2020 | A Multi-Scaled Receptive Field Learning Approach for Medical Image SegmentationabstractBiomedical image segmentation has been widely studied, and lots of methods have been proposed. Among these methods, attention U-Net has achieved a promising performance. However, it has drawbacks of extracting the multi-scaled receptive field features at the high-level feature maps, resulting in the degeneration when dealing with the lesions with apparent scale variations. To solve this problem, this paper integrates an atrous spatial pyramid pooling (ASPP) module in the contracting path of attention U-Net. This module employs multiple dilation rates for the purpose of obtaining several multi-scale receptive fields, which significantly improves the networks' ability to handle both large and small lesions. Evaluation experimental result shows that our approach significantly improves the performance of medical image segmentation and substantially outperforms the representative deep learning models on public datasets. Xiangdong Su, Feilong Bao |
ICASSP | 5 |
| 2020 | Masking and Inpainting: A Two-Stage Speech Enhancement Approach for Low SNR and Non-Stationary NoiseabstractCurrently, low signal-to-noise ratio (SNR) and non-stationary noise cause severe performance degradation for most of speech enhancement models. For better speech enhancement at the above scenarios, this paper proposes a two-stage approach that consists of binary masking and spectrogram inpainting. In the binary masking stage, we first obtain binary mask by hardening soft mask and then use it to remove time-frequency points that are dominated by severe noise. In the spectrogram inpainting stage, we use a CNN with partial convolution to perform inpainting on the masked spectrogram from the previous stage. We compared our approach with two powerful baselines, including Wave-U-Net and CRN, on a low SNR dataset containing lots of non-stationary noises. The experimental results show that our approach outperformed the baselines and achieved the state-of-the-art performance. Xiangdong Su, Shixue Wen, Yiqian Pan, Feilong Bao |
ICASSP | 6 |
| 2020 | Modeling Prosodic Phrasing With Multi-Task Learning in Tacotron-Based TTSabstractTacotron-based end-to-end speech synthesis has shown remarkable voice quality. However, the rendering of prosody in the synthesized speech remains to be improved, especially for long sentences, where prosodic phrasing errors can occur frequently. In this letter, we extend the Tacotron-based speech synthesis framework to explicitly model the prosodic phrase breaks. We propose a multi-task learning scheme for Tacotron training, that optimizes the system to predict both Mel spectrum and phrase breaks. To our best knowledge, this is the first implementation of multi-task learning for Tacotron based TTS with a prosodic phrasing model. Experiments show that our proposed training scheme consistently improves the voice quality for both Chinese and Mongolian systems. Rui Liu 0008, Berrak Sisman, Feilong Bao, Guanglai Gao, Haizhou Li 0001 |
IEEE Signal Process. Lett. | 3 |
| 2019 | Building Mongolian TTS Front-End with Encoder-Decoder Model by Using Bridge Method and Multi-view Features
Rui Liu 0008, Feilong Bao, Guanglai Gao |
ICONIP (5) | 2 |
| 2019 | Morphological Knowledge Guided Mongolian Constituent Parsing
Xiangdong Su, Guanglai Gao, Feilong Bao |
ICONIP (3) | 4 |
| 2019 | Learning an Adversarial Network for Speech Enhancement Under Extremely Low Signal-to-Noise Ratio Condition
Xiangdong Su, Huali Xu, Tongyang Liu, Guanglai Gao, Feilong Bao |
ICONIP (1) | 8 |
| 2019 | A Natural Scene Text Extraction Approach Based on Generative Adversarial Learning
Huali Xu, Xiangdong Su, Tongyang Liu, Guanglai Gao, Feilong Bao |
ICONIP (1) | 6 |
| 2019 | Neural Morphological Segmentation Model for MongolianabstractMorphological segmentation is useful for processing Mongolian. In this paper, we manually build a morphological segmentation data set for Mongolian. We then present a character-based encoder-decoder model with attention mechanism to perform the morphological segmentation task. We further investigate the influence of analogy features extracted from scratch and improve the performance of our model using multi languages setting. Experimental results show that our encoder-decoder model with attention mechanism provides a strong baseline for Mongolian morphological segmentation. The analogy features provide useful information to the model and improve the performance of the system. The use of multi languages data set shows the capability of our model to acquire knowledge through different languages and delivers the best result. Weihua Wang 0006, Rashel Fam, Feilong Bao, Yves Lepage, Guanglai Gao |
IJCNN | 3 |
| 2019 | An Automatic Spelling Correction Method for Classical Mongolian
Feilong Bao, Guanglai Gao, Weihua Wang 0006, Hui Zhang 0031 |
KSEM (2) | 2 |
| 2019 | A Context-Free Spelling Correction Method for Classical Mongolian
Feilong Bao, Guanglai Gao |
NLPCC (2) | 2 |
| 2019 | Research on Khalkha Dialect Mongolian Speech Recognition Acoustic Model Based on Weight Transfer
Linyan Shi, Feilong Bao, Yonghe Wang, Guanglai Gao |
NLPCC (2) | 2 |
| 2019 | End-to-End Model for Offline Handwritten Mongolian Word Recognition
Hongxi Wei, Hui Zhang 0031, Feilong Bao, Guanglai Gao |
NLPCC (2) | 4 |
| 2019 | Learning Morpheme Representation for Mongolian Named Entity Recognition
Weihua Wang 0006, Feilong Bao, Guanglai Gao |
Neural Process. Lett. | 2 |
| 2018 | A LSTM Approach with Sub-Word Embeddings for Mongolian Phrase Break PredictionabstractIn this paper, we first utilize the word embedding that focuses on sub-word units to the Mongolian Phrase Break (PB) prediction task by using Long-Short-Term-Memory (LSTM) model. Mongolian is an agglutinative language. Each root can be followed by several suffixes to form probably millions of words, but the existing Mongolian corpus is not enough to build a robust entire word embedding, thus it suffers a serious data sparse problem and brings a great difficulty for Mongolian PB prediction. To solve this problem, we look at sub-word units in Mongolian word, and encode their information to a meaningful representation, then fed it to LSTM to decode the best corresponding PB label. Experimental results show that the proposed model significantly outperforms traditional CRF model using manually features and obtains 7.49% F-Measure gain. Rui Liu 0008, Feilong Bao, Guanglai Gao, Hui Zhang 0031, Yonghe Wang |
COLING | 2 |
| 2018 | Mongolian Word Segmentation Based on Three Character Level Seq2Seq Models
Xiangdong Su, Guanglai Gao, Feilong Bao |
ICONIP (5) | 4 |
| 2018 | Improving Mongolian Phrase Break Prediction by Using Syllable and Morphological Embeddings with BiLSTM Model
Rui Liu 0008, Feilong Bao, Guanglai Gao, Hui Zhang 0031, Yonghe Wang |
INTERSPEECH | 2 |
| 2018 | Mongolian Grapheme to Phoneme Conversion by Using Hybrid Approach
Zhinan Liu, Feilong Bao, Guanglai Gao, Suburi |
NLPCC (1) | 2 |
| 2018 | Phonologically Aware BiLSTM Model for Mongolian Phrase Break Prediction with Attention Mechanism
Rui Liu 0008, Feilong Bao, Guanglai Gao, Hui Zhang 0031, Yonghe Wang |
PRICAI (1) | 2 |
| 2017 | Segmentation-Free Printed Traditional Mongolian OCR Using Sequence to Sequence with Attention ModelabstractMongolian Optical Character Recognition (OCR) systems are required for printed document digitization and Mongolian cultural resources utilization. Existing Mongolian OCR systems are based on segmentation. But, the Mongolian segmentation is more difficult than other languages. So, these methods are highly costly and error suffering. In this study, a segmentation-free based traditional Mongolian word recognition method is proposed. Specifically, we formalize the OCR task as a sequence to sequence mapping problem, in which the input Mongolian word image and the output textual string are treated as a sequence of image frames and a sequence of letters, respectively. A sequence to sequence with attention model is adopted to solve this problem. Experimental results on a dataset show the effectiveness of the proposed method. Hui Zhang 0031, Hongxi Wei, Feilong Bao, Guanglai Gao |
ICDAR | 3 |
| 2017 | Research on Mongolian Speech Recognition Based on FSMN
Yonghe Wang, Feilong Bao, Guanglai Gao |
NLPCC | 2 |
| 2016 | Mongolian Named Entity Recognition System with Rich FeaturesabstractIn this paper, we first build a manually annotated named entity corpus of Mongolian. Then, we propose three morphological processing methods and study comprehensive features, including syllable features, lexical features, context features, morphological features and semantic features in Mongolian named entity recognition. Moreover, we also evaluate the influence of word cluster features on the system and combine all features together eventually. The experimental result shows that segmenting each suffix into an individual token achieves better results than deleting suffixes or using the suffixes as feature. The system based on segmenting suffixes with all proposed features yields benchmark result of F-measure=84.65 on this corpus. Weihua Wang 0006, Feilong Bao, Guanglai Gao |
COLING | 2 |
| 2016 | Mongolian Named Entity Recognition with Bidirectional Recurrent Neural NetworksabstractTraditional approaches to Named Entity Recognition almost heavily rely on feature engineering. In this paper, we introduce a kind of bidirectional recurrent neural network with long short memory (BLSTM) to capture bidirectional and long dependencies in a sentence without any feature set. Our model combines BLSTM network with Conditional Random Field (CRF) layer to jointly decode the best output. Additionally, this model inputs the concatenation of Mongolian morpheme and character representation. Experimental results show that the bidirectional recurrent neural networks significantly outperform traditional CRF model using manual features. Weihua Wang 0006, Feilong Bao, Guanglai Gao |
ICTAI | 2 |
| 2016 | A knowledge-based recognition system for historical Mongolian documents
Xiangdong Su, Guanglai Gao, Hongxi Wei, Feilong Bao |
Int. J. Document Anal. Recognit. | 4 |
| 2015 | Enhancing the Mongolian Historical Document Recognition System with Multiple Knowledge-Based Strategies
Xiangdong Su, Guanglai Gao, Hongxi Wei, Feilong Bao |
ICONIP (2) | 4 |
| 2015 | Mongolian Inflection Suffix Processing in NLP: A Case StudyabstractInflection suffix is an important morphological characteristic of Mongolian words, since the suffixes express abundant syntactic and semantic meanings. In order to provide an informative introduction of it, this paper implements a case study on it. Through three Mongolian NLP tasks, we disclose the following information: (1) views of inflection suffix in NLP tasks, (2) Inflection suffix processing ways, (3) Inflection suffix effects on system performance and (4) some suffix related conclusion. Xiangdong Su, Guanglai Gao, Jing Wu 0011, Feilong Bao |
NLPCC | 5 |
| 2013 | Segmentation-based Mongolian LVCSR approachabstractMongolian is an agglutinative language. Each root can be followed by several suffixes to formulate new words. This special word formation characteristic results in probably millions of Mongolian words, which is far beyond the coverage of the pronunciation dictionary of any current Mongolian speech recognition system. Moreover, even if the pronunciation dictionary is large enough to cover all of the Mongolian words, the recognition system still cannot perform well due to the problem of sample sparseness. In this paper, we propose a segmentation-based Mongolian Large Vocabulary Continuous Speech Recognition (LVCSR) approach and rebuild the corresponding acoustic model and language model. Experimental results show that, by converting most of these words into their corresponding In-Vocabulary form, the proposed approach effectively recognizes most of the Mongolian words and greatly improves the sample sparseness problem in the language model. Feilong Bao, Guanglai Gao, Xueliang Yan, Weihua Wang 0006 |
ICASSP | 1 |
| 2013 | Language Model for Cyrillic Mongolian to Traditional Mongolian Conversion
Feilong Bao, Guanglai Gao, Xueliang Yan |
NLPCC | 1 |