VLDB 2026 Research / reviewers in the wild / expert
Guanglai Gao
dblp:72/6902
· DBLP profile ↗
123ranked-venue papers
2as first author
60since 2021 · last 2026
0009-0005-5513-1192ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 90 · 1 first-author · 44 since 2021Graphics, computer vision, multimedia, augmented reality and games · 37 · 18 since 2021Databases, data management, data science and information retrieval · 14 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 5 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FlorE: Integrating Full Lorentz Group and Directional Offsets for Effective Knowledge Graph EmbeddingabstractKnowledge Graph Embedding (KGE) aims to map entities and relationships into a continuous vector space to facilitate reasoning and downstream tasks. Although previous KGE methods based on Euclidean, complex spaces, or hyperbolic spaces have performed well, they still struggle to effectively model Z-Paradox relation patterns which account for a large proportion in each knowledge graph. To address this issue, we propose a novel KGE method **FlorE** which integrates full Lorentz Group and directional offset operation in hyperbolic space for KGE task. Specifically, we incorporates the full Lorentz Group to enable the same relation in knowledge graph (KG) to perform indefinite isometry, thus avoiding the overlapping of entities. Meanwhile, we implement directional offset operation via exponential mapping to transform the relations to the same Lorentz manifold of the entities, thus maintaining geometric consistency for the relations and entities in KG. By integrating these two techniques, FlorE can effectively model the Z-Paradox relation patterns and improve the representation learning ability for KGs. Experiments on the five benchmark datasets demonstrate that our method achieves state-of-the-art performance. For the Z-Paradox relation patterns, the improvement achieves **26.7%**, **15.6%**, **35.4%**, **33.7%**, and **31.5%** on FB15k-237, WN18RR, CoDEx-S, CoDEx-M and CoDEx-L, respectively. Zehua Duo, Jiang Li 0013, Xiangdong Su, Guanglai Gao |
AAAI | 4 |
| 2026 | CEDAR: A Chinese Evaluation Dataset for Computational ArgumentationabstractTian Lan, Jiang Li, Rong Yan, Feilong Bao, Weihua Wang, Guanglai Gao, Xiangdong Su. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jiang Li 0013, Feilong Bao, Weihua Wang 0006, Guanglai Gao, Xiangdong Su |
ACL (1) | 6 |
| 2026 | Who Wrote This Line? Evaluating the Detection of LLM-Generated Classical Chinese PoetryabstractJiang Li, Tian Lan, Shanshan Wang, Dongxing Zhang, Dianqing Lin, Guanglai Gao, Derek F. Wong, Xiangdong Su. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jiang Li 0013, Shanshan Wang 0009, Zdongxing, Dianqing Lin, Guanglai Gao, Derek F. Wong, Xiangdong Su |
ACL (1) | 6 |
| 2026 | Know Your Place: Diagnosing Implicit Social Adaptation Failures in Chinese Large Language ModelsabstractYu Tian, Jie Xing, Ziming Li, Jiang Li, Zehua Duo, Tian Lan, Xu Liu, Guanglai Gao, Xiangdong Su. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jiang Li 0013, Zehua Duo, Guanglai Gao, Xiangdong Su |
ACL (1) | 8 |
| 2026 | GSMP: Geometry-Structured Masked Pre-training with Multi-granularity Objectives and Curriculum Learning for Geometric Problem Solving
Xingxiang Zhou, Minzhi Zhang, Guanglai Gao, Xiangdong Su |
ICDAR (2) | 6 |
| 2026 | A knowledge prompt augmented lightweight multimodal language assistant for biomedicine
Lei Liu 0079, Xiangdong Su, Xingxiang Zhou, Guanglai Gao |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | Wisteria: A unified multi-scale feature learning framework for DNA language model
Weihua Wang 0006, Haoji Li, Feilong Bao, Guanglai Gao |
Pattern Recognit. | 5 |
| 2026 | Hyperbolic-Based Cross-Modal Semantic Remodeling Network for Zero-Shot Sketch-Based Image RetrievalabstractThe Zero-Shot Sketch-Based Image Retrieval (ZS-SBIR) task aims to retrieve images associated with sketches from unseen classes, bringing great convenience to the engineering field. To address the modality gap, most existing works project images and sketches into a shared Euclidean space. However, the hierarchical structure of image data makes the Euclidean space not the optimal choice as an embedding space for representing complex structured image data. Meanwhile, existing text and hierarchical models are not effective enough for addressing the problem of knowledge transfer. To address these issues, this article proposes an original Hyperbolic-Based Cross-Modal Semantic Remodeling Network (called HCMSN) for ZS-SBIR. Specifically, this article proposes to extract category-level word embeddings based on BERT model, then align image features and sketches with the word embeddings using adversarial methods. Meanwhile, this article further proposes a cross-modal retrieval feature reconstruction network for improving the informativeness and robustness of retrieval features. Moreover, this article presents a feature projection network that maps the retrieval features to the hyperbolic space to generate the hyperbolic retrieval features, thus effectively representing the data with hierarchical structure. Extensive experiments demonstrate that the mAP@all of our HCMSN model surpasses CNN-based models by 20.9% on the Sketchy dataset, 1.2% on the more difficult TU-Berlin dataset, and 13.6% on the more challenging QuickDraw dataset. Xiangdong Su, Feilong Bao, Guanglai Gao |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | SSAN: A Symbol Spatial-Aware Network for Handwritten Mathematical Expression RecognitionabstractThe great challenge of handwritten mathematical expression recognition (HMER) is the complex structures of the expressions, which are directly related to the symbol spatial positions. Existing HMER methods typically employ attention mechanisms in the decoder of their models to implicitly perceive the symbol positions, or employ symbol counting and tree-based strategies to model the symbol spatial relation. However, these methods still cannot effectively capture the structural information of formulas, thus negatively impacting the symbol decoding in HMER. To deal with this problem and enhance the HMER performance, this paper proposes a novel auxiliary task, namely predicting the symbol spatial distribution map of handwritten expression images. On such basis, this paper designs a symbol spatial-aware network (SSAN) for this task, which is jointly optimized with the HMER model. Specifically, considering the similarity of the symbol spatial positions between the handwritten mathematical expression images and their corresponding printed templates, we obtain the symbol spatial distribution map by first generating printed templates from LaTeX ground-truth for handwritten formula images and then replacing the connected components of printed templates with 2D Gaussian distribution maps of the same size. Meanwhile, due to the loose alignment of the symbol spatial positions between handwritten and printed formula images, and misclassification of similar symbols, we further propose a coarse-to-fine alignment strategy and an attention-guided symbol masking strategy in SSAN to tackle these issues. Extensive experiments demonstrate that SSAN significantly improves the recognition performance of the HMER models, and the proposed auxiliary tasks are more effective in enhancing HMER performance than existing auxiliary tasks. Xiangdong Su, Xingxiang Zhou, Guanglai Gao |
AAAI | 4 |
| 2025 | A Mutual Information Perspective on Knowledge Graph EmbeddingabstractKnowledge graph embedding techniques have emerged as a critical approach for addressing the issue of missing relations in knowledge graphs. However, existing methods often suffer from limitations, including high intra-group similarity, loss of semantic information, and insufficient inference capability, particularly in complex relation patterns such as 1-N and N-1 relations. To address these challenges, we introduce a novel KGE framework that leverages mutual information maximization to improve the semantic representation of entities and relations. By maximizing the mutual information between different components of triples, such as (h, r) and t, or (r, t) and h, the proposed method improves the model’s ability to preserve semantic dependencies while maintaining the relational structure of the knowledge graph. Extensive experiments on benchmark datasets demonstrate the effectiveness of our approach, with consistent performance improvements across various baseline models. Additionally, visualization analyses and case studies demonstrate the improved ability of the MI framework to capture complex relation patterns. Jiang Li 0013, Xiangdong Su, Zehua Duo, Xiaotao Guo, Guanglai Gao |
ACL (1) | 6 |
| 2025 | HiFusion-Pro: Geometric-Structural Aware Protein Function Prediction via Hierarchical Interaction Fusion and Multi-Task LearningabstractAccurate protein function prediction, such as assigning Enzyme Commission (EC) numbers and Gene Ontology (GO) terms, is a fundamental challenge in bioinformatics. We propose HiFusion-Pro, a deep learning framework that enhances prediction accuracy through two key innovations. First, we introduce a hierarchical fusion architecture that deeply integrates sequence and structure information. This is achieved by initializing a structural encoder (GearNet) with features from a protein language model (ESM) for early information sharing, and by aggregating multi-level intermediate representations from both encoders. This approach overcomes the limitations of conventional late-stage fusion. Second, we employ a multitask learning strategy that co-predicts protein function (primary task) with inter-residue distances (auxiliary task). This strategy regularizes the model, promoting robust feature learning that captures both global functional patterns and local geometric details. Experimental results demonstrate that HiFusion-Pro significantly outperforms state-of-the-art methods across multiple benchmarks. The codes for HiFusion-Pro and datasets are available at https://github.com/YuQing-cs/HiFusion-Pro. Zhongyu Hu, Juan Wang 0011, Guanglai Gao |
BIBM | 3 |
| 2025 | Distance-Adaptive Quaternion Knowledge Graph Embedding with Bidirectional RotationabstractQuaternion contains one real part and three imaginary parts, which provided a more expressive hypercomplex space for learning knowledge graph. Existing quaternion embedding models measure the plausibility of a triplet either through semantic matching or distance scoring functions. However, it appears that semantic matching diminishes the separability of entities, while the distance scoring function weakens the semantics of entities. To address this issue, we propose a novel quaternion knowledge graph embedding model. Our model combines semantic matching with entity’s geometric distance to better measure the plausibility of triplets. Specifically, in the quaternion space, we perform a right rotation on the head entity and a reverse rotation on the tail entity to learn the rich semantic features. Then, we utilize distance adaptive translations to learn the geometric distance between entities. Furthermore, we provide mathematical proofs to demonstrate our model can handle complex logical relationships. Extensive experimental results and analyses show our model significantly outperforms previous models on well-known knowledge graph completion benchmark datasets. Our code is available at https://anonymous.4open.science/r/l2730. Weihua Wang 0006, Qiuyu Liang, Feilong Bao, Guanglai Gao |
COLING | 4 |
| 2025 | Unifying Dual-Space Embedding for Entity Alignment via Contrastive LearningabstractEntity alignment (EA) aims to match identical entities across different knowledge graphs (KGs). Graph neural network-based entity alignment methods have achieved promising results in Euclidean space. However, KGs often contain complex local and hierarchical structures, which are hard to represent in a single space. In this paper, we propose a novel method named as UniEA, which unifies dual-space embedding to preserve the intrinsic structure of KGs. Specifically, we simultaneously learn graph structure embeddings in both Euclidean and hyperbolic spaces to maximize the consistency between embeddings in the two spaces. Moreover, we employ contrastive learning to mitigate the misalignment issues caused by similar entities, where embeddings of similar neighboring entities become too close. Extensive experiments on benchmark datasets demonstrate that our method achieves state-of-the-art performance in structure-based EA. Our code is available at https://github.com/wonderCS1213/UniEA. Cunda Wang, Weihua Wang 0006, Qiuyu Liang, Feilong Bao, Guanglai Gao |
COLING | 5 |
| 2025 | F²Bench: An Open-ended Fairness Evaluation Benchmark for LLMs with Factuality ConsiderationsabstractWarning: This paper contains content that may be offensive or harmful With the growing adoption of large language models (LLMs) in NLP tasks, concerns about their fairness have intensified.Yet, most existing fairness benchmarks rely on closed-ended evaluation formats, which diverge from realworld open-ended interactions.These formats are prone to position bias and introduce a "minimum score" effect, where models can earn partial credit simply by guessing.Moreover, such benchmarks often overlook factuality considerations rooted in historical, social, physiological, and cultural contexts, and rarely account for intersectional biases.To address these limitations, we propose F 2 Bench: an openended fairness evaluation benchmark for LLMs that explicitly incorporates factuality considerations.F 2 Bench comprises 2,568 instances across 10 demographic groups and two openended tasks.By integrating text generation, multi-turn reasoning, and factual grounding, F 2 Bench aims to more accurately reflect the complexities of real-world model usage.We conduct a comprehensive evaluation of several LLMs across different series and parameter sizes.Our results reveal that all models exhibit varying degrees of fairness issues.We further compare open-ended and closedended evaluations, analyze model-specific disparities, and provide actionable recommendations for future model development.Our code and dataset are publicly available at https: //github.com/VelikayaScarlet/F2Bench. Jiang Li 0013, Yemin Wang, Xiangdong Su, Guanglai Gao |
EMNLP | 6 |
| 2025 | Dynamic Structure Hypergraph for Document-level Event ExtractionabstractDocument-level Event Extraction (DEE) aims to identify event information from a given document. The two challenges of this task are the event arguments scattering across differrent sentences and the multiple events within a single document. In this paper, we propose a novel Dynamic Structure Hypergraph model to address the issue of limited global modeling capability in traditional graphs. Firstly, we construct a hypergraph to model the global interactions between different sentences and entities in a document. Then, new hyperedges are generated by constructing a mention-mention correlation matrix based on the updated node representations, which evolves the hypergraph into a dynamic structure. This will help the nodes to aware the contextual semantic information in time. Finally, extensive experiments and analysis demonstrate that our method has made significant improvements in addressing the two aforementioned challenges, which outperforms existing state-of-the-art models on two public datasets. Our code is available at https://github.com/1999rq/DSH. Qi Ren, Weihua Wang 0006, Jie Yu 0008, Guanglai Gao |
ICASSP | 4 |
| 2025 | Structural-Aware Disentangled Learning with CLIP for Hyperbolic Zero-Shot Sketch-Based Image RetrievalabstractThe zero-shot sketch-based image retrieval task faces two key challenges: domain gap and knowledge transfer. Our innovation is recognizing that directly aligning cross-domain features weakens the discriminative ability of the model, as it overlooks the asymmetry between sketches and images. Additionally, Euclidean space is inadequate for capturing the hierarchical structure, which limits the performance of the model on complex data. To address these issues, we propose a Structural-Aware Disentangled Learning network (termed SADLnet) that incorporates CLIP and hyperbolic geometry. Specifically, we use CLIP to extract visual features from each domain to enhance the domain generalization of the model. Furthermore, we design a structure-guided disentanglement strategy to decompose image representations into sketch-related and sketch-unrelated features, addressing the domain gap. Moreover, we project the retrieval features into hyperbolic space to capture hierarchical information, improving feature discrimination in retrieval tasks. Extensive experiments demonstrate that SADLnet establishes new state-of-the-art performance on three datasets. Feilong Bao, Xiangdong Su, Guanglai Gao |
ICASSP | 5 |
| 2025 | Task-Decoupled Bézier Surface Constraint for Uneven Low-Light Image Enhancement
Xingxiang Zhou, Xiangdong Su, Guanglai Gao |
ICCV | 5 |
| 2025 | Enhancing Mandarin Lip Reading with a Multimodal Conformer and Structured State-Space DecoderabstractLip reading is a visual recognition technology that interprets spoken content by decoding lip movements. Since speech perception is inherently a multimodal task, incorporating audio information during training is crucial to assist lip reading. This paper proposes a novel architecture that combines the Conformer network with a structured state space decoder under multimodal input to enhance Mandarin lip reading capabilities. As a tonal language, Mandarin benefits from audio cues that guide visual information learning, improving the accuracy of speech content recognition. Our approach leverages the Conformer to extract shared semantics from both audio and video, and employs a bidirectional structured state space decoder to decode, effectively capturing the temporal dynamics and complex dependencies of long sequences. This method achieved CERs of 54.97% and 12.53% in the CN-CVS and CMLR datasets, respectively. The research code is open source at: https://anonymous.4open.science/r/Lip-reading-model-D7B8. Meng Miao, Feilong Bao, Guanglai Gao |
IJCNN | 3 |
| 2025 | MedVSA: Medical Visual Spoken-Question AnsweringabstractWith the rapid advancement of technology, smart healthcare has made significant progress, particularly in medical visual question answering (MedVQA). However, current MedVQA primarily relies on text, whereas practical applications often involve spoken interactions, such as in medical consultations and mobile-based queries. To bridge this gap, we propose a novel task, medical visual spoken-question answering (MedVSA), extending the conventional medical image and text-based question-answering paradigm to include spoken interactions for enhanced applicability. We expand upon four commonly used MedVQA datasets, namely VQA-RAD, SLAKE, PathVQA, and OVQA, by leveraging Alibaba Cloud speech synthesis technology to convert text questions into spoken questions. Various strategies are incorporated to ensure diverse and realistic speech synthesis. The resulting dataset comprises images, corresponding text, and synthesized speech data. Subsequently, we design a single-stage model and a two-stage model to tackle the MedVSA task. For the single-stage model, we directly input the speech and images into our designed whisper self-distillation model to obtain the results. For the two-stage model, we first use the Whisper model to convert the speech into text, then input the converted text and medical images into the self-distillation model to obtain the results. We provide two solutions for the MedVSA task and establish two baselines. Experimental results show that the two-stage model significantly outperforms the single-stage model, indicating that text conversion is crucial for solving MedVSA. This study advances smart healthcare developments by proposing MedVSA and designing two baselines tailored to its specificities. Source code and MedVSA dataset are available at https://github.com/Alivelei/MedVSA. Lei Liu 0079, Xiangdong Su, Guanglai Gao |
ICMR | 3 |
| 2025 | Fourier Self-Adaptation for Transferring General Pretrained Models to Specific DomainsabstractWhile pre-trained models in the general domain have proliferated, existing methods for transferring these models to specific domains often depend on source domain data for distribution alignment and are typically tailored for single tasks. We propose a source data-free approach, Fourier Self-Adaptation (FSA), which effectively adapts general models to a wide range of specific domains. Our method leverages the distinct properties of Fourier phase and amplitude: phase contains high-level structural and positional information, which is less affected by domain shifts, while amplitude contains details and brightness information, which is more affected by domain shifts. FSA adjusts the image distribution by initializing a trainable adaptive image from a normal distribution. It then interpolates the amplitude of the target domain image with that of the adaptive image, where the interpolation ratio is dynamically controlled by learnable weight and bias. During training, the model captures advanced phase information of the target image and refines the data distribution through amplitude interpolation. Additionally, a dual regularization loss constrains the model representation, encouraging it to focus on the intrinsic relationships of the target domain data while discarding irrelevant knowledge. We evaluate FSA using general pre-trained models on 11 unimodal image classification datasets and 6 multimodal visual question answering datasets, covering specific domains such as radiology, pathology, remote sensing, and art. Our method consistently achieves state-of-the-art performance across multiple datasets, with performance improvements ranging from 1% to 8% compared to basic pre-trained models. Source code are available at https://github.com/Alivelei/FSA. Lei Liu 0079, Xiangdong Su, Guanglai Gao |
ACM Multimedia | 3 |
| 2025 | Mitigating Heterogeneity among Factor Tensors via Lie Group Manifolds for Tensor Decomposition Based Temporal Knowledge Graph EmbeddingabstractJiang Li, Xiangdong Su, Guanglai Gao. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Jiang Li 0013, Xiangdong Su, Guanglai Gao |
NAACL (Long Papers) | 3 |
| 2025 | Local and global structure-aware contrastive framework for entity alignment
Cunda Wang, Weihua Wang 0006, Qiuyu Liang, Guanglai Gao |
Neurocomputing | 4 |
| 2025 | Domain disentanglement and fusion based on hyperbolic neural networks for zero-shot sketch-based image retrieval
Xiangdong Su, Yonghe Wang, Feilong Bao, Guanglai Gao |
Inf. Process. Manag. | 6 |
| 2024 | Leveraging Convolutional Models as Backbone for Medical Visual Question AnsweringabstractConvolutional neural networks (CNNs) have made significant contributions to computer vision and offer the advantages of higher training efficiency and lower model complexity. However, their application as the backbone in medical visual question answering (MedVQA) remains an open question. To address this issue, we employ popular convolutional models, including ResNet, DenseNet, and ShuffleNet, as the foundation for MedVQA, achieving outstanding performance. Different backbones can be tailored to diverse real-world scenarios. The central challenge in utilizing CNNs for visual question answering is effectively managing textual features and integrating multi-modal information. To overcome this challenge, we design a novel global interaction attention (GIA) that facilitates efficient interactions between text and image features. Additionally, we utilize the dot product before the classifier output to enhance visual and textual modal fusions. To further enhance model performance, we propose a novel multi-modal hidden mixup (MHidMix) technique for data augmentation, which involves interpolating hidden states during model training. This data augmentation technique smoothes the decision boundary without the need for complex sample selection, further improving model performance. Experimental results underscore the versatility of our proposed framework across various convolutional models, leading to outstanding performance on four MedVQA datasets. Notably, we achieved an accuracy increase of 9.4% on the PathVQA dataset and 4.5% on the OVQA dataset. Lei Liu 0079, Xiangdong Su, Guanglai Gao |
BIBM | 3 |
| 2024 | Optimizing Transformer and MLP with Hidden States Perturbation for Medical Visual Question AnsweringabstractOptimizing model performance is a crucial objective in medical visual question answering (MedVQA), and a wide range of techniques have been developed to achieve this goal. In this paper, we propose a novel technique called network state perturbation, which distinguishes itself from existing research. Our study designs four innovative methods to modify the hidden states within the network and improve the performance of multimodal models in the MedVQA task. Specifically, we evaluate these methods on both general transformer and multilayer perceptron (MLP) models, which allow for hidden state adjustments at each layer. The four introduced methods are as follows: (1) Randomly Set Zero, which assigns zeros to the hidden states of different modalities in the network; (2) Randomly Replace Content, which performs interpolation between the hidden states of the text sequence and the image sequence; (3) Randomly Add Gaussian Noise, which adds Gaussian noise to the hidden states of different modalities; and (4) Pair Interpolation, which interpolates the hidden states of different modalities of the current sample with those of other samples and performs corresponding label interpolation to facilitate model training. Our experimental results demonstrate that the proposed hidden state perturbation methods significantly enhance the performance of various transformer and MLP models on multiple MedVQA datasets, without requiring additional data or computational resources during training. These findings highlight the potential of hidden state perturbation as a novel model improvement technique for multimodal models in the medical domain. Lei Liu 0079, Xiangdong Su, Guanglai Gao |
BIBM | 3 |
| 2024 | Learning Frequency Adaptation for Cross-domain Medical Image SegmentationabstractIn medical image segmentation, some recent methods improve domain adaptation performance through frequency domain adaptation and frequency mixup. However, these approaches have two limitations: (1) frequency domain adaptation ignores the adverse effects of high-frequency noise on model generalization, and (2) frequency mixup confuses semantic information. To address these issues, we propose a novel frequency adaptation approach for medical image segmentation including low-frequency component alignment (LFCA) and random amplitude cutmix (RAC). Since low-frequency contains major image information and high-frequency noise affects adaptation, we leverage discrete wavelet transform to decompose images into low and high-frequency components. LFCA aligns the domain distribution of low frequencies and high frequencies are passed to the decoder via skip connections. In addition, we design RAC to generate diverse augmented samples through amplitude cutmix while avoiding distortion of the original distributions. Experiments on benchmark datasets validate the efficacy of our proposed approach. Lei Liu 0079, Xiangdong Su, Guanglai Gao |
BIBM | 5 |
| 2024 | FSAM: Fine-tuning SAM encoder and decoder for Medical Image SegmentationabstractRecently, the Segment Anything Model (SAM), a large pre-training model, has achieved excellent results on natural image segmentation tasks and received extensive attention. However, SAM on medical image segmentation is unsatisfactory since there are significant differences between natural and medical images. How to extend the SAM’s powerful segmentation capabilities to the medical domain requires further exploration. To this end, we propose a simple and efficient fine-tuning approach for SAM that does not require large-scale data called FSAM. Specifically, FSAM simultaneously fine-tunes the encoder and decoder while freezing the prompt encoder. This allows the encoder to extract medical image features effectively and guide decoder segmentation predictions. FSAM achieves state-of-the-art results on eight public medical image datasets, outperforming SAM by +15.39%. Moreover, FSAM exhibits weaker sample scale dependence. The proposed framework further improves the segmentation capabilities of SAM in medical images. Lei Liu 0079, Xiangdong Su, Guanglai Gao |
BIBM | 5 |
| 2024 | TransERR: Translation-based Knowledge Graph Embedding via Efficient Relation RotationabstractThis paper presents a translation-based knowledge geraph embedding method via efficient relation rotation (TransERR), a straightforward yet effective alternative to traditional translation-based knowledge graph embedding models. Different from the previous translation-based models, TransERR encodes knowledge graphs in the hypercomplex-valued space, thus enabling it to possess a higher degree of translation freedom in mining latent information between the head and tail entities. To further minimize the translation distance, TransERR adaptively rotates the head entity and the tail entity with their corresponding unit quaternions, which are learnable in model training. We also provide mathematical proofs to demonstrate the ability of TransERR in modeling various relation patterns, including symmetry, antisymmetry, inversion, composition, and subrelation patterns. The experiments on 10 benchmark datasets validate the effectiveness and the generalization of TransERR. The results also indicate that TransERR can better encode large-scale datasets with fewer parameters than the previous translation-based models. Our code and datasets are available at https://github.com/dellixx/TransERR. Jiang Li 0013, Xiangdong Su, Fujun Zhang 0004, Guanglai Gao |
LREC/COLING | 4 |
| 2024 | L\²GC: Lorentzian Linear Graph Convolutional Networks for Node Classification
Qiuyu Liang, Weihua Wang 0006, Feilong Bao, Guanglai Gao |
LREC/COLING | 4 |
| 2024 | EpLSA: Synergy of Expert-prefix Mixtures and Task-Oriented Latent Space Adaptation for Diverse Generative ReasoningabstractExisting models for diverse generative reasoning still struggle to generate multiple unique and plausible results. Through an in-depth examination, we argue that it is critical to leverage a mixture of experts as prefixes to enhance the diversity of generated results and make task-oriented adaptation in the latent space of the generation models to improve the quality of the responses. At this point, we propose EpLSA, an innovative model based on the synergy of expert-prefix mixtures and task-oriented latent space adaptation for diverse generative reasoning. Specifically, we use expert-prefixes mixtures to encourage the model to create multiple responses with different semantics and design a loss function to address the problem that the semantics is interfered by the expert-prefixes. Meanwhile, we design a task-oriented adaptation block to make the pre-trained encoder within the generation model more effectively adapted to the pre-trained decoder in the latent space, thus further improving the quality of the generated text. Extensive experiments on three different types of generative reasoning tasks demonstrate that EpLSA outperforms existing baseline models in terms of both the quality and diversity of the generated outputs. Our code is publicly available at https://github.com/IMU-MachineLearningSXD/EpLSA. Fujun Zhang 0004, Xiangdong Su, Jiang Li 0013, Guanglai Gao |
LREC/COLING | 5 |
| 2024 | Exploring the Synergy of Dual-path Encoder and Alignment Module for Better Graph-to-Text GenerationabstractThe mainstream approaches view the knowledge graph-to-text (KG-to-text) generation as a sequence-to-sequence task and fine-tune the pre-trained model (PLM) to generate the target text from the linearized knowledge graph. However, the linearization of knowledge graphs and the structure of PLMs lead to the loss of a large amount of graph structure information. Moreover, PLMs lack an explicit graph-text alignment strategy because of the discrepancy between structural and textual information. To solve these two problems, we propose a synergetic KG-to-text model with a dual-path encoder, an alignment module, and a guidance module. The dual-path encoder consists of a graph structure encoder and a text encoder, which can better encode the structure and text information of the knowledge graph. The alignment module contains a two-layer Transformer block and an MLP block, which aligns and integrates the information from the dual encoder. The guidance module combines an improved pointer network and an MLP block to avoid error-generated entities and ensures the fluency and accuracy of the generated text. Our approach obtains very competitive performance on three benchmark datasets. Our code is available from https://github.com/IMu-MachineLearningsxD/G2T. Tianxin Zhao, Yingxin Liu, Xiangdong Su, Jiang Li 0013, Guanglai Gao |
LREC/COLING | 5 |
| 2024 | Fully Hyperbolic Rotation for Knowledge Graph EmbeddingabstractHyperbolic rotation is commonly used to effectively model knowledge graphs and their inherent hierarchies. However, existing hyperbolic rotation models rely on logarithmic and exponential mappings for feature transformation. These models only project data features into hyperbolic space for rotation, limiting their ability to fully exploit the hyperbolic space. To address this problem, we propose a novel fully hyperbolic model designed for knowledge graph embedding. Instead of feature mappings, we define the model directly in hyperbolic space with the Lorentz model. Our model considers each relation in knowledge graphs as a Lorentz rotation from the head entity to the tail entity. We adopt the Lorentzian version distance as the scoring function for measuring the plausibility of triplets. Extensive results on standard knowledge graph completion benchmarks demonstrated that our model achieves competitive results with fewer parameters. In addition, our model get the state-of-the-art performance on datasets of CoDEx-s and CoDEx-m, which are more diverse and challenging than before. Our code is available at https://github.com/llqy123/FHRE. Qiuyu Liang, Weihua Wang 0006, Feilong Bao, Guanglai Gao |
ECAI | 4 |
| 2024 | MEMix: Improving HMER with Diverse Formula Structure AugmentationabstractHandwritten Mathematical Expression Recognition (HMER) aims to transform images of mathematical expressions (MEs) into corresponding LaTeX sequences. However, the inherent complexity of 2D formula structures often misaligns with the 1D LaTeX sequences, resulting in decreased robustness in recognition models. A primary factor exacerbating this issue is the scarcity of annotated ME images with complex structures, which hinders the models to learning to good representation and adaptability for MEs. In this paper, drawing inspiration from Mixup, we introduce a data augmentation method called Mathematical Expression Mix (MEMix). This method is capable of generating typical structures in formulas, including radicals, fractions, and annotations, by employing straightforward matrix operations. Compared to alternative data augmentation methods, MEMix provides faster and more cost-effective computation, enabling online augmentation that improves training efficiency. Experiments demonstrate that MEMix significantly enhances the performance of the baseline model on the benchmark datasets. Xiangdong Su, Xingxiang Zhou, Guanglai Gao |
ICME | 4 |
| 2024 | Improving End-to-End Speech Recognition Through Conditional Cross-Modal Knowledge Distillation with Language ModelabstractRecently, cross-modal knowledge distillation methods for end-to-end automatic speech recognition (E2E-ASR) model training pointed out the potential help of text data for improving recognition performance. However, conventional optimization strategies can mislead student models to produce suboptimal performance due to erroneous predictions generated by teacher models. This paper addresses the issue by proposing a conditional cross-modal knowledge distillation strategy, a novel technique for selectively incorporating contextual linguistic information from language model into the E2E-ASR model for improving the recognition performance. We introduce a conditional selector to dynamically adjust the knowledge source of the student model to avoid knowledge distillation from erroneous predictions generated by teacher model. In pre-trained language model fine-tuning, we perform an analysis of the impact of unsupervised text data of varying scales on the quality of soft labels and the recognition performance of the E2E-ASR model. Our proposed method simultaneously improve two different non-autoregressive decoding approaches. Experiments on the Chinese speech datasets AISHELL-1 and AISHELL-2 show competitive performance. Yonghe Wang, Feilong Bao, Zhenjie Gao, Guanglai Gao |
IJCNN | 5 |
| 2024 | Pre-training Language Model for Mongolian with Agglutinative Linguistic Knowledge InjectionabstractBERT based Pre-training Language Model (PLM) has become a crucial step in achieving the best results in various natural language processing (NLP) tasks. However, the current progress, which mainly focuses on major languages such as English and Chinese, has not thoroughly investigated the low-resource languages, particularly agglutinative languages like Mongolian, due to the scarcity of large-scale data resources and the difficulty of understanding agglutinative knowledge. In this paper, we propose a novel PLM for the Mongolian language, that incorporates a novel three-stage agglutinative knowledge injection strategy. Specifically, early-stage injection aims to convert the Mongolia word sequence to the fine-grained sub-word token that comprises a stem and some suffixes; Middle-stage injection designed a morphological knowledge-based masking strategy to enhance the model's ability to learn agglutinative knowledge; Late-stage injection not only involves the model restoring the masked tokens but also predicting the order of suffixes. To address the issue of data scarcity, we create a large-scale Mongolian PLM dataset and three datasets for three downstream tasks, that are News Classification, Name Entity Recognition (NER), and Part-of-Speech (POS) prediction, etc. The experimental results on three downstream tasks demonstrate that our method surpasses the traditional BERT approach and successfully learns agglutinative language knowledge in Mongolian. Muhan Na, Rui Liu 0008, Feilong Bao, Guanglai Gao |
IJCNN | 4 |
| 2024 | Cross-Attention-Guided WaveNet for EEG-to-MEL Spectrogram Reconstruction
Hao Li 0046, Xueliang Zhang 0001, Fei Chen 0011, Guanglai Gao |
INTERSPEECH | 5 |
| 2024 | GSEA: Global Structure-Aware Graph Neural Networks for Entity Alignment
Cunda Wang, Weihua Wang 0006, Qiuyu Liang, Jie Yu 0008, Guanglai Gao |
NLPCC (2) | 5 |
| 2024 | The image and ground truth dataset of Mongolian movable-type newspapers for text recognition
Feilong Bao, Hui Zhang 0031, Guanglai Gao |
Int. J. Document Anal. Recognit. | 4 |
| 2024 | Text-to-Speech for Low-Resource Agglutinative Language With Morphology-Aware Language Model Pre-TrainingabstractText-to-Speech (TTS) aims to convert the input text to a human-like voice. With the development of deep learning, encoder-decoder based TTS models perform superior performance, in terms of naturalness, in mainstream languages such as Chinese, English, etc. Note that the linguistic information learning capability of the text encoder is the key. However, for TTS of low-resource agglutinative languages, the scale of the$< $text, speech$>$paired data is limited. Therefore, how to extract rich linguistic information from small-scale text data to enhance the naturalness of the synthesized speech, is an urgent issue that needs to be addressed. In this paper, we first collect a large unsupervised text data for BERT-like language model pre-training, and then adopt the trained language model to extract deep linguistic information for the input text of the TTS model to improve the naturalness of the final synthesized speech. It should be emphasized that in order to fully exploit the prosody-related linguistic information in agglutinative languages, we incorporated morphological information into the language model training and constructed a morphology-aware masking based BERT model (MAM-BERT). Experimental results based on various advanced TTS models validate the effectiveness of our approach. Further comparison of the various data scales also validates the effectiveness of our approach in low-resource scenarios. Rui Liu 0008, Yifan Hu 0004, Haolin Zuo, Zhaojie Luo, Longbiao Wang, Guanglai Gao |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2024 | Controllable Accented Text-to-Speech Synthesis With Fine and Coarse-Grained Intensity RenderingabstractAccented text-to-speech (TTS) synthesis seeks to generate speech with an accent (L2) as a variant of the standard version (L1), which is challenging as L2 is different from L1 in terms of phonetic rendering and prosody pattern (pitch, energy, and duration variance, etc.). Accented TTS has several significant real-world applications, such as language learning, preserving and documenting endangered languages and dialects, etc. that make it an important area of research and development. Moreover, changing the accent intensity of any conversational AI system has the potential to allow specific users to understand its produced speech better. However, there is no intuitive solution for the control of the accent intensity for an utterance at both fine and coarse-grained levels, that are phoneme and utterance levels respectively. In this work, we propose a neural TTS architecture that allows us to control the accent style and its intensity. This is achieved through two novel mechanisms: 1) the front-end and back-end accent knowledge injection mechanism to enhance the accent interpretability of TTS modeling; and 2) an automatic speech recognition (ASR) based accent intensity modeling strategy to quantify the accent intensity in both L2 phoneme and utterance levels. In the front-end, a newaccent variation adaptorseeks to project the accent-aware pitch, energy and duration features at a phoneme level, with the help of the fine-grained accent intensity information; In the back-end, a consistency constraint module that ensures the synthesized L2 speech manifests the expected accent intensity, is injected in the front-end, precisely. Experiments show that the proposed system attains superior performance to the baseline models in terms of accent rendering and intensity control. To our knowledge, this is the first study of accented TTS with explicit intensity control at both fine and coarse-grained levels. Rui Liu 0008, Berrak Sisman, Guanglai Gao, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | TeAST: Temporal Knowledge Graph Embedding via Archimedean Spiral TimelineabstractTemporal knowledge graph embedding (TKGE) models are commonly utilized to infer the missing facts and facilitate reasoning and decision-making in temporal knowledge graph based systems.However, existing methods fuse temporal information into entities, potentially leading to the evolution of entity information and limiting the link prediction performance of TKG.Meanwhile, current TKGE models often lack the ability to simultaneously model important relation patterns and provide interpretability, which hinders their effectiveness and potential applications.To address these limitations, we propose a novel TKGE model which encodes Temporal knowledge graph embeddings via Archimedean Spiral Timeline (TeAST), which maps relations onto the corresponding Archimedean spiral timeline and transforms the quadruples completion to 3th-order tensor completion problem.Specifically, the Archimedean spiral timeline ensures that relations that occur simultaneously are placed on the same timeline, and all relations evolve over time.Meanwhile, we present a novel temporal spiral regularizer to make the spiral timeline orderly.In addition, we provide mathematical proofs to demonstrate the ability of TeAST to encode various relation patterns.Experimental results show that our proposed model significantly outperforms existing TKGE methods. Jiang Li 0013, Xiangdong Su, Guanglai Gao |
ACL (1) | 3 |
| 2023 | TableSF: A Structural Bias Framework for Table-To-Text Generation
Weihua Wang 0006, Feilong Bao, Guanglai Gao |
ICANN (9) | 4 |
| 2023 | Exploiting Modality-Invariant Feature for Robust Multimodal Emotion Recognition with Missing ModalitiesabstractMultimodal emotion recognition leverages complementary information across modalities to gain performance. However, we cannot guarantee that the data of all modalities are always present in practice. In the studies to predict the missing data across modalities, the inherent difference between heterogeneous modalities, namely the modality gap, presents a challenge. To address this, we propose to use invariant features for a missing modality imagination network (IF-MMIN) which includes two novel mechanisms: 1) an invariant feature learning strategy that is based on the central moment discrepancy (CMD) distance under the full-modality scenario; 2) an invariant feature based imagination module (IF-IM) to alleviate the modality gap during the missing modalities prediction, thus improving the robustness of multimodal joint representation. Comprehensive experiments on the benchmark dataset IEMOCAP demonstrate that the proposed model outperforms all baselines and invariantly improves the overall emotion recognition performance under uncertain missing-modality conditions. We release the code at: https://github.com/ZhuoYulang/IF-MMIN. Haolin Zuo, Rui Liu 0008, Jinming Zhao, Guanglai Gao, Haizhou Li 0001 |
ICASSP | 4 |
| 2023 | Betray Oneself: A Novel Audio DeepFake Detection Model via Mono-to-Stereo Conversion
Rui Liu 0008, Guanglai Gao, Haizhou Li 0001 |
INTERSPEECH | 3 |
| 2023 | Explicit Intensity Control for Accented Text-to-speech
Rui Liu 0008, Haolin Zuo, De Hu, Guanglai Gao, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2023 | Few-Shot Table-to-Text Generation with Structural Bias Attention
Weihua Wang 0006, Feilong Bao, Guanglai Gao |
PRICAI (2) | 4 |
| 2023 | Noise-Separated Adaptive Feature Distillation for Robust Speech RecognitionabstractThis letter makes an improvement on feature-based knowledge distillation for robust speech recognition. The use of distillation techniques in speech recognition has been demonstrated to improve the robustness of the system. In this letter, we propose a noise-separated adaptive feature distillation method, including an adaptive distillation position selection strategy and a noise separation mechanism, assuming that there is a common network structure between the student and teacher. The proposed method has two improvements. First, distillation positions can be adaptively selected in each iteration by comparing loss values computed on intermediate representations of the student and the teacher, increasing the flexibility of knowledge transfer during distillation. Second, a noise separation module is proposed to constrain noise information elimination by explicitly separating the speech information and the noise information in noisy speech, which reduces the interference of noise information during distillation. Therefore, a better recognition performance is demonstrated with the proposed method compared to the standardized feature-based knowledge distillation method. Honglin Qu, Xiangdong Su, Yonghe Wang, Guanglai Gao |
IEEE Signal Process. Lett. | 5 |
| 2023 | A Comparative Study on Selecting Acoustic Modeling Units for WFST-based Mongolian Speech RecognitionabstractTraditional weighted finite-state transducer– (WFST) based Mongolian automatic speech recognition (ASR) systems use phonemes as pronunciation lexicon modeling units. However, Mongolian is an agglutinative, low-resource language, and building an ASR system based on the phoneme pronunciation lexicon remains a challenge for various reasons. First, the phoneme pronunciation lexicon manually constructed by Mongolian linguists is finite, which is usually used to build a grapheme-to-phoneme conversion (G2P) model to frequently expand new words. However, the data sparsity decreases the robustness of the G2P model and affects the performance of the final ASR system. Second, homophones and polysyllabic words are common in Mongolian, which has a certain impact on the construction of the Mongolian acoustic model. To address these problems, in this work, we first propose a grapheme-to-phoneme alignment model to obtain the mapping relationship between phonemes and subword units. Then, we construct an acoustic subword segmentation set to segment words directly instead of using the traditional G2P method to predict phoneme sequences to expand the pronunciation lexicon. Further, by analyzing the Mongolian encoding form, we also propose an acoustic subword modeling units construction method that removes control characters. Finally, we investigate various acoustic subword modeling units for pronunciation lexicon construction for the Mongolian ASR system. Experiments on a Mongolian dataset with 325 hours of training show that the pronunciation lexicon based on the acoustic subword modeling unit can effectively construct the WFST-based Mongolian ASR system. Further, removing the control characters when building the acoustic subword modeling unit can further improve the ASR system performance. Yonghe Wang, Feilong Bao, Guanglai Gao |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2022 | End-to-End Large-Scale Image Retrieval Network with Convolution and Vision Transformers
Feilong Bao, Xiangdong Su, Weihua Wang 0006, Guanglai Gao |
ICANN (4) | 5 |
| 2022 | Alignment-Learning Based Single-Step Decoding for Accurate and Fast Non-Autoregressive Speech RecognitionabstractNon-autoregressive transformer (NAT) based speech recognition models have gained more and more attention since they perform faster inference speed compared with autoregressive counterparts, especially when the single-step decoding is applied. However, the single-step decoding process with length prediction will suffer from the decoding stability problem and limited improvement for inference speed. To address this, in this paper, we propose an alignment learning based NAT model, named AL-NAT. Our idea is inspired by the fact that the encoder CTC output and the target sequence are monotonically related. Specifically, we design an alignment cost matrix between the CTC output tokens and the target tokens and define a novel alignment loss to minimize the distance between the alignment cost matrix and the ground truth monotonic alignment path. By eliminating the length prediction mechanism, our AL-NAT model achieves remarkable improvements in recognition accuracy and decoding speed. To learn the contextual knowledge to improve the decoding accuracy, we further add lightweight language model on both the encoder and decoder side. Our proposed method achieves WERs of 2.8%/6.3% and RTF of 0.011 on Librispeech test clean/other sets with a lightweight 3-gram LM, and a CER of 5.3% and RTF of 0.005 on Aishell1 without LM, respectively. Yonghe Wang, Rui Liu 0008, Feilong Bao, Hui Zhang 0031, Guanglai Gao |
ICASSP | 5 |
| 2022 | A Deep Investigation of RNN and Self-attention for the Cyrillic-Traditional Mongolian Bidirectional Conversion
Muhan Na, Rui Liu 0008, Feilong Bao, Guanglai Gao |
ICONIP (6) | 4 |
| 2022 | Accurate Emotion Strength Assessment for Seen and Unseen Speech Based on Data-Driven Deep LearningabstractEmotion classification of speech and assessment of the emotion strength are required in applications such as emotional text-to-speech and voice conversion. The emotion attribute ranking function based on Support Vector Machine (SVM) was proposed to predict emotion strength for emotional speech corpus. However, the trained ranking function doesn't generalize to new domains, which limits the scope of applications, especially for out-of-domain or unseen speech. In this paper, we propose a data-driven deep learning model, i.e. StrengthNet, to improve the generalization of emotion strength assessment for seen and unseen speech. This is achieved by the fusion of emotional data from various domains. We follow a multi-task learning network architecture that includes an acoustic encoder, a strength predictor, and an auxiliary emotion predictor. Experiments show that the predicted emotion strength of the proposed StrengthNet is highly correlated with ground truth scores for both seen and unseen speech. We release the source codes at: https://github.com/ttslr/StrengthNet. Rui Liu 0008, Berrak Sisman, Björn W. Schuller, Guanglai Gao, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2022 | QuatSE: Spherical Linear Interpolation of Quaternion for Knowledge Graph Embeddings
Jiang Li 0013, Xiangdong Su, Xinlan Ma, Guanglai Gao |
NLPCC (1) | 4 |
| 2022 | Decoding Knowledge Transfer for Neural Text-to-Speech TrainingabstractNeural end-to-end text-to-speech (TTS) is superior to conventional statistical methods in many ways. However, the exposure bias problem, that arises from the mismatch between the training and inference process in autoregressive models, remains an issue. It often leads to performance degradation in face of out-of-domain test data. To address this problem, we study a novel decoding knowledge transfer strategy, and propose a multi-teacher knowledge distillation (MT-KD) network for Tacotron2 TTS model. The idea is to pre-train two Tacotron2 TTS teacher models in teacher forcing and scheduled sampling modes, and transfer the pre-trained knowledge to a student model that performs free running decoding. We show that the MT-KD network provides an adequate platform for neural TTS training, where the student model learns to emulate the behaviors of the two teachers, at the same time, minimizing the mismatch between training and run-time inference. Experiments on both Chinese and English data show that MT-KD system consistently outperforms the competitive baselines in terms of naturalness, robustness and expressiveness for in-domain and out-of-domain test data. Furthermore, we show that knowledge distillation outperforms adversarial learning and data augmentation in addressing the exposure bias problem. Rui Liu 0008, Berrak Sisman, Guanglai Gao, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Joint Alignment Learning-Attention Based Model for Grapheme-to-Phoneme ConversionabstractSequence-to-sequence attention-based models for grapheme-to-phoneme (G2P) conversion have gained significant interests. The attention-based encoder-decoder framework learns the mapping of input to output tokens by selectively focusing on relevant information, and has been shown well performance. However, the attention mechanism can result in non-monotonic alignments, resulting in poor G2P conversion performance. In this paper, we present a novel approach to optimize the G2P conversion model directly alignment grapheme-phoneme sequence by using alignment learning (AL) as the loss function. Besides, we propose a multi-task learning method that uses a joint alignment learning model and attention model to predict the proper alignments and thus improve the accuracy of G2P conversion. Evaluations on Mongolian and CMUDict tasks show that alignment learning as the loss function can effectively train G2P conversion model. Further, our multi-task method can significantly outperform both the alignment learning-based model and attention-based model. Yonghe Wang, Feilong Bao, Hui Zhang 0031, Guanglai Gao |
ICASSP | 4 |
| 2021 | Panoptic-DLA: Document Layout Analysis of Historical Newspapers Based on Proposal-Free Panoptic Segmentation Model
Feilong Bao, Guanglai Gao |
KSEM | 3 |
| 2021 | Soft-BAC: Soft Bidirectional Alignment Cost for End-to-End Automatic Speech Recognition
Yonghe Wang, Hui Zhang 0031, Feilong Bao, Guanglai Gao |
PRICAI (2) | 4 |
| 2021 | Recurrent Neural Networks and Acoustic Features for Frame-Level Signal-to-Noise Ratio EstimationabstractIt is important to know the presence and the relative level of background noise for many speech processing tasks. Frame-level signal-to-noise ratio (SNR) provides a measure of instantaneous noise level of a noisy signal, and its estimation has been researched for decades. This problem can be approached from a supervised learning perspective by predicting SNR from features of noisy speech. In this study, we introduce a deep learning algorithm for frame-level SNR estimation. The proposed algorithm employs recurrent neural networks (RNNs) with long short-term memory (LSTM) to leverage contextual information. We also systematically examine a range of acoustic features and investigate feature combinations using Group Lasso and sequential floating forward selection (SFFS). The proposed algorithm naturally leads to an utterance-level SNR estimator. Systematical evaluations show that the proposed algorithm provides an accurate estimate of frame-level SNR, as well as utterance-level SNR, under different noise conditions, outperforming other estimators. Hao Li 0046, DeLiang Wang, Xueliang Zhang 0001, Guanglai Gao |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Exploiting Morphological and Phonological Features to Improve Prosodic Phrasing for Mongolian Speech SynthesisabstractProsodic phrasing is an important factor that affects naturalness and intelligibility in text-to-speech synthesis. Studies show that deep learning techniques improve prosodic phrasing when large text and speech corpus are available. However, for low-resource languages, such as Mongolian, prosodic phrasing remains a challenge for various reasons. First, the database suitable for system training is limited. Second, word composition knowledge that is prosody-informing has not been used in prosodic phrase modeling. To address these problems, in this article, we propose a feature augmentation method in conjunction with a self-attention neural classifier. We augment input text with morphological and phonological decompositions of words to enhance the text encoder. We study the use of self-attention classifier, that makes use of global context of a sentence, as a decoder for phrase break prediction. Both objective and subjective evaluations validate the effectiveness of the proposed phrase break prediction framework, that consistently improves voice quality in a Mongolian text-to-speech synthesis system. Rui Liu 0008, Berrak Sisman, Feilong Bao, Guanglai Gao, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2021 | Expressive TTS Training With Frame and Style Reconstruction LossabstractWe propose a novel training strategy for Tacotron-based text-to-speech (TTS) system that improves the speech styling at utterance level. One of the key challenges in prosody modeling is the lack of reference that makes explicit modeling difficult. The proposed technique doesn’t require prosody annotations from training data. It doesn’t attempt to model prosody explicitly either, but rather encodes the association between input text and its prosody styles using a Tacotron-based TTS framework. This study marks a departure from the style token paradigm where prosody is explicitly modeled by a bank of prosody embeddings. It adopts a combination of two objective functions: 1) frame level reconstruction loss, that is calculated between the synthesized and target spectral features; 2) utterance level style reconstruction loss, that is calculated between the deep style features of synthesized and target speech. The style reconstruction loss is formulated as a perceptual loss to ensure that utterance level speech style is taken into consideration during training. Experiments show that the proposed training strategy achieves remarkable performance and outperforms the state-of-the-art baseline in both naturalness and expressiveness. To our best knowledge, this is the first study to incorporate utterance level perceptual quality as a loss function into Tacotron training for improved expressiveness. Rui Liu 0008, Berrak Sisman, Guanglai Gao, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Incorporating Inner-word and Out-word Features for Mongolian Morphological SegmentationabstractMongolian morphological segmentation is regarded as a crucial preprocessing step in many Mongolian related NLP applications and has received extensive attention.Recently, end-to-end segmentation approaches with long short-term memory networks (LSTM) have achieved excellent results.However, the inner-word features among characters in the word and the out-word features from context are not well utilized in the segmentation process.In this paper, we propose a neural network incorporating inner-word and out-word features for Mongolian morphological segmentation.The network consists of two encoders and one decoder.The inner-word encoder uses the self-attention mechanisms to capture the inner-word features of the target word.The out-word encoder employs a two layers BiLSTM network to extract out-word features in the sentence.Then, the decoder adopts a multi-head double attention layer to fuse the inner-word features and out-word features and produces the segmentation result.The evaluation experiment compares the proposed network with the baselines and explores the effectiveness of the sub-modules. Xiangdong Su, Guanglai Gao, Feilong Bao |
COLING | 4 |
| 2020 | Teacher-Student Training For Robust Tacotron-Based TTSabstractWhile neural end-to-end text-to-speech (TTS) is superior to conventional statistical methods in many ways, the exposure bias problem in the autoregressive models remains an issue to be resolved. The exposure bias problem arises from the mismatch between the training and inference process, that results in unpredictable performance for out-of-domain test data at run-time. To overcome this, we propose a teacher-student training scheme for Tacotron-based TTS by introducing a distillation loss function in addition to the feature loss function. We first train a Tacotron2-based TTS model by always providing natural speech frames to the decoder, that serves as a teacher model. We then train another Tacotron2-based model as a student model, of which the decoder takes the predicted speech frames as input, similar to how the decoder works during run-time inference. With the distillation loss, the student model learns the output probabilities from the teacher model, that is called knowledge distillation. Experiments show that our proposed training scheme consistently improves the voice quality for out-of-domain test data both in Chinese and English systems. Rui Liu 0008, Berrak Sisman, Jingdong Li, Feilong Bao, Guanglai Gao, Haizhou Li 0001 |
ICASSP | 5 |
| 2020 | Beamformed Feature for Learning-based Dual-channel Speech SeparationabstractThis paper deals with the problem of separating target speech signal from reverberant and noisy environment with dual microphones, where the target speech comes from a predefined direction range. First, we apply two differential beamformers with opposite directions to dual-channel inputs. Then, the power spectra of beamforming outputs are used as input feature of deep learning architecture. As input features, the beamformer outputs reflect not only spectral information but also directional information by their power level difference. And the calculation is very simple. Systematic evaluation and comparison show that the proposed system achieves very good separation performance and substantially outperforms related algorithms under very challenging environments where both interfering speaker, noise and reverberations are present. Hao Li 0046, Xueliang Zhang 0001, Guanglai Gao |
ICASSP | 3 |
| 2020 | Snr-Based Teachers-Student Technique For Speech EnhancementabstractIt is very challenging for speech enhancement methods to achieves robust performance under both high signal-to-noise ratio (SNR) and low SNR simultaneously. In this paper, we propose a method that integrates an SNR-based teachers-student technique and time-domain U-Net to deal with this problem. Specifically, this method consists of multiple teacher models and a student model. We first train the teacher models under multiple small-range SNRs that do not coincide with each other so that they can perform speech enhancement well within the specific SNR range. Then, we choose different teacher models to supervise the training of the student model according to the SNR of the training data. Eventually, the student model can perform speech enhancement under both high SNR and low SNR. To evaluate the proposed method, we constructed a dataset with an SNR ranging from -20dB to 20dB based on the public dataset. We experimentally analyzed the effectiveness of the SNR-based teachers-student technique and compared the proposed method with several state-of-the-art methods. Xiangdong Su, Huali Xu, Guanglai Gao |
ICME | 6 |
| 2020 | An Edge Information and Mask Shrinking Based Image Inpainting ApproachabstractIn the image inpainting task, the ability to repair both high-frequency and low-frequency information in the missing regions has a substantial influence on the quality of the restored image. However, existing inpainting methods usually fail to consider both high-frequency and low-frequency information simultaneously. To solve this problem, this paper proposes edge information and mask shrinking based image inpainting approach, which consists of two models. The first model is an edge generation model used to generate complete edge information from the damaged image, and the second model is an image completion model used to fix the missing regions with the generated edge information and the valid contents of the damaged image. The mask shrinking strategy is employed in the image completion model to track the areas to be repaired. The proposed approach is evaluated qualitatively and quantitatively on the dataset Places2. The result shows our approach outperforms state-of-the-art methods. Huali Xu, Xiangdong Su, Guanglai Gao |
ICME | 5 |
| 2020 | Dataless Text Classification with Pseudo Topic RepresentationabstractAs for an automatic text classification approach, a large body of research on latent-topic based Dataless Text Classification (DTC) has been emerged in recent years. Perusing the candidate seed words or guaranteeing the quality of the category-topics is the core mission of this approach. However, few previous studies consider the quality of specific category-topics at the collection level instead at the document level, because not all topics are equally coherent or category sparsity. In this paper, we focus on alleviating the dilemma for the seed words selection problem in DTC by using pseudo text understanding. Differently from the existing latent-topic based DTC approach, we propose an unsupervised method named Pseudo Document Labeled Classification (PDLC). It extracts the most representative word list to capture the best latent semantic category-topic description. Experimental results indicate that our PDLC scheme achieves better classification accuracy without any labeled data or external resource. Guanglai Gao |
ICTAI | 3 |
| 2020 | Sub-Band Knowledge Distillation Framework for Speech EnhancementabstractIn single-channel speech enhancement, methods based on full-band spectral features have been widely studied. However, only a few methods pay attention to non-full-band spectral features. In this paper, we explore a knowledge distillation framework based on sub-band spectral mapping for single-channel speech enhancement. Specifically, we divide the full frequency band into multiple sub-bands and pre-train an elite-level sub-band enhancement model (teacher model) for each sub-band. These teacher models are dedicated to processing their own sub-bands. Next, under the teacher models' guidance, we train a general sub-band enhancement model (student model) that works for all sub-bands. Without increasing the number of model parameters and computational complexity, the student model's performance is further improved. To evaluate our proposed method, we conducted a large number of experiments on an open-source data set. The final experimental results show that the guidance from the elite-level teacher models dramatically improves the student model's performance, which exceeds the full-band model by employing fewer parameters. Shixue Wen, Xiangdong Su, Guanglai Gao |
INTERSPEECH | 5 |
| 2020 | Frame-Level Signal-to-Noise Ratio Estimation Using Deep Learning
Hao Li 0046, DeLiang Wang, Xueliang Zhang 0001, Guanglai Gao |
INTERSPEECH | 4 |
| 2020 | Topic Analysis by Exploring Headline Information
Guanglai Gao |
WISE (2) | 2 |
| 2020 | Modeling Prosodic Phrasing With Multi-Task Learning in Tacotron-Based TTSabstractTacotron-based end-to-end speech synthesis has shown remarkable voice quality. However, the rendering of prosody in the synthesized speech remains to be improved, especially for long sentences, where prosodic phrasing errors can occur frequently. In this letter, we extend the Tacotron-based speech synthesis framework to explicitly model the prosodic phrase breaks. We propose a multi-task learning scheme for Tacotron training, that optimizes the system to predict both Mel spectrum and phrase breaks. To our best knowledge, this is the first implementation of multi-task learning for Tacotron based TTS with a prosodic phrasing model. Experiments show that our proposed training scheme consistently improves the voice quality for both Chinese and Mongolian systems. Rui Liu 0008, Berrak Sisman, Feilong Bao, Guanglai Gao, Haizhou Li 0001 |
IEEE Signal Process. Lett. | 4 |
| 2019 | Sub-Word Based Mongolian Offline Handwriting RecognitionabstractMongolian is an agglutinative language, which re-sults in a large number of words derived from the same stems connecting different suffixes. This morphological richness leads to high out-of-vocabulary (OOV) rates and causes problems of data sparsity. In this paper, our proposed recognition system is composed of three parts: handwritten image preprocessing, mapping of images to grapheme sequences, and sub-word-based language model (LM) decoding. We present a sub-word-based n-gram LM to solve the high OOV rate problem. According to the characteristics of Mongolian, we modified the traditional token passing algorithm to improve decoding speed and to easy to combine with any n-gram LM. We evaluated the performance of sub-words at different levels on the open Mongolian offline handwriting dataset (MHW). The bi-syllable 2-gram LM showed the best performance, with 18.32% and 23.22% word-error rates (WERs) on two test sets. Our various experiments show that, this method can predict in vocabulary words with a higher accuracy rate and also predict OOV words with a certain accuracy rate. Daoerji Fan, Guanglai Gao, Huijuan Wu |
ICDAR | 2 |
| 2019 | Woodblock-Printing Mongolian Words Recognition by Bi-LSTM with Attention MechanismabstractWoodblock-printing Mongolian documents are seriously degraded due to aging. Therefore, it is difficult to segment woodblock-printing Mongolian words are into individual glyphs. In this paper, a holistic recognition approach based on sequence to sequence model has been proposed for the woodblock-printing Mongolian words. The input of the proposed model is the sequence of frames of a wood-block printing Mongolian word. In order to generating the corresponding sequence of frames, each word image should be normalized into the same sizes in advance. And then, each word image is segmented into several fragments with equal size along writing direction. The output of the proposed model is a sequence of letters. To be specific, the proposed model contains three parts: an encoder, a decoder and an attention network. The encoder consists of a deep neural network and a bi-directional Long Short-Term Memory (Bi-LSTM). The decoder consists of a Long Short-Term Memory (LSTM) with a softmax layer. The encoder and decoder are connected by an attention network, which can map multiple frames to one letter. Experimental results demonstrate that the proposed approach outperforms the segmentation based method. Yanke Kang, Hongxi Wei, Hui Zhang 0031, Guanglai Gao |
ICDAR | 4 |
| 2019 | Improving Text Image Resolution using a Deep Generative Adversarial Network for Optical Character RecognitionabstractOptical character recognition (OCR) has been widely studied in previous work. Except for the models used, the recognition accuracy depends most on the resolution of the image to be recognized. To enhance OCR performance, this paper proposes an approach based on a generative adversarial network to improve text image resolution. Our approach uses a perceptual loss function that consists of an adversarial loss, a content loss and an L1 loss. The adversarial loss and the L1 loss are used to ensure the generated super-resolved images are closer to the ground truth high-resolution images. Meanwhile, the content loss is used to ensure the generated super-resolved images and the input low-resolution images have similar features on the basis of perceptual instead of pixel similarity. To evaluate the proposed approach, we compare the recognition accuracies before and after improving the resolution of both English and Chinese text images. The results show that the recognition accuracies on the super-resolved text images obtained with our approach are significantly higher than those on the low-resolution images without processing. Xiangdong Su, Huali Xu, Ying Kang, Guanglai Gao |
ICDAR | 5 |
| 2019 | A Holistic Recognition Approach for Woodblock-Print Mongolian Words Based on Convolutional Neural NetworkabstractThis paper proposed a holistic recognition approach for woodblock-print Mongolian words using a convolutional neural network (CNN). To be specific, the whole word image is regarded as input of CNN. Hence, all the word images should be normalized into the same size before being inputted into CNN. By comparison, an appropriate normalization size has been determined in our study. Through the above manner, the woodblock-print Mongolian word images do not need to be segmented into glyphs. Thereby, the segmentation errors can be avoided under the circumstance. Furthermore, to solve the problem of imbalance distribution on our dataset, SMOTE technique is adopted to generate samples. In this way, the training procedure can be more efficient and the obtained CNN is more robust. Experimental results demonstrate that the proposed approach outperforms the segmentation based method and other baselines. Hongxi Wei, Guanglai Gao |
ICIP | 2 |
| 2019 | Building Mongolian TTS Front-End with Encoder-Decoder Model by Using Bridge Method and Multi-view Features
Rui Liu 0008, Feilong Bao, Guanglai Gao |
ICONIP (5) | 3 |
| 2019 | Morphological Knowledge Guided Mongolian Constituent Parsing
Xiangdong Su, Guanglai Gao, Feilong Bao |
ICONIP (3) | 3 |
| 2019 | Learning an Adversarial Network for Speech Enhancement Under Extremely Low Signal-to-Noise Ratio Condition
Xiangdong Su, Huali Xu, Tongyang Liu, Guanglai Gao, Feilong Bao |
ICONIP (1) | 7 |
| 2019 | A Natural Scene Text Extraction Approach Based on Generative Adversarial Learning
Huali Xu, Xiangdong Su, Tongyang Liu, Guanglai Gao, Feilong Bao |
ICONIP (1) | 5 |
| 2019 | Neural Morphological Segmentation Model for MongolianabstractMorphological segmentation is useful for processing Mongolian. In this paper, we manually build a morphological segmentation data set for Mongolian. We then present a character-based encoder-decoder model with attention mechanism to perform the morphological segmentation task. We further investigate the influence of analogy features extracted from scratch and improve the performance of our model using multi languages setting. Experimental results show that our encoder-decoder model with attention mechanism provides a strong baseline for Mongolian morphological segmentation. The analogy features provide useful information to the model and improve the performance of the system. The use of multi languages data set shows the capability of our model to acquire knowledge through different languages and delivers the best result. Weihua Wang 0006, Rashel Fam, Feilong Bao, Yves Lepage, Guanglai Gao |
IJCNN | 5 |
| 2019 | An Automatic Spelling Correction Method for Classical Mongolian
Feilong Bao, Guanglai Gao, Weihua Wang 0006, Hui Zhang 0031 |
KSEM (2) | 3 |
| 2019 | A Context-Free Spelling Correction Method for Classical Mongolian
Feilong Bao, Guanglai Gao |
NLPCC (2) | 3 |
| 2019 | Research on Khalkha Dialect Mongolian Speech Recognition Acoustic Model Based on Weight Transfer
Linyan Shi, Feilong Bao, Yonghe Wang, Guanglai Gao |
NLPCC (2) | 4 |
| 2019 | End-to-End Model for Offline Handwritten Mongolian Word Recognition
Hongxi Wei, Hui Zhang 0031, Feilong Bao, Guanglai Gao |
NLPCC (2) | 5 |
| 2019 | An End-to-End Preprocessor Based on Adversiarial Learning for Mongolian Historical Document OCR
Xiangdong Su, Huali Xu, Yanke Kang, Guanglai Gao, Batushiren |
PRICAI (3) | 5 |
| 2019 | Learning Morpheme Representation for Mongolian Named Entity Recognition
Weihua Wang 0006, Feilong Bao, Guanglai Gao |
Neural Process. Lett. | 3 |
| 2018 | A LSTM Approach with Sub-Word Embeddings for Mongolian Phrase Break PredictionabstractIn this paper, we first utilize the word embedding that focuses on sub-word units to the Mongolian Phrase Break (PB) prediction task by using Long-Short-Term-Memory (LSTM) model. Mongolian is an agglutinative language. Each root can be followed by several suffixes to form probably millions of words, but the existing Mongolian corpus is not enough to build a robust entire word embedding, thus it suffers a serious data sparse problem and brings a great difficulty for Mongolian PB prediction. To solve this problem, we look at sub-word units in Mongolian word, and encode their information to a meaningful representation, then fed it to LSTM to decode the best corresponding PB label. Experimental results show that the proposed model significantly outperforms traditional CRF model using manually features and obtains 7.49% F-Measure gain. Rui Liu 0008, Feilong Bao, Guanglai Gao, Hui Zhang 0031, Yonghe Wang |
COLING | 3 |
| 2018 | Training Supervised Speech Separation System to Improve STOI and PESQ DirectlyabstractSupervised speech separation methods train learning machine to cast the noisy speech to the target clean speech. Most of them use mean-square error (MSE) as loss function. However, MSE is not the perfect choice because it doesn't match the human auditory perception. Short-time objective intelligibility (STOI) and perceptual evaluation of speech quality (PESQ) are closely related to the human auditory perception and widely used in speech separation research as evaluation criteria. Therefore, STOI and PESQ may be better choices for the loss function. However, they are nondifferentiable functions which cannot be optimized by the conventional gradient descent algorithm. In this work, a gradient approximation method is used to calculate the gradients of the STOI and PESQ. Then the calculated gradients are used in the gradient descent algorithm to optimize the STOI and PESQ directly. Experimental results show the speech separation performance can be improved by the proposed method. Hui Zhang 0031, Xueliang Zhang 0001, Guanglai Gao |
ICASSP | 3 |
| 2018 | Mongolian Word Segmentation Based on Three Character Level Seq2Seq Models
Xiangdong Su, Guanglai Gao, Feilong Bao |
ICONIP (5) | 3 |
| 2018 | Convolutional Neural Network for Machine-Printed Traditional Mongolian Font Recognition
Hongxi Wei, Weiyuan Wang, Guanglai Gao |
ICONIP (5) | 4 |
| 2018 | Word Image Representation Based on Visual Embeddings and Spatial Constraints for Keyword Spotting on Historical DocumentsabstractThis paper proposed a visual embeddings approach to capturing semantic relatedness between visual words. To be specific, visual words are extracted and collected from a word image collection under the Bag-of-Visual-Words framework. And then, a deep learning procedure is used for mapping visual words into embedding vectors in a semantic space. To integrate spatial constraints into the representation of word images, one word image is segmented into several sub-regions with equal size along rows and columns. After that, each sub-region can be represented as an average of embedding vectors, which is the centroid of the embedding vectors of all visual words within the same sub-region. By this way, one word image can be converted into a fixed-length vector by concatenating the corresponding average embedding vectors from its all sub-regions. Euclidean distance can be calculated to measure similarity between word images. Experimental results demonstrate that the proposed representation approach outperforms Bag-of-Visual-Words, visual language model, spatial pyramid matching, latent Dirichlet allocation, average visual word embeddings and recurrent neural network. Hongxi Wei, Hui Zhang 0031, Guanglai Gao |
ICPR | 3 |
| 2018 | Improving Mongolian Phrase Break Prediction by Using Syllable and Morphological Embeddings with BiLSTM Model
Rui Liu 0008, Feilong Bao, Guanglai Gao, Hui Zhang 0031, Yonghe Wang |
INTERSPEECH | 3 |
| 2018 | Mongolian Grapheme to Phoneme Conversion by Using Hybrid Approach
Zhinan Liu, Feilong Bao, Guanglai Gao, Suburi |
NLPCC (1) | 3 |
| 2018 | Phonologically Aware BiLSTM Model for Mongolian Phrase Break Prediction with Attention Mechanism
Rui Liu 0008, Feilong Bao, Guanglai Gao, Hui Zhang 0031, Yonghe Wang |
PRICAI (1) | 3 |
| 2017 | Supervised Feature Learning via Within-Class ReconstructionabstractFeature representation of data is a key issue for recognition related tasks. Inspired by the creative ability of human beings, in this paper we propose a novel feature learning framework named within-class reconstruction (WCR). In WCR, the feature representation of the input sample are used to reconstruct all the samples within the same class. We minimize the mean squared error (MSE) cost function to update feature extracting functions. Furthermore, most unsupervised learning methods such as auto-encoders could embed in the proposed framework. To evaluate the effectiveness of the proposed framework, CNN is used to extract the feature representations and reconstruct the within-class samples. The experimental results demonstrate that the representations learned by the proposed WCR achieve better performance than that of auto-encoders. All the codes have been made publicly available at https://github.com/step123456789/wcr. Yunxue Shao, Jiantao Zhou 0002, Guanglai Gao |
ICDAR | 3 |
| 2017 | Segmentation-Free Printed Traditional Mongolian OCR Using Sequence to Sequence with Attention ModelabstractMongolian Optical Character Recognition (OCR) systems are required for printed document digitization and Mongolian cultural resources utilization. Existing Mongolian OCR systems are based on segmentation. But, the Mongolian segmentation is more difficult than other languages. So, these methods are highly costly and error suffering. In this study, a segmentation-free based traditional Mongolian word recognition method is proposed. Specifically, we formalize the OCR task as a sequence to sequence mapping problem, in which the input Mongolian word image and the output textual string are treated as a sequence of image frames and a sequence of letters, respectively. A sequence to sequence with attention model is adopted to solve this problem. Experimental results on a dataset show the effectiveness of the proposed method. Hui Zhang 0031, Hongxi Wei, Feilong Bao, Guanglai Gao |
ICDAR | 4 |
| 2017 | Representing word image using visual word embeddings and RNN for keyword spotting on historical document imagesabstractVisual words of Bag-of-Visual-Words (BoVW) framework are independent each other, which results in not only discarding spatial orders between visual words but also lacking semantic information. This study is inspired by word embeddings that a similar embedding procedure is applied to a large number of visual words. By this way, the corresponding embedding vectors of the visual words can be formulated. For a word image, the average of embedding vectors of all visual words within the word image is taken as its embedding vector. Moreover, Recurrent Neural Network (RNN) is utilized to encode each word image into embeddings like an auto-encoder. The RNN embeddings and the visual word embeddings are complementary. In this study, all word images are represented by combining visual word embeddings and RNN embeddings. Experimental results show that the proposed representation approach is superior to the traditional BoVW, spatial pyramid matching and latent Dirichlet allocation. Hongxi Wei, Hui Zhang 0031, Guanglai Gao |
ICME | 3 |
| 2017 | Using Word Mover's Distance with Spatial Constraints for Measuring Similarity Between Mongolian Word Images
Hongxi Wei, Hui Zhang 0031, Guanglai Gao, Xiangdong Su |
ICONIP (4) | 3 |
| 2017 | Pseudo-Based Relevance Analysis for Information RetrievalabstractThe existing strategies in PRF (Pseudo relevance feedback) have insufficient attention to the user real query intention and fail to model the user intricate activities, which are plagued by the feedback terms sensitive problem. The major challenge in PRF lies in how to get the reliability relevant content for the user query. For this purpose, we propose a novel PRF approach by diversifying the feedback documents. An abstract pseudo document is proposed to represent the content of each feedback document so as to cover as diverse aspects of the feedback set as possible. The experimental results on real data sets show that our method can obtain the better feedback source to improve the overall the performance of PRF. Guanglai Gao |
ICTAI | 2 |
| 2017 | Multi-Target Ensemble Learning for Monaural Speech Separation
Hui Zhang 0031, Xueliang Zhang 0001, Guanglai Gao |
INTERSPEECH | 3 |
| 2017 | Research on Mongolian Speech Recognition Based on FSMN
Yonghe Wang, Feilong Bao, Guanglai Gao |
NLPCC | 4 |
| 2016 | Mongolian Named Entity Recognition System with Rich FeaturesabstractIn this paper, we first build a manually annotated named entity corpus of Mongolian. Then, we propose three morphological processing methods and study comprehensive features, including syllable features, lexical features, context features, morphological features and semantic features in Mongolian named entity recognition. Moreover, we also evaluate the influence of word cluster features on the system and combine all features together eventually. The experimental result shows that segmenting each suffix into an individual token achieves better results than deleting suffixes or using the suffixes as feature. The system based on segmenting suffixes with all proposed features yields benchmark result of F-measure=84.65 on this corpus. Weihua Wang 0006, Feilong Bao, Guanglai Gao |
COLING | 3 |
| 2016 | Convolutional neural network for robust pitch determinationabstractPitch is an important characteristic of speech and is useful for many applications. However, pitch determination in noisy conditions is difficult. In this paper, we propose a supervised learning algorithm to estimate pitch using a convolutional neural network (CNN). Specifically, we use a CNN for pitch candidate selection, and dynamic programming for pitch tracking. Our experimental results show that the proposed method can obtain accurate pitch estimation and they show good generalization ability to new speakers and noisy conditions. We credit the success to the use of CNN, which is suitable for modeling the shift-invariant spectral feature for pitch detection. Hong Su, Hui Zhang 0031, Xueliang Zhang 0001, Guanglai Gao |
ICASSP | 4 |
| 2016 | A novel image classifier based on Gaussian mixture language modelabstractIn this paper, we propose a novel Gaussian Mixture Language Model to address the issues of the traditional bag of visual words (BoVW) based model. We firstly take full advantage of image semantic information to learn a new distance metric which can achieve the minimal loss of image information, and then we train Gaussian Mixture Models (GMM) using this distance metric. Given a test image, a visual document is firstly constructed using this codebook, and then its category is determined by estimating the maximum probability using the language model under a specific category. Experiments show that the codebook generated by our method can effectively reflect the image semantic information and highly suitable the language model, and confirm that the proposed method is satisfactory and competitive in comparison with the traditional BoVW based method as well as other state of the art methods. Wei Wu 0032, Guanglai Gao |
ICASSP | 2 |
| 2016 | DNN-HMM for Large Vocabulary Mongolian Offline Handwriting RecognitionabstractIn this paper, we propose a large vocabulary Mongolian offline handwriting recognition system, using hidden Markov models (HMMs)-deep neural networks (DNN) hybrid architectures which shows superior performance on auto speech recognize (ASR) tasks. We select 50 sub-characters from all shape of Mongolian letters as the smallest modeling unit. First, a set of intensity features are extracted from each of the segmented word, which is based on a sliding window moving across each word image. Then, Multiple contextdependent Gaussian mixture model (GMM)-HMMs are trained by the features. At last a DNN which have 4 hidden layers are trained as a frame classifier, where the class labels are state labels assigned to each input frame through forced alignment using the context-dependent model. In order to validate the proposed model, extensive experiments were carried out using the MHW database which contains 100,000 handwritten words in training set, 5,000 in test set I and 14,085 in Test set II. The DNN-HMM w hich is trained on raw image pixels yields best performance on Test set I with an accuracy of 97.61% and on Test set II with an accuracy of 94.14%. Daoerji Fan, Guanglai Gao |
ICFHR | 2 |
| 2016 | A Connection Reduced Network for Similar Handwritten Chinese Character DiscriminationabstractOne difficulty in handwritten Chinese character recognition (HCCR) is due to the large number of similar characters. In this study, we propose a connection reduced network (CRN) to discriminate similar pairs. Each hidden neuron in CRN is restricted to has one input signal and the strength of this input is set as a variable which is selected from the input of the network. Experimental results based on 100 similar pairs demonstrate that the proposed method yields highly competitive test recognition results compared to the state-of-the-art methods, while consuming less memory and time resources. Yunxue Shao, Guanglai Gao, Chunheng Wang |
ICFHR | 2 |
| 2016 | LDA-Based Word Image Representation for Keyword Spotting on Historical Mongolian Documents
Hongxi Wei, Guanglai Gao, Xiangdong Su |
ICONIP (4) | 2 |
| 2016 | Mongolian Named Entity Recognition with Bidirectional Recurrent Neural NetworksabstractTraditional approaches to Named Entity Recognition almost heavily rely on feature engineering. In this paper, we introduce a kind of bidirectional recurrent neural network with long short memory (BLSTM) to capture bidirectional and long dependencies in a sentence without any feature set. Our model combines BLSTM network with Conditional Random Field (CRF) layer to jointly decode the best output. Additionally, this model inputs the concatenation of Mongolian morpheme and character representation. Experimental results show that the bidirectional recurrent neural networks significantly outperform traditional CRF model using manual features. Weihua Wang 0006, Feilong Bao, Guanglai Gao |
ICTAI | 3 |
| 2016 | A spatial-temporal trajectory clustering algorithm for eye fixations identificationabstractEye movements mainly consist of fixations and saccades. The identification of eye fixations plays an important role in the process of eye-movement data research. At present, there is no standard method for identifying eye fixations. In this paper, eye movements are regarded as spatial-temporal traj ectories. Hence, we present a spatial-temporal trajectory clustering algorithm for eye fixations identification. The main idea of the algorithm is based on Density-Based Spatial Clustering Algorithm with Noise (DBSCAN), which is commonly used in spatial clustering data. In order to apply DBSCAN to our spatial-temporal clustering data, we modified its original concept and algorithm. In addition, the optimum dispersion threshold (Eps) is derived automatically from the data sets with the aid of the `gap statistic' theory. Using the confusion matrix measurement method, we compared the classification results obtained by our algorithm with four other expert algorithms for eye fixations identification show the proposed algorithm demonstrated an equal or better performance. Also, the robustness of our algorithm to additional noise in Points of Gaze (PoGs) data and changes in sampling rate has been verified. Mingxin Yu, Yingzi Lin, Jeffrey Breugelmans, Xiangzhou Wang, Guanglai Gao |
Intell. Data Anal. | 6 |
| 2016 | A knowledge-based recognition system for historical Mongolian documents
Xiangdong Su, Guanglai Gao, Hongxi Wei, Feilong Bao |
Int. J. Document Anal. Recognit. | 2 |
| 2016 | Nonlinear discriminant analysis based on vanishing component analysis
Yunxue Shao, Guanglai Gao, Chunheng Wang |
Neurocomputing | 2 |
| 2016 | A Pairwise Algorithm Using the Deep Stacking Network for Speech Separation and Pitch EstimationabstractSpeech separation and pitch estimation in noisy conditions are considered to be a “chicken-and-egg” problem. On one hand, pitch information is an important cue for speech separation. On the other hand, speech separation makes pitch estimation easier when background noise is removed. In this paper, we propose a supervised learning architecture to solve these two problems iteratively. The proposed algorithm is based on the deep stacking network (DSN), which provides a method for stacking simple processing modules to build deep architectures. Each module is a classifier whose target is the ideal binary mask (IBM), and the input vector includes spectral features, pitch-based features and the output from the previous module. During the testing stage, we estimate the pitch using the separation results and update the pitch-based features to the next module. When embedded into the DSN, pitch estimation and speech separation each run several times. We obtain the final results from the last module. Systematic evaluations show that the proposed system results in both a high quality estimated binary mask and accurate pitch estimation and outperforms recent systems in its generalization ability. Xueliang Zhang 0001, Hui Zhang 0031, Shuai Nie 0001, Guanglai Gao |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2015 | A pairwise algorithm for pitch estimation and speech separation using deep stacking networkabstractPitch information is an important cue for speech separation. However, pitch estimation in noisy condition is also a task as challenging as speech separation. In this paper, we propose a supervised learning architecture which combines these two problems concisely. The proposed algorithm is based on deep stacking network (DSN) which provides a method of stacking simple processing modules in building deep architecture. In the training stage, an ideal binary mask is used as target. The input vector includes the outputs of lower module and frame-level features which consist of spectral and pitch-based features. In the testing stage, each module provides an estimated binary mask which is employed to re-estimate pitch. Then we update the pitch-based features to the next module. This procedure is embedded iteratively in DSN, and we obtain the final separation results from the last module of DSN. Systematic evaluations show that the proposed approach produces high quality estimated binary mask and outperforms recent systems in generalization. Hui Zhang 0031, Xueliang Zhang 0001, Shuai Nie 0001, Guanglai Gao |
ICASSP | 4 |
| 2015 | A multiple instances approach to improving keyword spotting on historical Mongolian document imagesabstractFor keyword spotting of historical Mongolian document images, when user provides different instance image for the same query keyword, the performance will vary a lot. This paper proposed an approach to solving the above problem. Particularly, the whole procedure of keyword spotting is divided into two stages. The main task of the first stage is to generate multiple ranking lists for a query keyword. And the aim of the second stage is to merge the multiple ranking lists to form a final ranking. In the first stage, the ranking list of one query keyword is firstly returned by traditional image matching and then a number of instances for the query keyword are obtained using pseudo relevant feedback. Next, each instance of the query keyword can return the corresponding ranking list separately. In the second stage, the multiple ranking lists from the multiple instances of the query keyword are combined by the data fusion technique. The final ranking will be taken as the retrieval results of the query keyword. The experimental results show that the proposed approach can significantly improve the performance of keyword spotting for the historical Mongolian document images. Hongxi Wei, Guanglai Gao, Xiangdong Su |
ICDAR | 2 |
| 2015 | Enhancing the Mongolian Historical Document Recognition System with Multiple Knowledge-Based Strategies
Xiangdong Su, Guanglai Gao, Hongxi Wei, Feilong Bao |
ICONIP (2) | 2 |
| 2015 | Nearest Neighbor with Multi-feature Metric for Image Annotation
Wei Wu 0032, Guanglai Gao |
ICONIP (4) | 2 |
| 2015 | Mongolian Inflection Suffix Processing in NLP: A Case StudyabstractInflection suffix is an important morphological characteristic of Mongolian words, since the suffixes express abundant syntactic and semantic meanings. In order to provide an informative introduction of it, this paper implements a case study on it. Through three Mongolian NLP tasks, we disclose the following information: (1) views of inflection suffix in NLP tasks, (2) Inflection suffix processing ways, (3) Inflection suffix effects on system performance and (4) some suffix related conclusion. Xiangdong Su, Guanglai Gao, Jing Wu 0011, Feilong Bao |
NLPCC | 2 |
| 2014 | A keyword retrieval system for historical Mongolian document images
Hongxi Wei, Guanglai Gao |
Int. J. Document Anal. Recognit. | 2 |
| 2013 | Segmentation-based Mongolian LVCSR approachabstractMongolian is an agglutinative language. Each root can be followed by several suffixes to formulate new words. This special word formation characteristic results in probably millions of Mongolian words, which is far beyond the coverage of the pronunciation dictionary of any current Mongolian speech recognition system. Moreover, even if the pronunciation dictionary is large enough to cover all of the Mongolian words, the recognition system still cannot perform well due to the problem of sample sparseness. In this paper, we propose a segmentation-based Mongolian Large Vocabulary Continuous Speech Recognition (LVCSR) approach and rebuild the corresponding acoustic model and language model. Experimental results show that, by converting most of these words into their corresponding In-Vocabulary form, the proposed approach effectively recognizes most of the Mongolian words and greatly improves the sample sparseness problem in the language model. Feilong Bao, Guanglai Gao, Xueliang Yan, Weihua Wang 0006 |
ICASSP | 2 |
| 2013 | Word Spotting Application in Historical Mongolian Document Images
Hongxi Wei, Guanglai Gao |
ICIC (1) | 2 |
| 2013 | Language Model for Cyrillic Mongolian to Traditional Mongolian Conversion
Feilong Bao, Guanglai Gao, Xueliang Yan |
NLPCC | 2 |
| 2011 | Classical Mongolian Words Recognition in Historical DocumentabstractThere are many classical Mongolian historical documents which are reserved in image form, and as a result it is difficult for us to explore and retrieve them. In this paper, we investigate the peculiarities of classical Mongolian documents and propose an approach to recognize the words in them. We design an algorithm to segment the Mongolian words into several Glyph Units(Glyph Unit abbr. GU). Each GU is consisted of no more than three characters. Then we used a three-stage method to recognize the GUs. At the first stage, all the GUs are classified into nine groups by decision tree using three features of the GUs. At the second stage, the GUs in each group are classified individually by five independent BP Neutral Networks whose inputs are other five feature vectors of the GUs. At the last stage, the five results of each GU group from the above five classifiers are combined to provide the final recognized result. The recognition rate of the Mongolian words in our experiment achieves 71%, indicating that our method is effective. Guanglai Gao, Xiangdong Su, Hongxi Wei, Yeyun Gong |
ICDAR | 1 |
| 2011 | A Method for Removing Inflectional Suffixes in Word Spotting of Mongolian KanjurabstractAccording to characteristics of Mongolian word-formation, a method for removing inflectional suffixes from word images of the Mongolian Kanjur is proposed in this paper. By removing inflectional suffixes, the amount of clusters equivalent indexing terms might be reduced in word spotting. For the above purpose, we need to determine whether or not one word image contains inflectional suffix. If the word image contains inflectional suffix, the inflectional suffix would be segmented from the word image. The proposed method is as follows: first, many parts are segmented from the bottom of the word image according to the cutting positions of the inflectional suffixes. Then, the segmented parts are represented by a number of profile features and classified by multi-BP neural networks. Finally, the outputs of BP are confirmed by template matching using DTW. Experimental results on our data set prove the feasibility of the proposed method. Hongxi Wei, Guanglai Gao, Yulai Bao |
ICDAR | 2 |
| 2006 | A Mongolian Speech Recognition System Based on HMM
Guanglai Gao, Biligetu, Nabuqing, Shuwu Zhang |
ICIC (2) | 1 |