Xiangdong Su

dblp:25/10184 · DBLP profile ↗
← Back
70ranked-venue papers
7as first author
48since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 47 · 7 first-author · 30 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 16 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 FlorE: Integrating Full Lorentz Group and Directional Offsets for Effective Knowledge Graph Embedding
abstract
Knowledge Graph Embedding (KGE) aims to map entities and relationships into a continuous vector space to facilitate reasoning and downstream tasks. Although previous KGE methods based on Euclidean, complex spaces, or hyperbolic spaces have performed well, they still struggle to effectively model Z-Paradox relation patterns which account for a large proportion in each knowledge graph. To address this issue, we propose a novel KGE method **FlorE** which integrates full Lorentz Group and directional offset operation in hyperbolic space for KGE task. Specifically, we incorporates the full Lorentz Group to enable the same relation in knowledge graph (KG) to perform indefinite isometry, thus avoiding the overlapping of entities. Meanwhile, we implement directional offset operation via exponential mapping to transform the relations to the same Lorentz manifold of the entities, thus maintaining geometric consistency for the relations and entities in KG. By integrating these two techniques, FlorE can effectively model the Z-Paradox relation patterns and improve the representation learning ability for KGs. Experiments on the five benchmark datasets demonstrate that our method achieves state-of-the-art performance. For the Z-Paradox relation patterns, the improvement achieves **26.7%**, **15.6%**, **35.4%**, **33.7%**, and **31.5%** on FB15k-237, WN18RR, CoDEx-S, CoDEx-M and CoDEx-L, respectively.
Zehua Duo, Jiang Li 0013, Xiangdong Su, Guanglai Gao
AAAI3
2026 CEDAR: A Chinese Evaluation Dataset for Computational Argumentation
abstract
Tian Lan, Jiang Li, Rong Yan, Feilong Bao, Weihua Wang, Guanglai Gao, Xiangdong Su. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jiang Li 0013, Feilong Bao, Weihua Wang 0006, Guanglai Gao, Xiangdong Su
ACL (1)7
2026 Who Wrote This Line? Evaluating the Detection of LLM-Generated Classical Chinese Poetry
abstract
Jiang Li, Tian Lan, Shanshan Wang, Dongxing Zhang, Dianqing Lin, Guanglai Gao, Derek F. Wong, Xiangdong Su. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jiang Li 0013, Shanshan Wang 0009, Zdongxing, Dianqing Lin, Guanglai Gao, Derek F. Wong, Xiangdong Su
ACL (1)8
2026 Know Your Place: Diagnosing Implicit Social Adaptation Failures in Chinese Large Language Models
abstract
Yu Tian, Jie Xing, Ziming Li, Jiang Li, Zehua Duo, Tian Lan, Xu Liu, Guanglai Gao, Xiangdong Su. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jiang Li 0013, Zehua Duo, Guanglai Gao, Xiangdong Su
ACL (1)9
2026 GSMP: Geometry-Structured Masked Pre-training with Multi-granularity Objectives and Curriculum Learning for Geometric Problem Solving
Xingxiang Zhou, Minzhi Zhang, Guanglai Gao, Xiangdong Su
ICDAR (2)7
2026 G2I: A Progressive Structure-to-Detail Curriculum Training Strategy for Handwritten Mathematical Expression Recognition
Minzhi Zhang, Xingxiang Zhou, Xiangdong Su
ICDAR (2)6
2026 A knowledge prompt augmented lightweight multimodal language assistant for biomedicine
Lei Liu 0079, Xiangdong Su, Xingxiang Zhou, Guanglai Gao
Eng. Appl. Artif. Intell.2
2026 Hyperbolic-Based Cross-Modal Semantic Remodeling Network for Zero-Shot Sketch-Based Image Retrieval
abstract
The Zero-Shot Sketch-Based Image Retrieval (ZS-SBIR) task aims to retrieve images associated with sketches from unseen classes, bringing great convenience to the engineering field. To address the modality gap, most existing works project images and sketches into a shared Euclidean space. However, the hierarchical structure of image data makes the Euclidean space not the optimal choice as an embedding space for representing complex structured image data. Meanwhile, existing text and hierarchical models are not effective enough for addressing the problem of knowledge transfer. To address these issues, this article proposes an original Hyperbolic-Based Cross-Modal Semantic Remodeling Network (called HCMSN) for ZS-SBIR. Specifically, this article proposes to extract category-level word embeddings based on BERT model, then align image features and sketches with the word embeddings using adversarial methods. Meanwhile, this article further proposes a cross-modal retrieval feature reconstruction network for improving the informativeness and robustness of retrieval features. Moreover, this article presents a feature projection network that maps the retrieval features to the hyperbolic space to generate the hyperbolic retrieval features, thus effectively representing the data with hierarchical structure. Extensive experiments demonstrate that the mAP@all of our HCMSN model surpasses CNN-based models by 20.9% on the Sketchy dataset, 1.2% on the more difficult TU-Berlin dataset, and 13.6% on the more challenging QuickDraw dataset.
Xiangdong Su, Feilong Bao, Guanglai Gao
ACM Trans. Multim. Comput. Commun. Appl.3
2025 SSAN: A Symbol Spatial-Aware Network for Handwritten Mathematical Expression Recognition
abstract
The great challenge of handwritten mathematical expression recognition (HMER) is the complex structures of the expressions, which are directly related to the symbol spatial positions. Existing HMER methods typically employ attention mechanisms in the decoder of their models to implicitly perceive the symbol positions, or employ symbol counting and tree-based strategies to model the symbol spatial relation. However, these methods still cannot effectively capture the structural information of formulas, thus negatively impacting the symbol decoding in HMER. To deal with this problem and enhance the HMER performance, this paper proposes a novel auxiliary task, namely predicting the symbol spatial distribution map of handwritten expression images. On such basis, this paper designs a symbol spatial-aware network (SSAN) for this task, which is jointly optimized with the HMER model. Specifically, considering the similarity of the symbol spatial positions between the handwritten mathematical expression images and their corresponding printed templates, we obtain the symbol spatial distribution map by first generating printed templates from LaTeX ground-truth for handwritten formula images and then replacing the connected components of printed templates with 2D Gaussian distribution maps of the same size. Meanwhile, due to the loose alignment of the symbol spatial positions between handwritten and printed formula images, and misclassification of similar symbols, we further propose a coarse-to-fine alignment strategy and an attention-guided symbol masking strategy in SSAN to tackle these issues. Extensive experiments demonstrate that SSAN significantly improves the recognition performance of the HMER models, and the proposed auxiliary tasks are more effective in enhancing HMER performance than existing auxiliary tasks.
Xiangdong Su, Xingxiang Zhou, Guanglai Gao
AAAI2
2025 A Mutual Information Perspective on Knowledge Graph Embedding
abstract
Knowledge graph embedding techniques have emerged as a critical approach for addressing the issue of missing relations in knowledge graphs. However, existing methods often suffer from limitations, including high intra-group similarity, loss of semantic information, and insufficient inference capability, particularly in complex relation patterns such as 1-N and N-1 relations. To address these challenges, we introduce a novel KGE framework that leverages mutual information maximization to improve the semantic representation of entities and relations. By maximizing the mutual information between different components of triples, such as (h, r) and t, or (r, t) and h, the proposed method improves the model’s ability to preserve semantic dependencies while maintaining the relational structure of the knowledge graph. Extensive experiments on benchmark datasets demonstrate the effectiveness of our approach, with consistent performance improvements across various baseline models. Additionally, visualization analyses and case studies demonstrate the improved ability of the MI framework to capture complex relation patterns.
Jiang Li 0013, Xiangdong Su, Zehua Duo, Xiaotao Guo, Guanglai Gao
ACL (1)2
2025 C3LRSO: A Chinese Corpus for Complex Logical Reasoning in Sentence Ordering
abstract
Sentence ordering is the task of rearranging a set of unordered sentences into a coherent and logically consistent sequence. Recent work has primarily used pre-trained language models, achieving significant success in the task. However, existing sentence ordering corpora are predominantly in English, and comprehensive benchmark datasets for non-English languages are unavailable. Meanwhile, current datasets often insert specific markers into paragraphs, inadvertently making the logical sequence between sentences more apparent and reducing the models’ ability to handle genuinely unordered sentences in real applications. To address these limitations, we develop C3LRSO, a high-quality Chinese sentence ordering dataset that overcomes the aforementioned shortcomings by providing genuinely unordered sentences without artificial segmentation cues. Furthermore, given the outstanding performance of large language models on NLP tasks, we evaluate these models on our dataset for this task. Additionally, we propose a simple yet effective parameter-free approach that outperforms existing methods on this task. Experiments demonstrate the challenging nature of the dataset and the strong performance of our proposed method. These findings highlight the potential for further research in sentence ordering and the development of more robust language models. Our dataset is freely available at https://github.com/JasonGuo1/C3LRSO.
Xiaotao Guo, Jiang Li 0013, Xiangdong Su, Fujun Zhang 0004
COLING3
2025 F²Bench: An Open-ended Fairness Evaluation Benchmark for LLMs with Factuality Considerations
abstract
Warning: This paper contains content that may be offensive or harmful With the growing adoption of large language models (LLMs) in NLP tasks, concerns about their fairness have intensified.Yet, most existing fairness benchmarks rely on closed-ended evaluation formats, which diverge from realworld open-ended interactions.These formats are prone to position bias and introduce a "minimum score" effect, where models can earn partial credit simply by guessing.Moreover, such benchmarks often overlook factuality considerations rooted in historical, social, physiological, and cultural contexts, and rarely account for intersectional biases.To address these limitations, we propose F 2 Bench: an openended fairness evaluation benchmark for LLMs that explicitly incorporates factuality considerations.F 2 Bench comprises 2,568 instances across 10 demographic groups and two openended tasks.By integrating text generation, multi-turn reasoning, and factual grounding, F 2 Bench aims to more accurately reflect the complexities of real-world model usage.We conduct a comprehensive evaluation of several LLMs across different series and parameter sizes.Our results reveal that all models exhibit varying degrees of fairness issues.We further compare open-ended and closedended evaluations, analyze model-specific disparities, and provide actionable recommendations for future model development.Our code and dataset are publicly available at https: //github.com/VelikayaScarlet/F2Bench.
Jiang Li 0013, Yemin Wang, Xiangdong Su, Guanglai Gao
EMNLP5
2025 Multilingual Parameter-Sharing Adapters: A Method for Optimizing Low-Resource Neural Machine Translation
abstract
Adapter-based Multilingual Neural Machine Translation (MNMT) has become a significant approach in low-resource language translation by mitigating data imbalances between high-resource and low-resource language pairs and reducing training costs. However, existing adapter-based methods lack generalization in cross-lingual settings, particularly under low-resource conditions, where their scalability is limited. Additionally, current methods often introduce independent adapter modules for each language, leading to a linear increase in model parameters with the number of languages. To address these challenges, we propose a multilingual parameter-sharing adapter approach. Moreover, we introduce a neural architecture search (NAS)-based strategy to improve translation performance. Experimental results demonstrate that the multilingual parameter-sharing adapter exhibits competitive performance on both low-resource and high-resource datasets. The multilingual parameter-sharing adapter method has only 400K trainable parameters, which is 20× lower than the parameters of the traditional adapter method.
Yonghe Wang, Xiangdong Su, Feilong Bao
ICASSP4
2025 Structural-Aware Disentangled Learning with CLIP for Hyperbolic Zero-Shot Sketch-Based Image Retrieval
abstract
The zero-shot sketch-based image retrieval task faces two key challenges: domain gap and knowledge transfer. Our innovation is recognizing that directly aligning cross-domain features weakens the discriminative ability of the model, as it overlooks the asymmetry between sketches and images. Additionally, Euclidean space is inadequate for capturing the hierarchical structure, which limits the performance of the model on complex data. To address these issues, we propose a Structural-Aware Disentangled Learning network (termed SADLnet) that incorporates CLIP and hyperbolic geometry. Specifically, we use CLIP to extract visual features from each domain to enhance the domain generalization of the model. Furthermore, we design a structure-guided disentanglement strategy to decompose image representations into sketch-related and sketch-unrelated features, addressing the domain gap. Moreover, we project the retrieval features into hyperbolic space to capture hierarchical information, improving feature discrimination in retrieval tasks. Extensive experiments demonstrate that SADLnet establishes new state-of-the-art performance on three datasets.
Feilong Bao, Xiangdong Su, Guanglai Gao
ICASSP4
2025 Task-Decoupled Bézier Surface Constraint for Uneven Low-Light Image Enhancement
Xingxiang Zhou, Xiangdong Su, Guanglai Gao
ICCV2
2025 MedVSA: Medical Visual Spoken-Question Answering
abstract
With the rapid advancement of technology, smart healthcare has made significant progress, particularly in medical visual question answering (MedVQA). However, current MedVQA primarily relies on text, whereas practical applications often involve spoken interactions, such as in medical consultations and mobile-based queries. To bridge this gap, we propose a novel task, medical visual spoken-question answering (MedVSA), extending the conventional medical image and text-based question-answering paradigm to include spoken interactions for enhanced applicability. We expand upon four commonly used MedVQA datasets, namely VQA-RAD, SLAKE, PathVQA, and OVQA, by leveraging Alibaba Cloud speech synthesis technology to convert text questions into spoken questions. Various strategies are incorporated to ensure diverse and realistic speech synthesis. The resulting dataset comprises images, corresponding text, and synthesized speech data. Subsequently, we design a single-stage model and a two-stage model to tackle the MedVSA task. For the single-stage model, we directly input the speech and images into our designed whisper self-distillation model to obtain the results. For the two-stage model, we first use the Whisper model to convert the speech into text, then input the converted text and medical images into the self-distillation model to obtain the results. We provide two solutions for the MedVSA task and establish two baselines. Experimental results show that the two-stage model significantly outperforms the single-stage model, indicating that text conversion is crucial for solving MedVSA. This study advances smart healthcare developments by proposing MedVSA and designing two baselines tailored to its specificities. Source code and MedVSA dataset are available at https://github.com/Alivelei/MedVSA.
Lei Liu 0079, Xiangdong Su, Guanglai Gao
ICMR2
2025 Fourier Self-Adaptation for Transferring General Pretrained Models to Specific Domains
abstract
While pre-trained models in the general domain have proliferated, existing methods for transferring these models to specific domains often depend on source domain data for distribution alignment and are typically tailored for single tasks. We propose a source data-free approach, Fourier Self-Adaptation (FSA), which effectively adapts general models to a wide range of specific domains. Our method leverages the distinct properties of Fourier phase and amplitude: phase contains high-level structural and positional information, which is less affected by domain shifts, while amplitude contains details and brightness information, which is more affected by domain shifts. FSA adjusts the image distribution by initializing a trainable adaptive image from a normal distribution. It then interpolates the amplitude of the target domain image with that of the adaptive image, where the interpolation ratio is dynamically controlled by learnable weight and bias. During training, the model captures advanced phase information of the target image and refines the data distribution through amplitude interpolation. Additionally, a dual regularization loss constrains the model representation, encouraging it to focus on the intrinsic relationships of the target domain data while discarding irrelevant knowledge. We evaluate FSA using general pre-trained models on 11 unimodal image classification datasets and 6 multimodal visual question answering datasets, covering specific domains such as radiology, pathology, remote sensing, and art. Our method consistently achieves state-of-the-art performance across multiple datasets, with performance improvements ranging from 1% to 8% compared to basic pre-trained models. Source code are available at https://github.com/Alivelei/FSA.
Lei Liu 0079, Xiangdong Su, Guanglai Gao
ACM Multimedia2
2025 Mitigating Heterogeneity among Factor Tensors via Lie Group Manifolds for Tensor Decomposition Based Temporal Knowledge Graph Embedding
abstract
Jiang Li, Xiangdong Su, Guanglai Gao. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Jiang Li 0013, Xiangdong Su, Guanglai Gao
NAACL (Long Papers)2
2025 Domain disentanglement and fusion based on hyperbolic neural networks for zero-shot sketch-based image retrieval
Xiangdong Su, Yonghe Wang, Feilong Bao, Guanglai Gao
Inf. Process. Manag.3
2025 Fine-grained Automatic Augmentation for handwritten character recognition
Wei Chen 0166, Xiangdong Su, Hongxu Hou
Pattern Recognit.2
2024 Leveraging Convolutional Models as Backbone for Medical Visual Question Answering
abstract
Convolutional neural networks (CNNs) have made significant contributions to computer vision and offer the advantages of higher training efficiency and lower model complexity. However, their application as the backbone in medical visual question answering (MedVQA) remains an open question. To address this issue, we employ popular convolutional models, including ResNet, DenseNet, and ShuffleNet, as the foundation for MedVQA, achieving outstanding performance. Different backbones can be tailored to diverse real-world scenarios. The central challenge in utilizing CNNs for visual question answering is effectively managing textual features and integrating multi-modal information. To overcome this challenge, we design a novel global interaction attention (GIA) that facilitates efficient interactions between text and image features. Additionally, we utilize the dot product before the classifier output to enhance visual and textual modal fusions. To further enhance model performance, we propose a novel multi-modal hidden mixup (MHidMix) technique for data augmentation, which involves interpolating hidden states during model training. This data augmentation technique smoothes the decision boundary without the need for complex sample selection, further improving model performance. Experimental results underscore the versatility of our proposed framework across various convolutional models, leading to outstanding performance on four MedVQA datasets. Notably, we achieved an accuracy increase of 9.4% on the PathVQA dataset and 4.5% on the OVQA dataset.
Lei Liu 0079, Xiangdong Su, Guanglai Gao
BIBM2
2024 Optimizing Transformer and MLP with Hidden States Perturbation for Medical Visual Question Answering
abstract
Optimizing model performance is a crucial objective in medical visual question answering (MedVQA), and a wide range of techniques have been developed to achieve this goal. In this paper, we propose a novel technique called network state perturbation, which distinguishes itself from existing research. Our study designs four innovative methods to modify the hidden states within the network and improve the performance of multimodal models in the MedVQA task. Specifically, we evaluate these methods on both general transformer and multilayer perceptron (MLP) models, which allow for hidden state adjustments at each layer. The four introduced methods are as follows: (1) Randomly Set Zero, which assigns zeros to the hidden states of different modalities in the network; (2) Randomly Replace Content, which performs interpolation between the hidden states of the text sequence and the image sequence; (3) Randomly Add Gaussian Noise, which adds Gaussian noise to the hidden states of different modalities; and (4) Pair Interpolation, which interpolates the hidden states of different modalities of the current sample with those of other samples and performs corresponding label interpolation to facilitate model training. Our experimental results demonstrate that the proposed hidden state perturbation methods significantly enhance the performance of various transformer and MLP models on multiple MedVQA datasets, without requiring additional data or computational resources during training. These findings highlight the potential of hidden state perturbation as a novel model improvement technique for multimodal models in the medical domain.
Lei Liu 0079, Xiangdong Su, Guanglai Gao
BIBM2
2024 Learning Frequency Adaptation for Cross-domain Medical Image Segmentation
abstract
In medical image segmentation, some recent methods improve domain adaptation performance through frequency domain adaptation and frequency mixup. However, these approaches have two limitations: (1) frequency domain adaptation ignores the adverse effects of high-frequency noise on model generalization, and (2) frequency mixup confuses semantic information. To address these issues, we propose a novel frequency adaptation approach for medical image segmentation including low-frequency component alignment (LFCA) and random amplitude cutmix (RAC). Since low-frequency contains major image information and high-frequency noise affects adaptation, we leverage discrete wavelet transform to decompose images into low and high-frequency components. LFCA aligns the domain distribution of low frequencies and high frequencies are passed to the decoder via skip connections. In addition, we design RAC to generate diverse augmented samples through amplitude cutmix while avoiding distortion of the original distributions. Experiments on benchmark datasets validate the efficacy of our proposed approach.
Lei Liu 0079, Xiangdong Su, Guanglai Gao
BIBM3
2024 FSAM: Fine-tuning SAM encoder and decoder for Medical Image Segmentation
abstract
Recently, the Segment Anything Model (SAM), a large pre-training model, has achieved excellent results on natural image segmentation tasks and received extensive attention. However, SAM on medical image segmentation is unsatisfactory since there are significant differences between natural and medical images. How to extend the SAM’s powerful segmentation capabilities to the medical domain requires further exploration. To this end, we propose a simple and efficient fine-tuning approach for SAM that does not require large-scale data called FSAM. Specifically, FSAM simultaneously fine-tunes the encoder and decoder while freezing the prompt encoder. This allows the encoder to extract medical image features effectively and guide decoder segmentation predictions. FSAM achieves state-of-the-art results on eight public medical image datasets, outperforming SAM by +15.39%. Moreover, FSAM exhibits weaker sample scale dependence. The proposed framework further improves the segmentation capabilities of SAM in medical images.
Lei Liu 0079, Xiangdong Su, Guanglai Gao
BIBM3
2024 3D Blur Kernel on Gaussian Splatting
Yongchao Lin, Xiangdong Su
BMVC2
2024 Hyperbolic Representations for Prompt Learning
abstract
Continuous prompt tuning has gained significant attention for its ability to train only continuous prompts while freezing the language model. This approach greatly reduces the training time and storage for downstream tasks. In this work, we delve into the hierarchical relationship between the prompts and downstream text inputs. In prompt learning, the prefix prompt acts as a module to guide the downstream language model, establishing a hierarchical relationship between the prefix prompt and subsequent inputs. Furthermore, we explore the benefits of leveraging hyperbolic space for modeling hierarchical structures. We project representations of pre-trained models from Euclidean space into hyperbolic space using the Poincaré disk which effectively captures the hierarchical relationship between the prompt and input text. The experiments on natural language understanding (NLU) tasks illustrate that hyperbolic space can model the hierarchical relationship between prompt and text input. We release our code at https://github.com/myaxxxxx/Hyperbolic-Prompt-Learning.
Xiangdong Su, Feilong Bao
LREC/COLING2
2024 TransERR: Translation-based Knowledge Graph Embedding via Efficient Relation Rotation
abstract
This paper presents a translation-based knowledge geraph embedding method via efficient relation rotation (TransERR), a straightforward yet effective alternative to traditional translation-based knowledge graph embedding models. Different from the previous translation-based models, TransERR encodes knowledge graphs in the hypercomplex-valued space, thus enabling it to possess a higher degree of translation freedom in mining latent information between the head and tail entities. To further minimize the translation distance, TransERR adaptively rotates the head entity and the tail entity with their corresponding unit quaternions, which are learnable in model training. We also provide mathematical proofs to demonstrate the ability of TransERR in modeling various relation patterns, including symmetry, antisymmetry, inversion, composition, and subrelation patterns. The experiments on 10 benchmark datasets validate the effectiveness and the generalization of TransERR. The results also indicate that TransERR can better encode large-scale datasets with fewer parameters than the previous translation-based models. Our code and datasets are available at https://github.com/dellixx/TransERR.
Jiang Li 0013, Xiangdong Su, Fujun Zhang 0004, Guanglai Gao
LREC/COLING2
2024 APOLLO: An Optimized Training Approach for Long-form Numerical Reasoning
abstract
Long-form numerical reasoning aims to generate a reasoning program to calculate the answer for a given question. Previous work followed a retriever-generator framework, where the retriever selects key facts from a long-form document, and the generator generates a reasoning program based on the retrieved facts. However, they treated all facts equally without considering the different contributions of facts with and without numerical information. Furthermore, they ignored program consistency, leading to the wrong punishment of programs that differed from the ground truth. In order to address these issues, we proposed APOLLO (An optimized training aPproach fOr Long-form numericaL reasOning), to improve long-form numerical reasoning. APOLLO includes a number-aware negative sampling strategy for the retriever to discriminate key numerical facts, and a consistency-based reinforcement learning with target program augmentation for the generator to ultimately increase the execution accuracy. Experimental results on the FinQA and ConvFinQA leaderboards verify the effectiveness of our proposed methods, achieving the new state-of-the-art.
Jiashuo Sun, Hang Zhang 0029, Chen Lin 0001, Xiangdong Su, Yeyun Gong, Jian Guo 0016
LREC/COLING4
2024 EpLSA: Synergy of Expert-prefix Mixtures and Task-Oriented Latent Space Adaptation for Diverse Generative Reasoning
abstract
Existing models for diverse generative reasoning still struggle to generate multiple unique and plausible results. Through an in-depth examination, we argue that it is critical to leverage a mixture of experts as prefixes to enhance the diversity of generated results and make task-oriented adaptation in the latent space of the generation models to improve the quality of the responses. At this point, we propose EpLSA, an innovative model based on the synergy of expert-prefix mixtures and task-oriented latent space adaptation for diverse generative reasoning. Specifically, we use expert-prefixes mixtures to encourage the model to create multiple responses with different semantics and design a loss function to address the problem that the semantics is interfered by the expert-prefixes. Meanwhile, we design a task-oriented adaptation block to make the pre-trained encoder within the generation model more effectively adapted to the pre-trained decoder in the latent space, thus further improving the quality of the generated text. Extensive experiments on three different types of generative reasoning tasks demonstrate that EpLSA outperforms existing baseline models in terms of both the quality and diversity of the generated outputs. Our code is publicly available at https://github.com/IMU-MachineLearningSXD/EpLSA.
Fujun Zhang 0004, Xiangdong Su, Jiang Li 0013, Guanglai Gao
LREC/COLING2
2024 Exploring the Synergy of Dual-path Encoder and Alignment Module for Better Graph-to-Text Generation
abstract
The mainstream approaches view the knowledge graph-to-text (KG-to-text) generation as a sequence-to-sequence task and fine-tune the pre-trained model (PLM) to generate the target text from the linearized knowledge graph. However, the linearization of knowledge graphs and the structure of PLMs lead to the loss of a large amount of graph structure information. Moreover, PLMs lack an explicit graph-text alignment strategy because of the discrepancy between structural and textual information. To solve these two problems, we propose a synergetic KG-to-text model with a dual-path encoder, an alignment module, and a guidance module. The dual-path encoder consists of a graph structure encoder and a text encoder, which can better encode the structure and text information of the knowledge graph. The alignment module contains a two-layer Transformer block and an MLP block, which aligns and integrates the information from the dual encoder. The guidance module combines an improved pointer network and an MLP block to avoid error-generated entities and ensures the fluency and accuracy of the generated text. Our approach obtains very competitive performance on three benchmark datasets. Our code is available from https://github.com/IMu-MachineLearningsxD/G2T.
Tianxin Zhao, Yingxin Liu, Xiangdong Su, Jiang Li 0013, Guanglai Gao
LREC/COLING3
2024 Efficient Speech-to-Text Translation: Progressive Pruning for Accelerated Speech Pre-trained Model
abstract
Recently, speech pre-trained models based on the Transformer architecture have become very popular for speech-to-text translation tasks. However, computing representation outputs of speech pre-trained models is highly time-consuming, primarily due to the length of speech sequences far exceeding the corresponding texts, leading to quadratic computation costs for the self-attention module. To address this issue, we propose a novel pruning method that progressively reduces representation sequence length layer by layer. We leverage attention scores to calculate importance scores for all tokens. Additionally, we introduce both fixed and scheduled pruning rate strategies to determine which tokens should be retained. Experiments demonstrate that our approach reduces the output tokens of the speech pre-trained model by 55%, with only a 0.7% performance decrease, and improves in practice encoding speed up to 1.76 ×. Our method also is an orthogonal and complementary direction to efficient speech pre-trained models. We release our code at https://github.com/myaxxxxx/pruning.
Yonghe Wang, Xiangdong Su, Feilong Bao
ICME3
2024 MEMix: Improving HMER with Diverse Formula Structure Augmentation
abstract
Handwritten Mathematical Expression Recognition (HMER) aims to transform images of mathematical expressions (MEs) into corresponding LaTeX sequences. However, the inherent complexity of 2D formula structures often misaligns with the 1D LaTeX sequences, resulting in decreased robustness in recognition models. A primary factor exacerbating this issue is the scarcity of annotated ME images with complex structures, which hinders the models to learning to good representation and adaptability for MEs. In this paper, drawing inspiration from Mixup, we introduce a data augmentation method called Mathematical Expression Mix (MEMix). This method is capable of generating typical structures in formulas, including radicals, fractions, and annotations, by employing straightforward matrix operations. Compared to alternative data augmentation methods, MEMix provides faster and more cost-effective computation, enabling online augmentation that improves training efficiency. Experiments demonstrate that MEMix significantly enhances the performance of the baseline model on the benchmark datasets.
Xiangdong Su, Xingxiang Zhou, Guanglai Gao
ICME2
2024 Design and Optimization of a Dual-Rotor and Flat-Type-Stator Transverse Flux Machine
abstract
The development and application of transverse flux machines (TFMs) with permanent magnet (PM) excitation have received more and more attention in recent years because of the high torque density at low rotating speed applications. This paper proposes a novel topology structure of TFM with dual-rotor and flat-type-stator (DRFTS-TFM). One of the advantages of the proposed structure is that the sandwiched design greatly reduces the manufacturing and assembling difficulties. Another advantage is that the number of the armature phase can be easily changed by converting the mechanical angle of the stator. To investigate the best electromagnetic torque performance of the machine, the multi-objective particle swarm optimization (MOPSO) algorithm is adapted and validated by the finite element method (FEM). The corresponding electromagnetic performances in different conditions are analyzed.
Zhijun Ou, Hang Zhao 0010, Liyang Liu, Xiangdong Su, Hui Wang 0147
IECON4
2024 Ensuring Safe and High-Quality Outputs: A Guideline Library Approach for Language Models
abstract
Yi Luo, Zhenghao Lin, YuHao Zhang, Jiashuo Sun, Chen Lin, Chengjin Xu, Xiangdong Su, Yelong Shen, Jian Guo, Yeyun Gong. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Zhenghao Lin, Jiashuo Sun, Chen Lin 0001, Chengjin Xu, Xiangdong Su, Yelong Shen, Jian Guo 0016, Yeyun Gong
NAACL-HLT7
2024 Quat-DGNet: Enhancing 3D Dense Captioning with Quaternion-Based Spatial Offsets and Dynamic Neighborhood Graphs
Xiangdong Su, Jiang Li 0013, Fujun Zhang 0004
PRCV (6)2
2023 TeAST: Temporal Knowledge Graph Embedding via Archimedean Spiral Timeline
abstract
Temporal knowledge graph embedding (TKGE) models are commonly utilized to infer the missing facts and facilitate reasoning and decision-making in temporal knowledge graph based systems.However, existing methods fuse temporal information into entities, potentially leading to the evolution of entity information and limiting the link prediction performance of TKG.Meanwhile, current TKGE models often lack the ability to simultaneously model important relation patterns and provide interpretability, which hinders their effectiveness and potential applications.To address these limitations, we propose a novel TKGE model which encodes Temporal knowledge graph embeddings via Archimedean Spiral Timeline (TeAST), which maps relations onto the corresponding Archimedean spiral timeline and transforms the quadruples completion to 3th-order tensor completion problem.Specifically, the Archimedean spiral timeline ensures that relations that occur simultaneously are placed on the same timeline, and all relations evolve over time.Meanwhile, we present a novel temporal spiral regularizer to make the spiral timeline orderly.In addition, we provide mathematical proofs to demonstrate the ability of TeAST to encode various relation patterns.Experimental results show that our proposed model significantly outperforms existing TKGE methods.
Jiang Li 0013, Xiangdong Su, Guanglai Gao
ACL (1)2
2023 Multi-task Learning for Mongolian Morphological Analysis
Qing-Dao-Er-Ji Ren, Xiangdong Su, Yatu Ji, Aodengbala, Guiping Liu
ICANN (9)3
2023 FLPA: A fast label propagation algorithm for detecting overlapping community structure
Rong Yan 0001, Wei Yuan 0013, Xiangdong Su
Expert Syst. Appl.3
2023 Noise-Separated Adaptive Feature Distillation for Robust Speech Recognition
abstract
This letter makes an improvement on feature-based knowledge distillation for robust speech recognition. The use of distillation techniques in speech recognition has been demonstrated to improve the robustness of the system. In this letter, we propose a noise-separated adaptive feature distillation method, including an adaptive distillation position selection strategy and a noise separation mechanism, assuming that there is a common network structure between the student and teacher. The proposed method has two improvements. First, distillation positions can be adaptively selected in each iteration by comparing loss values computed on intermediate representations of the student and the teacher, increasing the flexibility of knowledge transfer during distillation. Second, a noise separation module is proposed to constrain noise information elimination by explicitly separating the speech information and the noise information in noisy speech, which reduces the interference of noise information during distillation. Therefore, a better recognition performance is demonstrated with the proposed method compared to the standardized feature-based knowledge distillation method.
Honglin Qu, Xiangdong Su, Yonghe Wang, Guanglai Gao
IEEE Signal Process. Lett.2
2022 How Well Apply Multimodal Mixup and Simple MLPs Backbone to Medical Visual Question Answering?
abstract
Although current methods have significantly improved the performance of medical visual question answering (Med-VQA), there are still two aspects worth exploring, namely the simplification of model structure and the effective model training on small-scale data. Different from the previous Med-VQA model, this paper only employs multi-layer perceptrons (MLPs) as the backbone network for feature extraction and modal fusion and designs a Med-VQA model on such basis, which achieves superior performance with a simple backbone network. To enhance model generalization, we design multimodal mixup (M-Mixup) to augment images and questions separately, which effectively alleviates the problem of insufficient training samples in the Med-VQA task. To prevent the destruction of the feature relationship when tokenizing the medical image, we design pooling tokens (PTs), a simple downsampling structure to capture fine-grained visual features without affecting the parameters and FLOPs of the entire model. Experimental results demonstrate that our model achieves state-of-the-art on the SLAKE, and obtains a remarkably competitive performance on the VQA-RAD. The source code and models are available at https://github.com/Alivelei/M-Mixup.
Lei Liu 0079, Xiangdong Su
BIBM2
2022 End-to-End Large-Scale Image Retrieval Network with Convolution and Vision Transformers
Feilong Bao, Xiangdong Su, Weihua Wang 0006, Guanglai Gao
ICANN (4)3
2022 Script-Level Word Sample Augmentation for Few-Shot Handwritten Text Recognition
Xiangdong Su
ICFHR2
2022 Medical Visual Question Answering via Targeted Choice Contrast and Multimodal Entity Matching
Lei Liu 0079, Xiangdong Su
ICONIP (2)3
2022 A Transformer-based Medical Visual Question Answering Model
abstract
While the Transformer architecture has been widely used in natural language processing tasks and computer vision tasks, its application in medical visual question answering is still limited. Most current methods rely on an image extractor to obtain visual features and a text extractor to capture semantic features, and then a fusion module to merge the information from the two modalities to predict the final result. In contrast, this paper proposes a novel Transformer-based medical vision question answering model, called MQAT, in which an improved Transformer structure is used for feature extraction and modal fusion to achieve better performance. Experimental results demonstrate that our Transformer structure not only ensures the stability of the model performance, but also accelerates its convergence, and the MQAT model outperforms the existing state-of-the-art methods.
Lei Liu 0079, Xiangdong Su, Daobin Zhu
ICPR2
2022 QuatSE: Spherical Linear Interpolation of Quaternion for Knowledge Graph Embeddings
Jiang Li 0013, Xiangdong Su, Xinlan Ma, Guanglai Gao
NLPCC (1)2
2021 Fullsubnet: A Full-Band and Sub-Band Fusion Model for Real-Time Single-Channel Speech Enhancement
abstract
This paper proposes a full-band and sub-band fusion model, named as FullSubNet, for single-channel real-time speech enhancement. Full-band and sub-band refer to the models that input full-band and sub-band noisy spectral feature, output full-band and sub-band speech target, respectively. The sub-band model processes each frequency independently. Its input consists of one frequency and several context frequencies. The output is the prediction of the clean speech target for the corresponding frequency. These two types of models have distinct characteristics. The full-band model can capture the global spectral context and the long-distance cross-band dependencies. However, it lacks the ability to modeling signal stationarity and attending the local spectral pattern. The sub-band model is just the opposite. In our proposed FullSubNet, we connect a pure full-band model and a pure sub-band model sequentially and use practical joint training to integrate these two types of models' advantages. We conducted experiments on the DNS challenge (INTERSPEECH 2020) dataset to evaluate the proposed method. Experimental results show that full-band and sub-band information are complementary, and the FullSubNet can effectively integrate them. Besides, the performance of the FullSubNet also exceeds that of the top-ranked methods in the DNS Challenge (INTERSPEECH 2020).
Xiangdong Su, Radu Horaud, Xiaofei Li 0001
ICASSP2
2021 An Efficient Local Word Augment Approach for Mongolian Handwritten Script Recognition
Xiangdong Su, Huali Xu
ICDAR (4)3
2021 MCDALNet: Multi-scale Contextual Dual Attention Learning Network for Medical Image Segmentation
abstract
Medical image segmentation has been widely studied, and many methods have been proposed. Among the existing methods, U-Net and its variants have achieved a promising performance. However, these methods miss certain areas because they only generate fixed-scale receptive fields in each layer of the encoder and cannot establish rich contextual dependencies on the fusion features in the decoder. To solve these problems, this paper proposes a multi-scale contextual dual attention learning network (named MCDALNet) to capture multi-scale information and the dependencies of spatial and channel features. MCDALNet contains two components: an encoder with three multi-scale contextual learning (MCL) modules and a decoder with three dual attention modules. The MCL module extracts multi-scale context information from low-level features through the split-transform-merge-residual architecture. The dual attention module consists of a position attention sub-module and a channel attention submodule, which improve the feature representation and help the medical image segmentation. The position attention submodule captures spatial dependencies by learning similar spatial features, and the channel attention sub-module captures channel dependencies by learning relevant features on the channel maps. Experiment results show that our approach achieves significant improvement in medical image segmentation and outperforms the representative deep learning models on public datasets.
Xiangdong Su, Feilong Bao
IJCNN2
2020 Incorporating Inner-word and Out-word Features for Mongolian Morphological Segmentation
abstract
Mongolian morphological segmentation is regarded as a crucial preprocessing step in many Mongolian related NLP applications and has received extensive attention.Recently, end-to-end segmentation approaches with long short-term memory networks (LSTM) have achieved excellent results.However, the inner-word features among characters in the word and the out-word features from context are not well utilized in the segmentation process.In this paper, we propose a neural network incorporating inner-word and out-word features for Mongolian morphological segmentation.The network consists of two encoders and one decoder.The inner-word encoder uses the self-attention mechanisms to capture the inner-word features of the target word.The out-word encoder employs a two layers BiLSTM network to extract out-word features in the sentence.Then, the decoder adopts a multi-head double attention layer to fuse the inner-word features and out-word features and produces the segmentation result.The evaluation experiment compares the proposed network with the baselines and explores the effectiveness of the sub-modules.
Xiangdong Su, Guanglai Gao, Feilong Bao
COLING2
2020 A Multi-Scaled Receptive Field Learning Approach for Medical Image Segmentation
abstract
Biomedical image segmentation has been widely studied, and lots of methods have been proposed. Among these methods, attention U-Net has achieved a promising performance. However, it has drawbacks of extracting the multi-scaled receptive field features at the high-level feature maps, resulting in the degeneration when dealing with the lesions with apparent scale variations. To solve this problem, this paper integrates an atrous spatial pyramid pooling (ASPP) module in the contracting path of attention U-Net. This module employs multiple dilation rates for the purpose of obtaining several multi-scale receptive fields, which significantly improves the networks' ability to handle both large and small lesions. Evaluation experimental result shows that our approach significantly improves the performance of medical image segmentation and substantially outperforms the representative deep learning models on public datasets.
Xiangdong Su, Feilong Bao
ICASSP2
2020 Masking and Inpainting: A Two-Stage Speech Enhancement Approach for Low SNR and Non-Stationary Noise
abstract
Currently, low signal-to-noise ratio (SNR) and non-stationary noise cause severe performance degradation for most of speech enhancement models. For better speech enhancement at the above scenarios, this paper proposes a two-stage approach that consists of binary masking and spectrogram inpainting. In the binary masking stage, we first obtain binary mask by hardening soft mask and then use it to remove time-frequency points that are dominated by severe noise. In the spectrogram inpainting stage, we use a CNN with partial convolution to perform inpainting on the masked spectrogram from the previous stage. We compared our approach with two powerful baselines, including Wave-U-Net and CRN, on a low SNR dataset containing lots of non-stationary noises. The experimental results show that our approach outperformed the baselines and achieved the state-of-the-art performance.
Xiangdong Su, Shixue Wen, Yiqian Pan, Feilong Bao
ICASSP2
2020 Snr-Based Teachers-Student Technique For Speech Enhancement
abstract
It is very challenging for speech enhancement methods to achieves robust performance under both high signal-to-noise ratio (SNR) and low SNR simultaneously. In this paper, we propose a method that integrates an SNR-based teachers-student technique and time-domain U-Net to deal with this problem. Specifically, this method consists of multiple teacher models and a student model. We first train the teacher models under multiple small-range SNRs that do not coincide with each other so that they can perform speech enhancement well within the specific SNR range. Then, we choose different teacher models to supervise the training of the student model according to the SNR of the training data. Eventually, the student model can perform speech enhancement under both high SNR and low SNR. To evaluate the proposed method, we constructed a dataset with an SNR ranging from -20dB to 20dB based on the public dataset. We experimentally analyzed the effectiveness of the SNR-based teachers-student technique and compared the proposed method with several state-of-the-art methods.
Xiangdong Su, Huali Xu, Guanglai Gao
ICME2
2020 An Edge Information and Mask Shrinking Based Image Inpainting Approach
abstract
In the image inpainting task, the ability to repair both high-frequency and low-frequency information in the missing regions has a substantial influence on the quality of the restored image. However, existing inpainting methods usually fail to consider both high-frequency and low-frequency information simultaneously. To solve this problem, this paper proposes edge information and mask shrinking based image inpainting approach, which consists of two models. The first model is an edge generation model used to generate complete edge information from the damaged image, and the second model is an image completion model used to fix the missing regions with the generated edge information and the valid contents of the damaged image. The mask shrinking strategy is employed in the image completion model to track the areas to be repaired. The proposed approach is evaluated qualitatively and quantitatively on the dataset Places2. The result shows our approach outperforms state-of-the-art methods.
Huali Xu, Xiangdong Su, Guanglai Gao
ICME2
2020 Sub-Band Knowledge Distillation Framework for Speech Enhancement
abstract
In single-channel speech enhancement, methods based on full-band spectral features have been widely studied. However, only a few methods pay attention to non-full-band spectral features. In this paper, we explore a knowledge distillation framework based on sub-band spectral mapping for single-channel speech enhancement. Specifically, we divide the full frequency band into multiple sub-bands and pre-train an elite-level sub-band enhancement model (teacher model) for each sub-band. These teacher models are dedicated to processing their own sub-bands. Next, under the teacher models' guidance, we train a general sub-band enhancement model (student model) that works for all sub-bands. Without increasing the number of model parameters and computational complexity, the student model's performance is further improved. To evaluate our proposed method, we conducted a large number of experiments on an open-source data set. The final experimental results show that the guidance from the elite-level teacher models dramatically improves the student model's performance, which exceeds the full-band model by employing fewer parameters.
Shixue Wen, Xiangdong Su, Guanglai Gao
INTERSPEECH3
2020 Differential Privacy Images Protection Based on Generative Adversarial Network
abstract
In recent years, as image data are widely used in data analysis tasks, the problem of privacy disclosure is becoming more and more serious. However, the privacy protection technology of image data is still immature. In this paper, we propose a privacy protection framework named dp-WGAN for image data. This framework uses differential privacy and generative adversarial network to train a generative model with privacy protection function. Using this generative model, synthetic data with similar characteristics to sensitive data can be obtained, and synthetic data is published instead sensitive data to complete all kinds of data analysis tasks. Through extensive empirical evaluation on benchmark datasets, we demonstrate that dp-WGAN can provide strong privacy protection for sensitive data and produce high-quality synthetic data.
Xuebin Ma, Xiangyu Bai, Xiangdong Su
TrustCom4
2019 Improving Text Image Resolution using a Deep Generative Adversarial Network for Optical Character Recognition
abstract
Optical character recognition (OCR) has been widely studied in previous work. Except for the models used, the recognition accuracy depends most on the resolution of the image to be recognized. To enhance OCR performance, this paper proposes an approach based on a generative adversarial network to improve text image resolution. Our approach uses a perceptual loss function that consists of an adversarial loss, a content loss and an L1 loss. The adversarial loss and the L1 loss are used to ensure the generated super-resolved images are closer to the ground truth high-resolution images. Meanwhile, the content loss is used to ensure the generated super-resolved images and the input low-resolution images have similar features on the basis of perceptual instead of pixel similarity. To evaluate the proposed approach, we compare the recognition accuracies before and after improving the resolution of both English and Chinese text images. The results show that the recognition accuracies on the super-resolved text images obtained with our approach are significantly higher than those on the low-resolution images without processing.
Xiangdong Su, Huali Xu, Ying Kang, Guanglai Gao
ICDAR1
2019 Morphological Knowledge Guided Mongolian Constituent Parsing
Xiangdong Su, Guanglai Gao, Feilong Bao
ICONIP (3)2
2019 Learning an Adversarial Network for Speech Enhancement Under Extremely Low Signal-to-Noise Ratio Condition
Xiangdong Su, Huali Xu, Tongyang Liu, Guanglai Gao, Feilong Bao
ICONIP (1)1
2019 A Natural Scene Text Extraction Approach Based on Generative Adversarial Learning
Huali Xu, Xiangdong Su, Tongyang Liu, Guanglai Gao, Feilong Bao
ICONIP (1)2
2019 UNetGAN: A Robust Speech Enhancement Approach in Time Domain for Extremely Low Signal-to-Noise Ratio Condition
abstract
Speech enhancement at extremely low signal-to-noise ratio (SNR) condition is a very challenging problem and rarely investigated in previous works. This paper proposes a robust speech enhancement approach (UNetGAN) based on U-Net and generative adversarial learning to deal with this problem. This approach consists of a generator network and a discriminator network, which operate directly in the time domain. The generator network adopts a U-Net like structure and employs dilated convolution in the bottleneck of it. We evaluate the performance of the UNetGAN at low SNR conditions (up to -20dB) on the public benchmark. The result demonstrates that it significantly improves the speech quality and substantially outperforms the representative deep learning models, including SEGAN, cGAN fo SE, Bidirectional LSTM using phase-sensitive spectrum approximation cost function (PSA-BLSTM) and Wave-U-Net regarding Short-Time Objective Intelligibility (STOI) and Perceptual evaluation of speech quality (PESQ).
Xiangdong Su, Hui Zhang 0031, Batushiren
INTERSPEECH2
2019 An End-to-End Preprocessor Based on Adversiarial Learning for Mongolian Historical Document OCR
Xiangdong Su, Huali Xu, Yanke Kang, Guanglai Gao, Batushiren
PRICAI (3)1
2018 Mongolian Word Segmentation Based on Three Character Level Seq2Seq Models
Xiangdong Su, Guanglai Gao, Feilong Bao
ICONIP (5)2
2018 Integrating Topic Information into VAE for Text Semantic Similarity
Xiangdong Su, Yujiao Fu
ICONIP (5)1
2017 Using Word Mover's Distance with Spatial Constraints for Measuring Similarity Between Mongolian Word Images
Hongxi Wei, Hui Zhang 0031, Guanglai Gao, Xiangdong Su
ICONIP (4)4
2016 LDA-Based Word Image Representation for Keyword Spotting on Historical Mongolian Documents
Hongxi Wei, Guanglai Gao, Xiangdong Su
ICONIP (4)3
2016 A knowledge-based recognition system for historical Mongolian documents
Xiangdong Su, Guanglai Gao, Hongxi Wei, Feilong Bao
Int. J. Document Anal. Recognit.1
2015 A multiple instances approach to improving keyword spotting on historical Mongolian document images
abstract
For keyword spotting of historical Mongolian document images, when user provides different instance image for the same query keyword, the performance will vary a lot. This paper proposed an approach to solving the above problem. Particularly, the whole procedure of keyword spotting is divided into two stages. The main task of the first stage is to generate multiple ranking lists for a query keyword. And the aim of the second stage is to merge the multiple ranking lists to form a final ranking. In the first stage, the ranking list of one query keyword is firstly returned by traditional image matching and then a number of instances for the query keyword are obtained using pseudo relevant feedback. Next, each instance of the query keyword can return the corresponding ranking list separately. In the second stage, the multiple ranking lists from the multiple instances of the query keyword are combined by the data fusion technique. The final ranking will be taken as the retrieval results of the query keyword. The experimental results show that the proposed approach can significantly improve the performance of keyword spotting for the historical Mongolian document images.
Hongxi Wei, Guanglai Gao, Xiangdong Su
ICDAR3
2015 Enhancing the Mongolian Historical Document Recognition System with Multiple Knowledge-Based Strategies
Xiangdong Su, Guanglai Gao, Hongxi Wei, Feilong Bao
ICONIP (2)1
2015 Mongolian Inflection Suffix Processing in NLP: A Case Study
abstract
Inflection suffix is an important morphological characteristic of Mongolian words, since the suffixes express abundant syntactic and semantic meanings. In order to provide an informative introduction of it, this paper implements a case study on it. Through three Mongolian NLP tasks, we disclose the following information: (1) views of inflection suffix in NLP tasks, (2) Inflection suffix processing ways, (3) Inflection suffix effects on system performance and (4) some suffix related conclusion.
Xiangdong Su, Guanglai Gao, Jing Wu 0011, Feilong Bao
NLPCC1
2011 Classical Mongolian Words Recognition in Historical Document
abstract
There are many classical Mongolian historical documents which are reserved in image form, and as a result it is difficult for us to explore and retrieve them. In this paper, we investigate the peculiarities of classical Mongolian documents and propose an approach to recognize the words in them. We design an algorithm to segment the Mongolian words into several Glyph Units(Glyph Unit abbr. GU). Each GU is consisted of no more than three characters. Then we used a three-stage method to recognize the GUs. At the first stage, all the GUs are classified into nine groups by decision tree using three features of the GUs. At the second stage, the GUs in each group are classified individually by five independent BP Neutral Networks whose inputs are other five feature vectors of the GUs. At the last stage, the five results of each GU group from the above five classifiers are combined to provide the final recognized result. The recognition rate of the Mongolian words in our experiment achieves 71%, indicating that our method is effective.
Guanglai Gao, Xiangdong Su, Hongxi Wei, Yeyun Gong
ICDAR2