Shengxiang Gao

dblp:47/10188 · DBLP profile ↗
← Back
58ranked-venue papers
8as first author
45since 2021 · last 2026
0000-0002-2980-8420ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 43 · 6 first-author · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorComputer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Consensus-Aligned Neuron Efficient Fine-Tuning Large Language Models for Multi-Domain Machine Translation
abstract
Multi-domain machine translation (MDMT) aims to build a unified model capable of translating content across diverse domains. Despite the impressive machine translation capabilities demonstrated by large language models (LLMs), domain adaptation still remains a challenge for LLMs. Existing MDMT methods such as in-context learning and parameter-efficient fine-tuning often suffer from domain shift, parameter interference and limited generalization. In this work, we propose a neuron-efficient fine-tuning framework for MDMT that identifies and updates consensus-aligned neurons within LLMs. These neurons are selected by maximizing the mutual information between neuron behavior and domain features, enabling LLMs to capture both generalizable translation patterns and domain-specific nuances. Our method then fine-tunes LLMs guided by these neurons, effectively mitigating parameter interference and domain-specific overfitting. Comprehensive experiments on three LLMs across ten German-English and Chinese-English translation domains evidence that our method consistently outperforms strong PEFT baselines on both seen and unseen domains, achieving state-of-the-art performance.
Shuting Jiang, Ran Song 0002, Yuxin Huang 0004, Yantuan Xian, Shengxiang Gao, Zhengtao Yu 0001
AAAI6
2026 Improving cross-lingual dependency parsing via LLM-based transferring and self-optimizing synthetic data augmentation
Jianjian Liu, Ying Li 0127, Zhengtao Yu 0001, Yuxin Huang 0004, Shengxiang Gao
Expert Syst. Appl.5
2026 A multi-dimensional instance weighting and dynamically supervised signal selection method for low-resource cross-lingual summarization
Yongbing Zhang 0004, Shengxiang Gao, Yuxin Huang 0004, Kaiwen Tan 0001, Zhengtao Yu 0001
Expert Syst. Appl.3
2026 Chinese and Vietnamese bilingual news topic discovery via association graph clustering
Xiao-Cong Wang, Peili Tang, Yuxin Huang 0004, Shengxiang Gao, Zhengtao Yu 0001
Frontiers Comput. Sci.4
2026 Improving cross-lingual dependency parsing via LLM progressive alignment
Jianghui He, Jianjian Liu, Ying Li 0127, Zhengtao Yu 0001, Yuxin Huang 0004, Shengxiang Gao, Cunli Mao
Pattern Recognit.6
2025 Dynamic Syntactic Feature Filtering and Injecting Networks for Cross-lingual Dependency Parsing
abstract
Pre-trained language models enhanced parsers have achieved outstanding performance in rich-resource languages. Cross-lingual dependency parsing aims to learn useful knowledge from high-resource languages to alleviate data scarcity in low-resource languages. However, effectively reducing the syntactic structure distributional bias and excavating the commonalities among languages is the key challenge for cross-lingual dependency parsing. To address this issue, we propose novel dynamic syntactic feature filtering and injecting networks based on the typical shared-private model that employs one shared and two private encoders to separate source and target language features. Concretely, a Language-Specific Filtering Network (LSFN) on private encoders emphasizes helpful information and ignores the irrelevant or harmful parts of it from the source language. Meanwhile, a Language-Invariant Injecting Network (LIIN) on the shared encoder integrates the advantages of BiLSTM and improved Transformer encoders to transcend language boundaries, thus amplifying syntactic commonalities across languages. Experiments on seven benchmark datasets show that our model achieves an average absolute gain of 1.84 UAS and 3.43 LAS compared with the shared-private model. Comparative experiments validate that both LSFN and LIIN components are complementary in transferring beneficial knowledge from source to target languages. Detailed analyses highlight that our model can effectively capture linguistic commonalities and mitigate the effect of distributional bias, showcasing its robustness and efficacy.
Jianjian Liu, Zhengtao Yu 0001, Ying Li 0127, Yuxin Huang 0004, Shengxiang Gao
AAAI5
2025 SECodec: Structural Entropy-based Compressive Speech Representation Codec for Speech Language Models
abstract
With the rapid advancement of large language models (LLMs), discrete speech representations have become crucial for integrating speech into LLMs. Existing methods for speech representation discretization rely on a predefined codebook size and Euclidean distance-based quantization. However, 1) the size of codebook is a critical parameter that affects both codec performance and downstream task training efficiency. 2) The Euclidean distance-based quantization may lead to audio distortion when the size of the codebook is controlled within a reasonable range. In fact, in the field of information compression, structural information and entropy guidance are crucial, but previous methods have largely overlooked these factors. Therefore, we address the above issues from an information-theoretic perspective, we present SECodec, a novel speech representation codec based on structural entropy (SE) for building speech language models. Specifically, we first model speech as a graph, clustering the speech features nodes within the graph and extracting the corresponding codebook by hierarchically and disentangledly minimizing 2D SE. Then, to address the issue of audio distortion, we propose a new quantization method. This method still adheres to the 2D SE minimization principle, adaptively selecting the most suitable token corresponding to the cluster for each incoming original speech node. Furthermore, we develop a Structural Entropy-based Speech Language Model (SESLM) that leverages SECodec. Experimental results demonstrate that SECodec performs comparably to EnCodec in speech reconstruction, and SESLM surpasses VALL-E in zero-shot text-to-speech tasks.
Linqin Wang, Zhengtao Yu 0001, Shengxiang Gao, Cunli Mao, Yuxin Huang 0004, Ling Dong
AAAI4
2025 Beyond Seen Data: Improving KBQA Generalization Through Schema-Guided Logical Form Generation
abstract
Knowledge Base Question Answering (KBQA) aims to answer user questions in natural language using rich human knowledge stored in large KBs.As current KBQA methods struggle with unseen knowledge base elements and their novel compositions at test time, we introduce SG-KBQA -a novel model that injects schema contexts into entity retrieval and logical form generation to tackle this issue.It exploits information about the semantics and structure of the knowledge base provided by schema contexts to enhance generalizability.We show that SG-KBQA achieves strong generalizability, outperforming state-of-the-art models on three commonly used benchmark datasets across a variety of test settings.Our source code is available at https://github. com/gaosx2000/SG_KBQA.
Shengxiang Gao, Jey Han Lau, Jianzhong Qi 0001
EMNLP1
2025 3R: Enhancing Sentence Representation Learning via Redundant Representation Reduction
abstract
Sentence representation learning (SRL) aims to learn sentence embeddings that conform to the semantic information of sentences.In recent years, fine-tuning methods based on pre-trained models and contrastive learning frameworks have significantly advanced the quality of sentence representations.However, within the semantic space of SRL models, both word embeddings and sentence representations derived from word embeddings exhibit substantial redundant information, which can adversely affect the precision of sentence representations.Existing approaches predominantly optimize training strategies to alleviate the redundancy problem, lacking fine-grained guidance on reducing redundant representations.This paper proposes a novel approach that dynamically identifies and reduces redundant information in a dimensional perspective, training the SRL model to redistribute semantics on different dimensions, and entailing better sentence representations.Extensive experiments across seven semantic text similarity benchmarks demonstrate the effectiveness and generality of the proposed method.A comprehensive analysis of the experimental results is conducted and the code/data will be released.
Longxuan Ma, Yuxin Huang 0004, Shengxiang Gao, Zhengtao Yu 0001
EMNLP4
2025 Universal Low-Resource Speech Synthesis Via Phoneme Fusion Coordinating Low-Rank Decomposition
abstract
Recent advancements in end-to-end text-to-speech models have made significant progress. However, these approaches based on high-resource languages, are inapplicable for low-resource languages, and existing low-resource speech synthesis methods are typically specific to single languages. Consequently, the pursuit of a universal methodology for low-resource speech synthesis emerges as a critical problem that requires immediate attention. Unlike previous works, we propose a novel and universal approach for multiple low-resource languages by leveraging phoneme fusion coordinating low-rank decomposition in a pre-trained multilingual model. Specifically, by establishing a multilingual phoneme dictionary through phoneme fusion and applying parameter freezing and matrix decomposition techniques for fine-tuning, the proposed method was extensively evaluated across four different languages, each utilizing approximately 3 hours of speech data. The experimental results demonstrate that the quality of the synthesized speech surpasses current mainstream methods for low-resource speech synthesis, achieving outcomes comparable to those trained with tens of hours of data. Audio samples are available at https://priankaan.github.io/Demo.github.io.
Yanliang Li, Zhengtao Yu 0001, Linqin Wang, Shengxiang Gao, Ling Dong
ICASSP4
2025 Voice Conversion via Structural Entropy
abstract
Voice conversion (VC) aims to transform a person’s voice to resemble that of another person while maintaining the original linguistic content. Existing methods suffer from the blurring of speech representations and the leakage of prosody information. To address this issue, this study introduces SEVC, a novel neural structural entropy-based VC framework. First, we extract self-supervised representations from both the source and reference speech. The representations of the reference speech are structured as a graph. Using two-dimensional (2D) structural entropy (SE), semantically similar representations are clustered together. For the speaker conversion process, each frame of the source speech is treated as a new node, and the most appropriate semantic cluster for each node is identified using SE. Each frame of the source representation is then replaced by the center representation of its corresponding semantic cluster from the reference speech. Finally, a pretrained vocoder synthesizes audio from the transformed representations. Quantitative and qualitative evaluations demonstrate that SEVC improves speaker similarity while maintaining intelligibility scores comparable to existing methods. Additionally, experimental results indicate that SEVC effectively disentangles paralinguistic features from the reference speech, allowing the model to better focus on the semantic content.
Linqin Wang, Zhengtao Yu 0001, Shengxiang Gao, Cunli Mao, Ling Dong, Yuxin Huang 0004
ICASSP3
2025 Unsupervised Speech Unit Boundary Detection Guided by Linguistic Rule Knowledge
abstract
Unsupervised Automatic Speech Recognition (UASR) aims to align speech signals with corresponding text without relying on annotated data. However, the unknown boundaries of speech units pose challenges for cross-modal alignment, and existing methods often struggle to accurately identify these boundaries and capture structural information. To address this, we propose a reward model based on knowledge of linguistic rules. By analyzing boundary structures at three levels—phonemes, syllables, and words—we design multi-granularity scoring functions, including Sequence Reasonability (SR), Syllable Legality Rules (SLR), Syllable-to-Word Evaluation (SWE), and Syntax Grammar Rules (SGR), to generate reward signals. These signals are utilized in reinforcement learning to train a segmentation model capable of effectively capturing the latent structural patterns in speech. Experimental evaluations on two publicly available English datasets demonstrate that the proposed method effectively improves speech unit boundary detection and cross-modal alignment accuracy, validating the feasibility of integrating linguistic knowledge into the UASR framework.
Sanlong Jiang, Ling Dong, Zhengtao Yu 0001, Shengxiang Gao
IJCNN5
2025 A Mixed-Language Multi-Document News Summarization Dataset and a Graphs-Based Extract-Generate Model
abstract
Shengxiang Gao, Fang Nan, Yongbing Zhang, Yuxin Huang, Kaiwen Tan, Zhengtao Yu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Shengxiang Gao, Yongbing Zhang 0004, Yuxin Huang 0004, Kaiwen Tan 0001, Zhengtao Yu 0001
NAACL (Long Papers)1
2025 SAGEC: Syntax-Aware Grammatical Error Correction with Retrieval-Augmented Generation
Shichang Zhu, Ying Li 0127, Zhengtao Yu 0001, Shengxiang Gao, Cunli Mao, Yuxin Huang 0004
NLPCC (2)5
2025 Fine-Grained Contrastive Learning for End-to-End Vietnamese Text Image Machine Translation
Cunli Mao, Ying Li 0127, Shengxiang Gao, Zhengtao Yu 0001
NLPCC (3)5
2025 Fine-Grained Prosody-Controllable Lao Speech Synthesis Guided by Natural Language
Zhenbei Guo, Cunli Mao, Linqin Wang, Zhengtao Yu 0001, Shengxiang Gao
NLPCC (2)6
2025 A Southeast Asian Language OCR Dataset and Evaluation for Large Multimodal Models
Cunli Mao, Ying Li 0127, Shengxiang Gao, Zhengtao Yu 0001
NLPCC (2)5
2025 Breaking barriers in hotspot mining: a novel approach to reflecting domain characteristics and correlations
Wei Chen 0120, Zhengtao Yu 0001, Shengxiang Gao, Yantuan Xian
Appl. Intell.3
2025 Element relational graph-augmented multi-granularity contextualized encoding for document-level event role filler extraction
Enchang Zhu, Zhengtao Yu 0001, Yuxin Huang 0004, Shengxiang Gao, Yantuan Xian
Frontiers Comput. Sci.4
2025 MKE-PLLM: A benchmark for multilingual knowledge editing on pretrained large language model
Ran Song 0002, Shengxiang Gao, Xiaofei Gao, Cunli Mao, Zhengtao Yu 0001
Neurocomputing2
2025 Salient event detection via hypergraph convolutional network with cross-view self-supervised learning
Enchang Zhu, Zhengtao Yu 0001, Yuxin Huang 0004, Shengxiang Gao, Yantuan Xian
Neurocomputing4
2024 Does Large Language Model Contain Task-Specific Neurons?
abstract
Large language models (LLMs) have demonstrated remarkable capabilities in comprehensively handling various types of natural language processing (NLP) tasks.However, there are significant differences in the knowledge and abilities required for different tasks.Therefore, it is important to understand whether the same LLM processes different tasks in the same way.Are there specific neurons in a LLM for different tasks?Inspired by neuroscience, this paper pioneers the exploration of whether distinct neurons are activated when a LLM handles different tasks.Compared with current research exploring the neurons of language and knowledge, task-specific neurons present a greater challenge due to their abstractness, diversity, and complexity.To address these challenges, this paper proposes a method for task-specific neuron localization based on Causal Gradient Variation with Special Tokens (CGVST).CGVST identifies task-specific neurons by concentrating on the most significant tokens during task processing, thereby eliminating redundant tokens and minimizing interference from non-essential neurons.Compared to traditional neuron localization methods, our approach can more effectively identify task-specific neurons.We conduct experiments across eight different public tasks.Experiments involving the inhibition and amplification of identified neurons demonstrate that our method can accurately locate task-specific neurons.
Ran Song 0002, Shizhu He, Shuting Jiang, Yantuan Xian, Shengxiang Gao, Kang Liu 0001, Zhengtao Yu 0001
EMNLP5
2024 DETS: End-to-End Single-Stage Text-to-Speech Via Hierarchical Diffusion Gan Models
abstract
End-to-end single-stage text-to-speech models have garnered significant attention in recent research, surpassing the performance of conventional two-stage pipeline systems. While prior single-stage models have made substantial advancements, there remains room for improvement in addressing intermittent issues related to unnaturalness and prosody diversity. Unlike previous works, we propose a novel single-stage TTS framework to tackle these problems via hierarchical denoising diffusion generative adversarial networks (GAN) modeling, which parameterizes the denoising model by directly predicting latent variables to improve the naturalness and diversity of the generated speech. Specifically, a conditional GAN is adopted as a non-Gaussian multimodal function to model the denoising distribution, which construct the duration predictor and speech decoder respectively. As such, it allows the TTS model learn the more natural one-to-many relationship in which a text input can be spoken in multiple ways with different pitches and rhythms. In addition, We show that DETS can generate high-fidelity speech waveform with only 1 denoising step. Extensive experimental results on the LJSpeech benchmark dataset demonstrate the favourable performance of the proposed method.
Linqin Wang, Zhengtao Yu 0001, Shengxiang Gao, Cunli Mao, Yuxin Huang 0004
ICASSP3
2024 Decoupling-Enhanced Vietnamese Speech Recognition Accent Adaptation Supervised by Prosodic Domain Information
abstract
The northern and southern accents of the Vietnamese exhibit pronunciation differences in terms of tones and rhythms. Existing Vietnamese pre-trained speech models show a bias in accent representation, leading to a lack of generalization capability for southern accents in Vietnamese speech recognition models and a noticeable decrease in recognition accuracy. In this paper, we propose a decoupled-enhanced adaptive representation strategy guided by prosody and accent domain label information. Through Domain Adversarial Training (DAT), the pre-trained speech model decouples domain-invariant content features. Simultaneously, by incorporating prosodic features, we enhance the pronunciation information of Vietnamese, enabling adaptive representation of the distinctive characteristics of northern and southern Vietnamese accents. Evaluating our proposed method on a dataset of Vietnamese accents, the results demonstrate its superiority over existing approaches, mitigating the performance degradation in Vietnamese speech recognition models caused by accent differences.
Yanwen Fang, Ling Dong, Shengxiang Gao, Hua Lai, Zhengtao Yu 0001
IJCNN4
2024 Knowledge-Guided Reinforcement Learning for Low-Resource Unsupervised Syllabification
abstract
Syllabification is a crucial task in natural language processing, and syllables also play a significant role as modeling units in speech processing. While deep learning methods have shown remarkable progress in syllabification, they face challenges in low-resource languages where ready-made segmentation datasets or rules are lacking. Large language models (LLMs) are mostly unsuitable for these low-resource languages as well. To address these challenges, this paper proposes an unsupervised syllabification approach that incorporates logical reasoning into the reinforcement learning training process, achieving knowledge-guided syllabification. By introducing logical reasoning knowledge and modeling the interaction between the Agent and the knowledge base (KB), the model gains a better understanding of language structures and patterns. The study primarily focuses on low-resource Lao language, with experiments conducted on a publicly available English dataset to validate the effectiveness of the proposed method.
Ling Dong, Zhengtao Yu 0001, Yuxin Huang 0004, Shengxiang Gao
IJCNN6
2024 Integrating Speech Self-Supervised Learning Models and Large Language Models for ASR
Ling Dong, Zhengtao Yu 0001, Yuxin Huang 0004, Shengxiang Gao, Guojiang Zhou
INTERSPEECH5
2024 Semantic-aware entity alignment for low resource language knowledge graph
Junfei Tang, Ran Song 0002, Yuxin Huang 0004, Shengxiang Gao, Zhengtao Yu 0001
Frontiers Comput. Sci.4
2024 Structure-guided feature and cluster contrastive learning for multi-view clustering
Zhenqiu Shu, Bin Li 0006, Cunli Mao, Shengxiang Gao, Zhengtao Yu 0001
Neurocomputing4
2024 A Cross-Lingual Summarization method based on cross-lingual Fact-relationship Graph Generation
Yongbing Zhang 0004, Shengxiang Gao, Yuxin Huang 0004, Kaiwen Tan 0001, Zhengtao Yu 0001
Pattern Recognit.2
2024 WDEA: The Structure and Semantic Fusion With Wasserstein Distance for Low-Resource Language Entity Alignment
abstract
Entity Alignment (EA) aims to identify pairs of entities from two distinct language knowledge graphs (KGs) that represent the same real-world objects. Current EA methods have exhibited impressive performance by leveraging both structural and semantic information. However, these approaches often falter when confronted with EA in low-resource languages. The primary challenge is that low-resource language KGs have sparse graph structures, resulting in difficulty in obtaining accurate entity representations. High-quality entity representation is the key to improving EA performance. Therefore, we propose augmenting entity representations with additional features derived from within the graph. In this paper, we introduce a novel approach: Structure and Semantic Fusion with WD for Low-Resource Language Entity Alignment (WDEA). Our method integrates structural and semantic information using the Wasserstein Distance. Specifically, we design a Wasserstein Graph Convolutional Network (WGCN), a GNN-based model that integrates multi-hop information using a message passing mechanism with WD. Additionally, our method adapts the semantic information from the pre-trained language model in the Wasserstein space to facilitate smooth integration. We also propose the Wasserstein Fusion Encoder (WFE), which effectively combines structural and semantic information in the Wasserstein space. To validate the efficacy of our proposed method, we construct low-resource language EA datasets, encompassing uncommon linguistic varieties with sparser structures compared to mainstream datasets. Experimental results show the superiority of our approach, demonstrating significant performance enhancements in low-resource language EA compared to prevailing baseline models across various information configurations.
Ran Song 0002, Hao Peng 0001, Shengxiang Gao, Zhengtao Yu 0001, Philip S. Yu
IEEE ACM Trans. Audio Speech Lang. Process.4
2024 Zero-Shot Text Normalization via Cross-Lingual Knowledge Distillation
abstract
Text normalization (TN) is a crucial preprocessing step in text-to-speech synthesis, which pertains to the accurate pronunciation of numbers and symbols within the text. Existing neural network-based TN methods have shown significant success in rich-resource languages. However, these methods are data-driven and highly rely on a large number of labeled datasets, which are not practical in zero-resource settings. Rule-based weighted finite-state transducers (WFST) are a common measure for zero-shot TN, but WFST-based TN approaches encounter challenges with ambiguous input, particularly in cases where the normalized form is context-dependent. On the other hand, conventional neural TN methods suffer from unrecoverable errors. In this paper, we propose ZSTN, a novel zero-shot TN framework based on cross-lingual knowledge distillation, which utilizes annotated data to train the teacher model on rich-resource language and unlabelled data to train the student model on zero-resource language. Furthermore, it incorporates expert knowledge from WFST into a knowledge distillation neural network. Concretely, a TN model with WFST pseudo-labels augmentation is trained as a teacher model in the source language. Subsequently, the student model is supervised by soft-labels from the teacher model and WFST pseudo-labels from the target language. By leveraging cross-lingual knowledge distillation, we address contextual ambiguity in the text, while WFST mitigates unrecoverable errors of the neural model. Additionally, ZSTN is adaptable to different zero-resource languages by using the joint loss function for the teacher model and WFST constraints. We also release a zero-shot text normalization dataset in five languages. We compare ZSTN with seven zero-shot TN benchmarks on public datasets in four languages for the teacher model and zero-shot datasets in five languages for the student model. The results demonstrate that the proposed ZSTN excels in performance without the need for labeled data.
Linqin Wang, Zhengtao Yu 0001, Hao Peng 0001, Shengxiang Gao, Cunli Mao, Yuxin Huang 0004, Ling Dong, Philip S. Yu
IEEE ACM Trans. Audio Speech Lang. Process.5
2024 Toward Cross-Lingual Social Event Detection with Hybrid Knowledge Distillation
abstract
Recently published graph neural networks (GNNs) show promising performance at social event detection tasks. However, most studies are oriented toward monolingual data in languages with abundant training samples. This has left the common lesser-spoken languages relatively unexplored. Thus, in this work, we present a GNN-based framework that integrates cross-lingual word embeddings into the process of graph knowledge distillation for detecting events in low-resource language data streams. To achieve this, a novel cross-lingual knowledge distillation framework, called CLKD, exploits prior knowledge learned from similar threads in English to make up for the paucity of annotated data. Specifically, to extract sufficient useful knowledge, we propose a hybrid distillation method that consists of both feature-wise and relation-wise information. To transfer both kinds of knowledge in an effective way, we add a cross-lingual module in the feature-wise distillation to eliminate the language gap and selectively choose beneficial relations in the relation-wise distillation to avoid distraction caused by teachers’ misjudgments. Our proposed CLKD framework also adopts different configurations to suit both offline and online situations. Experiments on real-world datasets show that the framework is highly effective at detection in languages where training samples are scarce.
Jiaqian Ren, Hao Peng 0001, Lei Jiang 0003, Zhifeng Hao 0005, Jia Wu 0001, Shengxiang Gao, Zhengtao Yu 0001
ACM Trans. Knowl. Discov. Data6
2024 Multi-Grained Representation Aggregating Transformer with Gating Cycle for Change Captioning
abstract
Change captioning aims to describe the difference within an image pair in natural language, which combines visual comprehension and language generation. Although significant progress has been achieved, it remains a key challenge of perceiving the object change from different perspectives, especially the severe situation with drastic viewpoint change. In this article, we propose a novel full-attentive network, namely Multi-grained Representation Aggregating Transformer (MURAT), to distinguish the actual change from viewpoint change. Specifically, the Pair Encoder first captures similar semantics between pairwise objects in a multi-level manner, which are regarded as the semantic cues of distinguishing the irrelevant change. Next, a novel Multi-grained Representation Aggregator (MRA) is designed to construct the reliable difference representation by employing both coarse- and fine-grained semantic cues. Finally, the language decoder generates a description of the change based on the output of MRA. Besides, the Gating Cycle Mechanism is introduced to facilitate the semantic consistency between difference representation learning and language generation with a reverse manipulation, so as to bridge the semantic gap between change features and text features. Extensive experiments demonstrate that the proposed MURAT can greatly improve the ability to describe the actual change in the distraction of irrelevant change and achieves state-of-the-art performance on three benchmarks, CLEVR-Change, CLEVR-DC, and Spot-the-Diff.
Shengbin Yue, Yunbin Tu, Liang Li 0003, Shengxiang Gao, Zhengtao Yu 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2023 CasSampling: Exploring Efficient Cascade Graph Learning for Popularity Prediction
Guixiang Cheng, Xin Yan 0005, Shengxiang Gao, Guangyi Xu 0001, Xianghua Miao
ECML/PKDD (3)3
2023 Multi-view clustering via label-embedded regularized NMF with dual-graph constraints
Bin Li 0006, Zhenqiu Shu, Cunli Mao, Shengxiang Gao, Zhengtao Yu 0001
Neurocomputing5
2023 Case Element Joint Extraction Based on Case Field Correlation and Dependency Graph Convolutional Network
Shengxiang Gao, Zhengtao Yu 0001, Chengding Zhao, Peilian Zhao, Peifu Han
Neural Process. Lett.2
2023 Relation-aware attention for video captioning via graph learning
Yunbin Tu, Huafeng Li 0001, Shengxiang Gao, Zhengtao Yu 0001
Pattern Recognit.5
2023 I3N: Intra- and Inter-Representation Interaction Network for Change Captioning
abstract
Change captioning aims to describe the disagreement of image pairs with a linguistic sentence. Compared with single image captioning, change captioning requires not only understanding the fine-grained information of each image, but also determining whether change occurs and further representing the differences of image pairs. Although much progress has been made, it remains a severe challenge of the precise difference representation in the distraction of viewpoint change, especially that of tiny difference. In this paper, we propose a novel Intra- and Inter-representation Interaction Network (I3N) to learn the fine difference representation and be immune to viewpoint change. In the Intra-representation Interaction stage, we design Geometry-Semantic Interaction Refining (GSIR) to explore the positional and semantic interactions of intra-image, which can be a prior knowledge of enduring viewpoint change and reinforce the cognition of semantic change. In the Inter-representation Interaction stage, to endow the model with the capability of pinpointing the latent difference in viewpoint change, Hierarchical Representation Interaction (HRI) models difference from coarse to fine representations through the Semantic Matcher and Change Amplifier module. The proposed approach outperforms the state-of-the-art methods with an encouraging performance on the existing change captioning benchmarks. Our code is available athttps://github.com/yueshengbin/I3N.
Shengbin Yue, Yunbin Tu, Liang Li 0003, Shengxiang Gao, Zhengtao Yu 0001
IEEE Trans. Multim.5
2022 Decoupling Mixture-of-Graphs: Unseen Relational Learning for Knowledge Graph Completion by Fusing Ontology and Textual Experts
abstract
Knowledge Graph Embedding (KGE) has been proposed and successfully utilized to knowledge Graph Completion (KGC). But classic KGE paradigm often fail in unseen relation representations. Previous studies mainly utilize the textual descriptions of relations and its neighbor relations to represent unseen relations. In fact, the semantics of a relation can be expressed by three kinds of graphs: factual graph, ontology graph, textual description graph, and they can complement each other. A more common scenario in the real world is that seen and unseen relations appear at the same time. In this setting, the training set (only seen relations) and testing set (both seen and unseen relations) own different distributions. And the train-test inconsistency problem will make KGE methods easiy overfit on seen relations and under-performance on unseen relations. In this paper, we propose decoupling mixture-of-graph experts (DMoG) for unseen relations learning, which could represent the unseen relations in the factual graph by fusing ontology and textual graphs, and decouple fusing space and reasoning space to alleviate overfitting for seen relations. The experiments on two unseen only public datasets and a mixture dataset verify the effectiveness of the proposed method, which improves the state-of-the-art methods by 6.84% in Hits@10 on average.
Ran Song 0002, Shizhu He, Suncong Zheng, Shengxiang Gao, Kang Liu 0001, Zhengtao Yu 0001, Jun Zhao 0001
COLING4
2022 Discrete asymmetric zero-shot hashing with application to cross-modal retrieval
Zhenqiu Shu, Kailing Yong, Jun Yu 0011, Shengxiang Gao, Cunli Mao, Zhengtao Yu 0001
Neurocomputing4
2022 Enhancing low-resource neural machine translation with syntax-graph guided self-attention
Longchao Gong, Zhengtao Yu 0001, Shengxiang Gao
Knowl. Based Syst.5
2022 I2Transformer: Intra- and Inter-Relation Embedding Transformer for TV Show Captioning
abstract
TV show captioning aims to generate a linguistic sentence based on the video and its associated subtitle. Compared to purely video-based captioning, the subtitle can provide the captioning model with useful semantic clues such as actors’ sentiments and intentions. However, the effective use of subtitle is also very challenging, because it is the pieces of scrappy information and has semantic gap with visual modality. To organize the scrappy information together and yield a powerful omni-representation for all the modalities, an efficient captioning model requires understanding video contents, subtitle semantics, and the relations in between. In this paper, we propose an Intra- and Inter-relation Embedding Transformer (I2Transformer), consisting of an Intra-relation Embedding Block (IAE) and an Inter-relation Embedding Block (IEE) under the framework of a Transformer. First, the IAE captures the intra-relation in each modality via constructing the learnable graphs. Then, IEE learns the cross attention gates, and selects useful information from each modality based on their inter-relations, so as to derive the omni-representation as the input to the Transformer. Experimental results on the public dataset show that the I2Transformer achieves the state-of-the-art performance. We also evaluate the effectiveness of the IAE and IEE on two other relevant tasks of video with text inputs,i.e., TV show retrieval and video-guided machine translation. The encouraging performance further validates that the IAE and IEE blocks have a good generalization ability. The code is available athttps://github.com/tuyunbin/I2Transformer.
Yunbin Tu, Liang Li 0003, Li Su 0003, Shengxiang Gao, Chenggang Yan 0001, Zhengjun Zha, Zhengtao Yu 0001, Qingming Huang
IEEE Trans. Image Process.4
2021 R\^3Net: Relation-embedded Representation Reconstruction Network for Change Captioning
abstract
Change captioning is to use a natural language sentence to describe the fine-grained disagreement between two similar images.Viewpoint change is the most typical distractor in this task, because it changes the scale and location of the objects and overwhelms the representation of real change.In this paper, we propose a Relation-embedded Representation Reconstruction Network (R 3 Net) to explicitly distinguish the real change from the large amount of clutter and irrelevant changes.Specifically, a relation-embedded module is first devised to explore potential changed objects in the large amount of clutter.Then, based on the semantic similarities of corresponding locations in the two images, a representation reconstruction module (RRM) is designed to learn the reconstruction representation and further model the difference representation.Besides, we introduce a syntactic skeleton predictor (SSP) to enhance the semantic interaction between change localization and caption generation.Extensive experiments show that the proposed method achieves the state-of-the-art results on two public datasets 1 .
Yunbin Tu, Liang Li 0003, Chenggang Yan 0001, Shengxiang Gao, Zhengtao Yu 0001
EMNLP (1)4
2021 Enhancing the alignment between target words and corresponding frames for video captioning
Yunbin Tu, Shengxiang Gao, Zhengtao Yu 0001
Pattern Recognit.4
2021 A Neural Joint Model with BERT for Burmese Syllable Segmentation, Word Segmentation, and POS Tagging
abstract
The smallest semantic unit of the Burmese language is called the syllable. In the present study, it is intended to propose the first neural joint learning model for Burmese syllable segmentation, word segmentation, and part-of-speech ( POS ) tagging with the BERT. The proposed model alleviates the error propagation problem of the syllable segmentation. More specifically, it extends the neural joint model for Vietnamese word segmentation, POS tagging, and dependency parsing [28] with the pre-training method of the Burmese character, syllable, and word embedding with BiLSTM-CRF-based neural layers. In order to evaluate the performance of the proposed model, experiments are carried out on Burmese benchmark datasets, and we fine-tune the model of multilingual BERT. Obtained results show that the proposed joint model can result in an excellent performance.
Cunli Mao, Zhibo Man, Zhengtao Yu 0001, Shengxiang Gao, Zhenhan Wang, Hongbin Wang 0002
ACM Trans. Asian Low Resour. Lang. Inf. Process.4
2020 Chinese-Vietnamese bilingual news event summarization based on distributed graph ranking
Shengxiang Gao, Zhengtao Yu 0001
J. Supercomput.1
2019 Multi-view low-rank matrix factorization using multiple manifold regularization
Shengxiang Gao, Zhengtao Yu 0001, Taisong Jin, Ming Yin 0002
Neurocomputing1
2019 Robust ℓ2-Hypergraph and its applications
Taisong Jin, Zhengtao Yu 0001, Yue Gao 0002, Shengxiang Gao, Xiaoshuai Sun, Cuihua Li
Inf. Sci.4
2019 Syntax-Based Chinese-Vietnamese Tree-to-Tree Statistical Machine Translation with Bilingual Features
abstract
Because of the scarcity of bilingual corpora, current Chinese--Vietnamese machine translation is far from satisfactory. Considering the differences between Chinese and Vietnamese, we investigate whether linguistic differences can be used to supervise machine translation and propose a method of syntax-based Chinese--Vietnamese tree-to-tree statistical machine translation with bilingual features. Analyzing the syntax differences between Chinese and Vietnamese, we define some linguistic difference-based rules, such as attributive position, time adverbial position, and locative adverbial position, and create rewards for similar rules. These rewards are integrated into the extraction of tree-to-tree translation rules, and we optimize the pruning of the search space during the decoding phase. The experiments on Chinese--Vietnamese bilingual sentence translation show that the proposed method performs better than several compared methods. Further, the results show that syntactic difference features, with search pruning, can improve the accuracy of machine translation without degrading the efficiency.
Shengxiang Gao, Jihao Huang, Mingya Xue, Zhengtao Yu 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2018 A Method to Review Expert Recommendation Using Topic Relevance and Expert Relationship
abstract
In the process of recommending review experts to projects, in order to effectively make use of the relevance among topics and the relationship among experts, a new method is proposed for review expert recommendation using topic relevance and expert relationship. In this method, firstly, the relevance among topics and the relationships among experts are used to respectively construct the Markov network of topics and the Markov network of experts. Next, the maximum topic clique is extracted from the topic Markov network and the maximum expert clique is extracted from the expert Markov network; then, with the information of the two maximum cliques, the relevance between experts and projects is calculated. After that, according to the descending order of the relevant degree, the candidates are ranked. Finally, the experts, who are the top N to projects, are recommended. The experiments on five domain datasets are made and the results show that the proposed method can improve the effect of review expert recommendation, and the F-value increases by an average of 5% than without considering the relevance among topics and the relationship among experts.
Shengxiang Gao, Zhengtao Yu 0001, Linbin Shi, Xin Yan 0005, Haixia Song
Int. J. Cooperative Inf. Syst.1
2017 Ensemble Learning for Countermeasure of Audio Replay Spoofing Attack in ASVspoof2017
Zhi-Yi Li, MaoBo An, Shengxiang Gao, Faru Zhao
INTERSPEECH5
2017 Detecting micro-blog user interest communities through the integration of explicit user relationship and implicit topic relations
Zhengtao Yu 0001, Shengxiang Gao, Linbin Shi
Sci. China Inf. Sci.4
2017 Combining paper cooperative network and topic model for expert topic analysis and extraction
Shengxiang Gao, Zhengtao Yu 0001
Neurocomputing1
2016 An expert disambiguation method based on attributed graph clustering
abstract
Leveraging expert attributes and their attribute-associated features, we propose an expert disambiguation method based on experts' attributed graph clustering model. In the method, firstly, the attributes and their co-occurrences are identified and extracted. Secondly, based on graph theory, the augmented expert attribute nodes are established and their correlations are connected to form a network of augmented expert attribute graph, which combines experts' attribute consistency and graph' structural consistency. Finally, we establish an entropy model to measure attribute information and structural information, and by minimizing the entropy of super nodes and super edges, we obtain the clustering partition for multiple expert nodes. The experimental results on real-world datasets show that the proposed method significantly outperforms the state-of-art spectral clustering method and the semi-supervised graph clustering method for the accuracy of disambiguation.
Shengxiang Gao, Zhengtao Yu 0001
IJCNN1
2016 Expert list-wise ranking method based on sparse learning
Liren Wang, Zhengtao Yu 0001, Taisong Jin, Xianhui Li, Shengxiang Gao
Neurocomputing5
2015 How to Detect Communities in Large Networks
Yasong Jiang, Shengxiang Gao, Yonghong Yan 0002
ICIC (1)4
2015 Research on the Extraction of Wikipedia-Based Chinese-Khmer Named Entity Equivalents
abstract
Named entity equivalent has been playing a significant role in the processing of cross-language information. However limited by the corpora resource, few in-depth studies have been made on the extraction of the bilingual Chinese-Khmer named entity equivalents. On account of this, this paper proposes a Wikipedia-based approach, utilizes the internal web links in Wikipedia and computes the feature similarity to extract the bilingual Chinese-Khmer named entity equivalents. The experimental result shows that good effect has been achieved when the entity equivalents are acquired through the internal web links in Wikipedia with F value up to 90.67%. Also it shows that the result is quite favorable when the bilingual Chinese-Khmer named entity equivalents are acquired through the computation of feature similarity, turning out that the method proposed in this paper is able to give better effect.
Xin Yan 0005, Zhengtao Yu 0001, Shengxiang Gao
NLPCC4
2014 Online Chinese-Vietnamese Bilingual Topic Detection Based on RCRP Algorithm with Event Elements
Wenxu Long, Jixun Gao, Zhengtao Yu 0001, Shengxiang Gao, Xudong Hong 0001
NLPCC4