VLDB 2026 Research / reviewers in the wild / expert
Yuxin Huang 0004
dblp:35/7536-4
· DBLP profile ↗
43ranked-venue papers
6as first author
42since 2021 · last 2026
0000-0003-1277-6212ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 4 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 12 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 8 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SageLM: A Multi-aspect and Explainable Large Language Model for Speech JudgementabstractSpeech-to-Speech (S2S) Large Language Models (LLMs) are foundational to natural human-computer interaction, enabling end-to-end spoken dialogue systems. However, evaluating these models remains a fundamental challenge. We propose SageLM, an end-to-end, multi-aspect, and explainable speech LLM for comprehensive S2S LLMs evaluation. First, unlike cascaded approaches that disregard acoustic features, SageLM jointly assesses both semantic and acoustic dimensions. Second, it leverages rationale-based supervision to enhance explainability and guide model learning, achieving superior alignment with evaluation outcomes compared to rule-based reinforcement learning methods. Third, we introduce SpeechFeedback, a synthetic preference dataset, and employ a two-stage training paradigm to mitigate the scarcity of speech preference data. Trained on both semantic and acoustic dimensions, SageLM achieves an 82.79% agreement rate with human evaluators, outperforming cascaded and SLM-based baselines by at least 7.42% and 26.20%, respectively. Yuan Ge 0001, Junxiang Zhang, Xiangnan Ma, Chenglong Wang 0002, Kaiyang Ye, Yangfan Du, Linfeng Zhang 0001, Yuxin Huang 0004, Tong Xiao 0001, Zhengtao Yu 0001 |
AAAI | 10 |
| 2026 | Consensus-Aligned Neuron Efficient Fine-Tuning Large Language Models for Multi-Domain Machine TranslationabstractMulti-domain machine translation (MDMT) aims to build a unified model capable of translating content across diverse domains. Despite the impressive machine translation capabilities demonstrated by large language models (LLMs), domain adaptation still remains a challenge for LLMs. Existing MDMT methods such as in-context learning and parameter-efficient fine-tuning often suffer from domain shift, parameter interference and limited generalization. In this work, we propose a neuron-efficient fine-tuning framework for MDMT that identifies and updates consensus-aligned neurons within LLMs. These neurons are selected by maximizing the mutual information between neuron behavior and domain features, enabling LLMs to capture both generalizable translation patterns and domain-specific nuances. Our method then fine-tunes LLMs guided by these neurons, effectively mitigating parameter interference and domain-specific overfitting. Comprehensive experiments on three LLMs across ten German-English and Chinese-English translation domains evidence that our method consistently outperforms strong PEFT baselines on both seen and unseen domains, achieving state-of-the-art performance. Shuting Jiang, Ran Song 0002, Yuxin Huang 0004, Yantuan Xian, Shengxiang Gao, Zhengtao Yu 0001 |
AAAI | 3 |
| 2026 | Improving cross-lingual dependency parsing via LLM-based transferring and self-optimizing synthetic data augmentation
Jianjian Liu, Ying Li 0127, Zhengtao Yu 0001, Yuxin Huang 0004, Shengxiang Gao |
Expert Syst. Appl. | 4 |
| 2026 | A multi-dimensional instance weighting and dynamically supervised signal selection method for low-resource cross-lingual summarization
Yongbing Zhang 0004, Shengxiang Gao, Yuxin Huang 0004, Kaiwen Tan 0001, Zhengtao Yu 0001 |
Expert Syst. Appl. | 4 |
| 2026 | Chinese and Vietnamese bilingual news topic discovery via association graph clustering
Xiao-Cong Wang, Peili Tang, Yuxin Huang 0004, Shengxiang Gao, Zhengtao Yu 0001 |
Frontiers Comput. Sci. | 3 |
| 2026 | Improving cross-lingual dependency parsing via LLM progressive alignment
Jianghui He, Jianjian Liu, Ying Li 0127, Zhengtao Yu 0001, Yuxin Huang 0004, Shengxiang Gao, Cunli Mao |
Pattern Recognit. | 5 |
| 2025 | Dynamic Syntactic Feature Filtering and Injecting Networks for Cross-lingual Dependency ParsingabstractPre-trained language models enhanced parsers have achieved outstanding performance in rich-resource languages. Cross-lingual dependency parsing aims to learn useful knowledge from high-resource languages to alleviate data scarcity in low-resource languages. However, effectively reducing the syntactic structure distributional bias and excavating the commonalities among languages is the key challenge for cross-lingual dependency parsing. To address this issue, we propose novel dynamic syntactic feature filtering and injecting networks based on the typical shared-private model that employs one shared and two private encoders to separate source and target language features. Concretely, a Language-Specific Filtering Network (LSFN) on private encoders emphasizes helpful information and ignores the irrelevant or harmful parts of it from the source language. Meanwhile, a Language-Invariant Injecting Network (LIIN) on the shared encoder integrates the advantages of BiLSTM and improved Transformer encoders to transcend language boundaries, thus amplifying syntactic commonalities across languages. Experiments on seven benchmark datasets show that our model achieves an average absolute gain of 1.84 UAS and 3.43 LAS compared with the shared-private model. Comparative experiments validate that both LSFN and LIIN components are complementary in transferring beneficial knowledge from source to target languages. Detailed analyses highlight that our model can effectively capture linguistic commonalities and mitigate the effect of distributional bias, showcasing its robustness and efficacy. Jianjian Liu, Zhengtao Yu 0001, Ying Li 0127, Yuxin Huang 0004, Shengxiang Gao |
AAAI | 4 |
| 2025 | SECodec: Structural Entropy-based Compressive Speech Representation Codec for Speech Language ModelsabstractWith the rapid advancement of large language models (LLMs), discrete speech representations have become crucial for integrating speech into LLMs. Existing methods for speech representation discretization rely on a predefined codebook size and Euclidean distance-based quantization. However, 1) the size of codebook is a critical parameter that affects both codec performance and downstream task training efficiency. 2) The Euclidean distance-based quantization may lead to audio distortion when the size of the codebook is controlled within a reasonable range. In fact, in the field of information compression, structural information and entropy guidance are crucial, but previous methods have largely overlooked these factors. Therefore, we address the above issues from an information-theoretic perspective, we present SECodec, a novel speech representation codec based on structural entropy (SE) for building speech language models. Specifically, we first model speech as a graph, clustering the speech features nodes within the graph and extracting the corresponding codebook by hierarchically and disentangledly minimizing 2D SE. Then, to address the issue of audio distortion, we propose a new quantization method. This method still adheres to the 2D SE minimization principle, adaptively selecting the most suitable token corresponding to the cluster for each incoming original speech node. Furthermore, we develop a Structural Entropy-based Speech Language Model (SESLM) that leverages SECodec. Experimental results demonstrate that SECodec performs comparably to EnCodec in speech reconstruction, and SESLM surpasses VALL-E in zero-shot text-to-speech tasks. Linqin Wang, Zhengtao Yu 0001, Shengxiang Gao, Cunli Mao, Yuxin Huang 0004, Ling Dong |
AAAI | 6 |
| 2025 | Unsupervised Timeline Summarization via Time-Aware Graph Structural Entropy Minimization
Fan Peng, Yantuan Xian, Hongbin Wang 0002, Yuxin Huang 0004, Ran Song 0002, Zhengtao Yu 0001 |
IEEE Big Data | 4 |
| 2025 | 3R: Enhancing Sentence Representation Learning via Redundant Representation ReductionabstractSentence representation learning (SRL) aims to learn sentence embeddings that conform to the semantic information of sentences.In recent years, fine-tuning methods based on pre-trained models and contrastive learning frameworks have significantly advanced the quality of sentence representations.However, within the semantic space of SRL models, both word embeddings and sentence representations derived from word embeddings exhibit substantial redundant information, which can adversely affect the precision of sentence representations.Existing approaches predominantly optimize training strategies to alleviate the redundancy problem, lacking fine-grained guidance on reducing redundant representations.This paper proposes a novel approach that dynamically identifies and reduces redundant information in a dimensional perspective, training the SRL model to redistribute semantics on different dimensions, and entailing better sentence representations.Extensive experiments across seven semantic text similarity benchmarks demonstrate the effectiveness and generality of the proposed method.A comprehensive analysis of the experimental results is conducted and the code/data will be released. Longxuan Ma, Yuxin Huang 0004, Shengxiang Gao, Zhengtao Yu 0001 |
EMNLP | 3 |
| 2025 | Voice Conversion via Structural EntropyabstractVoice conversion (VC) aims to transform a person’s voice to resemble that of another person while maintaining the original linguistic content. Existing methods suffer from the blurring of speech representations and the leakage of prosody information. To address this issue, this study introduces SEVC, a novel neural structural entropy-based VC framework. First, we extract self-supervised representations from both the source and reference speech. The representations of the reference speech are structured as a graph. Using two-dimensional (2D) structural entropy (SE), semantically similar representations are clustered together. For the speaker conversion process, each frame of the source speech is treated as a new node, and the most appropriate semantic cluster for each node is identified using SE. Each frame of the source representation is then replaced by the center representation of its corresponding semantic cluster from the reference speech. Finally, a pretrained vocoder synthesizes audio from the transformed representations. Quantitative and qualitative evaluations demonstrate that SEVC improves speaker similarity while maintaining intelligibility scores comparable to existing methods. Additionally, experimental results indicate that SEVC effectively disentangles paralinguistic features from the reference speech, allowing the model to better focus on the semantic content. Linqin Wang, Zhengtao Yu 0001, Shengxiang Gao, Cunli Mao, Ling Dong, Yuxin Huang 0004 |
ICASSP | 6 |
| 2025 | VFFG-CL: Virtual Fusion Feature Generation with Curriculum Learning for Missing-Modality Emotion RecognitionabstractMultimodal Emotion Recognition aims to identify emotions by analyzing data from multiple modalities. However, real-world scenarios often involve missing modalities, which significantly degrade recognition performance. Existing methods neglect the interaction between reconstructed and available modalities and fail to address varying degrees of modality missingness. To tackle these challenges, we propose a Virtual Fusion Feature Generation model with Curriculum Learning. Instead of reconstructing individual missing modalities, our method directly generates virtual fusion features by leveraging real fused multimodal features as guidance. Specifically, we design a residual autoencoder to generate virtual fusion features under missing conditions and introduce a KL-based Similarity Alignment mechanism to align these virtual features with the real fused features extracted from complete modalities. Additionally, a curriculum learning strategy is applied to gradually adapt the model to increasingly challenging missing modality scenarios. Experiments on IEMOCAP and MSP-IMPROV demonstrate that our approach outperforms state-of-the-art methods under various missing conditions. Xiaolan Tang, Zhengtao Yu 0001, Yuxin Huang 0004 |
ICME | 4 |
| 2025 | A Mixed-Language Multi-Document News Summarization Dataset and a Graphs-Based Extract-Generate ModelabstractShengxiang Gao, Fang Nan, Yongbing Zhang, Yuxin Huang, Kaiwen Tan, Zhengtao Yu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Shengxiang Gao, Yongbing Zhang 0004, Yuxin Huang 0004, Kaiwen Tan 0001, Zhengtao Yu 0001 |
NAACL (Long Papers) | 4 |
| 2025 | SAGEC: Syntax-Aware Grammatical Error Correction with Retrieval-Augmented Generation
Shichang Zhu, Ying Li 0127, Zhengtao Yu 0001, Shengxiang Gao, Cunli Mao, Yuxin Huang 0004 |
NLPCC (2) | 7 |
| 2025 | Concept-driven representation learning model for knowledge graph completion
Hongguang He, Zhengtao Yu 0001, Yuxin Huang 0004 |
Expert Syst. Appl. | 4 |
| 2025 | Element relational graph-augmented multi-granularity contextualized encoding for document-level event role filler extraction
Enchang Zhu, Zhengtao Yu 0001, Yuxin Huang 0004, Shengxiang Gao, Yantuan Xian |
Frontiers Comput. Sci. | 3 |
| 2025 | Salient event detection via hypergraph convolutional network with cross-view self-supervised learning
Enchang Zhu, Zhengtao Yu 0001, Yuxin Huang 0004, Shengxiang Gao, Yantuan Xian |
Neurocomputing | 3 |
| 2025 | Adversarial dual decision-based model for event-related opinion sentence recognition
Zhengtao Yu 0001, Yuxin Huang 0004 |
Multim. Tools Appl. | 4 |
| 2025 | Enhanced Chinese-Vietnamese Cross-Language Event Detection via Aligned Knowledge Event GraphabstractChinese-Vietnamese cross-language event detection aims to cluster texts that describe the same events in Chinese and Vietnamese into corresponding event clusters. However, because Vietnamese is a low-resource language, directly using multilingual pre-trained models to align event representations in Chinese and Vietnamese texts yields suboptimal results, leading to poor performance in cross-lingual event detection. To address this challenge, we propose a method to enhance cross-lingual event detection between Chinese and Vietnamese by utilizing an aligned knowledge event graph. By leveraging aligned event knowledge, such as personal and place names, to establish correlations between events in different languages, we construct a cross-lingual aligned knowledge event graph. Under the constraint of relational associations, we use contrastive learning to model the similarities and differences between various events, making the representations of the same events in different languages more compact. This approach improves the model’s ability to represent Chinese-Vietnamese cross-lingual event texts and enhances the effectiveness of cross-lingual event detection. Experimental results demonstrate that our method, on multiple multilingual pre-trained models, achieves significant improvements across evaluation metrics such as normalized mutual information, adjusted normalized mutual information, and the adjusted rand coefficient. Yuxin Huang 0004, Yuanlin Yang 0004, Zhengtao Yu 0001, Yantuan Xian |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2025 | Semantic Feature Graph Consistency with Contrastive Cluster Assignments for Multilingual Document ClusteringabstractMultilingual document clustering (MDC) aims to partition multilingual documents into distinct clusters based on topic categories in an unsupervised manner. However, existing MDC methods still suffer from several limitations in practice tasks. Firstly, most of them optimize multiple objectives within the same feature space, thereby leading to the conflict between learning consistently shared semantics and reconstructing inconsistent view-specific information. Secondly, several methods directly integrate information from multilingual documents during the fusion stage, thereby overlooking the semantic differences between different language features. To address the aforementioned problems, we propose a novel multi-view learning method, called Semantic Feature Graph Consistency with Contrastive Cluster Assignments (SFGC 3 A), for MDC. Specifically, the proposed SFGC 3 A method implements consistency objective and reconstruction objective in different feature spaces, thus effectively avoiding conflicts between consistency learning and inconsistency reconstruction. Subsequently, we design the semantic feature graph consistency and semantic label consistency modules to further explore consistent semantic information among multilingual documents, thereby reducing the semantic differences among different language views. Extensive experiments on several multilingual document datasets have shown the effectiveness of the proposed SFGC 3 A method in MDC tasks. The source codes for this work will be released later. Zhenqiu Shu, Yuxin Huang 0004, Hongbin Wang 0002, Zhengtao Yu 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2024 | DETS: End-to-End Single-Stage Text-to-Speech Via Hierarchical Diffusion Gan ModelsabstractEnd-to-end single-stage text-to-speech models have garnered significant attention in recent research, surpassing the performance of conventional two-stage pipeline systems. While prior single-stage models have made substantial advancements, there remains room for improvement in addressing intermittent issues related to unnaturalness and prosody diversity. Unlike previous works, we propose a novel single-stage TTS framework to tackle these problems via hierarchical denoising diffusion generative adversarial networks (GAN) modeling, which parameterizes the denoising model by directly predicting latent variables to improve the naturalness and diversity of the generated speech. Specifically, a conditional GAN is adopted as a non-Gaussian multimodal function to model the denoising distribution, which construct the duration predictor and speech decoder respectively. As such, it allows the TTS model learn the more natural one-to-many relationship in which a text input can be spoken in multiple ways with different pitches and rhythms. In addition, We show that DETS can generate high-fidelity speech waveform with only 1 denoising step. Extensive experimental results on the LJSpeech benchmark dataset demonstrate the favourable performance of the proposed method. Linqin Wang, Zhengtao Yu 0001, Shengxiang Gao, Cunli Mao, Yuxin Huang 0004 |
ICASSP | 5 |
| 2024 | Knowledge-Guided Reinforcement Learning for Low-Resource Unsupervised SyllabificationabstractSyllabification is a crucial task in natural language processing, and syllables also play a significant role as modeling units in speech processing. While deep learning methods have shown remarkable progress in syllabification, they face challenges in low-resource languages where ready-made segmentation datasets or rules are lacking. Large language models (LLMs) are mostly unsuitable for these low-resource languages as well. To address these challenges, this paper proposes an unsupervised syllabification approach that incorporates logical reasoning into the reinforcement learning training process, achieving knowledge-guided syllabification. By introducing logical reasoning knowledge and modeling the interaction between the Agent and the knowledge base (KB), the model gains a better understanding of language structures and patterns. The study primarily focuses on low-resource Lao language, with experiments conducted on a publicly available English dataset to validate the effectiveness of the proposed method. Ling Dong, Zhengtao Yu 0001, Yuxin Huang 0004, Shengxiang Gao |
IJCNN | 4 |
| 2024 | Integrating Speech Self-Supervised Learning Models and Large Language Models for ASR
Ling Dong, Zhengtao Yu 0001, Yuxin Huang 0004, Shengxiang Gao, Guojiang Zhou |
INTERSPEECH | 4 |
| 2024 | DGSRN: Noise-Robust Speech Recognition Method with Dual-Path Gated Spectral Refinement Network
Shangbin Mo, Ling Dong, Zhengtao Yu 0001, Yuxin Huang 0004 |
INTERSPEECH | 6 |
| 2024 | Predictive Score-Guided Mixup for Medical Text Classification
Yuhong Pang, Yantuan Xian, Yuxin Huang 0004 |
ISBRA (1) | 4 |
| 2024 | Enhancing Cross-Lingual Topic-Essay Generation with Knowledge and Topic Consistency Constraints
Huailing Gu, Yuxin Huang 0004, Zhengtao Yu 0001, Cunli Mao |
NLPCC (4) | 2 |
| 2024 | View-interactive attention information alignment-guided fusion for incomplete multi-view clustering
Zhenqiu Shu, Yunwei Luo, Yuxin Huang 0004, Cunli Mao, Zhengtao Yu 0001 |
Expert Syst. Appl. | 3 |
| 2024 | Semantic-aware entity alignment for low resource language knowledge graph
Junfei Tang, Ran Song 0002, Yuxin Huang 0004, Shengxiang Gao, Zhengtao Yu 0001 |
Frontiers Comput. Sci. | 3 |
| 2024 | Enhancing low-resource cross-lingual summarization from noisy data with fine-grained reinforcement learningabstractCross-lingual summarization (CLS) is the task of generating a summary in a target language from a document in a source language. Recently, end-to-end CLS models have achieved impressive results using large-scale, high-quality datasets typically constructed by translating monolingual summary corpora into CLS corpora. However, due to the limited performance of low-resource language translation models, translation noise can seriously degrade the performance of these models. In this paper, we propose a fine-grained reinforcement learning approach to address low-resource CLS based on noisy data. We introduce the source language summary as a gold signal to alleviate the impact of the translated noisy target summary. Specifically, we design a reinforcement reward by calculating the word correlation and word missing degree between the source language summary and the generated target language summary, and combine it with cross-entropy loss to optimize the CLS model. To validate the performance of our proposed model, we construct Chinese-Vietnamese and Vietnamese-Chinese CLS datasets. Experimental results show that our proposed model outperforms the baselines in terms of both the ROUGE score and BERTScore. Yuxin Huang 0004, Huailing Gu, Zhengtao Yu 0001, Yumeng Gao, Tong Pan, Jialong Xu |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2024 | A Cross-Lingual Summarization method based on cross-lingual Fact-relationship Graph Generation
Yongbing Zhang 0004, Shengxiang Gao, Yuxin Huang 0004, Kaiwen Tan 0001, Zhengtao Yu 0001 |
Pattern Recognit. | 3 |
| 2024 | Augmenting Low-Resource Cross-Lingual Summarization with Progression-Grounded Training and PromptingabstractCross-lingual summarization (CLS) , generating summaries in one language from source documents in another language, offers invaluable assistance in enabling global access to information for people worldwide. State-of-the-art neural summarization models typically train or fine-tune language models on large-scale corpora. However, this is difficult to achieve in realistic low-resource scenarios due to the lack of domain-specific annotated data. In this article, we present a novel cross-lingual summarization model that utilizes progressive training with mBART and employs reinforcement learning to optimize discrete prompts, which addresses low-resource cross-lingual summarization through a two-pronged approach. During training, we introduce a progressive approach based on mBART, which allows the pre-trained model to gradually acquire the ability to compress information, develop cross-lingual capabilities, and ultimately adapt to specific summarization tasks. During downstream summarization, we employ a discrete-prompts joint pre-trained model based on reinforcement learning optimization to achieve low-resource cross-lingual summarization. Experimental results on four cross-lingual summarization datasets demonstrate state-of-the-art performance and superiority compared to six baselines in low-resource scenarios. Jiushun Ma, Yuxin Huang 0004, Linqin Wang, Hao Peng 0001, Zhengtao Yu 0001, Philip S. Yu |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2024 | Zero-Shot Text Normalization via Cross-Lingual Knowledge DistillationabstractText normalization (TN) is a crucial preprocessing step in text-to-speech synthesis, which pertains to the accurate pronunciation of numbers and symbols within the text. Existing neural network-based TN methods have shown significant success in rich-resource languages. However, these methods are data-driven and highly rely on a large number of labeled datasets, which are not practical in zero-resource settings. Rule-based weighted finite-state transducers (WFST) are a common measure for zero-shot TN, but WFST-based TN approaches encounter challenges with ambiguous input, particularly in cases where the normalized form is context-dependent. On the other hand, conventional neural TN methods suffer from unrecoverable errors. In this paper, we propose ZSTN, a novel zero-shot TN framework based on cross-lingual knowledge distillation, which utilizes annotated data to train the teacher model on rich-resource language and unlabelled data to train the student model on zero-resource language. Furthermore, it incorporates expert knowledge from WFST into a knowledge distillation neural network. Concretely, a TN model with WFST pseudo-labels augmentation is trained as a teacher model in the source language. Subsequently, the student model is supervised by soft-labels from the teacher model and WFST pseudo-labels from the target language. By leveraging cross-lingual knowledge distillation, we address contextual ambiguity in the text, while WFST mitigates unrecoverable errors of the neural model. Additionally, ZSTN is adaptable to different zero-resource languages by using the joint loss function for the teacher model and WFST constraints. We also release a zero-shot text normalization dataset in five languages. We compare ZSTN with seven zero-shot TN benchmarks on public datasets in four languages for the teacher model and zero-shot datasets in five languages for the student model. The results demonstrate that the proposed ZSTN excels in performance without the need for labeled data. Linqin Wang, Zhengtao Yu 0001, Hao Peng 0001, Shengxiang Gao, Cunli Mao, Yuxin Huang 0004, Ling Dong, Philip S. Yu |
IEEE ACM Trans. Audio Speech Lang. Process. | 7 |
| 2024 | Generation and Recombination for Multifocus Image Fusion With Free Number of InputsabstractMultifocus image fusion is an effective method to overcome the limitations of optical lenses. The fused results can be obtained from some existing methods by generating decision maps. However, such methods assume that the focused areas of the two source images are complementary, making it impossible to achieve the simultaneous fusion of multiple images. Additionally, existing methods ignore the impact of hard pixels on the fusion performance, limiting the visual quality improvement of fusion images. To address these issues, a combined generation and recombination model called GRFusion is proposed. In GRFusion, the focus property detection of each source image can be independently implemented, enabling the simultaneous fusion of multiple source images and avoiding information loss caused by alternating fusion. It renders the GRFusion free from the limitation of the number of input images. Furthermore, GRFusion investigates the detection of hard pixels with ambiguous focus properties by analyzing the inconsistencies among the detection results of the focus areas in the source images. This allows the hard pixels to be distinguished from the source images. Besides, a multidirectional gradient embedding method is proposed for generating full-focus images. Subsequently, a hard-pixel-guided recombination mechanism for constructing the fused result is devised to integrate the complementary advantages of feature reconstruction-based and focused pixel recombination-based methods. Extensive experimental results demonstrate the effectiveness and superiority of the proposed method. The source code of the proposed method is available at: https://github.com/lhf12278/GRFusion. Huafeng Li 0001, Yuxin Huang 0004, Zhengtao Yu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Chinese-Vietnamese Cross-Lingual Event Causality Identification Based on Syntactic Graph Convolution
Enchang Zhu, Zhengtao Yu 0001, Yuxin Huang 0004, Yantuan Xian, Shuaishuai Zhou |
PRCV (7) | 3 |
| 2023 | Cross-lingual Sentence Embedding for Low-resource Chinese-Vietnamese Based on Contrastive LearningabstractCross-lingual sentence embedding’s goal is mapping sentences with similar semantics but in different languages close together and dissimilar sentences farther apart in the representation space. It is the basis of many downstream tasks such as cross-lingual document matching and cross-lingual summary extraction. At present, the works of cross-lingual sentence embedding tasks mainly focus on languages with large-scale corpus. But low-resource languages such as Chinese-Vietnamese are short of sentence-level parallel corpora and clear cross-lingual monitoring signals, and these works on low-resource languages have poor performances. Therefore, we propose a cross-lingual sentence embedding method based on contrastive learning and effectively fine-tune powerful pretraining mode by constructing sentence-level positive and negative samples to avoid the catastrophic forgetting problem of the traditional fine-tuning pre-trained model based only on small-scale aligned positive samples. First, we construct positive and negative examples by taking parallel Chinese Vietnamese sentences as positive examples and non-parallel sentences as negative examples. Second, we construct a siamese network to get contrastive loss by inputting positive and negative samples and fine-tuning our model. The experimental results show that our method can effectively improve the semantic alignment accuracy of cross-lingual sentence embedding in Chinese and Vietnamese contexts. Yuxin Huang 0004, Zhaoyuan Wu, Enchang Zhu, Zhengtao Yu 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2022 | Abstractive document summarization via multi-template decoding
Yuxin Huang 0004, Zhengtao Yu 0001, Yantuan Xian |
Appl. Intell. | 1 |
| 2022 | Exploiting comments information to improve legal public opinion news abstractive summarization
Yuxin Huang 0004, Zhengtao Yu 0001 |
Frontiers Comput. Sci. | 1 |
| 2022 | Linguistic feature template integration for Chinese-Vietnamese neural machine translation
Yantuan Xian, Zhengtao Yu 0001, Yuxin Huang 0004 |
Frontiers Comput. Sci. | 4 |
| 2022 | Event Graph Neural Network for Opinion Target Classification of Microblog CommentsabstractOpinion target classification of microblog comments is one of the most important tasks for public opinion analysis about an event. Due to the high cost of manual labeling, opinion target classification is generally considered as a weak-supervised task. This article attempts to address the opinion target classification of microblog comments through an event graph convolution network (EventGCN) in a weak-supervised manner. Specifically, we take microblog contents and comments as document nodes, and construct an event graph with three typical relationships of event microblogs, including the co-occurrence relationship of event keywords extracted from microblogs, the reply relationship of comments, and the document similarity. Finally, under the supervision of a small number of labels, both word features and comment features can be represented well to complete the classification. The experimental results on two event microblog datasets show that EventGCN can significantly improve the classification performance compared with other baseline models. Zhengtao Yu 0001, Yuxin Huang 0004, Yantuan Xian |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2022 | Improving Chinese-Vietnamese Neural Machine Translation with Linguistic DifferencesabstractWe present a simple, efficient data augmentation approach for boosting Chinese-Vietnamese neural machine translation performance by leveraging the linguistic difference between the two languages. We first define the formalized representation of modifier symmetry, which is one of the most representative linguistic differences between Chinese and Vietnamese. We then propose and test two data augmentation strategies for leveraging the linguistic difference, which can be integrated naturally with different translation models. Results indicate that both strategies can introduce linguistic rules to boost translation accuracy. Tests on Chinese-Vietnamese benchmarks show significant accuracy improvements. To facilitate studies in this domain, we also release an open-source toolkit 1 with flexible implementation for Chinese-Vietnamese linguistic difference tagging. Zhengtao Yu 0001, Yantuan Xian, Yuxin Huang 0004 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2021 | Hybrid node-based tensor graph convolutional network for aspect-category sentiment classification of microblog commentsabstractSummary Aspect‐category sentiment classification of microblog comments aims to identify the sentiment polarity of different opinion aspects in microblog comments, which is meaningful for the analysis of public opinion. At present, most of aspect‐category sentiment classification methods need much annotation data, and regard comments as independent samples, without using of the relationship between comments. This article proposes an aspect‐category sentiment classification method based on tensor graph convolutional networks. First, the combination of a comment and its aspect category is regarded as a hybrid node, and the original representation of a hybrid node is encoded by the Bert model. Second, sentiment graph and semantic graph are constructed according to the semantic similarity and sentimental relevance between hybrid nodes, and they are stacked into a tensor. Then two convolution operations, including intra‐graph convolution and inter‐graph convolution, are performed for each layer of graph tensor. In this way, hybrid nodes can learn and merge the heterogeneous information of different graphs. Finally, under the supervision of few labeled comments, the sentiment classification can be completed based on the features of the hybrid nodes. Experimental results on two microblog datasets show that the proposed model can significantly improve the performance of sentiment classification compared with other baseline models. Yantuan Xian, Yuxin Huang 0004, Zhengtao Yu 0001 |
Concurr. Comput. Pract. Exp. | 4 |
| 2021 | Element graph-augmented abstractive summarization for legal public opinion news with graph transformer
Yuxin Huang 0004, Zhengtao Yu 0001, Yantuan Xian |
Neurocomputing | 1 |
| 2020 | Efficient Low-Resource Neural Machine Translation with Reread and Feedback MechanismabstractHow to utilize information sufficiently is a key problem in neural machine translation (NMT), which is effectively improved in rich-resource NMT by leveraging large-scale bilingual sentence pairs. However, for low-resource NMT, lack of bilingual sentence pairs results in poor translation performance; therefore, taking full advantage of global information in the encoding-decoding process is effective for low-resource NMT. In this article, we propose a novel reread-feedback NMT architecture (RFNMT) for using global information. Our architecture builds upon the improved sequence-to-sequence neural network and consists of a double-deck attention-based encoder-decoder framework. In our proposed architecture, the information generated by the first-pass encoding and decoding process flows to the second-pass encoding process for more sufficient parameters initialization and information use. Specifically, we first propose a “reread” mechanism to transfer the outputs of the first-pass encoder to the second-pass encoder, and then the output is used for the initialization of the second-pass encoder. Second, we propose a “feedback” mechanism that transfers the first-pass decoder’s outputs to a second-pass encoder via an important weight model and an improved gated recurrent unit (GRU). Experiments on multiple datasets show that our approach achieves significant improvements over state-of-the-art NMT systems, especially in low-resource settings. Zhengtao Yu 0001, Yuxin Huang 0004, Yonghua Wen |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |