VLDB 2026 Research / reviewers in the wild / expert
Xiaodong Shi
dblp:73/5055 · also Xiao-Dong Shi
· DBLP profile ↗
76ranked-venue papers
4as first author
46since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 51 · 2 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 13 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3Security and privacy · 2Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient and Adaptive Simultaneous Speech Translation with Fully Unidirectional ArchitectureabstractSimultaneous speech translation (SimulST) produces translations incrementally while processing partial speech input. Although large language models (LLMs) have shown strong capabilities in offline translation tasks, applying them to SimulST poses notable challenges. Existing LLM-based SimulST approaches either incur significant computational overhead due to repeated encoding of bidirectional speech encoder, or they depend on a fixed read/write policy, limiting the efficiency and performance. In this work, we introduce Efficient and Adaptive Simultaneous Speech Translation (EASiST) with fully unidirectional architecture, including both speech encoder and LLM. EASiST includes a multi-latency data curation strategy to generate semantically aligned SimulST training samples and redefines SimulST as an interleaved generation task with explicit read/write tokens. To facilitate adaptive inference, we incorporate a lightweight policy head that dynamically predicts read/write actions. Additionally, we employ a multi-stage training strategy to align speech-text modalities and optimize both translation and policy behavior. Experiments on both in-domain (MuST-C) and out-of-domain (Europarl-ST) En-De and En-Es datasets demonstrate that EASiST offers superior latency-quality trade-offs compared to several strong baselines. Biao Fu, Donglei Yu, Minpeng Liao, Chengxi Li 0014, Xinjie Chen, Yidong Chen 0001, Kai Fan 0002, Xiaodong Shi |
AAAI | 8 |
| 2026 | Translation with Thought: Difficulty-Adaptive Reasoning via Reinforcement Learning for Multi-Domain Machine TranslationabstractMulti-domain machine translation (MDMT) poses a unique challenge due to varying levels of linguistic complexity across domains. Inspired by human translators’ ability to adapt reasoning effort based on difficulty, we propose TwT (Translation with Thought), a resource-rational framework that learns to modulate inference between intuitive and deliberate reasoning. TwT is trained in two stages: (1) supervised fine-tuning on difficulty-aware long chain-of-though traces distilled from DeepSeek-R1 and rewritten by GPT-4o to reflect human-like reasoning economy, and (2) reinforcement learning with a hybrid reward to optimize translation quality and reasoning efficiency. Evaluated on 15 benchmarks spanning in-domain and out-of-domain settings, as well as 3 seen and 59 unseen languages, with ablations across three backbone models, TwT-7B and TwT-14B outperform much larger SOTA reasoning models in translation quality, while reducing token usage by 32–60%. These results confirm that aligning translation behavior with cognitive principles enables robust generalization, high translation quality, and efficient reasoning in MDMT. Yongshi Ye, Biao Fu, Chongxuan Huang, Yidong Chen 0001, Xiaodong Shi |
ACL (1) | 5 |
| 2026 | Dual-Loop Multi-agent Framework and Policy Optimization for Numbered Musical Notation Digitization
Peng Bai 0006, Xiaodong Shi |
ICIC (16) | 3 |
| 2026 | Towards Chinese Gezi Opera Music Source Separation: A Dataset and Model Evaluation
Wujin Sun, Peng Bai 0006, Yue Zhou 0012, Zhicong Wu, Xiaodong Shi |
ICIC (14) | 5 |
| 2026 | UniVoice: a unified framework for text-to-speech, singing voice synthesis, and opera singing synthesis
Yue Zhou 0012, Peng Bai 0006, Xiaodong Shi |
Appl. Intell. | 3 |
| 2026 | Enhancing end-to-end speech translation via multi-stage knowledge distillation
Yue Zhou 0012, Yanyan Feng, Xiaodong Shi |
Neural Networks | 4 |
| 2025 | From Neurons to Semantics: Evaluating Cross-Linguistic Alignment Capabilities of Large Language Models via Neurons AlignmentabstractLarge language models (LLMs) have demonstrated remarkable multilingual capabilities, however, how to evaluate cross-lingual alignment remains underexplored.Existing alignment benchmarks primarily focus on sentence embeddings, but prior research has shown that neural models tend to induce a nonsmooth representation space, which impact of semantic alignment evaluation on low-resource languages.Inspired by neuroscientific findings that similar information activates overlapping neuronal regions, we propose a novel Neuron State-Based Cross-Lingual Alignment (NeuronXA) to assess the cross-lingual a lignment capabilities of LLMs, which offers a more semantically grounded approach to assess cross-lingual alignment.We evaluate NeuronXA on several prominent multilingual LLMs (LLaMA, Qwen, Mistral, GLM, and OLMo) across two transfer tasks and three multilingual benchmarks.The results demonstrate that with only 100 parallel sentence pairs, Neu-ronXA achieves a Pearson correlation of 0.9556 with downstream tasks performance and 0.8524 with transferability.These findings demonstrate NeuronXA's effectiveness in assessing both cross-lingual alignment and transferability, even with a small dataset.This highlights its potential to advance cross-lingual alignment research and to improve the semantic understanding of multilingual LLMs. Chongxuan Huang, Yongshi Ye, Biao Fu, Qifeng Su, Xiaodong Shi |
ACL (1) | 5 |
| 2025 | LLM×MapReduce: Simplified Long-Sequence Processing using Large Language ModelsabstractZihan Zhou, Chong Li, Xinyi Chen, Shuo Wang, Yu Chao, Zhili Li, Haoyu Wang, Qi Shi, Zhixing Tan, Xu Han, Xiaodong Shi, Zhiyuan Liu, Maosong Sun. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Shuo Wang 0013, Yu Chao, Zhili Li, Qi Shi 0002, Zhixing Tan, Xu Han 0007, Xiaodong Shi, Zhiyuan Liu 0001, Maosong Sun 0001 |
ACL (1) | 11 |
| 2025 | Representation Purification for End-to-End Speech TranslationabstractSpeech-to-text translation (ST) is a cross-modal task that involves converting spoken language into text in a different language. Previous research primarily focused on enhancing speech translation by facilitating knowledge transfer from machine translation, exploring various methods to bridge the gap between speech and text modalities. Despite substantial progress made, factors in speech that are not relevant to translation content, such as timbre and rhythm, often limit the efficiency of knowledge transfer. In this paper, we conceptualize speech representation as a combination of content-agnostic and content-relevant factors. We examine the impact of content-agnostic factors on translation performance through preliminary experiments and observe a significant performance deterioration when content-agnostic perturbations are introduced to speech signals. To address this issue, we propose a Speech Representation Purification with Supervision Enhancement (SRPSE) framework, which excludes the content-agnostic components within speech representations to mitigate their negative impact on ST. Experiments on MuST-C and CoVoST-2 datasets demonstrate that SRPSE significantly improves translation performance across all translation directions in three settings and achieves preeminent performance under a transcript-free setting. Yue Zhou 0012, Yidong Chen 0001, Xiaodong Shi |
COLING | 5 |
| 2025 | Generating Gezi Opera Scores with a Large Language Model and a High-Quality DatasetabstractDespite significant progress in music generation technology recently, covering various unique styles and genres, the generation of Chinese opera scores still urgently requires more attention, primarily due to the lack of a high-quality and lyric-melody alignment opera score dataset. In this study, we collect pictures of the jianpu of the Chinese Gezi opera and manually construct a high-quality standard data set of Chinese Gezi opera scores, which can be read by music notation software. This dataset includes 142 lyric-melody alignment Gezi opera scores, and it is intended to facilitate opera score generation and analysis. In the lyric-to-melody generation task, we design a triple data format (lyric, notes, and note lengths) specifically for Gezi opera scores. This data format aligns lyric and melody aims to enhance the model’s understanding and generation capabilities. Experiments show that our designed data format exhibits superior performance in the Gezi opera lyric-to-melody generation task, significantly outperforming strong baseline methods. Specifically, it surpasses the state-of-the-art model by 5. 04% in pitch distribution similarity, improves duration distribution similarity by 8. 62%, and achieves a lower melody distance of 0.49. Download information can be found at https://github.com/ZhenLEI96/GeziOperaDataset. Ke Gu 0003, Peng Bai 0006, Xiaodong Shi |
ICASSP | 4 |
| 2025 | A Cross-Font Image Retrieval Network for Recognizing Undeciphered Oracle Bone Inscriptions
Zhicong Wu, Qifeng Su, Ke Gu 0003, Xiaodong Shi |
ICIC (5) | 4 |
| 2025 | End-to-End Lyric-to-Melody Generation via Chord Integration and Bar-Level ModelingabstractThe lyric-to-melody generation task can support composers in their creative process, helping to improve their work efficiency. However, existing studies suffer from inadequate utilization of chord and a lack of bar-level information in melodies. To address the above issues, we propose a Chord-integrated and Bar-level modeled end-to-end encoder-decoder Lyric-to-Melody (CBL2M) generation model, effectively leveraging chord and bar information based on existing quantitative data. To enhance the representation capabilities of the encoder, we extract chord from the melody and encode the lyric and chord separately. In the decoder, we compute lyric-melody and chord-melody cross-attention. Furthermore, to accurately capture the regularity in bar-level duration within the melody, we design two specialized musical token sequences: the bar token sequence and the bar intra-position token sequence. We conduct extensive experiments on Chinese and English datasets, with both subjective and objective evaluation metrics showing that CBL2M generates accurate high-quality melodies that outperform existing strong models. Ke Gu 0003, Peng Bai 0006, Yue Zhou 0012, Zhicong Wu, Xiaodong Shi |
ICME | 6 |
| 2025 | Enhancing GPT-Based Input Method with Domain-Adaptive Pinyin Encoder(s)
Wenyao Peng, Xiaodong Shi |
NLPCC (2) | 2 |
| 2025 | LLM-Enhanced Translation for Low-Resource Languages: Cross-Lingual Alignment and Multi-domain Adaptation
Qifeng Su, Zhicong Wu, Xiaodong Shi |
NLPCC (3) | 3 |
| 2025 | CMiNER: Named entity recognition on imperfectly annotated data via confidence and meta weight adaptation
Shangfei Wei, Zhaohong Lai, Xiaodong Shi |
Expert Syst. Appl. | 3 |
| 2025 | Towards Simultaneous Sign Language Production: A Future-Context-Aware ApproachabstractSign Language Production (SLP) has achieved promising progress in offline settings, where full input text is available before generation. However, such methods are unsuitable for real-time applications requiring low latency. In this work, we introduce Simultaneous Sign Language Production (SimulSLP), a new task that generates sign pose sequences incrementally from streaming text input. We first formalize the SimulSLP task and adapt the Average Token Delay metric to quantify latency. Then, we benchmark this task using three strong baselines from offline SLP—an end-to-end system and two cascaded pipelines with neural and dictionary-based Gloss-to-Pose modules—under a wait-k policy. However, all baselines suffer from a mismatch between full-sequence training and partial-input inference. To mitigate this, we propose a Future-Context-Aware Inference (FCAI) strategy. FCAI enhances partial input representations by predicting a small number of future tokens using a large language model. Before decoding, speculative features from the predicted tokens are discarded to ensure alignment with the observed input. Experiments on PHOENIX2014T show that FCAI significantly improves the quality-latency trade-off, especially in low-latency settings, offering a promising step toward SimulSLP. Biao Fu, Xiaodong Shi, Yidong Chen 0001 |
IEEE Signal Process. Lett. | 3 |
| 2025 | Boosting Context-Aware Speech Translation With Large Language ModelsabstractWith the rise of large language models (LLMs), numerous studies have incorporated LLMs into the speech domain, yielding substantial improvements in sentence-level speech-to-text translation (ST) performance. However, when faced with complex context-aware speech translation tasks, the performance of LLMs often declines, sometimes even underperforming compared to existing context-aware ST models. This paper explores how to enhance the performance of LLMs in context-aware speech translation. Specifically, we optimize the LLM's context-aware speech translation capabilities through instruction tuning, while introducing multiple related tasks, such as automatic speech recognition and machine translation, to improve translation quality in complex contexts. Additionally, we propose a Task Distribution Regularization method to promote consistency across tasks and strengthen the model's understanding of context. We also design a multi-task hybrid learning strategy to ensure efficient fine-tuning of the LLM across tasks. Experimental results demonstrate that our approach achieves excellent performance on the MuST-C and CoVoST2 benchmarks, significantly improving context-aware ST performance. Yue Zhou 0012, Xiaodong Shi |
IEEE Signal Process. Lett. | 4 |
| 2024 | FT-GAN: Fine-Grained Tune Modeling for Chinese Opera SynthesisabstractAlthough singing voice synthesis (SVS) has made significant progress recently, with its unique styles and various genres, Chinese opera synthesis requires greater attention but is rarely studied for lack of training data and high expressiveness. In this work, we build a high-quality Gezi Opera (a type of Chinese opera popular in Fujian and Taiwan) audio-text alignment dataset and formulate specific data annotation methods applicable to Chinese operas. We propose FT-GAN, an acoustic model for fine-grained tune modeling in Chinese opera synthesis based on the empirical analysis of the differences between Chinese operas and pop songs. To further improve the quality of the synthesized opera, we propose a speech pre-training strategy for additional knowledge injection. The experimental results show that FT-GAN outperforms the strong baselines in SVS on the Gezi Opera synthesis task. Extensive experiments further verify that FT-GAN performs well on synthesis tasks of other operas such as Peking Opera. Audio samples, the dataset, and the codes are available at https://zhengmidon.github.io/FTGAN.github.io/. Meizhen Zheng, Peng Bai 0006, Xiaodong Shi, Yiting Yan |
AAAI | 3 |
| 2024 | Layer-Wise Representation Fusion for Compositional GeneralizationabstractExisting neural models are demonstrated to struggle with compositional generalization (CG), i.e., the ability to systematically generalize to unseen compositions of seen components. A key reason for failure on CG is that the syntactic and semantic representations of sequences in both the uppermost layer of the encoder and decoder are entangled. However, previous work concentrates on separating the learning of syntax and semantics instead of exploring the reasons behind the representation entanglement (RE) problem to solve it. We explain why it exists by analyzing the representation evolving mechanism from the bottom to the top of the Transformer layers. We find that the ``shallow'' residual connections within each layer fail to fuse previous layers' information effectively, leading to information forgetting between layers and further the RE problems. Inspired by this, we propose LRF, a novel Layer-wise Representation Fusion framework for CG, which learns to fuse previous layers' information back into the encoding and decoding process effectively through introducing a fuse-attention module at each encoder and decoder layer. LRF achieves promising results on two realistic benchmarks, empirically demonstrating the effectiveness of our proposal. Codes are available at https://github.com/thinkaboutzero/LRF. Yafang Zheng, Shuangtao Li, Zhaohong Lai, Biao Fu, Yidong Chen 0001, Xiaodong Shi |
AAAI | 9 |
| 2024 | Adaptive Simultaneous Sign Language Translation with Confident Translation Length EstimationabstractTraditional non-simultaneous Sign Language Translation (SLT) methods, while effective for pre-recorded videos, face challenges in real-time scenarios due to inherent inference delays. The emerging field of simultaneous SLT aims to address this issue by progressively translating incrementally received sign video. However, the sole existing work in simultaneous SLT adopts a fixed gloss-based policy, which suffer from limitations in boundary prediction and contextual comprehension. In this paper, we delve deeper into this area and propose an adaptive policy for simultaneous SLT. Our approach introduces the concept of “confident translation length”, denoting maximum accurate translation achievable from current input. An estimator measures this length for streaming sign video, enabling the model to make informed decisions on whether to wait for more input or proceed with translation. To train the estimator, we construct a training data of confident translation length based on the longest common prefix between translations of partial and complete inputs. Furthermore, we incorporate adaptive training, utilizing pseudo prefix pairs, to refine the offline translation model for optimal performance in simultaneous scenarios. Experimental results on PHOENIX2014T and CSL-Daily demonstrate the superiority of our adaptive policy over existing methods, particularly excelling in situations requiring extremely low latency. Biao Fu, Ruiquan Zhang, Xiaodong Shi, Jinsong Su, Yidong Chen 0001 |
LREC/COLING | 6 |
| 2024 | An Explicit Multi-Modal Fusion Method for Sign Language TranslationabstractSign Language Translation (SLT) aims to convert sign language videos into corresponding spoken text sequences. However, the inherent modality gap between sign language video and text hinders the development of SLT. Motivated by the linguistic consistency between gloss1and text, we propose EMF-SLT, an Explicit Multi-modal Fusion method for Sign Language Translation to mitigate the modality gap with the help of gloss. Specifically, EMF-SLT first leverages a vector quantizer and a fusion module to align and fuse sign language and gloss features, respectively, resulting in more informative multi-modal features for the decoder. Then, a multi-task mutual learning framework is introduced to regularize the output predictions from different modalities, which ensures the consistency of outputs across modalities and encourages different modalities to learn from each other. Experiments on two SLT benchmarks and further analyses show that our method achieves significant improvements over the baselines and effectively alleviates the modality gap. Biao Fu, Pei Yu, Xiaodong Shi, Yidong Chen 0001 |
ICASSP | 5 |
| 2024 | Memory-Augmented speech-to-text Translation with Multi-Scale Context Translation StrategyabstractEnd-to-end speech-to-text translation (ST) has demonstrated promising results on sentence-level translation. In real-world scenarios, audio is typically long and requires cross-sentence contextual connections for translation. Sentence-level ST models are facing challenges since they lack the ability to understand inter-sentential context. As context information has been proved to be effective for document-level machine translation, however, research on incorporating context information into ST remains under-explored. In this paper, we propose memory-augmented speech-to-text translation, which leverages a memory module to perform context-aware translation. To enhance the ability of the memory module to extract information from context, we develop Multi-Scale Context Translation Strategy (MSCTS) that translates segments with different size of context. Experiments on MuST-C benchmark show that our proposed method can significantly improve context-aware ST, outperforming the strong sentence-level baseline by +0.8 BLEU in average. Yue Zhou 0012, Xiaodong Shi |
ICASSP | 3 |
| 2024 | A multitask co-training framework for improving speech translation by leveraging speech recognition and machine translation tasks
Yue Zhou 0012, Xiaodong Shi |
Neural Comput. Appl. | 3 |
| 2024 | Improving few-shot relation extraction through semantics-guided learning
Hui Wu 0008, Yidong Chen 0001, Xiaodong Shi |
Neural Networks | 5 |
| 2023 | Exploring Self-Distillation Based Relational Reasoning Training for Document-Level Relation ExtractionabstractDocument-level relation extraction (RE) aims to extract relational triples from a document. One of its primary challenges is to predict implicit relations between entities, which are not explicitly expressed in the document but can usually be extracted through relational reasoning. Previous methods mainly implicitly model relational reasoning through the interaction among entities or entity pairs. However, they suffer from two deficiencies: 1) they often consider only one reasoning pattern, of which coverage on relational triples is limited; 2) they do not explicitly model the process of relational reasoning. In this paper, to deal with the first problem, we propose a document-level RE model with a reasoning module that contains a core unit, the reasoning multi-head self-attention unit. This unit is a variant of the conventional multi-head self-attention and utilizes four attention heads to model four common reasoning patterns, respectively, which can cover more relational triples than previous methods. Then, to address the second issue, we propose a self-distillation training framework, which contains two branches sharing parameters. In the first branch, we first randomly mask some entity pair feature vectors in the document, and then train our reasoning module to infer their relations by exploiting the feature information of other related entity pairs. By doing so, we can explicitly model the process of relational reasoning. However, because the additional masking operation is not used during testing, it causes an input gap between training and testing scenarios, which would hurt the model performance. To reduce this gap, we perform conventional supervised training without masking operation in the second branch and utilize Kullback-Leibler divergence loss to minimize the difference between the predictions of the two branches. Finally, we conduct comprehensive experiments on three benchmark datasets, of which experimental results demonstrate that our model consistently outperforms all competitive baselines. Our source code is available at https://github.com/DeepLearnXMU/DocRE-SD Jinsong Su, Zijun Min, Zhongjian Miao, Qingguo Hu, Biao Fu, Xiaodong Shi, Yidong Chen 0001 |
AAAI | 7 |
| 2023 | Consistent Prototype Learning for Few-Shot Continual Relation ExtractionabstractFew-shot continual relation extraction aims to continually train a model on incrementally fewshot data to learn new relations while avoiding forgetting old ones.However, current memory-based methods are prone to overfitting memory samples, resulting in insufficient activation of old relations and limited ability to handle the confusion of similar classes.In this paper, we design a new N-way-K-shot Continual Relation Extraction (NK-CRE) task and propose a novel few-shot continual relation extraction method with Consistent Prototype Learning (ConPL) to address the aforementioned issues.Our proposed ConPL is mainly composed of three modules: 1) a prototypebased classification module that provides primary relation predictions under few-shot continual learning; 2) a memory-enhanced module designed to select vital samples and refined prototypical representations as a novel multi-information episodic memory; 3) a consistent learning module to reduce catastrophic forgetting by enforcing distribution consistency.To effectively mitigate catastrophic forgetting, ConPL ensures that the samples and prototypes in the episodic memory remain consistent in terms of classification and distribution.Additionally, ConPL uses prompt learning to extract better representations and adopts a focal loss to alleviate the confusion of similar classes.Experimental results on two commonly-used datasets show that our model consistently outperforms other competitive baselines 1 . Xiudi Chen, Hui Wu 0008, Xiaodong Shi |
ACL (1) | 3 |
| 2023 | A Compact Phoneme-To-Audio Aligner for Singing Voice
Meizhen Zheng, Peng Bai 0006, Xiaodong Shi |
ADMA (2) | 3 |
| 2023 | Improving Chinese Pop Song and Hokkien Gezi Opera Singing Voice Synthesis by Enhancing Local ModelingabstractSinging Voice Synthesis (SVS) strives to synthesize pleasant vocals based on music scores and lyrics.The current acoustic models based on Transformer usually process the entire sequence globally and use a simple L1 loss.However, this approach overlooks the significance of local modeling within the sequence and the local optimization of the hard-to-synthesize parts in the predicted mel-spectrogram.Consequently, the synthesized audio exhibits local incongruities (e.g., local pronunciation jitter or noise).To address this problem, we propose two methods to enhance local modeling in the acoustic model.First, we devise a nearest neighbor local attention, where each phoneme token focuses only on the adjacent phoneme tokens located before and after it.Second, we propose a phoneme-level local adaptive weights loss function that enables the model to focus more on the hard-to-synthesize parts of the mel-spectrogram.We verify the universality of our methods on public Chinese pop song and Hokkien Gezi Opera datasets.Extensive experiments demonstrate the effectiveness of our methods, resulting in significant improvements in both objective and subjective evaluations when compared to the strong baselines.Our code and demonstration samples are available at https://github.com/baipeng1/SVSELM. Peng Bai 0006, Yue Zhou 0012, Meizhen Zheng, Wujin Sun, Xiaodong Shi |
EMNLP | 5 |
| 2023 | Adapting Offline Speech Translation Models for Streaming with Future-Aware Distillation and InferenceabstractA popular approach to streaming speech translation is to employ a single offline model with a wait-k policy to support different latency requirements, which is simpler than training multiple online models with different latency constraints.However, there is a mismatch problem in using a model trained with complete utterances for streaming inference with partial input.We demonstrate that speech representations extracted at the end of a streaming input are significantly different from those extracted from a complete utterance.To address this issue, we propose a new approach called Future-Aware Streaming Translation (FAST) that adapts an offline ST model for streaming input.FAST includes a Future-Aware Inference (FAI) strategy that incorporates future context through a trainable masked embedding, and a Future-Aware Distillation (FAD) framework that transfers future context from an approximation of full speech to streaming input.Our experiments on the MuST-C EnDe, EnEs, and EnFr benchmarks show that FAST achieves better trade-offs between translation quality and latency than strong baselines.Extensive analyses suggest that our methods effectively alleviate the aforementioned mismatch problem between offline training and online inference.1 Biao Fu, Minpeng Liao, Kai Fan 0002, Zhongqiang Huang, Boxing Chen, Yidong Chen 0001, Xiaodong Shi |
EMNLP | 7 |
| 2023 | A Token-Level Contrastive Framework for Sign Language TranslationabstractSign Language Translation (SLT) is a promising technology to bridge the communication gap between the deaf and the hearing people. Recently, researchers have adopted Neural Machine Translation (NMT) methods, which usually require large-scale corpus for training, to achieve SLT. However, the publicly available SLT corpus is very limited, which causes the collapse of the token representations and the inaccuracy of the generated tokens. To alleviate this issue, we propose Con-SLT, a novel token-level Contrastive learning framework for Sign Language Translation , which learns effective token representations by incorporating token-level contrastive learning into the SLT decoding process. Concretely, ConSLT treats each token and its counterpart generated by different dropout masks as positive pairs during decoding, and then randomly samples K tokens in the vocabulary that are not in the current sentence to construct negative examples. We conduct comprehensive experiments on two benchmarks (PHOENIX14T and CSL-Daily) for both end-to-end and cascaded settings. The experimental results demonstrate that ConSLT can achieve better translation quality than the strong baselines1. Biao Fu, Peigen Ye, Pei Yu, Xiaodong Shi, Yidong Chen 0001 |
ICASSP | 6 |
| 2023 | LEAPT: Learning Adaptive Prefix-to-Prefix Translation For Simultaneous Machine TranslationabstractSimultaneous machine translation, which aims at a realtime translation, is useful in many live scenarios but very challenging due to the trade-off between accuracy and latency. To achieve the balance for both, the model needs to wait for appropriate streaming text (READ policy) and then generates its translation (WRITE policy). However, WRITE policies of previous work either are specific to the method itself due to the end-to-end training or suffer from the input mismatch between training and decoding for the non-end-to-end training. Therefore, it is essential to learn a generic and better WRITE policy for simultaneous machine translation. Inspired by strategies utilized by human interpreters and "wait" policies, we propose a novel adaptive prefix-to-prefix training policy called LEAPT, which allows our machine translation model to learn how to translate source sentence prefixes and make use of the future context. Experiments show that our proposed methods greatly outperform competitive baselines and achieve promising results. Shuangtao Li, Xiaodong Shi |
ICASSP | 3 |
| 2023 | A High-Quality Melody-Aware Peking Opera Synthesizer Using Data AugmentationabstractThe performing art of Peking Opera places great demands on the singing skills of singers, including pronunciation, melody, role, personal style and emotional expression, which poses a great challenge to Peking Opera singing voice synthesis. In this paper, we propose OperaSinger, following the main architecture of FastSpeech 2, using the features from the musical score as input, while improving the encoder in FastSpeech 2 by employing a stack of melody-aware location-variable convolution blocks in parallel with feed-forward Transformer blocks to alleviate the lack of naturalness caused by ignoring relatively local features. Due to the limitation of publicly available opera data, we explore several novel data augmentation methods to boost the training of OperaSinger. Extensive experiment results have demonstrated that 1) OperaSinger can generate high-quality Peking Opera samples (MOS 3.80) with naturalness and expressiveness; 2) the proposed data augmentation methods effectively improve performance on both subjective and objective evaluations. Wujin Sun, Xiaodong Shi |
ICME | 3 |
| 2023 | Effective Guidance in Zero-Shot Multilingual Translation via Multiple Language Prototypes
Yafang Zheng, Xiaodong Shi |
ICONIP (6) | 4 |
| 2023 | COSYWA: Enhancing Semantic Integrity in Watermarking Natural Language Generation
Junjie Fang, Zhixing Tan, Xiaodong Shi |
NLPCC (1) | 3 |
| 2023 | A Novel POS-Guided Data Augmentation Method for Sign Language Gloss Translation
Yafang Zheng, Yidong Chen 0001, Xiaodong Shi |
NLPCC (2) | 5 |
| 2023 | MOPRD: A multidisciplinary open peer review dataset
Jialiang Lin 0001, Zhangping Zhou, Yidong Chen 0001, Xiaodong Shi |
Neural Comput. Appl. | 5 |
| 2022 | Adversarial Soft Prompt Tuning for Cross-Domain Sentiment AnalysisabstractCross-domain sentiment analysis has achieved promising results with the help of pre-trained language models.As GPT-3 appears, prompt tuning has been widely explored to enable better semantic modeling in many natural language processing tasks.However, directly using a fixed predefined template for crossdomain research cannot model different distributions of the [MASK] token in different domains, thus making underuse of the prompt tuning technique.In this paper, we propose a novel Adversarial Soft Prompt Tuning method (AdSPT) to better model cross-domain sentiment analysis.On the one hand, AdSPT adopts separate soft prompts instead of hard templates to learn different vectors for different domains, thus alleviating the domain discrepancy of the [MASK] token in the masked language modeling task.On the other hand, AdSPT uses a novel domain adversarial training strategy to learn domain-invariant representations between each source domain and the target domain.Experiments on a publicly available sentiment analysis dataset show that our model achieves new state-of-the-art results for both single-source domain adaptation and multi-source domain adaptation. Hui Wu 0008, Xiaodong Shi |
ACL (1) | 2 |
| 2022 | Towards Better Document-level Relation Extraction via Iterative InferenceabstractDocument-level relation extraction (RE) aimsto extract the relations between entities from the input document that usually containing many difficultly-predicted entity pairs whose relations can only be predicted through relational inference.Existing methods usually directly predict the relations of all entity pairs of input document in a one-pass manner, ignoring the fact that predictions of some entity pairs heavily depend on the predicted results of other pairs.To deal with this issue, in this paper, we propose a novel document-level RE model with iterative inference.Our model is mainly composed of two modules: 1) a base module expected to provide preliminary relation predictions on entity pairs; 2) an inference module introduced to refine these preliminary predictions by iteratively dealing with difficultlypredicted entity pairs depending on other pairs in an easy-to-hard manner.Unlike previous methods which only consider feature information of entity pairs, our inference module is equipped with two Extended Cross Attention units, allowing it to exploit both feature information and previous predictions of entity pairs during relational inference.Furthermore, we adopt a two-stage strategy to train our model.At the first stage, we only train our base module.During the second stage, we train the whole model, where contrastive learning is introduced to enhance the training of inference module.Experimental results on three commonly-used datasets show that our model consistently outperforms other competitive baselines.Our source code is available at https://github. com/DeepLearnXMU/DocRE-II. Jinsong Su, Yidong Chen 0001, Zhongjian Miao, Zijun Min, Qingguo Hu, Xiaodong Shi |
EMNLP | 7 |
| 2022 | FCH-TTS: Fast, Controllable and High-quality Non-Autoregressive Text-to-Speech SynthesisabstractInspired by the success of the non-autoregressive speech synthesis model FastSpeech, we propose FCH-TTS, a fast, controllable and universal neural text-to-speech (TTS) capable of generating high-quality spectrograms. The basic architecture of FCH-TTS is similar to that of FastSpeech, but FCH-TTS uses a simple yet effective attention-based soft alignment mechanism to replace the complex teacher model in FastSpeech, allowing the model to be better adapted to different languages. Specifically, in addition to the control of voice speed and prosody, a fusion module has been designed to better model speaker features in order to obtain the desired timbre. Meanwhile, several special loss functions were applied to ensure the quality of the output mel-spectrogram. Experimental results on the dataset LJSpeech show that FCH-TTS achieves the fastest inference speed compared to all baseline models, while also achieving the best speech quality. In addition, the controllability of the model with respect to prosody, voice speed and timbre was validated on several datasets, and the good performance on the low-resource Tibetan dataset demonstrates the universality of the model. Xiaodong Shi |
IJCNN | 3 |
| 2022 | Continuous Prompt Enhanced Biomedical Entity Normalization
Zhaohong Lai, Biao Fu, Shangfei Wei, Xiaodong Shi |
NLPCC (2) | 4 |
| 2022 | Automatic Analysis of Available Source Code of Top Artificial Intelligence Conference PapersabstractSource code is essential for researchers to reproduce the methods and replicate the results of artificial intelligence (AI) papers. Some organizations and researchers manually collect AI papers with available source code to contribute to the AI community. However, manual collection is a labor-intensive and time-consuming task. To address this issue, we propose a method to automatically identify papers with available source code and extract their source code repository URLs. With this method, we find that 20.5% of regular papers of 10 top AI conferences published from 2010 to 2019 are identified as papers with available source code and that 8.1% of these source code repositories are no longer accessible. We also create the XMU NLP Lab README Dataset, the largest dataset of labeled README files for source code document research. Through this dataset, we have discovered that quite a few README files have no installation instructions or usage tutorials provided. Further, a large-scale comprehensive statistical analysis is made for a general picture of the source code of AI conference papers. The proposed solution can also go beyond AI conference papers to analyze other scientific papers from both journals and conferences to shed light on more domains. Jialiang Lin 0001, Yingmin Wang, Yao Yu 0001, Yu Zhou 0007, Yidong Chen 0001, Xiaodong Shi |
Int. J. Softw. Eng. Knowl. Eng. | 6 |
| 2022 | A downsampling method enables robust clustering and integration of single-cell transcriptome dataabstractThe random noises, sampling biases, and batch effects often confound true biological variations in single-cell RNA-sequencing (scRNA-seq) data. Adjusting such biases is key to the robust discoveries in downstream analyses, such as cell clustering, gene selection and data integration. Here we propose a model-based downsampling algorithm based on minimal unbiased representative points (MURPXMBD). MURPXMBD is designed to retrieve a set of representative points by reducing gene-wise random independent errors, while retaining the covariance structure of biological origin hence provide an unbiased representation of the cell population. Subsequent validation using benchmark datasets shows that MURPXMBD can improve the quality and accuracy of clustering algorithms, and thus facilitate the discovery of new cell types. Besides, MURPXMBD also improves the performance of dataset integration algorithms. In summary, MURPXMBD serves as a useful noise-reduction method for single-cell sequencing analysis in biomedical studies. Yudi Hu, Xuejing Lyu, Hongkun Fang, Rongshan Yu, Xiaodong Shi |
J. Biomed. Informatics | 9 |
| 2021 | Synchronous Dual Network with Cross-Type Attention for Joint Entity and Relation ExtractionabstractJoint entity and relation extraction is challenging due to the complex interaction of interaction between named entity recognition and relation extraction.Although most existing works tend to jointly train these two tasks through a shared network, they fail to fully utilize the interdependence between entity types and relation types.In this paper, we design a novel synchronous dual network (SDN) with cross-type attention via separately and interactively considering the entity types and relation types.On the one hand, SDN adopts two isomorphic bi-directional type-attention LSTM to encode the entity type enhanced representations and the relation type enhanced representations, respectively.On the other hand, SDN explicitly models the interdependence between entity types and relation types via crosstype attention mechanism.In addition, we also propose a new multi-task learning strategy via modeling the interaction of two types of information.Experiments on NYT and WebNLG datasets verify the effectiveness of the proposed model, achieving state-of-the-art performance. Hui Wu 0008, Xiaodong Shi |
EMNLP (1) | 2 |
| 2021 | CTRD: A Chinese Theme-Rheme Discourse Dataset
Biao Fu, Yiqi Tong, Dawei Tian, Yidong Chen 0001, Xiaodong Shi |
NLPCC (1) | 5 |
| 2021 | Enhancing Neural Sign Language Translation by highlighting the facial expression information
Jiangbin Zheng 0002, Yidong Chen 0001, Chong Wu 0007, Xiaodong Shi, Suhail Muhammad Kamal |
Neurocomputing | 4 |
| 2021 | Knowledge Graph Embedding Based on Multi-View Clustering FrameworkabstractKnowledge representation is one of the critical problems in knowledge engineering and artificial intelligence, while knowledge embedding as a knowledge representation methodology indicates entities and relations in knowledge graph as low-dimensional, continuous vectors. In this way, knowledge graph is compatible with numerical machine learning models. Major knowledge embedding methods employ geometric translation to design score function, which is weak-semantic for natural language processing. To overcome this disadvantage, in this paper, we propose our model based on multi-view clustering framework, which could generate semantic representations of knowledge elements (i.e., entities/relations). With our semantic model, we also present an empowered solution to entity retrieval with entity description. Extensive experiments show that our model achieves substantial improvements against baselines on the task of knowledge graph completion, triple classification, entity classification, and entity retrieval. Yidong Chen 0001, Xiaodong Shi |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2020 | A Document-Level Neural Machine Translation Model with Dynamic Caching Guided by Theme-Rheme InformationabstractResearch on document-level Neural Machine Translation (NMT) models has attracted increasing attention in recent years.Although the proposed works have proved that the inter-sentence information is helpful for improving the performance of the NMT models, what information should be regarded as context remains ambiguous.To solve this problem, we proposed a novel cache-based document-level NMT model which conducts dynamic caching guided by theme-rheme information.The experiments on NIST evaluation sets demonstrate that our proposed model achieves substantial improvements over the state-of-the-art baseline NMT models.As far as we know, we are the first to introduce theme-rheme theory into the field of machine translation. Yiqi Tong, Jiangbin Zheng 0002, Hongkang Zhu, Yidong Chen 0001, Xiaodong Shi |
COLING | 5 |
| 2019 | Boosting implicit discourse relation recognition with connective-based word embeddings
Changxing Wu, Jinsong Su, Yidong Chen 0001, Xiaodong Shi |
Neurocomputing | 4 |
| 2019 | Multi-perspective neural architecture for recommendation system
Yidong Chen 0001, Xiaodong Shi |
Neural Networks | 3 |
| 2018 | Deep Semantic Role Labeling With Self-AttentionabstractSemantic Role Labeling (SRL) is believed to be a crucial step towards natural language understanding and has been widely studied. Recent years, end-to-end SRL with recurrent neural networks (RNN) has gained increasing attention. However, it remains a major challenge for RNNs to handle structural information and long range dependencies. In this paper, we present a simple and effective architecture for SRL which aims to address these problems. Our model is based on self-attention which can directly capture the relationships between two tokens regardless of their distance. Our single model achieves F1=83.4 on the CoNLL-2005 shared task dataset and F1=82.7 on the CoNLL-2012 shared task dataset, which outperforms the previous state-of-the-art results by 1.8 and 1.0 F1 score respectively. Besides, our model is computationally efficient, and the parsing speed is 50K tokens per second on a single Titan X GPU. Zhixing Tan, Mingxuan Wang, Yidong Chen 0001, Xiaodong Shi |
AAAI | 5 |
| 2018 | Lattice-to-sequence attentional Neural Machine Translation models
Zhixing Tan, Jinsong Su, Boli Wang, Yidong Chen 0001, Xiaodong Shi |
Neurocomputing | 5 |
| 2018 | Constructing and validating word similarity datasets by integrating methods from psychology, brain science and computational linguistics
Yu Wan 0004, Yidong Chen 0001, Xiaodong Shi, Changle Zhou |
Soft Comput. | 3 |
| 2017 | Lattice-Based Recurrent Neural Network Encoders for Neural Machine TranslationabstractNeural machine translation (NMT) heavily relies on word-level modelling to learn semantic representations of input sentences.However, for languages without natural word delimiters (e.g., Chinese) where input sentences have to be tokenized first,conventional NMT is confronted with two issues:1) it is difficult to find an optimal tokenization granularity for source sentence modelling, and2) errors in 1-best tokenizations may propagate to the encoder of NMT.To handle these issues, we propose word-lattice based Recurrent Neural Network (RNN) encoders for NMT,which generalize the standard RNN to word lattice topology.The proposed encoders take as input a word lattice that compactly encodes multiple tokenizations, and learn to generate new hidden states from arbitrarily many inputs and hidden states in preceding time steps.As such, the word-lattice based encoders not only alleviate the negative impact of tokenization errors but also are more expressive and flexible to embed input sentences.Experiment results on Chinese-English translation demonstrate the superiorities of the proposed encoders over the conventional encoder. Jinsong Su, Zhixing Tan, Deyi Xiong, Rongrong Ji, Xiaodong Shi, Yang Liu 0005 |
AAAI | 5 |
| 2017 | A synergetic semantic role labeling model with the introduction of fluctuating force accompanied with word sense informationabstractSemantic role labeling (SRL) is a key problem in natural language processing which goal is to find a sentence-level semantic representation. Word sense information plays an important role on the determination of semantic roles. The introduction of word sense in the process of semantic role labeling will hopefully lead to achieve better result. But how to better reflect the relationship between word sense information and semantic role information is a key task. Synergetic neural network (SNN) provides an opportunity for us to study how to use word sense for semantic role labeling. The role labeling process can be seen as a competition process of many roles chain order parameters with word sense, of which order parameter with the largest support will win, thereby obtaining desired pattern. There are three main contributions in this article: firstly, we introduce synergetic theory to semantic analysis and propose a semantic analysis method based on synergetic neural network, which can effectively use semantic information and word sense information. Secondly, fluctuating force is introduced into potential evolution function which can effectively make use of prior semantic knowledge. Finally, we use artificial fish swarm algorithm (AFSA) to realize the optimization of network parameter which has both global and local search ability, and not easy to fall into local extremism. Experiment results show the proposed model in this paper can further improve the performance of semantic role labeling, and thus provides an important reference value to future research. Zhehuang Huang, Yidong Chen 0001, Xiaodong Shi |
Intell. Data Anal. | 3 |
| 2017 | Leveraging bilingually-constrained synthetic data via multi-task neural networks for implicit discourse relation recognition
Changxing Wu, Xiaodong Shi, Yidong Chen 0001, Yanzhou Huang, Jinsong Su |
Neurocomputing | 2 |
| 2017 | Co-training for Implicit Discourse Relation Recognition Based on Manual and Distributed Features
Changxing Wu, Xiaodong Shi, Jinsong Su, Yidong Chen 0001, Yanzhou Huang |
Neural Process. Lett. | 2 |
| 2016 | Bilingually-constrained Synthetic Data for Implicit Discourse Relation RecognitionabstractTo alleviate the shortage of labeled data, we propose to use bilingually-constrained synthetic implicit data for implicit discourse relation recognition.These data are extracted from a bilingual sentence-aligned corpus according to the implicit/explicit mismatch between different languages.Incorporating these data via a multi-task neural network model achieves significant improvements over baselines, on both the English PDTB and Chinese CDTB data sets. Changxing Wu, Xiaodong Shi, Yidong Chen 0001, Yanzhou Huang, Jinsong Su |
EMNLP | 2 |
| 2016 | Sentiment analysis via integrating distributed representations of variable-length word sequence
Zhijian Cui, Xiaodong Shi, Yidong Chen 0001 |
Neurocomputing | 2 |
| 2016 | Adapted competitive learning on continuous semantic space for word sense induction
Yanzhou Huang, Deyi Xiong, Xiaodong Shi, Yidong Chen 0001, Changxing Wu, Guimin Huang |
Neurocomputing | 3 |
| 2016 | An SNN-Based Semantic Role Labeling Model with Its Network Parameters Optimized Using an Improved PSO Algorithm
Yidong Chen 0001, Zhehuang Huang, Xiaodong Shi |
Neural Process. Lett. | 3 |
| 2015 | Cross-Lingual Tense Tagging Based on Markov Tree Tagging ModelabstractIn this paper, we transform the issue of Chinese-English tense conversion into the issue of tagging a Chinese tense tree. And then we propose Markov Tree Tagging Model to tag nodes of the untagged tense tree with English tenses. Experimental results show that the method is much better than linear-based CRF tagging for the issue. Chang Su 0006, Xiaodong Shi |
NLPCC | 4 |
| 2015 | Unsupervised word sense induction using rival penalized competitive learning
Yanzhou Huang, Xiaodong Shi, Jinsong Su, Yidong Chen 0001, Guimin Huang |
Eng. Appl. Artif. Intell. | 2 |
| 2014 | Topic-aware pivot language approach for statisticalmachine translationabstractThe pivot language approach for statistical machine translation (SMT) is a good method to break the resource bottleneck for certain language pairs. However, in the implementation of conventional approaches, pivot-side context information is far from fully utilized, resulting in erroneous estimations of translation probabilities. In this study, we propose two topic-aware pivot language approaches to use different levels of pivot-side context. The first method takes advantage of document-level context by assuming that the bridged phrase pairs should be similar in the document-level topic distributions. The second method focuses on the effect of local context. Central to this approach are that the phrase sense can be reflected by local context in the form of probabilistic topics, and that bridged phrase pairs should be compatible in the latent sense distributions. Then, we build an interpolated model bringing the above methods together to further enhance the system performance. Experimental results on French-Spanish and French-German translations using English as the pivot language demonstrate the effectiveness of topic-based context in pivot-based SMT. Jinsong Su, Xiaodong Shi, Yanzhou Huang, Yang Liu 0005, Qingqiang Wu 0001, Yidong Chen 0001, Huailin Dong |
J. Zhejiang Univ. Sci. C | 2 |
| 2012 | Translation Model Adaptation for Statistical Machine Translation with Monolingual Topic Information
Jinsong Su, Hua Wu 0003, Haifeng Wang 0001, Yidong Chen 0001, Xiaodong Shi, Huailin Dong, Qun Liu 0001 |
ACL (1) | 5 |
| 2012 | Strip-oriented asynchronous prefetching for parallel disk systemsabstractSequential prefetching schemes are widely employed in storage servers to mask disk latency and improve system throughput. However, existing schemes cannot benefit parallel disk systems as expected due to the fact that they ignore the distinct internal characteristics of the parallel disk system, in particular, data striping. Moreover, their aggressive prefetching pattern suffers from premature evictions and prolonged request latencies. In this paper, we propose a strip-oriented asynchronous prefetching (SoAP) technique, which is dedicated to the parallel disk system. It settles the above-mentioned problems by providing multiple novel features, e.g., enhanced prediction accuracy, adaptive prefetching strength, physical data layout awareness, and timely prefetching. To validate SoAP, we implement a prototype by modifying the software redundant arrays of inexpensive disks (RAID) under Linux. Experimental results demonstrate that SoAP can consistently offer improved average response time and throughput to the parallel disk system under non-random workloads compared with STEP, SP, ASP, and Linux-like SEQPs. Yang Liu 0211, Jianzhong Huang 0001, Xiaodong Shi, Qiang Cao 0001, Changsheng Xie 0001 |
J. Zhejiang Univ. Sci. C | 3 |
| 2011 | Improving the Hierarchical Phrase-Based Translation Model
Xiaodong Shi, Yidong Chen 0001 |
MTSummit | 1 |
| 2009 | Metaphor Recognition: CHMETA, A Pattern-Based SystemabstractMetaphor recognition presents a computational challenge, in part due to metaphoric deviation from literal thinking, and also because of a metaphor's various linguistic expressions. This article forwards a new computational method, an integrated treatment of metaphor recognition from the computational perspective, which recent related studies have not entirely addressed. The authors differentiate metaphor recognition from complex metaphor inference and interpretation employing psychological clues. To accomplish this, we have developed a formalized system of metaphorical expression in metaphor role dependency schema, which specifically defines, classifies, and quantifies metaphorical anomalies, building a computable classification system for metaphors (incorporating 32 major patterns of metaphorical expressions) by providing a strategy to locate potential metaphorical anomalies in a target input sentence through a pattern recognition method and a metaphor components' tagging approach. This metaphor recognition and tagging system is named and implemented as “CHMeta.” Experiment results support the validity and efficiency of this metaphor recognition system. Compared with most metaphor computation systems, which work mainly on a few examples, this system classifies major metaphorical expressions from a computational perspective and is able to recognize a variety of different kinds of metaphors, including nested ones. Thus, this is the first integrated work in computable classification, recognition, and tagging of large‐scale metaphors in Chinese. Changle Zhou, Xiaodong Shi |
Comput. Intell. | 5 |
| 2009 | Discovering Event Evolution Graphs From News CorporaabstractGiven the advance of Internet technologies, we can now easily extract hundreds or thousands of news stories of any ongoing incidents from newswires such as CNN.com, but the volume of information is too large for us to capture the blueprint. Information retrieval techniques such as topic detection and tracking are able to organize news stories as events, in a flat hierarchical structure, within a topic. However, they are incapable of presenting the complex evolution relationships between the events. We are interested to learn not only what the major events are but also how they develop within the topic. It is beneficial to identify the seminal events, the intermediary and ending events, and the evolution of these events. In this paper, we propose to utilize the event timestamp, event content similarity, temporal proximity, and document distributional proximity to model the event evolution relationships between events in an incident. An event evolution graph is constructed to present the underlying structure of events for efficient browsing and extracting of information. Case study and experiments are presented to illustrate and show the performance of our proposed technique. It is found that our proposed technique outperforms the baseline technique and other comparable techniques in previous work. Christopher C. Yang, Xiaodong Shi, Chih-Ping Wei |
IEEE Trans. Syst. Man Cybern. Part A | 2 |
| 2007 | Dependency-Based Chinese-English Statistical Machine Translation
Xiaodong Shi, Yidong Chen 0001, Jianfeng Jia |
CICLing | 1 |
| 2007 | Translation Memory Sharing Models in XMCATabstractIn this paper, two Translation Memory (TM) sharing models adopted in XMCAT, a Computer Assisted Translation tool (CAT) supporting cooperated work in machine translation, was described in detail. One is Center-based TM sharing model, which is only fit for users in a local area network (LAN) and the other is a novel model called P2P-based TM sharing model, which could be used through Internet by geographically distributed users. With the two TM sharing models, a user may share data with other users through network, so that he/she may reduce the repeated work further and cooperate with others more easily. Besides, the methods used in XMCA T to deal with the problem of multi-translations arose in the cooperated memory sharing models, were also proposed in this paper. XMCAT system has been adopted and approved by some translation companies. Yidong Chen 0001, Xiaodong Shi, Changle Zhou, Tangqiu Li, Qingyang Hong |
CSCWD | 2 |
| 2007 | A VoIP Client System Framework and its ImplementationabstractThis paper describes the design of the whole framework of a VoIP Client System. In this design, we use the design idea of automaton to build the entire dialing and hang up processes, use multithread to implement event synchronous mechanism. Then we use a strategy to couple the Jitter-Buffer design with the automatic adjustment of the redundant packets. This strategy gets a good trade-off at the client between the methods of solving delay and those of solving dithering. The experiment result indicates that this framework keeps good QoS performance even in bad network conditions. Fengmei Zou, Yuhui Chen, Shaozi Li, Tangqiu Li, Xiaodong Shi |
CSCWD | 6 |
| 2007 | Mining related queries from Web search engine query logs using an improved association rule mining modelabstractAbstract With the overwhelming volume of information, the task of finding relevant information on a given topic on the Web is becoming increasingly difficult. Web search engines hence become one of the most popular solutions available on the Web. However, it has never been easy for novice users to organize and represent their information needs using simple queries. Users have to keep modifying their input queries until they get expected results. Therefore, it is often desirable for search engines to give suggestions on related queries to users. Besides, by identifying those related queries, search engines can potentially perform optimizations on their systems, such as query expansion and file indexing. In this work we propose a method that suggests a list of related queries given an initial input query. The related queries are based in the query log of previously submitted queries by human users, which can be identified using an enhanced model of association rules. Users can utilize the suggested related queries to tune or redirect the search process. Our method not only discovers the related queries, but also ranks them according to the degree of their relatedness. Unlike many other rival techniques, it also performs reasonably well on less frequent input queries. Xiaodong Shi, Christopher C. Yang |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2006 | Multi-document Summarization for Terrorism Information Extraction
Fu Lee Wang, Christopher C. Yang, Xiaodong Shi |
ISI | 3 |
| 2006 | Tracing the Event Evolution of Terror Attacks from On-Line News
Christopher C. Yang, Xiaodong Shi, Chih-Ping Wei |
ISI | 2 |
| 2006 | Mining related queries from search engine query logsabstractIn this work we propose a method that retrieves a list of related queries given an initial input query. The related queries are based on the query log of previously issued queries by human users, which can be discovered using our improved association rule mining model. Users can use the suggested related queries to tune or redirect the search process. Our method not only discovers the related queries, but also ranks them according to the degree of their relatedness. Unlike many other rival techniques, it exploits only limited query log information and performs relatively better on queries in all frequency divisions. Xiaodong Shi, Christopher C. Yang |
WWW | 1 |
| 2006 | Discovering event evolution graphs from newswiresabstractIn this paper, we propose an approach to automatically mine event evolution graphs from newswires on the Web. Event evolution graph is a directed graph in which the vertices and edges denote news events and the evolutions between events respectively, in a news affair. Our model utilizes the content similarity between events and incorporates temporal proximity and document distributional proximity as decaying functions. Our approach is effective in presenting the inside developments of news affairs along the timeline, which can facilitate users' information browsing tasks. Christopher C. Yang, Xiaodong Shi |
WWW | 2 |