EDBT 2026 Demo / reviewers in the wild / expert
Xinyu Duan
dblp:31/5936 · also Duan Xinyu
· DBLP profile ↗
26ranked-venue papers
3as first author
19since 2021 · last 2025
0000-0002-6803-7964ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 2 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multi-Branch Self-Drafting for LLM Inference AccelerationabstractThe autoregressive decoding paradigm endows large language models (LLMs) with superior language generation capabilities; however, its step-by-step decoding process inherently limits decoding speed. To mitigate these constraints, the prevalent “draft and validation” strategy enables parallel validation of candidate drafts, allowing LLMs to decode multiple tokens simultaneously during one model forward propagation. However, existing methodologies for obtaining drafts often incur additional overhead in communication or training process, or statistical biases from the corpus. To this end, we propose an innovative draft generation and maintenance approach that leverages the capabilities of LLM itself. Specifically, we extend the autoregressive decoding paradigm to a multi-branch drafting procedure, which can efficiently generate draft sequences without any additional models or training process, while preserving the quality of the generated content by maintaining LLM parameters. Experiments across various open-source benchmarks show that our method generates 2.0 to 3.2 tokens per forward step and achieves around 2 times improvement of end-to-end throughput compared to the autoregressive decoding strategy. Zipeng Gao, Qingrong Xia, Tong Xu 0001, Xinyu Duan, Zhi Zheng 0008, Zhefeng Wang 0001, Enhong Chen |
AAAI | 4 |
| 2025 | Accurate KV Cache Quantization with Outlier Tokens TracingabstractThe impressive capabilities of Large Language Models (LLMs) come at the cost of substantial computational resources during deployment. While KV Cache can significantly reduce recomputation during inference, it also introduces additional memory overhead. KV Cache quantization presents a promising solution, striking a good balance between memory usage and accuracy. Previous research has shown that the Keys are distributed by channel, while the Values are distributed by token. Consequently, the common practice is to apply channel-wise quantization to the Keys and token-wise quantization to the Values. However, our further investigation reveals that a small subset of unusual tokens exhibit unique characteristics that deviate from this pattern, which can substantially impact quantization accuracy. To address this, we develop a simple yet effective method to identify these tokens accurately during the decoding process and exclude them from quantization as outlier tokens, significantly improving overall accuracy. Extensive experiments show that our method achieves significant accuracy improvements under 2-bit quantization and can deliver a 6.4 times reduction in memory usage and a 2.3 times increase in throughput. Yi Su 0006, Yuechi Zhou, Quantong Qiu, Juntao Li 0005, Qingrong Xia, Ping Li 0016, Xinyu Duan, Zhefeng Wang 0001, Min Zhang 0005 |
ACL (1) | 7 |
| 2025 | Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow MatchingabstractJialong Zuo, Shengpeng Ji, Minghui Fang, Mingze Li, Ziyue Jiang, Xize Cheng, Xiaoda Yang, Chen Feiyang, Xinyu Duan, Zhou Zhao. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jialong Zuo, Shengpeng Ji, Minghui Fang 0002, Ziyue Jiang 0001, Xize Cheng, Xiaoda Yang, Feiyang Chen 0001, Xinyu Duan, Zhou Zhao 0001 |
ACL (1) | 9 |
| 2025 | Alignment-Augmented Speculative Decoding with Alignment Sampling and Conditional VerificationabstractRecent works have revealed the great potential of speculative decoding in accelerating the autoregressive generation process of large language models.The success of these methods relies on the alignment between draft candidates and the sampled outputs of the target model.Existing methods mainly achieve draft-target alignment with training-based methods, e.g., EAGLE, Medusa, involving considerable training costs.In this paper, we present a trainingfree alignment-augmented speculative decoding algorithm.We propose alignment sampling, which leverages output distribution obtained in the prefilling phase to provide more aligned draft candidates.To further benefit from highquality but non-aligned draft candidates, we also introduce a simple yet effective flexible verification strategy.Through an adaptive probability threshold, our approach can improve generation accuracy while further improving inference efficiency.Experiments on 8 datasets (including question answering, summarization and code completion tasks) show that our approach increases the average generation score by 3.3 points for the LLaMA3 model.Our method achieves a mean acceptance length up to 2.39 and speed up generation by 2.23×. Zhenxu Tian, Juntao Li 0005, Qingrong Xia, Xinyu Duan, Zhefeng Wang 0001, Baoxing Huai, Min Zhang 0005 |
EMNLP | 5 |
| 2025 | Beware of Calibration Data for Pruning Large Language ModelsabstractAs large language models (LLMs) are widely applied across various fields, model
compression has become increasingly crucial for reducing costs and improving
inference efficiency. Post-training pruning is a promising method that does not
require resource-intensive iterative training and only needs a small amount of
calibration data to assess the importance of parameters. Recent research has enhanced post-training pruning from different aspects but few of them systematically
explore the effects of calibration data, and it is unclear if there exist better calibration data construction strategies. We fill this blank and surprisingly observe that
calibration data is also crucial to post-training pruning, especially for high sparsity. Through controlled experiments on important influence factors of calibration
data, including the pruning settings, the amount of data, and its similarity with
pre-training data, we observe that a small size of data is adequate, and more similar data to its pre-training stage can yield better performance. As pre-training data
is usually inaccessible for advanced LLMs, we further provide a self-generating
calibration data synthesis strategy to construct feasible calibration data. Experimental results on recent strong open-source LLMs (e.g., DCLM, and LLaMA-3)
show that the proposed strategy can enhance the performance of strong pruning
methods (e.g., Wanda, DSnoT, OWL) by a large margin (up to 2.68%). Yixin Ji, Yang Xiang 0003, Juntao Li 0005, Qingrong Xia, Ping Li 0016, Xinyu Duan, Zhefeng Wang 0001, Min Zhang 0005 |
ICLR | 6 |
| 2025 | Taming the Titans: A Survey of Efficient LLM Inference ServingabstractLarge Language Models (LLMs) for Generative AI have achieved remarkable progress, evolving into sophisticated and versatile tools widely adopted across various domains and applications. However, the substantial memory overhead caused by their vast number of parameters, combined with the high computational demands of the attention mechanism, poses significant challenges in achieving low latency and high throughput for LLM inference services. Recent advancements, driven by groundbreaking research, have significantly accelerated progress in this field. This paper provides a comprehensive survey of these methods, covering fundamental instance-level approaches, in-depth cluster-level strategies, and emerging scenarios. At the instance level, we review model placement, request scheduling, decoding length prediction, storage management, and the disaggregation paradigm. At the cluster level, we explore GPU cluster deployment, multi-instance load balancing, and cloud service solutions. Additionally, we discuss specific tasks, modules, and auxiliary methods in emerging scenarios. Finally, we outline potential research directions to further advance the field of LLM inference serving. Ranran Zhen, Juntao Li 0005, Yixin Ji, Zhenlin Yang, Qingrong Xia, Xinyu Duan, Zhefeng Wang 0001, Baoxing Huai, Min Zhang 0005 |
INLG | 7 |
| 2025 | OPT-Tree: Speculative Decoding with Adaptive Draft Tree StructureabstractAbstract Autoregressive language models demonstrate excellent performance in various scenarios. However, the inference efficiency is limited by its one-step-one-word generation mode, which has become a pressing problem recently as the models become increasingly larger. Speculative decoding employs a “draft and then verify” mechanism to allow multiple tokens to be generated in one step, realizing lossless acceleration. Existing methods mainly adopt fixed heuristic draft structures, which do not adapt to different situations to maximize the acceptance length during verification. To alleviate this dilemma, we propose OPT-Tree, an algorithm to construct adaptive and scalable draft trees, which can be applied to any autoregressive draft model. It searches the optimal tree structure that maximizes the mathematical expectation of the acceptance length in each decoding step. Experimental results reveal that OPT-Tree outperforms the existing draft structures and achieves a speed-up ratio of up to 3.2 compared with autoregressive decoding. If the draft model is powerful enough and the node budget is sufficient, it can generate more than ten tokens in a single step. Our code is available at https://github.com/Jikai0Wang/OPT-Tree. Yi Su 0006, Juntao Li 0005, Qingrong Xia, Xinyu Duan, Zhefeng Wang 0001, Min Zhang 0005 |
Trans. Assoc. Comput. Linguistics | 6 |
| 2024 | Dialogues Are Not Just Text: Modeling Cognition for Dialogue Coherence EvaluationabstractThe generation of logically coherent dialogues by humans relies on underlying cognitive abilities. Based on this, we redefine the dialogue coherence evaluation process, combining cognitive judgment with the basic text to achieve a more human-like evaluation. We propose a novel dialogue evaluation framework based on Dialogue Cognition Graph (DCGEval) to implement the fusion by in-depth interaction between cognition modeling and text modeling. The proposed Abstract Meaning Representation (AMR) based graph structure called DCG aims to uniformly model four dialogue cognitive abilities. Specifically, core-semantic cognition is modeled by converting the utterance into an AMR graph, which can extract essential semantic information without redundancy. The temporal and role cognition are modeled by establishing logical relationships among the different AMR graphs. Finally, the commonsense knowledge from ConceptNet is fused to express commonsense cognition. Experiments demonstrate the necessity of modeling human cognition for dialogue evaluation, and our DCGEval presents stronger correlations with human judgments compared to other state-of-the-art evaluation metrics. Jia Su 0004, Yang Yang 0137, Zipeng Gao, Xinyu Duan, Yi Guan |
AAAI | 5 |
| 2024 | StyleSinger: Style Transfer for Out-of-Domain Singing Voice SynthesisabstractStyle transfer for out-of-domain (OOD) singing voice synthesis (SVS) focuses on generating high-quality singing voices with unseen styles (such as timbre, emotion, pronunciation, and articulation skills) derived from reference singing voice samples. However, the endeavor to model the intricate nuances of singing voice styles is an arduous task, as singing voices possess a remarkable degree of expressiveness. Moreover, existing SVS methods encounter a decline in the quality of synthesized singing voices in OOD scenarios, as they rest upon the assumption that the target vocal attributes are discernible during the training phase. To overcome these challenges, we propose StyleSinger, the first singing voice synthesis model for zero-shot style transfer of out-of-domain reference singing voice samples. StyleSinger incorporates two critical approaches for enhanced effectiveness: 1) the Residual Style Adaptor (RSA) which employs a residual quantization module to capture diverse style characteristics in singing voices, and 2) the Uncertainty Modeling Layer Normalization (UMLN) to perturb the style attributes within the content representation during the training phase and thus improve the model generalization. Our extensive evaluations in zero-shot style transfer undeniably establish that StyleSinger outperforms baseline models in both audio quality and similarity to the reference singing voice samples. Access to singing voice samples can be found at https://stylesinger.github.io/. Yu Zhang 0126, Rongjie Huang 0001, Ruiqi Li 0002, Jinzheng He, Yan Xia 0006, Feiyang Chen 0001, Xinyu Duan, Baoxing Huai, Zhou Zhao 0001 |
AAAI | 7 |
| 2024 | When and How to Grow? On Efficient Pre-training via Model Growth
Juntao Li 0005, Min Zhang 0005, Zechang Li, Qingrong Xia, Xinyu Duan, Zhefeng Wang 0001, Baoxing Huai |
ACML | 6 |
| 2024 | High-order Joint Constituency and Dependency ParsingabstractThis work revisits the topic of jointly parsing constituency and dependency trees, i.e., to produce compatible constituency and dependency trees simultaneously for input sentences, which is attractive considering that the two types of trees are complementary in representing syntax. The original work of Zhou and Zhao (2019) performs joint parsing only at the inference phase. They train two separate parsers under the multi-task learning framework (i.e., one shared encoder and two independent decoders). They design an ad-hoc dynamic programming-based decoding algorithm of O(n^5) time complexity for finding optimal compatible tree pairs. Compared to their work, we make progress in three aspects: (1) adopting a much more efficient decoding algorithm of O(n^4) time complexity, (2) exploring joint modeling at the training phase, instead of only at the inference phase, (3) proposing high-order scoring components to promote constituent-dependency interaction. We conduct experiments and analysis on seven languages, covering both rich-resource and low-resource scenarios. Results and analysis show that joint modeling leads to a modest overall performance boost over separate modeling, but substantially improves the complete matching ratio of whole trees, thanks to the explicit modeling of tree compatibility. Yanggang Gu, Yang Hou 0001, Zhefeng Wang 0001, Xinyu Duan, Zhenghua Li |
LREC/COLING | 4 |
| 2024 | TextrolSpeech: A Text Style Control Speech Corpus with Codec Language Text-to-Speech ModelsabstractRecently, there has been a growing interest in the field of controllable Text-to-Speech (TTS). While previous studies have relied on users providing specific style factor values based on acoustic knowledge or selecting reference speeches that meet certain requirements, generating speech solely from natural text prompts has emerged as a new challenge for researchers. This challenge arises due to the scarcity of high-quality speech datasets with natural text style prompt and the absence of advanced text-controllable TTS models. In light of this, 1) we propose TextrolSpeech, which is the first large-scale speech emotion dataset annotated with rich text attributes. The dataset comprises 236,203 pairs of style prompt in natural text descriptions with five style factors and corresponding speech samples. Through iterative experimentation, we introduce a multi-stage prompt programming approach that effectively utilizes the GPT model for generating natural style descriptions in large volumes. 2) Furthermore, to address the need for generating audio with greater style diversity, we propose an efficient architecture called Salle. This architecture treats text controllable TTS as a language model task, utilizing audio codec codes as an intermediate representation to replace the conventional mel-spectrogram. Finally, we successfully demonstrate the ability of the proposed model by showing a comparable performance in the controllable TTS task. Audio samples are available on the demo page https://sall-e.github.io/. Shengpeng Ji, Jialong Zuo, Minghui Fang 0002, Ziyue Jiang 0004, Feiyang Chen 0001, Xinyu Duan, Baoxing Huai, Zhou Zhao 0001 |
ICASSP | 6 |
| 2024 | AuG-KD: Anchor-Based Mixup Generation for Out-of-Domain Knowledge DistillationabstractDue to privacy or patent concerns, a growing number of large models are released without granting access to their training data, making transferring their knowledge inefficient and problematic. In response, Data-Free Knowledge Distillation (DFKD) methods have emerged as direct solutions. However, simply adopting models derived from DFKD for real-world applications suffers significant performance degradation, due to the discrepancy between teachers' training data and real-world scenarios (student domain). The degradation stems from the portions of teachers' knowledge that are not applicable to the student domain. They are specific to the teacher domain and would undermine students' performance. Hence, selectively transferring teachers' appropriate knowledge becomes the primary challenge in DFKD. In this work, we propose a simple but effective method AuG-KD. It utilizes an uncertainty-guided and sample-specific anchor to align student-domain data with the teacher domain and leverages a generative method to progressively trade off the learning process between OOD knowledge distillation and domain-specific information learning via mixup learning. Extensive experiments in 3 datasets and 8 settings demonstrate the stability and superiority of our approach. Zheqi Lv, Shengyu Zhang 0001, Xinyu Duan, Fei Wu 0001, Kun Kuang 0001 |
ICLR | 5 |
| 2024 | Are Bert Family Good Instruction Followers? A Study on Their Potential And LimitationsabstractLanguage modeling at scale has proven very effective and brought unprecedented success to natural language models. Many typical representatives, especially decoder-only models, e.g., BLOOM and LLaMA, and encoder-decoder models, e.g., Flan-T5 and AlexaTM, have exhibited incredible instruction-following capabilities while keeping strong task completion ability. These large language models can achieve superior performance in various tasks and even yield emergent capabilities, e.g., reasoning and universal generalization. Though the above two paradigms are mainstream and well explored, the potential of the BERT family, which are encoder-only based models and have ever been one of the most representative pre-trained models, also deserves attention, at least should be discussed. In this work, we adopt XML-R to explore the effectiveness of the BERT family for instruction following and zero-shot learning. We first design a simple yet effective strategy to utilize the encoder-only models for generation tasks and then conduct multi-task instruction tuning. Experimental results demonstrate that our fine-tuned model, Instruct-XMLR, outperforms Bloomz on all evaluation tasks and achieves comparable performance with mT0 on most tasks. Surprisingly, Instruct-XMLR also possesses strong task and language generalization abilities, indicating that Instruct-XMLR can also serve as a good instruction follower and zero-shot learner. Besides, Instruct-XMLR can accelerate decoding due to its non-autoregressive generation manner, achieving around 3 times speedup compared with current autoregressive large language models. Although we also witnessed several limitations through our experiments, such as the performance decline in long-generation tasks and the shortcoming of length prediction, Instruct-XMLR can still become a good member of the family of current large language models. Yisheng Xiao, Juntao Li 0005, Zechen Sun, Zechang Li, Qingrong Xia, Xinyu Duan, Zhefeng Wang 0001, Min Zhang 0005 |
ICLR | 6 |
| 2024 | Poisoning for Debiasing: Fair Recognition via Eliminating Bias Uncovered in Data PoisoningabstractNeural networks often tend to rely on bias features that have strong but spurious correlations with the target labels for decision-making, leading to poor performance on data that does not adhere to these correlations. Early debiasing methods typically construct an unbiased optimization objective based on the labels of bias features. Recent work assumes that bias label is unavailable and usually trains two models: a biased model to deliberately learn bias features for exposing data bias, and a target model to eliminate bias captured by the bias model. In this paper, we first reveal that previous biased models fit target labels, which resulted in failing to expose data bias. To tackle this issue, we propose poisoner, which utilizes data poisoning to embed the biases learned by biased models into the poisoned training data, thereby encouraging the models to learn more biases. Specifically, we couple data poisoning and model training to continuously prompt the biased model to learn more bias. By utilizing the biased model, we can identify samples in the data that contradict these biased correlations. Subsequently, we amplify the influence of these samples in the training of the target model to prevent the model from learning such biased correlations. Experiments show the superior debiasing performance of our method. Yi Zhang 0101, Zhefeng Wang 0001, Rui Hu 0011, Xinyu Duan, Yi Zheng 0007, Baoxing Huai, Jiarun Han, Jitao Sang 0001 |
ACM Multimedia | 4 |
| 2023 | OpenSR: Open-Modality Speech Recognition via Maintaining Multi-Modality AlignmentabstractSpeech Recognition builds a bridge between the multimedia streaming (audio-only, visualonly or audio-visual) and the corresponding text transcription.However, when training the specific model of new domain, it often gets stuck in the lack of new-domain utterances, especially the labeled visual utterances.To break through this restriction, we attempt to achieve zero-shot modality transfer by maintaining the multi-modality alignment in phoneme space learned with unlabeled multimedia utterances in the high resource domain during the pretraining (Shi et al., 2022), and propose a training system Open-modality Speech Recognition (OpenSR) that enables the models trained on a single modality (e.g., audio-only) applicable to more modalities (e.g., visual-only and audio-visual).Furthermore, we employ a cluster-based prompt tuning strategy to handle the domain shift for the scenarios with only common words in the new domain utterances.We demonstrate that OpenSR enables modality transfer from one to any in three different settings (zero-, few-and fullshot), and achieves highly competitive zeroshot performance compared to the existing fewshot and full-shot lip-reading methods.To the best of our knowledge, OpenSR achieves the state-of-the-art performance of word error rate in LRS2 on audio-visual speech recognition and lip-reading with 2.7% and 25.0%, respectively.The code and demo are available at https://github.com/Exgc/OpenSR. Xize Cheng, Tao Jin 0004, Linjun Li, Xinyu Duan, Zhou Zhao 0001 |
ACL (1) | 5 |
| 2022 | Revisiting Pre-trained Language Models and their Evaluation for Arabic Natural Language ProcessingabstractAbbas Ghaddar, Yimeng Wu, Sunyam Bagga, Ahmad Rashid, Khalil Bibi, Mehdi Rezagholizadeh, Chao Xing, Yasheng Wang, Xinyu Duan, Zhefeng Wang, Baoxing Huai, Xin Jiang, Qun Liu, Phillippe Langlais. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Abbas Ghaddar, Yimeng Wu, Sunyam Bagga, Ahmad Rashid, Khalil Bibi, Mehdi Rezagholizadeh, Yasheng Wang, Xinyu Duan, Zhefeng Wang 0001, Baoxing Huai, Xin Jiang 0002, Qun Liu 0001, Philippe Langlais |
EMNLP | 9 |
| 2022 | Industrial Internet of Things-enabled monitoring and maintenance mechanism for fully mechanized mining equipment
Chun-Hsien Chen, Xiangang Cao, Ray Y. Zhong, Xinyu Duan |
Adv. Eng. Informatics | 5 |
| 2021 | VLAD-VSA: Cross-Domain Face Presentation Attack Detection with Vocabulary Separation and AdaptationabstractFor face presentation attack detection (PAD), most of the spoofing cues are subtle, local image patterns (e.g., local image distortion, 3D mask edge and cut photo edges). The representations of existing PAD works with simple global pooling method, however, lose the local feature discriminability. In this paper, the VLAD aggregation method is adopted to quantize local features with visual vocabulary locally partitioning the feature space, and hence preserve the local discriminability. We further propose the vocabulary separation and adaptation method to modify VLAD for cross-domain PAD task. The proposed vocabulary separation method divides vocabulary into domain-shared and domain-specific visual words to cope with the diversity of live and attack faces under the cross-domain scenario.The proposed vocabulary adaptation method imitates the maximization step of the k-means algorithm in the end-to-end training, which guarantees the visual words be close to the center of assigned local features and thus brings robust similarity measurement. We give illustrations and extensive experiments to demonstrate the effectiveness of VLAD with the proposed vocabulary separation and adaptation method on standard cross-domain PAD benchmarks. The codes are available at https://github.com/Liubinggunzu/VLAD-VSA. Zhou Zhao 0001, Weike Jin, Xinyu Duan, Zhen Lei 0001, Baoxing Huai, Yiling Wu, Xiaofei He 0001 |
ACM Multimedia | 4 |
| 2019 | Legal Summarization for Multi-role Debate Dialogue via Controversy Focus Mining and Multi-task LearningabstractMulti-role court debate is a critical component in a civil trial where parties from different camps (plaintiff, defendant, witness, judge, etc.) actively involved. Unlike other types of dialogue, court debate can be lengthy, and important information, with respect to the controversy focus(es), often hides within the redundant and colloquial dialogue data. Summarizing court debate can be a novel but significant task to assist judge to effectively make the legal decision for the target trial. In this work, we propose an innovative end-to-end model to address this problem. Unlike prior summarization efforts, the proposed model projects the multi-role debate into the controversy focus space, which enables high-quality essential utterance(s) extraction in terms of legal knowledge and judicial factors. An extensive set of experiments with a large civil trial dataset shows that the proposed model can provide more accurate and readable summarization against several alternatives in the multi-role court debate scene. Xinyu Duan, Xiaozhong Liu 0001, Ruocheng Wang, Changlong Sun, Fei Wu 0001 |
CIKM | 1 |
| 2018 | Multi-Label Community-Based Question Classification via Personalized Sequence Memory Network LearningabstractMulti-label community-based question classification is a challenging problem in Community-based Question Answering (CQA) services, arising in many real applications such as question navigation and expert finding. Most of the existing approaches consider the problem as content-based tag suggestion task, which suffers from the textual sparsity issue. Unlike the previous studies, we consider the problem of multi-label community-based question classification from the viewpoint of personalized sequence learning. We introduce the personalized sequence memory network that leverages not only the semantics of questions but also the personalized information of askers to provide the sequence tag learning function to capture the high-order tag dependency. The experiment on real-world dataset shows the effectiveness of our method. Xinyu Duan, Shengyu Zhang 0001, Zhou Zhao 0001, Fei Wu 0001, Yueting Zhuang |
AAAI | 1 |
| 2018 | Temporality-enhanced knowledgememory network for factoid question answeringabstractQuestion answering is an important problem that aims to deliver specific answers to questions posed by humans in natural language. How to efficiently identify the exact answer with respect to a given question has become an active line of research. Previous approaches in factoid question answering tasks typically focus on modeling the semantic relevance or syntactic relationship between a given question and its corresponding answer. Most of these models suffer when a question contains very little content that is indicative of the answer. In this paper, we devise an architecture named the temporality-enhanced knowledge memory network (TE-KMN) and apply the model to a factoid question answering dataset from a trivia competition called quiz bowl. Unlike most of the existing approaches, our model encodes not only the content of questions and answers, but also the temporal cues in a sequence of ordered sentences which gradually remark the answer. Moreover, our model collaboratively uses external knowledge for a better understanding of a given question. The experimental results demonstrate that our method achieves better performance than several state-of-the-art methods. Xinyu Duan, Siliang Tang, Shengyu Zhang 0001, Yin Zhang 0006, Zhou Zhao 0001, Jianru Xue, Yueting Zhuang, Fei Wu 0001 |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2017 | An interpolation algorithm fitted for dynamic 3D face reconstruction
Dan Sui, Deheng Hou, Xinyu Duan |
Multim. Tools Appl. | 3 |
| 2017 | Temporal Interaction and Causal Influence in Community-Based Question AnsweringabstractDuring the last decade, community-based question answering (CQA) sites have accumulated a vast amount of questions and their crowdsourced answers over time. How to efficiently identify the quality of answers that are relevant to a given question has become an active line of research in CQA. The major challenge of CQA is the accurate selection of high-quality answers w.r.t given questions. Previous approaches tend to model the semantic matching between individual pair of one question and its corresponding answer (how fitting an answer is to a posted question). However, these works ignore the temporal interactions between answers (how previous answers influence the late posted answers). For example, a rational user likely adapts others' opinions, revises his inclinations, and posts a more appropriate answer after understanding the given question and previously posted answers. As a result, this paper devises an architecture named Temporal Interaction and Causal Influence LSTM (TC-LSTM) to effectively leverage not only the causal influence between question-answer (how appropriate an answer is for a given question) but also the temporal interactions between answers-answer (how a high-quality answer gradually forms). In particular, long short-term memory (LSTM) is used to capture the explicit question-answer influence and the implicit answers-answer interactions. Experiments are conducted on SemEval 2015 CQA dataset for answer classification task and Baidu Zhidao Dataset for answer ranking task. The experimental results show the advantage of our model comparing with other state-of-the-art methods. Fei Wu 0001, Xinyu Duan, Jun Xiao 0001, Zhou Zhao 0001, Siliang Tang, Yin Zhang 0006, Yueting Zhuang |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2016 | Community-Based Question Answering via Heterogeneous Social Network LearningabstractCommunity-based question answering (cQA) sites have accumulated vast amount of questions and corresponding crowdsourced answers over time. How to efficiently share the underlying information and knowledge from reliable (usually highly-reputable) answerers has become an increasingly popular research topic. A major challenge in cQA tasks is the accurate matching of high-quality answers w.r.t given questions. Many of traditional approaches likely recommend corresponding answers merely depending on the content similarity between questions and answers, therefore suffer from the sparsity bottleneck of cQA data. In this paper, we propose a novel framework which encodes not only the contents of question-answer(Q-A) but also the social interaction cues in the community to boost the cQA tasks. More specifically, our framework collaboratively utilizes the rich interaction among questions, answers and answerers to learn the relative quality rank of different answers w.r.t a same question. Moreover, the information in heterogeneous social networks is comprehensively employed to enhance the quality of question-answering (QA) matching by our deep random walk learning framework. Extensive experiments on a large-scale dataset from a real world cQA site show that leveraging the heterogeneous social information indeed achieves better performance than other state-of-the-art cQA methods. Hanyin Fang, Fei Wu 0001, Zhou Zhao 0001, Xinyu Duan, Yueting Zhuang, Martin Ester |
AAAI | 4 |
| 2008 | Research of Three Dimension Collision Technique in VRML Interactive Simulation EnvironmentabstractPaper discussed collision detection techniques theoretically, including polygon and non-polygon expression of collision model, collision model's bounding box tree optimization, and query of collision. Then introduced VRML's collision sensing technique and its Collision node's syntax and application, mainly about speeding collision detection by using a simple model substitute. Finally virtual simulated an auto-door collision sensing application, trigged by Collision grouping node. Jiang Pin, Xinyu Duan |
CW | 2 |