EDBT 2026 Demo / reviewers in the wild / expert
Yubin Ge
dblp:216/4408
· DBLP profile ↗
29ranked-venue papers
5as first author
21since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 4 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SAMULE: Self-Learning Agents Enhanced by Multi-level ReflectionabstractDespite the rapid advancements in LLM agents, they still face the challenge of generating meaningful reflections due to inadequate error analysis and a reliance on rare successful trajectories, especially in complex tasks.In this work, we propose SAMULE, a new framework for self-learning agents powered by a retrospective language model that is trained based on Multi-Level Reflection Synthesis.It first synthesizes high-quality reflections across three complementary levels: Single-Trajectory Learning (micro-level) for detailed error correction; Intra-Task Learning (meso-level) to build error taxonomies across multiple trials of the same task, and Inter-Task Learning (macrolevel) to extract transferable insights based on same typed errors from diverse task failures.Then we fine-tune a language model serving as the retrospective model to generate reflections during inference.We further extend our framework to interactive settings through a foresightbased reflection mechanism, enabling agents to proactively reflect and adapt during user interactions by comparing predicted and actual responses.Extensive experiments on three challenging benchmarks-TravelPlanner, NAT-URAL PLAN, and Tau-bench-demonstrate that our approach significantly outperforms reflection-based baselines.Our results highlight the critical role of well-designed reflection synthesis and failure-centric learning in building self-improving LLM agents. Yubin Ge, Salvatore Romeo, Jason Cai, Monica Sunkara |
EMNLP | 1 |
| 2025 | Examining Alignment of Large Language Models through Representative Heuristics: the case of political stereotypesabstractExamining the alignment of large language models (LLMs) has become increasingly important, e.g., when LLMs fail to operate as intended. This study examines the alignment of LLMs with human values for the domain of politics. Prior research has shown that LLM-generated outputs can include political leanings and mimic the stances of political parties on various issues. However, the extent and conditions under which LLMs deviate from empirical positions are insufficiently examined. To address this gap, we analyze the factors that contribute to LLMs' deviations from empirical positions on political issues, aiming to quantify these deviations and identify the conditions that cause them.
Drawing on findings from cognitive science about representativeness heuristics, i.e., situations where humans lean on representative attributes of a target group in a way that leads to exaggerated beliefs, we scrutinize LLM responses through this heuristics' lens. We conduct experiments to determine how LLMs inflate predictions about political parties, which results in stereotyping. We find that while LLMs can mimic certain political parties' positions, they often exaggerate these positions more than human survey respondents do. Also, LLMs tend to overemphasize representativeness more than humans. This study highlights the susceptibility of LLMs to representativeness heuristics, suggesting a potential vulnerability of LLMs that facilitates political stereotyping. We also test prompt-based mitigation strategies, finding that strategies that can mitigate representative heuristics in humans are also effective in reducing the influence of representativeness on LLM-generated responses. Sullam Jeoung, Yubin Ge, Haohan Wang, Jana Diesner |
ICLR | 2 |
| 2025 | Ordinal Unsupervised Domain Adaptation With Recursively Conditional Gaussian Imposed Variational DisentanglementabstractThere has been a growing interest in unsupervised domain adaptation (UDA) to alleviate the data scalability issue, while the existing works usually focus on classifying independently discrete labels. However, in many tasks (e.g., medical diagnosis), the labels are discrete and successively distributed. The UDA for ordinal classification requires inducing non-trivial ordinal distribution prior to the latent space. Target for this, the partially ordered set (poset) is defined for constraining the latent vector. Instead of the typically i.i.d. Gaussian latent prior, in this work, a recursively conditional Gaussian (RCG) set is proposed for ordered constraint modeling, which admits a tractable joint distribution prior. Furthermore, we are able to control the density of content vectors that violate the poset constraint by a simple "three-sigma rule." We explicitly disentangle the cross-domain images into a shared ordinal prior induced ordinal content space and two separate source/target ordinal-unrelated spaces, and the self-training is worked on the shared space exclusively for ordinal-aware domain alignment. Extensive experiments on UDA medical diagnoses and facial age estimation demonstrate its effectiveness. Xiaofeng Liu 0001, Site Li, Yubin Ge, Pengyi Ye, Jane You, Jun Lu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Contribution-based imbalanced hybrid resampling ensemble
Fei Han 0001, Yubin Ge, Qing Liu 0010, Henry Han |
Pattern Recognit. | 4 |
| 2024 | Extractive Summarization via Fine-grained Semantic Tuple ExtractionabstractTraditional extractive summarization treats the task as sentence-level classification and requires a fixed number of sentences for extraction.However, this rigid constraint on the number of sentences to extract may hinder model generalization due to varied summary lengths across datasets.In this work, we leverage the interrelation between information extraction (IE) and text summarization, and introduce a fine-grained autoregressive method for extractive summarization through semantic tuple extraction.Specifically, we represent each sentence as a set of semantic tuples, where tuples are predicate-argument structures derived from conducting IE.Then we adopt a Transformerbased autoregressive model to extract the tuples corresponding to the target summary given a source document.In inference, a greedy approach is proposed to select source sentences to cover extracted tuples, eliminating the need for a fixed number.Our experiments on CNN/DM and NYT demonstrate the method's superiority over strong baselines.Through the zero-shot setting for testing the generalization of models to diverse summary lengths across datasets, we further show our method outperforms baselines, including ChatGPT. Yubin Ge, Sullam Jeoung, Jana Diesner |
INLG | 1 |
| 2024 | Nuanced Multi-class Detection of Machine-Generated Scientific Text
Yubin Ge, Xiaofeng Liu 0001 |
PACLIC | 2 |
| 2023 | StereoMap: Quantifying the Awareness of Human-like Stereotypes in Large Language ModelsabstractLarge Language Models (LLMs) have been observed to encode and perpetuate harmful associations present in the training data.We propose a theoretically grounded framework called STEREOMAP to gain insights into their perceptions of how demographic groups have been viewed by society.The framework is grounded in the Stereotype Content Model (SCM); a well-established theory from psychology.According to SCM, stereotypes are not all alike.Instead, the dimensions of Warmth and Competence serve as the factors that delineate the nature of stereotypes.Based on the SCM theory, STEREOMAP maps LLMs' perceptions of social groups (defined by sociodemographic features) using the dimensions of Warmth and Competence.Furthermore, the framework enables the investigation of keywords and verbalizations of reasoning of LLMs' judgments to uncover underlying factors influencing their perceptions.Our results show that LLMs exhibit a diverse range of perceptions towards these groups, characterized by mixed evaluations along the dimensions of Warmth and Competence.Furthermore, analyzing the reasonings of LLMs, our findings indicate that LLMs demonstrate an awareness of social disparities, often stating statistical data and research findings to support their reasoning.This study contributes to the understanding of how LLMs perceive and represent social groups, shedding light on their potential biases and the perpetuation of harmful associations. Sullam Jeoung, Yubin Ge, Jana Diesner |
EMNLP | 2 |
| 2023 | What should I Ask: A Knowledge-driven Approach for Follow-up Questions Generation in Conversational Surveys
Yubin Ge, Ziang Xiao, Jana Diesner, Heng Ji 0001, Karrie Karahalios, Hari Sundaram |
PACLIC | 1 |
| 2023 | Inducing semantic hierarchy structure in empirical risk minimization with optimal transport measures
Wanqing Xie, Yubin Ge, Site Li, Zhenhua Guo 0001, Xiaofeng Liu 0001 |
Neurocomputing | 2 |
| 2022 | A Label Dependence-Aware Sequence Generation Model for Multi-Level Implicit Discourse Relation RecognitionabstractImplicit discourse relation recognition (IDRR) is a challenging but crucial task in discourse analysis. Most existing methods train multiple models to predict multi-level labels independently, while ignoring the dependence between hierarchically structured labels. In this paper, we consider multi-level IDRR as a conditional label sequence generation task and propose a Label Dependence-aware Sequence Generation Model (LDSGM) for it. Specifically, we first design a label attentive encoder to learn the global representation of an input instance and its level-specific contexts, where the label dependence is integrated to obtain better label embeddings. Then, we employ a label sequence decoder to output the predicted labels in a top-down manner, where the predicted higher-level labels are directly used to guide the label prediction at the current level. We further develop a mutual learning enhanced training method to exploit the label dependence in a bottom-up direction, which is captured by an auxiliary decoder introduced during training. Experimental results on the PDTB dataset show that our model achieves the state-of-the-art performance on multi-level IDRR. We release our code at https://github.com/nlpersECJTU/LDSGM. Changxing Wu, Liuwen Cao, Yubin Ge, Yang Liu 0005, Min Zhang 0005, Jinsong Su |
AAAI | 3 |
| 2022 | AAN+: Generalized Average Attention Network for Accelerating Neural TransformerabstractTransformer benefits from the high parallelization of attention networks in fast training, but it still suffers from slow decoding partially due to the linear dependency O(m) of the decoder self-attention on previous target words at inference. In this paper, we propose a generalized average attention network (AAN+) aiming at speeding up decoding by reducing the dependency from O(m) to O(1). We find that the learned self-attention weights in the decoder follow some patterns which can be approximated via a dynamic structure. Based on this insight, we develop AAN+, extending our previously proposed average attention (Zhang et al., 2018a, AAN) to support more general position- and content-based attention patterns. AAN+ only requires to maintain a small constant number of hidden states during decoding, ensuring its O(1) dependency. We apply AAN+ as a drop-in replacement of the decoder selfattention and conduct experiments on machine translation (with diverse language pairs), table-to-text generation and document summarization. With masking tricks and dynamic programming, AAN+ enables Transformer to decode sentences around 20% faster without largely compromising in the training speed and the generation performance. Our results further reveal the importance of the localness (neighboring words) in AAN+ and its capability in modeling long-range dependency. Biao Zhang 0002, Deyi Xiong, Yubin Ge, Junfeng Yao, Jinsong Su |
J. Artif. Intell. Res. | 3 |
| 2022 | An AST Structure Enhanced Decoder for Code GenerationabstractCurrently, the most dominant neural code generation modelsare often equipped with a tree-structured LSTM decoder, which outputs a sequence of actions to construct an Abstract Syntax Tree (AST) via pre-order traversal. However, such a decoder has two obvious drawbacks. First, except for the parent action, other faraway and important history actions rarely contribute to the current decision. Second, it also neglects future actions, which may be crucial for the prediction of the current action. To deal with these issues, in this paper, we propose a novel AST structure enhanced decoder for code generation, which significantly extends the decoder with respect to the above two aspects. First, we introduce an AST information enhanced attention mechanism to fully exploit history actions, of which impacts are further distinguished according to their syntactic distances, action types and relative positions; Second, we jointly model the predictions of current action and its important future action via multi-task learning, where the learned hidden state of the latter can be further leveraged to improve the former. Experimental results on commonly-used datasets demonstrate the effectiveness of our proposed decoder.1 Linfeng Song, Yubin Ge, Fandong Meng, Junfeng Yao, Jinsong Su |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Deep Verifier Networks: Verification of Deep Discriminative Models with Deep Generative ModelsabstractAI Safety is a major concern in many deep learning applications such as autonomous driving. Given a trained deep learning model, an important natural problem is how to reliably verify the model's prediction. In this paper, we propose a novel framework --- deep verifier networks (DVN) to detect unreliable inputs or predictions of deep discriminative models, using separately trained deep generative models. Our proposed model is based on conditional variational auto-encoders with disentanglement constraints to separate the label information from the latent representation. We give both intuitive and theoretical justifications for the model. Our verifier network is trained independently with the prediction model, which eliminates the need of retraining the verifier network for a new model. We test the verifier network on both out-of-distribution detection and adversarial example detection problems, as well as anomaly detection problems in structured prediction tasks such as image caption generation. We achieve state-of-the-art results in all of these problems. Tong Che, Xiaofeng Liu 0001, Site Li, Yubin Ge, Ruixiang Zhang, Caiming Xiong, Yoshua Bengio |
AAAI | 4 |
| 2021 | Improving Tree-Structured Decoder Training for Code Generation via Mutual LearningabstractCode generation aims to automatically generate a piece of code given an input natural language utterance. Currently, among dominant models, it is treated as a sequence-to-tree task, where a decoder outputs a sequence of actions corresponding to the pre-order traversal of an Abstract Syntax Tree. However, such a decoder only exploits the pre-order traversal based preceding actions, which are insufficient to ensure correct action predictions. In this paper, we first throughly analyze the context modeling difference between neural code generation models with different traversals based decodings (preorder traversal vs breadth-first traversal), and then propose to introduce a mutual learning framework to jointly train these models. Under this framework, we continuously enhance both two models via mutual distillation, which involves synchronous executions of two one-to-one knowledge transfers at each training step. More specifically, we alternately choose one model as the student and the other as its teacher, and require the student to fit the training data and the action prediction distributions of its teacher. By doing so, both models can fully absorb the knowledge from each other and thus could be improved simultaneously. Experimental results and in-depth analysis on several benchmark datasets demonstrate the effectiveness of our approach. We release our code at https://github.com/DeepLearnXMU/CGML. Binbin Xie, Jinsong Su, Yubin Ge, Xiang Li 0104, Jianwei Cui 0002, Junfeng Yao, Bin Wang 0004 |
AAAI | 3 |
| 2021 | BACO: A Background Knowledge- and Content-Based Framework for Citing Sentence GenerationabstractYubin Ge, Ly Dinh, Xiaofeng Liu, Jinsong Su, Ziyao Lu, Ante Wang, Jana Diesner. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yubin Ge, Ly Dinh, Xiaofeng Liu 0001, Jinsong Su, Ziyao Lu, Ante Wang, Jana Diesner |
ACL/IJCNLP (1) | 1 |
| 2021 | Improving Graph-based Sentence Ordering with Iteratively Predicted Pairwise OrderingsabstractShaopeng Lai, Ante Wang, Fandong Meng, Jie Zhou, Yubin Ge, Jiali Zeng, Junfeng Yao, Degen Huang, Jinsong Su. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Shaopeng Lai, Ante Wang, Fandong Meng, Jie Zhou 0016, Yubin Ge, Jiali Zeng, Junfeng Yao, Degen Huang, Jinsong Su |
EMNLP (1) | 5 |
| 2021 | Embedding Semantic Hierarchy in Discrete Optimal Transport for Risk MinimizationabstractThe widely-used cross-entropy (CE) loss-based deep networks achieved significant progress w.r.t. the classification accuracy. However, the CE loss can essentially ignore the risk of misclassification which is usually measured by the distance between the prediction and label in a semantic hierarchical tree. In this paper, we propose to incorporate the risk-aware inter-class correlation in a discrete optimal transport (DOT) training framework by configuring its ground distance matrix. The ground distance matrix can be pre-defined following a priori of hierarchical semantic risk. Specifically, we define the tree induced error (TIE) on a hierarchical semantic tree and extend it to its increasing function from the optimization perspective. The semantic similarity in each level of a tree is integrated with the information gain. We achieve promising results on several large scale image classification tasks with a semantic tree structure in a plug and play manner. Yubin Ge, Site Li, Wanqing Xie, Jane You, Xiaofeng Liu 0001 |
ICASSP | 1 |
| 2021 | Recursively Conditional Gaussian for Ordinal Unsupervised Domain AdaptationabstractThe unsupervised domain adaptation (UDA) has been widely adopted to alleviate the data scalability issue, while the existing works usually focus on classifying independently discrete labels. However, in many tasks (e.g., medical diagnosis), the labels are discrete and successively distributed. The UDA for ordinal classification requires inducing non-trivial ordinal distribution prior to the latent space. Target for this, the partially ordered set (poset) is defined for constraining the latent vector Instead of the typically i.i.d. Gaussian latent prior, in this work, a recursively conditional Gaussian (RCG) set is adapted for ordered constraint modeling, which admits a tractable joint distribution prior Furthermore, we are able to control the density of content vector that violates the poset constraints by a simple "three-sigma rule". We explicitly disentangle the cross-domain images into a shared ordinal prior induced ordinal content space and two separate source/target ordinal-unrelated spaces, and the self-training is worked on the shared space exclusively for ordinal-aware domain alignment. Extensive experiments on UDA medical diagnoses and facial age estimation demonstrate its effectiveness. Xiaofeng Liu 0001, Site Li, Yubin Ge, Pengyi Ye, Jane You, Jun Lu 0002 |
ICCV | 3 |
| 2021 | Enhanced aspect-based sentiment analysis models with progressive self-supervised attention learning
Jinsong Su, Jialong Tang, Ziyao Lu, Yubin Ge, Linfeng Song, Deyi Xiong, Le Sun 0001, Jiebo Luo 0001 |
Artif. Intell. | 5 |
| 2021 | Multi-modal neural machine translation with deep semantic interactions
Jinsong Su, Jinchang Chen, Chulun Zhou, Yubin Ge, Qingqiang Wu 0001, Yongxuan Lai |
Inf. Sci. | 6 |
| 2021 | Domain Adaptive Meta-Learning for Dialogue State TrackingabstractDomain adaptation for low-resource dialogue state tracking (DST) is of significance due to the growing diversity of conversation scenarios. In this paper, we propose a novel domain adaptive model-agnostic meta-learning (DAMAML) framework. Under this framework, we equip the DST model with two domain adaptors and a unified parameter generator. The parameter generator takes a domain embedding as input to produce parameters of domain adaptors, which modulate domain-shared initial parameters to the subspace of each domain. In this way, we simultaneously model multiple individual meta-learners with each covering the distribution of one domain, allowing more efficient adaptation. Compared with the conventional MAML, this framework not only is able to seek domain-shared initial parameters that facilitate fast adaptation, but also has better capability to fit a diversified domain distribution. Experimental results and in-depth analysis demonstrate the effectiveness of the proposed framework. Jiali Zeng, Yongjing Yin, Yang Liu 0005, Yubin Ge, Jinsong Su |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | Enhancing Pointer Network for Sentence Ordering with Pairwise Ordering PredictionsabstractDominant sentence ordering models use a pointer network decoder to generate ordering sequences in a left-to-right fashion. However, such a decoder only exploits the noisy left-side encoded context, which is insufficient to ensure correct sentence ordering. To address this deficiency, we propose to enhance the pointer network decoder by using two pairwise ordering prediction modules: The FUTURE module predicts the relative orientations of other unordered sentences with respect to the candidate sentence, and the HISTORY module measures the local coherence between several (e.g., 2) previously ordered sentences and the candidate sentence, without the influence of noisy left-side context. Using the pointer mechanism, we then incorporate this dynamically generated information into the decoder as a supplement to the left-side context for better predictions. On several commonly-used datasets, our model significantly outperforms other baselines, achieving the state-of-the-art performance. Further analyses verify that pairwise ordering predictions indeed provide extra useful context as expected, leading to better sentence ordering. We also evaluate our sentence ordering models on a downstream task, multi-document summarization, and the summaries reordered by our model achieve the best coherence scores. Our code is available at https://github.com/DeepLearnXMU/Pairwise.git. Yongjing Yin, Fandong Meng, Jinsong Su, Yubin Ge, Linfeng Song, Jie Zhou 0016, Jiebo Luo 0001 |
AAAI | 4 |
| 2020 | Structural Information Preserving for Graph-to-Text GenerationabstractThe task of graph-to-text generation aims at producing sentences that preserve the meaning of input graphs.As a crucial defect, the current state-of-the-art models may mess up or even drop the core structural information of input graphs when generating outputs.We propose to tackle this problem by leveraging richer training signals that can guide our model for preserving input information.In particular, we introduce two types of autoencoding losses, each individually focusing on different aspects (a.k.a.views) of input graphs.The losses are then back-propagated to better calibrate our model via multi-task training.Experiments on two benchmarks for graph-to-text generation show the effectiveness of our approach over a state-of-the-art baseline.Our code is available at http://github.com/ Soistesimmer/AMR-multiview. Linfeng Song, Ante Wang, Jinsong Su, Yue Zhang 0004, Kun Xu 0005, Yubin Ge, Dong Yu 0001 |
ACL | 6 |
| 2020 | An Iterative Multi-Source Mutual Knowledge Transfer Framework for Machine Reading ComprehensionabstractThe lack of sufficient training data in many domains, poses a major challenge to the construction of domain-specific machine reading comprehension (MRC) models with satisfying performance. In this paper, we propose a novel iterative multi-source mutual knowledge transfer framework for MRC. As an extension of the conventional knowledge transfer with one-to-one correspondence, our framework focuses on the many-to-many mutual transfer, which involves synchronous executions of multiple many-to-one transfers in an iterative manner.Specifically, to update a target-domain MRC model, we first consider other domain-specific MRC models as individual teachers, and employ knowledge distillation to train a multi-domain MRC model, which is differentially required to fit the training data and match the outputs of these individual models according to their domain-level similarities to the target domain. After being initialized by the multi-domain MRC model, the target-domain MRC model is fine-tuned to match both its training data and the output of its previous best model simultaneously via knowledge distillation. Compared with previous approaches, our framework can continuously enhance all domain-specific MRC models by enabling each model to iteratively and differentially absorb the domain-shared knowledge from others. Experimental results and in-depth analyses on several benchmark datasets demonstrate the effectiveness of our framework. Xin Liu 0066, Kai Liu 0023, Xiang Li 0104, Jinsong Su, Yubin Ge, Bin Wang 0004, Jiebo Luo 0001 |
IJCAI | 5 |
| 2020 | Dynamic Context-guided Capsule Network for Multimodal Machine TranslationabstractMultimodal machine translation (MMT), which mainly focuses on enhancing text-only translation with visual features, has attracted considerable attention from both computer vision and natural language processing communities. Most current MMT models resort to attention mechanism, global context modeling or multimodal joint representation learning to utilize visual features. However, the attention mechanism lacks sufficient semantic interactions between modalities while the other two provide fixed visual context, which is unsuitable for modeling the observed variability when generating translation. To address the above issues, in this paper, we propose a novel Dynamic Context-guided Capsule Network (DCCN) for MMT. Specifically, at each timestep of decoding, we first employ the conventional source-target attention to produce a timestep-specific source-side context vector. Next, DCCN takes this vector as input and uses it to guide the iterative extraction of related visual features via a context-guided dynamic routing mechanism. Particularly, we represent the input image with global and regional visual features, we introduce two parallel DCCNs to model multimodal context vectors with visual features at different granularities. Finally, we obtain two multimodal context vectors, which are fused and incorporated into the decoder for the prediction of the target word. Experimental results on the Multi30K dataset of English-to-German and English-to-French translation demonstrate the superiority of DCCN. Our code is available on https://github.com/DeepLearnXMU/MM-DCCN. Fandong Meng, Jinsong Su, Yongjing Yin, Zhengyuan Yang, Yubin Ge, Jie Zhou 0016, Jiebo Luo 0001 |
ACM Multimedia | 6 |
| 2019 | Progressive Self-Supervised Attention Learning for Aspect-Level Sentiment AnalysisabstractIn aspect-level sentiment classification (ASC), it is prevalent to equip dominant neural models with attention mechanisms, for the sake of acquiring the importance of each context word on the given aspect. However, such a mechanism tends to excessively focus on a few frequent words with sentiment polarities, while ignoring infrequent ones. In this paper, we propose a progressive self-supervised attention learning approach for neural ASC models, which automatically mines useful attention supervision information from a training corpus to refine attention mechanisms. Specifically, we iteratively conduct sentiment predictions on all training instances. Particularly, at each iteration, the context word with the maximum attention weight is extracted as the one with active/misleading influence on the correct/incorrect prediction of every instance, and then the word itself is masked for subsequent iterations. Finally, we augment the conventional training objective with a regularization term, which enables ASC models to continue equally focusing on the extracted active context words while decreasing weights of those misleading ones. Experimental results on multiple datasets show that our proposed approach yields better attention mechanisms, leading to substantial improvements over the two state-of-the-art neural ASC models. Source code and trained models are available at https://github.com/DeepLearnXMU/PSSAttention. Jialong Tang, Ziyao Lu, Jinsong Su, Yubin Ge, Linfeng Song, Le Sun 0001, Jiebo Luo 0001 |
ACL (1) | 4 |
| 2019 | Iterative Dual Domain Adaptation for Neural Machine TranslationabstractJiali Zeng, Yang Liu, Jinsong Su, Yubing Ge, Yaojie Lu, Yongjing Yin, Jiebo Luo. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Jiali Zeng, Yang Liu 0005, Jinsong Su, Yubin Ge, Yaojie Lu 0001, Yongjing Yin, Jiebo Luo 0001 |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Neural Collective Entity Linking Based on Recurrent Random Walk Network LearningabstractBenefiting from the excellent ability of neural networks on learning semantic representations, existing studies for entity linking (EL) have resorted to neural networks to exploit both the local mention-to-entity compatibility and the global interdependence between different EL decisions for target entity disambiguation. However, most neural collective EL methods depend entirely upon neural networks to automatically model the semantic dependencies between different EL decisions, which lack of the guidance from external knowledge. In this paper, we propose a novel end-to-end neural network with recurrent random-walk layers for collective EL, which introduces external knowledge to model the semantic interdependence between different EL decisions. Specifically, we first establish a model based on local context features, and then stack random-walk layers to reinforce the evidence for related EL decisions into high-probability decisions, where the semantic interdependence between candidate entities is mainly induced from an external knowledge base. Finally, a semantic regularizer that preserves the collective EL decisions consistency is incorporated into the conventional objective function, so that the external knowledge base can be fully exploited in collective EL decisions. Experimental results and in-depth analysis on various datasets show that our model achieves better performance than other state-of-the-art models. Our code and data are released at https://github.com/DeepLearnXMU/RRWEL. Mengge Xue, Weiming Cai, Jinsong Su, Linfeng Song, Yubin Ge, Bin Wang 0004 |
IJCAI | 5 |
| 2018 | Towards Automatic Generation of Peer-Targeted Science Talk in Curiosity-Evoking Virtual AgentabstractCuriosity is a critical skill that spurs learning, but is often found to decline with age and schooling. Recent research has shown that peer interaction may serve a special role in inducing curiosity through increased uncertainty and conceptual conflicts, since peers have similar authority in knowledge. For a virtual agent to stimulate curiosity, it should be able to generate curiosity-eliciting verbal behaviors such as hypothesis verbalization and argumentation, in the manner that simulates peer-like cognitive and behavioral abilities. In this paper, we design and implement a virtual peer that can carry out key curiosity-eliciting science talk during a dialog-based multi-party board game. We propose a child-centered and data-driven approach to simulate the latent reasoning process of young children and age-appropriate language during open-ended game play. In particular, we use a combination of child knowledge-graph construction and child-child interaction driven modeling to generate game appropriate behaviors that are compatible with 9-14 year old children. Encouraging human evaluation of the generated behaviors and generalizability of the generation framework to other tasks opens up new directions in incorporating open-endedness and science talk in virtual agents that will make them truly play a peer role in learning. Bhargavi Paranjape, Yubin Ge, Jessica Hammer, Justine Cassell |
IVA | 2 |