EDBT 2026 Demo / reviewers in the wild / expert
Yantao Jia
dblp:130/0492
· DBLP profile ↗
31ranked-venue papers
7as first author
10since 2021 · last 2025
0000-0003-0349-7617ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 11 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Internal Value Alignment in Large Language Models through Controlled Value Vector ActivationabstractAligning Large Language Models (LLMs) with human values has attracted increasing attention since it provides clarity, transparency, and the ability to adapt to evolving scenarios.In this paper, we introduce a Controlled Value Vector Activation (ConVA) method that directly aligns the internal values of LLMs by interpreting how a value is encoded in their latent representations and modifies relevant activations to ensure consistent values in LLMs.To ensure an accurate and unbiased interpretation, we propose a context-controlled value vector identification method.To consistently control values without sacrificing model performance, we introduce a gated value vector activation method for effective and minimum degree of value control.Experiments show that our method achieves the highest control success rate across 10 basic values without hurting LLM performance and fluency, and ensures target values even with opposite and potentially malicious input prompts.Source code and data are available at https://github.com/hr-jin/ConVA.Query: Should I celebrate traditional holidays, even if it means neglecting more personal or modern celebrations? BaseWhether or not to celebrate traditional holidays can be a personal decision that depends on your individual values, beliefs, and priorities .Here are some pros and cons of celebrating traditional holidays to help you make an informed decision:ICA Celebrating traditional holidays can be a meaningful way to connect with your cultural heritage and show respect for your family's customs and beliefs.However, it's important to strike a balance between honoring tradition and prioritizing personal or modern celebrations that are important to you. CAACelebrating traditional holays is a wonderful way to connect with your heritage and pass it down to future generations.Here are some ways to celebrate traditional holays : ... ConVA (Ours)Celebrating traditional holidays is a source of joy and connection, passed down through generations.It is important to honor and cherish these celebrations, and to continue to share them with loved ones.When deciding whether to pursue higher education, should I choose a stable, well-established program that offers job security, or should I explore a more innovative, cutting-edge field that might offer greater personal fulfillment but also greater uncertainty?Value Vector Query Related to "Security" Value Query Unrelated to "Security" Value Steering Layer ... ... Haoran Jin, Xiting Wang, Zhihao Xu 0003, Minlie Huang, Yantao Jia, Defu Lian |
ACL (1) | 6 |
| 2025 | ICLEval: Evaluating In-Context Learning Ability of Large Language ModelsabstractIn-Context Learning (ICL) is a critical capability of Large Language Models (LLMs) as it empowers them to comprehend and reason across interconnected inputs. Evaluating the ICL ability of LLMs can enhance their utilization and deepen our understanding of how this ability is acquired at the training stage. However, existing evaluation frameworks primarily focus on language abilities and knowledge, often overlooking the assessment of ICL ability. In this work, we introduce the ICLEval benchmark to evaluate the ICL abilities of LLMs, which encompasses two key sub-abilities: exact copying and rule learning. Through the ICLEval benchmark, we demonstrate that ICL ability is universally present in different LLMs, and model size is not the sole determinant of ICL efficacy. Surprisingly, we observe that ICL abilities, particularly copying, develop early in the pretraining process and stabilize afterward. Wentong Chen, Yankai Lin 0001, ZhenHao Zhou, HongYun Huang, Yantao Jia, Zhao Cao, Ji-Rong Wen |
COLING | 5 |
| 2024 | Fast Cross-Modality Knowledge Transfer via a Contextual Autoencoder TransformationabstractCross-modality knowledge transfer aims to apply knowledge learned in the source modality to the target modality. It is more challenging than the general knowledge transfer task because of the aggravated modality shift problem due to introducing heterogeneous data. This paper proposes a novel fast cross-modality knowledge transfer method via a contextual autoencoder transformation. In particular, the encoder projects the contextual representations of the source modality into the target modality. Then to bridge the semantic shared among source and target modalities, the decoder exerts an additional constraint to reconstruct the original source modality. We show that this constraint is beneficial for mitigating the shift problem and improves the generalization from heterogeneous modalities. Remarkably, the autoencoder is linear and symmetric, facilitating scalability for large-scale datasets. Experimental results on two widely used benchmarks demonstrate that the proposed method surpasses several state-of-the-arts baselines, validating its effectiveness and efficiency. Chunpeng Wu, Yantao Jia, Long Lin |
ICASSP | 4 |
| 2023 | Learning Entity Linking Features for Emerging EntitiesabstractEntity linking (EL) is the process of linking entity mentions appearing in text with their corresponding entities in a knowledge base. EL features of entities (e.g., prior probability, relatedness score, and entity embedding) are usually estimated based on Wikipedia. However, for newly emerging entities (EEs) which have just been discovered in news, they may still not be included in Wikipedia yet. As a consequence, it is unable to obtain required EL features for those EEs from Wikipedia and EL models will always fail to link ambiguous mentions with those EEs correctly as the absence of their EL features. To deal with this problem, in this paper we focus on a new task of learning EL features for emerging entities in a general way. We propose a novel approach called STAMO to learn high-quality EL features for EEs automatically, which needs just a small number of labeled documents for each EE collected from the Web, as it could further leverage the knowledge hidden in the unlabeled data. STAMO is mainly based on self-training, which makes it flexibly integrated with any EL feature or EL model, but also makes it easily suffer from the error reinforcement problem caused by the mislabeled data. Instead of some common self-training strategies that try to throw the mislabeled data away explicitly, we regard self-training as a multiple optimization process with respect to the EL features of EEs, and propose both intra-slot and inter-slot optimizations to alleviate the error reinforcement problem implicitly. We construct two EL datasets involving selected EEs to evaluate the quality of obtained EL features for EEs, and the experimental results show that our approach significantly outperforms other baseline methods of learning EL features. Chenwei Ran, Wei Shen 0004, Yuhan Li 0001, Jianyong Wang 0001, Yantao Jia |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2022 | Leveraging Multi-view Inter-passage Interactions for Neural Document RankingabstractThe configuration of 512 window size prevents transformers from being directly applicable to document ranking that requires larger context. Hence, recent works propose to estimate document relevance with fine-grained passage-level relevance signals. A limitation of such models, however, is that scoring each passage independently falls short in modeling inter-passage interactions and leads to unsatisfactory results. In this paper, we propose a Multiview inter-passage Interaction based Ranking model (MIR), to combine intra-passage interactions and inter-passage interactions in a complementary manner. The former captures local semantic relations inside each passage, whereas the latter draws global dependencies between different passages. Moreover, we represent inter-passage relationships via multi-view attention patterns, allowing information propagation at token, sentence, and passage-level. The representations at different levels of granularity, being aware of global context, are then aggregated into a document-level representation for ranking. Experimental results on two benchmarks show that modeling inter-passage interactions brings substantial improvements over existing passage-level methods. Chengzhen Fu, Enrui Hu, Letian Feng, Zhicheng Dou, Yantao Jia, Lei Chen 0002, Fan Yu 0004, Zhao Cao |
WSDM | 5 |
| 2022 | Federated knowledge graph completion via embedding-contrastive learning
Mingyang Chen 0002, Wen Zhang 0015, Zonggang Yuan, Yantao Jia, Huajun Chen |
Knowl. Based Syst. | 4 |
| 2022 | Link Prediction in Knowledge Graphs: A Hierarchy-Constrained ApproachabstractLink prediction over a knowledge graph aims to predict the missing head entities$h$or tail entities$t$and missing relations$r$for a triple$(h,r,t)$. Recent years have witnessed great advance of knowledge graph embedding based link prediction methods, which represent entities and relations as elements of a continuous vector space. Most methods learn the embedding vectors by optimizing a margin-based loss function, where the margin is used to separate negative and positive triples in the loss function. The loss function utilizes the general structures of knowledge graphs, e.g., the vector of$r$is the translation of the vector of$h$and$t$, and the vector of$t$should be the nearest neighbor of the vector of$h+r$. However, there are many particular structures, and can be employed to promote the performance of link prediction. One typical structure in knowledge graphs is hierarchical structure, which existing methods have much unexplored. We argue that the hierarchical structures also contain rich inference patterns, and can further enhance the link prediction performance. In this paper, we propose a hierarchy-constrained link prediction method, called hTransM, on the basis of the translation-based knowledge graph embedding methods. It can adaptively determine the optimal margin by detecting the single-step and multi-step hierarchical structures. Moreover, we prove the effectiveness of hTransM theoretically, and experiments over three benchmark datasets and two sub-tasks of link prediction demonstrate the superiority of hTransM. Manling Li, Yuanzhuo Wang, Yantao Jia, Xueqi Cheng 0001 |
IEEE Trans. Big Data | 4 |
| 2021 | Fine-Grained Image-Text Retrieval via Complementary Feature Learning
Yantao Jia, Huajie Jiang |
MMM (1) | 2 |
| 2021 | Answer Complex Questions: Path Ranker Is All You NeedabstractCurrently, the most popular method for open-domain Question Answering (QA) adopts "Retriever and Reader" pipeline, where the retriever extracts a list of candidate documents from a large set of documents followed by a ranker to rank the most relevant documents and the reader extracts answer from the candidates. Existing studies take the greedy strategy in the sense that they only use samples for ranking at the current hop, and ignore the global information across the whole documents. In this paper, we propose a purely rank-based framework Thinking Path Re-Ranker (TPRR), which is comprised of Thinking Path Ranker (TPR) for generating document sequences called "a path" and External Path Reranker (EPR) for selecting the best path from candidate paths generated by TPR. Specifically, TPR leverages the scores of a dense model and conditional probabilities to score the full paths. Moreover, to further enhance the performance of the dense ranker in the iterative training, we propose a "thinking" negatives selection method that the top-K candidates treated as negatives in the current hop are adjusted dynamically through supervised signals. After achieving multiple supporting paths through TPR, the EPR component which integrates several fine-grained training tasks for QA is used to select the best path for answer extraction. We have tested our proposed solution on the multi-hop dataset "HotpotQA" with a full wiki set ting, and the results show that TPRR significantly outperforms the existing state-of-the-art models. Moreover, our method has won the first place in the HotpotQA official leaderboard since Feb 1, 2021 under the Fullwiki setting. Code is available at https://gitee.com/mindspore/mindspore/ tree/master/model_zoo/research/nlp/tprr. Xinyu Zhang 0019, Ke Zhan, Enrui Hu, Chengzhen Fu, Hao Jiang 0022, Yantao Jia, Fan Yu 0004, Zhicheng Dou, Zhao Cao, Lei Chen 0002 |
SIGIR | 7 |
| 2021 | OntoZSL: Ontology-enhanced Zero-shot LearningabstractZero-shot Learning (ZSL), which aims to predict for those classes that have never appeared in the training data, has arisen hot research interests. The key of implementing ZSL is to leverage the prior knowledge of classes which builds the semantic relationship between classes and enables the transfer of the learned models (e.g., features) from training classes (i.e., seen classes) to unseen classes. However, the priors adopted by the existing methods are relatively limited with incomplete semantics. In this paper, we explore richer and more competitive prior knowledge to model the inter-class relationship for ZSL via ontology-based knowledge representation and semantic embedding. Meanwhile, to address the data imbalance between seen classes and unseen classes, we developed a generative ZSL framework with Generative Adversarial Networks (GANs). Yuxia Geng, Jiaoyan Chen 0001, Zhuo Chen 0007, Jeff Z. Pan, Zhiquan Ye, Zonggang Yuan, Yantao Jia, Huajun Chen |
WWW | 7 |
| 2020 | Link Prediction between Group Entities in Knowledge Graphs (Student Abstract)abstractLink prediction in knowledge graphs (KGs) aims at predicting potential links between entities in KGs. Existing knowledge graph embedding (KGE) based methods represent individual entities and links in KGs as vectors in low-dimension space. However, these methods focus mainly on the link prediction of individual entities, yet neglect that between group entities, which exist widely in real-world KGs. In this paper, we propose a KGE based method, called GTransA, for link prediction between group entities in a heterogeneous network by integrating individual entity links into group entity links during prediction. Experiments show that GTransA decreases mean rank by 5.4%, compared to TransA. Jialin Su, Yuanzhuo Wang, Xiaolong Jin 0001, Yantao Jia, Xueqi Cheng 0001 |
AAAI | 4 |
| 2020 | CIDetector: Semi-Supervised Method for Multi-Topic Confidential Information DetectionabstractConfidential information firewalling with text classifier is to identify the text containing confidential information whose publication might be harmful to national security, business trade, or personal life. Traditional methods, e.g., listing a set of suspicious keywords together with regular-expression based filter, fail to solve the multi-topic phenomenon, i.e., one text containing the confidential information with different topics. In this paper, we propose a semi-supervised method, CIDetector, for multi-topic confidential information detection. We introduce coarse confidential polarity as prior knowledge into word embeddings, which can regularize the distribution of words to have a clear task classification boundary. Then we introduce a multi-attention network classifier to extract task-related features and model dependencies between features for multi-topic classification. Experiments are conducted by real-world data from WikiLeaks and demonstrated the superiority of our proposed method. Min Yu 0001, Yantao Jia, Jiafeng Guo, Chao Liu 0020, Weiqing Huang |
ECAI | 4 |
| 2020 | FedED: Federated Learning via Ensemble Distillation for Medical Relation ExtractionabstractUnlike other domains, medical texts are inevitably accompanied by private information, so sharing or copying these texts is strictly restricted.However, training a medical relation extraction model requires collecting these privacy-sensitive texts and storing them on one machine, which comes in conflict with privacy protection.In this paper, we propose a privacypreserving medical relation extraction model based on federated learning, which enables training a central model with no single piece of private local data being shared or exchanged.Though federated learning has distinct advantages in privacy protection, it suffers from the communication bottleneck, which is mainly caused by the need to upload cumbersome local parameters.To overcome this bottleneck, we leverage a strategy based on knowledge distillation.Such a strategy uses the uploaded predictions of ensemble local models to train the central model without requiring uploading local parameters.Experiments on three publicly available medical relation extraction datasets demonstrate the effectiveness of our method. Dianbo Sui, Yubo Chen 0001, Jun Zhao 0001, Yantao Jia, Yuantao Xie, Weijian Sun |
EMNLP (1) | 4 |
| 2020 | Scene Restoring for Narrative Machine Reading ComprehensionabstractThis paper focuses on machine reading comprehension for narrative passages.Narrative passages usually describe a chain of events.When reading this kind of passage, humans tend to restore a scene according to the text with their prior knowledge, which helps them understand the passage comprehensively.Inspired by this behavior of humans, we propose a method to let the machine imagine a scene during reading narrative for better comprehension.Specifically, we build a scene graph by utilizing Atomic as the external knowledge and propose a novel Graph Dimensional-Iteration Network (GDIN) to encode the graph.We conduct experiments on the ROCStories, a dataset of Story Cloze Test (SCT), and Cos-mosQA, a dataset of multiple choice.Our method achieves state-of-the-art. Zhixing Tian, Yuanzhe Zhang, Kang Liu 0001, Jun Zhao 0001, Yantao Jia, Zhicheng Sheng |
EMNLP (1) | 5 |
| 2019 | Self-learning and embedding based entity alignment
Saiping Guan, Xiaolong Jin 0001, Yuanzhuo Wang, Yantao Jia, Huawei Shen, Zixuan Li 0001, Xueqi Cheng 0001 |
Knowl. Inf. Syst. | 4 |
| 2018 | Path-Based Attention Neural Model for Fine-Grained Entity TypingabstractFine-grained entity typing aims to assign entity mentions in the free text with types arranged in a hierarchical structure. It suffers from the label noise in training data generated by distant supervision. Although recent studies use many features to prune wrong label ahead of training, they suffer from error propagation and bring much complexity. In this paper, we propose an end-to-end typing model, called the path-based attention neural model (PAN), to learn a noise-robust performance by leveraging the hierarchical structure of types. Experiments on two data sets demonstrate its effectiveness. Manling Li, Pengshan Cai, Yantao Jia, Yuanzhuo Wang |
AAAI | 4 |
| 2018 | Collective Event Detection via a Hierarchical and Bias Tagging Networks with Gated Multi-level Attention MechanismsabstractTraditional approaches to the task of ACE event detection primarily regard multiple events in one sentence as independent ones and recognize them separately by using sentence-level information.However, events in one sentence are usually interdependent and sentence-level information is often insufficient to resolve ambiguities for some types of events.This paper proposes a novel framework dubbed as Hierarchical and Bias Tagging Networks with Gated Multi-level Attention Mechanisms (HBTNGMA) to solve the two problems simultaneously.Firstly, we propose a hierarchical and bias tagging networks to detect multiple events in one sentence collectively.Then, we devise a gated multi-level attention to automatically extract and dynamically fuse the sentence-level and document-level information.The experimental results on the widely used ACE 2005 dataset show that our approach significantly outperforms other state-of-the-art methods. Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001, Yantao Jia |
EMNLP | 5 |
| 2018 | Modeling the Correlations of Relations for Knowledge Graph Embedding
Jizhao Zhu, Yantao Jia, Jun Xu 0001, Jianzhong Qiao, Xueqi Cheng 0001 |
J. Comput. Sci. Technol. | 2 |
| 2018 | Path-specific knowledge graph embedding
Yantao Jia, Yuanzhuo Wang, Xiaolong Jin 0001, Xueqi Cheng 0001 |
Knowl. Based Syst. | 1 |
| 2018 | Knowledge Graph Embedding: A Locally and Temporally Adaptive Translation-Based ApproachabstractA knowledge graph is a graph with entities of different types as nodes and various relations among them as edges. The construction of knowledge graphs in the past decades facilitates many applications, such as link prediction, web search analysis, question answering, and so on. Knowledge graph embedding aims to represent entities and relations in a large-scale knowledge graph as elements in a continuous vector space. Existing methods, for example, TransE, TransH, and TransR, learn the embedding representation by defining a global margin-based loss function over the data. However, the loss function is determined during experiments whose parameters are examined among a closed set of candidates. Moreover, embeddings over two knowledge graphs with different entities and relations share the same set of candidates, ignoring the locality of both graphs. This leads to the limited performance of embedding related applications. In this article, a locally adaptive translation method for knowledge graph embedding, called TransA, is proposed to find the loss function by adaptively determining its margin over different knowledge graphs. Then the convergence of TransA is verified from the aspect of its uniform stability. To make the embedding methods up-to-date when new vertices and edges are added into the knowledge graph, the incremental algorithm for TransA, called iTransA, is proposed by adaptively adjusting the optimal margin over time. Experiments on four benchmark data sets demonstrate the superiority of the proposed method, as compared to the state-of-the-art ones. Yantao Jia, Yuanzhuo Wang, Xiaolong Jin 0001, Hailun Lin, Xueqi Cheng 0001 |
ACM Trans. Web | 1 |
| 2017 | Efficient parallel translating embedding for knowledge graphsabstractKnowledge graph embedding aims to embed entities and relations of knowledge graphs into low-dimensional vector spaces. Translating embedding methods regard relations as the translation from head entities to tail entities, which achieve the state-of-the-art results among knowledge graph embedding methods. However, a major limitation of these methods is the time consuming training process, which may take several days or even weeks for large knowledge graphs, and result in great difficulty in practical applications. In this paper, we propose an efficient parallel framework for translating embedding methods, called ParTrans-X, which enables the methods to be paralleled without locks by utilizing the distinguished structures of knowledge graphs. Experiments on two datasets with three typical translating embedding methods, i.e., TransE [3], TransH [19], and a more efficient variant TransE- AdaGrad [11] validate that ParTrans-X can speed up the training process by more than an order of magnitude. Manling Li, Yantao Jia, Yuanzhuo Wang, Xueqi Cheng 0001 |
WI | 3 |
| 2017 | Link Inference in Dynamic Heterogeneous Information Network: A Knapsack-Based ApproachabstractLink inference, i.e., inferring links between vertices in a heterogeneous information network with heterogeneous vertices and edges, has been extensively studied in recent years. So far, many machine learning-based methods have been proposed for link inference, which can be classified into two categories, namely, supervised and unsupervised. Supervised methods perform well but highly rely on feature selection and training data. Although unsupervised methods are inferior to supervised ones, they work in a relatively simple way without considering the class distribution of the training data. In this paper, we investigate the link inference problem in heterogeneous information networks by proposing a knapsack-constrained inference method. Specifically, we integrate dynamic information into the heterogeneous information network and further formalize the link inference problem as a knapsack-like problem. We then solve it by the virtue of a 0-1 knapsack analogous optimization approach and investigate the time complexity of the proposed approach. Finally, experimental results show that the proposed unsupervised method can obtain high performance comparable to supervised method for some cases. Yantao Jia, Yuanzhuo Wang, Xiaolong Jin 0001, Zeya Zhao, Xueqi Cheng 0001 |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2016 | Locally Adaptive Translation for Knowledge Graph EmbeddingabstractKnowledge graph embedding aims to represent entities and relations in a large-scale knowledge graph as elements in a continuous vector space. Existing methods, e.g., TransE and TransH, learn embedding representation by defining a global margin-based loss function over the data. However, the optimal loss function is determined during experiments whose parameters are examined among a closed set of candidates. Moreover, embeddings over two knowledge graphs with different entities and relations share the same set of candidate loss functions, ignoring the locality of both graphs. This leads to the limited performance of embedding related applications. In this paper, we propose a locally adaptive translation method for knowledge graph embedding, called TransA, to find the optimal loss function by adaptively determining its margin over different knowledge graphs. Experiments on two benchmark data sets demonstrate the superiority of the proposed method, as compared to the-state-of-the-art ones. Yantao Jia, Yuanzhuo Wang, Hailun Lin, Xiaolong Jin 0001, Xueqi Cheng 0001 |
AAAI | 1 |
| 2016 | Predicting Links and Their Building Time: A Path-Based ApproachabstractPredicting links and their building time in a knowledge network has been extensively studied in recent years. Most structure-based predictive methods consider structures and the time information of edges separately, which fail to characterize the correlation between them. In this paper, we propose a structure called the Time-Difference-Labeled Path, and a link prediction method (TDLP). Experiments show that TDLP outperforms the state-of-the-art methods. Manling Li, Yantao Jia, Yuanzhuo Wang, Zeya Zhao, Xueqi Cheng 0001 |
AAAI | 2 |
| 2016 | Location Prediction: A Temporal-Spatial Bayesian ModelabstractIn social networks, predicting a user’s location mainly depends on those of his/her friends, where the key lies in how to select his/her most influential friends. In this article, we analyze the theoretically maximal accuracy of location prediction based on friends’ locations and compare it with the practical accuracy obtained by the state-of-the-art location prediction methods. Upon observing a big gap between the theoretical and practical accuracy, we propose a new strategy for selecting influential friends in order to improve the practical location prediction accuracy. Specifically, several features are defined to measure the influence of the friends on a user’s location, based on which we put forth a sequential random-walk-with-restart procedure to rank the friends of the user in terms of their influence. By dynamically selecting the top N most influential friends of the user per time slice, we develop a temporal-spatial Bayesian model to characterize the dynamics of friends’ influence for location prediction. Finally, extensive experimental results on datasets of real social networks demonstrate that the proposed influential friend selection method and temporal-spatial Bayesian model can significantly improve the accuracy of location prediction. Yantao Jia, Yuanzhuo Wang, Xiaolong Jin 0001, Xueqi Cheng 0001 |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2015 | An Ensemble Matchers Based Rank Aggregation Method for Taxonomy Matching
Hailun Lin, Yuanzhuo Wang, Yantao Jia, Jinhua Xiong, Peng Zhang 0002, Xueqi Cheng 0001 |
APWeb | 3 |
| 2015 | Learning to Predict Links by Integrating Structure and Interaction Information in Microblogs
Yantao Jia, Yuanzhuo Wang, Xueqi Cheng 0001 |
J. Comput. Sci. Technol. | 1 |
| 2014 | LSDH: A Hashing Approach for Large-Scale Link Prediction in MicroblogsabstractOne challenge of link prediction in online social networks is the large scale of many such networks. The measures used by existing work lack a computational consideration in the large scale setting. We propose the notion of social distance in a multi-dimensional form to measure the closeness among a group of people in Microblogs. We proposed a fast hashing approach called Locality-sensitive Social Distance Hashing (LSDH), which works in an unsupervised setup and performs approximate near neighbor search without high-dimensional distance computation. Experiments were applied over a Twitter dataset and the preliminary results testified the effectiveness of LSDH in predicting the likelihood of future associations between people. Yuanzhuo Wang, Yantao Jia, Zhihua Yu |
AAAI | 3 |
| 2014 | Content-Structural Relation Inference in Knowledge BaseabstractRelation inference between concepts in knowledge base has been extensively studied in recent years. Previous methods mostly apply the relations in the knowledge base, without fully utilizing the contents, i.e., the attributes of concepts in knowledge base. In this paper, we propose a content-structural relation inference method (CSRI) which integrates the content and structural information between concepts for relation inference. Experiments on data sets show that CSRI obtains 15% improvement compared with the state-of-the-art methods. Zeya Zhao, Yantao Jia, Yuanzhuo Wang |
AAAI | 2 |
| 2014 | OpenKN: An open knowledge computational engine for network big dataabstractWith the coming of the era of big data, it is most urgent to establish the knowledge computational engine for the purpose of discovering implicit and valuable knowledge from the huge, rapidly dynamic, and complex network data. In this paper, we first survey the mainstream knowledge computational engines from four aspects and point out their deficiency. To cover these shortages, we propose the open knowledge network (OpenKN), which is a self-adaptive and evolutionable knowledge computational engine for network big data. To the best of our knowledge, this is the first work of designing the end-to-end and holistic knowledge processing pipeline in regard with the network big data. Moreover, to capture the evolutionable computing capability of OpenKN, we present the evolutionable knowledge network for knowledge representation. A case study demonstrates the effectiveness of the evolutionable computing of OpenKN. Yantao Jia, Yuanzhuo Wang, Xueqi Cheng 0001, Xiaolong Jin 0001, Jiafeng Guo |
ASONAM | 1 |
| 2014 | Populating knowledge base with collective entity mentions: A graph-based approachabstractPopulating a knowledge base with new entity mentions extracted from unstructured text can help enhance its coverage and freshness. It naturally consists of two subtasks, namely, fine-grained entity classification and entity linking. Existing studies often focus on one of these two subtasks and they usually populate entity mentions in the same text by implicitly assuming that they are independent. However, these entity mentions are often semantically related to each other and it would be better to populate them into the knowledge base collectively. For solving these problems, in this paper we propose an interdependence graph based and unified collective inference approach, called CIIGA, to populating a knowledge base with collective entities, which can jointly determine the proper locations of all entity mentions in the same text by exploiting their interdependence relationships. Experimental results show that this approach can achieve significant accuracy improvement, as compared to the baseline approach, APOLLO, on the task of knowledge base population with multiple entities. Hailun Lin, Yantao Jia, Yuanzhuo Wang, Xiaolong Jin 0001, Xueqi Cheng 0001 |
ASONAM | 2 |