Yue Kou

dblp:31/1140 · DBLP profile ↗
← Back
85ranked-venue papers in the field
7as first author
46since 2021 · last 2026
0000-0002-5307-4893ORCID · corroborated

Domains — venue-derived; a paper can count in several

Knowledge Engineering, Semantic Web & Information Systems · 38 (1 first)Database Systems & Data Management · 23 (4 first)Information Retrieval & Web Search · 20 (2 first)Data Mining & Knowledge Discovery · 2Big Data, Cloud & Distributed Data Systems · 1Other / Interdisciplinary · 1
YearPublicationVenuePosition
2026 Causal Cross-Domain Sequential Recommendation with Preference Evolution
Jiaxuan Ma, Yue Kou, Dong Li 0023, Derong Shen, Xiangmin Zhou, Tiezheng Nie
DASFAA (1)2
2026 Explainable Team Formation by Integrating Skill Evolution and High-Order Collaboration
Jiaming Pu, Yue Kou, Dong Li 0023, Derong Shen, Tiezheng Nie, Ge Yu 0001
DASFAA (2)2
2026 Workload-Aware DHB+Tree: A Dynamic B+-Tree for Persistent Memory
Jiarui Qi, Junbao Song, Derong Shen, Tiezheng Nie, Yue Kou
DASFAA (1)5
2026 FSEM: Few-Shot Entity Matching Using Multi-loss Adversarial Training with Multi-attention Masking
Mengfei Xiong, Huayan Ma, Derong Shen, Tiezheng Nie, Yue Kou
DASFAA (2)5
2026 Curious or Conservative: Dynamic Curiosity-aware Explainable Recommendation
abstract
Explainable recommendation has attracted great attention due to its capability of enhancing user trust and satisfaction. Users’ curiosities highly affect the recommendation accuracy and the effectiveness of explanations. Different target users have different levels of curiosities, while the curiosity of the same user changes dynamically. However, existing techniques cannot capture users’ dynamic curiosities from the historical user-item interactions for effective explainable recommendation. In this article, we propose a novel explainable recommendation approach for effective D ynamic C uriosity-aware E xplainable R ecommendation (DCER). Specifically, we first propose a novel multi-view representation learning to model the temporal user-item interactions. Then, we propose a new curiosity-enhanced recommendation to dynamically capture users’ curiosities, which improves the recommendation quality in a mutual promotion manner. Finally, we propose an adaptive rule-guided hybrid explanation generation strategy that enables more personalized explanations and well reflects the users’ dynamic psychological states behind the transactions. The experimental results demonstrate the high effectiveness of our proposed model.
Yue Kou, Dong Li 0023, Derong Shen, Xiangmin Zhou, Tiezheng Nie, Ge Yu 0001
Trans. Recomm. Syst.1
2025 MSAE-SQL: A Multi-layer Semantic-Aware Enhanced NL2SQL Model
Minxuan Li, Kexin Ding, Derong Shen, Tiezheng Nie, Yue Kou, Minghe Yu 0001
WISA5
2025 Dual-Space Relational Contrastive-Aware Neighbor-Hood Matching Method for Entity Alignment
Qiwen Tan, Tiezheng Nie, Derong Shen, Yue Kou
IEEE Big Data4
2025 LeadFairRec: LLM-enhanced Discriminative Counterfactual Debiasing for Two-sided Fairness in Recommendation
abstract
Fairness-aware recommendation has emerged as a pivotal research area in recent years. Current fairness studies primarily examine two independent dimensions: user-side fairness and item-side fairness. However, most approaches address each side's fairness in isolation while neglecting their complex interdependencies. In this paper, we propose an LLM-Enhanced DiscriminAtive Counterfactual Debiasing Model for Two-sided Fairness in Recommendation (LeadFairRec). Specifically, we first design a two-sided causal graph that jointly models provider-customer fairness interactions through their causal relationships. Then we propose a discriminative counterfactual debiasing method, which effectively removes spurious correlations while maintaining true user-item interactions. Finally, we propose an LLM-enhanced counterfactual inference method to derive noise-resistant user/item representations from interaction data, enhancing the robustness of causal debiasing. The experimental results demonstrate the high effectiveness of our proposed model. We provide our code at https://github.com/houyimin660/LeadFairRec.
Yue Kou, Derong Shen, Xiangmin Zhou, Dong Li 0023, Tiezheng Nie, Ge Yu 0001
CIKM2
2025 Experts2team: Task Relevance-Induced Team Formation by Combining Global Cohesion with Local Decoupling
Yue Kou, Yingxuan Du, Derong Shen, Xiangmin Zhou, Dong Li 0023, Tiezheng Nie, Ge Yu 0001
DASFAA (2)1
2025 Counterfactual Path Augmentation for Reinforcement Reasoning in Explainable Recommendation
Yue Kou, Eryu Jiang, Derong Shen, Xiangmin Zhou, Dong Li 0023, Tiezheng Nie, Ge Yu 0001
DASFAA (5)1
2025 Two-Stage Temporal Knowledge Graph Completion Based on Reinforcement Learning
Yong Wei 0002, Xinyi Dong, Jingyou Sun, LinLin Ding, Yue Kou
ECML/PKDD (6)6
2025 Multi-task self-supervised learning based fusion representation for Multi-view clustering
Tianlong Guo, Derong Shen, Yue Kou, Tiezheng Nie
Inf. Sci.3
2025 GARF+: self-supervised and interpretable data cleaning with sequence generative adversarial networks
Jinfeng Peng, Hanghai Cui, Derong Shen, Nan Tang 0001, Yue Kou, Tiezheng Nie, Hang Cui 0001, Ge Yu 0001
VLDB J.5
2024 A Hierarchical Structure Explanation Method for Complex Tables
Fangnuo Liu, Sainan Tong, Derong Shen, Tiezheng Nie, Yue Kou
WISA5
2024 DFCDR: Domain-Aware Feature Decoupling and Fusion for Cross-Domain Recommendation
Jinyue Wei, Yue Kou, Derong Shen, Tiezheng Nie, Dong Li 0023
WISA2
2024 High-Dimensional Nearest Neighbor Search-Based Blocking in Entity Resolution
Chenchen Sun, Derong Shen, Tiezheng Nie, Yue Kou
WISA5
2024 LE-NER: A Chinese NER Model Based on Lexical Enhancement
Dong Li 0023, Shumei Du, Baoyan Song, Zhicong Liu, Yue Kou
ADMA (5)6
2024 Enhancing Deep Entity Resolution with Integrated Blocker-Matcher Training: Balancing Consensus and Discrepancy
abstract
Deep entity resolution (ER) identifies matching entities across data sources using techniques based on deep learning. It involves two steps: a blocker for identifying the potential matches to generate the candidate pairs, and a matcher for accurately distinguishing the matches and non-matches among these candidate pairs. Recent deep ER approaches utilize pretrained language models (PLMs) to extract similarity features for blocking and matching, achieving state-of-the-art performance. However, they often fail to balance the consensus and discrepancy between the blocker and matcher, emphasizing the consensus while neglecting the discrepancy. This paper proposes MutualER, a deep entity resolution framework that integrates and jointly trains the blocker and matcher, balancing both the consensus and discrepancy between them. Specifically, we firstly introduce a lightweight PLM in siamese structure for the blocker and a heavier PLM in cross structure or an autoregressive large language model (LLM) for the matcher. Two optimization techniques named Mutual Sample Selection (MSS) and Similarity Knowledge Transferring (SKT) are designed to jointly train the blocker and matcher. MSS enables the blocker and matcher to mutually select the customized training samples for each other to maintain the discrepancy, while SKT allows them to share the similarity knowledge for improving their blocking and matching capabilities respectively to maintain the consensus. Extensive experiments on five datasets demonstrate that MutualER significantly outperforms existing PLM-based and LLM-based approaches, achieving leading performance in both effectiveness and efficiency.
Wenzhou Dou, Derong Shen, Xiangmin Zhou, Yue Kou, Tiezheng Nie, Hang Cui 0001, Ge Yu 0001
CIKM5
2024 GARF: A Self-supervised Data Cleaning System with SeqGAN
abstract
High-quality data is essential for data science and machine learning applications, but unfortunately, real-world data often contains significant amounts of errors, such as typos, missing values, and data inconsistencies. Despite all the efforts in cleaning data using either logical or learning-based methods, in practice, data cleaning still requires high human cost, for either manually providing data repairing rules or preparing labeled datasets for training machine learning models. In this paper, we introduce GARF, a novel data cleaning system based on sequence generative adversarial networks (SeqGAN). One key information GARF tries to learn is data repair rules. To automatically extracts data repair rules from dirty data, GARF employs a SeqGAN to capture the dependency relationships, and converts the information learned by machine to interpretable data repair rules for humans. Additionally, considering that both generated rules and data may not be fully trusted, GARF provides a co-cleaning process to iteratively update inaccurate rules and repair dirty data until there is no tuple violating rules. We have implemented and deployed GARF as an open-sourced system, and demonstrated its usability on data cleaning in real-world scenarios.
Jinfeng Peng, Hanghai Cui, Derong Shen, Yue Kou, Tiezheng Nie, Tianlong Guo
CIKM4
2024 Towards Long-Text Entity Resolution with Chain-of-Thought Knowledge Augmentation from Large Language Models
Jiakai Tang, Wenzhou Dou, Derong Shen, Tiezheng Nie, Yue Kou
DASFAA (5)5
2024 RLclean: An unsupervised integrated data cleaning framework based on deep reinforcement learning
Jinfeng Peng, Derong Shen, Tiezheng Nie, Yue Kou
Inf. Sci.4
2023 Exploiting Item Relationships with Dual-Channel Attention Networks for Session-Based Recommendation
Yue Kou, Derong Shen, Tiezheng Nie, Dong Li 0023
WISA2
2023 An Efficient Storage Optimization Scheme for Blockchain Based on Hash Slot
Jiquan Wang, Tiezheng Nie, Derong Shen, Yue Kou
WISA4
2023 A Blockchain Query Optimization Method Based on Hybrid Indexes
Derong Shen, Tiezheng Nie, Yue Kou
WISA4
2023 Rule-Enhanced Evolutional Dual Graph Convolutional Network for Temporal Knowledge Graph Link Prediction
Huichen Zhai, Xiaobo Cao, Derong Shen, Tiezheng Nie, Yue Kou
WISA6
2022 Bi-Directional Neighborhood-Aware Network for Entity Alignment in Knowledge Graphs
Jingwen Bai 0005, Tiezheng Nie, Derong Shen, Yue Kou, Ge Yu 0001
WISA4
2022 SAREM: Semi-supervised Active Heterogeneous Entity Matching Framework
Jinxiu Du, Tiezheng Nie, Wenzhou Dou, Derong Shen, Yue Kou
WISA5
2022 B-store, a General Block Storage and Retrieval System for Blockchain
Xiaofei Gao, Tiezheng Nie, Derong Shen, Yue Kou, Guangyu He
WISA4
2022 An Efficient Query Architecture for Permissioned Blockchain
Xiabin Huang, Derong Shen, Tiezheng Nie, Yue Kou, Guangyu He
WISA4
2022 Enabling Verifiable Single-Attribute Range Queries on Erasure-Coded Sharding-Based Blockchain Systems
Dongyang Pan, Derong Shen, Tiezheng Nie, Yue Kou, Guangyu He
WISA4
2022 Dual-level Hypergraph Representation Learning for Group Recommendation
Yue Kou, Derong Shen, Tiezheng Nie, Dong Li 0023
WISA2
2022 Multi-view Based Entity Frequency-Aware Graph Neural Network for Temporal Knowledge Graph Link Prediction
Derong Shen, Tiezheng Nie, Yue Kou
WISA4
2022 Sentiment-Aware Neural Recommendation with Opinion-Based Explanations
Lingyu Zhao, Yue Kou, Derong Shen, Tiezheng Nie, Dong Li 0023
WISA2
2022 MCQL: A Multi-node Consortium Blockchain Query Method Based on Node Dynamic Adjustment
Tiezheng Nie, Derong Shen, Yue Kou
WISA4
2022 Multi-task Generative Adversarial Network for Missing Mobility Data Imputation
abstract
Mobility data collected from location-based social networks are imperative for user movement behaviour analysis and marketing strategy customization. However, due to personal privacy and temporary failure of GPS devices, mobility data suffer from missing data issues. The missing mobility data hide beneficial information that can lead to distorted data analysis. To this end, we propose a multi-task generative adversarial network, termed as MDI-MG, to mitigate the negative impact of missing mobility data by imputing possible missing records. Specifically, in MDI-MG, we first introduce region-awareness modelling to fully capture sequential dependencies. Then, the generator is designed as a multi-task network, which unifies two highly pertinent tasks, including the primary task and the auxiliary missing POI region imputation task. The joint training on the two tasks enhances presentation capabilities and brings additional benefits. Besides, we adopt a discriminator to evaluate the generated sequences. The generator and the discriminator are optimized with a minimax two-player game. Experiments on two real-world datasets show that, MDI-MG achieves better performance in terms of both imputation accuracy and effectiveness, compared with state-of-the-art methods.
Meihui Shi, Derong Shen, Yue Kou, Tiezheng Nie, Ge Yu 0001
CIKM3
2022 Empowering Transformer with Hybrid Matching Knowledge for Entity Matching
Wenzhou Dou, Derong Shen, Tiezheng Nie, Yue Kou, Chenchen Sun, Hang Cui 0001, Ge Yu 0001
DASFAA (3)4
2022 A new shape-based clustering algorithm for time series
Derong Shen, Tiezheng Nie, Yue Kou
Inf. Sci.4
2022 Self-supervised and Interpretable Data Cleaning with Sequence Generative Adversarial Networks
abstract
We study the problem of self-supervised and interpretable data cleaning, which automatically extracts interpretable data repair rules from dirty data. In this paper, we propose a novel framework, namely Garf, based on sequence generative adversarial networks (SeqGAN). One key information Garf tries to capture is data repair rules (for example, if the city is "Dothan", then the county should be "Houston"). Garf employs a SeqGAN consisting of a generator G and a discriminator D that trains G to learn the dependency relationships ( e.g. , given a city value "Dothan" as input, the county can be determined as "Houston"). After training, the generator G can be used to generate data repair rules, but may contain both trusted and untrusted rules, especially when learning from dirty data. To mitigate this problem, Garf further updates the learned relationships with another discriminator D' to iteratively improve the quality of both rules and data. Garf takes advantages of both logical and learning-based methods, which allow cleaning dirty data with high interpretability and have no requirements for prior knowledge and training data. Extensive experiments on real-world and synthetic datasets demonstrate the effectiveness of Garf. Garf achieves new state-of-the-art data cleaning result with high accuracy, through learning from dirty datasets without human supervision.
Jinfeng Peng, Derong Shen, Nan Tang 0001, Tieying Liu, Yue Kou, Tiezheng Nie, Hang Cui 0001, Ge Yu 0001
Proc. VLDB Endow.5
2021 Entity Alignment of Knowledge Graph by Joint Graph Attention and Translation Representation
Shixian Jiang, Tiezheng Nie, Derong Shen, Yue Kou, Ge Yu 0001
WISA4
2021 A Method of MOBA Game Lineup Recommendation Based on NSGA-II
Kangwei Li, Zhaozhao Xu, Tiezheng Nie, Derong Shen, Yue Kou
WISA8
2021 Heterogeneous Embeddings for Relational Data Integration Tasks
Xuehui Li, Guangqi Wang, Derong Shen, Tiezheng Nie, Yue Kou
WISA5
2021 DualLink: Dual Domain Adaptation for User Identity Linkage Across Social Networks
Yue Kou, Guangqi Wang, Derong Shen, Tiezheng Nie
WISA2
2021 Explainable Recommendation via Neural Rating Regression and Fine-Grained Sentiment Perception
Ziyu Yin, Yue Kou, Guangqi Wang, Derong Shen, Tiezheng Nie
WISA2
2021 Missing POI Check-in Identification Using Generative Adversarial Networks
Meihui Shi, Derong Shen, Yue Kou, Tiezheng Nie, Ge Yu 0001
DASFAA (1)3
2021 MOBA Game Analysis System Based on Neural Networks
Kangwei Li, Xiaobo Cao, Tiezheng Nie, Yue Kou, Derong Shen
WISE (2)6
2021 A cluster-based oversampling algorithm combining SMOTE and k-means for imbalanced medical data
Zhaozhao Xu, Derong Shen, Tiezheng Nie, Yue Kou
Inf. Sci.4
2020 Link Prediction Based on Smooth Evolution of Network Embedding
Yue Kou, Derong Shen, Tiezheng Nie
WISA2
2020 Web Table Column Type Detection Using Deep Learning and Probability Graph Model
Derong Shen, Tiezheng Nie, Yue Kou
WISA4
2020 An Explainable Recommendation Method Based on Multi-timeslice Graph Embedding
Huiying Wang, Yue Kou, Derong Shen, Tiezheng Nie
WISA2
2020 An Approach for Progressive Set Similarity Join with GPU Accelerating
Lining Yu, Tiezheng Nie, Derong Shen, Yue Kou
WISA4
2020 Efficient Team Formation in Social Networks based on Constrained Pattern Graph
abstract
Finding a team that is both competent in performing the task and compatible in working together has been extensively studied. However, most methods for team formation tend to rely on a set of skills only. In order to solve this problem, we present an efficient team formation method based on Constrained Pattern Graph (called CPG). Unlike traditional methods, our method takes into account both structure constraints and communication constraints on team members, which can better meet the requirements of users. First, a CPG preprocessing method is proposed to normalize a CPG and represent it as a CoreCPG in order to establish the basis for efficient matching. Second, a Communication Cost Index (called CCI) is constructed to speed up the matching between a CPG and its corresponding social network. Third, a CCI-based node matching algorithm is proposed to minimize the total number of intermediate results. Moreover, a set of incremental maintenance strategies for the changes of social networks are proposed. We conduct experimental studies based on two real-world social networks. The experiments demonstrate the effectiveness and the efficiency of our proposed method in comparison with traditional methods.
Yue Kou, Derong Shen, Quinn Snell, Dong Li 0023, Tiezheng Nie, Ge Yu 0001, Shuai Ma 0001
ICDE1
2019 A Cross-Network User Identification Model Based on Two-Phase Expansion
Yue Kou, Shuo Feng 0004, Derong Shen, Tiezheng Nie
WISA1
2019 Link Prediction Based on Node Embedding and Personalized Time Interval in Temporal Multi-relational Network
Derong Shen, Yue Kou, Tiezheng Nie
WISA3
2017 User Identification across Social Networks Based on Global View Features
abstract
Nowadays, people prefer to take part in multiple social networks to enjoy different kinds of services. Consequently, a significant task is to identify users across networks. Most state-of-the-art works on this issue exploit user local structure features (e.g., friend, follow and followed). In this paper, we first proposes the notion of user global view features, which represent the location of users in the network. Then, we present an iterative two-stage algorithm (GAUI) using Global view features with user Attribute features to solve User Identification. In GAUI, we iteratively update pairwise similarity and predict new matching users. Certainly, we present a community based core anchor link filter strategy to reduce the computation cost, and present a stable matching based mapping strategy to improve the accuracy. At last, the experiments conducted on two real-world aligned networks demonstrate that our method has better performance on precision and recall.
Shuo Feng 0004, Derong Shen, Yue Kou, Tiezheng Nie, Ge Yu 0001
WISA4
2017 A Progressive Method for Detecting Duplication Entities Based on Bloom Filters
abstract
With the volume of data grows rapidly, the cost of detecting duplication entities has increased significantly in data cleaning. However, some real-time applications only need to identify as many duplicate entities as possible in a limited time, rather than all of them. The existing works adopt the sorting method to divide similar records into blocks, and arrange the processing order of blocks to detect duplicate entity progressively. However, this method only works well when the attributes of records are suitable for sorting. Therefore, this paper proposes a novel progressive de-duplicate method for records that can't be sorted by their attributes. The method distributes records into different blocks based on their features and generates a modified bloom filter index for each block. Then it uses the bloom filter to predict the probability of duplicate entities in this block, which determines the processing order of blocks to detect the duplicate entities more quickly. The comprehensive experiment shows that the number of duplicate detection by this algorithm in the finite time is far more efficient than other algorithms involved in the related works.
Yebing Luo, Tiezheng Nie, Derong Shen, Yue Kou, Ge Yu 0001
WISA4
2017 Determining Repairing Sequence of Inconsistencies in Content-Related Data
Yuefeng Du 0002, Derong Shen, Tiezheng Nie, Yue Kou, Ge Yu 0001
WISE (1)4
2017 Private Blocking Technique for Multi-party Privacy-Preserving Record Linkage
abstract
The process of matching and integrating records that relate to the same entity from one or more datasets is known as record linkage, and it has become an increasingly important subject in many application areas, including business, government and health system. The data from these areas often contain sensitive information. To prevent privacy breaches, ideally records should be linked in a private way such that no information other than the matching result is leaked in the process, and this technique is called privacy-preserving record linkage (PPRL). With the increasing data, scalability becomes the main challenge of PPRL, and many private blocking techniques have been developed for PPRL. They are aimed at reducing the number of record pairs to be compared in the matching process by removing obvious non-matching pairs without compromising privacy. However, most of them are designed for two databases and they vary widely in their ability to balance competing goals of accuracy, efficiency and security. In this paper, we propose a novel private blocking approach for PPRL based on dynamic k -anonymous blocking and Paillier cryptosystem which can be applied on two or multiple databases. In dynamic k -anonymous blocking, our approach dynamically generates blocks satisfying k -anonymity and more accurate values to represent the blocks with varying k . We also propose a novel similarity measure method which performs on the numerical attributes and combines with Paillier cryptosystem to measure the similarity of two or more blocks in security, which provides strong privacy guarantees that none information reveals even collusion. Experiments conducted on a public dataset of voter registration records validate that our approach is scalable to large databases and keeps a high quality of blocking. We compare our method with other techniques and demonstrate the increases in security and accuracy.
Shumin Han, Derong Shen, Tiezheng Nie, Yue Kou, Ge Yu 0001
Data Sci. Eng.4
2016 Scalable Private Blocking Technique for Privacy-Preserving Record Linkage
Shumin Han, Derong Shen, Tiezheng Nie, Yue Kou, Ge Yu 0001
APWeb (2)4
2016 A Graph Clustering Algorithm for Citation Networks
Tiezheng Nie, Derong Shen, Yue Kou, Ge Yu 0001
APWeb (2)4
2016 Anchor Link Prediction Using Topological Information in Social Networks
Shuo Feng 0004, Derong Shen, Yue Kou, Tiezheng Nie, Ge Yu 0001
WAIM (1)3
2015 AILabel: A Fast Interval Labeling Approach for Reachability Query on Very Large Graphs
Shuo Feng 0004, Ning Xie 0006, Derong Shen, Yue Kou, Ge Yu 0001
APWeb5
2015 An Efficient Approach of Overlapping Communities Search
Jing Shan, Derong Shen, Tiezheng Nie, Yue Kou, Ge Yu 0001
DASFAA (1)4
2015 GB-JER: A Graph-Based Model for Joint Entity Resolution
Chenchen Sun, Derong Shen, Yue Kou, Tiezheng Nie, Ge Yu 0001
DASFAA (1)3
2014 Discovering Condition-Combined Functional Dependency Rules
Yuefeng Du 0002, Derong Shen, Tiezheng Nie, Yue Kou, Ge Yu 0001
APWeb4
2014 Distributed Entity Resolution Based on Similarity Join for Large-Scale Data Clustering
Tiezheng Nie, Wang-Chien Lee, Derong Shen, Ge Yu 0001, Yue Kou
WAIM5
2013 Computing the Split Points for Learning Decision Tree in MapReduce
Mingdong Zhu, Derong Shen, Ge Yu 0001, Yue Kou, Tiezheng Nie
DASFAA (2)4
2012 Richly Semantical Keyword Searching over Relational Databases
abstract
With the development of keyword search over relational databases, how to improve the result quality is a popular problem. To solve it, existing work mainly are CN-based and graph-based. The CN-based approaches occupy little memory space and have high level of abstract. However, the defect of these approaches is not considering meaningful information in metadata of databases. In this paper, we propose a novel architecture. First, it provides richer semantics for keywords to generate more meaningful candidate networks. Second, it provides query templates to facilitate the query transforms of candidate networks, which contribute to generating more meaningful query results. Third, some properties of ranking are simply summarized. Finally, the experimental results demonstrate that the result quality is improved.
Jianzhao Zhai, Derong Shen, Yue Kou, Tiezheng Nie
WISA3
2012 A Multilayer Method of Schema Matching Based on Semantic and Functional Dependencies
abstract
Determing matching schemas enables queries on heterogeneous data space to be formulated and facilitates data integration. Current schema matching techniques most focus on mining mappings using elements' own information. This paper proposes to introduce semantic and functional dependencies into matching process to achieve multilayer schema matching results. It calculates semantic similarity with the help of Word Net and generates candidate mapping sets. By introducing functional dependency to formulize structural information, it can get structural similarities between element pairs. A probabilistic factor is considered to select mapping pairs. Through experimental evaluation on real data, the superiority of our method is verified.
Chenlu Zhao, Derong Shen, Yue Kou, Tiezheng Nie, Ge Yu 0001
WISA3
2012 An Entity Class Model Based Correlated Query Path Selection Method in Multiple Domains
Jing Shan, Derong Shen, Tiezheng Nie, Yue Kou, Ge Yu 0001
APWeb4
2012 The Equi-Join Processing and Optimization on Ring Architecture Key/Value Database
Xite Wang, Derong Shen, Tiezheng Nie, Yue Kou, Ge Yu 0001
APWeb4
2012 A Transparent Approach for Database Schema Evolution Using View Mechanism
Jianxin Xue, Derong Shen, Tiezheng Nie, Yue Kou, Ge Yu 0001
WAIM4
2012 An Adaptive Distributed Index for Similarity Queries in Metric Spaces
Mingdong Zhu, Derong Shen, Yue Kou, Tiezheng Nie, Ge Yu 0001
WAIM3
2011 A Bottom-up Approach of Web Data Extraction based on Entity Recognition and Integration
abstract
Nowadays, most popular methods for web data extraction (WDE) are top-down ones depending on structure. However, these techniques are not scalable enough when coming to complex pages. Consequently, we put forward a bottom-up approach for WDE based on entity recognition and integration to avoid over dependency to structure of web pages. The approach proposed focuses on primary text sequences labeling first and also gives consideration to repetitive patterns of them as well. We propose a Two-Level extraction model for entity recognition and repetitive pattern extraction algorithm for entity integration. Our approach can effectively reduce the attribute labeling mistakes. Also, we demonstrate our approach by scientifically experimental results. The conclusion is that our approach perform better than the traditional extraction techniques, especially on complex Web pages.
Derong Shen, Jing Shan, Tiezheng Nie, Yue Kou
WISA5
2011 An Entity Relation Extraction Model Based on Semantic Pattern Matching
abstract
This paper proposes a relation extraction model based on semantic pattern matching in Web environment. It consists of frequent pattern extraction, pattern clustering based on density, and pattern matching based on semantic similarity. First, based on the entities with known relations in a limited training set, we extract relation patterns containing these named entities from the web page. Then the relations between entities from the web page in specific areas can be extracted based on these relation patterns extracted. Experiments show the affectivity and the self-adaptive of our method on extracting relations between entities from dynamic web environment.
Tiezheng Nie, Derong Shen, Yue Kou, Ge Yu 0001, Dejun Yue
WISA3
2011 Layout Object Model for Extracting the Schema of Web Query Interfaces
Tiezheng Nie, Derong Shen, Ge Yu 0001, Yue Kou
APWeb4
2011 Personalized Web Search with User Geographic and Temporal Preferences
Tiezheng Nie, Derong Shen, Ge Yu 0001, Yue Kou
APWeb5
2011 A Self-adaptive Cross-Domain Query Approach on the Deep Web
Yingjun Li, Derong Shen, Tiezheng Nie, Ge Yu 0001, Jing Shan, Yue Kou
WAIM6
2011 Layered Graph Data Model for Data Management of DataSpace Support Platform
Derong Shen, Tiezheng Nie, Ge Yu 0001, Yue Kou
WAIM5
2010 An Effective and High-quality Query Relaxation Solution on the Deep Web
abstract
Because the amount of information contained on the Deep Web is much larger than the surface web, how to use it well has become a popular problem to research. When a query is sent to a deep web resource and the data sources return few results or even no result, a proper query relaxation solution should be adopted to get more satisfactory results to users. In this paper, such a query relaxation solution is presented. First, it solves the problem of relaxing attributes which contain multiple key words by value. That is, such attributes are not simply removed in the relaxation, but the query values of the attributes are modified. Second, when a data source returns many result pages, instead of getting all the pages, it evaluates the quality of the results in the current page to decide whether to send another query to fetch the next page. Thus, the number of query times is reduced. Finally, the experimental results demonstrate that both the result quality and the query efficiency are improved.
Jing Shan, Derong Shen, Tiezheng Nie, Yue Kou, Ge Yu 0001
APWeb4
2010 Potential Role Based Entity Matching for Dataspaces Search
Yue Kou, Derong Shen, Tiezheng Nie, Ge Yu 0001
WISE1
2008 An Effective Query Relaxation Solution for the Deep Web
Derong Shen, Yue Kou
APWeb3
2008 LG-ERM: An Entity-Level Ranking Mechanism for Deep Web Query
abstract
With the rapid growth of Web databases, it's necessary to extract and integrate large-scale data available in deep Web automatically. But current Web search engines conduct page-level ranking, which are becoming inadequate for entity-oriented vertical search. In this paper, we present an entity-level ranking mechanism called LG-ERM for deep Web query based on local scoring and global aggregation. Unlike traditional approaches, LG-ERM considers more rank influencing factors including the uncertainty of entity extraction, the style information of entities and the importance of Web sources, as well as the entity relationship. By combining local scoring and global aggregation in ranking, the query result can be more accurate and effective to meet users' needs. The experiments demonstrate the feasibility and effectiveness of the key techniques of LG-ERM.
Yue Kou, Derong Shen, Ge Yu 0001, Tiezheng Nie
WAIM1
2008 Subject-Oriented Classification Based on Scale Probing in the Deep Web
abstract
To access the large-scale data sources efficiently and automatically, it is necessary to classify these data sources into different domains and categories. In this paper, we propose a novel classification approach to classify data sources into detail domain subjects by query probing. In our approach, we train sample instances for each subject category and use them to probe the data scale of each source and category. And then we build a matrix to classify a data source into one or more subject categories and develop a decision algorithm based on probing iteration to rectify the classification result. Our experiments over real deep web sources show that our approach can achieve higher accuracy across a variety of data sources.
Tiezheng Nie, Derong Shen, Ge Yu 0001, Yue Kou
WAIM4
2008 Efficient Top-k Data Sources Ranking for Query on Deep Web
Derong Shen, Meifang Li, Ge Yu 0001, Yue Kou, Tiezheng Nie
WISE4
2006 An Effective Service Discovery Model for Highly Reliable Web Services Composition in a Specific Domain
Derong Shen, Ge Yu 0001, Tiezheng Nie, Yue Kou, Meifang Li
APWeb4