EDBT 2026 Demo / reviewers in the wild / expert
Derong Shen
dblp:42/3190
· DBLP profile ↗
107ranked-venue papers in the field
4as first author
56since 2021 · last 2026
0000-0003-0310-6372ORCID · corroborated
Domains — venue-derived; a paper can count in several
Knowledge Engineering, Semantic Web & Information Systems · 42Information Retrieval & Web Search · 31 (3 first)Database Systems & Data Management · 29 (1 first)Other / Interdisciplinary · 3Data Mining & Knowledge Discovery · 1Big Data, Cloud & Distributed Data Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Causal Cross-Domain Sequential Recommendation with Preference Evolution
Jiaxuan Ma, Yue Kou, Dong Li 0023, Derong Shen, Xiangmin Zhou, Tiezheng Nie |
DASFAA (1) | 4 |
| 2026 | Explainable Team Formation by Integrating Skill Evolution and High-Order Collaboration
Jiaming Pu, Yue Kou, Dong Li 0023, Derong Shen, Tiezheng Nie, Ge Yu 0001 |
DASFAA (2) | 4 |
| 2026 | Workload-Aware DHB+Tree: A Dynamic B+-Tree for Persistent Memory
Jiarui Qi, Junbao Song, Derong Shen, Tiezheng Nie, Yue Kou |
DASFAA (1) | 3 |
| 2026 | FSEM: Few-Shot Entity Matching Using Multi-loss Adversarial Training with Multi-attention Masking
Mengfei Xiong, Huayan Ma, Derong Shen, Tiezheng Nie, Yue Kou |
DASFAA (2) | 3 |
| 2026 | Curious or Conservative: Dynamic Curiosity-aware Explainable RecommendationabstractExplainable recommendation has attracted great attention due to its capability of enhancing user trust and satisfaction. Users’ curiosities highly affect the recommendation accuracy and the effectiveness of explanations. Different target users have different levels of curiosities, while the curiosity of the same user changes dynamically. However, existing techniques cannot capture users’ dynamic curiosities from the historical user-item interactions for effective explainable recommendation. In this article, we propose a novel explainable recommendation approach for effective D ynamic C uriosity-aware E xplainable R ecommendation (DCER). Specifically, we first propose a novel multi-view representation learning to model the temporal user-item interactions. Then, we propose a new curiosity-enhanced recommendation to dynamically capture users’ curiosities, which improves the recommendation quality in a mutual promotion manner. Finally, we propose an adaptive rule-guided hybrid explanation generation strategy that enables more personalized explanations and well reflects the users’ dynamic psychological states behind the transactions. The experimental results demonstrate the high effectiveness of our proposed model. Yue Kou, Dong Li 0023, Derong Shen, Xiangmin Zhou, Tiezheng Nie, Ge Yu 0001 |
Trans. Recomm. Syst. | 3 |
| 2025 | MSAE-SQL: A Multi-layer Semantic-Aware Enhanced NL2SQL Model
Minxuan Li, Kexin Ding, Derong Shen, Tiezheng Nie, Yue Kou, Minghe Yu 0001 |
WISA | 3 |
| 2025 | Dual-Space Relational Contrastive-Aware Neighbor-Hood Matching Method for Entity Alignment
Qiwen Tan, Tiezheng Nie, Derong Shen, Yue Kou |
IEEE Big Data | 3 |
| 2025 | LeadFairRec: LLM-enhanced Discriminative Counterfactual Debiasing for Two-sided Fairness in RecommendationabstractFairness-aware recommendation has emerged as a pivotal research area in recent years. Current fairness studies primarily examine two independent dimensions: user-side fairness and item-side fairness. However, most approaches address each side's fairness in isolation while neglecting their complex interdependencies. In this paper, we propose an LLM-Enhanced DiscriminAtive Counterfactual Debiasing Model for Two-sided Fairness in Recommendation (LeadFairRec). Specifically, we first design a two-sided causal graph that jointly models provider-customer fairness interactions through their causal relationships. Then we propose a discriminative counterfactual debiasing method, which effectively removes spurious correlations while maintaining true user-item interactions. Finally, we propose an LLM-enhanced counterfactual inference method to derive noise-resistant user/item representations from interaction data, enhancing the robustness of causal debiasing. The experimental results demonstrate the high effectiveness of our proposed model. We provide our code at https://github.com/houyimin660/LeadFairRec. Yue Kou, Derong Shen, Xiangmin Zhou, Dong Li 0023, Tiezheng Nie, Ge Yu 0001 |
CIKM | 3 |
| 2025 | Token-Fusion: A Sparse Expert Routing Method for Multi-task Data MatchingabstractMulti-task data matching-including entity matching, entity linking, and schema matching-is a fundamental task in data integration, yet remains challenging due to heterogeneous inputs and task-specific model designs. We propose Token-Fusion, a unified sparse expert method that integrates token-level dynamic expert routing, adaptive expert pool management, and a fusion strategy guided by both confidence and performance gain. Meanwhile, regularization losses are designed to encourage sparse and diverse expert activation for improved efficiency. Our extensive experimental evaluation on six public datasets demonstrates the effectiveness and efficiency of Token-Fusion in handling heterogeneous matching tasks, establishing it as a promising solution for unified and scalable multi-task data matching. Fangnuo Liu, Derong Shen |
CIKM | 2 |
| 2025 | Experts2team: Task Relevance-Induced Team Formation by Combining Global Cohesion with Local Decoupling
Yue Kou, Yingxuan Du, Derong Shen, Xiangmin Zhou, Dong Li 0023, Tiezheng Nie, Ge Yu 0001 |
DASFAA (2) | 3 |
| 2025 | Counterfactual Path Augmentation for Reinforcement Reasoning in Explainable Recommendation
Yue Kou, Eryu Jiang, Derong Shen, Xiangmin Zhou, Dong Li 0023, Tiezheng Nie, Ge Yu 0001 |
DASFAA (5) | 3 |
| 2025 | Multi-task self-supervised learning based fusion representation for Multi-view clustering
Tianlong Guo, Derong Shen, Yue Kou, Tiezheng Nie |
Inf. Sci. | 2 |
| 2025 | GARF+: self-supervised and interpretable data cleaning with sequence generative adversarial networks
Jinfeng Peng, Hanghai Cui, Derong Shen, Nan Tang 0001, Yue Kou, Tiezheng Nie, Hang Cui 0001, Ge Yu 0001 |
VLDB J. | 3 |
| 2024 | A Hierarchical Structure Explanation Method for Complex Tables
Fangnuo Liu, Sainan Tong, Derong Shen, Tiezheng Nie, Yue Kou |
WISA | 3 |
| 2024 | DFCDR: Domain-Aware Feature Decoupling and Fusion for Cross-Domain Recommendation
Jinyue Wei, Yue Kou, Derong Shen, Tiezheng Nie, Dong Li 0023 |
WISA | 3 |
| 2024 | High-Dimensional Nearest Neighbor Search-Based Blocking in Entity Resolution
Chenchen Sun, Derong Shen, Tiezheng Nie, Yue Kou |
WISA | 3 |
| 2024 | Enhancing Deep Entity Resolution with Integrated Blocker-Matcher Training: Balancing Consensus and DiscrepancyabstractDeep entity resolution (ER) identifies matching entities across data sources using techniques based on deep learning. It involves two steps: a blocker for identifying the potential matches to generate the candidate pairs, and a matcher for accurately distinguishing the matches and non-matches among these candidate pairs. Recent deep ER approaches utilize pretrained language models (PLMs) to extract similarity features for blocking and matching, achieving state-of-the-art performance. However, they often fail to balance the consensus and discrepancy between the blocker and matcher, emphasizing the consensus while neglecting the discrepancy. This paper proposes MutualER, a deep entity resolution framework that integrates and jointly trains the blocker and matcher, balancing both the consensus and discrepancy between them. Specifically, we firstly introduce a lightweight PLM in siamese structure for the blocker and a heavier PLM in cross structure or an autoregressive large language model (LLM) for the matcher. Two optimization techniques named Mutual Sample Selection (MSS) and Similarity Knowledge Transferring (SKT) are designed to jointly train the blocker and matcher. MSS enables the blocker and matcher to mutually select the customized training samples for each other to maintain the discrepancy, while SKT allows them to share the similarity knowledge for improving their blocking and matching capabilities respectively to maintain the consensus. Extensive experiments on five datasets demonstrate that MutualER significantly outperforms existing PLM-based and LLM-based approaches, achieving leading performance in both effectiveness and efficiency. Wenzhou Dou, Derong Shen, Xiangmin Zhou, Yue Kou, Tiezheng Nie, Hang Cui 0001, Ge Yu 0001 |
CIKM | 2 |
| 2024 | GARF: A Self-supervised Data Cleaning System with SeqGANabstractHigh-quality data is essential for data science and machine learning applications, but unfortunately, real-world data often contains significant amounts of errors, such as typos, missing values, and data inconsistencies. Despite all the efforts in cleaning data using either logical or learning-based methods, in practice, data cleaning still requires high human cost, for either manually providing data repairing rules or preparing labeled datasets for training machine learning models. In this paper, we introduce GARF, a novel data cleaning system based on sequence generative adversarial networks (SeqGAN). One key information GARF tries to learn is data repair rules. To automatically extracts data repair rules from dirty data, GARF employs a SeqGAN to capture the dependency relationships, and converts the information learned by machine to interpretable data repair rules for humans. Additionally, considering that both generated rules and data may not be fully trusted, GARF provides a co-cleaning process to iteratively update inaccurate rules and repair dirty data until there is no tuple violating rules. We have implemented and deployed GARF as an open-sourced system, and demonstrated its usability on data cleaning in real-world scenarios. Jinfeng Peng, Hanghai Cui, Derong Shen, Yue Kou, Tiezheng Nie, Tianlong Guo |
CIKM | 3 |
| 2024 | Towards Long-Text Entity Resolution with Chain-of-Thought Knowledge Augmentation from Large Language Models
Jiakai Tang, Wenzhou Dou, Derong Shen, Tiezheng Nie, Yue Kou |
DASFAA (5) | 3 |
| 2024 | Matching Feature Separation Network for Domain Adaptation in Entity MatchingabstractEntity matching (EM) determines whether two records from different data sources refer to the same real-world entity. It is a fundamental task in knowledge graph construction and data integration. Currently, deep learning (DL) based EM methods have achieved state-of-the-art (SOTA) results. However, apply-ing DL-based EM methods often costs a lot of human efforts to label the data. To address this challenge, we propose a new do-main adaptation (DA) framework for EM called Matching Fea-ture Separation Network (MFSN). We implement DA by sepa-rating private and common matching features. Briefly, MFSN first uses three encoders to explicitly model the private and common matching features in both the source and target do-mains. Then, it transfers the knowledge learned from the source common matching features to the target domain. We also pro-pose an enhanced variant called Feature Representation and Separation Enhanced MFSN (MFSN-FRSE). Compared with MFSN, it has superior feature representation and separation capabilities. We evaluate the effectiveness of MFSN and MFSN-FRSE on twelve DA in EM tasks. The results show that our framework is approximately 7% higher in F1 score on average than the previous SOTA methods. Then, we verify the effec-tiveness of each module in MFSN and MFSN-FRSE by ablation study. Finally, we explore the optimal strategy of each module in MFSN and MFSN-FRSE through detailed tests. Chenchen Sun, Yang Xu 0073, Derong Shen, Tiezheng Nie |
WWW | 3 |
| 2024 | Graph Neural Network-Based Short‑Term Load Forecasting with Temporal ConvolutionabstractAbstract An accurate short-term load forecasting plays an important role in modern power system’s operation and economic development. However, short-term load forecasting is affected by multiple factors, and due to the complexity of the relationships between factors, the graph structure in this task is unknown. On the other hand, existing methods do not fully aggregating data information through the inherent relationships between various factors. In this paper, we propose a short-term load forecasting framework based on graph neural networks and dilated 1D-CNN, called GLFN-TC. GLFN-TC uses the graph learning module to automatically learn the relationships between variables to solve problem with unknown graph structure. GLFN-TC effectively handles temporal and spatial dependencies through two modules. In temporal convolution module, GLFN-TC uses dilated 1D-CNN to extract temporal dependencies from historical data of each node. In densely connected residual convolution module, in order to ensure that data information is not lost, GLFN-TC uses the graph convolution of densely connected residual to make full use of the data information of each graph convolution layer. Finally, the predicted values are obtained through the load forecasting module. We conducted five studies to verify the outperformance of GLFN-TC. In short-term load forecasting, using MSE as an example, the experimental results of GLFN-TC decreased by 0.0396, 0.0137, 0.0358, 0.0213 and 0.0337 compared to the optimal baseline method on ISO-NE, AT, AP, SH and NCENT datasets, respectively. Results show that GLFN-TC can achieve higher prediction accuracy than the existing common methods. Chenchen Sun, Yan Ning, Derong Shen, Tiezheng Nie |
Data Sci. Eng. | 3 |
| 2024 | RLclean: An unsupervised integrated data cleaning framework based on deep reinforcement learning
Jinfeng Peng, Derong Shen, Tiezheng Nie, Yue Kou |
Inf. Sci. | 2 |
| 2023 | Exploiting Item Relationships with Dual-Channel Attention Networks for Session-Based Recommendation
Yue Kou, Derong Shen, Tiezheng Nie, Dong Li 0023 |
WISA | 3 |
| 2023 | Temporal Convolution and Multi-Attention Jointly Enhanced Electricity Load Forecasting
Chenchen Sun, Hongxin Guo, Derong Shen, Tiezheng Nie, Zhijiang Hou |
WISA | 3 |
| 2023 | An Efficient Storage Optimization Scheme for Blockchain Based on Hash Slot
Jiquan Wang, Tiezheng Nie, Derong Shen, Yue Kou |
WISA | 3 |
| 2023 | A Blockchain Query Optimization Method Based on Hybrid Indexes
Derong Shen, Tiezheng Nie, Yue Kou |
WISA | 2 |
| 2023 | Rule-Enhanced Evolutional Dual Graph Convolutional Network for Temporal Knowledge Graph Link Prediction
Huichen Zhai, Xiaobo Cao, Derong Shen, Tiezheng Nie, Yue Kou |
WISA | 4 |
| 2023 | Exploring the Design Space of Unsupervised Blocking with Pre-trained Language Models in Entity Resolution
Chenchen Sun, Yuyuan Jin, Yang Xu 0073, Derong Shen, Tiezheng Nie, Xite Wang |
ADMA (1) | 4 |
| 2023 | Enhancing Knowledge Graph Attention by Temporal Modeling for Entity Alignment with Sparse Seeds
Chenchen Sun, Yuyuan Jin, Derong Shen, Tiezheng Nie, Xite Wang, Yingyuan Xiao |
DASFAA (2) | 3 |
| 2022 | Bi-Directional Neighborhood-Aware Network for Entity Alignment in Knowledge Graphs
Jingwen Bai 0005, Tiezheng Nie, Derong Shen, Yue Kou, Ge Yu 0001 |
WISA | 3 |
| 2022 | SAREM: Semi-supervised Active Heterogeneous Entity Matching Framework
Jinxiu Du, Tiezheng Nie, Wenzhou Dou, Derong Shen, Yue Kou |
WISA | 4 |
| 2022 | B-store, a General Block Storage and Retrieval System for Blockchain
Xiaofei Gao, Tiezheng Nie, Derong Shen, Yue Kou, Guangyu He |
WISA | 3 |
| 2022 | Multi-party Privacy-Preserving Record Linkage Method Based on Trusted Execution Environment
Xuefei He, Haiping Wei, Shumin Han, Derong Shen |
WISA | 4 |
| 2022 | An Efficient Query Architecture for Permissioned Blockchain
Xiabin Huang, Derong Shen, Tiezheng Nie, Yue Kou, Guangyu He |
WISA | 2 |
| 2022 | Enabling Verifiable Single-Attribute Range Queries on Erasure-Coded Sharding-Based Blockchain Systems
Dongyang Pan, Derong Shen, Tiezheng Nie, Yue Kou, Guangyu He |
WISA | 2 |
| 2022 | Dual-level Hypergraph Representation Learning for Group Recommendation
Yue Kou, Derong Shen, Tiezheng Nie, Dong Li 0023 |
WISA | 3 |
| 2022 | Efficient Multi-party Privacy-Preserving Record Linkage Based on Blockchain
Haoshan Yao, Haiping Wei, Shumin Han, Derong Shen |
WISA | 4 |
| 2022 | Multi-view Based Entity Frequency-Aware Graph Neural Network for Temporal Knowledge Graph Link Prediction
Derong Shen, Tiezheng Nie, Yue Kou |
WISA | 2 |
| 2022 | Sentiment-Aware Neural Recommendation with Opinion-Based Explanations
Lingyu Zhao, Yue Kou, Derong Shen, Tiezheng Nie, Dong Li 0023 |
WISA | 3 |
| 2022 | MCQL: A Multi-node Consortium Blockchain Query Method Based on Node Dynamic Adjustment
Tiezheng Nie, Derong Shen, Yue Kou |
WISA | 3 |
| 2022 | Multi-task Generative Adversarial Network for Missing Mobility Data ImputationabstractMobility data collected from location-based social networks are imperative for user movement behaviour analysis and marketing strategy customization. However, due to personal privacy and temporary failure of GPS devices, mobility data suffer from missing data issues. The missing mobility data hide beneficial information that can lead to distorted data analysis. To this end, we propose a multi-task generative adversarial network, termed as MDI-MG, to mitigate the negative impact of missing mobility data by imputing possible missing records. Specifically, in MDI-MG, we first introduce region-awareness modelling to fully capture sequential dependencies. Then, the generator is designed as a multi-task network, which unifies two highly pertinent tasks, including the primary task and the auxiliary missing POI region imputation task. The joint training on the two tasks enhances presentation capabilities and brings additional benefits. Besides, we adopt a discriminator to evaluate the generated sequences. The generator and the discriminator are optimized with a minimax two-player game. Experiments on two real-world datasets show that, MDI-MG achieves better performance in terms of both imputation accuracy and effectiveness, compared with state-of-the-art methods. Meihui Shi, Derong Shen, Yue Kou, Tiezheng Nie, Ge Yu 0001 |
CIKM | 2 |
| 2022 | Empowering Transformer with Hybrid Matching Knowledge for Entity Matching
Wenzhou Dou, Derong Shen, Tiezheng Nie, Yue Kou, Chenchen Sun, Hang Cui 0001, Ge Yu 0001 |
DASFAA (3) | 2 |
| 2022 | Information Networks Based Multi-semantic Data Embedding for Entity Resolution
Chenchen Sun, Derong Shen, Tiezheng Nie |
DASFAA (3) | 2 |
| 2022 | A new symbolic representation method for time seriesabstractTime series symbolic representation methods have been a research hot issue. Among them, the most representative symbolic methods such as Symbol Aggregation Approximation(SAX) and Symbol Fourier Approximation(SFA) have been widely used in various scenarios. However, they all have some flaws in some way. For SAX, the converted data in the approximation phase displays a shrinkage distribution(standard deviation σ shrinkage) that does not meet the assumptions in the definition. For SFA, it only obtains global frequency domain information, which leads to poor recognition ability for reciprocating frequency conversion sequence. Simultaneously, the symbol distance defined has poor interpretability due to spanning two spaces. In this paper, we propose a novel symbolic method called Symbol Fractional Fourier Approximation(SFFA), which shows multivariate approximation capabilities by Fractional Fourier Transform (FrFT) and adds a new supervised strategy for symbol mapping based on the chi-square distribution. It not only effectively avoids the influence of shrinkage distribution, but also has a strong ability to distinguish special sequences. Moreover, the SFFA symbol distance is proved to satisfy the low boundary lemma. Furthermore, it can achieve the same effect as SFA when the appropriate parameters and strategies are selected. Finally, when combined with the Vector Space Model (VSM) to classify time series, a large number of experiments show that SFFA-VSM outperforms SFA on all open-source data sets. Derong Shen |
Inf. Sci. | 2 |
| 2022 | A new shape-based clustering algorithm for time series
Derong Shen, Tiezheng Nie, Yue Kou |
Inf. Sci. | 2 |
| 2022 | Self-supervised and Interpretable Data Cleaning with Sequence Generative Adversarial NetworksabstractWe study the problem of self-supervised and interpretable data cleaning, which automatically extracts interpretable data repair rules from dirty data. In this paper, we propose a novel framework, namely Garf, based on sequence generative adversarial networks (SeqGAN). One key information Garf tries to capture is data repair rules (for example, if the city is "Dothan", then the county should be "Houston"). Garf employs a SeqGAN consisting of a generator G and a discriminator D that trains G to learn the dependency relationships ( e.g. , given a city value "Dothan" as input, the county can be determined as "Houston"). After training, the generator G can be used to generate data repair rules, but may contain both trusted and untrusted rules, especially when learning from dirty data. To mitigate this problem, Garf further updates the learned relationships with another discriminator D' to iteratively improve the quality of both rules and data. Garf takes advantages of both logical and learning-based methods, which allow cleaning dirty data with high interpretability and have no requirements for prior knowledge and training data. Extensive experiments on real-world and synthetic datasets demonstrate the effectiveness of Garf. Garf achieves new state-of-the-art data cleaning result with high accuracy, through learning from dirty datasets without human supervision. Jinfeng Peng, Derong Shen, Nan Tang 0001, Tieying Liu, Yue Kou, Tiezheng Nie, Hang Cui 0001, Ge Yu 0001 |
Proc. VLDB Endow. | 2 |
| 2021 | Entity Alignment of Knowledge Graph by Joint Graph Attention and Translation Representation
Shixian Jiang, Tiezheng Nie, Derong Shen, Yue Kou, Ge Yu 0001 |
WISA | 3 |
| 2021 | A Method of MOBA Game Lineup Recommendation Based on NSGA-II
Kangwei Li, Zhaozhao Xu, Tiezheng Nie, Derong Shen, Yue Kou |
WISA | 7 |
| 2021 | Heterogeneous Embeddings for Relational Data Integration Tasks
Xuehui Li, Guangqi Wang, Derong Shen, Tiezheng Nie, Yue Kou |
WISA | 3 |
| 2021 | DualLink: Dual Domain Adaptation for User Identity Linkage Across Social Networks
Yue Kou, Guangqi Wang, Derong Shen, Tiezheng Nie |
WISA | 4 |
| 2021 | Explainable Recommendation via Neural Rating Regression and Fine-Grained Sentiment Perception
Ziyu Yin, Yue Kou, Guangqi Wang, Derong Shen, Tiezheng Nie |
WISA | 4 |
| 2021 | Missing POI Check-in Identification Using Generative Adversarial Networks
Meihui Shi, Derong Shen, Yue Kou, Tiezheng Nie, Ge Yu 0001 |
DASFAA (1) | 2 |
| 2021 | Entity Resolution with Hybrid Attention-Based Networks
Chenchen Sun, Derong Shen |
DASFAA (2) | 2 |
| 2021 | MOBA Game Analysis System Based on Neural Networks
Kangwei Li, Xiaobo Cao, Tiezheng Nie, Yue Kou, Derong Shen |
WISE (2) | 7 |
| 2021 | Scalable Multi-grained Cross-modal Similarity Query with InterpretabilityabstractAbstract Cross-modal similarity query has become a highlighted research topic for managing multimodal datasets such as images and texts. Existing researches generally focus on query accuracy by designing complex deep neural network models and hardly consider query efficiency and interpretability simultaneously, which are vital properties of cross-modal semantic query processing system on large-scale datasets. In this work, we investigate multi-grained common semantic embedding representations of images and texts and integrate interpretable query index into the deep neural network by developing a novel Multi-grained Cross-modal Query with Interpretability (MCQI) framework. The main contributions are as follows: (1) By integrating coarse-grained and fine-grained semantic learning models, a multi-grained cross-modal query processing architecture is proposed to ensure the adaptability and generality of query processing. (2) In order to capture the latent semantic relation between images and texts, the framework combines LSTM and attention mode, which enhances query accuracy for the cross-modal query and constructs the foundation for interpretable query processing. (3) Index structure and corresponding nearest neighbor query algorithm are proposed to boost the efficiency of interpretable queries. (4) A distributed query algorithm is proposed to improve the scalability of our framework. Comparing with state-of-the-art methods on widely used cross-modal datasets, the experimental results show the effectiveness of our MCQI approach. Mingdong Zhu, Derong Shen, Xianfang Wang |
Data Sci. Eng. | 2 |
| 2021 | A cluster-based oversampling algorithm combining SMOTE and k-means for imbalanced medical data
Zhaozhao Xu, Derong Shen, Tiezheng Nie, Yue Kou |
Inf. Sci. | 2 |
| 2020 | Link Prediction Based on Smooth Evolution of Network Embedding
Yue Kou, Derong Shen, Tiezheng Nie |
WISA | 3 |
| 2020 | Web Table Column Type Detection Using Deep Learning and Probability Graph Model
Derong Shen, Tiezheng Nie, Yue Kou |
WISA | 2 |
| 2020 | An Explainable Recommendation Method Based on Multi-timeslice Graph Embedding
Huiying Wang, Yue Kou, Derong Shen, Tiezheng Nie |
WISA | 3 |
| 2020 | An Approach for Progressive Set Similarity Join with GPU Accelerating
Lining Yu, Tiezheng Nie, Derong Shen, Yue Kou |
WISA | 3 |
| 2020 | Efficient Team Formation in Social Networks based on Constrained Pattern GraphabstractFinding a team that is both competent in performing the task and compatible in working together has been extensively studied. However, most methods for team formation tend to rely on a set of skills only. In order to solve this problem, we present an efficient team formation method based on Constrained Pattern Graph (called CPG). Unlike traditional methods, our method takes into account both structure constraints and communication constraints on team members, which can better meet the requirements of users. First, a CPG preprocessing method is proposed to normalize a CPG and represent it as a CoreCPG in order to establish the basis for efficient matching. Second, a Communication Cost Index (called CCI) is constructed to speed up the matching between a CPG and its corresponding social network. Third, a CCI-based node matching algorithm is proposed to minimize the total number of intermediate results. Moreover, a set of incremental maintenance strategies for the changes of social networks are proposed. We conduct experimental studies based on two real-world social networks. The experiments demonstrate the effectiveness and the efficiency of our proposed method in comparison with traditional methods. Yue Kou, Derong Shen, Quinn Snell, Dong Li 0023, Tiezheng Nie, Ge Yu 0001, Shuai Ma 0001 |
ICDE | 2 |
| 2019 | A Cross-Network User Identification Model Based on Two-Phase Expansion
Yue Kou, Shuo Feng 0004, Derong Shen, Tiezheng Nie |
WISA | 4 |
| 2019 | Link Prediction Based on Node Embedding and Personalized Time Interval in Temporal Multi-relational Network
Derong Shen, Yue Kou, Tiezheng Nie |
WISA | 2 |
| 2017 | User Identification across Social Networks Based on Global View FeaturesabstractNowadays, people prefer to take part in multiple social networks to enjoy different kinds of services. Consequently, a significant task is to identify users across networks. Most state-of-the-art works on this issue exploit user local structure features (e.g., friend, follow and followed). In this paper, we first proposes the notion of user global view features, which represent the location of users in the network. Then, we present an iterative two-stage algorithm (GAUI) using Global view features with user Attribute features to solve User Identification. In GAUI, we iteratively update pairwise similarity and predict new matching users. Certainly, we present a community based core anchor link filter strategy to reduce the computation cost, and present a stable matching based mapping strategy to improve the accuracy. At last, the experiments conducted on two real-world aligned networks demonstrate that our method has better performance on precision and recall. Shuo Feng 0004, Derong Shen, Yue Kou, Tiezheng Nie, Ge Yu 0001 |
WISA | 3 |
| 2017 | A Progressive Method for Detecting Duplication Entities Based on Bloom FiltersabstractWith the volume of data grows rapidly, the cost of detecting duplication entities has increased significantly in data cleaning. However, some real-time applications only need to identify as many duplicate entities as possible in a limited time, rather than all of them. The existing works adopt the sorting method to divide similar records into blocks, and arrange the processing order of blocks to detect duplicate entity progressively. However, this method only works well when the attributes of records are suitable for sorting. Therefore, this paper proposes a novel progressive de-duplicate method for records that can't be sorted by their attributes. The method distributes records into different blocks based on their features and generates a modified bloom filter index for each block. Then it uses the bloom filter to predict the probability of duplicate entities in this block, which determines the processing order of blocks to detect the duplicate entities more quickly. The comprehensive experiment shows that the number of duplicate detection by this algorithm in the finite time is far more efficient than other algorithms involved in the related works. Yebing Luo, Tiezheng Nie, Derong Shen, Yue Kou, Ge Yu 0001 |
WISA | 3 |
| 2017 | Determining Repairing Sequence of Inconsistencies in Content-Related Data
Yuefeng Du 0002, Derong Shen, Tiezheng Nie, Yue Kou, Ge Yu 0001 |
WISE (1) | 2 |
| 2017 | Private Blocking Technique for Multi-party Privacy-Preserving Record LinkageabstractThe process of matching and integrating records that relate to the same entity from one or more datasets is known as record linkage, and it has become an increasingly important subject in many application areas, including business, government and health system. The data from these areas often contain sensitive information. To prevent privacy breaches, ideally records should be linked in a private way such that no information other than the matching result is leaked in the process, and this technique is called privacy-preserving record linkage (PPRL). With the increasing data, scalability becomes the main challenge of PPRL, and many private blocking techniques have been developed for PPRL. They are aimed at reducing the number of record pairs to be compared in the matching process by removing obvious non-matching pairs without compromising privacy. However, most of them are designed for two databases and they vary widely in their ability to balance competing goals of accuracy, efficiency and security. In this paper, we propose a novel private blocking approach for PPRL based on dynamic k -anonymous blocking and Paillier cryptosystem which can be applied on two or multiple databases. In dynamic k -anonymous blocking, our approach dynamically generates blocks satisfying k -anonymity and more accurate values to represent the blocks with varying k . We also propose a novel similarity measure method which performs on the numerical attributes and combines with Paillier cryptosystem to measure the similarity of two or more blocks in security, which provides strong privacy guarantees that none information reveals even collusion. Experiments conducted on a public dataset of voter registration records validate that our approach is scalable to large databases and keeps a high quality of blocking. We compare our method with other techniques and demonstrate the increases in security and accuracy. Shumin Han, Derong Shen, Tiezheng Nie, Yue Kou, Ge Yu 0001 |
Data Sci. Eng. | 2 |
| 2016 | Scalable Private Blocking Technique for Privacy-Preserving Record Linkage
Shumin Han, Derong Shen, Tiezheng Nie, Yue Kou, Ge Yu 0001 |
APWeb (2) | 2 |
| 2016 | A Graph Clustering Algorithm for Citation Networks
Tiezheng Nie, Derong Shen, Yue Kou, Ge Yu 0001 |
APWeb (2) | 3 |
| 2016 | Anchor Link Prediction Using Topological Information in Social Networks
Shuo Feng 0004, Derong Shen, Yue Kou, Tiezheng Nie, Ge Yu 0001 |
WAIM (1) | 2 |
| 2016 | Uncertain top-k query processing in distributed environments
Xite Wang, Derong Shen, Ge Yu 0001 |
Distributed Parallel Databases | 2 |
| 2015 | AILabel: A Fast Interval Labeling Approach for Reachability Query on Very Large Graphs
Shuo Feng 0004, Ning Xie 0006, Derong Shen, Yue Kou, Ge Yu 0001 |
APWeb | 3 |
| 2015 | Hybrid-LSH for Spatio-Textual Similarity Queries
Mingdong Zhu, Derong Shen, Ling Liu 0001, Ge Yu 0001 |
APWeb | 2 |
| 2015 | An Efficient Approach of Overlapping Communities Search
Jing Shan, Derong Shen, Tiezheng Nie, Yue Kou, Ge Yu 0001 |
DASFAA (1) | 2 |
| 2015 | GB-JER: A Graph-Based Model for Joint Entity Resolution
Chenchen Sun, Derong Shen, Yue Kou, Tiezheng Nie, Ge Yu 0001 |
DASFAA (1) | 2 |
| 2014 | Discovering Condition-Combined Functional Dependency Rules
Yuefeng Du 0002, Derong Shen, Tiezheng Nie, Yue Kou, Ge Yu 0001 |
APWeb | 2 |
| 2014 | Distributed Entity Resolution Based on Similarity Join for Large-Scale Data Clustering
Tiezheng Nie, Wang-Chien Lee, Derong Shen, Ge Yu 0001, Yue Kou |
WAIM | 3 |
| 2013 | Computing the Split Points for Learning Decision Tree in MapReduce
Mingdong Zhu, Derong Shen, Ge Yu 0001, Yue Kou, Tiezheng Nie |
DASFAA (2) | 2 |
| 2012 | Richly Semantical Keyword Searching over Relational DatabasesabstractWith the development of keyword search over relational databases, how to improve the result quality is a popular problem. To solve it, existing work mainly are CN-based and graph-based. The CN-based approaches occupy little memory space and have high level of abstract. However, the defect of these approaches is not considering meaningful information in metadata of databases. In this paper, we propose a novel architecture. First, it provides richer semantics for keywords to generate more meaningful candidate networks. Second, it provides query templates to facilitate the query transforms of candidate networks, which contribute to generating more meaningful query results. Third, some properties of ranking are simply summarized. Finally, the experimental results demonstrate that the result quality is improved. Jianzhao Zhai, Derong Shen, Yue Kou, Tiezheng Nie |
WISA | 2 |
| 2012 | A Multilayer Method of Schema Matching Based on Semantic and Functional DependenciesabstractDeterming matching schemas enables queries on heterogeneous data space to be formulated and facilitates data integration. Current schema matching techniques most focus on mining mappings using elements' own information. This paper proposes to introduce semantic and functional dependencies into matching process to achieve multilayer schema matching results. It calculates semantic similarity with the help of Word Net and generates candidate mapping sets. By introducing functional dependency to formulize structural information, it can get structural similarities between element pairs. A probabilistic factor is considered to select mapping pairs. Through experimental evaluation on real data, the superiority of our method is verified. Chenlu Zhao, Derong Shen, Yue Kou, Tiezheng Nie, Ge Yu 0001 |
WISA | 2 |
| 2012 | An Entity Class Model Based Correlated Query Path Selection Method in Multiple Domains
Jing Shan, Derong Shen, Tiezheng Nie, Yue Kou, Ge Yu 0001 |
APWeb | 2 |
| 2012 | The Equi-Join Processing and Optimization on Ring Architecture Key/Value Database
Xite Wang, Derong Shen, Tiezheng Nie, Yue Kou, Ge Yu 0001 |
APWeb | 2 |
| 2012 | A Transparent Approach for Database Schema Evolution Using View Mechanism
Jianxin Xue, Derong Shen, Tiezheng Nie, Yue Kou, Ge Yu 0001 |
WAIM | 2 |
| 2012 | An Adaptive Distributed Index for Similarity Queries in Metric Spaces
Mingdong Zhu, Derong Shen, Yue Kou, Tiezheng Nie, Ge Yu 0001 |
WAIM | 2 |
| 2011 | A Bottom-up Approach of Web Data Extraction based on Entity Recognition and IntegrationabstractNowadays, most popular methods for web data extraction (WDE) are top-down ones depending on structure. However, these techniques are not scalable enough when coming to complex pages. Consequently, we put forward a bottom-up approach for WDE based on entity recognition and integration to avoid over dependency to structure of web pages. The approach proposed focuses on primary text sequences labeling first and also gives consideration to repetitive patterns of them as well. We propose a Two-Level extraction model for entity recognition and repetitive pattern extraction algorithm for entity integration. Our approach can effectively reduce the attribute labeling mistakes. Also, we demonstrate our approach by scientifically experimental results. The conclusion is that our approach perform better than the traditional extraction techniques, especially on complex Web pages. Derong Shen, Jing Shan, Tiezheng Nie, Yue Kou |
WISA | 2 |
| 2011 | An Entity Relation Extraction Model Based on Semantic Pattern MatchingabstractThis paper proposes a relation extraction model based on semantic pattern matching in Web environment. It consists of frequent pattern extraction, pattern clustering based on density, and pattern matching based on semantic similarity. First, based on the entities with known relations in a limited training set, we extract relation patterns containing these named entities from the web page. Then the relations between entities from the web page in specific areas can be extracted based on these relation patterns extracted. Experiments show the affectivity and the self-adaptive of our method on extracting relations between entities from dynamic web environment. Tiezheng Nie, Derong Shen, Yue Kou, Ge Yu 0001, Dejun Yue |
WISA | 2 |
| 2011 | Layout Object Model for Extracting the Schema of Web Query Interfaces
Tiezheng Nie, Derong Shen, Ge Yu 0001, Yue Kou |
APWeb | 2 |
| 2011 | Personalized Web Search with User Geographic and Temporal Preferences
Tiezheng Nie, Derong Shen, Ge Yu 0001, Yue Kou |
APWeb | 3 |
| 2011 | A Self-adaptive Cross-Domain Query Approach on the Deep Web
Yingjun Li, Derong Shen, Tiezheng Nie, Ge Yu 0001, Jing Shan, Yue Kou |
WAIM | 2 |
| 2011 | Layered Graph Data Model for Data Management of DataSpace Support Platform
Derong Shen, Tiezheng Nie, Ge Yu 0001, Yue Kou |
WAIM | 2 |
| 2010 | Domain-oriented Deep Web Data Sources' Discovery and IdentificationabstractAs Deep Web contains tremendous well-structured data sources, how to integrate data sources in Deep Web has become a hotspot in current research. Accurately discovering and identifying Deep Web data sources related to a specific domain become key issues. We propose a Domain-Oriented Deep Web data source Discovery method (DO-DWD) and a novel Domain Identification strategy of Deep Web data sources (DIDW). In the discovery stage, we use machine learning algorithms and some heuristic rules to find query interfaces of the data sources; In the identification stage, we identify Deep Web data sources associated with the domain by calculating the relevance between a query interface and the domain based on semantic similarity. Finally, we have extensive experiments on a real data set showing that DO-DWD and DIDW are of high correctness and accuracy. Yingjun Li, Tiezheng Nie, Derong Shen, Ge Yu 0001 |
APWeb | 3 |
| 2010 | An Effective and High-quality Query Relaxation Solution on the Deep WebabstractBecause the amount of information contained on the Deep Web is much larger than the surface web, how to use it well has become a popular problem to research. When a query is sent to a deep web resource and the data sources return few results or even no result, a proper query relaxation solution should be adopted to get more satisfactory results to users. In this paper, such a query relaxation solution is presented. First, it solves the problem of relaxing attributes which contain multiple key words by value. That is, such attributes are not simply removed in the relaxation, but the query values of the attributes are modified. Second, when a data source returns many result pages, instead of getting all the pages, it evaluates the quality of the results in the current page to decide whether to send another query to fetch the next page. Thus, the number of query times is reduced. Finally, the experimental results demonstrate that both the result quality and the query efficiency are improved. Jing Shan, Derong Shen, Tiezheng Nie, Yue Kou, Ge Yu 0001 |
APWeb | 2 |
| 2010 | Domain-Independent Classification for Deep Web Interfaces
Yingjun Li, Derong Shen, Tiezheng Nie, Ge Yu 0001 |
WAIM | 3 |
| 2010 | Potential Role Based Entity Matching for Dataspaces Search
Yue Kou, Derong Shen, Tiezheng Nie, Ge Yu 0001 |
WISE | 2 |
| 2008 | A Novel and Effective Method for Web System Tuning Based on Feature Selection
Shi Feng 0001, Daling Wang, Derong Shen |
APWeb | 4 |
| 2008 | An Effective Method Supporting Data Extraction and Schema Recognition on Deep Web
Derong Shen, Tiezheng Nie |
APWeb | 2 |
| 2008 | An Effective Query Relaxation Solution for the Deep Web
Derong Shen, Yue Kou |
APWeb | 2 |
| 2008 | A Hierarchical Replica Location Approach Based on Cache Mechanism and Load Balancing in Data Grid
Baoyan Song, Yanying Mao, Derong Shen |
APWeb | 4 |
| 2008 | LG-ERM: An Entity-Level Ranking Mechanism for Deep Web QueryabstractWith the rapid growth of Web databases, it's necessary to extract and integrate large-scale data available in deep Web automatically. But current Web search engines conduct page-level ranking, which are becoming inadequate for entity-oriented vertical search. In this paper, we present an entity-level ranking mechanism called LG-ERM for deep Web query based on local scoring and global aggregation. Unlike traditional approaches, LG-ERM considers more rank influencing factors including the uncertainty of entity extraction, the style information of entities and the importance of Web sources, as well as the entity relationship. By combining local scoring and global aggregation in ranking, the query result can be more accurate and effective to meet users' needs. The experiments demonstrate the feasibility and effectiveness of the key techniques of LG-ERM. Yue Kou, Derong Shen, Ge Yu 0001, Tiezheng Nie |
WAIM | 2 |
| 2008 | Subject-Oriented Classification Based on Scale Probing in the Deep WebabstractTo access the large-scale data sources efficiently and automatically, it is necessary to classify these data sources into different domains and categories. In this paper, we propose a novel classification approach to classify data sources into detail domain subjects by query probing. In our approach, we train sample instances for each subject category and use them to probe the data scale of each source and category. And then we build a matrix to classify a data source into one or more subject categories and develop a decision algorithm based on probing iteration to rectify the classification result. Our experiments over real deep web sources show that our approach can achieve higher accuracy across a variety of data sources. Tiezheng Nie, Derong Shen, Ge Yu 0001, Yue Kou |
WAIM | 2 |
| 2008 | Efficient Top-k Data Sources Ranking for Query on Deep Web
Derong Shen, Meifang Li, Ge Yu 0001, Yue Kou, Tiezheng Nie |
WISE | 1 |
| 2006 | An Effective Service Discovery Model for Highly Reliable Web Services Composition in a Specific Domain
Derong Shen, Ge Yu 0001, Tiezheng Nie, Yue Kou, Meifang Li |
APWeb | 1 |
| 2005 | A Common Application-Centric QoS Model for Selecting Optimal Grid Services
Derong Shen, Ge Yu 0001, Tiezheng Nie |
APWeb | 1 |
| 2004 | Modeling QoS for Semantic Equivalent Web Services
Derong Shen, Ge Yu 0001, Tiezheng Nie, Xiaochun Yang 0001 |
WAIM | 1 |
| 2003 | e_SWDL: An XML Based Workflow Definition Language for Complicated Applications in Web Environments
Baoyan Song, Derong Shen, Ge Yu 0001 |
APWeb | 3 |
| 2003 | An Efficient User Task Handling Mechanism Based on Dynamic Load-Balance for Workflow Systems
Baoyan Song, Ge Yu 0001, Dan Wang 0019, Derong Shen, Guoren Wang |
APWeb | 4 |
| 2003 | An Ant Algorithm Based Dynamic Routing Strategy for Mobile Agents
Dan Wang 0019, Ge Yu 0001, Mingsong Lv, Baoyan Song, Derong Shen, Guoren Wang |
APWeb | 5 |