VLDB 2026 Research / reviewers in the wild / expert
Hengjie Song
dblp:65/657
· DBLP profile ↗
33ranked-venue papers
12as first author
17since 2021 · last 2026
0000-0003-4121-9466ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 8 first-author · 9 since 2021Databases, data management, data science and information retrieval · 9 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DASR+: Training Domain Distance Aware Network for Unsupervised Image Super-Resolution
Xiaorui Zhao, Yunxuan Wei, Xin Deng 0002, Yawei Li 0001, Radu Timofte, Hengjie Song, Shuhang Gu |
Int. J. Comput. Vis. | 6 |
| 2026 | Semi-Supervised and Transfer Learning-Based Smart Contract Vulnerability DetectionabstractSmart contracts are crucial for managing sensitive financial transactions. However, these contracts are inherently vulnerable which can cause significant security risks. Traditional methods for detecting smart contract vulnerabilities depend on expert-defined rules, which are often complicated and limited by human experts' individual experience. In contrast, deep learning-based approaches automatically extract intricate feature representations, greatly improving detection efficiency and accuracy. Nevertheless, these approaches rely on access to large amounts of high-quality labeled data. They are less effective when facing new types of vulnerabilities where labeled data are scarce. To this end, we propose the Semi-supervised Tuning (SST) approach for smart contract vulnerability detection. It first leverages a source model trained on labeled source data to extract features for new vulnerabilities. Subsequently, it performs semi-supervised learning to explore the feature structure of unlabeled data. In particular, SST groups contract code features and constructs a shared feature queue containing labeled and unlabeled contracts to explore the complete feature structure and guide model training. Extensive experimental evaluations based on two real-world datasets containing eleven smart contract vulnerabilities demonstrate that SST is significantly more advantageous compared to eight state-of-the-art baseline methods, outperforming them by 24.12% in terms of F1 scores. Hengjie Song, Yangkai Wang, Han Yu 0001, Siyu Jiang |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2024 | Towards Robust and Efficient Cloud-Edge Elastic Model Adaptation via Selective Entropy DistillationabstractThe conventional deep learning paradigm often involves training a deep model on a server and then deploying the model or its distilled ones to resource-limited edge devices. Usually, the models shall remain fixed once deployed (at least for some period) due to the potential high cost of model adaptation for both the server and edge sides. However, in many real-world scenarios, the test environments may change dynamically (known as distribution shifts), which often results in degraded performance. Thus, one has to adapt the edge models promptly to attain promising performance. Moreover, with the increasing data collected at the edge, this paradigm also fails to further adapt the cloud model for better performance. To address these, we encounter two primary challenges: 1) the edge model has limited computation power and may only support forward propagation; 2) the data transmission budget between cloud and edge devices is limited in latency-sensitive scenarios. In this paper, we establish a Cloud-Edge Elastic Model Adaptation (CEMA) paradigm in which the edge models only need to perform forward propagation and the edge models can be adapted online. In our CEMA, to reduce the communication burden, we devise two criteria to exclude unnecessary samples from uploading to the cloud, i.e., dynamic unreliable and low-informative sample exclusion. Based on the uploaded samples, we update and distribute the affine parameters of normalization layers by distilling from the stronger foundation model to the edge model with a sample replay strategy. Extensive experimental results on ImageNet-C and ImageNet-R verify the effectiveness of our CEMA. Yaofo Chen, Shuaicheng Niu, Yaowei Wang 0001, Shoukai Xu, Hengjie Song, Mingkui Tan |
ICLR | 5 |
| 2024 | FedRMS: Privacy-Preserving Federated Knowledge Graph Embedding Through RandomizationabstractRecent years have witnessed a growing interest in Federated Knowledge Graph Embedding, driven by its potential to leverage knowledge from various data owners to improve link prediction performance without the need for data sharing. Existing works typically assume that the central Federated Learning (FL) server owns a table containing unique entities/relations for all FL clients. In addition, all clients are assumed to use the same knowledge graph embedding method. However, these methods are vulnerable to privacy leakage and do not fully explore the different contributions of local entity embeddings. To bridge this gap, we propose a randomized embedding method selection approach for privacy-preserving federated knowledge graph embedding (FedRMS). It selects a knowledge graph embedding method for each client during the local training process with randomness and employs an attention-based aggregator to derive the global entity embedding on the FL server. Extensive experiments on three real-world public datasets demonstrate that FedRMS achieves significant improvements in terms of both privacy preservation and link prediction against 5 state-of-the-art methods. Qianyu Li 0002, Xiaoli Tang 0001, Siyao Zhou 0004, Han Yu 0001, Hengjie Song, Li-Zhen Cui 0001, Xiaoxiao Li 0001 |
ICME | 5 |
| 2024 | Modeling Time Decay Effect in Temporal Knowledge Graphs via Multivariate Hawkes ProcessabstractKnowledge Graph Embedding (KGE) is attracting growing research interest because it offers great flexibility for the manipulation and application of Knowledge Graphs (KGs). However, most existing works focus on static KGE, while temporal KGE is still in its infancy. Recent temporal KGE methods attempt to obtain the long-term dependency of facts in consecutive timestamps by merging historical fact information. However, they ignore the different impacts of historical facts on the current facts due to the time decay effect and heterogeneity of historical facts. To bridge this gap, we formalize the concept of fact formation sequence to describe the evolution of an entity and propose the Modeling Time Decay Effect in Temporal Knowledge Graphs via Multivariate Hawkes Process method (TimeDE). TimeDE uses the Hawkes process to model the time decay effect of historical facts. It also incorporates the attention mechanism based on the score function of static KGE methods to better capture the impacts of heterogeneous historical facts on the current facts. Extensive experiments on the five commonly-used benchmark datasets demonstrate that TimeDE achieves significant improvements in terms of both Mean Reciprocal Rank and Hits@K compared to state-of-the-art methods. Qianyu Li 0002, Jiebin Chen, Xiaoli Tang 0001, Han Yu 0001, Hengjie Song |
IJCNN | 5 |
| 2024 | ConCPDP: A Cross-Project Defect Prediction Method Integrating Contrastive Pretraining and Category Boundary AdjustmentabstractSoftware defect prediction (SDP) is a crucial phase preceding the launch of software products. Cross‐project defect prediction (CPDP) is introduced for the anticipation of defects in novel projects lacking defect labels. CPDP can use defect information of mature projects to speed up defect prediction for new projects. So that developers can quickly get the defect information of the new project, so that they can test the software project pertinently. At present, the predominant approaches in CPDP rely on deep learning, and the performance of the ultimate model is notably affected by the quality of the training dataset. However, the dataset of CPDP not only has few samples but also has almost no label information in new projects, which makes the general deep‐learning‐based CPDP model not ideal. In addition, most of the current CPDP models do not fully consider the enrichment of classification boundary samples after cross‐domain, leading to suboptimal predictive capabilities of the model. To overcome these obstacles, we present contrastive learning pretraining for CPDP (ConCPDP), a CPDP method integrating contrastive pretraining and category boundary adjustment. We first perform data augmentation on the source and target domain code files and then extract the enhanced data as an abstract syntax tree (AST). The AST is then transformed into an integer sequence using specific mapping rules, serving as input for the subsequent neural network. A neural network based on bidirectional long short‐term memory (Bi‐LSTM) will receive an integer sequence and output a feature vector. Then, the feature vectors are input into the contrastive module to optimise the feature extraction network. The pretrained feature extractor can be fine‐tuned by the maximum mean discrepancy (MMD) between the feature distribution of the source domain and the target domain and the binary classification loss on the source domain. This paper conducts a large number of experiments on the PROMISE dataset, which is commonly used for CPDP, to validate ConCPDP’s efficacy, achieving superior results in terms of F 1 measure, area under curve (AUC), and Matthew’s correlation coefficient (MCC). Hengjie Song, Yufei Pan, Le Ma 0003, Siyu Jiang |
IET Softw. | 1 |
| 2024 | MuLAN: Multi-level attention-enhanced matching network for few-shot knowledge graph completion
Qianyu Li 0002, Bozheng Feng, Xiaoli Tang 0001, Han Yu 0001, Hengjie Song |
Neural Networks | 5 |
| 2024 | Automated Dominative Subspace Mining for Efficient Neural Architecture SearchabstractNeural Architecture Search (NAS) aims to automatically find effective architectures within a predefined search space. However, the search space is often extremely large. As a result, directly searching in such a large search space is non-trivial and also very time-consuming. To address the above issues, in each search step, we seek to limit the search space to a small but effective subspace to boost both the search performance and search efficiency. To this end, we propose a novel Neural Architecture Search method via Dominative Subspace Mining (DSM-NAS) that finds promising architectures in automatically mined subspaces. Specifically, we first perform a global search,i.e., dominative subspace mining, to find a good subspace from a set of candidates. Then, we perform a local search within the mined subspace to find effective architectures. More critically, we further boost search performance by taking well-designed/ searched architectures to initialize candidate subspaces. Experimental results demonstrate that DSM-NAS not only reduces the search cost but also discovers better architectures than state-of-the-art methods in various benchmark search spaces. Yaofo Chen, Daihai Liao, Fanbing Lv, Hengjie Song, James T. Kwok, Mingkui Tan |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Adversarial domain adaptation for cross-project defect prediction
Hengjie Song, Le Ma 0003, Yufei Pan, Qingan Huang, Siyu Jiang |
Empir. Softw. Eng. | 1 |
| 2023 | A multimodal approach for improving market price estimation in online advertising
Tengyun Wang, Haizhi Yang, Yang Liu 0165, Han Yu 0001, Hengjie Song |
Knowl. Based Syst. | 5 |
| 2023 | Capsule neural tensor networks with multi-aspect information for Few-shot Knowledge Graph Completion
Qianyu Li 0002, Jiale Yao, Xiaoli Tang 0001, Han Yu 0001, Siyu Jiang, Haizhi Yang, Hengjie Song |
Neural Networks | 7 |
| 2023 | Dynamically Optimizing Display Advertising Profits Under Diverse Budget SettingsabstractAs a revolutionary auction mechanism for display advertising, real-time bidding (RTB) allows advertisers to purchase individual ad impressions through real-time auctions. In RTB, the demand-side platform (DSP) acts as advertisers' bidding agent and aims at developing appropriate bidding strategies to maximize their specific key performance indicators (KPIs). Existing bidding strategies perform well for optimizing profits when the ad budget severely limited. However, when there is sufficient budget, their performance deteriorates. This results in added complexity for advertisers when applying these approaches in practice, hindering wider adoption. To address this challenging limitation, we propose the Adaptive ROI-Aware Bidding (ARAB) approach. It intelligently analyzes the budget setting and auction market conditions, and adjusts the bidding function accordingly to optimize profits. Different from previous studies that only bid based on the ad revenue, our proposed ROI-aware bidding function also takes into account the ad cost at impression-level. By doing so, ARAB dynamically allocates the budget on more cost-effective impressions to increase profits. Through extensive offline experiments on two real-world public datasets, we demonstrate that the proposed ARAB has achieved significant improvements in terms of both profit and ROI compared to state-of-the-art approaches. Haizhi Yang, Tengyun Wang, Xiaoli Tang 0001, Han Yu 0001, Fei Liu 0006, Hengjie Song |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2022 | A cross-project defect prediction method based on multi-adaptation and nuclear normabstractAbstract Cross‐project defect prediction (CPDP) is an important research direction in software defect prediction. Traditional CPDP methods based on hand‐crafted features ignore the semantic information in the source code. Existing CPDP methods based on the deep learning model may not fully consider the differences among projects. Additionally, these methods may not accurately classify the samples near the classification boundary. To solve these problems, the authors propose a model based on multi‐adaptation and nuclear norm (MANN) to deal with samples in projects. The feature of samples were embedded into the multi‐core Hilbert space for distribution and the multi‐kernel maximum mean discrepancy method was utilised to reduce differences among projects. More importantly, the nuclear norm module was constructed, which improved the discriminability and diversity of the target sample by calculating and maximizing the nuclear norm of the target sample in the process of domain adaptation, thus improving the performance of MANN. Finally, extensive experiments were conducted on 11 sizeable open‐source projects. The results indicate that the proposed method exceeds the state of the art under the widely used metrics. Qingan Huang, Le Ma 0003, Siyu Jiang, Hengjie Song, Libiao Jiang, Chunyun Zheng |
IET Softw. | 5 |
| 2022 | Kaplan-Meier Markov network: Learning the distribution of market price by censored data in online advertisingabstractWith the rapid development of real-time bidding (RTB) in online advertising, learning the distribution of market price has attracted wide attention, since it plays a critical role in designing bidding strategies. One important problem is the right-censored issue in which the true market price can only be observed by the winner of the auction. To address this, existing studies often use Kaplan–Meier estimation (KM), which is one of the best options for survival analysis. However, these approaches depend on counting sample segments and cannot provide accurate predictions for each individual bid request. To enhance the prediction ability, we propose an original method to build the KM for each bid request by predicting (1) the probability of winning an auction at a specific market price, and (2) the probability of losing an auction at a certain bid price. To deal with the high-dimensional sample data common in RTB scenarios, we design a Markov network to calculate these two probabilities. Extensive experiments on two public datasets demonstrate that the proposed approach significantly outperforms state-of-the-art baselines in terms of various metrics, including Wasserstein distance, KL-divergence, average negative log probability and mean squared error. Tengyun Wang, Haizhi Yang, Siyu Jiang, Yueyue Shi, Qianyu Li 0002, Xiaoli Tang 0001, Han Yu 0001, Hengjie Song |
Knowl. Based Syst. | 8 |
| 2022 | Time-Aware Graph Embedding: A Temporal Smoothness and Task-Oriented ApproachabstractKnowledge graph embedding, which aims at learning the low-dimensional representations of entities and relationships, has attracted considerable research efforts recently. However, most knowledge graph embedding methods focus on the structural relationships in fixed triples while ignoring the temporal information. Currently, existing time-aware graph embedding methods only focus on the factual plausibility, while ignoring the temporal smoothness, which models the interactions between a fact and its contexts, and thus can capture fine-granularity temporal relationships. This leads to the limited performance of embedding related applications. To solve this problem, this article presents a Robustly Time-aware Graph Embedding (RTGE) method by incorporating temporal smoothness. Two major innovations of our article are presented here. At first, RTGE integrates a measure of temporal smoothness in the learning process of the time-aware graph embedding. Via the proposed additional smoothing factor, RTGE can preserve both structural information and evolutionary patterns of a given graph. Secondly, RTGE provides a general task-oriented negative sampling strategy associated with temporally aware information, which further improves the adaptive ability of the proposed algorithm and plays an essential role in obtaining superior performance in various tasks. Extensive experiments conducted on multiple benchmark tasks show that RTGE can increase performance in entity/relationship/temporal scoping prediction tasks. Shengjie Sun 0001, Huiguo Zhang, Chang'an Yi, Yuan Miao 0001, Xiaonan Meng, Ke Wang 0001, Huaqing Min, Hengjie Song, Chuanyan Miao |
ACM Trans. Knowl. Discov. Data | 11 |
| 2021 | Multi-task Learning for Bias-Free Joint CTR Prediction and Market Price Modeling in Online AdvertisingabstractThe rapid rise of real-time bidding-based online advertising has brought significant economic benefits and attracted extensive research attention. From the perspective of an advertiser, it is crucial to perform accurate utility estimation and cost estimation for each individual auction in order to achieve cost-effective advertising. These problems are known as the click through rate (CTR) prediction task and the market price modeling task, respectively. However, existing approaches treat CTR prediction and market price modeling as two independent tasks to be optimized without regard to each other, thus resulting in suboptimal performance. Moreover, they do not make full use of unlabeled data from the losing bids during estimations, which makes them suffer from the sample selection bias issue. To address these limitations, we propose Multi-task Advertising Estimator (MTAE), an end-to-end joint optimization framework which performs both CTR prediction and market price modeling simultaneously. Through multi-task learning, both estimation tasks can take advantage of knowledge transfer to achieve improved feature representation and generalization abilities. In addition, we leverage the abundant bid price signals in the full-volume bid request data and introduce an auxiliary task of predicting the winning probability into the framework for unbiased learning. Through extensive experiments on two large-scale real-world public datasets, we demonstrate that our proposed approach has achieved significant improvements over the state-of-the-art models under various performance metrics. Haizhi Yang, Tengyun Wang, Xiaoli Tang 0001, Qianyu Li 0002, Yueyue Shi, Siyu Jiang, Han Yu 0001, Hengjie Song |
CIKM | 8 |
| 2021 | Unsupervised Real-World Image Super Resolution via Domain-Distance Aware TrainingabstractThese days, unsupervised super-resolution (SR) is soaring due to its practical and promising potential in real scenarios. The philosophy of off-the-shelf approaches lies in the augmentation of unpaired data, i.e. first generating synthetic low-resolution (LR) images ${\mathcal{Y}^g}$ corresponding to real-world high-resolution (HR) images ${\mathcal{X}^r}$ in the real-world LR domain ${\mathcal{Y}^r}$, and then utilizing the pseudo pairs $\left\{ {{\mathcal{Y}^g},{\mathcal{X}^r}} \right\}$ for training in a supervised manner. Unfortunately, since image translation itself is an extremely challenging task, the SR performance of these approaches is severely limited by the domain gap between generated synthetic LR images and real LR images. In this paper, we propose a novel domain-distance aware super-resolution (DASR) approach for unsupervised real-world image SR. The domain gap between training data (e.g. ${\mathcal{Y}^g}$) and testing data (e.g. ${\mathcal{Y}^r}$) is addressed with our domain-gap aware training and domain-distance weighted supervision strategies. Domain-gap aware training takes additional benefit from real data in the target domain while domain-distance weighted supervision brings forward the more rational use of labeled source domain data. The proposed method is validated on synthetic and real datasets and the experimental results show that DASR consistently outperforms state-of-the-art unsupervised SR approaches in generating SR outputs with more realistic and natural textures. Codes are available at https://github.com/ShuhangGu/DASR. Yunxuan Wei, Shuhang Gu, Yawei Li 0001, Radu Timofte, Longcun Jin, Hengjie Song |
CVPR | 6 |
| 2020 | Kernel-target alignment based non-linear metric learning
Chunyan Miao, Yong Liu 0020, Hengjie Song, Huaqing Min |
Neurocomputing | 4 |
| 2019 | AKUPM: Attention-Enhanced Knowledge-Aware User Preference Model for RecommendationabstractRecently, much attention has been paid to the usage of knowledge graph within the context of recommender systems to alleviate the data sparsity and cold-start problems. However, when incorporating entities from a knowledge graph to represent users, most existing works are unaware of the relationships between these entities and users. As a result, the recommendation results may suffer a lot from some unrelated entities. Tengyun Wang, Haizhi Yang, Hengjie Song |
KDD | 4 |
| 2018 | Multi-instance transfer metric learning by weighted distribution and consistent maximum likelihood estimation
Siyu Jiang, Hengjie Song, Qingyao Wu, Michael Kwok-Po Ng, Huaqing Min, Shaojian Qiu |
Neurocomputing | 3 |
| 2017 | A Unified Framework for Metric Transfer LearningabstractTransfer learning has been proven to be effective for the problems where training data from a source domain and test data from a target domain are drawn from different distributions. To reduce the distribution divergence between the source domain and the target domain, many previous studies have been focused on designing and optimizing objective functions with the Euclidean distance to measure dissimilarity between instances. However, in some real-world applications, the Euclidean distance may be inappropriate to capture the intrinsic similarity or dissimilarity between instances. To deal with this issue, in this paper, we propose a metric transfer learning framework (MTLF) to encode metric learning in transfer learning. In MTLF, instance weights are learned and exploited to bridge the distributions of different domains, while Mahalanobis distance is learned simultaneously to maximize the intra-class distances and minimize the inter-class distances for the target domain. Unlike previous work where instance weights and Mahalanobis distance are trained in a pipelined framework that potentially leads to error propagation across different components, MTLF attempts to learn instance weights and a Mahalanobis distance in a parallel framework to make knowledge transfer across domains more effective. Furthermore, we develop general solutions to both classification and regression problems on top of MTLF, respectively. We conduct extensive experiments on several real-world datasets on object recognition, handwriting recognition, and WiFi location to verify the effectiveness of MTLF compared with a number of state-of-the-art methods. Sinno Jialin Pan, Hui Xiong 0001, Qingyao Wu, Ronghua Luo, Huaqing Min, Hengjie Song |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2016 | ML-FOREST: A Multi-Label Tree Ensemble Method for Multi-Label ClassificationabstractMulti-label classification deals with the problem where each example is associated with multiple class labels. Since the labels are often dependent to other labels, exploiting label dependencies can significantly improve the multi-label classification performance. The label dependency in existing studies is often given as prior knowledge or learned from the labels only. However, in many real applications, such prior knowledge may not be available, or labeled information might be very limited. In this paper, we propose a new algorithm, called Ml-Forest , to learn an ensemble of hierarchical multi-label classifier trees to reveal the intrinsic label dependencies. In Ml-Forest, we construct a set of hierarchical trees, and develop a label transfer mechanism to identify the multiple relevant labels in a hierarchical way. In general, the relevant labels at higher levels of the trees capture more discriminable label concepts, and they will be transferred into lower level children nodes that are harder to discriminate. The relevant labels in the hierarchy are then aggregated to compute label dependency and make the final prediction. Our empirical study shows encouraging results of the proposed algorithm in comparison with the state-of-the-art multi-label classification algorithms under Friedman test and post-hoc Nemenyi test. Qingyao Wu, Mingkui Tan, Hengjie Song, Jian Chen 0011, Michael Kwok-Po Ng |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2016 | Individual Judgments Versus Consensus: Estimating Query-URL RelevanceabstractQuery-URL relevance, measuring the relevance of each retrieved URL with respect to a given query, is one of the fundamental criteria to evaluate the performance of commercial search engines. The traditional way to collect reliable and accurate query-URL relevance requires multiple annotators to provide their individual judgments based on their subjective expertise (e.g., understanding of user intents). In this case, the annotators’ subjectivity reflected in each annotator individual judgment (AIJ) inevitably affects the quality of the ground truth relevance (GTR). But to the best of our knowledge, the potential impact of AIJs on estimating GTRs has not been studied and exploited quantitatively by existing work. This article first studies how multiple AIJs and GTRs are correlated. Our empirical studies find that the multiple AIJs possibly provide more cues to improve the accuracy of estimating GTRs. Inspired by this finding, we then propose a novel approach to integrating the multiple AIJs with the features characterizing query-URL pairs for estimating GTRs more accurately. Furthermore, we conduct experiments in a commercial search engine—Baidu.com—and report significant gains in terms of the normalized discounted cumulative gains. Hengjie Song, Huaqing Min, Qingyao Wu, Wei Wei 0002, Jianshu Weng, Xiaogang Han, Qiang Yang 0001, Jialiang Shi, Jiaqian Gu, Chunyan Miao, Toyoaki Nishida |
ACM Trans. Web | 1 |
| 2015 | Synthetic Evidential Study as Augmented Collective Thought Process - Preliminary Report
Toyoaki Nishida, Masakazu Abe, Takashi Ookaki, Divesh Lala, Sutasinee Thovutikul, Hengjie Song, Yasser Mohammad, Christian Nitschke, Yoshimasa Ohmoto, Atsushi Nakazawa, Takaaki Shochi, Jean-Luc Rouas, Aurélie Bugeau, Fabien Lotte, Zuheng Ming, Geoffrey Letournel, Marine Guerry, Dominique Fourer |
ACIIDS (1) | 6 |
| 2013 | Utilizing URLs Position to Estimate Intrinsic Query-URL RelevanceabstractQuery-URL relevance (QUR) is an important criterion to measure the quality of commercial search engines. However, the traditional way to collect high-quality QURs is time-consuming and labor-intensive since it is primarily based on human judges. To address these issues, numerous models have been studied to automatically infer the QURs. Unlike the prior studies in this literature, we first empirically analyze the correlation between multiple annotators' judgments on QURs and URL position in ranking lists. By doing so, we reveal and justify the potential impacts of URL position on inferring intrinsic QURs. Inspired by this finding, a position-sensitive model (PSM) is proposed to infer QURs more accurately. In contrast with most existing approaches that attempt to construct the direct relationship between QURs and the features characterizing query-URL pairs, PSM assumes that the QUR is connected with the features through URL position. We conducted the experiments in real search engine Baidu.com, and compared the experimental results to those of the typical methods used in similar tasks, reporting significant gains over click-through rate and the normalized discounted cumulative gains (NDCGs). Xiaogang Han, Wenjun Zhou 0001, Xing Jiang 0001, Hengjie Song, Toyoaki Nishida |
ICDM | 4 |
| 2012 | A Mouse-Trajectory Based Model for Predicting Query-URL RelevanceabstractFor the learning-to-ranking algorithms used in commercial search engines, a conventional way to generate the training examples is to employ professional annotators to label the relevance of query-url pairs. Since label quality depends on the expertise of annotators to a large extent, this process is time-consuming and labor-intensive. Automatically generating labels from click-through data has been well studied to have comparable or better performance than human judges. Click-through data present users’ action and imply their satisfaction on search results, but exclude the interactions between users and search results beyond the page-view level (e.g., eye and mouse movements). This paper proposes a novel approach to comprehensively consider the information underlying mouse trajectory and click-through data so as to describe user behaviors more objectively and achieve a better understanding of the user experience. By integrating multi-sources data, the proposed approach reveals that the relevance labels of query-url pairs are related to positions of urls and users’ behavioral features. Based on their correlations, query-url pairs can be labeled more accurately and search results are more satisfactory to users. The experiments that are conducted on the most popular Chinese commercial search engine (Baidu) validated the rationality of our research motivation and proved that the proposed approach outperformed the state-of-the-art methods. Hengjie Song, Ruoxue Liao, Xiangliang Zhang 0001, Chunyan Miao, Qiang Yang 0001 |
AAAI | 1 |
| 2011 | Generating True Relevance Labels in Chinese Search Engine Using Clickthrough DataabstractIn current search engines, ranking functions are learned from a large number of labeled pairs in which the labels are assigned by human judges, describing how well the URLs match the different queries. However in commercial search engines, collecting high quality labels is time-consuming and labor-intensive. To tackle this issue, this paper studies how to produce the true relevance labels for pairs using clickthrough data. By analyzing the correlations between query frequency, true relevance labels and users’ behaviors, we demonstrate that the users who search the queries with similar frequency have similar search intents and behavioral characteristics. Based on such properties, we propose an efficient discriminative parameter estimation in a multiple instance learning algorithm (MIL) to automatically produce true relevance labels for pairs. Furthermore, we test our approach using a set of real world data extracted from a Chinese commercial search engine. Experimental results not only validate the effectiveness of the proposed approach, but also indicate that our approach is more likely to agree with the aggregation of the multiple judgments when strong disagreements exist in the panel of judges. In the event that the panel of judges is consensus, our approach provides more accurate automatic label results. In contrast with other models, our approach effectively improves the correlation between automatic labels and manual labels. Hengjie Song, Chunyan Miao, Zhiqi Shen 0001 |
AAAI | 1 |
| 2011 | A probabilistic fuzzy approach to modeling nonlinear systems
Hengjie Song, Chunyan Miao, Zhiqi Shen 0001, Roel Wuyts, Maja D'Hondt, Francky Catthoor |
Neurocomputing | 1 |
| 2010 | Design of fuzzy cognitive maps using neural networks for predicting chaotic time series
Hengjie Song, Chunyan Miao, Zhiqi Shen 0001, Roel Wuyts, Maja D'Hondt, Francky Catthoor |
Neural Networks | 1 |
| 2010 | Implementation of Fuzzy Cognitive Maps Based on Fuzzy Neural Network and Application in Prediction of Time SeriesabstractThe fuzzy cognitive map (FCM) has gradually emerged as a powerful paradigm for knowledge representation and a simulation mechanism that is applicable to numerous research and application fields. However, since efficient methods to determine the states of the investigated system and to quantify causalities that are the very foundations of FCM theory are lacking, constructing FCMs for complex causal systems greatly depends on expert knowledge. The manually developed models have a substantial shortcoming due to the model subjectivity and difficulties with assessing its reliability. In this paper, we proposed a fuzzy neural network to enhance the learning ability of FCMs. Our approach incorporates the inference mechanism of conventional FCMs with the determination of membership functions, as well as the quantification of causalities. In this manner, FCM models of the investigated systems can automatically be constructed from data and, therefore, operate with less human intervention. In the employed fuzzy neural network, the concept of mutual subsethood is used to describe the causalities, which provides more transparent interpretation for causalities in FCMs. The effectiveness of the proposed approach in handling the prediction of time series is demonstrated through many numerical simulations. Hengjie Song, Chunyan Miao, Roel Wuyts, Zhiqi Shen 0001, Francky Catthoor |
IEEE Trans. Fuzzy Syst. | 1 |
| 2009 | A fuzzy neural network with fuzzy impact grades
Hengjie Song, Chunyan Miao, Zhiqi Shen 0001, Yuan Miao 0001, Bu-Sung Lee |
Neurocomputing | 1 |
| 2007 | Fuzzy cognitive map learning based on multi-objectiveparticle swarm optimizationabstractIn order to eliminate the excessive subjective elements involved in construction of FCMs, this paper proposes a new FCM learning algorithm which is based on the application of multi-objective particle swarm optimization. The simulation results show that the novel method not only implements inference process and FCM learning in parallel, and improves the efficiency and robustness of FCMs. Hengjie Song, Chunyan Miao, Zhiqi Shen 0001 |
GECCO | 1 |
| 2006 | Probabilistic Fuzzy Cognitive MapabstractIn this paper, we present the probabilistic fuzzy cognitive map (PFCM) which is a novel extension of FCM theory. Each concept in PFCM is extended to a fuzzy event that models not only the fuzzy degree but also the fuzzy probability of both the cause and the effect concepts. PFCM enhances the capability of conventional FCMs to handle both randomness and fuzziness which are necessary to model the uncertainty involved in inference process of complex causal system. A formalized inference process of PFCMs is presented for adjustments on probability of fuzzy events and for dynamic update of causal weights. This enables PFCM to synthetically analyze the impacts of randomness and fuzziness on causal inference process. The simulation result shows a good match to the above features of PFCM. PFCM, as an initial attempt, provides a heuristic approach to model the uncertainty of complex causal systems and opens a collection of interesting research issues for further research. Hengjie Song, Zhiqi Shen 0001, Chunyan Miao, Yuan Miao 0001 |
FUZZ-IEEE | 1 |