VLDB 2026 Research / reviewers in the wild / expert
Jie Liu 0022
dblp:03/2134-22
· DBLP profile ↗
38ranked-venue papers
7as first author
29since 2021 · last 2026
0000-0001-5953-4566ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 4 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dual-perspective hypergraph learning network for multimodal entity and relation extraction
Jie Liu 0022, Mingying Xu, Baowen Wu, Linqi Song, Yinqiao Li, Lei Shi 0030, Feifei Kou |
Expert Syst. Appl. | 1 |
| 2026 | A self-modified hypergraph neural network for multimodal relation extraction
Mingying Xu, Jie Liu 0022, Linqi Song, Yinqiao Li, Lei Shi 0030 |
Inf. Process. Manag. | 3 |
| 2026 | Multimodal emotion recognition via large model guided dialogue state tracking with dynamic graph refinement
Yu Sui, Haoze Guo, Jie Liu 0022, Jianyong Duan, Hao Wang 0018, Linqi Song, Guizhong Xu |
Pattern Recognit. | 4 |
| 2026 | The uncertainty advantage: Enhancing large language models' reliability through chain of uncertainty reasoning
Zirong Peng, Xiaoming Liu 0020, Guan Yang, Jie Liu 0022, Xueping Peng, Yang Long 0001 |
Pattern Recognit. Lett. | 4 |
| 2026 | Contrastive Prompt Learning in Structured Graph Networks for Multimodal Fake News DetectionabstractWith social media now the primary channel for information dissemination, multimodal disinformation poses a severe challenge to cyberspace governance and public safety due to its heightened deceptiveness and viral potential. Current detection methods predominantly rely on simple feature concatenation or shallow interaction models, which struggle to capture semantic relationships and inconsistencies across modalities. Moreover, they also universally face the semantic gap issue caused by misalignment between pre-training tasks and downstream classification objectives. To address these issues, this paper proposes a Contrastive Prompt Learning in Structured Graph Networks (CPLS) model. CPLS comprises three core modules: 1) Structured graph network construction module that builds graph structures preserving key semantics through semantic-aware graph pruning and fusion strategies; 2) Structured graph network optimization module that uses data augmentation and contrastive learning to capture robust local node features and global graph structures; 3) Contrastive prompt learning module that innovatively integrates prompt learning with graph contrastive learning, reframing downstream classification tasks as contrastive tasks to unify optimization of pre-training and downstream objectives. Experiments on public datasets show that our method achieves accuracy improvements of 2.4% on Weibo and 4.5% on Twitter. Ying Guo 0004, Kexin Zhen, Jie Liu 0022 |
IEEE Trans. Big Data | 3 |
| 2025 | T-MES: Trait-Aware Mix-of-Experts Representation Learning for Multi-trait Essay ScoringabstractIn current research on automatic essay scoring, related work tends to focus more on evaluating the overall quality or a single trait of prompt-specific essays. However, when scoring essays in an educational context, it is essential not only to consider the overall score but also to provide feedback on various aspects of the writing. This helps students clearly identify areas for improvement, enabling them to engage in targeted practice. Although many methods have been proposed to address the scoring issue, they still suffer from insufficient learning of trait representations and overlook the diversity and correlations between trait scores in the scoring process. To address this problem, we propose a novel multi-trait essay scoring method based on Trait-Aware Mix-of-Experts Representation Learning. Our method obtains trait-specific essay representations using a Mix-of-Experts scoring architecture. Furthermore, based on this scoring architecture, we propose a diversified trait-expert method to learn distinguishable expert weights. And to facilitate multi-trait scoring, we introduce two trait correlation learning strategies that achieve learning the correlations among traits. Experimental results demonstrate the effectiveness of our method, and compared to existing methods, it achieves a further improvement in computational efficiency. Jie Liu 0022 |
COLING | 2 |
| 2025 | Determine-Then-Ensemble: Necessity of Top-k Union for Large Language Model EnsemblingabstractLarge language models (LLMs) exhibit varying strengths and weaknesses across different tasks, prompting recent studies to explore the benefits of ensembling models to leverage their complementary advantages. However, existing LLM ensembling methods often overlook model compatibility and struggle with inefficient alignment of probabilities across the entire vocabulary. In this study, we empirically investigate the factors influencing ensemble performance, identifying model performance, vocabulary size, and response style as key determinants, revealing that compatibility among models is essential for effective ensembling. This analysis leads to the development of a simple yet effective model selection strategy that identifies compatible models. Additionally, we introduce the \textsc{Uni}on \textsc{T}op-$k$ \textsc{E}nsembling (\textsc{UniTE}), a novel approach that efficiently combines models by focusing on the union of the top-k tokens from each model, thereby avoiding the need for full vocabulary alignment and reducing computational overhead. Extensive evaluations across multiple benchmarks demonstrate that \textsc{UniTE} significantly enhances performance compared to existing methods, offering a more efficient framework for LLM ensembling. Han Wu 0004, Sichun Luo, Xiongwei Han, Jie Liu 0022, Zhijiang Guo, Linqi Song |
ICLR | 6 |
| 2025 | Learning Hanzi Character Through VR-Based Mortise-Tenon
Conglin Ma, Sen-Zhe Xu 0001, Ju Dai, Jie Liu 0022, Feng Zhou 0007 |
ICXR | 5 |
| 2025 | Making meta-learning solve cross-prompt automatic essay scoring
Jie Liu 0022, Mingying Xu, Liguang Yang, Jianshe Zhou |
Expert Syst. Appl. | 3 |
| 2025 | HADT: Image super-resolution restoration using Hybrid Attention-Dense Connected Transformer Networks
Ying Guo 0004, Jie Liu 0022, Chong Di 0001, Keqing Ning |
Neurocomputing | 3 |
| 2025 | Causal-relationship representation enhanced joint extraction model for elements and relationships
Xiaoming Liu 0020, Guan Yang, Sensen Dong, Xingang Hu, Jie Liu 0022 |
Neurocomputing | 6 |
| 2025 | Fine-grained entity typing based on hyperbolic representation and label-context interaction
Mingying Xu, Jie Liu 0022, Weiping Ding 0001, Lei Shi 0030, Kaiyang Zhong |
Inf. Sci. | 3 |
| 2025 | Dmvae: a dual-stream multi-modal variational autoencoder for multi-task fake news detection
Ying Guo 0004, Shuting Hu, Chong Di 0001, Jie Liu 0022 |
Pattern Anal. Appl. | 5 |
| 2025 | HCT: image super-resolution restoration using hierarchical convolution transformer networks
Ying Guo 0004, Jie Liu 0022, Chong Di 0001, Keqing Ning |
Pattern Anal. Appl. | 4 |
| 2025 | Continuous Adaptive Knowledge Distillation for Few-Shot Relation ExtractionabstractThe goal of continuous few-shot relation extraction is to enable the model to continuously learn new relation types under conditions with limited labeled training data while avoiding the forgetting of previously learned relations. The primary challenges include catastrophic forgetting of old relations and overfitting due to data sparsity. To address these challenges, this article proposes a hierarchical distillation model that innovatively combines contrastive distillation with orthogonal adversarial distillation techniques. Specifically, we introduce a contrastive distillation approach in the feature distillation layer, integrating adaptive cosine techniques and negative sampling strategies to ensure that the model effectively retains and utilizes knowledge from previous tasks when learning new ones. Additionally, we employ orthogonal adversarial distillation in the hidden distillation layer to alleviate overfitting in low-resource scenarios. Experimental results demonstrate that our proposed method significantly outperforms the current state-of-the-art models for continuous few-shot relation extraction on two benchmark datasets, validating its effectiveness in handling few-shot data and knowledge transfer. Shuo Zhao 0008, Jianyong Duan, Hao Wang 0018, Jie Liu 0022 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 6 |
| 2024 | Combining Contrastive Learning and Sequence Learning for Automated Essay Scoring
Jie Liu 0022, Jianshe Zhou, Wang Jiong |
ICANN (9) | 2 |
| 2024 | Lifelong Sentiment Classification Based on Adaptive Parameter Updating
Kaifeng Nie, Jie Liu 0022 |
ICANN (7) | 5 |
| 2024 | A Multi-Polarization Framework for Enhanced RFI Suppression in Real SAR DataabstractSynthetic aperture radar (SAR) is a kind of active microwave remote sensing imaging radar, which can obtain high-resolution two-dimensional SAR images. As a multi-parameter, multi-channel SAR, polarimetric SAR (PolSAR) provides rich scattering information for topographic mapping, ocean exploration, polar observation, target identification, and many other fields. Compared to single-polarization SAR, multi-polarization SAR greatly improves the potential information of the data by extending the one-dimensional information. However, the above tasks cannot be carried out without clean SAR echo signal. The radio frequency interference (RFI) signals, seriously affect the subsequent tasks of PolSAR, and there is a great deal of potential information between polarized data. Therefore, this paper proposes a framework for combining multiple polarization data to improve low-rank based methods’ performance. Based on the proposed framework, one experiment is conducted on real PolSAR data, the experiment uses the PCA method to verify the applicability of the proposed framework in interference suppression. At last, the result verifies the framework achieves better suppression of low-rank based method. Yuan Mao, Xutao Yu, Zaichen Zhang, Hui Zhang 0071, Jie Liu 0022, Yan Huang 0018 |
IGARSS | 5 |
| 2024 | An Entailment Tree Generation Approach for Multimodal Multi-Hop Question Answering with Mixture-of-Experts and Iterative Feedback MechanismabstractWith the rise of large-scale language models (LLMs), it is currently popular and effective to convert multimodal information into text descriptions for multimodal multi-hop question answering. However, we argue that the current methods of multi-modal multi-hop question answering still mainly face two challenges: 1) The retrieved evidence containing a large amount of redundant information, inevitably leads to a significant drop in performance due to irrelevant information misleading the prediction. 2) The reasoning process without interpretable reasoning steps makes the model difficult to discover the logical errors for handling complex questions. To solve these problems, we propose a unified LLMs-based approach but without heavily relying on them due to the LLM's potential errors, and innovatively treat multimodal multi-hop question answering as a joint entailment tree generation and question answering problem. Specifically, we design a multi-task learning framework with a focus on facilitating common knowledge sharing across interpretability and prediction tasks while preventing task-specific errors from interfering with each other via mixture of experts. Afterward, we design an iterative feedback mechanism to further enhance both tasks by feeding back the results of the joint training to the LLM for regenerating entailment trees, aiming to iteratively refine the potential answer. Notably, our method has won the first place in the official leaderboard of WebQA (since April 10, 2024), and achieves competitive results on MultimodalQA. Haocheng Lv, Jie Liu 0022, Jianyong Duan, Hao Wang 0018, Mingying Xu |
ACM Multimedia | 3 |
| 2024 | Improved conversational recommender system based on dialog contextabstractAbstract Conversational recommender system (CRS) needs to be seamlessly integrated between the two modules of recommendation and dialog, aiming to recommend high-quality items to users through multiple rounds of interactive dialogs. Items can typically refer to goods, movies, news, etc. Through this form of interactive dialog, users can express their preferences in real time, and the system can fully understand the user’s thoughts and recommend corresponding items. Although mainstream dialog recommendation systems have improved the performance to some extent, there are still some key issues, such as insufficient consideration of the entity’s order in the dialog, the different contributions of items in the dialog history, and the low diversity of generated responses. To address these shortcomings, we propose an improved dialog context model based on time-series features. Firstly, we augment the semantic representation of words and items using two external knowledge graphs and align the semantic space using mutual information maximization techniques. Secondly, we add a retrieval model to the dialog recommendation system to provide auxiliary information for generating replies. We then utilize a deep timing network to serialize the dialog content and more accurately learn the feature relationship between users and items for recommendation. In this paper, the dialog recommendation system is divided into two components, and different evaluation indicators are used to evaluate the performance of the dialog component and the recommendation component. Experimental results on widely used benchmarks show that the proposed method is effective. Jie Liu 0022, Jianyong Duan |
Nat. Lang. Eng. | 2 |
| 2024 | Multi-label feature selection by strongly relevant label gain and label mutual aid
Jianhua Dai 0003, Chucai Zhang, Jie Liu 0022 |
Pattern Recognit. | 4 |
| 2024 | Seq2Set2Seq: A Two-stage Disentangled Method for Reply Keyword Generation in Social MediaabstractSocial media produces large amounts of content every day. How to predict the potential influences of the contents from a social reply feedback perspective is a key issue that has not been explored. Thus, we propose a novel task named reply keyword prediction in social media, which aims to predict the keywords in the potential replies in as many aspects as possible. One prerequisite challenge is that the accessible social media datasets labeling such keywords remain absent. To solve this issue, we propose a new dataset, 1 to study the reply keyword prediction in social media. This task could be seen as a single-turn dialogue keyword prediction for open-domain dialogue system. However, existing methods for dialogue keyword prediction cannot be adopted directly, which has two main drawbacks. First, they do not provide an explicit mechanism to model topic complementarity between keywords which is crucial in social media to controllably model all aspects of replies. Second, the collocations of keywords are not explicitly modeled, which also makes it less controllable to optimize for fine-grained prediction since the context information is much less than that in dialogue. To address these issues, we propose a two-stage disentangled framework, which can optimize the complementarity and collocation explicitly in a disentangled fashion. In the first stage, we use a sequence-to-set paradigm via multi-label prediction and determinantal point processes, to generate a set of keyword seeds satisfying the complementarity. In the second stage, we adopt a set-to-sequence paradigm via seq2seq model with the keyword seeds guidance from the set, to generate the more-fine-grained keywords with collocation. Experiments show that this method can generate not only a more diverse set of keywords but also more relevant and consistent keywords. Furthermore, the keywords obtained based on this method can achieve better reply generation results in the retrieval-based system than others. Jie Liu 0022, Shizhu He, Shun Wu, Kang Liu 0001, Shenping Liu |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2024 | Document-Level Relation Extraction Based on Machine Reading Comprehension and Hybrid Pointer-sequence LabelingabstractDocument-level relational extraction requires reading, memorization, and reasoning to discover relevant factual information in multiple sentences. It is difficult for the current hierarchical network and graph network methods to fully capture the structural information behind the document and make natural reasoning from the context. Different from the previous methods, this article reconstructs the relation extraction task into a machine reading comprehension task. Each pair of entities and relationships is characterized by a question template, and the extraction of entities and relationships is translated into identifying answers from the context. To enhance the context comprehension ability of the extraction model and achieve more precise extraction, we introduce large language models (LLMs) during question construction, enabling the generation of exemplary answers. Besides, to solve the multi-label and multi-entity problems in documents, we propose a new answer extraction model based on hybrid pointer-sequence labeling, which improves the reasoning ability of the model and realizes the extraction of zero or multiple answers in documents. Extensive experiments on three public datasets show that the proposed method is effective. Jie Liu 0022, Jianyong Duan, Guixia Guan, Jianshe Zhou |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2024 | Dynamic Correlation Learning and Regularization for Multi-Label Confidence CalibrationabstractModern visual recognition models often display overconfidence due to their reliance on complex deep neural networks and one-hot target supervision, resulting in unreliable confidence scores that necessitate calibration. While current confidence calibration techniques primarily address single-label scenarios, there is a lack of focus on more practical and generalizable multi-label contexts. This paper introduces the Multi-Label Confidence Calibration (MLCC) task, aiming to provide well-calibrated confidence scores in multi-label scenarios. Unlike single-label images, multi-label images contain multiple objects, leading to semantic confusion and further unreliability in confidence scores. Existing single-label calibration methods, based on label smoothing, fail to account for category correlations, which are crucial for addressing semantic confusion, thereby yielding sub-optimal performance. To overcome these limitations, we propose the Dynamic Correlation Learning and Regularization (DCLR) algorithm, which leverages multi-grained semantic correlations to better model semantic confusion for adaptive regularization. DCLR learns dynamic instance-level and prototype-level similarities specific to each category, using these to measure semantic correlations across different categories. With this understanding, we construct adaptive label vectors that assign higher values to categories with strong correlations, thereby facilitating more effective regularization. We establish an evaluation benchmark, re-implementing several advanced confidence calibration algorithms and applying them to leading multi-label recognition (MLR) models for fair comparison. Through extensive experiments, we demonstrate the superior performance of DCLR over existing methods in providing reliable confidence scores in multi-label scenarios. Tianshui Chen, Weihang Wang 0008, Tao Pu 0002, Jinghui Qin, Zhijing Yang, Jie Liu 0022, Liang Lin 0004 |
IEEE Trans. Image Process. | 6 |
| 2024 | Category-Adaptive Label Discovery and Noise Rejection for Multi-Label Recognition With Partial Positive LabelsabstractAs a cost-effective alternative to standard multi-label learning, the multi-label image recognition with partial positive labels (MLR-PPL) task attracts increasing attention, in which merely a portion of positive labels are given while the rest of positive labels and all negative labels are missing. To facilitate this task, we propose a novel framework that leverages semantic correlation among different images in a category-adaptive manner to complement unknown labels accurately. Specifically, the proposed framework consists of two complementary modules. 1) A category-adaptive label discovery (CALD) module is designed to measure the semantic similarity between positive samples and then complement unknown labels with high similarities. 2) A category-adaptive noise rejection (CANR) module is designed to compute the sample weights based on semantic similarities from different samples and discard noisy labels with low weights. Due to the various degrees of confidence calibration among different categories, searching appropriate thresholds for each category in the proposed framework is highly time-consuming. To avoid such a resource-intensive manual tuning, we introduce a category-adaptive threshold updating algorithm that introduces the category-specific positive and negative similarity to adjust the threshold adaptively. Extensive experiments on various benchmarks show that the proposed framework performs better than current state-of-the-art algorithms. Tao Pu 0002, Qianru Lao, Hefeng Wu, Tianshui Chen, Ling Tian, Jie Liu 0022, Liang Lin 0004 |
IEEE Trans. Multim. | 6 |
| 2023 | Hessian Non-negative Hypergraph
Lingling Li 0004, Taisong Jin, Jie Liu 0022 |
ICIC (5) | 5 |
| 2023 | Zero-Shot Relation Triplet Extraction via Retrieval-Augmented Synthetic Data Generation
Yuechen Yang, Hayilang Zhang, Zhengxin Gao, Hao Wang 0018, Jianyong Duan, Jie Liu 0022 |
ICONIP (15) | 8 |
| 2023 | Optimal Factored LT codes for Distributed Matrix MultiplicationabstractDue to the presence of slow or failed worker computers (called stragglers), distributed matrix multiplication over large clusters may encounter delays. To tackle this issue, Factored Luby Transform (FLT) codes have been proposed for edge computing scenarios involving distributed matrix multiplication of the form C = A * B. However, the decoding time of FLT codes becomes a bottleneck in coded distributed matrix multiplication, primarily due to the high computational complexity associated with peeling decoding. To overcome this limitation, this paper introduces the modified generalized degree distribution algorithm (MGDDA) for designing an optimal degree distribution in FLT codes with BP decoding. The MGDDA algorithm minimizes overhead and maximizes performance efficiency in coded distributed matrix multiplication scenarios. Additionally, we propose an optimal encoding algorithm that achieves a balanced probability of connecting input symbols from matrices A and B to encoded symbols. This algorithm further enhances the performance of FLT codes. Simulation results consistently demonstrate that our proposed Optimal FLT (OFTL) method outperforms other existing approaches in terms of average overhead, block error rate (BER) and computational delay. Jie Liu 0022 |
ICPADS | 2 |
| 2023 | Ridge-Regression-Induced Robust Graph Relational NetworkabstractGraph convolutional networks (GCNs) have attracted increasing research attention, which merits in its strong ability to handle graph data, such as the citation network or social network. Existing models typically use first-order neighborhood information to design specific convolution operations, which aggregate the features of all adjacent nodes. However, such models ignore the high-order spatial relationship among neighboring nodes in noisy data due to its modeling complexity. In this article, we propose a novel robust graph relational network to address this issue toward modeling high-order relationships in noisy data for graph convolution. Our key innovation lies in designing a generic relation network layer, which is used to infer the underlying relations among adjacent noisy nodes. Specifically, a fixed number of adjacent nodes for each node is chosen by solving the ridge regression problem, in which the regression coefficients are used to rank the adjacent nodes of each node in a graph. Furthermore, to mine the rich features, we extract high-order information from the nodes to significantly enhance the representation ability of the GCNs for extensive applications. We conduct extensive semisupervised node classification experiments on the noisy benchmark datasets, which clearly show that our model is superior to the existing methods and can achieve state-of-the-art performance. Taisong Jin, Jie Liu 0022, Huaqiang Dai, Lingling Li 0004, Fenlin Liu, Yongdong Zhang 0001 |
IEEE Trans. Cybern. | 2 |
| 2020 | Robust non-negative matrix factorization with multiple correntropy-induced hypergraph regularizer
Jie Liu 0022, Jianshe Zhou, Yongdong Zhang 0001 |
Signal Process. | 1 |
| 2018 | Light Field Foreground Matting Based on Defocus and Correspondence
Jianshe Zhou, Tuya Naren, Yike Ma, Jie Liu 0022 |
MMM (1) | 5 |
| 2018 | Study on Ontology Ranking Models Based on the Ensemble LearningabstractThis article describes how more knowledge appears on the Internet than in an ontological form. Displaying results to users precisely when searching is the key issue of the research on ontology retrieval. The considered factors of ontology ranking are not only limited to internal character-matching, but analysis of metadata, including the entities, structures and the relations in ontologies. Currently, existing single feature ranking algorithms focus on the structures, elements and the contents of a certain aspect in ontology, thus, the results are not satisfactory. Combining multiple single-featured models seems to achieve better results, but the objectivity and versatility of models' weights are debatable. Machine learning effectively solves the problem and putting advantages of ranking learning algorithms together is the pressing issue. So we propose ensemble learning strategies to combine different algorithms in ontology ranking. And the ranking result is more satisfied compared to Swoogle and base algorithms. Jie Liu 0022, Yuan Kerou, Jianshe Zhou, Jinsheng Shi |
Int. J. Semantic Web Inf. Syst. | 1 |
| 2018 | Hierarchical BoW with segmental sparse coding for large scale image classification and retrieval
Jianshe Zhou, Narentuya, Sheng Tang, Jie Liu 0022 |
Multim. Tools Appl. | 4 |
| 2017 | Collaborative Dictionary Learning and Soft Assignment for Sparse Coding of Image Features
Jie Liu 0022, Sheng Tang, Yu Li 0016 |
MMM (1) | 1 |
| 2017 | Kernelized product quantization
Jie Liu 0022, Jianshe Zhou, Jinsheng Shi, Yongdong Zhang 0001 |
Neurocomputing | 1 |
| 2017 | Creating Affective Autonomous Characters Using Planning in Partially Observable Stochastic DomainsabstractThe ability to reason about and respond to their own emotional states can enhance the believability of Non-Player Characters (NPCs). In this paper, we use a Partially Observable Markov Decision Process (POMDP)-based framework to model emotion over time. A two-level appraisal model, involving quick and reactive vs. slow and deliberate appraisals, is proposed for the creation of affective autonomous characters based on POMDPs, wherein the probability of goal satisfaction is used in an appraisal and reappraisal process for emotion generation. We not only extend Probabilistic Computation Tree Logic (PCTL) for reasoning about the properties of emotional states based on POMDPs but also illustrate how four reactive (primary) emotions and nine deliberate (secondary) emotions can be derived by combining PCTL with the belief-desire theory of emotion. The results of an empirical study suggest that the proposed model can be used to create characters that appear to be more believable and more intelligent. Xiangyang Huang, Shudong Zhang, Weigong Zhang, Jie Liu 0022 |
IEEE Trans. Comput. Intell. AI Games | 5 |
| 2016 | Ontology representation and mapping of common fuzzy knowledge
Jie Liu 0022, Bo-Ju Zheng, Jianshe Zhou, Zhengtao Yu 0001 |
Neurocomputing | 1 |
| 2016 | Web Image Search Re-Ranking With Click-Based Similarity and TypicalityabstractIn image search re-ranking, besides the well-known semantic gap, intent gap, which is the gap between the representation of users' query/demand and the real intent of the users, is becoming a major problem restricting the development of image retrieval. To reduce human effects, in this paper, we use image click-through data, which can be viewed as the implicit feedback from users, to help overcome the intention gap, and further improve the image search performance. Generally, the hypothesis-visually similar images should be close in a ranking list-and the strategy-images with higher relevance should be ranked higher than others-are widely accepted. To obtain satisfying search results, thus, image similarity and the level of relevance typicality are determinate factors correspondingly. However, when measuring image similarity and typicality, conventional re-ranking approaches only consider visual information and initial ranks of images, while overlooking the influence of click-through data. This paper presents a novel re-ranking approach, named spectral clustering re-ranking with click-based similarity and typicality. First, to learn an appropriate similarity measurement, we propose click-based multi-feature similarity learning algorithm, which conducts metric learning based on click-based triplets selection, and integrates multiple features into a unified similarity space via multiple kernel learning. Then, based on the learnt click-based image similarity measure, we conduct spectral clustering to group visually and semantically similar images into same clusters, and get the final re-rank list by calculating click-based clusters typicality and within-clusters click-based image typicality in descending order. Our experiments conducted on two real-world query-image data sets with diverse representative queries show that our proposed re-ranking approach can significantly improve initial search results, and outperform several existing re-ranking approaches. Tao Mei 0001, Yongdong Zhang 0001, Jie Liu 0022, Shin'ichi Satoh 0001 |
IEEE Trans. Image Process. | 4 |