Jialie Shen 0001

dblp:33/7046 · also Jerry Jialie Shen · DBLP profile ↗
← Back
153ranked-venue papers
30as first author
38since 2021 · last 2026
0000-0002-4560-8509ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 84 · 15 first-author · 23 since 2021Databases, data management, data science and information retrieval · 45 · 13 first-author · 6 since 2021Artificial intelligence and machine learning · 40 · 3 first-author · 14 since 2021Computer networks · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-author
YearPublicationVenuePosition
2026 Whole-Field Action Sensing via Wearable Single-Channel EMG Sensors and Resource-Efficient Motion Network
abstract
The proliferation of collaborative training and multi-person sports has underscored the necessity for concurrent whole-field action sensing. However, Electromyography (EMG) recognition, which plays a pivotal role in Wearable Human Activity Recognition (WHAR) for analyzing muscle activity and decoding action intent, still faces challenges in achieving a balance between performance, cost, and efficiency in multi-person scenarios. Unlike current channel-expansion solutions, we propose a wireless wearable Single-Dimensional Sparse EMG (2SEMG) Sensor for efficient personal sampling. These action-unaffected sensors leverage the proposed lightweight One-Dimensional Motion Network (OMONet) to facilitate concurrent action sensing. Experiments demonstrate that OMONet achieves leading performance and efficiency in action signal recognition, and two real-world badminton matches further confirm the performance, robustness, and real-time efficiency of the whole-field action sensing network constructed via 2SEMG Sensors and OMONet.
Xuanming Jiang, Dingyu Nie, Baoyi An 0001, Yuzhe Zheng, Yichuan Mao, Jialie Shen 0001, Xueming Qian, Zhiwen Jin, Guoshuai Zhao 0001
AAAI6
2026 Trustworthy Personalized Retrieval in the LLM Era: Fairness, Bias Awareness, Privacy, and Multimodality
abstract
Personalized search and recommendation are increasingly powered by large language models (LLMs) and multimodal deep learning, yet these advances raise practical concerns about bias, unfair outcomes, privacy risks, and limited transparency . At the same time, personalization pipelines are increasingly end-to-end and adaptive, meaning that small modeling or logging choices can propagate into downstream exposure effects and feedback loops that are hard to detect post hoc. As conversational and generative interfaces become the default access layer, users also face new challenges in judging provenance, uncertainty, and whether recommendations reflect their intent or the model's priors. This tutorial provides a unified, retrieval-centric view of trustworthy personalization: (1) how to measure and mitigate unfairness and bias in ranking, clustering, and classification; (2) how to design bias-aware retrieval experiences that improve user awareness; (3) how to support privacy-aware personal information retrieval (e.g., archived email, lifelogs) without sacrificing utility; and (4) how these concerns evolve for multimodal and retrieval-augmented generative recommendation, including LLM-based retriever/reranker settings. Participants will leave with concrete evaluation protocols, common failure modes, and a practical blueprint for building and auditing end-to-end IR/Rec pipelines for effectiveness, fairness, and privacy.
Frank Hopfgartner, Jialie Shen 0001
SIGIR2
2026 Dual transferable knowledge interaction for source-free domain adaptation
Mengmeng Zhan, Zongqian Wu, Jiaying Yang, Jialie Shen 0001, Xiaofeng Zhu 0001
Inf. Process. Manag.5
2026 A Plug-and-Play Model-Agnostic Embedding Enhancement Approach for Explainable Recommendation
abstract
Existing multimedia recommender systems provide users with suggestions of media by evaluating similarities, such as games and movies. To enhance the semantics and explainability of embeddings, it is a consensus to apply additional information (e.g., interactions, contexts, popularity). However, without systematically considering representativeness and value, the utility and explainability of embedding drop drastically. Hence, we introduceRVRec, a plug-and-play model-agnostic embedding enhancement approach that can improve both personality and explainability of existing systems. Specifically, we propose a probability-based embedding optimization method that uses a contrastive loss based on negative 2-Wasserstein distance to learn to enhance the representativeness of the embeddings. In addition, we introduce a reweighing method based on a multivariate Shapley values strategy to evaluate and explore the value of interactions and embeddings. Extensive experiments on multiple backbone recommenders and real-world datasets show that RVRec can improve the personalization and explainability of existing recommenders, outperforming state-of-the-art baselines.
Yunqi Mi, Boyang Yan, Guoshuai Zhao 0001, Jialie Shen 0001, Xueming Qian
IEEE Trans. Multim.4
2025 MTGA: Multi-View Temporal Granularity Aligned Aggregation for Event-Based Lip-Reading
abstract
Lip-reading is to utilize the visual information of the speaker’s lip movements to recognize words and sentences. Existing event-based lip-reading solutions integrate different frame rate branches to learn spatio-temporal features of varying granularities. However, aggregating events into event frames inevitably leads to the loss of fine-grained temporal information within frames. To remedy this drawback, we propose a novel framework termed Multi-view Temporal Granularity aligned Aggregation (MTGA). Specifically, we first present a novel event representation method, namely time-segmented voxel graph list, where the most significant local voxels are temporally connected into a graph list. Then we design a spatio-temporal fusion module based on temporal granularity alignment, where the global spatial features extracted from event frames, together with the local relative spatial and temporal features contained in voxel graph list are effectively aligned and integrated. Finally, we design a temporal aggregation module that incorporates positional encoding, which enables the capture of local absolute spatial and global temporal information. Experiments demonstrate that our method outperforms both the event-based and video-based lip-reading counterparts.
Yong Luo 0002, Wei Yu 0004, Zheng He 0001, Jialie Shen 0001
AAAI7
2025 Coupling the Generator with Teacher for Effective Data-Free Knowledge Distillation
Xu Chen 0053, Yang Li 0251, Yahong Han, Guangquan Xu, Jialie Shen 0001
ICCV5
2025 Decision Mixer: Integrating Long-term and Local Dependencies via Dynamic Token Selection for Decision-Making
abstract
The Conditional Sequence Modeling (CSM) paradigm, benefiting from the transformer’s powerful distribution modeling capabilities, has demonstrated considerable promise in offline Reinforcement Learning (RL) tasks. Depending on the task’s nature, it is crucial to carefully balance the interplay between inherent local features and long-term dependencies in Markov decision trajectories to mitigate potential performance degradation and unnecessary computational overhead. In this paper, we propose Decision Mixer (DM), which addresses the conflict between features of different scales in the modeling process from the perspective of dynamic integration. Drawing inspiration from conditional computation, we design a plug-and-play dynamic token selection mechanism to ensure the model can effectively allocate attention to different features based on task characteristics. Additionally, we employ an auxiliary predictor to alleviate the short-sightedness issue in the autoregressive sampling process. DM achieves state-of-the-art performance on various standard RL benchmarks while requiring significantly fewer computational resources, offering a viable solution for building efficient and scalable RL foundation models. Code is available at here.
Hongling Zheng, Li Shen 0008, Yong Luo 0002, Deheng Ye, Bo Du 0001, Jialie Shen 0001, Dacheng Tao
ICML6
2025 Ex Pede Herculem, Predicting Global Actionness Curve from Local Clips
abstract
Dense multi-label action detection in untrimmed long videos is a formidable task, with end-to-end training particularly challenging due to computational constraints, typically involving separate stages of off-the-shelf feature extraction and subsequent global modeling for action prediction.Existing methods fail to optimize all modules jointly for better performance. We introduce FreETAD, a Frequency-based End-to-end Temporal Action Detection approach, which shifts the focus from local actionness scores to frequency component estimation. Using the short-term Fourier Transform, FreETAD reconstructs the global action curve seamlessly. With a DETR-like decoder and frequency-encoded vectors for queries, it enhances multi-scale time-frequency interactions. FreETAD leverages end-to-end training effectively, boosting the mAP by 1.5% on Charades and 2.7% on MultiTHUMOS.
Xu Chen 0053, Yang Li 0251, Yahong Han, Jialie Shen 0001
ACM Multimedia4
2025 Ear with Eye: Lightweight Multimodal Audio-Visual Network Inspired by Bionic Structures
Xuanming Jiang, Baoyi An 0001, Zhengwei Zou, Dingyu Nie, Jialie Shen 0001, Xueming Qian, Guoshuai Zhao 0001
ACM Multimedia5
2025 Sequence-augmented Conversational Recommendation System Based on Diffusion Models for Personalized Cultural Exploration
abstract
Conversational Recommendation Systems (CRS) play a pivotal role in personalized cultural discovery by guiding user attention and mitigating information overload through interactive dialogue. However, existing Transformer-based CRS models predominantly focus on token-level generation, limiting their ability to capture sentence-level semantic interaction patterns. Furthermore, while user preferences are often inferred through interaction entities among these entities, current approaches typically overlook the sequential dependencies, either semantic or ID-based, which are crucial for accurate and context-aware recommendations. To overcome these limitations, we propose SDCRS, a sequence-augmented CRS based on diffusion models, which integrates sentence-level and entity-level sequential modeling to enhance the response generation and recommendation modules. By integrating diffusion mechanisms, SDCRS not only improves the diversity of generated responses but also enhances user preference modeling, particularly under cold-start conditions. Comprehensive experiments clearly demonstrate that SDCRS achieves superior performance over all baselines.
Lun Tan, Jiakui Shen, Yunqi Mi, Guoshuai Zhao 0001, Jialie Shen 0001, Xueming Qian
MMAsia6
2025 Value-Guided Decision Transformer: A Unified Reinforcement Learning Framework for Online and Offline Settings
abstract
The Conditional Sequence Modeling (CSM) paradigm, benefiting from the transformer's powerful distribution modeling capabilities, has demonstrated considerable promise in Reinforcement Learning (RL) tasks. However, much of the work has focused on applying CSM to single online or offline settings, with the general architecture rarely explored. Additionally, existing methods primarily focus on deterministic trajectory modeling, overlooking the randomness of state transitions and the diversity of future trajectory distributions. Fortunately, value-based methods offer a viable solution for CSM, further bridging the potential gap between offline and online RL. In this paper, we propose Value-Guided Decision Transformer (VDT), which leverages value functions to perform advantage-weighting and behavior regularization on the Decision Transformer (DT), guiding the policy toward upper-bound optimal decisions during the offline training phase. In the online tuning phase, VDT further integrates value-based policy improvement with behavior cloning under the CSM architecture through limited interaction and data collection, achieving performance improvement within minimal timesteps. The predictive capability of value functions for future returns is also incorporated into the sampling process. Our method achieves competitive performance on various standard RL benchmarks, providing a feasible solution for developing CSM architectures in general scenarios. Code is available at here.
Hongling Zheng, Li Shen 0008, Yong Luo 0002, Deheng Ye, Shuhan Xu, Bo Du 0001, Jialie Shen 0001, Dacheng Tao
NeurIPS7
2025 Optimizing domain-generalizable ReID through non-parametric normalization
Amran Bhuiyan, Aijun An, Jimmy Huang 0001, Jialie Shen 0001
Pattern Recognit.4
2025 Introduction to the Special Issue on Deep Multimodal Generation and Retrieval
abstract
This editorial introduces the Special Issue on Deep Multimodal Generation and Retrieval, hosted by the ACM Transactions on Multimedia Computing, Communications, and Applications in 2024. Information generation (IG) and information retrieval (IR) are two key representative approaches of information acquisition, i.e., producing content either via generation or via retrieval. While traditional IG and IR have achieved great success within the scope of languages, the under-utilization of varied data sources in different modalities (i.e., text, images, audio, and video) would hinder IG and IR techniques from giving the full advances and thus limit the applications in the real world. Knowing the fact that our world is replete with multimedia information, this special issue encourages the development of deep multimodal learning for the research of IG and IR. Benefiting from a variety of data types and modalities, some of the latest prevailing techniques are extensively invented to show great facilitation in multimodal IG and IR learning. With this special issue, we encourage explorations in Deep Multimodal Generation and Retrieval, providing a platform for researchers to share insights and advancements in this rapidly evolving domain. Each article provides novel insights into areas and challenges such as Multimodal Semantics Understanding, Generative Models for Vision Synthesis, Multimodal Information Retrieval, Explainable and Reliable Multimodal Learning. We summarize the main contributions of the included works and emphasize their role in advancing the field of multimodal generation or via retrieval. Finally, we discuss ongoing challenges and future opportunities in this rapidly evolving domain, particularly in the context of large foundational models.
Hao Fei 0001, Wei Ji 0008, Yinwei Wei, Zhedong Zheng, Jialie Shen 0001, Alan Hanjalic, Roger Zimmermann
ACM Trans. Multim. Comput. Commun. Appl.5
2025 GSyncCode: Geometry Synchronous Hidden Code for One-step Photography Decoding
abstract
Invisible hyperlinks and hidden barcodes have recently emerged as a hot topic in offline-to-online messaging, where an invisible message or barcode is embedded in an image and can be decoded via camera shooting. Current schemes involve a two-step decoding process: starting with vertex localization of the embedded region to correct the perspective distortion introduced by shooting, followed by decoding the message from the corrected region. However, vertex localization can be complex and time-consuming, which affects the efficiency and accuracy of message decoding. To address this issue, this article proposes a geometry synchronous decoding scheme called GSyncCode, allowing for one-step extraction of a Data Matrix code from the photograph. Instead of correction before decoding, GSyncCode directly decodes a geometry-transformed Data Matrix that is synchronized with the embedded region. A barcode scanner is then used to efficiently retrieve messages. We design a Haar transform-based encoder HaarUNet and a HaarLoss visual function to select the key component of the Data Matrix for embedding. They improve the visual quality of the embedded image by reducing redundant embedding signals. Extensive simulated and real-world experiments demonstrate the superiority of GSyncCode in both decoding efficiency and accuracy. Our codes are published at: https://github.com/zcx-language/GSyncCode .
Chengxin Zhao, Jialie Shen 0001, Han Fang 0004, Sijing Xie, Yaokun Fang, Zongyi Li, Ping Li 0021
ACM Trans. Multim. Comput. Commun. Appl.3
2024 On Which Nodes Does GCN Fail? Enhancing GCN From the Node Perspective
abstract
The label smoothness assumption is at the core of Graph Convolutional Networks (GCNs): nodes in a local region have similar labels. Thus, GCN performs local feature smoothing operation to adhere to this assumption. However, there exist some nodes whose labels obtained by feature smoothing conflict with the label smoothness assumption. We find that the label smoothness assumption and the process of feature smoothing are both problematic on these nodes, and call these nodes out of GCN's control (OOC nodes). In this paper, first, we design the corresponding algorithm to locate the OOC nodes, then we summarize the characteristics of OOC nodes that affect their representation learning, and based on their characteristics, we present DaGCN, an efficient framework that can facilitate the OOC nodes. Extensive experiments verify the superiority of the proposed method and demonstrate that current advanced GCNs are improvements specifically on OOC nodes; the remaining nodes under GCN's control (UC nodes) are already optimally represented by vanilla GCN on most datasets.
Jincheng Huang 0005, Jialie Shen 0001, Xiaoshuang Shi, Xiaofeng Zhu 0001
ICML2
2024 Decomposed Prompt Decision Transformer for Efficient Unseen Task Generalization
abstract
Multi-task offline reinforcement learning aims to develop a unified policy for diverse tasks without requiring real-time interaction with the environment. Recent work explores sequence modeling, leveraging the scalability of the transformer architecture as a foundation for multi-task learning. Given the variations in task content and complexity, formulating policies becomes a challenging endeavor, requiring careful parameter sharing and adept management of conflicting gradients to extract rich cross-task knowledge from multiple tasks and transfer it to unseen tasks. In this paper, we propose the Decomposed Prompt Decision Transformer (DPDT) that adopts a two-stage paradigm to efficiently learn prompts for unseen tasks in a parameter-efficient manner. We incorporate parameters from pre-trained language models (PLMs) to initialize DPDT, thereby providing rich prior knowledge encoded in language models. During the decomposed prompt tuning phase, we learn both cross-task and task-specific prompts on training tasks to achieve prompt decomposition. In the test time adaptation phase, the cross-task prompt, serving as a good initialization, were further optimized on unseen tasks through test time adaptation, enhancing the model's performance on these tasks. Empirical evaluation on a series of Meta-RL benchmarks demonstrates the superiority of our approach. The project is available at https://github.com/ruthless-man/DPDT.
Hongling Zheng, Li Shen 0008, Yong Luo 0002, Tongliang Liu, Jialie Shen 0001, Dacheng Tao
NeurIPS5
2024 LM-Metric: Learned pair weighting and contextual memory for deep metric learning
Shiyang Yan, Xinyao Shu, Jialie Shen 0001
Pattern Recognit.5
2024 Improving Conversational Recommendation System Through Personalized Preference Modeling and Knowledge Graph
abstract
Conversational recommendation systems (CRS) can actively discover users’ preferences and perform recommendations during conversations. The majority of works on CRS tend to focus on a single conversation and dig it using knowledge graphs, language models, etc. However, they often overlook the abundant and rich preference information that exists in the user's historical conversations. Meanwhile, end-to-end generation of recommendation results may lead to a decrease in recommendation quality. In this work, we propose a personalized conversational recommendation system infused with historical interaction information. This framework leverages users’ preferences extracted from their historical conversations and integrates them with the users’ preferences in current conversations. We find that this contributes to higher accuracy in recommendations and fewer recommendation turns. Moreover, we improve the interactive pattern between the recommendation module and the dialogue generation module by utilizing the slot filling method. This enables the results inferred by the recommendation module to be integrated into the conversation naturally and accurately. Our experiments on the benchmark dataset demonstrate that our model significantly outperforms the state-of-the-art methods in the evaluation of recommendations and dialogue generation.
Guoshuai Zhao 0001, Tengjiao Li, Jialie Shen 0001, Xueming Qian
IEEE Trans. Knowl. Data Eng.4
2024 Domain-Oriented Knowledge Transfer for Cross-Domain Recommendation
abstract
Cross-Domain Recommendation (CDR) aims to alleviate the cold-start problem by transferring knowledge from a data-rich domain (source domain) to a data-sparse domain (target domain), where knowledge needs to be transferred through a bridge connecting the two domains. Therefore, constructing a bridge connecting the two domains is fundamental for enabling cross-domain recommendation. However, existing CDR methods often overlook the valuable of natural relationships between items in connecting the two domains. To address this issue, we propose DKTCDR: a Domain-oriented Knowledge Transfer method for Cross-Domain Recommendation. In DKTCDR, We leverages the rich relationships between items in a cross-domain knowledge graph as bridges to facilitate both intra- and inter-domain knowledge transfer. Additionally, we design a cross-domain knowledge transfer strategy to enhance inter-domain knowledge transfer. Furthermore, we integrate the semantic modality information of items with the knowledge graph modality information to enhance item modeling. To support our investigation, we construct two high-quality cross-domain recommendation datasets, each containing a cross-domain knowledge graph. Our experimental results on these datasets validate the effectiveness of our proposed method. Source code is available athttps://github.com/zxxxl123/DKTCDR.
Guoshuai Zhao 0001, Jialie Shen 0001, Xueming Qian
IEEE Trans. Multim.4
2024 GRLC: Graph Representation Learning With Constraints
abstract
Contrastive learning has been successfully applied in unsupervised representation learning. However, the generalization ability of representation learning is limited by the fact that the loss of downstream tasks (e.g., classification) is rarely taken into account while designing contrastive methods. In this article, we propose a new contrastive-based unsupervised graph representation learning (UGRL) framework by 1) maximizing the mutual information (MI) between the semantic information and the structural information of the data and 2) designing three constraints to simultaneously consider the downstream tasks and the representation learning. As a result, our proposed method outputs robust low-dimensional representations. Experimental results on 11 public datasets demonstrate that our proposed method is superior over recent state-of-the-art methods in terms of different downstream tasks. Our code is available at https://github.com/LarryUESTC/GRLC.
Yujie Mo, Jie Xu 0044, Jialie Shen 0001, Xiaoshuang Shi, Xiaoxiao Li 0001, Heng Tao Shen, Xiaofeng Zhu 0001
IEEE Trans. Neural Networks Learn. Syst.4
2023 Disentangled Multiplex Graph Representation Learning
abstract
Unsupervised multiplex graph representation learning (UMGRL) has received increasing interest, but few works simultaneously focused on the common and private information extraction. In this paper, we argue that it is essential for conducting effective and robust UMGRL to extract complete and clean common information, as well as more-complementarity and less-noise private information. To achieve this, we first investigate disentangled representation learning for the multiplex graph to capture complete and clean common information, as well as design a contrastive constraint to preserve the complementarity and remove the noise in the private information. Moreover, we theoretically analyze that the common and private representations learned by our method are provably disentangled and contain more task-relevant and less task-irrelevant information to benefit downstream tasks. Extensive experiments verify the superiority of the proposed method in terms of different downstream tasks.
Yujie Mo, Yajie Lei, Jialie Shen 0001, Xiaoshuang Shi, Heng Tao Shen, Xiaofeng Zhu 0001
ICML3
2023 Rethinking the Localization in Weakly Supervised Object Localization
abstract
Weakly supervised object localization (WSOL) is one of the most popular and challenging tasks in computer vision. This task is to localize the objects in the images given only the image-level supervision. Recently, dividing WSOL into two parts (class-agnostic object localization and object classification) has become the state-of-the-art pipeline for this task. However, existing solutions under this pipeline usually suffer from the following drawbacks: 1) they are not flexible since they can only localize one object for each image due to the adopted single-class regression (SCR) for localization; 2) the generated pseudo bounding boxes may be noisy, but the negative impact of such noise is not well addressed. To remedy these drawbacks, we first propose to replace SCR with a binary-class detector (BCD) for localizing multiple objects, where the detector is trained by discriminating the foreground and background. Then we design a weighted entropy (WE) loss using the unlabeled data to reduce the negative impact of noisy bounding boxes. Extensive experiments on the popular CUB-200-2011 and ImageNet-1K datasets demonstrate the effectiveness of our method.
Rui Xu 0031, Yong Luo 0002, Han Hu 0003, Bo Du 0001, Jialie Shen 0001, Yonggang Wen 0001
ACM Multimedia5
2023 LGViT: Dynamic Early Exiting for Accelerating Vision Transformer
abstract
Recently, the efficient deployment and acceleration of powerful vision transformers (ViTs) on resource-limited edge devices for providing multimedia services have become attractive tasks. Although early exiting is a feasible solution for accelerating inference, most works focus on convolutional neural networks (CNNs) and transformer models in natural language processing (NLP). Moreover, the direct application of early exiting methods to ViTs may result in substantial performance degradation. To tackle this challenge, we systematically investigate the efficacy of early exiting in ViTs and point out that the insufficient feature representations in shallow internal classifiers and the limited ability to capture target semantic information in deep internal classifiers restrict the performance of these methods. We then propose an early exiting framework for general ViTs termed LGViT, which incorporates heterogeneous exiting heads, namely, local perception head and global aggregation head, to achieve an efficiency-accuracy trade-off. In particular, we develop a novel two-stage training scheme, including end-to-end training and self-distillation with the backbone frozen to generate early exiting ViTs, which facilitates the fusion of global and local information extracted by the two types of heads. We conduct extensive experiments using three popular ViT backbones on three vision datasets. Results demonstrate that our LGViT can achieve competitive performance with approximately 1.8 × speed-up.
Guanyu Xu, Li Shen 0008, Han Hu 0003, Yong Luo 0002, Jialie Shen 0001
ACM Multimedia7
2023 Fast LoG SIFT Keypoint Detector
abstract
Scale-invariant feature transform (SIFT) is a classical computer vision technique for scale-invariant keypoint detection and feature extraction. SIFT exhibits invariance to various transformations such as scale, rotation, noise, and illumination, making it applicable in a wide range of applications like object recognition, image matching and stitching, environment mapping, navigation, robotics, camera calibration, and more. A key contribution of SIFT is its utilization of the Difference-of-Gaussian (DoG) feature pyramid, which approximates the scale-space response of the Laplacian-of-Gaussian (LoG) filter. The DoG feature pyramid is computed by taking the separable Gaussian filtering and stacking the difference of Gaussian blurred images. In this paper, we propose a novel approach called “Fast LoG” filtering, which offers direct computation of the LoG filter to model the scale-space response solution. The “Fast LoG” filter is achieved by decomposing the LoG filter into two separable filters via SVD, and the scale-space response is computed by a direct polynomial fitting and differentiation, which is analytically more accurate. The polynomial fitting and differentiation only happen after the LoG peak strength thresholding, therefore the overall complexity is low compared with the DoG-based SIFT. The experimental results show that the keypoint generated by the Fast LoG method matches the SIFT keypoints, and per-pixel filtering complexity is lower.
Paras Maharjan, Lyle Vanfossan, Zhu Li 0001, Jialie Shen 0001
MMSP4
2023 Neighbor-Guided Consistent and Contrastive Learning for Semi-Supervised Action Recognition
abstract
Semi-supervised learning has been well established in the area of image classification but remains to be explored in video-based action recognition. FixMatch is a state-of-the-art semi-supervised method for image classification, but it does not work well when transferred directly to the video domain since it only utilizes the single RGB modality, which contains insufficient motion information. Moreover, it only leverages highly-confident pseudo-labels to explore consistency between strongly-augmented and weakly-augmented samples, resulting in limited supervised signals, long training time, and insufficient feature discriminability. To address the above issues, we propose neighbor-guided consistent and contrastive learning (NCCL), which takes both RGB and temporal gradient (TG) as input and is based on the teacher-student framework. Due to the limitation of labelled samples, we first incorporate neighbors information as a self-supervised signal to explore the consistent property, which compensates for the lack of supervised signals and the shortcoming of long training time of FixMatch. To learn more discriminative feature representations, we further propose a novel neighbor-guided category-level contrastive learning term to minimize the intra-class distance and enlarge the inter-class distance. We conduct extensive experiments on four datasets to validate the effectiveness. Compared with the state-of-the-art methods, our proposed NCCL achieves superior performance with much lower computational cost.
Jianlong Wu, Tian Gan 0002, Ning Ding 0006, Feijun Jiang, Jialie Shen 0001, Liqiang Nie
IEEE Trans. Image Process.6
2023 Efficient Query-based Black-box Attack against Cross-modal Hashing Retrieval
abstract
Deep cross-modal hashing retrieval models inherit the vulnerability of deep neural networks. They are vulnerable to adversarial attacks, especially for the form of subtle perturbations to the inputs. Although many adversarial attack methods have been proposed to handle the robustness of hashing retrieval models, they still suffer from two problems: (1) Most of them are based on the white-box settings, which is usually unrealistic in practical application. (2) Iterative optimization for the generation of adversarial examples in them results in heavy computation. To address these problems, we propose an Efficient Query-based Black-Box Attack (EQB 2 A) against deep cross-modal hashing retrieval, which can efficiently generate adversarial examples for the black-box attack. Specifically, by sending a few query requests to the attacked retrieval system, the cross-modal retrieval model stealing is performed based on the neighbor relationship between the retrieved results and the query, thus obtaining the knockoffs to substitute the attacked system. A multi-modal knockoffs-driven adversarial generation is proposed to achieve efficient adversarial example generation. While the entire network training converges, EQB 2 A can efficiently generate adversarial examples by forward-propagation with only given benign images. Experiments show that EQB 2 A achieves superior attacking performance under the black-box setting.
Lei Zhu 0002, Tianshi Wang 0001, Jingjing Li 0001, Zheng Zhang 0006, Jialie Shen 0001, Xinhua Wang 0003
ACM Trans. Inf. Syst.5
2023 Generative Metric Learning for Adversarially Robust Open-world Person Re-Identification
abstract
The vulnerability of re-identification (re-ID) models under adversarial attacks is of significant concern as criminals may use adversarial perturbations to evade surveillance systems. Unlike a closed-world re-ID setting (i.e., a fixed number of training categories), a reliable re-ID system in the open world raises the concern of training a robust yet discriminative classifier, which still shows robustness in the context of unknown examples of an identity. In this work, we improve the robustness of open-world re-ID models by proposing a generative metric learning approach to generate adversarial examples that are regularized to produce robust distance metric. The proposed approach leverages the expressive capability of generative adversarial networks to defend the re-ID models against feature disturbance attacks. By generating the target people variants and sampling the triplet units for metric learning, our learned distance metrics are regulated to produce accurate predictions in the feature metric space. Experimental results on the three re-ID datasets, i.e., Market-1501, DukeMTMC-reID, and MSMT17 demonstrate the robustness of our method.
Deyin Liu, Lin Wu 0001, Richang Hong, ZongYuan Ge, Jialie Shen 0001, Farid Boussaïd, Mohammed Bennamoun
ACM Trans. Multim. Comput. Commun. Appl.5
2022 Pseudo-Pair Based Self-Similarity Learning for Unsupervised Person Re-Identification
abstract
Person re-identification (re-ID) is of great importance to video surveillance systems by estimating the similarity between a pair of cross-camera person shorts. Current methods for estimating such similarity require a large number of labeled samples for supervised training. In this paper, we present a pseudo-pair based self-similarity learning approach for unsupervised person re-ID without human annotations. Unlike conventional unsupervised re-ID methods that use pseudo labels based on global clustering, we construct patch surrogate classes as initial supervision, and propose to assign pseudo labels to images through the pairwise gradient-guided similarity separation. This can cluster images in pseudo pairs, and the pseudos can be updated during training. Based on pseudo pairs, we propose to improve the generalization of similarity function via a novel self-similarity learning:it learns local discriminative features from individual images via intra-similarity, and discovers the patch correspondence across images via inter-similarity. The intra-similarity learning is based on channel attention to detect diverse local features from an image. The inter-similarity learning employs a deformable convolution with a non-local block to align patches for cross-image similarity. Experimental results on several re-ID benchmark datasets demonstrate the superiority of the proposed method over the state-of-the-arts.
Lin Wu 0001, Deyin Liu, Dapeng Chen, ZongYuan Ge, Farid Boussaïd, Mohammed Bennamoun, Jialie Shen 0001
IEEE Trans. Image Process.8
2022 Multimodal Marketing Intent Analysis for Effective Targeted Advertising
abstract
People’s daily information sharing and acquisition through the Internet has become more and more popular. The comprehensive multimodal marketing advertorial generated by ‘We Media’ accounts besides the normal social news is gaining its importance on social media platforms. In order to achieve effective advertising, the marketing intent understanding is a key step towards generating targeted advertising strategies (push advertorials to specific people at a specific time). However, advertorials in real are usually designed to pretend as normal social news with a wide range of contents. This poses big challenges to the platforms on accurately recognizing and analyzing the marketing intents behind the advertorials. As a pioneering study, we address this new problem of multimodal-based marketing intent analysis and answer three core questions: (1) does a piece of social news contain marketing intent? (2) what is the topic of marketing intent? (3) what is the extent of marketing intent? Towards this end, we propose a novel Multimodal-based Marketing Intent Analysis scheme (MMIA) to estimate the marketing intent embedded in the multimodal contents. Specifically, a novel supervised neural autoregressive model (SmiDocNADE) is proposed to enhance the discriminative capacity of the learned hidden features so that a single system is capable of solving the three questions. In order to effectively model inter-correlations between images and text in advertorials, we fuse multimodal data and extract features by Graph Convolution Networks as an enhancement to SmiDocNADE. The extensive evaluations demonstrate the advantages of our proposed system in multimodal-based marketing intent analysis from multiple aspects.
Lu Zhang 0062, Jialie Shen 0001, Jian Zhang 0002, Jingsong Xu, Zhibin Li 0002, Yazhou Yao, Litao Yu
IEEE Trans. Multim.2
2022 Unsupervised Image and Text Fusion for Travel Information Enhancement
abstract
With the explosive growth of the shared information on social media platforms, people are increasingly interested in sharing and making their travel plans by referring to others’ travel experiences. However, different social media sources render the heterogeneity of these valuable data, bringing difficulties for data collection and fusion. Thus, facing massive information online, one of the biggest challenges to enhance travel information is how to integrate and match these multi-source data without clear labels. In this paper, we propose an unsupervised method to fuse and match images and travelogues. We first use the three textual components (title, tag, and description) of the descriptive texts of images as three criteria to embed travelogues and the descriptive texts of images, and further introduce images into our method by joint embedding texts and images. Finally, a multiple kernel clustering approach is adopted for matching travelogues and images. Extensive experiments conducted on the real dataset crawled from two websites (Flickr and TripAdvisor) demonstrate the effectiveness and robustness of our proposed method.
Lu Zhang 0062, Jingsong Xu, Yongshun Gong, Litao Yu, Jian Zhang 0002, Jialie Shen 0001
IEEE Trans. Multim.6
2021 Inferring Emotion from Large-scale Internet Voice Data: A Semi-supervised Curriculum Augmentation based Deep Learning Approach
abstract
Effective emotion inference from user queries helps to give a more personified response for Voice Dialogue Applications(VDAs). The tremendous amounts of VDA users bring in diverse emotion expressions. How to achieve a high emotion inferring performance from large-scale Internet Voice Data in VDAs? Traditionally, researches on speech emotion recognition are based on acted voice datasets, which have limited speakers but strong and clear emotion expressions. Inspired by this, in this paper, we propose a novel approach to leverage acted voice data with strong emotion expressions to enhance large-scale unlabeled internet voice data with diverse emotion expressions for emotion inferring. Specifically, we propose a novel semi-supervised multi-modal curriculum augmentation deep learning framework. First, to learn more general emotion cues, we adopt a curriculum learning based epoch-wise training strategy, which trains our model guided by strong and balanced emotion samples from acted voice data and sub-sequently leverages weak and unbalanced emotion samples from internet voice data.Second, to employ more diverse emotion expressions, we design a Multi-path Mix-match Multimodal Deep Neural Network(MMMD), which effectively learns feature representations for multiple modalities and trains labeled and unlabeled data in hybrid semi-supervised methods for superior generalization and robustness. Experiments on an internet voice dataset with 500,000 utterances show our method outperforms (+10.09% in terms of F1) several alternative baselines, while an acted corpus with 2,397 utterances contributes 4.35%. To further compare our method with state-of-the-art techniques in traditionally acted voice datasets, we also conduct experiments on public dataset IEMOCAP. The results reveal the effectiveness of the proposed approach.
Suping Zhou, Jia Jia 0001, Zhiyong Wu 0001, Wei Chen 0071, Shuo Huang 0005, Jialie Shen 0001
AAAI9
2021 Incorporating Multimodal Cues for Advertorial Discovery
abstract
Commercial advertorials shared on websites are usually designed to pretend as normal social news for commercial benefits. The analysis of the commercial intents embedded in advertorials can greatly help media platforms personalize content. However, commercial intents are not only concealed in news texts but also conveyed by news images explicitly or implicitly. Consequently, how to effectively extract and incorporate the crucial cues of multiple modalities has been emerging as an important but challenging problem. Motivated by this observation, we propose a framework Multimodal Advertorial Discovery Model (MADM) to estimate the commercial intents embedded in the multimodal social news. Specifically, a novel Cross-graph Fusion (CGF) strategy is developed to achieve a soft assignment to incorporate images and text and generate comprehensive multimodal representations. The extensive evaluations demonstrate the superiority of our proposed system in multimodal-based advertorial detection and analysis.
Lu Zhang 0062, Jian Zhang 0002, Jialie Shen 0001, Jingsong Xu, Zhibin Li 0002, Litao Yu
ICME3
2021 Latent label mining for group activity recognition in basketball videos
abstract
Abstract Motion information has been widely exploited for group activity recognition in sports video. However, in order to model and extract the various motion information between the adjacent frames, existing algorithms only use the coarse video‐level labels as supervision cues. This may lead to the ambiguity of extracted features and the omission of changing rules of motion patterns that are also important sports video recognition. In this paper, a latent label mining strategy for group activity recognition in basketball videos is proposed. The authors' novel strategy allows them to obtain the latent labels set for marking different frames in an unsupervised way, and build the frame‐level and video‐level representations with two separate levels of supervision signal. Firstly, the latent labels of motion patterns are digged using the unsupervised hierarchical clustering technique. The generated latent labels are then taken as the frame‐level supervision signal to train a deep CNN for the frame‐level features extraction. Lastly, the frame‐level features are fed into an LSTM network to build the spatio‐temporal representation for group activity recognition. Experimental results on the public NCAA dataset demonstrate that the proposed algorithm achieves state‐of‐the‐art performance.
Lifang Wu, Ye Xiang, Meng Jian, Jialie Shen 0001
IET Image Process.5
2021 BBAS: Towards large scale effective ensemble adversarial attacks against deep neural network learning
Jialie Shen 0001, Neil Robertson 0001
Inf. Sci.1
2021 A new context-aware approach for automatic Chinese poetry generation
Shanliang Zhu, Jun Shen 0001, Jialie Shen 0001, Shuguo Yang, PengCheng Xiong
Knowl. Based Syst.5
2021 Global motion estimation with iterative optimization-based independent univariate model for action recognition
Lifang Wu, Meng Jian, Jialie Shen 0001, Xianglong Lang
Pattern Recognit.4
2021 Person Retrieval in Surveillance Videos Via Deep Attribute Mining and Reasoning
abstract
Person retrieval largely relies on the appearance features of pedestrians. This task is rather more difficult in surveillance videos due to the limitations of extracting robust appearance features brought by the cross-view and cross-camera data with lower image resolution, motion blur, occlusion and other kinds of image degradation. To build up a more reliable person retrieval system, recent works introduced appearance attribute models to describe and distinguish different persons with high-level semantic concepts. Despite the progress of previous works, the value of utilizing appearance attributes is still under-explored. On one hand, existing methods lack for concise and precise attribute representations that are specific for each attribute category and, in the meantime, are able to filter noisy information in irrelevant spatial locations and useless patterns. On the other hand, correlation and reasoning between different attributes are neglected, which could generate more useful information and add more robustness to the retrieval system. In this paper, we propose an Attribute Mining and Reasoning (AMR) framework which is capable to handle the issues in question. The AMR makes better use of appearance attributes with two main components. First, the AMR disentangles the representations of different attributes by localizing their spatial positions and identifying their effective patterns in a weakly supervised manner. To achieve more reliable localization, we propose the Attribute Localization Ensemble (ALE) module that is consisted of multiple localization heads and a voting mechanism. Second, we introduce the Attribute Reasoning (AR) module to correlate different attributes together with the global appearance features and discover their latent relations to generate more comprehensive descriptions of pedestrians. Extensive experiments on DukeMTMC-ReID and Market-1501 datasets demonstrate the effectiveness of the proposed AMR framework as well as its superiority over the existing state-of-the-art methods. The AMR model also shows great generalization ability on the unseen CUHK03 dataset when it is only trained on Market-1501 dataset.
Yuxuan Shi 0002, Zhen Wei 0001, Jialie Shen 0001, Ping Li 0021
IEEE Trans. Multim.5
2021 Adaptive and Robust Partition Learning for Person Retrieval With Policy Gradient
abstract
Person retrieval aims at effectively matching the pedestrian images over an extensive database given a specified identity. As extracting effective features is crucial in a high-performance retrieval system, recent significant progress was achieved by part-based models that have constructed robust local representations on top of vertically striped part features. However, this kind of models use predefined partitioning strategies, making the number and size of each partition identical even when input images vary a lot. This unchangeable setting usually leads to less flexibility and robustness in capturing visual variance. The primary reason for such a negative effect is that a fixed partitioning strategy is unable to deal with (a) the significant variance from pose, illumination and viewpoint which is common in a pedestrian image dataset, and (b) also the inference error and misalignment of human bodies introduced by the prepositive pedestrian detection module or human pose estimation module. In this paper, we tackle this problem via introducing the novel Adaptive Partition Network (APN). The APN utilizes deep reinforcement learning and applies an agent to generate optimal partitioning strategies dynamically for different input images. The agent inside the APN is optimized with the policy gradient algorithm and maximizes the reward of choosing the best partition setting. By leveraging the supervision cues from the objective partitioning strategies that are generated on a set of held-out training images, the agent is trained jointly with other parts of APN, which ensures the APN's robustness and generalization ability. Extensive experimental results on multiple datasets, including CUHK03, DukeMTMC and Market-1501, demonstrate the superiority of APN over the state-of-the-art models.
Yuxuan Shi 0002, Zhen Wei 0001, Pengfei Zhu 0001, Jialie Shen 0001, Ping Li 0021
IEEE Trans. Multim.6
2020 CaMR: Towards Connotation-aware Music Retrieval on Social Media with Visual Inputs
abstract
With the ubiquitous network connectivity and the proliferation of mobile devices, people are increasingly consuming digital contents from social media driven music sharing platforms (e.g., YouTube, Soundcloud). In this paper, we study a novel problem of connotation-aware music retrieval that focuses on the connotation which expresses the implicit feeling or emotion beyond the explicit content in artworks. Our goal is to automatically retrieve relevant music on social media based on the connotation of visual inputs (e.g., images, photos) provided by the users. The problem is challenging as it requires the accurate identification of the implicit connotation from both images and music pieces, and the precise matching of the identified connotation across different data modalities. We develop a connotation-aware music retrieval (CaMR) framework to address the above challenges. Evaluation results from a real-world social media dataset demonstrate that the CaMR framework can retrieve music that is highly relevant to the connotation of the input image.
Lanyu Shang, Daniel Yue Zhang, Siamul Karim Khan, Jialie Shen 0001, Dong Wang 0002
ASONAM4
2020 Inferring Emphasis for Real Voice Data: An Attentive Multimodal Neural Network Approach
Suping Zhou, Jia Jia 0001, Wei Chen 0071, Jialie Shen 0001
MMM (2)8
2020 A New Automatic Chinese Poetry Generation Model Based On Neural Network
abstract
Chinese ancient poetry has been a favorite literary genre for thousands of years. Nowadays, it is still read and ancient Chinese poets are honored. Recently, many websites provide automatic Chinese poetry generation service based on neural network. My research goal is to improve the quality of the poetry generation, i.e., making it as close as possible to the real poetry written by poets. To achieve this goal, I propose a new context-aware Chinese poetry generation method based on sequence-to-sequence framework. I generate a new concept called keyword team which is a combination of all the keywords to capture the context of the Chinese poetry. Then I use the keyword, the keyword team and the previously generated lines to generate the current line in a poem. I have already implemented the new context-aware Chinese poetry generation model called KPG (Keyword team based Poetry Generation). I have compared this method with state-of-the-art ones using loss function and human judgement. The comprehensive evaluation result show that our proposed method outperforms the state-of-the-art ones.
PengCheng Xiong, Jialie Shen 0001
SERVICES3
2020 Learning refined attribute-aligned network with attribute selection for person re-identification
Yuxuan Shi 0002, Lei Wu 0010, Jialie Shen 0001, Ping Li 0021
Neurocomputing4
2020 Semantics-aware influence maximization in social networks
Yipeng Chen, Qiang Qu 0001, Yuanxiang Ying, Hongyan Li 0002, Jialie Shen 0001
Inf. Sci.5
2020 Enhancing Sketch-Based Image Retrieval by CNN Semantic Re-ranking
abstract
This paper introduces a convolutional neural network (CNN) semantic re-ranking system to enhance the performance of sketch-based image retrieval (SBIR). Distinguished from the existing approaches, the proposed system can leverage category information brought by CNNs to support effective similarity measurement between the images. To achieve effective classification of query sketches and high-quality initial retrieval results, one CNN model is trained for classification of sketches, another for that of natural images. Through training dual CNN models, the semantic information of both the sketches and natural images is captured by deep learning. In order to measure the category similarity between images, a category similarity measurement method is proposed. Category information is then used for re-ranking. Re-ranking operation first infers the retrieval category of the query sketch and then uses the category similarity measurement to measure the category similarity between the query sketch and each initial retrieval result. Finally, the initial retrieval results are re-ranked. The experiments on different types of SBIR datasets demonstrate the effectiveness of the proposed re-ranking method. Comparisons with other re-ranking algorithms are also given to show the proposed method's superiority. Further, compared to the baseline systems, the proposed re-ranking approach achieves significantly higher precision in the top ten different SBIR methods and datasets.
Luo Wang, Xueming Qian, Yuting Zhang 0007, Jialie Shen 0001, Xiaochun Cao
IEEE Trans. Cybern.4
2019 Understanding the Teaching Styles by an Attention based Multi-task Cross-media Dimensional Modeling
abstract
Teaching style plays an influential role in helping students to achieve academic success. In this paper, we explore a new problem of effectively understanding teachers' teaching styles. Specifically, we study 1) how to quantitatively characterize various teachers' teaching styles for various teachers and 2) how to model the subtle relationship between cross-media teaching related data (speech, facial expressions and body motions, content et al.) and teaching styles. Using the adjectives selected from more than 10,000 feedback questionnaires provided by an educational enterprise, a novel concept called Teaching Style Semantic Space (TSSS) is developed based on the pleasure-arousal dimensional theory to describe teaching styles quantitatively and comprehensively. Then a multi-task deep learning based model, Attention-based Multi-path Multi-task Deep Neural Network (AMMDNN), is proposed to accurately and robustly capture the internal correlations between cross-media features and TSSS. Based on the benchmark dataset, we further develop a comprehensive data set including 4,541 full-annotated cross-modality teaching classes. Our experimental results demonstrate that the proposed AMMDNN outperforms (+0.0842% in terms of the concordance correlation coefficient (CCC) on average) baseline methods. To further demonstrate the advantages of the proposed TSSS and our model, several interesting case studies are carried out, such as teaching styles comparison among different teachers and courses, and leveraging the proposed method for teaching quality analysis.
Suping Zhou, Jia Jia 0001, Yufeng Yin 0002, Xiang Li 0105, Zeyang Ye, Kehua Lei, Jialie Shen 0001
ACM Multimedia10
2019 NAIRS: A Neural Attentive Interpretable Recommendation System
abstract
In this paper, we develop a neural attentive interpretable recommendation system, named NAIRS. A self-attention network, as a key component of the system, is designed to assign attention weights to interacted items of a user. This attention mechanism can distinguish the importance of the various interacted items in contributing to a user profile. %, and it also provides interpretable recommendations. Based on the user profiles obtained by the self-attention network, NAIRS offers personalized high-quality recommendation. Moreover, it develops visual cues to interpret recommendations. This demo application with the implementation of NAIRS enables users to interact with a recommendation system, and it persistently collects training data to improve the system. The demonstration and experimental results show the effectiveness of NAIRS.
Shuai Yu 0002, Min Yang 0007, Baocheng Li, Qiang Qu 0001, Jialie Shen 0001
WSDM6
2019 Fashion recommendations through cross-media information retrieval
Wei Zhou 0028, P. Y. Mok 0001, Yanghong Zhou, Yangping Zhou, Jialie Shen 0001, Qiang Qu 0001, K. P. Chau
J. Vis. Commun. Image Represent.5
2019 Toward efficient indexing structure for scalable content-based music retrieval
abstract
With advancement of various information processing and storage techniques, the scale of digital music collections has been growing at very fast speed during recent decades. To support high-quality content-based retrieval over such a large volume of music data, how to develop indexing structure with good effectiveness, efficiency and scalability becomes an important research issue. However, existing techniques mainly focus on improving query efficiency. Very few approaches have been proposed to address issues related to scalability and accuracy. In this study, we address the problem via introducing a novel indexing technique called effective music indexing framework (EMIF) to facilitate scalable and accurate music retrieval. It is designed based on a “classification-and-indexing” principle and consists of two main functionality modules: (1) music classification—a novel semantic-sensitive classification to identify an input song’s category and (2) indexing module—multiple local indexing structures, one for each semantic category to reduce query response time significantly. In particular, the classification model combining linear discriminative mixture model (LDMM) and advanced score fusion scheme has been applied to estimate category of music accurately. Layered architecture enables EMIF to enjoy superior scalability and efficiency. To evaluate the approach, a set of experimental studies has been carried out using two large music test collections and the results demonstrate various advantages of EMIF over state-of-the-art approaches including efficiency, scalability and effectiveness.
Jialie Shen 0001, Tao Mei 0001, Qiang Qu 0001, Dacheng Tao, Yong Rui
Multim. Syst.1
2018 Unsupervised Deep Hashing via Binary Latent Factor Models for Large-scale Cross-modal Retrieval
abstract
Despite its great success, matrix factorization based cross-modality hashing suffers from two problems: 1) there is no engagement between feature learning and binarization; and 2) most existing methods impose the relaxation strategy by discarding the discrete constraints when learning the hash function, which usually yields suboptimal solutions. In this paper, we propose a novel multimodal hashing framework, referred as Unsupervised Deep Cross-Modal Hashing (UDCMH), for multimodal data search in a self-taught manner via integrating deep learning and matrix factorization with binary latent factor models. On one hand, our unsupervised deep learning framework enables the feature learning to be jointly optimized with the binarization. On the other hand, the hashing system based on the binary latent factor models can generate unified binary codes by solving a discrete-constrained objective function directly with no need for a relaxation step. Moreover, novel Laplacian constraints are incorporated into the objective function, which allow to preserve not only the nearest neighbors that are commonly considered in the literature but also the farthest neighbors of data, even if the semantic labels are not available. Extensive experiments on multiple datasets highlight the superiority of the proposed framework over several state-of-the-art baselines.
Gengshen Wu, Zijia Lin, Jungong Han, Li Liu 0004, Guiguang Ding, Baochang Zhang 0001, Jialie Shen 0001
IJCAI7
2018 Learning to rank images for complex queries in concept-based search
Chaoran Cui, Jialie Shen 0001, Zhumin Chen, Shuaiqiang Wang, Jun Ma 0001
Neurocomputing2
2017 Exploiting Music Play Sequence for Music Recommendation
abstract
Users leave digital footprints when interacting with various music streaming services. Music play sequence, which contains rich information about personal music preference and song similarity, has been largely ignored in previous music recommender systems. In this paper, we explore the effects of music play sequence on developing effective personalized music recommender systems. Towards the goal, we propose to use word embedding techniques in music play sequences to estimate the similarity between songs. The learned similarity is then embedded into matrix factorization to boost the latent feature learning and discovery. Furthermore, the proposed method only considers the k-nearest songs (e.g., k = 5) in the learning process and thus avoids the increase of time complexity. Experimental results on two public datasets demonstrate that our methods could significantly improve the performance of both rating prediction and top-n recommendation tasks.
Zhiyong Cheng 0001, Jialie Shen 0001, Lei Zhu 0002, Mohan Kankanhalli, Liqiang Nie
IJCAI2
2017 Unsupervised Deep Video Hashing with Balanced Rotation
abstract
Recently, hashing video contents for fast retrieval has received increasing attention due to the enormous growth of online videos. As the extension of image hashing techniques, traditional video hashing methods mainly focus on seeking the appropriate video features but pay little attention to how the video-specific features can be leveraged to achieve optimal binarization. In this paper, an end-to-end hashing framework, namely Unsupervised Deep Video Hashing (UDVH), is proposed, where feature extraction, balanced code learning and hash function learning are integrated and optimized in a self-taught manner. Particularly, distinguished from previous work, our framework enjoys two novelties: 1) an unsupervised hashing method that integrates the feature clustering and feature binarization, enabling the neighborhood structure to be preserved in the binary space; 2) a smart rotation applied to the video-specific features that are widely spread in the low-dimensional space such that the variance of dimensions can be balanced, thus generating more effective hash codes. Extensive experiments have been performed on two real-world datasets and the results demonstrate its superiority, compared to the state-of-the-art video hashing methods. To bootstrap further developments, the source code will be made publically available.
Gengshen Wu, Li Liu 0004, Guiguang Ding, Jungong Han, Jialie Shen 0001, Ling Shao 0001
IJCAI6
2017 Dynamic Multi-View Hashing for Online Image Retrieval
abstract
Advanced hashing technique is essential to facilitate effective large scale online image organization and retrieval, where image contents could be frequently changed. Traditional multi-view hashing methods are developed based on batch-based learning, which leads to very expensive updating cost. Meanwhile, existing online hashing methods mainly focus on single-view data and thus can not achieve promising performance when searching real online images, which are multiple view based data. Further, both types of hashing methods can only produce hash code with fixed length. Consequently they suffer from limited capability to comprehensive characterization of streaming image data in the real world. In this paper, we propose dynamic multi-view hashing (DMVH), which can adaptively augment hash codes according to dynamic changes of image. Meanwhile, DMVH leverages online learning to generate hash codes. It can increase the code length when current code is not able to represent new images effectively. Moreover, to gain further improvement on overall performance, each view is assigned with a weight, which can be efficiently updated during the online learning process. In order to avoid the frequent updating of code length and view weights, an intelligent buffering scheme is also specifically designed to preserve significant data to maintain good effectiveness of DMVH. Experimental results on two real-world image datasets demonstrate superior performance of DWVH over several state-of-the-art hashing methods.
Liang Xie 0001, Jialie Shen 0001, Jungong Han, Lei Zhu 0002, Ling Shao 0001
IJCAI2
2017 Exploring User-Specific Information in Music Retrieval
abstract
With the advancement of mobile computing technology and cloud-based streaming music service, user-centered music retrieval has become increasingly important. User-specific information has a fundamental impact on personal music preferences and interests. However, existing research pays little attention to the modeling and integration of user-specific information in music retrieval algorithms/models to facilitate music search. In this paper, we propose a novel model, named User-Information-Aware Music Interest Topic (UIA-MIT) model. The model is able to effectively capture the influence of user-specific information on music preferences, and further associate users' music preferences and search terms under the same latent space. Based on this model, a user information aware retrieval system is developed, which can search and re-rank the results based on age- and/or gender-specific music preferences. A comprehensive experimental study demonstrates that our methods can significantly improve the search accuracy over existing text-based music retrieval methods.
Zhiyong Cheng 0001, Jialie Shen 0001, Liqiang Nie, Tat-Seng Chua, Mohan Kankanhalli
SIGIR2
2017 Version-sensitive mobile App recommendation
Da Cao, Liqiang Nie, Xiangnan He 0001, Xiaochi Wei, Jialie Shen 0001, Shunxiang Wu, Tat-Seng Chua
Inf. Sci.5
2017 Social tag relevance learning via ranking-oriented neighbor voting
Chaoran Cui, Jialie Shen 0001, Jun Ma 0001, Tao Lian
Multim. Tools Appl.2
2017 Exploring Representativeness and Informativeness for Active Learning
abstract
How can we find a general way to choose the most suitable samples for training a classifier? Even with very limited prior information? Active learning, which can be regarded as an iterative optimization procedure, plays a key role to construct a refined training set to improve the classification performance in a variety of applications, such as text analysis, image recognition, social network modeling, etc. Although combining representativeness and informativeness of samples has been proven promising for active sampling, state-of-the-art methods perform well under certain data structures. Then can we find a way to fuse the two active sampling criteria without any assumption on data? This paper proposes a general active learning framework that effectively fuses the two criteria. Inspired by a two-sample discrepancy problem, triple measures are elaborately designed to guarantee that the query samples not only possess the representativeness of the unlabeled data but also reveal the diversity of the labeled data. Any appropriate similarity measure can be employed to construct the triple measures. Meanwhile, an uncertain measure is leveraged to generate the informativeness criterion, which can be carried out in different ways. Rooted in this framework, a practical active learning algorithm is proposed, which exploits a radial basis function together with the estimated probabilities to construct the triple measures and a modified best-versus-second-best strategy to construct the uncertain measure, respectively. Experimental results on benchmark datasets demonstrate that our algorithm consistently achieves superior performance over the state-of-the-art active learning algorithms.
Bo Du 0001, Zengmao Wang, Lefei Zhang, Liangpei Zhang 0001, Wei Liu 0005, Jialie Shen 0001, Dacheng Tao
IEEE Trans. Cybern.6
2017 Unsupervised Topic Hypergraph Hashing for Efficient Mobile Image Retrieval
abstract
Hashing compresses high-dimensional features into compact binary codes. It is one of the promising techniques to support efficient mobile image retrieval, due to its low data transmission cost and fast retrieval response. However, most of existing hashing strategies simply rely on low-level features. Thus, they may generate hashing codes with limited discriminative capability. Moreover, many of them fail to exploit complex and high-order semantic correlations that inherently exist among images. Motivated by these observations, we propose a novel unsupervised hashing scheme, called topic hypergraph hashing (THH), to address the limitations. THH effectively mitigates the semantic shortage of hashing codes by exploiting auxiliary texts around images. In our method, relations between images and semantic topics are first discovered via robust collective non-negative matrix factorization. Afterwards, a unified topic hypergraph, where images and topics are represented with independent vertices and hyperedges, respectively, is constructed to model inherent high-order semantic correlations of images. Finally, hashing codes and functions are learned by simultaneously enforcing semantic consistence and preserving the discovered semantic relations. Experiments on publicly available datasets demonstrate that THH can achieve superior performance compared with several state-of-the-art methods, and it is more suitable for mobile image retrieval.
Lei Zhu 0002, Jialie Shen 0001, Liang Xie 0001, Zhiyong Cheng 0001
IEEE Trans. Cybern.2
2017 Augmented Collaborative Filtering for Sparseness Reduction in Personalized POI Recommendation
abstract
As mobile device penetration increases, it has become pervasive for images to be associated with locations in the form of geotags. Geotags bridge the gap between the physical world and the cyberspace, giving rise to new opportunities to extract further insights into user preferences and behaviors. In this article, we aim to exploit geotagged photos from online photo-sharing sites for the purpose of personalized Point-of-Interest (POI) recommendation. Owing to the fact that most users have only very limited travel experiences, data sparseness poses a formidable challenge to personalized POI recommendation. To alleviate data sparseness, we propose to augment current collaborative filtering algorithms along from multiple perspectives. Specifically, hybrid preference cues comprising user-uploaded and user-favored photos are harvested to study users’ tastes. Moreover, heterogeneous high-order relationship information is jointly captured from user social networks and POI multimodal contents with hypergraph models. We also build upon the matrix factorization algorithm to integrate the disparate sources of preference and relationship information, and apply our approach to directly optimize user preference rankings. Extensive experiments on a large and publicly accessible dataset well verified the potential of our approach for addressing data sparseness and offering quality recommendations to users, especially for those who have only limited travel experiences.
Chaoran Cui, Jialie Shen 0001, Liqiang Nie, Richang Hong, Jun Ma 0001
ACM Trans. Intell. Syst. Technol.2
2017 Unsupervised Visual Hashing with Semantic Assistant for Content-Based Image Retrieval
abstract
As an emerging technology to support scalable content-based image retrieval (CBIR), hashing has recently received great attention and became a very active research domain. In this study, we propose a novel unsupervised visual hashing approach called semantic-assisted visual hashing (SAVH). Distinguished from semi-supervised and supervised visual hashing, its core idea is to effectively extract the rich semantics latently embedded in auxiliary texts of images to boost the effectiveness of visual hashing without any explicit semantic labels. To achieve the target, a unified unsupervised framework is developed to learn hash codes by simultaneously preserving visual similarities of images, integrating the semantic assistance from auxiliary texts on modeling high-order relationships of inter-images, and characterizing the correlations between images and shared topics. Our performance study on three publicly available image collections: Wiki, MIR Flickr, and NUS-WIDE indicates that SAVH can achieve superior performance over several state-of-the-art techniques.
Lei Zhu 0002, Jialie Shen 0001, Liang Xie 0001, Zhiyong Cheng 0001
IEEE Trans. Knowl. Data Eng.2
2016 Online Cross-Modal Hashing for Web Image Retrieval
abstract
Cross-modal hashing (CMH) is an efficient technique for the fast retrieval of web image data, and it has gained a lot of attentions recently. However, traditional CMH methods usually apply batch learning for generating hash functions and codes. They are inefficient for the retrieval of web images which usually have streaming fashion. Online learning can be exploited for CMH. But existing online hashing methods still cannot solve two essential problems: efficient updating of hash codes and analysis of cross-modal correlation. In this paper, we propose Online Cross-modal Hashing (OCMH) which can effectively address the above two problems by learning the shared latent codes (SLC). In OCMH, hash codes can be represented by the permanent SLC and dynamic transfer matrix. Therefore, inefficient updating of hash codes is transformed to the efficient updating of SLC and transfer matrix, and the time complexity is irrelevant to the database size. Moreover, SLC is shared by all the modalities, and thus it can encode the latent cross-modal correlation, which further improves the overall cross-modal correlation between heterogeneous data. Experimental results on two real-world multi-modal web image datasets: MIR Flickr and NUS-WIDE, demonstrate the effectiveness and efficiency of OCMH for online cross-modal web image retrieval.
Liang Xie 0001, Jialie Shen 0001, Lei Zhu 0002
AAAI2
2016 Learning Compact Visual Representation with Canonical Views for Robust Mobile Landmark Search
Lei Zhu 0002, Jialie Shen 0001, Xiaobai Liu, Liang Xie 0001, Liqiang Nie
IJCAI2
2016 Smart Ambient Sound Analysis via Structured Statistical Modeling
Jialie Shen 0001, Liqiang Nie, Tat-Seng Chua
MMM (2)1
2016 Which Information Sources are More Effective and Reliable in Video Search
abstract
It is common that users are interested in finding video segments, which contain further information about the video contents in a segment of interest. To facilitate users to find and browse related video contents, video hyperlinking aims at constructing links among video segments with relevant information in a large video collection. In this study, we explore the effectiveness of various video features on the performance of video hyperlinking, including subtitle, metadata, content features (i.e., audio and visual), surrounding context, as well as the combinations of those features. Besides, we also test different search strategies over different types of queries, which are categorized according to their video contents. Comprehensive experimental studies have been conducted on the dataset of TRECVID 2015 video hyperlinking task. Results show that (1) text features play a crucial role in search performance, and the combination of audio and visual features cannot provide improvements; (2) the consideration of contexts cannot obtain better results; and (3) due to the lack of training examples, machine learning techniques cannot improve the performance.
Zhiyong Cheng 0001, Xuanchong Li, Jialie Shen 0001, Alex Hauptmann 0001
SIGIR3
2016 On Effective Personalized Music Retrieval by Exploring Online User Behaviors
abstract
In this paper, we study the problem of personalized text based music retrieval which takes users' music preferences on songs into account via the analysis of online listening behaviours and social tags. Towards the goal, a novel Dual-Layer Music Preference Topic Model (DL-MPTM) is proposed to construct latent music interest space and characterize the correlations among (user, song, term). Based on the DL-MPTM, we further develop an effective personalized music retrieval system. To evaluate the system's performance, extensive experimental studies have been conducted over two test collections to compare the proposed method with the state-of-the-art music retrieval methods. The results demonstrate that our proposed method significantly outperforms those approaches in terms of personalized search accuracy.
Zhiyong Cheng 0001, Jialie Shen 0001, Steven C. H. Hoi
SIGIR2
2016 The effects of multiple query evidences on social image retrieval
Zhiyong Cheng 0001, Jialie Shen 0001, Haiyan Miao
Multim. Syst.2
2016 Accurate online video tagging via probabilistic hybrid modeling
Jialie Shen 0001, Meng Wang 0001, Tat-Seng Chua
Multim. Syst.1
2016 On very large scale test collection for landmark image search benchmarking
Zhiyong Cheng 0001, Jialie Shen 0001
Signal Process.2
2016 Cast2Face: Assigning Character Names Onto Faces in Movie With Actor-Character Correspondence
abstract
Automatically identifying characters in movies has attracted researchers' interest and led to several significant and interesting applications. However, due to the vast variation in character appearance as well as the weakness and ambiguity of available annotation, it is still a challenging problem. In this paper, we investigate this problem with the supervision of actor-character name correspondence provided by the movie cast. Our proposed framework, namely, Cast2Face, is featured by: 1) we restrict the assigned names within the set of character names in the cast; 2) for each character, by using the corresponding actor and movie name as keywords, we retrieve from the Google image search and get a group of face images to form the gallery set; 3) the probe face tracks in the movie are then identified as one of the actors by a robust kernel multitask joint sparse representation and classification method; and 4) the conditional random field model with consideration of the constraints between face tracks is introduced to enhance the final labeling. Finally, the assigned actor name of a face track is then mapped to the character name based on the cast again. Besides face naming, we further apply the proposed method to spotlight the summarization of a particular actor in his/her movies. We conduct extensive experiments and empirical evaluations on several feature-length movies to demonstrate the satisfying performance of our method.
Guangyu Gao, Mengdi Xu, Jialie Shen 0001, Huadong Ma, Shuicheng Yan
IEEE Trans. Circuits Syst. Video Technol.3
2016 On Effective Location-Aware Music Recommendation
abstract
Rapid advances in mobile devices and cloud-based music service now allow consumers to enjoy music anytime and anywhere. Consequently, there has been an increasing demand in studying intelligent techniques to facilitate context-aware music recommendation. However, one important context that is generally overlooked is user’s venue, which often includes surrounding atmosphere, correlates with activities, and greatly influences the user’s music preferences. In this article, we present a novel venue-aware music recommender system called VenueMusic to effectively identify suitable songs for various types of popular venues in our daily lives. Toward this goal, a Location-aware Topic Model (LTM) is proposed to (i) mine the common features of songs that are suitable for a venue type in a latent semantic space and (ii) represent songs and venue types in the shared latent space, in which songs and venue types can be directly matched. It is worth mentioning that to discover meaningful latent topics with the LTM, a Music Concept Sequence Generation (MCSG) scheme is designed to extract effective semantic representations for songs. An extensive experimental study based on two large music test collections demonstrates the effectiveness of the proposed topic model and MCSG scheme. The comparisons with state-of-the-art music recommender systems demonstrate the superior performance of VenueMusic system on recommendation accuracy by associating venue and music contents using a latent semantic space. This work is a pioneering study on the development of a venue-aware music recommender system. The results show the importance of considering the influence of venue types in the development of context-aware music recommender systems.
Zhiyong Cheng 0001, Jialie Shen 0001
ACM Trans. Inf. Syst.2
2016 Landmark Reranking for Smart Travel Guide Systems by Combining and Analyzing Diverse Media
abstract
Advanced networking technologies and massive online social media have stimulated a booming growth of travel heterogeneous information in recent years. By employing such information, smart travel guide systems, such as landmark ranking systems, have been proposed to offer diverse online travel services. It is essential for a landmark ranking system to structure, analyze, and search the travel heterogeneous information to produce human-expected results. Therefore, currently the most fundamental yet challenging problems can be concluded: 1) how to fuse heterogeneous tourism information and 2) how to model landmark ranking. In this paper, a novel landmark search system is introduced based on a newly designed heterogeneous information fusion scheme and a query-dependent landmark ranking strategy. Different from the existing travel guide systems, the proposed system can effectively combine the heterogeneous information from multimodality media into a landmark reranking list via a user's query. Experimental results conducted on a large travel information collection illustrate the advantages of the proposed system in terms of both effectiveness and efficiency.
Junge Shen, Jialie Shen 0001, Tao Mei 0001, Xinbo Gao 0001
IEEE Trans. Syst. Man Cybern. Syst.2
2015 R2FP: Rich and Robust Feature Pooling for Mining Visual Data
abstract
The human visual system proves smart in extracting both global and local features. Can we design a similar way for unsupervised feature learning? In this paper, we propose anovel pooling method within an unsupervised feature learningframework, named Rich and Robust Feature Pooling (R2FP), to better explore rich and robust representation from sparsefeature maps of the input data. Both local and global poolingstrategies are further considered to instantiate such a methodand intensively studied. The former selects the most conductivefeatures in the sub-region and summarizes the joint distributionof the selected features, while the latter is utilized to extractmultiple resolutions of features and fuse the features witha feature balancing kernel for rich representation. Extensiveexperiments on several image recognition tasks demonstratethe superiority of the proposed techniques.
Wei Xiong 0008, Bo Du 0001, Lefei Zhang, Ruimin Hu, Wei Bian 0003, Jialie Shen 0001, Dacheng Tao
ICDM6
2015 EMIF: Towards a Scalable and Effective Indexing Framework for Large Scale Music Retrieval
abstract
In this article, we present a novel indexing technique called EMIF (Effective Music Indexing Framework) to facilitate scalable and accurate content based music retrieval. It is designed based on a "classification-and-indexing" principle and consists of two main functionality layers: 1) a novel semantic-sensitive classification to identify input music's category and 2) multiple indexing structures - one local indexing structure corresponds to one semantic category. EMIF's layered architecture not only enables superior search accuracy but also reduces query response time significantly. To evaluate the system, a set of comprehensive experimental studies have been carried out using large test collection and EMIF demonstrates promising performance over state-of-the-art approaches.
Jialie Shen 0001, Tao Mei 0001, Dacheng Tao, Xuelong Li 0001, Yong Rui
ICMR1
2015 Social Tag Relevance Estimation via Ranking-Oriented Neighbour Voting
abstract
User-generated tags associated with social images are frequently imprecise and incomplete. Therefore, a fundamental challenge in tag-based applications is the problem of tag relevance estimation, which concerns how to interpret and quantify the relevance of a tag with respect to the contents of an image. In this paper, we address the key problem from a new perspective of learning to rank, and develop a novel approach to facilitate tag relevance estimation to directly optimize the ranking performance of tag-based image search. A supervision step is introduced into the neighbour voting scheme, in which tag relevance is estimated by accumulating votes from visual neighbours. Through explicitly modelling the neighbour weights and tag correlations, the risk of making heuristic assumptions is effectively avoided for conventional methods. Extensive experiments on a benchmark dataset in comparison with the state-of-the-art methods demonstrate the promise of our approach.
Chaoran Cui, Jialie Shen 0001, Jun Ma 0001, Tao Lian
ACM Multimedia2
2015 Automatic Accident Detection and Alarm System
abstract
Accident detection and alarm system is very important to detect possible accidents or dangers for the peoples using their mobile devices while walking, i.e., distracted walking. In this paper, we introduce an automatic accident detection and alarm system, called AutoADAS, which is fully implemented and tested on the real mobile devices. The proposed system can be activated either manually or automatically when user walks. Under the manual mode, user activates the system before distracted walking while under the automatic mode, a "user behaviour profiling" module is used to recognize (distracted) walking behaviours and an "object detection" module is activated. Using image processing and camera field of view (FOV), the distance and angle between the user and detected objects are estimated and then applied to identify whether any potential accidents can happen. The "accident analysis and prediction" module includes: temporal alarm that inputs the user's walking speed and distance with respect to the detected objects and outputs temporal accident prediction; spatial alarm that inputs the user's walking direction and angle with respect to the detected objects and outputs spatial accident prediction. Once the proposed system positively predicts a potential accident, the "alarm and suggestion" module alerts the user with text, sound or vibration.
Zhuo Wei, Swee-Won Lo, Tieyan Li, Jialie Shen 0001, Robert H. Deng
ACM Multimedia5
2015 Topic Hypergraph Hashing for Mobile Image Retrieval
abstract
Hashing is one of the promising solutions to support efficient Mobile Image Retrieval (MIR). However, most of existing hashing strategies simply rely on low-level features, which inevitably makes the generated hashing codes less semantic. Moreover, many of them fail to exploit complex and high-order semantic correlations of images. Motivated by these observations, we propose a novel unsupervised hashing scheme, \emph{Topic Hypergraph Hashing} (THH), to address the limitations. A unified topic hypergraph, where images and topics are represented with independent vertices and hyperedges respectively, is first constructed to model latent semantics of images and their correlations. With topic hypergraph model, hashing codes and functions are then learned by simultaneously preserving similarity consistence and semantic correlation. Experiments on standard datasets demonstrate that THH can achieve superior performance compared with several state-of-the-art techniques, and it is more suitable for MIR.
Lei Zhu 0002, Jialie Shen 0001, Liang Xie 0001
ACM Multimedia2
2015 Travel Recommendation via Author Topic Model Based Collaborative Filtering
Shuhui Jiang, Xueming Qian, Jialie Shen 0001, Tao Mei 0001
MMM (2)3
2015 Multidimensional Context Awareness in Mobile Devices
Zhuo Wei, Robert H. Deng, Jialie Shen 0001, Jixiang Zhu, Kun Ouyang, Yongdong Wu
MMM (2)3
2015 VenueMusic: A Venue-Aware Music Recommender System
abstract
Users' music preferences can be greatly influenced by their location and environment nearby. In this demonstration, we present an intelligent music recommender system, called VenueMusic, to automatically identify suitable music for various popular venues in our daily lives. VenueMusic enjoys a set of nice features: i) music concept sequence generation scheme and Location-aware Topic Model (LTM) are proposed to map the characteristics of venues and music into a latent semantic space, where suitability of music for a venue can be directly measured, ii) a smart interface enabling user to smoothly interact with VenueMusic, and iii) high quality music playlist. The demonstration will show several interesting use-cases of VenueMusic, and illustrate its superiority on recommending music based on where user presents.
Zhiyong Cheng 0001, Jialie Shen 0001
SIGIR2
2015 On robust image spam filtering via comprehensive visual modeling
Jialie Shen 0001, Robert H. Deng, Zhiyong Cheng 0001, Liqiang Nie, Shuicheng Yan
Pattern Recognit.1
2015 Facilitating Image Search With a Scalable and Compact Semantic Mapping
abstract
This paper introduces a novel approach to facilitating image search based on a compact semantic embedding. A novel method is developed to explicitly map concepts and image contents into a unified latent semantic space for the representation of semantic concept prototypes. Then, a linear embedding matrix is learned that maps images into the semantic space, such that each image is closer to its relevant concept prototype than other prototypes. In our approach, the semantic concepts equated with query keywords and the images mapped into the vicinity of the prototype are retrieved by our scheme. In addition, a computationally efficient method is introduced to incorporate new semantic concept prototypes into the semantic space by updating the embedding matrix. This novelty improves the scalability of the method and allows it to be applied to dynamic image repositories. Therefore, the proposed approach not only narrows semantic gap but also supports an efficient image search process. We have carried out extensive experiments on various cross-modality image search tasks over three widely-used benchmark image datasets. Results demonstrate the superior effectiveness, efficiency, and scalability of our proposed approach.
Meng Wang 0001, Weisheng Li 0001, Dong Liu 0001, Bingbing Ni, Jialie Shen 0001, Shuicheng Yan
IEEE Trans. Cybern.5
2015 Content-Based Visual Landmark Search via Multimodal Hypergraph Learning
abstract
While content-based landmark image search has recently received a lot of attention and became a very active domain, it still remains a challenging problem. Among the various reasons, high diverse visual content is the most significant one. It is common that for the same landmark, images with a wide range of visual appearances can be found from different sources and different landmarks may share very similar sets of images. As a consequence, it is very hard to accurately estimate the similarities between the landmarks purely based on single type of visual feature. Moreover, the relationships between landmark images can be very complex and how to develop an effective modeling scheme to characterize the associations still remains an open question. Motivated by these concerns, we propose multimodal hypergraph (MMHG) to characterize the complex associations between landmark images. In MMHG, images are modeled as independent vertices and hyperedges contain several vertices corresponding to particular views. Multiple hypergraphs are firstly constructed independently based on different visual modalities to describe the hidden high-order relations from different aspects. Then, they are integrated together to involve discriminative information from heterogeneous sources. We also propose a novel content-based visual landmark search system based on MMHG to facilitate effective search. Distinguished from the existing approaches, we design a unified computational module to support query-specific combination weight learning. An extensive experiment study on a large-scale test collection demonstrates the effectiveness of our scheme over state-of-the-art approaches.
Lei Zhu 0002, Jialie Shen 0001, Hai Jin 0001, Liang Xie 0001
IEEE Trans. Cybern.2
2015 Dictionary Pair Learning on Grassmann Manifolds for Image Denoising
abstract
Image denoising is a fundamental problem in computer vision and image processing that holds considerable practical importance for real-world applications. The traditional patch-based and sparse coding-driven image denoising methods convert 2D image patches into 1D vectors for further processing. Thus, these methods inevitably break down the inherent 2D geometric structure of natural images. To overcome this limitation pertaining to the previous image denoising methods, we propose a 2D image denoising model, namely, the dictionary pair learning (DPL) model, and we design a corresponding algorithm called the DPL on the Grassmann-manifold (DPLG) algorithm. The DPLG algorithm first learns an initial dictionary pair (i.e., the left and right dictionaries) by employing a subspace partition technique on the Grassmann manifold, wherein the refined dictionary pair is obtained through a sub-dictionary pair merging. The DPLG obtains a sparse representation by encoding each image patch only with the selected sub-dictionary pair. The non-zero elements of the sparse representation are further smoothed by the graph Laplacian operator to remove the noise. Consequently, the DPLG algorithm not only preserves the inherent 2D geometric structure of natural images but also performs manifold smoothing in the 2D sparse coding space. We demonstrate that the DPLG algorithm also improves the structural SIMilarity values of the perceptual visual quality for denoised images using the experimental evaluations on the benchmark images and Berkeley segmentation data sets. Moreover, the DPLG also produces the competitive peak signal-to-noise ratio values from popular image denoising algorithms.
Wei Bian 0003, Wei Liu 0005, Jialie Shen 0001, Dacheng Tao
IEEE Trans. Image Process.4
2015 Bridging the Vocabulary Gap between Health Seekers and Healthcare Knowledge
abstract
The vocabulary gap between health seekers and providers has hindered the cross-system operability and the inter-user reusability. To bridge this gap, this paper presents a novel scheme to code the medical records by jointly utilizing local mining and global learning approaches, which are tightly linked and mutually reinforced. Local mining attempts to code the individual medical record by independently extracting the medical concepts from the medical record itself and then mapping them to authenticated terminologies. A corpus-aware terminology vocabulary is naturally constructed as a byproduct, which is used as the terminology space for global learning. Local mining approach, however, may suffer from information loss and lower precision, which are caused by the absence of key medical concepts and the presence of irrelevant medical concepts. Global learning, on the other hand, works towards enhancing the local medical coding via collaboratively discovering missing key terminologies and keeping off the irrelevant terminologies by analyzing the social neighbors. Comprehensive experiments well validate the proposed scheme and each of its component. Practically, this unsupervised scheme holds potential to large-scale data.
Liqiang Nie, Yi-Liang Zhao, Mohammad Akbari 0001, Jialie Shen 0001, Tat-Seng Chua
IEEE Trans. Knowl. Data Eng.4
2015 Author Topic Model-Based Collaborative Filtering for Personalized POI Recommendations
abstract
From social media has emerged continuous needs for automatic travel recommendations. Collaborative filtering (CF) is the most well-known approach. However, existing approaches generally suffer from various weaknesses. For example , sparsity can significantly degrade the performance of traditional CF. If a user only visits very few locations, accurate similar user identification becomes very challenging due to lack of sufficient information for effective inference. Moreover, existing recommendation approaches often ignore rich user information like textual descriptions of photos which can reflect users' travel preferences. The topic model (TM) method is an effective way to solve the “sparsity problem,” but is still far from satisfactory. In this paper, an author topic model-based collaborative filtering (ATCF) method is proposed to facilitate comprehensive points of interest (POIs) recommendations for social users. In our approach, user preference topics, such as cultural, cityscape, or landmark, are extracted from the geo-tag constrained textual description of photos via the author topic model instead of only from the geo-tags (GPS locations). Advantages and superior performance of our approach are demonstrated by extensive experiments on a large collection of data.
Shuhui Jiang, Xueming Qian, Jialie Shen 0001, Yun Fu 0001, Tao Mei 0001
IEEE Trans. Multim.3
2015 Landmark Classification With Hierarchical Multi-Modal Exemplar Feature
abstract
Landmark image classification attracts increasing research attention due to its great importance in real applications, ranging from travel guide recommendation to 3-D modelling and visualization of geolocation. While large amount of efforts have been invested, it still remains unsolved by academia and industry. One of the key reasons is the large intra-class variance rooted from the diverse visual appearance of landmark images. Distinguished from most existing methods based on scalable image search, we approach the problem from a new perspective and model landmark classification as multi-modal categorization , which enjoys advantages of low storage overhead and high classification efficiency. Toward this goal, a novel and effective feature representation, called hierarchical multi-modal exemplar (HMME) feature, is proposed to characterize landmark images. In order to compute HMME, training images are first partitioned into the regions with hierarchical grids to generate candidate images and regions. Then, at the stage of exemplar selection, hierarchical discriminative exemplars in multiple modalities are discovered automatically via iterative boosting and latent region label mining. Finally, HMME is generated via a region-based locality-constrained linear coding (RLLC), which effectively encodes semantics of the discovered exemplars into HMME. Meanwhile, dimension reduction is applied to reduce redundant information by projecting the raw HMME into lower-dimensional space. The final HMME enjoys advantages of discriminative and linearly separable. Experimental study has been carried out on real world landmark datasets, and the results demonstrate the superior performance of the proposed approach over several state-of-the-art techniques.
Lei Zhu 0002, Jialie Shen 0001, Hai Jin 0001, Liang Xie 0001
IEEE Trans. Multim.2
2014 Mobile visual search via hievarchical sparse coding
abstract
Mobile visual search is attracting much research attention recently. Existing works focus on addressing the limited capacity of wireless channel yet overlook its instability, thus is not adaptive to the change of channel capacity. In this paper, a novel image retrieval algorithm that is scalable to various channel condition is proposed. The proposed algorithm contains three contributions: (1) to achieve instant retrieval under various channel capacity, we adjust transmission load by sparseness instead of codebook size; (2) we introduce hierarchical sparse coding into our retrieval workflow, where original codebook is transformed into a tree-structured dictionary which implies elements' priority; (3) we propose transmission priority ranking schemes that is adaptive to specific query. Experiment results show that the proposed algorithm outperforms BoW and Lasso based algorithm under different parameter settings. Retrieval results under different channel limitation validate the scalability of our method.
Xiyu Yang, Lianli Liu, Xueming Qian, Tao Mei 0001, Jialie Shen 0001, Qi Tian 0001
ICME5
2014 Just-for-Me: An Adaptive Personalization System for Location-Aware Social Music Recommendation
abstract
The fast growth of online communities and increasing popularity of internet-accessing smart devices have significantly changed the way people consume and share music. As an emerging technology to facilitate effective music retrieval on the move, intelligent recommendation has been recently received great attentions in recent years. While a large amount of efforts have been invested in the field, the technology is still in its infancy. One of the major reasons for this stagnation is due to inability of the existing approaches to comprehensively take multiple kinds of contextual information into account. In the paper, we present a novel recommender system called Just-for-Me to facilitate effective social music recommendation by considering users' location related contexts as well as global music popularity trends. We also develop an unified recommendation model to integrate the contextual factors as well as music contents simultaneously. Furthermore, pseudo-observations are proposed to overcome the cold-start and sparsity problems. An extensive experimental study based on different test collections demonstrates that Just-for-Me system can significantly improve the recommendation performance at various geo-locations.
Zhiyong Cheng 0001, Jialie Shen 0001
ICMR2
2014 The Evolution of Research on Multimedia Travel Guide Search and Recommender Systems
Junge Shen, Zhiyong Cheng 0001, Jialie Shen 0001, Tao Mei 0001, Xinbo Gao 0001
MMM (2)3
2014 Just-for-me: an adaptive personalization system for location-aware social music recommendation
abstract
In recent years, location-aware music recommendation is increasing in popularity, as more and more users consume music on the move. In this demonstration, we present an intelligent system, called Just-for-Me, to facilitate accurate music recommendation based on where user presents. Our system is developed based on a novel probabilistic generative model, which can effectively integrate the location contexts and global music popularity trends. This approach allows us to gain more comprehensive modeling on user preference and thus significantly enhances the music recommendation performance.
Zhiyong Cheng 0001, Jialie Shen 0001, Tao Mei 0001
SIGIR2
2014 WenZher: comprehensive vertical search for healthcare domain
abstract
Online health seeking has transformed the way of health knowledge exchange and reusability. The existing general and vertical health search engines, however, just routinely return lists of matched documents or question answer (QA) pairs, which may overwhelm the seekers or not sufficiently meet the seekers' expectations. Instead, our multilingual system is able to return one multi-faceted answer that is well-structured and precisely extracted from multiple heterogeneous healthcare sources. Further, should the seekers not be satisfied with the returned search results, our system can automatically route the unsolved questions to the professionals with relevant expertise.
Liqiang Nie, Mohammad Akbari 0001, Jialie Shen 0001, Tat-Seng Chua
SIGIR4
2014 SoMeRA 2014: social media retrieval and analysis workshop
abstract
The SoMeRA workshop targets cutting edge research from all fields of retrieval, recommendation, and browsing in social media, as well as the analysis of user's multifaceted traces therein. Submissions to the workshop cover a broad range of topics including multimedia retrieval and exploration, user-aware recommender systems, network analysis, event detection, and computational linguistics.
Markus Schedl, Peter Knees, Jialie Shen 0001
SIGIR3
2014 A multi-dimensional image quality prediction model for user-generated images in social networks
You Yang 0002, Xu Wang 0006, Jialie Shen 0001, Li Yu 0003
Inf. Sci.4
2014 Representative Discovery of Structure Cues for Weakly-Supervised Image Segmentation
abstract
Weakly-supervised image segmentation is a challenging problem with multidisciplinary applications in multimedia content analysis and beyond. It aims to segment an image by leveraging its image-level semantics (i.e., tags). This paper presents a weakly-supervised image segmentation algorithm that learns the distribution of spatially structural superpixel sets from image-level labels. More specifically, we first extract graphlets from a given image, which are small-sized graphs consisting of superpixels and encapsulating their spatial structure. Then, an efficient manifold embedding algorithm is proposed to transfer labels from training images into graphlets. It is further observed that there are numerous redundant graphlets that are not discriminative to semantic categories, which are abandoned by a graphlet selection scheme as they make no contribution to the subsequent segmentation. Thereafter, we use a Gaussian mixture model (GMM) to learn the distribution of the selected post-embedding graphlets (i.e., vectors output from the graphlet embedding). Finally, we propose an image segmentation algorithm, termed representative graphlet cut, which leverages the learned GMM prior to measure the structure homogeneity of a test image. Experimental results show that the proposed approach outperforms state-of-the-art weakly-supervised image segmentation methods, on five popular segmentation data sets. Besides, our approach performs competitively to the fully-supervised segmentation models.
Yue Gao 0002, Yingjie Xia, Ke Lu 0002, Jialie Shen 0001, Rongrong Ji
IEEE Trans. Multim.5
2014 Learning to Recommend Descriptive Tags for Questions in Social Forums
abstract
Around 40% of the questions in the emerging social-oriented question answering forums have at most one manually labeled tag, which is caused by incomprehensive question understanding or informal tagging behaviors. The incompleteness of question tags severely hinders all the tag-based manipulations, such as feeds for topic-followers, ontological knowledge organization, and other basic statistics. This article presents a novel scheme that is able to comprehensively learn descriptive tags for each question. Extensive evaluations on a representative real-world dataset demonstrate that our scheme yields significant gains for question annotation, and more importantly, the whole process of our approach is unsupervised and can be extended to handle large-scale data.
Liqiang Nie, Yi-Liang Zhao, Jialie Shen 0001, Tat-Seng Chua
ACM Trans. Inf. Syst.4
2013 Towards efficient sparse coding for scalable image annotation
abstract
Nowadays, content-based retrieval methods are still the development trend of the traditional retrieval systems. Image labels, as one of the most popular approaches for the semantic representation of images, can fully capture the representative information of images. To achieve the high performance of retrieval systems, the precise annotation for images becomes inevitable. However, as the massive number of images in the Internet, one cannot annotate all the images without a scalable and flexible (i.e., training-free) annotation method. In this paper, we particularly investigate the problem of accelerating sparse coding based scalable image annotation, whose off-the-shelf solvers are generally inefficient on large-scale dataset. By leveraging the prior that most reconstruction coefficients should be zero, we develop a general and efficient framework to derive an accurate solution to the large-scale sparse coding problem through solving a series of much smaller-scale subproblems. In this framework, an active variable set, which expands and shrinks iteratively, is maintained, with each snapshot of the active variable set corresponding to a subproblem. Meanwhile, the convergence of our proposed framework to global optimum is theoretically provable. To further accelerate the proposed framework, a sub-linear time complexity hashing strategy, e.g. Locality-Sensitive Hashing, is seamlessly integrated into our framework. Extensive empirical experiments on NUS-WIDE and IMAGENET datasets demonstrate that the orders-of-magnitude acceleration is achieved by the proposed framework for large-scale image annotation, along with zero/negligible accuracy loss for the cases without/with hashing speed-up, compared to the expensive off-the-shelf solvers.
Junshi Huang, Hairong Liu, Jialie Shen 0001, Shuicheng Yan
ACM Multimedia3
2013 Towards next generation multimedia recommendation systems
abstract
Empowered by advances in information technology, such as social media network, digital library and mobile computing, there emerges an ever-increasing amounts of multimedia data. As the key technology to address the problem of information overload, multimedia recommendation system has been received a lot of attentions from both industry and academia. This course aims to 1) provide a series of detailed review of state-of-the-art in multimedia recommendation; 2) analyze key technical challenges in developing and evaluating next generation multimedia recommendation systems from different perspectives and 3) give some predictions about the road lies ahead of us.
Jialie Shen 0001, Xian-Sheng Hua 0001, Emre Sargin
ACM Multimedia1
2013 Building a Large Scale Test Collection for Effective Benchmarking of Mobile Landmark Search
Zhiyong Cheng 0001, Jialie Shen 0001, Haiyan Miao
MMM (2)3
2013 Multimedia recommendation: technology and techniques
abstract
In recent years, we have witnessed a rapid growth in the availability of digital multimedia on various application platforms and domains. Consequently, the problem of information overload has become more and more serious. In order to tackle the challenge, various multimedia recommendation technologies have been developed by different research communities (e.g., multimedia systems, information retrieval, machine learning and computer version). Meanwhile, many commercial web systems (e.g., Flick, YouTube, and Last.fm) have successfully applied recommendation techniques to provide users personalized content and services in a convenient and flexible way.
Jialie Shen 0001, Meng Wang 0001, Shuicheng Yan, Peng Cui 0001
SIGIR1
2013 Image collection summarization via dictionary learning for sparse representation
Chunlei Yang, Jialie Shen 0001, Jinye Peng 0001, Jianping Fan 0001
Pattern Recognit.2
2013 Robust Image Analysis With Sparse Representation on Quantized Visual Features
abstract
Recent techniques based on sparse representation (SR) have demonstrated promising performance in high-level visual recognition, exemplified by the highly accurate face recognition under occlusion and other sparse corruptions. Most research in this area has focused on classification algorithms using raw image pixels, and very few have been proposed to utilize the quantized visual features, such as the popular bag-of-words feature abstraction. In such cases, besides the inherent quantization errors, ambiguity associated with visual word assignment and misdetection of feature points, due to factors such as visual occlusions and noises, constitutes the major cause of dense corruptions of the quantized representation. The dense corruptions can jeopardize the decision process by distorting the patterns of the sparse reconstruction coefficients. In this paper, we aim to eliminate the corruptions and achieve robust image analysis with SR. Toward this goal, we introduce two transfer processes (ambiguity transfer and mis-detection transfer) to account for the two major sources of corruption as discussed. By reasonably assuming the rarity of the two kinds of distortion processes, we augment the original SR-based reconstruction objective with l(0) norm regularization on the transfer terms to encourage sparsity and, hence, discourage dense distortion/transfer. Computationally, we relax the nonconvex l(0) norm optimization into a convex l(1) norm optimization problem, and employ the accelerated proximal gradient method to optimize the convergence provable updating procedure. Extensive experiments on four benchmark datasets, Caltech-101, Caltech-256, Corel-5k, and CMU pose, illumination, and expression, manifest the necessity of removing the quantization corruptions and the various advantages of the proposed framework.
Bing-Kun Bao, Guangyu Zhu 0002, Jialie Shen 0001, Shuicheng Yan
IEEE Trans. Image Process.3
2013 Visual-Textual Joint Relevance Learning for Tag-Based Social Image Search
abstract
Due to the popularity of social media websites, extensive research efforts have been dedicated to tag-based social image search. Both visual information and tags have been investigated in the research field. However, most existing methods use tags and visual characteristics either separately or sequentially in order to estimate the relevance of images. In this paper, we propose an approach that simultaneously utilizes both visual and textual information to estimate the relevance of user tagged images. The relevance estimation is determined with a hypergraph learning approach. In this method, a social image hypergraph is constructed, where vertices represent images and hyperedges represent visual or textual terms. Learning is achieved with use of a set of pseudo-positive images, where the weights of hyperedges are updated throughout the learning process. In this way, the impact of different tags and visual words can be automatically modulated. Comparative results of the experiments conducted on a dataset including 370+images are presented, which demonstrate the effectiveness of the proposed approach.
Yue Gao 0002, Meng Wang 0001, Zhengjun Zha, Jialie Shen 0001, Xuelong Li 0001, Xindong Wu 0001
IEEE Trans. Image Process.4
2013 Query-Document-Dependent Fusion: A Case Study of Multimodal Music Retrieval
abstract
In recent years, multimodal fusion has emerged as a promising technology for effective multimedia retrieval. Developing the optimal fusion strategy for different modalities (e.g., content, metadata) has been the subject of intensive research. Given a query, existing methods derive a unified fusion strategy for all documents with the underlying assumption that the relative significance of a modality remains the same across all documents. However, this assumption is often invalid. We thus propose a general multimodal fusion framework, query-document-dependent fusion (QDDF), which derives the optimal fusion strategy for each query-document pair via intelligent content analysis of both queries and documents. By investigating multimodal fusion strategies adaptive to both queries and documents, we demonstrate that existing multimodal fusion approaches are special cases of QDDF and propose two QDDF approaches to derive fusion strategies. The dual-phase QDDF explicitly derives and fuses query- and document-dependent weights, and the regression-based QDDF determines the fusion weight for a query-document pair via a regression model derived from training data. To evaluate the proposed approaches, comprehensive experiments have been conducted using a multimedia data set with around 17 K full songs and over 236 K social queries. Results indicate that the regression-based QDDF is superior in handling single-dimension queries. In comparison, the dual-phase QDDF outperforms existing approaches for most query types. We found that document-dependent weights are instrumental in enhancing multimedia fusion performance. In addition, efficiency analysis demonstrates the scalability of QDDF over large data sets.
Bingjun Zhang, Yi Yu 0001, Jialie Shen 0001, Ye Wang 0007
IEEE Trans. Multim.4
2012 Obfuscating the Topical Intention in Enterprise Text Search
abstract
The text search queries in an enterprise can reveal the users' topic of interest, and in turn confidential staff or business information. To safeguard the enterprise from consequences arising from a disclosure of the query traces, it is desirable to obfuscate the true user intention from the search engine, without requiring it to be re-engineered. In this paper, we advocate a unique approach to profile the topics that are relevant to the user intention. Based on this approach, we introduce an (ε1, ε2)-privacy model that allows a user to stipulate that topics relevant to her intention at ε1level should appear to any adversary to be innocuous at ε2level. We then present a Top Priv algorithm to achieve the customized (ε1, ε2)-privacy requirement of individual users through injecting automatically formulated fake queries. The advantages of Top Priv over existing techniques are confirmed through benchmark queries on a real corpus, with experiment settings fashioned after an enterprise search application.
HweeHwa Pang, Xiaokui Xiao, Jialie Shen 0001
ICDE3
2012 Predicting Image Popularity in an Incomplete Social Media Community by a Weighted Bi-partite Graph
abstract
Popularity prediction is a key problem in networks to analyze the information diffusion, especially in social media communities. Recently, there have been some custom-build prediction models in Digg and YouTube. However, these models are hardly transplant to an incomplete social network site (e.g., Flickr) by their unique parameters. In addition, because of the large scale of the network in Flickr, it is difficult to get all of the photos and the whole network. Thus, we are seeking for a method which can be used in such incomplete network. Inspired by a collaborative filtering method-Network-based Inference (NBI), we devise a weighted bipartite graph with undetected users and items to represent the resource allocation process in an incomplete network. Instead of image analysis, we propose a modified interdisciplinary models, called Incomplete Network-based Inference (INI). Using the data from 30 months in Flickr, we show the proposed INI is able to increase prediction accuracy by over 58.1%, compared with traditional NBI. We apply our proposed INI approach to personalized advertising application and show that it is more attractive than traditional Flickr advertising.
Xiang Niu, Lusong Li, Tao Mei 0001, Jialie Shen 0001, Ke Xu 0001
ICME4
2012 Multimedia recommendation
abstract
Due to the rapid growth of online multimedia information, the problem of information overload has become more and more serious in recent decades. To address this problem, various multimedia recommendation technologies have been developed by different research communities (e.g., multimedia systems, information retrieval, and machine learning). Meanwhile, many commercial web systems (e.g., Flick, Youtube, and Last.fm) have successfully applied recommendation techniques to provide users personalized multimedia content and services in a convenient and flexible way. This tutorial focuses on exploring the state-of-the-art in multimedia recommendation. We also discuss the experience gained from developing existing systems and review key challenges associated with large-scale multimedia recommendation.
Jialie Shen 0001, Meng Wang 0001, Shuicheng Yan, Peng Cui 0001
ACM Multimedia1
2012 View-based 3D object retrieval by bipartite graph matching
abstract
Bipartite graph matching has been investigated in multiple view matching for 3D object retrieval. However, existing methods employ one-to-one vertex matching scheme while more than two views may share close semantic meanings in practice. In this work, we propose a bipartite graph matching method to measure the distance between two objects based on multiple views. In the proposed method, representative views are first selected by using view clustering for each object, and the corresponding weights are given based on the cluster results. A bipartite graph is constructed by using the two groups of representative views from two compared objects. To calculate the similarity between two objects, the bipartite graph is first partitioned to several subsets, and the views in the same sub-set are with high possibility to be with similar semantic meanings. The distances between two objects within individual subsets are then assembled through the graph to obtain the final similarity. Experimental results and comparison with the state-of-the-art methods demonstrate the effectiveness of the proposed algorithm.
Yue Wen, Yue Gao 0002, Richang Hong, Huan-Bo Luan, Qiong Liu 0001, Jialie Shen 0001, Rongrong Ji
ACM Multimedia6
2012 Modeling concept dynamics for large scale music search
abstract
Continuing advances in data storage and communication technologies have led to an explosive growth in digital music collections. To cope with their increasing scale, we need effective Music Information Retrieval (MIR) capabilities like tagging, concept search and clustering. Integral to MIR is a framework for modelling music documents and generating discriminative signatures for them. In this paper, we introduce a multimodal, layered learning framework called DMCM. Distinguished from the existing approaches that encode music as an ensemble of order-less feature vectors, our framework extracts from each music document a variety of acoustic features, and translates them into low-level encodings over the temporal dimension. From them, DMCM elucidates the concept dynamics in the music document, representing them with a novel music signature scheme called Stochastic Music Concept Histogram (SMCH) that captures the probability distribution over all the concepts. Experiment results with two large music collections confirm the advantages of the proposed framework over existing methods on various MIR tasks.
Jialie Shen 0001, HweeHwa Pang, Meng Wang 0001, Shuicheng Yan
SIGIR1
2012 k-Partite graph reinforcement and its application in multimedia information retrieval
Yue Gao 0002, Meng Wang 0001, Rongrong Ji, Zhengjun Zha, Jialie Shen 0001
Inf. Sci.5
2011 Efficient Subspace Segmentation via Quadratic Programming
abstract
We explore in this paper efficient algorithmic solutions to robustsubspace segmentation. We propose the SSQP, namely SubspaceSegmentation via Quadratic Programming, to partition data drawnfrom multiple subspaces into multiple clusters. The basic idea ofSSQP is to express each datum as the linear combination of otherdata regularized by an overall term targeting zero reconstructioncoefficients over vectors from different subspaces. The derivedcoefficient matrix by solving a quadratic programming problem istaken as an affinity matrix, upon which spectral clustering isapplied to obtain the ultimate segmentation result. Similar tosparse subspace clustering (SCC) and low-rank representation (LRR),SSQP is robust to data noises as validated by experiments on toydata. Experiments on Hopkins 155 database show that SSQP can achievecompetitive accuracy as SCC and LRR in segmenting affine subspaces,while experimental results on the Extended Yale Face Database Bdemonstrate SSQP's superiority over SCC and LRR. Beyond segmentationaccuracy, all experiments show that SSQP is much faster than bothSSC and LRR in the practice of subspace segmentation.
Shusen Wang, Xiao-Tong Yuan, Tiansheng Yao, Shuicheng Yan, Jialie Shen 0001
AAAI5
2011 Tag-based social image search with visual-text joint hypergraph learning
abstract
Tag-based social image search has attracted great interest and how to order the search results based on relevance level is a research problem. Visual content of images and tags have both been investigated. However, existing methods usually employ tags and visual content separately or sequentially to learn the image relevance. This paper proposes a tag-based image search with visual-text joint hypergraph learning. We simultaneously investigate the bag-of-words and bag-of-visual-words representations of images and accomplish the relevance estimation with a hypergraph learning approach. Each textual or visual word generates a hyperedge in the constructed hypergraph. We conduct experiments with a real-world data set and experimental results demonstrate the effectiveness of our approach.
Yue Gao 0002, Meng Wang 0001, Huan-Bo Luan, Jialie Shen 0001, Shuicheng Yan, Dacheng Tao
ACM Multimedia4
2011 Multimedia tagging: past, present and future
abstract
The tags have proved to be a very crucial mechanism to facilitate the effective sharing and organization of large scale of multimedia information. As a result, technical developments on intelligent multimedia tagging have attracted a substantial amount of efforts involving experts from information retrieval, multimedia computing and artificial intelligence (particularly computer vision). The truly interdisciplinary research has resulted in many algorithmic and methodological developments. Meanwhile, many commercial web systems (e.g., Youtube, Last.fm and Flickr) have successfully introduced a variety of toolkits to assist different users in discovering and exploring media content using tags. This tutorial aims to provide a comprehensive coverage on the evolution of research for developing multimedia tagging technologies and identify a range of major challenges for the further scholarly study in the coming years.
Jialie Shen 0001, Meng Wang 0001, Shuicheng Yan, Xian-Sheng Hua 0001
ACM Multimedia1
2011 Effective summarization of large-scale web images
abstract
In this paper, we present a novel framework to achieve effective summarization of large-scale web images by treating the problem of automatic image summarization as the problem of dictionary learning for sparse coding, e.g., the summary of a given image set can be treated as a sparse representation of the given image set (i.e., sparse dictionary for the given image set). For a given semantic category (i.e., certain object class or image concept), we build a sparsity model to reconstruct all its relevant images by using a subset of most representative images (i.e., image summary); and a stepwise basis selection algorithm is developed to learn such sparse dictionary (i.e., image summary) by minimizing an explicit optimization function. By investigating their reconstruction ability, the reconstruction Mean Square Error (MSE) is adapted to objectively measure the performance of various algorithms for automatic image summarization. Our experimental results demonstrate that our dictionary learning for sparse representation algorithm can obtain more accurate summary as compared with other baseline algorithms for automatic image summarization.
Chunlei Yang, Jialie Shen 0001, Jianping Fan 0001
ACM Multimedia2
2011 Optimizing multimodal reranking for web image search
abstract
In this poster, we introduce a web image search reranking approach with exploring multiple modalities. Diff erent from the conventional methods that build graph with one feature set for reranking, our approach integrates multiple feature sets that describe visual content from different aspects. We simultaneously integrate the learning of relevance scores, the weighting of different feature sets, the distance metric and the scaling for each feature set into a unified scheme. Experimental results on a large data set that contains more than 1,100 queries and 1 million images demonstrate the effectiveness of our approach.
Hao Li 0030, Meng Wang 0001, Zhisheng Li, Zhengjun Zha, Jialie Shen 0001
SIGIR5
2011 Personalized video similarity measure
Jialie Shen 0001, Zhiyong Cheng 0001
Multim. Syst.1
2010 Weakly-supervised hashing in kernel space
abstract
The explosive growth of the vision data motivates the recent studies on efficient data indexing methods such as locality-sensitive hashing (LSH). Most existing approaches perform hashing in an unsupervised way. In this paper we move one step forward and propose a supervised hashing method, i.e., the LAbel-regularized Max-margin Partition (LAMP) algorithm. The proposed method generates hash functions in weakly-supervised setting, where a small portion of sample pairs are manually labeled to be “similar” or “dissimilar”. We formulate the task as a Constrained Convex-Concave Procedure (CCCP), which can be relaxed into a series of convex sub-problems solvable with efficient Quadratic-Program (QP). The proposed hashing method possesses other characteristics including: 1) most existing LSH approaches rely on linear feature representation. Unfortunately, kernel tricks are often more natural to gauge the similarity between visual objects in vision research, which corresponds to probably infinite-dimensional Hilbert spaces. The proposed LAMP has a natural support for kernel-based feature representation. 2) traditional hashing methods assume uniform data distributions. Typically, the collision probability of two samples in hash buckets is only determined by pairwise similarity, unrelated to contextual data distribution. In contrast, we provide such a collision bound which is beyond pairwise data interaction based on Markov random fields theory. Extensive empirical evaluations are conducted on five widely-used benchmarks. It takes only several seconds to generate a new hashing function, and the adopted random supporting-vector scheme enables the LAMP algorithm scalable to large-scale problems. Experimental results well validate the superiorities of the LAMP algorithm over the state-of-the-art kernel-based hashing methods.
Yadong Mu, Jialie Shen 0001, Shuicheng Yan
CVPR2
2010 Intelligent query: open another door to 3d object retrieval
abstract
The increasing number of available 3D objects makes their efficient retrieval technology highly desired. Extensive research has been dedicated to view-based 3D object retrieval because of its advantage of 2D views for 3D object content representation. In this paradigm, typically the retrieval is accomplished based a set of different views of the query object, and focuses on the 3D object representation, matching and indexing. In this work, we present another aspect towards 3D object retrieval: intelligent query. Intelligent query includes query selection, query description and combination, and assistive query. We will show how this scheme is ideally suit for the 3D object retrieval problem. We conduct experiments on the National Taiwan University 3D Model database and results demonstrated that our approach can improve retrieval performance. Finally, we give insight into the future of the intelligent query for 3D object retrieval.
Yue Gao 0002, Meng Wang 0001, Jialie Shen 0001, Qionghai Dai, Naiyao Zhang
ACM Multimedia3
2010 Cast2Face: character identification in movie with actor-character correspondence
abstract
We investigate the problem of automatically identifying characters in a movie with the supervision of actor-character name correspondence provided by the movie cast. Our proposed framework, namely Cast2Face, is featured by: (i) we restrict the names to assign within the set of character names in the cast; (ii) for each character, by using the corresponding actor's name as a key word, we retrieve from Google image search a group of face images to form the gallery set; and (iii) the probe face tracks in the movie are then identified as one of the actors by robust multi-task joint sparse representation and classification method. The assigned actor name on a face track is then mapped to the character name based on the cast again. In addition to face naming, we further apply the proposed method to spotlights summarization of a particular actor in his/her movies. Empirical evaluations on several feature-length movies demonstrate the satisfying performance of our method.
Mengdi Xu, Xiao-Tong Yuan, Jialie Shen 0001, Shuicheng Yan
ACM Multimedia3
2010 Dual Phase Learning for Large Scale Video Gait Recognition
Jialie Shen 0001, HweeHwa Pang, Dacheng Tao, Xuelong Li 0001
MMM1
2010 Effective music tagging through advanced statistical modeling
abstract
Music information retrieval (MIR) holds great promise as a technology for managing large music archives. One of the key components of MIR that has been actively researched into is music tagging. While significant progress has been achieved, most of the existing systems still adopt a simple classification approach, and apply machine learning classifiers directly on low level acoustic features. Consequently, they suffer the shortcomings of (1) poor accuracy, (2) lack of comprehensive evaluation results and the associated analysis based on large scale datasets, and (3) incomplete content representation, arising from the lack of multimodal and temporal information integration.
Jialie Shen 0001, Wang Meng, Shuicheng Yan, HweeHwa Pang, Xian-Sheng Hua 0001
SIGIR1
2010 Privacy-preserving similarity-based text retrieval
abstract
Users of online services are increasingly wary that their activities could disclose confidential information on their business or personal activities. It would be desirable for an online document service to perform text retrieval for users, while protecting the privacy of their activities. In this article, we introduce a privacy-preserving, similarity-based text retrieval scheme that (a) prevents the server from accurately reconstructing the term composition of queries and documents, and (b) anonymizes the search results from unauthorized observers. At the same time, our scheme preserves the relevance-ranking of the search server, and enables accounting of the number of documents that each user opens. The effectiveness of the scheme is verified empirically with two real text corpora.
HweeHwa Pang, Jialie Shen 0001, Ramayya Krishnan
ACM Trans. Internet Techn.2
2009 Setting discrete bid levels adaptively in repeated auctions
abstract
The success of an auction design often hinges on its ability to set parameters such as reserve price and bid levels that will maximize an objective function such as the auctioneer revenue. Works on designing adaptive auction mechanisms have emerged recently, and the challenge is in learning different auction parameters by observing the bidding in previous auctions. In this paper, we propose a non-parametric method for determining discrete bid levels dynamically so as to maximize the auctioneer revenue. First, we propose a non-parametric kernel method for estimating the probabilities of closing price with past auction data. Then a greedy strategy has been devised to determine the discrete bid levels based on the estimated probability information of closing price. We show experimentally that our non-parametric method is robust to changes in parameters such as the distributions of participating bidders as well as the individual bidder evaluation, and it consistently outperforms different competitors with various settings with respect to auctioneer revenue maximization.
Jilian Zhang, Hoong Chuin Lau, Jialie Shen 0001
ICEC3
2009 Exploiting Intensity Inhomogeneity to Extract Textured Objects from Natural Scenes
Jundi Ding, Jialie Shen 0001, HweeHwa Pang, Songcan Chen, Jing-Yu Yang 0001
ACCV (3)2
2009 Comprehensive query-dependent fusion using regression-on-folksonomies: a case study of multimodal music search
abstract
The combination of heterogeneous knowledge sources has been widely regarded as an effective approach to boost retrieval accuracy in many information retrieval domains. While various technologies have been recently developed for information retrieval, multimodal music search has not kept pace with the enormous growth of data on the Internet. In this paper, we study the problem of integrating multiple online information sources to conduct effective query dependent fusion (QDF) of multiple search experts for music retrieval. We have developed a novel framework to construct a knowledge space of users' information need from online folksonomy data. With this innovation, a large number of comprehensive queries can be automatically constructed to train a better generalized QDF system against unseen user queries. In addition, our framework models QDF problem by regression of the optimal combination strategy on a query. Distinguished from the previous approaches, the regression model of QDF (RQDF) offers superior modeling capability with less constraints and more efficient computation. To validate our approach, a large scale test collection has been collected from different online sources, such as Last.fm, Wikipedia, and YouTube. All test data will be released to the public for better research synergy in multimodal music search. Our performance study indicates that the accuracy, efficiency, and robustness of the multimodal music search can be improved significantly by the proposed folksonomy-RQDF approach. In addition, since no human involvement is required to collect training examples, our approach offers great feasibility and practicality in system development.
Bingjun Zhang, Qiaoliang Xiang, Huanhuan Lu, Jialie Shen 0001, Ye Wang 0007
ACM Multimedia4
2009 CompositeMap: a novel music similarity measure for personalized multimodal music search
abstract
How to measure and model the similarity between different music items is one of the most fundamental yet challenging research problems in music information retrieval. This paper demonstrates a novel multimodal and adaptive music similarity measure (CompositeMap) with its application in a personalized multimodal music search system. CompositeMap can effectively combine music properties from different aspects into compact signatures via supervised learning, which lays the foundation for effective and efficient music search. In addition, an incremental Locality Sensitive Hashing algorithm is developed to support more efficient search processes. Experimental results based on two large music collections reveal various advantages in effectiveness, efficiency, adaptiveness, and scalability of the proposed music similarity measure and the music search system.
Bingjun Zhang, Qiaoliang Xiang, Ye Wang 0007, Jialie Shen 0001
ACM Multimedia4
2009 CompositeMap: a novel framework for music similarity measure
abstract
With the continuing advances in data storage and communication technology, there has been an explosive growth of music information from different application domains. As an effective technique for organizing, browsing, and searching large data collections, music information retrieval is attracting more and more attention. How to measure and model the similarity between different music items is one of the most fundamental yet challenging research problems. In this paper, we introduce a novel framework based on a multimodal and adaptive similarity measure for various applications. Distinguished from previous approaches, our system can effectively combine music properties from different aspects into a compact signature via supervised learning. In addition, an incremental Locality Sensitive Hashing algorithm has been developed to support efficient retrieval processes with different kinds of queries. Experimental results based on two large music collections reveal various advantages of the proposed framework including effectiveness, efficiency, adaptiveness, and scalability.
Bingjun Zhang, Jialie Shen 0001, Qiaoliang Xiang, Ye Wang 0007
SIGIR2
2009 Robust Semantic Concept Detection in Large Video Collections
abstract
With explosive amounts of video data emerging from the Internet, automatic video concept detection is becoming very important and has been received great attention. However, reported approaches mainly suffer from low identification accuracy and poor robustness over different concepts. One of the main reason is that the existing approaches typically isolate the video signature generation from the process of classifier training. Also, very few approaches consider effects of multiple video features. The paper describes a novel approach fusing different information from diverse knowledge sources to facilitate effective video concept detection. The system is designed based on CM*F scheme and its basic architecture contains two core components including 1) CM*F based video signature generation scheme and 2) CM*F based video concept detector. To evaluate the approach proposed, an extensive experimental study on two large video databases has been carried out. The results demonstrate the superiority of the method in terms of effectiveness and robustness.
Jialie Shen 0001, Dacheng Tao, Xuelong Li 0001
SMC1
2009 Stochastic modeling western paintings for effective classification
Jialie Shen 0001
Pattern Recognit.1
2009 Visual information analysis for security
Dacheng Tao, Yuan Yuan 0001, Jialie Shen 0001, Kaiqi Huang, Xuelong Li 0001
Signal Process.3
2009 QUC-Tree: Integrating Query Context Information for Efficient Music Retrieval
abstract
In this paper, we introduce a novel indexing scheme-query context tree (QUC-tree) to facilitate efficient query sensitive music search under different query contexts. Distinguished from the previous approaches, QUC-tree is a balanced multiway tree structure, where each level represents the data space at different dimensionality. Before the tree structure construction, principle component analysis (PCA) is applied for data analysis and transforming the raw composite features into a new feature space sorted by the importance of acoustic features. The PCA transformed data and reduced dimensions in the upper levels can alleviate suffering from dimensionality curse. To accurately mimic human perception, an extension called QUC +-tree is proposed, which further applies multivariate regression and EM based algorithm to estimate the weight of each individual feature. The comprehensive extensive experiments to evaluate the proposed structures against state-of-art techniques based on different datasets. The experimental results demonstrate the superiority of our technique.
Jialie Shen 0001, Dacheng Tao, Xuelong Li 0001
IEEE Trans. Multim.1
2009 A novel framework for efficient automated singer identification in large music databases
abstract
Over the past decade, there has been explosive growth in the availability of multimedia data, particularly image, video, and music. Because of this, content-based music retrieval has attracted attention from the multimedia database and information retrieval communities. Content-based music retrieval requires us to be able to automatically identify particular characteristics of music data. One such characteristic, useful in a range of applications, is the identification of the singer in a musical piece. Unfortunately, existing approaches to this problem suffer from either low accuracy or poor scalability. In this article, we propose a novel scheme, calledHybrid Singer Identifier(HSI), for efficient automated singer recognition. HSI uses multiple low-level features extracted from both vocal and nonvocal music segments to enhance the identification process; it achieves this via a hybrid architecture that builds profiles of individual singer characteristics based on statistical mixture models. An extensive experimental study on a large music database demonstrates the superiority of our method over state-of-the-art approaches in terms of effectiveness, efficiency, scalability, and robustness.
Jialie Shen 0001, John Shepherd 0001, Bin Cui 0001, Kian-Lee Tan
ACM Trans. Inf. Syst.1
2008 Bayesian tensor analysis
abstract
Vector data are normally used for probabilistic graphical models with Bayesian inference. However, tensor data, i.e., multidimensional arrays, are actually natural representations of a large amount of real data, in data mining, computer vision, and many other applications. Aiming at breaking the huge gap between vectors and tensors in conventional statistical tasks, e.g., automatic model selection, this paper proposes a decoupled probabilistic algorithm, named Bayesian tensor analysis (BTA). BTA automatically selects a suitable model for tensor data, as demonstrated by empirical studies.
Dacheng Tao, Jimeng Sun 0001, Jialie Shen 0001, Xindong Wu 0001, Xuelong Li 0001, Stephen J. Maybank, Christos Faloutsos
IJCNN3
2008 Distribution-based similarity measures for multi-dimensional point set retrieval applications
abstract
Effective and efficient method of similarity assessment continues to be one of the most fundamental problems in multimedia data analysis. In case of retrieving relevant items from a collection of objects based on series of multivariate observations (e.g., searching the similar video clips in a repository to a query example), satisfactory performance cannot be expected using many conventional similarity measures based on the aggregation of element pairwise comparisons. Some correlation information among the individual elements has also been investigated to characterize each set of multi-dimensional points for ranked retrieval, by making use of an unwarranted assumption that the underlying data distribution has a particular parametric form. Motivated by this observation, this paper introduces a novel collective gauge of relevance ranking by evaluating the probabilities that point sets are consistent with the same distribution of the query. Two non-parametric hypothesis tests in statistics are justified to exploit the distributional discrepancy of samples for assessing the similarity between two ensembles of points. While our methodology is mainly presented in the context of video similarity search, it enjoys great flexibility and can be easily adapted to other applications involving generic multi-dimensional point set representation for each object such as human gesture recognition.
Jie Shao 0001, Zi Huang, Heng Tao Shen, Jialie Shen 0001, Xiaofang Zhou 0001
ACM Multimedia4
2008 Effective video event detection via subspace projection
abstract
This paper describes a new video event detection framework based on subspace selection technique. With the approach, feature vectors presenting different kinds of video information can be easily projected from different modalities onto an unified subspace, on which recognition process can be performed. The approach is capable of discriminating different classes and preserving the intra-modal geometry of samples within an identical class. Distinguished from the existing multi-modal detection methods, the new system works well when some modalities are not available. Experimental results based on soccer video and TRECVID news video collections demonstrate the effectiveness, efficiency and robustness of the proposed method for individual recognition tasks in comparison to the existing approaches.
Jialie Shen 0001, Dacheng Tao, Xuelong Li 0001
MMSP1
2008 Towards Efficient and Flexible KNN Query Processing in Real-Life Road Networks
abstract
Along with the developments of mobile services, effectively modeling road networks and efficiently indexing and querying network constrained objects has become a challenging problem. In this paper, we first introduce a road network model which captures real-life road networks better than previous models. Then, based on the proposed model, we propose a novel index named the RNG (road network grid) index for accelerating KNN queries and continuous KNN queries over road network constrained data points. In contrast to conventional methods, speed limitations and blocking information of roads are included into the RNG index, which enables the index to support both distance-based and time-based KNN queries and continuous KNN queries. Our work extends previous ones by taking into account more practical scenarios, such as complexities in real-life road networks and time-based KNN queries. Extensive experimental study shows that our methods are efficient in terms of both CPU time and disk I/Os.
Bin Cui 0001, Jiakui Zhao, Hua Lu 0001, Jialie Shen 0001
WAIM5
2008 How to influence my customers?: the impact of electronic market design
abstract
This paper investigates the strategic decisions of online vendors for offering different mechanism, such as sampling and online reviews of information products, to increase their online sales. Focusing on measuring the effectiveness of electronic market design (offering reviews, sampling, or both), our study shows that online markets behavior as communication markets, and consumers learn product quality information both passively (reading online reviews) and actively but subjectively (listening to music sampling). Using data from Amazon, first we show that sampling along is a strong product quality signal that reduces the product uncertainty after controlling for halo effect. In general, products with sampling option enjoy a higher conversion rate (which leads to better sales) than those without sampling because sampling decreases the uncertainty of consuming experience goods. Second, the impact of online reviews on sales conversion rate is lower for experience goods with a sampling option than those without. Third, when the uncertainty of the societal reviews is higher, sampling plays a more important role because it mitigates such uncertainty introduced by online reviews.
Ling Liu 0001, Jialie Shen 0001
WWW4
2008 Modality Mixture Projections for Semantic Video Event Detection
abstract
Event detection is one of the most fundamental components for various kinds of domain applications of video information system. In recent years, it has gained a considerable interest of practitioners and academics from different areas. While detecting video event has been the subject of extensive research efforts recently, much less existing approach has considered multimodal information and related efficiency issues. In this paper, we use a subspace selection technique to achieve fast and accurate video event detection using a subspace selection technique. The approach is capable of discriminating different classes and preserving the intramodal geometry of samples within an identical class. With the method, feature vectors presenting different kind of multi data can be easily projected from different identities and modalities onto a unified subspace, on which recognition process can be performed. Furthermore, the training stage is carried out once and we have a unified transformation matrix to project different modalities. Unlike existing multimodal detection systems, the new system works well when some modalities are not available. Experimental results based on soccer video and TRECVID news video collections demonstrate the effectiveness, efficiency and robustness of the proposed MMP for individual recognition tasks in comparison to the existing approaches.
Jialie Shen 0001, Dacheng Tao, Xuelong Li 0001
IEEE Trans. Circuits Syst. Video Technol.1
2008 Bayesian Tensor Approach for 3-D Face Modeling
abstract
Effectively modeling a collection of three-dimensional (3-D) faces is an important task in various applications, especially facial expression-driven ones, e.g., expression generation, retargeting, and synthesis. These 3-D faces naturally form a set of second-order tensors-one modality for identity and the other for expression. The number of these second-order tensors is three times of that of the vertices for 3-D face modeling. As for algorithms, Bayesian data modeling, which is a natural data analysis tool, has been widely applied with great success; however, it works only for vector data. Therefore, there is a gap between tensor-based representation and vector-based data analysis tools. Aiming at bridging this gap and generalizing conventional statistical tools over tensors, this paper proposes a decoupled probabilistic algorithm, which is named Bayesian tensor analysis (BTA). Theoretically, BTA can automatically and suitably determine dimensionality for different modalities of tensor data. With BTA, a collection of 3-D faces can be well modeled. Empirical studies on expression retargeting also justify the advantages of BTA.
Dacheng Tao, Mingli Song, Xuelong Li 0001, Jialie Shen 0001, Jimeng Sun 0001, Xindong Wu 0001, Christos Faloutsos, Stephen J. Maybank
IEEE Trans. Circuits Syst. Video Technol.4
2007 Probabilistic Tensor Analysis with Akaike and Bayesian Information Criteria
Dacheng Tao, Jimeng Sun 0001, Xindong Wu 0001, Xuelong Li 0001, Jialie Shen 0001, Stephen J. Maybank, Christos Faloutsos
ICONIP (1)5
2007 QueST: querying music databases by acoustic and textual features
abstract
With continued growth of music content available on the Internet, music information retrieval has attracted increasing attention. An important challenge for music searching is its ability to support both keyword and content based queries efficiently and with high precision. In this paper, we present a music query system - QueST (Query by acouStic and Textual features) to support both keyword and content based retrieval in large music databases. QueST has two distinct features. First, it provides new index schemes that can efficiently handle various queries within a uniform architecture. Concretely, we propose a hybrid structure consisting of Inverted file and Signature file to support keyword search. For content based query, we introduce the notion of similarity to capture various music semantics like melody and genre. We extract acoustic features from a music object, and map it to multiple high-dimension spaces with respect to the similarity notion using PCA and RBF neural network. Second, we design a result fusion scheme, called the Quick Threshold Algorithm, to speed up the processing of complex queries involving both textual and multiple acoustic features. Our experimental results show that QueST offers higher accuracy and efficiency compared to existing algorithms.
Bin Cui 0001, Ling Liu 0001, Calton Pu, Jialie Shen 0001, Kian-Lee Tan
ACM Multimedia4
2006 An Agent-Based Approach for Cooperative Data Management
ChunYu Miao, Meilin Shi, Jialie Shen 0001
APWeb3
2006 HSI: A Novel Framework for Efficient Automated Singer Identification in Large Music Database
abstract
The singer’s information is essential in organising, browsing and exploring music data. As an important component of music database systems, the automated artist identification is gaining considerable momentum due to numerous potential applications including music indexing and retrieval, copy right management and music recommendation systems. Unfortunately, the most currently employed approaches are still in their infancy and the performance is by far less satisfactory. Indeed, they suffer from low effectiveness, less robustness and poor scalability to accommodate large scale of data. In this demo, we presents a novel system, called Hybrid Singer Identifier (HSI), for efficient and effective automated singer identification in large music databases.
Jialie Shen 0001, John Shepherd 0001, Bin Cui 0001, Kian-Lee Tan
ICDE1
2006 Exploring composite acoustic features for efficient music similarity query
abstract
Music similarity query based on acoustic content is becoming important with the ever-increasing growth of the music information from emerging applications such as digital libraries and WWW. However, relative techniques are still in their infancy and much less than satisfactory. In this paper, we present a novel index structure, called Composite Feature tree, CF-tree, to facilitate efficient content-based music search adopting multiple musical features. Before constructing the tree structure, we use PCA to transform the extracted features into a new space sorted by the importance of acoustic features. The CF-tree is a balanced multi-way tree structure where each level represents the data space at different dimensionalities. The PCA transformed data and reduced dimensions in the upper levels can alleviate suffering from dimensionality curse. To accurately mimic human perception, an extension, named CF+-tree, is proposed, which further applies multivariable regression to determine the weight of each individual feature. We conduct extensive experiments to evaluate the proposed structures against state-of-art techniques. The experimental results demonstrate superiority of our technique.
Bin Cui 0001, Jialie Shen 0001, Gao Cong, Heng Tao Shen, Cui Yu
ACM Multimedia2
2006 Efficient benchmarking of content-based image retrieval via resampling
abstract
While content-based image retrieval (CBIR) is an expanding field, and new approaches to ever more effective retrieval are frequently proposed, relatively little attention has so far been paid to the process of evaluating the effectiveness of CBIR methods. Most of the reported evaluations use standard IR evaluation methodologies, with little consideration of their statistical significance or appropriateness for CBIR, which makes it difficult to assess the precise impact of individual methods. In this paper, we present a new approach for evaluating CBIR systems which provides both efficient and statistically-sound performance evaluation. The approach is based on stratified sampling, and provides a significant improvement over existing evaluation approaches. Comprehensive experiments using our approach to evaluate a range of CBIR methods have shown that the approach reduces not only the estimation error, but also reduces the size of the test data set required to achieve specific estimation error levels.
Jialie Shen 0001, John Shepherd 0001
ACM Multimedia1
2006 Towards efficient automated singer identification in large music databases
abstract
Automated singer identification is important in organising, browsing and retrieving data in large music databases. In this paper, we propose a novel scheme, called Hybrid Singer Identifier (HSI), for automated singer recognition. HSI can effectively use multiple low-level features extracted from both vocal and non-vocal music segments to enhance the identification process with a hybrid architecture and build profiles of individual singer characteristics based on statistical mixture models. Extensive experimental results conducted on a large music database demonstrate the superiority of our method over state-of-the-art approaches. Categories and Subject Descriptors
Jialie Shen 0001, Bin Cui 0001, John Shepherd 0001, Kian-Lee Tan
SIGIR1
2006 InMAF: indexing music databases via multiple acoustic features
abstract
Music information processing has become very important due to the ever-growing amount of music data from emerging applications. In this demonstration,we present a novel approach for generating small but comprehensive music descriptors to facilitate efficient content music management (accessing and retrieval, in particular). Unlike previous approaches that rely on low-level spectral features adapted from speech analysis technology, our approach integrates human music perception to enhance the accuracy of the retrieval and classification process via PCA and neural networks. The superiority of our method is demonstrated by comparing it with state-of-the-art approaches in the areas of music classification query effectiveness, and robustness against various audio distortion/alternatives.
Jialie Shen 0001, John Shepherd 0001, Anne H. H. Ngu
SIGMOD Conference1
2006 Towards Effective Content-Based Music Retrieval With Multiple Acoustic Feature Combination
abstract
In this paper, we present a new approach to constructing music descriptors to support efficient content-based music retrieval and classification. The system applies multiple musical properties combined with a hybrid architecture based on principal component analysis (PCA) and a multilayer perceptron neural network. This architecture enables straightforward incorporation of multiple musical feature vectors, based on properties such as timbral texture, pitch, and rhythm structure, into a single low-dimensioned vector that is more effective for classification than the larger individual feature vectors. The use of supervised training enables incorporation of human musical perception that further enhances the classification process. We compare our approach with state of the art techniques and demonstrate its effectiveness on content-based music retrieval. In addition, extensive experimental study illustrates its effectiveness and robustness against various kinds of audio alteration.
Jialie Shen 0001, John Shepherd 0001, Anne H. H. Ngu
IEEE Trans. Multim.1
2005 On Efficient Music Genre Classification
Jialie Shen 0001, John Shepherd 0001, Anne H. H. Ngu
DASFAA1
2005 On Effective E-mail Classification via Neural Networks
Bin Cui 0001, Anirban Mondal, Jialie Shen 0001, Gao Cong, Kian-Lee Tan
DEXA3
2005 Semantic-Sensitive Classification for Large Image Libraries
abstract
With advances in multimedia technology, image data with various formats is is becoming available at an explosive rate from various domain applications. How to efficiently organise and access them has been an extremely important issue and enjoying growing attention. In this paper, we present results from experimental studies investigating performance of image classification for a novel dimension reduction scheme with hybrid architecture. We demonstrate that not only can the method provide superior quality of classification accuracy with various machine learning based classifier but also substantially speed up training and categorisation process. Moreover, it is fairly robust against various kinds of visual distortions and noises.
Jialie Shen 0001, John Shepherd 0001, Anne H. H. Ngu
MMM1
2004 Integrating heterogeneous reatures for efficient content based music retrieval
abstract
In this paper, we present a novel feature extraction method facilitating efficient content-based music retrieval and classification, called InMAF. The goal of our approach is to allow straightforward incorporation of multiple musical features, such as timbral texture, pitch and rhythm structure, into a single low dimensional vector that is effective for retrieval and classification. Unlike earlier approaches that used only acoustic properties as the basis for retrieval, our approach can easily incoporate human music perception to improve accuracy of retrieval and classification process. The superiority of our method is demonstrated by comparing it with state-of-the-art approaches in the areas of music classification (using a variety of machine learning algorithms), query effectiveness and robustness against audio distortion.
Jialie Shen 0001, John Shepherd 0001, Anne H. H. Ngu
CIKM1
2004 Improving Query Effectiveness for Large Image Databases with Multiple Visual Feature Combination
Jialie Shen 0001, John Shepherd 0001, Anne H. H. Ngu, Du Q. Huynh
DASFAA1
2003 CMVF: A Novel Dimension Reduction Scheme for Efficient Indexing in A Large Image Database
abstract
No abstract available.
Jialie Shen 0001, Anne H. H. Ngu, John Shepherd 0001, Du Q. Huynh, Quan Z. Sheng
SIGMOD Conference1