Zheng Qin 0003

dblp:95/6861-3 · DBLP profile ↗
← Back
52ranked-venue papers
0as first author
12since 2021 · last 2024
0000-0002-4637-2518ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 6 since 2021Databases, data management, data science and information retrieval · 20 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3Systems, architecture and hardware · 2Software engineering, systems software and programming languages · 2Computer networks · 1Security and privacy · 1
YearPublicationVenuePosition
2024 Improving Image Reconstruction and Synthesis by Balancing the Optimization from Frequency Perspective
abstract
Image reconstruction and synthesis tasks have been boosted by deep learning technologies in recent years. However, most existing deep learning-based transformation methods utilize loss functions in the spatial domain to guide the training process and having unavoidable defects such as image blurring, and checkerboard artifacts in certain cases. In order to improve the quality of generated images, we analyse the training process from frequency perspective, unveil the inherent optimization imbalance between low and high frequency components theoretically, and based upon which we propose the Gradient Balanced Frequency Loss (GBFL) to mitigate the imbalance by taking into account the historical accumulated gradients of each frequency component and reweighting the optimization process accordingly. Extensive quantitative and qualitative experiments demonstrate the versatility and efficiency of our proposed GBFL, as well as robustness to out-of-distribution test samples.
Xuan Dang, Guolong Wang 0001, Zheng Qin 0003
ICME4
2024 Routing Evidence for Unseen Actions in Video Moment Retrieval
abstract
Video moment retrieval (VMR) is a cutting-edge vision-language task locating a segment in a video according to the query. Though the methods have achieved significant performance, they assume that training and testing samples share the same action types, hindering real-world application. In this paper, we specifically consider a new problem: video moment retrieval by queries with unseen actions. We propose a plug-and-play structure, Routing Evidence (RE), with multiple evidence-learning heads and dynamically route one to locate a sentence with an unseen action. Each evidence-learning head estimates the uncertainty while regressing timestamps. We formulate the evidence distribution by a Normal-Inverse Gamma function and design a router to select the most appropriate distribution for a sample. Empirically, we study the efficacy of RE on three updated databases where training and testing samples contain different action types. We find that RE outperforms other state-of-the-art methods with a more robust predictor. Code and data will be available at https://github.com/dieuroi/Routing-Evidence.
Guolong Wang 0001, Zheng Qin 0003, Liangliang Shi
KDD3
2024 Element-Centered Multi-granularity Network for Dense Video Captioning
Xuan Dang, Guolong Wang 0001, Zheng Qin 0003
PRCV (10)4
2023 Instance-Aware Hierarchical Structured Policy for Prompt Learning in Vision-Language Models
abstract
In recent years, learnable prompts have emerged as a major prompt learning paradigm, enhancing the performance of large-scale vision-language pre-trained models in few-shot image classification. However, enhancing methods are often time-consuming and inflexible because 1) class-specific prompts are inefficient in certain situations; 2) instance-specific prompts are put in a fixed position. To address these issues, inspired by the coarse-to-fine decision-making paradigm of human, we propose an Instance-Aware Hierarchical-Structured Policy (IAHSP) that integrates instance-specific prompt selection and appropriate position selection using a reinforcement learning fashion. Specifically, IAHSP consists of two sub-policies: 1) the root policy selects the most suitable prompt from the prompts pool, and 2) the leaf policy identifies the optimal position for inserting the selected prompt. We train these two policies iteratively with rewards constraining the prompts while maintaining their diversity. Extensive experiments on 11 public benchmarks demonstrate that our IAHSP significantly boosts the few-shot image classification performance of vision-language pre-trained models, while also exhibiting superior generalization performance.
Guolong Wang 0001, Zhaoyuan Liu, Xuan Dang, Zheng Qin 0003
ICASSP5
2023 Reducing 0s bias in video moment retrieval with a circular competence-based captioner
Guolong Wang 0001, Zhaoyuan Liu, Zheng Qin 0003
Inf. Process. Manag.4
2022 Low-Pass Graph Convolutional Network for Recommendation
abstract
Spectral graph convolution is extremely time-consuming for large graphs, thus existing Graph Convolutional Networks (GCNs) reconstruct the kernel by a polynomial, which is (almost) fixed. To extract features from the graph data by learning kernels, Low-pass Collaborative Filter Network (LCFN) was proposed as a new paradigm with trainable kernels. However, there are two demerits of LCFN: (1) The hypergraphs in LCFN are constructed by mining 2-hop connections of the user-item bipartite graph, thus 1-hop connections are not used, resulting in serious information loss. (2) LCFN follows the general network structure of GCNs, which is suboptimal. To address these issues, we utilize the bipartite graph to define the graph space directly and explore the best network structure based on experiments. Comprehensive experiments on two real-world datasets demonstrate the effectiveness of the proposed model. Codes are available on https://github.com/Wenhui-Yu/LCFN.
Zheng Qin 0003
AAAI3
2022 MRM2: Multi-Relationship Modeling Module for Multivariate Time Series Classification
abstract
Multivariate Time Series Classification (MTSC) is a prevalent but challenging problem in data mining. With the development of Deep Neural Networks (DNN), hundreds of deep models for MTSC have been proposed. However, most prior works only explicitly model the relationship between time series and classes and ignore the diversity of the relationship, suffering from insufficient information exploitation. In this paper, we propose a novel module named Multi-Relationship Modeling Module(MRM2) for more effective MTSC. MRM2 uses the classified labels to explicitly model not only the relationship between time series and classes, but also the relationship among time series, enabling the backbone to generate distinguishable embeddings. In addition, MRM2 is versatile because it can be combined with the existing backbones of DNN for end-to-end training. Finally, we conduct a series of ablation studies and comparative experiments on the real multivariate time series archive UEA. Experimental results indicate that MRM2 can significantly improve classification performance in most cases. Codes are available on GitHub1.
Pengxiang Shi, Xuan Dang, Wenwen Ye, Zheng Qin 0003
ICDM5
2022 Self-Propagation Graph Neural Network for Recommendation
abstract
In recommendation tasks, we model user preferences by learning node representations (i.e., user and item embeddings) based on the observed user-item interaction data, which is a bipartite graph.GraphNeuralNetworks (GNNs) are widely used to refine the representations by exploring the topology of the graph: embeddings of neighbors are propagated to each node to reconstruct its embeddings. However, the propagation strategy in existing GNNs is empirical and defective: (1) a substantial proportion of links are missed in the sparse observed graph, which causes ineffective and biased propagation; and (2) the propagation weights are determined by a coarse pre-defined rule, which only takes the degree of nodes into consideration. In this paper, we propose a dense and data-driven propagation mechanism for GNNs. Considering the graph we use to propagate embeddings in recommendation tasks is extremely sparse, we complement it and use the predicted graph as the new propagation tool. We learn the propagation matrix from the data, and propose aSelf-propagationGraphNeuralNetwork (SGNN). Since it is very space- and time-consuming to maintain a large and dense propagation matrix, we factorize it for storing and updating. In this paper, we propose three methods to complete the sparse graph and construct the propagation matrix: (1) we complete the graph based on a recommendation model; (2) we measure the node distance based on spectral clustering; (3) we predict missing links of the graph based on predictive embeddings. In SGNN, the embeddings can be propagated to not only the observed neighbors, but also the potential yet unobserved neighbors, and the propagation weights are learned based on the connection strength. Comprehensive experiments on three real-world datasets demonstrate the effectiveness and efficiency of our proposed model: SGNN outperforms recent state-of-the-art GNNs significantly. Codes are available onhttps://github.com/Wenhui-Yu/LCFN.
Jinfei Liu, Junfeng Ge, Wenwu Ou, Zheng Qin 0003
IEEE Trans. Knowl. Data Eng.6
2021 Dense Video Captioning for Incomplete Videos
Xuan Dang, Guolong Wang 0001, Kun Xiong, Zheng Qin 0003
ICANN (5)4
2021 Dual Convolutional Neural Network for Lung Nodule Classification
abstract
In some computer vision tasks, in addition to images, we also have a graph indicating the relationship of the images. In this case, we can use Graph Convolutional Network (GCN) to get extra information from the graph. GCN is a widely-used deep neural network with graph convolutional layers extracting high-level feature maps from the graph. In this paper, we design a Dual Convolution Neural Network (Dual-CNN) with a CNN solving the images and a GCN solving the graph. We take the lung nodule classification task as an example to introduce our model. In this task, we have features including a CT image and several categorical attributes for each nodule, and the purpose is to classify if the nodule is malignant. Conventional classifiers take these features as the input and then give the prediction, yet we additionally construct a graph based on attributes and perform graph convolution on it. We also design experiments to demonstrate the effectiveness of our Dual-CNN models. Codes are available on https://github.com/temp2244/lung_nodules_classification.
Pengxiang Shi, Zheng Qin 0003
IJCNN4
2021 Self-Supervised Pre-training for Time Series Classification
abstract
Recently, significant progress has been made in time series classification with deep learning. However, using deep learning models to solve time series classification generally suffers from expensive calculations and difficulty of data labeling. In this work, we study self-supervised time series pre-training to overcome these challenges. Compared with the existing works, we focus on the universal and unlabeled time series pretraining. To this end, we propose a novel end-to-end neural network architecture based on self-attention, which is suitable for capturing long-term dependencies and extracting features from different time series. Then, we propose two different self-supervised pretext tasks for time series data type: Denoising and Similarity Discrimination based on DTW (Dynamic Time Warping). Finally, we carry out extensive experiments on 85 time series datasets (also known as UCR2015 [2]). Empirical results show that the time series model augmented with our proposed self-supervised pretext tasks achieves state-of-the-art / highly competitive results.
Pengxiang Shi, Wenwen Ye, Zheng Qin 0003
IJCNN3
2021 Visually aware recommendation with aesthetic features
Xiangnan He 0001, Jian Pei 0001, Xu Chen 0017, Li Xiong 0001, Jinfei Liu, Zheng Qin 0003
VLDB J.7
2020 Game Recommendation Based on Dynamic Graph Convolutional Network
Wenwen Ye, Zheng Qin 0003, Zhuoye Ding, Dawei Yin 0001
DASFAA (1)2
2020 Graph Convolutional Network for Recommendation with Low-pass Collaborative Filters
abstract
\textbf{G}raph \textbf{C}onvolutional \textbf{N}etwork (\textbf{GCN}) is widely used in graph data learning tasks such as recommendation. However, when facing a large graph, the graph convolution is very computationally expensive thus is simplified in all existing GCNs, yet is seriously impaired due to the oversimplification. To address this gap, we leverage the \emph{original graph convolution} in GCN and propose a \textbf{L}ow-pass \textbf{C}ollaborative \textbf{F}ilter (\textbf{LCF}) to make it applicable to the large graph. LCF is designed to remove the noise caused by exposure and quantization in the observed data, and it also reduces the complexity of graph convolution in an unscathed way. Experiments show that LCF improves the effectiveness and efficiency of graph convolution and our GCN outperforms existing GCNs significantly. Codes are available on \url{https://github.com/Wenhui-Yu/LCFN}.
Zheng Qin 0003
ICML2
2020 Towards Personalized Aesthetic Image Caption
abstract
Image captioning (IC) is a commonly-used technique for generating textual image description, which finds its applications on semantic image retrieval and multi-modal image understanding, among many others. This paper focuses on an important IC method specialized for generating aesthetic descriptions of images, i.e., aesthetic image captioning (AIC). Despite some effectiveness of initial work on AIC, their performances are inherently limited due to a lack of consideration of user preferences on aesthetics and better aesthetic feature, making it unusable for real-world applications where human users present a large variation on evaluating visual aesthetics of images. To tackle this, we propose a novel personalized aesthetic image caption (PAIC) approach for capturing and incorporating user preferences for AIC tasks. Our approach mainly contains Aesthetic feature Extraction Network(AEN), User Encoder network(UEN) and a personalized image caption model. AEN is designed to extract more expressive feature, UEN is introduced for learning the user vector from the limited information in our AVA-PCap dataset. Personalized image caption model is constructed to generate the caption when given the user id and photo pairs. The experimental results show that our methods outperform baselines by 10% , which is encouraging for a first step towards personalized aesthetic image caption.
Kun Xiong, Liu Jiang, Xuan Dang, Guolong Wang 0001, Wenwen Ye, Zheng Qin 0003
IJCNN6
2020 Semi-supervised Collaborative Filtering by Text-enhanced Domain Adaptation
abstract
Data sparsity is an inherent challenge in the recommender systems, where most of the data is collected from the implicit feedbacks of users. This causes two difficulties in designing effective algorithms: first, the majority of users only have a few interactions with the system and there is no enough data for learning; second, there are no negative samples in the implicit feedbacks and it is a common practice to perform negative sampling to generate negative samples. However, this leads to a consequence that many potential positive samples are mislabeled as negative ones and data sparsity would exacerbate the mislabeling problem.
Junfeng Ge, Wenwu Ou, Zheng Qin 0003
KDD5
2020 Learning to Select Elements for Graphic Design
abstract
Selecting elements for graphic design is essential for ensuring a correct understanding of clients' requirements as well as improving the efficiency of designers before a fine-designed process. Some semi-automatic design tools proposed layout templates where designers always select elements according to the rectangular boxes that specify how elements are placed. In practice, layout and element selection are complementary. Compared to the layout which can be readily obtained from pre-designed templates, it is generally time-consuming to mindfully pick out suitable elements, which calls for an automation of elements selection. To address this, we formulate element selection as a sequential decision-making process and develop a deep element selection network (DESN). Given a layout file with annotated elements, new graphical elements are selected to form graphic designs based on aesthetics and consistency criteria. To train our DESN, we propose an end-to-end, reinforcement learning based framework, where we design a novel reward function that jointly accounts for visual aesthetics and consistency. Based on this, visually readable and aesthetic drafts can be efficiently generated. We further contribute a layout-poster dataset with exhaustively labeled attributes of poster key elements. Qualitative and quantitative results indicate the efficacy of our approach.
Guolong Wang 0001, Zheng Qin 0003, Junchi Yan, Liu Jiang
ICMR2
2020 Time Matters: Sequential Recommendation with Complex Temporal Information
abstract
Incorporating temporal information into recommender systems has recently attracted increasing attention from both the industrial and academic research communities. Existing methods mostly reduce the temporal information of behaviors to behavior sequences for subsequently RNN-based modeling. In such a simple manner, crucial time-related signals have been largely neglected. This paper aims to systematically investigate the effects of the temporal information in sequential recommendations. In particular, we firstly discover two elementary temporal patterns of user behaviors: "absolute time patterns'' and "relative time patterns'', where the former highlights user time-sensitive behaviors, e.g., people may frequently interact with specific products at certain time point, and the latter indicates how time interval influences the relationship between two actions. For seamlessly incorporating these information into a unified model, we devise a neural architecture that jointly learns those temporal patterns to model user dynamic preferences. Extensive experiments on real-world datasets demonstrate the superiority of our model, comparing with the state-of-the-arts.
Wenwen Ye, Shuaiqiang Wang, Xu Chen 0017, Xuepeng Wang, Zheng Qin 0003, Dawei Yin 0001
SIGIR5
2020 Sampler Design for Implicit Feedback Data by Noisy-label Robust Learning
abstract
Implicit feedback data is extensively explored in recommendation as it is easy to collect and generally applicable. However, predicting users' preference on implicit feedback data is a challenging task since we can only observe positive (voted) samples and unvoted samples. It is difficult to distinguish between the negative samples and unlabeled positive samples from the unvoted ones. Existing works, such as Bayesian Personalized Ranking (BPR), sample unvoted items as negative samples uniformly, therefore suffer from a critical noisy-label issue. To address this gap, we design an adaptive sampler based on noisy-label robust learning for implicit feedback data.
Zheng Qin 0003
SIGIR2
2020 IMCFN: Image-based malware classification using fine-tuned convolutional neural network architecture
Danish Vasan, Mamoun Alazab, Sobia Wassan, Hamad Naeem, Babak Safaei, Zheng Qin 0003
Comput. Networks6
2020 Image-Based malware classification using ensemble of CNN architectures (IMCEC)
Danish Vasan, Mamoun Alazab, Sobia Wassan, Babak Safaei, Zheng Qin 0003
Comput. Secur.5
2020 A lightweight and aggregated system for indoor/outdoor detection using smart devices
Zheng Qin 0003, Houbing Song, Chengxiang Si, Renwei Zhang
Future Gener. Comput. Syst.2
2020 MTHAEL: Cross-Architecture IoT Malware Detection Based on Neural Network Advanced Ensemble Learning
abstract
The complexity, sophistication, and impact of malware evolve with industrial revolution and technology advancements. This article discusses and proposes a robust cross-architecture IoT malware threat hunting model based on advanced ensemble learning (MTHAEL). Our unique MTHAEL model using stacked ensemble of heterogeneous feature selection algorithms and state-of-the-art neural networks to learn different levels of semantic features demonstrates enhanced IoT malware detection than existing approaches. MTHAEL is the first of its kind that effectively optimizes recurrent neural network (RNN) and convolutional neural network (CNN) with high classification accuracy and consistently low computational overheads on different IoT architectures. Cross-architecture benchmarking is performed during the training with different architectures such as ARM, Intel80386, MIPS, and MIPS+Intel80386 individually. Two different hardware architectures were employed to analyze the architecture overhead, namely Raspberry Pi 4 (ARM-based architecture) and Core-i5 (Intel-based architecture). Our proposed MTHAEL is evaluated comprehensively with a large IoT cross-architecture dataset of 21,137 samples and has achieved 99.98 percent classification accuracy for ARM architecture samples, surpassing prior related works. Overall, MTHAEL has demonstrated practical suitability for cross-architecture IoT malware detection with low computational overheads requiring only 0.32 seconds to detect Any IoT malware.
Danish Vasan, Mamoun Alazab, Sitalakshmi Venkatraman, Junaid Akram, Zheng Qin 0003
IEEE Trans. Computers5
2020 Efficient Contour Computation of Group-Based Skyline
abstract
Skyline, aiming at finding a Pareto optimal subset of points in a multi-dimensional dataset, has gained great interest due to its extensive use for multi-criteria analysis and decision making. The skyline consists of all points that are not dominated by any other points. It is a candidate set of the optimal solution, which depends on a specific evaluation criterion for optimum. However, conventional skyline queries, which return individual points, are inadequate in group querying case since optimal combinations are required. To address this gap, we study the skyline computation in the group level and propose efficient methods to find the Group-based skyline (G-skyline). For computing the front l skyline layers, we lay out an efficient approach that does the search concurrently on each dimension and investigates each point in the subspace. After that, we present a novel structure to construct the G-skyline with a queue of combinations of the first-layer points. We further demonstrate that the G-skyline is a complete candidate set of top-l solutions, which is the main superiority over previous group-based skyline definitions. However, as G-skyline is complete, it contains a large number of groups which can make it impractical. To represent the “contour” of the G-skyline, we define the Representative G-skyline (RG-skyline). Then, we propose a Group-based clustering (G-clustering) algorithm to find out RG-skyline groups. Experimental results show that our algorithms are several orders of magnitude faster than the previous work.
Jinfei Liu, Jian Pei 0001, Li Xiong 0001, Xu Chen 0017, Zheng Qin 0003
IEEE Trans. Knowl. Data Eng.6
2019 Dynamic Explainable Recommendation Based on Neural Attentive Models
abstract
Providing explanations in a recommender system is getting more and more attention in both industry and research communities. Most existing explainable recommender models regard user preferences as invariant to generate static explanations. However, in real scenarios, a user’s preference is always dynamic, and she may be interested in different product features at different states. The mismatching between the explanation and user preference may degrade costumers’ satisfaction, confidence and trust for the recommender system. With the desire to fill up this gap, in this paper, we build a novel Dynamic Explainable Recommender (called DER) for more accurate user modeling and explanations. In specific, we design a time-aware gated recurrent unit (GRU) to model user dynamic preferences, and profile an item by its review information based on sentence-level convolutional neural network (CNN). By attentively learning the important review information according to the user current state, we are not only able to improve the recommendation performance, but also can provide explanations tailored for the users’ current preferences. We conduct extensive experiments to demonstrate the superiority of our model for improving recommendation performance. And to evaluate the explainability of our model, we first present examples to provide intuitive analysis on the highlighted review information, and then crowd-sourcing based evaluations are conducted to quantitatively verify our model’s superiority.
Xu Chen 0017, Yongfeng Zhang 0003, Zheng Qin 0003
AAAI3
2019 Intrusion Detection via Wide and Deep Model
Zheng Qin 0003, Pengbo Shen
ICANN (4)2
2019 Delving into Precise Attention in Image Captioning
Shaohan Hu, Shenglei Huang, Guolong Wang 0001, Zheng Qin 0003
ICONIP (5)5
2019 An Expert Validation Framework for Improving the Quality of Crowdsourced Clustering
Liu Jiang, Zheng Qin 0003, Pengbo Shen, Shaohan Hu
ICONIP (5)2
2019 Zero-Shot Learning for Intrusion Detection via Attribute Representation
Zheng Qin 0003, Pengbo Shen, Liu Jiang
ICONIP (1)2
2019 Intrusion Detection Using Temporal Convolutional Networks
Zheng Qin 0003, Pengbo Shen, Liu Jiang
ICONIP (4)2
2019 Personalized Fashion Recommendation with Visual Explanations based on Multimodal Attention Network: Towards Visually Explainable Recommendation
abstract
Fashion recommendation has attracted increasing attention from both industry and academic communities. This paper proposes a novel neural architecture for fashion recommendation based on both image region-level features and user review information. Our basic intuition is that: for a fashion image, not all the regions are equally important for the users, i.e., people usually care about a few parts of the fashion image. To model such human sense, we learn an attention model over many pre-segmented image regions, based on which we can understand where a user is really interested in on the image, and correspondingly, represent the image in a more accurate manner. In addition, by discovering such fine-grained visual preference, we can visually explain a recommendation by highlighting some regions of its image. For better learning the attention model, we also introduce user review information as a weak supervision signal to collect more comprehensive user preference. In our final framework, the visual and textual features are seamlessly coupled by a multimodal attention network. Based on this architecture, we can not only provide accurate recommendation, but also can accompany each recommended item with novel visual explanations. We conduct extensive experiments to demonstrate the superiority of our proposed model in terms of Top-N recommendation, and also we build a collectively labeled dataset for evaluating our provided visual explanations in a quantitative manner.
Xu Chen 0017, Hanxiong Chen, Hongteng Xu, Yongfeng Zhang 0003, Yixin Cao 0002, Zheng Qin 0003, Hongyuan Zha
SIGIR6
2019 Spectrum-enhanced Pairwise Learning to Rank
abstract
To enhance the performance of the recommender system, side information is extensively explored with various features (e.g., visual features and textual features). However, there are some demerits of side information: (1) the extra data is not always available in all recommendation tasks; (2) it is only for items, there is seldom high-level feature describing users. To address these gaps, we introduce the spectral features extracted from two hypergraph structures of the purchase records. Spectral features describe the similarity of users/items in the graph space, which is critical for recommendation. We leverage spectral features to model the users' preference and items' properties by incorporating them into a Matrix Factorization (MF) model.
Zheng Qin 0003
WWW2
2019 Adversarial Distillation for Efficient Recommendation with External Knowledge
abstract
Integrating external knowledge into the recommendation system has attracted increasing attention in both industry and academic communities. Recent methods mostly take the power of neural network for effective knowledge representation to improve the recommendation performance. However, the heavy deep architectures in existing models are usually incorporated in an embedded manner, which may greatly increase the model complexity and lower the runtime efficiency. To simultaneously take the power of deep learning for external knowledge modeling as well as maintaining the model efficiency at test time, we reformulate the problem of recommendation with external knowledge into a generalized distillation framework . The general idea is to free the complex deep architecture into a separate model, which is only used in the training phrase, while abandoned at test time. In particular, in the training phrase, the external knowledge is processed by a comprehensive teacher model to produce valuable information to teach a simple and efficient student model. Once the framework is learned, the teacher model is abandoned, and only the succinct yet enhanced student model is used to make fast predictions at test time. In this article, we specify the external knowledge as user review, and to leverage it in an effective manner, we further extend the traditional generalized distillation framework by designing a Selective Distillation Network (SDNet) with adversarial adaption and orthogonality constraint strategies to make it more robust to noise information. Extensive experiments verify that our model can not only improve the performance of rating prediction, but also can significantly reduce time consumption when making predictions as compared with several state-of-the-art methods.
Xu Chen 0017, Yongfeng Zhang 0003, Hongteng Xu, Zheng Qin 0003, Hongyuan Zha
ACM Trans. Inf. Syst.4
2018 Bridge Video and Text with Cascade Syntactic Structure
abstract
We present a video captioning approach that encodes features by progressively completing syntactic structure (LSTM-CSS). To construct basic syntactic structure (i.e., subject, predicate, and object), we use a Conditional Random Field to label semantic representations (i.e., motions, objects). We argue that in order to improve the comprehensiveness of the description, the local features within object regions can be used to generate complementary syntactic elements (e.g., attribute, adverbial). Inspired by redundancy of human receptors, we utilize a Region Proposal Network to focus on the object regions. To model the final temporal dynamics, Recurrent Neural Network with Path Embeddings is adopted. We demonstrate the effectiveness of LSTM-CSS on generating natural sentences: 42.3% and 28.5% in terms of BLEU@4 and METEOR. Superior performance when compared to state-of-the-art methods are reported on a large video description dataset (i.e., MSR-VTT-2016).
Guolong Wang 0001, Zheng Qin 0003, Kaiping Xu, Shuxiong Ye
COLING2
2018 Deep Collaborative Filtering Combined with High-Level Feature Generation on Latent Factor Model
Xu Chen 0017, Zheng Qin 0003
ICONIP (1)3
2018 A Semantic Parsing Based LSTM Model for Intrusion Detection
Zheng Qin 0003
ICONIP (4)2
2018 Deep Tag Recommendation Based on Discrete Tensor Factorization
Wenwen Ye, Zheng Qin 0003
ICONIP (1)2
2018 Collaborative and Attentive Learning for Personalized Image Aesthetic Assessment
abstract
The ever-increasing volume of visual images has stimulated the demand for organizing such data by aesthetic quality. Automatic and especially learning based aesthetic assessment methods have shown potential by recent works. Existing image aesthetic prediction is often user-agnostic which may ignore the fact that the rating to an image can be inherently individual. We fill this gap by formulating the personalized image aesthetic assessment problem with a novel learning method. Specifically, we collect user-image textual reviews in addition with visual images from the public dataset to organize a review-augmented benchmark. Using this enriched dataset, we devise a deep neural network with a user/image relation encoding input for collaborative filtering. Meanwhile an attentive mechanism is designed to capture the user-specific taste for image semantic tags and regions of interest by fusing the image and user's review. Extensive and promising experimental results on the review-augmented benchmark corroborate the efficacy of our approach.
Guolong Wang 0001, Junchi Yan, Zheng Qin 0003
IJCAI3
2018 Collision-Free LSTM for Human Trajectory Prediction
Kaiping Xu, Zheng Qin 0003, Guolong Wang 0001, Shuxiong Ye, Huidi Zhang
MMM (1)2
2018 Sequential Recommendation with User Memory Networks
abstract
User preferences are usually dynamic in real-world recommender systems, and a user»s historical behavior records may not be equally important when predicting his/her future interests. Existing recommendation algorithms -- including both shallow and deep approaches -- usually embed a user»s historical records into a single latent vector/representation, which may have lost the per item- or feature-level correlations between a user»s historical records and future interests. In this paper, we aim to express, store, and manipulate users» historical records in a more explicit, dynamic, and effective manner. To do so, we introduce the memory mechanism to recommender systems. Specifically, we design a memory-augmented neural network (MANN) integrated with the insights of collaborative filtering for recommendation. By leveraging the external memory matrix in MANN, we store and update users» historical records explicitly, which enhances the expressiveness of the model. We further adapt our framework to both item- and feature-level versions, and design the corresponding memory reading/writing operations according to the nature of personalized recommendation scenarios. Compared with state-of-the-art methods that consider users» sequential behavior for recommendation, e.g., sequential recommenders with recurrent neural networks (RNN) or Markov chains, our method achieves significantly and consistently better performance on four real-world datasets. Moreover, experimental analyses show that our method is able to extract the intuitive patterns of how users» future actions are affected by previous behaviors.
Xu Chen 0017, Hongteng Xu, Yongfeng Zhang 0003, Jiaxi Tang, Yixin Cao 0002, Zheng Qin 0003, Hongyuan Zha
WSDM6
2018 Aesthetic-based Clothing Recommendation
abstract
Recently, product images have gained increasing attention in clothing recommendation since the visual appearance of clothing products has a significant impact on consumers» decision. Most existing methods rely on conventional features to represent an image, such as the visual features extracted by convolutional neural networks (CNN features) and the scale-invariant feature transform algorithm (SIFT features), color histograms, and so on. Nevertheless, one important type of features, the aesthetic features, is seldom considered. It plays a vital role in clothing recommendation since a users» decision depends largely on whether the clothing is in line with her aesthetics, however the conventional image features cannot portray this directly. To bridge this gap, we propose to introduce the aesthetic information, which is highly relevant with user preference, into clothing recommender systems. To achieve this, we first present the aesthetic features extracted by a pre-trained neural network, which is a brain-inspired deep structure trained for the aesthetic assessment task. Considering that the aesthetic preference varies significantly from user to user and by time, we then propose a new tensor factorization model to incorporate the aesthetic features in a personalized manner. We conduct extensive experiments on real-world datasets, which demonstrate that our approach can capture the aesthetic preference of users and significantly outperform several state-of-the-art recommendation methods.
Huidi Zhang, Xiangnan He 0001, Xu Chen 0017, Li Xiong 0001, Zheng Qin 0003
WWW6
2017 Fast Algorithms for Pareto Optimal Group-based Skyline
abstract
Skyline, aiming at finding a Pareto optimal subset of points in a multi-dimensional dataset, has gained great interest due to its extensive use for multi-criteria analysis and decision making. Skyline consists of all points that are not dominated by, or not worse than other points. It is a candidate set of optimal solution, which depends on a specific evaluation criterion for optimum. However, conventional skyline queries, which return individual points, are inadequate in group querying case since optimal combinations are required. To address this gap, we study the skyline computation in group case and propose fast methods to find the group-based skyline (G-skyline), which contains Pareto optimal groups. For computing the front k skyline layers, we lay out an efficient approach that does the search concurrently on each dimension and investigates each point in subspace. After that, we present a novel structure to construct the G-skyline with a queue of combinations of the first-layer points. Experimental results show that our algorithms are several orders of magnitude faster than the previous work.
Zheng Qin 0003, Jinfei Liu, Li Xiong 0001, Xu Chen 0017, Huidi Zhang
CIKM2
2017 Intrusion Detection Using Convolutional Neural Networks for Representation Learning
Zheng Qin 0003, Shuxiong Ye
ICONIP (5)2
2017 Recognizing Emotions Based on Human Actions in Videos
Guolong Wang 0001, Zheng Qin 0003, Kaiping Xu
MMM (2)2
2017 Personalized Key Frame Recommendation
abstract
Key frames are playing a very important role for many video applications, such as on-line movie preview and video information retrieval. Although a number of key frame selection methods have been proposed in the past, existing technologies mainly focus on how to precisely summarize the video content, but seldom take the user preferences into consideration. However, in real scenarios, people may cast diverse interests on the contents even for the same video, and thus they may be attracted by quite different key frames, which makes the selection of key frames an inherently personalized process. In this paper, we propose and investigate the problem of personalized key frame recommendation to bridge the above gap. To do so, we make use of video images and user time-synchronized comments to design a novel key frame recommender that can simultaneously model visual and textual features in a unified framework. By user personalization based on her/his previously reviewed frames and posted comments, we are able to encode different user interests in a unified multi-modal space, and can thus select key frames in a personalized manner, which, to the best of our knowledge, is the first time in the research field of video content analysis. Experimental results show that our method performs better than its competitors on various measures.
Xu Chen 0017, Yongfeng Zhang 0003, Qingyao Ai, Hongteng Xu, Junchi Yan, Zheng Qin 0003
SIGIR6
2016 Human activities prediction by learning combinatorial sparse representations
abstract
Human activities prediction is to enable early recognition of unfinished activities from videos only containing the beginning parts, which is a challenge problem. Prediction of human activities is necessarily applied in particular scenes(e.g. surveillance systems, human-computer interfaces). To solve this problem, we propose a novel framework which classifies videos into activity classes by using combinatorial sparse representations (CSR). The major contributions of our work include: (1) dividing each video into multiple equal-length segments, where the local spatio-temporal features extracted; (2) concatenating combinatorial sparse activity dictionaries, formed by overcomplete dictionary of each segment; (3) computing combinatorial sparse coefficients of each segment, based on activity dictionaries above; (4) formulating probability of each activity to estimate the correct class. The results of our experiments show that the proposed method outperforms existing state-of-the-art comparison methods.
Kaiping Xu, Zheng Qin 0003, Guolong Wang 0001
ICIP2
2016 Recognize human activities from multi-part missing videos
abstract
Recognizing human activities from multi-part missing videos is a challenge problem. When the multiple missing parts are continuous, the problem is reduced to activity recognition in videos with single part missing at any position which is focused on by many researches. However, in many practical applications, some temporal gaps always appear in captured videos due to random frame loss(e.g. noise interfere). To solve this problem, we propose a novel framework: 1) dividing each video into multiple equal-length segments, where the local spatio-temporal features extracted; 2) concatenating combinatorial sparse activity dictionaries, formed by over-complete dictionary of each segment; 3) computing combinatorial sparse coefficients of each segment, based on activity dictionaries above; 4) formulating probability of each activity to estimate the correct class. Our experiments achieve superior performance not only in videos with single part missing at any position, but also in videos with multiple parts missing.
Kaiping Xu, Zheng Qin 0003, Guolong Wang 0001
ICME2
2016 Learning to Rank Features for Recommendation over Multiple Categories
abstract
Incorporating phrase-level sentiment analysis on users' textual reviews for recommendation has became a popular meth-od due to its explainable property for latent features and high prediction accuracy. However, the inherent limitations of the existing model make it difficult to (1) effectively distinguish the features that are most interesting to users, (2) maintain the recommendation performance especially when the set of items is scaled up to multiple categories, and (3) model users' implicit feedbacks on the product features. In this paper, motivated by these shortcomings, we first introduce a tensor matrix factorization algorithm to Learn to Rank user Preferences based on Phrase-level sentiment analysis across Multiple categories (LRPPM for short), and then by combining this technique with Collaborative Filtering (CF) method, we propose a novel model called LRPPM-CF to boost the performance of recommendation. Thorough experiments on two real-world datasets demonstrate that our proposed model is able to improve the performance in the tasks of capturing users' interested features and item recommendation by about 17%-24% and 7%-13%, respectively, as compared with several state-of-the-art methods.
Xu Chen 0017, Zheng Qin 0003, Yongfeng Zhang 0003
SIGIR2
2016 Composite-based conflict resolution in merging versions of UML models
abstract
Model-driven engineering is now playing an essential role in software development. Adequate model versioning systems are critical to enable efficient team-based development of models. The state-of-art model versioning systems are able to detect and help resolving basic conflicts which arise during the merging of different model versions. However, conflict resolution is typically conducted at the primitive operation level in operation-based system and user interaction is required to choose from the conflicting operations. In this study, we present an approach to resolve conflicts automatically at composite level in model versioning systems for Unified Modeling Language (UML). This approach has two main stages. During the merging stage, a temporary merged model is generated, which represent the central intention of model developers. And during the conflict resolution stage, our approach automatically finds and presents to the model developers all solutions for resolving all inconsistencies in the merged model. The approach was empirically evaluated on a range of test models and proved to be scalable to models of large size.
Hao Chong, Renwei Zhang, Zheng Qin 0003
SNPD3
2015 From LTL Formulae to Büchi Automata: A Direct Translation Using On-the-Fly De-Generalization
abstract
In this paper, we present a conversion algorithm to translate a linear temporal logic (LTL) formula to a Büchi automaton (BA) directly. A label, acceptance degree (AD), is presented to record acceptance conditions satisfied in each state or transition of an automaton. The AD for an automaton is a set of {U, F, R, G}-subformula of the given LTL formula. According to ADs attached to states and transitions, the on-the-fly de-generalization algorithm is presented. This on-the-fly de-generalization algorithm is used to transform a generalized Büchi automaton (GBA) into a Büchi automaton. It is different from the execution of the classic de-generalization algorithm that the on-the-fly de-generalization algorithm is performed during the expansion of the given LTL formula. A direct conversion algorithm based on the on-the-fly de-generalization algorithm is conceived and implemented. We compare the conversion algorithm presented in this paper with previous works, and show that it is more efficient for a series of formulae in usual use and random formulae generated by LBTT 1.2.1 (an LTL-to-BA translator testbench).
Lai-Xiang Shan, Zheng Qin 0003, Kaiping Xu, Xu Chen 0017
APSEC2
2011 Mining Uncertain Data Streams Using Clustering Feature Decision Trees
Wenhua Xu, Zheng Qin 0003
ADMA (2)2
2011 Clustering feature decision trees for semi-supervised classification from high-speed data streams
abstract
Most stream data classification algorithms apply the supervised learning strategy which requires massive labeled data. Such approaches are impractical since labeled data are usually hard to obtain in reality. In this paper, we build a clustering feature decision tree model, CFDT, from data streams having both unlabeled and a small number of labeled examples. CFDT applies a micro-clustering algorithm that scans the data only once to provide the statistical summaries of the data for incremental decision tree induction. Micro-clusters also serve as classifiers in tree leaves to improve classification accuracy and reinforce the any-time property. Our experiments on synthetic and real-world datasets show that CFDT is highly scalable for data streams while generating high classification accuracy with high speed.
Wenhua Xu, Zheng Qin 0003, Yang Chang
J. Zhejiang Univ. Sci. C2