Yan Li 0068

dblp:87/660-68 · DBLP profile ↗
← Back
21ranked-venue papers
2as first author
6since 2021 · last 2022
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 since 2021Artificial intelligence and machine learning · 9 · 4 since 2021Systems, architecture and hardware · 3 · 2 first-authorDatabases, data management, data science and information retrieval · 3Computer networks · 1
YearPublicationVenuePosition
2022 Neighborhood-Adaptive Structure Augmented Metric Learning
abstract
Most metric learning techniques typically focus on sample embedding learning, while implicitly assume a homogeneous local neighborhood around each sample, based on the metrics used in training ( e.g., hypersphere for Euclidean distance or unit hyperspherical crown for cosine distance). As real-world data often lies on a low-dimensional manifold curved in a high-dimensional space, it is unlikely that everywhere of the manifold shares the same local structures in the input space. Besides, considering the non-linearity of neural networks, the local structure in the output embedding space may not be homogeneous as assumed. Therefore, representing each sample simply with its embedding while ignoring its individual neighborhood structure would have limitations in Embedding-Based Retrieval (EBR). By exploiting the heterogeneity of local structures in the embedding space, we propose a Neighborhood-Adaptive Structure Augmented metric learning framework (NASA), where the neighborhood structure is realized as a structure embedding, and learned along with the sample embedding in a self-supervised manner. In this way, without any modifications, most indexing techniques can be used to support large-scale EBR with NASA embeddings. Experiments on six standard benchmarks with two kinds of embeddings, i.e., binary embeddings and real-valued embeddings, show that our method significantly improves and outperforms the state-of-the-art methods.
Pandeng Li, Yan Li 0068, Hongtao Xie 0001, Lei Zhang 0119
AAAI2
2021 Deep Metric Learning with Self-Supervised Ranking
abstract
Deep metric learning aims to learn a deep embedding space, where similar objects are pushed towards together and different objects are repelled against. Existing approaches typically use inter-class characteristics, e.g. class-level information or instance-level similarity, to obtain semantic relevance of data points and get a large margin between different classes in the embedding space. However, the intra-class characteristics, e.g. local manifold structure or relative relationship within the same class, are usually overlooked in the learning process. Hence the data structure cannot be fully exploited and the output embeddings have limitation in retrieval. More importantly, retrieval results lack in a good ranking. This paper presents a novel self-supervised ranking auxiliary framework, which captures intra-class characteristics as well as inter-class characteristics for better metric learning. Our method defines specific transform functions to simulates the local structure change of intra-class in the initial image domain, and formulates a self-supervised learning procedure to fully exploit this property and preserve it in the embedding space. Extensive experiments on three standard benchmarks show that our method significantly improves and outperforms the state-of-the-art methods on the performances of both retrieval and ranking by 2%-4%.
Zheren Fu, Yan Li 0068, Zhendong Mao 0001, Quan Wang 0002, Yongdong Zhang 0001
AAAI2
2021 Query-Memory Re-Aggregation for Weakly-supervised Video Object Segmentation
abstract
Weakly-supervised video object segmentation (WVOS) is an emerging video task that can track and segment the target given a simple bounding box label. However, existing WVOS methods are still unsatisfied in either speed or accuracy, since they only use the exemplar frame to guide the prediction while they neglect the reference from other frames. To solve the problem, we propose a novel Re-Aggregation based framework, which uses feature matching to efficiently find the target and capture the temporal dependencies from multiple frames to guide the segmentation. Based on a two-stage structure, our framework builds an information-symmetric matching process to achieve robust aggregation. In each stage, we design a Query-Memory Aggregation (QMA) module to gather features from the past frames and make bidirectional aggregation to adaptively weight the aggregated features, which relieves the latent misguidance in unidirectional aggregation. To further exploit the information from different aggregation stages, we propose a novel coarse-fine constraint by using the Cascaded Refinement Module (CRM) to combine the predictions from different stages and further boosts the performance. Experimental results on three benchmarks show that our method achieves the state-of-the-art performance in WVOS (e.g., an overall score of 84.7% on the DAVIS 2016 validation set).
Fanchao Lin, Hongtao Xie 0001, Yan Li 0068, Yongdong Zhang 0001
AAAI3
2021 Complementary Factorization towards Outfit Compatibility Modeling
abstract
Recently, outfit compatibility modeling, which aims to evaluate the compatibility of a given outfit that comprises a set of fashion items, has gained growing research attention. Although existing studies have achieved prominent progress, most of them overlook the essential global outfit representation learning, and the hidden complementary factors behind the outfit compatibility uncovering. Towards this end, we propose an Outfit Compatibility Modeling scheme via Complementary Factorization, termed as OCM-CF. In particular, OCM-CF consists of two key components: context-aware outfit representation modeling and hidden complementary factors modeling. The former works on adaptively learning the global outfit representation with graph convolutional networks and the multi-head attention mechanism, where the item context is fully explored. The latter targets at uncovering the latent complementary factors with multiple parallel networks, each of which corresponds to a factor-oriented context-aware outfit representation modeling. In this part, a new orthogonality-based complementarity regularization is proposed to encourage the learned factors to complement each other and better characterize the outfit compatibility. Finally, the outfit compatibility is obtained by summing all the hidden complementary factor-oriented outfit compatibility scores, each of which is derived from the corresponding outfit representation. Extensive experiments on two real-world datasets demonstrate the superiority of our OCM-CF over the state-of-the-art methods.
Xuemeng Song, Weili Guan, Yan Li 0068, Liqiang Nie
ACM Multimedia5
2021 Contrastive Learning for Cold-Start Recommendation
abstract
Recommending purely cold-start items is a long-standing and fundamental challenge in the recommender systems. Without any historical interaction on cold-start items, the collaborative filtering (CF) scheme fails to leverage collaborative signals to infer user preference on these items. To solve this problem, extensive studies have been conducted to incorporate side information of items (e.g. content features) into the CF scheme. Specifically, they employ modern neural network techniques (e.g., dropout, consistency constraint) to discover and exploit the coalition effect of content features and collaborative representations. However, we argue that these works less explore the mutual dependencies between content features and collaborative representations and lack sufficient theoretical supports, thus resulting in unsatisfactory performance on cold-start recommendation.
Yinwei Wei, Xiang Wang 0010, Liqiang Nie, Yan Li 0068, Tat-Seng Chua
ACM Multimedia5
2021 PRRNet: Pixel-Region relation network for face forgery detection
Zhihua Shang, Hongtao Xie 0001, Zhengjun Zha, Lingyun Yu 0002, Yan Li 0068, Yongdong Zhang 0001
Pattern Recognit.5
2020 PEIA: Personality and Emotion Integrated Attentive Model for Music Recommendation on Social Media Platforms
abstract
With the rapid expansion of digital music formats, it's indispensable to recommend users with their favorite music. For music recommendation, users' personality and emotion greatly affect their music preference, respectively in a long-term and short-term manner, while rich social media data provides effective feedback on these information. In this paper, aiming at music recommendation on social media platforms, we propose a Personality and Emotion Integrated Attentive model (PEIA), which fully utilizes social media data to comprehensively model users' long-term taste (personality) and short-term preference (emotion). Specifically, it takes full advantage of personality-oriented user features, emotion-oriented user features and music features of multi-faceted attributes. Hierarchical attention is employed to distinguish the important factors when incorporating the latent representations of users' personality and emotion. Extensive experiments on a large real-world dataset of 171,254 users demonstrate the effectiveness of our PEIA model which achieves an NDCG of 0.5369, outperforming the state-of-the-art methods. We also perform detailed parameter analysis and feature contribution analysis, which further verify our scheme and demonstrate the significance of co-modeling of user personality and emotion in music recommendation.
Tiancheng Shen, Jia Jia 0001, Yan Li 0068, Yihui Ma, Yaohua Bu, Hanjie Wang, Tat-Seng Chua, Wendy Hall 0001
AAAI3
2020 Find Objects and Focus on Highlights: Mining Object Semantics for Video Highlight Detection via Graph Neural Networks
abstract
With the increasing prevalence of portable computing devices, browsing unedited videos is time-consuming and tedious. Video highlight detection has the potential to significantly ease this situation, which discoveries moments of user's major or special interest in a video. Existing methods suffer from two problems. Firstly, most existing approaches only focus on learning holistic visual representations of videos but ignore object semantics for inferring video highlights. Secondly, current state-of-the-art approaches often adopt the pairwise ranking-based strategy, which cannot enjoy the global information to infer highlights. Therefore, we propose a novel video highlight framework, named VH-GNN, to construct an object-aware graph and model the relationships between objects from a global view. To reduce computational cost, we decompose the whole graph into two types of graphs: a spatial graph to capture the complex interactions of object within each frame, and a temporal graph to obtain object-aware representation of each frame and capture the global information. In addition, we optimize the framework via a proposed multi-stage loss, where the first stage aims to determine the highlight-probability and the second stage leverage the relationships between frames and focus on hard examples from the former stage. Extensive experiments on two standard datasets strongly evidence that VH-GNN obtains significant performance compared with state-of-the-arts.
Junyu Gao 0002, Xiaoshan Yang, Yan Li 0068, Changsheng Xu
AAAI5
2020 Multi-Modality Cross Attention Network for Image and Sentence Matching
abstract
The key of image and sentence matching is to accurately measure the visual-semantic similarity between an image and a sentence. However, most existing methods make use of only the intra-modality relationship within each modality or the inter-modality relationship between image regions and sentence words for the cross-modal matching task. Different from them, in this work, we propose a novel MultiModality Cross Attention (MMCA) Network for image and sentence matching by jointly modeling the intra-modality and inter-modality relationships of image regions and sentence words in a unified deep model. In the proposed MMCA, we design a novel cross-attention mechanism, which is able to exploit not only the intra-modality relationship within each modality, but also the inter-modality relationship between image regions and sentence words to complement and enhance each other for image and sentence matching. Extensive experimental results on two standard benchmarks including Flickr30K and MS-COCO demonstrate that the proposed model performs favorably against state-of-the-art image and sentence matching methods.
Tianzhu Zhang 0001, Yan Li 0068, Yongdong Zhang 0001, Feng Wu 0001
CVPR3
2020 Bilinear Graph Neural Network with Neighbor Interactions
abstract
Graph Neural Network (GNN) is a powerful model to learn representations and make predictions on graph data. Existing efforts on GNN have largely defined the graph convolution as a weighted sum of the features of the connected nodes to form the representation of the target node. Nevertheless, the operation of weighted sum assumes the neighbor nodes are independent of each other, and ignores the possible interactions between them. When such interactions exist, such as the co-occurrence of two neighbor nodes is a strong signal of the target node's characteristics, existing GNN models may fail to capture the signal. In this work, we argue the importance of modeling the interactions between neighbor nodes in GNN. We propose a new graph convolution operator, which augments the weighted sum with pairwise interactions of the representations of neighbor nodes. We term this framework as Bilinear Graph Neural Network (BGNN), which improves GNN representation ability with bilinear interactions between neighbor nodes. In particular, we specify two BGNN models named BGCN and BGAT, based on the well-known GCN and GAT, respectively. Empirical results on three public benchmarks of semi-supervised node classification verify the effectiveness of BGNN --- BGCN (BGAT) outperforms GCN (GAT) by 1.6% (1.5%) in classification accuracy. Codes are available at: https://github.com/zhuhm1996/bgnn.
Hongmin Zhu, Fuli Feng, Xiangnan He 0001, Xiang Wang 0010, Yan Li 0068, Kai Zheng 0001, Yongdong Zhang 0001
IJCAI5
2020 Enhancing Music Recommendation with Social Media Content: an Attentive Multimodal Autoencoder Approach
abstract
Music recommendation methods predict users' music preference primarily based on historical ratings. Meanwhile, manifold personal factors of users are also important for the problem, and research efforts have been made to improve the recommendation performance with auxiliary user information. As an important indicator of users' personal traits and states, the numerous social media content (e.g., texts, images and short videos), however, is still hardly exploited. In this work, we systematically study the utilization of multimodal social media content for music recommendation. We define groups of both targeted handcrafted features and generic deep features for each modality, and further propose an Attentive Multimodal Autoencoder approach (AMAE) to learn cross-modal latent representations from the extracted features. Attention mechanism is also employed to integrate users' global and contextual music preference with alterable weights. Experiments demonstrate remarkable improvement of recommendation performance (+2.40% in Hit Ratio and +3.30% in NDCG), manifesting the effectiveness of our AMAE approach, as well as the significance of incorporating social media content data in music recommendation.
Tiancheng Shen, Jia Jia 0001, Yan Li 0068, Hanjie Wang
IJCNN3
2020 CRNet: A Center-aware Representation for Detecting Text of Arbitrary Shapes
abstract
Existing scene text detection methods achieve state-of-the-art performance by designing elaborate anchors or complex post-processing. Nonetheless, most methods still face the dilemma of detecting adjacent texts as one instance and long text with large character spacing as multiple fragments. To tackle these problems, we propose an anchor-free scene text detector leveraging Center-aware Representation to achieve accurate arbitrary-shaped scene text detection namely CRNet. Firstly, we propose a center-aware location algorithm to explicitly learn center regions and center points of text instances, which is able to separate adjacent text instances effectively. Then, a multi-scale context extraction module capable of extracting local context, long-range dependencies and global context adaptively is designed to effectively perceive long text with large character spacing. Finally, a low-level features enhancement block is introduced to enhance the geometric information of text. Extensive experiments conducted on several benchmarks including SCUT-CTW1500, Total-Text, ICDAR2015, ICDAR2017 MLT, and MSRA-TD500 demonstrate the effectiveness of our method. Specifically, without any anchor and complicated post-processing, our CRNet achieves 84.2% and 85.1% on CTW1500 and MSRA-TD500 in F-measure, outperforming all state-of-the-art anchor-based and anchor-free methods.
Yu Zhou 0016, Hongtao Xie 0001, Shancheng Fang, Yan Li 0068, Yongdong Zhang 0001
ACM Multimedia4
2020 LightGCN: Simplifying and Powering Graph Convolution Network for Recommendation
abstract
Graph Convolution Network (GCN) has become new state-of-the-art for collaborative filtering. Nevertheless, the reasons of its effectiveness for recommendation are not well understood. Existing work that adapts GCN to recommendation lacks thorough ablation analyses on GCN, which is originally designed for graph classification tasks and equipped with many neural network operations. However, we empirically find that the two most common designs in GCNs -- feature transformation and nonlinear activation -- contribute little to the performance of collaborative filtering. Even worse, including them adds to the difficulty of training and degrades recommendation performance.
Xiangnan He 0001, Kuan Deng, Xiang Wang 0010, Yan Li 0068, Yongdong Zhang 0001, Meng Wang 0001
SIGIR4
2020 How to Retrain Recommender System?: A Sequential Meta-Learning Method
abstract
Practical recommender systems need be periodically retrained to refresh the model with new interaction data. To pursue high model fidelity, it is usually desirable to retrain the model on both historical and new data, since it can account for both long-term and short-term user preference. However, a full model retraining could be very time-consuming and memory-costly, especially when the scale of historical data is large. In this work, we study the model retraining mechanism for recommender systems, a topic of high practical values but has been relatively little explored in the research community.
Yang Zhang 0072, Fuli Feng, Chenxu Wang 0010, Xiangnan He 0001, Meng Wang 0001, Yan Li 0068, Yongdong Zhang 0001
SIGIR6
2020 A Joint Neural Model for User Behavior Prediction on Social Networking Platforms
abstract
Social networking services provide platforms for users to perform two kinds of behaviors: consumption behavior (e.g., recommending items of interest) and social link behavior (e.g., recommending potential social links). Accurately modeling and predicting users’ two kinds of behaviors are two core tasks in these platforms with various applications. Recently, with the advance of neural networks, many neural-based models have been designed to predict a single users’ behavior, i.e., social link behavior or consumption behavior. Compared to the classical shallow models, these neural-based models show better performance to drive a user’s behavior by modeling the complex patterns. However, there are few works exploiting whether it is possible to design a neural-based model to jointly predict users’ two kinds of behaviors to further enhance the prediction performance. In fact, social scientists have already shown that users’ two kinds of behaviors are not isolated; people trend to the consumption recommendation of friends on social platforms and would like to make new friends with like-minded users. While some previous works jointly model users’ two kinds of behaviors with shallow models, we argue that the correlation between users’ two kinds of behaviors are complex, which could not be well-designed with shallow linear models. To this end, in this article, we propose a neural joint behavior prediction model named Neural Joint Behavior Prediction Model (NJBP) to mutually enhance the prediction performance of these two tasks on social networking platforms. Specifically, there are two key characteristics of our proposed model: First, to model the correlation of users’ two kinds of behaviors, we design a fusion layer in the neural network to model the positive correlation of users’ two kinds of behaviors. Second, as the observed links in the social network are often very sparse, we design a new link-based loss function that could preserve the social network topology. After that, we design a joint optimization function to allow the two behaviors modeling tasks to be trained to mutually enhance each other. Finally, extensive experimental results on two real-world datasets show that our proposed method is on average 7.14% better than the best baseline on social link behavior while 6.21% on consumption behavior prediction. Compared with the pair-wise loss function on two datasets, our proposed link-based loss function improves at least 4.69% on the social link behavior prediction and 4.72% on the consumption behavior prediction.
Junwei Li 0011, Le Wu 0001, Richang Hong, Kun Zhang 0015, Yong Ge 0001, Yan Li 0068
ACM Trans. Intell. Syst. Technol.6
2019 Explainable Interaction-driven User Modeling over Knowledge Graph for Sequential Recommendation
abstract
Compared with the traditional recommendation system, sequential recommendation holds the ability of capturing the evolution of users' dynamic interests. Many previous studies in sequential recommendation focus on the accuracy of predicting the next item that a user might interact with, while generally ignore providing explanations why the item is recommended to the user. Appropriate explanations are critical to help users adopt the recommended item, and thus improve the transparency and trustworthiness of the recommendation system. In this paper, we propose a novel Explainable Interaction-driven User Modeling (EIUM) algorithm to exploit Knowledge Graph (KG) for constructing an effective and explainable sequential recommender. Qualified semantic paths between specific user-item pair are extracted from KG. Encoding those semantic paths and learning the importance scores for each path provides the path-wise explanation for the recommendation system. Different from traditional item- level sequential modeling methods, we capture the interaction-level user dynamic preferences by modeling the sequential interactions. It is a high- level representation which contains auxiliary semantic information from KG. Furthermore, we adopt a joint learning manner for better representation learning by employing multi-modal fusion, which benefits from the structural constraints in KG and involves three kinds of modalities. Extensive experiments on the large-scale dataset show the better performance of our approach in making sequential recommendations in terms of both accuracy and explainability.
Xiaowen Huang 0001, Quan Fang, Shengsheng Qian, Jitao Sang 0001, Yan Li 0068, Changsheng Xu
ACM Multimedia5
2019 High throughput hardware architecture for accurate semi-global matching
Yan Li 0068, Zhiwei Li 0006, Song Chen 0001
Integr.1
2019 Convolutional Attention Networks for Scene Text Recognition
abstract
In this article, we present Convoluitional Attention Networks (CAN) for unconstrained scene text recognition. Recent dominant approaches for scene text recognition are mainly based on Convolutional Neural Networks (CNN) and Recurrent Neural Networks (RNN), where the CNN encodes images and the RNN generates character sequences. Our CAN is different from these methods; our CAN is completely built on CNN and includes an attention mechanism. The distinctive characteristics of our method include (i) CAN follows encoder-decoder architecture, in which the encoder is a deep two-dimensional CNN and the decoder is a one-dimensional CNN; (ii) the attention mechanism is applied in every convolutional layer of the decoder, and we propose a novel spatial attention method using average pooling; and (iii) position embeddings are equipped in both a spatial encoder and a sequence decoder to give our networks a sense of location. We conduct experiments on standard datasets for scene text recognition, including Street View Text , IIIT5K, and ICDAR datasets. The experimental results validate the effectiveness of different components and show that our convolutional-based method achieves state-of-the-art or competitive performance over prior works, even without the use of RNN.
Hongtao Xie 0001, Shancheng Fang, Zhengjun Zha, Yating Yang, Yan Li 0068, Yongdong Zhang 0001
ACM Trans. Multim. Comput. Commun. Appl.5
2018 Temporal Hierarchical Attention at Category- and Item-Level for Micro-Video Click-Through Prediction
abstract
Micro-video sharing gains great popularity in recent years, which calls for effective recommendation algorithm to help user find their interested micro-videos. Compared with traditional online (e.g. YouTube) videos, micro-videos contributed by grass-root users and taken by smartphones are much shorter (tens of seconds) and more short of tags or descriptive text, making the recommendation of micro-videos a challenging task. In this paper, we investigate how to model user's historical behaviors so as to predict the user's click-through of micro-videos. Inspired by the recent deep network-based methods, we propose a Temporal Hierarchical Attention at Category- and Item-Level (THACIL) network for user behavior modeling. First, we use temporal windows to capture the short-term dynamics of user interests; Second, we leverage a category-level attention mechanism to characterize user's diverse interests, as well as an item-level attention mechanism for fine-grained profiling of user interests; Third, we adopt forward multi-head self-attention to capture the long-term correlation within user behaviors. Our proposed THACIL network was tested on MicroVideo-1.7M, a new dataset of 1.7 million micro-videos, coming from real data of a micro-video sharing service in China. Experimental results demonstrate the effectiveness of the proposed method in comparison with the state-of-the-art solutions.
Xusong Chen, Dong Liu 0002, Zhengjun Zha, Wengang Zhou 0001, Zhiwei Xiong, Yan Li 0068
ACM Multimedia6
2017 High throughput hardware architecture for accurate semi-global matching
abstract
As the most important step of a stereo vision system, stereo matching, which finds the correspondences in stereo image pairs, requires high-quality real-time depth computation. In this paper, a high accuracy and high throughput full-pipeline hardware architecture with disparity and row parallelism is proposed. In the semi-global aggregation stage, to improve the accuracy in discontinuous regions, adaptive weighted path costs are adopted, and, five aggregation paths are used without consuming external memory resources. The proposed hardware architecture is implemented on a Stratix V FPGA, which results in a throughput of 1280×960/197fps with 64 disparity levels at 156MHz.
Yan Li 0068, Zhiwei Li 0006, Song Chen 0001
ASP-DAC1
2016 Real-Time Hardware Stereo Matching Using Guided Image Filter
abstract
Stereo matching is a key step in stereo vision systems that require high accurate depth information and real-time processing of high definition image streams. This work presents a high-accuracy hardware implementation for the stereo matching based on the guided image filter, which is an edge-preserving filter and simplifies the adaptive support window algorithm. The coefficients in the guided image filter are calculated by the proposed mean filter tree structure, which saves hardware resources by sharing large amounts of additions among filter operations. The reference image is enhanced using Laplacian Filter, which improves the accuracy for the discontinuous disparity regions. Moreover, an 8×8 matching window and customized ping-pong caches are used to improve the whole throughputs. The proposed hardware architecture is implemented on a Cyclone IV FPGA resulting in a throughput of 1080p resolution images at 80fps with high accuracy of disparity.
Yan Li 0068, Song Chen 0001
ACM Great Lakes Symposium on VLSI2