Wenlong Xie

dblp:169/3287 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
0since 2021 · last 2020
0000-0001-8790-3007ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-authorArtificial intelligence and machine learning · 4 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Video understanding and tracking · 52% Knowledge representation and reasoning · 40% Image recognition and object detection · 8%
Computer graphics and multimedia
2 papers
Multimedia analysis and retrieval · 100%
Databases, data mining, and information retrieval
1 paper
Web and social media mining · 100%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking › event recognition
complex event detection
0.412019
Discovering Latent Discriminative Patterns for Multi-Mode Event Representation · IEEE Trans. Multim. 2019
Knowledge, reasoning and agents › Knowledge representation and reasoning › reasoning about action and change
event representation
0.412019
Discovering Latent Discriminative Patterns for Multi-Mode Event Representation · IEEE Trans. Multim. 2019
Web and social media mining › social media analysis
social image analysis
0.212016
Predicting Personalized Emotion Perceptions of Social Images · ACM Multimedia 2016
Multimedia analysis and retrieval › affective computing
affective image analysis
0.212016
Predicting Personalized Emotion Perceptions of Social Images · ACM Multimedia 2016
Multimedia analysis and retrieval
video recommendation
0.212015
"Clustering of Dancelets": Towards Video Recommendation Based on Dance Styles · ACM Multimedia 2015
Computer vision › Video understanding and tracking
video representation learning
0.112019
Discovering Latent Discriminative Patterns for Multi-Mode Event Representation · IEEE Trans. Multim. 2019
Computer vision › Image recognition and object detection
visual emotion recognition
0.112016
User-Centric Affective Computing of Image Emotion Perceptions · AAAI 2016

Methods — techniques the papers use, named apart from their topics

multi-task learning · 0.5hypergraph learning · 0.5latent discriminative pattern discovery · 0.4deep-visual-word encoding · 0.4rolling multi-task hypergraph learning · 0.2random forest · 0.2normalized cut · 0.2mid-level action representation · 0.2
YearPublicationVenuePosition
2020 Actionness-pooled Deep-convolutional Descriptor for fine-grained action recognition
Tingting Han 0003, Hongxun Yao, Xiaoshuai Sun, Wenlong Xie, Sicheng Zhao, Wei Yu 0004
Neurocomputing4
2020 TVENet: Temporal variance embedding network for fine-grained action representation
Tingting Han 0003, Hongxun Yao, Wenlong Xie, Xiaoshuai Sun, Sicheng Zhao, Jun Yu 0002
Pattern Recognit.3
2019 Discovering Latent Discriminative Patterns for Multi-Mode Event Representation
abstract
Representation of videos is essential since it conveys an understanding of video content and enables many higher level tasks to be tackled efficiently. However, it is challenging to propose a rational representation for complex event videos, as most video information is either noisy or redundant. In this paper, we propose a compact event representation method that can concisely describe the inner modes of events. We deem that an optimal event representation scheme should reflect the long-term and high-level visual semantics (visual topics) of events, so different from previous frame-level video semantics representation methods and concept-based video representation methods, we investigate the problem from the perspective of segment-level video representations. We then present three appealing properties of segment-level visual semantics. Based on the observation, we propose different algorithms that rely on a novel deep-visual-word-based video encoding method to discover latent discriminative patterns of events. Finally, our multi-mode event representation is obtained by concatenating the discovered patterns as inner modes. We adopt our event representation for representative event parts mining, which can highlight the visual topics of events and remarkably prune the raw videos. We validate our event representation method based on complex event detection task. Experimental results on two standard benchmarking datasets, MED11 and CCV Dataset, show that the proposed method can significantly outperform the state-of-the-art approaches.
Wenlong Xie, Hongxun Yao, Xiaoshuai Sun, Tingting Han 0003, Sicheng Zhao, Tat-Seng Chua
IEEE Trans. Multim.1
2018 Add: Actionness-Pooled Deep-Convolutional Descriptor
abstract
Recognition of general actions has achieved great breakthroughs in recent years. However, in real-world applications, finer-grained action classification is often needed. The major challenge is that fine-grained actions usually share high similarities in both appearance and motion pattern, making it difficult to distinguish them with existing general action representation. To solve this problem, we introduce visual attention mechanism into the proposed descriptor, termed as Actionness-pooled Deep-convolutional Descriptor (ADD). Instead of pooling features uniformly from the entire video, we aggregate features in sub-regions that are more likely to contain actions according to actionness maps, which endow ADD with the capability of capturing the subtle differences between fine-grained actions. We conduct experiments on HIT Dances dataset, one of the few existing datasets for fine-grained action analysis. Quantitative results have demonstrated that ADD remarkably outperforms traditional two-stream representation. Extensive experiments on two general action benchmarks, JHMDB and UCF101, have additionally proved that combining ADD with end-to-end ConvNet can further boost the recognition performance.
Tingting Han 0003, Hongxun Yao, Xiaoshuai Sun, Wenlong Xie, Yanhao Zhang 0001
ICME4
2018 Event patches: Mining effective parts for event detection and understanding
Wenlong Xie, Hongxun Yao, Sicheng Zhao, Xiaoshuai Sun, Tingting Han 0003
Signal Process.1
2017 Actor identification via mining representative actions
Wenlong Xie, Hongxun Yao, Xiaoshuai Sun, Sicheng Zhao, Wei Yu 0004, Shengping Zhang
Neurocomputing1
2016 User-Centric Affective Computing of Image Emotion Perceptions
abstract
We propose to predict the personalized emotion perceptions of images for each viewer. Different factors that may influence emotion perceptions, including visual content, social context, temporal evolution, and location influence are jointly investigated via the presented rolling multi-task hypergraph learning. For evaluation, we set up a large scale image emotion dataset from Flickr, named Image-Emotion-Social-Net, with over 1 million images and about 8,000 users. Experiments conducted on this dataset demonstrate the superiority of the proposed method, as compared to state-of-the-art.
Sicheng Zhao, Hongxun Yao, Wenlong Xie, Xiaolei Jiang
AAAI3
2016 Mining representative actions for actor identification
abstract
Previous works on actor identification mainly focused on static features based on face identification and costume detection, without considering the abundant dynamic information contained in videos. In this paper, we propose a novel method to mine representative actions of each actor, and show the remarkable power of such actions for actor identification task. Videos are firstly divided into shots and represented by BoW based on spatial-temporal features. Then we integrate the prototype theory with SVM to rank the shots and obtain the representative actions. Our method for actor identification combines representative actions with actors' appearance. We validate the method on episodes of the TV series "The Big Bang Theory". The experimental results show that the representative actions are consistent with human judgements and can greatly improve the matching performance as complementary to existing handcrafted static features for actor identification.
Wenlong Xie, Hongxun Yao, Xiaoshuai Sun, Sicheng Zhao, Tingting Han 0003
ICASSP1
2016 Predicting Personalized Emotion Perceptions of Social Images
abstract
Images can convey rich semantics and induce various emotions to viewers. Most existing works on affective image analysis focused on predicting the dominant emotions for the majority of viewers. However, such dominant emotion is often insufficient in real-world applications, as the emotions that are induced by an image are highly subjective and different with respect to different viewers. In this paper, we propose to predict the personalized emotion perceptions of images for each individual viewer. Different types of factors that may affect personalized image emotion perceptions, including visual content, social context, temporal evolution, and location influence, are jointly investigated. Rolling multi-task hypergraph learning is presented to consistently combine these factors and a learning algorithm is designed for automatic optimization. For evaluation, we set up a large scale image emotion dataset from Flickr, named Image-Emotion-Social-Net, on both dimensional and categorical emotion representations with over 1 million images and about 8,000 users. Experiments conducted on this dataset demonstrate that the proposed method can achieve significant performance gains on personalized emotion classification, as compared to several state-of-the-art approaches.
Sicheng Zhao, Hongxun Yao, Yue Gao 0002, Rongrong Ji, Wenlong Xie, Xiaolei Jiang, Tat-Seng Chua
ACM Multimedia5
2015 "Clustering of Dancelets": Towards Video Recommendation Based on Dance Styles
abstract
Dance is a special and important type of action, composed of abundant and various action elements. However, the recommendation of dance videos on the web are still not well studied. It is hard to realize it in the way of traditional methods using associated texts or static features of video content. In this paper, we study the problem focusing on extraction and representation of action information in dances. We propose to recommend dance videos based on the automatically discovered ``Dance Styles'', which play a significant role in characterizing different types of dances. To bridge the semantic gap of video content and mid-level concept, style, we take advantage of a mid-level action representation method, and extract representative patches as ``Dancelets'', a sort of intermediation between videos and the concepts. Furthermore, we propose to employ Motion Boundaries as saliency priors and sparsely extract patches containing more representative information to generate a set of dancelet candidates. Dancelets are then discovered by Normalized-cut method, which is superior in grouping visually similar patterns into the same clusters. For the fast and effective recommendation, a random forest-based index is built, and the ranking results are derived according to the matching results in all the leaf notes. Extensive experiments validated on the web dance videos demonstrate the effectiveness of the proposed methods for dance style discovery and video recommendation based on styles.
Tingting Han 0003, Hongxun Yao, Xiaoshuai Sun, Yanhao Zhang 0001, Sicheng Zhao, Xiusheng Lu, Yinghao Huang, Wenlong Xie
ACM Multimedia8