EDBT 2026 Demo / reviewers in the wild / expert
Wenlong Xie
dblp:169/3287
· DBLP profile ↗
10ranked-venue papers
4as first author
0since 2021 · last 2020
0000-0001-8790-3007ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-authorArtificial intelligence and machine learning · 4 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Video understanding and tracking · 52% Knowledge representation and reasoning · 40% Image recognition and object detection · 8% | |
| Computer graphics and multimedia
2 papers |
Multimedia analysis and retrieval · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Web and social media mining · 100% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Video understanding and tracking › event recognition
complex event detection |
0.4 | 1 | 2019 | Discovering Latent Discriminative Patterns for Multi-Mode Event Representation · IEEE Trans. Multim. 2019 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › reasoning about action and change
event representation |
0.4 | 1 | 2019 | Discovering Latent Discriminative Patterns for Multi-Mode Event Representation · IEEE Trans. Multim. 2019 |
Web and social media mining › social media analysis
social image analysis |
0.2 | 1 | 2016 | Predicting Personalized Emotion Perceptions of Social Images · ACM Multimedia 2016 |
Multimedia analysis and retrieval › affective computing
affective image analysis |
0.2 | 1 | 2016 | Predicting Personalized Emotion Perceptions of Social Images · ACM Multimedia 2016 |
Multimedia analysis and retrieval
video recommendation |
0.2 | 1 | 2015 | "Clustering of Dancelets": Towards Video Recommendation Based on Dance Styles · ACM Multimedia 2015 |
Computer vision › Video understanding and tracking
video representation learning |
0.1 | 1 | 2019 | Discovering Latent Discriminative Patterns for Multi-Mode Event Representation · IEEE Trans. Multim. 2019 |
Computer vision › Image recognition and object detection
visual emotion recognition |
0.1 | 1 | 2016 | User-Centric Affective Computing of Image Emotion Perceptions · AAAI 2016 |
Methods — techniques the papers use, named apart from their topics
multi-task learning · 0.5hypergraph learning · 0.5latent discriminative pattern discovery · 0.4deep-visual-word encoding · 0.4rolling multi-task hypergraph learning · 0.2random forest · 0.2normalized cut · 0.2mid-level action representation · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Actionness-pooled Deep-convolutional Descriptor for fine-grained action recognition
Tingting Han 0003, Hongxun Yao, Xiaoshuai Sun, Wenlong Xie, Sicheng Zhao, Wei Yu 0004 |
Neurocomputing | 4 |
| 2020 | TVENet: Temporal variance embedding network for fine-grained action representation
Tingting Han 0003, Hongxun Yao, Wenlong Xie, Xiaoshuai Sun, Sicheng Zhao, Jun Yu 0002 |
Pattern Recognit. | 3 |
| 2019 | Discovering Latent Discriminative Patterns for Multi-Mode Event RepresentationabstractRepresentation of videos is essential since it conveys an understanding of video content and enables many higher level tasks to be tackled efficiently. However, it is challenging to propose a rational representation for complex event videos, as most video information is either noisy or redundant. In this paper, we propose a compact event representation method that can concisely describe the inner modes of events. We deem that an optimal event representation scheme should reflect the long-term and high-level visual semantics (visual topics) of events, so different from previous frame-level video semantics representation methods and concept-based video representation methods, we investigate the problem from the perspective of segment-level video representations. We then present three appealing properties of segment-level visual semantics. Based on the observation, we propose different algorithms that rely on a novel deep-visual-word-based video encoding method to discover latent discriminative patterns of events. Finally, our multi-mode event representation is obtained by concatenating the discovered patterns as inner modes. We adopt our event representation for representative event parts mining, which can highlight the visual topics of events and remarkably prune the raw videos. We validate our event representation method based on complex event detection task. Experimental results on two standard benchmarking datasets, MED11 and CCV Dataset, show that the proposed method can significantly outperform the state-of-the-art approaches. Wenlong Xie, Hongxun Yao, Xiaoshuai Sun, Tingting Han 0003, Sicheng Zhao, Tat-Seng Chua |
IEEE Trans. Multim. | 1 |
| 2018 | Add: Actionness-Pooled Deep-Convolutional DescriptorabstractRecognition of general actions has achieved great breakthroughs in recent years. However, in real-world applications, finer-grained action classification is often needed. The major challenge is that fine-grained actions usually share high similarities in both appearance and motion pattern, making it difficult to distinguish them with existing general action representation. To solve this problem, we introduce visual attention mechanism into the proposed descriptor, termed as Actionness-pooled Deep-convolutional Descriptor (ADD). Instead of pooling features uniformly from the entire video, we aggregate features in sub-regions that are more likely to contain actions according to actionness maps, which endow ADD with the capability of capturing the subtle differences between fine-grained actions. We conduct experiments on HIT Dances dataset, one of the few existing datasets for fine-grained action analysis. Quantitative results have demonstrated that ADD remarkably outperforms traditional two-stream representation. Extensive experiments on two general action benchmarks, JHMDB and UCF101, have additionally proved that combining ADD with end-to-end ConvNet can further boost the recognition performance. Tingting Han 0003, Hongxun Yao, Xiaoshuai Sun, Wenlong Xie, Yanhao Zhang 0001 |
ICME | 4 |
| 2018 | Event patches: Mining effective parts for event detection and understanding
Wenlong Xie, Hongxun Yao, Sicheng Zhao, Xiaoshuai Sun, Tingting Han 0003 |
Signal Process. | 1 |
| 2017 | Actor identification via mining representative actions
Wenlong Xie, Hongxun Yao, Xiaoshuai Sun, Sicheng Zhao, Wei Yu 0004, Shengping Zhang |
Neurocomputing | 1 |
| 2016 | User-Centric Affective Computing of Image Emotion PerceptionsabstractWe propose to predict the personalized emotion perceptions of images for each viewer. Different factors that may influence emotion perceptions, including visual content, social context, temporal evolution, and location influence are jointly investigated via the presented rolling multi-task hypergraph learning. For evaluation, we set up a large scale image emotion dataset from Flickr, named Image-Emotion-Social-Net, with over 1 million images and about 8,000 users. Experiments conducted on this dataset demonstrate the superiority of the proposed method, as compared to state-of-the-art. Sicheng Zhao, Hongxun Yao, Wenlong Xie, Xiaolei Jiang |
AAAI | 3 |
| 2016 | Mining representative actions for actor identificationabstractPrevious works on actor identification mainly focused on static features based on face identification and costume detection, without considering the abundant dynamic information contained in videos. In this paper, we propose a novel method to mine representative actions of each actor, and show the remarkable power of such actions for actor identification task. Videos are firstly divided into shots and represented by BoW based on spatial-temporal features. Then we integrate the prototype theory with SVM to rank the shots and obtain the representative actions. Our method for actor identification combines representative actions with actors' appearance. We validate the method on episodes of the TV series "The Big Bang Theory". The experimental results show that the representative actions are consistent with human judgements and can greatly improve the matching performance as complementary to existing handcrafted static features for actor identification. Wenlong Xie, Hongxun Yao, Xiaoshuai Sun, Sicheng Zhao, Tingting Han 0003 |
ICASSP | 1 |
| 2016 | Predicting Personalized Emotion Perceptions of Social ImagesabstractImages can convey rich semantics and induce various emotions to viewers. Most existing works on affective image analysis focused on predicting the dominant emotions for the majority of viewers. However, such dominant emotion is often insufficient in real-world applications, as the emotions that are induced by an image are highly subjective and different with respect to different viewers. In this paper, we propose to predict the personalized emotion perceptions of images for each individual viewer. Different types of factors that may affect personalized image emotion perceptions, including visual content, social context, temporal evolution, and location influence, are jointly investigated. Rolling multi-task hypergraph learning is presented to consistently combine these factors and a learning algorithm is designed for automatic optimization. For evaluation, we set up a large scale image emotion dataset from Flickr, named Image-Emotion-Social-Net, on both dimensional and categorical emotion representations with over 1 million images and about 8,000 users. Experiments conducted on this dataset demonstrate that the proposed method can achieve significant performance gains on personalized emotion classification, as compared to several state-of-the-art approaches. Sicheng Zhao, Hongxun Yao, Yue Gao 0002, Rongrong Ji, Wenlong Xie, Xiaolei Jiang, Tat-Seng Chua |
ACM Multimedia | 5 |
| 2015 | "Clustering of Dancelets": Towards Video Recommendation Based on Dance StylesabstractDance is a special and important type of action, composed of abundant and various action elements. However, the recommendation of dance videos on the web are still not well studied. It is hard to realize it in the way of traditional methods using associated texts or static features of video content. In this paper, we study the problem focusing on extraction and representation of action information in dances. We propose to recommend dance videos based on the automatically discovered ``Dance Styles'', which play a significant role in characterizing different types of dances. To bridge the semantic gap of video content and mid-level concept, style, we take advantage of a mid-level action representation method, and extract representative patches as ``Dancelets'', a sort of intermediation between videos and the concepts. Furthermore, we propose to employ Motion Boundaries as saliency priors and sparsely extract patches containing more representative information to generate a set of dancelet candidates. Dancelets are then discovered by Normalized-cut method, which is superior in grouping visually similar patterns into the same clusters. For the fast and effective recommendation, a random forest-based index is built, and the ranking results are derived according to the matching results in all the leaf notes. Extensive experiments validated on the web dance videos demonstrate the effectiveness of the proposed methods for dance style discovery and video recommendation based on styles. Tingting Han 0003, Hongxun Yao, Xiaoshuai Sun, Yanhao Zhang 0001, Sicheng Zhao, Xiusheng Lu, Yinghao Huang, Wenlong Xie |
ACM Multimedia | 8 |