VLDB 2026 Research / reviewers in the wild / expert
Jia Chen 0001
dblp:99/6879-1
· DBLP profile ↗
25ranked-venue papers
9as first author
0since 2021 · last 2020
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 5 first-authorDatabases, data management, data science and information retrieval · 6 · 2 first-authorArtificial intelligence and machine learning · 3 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Vision and language · 73% Deep learning architectures and training · 12% Image recognition and object detection · 7% | |
| Computer graphics and multimedia
8 papers |
Multimedia analysis and retrieval · 89% Image and video processing · 11% | |
| Databases, data mining, and information retrieval
7 papers |
Information retrieval · 46% Recommender systems · 44% Data mining · 5% |
Topics — the 21 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
video captioning |
1.6 | 5 | 2020 | Better Captioning With Sequence-Level Exploration · CVPR 2020 Generating Video Descriptions With Latent Topic Guidance · IEEE Trans. Multim. 2019 Knowing Yourself: Improving Video Caption via In-depth Recap · ACM Multimedia 2017 |
Computer vision › Vision and language
image captioning |
0.4 | 1 | 2020 | Better Captioning With Sequence-Level Exploration · CVPR 2020 |
Machine learning › Deep learning architectures and training › sequence modeling
sequence-level training |
0.4 | 1 | 2020 | Better Captioning With Sequence-Level Exploration · CVPR 2020 |
Image and video processing › image reconstruction
event-based image reconstruction |
0.3 | 1 | 2017 | An Event Reconstruction Tool for Conflict Monitoring Using Social Media · AAAI 2017 |
Multimedia analysis and retrieval › image retrieval › instance-level image retrieval
landmark search |
0.3 | 2 | 2012 | Searching for diversified landmarks by photo · ACM Multimedia 2012 DLMSearch: diversified landmark search by photo · ACM Multimedia 2012 |
Multimedia analysis and retrieval › multimodal learning
multimodal topic modeling |
0.3 | 1 | 2017 | Video Captioning with Guidance of Multimodal Latent Topics · ACM Multimedia 2017 |
Multimedia analysis and retrieval › object tracking
person tracking |
0.3 | 1 | 2017 | An Event Reconstruction Tool for Conflict Monitoring Using Social Media · AAAI 2017 |
Multimedia analysis and retrieval
video analysis |
0.3 | 1 | 2017 | An Event Reconstruction Tool for Conflict Monitoring Using Social Media · AAAI 2017 |
Computer vision › Vision and language
multimodal fusion |
0.2 | 1 | 2016 | Describing Videos using Multi-modal Fusion · ACM Multimedia 2016 |
Information retrieval
multimedia analysis and retrieval |
0.2 | 1 | 2016 | History Rhyme: Searching Historic Events by Multimedia Knowledge · ACM Multimedia 2016 |
Recommender systems › e-commerce recommendation
product recommendation |
0.2 | 1 | 2016 | Boosting Recommendation in Unexplored Categories by User Price Preference · ACM Trans. Inf. Syst. 2016 |
Multimedia analysis and retrieval › cross-modal retrieval
image-text retrieval |
0.2 | 1 | 2016 | Semantic Image Profiling for Historic Events: Linking Images to Phrases · ACM Multimedia 2016 |
Computer vision › Face, body and person analysis
human pose estimation |
0.2 | 1 | 2014 | Unified Structured Learning for Simultaneous Human Pose Estimation and Garment Attribute Classification · IEEE Trans. Image Process. 2014 |
Recommender systems
cold-start recommendation |
0.2 | 1 | 2014 | Does product recommendation meet its waterloo in unexplored categories?: no, price comes to help · SIGIR 2014 |
Recommender systems
cross-domain recommendation |
0.2 | 1 | 2014 | Does product recommendation meet its waterloo in unexplored categories?: no, price comes to help · SIGIR 2014 |
Information retrieval
diversified retrieval |
0.2 | 2 | 2012 | DLMSearch: diversified landmark search by photo · ACM Multimedia 2012 Searching for diversified landmarks by photo · ACM Multimedia 2012 |
Multimedia analysis and retrieval › object recognition
landmark recognition |
0.2 | 1 | 2013 | Tell me what happened here in history · ACM Multimedia 2013 |
Performance modeling and evaluation › benchmarking
benchmark evaluation |
0.1 | 1 | 2017 | Knowing Yourself: Improving Video Caption via In-depth Recap · ACM Multimedia 2017 |
Natural language and speech › Language models and text generation › decoding
sequence decoding |
0.1 | 1 | 2016 | Describing Videos using Multi-modal Fusion · ACM Multimedia 2016 |
Data mining › predictive modeling › classification
noisy label learning |
0.1 | 1 | 2016 | Semantic Image Profiling for Historic Events: Linking Images to Phrases · ACM Multimedia 2016 |
Multimedia analysis and retrieval
visual summarization |
0.0 | 1 | 2012 | DLMSearch: diversified landmark search by photo · ACM Multimedia 2012 |
Methods — techniques the papers use, named apart from their topics
topic prediction · 0.6teacher-student learning · 0.6multi-task learning · 0.6ensemble · 0.6pseudo-label learning · 0.5instance-wise ranking loss · 0.5sequence-level exploration · 0.4reinforcement learning · 0.4temporal attention · 0.4multimodal ensemble · 0.4latent topic modeling · 0.4video synchronization · 0.3reranking · 0.3re-ranking · 0.3geolocalization · 0.33d reconstruction · 0.3utility function · 0.2regularization · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Better Captioning With Sequence-Level ExplorationabstractSequence-level learning objective has been widely used in captioning tasks to achieve the state-of-the-art performance for many models. In this objective, the model is trained by the reward on the quality of its generated captions (sequence-level). In this work, we show the limitation of the current sequence-level learning objective for captioning tasks from both theory and empirical result. In theory, we show that the current objective is equivalent to only optimizing the precision side of the caption set generated by the model and therefore overlooks the recall side. Empirical result shows that the model trained by this objective tends to get lower score on the recall side. We propose to add a sequence-level exploration term to the current objective to boost recall. It guides the model to explore more plausible captions in the training. In this way, the proposed objective takes both the precision and recall sides of generated captions into account. Experiments show the effectiveness of the proposed method on both video and image captioning datasets. Jia Chen 0001, Qin Jin |
CVPR | 1 |
| 2019 | Improving Captioning for Low-Resource Languages by Cycle ConsistencyabstractImproving the captioning performance on low-resource languages by leveraging English caption datasets has received increasing research interest in recent years. Existing works mainly fall into two categories: translation-based and alignment-based approaches. In this paper, we propose to combine the merits of both approaches in one unified architecture. Specifically, we use a pre-trained English caption model to generate high-quality English captions, and then take both the image and generated English captions to generate low-resource language captions. We improve the captioning performance by adding the cycle consistency constraint on the cycle of image regions, English words, and low-resource language words. Moreover, our architecture has a flexible design which enables it to benefit from large monolingual English caption datasets. Experimental results demonstrate that our approach outperforms the state-of-the-art methods on common evaluation metrics. The attention visualization also shows that the proposed approach really improves the fine-grained alignment between words and image regions. Yike Wu 0002, Shiwan Zhao, Jia Chen 0001, Ying Zhang 0015, Xiaojie Yuan, Zhong Su |
ICME | 3 |
| 2019 | Generating Video Descriptions With Latent Topic GuidanceabstractAutomatic video description generation (a.k.a video captioning) is one of the ultimate goals for video understanding. Despite the wide range of applications such as video indexing and retrieval etc., the video captioning task remains quite challenging due to the complexity and diversity of video content. First, open-domain videos cover a broad range of topics, which results in highly variable vocabularies and expression styles to describe the video contents. Second, videos naturally contain multiple modalities including image, motion, and acoustic media. The information provided by different modalities differs in different conditions. In this paper, we propose a novel topic-guided video captioning model to address the above-mentioned challenges in video captioning. Our model consists of two joint tasks, namely, latent topic generation and topic-guided caption generation. The topic generation task aims to automatically predict the latent topic of the video. Since there is no groundtruth topic information, we mine multimodal topics in an unsupervised fashion based on video contents and annotated captions, and then distill the topic distribution to a topic prediction model. In the topic-guided generation task, we employ the topic guidance for two purposes. The first is to narrow down the language complexity across topics, where we propose the topic-aware decoder to leverage the latent topics to induce topic-related language models. The decoder is also generic and can be integrated with a temporal attention mechanism. The second is to dynamically attend to important modalities by topics, where we propose a flexible topic-guided multimodal ensemble framework and use the topic gating network to determine the attention weights. The two tasks are correlated with each other, and they collaborate to generate more detailed and accurate video captions. Our extensive experiments on two public benchmark datasets MSR-VTT and Youtube2Text demonstrate the effectiveness of the proposed topic-guided video captioning system, which achieves state-of-the-art performance on both datasets. Shizhe Chen, Qin Jin, Jia Chen 0001, Alex Hauptmann 0001 |
IEEE Trans. Multim. | 3 |
| 2018 | Class-aware Self-Attention for Audio Event RecognitionabstractAudio event recognition (AER) has been an important research problem with a wide range of applications. However, it is very challenging to develop large scale audio event recognition models. On the one hand, usually there are only "weak" labeled audio training data available, which only contains labels of audio events without temporal boundaries. On the other hand, the distribution of audio events is generally long-tailed, with only a few positive samples for large amounts of audio events. These two issues make it hard to learn discriminative acoustic features to recognize audio events especially for long-tailed events. In this paper, we propose a novel class-aware self-attention mechanism with attention factor sharing to generate discriminative clip-level features for audio event recognition. Since a target audio event only occurs in part of an entire audio clip and its corresponding temporal interval varies, the proposed class-aware self-attention approach learns to highlight relevant temporal intervals and to suppress irrelevant noises at the same time. In order to learn attention patterns effectively for those long-tailed events, we combine both the domain knowledge and data driven strategies to share attention factors in the proposed attention mechanism, which transfers the common knowledge learned from other similar events to the rare events. The proposed attention mechanism is a pluggable component and can be trained end-to-end in the overall AER model. We evaluate our model on a large-scale audio event corpus "Audio Set" with both short-term and long-term acoustic features. The experimental results demonstrate the effectiveness of our model, which improves the overall audio event recognition performance with different acoustic features especially for events with low resources. Moreover, the experiments also show that our proposed model is able to learn new audio events with a few training examples effectively and efficiently without disturbing the previously learned audio events. Shizhe Chen, Jia Chen 0001, Qin Jin, Alex Hauptmann 0001 |
ICMR | 2 |
| 2017 | An Event Reconstruction Tool for Conflict Monitoring Using Social MediaabstractWhat happened during the Boston Marathon in 2013? Nowadays, at any major event, lots of people take videos and share them on social media. To fully understand exactly what happened in these major events, researchers and analysts often have to examine thousands of these videos manually. To reduce this manual effort, we present an investigative system that automatically synchronizes these videos to a global timeline and localizes them on a map. In addition to alignment in time and space, our system combines various functions for analysis, including gunshot detection, crowd size estimation, 3D reconstruction and person tracking. To our best knowledge, this is the first time a unified framework has been built for comprehensive event reconstruction for social media videos. Junwei Liang 0001, Desai Fan, Po-Yao Huang 0001, Jia Chen 0001, Lu Jiang 0004, Alex Hauptmann 0001 |
AAAI | 5 |
| 2017 | Synchronization for multi-perspective videos in the wildabstractIn the era of social media, a large number of user-generated videos are uploaded to the Internet every day, capturing events all over the world. Reconstructing the event truth based on information mined from these videos has been an emerging challenging task. Temporal alignment of videos “in the wild” which capture different moments at different positions with different perspectives is the critical step. In this paper, we propose a hierarchical approach to synchronize videos. Our system utilizes clustered audio-signatures to align video pairs. Global alignment for all videos is then achieved via forming alignable video groups with self-paced learning. Experiments on the Boston Marathon dataset show that the proposed method achieves excellent precision and robustness. Junwei Liang 0001, Po-Yao Huang 0001, Jia Chen 0001, Alex Hauptmann 0001 |
ICASSP | 3 |
| 2017 | Rewind to track: Parallelized apprenticeship learning with backward trackletsabstractData association, which could be categorized into offline approaches and the online counterparts, is a crucial part of a multi-object tracker in the tracking-by-detection framework. On the one hand, classical offline data association methods exploit all the video data and have high computation cost, which makes them unscalable to long-term offline video data. On the other hand, online approaches have much lower computation cost, but they suffer from ID-switches and tracklet drifting problem when directly applied to offline data as they are only aware of “past” observations. In this paper, we propose a mixed style tracker, which is not only as efficient as the online tracker but also aware of “future” observations in offline setting. We start from a Markov Decision Process (MDP) online tracker and design a parallelized apprenticeship learning algorithm to learn both the reward function and transition policy in MDP. By proposing a rewind to track strategy to generate backward tracklets, future detections in offline data are efficiently utilized to obtain a more stable similarity measurement for association. Experiment results show that our approach achieves the state-of-the-art performance on challenging datasets. Jiang Liu 0011, Jia Chen 0001, De Cheng, Chenqiang Gao, Alex Hauptmann 0001 |
ICME | 2 |
| 2017 | Generating Video Descriptions with Topic GuidanceabstractGenerating video descriptions in natural language (a.k.a. video captioning) is a more challenging task than image captioning as the videos are intrinsically more complicated than images in two aspects. First, videos cover a broader range of topics, such as news, music, sports and so on. Second, multiple topics could coexist in the same video. In this paper, we propose a novel caption model, topic-guided model (TGM), to generate topic-oriented descriptions for videos in the wild via exploiting topic information. In addition to predefined topics, i.e., category tags crawled from the web, we also mine topics in a data-driven way based on training captions by an unsupervised topic mining model. We show that data-driven topics reflect a better topic schema than the predefined topics. As for testing video topic prediction, we treat the topic mining model as teacher to train the student, the topic prediction model, by utilizing the full multi-modalities in the video especially the speech modality. We propose a series of caption models to exploit topic guidance, including implicitly using the topics as input features to generate words related to the topic and explicitly modifying the weights in the decoder with topics to function as an ensemble of topic-aware language decoders. Our comprehensive experimental results on the current largest video caption dataset MSR-VTT prove the effectiveness of our topic-guided model, which significantly surpasses the winning performance in the 2016 MSR video to language challenge. Shizhe Chen, Jia Chen 0001, Qin Jin |
ICMR | 2 |
| 2017 | Joint Saliency Estimation and Matching using Image Regions for Geo-Localization of Online VideoabstractIn this paper, we study automatic geo-localization of online event videos. Different from general image localization task through matching, the appearance of an environment during significant events varies greatly from its daily appearance, since there are usually crowds, decorations or even destruction when a major event happens. This introduces a major challenge: matching the event environment to the daily environment, e.g. as recorded by Google Street View. We observe that some regions in the image, as part of the environment, still preserve the daily appearance even though the whole image (environment) looks quite different. Based on this observation, we formulate the problem as joint saliency estimation and matching at the image region level, as opposed to the key point or whole-image level. As image-level labels of daily environment are easily generated with GPS information, we treat region based saliency estimation and matching as a weakly labeled learning problem over the training data. Our solution is to iteratively optimize saliency and the region-matching model. For saliency optimization, we derive a closed form solution, which has an intuitive explanation. For region matching model optimization, we use self-paced learning to learn from the pseudo labels generated by (sub-optimal) saliency values. We conduct extensive experiments on two challenging public datasets: Boston Marathon 2013 and Tokyo Time Machine. Experimental results show that our solution significantly improves over matching on whole images and the automatically learned saliency is a strong predictor of distinctive building areas. Freda Shi, Jia Chen 0001, Alex Hauptmann 0001 |
ICMR | 2 |
| 2017 | Video Captioning with Guidance of Multimodal Latent TopicsabstractThe topic diversity of open-domain videos leads to various vocabularies and linguistic expressions in describing video contents, and therefore, makes the video captioning task even more challenging. In this paper, we propose an unified caption framework, M&M TGM, which mines multimodal topics in unsupervised fashion from data and guides the caption decoder with these topics. Compared to pre-defined topics, the mined multimodal topics are more semantically and visually coherent and can reflect the topic distribution of videos better. We formulate the topic-aware caption generation as a multi-task learning problem, in which we add a parallel task, topic prediction, in addition to the caption task. For the topic prediction task, we use the mined topics as the teacher to train a student topic prediction model, which learns to predict the latent topics from multimodal contents of videos. The topic prediction provides intermediate supervision to the learning process. As for the caption task, we propose a novel topic-aware decoder to generate more accurate and detailed video descriptions with the guidance from latent topics. The entire learning procedure is end-to-end and it optimizes both tasks simultaneously. The results from extensive experiments conducted on the MSR-VTT and Youtube2Text datasets demonstrate the effectiveness of our proposed model. M&M TGM not only outperforms prior state-of-the-art methods on multiple evaluation metrics and on both benchmark datasets, but also achieves better generalization ability. Shizhe Chen, Jia Chen 0001, Qin Jin, Alex Hauptmann 0001 |
ACM Multimedia | 2 |
| 2017 | Knowing Yourself: Improving Video Caption via In-depth RecapabstractGenerating natural language descriptions for videos (a.k.a video captioning) has attracted much research attention in recent years, and a lot of models have been proposed to improve the caption performance. However, due to the rapid progress in dataset expansion and feature representation, newly proposed caption models have been evaluated on different settings, which makes it unclear about the contributions from either features or models. Therefore, in this work we aim to gain a deep understanding about "where are we" for the current development of video captioning. First, we carry out extensive experiments to identify the contribution from different components in video captioning task and make fair comparison among several state-of-the-art video caption models. Second, we discover that these state-of-the-art models are complementary so that we could benefit from "wisdom of the crowd" through ensembling and reranking. Finally, we give a preliminary answer to the question "how far are we from the human-level performance in general'' via a series of carefully designed experiments. In summary, our caption models achieve the state-of-the-art performance on the MSR-VTT 2017 challenge, and it is comparable with the average human-level performance on current caption metrics. However, our analysis also shows that we still have a long way to go, such as further improving the generalization ability of current caption models. Qin Jin, Shizhe Chen, Jia Chen 0001, Alex Hauptmann 0001 |
ACM Multimedia | 3 |
| 2016 | Semantic Image Profiling for Historic Events: Linking Images to PhrasesabstractAutomatically generating image profiles for historic events is desired for history knowledge preservation and curation. However, a simple profile with groups of related images lacks explicit semantic information, such as which images correspond to which aspects of the event. In this paper, we propose to add explicit semantic information to image profiling by linking images in the profile with related phrases in the event description. We measure the relevance of an image-phrase pair via a real-valued matching score. We exploit instance-wise ranking loss function to learn the matching score and we deal with two challenges: 1) how to automatically generate labeled positive data: we leverage out-of-domain labeled datasets to generate pseudo positive in-domain labels and propose a new algorithm (WIL4PPL) to robustly learn the model from the noisy pseudo positive labels; 2) how to automatically generate negative data: we propose a negative set generation algorithm to guide the model in learning which phrases and images to distinguish. We compare our model to three baselines and conduct detailed analysis and case studies to verify the quality of learnt semantic information. The extensive experiment results show the effectiveness of our proposed algorithms which significantly outperform the baselines. Jia Chen 0001, Qin Jin, Yifan Xiong 0001 |
ACM Multimedia | 1 |
| 2016 | Describing Videos using Multi-modal FusionabstractDescribing videos with natural language is one of the ultimate goals of video understanding. Video records multi-modal information including image, motion, aural, speech and so on. MSR Video to Language Challenge provides a good chance to study multi-modality fusion in caption task. In this paper, we propose the multi-modal fusion encoder and integrate it with text sequence decoder into an end-to-end video caption framework. Features from visual, aural, speech and meta modalities are fused together to represent the video contents. Long Short-Term Memory Recurrent Neural Networks (LSTM-RNNs) are then used as the decoder to generate natural language sentences. Experimental results show the effectiveness of multi-modal fusion encoder trained in the end-to-end framework, which achieved top performance in both common metrics evaluation and human evaluation. Qin Jin, Jia Chen 0001, Shizhe Chen, Yifan Xiong 0001, Alex Hauptmann 0001 |
ACM Multimedia | 2 |
| 2016 | History Rhyme: Searching Historic Events by Multimedia KnowledgeabstractThis demo shows a novel system "History Rhyme" which searches historic events by multimedia knowledge. Different from existing historic events related works, we focus on the retrieval of historic events by semantic related multimedia knowledge. Our system can not only search historic events based on keyword queries, but also retrieve similar historic events to a given event based on chosen facets. In both cases, the system returns top retrieval results and shows the image profiles of each historic event. To build such functions, we automatically mine knowledge from multimedia data and index each historic event. Our online demo is available at http://222.29.193.164/HistoryRhyme. Yifan Xiong 0001, Jia Chen 0001, Qin Jin |
ACM Multimedia | 2 |
| 2016 | Boosting Recommendation in Unexplored Categories by User Price PreferenceabstractState-of-the-art methods for product recommendation encounter a significant performance drop in categories where a user has no purchase history. This problem needs to be addressed since current online retailers are moving beyond single category and attempting to be diversified. In this article, we investigate the challenging problem of product recommendation in unexplored categories and discover that the price, a factor comparable across categories, can improve the recommendation performance significantly. We introduce the price utility concept to characterize users’ sense of price and propose three different utility functions. We show that user price preference in a category is a distribution and we mine typical user price preference patterns based on three different types of distance between distributions. We fuse user price preference through regularization and joint factorization to boost recommendation performance in both browsing and buying shopping orientations. Experimental results show that fusing user price preference improves performance in a series of recommendation tasks: unexplored category recommendation, product recommendation under a given unexplored category, and product recommendation under generic unexplored categories. Jia Chen 0001, Qin Jin, Shiwan Zhao, Shenghua Bao, Li Zhang 0007, Zhong Su, Yong Yu 0001 |
ACM Trans. Inf. Syst. | 1 |
| 2015 | Image Profiling for History Events on the FlyabstractHistory event related knowledge is precious and imagery is a powerful medium that records diverse information about the event. In this paper, we propose to automatically construct an image profile given a one sentence description of the historic event which contains where, when, who and what elements. Such a simple input requirement makes our solution easy to scale up and support a wide range of culture preservation and curation related applications ranging from wikipedia enrichment to history education. However, history relevant information on the web is available as "wild and dirty" data, which is quite different from clean, manually curated and structured information sources. There are two major challenges to build our proposed image profiles: 1) unconstrained image genre diversity. We categorize images into genres of documents/maps, paintings or photos. Image genre classification involves a full-spectrum of features from low-level color to high-level semantic concepts. 2) image content diversity. It can include faces, objects and scenes. Furthermore, even within the same event, the views and subjects of images are diverse and correspond to different facets of the event. To solve this challenge, we group images at two levels of granularity: iconic image grouping and facet image grouping. These require different types of features and analysis from near exact matching to soft semantic similarity. We develop a full-range feature analysis module which is composed of several levels, each suitable for different types of image analysis tasks. The wide range of features are based on both classical hand-crafted features and different layers of a convolutional neural network. We compare and study the performance of the different levels in the full-range features and show their effectiveness on handling such a wild, unconstrained dataset. Jia Chen 0001, Qin Jin, Yong Yu 0001, Alex Hauptmann 0001 |
ACM Multimedia | 1 |
| 2015 | Lead curve detection in drawings with complex cross-points
Jia Chen 0001, Min Li 0022, Qin Jin, Shenghua Bao, Zhong Su, Yong Yu 0001 |
Neurocomputing | 1 |
| 2015 | Exploitation and Exploration Balanced Hierarchical Summary for Landmark ImagesabstractWhile we have made significant progress over image understanding and search, how to meet the ultimate goal of satisfying both exploration and exploitation in one single system is still an open challenge. In the context of landmark images, it means that a system should not only be able to help users to quickly locate the photo they are interested in (exploitation), but also to discover different parts of the landmark which have never been seen before (exploration), which is a common request as evidenced by many recent multimedia studies. To the best of our knowledge, existing systems mainly focus on either exploration (e.g., photo browsing) or exploitation (e.g., representative photo identification), while users' need of exploration and exploitation is dynamically mixed. In this paper, we tackle the challenge by organizing landmark images into a hierarchical summary which gives user the flexibility of conducting both exploration and exploitation. In the hierarchical summary construction, we introduce two principles: the coherence principle and the diversity principle. Behind these two principles, the intrinsic concept is “detail-level,” which measures how much detail that an image reflects for a certain landmark. A new objective function is derived from the definition of both exploration and exploitation experience on detail-level. The problem of finding an optimal hierarchical summary is formulated as searching over a space of trees for the one that achieves the best objective score. Extensive quantitative experimental results and comprehensive user studies show that the optimized hierarchical summary is able to satisfy both experiences simultaneously. Jia Chen 0001, Qin Jin, Shenghua Bao, Zhong Su, Shimin Chen, Yong Yu 0001 |
IEEE Trans. Multim. | 1 |
| 2014 | Does product recommendation meet its waterloo in unexplored categories?: no, price comes to helpabstractState-of-the-art methods for product recommendation encounter significant performance drop in categories where a user has no purchase history. This problem needs to be addressed since current online retailers are moving beyond single category and attempting to be diversified. In this paper, we investigate the challenge problem of product recommendation in unexplored categories and discover that the price, a factor transferrable across categories, can improve the recommendation performance significantly. Through our investigation, we address four research questions progressively: 1) what is the impact of unexplored category on recommendation performance? 2) How to represent the price factor from the recommendation point of view? 3) What does price factor across categories mean to recommendation? 4) How to utilize price factor across categories for recommendation in unexplored categories? Based on a series of experiments and analysis conducted on a dataset collected from a leading E-commerce website, we discover valuable findings for the above four questions: first, unexplored categories cause performance drop by 40% relatively for current recommendation systems; second, the price factor can be represented as either a quantity for a product or a distribution for a user to improve performance; third, consumer behavior with respect to price factor across categories is complicated and needs to be carefully modeled; finally and most importantly, we propose a new method which encodes the two perspectives of the price factor. The proposed method significantly improves the recommendation performance in unexplored categories over the state-of-the-art baseline systems and shortens the performance gap by 43% relatively. Jia Chen 0001, Qin Jin, Shiwan Zhao, Shenghua Bao, Li Zhang 0007, Zhong Su, Yong Yu 0001 |
SIGIR | 1 |
| 2014 | Unified Structured Learning for Simultaneous Human Pose Estimation and Garment Attribute ClassificationabstractIn this paper, we utilize structured learning to simultaneously address two intertwined problems: 1) human pose estimation (HPE) and 2) garment attribute classification (GAC), which are valuable for a variety of computer vision and multimedia applications. Unlike previous works that usually handle the two problems separately, our approach aims to produce an optimal joint estimation for both HPE and GAC via a unified inference procedure. To this end, we adopt a preprocessing step to detect potential human parts from each image (i.e., a set of candidates) that allows us to have a manageable input space. In this way, the simultaneous inference of HPE and GAC is converted to a structured learning problem, where the inputs are the collections of candidate ensembles, outputs are the joint labels of human parts and garment attributes, and joint feature representation involves various cues such as pose-specific features, garment-specific features, and cross-task features that encode correlations between human parts and garment attributes. Furthermore, we explore the strong edge evidence around the potential human parts so as to derive more powerful representations for oriented human parts. Such evidences can be seamlessly integrated into our structured learning model as a kind of energy function, and the learning process could be performed by standard structured support vector machines algorithm. However, the joint structure of the two problems is a cyclic graph, which hinders efficient inference. To resolve this issue, we compute instead approximate optima using an iterative procedure, where in each iteration, the variables of one problem are fixed. In this way, satisfactory solutions can be efficiently computed by dynamic programming. Experimental results on two benchmark data sets show the state-of-the-art performance of our approach. Jie Shen 0005, Guangcan Liu, Jia Chen 0001, Yuqiang Fang, Jianbin Xie, Yong Yu 0001, Shuicheng Yan |
IEEE Trans. Image Process. | 3 |
| 2013 | Aggregation-Based Probing for Large-Scale Duplicate Image Detection
Ziming Feng, Jia Chen 0001, Xian Wu 0001, Yong Yu 0001 |
APWeb | 2 |
| 2013 | Tell me what happened here in historyabstractThis demo shows our system that takes a landmark image as input, recognizes the landmark from the image and returns historical events of the landmark with related photos. Different from existing landmark related researches, we focus on the temporal dimension of a landmark. Our system automatically recognizes the landmark, shows historical events chronologically and provides detailed photos for the events. To build these functions, we fuse information from multiple online resources. Jia Chen 0001, Qin Jin, Shenghua Bao, Zhong Su, Yong Yu 0001 |
ACM Multimedia | 1 |
| 2012 | DLMSearch: diversified landmark search by photoabstractThis paper focuses on the problem of searching for diversified landmarks with photos. More particularly, we propose a system called DLMSearch which handles image query, searches for diversified landmarks and provides representative visual summaries. DLMSearch allows a user to upload a query photo and searches for landmarks with high relevance and diversity in real time. Then DLMSearch presents a delicate photo summary for each returned landmark, considering both visual representativeness and diversity. Quantative evaluations on a web-scale landmark photo collection demonstrate the effectiveness of the DLMSearch system. Experimental results verify the merits of the proposed system. Junfeng Ye, Jia Chen 0001, Zejia Chen, Yihe Zhu, Shenghua Bao, Zhong Su, Yong Yu 0001 |
ACM Multimedia | 2 |
| 2012 | Searching for diversified landmarks by photoabstractThis demo focuses on the problem of searching for diversified landmarks with photos as input. More particularly, we propose a system called DLMSearch that allows a user to upload a photo as a query and searches for a diverse set of relevant landmarks in real time. It also presents a photo summary for each retrieved landmark, considering both visual representativeness and diversity. Our online demo is available at http://lm.apexlab.org/landmark/demo. Junfeng Ye, Jia Chen 0001, Zejia Chen, Yihe Zhu, Shenghua Bao, Zhong Su, Yong Yu 0001 |
ACM Multimedia | 2 |
| 2012 | Effective and Efficient Multi-Facet Web Image Annotation
Jia Chen 0001, Yihe Zhu, Haofen Wang, Wei Jin 0006, Yong Yu 0001 |
J. Comput. Sci. Technol. | 1 |