EDBT 2026 Demo / reviewers in the wild / expert
Yilin Wang 0002
dblp:47/3464-2
· DBLP profile ↗
35ranked-venue papers
8as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 6 first-author · 11 since 2021Artificial intelligence and machine learning · 23 · 3 first-author · 13 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 4 · 2 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Simple Edits: X-Planner for Complex Instruction-Based Image EditingabstractRecent diffusion-based image editing methods have made great strides in text-guided tasks but often struggle with complex, indirect instructions. Additionally, current models frequently exhibit poor identity preservation, unintended edits, or rely on manual masks. To overcome these limitations, we introduce X-Planner, a Multimodal Large Language Model (MLLM)-based planning system that bridges user intent with editing model capabilities. X-Planner uses chain-of-thought reasoning to systematically break down complex instructions into simpler sub-instructions. For each one, X-Planner automatically generates precise edit types and segmentation masks, enabling localized, identity-preserving edits without applying external tools or models during inference. To enable the training of such a planner, we also introduce a fully automated, reproducible pipeline to generate large-scale, high-quality training data. Our complete system achieves state-of-the-art results on both existing and newly proposed complex instruction-based editing benchmarks. Chun-Hsiao Yeh, Yilin Wang 0002, Nanxuan Zhao, Hao (Richard) Zhang, Krishna Kumar Singh |
AAAI | 2 |
| 2025 | UniReal: Universal Image Generation and Editing via Learning Real-world DynamicsabstractWe introduce UniReal, a unified framework designed to address various image generation and editing tasks. Existing solutions often vary by tasks, yet share fundamental principles: preserving consistency between inputs and outputs while capturing visual variations. Inspired by recent video generation models that effectively balance consistency and variation across frames, we propose a unifying approach that treats image-level tasks as discontinuous video generation. Specifically, we treat varying numbers of input and output images as frames, enabling seamless support for tasks such as image generation, editing, customization, composition, etc. Although designed for image-level tasks, we leverage videos as a scalable source for universal supervision. UniReal learns world dynamics from large-scale videos, demonstrating advanced capability in handling shadows, reflections, pose variation, and object interaction, while also exhibiting emergent capability for novel applications. Xi Chen 0119, He Zhang 0004, Yuqian Zhou, Soo Ye Kim, Qing Liu 0017, Yijun Li 0001, Jianming Zhang 0001, Nanxuan Zhao, Yilin Wang 0002, Zhe Lin 0001, Hengshuang Zhao |
CVPR | 10 |
| 2025 | FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any GranularityabstractThe advent of large Vision-Language Models (VLMs) has significantly advanced multimodal tasks, enabling more sophisticated and accurate reasoning across various applications, including image and video captioning, visual question answering, and cross-modal retrieval. Despite their superior capabilities, VLMs struggle with fine-grained image regional composition information perception. Specifically, they have difficulty accurately aligning the segmentation masks with the corresponding semantics and precisely describing the compositional aspects of the referred regions. However, compositionality – the ability to understand and generate novel combinations of known visual and textual components – is critical for facilitating coherent reasoning and understanding across modalities by VLMs. To address this issue, we propose FineCaption, a novel VLM that can recognize arbitrary masks as referential inputs and process high-resolution images for compositional image captioning at different granularity levels. To support this endeavor, we introduce CompositionCap, a new dataset for multi-grained region compositional image captioning, which introduces the task of compositional attribute-aware regional image captioning. Empirical results demonstrate the effectiveness of our proposed model compared to other state-of-the-art VLMs. Additionally, we analyze the capabilities of current VLMs in recognizing various visual prompts for compositional region image captioning, highlighting areas for improvement in VLM design and training. https://hanghuacs.github.io/FineCaption/ Hang Hua, Qing Liu 0017, Lingzhi Zhang, Jing Shi 0005, Soo Ye Kim, Yilin Wang 0002, Jianming Zhang 0001, Zhe Lin 0001, Jiebo Luo 0001 |
CVPR | 7 |
| 2024 | Amodal Scene Analysis via Holistic Occlusion Relation Inference and Generative Mask CompletionabstractAmodal scene analysis entails interpreting the occlusion relationship among scene elements and inferring the possible shapes of the invisible parts. Existing methods typically frame this task as an extended instance segmentation or a pair-wise object de-occlusion problem. In this work, we propose a new framework, which comprises a Holistic Occlusion Relation Inference (HORI) module followed by an instance-level Generative Mask Completion (GMC) module. Unlike previous approaches, which rely on mask completion results for occlusion reasoning, our HORI module directly predicts an occlusion relation matrix in a single pass. This approach is much more efficient than the pair-wise de-occlusion process and it naturally handles mutual occlusion, a common but often neglected situation. Moreover, we formulate the mask completion task as a generative process and use a diffusion-based GMC module for instance-level mask completion. This improves mask completion quality and provides multiple plausible solutions. We further introduce a large-scale amodal segmentation dataset with high-quality human annotations, including mutual occlusions. Experiments on our dataset and two public benchmarks demonstrate the advantages of our method. code public available at https://github.com/zbwxp/Amodal-AAAI. Bowen Zhang 0009, Qing Liu 0017, Jianming Zhang 0001, Yilin Wang 0002, Liyang Liu, Zhe Lin 0001, Yifan Liu 0001 |
AAAI | 4 |
| 2024 | UniHuman: A Unified Model For Editing Human Images in the WildabstractHuman image editing includes tasks like changing a person's pose, their clothing, or editing the image according to a text prompt. However, prior work often tackles these tasks separately, overlooking the benefit of mutual reinforcement from learning them jointly. In this paper, we propose UniHuman, a unified model that addresses multiple facets of human image editing in real-world settings. To enhance the model's generation quality and generalization capacity, we leverage guidance from human visual encoders and introduce a lightweight pose-warping module that can exploit different pose representations, accommodating unseen textures and patterns. Furthermore, to bridge the disparity between existing human editing benchmarks with real-world data, we curated 400K high-quality human image-text pairs for training and collected 2K human images for out-of-domain testing, both encompassing diverse clothing styles, backgrounds, and age groups. Experiments on both in-domain and out-of-domain test sets demonstrate that UniHuman outperforms task-specific models by a significant margin. In user studies, UniHuman is preferred by the users in an average of 77% of cases. Our project is available at this link. Nannan Li 0004, Qing Liu 0017, Krishna Kumar Singh, Yilin Wang 0002, Jianming Zhang 0001, Bryan A. Plummer, Zhe Lin 0001 |
CVPR | 4 |
| 2024 | SwapAnything: Enabling Arbitrary Object Swapping in Personalized Image Editing
Nanxuan Zhao, Wei Xiong 0008, Qing Liu 0017, He Zhang 0004, Jianming Zhang 0001, Hyunjoon Jung, Yilin Wang 0002, Xin Wang 0061 |
ECCV (32) | 9 |
| 2023 | LightPainter: Interactive Portrait Relighting with Freehand ScribbleabstractRecent portrait relighting methods have achieved realistic results of portrait lighting effects given a desired lighting representation such as an environment map. However, these methods are not intuitive for user interaction and lack precise lighting control. We introduce LightPainter, a scribble-based relighting system that allows users to interactively manipulate portrait lighting effect with ease. This is achieved by two conditional neural networks, a delighting module that recovers geometry and albedo optionally conditioned on skin tone, and a scribble-based module for re-lighting. To train the relighting module, we propose a novel scribble simulation procedure to mimic real user scribbles, which allows our pipeline to be trained without any human annotations. We demonstrate high-quality and flexible portrait lighting editing capability with both quantitative and qualitative experiments. User study comparisons with commercial lighting editing tools also demonstrate consistent user preference for our method. Yiqun Mei, He Zhang 0004, Xuaner Cecilia Zhang, Jianming Zhang 0001, Zhixin Shu, Yilin Wang 0002, Zijun Wei, Hyunjoon Jung, Vishal M. Patel |
CVPR | 6 |
| 2023 | Interactive Portrait Harmonization
Jeya Maria Jose Valanarasu, He Zhang 0004, Jianming Zhang 0001, Yilin Wang 0002, Zhe Lin 0001, Jose Echevarria, Yinglan Ma, Zijun Wei, Kalyan Sunkavalli, Vishal M. Patel |
ICLR | 4 |
| 2023 | PHOTOSWAP: Personalized Subject Swapping in ImagesabstractIn an era where images and visual content dominate our digital landscape, the ability to manipulate and personalize these images has become a necessity.
Envision seamlessly substituting a tabby cat lounging on a sunlit window sill in a photograph with your own playful puppy, all while preserving the original charm and composition of the image.
We present \emph{Photoswap}, a novel approach that enables this immersive image editing experience through personalized subject swapping in existing images.
\emph{Photoswap} first learns the visual concept of the subject from reference images and then swaps it into the target image using pre-trained diffusion models in a training-free manner. We establish that a well-conceptualized visual subject can be seamlessly transferred to any image with appropriate self-attention and cross-attention manipulation, maintaining the pose of the swapped subject and the overall coherence of the image.
Comprehensive experiments underscore the efficacy and controllability of \emph{Photoswap} in personalized subject swapping. Furthermore, \emph{Photoswap} significantly outperforms baseline methods in human ratings across subject swapping, background preservation, and overall quality, revealing its vast application potential, from entertainment to professional editing. Yilin Wang 0002, Nanxuan Zhao, Tsu-Jui Fu, Wei Xiong 0008, Qing Liu 0017, He Zhang 0004, Jianming Zhang 0001, Hyunjoon Jung, Xin Wang 0061 |
NeurIPS | 2 |
| 2022 | Lite Vision Transformer with Enhanced Self-AttentionabstractDespite the impressive representation capacity of vision transformer models, current light-weight vision transformer models still suffer from inconsistent and incorrect dense predictions at local regions. We suspect that the power of their self-attention mechanism is limited in shallower and thinner networks. We propose Lite Vision Transformer (LVT), a novel light-weight transformer network with two enhanced self-attention mechanisms to improve the model performances for mobile deployment. For the low-level features, we introduce Convolutional Self-Attention (CSA). Unlike previous approaches of merging convolution and self-attention, CSA introduces local self-attention into the convolution within a kernel of size$3\times 3$to enrich low-level features in the first stage of LVT. For the high-level features, we propose Recursive Atrous Self-Attention (RASA), which utilizes the multi-scale context when calculating the similarity map and a recursive mechanism to increase the representation capability with marginal extra parameter cost. The superiority of LVT is demonstrated on ImageNet recognition, ADE20K semantic segmentation, and COCO panoptic segmentation. The code is made publicly available11https://github.com/Chenglin-Yang/LVT. Yilin Wang 0002, Jianming Zhang 0001, He Zhang 0004, Zijun Wei, Zhe Lin 0001, Alan L. Yuille |
CVPR | 2 |
| 2021 | Mask Guided Matting via Progressive Refinement NetworkabstractWe propose Mask Guided (MG) Matting, a robust matting framework that takes a general coarse mask as guidance. MG Matting leverages a network (PRN) design which encourages the matting model to provide self-guidance to progressively refine the uncertain regions through the decoding process. A series of guidance mask perturbation operations are also introduced in the training to further enhance its robustness to external guidance. We show that PRN can generalize to unseen types of guidance masks such as trimap and low-quality alpha matte, making it suitable for various application pipelines. In addition, we revisit the foreground color prediction problem for matting and propose a surprisingly simple improvement to address the dataset issue. Evaluation on real and synthetic benchmarks shows that MG Matting achieves state-of-the-art performance using various types of guidance inputs. Code and models are available at https://github.com/yucornetto/MGMatting. Qihang Yu, Jianming Zhang 0001, He Zhang 0004, Yilin Wang 0002, Zhe Lin 0001, Ning Xu 0007, Yutong Bai, Alan L. Yuille |
CVPR | 4 |
| 2021 | Multimodal Contrastive Training for Visual Representation LearningabstractWe develop an approach to learning visual representations that embraces multimodal data, driven by a combination of intra- and inter-modal similarity preservation objectives. Unlike existing visual pre-training methods, which solve a proxy prediction task in a single domain, our method exploits intrinsic data properties within each modality and semantic information from cross-modal correlation simultaneously, hence improving the quality of learned visual representations. By including multimodal training in a unified framework with different types of contrastive losses, our method can learn more powerful and generic visual features. We first train our model on COCO and evaluate the learned visual representations on various downstream tasks including image classification, object detection, and instance segmentation. For example, the visual representations pre-trained on COCO by our method achieve state-of-the-art top-1 validation accuracy of 55.3% on ImageNet classification, under the common transfer protocol. We also evaluate our method on the large-scale Stock images dataset and show its effectiveness on multi-label image tagging, and cross-modal retrieval tasks. Zhe Lin 0001, Jason Kuen, Jianming Zhang 0001, Yilin Wang 0002, Michael Maire, Ajinkya Kale, Baldo Faieta |
CVPR | 5 |
| 2021 | SSH: A Self-Supervised Framework for Image HarmonizationabstractImage harmonization aims to improve the quality of image compositing by matching the "appearance" (e.g., color tone, brightness and contrast) between foreground and background images. However, collecting large-scale annotated datasets for this task requires complex professional retouching. Instead, we propose a novel Self-Supervised Harmonization framework (SSH) that can be trained using just "free" natural images without being edited. We reformulate the image harmonization problem from a representation fusion perspective, which separately processes the foreground and background examples, to address the background occlusion issue. This framework design allows for a dual data augmentation method, where diverse [foreground, background, pseudo GT] triplets can be generated by cropping an image with perturbations using 3D color lookup tables (LUTs). In addition, we build a real-world harmonization dataset as carefully created by expert users, for evaluation and benchmarking purposes. Our results show that the proposed self-supervised method outperforms previous state-of-the-art methods in terms of reference metrics, visual quality, and subject user study. Code and dataset are available at https://github.com/VITA-Group/SSHarmonization. Yifan Jiang 0001, He Zhang 0004, Jianming Zhang 0001, Yilin Wang 0002, Zhe Lin 0001, Kalyan Sunkavalli, Simon Chen, Sohrab Amirghodsi, Sarah Kong, Zhangyang Wang |
ICCV | 4 |
| 2020 | Incorporating Reinforced Adversarial Learning in Autoregressive Image Generation
Kenan E. Ak, Ning Xu 0007, Zhe Lin 0002, Yilin Wang 0002 |
ECCV (21) | 4 |
| 2020 | Shape Adaptor: A Learnable Resizing Module
Shikun Liu, Zhe Lin 0001, Yilin Wang 0002, Jianming Zhang 0001, Federico Perazzi, Edward Johns |
ECCV (12) | 3 |
| 2019 | Multimodal Style Transfer via Graph CutsabstractAn assumption widely used in recent neural style transfer methods is that image styles can be described by global statics of deep features like Gram or covariance matrices. Alternative approaches have represented styles by decomposing them into local pixel or neural patches. Despite the recent progress, most existing methods treat the semantic patterns of style image uniformly, resulting unpleasing results on complex styles. In this paper, we introduce a more flexible and general universal style transfer technique: multimodal style transfer (MST). MST explicitly considers the matching of semantic patterns in content and style images. Specifically, the style image features are clustered into sub-style components, which are matched with local content features under a graph cut formulation. A reconstruction network is trained to transfer each sub-style and render the final stylized result. We also generalize MST to improve some existing methods. Extensive experiments demonstrate the superior effectiveness, robustness, and flexibility of MST. Yulun Zhang 0001, Yilin Wang 0002, Zhe Lin 0001, Yun Fu 0001, Jimei Yang |
ICCV | 3 |
| 2018 | Generalizing Graph Matching beyond Quadratic Assignment ModelabstractGraph matching has received persistent attention over decades, which can be formulated as a quadratic assignment problem (QAP). We show that a large family of functions, which we define as Separable Functions, can approximate discrete graph matching in the continuous domain asymptotically by varying the approximation controlling parameters. We also study the properties of global optimality and devise convex/concave-preserving extensions to the widely used Lawler's QAP form. Our theoretical findings show the potential for deriving new algorithms and techniques for graph matching. We deliver solvers based on two specific instances of Separable Functions, and the state-of-the-art performance of our method is verified on popular benchmarks. Tianshu Yu 0001, Junchi Yan, Yilin Wang 0002, Wei Liu 0005, Baoxin Li |
NeurIPS | 3 |
| 2018 | Weakly Supervised Facial Attribute Manipulation via Deep Adversarial NetworkabstractAutomatically manipulating facial attributes is challenging because it needs to modify the facial appearances, while keeping not only the person's identity but also the realism of the resultant images. Unlike the prior works on the facial attribute parsing, we aim at an inverse and more challenging problem called attribute manipulation by modifying a facial image in line with a reference facial attribute. Given a source input image and reference images with a target attribute, our goal is to generate a new image (i.e., target image) that not only possesses the new attribute but also keeps the same or similar content with the source image. In order to generate new facial attributes, we train a deep neural network with a combination of a perceptual content loss and two adversarial losses, which ensure the global consistency of the visual content while implementing the desired attributes often impacting on local pixels. The model automatically adjusts the visual attributes on facial appearances and keeps the edited images as realistic as possible. The evaluation shows that the proposed model can provide a unified solution to both local and global facial attribute manipulation such as expression change and hair style transfer. Moreover, we further demonstrate that the learned attribute discriminator can be used for attribute localization. Yilin Wang 0002, Suhang Wang, Guo-Jun Qi, Jiliang Tang, Baoxin Li |
WACV | 1 |
| 2018 | CrossFire: Cross Media Joint Friend and Item RecommendationsabstractFriend and item recommendation on a social media site is an important task, which not only brings conveniences to users but also benefits platform providers. However, recommendation for newly launched social media sites is challenging because they often lack user historical data and encounter data sparsity and cold-start problem. Thus, it is important to exploit auxiliary information to help improve recommendation performances on these sites. Existing approaches try to utilize the knowledge transferred from other mature sites, which often require overlapped users or similar items to ensure an effective knowledge transfer. However, these assumptions may not hold in practice because 1) Overlapped user set is often unavailable and costly to identify due to the heterogeneous user profile, content and network data, and 2) Different schemes to show item attributes across sites cause the attribute values inconsistent, incomplete, and noisy. Thus, how to transfer knowledge when no direct bridge is given between two social media sites remains a challenge. In addition, another auxiliary information we can exploit is the mutual benefit between social relationships and rating preferences within the platform. User-user relationships are widely used as side information to improve item recommendation, whereas how to exploit user-item interactions for friend recommendation is rather limited. To tackle these challenges, we propose aCross media jointF riend andI temRe commendation framework (CrossFire ), which can capture both 1) cross-platform knowledge transfer, and 2) within-platform correlations among user-user relations and user-item interactions. Empirical results on real-world datasets demonstrate the effectiveness of the proposed framework. Kai Shu, Suhang Wang, Jiliang Tang, Yilin Wang 0002, Huan Liu 0001 |
WSDM | 4 |
| 2018 | Understanding and Predicting Delay in Reciprocal RelationsabstractReciprocity in directed networks points to user»s willingness to return favors in building mutual interactions. High reciprocity has been widely observed in many directed social media networks such as following relations in Twitter and Tumblr. Therefore, reciprocal relations between users are often regarded as a basic mechanism to create stable social ties and play a crucial role in the formation and evolution of networks. Each reciprocity relation is formed by two parasocial links in a back-and-forth manner with a time delay. Hence, understanding the delay can help us gain better insights into the underlying mechanisms of network dynamics. Meanwhile, the accurate prediction of delay has practical implications in advancing a variety of real-world applications such as friend recommendation and marketing campaign. For example, by knowing when will users follow back, service providers can focus on the users with a potential long reciprocal delay for effective targeted marketing. This paper presents the initial investigation of the time delay in reciprocal relations. Our study is based on a large-scale directed network from Tumblr that consists of 62.8 million users and 3.1 billion user following relations with a timespan of multiple years (from 31 Oct 2007 to 24 Jul 2013). We reveal a number of interesting patterns about the delay that motivate the development of a principled learning model to predict the delay in reciprocal relations. Experimental results on the above mentioned dynamic networks corroborate the effectiveness of the proposed delay prediction model. Jundong Li, Jiliang Tang, Yilin Wang 0002, Yali Wan, Yi Chang 0001, Huan Liu 0001 |
WWW | 3 |
| 2018 | Exploring Hierarchical Structures for Recommender SystemsabstractItems in real-world recommender systems exhibit certain hierarchical structures. Similarly, user preferences also present hierarchical structures. Recent studies show that incorporating the hierarchy of items or user preferences can improve the performance of recommender systems. However, hierarchical structures are often not explicitly available, especially those of user preferences. Thus, there's a gap between the importance of hierarchies and their availability. In this paper, we investigate the problem of exploring the implicit hierarchical structures for recommender systems when they are not explicitly available. We propose a novel recommendation framework to bridge the gap, which enables us to explore the implicit hierarchies of users and items simultaneously. We then extend the framework to integrate explicit hierarchies when they are available, which gives a unified framework for both explicit and implicit hierarchical structures. Experimental results on real-world datasets demonstrate the effectiveness of the proposed framework by incorporating implicit and explicit structures. Suhang Wang, Jiliang Tang, Yilin Wang 0002, Huan Liu 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2017 | CLARE: A Joint Approach to Label Classification and Tag RecommendationabstractData classification and tag recommendation are both important and challenging tasks in social media. These two tasks are often considered independently and most efforts have been made to tackle them separately. However, labels in data classification and tags in tag recommendation are inherently related. For example, a Youtube video annotated with NCAA, stadium, pac12 is likely to be labeled as football, while a video/image with the class label of coast is likely to be tagged with beach, sea, water and sand. The existence of relations between labels and tags motivates us to jointly perform classification and tag recommendation for social media data in this paper. In particular, we provide a principled way to capture the relations between labels and tags, and propose a novel framework CLARE, which fuses data CLAssification and tag REcommendation into a coherent model. With experiments on three social media datasets, we demonstrate that the proposed framework CLARE achieves superior performance on both tasks compared to the state-of-the-art methods. Yilin Wang 0002, Suhang Wang, Jiliang Tang, Guo-Jun Qi, Huan Liu 0001, Baoxin Li |
AAAI | 1 |
| 2017 | Exploiting Hierarchical Structures for Unsupervised Feature SelectionabstractFeature selection has been proven to be effective and efficient in preparing high-dimensional data for many mining and learning tasks. Features of real-world high-dimensional data such as words of documents, pixels of images and genes of microarray data, usually present inherent hierarchical structures. In a hierarchical structure, features could share certain properties. Such information has been exploited to help supervised feature selection but it is rarely investigated for unsupervised feature selection, which is challenging due to the lack of labels. Since real world data is often unlabeled, it is of practical importance to study the problem of feature selection with hierarchical structures in an unsupervised setting. In particular, we provide a principled method to exploit hierarchical structures of features and propose a novel framework HUFS, which utilizes the given hierarchical structures to help select features without labels. Experimental study on real-world datasets is conducted to assess the effectiveness of the proposed framework. Suhang Wang, Yilin Wang 0002, Jiliang Tang, Charu C. Aggarwal, Suhas Ranganath, Huan Liu 0001 |
SDM | 2 |
| 2017 | Understanding and Discovering Deliberate Self-harm Content in Social MediaabstractStudies suggest that self-harm users found it easier to discuss self-harm-related thoughts and behaviors using social media than in the physical world. Given the enormous and increasing volume of social media data, on-line self-harm content is likely to be buried rapidly by other normal content. To enable voices of self-harm users to be heard, it is important to distinguish self-harm content from other types of content. In this paper, we aim to understand self-harm content and provide automatic approaches to its detection. We first perform a comprehensive analysis on self-harm social media using different input cues. Our analysis, the first of its kind in large scale, reveals a number of important findings. Then we propose frameworks that incorporate the findings to discover self-harm content under both supervised and unsupervised settings. Our experimental results on a large social media dataset from Flickr demonstrate the effectiveness of the proposed frameworks and the importance of our findings in discovering self-harm content. Yilin Wang 0002, Jiliang Tang, Jundong Li, Baoxin Li, Yali Wan, Clayton Mellina, Neil O'Hare, Yi Chang 0001 |
WWW | 1 |
| 2017 | What Your Images Reveal: Exploiting Visual Contents for Point-of-Interest RecommendationabstractThe rapid growth of Location-based Social Networks (LBSNs) provides a vast amount of check-in data, which facilitates the study of point-of-interest (POI) recommendation. The majority of the existing POI recommendation methods focus on four aspects, i.e., temporal patterns, geographical influence, social correlations and textual content indications. For example, user's visits to locations have temporal patterns and users are likely to visit POIs near them. In real-world LBSNs such as Instagram, users can upload photos associating with locations. Photos not only reflect users' interests but also provide informative descriptions about locations. For example, a user who posts many architecture photos is more likely to visit famous landmarks; while a user posts lots of images about food has more incentive to visit restaurants. Thus, images have potentials to improve the performance of POI recommendation. However, little work exists for POI recommendation by exploiting images. In this paper, we study the problem of enhancing POI recommendation with visual contents. In particular, we propose a new framework Visual Content Enhanced POI recommendation (VPOI), which incorporates visual contents for POI recommendations. Experimental results on real-world datasets demonstrate the effectiveness of the proposed framework. Suhang Wang, Yilin Wang 0002, Jiliang Tang, Kai Shu, Suhas Ranganath, Huan Liu 0001 |
WWW | 2 |
| 2016 | PPP: Joint Pointwise and Pairwise Image Label PredictionabstractPointwise label and pairwise label are both widely used in computer vision tasks. For example, supervised image classification and annotation approaches use pointwise label, while attribute-based image relative learning often adopts pairwise labels. These two types of labels are often considered independently and most existing efforts utilize them separately. However, pointwise labels in image classification and tag annotation are inherently related to the pairwise labels. For example, an image labeled with "coast" and annotated with "beach, sea, sand, sky" is more likely to have a higher ranking score in terms of the attribute "open", while "men shoes" ranked highly on the attribute "formal" are likely to be annotated with "leather, lace up" than "buckle, fabric". The existence of potential relations between pointwise labels and pairwise labels motivates us to fuse them together for jointly addressing related vision tasks. In particular, we provide a principled way to capture the relations between class labels, tags and attributes, and propose a novel framework PPP(Pointwise and Pairwise image label Prediction), which is based on overlapped group structure extracted from the pointwise-pairwise-label bipartite graph. With experiments on benchmark datasets, we demonstrate that the proposed framework achieves superior performance on three vision tasks compared to the state-of-the-art methods. Yilin Wang 0002, Suhang Wang, Jiliang Tang, Huan Liu 0001, Baoxin Li |
CVPR | 1 |
| 2016 | Scale-adaptive EigenEye for fast eye detection in wild web imagesabstractDetecting eyes in images is fundamental for many computer vision applications including face detection, face recognition, and human-computer interaction. Most existing methods are designed and tested on datasets acquired under controlled lab settings (e.g., fixed scale, known poses, clean background, etc.), leaving their performance to be further examined on real-world, uncontrolled images, such as on-line images. This paper presents an effort on developing a fast and accurate eye detector for on-line images for which the acquisition condition is unknown and varies from one image to another, resulting in unpredictable background and variable scales for the eyes/faces. The key idea is to develop a scale-adaptive EigenEye approach, which employs an approximate scale estimated from face detection to modulate the pre-trained EigenEye basis in searching for the best match in a test image. The effort also includes building a 2845-image dataset with accurately-annotated eye locations and size, which will be made public to the community for future comparative study. Evaluation using this dataset, with comparison with a few leading state-of-the-art approaches, demonstrates the advantages of the proposed method. Yilin Wang 0002, Peng Zhang 0017, Baoxin Li |
ICIP | 2 |
| 2016 | Efficient unsupervised abnormal crowd activity detection based on a spatiotemporal saliency detectorabstractApproaches to abnormality detection in crowded scene largely rely on supervised methods using discriminative models. In this paper, we presents a novel and efficient unsupervised learning method for video analysis. We start from visual saliency, which has been used in several vision tasks, e.g., image classification, object detection, and foreground segmentation. To detect saliency regions in video sequences, we propose a new approach for detecting spatiotemporal visual saliency based on the phase spectrum of the videos, which is easy to implement and computationally efficient. With the proposed algorithm, we also study how the spatiotemporal saliency can be used in two important vision tasks, saliency prediction and abnormality detection. The proposed algorithm is evaluated on several benchmark datasets with comparison to the state-of-the-art methods from the literature. The experiments demonstrate the effectiveness of the proposed approach to spatiotemporal visual saliency detection and its application to the above vision tasks. Yilin Wang 0002, Qiang Zhang 0021, Baoxin Li |
WACV | 1 |
| 2015 | Real-time vehicle back-up warning system with a single cameraabstractIn this paper, we propose a real-time system using vehicle back-up camera to alert for potential back-up collisions. We developed a highly efficient algorithm, combining segmenting pedestrians and vehicles from moving background using local optical flow value, and a scale adaptive method using Deformable Part Model to detect objects at different distances. To test out algorithm, we created our own vehicle back-up dataset that contains rich scenes recorded from a back-up camera on moving/stationary vehicles with unique and challenging scenarios such as frequent occlusion with cluttered and moving background, and we made this dataset available to public for other researchers. Experiments on the dataset shows that our algorithm achieves high accuracy in near real-time, and it is about 10 times faster than the comparable state-of-the-art algorithm. Yilin Wang 0002, Baoxin Li |
ICIP | 2 |
| 2015 | Structure-preserving Image Quality AssessmentabstractPerceptual Image Quality Assessment (IQA) has many applications. Existing IQA approaches typically work only for one of three scenarios: full-reference, non-reference, or reduced-reference. Techniques that attempt to incorporate image structure information often rely on hand-crafted features, making them difficult to be extended to handle different scenarios. On the other hand, objective metrics like Mean Square Error (MSE), while being easy to compute, are often deemed ineffective for measuring perceptual quality. This paper presents a novel approach to perceptual quality assessment by developing an MSE-like metric, which enjoys the benefit of MSE in terms of inexpensive computation and universal applicability while allowing structural information of an image being taken into consideration. The latter was achieved through introducing structure-preserving kernelization into a MSE-like formulation. We show that the method can lead to competitive FR-IQA results. Further, by developing a feature coding scheme based on this formulation, we extend the model to improve the performance of NR-IQA methods. We report extensive experiments illustrating the results from both our FR-IQA and NR-IQA algorithms with comparison to existing state-of-the-art methods. Yilin Wang 0002, Qiang Zhang 0021, Baoxin Li |
ICME | 1 |
| 2015 | Inferring Sentiment from Web Images with Joint Inference on Visual and Social Cues: A Regulated Matrix Factorization Approach
Yilin Wang 0002, Yuheng Hu, Subbarao Kambhampati, Baoxin Li |
ICWSM | 1 |
| 2015 | Exploring Implicit Hierarchical Structures for Recommender Systems
Suhang Wang, Jiliang Tang, Yilin Wang 0002, Huan Liu 0001 |
IJCAI | 3 |
| 2015 | Unsupervised Sentiment Analysis for Social Media Images
Yilin Wang 0002, Suhang Wang, Jiliang Tang, Huan Liu 0001, Baoxin Li |
IJCAI | 1 |
| 2015 | Improving Vision-Based Self-Positioning in Intelligent Transportation Systems via Integrated Lane and Vehicle DetectionabstractTraffic congestion is a widespread problem. Dynamic traffic routing systems and congestion pricing are getting importance in recent research. Lane prediction and vehicle density estimation is an important component of such systems. We introduce a novel problem of vehicle self positioning which involves predicting the number of lanes on the road and vehicle's position in those lanes using videos captured by a dashboard camera. We propose an integrated closed-loop approach where we use the presence of vehicles to aid the task of self-positioning and vice versa. To incorporate multiple factors and high-level semantic knowledge into the solution, we formulate this problem as a Bayesian framework. In the framework, the number of lanes, the vehicle's position in those lanes and the presence of other vehicles are considered as parameters. We also propose a bounding box selection scheme to reduce the number of false detections and increase the computational efficiency. We show that the number of box proposals decreases by a factor of 6 using the selection approach. It also results in large reduction in the number of false detections. The entire approach is tested on real-world videos and is found to give acceptable results. Parag S. Chandakkar, Yilin Wang 0002, Baoxin Li |
WACV | 2 |
| 2014 | Image Cosegmentation via Multi-task Learning
Qiang Zhang 0021, Yilin Wang 0002, Jieping Ye, Baoxin Li |
BMVC | 3 |