VLDB 2026 Research / reviewers in the wild / expert
Xirong Li 0001
dblp:58/5856
· DBLP profile ↗
111ranked-venue papers
22as first author
50since 2021 · last 2026
0000-0002-0220-8310ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 89 · 21 first-author · 36 since 2021Artificial intelligence and machine learning · 34 · 1 first-author · 22 since 2021Databases, data management, data science and information retrieval · 16 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 5 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cross-Modal Fundus Image Registration Under Large FoV Disparity
Junyi Tao, Qijie Wei, Ningzhi Yang, Meng Wang 0001, Weihong Yu, Xirong Li 0001 |
MMM (1) | 7 |
| 2026 | Co-teaching for Unsupervised Domain Expansion
Hailan Lin, Qijie Wei, Kaibin Tian, Ruixiang Zhao, Xirong Li 0001 |
MMM (1) | 5 |
| 2026 | PRVR: Partially Relevant Video RetrievalabstractIn current text-to-video retrieval (T2VR), videos to be retrieved have been properly trimmed so that a correspondence between the videos and ad-hoc textual queries naturally exists. Note in practice that videos circulated on the Internet and social media platforms, while being relatively short, are typically rich in their content. Often, multiple scenes / actions / events are shown in a single video, leading to a more challenging T2VR setting wherein only part of the video content is relevant w.r.t. a given query. This paper presents a first study on this setting which we term Partially Relevant Video Retrieval (PRVR). Considering that a video typically consists of multiple moments, a video is regarded as partially relevant w.r.t. to a given query if it contains a query-related moment. We formulate the PRVR task as a multiple instance learning problem, and propose a Multi-Scale Similarity Learning (MS-SL++) network that jointly learns both clip-scale and frame-scale similarities to determine the partial relevance between video-query pairs. Extensive experiments on three diverse video-text datasets (TVshow Retrieval, ActivityNet-Captions and Charades-STA) demonstrate the viability of the proposed method. Xianke Chen, Daizong Liu, Xun Yang 0001, Xirong Li 0001, Jianfeng Dong, Meng Wang 0001, Xun Wang 0007 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | ASR-Enhanced Multimodal Representation Learning for Cross-Domain Product RetrievalabstractE-commerce is increasinglymultimedia-enriched, with products exhibited in a broad-domain manner as images, short videos, or live stream promotions. A unified and vectorized cross-domain production representation is essential. Due to large intra-product variance and high inter-product similarity in the broad-domain scenario, a visual-only representation is inadequate. While Automatic Speech Recognition (ASR) text derived from the short or live-stream videos is readily accessible, how to de-noise the excessively noisy text for multimodal representation learning is mostly untouched. We proposeASR-enhancedMultimodalProduct Representation Learning (AMPere). In order to extract product-specific information from the raw ASR text,AMPereuses an easy-to-implement LLM-based ASR text summarizer. The LLM-summarized text, together with visual data, is then fed into a multi-branch network to generate compact multimodal embeddings. Extensive experiments on a large-scale tri-domain dataset verify the effectiveness ofAMPerein obtaining a unified multimodal product representation that clearly improves cross-domain product retrieval. Ruixiang Zhao, Jian Jia, Yan Li 0043, Xuehan Bai, Quan Chen 0006, Han Li 0005, Peng Jiang 0002, Xirong Li 0001 |
IEEE Trans. Multim. | 8 |
| 2025 | D&M: Enriching E-commerce Videos with Sound Effects by Key Moment Detection and SFX MatchingabstractVideos showcasing specific products are increasingly important for E-commerce. Key moments naturally exist as the first appearance of a specific product, presentation of its distinctive features, the presence of a buying link, etc. Adding proper sound effects (SFX) to such moments, or video decoration with SFX (VDSFX), is crucial for enhancing user engaging experience. Previous work adds SFX to videos by video-to-SFX matching at a holistic level, lacking the ability of adding SFX to a specific moment. Meanwhile, previous studies on video highlight detection or video moment retrieval consider only moment localization, leaving moment to SFX matching untouched. By contrast, we propose in this paper D&M, a unified method that accomplishes key moment detection and moment-to-SFX matching simultaneously. Moreover, for the new VDSFX task we build a large-scale dataset SFX-Moment from an E-commerce video creation platform. For a fair comparison, we build competitive baselines by extending a number of current video moment detection methods to the new task. Extensive experiments on SFX-Moment show the superior performance of the proposed method over the baselines. Minquan Wang, Bo Wang 0011, Aozhu Chen, Quan Chen 0006, Peng Jiang 0002, Xirong Li 0001 |
AAAI | 8 |
| 2025 | PhD: A ChatGPT-Prompted Visual Hallucination Evaluation DatasetabstractMultimodal Large Language Models (MLLMs) hallucinate, resulting in an emerging topic of visual hallucination evaluation (VHE). This paper contributes a ChatGPT-Prompted visual hallucination evaluation Dataset (PhD) for objective VHE at a large scale. The essence of VHE is to ask an MLLM questions about specific images to assess its susceptibility to hallucination. Depending on what to ask (objects, attributes, sentiment, etc.) and how the questions are asked, we structure PhD along two dimensions, i.e. task and mode. Five visual recognition tasks, ranging from low-level (object/attribute recognition) to middle-level (sentiment/position recognition and counting), are considered. Besides a normal visual QA mode, which we term PhD-base, PhD also asks questions with specious context (PhD-sec) or with incorrect context (PhD-icc), or with AI-generated counter common sense images (PhD-ccs). We construct PhD by a ChatGPT-assisted semi-automated pipeline, encompassing four pivotal modules: task-specific hallucinatory item (hitem) selection, hitem-embedded question generation, specious/incorrect context generation, and counter-common-sense (CCS) image generation. With over 14k daily images, 750 CCS images and 102k VQA triplets in total, PhD reveals considerable variability in MLLMs’ performance across various modes and tasks, offering valuable insights into the nature of hallucination. As such, PhD stands as a potent tool not only for VHE but may also play a significant role in the refinement of MLLMs. Yuhan Fu, Ruobing Xie, Runquan Xie, Xingwu Sun, Fengzong Lian, Zhanhui Kang, Xirong Li 0001 |
CVPR | 8 |
| 2025 | Convolutional Prompting for Broad-Domain Retinal Vessel Segmentation
Qijie Wei, Weihong Yu, Xirong Li 0001 |
ICASSP | 3 |
| 2025 | Hybrid-Tower: Fine-Grained Pseudo-Query Interaction and Generation for Text-to-Video RetrievalabstractThe Text-to-Video Retrieval (T2VR) task aims to retrieve unlabeled videos by textual queries with the same semantic meanings. Recent CLIP-based approaches have explored two frameworks: Two-Tower versus Single-Tower framework, yet the former suffers from low effectiveness, while the latter suffers from low efficiency. In this study, we explore a new Hybrid-Tower framework that can hybridize the advantages of the Two-Tower and Single-Tower framework, achieving high effectiveness and efficiency simultaneously. We propose a novel hybrid method, Fine-grained Pseudo-query Interaction and Generation for T2VR, ie, PIG, which includes a new pseudo-query generator designed to generate a pseudo-query for each video. This enables the video feature and the textual features of pseudo-query to interact in a fine-grained manner, similar to the Single-Tower approaches to hold high effectiveness, even before the real textual query is received. Simultaneously, our method introduces no additional storage or computational overhead compared to the Two-Tower framework during the inference stage, thus maintaining high efficiency. Extensive experiments on five commonly used text-video retrieval benchmarks demonstrate that our method achieves a significant improvement over the baseline, with an increase of $1.6\% \sim 3.9\%$ in R@1. Furthermore, our method matches the efficiency of Two-Tower models while achieving near state-of-the-art performance, highlighting the advantages of the Hybrid-Tower framework. Bangxiang Lan, Ruobing Xie, Ruixiang Zhao, Xingwu Sun, Zhanhui Kang, Gang Yang 0001, Xirong Li 0001 |
ICCV | 7 |
| 2025 | Multi-Object Sketch Animation by Scene Decomposition and Motion PlanningabstractSketch animation, which brings static sketches to life by generating dynamic video sequences, has found widespread applications in GIF design, cartoon production, and daily entertainment. While current methods for sketch animation perform well in single-object sketch animation, they struggle in multi-object scenarios. By analyzing their failures, we identify two major challenges of transitioning from single-object to multi-object sketch animation: object-aware motion modeling and complex motion optimization. For multi-object sketch animation, we propose MoSketch based on iterative optimization through Score Distillation Sampling (SDS) and thus animating a multi-object sketch in a training-data free manner. To tackle the two challenges in a divide-and-conquer strategy, MoSketch has four novel modules, i.e., LLM-based scene decomposition, LLM-based motion planning, multi-grained motion refinement, and compositional SDS. Extensive qualitative and quantitative experiments demonstrate the superiority of our method over existing sketch animation approaches. MoSketch takes a pioneering step towards multi-object sketch animation, opening new avenues for future research and applications. Zijie Xin, Yuhan Fu, Ruixiang Zhao, Bangxiang Lan, Xirong Li 0001 |
ICCV | 6 |
| 2025 | Music Grounding by Short Video
Zijie Xin, Minquan Wang, Quan Chen 0006, Peng Jiang 0002, Xirong Li 0001 |
ICCV | 7 |
| 2025 | FunBench: Benchmarking Fundus Reading Skills of MLLMs
Qijie Wei, Kaiheng Qian, Xirong Li 0001 |
MICCAI (6) | 3 |
| 2025 | Learning Partially-Decorrelated Common Spaces for Ad-hoc Video SearchabstractAd-hoc Video Search (AVS) involves using a textual query to search for multiple relevant videos in a large collection of unlabeled short videos. The main challenge of AVS is the visual diversity of relevant videos. A simple query such as ''Find shots of a man and a woman dancing together indoors'' can span a multitude of environments, from brightly lit halls and shadowy bars to dance scenes in black-and-white animations. It is therefore essential to retrieve relevant videos as comprehensively as possible. Current solutions for the AVS task primarily fuse multiple features into one or more common spaces, yet overlook the need for diverse spaces. To fully exploit the expressive capability of individual features, we propose LPD, short for Learning Partially Decorrelated common spaces. LPD incorporates two key innovations: feature-specific common space construction and the de-correlation loss. Specifically, LPD learns a separate common space for each video and text feature, and employs de-correlation loss to diversify the ordering of negative samples across different spaces. To enhance the consistency of multi-space convergence, we designed an entropy-based fair multi-space triplet ranking loss. Extensive experiments on the TRECVID AVS benchmarks (2016-2023) justify the effectiveness of LPD. Moreover, diversity visualizations of LPD's spaces highlight its ability to enhance result diversity. Zijie Xin, Xirong Li 0001 |
ACM Multimedia | 3 |
| 2025 | AEFS: Adaptive Early Feature Selection for Deep Recommender SystemsabstractThe quality of features plays an important role in the performance of recommender systems. Recognizing this, feature selection has emerged as a crucial technique in refining recommender systems. Recent advancements leveraging Automated Machine Learning (AutoML) has drawn significant attention, particularly in two main categories: early feature selection and late feature selection, differentiated by whether the selection occurs before or after the embedding layer. The early feature selection selects a fixed subset of features and retrains the model, while the late feature selection, known as adaptive feature selection, dynamically adjusts feature choices for each data instance, recognizing the variability in feature significance. Although adaptive feature selection has shown remarkable improvements in performance, its main drawback lies in its post-embedding layer feature selection. This process often becomes cumbersome and inefficient in large-scale recommender systems with billions of ID-type features, leading to a highly sparse and parameter-heavy embedding layer. To overcome this, we introduce Adaptive Early Feature Selection (AEFS), a very simple method that not only adaptively selects informative features for each instance, but also significantly reduces the activated parameters of the embedding layer. AEFS employs a dual-model architecture, encompassing an auxiliary model dedicated to feature selection and a main model responsible for prediction. To ensure effective alignment between these two models, we incorporate two collaborative training loss constraints. Our extensive experiments on three benchmark datasets validate the efficiency and effectiveness of our approach. Notably, AEFS matches the performance of current state-of-theart Adaptive Late Feature Selection methods while achieving a significant reduction of 37. 5% in the activated parameters of the embedding layer. We believe that this work opens up new possibilities for feature selection. Gaofeng Lu, Chaonan Guo, Yuekui Yang, Xirong Li 0001 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2025 | Attribute guided adversarial editing for face privacy protectionabstractNowadays, the proliferation of portraits or photographs containing human faces on the internet has created significant risks of illegal privacy collection and analysis by intelligent systems. Previous attempts to protect against unauthorized identification by face recognition models have primarily involved manipulating or adding adversarial perturbations to photos. However, it remains a challenge to balance privacy protection effectiveness and maintaining image visual quality. That is, to successfully attack real-world black-box face recognition models, significant manipulation is required for the source image, which will obviously damage the image visual quality. To address these issues, we propose an attribute-guided face identity protection (AG-FIP) approach that can protect facial privacy effectively without introducing meaningless or conspicuous artifacts into the source image. The proposed method involves mapping the images to latent space and subsequently implementing an adversarial attack through attribute editing. An attribute selection module followed by an attribute adversarially editing module is proposed to enhance the efficiency and effectiveness of adversarial attacks. Experimental results demonstrate that our approach outperforms SOTAs in terms of confusing black-box face recognition models, commercial face recognition APIs, and image visual quality. Ziang Wang 0004, Fan Tang, Juan Cao 0001, Xirong Li 0001, Jintao Li 0001 |
Vis. Informatics | 5 |
| 2024 | Beyond Coarse-Grained Matching in Video-Text Retrieval
Aozhu Chen, Hazel Doughty, Xirong Li 0001, Cees Snoek |
ACCV (3) | 3 |
| 2024 | Tackling Long Code Search with Splitting, Encoding, and AggregatingabstractCode search with natural language helps us reuse existing code snippets. Thanks to the Transformer-based pretraining models, the performance of code search has been improved significantly. However, due to the quadratic complexity of multi-head self-attention, there is a limit on the input token length. For efficient training on standard GPUs like V100, existing pretrained code models, including GraphCodeBERT, CodeBERT, RoBERTa (code), take the first 256 tokens by default, which makes them unable to represent the complete information of long code that is greater than 256 tokens. To tackle the long code problem, we propose a new baseline SEA (Split, Encode and Aggregate), which splits long code into code blocks, encodes these blocks into embeddings, and aggregates them to obtain a comprehensive long code representation. With SEA, we could directly use Transformer-based pretraining models to model long code without changing their internal structure and re-pretraining. We also compare SEA with sparse Trasnformer methods. With GraphCodeBERT as the encoder, SEA achieves an overall mean reciprocal ranking score of 0.785, which is 10.1% higher than GraphCodeBERT on the CodeSearchNet benchmark, justifying SEA as a strong baseline for long code search. Yanlin Wang 0001, Lun Du, Hongyu Zhang 0002, Dongmei Zhang 0001, Xirong Li 0001 |
LREC/COLING | 6 |
| 2024 | Holistic Features are Almost Sufficient for Text-to-Video RetrievalabstractFor text-to-video retrieval (T2VR), which aims to retrieve unlabeled videos by ad-hoc textual queries, CLIP-based methods currently lead the way. Compared to CLIP4Clip which is efficient and compact, state-of-the-art models tend to compute video-text similarity through fine-grained cross-modal feature interaction and matching, putting their scalability for large-scale T2VR applications into doubt. We propose TeachCLIP, enabling a CLIP4Clip based student network to learn from more advanced yet computationally intensive models. In order to create a learning channel to convey fine-grained cross-modal knowledge from a heavy model to the student, we add to CLIP4Clip a simple Attentionalframe-Feature Aggregation (AFA) block, which by design adds no extra storage /computation overhead at the retrieval stage. Frame-text relevance scores calculated by the teacher network are used as soft labels to supervise the attentive weights produced by AFA. Extensive experiments on multiple public datasets justify the viability of the proposed method. TeachCLIP has the same efficiency and compact-ness as CLIP4Clip, yet has near-SOTA effectiveness. Kaibin Tian, Ruixiang Zhao, Zijie Xin, Bangxiang Lan, Xirong Li 0001 |
CVPR | 5 |
| 2024 | Cliprerank: An Extremely Simple Method For Improving Ad-Hoc Video SearchabstractAd-hoc Video Search (AVS) enables users to search for unlabeled video content using on-the-fly textual queries. Current deep learning-based models for AVS are trained to optimize holistic similarity between short videos and their associated descriptions. However, due to the diversity of ad-hoc queries, even for a short video, its truly relevant part w.r.t. a given query can be of shorter duration. In such a scenario, the holistic similarity becomes suboptimal. To remedy the issue, we propose in this paper CLIPRerank, a fine-grained re-scoring method. We compute cross-modal similarities between query and video frames using a pre-trained CLIP model, with multi-frame scores aggregated by max pooling. The fine-grained score is weightedly added to the initial score for search result reranking. As such, CLIPRerank is agnostic to the underlying video retrieval models and extremely simple, making it a handy plug-in for boosting AVS. Experiments on the challenging TRECVID AVS benchmarks (from 2016 to 2021) justify the effectiveness of the proposed strategy. CLIPRerank consistently improves the TRECVID top performers and multiple existing models including SEA, W2VV++, Dual Encoding, Dual Task, LAFF, CLIP2Video, TS2-Net and X-CLIP. Our method also works when substituting BLIP-2 for CLIP. Aozhu Chen, Fangming Zhou, Xirong Li 0001 |
ICASSP | 4 |
| 2023 | Supervised Domain Adaptation for Recognizing Retinal Diseases from Wide-Field Fundus ImagesabstractThis paper addresses the emerging task of recognizing multiple retinal diseases from wide-field (WF) and ultra-wide-field (UWF) fundus images. For an effective use of existing large amount of labeled color fundus photo (CFP) data and the relatively small amount of WF and UWF data, we propose a supervised domain adaptation method named Cross-domain Collaborative Learning (CdCL). Inspired by the success of fixed-ratio based mixup in unsupervised domain adaptation, we re-purpose this strategy for the current task. Due to the intrinsic disparity between the field-of-view of CFP and WF/UWF images, a scale bias naturally exists in a mixup sample that the anatomic structure from a CFP image will be considerably larger than its WF/UWF counterpart. The CdCL method resolves the issue by Scale-bias Correction, which employs Transformers for producing scale-invariant features. As demonstrated by extensive experiments on multiple datasets covering both WF and UWF images, the proposed method compares favorably against a number of competitive baselines. Qijie Wei, Jingyuan Yang 0004, Bo Wang 0011, Jinrui Wang, Jianchun Zhao, Niranchana Manivannan, Youxin Chen, Dayong Ding, Jing Zhou 0005, Xirong Li 0001 |
BIBM | 12 |
| 2023 | Towards Making a Trojan-Horse Attack on Text-to-Image RetrievalabstractWhile deep learning based image retrieval is reported to be vulnerable to adversarial attacks, existing works are mainly on image-to-image retrieval with their attacks performed at the front end via query modification. By contrast, we present in this paper the first study about a threat that occurs at the back end of a text-to-image retrieval (T2IR) system. Our study is motivated by the fact that the image collection indexed by the system will be regularly updated due to the arrival of new images from various sources such as web crawlers and advertisers. With malicious images indexed, it is possible for an attacker to indirectly interfere with the retrieval process, letting users see certain images that are completely irrelevant w.r.t. their queries. We put this thought into practice by proposing a novel Trojan-horse attack (THA). In particular, we construct a set of Trojan-horse images by first embedding word-specific adversarial information into a QR code and then putting the code on benign advertising images. A proof-of-concept evaluation, conducted on two popular T2IR datasets (Flickr30k and MS-COCO), shows the effectiveness of the proposed THA in a white-box mode. Aozhu Chen, Xirong Li 0001 |
ICASSP | 3 |
| 2023 | Geometrized Transformer for Self-Supervised Homography EstimationabstractFor homography estimation, we propose Geometrized Transformer (GeoFormer), a new detector-free feature matching method. Current detector-free methods, e.g. LoFTR, lack an effective mean to accurately localize small and thus computationally feasible regions for cross-attention diffusion. We resolve the challenge with an extremely simple idea: using the classical RANSAC geometry for attentive region search. Given coarse matches by LoFTR, a homography is obtained with ease. Such a homography allows us to compute cross-attention in a focused manner, where key/value sets required by Transformers can be reduced to small fix-sized regions rather than an entire image. Local features can thus be enhanced by standard Transformers. We integrate GeoFormer into the LoFTR framework. By minimizing a multi-scale cross-entropy based matching loss on auto-generated training data, the network is trained in a fully self-supervised manner. Extensive experiments are conducted on multiple real-world datasets covering natural images, heavily manipulated pictures and retinal images. The proposed method compares favorably against the state-of-the-art. Xirong Li 0001 |
ICCV | 2 |
| 2023 | SAFL-Net: Semantic-Agnostic Feature Learning Network with Auxiliary Plugins for Image Manipulation DetectionabstractSince image editing methods in real world scenarios cannot be exhausted, generalization is a core challenge for image manipulation detection, which could be severely weakened by semantically related features. In this paper we propose SAFL-Net, which constrains a feature extractor to learn semantic-agnostic features by designing specific modules with corresponding auxiliary tasks. Applying constraints directly to the features extracted by the encoder helps it learn semantic-agnostic manipulation trace features, which prevents the biases related to semantic information within the limited training data and improves generalization capabilities. The consistency of auxiliary boundary prediction task and original region prediction task is guaranteed by a feature transformation structure. Experiments on various public datasets and comparisons in multiple dimensions demonstrate that SAFL-Net is effective for image manipulation detection. Danding Wang, Xirong Li 0001, Juan Cao 0001 |
ICCV | 4 |
| 2023 | ChinaOpen: A Dataset for Open-world Multimodal LearningabstractThis paper introduces ChinaOpen, a dataset sourced from Bilibili, a popular Chinese video-sharing website, for open-world multimodal learning. While the state-of-the-art multimodal learning networks have shown impressive performance in automated video annotation and cross-modal video retrieval, their training and evaluation are primarily conducted on YouTube videos with English text. Their effectiveness on Chinese data remains to be verified. In order to support multimodal learning in the new context, we construct ChinaOpen-50k, a webly annotated training set of 50k Bilibili videos associated with user-generated titles and tags. Both text-based and content-based data cleaning are performed to remove low-quality videos in advance. For a multi-faceted evaluation, we build ChinaOpen-1k, a manually labeled test set of 1k videos. Each test video is accompanied with a manually checked user title and a manually written caption. Besides, each video is manually tagged to describe objects / actions / scenes shown in the visual content. The original user tags are also manually checked. Moreover, with all the Chinese text translated into English, ChinaOpen-1k is also suited for evaluating models trained on English data. In addition to ChinaOpen, we propose Generative Video-to-text Transformer (GVT) for Chinese video captioning. We conduct an extensive evaluation of the state-of-the-art single-task / multi-task models on the new dataset, resulting in a number of novel findings and insights. Aozhu Chen, Chengbo Dong, Kaibin Tian, Ruixiang Zhao, Xun Liang 0001, Zhanhui Kang, Xirong Li 0001 |
ACM Multimedia | 8 |
| 2023 | Revisiting Code Search in a Two-Stage ParadigmabstractWith a good code search engine, developers can reuse existing code snippets and accelerate software development process. Current code search methods can be divided into two categories: traditional information retrieval (IR) based and deep learning (DL) based approaches. DL-based approaches include the cross-encoder paradigm and the bi-encoder paradigm. However, both approaches have certain limitations. The inference of IR-based and bi-encoder models are fast; however, they are not accurate enough; while cross-encoder models can achieve higher search accuracy but consume more time. In this work, we propose TOSS, a two-stage fusion code search framework that can combine the advantages of different code search methods. TOSS first uses IR-based and bi-encoder models to efficiently recall a small number of top-K code candidates, and then uses fine-grained cross-encoders for finer ranking. Furthermore, we conduct extensive experiments on different code candidate volumes and multiple programming languages to verify the effectiveness of TOSS. We also compare TOSS with six data fusion methods. Experimental results show that TOSS is not only efficient, but also achieves state-of-the-art accuracy with an overall mean reciprocal ranking (MRR) score of 0.763, compared to the best baseline result on the CodeSearchNet benchmark of 0.713. Yanlin Wang 0001, Lun Du, Xirong Li 0001, Hongyu Zhang 0002, Shi Han, Dongmei Zhang 0001 |
WSDM | 4 |
| 2023 | Bias oriented unbiased data augmentation for cross-bias representation learning
Fan Tang, Juan Cao 0001, Xirong Li 0001, Danding Wang |
Multim. Syst. | 4 |
| 2023 | MVSS-Net: Multi-View Multi-Scale Supervised Networks for Image Manipulation DetectionabstractAs manipulating images by copy-move, splicing and/or inpainting may lead to misinterpretation of the visual content, detecting these sorts of manipulations is crucial for media forensics. Given the variety of possible attacks on the content, devising a generic method is nontrivial. Current deep learning based methods are promising when training and test data are well aligned, but perform poorly on independent tests. Moreover, due to the absence of authentic test images, their image-level detection specificity is in doubt. The key question is how to design and train a deep neural network capable of learning generalizable features sensitive to manipulations in novel data, whilst specific to prevent false alarms on the authentic. We propose multi-view feature learning to jointly exploit tampering boundary artifacts and the noise view of the input image. As both clues are meant to be semantic-agnostic, the learned features are thus generalizable. For effectively learning from authentic images, we train with multi-scale (pixel / edge / image) supervision. We term the new network MVSS-Net and its enhanced version MVSS-Net++. Experiments are conducted in both within-dataset and cross-dataset scenarios, showing that MVSS-Net++ performs the best, and exhibits better robustness against JPEG compression, Gaussian blur and screenshot based image re-capturing. Chengbo Dong, Xinru Chen, Ruohan Hu, Juan Cao 0001, Xirong Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | DRAG: Dynamic Region-Aware GCN for Privacy-Leaking Image DetectionabstractThe daily practice of sharing images on social media raises a severe issue about privacy leakage. To address the issue, privacy-leaking image detection is studied recently, with the goal to automatically identify images that may leak privacy. Recent advance on this task benefits from focusing on crucial objects via pretrained object detectors and modeling their correlation. However, these methods have two limitations: 1) they neglect other important elements like scenes, textures, and objects beyond the capacity of pretrained object detectors. 2) the correlation among objects is fixed, but a fixed correlation is not appropriate for all the images. To overcome the limitations, we propose the Dynamic Region-Aware Graph Convolutional Network (DRAG) that dynamically finds out crucial regions including objects and other important elements, and model their correlation adaptively for each input image. To find out crucial regions, we cluster spatially-correlated feature channels into several region-aware feature maps. Furthermore, we dynamically model the correlation with the self-attention mechanism and explore the interaction among the regions with a graph convolutional network. The DRAG achieved an accuracy of 87% on the largest dataset for privacy-leaking image detection, which is 10 percentage points higher than the state of the art. The further case study demonstrates that it found out crucial regions containing not only objects but other important elements like textures. The code and more details are in https://github.com/guang-yanng/DRAG. Guang Yang 0031, Juan Cao 0001, Qiang Sheng 0001, Peng Qi 0005, Xirong Li 0001, Jintao Li 0001 |
AAAI | 5 |
| 2022 | Deepfake Network Architecture AttributionabstractWith the rapid progress of generation technology, it has become necessary to attribute the origin of fake images. Existing works on fake image attribution perform multi-class classification on several Generative Adversarial Network (GAN) models and obtain high accuracies. While encouraging, these works are restricted to model-level attribution, only capable of handling images generated by seen models with a specific seed, loss and dataset, which is limited in real-world scenarios when fake images may be generated by privately trained models. This motivates us to ask whether it is possible to attribute fake images to the source models' architectures even if they are finetuned or retrained under different configurations. In this work, we present the first study on Deepfake Network Architecture Attribution to attribute fake images on architecture-level. Based on an observation that GAN architecture is likely to leave globally consistent fingerprints while traces left by model weights vary in different regions, we provide a simple yet effective solution named by DNA-Det for this problem. Extensive experiments on multiple cross-test setups and a large-scale dataset demonstrate the effectiveness of DNA-Det. Tianyun Yang, Ziyao Huang 0002, Juan Cao 0001, Xirong Li 0001 |
AAAI | 5 |
| 2022 | Lightweight Attentional Feature Fusion: A New Baseline for Text-to-Video Retrieval
Aozhu Chen, Fangming Zhou, Jianfeng Dong, Xirong Li 0001 |
ECCV (14) | 6 |
| 2022 | Semi-supervised Keypoint Detector and Descriptor for Retinal Image Matching
Xirong Li 0001, Qijie Wei, Jie Xu 0010, Dayong Ding |
ECCV (21) | 2 |
| 2022 | Semi-supervised Learning for Nerve Segmentation in Corneal Confocal Microscope Photography
Jun Wu 0022, Qi Pan, Jianchun Zhao, Gang Yang 0001, Xirong Li 0001, Dayong Ding |
MICCAI (4) | 10 |
| 2022 | Lesion Localization in OCT by Semi-Supervised Object DetectionabstractOver 300 million people worldwide are affected by various retinal diseases. By noninvasive Optical Coherence Tomography (OCT) scans, a number of abnormal structural changes in the retina, namely retinal lesions, can be identified. Automated lesion localization in OCT is thus important for detecting retinal diseases at their early stage. To conquer the lack of manual annotation for deep supervised learning, this paper presents a first study on utilizing semi-supervised object detection (SSOD) for lesion localization in OCT images. To that end, we develop a taxonomy to provide a unified and structured viewpoint of the current SSOD methods, and consequently identify key modules in these methods. To evaluate the influence of these modules in the new task, we build OCT-SS, a new dataset consisting of over 1k expert-labeled OCT B-scan images and over 13k unlabeled B-scans. Extensive experiments on OCT-SS identify Unbiased Teacher (UnT) as the best current SSOD method for lesion localization. Moreover, we improve over this strong baseline, with mAP increased from 49.34 to 50.86. Jianchun Zhao, Jingyuan Yang 0004, Weihong Yu, Youxin Chen, Xirong Li 0001 |
ICMR | 7 |
| 2022 | Partially Relevant Video RetrievalabstractCurrent methods for text-to-video retrieval (T2VR) are trained and tested on video-captioning oriented datasets such as MSVD, MSR-VTT and VATEX. A key property of these datasets is that videos are assumed to be temporally pre-trimmed with short duration, whilst the provided captions well describe the gist of the video content. Consequently, for a given paired video and caption, the video is supposed to be fully relevant to the caption. In reality, however, as queries are not known a priori, pre-trimmed video clips may not contain sufficient content to fully meet the query. This suggests a gap between the literature and the real world. To fill the gap, we propose in this paper a novel T2VR subtask termed Partially Relevant Video Retrieval (PRVR). An untrimmed video is considered to be partially relevant w.r.t. a given textual query if it contains a moment relevant to the query. PRVR aims to retrieve such partially relevant videos from a large collection of untrimmed videos. PRVR differs from single video moment retrieval and video corpus moment retrieval, as the latter two are to retrieve moments rather than untrimmed videos. We formulate PRVR as a multiple instance learning (MIL) problem, where a video is simultaneously viewed as a bag of video clips and a bag of video frames. Clips and frames represent video content at different time scales. We propose a Multi-Scale Similarity Learning (MS-SL) network that jointly learns clip-scale and frame-scale similarities for PRVR. Extensive experiments on three datasets (TVR, ActivityNet Captions, and Charades-STA) demonstrate the viability of the proposed method. We also show that our method can be used for improving video corpus moment retrieval. Jianfeng Dong, Xianke Chen, Minsong Zhang, Xun Yang 0001, Shujie Chen 0001, Xirong Li 0001, Xun Wang 0007 |
ACM Multimedia | 6 |
| 2022 | Learn to Understand Negation in Video RetrievalabstractNegation is a common linguistic skill that allows human to express what we do NOT want. Naturally, one might expect video retrieval to support natural-language queries with negation, e.g., finding shots of kids sitting on the floor and not playing with a dog. However, the state-of-the-art deep learning based video retrieval models lack such ability, as they are typically trained on video description datasets such as MSR-VTT and VATEX that lack negated descriptions. Their retrieved results basically ignore the negator in the sample query, incorrectly returning videos showing kids playing with dog. This paper presents the first study on learning to understand negation in video retrieval and make contributions as follows. By re-purposing two existing datasets (MSR-VTT and VATEX), we propose a new evaluation protocol for video retrieval with negation. We propose a learning based method for training a negation-aware video retrieval model. The key idea is to first construct a soft negative caption for a specific training video by partially negating its original caption, and then compute a bidirectionally constrained loss on the triplet. This auxiliary loss is weightedly added to a standard retrieval loss. Experiments on the re-purposed benchmarks show that re-training the CLIP (Contrastive Language-Image Pre-Training) model by the proposed method clearly improves its ability to handle queries with negation. In addition, the model performance on the original benchmarks is also improved. Aozhu Chen, Xirong Li 0001 |
ACM Multimedia | 4 |
| 2022 | Dual Encoding for Video Retrieval by TextabstractThis paper attacks the challenging problem of video retrieval by text. In such a retrieval paradigm, an end user searches for unlabeled videos by ad-hoc queries described exclusively in the form of a natural-language sentence, with no visual example provided. Given videos as sequences of frames and queries as sequences of words, an effective sequence-to-sequence cross-modal matching is crucial. To that end, the two modalities need to be first encoded into real-valued vectors and then projected into a common space. In this paper we achieve this by proposing a dual deep encoding network that encodes videos and queries into powerful dense representations of their own. Our novelty is two-fold. First, different from prior art that resorts to a specific single-level encoder, the proposed network performs multi-level encoding that represents the rich content of both modalities in a coarse-to-fine fashion. Second, different from a conventional common space learning algorithm which is either concept based or latent space based, we introduce hybrid space learning which combines the high performance of the latent space and the good interpretability of the concept space. Dual encoding is conceptually simple, practically effective and end-to-end trained with hybrid space learning. Extensive experiments on four challenging video datasets show the viability of the new method. Code and data are available at https://github.com/danieljf24/hybrid_space. Jianfeng Dong, Xirong Li 0001, Chaoxi Xu, Xun Yang 0001, Gang Yang 0001, Xun Wang 0007, Meng Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | BADet: Boundary-Aware 3D Object Detection from Point Clouds
Rui Qian 0002, Xirong Li 0001 |
Pattern Recognit. | 3 |
| 2022 | 3D Object Detection for Autonomous Driving: A Survey
Rui Qian 0002, Xirong Li 0001 |
Pattern Recognit. | 3 |
| 2022 | Reading-Strategy Inspired Visual Representation Learning for Text-to-Video RetrievalabstractThis paper aims for the task of text-to-video retrieval, where given a query in the form of a natural-language sentence, it is asked to retrieve videos which are semantically relevant to the given query, from a great number of unlabeled videos. The success of this task depends on cross-modal representation learning that projects both videos and sentences into common spaces for semantic similarity computation. In this work, we concentrate on video representation learning, an essential component for text-to-video retrieval. Inspired by the reading strategy of humans, we propose a Reading-strategy Inspired Visual Representation Learning (RIVRL) to represent videos, which consists of two branches: a previewing branch and an intensive-reading branch. The previewing branch is designed to briefly capture the overview information of videos, while the intensive-reading branch is designed to obtain more in-depth information. Moreover, the intensive-reading branch is aware of the video overview captured by the previewing branch. Such holistic information is found to be useful for the intensive-reading branch to extract more fine-grained features. Extensive experiments on three datasets are conducted, where our model RIVRL achieves a new state-of-the-art on TGIF and VATEX. Moreover, on MSR-VTT, our model using two video features shows comparable performance to the state-of-the-art using seven video features and even outperforms models pre-trained on the large-scale HowTo100M dataset. Code is available athttps://github.com/LiJiaBei-7/rivrl. Jianfeng Dong, Xianke Chen, Xiaoye Qu, Xirong Li 0001, Yuan He 0011, Xun Wang 0007 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Learning Two-Stream CNN for Multi-Modal Age-Related Macular Degeneration CategorizationabstractThis paper tackles automated categorization of Age-related Macular Degeneration (AMD), a common macular disease among people over 50. Previous research efforts mainly focus on AMD categorization with a single-modal input, let it be a color fundus photograph (CFP) or an OCT B-scan image. By contrast, we consider AMD categorization given a multi-modal input, a direction that is clinically meaningful yet mostly unexplored. Contrary to the prior art that takes a traditional approach of feature extraction plus classifier training that cannot be jointly optimized, we opt for end-to-end multi-modal Convolutional Neural Networks (MM-CNN). Our MM-CNN is instantiated by a two-stream CNN, with spatially-invariant fusion to combine information from the CFP and OCT streams. In order to visually interpret the contribution of the individual modalities to the final prediction, we extend the class activation mapping (CAM) technique to the multi-modal scenario. For effective training of MM-CNN, we develop two data augmentation methods. One is GAN-based CFP/OCT image synthesis, with our novel use of CAMs as conditional input of a high-resolution image-to-image translation GAN. The other method is Loose Pairing, which pairs a CFP image and an OCT image on the basis of their classes instead of eye identities. Experiments on a clinical dataset consisting of 1,094 CFP images and 1,289 OCT images acquired from 1,093 distinct eyes show that the proposed solution obtains better F1 and Accuracy than multiple baselines for multi-modal AMD categorization. Code and data are available at https://github.com/li-xirong/mmc-amd. Weisen Wang, Xirong Li 0001, Zhiyan Xu, Weihong Yu, Jianchun Zhao, Dayong Ding, Youxin Chen |
IEEE J. Biomed. Health Informatics | 2 |
| 2021 | Article Reranking by Memory-Enhanced Key Sentence Matching for Detecting Previously Fact-Checked ClaimsabstractQiang Sheng, Juan Cao, Xueyao Zhang, Xirong Li, Lei Zhong. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Qiang Sheng 0001, Juan Cao 0001, Xueyao Zhang, Xirong Li 0001 |
ACL/IJCNLP (1) | 4 |
| 2021 | Image Manipulation Detection by Multi-View Multi-Scale SupervisionabstractThe key challenge of image manipulation detection is how to learn generalizable features that are sensitive to manipulations in novel data, whilst specific to prevent false alarms on authentic images. Current research emphasizes the sensitivity, with the specificity overlooked. In this paper we address both aspects by multi-view feature learning and multi-scale supervision. By exploiting noise distribution and boundary artifact surrounding tampered regions, the former aims to learn semantic-agnostic and thus more generalizable features. The latter allows us to learn from authentic images which are nontrivial to be taken into account by current semantic segmentation network based methods. Our thoughts are realized by a new network which we term MVSS-Net. Extensive experiments on five benchmark sets justify the viability of MVSS-Net for both pixel-level and image-level manipulation detection. Xinru Chen, Chengbo Dong, Jiaqi Ji, Juan Cao 0001, Xirong Li 0001 |
ICCV | 5 |
| 2021 | Improving Fake News Detection by Using an Entity-enhanced Framework to Fuse Diverse Multimodal CluesabstractRecently, fake news with text and images have achieved more effective diffusion than text-only fake news, raising a severe issue of multimodal fake news detection. Current studies on this issue have made significant contributions to developing multimodal models, but they are defective in modeling the multimodal content sufficiently. Most of them only preliminarily model the basic semantics of the images as a supplement to the text, which limits their performance on detection. In this paper, we find three valuable text-image correlations in multimodal fake news: entity inconsistency, mutual enhancement, and text complementation. To effectively capture these multimodal clues, we innovatively extract visual entities (such as celebrities and landmarks) to understand the news-related high-level semantics of images, and then model the multimodal entity inconsistency and mutual enhancement with the help of visual entities. Moreover, we extract the embedded text in images as the complementation of the original text. All things considered, we propose a novel entity-enhanced multimodal fusion framework, which simultaneously models three cross-modal correlations to detect diverse multimodal fake news. Extensive experiments demonstrate the superiority of our model compared to the state of the art. Peng Qi 0005, Juan Cao 0001, Xirong Li 0001, Huan Liu 0031, Qiang Sheng 0001, Xiaoyue Mi, Yongbiao Lv, Chenyang Guo, Yingchao Yu |
ACM Multimedia | 3 |
| 2021 | Multi-Level Visual Representation with Semantic-Reinforced Learning for Video CaptioningabstractThis paper describes our bronze-medal solution for the video captioning task of the ACMMM2021 Pre-Training for Video Understanding Challenge. We depart from the Bottom-Up-Top-Down model, with technical improvements on both video content encoding and caption decoding. For encoding, we propose to extract multi-level video features that describe holistic scenes and fine-grained key objects, respectively. The scene-level and object-level features are enhanced separately by multi-head self-attention mechanisms before feeding them into the decoding module. Towards generating content-relevant and human-like captions, we train our network end-to-end by semantic-reinforced learning. Finally, in order to select the best caption from captions produced by distinct models, we perform caption reranking by cross-modal matching between a given video and each candidate caption. Both internal experiments on the MSR-VTT test set and external evaluations by the challenge organizers justify the viability of the proposed solution. Chengbo Dong, Xinru Chen, Aozhu Chen, Xirong Li 0001 |
ACM Multimedia | 6 |
| 2021 | Multi-Modal Multi-Instance Learning for Retinal Disease RecognitionabstractThis paper attacks an emerging challenge of multi-modal retinal disease recognition. Given a multi-modal case consisting of a color fundus photo (CFP) and an array of OCT B-scan images acquired during an eye examination, we aim to build a deep neural network that recognizes multiple vision-threatening diseases for the given case. As the diagnostic efficacy of CFP and OCT is disease-dependent, the network's ability of being both selective and interpretable is important. Moreover, as both data acquisition and manual labeling are extremely expensive in the medical domain, the network has to be relatively lightweight for learning from a limited set of labeled multi-modal samples. Prior art on retinal disease recognition focuses either on a single disease or on a single modality, leaving multi-modal fusion largely underexplored. We propose in this paper Multi-Modal Multi-Instance Learning (MM-MIL) for selectively fusing CFP and OCT modalities. Its lightweight architecture (as compared to current multi-head attention modules) makes it suited for learning from relatively small-sized datasets. For an effective use of MM-MIL, we propose to generate a pseudo sequence of CFPs by over sampling a given CFP. The benefits of this tactic include well balancing instances across modalities, increasing the resolution of the CFP input, and finding out regions of the CFP most relevant with respect to the final diagnosis. Extensive experiments on a real-world dataset consisting of 1,206 multi-modal cases from 1,193 eyes of 836 subjects demonstrate the viability of the proposed model. Xirong Li 0001, Hailan Lin, Jianchun Zhao, Dayong Ding, Weihong Yu, Youxin Chen |
ACM Multimedia | 1 |
| 2021 | Classifier Belief Optimization for Visual Categorization
Gang Yang 0001, Xirong Li 0001 |
MMM (1) | 2 |
| 2021 | Mining Dual Emotion for Fake News DetectionabstractEmotion plays an important role in detecting fake news online. When leveraging emotional signals, the existing methods focus on exploiting the emotions of news contents that conveyed by the publishers (i.e., publisher emotion). However, fake news often evokes high-arousal or activating emotions of people, so the emotions of news comments aroused in the crowd (i.e., social emotion) should not be ignored. Furthermore, it remains to be explored whether there exists a relationship between publisher emotion and social emotion (i.e., dual emotion), and how the dual emotion appears in fake news. In this paper, we verify that dual emotion is distinctive between fake and real news and propose Dual Emotion Features to represent dual emotion and the relationship between them for fake news detection. Further, we exhibit that our proposed features can be easily plugged into existing fake news detectors as an enhancement. Extensive experiments on three real-world datasets (one in English and the others in Chinese) show that our proposed feature set: 1) outperforms the state-of-the-art task-related emotional features; 2) can be well compatible with existing fake news detectors and effectively improve the performance of detecting fake news.1 2 Xueyao Zhang, Juan Cao 0001, Xirong Li 0001, Qiang Sheng 0001, Kai Shu |
WWW | 3 |
| 2021 | Detecting Adversarial Image Examples in Deep Neural Networks with Adaptive Noise ReductionabstractRecently, many studies have demonstrated deep neural network (DNN) classifiers can be fooled by the adversarial example, which is crafted via introducing some perturbations into an original sample. Accordingly, some powerful defense techniques were proposed. However, existing defense techniques often require modifying the target model or depend on the prior knowledge of attacks. In this paper, we propose a straightforward method for detecting adversarial image examples, which can be directly deployed into unmodified off-the-shelf DNN models. We consider the perturbation to images as a kind of noise and introduce two classic image processing techniques, scalar quantization and smoothing spatial filter, to reduce its effect. The image entropy is employed as a metric to implement an adaptive noise reduction for different kinds of images. Consequently, the adversarial example can be effectively detected by comparing the classification results of a given sample and its denoised version, without referring to any prior knowledge of attacks. More than 20,000 adversarial examples against some state-of-the-art DNN models are used to evaluate the proposed method, which are crafted with different attack techniques. The experiments show that our detection method can achieve a high overall F1 score of 96.39 percent and certainly raises the bar for defense-aware attacks. Bin Liang 0002, Miaoqiang Su, Xirong Li 0001, Wenchang Shi, XiaoFeng Wang 0001 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2021 | Feature Re-Learning with Data Augmentation for Video Relevance PredictionabstractPredicting the relevance between two given videos with respect to their visual content is a key component for content-based video recommendation and retrieval. Thanks to the increasing availability of pre-trained image and video convolutional neural network models, deep visual features are widely used for video content representation. However, as how two videos are relevant is task-dependent, such off-the-shelf features are not always optimal for all tasks. Moreover, due to varied concerns including copyright, privacy and security, one might have access to only pre-computed video features rather than original videos. We propose in this paper feature re-learning for improving video relevance prediction, with no need of revisiting the original video content. In particular, re-learning is realized by projecting a given deep feature into a new space by an affine transformation. We optimize the re-learning process by a novel negative-enhanced triplet ranking loss. In order to generate more training data, we propose a new data augmentation strategy which works directly on frame-level and video-level features. Extensive experiments in the context of the Hulu Content-based Video Relevance Prediction Challenge 2018 justify the effectiveness of the proposed method and its state-of-the-art performance for content-based video relevance prediction. Jianfeng Dong, Xun Wang 0007, Leimin Zhang, Chaoxi Xu, Gang Yang 0001, Xirong Li 0001 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2021 | SEA: Sentence Encoder Assembly for Video Retrieval by Textual QueriesabstractRetrieving unlabeled videos by textual queries, known as Ad-hoc Video Search (AVS), is a core theme in multimedia data management and retrieval. The success of AVS counts on cross-modal representation learning that encodes both query sentences and videos into common spaces for semantic similarity computation. Inspired by the initial success of previously few works in combining multiple sentence encoders, this paper takes a step forward by developing a new and general method for effectively exploiting diverse sentence encoders. The novelty of the proposed method, which we termSentence Encoder Assembly(SEA), is two-fold. First, different from prior art that uses only a single common space, SEA supports text-video matching in multiple encoder-specific common spaces. Such a property prevents the matching from being dominated by a specific encoder that produces an encoding vector much longer than other encoders. Second, in order to explore complementarities among the individual common spaces, we propose multi-space multi-loss learning. As extensive experiments on four benchmarks (MSR-VTT, TRECVID AVS 2016-2019, TGIF and MSVD) show, SEA surpasses the state-of-the-art. In addition, SEA is extremely ease to implement. All this makes SEA an appealing solution for AVS and promising for continuously advancing the task by harvesting new sentence encoders. Xirong Li 0001, Fangming Zhou, Chaoxi Xu, Jiaqi Ji, Gang Yang 0001 |
IEEE Trans. Multim. | 1 |
| 2021 | Unsupervised Domain Expansion for Visual CategorizationabstractExpanding visual categorization into a novel domain without the need of extra annotation has been a long-term interest for multimedia intelligence. Previously, this challenge has been approached by unsupervised domain adaptation (UDA). Given labeled data from a source domain and unlabeled data from a target domain, UDA seeks for a deep representation that is both discriminative and domain-invariant. While UDA focuses on the target domain, we argue that the performance on both source and target domains matters, as in practice which domain a test example comes from is unknown. In this article, we extend UDA by proposing a new task called unsupervised domain expansion (UDE), which aims to adapt a deep model for the target domain with its unlabeled data, meanwhile maintaining the model’s performance on the source domain. We propose Knowledge Distillation Domain Expansion (KDDE) as a general method for the UDE task. Its domain-adaptation module can be instantiated with any existing model. We develop a knowledge distillation-based learning mechanism, enabling KDDE to optimize a single objective wherein the source and target domains are equally treated. Extensive experiments on two major benchmarks, i.e., Office-Home and DomainNet, show that KDDE compares favorably against four competitive baselines, i.e., DDC, DANN, DAAN, and CDAN, for both UDA and UDE tasks. Our study also reveals that the current UDA models improve their performance on the target domain at the cost of noticeable performance loss on the source domain. Kaibin Tian, Dayong Ding, Gang Yang 0001, Xirong Li 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2020 | Deep Multiple Instance Learning with Spatial Attention for ROP Case Classification, Instance Selection and Abnormality LocalizationabstractThis paper tackles automated screening of Retinopathy of Prematurity (ROP), one of the most common causes of visual loss in childhood. Clinically, ROP screening per case requires multiple color fundus image instances that capture different zones of the (premature) retina. A desirable model shall not only make a decision at the case level, but also pinpoint which instances and what part of the instances are responsible for the decision. This paper makes the first attempt to accomplish three tasks, i.e. ROP case classification, instance selection and abnormality localization in a unified framework. To that end, we propose a new model that effectively combines instance-attention based deep multiple instance learning (MIL) and spatial attention (SA). The propose model, which we term MIL-SA, identifies positive instances in light of their contributions to case-level decision. Meanwhile, abnormal regions in the identified instances are automatically localized by the SA mechanism. Moreover, MIL-SA is learned from case-level binary labels exclusively, and in an end-to-end manner. Experiments on a large clinical dataset of 2,186 cases with 11,053 fundus images show the viability of the proposed model for all the three tasks. Xirong Li 0001, Wencui Wan, Jianchun Zhao, Qijie Wei, Junbo Rong, Pengyi Zhou, Limin Xu, Lijuan Lang, Chengzhi Niu, Dayong Ding, Xuemin Jin |
ICPR | 1 |
| 2020 | Learn to Segment Retinal Lesions and BeyondabstractTowards automated retinal screening, this paper makes an endeavor to simultaneously achieve pixel-level retinal lesion segmentation and image-level disease classification. Such a multi-task approach is crucial for accurate and clinically interpretable disease diagnosis. Prior art is insufficient due to three challenges, i.e., lesions lacking objective boundaries, clinical importance of lesions irrelevant to their size, and the lack of one-to-one correspondence between lesion and disease classes. This paper attacks the three challenges in the context of diabetic retinopathy (DR) grading. We propose Lesion-Net, a new variant of fully convolutional networks, with its expansive path redesigned to tackle the first challenge. A dual Dice loss that leverages both semantic segmentation and image classification losses is introduced to resolve the second challenge. Lastly, we build a multi-task network that employs Lesion-Net as a side-attention branch for both DR grading and result interpretation. A set of 12K fundus images is manually segmented by 45 ophthalmologists for 8 DR-related lesions, resulting in 290K manual segments in total. Extensive experiments on this large-scale dataset show that our proposed approach surpasses the prior art for multiple tasks including lesion segmentation, lesion classification and DR grading. Qijie Wei, Xirong Li 0001, Weihong Yu, Yongpeng Zhang, Bojie Hu, Bin Mo, Di Gong, Dayong Ding, Youxin Chen |
ICPR | 2 |
| 2020 | A GAN-based Domain Adaptation Method for Glaucoma DiagnosisabstractDomain adaptation is an important research topic in the field of computer vision, where the goal is to solve the difference of data distribution between different scenarios of the same task. In recent times, adversarial learning method becomes a mainstream approach to generate complicated images across diverse domains through optimizing deep networks, and it can also improve the recognition accuracy rate of deep networks despite existing domain shift or dataset bias. However, there are few effective efforts of domain adaptation for the disease diagnosis on fundus images. Fundus images are normally captured on different medical devices with different rules. When diagnosing glaucoma, there is a serious homogeneous domain shift, which means feature spaces between target domain and source domain images have a distribution shift although they are very similar. We propose a unified framework to solve this problem. Previous studies have shown that glaucoma can be monitored by analyzing the optic disc/cup and its surroundings. So we exploit a novel reconstruction loss which not only leverages unsupervised data to bring the source and target distributions closer but also keeps original target domain images label unchanged. The experimental results on several public and private datasets demonstrate that our method could increase the classification accuracy of glaucoma diagnosis. Yunzhe Sun, Gang Yang 0001, Dayong Ding, Gangwei Cheng, Jieping Xu, Xirong Li 0001 |
IJCNN | 6 |
| 2020 | High-Order Attention Networks for Medical Image Segmentation
Gang Yang 0001, Jun Wu 0022, Dayong Ding, Jie Xv, Gangwei Cheng, Xirong Li 0001 |
MICCAI (1) | 7 |
| 2020 | iCap: Interactive Image Captioning with Predictive TextabstractIn this paper we study a brand new topic of interactive image captioning with human in the loop. Different from automated image captioning where a given test image is the sole input in the inference stage, we have access to both the test image and a sequence of (incomplete) user-input sentences in the interactive scenario. We formulate the problem as Visually Conditioned Sentence Completion (VCSC). For VCSC, we propose ABD-Cap, asynchronous bidirectional decoding for image caption completion. With ABD-Cap as the core module, we build iCap, a web-based interactive image captioning system capable of predicting new text with respect to live input from a user. A number of experiments covering both automated evaluations and real user studies show the viability of our proposals. Zhengxiong Jia, Xirong Li 0001 |
ICMR | 2 |
| 2020 | A W2VV++ Case Study with Automated and Interactive Text-to-Video RetrievalabstractAs reported by respected evaluation campaigns focusing both on automated and interactive video search approaches, deep learning started to dominate the video retrieval area. However, the results are still not satisfactory for many types of search tasks focusing on high recall. To report on this challenging problem, we present two orthogonal task-based performance studies centered around the state-of-the-art W2VV++ query representation learning model for video retrieval. First, an ablation study is presented to investigate which components of the model are effective in two types of benchmark tasks focusing on high recall. Second, interactive search scenarios from the Video Browser Showdown are analyzed for two winning prototype systems implementing a selected variant of the model and providing additional querying and visualization components. The analysis of collected logs demonstrates that even with the state-of-the-art text search video retrieval model, it is still auspicious to integrate users into the search process for task types, where high recall is essential. Jakub Lokoc, Tomás Soucek, Patrik Veselý, Frantisek Mejzlík, Jiaqi Ji, Chaoxi Xu, Xirong Li 0001 |
ACM Multimedia | 7 |
| 2020 | Towards annotation-free evaluation of cross-lingual image captioningabstractCross-lingual image captioning, with its ability to caption an unlabeled image in a target language other than English, is an emerging topic in the multimedia field. In order to save the precious human resource from re-writing reference sentences per target language, in this paper we make a brave attempt towards annotation-free evaluation of cross-lingual image captioning. Depending on whether we assume the availability of English references, two scenarios are investigated. For the first scenario with the references available, we propose two metrics, i.e., WMDRel and CLinRel. WMDRel measures the semantic relevance between a model-generated caption and machine translation of an English reference using their Word Mover's Distance. By projecting both captions into a deep visual feature space, CLinRel is a visual-oriented cross-lingual relevance measure. As for the second scenario, which has zero reference and is thus more challenging, we propose CMedRel to compute a cross-media relevance between the generated caption and the image content, in the same visual feature space as used by CLinRel. We have conducted a number of experiments to evaluate the effectiveness of the three proposed metrics. The combination of WMDRel, CLinRel and CMedRel has a Spearman's rank correlation of 0.952 with the sum of BLEU-4, METEOR, ROUGE-L and CIDEr, four standard metrics computed using references in the target language. CMedRel alone has a Spearman's rank correlation of 0.786 with the standard metrics. The promising results show high potential of the new metrics for evaluation with no need of references in the target language. Aozhu Chen, Hailan Lin, Xirong Li 0001 |
MMAsia | 4 |
| 2020 | AttenNet: Deep Attention Based Retinal Disease Classification in OCT Images
Jun Wu 0022, Jianchun Zhao, Dayong Ding, Ningjiang Chen, Chunhui Jiang, Xuan Zou, Yuan Tian 0017, Zongjiang Shang, Kaiwei Wang, Xirong Li 0001, Gang Yang 0001, Jianping Fan 0001 |
MMM (2) | 16 |
| 2019 | Oval Shape Constraint based Optic Disc and Cup Segmentation in Fundus Photographs
Jun Wu 0022, Kaiwei Wang, Zongjiang Shang, Jie Xu 0010, Dayong Ding, Xirong Li 0001, Gang Yang 0001 |
BMVC | 6 |
| 2019 | Dual Encoding for Zero-Example Video RetrievalabstractThis paper attacks the challenging problem of zero-example video retrieval. In such a retrieval paradigm, an end user searches for unlabeled videos by ad-hoc queries described in natural language text with no visual example provided. Given videos as sequences of frames and queries as sequences of words, an effective sequence-to-sequence cross-modal matching is required. The majority of existing methods are concept based, extracting relevant concepts from queries and videos and accordingly establishing associations between the two modalities. In contrast, this paper takes a concept-free approach, proposing a dual deep encoding network that encodes videos and queries into powerful dense representations of their own. Dual encoding is conceptually simple, practically effective and end-to-end. As experiments on three benchmarks, i.e. MSR-VTT, TRECVID 2016 and 2017 Ad-hoc Video Search show, the proposed solution establishes a new state-of-the-art for zero-example video retrieval. Jianfeng Dong, Xirong Li 0001, Chaoxi Xu, Shouling Ji, Yuan He 0011, Gang Yang 0001, Xun Wang 0007 |
CVPR | 2 |
| 2019 | Two-Stream CNN with Loose Pair Training for Multi-modal AMD Categorization
Weisen Wang, Zhiyan Xu, Weihong Yu, Jianchun Zhao, Jingyuan Yang 0004, Zhikun Yang, Dayong Ding, Youxin Chen, Xirong Li 0001 |
MICCAI (1) | 11 |
| 2019 | Fully Deep Learning for Slit-Lamp Photo Based Nuclear Cataract Grading
Chaoxi Xu, Xiangjia Zhu, Wenwen He, Xixi He, Zongjiang Shang, Jun Wu 0022, Yinglei Zhang, Xianfang Rong, Zhennan Zhao, Dayong Ding, Xirong Li 0001 |
MICCAI (4) | 14 |
| 2019 | W2VV++: Fully Deep Learning for Ad-hoc Video SearchabstractAd-hoc video search (AVS) is an important yet challenging problem in multimedia retrieval. Different from previous concept-based methods, we propose a fully deep learning method for query representation learning. The proposed method requires no explicit concept modeling, matching and selection. The backbone of our method is the proposed W2VV++ model, a super version of Word2VisualVec (W2VV) previously developed for visual-to-text matching. W2VV++ is obtained by tweaking W2VV with a better sentence encoding strategy and an improved triplet ranking loss. With these simple yet important changes, W2VV++ brings in a substantial improvement. As our participation in the TRECVID 2018 AVS task and retrospective experiments on the TRECVID 2016 and 2017 data show, our best single model, with an overall inferred average precision (infAP) of 0.157, outperforms the state-of-the-art. The performance can be further boosted by model ensemble using late average fusion, reaching a higher infAP of 0.163. With W2VV++, we establish a new baseline for ad-hoc video search. Xirong Li 0001, Chaoxi Xu, Gang Yang 0001, Zhineng Chen, Jianfeng Dong |
ACM Multimedia | 1 |
| 2019 | Exploring Content-based Video Relevance for Video Click-Through Rate PredictionabstractThis paper describes our solution for the Hulu Challenge. To answer the challenge, we introduce two content-based models, namely, Cascading Mapping Network (CMN) and Relevant-Enhanced Deep Interest Network (REDIN). CMN predicts video Click-Through Rate (CTR) by predicting content-based video relevance. REDIN mainly improves the popular Deep Interest Network by adding explicit video relevance constraint, which provides guidance for low-level video feature learning thus helpful for CTR prediction. Based on the two models, our solution obtains Area Under Curve (AUC) score of 0.6022 and 0.6155 on the TV-shows and Movie track respectively. What is more, we are one of the only two teams giving scores of over 0.6 on both tracks. The results justify the effectiveness and stability of our proposed solution. Xun Wang 0007, Yali Du 0001, Leimin Zhang, Xirong Li 0001, Jianfeng Dong |
ACM Multimedia | 4 |
| 2019 | Four Models for Automatic Recognition of Left and Right Eye in Fundus Images
Xirong Li 0001, Rui Qian 0002, Dayong Ding, Jun Wu 0022, Jieping Xu |
MMM (1) | 2 |
| 2019 | A Coarse-to-fine Cascading Model for Cataract Nuclear Segmentation in Slit-lamp PhotographsabstractA nuclear cataract is an age-related chronic and priority ophthalmic disease in which a clouding of the lens in the human eye affects vision. Automatic segmentation of nuclear region based on slit-lamp photographs is a basic step for computer-aided diagnosis such as nuclear cataract grading. However, slit-lamp photographs collected from a clinic scenario often have complex background containing the eyelids, sclera and cornea with spectral highlights. The existing efforts using traditional image processing that have unsatisfactory results, and the deep learning method using standard Faster R-CNN tends to obtain a bigger nuclear contour. In this paper, we propose a coarse-to-fine deep learning solution to localize nuclear regions by cascading the Faster R-CNN in a two-stage framework. First, a nuclear ROI (region of interest) predictor is pre-trained to localize a rough position and remove complex backgrounds. Then, a fine nuclear locator is applied to predict a more compact nuclear bounding box. Finally, an ellipse-like nuclear contour is fitted based on its bounding box. Evaluated on a clinical dataset of 884 slit-lamp photographs, the proposed method outperforms the state-of-the-art, improving the overlapping rate (IoU) by 0.33% from 67.98% to 68.31%, and increasing the success rate by 2.55% from 85.71% to 88.26%. Jun Wu 0022, Xianfang Rong, Zhennan Zhao, Dayong Ding, Xirong Li 0001, Zongjiang Shang, Kaiwei Wang, Xixi He, Xiangjia Zhu, Wenwen He, Yinglei Zhang |
VCIP | 6 |
| 2019 | COCO-CN for Cross-Lingual Image Tagging, Captioning, and RetrievalabstractThis paper contributes to cross-lingual image annotation and retrieval in terms of data and baseline methods. We propose COCO-CN, a novel dataset enriching MS-COCO with manually written Chinese sentences and tags. For effective annotation acquisition, we develop a recommendation-assisted collective annotation system, automatically providing an annotator with several tags and sentences deemed to be relevant with respect to the pictorial content. Having 20 342 images annotated with 27 218 Chinese sentences and 70 993 tags, COCO-CN is currently the largest Chinese-English dataset that provides a unified and challenging platform for cross-lingual image tagging, captioning, and retrieval. We develop conceptually simple yet effective methods per task for learning from cross-lingual resources. Extensive experiments on the three tasks justify the viability of the proposed dataset and methods. Data and code are publicly available at https://github.com/li-xirong/coco-cn. Xirong Li 0001, Chaoxi Xu, Weiyu Lan, Zhengxiong Jia, Gang Yang 0001, Jieping Xu |
IEEE Trans. Multim. | 1 |
| 2018 | Laser Scar Detection in Fundus Images Using Convolutional Neural Networks
Qijie Wei, Xirong Li 0001, Dayong Ding, Weihong Yu, Youxin Chen |
ACCV (4) | 2 |
| 2018 | Cross-Class Sample Synthesis for Zero-shot Learning
Jinlu Liu, Xirong Li 0001, Gang Yang 0001 |
BMVC | 2 |
| 2018 | Deep Text Classification Can be FooledabstractIn this paper, we present an effective method to craft text adversarial samples, revealing one important yet underestimated fact that DNN-based text classifiers are also prone to adversarial sample attack. Specifically, confronted with different adversarial scenarios, the text items that are important for classification are identified by computing the cost gradients of the input (white-box attack) or generating a series of occluded test samples (black-box attack). Based on these items, we design three perturbation strategies, namely insertion, modification, and removal, to generate adversarial samples. The experiment results show that the adversarial samples generated by our method can successfully fool both state-of-the-art character-level and word-level DNN-based text classifiers. The adversarial samples can be perturbed to any desirable classes without compromising their utilities. At the same time, the introduced perturbation is difficult to be perceived. Bin Liang 0002, Miaoqiang Su, Pan Bian, Xirong Li 0001, Wenchang Shi |
IJCAI | 5 |
| 2018 | Feature Re-Learning with Data Augmentation for Content-based Video RecommendationabstractThis paper describes our solution for the Hulu Content-based Video Relevance Prediction Challenge. Noting the deficiency of the original features, we propose feature re-learning to improve video relevance prediction. To generate more training instances for supervised learning, we develop two data augmentation strategies, one for frame-level features and the other for video-level features. In addition, late fusion of multiple models is employed to further boost the performance. Evaluation conducted by the organizers shows that our best run outperforms the Hulu baseline, obtaining relative improvements of 26.2% and 30.2% on the TV-shows track and the Movies track, respectively, in terms of [email protected] The results clearly justify the effectiveness of the proposed solution. Jianfeng Dong, Xirong Li 0001, Chaoxi Xu, Gang Yang 0001, Xun Wang 0007 |
ACM Multimedia | 2 |
| 2018 | Dissimilarity Representation Learning for Generalized Zero-Shot RecognitionabstractGeneralized zero-shot learning (GZSL) aims to recognize any test instance coming either from a known class or from a novel class that has no training instance. To synthesize training instances for novel classes and thus resolving GZSL as a common classification problem, we propose a Dissimilarity Representation Learning (DSS) method. Dissimilarity representation is to represent a specific instance in terms of its (dis)similarity to other instances in a visual or attribute based feature space. In the dissimilarity space, instances of the novel classes are synthesized by an end-to-end optimized neural network. The neural network realizes two-level feature mappings and domain adaptions in the dissimilarity space and the attribute based feature space. Experimental results on five benchmark datasets, i.e., AWA, AWA$_2$, SUN, CUB, and aPY, show that the proposed method improves the state-of-the-art with a large margin, approximately 10% gain in terms of the harmonic mean of the top-1 accuracy. Consequently, this paper establishes a new baseline for GZSL. Gang Yang 0001, Jinlu Liu, Jieping Xu, Xirong Li 0001 |
ACM Multimedia | 4 |
| 2018 | Imagination Based Sample Construction for Zero-Shot LearningabstractZero-shot learning (ZSL) which aims to recognize unseen classes with no labeled training sample, efficiently tackles the problem of missing labeled data in image retrieval. Nowadays there are mainly two types of popular methods for ZSL to recognize images of unseen classes: probabilistic reasoning and feature projection. Different from these existing types of methods, we propose a new method: sample construction to deal with the problem of ZSL. Our proposed method, called Imagination Based Sample Construction (IBSC), innovatively constructs image samples of target classes in feature space by mimicking human associative cognition process. Based on an association between attribute and feature, target samples are constructed from different parts of various samples. Furthermore, dissimilarity representation is employed to select high-quality constructed samples which are used as labeled data to train a specific classifier for those unseen classes. In this way, zero-shot learning is turned into a supervised learning problem. As far as we know, it is the first work to construct samples for ZSL thus, our work is viewed as a baseline for future sample construction methods. Experiments on four benchmark datasets show the superiority of our proposed method. Gang Yang 0001, Jinlu Liu, Xirong Li 0001 |
SIGIR | 3 |
| 2018 | Predicting Visual Features From Text for Image and Video Caption RetrievalabstractThis paper strives to find amidst a set of sentences the one best describing the content of a given image or video. Different from existing works, which rely on a joint subspace for their image and video caption retrieval, we propose to do so in a visual space exclusively. Apart from this conceptual novelty, we contribute Word2VisualVec , a deep neural network architecture that learns to predict a visual feature representation from textual input. Example captions are encoded into a textual embedding based on multiscale sentence vectorization and further transferred into a deep visual feature of choice via a simple multilayer perceptron. We further generalize Word2VisualVec for video caption retrieval, by predicting from text both three-dimensional convolutional neural network features as well as a visual-audio representation. Experiments on Flickr8k, Flickr30k, the Microsoft Video Description dataset, and the very recent NIST TrecVid challenge for video caption retrieval detail Word2VisualVec's properties, its benefit over textual embeddings, the potential for multimodal query composition, and its state-of-the-art results. Jianfeng Dong, Xirong Li 0001, Cees Snoek |
IEEE Trans. Multim. | 2 |
| 2018 | Cross-Media Similarity Evaluation for Web Image Retrieval in the WildabstractIn order to retrieve unlabeled images by textual queries, cross-media similarity computation is a key ingredient. Although novel methods are continuously introduced, little has been done to evaluate these methods together with large-scale query log analysis. Consequently, how far have these methods brought us in answering real-user queries is unclear. Given baseline methods that use relatively simple text/image matching, how much progress have advanced models made is also unclear. This paper takes a pragmatic approach to answering the two questions. Queries are automatically categorized according to the proposed query visualness measure and later connected to the evaluation of multiple cross-media similarity models on three test sets. Such a connection reveals that the success of the state of the art is mainly attributed to their good performance on visual-oriented queries, which account for only a small part of real-user queries. To quantify the current progress, we propose a simple text2image method, representing a novel query by a set of images selected from large-scale query log. Consequently, computing cross-media similarity between the query and a given image boils down to comparing the visual similarity between the given image and the selected images. Image retrieval experiments on the challenging Clickture dataset show that the proposed text2image is a strong baseline, comparing favorably to recent deep learning alternatives. Jianfeng Dong, Xirong Li 0001, Duanqing Xu |
IEEE Trans. Multim. | 2 |
| 2017 | Fluency-Guided Cross-Lingual Image CaptioningabstractImage captioning has so far been explored mostly in English, as most available datasets are in this language. However, the application of image captioning should not be restricted by language. Only few studies have been conducted for image captioning in a cross-lingual setting. Different from these works that manually build a dataset for a target language, we aim to learn a cross-lingual captioning model fully from machine-translated sentences. To conquer the lack of fluency in the translated sentences, we propose in this paper a fluency-guided learning framework. The framework comprises a module to automatically estimate the fluency of the sentences and another module to utilize the estimated fluency scores to effectively train an image captioning model for the target language. As experiments on two bilingual (English-Chinese) datasets show, our approach improves both fluency and relevance of the generated captions in Chinese, but without using any manually written sentences from the target language. Weiyu Lan, Xirong Li 0001, Jianfeng Dong |
ACM Multimedia | 2 |
| 2017 | Tag relevance fusion for social image retrieval
Xirong Li 0001 |
Multim. Syst. | 1 |
| 2016 | Adding Chinese Captions to ImagesabstractThis paper extends research on automated image captioning in the dimension of language, studying how to generate Chinese sentence descriptions for unlabeled images. To evaluate image captioning in this novel context, we present Flickr8k-CN, a bilingual extension of the popular Flickr8k set. The new multimedia dataset can be used to quantitatively assess the performance of Chinese captioning and English-Chinese machine translation. The possibility of re-using existing English data and models via machine translation is investigated. Our study reveals to some extent that a computer can master two distinct languages, English and Chinese, at a similar level for describing the visual world. Data is publicly available at http://tinyurl.com/flickr8kcn Xirong Li 0001, Weiyu Lan, Jianfeng Dong |
ICMR | 1 |
| 2016 | Early Embedding and Late Reranking for Video CaptioningabstractThis paper describes our solution for the MSR Video to Language Challenge. We start from the popular ConvNet + LSTM model, which we extend with two novel modules. One is early embedding, which enriches the current low-level input to LSTM by tag embeddings. The other is late reranking, for re-scoring generated sentences in terms of their relevance to a specific video. The modules are inspired by recent works on image captioning, repurposed and redesigned for video. As experiments on the MSR-VTT validation set show, the joint use of these two modules add a clear improvement over a non-trivial ConvNet + LSTM baseline under four performance metrics. The viability of the proposed solution is further confirmed by the blind test by the organizers. Our system is ranked at the 4th place in terms of overall performance, while scoring the best CIDEr-D, which measures the human-likeness of generated captions. Jianfeng Dong, Xirong Li 0001, Weiyu Lan, Cees Snoek |
ACM Multimedia | 2 |
| 2016 | Detecting Violence in Video using SubclassesabstractThis paper attacks the challenging problem of violence detection in videos. Different from existing works focusing on combining multi-modal features, we go one step further by adding and exploiting subclasses visually related to violence. We enrich the MediaEval 2015 violence dataset by manually labeling violence videos with respect to the subclasses. Such fine-grained annotations not only help understand what have impeded previous efforts on learning to fuse the multi-modal features, but also enhance the generalization ability of the learned fusion to novel test data. The new subclass based solution, with AP of 0.303 and P100 of 0.55 on the MediaEval 2015 test set, outperforms the state-of-the-art. Notice that our solution does not require fine-grained annotations on the test set, so it can be directly applied on novel and fully unlabeled videos. Interestingly, our study shows that motion related features (MBH, HOG and HOF), though being essential part in previous systems, are seemingly dispensable. Data is available at http://lixirong.net/datasets/mm2016vsd Xirong Li 0001, Qin Jin, Jieping Xu |
ACM Multimedia | 1 |
| 2016 | TagBook: A Semantic Video Representation Without Supervision for Event DetectionabstractWe consider the problem of event detection in video for scenarios where only a few, or even zero, examples are available for training. For this challenging setting, the prevailing solutions in the literature rely on a semantic video representation obtained from thousands of pretrained concept detectors. Different from existing work, we propose a new semantic video representation that is based on freely available social tagged videos only, without the need for training any intermediate concept detectors. We introduce a simple algorithm that propagates tags from a video's nearest neighbors, similar in spirit to the ones used for image retrieval, but redesign it for video event detection by including video source set refinement and varying the video tag assignment. We call our approach TagBook and study its construction, descriptiveness, and detection performance on the TRECVID 2013 and 2014 multimedia event detection datasets and the Columbia Consumer Video dataset. Despite its simple nature, the proposed TagBook video representation is remarkably effective for few-example and zero-example event detection, even outperforming very recent state-of-the-art alternatives building on supervised representations. Masoud Mazloom, Xirong Li 0001, Cees Snoek |
IEEE Trans. Multim. | 2 |
| 2015 | Detecting semantic concepts in consumer videos using audioabstractWith the increasing use of audio sensors in user generated content collection, how to detect semantic concepts using audio streams has become an important research problem. In this paper, we present a semantic concept annotation system using soundtracks/ audio of the video. We investigate three different acoustic feature representations for audio semantic concept annotation and explore fusion of audio annotation with visual annotation systems. We test our system on the data collection from HUAWEI Accurate and Fast Mobile Video Annotation Grand Challenge 2014. The experimental results show that our audio-only concept annotation system can detect semantic concepts significantly better than random guess. It can also provide significant complementary information to the visual-based concept annotation system for performance boost. Further detailed analysis shows that for interpreting a semantic concept both visually and acoustically, it is better to train concept models for the visual system and audio system using visual-driven and audio-driven ground truth separately. Junwei Liang 0001, Qin Jin, Xixi He, Gang Yang 0001, Jieping Xu, Xirong Li 0001 |
ICASSP | 6 |
| 2015 | Semantic Concept Annotation For User Generated Videos Using SoundtracksabstractWith the increasing use of audio sensors in user generated content (UGC) collections, semantic concept annotation from video soundtracks has become an important research problem. In this paper, we investigate reducing the semantic gap of the traditional data-driven bag-of-audio-words based audio annotation approach by utilizing the large-amount of wild audio data and their rich user tags, from which we propose a new feature representation based on semantic class model distance. We conduct experiments on the data collection from HUAWEI Accurate and Fast Mobile Video Annotation Grand Challenge 2014. We also fuse the audio-only annotation system with a visual-only system. The experimental results show that our audio-only concept annotation system can detect semantic concepts significantly better than does random guessing. The new feature representation achieves comparable annotation performance with the bag-of-audio-words feature. In addition, it can provide more semantic interpretation in the output. The experimental results also prove that the audio-only system can provide significant complementary information to the visual-only concept annotation system for performance boost and for better interpretation of semantic concepts both visually and acoustically. Qin Jin, Junwei Liang 0001, Xixi He, Gang Yang 0001, Jieping Xu, Xirong Li 0001 |
ICMR | 6 |
| 2015 | Music Positioning and Annotation For Television VideosabstractThis paper proposed a framework to assist highlighting and annotating music utilization situation automatically within videos, further to supervise and protect music copyrights. Nowadays, music copyrighters pay attention to their rights increasingly, thus music embedded in TV channel videos should be validated to avoid infringing. Our framework supports Music Copyrighter Society of China(MCSC) to do statistic works to protect the copyright owners. In our framework, through AV separation, feature extractor, classification and assemblage functions, music positioning could be confirmed effectively. Then applying music fingerprint retrieving, music could be annotated automatically with high accuracy. Moreover, our framework is a closed loop self-adaptation system as it can be re-trained regularly to expand annotation database and enhance classifier's efficiency. The system based on our framework has been implemented in MCSC and its effectiveness has been evaluated in a real-life scenario. The results, on experiments of the real-life TV stations and comparisons of former works, show that the music positioning and annotation completed automatically by our system have significant improvement about over 30 times enhancement on the working efficiency. Gang Yang 0001, Jieping Xu, Xirong Li 0001 |
ICMR | 3 |
| 2015 | Image Retrieval by Cross-Media Relevance FusionabstractHow to estimate cross-media relevance between a given query and an unlabeled image is a key question in the MSR-Bing Image Retrieval Challenge. We answer the question by proposing cross-media relevance fusion, a conceptually simple framework that exploits the power of individual methods for cross-media relevance estimation. Four base cross-media relevance functions are investigated, and later combined by weights optimized on the development set. With DCG25 of 0.5200 on the test dataset, the proposed image retrieval system secures the first place in the evaluation. Jianfeng Dong, Xirong Li 0001, Shuai Liao, Jieping Xu, Duanqing Xu, Xiaoyong Du 0001 |
ACM Multimedia | 2 |
| 2015 | Image Tag Assignment, Refinement and RetrievalabstractThis tutorial focuses on challenges and solutions for content-based image annotation and retrieval in the context of online image sharing and tagging. We present a unified review on three closely linked problems, i.e., tag assignment, tag refinement, and tag-based image retrieval. We introduce a taxonomy to structure the growing literature, understand the ingredients of the main works, clarify their connections and difference, and recognize their merits and limitations. Moreover, we present an open-source testbed, with training sets of varying sizes and three test datasets, to evaluate methods of varied learning complexity. A selected set of eleven representative works have been implemented and evaluated. During the tutorial we provide a practice session for hands on experience with the methods, software and datasets. For repeatable experiments all data and code are online at http://www.micc.unifi.it/tagsurvey Xirong Li 0001, Tiberio Uricchio, Lamberto Ballan, Marco Bertini 0001, Cees Snoek, Alberto Del Bimbo |
ACM Multimedia | 1 |
| 2015 | Zero-shot Image Tagging by Hierarchical Semantic EmbeddingabstractGiven the difficulty of acquiring labeled examples for many fine-grained visual classes, there is an increasing interest in zero-shot image tagging, aiming to tag images with novel labels that have no training examples present. Using a semantic space trained by a neural language model, the current state-of-the-art embeds both images and labels into the space, wherein cross-media similarity is computed. However, for labels of relatively low occurrence, its similarity to images and other labels can be unreliable. This paper proposes Hierarchical Semantic Embedding (HierSE), a simple model that exploits the WordNet hierarchy to improve label embedding and consequently image embedding. Moreover, we identify two good tricks, namely training the neural language model using Flickr tags instead of web documents, and using partial match instead of full match for vectorizing a WordNet node. All this lets us outperform the state-of-the-art. On a test set of over 1,500 visual object classes and 1.3 million images, the proposed model beats the current best results (18.3% versus 9.4% in hit@1). Xirong Li 0001, Shuai Liao, Weiyu Lan, Xiaoyong Du 0001, Gang Yang 0001 |
SIGIR | 1 |
| 2015 | Best practices for learning video concept detectors from social media examples
Svetlana Kordumova, Xirong Li 0001, Cees Snoek |
Multim. Tools Appl. | 2 |
| 2015 | Tag Features for Geo-Aware Image ClassificationabstractThe use of geo tags in recording the location at which a picture was taken is becoming part of image metadata . Therefore , studying approaches to image classification that can favorably exploit both geo tags and the underlying geo context has become an emerging topic. This paper contributes to geo-aware image classification by studying how to encode geo information into image representation. Given a geo-tagged image, we propose to extract geo-aware tag features by tag propagation from the geo and visual neighbors of the given image. Depending on what neighbors are used and how they are weighted, we present and compare eight variants of geo-aware tag features. Using millions of Flickr images as source data for tag feature extraction, experiments on a popular benchmark set justify the effectiveness and robustness of the proposed tag features for geo-aware image classification. Shuai Liao, Xirong Li 0001, Heng Tao Shen, Yang Yang 0002, Xiaoyong Du 0001 |
IEEE Trans. Multim. | 2 |
| 2014 | Structure Perturbation Optimization for Hopfield-Type Neural Networks
Gang Yang 0001, Xirong Li 0001, Jieping Xu, Qin Jin |
ICANN | 2 |
| 2014 | Building geo-aware tag features for image classificationabstractGiven the proliferation of geo-tagged images, geo-aware image classification is an emerging topic. To derive a better image representation, tag features which represents an image as a histogram of tags are recently introduced. However, it is unclear whether geo tags can improve the tag features. To resolve the uncertainty, this paper studies geo-aware tag features. Our work is based on previous work which builds tag features by propagating tags from visual neighbors retrieved from many user-tagged images. What is different is that we build tag features by tag propagation from the union of visual and geo neighbors. This simple modification makes the new tag feature both content-aware and geo-aware. Using 1M Flickr images as a source set to construct the tag feature, experiments on the public NUS-WIDE set justify our proposal. The geo-aware tag feature outperforms the previous tag feature and a standard bag of visual words feature. Our geo-aware image classification system beats a recent alternative. For its simplicity and effectiveness, we consider the proposed tag feature promising for geo-aware image classification. Shuai Liao, Xirong Li 0001, Xiaoyong Du 0001 |
ICME | 2 |
| 2014 | Few-Example Video Event Retrieval using Tag PropagationabstractAn emerging topic in multimedia retrieval is to detect a complex event in video using only a handful of video examples. Different from existing work which learns a ranker from positive video examples and hundreds of negative examples, we aim to query web video for events using zero or only a few visual examples. To that end, we propose in this paper a tag-based video retrieval system which propagates tags from a tagged video source to an unlabeled video collection without the need of any training examples. Our algorithm is based on weighted frequency neighbor voting using concept vector similarity. Once tags are propagated to unlabeled video we can rely on off-the-shelf language models to rank these videos by the tag similarity. We study the behavior of our tag-based video event retrieval system by performing three experiments on web videos from the TRECVID multimedia event detection corpus, with zero, one and multiple query examples that beats a recent alternative. Masoud Mazloom, Xirong Li 0001, Cees Snoek |
ICMR | 2 |
| 2014 | Source Separation Improves Music Emotion RecognitionabstractDespite the impressive progress in music emotion recognition, it remains unclear what aspect of a song, i.e., singing voice and accompanied music, carries more emotional information. As an initial attempt to answer the question, we introduce source separation into a standard music emotion recognition system. This allows us to compare systems with and without source separation, and consequently reveal the influence of singing voice and accompanied music on emotion recognition. Classification experiments on a set of 267 songs with last.fm annotations verify the new finding that source separation improves song music emotion recognition. Jieping Xu, Xirong Li 0001, Gang Yang 0001 |
ICMR | 2 |
| 2014 | A guided Hopfield evolutionary algorithm with local search for maximum clique problemabstractIn this paper, a novel hybrid evolutionary algorithm combining a Hopfield net and a local search strategy is proposed to solve maximum clique problem. The algorithm makes full use of powerful searching capability of Hopfield net and probabilistic statistic feature of estimation of distribution algorithm to produce wider search in global solution domain. In particular, a possible extension way correlated with local search optimization is introduced to affect the mutation probability thus to produce guided evolution. Experiments on the popular DIMACS benchmark demonstrate that the hybrid evolutionary algorithm produces comparable and better results than other compared algorithms, including EA/G which is a state-of-the-art algorithm in the field of evolutionary computation. Gang Yang 0001, Xirong Li 0001, Jieping Xu, Qin Jin |
SMC | 2 |
| 2013 | SCH-EGA: An Efficient Hybrid Algorithm for the Frequency Assignment Problem
Shaohui Wu, Gang Yang 0001, Jieping Xu, Xirong Li 0001 |
EANN (1) | 4 |
| 2013 | A Novel Hybrid SCH-ABC Approach for the Frequency Assignment Problem
Gang Yang 0001, Shaohui Wu, Jieping Xu, Xirong Li 0001 |
ICONIP (2) | 4 |
| 2013 | Classifying tag relevance with relevant positive and negative examplesabstractImage tag relevance estimation aims to automatically determine what people label about images is factually present in the pictorial content. Different from previous works, which either use only positive examples of a given tag or use positive and random negative examples, we argue the importance of relevant positive and relevant negative examples for tag relevance estimation. We propose a system that selects positive and negative examples, deemed most relevant with respect to the given tag from crowd-annotated images. While applying models for many tags could be cumbersome, our system trains efficient ensembles of Support Vector Machines per tag, enabling fast classification. Experiments on two benchmark sets show that the proposed system compares favorably against five present day methods. Given extracted visual features, for each image our system can process up to 3,787 tags per second. The new system is both effective and efficient for tag relevance estimation. Xirong Li 0001, Cees Snoek |
ACM Multimedia | 1 |
| 2013 | Bootstrapping Visual Categorization With Relevant NegativesabstractLearning classifiers for many visual concepts are important for image categorization and retrieval. As a classifier tends to misclassify negative examples which are visually similar to positive ones, inclusion of such misclassified and thus relevant negatives should be stressed during learning. User-tagged images are abundant online, but which images are the relevant negatives remains unclear. Sampling negatives at random is the de facto standard in the literature. In this paper, we go beyond random sampling by proposing Negative Bootstrap. Given a visual concept and a few positive examples, the new algorithm iteratively finds relevant negatives. Per iteration, we learn from a small proportion of many user-tagged images, yielding an ensemble of meta classifiers. For efficient classification, we introduce Model Compression such that the classification time is independent of the ensemble size. Compared with the state of the art, we obtain relative gains of 14% and 18% on two present-day benchmarks in terms of mean average precision. For concept search in one million images, model compression reduces the search time from over 20 h to approximately 6 min. The effectiveness and efficiency, without the need of manually labeling any negatives, make negative bootstrap appealing for learning better visual concept classifiers. Xirong Li 0001, Cees Snoek, Marcel Worring, Dennis C. Koelma, Arnold W. M. Smeulders |
IEEE Trans. Multim. | 1 |
| 2012 | Fusing concept detection and geo context for visual searchabstractGiven the proliferation of geo-tagged images, the question of how to exploit geo tags and the underlying geo context for visual search is emerging. Based on the observation that the importance of geo context varies over concepts, we propose a concept-based image search engine which fuses visual concept detection and geo context in a concept-dependent manner. Compared to individual content-based and geo-based concept detectors and their uniform combination, concept-dependent fusion shows improvements. Moreover, since the proposed search engine is trained on social-tagged images alone without the need of human interaction, it is flexible to cope with many concepts. Search experiments on 101 popular visual concepts justify the viability of the proposed solution. In particular, for 79 out of the 101 concepts, the learned weights yield improvements over the uniform weights, with a relative gain of at least 5% in terms of average precision. Xirong Li 0001, Cees Snoek, Marcel Worring, Arnold W. M. Smeulders |
ICMR | 1 |
| 2012 | Harvesting Social Images for Bi-Concept SearchabstractSearching for the co-occurrence of two visual concepts in unlabeled images is an important step towards answering complex user queries. Traditional visual search methods use combinations of the confidence scores of individual concept detectors to tackle such queries. In this paper we introduce the notion of bi-concepts, a new concept-based retrieval method that is directly learned from social-tagged images. As the number of potential bi-concepts is gigantic, manually collecting training examples is infeasible. Instead, we propose a multimedia framework to collect de-noised positive as well as informative negative training examples from the social web, to learn bi-concept detectors from these examples, and to apply them in a search engine for retrieving bi-concepts in unlabeled images. We study the behavior of our bi-concept search engine using 1.2 M social-tagged images as a data source. Our experiments indicate that harvesting examples for bi-concepts differs from traditional single-concept methods, yet the examples can be collected with high accuracy using a multi-modal approach. We find that directly learning bi-concepts is better than oracle linear fusion of single-concept detectors, with a relative improvement of 100%. This study reveals the potential of learning high-order semantics from social images, for free, suggesting promising new lines of research. Xirong Li 0001, Cees Snoek, Marcel Worring, Arnold W. M. Smeulders |
IEEE Trans. Multim. | 1 |
| 2011 | Social negative bootstrapping for visual categorizationabstractTo learn classifiers for many visual categories, obtaining labeled training examples in an efficient way is crucial. Since a classifier tends to misclassify negative examples which are visually similar to positive examples, inclusion of such informative negatives should be stressed in the learning process. However, they are unlikely to be hit by random sampling, the de facto standard in literature. In this paper, we go beyond random sampling by introducing a novel social negative bootstrapping approach. Given a visual category and a few positive examples, the proposed approach adaptively and iteratively harvests informative negatives from a large amount of social-tagged images. To label negative examples without human interaction, we design an effective virtual labeling procedure based on simple tag reasoning. Virtual labeling, in combination with adaptive sampling, enables us to select the most misclassified negatives as the informative samples. Learning from the positive set and the informative negative sets results in visual classifiers with higher accuracy. Experiments on two present-day image benchmarks employing 650K virtually labeled negative examples show the viability of the proposed approach. On a popular visual categorization benchmark our precision at 20 increases by 34%, compared to baselines trained on randomly sampled negatives. We achieve more accurate visual categorization without the need of manually labeling any negatives. Xirong Li 0001, Cees Snoek, Marcel Worring, Arnold W. M. Smeulders |
ICMR | 1 |
| 2011 | Image search 2.0abstractThis abstract sketches my PhD research towards establishing a generic mechanism for exploiting social intelligence for next-generation image search. Xirong Li 0001 |
ACM Multimedia | 1 |
| 2011 | Personalizing automated image annotation using cross-entropyabstractAnnotating the increasing amounts of user-contributed images in a personalized manner is in great demand. However, this demand is largely ignored by the mainstream of automated image annotation research. In this paper we aim for personalizing automated image annotation by jointly exploiting personalized tag statistics and content-based image annotation. We propose a cross-entropy based learning algorithm which personalizes a generic annotation model by learning from a user's multimedia tagging history. Using cross-entropy-minimization based Monte Carlo sampling, the proposed algorithm optimizes the personalization process in terms of a performance measurement which can be flexibly chosen. Automatic image annotation experiments with 5,315 realistic users in the social web show that the proposed method compares favorably to a generic image annotation method and a method using personalized tag statistics only. For 4,442 users the performance improves, where for 1,088 users the absolute performance gain is at least 0.05 in terms of average precision. The results show the value of the proposed method. Xirong Li 0001, Efstratios Gavves, Cees Snoek, Marcel Worring, Arnold W. M. Smeulders |
ACM Multimedia | 1 |
| 2009 | Annotating images by harnessing worldwide user-tagged photosabstractAutomatic image tagging is important yet challenging due to the semantic gap and the lack of learning examples to model a tag's visual diversity. Meanwhile, social user tagging is creating rich multimedia content on the Web. In this paper, we propose to combine the two tagging approaches in a search-based framework. For an unlabeled image, we first retrieve its visual neighbors from a large user-tagged image database. We then select relevant tags from the result images to annotate the unlabeled image. To tackle the unreliability and sparsity of user tagging, we introduce a joint-modality tag relevance estimation method which efficiently addresses both textual and visual clues. Experiments on 1.5 million Flickr photos and 10 000 Corel images verify the proposed method. Xirong Li 0001, Cees Snoek, Marcel Worring |
ICASSP | 1 |
| 2009 | Visual categorization with negative examples for freeabstractAutomatic visual categorization is critically dependent on labeled examples for supervised learning. As an alternative to traditional expert labeling, social-tagged multimedia is becoming a novel yet subjective and inaccurate source of learning examples. Different from existing work focusing on collecting positive examples, we study in this paper the potential of substituting social tagging for expert labeling for creating negative examples. We present an empirical study using 6.5 million Flickr photos as a source of social tagging. Our experiments on the PASCAL VOC challenge 2008 show that with a relative loss of only 4.3% in terms of mean average precision, expert-labeled negative examples can be completely replaced by social-tagged negative examples for consumer photo categorization. Xirong Li 0001, Cees Snoek |
ACM Multimedia | 1 |
| 2009 | Query representation by structured concept threads with application to interactive video retrieval
Dong Wang 0022, Jianmin Li 0001, Bo Zhang 0010, Xirong Li 0001 |
J. Vis. Commun. Image Represent. | 5 |
| 2009 | Learning Social Tag Relevance by Neighbor VotingabstractSocial image analysis and retrieval is important for helping people organize and access the increasing amount of user tagged multimedia. Since user tagging is known to be uncontrolled, ambiguous, and overly personalized, a fundamental problem is how to interpret the relevance of a user-contributed tag with respect to the visual content the tag is describing. Intuitively, if different persons label visually similar images using the same tags, these tags are likely to reflect objective aspects of the visual content. Starting from this intuition, we propose in this paper a neighbor voting algorithm which accurately and efficiently learns tag relevance by accumulating votes from visual neighbors. Under a set of well-defined and realistic assumptions, we prove that our algorithm is a good tag relevance measurement for both image ranking and tag ranking. Three experiments on 3.5 million Flickr photos demonstrate the general applicability of our algorithm in both social image retrieval and image tag suggestion. Our tag relevance learning algorithm substantially improves upon baselines for all the experiments. The results suggest that the proposed algorithm is promising for real-world applications. Xirong Li 0001, Cees Snoek, Marcel Worring |
IEEE Trans. Multim. | 1 |
| 2008 | Annotating Images by Mining Image Search ResultsabstractAlthough it has been studied for years by the computer vision and machine learning communities, image annotation is still far from practical. In this paper, we propose a novel attempt at model-free image annotation, which is a data-driven approach that annotates images by mining their search results. Some 2.4 million images with their surrounding text are collected from a few photo forums to support this approach. The entire process is formulated in a divide-and-conquer framework where a query keyword is provided along with the uncaptioned image to improve both the effectiveness and efficiency. This is helpful when the collected data set is not dense everywhere. In this sense, our approach contains three steps: 1) the search process to discover visually and semantically similar search results, 2) the mining process to identify salient terms from textual descriptions of the search results, and 3) the annotation rejection process to filter out noisy terms yielded by Step 2. To ensure real-time annotation, two key techniques are leveraged-one is to map the high-dimensional image visual features into hash codes, the other is to implement it as a distributed system, of which the search and mining processes are provided as Web services. As a typical result, the entire process finishes in less than 1 second. Since no training data set is required, our approach enables annotating with unlimited vocabulary and is highly scalable and robust to outliers. Experimental results on both real Web images and a benchmark image data set show the effectiveness and efficiency of the proposed algorithm. It is also worth noting that, although the entire approach is illustrated within the divide-and conquer framework, a query keyword is not crucial to our current implementation. We provide experimental results to prove this. Xin-Jing Wang, Lei Zhang 0001, Xirong Li 0001, Wei-Ying Ma |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2007 | SBIA: search-based image annotation by leveraging web-scale imagesabstractIn this technical demonstration, we showcase the SBIA system - a search-based image annotation system. At the heart of the system lies a very large-scale image search engine which indexed three million Web images and supports both text and visual queries. Given an image (with initial annotations), SBIA first finds semantically/visually similar images via the search engine, and then mines representative keywords from the retrieved images. These keywords, after annotation rejection and relevance ranking, are finally used to annotate the query image. Xirong Li 0001, Xin-Jing Wang, Changhu Wang, Lei Zhang 0001 |
ACM Multimedia | 1 |
| 2007 | The importance of query-concept-mapping for automatic video retrievalabstractA new video retrieval paradigm of query-by-concept emerges recently. However, it remains unclear how to exploit the detected concepts in retrieval given a multimedia query. In this paper, we point out that it is important to map the query to a few relevant concepts instead of search with all concepts. In addition, we show that solving this problem through both text and image inputs are effective for search, and it is possible to determine the number of related concepts by a language modeling approach. Experimental evidence is obtained on the automatic search task of TRECVID 2006 using a large lexicon of 311 learned semantic concept detectors. Dong Wang 0022, Xirong Li 0001, Jianmin Li 0001, Bo Zhang 0010 |
ACM Multimedia | 2 |
| 2006 | Image annotation by large-scale content-based image retrievalabstractImage annotation has been an active research topic in recent years due to its potentially large impact on both image understanding and Web image search. In this paper, we target at solving the automatic image annotation problem in a novel search and mining framework. Given an uncaptioned image, first in the search stage, we perform content-based image retrieval (CBIR) facilitated by high-dimensional indexing to find a set of visually similar images from a large-scale image database. The database consists of images crawled from the World Wide Web with rich annotations, e.g. titles and surrounding text. Then in the mining stage, a search result clustering technique is utilized to find most representative keywords from the annotations of the retrieved image subset. These keywords, after salience ranking, are finally used to annotate the uncaptioned image. Based on search technologies, this framework does not impose an explicit training stage, but efficiently leverages large-scale and well-annotated images, and is potentially capable of dealing with unlimited vocabulary. Based on 2.4 million real Web images, comprehensive evaluation of image annotation on Corel and U. Washington image databases show the effectiveness and efficiency of the proposed approach. Xirong Li 0001, Lei Zhang 0001, Fuzong Lin, Wei-Ying Ma |
ACM Multimedia | 1 |