Jin Yuan 0002

dblp:98/609-2 · DBLP profile ↗
← Back
41ranked-venue papers
8as first author
23since 2021 · last 2026
0000-0002-9600-7789ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 29 · 6 first-author · 15 since 2021Artificial intelligence and machine learning · 18 · 2 first-author · 12 since 2021Computer networks · 4 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 ViMAEdit: Vision-Guided and Mask-Enhanced Adaptive Editing Algorithm for Prompt-Based Image Editing
Kejie Wang, Xuemeng Song, Meng Liu 0006, Jin Yuan 0002, Weili Guan
IEEE Trans. Circuits Syst. Video Technol.4
2026 Scene Graph-Guided SegCaptioning Transformer With Fine-Grained Alignment for Controllable Video Segmentation and Captioning
abstract
Recent advancements in multimodal large models have significantly bridged the representation gap between diverse modalities, catalyzing the evolution of video multimodal interpretation, which enhances users' understanding of video content by generating correlated modalities. However, most existing video multimodal interpretation methods primarily concentrate on global comprehension with limited user interaction. To address this, we propose a novel task, Controllable Video Segmentation and Captioning (SegCaptioning), which empowers users to provide specific prompts, such as a bounding box around an object of interest, to simultaneously generate correlated masks and captions that precisely embody user intent. An innovative framework, Scene Graph-guided Fine-grained SegCaptioning Transformer (SG-FSCFormer), is designed to integrate a Prompt-guided Temporal Graph Former to effectively capture and represent user intent through an adaptive prompt adaptor, ensuring that the generated content aligns well with the user's requirements. Furthermore, our model introduces a Fine-grained Mask-linguistic Decoder to collaboratively predict high-quality caption-mask pairs using a Multi-entity Contrastive loss, while providing fine-grained alignment between each mask and its corresponding caption tokens, thereby enhancing the user's comprehension of videos. Comprehensive experiments conducted on two benchmark datasets demonstrate that SG-FSCFormer achieves remarkable performance, effectively capturing user intent and generating precise multimodal outputs tailored to user specifications. Our code is available at https://github.com/XuZhang1211/SG-FSCFormer.
Xu Zhang 0025, Jin Yuan 0002, BinHong Yang, Xuan Liu 0001, Qianjun Zhang, Yuyi Wang 0001, Zhiyong Li 0001, Hanwang Zhang
IEEE Trans. Image Process.2
2025 SGDiff: Scene Graph Guided Diffusion Model for Image Collaborative SegCaptioning
abstract
Controllable image semantic understanding tasks, such as captioning or segmentation, necessitate users to input a prompt (e.g., text or bounding boxes) to predict a unique outcome, presenting challenges such as high-cost prompt input or limited information output. This paper introduces a new task ``Image Collaborative Segmentation and Captioning'' (SegCaptioning), which aims to translate a straightforward prompt, like a bounding box around an object, into diverse semantic interpretations represented by (caption, masks) pairs, allowing flexible result selection by users. This task poses significant challenges, including accurately capturing a user's intention from a minimal prompt while simultaneously predicting multiple semantically aligned caption words and masks. Technically, we propose a novel Scene Graph Guided Diffusion Model that leverages structured scene graph features for correlated mask-caption prediction. Initially, we introduce a Prompt-Centric Scene Graph Adaptor to map a user's prompt to a scene graph, effectively capturing his intention. Subsequently, we employ a diffusion process incorporating a Scene Graph Guided Bimodal Transformer to predict correlated caption-mask pairs by uncovering intricate correlations between them. To ensure accurate alignment, we design a Multi-Entities Contrastive Learning loss to explicitly align visual and textual entities by considering inter-modal similarity, resulting in well-aligned caption-mask pairs. Extensive experiments conducted on two datasets demonstrate that SGDiff achieves superior performance in SegCaptioning, yielding promising results for both captioning and segmentation tasks with minimal prompt input.
Xu Zhang 0025, Jin Yuan 0002, Hanwang Zhang, Guojin Zhong, Yongsheng Zang, Jiacheng Lin, Zhiyong Li 0001
AAAI2
2025 DECIDER: Difference-aware Contrastive Diffusion Model with Adversarial Perturbations for Image Change Captioning
abstract
Image change captioning (ICC) poses great challenges stemming from describing subtle differences between two similar images in natural language, significantly increasing the complexity of feature extraction and cross-modal learning compared to the image captioning task. Existing ICC methods often suffer from two key challenges: 1) Massive irrelevant information of uni-image features leads to suboptimal visual difference representations; 2) Imprecise inter-modality correspondence degrades the quality of generated captions. This paper proposes a Difference-aware Contrastive Diffusion Model with Adversarial Perturbations (DECIDER) for ICC due to the excellent performance of diffusion models in image/text generation. Technically, difference-aware cross-modal learning is developed to suppress irrelevant information and learn compact yet robust visual difference representations. This is achieved by optimizing a novel objective mathematically derived from the information bottleneck principle that excels in filtering redundant features and highlighting differences. Furthermore, we propose to dynamically generate ``hard'' positive and negative samples via adversarial perturbations, which are involved in contrastive diffusion training with a tighter variational bound. This design encourages our DECIDER to excavate and construct complex correspondences between visual differences and captions, thereby improving generalization performance. Extensive experiments on four datasets demonstrate that DECIDER significantly exceeds state-of-the-art performance.
Guojin Zhong, Jinhong Hu, Jin Yuan 0002
AAAI4
2025 AVAM: A Universal Training-Free Adaptive Visual Anchoring Embedded into Multimodal Large Language Model for Multi-Image Question Answering
Kang Zeng, Guojin Zhong, Jintao Cheng, Jin Yuan 0002, Zhiyong Li 0001
ICCV4
2025 Multi-Resolution Decomposable Diffusion Model for Non-Stationary Time Series Anomaly Detection
abstract
Recently, generative models have shown considerable promise in unsupervised time series anomaly detection. Nonetheless, the task of effectively capturing complex temporal patterns and minimizing false alarms becomes increasingly challenging when dealing with non-stationary time series, characterized by continuously fluctuating statistical attributes and joint distributions. To confront these challenges, we underscore the benefits of multi-resolution modeling, which improves the ability to distinguish between anomalies and non-stationary behaviors by leveraging correlations across various resolution scales. In response, we introduce a **M**ulti-Res**o**lution **De**composable Diffusion **M**odel (MODEM), which integrates a coarse-to-fine diffusion paradigm with a frequency-enhanced decomposable network to adeptly navigate the intricacies of non-stationarity. Technically, the coarse-to-fine diffusion model embeds cross-resolution correlations into the forward process to optimize diffusion transitions mathematically. It then innovatively employs low-resolution recovery to guide the reverse trajectories of high-resolution series in a coarse-to-fine manner, enhancing the model's ability to learn and elucidate underlying temporal patterns. Furthermore, the frequency-enhanced decomposable network operates in the frequency domain to extract globally shared time-invariant information and time-variant temporal dynamics for accurate series reconstruction. Extensive experiments conducted across five real-world datasets demonstrate that our proposed MODEM achieves state-of-the-art performance and can be generalized to other time series tasks.
Guojin Zhong, Pan Wang 0011, Jin Yuan 0002, Zhiyong Li 0001, Long Chen 0016
ICLR3
2025 MCT-CCDiff: Context-Aware Contrastive Diffusion Model With Mediator-Bridging Cross-Modal Transformer for Image Change Captioning
abstract
Recent advancements in diffusion models (DMs) have showcased superior capabilities in generating images and text. This paper first introduces DMs for image change captioning (ICC) and proposes a novel Context-aware Contrastive Diffusion model with Mediator-bridging Cross-modal Transformer (MCT-CCDiff) to accurately predict visual difference descriptions conditioned on two similar images. Technically, MCT-CCDiff develops a Text Embedding Contrastive Loss (TECL) that leverages both positive and negative samples to more effectively distinguish text embeddings, thus generating more discriminative text representations for ICC. To accurately predict visual difference descriptions, MCT-CCDiff introduces a Mediator-bridging Cross-modal Transformer (MCTrans) designed to efficiently explore the cross-modal correlations between visual differences and corresponding text by using a lightweight mediator, mitigating interference from visual redundancy and reducing interaction overhead. Additionally, it incorporates context-augmented denoising to further understand the contextual relationships within caption words implemented by a revised diffusion loss, which provides a tighter optimization bound, leading to enhanced optimization effects for high-quality text generation. Extensive experiments conducted on four benchmark datasets clearly demonstrate that our MCT-CCDiff significantly outperforms state-of-the-art methods in the field of ICC, marking a notable advancement in the generation of precise visual difference descriptions.
Jinhong Hu, Guojin Zhong, Jin Yuan 0002
IEEE Trans. Image Process.3
2024 CF-Deformable DETR: An End-to-End Alignment-Free Model for Weakly Aligned Visible-Infrared Object Detection
Haolong Fu, Jin Yuan 0002, Guojin Zhong, Jiacheng Lin, Zhiyong Li 0001
IJCAI2
2024 PROMOTE: Prior-Guided Diffusion Model with Global-Local Contrastive Learning for Exemplar-Based Image Translation
abstract
Exemplar-based image translation has garnered significant interest from researchers due to its broad applications in multimedia/multimodal processing. Existing methods primarily employ Euclidean-based losses to implicitly establish cross-domain correspondences between exemplar and conditional images, aiming to produce high-fidelity images. However, these methods often suffer from two challenges: 1) Insufficient excavation of domain-invariant features leads to low-quality cross-domain correspondences, and 2) Inaccurate correspondences result in errors propagated during the translation process due to a lack of reliable prior guidance. To tackle these issues, we propose a novel prior-guided diffusion model with global-local contrastive learning (PROMOTE), which is trained in a self-supervised manner. Technically, global-local contrastive learning is designed to align two cross-domain images within hyperbolic space and reduce the gap between their semantic correlation distributions using the Fisher-Rao metric, allowing the visual encoders to extract domain-invariant features more effectively. Moreover, a prior-guided diffusion model is developed that propagates the structural prior to all timesteps in the diffusion process. It is optimized by a novel prior denoising loss, mathematically derived from the transitions modified by prior information in a self-supervised manner, successfully alleviating the impact of inaccurate correspondences on image translation. Extensive experiments conducted across seven datasets demonstrate that our proposed PROMOTE significantly exceeds state-of-the-art performance in diverse exemplar-based image translation tasks. The source code is publicly available at http://github.com/zgj77/PROMOTE.
Guojin Zhong, Yihu Guo, Jin Yuan 0002, Qianjun Zhang, Weili Guan, Long Chen 0016
ACM Multimedia3
2024 Privacy preservation network with global-aware focal loss for Interactive Personal Visual Privacy Preservation
Jiacheng Lin, Haolong Fu, Yifan Li 0005, Jin Yuan 0002, Zhiyong Li 0001
Neurocomputing6
2024 PVPUFormer: Probabilistic Visual Prompt Unified Transformer for Interactive Image Segmentation
abstract
Integration of diverse visual prompts like clicks, scribbles, and boxes in interactive image segmentation significantly facilitates users' interaction as well as improves interaction efficiency. However, existing studies primarily encode the position or pixel regions of prompts without considering the contextual areas around them, resulting in insufficient prompt feedback, which is not conducive to performance acceleration. To tackle this problem, this paper proposes a simple yet effective Probabilistic Visual Prompt Unified Transformer (PVPUFormer) for interactive image segmentation, which allows users to flexibly input diverse visual prompts with the probabilistic prompt encoding and feature post-processing to excavate sufficient and robust prompt features for performance boosting. Specifically, we first propose a Probabilistic Prompt-unified Encoder (PPuE) to generate a unified one-dimensional vector by exploring both prompt and non-prompt contextual information, offering richer feedback cues to accelerate performance improvement. On this basis, we further present a Prompt-to-Pixel Contrastive (P2C) loss to accurately align both prompt and pixel features, bridging the representation gap between them to offer consistent feature representations for mask prediction. Moreover, our approach designs a Dual-cross Merging Attention (DMA) module to implement bidirectional feature interaction between image and prompt features, generating notable features for performance improvement. A comprehensive variety of experiments on several challenging datasets demonstrates that the proposed components achieve consistent improvements, yielding state-of-the-art interactive segmentation performance. Our code is available at https://github.com/XuZhang1211/PVPUFormer.
Xu Zhang 0025, Kailun Yang 0001, Jiacheng Lin, Jin Yuan 0002, Zhiyong Li 0001, Shutao Li 0001
IEEE Trans. Image Process.4
2023 Contrast-augmented Diffusion Model with Fine-grained Sequence Alignment for Markup-to-Image Generation
abstract
The recently rising markup-to-image generation poses greater challenges as compared to natural image generation, due to its low tolerance for errors as well as the complex sequence and context correlations between markup and rendered image. This paper proposes a novel model named "Contrast-augmented Diffusion Model with Fine-grained Sequence Alignment'' (FSA-CDM), which introduces contrastive positive/negative samples into the diffusion model to boost performance for markup-to-image generation. Technically, we design a fine-grained cross-modal alignment module to well explore the sequence similarity between the two modalities for learning robust feature representations. To improve the generalization ability, we propose a contrast-augmented diffusion model to explicitly explore positive and negative samples by maximizing a novel contrastive variational objective, which is mathematically inferred to provide a tighter bound for the model's optimization. Moreover, the context-aware cross attention module is developed to capture the contextual information within markup language during the denoising process, yielding better noise prediction results. Extensive experiments are conducted on four benchmark datasets from different domains, and the experimental results demonstrate the effectiveness of the proposed components in FSA-CDM, significantly exceeding state-of-the-art performance by about 2% ~ 12% DTW improvements.
Guojin Zhong, Jin Yuan 0002, Pan Wang 0011, Kailun Yang 0001, Weili Guan, Zhiyong Li 0001
ACM Multimedia2
2023 A Text-Specific Domain Adaptive Network for Scene Text Detection in the Wild
Jin Yuan 0002, Zhiyong Li 0001
Appl. Intell.2
2023 Domain adaptive multigranularity proposal network for text detection under extreme traffic scenes
Zhiyong Li 0001, Jiacheng Lin, Ke Nai, Jin Yuan 0002, Yifan Li 0005
Comput. Vis. Image Underst.5
2023 BRPPNet: Balanced privacy protection network for referring personal image privacy protection
Jiacheng Lin, Xianwen Dai, Ke Nai, Jin Yuan 0002, Zhiyong Li 0001, Xu Zhang 0025, Shutao Li 0001
Expert Syst. Appl.4
2023 DO-SA&R: Distant Object Augmented Set Abstraction and Regression for Point-Based 3D Object Detection
abstract
Point-based 3D detection approaches usually suffer from the severe point sampling imbalance problem between foreground and background. We observe that prior works have attempted to alleviate this imbalance by emphasizing foreground sampling. However, even adequate foreground sampling may be extremely unbalanced between nearby and distant objects, yielding unsatisfactory performance in detecting distant objects. To tackle this issue, this paper first proposes a novel method named Distant Object Augmented Set Abstraction and Regression (DO-SA&R) to enhance distant object detection, which is vital for the timely response of decision-making systems like autonomous driving. Technically, our approach first designs DO-SA with novel distant object augmented farthest point sampling (DO-FPS) to emphasize sampling on distant objects by leveraging both object-dependent and depth-dependent information. Then, we propose distant object augmented regression to reweight all the instance boxes for strengthening regression training on distant objects. In practice, the proposed DO-SA&R can be easily embedded into the existing modules, yielding consistent performance improvements, especially on detecting distant objects. Extensive experiments are conducted on the popular KITTI, nuScenes and Waymo datasets, and DO-SA&R demonstrates superior performance, especially for distant object detection. Our code is available at https://github.com/mikasa3lili/DO-SAR.
Jiacheng Lin, Ke Nai, Jin Yuan 0002, Zhiyong Li 0001
IEEE Trans. Image Process.5
2023 TP-FER: An Effective Three-phase Noise-tolerant Recognizer for Facial Expression Recognition
abstract
Single-label facial expression recognition (FER), which aims to classify single expression for facial images, usually suffers from the label noisy and incomplete problem, where manual annotations for partial training images exist wrong or incomplete labels, resulting in performance decline. Although prior work has attempted to leverage external sources or manual annotations to handle this problem, it usually requires extra costs. This article explores a simple yet effective three-phase paradigm (“warm-up,” “selection,” and “relabeling”) for FER task. First, the warm-up phase attempts to build an initial recognition network based on noisy samples for discriminative feature extractions and facial expression predictions. Then, the second selection phase defines several rules to choose high confident samples according to prediction scores, and the third relabeling phase assigns two potential labels to those samples for network updating according to a composite two-label loss. Compared with the previous studies, the three-phase learning could effectively correct noisy labels in the ground truth without extra information and automatically assign two potential labels to single-label samples without manual annotations. As a result, the label information is purified and supplemented with few cost, yielding significant performance improvement. Extensive experiments are conducted on three datasets, and the experimental results demonstrate that our approach is robust to noisy training samples and outperforms several state-of-the-art methods.
Jin Yuan 0002, Zhiyong Li 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2023 JDAN: Joint Detection and Association Network for Real-Time Online Multi-Object Tracking
abstract
In the last few years, enormous strides have been made for object detection and data association, which are vital subtasks for one-stage online multi-object tracking (MOT). However, the two separated submodules involved in the whole MOT pipeline are processed or optimized separately, resulting in a complex method design and requiring manual settings. In addition, few works integrate the two subtasks into a single end-to-end network to optimize the overall task. In this study, we propose an end-to-end MOT network called joint detection and association network (JDAN) that is trained and inferred in a single network. All layers in JDAN are differentiable, and can be optimized jointly to detect targets and output an association matrix for robust multi-object tracking. What’s more, we generate suitable pseudo-labels to address the data inconsistency between object detection and association. The detection and association submodules could be optimized by the composite loss function that is derived from the detection results and the generated pseudo association labels, respectively. The proposed approach is evaluated on two MOT challenge datasets, and achieves promising performance compared with classic and latest methods.
Zhiyong Li 0001, Jin Yuan 0002, Shutao Li 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2022 BTN: Neuroanatomical aligning between visual object tracking in deep neural network and smooth pursuit in brain
Zhiyong Li 0001, Ke Nai, Jin Yuan 0002, Shutao Li 0001, Xianghua Li
Neurocomputing4
2022 Discriminative Style Learning for Cross-Domain Image Captioning
abstract
The cross-domain image captioning, which is trained on a source domain and generalized to other domains, usually faces the large domain shift problem. Although prior work has attempted to leverage both paired source and unpaired target data to minimize this shift, the performance is still unsatisfactory. One main reason lies in the large discrepancy in language expression between two domains, where diverse language styles are adopted to describe an image from different views, resulting in different semantic descriptions for an image. To tackle this problem, this paper proposes a Style-based Cross-domain Image Captioner (SCIC) which incorporates the discriminative style information into the encoder-decoder framework, and interprets an image as a special sentence according to external style instructions. Technically, we design a novel "Instruction-based LSTM", which adds the instruct gate to collect a style instruction, and then outputs a specified format according to that instruction. Two objectives are designed to train I-LSTM: 1) generating correct image descriptions and 2) generating correct styles, thus the model is expected to accurately capture the semantic meanings of an image by the special caption as well as understand the syntactic structure of the caption. We use MS-COCO as the source domain, and Oxford-102, CUB-200, Flickr30k as the target domains. Experimental results demonstrate that our model consistently outperforms the previous methods, and the style information incorporating with I-LSTM significantly improves the performance, with 5% CIDEr improvements at least on all datasets.
Jin Yuan 0002, Shuai Zhu, Shuyin Huang, Hanwang Zhang, Yaoqiang Xiao, Zhiyong Li 0001, Meng Wang 0001
IEEE Trans. Image Process.1
2021 Person re-identification with part prediction alignment
Zhiyong Li 0001, Jingyi Lv, Jin Yuan 0002
Comput. Vis. Image Underst.4
2021 Siamese target estimation network with AIoU loss for real-time visual tracking
Zhiyong Li 0001, Chenming Hu, Ke Nai, Jin Yuan 0002
J. Vis. Commun. Image Represent.4
2021 History-based attention in Seq2Seq model for multi-label text classification
Yaoqiang Xiao, Jin Yuan 0002, Songrui Guo, Yi Xiao 0004, Zhiyong Li 0001
Knowl. Based Syst.3
2020 Document and Word Representations Generated by Graph Convolutional Network and BERT for Short Text Classification
abstract
In many studies, the graph convolution neural networks were used to solve different natural language processing (NLP) problems. However, few researches employ graph convolutional network for text classification, especially for short text classification. In this work, a special text graph of the short-text corpus is created, and then a short-text graph convolutional network (STGCN) is developed. Specifically, different topic models for short text are employed, and a short text short-text graph based on the word co-occurrence, document word relations, and text topic information, is developed. The word and sentence representations generated by the STGCN are considered as the classification feature. In addition, a pre-trained word vector obtained by the BERTs hidden layer is employed, which greatly improves the classification effect of our model. The experimental results show that our model outperforms the state-of-the-art models on multiple short text datasets.
Zhihao Ye, Gongyao Jiang, Zhiyong Li 0001, Jin Yuan 0002
ECAI5
2020 Person re-identification with expanded neighborhoods distance re-ranking
Jingyi Lv, Zhiyong Li 0001, Ke Nai, Jin Yuan 0002
Image Vis. Comput.5
2020 Real-time traffic sign detection and classification towards real traffic scene
Yiqiang Wu, Zhiyong Li 0001, Ke Nai, Jin Yuan 0002
Multim. Tools Appl.5
2020 Stroke classification for sketch segmentation by fine-tuning a developmental VGGNet16
Xianyi Zhu, Jin Yuan 0002, Yi Xiao 0004, Yan Zheng 0003, Zheng Qin 0001
Multim. Tools Appl.2
2020 Gated CNN: Integrating multi-scale feature layers for object detection
Jin Yuan 0002, Heng-Chang Xiong, Yi Xiao 0004, Weili Guan, Meng Wang 0001, Richang Hong, Zhiyong Li 0001
Pattern Recognit.1
2020 Image Captioning with a Joint Attention Mechanism by Visual Concept Samples
abstract
The attention mechanism has been established as an effective method for generating caption words in image captioning; it explores one noticed subregion in an image to predict a related caption word. However, even though the attention mechanism could offer accurate subregions to train a model, the learned captioner may predict wrong, especially for visual concept words, which are the most important parts to understand an image. To tackle the preceding problem, in this article we propose Visual Concept Enhanced Captioner, which employs a joint attention mechanism with visual concept samples to strengthen prediction abilities for visual concepts in image captioning. Different from traditional attention approaches that adopt one LSTM to explore one noticed subregion each time, Visual Concept Enhanced Captioner introduces multiple virtual LSTMs in parallel to simultaneously receive multiple subregions from visual concept samples. Then, the model could update parameters by jointly exploring these subregions according to a composite loss function. Technically, this joint learning is helpful in finding the common characters of a visual concept, and thus it enhances the prediction accuracy for visual concepts. Moreover, by integrating diverse visual concept samples from different domains, our model can be extended to bridge visual bias in cross-domain learning for image captioning, which saves the cost for labeling captions. Extensive experiments have been conducted on two image datasets (MSCOCO and Flickr30K), and superior results are reported when comparing to state-of-the-art approaches. It is impressive that our approach could significantly increase BLUE-1 and F1 scores, which demonstrates an accuracy improvement for visual concepts in image captioning.
Jin Yuan 0002, Songrui Guo, Yi Xiao 0004, Zhiyong Li 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2019 Person re-identification based on re-ranking with expanded k-reciprocal nearest neighbors
Jin Yuan 0002, Zhiyong Li 0001, Yiqiang Wu, Mourad Nouioua, Guoqi Xie
J. Vis. Commun. Image Represent.2
2019 Joint residual pyramid for joint image super-resolution
Yan Zheng 0003, Yi Xiao 0004, Xianyi Zhu, Jin Yuan 0002
J. Vis. Commun. Image Represent.5
2019 Multi-criteria active deep learning for image classification
Jin Yuan 0002, Xingxing Hou, Yaoqiang Xiao, Da Cao, Weili Guan, Liqiang Nie
Knowl. Based Syst.1
2018 Gradient-Guided DCNN for Inverse Halftoning and Image Expanding
Yi Xiao 0004, Chao Pan 0009, Yan Zheng 0003, Xianyi Zhu, Zheng Qin 0001, Jin Yuan 0002
ACCV (4)6
2014 Memory recall based video search: Finding videos you have seen before based on your memory
abstract
We often remember images and videos that we have seen or recorded before but cannot quite recall the exact venues or details of the contents. We typically have vague memories of the contents, which can often be expressed as a textual description and/or rough visual descriptions of the scenes. Using these vague memories, we then want to search for the corresponding videos of interest. We call this “Memory Recall based Video Search” (MRVS). To tackle this problem, we propose a video search system that permits a user to input his/her vague and incomplete query as a combination of text query, a sequence of visual queries, and/or concept queries. Here, a visual query is often in the form of a visual sketch depicting the outline of scenes within the desired video, while each corresponding concept query depicts a list of visual concepts that appears in that scene. As the query specified by users is generally approximate or incomplete, we need to develop techniques to handle this inexact and incomplete specification by also leveraging on user feedback to refine the specification. We utilize several innovative approaches to enhance the automatic search. First, we employ a visual query suggestion model to automatically suggest potential visual features to users as better queries. Second, we utilize a color similarity matrix to help compensate for inexact color specification in visual queries. Third, we leverage on the ordering of visual queries and/or concept queries to rerank the results by using a greedy algorithm. Moreover, as the query is inexact and there is likely to be only one or few possible answers, we incorporate an interactive feedback loop to permit the users to label related samples which are visually similar or semantically close to the relevant sample. Based on the labeled samples, we then propose optimization algorithms to update visual queries and concept weights to refine the search results. We conduct experiments on two large-scale video datasets: TRECVID 2010 and YouTube. The experimental results demonstrate that our proposed system is effective for MRVS tasks.
Jin Yuan 0002, Yi-Liang Zhao, Huan-Bo Luan, Meng Wang 0001, Tat-Seng Chua
ACM Trans. Multim. Comput. Commun. Appl.1
2013 Video recommendation over multiple information sources
Xiaojian Zhao, Jin Yuan 0002, Meng Wang 0001, Guangda Li, Richang Hong, Zhoujun Li 0001, Tat-Seng Chua
Multim. Syst.2
2012 Video Browser Showdown by NUS
Jin Yuan 0002, Huan-Bo Luan, Dejun Hou, Han Zhang 0010, Yantao Zheng, Zhengjun Zha, Tat-Seng Chua
MMM1
2012 On Video Recommendation over Social Network
Xiaojian Zhao, Jin Yuan 0002, Richang Hong, Meng Wang 0001, Zhoujun Li 0001, Tat-Seng Chua
MMM2
2012 Relationship strength estimation for online social networks with the study on Facebook
Xiaojian Zhao, Jin Yuan 0002, Guangda Li, Xiaoming Chen 0007, Zhoujun Li 0001
Neurocomputing2
2011 Learning concept bundles for video search with complex queries
abstract
Classifiers for primitive visual concepts like "car", "sky" have been well developed and widely used to support video search on simple queries. However, it is usually ineffective for complex queries like "one or more people at a table or desk with a computer visible", as they carry semantics far more complex and different from simply aggregating the meanings of their constituent primitive concepts. To facilitate video search of complex queries, we propose a higher-level semantic descriptor named "concept bundle", which integrates multiple primitive concepts, such as "(soccer, fighting)", "(lion, hunting, zebra)" etc, to describe the visual representation of the complex semantics. The proposed approach first automatically selects informative concept bundles. It then builds a novel concept bundle classifier based on multi-task learning by exploiting the relatedness between concept bundle and its primitive concepts. To model a complex query, it proposes an optimal selection strategy to select related primitive concepts and concept bundles by considering both their classifier performance and semantic relatedness with respect to the query. The final results are generated by fusing the individual results from these selected primitive concepts and concept bundles. Extensive experiments are conducted on two video datasets: TRECVID 2008 and YouTube datasets. The experimental results indicate that: (a) our concept bundle learning approach outperforms the state-of-the-art methods by at least 19% and 29% on TRECVID 2008 and YouTube datasets, respectively; and (b) the use of concept bundles can improve the search performance for complex queries by at least 37.5% on TRECVID 2008 and 52% on YouTube datasets.
Jin Yuan 0002, Zhengjun Zha, Yantao Zheng, Meng Wang 0001, Tat-Seng Chua
ACM Multimedia1
2011 Integrating rich information for video recommendation with multi-task rank aggregation
abstract
Video recommendation is an important approach for helping people to access interesting videos. In this paper, we propose a scheme to integrate rich information for video recommendation. We regard video recommendation as a ranking problem and generate multiple ranking lists by exploring different information sources. A multi-task rank aggregation approach is proposed to integrate the ranking lists for different users in a joint manner. Our scheme is flexible and can easily incorporate other methods by adding their generated ranking lists into our multi-task learning algorithm. We conduct experiments with 76 users and more than 10,000 videos. The results demonstrate the feasibility and effectiveness of our approach.
Xiaojian Zhao, Guangda Li, Meng Wang 0001, Jin Yuan 0002, Zhengjun Zha, Zhoujun Li 0001, Tat-Seng Chua
ACM Multimedia4
2011 Utilizing Related Samples to Enhance Interactive Concept-Based Video Search
abstract
One of the main challenges in interactive concept-based video search is the problem of insufficient relevant samples, especially for queries with complex semantics. In this paper, “related samples” are exploited to enhance interactive video search. The related samples refer to those video segments that are relevant to part of the query rather than the entire query. Compared to the relevant samples which may be rare, the related samples are usually plentiful and easy to find in search results. Generally, the related samples are visually similar and temporally neighboring to the relevant samples. Based on these two characters, we develop a visual ranking model that simultaneously exploits the relevant, related, and irrelevant samples, as well as a temporal ranking model to leverage the temporal relationship between related and relevant samples. An adaptive fusion method is then proposed to optimally explore these two ranking models to generate search results. We conduct extensive experiments on two real-world video datasets: TRECVID 2008 and YouTube datasets. As the experimental results show, our approach achieves at least 96% and 167% performance improvements against the state-of-the-art approaches on the TRECVID 2008 and YouTube datasets, respectively.
Jin Yuan 0002, Zhengjun Zha, Yantao Zheng, Meng Wang 0001, Tat-Seng Chua
IEEE Trans. Multim.1