Guangyu Gao

dblp:33/7626 · DBLP profile ↗
← Back
41ranked-venue papers
16as first author
24since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 31 · 12 first-author · 17 since 2021Artificial intelligence and machine learning · 15 · 5 first-author · 13 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 ReBac: Current-to-Past Residual Background Correction for Class-Incremental Semantic Segmentation
Guangyu Gao, Anqi Zhang 0002, Jianbo Jiao, Fangkunhan Liu, Chi Harold Liu, Yunchao Wei
Int. J. Comput. Vis.1
2026 SynerNet: Broad-to-precise CAM synergy for weakly supervised semantic segmentation
Zhonggai Wang, Guangyu Gao, Zhuoshu Li, A. K. Qin 0001
Neural Networks2
2025 CoMBO: Conflict Mitigation via Branched Optimization for Class Incremental Segmentation
abstract
Effective Class Incremental Segmentation (CIS) requires simultaneously mitigating catastrophic forgetting and ensuring sufficient plasticity to integrate new classes. The inherent conflict above often leads to a back-and-forth, which turns the objective into finding the balance between the performance of previous (old) and incremental (new) classes. To address this conflict, we introduce a novel approach, Conflict Mitigation via Branched Optimization (CoMBO). Within this approach, we present the Query Conflict Reduction module, designed to explicitly refine queries for new classes through lightweight, class-specific adapters. This module provides an additional branch for the acquisition of new classes while preserving the original queries for distillation. Moreover, we develop two strategies to further mitigate the conflict following the branched structure, i.e., the Half-Learning Half-Distillation (HDHL) over classification probabilities, and the Importance-Based Knowledge Distillation (IKD) over query features. HDHL selectively engages in learning for classification probabilities of queries that match the ground truth of new classes, while aligning unmatched ones to the corresponding old probabilities, thus ensuring retention of old knowledge while absorbing new classes via learning negative samples. Meanwhile, IKD assesses the importance of queries based on their matching degree to old classes, prioritizing the distillation of important features and allowing less critical features to evolve. Extensive experiments in Class Incremental Panoptic and Semantic Segmentation settings have demonstrated the superior performance of CoMBO. Project page: https://guangyu-ryan.github.io/CoMBO.
Anqi Zhang 0002, Guangyu Gao, Jianbo Jiao, Chi Harold Liu, Yunchao Wei
CVPR3
2025 MKSNet: Advanced Small Object Detection in Remote Sensing Imagery with Multi-Kernel and Dual Attention Mechanisms
Guangyu Gao
MMM (2)3
2025 Advertising or adversarial? AdvSign: Artistic advertising sign camouflage for target physical attacking to object detector
Guangyu Gao, Zhuocheng Lv, A. K. Qin 0001
Neural Networks1
2025 PRFormer: Matching Proposal and Reference Masks by Semantic and Spatial Similarity for Few-Shot Semantic Segmentation
abstract
Few-shot Semantic Segmentation (FSS) aims to accurately segment query images with guidance from only a few annotated support images. Previous methods typically rely on pixel-level feature correlations, denoted as the many-to-many (pixels-to-pixels) or few-to-many (prototype-to-pixels) manners. Recent mask proposals classification pipeline in semantic segmentation enables more efficient few-to-few (prototype-to-prototype) correlation between masks of query proposals and support reference. However, these methods still involve intermediate pixel-level feature correlation, resulting in lower efficiency. In this paper, we introduce the Proposal and Reference masks matching transFormer (PRFormer), designed to rigorously address mask matching in both spatial and semantic aspects in a thorough few-to-few manner. Following the mask-classification paradigm, PRFormer starts with a class-agnostic proposal generator to partition the query image into proposal masks. It then evaluates the features corresponding to query proposal masks and support reference masks using two strategies: semantic matching based on feature similarity across prototypes and spatial matching through mask intersection ratio. These strategies are implemented as the Prototype Contrastive Correlation (PrCC) and Prior-Proposals Intersection (PPI) modules, respectively. These strategies enhance matching precision and efficiency while eliminating dependence on pixel-level feature correlations. Additionally, we propose the category discrimination NCE (cdNCE) loss and IoU-KLD loss to constrain the adapted prototypes and align the similarity vector with the corresponding IoU between proposals and ground truth. Given that class-agnostic proposals tend to be more accurate for training classes than for novel classes in FSS, we introduce the Weighted Proposal Refinement (WPR) to refine the most confident masks with detailed features, yielding more precise predictions. Experiments on the popular Pascal-5i and COCO-20i benchmarks show that our Few-to-Few approach, PRFormer, outperforms previous methods, achieving mIoU scores of 70.4% and 49.4%, respectively, on 1-shot segmentation. Code is available athttps://github.com/ANDYZAQ/PRFormer.
Guangyu Gao, Anqi Zhang 0002, Jianbo Jiao, Chi Harold Liu, Yunchao Wei
IEEE Trans. Circuits Syst. Video Technol.1
2025 Source-Free Active Domain Adaptation via Augmentation-Based Sample Query and Progressive Model Adaptation
abstract
Active domain adaptation (ADA), which enormously improves the performance of unsupervised domain adaptation (UDA) at the expense of annotating limited target data, has attracted a surge of interest. However, in real-world applications, the source data in conventional ADA are not always accessible due to data privacy and security issues. To alleviate this dilemma, we introduce a more practical and challenging setting, dubbed as source-free ADA (SFADA), where one can select a small quota of target samples for label query to assist the model learning, but labeled source data are unavailable. Therefore, how to query the most informative target samples and mitigate the domain gap without the aid of source data are two key challenges in SFADA. To address SFADA, we propose a unified method SQAdapt via augmentation-based ample uery and progressive model Adapt ation. In specific, an active selection module (ASM) is built for target label query, which exploits data augmentation to select the most informative target samples with high predictive sensitivity and uncertainty. Then, we further introduce a classifier adaptation module (CAM) to leverage both the labeled and unlabeled target data for progressively calibrating the classifier weights. Meanwhile, the source-like target samples with low selection scores are taken as source surrogates to realize the distribution alignment in the source-free scenario by the proposed distribution alignment module (DAM). Moreover, as a general active label query method, SQAdapt can be easily integrated into other source-free UDA (SFUDA) methods, and improve their performance. Comprehensive experiments on multiple benchmarks have shown that SQAdapt can achieve superior performance and even surpass most of the ADA methods.
Shuang Li 0008, Rui Zhang 0113, Kaixiong Gong, Mixue Xie, Wenxuan Ma 0001, Guangyu Gao
IEEE Trans. Neural Networks Learn. Syst.6
2024 Background Adaptation with Residual Modeling for Exemplar-Free Class-Incremental Semantic Segmentation
Anqi Zhang 0002, Guangyu Gao
ECCV (52)2
2024 Audio-Visual Segmentation by Leveraging Multi-scaled Features Learning
Sze An Peter Tan, Guangyu Gao
MMM (2)2
2024 "Car or Bus?" CLearSeg: CLIP-Enhanced Discrimination Among Resembling Classes for Few-Shot Semantic Segmentation
Anqi Zhang 0002, Guangyu Gao, Zhuocheng Lv, Yukun An
MMM (1)2
2024 Bridge the Points: Graph-based Few-shot Segment Anything Semantically
abstract
The recent advancements in large-scale pre-training techniques have significantly enhanced the capabilities of vision foundation models, notably the Segment Anything Model (SAM), which can generate precise masks based on point and box prompts. Recent studies extend SAM to Few-shot Semantic Segmentation (FSS), focusing on prompt generation for SAM-based automatic semantic segmentation. However, these methods struggle with selecting suitable prompts, require specific hyperparameter settings for different scenarios, and experience prolonged one-shot inference times due to the overuse of SAM, resulting in low efficiency and limited automation ability. To address these issues, we propose a simple yet effective approach based on graph analysis. In particular, a Positive-Negative Alignment module dynamically selects the point prompts for generating masks, especially uncovering the potential of the background context as the negative reference. Another subsequent Point-Mask Clustering module aligns the granularity of masks and selected points as a directed graph, based on mask coverage over points. These points are then aggregated by decomposing the weakly connected components of the directed graph in an efficient manner, constructing distinct natural clusters. Finally, the positive and overshooting gating, benefiting from graph-based granularity alignment, aggregates high-confident masks and filters the false-positive masks for final prediction, reducing the usage of additional hyperparameters and redundant mask generation. Extensive experimental analysis across standard FSS, One-shot Part Segmentation, and Cross Domain FSS datasets validate the effectiveness and efficiency of the proposed approach, surpassing state-of-the-art generalist models with a mIoU of 58.7% on COCO-20i and 35.2% on LVIS-92i. The project page of this work is https://andyzaq.github.io/GF-SAM/.
Anqi Zhang 0002, Guangyu Gao, Jianbo Jiao, Yunchao Wei
NeurIPS2
2023 CoinSeg: Contrast Inter- and Intra- Class Representations for Incremental Segmentation
abstract
Class incremental semantic segmentation aims to strike a balance between the model’s stability and plasticity by maintaining old knowledge while adapting to new concepts. However, most state-of-the-art methods use the freeze strategy for stability, which compromises the model’s plasticity. In contrast, releasing parameter training for plasticity could lead to the best performance for all categories, but this requires discriminative feature representation. Therefore, we prioritize the model’s plasticity and propose the Contrast inter- and intra-class representations for Incremental Segmentation (CoinSeg), which pursues discriminative representations for flexible parameter tuning. Inspired by the Gaussian mixture model that samples from a mixture of Gaussian distributions, CoinSeg emphasizes intra-class diversity with multiple contrastive representation centroids. Specifically, we use mask proposals to identify regions with strong objectness that are likely to be diverse instances/centroids of a category. These mask proposals are then used for contrastive representations to reinforce intra-class diversity. Meanwhile, to avoid bias from intra-class diversity, we also apply category-level pseudo-labels to enhance category-level consistency and inter-category diversity. Additionally, CoinSeg ensures the model’s stability and alleviates forgetting through a specific flexible tuning strategy. We validate CoinSeg on Pascal VOC 2012 and ADE20K datasets with multiple incremental scenarios and achieve superior results compared to previous state-of-the-art methods, especially in more challenging and realistic long-term scenarios. Code is available at https://github.com/zkzhang98/CoinSeg.
Zekang Zhang, Guangyu Gao, Jianbo Jiao, Chi Harold Liu, Yunchao Wei
ICCV2
2023 DecTrans: Person Re-identification with Multifaceted Part Features via Decomposed Transformer
Yan Zhang 0134, Guangyu Gao, Qianxiang Wang
PRCV (12)2
2023 Hierarchical context-agnostic network with contrastive feature diversity for one-shot semantic segmentation
Zhiyuan Fang, Guangyu Gao, Zekang Zhang, Anqi Zhang 0002
J. Vis. Commun. Image Represent.2
2023 Hardest and semi-hard negative pairs mining for text-based person search with visual-textual attention
Qianxiang Wang, Guangyu Gao
Multim. Syst.3
2023 DANet: Semi-supervised differentiated auxiliaries guided network for video action recognition
Guangyu Gao, Ziming Liu 0003, Jinyang Li 0007, A. K. Qin 0001
Neural Networks1
2023 A Fast Data-Driven Iteratively Regularized Method with Convex Penalty for Solving Ill-Posed Problems
abstract
Abstract. We propose a new iterative regularization method for solving inverse problems in Hilbert spaces. The iterative process of the proposed method combines classical iterative regularization format and Data-Driven approach. Data-Driven technique is based on the idea of deep learning to estimate the interior of a black box through a training set, so as to solve problems better and faster in some cases. In order to capture the special feature of solutions, convex functions are utilized to be penalty terms. Algorithmically, the two-point gradient acceleration strategy based on homotopy perturbation method is applied to the iterative scheme, which makes the method have satisfactory acceleration effect. We provide convergence analysis of the method under standard assumptions for iterative regularization methods. Finally, several numerical experiments are presented to show the effectiveness and acceleration effect of our method.
Guangyu Gao, Bo Han 0007, Zhenwu Fu, Shanshan Tong
SIAM J. Imaging Sci.1
2022 AONet: Attentional Occlusion-Aware Network for Occluded Person Re-identification
Guangyu Gao, Qianxiang Wang, Yan Zhang 0134
ACCV (5)1
2022 Adaptive Spatial-BCE Loss for Weakly Supervised Semantic Segmentation
Tong Wu 0014, Guangyu Gao, Junshi Huang, Xiaolin Wei, Xiaoming Wei, Chi Harold Liu
ECCV (29)2
2022 Mining Unseen Classes via Regional Objectness: A Simple Baseline for Incremental Segmentation
abstract
Incremental or continual learning has been extensively studied for image classification tasks to alleviate catastrophic forgetting, a phenomenon in which earlier learned knowledge is forgotten when learning new concepts. For class incremental semantic segmentation, such a phenomenon often becomes much worse due to the semantic shift of the background class, \ie, some concepts learned at previous stages are assigned to the background class at the current training stage, therefore, significantly reducing the performance of these old concepts. To address this issue, we propose a simple yet effective method in this paper, named Mining unseen Classes via Regional Objectness (MicroSeg). Our MicroSeg is based on the assumption that \emph{background regions with strong objectness possibly belong to those concepts in the historical or future stages}. Therefore, to avoid forgetting old knowledge at the current training stage, our MicroSeg first splits the given image into hundreds of segment proposals with a proposal generator. Those segment proposals with strong objectness from the background are then clustered and assigned new defined labels during the optimization. In this way, the distribution characterizes of old concepts in the feature space could be better perceived, relieving the catastrophic forgetting caused by the semantic shift of the background class accordingly. We conduct extensive experiments on Pascal VOC and ADE20K, and competitive results well demonstrate the effectiveness of our MicroSeg. Code is available at \href{https://github.com/zkzhang98/MicroSeg}{\textcolor{orange}{\texttt{https://github.com/zkzhang98/MicroSeg}}}.
Zekang Zhang, Guangyu Gao, Zhiyuan Fang, Jianbo Jiao, Yunchao Wei
NeurIPS2
2022 Perceiving informative key-points: A self-attention approach for person search
Guangyu Gao, Cen Han
Signal Process. Image Commun.1
2022 DRNet: Double Recalibration Network for Few-Shot Semantic Segmentation
abstract
Few-shot segmentation aims at learning to segment query images guided by only a few annotated images from the support set. Previous methods rely on mining the feature embedding similarity across the query and the support images to achieve successful segmentation. However, these models tend to perform badly in cases where the query instances have a large variance from the support ones. To enhance model robustness against such intra-class variance, we propose a Double Recalibration Network (DRNet) with two recalibration modules, i.e., the Self-adapted Recalibration (SR) module and the Cross-attended Recalibration (CR) module. In particular, beyond learning robust feature embedding for pixel-wise comparison between support and query as in conventional methods, the DRNet further exploits semantic-aware knowledge embedded in the query image to help segment itself, which we call ‘self-adapted recalibration’. More specifically, DRNet first employs guidance from the support set to roughly predict an incomplete but correct initial object region for the query image, and then reversely uses the feature embedding extracted from the incomplete object region to segment the query image. Also, we devise a CR module to refine the feature representation of the query image by propagating the underlying knowledge embedded in the support image’s foreground to the query. Instead of foreground global pooling, we refine the response at each pixel in the query feature map by attending to all foreground pixels in the support feature map and taking the weighted average by their similarity; meanwhile, feature maps of the query image are also added back to weighted feature maps as a residual connection. Our DRNet can effectively address the intra-class variance under the few-shot setting with such two recalibration modules, and mine more accurate target regions for query images. We conduct extensive experiments on the popular benchmarks PASCAL-$5^{i}$and COCO-$20^{i}$. The DRNet with the best configuration achieves the mIoU of$\textbf {63.6}\%$and$\textbf {64.9}\%$on PASCAL-$5^{i}$and$\textbf {44.7}\%$and$\textbf {49.6}\%$on COCO-$20^{i}$for 1-shot and 5-shot settings respectively, significantly outperforming the state-of-the-arts without any bells and whistles. Code is available at:https://github.com/fangzy97/drnet.
Guangyu Gao, Zhiyuan Fang, Cen Han, Yunchao Wei, Chi Harold Liu, Shuicheng Yan
IEEE Trans. Image Process.1
2021 Embedded Discriminative Attention Mechanism for Weakly Supervised Semantic Segmentation
abstract
Weakly Supervised Semantic Segmentation (WSSS) with image-level annotation uses class activation maps from the classifier as pseudo-labels for semantic segmentation. However, such activation maps usually highlight the local discriminative regions rather than the whole object, which deviates from the requirement of semantic segmentation. To explore more comprehensive class-specific activation maps, we propose an Embedded Discriminative Attention Mechanism (EDAM) by integrating the activation map generation into the classification network directly for WSSS. Specifically, a Discriminative Activation (DA) layer is designed to explicitly produce a series of normalized class-specific masks, which are then used to generate class-specific pixel-level pseudo-labels demanded in segmentation. For learning the pseudo-labels, the masks are multiplied with the feature maps after the backbone to generate the discriminative activation maps, each of which encodes the specific information of the corresponding category in the input images. Given such class-specific activation maps, a Collaborative Multi-Attention (CMA) module is proposed to extract the collaborative information of each given category from images in a batch. In inference, we directly use the activation masks from the DA layer as pseudo-labels for segmentation. Based on the generated pseudo-labels, we achieve the mIoU of 70.60% on PASCAL VOC 2012 segmentation testset, which is the new state-of-the-art, to our best knowledge. Code and pre-trained models are available online soon.
Tong Wu 0014, Junshi Huang, Guangyu Gao, Xiaoming Wei, Xiaolin Wei, Chi Harold Liu
CVPR3
2021 HRDNet: High-Resolution Detection Network for Small Objects
abstract
Small object detection is a very challenging yet practical vision task. With deep network-based methods, the contextual information of small objects may disappear when the network goes deeper. An intuitive solution to alleviate this issue is to increase the input resolution, however, it will aggravate the large variant of object scale and introduce unbearable computation cost. To leverage the benefits of high-resolution images without bringing up new problems, we propose a High-Resolution Detection Network (HRDNet) which takes multiple resolution inputs with multi-depth backbones. Meanwhile, we propose the Multi-Depth Image Pyramid Network (MD-IPN) and Multi-Scale Feature Pyramid Network (MS-FPN). The MD-IPN maintains multiple position information using multiple depth backbones. Specifically, high-resolution input will be fed into a shallow network to reserve more positional information and reduce computational costs, while low-resolution input will be fed into a deep network to extract more semantics. By extracting various features from high to low resolutions, the MD-IPN can improve the performance of small object detection and maintain the performance of middle and large objects. Additionally, MS-FPN is introduced to align and fuse multi-scale feature groups generated by MD-IPN to reduce the information imbalance. Extensive experiments are conducted on the COCO2017 and the typical small object dataset, VisDrone 2019. Notably, our HRDNet achieves the state-of-the-art on these two datasets with significant improvements on small objects.
Ziming Liu 0003, Guangyu Gao, Lin Sun 0004, Zhiyuan Fang
ICME2
2020 Anti-distractors: two-branch siamese tracker with both static and dynamic filters for object tracking
Hao Shen 0010, Tao Song 0004, Guangyu Gao
Multim. Syst.4
2019 Action Recognition with Bootstrapping based Long-range Temporal Context Attention
abstract
Actions always refer to complex vision variations in a long-range redundant video sequence. Instead of focusing on limited range sequence, i.e. convolution on adjacent frames, in this paper, we proposed an action recognition approach with bootstrapping based long-range temporal context attention. Specifically, due to vision variations of the local region across frames, we target at capturing temporal context by proposing the Temporal Pixels based Parallel-head Attention (TPPA) block. In TPPA, we apply the self-attention mechanism between local regions at the same position across temporal frames to capture the interaction impacts. Meanwhile, to deal with video redundancy and capture long-range context, the TPPA is extended to the Random Frames based Bootstrapping Attention (RFBA) framework. While the bootstrapping sampling frames have the same distribution of the whole video sequence, the RFBA not only captures longer temporal context with only a few sampling frames but also has comprehensive representation through multiple sampling. Furthermore, we also try to apply this temporal context attention to image-based action recognition, by transforming the image into "pseudo video" with the spatial shift. Finally, we conduct extensive experiments and empirical evaluations on two most popular datasets:UCF101 for videos andStanford40 for images. In particular, our approach achieves top-1 accuracy of $91.7%$ in UCF101 and mAP of $90.9%$ in Stanford40.
Ziming Liu 0003, Guangyu Gao, A. K. Qin 0001, Tong Wu 0014, Chi Harold Liu
ACM Multimedia2
2019 Fashion clothes matching scheme based on Siamese Network and AutoEncoder
Guangyu Gao, Liling Liu
Multim. Syst.1
2019 Real-time small traffic sign detection with revised faster-RCNN
Cen Han, Guangyu Gao
Multim. Tools Appl.2
2018 Optimal feature combination analysis for crowd saliency prediction
Guangyu Gao, Cen Han, Chi Harold Liu, Erwu Liu
J. Vis. Commun. Image Represent.1
2017 Novel evaluation metrics for seam carving based image retargeting
abstract
Image retargeting effectively resizes images by preserving the recognizability of important image regions. Most of retargeting methods rely on good importance maps as a cue to retain or remove certain regions in the input image. In addition, the traditional evaluation exhaustively depends on user ratings. There is a legitimate need for a methodological approach for evaluating retargeted results. Therefore, in this paper, we conduct a study and analysis on the prominent method in image retargeting, Seam Carving. First, we introduce two novel evaluation metrics which can be considered as the proxy of user ratings. Second, we exploit salient object dataset as a benchmark for this task. We then investigate different types of importance maps for this particular problem. The experiments show that humans in general agree with the evaluation metrics on the retargeted results and some importance map methods are consistently more favorable than others.
Tam V. Nguyen 0002, Guangyu Gao
ICIP2
2016 Cast2Face: Assigning Character Names Onto Faces in Movie With Actor-Character Correspondence
abstract
Automatically identifying characters in movies has attracted researchers' interest and led to several significant and interesting applications. However, due to the vast variation in character appearance as well as the weakness and ambiguity of available annotation, it is still a challenging problem. In this paper, we investigate this problem with the supervision of actor-character name correspondence provided by the movie cast. Our proposed framework, namely, Cast2Face, is featured by: 1) we restrict the assigned names within the set of character names in the cast; 2) for each character, by using the corresponding actor and movie name as keywords, we retrieve from the Google image search and get a group of face images to form the gallery set; 3) the probe face tracks in the movie are then identified as one of the actors by a robust kernel multitask joint sparse representation and classification method; and 4) the conditional random field model with consideration of the constraints between face tracks is introduced to enhance the final labeling. Finally, the assigned actor name of a face track is then mapped to the character name based on the cast again. Besides face naming, we further apply the proposed method to spotlight the summarization of a particular actor in his/her movies. We conduct extensive experiments and empirical evaluations on several feature-length movies to demonstrate the satisfying performance of our method.
Guangyu Gao, Mengdi Xu, Jialie Shen 0001, Huadong Ma, Shuicheng Yan
IEEE Trans. Circuits Syst. Video Technol.1
2016 Cloud-Based Actor Identification With Batch-Orthogonal Local-Sensitive Hashing and Sparse Representation
abstract
Recognizing and retrieving multimedia content with movie/TV series actors, especially querying actor-specific videos in large scale video datasets, has attracted much attention in both the video processing and computer vision research field. However, many existing methods have low efficiency both in training and testing processes and also a less than satisfactory performance. Considering these challenges, in this paper, we propose an efficient cloud-based actor identification approach with batch-orthogonal local-sensitive hashing (BOLSH) and multi-task joint sparse representation classification. Our approach is featured by the following: 1) videos from movie/TV series are segmented into shots with the cloud-based shot boundary detection; 2) while faces in each shot are detected and tracked, the cloud-based BOLSH is then implemented on these faces for feature description; 3) the sparse representation is then adopted for actor identification in each shot; and 4) finally, a simple application, actor-specific shots retrieval is realized to verify our approach. We conduct extensive experiments and empirical evaluations on a large scale dataset, to demonstrate the satisfying performance of our approach considering both accuracy and efficiency.
Guangyu Gao, Chi Harold Liu, Min Chen 0003, Song Guo 0001, Kin K. Leung
IEEE Trans. Multim.1
2015 The Optimization and Improvement of MapReduce in Web Data Mining
Changqing Yin, Shukun Liu, Shangwei Song, Guangyu Gao, Xiyuan Zhou
ICA3PP (1)5
2014 Movie Scene Recognition Using Panoramic Frame and Representative Feature Patches
Guangyu Gao, Hua-Dong Ma
J. Comput. Sci. Technol.1
2014 To accelerate shot boundary detection by reducing detection region and scope
Guangyu Gao, Huadong Ma
Multim. Tools Appl.1
2014 Detecting both superimposed and scene text with multiple languages and multiple alignments in video
Xiaodong Huang 0005, Huadong Ma, Charles Ling 0001, Guangyu Gao
Multim. Tools Appl.4
2014 Batch-Orthogonal Locality-Sensitive Hashingfor Angular Similarity
abstract
Sign-random-projection locality-sensitive hashing (SRP-LSH) is a widely used hashing method, which provides an unbiased estimate of pairwise angular similarity, yet may suffer from its large estimation variance. We propose in this work batch-orthogonal locality-sensitive hashing (BOLSH), as a significant improvement of SRP-LSH. Instead of independent random projections, BOLSH makes use of batch-orthogonalized random projections, i.e, we divide random projection vectors into several batches and orthogonalize the vectors in each batch respectively. These batch-orthogonalized random projections partition the data space into regular regions, and thus provide a more accurate estimator. We prove theoretically that BOLSH still provides an unbiased estimate of pairwise angular similarity, with a smaller variance for any angle in (0, π), compared with SRP-LSH. Furthermore, we give a lower bound on the reduction of variance. The extensive experiments on real data well validate that with the same length of binary code, BOLSH may achieve significant mean squared error reduction in estimating pairwise angular similarity. Moreover, BOLSH shows the superiority in extensive approximate nearest neighbor (ANN) retrieval experiments.
Jianqiu Ji, Shuicheng Yan, Jianmin Li 0001, Guangyu Gao, Qi Tian 0001, Bo Zhang 0010
IEEE Trans. Pattern Anal. Mach. Intell.4
2014 Adaptive Learning for Celebrity Identification With Video Context
abstract
In this paper, we propose a novel semi-supervised learning strategy to address the problem of celebrity identification. The video context information is explored to facilitate the learning process based on the assumption that faces in the same video track share the same identity. Once a frame within a track is recognized confidently, the label can be propagated through the whole track, referred to as the confident track. More specifically, given a few static images and vast face videos, an initial weak classifier is trained and gradually evolves by iteratively promoting the confident tracks into the “labeled” set. The iterative selection process enriches the diversity of the “labeled” set such that the performance of the classifier is gradually improved. This learning theme may suffer from semantic drifting caused by errors in selecting the confident tracks. To address this issue, we propose to treat the selected frames as related samples-an intermediate state between labeled and unlabeled instead of labeled as in the traditional approach. To evaluate the performance, we construct a new dataset, which includes 3000 static images and 2700 face tracks of 30 celebrities. Comprehensive evaluations on this dataset and a public video dataset indicate significant improvement of our approach over established baseline methods.
Guangyu Gao, Zhengjun Zha, Shuicheng Yan, Huadong Ma, Tae-Kyun Kim 0001
IEEE Trans. Multim.2
2013 Static saliency vs. dynamic saliency: a comparative study
abstract
Recently visual saliency has attracted wide attention of researchers in the computer vision and multimedia field. However, most of the visual saliency-related research was conducted on still images for studying static saliency. In this paper, we give a comprehensive comparative study for the first time of dynamic saliency (video shots) and static saliency (key frames of the corresponding video shots), and two key observations are obtained: 1) video saliency is often different from, yet quite related with, image saliency, and 2) camera motions, such as tilting, panning or zooming, affect dynamic saliency significantly. Motivated by these observations, we propose a novel camera motion and image saliency aware model for dynamic saliency prediction. The extensive experiments on two static-vs-dynamic saliency datasets collected by us show that our proposed method outperforms the state-of-the-art methods for dynamic saliency prediction. Finally, we also introduce the application of dynamic saliency prediction for dynamic video captioning, assisting people with hearing impairments to better entertain videos with only off-screen voices, e.g., documentary films, news videos and sports videos.
Tam V. Nguyen 0002, Mengdi Xu, Guangyu Gao, Mohan Kankanhalli, Qi Tian 0001, Shuicheng Yan
ACM Multimedia3
2012 Multi-modality movie scene detection using Kernel Canonical Correlation Analysis
Guangyu Gao, Huadong Ma
ICPR1
2011 Accelerating Shot Boundary Detection by reducing spatial and temporal redundant information
abstract
Shot Boundary Detection (SBD) is the fundamental process in video processing area. However, according to the most existing SBD methods [1], [2], researchers have paid much attention to detect the shot boundaries as accurately as possible with expensive computation cost. In this paper, we propose an approach to accelerate the SBD process by reducing both spatial pixels and temporal frames while keeping satisfactory performance. Our method accelerates the SBD process mainly from two points of view. In spatial domain, the proposed approach only uses the pixels falling into the defined focus region rather than all pixels in a frame; in temporal domain, a step-skip method is employed to reduce the processed frames. Through reducing both detection region and scope, using the corner distribution of frames to remove most of false boundaries, the speed of our SBD approach is improved obviously. We conduct extensive experiments to evaluate the proposed approach, and the results show that our approach can obviously accelerate the SBD process by reducing the computation complexity while keeping satisfactory performance.
Guangyu Gao, Huadong Ma
ICME1