Laiyun Qing

dblp:67/5325 · DBLP profile ↗
← Back
56ranked-venue papers
3as first author
17since 2021 · last 2026
0000-0001-9923-5034ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 34 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 1 first-author · 9 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Mamba-Based Temporally Guided Spatial Alignment for Unaligned RGB-T Tracking
Guorong Li, Chen Zhang 0013, Wentao Cao, Laiyun Qing
ICIC (1)6
2026 Synergistic Dual-Graph Co-Evolutionary Network for point-supervised temporal action localization
Laiyun Qing, Guorong Li, Qingming Huang
Comput. Vis. Image Underst.3
2026 Dynamic example network for class-agnostic object counting
Xinyan Liu 0008, Guorong Li, Yuankai Qi, Ziheng Yan, Weigang Zhang, Laiyun Qing, Qingming Huang
Pattern Recognit.6
2026 RETTA: Retrieval-enhanced test-time adaptation for zero-shot video captioning
abstract
Despite the significant progress of fully-supervised video captioning, zero-shot methods remain much less explored. In this paper, we propose a novel zero-shot video captioning framework named R etrieval- E nhanced T est- T ime A daptation (RETTA), which takes advantage of existing pre-trained large-scale vision and language models to directly generate captions with test-time adaptation. Specifically, we bridge video and text using four key models: a general video-text retrieval model XCLIP, a general image-text matching model CLIP, a text alignment model AnglE, and a text generation model GPT-2, due to their source-code availability. The main challenge is how to enable the text generation model to be sufficiently aware of the content in a given video so as to generate corresponding captions. To address this problem, we propose using learnable tokens as a communication medium among these four frozen models GPT-2, XCLIP, CLIP, and AnglE. Different from the conventional way that trains these tokens with training data, we propose to learn these tokens with soft targets of the inference data under several carefully crafted loss functions, which enable the tokens to absorb video information catered for GPT-2. This adaptation requires only a few iterations ( e.g. , 16) and does not require ground truth data. Extensive experimental on MSR-VTT, MSVD, and VATEX, show absolute 5.1 % ∼ 32.4 % improvements in CIDEr scores compared to several state-of-the-art zero-shot video captioning methods.
Yunchuan Ma, Laiyun Qing, Guorong Li, Yuankai Qi, Amin Beheshti, Quan Z. Sheng, Qingming Huang
Pattern Recognit.2
2025 SDVPT: Semantic-Driven Visual Prompt Tuning for Open-world Object Counting
abstract
Open-world object counting leverages the robust text-image alignment of pre-trained vision-language models (VLMs) to enable counting of arbitrary categories in images specified by textual queries. However, widely adopted naive fine-tuning strategies concentrate exclusively on text-image consistency for categories contained in, which leads to limited generalizability for unseen categories. In this work, we propose a plug-and-play Semantic-Driven Visual Prompt Tuning framework (SDVPT) that transfers knowledge from the training set to unseen categories with minimal overhead in parameters and inference time. First, we introduce a two-stage visual prompt learning strategy composed of Category-Specific Prompt Initialization (CSPI) and Topology-Guided Prompt Refinement (TGPR). The CSPI generates category-specific visual prompts, and then TGPR distills latent structural patterns from the VLM's text encoder to refine these prompts. During inference, we dynamically synthesize the visual prompts for unseen categories based on the semantic correlation between unseen and training categories, facilitating robust text-image alignment for unseen categories. Extensive experiments integrating SDVPT with all available open-world object counting models demonstrate its effectiveness and adaptability across three widely used datasets: FSC-147, CARPK, and PUCPR+. Code is available https://github.com/Eamon-0v0/SDVPT
Guorong Li, Laiyun Qing, Amin Beheshti, Jian Yang 0001, Quan Z. Sheng, Yuankai Qi, Qingming Huang
ACM Multimedia3
2025 MMA: Video Reconstruction for Spike Camera Based on Multiscale Temporal Modeling and Fine-Grained Attention
abstract
This paper presents a Multiscale Temporal Correlation Learning with the Mamba-Fused Attention Model (MMA), an efficient and effective method for reconstructing a video clip from a spike stream. Spike cameras offer unique advantages for capturing rapid scene changes with high temporal resolution. A spike stream contains sufficient information for multiple image reconstructions. However, existing methods generate only a single image at a time for a given spike stream, which results in excessive redundant computations between consecutive frames when aiming at restoring a video clip, thereby increasing computational costs significantly. The proposed MMA addresses such challenges by constructing a spike-to-video model, directly producing an image sequence at a time. Specifically, we propose a U-shaped Multiscale Temporal Correlation Learning (MTCL) to fuse the features at different temporal resolutions for clear video reconstruction. At each scale, we introduce a Fine-Grained Attention (FGA) module for fine-spatial context modeling within a patch and a Mamba module for integrating features across patches. Adopting a lightweight U-shaped structure and fine-grained feature extraction at each level, our method reconstructs high-quality image sequences quickly. The experimental results show that the proposed MMA surpasses current state-of-the-art methods in image quality, computation cost, and model size.
Dilmurat Alim, Chen Yang 0034, Laiyun Qing, Guorong Li, Qingming Huang
IEEE Signal Process. Lett.3
2025 Boosting UAV Detection via Memory-Enhanced Attention and Contrastive Learning
abstract
With unmanned aerial vehicles (UAVs) having emerged in diverse application domains, visual detection of UAVs has become a critical research focus in recent years. However, most existing methods are limited in capturing small UAVs and may not perform well in complex backgrounds. To address these challenges, we propose a novel detection framework that integrates newly designed memory mechanism and contrastive loss to improve UAV detection. Specifically, we first utilize a clustering algorithm to gather representative UAV prototypes, which are then utilized to construct a reliable memory bank. Then, we design a UAV Memory-Enhanced Attention (UMEA) module to propagate high-confidence prototypes from the memory bank, thereby enhancing the appearance features of UAVs in the input frame. Furthermore, we introduce a Memory-Driven Contrastive Learning (MDCL) loss function to pull UAVs closer in the feature space while pushing them further away from the background. Extensive experiments conducted on three challenging datasets, NPS-Drones, ARD-MAV and Drone-vs-Bird demonstrate that the proposed method outperforms several state-of-the-art models in terms of the main metric AP with a large absolute margin, 2.1%, 3.6%, and 4.4%, respectively.
Yunchuan Ma, Yuankai Qi, Laiyun Qing, Guorong Li
IEEE Signal Process. Lett.4
2025 Boost Tracking by Natural Language With Prompt-Guided Grounding
abstract
TNL (Tracking by Natural Language) aims to locate the target described by a natural language sentence in a video. Most existing TNL methods are typically composed of three modules: object grounding, object tracking, and switching module, and their performance is limited by the poor performance of the grounding and switching modules due to the complex backgrounds and inaccurate information stored in the memory. This paper presents a global-local framework to address these issues, which includes a prompt-guided grounding module, a trained local tracking module, and a memory-based switcher module. The prompt-guided grounding module uses noun prompts to guide the CLIP model in focusing more on target regions and aligning visual features semantically with linguistic features, avoiding being misled by distractors and background. The memory-based switch module stores historical information with higher-quality memory, allowing the model to make more accurate decisions based on reliable data, thus improving the overall performance. Experiments on TNL2K, LaSOT, and OTB-Lang demonstrate the effectiveness and generalizability of the proposed framework.
Hengyou Li, Xinyan Liu 0008, Guorong Li, Shuhui Wang, Laiyun Qing, Qingming Huang
IEEE Trans. Intell. Transp. Syst.5
2025 Dynamic Erasing Network With Adaptive Temporal Modeling for Weakly Supervised Video Anomaly Detection
abstract
The weakly supervised video anomaly detection aims to learn a detection model using only video-level labeled data. Prior studies ignore the complexity or duration of anomalies present in abnormal videos during temporal modeling. Moreover, existing works usually detect the most abnormal segments, potentially overlooking the completeness of anomalies. We propose a dynamic erasing network (DE-Net) for weakly supervised video anomaly detection, which learns video-specific temporal features via adaptive temporal modeling (ATM) to address these limitations. Specifically, to handle duration variations of abnormal events, we propose an ATM module capable of adaptively selecting and aggregating the most appropriate K temporal scale features for each video. Then, we design a dynamic erasing (DE) strategy that dynamically assesses the completeness of the detected anomalies and erases prominent abnormal segments to encourage the model to discover gentle abnormal segments. The proposed method achieves favorable performance compared to several state-of-the-art approaches on the widely used XD-Violence, TAD, and UCF-Crime datasets.
Chen Zhang 0013, Guorong Li, Yuankai Qi, Hanhua Ye, Laiyun Qing, Ming-Hsuan Yang 0001, Qingming Huang
IEEE Trans. Neural Networks Learn. Syst.5
2024 Directly Locating Actions in Video with Single Frame Annotation
abstract
We propose a novel method for point-supervised action localization.Differs from the common practice of locating actions by first categorizing each video frame, our method directly predicts actions' positions and length. Specifically, point-supervised action localization is achieved by a series of fully supervised action location iteratively. In each iteration, the input video are used as input tokens and fed into a transformer, where the encoder extracts global context of the clips, and the decoder generates queries containing information for action localization. Three MLP heads are built on each query to obtain the probability, the center, and the length of each action instance respectively. Experiments on three popular datasets prove the potential of our method.
Haoran Tong, Xinyan Liu 0008, Guorong Li, Laiyun Qing
ICMR4
2024 Point-Supervised Temporal Action Detection with Label Supplementation Based on Transformer
Cui Xu, Laiyun Qing
MMAsia2
2024 Style-aware two-stage learning framework for video captioning
abstract
Significant progress has been made in video captioning in recent years. However, most existing methods directly learn from all given captions without distinguishing the styles of captions. The large diversity in these captions might bring ambiguity to the model learning. To address this issue, we propose a style-aware two-stage learning framework. In the first stage, the model is trained with captions of separate styles, including length style (short, medium, long), action style (single action or multiple actions), and object style (one object or more). For efficiency, a shared model with multiple individual style vectors is learned. In the second stage, a video style encoder is devised to capture style information from the input video, and it outputs a guidance signal of how to utilize the style vectors for the final caption generation. Without whistles and bells, our method achieves state-of-the-art performance on three widely-used public datasets, MSVD, MSR-VTT and VATEX. The source code and trained models will be made available to the public.
Yunchuan Ma, Yuankai Qi, Amin Beheshti, Laiyun Qing, Guorong Li
Knowl. Based Syst.6
2024 Learning Hierarchical Modular Networks for Video Captioning
abstract
Video captioning aims to generate natural language descriptions for a given video clip. Existing methods mainly focus on end-to-end representation learning via word-by-word comparison between predicted captions and ground-truth texts. Although significant progress has been made, such supervised approaches neglect semantic alignment between visual and linguistic entities, which may negatively affect the generated captions. In this work, we propose a hierarchical modular network to bridge video representations and linguistic semantics at four granularities before generating captions: entity, verb, predicate, and sentence. Each level is implemented by one module to embed corresponding semantics into video representations. Additionally, we present a reinforcement learning module based on the scene graph of captions to better measure sentence similarity. Extensive experimental results show that the proposed method performs favorably against the state-of-the-art models on three widely-used benchmark datasets, including microsoft research video description corpus (MSVD), MSR-video to text (MSR-VTT), and video-and-TEXt (VATEX).
Guorong Li, Hanhua Ye, Yuankai Qi, Shuhui Wang, Laiyun Qing, Qingming Huang, Ming-Hsuan Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 SpikeODE: Image Reconstruction for Spike Camera With Neural Ordinary Differential Equation
abstract
The recently invented retina-inspired spike camera has shown great potential for capturing dynamic scenes. However, reconstructing high-quality images from the binary spike data remains a challenge due to the existence of noises in the camera. This paper proposes SpikeODE, a novel approach to reconstructing clear images by exploring temporal-spatial correlation to depress noises. The main idea of our method is to restore the continuous dynamic process of real scenes in a latent space and learn the temporal correlations in a fine-grained manner. Furthermore, to model the dynamic process more effectively, we design a conditional ODE where the latent state of each timestamp is conditioned on the observed spike data. Subsequently, forward and backward inferences are conducted through the ODE to investigate the correlations between the representation of the target timestamp and the information from both past and future contexts. Additionally, we incorporate a Unet structure with a pixel-wise attention mechanism at each level to learn spatial correlations. Experimental results demonstrate that our method outperforms state-of-the-art methods across several metrics.
Chen Yang 0034, Guorong Li, Shuhui Wang, Li Su 0003, Laiyun Qing, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.5
2023 Exploiting Completeness and Uncertainty of Pseudo Labels for Weakly Supervised Video Anomaly Detection
abstract
Weakly supervised video anomaly detection aims to identify abnormal events in videos using only video-level labels. Recently, two-stage self-training methods have achieved significant improvements by self-generating pseudo labels and self-refining anomaly scores with these labels. As the pseudo labels play a crucial role, we propose an enhancement framework by exploiting completeness and uncertainty properties for effective self-training. Specifically, we first design a multi-head classification module (each head serves as a classifier) with a diversity loss to maximize the distribution differences of predicted pseudo labels across heads. This encourages the generated pseudo labels to cover as many abnormal events as possible. We then devise an iterative uncertainty pseudo label refinement strategy, which improves not only the initial pseudo labels but also the updated ones obtained by the desired classifier in the second stage. Extensive experimental results demonstrate the proposed method performs favorably against state-of-the-art approaches on the UCF-Crime, TAD, and XD-Violence benchmark datasets.
Chen Zhang 0013, Guorong Li, Yuankai Qi, Shuhui Wang, Laiyun Qing, Qingming Huang, Ming-Hsuan Yang 0001
CVPR5
2022 Enhanced Semantic Head for Cascade Instance Segmentation
abstract
Recently, cascade instance segmentation inspired by cascade object detection has achieved notable performance. Due to the lack of global information, many methods suffer from incomplete segmentation such as missing edge regions and discontinuities within instances. To solve this problem, we proposed an effective and flexible semantic head to extract enhanced spatial context information. A vision transformer is utilized to generate global context features, and a convolution network is adopted to generate spatial context features. After combining the two modules, we obtain enhanced semantic segmentation features for segmentation. Extensive experiments show that the enhanced semantic head achieves 40.6% and 42.3% mask AP for cascade predictor HTC and DSC, which surpass about 0.9 and 1.4 percentage points respectively. The enhanced semantic head is universal and effective to improve the performance of different cascade predictors.
Xuerong Huang, Li Su 0003, Guorong Li, Xinfeng Zhang 0001, Laiyun Qing, Qingming Huang
ICME5
2022 Multi-Attention Network for Compressed Video Referring Object Segmentation
abstract
Referring video object segmentation aims to segment the object referred by a given language expression. Existing works typically require compressed video bitstream to be decoded to RGB frames before being segmented, which increases computation and storage requirements and ultimately slows the inference down. This may hamper its application in real-world computing resource limited scenarios, such as autonomous cars and drones. To alleviate this problem, in this paper, we explore the referring object segmenta- tion task on compressed videos, namely on the original video data flow. Besides the inherent difficulty of the video referring object segmentation task itself, obtaining discriminative representation from compressed video is also rather challenging. To address this problem, we propose a multi-attention network which consists of dual-path dual-attention module and a query-based cross-modal Transformer module. Specifically, the dual-path dual-attention module is designed to extract effective representation from compressed data in three modalities, i.e., I-frame, Motion Vector and Residual. The query-based cross-modal Transformer firstly models the corre- lation between linguistic and visual modalities, and then the fused multi-modality features are used to guide object queries to generate a content-aware dynamic kernel and to predict final segmentation masks. Different from previous works, we propose to learn just one kernel, which thus removes the complicated post mask-matching procedure of existing methods. Extensive promising experimental results on three challenging datasets show the effectiveness of our method compared against several state-of-the-art methods which are proposed for processing RGB data. Source code is available at: https://github.com/DexiangHong/MANet.
Weidong Chen 0013, Dexiang Hong, Yuankai Qi, Zhenjun Han, Shuhui Wang, Laiyun Qing, Qingming Huang, Guorong Li
ACM Multimedia6
2020 Action prediction via deep residual feature learning and weighted loss
Shuangshuang Guo, Laiyun Qing, Lijuan Duan
Multim. Tools Appl.2
2019 Temporal Convolutional Network with Complementary Inner Bag Loss for Weakly Supervised Anomaly Detection
abstract
Weakly supervised anomaly detection (WSAD) is a challenging task with only normal and anomaly video label supervision but required to localize intervals where anomalies take place. We employ multiple instance learning (MIL) for weakly supervised anomaly detection and define a novel inner bag loss (IBL) for MIL to constrain the function space of weakly supervised problem, which consider the lowest anomaly instance score and highest score in each bag. More specifically, the gap between the lowest score and highest score in a positive bag should be large, and that of a negative bag should be small. In order to model the temporal structure of video, we encode preceding adjacent instance by temporal convolutional network (TCN) explicitly. Experimental results show that temporal convolutional network with complementary inner bag loss outperforms the state-of-the-art on the Crime dataset.
Jiangong Zhang, Laiyun Qing
ICIP2
2019 Deep feature representation based on privileged knowledge transfer
Lijuan Duan, Qing En, Yuanhua Qiao, Laiyun Qing
Pattern Recognit. Lett.5
2018 Stereoscopic saliency model using contrast and depth-guided-background prior
Fangfang Liang, Lijuan Duan, Wei Ma 0008, Yuanhua Qiao, Zhi Cai, Laiyun Qing
Neurocomputing6
2017 Rotative maximal pattern: A local coloring descriptor for object classification and recognition
Junbiao Pang, Weigang Zhang, Laiyun Qing, Qingming Huang
Inf. Sci.5
2016 Human Interaction Recognition by Mining Discriminative Patches on Key Frames
Dingyi Shan, Laiyun Qing
ACCV (2)2
2016 Gaze Movement Control Neural Network Based on Multidimensional Topographic Class Grouping
Wenqi Zhong, Laiyun Qing
ICONIP (2)3
2015 Activity Auto-Completion: Predicting Human Activities from Partial Videos
abstract
In this paper, we propose an activity auto-completion (AAC) model for human activity prediction by formulating activity prediction as a query auto-completion (QAC) problem in information retrieval. First, we extract discriminative patches in frames of videos. A video is represented based on these patches and divided into a collection of segments, each of which is regarded as a character typed in the search box. Then a partially observed video is considered as an activity prefix, consisting of one or more characters. Finally, the missing observation of an activity is predicted as the activity candidates provided by the auto-completion model. The candidates are matched against the activity prefix on-the-fly and ranked by a learning-to-rank algorithm. We validate our method on UT-Interaction Set #1 and Set #2 [19]. The experimental results show that the proposed activity auto-completion model achieves promising performance.
Laiyun Qing
ICCV2
2015 Hierarchical Extreme Learning Machine for unsupervised representation learning
abstract
Learning representations from massive unlabeled data is a hot topic for high-level tasks in many applications. The recent great improvements on benchmark data sets, which are achieved by increasingly complex unsupervised learning methods and deep learning models with lots of parameters, usually require many tedious tricks and much expertise to tune. However, filters learned by these complex architectures are quite similar to standard hand-crafted features visually, and training the deep models costs quite long time to fine-tune their weights. In this paper, Extreme Learning Machine-Autoencoder (ELM-AE) is employed as the learning unit to learn local receptive fields at each layer, and the lower layer responses are transferred to the last layer (trans-layer) to form a more complete representation to retain more information. In addition, some beneficial methods in deep learning architectures such as local contrast normalization and whitening are added to the proposed hierarchical Extreme Learning Machine networks to further boost the performance. The obtained trans-layer representations are followed by block histograms with binary hashing to learn translation and rotation invariant representations, which are utilized to do high-level tasks such as recognition and detection. Compared to traditional deep learning methods, the proposed trans-layer representation method with ELM-AE based learning of local receptive filters has much faster learning speed and is validated in several typical experiments, such as digit recognition on MNIST and MNIST variations, object recognition on Caltech 101. State-of-the-art performances are achieved on both Caltech 101 15 samples per class task and 4 of 6 MNIST variations data sets, and highly impressive results are obtained on MNIST data set and other tasks.
Laiyun Qing, Guang-Bin Huang
IJCNN3
2015 Online dictionary learning for Local Coordinate Coding with Locality Coding Adaptors
Junbiao Pang, Chunjie Zhang 0001, Weigang Zhang, Laiyun Qing, Qingming Huang
Neurocomputing5
2014 Constrained Extreme Learning Machine: A novel highly discriminative random feedforward neural network
abstract
In this paper, a novel single hidden layer feedforward neural network, called Constrained Extreme Learning Machine (CELM), is proposed based on Extreme Learning Machine (ELM). In CELM, the connection weights between the input layer and hidden neurons are randomly drawn from a constrained set of difference vectors of between-class samples, rather than an open set of arbitrary vectors. Therefore, the CELM is expected to be more suitable for discriminative tasks, whilst retaining other advantages of ELM. The experimental results are presented to show the high efficiency of the CELM, compared with ELM and some other related learning machines.
Laiyun Qing
IJCNN3
2014 Vehicle detection in driving simulation using extreme learning machine
Jiangbi Hu, Laiyun Qing
Neurocomputing4
2014 Robust regression with extreme support vectors
Laiyun Qing
Pattern Recognit. Lett.3
2013 Salient region detection via texture-suppressed background contrast
abstract
We propose a novel salient region detection algorithm by texture-suppressed background contrast. We employ a structure extraction algorithm to suppress the small scale textures which are supposed to be not sensitive for human vision system. Then the texture-suppressed image is segmented into homogeneous superpixels. Motivated by the observation that the spatial distribution of the background has a high probability on the boundaries of images, we estimate the background as superpixels near the image boundaries. The saliency of each superpixel is then defined as the summation of its k minimum color distances to the estimated background superpixels. Finally a post-processing process involving spatial and color adjacency is employed to generate a per-pixel saliency map. Experimental results demonstrate that the proposed method outperforms the state-of-the-art approaches.
Jiamei Shuai, Laiyun Qing, Zhiguo Ma, Xilin Chen 0001
ICIP2
2013 Robust multi-patch tracking
abstract
In this paper, we propose a robust and fast multi-patch visual tracking algorithm within the Bayesian inference framework. The target template is initialized by selecting the object in the first frame manually and dividing it into small patches. For one certain frame, target candidates are sampled with the state transition model. Each candidate is divided into patches in the same way as the target template. By comparing the candidate's patches with the corresponding template patches, we can get the candidate's likelihood. The tracking result is the candidate which Maximum a Posteriori estimation. After that, tracking is continued using the Bayesian state inference and template update. Our approach can handle appearance variation, occlusion, illumination change, scale variation, rotation and cluttered background. The tracker is fast and performs favorably against several state-of-the-art trackers on challenging sequences.
Shanxin Yuan, Laiyun Qing
ICIP3
2012 Activity recognition based on semantic spatial relation
Lingxun Meng, Laiyun Qing, Peng Yang 0001, Xilin Chen 0001, Dimitris N. Metaxas
ICPR2
2011 Visual saliency detection by spatially weighted dissimilarity
abstract
In this paper, a new visual saliency detection method is proposed based on the spatially weighted dissimilarity. We measured the saliency by integrating three elements as follows: the dissimilarities between image patches, which were evaluated in the reduced dimensional space, the spatial distance between image patches and the central bias. The dissimilarities were inversely weighted based on the corresponding spatial distance. A weighting mechanism, indicating a bias for human fixations to the center of the image, was employed. The principal component analysis (PCA) was the dimension reducing method used in our system. We extracted the principal components (PCs) by sampling the patches from the current image. Our method was compared with four saliency detection approaches using three image datasets. Experimental results show that our method outperforms current state-of-the-art methods on predicting human fixations.
Lijuan Duan, Chunpeng Wu, Laiyun Qing
CVPR4
2011 Attention driven face recognition: A combination of spatial variant fixations and glance
Chongxiu Wang, Laiyun Qing, Fang Fang 0003, Xilin Chen 0001
FG2
2011 Bio-inspired Visual Saliency Detection and Its Application on Image Retargeting
Lijuan Duan, Chunpeng Wu, Haitao Qiao, Jili Gu, Laiyun Qing, Zhen Yang 0004
ICONIP (1)6
2011 Saliency Detection Based on Scale Selectivity of Human Visual System
Fang Fang 0003, Laiyun Qing, Xilin Chen 0001, Wen Gao 0001
ICONIP (1)2
2011 An improved neural architecture for gaze movement control in target searching
abstract
This paper presents an improved neural architecture for gaze movement control in target searching. Compared with the four-layer neural structure proposed in [14], a new movement coding neuron layer is inserted between the third layer and the fourth layer in previous structure for finer gaze motion estimation and control. The disadvantage of the previous structure is that all the large responding neurons in the third layer were involved in gaze motion synthesis by transmitting weighted responses to the movement control neurons in the fourth layer. However, these large responding neurons may produce different groups of movement estimation. To discriminate and group these neurons' movement estimation in terms of grouped connection weights form them to the movement control neurons in the fourth layer is necessary. Adding a new neuron layer between the third layer and the fourth lay is the measure that we solve this problem. Comparing experiments on target locating showed that the new architecture made the significant improvement.
Lijuan Duan, Laiyun Qing, Yuanhua Qiao
IJCNN3
2010 Lighting Aware Preprocessing for Face Recognition across Varying Illumination
Hu Han 0001, Shiguang Shan, Laiyun Qing, Xilin Chen 0001, Wen Gao 0001
ECCV (2)3
2010 Learning Internal Representation of Visual Context in a Neural Coding Network
Baixian Zou, Laiyun Qing, Lijuan Duan
ICANN (1)3
2009 A hybrid text segmentation approach
abstract
In this paper, we present a hybrid text segmentation approach for embedded text in images, aiming to combining the advantages of the difference-based methods and the similarity-based methods together. First a new stroke edge filter is applied to obtain stroke edge map. Then a two-threshold method based on the improved Niblack thresholding technique is utilized to identify stroke edges. Those pixels between the edge pairs above the high threshold are collected to estimate the representative of stroke color, so that stroke pixels are further extracted by computing the color similarity. Finally some heuristic rules are devised to integrate stroke edge and stroke region information to obtain better segmentation results. The experimental results show that our approach can effectively segment text from background.
Weiqiang Wang 0001, Qingming Huang, Wen Gao 0001, Laiyun Qing
ICME5
2009 Advertisement evaluation using visual saliency based on foveated image
abstract
This paper proposes a novel approach to advertisement evaluation using automatic salient regions. The salient regions are detected using a predicting model, in which the estimation are obtained by the space variant foveated image. The saliency is defined as the difference between the input image and its estimation. Then an advertisement is determined as attractive if the detected salient regions are overlapped with the interested regions of the advertisement. The experimental results on the advertisements data set are encouraging.
Zhiguo Ma, Laiyun Qing, Xilin Chen 0001
ICME2
2009 Single vs. population cell coding: Gaze movement control in target search
abstract
Gaze movement plays an important role in human visual search system. In literature, the winner-take-all method is wildly used to simulate the controlling of the gaze movement. The winner-take-all is a type of single-cell coding method, which uses one cell (grandmother cell) or one response to represent an object. However, eye movement is affected by the visual context which includes more than one object in images, especially in target search. Therefore, we propose to use the population coding with more than one response rather than the single-cell coding on gaze movement control. The proposed method is supported by the theoretical analysis and experiments on a real image database which show the population-cell-coding improves the target locating accuracy by 44.4% only at the cost of coding 22.4% more information than that of single-cell-coding.
Laiyun Qing, Lijuan Duan, Baixian Zou
IJCNN2
2009 Are Gabor phases really useless for face recognition?
Shiguang Shan, Laiyun Qing, Xilin Chen 0001, Wen Gao 0001
Pattern Anal. Appl.3
2008 Unified Principal Component Analysis with generalized Covariance Matrix for face recognition
abstract
Recently, 2DPCA and its variants have attracted much attention in face recognition area. In this paper, some efforts are made to discover the underlying fundaments of these methods, and a novel framework called Unified Principal Component Analysis (UPCA) is proposed. First, we introduce a novel concept, named Generalized Covariance Matrix (GCM), which is naturally derived from the traditional Covariance Matrix (CM). Each element of GCM is a generalized covariance of two random vectors rather than two scalar variables in CM. Based on GCM, the UPCA framework is proposed, from which the traditional PCA and its 2D counterparts can be deduced as special cases. Furthermore, under the UPCA framework, we not only revisit the existing 2D PCA methods and their limitations, but also propose two new methods: the grid-sampling method (GridPCA) and the intra-group correlation reduction method. Extensive experimental results on the FERET face database support the theoretical analysis and validate the feasibility of the proposed methods.
Shiguang Shan, Yu Su 0010, Laiyun Qing, Xilin Chen 0001, Wen Gao 0001
CVPR4
2008 Face reconstruction using fixation positions and foveated imaging
abstract
Face representation is important for face recognition system. Though most popular face representations are based on uniform grid sampling, some recent face recognition systems adopt weighted sampling on the different regions of a face. Psychological analysis of visual attention or human eye fixations on human face images may suggest some cues for face representation. human visual system (HVS) gives different weight to different region of human face via space-variant sampling on fovea and non-uniform distribution of fixations. This paper focuses on the problem of simulation of the foveated imaging phenomenon in HVS, and introduction of foveated imaging method into reconstruction of face in region of interest (ROI) using different fixation sources. We compare the effectiveness of actual fixation on reconstruction of face in ROI with uniform, random distribution fixation as well as fixation generated by artificial model. The experimental results on 100 face images from FRGC [7] data set show that actual fixation positions and model-generated fixation positions reconstruct the face in ROI with considerably better quality. A further analysis on the statistics of fixation positions also shows that the distribution of the fixation points is consistent with the weights of different regions on face images used in some other face recognition systems.
Fang Fang 0003, Zhiguo Ma, Laiyun Qing, Xilin Chen 0001, Wen Gao 0001
FG3
2008 Visual context representation using a combination of feature-driven and object-driven mechanisms
abstract
Visual context between objects is an important cue for object position perception. How to effectively represent the visual context is a key issue to study. Some past work introduced task-driven methods for object perception, which led a large coding quantity. This paper proposes an approach that incorporates feature-driven mechanism into object-driven context representation for object locating. As an example, the paper discusses how a neuronal network encodes the visual context between feature salient regions and human eye centers with as little coding quantity as possible. A group of experiments on efficiency of visual context coding and object searching are analyzed and discussed, which show that the proposed method decreases the coding quantity and improve the object searching accuracy effectively.
Lijuan Duan, Laiyun Qing, Xilin Chen 0001, Wen Gao 0001
IJCNN3
2008 Spatial relationship representation for visual object searching
Lijuan Duan, Laiyun Qing, Wen Gao 0001, Xilin Chen 0001, Yuan Yuan 0001
Neurocomputing3
2007 Monocular Tracking 3D People By Gaussian Process Spatio-Temporal Variable Model
abstract
Tracking 3D people from monocular video is often poorly constrained. To mitigate this problem, prior knowledge should be exploited. In this paper, the Gaussian process spatio-temporal variable model (GPSTVM), a novel dynamical system modeling method is proposed for learning human pose and motion priors. The GPSTVM provides a low dimensional embedding of human motion data, with a smooth density function that provides higher probability to the poses and motions close to the training data. The low dimensional latent space is optimized directly to retain the spatio-temporal structure of the high dimensional pose space. After the prior on human pose is learned, the particle filtering can be used tracking articulated human pose; particle filtering propagates over time in the embedding space, avoiding the curse of dimensionality. Experiments demonstrate that our approach tracks 3D people accurately.
Junbiao Pang, Laiyun Qing, Qingming Huang, Shuqiang Jiang, Wen Gao 0001
ICIP (5)2
2007 A Fast Approach for Natural Image Matting using Structure Information
abstract
Natural image matting plays an important role in image and video editing. It has been addressed hotly because it is inherently an under-constrained problem - we must estimate both the foreground and background colors (we call them component colors) at each pixel to calculate its opacity, according to the only known observed color. Prior assumption such as statistics and smoothness are utilized for estimation, but these methods either are simple to handle situations such as complex background and large interaction region or have high computational complexity. This paper proposes a fast technique to estimate the component colors based on structure information. Our approach exploits a simple convolution operation to detect structure information in images. It then uses two kinds of estimation methods to propagate colors based on the structure types. Experimental results show that our method is fast and efficient to handle objects with strong structures and large part of interaction region with the background.
Qianhui Ning, Weiqiang Wang 0001, Caifeng Zhu, Laiyun Qing, Qingming Huang
ICME4
2007 Learning and Memory of Spatial Relationship by a Neural Network with Sparse Features
abstract
Research on efficiency of learning and memory is very important for theoretic exploration and practical application. This paper gives a discussion on learning and memory of spatial relationships between initial positions and object positions by a neural network with sparse features. As an example, the paper discusses how the neural network learns the visual contexts between human eye centers and random initial positions surrounding the eye centers in images with as little memory as possible. Some sparse features are designed and distances between initial positions and the labeled eye centers in horizontal and vertical directions are learned and memorized respectively. Such a system could predict object positions from a new initial position according to the contexts that the neural network learned. A group of experiments on efficiency of learning and memory with sparse features in several single and integrated scales are analyzed and discussed.
Lijuan Duan, Laiyun Qing, Wen Gao 0001, Yiqiang Chen 0001
IJCNN3
2007 Searching Eye Centers Using a Context-Based Neural Network
Laiyun Qing, Lijuan Duan, Wen Gao 0001
ISNN (2)2
2005 Empirical comparisons of several preprocessing methods for illumination insensitive face recognition
abstract
Illumination variation is one of the bottlenecks of face recognition systems. Many approaches to coping with illumination variations have been proposed. They can be categorized into model-based and preprocessing-based. Although the model-based approaches seem better in theory, they commonly introduce more constraints, which make them not practical enough for real applications. On the other hand, the preprocessing approaches commonly exploit simple and efficient image processing techniques. Typical approaches based on image processing include histogram equalization (HE), histogram specification (HS), logarithm transform (Log), gamma intensity correction (GIC), and self-quotient image (SQI). We have performed extensive experiments to analyze and compare these methods empirically by evaluating them on three large-scale face databases: CMU-PIE database, FERET database; CAS-PEAL database. Our experimental results show that HE, HS and GIC can improve recognition performance for images both with and without illumination variation, while Log and SQI may decrease the recognition rate for face images without much illumination variation, though they may facilitate the recognition of face images with illumination variation.
Bo Du 0007, Shiguang Shan, Laiyun Qing, Wen Gao 0001
ICASSP (2)3
2005 Face recognition under generic illumination based on harmonic relighting
abstract
The performances of the current face recognition systems suffer heavily from the variations in lighting. To deal with this problem, this paper presents an illumination normalization approach by relighting face images to a canonical illumination based on the harmonic images model. Benefiting from the observations that human faces share similar shape, and the albedos of the face surfaces are quasi-constant, we first estimate the nine low-frequency components of the illumination from the input facial image. The facial image is then normalized to the canonical illumination by re-rendering it using the illumination ratio image technique. For the purpose of face recognition, two kinds of canonical illuminations, the uniform illumination and a frontal flash with the ambient lights, are considered, among which the former encodes merely the texture information, while the latter encodes both the texture and shading information. Our experiments on the CMU-PIE face database and the Yale B face database have shown that the proposed relighting normalization can significantly improve the performance of a face recognition system when the probes are collected under varying lighting conditions.
Laiyun Qing, Shiguang Shan, Wen Gao 0001, Bo Du 0007
Int. J. Pattern Recognit. Artif. Intell.1
2004 Face relighting for face recognition under generic illumination
abstract
The performance of current face recognition systems suffers heavily from variations in lighting. To deal with this problem, this paper presents a novel illumination normalization approach, by relighting faces to a canonical illumination, based on harmonic images. Benefiting from the observation that human faces share similar shape, and the albedo of the face surface is quasi-constant, we first estimate the nine low-frequency components of the lighting from the input face image. Then, the face image is normalized by relighting it to a canonical illumination, based on the illumination ratio image. For face recognition purposes, two kinds of canonical illumination, uniform and frontal point lighting, are considered, between which the former encodes merely texture information, while the latter encodes both texture and shading information. Our experimental results show that the proposed relighting normalization can significantly improve the performance of a face recognition system.
Laiyun Qing, Shiguang Shan, Xilin Chen 0001
ICASSP (5)1
2003 Illumination Invariant Shot Boundary Detection
Laiyun Qing, Weiqiang Wang 0001, Wen Gao 0001
IDEAL1