Yingying Chen 0003

dblp:18/2343-3 · DBLP profile ↗
← Back
55ranked-venue papers
5as first author
30since 2021 · last 2026
0000-0002-5049-8092ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 45 · 5 first-author · 25 since 2021Artificial intelligence and machine learning · 26 · 15 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 AnomalyMoE: Towards a Language-free Generalist Model for Unified Visual Anomaly Detection
abstract
Anomaly detection is a critical task across numerous domains and modalities, yet existing methods are often highly specialized, limiting their generalizability. These specialized models, tailored for specific anomaly types like textural defects or logical errors, typically exhibit limited performance when deployed outside their designated contexts. To overcome this limitation, we propose AnomalyMoE, a novel and universal anomaly detection framework based on a Mixture-of-Experts (MoE) architecture. Our key insight is to decompose the complex anomaly detection problem into three distinct semantic hierarchies: local structural anomalies, component-level semantic anomalies, and global logical anomalies. AnomalyMoE correspondingly employs three dedicated expert networks at the patch, component, and global levels, and is specialized in reconstructing features and identifying deviations at its designated semantic level. This hierarchical design allows a single model to concurrently understand and detect a wide spectrum of anomalies. Furthermore, we introduce an Expert Information Repulsion (EIR) module to promote expert diversity and an Expert Selection Balancing (ESB) module to ensure the comprehensive utilization of all experts. Experiments on 8 challenging datasets spanning industrial imaging, 3D point clouds, medical imaging, video surveillance, and logical anomaly detection demonstrate that AnomalyMoE establishes new state-of-the-art performance, significantly outperforming specialized methods in their respective domains.
Zhaopeng Gu, Bingke Zhu, Guibo Zhu, Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang
AAAI4
2026 Quality-Aware Language-Conditioned Local Auto-Regressive Anomaly Synthesis and Detection
abstract
Despite substantial progress in anomaly synthesis, existing diffusion-based and coarse inpainting pipelines commonly suffer from structural deficiencies such as micro-structural discontinuities, limited semantic controllability, and inefficient generation. To overcome these limitations, we introduce ARAS, a language-conditioned, auto-regressive anomaly synthesis approach that precisely injects local, text-specified defects into normal images via token-anchored latent editing. Leveraging a hard-gated auto-regressive operator and a training-free, context-preserving masked sampling kernel, ARAS significantly enhances defect realism, preserves fine-grained material textures, and provides continuous semantic control over synthesized anomalies. Integrated within our Quality-Aware Re-weighted Anomaly Detection (QARAD) framework, we propose a dynamic weighting strategy that emphasizes high-quality synthetic samples by computing an image-text similarity score with a dual-encoder model. Extensive experiments across three datasets, MVTec AD, VisA, and BTAD, demonstrate that our QARAD outperforms SOTA methods in both image- and pixel-level anomaly detection tasks, achieving improved accuracy, robustness, and a 5× synthesis speedup compared to diffusion-based alternatives.
Bingke Zhu, Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang
AAAI3
2026 MiMMamba: A Motion in Motion Mamba network for human motion forecasting
Yingying Chen 0003, Jinqiao Wang
Comput. Vis. Image Underst.2
2026 Compositional Gamba for 3D human pose estimation
Yingying Chen 0003, Jinqiao Wang
Image Vis. Comput.2
2026 FiLo++: Zero-/Few-Shot Anomaly Detection by Fused Fine-Grained Descriptions and Deformable Localization
abstract
Anomaly detection methods typically require extensive normal samples from the target class for training, limiting their applicability in scenarios that require rapid adaptation, such as cold start. Zero-shot and few-shot anomaly detection do not require labeled samples from the target class in advance, making them a promising research direction. Existing zero-shot and few-shot approaches often leverage powerful multimodal models to detect and localize anomalies by comparing image-text similarity. However, their handcrafted generic descriptions fail to capture the diverse range of anomalies that may emerge in different objects, and simple patch-level image-text matching often struggles to localize anomalous regions of varying shapes and sizes. To address these issues, this paper proposes the FiLo++ method, which consists of two key components. The first component, Fused Fine-Grained Descriptions (FusDes), utilizes large language models to generate anomaly descriptions for each object category, combines both fixed and learnable prompt templates and applies a runtime prompt filtering method, producing more accurate and task-specific textual descriptions. The second component, Deformable Localization (DefLoc), integrates the vision foundation model Grounding DINO with position-enhanced text descriptions and a Multi-scale Deformable Cross-modal Interaction (MDCI) module, enabling accurate localization of anomalies with various shapes and sizes. In addition, we design a position-enhanced patch matching approach to improve few-shot anomaly detection performance. Experiments on multiple datasets demonstrate that FiLo++ achieves significant performance improvements compared with existing methods. Code will be available at https://github.com/CASIA-IVA-Lab/FiLo.
Zhaopeng Gu, Bingke Zhu, Guibo Zhu, Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang
IEEE Trans. Circuits Syst. Video Technol.4
2026 Adapting CLIP for 3D Human Pose Estimation
abstract
Two-stage 3D human pose estimation has garnered remarkable progress, whereas ill-posed issue still poses great deterioration to the performance. During the 2D-3D projection, comprehension of the human topology serves as a significant cue to remove the ambiguities. However, limited feature representation restricts the grasp of human knowledge. In this light, we advocate for enriching the feature representation via integrating information from different sources. Contrastive imagetext pretrained models like CLIP demonstrate the capability of generalized representation. Exploiting the strong representation power of these models improves the performance of massive downstream tasks. Inspired by this, we resort to CLIP and delve into the integration of multi-source information with the distillation technique to reduce the location error of the regression task. Firstly, we endeavor to adapt the relation knowledge from CLIP text encoder to the transformer block of joint relation modeling module. Embedding of external language knowledge enriches the feature with the tailored adjuster and conveys the relation information from a quite different aspect. Secondly, we employ the CLIP visual encoder as another source to attain the knowledge from the vision aspect.We transfer the discrete human joints into the sketch image where human skeleton is depicted on a white board. Afterwards, the sketch drawing serves as the input of CLIP visual encoder, leading to a full understanding of human anatomy. We inject this representation into 3D human pose estimation network with the distillation technique where a dual-path distillation module is advanced to retain multimodal information. Extensive experiments on Human3.6M and MPIINF- 3DHP showcase that proposed method achieves competitive results over the previous approaches and improves the location precision, revealing the efficacy of our approach.
Yingying Chen 0003, Jinqiao Wang
IEEE Trans. Multim.2
2025 UniVAD: A Training-free Unified Model for Few-shot Visual Anomaly Detection
abstract
Visual Anomaly Detection (VAD) aims to identify abnormal samples in images that deviate from normal patterns, covering multiple domains, including industrial, logical, and medical fields. Due to the domain gaps between these fields, existing VAD methods are typically tailored to each domain, with specialized detection techniques and model architectures that are difficult to generalize across different domains. Moreover, even within the same domain, current VAD approaches often require large amounts of normal samples to train class-specific models, resulting in poor generalizability and hindering unified evaluation across domains. To address this issue, we propose a generalized few-shot VAD method, UniVAD, capable of detecting anomalies across various domains, with a training-free unified model. UniVAD only needs few normal samples as references during testing to detect anomalies in previously unseen objects, without training on the specific domain. Specifically, UniVAD employs a Contextual Component Clustering (C3) module based on clustering and vision foundation models to segment components within the image accurately, and leverages Component-Aware Patch Matching (CAPM) and Graph-Enhanced Component Modeling (GECM) modules to detect anomalies at different semantic levels, which are aggregated to produce the final detection result. We conduct experiments on nine datasets spanning industrial, logical, and medical fields, and the results demonstrate that UniVAD achieves state-of-the-art performance in few-shot anomaly detection tasks across multiple domains, outperforming domain-specific anomaly detection models. Code is available at https://github.com/FantasticGNU/UniVAD.
Zhaopeng Gu, Bingke Zhu, Guibo Zhu, Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang
CVPR4
2025 LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing
abstract
Audio-visual video parsing focuses on classifying videos through weak labels while identifying events as either visible, audible, or both, alongside their respective temporal boundaries. Many methods ignore that different modalities often lack alignment, thereby introducing extra noise during modal interaction. In this work, we introduce a Learning Interaction method for Non-aligned Knowledge (LINK), designed to equilibrate the contributions of distinct modalities by dynamically adjusting their input during event prediction. Additionally, we leverage the semantic information of pseudo-labels as a priori knowledge to mitigate noise from other modalities. Our experimental findings demonstrate that our model outperforms existing methods on the LLP dataset.
Langyu Wang, Bingke Zhu, Yingying Chen 0003, Jinqiao Wang
ICASSP3
2025 MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing
abstract
The weakly-supervised audio-visual video parsing (AVVP) aims to predict all modality-specific events and locate their temporal boundaries. Despite significant progress, due to the limitations of the weakly-supervised and the deficiencies of the model architecture, existing methods are lacking in simultaneously improving both the segment-level prediction and the event-level prediction. In this work, we propose a audio-visual Mamba network with pseudo labeling aUGmentation (MUG) for emphasising the uniqueness of each segment and excluding the noise interference from the alternate modalities. Specifically, we annotate some of the pseudo-labels based on previous work. Using unimodal pseudo-labels, we perform cross-modal random combinations to generate new data, which can enhance the model's ability to parse various segment-level event combinations. For feature processing and interaction, we employ a audio-visual mamba network. The AV-Mamba enhances the ability to perceive different segments and excludes additional modal noise while sharing similar modal information. Our extensive experiments demonstrate that MUG improves state-of-the-art results on LLP dataset in all metrics (e.g,, gains of 2.1% and 1.2% in terms of visual Segment-level and audio Segment-level metrics). Our code is available at https://github.com/WangLY136/MUG.
Langyu Wang, Bingke Zhu, Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang
ICCV3
2025 FLARE: A Framework for Stellar Flare Forecasting Using Stellar Physical Properties and Historical Records
abstract
Stellar flare events are critical observational samples for astronomical research; however, recorded flare events remain limited. Stellar flare forecasting can provide additional flare event samples to support research efforts. Despite this potential, no specialized models for stellar flare forecasting have been proposed to date. In this paper, we present extensive experimental evidence demonstrating that both stellar physical properties and historical flare records are valuable inputs for flare forecasting tasks. We then introduce FLARE (Forecasting Light-curve-based Astronomical Records via features Ensemble), the first-of-its-kind large model specifically designed for stellar flare forecasting. FLARE integrates stellar physical properties and historical flare records through a novel Soft Prompt Module and Residual Record Fusion Module. Experiments on the Kepler light curve dataset demonstrate that FLARE achieves superior performance compared to other methods across all evaluation metrics. Finally, we validate the forecast capability of our model through a comprehensive case study.
Bingke Zhu, Minghui Jia, Yihan Tao, A-Li Luo, Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang
IJCAI7
2025 DSTA: Reinforcing Vision-Language Understanding for Scene-Text VQA With Dual-Stream Training Approach
abstract
Scene-Text Visual Question Answering (STVQA) is a comprehensive task that requires reading and understanding the text in images to answer the question. Existing methods of exploring the vision-language relationships between questions, images, and scene text have achieved impressive results. However, these studies heavily rely on auxiliary modules, such as external OCR systems and object detection networks, making the question-answering process cumbersome and highly dependent. In addition, OCR text is treated as textual content only in these approaches, while its visual learning is ignored. To alleviate the above problems, we propose a novel end-to-end dual-stream multi-loss training approach called DSTA. Our model first integrates a text spotter into multimodal learning to incorporate overall textual and visual OCR features. Specifically, we propose a novel dual-stream multi-loss training strategy that improves multimodal understanding while training question-answering. In addition, we design OCR Contrastive Learning (OCL) to enhance vision-language understanding by exploring the multimodal features of OCR text in depth. Experiments show that DSTA outperforms previous state-of-the-art methods on two STVQA benchmarks without any extra training data.
Yingtao Tan, Yingying Chen 0003, Jinqiao Wang
IEEE Signal Process. Lett.2
2025 AMITA: Attribute-Guided Masked Image-Text Alignment for Multi-Label Image Representation
abstract
Multi-label image classification, which involves recognizing multiple objects within a single image, is a fundamental task in computer vision. Recently, Visual-Language Models (VLMs) have made remarkable progress in this area. Many approaches combine textual and visual modalities to understand the entire image. In this paper, we find that there is a direct correlation between the accurate localization of objects and the accuracy of multi-label classification. However, previous research methods did not specifically address localization accuracy, resulting in sub-optimal accuracy. Therefore, we propose the AMITA, namely Attribute-guided Masked Image-Text Alignment for multi-label image representation. AMITA improves localization accuracy by segmenting object masks, thereby enhancing the accuracy of multi-label image classification. Additionally, AMITA introduces an AutoFocus method to handle the localization problem of small objects. AutoFocus conducts recognition by resizing and cropping the image respectively, and automatically selects the images useful for the classification target. Moreover, AMITA incorporates Attribute-guided Prompting to strengthen the semantic distinction among different categories. It uses large language models to obtain the attributes of different categories and carefully designs prompts to enhance the attribute differences among different categories. Finally, extensive experiments on three popular datasets, including MS-COCO, Pascal VOC 2007, and NUS-WIDE, demonstrate the superiority of AMITA.
Jinyi Fang, Bingke Zhu, Jingling Yuan, Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang
IEEE Trans. Circuits Syst. Video Technol.4
2025 Optimization of Prompt Learning via Multi-Knowledge Representation for Vision-Language Models
abstract
Vision-language models (VLMs), such as CLIP, play a foundational role in various cross-modal applications. To fully leverage the potential of VLMs in adapting to downstream tasks, context optimization methods such as prompt tuning are essential. However, one key limitation is the lack of diversity in prompt templates, whether they are hand-crafted or learned through additional modules. This limitation restricts the capabilities of pretrained VLMs and can result in incorrect predictions in downstream tasks. To address this challenge, we propose context optimization with multi-knowledge representation (CoKnow), a framework that enhances prompt learning for VLMs with rich contextual knowledge. To facilitate CoKnow during inference, we train lightweight semantic knowledge mappers, which are capable of generating multi-knowledge representations for an input image without requiring additional priors. Experimentally, we conduct extensive experiments on 11 publicly available datasets, demonstrating that CoKnow outperforms a series of previous methods.
Enming Zhang, Bingke Zhu, Yingying Chen 0003, Qinghai Miao, Ming Tang 0001, Jinqiao Wang
IEEE Trans. Multim.3
2024 AnomalyGPT: Detecting Industrial Anomalies Using Large Vision-Language Models
abstract
Large Vision-Language Models (LVLMs) such as MiniGPT-4 and LLaVA have demonstrated the capability of understanding images and achieved remarkable performance in various visual tasks. Despite their strong abilities in recognizing common objects due to extensive training datasets, they lack specific domain knowledge and have a weaker understanding of localized details within objects, which hinders their effectiveness in the Industrial Anomaly Detection (IAD) task. On the other hand, most existing IAD methods only provide anomaly scores and necessitate the manual setting of thresholds to distinguish between normal and abnormal samples, which restricts their practical implementation. In this paper, we explore the utilization of LVLM to address the IAD problem and propose AnomalyGPT, a novel IAD approach based on LVLM. We generate training data by simulating anomalous images and producing corresponding textual descriptions for each image. We also employ an image decoder to provide fine-grained semantic and design a prompt learner to fine-tune the LVLM using prompt embeddings. Our AnomalyGPT eliminates the need for manual threshold adjustments, thus directly assesses the presence and locations of anomalies. Additionally, AnomalyGPT supports multi-turn dialogues and exhibits impressive few-shot in-context learning capabilities. With only one normal shot, AnomalyGPT achieves the state-of-the-art performance with an accuracy of 86.1%, an image-level AUC of 94.1%, and a pixel-level AUC of 95.3% on the MVTec-AD dataset.
Zhaopeng Gu, Bingke Zhu, Guibo Zhu, Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang
AAAI4
2024 FiLo: Zero-Shot Anomaly Detection by Fine-Grained Description and High-Quality Localization
abstract
Zero-shot anomaly detection (ZSAD) methods detect anomalies without prior access to known normal or abnormal samples within target categories. Existing methods typically rely on pretrained multimodal models, computing similarities between manually crafted textual features representing ''normal'' or ''abnormal'' semantics and image patch features to detect anomalies. However, the generic descriptions of ''abnormal'' often fail to precisely match diverse types of anomalies across different object categories. Additionally, computing feature similarities for single patches struggles to pinpoint specific locations of anomalies with various sizes and scales. To address these issues, we propose a novel ZSAD method called FiLo, comprising two components: adaptively learned Fine-Grained Description (FG-Des) and position-enhanced High-Quality Localization (HQ-Loc). FG-Des introduces fine-grained anomaly descriptions for each category using Large Language Models (LLMs) and employs adaptively learned textual templates to enhance the accuracy and interpretability of anomaly detection. HQ-Loc, utilizing Grounding DINO for preliminary localization, position-enhanced text prompts, and Multi-scale Multi-shape Cross-modal Interaction (MMCI) module, facilitates more accurate localization of anomalies of different sizes and shapes. Experimental results on datasets like MVTec and VisA demonstrate that FiLo significantly improves the performance of ZSAD in both detection and localization, achieving state-of-the-art performance with an image-level AUC of 83.9% and a pixel-level AUC of 95.9% on the VisA dataset. Code is available at https://github.com/CASIA-IVA-Lab/FiLo.
Zhaopeng Gu, Bingke Zhu, Guibo Zhu, Yingying Chen 0003, Hao Li 0115, Ming Tang 0001, Jinqiao Wang
ACM Multimedia4
2024 SlowFastFormer for 3D human pose estimation
Yingying Chen 0003, Jinqiao Wang
Comput. Vis. Image Underst.2
2024 EFCPose: End-to-End Multi-Person Pose Estimation With Fully Convolutional Heads
abstract
Mainstream methods of multi-person pose estimation are not end-to-end. Recently, some methods build an end-to-end framework based on the DETR framework, aiming to eliminate the need for hand-crafted modules like heuristic grouping and NMS post-processing. However, these DETR-based methods suffer from a heavy memory burden of processing the high-resolution backbone feature maps with transformers. In this paper, we propose an end-to-end multi-person pose estimation method with a fully convolutional network, termed EFCPose. Different from DETR-based methods, it directly predicts instance-aware poses in a pixel-wise manner with lightweight convolutional heads, avoiding the heavy memory burden. Overall, our method adopts the center-offset formulation and a one-to-one label assignment strategy to achieve the multi-person pose estimation in an end-to-end manner. The main contribution of our fully convolutional heads includes two aspects. On the one hand, we propose an unaligned center-offset representation to learn more reliable semantic centers to replace the inconsistent geometric centers, improving the performance of instance detection. On the other hand, we propose a novel regression strategy named limb-aware adaptive regression, which leverages separate adaptive points to convert challenging long-range offsets into simplified short-range offsets and incorporates limb constraints to elevate the regression quality of joint offsets. Compared with current DETR-based end-to-end methods, EFCPose avoids high computational complexity and achieves higher accuracy. Extensive experiments on COCO Keypoint and CrowdPose benchmarks show that EFCPose outperforms other state-of-the-art bottom-up and single-stage methods without flipping augmentation.
Yingying Chen 0003, Zhiyang Chen 0002, Ming Tang 0001, Jinqiao Wang
IEEE Trans. Circuits Syst. Video Technol.3
2024 Dual-Path Transformer for 3D Human Pose Estimation
abstract
Video-based 3D human pose estimation has achieved great progress, however, it is still difficult to learn precise 2D-3D projection under some hard cases. Multi-level human knowledge and motion information serve as two key elements in the field to conquer the challenges caused by various factors, where the former encodes various human structure information spatially and the latter captures the motion change temporally. Inspired by this, we propose a DualFormer (dual-path transformer) network which encodes multiple human contexts and motion detail to perform the spatial-temporal modeling. Firstly, motion information which depicts the movement change of human body is embedded to provide explicit motion prior for the transformer module. Secondly, a dual-path transformer framework is proposed to model long-range dependencies of both joint sequence and limb sequence. Parallel context embedding is performed initially and a cross transformer block is then appended to promote the interaction of the dual paths which improves the feature robustness greatly. Specifically, predictions of multiple levels can be acquired simultaneously. Lastly, we employ the weighted distillation technique to accelerate the convergence of the dual-path framework. We conduct extensive experiments on three different benchmarks, i.e., Human 3.6M, MPI-INF-3DHP and HumanEva-I. We mainly compute the MPJPE, P-MPJPE, PCK and AUC to evaluate the effectiveness of proposed approach and our work achieves competitive results compared with state-of-the-art approaches. Specifically, the MPJPE is reduced to 42.8mm which is 1.5mm lower than PoseFormer on Human3.6M, which proves the efficacy of the proposed approach.
Yingying Chen 0003, Jinqiao Wang
IEEE Trans. Circuits Syst. Video Technol.2
2023 Explicit Attention Modeling for Pedestrian Attribute Recognition
abstract
Recent studies on pedestrian attribute recognition have achieved significant improvements by utilizing complex networks and attention mechanisms. However, most of these studies learn the attention map implicitly through the class activation map. In this paper, we propose an explicit attention modeling approach for pedestrian attribute recognition. We construct a mask branch to learn the attention maps with a lightweight feature pyramid network. The features inside the specific mask are then averaged to obtain the scores for attribute recognition. Additionally, we introduce spatial and semantic distillation to improve the consistency of attention masks and attribute scores. Our experiments demonstrate that the proposed explicit attention modeling can achieve state-of-the-art performance on PA100K, PETA, and PAR datasets with negligible parameters.
Jinyi Fang, Bingke Zhu, Yingying Chen 0003, Jinqiao Wang, Ming Tang 0001
ICME3
2023 Uncertainty-Aware Boundary Attention Network for Real-Time Semantic Segmentation
Yuanbing Zhu, Bingke Zhu, Yingying Chen 0003, Jinqiao Wang
PRCV (3)3
2023 Human Parsing With Part-Aware Relation Modeling
abstract
In this paper, a Part-aware Relation Modeling (PRM) is developed to handle the task of human parsing. For pixel-level recognition, it is essential to generate features with adaptive context for various sizes and shapes of human parts. To address the issue, we adaptively capture contexts based on the part-aware relation mechanism. PRM mainly consists of three modules, including a part class module, a part-relation aggregation module, and a part-relation dispersion module. The part class module selectively enhances spatial details of the high-level features to obtain enhanced original features, and then extracts the high-level representations of every human part from a categorical perspective. The part-relation aggregation module is developed to extract the representative global context by exploring associated semantics of human parts, adaptively augmenting the context for human parts. The part-relation dispersion module is designed to generate the discriminative and effective local context and neglect the distracting one by making the affinity of human parts disperse. It ensures that features of the same class will be close to each other and away from those of different classes. By fusing the outputs of the two part-relation modules and the first outputs of the part class module, our PRM produces adaptive contextual features for diverse sizes of human parts, boosting the parsing accuracy. Extensive experiments are conducted to validate the effectiveness of our network, and a new state-of-the-art segmentation performance is achieved on three challenging human parsing datasets,i.e., PASCAL-Person-Part, LIP, and CIHP. PRM is also extended to other tasks like animal parsing, and exhibits its generality.
Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang, Xiangyu Zhu 0001, Zhen Lei 0001
IEEE Trans. Multim.2
2022 UniVIP: A Unified Framework for Self-Supervised Visual Pre-training
abstract
Self-supervised learning (SSL) holds promise in leveraging large amounts of unlabeled data. However, the success of popular SSL methods has limited on single-centric-object images like those in ImageNet and ignores the correlation among the scene and instances, as well as the semantic difference of instances in the scene. To address the above problems, we propose a Unified Self-supervised Visual Pre-training (UniVIP), a novel self-supervised framework to learn versatile visual representations on either single-centric-object or non-iconic dataset. The framework takes into account the representation learning at three levels: 1) the similarity of scene-scene, 2) the correlation of scene-instance, 3) the discrimination of instance-instance. During the learning, we adopt the optimal transport algorithm to automatically measure the discrimination of instances. Massive experiments show that Uni-VIP pre-trained on non-iconic COCO achieves state-of-the-art transfer performance on a variety of downstream tasks, such as image classification, semi-supervised learning, object detection and segmentation. Furthermore, our method can also exploit single-centric-object dataset such as ImageNet and outperforms BYOL by 2.5% with the same pre-training epochs in linear probing, and surpass current self-supervised object detection methods on COCO dataset, demonstrating its universality and potential.
Zhaowen Li, Yousong Zhu, Fan Yang 0089, Wei Li 0314, Chaoyang Zhao, Yingying Chen 0003, Zhiyang Chen 0002, Jiahao Xie 0002, Rui Zhao 0001, Ming Tang 0001, Jinqiao Wang
CVPR6
2022 C2AM Loss: Chasing a Better Decision Boundary for Long-Tail Object Detection
abstract
Long-tail object detection suffers from poor performance on tail categories. We reveal that the real culprit lies in the extremely imbalanced distribution of the classifier's weight norm. For conventional softmax cross-entropy loss, such imbalanced weight norm distribution yields ill conditioned decision boundary for categories which have small weight norms. To get rid of this situation, we choose to maxi-mize the cosine similarity between the learned feature and the weight vector of target category rather than the inner-product of them. The decision boundary between any two categories is the angular bisector of their weight vectors. Whereas, the absolutely equal decision boundary is sub-optimal because it reduces the model's sensitivity to vari-ous categories. Intuitively, categories with rich data diver-sity should occupy a larger area in the classification space while categories with limited data diversity should occupy a slightly small space. Hence, we devise a Category-Aware Angular Margin Loss (C2AM Loss) to introduce an adaptive angular margin between any two categories. Specif-ically, the margin between two categories is proportional to the ratio of their classifiers' weight norms. As a result, the decision boundary is slightly pushed towards the cat-egory which has a smaller weight norm. We conduct comprehensive experiments on LVIS dataset. C2AM Loss brings 4.9~5.2 AP improvements on different detectors and back-bones compared with baseline.
Tong Wang 0015, Yousong Zhu, Yingying Chen 0003, Chaoyang Zhao, Jinqiao Wang, Ming Tang 0001
CVPR3
2022 Regularizing Vector Embedding in Bottom-Up Human Pose Estimation
Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang
ECCV (6)3
2022 When Skeleton Meets Appearance: Adaptive Appearance Information Enhancement for Skeleton Based Action Recognition
abstract
Skeleton-based action recognition methods which utilize graph convolution networks (GCNs) have achieved remark-able success in recent years. However, action recognizer can be easily confused by the ambiguity caused by different actions with similar skeleton sequences when only skeleton data is trained. Introducing appearance information can effectively eliminate the ambiguity. Based on this, we introduce a two-stream network for action recognition. One trained on RGB images extracts appearance information. The other trained on skeleton data models motion information and adaptively captures appearance information of action areas at action-related intervals via a specially tailored attention mechanism. Our architecture is trained and evaluated on two large-scale datasets: NTU RGB+D and NTU RGB+D 120, and a small scale human-object interaction dataset Northwestern-UCLA. Experiment results verify the effectiveness of our method and the performance of our method exceeds the state-of-the-art with a significant margin.
Suqin Wang, Yingying Chen 0003, Jiangtao Huo, Jinqiao Wang
ICME3
2022 Grammar-Induced Wavelet Network for Human Parsing
abstract
Most existing methods of human parsing still face a challenge: how to extract the accurate foreground from similar or cluttered scenes effectively. In this paper, we propose a Grammar-induced Wavelet Network (GWNet), to deal with the challenge. GWNet mainly consists of two modules, including a blended grammar-induced module and a wavelet prediction module. We design the blended grammar-induced module to exploit the relationship of different human parts and the inherent hierarchical structure of a human body by means of grammar rules in both cascaded and paralleled manner. In this way, conspicuous parts, which are easily distinguished from the background, can amend the segmentation of inconspicuous ones, improving the foreground extraction. We also design a Part-aware Convolutional Recurrent Neural Network (PCRNN) to pass messages which are generated by grammar rules. To further improve the performance, we propose a wavelet prediction module to capture the basic structure and the edge details of a person by decomposing the low-frequency and high-frequency components of features. The low-frequency component can represent the smooth structures and the high-frequency components can describe the fine details. We conduct extensive experiments to evaluate GWNet on PASCAL-Person-Part, LIP, and PPSS datasets. GWNet obtains state-of-the-art performance on these human parsing datasets.
Yingying Chen 0003, Ming Tang 0001, Zhen Lei 0001, Jinqiao Wang
IEEE Trans. Image Process.2
2021 Improving Multiple Object Tracking With Single Object Tracking
abstract
Despite considerable similarities between multiple object tracking (MOT) and single object tracking (SOT) tasks, modern MOT methods have not benefited from the development of SOT ones to achieve satisfactory performance. The major reason for this situation is that it is inappropriate and inefficient to apply multiple SOT models directly to the MOT task, although advanced SOT methods are of the strong discriminative power and can run at fast speeds.In this paper, we propose a novel and end-to-end trainable MOT architecture that extends CenterNet by adding an SOT branch for tracking objects in parallel with the existing branch for object detection, allowing the MOT task to benefit from the strong discriminative power of SOT methods in an effective and efficient way. Unlike most existing SOT methods which learn to distinguish the target object from its local backgrounds, the added SOT branch trains a separate SOT model per target online to distinguish the target from its surrounding targets, assigning SOT models the novel discrimination. Moreover, similar to the detection branch, the SOT branch treats objects as points, making its online learning efficient even if multiple targets are processed simultaneously. Without tricks, the proposed tracker achieves MOTAs of 0.710 and 0.686, IDF1s of 0.719 and 0.714, on MOT17 and MOT20 benchmarks, respectively, while running at 16 FPS on MOT17.
Linyu Zheng, Ming Tang 0001, Yingying Chen 0003, Guibo Zhu, Jinqiao Wang, Hanqing Lu
CVPR3
2021 Macro-micro mutual learning inside compositional model for human pose estimation
Yingying Chen 0003, Congqi Cao, Yakui Chu, Jinqiao Wang, Hanqing Lu
Neurocomputing2
2021 STN-enhanced message passing guided by adversarial learning for human pose estimation
Yingying Chen 0003, Congqi Cao, Jinqiao Wang, Hanqing Lu
Neurocomputing2
2021 Semi-Supervised Scene Text Recognition
abstract
Scene text recognition has been widely researched with supervised approaches. Most existing algorithms require a large amount of labeled data and some methods even require character-level or pixel-wise supervision information. However, labeled data is expensive, unlabeled data is relatively easy to collect, especially for many languages with fewer resources. In this paper, we propose a novel semi-supervised method for scene text recognition. Specifically, we design two global metrics, i.e., edit reward and embedding reward, to evaluate the quality of generated string and adopt reinforcement learning techniques to directly optimize these rewards. The edit reward measures the distance between the ground truth label and the generated string. Besides, the image feature and string feature are embedded into a common space and the embedding reward is defined by the similarity between the input image and generated string. It is natural that the generated string should be the nearest with the image it is generated from. Therefore, the embedding reward can be obtained without any ground truth information. In this way, we can effectively exploit a large number of unlabeled images to improve the recognition performance without any additional laborious annotations. Extensive experimental evaluations on the five challenging benchmarks, the Street View Text, IIIT5K, and ICDAR datasets demonstrate the effectiveness of the proposed approach, and our method significantly reduces annotation effort while maintaining competitive recognition performance.
Yunze Gao, Yingying Chen 0003, Jinqiao Wang, Hanqing Lu
IEEE Trans. Image Process.2
2020 Progressive Bi-C3D Pose Grammar for Human Pose Estimation
abstract
In this paper, we propose a progressive pose grammar network learned with Bi-C3D (Bidirectional Convolutional 3D) for human pose estimation. Exploiting the dependencies among the human body parts proves effective in solving the problems such as complex articulation, occlusion and so on. Therefore, we propose two articulated grammars learned with Bi-C3D to build the relationships of the human joints and exploit the contextual information of human body structure. Firstly, a local multi-scale Bi-C3D kinematics grammar is proposed to promote the message passing process among the locally related joints. The multi-scale kinematics grammar excavates different levels human context learned by the network. Moreover, a global sequential grammar is put forward to capture the long-range dependencies among the human body joints. The whole procedure can be regarded as a local-global progressive refinement process. Without bells and whistles, our method achieves competitive performance on both MPII and LSP benchmarks compared with previous methods, which confirms the feasibility and effectiveness of C3D in information interactions.
Yingying Chen 0003, Jinqiao Wang, Hanqing Lu
AAAI2
2020 Part-Aware Context Network for Human Parsing
abstract
Recent works have made significant progress in human parsing by exploiting rich contexts. However, human parsing still faces a challenge of how to generate adaptive contextual features for the various sizes and shapes of human parts. In this work, we propose a Part-aware Context Network (PCNet), a novel and effective algorithm to deal with the challenge. PCNet mainly consists of three modules, including a part class module, a relational aggregation module, and a relational dispersion module. The part class module extracts the high-level representations of every human part from a categorical perspective. We design a relational aggregation module to capture the representative global context by mining associated semantics of human parts, which adaptively augments the context for human parts. We propose a relational dispersion module to generate the discriminative and effective local context and neglect disturbing one by making the affinity of human parts dispersed. The relational dispersion module ensures that features in the same class will be close to each other and away from those of different classes. By fusing the outputs of the relational aggregation module, the relational dispersion module and the backbone network, our PCNet generates adaptive contextual features for various sizes of human parts, improving the parsing accuracy. We achieve a new state-of-the-art segmentation performance on three challenging human parsing datasets, i.e., PASCAL-Person-Part, LIP, and CIHP.
Yingying Chen 0003, Bingke Zhu, Jinqiao Wang, Ming Tang 0001
CVPR2
2020 Blended Grammar Network for Human Parsing
Yingying Chen 0003, Bingke Zhu, Jinqiao Wang, Ming Tang 0001
ECCV (24)2
2020 Learning Feature Embeddings for Discriminant Model Based Tracking
Linyu Zheng, Ming Tang 0001, Yingying Chen 0003, Jinqiao Wang, Hanqing Lu
ECCV (15)3
2020 Occlusion-Aware Siamese Network for Human Pose Estimation
Yingying Chen 0003, Yunze Gao, Jinqiao Wang, Hanqing Lu
ECCV (20)2
2020 High-Speed And Accurate Scale Estimation For Visual Tracking With Gaussian Process Regression
abstract
Recent years have seen remarkable progress in the visual tracking domain. However, it remains a challenging task to estimate the scale of target efficiently and accurately. In this paper, we present a novel and high-performance scale estimation approach for tracking-by-detection framework. The proposed approach, named GPAS, formulates the scale estimation as a Gaussian process regression problem based on scale pyramid representation. In general, it enjoys the following there advantages. (i) Efficient. It only takes 2ms to estimate the scale of a target on a single CPU. (ii) Accurate. Without bells and whistles, its accuracy surpasses all previous hand-crafted features based scale estimation methods by large margins. (iii) Generic. It can be incorporated into any tracking-by-detection framework based trackers easily. Experiment results show that compared to the latest and classical scale estimation method, fDSST, our GPAS significantly improves the performance by 6.2% in mean distance precision, 8.9% in mean overlap precision, and 5.5% in mean AUC on 28 sequences of OTB2013 with significant scale variations.
Linyu Zheng, Ming Tang 0001, Yingying Chen 0003, Jinqiao Wang, Hanqing Lu
ICME3
2020 Progressive rectification network for irregular text recognition
Yunze Gao, Yingying Chen 0003, Jinqiao Wang, Hanqing Lu
Sci. China Inf. Sci.2
2020 Semantic-spatial fusion network for human parsing
Yingying Chen 0003, Bingke Zhu, Jinqiao Wang, Ming Tang 0001
Neurocomputing2
2020 Siamese Deformable Cross-Correlation Network for Real-Time Visual Tracking
Linyu Zheng, Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang, Hanqing Lu
Neurocomputing2
2019 Gate-based Bidirectional Interactive Decoding Network for Scene Text Recognition
abstract
Scene text recognition has attracted rapidly increasing attention from the research community. Recent dominant approaches typically follow an attention-based encoder-decoder framework that uses a unidirectional decoder to perform decoding in a left-to-right manner, but ignoring equally important right-to-left grammar information. In this paper, we propose a novel Gate-based Bidirectional Interactive Decoding Network (GBIDN) for scene text recognition. Firstly, the backward decoder performs decoding from right to left and generates the reverse language context. After that, the forward decoder simultaneously utilizes the visual context from image encoder and the reverse language context from backward decoder through two attention modules. In this way, the bidirectional decoders perform effective interaction to fully fuse the bidirectional grammar information and further improve the decoding quality. Besides, in order to relieve the adverse effect of noises, we devise a gated context mechanism to adaptively make use of the visual context and reverse language context. Extensive experiments on various challenging benchmarks demonstrate the effectiveness of our method.
Yunze Gao, Yingying Chen 0003, Jinqiao Wang, Hanqing Lu
CIKM2
2019 Fast-deepKCF Without Boundary Effect
abstract
In recent years, correlation filter based trackers (CF trackers) have received much attention because of their top performance. Most CF trackers, however, suffer from low frame-per-second (fps) in pursuit of higher localization accuracy by relaxing the boundary effect or exploiting the high-dimensional deep features. In order to achieve real-time tracking speed while maintaining high localization accuracy, in this paper, we propose a novel CF tracker, fdKCF*, which casts aside the popular acceleration tool, i.e., fast Fourier transform, employed by all existing CF trackers, and exploits the inherent high-overlap among real (i.e., noncyclic) and dense samples to efficiently construct the kernel matrix. Our fdKCF* enjoys the following three advantages. (i) It is efficiently trained in kernel space and spatial domain without the boundary effect. (ii) Its fps is almost independent of the number of feature channels. Therefore, it is almost real-time, i.e., 24 fps on OTB-2015, even though the high-dimensional deep features are employed. (iii) Its localization accuracy is state-of-the-art. Extensive experiments on four public benchmarks, OTB-2013, OTB-2015, VOT2016, and VOT2017, show that the proposed fdKCF* achieves the state-of-the-art localization performance with remarkably faster speed than C-COT and ECO.
Linyu Zheng, Ming Tang 0001, Yingying Chen 0003, Jinqiao Wang, Hanqing Lu
ICCV3
2019 Pose-Weighted Gan for Photorealistic Face Frontalization
abstract
Face recognition methods have achieved high accuracy when faces are captured in frontal pose and constrained scenes. However, severe drop in accuracy is observed when large pose variations exist. The main reason is that the large yaw angle leads to ID information loss. In this paper, we intend to solve the large pose variations in a generation manner. Specifically, we propose a Pose-Weighted Generative Adversarial Network (PW-GAN) for photorealistic frontal view synthesis. We find frontalizing the faces in large poses (yaw angle larger than 60°) is so difficult that the results are not photorealistic and the ID information is lost. To simplify the problem, we first frontalize the face image through 3D face model, which is then used to guide the network predicting. Second, we refine the pose code in the loss function to make the network pay more attention to large poses. Quantitative and qualitative experimental results on the Multi-PIE and LFW demonstrate our method achieves state of the art.
Su-Fang Zhang, Qinghai Miao, Min Huang 0009, Xiangyu Zhu 0001, Yingying Chen 0003, Zhen Lei 0001, Jinqiao Wang
ICIP5
2019 Bi-Directional Message Passing Based Scanet for Human Pose Estimation
abstract
Articulated human pose estimation is one of the fundamental computer vision problems. In this paper, a Bi-directional Message Passing(BDMP) module is proposed to fuse convolutional features of different scales in the up-sampling process of the hourglass model for human pose estimation. Moreover, a novel module which integrates Spatial and Channelwise Attention Network(SCANet) is proposed to refine the features obtained from the message passing stage. We design a Semantics-aware Channel-wise Attention(SACWA) module to reduce the feature redundancy and enrich the semantic information simultaneously. A Sharper Spatial Attention(SSA) module based on the Gumbel-Softmax sampling is proposed to exclude the interference from cluttered background and overcomes the gradient degradation induced by the softmax normalization. The proposed framework achieves leading position on MPII benchmark against the state-of-the-arts methods with much less parameters.
Yingying Chen 0003, Jinqiao Wang, Ming Tang 0001, Hanqing Lu
ICME2
2019 Reading scene text with fully convolutional sequence modeling
Yunze Gao, Yingying Chen 0003, Jinqiao Wang, Ming Tang 0001, Hanqing Lu
Neurocomputing2
2019 Pixelwise Deep Sequence Learning for Moving Object Detection
abstract
Moving object detection is an essential, well-studied but still open problem in computer vision and plays a fundamental role in many applications. Traditional approaches usually reconstruct background images with hand-crafted visual features, such as color, texture, and edge. Due to lack of prior knowledge or semantic information, it is difficult to deal with complicated and rapid changing scenes. To exploit the temporal structure of the pixel-level semantic information, in this paper, we propose an end-to-end deep sequence learning architecture for moving object detection. First, the video sequences are input into a deep convolutional encoder-decoder network for extracting pixel-wise semantic features. Then, to exploit the temporal context, we propose a novel attention long short-term memory (Attention ConvLSTM) to model pixelwise changes over time. A spatial transformer network and a conditional random field layer are finally appended to reduce the sensitivity to camera motion and smooth the foreground boundaries. A multi-task loss is proposed to jointly optimization for frame-based classification and temporal prediction in an end-to-end network. Experimental results on CDnet 2014 and LASIESTA show 12.15% and 16.71% improvement to the state of the art, respectively.
Yingying Chen 0003, Jinqiao Wang, Bingke Zhu, Ming Tang 0001, Hanqing Lu
IEEE Trans. Circuits Syst. Video Technol.1
2018 Progressive Cognitive Human Parsing
abstract
Human parsing is an important task for human-centric understanding. Generally, two mainstreams are used to deal with this challenging and fundamental problem. The first one is employing extra human pose information to generate hierarchical parse graph to deal with human parsing task. Another one is training an end-to-end network with the semantic information in image level. In this paper, we develop an end-to-end progressive cognitive network to segment human parts. In order to establish a hierarchical relationship, a novel component-aware region convolution structure is proposed. With this structure, latter layers inherit prior component information from former layers and pay its attention to a finer component. In this way, we deal with human parsing as a progressive recognition task, that is, we first locate the whole human and then segment the hierarchical components gradually. The experiments indicate that our method has a better location capacity for the small objects and a better classification capacity for the large objects. Moreover, our framework can be embedded into any fully convolutional network to enhance the performance significantly.
Bingke Zhu, Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang
AAAI2
2018 Dense Chained Attention Network for Scene Text Recognition
abstract
Reading text in the wild is a challenging task in computer vision. Scene text suffers from various background noise, including shadow, irrelevant symbols and background texture. In order to reduce the disturbance of background noise, we propose a dense chained attention network with stacked attention modules for scene text recognition. Each attention module learns the attention map that is adapted to corresponding features to enhance the foreground text and suppress the background noise. Besides, the attention branch is designed with the convolution-deconvolution structure which rapidly captures global information to guide the discriminative feature selection. We stack multiple attention modules to gradually refine the attention maps and capture both the low-level appearance feature and the high-level semantic information. Extensive experiments on the standard benchmarks, the Street View Text, IIIT5K, and ICDAR datasets validate the superiority of the proposed method. The dense chained attention network achieves state-of-the-art or highly competitive recognition performance.
Yunze Gao, Yingying Chen 0003, Jinqiao Wang, Ming Tang 0001, Hanqing Lu
ICIP2
2018 Tree Hierarchical CNNs for Object Parsing
abstract
Object parsing is a challenging topic in computer vision, which is to distinguish all parts of visual objects. Although lots of works have been proposed, it is difficult to segment complicated objects from complex scenes. Therefore, in this paper we propose a tree hierarchical CNNs for object parsing. Rather than segment all parts of objects at once, we segment object parts step by step in a tree hierarchy and then merge the results together with a full convolutional network. In the tree hierarchy, the segmentation errors of the previous layers of the network outputs could be passed down to following layers and result in accumulated errors. In order to reduce the accumulated errors, we adopt a new part-aware fusion strategy, which fuses global-level feature maps from fully convolutional networks as well as the part-level object feature maps from the output of previous layer. It also contributes to improve the integrity and robustness of object parsing. Finally, the experiments on published datasets show the superiority of the proposed approach, especially for neighboring objects in complex scene.
Yingying Chen 0003, Bingke Zhu, Jinqiao Wang, Ming Tang 0001, Hanqing Lu
ICIP2
2017 Joint background reconstruction and foreground segmentation via a two-stage convolutional neural network
abstract
Foreground segmentation in video sequences is a classic topic in computer vision. Due to the lack of semantic and prior knowledge, it is difficult for existing methods to deal with sophisticated scenes well. Therefore, in this paper, we propose an end-to-end two-stage deep convolutional neural network (CNN) framework for foreground segmentation in video sequences. In the first stage, a convolutional encoder-decoder sub-network is employed to reconstruct the background images and encode rich prior knowledge of background scenes. In the second stage, the reconstructed background and current frame are input into a multi-channel fully-convolutional sub-network (MCFCN) for accurate foreground segmentation. In the two-stage CNN, the reconstruction loss and segmentation loss are jointly optimized. The background images and foreground objects are output simultaneously in an end-to-end way. Moreover, by incorporating the prior semantic knowledge of foreground and background in the pre-training process, our method could restrain the background noise and keep the integrity of foreground objects at the same time. Experiments on CDNet 2014 show that our method outperforms the state-of-the-art by 4.9%.
Xu Zhao 0003, Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang
ICME2
2017 Fast Deep Matting for Portrait Animation on Mobile Phone
abstract
Image matting plays an important role in image and video editing. However, the formulation of image matting is inherently ill-posed. Traditional methods usually employ interaction to deal with the image matting problem with trimaps and strokes, and cannot run on the mobile phone in real-time. In this paper, we propose a real-time automatic deep matting approach for mobile devices. By leveraging the densely connected blocks and the dilated convolution, a light full convolutional network is designed to predict a coarse binary mask for portrait image. And a feathering block, which is edge-preserving and matting adaptive, is further developed to learn the guided filter and transform the binary mask into alpha matte. Finally, an automatic portrait animation system based on fast deep matting is built on mobile devices, which does not need any interaction and can realize real-time matting with 15 fps. The experiments show that the proposed approach achieves comparable results with the state-of-the-art matting solvers.
Bingke Zhu, Yingying Chen 0003, Jinqiao Wang, Si Liu 0001, Bo Zhang 0069, Ming Tang 0001
ACM Multimedia2
2016 A unified model sharing framework for moving object detection
Yingying Chen 0003, Jinqiao Wang, Min Xu 0001, Xiangjian He, Hanqing Lu
Signal Process.1
2016 Adaptive Content Condensation Based on Grid Optimization for Thumbnail Image Generation
abstract
An ideal thumbnail generator should effectively condense unimportant regions and keep the important content undeformed, completed, and at a proper scale, i.e., accuracy, completeness, and sufficiency. Each retargeting method has its own advantage for resizing arbitrary images. However, they often ignore the completeness and sufficiency for information presentation in thumbnails. In this paper, we formulate thumbnail generation as an image content condensation problem and propose a unified grid optimization framework to fuse multiple operators. From the view of accuracy, completeness, and sufficiency for information presentation, we exploit complementary relationships among three condensation operators and fuse them into a unified grid-based convex programming problem, which could be solved simultaneously and efficiently through numerical optimization. Besides warping energy to preserve the geometric structure of important objects, we put forward two grid-based energy terms to keep the completeness of important objects and retain them at a proper size. Finally, an adaptive procedure is proposed to dynamically adjust the contribution of loss functions for achieving optimal content condensation. Both qualitative and quantitative comparison results demonstrate that the proposed method achieves an excellent tradeoff among accuracy, completeness, and sufficiency of information preservation. The experimental results show that our approach is obviously superior to the state-of-the-art techniques.
Jinqiao Wang, Yingying Chen 0003, Tao Mei 0001, Min Xu 0001, La Zhang, Hanqing Lu
IEEE Trans. Circuits Syst. Video Technol.3
2015 Multiple features based shared models for background subtraction
abstract
Background modeling is a fundamental problem in computer vision and usually as the first step for high-level applications. Pixel based approaches usually ignore the spatial coherence, while region based approaches are sensitive to region size and scene complexity. In this paper, we propose a robust background subtraction approach via multiple features based shared models. Each shared model is represented by a sequence of samples based on sample consensus. Each pixel dynamically searches a matched model around the neighborhood. This shared mechanism not only enhances the robustness for background noise and jitter but also significantly reduces the number of models and samples for each model. Besides, we concatenate color and texture features as multiple features according to the discriminability and complementarity, so that each pixel can find a proper model more easily. Finally, the shared models are updated by random selecting a pixel matched the model with an adaptive update rate. Experiments on ChangeDetection benchmark 2014 show that the proposed approach outperforms the state-of-the-art methods.
Yingying Chen 0003, Jinqiao Wang, Jianqiang Li 0002, Hanqing Lu
ICIP1
2015 Learning sharable models for robust background subtraction
abstract
Background modeling and subtraction is a classical topic in compute vision. Gaussian mixture modeling (GMM) is a popular choice for its capability of adaptation to background variations. Lots of improvements have been made to enhance the robustness by considering spatial consistency and temporal correlation. In this paper, we propose a sharable GMM based background subtraction approach. Firstly, a sharable mechanism is presented to model the many-to-one relationship between pixels and models. Each pixel dynamically searches the best matched model in the neighborhood. This kind of space-sharing way is robust to camera jitter, dynamic background, etc. Secondly, the sharable models are built for both background and foreground. The noises resulted by local small movements could be effectively eliminated through the background sharable models, while the integrity of moving objects is enhanced by the foreground sharable models, especially for small objects. Finally, each sharable model is updated through randomly selecting a pixel which matches this model. And a flexible mechanism is added for switching between background and foreground models. Experiments on ChangeDetection benchmark dataset demonstrate the effectiveness of our approach.
Yingying Chen 0003, Jinqiao Wang, Hanqing Lu
ICME1
2015 Mobile Media Thumbnailing
abstract
With the development of Multimedia and Internet techniques, massively increasing visual data, such as image and video, need to be shown and browsed as thumbnails in various digital display platforms, like PC, cell phone, etc. This demonstration presents a grid based adaptive media thumb-nailing approach to maximize user experience in mobile image and video browsing. After representative frame extraction by spectral clustering and salient region detection, we obtain thumbnails with three resizing operators: cropping, warping and scaling, and adaptively fuse them into a unified grid based convex programming problem which could be solved simultaneously and efficiently through numerical optimization. Extensive experiments and comparisons on HUAWEI Honor 6 and Samsung S5 demonstrate that the proposed method achieves an excellent information preservation for thumbnails in mobile devices.
Yingying Chen 0003, Jinqiao Wang, Jing Liu 0001, Hanqing Lu
ICMR1