Bingke Zhu

dblp:203/8553 · DBLP profile ↗
← Back
25ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0001-6429-1773ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 20 · 4 first-author · 14 since 2021Artificial intelligence and machine learning · 14 · 2 first-author · 10 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 AnomalyMoE: Towards a Language-free Generalist Model for Unified Visual Anomaly Detection
abstract
Anomaly detection is a critical task across numerous domains and modalities, yet existing methods are often highly specialized, limiting their generalizability. These specialized models, tailored for specific anomaly types like textural defects or logical errors, typically exhibit limited performance when deployed outside their designated contexts. To overcome this limitation, we propose AnomalyMoE, a novel and universal anomaly detection framework based on a Mixture-of-Experts (MoE) architecture. Our key insight is to decompose the complex anomaly detection problem into three distinct semantic hierarchies: local structural anomalies, component-level semantic anomalies, and global logical anomalies. AnomalyMoE correspondingly employs three dedicated expert networks at the patch, component, and global levels, and is specialized in reconstructing features and identifying deviations at its designated semantic level. This hierarchical design allows a single model to concurrently understand and detect a wide spectrum of anomalies. Furthermore, we introduce an Expert Information Repulsion (EIR) module to promote expert diversity and an Expert Selection Balancing (ESB) module to ensure the comprehensive utilization of all experts. Experiments on 8 challenging datasets spanning industrial imaging, 3D point clouds, medical imaging, video surveillance, and logical anomaly detection demonstrate that AnomalyMoE establishes new state-of-the-art performance, significantly outperforming specialized methods in their respective domains.
Zhaopeng Gu, Bingke Zhu, Guibo Zhu, Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang
AAAI2
2026 Quality-Aware Language-Conditioned Local Auto-Regressive Anomaly Synthesis and Detection
abstract
Despite substantial progress in anomaly synthesis, existing diffusion-based and coarse inpainting pipelines commonly suffer from structural deficiencies such as micro-structural discontinuities, limited semantic controllability, and inefficient generation. To overcome these limitations, we introduce ARAS, a language-conditioned, auto-regressive anomaly synthesis approach that precisely injects local, text-specified defects into normal images via token-anchored latent editing. Leveraging a hard-gated auto-regressive operator and a training-free, context-preserving masked sampling kernel, ARAS significantly enhances defect realism, preserves fine-grained material textures, and provides continuous semantic control over synthesized anomalies. Integrated within our Quality-Aware Re-weighted Anomaly Detection (QARAD) framework, we propose a dynamic weighting strategy that emphasizes high-quality synthetic samples by computing an image-text similarity score with a dual-encoder model. Extensive experiments across three datasets, MVTec AD, VisA, and BTAD, demonstrate that our QARAD outperforms SOTA methods in both image- and pixel-level anomaly detection tasks, achieving improved accuracy, robustness, and a 5× synthesis speedup compared to diffusion-based alternatives.
Bingke Zhu, Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang
AAAI2
2026 DICE: Disentangling Causal Evidence for multimodal textbook question answering via attentive embedding fusion
Bingke Zhu, Jinqiao Wang, Xiaolin Qin
Pattern Recognit.3
2026 TRIS: A multimodal and multitask framework for unifying text-image retrieval and referring image segmentation
Zengzhi Qian, Weide Kang, Bingke Zhu, Jinqiao Wang
Pattern Recognit. Lett.4
2026 FiLo++: Zero-/Few-Shot Anomaly Detection by Fused Fine-Grained Descriptions and Deformable Localization
abstract
Anomaly detection methods typically require extensive normal samples from the target class for training, limiting their applicability in scenarios that require rapid adaptation, such as cold start. Zero-shot and few-shot anomaly detection do not require labeled samples from the target class in advance, making them a promising research direction. Existing zero-shot and few-shot approaches often leverage powerful multimodal models to detect and localize anomalies by comparing image-text similarity. However, their handcrafted generic descriptions fail to capture the diverse range of anomalies that may emerge in different objects, and simple patch-level image-text matching often struggles to localize anomalous regions of varying shapes and sizes. To address these issues, this paper proposes the FiLo++ method, which consists of two key components. The first component, Fused Fine-Grained Descriptions (FusDes), utilizes large language models to generate anomaly descriptions for each object category, combines both fixed and learnable prompt templates and applies a runtime prompt filtering method, producing more accurate and task-specific textual descriptions. The second component, Deformable Localization (DefLoc), integrates the vision foundation model Grounding DINO with position-enhanced text descriptions and a Multi-scale Deformable Cross-modal Interaction (MDCI) module, enabling accurate localization of anomalies with various shapes and sizes. In addition, we design a position-enhanced patch matching approach to improve few-shot anomaly detection performance. Experiments on multiple datasets demonstrate that FiLo++ achieves significant performance improvements compared with existing methods. Code will be available at https://github.com/CASIA-IVA-Lab/FiLo.
Zhaopeng Gu, Bingke Zhu, Guibo Zhu, Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang
IEEE Trans. Circuits Syst. Video Technol.2
2025 UniVAD: A Training-free Unified Model for Few-shot Visual Anomaly Detection
abstract
Visual Anomaly Detection (VAD) aims to identify abnormal samples in images that deviate from normal patterns, covering multiple domains, including industrial, logical, and medical fields. Due to the domain gaps between these fields, existing VAD methods are typically tailored to each domain, with specialized detection techniques and model architectures that are difficult to generalize across different domains. Moreover, even within the same domain, current VAD approaches often require large amounts of normal samples to train class-specific models, resulting in poor generalizability and hindering unified evaluation across domains. To address this issue, we propose a generalized few-shot VAD method, UniVAD, capable of detecting anomalies across various domains, with a training-free unified model. UniVAD only needs few normal samples as references during testing to detect anomalies in previously unseen objects, without training on the specific domain. Specifically, UniVAD employs a Contextual Component Clustering (C3) module based on clustering and vision foundation models to segment components within the image accurately, and leverages Component-Aware Patch Matching (CAPM) and Graph-Enhanced Component Modeling (GECM) modules to detect anomalies at different semantic levels, which are aggregated to produce the final detection result. We conduct experiments on nine datasets spanning industrial, logical, and medical fields, and the results demonstrate that UniVAD achieves state-of-the-art performance in few-shot anomaly detection tasks across multiple domains, outperforming domain-specific anomaly detection models. Code is available at https://github.com/FantasticGNU/UniVAD.
Zhaopeng Gu, Bingke Zhu, Guibo Zhu, Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang
CVPR2
2025 LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing
abstract
Audio-visual video parsing focuses on classifying videos through weak labels while identifying events as either visible, audible, or both, alongside their respective temporal boundaries. Many methods ignore that different modalities often lack alignment, thereby introducing extra noise during modal interaction. In this work, we introduce a Learning Interaction method for Non-aligned Knowledge (LINK), designed to equilibrate the contributions of distinct modalities by dynamically adjusting their input during event prediction. Additionally, we leverage the semantic information of pseudo-labels as a priori knowledge to mitigate noise from other modalities. Our experimental findings demonstrate that our model outperforms existing methods on the LLP dataset.
Langyu Wang, Bingke Zhu, Yingying Chen 0003, Jinqiao Wang
ICASSP2
2025 MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing
abstract
The weakly-supervised audio-visual video parsing (AVVP) aims to predict all modality-specific events and locate their temporal boundaries. Despite significant progress, due to the limitations of the weakly-supervised and the deficiencies of the model architecture, existing methods are lacking in simultaneously improving both the segment-level prediction and the event-level prediction. In this work, we propose a audio-visual Mamba network with pseudo labeling aUGmentation (MUG) for emphasising the uniqueness of each segment and excluding the noise interference from the alternate modalities. Specifically, we annotate some of the pseudo-labels based on previous work. Using unimodal pseudo-labels, we perform cross-modal random combinations to generate new data, which can enhance the model's ability to parse various segment-level event combinations. For feature processing and interaction, we employ a audio-visual mamba network. The AV-Mamba enhances the ability to perceive different segments and excludes additional modal noise while sharing similar modal information. Our extensive experiments demonstrate that MUG improves state-of-the-art results on LLP dataset in all metrics (e.g,, gains of 2.1% and 1.2% in terms of visual Segment-level and audio Segment-level metrics). Our code is available at https://github.com/WangLY136/MUG.
Langyu Wang, Bingke Zhu, Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang
ICCV2
2025 FLARE: A Framework for Stellar Flare Forecasting Using Stellar Physical Properties and Historical Records
abstract
Stellar flare events are critical observational samples for astronomical research; however, recorded flare events remain limited. Stellar flare forecasting can provide additional flare event samples to support research efforts. Despite this potential, no specialized models for stellar flare forecasting have been proposed to date. In this paper, we present extensive experimental evidence demonstrating that both stellar physical properties and historical flare records are valuable inputs for flare forecasting tasks. We then introduce FLARE (Forecasting Light-curve-based Astronomical Records via features Ensemble), the first-of-its-kind large model specifically designed for stellar flare forecasting. FLARE integrates stellar physical properties and historical flare records through a novel Soft Prompt Module and Residual Record Fusion Module. Experiments on the Kepler light curve dataset demonstrate that FLARE achieves superior performance compared to other methods across all evaluation metrics. Finally, we validate the forecast capability of our model through a comprehensive case study.
Bingke Zhu, Minghui Jia, Yihan Tao, A-Li Luo, Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang
IJCAI1
2025 AMITA: Attribute-Guided Masked Image-Text Alignment for Multi-Label Image Representation
abstract
Multi-label image classification, which involves recognizing multiple objects within a single image, is a fundamental task in computer vision. Recently, Visual-Language Models (VLMs) have made remarkable progress in this area. Many approaches combine textual and visual modalities to understand the entire image. In this paper, we find that there is a direct correlation between the accurate localization of objects and the accuracy of multi-label classification. However, previous research methods did not specifically address localization accuracy, resulting in sub-optimal accuracy. Therefore, we propose the AMITA, namely Attribute-guided Masked Image-Text Alignment for multi-label image representation. AMITA improves localization accuracy by segmenting object masks, thereby enhancing the accuracy of multi-label image classification. Additionally, AMITA introduces an AutoFocus method to handle the localization problem of small objects. AutoFocus conducts recognition by resizing and cropping the image respectively, and automatically selects the images useful for the classification target. Moreover, AMITA incorporates Attribute-guided Prompting to strengthen the semantic distinction among different categories. It uses large language models to obtain the attributes of different categories and carefully designs prompts to enhance the attribute differences among different categories. Finally, extensive experiments on three popular datasets, including MS-COCO, Pascal VOC 2007, and NUS-WIDE, demonstrate the superiority of AMITA.
Jinyi Fang, Bingke Zhu, Jingling Yuan, Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang
IEEE Trans. Circuits Syst. Video Technol.2
2025 Optimization of Prompt Learning via Multi-Knowledge Representation for Vision-Language Models
abstract
Vision-language models (VLMs), such as CLIP, play a foundational role in various cross-modal applications. To fully leverage the potential of VLMs in adapting to downstream tasks, context optimization methods such as prompt tuning are essential. However, one key limitation is the lack of diversity in prompt templates, whether they are hand-crafted or learned through additional modules. This limitation restricts the capabilities of pretrained VLMs and can result in incorrect predictions in downstream tasks. To address this challenge, we propose context optimization with multi-knowledge representation (CoKnow), a framework that enhances prompt learning for VLMs with rich contextual knowledge. To facilitate CoKnow during inference, we train lightweight semantic knowledge mappers, which are capable of generating multi-knowledge representations for an input image without requiring additional priors. Experimentally, we conduct extensive experiments on 11 publicly available datasets, demonstrating that CoKnow outperforms a series of previous methods.
Enming Zhang, Bingke Zhu, Yingying Chen 0003, Qinghai Miao, Ming Tang 0001, Jinqiao Wang
IEEE Trans. Multim.2
2024 AnomalyGPT: Detecting Industrial Anomalies Using Large Vision-Language Models
abstract
Large Vision-Language Models (LVLMs) such as MiniGPT-4 and LLaVA have demonstrated the capability of understanding images and achieved remarkable performance in various visual tasks. Despite their strong abilities in recognizing common objects due to extensive training datasets, they lack specific domain knowledge and have a weaker understanding of localized details within objects, which hinders their effectiveness in the Industrial Anomaly Detection (IAD) task. On the other hand, most existing IAD methods only provide anomaly scores and necessitate the manual setting of thresholds to distinguish between normal and abnormal samples, which restricts their practical implementation. In this paper, we explore the utilization of LVLM to address the IAD problem and propose AnomalyGPT, a novel IAD approach based on LVLM. We generate training data by simulating anomalous images and producing corresponding textual descriptions for each image. We also employ an image decoder to provide fine-grained semantic and design a prompt learner to fine-tune the LVLM using prompt embeddings. Our AnomalyGPT eliminates the need for manual threshold adjustments, thus directly assesses the presence and locations of anomalies. Additionally, AnomalyGPT supports multi-turn dialogues and exhibits impressive few-shot in-context learning capabilities. With only one normal shot, AnomalyGPT achieves the state-of-the-art performance with an accuracy of 86.1%, an image-level AUC of 94.1%, and a pixel-level AUC of 95.3% on the MVTec-AD dataset.
Zhaopeng Gu, Bingke Zhu, Guibo Zhu, Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang
AAAI2
2024 A Set of Effective Strategies for Optimized Road Damage Detection
abstract
In this paper, we propose an optimized method for road damage detection using a lightweight YOLO model as the baseline. Our approach incorporates a set of effective strategies, including lightweight attention mechanisms, data augmentation, dynamic sampling, weight averaging, and multi-step knowledge distillation. Our method significantly improves inference speed while maintaining high accuracy compared to previous mainstream ensemble-based methods. Our approach achieves notable success in the IEEE Big Data 2024 Optimized Road Damage Detection Challenge (ORDDC’2024), securing second place with an F1 score of 0.7013 and an inference time of 0.0328 s per image. These results demonstrate a strong balance between accuracy and efficiency. Extensive experiments confirm that our method boosts detection accuracy and greatly accelerates inference, making it highly suitable for real-world applications. The source code and trained model are available at https://github.com/YinglongDu/ShiYu_Kunchuan_ORDDC2024.
Yinglong Du, Xu Zhao 0003, Bailin He, Bingke Zhu, Shuaihua Zhao, Guibo Zhu, Jinqiao Wang
IEEE Big Data4
2024 Estate: Expert-Guided State Text Enhancement for Zero-Shot Industrial Anomaly Detection
abstract
The Expert-Guided State Text Enhancement Anomaly Detection (ESTATE) framework addresses the challenges in industrial anomaly detection arising from diverse product categories and limited defective samples. This framework, integrating expert insights through comparative state prompts, leverages two innovative text-guided networks, CLS-Refiner and SEG-Refiner, enhancing model training. These networks, connected to residual textual features of standard vision-language pre-trained models, focus on amplifying adjectives’ significance in text for improved image block and pixel-level alignment. ESTATE’s effectiveness is demonstrated through evaluations on MVTecAD and VisA datasets, achieving AUROC scores of 89.6%/89.6% for classification and 95.1%/85.0% for segmentation tasks, alongside setting new benchmarks in F1Max and PRO metrics. The AUC-cls on MVTecAD and VisA demonstrated an enhancement of 5.06% and 8.97%, respectively, compared to the APRIL-GAN approach.
Bingke Zhu, Hao Li 0115, Changlin Chen, Liujie Hua, Jinqiao Wang
ICIP1
2024 FiLo: Zero-Shot Anomaly Detection by Fine-Grained Description and High-Quality Localization
abstract
Zero-shot anomaly detection (ZSAD) methods detect anomalies without prior access to known normal or abnormal samples within target categories. Existing methods typically rely on pretrained multimodal models, computing similarities between manually crafted textual features representing ''normal'' or ''abnormal'' semantics and image patch features to detect anomalies. However, the generic descriptions of ''abnormal'' often fail to precisely match diverse types of anomalies across different object categories. Additionally, computing feature similarities for single patches struggles to pinpoint specific locations of anomalies with various sizes and scales. To address these issues, we propose a novel ZSAD method called FiLo, comprising two components: adaptively learned Fine-Grained Description (FG-Des) and position-enhanced High-Quality Localization (HQ-Loc). FG-Des introduces fine-grained anomaly descriptions for each category using Large Language Models (LLMs) and employs adaptively learned textual templates to enhance the accuracy and interpretability of anomaly detection. HQ-Loc, utilizing Grounding DINO for preliminary localization, position-enhanced text prompts, and Multi-scale Multi-shape Cross-modal Interaction (MMCI) module, facilitates more accurate localization of anomalies of different sizes and shapes. Experimental results on datasets like MVTec and VisA demonstrate that FiLo significantly improves the performance of ZSAD in both detection and localization, achieving state-of-the-art performance with an image-level AUC of 83.9% and a pixel-level AUC of 95.9% on the VisA dataset. Code is available at https://github.com/CASIA-IVA-Lab/FiLo.
Zhaopeng Gu, Bingke Zhu, Guibo Zhu, Yingying Chen 0003, Hao Li 0115, Ming Tang 0001, Jinqiao Wang
ACM Multimedia2
2023 Explicit Attention Modeling for Pedestrian Attribute Recognition
abstract
Recent studies on pedestrian attribute recognition have achieved significant improvements by utilizing complex networks and attention mechanisms. However, most of these studies learn the attention map implicitly through the class activation map. In this paper, we propose an explicit attention modeling approach for pedestrian attribute recognition. We construct a mask branch to learn the attention maps with a lightweight feature pyramid network. The features inside the specific mask are then averaged to obtain the scores for attribute recognition. Additionally, we introduce spatial and semantic distillation to improve the consistency of attention masks and attribute scores. Our experiments demonstrate that the proposed explicit attention modeling can achieve state-of-the-art performance on PA100K, PETA, and PAR datasets with negligible parameters.
Jinyi Fang, Bingke Zhu, Yingying Chen 0003, Jinqiao Wang, Ming Tang 0001
ICME2
2023 Uncertainty-Aware Boundary Attention Network for Real-Time Semantic Segmentation
Yuanbing Zhu, Bingke Zhu, Yingying Chen 0003, Jinqiao Wang
PRCV (3)2
2022 An Ensemble of One-Stage and Two-Stage Detectors Approach for Road Damage Detection
abstract
With the growth of the city and the increase in the number of cars, the maintenance and management of roads attract more attention. Road damage detection of road images is the basic step of road maintenance. To reduce the cost of labor, it is crucial to make the best use of road damage images from different geographical environments and capturing devices. This paper describes our 1-st place solution used in the Crowd sensing-based Road Damage Detection Challenge of the 2022 IEEE International Conference on Big Data. We use YOLO-series models and Faster RCNN as our one-stage and two-stage baseline models respectively. Our model only needs to be trained directly on the datasets of the overall six countries. Besides, with ensemble learning and test time augmentation, our ensemble model achieves the best results on the learderboard of each single country (India, Japan, United States, and Norway) without fine-tuning. Our ensemble model achieves the F1 scores of 0.7699 and 0.7160 on Overall and Average leaderboard, which significantly outperformed the 2-nd p lace F1 s cores of 0.7432 and 0.6744. The source code and trained model are available at https://github.com/berry-ding/ShiYu_SeaView_GRDDC2022.
Wenchao Ding 0004, Xu Zhao 0003, Bingke Zhu, Yinglong Du, Guibo Zhu, Tao Yu 0013, Jinqiao Wang
IEEE Big Data3
2020 Part-Aware Context Network for Human Parsing
abstract
Recent works have made significant progress in human parsing by exploiting rich contexts. However, human parsing still faces a challenge of how to generate adaptive contextual features for the various sizes and shapes of human parts. In this work, we propose a Part-aware Context Network (PCNet), a novel and effective algorithm to deal with the challenge. PCNet mainly consists of three modules, including a part class module, a relational aggregation module, and a relational dispersion module. The part class module extracts the high-level representations of every human part from a categorical perspective. We design a relational aggregation module to capture the representative global context by mining associated semantics of human parts, which adaptively augments the context for human parts. We propose a relational dispersion module to generate the discriminative and effective local context and neglect disturbing one by making the affinity of human parts dispersed. The relational dispersion module ensures that features in the same class will be close to each other and away from those of different classes. By fusing the outputs of the relational aggregation module, the relational dispersion module and the backbone network, our PCNet generates adaptive contextual features for various sizes of human parts, improving the parsing accuracy. We achieve a new state-of-the-art segmentation performance on three challenging human parsing datasets, i.e., PASCAL-Person-Part, LIP, and CIHP.
Yingying Chen 0003, Bingke Zhu, Jinqiao Wang, Ming Tang 0001
CVPR3
2020 Blended Grammar Network for Human Parsing
Yingying Chen 0003, Bingke Zhu, Jinqiao Wang, Ming Tang 0001
ECCV (24)3
2020 Semantic-spatial fusion network for human parsing
Yingying Chen 0003, Bingke Zhu, Jinqiao Wang, Ming Tang 0001
Neurocomputing3
2019 Pixelwise Deep Sequence Learning for Moving Object Detection
abstract
Moving object detection is an essential, well-studied but still open problem in computer vision and plays a fundamental role in many applications. Traditional approaches usually reconstruct background images with hand-crafted visual features, such as color, texture, and edge. Due to lack of prior knowledge or semantic information, it is difficult to deal with complicated and rapid changing scenes. To exploit the temporal structure of the pixel-level semantic information, in this paper, we propose an end-to-end deep sequence learning architecture for moving object detection. First, the video sequences are input into a deep convolutional encoder-decoder network for extracting pixel-wise semantic features. Then, to exploit the temporal context, we propose a novel attention long short-term memory (Attention ConvLSTM) to model pixelwise changes over time. A spatial transformer network and a conditional random field layer are finally appended to reduce the sensitivity to camera motion and smooth the foreground boundaries. A multi-task loss is proposed to jointly optimization for frame-based classification and temporal prediction in an end-to-end network. Experimental results on CDnet 2014 and LASIESTA show 12.15% and 16.71% improvement to the state of the art, respectively.
Yingying Chen 0003, Jinqiao Wang, Bingke Zhu, Ming Tang 0001, Hanqing Lu
IEEE Trans. Circuits Syst. Video Technol.3
2018 Progressive Cognitive Human Parsing
abstract
Human parsing is an important task for human-centric understanding. Generally, two mainstreams are used to deal with this challenging and fundamental problem. The first one is employing extra human pose information to generate hierarchical parse graph to deal with human parsing task. Another one is training an end-to-end network with the semantic information in image level. In this paper, we develop an end-to-end progressive cognitive network to segment human parts. In order to establish a hierarchical relationship, a novel component-aware region convolution structure is proposed. With this structure, latter layers inherit prior component information from former layers and pay its attention to a finer component. In this way, we deal with human parsing as a progressive recognition task, that is, we first locate the whole human and then segment the hierarchical components gradually. The experiments indicate that our method has a better location capacity for the small objects and a better classification capacity for the large objects. Moreover, our framework can be embedded into any fully convolutional network to enhance the performance significantly.
Bingke Zhu, Yingying Chen 0003, Ming Tang 0001, Jinqiao Wang
AAAI1
2018 Tree Hierarchical CNNs for Object Parsing
abstract
Object parsing is a challenging topic in computer vision, which is to distinguish all parts of visual objects. Although lots of works have been proposed, it is difficult to segment complicated objects from complex scenes. Therefore, in this paper we propose a tree hierarchical CNNs for object parsing. Rather than segment all parts of objects at once, we segment object parts step by step in a tree hierarchy and then merge the results together with a full convolutional network. In the tree hierarchy, the segmentation errors of the previous layers of the network outputs could be passed down to following layers and result in accumulated errors. In order to reduce the accumulated errors, we adopt a new part-aware fusion strategy, which fuses global-level feature maps from fully convolutional networks as well as the part-level object feature maps from the output of previous layer. It also contributes to improve the integrity and robustness of object parsing. Finally, the experiments on published datasets show the superiority of the proposed approach, especially for neighboring objects in complex scene.
Yingying Chen 0003, Bingke Zhu, Jinqiao Wang, Ming Tang 0001, Hanqing Lu
ICIP3
2017 Fast Deep Matting for Portrait Animation on Mobile Phone
abstract
Image matting plays an important role in image and video editing. However, the formulation of image matting is inherently ill-posed. Traditional methods usually employ interaction to deal with the image matting problem with trimaps and strokes, and cannot run on the mobile phone in real-time. In this paper, we propose a real-time automatic deep matting approach for mobile devices. By leveraging the densely connected blocks and the dilated convolution, a light full convolutional network is designed to predict a coarse binary mask for portrait image. And a feathering block, which is edge-preserving and matting adaptive, is further developed to learn the guided filter and transform the binary mask into alpha matte. Finally, an automatic portrait animation system based on fast deep matting is built on mobile devices, which does not need any interaction and can realize real-time matting with 15 fps. The experiments show that the proposed approach achieves comparable results with the state-of-the-art matting solvers.
Bingke Zhu, Yingying Chen 0003, Jinqiao Wang, Si Liu 0001, Bo Zhang 0069, Ming Tang 0001
ACM Multimedia1