Yongtao Wang

dblp:48/4720 · DBLP profile ↗
← Back
84ranked-venue papers
8as first author
39since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 51 · 2 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 50 · 6 first-author · 20 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 2 since 2021Computer networks · 2 · 1 since 2021Systems, architecture and hardware · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DrivingGaussian++: Toward Realistic Reconstruction and Editable Simulation for Surrounding Dynamic Driving Scenes
abstract
We present DrivingGaussian++, an efficient and effective framework for realistic reconstruction and controllable editing of surrounding dynamic autonomous driving scenes. DrivingGaussian++ models the static background with incremental 3D Gaussians and reconstructs moving objects with a composite dynamic Gaussian graph, ensuring accurate positions and occlusions. By integrating a LiDAR prior, it achieves detailed and consistent scene reconstruction, outperforming existing methods in dynamic scene reconstruction and photorealistic surround-view synthesis. DrivingGaussian++ supports training-free controllable editing for dynamic driving scenes, including texture modification, weather simulation, and object manipulation, leveraging multi-view images and depth priors. By integrating large language models (LLMs) and controllable editing, our method can automatically generate dynamic object motion trajectories and enhance their realism during the optimization process. DrivingGaussian++ demonstrates consistent and realistic editing results and generates dynamic multi-view driving scenarios, while significantly enhancing scene diversity.
Yajiao Xiong, Yongtao Wang, Deqing Sun, Ming-Hsuan Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 CMUA++: A cross-model universal active defense framework for combating deepfake
Xiaoyu Ye, Yongtao Wang, Kai-Kuang Ma
Pattern Recognit.4
2025 Multi-representation Adapter with Neural Architecture Search for Efficient Range-Doppler Radar Object Detection
Weicheng Zheng, Yongtao Wang
ICANN (2)3
2025 AutoOcc: Automatic Open-Ended Semantic Occupancy Annotation via Vision-Language Guided Gaussian Splatting
Yongtao Wang, Yufei Wei, Nan Dong, Ming-Hsuan Yang 0001
ICCV3
2025 RobuRCDet: Enhancing Robustness of Radar-Camera Fusion in Bird's Eye View for 3D Object Detection
abstract
While recent low-cost radar-camera approaches have shown promising results in multi-modal 3D object detection, both sensors face challenges from environmen- tal and intrinsic disturbances. Poor lighting or adverse weather conditions de- grade camera performance, while radar suffers from noise and positional ambigu- ity. Achieving robust radar-camera 3D object detection requires consistent perfor- mance across varying conditions, a topic that has not yet been fully explored. In this work, we first conduct a systematic analysis of robustness in radar-camera de- tection on five kinds of noises and propose RobuRCDet, a robust object detection model in bird’s eye view (BEV). Specifically, we design a 3D Gaussian Expan- sion (3DGE) module to mitigate inaccuracies in radar points, including position, Radar Cross-Section (RCS), and velocity. The 3DGE uses RCS and velocity priors to generate a deformable kernel map and variance for kernel size adjustment and value distribution. Additionally, we introduce a weather-adaptive fusion module, which adaptively fuses radar and camera features based on camera signal confi- dence. Extensive experiments on the popular benchmark, nuScenes, show that our RobuRCDet achieves competitive results in regular and noisy conditions. The source codes and trained models will be made available.
Jingtong Yue, Xiangtai Li, Lu Qi 0001, Yongtao Wang, Ming-Hsuan Yang 0001
ICLR7
2025 Securing Millions of Decentralized Identities in Alipay Super App with End-to-End Formal Verification
abstract
Decentralized Identity (DID) enhances authentication and privacy by empowering individuals to control their own digital identities, which has gained traction globally. To our knowledge, this paper presents the first end-to-end verification effort (from design to implementation) of a real-world Decentralized Identity (DID) protocol following the IIFAA DID standard, which has been deployed within the widely used super app Alipay and issued millions of DIDs in practice. We integrate formal verification into the development lifecycle of such industrial security protocol to systematically enhance its reliability from two levels: (1) At the design level, we utilized state-of-the-art protocol design verifier Tamarin to formally model the IIFAA DID standard under a realistic threat model tailored for super apps. We then formulated and performed automated verification of desired security properties using Tamarin. We identified several design flaws that could lead to a security breach. These issues were reported to the design team and have been addressed in the updated design. (2) At the implementation level, we first extract the desired specification derived from the verified symbolic model of protocol design in the form of a set of intermediate I/O specifications. Subsequently, we translate the I/O specifications into a set of functional specifications at the implementation level, which can then be verified by the automated tool VeriFast. We identified several inconsistencies between the implementation and the verified design which are fixed by the development team and led to verified implementation faithfully obeying the verified design, together offering an end-to-end verified secure DID protocol in Alipay super app. Our work showcases how an industrial security protocol development team can design and implement a practical verified secure Decentralized Identity (DID) protocol with the help of end-to-end formal verification.
Ziyu Mao, Xiaolin Ma, Lin Huang 0005, Weichao Sun, Yongtao Wang, Jingling Xue, Jingyi Wang 0004
ASE7
2025 VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion
abstract
Current perception models have achieved remarkable success by leveraging large-scale labeled datasets, but still face challenges in open-world environments with novel objects. To address this limitation, researchers introduce open-set perception models to detect or segment arbitrary test-time user-input categories. However, open-set models rely on human involvement to provide predefined object categories as input during inference. More recently, researchers have framed a more realistic and challenging task known as open-ended perception that aims to discover unseen objects without requiring any category-level input from humans at inference time. Nevertheless, open-ended models suffer from low performance compared to open-set models. In this paper, we present VL-SAM-V2, an open-world object detection framework that is capable of discovering unseen objects while achieving favorable performance. To achieve this, we combine queries from open-set and open-ended models and propose a general and specific query fusion module to allow different queries to interact. By adjusting queries from open-set models, we enable VL-SAM-V2 to be evaluated in the open-set or open-ended mode. In addition, to learn more diverse queries, we introduce ranked learnable queries to match queries with proposals from open-ended models by sorting. Moreover, we design a denoising point training strategy to facilitate the training process. Experimental results on LVIS show that our method surpasses the previous open-set and open-ended methods, especially on rare objects.
Yongtao Wang
NeurIPS2
2025 OpenAD: Open-World Autonomous Driving Benchmark for 3D Object Detection
abstract
Open-world perception aims to develop a model adaptable to novel domains and various sensor configurations and can understand uncommon objects and corner cases. However, current research lacks sufficiently comprehensive open-world 3D perception benchmarks and robust generalizable methodologies. This paper introduces OpenAD, the first real open-world autonomous driving benchmark for 3D object detection. OpenAD is built upon a corner case discovery and annotation pipeline that integrates with a multimodal large language model (MLLM). The proposed pipeline annotates corner case objects in a unified format for five autonomous driving perception datasets with 2000 scenarios. In addition, we devise evaluation methodologies and evaluate various open-world and specialized 2D and 3D models. Moreover, we propose a vision-centric 3D open-world object detection baseline and further introduce an ensemble method by fusing general and specialized models to address the issue of lower precision in existing open-world methods for the OpenAD benchmark. We host an online challenge on EvalAI. Data, toolkit codes, and evaluation codes are available at https://github.com/VDIGPKU/OpenAD.
Zhongyu Xia, Jishuo Li, Yongtao Wang, Ming-Hsuan Yang 0001
NeurIPS5
2025 EA3D: Online Open-World 3D Object Extraction from Streaming Videos
abstract
Current 3D scene understanding methods are limited by offline-collected multi-view data or pre-constructed 3D geometry. In this paper, we present ExtractAnything3D (EA3D), a unified online framework for open-world 3D object extraction that enables simultaneous geometric reconstruction and holistic scene understanding. Given a streaming video, EA3D dynamically interprets each frame using vision-language and 2D vision foundation encoders to extract object-level knowledge. This knowledge is integrated and embedded into a Gaussian feature map via a feed-forward online update strategy. We then iteratively estimate visual odometry from historical frames and incrementally update online Gaussian features with new observations. A recurrent joint optimization module directs the model's attention to regions of interest, simultaneously enhancing both geometric reconstruction and semantic understanding. Extensive experiments across diverse benchmarks and tasks, including photo-realistic rendering, semantic and instance segmentation, 3D bounding box and semantic occupancy estimation, and 3D mesh generation, demonstrate the effectiveness of EA3D. Our method establishes a unified and efficient framework for joint online 3D reconstruction and holistic scene understanding, enabling a broad range of downstream tasks. The project webpage is available at \url{https://github.com/VDIGPKU/EA3D}.
Yuang Jia, Yongtao Wang, Deqing Sun, Ming-Hsuan Yang 0001
NeurIPS4
2025 SAR-NAS: Lightweight SAR Object Detection with Neural Architecture Search
Yongtao Wang
PRCV (16)3
2025 FASTCC: A lightweight human pose detection method leveraging SimCC
abstract
As a focal area within computer vision algorithms, human pose estimation algorithms find applications in diverse fields such as security and virtual reality. Achieving a balance between speed and accuracy is imperative for practical applications. Existing methods often present a trade-off between high accuracy and real-time performance. In response, this paper introduces the Fast Coordinate Classification (FastCC) detection head. It employs a shared fully connected Transformer for global self-attention operations on feature layers from the backbone network. The spatial attention coordinate encoder then outputs the coordinates of the horizontal and vertical axes of the keypoints, which are subsequently combined to derive the actual keypoint positions. Experimental evaluations conducted on the COCO and MPII datasets demonstrate that our detection head enhances the accuracy of human pose estimation algorithms while maintaining a lightweight design, outperforming the traditional heatmap method.
Yi Li 0054, Yongtao Wang, Dou Quan, Yabo Yan, Qinghai Yang
Neurocomputing3
2025 NAS-BNN: Neural Architecture Search for Binary Neural Networks
Yongtao Wang, Jinhe Zhang, Xiaojie Chu, Haibin Ling
Pattern Recognit.2
2024 BEV-MAE: Bird's Eye View Masked Autoencoders for Point Cloud Pre-training in Autonomous Driving Scenarios
abstract
Existing LiDAR-based 3D object detection methods for autonomous driving scenarios mainly adopt the training-from-scratch paradigm. Unfortunately, this paradigm heavily relies on large-scale labeled data, whose collection can be expensive and time-consuming. Self-supervised pre-training is an effective and desirable way to alleviate this dependence on extensive annotated data. In this work, we present BEV-MAE, an efficient masked autoencoder pre-training framework for LiDAR-based 3D object detection in autonomous driving. Specifically, we propose a bird's eye view (BEV) guided masking strategy to guide the 3D encoder learning feature representation in a BEV perspective and avoid complex decoder design during pre-training. Furthermore, we introduce a learnable point token to maintain a consistent receptive field size of the 3D encoder with fine-tuning for masked point cloud inputs. Based on the property of outdoor point clouds in autonomous driving scenarios, i.e., the point clouds of distant objects are more sparse, we propose point density prediction to enable the 3D encoder to learn location information, which is essential for object detection. Experimental results show that BEV-MAE surpasses prior state-of-the-art self-supervised methods and achieves a favorably pre-training efficiency. Furthermore, based on TransFusion-L, BEV-MAE achieves new state-of-the-art LiDAR-based 3D object detection results, with 73.6 NDS and 69.6 mAP on the nuScenes benchmark. The source code will be released at https://github.com/VDIGPKU/BEV-MAE.
Yongtao Wang, Shengxiang Qi, Nan Dong, Ming-Hsuan Yang 0001
AAAI2
2024 RCBEVDet: Radar-Camera Fusion in Bird's Eye View for 3D Object Detection
abstract
Three-dimensional object detection is one of the key tasks in autonomous driving. To reduce costs in practice, low-cost multi-view cameras for 3D object detection are proposed to replace the expansive LiDAR sensors. However, relying solely on cameras is difficult to achieve highly accurate and robust 3D object detection. An effective solution to this issue is combining multi-view cameras with the economical millimeter-wave radar sensor to achieve more reliable multi-modal 3D object detection. In this paper, we introduce RCBEVDet, a radar-camera fusion 3D object detection method in the bird's eye view (BEV). Specifically, we first design RadarBEVNet for radar BEV feature extraction. RadarBEVNet consists of a dual-stream radar backbone and a Radar Cross-Section (RCS) aware BEV encoder. In the dual-stream radar backbone, a point-based encoder and a transformer-based encoder are proposed to extract radar features, with an injection and extraction module to facilitate communication between the two encoders. The RCS-aware BEV encoder takes RCS as the object size prior to scattering the point feature in BEV. Besides, we present the Cross-Attention Multi-layer Fusion module to automatically align the multi-modal BEV feature from radar and camera with the deformable attention mechanism, and then fuse the feature with channel and spatial fusion layers. Experimental results show that RCBEVDet achieves new state-of-the-art radar-camera fusion results on nuScenes and view-of-delft (VoD) 3D object detection benchmarks. Furthermore, RCBEVDet achieves better 3D detection results than all real-time camera-only and radar-camera 3D object detectors with a faster inference speed at 21∼28 FPS. The source code will be released at https://github.com/VDIGPKU/RCBEVDet.
Zhongyu Xia, Yongtao Wang, Shengxiang Qi, Nan Dong, Ce Zhu
CVPR5
2024 DrivingGaussian: Composite Gaussian Splatting for Surrounding Dynamic Autonomous Driving Scenes
abstract
We present DrivingGaussian, an efficient and effective framework for surrounding dynamic autonomous driving scenes. For complex scenes with moving objects, we first sequentially and progressively model the static background of the entire scene with incremental static 3D Gaussians. We then leverage a composite dynamic Gaussian graph to handle multiple moving objects, individually reconstructing each object and restoring their accurate positions and occlusion relationships within the scene. We further use a LiDAR prior for Gaussian Splatting to reconstruct scenes with greater details and maintain panoramic consistency. DrivingGaussian outperforms existing methods in dynamic driving scene reconstruction and enables photorealistic surround-view synthesis with high-fidelity and multi-camera consistency. Our project page is at: https:/github.com/VDIGPKU/DrivingGaussian.
Xiaojun Shan, Yongtao Wang, Deqing Sun, Ming-Hsuan Yang 0001
CVPR4
2024 TEOcc: Radar-Camera Multi-Modal Occupancy Prediction via Temporal Enhancement
abstract
As a novel 3D scene representation, semantic occupancy has gained much attention in autonomous driving. However, existing occupancy prediction methods mainly focus on designing better occupancy representations, such as tri-perspective view or neural radiance fields, while ignoring the advantages of using long-temporal information. In this paper, we propose a radar-camera multi-modal temporal enhanced occupancy prediction network, dubbed TEOcc. Our method is inspired by the success of utilizing temporal information in 3D object detection. Specifically, we introduce a temporal enhancement branch to learn temporal occupancy prediction. In this branch, we randomly discard the t−k input frame of the multi-view camera and predict its 3D occupancy by long-term and short-term temporal decoders separately with the information from other adjacent frames and multi-modal inputs. Besides, to reduce computational costs and incorporate multi-modal inputs, we specially designed 3D convolutional layers for long-term and short-term temporal decoders. Furthermore, since the lightweight occupancy prediction head is a dense classification head, we propose to use a shared occupancy prediction head for the temporal enhancement and main branches. It is worth noting that the temporal enhancement branch is only performed during training and is discarded during inference. Experiment results demonstrate that TEOcc achieves state-of-the-art occupancy prediction on nuScenes benchmarks. In addition, the proposed temporal enhancement branch is a plug-and-play module that can be easily integrated into existing occupancy prediction methods to improve the performance of occupancy prediction. The source code and models will be released at https://github.com/VDIGPKU/TEOcc.
Hongbo Jin, Yongtao Wang, Yufei Wei, Nan Dong
ECAI3
2024 HENet: Hybrid Encoding for End-to-End Multi-task 3D Perception from Multi-view Cameras
Zhongyu Xia, Yongtao Wang, Shengxiang Qi, Nan Dong, Ming-Hsuan Yang 0001
ECCV (50)4
2024 GALA3D: Towards Text-to-3D Complex Scene Generation via Layout-guided Generative Gaussian Splatting
abstract
We present GALA3D, generative 3D GAussians with LAyout-guided control, for effective compositional text-to-3D generation. We first utilize large language models (LLMs) to generate the initial layout and introduce a layout-guided 3D Gaussian representation for 3D content generation with adaptive geometric constraints. We then propose an instance-scene compositional optimization mechanism with conditioned diffusion to collaboratively generate realistic 3D scenes with consistent geometry, texture, scale, and accurate interactions among multiple objects while simultaneously adjusting the coarse layout priors extracted from the LLMs to align with the generated scene. Experiments show that GALA3D is a user-friendly, end-to-end framework for state-of-the-art scene-level 3D content generation and controllable editing while ensuring the high fidelity of object-level entities within the scene. The source codes and models will be available at gala3d.github.io.
Xingjian Ran, Yajiao Xiong, Jinlin He, Yongtao Wang, Deqing Sun, Ming-Hsuan Yang 0001
ICML6
2024 Training-Free Open-Ended Object Detection and Segmentation via Attention as Prompts
abstract
Existing perception models achieve great success by learning from large amounts of labeled data, but they still struggle with open-world scenarios. To alleviate this issue, researchers introduce open-set perception tasks to detect or segment unseen objects in the training set. However, these models require predefined object categories as inputs during inference, which are not available in real-world scenarios. Recently, researchers pose a new and more practical problem, i.e., open-ended object detection, which discovers unseen objects without any object categories as inputs. In this paper, we present VL-SAM, a training-free framework that combines the generalized object recognition model (i.e., Vision-Language Model) with the generalized object localization model (i.e., Segment-Anything Model), to address the open-ended object detection and segmentation task. Without additional training, we connect these two generalized models with attention maps as the prompts. Specifically, we design an attention map generation module by employing head aggregation and a regularized attention flow to aggregate and propagate attention maps across all heads and layers in VLM, yielding high-quality attention maps. Then, we iteratively sample positive and negative points from the attention maps with a prompt generation module and send the sampled points to SAM to segment corresponding objects. Experimental results on the long-tail instance segmentation dataset (LVIS) show that our method surpasses the previous open-ended method on the object detection task and can provide additional instance segmentation masks. Besides, VL-SAM achieves favorable performance on the corner case object detection dataset (CODA), demonstrating the effectiveness of VL-SAM in real-world applications. Moreover, VL-SAM exhibits good model generalization that can incorporate various VLMs and SAMs.
Yongtao Wang, Zhi Tang 0001
NeurIPS2
2024 FlowNAS: Neural Architecture Search for Optical Flow Estimation
Tingting Liang, Taihong Xiao, Yongtao Wang, Ming-Hsuan Yang 0001
Int. J. Comput. Vis.4
2023 T-SEA: Transfer-Based Self-Ensemble Attack on Object Detection
abstract
Compared to query-based black-box attacks, transferbased black-box attacks do not require any information of the attacked models, which ensures their secrecy. However, most existing transfer-based approaches rely on ensembling multiple models to boost the attack transferability, which is time- and resource-intensive, not to mention the difficulty of obtaining diverse models on the same task. To address this limitation, in this work, we focus on the single-model transfer-based black-box attack on object detection, utilizing only one model to achieve a high-transferability adversarial attack on multiple black-box detectors. Specifically, we first make observations on the patch optimization process of the existing method and propose an enhanced attack framework by slightly adjusting its training strategies. Then, we analogize patch optimization with regular model optimization, proposing a series of self-ensemble approaches on the input data, the attacked model, and the adversarial patch to efficiently make use of the limited information and prevent the patch from overfitting. The experimental results show that the proposed framework can be applied with multiple classical base attack methods (e.g., PGD and MIM) to greatly improve the black-box transferability of the well-optimized patch on multiple mainstream detectors, meanwhile boosting white-box performance. Our code is available at https://github.com/VDIGPKU/TSEA.
Huanran Chen, Yongtao Wang
CVPR4
2023 DynamicDet: A Unified Dynamic Architecture for Object Detection
abstract
Dynamic neural network is an emerging research topic in deep learning. With adaptive inference, dynamic models can achieve remarkable accuracy and computational efficiency. However, it is challenging to design a powerful dynamic detector, because of no suitable dynamic architecture and exiting criterion for object detection. To tackle these difficulties, we propose a dynamic framework for object detection, named DynamicDet. Firstly, we carefully design a dynamic architecture based on the nature of the object detection task. Then, we propose an adaptive router to analyze the multi-scale information and to decide the inference route automatically. We also present a novel optimization strategy with an exiting criterion based on the detection losses for our dynamic detectors. Last, we present a variable-speed inference strategy, which helps to realize a wide range of accuracy-speed trade-offs with only one dynamic detector. Extensive experiments conducted on the COCO benchmark demonstrate that the proposed DynamicDet achieves new state-of-the-art accuracy-speed trade-offs. For instance, with comparable accuracy, the inference speed of our dynamic detector Dy-YOLOv7-W6 surpasses YOLOv7-E6 by 12%, YOLOv7-D6 by 17%, and YOLOv7-E6E by 39%. The code is available at https://github.com/VDIGPKU/DynamicDet.
Yongtao Wang, Jinhe Zhang, Xiaojie Chu
CVPR2
2023 Differentiable Architecture Search with Random Features
abstract
Differentiable architecture search (DARTS) has signif-icantly promoted the development of NAS techniques because of its high search efficiency and effectiveness but suf-fers from performance collapse. In this paper, we make efforts to alleviate the performance collapse problem for DARTS from two aspects. First, we investigate the expres-sive power of the supernet in DARTS and then derive a new setup of DARTS paradigm with only training Batch-Norm. Second, we theoretically find that random features dilute the auxiliary connection role of skip-connection in supernet optimization and enable search algorithm focus on fairer operation selection, thereby solving the performance collapse problem. We instantiate DARTS and PC-DARTS with random features to build an improved version for each named RF-DARTS and RF-PCDARTS respectively. Experimental results show that RF-DARTS obtains 94.36% test accuracy on CIFAR-10 (which is the nearest optimal result in NAS-Bench-201), and achieves the newest state-of-the-art top-1 test error of 24.0% on ImageNet when transferring from CIFAR-10. Moreover, RF-DARTS performs robustly across three datasets (CIFAR-10, CIFAR-100, and SVHN) and four search spaces (S1-S4). Besides, RF-PCDARTS achieves even better results on ImageNet, that is, 23.9% top-1 and 7.1% top-5 test error, surpassing representative methods like single-path, training-free, and partial-channel paradigms directly searched on ImageNet.
Xuanyang Zhang, Yonggang Li 0001, Xiangyu Zhang 0005, Yongtao Wang, Jian Sun 0001
CVPR4
2023 Towards Fair and Comprehensive Comparisons for Image-Based 3D Object Detection
abstract
In this work, we build a modular-designed codebase, formulate strong training recipes, design an error diagnosis toolbox, and discuss current methods for image-based 3D object detection. In particular, different from other highly mature tasks, e.g., 2D object detection, the community of image-based 3D object detection is still evolving, where methods often adopt different training recipes and tricks resulting in unfair evaluations and comparisons. What is worse, these tricks may overwhelm their proposed designs in performance, even leading to wrong conclusions. To address this issue, we build a module-designed code-base and formulate unified training standards for the community. Furthermore, we also design an error diagnosis toolbox to measure the detailed characterization of detection models. Using these tools, we analyze current methods in-depth under varying settings and provide discussions for some open questions, e.g., discrepancies in conclusions on KITTI-3D and nuScenes datasets, which have led to different dominant methods for these datasets. We hope that this work will facilitate future research in image-based 3D object detection. Our codes will be released at https://github.com/OpenGVLab/3dodi.
Xinzhu Ma, Yongtao Wang, Yinmin Zhang, Zhiyi Xia, Zhihui Wang 0001, Wanli Ouyang
ICCV2
2023 SAMPLING: Scene-adaptive Hierarchical Multiplane Images Representation for Novel View Synthesis from a Single Image
abstract
Recent novel view synthesis methods obtain promising results for relatively small scenes, e.g., indoor environments and scenes with a few objects, but tend to fail for unbounded outdoor scenes with a single image as input. In this paper, we introduce SAMPLING, a Scene-adaptive Hierarchical Multiplane Images Representation for Novel View Synthesis from a Single Image based on improved multiplane images (MPI). Observing that depth distribution varies significantly for unbounded outdoor scenes, we employ an adaptive-bins strategy for MPI to arrange planes in accordance with each scene image. To represent intricate geometry and multi-scale details, we further introduce a hierarchical refinement branch, which results in high-quality synthesized novel views. Our method demonstrates considerable performance gains in synthesizing large-scale unbounded outdoor scenes using a single image on the KITTI dataset and generalizes well to the unseen Tanks and Temples dataset. The code and models will be made available at https://pkuvdig.github.io/SAMPLING/.
Xiaojun Shan, Yongtao Wang, Deqing Sun, Ming-Hsuan Yang 0001
ICCV4
2023 Foreground Guidance and Multi-Layer Feature Fusion for Unsupervised Object Discovery with Transformers
abstract
Unsupervised object discovery (UOD) has recently shown encouraging progress with the adoption of pre-trained Transformer features. However, current methods based on Transformers mainly focus on designing the localization head (e.g., seed selection-expansion and normalized cut) and overlook the importance of improving Transformer features. In this work, we handle UOD task from the perspective of feature enhancement and propose FOReground guidance and MUlti-LAyer feature fusion for unsupervised object discovery, dubbed FORMULA. Firstly, we present a foreground guidance strategy with an off-the-shelf UOD detector to highlight the foreground regions on the feature maps and then refine object locations in an iterative fashion. Moreover, to solve the scale variation issues in object detection, we design a multi-layer feature fusion module that aggregates features responding to objects at different scales. The experiments on VOC07, VOC12, and COCO 20k show that the proposed FORMULA achieves new state-of-the-art results on unsupervised object discovery. The code will be released at https://github.com/VDIGPKU/FORMULA.
Zengyu Yang, Yongtao Wang
WACV3
2023 A cascaded refined rgb-d salient object detection network based on the attention mechanism
Guanyu Zong, Longsheng Wei, Yongtao Wang
Appl. Intell.4
2023 An Intelligent MT Data Inversion Method With Seismic Attribute Enhancement
abstract
Magnetotelluric (MT) data inversion reconstructs an electrical resistivity structure most compatible with the observed MT data, and static correction can remove the undesired static shift effect in MT data. Conventional MT data static shift correction often faces the challenge of demanding requirements, such as large data amount, additional types of data, or a deep understanding of the research area. MT inversion constrained by seismic data often has better resolution and model consistency compared with independent MT inversion. However, valuable inversion knowledge contained in geophysicists’ expertise is not effectively incorporated. In this work, we present an intelligent MT data inversion method leveraging data- and physics-driven techniques based on deep learning. A novel MT data static shift correction method is introduced based on a neural network (NN). An MT data inversion method is formulated with the constraint of the extracted seismic reflection image based on two different NNs. Experiments on synthetic and field data verify the effectiveness of the proposed method.
Rui Guo 0017, Maokun Li, Fan Yang 0027, Shenheng Xu, Maoshan Chen, Yongtao Wang, Deqiang Tao, Zuzhi Hu, Xianwen Cui, Qinian Wang, Jiangbo Zhu, Suhe Huang
IEEE Trans. Geosci. Remote. Sens.7
2022 CMUA-Watermark: A Cross-Model Universal Adversarial Watermark for Combating Deepfakes
abstract
Malicious applications of deepfakes (i.e., technologies generating target facial attributes or entire faces from facial images) have posed a huge threat to individuals' reputation and security. To mitigate these threats, recent studies have proposed adversarial watermarks to combat deepfake models, leading them to generate distorted outputs. Despite achieving impressive results, these adversarial watermarks have low image-level and model-level transferability, meaning that they can protect only one facial image from one specific deepfake model. To address these issues, we propose a novel solution that can generate a Cross-Model Universal Adversarial Watermark (CMUA-Watermark), protecting a large number of facial images from multiple deepfake models. Specifically, we begin by proposing a cross-model universal attack pipeline that attacks multiple deepfake models iteratively. Then, we design a two-level perturbation fusion strategy to alleviate the conflict between the adversarial watermarks generated by different facial images and models. Moreover, we address the key problem in cross-model optimization with a heuristic approach to automatically find the suitable attack step sizes for different models, further weakening the model-level conflict. Finally, we introduce a more reasonable and comprehensive evaluation method to fully test the proposed method and compare it with existing ones. Extensive experimental results demonstrate that the proposed CMUA-Watermark can effectively distort the fake facial images generated by multiple deepfake models while achieving a better performance than existing methods. Our code is available at https://github.com/VDIGPKU/CMUA-Watermark.
Yongtao Wang, Zhaoyu Chen 0001, Zhi Tang 0001, Jingdong Chen, Weisi Lin, Kai-Kuang Ma
AAAI2
2022 Continual Contrastive Learning for Image Classification
abstract
Recently, self-supervised representation learning gives further development in multimedia technology. Most existing self-supervised learning methods are applicable to packaged data. However, when it comes to streamed data, they are suffering from catastrophic forgetting problem, which is not studied extensively. In this paper, we make the first attempt to tackle the catastrophic forgetting problem in the mainstream self-supervised methods, i.e., contrastive learning methods. Specifically, we first develop a rehearsal-based framework combined with a novel sampling strategy and a self-supervised knowledge distillation to transfer information over time efficiently. Then, we propose an extra sample queue to help the network separate the feature representations of old and new data in the embedding space. Experimental results show that compared with the naive self-supervised baseline, which learns tasks one by one without taking any technique, we improve the image classification accuracy by 1.60% on CIFAR-100, 2.86% on ImageNet-Sub, and 1.29% on ImageNet-Full under 10 incremental steps setting. Our code will be available at https://github.com/VDIGPKU/ContinualContrastiveLearning.
Yongtao Wang, Hongxiang Lin
ICME2
2022 IterVM: Iterative Vision Modeling Module for Scene Text Recognition
abstract
Scene text recognition (STR) is a challenging problem due to the imperfect imagery conditions in natural images. State-of-the-art methods utilize both visual cues and linguistic knowledge to tackle this challenging problem. Specifically, they propose iterative language modeling module (IterLM) to repeatedly refine the output sequence from the visual modeling module (VM). Though achieving promising results, the vision modeling module has become the performance bottleneck of these methods. In this paper, we newly propose iterative vision modeling module (IterVM) to further improve the STR accuracy. Specifically, the first VM directly extracts multi-level features from the input image, and the following VMs re-extract multi-level features from the input image and fuse them with the high-level (i.e., the most semantic one) feature extracted by the previous VM. By combining the proposed IterVM with iterative language modeling module, we further propose a powerful scene text recognizer called IterNet. Extensive experiments demonstrate that the proposed IterVM can significantly improve the scene text recognition accuracy, especially on low-quality scene text images. Moreover, the proposed scene text recognizer IterNet achieves new state-of-the-art results on several public benchmarks. Codes and models are available at https://github.com/VDIGPKU/IterNet.
Xiaojie Chu, Yongtao Wang
ICPR2
2022 BEVFusion: A Simple and Robust LiDAR-Camera Fusion Framework
abstract
Fusing the camera and LiDAR information has become a de-facto standard for 3D object detection tasks. Current methods rely on point clouds from the LiDAR sensor as queries to leverage the feature from the image space. However, people discovered that this underlying assumption makes the current fusion framework infeasible to produce any prediction when there is a LiDAR malfunction, regardless of minor or major. This fundamentally limits the deployment capability to realistic autonomous driving scenarios. In contrast, we propose a surprisingly simple yet novel fusion framework, dubbed BEVFusion, whose camera stream does not depend on the input of LiDAR data, thus addressing the downside of previous methods. We empirically show that our framework surpasses the state-of-the-art methods under the normal training settings. Under the robustness training settings that simulate various LiDAR malfunctions, our framework significantly surpasses the state-of-the-art methods by 15.7% to 28.9% mAP. To the best of our knowledge, we are the first to handle realistic LiDAR malfunction and can be deployed to realistic scenarios without any post-processing procedure.
Tingting Liang, Hongwei Xie, Kaicheng Yu, Zhongyu Xia, Yongtao Wang, Zhi Tang 0001
NeurIPS6
2022 Application of U-Net for the Recognition of Regional Features in Geophysical Inversion Results
abstract
Engineering geophysical prospecting is used to identify anomalous underground bodies, which are typically identified by inversion results. In recent years, machine learning has been applied to many geophysics studies, including geophysical data processing and fault prediction. Machine learning methods can also be used to identify complex anomalous bodies as direct identification can be difficult due to the complexity of local anomalous bodies. We propose a U-Net network to extract anomalous bodies from inverse results. Our model uses the regional similarity of the stratigraphic structure around the anomalous bodies, and then, the neural network extracts the regional background from the geophysical inversion results. Finally, the local anomalous bodies can be subtracted from the predicted regional results and the original inversion results. We designated a large number of datasets for training and designed triangular models for testing. Resistivity data for contaminated soil in a crane storage yard in Shanghai, China, were then processed and the resistivity range was calculated. We processed resistivity and polarizability data from Yantong Mountain, Heibei, China, to determine the degree of mineralization in the area.
LuoLei Zhang, Chongjin Zhao, Yongtao Wang, Yuhan Xu
IEEE Trans. Geosci. Remote. Sens.5
2022 CBNet: A Composite Backbone Network Architecture for Object Detection
abstract
top-performing object detectors depend heavily on backbone networks, whose advances bring consistent performance gains through exploring more effective network structures. In this paper, we propose a novel and flexible backbone framework, namely CBNet, to construct high-performance detectors using existing open-source pre-trained backbones under the pre-training fine-tuning paradigm. In particular, CBNet architecture groups multiple identical backbones, which are connected through composite connections. Specifically, it integrates the high- and low-level features of multiple identical backbone networks and gradually expands the receptive field to more effectively perform object detection. We also propose a better training strategy with auxiliary supervision for CBNet-based detectors. CBNet has strong generalization capabilities for different backbones and head designs of the detector architecture. Without additional pre-training of the composite backbone, CBNet can be adapted to various backbones (i.e., CNN-based vs. Transformer-based) and head designs of most mainstream detectors (i.e., one-stage vs. two-stage, anchor-based vs. anchor-free-based). Experiments provide strong evidence that, compared with simply increasing the depth and width of the network, CBNet introduces a more efficient, effective, and resource-friendly way to build high-performance backbone networks. Particularly, our CB-Swin-L achieves 59.4% box AP and 51.6% mask AP on COCO test-dev under the single-model and single-scale testing protocol, which are significantly better than the state-of-the-art results (i.e., 57.7% box AP and 50.2% mask AP) achieved by Swin-L, while reducing the training time by 6×. With multi-scale testing, we push the current best single model result to a new record of 60.1% box AP and 52.3% mask AP without using extra training data. Code is available at https://github.com/VDIGPKU/CBNetV2.
Ting-Ting Liang, Xiaojie Chu, Yongtao Wang, Zhi Tang 0001, Jingdong Chen, Haibin Ling
IEEE Trans. Image Process.4
2022 Screen Content Video Quality Assessment Model Using Hybrid Spatiotemporal Features
abstract
In this paper, a full-reference video quality assessment (VQA) model is designed for the perceptual quality assessment of the screen content videos (SCVs), called the hybrid spatiotemporal feature-based model (HSFM). The SCVs are of hybrid structure including screen and natural scenes, which are perceived by the human visual system (HVS) with different visual effects. With this consideration, the three dimensional Laplacian of Gaussian (3D-LOG) filter and three dimensional Natural Scene Statistics (3D-NSS) are exploited to extract the screen and natural spatiotemporal features, based on the reference and distorted SCV sequences separately. The similarities of these extracted features are then computed independently, followed by generating the distorted screen and natural quality scores for screen and natural scenes. After that, an adaptive screen and natural quality fusion scheme through the local video activity is developed to combine them for arriving at the final VQA score of the distorted SCV under evaluation. The experimental results on the Screen Content Video Database (SCVD) and Compressed Screen Content Video Quality (CSCVQ) databases have shown that the proposed HSFM is more in line with the perceptual quality assessment of the SCVs perceived by the HVS, compared with a variety of classic and latest IQA/VQA models.
Huanqiang Zeng, Hailiang Huang 0002, Junhui Hou, Jiuwen Cao, Yongtao Wang, Kai-Kuang Ma
IEEE Trans. Image Process.5
2021 OPANAS: One-Shot Path Aggregation Network Architecture Search for Object Detection
abstract
Recently, neural architecture search (NAS) has been exploited to design feature pyramid networks (FPNs) and achieved promising results for visual object detection. Encouraged by the success, we propose a novel One-Shot Path Aggregation Network Architecture Search (OPANAS) algorithm, which significantly improves both searching efficiency and detection accuracy. Specifically, we first introduce six heterogeneous information paths to build our search space, namely top-down, bottom-up, fusing-splitting, scale-equalizing, skip-connect and none. Second, we propose a novel search space of FPNs, in which each FPN candidate is represented by a densely-connected directed acyclic graph (each node is a feature pyramid and each edge is one of the six heterogeneous information paths). Third, we propose an efficient one-shot search method to find the optimal path aggregation architecture; specifically, we first train a super-net and then find the optimal candidate with an evolutionary algorithm. Experimental results demonstrate the efficacy of the proposed OPANAS for object detection: (1) OPANAS is more efficient than state-of-the-art methods (e.g., NAS-FPN and Auto-FPN) at significantly smaller searching cost (e.g., only 4 GPU days on MS-COCO); (2) the optimal architecture found by OPANAS significantly improves main-stream detectors including RetinaNet, Faster R-CNN and Cascade R-CNN, by 2.3∼3.2 % mAP compared to their FPN counterparts; and (3) a new state-of-the-art accuracy-speed trade-off (52.2 % mAP at 7.6 FPS) is achieved at smaller training costs than comparable recent arts. Code will be released at https://github.com/VDIGPKU/OPANAS.
Tingting Liang, Yongtao Wang, Zhi Tang 0001, Guosheng Hu, Haibin Ling
CVPR2
2021 Geometric Object 3D Reconstruction from Single Line Drawing Image Based on a Network for Classification and Sketch Extraction
Zhuoying Wang, Qingkai Fang, Yongtao Wang
ICDAR (1)3
2021 Rpattack: Refined Patch Attack on General Object Detectors
abstract
Nowadays, general object detectors like YOLO and Faster R-CNN as well as their variants are widely exploited in many applications. Many works have revealed that these detectors are extremely vulnerable to adversarial patch attacks. The perturbed regions generated by previous patch-based attack works on object detectors are very large which are not necessary for attacking and perceptible for human eyes. To generate much less but more efficient perturbation, we propose a novel patch-based method for attacking general object detectors. Firstly, we propose a patch selection and refining scheme to find the pixels which have the greatest importance for attack and remove the inconsequential perturbations gradually. Then, for a stable ensemble attack, we balance the gradients of detectors to avoid over-optimizing one of them during the training phase. Our RPAttack can achieve an amazing missed detection rate of 100% for both Yolo v4 and Faster R-CNN while only modifies 0.32% pixels on VOC 2007 test set. Our code is available at https://github.com/VDIGPKU/RPAttack.
Yongtao Wang, Zhaoyu Chen 0001, Zhi Tang 0001, Kai-Kuang Ma
ICME2
2021 Cascading Scene and Viewpoint Feature Learning for Pedestrian Gender Recognition
abstract
Pedestrian gender recognition plays an important role in smart city. To effectively improve the pedestrian gender recognition performance, a new method, called cascading scene and viewpoint feature learning (CSVFL), is proposed in this article. The novelty of the proposed CSVFL lies on the joint consideration of two crucial challenges in pedestrian gender recognition, namely, scene and viewpoint variation. For that, the proposed CSVFL starts with the scene transfer (ST) scheme, followed by the viewpoint adaptation (VA) scheme in a cascading manner. Specifically, the ST scheme exploits the key pedestrian segmentation network to extract the key pedestrian masks for the subsequent key pedestrian transfer generative adversarial network, with the goal of encouraging the input pedestrian image to have the similar style to the target scene while preserving the image details of the key pedestrian as much as possible. Afterward, the obtained scene-transferred pedestrian images are fed to train the deep feature learning network with the VA scheme, in which each neuron will be enabled/disabled for different viewpoints depending on whether it has contribution on the corresponding viewpoint. Extensive experiments conducted on the commonly used pedestrian attribute data sets have demonstrated that the proposed CSVFL approach outperforms multiple recently reported pedestrian gender recognition methods.
Huanqiang Zeng, Jianqing Zhu, Jiuwen Cao, Yongtao Wang, Kai-Kuang Ma
IEEE Internet Things J.5
2020 CBNet: A Novel Composite Backbone Network Architecture for Object Detection
abstract
In existing CNN based detectors, the backbone network is a very important component for basic feature1 extraction, and the performance of the detectors highly depends on it. In this paper, we aim to achieve better detection performance by building a more powerful backbone from existing ones like ResNet and ResNeXt. Specifically, we propose a novel strategy for assembling multiple identical backbones by composite connections between the adjacent backbones, to form a more powerful backbone named Composite Backbone Network (CBNet). In this way, CBNet iteratively feeds the output features of the previous backbone, namely high-level features, as part of input features to the succeeding backbone, in a stage-by-stage fashion, and finally the feature maps of the last backbone (named Lead Backbone) are used for object detection. We show that CBNet can be very easily integrated into most state-of-the-art detectors and significantly improve their performances. For example, it boosts the mAP of FPN, Mask R-CNN and Cascade R-CNN on the COCO dataset by about 1.5 to 3.0 points. Moreover, experimental results show that the instance segmentation results can be improved as well. Specifically, by simply integrating the proposed CBNet into the baseline detector Cascade Mask R-CNN, we achieve a new state-of-the-art result on COCO dataset (mAP of 53.3) with a single model, which demonstrates great effectiveness of the proposed CBNet architecture. Code will be made available at https://github.com/PKUbahuangliuhe/CBNet.
Yongtao Wang, Siwei Wang 0001, Tingting Liang, Qijie Zhao, Zhi Tang 0001, Haibin Ling
AAAI2
2020 Differentiable Automatic Data Augmentation
Yonggang Li 0001, Guosheng Hu, Yongtao Wang, Timothy M. Hospedales, Neil Robertson 0002, Yongxin Yang
ECCV (22)3
2020 Dual Loss for Manga Character Recognition with Imbalanced Training Data
abstract
Manga character recognition is a key technology for manga character retrieval and verification. This task is very challenging since the manga character images have a long-tailed distribution and large quality variations. Training models with cross-entropy softmax loss on such imbalanced data would introduce biases to feature and class weight norms. To handle this problem, we propose a novel dual loss which is the sum of two losses: dual ring loss and dual adaptive re-weighting loss. Dual ring loss combines weight and feature soft normalization and serves as a regularization term to softmax loss. Dual adaptive re-weighting loss re-weights softmax loss according to the norm of both feature and class weight. With the proposed losses, we have achieved encouraging results on the Manga109 benchmark. Specifically, compared with the baseline softmax loss, our method improves the character retrieval mAP from 35.72% to 38.88% and the character verification accuracy from 87.00% to 88.50%.
Yonggang Li 0001, Yafeng Zhou, Yongtao Wang, Xiaoran Qin, Zhi Tang 0001
ICPR3
2020 MixTConv: Mixed Temporal Convolutional Kernels for Efficient Action Recognition
abstract
To efficiently extract spatiotemporal features of video for action recognition, most state-of-the-art methods integrate 1D temporal convolutional filters into 2D CNN backbones. However, they all exploit 1D temporal convolutional filters of fixed kernel size (i.e., 3) in their network building block, thus have suboptimal temporal modeling capability to handle both long-term and short-term actions. To address this problem, we first investigate the impacts of different kernel sizes for the 1D temporal convolutional filters. Then, we propose a simple yet efficient operation called Mixed Temporal Convolution (MixTConv), which consists of multiple depthwise 1D convolutional filters with different kernel sizes. By plugging MixTConv into the conventional 2D CNN backbone ResNet-50, we further propose an efficient and effective network architecture named MSTNet for action recognition, and achieve state-of-the-art results on multiple large-scale benchmarks.
Kaiyu Shan, Yongtao Wang, Zhi Tang 0001, Yangyan Li
ICPR2
2020 GSTO: Gated Scale-Transfer Operation for Multi-Scale Feature Learning in Semantic Segmentation
abstract
Existing CNN-based methods for semantic segmentation heavily depend on multi-scale features to meet the requirements of both semantic comprehension and detail preservation. State-of-the-art segmentation networks widely exploit conventional scale-transfer operations, i.e., up-sampling and down-sampling to learn multi-scale features. In this work, we find that these operations lead to scale-confused features and suboptimal performance because they are spatial-invariant and directly transit all feature information cross scales without spatial selection. To address this issue, we propose the Gated Scale-Transfer Operation (GSTO) to properly transit spatial-filtered features to another scale. Specifically, GSTO can work either with or without extra supervision. Unsupervised GSTO is learned from the feature itself while the supervised one is guided by the supervised probability matrix. Both forms of GSTO are lightweight and plug-and-play, which can be flexibly integrated into networks or modules for learning better multi-scale features. In particular, by plugging GSTO into HRNet, we get a more powerful backbone (namely GSTO-HRNet) for pixel labeling, and it achieves new state-of-the-art results on multiple benchmarks for semantic segmentation including Cityscapes, LIP, and Pascal Context, with a negligible extra computational cost. Moreover, experiment results demonstrate that GSTO can also significantly boost the performance of multi-scale feature aggregation modules like PPM and ASPP.
Zhuoying Wang, Yongtao Wang, Zhi Tang 0001, Yangyan Li, Haibin Ling, Weisi Lin
ICPR2
2020 Towards Accurate Panel Detection in Manga: A Combined Effort of CNN and Heuristics
Yafeng Zhou, Yongtao Wang, Zheqi He, Zhi Tang 0001, Ching Y. Suen
MMM (1)2
2020 A quadrilateral scene text detector with two-stage network architecture
Siwei Wang 0006, Zheqi He, Yongtao Wang, Zhi Tang 0001
Pattern Recognit.4
2020 Learning Local and Global Priors for JPEG Image Artifacts Removal
abstract
Lossy compression will inevitably introduce image artifacts in the decoded image and degrade the image quality. In recent years, convolutional neural network (CNN) has been exploited for removing compression artifacts with great success. However, most existing CNN-based methods only utilize image's local prior without considering the global prior on the training of their networks. In this letter, a novel CNN, called the local and global priors network (LGPNet), is proposed that simultaneously learns both the local and the global priors for removing compression image artifacts. To achieve this goal, a dual-attention unit (DAU) is developed and incorporated into the well-known U-Net architecture for learning a better local prior. Meanwhile, the global prior is also learned from the entire image via our proposed global prior network. Extensive experimental results have clearly demonstrated that our proposed LGPNet is able to effectively remove image artifacts and greatly improve the image quality of JPEG-compressed images.
Yongtao Wang, Haihua Xie, Kai-Kuang Ma
IEEE Signal Process. Lett.2
2020 Learning a Single Model With a Wide Range of Quality Factors for JPEG Image Artifacts Removal
abstract
Lossy compression brings artifacts into the compressed image and degrades the visual quality. In recent years, many compression artifacts removal methods based on convolutional neural network (CNN) have been developed with great success. However, these methods usually train a model based on one specific value or a small range of quality factors. Obviously, if the test images quality factor does not match to the assumed value range, then degraded performance will be resulted. With this motivation and further consideration of practical usage, a highly robust compression artifacts removal network is proposed in this paper. Our proposed network is a single model approach that can be trained for handling a wide range of quality factors while consistently delivering superior or comparable image artifacts removal performance. To demonstrate, we focus on the JPEG compression with quality factors, ranging from 1 to 60. Note that a turnkey success of our proposed network lies in the novel utilization of the quantization tables as part of the training data. Furthermore, it has two branches in parallel-i.e., the restoration branch and the global branch. The former effectively removes the local artifacts, such as ringing artifacts removal. On the other hand, the latter extracts the global features of the entire image that provides highly instrumental image quality improvement, especially effective on dealing with the global artifacts, such as blocking, color shifting. Extensive experimental results performed on color and grayscale images have clearly demonstrated the effectiveness and efficacy of our proposed single-model approach on the removal of compression artifacts from the decoded image.
Yongtao Wang, Haihua Xie, Kai-Kuang Ma
IEEE Trans. Image Process.2
2019 M2Det: A Single-Shot Object Detector Based on Multi-Level Feature Pyramid Network
abstract
Feature pyramids are widely exploited by both the state-of-the-art one-stage object detectors (e.g., DSSD, RetinaNet, RefineDet) and the two-stage object detectors (e.g., Mask RCNN, DetNet) to alleviate the problem arising from scale variation across object instances. Although these object detectors with feature pyramids achieve encouraging results, they have some limitations due to that they only simply construct the feature pyramid according to the inherent multiscale, pyramidal architecture of the backbones which are originally designed for object classification task. Newly, in this work, we present Multi-Level Feature Pyramid Network (MLFPN) to construct more effective feature pyramids for detecting objects of different scales. First, we fuse multi-level features (i.e. multiple layers) extracted by backbone as the base feature. Second, we feed the base feature into a block of alternating joint Thinned U-shape Modules and Feature Fusion Modules and exploit the decoder layers of each Ushape module as the features for detecting objects. Finally, we gather up the decoder layers with equivalent scales (sizes) to construct a feature pyramid for object detection, in which every feature map consists of the layers (features) from multiple levels. To evaluate the effectiveness of the proposed MLFPN, we design and train a powerful end-to-end one-stage object detector we call M2Det by integrating it into the architecture of SSD, and achieve better detection performance than state-of-the-art one-stage detectors. Specifically, on MSCOCO benchmark, M2Det achieves AP of 41.0 at speed of 11.8 FPS with single-scale inference strategy and AP of 44.2 with multi-scale inference strategy, which are the new stateof-the-art results among one-stage detectors. The code will be made available on https://github.com/qijiezhao/M2Det.
Qijie Zhao, Yongtao Wang, Zhi Tang 0001, Haibin Ling
AAAI3
2019 Scene Text Recognition via Gated Cascade Attention
abstract
Scene text recognition is very challenging due to the complex background, low resolution, perspective distortion and curved placement, etc. Most of the state-of-the-art methods adopt the attention-based encoder-decoder framework, and usually get suboptimal recognition performance for challenging text images due to the misalignment between attention region and target character region. In this paper, a novel module, named Gated Cascade Attention Module (GCAM), is proposed to increase the alignment precision of attention in a cascade way. Moreover, a channel and spatial attention module is introduced into the encoder to extract more discriminative features for text recognition. By assembling these two modules, a novel scene text recognizer is developed, and extensive experiments demonstrate it can achieve state-of-the-art results on multiple benchmarks of regular and irregular text images.
Siwei Wang 0006, Yongtao Wang, Xiaoran Qin, Qijie Zhao, Zhi Tang 0001
ICME2
2019 Drone Image Stitching Guided by Robust Elastic Warping and Locality Preserving Matching
abstract
Image stitching stitches multiple overlapping images into a seamless image according to the corresponding geometric relationship between the reference and source images. In this study, the parallax-tolerant image stitching method based on robust elastic warping is applied to the stitching of drone images, and locality-preserving feature matching is used to effectively remove outliers from the drone images. The method can be divided into three stages, namely, locality-preserving feature matching, robust elastic warping, and global projectivity preservation. First, a set of high- precision point matching is provided for a drone image, and local matching is used. Second, the robust elastic warping function eliminates the parallax error, and the input image is distorted according to the calculated deformation on the grid plane. Finally, the global projectivity-preserving method is applied to obtain high-precision result panoramas. Experiments on several sets of drone images demonstrate that our method can generate better panoramas over the competitors.
Linbo Luo 0002, Qi Wan, Jun Chen 0019, Yongtao Wang, Xiaoguang Mei
IGARSS4
2019 High Performance Gesture Recognition via Effective and Efficient Temporal Modeling
abstract
State-of-the-art hand gesture recognition methods have investigated the spatiotemporal features based on 3D convolutional neural networks (3DCNNs) or convolutional long short-term memory (ConvLSTM). However, they often suffer from the inefficiency due to the high computational complexity of their network structures. In this paper, we focus instead on the 1D convolutional neural networks and propose a simple and efficient architectural unit, Multi-Kernel Temporal Block (MKTB), that models the multi-scale temporal responses by explicitly applying different temporal kernels. Then, we present a Global Refinement Block (GRB), which is an attention module for shaping the global temporal features based on the cross-channel similarity. By incorporating the MKTB and GRB, our architecture can effectively explore the spatiotemporal features within tolerable computational cost. Extensive experiments conducted on public datasets demonstrate that our proposed model achieves the state-of-the-art with higher efficiency. Moreover, the proposed MKTB and GRB are plug-and-play modules and the experiments on other tasks, like video understanding and video-based person re-identification, also display their good performance in efficiency and capability of generalization.
Feng Ni, Yuexin Ma, Xinge Zhu, Yuankai Qi, Riming Qiu, Yongtao Wang
IJCAI9
2019 SGDNet: An End-to-End Saliency-Guided Deep Neural Network for No-Reference Image Quality Assessment
abstract
We propose an end-to-end saliency-guided deep neural network (SGDNet) for no-reference image quality assessment (NR-IQA). Our SGDNet is built on an end-to-end multi-task learning framework in which two sub-tasks including visual saliency prediction and image quality prediction are jointly optimized with a shared feature extractor. The existing multi-task CNN-based NR-IQA methods which usually consider distortion identification as the auxiliary sub-task cannot accurately identify the complex mixtures of distortions exist in authentically distorted images. By contrast, our saliency prediction sub-task is more universal because visual attention always exists when viewing every image, regardless of its distortion type. More importantly, related works have reported that saliency information is highly correlated with image quality while this property is fully utilized in our proposed SGNet by training the model with more informative labels including saliency maps and quality scores simultaneously. In addition, the outputs of the saliency prediction sub-task are transparent to the primary quality regression sub-task by providing a kind of spatial attention masks for a more perceptually-consistent feature fusion. By training the whole network with the two sub-tasks together, more discriminant features can be learned and a more accurate mapping from feature representations to quality scores can be established. Experimental results on both authentically and synthetically distorted IQA datasets demonstrate the superiority of our SGDNet, as compared to the state-of-the-art approaches.
Sheng Yang 0006, Qiuping Jiang, Weisi Lin, Yongtao Wang
ACM Multimedia4
2019 Synthesizing data for text recognition with style transfer
Siwei Wang 0006, Yongtao Wang, Zhi Tang 0001
Multim. Tools Appl.3
2018 Comprehensive Feature Enhancement Module for Single-Shot Object Detector
Qijie Zhao, Yongtao Wang, Zhi Tang 0001
ACCV (5)2
2018 An End-to-End Quadrilateral Regression Network for Comic Panel Extraction
abstract
Comic panel extraction, i.e., decomposing a comic page image into panels, has become a fundamental technique for meeting many practical needs of mobile comic reading such as comic content adaptation and comic animating. Most of existing approaches are based on handcrafted low-level visual patterns and heuristics rules, thus having limited ability to deal with irregular comic panels. Only one existing method is based on deep learning and achieves better experimental results, but its architecture is redundant and its time efficiency is not good. To address these problems, we propose an end-to-end, two-stage quadrilateral regressing network architecture for comic panel detection, which inherits the architecture of Faster R-CNN. At the first stage, we propose a quadrilateral region proposal network for generating panel proposals, based on a newly proposed quadrilateral regression method. At the second stage, we classify the proposals and refine their shapes with the proposed quadrilateral regression method again. Extensive experimental results demonstrate that the proposed method significantly outperforms the existing comic panel detection methods on multiple datasets by F1-score and page accuracy.
Zheqi He, Yafeng Zhou, Yongtao Wang, Siwei Wang 0006, Xiaoqing Lu, Zhi Tang 0001
ACM Multimedia3
2017 SReN: Shape Regression Network for Comic Storyboard Extraction
abstract
The goal of storyboard extraction is to decompose the comic image into several storyboards(or frames), which is the fundamental step of comic image understanding and producing digital comic documents suitable for mobile reading. Most of existing approaches are based on hand crafted low-level visual patters like edge segments and line segments, which do not capture high-level vision. To overcome shortcomings of the existing approaches, we propose a novel architecture based on deep convolutional neural network, namely Shape Regression Network(SReN), to detect storyboards within comic images. Firstly, we use Fast R-CNN to generate rectangle bounding boxes as storyboard proposals. Then we train a deep neural network to predict quadrangles for these propos- als. Unlike existing object detection methods which only output rectangle bounding boxes, SReN can produce more precise quadrangle bounding boxes. Experimental results, evaluating on 7382 comic pages, demonstrate that SReN outperforms the state-of-the-art methods by more than 10% in terms of F1-score and page correction rate.
Zheqi He, Yafeng Zhou, Yongtao Wang, Zhi Tang 0001
AAAI3
2017 Mutual Enhancement for Detection of Multiple Logos in Sports Videos
abstract
Detecting logo frequency and duration in sports videos provides sponsors an effective way to evaluate their advertising efforts. However, general-purposed object detection methods cannot address all the challenges in sports videos. In this paper, we propose a mutual-enhanced approach that can improve the detection of a logo through the information obtained from other simultaneously occurred logos. In a Fast-RCNN-based framework, we first introduce a homogeneity-enhanced re-ranking method by analyzing the characteristics of homogeneous logos in each frame, including type repetition, color consistency, and mutual exclusion. Different from conventional enhance mechanism that improves the weak proposals with the dominant proposals, our mutual method can also enhance the relatively significant proposals with weak proposals. Mutual enhancement is also included in our frame propagation mechanism that improves logo detection by utilizing the continuity of logos across frames. We use a tennis video dataset and an associated logo collection for detection evaluation. Experiments show that the proposed method outperforms existing methods with a higher accuracy.
Xiaoqing Lu, Chengcui Zhang, Yongtao Wang, Zhi Tang 0001
ICCV4
2017 Geometric Object 3D Reconstruction from Single Line Drawing Image with Bottom-Up and Top-Down Classification and Sketch Generation
abstract
Since geometric objects are usually illustrated as 2D line drawings with the loss of depth information, it is difficult for people to fully understand the geometric structure of them. Most of the existing 3D reconstruction methods require the input to be perfect sketch of the line drawing, therefore they are not suitable when only line drawing image is offered. To address this problem, we propose a robust reconstruction method that can directly conduct reconstruction from the input single line drawing image. Our main contribution is that we introduce a bottom-up and top-down scheme to the task of geometric object 3D reconstruction. First, for the sub-task of geometric object classification, we perform a bottom-up process of geometric features extraction together with a top-down process of additional feature searching to classify the geometric object. Second, for the sub-task of sketch completion, we complete the sketch extracted from the geometric object's contour based on the class of the geometric object, in a top-down manner. Extensive experimental results demonstrate that our approach can improve in both accuracy and efficiency compared with the existing methods.
Yongtao Wang, Yafeng Zhou, Zheqi He, Zhi Tang 0001
ICDAR2
2017 A Faster R-CNN Based Method for Comic Characters Face Detection
abstract
Face detection of comic characters is a necessary step in most applications, such as comic character retrieval, automatic character classification and comic analysis. However, the existing methods were developed for simple cartoon images or small size comic datasets, and detection performance remains to be improved. In this paper, we propose a Faster R-CNN based method for face detection of comic characters. Our contribution is twofold. First, for the binary classification task of face detection, we empirically find that the sigmoid classifier shows a slightly better performance than the softmax classifier. Second, we build two comic datasets, JC2463 and AEC912, consisting of 3375 comic pages in total for characters face detection evaluation. Experimental results have demonstrated that the proposed method not only performs better than existing methods, but also works for comic images with different drawing styles.
Xiaoran Qin, Yafeng Zhou, Zheqi He, Yongtao Wang, Zhi Tang 0001
ICDAR4
2016 Design of the arm-wrestling robot's force acquisition system based on Qt
abstract
As a collection of entertainment and medical rehabilitation in a robot, the research on the arm-wrestling robot is of great significance. In order to achieve the collection of the arm-wrestling robot’s force signals, the design and implementation of arm-wrestling robot’s force acquisition system is introduced in this paper. The system is based on MP4221 data acquisition card and is programmed by Qt. It runs successfully in collecting the analog signals on PC. The interface of the system is simple and the real-time performance is good. The result of the test shows the feasibility in arm-wrestling robot.
Zhixiang Huo, Yongtao Wang
ICMV3
2016 Context-aware Geometric Object Reconstruction for Mobile Education
abstract
The solid geometric objects in the educational geometric books are usually illustrated as 2D line drawings accompanied with description text. In this paper, we present a method to recover the geometric objects from 2D to 3D. Unlike the previous methods, we not only use the geometric information from the line drawing itself, but also the textual information extracted from its context. The essential of our method is a cost function to mix the two types of information, and we optimize the cost function to identify the geometric object and recover its 3D information. Our method can recover various types of solid geometric objects including straight-edge manifolds and curved objects such as cone, cylinder and sphere. We show that our method performs significantly better compared to the previous ones.
Jinxin Zheng, Yongtao Wang, Zhi Tang 0001
ACM Multimedia2
2016 An R-CNN Based Method to Localize Speech Balloons in Comics
Yongtao Wang, Xicheng Liu, Zhi Tang 0001
MMM (1)1
2016 Comic storyboard extraction via edge segment analysis
Yongtao Wang, Yafeng Zhou, Zhi Tang 0001
Multim. Tools Appl.1
2016 Recovering solid geometric object from single line drawing image
Jinxin Zheng, Yongtao Wang, Zhi Tang 0001
Multim. Tools Appl.2
2016 Improving retrieval of plane geometry figure with learning to rank
Lu Liu 0018, Xiaoqing Lu, Yongtao Wang, Zhi Tang 0001
Pattern Recognit. Lett.4
2015 A fast and robust ellipse detector based on top-down least-square fitting
Yongtao Wang, Zheqi He, Xicheng Liu, Zhi Tang 0001, Luyuan Li
BMVC1
2015 A clump splitting based method to localize speech balloons in comics
abstract
Comic books enjoy great popularity across the world and people tend to read comics on digital devices nowadays. Localizing speech balloons is a necessary step in the process of digitalizing comic books, not only because they are one of the basic elements constituting a comic page but also because they are helpful to the localization of other comic elements. However, only a few studies have been done in this direction. In this paper, we propose a clump splitting based speech balloon localization method which can detect both closed and unclosed speech balloons. The contour generated by our method is accurate at pixel level and our method works for comic books of different drawing styles. Experimental results have demonstrated that our speech balloon localization method achieves better performance than existing methods on the common evaluation measures of recall, precision and F1score.
Xicheng Liu, Yongtao Wang, Zhi Tang 0001
ICDAR2
2015 Comic frame extraction via line segments combination
abstract
Automatic comic frame extraction is the core technique for comic content adaptation. Typical algorithms segment a comic page into a set of frames using connected components or division lines, but they cannot produce frames without blank margins and there is still room for improvement for complex comic images. We present a method that identifies frame polygons via connected component labeling and line segments combination. We analyze lines within each component and preliminarily judge the frame type, and then we optimize an energy-like score function constrained by several rules to choose frames. Experimental results indicate an increase of both accuracy and F-score compared with previous methods.
Yongtao Wang, Yafeng Zhou, Zhi Tang 0001
ICDAR1
2015 Aesthetic QR Codes Based on Two-Stage Image Blending
Yongtai Zhang, Shihong Deng, Yongtao Wang
MMM (2)4
2015 A tree conditional random field model for panel detection in comic images
Luyuan Li, Yongtao Wang, Ching Y. Suen, Zhi Tang 0001
Pattern Recognit.2
2014 Logical Labeling of Fixed Layout PDF Documents Using Multiple Contexts
abstract
The task of logical structure recovery is known to be of crucial importance, yet remains unsolved not only for image based document but also for born-digital document system. In this work, the modeling of contextual information based on 2D Conditional Random Fields is proposed to learn page structure for born-digital fixed-layout documents. Heuristic prior knowledge of Portable Document Format (PDF) content and layout are interpreted to construct neighborhood graphs and various pair wise clique templates for the modeling of multiple contexts. By integrating local and contextual observations obtained from PDF attributes, the ambiguities of semantic labels are better resolved. Experimental comparisons for six types of clique templates has demonstrated the benefits of contextual information in logical labeling of 16 finely defined categories.
Zhi Tang 0001, Canhui Xu 0002, Yongtao Wang
Document Analysis Systems4
2014 A Method of Density Analysis for Chinese Characters
Jingwei Qu, Xiaoqing Lu, Lu Liu 0018, Zhi Tang 0001, Yongtao Wang
NLPCC5
2014 Multiple-access wiretap channel with common channel state information at the encoders
abstract
The multiple‐access wiretap channel (MAC‐WTC) with common channel state information (CSI) at the encoders is studied, where two transmitters wish to send their confidential messages (no common message) to a legitimate receiver, while keeping a wiretapper as ignorant of the confidential messages as possible. Meanwhile, the channel is controlled by CSI, and it is available at the transmitters in a non‐causal manner or causal manner. This model can be viewed as a MAC extension of the WTC with CSI. Both the situation that the encoders can cooperate with each other and the situation that cooperation is not allowed between the encoders are investigated. More specifically, first, the cooperative MAC‐WTC with common CSI at the encoders is investigated, inner bounds on the secrecy capacity regions are provided for both non‐causal and causal manners. Secondly, the non‐cooperative MAC‐WTC with common CSI at the encoders is investigated, and also provide inner bounds on the secrecy capacity regions for both non‐causal and causal manners. Finally, by calculating the binary examples, the authors find that the cooperation between the encoders helps to obtain a larger secrecy rate region, that is, cooperation enhances the security of the MAC‐WTC with common CSI at the encoders.
Bin Dai 0003, Yongtao Wang, Zhuojun Zhuang
IET Commun.2
2014 Automatic comic page segmentation based on polygon detection
Luyuan Li, Yongtao Wang, Zhi Tang 0001, Liangcai Gao
Multim. Tools Appl.2
2014 Epipolar geometry estimation for wide baseline stereo by Clustering Pairing Consensus
Dazhi Zhang, Yongtao Wang, Wenbing Tao, Chengyi Xiong
Pattern Recognit. Lett.2
2013 Unsupervised Speech Text Localization in Comic Images
abstract
Localizing speech texts in comic images is a crucial step for catering the growing needs of reading comics on mobile devices. For example, automatically reading speech texts while adding sound effects alongside can not only render comic contents vividly but also help visually impaired readers. Unlike conventional text localization methods, we present an effective unsupervised speech text localization method in this paper that is free of training data. The proposed method consists of two major stages: (1) based on the concurrence of characters, the first stage of our method is to generate some of the character strings (a row or column of characters that align horizontally or vertically) from the comic images while the fonts and gaps of the adjacent characters within the character string are also obtained, (2) in the second stage, the obtained fonts and gaps of adjacent characters are used to detect rest of the character strings within the comic image via Bayesian classifier. The proposed method is tested on a dataset consists of 1000 comic images from ten printed comic series and provide satisfactory results.
Luyuan Li, Yongtao Wang, Zhi Tang 0001, Xiaoqing Lu, Liangcai Gao
ICDAR2
2013 Infrared Patch-Image Model for Small Target Detection in a Single Image
abstract
The robust detection of small targets is one of the key techniques in infrared search and tracking applications. A novel small target detection method in a single infrared image is proposed in this paper. Initially, the traditional infrared image model is generalized to a new infrared patch-image model using local patch construction. Then, because of the non-local self-correlation property of the infrared background image, based on the new model small target detection is formulated as an optimization problem of recovering low-rank and sparse matrices, which is effectively solved using stable principle component pursuit. Finally, a simple adaptive segmentation method is used to segment the target image and the segmentation result can be refined by post-processing. Extensive synthetic and real data experiments show that under different clutter backgrounds the proposed method not only works more stably for different target sizes and signal-to-clutter ratio values, but also has better detection performance compared with conventional baseline methods.
Chenqiang Gao, Deyu Meng, Yi Yang 0001, Yongtao Wang, Xiaofang Zhou 0001, Alex Hauptmann 0001
IEEE Trans. Image Process.4
2012 A graph-based method of newspaper article reconstruction
Liangcai Gao, Zhi Tang 0001, Xiaoyan Lin, Yongtao Wang
ICPR4
2012 Accountable authority key policy attribute-based encryption
Yongtao Wang, Kefei Chen, Yu Long 0001
Sci. China Inf. Sci.1
2011 Large Disparity Motion Layer Extraction via Topological Clustering
abstract
In this paper, we present a robust and efficient approach to extract motion layers from a pair of images with large disparity motion. First, motion models are established as: 1) initial SIFT matches are obtained and grouped into a set of clusters using our developed topological clustering algorithm; 2) for each cluster with no less than three matches, an affine transformation is estimated with least-square solution as tentative motion model; and 3) the tentative motion models are refined and the invalid models are pruned. Then, with the obtained motion models, a graph cuts based layer assignment algorithm is employed to segment the scene into several motion layers. Experimental results demonstrate that our method can successfully segment scenes containing objects with large interframe motion or even with significant interframe scale and pose changes. Furthermore, compared with the previous method invented by Wills and its modified version, our method is much faster and more robust.
Yongtao Wang, Junbin Gong, Dazhi Zhang, Chenqiang Gao, Jinwen Tian, Huanqiang Zeng
IEEE Trans. Image Process.1
2005 Design of sigma-delta modulators with arbitrary transfer functions
abstract
This paper addresses the design of sigma-delta modulators with arbitrary signal and noise transfer functions by presenting a genetic algorithm (GA) based search method. The objective function is defined to include the difference D between the magnitude of the frequency responses of the designed transfer functions and the ideal one, the quantizer gain /spl lambda//sub critical/ for which the poles of the modulator start moving out of the unit circle, and the spread of the coefficients S. Stability can be improved by reducing /spl lambda//sub critical/ while a smaller S reduces the implementation complexity. A genetic algorithm (GA) searches for poles/zeros of the transfer functions to minimize the objective function D+w/sub 1/*/spl lambda//sub critical/+w2*S, where w/sub 1/ and w/sub 2/ are two weighing factors. Numerical results demonstrate the effectiveness of the proposed method.
Yongtao Wang, Khurram Muhammad, Kaushik Roy 0001
ICASSP (5)1
2004 Hardware architecture and VLSI implementation of a low-power high-performance polyphase channelizer with applications to subband adaptive filtering
abstract
The polyphase channelizer is an important component of a subband adaptive filtering system. This paper presents an efficient hardware architecture and VLSI implementation of a low-power high-performance polyphase channelizer, integrating optimizations at algorithmic, architectural and circuit level. At the algorithm level, a computationally efficient structure is derived. Tradeoffs between hardware complexity and system performance are explored during the fixed-point modeling of the system. A computational complexity reduction technique is also employed to reduce the complexity of the hardware architecture. Circuit-level optimizations, including an efficient commutator implementation, dual-VDD scheme and novel level-converting flip-flops, are also integrated. Simulation results show that the design consumes 352 mW power with system throughput of 480 million samples per second (MSPS). A test chip has been submitted for fabrication to validate the proposed hardware architecture and VLSI design techniques.
Yongtao Wang, Hamid Mahmoodi, Lih-Yih Chiou, Hunsoo Choo, Jongsun Park 0001, Woopyo Jeong, Kaushik Roy 0001
ICASSP (5)1
2002 High performance and low power FIR filter design based on sharing multiplication
abstract
We present a high performance and low power FIR filter design, which is based on computation sharing multiplier (CSHM). CSHM specifically targets computation re-use in vector-scalar products and is effectively used in our FIR filter design. Efficient circuit level techniques: a new carry select adder and conditional capture flip-flop (CCFF), are also used to further improve power and performance. The proposed FIR filter architecture was implemented in 0.25 μm technology. Experimental results on a 10 tap low pass CSHM FIR filter show speed and power improvement of 19% and 17%, respectively, with respect to an FIR filter based on Wallace tree multiplier.
Jongsun Park 0001, Woopyo Jeong, Hunsoo Choo, Hamid Mahmoodi, Yongtao Wang, Kaushik Roy 0001
ISLPED5