Song Wang 0002

dblp:62/3151-2 · DBLP profile ↗
← Back
183ranked-venue papers
7as first author
79since 2021 · last 2026
0000-0003-4152-5295ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 124 · 2 first-author · 52 since 2021Artificial intelligence and machine learning · 109 · 6 first-author · 50 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 5 since 2021Security and privacy · 5 · 1 since 2021Databases, data management, data science and information retrieval · 2Theory of computation · 1
YearPublicationVenuePosition
2026 From indoor to outdoor: Unsupervised domain adaptive gait recognition
Likai Wang 0002, Wei Feng 0005, Rui-Ze Han, Xiangqun Zhang 0003, Yanjie Wei, Song Wang 0002
Pattern Recognit.6
2026 Shadow vanishing point detection via combined human/shadow adaptive modulation
Jin Wan, Hui Yin 0002, Zhenyao Wu, Xinyi Wu 0002, Song Wang 0002
Signal Process. Image Commun.6
2025 Acquire and then Adapt: Squeezing out Text-to-Image Model for Image Restoration
abstract
Recently, pre-trained text-to-image (T2I) models have been extensively adopted for real-world image restoration because of their powerful generative prior. However, controlling these large models for image restoration usually requires a large number of high-quality images and immense computational resources for training, which is costly and not privacy-friendly. In this paper, we find that the well-trained large T2I model (i.e., Flux) is able to produce a variety of high-quality images aligned with real-world distributions, offering an unlimited supply of training samples to mitigate the above issue. Specifically, we proposed a training data construction pipeline for image restoration, namely FluxGen, which includes unconditional image generation, image selection, and degraded image simulation. A novel light-weighted adapter (FluxIR) with squeeze-and-excitation layers is also carefully designed to control the large Diffusion Transformer (DiT)-based T2I model so that reasonable details can be restored. Experiments demonstrate that our proposed method enables the Flux model to adapt effectively to real-world image restoration tasks, achieving superior scores and visual quality on both synthetic and real-world degradation datasets - at only about 8.5% of the training cost compared to current approaches.
Junyuan Deng, Xinyi Wu 0002, Yongxing Yang, Congchao Zhu, Song Wang 0002, Zhenyao Wu
CVPR5
2025 Occluded person re-identification with feature complement and dual attention
Wei Xiong 0004, Zixin Tian, Lirong Li, Qin Zou 0001, Song Wang 0002
Expert Syst. Appl.5
2025 EfficientDeRain+: Learning Uncertainty-Aware Filtering via RainMix Augmentation for High-Efficiency Deraining
Qing Guo 0005, Hua Qi, Jingyang Sun, Felix Juefei-Xu, Lei Ma 0003, Di Lin 0002, Wei Feng 0005, Song Wang 0002
Int. J. Comput. Vis.8
2025 Coarse-to-fine text injecting for realistic image super-resolution
Chao Bai, Zhenyao Wu, Xinyi Wu 0002, Qi Zou 0001, Song Wang 0002
Neurocomputing7
2025 Unveiling the Power of Self-Supervision for Multi-View Multi-Human Association and Tracking
abstract
Multi-view multi-human association and tracking (MvMHAT), is an emerging yet important problem for multi-person scene video surveillance, aiming to track a group of people over time in each view, as well as to identify the same person across different views at the same time, which is different from previous MOT and multi-camera MOT tasks only considering the over-time human tracking. This way, the videos for MvMHAT require more complex annotations while containing more information for self-learning. In this work, we tackle this problem with an end-to-end neural network in a self-supervised learning manner. Specifically, we propose to take advantage of the spatial-temporal self-consistency rationale by considering three properties of reflexivity, symmetry, and transitivity. Besides the reflexivity property that naturally holds, we design the self-supervised learning losses based on the properties of symmetry and transitivity, for both appearance feature learning and assignment matrix optimization, to associate multiple humans over time and across views. Furthermore, to promote the research on MvMHAT, we build two new large-scale benchmarks for the network training and testing of different algorithms. Extensive experiments on the proposed benchmarks verify the effectiveness of our method. We have released the benchmark and code to the public.
Wei Feng 0005, Fei Wang 0032, Rui-Ze Han, Yiyang Gan, Zekun Qian, Junhui Hou, Song Wang 0002
IEEE Trans. Pattern Anal. Mach. Intell.7
2025 Learning with privileged stereo knowledge for monocular absolute 3D human pose estimation
Cunling Bian, Weigang Lu 0002, Wei Feng 0005, Song Wang 0002
Pattern Recognit. Lett.4
2025 A New Benchmark and Algorithm for Clothes-Changing Video Person Re-Identification
abstract
Person re-identification (Re-ID) is a classical computer vision task and has significant applications for public security and information forensics. Recently, long-term Re-ID with clothes-changing has attracted increasing attention. However, existing methods mainly focus on image-based setting, where richer temporal information is overlooked. In this paper, we focus on the relatively new yet practical problem of Clothes-Changing Video-based Re-ID (CCVReID), which is less studied. First, given the dataset shortage, we build two new benchmark datasets for CCVReID problem, including a large-scale synthetic video dataset and a real-world one, both containing human sequences with various clothing changes. Moreover, we systematically study this problem by simultaneously considering the classical appearance feature and temporal feature contained in the video. We develop a dual-branch fusion framework that makes use of the information from both clothes-aware appearance feature and clothes-free gait feature. For better information fusion, a confidence-guided re-ranking strategy is proposed to adaptively balance the weight of these two categories of features. We have released the benchmark and code proposed in this work to the public athttps://github.com/kkw98/CCVReID.
Likai Wang 0002, Xiangqun Zhang 0003, Rui-Ze Han, Yanjie Wei, Song Wang 0002, Wei Feng 0005
IEEE Trans. Inf. Forensics Secur.5
2024 From a Bird's Eye View to See: Joint Camera and Subject Registration without the Camera Calibration
abstract
We tackle a new problem of multi-view camera and sub-ject registration in the bird’ s eye view (BEV) without pregiven camera calibration, which promotes the multi-view subject registration problem to a new calibration-free stage. This greatly alleviates the limitation in many practical applications. However, this is a very challenging problem since its only input is several RGB images from different first-person views (FPVs), without the BEV image and the calibration of the FPVs, while the output is a unified plane aggregated from all views with the positions and orientations of both the subjects and cameras in a BEV. For this purpose, we propose an end-to-end framework solving cam-era and subject registration together by taking advantage of their mutual dependence, whose main idea is as below: i) creating a subject view-transform module (VTM) to project each pedestrian from FPV to a virtual BEV, ii) deriving a multi-view geometry-based spatial alignment module (SAM) to estimate the relative camera pose in a unified BEV, iii) selecting and refining the subject and camera registration results within the unified BEV. We collect a new large-scale synthetic dataset with rich annotations for training and evaluation. Additionally, we also collect a real dataset for cross-domain evaluation. The experimental results show the remarkable effectiveness of our method. The code and proposed datasets are available at BEVSee.
Zekun Qian, Rui-Ze Han, Wei Feng 0005, Song Wang 0002
CVPR4
2024 Bidirectional Autoregressive Diffusion Model for Dance Generation
abstract
Dance serves as a powerful medium for expressing human emotions, but the lifelike generation of dance is still a considerable challenge. Recently, diffusion models have showcased remarkable generative abilities across various domains. They hold promise for human motion generation due to their adaptable many-to-many nature. Nonetheless, current diffusion-based motion generation models often create entire motion sequences directly and unidirectionally, lacking focus on the motion with local and bidirectional enhancement. When choreographing high-quality dance movements, people need to take into account not only the musical context but also the nearby music-aligned dance motions. To authentically capture human behavior, we propose a Bidirectional Autoregressive Diffusion Model (BADM) for music-to-dance generation, where a bidirectional encoder is built to enforce that the generated dance is harmonious in both the forward and backward directions. To make the generated dance motion smoother, a local information decoder is built for local motion enhancement. The proposed framework is able to generate new motions based on the input conditions and nearby motions, which foresees individual motion slices iteratively and con-solidates all predictions. To further refine the synchronicity between the generated dance and the beat, the beat information is incorporated as an input to generate better music-aligned dance movements. Experimental results demonstrate that the proposed model achieves state-of-the-art performance compared to existing unidirectional approaches on the prominent benchmark for music-to-dance generation. The code and models are available: https://github.com/czzhang179/BADM.
Canyu Zhang 0002, Youbao Tang, Ruei-Sung Lin, Jing Xiao 0006, Song Wang 0002
CVPR7
2024 EINet: Point Cloud Completion via Extrapolation and Interpolation
Pingping Cai, Canyu Zhang 0002, Lingjia Shi, Nasrin Imanpour, Song Wang 0002
ECCV (40)6
2024 SAIR: Learning Semantic-Aware Implicit Representation
Canyu Zhang 0002, Qing Guo 0005, Song Wang 0002
ECCV (4)4
2024 Crossmodal Few-shot 3D Point Cloud Semantic Segmentation via View Synthesis
abstract
Cross-modal 2D-3D point cloud semantic segmentation using few-shot-based learning provides a practical approach for borrowing matured 2D domain knowledge into the 3D segmentation model, which reduces the reliance on laborious 3D annotation work and improves generalization to new categories. However, previous methods use single-view point cloud generation algorithms to bridge the gap between 2D images and 3D point clouds, leaving the incomplete geometry of an object or scene due to occlusions. To address this issue, we propose a novel view synthesis cross-modal few-shot point cloud semantic segmentation network. It introduces the color and depth inpainting to generate multi-view images and masks, which compensate for the absent depth information of generated point clouds. Additionally, we propose a Co-embedding Network to bridge the domain features between synthesized and original, collected 3D data, and a weighted prototype network is employed to balance the impact of multi-view images and enhance the segmentation performance. Extensive experiments on two benchmarks show the superiority of our method by outperforming the existing cross-modal few-shot 3D segmentation methods.
Pingping Cai, Canyu Zhang 0002, Song Wang 0002
ACM Multimedia5
2024 Estimating intrinsic characteristics of images for shadow removal
Hui Yin 0002, Jin Wan, Zhenyao Wu, Xinyi Wu 0002, Song Wang 0002
Comput. Graph.7
2024 Contactless interaction recognition and interactor detection in multi-person scenes
Rui-Ze Han, Wei Feng 0005, Haomin Yan, Song Wang 0002
Frontiers Comput. Sci.5
2024 Benchmarking the Complementary-View Multi-human Association and Tracking
Rui-Ze Han, Wei Feng 0005, Zekun Qian, Haomin Yan, Song Wang 0002
Int. J. Comput. Vis.6
2024 CRFormer: A cross-region transformer for shadow removal
Jin Wan, Hui Yin 0002, Zhenyao Wu, Xinyi Wu 0002, Song Wang 0002
Image Vis. Comput.6
2024 Asynchronous Multimodal Video Sequence Fusion via Learning Modality-Exclusive and -Agnostic Representations
abstract
Understanding human intentions (e.g., emotions) from videos has received considerable attention recently. Video streams generally constitute a blend of temporal data stemming from distinct modalities, including natural language, facial expressions, and auditory clues. Despite the impressive advancements of previous works via attention-based paradigms, the inherent temporal asynchrony and modality heterogeneity challenges remain in multimodal sequence fusion, causing adverse performance bottlenecks. To tackle these issues, we propose a Multimodal fusion approach for learning modality-Exclusive and modality-Agnostic representations (MEA) to refine multimodal features and leverage the complementarity across distinct modalities. On the one hand, MEA introduces a predictive self-attention module to capture reliable context dynamics within modalities and reinforce unique features over the modality-exclusive spaces. On the other hand, a hierarchical cross-modal attention module is designed to explore valuable element correlations among modalities over the modality-agnostic space. Meanwhile, a double-discriminator strategy is presented to ensure the production of distinct representations in an adversarial manner. Eventually, we propose a decoupled graph fusion mechanism to enhance knowledge exchange across heterogeneous modalities and learn robust multimodal representations for downstream tasks. Numerous experiments are implemented on three multimodal datasets with asynchronous sequences. Systematic analyses show the necessity of our approach.
Dingkang Yang, Mingcheng Li, Linhao Qu, Kun Yang 0010, Peng Zhai, Song Wang 0002, Lihua Zhang 0002
IEEE Trans. Circuits Syst. Video Technol.6
2024 MS-DETR: Multispectral Pedestrian Detection Transformer With Loosely Coupled Fusion and Modality-Balanced Optimization
abstract
Multispectral pedestrian detection is an important task for many around-the-clock applications, since the visible and thermal modalities can provide complementary information especially under low light conditions. Due to the presence of two modalities, misalignment and modality imbalance are the most significant issues in multispectral pedestrian detection. In this paper, we propose MultiSpectral pedestrian DEtection TRansformer (MS-DETR) to fix above issues. MS-DETR consists of two modality-specific backbones and Transformer encoders, followed by a multi-modal Transformer decoder, and the visible and thermal features are fused in the multi-modal Transformer decoder. To well resist the misalignment between multi-modal images, we design a loosely coupled fusion strategy by sparsely sampling some keypoints from multi-modal features independently and fusing them with adaptively learned attention weights. Moreover, based on the insight that not only different modalities, but also different pedestrian instances tend to have different confidence scores to final detection, we further propose an instance-aware modality-balanced optimization strategy, which preserves visible and thermal decoder branches and aligns their predicted slots through an instance-wise dynamic loss. Our end-to-end MS-DETR shows superior performance on the challenging KAIST, CVC-14 and LLVIP benchmark datasets. The source code is available athttps://github.com/YinghuiXing/MS-DETR.
Yinghui Xing, Song Wang 0002, Shizhou Zhang, Guoqiang Liang 0001, Xiuwei Zhang 0001, Yanning Zhang 0001
IEEE Trans. Intell. Transp. Syst.3
2023 Parametric Surface Constrained Upsampler Network for Point Cloud
abstract
Designing a point cloud upsampler, which aims to generate a clean and dense point cloud given a sparse point representation, is a fundamental and challenging problem in computer vision. A line of attempts achieves this goal by establishing a point-to-point mapping function via deep neural networks. However, these approaches are prone to produce outlier points due to the lack of explicit surface-level constraints. To solve this problem, we introduce a novel surface regularizer into the upsampler network by forcing the neural network to learn the underlying parametric surface represented by bicubic functions and rotation functions, where the new generated points are then constrained on the underlying surface. These designs are integrated into two different networks for two tasks that take advantages of upsampling layers -- point cloud upsampling and point cloud completion for evaluation. The state-of-the-art experimental results on both tasks demonstrate the effectiveness of the proposed method. The implementation code will be available at https://github.com/corecai163/PSCU.
Pingping Cai, Zhenyao Wu, Xinyi Wu 0002, Song Wang 0002
AAAI4
2023 Few-Shot 3D Point Cloud Semantic Segmentation via Stratified Class-Specific Attention Based Transformer Network
abstract
3D point cloud semantic segmentation aims to group all points into different semantic categories, which benefits important applications such as point cloud scene reconstruction and understanding. Existing supervised point cloud semantic segmentation methods usually require large-scale annotated point clouds for training and cannot handle new categories. While a few-shot learning method was proposed recently to address these two problems, it suffers from high computational complexity caused by graph construction and inability to learn fine-grained relationships among points due to the use of pooling operations. In this paper, we further address these problems by developing a new multi-layer transformer network for few-shot point cloud semantic segmentation. In the proposed network, the query point cloud features are aggregated based on the class-specific support features in different scales. Without using pooling operations, our method makes full use of all pixel-level features from the support samples. By better leveraging the support features for few-shot learning, the proposed method achieves the new state-of-the-art performance, with 15% less inference time, over existing few-shot 3D point cloud segmentation models on the S3DIS dataset and the ScanNet dataset. Our code is available at https://github.com/czzhang179/SCAT.
Canyu Zhang 0002, Zhenyao Wu, Xinyi Wu 0002, Song Wang 0002
AAAI5
2023 CDUL: CLIP-Driven Unsupervised Learning for Multi-Label Image Classification
abstract
This paper presents a CLIP-based unsupervised learning method for annotation-free multi-label image classification, including three stages: initialization, training, and inference. At the initialization stage, we take full advantage of the powerful CLIP model and propose a novel approach to extend CLIP for multi-label predictions based on globallocal image-text similarity aggregation. To be more specific, we split each image into snippets and leverage CLIP to generate the similarity vector for the whole image (global) as well as each snippet (local). Then a similarity aggregator is introduced to leverage the global and local similarity vectors. Using the aggregated similarity scores as the initial pseudo labels at the training stage, we propose an optimization framework to train the parameters of the classification network and refine pseudo labels for unobserved labels. During inference, only the classification network is used to predict the labels of the input image. Extensive experiments show that our method outperforms state-of-the-art unsupervised methods on MS-COCO, PASCAL VOC 2007, PASCAL VOC 2012, and NUS datasets and even achieves comparable results to weakly supervised classification methods.
Rabab Abdelfattah, Qing Guo 0005, Xiaofeng Wang 0007, Song Wang 0002
ICCV5
2023 Leveraging Inpainting for Single-Image Shadow Removal
abstract
Fully-supervised shadow removal methods achieve the best restoration qualities on public datasets but still generate some shadow remnants. One of the reasons is the lack of large-scale shadow & shadow-free image pairs. Unsupervised methods can alleviate the issue but their restoration qualities are much lower than those of fully-supervised methods. In this work, we find that pretraining shadow removal networks on the image inpainting dataset can reduce the shadow remnants significantly: a naive encoder-decoder network gets competitive restoration quality w.r.t. the state-of-the-art methods via only 10% shadow & shadow-free image pairs. After analyzing networks with/without inpainting pretraining via the information stored in the weight (IIW), we find that inpainting pretraining improves restoration quality in non-shadow regions and enhances the generalization ability of networks significantly. Additionally, shadow removal fine-tuning enables networks to fill in the details of shadow regions. Inspired by these observations we formulate shadow removal as an adaptive fusion task that takes advantage of both shadow removal and image inpainting. Specifically, we develop an adaptive fusion network consisting of two encoders, an adaptive fusion block, and a decoder. The two encoders are responsible for extracting the features from the shadow image and the shadow-masked image respectively. The adaptive fusion block is responsible for combining these features in an adaptive manner. Finally, the decoder converts the adaptive fused features to the desired shadow-free result. The extensive experiments show that our method empowered with inpainting outperforms all state-of-the-art methods. We have realized codes and models in https://github.com/tsingqguo/inpaint4shadow
Qing Guo 0005, Rabab Abdelfattah, Di Lin 0002, Wei Feng 0005, Ivor W. Tsang, Song Wang 0002
ICCV7
2023 An end-to-end network for co-saliency detection in one single image
Yuanhao Yue, Qin Zou 0001, Hongkai Yu, Qian Wang 0002, Zhongyuan Wang 0001, Song Wang 0002
Sci. China Inf. Sci.6
2023 Global-local contrastive multiview representation learning for skeleton-based action recognition
Cunling Bian, Wei Feng 0005, Song Wang 0002
Comput. Vis. Image Underst.4
2023 Relating View Directions of Complementary-View Mobile Cameras via the Human Shadow
Rui-Ze Han, Yiyang Gan, Likai Wang 0002, Nan Li 0048, Wei Feng 0005, Song Wang 0002
Int. J. Comput. Vis.6
2023 A One-Stage Domain Adaptation Network With Image Alignment for Unsupervised Nighttime Semantic Segmentation
abstract
In this paper, we tackle the problem of semantic segmentation for nighttime images that plays an equally important role as that for daytime images in autonomous driving, but is also much more challenging due to very poor illuminations and scarce annotated datasets. It can be treated as an unsupervised domain adaptation (UDA) problem, i.e., applying other labeled dataset taken in the daytime to guide the network training meanwhile reducing the domain shift, so that the trained model can generalize well to the desired domain of nighttime images. However, current general-purpose UDA approaches are insufficient to address the significant appearance difference between the day and night domains. To overcome such a large domain gap, we propose a novel domain adaptation network "DANIA" for nighttime semantic image segmentation by leveraging a labeled daytime dataset (the source domain) and an unlabeled dataset that contains coarsely aligned day-night image pairs (the target daytime and nighttime domains). These three domains are used to perform a multi-target adaptation via adversarial training in the network. Specifically, for the unlabeled day-night image pairs, we use the pixel-level predictions of static object categories on a daytime image as a pseudo supervision to segment its counterpart nighttime image. We also include a step of image alignment to relieve the inaccuracy caused by the misalignment between day-night image pairs by estimating a flow to refine the pseudo supervision produced by daytime images. Finally, a re-weighting strategy is applied to further improve the predictions, especially boosting the prediction accuracy of small objects. The proposed DANIA is a one-stage adaptation framework for nighttime semantic segmentation, which does not train additional day-night image transfer models as a separate pre-processing stage. Extensive experiments on Dark Zurich and Nighttime Driving datasets show that our DANIA achieves state-of-the-art performance for nighttime semantic segmentation.
Xinyi Wu 0002, Zhenyao Wu, Lili Ju, Song Wang 0002
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Capture the Moment: High-Speed Imaging With Spiking Cameras Through Short-Term Plasticity
abstract
High-speed imaging can help us understand some phenomena that are too fast to be captured by our eyes. Although ultra-fast frame-based cameras (e.g., Phantom) can record millions of fps at reduced resolution, they are too expensive to be widely used. Recently, a retina-inspired vision sensor, spiking camera, has been developed to record external information at 40, 000 Hz. The spiking camera uses the asynchronous binary spike streams to represent visual information. Despite this, how to reconstruct dynamic scenes from asynchronous spikes remains challenging. In this paper, we introduce novel high-speed image reconstruction models based on the short-term plasticity (STP) mechanism of the brain, termed TFSTP and TFMDSTP. We first derive the relationship between states of STP and spike patterns. Then, in TFSTP, by setting up the STP model at each pixel, the scene radiance can be inferred by the states of the models. In TFMDSTP, we use the STP to distinguish the moving and stationary regions, and then use two sets of STP models to reconstruct them respectively. In addition, we present a strategy for correcting error spikes. Experimental results show that the STP-based reconstruction methods can effectively reduce noise with less computing time, and achieve the best performances on both real-world and simulated datasets.
Yajing Zheng, Lingxiao Zheng, Zhaofei Yu, Tiejun Huang 0001, Song Wang 0002
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 PLGAN: Generative Adversarial Networks for Power-Line Segmentation in Aerial Images
abstract
Accurate segmentation of power lines in various aerial images is very important for UAV flight safety. The complex background and very thin structures of power lines, however, make it an inherently difficult task in computer vision. This paper presents PLGAN, a simple yet effective method based on generative adversarial networks, to segment power lines from aerial images with different backgrounds. Instead of directly using the adversarial networks to generate the segmentation, we take their certain decoding features and embed them into another semantic segmentation network by considering more context, geometry, and appearance information of power lines. We further exploit the appropriate form of the generated images for high-quality feature embedding and define a new loss function in the Hough-transform parameter space to enhance the segmentation of very thin power lines. Extensive experiments and comprehensive analysis demonstrate that our proposed PLGAN outperforms the prior state-of-the-art methods for semantic segmentation and line detection.
Rabab Abdelfattah, Xiaofeng Wang 0007, Song Wang 0002
IEEE Trans. Image Process.3
2023 Spike-Based Motion Estimation for Object Tracking Through Bio-Inspired Unsupervised Learning
abstract
Neuromorphic vision sensors, whose pixels output events/spikes asynchronously with a high temporal resolution according to the scene radiance change, are naturally appropriate for capturing high-speed motion in the scenes. However, how to utilize the events/spikes to smoothly track high-speed moving objects is still a challenging problem. Existing approaches either employ time-consuming iterative optimization, or require large amounts of labeled data to train the object detector. To this end, we propose a bio-inspired unsupervised learning framework, which takes advantage of the spatiotemporal information of events/spikes generated by neuromorphic vision sensors to capture the intrinsic motion patterns. Without off-line training, our models can filter the redundant signals with dynamic adaption module based on short-term plasticity, and extract the motion patterns with motion estimation module based on the spike-timing-dependent plasticity. Combined with the spatiotemporal and motion information of the filtered spike stream, the traditional DBSCAN clustering algorithm and Kalman filter can effectively track multiple targets in extreme scenes. We evaluate the proposed unsupervised framework for object detection and tracking tasks on synthetic data, publicly available event-based datasets, and spiking camera datasets. The experiment results show that the proposed model can robustly detect and smoothly track the moving targets on various challenging scenarios and outperforms state-of-the-art approaches.
Yajing Zheng, Zhaofei Yu, Song Wang 0002, Tiejun Huang 0001
IEEE Trans. Image Process.3
2023 Information Maximizing Adaptation Network With Label Distribution Priors for Unsupervised Domain Adaptation
abstract
Unsupervised domain adaptation, which transfers knowledge from the source domain to the target domain, has still been a challenging problem. However, previous domain adaptation methods typically minimize the domain discrepancy by using the pseudo target labels. Since the pseudo labels can be noisy, which may cause misalignment and unsatisfying adaptation performance. To address the above challenges, we propose an information maximization adaptation network with label distribution priors. We revisit feature alignment in unsupervised domain adaptation from the perspective of distribution alignment, and find that learning discriminant feature representation requires to minimizing distribution discrepancy and maximizing source mutual information between the outputs of the classifier and feature representations. Due to domain shift, maximizing target mutual information may align features to incorrect class directly. We propose a weighted target mutual information by re-weighting the estimated mutual information via the mean prediction confidence in mini-batch, which can eliminate the negative impact of inaccurate estimation. In addition, we introduce a regularization term of label priors distribution to encourage the similarity to the real label distribution. Extensive experimental results on three benchmark datasets show that our proposed method can achieve remarkable results compared with previous methods.
Pei Wang 0016, Yun Yang 0003, Yuelong Xia, Xingyi Zhang 0001, Song Wang 0002
IEEE Trans. Multim.6
2022 Style Mixing and Patchwise Prototypical Matching for One-Shot Unsupervised Domain Adaptive Semantic Segmentation
abstract
In this paper, we tackle the problem of one-shot unsupervised domain adaptation (OSUDA) for semantic segmentation where the segmentors only see one unlabeled target image during training. In this case, traditional unsupervised domain adaptation models usually fail since they cannot adapt to the target domain with over-fitting to one (or few) target samples. To address this problem, existing OSUDA methods usually integrate a style-transfer module to perform domain randomization based on the unlabeled target sample, with which multiple domains around the target sample can be explored during training. However, such a style-transfer module relies on an additional set of images as style reference for pre-training and also increases the memory demand for domain adaptation. Here we propose a new OSUDA method that can effectively relieve such computational burden. Specifically, we integrate several style-mixing layers into the segmentor which play the role of style-transfer module to stylize the source images without introducing any learned parameters. Moreover, we propose a patchwise prototypical matching (PPM) method to weighted consider the importance of source pixels during the supervised training to relieve the negative adaptation. Experimental results show that our method achieves new state-of-the-art performance on two commonly used benchmarks for domain adaptive semantic segmentation under the one-shot setting and is more efficient than all comparison approaches.
Xinyi Wu 0002, Zhenyao Wu, Lili Ju, Song Wang 0002
AAAI5
2022 G2Net: Generic Game-Theoretic Network for Partial-Label Image Classification
Rabab Abdelfattah, Mostafa Fouda, Xiaofeng Wang 0007, Song Wang 0002
BMVC5
2022 Can You Spot the Chameleon? Adversarially Camouflaging Images from Co-Salient Object Detection
abstract
Co-salient object detection (CoSOD) has recently achieved significant progress and played a key role in retrieval-related tasks. However, it inevitably poses an entirely new safety and security issue, i.e., highly personal and sensitive content can potentially be extracting by powerful CoSOD methods. In this paper, we address this problem from the perspective of adversarial attacks and identify a novel task: adversarial co-saliency attack. Specially, given an image selected from a group of images containing some common and salient objects, we aim to generate an adversarial version that can mislead CoSOD methods to predict incorrect co-salient regions. Note that, compared with general white-box adversarial attacks for classification, this new task faces two additional challenges: (1) low success rate due to the diverse appearance of images in the group; (2) low transferability across CoSOD methods due to the considerable difference between CoSOD pipelines. To address these challenges, we propose the very first blackbox joint adversarial exposure and noise attack (Jadena), where we jointly and locally tune the exposure and additive perturbations of the image according to a newly designed high-feature-level contrast-sensitive loss function. Our method, without any information on the state-of-the-art CoSOD methods, leads to significant performance degradation on various co-saliency detection datasets and makes the co-salient objects undetectable. This can have strong practical benefits in properly securing the large number of personal photos currently shared on the Internet. Moreover, our method is potential to be utilized as a metric for evaluating the robustness of CoSOD methods.
Ruijun Gao, Qing Guo 0005, Felix Juefei-Xu, Hongkai Yu, Huazhu Fu, Wei Feng 0005, Yang Liu 0003, Song Wang 0002
CVPR8
2022 Connecting the Complementary-view Videos: Joint Camera Identification and Subject Association
abstract
We attempt to connect the data from complementary views, i.e., top view from drone-mounted cameras in the air, and side view from wearable cameras on the ground. Collaborative analysis of such complementary-view data can facilitate to build the air-ground cooperative visual system for various kinds of applications. This is a very challenging problem due to the large view difference between top and side views. In this paper, we develop a new approach that can simultaneously handle three tasks: i) localizing the side-view camera in the top view; ii) estimating the view direction of the side-view camera; iii) detecting and associating the same subjects on the ground across the complementary views. Our main idea is to explore the spatial position layout of the subjects in two views. In particular, we propose a spatial-aware position representation method to embed the spatial-position distribution of the subjects in different views. We further design a cross-view video collaboration framework composed of a camera identification module and a subject association module to simultaneously perform the above three tasks. We collect a new synthetic dataset consisting of top-view and side-view video sequence pairs for performance evaluation and the experimental results show the effectiveness of the proposed method.
Rui-Ze Han, Yiyang Gan, Wei Feng 0005, Song Wang 0002
CVPR6
2022 MISF: Multi-level Interactive Siamese Filtering for High-Fidelity Image Inpainting
abstract
Although achieving significant progress, existing deep generative inpainting methods still show low generalization across different scenes. As a result, the generated images usually contain artifacts or the filled pixels differ greatly from the ground truth, making them far from real-world applications. Image-level predictive filtering is a widely used restoration technique by predicting suitable kernels adaptively according to different input scenes. Inspired by this inherent advantage, we explore the possibility of addressing image inpainting as a filtering task. To this end, we first study the advantages and challenges of the image-level predictive filtering for inpainting: the method can preserve local structures and avoid artifacts but fails to fill large missing areas. Then, we propose the semantic filtering by conducting filtering on deep feature level, which fills the missing semantic information but fails to recover the details. To address the issues while adopting the respective advantages, we propose a novel filtering technique, i.e., Multi-level Interactive Siamese Filtering (MISF) containing two branches: kernel prediction branch (KPB) and semantic & image filtering branch (SIFB). These two branches are interactively linked: SIFB provides multi-level features for KPB while KPB predicts dynamic kernels for SIFB. As a result, the final method takes the advantage of effective semantic & image-level filling for high-fidelity inpainting. Moreover, we discuss the relationship between MISF and the naive encoder-decoder-based inpainting, inferring that MISF provides novel dynamic convolutional operations to enhance the high generalization capability across scenes. We validate our method on three challenging datasets, i.e., Dunhuang, Places2, and CelebA. Our method outperforms state-of-the-art baselines on four metrics, i.e.,$L_{1}$, PSNR, SSIM, and LPIPS.
Qing Guo 0005, Di Lin 0002, Ping Li 0016, Wei Feng 0005, Song Wang 0002
CVPR6
2022 Panoramic Human Activity Recognition
Rui-Ze Han, Haomin Yan, Song Wang 0002, Wei Feng 0005
ECCV (4)4
2022 Self-supervised Social Relation Representation for Human Group Detection
Rui-Ze Han, Haomin Yan, Zekun Qian, Wei Feng 0005, Song Wang 0002
ECCV (35)6
2022 Style-Guided Shadow Removal
Jin Wan, Hui Yin 0002, Zhenyao Wu, Xinyi Wu 0002, Song Wang 0002
ECCV (19)6
2022 Rethinking Video Rain Streak Removal: A New Synthesis Model and a Deraining Network with Video Rain Prior
Lei Zhu 0003, Huazhu Fu, Harry Qin, Carola-Bibiane Schönlieb, Wei Feng 0005, Song Wang 0002
ECCV (19)7
2022 Is It Necessary to Transfer Temporal Knowledge for Domain Adaptive Video Semantic Segmentation?
Xinyi Wu 0002, Zhenyao Wu, Jin Wan, Lili Ju, Song Wang 0002
ECCV (27)5
2022 SiamDoGe: Domain Generalizable Semantic Segmentation Using Siamese Network
Zhenyao Wu, Xinyi Wu 0002, Xiaoping Zhang 0005, Lili Ju, Song Wang 0002
ECCV (38)5
2022 Background-Insensitive Scene Text Recognition with Text Semantic Segmentation
Zhenyao Wu, Xinyi Wu 0002, Greg Wilsbacher, Song Wang 0002
ECCV (25)5
2022 Self-Supervised Representation Learning for Skeleton-Based Group Activity Recognition
abstract
Group activity recognition (GAR) is a challenging task for discerning the behavior of a group of actors. This paper aims at learning discriminative representation for GAR in a self-supervised manner based on human skeletons. As modeling relations between actors lie at the center of GAR, we propose a valid self-supervised learning pretext task with a matching framework, where a representation model is driven to identify subgroups in a synthetic group based on actors' skeleton sequences. For backbone networks, while spatial-temporal graph convolution networks have dominated the skeleton-based action recognition, they under-explore the group relevant interactions among actors. To address this issue, we come up with a novel plug-in Actor-Association Graph Convolution Module (AAGCM) based on inductive graph convolution, which can be integrated into many common backbones. It can not only model the interactions at different levels but also adapt to variable group sizes. The effectiveness of our approaches is demonstrated by extensive experiments on three benchmark datasets: Volleyball, Collective Activity, and Mutual NTU.
Cunling Bian, Wei Feng 0005, Song Wang 0002
ACM Multimedia3
2022 Video Instance Lane Detection via Deep Temporal and Geometry Consistency Constraints
abstract
Video instance lane detection is one of the most important tasks in autonomous driving.Due to the very sparse region and weak context in lane annotations, accurately detecting instance-level lanes in real-world traffic scenarios is challenging, especially for scenes with occlusion, bad weather conditions, dim or dazzling lights.Current methods mainly address this problem by integrating features of adjacent video frames to simply encourage temporal constancy for image-level lane detectors. However, most of them ignore lane shape constraint of adjacent frames and geometry consistency of individual lanes, thereby harming the performance of video instance lane detection. In this paper, we propose TGC-Net via temporal and geometry consistency constraints for reliable video instance lane detection. Specifically, we devise a temporal recurrent feature-shift aggregation module (T-RESA) to learn spatio-temporal lane features along horizontal, vertical, and temporal directions of the feature tensor. We further impose temporal consistency constraint by encouraging spatial distribution consistency among the lane features of adjacent frames. Besides, we devise two effective geometry constraints to ensure the integrity and continuity of lane predictions by leveraging pairwise point affinity loss and vanishing point guided geometric context, respectively. Extensive experiments on public benchmark dataset show that our TGC-Net quantitatively and qualitatively outperforms state-of-the-art video instance lane detectors and video object segmentation competitors. Our code and our results have been released at https://github.com/wmq12345/TGC-Net.
Yujun Zhang 0002, Wei Feng 0005, Lei Zhu 0003, Song Wang 0002
ACM Multimedia5
2022 Self-Supervised Human Pose based Multi-Camera Video Synchronization
abstract
Multi-view video collaborative analysis is an important task and has many applications in multimedia community. However, it always requires the given multiple videos to be temporally synchronized. Existing methods commonly synchronize the videos by the wired communication, which may hinder the practical application in real world, especially for moving cameras. In this paper, we focus on the human-centric video analysis and propose a self-supervised framework for the automatic multi-camera video synchronization. Specifically, we develop SeSyn-Net with the 2D human pose as input for feature embedding and design a series of self-supervised losses to effectively extract the view-invariant but time-discriminative representation for video synchronization. We also build two new datasets for the performance evaluation. Extensive experimental results verify the effectiveness of our method, which achieves the superior performance compared to both the classical and state-of-the-art methods.
Liqiang Yin, Rui-Ze Han, Wei Feng 0005, Song Wang 0002
ACM Multimedia4
2022 Crossmodal Few-shot 3D Point Cloud Semantic Segmentation
abstract
Recently, few-shot 3D point cloud semantic segmentation methods have been introduced to mitigate the limitations of existing fully supervised approaches, i.e., heavy dependence on labeled 3D data and poor capacity to generalize to new categories. However, those few-shot learning methods need one or few labeled data as support for testing. In practice, such data labeling usually requires manual annotation of large-scale points in 3D space, which can be very difficult and laborious. To address this problem, in this paper we introduce a novel crossmodal few-shot learning approach for 3D point cloud semantic segmentation. In this approach, the point cloud to be segmented is taken as query while one or few labeled 2D RGB images are taken as support to guide the segmentation of query. This way, we only need to annotate on a few 2D support images for the categories of interest. Specifically, we first convert the 2D support images into 3D point cloud format based on both appearance and the estimated depth information. We then introduce a co-embedding network for extracting the features of support and query, both from 3D point cloud format, to fill their domain gap. Finally, we compute the prototypes of support and employ cosine similarity between the prototypes and the query features for final segmentation. Experimental results on two widely-used benchmarks show that, with one or few labeled 2D images as support, our proposed method achieves competitive results against existing few-shot 3D point cloud semantic segmentation methods.
Zhenyao Wu, Xinyi Wu 0002, Canyu Zhang 0002, Song Wang 0002
ACM Multimedia5
2022 Visual Attention Consistency for Human Attribute Recognition
Hao Guo 0002, Xiaochuan Fan, Song Wang 0002
Int. J. Comput. Vis.3
2022 Snowvision: Segmenting, Identifying, and Discovering Stamped Curve Patterns from Fragments of Pottery
Sam T. McDorman, Canyu Zhang 0002, Deja Scott, Jake Bukuts, Colin Wilder, Karen Y. Smith, Song Wang 0002
Int. J. Comput. Vis.9
2022 Multiple Human Association and Tracking From Egocentric and Complementary Top Views
abstract
Crowded scene surveillance can significantly benefit from combining egocentric-view and its complementary top-view cameras. A typical setting is an egocentric-view camera, e.g., a wearable camera on the ground capturing rich local details, and a top-view camera, e.g., a drone-mounted one from high altitude providing a global picture of the scene. To collaboratively analyze such complementary-view videos, an important task is to associate and track multiple people across views and over time, which is challenging and differs from classical human tracking, since we need to not only track multiple subjects in each video, but also identify the same subjects across the two complementary views. This paper formulates it as a constrained mixed integer programming problem, wherein a major challenge is how to effectively measure subjects similarity over time in each video and across two views. Although appearance and motion consistencies well apply to over-time association, they are not good at connecting two highly different complementary views. To this end, we present a spatial distribution based approach to reliable cross-view subject association. We also build a dataset to benchmark this new challenging task. Extensive experiments verify the effectiveness of our method.
Rui-Ze Han, Wei Feng 0005, Yujun Zhang 0002, Jiewen Zhao, Song Wang 0002
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Recognizing micro actions in videos by learning multi-layer local features
Yang Mi, Song Wang 0002
Pattern Recognit. Lett.4
2022 Let There Be Light: Improved Traffic Surveillance via Detail Preserving Night-to-Day Transfer
abstract
In recent years, image and video surveillance have made considerable progresses to the Intelligent Transportation Systems (ITS) with the help of deep Convolutional Neural Networks (CNNs). As one of the state-of-the-art perception approaches, detecting the interested objects in each frame of video surveillance is widely desired by ITS. Currently, object detection shows remarkable efficiency and reliability in standard scenarios such as daytime scenes with favorable illumination conditions. However, in face of adverse conditions such as the nighttime, object detection loses its accuracy significantly. One of the main causes of the problem is the lack of sufficient annotated detection datasets of nighttime scenes. In this paper, we propose a framework to alleviate the accuracy decline when object detection is taken to adverse conditions by using image translation method. We propose to utilize style translation based StyleMix method to acquire pairs of day time image and nighttime image as training data for following nighttime to daytime image translation. To alleviate the detail corruptions caused by Generative Adversarial Networks (GANs), we propose to utilize Kernel Prediction Network (KPN) based method to refine the nighttime to daytime image translation. The KPN network is trained with object detection task together to adapt the trained daytime model to nighttime vehicle detection directly. Experiments on vehicle detection verified the accuracy and effectiveness of the proposed approach.
Lan Fu, Hongkai Yu, Felix Juefei-Xu, Qing Guo 0005, Song Wang 0002
IEEE Trans. Circuits Syst. Video Technol.6
2022 Two-Stage Selective Ensemble of CNN via Deep Tree Training for Medical Image Classification
abstract
Medical image classification is an important task in computer-aided diagnosis systems. Its performance is critically determined by the descriptiveness and discriminative power of features extracted from images. With rapid development of deep learning, deep convolutional neural networks (CNNs) have been widely used to learn the optimal high-level features from the raw pixels of images for a given classification task. However, due to the limited amount of labeled medical images with certain quality distortions, such techniques crucially suffer from the training difficulties, including overfitting, local optimums, and vanishing gradients. To solve these problems, in this article, we propose a two-stage selective ensemble of CNN branches via a novel training strategy called deep tree training (DTT). In our approach, DTT is adopted to jointly train a series of networks constructed from the hidden layers of CNN in a hierarchical manner, leading to the advantage that vanishing gradients can be mitigated by supplementing gradients for hidden layers of CNN, and intrinsically obtain the base classifiers on the middle-level features with minimum computation burden for an ensemble solution. Moreover, the CNN branches as base learners are combined into the optimal classifier via the proposed two-stage selective ensemble approach based on both accuracy and diversity criteria. Extensive experiments on CIFAR-10 benchmark and two specific medical image datasets illustrate that our approach achieves better performance in terms of accuracy, sensitivity, specificity, and F1 score measurement.
Yun Yang 0003, Xingyi Zhang 0001, Song Wang 0002
IEEE Trans. Cybern.4
2022 Multi-View Multi-Human Association With Deep Assignment Network
abstract
Identifying the same persons across different views plays an important role in many vision applications. In this paper, we study this important problem, denoted as Multi-view Multi-Human Association (MvMHA), on multi-view images that are taken by different cameras at the same time. Different from previous works on human association across two views, this paper is focused on more general and challenging scenarios of more than two views, and none of these views are fixed or priorly known. In addition, each involved person may be present in all the views or only a subset of views, which are also not priorly known. We develop a new end-to-end deep-network based framework to address this problem. First, we use an appearance-based deep network to extract the feature of each detected subject on each image. We then compute pairwise-similarity scores between all the detected subjects and construct a comprehensive affinity matrix. Finally, we propose a Deep Assignment Network (DAN) to transform the affinity matrix into an assignment matrix, which provides a binary assignment result for MvMHA. We build both a synthetic dataset and a real image dataset to verify the effectiveness of the proposed method. We also test the trained network on other three public datasets, resulting in very good cross-domain performance.
Rui-Ze Han, Haomin Yan, Wei Feng 0005, Song Wang 0002
IEEE Trans. Image Process.5
2022 Deep Domain Adaptation Based Multi-Spectral Salient Object Detection
abstract
Salient Object Detection (SOD) plays an important role in many image-related multimedia applications. Although there are many existing research works about the salient object detection in traditional RGB (visible-light spectrum) images, there are still many complex situations that regular RGB images cannot provide enough cues for the accurate SOD, such as the shadow effect, similar appearance between background and foreground, strong or insufficient illumination, etc. Because of the success of near-infrared spectrum in many computer vision tasks, we explore the multi-spectral SOD in the synchronized RGB images and near-infrared (NIR) images for the both simple and complex situations. We assume that the RGB SOD in the existing RGB image datasets could provide references for the multi-spectral SOD problem. In this paper, we mainly model this research problem as a deep learning based domain adaptation from the traditional RGB image data (source domain) to the multi-spectral data (target domain), and an adversarial deep domain adaptation model is proposed. We first collect and will publicize a large multi-spectral dataset, RGBN-SOD dataset, including 780 synchronized RGB and NIR image pairs for the multi-spectral SOD problem in the simple and complex situations. Intensive experimental results show the effectiveness and accuracy of the proposed deep domain adaptation for the multi-spectral SOD. Besides, due to the absence of research on the field of multi-spectral co-saliency detection, we also collect 200 synchronized RGB and NIR image pairs in addition to explore the multi-spectral co-saliency detection.
Shaoyue Song, Zhenjiang Miao, Hongkai Yu, Jianwu Fang, Cong Ma 0004, Song Wang 0002
IEEE Trans. Multim.7
2022 Transductive Zero-Shot Hashing for Multilabel Image Retrieval
abstract
Hash coding has been widely used in the approximate nearest neighbor search for large-scale image retrieval. Given semantic annotations such as class labels and pairwise similarities of the training data, hashing methods can learn and generate effective and compact binary codes. While some newly introduced images may contain undefined semantic labels, which we call unseen images, zero-shot hashing (ZSH) techniques have been studied for retrieval. However, existing ZSH methods mainly focus on the retrieval of single-label images and cannot handle multilabel ones. In this article, for the first time, a novel transductive ZSH method is proposed for multilabel unseen image retrieval. In order to predict the labels of the unseen/target data, a visual-semantic bridge is built via instance-concept coherence ranking on the seen/source data. Then, pairwise similarity loss and focal quantization loss are constructed for training a hashing model using both the seen/source and unseen/target data. Extensive evaluations on three popular multilabel data sets demonstrate that the proposed hashing method achieves significantly better results than the comparison methods.
Qin Zou 0001, Ling Cao, Zheng Zhang 0036, Long Chen 0005, Song Wang 0002
IEEE Trans. Neural Networks Learn. Syst.5
2021 Multi-Domain Multi-Task Rehearsal for Lifelong Learning
abstract
Rehearsal, seeking to remind the model by storing old knowledge in lifelong learning, is one of the most effective ways to mitigate catastrophic forgetting, i.e., biased forgetting of previous knowledge when moving to new tasks. However, the old tasks of the most previous rehearsal-based methods suffer from the unpredictable domain shift when training the new task. This is because these methods always ignore two significant factors. First, the Data Imbalance between the new task and old tasks that makes the domain of old tasks prone to shift. Second, the Task Isolation among all tasks will make the domain shift toward unpredictable directions; To address the unpredictable domain shift, in this paper, we propose Multi-Domain Multi-Task (MDMT) rehearsal to train the old tasks and new task parallelly and equally to break the isolation among tasks. Specifically, a two-level angular margin loss is proposed to encourage the intra-class/task compactness and inter-class/task discrepancy, which keeps the model from domain chaos. In addition, to further address domain shift of the old tasks, we propose an optional episodic distillation loss on the memory to anchor the knowledge for each old task. Experiments on benchmark datasets validate the proposed approach can effectively mitigate the unpredictable domain shift.
Fan Lyu, Wei Feng 0005, Zihan Ye, Fuyuan Hu, Song Wang 0002
AAAI6
2021 Binaural Audio-Visual Localization
abstract
Localizing sound sources in a visual scene has many important applications and quite a few traditional or learning-based methods have been proposed for this task. Humans have the ability to roughly localize sound sources within or beyond the range of the vision using their binaural system. However most existing methods use monaural audio, instead of binaural audio, as a modality to help the localization. In addition, prior works usually localize sound sources in the form of object-level bounding boxes in images or videos and evaluate the localization accuracy by examining the overlap between the ground-truth and predicted bounding boxes. This is too rough since a real sound source is often only a part of an object. In this paper, we propose a deep learning method for pixel-level sound source localization by leveraging both binaural recordings and the corresponding videos. Specifically, we design a novel Binaural Audio-Visual Network (BAVNet), which concurrently extracts and integrates features from binaural recordings and videos. We also propose a point-annotation strategy to construct pixel-level ground truth for network training and performance evaluation. Experimental results on Fair-Play and YT-Music datasets demonstrate the effectiveness of the proposed method and show that binaural audio can greatly improve the performance of localizing the sound sources, especially when the quality of the visual information is limited.
Xinyi Wu 0002, Zhenyao Wu, Lili Ju, Song Wang 0002
AAAI4
2021 Long-Tailed Multi-Label Visual Recognition by Collaborative Training on Uniform and Re-Balanced Samplings
abstract
Long-tailed data distribution is common in many multi-label visual recognition tasks and the direct use of these data for training usually leads to relatively low performance on tail classes. While re-balanced data sampling can improve the performance on tail classes, it may also hurt the performance on head classes in training due to label co-occurrence. In this paper, we propose a new approach to train on both uniform and re-balanced samplings in a collaborative way, resulting in performance improvement on both head and tail classes. More specifically, we design a visual recognition network with two branches: one takes the uniform sampling as input while the other takes the re-balanced sampling as the input. For each branch, we conduct visual recognition using a binary-cross-entropy-based classification loss with learnable logit compensation. We further define a new cross-branch loss to enforce the consistency when the same input image goes through the two branches. We conduct extensive experiments on VOC-LT and COCO-LT datasets. The results show that the proposed method significantly outperforms previous state-of-the-art methods on long-tailed multi-label visual recognition.
Hao Guo 0002, Song Wang 0002
CVPR2
2021 Auto-Exposure Fusion for Single-Image Shadow Removal
abstract
Shadow removal is still a challenging task due to its inherent background-dependent1and spatial-variant properties, leading to unknown and diverse shadow patterns. Even powerful deep neural networks could hardly recover traceless shadow-removed background. This paper proposes a new solution for this task by formulating it as an exposure fusion problem to address the challenges. Intuitively, we first estimate multiple over-exposure images w.r.t. the input image to let the shadow regions in these images have the same color with shadow-free areas in the input image. Then, we fuse the original input with the over-exposure images to generate the final shadow-free counterpart. Nevertheless, the spatial-variant property of the shadow requires the fusion to be sufficiently ‘smart’, that is, it should automatically select proper over-exposure pixels from different images to make the final output natural. To address this challenge, we propose the shadow-aware FusionNet that takes the shadow image as input to generate fusion weight maps across all the over-exposure images. Moreover, we propose the boundary-aware RefineNet to eliminate the remaining shadow trace further. We conduct extensive experiments on the ISTD, ISTD+, and SRD datasets to validate our method’s effectiveness and show better performance in shadow regions and comparable performance in non-shadow regions over the state-of-the-art methods. We release the code in https://github.com/tsingqguo/exposure-fusion-shadow-removal.
Lan Fu, Changqing Zhou, Qing Guo 0005, Felix Juefei-Xu, Hongkai Yu, Wei Feng 0005, Yang Liu 0003, Song Wang 0002
CVPR8
2021 From Shadow Generation To Shadow Removal
abstract
Shadow removal is a computer-vision task that aims to restore the image content in shadow regions. While almost all recent shadow-removal methods require shadow-free images for training, in ECCV 2020 Le and Samaras introduces an innovative approach without this requirement by cropping patches with and without shadows from shadow images as training samples. However, it is still laborious and time-consuming to construct a large amount of such unpaired patches. In this paper, we propose a new G2R-ShadowNet which leverages shadow generation for weakly-supervised shadow removal by only using a set of shadow images and their corresponding shadow masks for training. The proposed G2R-ShadowNet consists of three sub-networks for shadow generation, shadow removal and refinement, respectively and they are jointly trained in an end-to-end fashion. In particular, the shadow generation sub-net stylises non-shadow regions to be shadow ones, leading to paired data for training the shadow-removal sub-net. Extensive experiments on the ISTD dataset and the Video Shadow Removal dataset show that the proposed G2R-ShadowNet achieves competitive performances against the current state of the arts and outperforms Le and Samaras’ patch-based shadow-removal method.
Hui Yin 0002, Xinyi Wu 0002, Zhenyao Wu, Yang Mi, Song Wang 0002
CVPR6
2021 DANNet: A One-Stage Domain Adaptation Network for Unsupervised Nighttime Semantic Segmentation
abstract
Semantic segmentation of nighttime images plays an equally important role as that of daytime images in autonomous driving, but the former is much more challenging due to poor illuminations and arduous human annotations. In this paper, we propose a novel domain adaptation network (DANNet) for nighttime semantic segmentation without using labeled nighttime image data. It employs an adversarial training with a labeled daytime dataset and an unlabeled dataset that contains coarsely aligned day-night image pairs. Specifically, for the unlabeled day-night image pairs, we use the pixel-level predictions of static object categories on a daytime image as a pseudo supervision to segment its counterpart nighttime image. We further design a re-weighting strategy to handle the inaccuracy caused by misalignment between day-night image pairs and wrong predictions of daytime images, as well as boost the prediction accuracy of small objects. The proposed DANNet is the first one-stage adaptation framework for nighttime semantic segmentation, which does not train additional day-night image transfer models as a separate pre-processing stage. Extensive experiments on Dark Zurich and Nighttime Driving datasets show that our method achieves state-of-the-art performance for nighttime semantic segmentation.
Xinyi Wu 0002, Zhenyao Wu, Hao Guo 0002, Lili Ju, Song Wang 0002
CVPR5
2021 Multiple Human Tracking in Non-Specific Coverage with Wearable Cameras
abstract
Compared to fixed cameras, wearable cameras have time-varying non-specific view coverage and can be used to alternately observe people at different sites by varying the camera views. However, such view change of wearable cameras may introduce intervals of transitional frames without useful information, which brings new challenge for the important multiple object tracking (MOT) task – existing MOT methods can not handle well frequent disappearing/reappearing targets in the field of view, especially in the presence of informationless transitional sequences of frames. To address this problem, in this paper we propose a Markov Decision Process with jump state (JMDP) to model the target’s lifetime in tracking, and use optical flow of the camera motion and the statistical information of the targets to model the camera state transition. We further develop a frame-level classification algorithm to locate the transitional sequence. By combining all of them, we formulate the proposed non-specific-coverage MOT problem as a joint state transition problem, which can be solved by the state transfer mechanism of the targets and the camera. We collect a new dataset for performance evaluation and the experimental results show the effectiveness of the proposed method.
Sibo Wang 0008, Rui-Ze Han, Wei Feng 0005, Song Wang 0002
ICASSP4
2021 VIL-100: A New Dataset and A Baseline Model for Video Instance Lane Detection
abstract
Lane detection plays a key role in autonomous driving. While car cameras always take streaming videos on the way, current lane detection works mainly focus on individual images (frames) by ignoring dynamics along the video. In this work, we collect a new video instance lane detection (VIL-100) dataset, which contains 100 videos with in total 10,000 frames, acquired from different real traffic scenarios. All the frames in each video are manually annotated to a high-quality instance-level lane annotation, and a set of frame-level and video-level metrics are included for quantitative performance evaluation. Moreover, we propose a new baseline model, named multi-level memory aggregation network (MMA-Net), for video instance lane detection. In our approach, the representation of current frame is enhanced by attentively aggregating both local and global memory features from other frames. Experiments on the new collected dataset show that the proposed MMA-Net outperforms state-of-the-art lane detection methods and video object segmentation methods. We release our dataset and code at https://github.com/yujun0-0/MMA-Net.
Yujun Zhang 0002, Lei Zhu 0003, Wei Feng 0005, Huazhu Fu, Qingxia Li, Song Wang 0002
ICCV8
2021 Learning Depth from Single Image Using Depth-Aware Convolution and Stereo Knowledge
abstract
Estimating depth from a monocular image has become a very popular task in computer vision for identifying important geometric information of the scene. While its performance has been significantly improved by convolutional neural networks (CNNs) in recent years, depth-estimation accuracy is still unsatisfactory at locations with abrupt depth changes. This is mainly caused by the use of spatially consistent filters in CNNs which directly mix the features of different objects when applied to the pixels near the object borders. Moreover, the performance gap between depth estimation from single image and that from a stereo pair remains quite large due to the ill-posed nature of the former one. In this paper, we propose a new depth-aware convolutional neural network (DACNN) to address these issues. We first design a novel depth-aware convolution operation for DACNN, that can adaptively choose subsets of relevant features for convolutions at each location. Specifically, we compute hierarchical depth features as the guidance, and then estimate the depth map using such depth-aware convolution which can leverage the guidance to adapt the filters. In addition, we also introduce a pre-trained stereo network into DACNN as the teacher to carry out knowledge distillation on the student monocular network with a specially designed loss function. Experimental results on the KITTI online benchmark and Eigen split datasets show that the proposed method achieves the state-of-the-art performance for single-image depth estimation.
Zhenyao Wu, Xinyi Wu 0002, Xiaoping Zhang 0005, Song Wang 0002, Lili Ju
ICME4
2021 CPNet: Cycle Prototype Network for Weakly-Supervised 3D Renal Compartments Segmentation on CT Images
Song Wang 0002, Yuting He 0001, Youyong Kong, Xiaomei Zhu, Shaobo Zhang 0008, Jean-Louis Dillenseger, Jean-Louis Coatrieux, Shuo Li 0001, Guanyu Yang 0001
MICCAI (2)1
2021 Self-supervised Multi-view Multi-Human Association and Tracking
abstract
Multi-view Multi-human association and tracking (MvMHAT) aims to track a group of people over time in each view, as well as to identify the same person across different views at the same time. This is a relatively new problem but is very important for multi-person scene video surveillance. Different from previous multiple object tracking (MOT) and multi-target multi-camera tracking (MTMCT) tasks, which only consider the over-time human association, MvMHAT requires to jointly achieve both cross-view and over-time data association. In this paper, we model this problem with a self-supervised learning framework and leverage an end-to-end network to tackle it. Specifically, we propose a spatial-temporal association network with two designed self-supervised learning losses, including a symmetric-similarity loss and a transitive-similarity loss, at each time to associate the multiple humans over time and across views. Besides, to promote the research on MvMHAT, we build a new large-scale benchmark for the training and testing of different algorithms. Extensive experiments on the proposed benchmark verify the effectiveness of our method. We have released the benchmark and code to the public.
Yiyang Gan, Rui-Ze Han, Liqiang Yin, Wei Feng 0005, Song Wang 0002
ACM Multimedia5
2021 JPGNet: Joint Predictive Filtering and Generative Network for Image Inpainting
abstract
Image inpainting aims to restore the missing regions of corrupted images and make the recovery result identical to the originally complete image, which is different from the common generative task emphasizing the naturalness or realism of generated images. Nevertheless, existing works usually regard it as a pure generation problem and employ cutting-edge deep generative techniques to address it. The generative networks can fill the main missing parts with realistic contents but usually distort the local structures or introduce obvious artifacts. In this paper, for the first time, we formulate image inpainting as a mix of two problems, i.e., predictive filtering and deep generation. Predictive filtering is good at preserving local structures and removing artifacts but falls short to complete the large missing regions. The deep generative network can fill the numerous missing pixels based on the understanding of the whole scene but hardly restores the details identical to the original ones. To make use of their respective advantages, we propose the joint predictive filtering and generative network (JPGNet) that contains three branches: predictive filtering & uncertainty network (PFUNet), deep generative network, and uncertainty-aware fusion network (UAFNet). The PFUNet can adaptively predict pixel-wise kernels for filtering-based inpainting according to the input image and output an uncertainty map. This map indicates the pixels should be processed by filtering or generative networks, which is further fed to the UAFNet for a smart combination between filtering and generative results. Note that, our method as a novel framework for the image inpainting problem can benefit any existing generation-based methods. We validate our method on three public datasets, i.e., Dunhuang, Places2, and CelebA, and demonstrate that our method can enhance three state-of-the-art generative methods (i.e., StructFlow, EdgeConnect, and RFRNet) significantly with slightly extra time costs. We have released the code at https://github.com/tsingqguo/jpgnet.
Qing Guo 0005, Felix Juefei-Xu, Hongkai Yu, Yang Liu 0003, Song Wang 0002
ACM Multimedia6
2021 Hypergraph Convolutional Network with Hybrid Higher-Order Neighbors
Fangyuan Lei, Senhong Wang, Song Wang 0002
PRCV (4)4
2021 Deep Poisoning: Towards Robust Image Data Sharing against Visual Disclosure
abstract
Due to respectively limited training data, different entities addressing the same vision task based on certain sensitive images may not train a robust deep network. This paper introduces a new vision task where various entities share task-specific image data to enlarge each other's training data volume without visually disclosing sensitive contents (e.g. illegal images). Then, we present a new structure-based training regime to enable different entities learn task-specific and reconstruction-proof image representations for image data sharing. Specifically, each entity learns a private Deep Poisoning Module (DPM) and insert it to a pre-trained deep network, which is designed to perform the specific vision task. The DPM deliberately poisons convolutional image features to prevent image reconstructions, while ensuring that the altered image data is functionally equivalent to the non-poisoned data for the specific vision task. Given this equivalence, the poisoned features shared from one entity could be used by another entity for further model refinement. Experimental results on image classification prove the efficacy of the proposed method.
Hao Guo 0002, Brian Dolhansky, Eric Hsin, Phong Dinh, Cristian Canton, Song Wang 0002
WACV6
2021 Video sketch: A middle-level representation for action recognition
Xingyuan Zhang, Yang Mi, Yanting Pei, Qi Zou 0001, Song Wang 0002
Appl. Intell.6
2021 Guest editorial: Graph learning for computer vision
abstract
Many fields in the real world involve a lot of structured data, such as social networks, transportation networks, communication networks etc., and their structures carry important information about the characteristics of the data. However, how to use its structural information to analyse and process the data efficiently has caused continuous research in the field. A graph provides an important means for dealing with structured data. It can describe the geometric structure of data intuitively and flexibly, especially in the representation of spatial irregular data. Graph learning refers to machine learning on graphs, which mainly utilises machine learning algorithms to extract the relevant features of graphs. In recent years, combined with specific applications, researchers have conducted in-depth research on graph learning and proposed various approaches. This Special Issue aims to introduce the latest studies in graph learning for computer vision and proposes new theories and approaches to solve the existing problems. It received a number of submissions from researchers in the field, which all went through a rigorous review process. After several rounds of review, six papers were accepted. These papers cover a variety of fields, such as medicine, remote sensing, and data mining. Specific tasks include image segmentation, knowledge graph reasoning, clustering, and image classification. These accepted papers are mainly divided into two categories. The first category covers the graph learning method guided by optimisation, which obtains the graph structure by establishing a clear model and solving the corresponding optimisation problem. The second category is the deep learning-oriented graph learning method, which combines convolutional neural network and graph neural network to construct the model. In the first paper of the Special Issue by Wang et al. entitled An Enhanced 3D U-Net with Graph-based Refining for Segmentation of Gastrointestinal Stromal Tumours, the authors propose a segmentation algorithm using an improved 3D U-Net to segment gastrointestinal stromal tumours. To enhance information transmission, multiple skip connections are attached into same size feature maps between an encoder and a decoder. Due to difficulties in tumour labelling and other reasons, the author transforms the small intestinal segmentation model into a gastrointestinal stromal tumour segmentation model. Since fully convolutional networks typically suffer from inaccuracies around the boundaries of small structures, the graph neural network is introduced to refine segmentation results. Experiments demonstrate that the proposed method presents superior performance over traditional U-Net. The second paper by Ma et al. entitled Hybrid Attention Mechanism for Few-Shot Relational Learning of Knowledge Graphs, develops a few-shot relationship learning framework. The authors first design an entity-enhanced encoder with weak attention networks and self-attention mechanisms to explore the influence of different levels for source entities. The local graph structure is then utilised to enhance the embedding of the source entity by combining explicit and implicit features. Finally, the model parameters are optimised to infer real entities in the candidate set of similar entities obtained by a loop-processing matching processor. The authors provide extensive experiments and confirm the excellent accuracy of the proposed model. The third paper by Zhao et al. entitled Incremental Multi-View Correlated Feature Learning Based on Non-Negative Matrix Factorization, studies multi-view data. The authors present an incremental multi-view correlated feature learning approach based on non-negative matrix factorization to analyse the uncorrelated items in each view. The algorithm separates uncorrelateditems across views and constructs incremental joint learning with uncorrelated and correlated features to study the common features for multi-view data. Subsequently, the authors design an incremental objective function and derive an effective updating scheme. The proposed method is proved to converge effectively, and its complexity is discussed. They evaluate the proposed solution on real-world datasets and report excellent performance in comparison with the existing state-of-the-art solutions. The fourth paper by Hu et al. entitled Complete/Incomplete Multi-view Subspace Clustering via Soft Block-Diagonal-Induced Regularizer, concentrates on complete and incomplete multi-view clustering problems. The proposed method adopts the self-representation model to individually construct the similarity graphs for each view. To fuse a shared affinity matrix for all views, the authors design the soft block-diagonal-induced regulariser to encourage the generation of a matrix with K diagonal blocks. Considering the incomplete multi-view data, the proposed method effectively utilises some indicator matrices to accurately mark the missing instances in each view. The authors analyse the complexity and convergence of the proposed method on four public datasets and demonstrate that it is better than the most advanced complete/incomplete clustering methods. The fifth paper by Gong et al. entitled Few-shot Learning with Relation Propagation and Constraint, aims to extract valuable information of pair-wise correlation between sparse training samples. The authors state that transductive relation propagation simply propagates the pair-wise relation without relation constraints. Thus, the paper develops a constrained relation–propagation network to capture the accurate relation so as to generate discriminative relational representations. To constrain the pair-wise relation, the proposed method introduces a relation constraint module to regularise the distilled relations between samples, which helps to calibrate the propagated correlation information. Extensive experiments conducted on several benchmark datasets indicate that the proposed method achieves remarkable performance compared to few-shot learning methods. The last paper by Guo et al. entitled CNN-Combined Graph Residual Network with Multilevel Feature Fusion for Hyperspectral Image Classification introduces graph convolutional networks to obtain more superpixel-level features with a topological structure. This paper develops an effective CNN-combined graph residual network with a multilevel feature fusion strategy. The main idea is to learn superpixeltopological information by using the graph residual network and pixel information by using the convolutional neural network. The strategy can adequately leverage the superpixel level and pixel-level features and capture the class boundary features, which further enhances the generalisation performance. Experiments report highly competitive performance in comparison to existing hyperspectral image classification approaches. The papers selected in this Special Issue highlight the extensive study of graph learning in computer vision. We hope that these papers can promote the theoretical study of graph learning as well as provide new ideas for more researchers who are committed to graph learning.
Qi Wang 0009, Hongkai Yu, Song Wang 0002, Jianzhe Lin
IET Comput. Vis.3
2021 Effects of Image Degradation and Degradation Removal to CNN-Based Image Classification
abstract
Just like many other topics in computer vision, image classification has achieved significant progress recently by using deep learning neural networks, especially the Convolutional Neural Networks (CNNs). Most of the existing works focused on classifying very clear natural images, evidenced by the widely used image databases, such as Caltech-256, PASCAL VOCs, and ImageNet. However, in many real applications, the acquired images may contain certain degradations that lead to various kinds of blurring, noise, and distortions. One important and interesting problem is the effect of such degradations to the performance of CNN-based image classification and whether degradation removal helps CNN-based image classification. More specifically, we wonder whether image classification performance drops with each kind of degradation, whether this drop can be avoided by including degraded images into training, and whether existing computer vision algorithms that attempt to remove such degradations can help improve the image classification performance. In this article, we empirically study those problems for nine kinds of degraded images-hazy images, motion-blurred images, fish-eye images, underwater images, low resolution images, salt-and-peppered images, images with white Gaussian noise, Gaussian-blurred images, and out-of-focus images. We expect this article can draw more interests from the community to study the classification of degraded images.
Yanting Pei, Qi Zou 0001, Xingyuan Zhang, Song Wang 0002
IEEE Trans. Pattern Anal. Mach. Intell.5
2021 Topological optimization of the DenseNet with pretrained-weights inheritance and genetic channel selection
Zhenyu Fang, Jinchang Ren, Stephen Marshall, Huimin Zhao 0001, Song Wang 0002, Xuelong Li 0001
Pattern Recognit.5
2021 Structural Knowledge Distillation for Efficient Skeleton-Based Action Recognition
abstract
Skeleton data have been extensively used for action recognition since they can robustly accommodate dynamic circumstances and complex backgrounds. To guarantee the action-recognition performance, we prefer to use advanced and time-consuming algorithms to get more accurate and complete skeletons from the scene. However, this may not be acceptable in time- and resource-stringent applications. In this paper, we explore the feasibility of using low-quality skeletons, which can be quickly and easily estimated from the scene, for action recognition. While the use of low-quality skeletons will surely lead to degraded action-recognition accuracy, in this paper we propose a structural knowledge distillation scheme to minimize this accuracy degradations and improve recognition model's robustness to uncontrollable skeleton corruptions. More specifically, a teacher which observes high-quality skeletons obtained from a scene is used to help train a student which only sees low-quality skeletons generated from the same scene. At inference time, only the student network is deployed for processing low-quality skeletons. In the proposed network, a graph matching loss is proposed to distill the graph structural knowledge at an intermediate representation level. We also propose a new gradient revision strategy to seek a balance between mimicking the teacher model and directly improving the student model's accuracy. Experiments are conducted on Kenetics400, NTU RGB+D and Penn action recognition datasets and the comparison results demonstrate the effectiveness of our scheme.
Cunling Bian, Wei Feng 0005, Song Wang 0002
IEEE Trans. Image Process.4
2021 Exploring the Effects of Blur and Deblurring to Visual Object Tracking
abstract
The existence of motion blur can inevitably influence the performance of visual object tracking. However, in contrast to the rapid development of visual trackers, the quantitative effects of increasing levels of motion blur on the performance of visual trackers still remain unstudied. Meanwhile, although image-deblurring can produce visually sharp videos for pleasant visual perception, it is also unknown whether visual object tracking can benefit from image deblurring or not. In this paper, we present a Blurred Video Tracking (BVT) benchmark to address these two problems, which contains a large variety of videos with different levels of motion blurs, as well as ground-truth tracking results. To explore the effects of blur and deblurring to visual object tracking, we extensively evaluate 25 trackers on the proposed BVT benchmark and obtain several new interesting findings. Specifically, we find that light motion blur may improve the accuracy of many trackers, but heavy blur usually hurts the tracking performance. We also observe that image deblurring is helpful to improve tracking accuracy on heavily-blurred videos but hurts the performance of lightly-blurred videos. According to these observations, we then propose a new general GAN-based scheme to improve a tracker's robustness to motion blur. In this scheme, a fine-tuned discriminator can effectively serve as an adaptive blur assessor to enable selective frames deblurring during the tracking process. We use this scheme to successfully improve the accuracy of 6 state-of-the-art trackers on motion-blurred videos.
Qing Guo 0005, Wei Feng 0005, Ruijun Gao, Yang Liu 0003, Song Wang 0002
IEEE Trans. Image Process.5
2021 Shadow Removal by a Lightness-Guided Network With Training on Unpaired Data
abstract
Shadow removal can significantly improve the image visual quality and has many applications in computer vision. Deep learning methods based on CNNs have become the most effective approach for shadow removal by training on either paired data, where both the shadow and underlying shadow-free versions of an image are known, or unpaired data, where shadow and shadow-free training images are totally different with no correspondence. In practice, CNN training on unpaired data is more preferred given the easiness of training data collection. In this paper, we present a new Lightness-Guided Shadow Removal Network (LG-ShadowNet) for shadow removal by training on unpaired data. In this method, we first train a CNN module to compensate for the lightness and then train a second CNN module with the guidance of lightness information from the first CNN module for final shadow removal. We also introduce a loss function to further utilise the colour prior of existing data. Extensive experiments on widely used ISTD, adjusted ISTD and USR datasets demonstrate that the proposed method outperforms the state-of-the-art methods with training on unpaired data.
Hui Yin 0002, Yang Mi, Mengyang Pu, Song Wang 0002
IEEE Trans. Image Process.5
2021 Contour Transformer Network for One-Shot Segmentation of Anatomical Structures
abstract
Accurate segmentation of anatomical structures is vital for medical image analysis. The state-of-the-art accuracy is typically achieved by supervised learning methods, where gathering the requisite expert-labeled image annotations in a scalable manner remains a main obstacle. Therefore, annotation-efficient methods that permit to produce accurate anatomical structure segmentation are highly desirable. In this work, we present Contour Transformer Network (CTN), a one-shot anatomy segmentation method with a naturally built-in human-in-the-loop mechanism. We formulate anatomy segmentation as a contour evolution process and model the evolution behavior by graph convolutional networks (GCNs). Training the CTN model requires only one labeled image exemplar and leverages additional unlabeled data through newly introduced loss functions that measure the global shape and appearance consistency of contours. On segmentation tasks of four different anatomies, we demonstrate that our one-shot learning method significantly outperforms non-learning-based methods and performs competitively to the state-of-the-art fully supervised deep learning methods. With minimal human-in-the-loop editing feedback, the segmentation performance can be further improved to surpass the fully supervised methods.
Weijian Li 0001, Yirui Wang 0002, Adam P. Harrison, Chihung Lin, Song Wang 0002, Jing Xiao 0006, Le Lu 0001, Chang-Fu Kuo, Shun Miao
IEEE Trans. Medical Imaging7
2020 Complementary-View Multiple Human Tracking
abstract
The global trajectories of targets on ground can be well captured from a top view in a high altitude, e.g., by a drone-mounted camera, while their local detailed appearances can be better recorded from horizontal views, e.g., by a helmet camera worn by a person. This paper studies a new problem of multiple human tracking from a pair of top- and horizontal-view videos taken at the same time. Our goal is to track the humans in both views and identify the same person across the two complementary views frame by frame, which is very challenging due to very large field of view difference. In this paper, we model the data similarity in each view using appearance and motion reasoning and across views using appearance and spatial reasoning. Combing them, we formulate the proposed multiple human tracking as a joint optimization problem, which can be solved by constrained integer programming. We collect a new dataset consisting of top- and horizontal-view video pairs for performance evaluation and the experimental results show the effectiveness of the proposed method.
Rui-Ze Han, Wei Feng 0005, Jiewen Zhao, Zicheng Niu, Yujun Zhang 0002, Song Wang 0002
AAAI7
2020 Multi-Spectral Salient Object Detection by Adversarial Domain Adaptation
abstract
Although there are many existing research works about the salient object detection (SOD) in RGB images, there are still many complex situations that regular RGB images cannot provide enough cues for the accurate SOD, such as the shadow effect, similar appearance between background and foreground, strong or insufficient illumination, etc. Because of the success of near-infrared spectrum in many computer vision tasks, we explore the multi-spectral SOD in the synchronized RGB images and near-infrared (NIR) images for the both simple and complex situations. We assume that the RGB SOD in the existing RGB image datasets could provide references for the multi-spectral SOD problem. In this paper, we first collect and will publicize a large multi-spectral dataset including 780 synchronized RGB and NIR image pairs for the multi-spectral SOD problem in the simple and complex situations. We model this research problem as an adversarial domain adaptation from the existing RGB image dataset (source domain) to the collected multi-spectral dataset (target domain). Experimental results show the effectiveness and accuracy of the proposed adversarial domain adaptation for the multi-spectral SOD.
Shaoyue Song, Hongkai Yu, Zhenjiang Miao, Jianwu Fang, Cong Ma 0004, Song Wang 0002
AAAI7
2020 SalSAC: A Video Saliency Prediction Model with Shuffled Attentions and Correlation-Based ConvLSTM
abstract
The performance of predicting human fixations in videos has been much enhanced with the help of development of the convolutional neural networks (CNN). In this paper, we propose a novel end-to-end neural network “SalSAC” for video saliency prediction, which uses the CNN-LSTM-Attention as the basic architecture and utilizes the information from both static and dynamic aspects. To better represent the static information of each frame, we first extract multi-level features of same size from different layers of the encoder CNN and calculate the corresponding multi-level attentions, then we randomly shuffle these attention maps among levels and multiply them to the extracted multi-level features respectively. Through this way, we leverage the attention consistency across different layers to improve the robustness of the network. On the dynamic aspect, we propose a correlation-based ConvLSTM to appropriately balance the influence of the current and preceding frames to the prediction. Experimental results on the DHF1K, Hollywood2 and UCF-sports datasets show that SalSAC outperforms many existing state-of-the-art methods.
Xinyi Wu 0002, Zhenyao Wu, Lili Ju, Song Wang 0002
AAAI5
2020 Multi-Type Self-Attention Guided Degraded Saliency Detection
abstract
Existing saliency detection techniques are sensitive to image quality and perform poorly on degraded images. In this paper, we systematically analyze the current status of the research on detecting salient objects from degraded images and then propose a new multi-type self-attention network, namely MSANet, for degraded saliency detection. The main contributions include: 1) Applying attention transfer learning to promote semantic detail perception and internal feature mining of the target network on degraded images; 2) Developing a multi-type self-attention mechanism to achieve the weight recalculation of multi-scale features. By computing global and local attention scores, we obtain the weighted features of different scales, effectively suppress the interference of noise and redundant information, and achieve a more complete boundary extraction. The proposed MSANet converts low-quality inputs to high-quality saliency maps directly in an end-to-end fashion. Experiments on seven widely-used datasets show that our approach produces good performance on both clear and degraded images.
Ziqi Zhou 0002, Zheng Wang 0008, Huchuan Lu, Song Wang 0002, Meijun Sun
AAAI4
2020 TTPLA: An Aerial-Image Dataset for Detection and Segmentation of Transmission Towers and Power Lines
Rabab Abdelfattah, Xiaofeng Wang 0007, Song Wang 0002
ACCV (6)3
2020 A Multi-Task Mean Teacher for Semi-Supervised Shadow Detection
abstract
Existing shadow detection methods suffer from an intrinsic limitation in relying on limited labeled datasets, and they may produce poor results in some complicated situations. To boost the shadow detection performance, this paper presents a multi-task mean teacher model for semi-supervised shadow detection by leveraging unlabeled data and exploring the learning of multiple information of shadows simultaneously. To be specific, we first build a multi-task baseline model to simultaneously detect shadow regions, shadow edges, and shadow count by leveraging their complementary information and assign this baseline model to the student and teacher network. After that, we encourage the predictions of the three tasks from the student and teacher networks to be consistent for computing a consistency loss on unlabeled data, which is then added to the supervised loss on the labeled data from the predictions of the multi-task baseline model. Experimental results on three widely-used benchmark datasets show that our method consistently outperforms all the compared state-of- the-art methods, which verifies that the proposed network can effectively leverage additional unlabeled data to boost the shadow detection performance.
Zhihao Chen 0004, Lei Zhu 0003, Song Wang 0002, Wei Feng 0005, Pheng-Ann Heng
CVPR4
2020 Modeling Cross-View Interaction Consistency for Paired Egocentric Interaction Recognition
abstract
With the development of Augmented Reality (AR), egocentric action recognition (EAR) plays an important role in accurately understanding demands from the user. However, EAR is designed to help recognize human-machine interaction in single egocentric view, thus difficult to capture interactions between two face-to-face AR users. Paired egocentric interaction recognition (PEIR) is the task to collaboratively recognize the interactions between two persons with the videos in their corresponding views. Unfortunately, existing PEIR methods always directly use linear decision function to fuse the features extracted from two corresponding egocentric videos, which ignore the consistency of interaction in paired egocentric videos. The consistency of interactions in paired videos, and features extracted from them, are correlated to each other. On top of that, we propose to derive the relevance between two views using bilinear pooling, which captures the consistency of two views in feature-level. Specifically, each neuron in the feature maps from one view connects to the neurons from the other view, which enforces the compact consistency between two views and then all possible paired neurons are used for PEIR. To be efficient, we use compact bilinear pooling with Count Sketch to avoid directly computing outer product. Experimental results on the PEV dataset shows the superiority of the proposed methods on the task PEIR.
Zhongguo Li, Fan Lyu, Wei Feng 0005, Song Wang 0002
ICME4
2020 Mutatt: Visual-Textual Mutual Guidance For Referring Expression Comprehension
abstract
Referring expression comprehension (REC) aims to localize a text-related region in a given image by a referring expression in natural language. Existing methods focus on how to build convincing visual and language representations independently, which may significantly isolate visual and language information. In this paper, we argue that for REC the referring expression and the target region are semantically correlated and subject, location and relationship consistency exist between vision and language. On top of this, we propose a novel approach called MutAtt to construct mutual guidance between vision and language, which treats vision and language equally thus yields compact information matching. Specifically, for each module of subject, location and relationship, MutAtt builds two kinds of attention-based mutual guidance strategies. One strategy is to generate vision-guided language embedding for the sake of matching relevant visual features. The other reversely generates language-guided visual features to match relevant language embedding. This mutual guidance strategy can effectively enforce the vision-language consistency in three modules. Experiments on three popular REC datasets demonstrate that the proposed approach outperforms the current state-of-the-art methods.
Fan Lyu, Wei Feng 0005, Song Wang 0002
ICME4
2020 Learning to Segment Anatomical Structures Accurately from One Exemplar
Weijian Li 0001, Yirui Wang 0002, Adam P. Harrison, Chihung Lin, Song Wang 0002, Jing Xiao 0006, Le Lu 0001, Chang-Fu Kuo, Shun Miao
MICCAI (1)7
2020 Complementary-View Co-Interest Person Detection
abstract
Fast and accurate identification of the co-interest persons, who draw joint interest of the surrounding people, plays an important role in social scene understanding and surveillance. Previous study mainly focuses on detecting co-interest persons from a single-view video. In this paper, we study a much more realistic and challenging problem, namely co-interest person~(CIP) detection from multiple temporally-synchronized videos taken by the complementary and time-varying views. Specifically, we use a top-view camera, mounted on a flying drone at a high altitude to obtain a global view of the whole scene and all subjects on the ground, and multiple horizontal-view cameras, worn by selected subjects, to obtain a local view of their nearby persons and environment details. We present an efficient top- and horizontal-view data fusion strategy to map multiple horizontal views into the global top view. We then propose a spatial-temporal CIP potential energy function that jointly considers both intra-frame confidence and inter-frame consistency, thus leading to an effective Conditional Random Field~(CRF) formulation. We also construct a complementary-view video dataset, which provides a benchmark for the study of multi-view co-interest person detection. Extensive experiments validate the effectiveness and superiority of the proposed method.
Rui-Ze Han, Jiewen Zhao, Wei Feng 0005, Yiyang Gan, Song Wang 0002
ACM Multimedia6
2020 Human Identification and Interaction Detection in Cross-View Multi-Person Videos with Wearable Cameras
abstract
Compared to a single fixed camera, multiple moving cameras, e.g., those worn by people, can better capture the human interactive and group activities in a scene, by providing multiple, flexible and possibly complementary views of the involved people. In this setting the actual promotion of activity detection is highly dependent on the effective correlation and collaborative analysis of multiple videos taken by different wearable cameras, which is highly challenging given the time-varying view differences across different cameras and mutual occlusion of people in each video. By focusing on two wearable cameras and the interactive activities that involve only two people, in this paper we develop a new approach that can simultaneously: (i) identify the same persons across the two videos, (ii) detect the interactive activities of interest, including their occurrence intervals and involved people, and (iii) recognize the category of each interactive activity. Specifically, we represent each video by a graph, with detected persons as nodes, and propose a unified Graph Neural Network (GNN) based framework to jointly solve the above three problems. A graph matching network is developed for identifying the same persons across the two videos and a graph inference network is then used for detecting the human interactions. We also build a new video dataset, which provides a benchmark for this study, and conduct extensive experiments to validate the effectiveness and superiority of the proposed method.
Jiewen Zhao, Rui-Ze Han, Yiyang Gan, Wei Feng 0005, Song Wang 0002
ACM Multimedia6
2020 vtGraphNet: Learning weakly-supervised scene graph for complex visual grounding
Fan Lyu, Wei Feng 0005, Song Wang 0002
Neurocomputing3
2020 Weakly supervised easy-to-hard learning for object detection in image sequences
Hongkai Yu, Dazhou Guo, Zhipeng Yan, Lan Fu, Jeff P. Simmons, Craig Przybyla, Song Wang 0002
Neurocomputing7
2020 A Hybrid convolutional neural network for sketch recognition
Xingyuan Zhang, Qi Zou 0001, Yanting Pei, Runsheng Zhang, Song Wang 0002
Pattern Recognit. Lett.6
2020 Challenge-Response Authentication Using In-Air Handwriting Style Verification
abstract
Challenge-response (CR) is an effective way to authenticate users even if the communication channel is insecure. Traditionally CR authentication relies on one-way hashes and shared secrets to verify the identities of users. Such a method cannot cope with an insider attack, where a user can obtained the secret (i.e., the response) from a legitimate user. To cope with it, we design a biometric-based CR authentication scheme (hereafter MoCRA), which is derived from the motions as a user operates emerging depth-sensorbased input devices, such as a Leap Motion controller. We envision that to authenticate a user, MoCRA randomly chooses a string (e.g., a few words), and the user has to write the string in the air. Using Leap Motion, MoCRA captures the user's writing movements and then extracts his / her handwriting style. After verifying that what the user writes matches what is asked for, MoCRA leverages a Support Vecter Machine (SVM) with co-occurrence matrices to model the handwriting styles and can reliably authenticate users, even if what they write is completely different every time. Evaluated on data from 24 subjects over 7 months, MoCRA managed to verify a user with an average of 1.18% (Equal Error Rate) EER and to reject impostors with 2.45% EER.
Wenyuan Xu 0001, Yu Cao 0003, Song Wang 0002
IEEE Trans. Dependable Secur. Comput.4
2020 Deep Domain Adaptation With Differential Privacy
abstract
Nowadays, it usually requires a massive amount of labeled data to train a deep neural network. When no labeled data is available in some application scenarios, domain adaption can be employed to transfer a learner from one or more source domains with labeled data to a target domain with unlabeled data. However, due to the exposure of the trained model to the target domain, the user privacy may potentially be compromised. Nevertheless, the private information may be encoded into the representations in different stages of the deep neural networks, i.e., hierarchical convolutional feature maps, which poses a great challenge for a full-fledged privacy protection. In this paper, we propose a novel differentially private domain adaptation framework called DPDA to achieve domain adaptation with privacy assurance. Specifically, we perform domain adaptation in an adversarial-learning manner and embed the differentially private design into specific layers and learning processes. Although applying differential privacy techniques directly will undermine the performance of deep neural networks, DPDA can increase the classification accuracy for the unlabeled target data compared to the prior arts. We conduct extensive experiments on standard benchmark datasets, and the results show that our proposed DPDA can indeed achieve high accuracy in many domain adaptation tasks with only a modest privacy loss.
Qian Wang 0002, Qin Zou 0001, Lingchen Zhao, Song Wang 0002
IEEE Trans. Inf. Forensics Secur.5
2020 Degraded Image Semantic Segmentation With Dense-Gram Networks
abstract
Degraded image semantic segmentation is of great importance in autonomous driving, highway navigation systems, and many other safety-related applications and it was not systematically studied before. In general, image degradations increase the difficulty of semantic segmentation, usually leading to decreased semantic segmentation accuracy. Therefore, performance on the underlying clean images can be treated as an upper bound of degraded image semantic segmentation. While the use of supervised deep learning has substantially improved the state of the art of semantic image segmentation, the gap between the feature distribution learned using the clean images and the feature distribution learned using the degraded images poses a major obstacle in improving the degraded image semantic segmentation performance. The conventional strategies for reducing the gap include: 1) Adding image-restoration based pre-processing modules; 2) Using both clean and the degraded images for training; 3) Fine-tuning the network pre-trained on the clean image. In this paper, we propose a novel Dense-Gram Network to more effectively reduce the gap than the conventional strategies and segment degraded images. Extensive experiments demonstrate that the proposed Dense-Gram Network yields stateof-the-art semantic segmentation performance on degraded images synthesized using PASCAL VOC 2012, SUNRGBD, CamVid, and CityScapes datasets.
Dazhou Guo, Yanting Pei, Hongkai Yu, Song Wang 0002
IEEE Trans. Image Process.6
2020 Fast Learning of Spatially Regularized and Content Aware Correlation Filter for Visual Tracking
abstract
With a good balance between accuracy and speed, correlation filter (CF) has become a popular and dominant visual object tracking scheme. It implicitly extends the training samples by circular shifts of a given target patch, which serve as negative samples for fast online learning of the filters. Since all these shifted patches are not real negative samples of the target, CF tracking scheme suffers from the annoying boundary effects that can greatly harm the tracking performance, especially under challenging situations, like occlusion and fast temporal variation. Spatial regularization is known as a potent way to alleviate such boundary effects, but with the cost of highly increased time complexity, caused by complex optimization imported by spatial regularization. In this paper, we propose a new fast learning approach to content-aware spatial regularization, namely weighted sample based CF tracking (WSCF). In WSCF, specifically, we present a simple yet effective energy function that implicitly weighs different training samples by spatial deviations. With the energy function, the learning of correlation filters is composed of two subproblems with closed-form solution and can be efficiently solved in an alternate way. We further develop a content-aware updating strategy to dynamically refine the weight distribution to well adapt to the temporal variations of the target and background. Finally, the proposed WSCF is used to enhance two state-of-the-art CF trackers to significantly boost their tracking accuracy, with little sacrifice on the tracking speed. Extensive experiments on five benchmarks validate the effectiveness of the proposed approach.
Rui-Ze Han, Wei Feng 0005, Song Wang 0002
IEEE Trans. Image Process.3
2020 Dual-Branch Network With a Subtle Motion Detector for Microaction Recognition in Videos
abstract
By involving only subtle motions of body parts, video-based microaction recognition is a very important but challenging problem. Most existing action recognition methods are developed for general actions, and the current state-of-the-art methods usually largely rely on high-layer features learned from convolutional neural networks (CNNs). High-layer CNN features usually contain more semantic information but less detailed information. However, detailed information can be important for microactions due to the motion subtleness of such actions. In this paper, we propose to more effectively learn midlayer CNN features for enhancing microaction recognition. More specifically, we develop a new dual-branch network for microaction recognition: one branch uses the high-layer CNN features for classification, and the second branch further explores the midlayer CNN features for classification. In the second branch, we introduce a novel subtle motion detector consisting of three modules: 1) a discriminative spatial-temporal feature learning module, which further learns the subtle motion features corresponding to the discriminative spatial-temporal regions, 2) a parallel multiplier attention module, which further refines the features learned in channels and spatial-temporal domains, and 3) an activation fusion module, which fuses the max and average activations from midlayer CNN features for classification. In the experiments, we build a new microaction video dataset, where the micromotions of interest are mixed with other larger general motions such as walking. Comprehensive experimental results verify that the proposed method yields new state-of-the-art performance in two microaction video datasets, while its performance on two generalaction video datasets is also very promising.
Yang Mi, Xingyuan Zhang, Zhongguo Li, Song Wang 0002
IEEE Trans. Image Process.4
2020 A New Method and Benchmark for Detecting Co-Saliency Within a Single Image
abstract
Recently, saliency detection in a single image and co-saliency detection in multiple images have drawn extensive research interest in the vision and multimedia communities. In this paper, we investigate a new problem of co-saliency detection within a single image, i.e., detecting within-image co-saliency. By identifying common saliency within an image, e.g., highlighting multiple occurrences of an object class with similar appearance, this work can benefit many important applications, such as the detection of objects of interest, more robust object recognition, reduction of information redundancy, and animation synthesis. We propose a new bottom-up method to address this problem. Specifically, a large number of object proposals are first detected from the image. Then we develop an optimization algorithm to derive a set of proposal groups, each of which contains multiple proposals showing good common saliency in the image. For each proposal group, we calculate a co-saliency map and then use a low-rank based algorithm to fuse the maps calculated from all the proposal groups for the final co-saliency map in the image. In the experiment, we collect a new benchmark dataset of 664 color images (two subsets) for within-image co-saliency detection. Experiment results show that the proposed method can better detect the within-image co-saliency than existing algorithms. The experimental results also show that the proposed method can be applied to detect the repetitive patterns in a single image and detect the co-saliency in multiple images.
Hongkai Yu, Jianwu Fang, Hao Guo 0002, Song Wang 0002
IEEE Trans. Multim.5
2020 Improved Deep Hashing With Soft Pairwise Similarity for Multi-Label Image Retrieval
abstract
Hash coding has been widely used in the approximate nearest neighbor search for large-scale image retrieval. Recently, many deep hashing methods have been proposed and shown largely improved performance over traditional feature-learning methods. Most of these methods examine the pairwise similarity on the semantic-level labels, where the pairwise similarity is generally defined in a hard-assignment way. That is, the pairwise similarity is “1” if they share no less than one class label and “0” if they do not share any. However, such similarity definition cannot reflect the similarity ranking for pairwise images that hold multiple labels. In this paper, an improved deep hashing method is proposed to enhance the ability of multi-label image retrieval. We introduce a pairwise quantified similarity calculated on the normalized semantic labels. Based on this, we divide the pairwise similarity into two situations-“hard similarity” and “soft similarity,” where cross-entropy loss and mean square error loss are adapted respectively for more robust feature learning and hash coding. Experiments on four popular datasets demonstrate that the proposed method outperforms the competing methods and achieves the state-of-the-art performance in multi-label image retrieval.
Zheng Zhang 0036, Qin Zou 0001, Yuewei Lin, Long Chen 0005, Song Wang 0002
IEEE Trans. Multim.5
2019 Visual Attention Consistency Under Image Transforms for Multi-Label Image Classification
abstract
Human visual perception shows good consistency for many multi-label image classification tasks under certain spatial transforms, such as scaling, rotation, flipping and translation. This has motivated the data augmentation strategy widely used in CNN classifier training -- transformed images are included for training by assuming the same class labels as their original images. In this paper, we further propose the assumption of perceptual consistency of visual attention regions for classification under such transforms, i.e., the attention region for a classification follows the same transform if the input image is spatially transformed. While the attention regions of CNN classifiers can be derived as an attention heatmap in middle layers of the network, we find that their consistency under many transforms are not preserved. To address this problem, we propose a two-branch network with an original image and its transformed image as inputs and introduce a new attention consistency loss that measures the attention heatmap consistency between two branches. This new loss is then combined with multi-label image classification loss for network training. Experiments on three datasets verify the superiority of the proposed network by achieving new state-of-the-art classification performance.
Hao Guo 0002, Xiaochuan Fan, Hongkai Yu, Song Wang 0002
CVPR5
2019 Semantic Stereo Matching With Pyramid Cost Volumes
abstract
The accuracy of stereo matching has been greatly improved by using deep learning with convolutional neural networks. To further capture the details of disparity maps, in this paper, we propose a novel semantic stereo network named SSPCV-Net, which includes newly designed pyramid cost volumes for describing semantic and spatial information on multiple levels. The semantic features are inferred by a semantic segmentation subnetwork while the spatial features are derived by hierarchical spatial pooling. In the end, we design a 3D multi-cost aggregation module to integrate the extracted multilevel features and perform regression for accurate disparity maps. We conduct comprehensive experiments and comparisons with some recent stereo matching networks on Scene Flow, KITTI 2015 and 2012, and Cityscapes benchmark datasets, and the results show that the proposed SSPCV-Net significantly promotes the state-of-the-art stereo-matching performance.
Zhenyao Wu, Xinyi Wu 0002, Xiaoping Zhang 0005, Song Wang 0002, Lili Ju
ICCV4
2019 Spatial Correspondence With Generative Adversarial Network: Learning Depth From Monocular Videos
abstract
Depth estimation from monocular videos has important applications in many areas such as autonomous driving and robot navigation. It is a very challenging problem without knowing the camera pose since errors in camera-pose estimation can significantly affect the video-based depth estimation accuracy. In this paper, we present a novel SC-GAN network with end-to-end adversarial training for depth estimation from monocular videos without estimating the camera pose and pose change over time. To exploit cross-frame relations, SC-GAN includes a spatial correspondence module which uses Smolyak sparse grids to efficiently match the features across adjacent frames, and an attention mechanism to learn the importance of features in different directions. Furthermore, the generator in SC-GAN learns to estimate depth from the input frames, while the discriminator learns to distinguish between the ground-truth and estimated depth map for the reference frame. Experiments on the KITTI and Cityscapes datasets show that the proposed SC-GAN can achieve much more accurate depth maps than many existing state-of-the-art methods on monocular videos.
Zhenyao Wu, Xinyi Wu 0002, Xiaoping Zhang 0005, Song Wang 0002, Lili Ju
ICCV4
2019 Recognizing Micro Actions in Videos: Learning Motion Details via Segment-Level Temporal Pyramid
abstract
Recognizing micro actions from videos is a very challenging problem since they involve only subtle motions of body parts. In this paper, we propose a new deep-learning based method for micro action recognition by building a segment-level temporal pyramid to better capture the motion details. More specifically, we first temporally sample the input video for short segments and for each of the video segment, we employ a two-stream convolutional neural networks (CNNs) followed by a temporal pyramid for extracting deep features. Finally, the features derived from all the video segments are combined for action classification. We evaluate the proposed method on a micro-action video dataset, as well as a general-action video dataset, with very promising results.
Yang Mi, Song Wang 0002
ICME2
2019 An easy-to-hard learning strategy for within-image co-saliency detection
Shaoyue Song, Hongkai Yu, Zhenjiang Miao, Dazhou Guo, Wei Ke 0001, Cong Ma 0004, Song Wang 0002
Neurocomputing7
2019 Domain Adaptation for Convolutional Neural Networks-Based Remote Sensing Scene Classification
abstract
Remote sensing (RS) scene classification plays an important role in the field of earth observation. With the rapid development of the RS techniques, a large number of RS scene images are available. As manually labeling large-scale RS scene images is both labor and time consuming, when a new unlabeled data set is obtained, how to use the existing labeled data sets to classify the new unlabeled images is an important research direction. Different RS scene image data sets may be taken from different type of sensors, and the images may vary from imaging modalities, spatial resolutions, and image scales, so the distribution discrepancy exists among different image data sets. As a result, simply applying convolutional neural networks (CNN) trained on source domain cannot accurately classify the images on target domain. Domain adaptation (DA) can be helpful to solve this problem. In this letter, we design a subspace alignment (SA) and CNN-based framework to solve the DA problem in RS scene image classification. A new SA layer is proposed and added into CNN models for DA, which could align the source and target domains in some feature subspace. Fine-tuning the modified CNN model with the added SA layer makes the CNN model adapt to the aligned feature subspace and helps to relieve the domain distribution discrepancy. The experiments conducted on two public data sets show that adding the SA layer into CNN improves the scene classification on the target domain.
Shaoyue Song, Hongkai Yu, Zhenjiang Miao, Qiang Zhang 0030, Yuewei Lin, Song Wang 0002
IEEE Geosci. Remote. Sens. Lett.6
2019 Dynamic Saliency-Aware Regularization for Correlation Filter-Based Object Tracking
abstract
With a good balance between tracking accuracy and speed, correlation filter (CF) has become one of the best object tracking frameworks, based on which many successful trackers have been developed. Recently, spatially regularized CF tracking (SRDCF) has been developed to remedy the annoying boundary effects of CF tracking, thus further boosting the tracking performance. However, SRDCF uses a fixed spatial regularization map constructed from a loose bounding box and its performance inevitably degrades when the target or background show significant variations, such as object deformation or occlusion. To address this problem, we propose a new dynamic saliency-aware regularized CF tracking (DSAR-CF) scheme. In DSAR-CF, a simple yet effective energy function, which reflects the object saliency and tracking reliability in the spatial-temporal domain, is defined to guide the online updating of the regularization weight map using an efficient level-set algorithm. Extensive experiments validate that the proposed DSAR-CF leads to better performance in terms of accuracy and speed than the original SRDCF.
Wei Feng 0005, Rui-Ze Han, Qing Guo 0005, Jianke Zhu, Song Wang 0002
IEEE Trans. Image Process.5
2019 Small Object Sensitive Segmentation of Urban Street Scene With Spatial Adjacency Between Object Classes
abstract
Recent advancements in deep learning have shown exciting promise in the urban street scene segmentation. However, many objects, such as poles and sign symbols, are relatively small and they usually cannot be accurately segmented since the larger objects usually contribute more to the segmentation loss. In this paper, we propose a new boundary-based metric that measures the level of spatial adjacency between each pair of object classes and find that this metric is robust against object size induced biases. We develop a new method to enforce this metric into the segmentation loss. We propose a network, which starts with a segmentation network, followed by a new encoder to compute the proposed boundary-based metric, and then trains this network in an end-to-end fashion. In deployment, we only use the trained segmentation network, without the encoder, to segment new unseen images. Experimentally, we evaluate the proposed method using CamVid and CityScapes datasets and achieve a favorable overall performance improvement and a substantial improvement in segmenting small objects.
Dazhou Guo, Ligeng Zhu, Hongkai Yu, Song Wang 0002
IEEE Trans. Image Process.5
2019 Cross-View Person Identification Based on Confidence-Weighted Human Pose Matching
abstract
Cross-view person identification (CVPI) from multiple temporally synchronized videos taken by multiple wearable cameras from different, varying views is a very challenging but important problem, which has attracted more interest recently. Current state-of-the-art performance of CVPI is achieved by matching appearance and motion features across videos, while the matching of pose features does not work effectively given the high inaccuracy of the 3D pose estimation on videos/images collected in the wild. To address this problem, we first introduce a new metric of confidence to the estimated location of each human-body joint in 3D human pose estimation. Then, a mapping function, which can be hand-crafted or learned directly from the datasets, is proposed to combine the inaccurately estimated human pose and the inferred confidence metric to accomplish CVPI. Specifically, the joints with higher confidence are weighted more in the pose matching for CVPI. Finally, the estimated pose information is integrated into the appearance and motion features to boost the CVPI performance. In the experiments, we evaluate the proposed method on three wearable-camera video datasets and compare the performance against several other existing CVPI methods. The experimental results show the effectiveness of the proposed confidence metric, and the integration of pose, appearance, and motion produces a new state-of-the-art CVPI performance.
Guoqiang Liang 0001, Xuguang Lan, Xingyu Chen 0001, Song Wang 0002, Nanning Zheng 0001
IEEE Trans. Image Process.5
2019 DeepCrack: Learning Hierarchical Convolutional Features for Crack Detection
abstract
Cracks are typical line structures that are of interest in many computer-vision applications. In practice, many cracks, e.g., pavement cracks, show poor continuity and low contrast, which brings great challenges to image-based crack detection by using low-level features. In this paper, we propose DeepCrack - an end-to-end trainable deep convolutional neural network for automatic crack detection by learning high-level features for crack representation. In this method, multi-scale deep convolutional features learned at hierarchical convolutional stages are fused together to capture the line structures. More detailed representations are made in larger-scale feature maps and more holistic representations are made in smaller-scale feature maps. We build DeepCrack net on the encoder-decoder architecture of SegNet, and pairwisely fuse the convolutional features generated in the encoder network and in the decoder network at the same scale. We train DeepCrack net on one crack dataset and evaluate it on three others. The experimental results demonstrate that DeepCrack achieves F-Measure over 0.87 on the three challenging datasets in average and outperforms the current state-of-the-art methods.
Qin Zou 0001, Zheng Zhang 0036, Qingquan Li 0001, Xianbiao Qi, Qian Wang 0002, Song Wang 0002
IEEE Trans. Image Process.6
2018 Cross-View Person Identification by Matching Human Poses Estimated With Confidence on Each Body Joint
abstract
Cross-view person identification (CVPI) from multiple temporally synchronized videos taken by multiple wearable cameras from different, varying views is a very challenging but important problem, which has attracted more interests recently. Current state-of-the-art performance of CVPI is achieved by matching appearance and motion features across videos, while the matching of pose features does not work effectively given the high inaccuracy of the 3D human pose estimation on videos/images collected in the wild. In this paper, we introduce a new metric of confidence to the 3D human pose estimation and show that the combination of the inaccurately estimated human pose and the inferred confidence metric can be used to boost the CVPI performance---the estimated pose information can be integrated to the appearance and motion features to achieve the new state-of-the-art CVPI performance. More specifically, the estimated confidence metric is measured at each human-body joint and the joints with higher confidence are weighted more in the pose matching for CVPI. In the experiments, we validate the proposed method on three wearable-camera video datasets and compare the performance against several other existing CVPI methods.
Guoqiang Liang 0001, Xuguang Lan, Song Wang 0002, Nanning Zheng 0001
AAAI4
2018 Curve-Structure Segmentation From Depth Maps: A CNN-Based Approach and Its Application to Exploring Cultural Heritage Objects
abstract
Motivated by the important archaeological application of exploring cultural heritage objects, in this paper we study the challenging problem of automatically segmenting curve structures that are very weakly stamped or carved on an object surface in the form of a highly noisy depth map. Different from most classical low-level image segmentation methods that are known to be very sensitive to the noise and occlusions, we propose a new supervised learning algorithm based on Convolutional Neural Network (CNN) to implicitly learn and utilize more curve geometry and pattern information for addressing this challenging problem. More specifically, we first propose a Fully Convolutional Network (FCN) to estimate the skeleton of curve structures and at each skeleton pixel, a scale value is estimated to reflect the local curve width. Then we propose a dense prediction network to refine the estimated curve skeletons. Based on the estimated scale values, we finally develop an adaptive thresholding algorithm to achieve the final segmentation of curve structures. In the experiment, we validate the performance of the proposed method on a dataset of depth images scanned from unearthed pottery shards dating to the Woodland period of Southeastern North America.
Jing Wang 0129, Karen Y. Smith, Colin Wilder, Song Wang 0002
AAAI7
2018 Co-Saliency Detection Within a Single Image
abstract
Recently, saliency detection in a single image and co-saliency detection in multiple images have drawn extensive research interest in the vision community. In this paper, we investigate a new problem of co-saliency detection within a single image, i.e., detecting within-image co-saliency. By identifying common saliency within an image, e.g., highlighting multiple occurrences of an object class with similar appearance, this work can benefit many important applications, such as the detection of objects of interest, more robust object recognition, reduction of information redundancy, and animation synthesis. We propose a new bottom-up method to address this problem. Specifically, a large number of object proposals are first detected from the image. Then we develop an optimization algorithm to derive a set of proposal groups, each of which contains multiple proposals showing good common saliency in the original image. For each proposal group, we calculate a co-saliency map and then use a low-rank based algorithm to fuse the maps calculated from all the proposal groups for the final co-saliency map in the image. In the experiment, we collect a new dataset of 364 color images with within-image cosaliency. Experiment results show that the proposed method can better detect the within-image co-saliency than existing algorithms.
Hongkai Yu, Jianwu Fang, Hao Guo 0002, Wei Feng 0005, Song Wang 0002
AAAI6
2018 Does Haze Removal Help CNN-Based Image Classification?
Yanting Pei, Qi Zou 0001, Song Wang 0002
ECCV (10)5
2018 Recognizing Actions in Wearable-Camera Videos by Training Classifiers on Fixed-Camera Videos
abstract
Recognizing human actions in wearable camera videos, such as videos taken by GoPro or Google Glass, can benefit many multimedia applications. By mixing the complex and non-stop motion of the camera, motion features extracted from videos of the same action may show very large variation and inconsistency. It is very difficult to collect sufficient videos to cover all such variations and use them to train action classifiers with good generalization ability. In this paper, we develop a new approach to train action classifiers on a relatively smaller set of fixed-camera videos with different views, and then apply them to recognize actions in wearable-camera videos. In this approach, we temporally divide the input video into many shorter video segments and transform the motion features to stable ones in each video segment, in terms of a fixed view defined by an anchor frame in the segment. Finally, we use sparse coding to estimate the action likelihood in each segment, followed by combining the likelihoods from all the video segments for action recognition. We conduct experiments by training on a set of fixed-camera videos and testing on a set of wearable-camera videos, with very promising results.
Yang Mi, Song Wang 0002
ICMR3
2018 Dating ancient paintings of Mogao Grottoes using deeply learnt visual codes
Qingquan Li 0001, Qin Zou 0001, De Ma, Qian Wang 0002, Song Wang 0002
Sci. China Inf. Sci.5
2018 Multiple human tracking in wearable camera videos with informationless intervals
Hongkai Yu, Haozhou Yu, Hao Guo 0002, Jeff P. Simmons, Qin Zou 0001, Wei Feng 0005, Song Wang 0002
Pattern Recognit. Lett.7
2018 Robust Gait Recognition by Integrating Inertial and RGBD Sensors
abstract
Gait has been considered as a promising and unique biometric for person identification. Traditionally, gait data are collected using either color sensors, such as a CCD camera, depth sensors, such as a Microsoft Kinect, or inertial sensors, such as an accelerometer. However, a single type of sensors may only capture part of the dynamic gait features and make the gait recognition sensitive to complex covariate conditions, leading to fragile gait-based person identification systems. In this paper, we propose to combine all three types of sensors for gait data collection and gait recognition, which can be used for important identification applications, such as identity recognition to access a restricted building or area. We propose two new algorithms, namely EigenGait and TrajGait, to extract gait features from the inertial data and the RGBD (color and depth) data, respectively. Specifically, EigenGait extracts general gait dynamics from the accelerometer readings in the eigenspace and TrajGait extracts more detailed subdynamics by analyzing 3-D dense trajectories. Finally, both extracted features are fed into a supervised classifier for gait recognition and person identification. Experiments on 50 subjects, with comparisons to several other state-of-the-art gait-recognition approaches, show that the proposed approach can achieve higher recognition accuracy and robustness.
Qin Zou 0001, Lihao Ni, Qian Wang 0002, Qingquan Li 0001, Song Wang 0002
IEEE Trans. Cybern.5
2017 Are You Lying: Validating the Time-Location of Outdoor Images
Xiaopeng Li 0001, Wenyuan Xu 0001, Song Wang 0002, Xianshan Qu
ACNS3
2017 Learning Dynamic Siamese Network for Visual Object Tracking
abstract
How to effectively learn temporal variation of target appearance, to exclude the interference of cluttered background, while maintaining real-time response, is an essential problem of visual object tracking. Recently, Siamese networks have shown great potentials of matching based trackers in achieving balanced accuracy and beyond realtime speed. However, they still have a big gap to classification & updating based trackers in tolerating the temporal changes of objects and imaging conditions. In this paper, we propose dynamic Siamese network, via a fast transformation learning model that enables effective online learning of target appearance variation and background suppression from previous frames. We then present elementwise multi-layer fusion to adaptively integrate the network outputs using multi-level deep features. Unlike state-of-theart trackers, our approach allows the usage of any feasible generally- or particularly-trained features, such as SiamFC and VGG. More importantly, the proposed dynamic Siamese network can be jointly trained as a whole directly on the labeled video sequences, thus can take full advantage of the rich spatial temporal information of moving objects. As a result, our approach achieves state-of-the-art performance on OTB-2013 and VOT-2015 benchmarks, while exhibits superiorly balanced accuracy and real-time response over state-of-the-art competitors.
Qing Guo 0005, Wei Feng 0005, Ce Zhou, Rui Huang 0006, Song Wang 0002
ICCV6
2017 Learning View-Invariant Features for Person Identification in Temporally Synchronized Videos Taken by Wearable Cameras
abstract
In this paper, we study the problem of Cross-View Person Identification (CVPI), which aims at identifying the same person from temporally synchronized videos taken by different wearable cameras. Our basic idea is to utilize the human motion consistency for CVPI, where human motion can be computed by optical flow. However, optical flow is view-variant - the same person's optical flow in different videos can be very different due to view angle change. In this paper, we attempt to utilize 3D human-skeleton sequences to learn a model that can extract view-invariant motion features from optical flows in different views. For this purpose, we use 3D Mocap database to build a synthetic optical flow dataset and train a Triplet Network (TN) consisting of three sub-networks: two for optical flow sequences from different views and one for the underlying 3D Mocap skeleton sequence. Finally, sub-networks for optical flows are used to extract view-invariant features for CVPI. Experimental results show that, using only the motion information, the proposed method can achieve comparable performance with the state-of-the-art methods. Further combination of the proposed method with an appearance-based method achieves new state-of-the-art performance.
Xiaochuan Fan, Yuewei Lin, Hao Guo 0002, Hongkai Yu, Dazhou Guo, Song Wang 0002
ICCV7
2017 Lesion detection using T1-weighted MRI: A new approach based on functional cortical ROIs
abstract
Accurate and precise detection of brain lesions on MR images is important for relating lesion locations to impaired behaviors. In this paper, we propose a method to detect lesion voxels on each functional cortical ROI (Region of Interest) independently using only T1-weighted MR images (T1-MRI). In contrast to existing automatic lesion detection methods, which typically detect lesion voxels on the whole MR image or on gray matter (GM)/white matter (WM), we show that the proposed functional cortical ROI based method can lead to better lesion-detection performance. We evaluate the proposed method using an in-house dataset with 60 chronic stroke patients. Using leave-one-subject-out cross validation, the proposed method can achieve an average Dice coefficient of 0.74 ± 0.11 and outperform three state-of-the-art methods by more than 0.05.
Dazhou Guo, Song Wang 0002
ICIP3
2017 Loosecut: Interactive image segmentation with loosely bounded boxes
abstract
One popular approach to interactively segment an object of interest from an image is to annotate a bounding box that covers the object, followed by a binary labeling. However, the existing algorithms for such interactive image segmentation prefer a bounding box that tightly encloses the object. This increases the annotation burden, and prevents these algorithms from utilizing automatically detected bounding boxes. In this paper, we develop a new LooseCut algorithm that can handle cases where the bounding box only loosely covers the object. We propose a new Markov Random Fields (MRF) model for segmentation with loosely bounded boxes, including an additional energy term to encourage consistent labeling of similar-appearance pixels and a global similarity constraint to better distinguish the foreground and background. This MRF model is then solved by an iterated max-flow algorithm. We evaluate LooseCut in three public image datasets, and show its better performance against several state-of-the-art methods when increasing the bounding-box size.
Hongkai Yu, Youjie Zhou, Hui Qian 0001, Min Xian, Song Wang 0002
ICIP5
2017 Feature sampling strategies for action recognition
abstract
Although dense local spatial-temporal features with bag-of-features representation achieve state-of-the-art performance for action recognition, the huge feature number and feature size prevent current methods from scaling up to real size problems. In this work, we investigate different types of feature sampling strategies for action recognition, namely dense sampling, uniformly random sampling and selective sampling. We propose two effective selective sampling methods using object proposal techniques. Experiments conducted on a large video dataset show that we are able to achieve better average recognition accuracy using 25% less features, through one of the proposed selective sampling methods, and even maintain comparable accuracy while discarding 70% features.
Youjie Zhou, Hongkai Yu, Song Wang 0002
ICIP3
2017 Human attribute recognition by refining attention heat map
Hao Guo 0002, Xiaochuan Fan, Song Wang 0002
Pattern Recognit. Lett.3
2017 Visual-Attention-Based Background Modeling for Detecting Infrequently Moving Objects
abstract
Motion is one of the most important cues to separate foreground objects from the background in a video. Using a stationary camera, it is usually assumed that the background is static, while the foreground objects are moving most of the time. However, in practice, the foreground objects may show infrequent motions, such as abandoned objects and sleeping persons. Meanwhile, the background may contain frequent local motions, such as waving trees and/or grass. Such complexities may prevent the existing background subtraction algorithms from correctly identifying the foreground objects. In this paper, we propose a new approach that can detect the foreground objects with frequent and/or infrequent motions. Specifically, we use a visual-attention mechanism to infer a complete background from a subset of frames and then propagate it to the other frames for accurate background subtraction. Furthermore, we develop a feature-matching-based local motion stabilization algorithm to identify frequent local motions in the background for reducing false positives in the detected foreground. The proposed approach is fully unsupervised, without using any supervised learning for object detection and tracking. Extensive experiments on a large number of videos have demonstrated that the proposed approach outperforms the state-of-the-art motion detection and background subtraction methods in comparison.
Yuewei Lin, Yu Cao 0003, Youjie Zhou, Song Wang 0002
IEEE Trans. Circuits Syst. Video Technol.5
2017 Cross-Domain Recognition by Identifying Joint Subspaces of Source Domain and Target Domain
abstract
This paper introduces a new method to solve the cross-domain recognition problem. Different from the traditional domain adaption methods which rely on a global domain shift for all classes between the source and target domains, the proposed method is more flexible to capture individual class variations across domains. By adopting a natural and widely used assumption that the data samples from the same class should lay on an intrinsic low-dimensional subspace, even if they come from different domains, the proposed method circumvents the limitation of the global domain shift, and solves the cross-domain recognition by finding the joint subspaces of the source and target domains. Specifically, given labeled samples in the source domain, we construct a subspace for each of the classes. Then we construct subspaces in the target domain, called anchor subspaces, by collecting unlabeled samples that are close to each other and are highly likely to belong to the same class. The corresponding class label is then assigned by minimizing a cost function which reflects the overlap and topological structure consistency between subspaces across the source and target domains, and within the anchor subspaces, respectively. We further combine the anchor subspaces to the corresponding source subspaces to construct the joint subspaces. Subsequently, one-versus-rest support vector machine classifiers are trained using the data samples belonging to the same joint subspaces and applied to unlabeled data in the target domain. We evaluate the proposed method on two widely used datasets: 1) object recognition dataset for computer vision tasks and 2) sentiment classification dataset for natural language processing tasks. Comparison results demonstrate that the proposed method outperforms the comparison methods on both datasets.
Yuewei Lin, Jing Chen 0008, Yu Cao 0003, Youjie Zhou, Lingfeng Zhang 0001, Yuan Yan Tang, Song Wang 0002
IEEE Trans. Cybern.7
2017 Local Pattern Collocations Using Regional Co-occurrence Factorization
abstract
Human vision benefits a lot from pattern collocations in visual activities such as object detection and recognition. Usually, pattern collocations display as the co-occurrences of visual primitives, e.g., colors, gradients, or textons, in neighboring regions. In the past two decades, many sophisticated local feature descriptors have been developed to describe visual primitives, and some of them even take into account the co-occurrence information for improving their discriminative power. However, most of these descriptors only consider feature co-occurrence within a very small neighborhood, e.g., 8-connected or 16-connected area, which would fall short in describing pattern collocations built up by feature co-occurrences in a wider neighborhood. In this paper, we propose to describe local pattern collocations by using a new and general regional co-occurrence approach. In this approach, an input image is first partitioned into a set of homogeneous superpixels. Then, features in each superpixel are extracted by a variety of local feature descriptors, based on which a number of patterns are computed. Finally, pattern co-occurrences within the superpixel and between the neighboring superpixels are calculated and factorized into a final descriptor for local pattern collocation. The proposed regional co-occurrence framework is extensively tested on a wide range of popular shape, color, and texture descriptors in terms of image and object categorizations. The experimental results have shown significant performance improvements by using the proposed framework over the existing popular descriptors.
Qin Zou 0001, Lihao Ni, Qian Wang 0002, Zhongwen Hu, Qingquan Li 0001, Song Wang 0002
IEEE Trans. Multim.6
2016 Groupwise Tracking of Crowded Similar-Appearance Targets from Low-Continuity Image Sequences
abstract
Automatic tracking of large-scale crowded targets are of particular importance in many applications, such as crowded people/vehicle tracking in video surveillance, fiber tracking in materials science, and cell tracking in biomedical imaging. This problem becomes very challenging when the targets show similar appearance and the interslice/ inter-frame continuity is low due to sparse sampling, camera motion and target occlusion. The main challenge comes from the step of association which aims at matching the predictions and the observations of the multiple targets. In this paper we propose a new groupwise method to explore the target group information and employ the within-group correlations for association and tracking. In particular, the within-group association is modeled by a nonrigid 2D Thin-Plate transform and a sequence of group shrinking, group growing and group merging operations are then developed to refine the composition of each group. We apply the proposed method to track large-scale fibers from microscopy material images and compare its performance against several other multi-target tracking methods. We also apply the proposed method to track crowded people from videos with poor inter-frame continuity.
Hongkai Yu, Youjie Zhou, Jeff P. Simmons, Craig Przybyla, Yuewei Lin, Xiaochuan Fan, Yang Mi, Song Wang 0002
CVPR8
2016 Large-Scale Fiber Tracking Through Sparsely Sampled Image Sequences of Composite Materials
abstract
Fast and accurate characterization of fiber micro-structures plays a central role for material scientists to analyze physical properties of continuous fiber reinforced composite materials. In materials science, this is usually achieved by continuously cross-sectioning a 3D material sample for a sequence of 2D microscopic images, followed by a fiber detection/tracking algorithm through the obtained image sequence. To speed up this process and be able to handle larger size material samples, this paper proposes sparse sampling with larger inter-slice distance in cross sectioning and develops a new algorithm that can robustly track large-scale fibers from such a sparsely sampled image sequence. In particular, the problem is formulated as multi-target tracking, and the Kalman filters are applied to track each fiber along the image sequence. One main challenge in this tracking process is to correctly associate each fiber to its observation given that: fiber observations are of large scale, crowded, and show very similar appearances in a 2D slice and there may be a large gap between the predicted location of a fiber and its observation in the sparse sampling. To address this challenge, a novel group-wise association algorithm is developed by leveraging the fact that fibers are implanted in bundles and the fibers in the same bundle are highly correlated through the image sequence. In experiments, the proposed algorithm is tested on three tiles of 100-slice S200 material samples and the tracking performance is evaluated using 1136 human annotated ground-truth fiber tracks. Both quantitative and qualitative results show that the proposed algorithm clearly outperforms the state-of-the-art multiple-target tracking algorithms on sparsely sampled image sequences.
Youjie Zhou, Hongkai Yu, Jeff P. Simmons, Craig Przybyla, Song Wang 0002
IEEE Trans. Image Process.5
2015 Combining local appearance and holistic view: Dual-Source Deep Neural Networks for human pose estimation
abstract
We propose a new learning-based method for estimating 2D human pose from a single image, using Dual-Source Deep Convolutional Neural Networks (DS-CNN). Recently, many methods have been developed to estimate human pose by using pose priors that are estimated from physiologically inspired graphical models or learned from a holistic perspective. In this paper, we propose to integrate both the local (body) part appearance and the holistic view of each local part for more accurate human pose estimation. Specifically, the proposed DS-CNN takes a set of image patches (category-independent object proposals for training and multi-scale sliding windows for testing) as the input and then learns the appearance of each local part by considering their holistic views in the full body. Using DS-CNN, we achieve both joint detection, which determines whether an image patch contains a body joint, and joint localization, which finds the exact location of the joint in the image patch. Finally, we develop an algorithm to combine these joint detection/localization results from all the image patches for estimating the human pose. The experimental results show the effectiveness of the proposed method by comparing to the state-of-the-art human-pose estimation methods based on pose priors that are estimated from physiologically inspired graphical models or learned from a holistic perspective.
Xiaochuan Fan, Yuewei Lin, Song Wang 0002
CVPR4
2015 Co-Interest Person Detection from Multiple Wearable Camera Videos
abstract
Wearable cameras, such as Google Glass and Go Pro, enable video data collection over larger areas and from different views. In this paper, we tackle a new problem of locating the co-interest person (CIP), i.e., the one who draws attention from most camera wearers, from temporally synchronized videos taken by multiple wearable cameras. Our basic idea is to exploit the motion patterns of people and use them to correlate the persons across different videos, instead of performing appearance-based matching as in traditional video co-segmentation/localization. This way, we can identify CIP even if a group of people with similar appearance are present in the view. More specifically, we detect a set of persons on each frame as the candidates of the CIP and then build a Conditional Random Field (CRF) model to select the one with consistent motion patterns in different videos and high spacial-temporal consistency in each video. We collect three sets of wearable-camera videos for testing the proposed algorithm. All the involved people have similar appearances in the collected videos and the experiments demonstrate the effectiveness of the proposed algorithm.
Yuewei Lin, Kareem Abdelfatah, Youjie Zhou, Xiaochuan Fan, Hongkai Yu, Hui Qian 0001, Song Wang 0002
ICCV7
2015 Cross-domain recognition by identifying compact joint subspaces
abstract
This paper introduces a new method to solve the cross-domain recognition problem. Different from the traditional domain adaption methods which rely on a global domain shift for all classes between source and target domain, the proposed method is more flexible to capture individual class variations across domains. We propose to solves the problem by finding the compact joint subspaces of source and target domain. We evaluate the proposed method on two widely used datasets and comparison results demonstrates that the proposed method outperforms the comparison methods.
Yuewei Lin, Jing Chen 0008, Yu Cao 0003, Youjie Zhou, Lingfeng Zhang 0001, Song Wang 0002
ICIP6
2015 Discriminative regional color co-occurrence descriptor
abstract
Traditional color feature descriptors are focused on color-value distributions in the color space, e.g., color histograms, color bag-of-words, which ignore the spatial location and contextual information of different colors. In this paper, a new regional color co-occurrence feature descriptor (RCC) is proposed to reflect spatial relations of colors in an image. First, we partition an image into a number of disjoint regions using superpixel techniques. Then, we construct a color histogram for each region, based on which we construct a color co-occurrence matrix for each pair of neighboring regions. Finally, all the constructed co-occurrence matrices from an image are summed up and normalized as a color descriptor to represent this image. This new color descriptor reflects the color-collocation patterns in the image. We use this new color descriptor for image/object classification and find that it leads to higher classification accuracies than other competing color descriptors.
Qin Zou 0001, Xianbiao Qi, Qingquan Li 0001, Song Wang 0002
ICIP4
2015 Simple Atom Selection Strategy for Greedy Matrix Completion
Zebang Shen, Hui Qian 0001, Song Wang 0002
IJCAI4
2015 Distance Transform Based Active Contour Approach for Document Image Rectification
abstract
Digitization of document images using OCR based systems is adversely affected if the image of the document contains distortion (warping). Often, costly and precisely calibrated special hardware such as stereo cameras, laser scanners, etc. are used to infer the 3D model of the distorted image which is used to remove the distortion. Recent methods focus on creating a 3D shape model based on the 2D document image. The performance of these methods is highly dependent on estimating an accurate 2D distortion grid. In the domain of printed document images, the white space between the text lines carries as much information about the 2D distortion as the text lines themselves. Based on this intuitive idea, we build a 2D distortion grid from white space lines, which can be used to rectify a printed document image by a Dewar ping algorithm. These white space lines are extracted using a propagation technique on the distance transform of the binarized document image, guided by an open active contour algorithm. We compare our proposed method against a state-of-the-art 2D distortion grid construction method and obtain better results. We also present qualitative and quantitative evaluations for the proposed method.
Dhaval Salvi, Youjie Zhou, Song Wang 0002
WACV4
2015 Topology-Preserving Multi-label Image Segmentation
abstract
Enforcing a specific topology in image segmentation is a very important but challenging problem, which has attracted much attention in the computer vision community. Most recent works on topology-constrained image segmentation focus on binary segmentation, where the topology is often described by the connectivity of both foreground and background. In this paper, we develop a new multi-labeling method to enforce topology in multi-label image segmentation. In this case, we not only require each segment to be a connected region (intra-segment topology), but also require specific adjacency relations between each pair of segments (inter-segment topology). We develop our method in the context of segmentation propagation, where a segmented template image defines the topology, and our goal is to propagate the segmentation to a target image while preserving the topology. Our method requires good spatial structure continuity between the template and the target such that the template segmentation can be used as a good initialization for segmenting the target. In addition, we focus on multi-label segmentation where a segment and its adjacent segments form a ring structure, which is among the most complex type of inter-segment topology for 2D structures. We apply the proposed method to segment 3D metallic image volumes for the underlying grain structures and achieve better results than several comparison methods. Finally, we also apply the proposed method to interactive segmentation and stereo matching applications.
Jarrell W. Waggoner, Youjie Zhou, Jeff P. Simmons, Marc De Graef, Song Wang 0002
WACV5
2015 Multiscale Superpixels and Supervoxels Based on Hierarchical Edge-Weighted Centroidal Voronoi Tessellation
abstract
Super pixels and super voxels play an important role in many computer vision applications, such as image segmentation, object recognition and video analysis. In this paper, we propose a hierarchical edge-weighted centroidal Voronoi tessellation (HEWCVT) method for generating superpixels/supervoxels in multiple scales. In this method we model the problem as a multilevel clustering process: superpixels/supervoxels in one level are clustered to obtain larger size superpixels/supervoxels in the next level. In the finest scale, the initial clustering is directly conducted on pixels/voxels. The clustering energy involves both color/feature similarities and the proposed boundary smoothness of superpixels/supervoxels. The resulting superpixels/supervoxels can be easily represented by a hierarchical tree which describes the nesting relation of superpixels/supervoxels across different scales. We evaluate and compare the proposed method with several state-of-the-art superpixel/supervoxel methods on standard image and video datasets. Both quantitative and qualitative results show that the proposed HEWCVT method achieves superior or comparable performances to other methods.
Youjie Zhou, Lili Ju, Song Wang 0002
WACV3
2015 Multiscale Superpixels and Supervoxels Based on Hierarchical Edge-Weighted Centroidal Voronoi Tessellation
abstract
Superpixels and supervoxels play an important role in many computer vision applications, such as image segmentation, object recognition, and video analysis. In this paper, we propose a new hierarchical edge-weighted centroidal Voronoi tessellation (HEWCVT) method for generating superpixels/supervoxels in multiple scales. In this method, we model the problem as a multilevel clustering process: superpixels/supervoxels in one level are clustered to obtain larger size superpixels/supervoxels in the next level. In the finest scale, the initial clustering is directly conducted on pixels/voxels. The clustering energy involves both color similarities and boundary smoothness of superpixels/supervoxels. The resulting superpixels/supervoxels can be easily represented by a hierarchical tree which describes the nesting relation of superpixels/supervoxels across different scales. We first investigate the performance of obtained superpixels/supervoxels under different parameter settings, then we evaluate and compare the proposed method with several state-of-the-art superpixel/supervoxel methods on standard image and video data sets. Both quantitative and qualitative results show that the proposed HEWCVT method achieves superior or comparable performances with other methods.
Youjie Zhou, Lili Ju, Song Wang 0002
IEEE Trans. Image Process.3
2014 Pose Locality Constrained Representation for 3D Human Pose Reconstruction
Xiaochuan Fan, Youjie Zhou, Song Wang 0002
ECCV (1)4
2014 Graph-cut based interactive segmentation of 3D materials-science images
Jarrell W. Waggoner, Youjie Zhou, Jeff P. Simmons, Marc De Graef, Song Wang 0002
Mach. Vis. Appl.5
2014 Automatic inpainting by removing fence-like structures in RGBD images
Qin Zou 0001, Yu Cao 0003, Qingquan Li 0001, Qingzhou Mao, Song Wang 0002
Mach. Vis. Appl.5
2014 Chronological classification of ancient paintings using appearance and shape features
Qin Zou 0001, Yu Cao 0003, Qingquan Li 0001, Chuanhe Huang, Song Wang 0002
Pattern Recognit. Lett.5
2013 Recognize Human Activities from Partially Observed Videos
abstract
Recognizing human activities in partially observed videos is a challenging problem and has many practical applications. When the unobserved subsequence is at the end of the video, the problem is reduced to activity prediction from unfinished activity streaming, which has been studied by many researchers. However, in the general case, an unobserved subsequence may occur at any time by yielding a temporal gap in the video. In this paper, we propose a new method that can recognize human activities from partially observed videos in the general case. Specifically, we formulate the problem into a probabilistic framework: 1) dividing each activity into multiple ordered temporal segments, 2) using spatiotemporal features of the training video samples in each segment as bases and applying sparse coding (SC) to derive the activity likelihood of the test video sample at each segment, and 3) finally combining the likelihood at each segment to achieve a global posterior for the activities. We further extend the proposed method to include more bases that correspond to a mixture of segments with different temporal lengths (MSSC), which can better represent the activities with large intra-class variations. We evaluate the proposed methods (SC and MSSC) on various real videos. We also evaluate the proposed methods on two special cases: 1) activity prediction where the unobserved subsequence is at the end of the video, and 2) human activity recognition on fully observed videos. Experimental results show that the proposed methods outperform existing state-of-the-art comparison methods.
Yu Cao 0003, Daniel Paul Barrett, Andrei Barbu, N. Siddharth 0001, Haonan Yu, Aaron Michaux, Yuewei Lin, Sven J. Dickinson, Jeffrey Mark Siskind, Song Wang 0002
CVPR10
2013 KinWrite: Handwriting-Based Authentication Using Kinect
Chengzhang Qu, Wenyuan Xu 0001, Song Wang 0002
NDSS4
2013 A graph-based algorithm for multi-target tracking with occlusion
abstract
Multi-target tracking plays a key role in many computer vision applications including robotics, human-computer interaction, event recognition, etc., and has received increasing attention in past several years. Starting with an object detector is one of many approaches used by existing multi-target tracking methods to create initial short tracks called tracklets. These tracklets are then gradually grouped into longer final tracks in a heirarchical framework. Although object detectors have greatly improved in recent years, these detectors are far from perfect and can fail to detect the object of interest or identify a false positive as the desired object. Due to the presence of false positives or mis-detections from the object detector, these tracking methods can suffer from track fragmentations and identity switches. To address this problem, we formulate multi-target tracking as a min-cost flow graph problem which we call the average shortest path. This average shortest path is designed to be less biased towards the track length. In our average shortest path framework, object misdetection is treated as an occlusion and is represented by the edges between track-let nodes across non consecutive frames. We evaluate our method on the publicly available ETH dataset. Camera motion and long occlusions in a busy street scene make ETH a challenging dataset. We achieve competitive results with lower identity switches on this dataset as compared to the state of the art methods.
Dhaval Salvi, Jarrell W. Waggoner, Andrew Temlyakov, Song Wang 0002
WACV4
2013 Handwritten text segmentation using average longest path algorithm
abstract
Offline handwritten text recognition is a very challenging problem. Aside from the large variation of different handwriting styles, neighboring characters within a word are usually connected, and we may need to segment a word into individual characters for accurate character recognition. Many existing methods achieve text segmentation by evaluating the local stroke geometry and imposing constraints on the size of each resulting character, such as the character width, height and aspect ratio. These constraints are well suited for printed texts, but may not hold for handwritten texts. Other methods apply holistic approach by using a set of lexicons to guide and correct the segmentation and recognition. This approach may fail when the lexicon domain is insufficient. In this paper, we present a new global non-holistic method for handwritten text segmentation, which does not make any limiting assumptions on the character size and the number of characters in a word. Specifically, the proposed method finds the text segmentation with the maximum average likeliness for the resulting characters. For this purpose, we use a graph model that describes the possible locations for segmenting neighboring characters, and we then develop an average longest path algorithm to identify the globally optimal segmentation. We conduct experiments on real images of handwritten texts taken from the IAM handwriting database and compare the performance of the proposed method against an existing text segmentation algorithm that uses dynamic programming.
Dhaval Salvi, Jarrell W. Waggoner, Song Wang 0002
WACV4
2013 Shape and image retrieval by organizing instances using population cues
abstract
Reliably measuring the similarity of two shapes or images (instances) is an important problem for various computer vision applications such as classification, recognition, and retrieval. While pairwise measures take advantage of the geometric differences between two instances to quantify their similarity, recent advances use relationships among the population of instances when quantifying pairwise measures. In this paper, we propose a novel method which refines pairwise similarity measures using population cues by examining the most similar instances shared by the compared shapes or images. We then use this refined measure to organize instances into disjoint components that consist of similar instances. Connectivity is then established between components to avoid hard constraints on what instances can be retrieved, improving retrieval performance. To evaluate the proposed method we conduct experiments on the well-known MPEG-7 and Swedish Leaf shape datasets as well as the Nister and Stewenius image dataset. We show that the proposed method is versatile, performing very well on its own or in concert with existing methods.
Andrew Temlyakov, Pahal Dalal, Jarrell W. Waggoner, Dhaval Salvi, Song Wang 0002
WACV5
2013 Special issue on Shape Modeling in Medical Image Analysis
Wiro J. Niessen, Shuo Li 0001, Song Wang 0002
Comput. Vis. Image Underst.3
2013 A new criterion for choosing planar subproblems in MAP-MRF inference
Jianwu Dong, Feng Chen 0007, Qiang Shawn Cheng, Song Wang 0002
Neurocomputing4
2013 A Visual-Attention Model Using Earth Mover's Distance-Based Saliency Measurement and Nonlinear Feature Combination
abstract
This paper introduces a new computational visual-attention model for static and dynamic saliency maps. First, we use the Earth Mover's Distance (EMD) to measure the center-surround difference in the receptive field, instead of using the Difference-of-Gaussian filter that is widely used in many previous visual-attention models. Second, we propose to take two steps of biologically inspired nonlinear operations for combining different features: combining subsets of basic features into a set of super features using the Lm-norm and then combining the super features using the Winner-Take-All mechanism. Third, we extend the proposed model to construct dynamic saliency maps from videos by using EMD for computing the center-surround difference in the spatiotemporal receptive field. We evaluate the performance of the proposed model on both static image data and video data. Comparison results show that the proposed model outperforms several existing models under a unified evaluation setting.
Yuewei Lin, Yuan Yan Tang, Bin Fang 0001, Zhaowei Shang, Song Wang 0002
IEEE Trans. Pattern Anal. Mach. Intell.6
2013 3D Superalloy Grain Segmentation Using a Multichannel Edge-Weighted Centroidal Voronoi Tessellation Algorithm
abstract
Accurate grain segmentation on 3D superalloy images is very important in materials science and engineering. From grain segmentation, we can derive the underlying superalloy grains' micro-structures, based on how many important physical, mechanical, and chemical properties of the superalloy samples can be evaluated. Grain segmentation is, however, usually a very challenging problem because: 1) even a small 3D superalloy sample may contain hundreds of grains; 2) carbides and noises may degrade the imaging quality; and 3) the intensity within a grain may not be homogeneous. In addition, the same grain may present different appearances, e.g., different intensities, under different microscope settings. In practice, a 3D superalloy image may contain multichannel information where each channel corresponds to a specific microscope setting. In this paper, we develop a multichannel edge-weighted centroidal Voronoi tessellation (MCEWCVT) algorithm to effectively and robustly segment the superalloy grains from 3D multichannel superalloy images. MCEWCVT performs segmentation by minimizing an energy function, which encodes both the multichannel voxel-intensity similarity within each cluster in the intensity domain and the smoothness of segmentation boundaries in the 3D image domain. In the experiment, we first quantitatively evaluate the proposed MCEWCVT algorithm on a four-channel Ni-based 3D superalloy data set (IN100) against the manually annotated ground-truth segmentation. We further evaluate the MCEWCVT algorithm on two synthesized four-channel superalloy data sets. The qualitative and quantitative comparisons of 18 existing image segmentation algorithms demonstrate the effectiveness and robustness of the proposed MCEWCVT algorithm.
Yu Cao 0003, Lili Ju, Youjie Zhou, Song Wang 0002
IEEE Trans. Image Process.4
2013 3D Materials Image Segmentation by 2D Propagation: A Graph-Cut Approach Considering Homomorphism
abstract
Segmentation propagation, similar to tracking, is the problem of transferring a segmentation of an image to a neighboring image in a sequence. This problem is of particular importance to materials science, where the accurate segmentation of a series of 2D serial-sectioned images of multiple, contiguous 3D structures has important applications. Such structures may have distinct shape, appearance, and topology, which can be considered to improve segmentation accuracy. For example, some materials images may have structures with a specific shape or appearance in each serial section slice, which only changes minimally from slice to slice, and some materials may exhibit specific inter-structure topology that constrains their neighboring relations. Some of these properties have been individually incorporated to segment specific materials images in prior work. In this paper, we develop a propagation framework for materials image segmentation where each propagation is formulated as an optimal labeling problem that can be efficiently solved using the graph-cut algorithm. Our framework makes three key contributions: 1) a homomorphic propagation approach, which considers the consistency of region adjacency in the propagation; 2) incorporation of shape and appearance consistency in the propagation; and 3) a local non-homomorphism strategy to handle newly appearing and disappearing substructures during this propagation. To show the effectiveness of our framework, we conduct experiments on various 3D materials images, and compare the performance against several existing image segmentation methods.
Jarrell W. Waggoner, Youjie Zhou, Jeff P. Simmons, Marc De Graef, Song Wang 0002
IEEE Trans. Image Process.5
2012 Superedge grouping for object localization by combining appearance and shape information
abstract
Both appearance and shape play important roles in object localization and object detection. In this paper, we propose a new superedge grouping method for object localization by incorporating both boundary shape and appearance information of objects. Compared with the previous edge grouping methods, the proposed method does not subdivide detected edges into short edgels before grouping. Such long, unsubdivided superedges not only facilitate the incorporation of object shape information into localization, but also increase the robustness against image noise and reduce computation. We identify and address several important problems in achieving the proposed superedge grouping, including gap filling for connecting superedges, accurate encoding of region-based information into individual edges, and the incorporation of object-shape information into object localization. In this paper, we use the bag of visual words technique to quantify the region-based appearance features of the object of interest. We find that the proposed method, by integrating both boundary and region information, can produce better localization performance than previous subwindow search and edge grouping methods on most of the 20 object categories from the VOC 2007 database. Experiments also show that the proposed method is roughly 50 times faster than the previous edge grouping method.
Sanja Fidler, Jarrell W. Waggoner, Yu Cao 0003, Sven J. Dickinson, Jeffrey Mark Siskind, Song Wang 0002
CVPR7
2012 Grain Segmentation of 3D Superalloy Images Using Multichannel EWCVT under Human Annotation Constraints
Yu Cao 0003, Lili Ju, Song Wang 0002
ECCV (3)3
2012 Video In Sentences Out
Andrei Barbu, Alexander Bridge, Zachary Burchill, Dan Coroian, Sven J. Dickinson, Sanja Fidler, Aaron Michaux, Sam Mussman, N. Siddharth 0001, Dhaval Salvi, Lara Schmidt, Jiangnan Shangguan, Jeffrey Mark Siskind, Jarrell W. Waggoner, Song Wang 0002, Jinlian Wei
UAI15
2012 Pre-organizing Shape Instances for Landmark-Based Shape Correspondence
Brent C. Munsell, Andrew Temlyakov, Martin Styner, Song Wang 0002
Int. J. Comput. Vis.4
2012 Recursive sum-product algorithm for generalized outer-planar graphs
Qiang Shawn Cheng, Feng Chen 0007, Wenli Xu, Song Wang 0002
Inf. Process. Lett.4
2012 CrackTree: Automatic crack detection from pavement images
Qin Zou 0001, Yu Cao 0003, Qingquan Li 0001, Qingzhou Mao, Song Wang 0002
Pattern Recognit. Lett.5
2011 A Multichannel Edge-Weighted Centroidal Voronoi Tessellation algorithm for 3D super-alloy image segmentation
abstract
In material science and engineering, the grain structure inside a super-alloy sample determines its mechanical and physical properties. In this paper, we develop a new Multichannel Edge-Weighted Centroidal Voronoi Tessellation (MCEWCVT) algorithm to automatically segment all the 3D grains from microscopic images of a super-alloy sample. Built upon the classical k-means/CVT algorithm, the proposed algorithm considers both the voxel-intensity similarity within each cluster and the compactness of each cluster. In addition, the same slice of a super-alloy sample can produce multiple images with different grain appearances using different settings of the microscope. We call this multichannel imaging and in this paper, we further adapt the proposed segmentation algorithm to handle such multichannel images to achieve higher grain-segmentation accuracy. We test the proposed MCEWCVT algorithm on a 4-channel Ni-based 3D super-alloy image consisting of 170 slices. The segmentation performance is evaluated against the manually annotated ground-truth segmentation and quantitatively compared with other six image segmentation/edge-detection methods. The experimental results demonstrate the higher accuracy of the proposed algorithm than the comparison methods.
Yu Cao 0003, Lili Ju, Qin Zou 0001, Chengzhang Qu, Song Wang 0002
CVPR5
2011 2D nonrigid partial shape matching using MCMC and contour subdivision
abstract
Shape matching has many applications in computer vision, such as shape classification, object recognition, object detection, and localization. In 2D cases, shape instances are 2D closed contours and matching two shape contours can usually be formulated as finding a one-to-one dense point correspondence between them. However, in practice, many shape contours are extracted from real images and may contain partial occlusions. This leads to the challenging partial shape matching problem, where we need to identify and match a subset of segments of the two shape contours. In this paper, we propose a new MCMC (Markov chain Monte Carlo) based algorithm to handle partial shape matching with mildly non-rigid deformations. Specifically, we represent each shape contour by a set of ordered landmark points. The selection of a subset of these landmark points into the shape matching is evaluated and updated by a posterior distribution, which is composed of both a matching likelihood and a prior distribution. This prior distribution favors the inclusion of more and consecutive landmark points into the matching. To better describe the matching likelihood, we develop a contour-subdivision technique to highlight the contour segment with highest matching cost from the selected subsequences of the points. In our experiments, we construct 1,600 test shape instances by introducing partial occlusions to the 40 shapes chosen from different categories in MPEG-7 dataset. We evaluate the performance of the proposed algorithm by comparing with three well-known partial shape matching methods.
Yu Cao 0003, Irina Czogiel, Ian L. Dryden, Song Wang 0002
CVPR5
2011 Object tracking via appearance modeling and sparse representation
Feng Chen 0007, Qing Wang 0017, Song Wang 0002, Wenli Xu
Image Vis. Comput.3
2010 Two perceptually motivated strategies for shape classification
abstract
In this paper, we propose two new, perceptually motivated strategies to better measure the similarity of 2D shape instances that are in the form of closed contours. The first strategy handles shapes that can be decomposed into a base structure and a set of inward or outward pointing “strand” structures, where a strand structure represents a very thin, elongated shape part attached to the base structure. The similarity of two such shape contours can be better described by measuring the similarity of their base structures and strand structures in different ways. The second strategy handles shapes that exhibit good bilateral symmetry. In many cases, such shapes are invariant to a certain level of scaling transformation along their symmetry axis. In our experiments, we show that these two strategies can be integrated into available shape matching methods to improve the performance of shape classification on several widely-used shape data sets.
Andrew Temlyakov, Brent C. Munsell, Jarrell W. Waggoner, Song Wang 0002
CVPR4
2010 Free-shape subwindow search for object localization
abstract
Object localization in an image is usually handled by searching for an optimal subwindow that tightly covers the object of interest. However, the subwindows considered in previous work are limited to rectangles or other specified, simple shapes. With such specified shapes, no subwindow can cover the object of interest tightly. As a result, the desired subwindow around the object of interest may not be optimal in terms of the localization objective function, and cannot be detected by a subwindow search algorithm. In this paper, we propose a new graph-theoretic approach for object localization by searching for an optimal subwindow without pre-specifying its shape. Instead, we require the resulting subwindow to be well aligned with edge pixels that are detected from the image. This requirement is quantified and integrated into the localization objective function based on the widely-used bag of visual words technique. We show that the ratio-contour graph algorithm can be adapted to find the optimal free-shape subwindow in terms of the new localization objective function. In the experiment, we test the proposed approach on the PASCAL VOC 2006 and VOC 2007 databases for localizing several categories of animals. We find that its performance is better than the previous efficient subwindow search algorithm.
Yu Cao 0003, Dhaval Salvi, Kenton Oliver, Jarrell W. Waggoner, Song Wang 0002
CVPR6
2010 Multiple Cortical Surface Correspondence Using Pairwise Shape Similarity
Pahal Dalal, Feng Shi 0001, Dinggang Shen, Song Wang 0002
MICCAI (1)4
2009 Fast multiple shape correspondence by pre-organizing shape instances
abstract
Accurately identifying corresponded landmarks from a population of shape instances is the major challenge in constructing statistical shape models. In general, shape-correspondence methods can be grouped into one of two categories: global methods and pair-wise methods. In this paper, we develop a new method that attempts to address the limitations of both the global and pair-wise methods. In particular, we reorganize the input population into a tree structure that incorporates global information about the population of shape instances, where each node in the tree represents a shape instance and each edge connects two very similar shape instances. Using this organized tree, neighboring shape instances can be corresponded efficiently and accurately by a pair-wise method. In the experiments, we evaluate the proposed method and compare its performance to five available shape correspondence methods and show the proposed method achieves the accuracy of a global method with speed of a pair-wise method.
Brent C. Munsell, Andrew Temlyakov, Song Wang 0002
CVPR3
2009 3D open-surface shape correspondence for statistical shape modeling: Identifying topologically consistent landmarks
abstract
Shape correspondence, which aims at accurately identifying corresponding landmarks from a given population of shape instances, is a very challenging step in constructing a statistical shape model such as the Point Distribution Model. The state-of-the-art methods such as MDL and SPHARM are primarily focused on closed-surface shape correspondence. In this paper we develop a novel method aimed at identifying accurately corresponding landmarks on 3D open-surfaces with a closed boundary. In particular, we enforce explicit topology consistency on the identified landmarks to ensure that they form a simple, consistent triangle mesh to more accurately model the correspondence of the underlying continuous shape instances. The proposed method also ensures the correspondence of the boundary of the open surfaces. For our experiments, we test the proposed method by constructing a statistical shape model of the human diaphragm from 26 shape instances.
Pahal Dalal, Lili Ju, Michael McLaughlin, Xiangrong Zhou, Hiroshi Fujita 0001, Song Wang 0002
ICCV6
2008 Evaluating Shape Correspondence for Statistical Shape Analysis: A Benchmark Study
abstract
This paper introduces a new benchmark study to evaluate the performance of landmark-based shape correspondence used for statistical shape analysis. Different from previous shape-correspondence evaluation methods, the proposed benchmark first generates a large set of synthetic shape instances by randomly sampling a given statistical shape model that defines a ground-truth shape space. We then run a test shape-correspondence algorithm on these synthetic shape instances to identify a set of corresponded landmarks. According to the identified corresponded landmarks, we construct a new statistical shape model, which defines a new shape space. We finally compare this new shape space against the ground-truth shape space to determine the performance of the test shape-correspondence algorithm. In this paper, we introduce three new performance measures that are landmark independent to quantify the difference between the ground-truth and the newly derived shape spaces. By introducing a ground-truth shape space that is defined by a statistical shape model and three new landmark-independent performance measures, we believe the proposed benchmark allows for a more objective evaluation of shape correspondence than previous methods. In this paper, we focus on developing the proposed benchmark for 2D shape correspondence. However it can be easily extended to 3D cases.
Brent C. Munsell, Pahal Dalal, Song Wang 0002
IEEE Trans. Pattern Anal. Mach. Intell.3
2008 Globally Optimal Grouping for Symmetric Closed Boundaries by Combining Boundary and Region Information
abstract
Many natural and man-made structures have a boundary that shows a certain level of bilateral symmetry, a property that plays an important role in both human and computer vision. In this paper, we present a new grouping method for detecting closed boundaries with symmetry. We first construct a new type of grouping token in the form of symmetric trapezoids by pairing line segments detected from the image. A closed boundary can then be achieved by connecting some trapezoids with a sequence of gap-filling quadrilaterals. For such a closed boundary, we define a unified grouping cost function in a ratio form: the numerator reflects the boundary information of proximity and symmetry and the denominator reflects the region information of the enclosed area. The introduction of the region-area information in the denominator is able to avoid a bias toward shorter boundaries. We then develop a new graph model to represent the grouping tokens. In this new graph model, the grouping cost function can be encoded by carefully designed edge weights and the desired optimal boundary corresponds to a special cycle with a minimum ratio-form cost. We finally show that such a cycle can be found in polynomial time using a previous graph algorithm. We implement this symmetry-grouping method and test it on a set of synthetic data and real images. The performance is compared to two previous grouping methods that do not consider symmetry in their grouping cost functions.
Joachim S. Stahl, Song Wang 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2007 A Fast 3D Correspondence Method for Statistical Shape Modeling
abstract
Accurately identifying corresponded landmarks from a population of shape instances is the major challenge in constructing statistical shape models. In this paper, we address this landmark-based shape-correspondence problem for 3D cases by developing a highly efficient landmark-sliding algorithm. This algorithm is able to quickly refine all the landmarks in a parallel fashion by sliding them on the 3D shape surfaces. We use 3D thin-plate splines to model the shape-correspondence error so that the proposed algorithm is invariant to affine transformations and more accurately reflects the nonrigid biological shape deformations between different shape instances. In addition, the proposed algorithm can handle both open-and closed-surface shape, while most of the current 3D shape-correspondence methods can only handle genus-0 closed surfaces. We conduct experiments on 3D hippocampus data and compare the performance of the proposed algorithm to the state-of-the-art MDL and SPHARM methods. We find that, while the proposed algorithm produces a shape correspondence with a better or comparable quality to the other two, it takes substantially less CPU time. We also apply the proposed algorithm to correspond 3D diaphragm data which have an open-surface shape.
Pahal Dalal, Brent C. Munsell, Song Wang 0002, Jijun Tang, Kenton Oliver, Hiroaki Ninomiya, Xiangrong Zhou, Hiroshi Fujita 0001
CVPR3
2007 A New Benchmark for Shape Correspondence Evaluation
Brent C. Munsell, Pahal Dalal, Song Wang 0002
MICCAI (1)3
2007 Global Detection of Salient Convex Boundaries
Song Wang 0002, Joachim S. Stahl, Adam Bailey, Michael Dropps
Int. J. Comput. Vis.1
2007 Edge Grouping Combining Boundary and Region Information
abstract
This paper introduces a new edge-grouping method to detect perceptually salient structures in noisy images. Specifically, we define a new grouping cost function in a ratio form, where the numerator measures the boundary proximity of the resulting structure and the denominator measures the area of the resulting structure. This area term introduces a preference towards detecting larger-size structures and, therefore, makes the resulting edge grouping more robust to image noise. To find the optimal edge grouping with the minimum grouping cost, we develop a special graph model with two different kinds of edges and then reduce the grouping problem to finding a special kind of cycle in this graph with a minimum cost in ratio form. This optimal cycle-finding problem can be solved in polynomial time by a previously developed graph algorithm. We implement this edge-grouping method, test it on both synthetic data and real images, and compare its performance against several available edge-grouping and edge-linking methods. Furthermore, we discuss several extensions of the proposed method, including the incorporation of the well-known grouping cues of continuity and intensity homogeneity, introducing a factor to balance the contributions from the boundary and region information, and the prevention of detecting self-intersecting boundaries.
Joachim S. Stahl, Song Wang 0002
IEEE Trans. Image Process.2
2006 Image-Segmentation Evaluation From the Perspective of Salient Object Extraction
abstract
Image segmentation and its performance evaluation are very difficult but important problems in computer vision. A major challenge in segmentation evaluation comes from the fundamental conflict between generality and objectivity: For general-purpose segmentation, the ground truth and segmentation accuracy may not be well defined, while embedding the evaluation in a specific application, the evaluation results may not be extended to other applications. We present in this paper a new benchmark for evaluating image segmentation. Specifically, we formulate image segmentation as identifying the single most perceptually salient structure from an image. We collect a large variety of test images that conforms to this specific formulation, construct unambiguous ground truth for each image, and define a reliable way to measure the segmentation accuracy. We then present two special strategies to further address two important issues: (a) the most salient structures in some real images may not be unique or unambiguously defined, and (b) many available image-segmentation methods are not developed to directly extract a single salient structure. Finally, we apply this benchmark to evaluate and compare the performance of several state-of-the-art image-segmentation methods, including the normalized-cut method, the level-set method, the efficient graph-based method, the mean-shift method, and the ratio-contour method.
Feng Ge, Song Wang 0002, Tiecheng Liu
CVPR (1)2
2006 Globally Optimal Grouping for Symmetric Boundaries
abstract
Many natural and man-made structures have a boundary that shows certain level of bilateral symmetry, a property that has been used to solve many computer-vision tasks. In this paper, we present a new grouping method for detecting closed boundaries with symmetry. We first construct a new type of grouping token in the form of a symmetric trapezoid, with which we can flexibly incorporate various boundary and region information into a unified grouping cost function. Particularly, this grouping cost function integrates Gestalt laws of proximity, closure, and continuity, besides the desirable boundary symmetry. We then develop a graph algorithm to find the boundary that minimizes this grouping cost function in a globally optimal fashion. Finally, we test this method by some experiments on a set of natural and medical images.
Joachim S. Stahl, Song Wang 0002
CVPR (1)2
2006 Open-Curve Shape Correspondence Without Endpoint Correspondence
Theodor Richardson, Song Wang 0002
MICCAI (1)2
2005 Convex Grouping Combining Boundary and Region Information
abstract
Convexity is an important geometric property of many natural and man-made structures. Prior research has shown that it is imperative to many perceptual-organization and image-understanding tasks. This paper presents a new grouping method for detecting convex structures from noisy images in a globally optimal fashion. Particularly, this method combines both region and boundary information: the detected structural boundary is closed and well aligned with detected edges while the enclosed region has good intensity homogeneity. We introduce a ratio-form cost function for measuring the structural desirability, which avoids a possible bias to detect small structures. A new fragment-pruning algorithm is developed to achieve the structural convexity. The proposed method can also be extended to detect open boundaries, which correspond to the structures that are partially cropped by the image perimeter and incorporate a human-computer interaction for detecting a convex boundary around a specified point. We test the proposed method on a set of real images and compare it with the Jacobs'convex-grouping method.
Joachim S. Stahl, Song Wang 0002
ICCV2
2005 Nonrigid Shape Correspondence Using Landmark Sliding, Insertion and Deletion
Theodor Richardson, Song Wang 0002
MICCAI (2)2
2005 Salient Closed Boundary Extraction with Ratio Contour
abstract
We present ratio contour, a novel graph-based method for extracting salient closed boundaries from noisy images. This method operates on a set of boundary fragments that are produced by edge detection. Boundary extraction identifies a subset of these fragments and connects them sequentially to form a closed boundary with the largest saliency. We encode the Gestalt laws of proximity and continuity in a novel boundary-saliency measure based on the relative gap length and average curvature when connecting fragments to form a closed boundary. This new measure attempts to remove a possible bias toward short boundaries. We present a polynomial-time algorithm for finding the most-salient closed boundary. We also present supplementary preprocessing steps that facilitate the application of ratio contour to real images. We compare ratio contour to two closely related methods for extracting closed boundaries: Elder and Zucker's method based on the shortest-path algorithm and Williams and Thornber's method based on spectral analysis and a strongly-connected-components algorithm. This comparison involves both theoretic analysis and experimental evaluation on both synthesized data and real images.
Song Wang 0002, Toshiro Kubota, Jeffrey Mark Siskind
IEEE Trans. Pattern Anal. Mach. Intell.1
2003 Salient Boundary Detection using Ratio Contour
abstract
This paper presents a novel graph-theoretic approach, named ratio con- tour, to extract perceptually salient boundaries from a set of noisy bound- ary fragments detected in real images. The boundary saliency is defined using the Gestalt laws of closure, proximity, and continuity. This pa- per first constructs an undirected graph with two different sets of edges: solid edges and dashed edges. The weights of solid and dashed edges measure the local saliency in and between boundary fragments, respec- tively. Then the most salient boundary is detected by searching for an optimal cycle in this graph with minimum average weight. The proposed approach guarantees the global optimality without introducing any biases related to region area or boundary length. We collect a variety of images for testing the proposed approach with encouraging results.
Song Wang 0002, Toshiro Kubota, Jeffrey Mark Siskind
NIPS1
2003 Image Segmentation with Ratio Cut
abstract
This paper proposes a new cost function, cut ratio, for segmenting images using graph-based methods. The cut ratio is defined as the ratio of the corresponding sums of two different weights of edges along the cut boundary and models the mean affinity between the segments separated by the boundary per unit boundary length. This new cost function allows the image perimeter to be segmented, guarantees that the segments produced by bipartitioning are connected, and does not introduce a size, shape, smoothness, or boundary-length bias. The latter allows it to produce segmentations where boundaries are aligned with image edges. Furthermore, the cut-ratio cost function allows efficient iterated region-based segmentation as well as pixel-based segmentation. These properties may be useful for some image-segmentation applications. While the problem of finding a minimum ratio cut in an arbitrary graph is NP-hard, one can find a minimum ratio cut in the connected planar graphs that arise during image segmentation in polynomial time. While the cut ratio, alone, is not sufficient as a baseline method for image segmentation, it forms a good basis for an extended method of image segmentation when combined with a small number of standard techniques. We present an implemented algorithm for finding a minimum ratio cut, prove its correctness, discuss its application to image segmentation, and present the results of segmenting a number of medical and natural images using our techniques.
Song Wang 0002, Jeffrey Mark Siskind
IEEE Trans. Pattern Anal. Mach. Intell.1
2003 Image Segmentation with Ratio Cut - Supplemental Material
Song Wang 0002, Jeffrey Mark Siskind
IEEE Trans. Pattern Anal. Mach. Intell.1
2001 Image Segmentation with Minimum Mean Cut
Song Wang 0002, Jeffrey Mark Siskind
ICCV1