Wei Feng 0005

dblp:17/1152-5 · DBLP profile ↗
← Back
186ranked-venue papers
17as first author
112since 2021 · last 2026
0000-0003-3809-1086ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 121 · 10 first-author · 68 since 2021Artificial intelligence and machine learning · 104 · 13 first-author · 69 since 2021Security and privacy · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 UV-RGS: Relightable 3D Gaussian Splatting from Unposed Views Under Varied Illuminations
abstract
The latest advancements in scene relighting have been predominantly driven by inverse rendering with 3D Gaussian Splatting (3DGS). However, existing methods remain overly reliant on precise camera parameters under static illumination conditions, which is prohibitively expensive and even impractical in real-world scenarios. In this paper, we propose a novel learning from Unposed views under Varied illuminations Relightable 3D Gaussian Splatting (dubbed UV-RGS), to address this challenge by jointly optimizing camera poses, 3DGS representations, surface materials, and environment illuminations (i.e., unknown and varied lighting conditions in training) using only unposed views under varied lightings. Firstly, UV-RGS presents a viewpoint dividing strategy to group inputs into constituent units, enabling each unit can perform similar poses and illuminations. Next, for each unit, to get the constituent model, UV-RGS establishes an incrementally pose learning module to estimate coarse camera parameters, which also enjoy a proxy-view refinement to alleviate the sparse view learning. Additionally, for all constituent unit models, we introduce a holistic model learning strategy that integrates progressive unit aggregation component and the 3DGS coupled with camera poses joint optimization, which realizes the scene high-fidelity perception by the physical-based rendering. Extensive experiments on both real-world and synthetic challenging datasets demonstrate the effectiveness of UV-RGS, achieving the state-of-the-art performance for scene inverse rendering by learning 3DGS from only unposed views under varied illuminations.
Wei Feng 0005, Chi Huang, Qi Zhang 0071, Qian Zhang 0051, Nan Li 0048
AAAI1
2026 E-Logic Prompt: Unified Energy-Logic Framework for Continual Visual Question Answering
abstract
Prompt tuning has shown promise for continual visual question answering (CVQA), facilitating modular and transferable knowledge across tasks. However, existing approaches often overlook the guiding role of prompts in the model’s implicit reasoning process. This oversight can lead to inconsistent reasoning paths and performance degradation across tasks. To address this issue, we propose the E Logic Prompt framework, which employs energy-based models (EBMs) to model the semantic compatibility between prompts and queries. In this framework, prompts function not only as adapters but also as reasoning guides that help maintain coherence throughout the inference process. The framework enforces logical consistency at three levels. At the input level, it selects semantically aligned prompts by minimizing the energy between queries and prompts. Within the model, it aligns intermediate representations with prompts across layers to preserve step-by-step reasoning. Across tasks, it applies energy-based constraints to regulate prompt behavior, effectively suppressing semantic drift and enabling prompt reuse. These three levels of consistency together enhance the guiding capacity of prompts, allowing them to steer the model toward more stable and coherent reasoning. Extensive experiments show that E Logic Prompt outperforms existing methods in both accuracy and knowledge retention, while effectively maintaining balanced cross-modal reasoning throughout continual learning.
Jiayao Tan, Fuyuan Hu, Wei Feng 0005
AAAI4
2026 Imaging-sensitive defect detection for high-surface-quality products
Nan Li 0048, Qian Zhang 0051, Qi Zhang 0071, Wei Feng 0005
Expert Syst. Appl.5
2026 SAVSR++: An All-Stage Scale-Aware and Temporal Omniscient Framework for Arbitrary-Scale Video Super-Resolution
Hongying Liu 0001, Zekun Li 0014, Wenbin Zuo, Fanhua Shang, Liang Wang 0001, Wei Feng 0005
Int. J. Comput. Vis.6
2026 Video Shadow Detection with Intra-and Inter-video Cooperation
Zhihao Chen 0004, Junting Zhao, Lei Zhu 0003, Huazhu Fu, Wei Feng 0005
Int. J. Comput. Vis.6
2026 Adversarial rain attack and defensive deraining for DNN perception
Liming Zhai, Qing Guo 0003, Felix Juefei-Xu, Xiaofei Xie, Lei Ma 0003, Wei Feng 0005, Shengchao Qin, Yang Liu 0003
Neural Networks6
2026 Causal Inference via Style Bias Deconfounding for Domain Generalization
abstract
Deep neural networks (DNNs) often struggle with out-of-distribution data, limiting their reliability in real-world visual applications. To address this issue, domain generalization methods have been developed to learn domain-invariant features from single or multiple training domains, enabling generalization to unseen testing domains. However, existing approaches usually overlook the impact of style frequency within the training set. This oversight predisposes models to capture spurious visual correlations caused by style confounding factors, rather than learning truly causal representations, thereby undermining inference reliability. In this work, we introduce Style Deconfounding Causal Learning (SDCL), a novel causal inference-based framework that explicitly addresses style as a confounding factor to enhance domain generalization in image modalities. Our approaches begins with constructing a structural causal model (SCM) tailored to the domain generalization problem and applies a backdoor adjustment strategy to account for style influence. Building on this foundation, we design a style-guided expert module (SGEM) to adaptively clusters style distributions during training, capturing the global confounding style. Additionally, a backdoor causal learning module (BDCL) performs causal interventions during feature extraction, ensuring fair integration of global confounding styles into sample predictions, effectively reducing style bias. The SDCL framework is highly versatile and can be seamlessly integrated with state-of-the-art data augmentation techniques. Extensive experiments across diverse natural and medical image recognition tasks validate its efficacy, demonstrating superior performance in both multi-domain and the more challenging single-domain generalization scenarios.
Di Lin 0002, Hao Chen 0011, Hongying Liu 0001, Wei Feng 0005
IEEE Trans. Pattern Anal. Mach. Intell.6
2026 BEVTrack: Multi-View Multi-Human Registration and Tracking in the Bird's Eye View
abstract
We handle a new problem of multi-view multi-human tracking in the bird's eye view (BEV). Different from previous works, we require neither the calibration among the multi-view cameras nor the actually captured BEV video. This makes the studied problem closer to real-world applications, however, more challenging. For this purpose, in this work, we propose a novel BEVTrack scheme. Specifically, given multi-view videos, we first use a virtual BEV transform module to obtain the BEV for each view. Then, we propose a unified BEV alignment module to fuse the respectively generated BEVs, in which we specifically design the self-supervised losses by considering both the spatial consistency and the temporal continuity. During the inference, we design the camera-subject collaborative registration and tracking strategy to make use of the mutual dependence between the multi-view cameras and the multiple targets, to achieve the desired BEV tracking. We also build a new benchmark for training and evaluation, the experimental results on which have verified the rationality of the problem and the effectiveness of our method.
Zekun Qian, Wei Feng 0005, Rui-Ze Han
IEEE Trans. Pattern Anal. Mach. Intell.2
2026 SSL-SSAW: Self-supervised Learning with Sigmoid Self-attention Weighting for question-based Sign Language Translation
Zekang Liu, Wei Feng 0005, Fanhua Shang, Lianyu Hu 0003, Jichao Feng, Liqing Gao
Pattern Recognit.2
2026 From indoor to outdoor: Unsupervised domain adaptive gait recognition
Likai Wang 0002, Wei Feng 0005, Rui-Ze Han, Xiangqun Zhang 0003, Yanjie Wei, Song Wang 0002
Pattern Recognit.2
2026 Category-agnostic object re-identification
Likai Wang 0002, Rui-Ze Han, Bingliang Jiao, Wei Feng 0005
Pattern Recognit.4
2026 Dual Domain-Attribute Learning Framework With Asynchronous Adapters for Continual Test-Time Adaptation
abstract
Continual test-time domain adaptation (CTTA) aims to adapt a pre-trained source model to a stream of continually evolving unlabeled target domains, facilitating model deployment in dynamic and non-stationary environments. Contemporary works usually encode domain-specific (DS) style information in a domain-agnostic manner, synchronizing with the learning of domain-invariant (DI) semantic information. This scheme forces DS information to be optimized using the weights of the previous domain, corrupted by cross-domain discrepancies, and hence leads to error accumulation and catastrophic forgetting issues. Inspired by the Attribute Memory Model (AMM) in brain neuroscience, we propose a dual domain-attribute learning framework based on independent asynchronous updates, aiming to imitate how brain learns new knowledge without forgetting. Concretely, we explicitly decompose the continual adaptation process into two complementary systems: an event-based learning system (ELS) that captures DS style representations and a knowledge-based learning system (KLS) that concentrates on the DI structural characteristics. The ELS first detects differences in the distribution of data streams, and actively builds an adapter pool for new latent domains. The KLS adopts a cross-domain shared adapter emphasizing general knowledge, and cooperates with the adapter from ELS to jointly guide adaptation. To make DS and DI knowledge collaboratively working, we exploit a gradient conflict solver to ease the conflict between the past and current DI knowledge, realizing a win-win game (i.e., no interference adaptation) across evolving domains. Our framework have been extensively evaluated on four benchmarks and outperformed the state-of-the-art approaches on both segmentation and classification CTTA tasks.
Yuntong Tian, Kang Li 0007, Tianyang He, Pheng-Ann Heng, Wei Feng 0005
IEEE Trans. Image Process.6
2025 Dual Semantic Guidance for Open Vocabulary Semantic Segmentation
abstract
Open-vocabulary semantic segmentation aims to enable models to segment arbitrary categories. Currently, though pre-trained Vision-Language Models (VLMs) like CLIP have established a robust foundation for this task by learning to match text and image representations from large-scale data, their lack of pixel-level recognition necessitates further fine-tuning. Most existing methods leverage text as a guide to achieve pixel-level recognition. However, the inherent biases in text semantic descriptions and the lack of pixel-level supervisory information make it challenging to fine-tune CLIP-based models effectively. This paper considers leveraging image-text data to simultaneously capture the semantic information contained in both image and text, thereby constructing Dual Semantic Guidance and corresponding pixel-level pseudo annotations. Particularly, the visual semantic guidance is enhanced via explicitly exploring foreground regions and minimizing the influence of background. The dual semantic guidance is then jointly utilized to fine-tune CLIP-based segmentation models, achieving decent fine-grained recognition capabilities. As the comprehensive evaluation shows, our method outperforms state-of-art results with large margins, on eight commonly used datasets with/without background.
Tingliang Feng, Fan Lyu, Fanhua Shang, Wei Feng 0005
CVPR5
2025 Beyond Background Shift: Rethinking Instance Replay in Continual Semantic Segmentation
abstract
In this work, we focus on continual semantic segmentation (CSS), where segmentation networks are required to continuously learn new classes without erasing knowledge of previously learned ones. Although storing images of old classes and directly incorporating them into the training of new models has proven effective in mitigating catastrophic forgetting in classification tasks, this strategy presents notable limitations in CSS. Specifically, the stored and new images with partial category annotations leads to confusion between unannotated categories and the background, complicating model fitting. To tackle this issue, this paper proposes a novel Enhanced Instance Replay (EIR) method, which not only preserves knowledge of old classes while simultaneously eliminating background confusion by instance storage of old classes, but also mitigates background shifts in the new images by integrating stored instances with new images. By effectively resolving background shifts in both stored and new images, EIR alleviates catastrophic forgetting in the CSS task, thereby enhancing the model’s capacity for CSS. Experimental results validate the efficacy of our approach, which significantly outperforms state-of-the-art CSS methods. The code is available at https://github.com/YikeYin97/EIR.
Hongmei Yin, Tingliang Feng, Fan Lyu, Fanhua Shang, Hongying Liu 0001, Wei Feng 0005
CVPR6
2025 Generative Hard Example Augmentation for Semantic Point Cloud Segmentation
abstract
The recent progress in semantic point cloud segmentation is attributed to deep networks, which require a large amount of point cloud data for training. However, how to collect substantial point-wise annotations of the point clouds at affordable cost for the end-to-end network training still needs to be solved. In this paper, we propose Generative Hard Example Augmentation (GHEA) to achieve novel examples of point clouds, which enrich the data for training the segmentation network. Firstly, GHEA employs the generative network to embed the discrepancy between the point clouds into the latent space. From the latent space, we sample multiple discrepancies for reshaping a point cloud to various examples, contributing to the richness of the training data. Secondly, GHEA mixes the reshaped point clouds by respecting their segmentation errors. This mixup allows the reshaped point clouds, which are difficult to segment, to join as the challenging example for network training. We evaluate the effectiveness of GHEA, which helps the popular segmentation networks to improve the performances.
Qi Zhang 0071, Jibin Peng, Wei Feng 0005, Di Lin 0002
CVPR4
2025 VOVTrack: Exploring the Potentiality in Raw Videos for Open-Vocabulary Multi-Object Tracking
Zekun Qian, Rui-Ze Han, Junhui Hou, Linqi Song, Wei Feng 0005
ICCV5
2025 COVTrack: Continuous Open-Vocabulary Tracking via Adaptive Multi-Cue Fusion
Zekun Qian, Rui-Ze Han, Junhui Hou, Wei Feng 0005
ICCV5
2025 Greg: GEometry-Aware RegIon Refinement for Sign Language Video Generation
Tongkai Shi, Lianyu Hu 0003, Fanhua Shang, Liqing Gao, Wei Feng 0005
ICCV5
2025 SU-RGS: Relightable 3D Gaussian Splatting from Sparse Views Under Unconstrained Illuminations
Qi Zhang 0071, Chi Huang, Qian Zhang 0051, Nan Li 0048, Wei Feng 0005
ICCV5
2025 FedAGC: Federated Continual Learning with Asymmetric Gradient Correction
Chengchao Zhang, Fanhua Shang, Hongyin Liu, Wei Feng 0005
ICCV5
2025 Controllable Continual Test-Time Adaptation
abstract
Continual Test-Time Adaptation (CTTA) is an emerging and challenging task where a model trained in a source domain must adapt to continuously changing conditions during testing, without access to the original source data. CTTA is prone to error accumulation due to uncontrollable domain shifts, leading to blurred decision boundaries between categories. Existing CTTA methods primarily focus on suppressing domain shifts, which proves inadequate during the unsupervised test phase. In contrast, we introduce a novel approach that guides rather than suppresses these shifts. Specifically, we propose Controllable Continual Test-Time Adaptation (C-CoTTA), which explicitly prevents any single category from encroaching on others, thereby mitigating the mutual influence between categories caused by uncontrollable shifts. Moreover, our method reduces the sensitivity of model to domain transformations, thereby minimizing the magnitude of category shifts. Extensive quantitative experiments demonstrate the effectiveness of our method, while qualitative analyses, such as t-SNE plots, confirm the theoretical validity of our approach. Our code is available at https://github.com/RenshengJi/C-CoTTA.
Ziqi Shi, Fan Lyu, Fanhua Shang, Fuyuan Hu, Wei Feng 0005, Zhang Zhang 0001, Liang Wang 0001
ICME6
2025 Improving Generalization in Federated Learning with Highly Heterogeneous Data via Momentum-Based Stochastic Controlled Weight Averaging
abstract
For federated learning (FL) algorithms such as FedSAM, their generalization capability is crucial for real-word applications. In this paper, we revisit the generalization problem in FL and investigate the impact of data heterogeneity on FL generalization. We find that FedSAM usually performs worse than FedAvg in the case of highly heterogeneous data, and thus propose a novel and effective federated learning algorithm with Stochastic Weight Averaging (called \texttt{FedSWA}), which aims to find flatter minima in the setting of highly heterogeneous data. Moreover, we introduce a new momentum-based stochastic controlled weight averaging FL algorithm (\texttt{FedMoSWA}), which is designed to better align local and global models. Theoretically, we provide both convergence analysis and generalization bounds for \texttt{FedSWA} and \texttt{FedMoSWA}. We also prove that the optimization and generalization errors of \texttt{FedMoSWA} are smaller than those of their counterparts, including FedSAM and its variants. Empirically, experimental results on CIFAR10/100 and Tiny ImageNet demonstrate the superiority of the proposed algorithms compared to their counterparts.
Junkang Liu, Yuanyuan Liu 0001, Fanhua Shang, Hongying Liu 0001, Wei Feng 0005
ICML6
2025 2D Gaussian Splatting for Outdoor Scene Decomposition and Relighting
abstract
Gaussian splatting techniques have recently revolutionized outdoor scene decomposition and relighting through multi-view images. However, achieving high rendering quality still requires a fixed lighting condition among all input views, which is costly or even impractical to capture in outdoor scenes. In this paper, we propose outdoor scene decomposition and relighting with 2D Gaussian splatting (OSDR-GS), a novel inverse rendering strategy under outdoor changing and unknown lighting conditions. Firstly, we present a lighting-based group learning framework that categorizes input images into multiple lighting groups, to learn the separate lighting from each group individually. Secondly, OSDR-GS introduces a fine-grained outdoor lighting component to represent sun-light and sky-light, respectively, which are also adjusted via the correlative exposure factors adaptively. Finally, we construct a visibility-driven shadow module to characterize the nuanced interplay of light and occlusion realistically, for eliminating the uncertainty of dark pixels on lighting-based group learning. Extensive experiments on multiple challenging outdoor datasets validate the effectiveness of OSDR-GS, which achieves the state-of-the-art performance in changing lighting scene inverse rendering.
Wei Feng 0005, Kangrui Ye, Qi Zhang 0071, Qian Zhang 0051, Nan Li 0048
IJCAI1
2025 TriGS: Tri-consistency 3D Gaussian Splatting from Sparse and Unposed Views
abstract
Recent advances in 3D scene representation, particularly 3D Gaussian Splatting (3DGS), have demonstrated remarkable photorealistic rendering capabilities. However, the heavy reliance on dense and precisely calibrated camera configurations limits effectiveness in sparse view and unposed scenarios. In this paper, we present Tri-consistency 3D Gaussian Splatting (dubbed TriGS), a novel framework that jointly optimizes 3DGS parameters and camera poses only from sparse and unposed images via triple consistency supervisions coupled with the adaptive regularization strategy. We first estimate coarse camera poses by exploiting 3DGS's anisotropic properties through iterative relative pose optimization. Building upon this foundation, we introduce cross-view consistency enforcement through synchronized photometric color, geometric structure, and deep feature, effectively resolving rendering ambiguities with auxiliary supervisions. A unified rendering paradigm is also proposed to jointly refine Gaussian primitives and camera poses by transforming positions, covariances, and spherical harmonics. To combat overfitting inherent in joint optimization, we devise an adaptive regularization mechanism that strategically samples hard viewpoints based on baseline distances and training dynamics, enforcing projection consistency through deep feature priors. Extensive experiments on multiple challenging real-world datasets validate the effectiveness of TriGS, which achieves satisfactory results to set a new state-of-the-art without the reliance on external pose priors only under sparse and unposed view inputs.
Chi Huang, Qi Zhang 0071, Qian Zhang 0051, Nan Li 0048, Yipu Gong, Wei Feng 0005
ACM Multimedia7
2025 Gloss-Free Sign Language Translation With Optical-Flow Guided Two-Stream Network
abstract
Sign Language Translation (SLT) is a challenging task, with existing approaches often constrained by the necessity of gloss annotations (sign language lexemes). Since gloss annotations are expensive and difficult to obtain, they severely limit the scalability of SLT in practical applications. To address this, we propose leveraging inter-frame optical flow as crucial prior information to enhance visual feature extraction, which naturally corresponds to signers’ semantic movements. We design a novel gloss-free optical-Flow Guided Two-Stream Network architecture (FGTSN): The Image Encoder Stream employs optical flow as a static prior to guide spatial attention while the Skeleton Encoder Stream integrates it to enrich the keypoint features and enhance the representation of local temporal dynamics and structural changes. To tackle the alignment challenge inherent in gloss-free SLT, we introduce a novel auxiliary block prediction loss that strengthens the local temporal relationships of the visual features. Experiments on three major datasets demonstrate that our FGTSN achieves state-of-the-art performance with significant improvements over existing methods, providing a novel and lightweight perspective for gloss-free sign language translation.
Lianyu Hu 0003, Tongkai Shi, Fanhua Shang, Jichao Feng, Wei Feng 0005
MMAsia6
2025 DOVTrack: Data-Efficient Open-Vocabulary Tracking
abstract
Open-Vocabulary Multi-Object Tracking (OVMOT) aims to detect and track multi-category objects including both seen and unseen categories during training. Currently, a significant challenge in this domain is the lack of large-scale annotated video data for training. To address this challenge, this work aims to effectively train the OV tracker using only the existing limited and sparsely annotated video data. We propose a comprehensive training sample space expansion strategy that addresses the fundamental limitation of sparse annotations in OVMOT training. Specifically, for the association task, we develop a diffusion-based feature generation framework that synthesizes intermediate object features between sparsely annotated frames, effectively expanding the training sample space by approximately 3× and enabling robust association learning from temporally continuous features. For the detection task, we introduce a dynamic group contrastive learning approach that generates diverse sample groups through affinity, dispersion, and adversarial grouping strategies, tripling the effective training samples for classification while maintaining sample quality. Additionally, we propose an adaptive localization loss that expands positive sample coverage by lowering IoU thresholds while mitigating noise through confidence-based weighting. Extensive experiments demonstrate that our method achieves state-of-the-art performance on the OVMOT benchmark, surpassing existing methods by 3.8\% in TETA metric, without requiring additional data or annotations. The code will be available at https://github.com/zekunqian/DOVTrack.
Zekun Qian, Rui-Ze Han, Junhui Hou, Wei Feng 0005
NeurIPS5
2025 QBasicVSR: Temporal Awareness Adaptation Quantization for Video Super-Resolution
abstract
While model quantization has become pivotal for deploying super-resolution (SR) networks on mobile devices, existing works focus on quantization methods only for image super-resolution. Different from image SR quantization, the temporal error propagation, shared temporal parameterization, and temporal metric mismatch significantly degrade the quantization performance of a video SR model. To address these issues, we propose the first quantization method, QBasicVSR, for video super-resolution. A novel temporal awareness adaptation post-training quantization (PTQ) framework for video super-resolution with the flow-gradient video bit adaptation and temporal shared layer bit adaptation is presented. Moreover, we put forward a novel fine-tuning method for VSR with the supervision of the full-precision model. Our method achieves extraordinary performance with state-of-the-art efficient VSR approaches, delivering up to $\times$200 faster processing speed while utilizing only 1/8 of the GPU resources. Additionally, extensive experiments demonstrate that the proposed method significantly outperforms existing PTQ algorithms on various datasets. For instance, it attains a 2.53 dB increase on the UDM10 benchmark when quantizing BasicVSR to 4-bit with 100 unlabeled video clips. The code and models will be released on GitHub.
Fanhua Shang, Hongying Liu 0001, Liang Wang 0001, Wei Feng 0005, Yanming Hui
NeurIPS5
2025 Weakly Supervised Instance Action Recognition
abstract
We study the novel problem of weakly supervised instance action recognition (WSiAR) in multi-person (crowd) scenes. We specifically aim to recognize the action of each subject in the crowd, for which we propose the use of a weakly supervised method, considering the expense of large-scale annotations for training. This problem is of great practical value for video surveillance and sports scene analysis. To this end, we investigated and designed a series of weak annotations for the supervision of weakly supervised instance action recognition (WSiAR). We propose two categories of weak label settings, bag labels and sparse labels, to significantly reduce the number of labels. Based on the former, we propose a novel sub-block-aware multi-instance learning (MIL) loss to obtain more effective information from weak labels during training. With respect to the latter, we propose a pseudo label generation strategy for extending sparse labels. This enables our method to achieve results comparable to those of fully supervised methods but with significantly fewer annotations. The experimental results on two benchmarks verified the rationality of the problem definition and effectiveness of the proposed weakly supervised training method in solving our problem.
Haomin Yan, Rui-Ze Han, Wei Feng 0005, Jiewen Zhao, Songmiao Wang
Comput. Vis. Media3
2025 A structure-based disentangled network with contrastive regularization for sign language recognition
Liqing Gao, Lei Zhu 0003, Lianyu Hu 0003, Liang Wang 0001, Wei Feng 0005
Expert Syst. Appl.6
2025 EfficientDeRain+: Learning Uncertainty-Aware Filtering via RainMix Augmentation for High-Efficiency Deraining
Qing Guo 0005, Hua Qi, Jingyang Sun, Felix Juefei-Xu, Lei Ma 0003, Di Lin 0002, Wei Feng 0005, Song Wang 0002
Int. J. Comput. Vis.7
2025 Concept-Guided Open-Vocabulary Temporal Action Detection
Songmiao Wang, Rui-Ze Han, Wei Feng 0005
J. Comput. Sci. Technol.3
2025 Unveiling the Power of Self-Supervision for Multi-View Multi-Human Association and Tracking
abstract
Multi-view multi-human association and tracking (MvMHAT), is an emerging yet important problem for multi-person scene video surveillance, aiming to track a group of people over time in each view, as well as to identify the same person across different views at the same time, which is different from previous MOT and multi-camera MOT tasks only considering the over-time human tracking. This way, the videos for MvMHAT require more complex annotations while containing more information for self-learning. In this work, we tackle this problem with an end-to-end neural network in a self-supervised learning manner. Specifically, we propose to take advantage of the spatial-temporal self-consistency rationale by considering three properties of reflexivity, symmetry, and transitivity. Besides the reflexivity property that naturally holds, we design the self-supervised learning losses based on the properties of symmetry and transitivity, for both appearance feature learning and assignment matrix optimization, to associate multiple humans over time and across views. Furthermore, to promote the research on MvMHAT, we build two new large-scale benchmarks for the network training and testing of different algorithms. Extensive experiments on the proposed benchmarks verify the effectiveness of our method. We have released the benchmark and code to the public.
Wei Feng 0005, Fei Wang 0032, Rui-Ze Han, Yiyang Gan, Zekun Qian, Junhui Hou, Song Wang 0002
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 A large-scale combinatorial benchmark for sign language recognition
Liqing Gao, Liang Wang 0001, Lianyu Hu 0003, Rui-Ze Han, Zekang Liu, Fanhua Shang, Wei Feng 0005
Pattern Recognit.8
2025 Learning with privileged stereo knowledge for monocular absolute 3D human pose estimation
Cunling Bian, Weigang Lu 0002, Wei Feng 0005, Song Wang 0002
Pattern Recognit. Lett.3
2025 NUC-Net: Non-Uniform Cylindrical Partition Network for Efficient LiDAR Semantic Segmentation
abstract
LiDAR semantic segmentation plays a vital role in autonomous driving. Existing voxel-based methods for LiDAR semantic segmentation apply uniform partition to the 3D LiDAR point cloud to form a structured representation based on cartesian/cylindrical coordinates. Although these methods show impressive performance, the drawback of existing voxel-based methods remains in two aspects: 1) it requires a large enough input voxel resolution, which brings a large amount of computation cost and memory consumption. 2) it does not well handle the unbalanced point distribution of LiDAR point cloud. In this paper, we propose a non-uniform cylindrical partition network named NUC-Net to tackle the above challenges. Specifically, we propose the Arithmetic Progression of Interval (API) method to non-uniformly partition the radial axis and generate the voxel representation which is representative and efficient. Moreover, we propose a non-uniform multi-scale aggregation method to improve contextual information. Our method achieves state-of-the-art performance on SemanticKITTI and nuScenes datasets with much faster speed and much less training time. And our method can be a general component for LiDAR semantic segmentation, which significantly improves both the accuracy and efficiency of the uniform counterpart by$4 \times $training faster and$2 \times $GPU memory reduction and$3 \times $inference speedup. We further provide theoretical analysis towards understanding why NUC is effective and how point distribution affects performance. Code is available athttps://github.com/alanWXZ/NUC-Net.
Xuzhi Wang, Wei Feng 0005, Lingdong Kong
IEEE Trans. Circuits Syst. Video Technol.2
2025 A New Benchmark and Algorithm for Clothes-Changing Video Person Re-Identification
abstract
Person re-identification (Re-ID) is a classical computer vision task and has significant applications for public security and information forensics. Recently, long-term Re-ID with clothes-changing has attracted increasing attention. However, existing methods mainly focus on image-based setting, where richer temporal information is overlooked. In this paper, we focus on the relatively new yet practical problem of Clothes-Changing Video-based Re-ID (CCVReID), which is less studied. First, given the dataset shortage, we build two new benchmark datasets for CCVReID problem, including a large-scale synthetic video dataset and a real-world one, both containing human sequences with various clothing changes. Moreover, we systematically study this problem by simultaneously considering the classical appearance feature and temporal feature contained in the video. We develop a dual-branch fusion framework that makes use of the information from both clothes-aware appearance feature and clothes-free gait feature. For better information fusion, a confidence-guided re-ranking strategy is proposed to adaptively balance the weight of these two categories of features. We have released the benchmark and code proposed in this work to the public athttps://github.com/kkw98/CCVReID.
Likai Wang 0002, Xiangqun Zhang 0003, Rui-Ze Han, Yanjie Wei, Song Wang 0002, Wei Feng 0005
IEEE Trans. Inf. Forensics Secur.6
2025 Synthetic-to-Real Video Person Re-ID
abstract
Person re-identification (Re-ID) is an important task and has significant applications for public security and information forensics, which has progressed rapidly with the development of deep learning. In this work, we investigate a novel and challenging setting of Re-ID, i.e., cross-domain video-based person Re-ID. Specifically, we utilize synthetic video datasets as the source domain for training and real-world videos for testing, notably reducing the reliance on expensive real data acquisition and annotation. To harness the potential of synthetic data, we first propose a self-supervised domain-invariant feature learning strategy for both static and dynamic (temporal) features. Additionally, to enhance person identification accuracy in the target domain, we propose a mean-teacher scheme incorporating a self-supervised ID consistency loss. Experimental results across five real datasets validate the rationale behind cross-synthetic-real domain adaptation and demonstrate the efficacy of our method. Notably, the discovery that synthetic data outperforms real data in the cross-domain scenario is a surprising outcome. The code and data are publicly available at https://github.com/XiangqunZhang/UDA_Video_ReID.
Xiangqun Zhang 0003, Rui-Ze Han, Likai Wang 0002, Linqi Song, Junhui Hou, Wei Feng 0005
IEEE Trans. Inf. Forensics Secur.6
2025 CarveNet: Carving Point-Block for Complex 3D Shape Completion
abstract
3D point cloud completion is very challenging because it relies on accurately understanding the complex 3D shapes (e.g., high-curvature, concave/convex, and hollowed-out 3D shapes) and the unknown & diverse patterns of the partially available point clouds. In this paper, we propose a novel solution, i.e.,Point-block Carving(PC), for completing the complex 3D point cloud completion. Given the partial point cloud as the guidance, we carve a 3D block that contains the uniformly distributed 3D points, yielding the entire point cloud. We propose a new network architecture to achieve PC, i.e.,CarveNet. This network conducts the exclusive convolution on each block point, where the convolutional kernels are trained on the 3D shape data. CarveNet determines which point should be carved to recover the complete shapes' details effectively. Furthermore, we propose a sensor-aware method for data augmentation, i.e.,SensorAug, for training CarveNet on richer patterns of partial point clouds, thus enhancing the completion power of the network. The extensive evaluations on the ShapeNet, ShapNet-55/34 and KITTI datasets demonstrate the generality of our approach on the partial point clouds with diverse patterns. On these datasets, CarveNet successfully outperforms the state-of-the-art methods.
Qing Guo 0005, Zhijie Wang 0014, Lubo Wang, Haotian Dong, Felix Juefei-Xu, Di Lin 0002, Lei Ma 0003, Wei Feng 0005, Yang Liu 0003
IEEE Trans. Multim.8
2025 DBL-SC: background-independent sign language recognition based on spatial channel separation computation
Zekang Liu, Wei Feng 0005, Liqing Gao, Lianyu Hu 0003
Vis. Comput.2
2024 SAVSR: Arbitrary-Scale Video Super-Resolution via a Learned Scale-Adaptive Network
abstract
Deep learning-based video super-resolution (VSR) networks have gained significant performance improvements in recent years. However, existing VSR networks can only support a fixed integer scale super-resolution task, and when we want to perform VSR at multiple scales, we need to train several models. This implementation certainly increases the consumption of computational and storage resources, which limits the application scenarios of VSR techniques. In this paper, we propose a novel Scale-adaptive Arbitrary-scale Video Super-Resolution network (SAVSR), which is the first work focusing on spatial VSR at arbitrary scales including both non-integer and asymmetric scales. We also present an omni-dimensional scale-attention convolution, which dynamically adapts according to the scale of the input to extract inter-frame features with stronger representational power. Moreover, the proposed spatio-temporal adaptive arbitrary-scale upsampling performs VSR tasks using both temporal features and scale information. And we design an iterative bi-directional architecture for implicit feature alignment. Experiments at various scales on the benchmark datasets show that the proposed SAVSR outperforms state-of-the-art (SOTA) methods at non-integer and asymmetric scales. The source code is available at https://github.com/Weepingchestnut/SAVSR.
Zekun Li 0014, Hongying Liu 0001, Fanhua Shang, Yuanyuan Liu 0001, Wei Feng 0005
AAAI6
2024 COMMA: Co-articulated Multi-Modal Learning
abstract
Pretrained large-scale vision-language models such as CLIP have demonstrated excellent generalizability over a series of downstream tasks. However, they are sensitive to the variation of input text prompts and need a selection of prompt templates to achieve satisfactory performance. Recently, various methods have been proposed to dynamically learn the prompts as the textual inputs to avoid the requirements of laboring hand-crafted prompt engineering in the fine-tuning process. We notice that these methods are suboptimal in two aspects. First, the prompts of the vision and language branches in these methods are usually separated or uni-directionally correlated. Thus, the prompts of both branches are not fully correlated and may not provide enough guidance to align the representations of both branches. Second, it's observed that most previous methods usually achieve better performance on seen classes but cause performance degeneration on unseen classes compared to CLIP. This is because the essential generic knowledge learned in the pretraining stage is partly forgotten in the fine-tuning process. In this paper, we propose Co-Articulated Multi-Modal Learning (COMMA) to handle the above limitations. Especially, our method considers prompts from both branches to generate the prompts to enhance the representation alignment of both branches. Besides, to alleviate forgetting about the essential knowledge, we minimize the feature discrepancy between the learned prompts and the embeddings of hand-crafted prompts in the pre-trained CLIP in the late transformer layers. We evaluate our method across three representative tasks of generalization to novel classes, new target datasets and unseen domain shifts. Experimental results demonstrate the superiority of our method by exhibiting a favorable performance boost upon all tasks with high efficiency. Code is available at https://github.com/hulianyuyy/COMMA.
Lianyu Hu 0003, Liqing Gao, Zekang Liu, Chi-Man Pun, Wei Feng 0005
AAAI5
2024 Long-Tailed Learning as Multi-Objective Optimization
abstract
Real-world data is extremely imbalanced and presents a long-tailed distribution, resulting in models biased towards classes with sufficient samples and performing poorly on rare classes. Recent methods propose to rebalance classes but they undertake the seesaw dilemma (what is increasing performance on tail classes may decrease that of head classes, and vice versa). In this paper, we argue that the seesaw dilemma is derived from the gradient imbalance of different classes, in which gradients of inappropriate classes are set to important for updating, thus prone to overcompensation or undercompensation on tail classes. To achieve ideal compensation, we formulate long-tailed recognition as a multi-objective optimization problem, which fairly respects the contributions of head and tail classes simultaneously. For efficiency, we propose a Gradient-Balancing Grouping (GBG) strategy to gather the classes with similar gradient directions, thus approximately making every update under a Pareto descent direction. Our GBG method drives classes with similar gradient directions to form a more representative gradient and provides ideal compensation to the tail classes. Moreover, we conduct extensive experiments on commonly used benchmarks in long-tailed learning and demonstrate the superiority of our method over existing SOTA methods. Our code is released at https://github.com/WickyLee1998/GBG_v1.
Fan Lyu, Fanhua Shang, Wei Feng 0005
AAAI5
2024 Dynamic Spatial-Temporal Aggregation for Skeleton-Aware Sign Language Recognition
abstract
Skeleton-aware sign language recognition (SLR) has gained popularity due to its ability to remain unaffected by background information and its lower computational requirements. Current methods utilize spatial graph modules and temporal modules to capture spatial and temporal features, respectively. However, their spatial graph modules are typically built on fixed graph structures such as graph convolutional networks or a single learnable graph, which only partially explore joint relationships. Additionally, a simple temporal convolution kernel is used to capture temporal information, which may not fully capture the complex movement patterns of different signers. To overcome these limitations, we propose a new spatial architecture consisting of two concurrent branches, which build input-sensitive joint relationships and incorporates specific domain knowledge for recognition, respectively. These two branches are followed by an aggregation process to distinguishe important joint connections. We then propose a new temporal module to model multi-scale temporal information to capture complex human dynamics. Our method achieves state-of-the-art accuracy compared to previous skeleton-aware methods on four large-scale SLR benchmarks. Moreover, our method demonstrates superior accuracy compared to RGB-based methods in most cases while requiring much fewer computational resources, bringing better accuracy-computation trade-off. Code is available at https://github.com/hulianyuyy/DSTA-SLR.
Lianyu Hu 0003, Liqing Gao, Zekang Liu, Wei Feng 0005
LREC/COLING4
2024 Silhouette-Based 6D Object Pose Estimation
Nan Li 0048, Qian Zhang 0051, Wei Feng 0005
CVM (2)5
2024 From a Bird's Eye View to See: Joint Camera and Subject Registration without the Camera Calibration
abstract
We tackle a new problem of multi-view camera and sub-ject registration in the bird’ s eye view (BEV) without pregiven camera calibration, which promotes the multi-view subject registration problem to a new calibration-free stage. This greatly alleviates the limitation in many practical applications. However, this is a very challenging problem since its only input is several RGB images from different first-person views (FPVs), without the BEV image and the calibration of the FPVs, while the output is a unified plane aggregated from all views with the positions and orientations of both the subjects and cameras in a BEV. For this purpose, we propose an end-to-end framework solving cam-era and subject registration together by taking advantage of their mutual dependence, whose main idea is as below: i) creating a subject view-transform module (VTM) to project each pedestrian from FPV to a virtual BEV, ii) deriving a multi-view geometry-based spatial alignment module (SAM) to estimate the relative camera pose in a unified BEV, iii) selecting and refining the subject and camera registration results within the unified BEV. We collect a new large-scale synthetic dataset with rich annotations for training and evaluation. Additionally, we also collect a real dataset for cross-domain evaluation. The experimental results show the remarkable effectiveness of our method. The code and proposed datasets are available at BEVSee.
Zekun Qian, Rui-Ze Han, Wei Feng 0005, Song Wang 0002
CVPR3
2024 Pose-Guided Fine-Grained Sign Language Video Generation
Tongkai Shi, Lianyu Hu 0003, Fanhua Shang, Jichao Feng, Wei Feng 0005
ECCV (77)6
2024 LRR: Language-Driven Resamplable Continuous Representation against Adversarial Tracking Attacks
abstract
Visual object tracking plays a critical role in visual-based autonomous systems, as it aims to estimate the position and size of the object of interest within a live video. Despite significant progress made in this field, state-of-the-art (SOTA) trackers often fail when faced with adversarial perturbations in the incoming frames. This can lead to significant robustness and security issues when these trackers are deployed in the real world. To achieve high accuracy on both clean and adversarial data, we propose building a spatial-temporal continuous representation using the semantic text guidance of the object of interest. This novel continuous representation enables us to reconstruct incoming frames to maintain semantic and appearance consistency with the object of interest and its clean counterparts. As a result, our proposed method successfully defends against different SOTA adversarial tracking attacks while maintaining high accuracy on clean data. In particular, our method significantly increases tracking accuracy under adversarial attacks with around 90% relative improvement on UAV123, which is even higher than the accuracy on clean data.
Jianlang Chen, Xuhong Ren, Qing Guo 0005, Felix Juefei-Xu, Di Lin 0002, Wei Feng 0005, Lei Ma 0003, Jianjun Zhao 0001
ICLR6
2024 Learning Geometry Consistent Neural Radiance Fields from Sparse and Unposed Views
abstract
The latest progress in novel view synthesis can be attributed to the Neural Radiance Field (NeRF), which requires densely sampled images with precise camera poses. However, collecting dense input images for a NeRF with accurate camera poses is highly expensive in many real-world scenarios. In this paper, we propose to learn Geometry Consistent Neural Radiance Field (GC-NeRF), to tackle this challenge by jointly optimizing a NeRF and its corresponding camera poses with sparse (as low as 2) and unposed views. First, the proposed GC-NeRF establishes image-level geometric consistencies, by producing photometric constraints from inter- and intra-views to update the NeRF and the camera poses in a fine-grained manner. Then, we adopt geometry projection with camera extrinsic parameters to further provide region-level consistency supervisions, which constructs pseudo-pixel labels to capture critical matching correlations. Moreover, we present an adaptive high-frequency mapping function to augment the geometry and texture information of the 3D scene. Extensive experiments on multiple challenging real-world datasets validate the effectiveness of the proposed GC-NeRF, which sets a new state-of-the-art for effectively learning NeRF with sparse and unposed views.
Qi Zhang 0071, Chi Huang, Qian Zhang 0051, Nan Li 0048, Wei Feng 0005
ACM Multimedia5
2024 Rethinking the One-shot Object Detection: Cross-Domain Object Search
abstract
One-shot object detection (OSOD) uses a query patch to identify the same category of object in a target image. As the OSOD setting, the target images are required to contain the object category of the query patch, and the image styles (domains) of the query patch and target images are always similar. However, in practical application, the above requirements are not commonly satisfied. Therefore, we propose a new problem namely Cross-Domain Object Search (CDOS), where the object categories of the query patch and target image are decoupled, and the image styles between them may also be significantly different. For this problem, we develop a new method, which incorporates both foreground-background contrastive learning heads and a domain-generalized feature augmentation technique. This makes our method effectively handle the object category gap and domain distribution gap, between the query patch and target image in the training and testing datasets. We further build a new benchmark for the proposed CDOS problem, on which our method shows significant performance improvements over the comparison methods.
Shuqi Zheng, Rui-Ze Han, Yuzhong Feng, Junhui Hou, Linqi Song, Wei Feng 0005
ACM Multimedia7
2024 Deep Correlated Prompting for Visual Recognition with Missing Modalities
abstract
Large-scale multimodal models have shown excellent performance over a series of tasks powered by the large corpus of paired multimodal training data. Generally, they are always assumed to receive modality-complete inputs. However, this simple assumption may not always hold in the real world due to privacy constraints or collection difficulty, where models pretrained on modality-complete data easily demonstrate degraded performance on missing-modality cases. To handle this issue, we refer to prompt learning to adapt large pretrained multimodal models to handle missing-modality scenarios by regarding different missing cases as different types of input. Instead of only prepending independent prompts to the intermediate layers, we present to leverage the correlations between prompts and input features and excavate the relationships between different layers of prompts to carefully design the instructions. We also incorporate the complementary semantics of different modalities to guide the prompting design for each modality. Extensive experiments on three commonly-used datasets consistently demonstrate the superiority of our method compared to the previous approaches upon different missing scenarios. Plentiful ablations are further given to show the generalizability and reliability of our method upon different modality-missing ratios and types.
Lianyu Hu 0003, Tongkai Shi, Wei Feng 0005, Fanhua Shang
NeurIPS3
2024 Combinational sign language recognition
Liqing Gao, Wei Feng 0005, Fan Lyu, Liang Wang 0001
Comput. Vis. Image Underst.2
2024 Multi-modal fusion architecture search for camera-based semantic scene completion
Xuzhi Wang, Wei Feng 0005
Expert Syst. Appl.2
2024 Contactless interaction recognition and interactor detection in multi-person scenes
Rui-Ze Han, Wei Feng 0005, Haomin Yan, Song Wang 0002
Frontiers Comput. Sci.3
2024 Benchmarking the Complementary-View Multi-human Association and Tracking
Rui-Ze Han, Wei Feng 0005, Zekun Qian, Haomin Yan, Song Wang 0002
Int. J. Comput. Vis.2
2024 Sign language translation with hierarchical memorized context in question answering scenarios
Liqing Gao, Wei Feng 0005, Rui-Ze Han, Di Lin 0002, Liang Wang 0001
Neural Comput. Appl.2
2024 Cross-modal knowledge distillation for continuous sign language recognition
Liqing Gao, Lianyu Hu 0003, Jichao Feng, Lei Zhu 0003, Liang Wang 0001, Wei Feng 0005
Neural Networks7
2024 Scalable frame resolution for efficient continuous sign language recognition
Lianyu Hu 0003, Liqing Gao, Zekang Liu, Wei Feng 0005
Pattern Recognit.4
2024 Overcoming Modality Bias in Question-Driven Sign Language Video Translation
abstract
Question-Driven Sign Language Translation (QSLT) addresses the challenge of translating sign language using pertinent questions in question-answering contexts. However, the pronounced modality complexity between question text and sign video poses a predicament: the model tends to overly depend on questions to generate translations, thereby neglecting the value of visual cues. To tackle this issue, the paper presents a Gloss-Bridged Translator (GBT), which introduces sign gloss as an intermediary conduit to establish semantic connections between questions and videos. By leveraging gloss, visual features are transformed into textual counterparts, mitigating the modality imbalance between these representations. Moreover, a cross-modal contrastive learning strategy is implemented, bolstering the global contextual relevance and local semantic alignment between questions and sign language. The proposed methodology is validated through extensive experiments on the proposed QSL dataset and other public sign language datasets. The results show the efficacy of integrating questions into sign language translation. The GBT yields remarkable improvements over prevailing SLT methods, attesting to its effectiveness and rationale. Our code and dataset is available athttps://github.com/glq-1992/QSL.
Liqing Gao, Fan Lyu, Lei Zhu 0003, Junfu Pu, Liang Wang 0001, Wei Feng 0005
IEEE Trans. Circuits Syst. Video Technol.7
2024 Adversarial Relighting Against Face Recognition
abstract
Deep face recognition (FR) has achieved significantly high accuracy on several challenging datasets and fosters successful real-world applications, even showing high robustness to the illumination variation that is usually regarded as a main threat to the FR system. However, in the real world, illumination variation caused by diverse lighting conditions cannot be fully covered by the limited face dataset. In this paper, we study the threat of lighting against FR from a new angle,i.e.,adversarial attack, and identify a new task,i.e.,adversarial relighting. Given a face image, adversarial relighting aims to produce a naturally relighted counterpart while fooling the state-of-the-art deep FR methods. To this end, we first propose the physical model-based adversarial relighting attack (ARA) denoted asalbedo-quotient-based adversarial relighting attack (AQ-ARA). It generates natural adversarial lighting under the guidance of FR systems and synthesizes adversarially relighted face images. Moreover, we propose theauto-predictive adversarial relighting attack (AP-ARA)by training an adversarial relighting network (ARNet) to automatically predict the adversarial lighting in a one-step manner according to different input faces, allowing efficiency-sensitive applications. More importantly, we propose to transfer the above digital attacks tophysical ARA (Phy-ARA)through a precise relighting device, making the estimated adversarial lighting condition reproducible in the real world. We validate our methods on several state-of-the-art deep FR methods on two public datasets. The extensive and insightful results demonstrate our work can generate realistic adversarial relighted face images fooling face recognition tasks easily, revealing the threat of specific light directions and strengths.
Qian Zhang 0051, Qing Guo 0005, Ruijun Gao, Felix Juefei-Xu, Hongkai Yu, Wei Feng 0005
IEEE Trans. Inf. Forensics Secur.6
2024 Multi-Label Continual Learning Using Augmented Graph Convolutional Network
abstract
Multi-Label Continual Learning (MLCL) is a framework designed for class-incremental multi-label image recognition. However, MLCL faces two critical challenges: the construction of label relationships onpast-missing and future-missing partial labelsof training data, and the problem ofcatastrophic forgetting, which leads to poor generalization. To address these challenges, this study proposes an enhanced version of the Augmented Graph Convolutional Network (AGCN++), capable of constructing cross-task label relationships and mitigating catastrophic forgetting. First, an Augmented Correlation Matrix (ACM) is constructed across all observed classes, incorporating intra-task relationships derived from hard label statistics. Additionally, inter-task relationships are established by leveraging both hard and soft labels obtained from the data, as well as a constructed expert network. Next, a novel partial label encoder (PLE) is introduced for MLCL, enabling the extraction of dynamic class representations for each partial label image as graph nodes. This PLE also facilitates the generation of soft labels, which contribute to the creation of a more persuasive ACM and effectively mitigate forgetting. Lastly, a relationship-preserving constrainter is proposed to address the issue of forgetting label dependencies across old tasks. In the AGCN++, the label relationships topology can be augmented automatically, thereby generating efficient class representations. The effectiveness of the proposed method is evaluated using two multi-label image benchmarks. The experimental results demonstrate that the proposed approach is highly effective in the context of MLCL image recognition. It can establish compelling correlations across tasks, even in scenarios where the old task labels are missing.
Kaile Du, Fan Lyu, Fuyuan Hu, Wei Feng 0005, Fenglei Xu, Hanjing Cheng
IEEE Trans. Multim.5
2024 MFDNet: Multi-Frequency Deflare Network for efficient nighttime flare removal
Yiguo Jiang, Xuhang Chen 0002, Chi-Man Pun, Shuqiang Wang, Wei Feng 0005
Vis. Comput.5
2024 Multi-scale context-aware network for continuous sign language recognition
abstract
The hands and face are the most important parts for expressing sign language morphemes in sign language videos. However, we find that existing Continuous Sign Language Recognition (CSLR) methods lack the mining of hand and face information in visual backbones or use expensive and time-consuming external extractors to explore this information. In addition, the signs have different lengths, whereas previous CSLR methods typically use a fixed-length window to segment the video to capture sequential features and then perform global temporal modeling, which disturbs the perception of complete signs. In this study, we propose a Multi-Scale Context-Aware network (MSCA-Net) to solve the aforementioned problems. Our MSCA-Net contains two main modules: (1) Multi-Scale Motion Attention (MSMA), which uses the differences among frames to perceive information of the hands and face in multiple spatial scales, replacing the heavy feature extractors; and (2) Multi-Scale Temporal Modeling (MSTM), which explores crucial temporal information in the sign language video from different temporal scales. We conduct extensive experiments using three widely used sign language datasets, i.e., RWTH-PHOENIX-Weather-2014, RWTH-PHOENIX-Weather-2014T, and CSL-Daily. The proposed MSCA-Net achieve state-of-the-art performance, demonstrating the effectiveness of our approach.
Senhua Xue, Liqing Gao, Wei Feng 0005
Virtual Real. Intell. Hardw.4
2023 Self-Emphasizing Network for Continuous Sign Language Recognition
abstract
Hand and face play an important role in expressing sign language. Their features are usually especially leveraged to improve system performance. However, to effectively extract visual representations and capture trajectories for hands and face, previous methods always come at high computations with increased training complexity. They usually employ extra heavy pose-estimation networks to locate human body keypoints or rely on additional pre-extracted heatmaps for supervision. To relieve this problem, we propose a self-emphasizing network (SEN) to emphasize informative spatial regions in a self-motivated way, with few extra computations and without additional expensive supervision. Specifically, SEN first employs a lightweight subnetwork to incorporate local spatial-temporal features to identify informative regions, and then dynamically augment original features via attention maps. It's also observed that not all frames contribute equally to recognition. We present a temporal self-emphasizing module to adaptively emphasize those discriminative frames and suppress redundant ones. A comprehensive comparison with previous methods equipped with hand and face features demonstrates the superiority of our method, even though they always require huge computations and rely on expensive extra supervision. Remarkably, with few extra computations, SEN achieves new state-of-the-art accuracy on four large-scale datasets, PHOENIX14, PHOENIX14-T, CSL-Daily, and CSL. Visualizations verify the effects of SEN on emphasizing informative spatial and temporal features. Code is available at https://github.com/hulianyuyy/SEN_CSLR
Lianyu Hu 0003, Liqing Gao, Zekang Liu, Wei Feng 0005
AAAI4
2023 Spatial-Temporal Consistency Constraints for Chinese Sign Language Synthesis
Liqing Gao, Wei Feng 0005
CAD/Graphics4
2023 Continuous Sign Language Recognition with Correlation Network
abstract
Human body trajectories are a salient cue to identify actions in the video. Such body trajectories are mainly conveyed by hands and face across consecutive frames in sign language. However, current methods in continuous sign language recognition (CSLR) usually process frames independently, thus failing to capture cross-frame trajectories to effectively identify a sign. To handle this limitation, we propose correlation network (CorrNet) to explicitly capture and leverage body trajectories across frames to identify signs. In specific, a correlation module is first proposed to dynamically compute correlation maps between the current frame and adjacent frames to identify trajectories of all spatial patches. An identification module is then presented to dynamically emphasize the body trajectories within these correlation maps. As a result, the generated features are able to gain an overview of local temporal movements to identify a sign. Thanks to its special attention on body trajectories, CorrNet achieves new state-of-the-art accuracy on four largescale datasets, i.e., PHOENIX14, PHOENIX14-T, CSL-Daily, and CSL. A comprehensive comparison with previous spatial-temporal reasoning methods verifies the effectiveness of CorrNet. Visualizations demonstrate the effects of CorrNet on emphasizing human body trajectories across adjacent frames.
Lianyu Hu 0003, Liqing Gao, Zekang Liu, Wei Feng 0005
CVPR4
2023 Combining the Silhouette and Skeleton Data for Gait Recognition
abstract
Gait recognition, a long-distance biometric technology, has aroused intense interest recently. Currently, the two dominant gait recognition works are appearance-based and model-based, which extract features from silhouettes and skeletons, respectively. However, appearance-based methods are greatly affected by clothes-changing and carrying conditions, while model-based methods are limited by the accuracy of pose estimation. To tackle this challenge, a simple yet effective two-branch network is proposed in this paper, which contains a CNN-based branch taking silhouettes as input and a GCN-based branch taking skeletons as input. In addition, for better gait representation in the GCN-based branch, we present a fully connected graph convolution operator to integrate multi-scale graph convolutions and alleviate the dependence on natural joint connections. Also, we deploy a multi-dimension attention module named STC-Att to learn spatial, temporal and channel-wise attention simultaneously. The experimental results on CASIA-B and OUMVLP show that our method achieves state-of-the-art performance in various conditions.
Likai Wang 0002, Rui-Ze Han, Wei Feng 0005
ICASSP3
2023 Leveraging Inpainting for Single-Image Shadow Removal
abstract
Fully-supervised shadow removal methods achieve the best restoration qualities on public datasets but still generate some shadow remnants. One of the reasons is the lack of large-scale shadow & shadow-free image pairs. Unsupervised methods can alleviate the issue but their restoration qualities are much lower than those of fully-supervised methods. In this work, we find that pretraining shadow removal networks on the image inpainting dataset can reduce the shadow remnants significantly: a naive encoder-decoder network gets competitive restoration quality w.r.t. the state-of-the-art methods via only 10% shadow & shadow-free image pairs. After analyzing networks with/without inpainting pretraining via the information stored in the weight (IIW), we find that inpainting pretraining improves restoration quality in non-shadow regions and enhances the generalization ability of networks significantly. Additionally, shadow removal fine-tuning enables networks to fill in the details of shadow regions. Inspired by these observations we formulate shadow removal as an adaptive fusion task that takes advantage of both shadow removal and image inpainting. Specifically, we develop an adaptive fusion network consisting of two encoders, an adaptive fusion block, and a decoder. The two encoders are responsible for extracting the features from the shadow image and the shadow-masked image respectively. The adaptive fusion block is responsible for combining these features in an adaptive manner. Finally, the decoder converts the adaptive fused features to the desired shadow-free result. The extensive experiments show that our method empowered with inpainting outperforms all state-of-the-art methods. We have realized codes and models in https://github.com/tsingqguo/inpaint4shadow
Qing Guo 0005, Rabab Abdelfattah, Di Lin 0002, Wei Feng 0005, Ivor W. Tsang, Song Wang 0002
ICCV5
2023 Measuring Asymmetric Gradient Discrepancy in Parallel Continual Learning
abstract
In Parallel Continual Learning (PCL), the parallel multiple tasks start and end training unpredictably, thus suffering from both training conflict and catastrophic forgetting issues. The two issues are raised because the gradients from parallel tasks differ in directions and magnitudes. Thus, in this paper, we formulate the PCL into a minimum distance optimization problem among gradients and propose an explicit Asymmetric Gradient Distance (AGD) to evaluate the gradient discrepancy in PCL. AGD considers both gradient magnitude ratios and directions, and has a tolerance when updating with a small gradient of inverse direction, which reduces the imbalanced influence of gradients on parallel task training. Moreover, we present a novel Maximum Discrepancy Optimization (MaxDO) strategy to minimize the maximum discrepancy among multiple gradients. Solving by MaxDO with AGD, parallel training reduces the influence of the training conflict and suppresses the catastrophic forgetting of finished tasks. Extensive experiments validate the effectiveness of our approach on three image recognition datasets in task-incremental and class-incremental PCL. Our code is available at https://github.com/fanlyu/maxdo.
Fan Lyu, Fanhua Shang, Wei Feng 0005
ICCV5
2023 AdaBrowse: Adaptive Video Browser for Efficient Continuous Sign Language Recognition
abstract
Raw videos have been proven to own considerable feature redundancy where in many cases only a portion of frames can already meet the requirements for accurate recognition. In this paper, we are interested in whether such redundancy can be effectively leveraged to facilitate efficient inference in continuous sign language recognition (CSLR). We propose a novel adaptive model (AdaBrowse) to dynamically select a most informative subsequence from input video sequences by modelling this problem as a sequential decision task. In specific, we first utilize a lightweight network to quickly scan input videos to extract coarse features. Then these features are fed into a policy network to intelligently select a subsequence to process. The corresponding subsequence is finally inferred by a normal CSLR model for sentence prediction. As only a portion of frames are processed in this procedure, the total computations can be considerably saved. Besides temporal redundancy, we are also interested in whether the inherent spatial redundancy can be seamlessly integrated together to achieve further efficiency, i.e., dynamically selecting a lowest input resolution for each sample, whose model is referred to as AdaBrowse+. Extensive experimental results on four large-scale CSLR datasets, i.e., PHOENIX14, PHOENIX14-T, CSL-Daily and CSL, demonstrate the effectiveness of AdaBrowse and AdaBrowse+ by achieving comparable accuracy with state-of-the-art methods with 1.44X throughput and 2.12X fewer FLOPs. Comparisons with other commonly-used 2D CNNs and adaptive efficient methods verify the effectiveness of AdaBrowse. Code is available at https://github.com/hulianyuyy/AdaBrowse.
Lianyu Hu 0003, Liqing Gao, Zekang Liu, Chi-Man Pun, Wei Feng 0005
ACM Multimedia5
2023 Improving the Transferability of Adversarial Examples with Arbitrary Style Transfer
abstract
Deep neural networks are vulnerable to adversarial examples crafted by applying human-imperceptible perturbations on clean inputs. Although many attack methods can achieve high success rates in the white-box setting, they also exhibit weak transferability in the black-box setting. Recently, various methods have been proposed to improve adversarial transferability, in which the input transformation is one of the most effective methods. In this work, we notice that existing input transformation-based works mainly adopt the transformed data in the same domain for augmentation. Inspired by domain generalization, we aim to further improve the transferability using the data augmented from different domains. Specifically, a style transfer network can alter the distribution of low-level visual features in an image while preserving semantic content for humans. Hence, we propose a novel attack method named Style Transfer Method (STM) that utilizes a proposed arbitrary style transfer network to transform the images into different domains. To avoid inconsistent semantic information of stylized images for the classification network, we fine-tune the style transfer network and mix up the generated images added by random noise with the original images to maintain semantic consistency and boost input diversity. Extensive experimental results on the ImageNet-compatible dataset show that our proposed method can significantly improve the adversarial transferability on either normally trained models or adversarially trained models than state-of-the-art input transformation-based attacks. Code is available at: https://github.com/Zhijin-Ge/STM.
Zhijin Ge, Fanhua Shang, Hongying Liu 0001, Yuanyuan Liu 0001, Wei Feng 0005, Xiaosen Wang
ACM Multimedia6
2023 Open Compound Domain Adaptation with Object Style Compensation for Semantic Segmentation
abstract
Many methods of semantic image segmentation have borrowed the success of open compound domain adaptation. They minimize the style gap between the images of source and target domains, more easily predicting the accurate pseudo annotations for target domain's images that train segmentation network. The existing methods globally adapt the scene style of the images, whereas the object styles of different categories or instances are adapted improperly. This paper proposes the Object Style Compensation, where we construct the Object-Level Discrepancy Memory with multiple sets of discrepancy features. The discrepancy features in a set capture the style changes of the same category's object instances adapted from target to source domains. We learn the discrepancy features from the images of source and target domains, storing the discrepancy features in memory. With this memory, we select appropriate discrepancy features for compensating the style information of the object instances of various categories, adapting the object styles to a unified style of source domain. Our method enables a more accurate computation of the pseudo annotations for target domain's images, thus yielding state-of-the-art results on different datasets.
Tingliang Feng, Xueyang Liu, Wei Feng 0005, Di Lin 0002
NeurIPS4
2023 Global-local contrastive multiview representation learning for skeleton-based action recognition
Cunling Bian, Wei Feng 0005, Song Wang 0002
Comput. Vis. Image Underst.2
2023 Skeleton-based action recognition with local dynamic spatial-temporal aggregation
Lianyu Hu 0003, Shenglan Liu 0001, Wei Feng 0005
Expert Syst. Appl.3
2023 Phase-based fine-grained change detection
Xuzhi Wang, Di Lin 0002, Wei Feng 0005
Expert Syst. Appl.4
2023 Relating View Directions of Complementary-View Mobile Cameras via the Human Shadow
Rui-Ze Han, Yiyang Gan, Likai Wang 0002, Nan Li 0048, Wei Feng 0005, Song Wang 0002
Int. J. Comput. Vis.5
2023 TAGNet: Learning Configurable Context Pathways for Semantic Segmentation
abstract
State-of-the-art semantic segmentation methods capture the relationship between pixels to facilitate contextual information exchange. Advanced methods utilize fixed pathways for context exchange, lacking the flexibility to harness the most relevant context for each pixel. In this paper, we present Configurable Context Pathways (CCPs), a novel model for establishing pathways for augmenting contextual information. In contrast to previous pathway models, CCPs are learned, leveraging configurable regions to form information flows between pairs of pixels. We propose TAGNet to adaptively configure the regions, which span over the entire image space, driven by the relationships between the remote pixels. Subsequently, the information flows along the pathways are updated gradually by the information provided by sequences of configurable regions, forming more powerful contextual information. We extensively evaluate the traveling, adaption, and gathering (TAG) stages of our network on the public benchmarks, demonstrating that all of the stages successfully improve the segmentation accuracy and help to surpass the state-of-the-art results. The code package is available at: https://github.com/dilincv/TAGNet.
Di Lin 0002, Dingguo Shen, Yuanfeng Ji, Siting Shen, Mingrui Xie, Wei Feng 0005, Hui Huang 0004
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 Multi-semantic hypergraph neural network for effective few-shot learning
Hao Chen 0011, Fuyuan Hu, Fan Lyu, Liuqing Zhao, Kaizhu Huang, Wei Feng 0005, Zhenping Xia
Pattern Recognit.7
2023 Structure-Informed Shadow Removal Networks
abstract
Existing deep learning-based shadow removal methods still produce images with shadow remnants. These shadow remnants typically exist in homogeneous regions with low-intensity values, making them untraceable in the existing image-to-image mapping paradigm. We observe that shadows mainly degrade images at the image-structure level (in which humans perceive object shapes and continuous colors). Hence, in this paper, we propose to remove shadows at the image structure level. Based on this idea, we propose a novel structure-informed shadow removal network (StructNet) to leverage the image-structure information to address the shadow remnant problem. Specifically, StructNet first reconstructs the structure information of the input image without shadows and then uses the restored shadow-free structure prior to guiding the image-level shadow removal. StructNet contains two main novel modules: 1) a mask-guided shadow-free extraction (MSFE) module to extract image structural features in a non-shadow-to-shadow directional manner; and 2) a multi-scale feature & residual aggregation (MFRA) module to leverage the shadow-free structure information to regularize feature consistency. In addition, we also propose to extend StructNet to exploit multi-level structure information (MStructNet), to further boost the shadow removal performance with minimum computational overheads. Extensive experiments on three shadow removal benchmarks demonstrate that our method outperforms existing shadow removal methods, and our StructNet can be integrated with existing methods to improve them further.
Yuhao Liu 0001, Qing Guo 0005, Lan Fu, Zhanghan Ke, Ke Xu 0010, Wei Feng 0005, Ivor W. Tsang, Rynson W. H. Lau
IEEE Trans. Image Process.6
2023 Difference-guided multi-scale spatial-temporal representation for sign language recognition
Liqing Gao, Lianyu Hu 0003, Fan Lyu, Lei Zhu 0003, Chi-Man Pun, Wei Feng 0005
Vis. Comput.7
2022 Can You Spot the Chameleon? Adversarially Camouflaging Images from Co-Salient Object Detection
abstract
Co-salient object detection (CoSOD) has recently achieved significant progress and played a key role in retrieval-related tasks. However, it inevitably poses an entirely new safety and security issue, i.e., highly personal and sensitive content can potentially be extracting by powerful CoSOD methods. In this paper, we address this problem from the perspective of adversarial attacks and identify a novel task: adversarial co-saliency attack. Specially, given an image selected from a group of images containing some common and salient objects, we aim to generate an adversarial version that can mislead CoSOD methods to predict incorrect co-salient regions. Note that, compared with general white-box adversarial attacks for classification, this new task faces two additional challenges: (1) low success rate due to the diverse appearance of images in the group; (2) low transferability across CoSOD methods due to the considerable difference between CoSOD pipelines. To address these challenges, we propose the very first blackbox joint adversarial exposure and noise attack (Jadena), where we jointly and locally tune the exposure and additive perturbations of the image according to a newly designed high-feature-level contrast-sensitive loss function. Our method, without any information on the state-of-the-art CoSOD methods, leads to significant performance degradation on various co-saliency detection datasets and makes the co-salient objects undetectable. This can have strong practical benefits in properly securing the large number of personal photos currently shared on the Internet. Moreover, our method is potential to be utilized as a metric for evaluating the robustness of CoSOD methods.
Ruijun Gao, Qing Guo 0005, Felix Juefei-Xu, Hongkai Yu, Huazhu Fu, Wei Feng 0005, Yang Liu 0003, Song Wang 0002
CVPR6
2022 Connecting the Complementary-view Videos: Joint Camera Identification and Subject Association
abstract
We attempt to connect the data from complementary views, i.e., top view from drone-mounted cameras in the air, and side view from wearable cameras on the ground. Collaborative analysis of such complementary-view data can facilitate to build the air-ground cooperative visual system for various kinds of applications. This is a very challenging problem due to the large view difference between top and side views. In this paper, we develop a new approach that can simultaneously handle three tasks: i) localizing the side-view camera in the top view; ii) estimating the view direction of the side-view camera; iii) detecting and associating the same subjects on the ground across the complementary views. Our main idea is to explore the spatial position layout of the subjects in two views. In particular, we propose a spatial-aware position representation method to embed the spatial-position distribution of the subjects in different views. We further design a cross-view video collaboration framework composed of a camera identification module and a subject association module to simultaneously perform the above three tasks. We collect a new synthetic dataset consisting of top-view and side-view video sequence pairs for performance evaluation and the experimental results show the effectiveness of the proposed method.
Rui-Ze Han, Yiyang Gan, Wei Feng 0005, Song Wang 0002
CVPR5
2022 MISF: Multi-level Interactive Siamese Filtering for High-Fidelity Image Inpainting
abstract
Although achieving significant progress, existing deep generative inpainting methods still show low generalization across different scenes. As a result, the generated images usually contain artifacts or the filled pixels differ greatly from the ground truth, making them far from real-world applications. Image-level predictive filtering is a widely used restoration technique by predicting suitable kernels adaptively according to different input scenes. Inspired by this inherent advantage, we explore the possibility of addressing image inpainting as a filtering task. To this end, we first study the advantages and challenges of the image-level predictive filtering for inpainting: the method can preserve local structures and avoid artifacts but fails to fill large missing areas. Then, we propose the semantic filtering by conducting filtering on deep feature level, which fills the missing semantic information but fails to recover the details. To address the issues while adopting the respective advantages, we propose a novel filtering technique, i.e., Multi-level Interactive Siamese Filtering (MISF) containing two branches: kernel prediction branch (KPB) and semantic & image filtering branch (SIFB). These two branches are interactively linked: SIFB provides multi-level features for KPB while KPB predicts dynamic kernels for SIFB. As a result, the final method takes the advantage of effective semantic & image-level filling for high-fidelity inpainting. Moreover, we discuss the relationship between MISF and the naive encoder-decoder-based inpainting, inferring that MISF provides novel dynamic convolutional operations to enhance the high generalization capability across scenes. We validate our method on three challenging datasets, i.e., Dunhuang, Places2, and CelebA. Our method outperforms state-of-the-art baselines on four metrics, i.e.,$L_{1}$, PSNR, SSIM, and LPIPS.
Qing Guo 0005, Di Lin 0002, Ping Li 0016, Wei Feng 0005, Song Wang 0002
CVPR5
2022 Panoramic Human Activity Recognition
Rui-Ze Han, Haomin Yan, Song Wang 0002, Wei Feng 0005
ECCV (4)5
2022 Temporal Lift Pooling for Continuous Sign Language Recognition
Lianyu Hu 0003, Liqing Gao, Zekang Liu, Wei Feng 0005
ECCV (35)4
2022 Self-supervised Social Relation Representation for Human Group Detection
Rui-Ze Han, Haomin Yan, Zekun Qian, Wei Feng 0005, Song Wang 0002
ECCV (35)5
2022 Rethinking Video Rain Streak Removal: A New Synthesis Model and a Deraining Network with Video Rain Prior
Lei Zhu 0003, Huazhu Fu, Harry Qin, Carola-Bibiane Schönlieb, Wei Feng 0005, Song Wang 0002
ECCV (19)6
2022 AGCN: Augmented Graph Convolutional Network for Lifelong Multi-Label Image Recognition
abstract
The Lifelong Multi-Label (LML) image recognition builds an online class-incremental classifier in a sequential multilabel image recognition data stream. However, training on the data with different Partial Labels may result in more serious Catastrophic Forgetting in old classes. To solve the problem, the study proposes an Augmented Graph Convolutional Network (AGCN)to build an Augmented Correlation Matrix (ACM) across the sequential partial-label tasks and sustain the catastrophic forgetting. First, in ACM, the intra-task relations derive from the hard label statistics, while the inter-task relations further leverage the soft labels from a stored expert network. Then, based on the ACM, AGCN captures label dependencies with dynamic augmented structure and yields effective class representations. Our method is evaluated on two multi-label image benchmarks and the results show that the proposed method is effective for LML image recognition.
Kaile Du, Fan Lyu, Fuyuan Hu, Wei Feng 0005, Fenglei Xu, Qiming Fu 0001
ICME5
2022 Self-Supervised Representation Learning for Skeleton-Based Group Activity Recognition
abstract
Group activity recognition (GAR) is a challenging task for discerning the behavior of a group of actors. This paper aims at learning discriminative representation for GAR in a self-supervised manner based on human skeletons. As modeling relations between actors lie at the center of GAR, we propose a valid self-supervised learning pretext task with a matching framework, where a representation model is driven to identify subgroups in a synthetic group based on actors' skeleton sequences. For backbone networks, while spatial-temporal graph convolution networks have dominated the skeleton-based action recognition, they under-explore the group relevant interactions among actors. To address this issue, we come up with a novel plug-in Actor-Association Graph Convolution Module (AAGCM) based on inductive graph convolution, which can be integrated into many common backbones. It can not only model the interactions at different levels but also adapt to variable group sizes. The effectiveness of our approaches is demonstrated by extensive experiments on three benchmark datasets: Volleyball, Collective Activity, and Mutual NTU.
Cunling Bian, Wei Feng 0005, Song Wang 0002
ACM Multimedia2
2022 Video Instance Lane Detection via Deep Temporal and Geometry Consistency Constraints
abstract
Video instance lane detection is one of the most important tasks in autonomous driving.Due to the very sparse region and weak context in lane annotations, accurately detecting instance-level lanes in real-world traffic scenarios is challenging, especially for scenes with occlusion, bad weather conditions, dim or dazzling lights.Current methods mainly address this problem by integrating features of adjacent video frames to simply encourage temporal constancy for image-level lane detectors. However, most of them ignore lane shape constraint of adjacent frames and geometry consistency of individual lanes, thereby harming the performance of video instance lane detection. In this paper, we propose TGC-Net via temporal and geometry consistency constraints for reliable video instance lane detection. Specifically, we devise a temporal recurrent feature-shift aggregation module (T-RESA) to learn spatio-temporal lane features along horizontal, vertical, and temporal directions of the feature tensor. We further impose temporal consistency constraint by encouraging spatial distribution consistency among the lane features of adjacent frames. Besides, we devise two effective geometry constraints to ensure the integrity and continuity of lane predictions by leveraging pairwise point affinity loss and vanishing point guided geometric context, respectively. Extensive experiments on public benchmark dataset show that our TGC-Net quantitatively and qualitatively outperforms state-of-the-art video instance lane detectors and video object segmentation competitors. Our code and our results have been released at https://github.com/wmq12345/TGC-Net.
Yujun Zhang 0002, Wei Feng 0005, Lei Zhu 0003, Song Wang 0002
ACM Multimedia3
2022 Self-Supervised Human Pose based Multi-Camera Video Synchronization
abstract
Multi-view video collaborative analysis is an important task and has many applications in multimedia community. However, it always requires the given multiple videos to be temporally synchronized. Existing methods commonly synchronize the videos by the wired communication, which may hinder the practical application in real world, especially for moving cameras. In this paper, we focus on the human-centric video analysis and propose a self-supervised framework for the automatic multi-camera video synchronization. Specifically, we develop SeSyn-Net with the 2D human pose as input for feature embedding and design a series of self-supervised losses to effectively extract the view-invariant but time-discriminative representation for video synchronization. We also build two new datasets for the performance evaluation. Extensive experimental results verify the effectiveness of our method, which achieves the superior performance compared to both the classical and state-of-the-art methods.
Liqiang Yin, Rui-Ze Han, Wei Feng 0005, Song Wang 0002
ACM Multimedia3
2022 Cross-Image Context for Single Image Inpainting
abstract
Visual context is of crucial importance for image inpainting. The contextual information captures the appearance and semantic correlation between the image regions, helping to propagate the information of the complete regions for reasoning the content of the corrupted regions. Many inpainting methods compute the visual context based on the regions within the single image. In this paper, we propose the Cross-Image Context Memory (CICM) for learning and using the cross-image context to recover the corrupted regions. CICM consists of multiple sets of the cross-image representations learned from the image regions with different visual patterns. The regional representations are learned across different images, thus providing richer context that benefit the inpainting task. The experimental results demonstrate the effectiveness and generalization of CICM, which achieves state-of-the-art performances on various datasets for single image inpainting.
Tingliang Feng, Wei Feng 0005, Di Lin 0002
NeurIPS2
2022 Exploring Example Influence in Continual Learning
abstract
Continual Learning (CL) sequentially learns new tasks like human beings, with the goal to achieve better Stability (S, remembering past tasks) and Plasticity (P, adapting to new tasks). Due to the fact that past training data is not available, it is valuable to explore the influence difference on S and P among training examples, which may improve the learning pattern towards better SP. Inspired by Influence Function (IF), we first study example influence via adding perturbation to example weight and computing the influence derivation. To avoid the storage and calculation burden of Hessian inverse in neural networks, we propose a simple yet effective MetaSP algorithm to simulate the two key steps in the computation of IF and obtain the S- and P-aware example influence. Moreover, we propose to fuse two kinds of example influence by solving a dual-objective optimization problem, and obtain a fused influence towards SP Pareto optimality. The fused influence can be used to control the update of model and optimize the storage of rehearsal. Empirical results show that our algorithm significantly outperforms state-of-the-art methods on both task- and class-incremental benchmark CL datasets.
Fan Lyu, Fanhua Shang, Wei Feng 0005
NeurIPS4
2022 Harnessing Multi-Semantic Hypergraph for Few-Shot Learning
Hao Chen 0011, Zhenping Xia, Fan Lyu, Liuqing Zhao, Kaizhu Huang, Wei Feng 0005, Fuyuan Hu
PRCV (1)7
2022 Fast and robust active camera relocalization in the wild for fine-grained change detection
Qian Zhang 0051, Wei Feng 0005, Yi-Bo Shi, Di Lin 0002
Neurocomputing2
2022 Multiple Human Association and Tracking From Egocentric and Complementary Top Views
abstract
Crowded scene surveillance can significantly benefit from combining egocentric-view and its complementary top-view cameras. A typical setting is an egocentric-view camera, e.g., a wearable camera on the ground capturing rich local details, and a top-view camera, e.g., a drone-mounted one from high altitude providing a global picture of the scene. To collaboratively analyze such complementary-view videos, an important task is to associate and track multiple people across views and over time, which is challenging and differs from classical human tracking, since we need to not only track multiple subjects in each video, but also identify the same subjects across the two complementary views. This paper formulates it as a constrained mixed integer programming problem, wherein a major challenge is how to effectively measure subjects similarity over time in each video and across two views. Although appearance and motion consistencies well apply to over-time association, they are not good at connecting two highly different complementary views. To this end, we present a spatial distribution based approach to reliable cross-view subject association. We also build a dataset to benchmark this new challenging task. Extensive experiments verify the effectiveness of our method.
Rui-Ze Han, Wei Feng 0005, Yujun Zhang 0002, Jiewen Zhao, Song Wang 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Laplacian Smoothing Stochastic ADMMs With Differential Privacy Guarantees
abstract
Many machine learning tasks such as structured sparse coding and multi-task learning can be converted into an equality constrained optimization problem. The stochastic alternating direction method of multipliers (SADMM) is a popular algorithm to solve such large-scale problems, and has been successfully used in many real-world applications. However, existing SADMMs fail to take into consideration an important issue in their designs, i.e., protecting sensitive information. To address this challenging issue, this paper proposes a novel differential privacy stochastic ADMM framework for solving equality constrained machine learning problems. In particular, to further lift the utility in privacy-preserving equality constrained optimization, a Laplacian smoothing operation is also introduced into our differential privacy ADMM framework, and it can smooth out the Gaussian noise used in the Gaussian mechanism. Then we propose an efficient differentially private variance reduced stochastic ADMM (DP-VRADMM) algorithm with Laplacian smoothing for both strongly convex and general convex objectives. As a by-product, we also present a new differentially private stochastic ADMM algorithm with DP guarantees. In theory, we provide both private guarantees and utility guarantees for the proposed algorithms, which show that Laplacian smoothing can improve the utility bounds of our algorithms. Experimental results on real-world datasets verify our theoretical results and the effectiveness of our algorithms.
Yuanyuan Liu 0001, Jiacheng Geng, Fanhua Shang, Weixin An, Hongying Liu 0001, Wei Feng 0005
IEEE Trans. Inf. Forensics Secur.7
2022 Multi-View Multi-Human Association With Deep Assignment Network
abstract
Identifying the same persons across different views plays an important role in many vision applications. In this paper, we study this important problem, denoted as Multi-view Multi-Human Association (MvMHA), on multi-view images that are taken by different cameras at the same time. Different from previous works on human association across two views, this paper is focused on more general and challenging scenarios of more than two views, and none of these views are fixed or priorly known. In addition, each involved person may be present in all the views or only a subset of views, which are also not priorly known. We develop a new end-to-end deep-network based framework to address this problem. First, we use an appearance-based deep network to extract the feature of each detected subject on each image. We then compute pairwise-similarity scores between all the detected subjects and construct a comprehensive affinity matrix. Finally, we propose a Deep Assignment Network (DAN) to transform the affinity matrix into an assignment matrix, which provides a binary assignment result for MvMHA. We build both a synthetic dataset and a real image dataset to verify the effectiveness of the proposed method. We also test the trained network on other three public datasets, resulting in very good cross-domain performance.
Rui-Ze Han, Haomin Yan, Wei Feng 0005, Song Wang 0002
IEEE Trans. Image Process.4
2022 Pasadena: Perceptually Aware and Stealthy Adversarial Denoise Attack
abstract
Image denoising can remove natural noise that widely exists in images captured by multimedia devices due to low-quality imaging sensors, unstable image transmission processes, or low light conditions. Recent works also find that image denoising benefits the high-level vision tasks,e.g., image classification. In this work, we try to challenge this common sense and explore a totally new problem,i.e., whether the image denoising can be given the capability of fooling the state-of-the-art deep neural networks (DNNs) while enhancing the image quality. To this end, we initiate the very first attempt to study this problem from the perspective of adversarial attack and propose theadversarial denoise attack. More specifically, our main contributions are three-fold:First, we identify a new task that stealthily embeds attacks inside the image denoising module widely deployed in multimedia devices as an image post-processing operation to simultaneously enhance the visual image quality and fool DNNs.Second, we formulate this new task as a kernel prediction problem for image filtering and propose theadversarial-denoising kernel predictionthat can produce adversarial-noiseless kernels for effective denoising and adversarial attacking simultaneously.Third, we implement an adaptiveperceptual region localizationto identify semantic-related vulnerability regions with which the attack can be more effective while not doing too much harm to the denoising. We name the proposed method asPasadena(Perceptually Aware and Stealthy Adversarial DENoise Attack) and validate our method on the NeurIPS’17 adversarial competition dataset, CVPR2021-AIC-VI: unrestricted adversarial attacks on ImageNet, and Tiny-ImageNet-C dataset. The comprehensive evaluation and analysis demonstrate that our method not only realizes denoising but also achieves a significantly higher success rate and transferability over state-of-the-art attacks.
Yupeng Cheng, Qing Guo 0005, Felix Juefei-Xu, Shangwei Lin 0001, Wei Feng 0005, Weisi Lin, Yang Liu 0003
IEEE Trans. Multim.5
2022 Fine-grained scale space learning for single image super-resolution
Fan Lyu, Wei Feng 0005
Vis. Comput.4
2021 EfficientDeRain: Learning Pixel-wise Dilation Filtering for High-Efficiency Single-Image Deraining
abstract
Single-image deraining is rather challenging due to the unknown rain model. Existing methods often make specific assumptions of the rain model, which can hardly cover many diverse circumstances in the real world, compelling them to employ complex optimization or progressive refinement. This, however, significantly affects these methods' efficiency and effectiveness for many efficiency-critical applications. To fill this gap, in this paper, we regard the single-image deraining as a general image-enhancing problem and originally propose a model-free deraining method, i.e., EfficientDeRain, which is able to process a rainy image within 10 ms (i.e., around 6 ms on average), over 80 times faster than the state-of-the-art method (i.e., RCDNet), while achieving similar de-rain effects. We first propose novel pixel-wise dilation filtering. In particular, a rainy image is filtered with the pixel-wise kernels estimated from a kernel prediction network, by which suitable multi-scale kernels for each pixel can be efficiently predicted. Then, to eliminate the gap between synthetic and real data, we further propose an effective data augmentation method (i.e., RainMix) that helps to train the network for handling real rainy images. We perform a comprehensive evaluation on both synthetic and real-world rainy datasets to demonstrate the effectiveness and efficiency of our method. We release the model and code in https://github.com/tsingqguo/efficientderain.git.
Qing Guo 0005, Jingyang Sun, Felix Juefei-Xu, Lei Ma 0003, Xiaofei Xie, Wei Feng 0005, Yang Liu 0003, Jianjun Zhao 0001
AAAI6
2021 Multi-Domain Multi-Task Rehearsal for Lifelong Learning
abstract
Rehearsal, seeking to remind the model by storing old knowledge in lifelong learning, is one of the most effective ways to mitigate catastrophic forgetting, i.e., biased forgetting of previous knowledge when moving to new tasks. However, the old tasks of the most previous rehearsal-based methods suffer from the unpredictable domain shift when training the new task. This is because these methods always ignore two significant factors. First, the Data Imbalance between the new task and old tasks that makes the domain of old tasks prone to shift. Second, the Task Isolation among all tasks will make the domain shift toward unpredictable directions; To address the unpredictable domain shift, in this paper, we propose Multi-Domain Multi-Task (MDMT) rehearsal to train the old tasks and new task parallelly and equally to break the isolation among tasks. Specifically, a two-level angular margin loss is proposed to encourage the intra-class/task compactness and inter-class/task discrepancy, which keeps the model from domain chaos. In addition, to further address domain shift of the old tasks, we propose an optional episodic distillation loss on the memory to anchor the knowledge for each old task. Experiments on benchmark datasets validate the proposed approach can effectively mitigate the unpredictable domain shift.
Fan Lyu, Wei Feng 0005, Zihan Ye, Fuyuan Hu, Song Wang 0002
AAAI3
2021 Auto-Exposure Fusion for Single-Image Shadow Removal
abstract
Shadow removal is still a challenging task due to its inherent background-dependent1and spatial-variant properties, leading to unknown and diverse shadow patterns. Even powerful deep neural networks could hardly recover traceless shadow-removed background. This paper proposes a new solution for this task by formulating it as an exposure fusion problem to address the challenges. Intuitively, we first estimate multiple over-exposure images w.r.t. the input image to let the shadow regions in these images have the same color with shadow-free areas in the input image. Then, we fuse the original input with the over-exposure images to generate the final shadow-free counterpart. Nevertheless, the spatial-variant property of the shadow requires the fusion to be sufficiently ‘smart’, that is, it should automatically select proper over-exposure pixels from different images to make the final output natural. To address this challenge, we propose the shadow-aware FusionNet that takes the shadow image as input to generate fusion weight maps across all the over-exposure images. Moreover, we propose the boundary-aware RefineNet to eliminate the remaining shadow trace further. We conduct extensive experiments on the ISTD, ISTD+, and SRD datasets to validate our method’s effectiveness and show better performance in shadow regions and comparable performance in non-shadow regions over the state-of-the-art methods. We release the code in https://github.com/tsingqguo/exposure-fusion-shadow-removal.
Lan Fu, Changqing Zhou, Qing Guo 0005, Felix Juefei-Xu, Hongkai Yu, Wei Feng 0005, Yang Liu 0003, Song Wang 0002
CVPR6
2021 Multiple Human Tracking in Non-Specific Coverage with Wearable Cameras
abstract
Compared to fixed cameras, wearable cameras have time-varying non-specific view coverage and can be used to alternately observe people at different sites by varying the camera views. However, such view change of wearable cameras may introduce intervals of transitional frames without useful information, which brings new challenge for the important multiple object tracking (MOT) task – existing MOT methods can not handle well frequent disappearing/reappearing targets in the field of view, especially in the presence of informationless transitional sequences of frames. To address this problem, in this paper we propose a Markov Decision Process with jump state (JMDP) to model the target’s lifetime in tracking, and use optical flow of the camera motion and the statistical information of the targets to model the camera state transition. We further develop a frame-level classification algorithm to locate the transitional sequence. By combining all of them, we formulate the proposed non-specific-coverage MOT problem as a joint state transition problem, which can be solved by the state transfer mechanism of the targets and the camera. We collect a new dataset for performance evaluation and the experimental results show the effectiveness of the proposed method.
Sibo Wang 0008, Rui-Ze Han, Wei Feng 0005, Song Wang 0002
ICASSP3
2021 VIL-100: A New Dataset and A Baseline Model for Video Instance Lane Detection
abstract
Lane detection plays a key role in autonomous driving. While car cameras always take streaming videos on the way, current lane detection works mainly focus on individual images (frames) by ignoring dynamics along the video. In this work, we collect a new video instance lane detection (VIL-100) dataset, which contains 100 videos with in total 10,000 frames, acquired from different real traffic scenarios. All the frames in each video are manually annotated to a high-quality instance-level lane annotation, and a set of frame-level and video-level metrics are included for quantitative performance evaluation. Moreover, we propose a new baseline model, named multi-level memory aggregation network (MMA-Net), for video instance lane detection. In our approach, the representation of current frame is enhanced by attentively aggregating both local and global memory features from other frames. Experiments on the new collected dataset show that the proposed MMA-Net outperforms state-of-the-art lane detection methods and video object segmentation methods. We release our dataset and code at https://github.com/yujun0-0/MMA-Net.
Yujun Zhang 0002, Lei Zhu 0003, Wei Feng 0005, Huazhu Fu, Qingxia Li, Song Wang 0002
ICCV3
2021 DGD-net: Local Descriptor Guided Keypoint Detection Network
abstract
In recent years, learning-based feature detection network has greatly improved the performance of keypoints matching. However, existing approaches have not fully utilized the representational ability of learned descriptors for feature detection. We propose a novel keypoint detection and description approach to make more use of the reliability of descriptors in matching on keypoint detection. We utilize the descriptor training loss to construct a guided score, which depicts the matching reliability of the descriptors, to supervise the detectors learning. We also propose a descriptor training loss to adapt to our detector training. In order to improve localization accuracy from the low-resolution feature map, we propose a method based on the idea of backtracking. Extensive experiments show the effectiveness of the proposed approach.
Fei-Peng Tian, Wei Feng 0005
ICME4
2021 Self-supervised Multi-view Multi-Human Association and Tracking
abstract
Multi-view Multi-human association and tracking (MvMHAT) aims to track a group of people over time in each view, as well as to identify the same person across different views at the same time. This is a relatively new problem but is very important for multi-person scene video surveillance. Different from previous multiple object tracking (MOT) and multi-target multi-camera tracking (MTMCT) tasks, which only consider the over-time human association, MvMHAT requires to jointly achieve both cross-view and over-time data association. In this paper, we model this problem with a self-supervised learning framework and leverage an end-to-end network to tackle it. Specifically, we propose a spatial-temporal association network with two designed self-supervised learning losses, including a symmetric-similarity loss and a transitive-similarity loss, at each time to associate the multiple humans over time and across views. Besides, to promote the research on MvMHAT, we build a new large-scale benchmark for the training and testing of different algorithms. Extensive experiments on the proposed benchmark verify the effectiveness of our method. We have released the benchmark and code to the public.
Yiyang Gan, Rui-Ze Han, Liqiang Yin, Wei Feng 0005, Song Wang 0002
ACM Multimedia4
2021 From Synthetic to Real: Image Dehazing Collaborating with Unlabeled Real Data
abstract
Single image dehazing is a challenging task, for which the domain shift between synthetic training data and real-world testing images usually leads to degradation of existing methods. To address this issue, we propose a novel image dehazing framework collaborating with unlabeled real data. First, we develop a disentangled image dehazing network (DID-Net), which disentangles the feature representations into three component maps, i.e. the latent haze-free image, the transmission map, and the global atmospheric light estimate, respecting the physical model of a haze process. Our DID-Net predicts the three component maps by progressively integrating features across scales, and refines each map by passing an independent refinement network. Then a disentangled-consistency mean-teacher network (DMT-Net) is employed to collaborate unlabeled real data for boosting single image dehazing. Specifically, we encourage the coarse predictions and refinements of each disentangled component to be consistent between the student and teacher networks by using a consistency loss on unlabeled real data. We make comparison with 13 state-of-the-art dehazing methods on a new collected dataset (Haze4K) and two widely-used dehazing datasets (i.e., SOTS and HazeRD), as well as on real-world hazy images. Experimental results demonstrate that our method has obvious quantitative and qualitative improvements over the existing methods.
Lei Zhu 0003, Shunda Pei, Huazhu Fu, Harry Qin, Qing Zhang 0006, Wei Feng 0005
ACM Multimedia8
2021 Graph Matching Based Robust Line Segment Correspondence for Active Camera Relocalization
Mengyu Pan, Fei-Peng Tian, Wei Feng 0005
PRCV (2)4
2021 RNN-Transducer based Chinese Sign Language Recognition
Liqing Gao, Zekang Liu, Wei Feng 0005
Neurocomputing6
2021 Structural Knowledge Distillation for Efficient Skeleton-Based Action Recognition
abstract
Skeleton data have been extensively used for action recognition since they can robustly accommodate dynamic circumstances and complex backgrounds. To guarantee the action-recognition performance, we prefer to use advanced and time-consuming algorithms to get more accurate and complete skeletons from the scene. However, this may not be acceptable in time- and resource-stringent applications. In this paper, we explore the feasibility of using low-quality skeletons, which can be quickly and easily estimated from the scene, for action recognition. While the use of low-quality skeletons will surely lead to degraded action-recognition accuracy, in this paper we propose a structural knowledge distillation scheme to minimize this accuracy degradations and improve recognition model's robustness to uncontrollable skeleton corruptions. More specifically, a teacher which observes high-quality skeletons obtained from a scene is used to help train a student which only sees low-quality skeletons generated from the same scene. At inference time, only the student network is deployed for processing low-quality skeletons. In the proposed network, a graph matching loss is proposed to distill the graph structural knowledge at an intermediate representation level. We also propose a new gradient revision strategy to seek a balance between mimicking the teacher model and directly improving the student model's accuracy. Experiments are conducted on Kenetics400, NTU RGB+D and Penn action recognition datasets and the comparison results demonstrate the effectiveness of our scheme.
Cunling Bian, Wei Feng 0005, Song Wang 0002
IEEE Trans. Image Process.2
2021 Exploring the Effects of Blur and Deblurring to Visual Object Tracking
abstract
The existence of motion blur can inevitably influence the performance of visual object tracking. However, in contrast to the rapid development of visual trackers, the quantitative effects of increasing levels of motion blur on the performance of visual trackers still remain unstudied. Meanwhile, although image-deblurring can produce visually sharp videos for pleasant visual perception, it is also unknown whether visual object tracking can benefit from image deblurring or not. In this paper, we present a Blurred Video Tracking (BVT) benchmark to address these two problems, which contains a large variety of videos with different levels of motion blurs, as well as ground-truth tracking results. To explore the effects of blur and deblurring to visual object tracking, we extensively evaluate 25 trackers on the proposed BVT benchmark and obtain several new interesting findings. Specifically, we find that light motion blur may improve the accuracy of many trackers, but heavy blur usually hurts the tracking performance. We also observe that image deblurring is helpful to improve tracking accuracy on heavily-blurred videos but hurts the performance of lightly-blurred videos. According to these observations, we then propose a new general GAN-based scheme to improve a tracker's robustness to motion blur. In this scheme, a fine-tuned discriminator can effectively serve as an adaptive blur assessor to enable selective frames deblurring during the tracking process. We use this scheme to successfully improve the accuracy of 6 state-of-the-art trackers on motion-blurred videos.
Qing Guo 0005, Wei Feng 0005, Ruijun Gao, Yang Liu 0003, Song Wang 0002
IEEE Trans. Image Process.2
2021 Learning Guided Convolutional Network for Depth Completion
abstract
Dense depth perception is critical for autonomous driving and other robotics applications. However, modern LiDAR sensors only provide sparse depth measurement. It is thus necessary to complete the sparse LiDAR data, where a synchronized guidance RGB image is often used to facilitate this completion. Many neural networks have been designed for this task. However, they often naïvely fuse the LiDAR data and RGB image information by performing feature concatenation or element-wise addition. Inspired by the guided image filtering, we design a novel guided network to predict kernel weights from the guidance image. These predicted kernels are then applied to extract the depth image features. In this way, our network generates content-dependent and spatially-variant kernels for multi-modal feature fusion. Dynamically generated spatially-variant kernels could lead to prohibitive GPU memory consumption and computation overhead. We further design a convolution factorization to reduce computation and memory consumption. The GPU memory reduction makes it possible for feature fusion to work in multi-stage scheme. We conduct comprehensive experiments to verify our method on real-world outdoor, indoor and synthetic datasets. Our method produces strong results. It outperforms state-of-the-art methods on the NYUv2 dataset and ranks 1st on the KITTI depth completion benchmark at the time of submission. It also presents strong generalization capability under different 3D point densities, various lighting and weather conditions as well as cross-dataset evaluations. The code will be released for reproduction.
Jie Tang 0015, Fei-Peng Tian, Wei Feng 0005, Jian Li 0003, Ping Tan 0002
IEEE Trans. Image Process.3
2020 Complementary-View Multiple Human Tracking
abstract
The global trajectories of targets on ground can be well captured from a top view in a high altitude, e.g., by a drone-mounted camera, while their local detailed appearances can be better recorded from horizontal views, e.g., by a helmet camera worn by a person. This paper studies a new problem of multiple human tracking from a pair of top- and horizontal-view videos taken at the same time. Our goal is to track the humans in both views and identify the same person across the two complementary views frame by frame, which is very challenging due to very large field of view difference. In this paper, we model the data similarity in each view using appearance and motion reasoning and across views using appearance and spatial reasoning. Combing them, we formulate the proposed multiple human tracking as a joint optimization problem, which can be solved by constrained integer programming. We collect a new dataset consisting of top- and horizontal-view video pairs for performance evaluation and the experimental results show the effectiveness of the proposed method.
Rui-Ze Han, Wei Feng 0005, Jiewen Zhao, Zicheng Niu, Yujun Zhang 0002, Song Wang 0002
AAAI2
2020 A Multi-Task Mean Teacher for Semi-Supervised Shadow Detection
abstract
Existing shadow detection methods suffer from an intrinsic limitation in relying on limited labeled datasets, and they may produce poor results in some complicated situations. To boost the shadow detection performance, this paper presents a multi-task mean teacher model for semi-supervised shadow detection by leveraging unlabeled data and exploring the learning of multiple information of shadows simultaneously. To be specific, we first build a multi-task baseline model to simultaneously detect shadow regions, shadow edges, and shadow count by leveraging their complementary information and assign this baseline model to the student and teacher network. After that, we encourage the predictions of the three tasks from the student and teacher networks to be consistent for computing a consistency loss on unlabeled data, which is then added to the supervised loss on the labeled data from the predictions of the multi-task baseline model. Experimental results on three widely-used benchmark datasets show that our method consistently outperforms all the compared state-of- the-art methods, which verifies that the proposed network can effectively leverage additional unlabeled data to boost the shadow detection performance.
Zhihao Chen 0004, Lei Zhu 0003, Song Wang 0002, Wei Feng 0005, Pheng-Ann Heng
CVPR5
2020 SPARK: Spatial-Aware Online Incremental Attack Against Visual Tracking
Qing Guo 0005, Xiaofei Xie, Felix Juefei-Xu, Lei Ma 0003, Zhongguo Li, Wanli Xue, Wei Feng 0005, Yang Liu 0003
ECCV (25)7
2020 Key Action and Joint CTC-Attention based Sign Language Recognition
abstract
Sign Language Recognition (SLR) translates sign language video into natural language. In practice, sign language video, owning a large number of redundant frames, is necessary to be selected the essential. However, unlike common video that describes actions, sign language video is characterized as continuous and dense action sequence, which is difficult to capture key actions corresponding to meaningful sentence. In this paper, we propose to hierarchically search key actions by a pyramid BiLSTM. Specifically, we first construct three BiL-STMs to produce temporal relationships among input video sequence. Then, we associate these BiLSTMs by searching the salient responses in two groups of fixed-scale sliding window and capture key actions. Additionally, in order to balance the sequence alignment and dependency, we propose to jointly train Connectionist Temporal Classification (CTC) and Long Short-Term Memory (LSTM). Experimental results demonstrate the effectiveness of the proposed method.
Liqing Gao, Rui-Ze Han, Wei Feng 0005
ICASSP5
2020 Modeling Cross-View Interaction Consistency for Paired Egocentric Interaction Recognition
abstract
With the development of Augmented Reality (AR), egocentric action recognition (EAR) plays an important role in accurately understanding demands from the user. However, EAR is designed to help recognize human-machine interaction in single egocentric view, thus difficult to capture interactions between two face-to-face AR users. Paired egocentric interaction recognition (PEIR) is the task to collaboratively recognize the interactions between two persons with the videos in their corresponding views. Unfortunately, existing PEIR methods always directly use linear decision function to fuse the features extracted from two corresponding egocentric videos, which ignore the consistency of interaction in paired egocentric videos. The consistency of interactions in paired videos, and features extracted from them, are correlated to each other. On top of that, we propose to derive the relevance between two views using bilinear pooling, which captures the consistency of two views in feature-level. Specifically, each neuron in the feature maps from one view connects to the neurons from the other view, which enforces the compact consistency between two views and then all possible paired neurons are used for PEIR. To be efficient, we use compact bilinear pooling with Count Sketch to avoid directly computing outer product. Experimental results on the PEV dataset shows the superiority of the proposed methods on the task PEIR.
Zhongguo Li, Fan Lyu, Wei Feng 0005, Song Wang 0002
ICME3
2020 Mutatt: Visual-Textual Mutual Guidance For Referring Expression Comprehension
abstract
Referring expression comprehension (REC) aims to localize a text-related region in a given image by a referring expression in natural language. Existing methods focus on how to build convincing visual and language representations independently, which may significantly isolate visual and language information. In this paper, we argue that for REC the referring expression and the target region are semantically correlated and subject, location and relationship consistency exist between vision and language. On top of this, we propose a novel approach called MutAtt to construct mutual guidance between vision and language, which treats vision and language equally thus yields compact information matching. Specifically, for each module of subject, location and relationship, MutAtt builds two kinds of attention-based mutual guidance strategies. One strategy is to generate vision-guided language embedding for the sake of matching relevant visual features. The other reversely generates language-guided visual features to match relevant language embedding. This mutual guidance strategy can effectively enforce the vision-language consistency in three modules. Experiments on three popular REC datasets demonstrate that the proposed approach outperforms the current state-of-the-art methods.
Fan Lyu, Wei Feng 0005, Song Wang 0002
ICME3
2020 Local and Global Structure-Aware Entropy Regularized Mean Teacher Model for 3D Left Atrium Segmentation
Wenlong Hang, Wei Feng 0005, Shuang Liang 0015, Lequan Yu, Qiong Wang 0001, Kup-Sze Choi, Harry Qin
MICCAI (1)2
2020 Complementary-View Co-Interest Person Detection
abstract
Fast and accurate identification of the co-interest persons, who draw joint interest of the surrounding people, plays an important role in social scene understanding and surveillance. Previous study mainly focuses on detecting co-interest persons from a single-view video. In this paper, we study a much more realistic and challenging problem, namely co-interest person~(CIP) detection from multiple temporally-synchronized videos taken by the complementary and time-varying views. Specifically, we use a top-view camera, mounted on a flying drone at a high altitude to obtain a global view of the whole scene and all subjects on the ground, and multiple horizontal-view cameras, worn by selected subjects, to obtain a local view of their nearby persons and environment details. We present an efficient top- and horizontal-view data fusion strategy to map multiple horizontal views into the global top view. We then propose a spatial-temporal CIP potential energy function that jointly considers both intra-frame confidence and inter-frame consistency, thus leading to an effective Conditional Random Field~(CRF) formulation. We also construct a complementary-view video dataset, which provides a benchmark for the study of multi-view co-interest person detection. Extensive experiments validate the effectiveness and superiority of the proposed method.
Rui-Ze Han, Jiewen Zhao, Wei Feng 0005, Yiyang Gan, Song Wang 0002
ACM Multimedia3
2020 DeepRhythm: Exposing DeepFakes with Attentional Visual Heartbeat Rhythms
abstract
As the GAN-based face image and video generation techniques, widely known as DeepFakes, have become more and more matured and realistic, there comes a pressing and urgent demand for effective DeepFakes detectors. Motivated by the fact that remote visual photoplethysmography (PPG) is made possible by monitoring the minuscule periodic changes of skin color due to blood pumping through the face, we conjecture that normal heartbeat rhythms found in the real face videos will be disrupted or even entirely broken in a DeepFake video, making it a potentially powerful indicator for DeepFake detection. In this work, we propose DeepRhythm, a DeepFake detection technique that exposes DeepFakes by monitoring the heartbeat rhythms. DeepRhythm utilizes dual-spatial-temporal attention to adapt to dynamically changing face and fake types. Extensive experiments on FaceForensics++ and DFDC-preview datasets have confirmed our conjecture and demonstrated not only the effectiveness, but also the generalization capability of DeepRhythm over different datasets by various DeepFakes generation techniques and multifarious challenging degradations.
Hua Qi, Qing Guo 0005, Felix Juefei-Xu, Xiaofei Xie, Lei Ma 0003, Wei Feng 0005, Yang Liu 0003, Jianjun Zhao 0001
ACM Multimedia6
2020 Human Identification and Interaction Detection in Cross-View Multi-Person Videos with Wearable Cameras
abstract
Compared to a single fixed camera, multiple moving cameras, e.g., those worn by people, can better capture the human interactive and group activities in a scene, by providing multiple, flexible and possibly complementary views of the involved people. In this setting the actual promotion of activity detection is highly dependent on the effective correlation and collaborative analysis of multiple videos taken by different wearable cameras, which is highly challenging given the time-varying view differences across different cameras and mutual occlusion of people in each video. By focusing on two wearable cameras and the interactive activities that involve only two people, in this paper we develop a new approach that can simultaneously: (i) identify the same persons across the two videos, (ii) detect the interactive activities of interest, including their occurrence intervals and involved people, and (iii) recognize the category of each interactive activity. Specifically, we represent each video by a graph, with detected persons as nodes, and propose a unified Graph Neural Network (GNN) based framework to jointly solve the above three problems. A graph matching network is developed for identifying the same persons across the two videos and a graph inference network is then used for detecting the human interactions. We also build a new video dataset, which provides a benchmark for this study, and conduct extensive experiments to validate the effectiveness and superiority of the proposed method.
Jiewen Zhao, Rui-Ze Han, Yiyang Gan, Wei Feng 0005, Song Wang 0002
ACM Multimedia5
2020 Watch out! Motion is Blurring the Vision of Your Deep Neural Networks
abstract
The state-of-the-art deep neural networks (DNNs) are vulnerable against adversarial examples with additive random-like noise perturbations. While such examples are hardly found in the physical world, the image blurring effect caused by object motion, on the other hand, commonly occurs in practice, making the study of which greatly important especially for the widely adopted real-time image processing tasks (e.g., object detection, tracking). In this paper, we initiate the first step to comprehensively investigate the potential hazards of blur effect for DNN, caused by object motion. We propose a novel adversarial attack method that can generate visually natural motion-blurred adversarial examples, named motion-based adversarial blur attack (ABBA). To this end, we first formulate the kernel-prediction-based attack where an input image is convolved with kernels in a pixel-wise way, and the misclassification capability is achieved by tuning the kernel weights. To generate visually more natural and plausible examples, we further propose the saliency-regularized adversarial kernel prediction, where the salient region serves as a moving object, and the predicted kernel is regularized to achieve naturally visual effects. Besides, the attack is further enhanced by adaptively tuning the translations of object and background. A comprehensive evaluation on the NeurIPS'17 adversarial competition dataset demonstrates the effectiveness of ABBA by considering various kernel sizes, translations, and regions. The in-depth study further confirms that our method shows a more effective penetrating capability to the state-of-the-art GAN-based deblurring mechanisms compared with other blurring methods. We release the code to \url{https://github.com/tsingqguo/ABBA}.
Qing Guo 0005, Felix Juefei-Xu, Xiaofei Xie, Lei Ma 0003, Jian Wang 0067, Wei Feng 0005, Yang Liu 0003
NeurIPS7
2020 Exemplar-based image saliency and co-saliency detection
Rui Huang 0006, Wei Feng 0005, Yaobin Zou
Neurocomputing2
2020 vtGraphNet: Learning weakly-supervised scene graph for complex visual grounding
Fan Lyu, Wei Feng 0005, Song Wang 0002
Neurocomputing2
2020 Selective Spatial Regularization by Reinforcement Learned Decision Making for Object Tracking
abstract
Spatial regularization (SR) is known as an effective tool to alleviate the boundary effect of correlation filter (CF), a successful visual object tracking scheme, from which a number of state-of-the-art visual object trackers can be stemmed. Nevertheless, SR highly increases the optimization complexity of CF and its target-driven nature makes spatially-regularized CF trackers may easily lose the occluded targets or the targets surrounded by other similar objects. In this paper, we propose selective spatial regularization (SSR) for CF-tracking scheme. It can achieve not only higher accuracy and robustness, but also higher speed compared with spatially-regularized CF trackers. Specifically, rather than simply relying on foreground information, we extend the objective function of CF tracking scheme to learn the target-context-regularized filters using target-context-driven weight maps. We then formulate the online selection of these weight maps as a decision making problem by a Markov Decision Process (MDP), where the learning of weight map selection is equivalent to policy learning of the MDP that is solved by a reinforcement learning strategy. Moreover, by adding a special state, representing not-updating filters, in the MDP, we can learn when to skip unnecessary or erroneous filter updating, thus accelerating the online tracking. Finally, the proposed SSR is used to equip three popular spatially-regularized CF trackers to significantly boost their tracking accuracy, while achieving much faster online tracking speed. Besides, extensive experiments on five benchmarks validate the effectiveness of SSR.
Qing Guo 0005, Rui-Ze Han, Wei Feng 0005, Zhihao Chen 0004
IEEE Trans. Image Process.3
2020 Fast Learning of Spatially Regularized and Content Aware Correlation Filter for Visual Tracking
abstract
With a good balance between accuracy and speed, correlation filter (CF) has become a popular and dominant visual object tracking scheme. It implicitly extends the training samples by circular shifts of a given target patch, which serve as negative samples for fast online learning of the filters. Since all these shifted patches are not real negative samples of the target, CF tracking scheme suffers from the annoying boundary effects that can greatly harm the tracking performance, especially under challenging situations, like occlusion and fast temporal variation. Spatial regularization is known as a potent way to alleviate such boundary effects, but with the cost of highly increased time complexity, caused by complex optimization imported by spatial regularization. In this paper, we propose a new fast learning approach to content-aware spatial regularization, namely weighted sample based CF tracking (WSCF). In WSCF, specifically, we present a simple yet effective energy function that implicitly weighs different training samples by spatial deviations. With the energy function, the learning of correlation filters is composed of two subproblems with closed-form solution and can be efficiently solved in an alternate way. We further develop a content-aware updating strategy to dynamically refine the weight distribution to well adapt to the temporal variations of the target and background. Finally, the proposed WSCF is used to enhance two state-of-the-art CF trackers to significantly boost their tracking accuracy, with little sacrifice on the tracking speed. Extensive experiments on five benchmarks validate the effectiveness of the proposed approach.
Rui-Ze Han, Wei Feng 0005, Song Wang 0002
IEEE Trans. Image Process.2
2019 Learned Map Prediction for Enhanced Mobile Robot Exploration
abstract
We demonstrate an autonomous ground robot capable of exploring unknown indoor environments for reconstructing their 2D maps. This problem has been traditionally tackled by geometric heuristics and information theory. More recently, deep learning and reinforcement learning based approaches have been proposed to learn exploration behavior in an end-to-end manner. We present a method that combines the strengths of these different approaches. Specifically, we employ a state-of-the-art generative neural network to predict unknown regions of a partially explored map, and use the prediction to enhance the exploration in an information-theoretic manner. We evaluate our system in simulation using floor plans of real buildings. We also present comparisons with traditional methods which demonstrate the advantage of our method in terms of exploration efficiency. We retain an advantage over end-to-end learned exploration methods in that the robot's behavior is easily explicable in terms of the predicted map.
Rakesh Shrestha, Fei-Peng Tian, Wei Feng 0005, Ping Tan 0002, Richard Vaughan 0001
ICRA3
2019 Story co-segmentation of Chinese broadcast news using weakly-supervised semantic similarity
Wei Feng 0005, Xuecheng Nie, Yujun Zhang 0002, Jianwu Dang 0001
Neurocomputing1
2019 Fast and object-adaptive spatial regularization for correlation filters based tracking
Qing Guo 0005, Wei Feng 0005
Neurocomputing3
2019 Woodblock image decomposition of Chinese new year paintings
Haipeng Dai 0002, Wei Feng 0005, Jiawan Zhang
Multim. Tools Appl.4
2019 Active Camera Relocalization from a Single Reference Image without Hand-Eye Calibration
abstract
This paper studies active relocalization of 6D camera pose from a single reference image, a new and challenging problem in computer vision and robotics. Straightforward active camera relocalization (ACR) is a tricky and expensive task that requires elaborate hand-eye calibration on precision robotic platforms. In this paper, we show that high-quality camera relocalization can be achieved in an active and much easier way. We propose a hand-eye calibration free approach to actively relocating the camera to the same 6D pose that produces the input reference image. We theoretically prove that, given bounded unknown hand-eye pose displacement, this approach is able to rapidly reduce both 3D relative rotational and translational pose between current camera and the reference one to an identical matrix and a zero vector, respectively. Based on these findings, we develop an effective ACR algorithm with fast convergence rate, reliable accuracy and robustness. Extensive experiments validate the effectiveness and feasibility of our approach on both laboratory tests and challenging real-world applications in fine-grained change monitoring of cultural heritages.
Fei-Peng Tian, Wei Feng 0005, Qian Zhang 0051, Vincenzo Loia
IEEE Trans. Pattern Anal. Mach. Intell.2
2019 Dynamic Saliency-Aware Regularization for Correlation Filter-Based Object Tracking
abstract
With a good balance between tracking accuracy and speed, correlation filter (CF) has become one of the best object tracking frameworks, based on which many successful trackers have been developed. Recently, spatially regularized CF tracking (SRDCF) has been developed to remedy the annoying boundary effects of CF tracking, thus further boosting the tracking performance. However, SRDCF uses a fixed spatial regularization map constructed from a loose bounding box and its performance inevitably degrades when the target or background show significant variations, such as object deformation or occlusion. To address this problem, we propose a new dynamic saliency-aware regularized CF tracking (DSAR-CF) scheme. In DSAR-CF, a simple yet effective energy function, which reflects the object saliency and tracking reliability in the spatial-temporal domain, is defined to guide the online updating of the regularization weight map using an efficient level-set algorithm. Extensive experiments validate that the proposed DSAR-CF leads to better performance in terms of accuracy and speed than the original SRDCF.
Wei Feng 0005, Rui-Ze Han, Qing Guo 0005, Jianke Zhu, Song Wang 0002
IEEE Trans. Image Process.1
2018 Co-Saliency Detection Within a Single Image
abstract
Recently, saliency detection in a single image and co-saliency detection in multiple images have drawn extensive research interest in the vision community. In this paper, we investigate a new problem of co-saliency detection within a single image, i.e., detecting within-image co-saliency. By identifying common saliency within an image, e.g., highlighting multiple occurrences of an object class with similar appearance, this work can benefit many important applications, such as the detection of objects of interest, more robust object recognition, reduction of information redundancy, and animation synthesis. We propose a new bottom-up method to address this problem. Specifically, a large number of object proposals are first detected from the image. Then we develop an optimization algorithm to derive a set of proposal groups, each of which contains multiple proposals showing good common saliency in the original image. For each proposal group, we calculate a co-saliency map and then use a low-rank based algorithm to fuse the maps calculated from all the proposal groups for the final co-saliency map in the image. In the experiment, we collect a new dataset of 364 color images with within-image cosaliency. Experiment results show that the proposed method can better detect the within-image co-saliency than existing algorithms.
Hongkai Yu, Jianwu Fang, Hao Guo 0002, Wei Feng 0005, Song Wang 0002
AAAI5
2018 Spherical Superpixels: Benchmark and Evaluation
Xiaorui Xu, Qiang Zhao 0005, Wei Feng 0005
ACCV (6)4
2018 Active Camera Relocalization with RGBD Camera from a Single 2D Image
abstract
Active camera relocalization (ACR) focuses on dynamically and physically relocating camera to a previous pose, effectively supporting many applications in computer vision and robotics, such as automated picking and stowing, fine-grained change detection. Previous work [1] uses barely 2D images to realize ACR, bringing about unknown translation scale problem. To solve this problem, they use bisection approach to guess translation scale, which leads to reciprocating motion and slows down the convergence process. In this paper, we utilize additional depth information from an RGBD camera to solve the real translation scale problem. Via iteratively and sequentially adjusting 3D translation and rotation, our ACR approach greatly reduces the iteration number and speeding up the process. To cope with imprecise pose estimation and achieve high relocalization accuracy, we propose a bounding strategy to restrict camera motion. Experiments validate the proposed method is much efficient and its accuracy is on par with previous ACR method.
Dongxu Miao, Fei-Peng Tian, Wei Feng 0005
ICASSP3
2018 Background-Suppressed Correlation Filters for Visual Tracking
abstract
Correlation filters (CF) visual object tracking is a powerful framework, with excellent tracking accuracy and beyond real-time frame rate. Its performance, however, can be severely degraded in cluttered background images. In this paper, we propose background-suppressed correlation filters (BSCF), a better CF tracking scheme, which can significantly improve the reliability and accuracy of CF trackers, without harming their beyond real-time speed. Specifically, we present a unified BSCF object function. We show that both the correlation filters and BS weight map can be efficiently and jointly solved in frequency domain. Extensive experiments on OTB-100 benchmark validate the effectiveness and generality of BS in improving multiple CF trackers with higher accuracy and robustness while maintaining their fast tracking speed. We also show BS boosted CF tracker can achieve comparable accuracy of the state-of-the-art spatially-regularized CF tracker but is 14 times faster.
Zhihao Chen 0004, Qing Guo 0005, Wei Feng 0005
ICME4
2018 Content-Related Spatial Regularization for Visual Object Tracking
abstract
Spatial regularization (SR), being an effective tool to alleviate the boundary effects, can significantly improve the accuracy and robustness of correlation filters (CF) based visual object tracking. The core of SR is a spatially variant weight map that is used to regularize the online learned correlation filters by selecting more meaningful samples. However, most existing trackers apply a data-independent SR weight map. In this paper, we show that a content-related spatial regularization (CRSR) can help to further boost both the tracking accuracy and robustness. Specifically, we present to consider both frame saliency and spatial preference to online generate the CRSR weight map and propose a simple yet effective saliency-embedded CF objective function to simultaneously optimize the filters and CRSR weight map in spatial-temporal domain. Extensive experiments validate that our content-related SR outperforms the classical SR, with higher tracking accuracy and almost two times faster speed.
Rui-Ze Han, Qing Guo 0005, Wei Feng 0005
ICME3
2018 Soft Clustering Guided Image Smoothing
abstract
Image smoothing, which aims to remove unwanted textures and preserve desired structures, plays an important role in many multimedia and computer vision tasks. The key to image smoothing, despite different applications, is to distinguish the structures from the textures. This paper presents a novel image smoothing method, following the principle that, for a certain pixel, its neighbors in both space and intensity should contribute more on smoothing, while the distant ones be insulated for avoiding over-smoothing. Intuitively, clustering is a good candidate to achieve the goal. However, due to rich textures and clutters within images, simply performing the clustering on the input likely obtains inaccurate results, and thus leads to unsatisfied smoothing results. In addition, for our task, using traditional hard clustering techniques is at high risk of generating staircase artifacts. For addressing these issues, an algorithm is customized, which on the one hand adopts the soft clustering to more faithfully assign pixels, on the other hand iterates the soft clustering and smoothing, expecting to improve each other. Experiments on several challenging images are provided to show the efficacy of our method, and its superiority over other prevailing approaches.
Liang Li 0039, Xiaojie Guo 0001, Wei Feng 0005, Jiawan Zhang
ICME3
2018 Fast and Reliable Computational Rephotography on Mobile Device
abstract
Rephotography, aiming to take a well-aligned image from the same pose of a reference image, is a very useful technique to study history, monitor environment and detect minute changes. Previous works either rely on human eye judgement that is very challenging and tedious or depend on high-precision robot platform that is non-portable and inapplicable to wild scenes. In this paper, we propose a fast and reliable computational rephotography approach that works well on mobile devices. Through well designed fast matching, our approach can generate faithful visual navigation rectangles and provide reliable navigation information in nearly real-time, thus is able to easily guide the rephotography process. To improve robustness and speed, we present a series of techniques including effective keyframe decision, feature substitution' flow-based fast matching and adaptive matching mode switching. As a result, our approach can be well applied on wild environments and robust to large illumination and target changes. Extensive experiments verify its accuracy, effectiveness and generality on both planar and 3D scenes.
Yi-Bo Shi, Fei-Peng Tian, Dongxu Miao, Wei Feng 0005
ICME4
2018 Active Recurrence of Lighting Condition for Fine-Grained Change Detection
abstract
This paper addresses active lighting recurrence (ALR), a new problem that actively relocalizes a light source to physically reproduce the lighting condition for a same scene from single reference image. ALR is of great importance for fine-grained visual monitoring and change detection, because some phenomena or minute changes can only be clearly observed under particular lighting conditions. Hence, effective ALR should be able to online navigate a light source toward the target pose, which is challenging due to the complexity and diversity of real-world lighting \& imaging processes. We propose to use the simple parallel lighting as an analogy model and based on Lambertian law to compose an instant navigation ball for this purpose. We theoretically prove the feasibility of this ALR strategy for realistic near point light sources and its invariance to the ambiguity of normal \& lighting decomposition. Extensive quantitative experiments and challenging real-world tasks on fine-grained change monitoring of cultural heritages verify the effectiveness of our approach. We also validate its generality to non-Lambertian scenes.
Qian Zhang 0051, Wei Feng 0005, Fei-Peng Tian, Ping Tan 0002
IJCAI2
2018 Fast Spatially-Regularized Correlation Filters for Visual Object Tracking
Qing Guo 0005, Wei Feng 0005
PRICAI (1)3
2018 High-Resolution Depth Refinement by Photometric and Multi-shading Constraints
Yujun Zhang 0002, Qian Zhang 0051, Wei Feng 0005
PRICAI3
2018 Unsupervised measure of Chinese lexical semantic similarity using correlated graph model for news story segmentation
Wei Feng 0005, Xuecheng Nie, Yujun Zhang 0002, Lei Xie 0001, Jianwu Dang 0001
Neurocomputing1
2018 Frequency-tuned active contour model
Qing Guo 0005, Shuifa Sun, Xuhong Ren, Fangmin Dong, Bruce Zhi Gao, Wei Feng 0005
Neurocomputing6
2018 Multiscale blur detection by learning discriminative deep features
Rui Huang 0006, Wei Feng 0005, Mingyuan Fan 0001
Neurocomputing2
2018 Multiple human tracking in wearable camera videos with informationless intervals
Hongkai Yu, Haozhou Yu, Hao Guo 0002, Jeff P. Simmons, Qin Zou 0001, Wei Feng 0005, Song Wang 0002
Pattern Recognit. Lett.6
2017 Frequency-tuned ACM for biomedical image segmentation
abstract
Biomedical images are usually corrupted by strong noise and intensity inhomogeneity simultaneously. Existing region-based active contour models (RACMs) easily fail when segmenting such images. In the frequency domain, we propose a generalized RACM that presents a new way to understand the essence of classical RACMs whose segmentation results are determined by a frequency filter to extract the proposed frequency boundary energy. Then, we introduce the difference of Gaussians as the optimal filter to exclude strong noise and intensity inhomogeneity effectively. We show superior performance of the model by comparing with six state-of-the-art methods on challenge biomedical images and segmenting an optical coherence tomography image sequence.
Qing Guo 0005, Shuifa Sun, Fangmin Dong, Wei Feng 0005, Bruce Zhi Gao, Siyu Ma
ICASSP4
2017 Selective object and context tracking
abstract
Robust appearance model is significantly important to state-of-the-art trackers. However, such trackers highly rely on the reliability of foreground appearance model. When the foreground is seriously occluded or the scene contains multiple objects with similar appearance, such foundation is destroyed. To extend the ability of trackers to handle these difficulties, we propose selective object and context tracking to locate the target according to the reliability of the foreground appearance model which is determined by two measures about whether the target is occluded or surrounded by similar objects. Extensive experiments show that our method achieves better performance than state-of-the-art trackers on VOT TIR-2015 dataset and is able to track the target even when the foreground appearance is completely unreliable.
Ce Zhou, Qing Guo 0005, Wei Feng 0005
ICASSP4
2017 Learning Dynamic Siamese Network for Visual Object Tracking
abstract
How to effectively learn temporal variation of target appearance, to exclude the interference of cluttered background, while maintaining real-time response, is an essential problem of visual object tracking. Recently, Siamese networks have shown great potentials of matching based trackers in achieving balanced accuracy and beyond realtime speed. However, they still have a big gap to classification & updating based trackers in tolerating the temporal changes of objects and imaging conditions. In this paper, we propose dynamic Siamese network, via a fast transformation learning model that enables effective online learning of target appearance variation and background suppression from previous frames. We then present elementwise multi-layer fusion to adaptively integrate the network outputs using multi-level deep features. Unlike state-of-theart trackers, our approach allows the usage of any feasible generally- or particularly-trained features, such as SiamFC and VGG. More importantly, the proposed dynamic Siamese network can be jointly trained as a whole directly on the labeled video sequences, thus can take full advantage of the rich spatial temporal information of moving objects. As a result, our approach achieves state-of-the-art performance on OTB-2013 and VOT-2015 benchmarks, while exhibits superiorly balanced accuracy and real-time response over state-of-the-art competitors.
Qing Guo 0005, Wei Feng 0005, Ce Zhou, Rui Huang 0006, Song Wang 0002
ICCV2
2017 Near-surface lighting estimation and reconstruction
abstract
In this paper, we propose an effective approach to estimating a near-surface lighting function from a limited number of images captured under different illuminations. Unlike classical methods relying on simplified parallel lighting model or near-point lighting model, our approach directly focuses on the much more realistic near-surface light source and formulates it as a regular grid of near-point light sources. We present an iterative joint optimization strategy to solve the scene normal, reflectance and near-point light source positions. Based on such new model, reliable relighting under arbitrary new illuminations can be faithfully reconstructed by applying the given lighting condition to the same scene. Experiments show that the proposed approach can generate more accurate re-lighting results than state-of-the-art competitors.
Qian Zhang 0051, Fei-Peng Tian, Rui-Ze Han, Wei Feng 0005
ICME4
2017 Color Feature Reinforcement for Cosaliency Detection Without Single Saliency Residuals
abstract
Cosaliency detects the common salient objects within a group of images. Hence, those objects that are salient only in individual image or small portion of the image group should conceptually be treated as background. However, most state-of-the-art methods cannot do this well because they measure cosaliency as an explicit combination of single-image saliency and interimage similarity, thus inevitably leaving single saliency residuals into the cosaliency maps. In this letter, we show such problem can be solved by color feature reinforcement, based on a simple observation that cosalient objects usually have similar color distributions in an abundant color feature space. Since we model the cosaliency of an image w.r.t. another one as a reinforced product of the foreground dictionary and sparsely coded saliency coefficients of the two images, respectively, within a same rich feature space, we can effectively eliminate the single saliency residual effect in cosaliency detection. Extensive experiments validate the superior performance of the proposed approach on benchmark datasets.
Rui Huang 0006, Wei Feng 0005
IEEE Signal Process. Lett.2
2017 Structure-Regularized Compressive Tracking With Online Data-Driven Sampling
abstract
Being a powerful appearance model, compressive random projection derives effective Haar-like features from non-rotated 4-D-parameterized rectangles, thus supporting fast and reliable object tracking. In this paper, we show that such successful fast compressive tracking scheme can be further significantly improved by structural regularization and online data-driven sampling. Our major contribution is threefold. First, we find that superpixel-guided compressive projection can generate more discriminative features by sufficiently capturing rich local structural information of images. Second, we propose fast directional integration that enables low-cost extraction of feasible Haar-like features from arbitrarily rotated 5-D-parameterized rectangles to realize more accurate object localization. Third, beyond naive dense uniform sampling, we present two practical online data-driven sampling strategies to produce less yet more effective candidate and training samples for object detection and classifier updating, respectively. Extensive experiments on real-world benchmark data sets validate the superior performance, i.e., much better object localization ability and robustness, of the proposed approach over state-of-the-art trackers.
Qing Guo 0005, Wei Feng 0005, Ce Zhou, Chi-Man Pun
IEEE Trans. Image Process.2
2016 6D Dynamic Camera Relocalization from Single Reference Image
abstract
Dynamic relocalization of 6D camera pose from single reference image is a costly and challenging task that requires delicate hand-eye calibration and precision positioning platform to do 3D mechanical rotation and translation. In this paper, we show that high-quality camera relocalization can be achieved in a much less expensive way. Based on inexpensive platform with unreliable absolute repositioning accuracy (ARA), we propose a hand-eye calibration free strategy to actively relocate camera into the same 6D pose that produces the input reference image, by sequentially correcting 3D relative rotation and translation. We theoretically prove that, by this strategy, both rotational and translational relative pose can be effectively reduced to zero, with bounded unknown hand-eye pose displacement. To conquer 3D rotation and translation ambiguity, this theoretical strategy is further revised to a practical relocalization algorithm with faster convergence rate and more reliability by jointly adjusting 3D relative rotation and translation. Extensive experiments validate the effectiveness and superior accuracy of the proposed approach on laboratory tests and challenging real-world applications.
Wei Feng 0005, Fei-Peng Tian, Qian Zhang 0051
CVPR1
2016 Structure-regularized compressive tracking
abstract
Compressive random projection is a powerful appearance model to derive effective Haar-like features from non-rotated 4D rectangles, which can support fast and reliable object tracking. In this paper, we show that such successful compressive tracking scheme can be further significantly improved by structural regularization. Specifically, we propose two effective structural regularizations. First, we find that, guided by superpixels, compressive random projection can always generate more discriminative features by sufficiently capturing the rich local structure information of images. Second, we present fast directional integration to enable low-cost extraction of feasible Haar-like features from arbitrarily rotated 5D rectangles to realize more accurate object localization. We compare the proposed structure-regularized compressive tracker with a number of state-of-the-art methods. Extensive experiments on challenging benchmark dataset validate the superior performance and comparable real-time speed of the proposed approach.
Qing Guo 0005, Wei Feng 0005, Ce Zhou
ICME2
2016 Tone-Mapped Mean-Shift Based Environment Map Sampling
abstract
In this paper, we present a novel approach for environment map sampling, which is an effective and pragmatic technique to reduce the computational cost of realistic rendering and get plausible rendering images. The proposed approach exploits the advantage of adaptive mean-shift image clustering with aid of tone-mapping, yielding oversegmented strata that have uniform intensities and capture shapes of light regions. The resulted strata, however, have unbalanced importance metric values for rendering, and the strata number is not user-controlled. To handle these issues, we develop an adaptive split-and-merge scheme that refines the strata and obtains a better balanced strata distribution. Compared to the state-of-the-art methods, our approach achieves comparable and even better rendering quality in terms of SSIM, RMSE and HDRVDP2 image quality metrics. Experimental results further show that our approach is more robust to the variation of viewpoint, environment rotation, and sample number.
Wei Feng 0005, Changguo Yu
IEEE Trans. Vis. Comput. Graph.1
2015 Fine-Grained Change Detection of Misaligned Scenes with Varied Illuminations
abstract
Detecting fine-grained subtle changes among a scene is critically important in practice. Previous change detection methods, focusing on detecting large-scale significant changes, cannot do this well. This paper proposes a feasible end-to-end approach to this challenging problem. We start from active camera relocation that quickly relocates camera to nearly the same pose and position of the last time observation. To guarantee detection sensitivity and accuracy of minute changes, in an observation, we capture a group of images under multiple illuminations, which need only to be roughly aligned to the last time lighting conditions. Given two times observations, we formulate fine-grained change detection as a joint optimization problem of three related factors, i.e., normal-aware lighting difference, camera geometry correction flow, and real scene change mask. We solve the three factors in a coarse-to-fine manner and achieve reliable change decision by rank minimization. We build three real-world datasets to benchmark fine-grained change detection of misaligned scenes under varied multiple lighting conditions. Extensive experiments show the superior performance of our approach over state-of-the-art change detection methods and its ability to distinguish real scene changes from false ones caused by lighting variations.
Wei Feng 0005, Fei-Peng Tian, Qian Zhang 0051
ICCV1
2015 Saliency and co-saliency detection by low-rank multiscale fusion
abstract
To facilitate efficiency, most recent successful saliency detection methods are built on superpixel level. However, saliency detection with single-scale superpixel segmentation may fail in capturing the intrinsic salient objects in complex natural scenes with small-scale high-contrast backgrounds. To tackle this problem and realize more reliable saliency detection, we present a simple strategy using multiscale superpixels to jointly detect salient object via low-rank analysis. Specifically, we construct a multiscale superpixel pyramid and derive the corresponding saliency map using multiple saliency features and priors for each single scale at first. Then, we show that by joint low-rank analysis of multiscale saliency maps, we can obtain a more reliable adaptively fused saliency map that takes all scales saliency results into account. We further propose a GMM-based co-saliency prior to enable the above approach to detecting co-salient objects from multiple images. Extensive experiments on benchmark datasets validate the effectiveness and superiority of the proposed approach over state-of-the-art methods.
Rui Huang 0006, Wei Feng 0005
ICME2
2015 SPHORB: A Fast and Robust Binary Feature on the Sphere
Qiang Zhao 0005, Wei Feng 0005, Jiawan Zhang
Int. J. Comput. Vis.2
2015 Topic segmentation on spoken documents using self-validated acoustic cuts
Hongjie Chen 0001, Lei Xie 0001, Wei Feng 0005, Lilei Zheng, Yanning Zhang 0001
Soft Comput.3
2015 NestDE: generic parameters tuning for automatic story segmentation
Wei Feng 0005, Xuefei Yin, Lei Xie 0001
Soft Comput.1
2015 Layered modeling and generation of Pollock's drip style
Yan Zheng 0002, Xuecheng Nie, Zhaopeng Meng, Wei Feng 0005, Kang Zhang 0001
Vis. Comput.4
2014 L0 co-intrinsic images decomposition
abstract
In this paper, we focus on co-intrinsic decomposition, a new problem that performs intrinsic decomposition on a pair of images simultaneously, which share the same foreground with arbitrarily different illuminations and backgrounds. We specifically demand the common foreground across different images to share same reflectance values. For the purpose of efficiency and feasibility, we perform the co-intrinsic decomposition at superpixel-level and propose a uniform approach to automatically derive non-local reflectance relationships via unsupervised L0sparsity between superpixels from intra-and inter-images. We present a unicolor-light-based intrinsic model, from which we construct a non-local L0sparse co-Retinex model that imposes feasible constraints on shading, reflectance and environment light, respectively. The co-intrinsic decomposition is finally modeled as a quadratic minimization problem that leads to a fast closed form solution. Extensive experiments show plausible results of our approach in extracting common reflectance components from multiple images. We also validate the benefits of our results in boosting the accuracy of image co-saliency detection.
Haipeng Dai 0002, Wei Feng 0005, Xuecheng Nie
ICME2
2014 Contrast enhancement based single image dehazing VIA TV-l1 minimization
abstract
In this paper, we propose a general algorithm to removing haze from single images using total variation minimization. Our approach stems from two simple yet fundamental observations about haze-free images and the haze itself. First, clear-day images usually have stronger contrast than images plagued by bad weather; and second, the variations in natural atmospheric veil, which highly depends on the depth of objects, always tend to be smooth. Integrating these two criteria together leads to a new effective dehazing model, which encourages the gradient ℓ1sparsity of atmospheric veil and implicitly maximizes the global contrast of haze-free image in the meanwhile. We also show that the proposed dehazing model can be efficiently solved using the TV-ℓ1minimization. Compared to alternative state-of-the-art methods, our approach is physically plausible and works well for all types of hazy situations. Comparative study and quantitative evaluation on both synthetic and natural images validate the superior performance and the generality of our approach.
Liang Li 0039, Wei Feng 0005, Jiawan Zhang
ICME2
2014 Intrinsic image decomposition by hierarchical L0 sparsity
abstract
This paper presents a hierarchical approach to single image intrinsic decomposition based on non-local L0sparsity. In contrast to previous studies using heuristic methods to well-define the ill-posed problem, our approach is able to effectively construct sparse, non-local and multiscale reflectance dependencies in an unsupervised manner, thus is less dependent on the chromaticity feature and more accurately captures the global reflectance correlations. Besides, we impose homogenous smoothness prior and scale constraint in our model to further improve the decomposition accuracy. We formulate the decomposition as a quadratic minimization problem, which can be efficiently solved in closed form. Extensive experiments show that our approach can successfully extract the shading and reflectance components from a single image, and outperforms state-of-the-art methods on benchmark dataset. Besides, our approach can achieve comparable results with user-assisted methods on natural scenes.
Xuecheng Nie, Wei Feng 0005, Haipeng Dai 0002, Chi-Man Pun
ICME2
2014 Bag of squares: A reliable model of measuring superpixel similarity
abstract
As the increasing popularity of superpixel-based applications, measuring superpixel-level similarity becomes an important and commonly required problem. In this paper, we propose a general bag of squares (BoS) model for such particular purpose. Compared to existing methods, our approach provides a full scheme to both invariantly represent superpixels and accurately measure their pairwise similarities. In order to handle the split-and-merge variety of superpixels of same objects in different scenes, our model is based on superpixel pyramid. As a result, the BoS model of a superpixel is built upon a group of subregions consisting of the superpixel itself and its children subregions in the pyramid. For each subregion, we extract a proper number of maximum squares via distance transform, and then use a fast self-validated approach to clustering them into a small number of dominant squares, which together with a rotation and scale invariant square descriptor, jointly compose the BoS model for the particular superpixel. Finally, we measure the similarity between a pair of superpixels by the closeness of their BoS models. Experiments on interactive object segmentation and co-saliency detection show that the proposed BoS model can reliably capture the delicate differences among superpixels, thus always producing better segmentation results, especially for segmenting highly variant objects in clutter scenes.
Wei Feng 0005, Jiawan Zhang, Chi-Man Pun
ICME2
2014 Self-Adaptively Weighted Co-Saliency Detection via Rank Constraint
abstract
Co-saliency detection aims at discovering the common salient objects existing in multiple images. Most existing methods combine multiple saliency cues based on fixed weights, and ignore the intrinsic relationship of these cues. In this paper, we provide a general saliency map fusion framework, which exploits the relationship of multiple saliency cues and obtains the self-adaptive weight to generate the final saliency/cosaliency map. Given a group of images with similar objects, our method firstly utilizes several saliency detection algorithms to generate a group of saliency maps for all the images. The feature representation of the co-salient regions should be both similar and consistent. Therefore, the matrix jointing these feature histograms appears low rank. We formalize this general consistency criterion as the rank constraint, and propose two consistency energy to describe it, which are based on low rank matrix approximation and low rank matrix recovery, respectively. By calculating the self-adaptive weight based on the consistency energy, we highlight the common salient regions. Our method is valid for more than two input images and also works well for single image saliency detection. Experimental results on a variety of benchmark datasets demonstrate that the proposed method outperforms the state-of-the-art methods.
Xiaochun Cao, Zhiqiang Tao, Huazhu Fu, Wei Feng 0005
IEEE Trans. Image Process.5
2013 Maximum Cohesive Grid of Superpixels for Fast Object Localization
abstract
This paper addresses a challenging problem of regularizing arbitrary super pixels into an optimal grid structure, which may significantly extend current low-level vision algorithms by allowing them to use super pixels (SPs) conveniently as using pixels. For this purpose, we aim at constructing maximum cohesive SP-grid, which is composed of real nodes, i.e SPs, and dummy nodes that are meaningless in the image with only position-taking function in the grid. For a given formation of image SPs and proper number of dummy nodes, we first dynamically align them into a grid based on the centroid localities of SPs. We then define the SP-grid coherence as the sum of edge weights, with SP locality and appearance encoded, along all direct paths connecting any two nearest neighboring real nodes in the grid. We finally maximize the SP-grid coherence via cascade dynamic programming. Our approach can take the regional objectness as an optional constraint to produce more semantically reliable SP-grids. Experiments on object localization show that our approach outperforms state-of-the-art methods in terms of both detection accuracy and speed. We also find that with the same searching strategy and features, object localization at SP-level is about 100-500 times faster than pixel-level, with usually better detection accuracy.
Liang Li 0039, Wei Feng 0005, Jiawan Zhang
CVPR2
2013 Measuring semantic similarity by contextualword connections in Chinese news story segmentation
abstract
A lot of recent work in story segmentation focuses on developing better partitioning criteria to segment news transcripts into sequences of topically coherent stories, while simply relying on the repetition based hard word-level similarities and ignoring the semantic correlations between different words. In this paper, we propose a purely data-driven approach to measuring soft semantic word- and sentence-level similarity from a given corpus, without the guidance of linguistic knowledge, ground-truth topic labeling or story boundaries. We show that contextual word connections can help to produce semantically meaningful similarity measurement between any pair of Chinese words. Based on this, we further use a parallel all-pair SimRank algorithm to propagate such contextual similarities throughout the whole vocabulary. The resultant word semantic similarity matrix is then used to refine the classical cosine similarity measurement of sentences. Experiments on benchmark Chinese news corpora show that, story segmentation using the proposed soft semantic similarity measurement can always produce better segmentation accuracy than using the hard similarity. Specifically, we can achieve 3%-10% average F1-measure improvement to state-of-the-art NCuts based story segmentation.
Xuecheng Nie, Wei Feng 0005, Lei Xie 0001
ICASSP2
2013 Image co-saliency detection by propagating superpixel affinities
abstract
Image co-saliency detection is a valuable technique to highlight perceptually salient regions in image pairs. In this paper, we propose a self-contained co-saliency detection algorithm based on superpixel affinity matrix. We first compute both intra and inter similarities of superpixels of image pairs. Bipartite graph matching is applied to determine most reliable inter similarities. To update the similarity score between every two superpixels, we next employ a GPU-based all-pair SimRank algorithm to do propagation on the affinity matrix. Based on the inter superpixel affinities we derive a co-saliency measure that evaluates the foreground cohesiveness and locality compactness of superpixels within one image. The effectiveness of our method is demonstrated in experimental evaluation.
Zhiyu Tan, Wei Feng 0005, Chi-Man Pun
ICASSP3
2013 A tighter lower bound estimate for dynamic time warping
abstract
In this paper, we propose a new lower-bound estimate for speeding up dynamic time warping (DTW) on multivariate time sequences. It has several advantages as compared with the inner-product lower bound [1] recently proposed to eliminate a large number of DTW computations. First, we prove that it is tighter than the inner product lower bound while the computational complexity remains comparable. Second, the inner product lower bound is specifically designed for the inner product distance while the proposed lower bound is valid for any distance measure. Third, DTW search can be further speeded up since the distance matrix is calculated in advance at the lower bound estimation stage. Spoken term detection experiments on the TIMIT corpus show that the proposed lower bound estimate is able to reduce the computational requirements for DTW-KNN search by 54% as compared with the inner-product lower bound. in black ink.
Lei Xie 0001, Qiao Luan, Wei Feng 0005
ICASSP4
2013 A spectral-multiplicity-tolerant approach to robust graph matching
Wei Feng 0005, Chi-Man Pun, Jianmin Jiang
Pattern Recognit.1
2013 Efficient Semisupervised MEDLINE Document Clustering With MeSH-Semantic and Global-Content Constraints
abstract
For clustering biomedical documents, we can consider three different types of information: the local-content (LC) information from documents, the global-content (GC) information from the whole MEDLINE collections, and the medical subject heading (MeSH)-semantic (MS) information. Previous methods for clustering biomedical documents are not necessarily effective for integrating different types of information, by which only one or two types of information have been used. Recently, the performance of MEDLINE document clustering has been enhanced by linearly combining both the LC and MS information. However, the simple linear combination could be ineffective because of the limitation of the representation space for combining different types of information (similarities) with different reliability. To overcome the limitation, we propose a new semisupervised spectral clustering method, i.e., SSNCut, for clustering over the LC similarities, with two types of constraints: must-link (ML) constraints on document pairs with high MS (or GC) similarities and cannot-link (CL) constraints on those with low similarities. We empirically demonstrate the performance of SSNCut on MEDLINE document clustering, by using 100 data sets of MEDLINE records. Experimental results show that SSNCut outperformed a linear combination method and several well-known semisupervised clustering methods, being statistically significant. Furthermore, the performance of SSNCut with constraints from both MS and GC similarities outperformed that from only one type of similarities. Another interesting finding was that ML constraints more effectively worked than CL constraints, since CL constraints include around 10% incorrect ones, whereas this number was only 1% for ML constraints.
Wei Feng 0005, Hiroshi Mamitsuka, Shanfeng Zhu
IEEE Trans. Cybern.2
2013 Cube2Video: Navigate Between Cubic Panoramas in Real-Time
abstract
Online virtual navigation systems enable users to hop from one 360° panorama to another, which belong to a sparse point-to-point collection, resulting in a less pleasant viewing experience. In this paper, we present a novel method, namely Cube2Video, to support navigating between cubic panoramas in a video-viewing mode. Our method circumvents the intrinsic challenge of cubic panoramas, i.e., the discontinuities between cube faces, in an efficient way. The proposed method extends the matching-triangulation-interpolation procedure with special considerations of the spherical domain. A triangle-to-triangle homography-based warping is developed to achieve physically plausible and visually pleasant interpolation results. The temporal smoothness of the synthesized video sequence is improved by means of a compensation transformation. As experimental results demonstrate, our method can synthesize pleasant video sequences in real time, thus mimicking walking or driving navigation.
Qiang Zhao 0005, Wei Feng 0005, Jiawan Zhang, Tien-Tsin Wong
IEEE Trans. Multim.3
2012 Scalable image co-segmentation using color and covariance features
Wei Feng 0005, Jiawan Zhang, Jianmin Jiang
ICPR2
2012 Lexical Story Co-Segmentation of Chinese Broadcast News
abstract
We present an unsupervised technique, namely story co-segmentation, to automatically extract the common sto-ries on the same topic within a pair of Chinese broadcast news transcripts. Unlike classical topic tracking that usu-ally relies on previously trained topic models, our method is purely data-driven and is able to simultaneously deter-mine the common stories of the input texts. Specifical-ly, we propose an iterative four-step MRF solution to the problem of story co-segmentation using lexical cues only. We first construct a sentence-level graph formulation of the input news transcripts, and initialize foreground and background labeling by lexical clustering. We then up-date both foreground and background models based on the current labeling. We formalize story co-segmentation as a Gibbs energy minimization problem that balances the optimal objectives of foreground/background likeli-hood, intra-doc coherence, and inter-doc similarity. Fi-nally, the labeling refinement is obtained by hybrid op-timization with QPBO and BP. The effectiveness of our method has been validated on real-world CCTV corpus. Index Terms: story co-segmentation, foreground and background story modeling, lexical clustering, MRF, QP-
Wei Feng 0005, Xuecheng Nie, Lei Xie 0001, Jianmin Jiang
INTERSPEECH1
2011 Pitch-density-based features and an SVM binary tree approach for multi-class audio classification in broadcast news
Lei Xie 0001, Zhong-Hua Fu, Wei Feng 0005
Multim. Syst.3
2010 Maximum lexical cohesion for fine-grained news story segmentation
abstract
We propose a maximum lexical cohesion (MLC) approach to news story segmentation. Unlike sentence-dependent lexical methods, our approach is able to detect story boundaries at finer word/subword granularity, and thus is more suitable for speech recognition transcripts which have no sentence delimiters. The proposed segmentation goodness measure takes account of both lexical cohesion and a prior preference of story length. We mea-sure the lexical cohesion of a segment by the KL-divergence from its word distribution to an associated piecewise uniform distribution. Taking account of the uneven contributions of dif-ferent words to a story, the cohesion measure is further refined by two word weighting schemes, i.e. the inverse document fre-quency (IDF) and a new weighting method called difference from expectation (DFE). We then propose a dynamic program-ming solution to exactly maximize the segmentation goodness and efficiently locate story boundaries in polynomial time. Ex-perimental results show that our MLC approach outperforms several state-of-the-art lexical methods. Index Terms: story segmentation, KL-divergence, lexical co-hesion, word weighting, dynamic programming, spoken docu-ment segmentation, spoken document retrieval 1.
Lei Xie 0001, Wei Feng 0005
INTERSPEECH3
2010 Cascade Markov random fields for stroke extraction of Chinese characters
Wei Feng 0005, Lei Xie 0001
Inf. Sci.2
2010 Self-Validated Labeling of Markov Random Fields for Image Segmentation
abstract
This paper addresses the problem of self-validated labeling of Markov random fields (MRFs), namely to optimize an MRF with unknown number of labels. We present graduated graph cuts (GGC), a new technique that extends the binary s-t graph cut for self-validated labeling. Specifically, we use the split-and-merge strategy to decompose the complex problem to a series of tractable subproblems. In terms of Gibbs energy minimization, a suboptimal labeling is gradually obtained based upon a set of cluster-level operations. By using different optimization structures, we propose three practical algorithms: tree-structured graph cuts (TSGC), net-structured graph cuts (NSGC), and hierarchical graph cuts (HGC). In contrast to previous methods, the proposed algorithms can automatically determine the number of labels, properly balance the labeling accuracy, spatial coherence, and the labeling cost (i.e., the number of labels), and are computationally efficient, independent to initialization, and able to converge to good local minima of the objective energy function. We apply the proposed algorithms to natural image segmentation. Experimental results show that our algorithms produce generally feasible segmentations for benchmark data sets, and outperform alternative methods in terms of robustness to noise, speed, and preservation of soft boundaries.
Wei Feng 0005, Jiaya Jia
IEEE Trans. Pattern Anal. Mach. Intell.1
2009 Multicue Graph Mincut for Image Segmentation
Wei Feng 0005, Lei Xie 0001
ACCV (2)1
2009 Minimizing sparse higher order energy functions of discrete variables
abstract
Higher order energy functions have the ability to encode high level structural dependencies between pixels, which have been shown to be extremely powerful for image labeling problems. Their use, however, is severely hampered in practice by the intractable complexity of representing and minimizing such functions. We observed that higher order functions encountered in computer vision are very often “sparse”, i.e. many labelings of a higher order clique are equally unlikely and hence have the same high cost. In this paper, we address the problem of minimizing such sparse higher order energy functions. Our method works by transforming the problem into an equivalent quadratic function minimization problem. The resulting quadratic function can be minimized using popular message passing or graph cut based algorithms for MAP inference. Although this is primarily a theoretical paper, it also shows how higher order functions can be used to obtain impressive results for the binary texture restoration problem.
Carsten Rother, Pushmeet Kohli, Wei Feng 0005, Jiaya Jia
CVPR3
2008 Perceptual image preview
Wei Feng 0005, Zhouchen Lin, Tien-Tsin Wong
Multim. Syst.2
2008 Region-Level Image Authentication Using Bayesian Structural Content Abstraction
abstract
Image authentication (IA) verifies the integrity of image content by detecting malicious modifications. A good IA system should be able to tolerate noncontent-changing operations (NCOs) robustly, and detect content-changing operations (COs) sensitively. Most existing IA methods realize either bit-level or pixel-level authentication; thus, they can tolerate only particular and limited kinds of NCOs. In this paper, we propose an unsupervised region-level IA scheme named Bayesian structural content abstraction (BaSCA), which is capable of tolerating a wide and dynamic range of NCOs and can sensitively detect real COs. We model image structural content using the net-structured Markov Pixon random field (NS-MPRF), from which we derive the size-controllable BaSCA signature. Furthermore, to support dynamic NCO/CO partition, we present an analogous mean-shift algorithm to iteratively optimize the BaSCA signature in the user-defined NCO space. Both theoretical analysis and experimental results demonstrate that our BaSCA scheme has much less false positive and comparable false negative probability, as compared to state-of-the-art IA methods.
Wei Feng 0005
IEEE Trans. Image Process.1
2006 Spectral Multiplicity Tolerant Inexact Graph Matching
abstract
A graph can be exactly specified by the spectrum and corresponding eigenvectors of its adjacency matrix. This provides a solid foundation for spectrum based graph matching. However, most previous methods ignore the spectral multiplicity, which may significantly affect the matching accuracy. In this paper, we address the problem of minimizing the matching error when graph spectral multiplicity is involved. We first model spectral multiplicity by the sub-eigenspace rotation matrix R, and integrate R into the spectrum based graph matching model. We then focus on the exact graph matching problem, and show how to establish the vertex-to-vertex correspondence by iteratively optimizing the sub-eigenspace rotation matrix R and the permutation matrix P. A reliable matching initialization method is proposed to make this process converge rapidly. Finally, we extend the approach to the inexact graph matching problem by optimally warping two graphs to the same size. The proposed approach is robust and efficient. We support our approach with numerical experiments and demonstrate its effectiveness in the practical application of uncalibrated stereo matching.
Wei Feng 0005
SMC1
2005 Bayesian Structural Content Abstraction for Region-Level Image Authentication
abstract
We present a hierarchical representation of image structure and use it for image content authentication. Firstly, we model the image with the Markov pixon random field. Within the Bayesian framework, the optimal label map and regional pixon map can be obtained, based on which we define an undirected graph, namely Bayesian structural content abstraction (BaSCA). This representation captures the spatial topology information of homogeneous regions as well as their finest scale and interactions. Then, an efficient optimization scheme has been proposed to iteratively minimize the distance (or learning error) to all content-identical image samples generated by an acceptable operation set defined by the user. In addition, we use the regional pixon map to remove spurious vertices and thus to establish a BaSCA hierarchy naturally The BaSCA itself and its features can act as the signature of the protected image. Our experimental results show that the proposed approach has much less false positive and comparable false negative probability compared with the existing methods.
Wei Feng 0005
ICCV1