Dan Zeng 0001

dblp:06/6575-1 · DBLP profile ↗
← Back
115ranked-venue papers
8as first author
89since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 74 · 2 first-author · 58 since 2021Artificial intelligence and machine learning · 36 · 5 first-author · 27 since 2021Computer networks · 11 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 9 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Beyond Predictive Resampling: Learning Input-Agnostic Downsampling for Efficient Aligned Vision Recognition
abstract
Images are typically sampled on a uniform grid,despite their non-uniform information distribution—some regions are rich in content while others are not. The mismatch leads to inefficient computation allocation in deep learning models. To address this, recent studies have proposed predictive downsampling methodsthat adaptively downsample images based on predicted per-pixel importance, allocating more pixels to informative areas. However,these methods require high-resolution processing to accurately estimate importance, which undermines their efficiency:the prediction itself must process the full-resolution image,consuming most of the computational budget. This high-resolution importance prediction is necessary because each input may differ significantly in structure and content. In this paper, we take a different approach and introduce a learn-to-downsample paradigmtailored for aligned vision recognition tasks, such as face recognition and palmprint recognition, where input alignment ensures consistent spatial structure across images. This alignment ensures structural consistency across images, allowing a shared, input-agnostic downsampling template applicable to all inputs. Furthermore, instead of relying on implicit importance maps, we introduce a flow-based representation that explicitly models the spatial warping from the original image to the downsampled version. The flow representation is not only more efficient but also more controllable: we regularize the flow using its Jacobian determinant to precisely control the sampling density and coverage,enabling interpretable and tunable sampling patterns. Extensive experiments on two aligned recognition tasks, face and palmprint recognition, demonstrate that our method substantially reduces computational cost with minimal accuracy degradation, achieving a significantly better performance-efficiency trade-off than existing predictive downsampling methods.
Kai Zhao 0012, Liting Ruan, Xiaoqiang Zhu, Xianchao Zhang 0002, Dan Zeng 0001
AAAI6
2026 Open-Vocabulary Camouflaged Object Segmentation with Cascaded Vision Language Models
abstract
Open-vocabulary camouflaged object segmentation (OVCOS) seeks to segment and classify camouflaged objects in arbitrary categories, presenting unique challenges due to visual ambiguity and unseen categories. Recent approaches typically adopt a two-stage paradigm: they first segment objects, and then classify the segmented regions using vision language models (VLMs). However, such methods (i) suffer from a domain gap caused by the mismatch between VLMs' full-image training and cropped-region inferencing, and (ii) depend on generic segmentation models optimized for well-delineated objects which are less effective for camouflaged objects. Without explicit guidance, generic segmentation models often overlook subtle boundaries, leading to imprecise segmentation. In this paper, we introduce a novel VLM-guided cascaded framework to address these issues in OVCOS. For segmentation, we leverage the segment anything model (SAM), guided by the VLM. Our framework uses VLM-derived features as explicit prompts to SAM, effectively directing attention to camouflaged regions and significantly improving localization accuracy. For classification, we avoid the domain gap introduced by hard cropping. Instead, we treat the segmentation output as a soft spatial prior using the alpha channel. This retains the full image context while providing precise spatial guidance, leading to more accurate and context-aware classification of camouflaged objects. The same VLM is shared between segmentation and classification to ensure efficiency and semantic consistency. Extensive experiments on both OVCOS and conventional camouflaged object segmentation benchmarks demonstrate the clear superiority of our method, highlighting the effectiveness of leveraging rich VLM semantics for both segmentation and classification of camouflaged objects. Our code and models are open-sourced at https://github.com/intcomp/camouflaged-vlm.
Kai Zhao 0012, Wubang Yuan, Zheng Wang 0059, Guanyi Li, Xiaoqiang Zhu, Deng-Ping Fan, Dan Zeng 0001
Comput. Vis. Media7
2026 Few-shot Class-Incremental Learning via Generative Co-Memory Regularization
Kexin Bao, Dan Zeng 0001, Shiming Ge
Int. J. Comput. Vis.3
2026 IPDiff: Diffusion-driven ORSI Salient Object Detection with Information Reconstruction and Multi-Prior Guidance
Gongyang Li, Zhen Bai 0001, Runmin Cong, Dan Zeng 0001, Weisi Lin
Int. J. Comput. Vis.4
2026 Few-shot medical image segmentation via CLIP-driven dual contrastive learning
Meiqi Zou, Wubang Yuan, Yiqing Liu 0001, Dan Zeng 0001
Neurocomputing4
2026 Learning-Based Resource Allocation for Integrated Sensing, Communication, and Computation Networks: A Delay-Aware Approach
abstract
Integrated sensing, communication, and computation (ISCC) network has been recognized as a key enabler to realize the vision of Internet-of-Things. In this paper, we explore the resource allocation problem in ISCC networks, where the task execution workflow consists of multiple dependent processes, i.e., wireless sensing, signal processing, data delivery, and data processing. To this end, a tandem-parallel queuing model is first proposed to characterize the end-to-end (E2E) task execution process. Given the model, the E2E delay upper bound is derived according to the stochastic network calculus theory. Based on the analytical results, the joint allocation problem of the sensing, communication, and computation (SCC) resources is formulated to minimize the E2E delay while satisfying the constraints of network resources, tolerable delay, and sensing mutual information, etc. Further, this non-convex optimization problem is parameterized to enable a learning-based optimization approach. Next, we design the unsupervised learning (UL) framework based on multilevel decomposition architecture (MDA) and residual network (RN) to accelerate training speed and ensure effective primal-dual learning. Numerical results demonstrate that the proposed UL-MDA-RN framework is superior to existing baselines with excellent convergence efficiency and lower achieved E2E delay. In addition, our results analyze the impacts of the network parameters on the E2E delay performance to guide the design of appropriate SCC resource provisioning patterns.
Mengxin Yang, Yixiao Gu, Han Hu 0003, Dan Zeng 0001
IEEE Internet Things J.4
2026 Hierarchical prior-guided and channel-wise adaptation fusion network for RGB-D railway surface defect inspection
Jianlin Chen, Gongyang Li, Zhijiang Zhang, Dan Zeng 0001
J. Vis. Commun. Image Represent.4
2026 MIST: A Benchmark and Baseline for Multi-Frame Infrared Small Target Detection in Complex Motion
abstract
Motion cues play a vital role in multi-frame infrared small target detection (MISTD). However, most targets in existing datasets exhibit regular and slow motion, which cannot reflect the complex and diverse motion patterns in real-world scenarios. This biased data distribution makes recent data-driven methods highly rely on simplified motion assumptions that tend to fail in irregular or fast motion, resulting in noisy feature representations cluttered with target-irrelevant factors. Hence, we stress that methods for MISTD should also work when targets are in complex motion. To enable this research, we propose a large-scale dataset called MIST for airborne infrared detection scenarios. The dataset is built on a synthetic data engine that models variations in pose, size, and intensity of moving targets while seamlessly blending them into real backgrounds for physical, geometric, and visual realism. Targets in MIST exhibit low signal-to-clutter ratios and complex motion, making it a promising yet challenging benchmark for developing algorithms focused on motion analysis. To tackle the challenges of MIST, we develop MISTNet, a robust baseline based on the Information Bottleneck theory. To handle irregular and fast motion, we propose a shifted neighborhood compensation block to efficiently model multi-scale correspondences for implicit motion compensation. To distill compact representations free from irrelevant cues, we design a progressive distillation decoder to hierarchically filter out redundancy while preserving target-relevant information. We benchmark 31 state-of-the-art methods and find that their performance on MIST drops significantly compared with that on the widely used NUDT-MIRSDT dataset. Our MISTNet outperforms all other methods by a large margin, with an over 6% gain in the IoU metric, demonstrating its superiority. The dataset, code, and model weights are available at https://github.com/GR-ray/MIST.
Meihong Zhang, Gongyang Li, Guanyi Li, Kai Zhao 0012, Xianchao Zhang 0002, Dan Zeng 0001
IEEE Trans. Image Process.7
2026 ASDTracker: Adaptively Sparse Detection With Attention-Guided Refinement for Efficient Multi-Object Tracking
abstract
Tracking-by-Detection paradigms shine in generic multi-object tracking (MOT), while their compact construction hinders the real-time applications. In this work, we attribute the substantial computational burden to two expensive components, i.e. detection and re-identification. Building upon the principle of adaptively maintaining acceptable inference efficiency, we present Adaptively Sparse Detection with attention-guided refinement (ASDTracker) for efficient tracking. In specific, our ASDTracker rapidly assess the short-term and long-term occlusion, dynamically determining the usage of the expensive detector. For non-key frames, we efficiently refine small-size crops out of Kalman Filter predictions and introduce the noisy shadow labels to robustly train this refinement network. Additionally, we substitute the lightweight appearance representation for the heavy ReID network, which efficiently extracts sufficient appearance cues in the coarsely quantized color spaces. Extensive experiments on four benchmarks demonstrate that ASDTracker achieves competitive performance in generalization and robustness under favorable inference speed. Moreover, the efficient tracking deployment is further implemented to an unmanned surface vehicle with high accuracy and low latency in real-world scenarios.
Yueying Wang, Chenyang Yan, Cairong Zhao, Weidong Zhang 0004, Dan Zeng 0001
IEEE Trans. Image Process.5
2026 Learning Compact Representations With an Information Bottleneck for Camouflaged Object Detection
abstract
Frequency domain-based methods have demonstrated promising performance in Camouflaged Object Detection (COD) tasks because of their enhanced power for distinguishing between objects and the background in the frequency domain. However, these methods often overlook the interference caused by task-irrelevant cues such as background textures. These extraneous factors are learned alongside task-relevant features by the employed network, increasing the number of false positives. Therefore, we propose a camouflaged object detection method based on the Information Bottleneck (IB) theory. The aim is to obtain a robust representation that retains the essential features needed for prediction while minimizing the redundant information derived from both the RGB and frequency domains. Specifically, we propose a Feature Selection Information Bottleneck Module (FSIBM). By explicit supervision, this module minimizes the mutual information between the fused feature from two domains and the predictive features, thereby weakening task-irrelated information. Simultaneously, the FSIBM maximizes the mutual information between the predictive features and the ground truth (i.e., emphasizing task-related elements). Additionally, we introduce a Cross-Domain Awareness Interaction Module (CDAIM), which establishes self-reinforcement for the object attributes within each domain and facilitates cross-domain complementarity. This enables the capture of sufficient discriminative features from both domains. To verify the generalization ability of the proposed method, we applied it to three benchmark datasets, on which our method outperformed the corresponding state-of-the-art methods. Our code is released athttps://github.com/KwunYat/CODIB.
Guanyi Li, Junjie Zhang 0002, Wubang Yuan, Gloria Jin, Dan Zeng 0001
IEEE Trans. Multim.6
2026 EEformer: Early Exiting for Transformer With Global-Local Exits and Progressive Fine-Tuning
abstract
Recently, the efficient deployment and acceleration of transformer-based pre-trained models (TPMs) on resource-constrained edge devices for multimedia services have gained significant interest. Although early exiting is a feasible solution, it may lead to extra computational cost and substantial performance degradation compared to the original models. To tackle these issues, we propose a framework termed EEformer, which incorporates global-local heads (GLHs) into intermediate layers to construct the early exiting dynamic neural network (EDNN). The GLH can efficiently extract global and local information from hidden states produced by the backbone layer, thereby achieving a better performance-efficiency trade-off for the EDNN. Moreover, we propose a novel progressive fine-tuning strategy to steadily improve the efficiency of the EDNN while maintaining its performance comparable to the original mode through three fine-tuning stages. We conduct extensive experiments on image classification and natural language processing tasks, demonstrating the superiority of the proposed framework. In particular, the proposed framework achieves 1.87× speed-up while maintaining 99.0% performance on the CIFAR-100 dataset, and 3.05× speed-up while maintaining 98.5% performance on the SST-2 dataset.
Guanyu Xu, Yong Luo 0002, Li Shen 0008, Han Hu 0003, Dan Zeng 0001
IEEE Trans. Multim.6
2026 COP: CrOss-View Attention Prompt for Zero-Shot Sketch-Based Image Retrieval
abstract
Zero-shot Sketch-based Image Retrieval (ZS-SBIR) is a challenging yet rewarding task, as it demands models to possess both brain- like zero-shot learning and cross-view alignment capabilities. Recent advances suggest that powerful pre-trained vision encoders, such as CLIP, offer a promising alternative for addressing the ZS-SBIR task. However, the problem of simultaneously evoking the zero-shot learning capability and cross-view alignment capability of pre-trained vision encoders has barely been discussed. To this end, we propose the CrOss-view Attention Prompt (COP) framework, which is composed of an Attention Prompt module and a Cross-view Query module. Specifically, we formulate prompt construction as a retrieval problem by introducing a prompt pool and attention mechanism, thereby constructing attention prompts with fine granularity to enhance the zero-shot learning capability. Furthermore, to endow COP with cross-view alignment capabilities, we replace single-view queries with carefully designed cross-view queries, which can be smoothly inserted into the Attention Prompt module. The proposed COP is scenario-agnostic and supports vision encoders with diverse pre-training schemes. Comprehensive experiments show that COP achieves competitive performance in ZS-SBIR, Generalized ZS-SBIR, and Cross-data ZS-SBIR scenarios, regardless of whether it is based on the ImageNet pre-trained vision encoder or the CLIP pre-trained vision encoder.
Jiahao Zheng 0001, Yongcan Luo, Ning Chen 0008, Dan Zeng 0001, Dapeng Oliver Wu
IEEE Trans. Multim.5
2026 MCFINet: A Cost-Efficient Multi-Channel Feature Integration Network for Surface Scenarios Image Super-Resolution
abstract
Convolutional Neural Network (CNN) and Vision Transformer (ViT) have revolutionized the field of image super-resolution (SR). However, their complexity poses challenges for resource—constrained scenarios, particularly due to the high computational demands of Transformers and their excessive reliance on global information. To tackle these challenges, we propose a Multi-Channel Feature Integration Network (MCFINet), designed to maximize input pixel utilization while minimizing computational overhead. It integrates both local and global features within the channels, thereby exploiting their complementary advantages. First, the designed Feature Integration Block (FIB) effectively captures local information and improves visual quality by enhancing the mapping of non-local features. Subsequently, we utilize the Adaptive Channel Fusion Block (ACFB), which strengthens the interaction between features and channels while maintaining computational efficiency. Finally, for SR task on resource-constrained surface scenarios, we propose a more suitable pre-training method, which further boosts the model’s learning ability. Evaluation results indicate that the proposed MCFINet achieves a better balance between lightweight design and high-quality restoration on both standard evaluation datasets and water surface target datasets. Specifically, compared to the traditional SwinIR-L, MCFINet reduces model training time and runtime by 12% on the test set, while also decreasing model complexity by 43%. Our codes are available at https://github.com/Lcasjz/MCFINet .
Liangcheng Zhao, Yueying Wang, Yuhao Qing, Dan Zeng 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2026 Modeling and Performance Analysis for Clustered Integrated Sensing and Communication Networks
abstract
The stochastic geometry-based modeling and analysis of large-scale Integrated Sensing and Communication (ISAC) networks are vital for providing useful ISAC design insights. One important ISAC network characteristic is the sensing and communication (S&C) spatial correlations since the communication users (CUs) are more interested in the sensing target (STs) around them and the base stations prefer to utilize one ISAC signal to serve the CUs and STs close to each other to enable effective S&C coverage. However, most existing works focused on ISAC systems where the locations of CUs and STs are assumed to be independent. This paper bridges this gap by proposing an analytical framework for the ISAC networks where the unified ISAC waveform is modeled with limited main lobe beamwidth and the CUs and STs served by one signal are assumed to be correlatively distributed in spatial domain. Given the model, we first derive some prerequisite auxiliary quantities (i.e., the probability that an ST is served by the main lobe or side lobe, the link distance distribution, etc.) to analyze the network characteristics. Further, the communication/sensing coverage probability, as well as the joint and conditional ISAC coverage probability are analyzed to provide the comprehensive analysis results. Combining the theoretical analysis and simulation results, it verifies the accuracy of the analytical framework and quantifies the impact of the network parameters and spatial correlations on the S&C performance. Moreover, our results reveal how the sensing performance and communication performance are mutually restricted to describe the tradeoffs of S&C performance.
Yixiao Gu, Yinghong Guo, Bin Xia 0001, Dan Zeng 0001
IEEE Trans. Wirel. Commun.5
2026 Cooperative ISAC Systems With Extended Targets: Performance Analysis and Beamforming Design
abstract
This paper investigates a cooperative integrated sensing and communication (ISAC) system, where multiple base stations (BSs) employ coordinated transmit beamforming to communicate with their respective users and jointly sense a set of extended targets (ETs). Different from prior cooperative ISAC works considering point targets (PTs), the visible scatterers on the same extended target (ET) and the radar cross section (RCS) of the same scatterers observed by multiple BSs are considered to be different. Given the model, we first derive the Cramér-Rao bound (CRB) for the BSs to cooperatively estimate the ET’s parameters, thus quantifying the cooperative sensing gains. Next, based on the derived CRB, we formulate a joint node selection and coordinated transmit beamforming design problem with the object of minimizing the average trace of sensing CRB, while satisfying the minimum communication rate constraint, the maximum transmit power constraint, and the node selection constraints. To solve this non-convex optimization problem, we first utilize the block coordinate descent (BCD) method to decompose it into node selection sub-problem and beamforming design sub-problem. Next, the continuous relaxation and linear programming (LP) approach are employed to handle the node selection sub-problem, and a decentralized augmented Lagrangian manifold optimization algorithm is developed to solve the beamforming sub-problem with reduced computation complexity. Numerical simulations demonstrate that the proposed design outperforms benchmark designs with larger CRB-rate region. Moreover, our results show the impacts of the ET’s state and the number of network nodes on network performance to enable valuable ISAC beamforming design insights.
Yixiao Gu, Han Hu 0003, Jie Xu 0002, Dan Zeng 0001
IEEE Trans. Wirel. Commun.5
2026 CRB-Rate Bound and Bound-Achieving Inputs for ISAC Systems With Amplitude Constraints
Yinghong Guo, Yixiao Gu, Dan Zeng 0001, Bin Xia 0001
IEEE Trans. Wirel. Commun.4
2025 Integrating Low-Level Visual Cues for Enhanced Unsupervised Semantic Segmentation
abstract
Unsupervised semantic segmentation algorithms aim to identify meaningful semantic groups without annotations. Recent approaches leveraging self-supervised transformers as pre-training backbones have successfully obtained high-level dense features that effectively express semantic coherence. However, these methods often overlook local semantic coherence and low-level features such as color and texture. We propose integrating low-level visual cues to complement high-level visual cues derived from self-supervised pre-training branches. Our findings indicate that low-level visual cues provide a more coherent recognition of color-texture aspects, ensuring the continuity of spatial structures within classes. This insight led us to develop IL2Vseg, an unsupervised semantic segmentation method that leverages the complementation of low-level visual cues. The core of IL2Vseg is a spatially-constrained fuzzy clustering algorithm based on color affinities, which preserves the intra-class affinity of spatially-adjacent and similarly-colored pixels in low-level visual cues. Additionally, to effectively couple low-level and high-level visual cues, we introduce a feature similarity loss function to optimize the feature representation of fused visual cues. To further enhance consistent feature learning, we incorporate contrast loss functions based on color invariance and luminosity invariance, which improve the learning of features from different semantic categories. Extensive experiments on multiple datasets, including COCO-Stuff-27, Cityscapes, Potsdam, and MaSTr1325, demonstrate that IL2Vseg achieves state-of-the-art results.
Yuhao Qing, Dan Zeng 0001, Shaorong Xie, Kaer Huang, Yueying Wang
AAAI2
2025 3CAD: A Large-Scale Real-World 3C Product Dataset for Unsupervised Anomaly Detection
abstract
Industrial anomaly detection achieves progress thanks to datasets such as MVTec-AD and VisA. However, they suffer from limitations in terms of the number of defect samples, types of defects, and availability of real-world scenes. These constraints inhibit researchers from further exploring the performance of industrial detection with higher accuracy. To this end, we propose a new large-scale anomaly detection dataset called 3CAD, which is derived from real 3C production lines. Specifically, the proposed 3CAD includes eight different types of manufactured parts, totaling 27,039 high-resolution images labeled with pixel-level anomalies. The key features of 3CAD are that it covers anomalous regions of different sizes, multiple anomaly types, and the possibility of multiple anomalous regions and multiple anomaly types per anomaly image. This is the largest and first anomaly detection dataset dedicated to 3C product quality control for community exploration and development. Meanwhile, we introduce a simple yet effective framework for unsupervised anomaly detection: a Coarse-to-Fine detection paradigm with Recovery Guidance (CFRG). To detect small defect anomalies, the proposed CFRG utilizes a coarse-to-fine detection paradigm. Specifically, we utilize a heterogeneous distillation model for coarse localization and then fine localization through a segmentation model. In addition, to better capture normal patterns, we introduce recovery features as guidance. Finally, we report the results of our CFRG framework and popular anomaly detection methods on the 3CAD dataset, demonstrating strong competitiveness and providing a highly challenging benchmark to promote the development of the anomaly detection field.
Enquan Yang, Peng Xing, Hanyang Sun, Wenbo Guo 0017, Yuanwei Ma, Zechao Li, Dan Zeng 0001
AAAI7
2025 Divide and Conquer: Static-Dynamic Collaboration for Few-Shot Class-Incremental Learning
abstract
Continual learning systems suffer from catastrophic forgetting, where updates for new tasks destructively interfere with previously acquired knowledge. Recent empirical advances—including flatness-based optimization, static–dynamic architectural decomposition, and probabilistic reg- ularization— have demonstrated strong mitigation of forgetting. However, a unified structural explanation for why these methods succeed remains underdeveloped. This paper proposes a constraint geometry perspective on representation updates in continual learning. We argue that catastrophic forgetting can be interpreted as a curvature-induced vio- lation of constraint-preserving update dynamics. Under this view, successful continual learning methods implicitly regulate update directions in high-curvature regions of the loss landscape. Rather than introducing a new algorithm, this work provides a structural interpretation that clarifies why diverse empirical strategies succeed. Identifying and preserving geometric constraints during gradient-based updates may serve as a guiding principle for future continual learning research.
Kexin Bao, Daichi Zhang, Dan Zeng 0001, Shiming Ge
ICMR4
2025 MixPrompt: Efficient Mixed Prompting for Multimodal Semantic Segmentation
abstract
Recent advances in multimodal semantic segmentation show that incorporating auxiliary inputs—such as depth or thermal images—can significantly improve performance over single-modality (RGB-only) approaches. However, most existing solutions rely on parallel backbone networks and complex fusion modules, greatly increasing model size and computational demands. Inspired by prompt tuning in large language models, we introduce \textbf{MixPrompt}: a prompting-based framework that integrates auxiliary modalities into a pretrained RGB segmentation model without modifying its architecture. MixPrompt uses a lightweight prompting module to extract and fuse information from auxiliary inputs into the main RGB backbone. This module is initialized using the early layers of a pretrained RGB feature extractor, ensuring a strong starting point. At each backbone layer, MixPrompt aligns RGB and auxiliary features in multiple low-rank subspaces, maximizing information use with minimal parameter overhead. An information mixing scheme enables cross-subspace interaction for further performance gains. During training, only the prompting module and segmentation head are updated, keeping the RGB backbone frozen for parameter efficiency. Experiments across NYU Depth V2, SUN-RGBD, MFNet, and DELIVER datasets show that MixPrompt achieves improvements of 4.3, 1.1, 0.4, and 1.1 mIoU, respectively, over two-branch baselines, while using nearly half the parameters. MixPrompt also outperforms recent prompting-based methods under similar compute budgets.
Zhiwei Hao 0001, Zhongyu Xiao, Jianyuan Guo, Li Shen 0008, Yong Luo 0002, Han Hu 0003, Dan Zeng 0001
NeurIPS7
2025 A User-Centric Cooperative Offloading Scheme for Stochastic MEC Networks
abstract
The modeling and analysis of large-scale stochastic MEC systems are of great significance in providing useful design guidelines for practical MEC networks. In most prior works, the users generally adopt the same strategy to select appropriate MEC access points (MAPs), where the fact that the available mobile computing services are different among the randomly distributed users is ignored. To this end, this paper proposes a user-centric cooperative offloading scheme to enable more flexible and efficient user task offloading. Specifically, each user can be served by one or two MAPs based on both the communication performance and computing performance. To evaluate the performance gains acquired from the proposed task offloading scheme, we first derive the service mode assignment probability, link distance distribution, and interference intensity to capture the network characteristics. Further, the moment and the meta distribution of the task transmission performance are analyzed. Based on the above results, we focus on the distribution of the computation workload to evaluate the edge computing service capability. From the analytical and simulation results, it is demonstrated that compared with the non-cooperative offloading scheme, the proposed task offloading scheme not only achieves more reliable and fair task offloading but also increases the computing service capacity.
Yixiao Gu, Dan Zeng 0001, Yinghong Guo, Bin Xia 0001, Zhiyong Chen 0002, Jiangzhou Wang
IEEE Internet Things J.2
2025 Learning Instructive Frequency Spectral and Curvature Features for Cloud Detection
abstract
Current cloud detection methods often treat all spectral bands equally, which limits their ability to capture instructive clues necessary for accurate detection. As a result, distinguishing clouds from snow in coexisting environments remains challenging. Moreover, most approaches struggle to adaptively model the boundaries of clouds, which is crucial for detecting thin clouds with ambiguous edges. To address these challenges, we propose a novel approach for cloud detection called FSCFNet, which captures guiding visual features from frequency and curvature computations. FSCFNet comprises two key modules: the Frequency Spectral Feature Enhancement Module (FSFEM) and the Curvature-based Edge-Awareness Module (CEAM). The FSFEM leverages the distinct characteristics of spectral bands to extract instructive visual cues, enabling the network to learn robust discriminative features for ice, snow, and clouds. In contrast, the CEAM adaptively identifies texture-rich regions using curvature, enhancing the ability to delineate thin cloud boundaries. Comprehensive quantitative and qualitative experiments on the Landsat 8 and MODIS datasets demonstrate that FSCFNet consistently outperforms state-of-the-art methods. Our code is publicly available at https://github.com/wanjuanhu/FSCFNet/tree/main.
Wanjuan Hu, Guanyi Li, Guoguo Zhang, Dan Zeng 0001
IEEE Geosci. Remote. Sens. Lett.5
2025 PKI: Prior knowledge-infused neural network for few-shot class-incremental learning
Kexin Bao, Fanzhao Lin, Dan Zeng 0001, Shiming Ge
Neural Networks5
2025 3DBench: A scalable benchmark for object and scene-level instruction-tuning of 3D large language models
Tianci Hu, Junjie Zhang 0002, Yutao Rao, Dan Zeng 0001, Hongwen Yu, Xiaoshui Huang
Neural Networks4
2025 DiffusionAD: Norm-Guided One-Step Denoising Diffusion for Anomaly Detection
abstract
Anomaly detection has garnered extensive applications in real industrial manufacturing due to its remarkable effectiveness and efficiency. However, previous generative-based models have been limited by suboptimal reconstruction quality, hampering their overall performance. We introduce DiffusionAD, a novel anomaly detection pipeline comprising a reconstruction sub-network and a segmentation sub-network. A fundamental enhancement lies in our reformulation of the reconstruction process using a diffusion model into a noise-to-norm paradigm. Here, the anomalous region loses its distinctive features after being disturbed by Gaussian noise and is subsequently reconstructed into an anomaly-free one. Afterward, the segmentation sub-network predicts pixel-level anomaly scores based on the similarities and discrepancies between the input image and its anomaly-free reconstruction. Additionally, given the substantial decrease in inference speed due to the iterative denoising nature of diffusion models, we revisit the denoising process and introduce a rapid one-step denoising paradigm. This paradigm achieves hundreds of times acceleration while preserving comparable reconstruction quality. Furthermore, considering the diversity in the manifestation of anomalies, we propose a norm-guided paradigm to integrate the benefits of multiple noise scales, enhancing the fidelity of reconstructions. Comprehensive evaluations on four standard and challenging benchmarks reveal that DiffusionAD outperforms current state-of-the-art approaches and achieves comparable inference speed, demonstrating the effectiveness and broad applicability of the proposed pipeline.
Hui Zhang 0090, Zheng Wang 0059, Dan Zeng 0001, Zuxuan Wu, Yu-Gang Jiang 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 DSENet++: A Coarse-to-Fine Framework for Enhanced Sub-Region Detection in Aerial Images
Xiangjie Wang, Liang Chen 0004, Junjie Zhang 0002, Jian Zhang 0002, Shiming Ge, Dan Zeng 0001
IEEE Trans. Multim.7
2025 EMS: A Large-Scale Eye Movement Dataset, Benchmark, and New Model for Schizophrenia Recognition
abstract
Schizophrenia (SZ) is a common and disabling mental illness, and most patients encounter cognitive deficits. The eye-tracking technology has been increasingly used to characterize cognitive deficits for its reasonable time and economic costs. However, there is no large-scale and publicly available eye movement dataset and benchmark for SZ recognition. To address these issues, we release a large-scale Eye Movement dataset for SZ recognition (EMS), which consists of eye movement data from 104 schizophrenics and 104 healthy controls (HCs) based on the free-viewing paradigm with 100 stimuli. We also conduct the first comprehensive benchmark, which has been absent for a long time in this field, to compare the related 13 psychosis recognition methods using six metrics. Besides, we propose a novel mean-shift-based network (MSNet) for eye movement-based SZ recognition, which elaborately combines the mean shift algorithm with convolution to extract the cluster center as the subject feature. In MSNet, first, a stimulus feature branch (SFB) is adopted to enhance each stimulus feature with similar information from all stimulus features, and then, the cluster center branch (CCB) is utilized to generate the cluster center as subject feature and update it by the mean shift vector. The performance of our MSNet is superior to prior contenders, thus, it can act as a powerful baseline to advance subsequent study. To pave the road in this research field, the EMS dataset, the benchmark results, and the code of MSNet are publicly available at https://github.com/YingjieSong1/EMS.
Zhi Liu 0003, Gongyang Li, Qiang Wu 0001, Dan Zeng 0001, Lihua Xu, Tianhong Zhang, Jijun Wang 0003
IEEE Trans. Neural Networks Learn. Syst.6
2025 Precision in pursuit: a multi-consistency joint approach for infrared anti-UAV tracking
Junjie Zhang 0002, Pangrong Shi, Xiaoqiang Zhu, Dan Zeng 0001
Vis. Comput.6
2024 Coupled Confusion Correction: Learning from Crowds with Sparse Annotations
abstract
As the size of the datasets getting larger, accurately annotating such datasets is becoming more impractical due to the expensiveness on both time and economy. Therefore, crowd-sourcing has been widely adopted to alleviate the cost of collecting labels, which also inevitably introduces label noise and eventually degrades the performance of the model. To learn from crowd-sourcing annotations, modeling the expertise of each annotator is a common but challenging paradigm, because the annotations collected by crowd-sourcing are usually highly-sparse. To alleviate this problem, we propose Coupled Confusion Correction (CCC), where two models are simultaneously trained to correct the confusion matrices learned by each other. Via bi-level optimization, the confusion matrices learned by one model can be corrected by the distilled data from the other. Moreover, we cluster the ``annotator groups'' who share similar expertise so that their confusion matrices could be corrected together. In this way, the expertise of the annotators, especially of those who provide seldom labels, could be better captured. Remarkably, we point out that the annotation sparsity not only means the average number of labels is low, but also there are always some annotators who provide very few labels, which is neglected by previous works when constructing synthetic crowd-sourcing annotations. Based on that, we propose to use Beta distribution to control the generation of the crowd-sourcing labels so that the synthetic annotations could be more consistent with the real-world ones. Extensive experiments are conducted on two types of synthetic datasets and three real-world datasets, the results of which demonstrate that CCC significantly outperforms state-of-the-art approaches. Source codes are available at: https://github.com/Hansong-Zhang/CCC.
Hansong Zhang 0003, Shikun Li, Dan Zeng 0001, Chenggang Yan 0001, Shiming Ge
AAAI3
2024 M3D: Dataset Condensation by Minimizing Maximum Mean Discrepancy
abstract
Training state-of-the-art (SOTA) deep models often requires extensive data, resulting in substantial training and storage costs. To address these challenges, dataset condensation has been developed to learn a small synthetic set that preserves essential information from the original large-scale dataset. Nowadays, optimization-oriented methods have been the primary method in the field of dataset condensation for achieving SOTA results. However, the bi-level optimization process hinders the practical application of such methods to realistic and larger datasets. To enhance condensation efficiency, previous works proposed Distribution-Matching (DM) as an alternative, which significantly reduces the condensation cost. Nonetheless, current DM-based methods still yield less comparable results to SOTA optimization-oriented methods. In this paper, we argue that existing DM-based methods overlook the higher-order alignment of the distributions, which may lead to sub-optimal matching results. Inspired by this, we present a novel DM-based method named M3D for dataset condensation by Minimizing the Maximum Mean Discrepancy between feature representations of the synthetic and real images. By embedding their distributions in a reproducing kernel Hilbert space, we align all orders of moments of the distributions of real and synthetic images, resulting in a more generalized condensed set. Notably, our method even surpasses the SOTA optimization-oriented method IDC on the high-resolution ImageNet dataset. Extensive analysis is conducted to verify the effectiveness of the proposed method. Source codes are available at https://github.com/Hansong-Zhang/M3D.
Hansong Zhang 0003, Shikun Li, Dan Zeng 0001, Shiming Ge
AAAI4
2024 DiffMOT: A Real-time Diffusion-based Multiple Object Tracker with Non-linear Prediction
abstract
In Multiple Object Tracking, objects often exhibit nonlinear motion of acceleration and deceleration, with irregular direction changes. Tacking-by-detection (TBD) trackers with Kalman Filter motion prediction work well in pedestrian-dominant scenarios but fall short in complex situations when multiple objects perform non-linear and diverse motion simultaneously. To tackle the complex non-linear motion, we propose a real-time diffusion-based MOT approach named DiffMOT. Specifically, for the motion predictor component, we propose a novel Decoupled Diffusion based Motion Predictor (D2MP). It models the entire distribution of various motion presented by the data as a whole. It also predicts an individual object's motion conditioning on an individual's historical motion information. Furthermore, it optimizes the diffusion process with much fewer sampling steps. As a MOT tracker, the DiffMOT is real-time at 22.7FPS, and also outperforms the state-of-the-art on DanceTrack[30] and SportsMOT[6] datasets with 62.3% and 76.2% in HOTA metrics, respectively. To the best of our knowledge, DiffMOT is the first to introduce a diffusion probabilistic model into the MOT to tackle non-linear motion prediction.
Weiyi Lv, Yuhang Huang 0006, Ning Zhang 0023, Ruei-Sung Lin, Dan Zeng 0001
CVPR6
2024 DSENet: An Object-Wise Density-Informed Coarse-to-Fine Object Detector for Aerial Image
abstract
Object detection in aerial images remains formidable due to substantial object scale variations, and uneven object distributions. Previous methods widely adopt the coarse-to-fine methodology where detectors focus on large-scale objects coarsely. Sub-regions that contain densely distributed small ones are captured and detected finely. However, two pivotal assessment factors of sub-regions, positional precision, and detection difficulty, deserve further consideration. In this paper, we propose an object-wise density-informed DSENet including consecutive stages termed "Discernment, Selection, Elevation ". Specifically, the sophisticated object-wise density map that considers both object scales and angles, helps discern more positional-precise sub-regions. Then sub-regions with high detection difficulty are selected based on density intensities and coarse detections collaboratively. Finally, the fine detector head instead of the full detector, fine-tuned with selected sub-regions efficiently, elevates what and where coarse detections are mediocre. Extensive experiments show that DSENet achieves state-of-the-art performance on two popular aerial image datasets, VisDrone and DOTA-V1.5.
Xiangjie Wang, Junjie Zhang 0002, Jian Zhang 0002, Dan Zeng 0001
ICME5
2024 Densely Connected Transformer with Frequency Awareness and Sam Guidance for Semi-Supervised Hyperspectral Image Classification
abstract
Advancements in Hyperspectral Image (HSI) spatial resolution pose challenges in pixel-wise classification. Semi-supervised self-training shows potential by using pseudo-labels from unlabeled samples. However, the Hughes phenomenon and environmental factors often lead to spectral variability and undermine pseudo-label credibility. To address above issues, we propose a densely connected Transformer leveraging Discrete Wavelet Transform for extracting nuanced spatial-spectral features and redundancy removal, and we design a filtering strategy guided by the Segment Anything Model (SAM) to retain reliable pseudo labeled samples given the spatial and semantic consistency of HSI regions. Experiments show promising performance of proposed model on high-resolution HSIs compared to trending methods under limited supervision.
Yutao Rao, Liwei Sun, Junjie Zhang 0002, Jian Zhang 0002, Dan Zeng 0001
ICME6
2024 Masked Face Recognition with Generative-to-Discriminative Representations
abstract
Masked face recognition is important for social good but challenged by diverse occlusions that cause insufficient or inaccurate representations. In this work, we propose a unified deep network to learn generative-to-discriminative representations for facilitating masked face recognition. To this end, we split the network into three modules and learn them on synthetic masked faces in a greedy module-wise pretraining manner. First, we leverage a generative encoder pretrained for face inpainting and finetune it to represent masked faces into category-aware descriptors. Attribute to the generative encoder’s ability in recovering context information, the resulting descriptors can provide occlusion-robust representations for masked faces, mitigating the effect of diverse masks. Then, we incorporate a multi-layer convolutional network as a discriminative reformer and learn it to convert the category-aware descriptors into identity-aware vectors, where the learning is effectively supervised by distilling relation knowledge from off-the-shelf face recognition model. In this way, the discriminative reformer together with the generative encoder serves as the pretrained backbone, providing general and discriminative representations towards masked faces. Finally, we cascade one fully-connected layer following by one softmax layer into a feature classifier and finetune it to identify the reformed identity-aware vectors. Extensive experiments on synthetic and realistic datasets demonstrate the effectiveness of our approach in recognizing masked faces.
Shiming Ge, Weijia Guo, Chenyu Li 0001, Junzheng Zhang, Dan Zeng 0001
ICML6
2024 3DBench: A Scalable 3D Benchmark and Instruction-Tuning Dataset
Junjie Zhang 0002, Tianci Hu, Xiaoshui Huang, Yongshun Gong, Dan Zeng 0001
IJCAI5
2024 EFDCNet: Encoding fusion and decoding correction network for RGB-D indoor semantic segmentation
Jianlin Chen, Gongyang Li, Zhijiang Zhang, Dan Zeng 0001
Image Vis. Comput.4
2024 A streamlined framework for BEV-based 3D object detection with prior masking
Qinglin Tong, Junjie Zhang 0002, Chenggang Yan 0001, Dan Zeng 0001
Image Vis. Comput.4
2024 TMSDNet: Transformer with multi-scale dense network for single and multi-view 3D reconstruction
abstract
Abstract 3D reconstruction is a long‐standing problem. Recently, a number of studies have emerged that utilize transformers for 3D reconstruction, and these approaches have demonstrated strong performance. However, transformer‐based 3D reconstruction methods tend to establish the transformation relationship between the 2D image and the 3D voxel space directly using transformers or rely solely on the powerful feature extraction capabilities of transformers. They ignore the crucial role played by deep multi‐scale representation of the object in the voxel feature domain, which can provide extensive global shape and local detail information about the object in a multi‐scale manner. In this article, we propose a novel framework TMSDNet (transformer with multi‐scale dense network) for single‐view and multi‐view 3D reconstruction with transformer to solve this problem. Based on our well‐designed combined‐transformer Block, which is canonical encoder–decoder architecture, voxel features with spatial order can be extracted from the input image, which are used to further extract multi‐scale global features in parallel using a multi‐scale residual attention module. Furthermore, a residual dense attention block is introduced for deep local features extraction and adaptive fusion. Finally, the reconstructed objects are produced with the voxel reconstruction block. Experiment results on the benchmarks such as ShapeNet and Pix3D datasets demonstrate that TMSDNet outperforms the existing state‐of‐the‐art reconstruction methods substantially.
Xiaoqiang Zhu, Xinsheng Yao, Junjie Zhang 0002, Lihua You, Xiaosong Yang, Jian J. Zhang 0001, Dan Zeng 0001
Comput. Animat. Virtual Worlds9
2024 DCTracker: Rethinking MOT in soccer events under dual views via cascade association
Long Hu, Junjie Zhang 0002, Weiyi Lv, Yongshun Gong, Jingya Wang 0001, Jian Zhang 0002, Dan Zeng 0001
Knowl. Based Syst.7
2024 Leveraging Frequency-Guided Mixer and Target-Aware Attention for Ground-Based Cloud Detection
abstract
Compared to satellite imagery, ground-based cameras capture cloud data (ground-to-sky data) with higher temporal and spatial resolutions, providing more detailed cloud information. However, the spectral information available in ground-to-sky data is limited. Therefore, extracting features with strong discrimination from optical remote sensing images (ORSIs) is challenging. Currently, deep learning-based cloud detection methods face two main challenges. Firstly, although Convolutional Neural Networks (CNNs) effectively extract high-frequency (HF) components from images through convolutions, they struggle to capture low-frequency (LF) components, which are capable of representing global features and target structures. Secondly, in ORSIs, the spectral characteristics of thin clouds and the sky are similar, making it difficult to distinguish cloud regions from the background. To address these challenges, we propose a network consisting of two main modules: the Mixer Module (MM) and the Cloud Aware Attention Module (CAAM). The MM comprises a HF and a LF components extraction branch. The HF branch extracts local textures through max-pooling and parallel convolution operations. The LF branch captures long-range dependency by decomposing a large kernel convolution. It leverages the advantages of both convolution and self-attention to effectively capture global features. In addition, we introduce the CAAM, which quantifies images into histograms to separate clouds from the background and enhances the perception of clouds using attention mechanism. We conducted experiments using both daytime and nighttime cloud image data from the SWINySeg dataset with mIoU reaching 88.93% and OA reaching 93.97%. The results demonstrate that our proposed method achieves promising performance compared to state-of-the-art cloud detection methods.
Chenyu Dong, Guanyi Li, Yixiao Gu, Junjie Zhang 0002, Dan Zeng 0001
IEEE Geosci. Remote. Sens. Lett.5
2024 Multi-Level Information Fusion Network With Edge Information Injection for Single-Band Cloud Detection
abstract
Current cloud detection methods have demonstrated effectiveness by utilizing the rich spectral features of multi-spectral images. Compared to multispectral images, single-band infrared images offer higher efficiency in terms of sampling and processing speed. However, single-band cloud detection methods have not been fully developed, and existing methods based on multispectral cloud detection have some limitations when applied directly to single-band images: Firstly, they often blend shallow features containing spatial details with deep features providing high-level semantic information, yet struggle to disentangle features with strong discrimination representing cloud edges and bodies from limited information. Additionally, the correlation between features at different aspects is not fully reasoned, resulting in blurred boundary segmentation. To address these issues, we introduce a Multi-level Information Fusion Network (MIFNet) with an integrated edge information injection strategy. Our method effectively decouples clouds into their fundamental components: body and edge (Low-Frequency (LF) and High-Frequency (HF) components), enabling the comprehensive acquisition of strong discriminative features. Specifically, we propose an Edge Feature Extraction Module (EFEM) that isolates the cloud body through low-pass filtering, while the cloud’s edge is extracted by subtracting lower-level features from LF components. Furthermore, we employ a Feature Refinement Module (FRM) to locate the cloud body’s position precisely. Building upon this foundation, we devise a Graph Reasoning Module (GRM) to facilitate the full inference of feature correlations at different levels and to model the global interdependence between edges and semantics. Through comprehensive evaluations on benchmark datasets comprising infrared band images from Landsat 8 and MODIS satellites, we demonstrate that our proposed MIFNet outperforms state-of-the-art methods, yielding promising results in cloud detection accuracy. Our code is publicly available at https://github.com/KwunYat/MIFNet.
Guanyi Li, Junjie Zhang 0002, Enquan Yang, Dan Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 One-Shot Multiple Object Tracking With Robust ID Preservation
abstract
Maintaining identity consistency and avoiding ID-switch during tracking is one of the primary focuses of multiple object tracking (MOT). One-shot MOT methods which jointly learn the detection and tracking models in one single network (hence namely, one-shot) have achieved promising results in tracking accuracy and speed. However, their capabilities of maintaining ID consistency are somehow weakened. The reason for this weakened ID consistency is two-fold: (1) the ID features learned by one-shot methods are not discriminative enough due to their heatmap-based single-location representation. (2) severe occlusion in the MOT scene leads to feature ambiguity and high ID-switch. In this paper, we propose a one-shot MOT system with strong ID consistency called PID-MOT (Preserved ID MOT). Specifically, we devise a visibility branch to predict the object occlusion level, and a predicted visibility map will be used in both Feature Refinement Model (FRM) and a visibility-guided two-stage association strategy (VGTAS). FRM is designed to strengthen the location-based features and enrich the identity information. VGTAS is proposed for tackling objects with high and low visibility separately. In addition, we initialize the parameters of our model by training on the recently emerged abundant synthetic MOTSynth dataset from scratch rather than the commonly used COCO dataset for full training. Finally, we carry out our method on the commonly used MOT datasets and the experimental results demonstrate that the proposed PID-MOT achieves especially good performances in ID F1 score (IDF1) and ID-Switch (IDS) compared with other state-of-the-art one-shot trackers, with comparable overall HOTA/MOTA performance. The code is available at https://github.com/Kroery/PIDMOT.
Weiyi Lv, Ning Zhang 0023, Junjie Zhang 0002, Dan Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 Normal Image Guided Segmentation Framework for Unsupervised Anomaly Detection
abstract
Unsupervised anomaly detection is required to detect/segment anomalous samples/regions that deviate from the normal pattern while learning only through the normal sample category. Towards this end, this paper proposes a novel framework for anomaly detection by introducing normal images as guidance called Normal Image Guided Segmentation Framework (NIGSF). It consists of a Normal Guided Network (NGN) and a Saliency Augmentation Module (SAM). NGN constructs the contrast set, which is a candidate set for extracting normal sample features. Then, a normal feature extractor is developed to extract detailed and complete features containing normal semantic information as guidance features. Meanwhile, the guidance feature fusion module is introduced to realize normal semantic guidance in the feature space, and then the segmentation module discriminates the features that are different from the normal guidance features as anomalies. SAM aims to generate forged anomaly samples utilizing available normal samples. It introduces saliency maps and random Perlin noise to generate saliency Perlin noise maps and then to generate diverse forged anomaly samples. Extensive experiments are conducted to evaluate the performance of NIGSF on three anomaly detection benchmark datasets. The results demonstrate the effectiveness of each proposed module and the superiority of the proposed method. Specifically, NIGSF outperforms the runner-up by 5.4% in terms of anomaly segmentation AP metric.
Peng Xing, Yanpeng Sun, Dan Zeng 0001, Zechao Li
IEEE Trans. Circuits Syst. Video Technol.3
2024 Low-Resolution Object Recognition With Cross-Resolution Relational Contrastive Distillation
abstract
Recognizing objects in low-resolution images is a challenging task due to the lack of informative details. Recent studies have shown that knowledge distillation approaches can effectively transfer knowledge from a high-resolution teacher model to a low-resolution student model by aligning cross-resolution representations. However, these approaches still face limitations in adapting to the situation where the recognized objects exhibit significant representation discrepancies between training and testing images. In this study, we propose a cross-resolution relational contrastive distillation approach to facilitate low-resolution object recognition. Our approach enables the student model to mimic the behavior of a well-trained teacher model which delivers high accuracy in identifying high-resolution objects. To extract sufficient knowledge, the student learning is supervised with contrastive relational distillation loss, which preserves the similarities in various relational structures in contrastive representation space. In this manner, the capability of recovering missing details of familiar low-resolution objects can be effectively enhanced, leading to a better knowledge transfer. Extensive experiments on low-resolution object classification and low-resolution face recognition clearly demonstrate the effectiveness and adaptability of our approach.
Kangkai Zhang, Shiming Ge, Ruixin Shi, Dan Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 DFIE3D: 3D-Aware Disentangled Face Inversion and Editing via Facial-Contrastive Learning
abstract
Recent advances in NeRF-based 3D-aware GANs have achieved outstanding performance, especially in the realm of human facial representations, making projection of facial images back into their latent space superior and preferable compared to 2D GAN inversion. However, the direct application of 2DGAN inversion techniques to 3DGAN raises challenges due to potential appearance distortions and geometric inconsistences. To tackle these issues, this work presents a novel integrated framework that combines a composite inversion pipeline in both the SS and W+ spaces and integrates a contrastive-based training strategy, ensuring proficient disentanglement within the module. Moreover, we design a facial semantic manipulation technique based on dimensional analysis of the latent code, which is fully compatible with the proposed 3DGAN inversion pipeline. Comprehensive experimental validations substantiate the effectiveness of the proposed approach in executing 3d-aware face inversion and semantic editing tasks, presenting a robust technological solution for a diverse array of digital human modeling applications in the downstream.
Xiaoqiang Zhu, Lihua You, Xiaosong Yang, Jian Chang 0001, Jian J. Zhang 0001, Dan Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.7
2024 Single-Band Stripe Noise Removal in Multispectral Remote Sensing Images Based on Semisupervised Disentangled Transformation Network
abstract
The presence of stripe noise in multispectral data is a common issue caused by various factors during the imaging process. This noise severely degrades image quality and imposes limitations on downstream tasks. Although deep learning-based methods have demonstrated promising results in destriping, they often encounter challenges due to the disparity between the stripe noise distribution in simulated and real images. As a result, their destriping performance on real data is significantly hindered. To address this challenge, we propose a semisupervised disentangled transformation network (SDTNet) that encourages the model to learn the real stripe noise distribution through image decoupling and noise transformation using simulated and real data. SDTNet consists of the simulated and real image branches, which are subject to supervised and unsupervised constraints, respectively. They are jointly trained to enhance each other mutually. Furthermore, we introduce a decoupling strategy that effectively preserves the clean background component using self-consistency and adversarial losses. Instead of directly converting the whole image from the simulated domain to the real domain, SDTNet focuses on the relatively simpler task of converting the stripe noise component while maintaining the consistency of the image background. Extensive experimental evaluations on various datasets demonstrate the superior destriping performance of the proposed SDTNet compared to other methods, particularly in effectively removing stripe noise from real images.
Jia Li 0032, Xianying He, Fangfang Cui, Dan Zeng 0001
IEEE Trans. Geosci. Remote. Sens.7
2024 Weakly Supervised Semantic Segmentation With Consistency-Constrained Multiclass Attention for Remote Sensing Scenes
abstract
Obtaining image-level class labels for Remote Sensing (RS) images is a relatively straightforward process, sparking significant interest in Weakly Supervised Semantic Segmentation (WSSS). However, RS images present challenges beyond those encountered in generic WSSS, including complex backgrounds, densely distributed small objects, and considerable scale variations. To address above issues, we introduce a COnsistency-COnstrained Multi-Class Attention model, noted asCocoaNet. Specifically, CocoaNet endeavors to capture both semantic correlation and class distinctiveness using a Global-Local Adaptive Attention mechanism, which integrates the self-attention to model global correlation, complemented by a Local Perception branch that intensifies focus on local regions. The resulting class-specific attention weights and patch-level pairwise affinity weights are employed to optimize the initial CAMs. This mechanism proves highly effective in mitigating inter-class interference and managing the distribution of densely clustered small objects. Moreover, we invoke a Consistency Constraint to rectify activation inaccuracy. By utilizing a Siamese structure for the mutual supervision of features extracted from images at different scales, we address substantial scale variations in RS scenes. Simultaneously, a Class Contrast Loss is adopted to enhance the discriminativeness of class-specific features. Departing from the conventional CAM optimization, which is rather complex and time-consuming, we harness the prior knowledge from generic Segment Anything model to design a joint optimization strategy that refines target boundaries and further promotes discriminative visual features. We validate the effectiveness of our proposed approach on three benchmark datasets in multi-class RS scenarios, experimental results demonstrate that our model yield promising advancements compared to state-of-the-art methods.
Junjie Zhang 0002, Yongshun Gong, Jian Zhang 0002, Liang Chen 0004, Dan Zeng 0001
IEEE Trans. Geosci. Remote. Sens.6
2024 Learning Contrast-Enhanced Shape-Biased Representations for Infrared Small Target Detection
abstract
Detecting infrared small targets under cluttered background is mainly challenged by dim textures, low contrast and varying shapes. This paper proposes an approach to facilitate infrared small target detection by learning contrast-enhanced shape-biased representations. The approach cascades a contrast-shape encoder and a shape-reconstructable decoder to learn discriminative representations that can effectively identify target objects. The contrast-shape encoder applies a stem of central difference convolutions and a few large-kernel convolutions to extract shape-preserving features from input infrared images. This specific design in convolutions can effectively overcome the challenges of low contrast and varying shapes in a unified way. Meanwhile, the shape-reconstructable decoder accepts the edge map of input infrared image and is learned by simultaneously optimizing two shape-related consistencies: the internal one decodes the encoder representations by upsampling reconstruction and constraints segmentation consistency, whilst the external one cascades three gated ResNet blocks to hierarchically fuse edge maps and decoder representations and constrains contour consistency. This decoding way can bypass the challenge of dim texture and varying shapes. In our approach, the encoder and decoder are learned in an end-to-end manner, and the resulting shape-biased encoder representations are suitable for identifying infrared small targets. Extensive experimental evaluations are conducted on public benchmarks and the results demonstrate the effectiveness of our approach.
Fanzhao Lin, Kexin Bao, Dan Zeng 0001, Shiming Ge
IEEE Trans. Image Process.4
2024 VWP:An Efficient DRL-Based Autonomous Driving Model
abstract
In this paper, a novel DRL-based model (VWP, VAE-WGAN-PPOE) is proposed to solve the problem of long training time and unsatisfactory training effect in the end-to-end autonomous driving. The model is optimized from feature extraction and algorithm decision. In feature extraction, we encode the input video by combining variational auto encoder (VAE) with wasserstein generative adversarial network (WGAN). The state dimension is reduced and the problem of mode collapse and gradient disappearance caused by generative adversarial network (GAN) training is solved. In decision algorithm, we formulate a new reward function by analyzing the factors affecting driving performance. Furthermore, we propose an enhanced algorithm PPOE based on the proximal policy optimization (PPO). In the CARLA simulator, compared with CNN and ResNet34, the convergence speed of the DRL model based on VAE-WGAN increases by 26.1% and 20.3%, the navigation task completion rate increases by 18.5% and 9.2%, and the collision rate decreases by 13.6% and 9.4%. Compared with deep deterministic policy gradient (DDPG) decision algorithm, the convergence speed of the DRL model based on PPOE increases by 23.3%, the navigation task completion rate increases by 5.0% in sunny days and 8.4% in severe weather, the collision rate decreases by 3.5% in sunny days and 6.6% in severe weather. Extensive experiments show that the proposed model enables the agent to drive safely along the navigational route in the complex environment with pedestrian and vehicle interaction, even in severe weather.
Yanliang Jin, Ze-Yu Ji, Dan Zeng 0001, Xiao-Ping Zhang 0002
IEEE Trans. Multim.3
2024 Learning Shape-Biased Representations for Infrared Small Target Detection
abstract
Typically, infrared small target detection aims to accurately localize objects from complex backgrounds where the object textures are often dim and the object shapes are varying. A feasible solution is learning discriminative representations with deep convolutional neural networks (CNNs). However, the representations learned by traditional deep CNNs often suffer from low shape bias. In this work, we propose a unified framework to learn shape-biased representations for facilitating infrared small target detection by explicitly incorporating shape information into model learning. The framework cascades a large-kernel encoder and a shape-guided decoder to learn discriminative shape-biased representations in an end-to-end manner. The large-kernel encoder describes infrared images into shape-preserving representations by using a few convolutions whose kernel size is as large as$9\times 9$, in contrast to commonly used$3\times 3$. The shape-guided decoder simultaneously addresses two tasks: decodes the encoder representations via upsampling reconstruction to reconstruct the segmentation, and hierarchically fuses the decoder representations and edge information via cascaded gated ResNet blocks to reconstruct the contour. In this way, the learned shape-biased representations are effective for identifying infrared small targets. Extensive experiments show our approach outperforms 18 state-of-the-arts.
Fanzhao Lin, Shiming Ge, Kexin Bao, Chenggang Yan 0001, Dan Zeng 0001
IEEE Trans. Multim.5
2024 Cross-Modal Quantization for Co-Speech Gesture Generation
abstract
Learning proper representations for speech and gesture is essential for co-speech gesture generation. Existing approaches either utilize direct representations or independently encode the speech and gesture, which neglect the joint representation to highlight the interplay between these two modalities. In this work, we propose a novel Cross-modal Quantization (CMQ) to jointly learn the quantized codes for speech and gesture together. Such representation highlights the speech-gesture interaction before actually learning the complex mapping, and thus better suits the intricate mapping between speech and gesture. Specifically, the Cross-modal Quantizer jointly encodes speech and gesture as discrete codebooks, enabling better cross-modal interaction. Cross-modal Predictor subsequently utilizes the learned codebooks to autoregressively predict the next-step gesture. With cross-modal quantization, our approach yields much higher codebook usage and generates more realistic and diverse gestures in practice. Extensive experiments are conducted on both 3D and 2D datasets as well as the subjective user study, demonstrating a clear performance gain compared to several baseline models in terms of audio-visual alignment and gesture diversity. In particular, our method demonstrates a three-fold improvement in diversity compared to baseline models, while simultaneously maintaining high motion fidelity.
Zheng Wang 0059, Wei Zhang 0031, Long Ye, Dan Zeng 0001, Tao Mei 0001
IEEE Trans. Multim.4
2024 Frequency-Aware Multi-Modal Fine-Tuning for Few-Shot Open-Set Remote Sensing Scene Classification
abstract
Few-shot open-set recognition, as a new paradigm, leveraging a limited amount of supervised data to identify specific Remote Sensing (RS) scene categories and generalize to novel ones. However, the data bias induced by the small sample size not only causes severe overfitting within base classes, but also impairs the capacity for inference to identify RS scenes in hitherto unobserved categories. Furthermore, owing to environmental influences, RS images frequently manifest notable intra-class disparities and comparatively low inter-class distinctions, intensifying the challenge in obtaining suitable classifiers. To address above issues, we investigate the utilization of a Multi-modal Foundational Model (MFM) infused with essential domain knowledge to mitigate the generalization limitations encountered in few-shot scenarios. Recognizing that existing MFMs with a visual-text dual-branch structure are primarily tailored for natural scenes, we propose a custom Frequency Distribution-based Multi-modal Fine-Tuning strategy (FreqDiMFT) in a parameter-efficient manner. More specifically, within the vision branch, we address the high inter-class similarity and intra-class diversity in RS images by embedding the local-global frequency distribution information to facilitate the recognition of RS scenes. To further amplify the model's generalization ability post transfer, we introduce an adaptive feature refinement module designed for Transformers, proficient in filtering redundant features resulting from domain disparities. To mitigate the domain drift on the textual branch, we adopt an input format that combines basic templates with domain expertise from RS end to generate more discriminative class prototypes. To fully verify the effectiveness of our FreqDiMFT in a more practical setting, we collect a Large-Scale hybrid dataset (LSRS). Extensive experiments demonstrate that, even with a scant number of training samples, our strategy yields advanced performances compared to state-of-the-art models.
Junjie Zhang 0002, Yutao Rao, Xiaoshui Huang, Guanyi Li, Dan Zeng 0001
IEEE Trans. Multim.6
2024 STAT: Multi-Object Tracking Based on Spatio-Temporal Topological Constraints
abstract
The mainstream tracking-by-detection paradigm for multi-object tracking generally conducts detection first, followed by Re-IDentification (Re-ID) and motion estimation. The associations between the predicted boxes and existing tracks are then performed via visual and motion association. However, challenges such as irregular motion patterns, similar appearances, and frequent occlusions often arise, making object tracking a nontrivial task. In this article, we propose a multi-object tracker based on Spatio-TemporAl Topological (STAT) constraints to address the above issues. More specifically, we design the Feature Adaptive Association Module (FAAM) to establish the association between motion and appearance regionally, completing a complementary combination of appearance and motion features. Among these, the Appearance Feature Update Module (AFUM) is proposed to manage the appearance updates of tracked objects by imposing constraints based on the spatial locations and the degree of object occlusion, while temporal consistency is adopted to smooth the appearance states of tracks to mitigate the accumulation of appearance noise. Moreover, the Robust Motion Tracking Module (RMTM) is established to reduce the impact of irregular motions and certain unreliable detection results. The proposed module includes a higher weighted momentum term to accommodate the excessive motion amplitude and considers low-confidence boxes accompanied by the stage-wise association strategy for high-confidence boxes. Extensive experiments on DanceTrack and benchmark MOT datasets verify the effectiveness of our STAT tracker, especially the state-of-the-art results on DanceTrack, which is characterized by irregular motion and indistinguishable appearance attributes.
Junjie Zhang 0002, Xinyu Zhang 0015, Chenggang Yan 0001, Dan Zeng 0001
IEEE Trans. Multim.6
2024 GLCSA-Net: global-local constraints-based spectral adaptive network for hyperspectral image inpainting
Jia Li 0032, Junjie Zhang 0002, Chenggang Yan 0001, Dan Zeng 0001
Vis. Comput.6
2024 Region-guided network with visual cues correction for infrared small target detection
Junjie Zhang 0002, Dan Zeng 0001
Vis. Comput.4
2023 Bootstrapping Multi-View Representations for Fake News Detection
abstract
Previous researches on multimedia fake news detection include a series of complex feature extraction and fusion networks to gather useful information from the news. However, how cross-modal consistency relates to the fidelity of news and how features from different modalities affect the decision-making are still open questions. This paper presents a novel scheme of Bootstrapping Multi-view Representations (BMR) for fake news detection. Given a multi-modal news, we extract representations respectively from the views of the text, the image pattern and the image semantics. Improved Multi-gate Mixture-of-Expert networks (iMMoE) are proposed for feature refinement and fusion. Representations from each view are separately used to coarsely predict the fidelity of the whole news, and the multimodal representations are able to predict the cross-modal consistency. With the prediction scores, we reweigh each view of the representations and bootstrap them for fake news detection. Extensive experiments conducted on typical fake news detection datasets prove that BMR outperforms state-of-the-art schemes.
Qichao Ying, Xiaoxiao Hu, Yangming Zhou, Zhenxing Qian, Dan Zeng 0001, Shiming Ge
AAAI5
2023 Model Conversion via Differentially Private Data-Free Distillation
abstract
While massive valuable deep models trained on large-scale data have been released to facilitate the artificial intelligence community, they may encounter attacks in deployment which leads to privacy leakage of training data. In this work, we propose a learning approach termed differentially private data-free distillation (DPDFD) for model conversion that can convert a pretrained model (teacher) into its privacy-preserving counterpart (student) via an intermediate generator without access to training data. The learning collaborates three parties in a unified way. First, massive synthetic data are generated with the generator. Then, they are fed into the teacher and student to compute differentially private gradients by normalizing the gradients and adding noise before performing descent. Finally, the student is updated with these differentially private gradients and the generator is updated by taking the student as a fixed discriminator in an alternate manner. In addition to a privacy-preserving student, the generator can generate synthetic data in a differentially private way for other down-stream tasks. We theoretically prove that our approach can guarantee differential privacy and well convergence. Extensive experiments that significantly outperform other differentially private generative approaches demonstrate the effectiveness of our approach.
Bochao Liu, Shikun Li, Dan Zeng 0001, Shiming Ge
IJCAI4
2023 Personalized Federated Learning via Backbone Self-Distillation
abstract
In practical scenarios, federated learning frequently necessitates training personalized models for each client using heterogeneous data. This paper proposes a backbone self-distillation approach to facilitate personalized federated learning. In this approach, each client trains its local model and only sends the backbone weights to the server. These weights are then aggregated to create a global backbone, which is returned to each client for updating. However, the client’s local backbone lacks personalization because of the common representation. To solve this problem, each client further performs backbone self-distillation by using the global backbone as a teacher and transferring knowledge to update the local backbone. This process involves learning two components: the shared backbone for common representation and the private head for local personalization, which enables effective global knowledge transfer. Extensive experiments and comparisons with 12 state-of-the-art approaches demonstrate the effectiveness of our approach.
Bochao Liu, Dan Zeng 0001, Chenggang Yan 0001, Shiming Ge
MMAsia3
2023 A Pyramid Attention Network With Edge Information Injection for Remote-Sensing Object Detection
abstract
Remote sensing images (RSIs) are often characterized by the high spatial resolution, strong object scale effects, and complex scenes, which poses great challenges to the object detection. Although mainstream neural network-based methods work well in detecting common objects, they often fail to fully exploit the detailed structural information in the spatial domain, leading to the poor performance for objects with diverse scales and distributions under complicated backgrounds. To address the above issue, we propose a pyramid attention network with edge information injection for remote sensing object detection. Considering each object is composed of the inner body and outer profile parts that corresponding to the low and high frequency components of image respectively, the difference between the original image and its low frequency component is beneficial for obtaining the high frequency counterpart. We design the Edge Information Extraction Module (EIEM) to mine the detailed edge features at multiple scales, and subsequently inject them into features at corresponding scales in the backbone network. As for promoting the performance in complex scenes, we introduce a Pyramid Feature Fusion (PFF) module, which leverages both local and global attention for establishing the long-range channel dependency, thereby highlighting objects that need to be concentrated on. To verify the effectiveness of our proposed method, we conduct extensive experiments on DIOR and RSOD datasets with mean Average Precision (mAP) reaching 74.93% and 96.44% respectively, demonstrating that our model achieved SOTA performance compared to mainstream methods.
Junjie Zhang 0002, Anqi Ding, Guanyi Li, Liangang Zhang, Dan Zeng 0001
IEEE Geosci. Remote. Sens. Lett.5
2023 CaCo: Both Positive and Negative Samples are Directly Learnable via Cooperative-Adversarial Contrastive Learning
abstract
As a representative self-supervised method, contrastive learning has achieved great successes in unsupervised training of representations. It trains an encoder by distinguishing positive samples from negative ones given query anchors. These positive and negative samples play critical roles in defining the objective to learn the discriminative encoder, avoiding it from learning trivial features. While existing methods heuristically choose these samples, we present a principled method where both positive and negative samples are directly learnable end-to-end with the encoder. We show that the positive and negative samples can be cooperatively and adversarially learned by minimizing and maximizing the contrastive loss, respectively. This yields cooperative positives and adversarial negatives with respect to the encoder, which are updated to continuously track the learned representation of the query anchors over mini-batches. The proposed method achieves 71.3% and 75.3% in top-1 accuracy respectively over 200 and 800 epochs of pre-training ResNet-50 backbone on ImageNet1K without tricks such as multi-crop or stronger augmentations. With Multi-Crop, it can be further boosted into 75.7%. The source code and pre-trained model are released in https://github.com/maple-research-lab/caco.
Xiao Wang 0013, Yuhang Huang 0006, Dan Zeng 0001, Guo-Jun Qi
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 RGB-T Semantic Segmentation With Location, Activation, and Sharpening
abstract
Semantic segmentation is important for scene understanding. To address the scenes of adverse illumination conditions of natural images, thermal infrared (TIR) images are introduced. Most existing RGB-T semantic segmentation methods follow three cross-modal fusion paradigms, i. e., encoder fusion, decoder fusion, and feature fusion. Some methods, unfortunately, ignore the properties of RGB and TIR features or the properties of features at different levels. In this paper, we propose a novel feature fusion-based network for RGB-T semantic segmentation, named LASNet, which follows three steps of location, activation, and sharpening. The highlight of LASNet is that we fully consider the characteristics of cross-modal features at different levels, and accordingly propose three specific modules for better segmentation. Concretely, we propose a Collaborative Location Module (CLM) for high-level semantic features, aiming to locate all potential objects. We propose a Complementary Activation Module for middle-level features, aiming to activate exact regions of different objects. We propose an Edge Sharpening Module (ESM) for low-level texture features, aiming to sharpen the edges of objects. Furthermore, in the training phase, we attach a location supervision and an edge supervision after CLM and ESM, respectively, and impose two semantic supervisions in the decoder part to facilitate network convergence. Experimental results on two public datasets demonstrate that the superiority of our LASNet over relevant state-of-the-art methods. The code and results of our method are available athttps://github.com/MathLee/LASNet.
Gongyang Li, Yike Wang 0003, Zhi Liu 0003, Xinpeng Zhang 0001, Dan Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.5
2023 Adjacent Context Coordination Network for Salient Object Detection in Optical Remote Sensing Images
abstract
Salient object detection (SOD) in optical remote sensing images (RSIs), or RSI-SOD, is an emerging topic in understanding optical RSIs. However, due to the difference between optical RSIs and natural scene images (NSIs), directly applying NSI-SOD methods to optical RSIs fails to achieve satisfactory results. In this article, we propose a novel adjacent context coordination network (ACCoNet) to explore the coordination of adjacent features in an encoder-decoder architecture for RSI-SOD. Specifically, ACCoNet consists of three parts: 1) an encoder; 2) adjacent context coordination modules (ACCoMs); and 3) a decoder. As the key component of ACCoNet, ACCoM activates the salient regions of output features of the encoder and transmits them to the decoder. ACCoM contains a local branch and two adjacent branches to coordinate the multilevel features simultaneously. The local branch highlights the salient regions in an adaptive way, while the adjacent branches introduce global information of adjacent levels to enhance salient regions. In addition, to extend the capabilities of the classic decoder block (i.e., several cascaded convolutional layers), we extend it with two bifurcations and propose a bifurcation-aggregation block (BAB) to capture the contextual information in the decoder. Extensive experiments on two benchmark datasets demonstrate that the proposed ACCoNet outperforms 22 state-of-the-art methods under nine evaluation metrics, and runs up to 81 fps on a single NVIDIA Titan X GPU. The code and results of our method are available at https://github.com/MathLee/ACCoNet.
Gongyang Li, Zhi Liu 0003, Dan Zeng 0001, Weisi Lin, Haibin Ling
IEEE Trans. Cybern.3
2023 Progressive Recurrent Neural Network for Multispectral Remote Sensing Image Destriping
abstract
An unstable imaging system often introduces additional stripe noise in multispectral remote sensing images during the data acquisition process given a variety of factors. The complicated stripe distributions lead to the residual stripe in the results of existing methods, thus increasing the difficulty of destriping in practice. Mainstream deep learning-based methods show the encouraging destriping performance on multispectral remote sensing images. However, they often require the model to handle the varying degrees of stripe noise in a single shot for each image, which results in the poor destriping performance when facing practical cases with diverse stripe distributions. To address the above issue, we propose a Progressive Recurrent Neural Network (PRNet) to remove the stripe noise for each degraded image in an iterative manner. More specifically, a progressive destriping strategy is designed to gradually restore the clean image, in which the Main Recurrent Module (MRM) is introduced to iteratively process the stripe removal results generated from previous timesteps until the clean image is obtained. Furthermore, since the uniformity of the entire image is supposed to be significantly enhanced after the destriping, it is necessary to take the local spatial correlation into account during the destriping. Therefore, we present the Patch-based Sequence Module (PSM) to leverage the local spatial correlation by splitting the image into multi-scale patch sequences and capturing the relationship among different patches. Extensive experimental results on different datasets demonstrate that the proposed model yields superior destriping performance compared to other methods, especially for removing the stripe noise with complex distributions.
Jia Li 0032, Junjie Zhang 0002, Jungong Han, Chenggang Yan 0001, Dan Zeng 0001
IEEE Trans. Geosci. Remote. Sens.5
2023 Learning Privacy-Preserving Student Networks via Discriminative-Generative Distillation
abstract
While deep models have proved successful in learning rich knowledge from massive well-annotated data, they may pose a privacy leakage risk in practical deployment. It is necessary to find an effective trade-off between high utility and strong privacy. In this work, we propose a discriminative-generative distillation approach to learn privacy-preserving deep models. Our key idea is taking models as bridge to distill knowledge from private data and then transfer it to learn a student network via two streams. First, discriminative stream trains a baseline classifier on private data and an ensemble of teachers on multiple disjoint private subsets, respectively. Then, generative stream takes the classifier as a fixed discriminator and trains a generator in a data-free manner. After that, the generator is used to generate massive synthetic data which are further applied to train a variational autoencoder (VAE). Among these synthetic data, a few of them are fed into the teacher ensemble to query labels via differentially private aggregation, while most of them are embedded to the trained VAE for reconstructing synthetic data. Finally, a semi-supervised student learning is performed to simultaneously handle two tasks: knowledge transfer from the teachers with distillation on few privately labeled synthetic data, and knowledge enhancement with tangent-normal adversarial regularization on many triples of reconstructed synthetic data. In this way, our approach can control query cost over private data and mitigate accuracy degradation in a unified manner, leading to a privacy-preserving student model. Extensive experiments and analysis clearly show the effectiveness of the proposed approach.
Shiming Ge, Bochao Liu, Dan Zeng 0001
IEEE Trans. Image Process.5
2023 Trustable Co-Label Learning From Multiple Noisy Annotators
abstract
Supervised deep learning depends on massive accurately annotated examples, which is usually impractical in many real-world scenarios. A typical alternative is learning from multiple noisy annotators. Numerous earlier works assume that all labels are noisy, while it is usually the case that a few trusted samples with clean labels are available. This raises the following important question: how can we effectively use a small amount of trusted data to facilitate robust classifier learning from multiple annotators? This paper proposes a data-efficient approach, calledTrustable Co-label Learning(TCL), to learn deep classifiers from multiple noisy annotators when a small set of trusted data is available. This approach follows the coupled-view learning manner, which jointly learns the data classifier and the label aggregator. It effectively uses trusted data as a guide to generate trustable soft labels (termed co-labels). A co-label learning can then be performed by alternately reannotating the pseudo labels and refining the classifiers. In addition, we further improve TCL for a special complete data case, where each instance is labeled by all annotators and the label aggregator is represented by multilayer neural networks to enhance model capacity. Extensive experiments on synthetic and real datasets clearly demonstrate the effectiveness and robustness of the proposed approach. Source code is available athttps://github.com/ShikunLi/TCL.
Shikun Li, Tongliang Liu, Jiyong Tan, Dan Zeng 0001, Shiming Ge
IEEE Trans. Multim.4
2023 RINet: Relative Importance-Aware Network for Fixation Prediction
abstract
Fixation prediction aims to simulate human visual selection mechanism and estimate the visual saliency degree of regions in a scene. In semantically rich scenes, there are generally multiple salient regions. This condition requires a fixation prediction model to understand the relative importance relationship of multiple salient regions, that is, to identify which region is more important. In practice, existing fixation prediction models implicitly explore the relative importance relationship in the end-to-end training process while they do not work well. In this article, we propose a novel Relative Importance-aware Network (RINet) to explicitly explore the modeling of relative importance in fixation prediction. RINet perceives multi-scale local and global relative importance through the Hierarchical Relative Importance Enhancement (HRIE) module. Within a single scale subspace, on the one hand, HRIE module regards the similarity matrix as the local relative importance map to weight the input feature. On the other hand, HRIE module integrates a set of local relative importance maps into one map, defined as the global relative importance map, to grasp global relative importance. Moreover, we propose a Complexity-Relevant Focal (CRF) loss for network training. As such, we can progressively emphasize learning difficult samples for better handling the complicated scenarios, further improving the performance. The ablation studies confirm the contributions of key components of our RINet, and extensive experiments on five datasets demonstrate our RINet is superior to 28 relevant state-of-the-art models.
Zhi Liu 0003, Gongyang Li, Dan Zeng 0001, Tianhong Zhang, Lihua Xu, Jijun Wang 0003
IEEE Trans. Multim.4
2023 Boosting Relationship Detection in Images with Multi-Granular Self-Supervised Learning
abstract
Visual and spatial relationship detection in images has been a fast-developing research topic in the multimedia field, which learns to recognize the semantic/spatial interactions between objects in an image, aiming to compose a structured semantic understanding of the scene. Most of the existing techniques directly encapsulate the holistic image feature plus the semantic and spatial features of the given two objects for predicting the relationship, but leave the inherent supervision derived from such structured and thorough image understanding under-exploited. Specifically, the inherent supervision among objects or relations within an image can span different granularities in this hierarchy including, from simple to comprehensive, (1) the object-based supervision that captures the interaction between the semantic and spatial features of each individual object, (2) the inter-object supervision that characterizes the dependency within the relationship triplet ( ), and (3) the inter-relation supervision that exploits contextual information among all relationship triplets in an image. These inherent multi-granular supervisions offer a fertile ground for building self-supervised proxy tasks. In this article, we compose a trilogy of exploring the multi-granular supervision in the sequence from object-based, inter-object, and inter-relation perspectives. We integrate the standard relationship detection objective with a series of proposed self-supervised proxy tasks, which is named as Multi-Granular Self-Supervised learning (MGS). Our MGS is appealing in view that it is pluggable to any neural relationship detection models by simply including the proxy tasks during training, without increasing the computational cost at inference. Through extensive experiments conducted on the SpatialSense and VRD datasets, we demonstrate the superiority of MGS for both spatial and visual relationship detection tasks.
Xuewei Ding, Yingwei Pan, Yehao Li, Ting Yao 0003, Dan Zeng 0001, Tao Mei 0001
ACM Trans. Multim. Comput. Commun. Appl.5
2023 Introduction to the Special Issue on Trustworthy Multimedia Computing and Applications in Urban Scenes
abstract
Special Issue Part 1 (Issue 3) and Part 2 (Issue 4) of AIEDAM are based on a workshop on Learning and Creativity held at the 2002 conference on Artificial Intelligence in Design, AID '02 (www.cad.strath.ac.uk/AID02_workshop/Workshop_webpage.html; Gero, ...
Wu Liu 0005, Hailin Shi, Yunchao Wei, Dan Zeng 0001, Nicu Sebe, Jiebo Luo 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2023 Multi-Agent Semi-Siamese Training for Long-Tail and Shallow Face Learning
abstract
With the recent development of deep convolutional neural networks and large-scale datasets, deep face recognition has made remarkable progress and been widely used in various applications. However, unlike the existing public face datasets, in many real-world scenarios of face recognition, the depth of the training dataset is shallow, which means that only two face images are available for each ID. With the non-uniform increase of samples, such issue is converted to a more general case, known as long-tail face learning, which suffers from data imbalance and intra-class diversity dearth simultaneously. These adverse conditions damage the training and result in the decline of model performance. Based on Semi-Siamese Training, we introduce an advanced solution, namedMulti-Agent Semi-Siamese Training(MASST), to address these problems. MASST includes a probe network and multiple gallery agents—the former aims to encode the probe features, and the latter constitutes a stack of networks that encode the prototypes (gallery features). For each training iteration, the gallery network, which is sequentially rotated from the stack, and the probe network form a pair of Semi-Siamese networks. We give the theoretical and empirical analysis that, given the long-tail (or shallow) data and training loss, MASST smooths the loss landscape and satisfies the Lipschitz continuity with the help of multiple agents and the updating gallery queue. The proposed method is out of extra-dependency, and thus can be easily integrated with the existing loss functions and network architectures. It is worth noting that although multiple gallery agents are employed for training, only the probe network is needed for inference, without increasing the inference cost. Extensive experiments and comparisons demonstrate the advantages of MASST for long-tail and shallow face learning.
Yichun Tai, Hailin Shi, Dan Zeng 0001, Yibo Hu 0003, Zhijiang Zhang, Tao Mei 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2022 Genre-Conditioned Long-Term 3D Dance Generation Driven by Music
abstract
Dancing to music is an artistic behavior of humans, however, letting machines generate dances from music is still challenging. Most existing works have been made progress in tackling the problem of motion prediction conditioned by music, yet they rarely consider the importance of the musical genre. In this paper, we focus on generating long-term 3D dance from music with a specific genre. Specifically, we construct a pure transformer-based architecture to correlate motion features and music features. To utilize the genre information, we propose to embed the genre categories into the transformer decoder so that it can guide every frame. Moreover, different from previous inference schemes, we introduce the motion queries to output the dance sequence in parallel that significantly improves the efficiency. Extensive experiments on AIST++[1] dataset show that our model outperforms state-of-the-art methods with a much faster inference speed.
Yuhang Huang 0006, Junjie Zhang 0002, Qian Bao, Dan Zeng 0001, Zhineng Chen, Wu Liu 0005
ICASSP5
2022 Deepfake Video Detection with Spatiotemporal Dropout Transformer
abstract
While the abuse of deepfake technology has caused serious concerns recently, how to detect deepfake videos is still a challenge due to the high photo-realistic synthesis of each frame. Existing image-level approaches often focus on single frame and ignore the spatiotemporal cues hidden in deepfake videos, resulting in poor generalization and robustness. The key of a video-level detector is to fully exploit the spatiotemporal inconsistency distributed in local facial regions across different frames in deepfake videos. Inspired by that, this paper proposes a simple yet effective patch-level approach to facilitate deepfake video detection via spatiotemporal dropout transformer. The approach reorganizes each input video into bag of patches that is then fed into a vision transformer to achieve robust representation. Specifically, a spatiotemporal dropout operation is proposed to fully explore patch-level spatiotemporal cues and serve as effective data augmentation to further enhance model's robustness and generalization ability. The operation is flexible and can be easily plugged into existing vision transformers. Extensive experiments demonstrate the effectiveness of our approach against 25 state-of-the-arts with impressive robustness, generalizability, and representation ability.
Daichi Zhang, Fanzhao Lin, Yingying Hua, Dan Zeng 0001, Shiming Ge
ACM Multimedia5
2022 Privacy-Preserving Student Learning with Differentially Private Data-Free Distillation
abstract
Deep learning models can achieve high inference accuracy by extracting rich knowledge from massive well-annotated data, but may pose the risk of data privacy leakage in practical deployment. In this paper, we present an effective teacher-student learning approach to train privacy-preserving deep learning models via differentially private data-free distillation. The main idea is generating synthetic data to learn a student that can mimic the ability of a teacher well-trained on private data. In the approach, a generator is first pretrained in a data-free manner by incorporating the teacher as a fixed discriminator. With the generator, massive synthetic data can be generated for model training without exposing data privacy. Then, the synthetic data is fed into the teacher to generate private labels. Towards this end, we propose a label differential privacy algorithm termed selective randomized response to protect the label information. Finally, a student is trained on the synthetic data with the supervision of private labels. In this way, both data privacy and label privacy are well protected in a unified framework, leading to privacy-preserving models. Extensive experiments and analysis clearly demonstrate the effectiveness of our approach.
Bochao Liu, Jianghu Lu, Junjie Zhang 0002, Dan Zeng 0001, Zhenxing Qian, Shiming Ge
MMSP5
2022 Singing Voice Synthesis with Vibrato Modeling and Latent Energy Representation
abstract
This paper proposes an expressive singing voice synthesis system by introducing explicit vibrato modeling and latent energy representation. Vibrato is essential to the naturalness of synthesized sound, due to the inherent characteristics of human singing. Hence, a deep learning-based vibrato model is introduced in this paper to control the vibrato's likeliness, rate, depth and phase in singing, where the vibrato likeliness represents the existence probability of vibrato and it would help improve the singing voice's naturalness. Actually, there is no annotated label about vibrato likeliness in existing singing corpus. We adopt a novel vibrato likeliness labeling method to label the vibrato likeliness automatically. Meanwhile, the power spectrogram of audio contains rich information that can improve the expressiveness of singing. An autoencoder-based latent energy bottleneck feature is proposed for expressive singing voice synthesis. Experimental results on the open dataset NUS48E show that both the vibrato modeling and the latent energy representation could significantly improve the expressiveness of singing voice. The audio samples are shown in the demo website11https://mango321321.github.io/ExpressiveSing/.
Wei Zhang 0031, Zhengchen Zhang, Dan Zeng 0001, Zhi Liu 0003
MMSP5
2022 Analysis of Formations and Game Styles in Soccer
abstract
An appropriate tactic will affect the situation of a soccer game, so corresponding tactical analysis is very important. The analysis of formations and the analysis of game styles are two significant aspects. In this paper, we find the most suitable method of team formation extraction and design a game style analysis network. Specifically, based on the method of minimum entropy data partitioning, we first conduct a series of experiments to identify the formations utilized by different teams and find an effective way to identify them through the comparison of experimental results. Then, we utilize the extracted formation information and a role-based passing network to design a novel team style analysis network. The experimental results show that our method has a good effect on game style analysis.
Qinglin Tong, Wendi Yao, Weiyi Lv, Dan Zeng 0001
MMSP4
2022 Controllable blending of line and polygon skeleton-based convolution surfaces with finite support kernels
Xiaoqiang Zhu, Sihu Liu, Chenjie Fan, Chenze Song, Junjie Zhang 0002, Dan Zeng 0001, Xiaogang Jin 0001
Comput. Graph.7
2022 Expression-tailored talking face generation with adaptive cross-modal weighting
Dan Zeng 0001, Shuaitao Zhao, Junjie Zhang 0002
Neurocomputing1
2022 Hyperspectral Anomaly Detection via Low-Rank Decomposition and Morphological Filtering
abstract
To effectively detect anomalies and eliminate the influence of noise on anomaly detection (AD), we propose a hyperspectral AD method based on low-rank decomposition and morphological filtering (LRDMF). For one thing, given the different ways in which anomalies and noise occur in the spectral bands, a low-rank decomposition model is proposed to decompose the original hyperspectral image (HSI) into the background, anomaly, and noise components, where a superpixel segmentation method and the sparse representation (SR) model are used to construct a robust background dictionary. For another thing, considering that the anomalies in HSI possess small area characteristics, a morphological filtering method is applied to preserve the small connected components. Finally, anomalies are detected by jointly considering the LRDMF results. The experimental results conducted on two real hyperspectral datasets demonstrate that the proposed method outperforms some of the state-of-the-art methods.
Yating Xu, Junjie Zhang 0002, Dan Zeng 0001
IEEE Geosci. Remote. Sens. Lett.4
2022 Dual Spoof Disentanglement Generation for Face Anti-Spoofing With Depth Uncertainty Learning
abstract
Face anti-spoofing (FAS) plays a vital role in preventing face recognition systems from presentation attacks. Existing face anti-spoofing datasets lack diversity due to the insufficient identity and insignificant variance, which limits the generalization ability of FAS model. In this paper, we propose Dual Spoof Disentanglement Generation (DSDG) framework to tackle this challenge by “anti-spoofing via generation”. Depending on the interpretable factorized latent disentanglement in Variational Autoencoder (VAE), DSDG learns a joint distribution of the identity representation and the spoofing pattern representation in the latent space. Then, large-scale paired live and spoofing images can be generated from random noise to boost the diversity of the training set. However, some generated face images are partially distorted due to the inherent defect of VAE. Such noisy samples are hard to predict precise depth values, thus may obstruct the widely-used depth supervised optimization. To tackle this issue, we further introduce a lightweight Depth Uncertainty Module (DUM), which alleviates the adverse effects of noisy samples by depth uncertainty learning. DUM is developed without extra-dependency, thus can be flexibly integrated with any depth supervised network for face anti-spoofing. We evaluate the effectiveness of the proposed method on five popular benchmarks and achieve state-of-the-art results under both intra- and inter- test settings. The codes are available athttps://github.com/JDAI-CV/FaceX-Zoo/tree/main/addition_module/DSDG.
Hangtong Wu, Dan Zeng 0001, Yibo Hu 0003, Hailin Shi, Tao Mei 0001
IEEE Trans. Circuits Syst. Video Technol.2
2022 Student Network Learning via Evolutionary Knowledge Distillation
abstract
Knowledge distillation provides an effective way to transfer knowledge via teacher-student learning, where most existing distillation approaches apply a fixed pre-trained model as teacher to supervise the learning of student network. This manner usually brings in a big capability gap between teacher and student networks during learning. Recent researches have observed that a small teacher-student capability gap can facilitate knowledge transfer. Inspired by that, we propose an evolutionary knowledge distillation approach to improve the transfer effectiveness of teacher knowledge. Instead of a fixed pre-trained teacher, an evolutionary teacher is learned online and consistently transfers intermediate knowledge to supervise student network learning on-the-fly. To enhance intermediate knowledge representation and mimicking, several simple guided modules are introduced between corresponding teacher-student blocks. In this way, the student can simultaneously obtain rich internal knowledge and capture its growth process, leading to effective student network learning. Extensive experiments clearly demonstrate the effectiveness of our approach as well as good adaptability in the low-resolution and few-sample scenarios.
Kangkai Zhang, Chunhui Zhang 0001, Shikun Li, Dan Zeng 0001, Shiming Ge
IEEE Trans. Circuits Syst. Video Technol.4
2022 Adaptive Material Matching for Hyperspectral Imagery Destriping
abstract
Due to instrument instability, slit contamination, and light interference, hyperspectral images often suffer from striping artifacts, which greatly impairs the data quality. Real hyperspectral data are usually characterized by a small amount of historical data, complex material distribution, insignificant periodicity of noise, and so on, which brings significant challenges for the destriping task. However, the assumptions made by traditional destriping methods are often inconsistent with these characteristics. To this end, we propose a novel destriping method based on adaptive material matching (MAM) without making explicit assumptions of hyperspectral data. Specifically, to identify pixels that belong to the same material, we propose a principal material analysis (PMA) to adaptively generate thresholds within each superpixel. The pixels are matched by thresholding their vertical gradients and leveraging both inner stripe gradient feature (ISGF) and neighbor-stripe geometry feature (NSGF). Correction pixels selected from the same material can then be used to calculate the offsets and gains of pixels to adjust adjacent columns. To further improve the stability of the destriping process, we generate a set of correction candidates for each column and select the optimal candidate by considering the prior distribution and destriping nonuniformity. The stripe noise within the whole image is finally removed by iteratively performing the correction between adjacent columns. We compare the proposed model against traditional and deep learning methods on both synthetic and real hyperspectral images. The promising results indicate that MAM can effectively remove the image stripes, retain original image information, and improve the nonuniformity.
Jia Li 0032, Junjie Zhang 0002, Kai Zhao 0012, Dan Zeng 0001
IEEE Trans. Geosci. Remote. Sens.5
2022 Guest Editorial: Learning From Noisy Multimedia Data
abstract
This special issue provides a premier forum for researchers in multimedia big data to share challenges and recent advancements in learning from noisy multimedia data. The multimedia age and its proliferation of devices and platforms is fueling exponential data growth. As computational power and deep learning algorithms rapidly evolve, the web has become a rich source of potential training data for robust machine learning, with search engines such as Google and Bing, Twitter, TikTok, Instagram, and short video sharing platforms offering large-scale data points in the hundreds of millions. The concurrent shift in the Internet to richer web data modalities such as text, audio, image, and video reveal further opportunities to leverage large-scale data for the automatic construction of a variety of datasets for model training and testing. However, the ubiquity of multimedia data means noise is a fundamental challenge, with ‘label noise’ and ‘domain mismatch’ the most critical issues in automatically collected datasets. Learning from noisy multimedia data tends towards poor performance, making it increasingly essential to address these challenges.
Jian Zhang 0002, Alan Hanjalic, Ramesh Jain 0001, Xian-Sheng Hua 0001, Shin'ichi Satoh 0001, Yazhou Yao, Dan Zeng 0001
IEEE Trans. Multim.7
2022 Deepfake Video Detection via Predictive Representation Learning
abstract
Increasingly advanced deepfake approaches have made the detection of deepfake videos very challenging. We observe that the general deepfake videos often exhibit appearance-level temporal inconsistencies in some facial components between frames, resulting in discriminative spatiotemporal latent patterns among semantic-level feature maps. Inspired by this finding, we propose a predictive representative learning approach termed Latent Pattern Sensing to capture these semantic change characteristics for deepfake video detection. The approach cascades a Convolution Neural Network-based encoder, a ConvGRU-based aggregator, and a single-layer binary classifier. The encoder and aggregator are pretrained in a self-supervised manner to form the representative spatiotemporal context features. Then, the classifier is trained to classify the context features, distinguishing fake videos from real ones. Finally, we propose a selective self-distillation fine-tuning method to further improve the robustness and performance of the detector. In this manner, the extracted features can simultaneously describe the latent patterns of videos across frames spatially and temporally in a unified way, leading to an effective and robust deepfake video detector. Extensive experiments and comprehensive analysis prove the effectiveness of our approach, e.g., achieving a very highest Area Under Curve (AUC) score of 99.94% on FaceForensics++ benchmark and surpassing 12 states of the art at least 7.90%@AUC and 8.69%@AUC on challenging DFDC and Celeb-DF(v2) benchmarks, respectively.
Shiming Ge, Fanzhao Lin, Chenyu Li 0001, Daichi Zhang, Weiping Wang 0005, Dan Zeng 0001
ACM Trans. Multim. Comput. Commun. Appl.6
2021 Neural Architecture Search for Joint Human Parsing and Pose Estimation
abstract
Human parsing and pose estimation are crucial for the understanding of human behaviors. Since these tasks are closely related, employing one unified model to perform two tasks simultaneously allows them to benefit from each other. However, since human parsing is a pixel-wise classification process while pose estimation is usually a regression task, it is non-trivial to extract discriminative features for both tasks while modeling their correlation in the joint learning fashion. Recent studies have shown that Neural Architecture Search (NAS) has the ability to allocate efficient feature connections for specific tasks automatically. With the spirit of NAS, we propose to search for an efficient network architecture (NPPNet) to tackle two tasks at the same time. On the one hand, to extract task-specific features for the two tasks and lay the foundation for the further searching of feature interaction, we propose to search their encoder-decoder architectures, respectively. On the other hand, to ensure two tasks fully communicate with each other, we propose to embed NAS units in both multi-scale feature interaction and high-level feature fusion to establish optimal connections between two tasks. Experimental results on both parsing and pose estimation benchmark datasets have demonstrated that the searched model achieves state-of-the-art performances on both tasks.1
Dan Zeng 0001, Yuhang Huang 0006, Qian Bao, Junjie Zhang 0002, Chi Su, Wu Liu 0005
ICCV1
2021 Detecting Deepfake Videos with Temporal Dropout 3DCNN
abstract
While the abuse of deepfake technology has brought about a serious impact on human society, the detection of deepfake videos is still very challenging due to their highly photorealistic synthesis on each frame. To address that, this paper aims to leverage the possible inconsistent cues among video frames and proposes a Temporal Dropout 3-Dimensional Convolutional Neural Network (TD-3DCNN) to detect deepfake videos. In the approach, the fixed-length frame volumes sampled from a video are fed into a 3-Dimensional Convolutional Neural Network (3DCNN) to extract features across different scales and identified whether they are real or fake. Especially, a temporal dropout operation is introduced to randomly sample frames in each batch. It serves as a simple yet effective data augmentation and can enhance the representation and generalization ability, avoiding model overfitting and improving detecting accuracy. In this way, the resulting video-level classifier is accurate and effective to identify deepfake videos. Extensive experiments on benchmarks including Celeb-DF(v2) and DFDC clearly demonstrate the effectiveness and generalization capacity of our approach.
Daichi Zhang, Chenyu Li 0001, Fanzhao Lin, Dan Zeng 0001, Shiming Ge
IJCAI4
2021 Some Results on Density Evolution of Nonbinary SC-LDPC Ensembles Over the BEC
abstract
In this paper, we present several theoretic results on density evolution (DE) for nonbinary spatially-coupled low-density parity-check (SC-LDPC) ensembles when the transmission takes place on the binary erasure channel (BEC). In specific, we establish the duality rule for entropy for the nonbinary variable-node (VN) and check-node (CN) operators in such a scenario. We define the partial order between densities and show that the VN and CN operators exhibit the property of partial order preservation. More importantly, we explicitly construct the potential functions for uncoupled and coupled DE recursions, the forms of which almost coincide with those for binary LDPC and SC-LDPC ensembles over general binary memoryless symmetric (BMS) channels. These theoretic findings greatly facilitate the proof of threshold saturation for nonbinary SC-LDPC ensembles on the BEC. Finally, we develop the threshold saturation theorem and its converse, following the lines established by S. Kumar et al.
Mengnan Xu, Dan Zeng 0001, Zhichao Sheng, Chongbin Xu
ISIT2
2021 Latent Pattern Sensing: Deepfake Video Detection via Predictive Representation Learning
abstract
Increasingly advanced deepfake approaches have made the detection of deepfake videos very challenging. We observe that the general deepfake videos often exhibit appearance-level temporal inconsistencies in some facial components between frames, resulting in discriminable spatiotemporal latent patterns among semantic-level feature maps. Inspired by this finding, we propose a predictive representative learning approach termed Latent Pattern Sensing to capture these semantic change characteristics for deepfake video detection. The approach cascades a CNN-based encoder, a ConvGRU-based aggregator and a single-layer binary classifier. The encoder and aggregator are pre-trained in a self-supervised manner to form the representative spatiotemporal context features. Finally, the classifier is trained to classify the context features, distinguishing fake videos from real ones. In this manner, the extracted features can simultaneously describe the latent patterns of videos across frames spatially and temporally in a unified way, leading to an effective deepfake video detector. Extensive experiments prove our approach’s effectiveness, e.g., surpassing 10 state-of-the-arts at least 7.92%@AUC on challenging Celeb-DF(v2) benchmark.
Shiming Ge, Fanzhao Lin, Chenyu Li 0001, Daichi Zhang, Jiyong Tan, Weiping Wang 0005, Dan Zeng 0001
MMAsia7
2021 Towards NIR-VIS Masked Face Recognition
abstract
Near-infrared to visible (NIR-VIS) face recognition is the most common case in heterogeneous face recognition, which aims to match a pair of face images captured from two different modalities. Existing deep learning based methods have made remarkable progress in NIR-VIS face recognition, while it encounters certain newly-emerged difficulties during the pandemic of COVID-19, since people are supposed to wear facial masks to cut off the spread of the virus. We define this task as NIR-VIS masked face recognition, and find it problematic with the masked face in the NIR probe image. First, the lack of masked face data is a challenging issue for the network training. Second, most of the facial parts (cheeks, mouth, nose etc.) are fully occluded by the mask, which leads to a large amount of loss of information. Third, the domain gap still exists in the remaining facial parts. In such scenario, the existing methods suffer from significant performance degradation caused by the above issues. In this paper, we aim to address the challenge of NIR-VIS masked face recognition from the perspectives of training data and training method. Specifically, we propose a novel heterogeneous training method to maximize the mutual information shared by the face representation of two domains with the help of semi-siamese networks. In addition, a 3D face reconstruction based approach is employed to synthesize masked face from the existing NIR image. Resorting to these practices, our solution provides the domain-invariant face representation which is also robust to the mask occlusion. Extensive experiments on three NIR-VIS face datasets demonstrate the effectiveness and cross-dataset-generalization capacity of our method.
Hailin Shi, Yinglu Liu, Dan Zeng 0001, Tao Mei 0001
IEEE Signal Process. Lett.4
2021 Flexible Auto-Weighted Local-Coordinate Concept Factorization: A Robust Framework for Unsupervised Clustering
abstract
Concept Factorization (CF) and its variants may produce inaccurate representation and clustering results due to the sensitivity to noise, hard constraint on the reconstruction error, and pre-obtained approximate similarities. To improve the representation ability, a novel unsupervised Robust Flexible Auto-weighted Local-coordinate Concept Factorization (RFA-LCF) framework is proposed for clustering high-dimensional data. Specifically, RFA-LCF integrates the robust flexible CF by clean data space recovery, robust sparse local-coordinate coding, and adaptive weighting into a unified model. RFA-LCF improves the representations by enhancing the robustness of CF to noise and errors, providing a flexible constraint on the reconstruction error and optimizing the locality jointly. For robust learning, RFA-LCF clearly learns a sparse projection to recover the underlying clean data space, and then the flexible CF is performed in the projected feature space. RFA-LCF also uses a L2,1-norm based flexible residue to encode the mismatch between the recovered data and its reconstruction, and uses the robust sparse local-coordinate coding to represent data using a few nearby basis concepts. For auto-weighting, RFA-LCF jointly preserves the manifold structures in the basis concept space and new coordinate space in an adaptive manner by minimizing the reconstruction errors on clean data, anchor points and coordinates. By updating the local-coordinate preserving data, basis concepts and new coordinates alternately, the representation abilities can be potentially improved. Extensive results on public databases show that RFA-LCF delivers the state-of-the-art clustering results compared with other related methods.
Zhao Zhang 0001, Yan Zhang 0053, Sheng Li 0001, Guangcan Liu, Dan Zeng 0001, Shuicheng Yan, Meng Wang 0001
IEEE Trans. Knowl. Data Eng.5
2021 Cascaded Correlation Refinement for Robust Deep Tracking
abstract
Recent deep trackers have shown superior performance in visual tracking. In this article, we propose a cascaded correlation refinement approach to facilitate the robustness of deep tracking. The core idea is to address accurate target localization and reliable model update in a collaborative way. To this end, our approach cascades multiple stages of correlation refinement to progressively refine target localization. Thus, the localized object could be used to learn an accurate on-the-fly model for improving the reliability of model update. Meanwhile, we introduce an explicit measure to identify the tracking failure and then leverage a simple yet effective look-back scheme to adaptively incorporate the initial model and on-the-fly model to update the tracking model. As a result, the tracking model can be used to localize the target more accurately. Extensive experiments on OTB2013, OTB2015, VOT2016, VOT2018, UAV123, and GOT-10k demonstrate that the proposed tracker achieves the best robustness against the state of the arts.
Shiming Ge, Chunhui Zhang 0001, Shikun Li, Dan Zeng 0001, Dacheng Tao
IEEE Trans. Neural Networks Learn. Syst.4
2020 Semi-Siamese Training for Shallow Face Learning
Hailin Shi, Yuchi Liu, Jun Wang 0127, Zhen Lei 0001, Dan Zeng 0001, Tao Mei 0001
ECCV (4)6
2020 Robust Visual Object Tracking with Two-Stream Residual Convolutional Networks
abstract
The current deep learning based visual tracking approaches have been very successful by learning the target classification and/or estimation model from a large amount of supervised training data in offline mode. However, most of them can still fail in tracking objects due to some more challenging issues such as dense distractor objects, confusing background, motion blurs, and so on. Inspired by the human “visual tracking” capability which leverages motion cues to distinguish the target from the background, we propose a Two-Stream Residual Convolutional Network (TS-RCN) for visual tracking, which successfully exploits both appearance and motion features for model update. Our TS-RCN can be integrated with existing deep learning based visual trackers. To further improve the tracking performance, we adopt a “wider” residual network ResNeXt as its feature extraction backbone. To the best of our knowledge, TS-RCN is the first end-to-end trainable two-stream visual tracking system, which makes full use of both appearance and motion features of the target. We have extensively evaluated the TS-RCN on most widely used benchmark datasets including VOT2018, VOT2019, and GOT-10K. The experiment results have successfully demonstrated that our two-stream model can greatly outperform the appearance-based tracker, and achieves state-of-the-art performance. The tracking system can run at up to 38.1 FPS.
Ning Zhang 0023, Jingen Liu, Dan Zeng 0001, Tao Mei 0001
ICPR4
2020 Talking Face Generation with Expression-Tailored Generative Adversarial Network
abstract
A key of automatically generating vivid talking faces is to synthesize identity-preserving natural facial expressions beyond audio-lip synchronization, which usually need to disentangle the informative features from multiple modals and then fuse them together. In this paper, we propose an end-to-end Expression-Tailored Generative Adversarial Network (ET-GAN) to generate an expression enriched talking face video of arbitrary identity. Different from talking face generation based on identity image and audio, an expressional video of arbitrary identity serves as the expression source in our approach. Expression encoder is proposed to disentangle expression-tailored representation from the guiding expressional video, while audio encoder disentangles audio-lip representation. Instead of using single image as identity input, multi-image identity encoder is proposed by learning different views of faces and merging a unified representation. Multiple discriminators are exploited to keep both image-aware and the video-aware realistic details, including a spatial-temporal discriminator for visual continuity of expression synthesis and facial movements. We conduct extensive experimental evaluations on quantitative metrics, expression retention quality and audio-visual synchronization. The results show the effectiveness of our ET-GAN in generating high quality expressional talking face videos against existing state-of-the-arts.
Dan Zeng 0001, Shiming Ge
ACM Multimedia1
2020 Accurate UAV Tracking with Distance-Injected Overlap Maximization
abstract
UAV tracking is usually challenged by the dual-dynamic disturbances that arise from not only diverse moving target but also motion camera, leading to a more serious model drift issue than traditional visual tracking. In this work, we propose to alleviate this issue with distance-injected overlap maximization. Our idea is improving the accuracy of target localization by deriving a conceptually simple target localization loss and a global feature recalibration scheme in a mutual reinforced way. In particular, the target localization loss is designed by simply incorporating the normalized distance of target offset and generic semantic IoU loss, resulting in the distance-injected semantic IoU loss, and its minimal solution can alleviate the drift problem caused by camera motion. Moreover, the deep feature extractor is reconstructed and alternated with a feature recalibration network, which can leverage the global information to recalibrate significant features and suppress negligible features. Following by multi-scale feature concat, the proposed tracker can improve the discriminative capability of feature representation for UAV targets on the fly. Extensive experimental results on four benchmarks, i.e. UAV123, UAVDT, DTB70, and VisDrone, demonstrate the superiority of the proposed tracker against existing state-of-the-arts on UAV tracking.
Chunhui Zhang 0001, Shiming Ge, Kangkai Zhang, Dan Zeng 0001
ACM Multimedia4
2020 AI-SAS: Automated In-match Soccer Analysis System
abstract
Real-time in-match soccer statistics provide continuous tracking of soccer ball and player positions and speeds, enabling advanced analytics. Currently, only elite soccer leagues have the luxury of tracking in-match soccer statistics operated with a large number of trained personnel. In this work, we present an Automated In-match Soccer Analysis System (AI-SAS), using a domain-knowledge-based multi-view global tracking. This system tracks player team, position, and speed automatically, providing real-time in-match team- and individual-level statistics and analyses. In comparison with the latest soccer analysis systems, AI-SAS is more scalable in streaming multiple video sources for real-time process and more flexible in hosting plug-and-play deep-learning-based tracking-by-detection algorithms. The global multi-view tracking also overcomes the single-view limitation and improves the tracking accuracy.
Ning Zhang 0023, Wei Zhang 0031, Dan Zeng 0001, Jingen Liu, Tao Mei 0001
ACM Multimedia5
2020 Velocity Analysis of BP Decoding Waves for SC-LDPC Ensembles on BMS Channels: An Interpolation-Based Approach
abstract
This paper is concerned with the dynamics of spatially-coupled low-density parity-check (SC-LDPC) ensembles for transmission on general binary memoryless symmetric (BMS) channels under belief-propagation (BP) decoding. The decoding waves of such ensembles are found to exhibit solitonic behavior, propagating along the Tanner graphs at asymptotically constant velocities. A low-complexity approach termed interpolated density evolution (IDE) is proposed to predict the decoding wave velocities. In this approach, the densities of a decoding wave are approximated by interpolating between some fixed points of an uncoupled DE recursion with one-dimensional functions. Two transfer functions for updating these interpolation functions in the IDE recursion are established and a simple strategy is introduced to deal with the coexistence of multiple transition regions. In addition, a threshold analysis is developed based on two ansatzes, explaining why our approach can achieve a good trade-off between computational cost and accuracy, as illustrated with some numerical examples at the end of this paper.
Zhangyou Peng, Dan Zeng 0001, Mingjun Dai
IEEE Trans. Commun.4
2020 Occluded Face Recognition in the Wild by Identity-Diversity Inpainting
abstract
Face recognition has achieved advanced development by using convolutional neural network (CNN) based recognizers. Existing recognizers typically demonstrate powerful capacity in recognizing un-occluded faces, but often suffer from accuracy degradation when directly identifying occluded faces. This is mainly due to insufficient visual and identity cues caused by occlusions. On the other hand, generative adversarial network (GAN) is particularly suitable when it needs to reconstruct visually plausible occlusions by face inpainting. Motivated by these observations, this paper proposes identity-diversity inpainting to facilitate occluded face recognition. The core idea is integrating GAN with an optimized pre-trained CNN recognizer which serves as the third player to compete with the generator by distinguishing diversity within the same identity class. To this end, a collect of identity-centered features is applied in the recognizer as supervision to enable the inpainted faces clustering towards their identity centers. In this way, our approach can benefit from GAN for reconstruction and CNN for representation, and simultaneously addresses two challenging tasks, face inpainting and face recognition. Experimental results compared with 4 state-of-the-arts prove the efficacy of the proposed approach.
Shiming Ge, Chenyu Li 0001, Shengwei Zhao, Dan Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.4
2019 Robust Deep Tracking with Two-step Augmentation Discriminative Correlation Filters
abstract
Recently, deep trackers have proven success in visual tracking due to their powerful feature representation. Among them, discriminative correlation filter (DCF) paradigm is widely used. However, these trackers are still difficult to learn an adaptive appearance model of the object due to the limited data available. To address that, this paper proposes a two-step augmentation discriminative correlation filters (TADCF) approach to improve robustness. Firstly, we propose an online frame augmentation scheme to obtain rich and robust deep features which can effectively alleviate background distractors, leading to better generalization and adaptation of the learned model. Secondly, an object augmentation mechanism is implemented by exploiting rotation continuity restriction, which simultaneously models target appearance changes from rotation and scale variations. Extensive experiments on four benchmarks illustrate that the proposed approach performs favorably against state-of-the-art trackers.
Chunhui Zhang 0001, Shiming Ge, Yingying Hua, Dan Zeng 0001
ICME4
2019 Defending Against Adversarial Examples via Soft Decision Trees Embedding
abstract
Convolutional neural networks (CNNs) have shown vulnerable to adversarial examples which contain imperceptible perturbations. In this paper, we propose an approach to defend against adversarial examples with soft decision trees embedding. Firstly, we extract the semantic features of adversarial examples with a feature extraction network. Then, a specific soft decision tree is trained and embedded to select the key semantic features for each feature map from convolutional layers and the selected features are fed to a light-weight classification network. To this end, we use the probability distributions of each tree node to quantify the semantic features. In this way, some small perturbations can be effectively removed and the selected features are more discriminative in identifying adversarial examples. Moreover, the influence of adversarial perturbations on classification can be reduced by migrating the interpretability of soft decision trees into the black-box neural networks. We conduct experiments to defend the state-of-the-art adversarial attacks. The experimental results demonstrate that our proposed approach can effectively defend against these attacks and improve the robustness of deep neural networks.
Yingying Hua, Shiming Ge, Xindi Gao, Xin Jin 0015, Dan Zeng 0001
ACM Multimedia5
2019 MetaAdvDet: Towards Robust Detection of Evolving Adversarial Attacks
abstract
Deep neural networks (DNNs) are vulnerable to the adversarial attack which is maliciously implemented by adding human-imperceptible perturbation to images and thus leads to incorrect prediction. Existing studies have proposed various methods to detect the new adversarial attacks. However, new attack methods keep evolving constantly and yield new adversarial examples to bypass the existing detectors. It needs to collect tens of thousands samples to train detectors, while the new attacks evolve much more frequently than the high-cost data collection. Thus, this situation leads the newly evolved attack samples to remain in small scales. To solve such few-shot problem with the evolving attacks, we propose a meta-learning based robust detection method to detect new adversarial attacks with limited examples. Specifically, the learning consists of a double-network framework: a task-dedicated network and a master network which alternatively learn the detection capability for either seen attack or a new attack. To validate the effectiveness of our approach, we construct the benchmarks with few-shot-fashion protocols based on three conventional datasets, i.e. CIFAR-10, MNIST and Fashion-MNIST. Comprehensive experiments are conducted on them to verify the superiority of our approach with respect to the traditional adversarial attack detection methods. The implementation code is available online.
Chen Ma 0003, Hailin Shi, Li Chen 0031, Jun-Hai Yong, Dan Zeng 0001
ACM Multimedia6
2019 Proposal pyramid networks for fast face detection
Dan Zeng 0001, Fan Zhao 0004, Shiming Ge, Wei Shen 0002, Zhijiang Zhang
Inf. Sci.1
2019 Cloud Detection Using Super Pixel Classification and Semantic Segmentation
Dan Zeng 0001, Qi Tian 0001
J. Comput. Sci. Technol.3
2019 Fast cascade face detection with pyramid network
Dan Zeng 0001, Fan Zhao 0004, Shiming Ge, Wei Shen 0002
Pattern Recognit. Lett.1
2018 Volumeter: 3D human body parameters measurement with a single Kinect
abstract
3D human body parameters measurement is a challenging task due to two main reasons: (i) it is difficult to reconstruct 3D human model due to flexible deformation of non‐rigid body during images capturing process and (ii) there lies a gap between 3D model and body parameters. To address these two issues, a 3D human body parameters measurement system is represented. With the object freely spinning in front of a Kinect, body parameters are calculated. To reduce registration errors caused by body deformation while rotating, a piecewise tracking and mapping algorithm based on KinectFusion framework is proposed. Then model–model iterative closest point and non‐rigid constraints are introduced to optimise alignments and disambiguate different surfaces caused by aliasing in the piecewise strategy. Finally, a novel method is presented to measure the volume and perimeter of human body with the truncated signed distance function values of voxels. Extensive experimental results show that the proposed method achieves comparable accuracy to the state of the arts, and the error of volume and perimeter measurements are 2.0 and 5.8%, respectively.
Qinzhu He, Yijun Ji, Dan Zeng 0001, Zhijiang Zhang
IET Comput. Vis.3
2018 Bag of Shape Features with a learned pooling function for shape recognition
Wei Shen 0002, Chenting Du, Yuan Jiang 0002, Dan Zeng 0001, Zhijiang Zhang
Pattern Recognit. Lett.4
2018 Multi-oriented text detection from natural scene images based on a CNN and pruning non-adjacent graph edges
Yuanwang Wei, Wei Shen 0002, Dan Zeng 0001, Lihua Ye, Zhijiang Zhang
Signal Process. Image Commun.3
2017 Shape recognition by bag of contour fragments with a learned pooling function
abstract
Bag of Contour Fragments (BoCF), derived from the well-known Bag-of-Features (BoF), is an effective framework for shape representation. The feature pooling in this framework is a critical step, while either max pooling or average pooling is not a learnable process. In this paper, we aim at learning a pooling function which is adaptive to the input contour fragment features instead. Towards this end, we formulate our pooling function as a weighted sum of max pooling and average pooling, where the weight is expressed by an activation function of the input contour fragment features. To automatically learn this weight, the output of the pooling function is fed into a SVM classifier and they are trained jointly to minimize a shape classification loss. Experimental results on several standard shape datasets demonstrate the effectiveness of the proposed learned pooling function, which can achieve considerable improvements compared with BoCF.
Wei Shen 0002, Wenjing Gao, Yuan Jiang 0002, Dan Zeng 0001, Zhijiang Zhang
ICIP4
2017 Neighborhood geometry based feature matching for geostationary satellite remote sensing image
Dan Zeng 0001, Wei Shen 0002, Qi Tian 0001
Neurocomputing1
2017 Text detection in scene images based on exhaustive segmentation
Yuanwang Wei, Zhijiang Zhang, Wei Shen 0002, Dan Zeng 0001, Mei Fang, Shifu Zhou
Signal Process. Image Commun.4
2016 Face database generation based on text-video correlation
Dan Zeng 0001, Yixin Bao, Fan Zhao 0004, Qi Tian 0001
Neurocomputing1
2016 Shape recognition by bag of skeleton-associated contour parts
Wei Shen 0002, Yuan Jiang 0002, Wenjing Gao, Dan Zeng 0001, Xinggang Wang
Pattern Recognit. Lett.4
2016 Spatial-temporal convolutional neural networks for anomaly detection and localization in crowded scenes
Shifu Zhou, Wei Shen 0002, Dan Zeng 0001, Mei Fang, Yuanwang Wei, Zhijiang Zhang
Signal Process. Image Commun.3
2015 Unusual event detection in crowded scenes by trajectory analysis
abstract
Anomaly detection in crowded scenes is a challenge task due to variation of the definitions for both abnormality and normality, the low resolution on the target, ambiguity of appearance, and severe occlusions of inter-object. In this paper, we propose a novel statistical framework to detect abnormal behaviors of the crowded scene by modeling trajectories of pedestrians. First, the trajectories are acquired by Kanade-Lucas-Tomasi Feature Tracker (KLT). Then trajectories are grouped to form representative trajectories, which characterize the underlying motion patterns of the crowd. Finally, trajectories are modeled by Multi-Observation Hidden Markov Model (MOHMM) to determine whether frames are normal or abnormal. The experiments are conducted on a well-known crowded scene dataset. Experimental results show that the proposed method can capture abnormal crowd behaviors successfully and achieves state-of-the-art performances.
Shifu Zhou, Wei Shen 0002, Dan Zeng 0001, Zhijiang Zhang
ICASSP3
2015 Augmented Feature Fusion for Image Retrieval System
abstract
The performance of current image retrieval system is largely determined by the quality and discriminative capability of features. Therefore, using what features and how to effectively combine the power of appropriate features are important in the system. We adopt the reciprocal neighbor based graph fusion approach for feature fusion. More importantly, we explicitly augment the original approach with the following two strategies: 1) we investigate the most suitable feature combinations on various datasets, including the deep learning feature, which has been popular for image retrieval recently; 2) we further improve the robustness of original graph fusion approach by the SVM prediction strategy.
Yang Zhou 0017, Dan Zeng 0001, Shiliang Zhang, Qi Tian 0001
ICMR2
2014 Regularity Guaranteed Human Pose Correction
Wei Shen 0002, Rui Lei, Dan Zeng 0001, Zhijiang Zhang
ACCV (2)3
2012 Three-dimensional deformation in curl vector field
abstract
Deformation is an important research topic in graphics. There are two key issues in mesh deformation: (1) self-intersection and (2) volume preserving. In this paper, we present a new method to construct a vector field for volume-preserving mesh deformation of free-form objects. Volume-preserving is an inherent feature of a curl vector field. Since the field lines of the curl vector field will never intersect with each other, a mesh deformed under a curl vector field can avoid self-intersection between field lines. Designing the vector field based on curl is useful in preserving graphic features and preventing self-intersection. Our proposed algorithm introduces distance field into vector field construction; as a result, the shape of the curl vector field is closely related to the object shape. We define the construction of the curl vector field for translation and rotation and provide some special effects such as twisting and bending. Taking into account the information of the object, this approach can provide easy and intuitive construction for free-form objects. Experimental results show that the approach works effectively in real-time animation.
Dan Zeng 0001, Da-yue Zheng
J. Zhejiang Univ. Sci. C1