Le Zhang 0001

dblp:03/4043-1 · DBLP profile ↗
← Back
102ranked-venue papers
25as first author
64since 2021 · last 2026
0000-0002-6930-8674ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 55 · 13 first-author · 39 since 2021Graphics, computer vision, multimedia, augmented reality and games · 42 · 1 first-author · 31 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 3 since 2021Systems, architecture and hardware · 5 · 4 first-authorComputer networks · 5 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 1 since 2021Security and privacy · 2 · 2 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Enhancing Generalization of Depth Estimation Foundation Model via Weakly-Supervised Adaptation with Regularization
abstract
The emergence of foundation models has substantially advanced zero-shot generalization in monocular depth estimation (MDE), as exemplified by the Depth Anything series. However, given access to some data from downstream tasks, a natural question arises: can the performance of these models be further improved? To this end, we propose WeSTAR, a parameter-efficient framework that performs \textbf{We}akly supervised \textbf{S}elf-\textbf{T}raining \textbf{A}daptation with \textbf{R}egularization, designed to enhance the robustness of MDE foundation models in unseen and diverse domains. We first adopt a dense self-training objective as the primary source of structural self-supervision. To further improve robustness, we introduce semantically-aware hierarchical normalization, which exploits instance-level segmentation maps to perform more stable and multi-scale structural normalization. Beyond dense supervision, we introduce a cost-efficient weak supervision in the form of pairwise ordinal depth annotations to further guide the adaptation process, which enforces informative ordinal constraints to mitigate local topological errors. Finally, a weight regularization loss is employed to anchor the LoRA updates, ensuring training stability and preserving the model's generalizable knowledge. Extensive experiments on both realistic and corrupted out-of-distribution datasets under diverse and challenging scenarios demonstrate that WeSTAR consistently improves generalization and achieves state-of-the-art performance across a wide range of benchmarks.
Yongyi Su, Le Zhang 0001, Xun Xu 0002
AAAI4
2026 Graph Meets Deep Unfolding: An Interpretable Mutual-benefit Multi-view Learning Network
abstract
Significant efforts have been focused on enhancing the utilization of multiple node features and topological structures in multi-view graph learning through explicit model-driven and implicit deep learning-based methodologies. The former excels in embedding prior knowledge, thereby offering theoretical interpretability but is limited in application flexibility due to manual parameter selection. In contrast, the latter leverages automatic differentiation, providing greater flexibility but lacking theoretical interpretability due to their opaque nature. Motivated by these observations, we propose an interpretable deep unfolding network for mutual-benefit multi-view graph learning, aiming to combine the strengths of both approaches. Specifically, we employ the Alternating Direction Method of Multipliers (ADMM) to solve a multi-view graph learning model with sparse and low-rank constraints. This solution is then integrated into deep unfolding networks to enhance interpretability. Furthermore, we convert optimization conditions into implicit losses and utilize automatic differentiation to update parameters, reducing the need for manual tuning and increasing flexibility. This integration optimizes multi-view learning for a graph representation that balances interpretability and flexibility. Empirical evaluations on six diverse datasets demonstrate the effectiveness and superiority of the proposed method over state-of-the-art approaches.
Renjie Lin, Hongzhi He, Yilin Wu 0001, Shide Du, Le Zhang 0001
AAAI5
2026 Coupled tensor train decomposition in federated learning
Xiangtao Zhang, Eleftherios Kofidis, Ruituo Wu, Ce Zhu, Le Zhang 0001, Yipeng Liu 0001
Pattern Recognit.5
2026 FunOTTA: On-the-Fly Adaptation on Cross-Domain Fundus Image via Stable Test-Time Training
abstract
Fundus images are essential for the early screening and detection of eye diseases. While deep learning models using fundus images have significantly advanced the diagnosis of multiple eye diseases, variations in images from different imaging devices and locations (known as domain shifts) pose challenges for deploying pre-trained models in real-world applications. To address this, we propose a novel Fundus On-the-fly Test-Time Adaptation (FunOTTA) framework that effectively generalizes a fundus image diagnosis model to unseen environments, even under strong domain shifts. FunOTTA stands out for its stable adaptation process by performing dynamic disambiguation in the memory bank while minimizing harmful prior knowledge bias. We also introduce a new training objective during adaptation that enables the classifier to incrementally adapt to target patterns with reliable class conditional estimation and consistency regularization. We compare our method with several state-of-the-art test-time adaptation (TTA) pipelines. Experiments on cross-domain fundus image benchmarks across two diseases demonstrate the superiority of the overall framework and individual components under different backbone networks. Code is available at https://github.com/Casperqian/FunOTTA.
Le Zhang 0001, Yipeng Liu 0001, Ce Zhu, Fan Zhang 0013
IEEE Trans. Medical Imaging2
2026 Interpretable Multi-View Feature Representation via Physical Partial Differential Equation
abstract
Graph Neural Networks (GNNs) have become a powerful tool for learning representations from graph-structured data, leveraging the relationships between nodes and their features. Despite their success, they often lack interpretability due to the black-box nature of neural networks, and further development may be limited. Moreover, previous GNN-based multi-view methods typically rely on simple feature fusion techniques such as weighted averaging or concatenation, which fail to capture the complex dependencies between views. In this paper, we propose a novel framework, namely Interpretable Multi-View Feature Representation via physical partial differential equation (IMvFR), to address these limitations in the context of multi-view semi-supervised learning. By integrating GNNs with partial differential equations (PDEs), we model the evolution of multi-view feature representations as a dynamic process. This provides a natural and interpretable framework for understanding how information flows between different views, overcoming the black-box nature of traditional GNNs. Additionally, we formulate multi-view feature representations as an initial-value problem within the framework of PDEs, providing a clear and interpretable mechanism for label propagation and feature fusion, thus facilitating the acquisition of global and local information between views. Comprehensive experimental results on eight datasets demonstrate that the proposed method achieves superior performance compared with state-of-the-art methods.
Renjie Lin, Li Xu 0002, Le Zhang 0001, Shiping Wang
IEEE Trans. Multim.5
2025 WiFi CSI Based Temporal Activity Detection via Dual Pyramid Network
abstract
We address the challenge of WiFi-based temporal activity detection and propose an efficient Dual Pyramid Network that integrates Temporal Signal Semantic Encoders and Local Sensitive Response Encoders. The Temporal Signal Semantic Encoder splits feature learning into high and low-frequency components, using a novel Signed Mask-Attention mechanism to emphasize important areas and downplay unimportant ones, with the features fused using ContraNorm. The Local Sensitive Response Encoder captures fluctuations without learning. These feature pyramids are then combined using a new cross-attention fusion mechanism. We also introduce a dataset with over 2,114 activity segments across 553 WiFi CSI samples, each lasting around 85 seconds. Extensive experiments show our method outperforms challenging baselines.
Le Zhang 0001, Bing Li 0002, Yingjie Zhou 0001, Zhenghua Chen, Ce Zhu
AAAI2
2025 Rethinking Token Reduction with Parameter-Efficient Fine-Tuning in ViT for Pixel-Level Tasks
abstract
Parameter-Efficient fine-tuning (PEFT) adapts pre-trained models to new tasks by updating only a small subset of parameters, achieving efficiency but still facing significant inference costs driven by input token length. This challenge is even more pronounced in pixel-level tasks, which require longer input sequences compared to image-level tasks. Although token reduction (TR) techniques can help reduce computational demands, they often lead to homogeneous attention patterns that compromise performance in pixel-level scenarios. This study underscores the importance of maintaining attention diversity for these tasks and proposes to enhance attention diversity while ensuring the completeness of token sequences. Our approach effectively reduces the number of tokens processed within transformer blocks, improving computational efficiency without sacrificing performance on several pixel-level tasks. We also demonstrate the superior generalization capability of our proposed method compared to challenging baseline models. The source code will be made available at https://github.com/AVC2-UESTC/DAR-TR-PEFT.
Ao Li 0007, Hu Yao, Ce Zhu, Le Zhang 0001
CVPR5
2025 Handling Spatial-Temporal Data Heterogeneity for Federated Continual Learning via Tail Anchor
abstract
Federated continual learning (FCL) allows each client to continually update its knowledge from task streams, enhancing the applicability of federated learning in real-world scenarios. However, FCL needs to address not only spatial data heterogeneity between clients but also temporal data heterogeneity between tasks. In this paper, empirical experiments demonstrate that such input-level heterogeneity significantly affects the model’s internal parameters and outputs, leading to severe spatial-temporal catastrophic forgetting of local and previous knowledge. To this end, we propose Federated Tail Anchor (FedTA) to mix trainable Tail Anchor with the frozen output features to adjust their position in the feature space, thereby overcoming parameter-forgetting and output-forgetting. Three novel components are also included: Input Enhancement for improving the performance of pre-trained models on downstream tasks; Selective Input Knowledge Fusion for fusion of heterogeneous local knowledge on the server; and Best Global Prototype Selection for finding the best anchor point for each class in the feature space. Extensive experiments demonstrate that FedTA not only outperforms existing FCL methods but also effectively preserves the relative positions of features. Code is available at: https://github.com/SkyOfBeginning/FedTA_CVPR2025.
Hao Yu 0023, Xin Yang 0012, Le Zhang 0001, Hanlin Gu, Tianrui Li 0001, Lixin Fan, Qiang Yang 0001
CVPR3
2025 Subspace Constraint and Contribution Estimation for Heterogeneous Federated Learning
abstract
Heterogeneous Federated Learning (HFL) has received widespread attention due to its adaptability to different models and data. The HFL approach utilizing auxiliary models for knowledge transfer can further enhance flexibility. However, existing frameworks face the challenges of local overfitting and aggregation bias. To address these issues, we propose FedSCE. By restricting specific layers of the local model updates to a subspace, FedSCE reduces the degrees of freedom of the update, enhances generalization, and mitigates the risk of overfitting. The subspace is dynamically updated to ensure coverage of the latest model update trajectory. Additionally, FedSCE evaluates client contributions based on the update distance of the auxiliary model in feature space and parameter space, achieving adaptive weighted aggregation. We validate our approach in both feature-skewed and label-skewed scenarios, demonstrating that on Office10, our method exceeds the best baseline by 3.87%. The code will be available at https://github.com/AVC2-UESTC/FedSCE.git.
Xiangtao Zhang, Ao Li 0007, Yipeng Liu 0001, Fan Zhang 0013, Ce Zhu, Le Zhang 0001
CVPR7
2025 Codar: Complex-valued Neural Network for Crossing-Floor Intrusion Detection via WiFi
abstract
WiFi systems offer enormous potential for device-free human intrusion detection. Current methods often require routers to be deployed in multiple adjacent rooms on the same floor, which is redundant and costly. To solve this, we introduce the first work on intrusion detection in the crossing-floor scenario via WiFi. Routers on different floors are utilized without major modifications to the existing router layout. Many previous works require a high sample rate and ignore the phase information. In this paper, we propose Codar, a complex-valued LSTM-CNN neural network. The LSTM effectively captures temporal dependencies at a low sample rate in harsh propagation environments. Moreover, amplitude and phase features are explored jointly by complex-valued operations. Experimental results demonstrate Codar achieves 95%, 94.5%, and 99% accuracy for intrusion detection, user identification, and intruded floor identification, surpassing competitive methods. The code and dataset are available at https://github.com/ouweiting/Codar.
Weiting Ou, Yipeng Liu 0001, Bing Li 0002, Le Zhang 0001, Ce Zhu
ICASSP5
2025 MobileIE: An Extremely Lightweight and Effective ConvNet for Real-Time Image Enhancement on Mobile Devices
abstract
Recent advancements in deep neural networks have driven significant progress in image enhancement (IE). However, deploying deep learning models on resource-constrained platforms, such as mobile devices, remains challenging due to high computation and memory demands. To address these challenges and facilitate real-time IE on mobile, we introduce an extremely lightweight Convolutional Neural Network (CNN) framework with around 4K parameters. Our approach integrates reparameterization with an Incremental Weight Optimization strategy to ensure efficiency. Additionally, we enhance performance with a Feature Self-Transform module and a Hierarchical Dual-Path Attention mechanism, optimized with a Local Variance-Weighted loss. With this efficient framework, we are the first to achieve real-time IE inference at up to 1,100 frames per second (FPS) while delivering competitive image quality, achieving the best trade-off between speed and performance across multiple IE tasks. The code will be available at https://github.com/AVC2-UESTC/MobileIE.git.
Hailong Yan, Ao Li 0007, Xiangtao Zhang, Zhe Liu 0019, Zenglin Shi, Ce Zhu, Le Zhang 0001
ICCV7
2025 TRR-LGF: a Simple yet Efficient Classification Network
abstract
Hybrid models that combine convolution and self attention are popular for efficient local feature extraction and capturing long-range dependencies. However, these models often:1) only explore local and global features; 2) flatten high-order features at the output layer, which limit feature hierarchy exploration and the feature utility in the output layer. To address these issues, this paper introduces Tensor Ring Regression with Local-to-Global Features (TRR-LGF), a simple and effective classification network. It uses a local-to-global learning framework to capture diverse features at multiple scales. Additionally, a tensor ring regression layer replaces the linear output layer, preserving high-order feature structure and reducing parameters. Experimental results show that TRR-LGF outperforms existing state-of-the-art methods on various datasets, especially in noisy and sample-imbalanced settings. Furthermore, the model utilizes 7.6M parameters, reducing computational requirements by about 50% compared to the multilayer perceptron output layer. The code is available at https://github.com/Calcium-Oxide/TRR-LGF.
Zhen Long, Hu Yao, Yipeng Liu 0001, Le Zhang 0001, Ce Zhu
ICME5
2025 OIMGC-Net: Optimization-inspired Interpretable Multi-view Graph Clustering Network
abstract
Deep multi-view graph clustering seeks to integrate diverse graph feature sets and uncover consistent information across multiple views. While extensive prior research has utilized various neural network architectures to address multi-view graph clustering challenges, these approaches exhibit notable limitations: 1) The ''black-box'' nature of deep learning models, which obscures their internal mechanisms and impedes interpretability; 2) Insufficient efforts aim to capture low-dimensional representations through graphs that reflect intuitive clustering structures and reduce computational cost. To address these limitations, this paper introduces an interpretable multi-view graph clustering framework constructed with optimization-inspired modules. The proposed approach formulates low-dimensional clustering representation learning from graph matrices as an optimization problem, deriving an iterative solution rooted in this formulation. By seamlessly bridging this optimization process to a deep network architecture, the model learns a low-dimensional clustering representation for graph-structured data across multiple views while adhering to the iterative optimization principles and reducing computational costs. This transparent network design enhances the interpretability of multi-view clustering, enabling intuitive and human-understandable learning of clustering structures. Extensive experimental evaluations validate the proposed framework's superiority over state-of-the-art methods in multi-view clustering tasks while ensuring interpretability and reducing computational costs.
Renjie Lin, Shide Du, Shiping Wang, Le Zhang 0001
ACM Multimedia5
2025 SlimHead: Rethinking the Efficiency Bottleneck in Dense Object Detection
Zhaohui Zheng 0003, Ping Wang 0072, Le Zhang 0001, Xiang Li 0041, Qibin Hou, Ming-Ming Cheng
PRCV (16)4
2025 Correction to: Deep negative correlation classification
Le Zhang 0001, Qibin Hou, Yun Liu 0011, Jiawang Bian, Xun Xu 0002, Joey Tianyi Zhou, Ce Zhu
Mach. Learn.1
2025 Towards Real Zero-Shot Camouflaged Object Segmentation Without Camouflaged Annotations
abstract
Camouflaged Object Segmentation (COS) faces significant challenges due to the scarcity of annotated data, where meticulous pixel-level annotation is both labor-intensive and costly, primarily due to the intricate object-background boundaries. Addressing the core question, "Can COS be effectively achieved in a zero-shot manner without manual annotations for any camouflaged object?", we propose an affirmative solution. We examine the learned attention patterns for camouflaged objects and introduce a robust zero-shot COS framework. Our findings reveal that while transformer models for salient object segmentation (SOS) prioritize global features in their attention mechanisms, camouflaged object segmentation exhibits both global and local attention biases. Based on these findings, we design a framework that adapts with the inherent local pattern bias of COS while incorporating global attention patterns and a broad semantic feature space derived from SOS. This enables efficient zero-shot transfer for COS. Specifically, We incorporate a Masked Image Modeling (MIM) based image encoder optimized for Parameter-Efficient Fine-Tuning (PEFT), a Multimodal Large Language Model (M-LLM), and a Multi-scale Fine-grained Alignment (MFA) mechanism. The MIM encoder captures essential local features, while the PEFT module learns global and semantic representations from SOS datasets. To further enhance semantic granularity, we leverage the M-LLM to generate caption embeddings conditioned on visual cues, which are meticulously aligned with multi-scale visual features via MFA. This alignment enables precise interpretation of complex semantic contexts. Moreover, we introduce a learnable codebook to represent the M-LLM during inference, significantly reducing computational demands while maintaining performance. Our framework demonstrates its versatility and efficacy through rigorous experimentation, achieving state-of-the-art performance in zero-shot COS with $F_{\beta }^{w}$Fβw scores of 72.9% on CAMO and 71.7% on COD10K. By removing the M-LLM during inference, we achieve an inference speed comparable to that of traditional end-to-end models, reaching 18.1 FPS. Additionally, our method excels in polyp segmentation, and underwater scene segmentation, outperforming challenging baselines in both zero-shot and supervised settings, thereby implying its potentiality in various segmentation tasks.
Tian-Zhu Xiang, Ao Li 0007, Ce Zhu, Le Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.7
2025 STADe: Sensory Temporal Action Detection via Temporal-Spectral Representation Learning
abstract
Temporal action detection (TAD) is a vital challenge in computer vision and the Internet of Things, aiming to detect and identify actions within temporal sequences. While TAD has primarily been associated with video data, its applications can also be extended to sensor data, opening up opportunities for various real-world applications. However, applying existing TAD models to sensory signals presents distinct challenges such as varying sampling rates, intricate pattern structures, and subtle, noise-prone patterns. In response to these challenges, we propose a Sensory Temporal Action Detection (STADe) model. STADe leverages Fourier kernels and adaptive frequency filtering to adaptively capture the nuanced interplay of temporal and frequency features underlying complex patterns. Moreover, STADe embraces adaptability by employing deep fusion at varying resolutions and scales, making it versatile enough to accommodate diverse data characteristics, such as the wide spectrum of sampling rates and action durations encountered in sensory signals. Unlike conventional models with unidirectional category-to-proposal dependencies, STADe adopts a cross-cascade predictor to introduce bidirectional and temporal dependencies within categories. To extensively evaluate STADe and promote future research in sensory TAD, we establish three diverse datasets using various sensors, featuring diverse sensor types, action categories, and sampling rates. Experiments across one public and our three new datasets demonstrate STADe's superior performance over state-of-the-art TAD models in sensory TAD tasks.
Bing Li 0002, Haotian Duan, Yun Liu 0011, Le Zhang 0001, Wei Cui 0002, Joey Tianyi Zhou
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Exploring Frequency-Inspired Optimization in Transformer for Efficient Single Image Super-Resolution
abstract
Transformer-based methods have exhibited remarkable potential in single image super-resolution (SISR) by effectively extracting long-range dependencies. However, most of the current research in this area has prioritized the design of transformer blocks to capture global information, while overlooking the importance of incorporating high-frequency priors, which we believe could be beneficial. In our study, we conducted a series of experiments and found that transformer structures are more adept at capturing low-frequency information, but have limited capacity in constructing high-frequency representations when compared to their convolutional counterparts. Our proposed solution, the cross-refinement adaptive feature modulation transformer (CRAFT), integrates the strengths of both convolutional and transformer structures. It comprises three key components: the high-frequency enhancement residual block (HFERB) for extracting high-frequency information, the shift rectangle window attention block (SRWAB) for capturing global information, and the hybrid fusion block (HFB) for refining the global representation. To tackle the inherent intricacies of transformer structures, we introduce a frequency-guided post-training quantization (PTQ) method aimed at enhancing CRAFT's efficiency. These strategies incorporate adaptive dual clipping and boundary refinement. To further amplify the versatility of our proposed approach, we extend our PTQ strategy to function as a general quantization method for transformer-based SISR techniques. Our experimental findings showcase CRAFT's superiority over current state-of-the-art methods, both in full-precision and quantization scenarios. These results underscore the efficacy and universality of our PTQ strategy.
Ao Li 0007, Le Zhang 0001, Yun Liu 0011, Ce Zhu
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Multi-Channel Equilibrium Graph Neural Network for Multi-View Semi-Supervised Learning
abstract
In practical applications, the difficulty of multi-view data annotation poses a challenge for multi-view semi-supervised learning. Although some graph-based approaches have been proposed for this task, they often struggle with capturing long-range information and memory bottlenecks, and usually encounter over-smoothing. To address these issues, this paper proposes an implicit model, named multi-channel Equilibrium Graph Neural Network (MEGNN). Through an equilibrium point iterative process, the proposed MEGNN naturally captures long-range information and effectively reduces the consumption of memory compared with explicit models. Furthermore, the proposed method deals with the issue of over-smoothing in deep graph convolutional networks by residual connection and shrinkage factor. We analyze the effect of the shrinkage factor on the information capturing capability of the model, and demonstrate that the proposed method does not encounter over-smoothing. Comprehensive experimental results demonstrate that the proposed method outperforms the state-of-the-art methods.
Shiping Wang, Yueyang Pi, Fuhai Chen, Le Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Low-Resolution Self-Attention for Semantic Segmentation
abstract
Semantic segmentation tasks naturally require high-resolution information for pixel-wise segmentation and global context information for class prediction. While existing vision transformers demonstrate promising performance, they often utilize high-resolution context modeling, resulting in a computational bottleneck. In this work, we challenge conventional wisdom and introduce the Low-Resolution Self-Attention (LRSA) mechanism to capture global context at a significantly reduced computational cost, i.e., FLOPs. Our approach involves computing self-attention in a fixed low-resolution space, regardless of the input image's resolution, with additional $\text{3}\times \text{3}$3×3 depth-wise convolutions to capture fine details in the high-resolution space. We demonstrate the effectiveness of our LRSA approach by building the LRFormer, a vision transformer with an encoder-decoder structure. Extensive experiments on the ADE20 K, COCO-Stuff, and CityScapes datasets demonstrate that LRFormer outperforms state-of-the-art models.
Yu-Huan Wu, Shi-Chen Zhang, Yun Liu 0011, Le Zhang 0001, Xin Zhan, Daquan Zhou, Jiashi Feng, Ming-Ming Cheng, Liangli Zhen
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Deep-Learning-Empowered Super Resolution: A Comprehensive Survey and Future Prospects
abstract
Super-resolution (SR) has garnered significant attention within the computer vision community, driven by advances in deep learning (DL) techniques and the growing demand for high-quality visual applications. With the expansion of this field, numerous surveys have emerged. Most existing surveys focus on specific domains, lacking a comprehensive overview of this field. Here, we present an in-depth review of diverse SR methods, encompassing single-image SR (SISR), video SR (VSR), stereo SR (SSR), and light field SR (LFSR). We extensively cover over 150 SISR methods, nearly 70 VSR approaches, and approximately 30 techniques for SSR and LFSR. We analyze methodologies, datasets, evaluation protocols, empirical results, and complexity. In addition, we conducted a taxonomy based on each backbone structure according to the diverse purposes. We also explore valuable yet understudied open issues in the field. We believe that this work will serve as a valuable resource and offer guidance to researchers in this domain. To facilitate access to related work, we created a dedicated repository available at https://github.com/AVC2-UESTC/Holistic-Super-Resolution-Review
Le Zhang 0001, Ao Li 0007, Qibin Hou, Ce Zhu, Yonina C. Eldar
Proc. IEEE1
2025 Prompt to Restore, Restore to Prompt: Cyclic Prompting for Universal Adverse Weather Removal
abstract
Universal adverse weather removal (UAWR) seeks to address various weather degradations within a unified framework. Recent methods are inspired by prompt learning using pre-trained vision-language models (e.g., CLIP), leveraging degradation-aware prompts to facilitate weather-free image restoration, yielding significant improvements. In this work, we propose CyclicPrompt, an innovative cyclic prompt approach designed to enhance the effectiveness, adaptability, and generalizability of UAWR. CyclicPrompt comprises two key components: 1) a composite context prompt that integrates weather-related information and context-aware representations into the network to guide restoration. This prompt differs from previous methods by marrying learnable input-conditional vectors with weather-specific knowledge, thereby improving adaptability across various degradations and 2) the erase-and-paste mechanism, after the initial guided restoration, substitutes weather-specific knowledge with constrained restoration priors, inducing high-quality weather-free concepts into the composite prompt to further fine-tune the restoration process. Therefore, we can form a cyclic "Prompt-Restore-Prompt" pipeline that adeptly harnesses weather-specific knowledge, textual contexts, and reliable textures. Extensive experiments on synthetic and real-world datasets validate the superior performance of CyclicPrompt. The code is available at: https://github.com/RongxinL/CyclicPrompt.
Rongxin Liao, Feng Li 0037, Yanyan Wei, Zenglin Shi, Le Zhang 0001, Huihui Bai 0001, Meng Wang 0001
IEEE Trans. Image Process.5
2025 Towards Open-Vocabulary Video Semantic Segmentation
abstract
Semantic segmentation in videos has been a focal point of recent research. However, existing models encounter challenges when faced with unfamiliar categories. To address this, we introduce the Open Vocabulary Video Semantic Segmentation (OV-VSS) task, designed to accurately segment every pixel across a wide range of open-vocabulary categories, including those that are novel or previously unexplored. To enhance OV-VSS performance, we propose a robust baseline, OV2VSS, which integrates a spatial-temporal fusion module, allowing the model to utilize temporal relationships across consecutive frames. Additionally, we incorporate a random frame enhancement module, broadening the model's understanding of semantic context throughout the entire video sequence. Our approach also includes video text encoding, which strengthens the model's capability to interpret textual information within the video context. Comprehensive evaluations on benchmark datasets such as VSPW and Cityscapes highlight OV-VSS's zero-shot generalization capabilities, especially in handling novel categories. The results validate OV2VSS's effectiveness, demonstrating improved performance in semantic segmentation tasks across diverse video datasets.
Yun Liu 0011, Guolei Sun, Min Wu 0008, Le Zhang 0001, Ce Zhu
IEEE Trans. Multim.5
2025 Guest Editorial: Special Issue on Trustworthy Federated Learning
Qiang Yang 0001, Han Yu 0001, Sin G. Teo, Bo Li 0001, Guodong Long, Lixin Fan, Yang Liu 0165, Le Zhang 0001
IEEE Trans. Neural Networks Learn. Syst.9
2024 Learning Triangular Distribution in Visual World
abstract
Convolution neural network is successful in pervasive vision tasks, including label distribution learning, which usually takes the form of learning an injection from the nonlinear visual features to the well-defined labels. However, how the discrepancy between features is mapped to the label discrepancy is ambient, and its correctness is not guaranteed. To address these problems, we study the mathematical connection between feature and its label, presenting a general and simple framework for label distribution learning. We propose a so-called Triangular Distribution Transform (TDT) to build an injective function between feature and label, guaranteeing that any symmetric feature discrepancy linearly reflects the difference between labels. The proposed TDT can be used as a plug-in in mainstream backbone networks to address different label distribution learning tasks. Experiments on Facial Age Recognition, Illumination Chromaticity Estimation, and Aesthetics assessment show that TDT achieves on-par or better results than the prior arts. Code is available at https://github.com/redcping/TDT.
Xingpeng Zhang, Chengtao Zhou, Dichao Fan, Peng Tu, Le Zhang 0001, Yanlin Qian
CVPR6
2024 CorrMatch: Label Propagation via Correlation Matching for Semi-Supervised Semantic Segmentation
abstract
This paper presents a simple but performant semi-supervised semantic segmentation approach, called CorrMatch. Previous approaches mostly employ complicated training strategies to leverage unlabeled data but overlook the role of correlation maps in modeling the relationships between pairs of locations. We observe that the correlation maps not only enable clustering pixels of the same category easily but also contain good shape information, which previous works have omitted. Motivated by these, we aim to improve the use efficiency of unlabeled data by designing two novel label propagation strategies. First, we propose to conduct pixel propagation by modeling the pairwise similarities of pixels to spread the high-confidence pixels and dig out more. Then, we perform region propagation to enhance the pseudo labels with accurate class-agnostic masks extracted from the correlation maps. CorrMatch achieves great performance on popular segmentation benchmarks. Taking the DeepLabV3+ with ResNet-101 backbone as our segmentation model, we receive a 76%+ mIoU score on the Pascal VOC 2012 dataset with only 92 annotated images. Code is available at https://github.com/BBBBchan/CorrMatch.
Boyuan Sun 0001, Le Zhang 0001, Ming-Ming Cheng, Qibin Hou
CVPR3
2024 Deep negative correlation classification
Le Zhang 0001, Qibin Hou, Yun Liu 0011, Jiawang Bian, Xun Xu 0002, Joey Tianyi Zhou, Ce Zhu
Mach. Learn.1
2024 GSB: Group superposition binarization for vision transformer with limited training samples
Tian Gao 0004, Cheng-Zhong Xu 0001, Le Zhang 0001, Hui Kong 0001
Neural Networks3
2024 Multi-view heterogeneous graph learning with compressed hypergraph neural networks
Aiping Huang, Zihan Fang 0002, Zhihao Wu 0003, Yanchao Tan, Peng Han 0005, Shiping Wang, Le Zhang 0001
Neural Networks7
2024 Multi-View Graph Embedding Learning for Image Co-Segmentation and Co-Localization
abstract
Image co-segmentation and co-localization exploit inter-image information to identify and extract foreground objects with a batch mode. However, they remain challenging when confronted with large object variations or complex backgrounds. This paper proposes a multi-view graph embedding (MV-Gem) learning scheme which integrates diversity, robustness and discernibility of object features to alleviate this phenomenon. To encourage the diversity, the deep co-information containing both low-layer general representations and high-layer semantic information is generated to form a multi-view feature pool for comprehensive co-object description. To enhance the robustness, a multi-view adaptive weighted learning is formulated to fuse the deep co-information for feature complementation. To ensure the discernibility, the graph embedding and sparse constraint are embedded into the fusion formulation for feature selection. The former aims to inherit important structures from multiple views, and the latter further selects important features to restrain irrelevant backgrounds. With these techniques, MV-Gem gradually recovers all co-objects through optimization iterations. Extensive experimental results on real-world datasets demonstrate that MV-Gem is capable of locating and delineating co-objects in an image group.
Aiping Huang, Lijian Li 0004, Le Zhang 0001, Yuzhen Niu, Tiesong Zhao, Chia-Wen Lin
IEEE Trans. Circuits Syst. Video Technol.3
2024 Boosting Salient Object Detection With Transformer-Based Asymmetric Bilateral U-Net
abstract
Existing salient object detection (SOD) methods mainly rely on U-shaped convolution neural networks (CNNs) with skip connections to combine the global contexts and local spatial details that are crucial for locating salient objects and refining object details, respectively. Despite great successes, the ability of CNNs in learning global contexts is limited. Recently, the vision transformer has achieved revolutionary progress in computer vision owing to its powerful modeling of global dependencies. However, directly applying the transformer to SOD is suboptimal because the transformer lacks the ability to learn local spatial representations. To this end, this paper explores the combination of transformers and CNNs to learn both global and local representations for SOD. We propose a transformer-based Asymmetric Bilateral U-Net (ABiU-Net). The asymmetric bilateral encoder has a transformer path and a lightweight CNN path, where the two paths communicate at each encoder stage to learn complementary global contexts and local spatial details, respectively. The asymmetric bilateral decoder also consists of two paths to process features from the transformer and CNN encoder paths, with communication at each decoder stage for decoding coarse salient object locations and fine-grained object details, respectively. Such communication between the two encoder/decoder paths enables AbiU-Net to learn complementary global and local representations, taking advantage of the natural merits of transformers and CNNs, respectively. Hence, ABiU-Net provides a new perspective for transformer-based SOD. Extensive experiments demonstrate that ABiU-Net performs favorably against previous state-of-the-art SOD methods. The code is available athttps://github.com/yuqiuyuqiu/ABiU-Net.
Yun Liu 0011, Le Zhang 0001, Haotian Lu 0003, Jing Xu 0008
IEEE Trans. Circuits Syst. Video Technol.3
2024 Democratizing Federated WiFi-Based Human Activity Recognition Using Hypothesis Transfer
abstract
Human activity recognition (HAR) is a crucial task in IoT systems with applications ranging from surveillance and intruder detection to home automation and more. Recently, non-invasive HAR utilizing WiFi signals has gained considerable attention due to advancements in ubiquitous WiFi technologies. However, recent studies have revealed significant privacy risks associated with WiFi signals, raising concerns about bio-information leakage. To address these concerns, the decentralized paradigm, particularly federated learning (FL), has emerged as a promising approach for training HAR models while preserving data privacy. Nevertheless, FL models may struggle in end-user environments due to substantial domain discrepancies between the source training data and the target end-user environment. This discrepancy arises from the sensitivity of WiFi signals to environmental changes, resulting in notable domain shifts. As a consequence, FL-based HAR approaches often face challenges when deployed in real-world WiFi environments. Albeit there are pioneer attempts on federated domain adaptation, they typically require non-trivial communication and computation cost, which is prohibitively expensive especially considering edge-based hardware equipment of end-user environment. In this paper, we propose a model to democratize the WiFi-based HAR system by enhancing recognition accuracy in unannotated end-user environments while prioritizing data privacy. Our model leverages the hypothesis transfer and a lightweight hypothesis ensemble to mitigate negative transfer. We prove a tighter theoretical upper bound compared to existing multi-source federated domain adaptation models. Extensive experiments shows our model improves the average accuracy by approximately 10 absolute percentage points in both cross-person and cross-environment settings comparing several state-of-the-art baselines.
Bing Li 0002, Wei Cui 0002, Le Zhang 0001, Min Wu 0008, Joey Tianyi Zhou
IEEE Trans. Mob. Comput.3
2024 GrapHAR: A Lightweight Human Activity Recognition Model by Exploring the Sub-Carrier Correlations
abstract
Human activity recognition (HAR) is an important task due to its far-reaching applications, such as surveillance, healthcare systems, and human-computer interaction. Recently, Channel State Information (CSI)-based HAR has attracted increasing attention in the research community due to its ubiquitous availability, good user privacy, and fewer constraints on working conditions. Most of the existing methods for CSI-based HAR use various deep learning models, such as Convolutional Neural Networks (CNNs), Long Short-Term Memory (LSTM), and Transformers, to distinguish activities based on their temporal patterns. Despite their remarkable effectiveness, these methods solely focus on temporal patterns while ignoring the correlations among sub-carriers. This limitation prevents them from achieving further performance improvement. Moreover, recent works often involve advanced yet massive and inefficient neural architectures, like Transformers, to obtain satisfactory recognition accuracy. The performance gain is traded off with a steep increase in model complexity, which leads to low efficacy and high training/inference costs outsides the small time window. To address these issues, we propose a lightweight CSI-based HAR model. Our model makes the first effort to explore the graphical correlations of CSI sub-carriers, working in conjunction with a temporal causal convolution module. The high efficacy design enables our model to be highly effective without requiring excessive model complexity. Extensive experiments conducted on four real-world datasets demonstrate that our model outperforms state-of-the-art methods, including a strong Transformer-based baseline. It achieves an average improvement of 8 percentage points in recognition accuracy, with only 10% of the parameters compared to the Transformer-based method (4.95M vs. 49.24M). Additionally, our model is significantly faster, with empirical training and execution times at least 2.07 times faster than the baseline.
Wei Meng 0002, Zhicong Liu, Bing Li 0002, Wei Cui 0002, Joey Tianyi Zhou, Le Zhang 0001
IEEE Trans. Wirel. Commun.6
2023 Feature Modulation Transformer: Cross-Refinement of Global Representation via High-Frequency Prior for Image Super-Resolution
abstract
Transformer-based methods have exhibited remarkable potential in single image super-resolution (SISR) by effectively extracting long-range dependencies. However, most of the current research in this area has prioritized the design of transformer blocks to capture global information, while overlooking the importance of incorporating high-frequency priors, which we believe could be beneficial. In our study, we conducted a series of experiments and found that transformer structures are more adept at capturing low-frequency information, but have limited capacity in constructing high-frequency representations when compared to their convolutional counterparts. Our proposed solution, the cross-refinement adaptive feature modulation transformer (CRAFT), integrates the strengths of both convolutional and transformer structures. It comprises three key components: the high-frequency enhancement residual block (HFERB) for extracting high-frequency information, the shift rectangle window attention block (SRWAB) for capturing global information, and the hybrid fusion block (HFB) for refining the global representation. Our experiments on multiple datasets demonstrate that CRAFT outperforms state-of-the-art methods by up to 0.29dB while using fewer parameters. The source code will be made available at: https://github.com/AVC2-UESTC/CRAFT-SR.git.
Ao Li 0007, Le Zhang 0001, Yun Liu 0011, Ce Zhu
ICCV2
2023 Overlap Loss: Rethinking Weakly Supervised Instance Segmentation in Crowded Scenes
abstract
Weakly supervised instance segmentation (WSIS) has gained increasing popularity in recent years due to low labelling cost. However, its performance deteriorates dramatically in more challenging crowded scenario, which is caused by overlapping among similar objects. To ameliorate the negative effects of instance overlapping, we propose a new loss, i.e., OverlapLoss, which achieves instance disentanglement between masks according to the degree of overlapping among instances. Besides, a new dataset of CrowdHuman Instance Segmentation (CIS) is presented to bridge the gap in crowded scenes. Experiments on the CIS and COCO datasets validate that the proposed loss can improve the baseline in typical crowed scenes by at least 2% and in uncrowded scenes by more than 0.3% w.r.t. absolute AP. The code and dataset are available at: https://github.com/shanghangjiang/CIS.
Shanghang Jiang, Shichao Zhao, Le Zhang 0001
ICIP4
2023 Efficient Lightweight Attention Based Learned Image Compression
abstract
The CNN-based end-to-end learned image compression methods have already achieved a significant improvement in terms of coding efficiency. Moreover, with the capability of modeling long-range global correlation, the transformer-based image compression has further elevated the coding efficiency to outperform the latest Versatile Video Coding (VVC) standard. Nonetheless, the high computational burden of the self-attention mechanism in the transformer design present a significant obstacle for practical applications. To address this concern, we propose a relatively low-complexity end-to-end learned image compression approach by integrating an efficient lightweight attention module, which effectively mitigates the computational overhead associated with self-attention in the transformer design. The experimental results demonstrate that the proposed method achieves better RD performance than VVC. Furthermore, as compared to the prevailing state-of-the-art transformer-based approach, our method accelerates the coding speed by over 6 times while maintaining comparable coding efficiency.
Lei Luo 0003, Le Zhang 0001, Hongwei Guo 0001, Ce Zhu
VCIP3
2023 Distributed representation learning with skip-gram model for trained random forests
Chao Ma 0006, Le Zhang 0001, Zhiguang Cao, Yue Huang 0001, Xinghao Ding
Neurocomputing3
2023 Diverse and consistent multi-view networks for semi-supervised regression
Cuong Manh Nguyen, Arun Raja, Le Zhang 0001, Xun Xu 0002, Balagopal Unnikrishnan, Mohamed Ragab 0002, Kangkang Lu 0001, Chuan-Sheng Foo
Mach. Learn.3
2023 DifFormer: Multi-Resolutional Differencing Transformer With Dynamic Ranging for Time Series Analysis
abstract
Time series analysis is essential to many far-reaching applications of data science and statistics including economic and financial forecasting, surveillance, and automated business processing. Though being greatly successful of Transformer in computer vision and natural language processing, the potential of employing it as the general backbone in analyzing the ubiquitous times series data has not been fully released yet. Prior Transformer variants on time series highly rely on task-dependent designs and pre-assumed "pattern biases", revealing its insufficiency in representing nuanced seasonal, cyclic, and outlier patterns which are highly prevalent in time series. As a consequence, they can not generalize well to different time series analysis tasks. To tackle the challenges, we propose DifFormer, an effective and efficient Transformer architecture that can serve as a workhorse for a variety of time-series analysis tasks. DifFormer incorporates a novel multi-resolutional differencing mechanism, which is able to progressively and adaptively make nuanced yet meaningful changes prominent, meanwhile, the periodic or cyclic patterns can be dynamically captured with flexible lagging and dynamic ranging operations. Extensive experiments demonstrate DifFormer significantly outperforms state-of-the-art models on three essential time-series analysis tasks, including classification, regression, and forecasting. In addition to its superior performances, DifFormer also excels in efficiency - a linear time/memory complexity with empirically lower time consumption.
Bing Li 0002, Wei Cui 0002, Le Zhang 0001, Ce Zhu, Wei Wang 0011, Ivor W. Tsang, Joey Tianyi Zhou
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Privacy-Preserving Cross-Environment Human Activity Recognition
abstract
Recent studies have demonstrated the success of using the channel state information (CSI) from the WiFi signal to analyze human activities in a fixed and well-controlled environment. Those systems usually degrade when being deployed in new environments. A straightforward solution to solve this limitation is to collect and annotate data samples from different environments with advanced learning strategies. Although workable as reported, those methods are often privacy sensitive because the training algorithms need to access the data from different environments, which may be owned by different organizations. We present a practical method for the WiFi-based privacy-preserving cross-environment human activity recognition (HAR). It collects and shares information from different environments, while maintaining the privacy of individual person being involved. At the core of our approach is the utilization of the Johnson-Lindenstrauss transform, which is theoretically shown to be differentially private. Based on that, we further design an adversarial learning strategy to generate environment-invariant representations for HAR. We demonstrate the effectiveness of the proposed method with different data modalities from two real-life environments. More specifically, on the raw CSI dataset, it shows 2.18% and 1.24% improvements over challenging baselines for two environments, respectively. Moreover, with the discrete wavelet transform features, it further yields 5.71% and 1.55% improvements, respectively.
Le Zhang 0001, Wei Cui 0002, Bing Li 0002, Zhenghua Chen, Min Wu 0008, Sin G. Teo
IEEE Trans. Cybern.1
2023 Dual-Stream Contrastive Learning for Channel State Information Based Human Activity Recognition
abstract
WiFi-based human activity recognition (HAR) has been extensively studied due to its far-reaching applications in health domains, including elderly monitoring, exercise supervision and rehabilitation monitoring, etc. Although existing supervised deep learning techniques have achieved remarkable performances for these tasks, they are however data-hungry and hence are notoriously difficult due to the privacy and incomprehensibility of WiFi-based HAR data. Existing contrastive learning models, mainly designed for computer vision, cannot guarantee their performance on channel state information (CSI) data. To this end, we propose a new dual-stream contrastive learning model that can process and learn the raw WiFi CSI data in a self-supervised manner. More specifically, our proposed method, coined as DualConFi, takes raw WiFI CSI data as input and incorporates channel and temporal streams to learn highly-discriminative spatiotemporal features under a mutual information constraint using unlabeled data. We exhibit the effectiveness of our model on three publicly available CSI data sets in various experiment settings, including linear evaluation, semi-supervised, and transfer learning. We show that DualConFi is able to perform favourably against challenging baselines in each setting. Moreover, by studying the effects of different transform functions on CSI data, we finally verify the effectiveness of highly-discriminative features.
Jiangtao Wang 0001, Le Zhang 0001, Hongyuan Zhu 0002, Dingchang Zheng
IEEE J. Biomed. Health Informatics3
2023 An Improved Artificial Bee Colony Algorithm With Q-Learning for Solving Permutation Flow-Shop Scheduling Problems
abstract
A permutation flow-shop scheduling problem (PFSP) has been studied for a long time due to its significance in real-life applications. This work proposes an improved artificial bee colony (ABC) algorithm with$Q$-learning, named QABC, for solving it with minimizing the maximum completion time (makespan). First, the Nawaz–Enscore–Ham (NEH) heuristic is employed to initialize the population of ABC. Second, a set of problem-specific and knowledge-based neighborhood structures are designed in the employ bee phase.$Q$-learning is employed to favorably choose the premium neighborhood structures. Next, an all-round search strategy is proposed to further enhance the quality of individuals in the onlooker bee phase. Moreover, an insert-based method is applied to avoid local optima. Finally, QABC is used to solve 151 well-known benchmark instances. Its performance is verified by comparing it with the state-of-the-art algorithms. Experimental and statistical results demonstrate its superiority over its peers in solving the concerned problems.
Kai-Zhou Gao, Peiyong Duan, Junqing Li 0001, Le Zhang 0001
IEEE Trans. Syst. Man Cybern. Syst.5
2022 Mining Relations Among Cross-Frame Affinities for Video Semantic Segmentation
Guolei Sun, Yun Liu 0011, Hao Tang 0005, Ajad Chhatkuli, Le Zhang 0001, Luc Van Gool
ECCV (34)5
2022 FMNet: Frequency-Aware Modulation Network for SDR-to-HDR Translation
abstract
High-dynamic-range (HDR) media resources that preserve high contrast and more details in shadow and highlight areas in television are becoming increasingly popular for modern display technology compared to the widely available standard-dynamic-range (SDR) media resources. However, due to the exorbitant price of HDR cameras, researchers have attempted to develop the SDR-to-HDR techniques to convert the abundant SDR media resources to the HDR versions for cost-saving. Recent SDR-to-HDR methods mostly apply the image-adaptive modulation scheme to dynamically modulate the local contrast. However, these methods often fail to properly capture the low-frequency cues, resulting in artifacts in the low-frequency regions and low visual quality. Motivated by the Discrete Cosine Transform (DCT), in this paper, we propose a Frequency-aware Modulation Network (FMNet) to enhance the contrast in a frequency-adaptive way for SDR-to-HDR translation. Specifically, we design a frequency-aware modulation block that can dynamically modulate the features according to its frequency-domain responses. This allows us to reduce the structural distortions and artifacts in the translated low-frequency regions and reconstruct high-quality HDR content in the translated results. Experimental results on the HDRTV1K dataset show that our FMNet outperforms previous methods and the perceptual quality of the generated HDR images can be largely improved. Our code is available at https://github.com/MCG-NKU/FMNet.
Qibin Hou, Le Zhang 0001, Ming-Ming Cheng
ACM Multimedia3
2022 Semantic Edge Detection with Diverse Deep Supervision
Yun Liu 0011, Ming-Ming Cheng, Deng-Ping Fan, Le Zhang 0001, Jiawang Bian, Dacheng Tao
Int. J. Comput. Vis.4
2022 Locality-Aware Crowd Counting
abstract
Imbalanced data distribution in crowd counting datasets leads to severe under-estimation and over-estimation problems, which has been less investigated in existing works. In this paper, we tackle this challenging problem by proposing a simple but effective locality-based learning paradigm to produce generalizable features by alleviating sample bias. Our proposed method is locality-aware in two aspects. First, we introduce a locality-aware data partition (LADP) approach to group the training data into different bins via locality-sensitive hashing. As a result, a more balanced data batch is then constructed by LADP. To further reduce the training bias and enhance the collaboration with LADP, a new data augmentation method called locality-aware data augmentation (LADA) is proposed where the image patches are adaptively augmented based on the loss. The proposed method is independent of the backbone network architectures, and thus could be smoothly integrated with most existing deep crowd counting approaches in an end-to-end paradigm to boost their performance. We also demonstrate the versatility of the proposed method by applying it for adversarial defense. Extensive experiments verify the superiority of the proposed method over the state of the arts.
Joey Tianyi Zhou, Le Zhang 0001, Jiawei Du 0002, Xi Peng 0001, Zhiwen Fang, Hongyuan Zhu 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 PersEmoN: A Deep Network for Joint Analysis of Apparent Personality, Emotion and Their Relationship
abstract
Apparent personality and emotion analysis are both central to affective computing. Existing works solve them individually. In this paper we investigate if such high-level affect traits and their relationship can be jointly learned from face images in the wild. To this end, we introducePersEmoN, an end-to-end trainable and deep Siamese-like network. It consists of two convolutional network branches, one for emotion and the other for apparent personality. Both networks share their bottom feature extraction module and are optimized within a multi-task learning framework. Emotion and personality networks are dedicated to their own annotated dataset. Furthermore, an adversarial-like loss function is employed to promote representation coherence among heterogeneous dataset sources. Based on this, we also explore the emotion-to-apparent-personality relationship. Extensive experiments demonstrate the effectiveness ofPersEmoN.
Le Zhang 0001, Songyou Peng, Stefan Winkler 0001
IEEE Trans. Affect. Comput.1
2022 Learning Clustering for Motion Segmentation
abstract
Subspace clustering has been extensively studied from the hypothesis-and-test, algebraic, and spectral clustering-based perspectives. Most assume that only a single type/class of subspace is present. Generalizations to multiple types are non-trivial, plagued by challenges such as choice of types and numbers of models, sampling imbalance and parameter tuning. In many real world problems, data may not lie perfectly on a linear subspace and hand designed linear subspace models may not fit into these situations. In this work, we formulate the multi-type subspace clustering problem as one of learning non-linear subspace filters via deep multi-layer perceptrons (mlps). The response to the learnt subspace filters serve as the feature embedding that is clustering-friendly, i.e., points of the same clusters will be embedded closer together through the network. For inference, we apply K-means to the network output to cluster the data. Experiments are carried out on synthetic data and real world motion segmentation problems, producing state-of-the-art results.
Xun Xu 0002, Le Zhang 0001, Loong Fah Cheong, Zhuwen Li, Ce Zhu
IEEE Trans. Circuits Syst. Video Technol.2
2022 Exploring Structural Knowledge for Automated Visual Inspection of Moving Trains
abstract
Deep learning methods are becoming the de-facto standard for generic visual recognition in the literature. However, their adaptations to industrial scenarios, such as visual recognition for machines, product streamlines, etc., which consist of countless components, have not been investigated well yet. Compared with the generic object detection, there is some strong structural knowledge in these scenarios (e.g., fixed relative positions of components, component relationships, etc.). A case worth exploring could be automated visual inspection for trains, where there are various correlated components. However, the dominant object detection paradigm is limited by treating the visual features of each object region separately without considering common sense knowledge among objects. In this article, we propose a novel automated visual inspection framework for trains exploring structural knowledge for train component detection, which is called SKTCD. SKTCD is an end-to-end trainable framework, in which the visual features of train components and structural knowledge (including hierarchical scene contexts and spatial-aware component relationships) are jointly exploited for train component detection. We propose novel residual multiple gated recurrent units (Res-MGRUs) that can optimally fuse the visual features of train components and messages from the structural knowledge in a weighted-recurrent way. In order to verify the feasibility of SKTCD, a dataset that contains high-resolution images captured from moving trains has been collected, in which 18 590 critical train components are manually annotated. Extensive experiments on this dataset and on the PASCAL VOC dataset have demonstrated that SKTCD outperforms the existing challenging baselines significantly. The dataset as well as the source code can be downloaded online (https://github.com/smartprobe/SKCD).
Cen Chen 0002, Xiaofeng Zou, Zeng Zeng, Zhongyao Cheng, Le Zhang 0001, Steven C. H. Hoi
IEEE Trans. Cybern.5
2022 MA-GANet: A Multi-Attention Generative Adversarial Network for Defocus Blur Detection
abstract
Background clutters pose challenges to defocus blur detection. Existing approaches often produce artifact predictions in background areas with clutter and relatively low confident predictions in boundary areas. In this work, we tackle the above issues from two perspectives. Firstly, inspired by the recent success of self-attention mechanism, we introduce channel-wise and spatial-wise attention modules to attentively aggregate features at different channels and spatial locations to obtain more discriminative features. Secondly, we propose a generative adversarial training strategy to suppress spurious and low reliable predictions. This is achieved by utilizing a discriminator to identify predicted defocus map from ground-truth ones. As such, the defocus network (generator) needs to produce ‘realistic’ defocus map to minimize discriminator loss. We further demonstrate that the generative adversarial training allows exploiting additional unlabeled data to improve performance, a.k.a. semi-supervised learning, and we provide the first benchmark on semi-supervised defocus detection. Finally, we demonstrate that the existing evaluation metrics for defocus detection generally fail to quantify the robustness with respect to thresholding. For a fair and practical evaluation, we introduce an effective yet efficient$AUF_\beta $metric. Extensive experiments on three public datasets verify the superiority of the proposed methods compared against state-of-the-art approaches.
Xun Xu 0002, Le Zhang 0001, Chao Zhang 0072, Chuan-Sheng Foo, Ce Zhu
IEEE Trans. Image Process.3
2022 EDN: Salient Object Detection via Extremely-Downsampled Network
abstract
Recent progress on salient object detection (SOD) mainly benefits from multi-scale learning, where the high-level and low-level features collaborate in locating salient objects and discovering fine details, respectively. However, most efforts are devoted to low-level feature learning by fusing multi-scale features or enhancing boundary representations. High-level features, which although have long proven effective for many other tasks, yet have been barely studied for SOD. In this paper, we tap into this gap and show that enhancing high-level features is essential for SOD as well. To this end, we introduce an Extremely-Downsampled Network (EDN), which employs an extreme downsampling technique to effectively learn a global view of the whole image, leading to accurate salient object localization. To accomplish better multi-level feature fusion, we construct the Scale-Correlated Pyramid Convolution (SCPC) to build an elegant decoder for recovering object details from the above extreme downsampling. Extensive experiments demonstrate that EDN achieves state-of-the-art performance with real-time speed. Our efficient EDN-Lite also achieves competitive performance with a speed of 316fps. Hence, this work is expected to spark some new thinking in SOD. Code is available at https://github.com/yuhuan-wu/EDN.
Yu-Huan Wu, Yun Liu 0011, Le Zhang 0001, Ming-Ming Cheng, Bo Ren 0003
IEEE Trans. Image Process.3
2021 Two-Stream Convolution Augmented Transformer for Human Activity Recognition
abstract
Recognition of human activities is an important task due to its far-reaching applications such as healthcare system, context-aware applications, and security monitoring. Recently, WiFi based human activity recognition (HAR) is becoming ubiquitous due to its non-invasiveness. Existing WiFi-based HAR methods regard WiFi signals as a temporal sequence of channel state information (CSI), and employ deep sequential models (e.g., RNN, LSTM) to automatically capture channel-over-time features. Although being remarkably effective, they suffer from two major drawbacks. Firstly, the granularity of a single temporal point is blindly elementary for representing meaningful CSI patterns. Secondly, the time-over-channel features are also important, and could be a natural data augmentation. To address the drawbacks, we propose a novel Two-stream Convolution Augmented Human Activity Transformer (THAT) model. Our model proposes to utilize a two-stream structure to capture both time-over-channel and channel-over-time features, and use the multi-scale convolution augmented transformer to capture range-based patterns. Extensive experiments on four real experiment datasets demonstrate that our model outperforms state-of-the-art models in terms of both effectiveness and efficiency.
Bing Li 0002, Wei Cui 0002, Wei Wang 0011, Le Zhang 0001, Zhenghua Chen, Min Wu 0008
AAAI4
2021 On Automatic Data Augmentation for 3D Point Cloud Classification
Wanyue Zhang, Xun Xu 0002, Fayao Liu, Le Zhang 0001, Chuan-Sheng Foo
BMVC4
2021 A Multi-Stage Progressive Learning Strategy for Covid-19 Diagnosis Using Chest Computed Tomography with Imbalanced Data
abstract
In this paper, a multi-stage progressive learning strategy is investigated to train classifiers for COVID-19 Diagnosis using imbalanced Chest Computed Tomography Data acquired from patients infected with COVID-19 Pneumonia, Community Acquired Pneumonia (CAP) and from normal healthy subjects. In the first learning stage, pre-processed volumetric CT data together with the segmented lung masks are fed into a 3D ResNet module, and an initial classification result can be obtained. However, due to categorical data imbalance, we observe large differences in sensitivity between COVID-19 and CAP cases. In the second stage, five learning models are independently trained over data with only COVID-19 and CAP cases, and are then ensembled to further discriminate the two classes. The final classification results are obtained by combining the predictions from both stages. Based on the validation dataset, we have evaluated our method and compared it with up-to-date methods in terms of overall accuracy and sensitivity for each class. The validation results validate the accuracy of the proposed multi-stage learning strategy. The overall accuracy of the validation dataset is 88.8%, and the sensitivities are 0.873, 0.789 and 1 for COVID-19, CAP and normal cases, respectively.
Zaifeng Yang, Yubo Hou, Zhenghua Chen, Le Zhang 0001, Jie Chen 0026
ICASSP4
2021 Learning to Iteratively Solve Routing Problems with Dual-Aspect Collaborative Transformer
abstract
Recently, Transformer has become a prevailing deep architecture for solving vehicle routing problems (VRPs). However, it is less effective in learning improvement models for VRP because its positional encoding (PE) method is not suitable in representing VRP solutions. This paper presents a novel Dual-Aspect Collaborative Transformer (DACT) to learn embeddings for the node and positional features separately, instead of fusing them together as done in existing ones, so as to avoid potential noises and incompatible correlations. Moreover, the positional features are embedded through a novel cyclic positional encoding (CPE) method to allow Transformer to effectively capture the circularity and symmetry of VRP solutions (i.e., cyclic sequences). We train DACT using Proximal Policy Optimization and design a curriculum learning strategy for better sample efficiency. We apply DACT to solve the traveling salesman problem (TSP) and capacitated vehicle routing problem (CVRP). Results show that our DACT outperforms existing Transformer based improvement models, and exhibits much better generalization performance across different problem sizes on synthetic and benchmark instances, respectively.
Yining Ma 0001, Zhiguang Cao, Wen Song 0004, Le Zhang 0001, Zhenghua Chen, Jing Tang 0004
NeurIPS5
2021 Unsupervised Scale-Consistent Depth Learning from Video
Jiawang Bian, Huangying Zhan, Naiyan Wang, Le Zhang 0001, Chunhua Shen, Ming-Ming Cheng, Ian D. Reid 0001
Int. J. Comput. Vis.5
2021 Deep learning for human activity recognition
Xiaoli Li 0001, Peilin Zhao, Min Wu 0008, Zhenghua Chen, Le Zhang 0001
Neurocomputing5
2021 Bridge health anomaly detection using deep support vector data description
Jianxi Yang, Shixin Jiang, Guiping Wang, Le Zhang 0001, Zeng Zeng
Neurocomputing7
2021 Visual relationship detection with region topology structure
Le Zhang 0001, Ying Wang 0007, HaiShun Chen, Jie Li 0001
Inf. Sci.1
2021 Ordered or Orderless: A Revisit for Video Based Person Re-Identification
abstract
Is recurrent network really necessary for learning a good visual representation for video based person re-identification (VPRe-id)? In this paper, we first show that the common practice of employing recurrent neural networks (RNNs) to aggregate temporal-spatial features may not be optimal. Specifically, with a diagnostic analysis, we show that the recurrent structure may not be effective learn temporal dependencies than what we expected and implicitly yields an orderless representation. Based on this observation, we then present a simple yet surprisingly powerful approach for VPRe-id, where we treat VPRe-id as an efficient orderless ensemble of image based person re-identification problem. More specifically, we divide videos into individual images and re-identify person with ensemble of image based rankers. Under the i.i.d. assumption, we provide an error bound that sheds light upon how could we improve VPRe-id. Our work also presents a promising way to bridge the gap between video and image based person re-identification. Comprehensive experimental evaluations demonstrate that the proposed solution achieves state-of-the-art performances on multiple widely used datasets (iLIDS-VID, PRID 2011, and MARS).
Le Zhang 0001, Zenglin Shi, Joey Tianyi Zhou, Ming-Ming Cheng, Yun Liu 0011, Jiawang Bian, Zeng Zeng, Chunhua Shen
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 Nonlinear Regression via Deep Negative Correlation Learning
abstract
Nonlinear regression has been extensively employed in many computer vision problems (e.g., crowd counting, age estimation, affective computing). Under the umbrella of deep learning, two common solutions exist i) transforming nonlinear regression to a robust loss function which is jointly optimizable with the deep convolutional network, and ii) utilizing ensemble of deep networks. Although some improved performance is achieved, the former may be lacking due to the intrinsic limitation of choosing a single hypothesis and the latter may suffer from much larger computational complexity. To cope with those issues, we propose to regress via an efficient "divide and conquer" manner. The core of our approach is the generalization of negative correlation learning that has been shown, both theoretically and empirically, to work well for non-deep regression problems. Without extra parameters, the proposed method controls the bias-variance-covariance trade-off systematically and usually yields a deep regression ensemble where each base model is both "accurate" and "diversified." Moreover, we show that each sub-problem in the proposed method has less Rademacher Complexity and thus is easier to optimize. Extensive experiments on several diverse and challenging tasks including crowd counting, personality analysis, age estimation, and image super-resolution demonstrate the superiority over challenging baselines as well as the versatility of the proposed method. The source code and trained models are available on our project page: https://mmcheng.net/dncl/.
Le Zhang 0001, Zenglin Shi, Ming-Ming Cheng, Yun Liu 0011, Jiawang Bian, Joey Tianyi Zhou, Guoyan Zheng, Zeng Zeng
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 Correction to "Nonlinear Regression via Deep Negative Correlation Learning"
abstract
Reports on changes to the author information presented in the above named paper.
Le Zhang 0001, Zenglin Shi, Ming-Ming Cheng, Yun Liu 0011, Jiawang Bian, Joey Tianyi Zhou, Guoyan Zheng, Zeng Zeng
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 SAMNet: Stereoscopically Attentive Multi-Scale Network for Lightweight Salient Object Detection
abstract
Recent progress on salient object detection (SOD) mostly benefits from the explosive development of Convolutional Neural Networks (CNNs). However, much of the improvement comes with the larger network size and heavier computation overhead, which, in our view, is not mobile-friendly and thus difficult to deploy in practice. To promote more practical SOD systems, we introduce a novel Stereoscopically Attentive Multi-scale (SAM) module, which adopts a stereoscopic attention mechanism to adaptively fuse the features of various scales. Embarking on this module, we propose an extremely lightweight network, namely SAMNet, for SOD. Extensive experiments on popular benchmarks demonstrate that the proposed SAMNet yields comparable accuracy with state-of-the-art methods while running at a GPU speed of 343fps and a CPU speed of 5fps for 336 ×336 inputs with only 1.33M parameters. Therefore, SAMNet paves a new path towards SOD. The source code is available on the project page https://mmcheng.net/SAMNet/.
Yun Liu 0011, Xinyu Zhang 0023, Jiawang Bian, Le Zhang 0001, Ming-Ming Cheng
IEEE Trans. Image Process.4
2021 Regularized Densely-Connected Pyramid Network for Salient Instance Segmentation
abstract
Much of the recent efforts on salient object detection (SOD) have been devoted to producing accurate saliency maps without being aware of their instance labels. To this end, we propose a new pipeline for end-to-end salient instance segmentation (SIS) that predicts a class-agnostic mask for each detected salient instance. To better use the rich feature hierarchies in deep networks and enhance the side predictions, we propose the regularized dense connections, which attentively promote informative features and suppress non-informative ones from all feature pyramids. A novel multi-level RoIAlign based decoder is introduced to adaptively aggregate multi-level features for better mask predictions. Such strategies can be well-encapsulated into the Mask R-CNN pipeline. Extensive experiments on popular benchmarks demonstrate that our design significantly outperforms existing state-of-the-art competitors by 6.3% (58.6% vs. 52.3%) in terms of the AP metric. The code is available at https://github.com/yuhuan-wu/RDPNet.
Yu-Huan Wu, Yun Liu 0011, Le Zhang 0001, Wang Gao 0001, Ming-Ming Cheng
IEEE Trans. Image Process.3
2020 GMS: Grid-Based Motion Statistics for Fast, Ultra-robust Feature Correspondence
abstract
Abstract Feature matching aims at generating correspondences across images, which is widely used in many computer vision tasks. Although considerable progress has been made on feature descriptors and fast matching for initial correspondence hypotheses, selecting good ones from them is still challenging and critical to the overall performance. More importantly, existing methods often take a long computational time, limiting their use in real-time applications. This paper attempts to separate true correspondences from false ones at high speed. We term the proposed method (GMS) grid-based motion Statistics, which incorporates the smoothness constraint into a statistic framework for separation and uses a grid-based implementation for fast calculation. GMS is robust to various challenging image changes, involving in viewpoint, scale, and rotation. It is also fast, e.g., take only 1 or 2 ms in a single CPU thread, even when 50K correspondences are processed. This has important implications for real-time applications. What’s more, we show that incorporating GMS into the classic feature matching and epipolar geometry estimation pipeline can significantly boost the overall performance. Finally, we integrate GMS into the well-known ORB-SLAM system for monocular initialization, resulting in a significant improvement.
Jiawang Bian, Wen-Yan Lin, Yun Liu 0011, Le Zhang 0001, Sai-Kit Yeung, Ming-Ming Cheng, Ian D. Reid 0001
Int. J. Comput. Vis.4
2020 WiFi-Based Indoor Robot Positioning Using Deep Fuzzy Forests
abstract
Addressing the positioning problem of a mobile robot remains challenging to date despite many years of research. Indoor robot positioning strategies developed in the literature either rely on sophisticated computer vision techniques to handle visual inputs or require strong domain knowledge for nonvisual sensors. Although some systems have been deployed, the former may be lacking due to the intrinsic limitation of cameras (such as calibration, data association, system initialization, etc.) and the latter usually only works under certain environment layouts and additional equipment. To cope with those issues, we design a lightweight indoor robot positioning system which operates on cost-effective WiFi-based received signal strength (RSS) and could be readily pluggable into any existing WiFi network infrastructures. Moreover, a novel deep fuzzy forest is proposed to inherit the merits of decision trees and deep neural networks within an end-to-end trainable architecture. Real-world indoor localization experiments are conducted and results demonstrate the superiority of the proposed method over the existing approaches.
Le Zhang 0001, Zhenghua Chen, Wei Cui 0002, Bing Li 0002, Cen Chen 0002, Zhiguang Cao, Kai-Zhou Gao
IEEE Internet Things J.1
2020 Heterogeneous oblique random forest
abstract
Decision trees in random forests use a single feature in non-leaf nodes to split the data. Such splitting results in axis-parallel decision boundaries which may fail to exploit the geometric structure in the data. In oblique decision trees, an oblique hyperplane is employed instead of an axis-parallel hyperplane . Trees with such hyperplanes can better exploit the geometric structure to increase the accuracy of the trees and reduce the depth. The present realizations of oblique decision trees do not evaluate many promising oblique splits to select the best. In this paper, we propose a random forest of heterogeneous oblique decision trees that employ several linear classifiers at each non-leaf node on some top ranked partitions which are obtained via one-vs-all and two-hyperclasses based approaches and ranked based on ideal Gini scores and cluster separability. The oblique hyperplane that optimizes the impurity criterion is then selected as the splitting hyperplane for that node. We benchmark 190 classifiers on 121 UCI datasets. The results show that the oblique random forests proposed in this paper are the top 3 ranked classifiers with the heterogeneous oblique random forest being statistically better than all 189 classifiers in the literature.
Rakesh Katuwal, Ponnuthurai N. Suganthan, Le Zhang 0001
Pattern Recognit.3
2020 Attention-Driven Loss for Anomaly Detection in Video Surveillance
abstract
Recent video anomaly detection methods focus on reconstructing or predicting frames. Under this umbrella, the long-standing inter-class data-imbalance problem resorts to the imbalance between foreground and stationary background objects in video anomaly detection and this has been less investigated by existing solutions. Naively optimizing the reconstructing loss yields a biased optimization towards background reconstruction rather than the objects of interest in the foreground. To solve this, we proposed a simple yet effective solution, termed attention-driven loss to alleviate the foreground-background imbalance problem in anomaly detection. Specifically, we compute a single mask map that summarizes the frame evolution of moving foreground regions and suppresses the background in the training video clips. After that, we construct an attention map through the combination of the mask map and background to give different weights to the foreground and background region respectively. The proposed attention-driven loss is independent of backbone networks and can be easily augmented in most existing anomaly detection models. Augmented with attention-driven loss, the model is able to achieve AUC 86.0% on Avenue, 83.9% on Ped1, 96% on Ped2 datasets. Extensive experimental results and ablation studies further validate the effectiveness of our model.
Joey Tianyi Zhou, Le Zhang 0001, Zhiwen Fang, Jiawei Du 0002, Xi Peng 0001, Yang Xiao 0007
IEEE Trans. Circuits Syst. Video Technol.2
2019 An Evaluation of Feature Matchers for Fundamental Matrix Estimation
Jiawang Bian, Yu-Huan Wu, Ji Zhao 0001, Yun Liu 0011, Le Zhang 0001, Ming-Ming Cheng, Ian D. Reid 0001
BMVC5
2019 Contrast Prior and Fluid Pyramid Integration for RGBD Salient Object Detection
abstract
The large availability of depth sensors provides valuable complementary information for salient object detection (SOD) in RGBD images. However, due to the inherent difference between RGB and depth information, extracting features from the depth channel using ImageNet pre-trained backbone models and fusing them with RGB features directly are sub-optimal. In this paper, we utilize contrast prior, which used to be a dominant cue in none deep learning based SOD approaches, into CNNs-based architecture to enhance the depth information. The enhanced depth cues are further integrated with RGB features for SOD, using a novel fluid pyramid integration, which can make better use of multi-scale cross-modal features. Comprehensive experiments on 5 challenging benchmark datasets demonstrate the superiority of the architecture CPFP over 9 state-of-the-art alternative methods.
Jiaxing Zhao, Yang Cao 0017, Deng-Ping Fan, Ming-Ming Cheng, Xuan-Yi Li, Le Zhang 0001
CVPR6
2019 Using FTOC to track shuttlecock for the badminton robot
Tingbo Liao, Zhihang Li, Haozhi Lin, Le Zhang 0001, Jing Guo 0007, Zhiguang Cao
Neurocomputing6
2019 Richer Convolutional Features for Edge Detection
abstract
Edge detection is a fundamental problem in computer vision. Recently, convolutional neural networks (CNNs) have pushed forward this field significantly. Existing methods which adopt specific layers of deep CNNs may fail to capture complex data structures caused by variations of scales and aspect ratios. In this paper, we propose an accurate edge detector using richer convolutional features (RCF). RCF encapsulates all convolutional features into more discriminative representation, which makes good usage of rich feature hierarchies, and is amenable to training via backpropagation. RCF fully exploits multiscale and multilevel information of objects to perform the image-to-image prediction holistically. Using VGG16 network, we achieve state-of-the-art performance on several available datasets. When evaluating on the well-known BSDS500 benchmark, we achieve ODS F-measure of 0.811 while retaining a fast speed (8 FPS). Besides, our fast version of RCF achieves ODS F-measure of 0.806 with 30 FPS. We also demonstrate the versatility of the proposed method by applying RCF edges for classical image segmentation.
Yun Liu 0011, Ming-Ming Cheng, Xiaowei Hu 0003, Jiawang Bian, Le Zhang 0001, Xiang Bai, Jinhui Tang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2019 SeSe-Net: Self-Supervised deep learning for segmentation
Zeng Zeng, Xulei Yang, Yu Qiyun, Le Zhang 0001
Pattern Recognit. Lett.5
2019 WiFi CSI Based Passive Human Activity Recognition Using Attention Based BLSTM
abstract
Human activity recognition can benefit various applications including healthcare services and context awareness. Since human actions will influence WiFi signals, which can be captured by the channel state information (CSI) of WiFi, WiFi CSI based human activity recognition has gained more and more attention. Due to the complex relationship between human activities and WiFi CSI measurements, the accuracies of current recognition systems are far from satisfactory. In this paper, we propose a new deep learning based approach, i.e., attention based bi-directional long short-term memory (ABLSTM), for passive human activity recognition using WiFi CSI signals. The BLSTM is employed to learn representative features in two directions from raw sequential CSI measurements. Since the learned features may have different contributions for final activity recognition, we leverage on an attention mechanism to assign different weights for all the learned features. Real experiments have been carried out to evaluate the performance of the proposed ABLSTM for human activity recognition. The experimental results show that our proposed ABLSTM is able to achieve the best recognition performance for all activities when compared with some benchmark approaches.
Zhenghua Chen, Le Zhang 0001, Chaoyang Jiang, Zhiguang Cao, Wei Cui 0002
IEEE Trans. Mob. Comput.2
2019 Understanding the Dynamics of Social Interactions: A Multi-Modal Multi-View Approach
abstract
In this article, we deal with the problem of understanding human-to-human interactions as a fundamental component of social events analysis. Inspired by the recent success of multi-modal visual data in many recognition tasks, we propose a novel approach to model dyadic interaction by means of features extracted from synchronized 3D skeleton coordinates, depth, and Red Green Blue (RGB) sequences. From skeleton data, we extract new view-invariant proxemic features, named Unified Proxemic Descriptor (UProD), which is able to incorporate intrinsic and extrinsic distances between two interacting subjects. A novel key frame selection method is introduced to identify salient instants of the interaction sequence based on the joints’ energy. From Red Green Blue Depth (RGBD) videos, more holistic CNN features are extracted by applying an adaptive pre-trained Convolutional Neural Networks (CNNs) on optical flow frames. For better understanding the dynamics of interactions, we expand the boundaries of dyadic interactions analysis by proposing a fundamentally new modeling for non-treated problem aiming to discern the active from the passive interactor. Extensive experiments have been carried out on four multi-modal and multi-view interactions datasets. The experimental results demonstrate the superiority of our proposed techniques against the state-of-the-art approaches.
Rim Trabelsi, Jagannadan Varadarajan, Le Zhang 0001, Issam Jabri, Yong Pei, Fethi Smach, Ammar Bouallègue, Pierre Moulin
ACM Trans. Multim. Comput. Commun. Appl.3
2018 Kernel Cross-Correlator
abstract
Cross-correlator plays a significant role in many visual perception tasks, such as object detection and tracking. Beyond the linear cross-correlator, this paper proposes a kernel cross-correlator (KCC) that breaks traditional limitations. First, by introducing the kernel trick, the KCC extends the linear cross-correlation to non-linear space, which is more robust to signal noises and distortions. Second, the connection to the existing works shows that KCC provides a unified solution for correlation filters. Third, KCC is applicable to any kernel function and is not limited to circulant structure on training data, thus it is able to predict affine transformations with customized properties. Last, by leveraging the fast Fourier transform (FFT), KCC eliminates direct calculation of kernel vectors, thus achieves better performance yet still with a reasonable computational cost. Comprehensive experiments on visual tracking and human activity recognition using wearable devices demonstrate its robustness, flexibility, and efficiency. The source codes of both experiments are released at https://github.com/wang-chen/KCC.
Chen Wang 0033, Le Zhang 0001, Lihua Xie 0001, Junsong Yuan 0001
AAAI2
2018 Crowd Counting With Deep Negative Correlation Learning
abstract
Deep convolutional networks (ConvNets) have achieved unprecedented performances on many computer vision tasks. However, their adaptations to crowd counting on single images are still in their infancy and suffer from severe over-fitting. Here we propose a new learning strategy to produce generalizable features by way of deep negative correlation learning (NCL). More specifically, we deeply learn a pool of decorrelated regressors with sound generalization capabilities through managing their intrinsic diversities. Our proposed method, named decorrelated ConvNet (D-ConvNet), is end-to-end-trainable and independent of the backbone fully-convolutional network architectures. Extensive experiments on very deep VGGNet as well as our customized network structure indicate the superiority of D-ConvNet when compared with several state-of-the-art methods. Our implementation will be released at https://github.com/shizenglin/Deep-NCL.
Zenglin Shi, Le Zhang 0001, Yun Liu 0011, Yangdong Ye, Ming-Ming Cheng, Guoyan Zheng
CVPR2
2018 DEL: Deep Embedding Learning for Efficient Image Segmentation
abstract
Image segmentation has been explored for many years and still remains a crucial vision problem. Some efficient or accurate segmentation algorithms have been widely used in many vision applications. However, it is difficult to design a both efficient and accurate image segmenter. In this paper, we propose a novel method called DEL (deep embedding learning) which can efficiently transform superpixels into image segmentation. Starting with the SLIC superpixels, we train a fully convolutional network to learn the feature embedding space for each superpixel. The learned feature embedding corresponds to a similarity measure that measures the similarity between two adjacent superpixels. With the deep similarities, we can directly merge the superpixels into large segments. The evaluation results on BSDS500 and PASCAL Context demonstrate that our approach achieves a good trade-off between efficiency and effectiveness. Specifically, our DEL algorithm can achieve comparable segments when compared with MCG but is much faster than it, i.e. 11.4fps vs. 0.07fps.
Yun Liu 0011, Peng-Tao Jiang, Vahan Petrosyan, Shijie Li 0006, Jiawang Bian, Le Zhang 0001, Ming-Ming Cheng
IJCAI6
2018 Bayesian VoxDRN: A Probabilistic Deep Voxelwise Dilated Residual Network for Whole Heart Segmentation from 3D MR Images
Zenglin Shi, Guodong Zeng, Le Zhang 0001, Xiahai Zhuang, Lei Li 0020, Guang Yang 0006, Guoyan Zheng
MICCAI (4)3
2018 Give Me One Portrait Image, I Will Tell You Your Emotion and Personality
abstract
Personality and emotion are both central to affective computing. Existing works address them individually. In this demo we investigate if such high-level affect traits and their relationship can be jointly learned from face images in the wild. To this end, we introduce an end-to-end trainable and deep Siamese-like network. At inference time, our system can take one portrait photo as input and predict one's Big-Five apparent personality as well as emotion attributes. With such a system, we also demonstrate the feasibility of inferring the apparent personality directly fro emotion.
Songyou Peng, Le Zhang 0001, Stefan Winkler 0001, Marianne Winslett
ACM Multimedia2
2018 Historical Context-based Style Classification of Painting Images via Label Distribution Learning
abstract
Analyzing and categorizing the style of visual art images, especially paintings, is gaining popularity owing to its importance in understanding and appreciating the art. The evolution of painting style is both continuous, in a sense that new styles may inherit, develop or even mutate from their predecessors and multi-modal because of various issues such as the visual appearance, the birthplace, the origin time and the art movement. Motivated by this peculiarity, we introduce a novel knowledge distilling strategy to assist visual feature learning in the convolutional neural network for painting style classification. More specifically, a multi-factor distribution is employed as soft-labels to distill complementary information with visual input, which extracts from different historical context via label distribution learning. The proposed method is well-encapsulated in a multi-task learning framework which allows end-to-end training. We demonstrate the superiority of the proposed method over the state-of-the-art approaches on Painting91, OilPainting, and Pandora datasets.
Jufeng Yang, Liyi Chen 0003, Le Zhang 0001, Xiaoxiao Sun 0002, Dongyu She, Shao-Ping Lu, Ming-Ming Cheng
ACM Multimedia3
2018 Distilling the Knowledge From Handcrafted Features for Human Activity Recognition
abstract
Human activity recognition is a core problem in intelligent automation systems due to its far-reaching applications including ubiquitous computing, health-care services, and smart living. Due to the nonintrusive property of smartphones, smartphone sensors are widely used for the identification of human activities. However, unlike applications in vision or data mining domain, feature embedding from deep neural networks performs much worse in terms of recognition accuracy than properly designed handcrafted features. In this paper, we posit that feature embedding from deep neural networks may convey complementary information and propose a novel knowledge distilling strategy to improve its performance. More specifically, an efficient shallow network, i.e., single-layer feedforward neural network (SLFN), with handcrafted features is utilized to assist a deep long short-term memory (LSTM) network. On the one hand, the deep LSTM network is able to learn features from raw sensory data to encode temporal dependencies. On the other hand, the deep LSTM network can also learn from SLFN to mimic how it generalizes. Experimental results demonstrate the superiority of the proposed method in terms of recognition accuracy against several state-of-the-art methods in the literature.
Zhenghua Chen, Le Zhang 0001, Zhiguang Cao, Jing Guo 0007
IEEE Trans. Ind. Informatics2
2018 Received Signal Strength Based Indoor Positioning Using a Random Vector Functional Link Network
abstract
Fingerprinting based indoor positioning system is gaining more research interest under the umbrella of location-based services. However, existing works have certain limitations in addressing issues such as noisy measurements, high computational complexity, and poor generalization ability. In this work, a random vector functional link network based approach is introduced to address these issues. In the proposed system, a subset of informative features from many randomized noisy features is selected to both reduce the computational complexity and boost the generalization ability. Moreover, the feature selector and predictor are jointly learned iteratively in a single framework based on an augmented Lagrangian method. The proposed system is appealing as it can be naturally fit into parallel or distributed computing environment. Extensive real-world indoor localization experiments are conducted on users with smartphone devices and results demonstrate the superiority of the proposed method over the existing approaches.
Wei Cui 0002, Le Zhang 0001, Bing Li 0002, Jing Guo 0007, Wei Meng 0002, Haixia Wang 0003, Lihua Xie 0001
IEEE Trans. Ind. Informatics2
2018 Multiscale Multitask Deep NetVLAD for Crowd Counting
abstract
Deep convolutional networks (CNNs) reign undisputed as the new de-facto method for computer vision tasks owning to their success in visual recognition task on still images. However, their adaptations to crowd counting have not clearly established their superiority over shallow models. Existing CNNs turn out to be self-limiting in challenging scenarios such as camera illumination changing, partial occlusions, diverse crowd distributions, and perspective distortions for crowd counting because of their shallow structure. In this paper, we introduce a dynamic augmentation technique to train a much deeper CNN for crowd counting. In order to decrease overfitting caused by limited number of training samples, multitask learning is further employed to learn generalizable representations across similar domains. We also propose to aggregate multiscale convolutional features extracted from the entire image into a compact single vector representation amenable to efficient and accurate counting by way of “Vector of Locally Aggregated Descriptors” (VLAD). The “deeply supervised” strategy is employed to provide additional supervision signal for bottom layers for further performance improvement. Experimental results on three benchmark crowd datasets show that our method achieves better performance than the existing methods. Our implementation will be released at https://github.com/shizenglin/Multitask-Multiscale-Deep-NetVLAD.
Zenglin Shi, Le Zhang 0001, Yangdong Ye
IEEE Trans. Ind. Informatics2
2017 Robust Visual Tracking Using Oblique Random Forests
abstract
Random forest has emerged as a powerful classification technique with promising results in various vision tasks including image classification, pose estimation and object detection. However, current techniques have shown little improvements in visual tracking as they mostly rely on piece wise orthogonal hyperplanes to create decision nodes and lack a robust incremental learning mechanism that is much needed for online tracking. In this paper, we propose a discriminative tracker based on a novel incremental oblique random forest. Unlike conventional orthogonal decision trees that use a single feature and heuristic measures to obtain a split at each node, we propose to use a more powerful proximal SVM to obtain oblique hyperplanes to capture the geometric structure of the data better. The resulting decision surface is not restricted to be axis aligned and hence has the ability to represent and classify the input data better. Furthermore, in order to generalize to online tracking scenarios, we derive incremental update steps that enable the hyperplanes in each node to be updated recursively, efficiently and in a closed-form fashion. We demonstrate the effectiveness of our method using two large scale benchmark datasets (OTB-51 and OTB-100) and show that our method gives competitive results on several challenging cases by relying on simple HOG features as well as in combination with more sophisticated deep neural network based models. The implementations of the proposed random forest are available at https://github.com/ZhangLeUestc/ Incremental-Oblique-Random-Forest.
Le Zhang 0001, Jagannadan Varadarajan, Ponnuthurai N. Suganthan, Narendra Ahuja, Pierre Moulin
CVPR1
2017 Oblique random forest ensemble via Least Square Estimation for time series forecasting
Xueheng Qiu, Le Zhang 0001, Ponnuthurai N. Suganthan, Gehan A. J. Amaratunga
Inf. Sci.2
2017 Robust visual tracking via co-trained Kernelized correlation filters
Le Zhang 0001, Ponnuthurai N. Suganthan
Pattern Recognit.1
2017 Visual Tracking With Convolutional Random Vector Functional Link Network
abstract
Deep neural network-based methods have recently achieved excellent performance in visual tracking task. As very few training samples are available in visual tracking task, those approaches rely heavily on extremely large auxiliary dataset such as ImageNet to pretrain the model. In order to address the discrepancy between the source domain (the auxiliary data) and the target domain (the object being tracked), they need to be finetuned during the tracking process. However, those methods suffer from sensitivity to the hyper-parameters such as learning rate, maximum number of epochs, size of mini-batch, and so on. Thus, it is worthy to investigate whether pretraining and fine tuning through conventional back-prop is essential for visual tracking. In this paper, we shed light on this line of research by proposing convolutional random vector functional link (CRVFL) neural network, which can be regarded as a marriage of the convolutional neural network and random vector functional link network, to simplify the visual tracking system. The parameters in the convolutional layer are randomly initialized and kept fixed. Only the parameters in the fully connected layer need to be learned. We further propose an elegant approach to update the tracker. In the widely used visual tracking benchmark, without any auxiliary data, a single CRVFL model achieves 79.0% with a threshold of 20 pixels for the precision plot. Moreover, an ensemble of CRVFL yields comparatively the best result of 86.3%.
Le Zhang 0001, Ponnuthurai N. Suganthan
IEEE Trans. Cybern.1
2017 Robust Human Activity Recognition Using Smartphone Sensors via CT-PCA and Online SVM
abstract
Human activity recognition using either wearable devices or smartphones can benefit various applications including healthcare, fitness, smart home, etc. Instead of using wearable devices which are intrusive and require extra cost, we shall leverage on modern smartphones embedded with a variety of sensors. Due to the flexibility of using smartphones, the recognition accuracy will degrade with orientation, placement, and subject variations. In this paper, we propose a robust human activity recognition system in terms of orientation, placement, and subject variations based on coordinate transformation and principal component analysis (CT-PCA) and online support vector machine (OSVM). The proposed CT-PCA scheme is utilized to eliminate the effect of orientation variations. Experiments show that the proposed scheme significantly improves the activity recognition accuracy and outperforms the state-of-the-art methods on leave one orientation out experiments, which demonstrates the generalization ability of the proposed scheme on the data from unseen orientations. We also show the effectiveness of this scheme on placement and subject variations. However, the inherent difference of signal properties for different placement and subject dramatically reduces the recognition accuracy, especially for different placement. Thus, we present an efficient OSVM algorithm, that is, online-independent support vector machine (OISVM), which utilizes a small portion of data from the unseen placement or subject to online update the parameters of the SVM algorithm. The experimental results demonstrate the effectiveness of this OISVM algorithm on placement and subject variations.
Zhenghua Chen, Qingchang Zhu, Yeng Chai Soh, Le Zhang 0001
IEEE Trans. Ind. Informatics4
2016 A survey of randomized algorithms for training neural networks
Le Zhang 0001, Ponnuthurai N. Suganthan
Inf. Sci.1
2016 A comprehensive evaluation of random vector functional link networks
Le Zhang 0001, Ponnuthurai N. Suganthan
Inf. Sci.1
2015 Statistical analysis and design of 6T SRAM cell for physical unclonable function with dual application modes
abstract
Apart from performance and power efficiency, security is another critical concern in the modern memory sub-system design. SRAM, which is routinely used as a data preservation component, has now been developed into an effective primitive known as Physical Unclonable Function (PUF) for cryptographic key generation to protect the sensitive local information. Considering the constraints of hardware resource on embedded systems, it is desirable to have an SRAM used both as a regular memory and a PUF to save the overheads of having these two functions implemented independently. Unfortunately, while process variations are the entropy sources for secure key generation, it impacts failure rates in memory-mode operations. This paper presents a statistical analysis on SRAM and provides an insight into how the SRAM cell geometry can be optimized to qualify it for both modes of operation simultaneously.
Le Zhang 0001, Chip-Hong Chang, Zhi-Hui Kong, Chao Qun Liu
ISCAS1
2015 Visual Tracking with Convolutional Neural Network
abstract
Visual Tracking is a fundamental task in computer vision which has been extensively researched. Though much progress exists in literature, it is still very challenging due to factors such as partial occlusions, pose variations, viewpoint variations and so on. In this paper, we address the visual tracking problem in a discriminant manner where a simple convolutional neural network (CNN) is employed to extract discriminant features and simultaneously classify the object from the background. The effectiveness of the proposed method is validated on a comprehensive evaluation involving 10 challenging video sequences and five state-of-the-art trackers.
Le Zhang 0001, Ponnuthurai N. Suganthan
SMC1
2015 A Low-Power Hybrid RO PUF With Improved Thermal Stability for Lightweight Applications
abstract
Ring oscillator (RO)-based physical unclonable function (PUF) is resilient against noise impacts, but its response is susceptible to temperature variations. This paper presents a low-power and small footprint hybrid RO PUF with a very high temperature stability, which makes it an ideal candidate for lightweight applications. The negative temperature coefficient of the low-power subthreshold operation of current starved inverters is exploited to mitigate the variations of differential RO frequencies with temperature. The new architecture uses conspicuously simplified circuitries to generate and compare a large number of pairs of RO frequencies. The proposed nine-stage hybrid RO PUF was fabricated using global foundry 65-nm CMOS technology. The PUF occupies only 250 μm2of chip area and consumes only 32.3 μW per challenge response pair at 1.2 V and 230 MHz. The measured average and worst-case reliability of its responses are 99.84% and 97.28%, respectively, over a wide range of temperature from -40 to 120 °C.
Yuan Cao 0003, Le Zhang 0001, Chip-Hong Chang, Shoushun Chen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2015 Optimizating Emerging Nonvolatile Memories for Dual-Mode Applications: Data Storage and Key Generator
abstract
Memory-based physical unclonable functions (PUFs) have been studied and developed as powerful primitives to generate device-specific random keys, which can be used for various security applications. However, the existing memory-based PUFs need to safely buffer the data bits in the memory before it is used to produce random bits, resulting in additional area/energy consumption and potential data security issues. In this paper, we propose a new memory-based PUF that exploits the nonvolatility and random variability of emerging memory technologies to produce random bits. Unlike conventional implementations, the random bit generation process of our proposed PUF does not disturb the data bits already stored in the memory. To satisfy the quality requirements for both memory and PUF applications, we also propose a general method to find the optimal design point of emerging nonvolatile memory (eNVM)-based PUF. An illustrative design using spin-transfer torque magnetic RAM exhibits desirable results using our method. Compared to the conventional types of memory-based PUFs, eNVM-based PUFs features enhanced security as cryptographic primitives and lower area and energy cost as data storage.
Le Zhang 0001, Xuanyao Fong, Chip-Hong Chang, Zhi-Hui Kong, Kaushik Roy 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2015 Oblique Decision Tree Ensemble via Multisurface Proximal Support Vector Machine
abstract
A new approach to generate oblique decision tree ensemble is proposed wherein each decision hyperplane in the internal node of tree classifier is not always orthogonal to a feature axis. All training samples in each internal node are grouped into two hyper-classes according to their geometric properties based on a randomly selected feature subset. Then multisurface proximal support vector machine is employed to obtain two clustering hyperplanes where each hyperplane is generated such that it is closest to one group of the data and as far as possible from the other group. Then, one of the bisectors of these two hyperplanes is regarded as the test hyperplane for this internal node. Several regularization methods have been applied to handle the small sample size problem as the tree grows. The effectiveness of the proposed method is demonstrated by 44 real-world benchmark classification data sets from various research fields. These classification results show the advantage of the proposed approach in both computation time and classification accuracy.
Le Zhang 0001, Ponnuthurai N. Suganthan
IEEE Trans. Cybern.1
2015 Highly Reliable Spin-Transfer Torque Magnetic RAM-Based Physical Unclonable Function With Multi-Response-Bits Per Cell
abstract
Memory-based physical unclonable function (MemPUF) has gained tremendous popularity in the recent years to securely preserve secret information in computing systems. Most MemPUFs in the literature have unreliable bit generation and/or are incapable of generating more than one response-bit per cell. Hence, we propose a novel MemPUF exploiting the unique characteristics of spin-transfer torque magnetic RAM (STT-MRAM) that can overcome these issues. Bit generation in our STT-MRAM-based MemPUF is stabilized using a novel automatic write-back technique. In addition, the alterability of the magnetic tunneling junction state is exploited to expand the response-bit capacity per cell. Our analysis demonstrated the advantage of our scheme in reliability enhancement (bit-error rate from ~10-1to ~10-6in the worst case under varying conditions) and response-bit capacity per cell improvement (from 1 to 1.48 bit). In comparison with the conventional MemPUFs, our approach is also better in terms of the average chip area and energy for producing a response-bit.
Le Zhang 0001, Xuanyao Fong, Chip-Hong Chang, Zhi-Hui Kong, Kaushik Roy 0001
IEEE Trans. Inf. Forensics Secur.1
2014 Towards generating random forests via extremely randomized trees
abstract
The classification error of a specified classifier can be decomposed into bias and variance. Decision tree based classifier has very low bias and extremely high variance. Ensemble methods such as bagging can significantly reduce the variance of such unstable classifiers and thus return an ensemble classifier with promising generalized performance. In this paper, we compare different tree-induction strategies within a uniform ensemble framework. The results on several public datasets show that random partition (cut-point for univariate decision tree or both coefficients and cut-point for multivariate decision tree) without exhaustive search at each node of a decision tree can yield better performance with less computational complexity.
Le Zhang 0001, Ye Ren, Ponnuthurai N. Suganthan
IJCNN1
2014 Highly reliable memory-based Physical Unclonable Function using Spin-Transfer Torque MRAM
abstract
In recent years, Physical Unclonable Function (PUF) based on the inimitable and unpredictable disorder of physical devices has emerged to address security issues related to cryptographic key generation. In this paper, a novel memory-based PUF based on Spin-Transfer Torque (STT) Magnetic RAM, named as STT-PUF, is proposed as a key generation primitive for embedded computing systems. By comparing the resistances of STT-MRAM memory cells which are initialized to the same state, response bits can be generated by exploiting the inherent random mismatches between them. To enhance the robustness of response bits regeneration, an Automatic Write-Back (AWB) technique is proposed without compromising the resilience of STT-PUF against possible attacks. Simulations show that the proposed STT-PUF is able to produce raw response bits with uniqueness of 50.1% and entropy of 0.985 bit per cell. The worst-case Bit-Error Rate (BER) under varying operating conditions is 6.6 × 10-6.
Le Zhang 0001, Xuanyao Fong, Chip-Hong Chang, Zhi-Hui Kong, Kaushik Roy 0001
ISCAS1
2014 Random Forests with ensemble of feature spaces
Le Zhang 0001, Ponnuthurai N. Suganthan
Pattern Recognit.1
2014 Exploiting Process Variations and Programming Sensitivity of Phase Change Memory for Reconfigurable Physical Unclonable Functions
abstract
Physical unclonable function (PUF) leverages the immensely complex and irreproducible nature of physical structures to achieve device authentication and secret information storage. To enhance the security and robustness of conventional PUFs, reconfigurable physical unclonable functions (RPUFs) with dynamically refreshable challenge-response pairs (CRPs) have emerged recently. In this paper, we propose two novel physically reconfigurable PUF (P-RPUF) schemes that exploit the process parameter variability and programming sensitivity of phase change memory (PCM) for CRP reconfiguration and evaluation. The first proposed PCM-based P-RPUF scheme extracts its CRPs from the measurable differences of the PCM cell resistances programmed by randomly varying pulses. An imprecisely controlled regulator is used to protect the privacy of the CRP in case the configuration state of the RPUF is divulged. The second proposed PCM-based RPUF scheme produces the random response by counting the number of programming pulses required to make the cell resistance converge to a predetermined target value. The merging of CRP reconfiguration and evaluation overcomes the inherent vulnerability of P-RPUF devices to malicious prediction attacks by limiting the number of accessible CRPs between two consecutive reconfigurations to only one. Both schemes were experimentally evaluated on 180-nm PCM chips. The obtained results demonstrated their quality for refreshable key generation when appropriate fuzzy extractor algorithms are incorporated.
Le Zhang 0001, Zhi-Hui Kong, Chip-Hong Chang, Alessandro Cabrini, Guido Torelli
IEEE Trans. Inf. Forensics Secur.1
2013 PCKGen: A Phase Change Memory based cryptographic key generator
abstract
Physical Unclonable Function (PUF) is widely known as an effective countermeasure to withstand non-invasive computational attacks as well as invasive tempering attacks on trusted computing systems. However, vast majority of the PUFs reported to-date are defined by static Challenge-Response Pairs (CRPs) with inferior security. In this paper, we propose a novel design of dynamically reconfigurable PUF based on Phase Change Memory (PCM) technology to yield refreshed cryptographic keys whenever the need arises to achieve enhanced security. A dedicated circuit framework is also introduced to reinforce the diversity of the CRP sets and improve the stability of the proposed PUF. Extensive simulation results show that our proposed work promises a clean delineation from the security bottlenecks faced by the state-of-the-art PUF designs.
Le Zhang 0001, Zhi-Hui Kong, Chip-Hong Chang
ISCAS1