EDBT 2026 Demo / reviewers in the wild / expert
Hancheng Zhu
dblp:145/1250
· DBLP profile ↗
60ranked-venue papers
13as first author
50since 2021 · last 2026
0000-0002-5418-9879ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 37 · 7 first-author · 29 since 2021Artificial intelligence and machine learning · 20 · 6 first-author · 17 since 2021Computer networks · 8 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DTTNet: Improving Video Shadow Detection via Dark-Aware Guidance and Tokenized Temporal ModelingabstractVideo shadow detection confronts two entwined difficulties: distinguishing shadows from complex backgrounds and modeling dynamic shadow deformations under varying illumination. To address shadow-background ambiguity, we leverage linguistic priors through the proposed Vision-language Match Module (VMM) and a Dark-aware Semantic Block (DSB), extracting text-guided features to explicitly differentiate shadows from dark objects. Furthermore, we introduce adaptive mask reweighting to downweight penumbra regions during training and apply edge masks at the final decoder stage for better supervision. For temporal modeling of variable shadow shapes, we propose a Tokenized Temporal Block (TTB) that decouples spatiotemporal learning. TTB summarizes cross-frame shadow semantics into learnable temporal tokens, enabling efficient sequence encoding with minimal computation overhead. Comprehensive Experiments on multiple benchmark datasets demonstrate state-of-the-art accuracy and real-time inference efficiency. Kunyang Sun, Rui Yao 0006, Hancheng Zhu, Fuyuan Hu, Jiaqi Zhao 0001, Zhiwen Shao, Yong Zhou 0003 |
AAAI | 4 |
| 2026 | Causal Decoupling Domain Generalization for Remote Sensing Change DetectionabstractWhile current state-of-the-art Remote Sensing Change Detection (RSCD) methods can achieve impressive results on individual datasets, they become unreliable in unseen environments and imaging conditions, with performance metrics declining by as much as 60% to 80%. Simultaneously, variable environments and complex imaging conditions are the main characteristics of remote sensing data, calling for generalizable RSCD methods. To address this issue, we propose a novel RSCD method capable of domain generalization—CDDGNet. This method is based on causal decoupling theory, which progressively decouples invariant change features from variable domain features to extract generalizable characteristics. This enables a network trained on a single domain to accurately identify change regions in other domains. Specifically, firstly, the Causal Feature Adaptation Module is proposed to preliminarily decouple and simplify feature information during the encoding process by using wavelet transformation and feature energy spectralization methods. Secondly, the Causal Feature Fusion Module is presented to fully decouple features and aggregate significant change features during the decoding process through frequency domain processing and feature re-attention mechanisms. Thirdly, the Decoupling Effect Loss Function is proposed to optimize the process by evaluating the effectiveness of causal decoupling. Extensive experiments have shown that our model significantly outperforms existing methods across multiple groups of generalization tasks with varying levels of difficulty. Jiaqi Zhao 0001, Jianpeng Xie 0001, Yong Zhou 0003, Wen-Liang Du 0002, Hancheng Zhu, Rui Yao 0006 |
AAAI | 5 |
| 2026 | Unified Framework for Outage-Constrained Rate Maximization in Secure ISAC Under Various Sensing MetricsabstractIntegrated sensing and communication (ISAC) is poised to redefine the landscape of wireless networks by seamlessly combining data transmission and environmental sensing. However, ISAC systems remain susceptible to eavesdropping, especially under uncertainty in eavesdroppers’ channel state information, which can lead to secrecy outages. On the other hand, diverse and complex sensing performance requirements further complicate resource optimization, often requiring custom solutions for each scenario. To this end, this paper introduces a unified optimization framework that holistically addresses both the worst-case user secrecy rate and the sum secrecy rate across multiple users. Besides putting the two commonly used objectives into a single but flexible objective function, the framework accurately controls secrecy outage probabilities while accommodating a broad spectrum of sensing constraints. To solve such a general problem, we integrate the sensing requirements into the objective function through an auxiliary variable. This enables efficient alternating optimization and the proposed approach is theoretically guaranteed to converge to at least a stationary point of the original problem. Extensive simulation results show that the proposed framework consistently achieves higher optimized secrecy rates under various sensing constraints compared to existing methods. These results underscore the proposed unified framework’s superiority and versatility in secure ISAC systems. Hancheng Zhu, Zongze Li 0002, Yik-Chung Wu |
IEEE J. Sel. Areas Commun. | 1 |
| 2026 | A unified multi-stream diffusion framework for robust video camouflaged object detection
Yuyao Ke, Rui Yao 0006, Kunyang Sun, Hancheng Zhu, Jiaqi Zhao 0001, Bing Liu 0016 |
Neural Networks | 4 |
| 2026 | Dynamic Prompt Memory Network for video shadow detection
Rui Yao 0006, Hancheng Zhu, Kunyang Sun, Jiaqi Zhao 0001, Zhiwen Shao, Abdulmotaleb El Saddik |
Pattern Recognit. | 3 |
| 2026 | TextRSR: Enhanced Arbitrary-Shaped Scene Text Representation via Robust Subspace RecoveryabstractIn recent years, scene text detection research has increasingly focused on arbitrary-shaped texts, where text representation is a fundamental problem. However, most existing methods still struggle to separate adjacent or overlapping texts due to ambiguous spatial positions of points or segmentation masks. Besides, the time efficiency of the entire pipeline is often neglected, resulting in sub-optimal inference speed. To tackle these problems, we first propose a novel text representation method based on robust subspace recovery, which robustly represents complex text shapes by combining orthogonal basis vectors learned from labeled text contours. These basis vectors capture basis contour patterns with distinct information, enabling clearer boundaries even in densely populated text scenarios. Moreover, we propose a dynamic sparse assignment scheme for positive samples that adaptively adjusts their weights during training, which not only accelerates inference speed by eliminating redundant predictions but also enhances feature learning by providing sufficient supervision signals. Building on these innovations, we present TextRSR, an accurate and efficient scene text detection network. Extensive experiments on challenging benchmarks demonstrate the superior accuracy and efficiency of TextRSR compared to state-of-the-art methods. Particularly, TextRSR achieves an F-measure of 88.5% at 37.8 frames per second (FPS) for CTW1500 dataset and an F-measure of 89.1% at 23.1 FPS for Total-Text dataset. Zhiwen Shao, Shengtian Jiang, Hancheng Zhu, Xuehuai Shi, Canlin Li, Lizhuang Ma, Dit-Yan Yeung |
IEEE Trans. Multim. | 3 |
| 2026 | Dual Sparse Long-Short Term Transformer for Video Shadow DetectionabstractVideo Shadow Detection (VSD) is critical yet challenging, primarily due to ambiguous shadow boundaries and the presence of confusing shadow-like non-shadow regions, which existing methods struggle to resolve effectively by limited temporal modeling. We propose the Dual Sparse Long-Short Term Transformer Network (DSLSTT-Net), a novel framework designed to enhance feature learning by integrating robust temporal consistency and detailed local context. DSLSTT-Net utilizes a dual-stream architecture to concurrently process global temporal information and local shadow feature refinement, enabling effective discrimination between true shadows and confusing areas. At its core, the Sparse Long-Short Term Attention Module (Sparse LSTAM) is introduced to efficiently propagate only high-confidence shadow features from memory, significantly enhancing feature discriminability and computational efficiency. Furthermore, an Adaptive Fusion Module (AFM) dynamically merges purified long-term features with short-term details, optimizing final segmentation. Experimental results confirm that DSLSTT-Net significantly outperforms state-of-the-art methods on VSD benchmarks, validating our approach of dual-stream architecture and sparse temporal modeling. The source code is available at https://github.com/rayyao/DSLSTTNet . Rui Yao 0006, Huili Hao, Hancheng Zhu, Jiaqi Zhao 0001, Yong Zhou 0003 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2026 | Spatio-Temporal Disentanglement and Constrained Self-Attention for Multi-Modal Deception DetectionabstractMulti-modal deception detection is a challenging yet important task, having pivotal applications in many fields such as business credibility assessment and multimedia anti-frauds. Previous methods either rely solely on spatial features or overemphasize only temporal information within or across modalities, which may overlook potential critical clues. Motivated by these observations, we propose a Spatio-Temporal Representation Disentanglement (STRD) framework for multi-modal deception detection, which uses a dual-encoder structure to learn spatial and temporal representations for each modality. Specifically, we introduce a pre-trained foundation model to act as the spatial encoder and design a lightweight network as the temporal encoder, extracting spatial semantics and capturing dynamic temporal patterns. Then, we propose a Constrained Self-Attention Block (CSAB), in which self-attention distribution of each head is regarded as spatial distribution and is constrained to attend a certain facial local region. Furthermore, we present a Cross-Modal Correlation Fusion Block (CCFB) to achieve temporal synchronization across modalities by measuring the correlations between visual and audio features. Extensive experiments show that our STRD outperforms the state-of-the-art methods on challenging DOLOS, BOL, BgOL, and RLtrial benchmarks. Particularly, STRD improves by 2.12% and 1.88% over the previous best results in terms of ACC on the DOLOS and BOL datasets, respectively. Additionally, STRD outperforms previous methods in cross-dataset testing, highlighting its superior generalization ability. Zhiwen Shao, Hancheng Zhu, Rui Yao 0006, Lixin Zou, Mengtian Li 0002, Bin Sheng 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2026 | A General Optimization Framework for Tackling Distance Constraints in Movable Antenna-Aided SystemsabstractThe recently emerged movable antenna (MA) shows great potential in leveraging spatial degrees of freedom for enhancing the performance of wireless systems. However, resource allocation in MA-aided systems faces unique challenges due to the non-convex and coupled constraints on antenna positions. This paper systematically reveals the challenges brought by the minimum MA separation constraints, and proposes a penalty framework for resource allocation under such new constraints in MA-aided systems. By introducing auxiliary variables, the proposed framework separates the non-convex and coupled antenna distance constraints from the movable region constraint. This enables the resulting problem be efficiently solved by alternating optimization, where the optimization of the original variables resembles that in conventional resource allocation problem while the optimization with respect to the auxiliary variables is achieved in closed-form solutions. To illustrate the effectiveness of the proposed framework, we present three case studies: capacity maximization, latency minimization, and regularized zero-forcing precoding. Simulation results demonstrate that the proposed optimization framework consistently outperforms state-of-the-art schemes. Yichen Jin, Qingfeng Lin, Yang Li 0035, Hancheng Zhu, Bingyang Cheng, Yik-Chung Wu, Rui Zhang 0006 |
IEEE Trans. Wirel. Commun. | 4 |
| 2025 | Facial Action Unit Detection with Iterative Rank Reduction Adapter and Directional Attention
Zhiwen Shao, Hancheng Zhu, Rui Yao 0006, Bing Liu 0016 |
CGI (1) | 4 |
| 2025 | ReDiffDet: Rotation-equivariant Diffusion Model for Oriented Object DetectionabstractThe diffusion model has been successfully applied to various detection tasks. However, it still faces several challenges when used for oriented object detection: objects that are arbitrarily rotated require the diffusion model to encode their orientation information; uncontrollable random boxes inaccurately locate objects with dense arrangements and extreme aspect ratios; oriented boxes result in the misalignment between them and image features. To overcome these limitations, we propose ReDiffDet, a framework that formulates oriented object detection as a rotation-equivariant denoising diffusion process. First, we represent an oriented box as a 2D Gaussian distribution, forming the basis of the denoising paradigm. The reverse process can be proven to be rotation-equivariant within this representation and model framework. Second, we design a conditional encoder with conditional boxes to prevent boxes from being randomly placed across the entire image. Third, we propose an aligned decoder for alignment between oriented boxes and image features. The extensive experiments demonstrate ReDiffDet achieves promising performance and significantly outperforms the diffusion-based baseline detector. Codes are available at https://github.com/wokaikaixinxin/ReDiffDet. Jiaqi Zhao 0001, Zeyu Ding 0010, Yong Zhou 0003, Hancheng Zhu, Wen-Liang Du 0002, Rui Yao 0006 |
CVPR | 4 |
| 2025 | GSDet: Gaussian Splatting for Oriented Object DetectionabstractOriented object detection has advanced with the development of convolutional neural networks (CNNs) and transformers. However, modern detectors still rely on predefined object candidates, such as anchors in CNN-based methods or queries in transformer-based methods, which struggle to capture spatial information effectively. To address the limitations, we propose GSDet, a novel framework that formulates oriented object detection as Gaussian splatting. Specifically, our approach performs detection within a 3D feature space constructed from image features, where 3D Gaussians are employed to represent oriented objects. These 3D Gaussians are projected onto the image plane to form 2D Gaussians, which are then transformed into oriented boxes. Furthermore, we optimize the mean, anisotropic covariance, and confidence scores of these randomly initialized 3D Gaussians, using a decoder that incorporates 3D Gaussian sampling. Moreover, our method exhibits flexibility, enabling adaptive control and a dynamic number of Gaussians during inference. Experiments on 3 datasets indicate that GSDet achieves AP50 gains of 0.7% on DIOR-R, 0.3% on DOTA-v1.0, and 0.55% on DOTA-v1.5 when evaluated with adaptive control and outperforms mainstream detectors. Zeyu Ding 0010, Jiaqi Zhao 0001, Yong Zhou 0003, Wen-Liang Du 0002, Hancheng Zhu, Rui Yao 0006 |
IJCAI | 5 |
| 2025 | Beyond Individual and Point: Next POI Recommendation via Region-aware Dynamic Hypergraph with Dual-level ModelingabstractNext POI recommendation contributes to the prosperity of various intelligent location-based services. Existing studies focus on exploring sequential patterns and POI interactions using sequential and graph-based methods to enhance recommendation performance. However, they don't effectively exploit geographical information. In addition, methods that focus on modeling mobility patterns using individual limited data may suffer from data sparsity and the information cocoons problem. Moreover, most graph structures focus on adjacent nodes, failing to capture potential high-order associations among POIs. To address these challenges, we propose the Region-aware dynamic Hypergraph learning method with Dual-level interaction Modeling (ReHDM), which exploits users' dynamic mobility beyond individual and point. Specifically, ReHDM utilizes regional encoding to mine the potential spatial relationships among POIs with coarse-grained geographical information. By incorporating POI-level and trajectory-level associations within a hypergraph convolutional network, ReHDM comprehensively captures cross-user collaborative information. Furthermore, ReHDM captures not only dependencies among POIs within each trajectory for a single user, but also the high-order collaborative information across individual user trajectories and associated users' trajectories. Experimental results on three public datasets demonstrate the superiority of ReHDM to the state-of-the-art. Zhuo Gu, Rui Yao 0006, Yong Zhou 0003, Hancheng Zhu, Jiaqi Zhao 0001, Wen-Liang Du 0002 |
IJCAI | 5 |
| 2025 | Modality-Guided Dynamic Graph Fusion and Temporal Diffusion for Self-Supervised RGB-T TrackingabstractTo reduce the reliance on large-scale annotations, self-supervised RGB-T tracking approaches have garnered significant attention. However, the omission of the object region by erroneous pseudo-label or the introduction of background noise affects the efficiency of modality fusion, while pseudo-label noise triggered by similar object noise can further affect the tracking performance. In this paper, we propose GDSTrack, a novel approach that introduces dynamic graph fusion and temporal diffusion to address the above challenges in self-supervised RGB-T tracking. GDSTrack dynamically fuses the modalities of neighboring frames, treats them as distractor noise, and leverages the denoising capability of a generative model. Specifically, by constructing an adjacency matrix via an Adjacency Matrix Generator (AMG), the proposed Modality-guided Dynamic Graph Fusion (MDGF) module uses a dynamic adjacency matrix to guide graph attention, focusing on and fusing the object’s coherent regions. Temporal Graph-Informed Diffusion (TGID) models MDGF features from neighboring frames as interference, and thus improving robustness against similar-object noise. Extensive experiments conducted on four public RGB-T tracking datasets demonstrate that GDSTrack outperforms the existing state-of-the-art methods. The source code is available at https://github.com/LiShenglana/GDSTrack. Shenglan Li, Rui Yao 0006, Yong Zhou 0003, Hancheng Zhu, Kunyang Sun, Bing Liu 0016, Zhiwen Shao, Jiaqi Zhao 0001 |
IJCAI | 4 |
| 2025 | RQFormer: Rotated Query Transformer for end-to-end oriented object detection
Jiaqi Zhao 0001, Zeyu Ding 0010, Yong Zhou 0003, Hancheng Zhu, Wen-Liang Du 0002, Rui Yao 0006, Abdulmotaleb El Saddik |
Expert Syst. Appl. | 4 |
| 2025 | Facial Action Unit Detection by Adaptively Constraining Self-Attention and Causally Deconfounding Sample
Zhiwen Shao, Hancheng Zhu, Yong Zhou 0003, Xiang Xiang 0001, Bing Liu 0016, Rui Yao 0006, Lizhuang Ma |
Int. J. Comput. Vis. | 2 |
| 2025 | Revisiting trace norm minimization for tensor Tucker completion: A direct multilinear rank learning approach
Xueke Tong, Hancheng Zhu, Lei Cheng 0003, Yik-Chung Wu |
Pattern Recognit. | 2 |
| 2025 | Progressively Generated Text-Assisted Image Aesthetic Quality AssessmentabstractImage Aesthetic Quality Assessment (IAQA) aims to simulate user perceptions to judge the aesthetic quality of images. Due to the high subjectivity of users and the complexity of image aesthetics, modeling IAQA solely at the image level is a compromise. Consequently, existing methods mainly focus on multimodal-based models and achieve effective performance. These methods explore aesthetic comments on images to characterize users and serve as auxiliary text information for multimodal modeling. Unfortunately, this may suffer from two limitations. One limitation is that aesthetic comments are often unavailable for an unknown image in the test phase, and another limitation is that the semantic information of these comments may be uncertain and fuzzy. Therefore, this paper proposes a progressively generated text-assisted image aesthetic quality assessment method, aiming to address the lack of aesthetic comments and the fuzziness of aesthetic judgments in these comments. Specifically, we first adopt a Multimodal Large Language Model (MLLM) to generate aesthetic comments on images by simulating user perceptions and utilize the generated comments to characterize their aesthetic perception to assist in the pre-training of our multimodal-based IAQA model. Then, we design an attribute prediction module to determine the attribute levels of aesthetic judgments and utilize text template construction to further generate explicit descriptions of image aesthetics. Finally, we leverage the generated attribute descriptions to further assist in training our IAQA model. By progressively generating textual auxiliary descriptions of aesthetics for images, the proposed model can gradually determine the aesthetic quality of the images. Massive experimental results indicate that the proposed method outperforms existing mainstream methods on multiple IAQA datasets. Hancheng Zhu, Ju Shi, Zhiwen Shao, Rui Yao 0006, Kunyang Sun, Leida Li |
IEEE Trans. Fuzzy Syst. | 1 |
| 2025 | Hyperspectral Object Tracking With Dual-Stream PromptabstractHyperspectral images, rich in spectral details, offeradvantages for object tracking across diverse scenarios. Current hyperspectral tracking often fine-tunes parameters using pretrained RGB trackers, but this manner is suboptimal due to redundancy in spectral bands and limited training data. Existing hyperspectral trackers also underuse temporal information. To address these issues, we propose a unified spectral-spatiotemporal multimodal dual-stream prompt hyperspectral object tracking, named HDSP. We design a density clustering-based band selection module (BSM) to preserve spectral prompt information efficiently. Using the generated bands and temporal data as multimodal prompts, a dual-stream visual prompter is proposed. Designed multimodal dual-stream visual prompter (MDVP) transforms the multimodal input into a single modality, enhancing the foundational modality’s representation capabilities for hyperspectral tracking. Experiments on hyperspectral videos (HSVs) tracking datasets demonstrate that the proposed tracker achieves state-of-the-art performance. The source code is available athttps://github.com/rayyao/HDSP. Rui Yao 0006, Yong Zhou 0003, Hancheng Zhu, Jiaqi Zhao 0001, Zhiwen Shao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | CCFL: Customized Client Federated Learning for Unsupervised Person Re-identificationabstractFederated learning-based person re-identification (Re-ID) aims to address the issue of data silos in surveillance systems caused by increasingly stringent regulations on sensitive data. However, due to differences in data collection locations, times, and scales, severe non-independent and identically distributed (non-IID) characteristics exist across different Re-ID datasets. Existing federated learning-based Re-ID methods often adopt a unified model structure, which prevents the model from adapting well to diverse data environments, thereby significantly degrading the overall Re-ID performance. To address the challenges of training neural networks on non-IID data across different datasets, we propose a customizable federated learning framework. First, customizable clients allow each organization to freely select suitable neural network training methods and model architectures based on local data scales and prior knowledge, thus improving training outcomes. Second, since traditional federated learning frameworks cannot achieve knowledge fusion through parameter exchange between models with different architectures, we introduce an independent model, referred to as the interaction model, specifically designed for knowledge exchange among clients. The interaction model learns parameters (knowledge) from local models on each client through distillation learning. Subsequently, the interaction model is uploaded to the server, where it undergoes parameter fusion (knowledge exchange) with interaction models from other clients. Finally, the interaction model, enriched with knowledge from other clients, guides local model training through knowledge distillation. It is worth noting that selecting a lightweight interaction model, while potentially impacting Re-ID performance, can significantly reduce communication costs between the server and clients. Yong Zhou 0003, Fayao Liu, Jiaqi Zhao 0001, Hancheng Zhu, Wen-Liang Du 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | Image Cropping with Content and Composition Attribute-aware Global Relation ReasoningabstractImage cropping aims to find visually pleasing content in an image, which will enhance its aesthetic quality. Existing image cropping approaches mainly emphasize the geometric properties of images, such as composition and layout, neglecting the rich aesthetic information available from the physical attributes (e.g., content and themes), and background information beyond the foreground in images. Consequently, this article proposes an image cropping method based on the content and composition attribute-aware global relation reasoning, which aims at guiding the generation of cropped sub-images by exploring critical attributes based on content and composition as well as global object correlations that affect aesthetics in images. Particularly, to comprehensively introduce aesthetic information into image cropping, we capture feature representations reinforced by content and composition attributes simultaneously. The feature representations can strengthen the visual aesthetics of cropped sub-images. To make the cropped sub-images amply contain more global information, we introduce a global relation reasoning branch in the proposed cropping module, which can fully exploit the dependency relationship between the foreground and background in images. Extensive experiments on image cropping benchmarks demonstrate that our approach is superior to state-of-the-art image cropping methods. Hancheng Zhu, Yong Zhou 0003, Rui Yao 0006, Zhiwen Shao, Jiaqi Zhao 0001, Leida Li |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Attribute-Driven Multimodal Hierarchical Prompts for Image Aesthetic Quality AssessmentabstractImage Aesthetic Quality Assessment (IAQA) aims to simulate users' visual perception to judge the aesthetic quality of images. In social media, users' aesthetic experiences are often reflected in their textual comments regarding the aesthetic attributes of images. To fully explore the attribute information perceived by users for evaluating image aesthetic quality, this paper proposes an image aesthetic quality assessment method based on attribute-driven multimodal hierarchical prompts. Unlike existing IAQA methods that utilize multimodal pre-training or straightforward prompts for model learning, the proposed method leverages attribute comments and quality-level text templates to hierarchically learn the aesthetic attributes and quality of images. Specifically, we first leverage users' aesthetic attribute comments to perform prompt learning on images. The learned attribute-driven multimodal features can comprehensively capture the semantic information of image aesthetic attributes perceived by users. Then, we construct text templates for different aesthetic quality levels to further facilitate prompt learning through semantic information related to the aesthetic quality of images. The proposed method can explicitly simulate users' aesthetic judgment of images to obtain more precise aesthetic quality. Experimental results demonstrate that the proposed IAQA method based on hierarchical prompts outperforms existing methods significantly on multiple IAQA databases. Our source code is public at https://github.com/GitHub-Ju/AMHP. Hancheng Zhu, Ju Shi, Zhiwen Shao, Rui Yao 0006, Yong Zhou 0003, Jiaqi Zhao 0001, Leida Li |
ACM Multimedia | 1 |
| 2024 | A Mamba-Diffusion Framework for Multimodal Remote Sensing Image Semantic SegmentationabstractRecent advances in deep learning have made significant progress in multimodal remote sensing semantic segmentation. However, current methods face challenges in maintaining geometric consistency, particularly when dealing with large objects, resulting in fragmented segmentation masks. We propose a Mamba-diffusion framework to preserve geometric consistency in segmentation masks. This framework preserves geometric consistency by introducing a generative diffusion-based semantic segmentation pipeline and developing a Mamba-based multimodal fusion model. The fusion model fuses the multimodal images in multiple scales and scanning mechanisms by a double cross-fusion (DCF) module. Then, the cross-modal information is further integrated by a dual-splitting structured state-space (DS-S4) model. Finally, the diffusion-based segmentation pipeline predicts semantic masks by progressively refining random Gaussian noise, guided by fused multimodal features. Our experimental results, verified on WHU-OPT-SAR and Hunan datasets, demonstrate that the proposed framework surpasses state-of-the-art (SOTA) methods by a considerable margin. Our codes are available athttps://github.com/WenliangDu/MambaDiffusion. Wen-Liang Du 0002, Yang Gu 0005, Jiaqi Zhao 0001, Hancheng Zhu, Rui Yao 0006, Yong Zhou 0003 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | Multi-level self attention for unsupervised learning person re-identification
Jiaqi Zhao 0001, Yong Zhou 0003, Fayao Liu, Rui Yao 0006, Hancheng Zhu, Abdulmotaleb El Saddik |
Multim. Tools Appl. | 6 |
| 2024 | Joint facial action unit recognition and self-supervised optical flow estimation
Zhiwen Shao, Yong Zhou 0003, Feiran Li, Hancheng Zhu, Bing Liu 0016 |
Pattern Recognit. Lett. | 4 |
| 2024 | CT-Net: Arbitrary-Shaped Text Detection via Contour TransformerabstractContour based scene text detection methods have rapidly developed recently, but still suffer from inaccurate front-end contour initialization, multi-stage error accumulation, or deficient local information aggregation. To tackle these limitations, we propose a novel arbitrary-shaped scene text detection framework named CT-Net by progressive contour regression with contour transformers. Specifically, we first employ a contour initialization module that generates coarse text contours without any post-processing. Then, we adopt contour refinement modules to adaptively refine text contours in an iterative manner, which are beneficial for context information capturing and progressive global contour deformation. Besides, we propose an adaptive training strategy to enable the contour transformers to learn more potential deformation paths, and introduce a re-score mechanism that can effectively suppress false positives. Extensive experiments are conducted on four challenging datasets, which demonstrate the accuracy and efficiency of our CT-Net over state-of-the-art methods. Particularly, CT-Net achieves F-measure of 86.1 at 11.2 frames per second (FPS) and F-measure of 87.8 at 10.1 FPS for CTW1500 and Total-Text datasets, respectively. Zhiwen Shao, Yong Zhou 0003, Hancheng Zhu, Bing Liu 0016, Rui Yao 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Information Gap Narrowing for Point Cloud Few-Shot SegmentationabstractPoint-by-point labeling of point clouds is a very costly task. Previous meta-learning-based few-shot methods predict categories by calculating the distance between unlabeled data (query set) and the prototype calculated by a few of data with the label (support set), which can reduce the dependence of point cloud segmentation algorithms on large amounts of labeled data. But it ignores the category information gap caused by object diversity between the two types of data and forcing information transfer is ineffective. To address this issue, we propose a co-occurrent object mining module for mining co-occurring object information from support and query sets. Specifically, the capture of co-occurrent information is used to activate the feature that co-occurs between the support and query set in the high-dimensional feature space so that the prototype generated by computing the mean of support features is more similar to the query set. By reducing the object diversity within the same category, the information gap problem is gradually improved. In addition, we propose a point-attention module to refine the support set features before mining co-occurrent features. It can be widely embedded in the point cloud backbone network. The experimental results on two semantic segmentation datasets demonstrate that our method obtains an average 19.43% lead over the state-of-the-art methods in 4 different few-shot tasks, while inference is around 45 times faster. Guanyu Zhu, Yong Zhou 0003, Rui Yao 0006, Hancheng Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | OrientedFormer: An End-to-End Transformer-Based Oriented Object Detector in Remote Sensing ImagesabstractOriented object detection in remote sensing images is a challenging task due to objects being distributed in multiorientation. Recently, end-to-end transformer-based methods have achieved success by eliminating the need for post-processing operators compared to traditional convolutional neural network (CNN)-based methods. However, directly extending transformers to oriented object detection presents three main issues: 1) objects rotate arbitrarily, necessitating the encoding of angles along with position and size; 2) the geometric relations of oriented objects are lacking in self-attention, due to the absence of interaction between content and positional queries; and 3) oriented objects cause misalignment, mainly between values and positional queries in cross-attention, making accurate classification and localization difficult. In this article, we propose an end-to-end transformer-based oriented object detector, consisting of three dedicated modules to address these issues. First, Gaussian positional encoding (PE) is proposed to encode the angle, position, and size of oriented boxes using Gaussian distributions. Second, Wasserstein self-attention is proposed to introduce geometric relations and facilitate interaction between content and positional queries by utilizing Gaussian Wasserstein distance scores. Third, oriented cross-attention is proposed to align values and positional queries by rotating sampling points around the positional query according to their angles. Experiments on six datasets DIOR-R, a series of DOTA, HRSC2016, and ICDAR2015 show the effectiveness of our approach. Compared with previous end-to-end detectors, the OrientedFormer gains 1.16 and 1.21 AP50 on DIOR-R and DOTA-v1.0, respectively, while reducing training epochs from$3\times $to$1\times $. The code is available athttps://github.com/wokaikaixinxin/OrientedFormer. Jiaqi Zhao 0001, Zeyu Ding 0010, Yong Zhou 0003, Hancheng Zhu, Wen-Liang Du 0002, Rui Yao 0006, Abdulmotaleb El Saddik |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Motion-Aware Self-Supervised RGBT Tracking with Multi-Modality Hierarchical TransformersabstractSupervised RGBT (SRGBT) tracking tasks need both expensive and time-consuming annotations. Therefore, the implementation of Self-Supervised RGBT (SSRGBT) tracking methods has become increasingly important. Straightforward SSRGBT tracking methods use pseudo-labels for tracking, but inaccurate pseudo-labels can lead to object drift, which severely affects tracking performance. This article proposes a self-supervised RGBT object tracking method (S2OTFormer) to bridge the gap between tracking methods supervised under pseudo-labels and ground truth labels. Firstly, to provide more robust appearance features for motion cues, we introduce a multi-modality hierarchical transformer (MHT) module for feature fusion. This module allocates weights to both modalities and strengthens the expressive capability of the MHT module through multiple nonlinear layers to fully utilize the complementary information of the two modalities. Secondly, in order to solve the problems of motion blur caused by camera motion and inaccurate appearance information caused by pseudo-labels, we introduce a motion-aware mechanism (MAM). The MAM extracts the average motion vectors from the previous multi-frame search frame features and constructs the consistency loss with the motion vectors of the current search frame features. The motion vectors of inter-frame objects are obtained by reusing the inter-frame attention map to predict coordinate positions. Finally, to further reduce the effect of inaccurate pseudo-labels, we propose an Attention-Based Multi-Scale Enhancement Module. By introducing cross-attention to achieve more precise and accurate object tracking, this module overcomes the receptive field limitations of traditional CNN tracking heads. We demonstrate the effectiveness of S2OTFormer on four large-scale public datasets through extensive comparisons as well as numerous ablation experiments. The source code is available at https://github.com/LiShenglana/S2OTFormer . Shenglan Li, Rui Yao 0006, Yong Zhou 0003, Hancheng Zhu, Jiaqi Zhao 0001, Zhiwen Shao, Abdulmotaleb El Saddik |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Personalized Image Aesthetics Assessment with Attribute-guided Fine-grained Feature RepresentationabstractPersonalized image aesthetics assessment (PIAA) has gained increasing attention from researchers due to its ability to measure individual users' specific aesthetic experiences. However, most existing PIAA methods rely on holistic features or simplistic coding to characterize users' aesthetic preferences for images, and we believe that more rich explicit features are needed in modeling PIAA. Consequently, we propose an attribute-guided fine-grained feature-aware personalized image aesthetics assessment method, which can fully capture fine-grained features from multiple attributes to represent users' aesthetic preferences for images. To achieve this, we first build a fine-grained feature extraction (FFE) module to obtain the refined local features of image attributes to compensate for holistic features. The FFE module is then used to generate user-level features, which are combined with the image-level features to obtain user-preferred fine-grained feature representations. By training extensive users' PIAA tasks, the aesthetic distribution of most users can be transferred to the personalized scores of individual users. To enable our proposed model to learn more generalizable aesthetics among individual users, we incorporate the degree of dispersion between users' personalized scores and image aesthetic distribution as a coefficient in the loss function during model training. Experimental results on several PIAA databases show that our method outperforms existing mainstream PIAA methods, and can effectively infer users' personalized aesthetics of images. Hancheng Zhu, Zhiwen Shao, Yong Zhou 0003, Guangcheng Wang, Pengfei Chen 0003, Leida Li |
ACM Multimedia | 1 |
| 2023 | Identity-invariant representation and transformer-style relation for micro-expression recognition
Zhiwen Shao, Feiran Li, Yong Zhou 0003, Hancheng Zhu, Rui Yao 0006 |
Appl. Intell. | 5 |
| 2023 | Visible and Infrared Object Tracking via Convolution-Transformer Network With Joint Multimodal Feature LearningabstractThe existing Transformer-based RGBT tracker mainly focus on the enhancement of features extracted by Convolutional Neural Network (CNN). The potential of Transformer in representation learning remains under-explored. In this paper, we propose a Convolution-Transformer network with joint multimodal feature learning, in which both representation learning and feature fusion leverage Transformer. Specifically, we use the multi-branch Convolution-Transformer feature extraction network to process the extraction task of local modality-independent features and global modality-shared features respectively. Several simplified Transformer encoder layers form the Transformer backbone network, which is more suitable for the real-time object tracking. Besides, we found that inter-modality correlation is an important factor for modality interactions and mutual exploitation. Therefore, we propose a Joint Multimodal Feature Learning (JMFL) module, which uses cross-attention to capture the dependencies of cross-modal and enhance multimodal fusion by bidirectional guidance of multimodal information. The proposed method is fully experimented on two large benchmark datasets and compared with some current well-performing methods. The experimental results show that the proposed method performs well in terms of tracking accuracy and speed. Jiazhu Qiu, Rui Yao 0006, Yong Zhou 0003, Peng Wang 0015, Yanning Zhang 0001, Hancheng Zhu |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2023 | Unsupervised RGB-T object tracking with attentional multi-modal feature fusion
Shenglan Li, Rui Yao 0006, Yong Zhou 0003, Hancheng Zhu, Bing Liu 0016, Jiaqi Zhao 0001, Zhiwen Shao |
Multim. Tools Appl. | 4 |
| 2023 | Semi-supervised transformable architecture search for feature distillation
Man Zhang 0006, Yong Zhou 0003, Bing Liu 0016, Jiaqi Zhao 0001, Rui Yao 0006, Zhiwen Shao, Hancheng Zhu |
Pattern Anal. Appl. | 7 |
| 2023 | Facial Action Unit Detection via Adaptive Attention and RelationabstractFacial action unit (AU) detection is challenging due to the difficulty in capturing correlated information from subtle and dynamic AUs. Existing methods often resort to the localization of correlated regions of AUs, in which predefining local AU attentions by correlated facial landmarks often discards essential parts, or learning global attention maps often contains irrelevant areas. Furthermore, existing relational reasoning methods often employ common patterns for all AUs while ignoring the specific way of each AU. To tackle these limitations, we propose a novel adaptive attention and relation (AAR) framework for facial AU detection. Specifically, we propose an adaptive attention regression network to regress the global attention map of each AU under the constraint of attention predefinition and the guidance of AU detection, which is beneficial for capturing both specified dependencies by landmarks in strongly correlated regions and facial globally distributed dependencies in weakly correlated regions. Moreover, considering the diversity and dynamics of AUs, we propose an adaptive spatio-temporal graph convolutional network to simultaneously reason the independent pattern of each AU, the inter-dependencies among AUs, as well as the temporal dependencies. Extensive experiments show that our approach (i) achieves competitive performance on challenging benchmarks including BP4D, DISFA, and GFT in constrained scenarios and Aff-Wild2 in unconstrained scenarios, and (ii) can precisely learn the regional correlation distribution of each AU. Zhiwen Shao, Yong Zhou 0003, Jianfei Cai 0001, Hancheng Zhu, Rui Yao 0006 |
IEEE Trans. Image Process. | 4 |
| 2023 | TextDCT: Arbitrary-Shaped Text Detection via Discrete Cosine Transform MaskabstractArbitrary-shaped scene text detection is a challenging task due to the variety of text changes in font, size, color, and orientation. Most existing regression based methods resort to regress the masks or contour points of text regions to model the text instances. However, regressing the complete masks requires high training complexity, and contour points are not sufficient to capture the details of highly curved texts. To tackle the above limitations, we propose a novel light-weight anchor-free text detection framework called TextDCT, which adopts the discrete cosine transform (DCT) to encode the text masks as compact vectors. Further, considering the imbalanced number of training samples among pyramid layers, we only employ a single-level head for top-down prediction. To model the multi-scale texts in a single-level head, we introduce a novel positive sampling strategy by treating the shrunk text region as positive samples, and design a feature awareness module (FAM) for spatial-awareness and scale-awareness by fusing rich contextual information and focusing on more significant features. Moreover, we propose a segmented non-maximum suppression (S-NMS) method that can filter low-quality mask regressions. Extensive experiments are conducted on four challenging datasets, which demonstrate our TextDCT obtains competitive performance on both accuracy and efficiency. Specifically, TextDCT achieves F-measure of 85.1 at 17.2 frames per second (FPS) and F-measure of 84.9 at 15.1 FPS for CTW1500 and Total-Text datasets, respectively. Zhiwen Shao, Yong Zhou 0003, Hancheng Zhu, Bing Liu 0016, Rui Yao 0006 |
IEEE Trans. Multim. | 5 |
| 2023 | Weakly Supervised Few-Shot Semantic Segmentation via Pseudo Mask Enhancement and Meta LearningabstractFew shot semantic segmentation has been proposed to enhance the generalization ability of traditional models with limited data. Previous works mainly focus on the supervised tasks, while limited amount of work is explored for the weakly supervised tasks. Weakly supervised semantic segmentation has become an active research area because weakly supervised labels effectively reduce the annotation cost of visual tasks. To this end, we propose a weakly supervised few-shot semantic segmentation model based on the meta learning framework, which utilizes prior knowledge and adjusts itself according to new tasks. Thereupon then, the proposed network is capable of both high efficiency and generalization ability to new tasks. In the pseudo mask generation stage, we develop a WRCAM method with the channel-spatial attention mechanism to refine the coverage size of targets in pseudo masks. In the few-shot semantic segmentation stage, the optimization based meta learning method is used to realize few-shot semantic segmentation by virtue of the refined pseudo masks. The experimental results show that the proposed method not only significantly outperforms weakly supervised SOTA methods, but also could be comparative to some supervised SOTA methods. Man Zhang 0006, Yong Zhou 0003, Bing Liu 0016, Jiaqi Zhao 0001, Rui Yao 0006, Zhiwen Shao, Hancheng Zhu |
IEEE Trans. Multim. | 7 |
| 2023 | Learning Personalized Image Aesthetics From Subjective and Objective AttributesabstractDue to the widespread popularity of social media, researchers have developed a strong interest in learning the personalized image aesthetics of online users. Personalized image aesthetics assessment (PIAA) aims to study the aesthetic preferences of individual users for images, which should be affected by the properties of both users and images. Existing PIAA approaches usually use the generic aesthetics learned from images as a prior model and adapt it to PIAA models through a small number of data annotated by individual users. However, the prior model merely learns the objective attributes of images, which is agnostic to the subjective attributes of users, complicating efficient learning of the personalized image aesthetics of individual users. Therefore, we propose a personalized image aesthetics assessment method that integrates the subjective attributes of users and objective attributes of images simultaneously. To characterize these two attributes jointly, an attribute extraction module is introduced to learn users’ personality traits and image aesthetic attributes. Then, an aesthetic prior model is built from numerous individual users’ annotated data, which leverages the personality traits of users and the aesthetic attributes of rated images as prior knowledge to model both the image aesthetic distribution and users’ residual scores relative to generic aesthetics simultaneously. Finally, a PIAA model is obtained by fine-tuning the aesthetic prior model with an individual user’s annotated data. Experiments demonstrate that the proposed method is superior to existing PIAA methods in learning individual users’ personalized image aesthetics. Hancheng Zhu, Yong Zhou 0003, Leida Li, Yandong Guo |
IEEE Trans. Multim. | 1 |
| 2023 | Cross-Class Bias Rectification for Point Cloud Few-Shot SegmentationabstractThe point cloud is a densely distributed 3D (three-dimensional) data, and annotating the point cloud is a time-consuming and labor-intensive work. The existing semantics segmentation work adopts few-shot learning to reduce the dependence on labeling samples while improving the generalization of the model to new categories. Since point clouds are 3D structures with rich geometric features, even objects of the same category have feature differences that cannot be ignored. Therefore, a few samples (support set) used to train the model do not cover all the features of this category. There is a distribution difference between the support samples and the samples used to verify the model performance (query set). In this paper, we propose an efficient point cloud few-shot segmentation method based on prototypes for bias rectification. A prototype is a vector representation of a category in the metric space. To make the prototype representation of the support set closer to the query set features, we define a feature bias term and reduce the distribution distance between the two sets by fusing the support set features and the bias term. On this basis, we design a feature cross-reference module. By mining the co-occurring features of the support and query sets, it can generate a more representative prototype which captures the overall features of the point cloud. Extensive experiments on two challenging datasets demonstrate that our method outperforms the state-of-the-art method by an average of 3.31$\%$in several N-way K-shot tasks, and achieves approximately 200 times faster reasoning speed. Our code is available athttps://github.com/964918993/2CBR. Guanyu Zhu, Yong Zhou 0003, Rui Yao 0006, Hancheng Zhu |
IEEE Trans. Multim. | 4 |
| 2023 | Cyclic Self-attention for Point Cloud RecognitionabstractPoint clouds provide a flexible geometric representation for computer vision research. However, the harsh demands for the number of input points and computer hardware are still significant challenges, which hinder their deployment in real applications. To address these challenges, we design a simple and effective module named cyclic self-attention module (CSAM). Specifically, three attention maps of the same input are obtained by cyclically pairing the feature maps, thus exploring the features sufficiently of the attention space of the original input. CSAM can adequately explore the correlation between points to obtain sufficient feature information despite the multiplicative decrease in inputs. Meanwhile, it can direct the computational power to the more essential features, relieving the burden on the computer hardware. We build a point cloud classification network by simply stacking CSAM called cyclic self-attention network (CSAN). We also propose a novel framework for point cloud semantic segmentation called full cyclic self-attention network (FCSAN). By adaptively fusing the original mapping features and the CSAM extracted features, it can better capture the context information of point clouds. Extensive experiments on several benchmark datasets show that our methods can achieve competitive performance in classification and segmentation tasks. Guanyu Zhu, Yong Zhou 0003, Rui Yao 0006, Hancheng Zhu, Jiaqi Zhao 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2022 | Two-stage unsupervised facial image quality measurement
Guangcheng Wang, Zhongyuan Wang 0001, Baojin Huang, Kui Jiang, Zheng He 0001, Hancheng Zhu, Jinsheng Xiao, Xin Tian 0006 |
Inf. Sci. | 6 |
| 2022 | Personality modeling from image aesthetic attribute-aware graph representation learning
Hancheng Zhu, Yong Zhou 0003, Qiaoyue Li, Zhiwen Shao |
J. Vis. Commun. Image Represent. | 1 |
| 2022 | A Semi-Supervised Image-to-Image Translation Framework for SAR-Optical Image MatchingabstractSynthetic Aperture Radar (SAR) and optical image matching aims to acquire correspondences from a certain pair of SAR and optical images. Recent advances in the image-to-image translation provided a way to simplify the SAR-optical image matching into the SAR-SAR or optical-optical image matchings. Existing image-to-image translations mainly focus on supervised or unsupervised learning. However, gathering sufficient amounts of aligned training data for supervised learning is challenging, while unsupervised learning cannot guarantee enough correct correspondences. In this work, we investigate the applicability of semi-supervised image-to-image translation for SAR-optical image matching such that both aligned and unaligned SAR-optical images could be used. To this end, we combine the benefits of both supervised and unsupervised well-known image-to-image translation methods, i.e., Pix2pix and CycleGAN, and propose a simple yet effective semi-supervised image-to-image translation framework. Through extensive experimental comparisons to baseline methods, we verify the effectiveness of the proposed framework in both semi-supervised and fully-supervised settings. Our codes are available at https://github.com/WenliangDu/Semi-I2I. Wen-Liang Du 0002, Yong Zhou 0003, Hancheng Zhu, Jiaqi Zhao 0001, Zhiwen Shao, Xiaolin Tian 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Few-Shot Object Detection via Context-Aware Aggregation for Remote Sensing ImagesabstractFew-shot object detection methods have made prodigious progress in recent years. However, these methods are designed for optical images at a single scale, which leads to significantly degraded detection performance due to object scale variation of remote sensing images. In this letter, we propose a few-shot object detection method for the problem of scale variation in remote sensing images. More specifically, our model contains two main components: a context-aware pixel aggregation (CPA) that allows the model to adapt to objects at different scales through different scale convolution and a context-aware feature aggregation (CFA) that enhances context awareness to obtain more semantic information through a graph convolution network (GCN). Experiments on the DIOR dataset demonstrate that our model can achieve a satisfying detection performance on remote sensing images, and our model performs significantly better than the state-of-the-art model. Yong Zhou 0003, Jiaqi Zhao 0001, Hancheng Zhu, Rui Yao 0006, Wen-Liang Du 0002 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Fine-Grained Feature Enhancement for Object Detection in Remote Sensing ImagesabstractRecently, object detection in aerial images has ushered in a new challenge—a new benchmark for fine-grained object recognition in high-resolution remote sensing imagery called FAIR1M has been proposed. Fine-grained categories usually have smaller inter class differences and intra-class similarities, which is more difficult to classify with existing object detectors. To address this problem, we propose two enhanced strategies on the current two-stage object detection algorithm. The first strategy uses attention-based group feature enhancement called group enhance module (GEM). By extending and grouping feature channels, the model can improve the ability to extract various discriminative features. The second strategy is to emphasize the sub-saliency feature learning, avoiding the network only focusing on the most significant part of the feature and ignoring the other parts. Our method is easy to implement and effective, and experiments show that our method can improve the Oriented regions with convolutional neural networks features (R-CNN) by about 1.45 mAP on the FAIR1M benchmark. Yong Zhou 0003, Sifan Wang, Jiaqi Zhao 0001, Hancheng Zhu, Rui Yao 0006 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Survey for person re-identification based on coarse-to-fine feature learning
Minjie Liu, Jiaqi Zhao 0001, Yong Zhou 0003, Hancheng Zhu, Rui Yao 0006, Ying Chen 0005 |
Multim. Tools Appl. | 4 |
| 2022 | Learning image aesthetic subjectivity from attribute-aware relational reasoning network
Hancheng Zhu, Yong Zhou 0003, Rui Yao 0006, Guangcheng Wang, Yuzhe Yang 0001 |
Pattern Recognit. Lett. | 1 |
| 2022 | Generalizable No-Reference Image Quality Assessment via Deep Meta-LearningabstractRecently, researchers have shown great interest in using convolutional neural networks (CNNs) for no-reference image quality assessment (NR-IQA). Due to the lack of big training data, the efforts of existing metrics in optimizing CNN-based NR-IQA models remain limited. Furthermore, the diversity of distortions in images result in the generalization problem of NR-IQA models when trained with known distortions and tested on unseen distortions, which is an easy task for human. Hence, we propose a NR-IQA metric via deep meta-learning, which is highly generalizable in the face of unseen distortions. The fundamental idea is to learn the meta-knowledge shared by human when evaluating the quality of images with diversified distortions. Specifically, we define NR-IQA of different distortions as a series of tasks and propose a task selection strategy to build two task sets, which are characterized by synthetic to synthetic and synthetic to authentic distortions, respectively. Based on these two task sets, an optimization-based meta-learning is proposed to learn the generalized NR-IQA model, which can be directly used to evaluate the quality of images with unseen distortions. Extensive experiments demonstrate that our NR-IQA metric outperforms the state-of-the-arts in terms of both evaluation performance and generalization ability. Hancheng Zhu, Leida Li, Jinjian Wu, Weisheng Dong, Guangming Shi |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Personalized Image Aesthetics Assessment via Meta-Learning With Bilevel Gradient OptimizationabstractTypical image aesthetics assessment (IAA) is modeled for the generic aesthetics perceived by an "average" user. However, such generic aesthetics models neglect the fact that users' aesthetic preferences vary significantly depending on their unique preferences. Therefore, it is essential to tackle the issue for personalized IAA (PIAA). Since PIAA is a typical small sample learning (SSL) problem, existing PIAA models are usually built by fine-tuning the well-established generic IAA (GIAA) models, which are regarded as prior knowledge. Nevertheless, this kind of prior knowledge based on "average aesthetics" fails to incarnate the aesthetic diversity of different people. In order to learn the shared prior knowledge when different people judge aesthetics, that is, learn how people judge image aesthetics, we propose a PIAA method based on meta-learning with bilevel gradient optimization (BLG-PIAA), which is trained using individual aesthetic data directly and generalizes to unknown users quickly. The proposed approach consists of two phases: 1) meta-training and 2) meta-testing. In meta-training, the aesthetics assessment of each user is regarded as a task, and the training set of each task is divided into two sets: 1) support set and 2) query set. Unlike traditional methods that train a GIAA model based on average aesthetics, we train an aesthetic meta-learner model by bilevel gradient updating from the support set to the query set using many users' PIAA tasks. In meta-testing, the aesthetic meta-learner model is fine-tuned using a small amount of aesthetic data of a target user to obtain the PIAA model. The experimental results show that the proposed method outperforms the state-of-the-art PIAA metrics, and the learned prior model of BLG-PIAA can be quickly adapted to unseen PIAA tasks. Hancheng Zhu, Leida Li, Jinjian Wu, Sicheng Zhao, Guiguang Ding, Guangming Shi |
IEEE Trans. Cybern. | 1 |
| 2022 | Facial action unit detection via hybrid relational reasoning
Zhiwen Shao, Yong Zhou 0003, Bing Liu 0016, Hancheng Zhu, Wen-Liang Du 0002, Jiaqi Zhao 0001 |
Vis. Comput. | 4 |
| 2020 | MetaIQA: Deep Meta-Learning for No-Reference Image Quality AssessmentabstractRecently, increasing interest has been drawn in exploiting deep convolutional neural networks (DCNNs) for no-reference image quality assessment (NR-IQA). Despite of the notable success achieved, there is a broad consensus that training DCNNs heavily relies on massive annotated data. Unfortunately, IQA is a typical small sample problem. Therefore, most of the existing DCNN-based IQA metrics operate based on pre-trained networks. However, these pre-trained networks are not designed for IQA task, leading to generalization problem when evaluating different types of distortions. With this motivation, this paper presents a no-reference IQA metric based on deep meta-learning. The underlying idea is to learn the meta-knowledge shared by human when evaluating the quality of images with various distortions, which can then be adapted to unknown distortions easily. Specifically, we first collect a number of NR-IQA tasks for different distortions. Then meta-learning is adopted to learn the prior knowledge shared by diversified distortions. Finally, the quality prior model is fine-tuned on a target NR-IQA task for quickly obtaining the quality model. Extensive experiments demonstrate that the proposed metric outperforms the state-of-the-arts by a large margin. Furthermore, the meta-model learned from synthetic distortions can also be easily generalized to authentic distortions, which is highly desired in real-world applications of IQA metrics. Hancheng Zhu, Leida Li, Jinjian Wu, Weisheng Dong, Guangming Shi |
CVPR | 1 |
| 2020 | Inferring Personality Traits from Attentive Regions of User Liked Images Via Weakly Supervised Dual Convolutional Network
Hancheng Zhu, Leida Li, Allen Tan |
Neural Process. Lett. | 1 |
| 2020 | Blind Quality Index of Depth Images Based on Structural Statistics for View SynthesisabstractThe quality of depth images is crucial for virtual view synthesis. However, the quality assessment of depth images is still largely unexplored. This letter presents a blind quality metric of Depth image based on Structural Statistics (DSS). The design philosophy is inspired by the fact that structural distortion in the depth images usually leads to geometric distortion, which is the main cause for degraded quality of synthesized views. Specifically, the statistical features for shape and orientation are calculated based on discrete orthogonal moments and gradients, generating two groups of quality-aware features. Then, the quality model is built from the extracted statistical features using a regression module. The experimental results demonstrate the effectiveness of the proposed metric. Yipo Huang, Leida Li, Hancheng Zhu, Bo Hu 0008 |
IEEE Signal Process. Lett. | 3 |
| 2020 | Personality-Assisted Multi-Task Learning for Generic and Personalized Image Aesthetics AssessmentabstractTraditional image aesthetics assessment (IAA) approaches mainly predict the average aesthetic score of an image. However, people tend to have different tastes on image aesthetics, which is mainly determined by their subjective preferences. As an important subjective trait, personality is believed to be a key factor in modeling individual's subjective preference. In this paper, we present a personality-assisted multi-task deep learning framework for both generic and personalized image aesthetics assessment. The proposed framework comprises two stages. In the first stage, a multi-task learning network with shared weights is proposed to predict the aesthetics distribution of an image and Big-Five (BF) personality traits of people who like the image. The generic aesthetics score of the image can be generated based on the predicted aesthetics distribution. In order to capture the common representation of generic image aesthetics and people's personality traits, a Siamese network is trained using aesthetics data and personality data jointly. In the second stage, based on the predicted personality traits and generic aesthetics of an image, an inter-task fusion is introduced to generate individual's personalized aesthetic scores on the image. The performance of the proposed method is evaluated using two public image aesthetics databases. The experimental results demonstrate that the proposed method outperforms the state-of-the-arts in both generic and personalized IAA tasks. Leida Li, Hancheng Zhu, Sicheng Zhao, Guiguang Ding, Weisi Lin |
IEEE Trans. Image Process. | 2 |
| 2019 | Personality Driven Multi-task Learning for Image Aesthetic AssessmentabstractWith the prevalence of convolutional neural networks (CNNs), assessing the aesthetics of an image has gained great advances recently. Individual users often have different aesthetic preferences on images, which we believe are mainly affected by their personality traits. However, most of the current aesthetics models predict a generic aesthetic score based on handcrafted and/or learned feature representations, which are unified and thus cannot reflect the individual differences during image aesthetic rating. In this paper, we propose an end-to-end personality driven multi-task deep learning model to address this problem. Firstly, both image aesthetics and personality traits are learned from the proposed multi-task model. Then the personality features are employed to modulate the aesthetics features, producing the optimal generic image aesthetics scores. The experimental results on two public databases show that the proposed method is superior to the state-of-the-art approaches. Leida Li, Hancheng Zhu, Sicheng Zhao, Guiguang Ding, Allen Tan |
ICME | 2 |
| 2019 | Naturalness Preserved Image Aesthetic Enhancement with Perceptual Encoder ConstraintabstractTypical supervised image enhancement pipeline is to minimize the distance between the enhanced image and the reference one. Pixel-wise and perceptual-wise loss functions could help to improve the general image quality, however are not very efficient in improving the image aesthetic quality. In this paper, we propose a novel Residual connected Dilated U-Net (RDU-Net) for improving the image aesthetic quality. By using different dilation rates, the RDU-Net can extract multiple receptive-field features and merge the maximum information from local to global, which are highly desired in image enhancement. Also, we propose an encoder constraint perceptual loss, which could teach the enhancement network to dig out the latent aesthetic factors and make the enhanced image more natural and aesthetically appealing. The proposed approach can alleviate the over-enhancement phenomenons. The experimental results show that the proposed perceptual loss function could give a steady back propagation and the proposed method outperforms the state-of-the-arts. Leida Li, Yuzhe Yang 0001, Hancheng Zhu |
ICMR | 3 |
| 2019 | No-reference quality assessment for contrast-distorted images based on multifaceted statistical representation of structure
Yu Zhou 0009, Leida Li, Hancheng Zhu, Hantao Liu, Shiqi Wang 0001, Yao Zhao 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2018 | Evaluating attributed personality traits from scene perception probability
Hancheng Zhu, Leida Li, Sicheng Zhao |
Pattern Recognit. Lett. | 1 |
| 2014 | Learning Structural Regularity for Evaluating Blocking Artifacts in JPEG ImagesabstractImage degradation damages genuine visual structures and causes pseudo structures. Pseudo structures are usually present with regularities. This letter proposes a machine learning based blocking artifacts metric for JPEG images by measuring the regularities of pseudo structures. Image corner, block boundary and color change properties are used to differentiate the blocking artifacts. A support vector regression (SVR) model is adopted to learn the underlying relations between these features and perceived blocking artifacts. The blocking artifacts score of a test image is predicted using the trained model. Extensive experiments demonstrate the effectiveness of the method. Leida Li, Weisi Lin, Hancheng Zhu |
IEEE Signal Process. Lett. | 3 |
| 2014 | Referenceless Measure of Blocking Artifacts by Tchebichef Kernel AnalysisabstractThis letter presents a Referenceless quality Measure of Blocking artifacts (RMB) using Tchebichef moments. It is based on the observation that Tchebichef kernels with different orders have varying abilities to capture blockiness. In a block manner, high-odd-order moments are computed to score the blocking artifacts. The blockiness scores are further weighted to incorporate the characteristic of Human Visual System (HVS), which is achieved by classifying the blocks into smooth and textured. Experimental results and comparisons demonstrate the advantage of the proposed method. Leida Li, Hancheng Zhu, Gaobo Yang, Jiansheng Qian |
IEEE Signal Process. Lett. | 2 |