Zhiwen Shao

dblp:180/2845 · DBLP profile ↗
← Back
64ranked-venue papers
23as first author
50since 2021 · last 2026
0000-0002-9383-8384ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 41 · 13 first-author · 32 since 2021Artificial intelligence and machine learning · 26 · 11 first-author · 21 since 2021Computer networks · 6 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 DTTNet: Improving Video Shadow Detection via Dark-Aware Guidance and Tokenized Temporal Modeling
abstract
Video shadow detection confronts two entwined difficulties: distinguishing shadows from complex backgrounds and modeling dynamic shadow deformations under varying illumination. To address shadow-background ambiguity, we leverage linguistic priors through the proposed Vision-language Match Module (VMM) and a Dark-aware Semantic Block (DSB), extracting text-guided features to explicitly differentiate shadows from dark objects. Furthermore, we introduce adaptive mask reweighting to downweight penumbra regions during training and apply edge masks at the final decoder stage for better supervision. For temporal modeling of variable shadow shapes, we propose a Tokenized Temporal Block (TTB) that decouples spatiotemporal learning. TTB summarizes cross-frame shadow semantics into learnable temporal tokens, enabling efficient sequence encoding with minimal computation overhead. Comprehensive Experiments on multiple benchmark datasets demonstrate state-of-the-art accuracy and real-time inference efficiency.
Kunyang Sun, Rui Yao 0006, Hancheng Zhu, Fuyuan Hu, Jiaqi Zhao 0001, Zhiwen Shao, Yong Zhou 0003
AAAI7
2026 Constrained and directional ensemble attention for facial action unit detection
Zhiwen Shao, Bikuan Chen, Yong Zhou 0003, Xuehuai Shi, Canlin Li, Lizhuang Ma, Dit-Yan Yeung
Pattern Recognit.1
2026 CurvNet: Latent contour representation and iterative data engine for curvature angle estimation
Zhiwen Shao, Lizhuang Ma, Xiaojia Zhu
Pattern Recognit.1
2026 Dynamic Prompt Memory Network for video shadow detection
Rui Yao 0006, Hancheng Zhu, Kunyang Sun, Jiaqi Zhao 0001, Zhiwen Shao, Abdulmotaleb El Saddik
Pattern Recognit.6
2026 TextRSR: Enhanced Arbitrary-Shaped Scene Text Representation via Robust Subspace Recovery
abstract
In recent years, scene text detection research has increasingly focused on arbitrary-shaped texts, where text representation is a fundamental problem. However, most existing methods still struggle to separate adjacent or overlapping texts due to ambiguous spatial positions of points or segmentation masks. Besides, the time efficiency of the entire pipeline is often neglected, resulting in sub-optimal inference speed. To tackle these problems, we first propose a novel text representation method based on robust subspace recovery, which robustly represents complex text shapes by combining orthogonal basis vectors learned from labeled text contours. These basis vectors capture basis contour patterns with distinct information, enabling clearer boundaries even in densely populated text scenarios. Moreover, we propose a dynamic sparse assignment scheme for positive samples that adaptively adjusts their weights during training, which not only accelerates inference speed by eliminating redundant predictions but also enhances feature learning by providing sufficient supervision signals. Building on these innovations, we present TextRSR, an accurate and efficient scene text detection network. Extensive experiments on challenging benchmarks demonstrate the superior accuracy and efficiency of TextRSR compared to state-of-the-art methods. Particularly, TextRSR achieves an F-measure of 88.5% at 37.8 frames per second (FPS) for CTW1500 dataset and an F-measure of 89.1% at 23.1 FPS for Total-Text dataset.
Zhiwen Shao, Shengtian Jiang, Hancheng Zhu, Xuehuai Shi, Canlin Li, Lizhuang Ma, Dit-Yan Yeung
IEEE Trans. Multim.1
2026 Spatio-Temporal Disentanglement and Constrained Self-Attention for Multi-Modal Deception Detection
abstract
Multi-modal deception detection is a challenging yet important task, having pivotal applications in many fields such as business credibility assessment and multimedia anti-frauds. Previous methods either rely solely on spatial features or overemphasize only temporal information within or across modalities, which may overlook potential critical clues. Motivated by these observations, we propose a Spatio-Temporal Representation Disentanglement (STRD) framework for multi-modal deception detection, which uses a dual-encoder structure to learn spatial and temporal representations for each modality. Specifically, we introduce a pre-trained foundation model to act as the spatial encoder and design a lightweight network as the temporal encoder, extracting spatial semantics and capturing dynamic temporal patterns. Then, we propose a Constrained Self-Attention Block (CSAB), in which self-attention distribution of each head is regarded as spatial distribution and is constrained to attend a certain facial local region. Furthermore, we present a Cross-Modal Correlation Fusion Block (CCFB) to achieve temporal synchronization across modalities by measuring the correlations between visual and audio features. Extensive experiments show that our STRD outperforms the state-of-the-art methods on challenging DOLOS, BOL, BgOL, and RLtrial benchmarks. Particularly, STRD improves by 2.12% and 1.88% over the previous best results in terms of ACC on the DOLOS and BOL datasets, respectively. Additionally, STRD outperforms previous methods in cross-dataset testing, highlighting its superior generalization ability.
Zhiwen Shao, Hancheng Zhu, Rui Yao 0006, Lixin Zou, Mengtian Li 0002, Bin Sheng 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2026 MOA: Efficient Scene-Aware Multi-Object Arrangement in VR
abstract
3D multi-object arrangement is a fundamental task in VR that relies on accurate and natural initial selection alongside rapid and convenient subsequent manipulation to ensure high efficiency. However, existing methods fail to support efficient multi-object arrangement in highly occluded scenes with densely packed candidate objects through controller-free natural interactions. In this article, we propose an efficient, scene-aware multi-object arrangement method (MOA) designed for fast, precise, and convenient object arrangement. First, MOA introduces an importance-driven multi-object initial selection algorithm that assigns higher spatiotemporally correlated object importance (IMP) to target objects, establishing a natural multi-object initial selection mode that enables quick and accurate selection of high-IMP objects. Subsequently, it presents an auxiliary-structure-guided multi-object manipulation algorithm that constructs an auxiliary manipulation structure to assist subsequent multi-object manipulation, alongside a multi-modal interaction mode that facilitates swift and natural manipulation. Compared to state-of-the-art controller-free and controller-based methods, MOA significantly improves task performance, reduces task load, and enhances convenience in complex multi-object arrangement scenes involving hundreds of highly occluded objects need to be arranged.
Xuehuai Shi, Yuhan Duan, Ziteng Wang 0002, Jian Wu 0033, Zhiwen Shao, Jieming Yin, Lili Wang 0006
IEEE Trans. Vis. Comput. Graph.5
2025 G-VEval: A Versatile Metric for Evaluating Image and Video Captions Using GPT-4o
abstract
Evaluation metric of visual captioning is important yet not thoroughly explored. Traditional metrics like BLEU, METEOR, CIDEr, and ROUGE often miss semantic depth, while trained metrics such as CLIP-Score, PAC-S, and Polos are limited in zero-shot scenarios. Advanced Language Model-based metrics also struggle with aligning to nuanced human preferences. To address these issues, we introduce G-VEval, a novel metric inspired by G-Eval and powered by the new GPT-4o. G-VEval uses chain-of-thought reasoning in large multimodal models and supports three modes: reference-free, reference-only, and combined, accommodating both video and image inputs. We also propose MSVD-Eval, a new dataset for video captioning evaluation, to establish a more transparent and consistent framework for both human experts and evaluation metrics. It is designed to address the lack of clear criteria in existing datasets by introducing distinct dimensions of Accuracy, Completeness, Conciseness, and Relevance (ACCR). Extensive results show that G-VEval outperforms existing methods in correlation with human annotations, as measured by Kendall tau-b and Kendall tau-c. This provides a flexible solution for diverse captioning tasks and suggests a straightforward yet effective approach for large language models to understand video content, paving the way for advancements in automated captioning.
Tony Cheng Tong, Zhiwen Shao, Dit-Yan Yeung
AAAI3
2025 Facial Action Unit Detection with Iterative Rank Reduction Adapter and Directional Attention
Zhiwen Shao, Hancheng Zhu, Rui Yao 0006, Bing Liu 0016
CGI (1)1
2025 Fast and Slow Streams for Online Time Series Forecasting Without Information Leakage
abstract
Current research in online time series forecasting (OTSF) faces two significant issues. The first is information leakage, where models make predictions and are then evaluated on historical time steps that have already been used in backpropagation for parameter updates. The second is practicality: while forecasting in real-world applications typically emphasizes looking ahead and anticipating future uncertainties, prediction sequences in this setting include only one future step with the remaining being observed time points. This necessitates a redefinition of the OTSF setting, focusing on predicting unknown future steps and evaluating unobserved data points. Following this new setting, challenges arise in leveraging incomplete pairs of ground truth and predictions for backpropagation, as well as in generalizing accurate information without overfitting to noise from recent data streams. To address these challenges, we propose a novel dual-stream framework for online forecasting (DSOF): a slow stream that updates with complete data using experience replay, and a fast stream that adapts to recent data through temporal difference learning. This dual-stream approach updates a teacher-student model learned through a residual learning strategy, generating predictions in a coarse-to-fine manner. Extensive experiments demonstrate its improvement in forecasting performance in changing environments. Our code is publicly available at https://github.com/yyalau/iclr2025_dsof.
Ying-yee Ava Lau, Zhiwen Shao, Dit-Yan Yeung
ICLR2
2025 SU-SAM: A Simple Unified Framework for Adapting SAM in Underperformed Scene
abstract
Segment Anything Model (SAM) excels in common vision tasks but struggles with specialized data. Recent methods fine-tune SAM using parameter-efficient techniques and task-specific designs, but they rely heavily on handcrafting and pre/post-processing, limiting the generalizability. In this paper, we propose SU-SAM, a simple and unified framework that adapts SAM efficiently without task-specific designs, improving its adaptability to underperforming scenes. SU-SAM abstracts parameter-efficient modules into basic design elements, offering four variants: series, parallel, mixed, and LoRA structures. Experiments across nine datasets and six tasks, including medical and defect segmentation, demonstrate SU-SAM’s superior performance. We analyze the effectiveness of different parameter-efficient designs and present a generalized model and benchmark, highlighting SU-SAM’s adaptability across diverse datasets.
Yiran Song, Qianyu Zhou 0001, Xuequan Lu, Zhiwen Shao, Lizhuang Ma
ICME4
2025 Modality-Guided Dynamic Graph Fusion and Temporal Diffusion for Self-Supervised RGB-T Tracking
abstract
To reduce the reliance on large-scale annotations, self-supervised RGB-T tracking approaches have garnered significant attention. However, the omission of the object region by erroneous pseudo-label or the introduction of background noise affects the efficiency of modality fusion, while pseudo-label noise triggered by similar object noise can further affect the tracking performance. In this paper, we propose GDSTrack, a novel approach that introduces dynamic graph fusion and temporal diffusion to address the above challenges in self-supervised RGB-T tracking. GDSTrack dynamically fuses the modalities of neighboring frames, treats them as distractor noise, and leverages the denoising capability of a generative model. Specifically, by constructing an adjacency matrix via an Adjacency Matrix Generator (AMG), the proposed Modality-guided Dynamic Graph Fusion (MDGF) module uses a dynamic adjacency matrix to guide graph attention, focusing on and fusing the object’s coherent regions. Temporal Graph-Informed Diffusion (TGID) models MDGF features from neighboring frames as interference, and thus improving robustness against similar-object noise. Extensive experiments conducted on four public RGB-T tracking datasets demonstrate that GDSTrack outperforms the existing state-of-the-art methods. The source code is available at https://github.com/LiShenglana/GDSTrack.
Shenglan Li, Rui Yao 0006, Yong Zhou 0003, Hancheng Zhu, Kunyang Sun, Bing Liu 0016, Zhiwen Shao, Jiaqi Zhao 0001
IJCAI7
2025 Symmetric perception and ordinal regression for detecting scoliosis natural image
Xiaojia Zhu, Zhiwen Shao, Yuhu Dai, Chuandong Lang
Appl. Intell.4
2025 Facial Action Unit Detection by Adaptively Constraining Self-Attention and Causally Deconfounding Sample
Zhiwen Shao, Hancheng Zhu, Yong Zhou 0003, Xiang Xiang 0001, Bing Liu 0016, Rui Yao 0006, Lizhuang Ma
Int. J. Comput. Vis.1
2025 MOL: Joint Estimation of Micro-Expression, Optical Flow, and Landmark via Transformer-Graph-Style Convolution
abstract
Facial micro-expression recognition (MER) is a challenging problem, due to transient and subtle micro-expression (ME) actions. Most existing methods depend on hand-crafted features, key frames like onset, apex, and offset frames, or deep networks limited by small-scale and low-diversity datasets. In this paper, we propose an end-to-end micro-action-aware deep learning framework with advantages from transformer, graph convolution, and vanilla convolution. In particular, we propose a novel F5C block composed of fully-connected convolution and channel correspondence convolution to directly extract local-global features from a sequence of raw frames, without the prior knowledge of key frames. The transformer-style fully-connected convolution is proposed to extract local features while maintaining global receptive fields, and the graph-style channel correspondence convolution is introduced to model the correlations among feature patterns. Moreover, MER, optical flow estimation, and facial landmark detection are jointly trained by sharing the local-global features. The two latter tasks contribute to capturing facial subtle action information for MER, which can alleviate the impact of insufficient training data. Extensive experiments demonstrate that our framework (i) outperforms the state-of-the-art MER methods on CASME II, SAMM, and SMIC benchmarks, (ii) works well for optical flow estimation and facial landmark detection, and (iii) can capture facial subtle muscle actions in local regions associated with MEs.
Zhiwen Shao, Feiran Li, Yong Zhou 0003, Xuequan Lu, Yuan Xie 0006, Lizhuang Ma
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Mirror Detection via Multi-Directional Similarity Perception and Spectral Saliency Enhancement
abstract
Mirror detection is a challenging task, due to the reflective properties of mirrors. Most existing approaches rely on exploiting the relationship between the content inside the mirror and the surrounding environment to aid in locating mirrors. A typical solution is to utilize contextual contrasted features. However, the discontinuity in content at the edges of mirrors may not always be prominent. To overcome this limitation, we propose a novel mirror detection framework called S2MD including two main modules, multi-directional similarity perception module (MSPM) and spectral saliency enhancement decoder module (SSEDM). Specifically, we employ a backbone network to extract multi-scale global information from images using a dual-path approach. Then, we feed these high-level dual-path features into MSPMs to generate direction-sensitive similarity-consistent features. MSPM utilizes active rotating filters and oriented response pooling to model the similarity relations in different orientations. Moreover, the SSEDM is utilized to enhance the spatial contextual contrasted features using feature spectral residuals and fuse the dual-path features to obtain the final predicted mirror mask. Extensive experiments demonstrate that our method achieves state-of-the-art performance on challenging MSD, PMD, and RGBD-Mirror benchmarks. The code is available at https://github.com/RuiChen-stack/M2SD.
Zhiwen Shao, Xuehuai Shi, Bing Liu 0016, Canlin Li, Lizhuang Ma, Dit-Yan Yeung
IEEE Trans. Circuits Syst. Video Technol.1
2025 Progressively Generated Text-Assisted Image Aesthetic Quality Assessment
abstract
Image Aesthetic Quality Assessment (IAQA) aims to simulate user perceptions to judge the aesthetic quality of images. Due to the high subjectivity of users and the complexity of image aesthetics, modeling IAQA solely at the image level is a compromise. Consequently, existing methods mainly focus on multimodal-based models and achieve effective performance. These methods explore aesthetic comments on images to characterize users and serve as auxiliary text information for multimodal modeling. Unfortunately, this may suffer from two limitations. One limitation is that aesthetic comments are often unavailable for an unknown image in the test phase, and another limitation is that the semantic information of these comments may be uncertain and fuzzy. Therefore, this paper proposes a progressively generated text-assisted image aesthetic quality assessment method, aiming to address the lack of aesthetic comments and the fuzziness of aesthetic judgments in these comments. Specifically, we first adopt a Multimodal Large Language Model (MLLM) to generate aesthetic comments on images by simulating user perceptions and utilize the generated comments to characterize their aesthetic perception to assist in the pre-training of our multimodal-based IAQA model. Then, we design an attribute prediction module to determine the attribute levels of aesthetic judgments and utilize text template construction to further generate explicit descriptions of image aesthetics. Finally, we leverage the generated attribute descriptions to further assist in training our IAQA model. By progressively generating textual auxiliary descriptions of aesthetics for images, the proposed model can gradually determine the aesthetic quality of the images. Massive experimental results indicate that the proposed method outperforms existing mainstream methods on multiple IAQA datasets.
Hancheng Zhu, Ju Shi, Zhiwen Shao, Rui Yao 0006, Kunyang Sun, Leida Li
IEEE Trans. Fuzzy Syst.3
2025 Hyperspectral Object Tracking With Dual-Stream Prompt
abstract
Hyperspectral images, rich in spectral details, offeradvantages for object tracking across diverse scenarios. Current hyperspectral tracking often fine-tunes parameters using pretrained RGB trackers, but this manner is suboptimal due to redundancy in spectral bands and limited training data. Existing hyperspectral trackers also underuse temporal information. To address these issues, we propose a unified spectral-spatiotemporal multimodal dual-stream prompt hyperspectral object tracking, named HDSP. We design a density clustering-based band selection module (BSM) to preserve spectral prompt information efficiently. Using the generated bands and temporal data as multimodal prompts, a dual-stream visual prompter is proposed. Designed multimodal dual-stream visual prompter (MDVP) transforms the multimodal input into a single modality, enhancing the foundational modality’s representation capabilities for hyperspectral tracking. Experiments on hyperspectral videos (HSVs) tracking datasets demonstrate that the proposed tracker achieves state-of-the-art performance. The source code is available athttps://github.com/rayyao/HDSP.
Rui Yao 0006, Yong Zhou 0003, Hancheng Zhu, Jiaqi Zhao 0001, Zhiwen Shao
IEEE Trans. Geosci. Remote. Sens.6
2025 Micro-Expression Recognition via Fine-Grained Dynamic Perception
abstract
Facial micro-expression recognition (MER) is a challenging task, due to the transience, subtlety, and dynamics of micro-expressions (MEs). Most existing methods resort to hand-crafted features or deep networks, in which the former often additionally requires key frames, and the latter suffers from small-scale and low-diversity training data. In this article, we develop a novel fine-grained dynamic perception (FDP) framework for MER. We propose to rank frame-level features of a sequence of raw frames in chronological order, in which the rank process encodes the dynamic information of both ME appearances and motions. Specifically, a novel local-global feature-aware transformer is proposed for frame representation learning. A rank scorer is further adopted to calculate rank scores of each frame-level feature. Afterwards, the rank features from rank scorer are pooled in temporal dimension to capture dynamic representation. Finally, the dynamic representation is shared by a MER module and a dynamic image construction module, in which the former predicts the ME category, and the latter uses an encoder-decoder structure to construct the dynamic image. The design of dynamic image construction task is beneficial for capturing facial subtle actions associated with MEs and alleviating the data scarcity issue. Extensive experiments show that our method (i) significantly outperforms the state-of-the-art MER methods, and (ii) works well for dynamic image construction. Particularly, our FDP improves by 4.05%, 2.50%, 7.71%, and 2.11% over the previous best results in terms of F1-score on the CASME II, SAMM, CAS(ME) 2 , and CAS(ME) 3 datasets, respectively. The code is available at https://github.com/CYF-cuber/FDP .
Zhiwen Shao, Xuehuai Shi, Canlin Li, Lizhuang Ma, Dit-Yan Yeung
ACM Trans. Multim. Comput. Commun. Appl.1
2025 Historical Object-Aware Prompt Learning for Universal Hyperspectral Object Tracking
abstract
Hyperspectral Object Tracking (HOT), utilizing rich spectral information from hyperspectral video (HSV), holds significant importance for object tracking. We identify that a major obstacle in improving HOT performance lies in effectively leveraging spectral and historical information. Furthermore, due to the mismatch in band dimensions between hyperspectral and RGB images, state-of-the-art RGB-based trackers struggle to adapt to unified HOT tasks. To address this, we propose a Historical Object-Aware Prompt Learning (HOPL) method for universal hyperspectral object tracking. Initially, we transform hyperspectral image ( \( N \) bands) into multiple sets of three bands with different combinations and feed them into a backbone network to generate base features. Subsequently, we introduce a historical object-aware prompter, where historical object-aware images are input to generate prompt features that enhance the representation of object information when combined with base features. Additionally, we design a band information fusion module to integrate the multiple sets of base features. By introducing historical object-aware prompts, HOPL significantly enhances tracking performance without retraining the backbone network. Experimental results on the HOT2023 dataset (comprising HSV with 25-band, 16-band, and 15-band wavelength ranges) and HOT2022 dataset validate the superiority of HOPL over state-of-the-art methods. The source code is available at https://github.com/rayyao/HOPL .
Rui Yao 0006, Yong Zhou 0003, Fuyuan Hu, Jiaqi Zhao 0001, Zhiwen Shao
ACM Trans. Multim. Comput. Commun. Appl.7
2025 Image Cropping with Content and Composition Attribute-aware Global Relation Reasoning
abstract
Image cropping aims to find visually pleasing content in an image, which will enhance its aesthetic quality. Existing image cropping approaches mainly emphasize the geometric properties of images, such as composition and layout, neglecting the rich aesthetic information available from the physical attributes (e.g., content and themes), and background information beyond the foreground in images. Consequently, this article proposes an image cropping method based on the content and composition attribute-aware global relation reasoning, which aims at guiding the generation of cropped sub-images by exploring critical attributes based on content and composition as well as global object correlations that affect aesthetics in images. Particularly, to comprehensively introduce aesthetic information into image cropping, we capture feature representations reinforced by content and composition attributes simultaneously. The feature representations can strengthen the visual aesthetics of cropped sub-images. To make the cropped sub-images amply contain more global information, we introduce a global relation reasoning branch in the proposed cropping module, which can fully exploit the dependency relationship between the foreground and background in images. Extensive experiments on image cropping benchmarks demonstrate that our approach is superior to state-of-the-art image cropping methods.
Hancheng Zhu, Yong Zhou 0003, Rui Yao 0006, Zhiwen Shao, Jiaqi Zhao 0001, Leida Li
ACM Trans. Multim. Comput. Commun. Appl.5
2025 Scene-Aware Foveated Neural Radiance Fields
abstract
Foveated rendering provides an idea for improving the image synthesis performance of neural radiance fields (NeRF) methods. In this article, we propose a scene-aware foveated neural radiance fields method to synthesize high-quality foveated images in complex VR scenes at high frame rates. First, we construct a multi-ellipsoidal neural representation to enhance the neural radiance field's representation capability in salient regions of complex VR scenes based on the scene content. Then, we introduce a uniform sampling based foveated neural radiance field framework to improve the foveated image synthesis performance with one-pass color inference, and improve the synthesis quality by leveraging the foveated scene-aware objective function. Our method synthesizes high-quality binocular foveated images at the average frame rate of 66 frames per second ($FPS$FPS) in complex scenes with high occlusion, intricate textures, and sophisticated geometries. Compared with the state-of-the-art foveated NeRF method, our method achieves significantly higher synthesis quality in both the foveal and peripheral regions with 1.41-1.46× speedup. We also conduct a user study to prove that the perceived quality of our method has a high visual similarity with the ground truth.
Xuehuai Shi, Lili Wang 0006, Xinda Liu, Jian Wu 0033, Zhiwen Shao
IEEE Trans. Vis. Comput. Graph.5
2025 High-level LoRA and hierarchical fusion for enhanced micro-expression recognition
Zhiwen Shao, Yong Zhou 0003, Xiang Xiang 0001, Jian Li 0054, Bing Liu 0016, Dit-Yan Yeung
Vis. Comput.1
2025 Self-similarity guided regression with contrast enhancement for spine segmentation
abstract
Accurate spine segmentation is critical for scoliosis diagnosis and treatment. For instance, automatic Cobb angle measurement for scoliosis relies on precisely localized vertebral masks. However, it remains a challenging task due to low tissue contrast, blurred vertebral edges, and overlapping anatomical structures. In this paper, we propose SRNet, a pure segmentation network that produces binary masks of each vertebra. SRNet integrates two novel components, a Self-similarity Guided Dynamic Convolution (SGDC) module and a Contrast-Enhanced Boundary Decoder (CEBD). SGDC exploits the repetitive structure of vertebrae by leveraging non-local attention to compute self-similarity across feature maps and dynamic convolution to combine multiple convolution kernels adaptively. CEBD sharpens segmentation boundaries via a reverse-attention mechanism that erases the coarse prediction and focuses on missing edge details, combined with a spectral-residual filter that amplifies high-frequency edge information. Extensive experiments on the AASCE spine X-ray dataset show that our SRNet achieves a high Dice score of 92.37%, outperforming state-of-the-art approaches. While our primary focus here is mask segmentation, the accurate vertebral masks produced by SRNet could readily support future tasks such as scoliosis Cobb angle estimation.
Xiaojia Zhu, Zhiwen Shao
Vis. Informatics4
2024 LRANet: Towards Accurate and Efficient Scene Text Detection with Low-Rank Approximation Network
abstract
Recently, regression-based methods, which predict parameterized text shapes for text localization, have gained popularity in scene text detection. However, the existing parameterized text shape methods still have limitations in modeling arbitrary-shaped texts due to ignoring the utilization of text-specific shape information. Moreover, the time consumption of the entire pipeline has been largely overlooked, leading to a suboptimal overall inference speed. To address these issues, we first propose a novel parameterized text shape method based on low-rank approximation. Unlike other shape representation methods that employ data-irrelevant parameterization, our approach utilizes singular value decomposition and reconstructs the text shape using a few eigenvectors learned from labeled text contours. By exploring the shape correlation among different text contours, our method achieves consistency, compactness, simplicity, and robustness in shape representation. Next, we propose a dual assignment scheme for speed acceleration. It adopts a sparse assignment branch to accelerate the inference speed, and meanwhile, provides ample supervised signals for training through a dense assignment branch. Building upon these designs, we implement an accurate and efficient arbitrary-shaped text detector named LRANet. Extensive experiments are conducted on several challenging benchmarks, demonstrating the superior accuracy and efficiency of LRANet compared to state-of-the-art methods. Code is available at: https://github.com/ychensu/LRANet.git
Zhineng Chen, Zhiwen Shao, Yuning Du, Zhilong Ji, Jinfeng Bai, Yong Zhou 0003, Yu-Gang Jiang 0001
AAAI3
2024 Attribute-Driven Multimodal Hierarchical Prompts for Image Aesthetic Quality Assessment
abstract
Image Aesthetic Quality Assessment (IAQA) aims to simulate users' visual perception to judge the aesthetic quality of images. In social media, users' aesthetic experiences are often reflected in their textual comments regarding the aesthetic attributes of images. To fully explore the attribute information perceived by users for evaluating image aesthetic quality, this paper proposes an image aesthetic quality assessment method based on attribute-driven multimodal hierarchical prompts. Unlike existing IAQA methods that utilize multimodal pre-training or straightforward prompts for model learning, the proposed method leverages attribute comments and quality-level text templates to hierarchically learn the aesthetic attributes and quality of images. Specifically, we first leverage users' aesthetic attribute comments to perform prompt learning on images. The learned attribute-driven multimodal features can comprehensively capture the semantic information of image aesthetic attributes perceived by users. Then, we construct text templates for different aesthetic quality levels to further facilitate prompt learning through semantic information related to the aesthetic quality of images. The proposed method can explicitly simulate users' aesthetic judgment of images to obtain more precise aesthetic quality. Experimental results demonstrate that the proposed IAQA method based on hierarchical prompts outperforms existing methods significantly on multiple IAQA databases. Our source code is public at https://github.com/GitHub-Ju/AMHP.
Hancheng Zhu, Ju Shi, Zhiwen Shao, Rui Yao 0006, Yong Zhou 0003, Jiaqi Zhao 0001, Leida Li
ACM Multimedia3
2024 Joint facial action unit recognition and self-supervised optical flow estimation
Zhiwen Shao, Yong Zhou 0003, Feiran Li, Hancheng Zhu, Bing Liu 0016
Pattern Recognit. Lett.1
2024 CT-Net: Arbitrary-Shaped Text Detection via Contour Transformer
abstract
Contour based scene text detection methods have rapidly developed recently, but still suffer from inaccurate front-end contour initialization, multi-stage error accumulation, or deficient local information aggregation. To tackle these limitations, we propose a novel arbitrary-shaped scene text detection framework named CT-Net by progressive contour regression with contour transformers. Specifically, we first employ a contour initialization module that generates coarse text contours without any post-processing. Then, we adopt contour refinement modules to adaptively refine text contours in an iterative manner, which are beneficial for context information capturing and progressive global contour deformation. Besides, we propose an adaptive training strategy to enable the contour transformers to learn more potential deformation paths, and introduce a re-score mechanism that can effectively suppress false positives. Extensive experiments are conducted on four challenging datasets, which demonstrate the accuracy and efficiency of our CT-Net over state-of-the-art methods. Particularly, CT-Net achieves F-measure of 86.1 at 11.2 frames per second (FPS) and F-measure of 87.8 at 10.1 FPS for CTW1500 and Total-Text datasets, respectively.
Zhiwen Shao, Yong Zhou 0003, Hancheng Zhu, Bing Liu 0016, Rui Yao 0006
IEEE Trans. Circuits Syst. Video Technol.1
2024 Motion-Aware Self-Supervised RGBT Tracking with Multi-Modality Hierarchical Transformers
abstract
Supervised RGBT (SRGBT) tracking tasks need both expensive and time-consuming annotations. Therefore, the implementation of Self-Supervised RGBT (SSRGBT) tracking methods has become increasingly important. Straightforward SSRGBT tracking methods use pseudo-labels for tracking, but inaccurate pseudo-labels can lead to object drift, which severely affects tracking performance. This article proposes a self-supervised RGBT object tracking method (S2OTFormer) to bridge the gap between tracking methods supervised under pseudo-labels and ground truth labels. Firstly, to provide more robust appearance features for motion cues, we introduce a multi-modality hierarchical transformer (MHT) module for feature fusion. This module allocates weights to both modalities and strengthens the expressive capability of the MHT module through multiple nonlinear layers to fully utilize the complementary information of the two modalities. Secondly, in order to solve the problems of motion blur caused by camera motion and inaccurate appearance information caused by pseudo-labels, we introduce a motion-aware mechanism (MAM). The MAM extracts the average motion vectors from the previous multi-frame search frame features and constructs the consistency loss with the motion vectors of the current search frame features. The motion vectors of inter-frame objects are obtained by reusing the inter-frame attention map to predict coordinate positions. Finally, to further reduce the effect of inaccurate pseudo-labels, we propose an Attention-Based Multi-Scale Enhancement Module. By introducing cross-attention to achieve more precise and accurate object tracking, this module overcomes the receptive field limitations of traditional CNN tracking heads. We demonstrate the effectiveness of S2OTFormer on four large-scale public datasets through extensive comparisons as well as numerous ablation experiments. The source code is available at https://github.com/LiShenglana/S2OTFormer .
Shenglan Li, Rui Yao 0006, Yong Zhou 0003, Hancheng Zhu, Jiaqi Zhao 0001, Zhiwen Shao, Abdulmotaleb El Saddik
ACM Trans. Multim. Comput. Commun. Appl.6
2024 Diverse Image Captioning via Conditional Variational Autoencoder and Dual Contrastive Learning
abstract
Diverse image captioning has achieved substantial progress in recent years. However, the discriminability of generative models and the limitation of cross entropy loss are generally overlooked in the traditional diverse image captioning models, which seriously hurts both the diversity and accuracy of image captioning. In this article, aiming to improve diversity and accuracy simultaneously, we propose a novel Conditional Variational Autoencoder (DCL-CVAE) framework for diverse image captioning by seamlessly integrating sequential variational autoencoder with contrastive learning. In the encoding stage, we first build conditional variational autoencoders to separately learn the sequential latent spaces for a pair of captions. Then, we introduce contrastive learning in the sequential latent spaces to enhance the discriminability of latent representations for both image-caption pairs and mismatched pairs. In the decoding stage, we leverage the captions sampled from the pre-trained Long Short-Term Memory (LSTM), LSTM decoder as the negative examples and perform contrastive learning with the greedily sampled positive examples, which can restrain the generation of common words and phrases induced by the cross entropy loss. By virtue of dual constrastive learning, DCL-CVAE is capable of encouraging the discriminability and facilitating the diversity, while promoting the accuracy of the generated captions. Extensive experiments are conducted on the challenging MSCOCO dataset, showing that our proposed methods can achieve a better balance between accuracy and diversity compared to the state-of-the-art diverse image captioning models.
Bing Liu 0016, Yong Zhou 0003, Rui Yao 0006, Zhiwen Shao
ACM Trans. Multim. Comput. Commun. Appl.6
2024 Boundary-aware small object detection with attention and interaction
Qihan Feng, Zhiwen Shao
Vis. Comput.2
2023 IterativePFN: True Iterative Point Cloud Filtering
abstract
The quality of point clouds is often limited by noise introduced during their capture process. Consequently, a fundamental 3D vision task is the removal of noise, known as point cloud filtering or denoising. State-of-the-art learning based methods focus on training neural networks to infer filtered displacements and directly shift noisy points onto the underlying clean surfaces. In high noise conditions, they iterate the filtering process. However, this iterative filtering is only done at test time and is less effective at ensuring points converge quickly onto the clean surfaces. We propose IterativePFN (iterative point cloud filtering network), which consists of multiple IterationModules that model the true iterative filtering process internally, within a single network. We train our IterativePFn network using a novel loss function that utilizes an adaptive ground truth target at each iteration to capture the relationship between intermediate filtering results during training. This ensures that the filtered results converge faster to the clean surfaces. Our method is able to obtain better performance compared to state-of-the-art methods. The source code can be found at: https://github.com/ddsediri/IterativePFN
Dasith de Silva Edirimuni, Xuequan Lu, Zhiwen Shao, Gang Li 0009, Antonio Robles-Kelly, Ying He 0001
CVPR3
2023 Personalized Image Aesthetics Assessment with Attribute-guided Fine-grained Feature Representation
abstract
Personalized image aesthetics assessment (PIAA) has gained increasing attention from researchers due to its ability to measure individual users' specific aesthetic experiences. However, most existing PIAA methods rely on holistic features or simplistic coding to characterize users' aesthetic preferences for images, and we believe that more rich explicit features are needed in modeling PIAA. Consequently, we propose an attribute-guided fine-grained feature-aware personalized image aesthetics assessment method, which can fully capture fine-grained features from multiple attributes to represent users' aesthetic preferences for images. To achieve this, we first build a fine-grained feature extraction (FFE) module to obtain the refined local features of image attributes to compensate for holistic features. The FFE module is then used to generate user-level features, which are combined with the image-level features to obtain user-preferred fine-grained feature representations. By training extensive users' PIAA tasks, the aesthetic distribution of most users can be transferred to the personalized scores of individual users. To enable our proposed model to learn more generalizable aesthetics among individual users, we incorporate the degree of dispersion between users' personalized scores and image aesthetic distribution as a coefficient in the loss function during model training. Experimental results on several PIAA databases show that our method outperforms existing mainstream PIAA methods, and can effectively infer users' personalized aesthetics of images.
Hancheng Zhu, Zhiwen Shao, Yong Zhou 0003, Guangcheng Wang, Pengfei Chen 0003, Leida Li
ACM Multimedia2
2023 Identity-invariant representation and transformer-style relation for micro-expression recognition
Zhiwen Shao, Feiran Li, Yong Zhou 0003, Hancheng Zhu, Rui Yao 0006
Appl. Intell.1
2023 Unsupervised RGB-T object tracking with attentional multi-modal feature fusion
Shenglan Li, Rui Yao 0006, Yong Zhou 0003, Hancheng Zhu, Bing Liu 0016, Jiaqi Zhao 0001, Zhiwen Shao
Multim. Tools Appl.7
2023 Semi-supervised transformable architecture search for feature distillation
Man Zhang 0006, Yong Zhou 0003, Bing Liu 0016, Jiaqi Zhao 0001, Rui Yao 0006, Zhiwen Shao, Hancheng Zhu
Pattern Anal. Appl.6
2023 Facial Action Unit Detection via Adaptive Attention and Relation
abstract
Facial action unit (AU) detection is challenging due to the difficulty in capturing correlated information from subtle and dynamic AUs. Existing methods often resort to the localization of correlated regions of AUs, in which predefining local AU attentions by correlated facial landmarks often discards essential parts, or learning global attention maps often contains irrelevant areas. Furthermore, existing relational reasoning methods often employ common patterns for all AUs while ignoring the specific way of each AU. To tackle these limitations, we propose a novel adaptive attention and relation (AAR) framework for facial AU detection. Specifically, we propose an adaptive attention regression network to regress the global attention map of each AU under the constraint of attention predefinition and the guidance of AU detection, which is beneficial for capturing both specified dependencies by landmarks in strongly correlated regions and facial globally distributed dependencies in weakly correlated regions. Moreover, considering the diversity and dynamics of AUs, we propose an adaptive spatio-temporal graph convolutional network to simultaneously reason the independent pattern of each AU, the inter-dependencies among AUs, as well as the temporal dependencies. Extensive experiments show that our approach (i) achieves competitive performance on challenging benchmarks including BP4D, DISFA, and GFT in constrained scenarios and Aff-Wild2 in unconstrained scenarios, and (ii) can precisely learn the regional correlation distribution of each AU.
Zhiwen Shao, Yong Zhou 0003, Jianfei Cai 0001, Hancheng Zhu, Rui Yao 0006
IEEE Trans. Image Process.1
2023 Attention-guided Adversarial Attack for Video Object Segmentation
abstract
Video Object Segmentation (VOS) methods have made many breakthroughs with the help of the continuous development and advancement of deep learning. However, the deep learning model is vulnerable to malicious adversarial attacks, which mislead the model to make wrong decisions by adding adversarial perturbation that humans cannot perceive to the input image. Threats to deep learning models remind us that video object segmentation methods are also vulnerable to attacks, thereby threatening their security. Therefore, we study adversarial attacks on the VOS task to better identify the vulnerabilities of the VOS method, which in turn provides an opportunity to improve its robustness. In this paper, we propose an attention-guided adversarial attack method, which uses spatial attention blocks to capture features with global dependencies to construct correlations between consecutive video frames, and performs multipath aggregation to effectively integrate spatial-temporal perturbation, thereby guiding the deconvolution network to generate adversarial examples with strong attack capability. Specifically, the class loss function is designed to enable the deconvolution network to better activate noise in other regions and suppress the activation related to the object class based on the enhanced feature map of the object class. At the same time, attentional feature loss is designed to enhance the transferability against attack. The experimental results on the DAVIS dataset show that the proposed attention-guided adversarial attack method can significantly reduce the segmentation accuracy of OSVOS, and the J & F mean on DAVIS 2016 can reach 73.6% drop rate. The generated adversarial examples are also highly transferable to other video object segmentation models.
Rui Yao 0006, Ying Chen 0005, Yong Zhou 0003, Fuyuan Hu, Jiaqi Zhao 0001, Bing Liu 0016, Zhiwen Shao
ACM Trans. Intell. Syst. Technol.7
2023 TextDCT: Arbitrary-Shaped Text Detection via Discrete Cosine Transform Mask
abstract
Arbitrary-shaped scene text detection is a challenging task due to the variety of text changes in font, size, color, and orientation. Most existing regression based methods resort to regress the masks or contour points of text regions to model the text instances. However, regressing the complete masks requires high training complexity, and contour points are not sufficient to capture the details of highly curved texts. To tackle the above limitations, we propose a novel light-weight anchor-free text detection framework called TextDCT, which adopts the discrete cosine transform (DCT) to encode the text masks as compact vectors. Further, considering the imbalanced number of training samples among pyramid layers, we only employ a single-level head for top-down prediction. To model the multi-scale texts in a single-level head, we introduce a novel positive sampling strategy by treating the shrunk text region as positive samples, and design a feature awareness module (FAM) for spatial-awareness and scale-awareness by fusing rich contextual information and focusing on more significant features. Moreover, we propose a segmented non-maximum suppression (S-NMS) method that can filter low-quality mask regressions. Extensive experiments are conducted on four challenging datasets, which demonstrate our TextDCT obtains competitive performance on both accuracy and efficiency. Specifically, TextDCT achieves F-measure of 85.1 at 17.2 frames per second (FPS) and F-measure of 84.9 at 15.1 FPS for CTW1500 and Total-Text datasets, respectively.
Zhiwen Shao, Yong Zhou 0003, Hancheng Zhu, Bing Liu 0016, Rui Yao 0006
IEEE Trans. Multim.2
2023 Weakly Supervised Few-Shot Semantic Segmentation via Pseudo Mask Enhancement and Meta Learning
abstract
Few shot semantic segmentation has been proposed to enhance the generalization ability of traditional models with limited data. Previous works mainly focus on the supervised tasks, while limited amount of work is explored for the weakly supervised tasks. Weakly supervised semantic segmentation has become an active research area because weakly supervised labels effectively reduce the annotation cost of visual tasks. To this end, we propose a weakly supervised few-shot semantic segmentation model based on the meta learning framework, which utilizes prior knowledge and adjusts itself according to new tasks. Thereupon then, the proposed network is capable of both high efficiency and generalization ability to new tasks. In the pseudo mask generation stage, we develop a WRCAM method with the channel-spatial attention mechanism to refine the coverage size of targets in pseudo masks. In the few-shot semantic segmentation stage, the optimization based meta learning method is used to realize few-shot semantic segmentation by virtue of the refined pseudo masks. The experimental results show that the proposed method not only significantly outperforms weakly supervised SOTA methods, but also could be comparative to some supervised SOTA methods.
Man Zhang 0006, Yong Zhou 0003, Bing Liu 0016, Jiaqi Zhao 0001, Rui Yao 0006, Zhiwen Shao, Hancheng Zhu
IEEE Trans. Multim.6
2022 Show, Deconfound and Tell: Image Captioning with Causal Inference
abstract
The transformer-based encoder-decoder framework has shown remarkable performance in image captioning. However, most transformer-based captioning methods ever overlook two kinds of elusive confounders: the visual confounder and the linguistic confounder, which generally lead to harmful bias, induce the spurious correlations during training, and degrade the model generalization. In this paper, we first use Structural Causal Models (SCMs) to show how two confounders damage the image captioning. Then we apply the backdoor adjustment to propose a novel causal inference based image captioning (CIIC) framework, which consists of an interventional object detector (IOD) and an interventional transformer decoder (ITD) to jointly confront both confounders. In the encoding stage, the IOD is able to disentangle the region-based visual features by deconfounding the visual confounder. In the decoding stage, the ITD introduces causal intervention into the transformer decoder and deconfounds the visual and linguistic confounders simultaneously. Two modules collaborate with each other to alleviate the spurious correlations caused by the unobserved confounders. When tested on MSCOCO, our proposal significantly outperforms the state-of-the-art encoder-decoder models on Karpathy split and online test split. Code is published in https://github.com/CUMTGG/CIIC.
Bing Liu 0016, Xu Yang 0021, Yong Zhou 0003, Rui Yao 0006, Zhiwen Shao, Jiaqi Zhao 0001
CVPR6
2022 Personality modeling from image aesthetic attribute-aware graph representation learning
Hancheng Zhu, Yong Zhou 0003, Qiaoyue Li, Zhiwen Shao
J. Vis. Commun. Image Represent.4
2022 A Semi-Supervised Image-to-Image Translation Framework for SAR-Optical Image Matching
abstract
Synthetic Aperture Radar (SAR) and optical image matching aims to acquire correspondences from a certain pair of SAR and optical images. Recent advances in the image-to-image translation provided a way to simplify the SAR-optical image matching into the SAR-SAR or optical-optical image matchings. Existing image-to-image translations mainly focus on supervised or unsupervised learning. However, gathering sufficient amounts of aligned training data for supervised learning is challenging, while unsupervised learning cannot guarantee enough correct correspondences. In this work, we investigate the applicability of semi-supervised image-to-image translation for SAR-optical image matching such that both aligned and unaligned SAR-optical images could be used. To this end, we combine the benefits of both supervised and unsupervised well-known image-to-image translation methods, i.e., Pix2pix and CycleGAN, and propose a simple yet effective semi-supervised image-to-image translation framework. Through extensive experimental comparisons to baseline methods, we verify the effectiveness of the proposed framework in both semi-supervised and fully-supervised settings. Our codes are available at https://github.com/WenliangDu/Semi-I2I.
Wen-Liang Du 0002, Yong Zhou 0003, Hancheng Zhu, Jiaqi Zhao 0001, Zhiwen Shao, Xiaolin Tian 0001
IEEE Geosci. Remote. Sens. Lett.5
2022 GeoConv: Geodesic guided convolution for facial action unit recognition
Yuedong Chen, Guoxian Song, Zhiwen Shao, Jianfei Cai 0001, Tat-Jen Cham, Jianmin Zheng
Pattern Recognit.3
2022 Unconstrained Facial Action Unit Detection via Latent Feature Domain
abstract
Facial action unit (AU) detection in the wild is a challenging problem, due to the unconstrained variability in facial appearances and the lack of accurate annotations. Most existing methods depend on either impractical labor-intensive labeling or inaccurate pseudo labels. In this paper, we propose an end-to-end unconstrained facial AU detection framework based on domain adaptation, which transfers accurate AU labels from a constrained source domain to an unconstrained target domain by exploiting labels of AU-related facial landmarks. Specifically, we map a source image with label and a target image without label into a latent feature domain by combining source landmark-related feature with target landmark-free feature. Due to the combination of source AU-related information and target AU-free information, the latent feature domain with transferred source label can be learned by maximizing the target-domain AU detection performance. Moreover, we introduce a novel landmark adversarial loss to disentangle the landmark-free feature from the landmark-related feature by treating the adversarial learning as a multi-player minimax game. Our framework can also be naturally extended for use with target-domain pseudo AU labels. Extensive experiments show that our method soundly outperforms lower-bounds and upper-bounds of the basic model, as well as state-of-the-art approaches on the challenging in-the-wild benchmarks. The code is available athttps://github.com/ZhiwenShao/ADLD.
Zhiwen Shao, Jianfei Cai 0001, Tat-Jen Cham, Xuequan Lu, Lizhuang Ma
IEEE Trans. Affect. Comput.1
2022 Facial Action Unit Detection Using Attention and Relation Learning
abstract
Attention mechanism has recently attracted increasing attentions in the field of facial action unit (AU) detection. By finding the region of interest of each AU with the attention mechanism, AU-related local features can be captured. Most of the existing attention based AU detection works use prior knowledge to predefine fixed attentions or refine the predefined attentions within a small range, which limits their capacity to model various AUs. In this paper, we propose an end-to-end deep learning based attention and relation learning framework for AU detection with only AU labels, which has not been explored before. In particular, multi-scale features shared by each AU are learned first, and then both channel-wise and spatial attentions are adaptively learned to select and extract AU-related local features. Moreover, pixel-level relations for AUs are further captured to refine spatial attentions so as to extract more relevant local features. Without changing the network architecture, our framework can be easily extended for AU intensity estimation. Extensive experiments show that our framework (i) soundly outperforms the state-of-the-art methods for both AU detection and AU intensity estimation on the challenging BP4D, DISFA, FERA 2015, and BP4D+ benchmarks, (ii) can adaptively capture the correlated regions of each AU, and (iii) also works well under severe occlusions and large poses.
Zhiwen Shao, Zhilei Liu, Jianfei Cai 0001, Yunsheng Wu, Lizhuang Ma
IEEE Trans. Affect. Comput.1
2022 Sketch-to-photo face generation based on semantic consistency preserving and similar connected component refinement
Junshu Tang, Zhiwen Shao, Xin Tan 0002, Lizhuang Ma
Vis. Comput.3
2022 Facial action unit detection via hybrid relational reasoning
Zhiwen Shao, Yong Zhou 0003, Bing Liu 0016, Hancheng Zhu, Wen-Liang Du 0002, Jiaqi Zhao 0001
Vis. Comput.1
2021 JÂA-Net: Joint Facial Action Unit Detection and Face Alignment Via Adaptive Attention
Zhiwen Shao, Zhilei Liu, Jianfei Cai 0001, Lizhuang Ma
Int. J. Comput. Vis.1
2021 Explicit Facial Expression Transfer via Fine-Grained Representations
abstract
Facial expression transfer between two unpaired images is a challenging problem, as fine-grained expression is typically tangled with other facial attributes. Most existing methods treat expression transfer as an application of expression manipulation, and use predicted global expression, landmarks or action units (AUs) as a guidance. However, the prediction may be inaccurate, which limits the performance of transferring fine-grained expression. Instead of using an intermediate estimated guidance, we propose to explicitly transfer facial expression by directly mapping two unpaired input images to two synthesized images with swapped expressions. Specifically, considering AUs semantically describe fine-grained expression details, we propose a novel multi-class adversarial training method to disentangle input images into two types of fine-grained representations: AU-related feature and AU-free feature. Then, we can synthesize new images with preserved identities and swapped expressions by combining AU-free features with swapped AU-related features. Moreover, to obtain reliable expression transfer results of the unpaired input, we introduce a swap consistency loss to make the synthesized images and self-reconstructed images indistinguishable. Extensive experiments show that our approach outperforms the state-of-the-art expression manipulation methods for transferring fine-grained expressions while preserving other attributes including identity and pose.
Zhiwen Shao, Hengliang Zhu, Junshu Tang, Xuequan Lu, Lizhuang Ma
IEEE Trans. Image Process.1
2020 "Forget" the Forget Gate: Estimating Anomalies in Videos Using Self-contained Long Short-Term Memory Networks
Habtamu Fanta, Zhiwen Shao, Lizhuang Ma
CGI2
2020 Fine-Grained Expression Manipulation Via Structured Latent Space
abstract
Fine-grained facial expression manipulation is a challenging problem, as fine-grained expression details are difficult to be captured. Most existing expression manipulation methods resort to discrete expression labels, which mainly edit global expressions and ignore the manipulation of fine details. To tackle this limitation, we propose an end-to-end expression-guided generative adversarial network (EGGAN), which utilizes structured latent codes and continuous expression labels as input to generate images with expected expressions. Specifically, we adopt an adversarial autoencoder to map a source image into a structured latent space. Then, given the source latent code and the target expression label, we employ a conditional GAN to generate a new image with the target expression. Moreover, we introduce a perceptual loss and a multi-scale structural similarity loss to preserve identity and global shape during generation. Extensive experiments show that our method can manipulate fine-grained expressions, and generate continuous intermediate expressions between source and target expressions.
Junshu Tang, Zhiwen Shao, Lizhuang Ma
ICME2
2020 CPCS: Critical Points Guided Clustering and Sampling for Point Cloud Analysis
Zhiwen Shao, Wencai Zhong, Lizhuang Ma
ICONIP (4)2
2020 Deep multi-center learning for face alignment
Zhiwen Shao, Hengliang Zhu, Xin Tan 0002, Yangyang Hao, Lizhuang Ma
Neurocomputing1
2020 SiTGRU: Single-Tunnelled Gated Recurrent Unit for Abnormality Detection
Habtamu Fanta, Zhiwen Shao, Lizhuang Ma
Inf. Sci.2
2019 Feedback cascade regression model for face alignment
abstract
Face alignment has made great progress in recent years and the cascade regression framework is one of the main contributors. However, the performance of this framework is unsatisfactory on heavily occluded faces or those far from the frontal pose. This is because regression is sensitive to hidden landmarks and unified initialisation can often lead to the method falling into local minima. The authors propose a new pipeline of salient‐to‐inner‐to‐all to progressively compute the locations of landmarks. Additionally, a feedback process is utilised to improve the robustness of regression. They bring out a pose‐invariant shape retrieval method to generate the discriminative initialisation. Experiments are performed on two benchmarks, and the experimental results demonstrate that the proposed method has a considerable improvement on the cascade regression model, and achieves favourable results compared with the state‐of‐the‐art deep learning‐based methods.
Yangyang Hao, Hengliang Zhu, Zhiwen Shao, Lizhuang Ma
IET Comput. Vis.3
2018 Deep Adaptive Attention for Joint Facial Action Unit Detection and Face Alignment
Zhiwen Shao, Zhilei Liu, Jianfei Cai 0001, Lizhuang Ma
ECCV (13)1
2018 Saliency Detection by Deep Network with Boundary Refinement and Global Context
abstract
A novel end-to-end fully convolutional neural network for saliency detection is proposed in this paper, aiming at refining the boundary and covering the global context (GBR-Net). Previous CNN based methods for saliency detection are universally accompanied with blurring edge and ambiguous salient object. To tackle this problem, we propose to embed the boundary enhancement block (BEB) into the network to refine edge. It keeps the details by the mutual-coupling con-volutionallayers. Besides, we employ a pooling pyramid that utilizes the multi-level feature informations to search global context, and it also contributes as an auxiliary supervision. The final saliency map is obtained by fusing the edge refinement with global context extraction. Experiments on four benchmark datasets prove that the proposed saliency detection model gains an edge over the state-of-the-art approaches.
Xin Tan 0002, Hengliang Zhu, Zhiwen Shao, Xiao-Nan Hou, Yangyang Hao, Lizhuang Ma
ICME3
2018 Multi-Path Feature Fusion Network for Saliency Detection
abstract
Recent saliency detection methods have made great progress with the fully convolutional network. However, we find that the saliency maps are usually coarse and fuzzy, especially near the boundary of salient object. To deal with this problem, in this paper, we exploit a multi-path feature fusion model for saliency detection. The proposed model is a fully convolutional network with raw images as input and saliency maps as output. In particular, we propose a multi-path fusion strategy for deriving the intrinsic features of salient objects. The structure has the ability of capturing the low-level visual features and generating the boundary-preserving saliency maps. Moreover, a coupled structure module is proposed in our model, which helps to explore the high-level semantic properties of salient objects. Extensive experiments on four public benchmarks indicate that our saliency model is effective and outperforms state-of-the-art methods.
Hengliang Zhu, Xin Tan 0002, Zhiwen Shao, Yangyang Hao, Lizhuang Ma
ICME3
2018 Facial Landmark Detection Under Large Pose
Yangyang Hao, Hengliang Zhu, Zhiwen Shao, Xin Tan 0002, Lizhuang Ma
ICONIP (4)3
2018 Better initialization for regression-based face alignment
Hengliang Zhu, Bin Sheng 0001, Zhiwen Shao, Yangyang Hao, Xiao-Nan Hou, Lizhuang Ma
Comput. Graph.3
2017 Learning a multi-center convolutional network for unconstrained face alignment
abstract
In this paper, we propose a novel multi-center convolutional neural network for unconstrained face alignment. To utilize structural correlations among different facial landmarks, we determine several clusters based on their spatial position. We pre-train our network to learn generic feature representations. We further fine-tune the pre-trained model to emphasize on locating a certain cluster of landmarks respectively. Fine-tuning contributes to searching an optimal solution smoothly without deviating from the pre-trained model excessively. We obtain an excellent solution by combining multiple fine-tuned models. Extensive experiments demonstrate that our method possesses superior capability of handling extreme occlusions and complex variations of pose, expression, illumination. The code for our method is available at https://github.com/ZhiwenShao/MCNet.
Zhiwen Shao, Hengliang Zhu, Yangyang Hao, Min Wang 0024, Lizhuang Ma
ICME1
2016 Face alignment by deep convolutional network with adaptive learning rate
abstract
Deep convolutional network has been widely used in face recognition while not often used in face alignment. One of the most important reasons of this is the lack of training images annotated with landmarks due to fussy and time-consuming annotation work. To overcome this problem, we propose a novel data augmentation strategy. And we design an innovative training algorithm with adaptive learning rate for two iterative procedures, which helps the network to search an optimal solution. Our convolutional network can learn global high-level features and directly predict the coordinates of facial landmarks. Extensive evaluations show that our approach outperforms state-of-the-art methods especially in the condition of complex occlusion, pose, illumination and expression variations.
Zhiwen Shao, Shouhong Ding, Hengliang Zhu, Chengjie Wang 0001, Lizhuang Ma
ICASSP1
2016 LSOD: Local Sparse Orthogonal Descriptor for Image Matching
abstract
We propose a novel method for feature description used for image matching in this paper. Our method is inspired by the autoencoder, an artificial neural network designed for learning efficient codings. Sparse and orthogonal constraints are imposed on the autoencoder and make it a highly discriminative descriptor. It is shown that the proposed descriptor is not only invariant to geometric and photometric transformations (such as viewpoint change, intensity change, noise, image blur and JPEG compression), but also highly efficient. We compare it with existing state-of-the-art descriptors on standard benchmark datasets, the experimental results show that our LSOD method yields better performance both in accuracy and efficiency.
Yiru Zhao, Yaoyi Li, Zhiwen Shao, Hongtao Lu 0001
ACM Multimedia3