Zibo Meng

dblp:126/8042 · DBLP profile ↗
← Back
22ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0001-7299-7290ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 F2T2-HIT: A U-Shaped FFT Transformer and Hierarchical Transformer for Reflection Removal
abstract
Single Image Reflection Removal (SIRR) technique plays a crucial role in image processing by eliminating unwanted reflections from the background. These reflections, often caused by photographs taken through glass surfaces, can significantly degrade image quality. SIRR remains a challenging problem due to the complex and varied reflections encountered in real-world scenarios. These reflections vary significantly in intensity, shapes, light sources, sizes, and coverage areas across the image, posing challenges for most existing methods to effectively handle all cases. To address these challenges, this paper introduces a U-shaped Fast Fourier Transform Transformer and Hierarchical Transformer (F2T2-HiT) architecture, an innovative Transformer-based design for SIRR. Our approach uniquely combines Fast Fourier Transform (FFT) Transformer blocks and Hierarchical Transformer blocks within a UNet framework. The FFT Transformer blocks leverage the global frequency domain information to effectively capture and separate reflection patterns, while the Hierarchical Transformer blocks utilize multi-scale feature extraction to handle reflections of varying sizes and complexities. Extensive experiments conducted on three publicly available testing datasets demonstrate state-of-the-art performance, validating the effectiveness of our approach.
Jie Cai 0001, Kangning Yang, Ling Ouyang, Lan Fu, Jiaming Ding, Huiming Sun, Chiu Man Ho, Zibo Meng
ICIP8
2025 OpenRR-1k: A Scalable Dataset for Real-World Reflection Removal
abstract
Reflection removal technology plays a crucial role in photography and computer vision applications. However, existing techniques are hindered by the lack of high-quality in-the-wild datasets. In this paper, we propose a novel paradigm for collecting reflection datasets from a fresh perspective. Our approach is convenient, cost-effective, and scalable, while ensuring that the collected data pairs are of high quality, perfectly aligned, and represent natural and diverse scenarios. Following this paradigm, we collect a Real-world, Diverse, and Pixel-aligned dataset (named OpenRR-1k dataset), which contains 1,000 high-quality transmission-reflection image pairs collected in the wild. Through the analysis of several reflection removal methods and benchmark evaluation experiments on our dataset, we demonstrate its effectiveness in improving robustness in challenging real-world environments. Our dataset is available at https://github.com/caijie0620/OpenRR-1k.
Kangning Yang, Ling Ouyang, Huiming Sun, Jie Cai 0001, Lan Fu, Jiaming Ding, Chiu Man Ho, Zibo Meng
ICIP8
2024 EVD4UAV: An Altitude-Sensitive Benchmark to Evade Vehicle Detection in UAV
abstract
Vehicle detection in Unmanned Aerial Vehicle (UAV) captured images has wide applications in aerial photography and remote sensing. There are many public benchmark datasets proposed for the vehicle detection and tracking in UAV images. Recent studies show that adding an adversarial patch on objects can fool the well-trained deep neural networks based object detectors, posing security concerns to the downstream tasks. However, the current public UAV datasets might ignore the diverse altitudes, vehicle attributes, fine-grained instance-level annotation in mostly side view with blurred vehicle roof, so none of them is good to study the adversarial patch based vehicle detection attack problem. In this paper, we propose a new dataset named EVD4UAV as an altitude-sensitive benchmark to evade vehicle detection in UAV with 6,284 images and 90,886 fine-grained annotated vehicles. The EVD4UAV dataset has diverse altitudes (50m, 70m, 90m), vehicle attributes (color, type), fine-grained annotation (horizontal and rotated bounding boxes, instance-level mask) in top view with clear vehicle roof. One white-box and two black-box patch based attack methods are implemented to attack three classic deep neural networks based object detectors on EVD4UAV. The experimental results show that these representative attack methods could not achieve the robust altitude-insensitive attack performance.
Huiming Sun, Jiacheng Guo, Zibo Meng, Tianyun Zhang, Jianwu Fang, Yuewei Lin, Hongkai Yu
IV3
2024 Defense against Adversarial Cloud Attack on Remote Sensing Salient Object Detection
abstract
Detecting the salient objects in a remote sensing image has wide applications. Many existing deep learning methods have been proposed for Salient Object Detection (SOD) in remote sensing images with remarkable results. However, the recent adversarial attack examples, generated by changing a few pixel values on the original image, could result in a collapse for the well-trained deep learning model. Different with existing methods adding perturbation to original images, we propose to jointly tune adversarial exposure and additive perturbation for attack and constrain image close to cloudy image as Adversarial Cloud. Cloud is natural and common in remote sensing images, however, camouflaging cloud based adversarial attack and defense for remote sensing images are not well studied before. Furthermore, we design DefenseNet as a learnable pre-processing to the adversarial cloudy images to preserve the performance of the deep learning based remote sensing SOD model, without tuning the already deployed deep SOD model. By considering both regular and generalized adversarial examples, the proposed DefenseNet can defend the proposed Adversarial Cloud in white-box setting and other attack methods in black-box setting. Experimental results on a synthesized benchmark from the public remote sensing dataset (EORSSD) show the promising defense against adversarial cloud attacks.
Huiming Sun, Lan Fu, Qing Guo 0005, Zibo Meng, Tianyun Zhang, Yuewei Lin, Hongkai Yu
WACV5
2023 Pik-Fix: Restoring and Colorizing Old Photos
abstract
Restoring and inpainting the visual memories that are present, but often impaired, in old photos remains an intriguing but unsolved research topic. Decades-old photos often suffer from severe and commingled degradation such as cracks, defocus, and color-fading, which are difficult to treat individually and harder to repair when they interact. Deep learning presents a plausible avenue, but the lack of large-scale datasets of old photos makes addressing this restoration task very challenging. Here we present a novel reference-based end-to-end learning framework that is able to both repair and colorize old, degraded pictures. Our proposed framework consists of three modules: a restoration sub-network that conducts restoration from degradations, a similarity network that performs color histogram matching and color transfer, and a colorization subnet that learns to predict the chroma elements of images conditioned on chromatic reference signals. The overall system makes uses of color histogram priors from reference images, which greatly reduces the need for large-scale training data. We have also created a first-of-a-kind public dataset of real old photos that are paired with ground truth "pristine" photos that have been manually restored by PhotoShop experts. We conducted extensive experiments on this dataset and synthetic datasets, and found that our method significantly outperforms previous state-of-the-art models using both qualitative comparisons and quantitative measurements. The code is available at https://github.com/DerrickXuNu/Pik-Fix.
Runsheng Xu, Zhengzhong Tu, Yuanqi Du, Zibo Meng, Jiaqi Ma 0003, Alan C. Bovik, Hongkai Yu
WACV6
2023 Probabilistic Attribute Tree Structured Convolutional Neural Networks for Facial Expression Recognition in the Wild
abstract
Very recent work has demonstrated tremendous improvements in facial expression recognition (FER) on laboratory-controlled datasets. However, recognizing facial expressions under in-the-wild conditions still remains challenging, especially on unseen subjects due to high inter-subject variations. In this paper, we propose a novel Probabilistic Attribute Tree Convolutional Neural Network (PAT-CNN) to explicitly deal with large intra-class variations caused by identity-related attributes, e.g., age, race, and gender. Specifically, a PAT module with an associated PAT loss is proposed to learn features in a hierarchical tree structure organized according to identity-related attributes, where the final features are less affected by the attributes. Then, expression-related features are extracted from leaf nodes. Samples are probabilistically assigned to tree nodes at different levels such that expression-related features can be learned from all samples weighted by probabilities. Furthermore, the proposed PAT-CNN can be learned from limited attribute-annotated samples to make the best use of available data. Experimental results on four spontaneous facial expression datasets, i.e., RAF-DB, SFEW, ExpW, and FER-2013, have demonstrated that the proposed PAT-CNN achieves the best performance when compared to state-of-the-art methods by explicitly modeling attributes. Impressively, a single model PAT-CNN achieves the best performance on the SFEW test dataset when compared to the state-of-the-art methods using an ensemble of hundreds of CNNs.
Jie Cai 0001, Zibo Meng, Ahmed-Shehab Khan, Zhiyuan Li 0006, James O'Reilly
IEEE Trans. Affect. Comput.2
2022 Point Adversarial Self-Mining: A Simple Method for Facial Expression Recognition
abstract
In this article, we propose a simple yet effective approach, called point adversarial self mining (PASM), to improve the recognition accuracy in facial expression recognition (FER). Unlike previous works focusing on designing specific architectures or loss functions to solve this problem, PASM boosts the network capability by simulating human learning processes: providing updated learning materials and guidance from more capable teachers. Specifically, to generate new learning materials, PASM leverages a point adversarial attack method and a trained teacher network to locate the most informative position related to the target task, generating harder learning samples to refine the network. The searched position is highly adaptive since it considers both the statistical information of each sample and the teacher network capability. Other than being provided new learning materials, the student network also receives guidance from the teacher network. After the student network finishes training, the student network changes its role and acts as a teacher, generating new learning materials and providing stronger guidance to train a better student network. The adaptive learning materials generation and teacher/student update can be conducted more than one time, improving the network capability iteratively. Extensive experimental results validate the efficacy of our method over the existing state of the arts for FER.
Ping Liu 0004, Yuewei Lin, Zibo Meng, Weihong Deng, Joey Tianyi Zhou, Yi Yang 0001
IEEE Trans. Cybern.3
2021 Identity-Free Facial Expression Recognition Using Conditional Generative Adversarial Network
abstract
A novel Identity-Free conditional Generative Adversarial Network (IF-GAN) was proposed for Facial Expression Recognition (FER) to explicitly reduce high inter-subject variations caused by identity-related facial attributes, e.g., age, race, and gender. As part of an end-to-end system, a cGAN was designed to transform a given input facial expression to an “average” identity face with the same expression as the input. Then, identity-free FER is possible since the generated images have the same synthetic “average” identity and differ only in their displayed expressions. Experiments on four facial expression datasets, one with spontaneous expressions, show that IF-GAN outperforms the baseline CNN and achieves state-of-the-art performance for FER.
Jie Cai 0001, Zibo Meng, Ahmed-Shehab Khan, James O'Reilly, Zhiyuan Li 0006, Shizhong Han
ICIP2
2019 Pooling Map Adaptation in Convolutional Neural Network for Facial Expression Recognition
abstract
In this work, we proposed adaptive pooling maps (APMs) for CNNs to aid facial expression recognition. Inspired by superpixels, which represent the image content more naturally, pooling maps consisting of irregular pooling regions are learned from training images as part of training a CNN model. The APMs preserve the local structural information and thus are more capable of capturing subtle facial appearance and geometrical changes caused by facial expression. Furthermore, we developed an efficient algorithm to learn the APMs efficiently. Experiments on three benchmark datasets have shown that the proposed APM-based CNN model outperforms the one with the standard pooling map and achieves state-of-the-art recognition performance for facial expression recognition in the wild.
Zhiyuan Li 0006, Shizhong Han, Ahmed-Shehab Khan, Jie Cai 0001, Zibo Meng, James O'Reilly
ICME5
2019 Listen to Your Face: Inferring Facial Action Units from Audio Channel
abstract
Extensive efforts have been devoted to recognizing facial action units (AUs). However, it is still challenging to recognize AUs from spontaneous facial displays especially when they are accompanied by speech. Different from all prior work that utilized visual observations for facial AU recognition, this paper presents a novel approach that recognizes speech-related AUs exclusively from audio signals based on the fact that facial activities are highly correlated with voice during speech. Specifically, dynamic and physiological relationships between AUs and phonemes are modeled through a continuous time Bayesian network (CTBN); then AU recognition is performed by probabilistic inference via the CTBN model. A pilot audiovisual AU-coded database has been constructed to evaluate the proposed audio-based AU recognition framework. The database consists of a “clean” subset with frontal and neutral faces and a challenging subset collected with large head movements and occlusions. Experimental results on this database show that the proposed CTBN model achieves promising recognition performance for 7 speech-related AUs and outperforms both the state-of-the-art visual-based and audio-based methods especially for those AUs that are activated at low intensities or “hardly visible” in the visual channel. The improvement is more impressive on the challenging subset, where the visual-based approaches suffer significantly.
Zibo Meng, Shizhong Han
IEEE Trans. Affect. Comput.1
2019 Improving Speech Related Facial Action Unit Recognition by Audiovisual Information Fusion
abstract
It is challenging to recognize facial action unit (AU) from spontaneous facial displays, especially when they are accompanied by speech. The major reason is that the information is extracted from a single source, i.e., the visual channel, in the current practice. However, facial activity is highly correlated with voice in natural human communications. Instead of solely improving visual observations, this paper presents a novel audiovisual fusion framework, which makes the best use of visual and acoustic cues in recognizing speech-related facial AUs. In particular, a dynamic Bayesian network is employed to explicitly model the semantic and dynamic physiological relationships between AUs and phonemes as well as measurement uncertainty. Experiments on a pilot audiovisual AU-coded database have demonstrated that the proposed framework significantly outperforms the state-of-the-art visual-based methods in terms of recognizing speech-related AUs, especially for those AUs whose visual observations are impaired during speech, and more importantly is also superior to audio-based methods and feature-level fusion methods, which employ low-level audio features, by explicitly modeling and exploiting physiological relationships between AUs and phonemes.
Zibo Meng, Shizhong Han, Ping Liu 0004
IEEE Trans. Cybern.1
2018 Optimizing Filter Size in Convolutional Neural Networks for Facial Action Unit Recognition
abstract
Recognizing facial action units (AUs) during spontaneous facial displays is a challenging problem. Most recently, Convolutional Neural Networks (CNNs) have shown promise for facial AU recognition, where predefined and fixed convolution filter sizes are employed. In order to achieve the best performance, the optimal filter size is often empirically found by conducting extensive experimental validation. Such a training process suffers from expensive training cost, especially as the network becomes deeper. This paper proposes a novel Optimized Filter Size CNN (OFS-CNN), where the filter sizes and weights of all convolutional layers are learned simultaneously from the training data along with learning convolution filters. Specifically, the filter size is defined as a continuous variable, which is optimized by minimizing the training loss. Experimental results on two AU-coded spontaneous databases have shown that the proposed OFS-CNN is capable of estimating optimal filter size for varying image resolution and outperforms traditional CNNs with the best filter size obtained by exhaustive search. The OFS-CNN also beats the CNN using multiple filter sizes and more importantly, is much more efficient during testing with the proposed forward-backward propagation algorithm.
Shizhong Han, Zibo Meng, Zhiyuan Li 0006, James O'Reilly, Jie Cai 0001
CVPR2
2018 Island Loss for Learning Discriminative Features in Facial Expression Recognition
abstract
Over the past few years, Convolutional Neural Networks (CNNs) have shown promise on facial expression recognition. However, the performance degrades dramatically under real-world settings due to variations introduced by subtle facial appearance changes, head pose variations, illumination changes, and occlusions. In this paper, a novel island loss is proposed to enhance the discriminative power of deeply learned features. Specifically, the island loss is designed to reduce the intra-class variations while enlarging the inter-class differences simultaneously. Experimental results on four benchmark expression databases have demonstrated that the CNN with the proposed island loss (IL-CNN) outperforms the baseline CNN models with either traditional softmax loss or center loss and achieves comparable or better performance compared with the state-of-the-art methods for facial expression recognition.
Jie Cai 0001, Zibo Meng, Ahmed-Shehab Khan, Zhiyuan Li 0006, James O'Reilly
FG2
2018 Group-Level Emotion Recognition using Deep Models with A Four-stream Hybrid Network
abstract
Group-level Emotion Recognition (GER) in the wild is a challenging task gaining lots of attention. Most recent works utilized two channels of information, a channel involving only faces and a channel containing the whole image, to solve this problem. However, modeling the relationship between faces and scene in a global image remains challenging. In this paper, we proposed a novel face-location aware global network, capturing the face location information in the form of an attention heatmap to better model such relationships. We also proposed a multi-scale face network to infer the group-level emotion from individual faces, which explicitly handles high variance in image and face size, as images in the wild are collected from different sources with different resolutions. In addition, a global blurred stream was developed to explicitly learn and extract the scene-only features. Finally, we proposed a four-stream hybrid network, consisting of the face-location aware global stream, the multi-scale face stream, a global blurred stream, and a global stream, to address the GER task, and showed the effectiveness of our method in GER sub-challenge, a part of the six Emotion Recognition in the Wild (EmotiW 2018) [10] Challenge. The proposed method achieved 65.59% and 78.39% accuracy on the testing and validation sets, respectively, and is ranked the third place on the leaderboard.
Ahmed-Shehab Khan, Zhiyuan Li 0006, Jie Cai 0001, Zibo Meng, James O'Reilly
ICMI4
2017 Down syndrome prediction/screening model based on deep learning and illumina genotyping array
abstract
Down syndrome (DS) is a genetic disorder with genome dosage imbalances and micro-duplications of human chromosome 21. It is usually associated with a group of serious diseases, including intellectual disabilities, cardiac diseases, physical abnormalities, and other abnormalities. Currently, since there is no cure for human DS, screening and early detection have become the most efficient way for DS prevention. In this study, we used deep learning techniques to build accurate DS prediction/screening models based on the analysis of newly introduced Illumina genotyping array. Specifically, we built chromosome SNP maps based on clinical genotyping data collected by Vanderbilt University Medical Center. Then we proposed a convolutional neural network (CNN) architecture with ten layers and two merged CNN models, which took two input chromosome SNP maps in combination. Our CNN DS prediction/screening model achieved over 99.3% average accuracy, as well as very low false positive and false negative rate, which are critical to disease prediction and screening in medical practice. It also had better performances in terms of all evaluating metrics when compared with three conventional machine-learning algorithms. Finally, we visualized the feature maps and the trained filter weights from intermediate layers of our trained CNN model. We further discussed the advantages of our method and the underlying reasons for its robust performance.
Bing Feng, David C. Samuels, William Hoskins, Jijun Tang, Zibo Meng
BIBM7
2017 Identity-Aware Convolutional Neural Network for Facial Expression Recognition
abstract
Facial expression recognition suffers under realworldconditions, especially on unseen subjects due to highinter-subject variations. To alleviate variations introduced bypersonal attributes and achieve better facial expression recognitionperformance, a novel identity-aware convolutional neuralnetwork (IACNN) is proposed. In particular, a CNN with a newarchitecture is employed as individual streams of a bi-streamidentity-aware network. An expression-sensitive contrastive lossis developed to measure the expression similarity to ensure thefeatures learned by the network are invariant to expressionvariations. More importantly, an identity-sensitive contrastiveloss is proposed to learn identity-related information from identitylabels to achieve identity-invariant expression recognition.Extensive experiments on three public databases including aspontaneous facial expression database have shown that theproposed IACNN achieves promising results in real world.
Zibo Meng, Ping Liu 0004, Jie Cai 0001, Shizhong Han
FG1
2016 Incremental Boosting Convolutional Neural Network for Facial Action Unit Recognition
abstract
Recognizing facial action units (AUs) from spontaneous facial expressions is still a challenging problem. Most recently, CNNs have shown promise on facial AU recognition. However, the learned CNNs are often overfitted and do not generalize well to unseen subjects due to limited AU-coded training images. We proposed a novel Incremental Boosting CNN (IB-CNN) to integrate boosting into the CNN via an incremental boosting layer that selects discriminative neurons from the lower layer and is incrementally updated on successive mini-batches. In addition, a novel loss function that accounts for errors from both the incremental boosted classifier and individual weak classifiers was proposed to fine-tune the IB-CNN. Experimental results on four benchmark AU databases have demonstrated that the IB-CNN yields significant improvement over the traditional CNN and the boosting CNN without incremental learning, as well as outperforming the state-of-the-art CNN-based methods in AU recognition. The improvement is more impressive for the AUs that have the lowest frequencies in the databases.
Shizhong Han, Zibo Meng, Ahmed-Shehab Khan
NIPS2
2015 Feature Level Fusion for Bimodal Facial Action Unit Recognition
abstract
Recognizing facial actions from spontaneous facial displays suffers from subtle and complex facial deformation, frequent head movements, and partial occlusions. It is especially challenging when the facial activities are accompanied with speech. Instead of employing information solely from the visual channel, this paper presents a novel fusion framework, which exploits information from both visual and audio channels in recognizing speech-related facial action units (AUs). In particular, features are first extracted from visual and audio channels, independently. Then, the audio features are aligned with the visual features in order to handle the difference in time scales and the time shift between the two signals. Finally, these aligned audio and visual features are integrated via a feature-level fusion framework and utilized in recognizing AUs. Experimental results on a new audiovisual AU-coded dataset have demonstrated that the proposed feature-level fusion framework outperforms a state-of-the-art visual-based method in recognizing speech-related AUs, especially for those AUs that are "invisible" in the visual channel during speech. The improvement is more impressive with occlusions on the facial images, which, fortunately, would not affect the audio channel.
Zibo Meng, Shizhong Han, Min Chen 0009
ISM1
2014 Facial Expression Recognition via a Boosted Deep Belief Network
abstract
A training process for facial expression recognition is usually performed sequentially in three individual stages: feature learning, feature selection, and classifier construction. Extensive empirical studies are needed to search for an optimal combination of feature representation, feature set, and classifier to achieve good recognition performance. This paper presents a novel Boosted Deep Belief Network (BDBN) for performing the three training stages iteratively in a unified loopy framework. Through the proposed BDBN framework, a set of features, which is effective to characterize expression-related facial appearance/shape changes, can be learned and selected to form a boosted strong classifier in a statistical way. As learning continues, the strong classifier is improved iteratively and more importantly, the discriminative capabilities of selected features are strengthened as well according to their relative importance to the strong classifier via a joint fine-tune process in the BDBN framework. Extensive experiments on two public databases showed that the BDBN framework yielded dramatic improvements in facial expression analysis.
Ping Liu 0004, Shizhong Han, Zibo Meng
CVPR3
2014 Feature Disentangling Machine - A Novel Approach of Feature Selection and Disentangling in Facial Expression Analysis
Ping Liu 0004, Joey Tianyi Zhou, Ivor W. Tsang, Zibo Meng, Shizhong Han
ECCV (4)4
2014 Facial grid transformation: A novel face registration approach for improving facial action unit recognition
abstract
Face registration is a major and critical step for face analysis. Existing facial activity recognition systems often employ coarse face alignment based on a few fiducial points such as eyes and extract features from equal-sized grid. Such extracted features are susceptible to variations in face pose, facial deformation, and person-specific geometry. In this work, we propose a novel face registration method named facial grid transformation to improve feature extraction for recognizing facial Action Units (AUs). Based on the transformed grid, novel grid edge features are developed to capture local facial motions related to AUs. Extensive experiments on two well-known AU-coded databases have demonstrated that the proposed method yields significant improvements over the methods based on equal-sized grid on both posed and more importantly, spontaneous facial displays. Furthermore, the proposed method also outperforms the state-of-the-art methods using either coarse alignment or mesh-based face registration.
Shizhong Han, Zibo Meng, Ping Liu 0004
ICIP2
2013 Coordinated transceiver in MIMO heterogeneous network with physical-layer network coding
abstract
In this contribution, a multiple-input multiple-output heterogeneous network with physical-layer network coding is considered. The article proposes a coordinated joint transmitter and receiver design to tackle the interference problem between two-way relaying channel in small cell and uplink channel in macro cell. We design the transceiver on the basis that the data streams from two-way communication nodes are aligned to the same direction to perform multi-stream decode-and-forward physical-layer network coding, while other streams are adjusted to orthogonal direction for spatial multiplexing. Coordinated beam-forming and joint multi-cell signal processing are both considered. Based on the criterion that minimizes system-wide mean square error, we formulate the transceiver optimization problem with transmit power constraints on each user and provide an approximate optimal solution for the non-convex problem through iterative algorithm. During the iteration, each of the transmitters and receivers is solved by Lagrange multiplier method and Karush-Kuhn-Tucker condition. The scheme reduces interferences when all the nodes transmit multi-stream symbols simultaneously and significantly improves the system performance. Numerical results such as BER and MSE performances are provided to support the proposed scheme and show that a below 10-3BER could be reached when SNR≥8dB which is close to the lower bound.
Zhigang Wen, Zibo Meng, Dongjian Chen, Chunxiao Fan 0001
PIMRC3