Bolun Cai

dblp:124/2077 · DBLP profile ↗
← Back
24ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0002-9394-7820ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
YearPublicationVenuePosition
2026 EETalk: Expression Enhancement in Speech-Driven 3D Facial Animation
abstract
Speech-driven 3D facial animation aims to generate natural and expressive facial movements from speech. Although significant progress has been made, existing methods still face challenges in generating realistic upper facial expressions. Specifically, existing methods that jointly optimize the holistic face tend to overlook fine-grained spatial movements due to motion differences across facial regions. Pre-trained speech feature extractors, which emphasize long-term dependencies, provide limited fine-grained temporal cues. In this work, we propose a novel framework, EETalk, to enhance the realism of facial expressions, which captures fine-grained spatial information and f ine-grained temporal dynamics from speech. To alleviate loss of fine-grained spatial information, we propose a novel disassemble and-reassemble modeling strategy. This strategy constructs two independent motion representation spaces for the upper and lower faces, allowing for the capture of weak upper-face movements while preserving motion diversity. Then, we propose a Cross-Region Coordination Module to ensure the synchronization and coordination of the movements from the independent upper and lower faces. To effectively capture facial micro-expressions, we incorporate fine-grained time-varying features to compensate for the short-timescale details underrepresented in the long term semantic features extracted by pre-trained self-supervised models, thereby improving the generation of fast, subtle facial expressions. Experimental results demonstrate that our approach significantly improves motion accuracy, expression consistency, and perceptual quality compared to existing methods.
Zhaojie Chu, Kailing Guo, Xiaofen Xing, Bolun Cai, Lin Wang 0004, Xiangmin Xu 0001
IEEE Trans. Multim.4
2025 Alleviating One-to-Many Mapping in Talking Head Synthesis With Dynamic Adaptation Context and Style Adapter
abstract
Speech-driven talking head synthesis technology has made remarkable progress, but it still faces the challenge of one-to-many pathological mapping. The challenge results in inaccurate lip movements, ambiguity in facial expressions, and a lack of coherence during transitions between facial motions. The phenomenon is primarily caused by: (1) for one speaker, the same phoneme corresponds to a wide range of mouth shapes and facial expressions due to contextual variations, and (2) for the same spoken content, different speakers exhibit diverse facial motions as a result of unique speaking styles. In this work, we propose a novel framework, called AllTalk, to alleviate one-to-many pathological mapping, which enables a more vivid and natural talking head. Specifically, considering the asymmetry and dynamic nature of mouth shapes’ dependence on phoneme context, we propose a Dynamic Adaptive Context encoder to capture the context around the phoneme and its dynamics, thereby reducing the ambiguity in mapping speech to facial movements. Moreover, to alleviate the uncertainty caused by differences in speaking style, we propose a Style Adapter that expands a generic discrete motion space for the target speaker. The Style Adapter not only effectively represents general facial motions but also captures the personalized nuances of facial movements. To further enhance the fidelity of output, we introduce a Dynamic Gaussian Renderer based on 3D Gaussian Splatting, capable of producing stable and realistic rendering videos. Extensive qualitative and quantitative experiments demonstrate that AllTalk surpasses existing state-of-the-art methods, providing an effective solution to the challenge of one-to-many mapping. Project page: https://zjchu.github.io/projects/AllTalk.
Zhaojie Chu, Kailing Guo, Xiaofen Xing, Bolun Cai, Xiangmin Xu 0001
IEEE Trans. Circuits Syst. Video Technol.4
2025 DCPTalk: Speech-Driven 3D Face Animation With Personalized Facial Dynamic Coupling Properties
abstract
Speech-driven 3D facial animation has emerged as a hot topic. During this process, movements in different facial regions are interdependent, influenced by the intricate interactions among facial muscles, and manifest personalized differences. The existing methods typically simplify the facial animation generation task to an infinitely thin surface skin deformation without an underlying structure, thereby ignoring the intricate and personalized dynamics of facial muscle activity. These methods tend to produce static or weak upper-face animations with an average facial movement style. In this work, we propose a novel framework, called DCPTalk, to mimic the intricate dynamics of facial muscle activity and portray personalized facial animations. Based on facial dynamic coupling properties, we propose Mouth2Face to simulate the facial muscle control system, yielding realistic and coordinated facial animations evoked by mouth movements. Mouth movements are easily synthesized from speech signals due to their direct correlation with phonetic articulation and vocal tract dynamics. To further enhance the detail of facial movements, we employ surface skin deformation to refine the facial animation derived from Mouth2Face. Furthermore, personal factors, including inherent physical traits and acquired speaking styles, directly determine the uniqueness and realism of facial animations. Inherent physical traits are embedded into Mouth2Face for constructing personalized facial muscle control system, while acquired speaking styles are employed to modulate external driving signals. Extensive qualitative and quantitative experiments as well as a user study indicate that DCPTalk outperforms the existing state-of-the-art methods.
Zhaojie Chu, Kailing Guo, Xiaofen Xing, Pengsheng Liu, Bolun Cai, Xiangmin Xu 0001
IEEE Trans. Multim.5
2024 CorrTalk: Correlation Between Hierarchical Speech and Facial Activity Variances for 3D Animation
abstract
Speech-driven 3D facial animation is a challenging cross-modal task that has attracted growing research interest. During speaking activities, the mouth displays strong motions, while the other facial regions typically demonstrate comparatively weak activity levels. Existing approaches often simplify the process by directly mapping single-level speech features to the entire facial animation, which overlook the differences in facial activity intensity leading to overly smoothed facial movements. In this study, we propose a novel framework, CorrTalk, which effectively establishes the temporal correlation between hierarchical speech features and facial activities of different intensities across distinct regions. A novel facial activity intensity prior is defined to distinguish between strong and weak facial activity, obtained by statistically analyzing facial animations. Based on the facial activity intensity prior, we propose a dual-branch decoding framework to synchronously synthesize strong and weak facial activity, which guarantees wider intensity facial animation synthesis. Furthermore, a weighted hierarchical feature encoder is proposed to establish temporal correlation between hierarchical speech features and facial activity at different intensities, which ensures lip-sync and plausible facial expressions. Extensive qualitatively and quantitatively experiments as well as a user study indicate that our CorrTalk outperforms existing state-of-the-art methods. The source code and supplementary video are publicly available at: https://zjchu.github.io/projects/CorrTalk/.
Zhaojie Chu, Kailing Guo, Xiaofen Xing, Yilin Lan, Bolun Cai, Xiangmin Xu 0001
IEEE Trans. Circuits Syst. Video Technol.5
2023 MGAT: Multi-Granularity Attention Based Transformers for Multi-Modal Emotion Recognition
abstract
Multi-modal emotion recognition is crucial for human-computer interaction. Many existing algorithms attempt to achieve multi-modal interactions through a cross-attention mechanism. Due to the problems of noise introduction and heavy computation in the original attention mechanism, window attention has become a new trend. However, emotions are presented asynchronously between different modalities, which makes it difficult to interact with emotional information between windows. Furthermore, multi-modal data are temporally misaligned, so single fixed window size is hard to describe cross-modal information. In this paper, we put these two issues into a unified framework and propose the multi-granularity attention based Transformers (MGAT). It addresses the emotional asynchrony and modality misalignment issues through a multi-granularity attention mechanism. Experimental results confirm the effectiveness of our method and the state-of-the-art performance is achieved on IEMOCAP.
Weiquan Fan, Xiaofen Xing, Bolun Cai, Xiangmin Xu 0001
ICASSP3
2023 Unsupervised Pre-Training for Detection Transformers
abstract
DEtection TRansformer (DETR) for object detection reaches competitive performance compared with Faster R-CNN via a transformer encoder-decoder architecture. However, trained with scratch transformers, DETR needs large-scale training data and an extreme long training schedule even on COCO dataset. Inspired by the great success of pre-training transformers in natural language processing, we propose a novel pretext task named random query patch detection in Unsupervised Pre-training DETR (UP-DETR). Specifically, we randomly crop patches from the given image and then feed them as queries to the decoder. The model is pre-trained to detect these query patches from the input image. During the pre-training, we address two critical issues: multi-task learning and multi-query localization. (1) To trade off classification and localization preferences in the pretext task, we find that freezing the CNN backbone is the prerequisite for the success of pre-training transformers. (2) To perform multi-query localization, we develop UP-DETR with multi-query patch detection with attention mask. Besides, UP-DETR also provides a unified perspective for fine-tuning object detection and one-shot detection tasks. In our experiments, UP-DETR significantly boosts the performance of DETR with faster convergence and higher average precision on object detection, one-shot detection and panoptic segmentation. Code and pre-training models: https://github.com/dddzg/up-detr.
Zhigang Dai, Bolun Cai, Yugeng Lin
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Reflective Learning With Label Noise
abstract
Learning with noisy labels is one of the most challenging tasks in semi-supervised learning, and it poses significant problems in various practical applications. In the network learning process, the noisy labels concealed in the training dataset are easy to remember, resulting in poor generalization performance. To overcome this problem, inspired by the correction ability of humans – “think and learn from the past,” an end-to-end dynamic correction framework against label noise called Reflective Learning (RL) is proposed. This solution incorporates valuable knowledge from the past network training process to assist in correcting noisy labels. Specifically, during network training, a dynamic iterative function is implemented to adaptively correct noisy labels by employing the network’s predictive distribution information of all training epochs. This dynamic iterative function takes the form of a Standard Normal Distribution function to effectively match the changes of noisy label correction information contained in the network’s predictive probabilities. The proposed method is general and applicable to any backbone network and different types of noise without auxiliary information. Experiments are conducted on datasets with synthetic and real-world label noise datasets, including CIFAR-10, CIFAR-100, Tiny-ImageNet, and Clothing1M. They demonstrate that the proposed method is superior to the state-of-the-art results.
Lin Wang 0004, Xiangmin Xu 0001, Kailing Guo, Bolun Cai, Fang Liu 0030
IEEE Trans. Circuits Syst. Video Technol.4
2022 UniMoCo: Unsupervised, Semi-Supervised and Fully-Supervised Visual Representation Learning
abstract
Momentum Contrast (MoCo) achieves great success for unsupervised visual representation learning. However, there are a lot of supervised and semi-supervised datasets, which are already labeled. To fully utilize the label annotations, we propose Unified Momentum Contrast (UniMoCo), which extends MoCo to support arbitrary ratios of labeled data and unlabeled data training. Compared with MoCo, UniMoCo has two modifications as follows: (1) Different from a single positive pair in MoCo, we maintain multiple positive pairs on-the-fly by comparing the query label to a label queue. (2) We propose a Unified Contrastive (UniCon) loss to support an arbitrary number of positives and negatives in a unified pair-wise optimization perspective. Our UniCon is more reasonable and powerful than the supervised contrastive loss in theory and practice. In our experiments, we pre-train multiple UniMoCo models with different ratios of ImageNet labels and evaluate the performance on various downstream tasks. Experiment results show that UniMoCo generalizes well for unsupervised, semi-supervised and fully-supervised visual representation learning. Besides, we surprisingly find that UniMoCo performs best with 60% ImageNet labels for COCO and VOC transfer learning. The code is available: https://github.com/dddzg/unimoco.
Zhigang Dai, Bolun Cai
SMC2
2022 ISNet: Individual Standardization Network for Speech Emotion Recognition
abstract
Speech emotion recognition plays an essential role in human-computer interaction. However, cross-individual representation learning and individual-agnostic systems are challenging due to the distribution deviation caused by individual differences. The existing related approaches mostly use the auxiliary task of speaker recognition to eliminate individual differences. Unfortunately, although these methods can reduce interindividual voiceprint differences, it is difficult to dissociate interindividual expression differences since each individual has its unique expression habits. In this paper, we propose an individual standardization network (ISNet) for speech emotion recognition to alleviate the problem of interindividual emotion confusion caused by individual differences. Specifically, we model individual benchmarks as representations of nonemotional neutral speech, and ISNet realizes individual standardization using the automatically generated benchmark, which improves the robustness of individual-agnostic emotion representations. In response to individual differences, we also propose more comprehensive and meaningful individual-level evaluation metrics. In addition, we continue our previous work to construct a challenging large-scale speech emotion dataset (LSSED). We propose a more reasonable division method of the training set and testing set to prevent individual information leakage. Experimental results on datasets of both large and small scales have proven the effectiveness of ISNet, and the new state-of-the-art performance is achieved under the same experimental conditions on IEMOCAP and LSSED.
Weiquan Fan, Xiangmin Xu 0001, Bolun Cai, Xiaofen Xing
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 UP-DETR: Unsupervised Pre-Training for Object Detection With Transformers
abstract
Object detection with transformers (DETR) reaches competitive performance with Faster R-CNN via a transformer encoder-decoder architecture. Inspired by the great success of pre-training transformers in natural language processing, we propose a pretext task named random query patch detection to Unsupervisedly Pre-train DETR (UP-DETR) for object detection. Specifically, we randomly crop patches from the given image and then feed them as queries to the decoder. The model is pre-trained to detect these query patches from the original image. During the pre-training, we address two critical issues: multi-task learning and multi-query localization. (1) To trade off classification and localization preferences in the pretext task, we freeze the CNN backbone and propose a patch feature reconstruction branch which is jointly optimized with patch detection. (2) To perform multi-query localization, we introduce UP-DETR from single-query patch and extend it to multi-query patches with object query shuffle and attention mask. In our experiments, UP-DETR significantly boosts the performance of DETR with faster convergence and higher average precision on object detection, one-shot detection and panoptic segmentation. Code and pre-training models: https://github.com/dddzg/up-detr.
Zhigang Dai, Bolun Cai, Yugeng Lin
CVPR2
2021 Spatiotemporal and frequential cascaded attention networks for speech emotion recognition
Shuzhen Li, Xiaofen Xing, Weiquan Fan, Bolun Cai, Perry Fordson, Xiangmin Xu 0001
Neurocomputing4
2021 Attention-aware concentrated network for saliency prediction
Pengqian Li, Xiaofen Xing, Xiangmin Xu 0001, Bolun Cai
Neurocomputing4
2021 Hierarchical Lifelong Learning by Sharing Representations and Integrating Hypothesis
abstract
In lifelong machine learning (LML) systems, consecutive new tasks from changing circumstances are learned and added to the system. However, sufficiently labeled data are indispensable for extracting intertask relationships before transferring knowledge in classical supervised LML systems. Inadequate labels may deteriorate the performance due to the poor initial approximation. In order to extend the typical LML system, we propose a novel hierarchical lifelong learning algorithm (HLLA) consisting of two following layers: 1) the knowledge layer consisted of shared representations and integrated knowledge basis at the bottom and 2) parameterized hypothesis functions with features at the top. Unlabeled data is leveraged in HLLA for pretraining of the shared representations. We also have considered a selective inherited updating method to deal with intertask distribution shifting. Experiments show that our HLLA method outperforms many other recent LML algorithms, especially when dealing with higher dimensional, lower correlation, and fewer labeled data problems.
Tong Zhang 0015, Guoxi Su, Chunmei Qing, Xiangmin Xu 0001, Bolun Cai, Xiaofen Xing
IEEE Trans. Syst. Man Cybern. Syst.5
2018 Perception Preserving Decolorization
abstract
Decolorization is a basic tool to transform a color image into a grayscale image, which is used in digital printing, stylized black-and-white photography, and in many single-channel image processing applications. While recent researches focus on retaining as much as possible meaningful visual features and color contrast. In this paper, we explore how to use deep neural networks for decolorization, and propose an optimization approach aiming at perception preserving. The system uses deep representations to extract content information based on human visual perception, and automatically selects suitable grayscale for decolorization. The evaluation experiments show the effectiveness of the proposed method.
Bolun Cai, Xiangmin Xu 0001, Xiaofen Xing
ICIP1
2018 Learning Adaptive Selection Network for Real-Time Visual Tracking
abstract
Offline-trained trackers based on convolutional neural networks (CNNs) have shown great potential in achieving balanced accuracy and real-time speed. However, offline-trained trackers are prone to drift to background clutters. In this paper, we present an adaptive selection network tracker (ASNT) to address the tracking drift problem. Inspired by feature selection technique used in other vision problems, we introduce a learnable selection unit for Siamese network based trackers. The selection unit enables the tracker to select relevant feature map automatically for the target. Channel dropout is applied in the selection unit to improve generalization performance for convolutional layers. To further improve the discrimination between background clutters and the target, an adaptive method is used to initialize the tracker for each video sequence. Experiments on OTB-2013 and VOT2014 datasets demonstrate that our ASNT tracker has a comparable performance against state-of-the-art methods, yet can run at a speed of over 100 fps.
Jiangfeng Xiong, Xiangmin Xu 0001, Bolun Cai, Xiaofen Xing, Kailing Guo
ICME3
2018 FReLU: Flexible Rectified Linear Units for Improving Convolutional Neural Networks
abstract
Rectified linear unit (ReLU) is a widely used activation function for deep convolutional neural networks. However, because of the zero-hard rectification, ReLU networks lose the benefits from negative values. In this paper, we propose a novel activation function called flexible rectified linear unit (FReLU) to further explore the effects of negative values. By redesigning the rectified point of ReLU as a learnable parameter, FReLU expands the states of the activation output. When a network is successfully trained, FReLU tends to converge to a negative value, which improves the expressiveness and thus the performance. Furthermore, FReLU is designed to be simple and effective without exponential functions to maintain low-cost computation. For being able to easily used in various network architectures, FReLU does not rely on strict assumptions by self-adaption. We evaluate FReLU on three standard image classification datasets, including CIFAR-10, CIFAR-100, and ImageNet. Experimental results show that FReLU achieves fast convergence and competitive performance on both plain and residual networks.
Suo Qiu, Xiangmin Xu 0001, Bolun Cai
ICPR3
2018 Visual Sentiment Analysis with Noisy Labels by Reweighting Loss
abstract
Visual sentiment analysis of online user generated content is important for many social media analysis tasks. However, label noise is common in sentiment analysis datasets, which deteriorate classification performance. To address this issue, we propose a novel visual sentiment analysis method based on loss reweighting to improve model robustness for label noise. First, a CNN is pre-trained with softmax loss on noisy labels datasets. Second, noise matrix is estimated by resorting and repositioning predicted probability, which is predicted by the pre-trained CNN. Third, converting noise estimation to the loss weight, the degeneration of sentiment classifiers performance caused by noisy labels can be compensated by re-training neural network with this reweighing loss. We conduct experiments on public sentiment datasets including Sentibank and Twitter datasets, and demonstrate that the proposed method outperforms state-of-the-art results.
Kailing Guo, Xiangmin Xu 0001, Lin Wang 0004, Bolun Cai
SMC4
2017 A Joint Intrinsic-Extrinsic Prior Model for Retinex
Bolun Cai, Xianming Xu, Kailing Guo, Kui Jia, Dacheng Tao
ICCV1
2017 Edge/structure preserving smoothing via relativity-of-Gaussian
abstract
This paper presents a novel edge/structure-preserving image smoothing via relativity-of-Gaussian. As a simple local regularization, it performs the local analysis of scale features and globally optimizes its results into a piecewise smooth. The central idea to ensure proper texture smoothing is based on cross-scale relative that captures the weak textures from the most prominent edges/structures. Our method outperforms the previous methods in removing the detail information while preserving main image content.
Bolun Cai, Xiaofen Xing, Xiangmin Xu 0001
ICIP1
2017 Multi-scale convolutional neural networks for crowd counting
abstract
Crowd counting on static images is a challenging problem due to scale variations. Recently deep neural networks have been shown to be effective in this task. However, existing neural-networks-based methods often use the multi-column or multi-network model to extract the scale-relevant features, which is more complicated for optimization and computation wasting. To this end, we propose a novel multi-scale convolutional neural network (MSCNN) for single image crowd counting. Based on the multi-scale blobs, the network is able to generate scale-relevant features for higher crowd counting performances in a single-column architecture, which is both accuracy and cost effective for practical applications. Complemental results show that our method outperforms the state-of-the-art methods on both accuracy and robustness with far less number of parameters.
Lingke Zeng, Xiangmin Xu 0001, Bolun Cai, Suo Qiu, Tong Zhang 0015
ICIP3
2016 Image and video dehazing using view-based cluster segmentation
abstract
To avoid distortion in sky regions and make the sky and white objects clear, in this paper we propose a new image and video dehazing method utilizing the view-based cluster segmentation. Firstly, GMM(Gaussian Mixture Model)is utilized to cluster the depth map based on the distant view to estimate the sky region and then the transmission estimation is modified to reduce distortion. Secondly, we present to use GMM based on Color Attenuation Prior to divide a single hazy image into K classifications, so that the atmospheric light estimation is refined to improve global contrast. Finally, online GMM cluster is applied to video dehazing. Extensive experimental results demonstrate that the proposed algorithm can have superior haze removing and color balancing capabilities.
Chunmei Qing, Xiangmin Xu 0001, Bolun Cai
VCIP4
2016 DehazeNet: An End-to-End System for Single Image Haze Removal
abstract
Single image haze removal is a challenging ill-posed problem. Existing methods use various constraints/priors to get plausible dehazing solutions. The key to achieve haze removal is to estimate a medium transmission map for an input hazy image. In this paper, we propose a trainable end-to-end system called DehazeNet, for medium transmission estimation. DehazeNet takes a hazy image as input, and outputs its medium transmission map that is subsequently used to recover a haze-free image via atmospheric scattering model. DehazeNet adopts convolutional neural network-based deep architecture, whose layers are specially designed to embody the established assumptions/priors in image dehazing. Specifically, the layers of Maxout units are used for feature extraction, which can generate almost all haze-relevant features. We also propose a novel nonlinear activation function in DehazeNet, called bilateral rectified linear unit, which is able to improve the quality of recovered haze-free image. We establish connections between the components of the proposed DehazeNet and those used in existing methods. Experiments on benchmark images show that DehazeNet achieves superior performance over existing methods, yet keeps efficient and easy to use.
Bolun Cai, Xiangmin Xu 0001, Kui Jia, Chunmei Qing, Dacheng Tao
IEEE Trans. Image Process.1
2016 BIT: Biologically Inspired Tracker
abstract
Visual tracking is challenging due to image variations caused by various factors, such as object deformation, scale change, illumination change, and occlusion. Given the superior tracking performance of human visual system (HVS), an ideal design of biologically inspired model is expected to improve computer visual tracking. This is, however, a difficult task due to the incomplete understanding of neurons' working mechanism in the HVS. This paper aims to address this challenge based on the analysis of visual cognitive mechanism of the ventral stream in the visual cortex, which simulates shallow neurons (S1 units and C1 units) to extract low-level biologically inspired features for the target appearance and imitates an advanced learning mechanism (S2 units and C2 units) to combine generative and discriminative models for target location. In addition, fast Gabor approximation and fast Fourier transform are adopted for real-time learning and detection in this framework. Extensive experiments on large-scale benchmark data sets show that the proposed biologically inspired tracker performs favorably against the state-of-the-art methods in terms of efficiency, accuracy, and robustness. The acceleration technique in particular ensures that biologically inspired tracker maintains a speed of approximately 45 frames/s.
Bolun Cai, Xiangmin Xu 0001, Xiaofen Xing, Kui Jia, Jie Miao, Dacheng Tao
IEEE Trans. Image Process.1
2015 BIT: Bio-inspired tracker
abstract
Visual tracking is a challenging problem due to various factors such as deformation, rotation and illumination. As is well known, given the superior tracking performance of human vision, bio-inspired model is expected to improve the computer visual tracking. However, the design of bio-inspired tracking framework is challenging, due to the incomplete comprehension and hyper-scale of senior neurons, which will influence the effectiveness and real-time performance of the tracker. According to the ventral stream in visual cortex, a novel bio-inspired tracker (BIT) is proposed, which simulates shallow neurons (S1 and C1) to extract low-level bio-inspired feature for target appearance and imitates senior learning mechanism (S2 and C2) to combine generative and discriminative model for position estimation. In addition, Fast Fourier Transform (FFT) is adopted for real-time learning and detection in this framework. On the recent benchmark[1], extensive experimental results show BIT performs favorably against state-of-the-art methods in terms of accuracy and robustness.
Bolun Cai, Xiangmin Xu 0001, Xiaofen Xing, Chunmei Qing
ICIP1