EDBT 2026 Demo / reviewers in the wild / expert
Bolun Cai
dblp:124/2077
· DBLP profile ↗
24ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0002-9394-7820ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EETalk: Expression Enhancement in Speech-Driven 3D Facial AnimationabstractSpeech-driven 3D facial animation aims to generate natural and expressive facial movements from speech. Although significant progress has been made, existing methods still face challenges in generating realistic upper facial expressions. Specifically, existing methods that jointly optimize the holistic face tend to overlook fine-grained spatial movements due to motion differences across facial regions. Pre-trained speech feature extractors, which emphasize long-term dependencies, provide limited fine-grained temporal cues. In this work, we propose a novel framework, EETalk, to enhance the realism of facial expressions, which captures fine-grained spatial information and f ine-grained temporal dynamics from speech. To alleviate loss of fine-grained spatial information, we propose a novel disassemble and-reassemble modeling strategy. This strategy constructs two independent motion representation spaces for the upper and lower faces, allowing for the capture of weak upper-face movements while preserving motion diversity. Then, we propose a Cross-Region Coordination Module to ensure the synchronization and coordination of the movements from the independent upper and lower faces. To effectively capture facial micro-expressions, we incorporate fine-grained time-varying features to compensate for the short-timescale details underrepresented in the long term semantic features extracted by pre-trained self-supervised models, thereby improving the generation of fast, subtle facial expressions. Experimental results demonstrate that our approach significantly improves motion accuracy, expression consistency, and perceptual quality compared to existing methods. Zhaojie Chu, Kailing Guo, Xiaofen Xing, Bolun Cai, Lin Wang 0004, Xiangmin Xu 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | Alleviating One-to-Many Mapping in Talking Head Synthesis With Dynamic Adaptation Context and Style AdapterabstractSpeech-driven talking head synthesis technology has made remarkable progress, but it still faces the challenge of one-to-many pathological mapping. The challenge results in inaccurate lip movements, ambiguity in facial expressions, and a lack of coherence during transitions between facial motions. The phenomenon is primarily caused by: (1) for one speaker, the same phoneme corresponds to a wide range of mouth shapes and facial expressions due to contextual variations, and (2) for the same spoken content, different speakers exhibit diverse facial motions as a result of unique speaking styles. In this work, we propose a novel framework, called AllTalk, to alleviate one-to-many pathological mapping, which enables a more vivid and natural talking head. Specifically, considering the asymmetry and dynamic nature of mouth shapes’ dependence on phoneme context, we propose a Dynamic Adaptive Context encoder to capture the context around the phoneme and its dynamics, thereby reducing the ambiguity in mapping speech to facial movements. Moreover, to alleviate the uncertainty caused by differences in speaking style, we propose a Style Adapter that expands a generic discrete motion space for the target speaker. The Style Adapter not only effectively represents general facial motions but also captures the personalized nuances of facial movements. To further enhance the fidelity of output, we introduce a Dynamic Gaussian Renderer based on 3D Gaussian Splatting, capable of producing stable and realistic rendering videos. Extensive qualitative and quantitative experiments demonstrate that AllTalk surpasses existing state-of-the-art methods, providing an effective solution to the challenge of one-to-many mapping. Project page: https://zjchu.github.io/projects/AllTalk. Zhaojie Chu, Kailing Guo, Xiaofen Xing, Bolun Cai, Xiangmin Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | DCPTalk: Speech-Driven 3D Face Animation With Personalized Facial Dynamic Coupling PropertiesabstractSpeech-driven 3D facial animation has emerged as a hot topic. During this process, movements in different facial regions are interdependent, influenced by the intricate interactions among facial muscles, and manifest personalized differences. The existing methods typically simplify the facial animation generation task to an infinitely thin surface skin deformation without an underlying structure, thereby ignoring the intricate and personalized dynamics of facial muscle activity. These methods tend to produce static or weak upper-face animations with an average facial movement style. In this work, we propose a novel framework, called DCPTalk, to mimic the intricate dynamics of facial muscle activity and portray personalized facial animations. Based on facial dynamic coupling properties, we propose Mouth2Face to simulate the facial muscle control system, yielding realistic and coordinated facial animations evoked by mouth movements. Mouth movements are easily synthesized from speech signals due to their direct correlation with phonetic articulation and vocal tract dynamics. To further enhance the detail of facial movements, we employ surface skin deformation to refine the facial animation derived from Mouth2Face. Furthermore, personal factors, including inherent physical traits and acquired speaking styles, directly determine the uniqueness and realism of facial animations. Inherent physical traits are embedded into Mouth2Face for constructing personalized facial muscle control system, while acquired speaking styles are employed to modulate external driving signals. Extensive qualitative and quantitative experiments as well as a user study indicate that DCPTalk outperforms the existing state-of-the-art methods. Zhaojie Chu, Kailing Guo, Xiaofen Xing, Pengsheng Liu, Bolun Cai, Xiangmin Xu 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | CorrTalk: Correlation Between Hierarchical Speech and Facial Activity Variances for 3D AnimationabstractSpeech-driven 3D facial animation is a challenging cross-modal task that has attracted growing research interest. During speaking activities, the mouth displays strong motions, while the other facial regions typically demonstrate comparatively weak activity levels. Existing approaches often simplify the process by directly mapping single-level speech features to the entire facial animation, which overlook the differences in facial activity intensity leading to overly smoothed facial movements. In this study, we propose a novel framework, CorrTalk, which effectively establishes the temporal correlation between hierarchical speech features and facial activities of different intensities across distinct regions. A novel facial activity intensity prior is defined to distinguish between strong and weak facial activity, obtained by statistically analyzing facial animations. Based on the facial activity intensity prior, we propose a dual-branch decoding framework to synchronously synthesize strong and weak facial activity, which guarantees wider intensity facial animation synthesis. Furthermore, a weighted hierarchical feature encoder is proposed to establish temporal correlation between hierarchical speech features and facial activity at different intensities, which ensures lip-sync and plausible facial expressions. Extensive qualitatively and quantitatively experiments as well as a user study indicate that our CorrTalk outperforms existing state-of-the-art methods. The source code and supplementary video are publicly available at: https://zjchu.github.io/projects/CorrTalk/. Zhaojie Chu, Kailing Guo, Xiaofen Xing, Yilin Lan, Bolun Cai, Xiangmin Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | MGAT: Multi-Granularity Attention Based Transformers for Multi-Modal Emotion RecognitionabstractMulti-modal emotion recognition is crucial for human-computer interaction. Many existing algorithms attempt to achieve multi-modal interactions through a cross-attention mechanism. Due to the problems of noise introduction and heavy computation in the original attention mechanism, window attention has become a new trend. However, emotions are presented asynchronously between different modalities, which makes it difficult to interact with emotional information between windows. Furthermore, multi-modal data are temporally misaligned, so single fixed window size is hard to describe cross-modal information. In this paper, we put these two issues into a unified framework and propose the multi-granularity attention based Transformers (MGAT). It addresses the emotional asynchrony and modality misalignment issues through a multi-granularity attention mechanism. Experimental results confirm the effectiveness of our method and the state-of-the-art performance is achieved on IEMOCAP. Weiquan Fan, Xiaofen Xing, Bolun Cai, Xiangmin Xu 0001 |
ICASSP | 3 |
| 2023 | Unsupervised Pre-Training for Detection TransformersabstractDEtection TRansformer (DETR) for object detection reaches competitive performance compared with Faster R-CNN via a transformer encoder-decoder architecture. However, trained with scratch transformers, DETR needs large-scale training data and an extreme long training schedule even on COCO dataset. Inspired by the great success of pre-training transformers in natural language processing, we propose a novel pretext task named random query patch detection in Unsupervised Pre-training DETR (UP-DETR). Specifically, we randomly crop patches from the given image and then feed them as queries to the decoder. The model is pre-trained to detect these query patches from the input image. During the pre-training, we address two critical issues: multi-task learning and multi-query localization. (1) To trade off classification and localization preferences in the pretext task, we find that freezing the CNN backbone is the prerequisite for the success of pre-training transformers. (2) To perform multi-query localization, we develop UP-DETR with multi-query patch detection with attention mask. Besides, UP-DETR also provides a unified perspective for fine-tuning object detection and one-shot detection tasks. In our experiments, UP-DETR significantly boosts the performance of DETR with faster convergence and higher average precision on object detection, one-shot detection and panoptic segmentation. Code and pre-training models: https://github.com/dddzg/up-detr. Zhigang Dai, Bolun Cai, Yugeng Lin |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Reflective Learning With Label NoiseabstractLearning with noisy labels is one of the most challenging tasks in semi-supervised learning, and it poses significant problems in various practical applications. In the network learning process, the noisy labels concealed in the training dataset are easy to remember, resulting in poor generalization performance. To overcome this problem, inspired by the correction ability of humans – “think and learn from the past,” an end-to-end dynamic correction framework against label noise called Reflective Learning (RL) is proposed. This solution incorporates valuable knowledge from the past network training process to assist in correcting noisy labels. Specifically, during network training, a dynamic iterative function is implemented to adaptively correct noisy labels by employing the network’s predictive distribution information of all training epochs. This dynamic iterative function takes the form of a Standard Normal Distribution function to effectively match the changes of noisy label correction information contained in the network’s predictive probabilities. The proposed method is general and applicable to any backbone network and different types of noise without auxiliary information. Experiments are conducted on datasets with synthetic and real-world label noise datasets, including CIFAR-10, CIFAR-100, Tiny-ImageNet, and Clothing1M. They demonstrate that the proposed method is superior to the state-of-the-art results. Lin Wang 0004, Xiangmin Xu 0001, Kailing Guo, Bolun Cai, Fang Liu 0030 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | UniMoCo: Unsupervised, Semi-Supervised and Fully-Supervised Visual Representation LearningabstractMomentum Contrast (MoCo) achieves great success for unsupervised visual representation learning. However, there are a lot of supervised and semi-supervised datasets, which are already labeled. To fully utilize the label annotations, we propose Unified Momentum Contrast (UniMoCo), which extends MoCo to support arbitrary ratios of labeled data and unlabeled data training. Compared with MoCo, UniMoCo has two modifications as follows: (1) Different from a single positive pair in MoCo, we maintain multiple positive pairs on-the-fly by comparing the query label to a label queue. (2) We propose a Unified Contrastive (UniCon) loss to support an arbitrary number of positives and negatives in a unified pair-wise optimization perspective. Our UniCon is more reasonable and powerful than the supervised contrastive loss in theory and practice. In our experiments, we pre-train multiple UniMoCo models with different ratios of ImageNet labels and evaluate the performance on various downstream tasks. Experiment results show that UniMoCo generalizes well for unsupervised, semi-supervised and fully-supervised visual representation learning. Besides, we surprisingly find that UniMoCo performs best with 60% ImageNet labels for COCO and VOC transfer learning. The code is available: https://github.com/dddzg/unimoco. Zhigang Dai, Bolun Cai |
SMC | 2 |
| 2022 | ISNet: Individual Standardization Network for Speech Emotion RecognitionabstractSpeech emotion recognition plays an essential role in human-computer interaction. However, cross-individual representation learning and individual-agnostic systems are challenging due to the distribution deviation caused by individual differences. The existing related approaches mostly use the auxiliary task of speaker recognition to eliminate individual differences. Unfortunately, although these methods can reduce interindividual voiceprint differences, it is difficult to dissociate interindividual expression differences since each individual has its unique expression habits. In this paper, we propose an individual standardization network (ISNet) for speech emotion recognition to alleviate the problem of interindividual emotion confusion caused by individual differences. Specifically, we model individual benchmarks as representations of nonemotional neutral speech, and ISNet realizes individual standardization using the automatically generated benchmark, which improves the robustness of individual-agnostic emotion representations. In response to individual differences, we also propose more comprehensive and meaningful individual-level evaluation metrics. In addition, we continue our previous work to construct a challenging large-scale speech emotion dataset (LSSED). We propose a more reasonable division method of the training set and testing set to prevent individual information leakage. Experimental results on datasets of both large and small scales have proven the effectiveness of ISNet, and the new state-of-the-art performance is achieved under the same experimental conditions on IEMOCAP and LSSED. Weiquan Fan, Xiangmin Xu 0001, Bolun Cai, Xiaofen Xing |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | UP-DETR: Unsupervised Pre-Training for Object Detection With TransformersabstractObject detection with transformers (DETR) reaches competitive performance with Faster R-CNN via a transformer encoder-decoder architecture. Inspired by the great success of pre-training transformers in natural language processing, we propose a pretext task named random query patch detection to Unsupervisedly Pre-train DETR (UP-DETR) for object detection. Specifically, we randomly crop patches from the given image and then feed them as queries to the decoder. The model is pre-trained to detect these query patches from the original image. During the pre-training, we address two critical issues: multi-task learning and multi-query localization. (1) To trade off classification and localization preferences in the pretext task, we freeze the CNN backbone and propose a patch feature reconstruction branch which is jointly optimized with patch detection. (2) To perform multi-query localization, we introduce UP-DETR from single-query patch and extend it to multi-query patches with object query shuffle and attention mask. In our experiments, UP-DETR significantly boosts the performance of DETR with faster convergence and higher average precision on object detection, one-shot detection and panoptic segmentation. Code and pre-training models: https://github.com/dddzg/up-detr. Zhigang Dai, Bolun Cai, Yugeng Lin |
CVPR | 2 |
| 2021 | Spatiotemporal and frequential cascaded attention networks for speech emotion recognition
Shuzhen Li, Xiaofen Xing, Weiquan Fan, Bolun Cai, Perry Fordson, Xiangmin Xu 0001 |
Neurocomputing | 4 |
| 2021 | Attention-aware concentrated network for saliency prediction
Pengqian Li, Xiaofen Xing, Xiangmin Xu 0001, Bolun Cai |
Neurocomputing | 4 |
| 2021 | Hierarchical Lifelong Learning by Sharing Representations and Integrating HypothesisabstractIn lifelong machine learning (LML) systems, consecutive new tasks from changing circumstances are learned and added to the system. However, sufficiently labeled data are indispensable for extracting intertask relationships before transferring knowledge in classical supervised LML systems. Inadequate labels may deteriorate the performance due to the poor initial approximation. In order to extend the typical LML system, we propose a novel hierarchical lifelong learning algorithm (HLLA) consisting of two following layers: 1) the knowledge layer consisted of shared representations and integrated knowledge basis at the bottom and 2) parameterized hypothesis functions with features at the top. Unlabeled data is leveraged in HLLA for pretraining of the shared representations. We also have considered a selective inherited updating method to deal with intertask distribution shifting. Experiments show that our HLLA method outperforms many other recent LML algorithms, especially when dealing with higher dimensional, lower correlation, and fewer labeled data problems. Tong Zhang 0015, Guoxi Su, Chunmei Qing, Xiangmin Xu 0001, Bolun Cai, Xiaofen Xing |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2018 | Perception Preserving DecolorizationabstractDecolorization is a basic tool to transform a color image into a grayscale image, which is used in digital printing, stylized black-and-white photography, and in many single-channel image processing applications. While recent researches focus on retaining as much as possible meaningful visual features and color contrast. In this paper, we explore how to use deep neural networks for decolorization, and propose an optimization approach aiming at perception preserving. The system uses deep representations to extract content information based on human visual perception, and automatically selects suitable grayscale for decolorization. The evaluation experiments show the effectiveness of the proposed method. Bolun Cai, Xiangmin Xu 0001, Xiaofen Xing |
ICIP | 1 |
| 2018 | Learning Adaptive Selection Network for Real-Time Visual TrackingabstractOffline-trained trackers based on convolutional neural networks (CNNs) have shown great potential in achieving balanced accuracy and real-time speed. However, offline-trained trackers are prone to drift to background clutters. In this paper, we present an adaptive selection network tracker (ASNT) to address the tracking drift problem. Inspired by feature selection technique used in other vision problems, we introduce a learnable selection unit for Siamese network based trackers. The selection unit enables the tracker to select relevant feature map automatically for the target. Channel dropout is applied in the selection unit to improve generalization performance for convolutional layers. To further improve the discrimination between background clutters and the target, an adaptive method is used to initialize the tracker for each video sequence. Experiments on OTB-2013 and VOT2014 datasets demonstrate that our ASNT tracker has a comparable performance against state-of-the-art methods, yet can run at a speed of over 100 fps. Jiangfeng Xiong, Xiangmin Xu 0001, Bolun Cai, Xiaofen Xing, Kailing Guo |
ICME | 3 |
| 2018 | FReLU: Flexible Rectified Linear Units for Improving Convolutional Neural NetworksabstractRectified linear unit (ReLU) is a widely used activation function for deep convolutional neural networks. However, because of the zero-hard rectification, ReLU networks lose the benefits from negative values. In this paper, we propose a novel activation function called flexible rectified linear unit (FReLU) to further explore the effects of negative values. By redesigning the rectified point of ReLU as a learnable parameter, FReLU expands the states of the activation output. When a network is successfully trained, FReLU tends to converge to a negative value, which improves the expressiveness and thus the performance. Furthermore, FReLU is designed to be simple and effective without exponential functions to maintain low-cost computation. For being able to easily used in various network architectures, FReLU does not rely on strict assumptions by self-adaption. We evaluate FReLU on three standard image classification datasets, including CIFAR-10, CIFAR-100, and ImageNet. Experimental results show that FReLU achieves fast convergence and competitive performance on both plain and residual networks. Suo Qiu, Xiangmin Xu 0001, Bolun Cai |
ICPR | 3 |
| 2018 | Visual Sentiment Analysis with Noisy Labels by Reweighting LossabstractVisual sentiment analysis of online user generated content is important for many social media analysis tasks. However, label noise is common in sentiment analysis datasets, which deteriorate classification performance. To address this issue, we propose a novel visual sentiment analysis method based on loss reweighting to improve model robustness for label noise. First, a CNN is pre-trained with softmax loss on noisy labels datasets. Second, noise matrix is estimated by resorting and repositioning predicted probability, which is predicted by the pre-trained CNN. Third, converting noise estimation to the loss weight, the degeneration of sentiment classifiers performance caused by noisy labels can be compensated by re-training neural network with this reweighing loss. We conduct experiments on public sentiment datasets including Sentibank and Twitter datasets, and demonstrate that the proposed method outperforms state-of-the-art results. Kailing Guo, Xiangmin Xu 0001, Lin Wang 0004, Bolun Cai |
SMC | 4 |
| 2017 | A Joint Intrinsic-Extrinsic Prior Model for Retinex
Bolun Cai, Xianming Xu, Kailing Guo, Kui Jia, Dacheng Tao |
ICCV | 1 |
| 2017 | Edge/structure preserving smoothing via relativity-of-GaussianabstractThis paper presents a novel edge/structure-preserving image smoothing via relativity-of-Gaussian. As a simple local regularization, it performs the local analysis of scale features and globally optimizes its results into a piecewise smooth. The central idea to ensure proper texture smoothing is based on cross-scale relative that captures the weak textures from the most prominent edges/structures. Our method outperforms the previous methods in removing the detail information while preserving main image content. Bolun Cai, Xiaofen Xing, Xiangmin Xu 0001 |
ICIP | 1 |
| 2017 | Multi-scale convolutional neural networks for crowd countingabstractCrowd counting on static images is a challenging problem due to scale variations. Recently deep neural networks have been shown to be effective in this task. However, existing neural-networks-based methods often use the multi-column or multi-network model to extract the scale-relevant features, which is more complicated for optimization and computation wasting. To this end, we propose a novel multi-scale convolutional neural network (MSCNN) for single image crowd counting. Based on the multi-scale blobs, the network is able to generate scale-relevant features for higher crowd counting performances in a single-column architecture, which is both accuracy and cost effective for practical applications. Complemental results show that our method outperforms the state-of-the-art methods on both accuracy and robustness with far less number of parameters. Lingke Zeng, Xiangmin Xu 0001, Bolun Cai, Suo Qiu, Tong Zhang 0015 |
ICIP | 3 |
| 2016 | Image and video dehazing using view-based cluster segmentationabstractTo avoid distortion in sky regions and make the sky and white objects clear, in this paper we propose a new image and video dehazing method utilizing the view-based cluster segmentation. Firstly, GMM(Gaussian Mixture Model)is utilized to cluster the depth map based on the distant view to estimate the sky region and then the transmission estimation is modified to reduce distortion. Secondly, we present to use GMM based on Color Attenuation Prior to divide a single hazy image into K classifications, so that the atmospheric light estimation is refined to improve global contrast. Finally, online GMM cluster is applied to video dehazing. Extensive experimental results demonstrate that the proposed algorithm can have superior haze removing and color balancing capabilities. Chunmei Qing, Xiangmin Xu 0001, Bolun Cai |
VCIP | 4 |
| 2016 | DehazeNet: An End-to-End System for Single Image Haze RemovalabstractSingle image haze removal is a challenging ill-posed problem. Existing methods use various constraints/priors to get plausible dehazing solutions. The key to achieve haze removal is to estimate a medium transmission map for an input hazy image. In this paper, we propose a trainable end-to-end system called DehazeNet, for medium transmission estimation. DehazeNet takes a hazy image as input, and outputs its medium transmission map that is subsequently used to recover a haze-free image via atmospheric scattering model. DehazeNet adopts convolutional neural network-based deep architecture, whose layers are specially designed to embody the established assumptions/priors in image dehazing. Specifically, the layers of Maxout units are used for feature extraction, which can generate almost all haze-relevant features. We also propose a novel nonlinear activation function in DehazeNet, called bilateral rectified linear unit, which is able to improve the quality of recovered haze-free image. We establish connections between the components of the proposed DehazeNet and those used in existing methods. Experiments on benchmark images show that DehazeNet achieves superior performance over existing methods, yet keeps efficient and easy to use. Bolun Cai, Xiangmin Xu 0001, Kui Jia, Chunmei Qing, Dacheng Tao |
IEEE Trans. Image Process. | 1 |
| 2016 | BIT: Biologically Inspired TrackerabstractVisual tracking is challenging due to image variations caused by various factors, such as object deformation, scale change, illumination change, and occlusion. Given the superior tracking performance of human visual system (HVS), an ideal design of biologically inspired model is expected to improve computer visual tracking. This is, however, a difficult task due to the incomplete understanding of neurons' working mechanism in the HVS. This paper aims to address this challenge based on the analysis of visual cognitive mechanism of the ventral stream in the visual cortex, which simulates shallow neurons (S1 units and C1 units) to extract low-level biologically inspired features for the target appearance and imitates an advanced learning mechanism (S2 units and C2 units) to combine generative and discriminative models for target location. In addition, fast Gabor approximation and fast Fourier transform are adopted for real-time learning and detection in this framework. Extensive experiments on large-scale benchmark data sets show that the proposed biologically inspired tracker performs favorably against the state-of-the-art methods in terms of efficiency, accuracy, and robustness. The acceleration technique in particular ensures that biologically inspired tracker maintains a speed of approximately 45 frames/s. Bolun Cai, Xiangmin Xu 0001, Xiaofen Xing, Kui Jia, Jie Miao, Dacheng Tao |
IEEE Trans. Image Process. | 1 |
| 2015 | BIT: Bio-inspired trackerabstractVisual tracking is a challenging problem due to various factors such as deformation, rotation and illumination. As is well known, given the superior tracking performance of human vision, bio-inspired model is expected to improve the computer visual tracking. However, the design of bio-inspired tracking framework is challenging, due to the incomplete comprehension and hyper-scale of senior neurons, which will influence the effectiveness and real-time performance of the tracker. According to the ventral stream in visual cortex, a novel bio-inspired tracker (BIT) is proposed, which simulates shallow neurons (S1 and C1) to extract low-level bio-inspired feature for target appearance and imitates senior learning mechanism (S2 and C2) to combine generative and discriminative model for position estimation. In addition, Fast Fourier Transform (FFT) is adopted for real-time learning and detection in this framework. On the recent benchmark[1], extensive experimental results show BIT performs favorably against state-of-the-art methods in terms of accuracy and robustness. Bolun Cai, Xiangmin Xu 0001, Xiaofen Xing, Chunmei Qing |
ICIP | 1 |