VLDB 2026 Research / reviewers in the wild / expert
Kailing Guo
dblp:141/9976
· DBLP profile ↗
34ranked-venue papers
5as first author
27since 2021 · last 2026
0000-0003-4753-9022ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 2 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 1 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DAVID: Dual-stage Adaptive Vision-text Integrated Decoupling for Multimodal KV Cache EvictionabstractWith the rapid development of multimodal large language models (MLLMs), deploying them on low-resource devices remains challenging. Beyond the model size, long multimodal inputs cause substantial memory overhead in the KV cache, making efficient cache management critical. In this paper, we propose DAVID, a KV cache eviction strategy that adapts to the degree of modality fusion across layers. By analyzing the feature distributions of vision and text tokens, we observe low fusion in early layers and high fusion in deeper layers. Based on this observation, DAVID adopts a decoupled eviction strategy in shallow layers and a super-modal eviction strategy in deeper layers. To support this dynamic switching, we design a lightweight metric that quantifies cross-modal fusion and uses a threshold to determine which layers require decoupling. Experimental results show that DAVID achieves state-of-the-art performance on multiple benchmarks and offers a new perspective on KV cache eviction for MLLMs. Yifeng Gu, Jianxiu Jin, Kailing Guo, Xiangmin Xu 0001 |
AAAI | 3 |
| 2026 | Random graph construction-based hierarchical attention multi-task graph convolution network for sEEG SOZ identification
Huachao Yan, Kailing Guo, Shiwei Song, Xiaofen Xing, Xiangmin Xu 0001 |
Neurocomputing | 2 |
| 2026 | Knowledge-distillation based personalized federated learning with distribution constraints
Chang Mu, Kailing Guo, Xiang Tian 0003, Xiangmin Xu 0001 |
Neural Networks | 3 |
| 2026 | Soft local reactivation for communication efficient federated learning
Chang Mu, Kailing Guo, Xiang Tian 0003, Xiangmin Xu 0001 |
Pattern Recognit. | 3 |
| 2026 | EETalk: Expression Enhancement in Speech-Driven 3D Facial AnimationabstractSpeech-driven 3D facial animation aims to generate natural and expressive facial movements from speech. Although significant progress has been made, existing methods still face challenges in generating realistic upper facial expressions. Specifically, existing methods that jointly optimize the holistic face tend to overlook fine-grained spatial movements due to motion differences across facial regions. Pre-trained speech feature extractors, which emphasize long-term dependencies, provide limited fine-grained temporal cues. In this work, we propose a novel framework, EETalk, to enhance the realism of facial expressions, which captures fine-grained spatial information and f ine-grained temporal dynamics from speech. To alleviate loss of fine-grained spatial information, we propose a novel disassemble and-reassemble modeling strategy. This strategy constructs two independent motion representation spaces for the upper and lower faces, allowing for the capture of weak upper-face movements while preserving motion diversity. Then, we propose a Cross-Region Coordination Module to ensure the synchronization and coordination of the movements from the independent upper and lower faces. To effectively capture facial micro-expressions, we incorporate fine-grained time-varying features to compensate for the short-timescale details underrepresented in the long term semantic features extracted by pre-trained self-supervised models, thereby improving the generation of fast, subtle facial expressions. Experimental results demonstrate that our approach significantly improves motion accuracy, expression consistency, and perceptual quality compared to existing methods. Zhaojie Chu, Kailing Guo, Xiaofen Xing, Bolun Cai, Lin Wang 0004, Xiangmin Xu 0001 |
IEEE Trans. Multim. | 2 |
| 2026 | PEGCL: Pseudo-Entropy Guided Complementary Learning for Robust Facial Expression Recognition Under Label NoiseabstractFacial Expression Recognition (FER) has recently plays a crucial role in advancing human-computer interaction systems, aiming to understand users' inner states and underlying intentions. However, FER in real-world scenarios remains challenging due to significant label noise, caused by ambiguous facial expressions in low-quality images and annotation bias. To tackle this issue, this paper proposes a novel framework, Pseudo-Entropy Guided Complementary Learning (PEGCL), designed to robustly handle noisy labels by leveraging complementary information, which trains networks using all complementary labels defined as “facial expression images that do not belong to complementary emotion labels.” This approach effectively utilizes non-target emotion labels to mitigate the impact of label noise, rather than relying solely on annotated emotion labels. Specifically, the proposed PEGCL framework consists of three components: logit normalization to stabilize predicted probabilities and prevent gradient explosions, transformed complementary learning to redistribute the optimization focus across complementary categories by leveraging pseudo-entropy guided, and random complementary label dropping to dynamically exclude subsets of complementary labels, enhancing generalization and preventing overfitting. These components collectively ensure robust and efficient optimization under noisy label conditions. Importantly, the proposed PEGCL does not require explicit noise estimation or complex label correction mechanisms, making it a simple and effective solution for real-world FER tasks. Extensive experiments on benchmark FER datasets demonstrate that PEGCL consistently outperforms existing methods, achieving the state-of-the-art robustness against label noise while maintaining high classification accuracy. Lin Wang 0004, Dan Liao, Fang Liu 0030, Xiangmin Xu 0001, Kailing Guo, Zhanpeng Jin |
IEEE Trans. Multim. | 5 |
| 2026 | Topology-Aware Modeling for Unsupervised Simulation-to-Reality Point Cloud RecognitionabstractLearning semantic representations from point sets of 3D object shapes is often challenged by significant geometric variations, primarily due to differences in data acquisition methods. Typically, training data is generated using point simulators, while testing data is collected with distinct 3D sensors, leading to a simulation-to-reality (Sim2Real) domain gap that limits the generalization ability of point classifiers. Current unsupervised domain adaptation (UDA) techniques struggle with this gap, as they often lack robust, domain-insensitive descriptors capable of capturing global topological information, resulting in overfitting to the limited semantic patterns of the source domain. To address this issue, we introduce a novel Topology-Aware Modeling (TAM) framework for Sim2Real UDA on object point clouds. Our approach mitigates the domain gap by leveraging global spatial topology, characterized by low-level, high-frequency 3D structures, and by modeling the topological relations of local geometric features through a novel self-supervised learning task. Additionally, we propose an advanced self-training strategy that combines cross-domain contrastive learning with self-training, effectively reducing the impact of noisy pseudo-labels and enhancing the robustness of the adaptation process. Experimental results on three public Sim2Real benchmarks validate the effectiveness of our TAM framework, showing consistent improvements over state-of-the-art methods across all evaluated tasks. The source code of this work will be available athttps://github.com/zou-longkun/TAG.git. Longkun Zou, Kangjun Liu, Ke Chen 0004, Kailing Guo, Kui Jia, Yaowei Wang 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | Facial expression recognition based on multi-task self-distillation with coarse and fine grained labels
Kailing Guo, Xiangmin Xu 0001 |
Expert Syst. Appl. | 3 |
| 2025 | Tensor self-representation network for subspace clustering via alternating direction method of multipliers
Kailing Guo, Xiangmin Xu 0001 |
Knowl. Based Syst. | 2 |
| 2025 | rPPG-TFCL: Time-frequency consistency learning for robust remote physiological measurement
Kailing Guo, Fang Liu 0030, Xiaofen Xing, Lin Wang 0004, Xiangmin Xu 0001, Zhanpeng Jin |
Knowl. Based Syst. | 2 |
| 2025 | GUS-IR: Gaussian Splatting With Unified Shading for Inverse RenderingabstractRecovering the intrinsic physical attributes of a scene from images, generally termed as the inverse rendering problem, has been a central and challenging task in computer vision and computer graphics. In this paper, we present GUS-IR, a novel framework designed to address the inverse rendering problem for complicated scenes featuring rough and glossy surfaces. This paper starts by analyzing and comparing two prominent shading techniques popularly used for inverse rendering, forward shading and deferred shading, effectiveness in handling complex materials. More importantly, we propose a unified shading solution that combines the advantages of both techniques for better decomposition. In addition, we analyze the normal modeling in 3D Gaussian Splatting (3DGS) and utilize the shortest axis as normal for each particle in GUS-IR, along with a depth-related regularization, resulting in improved geometric representation and better shape reconstruction. Furthermore, we enhance the probe-based baking scheme proposed by GS-IR to achieve more accurate ambient occlusion modeling to better handle indirect illumination. Extensive experiments have demonstrated the superior performance of GUS-IR in achieving precise intrinsic decomposition and geometric representation, supporting many downstream tasks (such as relighting, retouching) in computer vision, graphics, and extended reality. Zhihao Liang 0002, Hongdong Li, Kui Jia, Kailing Guo, Qi Zhang 0029 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | CSE-GResNet: A Simple and Highly Efficient Network for Facial Expression RecognitionabstractFacial expression recognition (FER) has recently attracted extensive attention in computer vision. However, existing methods mostly focus on the explicit performance and overlook their computational resources. Hence, achieving competitive performance while maintaining the model efficiency is still a huge challenge. To tackle these issues, we propose a highly lightweight yet effective Channel Shift-Enhancement Gabor-ResNet (CSE-GResNet) to capture the crucial visual properties in facial images. Concretely, we incorporate the Gabor Convolution (GConv) into ResNet to produce the robust GResNet as our backbone with limited memory cost. Furthermore, we propose extremely efficient Channel-Shift Module and Channel-Enhancement Module to insert in the GResNet in cascade. They are adopted to obtain and aggregate the facial informative representation from adjacent channels for extracting the subtle facial expression representation. We conduct extensive experiments on three wild datasets: RAF-DB, FER2013 and SFEW. The results show that the proposed CSE-GResNet achieves superior performance against the state-of-the-art methods with less computational and memory cost. Shaoping Jiang, Xiaofen Xing, Fang Liu 0030, Xiangmin Xu 0001, Lin Wang 0004, Kailing Guo |
IEEE Trans. Affect. Comput. | 6 |
| 2025 | Alleviating One-to-Many Mapping in Talking Head Synthesis With Dynamic Adaptation Context and Style AdapterabstractSpeech-driven talking head synthesis technology has made remarkable progress, but it still faces the challenge of one-to-many pathological mapping. The challenge results in inaccurate lip movements, ambiguity in facial expressions, and a lack of coherence during transitions between facial motions. The phenomenon is primarily caused by: (1) for one speaker, the same phoneme corresponds to a wide range of mouth shapes and facial expressions due to contextual variations, and (2) for the same spoken content, different speakers exhibit diverse facial motions as a result of unique speaking styles. In this work, we propose a novel framework, called AllTalk, to alleviate one-to-many pathological mapping, which enables a more vivid and natural talking head. Specifically, considering the asymmetry and dynamic nature of mouth shapes’ dependence on phoneme context, we propose a Dynamic Adaptive Context encoder to capture the context around the phoneme and its dynamics, thereby reducing the ambiguity in mapping speech to facial movements. Moreover, to alleviate the uncertainty caused by differences in speaking style, we propose a Style Adapter that expands a generic discrete motion space for the target speaker. The Style Adapter not only effectively represents general facial motions but also captures the personalized nuances of facial movements. To further enhance the fidelity of output, we introduce a Dynamic Gaussian Renderer based on 3D Gaussian Splatting, capable of producing stable and realistic rendering videos. Extensive qualitative and quantitative experiments demonstrate that AllTalk surpasses existing state-of-the-art methods, providing an effective solution to the challenge of one-to-many mapping. Project page: https://zjchu.github.io/projects/AllTalk. Zhaojie Chu, Kailing Guo, Xiaofen Xing, Bolun Cai, Xiangmin Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | CLIP-Vision Guided Few-Shot Metal Surface Defect RecognitionabstractMetal surface defect recognition (MSDR) based on deep learning encounters the challenge of few-shot expert-labeled data. In this study, we proposed a CLIP-vision guided self supervised learning (CVGSSL) framework for representation learning of unlabeled data, completing MSDR using few-shot labeled data. This framework initially generates rich and diverse representation information through multiple CLIP-Vs to ensure effective SSL pretraining, followed by the design of an MLP-adapter to distill knowledge and adapt these representations to recognition tasks. In addition, we constructed a self-constrained loss to address the inherent problem of intraclass and interclass distance ambiguity that causes the representation to fall into an equivocal decision margin. Following label-free pretraining of CVGSSL, the downstream model adapts to one-shot to four-shot defect recognition tasks through fine-tuning. Experimental results demonstrate that CVGSSL outperforms state-of-the-art SSL methods across three public metal surface defect datasets, with the efficacy of the approach validated through extensive ablation experiments. Tianlei Wang, Zeliang Li, Ying Xu 0005, Yikui Zhai, Xiaofen Xing, Kailing Guo, Pasquale Coscia, Angelo Genovese, Vincenzo Piuri, Fabio Scotti |
IEEE Trans. Ind. Informatics | 6 |
| 2025 | DCPTalk: Speech-Driven 3D Face Animation With Personalized Facial Dynamic Coupling PropertiesabstractSpeech-driven 3D facial animation has emerged as a hot topic. During this process, movements in different facial regions are interdependent, influenced by the intricate interactions among facial muscles, and manifest personalized differences. The existing methods typically simplify the facial animation generation task to an infinitely thin surface skin deformation without an underlying structure, thereby ignoring the intricate and personalized dynamics of facial muscle activity. These methods tend to produce static or weak upper-face animations with an average facial movement style. In this work, we propose a novel framework, called DCPTalk, to mimic the intricate dynamics of facial muscle activity and portray personalized facial animations. Based on facial dynamic coupling properties, we propose Mouth2Face to simulate the facial muscle control system, yielding realistic and coordinated facial animations evoked by mouth movements. Mouth movements are easily synthesized from speech signals due to their direct correlation with phonetic articulation and vocal tract dynamics. To further enhance the detail of facial movements, we employ surface skin deformation to refine the facial animation derived from Mouth2Face. Furthermore, personal factors, including inherent physical traits and acquired speaking styles, directly determine the uniqueness and realism of facial animations. Inherent physical traits are embedded into Mouth2Face for constructing personalized facial muscle control system, while acquired speaking styles are employed to modulate external driving signals. Extensive qualitative and quantitative experiments as well as a user study indicate that DCPTalk outperforms the existing state-of-the-art methods. Zhaojie Chu, Kailing Guo, Xiaofen Xing, Pengsheng Liu, Bolun Cai, Xiangmin Xu 0001 |
IEEE Trans. Multim. | 2 |
| 2025 | Compact Model Training by Low-Rank Projection With Energy TransferabstractLow-rankness plays an important role in traditional machine learning but is not so popular in deep learning. Most previous low-rank network compression methods compress networks by approximating pretrained models and retraining. However, the optimal solution in the Euclidean space may be quite different from the one with low-rank constraint. A well-pretrained model is not a good initialization for the model with low-rank constraints. Thus, the performance of a low-rank compressed network degrades significantly. Compared with other network compression methods such as pruning, low-rank methods attract less attention in recent years. In this article, we devise a new training method, low-rank projection with energy transfer (LRPET), that trains low-rank compressed networks from scratch and achieves competitive performance. We propose to alternately perform stochastic gradient descent training and projection of each weight matrix onto the corresponding low-rank manifold. Compared to retraining on the compact model, this enables full utilization of model capacity since solution space is relaxed back to Euclidean space after projection. The matrix energy (the sum of squares of singular values) reduction caused by projection is compensated by energy transfer. We uniformly transfer the energy of the pruned singular values to the remaining ones. We theoretically show that energy transfer eases the trend of gradient vanishing caused by projection. In modern networks, a batch normalization (BN) layer can be merged into the previous convolution layer for inference, thereby influencing the optimal low-rank approximation (LRA) of the previous layer. We propose BN rectification to cut off its effect on the optimal LRA, which further improves the performance. Comprehensive experiments on CIFAR-10 and ImageNet have justified that our method is superior to other low-rank compression methods and also outperforms recent state-of-the-art pruning methods. For object detection and semantic segmentation, our method still achieves good compression results. In addition, we combine LRPET with quantization and hashing methods and achieve even better compression than the original single method. We further apply it in Transformer-based models to demonstrate its transferability. Our code is available at https://github.com/BZQLin/LRPET. Kailing Guo, Zhenquan Lin, Canyang Chen, Xiaofen Xing, Fang Liu 0030, Xiangmin Xu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Local Reactivation for Communication Efficient Federated Learning Based on Sparse Gradient Deviation
Chang Mu, Xiang Tian 0003, Kailing Guo, Xiangmin Xu 0001 |
PRCV (4) | 4 |
| 2024 | Bridge Graph Attention Based Graph Convolution Network With Multi-Scale Transformer for EEG Emotion RecognitionabstractIn multichannel electroencephalograph (EEG) emotion recognition, most graph-based studies employ shallow graph model for spatial characteristics learning due to node over-smoothing caused by an increase in network depth. To address over-smoothing, we propose the bridge graph attention-based graph convolution network (BGAGCN). It bridges previous graph convolution layers to attention coefficients of the final layer by adaptively combining each graph convolution output based on the graph attention network, thereby enhancing feature distinctiveness. Considering that graph-based networks primarily focus on local EEG channel relationships, we introduce a transformer for global dependency. Inspired by the neuroscience finding that neural activities of different timescales reflect distinct spatial connectivities, we modify the transformer to a multi-scale transformer (MT) by applying multi-head attention to multichannel EEG signals after 1D convolutions at different scales. MT learns spatial features more elaborately to enhance feature representation ability. By combining BGAGCN and MT, our model BGAGCN-MT achieves state-of-the-art accuracy under subject-dependent and subject-independent protocols across three benchmark EEG emotion datasets (SEED, SEED-IV and DREAMER). Notably, our model effectively addresses over-smoothing in graph neural networks and provides an efficient solution to learning spatial relationships of EEG features at different scales. Our code is available athttps://github.com/LogzZ. Huachao Yan, Kailing Guo, Xiaofen Xing, Xiangmin Xu 0001 |
IEEE Trans. Affect. Comput. | 2 |
| 2024 | CorrTalk: Correlation Between Hierarchical Speech and Facial Activity Variances for 3D AnimationabstractSpeech-driven 3D facial animation is a challenging cross-modal task that has attracted growing research interest. During speaking activities, the mouth displays strong motions, while the other facial regions typically demonstrate comparatively weak activity levels. Existing approaches often simplify the process by directly mapping single-level speech features to the entire facial animation, which overlook the differences in facial activity intensity leading to overly smoothed facial movements. In this study, we propose a novel framework, CorrTalk, which effectively establishes the temporal correlation between hierarchical speech features and facial activities of different intensities across distinct regions. A novel facial activity intensity prior is defined to distinguish between strong and weak facial activity, obtained by statistically analyzing facial animations. Based on the facial activity intensity prior, we propose a dual-branch decoding framework to synchronously synthesize strong and weak facial activity, which guarantees wider intensity facial animation synthesis. Furthermore, a weighted hierarchical feature encoder is proposed to establish temporal correlation between hierarchical speech features and facial activity at different intensities, which ensures lip-sync and plausible facial expressions. Extensive qualitatively and quantitatively experiments as well as a user study indicate that our CorrTalk outperforms existing state-of-the-art methods. The source code and supplementary video are publicly available at: https://zjchu.github.io/projects/CorrTalk/. Zhaojie Chu, Kailing Guo, Xiaofen Xing, Yilin Lan, Bolun Cai, Xiangmin Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Boosting Cross-Domain Point Classification via Distilling Relational Priors From 2D TransformersabstractSemantic pattern of an object point cloud is determined by its topological configuration of local geometries. Learning discriminative representations can be challenging due to large shape variations of point sets in local regions and incomplete surface in a global perspective, which can be made even more severe in the context of unsupervised domain adaptation (UDA). In specific, traditional 3D networks mainly focus on local geometric details and ignore the topological structure between local geometries, which greatly limits their cross-domain generalization. Recently, the transformer-based models have achieved impressive performance gain in a range of image-based tasks, benefiting from its strong generalization capability and scalability stemming from capturing long range correlation across local patches. Inspired by such successes of visual transformers, we propose a novel Relational Priors Distillation (RPD) method to extract relational priors from the well-trained transformers on massive images, which can significantly empower cross-domain representations with consistent topological priors of objects. To this end, we establish a parameter-frozen pre-trained transformer module shared between 2D teacher and 3D student models, complemented by an online knowledge distillation strategy for semantically regularizing the 3D student model. Furthermore, we introduce a novel self-supervised task centered on reconstructing masked point cloud patches using corresponding masked multi-view image features, thereby empowering the model with incorporating 3D geometric information. Experiments on the PointDA-10 and the Sim-to-Real datasets verify that the proposed method consistently achieves the state-of-the-art performance of UDA for point cloud classification. The source code of this work is available athttps://github.com/zou-longkun/RPD.git. Longkun Zou, Wanru Zhu, Ke Chen 0004, Lihua Guo, Kailing Guo, Kui Jia, Yaowei Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Label-Guided Dynamic Spatial-Temporal Fusion for Video-Based Facial Expression RecognitionabstractVideo-based facial expression recognition (FER) in the wild is a common yet challenging task. Extracting spatial and temporal features simultaneously is a common approach but may not always yield optimal results due to the distinct nature of spatial and temporal information. Extracting spatial and temporal features cascadingly has been proposed as an alternative approach However, the results of video-based FER sometimes fall short compared to image-based FER, indicating underutilization of spatial information of each frame and suboptimal modeling of frame relations in spatial-temporal fusion strategies. Although frame label is highly related to video label, it is overlooked in previous video-based FER methods. This paper proposes label-guided dynamic spatial-temporal fusion (LG-DSTF) that adopts frame labels to enhance the discriminative ability of spatial features and guide temporal fusion. By assigning each frame a video label, two auxiliary classification loss functions are constructed to steer discriminative spatial feature learning at different levels. The cross entropy between a uniform distribution and label distribution of spatial features is utilized to measure the classification confidence of each frame. The confidence values serve as dynamic weights to emphasize crucial frames during temporal fusion of spatial features. Our LG-DSTF achieves state-of-the-art results on FER benchmarks. Xiang Tian 0003, Kailing Guo, Xiangmin Xu 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Cross Range Quantization for Network CompressionabstractQuantization is effective in reducing model memory and accelerating inference, and is an important way to deploy deep neural networks on mobile smart devices. However, current popular learnable quantization functions often take simple truncation operations for values beyond the quantization range. We find that the truncation operation cause information loss and and restricts the update of values out of quantization range. To address this problem, we propose a universal cross range quantization (CRQ) method to reduce the information loss caused by the conventional truncation operation. CRQ splits the values exceeding the quantization range into two parts for separate quantization, and thus retain the information efficiently for performance improvement. In addition, we define a new metric named performance improvement efficiency (PIE) to measure the relationship between increased computation and performance improvement. Experiments on public benchmark image classification datasets show that CRQ achieves a significant accuracy gain with only a small increase in computation compared to the original learnable quantization method, and also outperforms many sophisticated designed state-of-the-art quantization methods in terms of accuracy and PIE. Yicai Yang, Xiaofen Xing, Kailing Guo, Xiangmin Xu 0001, Fang Liu 0030 |
IJCNN | 4 |
| 2023 | Enhanced discriminative global-local feature learning with priority for facial expression recognition
Xiang Tian 0003, Kailing Guo, Xiangmin Xu 0001 |
Inf. Sci. | 4 |
| 2023 | Reflective Learning With Label NoiseabstractLearning with noisy labels is one of the most challenging tasks in semi-supervised learning, and it poses significant problems in various practical applications. In the network learning process, the noisy labels concealed in the training dataset are easy to remember, resulting in poor generalization performance. To overcome this problem, inspired by the correction ability of humans – “think and learn from the past,” an end-to-end dynamic correction framework against label noise called Reflective Learning (RL) is proposed. This solution incorporates valuable knowledge from the past network training process to assist in correcting noisy labels. Specifically, during network training, a dynamic iterative function is implemented to adaptively correct noisy labels by employing the network’s predictive distribution information of all training epochs. This dynamic iterative function takes the form of a Standard Normal Distribution function to effectively match the changes of noisy label correction information contained in the network’s predictive probabilities. The proposed method is general and applicable to any backbone network and different types of noise without auxiliary information. Experiments are conducted on datasets with synthetic and real-world label noise datasets, including CIFAR-10, CIFAR-100, Tiny-ImageNet, and Clothing1M. They demonstrate that the proposed method is superior to the state-of-the-art results. Lin Wang 0004, Xiangmin Xu 0001, Kailing Guo, Bolun Cai, Fang Liu 0030 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Simple-action-guided dictionary learning for complex action recognition
Fang Liu 0030, Xiangmin Xu 0001, Xiaofen Xing, Kailing Guo, Lin Wang 0004 |
Neurocomputing | 4 |
| 2021 | Pruning the Unimportant or Redundant Filters? Synergy Makes BetterabstractFilter pruning is a hot topic in convolutional neural network compression due to its friendliness to hardware implementation. Most pruning methods prune filters according to their importance, i.e., removing the filters that have little effect on the final performance of the network. While from another perspective, some recent works propose to prune upon the redundancy of filters. Filters pruned in this way usually have non-negligible effects on the final performance, whereas those effects could be compensated by the remaining filters. Since importance and redundancy pruning respectively captures the local and global information of a convolutional layer, which are mutually complementary to some extent, we propose a new pruning criterion that synergies both of those previous pruning criteria to make full use of the filter information. Comprehensive experiments on benchmark image classification datasets show the effectiveness of our proposed pruning criterion. Yucheng Cai, Zhuowen Yin, Kailing Guo, Xiangmin Xu 0001 |
IJCNN | 3 |
| 2021 | Weight Evolution: Improving Deep Neural Networks Training through Evolving Inferior Weight ValuesabstractTo obtain good performance, convolutional neural networks are usually over-parameterized. This phenomenon has stimulated two interesting topics: pruning the unimportant weights for compression and reactivating the unimportant weights to make full use of network capability. However, current weight reactivation methods usually reactivate the entire filters, which may not be precise enough. Looking back in history, the prosperity of filter pruning is mainly due to its friendliness to hardware implementation, but pruning at a finer structure level, i.e., weight elements, usually leads to better network performance. We study the problem of weight element reactivation in this paper. Motivated by evolution, we select the unimportant filters and update their unimportant elements by combining them with the important elements of important filters, just like gene crossover to produce better offspring, and the proposed method is called weight evolution (WE). WE is mainly composed of four strategies. We propose a global selection strategy and a local selection strategy and combine them to locate the unimportant filters. A forward matching strategy is proposed to find the matched important filters and a crossover strategy is proposed to utilize the important elements of the important filters for updating unimportant filters. WE is plug-in to existing network architectures. Comprehensive experiments show that WE outperforms the other reactivation methods and plug-in training methods with typical convolutional neural networks, especially lightweight networks. Our code is available at https://github.com/BZQLin/Weight-evolution. Zhenquan Lin, Kailing Guo, Xiaofen Xing, Xiangmin Xu 0001 |
ACM Multimedia | 2 |
| 2020 | Exploring privileged information from simple actions for complex action recognition
Fang Liu 0030, Xiangmin Xu 0001, Tong Zhang 0015, Kailing Guo, Lin Wang 0004 |
Neurocomputing | 4 |
| 2018 | Learning Adaptive Selection Network for Real-Time Visual TrackingabstractOffline-trained trackers based on convolutional neural networks (CNNs) have shown great potential in achieving balanced accuracy and real-time speed. However, offline-trained trackers are prone to drift to background clutters. In this paper, we present an adaptive selection network tracker (ASNT) to address the tracking drift problem. Inspired by feature selection technique used in other vision problems, we introduce a learnable selection unit for Siamese network based trackers. The selection unit enables the tracker to select relevant feature map automatically for the target. Channel dropout is applied in the selection unit to improve generalization performance for convolutional layers. To further improve the discrimination between background clutters and the target, an adaptive method is used to initialize the tracker for each video sequence. Experiments on OTB-2013 and VOT2014 datasets demonstrate that our ASNT tracker has a comparable performance against state-of-the-art methods, yet can run at a speed of over 100 fps. Jiangfeng Xiong, Xiangmin Xu 0001, Bolun Cai, Xiaofen Xing, Kailing Guo |
ICME | 5 |
| 2018 | Visual Sentiment Analysis with Noisy Labels by Reweighting LossabstractVisual sentiment analysis of online user generated content is important for many social media analysis tasks. However, label noise is common in sentiment analysis datasets, which deteriorate classification performance. To address this issue, we propose a novel visual sentiment analysis method based on loss reweighting to improve model robustness for label noise. First, a CNN is pre-trained with softmax loss on noisy labels datasets. Second, noise matrix is estimated by resorting and repositioning predicted probability, which is predicted by the pre-trained CNN. Third, converting noise estimation to the loss weight, the degeneration of sentiment classifiers performance caused by noisy labels can be compensated by re-training neural network with this reweighing loss. We conduct experiments on public sentiment datasets including Sentibank and Twitter datasets, and demonstrate that the proposed method outperforms state-of-the-art results. Kailing Guo, Xiangmin Xu 0001, Lin Wang 0004, Bolun Cai |
SMC | 1 |
| 2018 | GoDec+: Fast and Robust Low-Rank Matrix Decomposition Based on Maximum CorrentropyabstractGoDec is an efficient low-rank matrix decomposition algorithm. However, optimal performance depends on sparse errors and Gaussian noise. This paper aims to address the problem that a matrix is composed of a low-rank component and unknown corruptions. We introduce a robust local similarity measure called correntropy to describe the corruptions and, in doing so, obtain a more robust and faster low-rank decomposition algorithm: GoDec+. Based on half-quadratic optimization and greedy bilateral paradigm, we deliver a solution to the maximum correntropy criterion (MCC)-based low-rank decomposition problem. Experimental results show that GoDec+ is efficient and robust to different corruptions including Gaussian noise, Laplacian noise, salt & pepper noise, and occlusion on both synthetic and real vision data. We further apply GoDec+ to more general applications including classification and subspace clustering. For classification, we construct an ensemble subspace from the low-rank GoDec+ matrix and introduce an MCC-based classifier. For subspace clustering, we utilize GoDec+ values low-rank matrix for MCC-based self-expression and combine it with spectral clustering. Face recognition, motion segmentation, and face clustering experiments show that the proposed methods are effective and robust. In particular, we achieve the state-of-the-art performance on the Hopkins 155 data set and the first 10 subjects of extended Yale B for subspace clustering. Kailing Guo, Liu Liu 0014, Xiangmin Xu 0001, Dong Xu 0001, Dacheng Tao |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | A Joint Intrinsic-Extrinsic Prior Model for Retinex
Bolun Cai, Xianming Xu, Kailing Guo, Kui Jia, Dacheng Tao |
ICCV | 3 |
| 2015 | Particle swarm optimization based multi-domain virtual network embeddingabstractMulti-domain virtual network embedding (MVNE) aims to embed a virtual network (VN) across multiple physical domains while minimizing the embedding cost. A key phrase of MVNE is VN partitioning which partitions a VN into multiple physical domains. Since the MVNE problem is NP-hard, we provide a heuristic VN partitioning approach named VNP-PSO based on the Particle Swarm Optimization (PSO) to increase the efficiency of VN partitioning. The VNP-PSO algorithm generates a near-optimal solution of VN partitioning through the evolution process of the particles. The simulation results show that our proposal can increase the efficiency of VN partitioning and decrease the embedding cost of MVNE. Kailing Guo, Ying Wang 0002, Xuesong Qiu 0001, Wenjing Li 0001, Ailing Xiao |
IM | 1 |
| 2013 | A novel incremental weighted PCA algorithm for visual trackingabstractThis paper addresses the drifting problem in online visual tracking. The tracking result is usually described by a bounding box, which inevitably contains background in the box and causes drifting. This paper tries to treat the background part and the truly target part discriminatively to reduce the effect of background. A novel incremental weighted PCA (IWPCA) algorithm is proposed. The most important contribution of this paper is an approximation method which limits the great and increasing computational cost, caused by the weighted form, to a constant. Therefore it is suitable for online tracking. Combined with particle filter, the proposed algorithm achieves superior results in several challenging video sequences in terms of stableness and accuracy, and greatly alleviates the drifting problem. Kailing Guo, Xiangmin Xu 0001, Fuhao Qiu, Jiayong Chen |
ICIP | 1 |