Baosheng Yu

dblp:178/8725 · DBLP profile ↗
← Back
77ranked-venue papers
7as first author
68since 2021 · last 2026
0000-0002-0761-7893ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 48 · 6 first-author · 42 since 2021Graphics, computer vision, multimedia, augmented reality and games · 44 · 5 first-author · 37 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Computer networks · 3 · 2 since 2021Security and privacy · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Cross-Sample Augmented Test-Time Adaptation for Personalized Intraoperative Hypotension Prediction
abstract
Intraoperative hypotension (IOH) poses significant surgical risks, but accurate prediction remains challenging due to patient-specific variability. While test-time adaptation (TTA) offers a promising approach for personalized prediction, the rarity of IOH events often leads to unreliable test-time training. To address this, we propose CSA-TTA, a novel cross-sample augmented test-time adaptation framework that enhances training by incorporating hypotension events from other individuals. Specifically, we first construct a cross-sample bank by segmenting historical data into hypotensive and non-hypotensive samples. Then, we introduce a coarse-to-fine retrieval strategy for building test-time training data: we initially apply K-Shape clustering to identify representative cluster centers and subsequently retrieve the top-K semantically similar samples based on the current patient signal. Additionally, we integrate both self-supervised masked reconstruction and retrospective sequence forecasting signals during training to enhance model adaptability to rapid and subtle intraoperative dynamics. We evaluate the proposed CSA-TTA on both the VitalDB dataset and a real-world in-hospital dataset by integrating it with state-of-the-art time series forecasting models, including TimesFM and UniTS. CSA-TTA consistently enhances performance across settings—for instance, on VitalDB, it improves Recall and F1 scores by +1.33% and +1.13%, respectively, under fine-tuning, and by +7.46% and +5.07% in zero-shot scenarios—demonstrating strong robustness and generalization.
Kanxue Li, Yibing Zhan, Chongchong Qi, Baosheng Yu
AAAI6
2026 Remodeling Semantic Relationships in Vision-Language Fine-Tuning
abstract
Vision-language fine-tuning has emerged as an efficient paradigm for constructing multimodal foundation models. While textual context often highlights semantic relationships within an image, existing fine-tuning methods typically overlook this information when aligning vision and language, thus leading to suboptimal performance. Toward solving this problem, we propose a method that can improve multimodal alignment and fusion based on both semantics and relationships.Specifically, we first extract multilevel semantic features from different vision encoder to capture more visual cues of the relationships. Then, we learn to project the vision features to group related semantics, among which are more likely to have relationships. Finally, we fuse the visual features with the textual by using inheritable cross-attention, where we globally remove the redundant visual relationships by discarding visual-language feature pairs with low correlation. We evaluate our proposed method on eight foundation models and two downstream tasks, visual question answering and image captioning, and show that it outperforms all existing methods.
Liu Liu 0014, Baosheng Yu, Jiayan Qiu
AAAI3
2026 Controllable Contamination Detection for Reliable LLM Evaluation with Statistical Guarantees
abstract
Zheng Zhang, Qi Liu, Siyuan Liang, Ning Li, Zirui Hu, Weibo Gao, Rui Li, Zhenya Huang, Leszek Rutkowski, Baosheng Yu, Dacheng Tao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zheng Zhang 0048, Qi Liu 0003, Siyuan Liang 0004, Ning Li 0055, Zirui Hu, Weibo Gao, Rui Li 0093, Zhenya Huang, Leszek Rutkowski, Baosheng Yu, Dacheng Tao
ACL (1)10
2026 Deep Learning-Based Point Cloud Registration: A Comprehensive Survey and Taxonomy
Yu-Xin Zhang 0004, Jie Gui, Baosheng Yu, Xiaofeng Cong, Xin Gong 0001, Wenbing Tao, Dacheng Tao
Int. J. Comput. Vis.3
2026 Evaluating large language models for real-world perioperative clinical consultation
Yibing Zhan, Baosheng Yu, Pingbo Xu, Lijing Chen, Chong Zhang 0013, Chengli Zhou, Xiongbin Wang, Dapeng Tao
Neurocomputing3
2026 Weighted sample correlation knowledge distillation for visual recognition
Daidai Liu, Jianping Gou, Baosheng Yu, Zhang Yi 0001
Pattern Recognit.5
2026 Mind the data: Evaluating data quality sensitivity in medical LLMs
Xiaodong Han, Yibing Zhan, Baosheng Yu, Dapeng Tao
Pattern Recognit. Lett.4
2026 Motion-Guided Disentanglement for Point Cloud Masked Autoencoders
abstract
Masked autoencoders have been extended beyond images, but random masking often fails to capture non-uniform motion regions in inherently disordered and irregular data like point cloud videos. In this paper, we propose a motion-guided disentanglement method (MGD) to improve masked autoencoders for point cloud video representation learning. Specifically, we begin by estimating motion intensity using an optimal transport approach, which guides the separate masking of dynamic and static regions. This motion-guided masking ensures balanced coverage, addressing the limitations of random masking in capturing non-uniformly distributed motion regions. Furthermore, we disentangle the prediction tasks into motion prediction for high-motion point tubes and appearance reconstruction for low-motion ones. This disentanglement enables the model to more effectively capture both motion and appearance in point cloud videos. We conducted experiments on four widely used point cloud video datasets—NTU RGB+D, MSR-Action3D, NvGesture, and SHREC’17—which demonstrate that our approach consistently improves masked autoencoders for point cloud video representation learning, achieving new state-of-the-art results. Code will be publicly available on GitHub.
Haoran Wang 0001, Shaqing Song, Baosheng Yu, Tong Jia 0001, Dongyue Chen 0001, Chunfeng Yuan, Weiming Hu 0004, Haibin Ling
IEEE Trans. Circuits Syst. Video Technol.3
2026 Axial-View-Oriented Contrastive Adversarial Training for Robust Point Cloud Recognition
abstract
Contrastive adversarial training emerges as an effective approach to enhancing model robustness in safety-critical applications, particularly point cloud recognition for autonomous driving and medical imaging. However, existing point cloud adversarial training methods mainly emphasize global contrastive learning while overlooking local geometric variations induced by adversarial perturbations. Motivated by the spatial and intensity variations of perturbations across axial views, we propose AVOC, a novel local-global adversarial training framework that utilizes axial-view-oriented contrastive learning. This framework leverages the smallest axial view for local contrastive learning, as it exhibits the highest perturbation differences, and utilizes the largest axial view for global contrastive learning, as it preserves global structural consistency. We conduct comprehensive experiments across four representative architectures, demonstrating significant robustness improvements on widely-adopted recognition benchmarks, including ModelNet40, ShapeNetPart, ModelNet40-C, and ScanObjectNN-C, and further validate its effectiveness on the large-scale KITTI benchmark for 3D object detection. Our results across diverse perturbation scenarios, encompassing white-box attacks, black-box attacks, and natural perturbations, demonstrate the consistent and significant model robustness enhancement of our proposed method.
Jie Gui, Yu-Xin Zhang 0004, Xiaofeng Cong, Baosheng Yu, Zhipeng Gui, Yuan Yan Tang, James T. Kwok
IEEE Trans. Inf. Forensics Secur.4
2026 The Other Side of the Coin: Exploring Fairness in Retrieval-Augmented Generation
abstract
Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by retrieving relevant document from external knowledge sources. By referencing this external knowledge, RAG effectively reduces the generation of factually incorrect content and addresses hallucination issues within LLMs. Recently, there has been growing attention to improving the performance and efficiency of RAG systems from various perspectives. While these advancements have yielded significant results, the application of RAG in domains with considerable societal implications raises a critical question about fairness: What impact does the introduction of the RAG paradigm have on the fairness of LLMs? To address this question, we conduct extensive experiments by varying the LLMs, retrievers, and retrieval sources. Our experimental analysis reveals that the scale of the LLMs plays a significant role in influencing fairness outcomes within the RAG framework. When the model scale is smaller than 8B, the integration of retrieval mechanisms often exacerbates unfairness in small-scale LLMs (e.g., LLaMA3.2-1B, Mistral-7B, and LLaMA3-8B). To mitigate the fairness issues introduced by RAG for small-scale LLMs, we propose two approaches, FairFT and FairFilter. Specifically, in FairFT, we align the retriever with the LLM in terms of fairness, enabling it to retrieve documents that facilitate fairer model outputs. In FairFilter, we propose a fairness filtering mechanism to filter out biased content after retrieval. Finally, we validate our proposed approaches on real-world datasets, demonstrating their effectiveness in improving fairness while maintaining performance.
Zheng Zhang 0048, Ning Li 0055, Qi Liu 0003, Rui Li 0093, Weibo Gao, Qingyang Mao, Zhenya Huang, Baosheng Yu, Dacheng Tao
IEEE Trans. Knowl. Data Eng.8
2025 Self-Supervised Learning for Detecting AI-Generated Faces as Anomalies
abstract
The detection of AI-generated faces is commonly approached as a binary classification task. Nevertheless, the resulting detectors frequently struggle to adapt to novel AI face generators, which evolve rapidly. In this paper, we describe an anomaly detection method for AI-generated faces by leveraging self-supervised learning of camera-intrinsic and face-specific features purely from photographic face images. The success of our method lies in designing a pretext task that trains a feature extractor to rank four ordinal exchangeable image file format (EXIF) tags and classify artificially manipulated face images. Subsequently, we model the learned feature distribution of photographic face images using a Gaussian mixture model. Faces with low likelihoods are flagged as AI-generated. Both quantitative and qualitative experiments validate the effectiveness of our method. Our code is available at https://github.com/MZMMSEC/AIGFD_EXIF.git.
Mian Zou, Baosheng Yu, Yibing Zhan, Kede Ma
ICASSP2
2025 Bi-Level Optimization for Self-Supervised AI-Generated Face Detection
Mian Zou, Nan Zhong, Baosheng Yu, Yibing Zhan, Kede Ma
ICCV3
2025 SkipNode: On Alleviating Performance Degradation for Deep Graph Convolutional Networks (Extended Abstract)
abstract
Graph Convolutional Networks (GCNs) are powerful tools for learning representations in graph-structured data. However, their performance tends to degrade with increased model depth due to over-smoothing. Although previous studies attribute degradation to over-smoothing, this work identifies the mutually reinforcing effects of over-smoothing and gradient vanishing as the root cause. In this paper, we propose SkipNode, a plug-and-play module that mitigates degradation in deep GCNs. SkipNode introduces node-sampling in each convolutional layer to selectively skip convolutions, preventing over-smoothing by reducing the depth experienced by specific nodes and facilitating gradient backpropagation. We demonstrate both theoretically and experimentally that SkipNode effectively curtails over-smoothing and gradient vanishing, improving deep GCN performance across diverse tasks. Extensive evaluations show SkipNode's robustness and superior performance over state-of-the-art (SOTA) baselines, establishing it as a practical solution for training deep GCNs.
Weigang Lu 0001, Yibing Zhan, Binbin Lin 0001, Ziyu Guan, Liu Liu 0014, Baosheng Yu, Wei Zhao 0019, Yaming Yang 0002, Dacheng Tao
ICDE6
2025 FingerVeinSyn-5M: A Million-Scale Dataset and Benchmark for Finger Vein Recognition
abstract
A major challenge in finger vein recognition is the lack of large-scale public datasets. Existing datasets contain few identities and limited samples per finger, restricting the advancement of deep learning-based methods. To address this, we introduce FVeinSyn, a synthetic generator capable of producing diverse finger vein patterns with rich intra-class variations. Using FVeinSyn, we created FingerVeinSyn-5M -- the largest available finger vein dataset -- containing 5 million samples from 50,000 unique fingers, each with 100 variations including shift, rotation, scale, roll, varying exposure levels, skin scattering blur, optical blur, and motion blur. FingerVeinSyn-5M is also the first to offer fully annotated finger vein images, supporting deep learning applications in this field. Models pretrained on FingerVeinSyn-5M and fine-tuned with minimal real data achieve an average 53.91% performance gain across multiple benchmarks. The dataset is publicly available at: https://github.com/EvanWang98/FingerVeinSyn-5M.
Yifan Wang 0036, Jie Gui, Baosheng Yu, Qi Li 0005, Zhenan Sun, Juho Kannala, Guoying Zhao 0001
ACM Multimedia3
2025 SIGMA: Refining Large Language Model Reasoning via Sibling-Guided Monte Carlo Augmentation
abstract
Enhancing large language models by simply scaling up datasets has begun to yield diminishing returns, shifting the spotlight to data quality. Monte Carlo Tree Search (MCTS) has emerged as a powerful technique for generating high-quality chain-of-thought data, yet conventional approaches typically retain only the top-scoring trajectory from the search tree, discarding sibling nodes that often contain valuable partial insights, recurrent error patterns, and alternative reasoning strategies. This unconditional rejection of non-optimal reasoning branches may waste vast amounts of informative data in the whole search tree. We propose SIGMA (Sibling Guided Monte Carlo Augmentation), a novel framework that reintegrates these discarded sibling nodes to refine LLM reasoning. SIGMA forges semantic links among sibling nodes along each search path and applies a two-stage refinement: a critique model identifies overlooked strengths and weaknesses across the sibling set, and a revision model conducts text-based backpropagation to refine the top-scoring trajectory in light of this comparative feedback. By recovering and amplifying the underutilized but valuable signals from non-optimal reasoning branches, SIGMA substantially improves reasoning trajectories. On the challenging MATH benchmark, our SIGMA-tuned 7B model achieves 54.92\% accuracy using only 30K samples, outperforming state-of-the-art models trained on 590K samples. This result highlights that our sibling-guided optimization not only significantly reduces data usage but also significantly boosts LLM reasoning.
Yanwei Ren, Fuxiang Wu, Jiayan Qiu, Jiaxing Huang 0001, Baosheng Yu, Liu Liu 0014
NeurIPS6
2025 Pseudo Contrastive Learning for graph-based semi-supervised learning
Weigang Lu 0001, Ziyu Guan, Wei Zhao 0019, Yaming Yang 0002, Yuanhai Lv, Baosheng Yu, Dacheng Tao
Neurocomputing7
2025 Hypnos: A domain-specific large language model for anesthesiology
Zhonghai Wang, Yibing Zhan, Bohao Zhou, Chong Zhang 0013, Baosheng Yu, Liang Ding 0006, Weifeng Liu 0001
Neurocomputing7
2025 Intra-class progressive and adaptive self-distillation
Jianping Gou, Jiaye Lin, Weihua Ou, Baosheng Yu, Zhang Yi 0001
Neural Networks5
2025 Neighborhood relation-based knowledge distillation for image classification
Jianping Gou, Xiaomeng Xin, Baosheng Yu, Heping Song, Weiyong Zhang, Shaohua Wan 0001
Neural Networks3
2025 Learning to Explore Sample Relationships
abstract
Despite the great success achieved, deep learning technologies usually suffer from data scarcity issues in real-world applications, where existing methods mainly explore sample relationships in a vanilla way from the perspectives of either the input or the loss function. In this paper, we propose a batch transformer module, BatchFormerV1, to equip deep neural networks themselves with the abilities to explore sample relationships in a learnable way. Basically, the proposed method enables data collaboration, e.g., head-class samples will also contribute to the learning of tail classes. Considering that exploring instance-level relationships has very limited impacts on dense prediction, we generalize and refer to the proposed module as BatchFormerV2, which further enables exploring sample relationships for pixel-/patch-level dense representations. In addition, to address the train-test inconsistency where a mini-batch of data samples are neither necessary nor desirable during inference, we also devise a two-stream training pipeline, i.e., a shared model is first jointly optimized with and without BatchFormerV2 which is then removed during testing. The proposed module is plug-and-play without requiring any extra inference cost. Lastly, we evaluate the proposed method on over ten popular datasets, including 1) different data scarcity settings such as long-tailed recognition, zero-shot learning, domain generalization, and contrastive learning; and 2) different visual recognition tasks ranging from image classification to object detection and panoptic segmentation.
Zhi Hou, Baosheng Yu, Yibing Zhan, Dacheng Tao
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Distilling interaction knowledge for semi-supervised egocentric action recognition
Haoran Wang 0001, Baosheng Yu, Yibing Zhan, Dapeng Tao, Haibin Ling
Pattern Recognit.3
2025 Graph Convolutional Networks With Collaborative Feature Fusion for Sequential Recommendation
abstract
Sequential recommendation seeks to understand user preferences based on their past actions and predict future interactions with items. Recently, several techniques for sequential recommendation have emerged, primarily leveraging graph convolutional networks (GCNs) for their ability to model relationships effectively. However, real-world scenarios often involve sparse interactions, where early and recent short-term preferences play distinct roles in the recommendation process. Consequently, vanilla GCNs struggle to effectively capture the explicit correlations between these early and recent short-term preferences. To address these challenges, we introduce a novel approach termed Graph Convolutional Networks with Collaborative Feature Fusion (COFF). Specifically, our method addresses the issue by initially dividing each user interaction sequence into two segments. We then construct two separate graphs for these segments, aiming to capture the user's early and recent short-term preferences independently. To obtain robust prediction, we employ multiple GCNs in a collaborative distillation manner, incorporating a feature fusion module to establish connections between the early and recent short-term preferences. This approach enables a more precise representation of user preferences. Experimental evaluations conducted on five popular sequential recommendation datasets demonstrate that our COFF model outperforms recent state-of-the-art methods in terms of recommendation accuracy.
Jianping Gou, Youhui Cheng, Yibing Zhan, Baosheng Yu, Weihua Ou, Zhang Yi 0001
IEEE Trans. Big Data4
2025 Semantics-Oriented Multitask Learning for DeepFake Detection: A Joint Embedding Approach
abstract
In recent years, the multimedia forensics and security community has seen remarkable progress in multitask learning for DeepFake (i.e., face forgery) detection. The prevailing approach has been to frame DeepFake detection as a binary classification problem augmented by manipulation-oriented auxiliary tasks. This scheme focuses on learning features specific to face manipulations with limited generalizability. In this paper, we delve deeper into semantics-oriented multitask learning for DeepFake detection, capturing the relationships among face semantics via joint embedding. We first propose an automated dataset expansion technique that broadens current face forgery datasets to support semantics-oriented DeepFake detection tasks at both the global face attribute and local face region levels. Furthermore, we resort to the joint embedding of face images and labels (depicted by text descriptions) for prediction. This approach eliminates the need for manually setting task-agnostic and task-specific parameters, which is typically required when predicting multiple labels directly from images. In addition, we employ bi-level optimization to dynamically balance the fidelity loss weightings of various tasks, making the training process fully automated. Extensive experiments on six DeepFake datasets show that our method improves the generalizability of DeepFake detection and renders some degree of model interpretation by providing human-understandable explanations.
Mian Zou, Baosheng Yu, Yibing Zhan, Siwei Lyu, Kede Ma
IEEE Trans. Circuits Syst. Video Technol.2
2025 Semantic Contextualization of Face Forgery: A New Definition, Dataset, and Detection Method
abstract
In recent years, deep learning has greatly streamlined the process of manipulating photographic face images. Aware of the potential dangers, researchers have developed various tools to spot these counterfeits. Yet, none asks the fundamental question:What digital manipulations make a real photographic face image fake, while others do not? In this paper, we put face forgery in a semantic context and define thatcomputational methods that alter semantic face attributes to exceed human discrimination thresholds are sources of face forgery. Following our definition, we construct a large face forgery image dataset, where each image is associated with a set of labels organized in a hierarchical graph. Our dataset enables two new testing protocols to probe the generalizability of face forgery detectors. Moreover, we propose a semantics-oriented face forgery detection method that captures label relations and prioritizes the primary task (i.e., real or fake face detection). We show that the proposed dataset successfully exposes the weaknesses of current detectors as the test set and consistently improves their generalizability as the training set. Additionally, we demonstrate the superiority of our semantics-oriented method over traditional binary and multi-class classification-based detectors.
Mian Zou, Baosheng Yu, Yibing Zhan, Siwei Lyu, Kede Ma
IEEE Trans. Inf. Forensics Secur.2
2025 MIFNet: Learning Modality-Invariant Features for Generalizable Multimodal Image Matching
abstract
Many keypoint detection and description methods have been proposed for image matching or registration. While these methods demonstrate promising performance for single-modality image matching, they often struggle with multimodal data because the descriptors trained on single-modality data tend to lack robustness against the non-linear variations present in multimodal data. Extending such methods to multimodal image matching often requires well-aligned multimodal data to learn modality-invariant descriptors. However, acquiring such data is often costly and impractical in many real-world scenarios. To address this challenge, we propose a modality-invariant feature learning network (MIFNet) to compute modality-invariant features for keypoint descriptions in multimodal image matching using only single-modality training data. Specifically, we propose a novel latent feature aggregation module and a cumulative hybrid aggregation module to enhance the base keypoint descriptors trained on single-modality data by leveraging pre-trained features from Stable Diffusion models. We validate our method with recent keypoint detection and description methods in three multimodal retinal image datasets (CF-FA, CF-OCT, EMA-OCTA) and two remote sensing datasets (Optical-SAR and Optical-NIR). Extensive experiments demonstrate that the proposed MIFNet is able to learn modality-invariant feature for multimodal image matching without accessing the targeted modality and has good zero-shot generalization ability. The code will be released at https://github.com/lyp-deeplearning/MIFNet.
Yepeng Liu 0002, Zhichao Sun 0004, Baosheng Yu, Yitian Zhao, Bo Du 0001, Yongchao Xu, Jun Cheng 0003
IEEE Trans. Image Process.3
2025 PointWavelet: Learning in Spectral Domain for 3-D Point Cloud Analysis
abstract
With recent success of deep learning in 2-D visual recognition, deep-learning-based 3-D point cloud analysis has received increasing attention from the community, especially due to the rapid development of autonomous driving technologies. However, most existing methods directly learn point features in the spatial domain, leaving the local structures in the spectral domain poorly investigated. In this article, we introduce a new method, PointWavelet, to explore local graphs in the spectral domain via a learnable graph wavelet transform. Specifically, we first introduce the graph wavelet transform to form multiscale spectral graph convolution to learn effective local structural representations. To avoid the time-consuming spectral decomposition, we then devise a learnable graph wavelet transform, which significantly accelerates the overall training process. Extensive experiments on four popular point cloud datasets, ModelNet40, ScanObjectNN, ShapeNet-Part, and S3DIS, demonstrate the effectiveness of the proposed method on point cloud classification and segmentation.
Cheng Wen 0001, Jianzhi Long, Baosheng Yu, Dacheng Tao
IEEE Trans. Neural Networks Learn. Syst.3
2024 Progressive Retinal Image Registration via Global and Local Deformable Transformations
abstract
Retinal image registration plays an important role in the ophthalmological diagnosis process. Since there exist variances in viewing angles and anatomical structures across different retinal images, keypoint-based approaches become the mainstream methods for retinal image registration thanks to their robustness and low latency. These methods typically assume the retinal surfaces are planar, and adopt feature matching to obtain the homography matrix that represents the global transformation between images. Yet, such a planar hypothesis inevitably introduces registration errors since retinal surface is approximately curved. This limitation is more prominent when registering image pairs with significant differences in viewing angles. To address this problem, we propose a hybrid registration framework called HybridRetina, which progressively registers retinal images with global and local deformable transformations. For that, we use a keypoint detector and a deformation network called GAMorph to estimate the global transformation and local deformable transformation, respectively. Specifically, we integrate multi-level pixel relation knowledge to guide the training of GAMorph. Additionally, we utilize an edge attention module that includes the geometric priors of the images, ensuring the deformation field focuses more on the vascular regions of clinical interest. Experiments on two widely-used datasets, FIRE and FLoRI21, show that our proposed HybridRetina significantly outperforms some state-of-the-art methods. The code is available at https://github.com/lyp-deeplearning/awesome-retinal-registration.
Yepeng Liu 0002, Baosheng Yu, Yuliang Gu, Bo Du 0001, Yongchao Xu, Jun Cheng 0003
BIBM2
2024 Visual Relationship Transformation
Jiayan Qiu, Baosheng Yu
ECCV (65)3
2024 MuEP: A Multimodal Benchmark for Embodied Planning with Foundation Models
Kanxue Li, Baosheng Yu, Yibing Zhan, Qiong Cao, Li Shen 0008, Lusong Li, Dapeng Tao, Xiaodong He 0001
IJCAI2
2024 Self-Distillation via Intra-Class Compactness
Jiaye Lin, Baosheng Yu, Weihua Ou, Jianping Gou
PRCV (1)3
2024 Free-Form Composition Networks for Egocentric Action Recognition
abstract
Egocentric action recognition is gaining significant attention in the field of human action recognition. In this paper, we address data scarcity issue in egocentric action recognition from a compositional generalization perspective. To tackle this problem, we propose a free-form composition network (FFCN) that can simultaneously learn disentangled verb, preposition, and noun representations, and then use them to compose new samples in the feature space for rare classes of action videos. First, we use a graph to capture the spatial-temporal relations among different hand/object instances in each action video. We thus decompose each action into a set of verb and preposition spatial-temporal representations using the edge features in the graph. The temporal decomposition extracts verb and preposition representations from different video frames, while the spatial decomposition adaptively learns verb and preposition representations from action-related instances in each frame. With these spatial-temporal representations of verbs and prepositions, we can compose new samples for those rare classes in a free-form manner, which is not restricted to a rigid form of a verb and a noun. The proposed FFCN can directly generate new training data samples for rare classes, hence significantly improve action recognition performance. We evaluated our method on three popular egocentric action recognition datasets, Something-Something V2, H2O, and EPIC-KITCHENS-100, and the experimental results demonstrate the effectiveness of the proposed method for handling data scarcity problems, including long-tailed and few-shot egocentric action recognition.
Haoran Wang 0001, Qinghua Cheng, Baosheng Yu, Yibing Zhan, Dapeng Tao, Liang Ding 0006, Haibin Ling
IEEE Trans. Circuits Syst. Video Technol.3
2024 SkipNode: On Alleviating Performance Degradation for Deep Graph Convolutional Networks
abstract
Graph Convolutional Networks (GCNs) suffer from performance degradation when models go deeper. However, earlier works only attributed the performance degeneration to over-smoothing. In this paper, we conduct theoretical and experimental analysis to explore the fundamental causes of performance degradation in deep GCNs: over-smoothing and gradient vanishing have a mutually reinforcing effect that causes the performance to deteriorate more quickly in deep GCNs. On the other hand, existing anti-over-smoothing methods all perform full convolutions up to the model depth. They could not well resist the exponential convergence of over-smoothing due to model depth increasing. In this work, we propose a simple yet effective plug-and-play module,SkipNode, to overcome the performance degradation of deep GCNs. It samples graph nodes in each convolutional layer to skip the convolution operation. In this way, both over-smoothing and gradient vanishing can be effectively suppressed since (1) not all nodes'features propagate through full layers and, (2) the gradient can be directly passed back through “skipped” nodes. We provide both theoretical analysis and empirical evaluation to demonstrate the efficacy ofSkipNodeand its superiority over SOTA baselines.
Weigang Lu 0001, Yibing Zhan, Binbin Lin 0001, Ziyu Guan, Liu Liu 0014, Baosheng Yu, Wei Zhao 0019, Yaming Yang 0002, Dacheng Tao
IEEE Trans. Knowl. Data Eng.6
2024 Reciprocal Teacher-Student Learning via Forward and Feedback Knowledge Distillation
abstract
Knowledge distillation (KD) is a prevalent model compression technique in deep learning, aiming to leverage knowledge from a large teacher model to enhance the training of a smaller student model. It has found success in deploying compact deep models in intelligent applications like intelligent transportation, smart health, and distributed intelligence. Current knowledge distillation methods primarily fall into two categories: offline and online knowledge distillation. Offline methods involve a one-way distillation process, transferring unvaried knowledge from teacher to student, while online methods enable the simultaneous training of multiple peer students. However, existing knowledge distillation methods often face challenges where the student may not fully comprehend the teacher's knowledge due to model capacity gaps, and there might be knowledge incongruence among outputs of multiple students without teacher guidance. To address these issues, we propose a novel reciprocal teacher-student learning inspired by human teaching and examining through forward and feedback knowledge distillation (FFKD). Forward knowledge distillation operates offline, while feedback knowledge distillation follows an online scheme. The rationale is that feedback knowledge distillation enables the pre-trained teacher model to receive feedback from students, allowing the teacher to refine its teaching strategies accordingly. To achieve this, we introduce a new weighting constraint to gauge the extent of students' understanding of the teacher's knowledge, which is then utilized to enhance teaching strategies. Experimental results on five visual recognition datasets demonstrate that the proposed FFKD outperforms current state-of-the-art knowledge distillation methods.
Jianping Gou, Baosheng Yu, Jinhua Liu 0001, Lan Du 0002, Shaohua Wan 0001, Zhang Yi 0001
IEEE Trans. Multim.3
2024 Hierarchical Locality-Aware Deep Dictionary Learning for Classification
abstract
Deep dictionary learning (DDL) shows good performance in visual classification tasks. However, almost all existing DDL methods ignore the locality relationships between the input data representations and the learned dictionary atoms, and learn sub-optimal representations in the feature coding stage, which are less conducive to classification. To this end, we propose a hierarchical locality-aware deep dictionary learning (HILADLE) framework for classification, which can learn locality-constrained dictionaries at different abstract levels through hierarchical dictionary learning. The locality constraints play an important role in learning informative dictionary atoms while preserving the data structure in the original input feature space. Moreover, instead of using an identity activation function like existing DDL methods, we further boost the generalization performance of our HILADLE method with a ReLU activation function to deal with the overfitting issue caused by over-parameterization, inspired by its effectiveness in deep neural networks. Finally, the concatenation of all feature representations learned at different layers is used as input to the final classifier. We demonstrate, through an extensive set of experiments on several benchmark face recognition, image classification, and age estimation datasets, that our method is able to surpass several dictionary learning, deep dictionary learning and deep learning methods.
Jianping Gou, Xin He 0034, Lan Du 0002, Baosheng Yu, Zhang Yi 0001
IEEE Trans. Multim.4
2024 Bounding Box Vectorization for Oriented Object Detection With Tanimoto Coefficient Regression
abstract
Current oriented object detection methods mainly utilize a vanilla coordinate-angle representation for bounding box regression, which usually suffers from inconsistency between the bounding box regression losses and prediction errors induced with respect to different rotation angles, aspect ratios, and scales. Therefore, although the existing oriented object detectors have achieved very good performances under coarse evaluation metrics such as AP50, their performance significantly degrades when using stricter evaluation metric such as AP75. To address the abovementioned issues, we propose a new regression method with bounding box vectorization that implicitly represents the shape and orientation of an object with a set of orthogonal vectors. By doing this, the proposed method delicately avoids the inconsistency issues encountered in oriented bounding box regression. During training, we introduce the Tanimoto coefficient to evaluate the similarity of the bounding box vector in a shape- and orientation-aware manner, and we refer to the proposed box-to-vector loss as the B2V loss. In addition to 2D object detection, the proposed method can be easily generalized to 3D scenarios involving orientation estimation, such as autonomous driving. We evaluate the proposed method through extensive experiments conducted on four popular oriented object detection datasets, including both 2D and 3D datasets, where the proposed method significantly outperforms the recently developed state-of-the-art methods when using a more accurate evaluation metric.
Linfei Wang, Yibing Zhan, Wei Liu 0005, Baosheng Yu, Dapeng Tao
IEEE Trans. Multim.4
2024 Collaborative Knowledge Distillation via Multiknowledge Transfer
abstract
Knowledge distillation (KD), as an efficient and effective model compression technique, has received considerable attention in deep learning. The key to its success is about transferring knowledge from a large teacher network to a small student network. However, most existing KD methods consider only one type of knowledge learned from either instance features or relations via a specific distillation strategy, failing to explore the idea of transferring different types of knowledge with different distillation strategies. Moreover, the widely used offline distillation also suffers from a limited learning capacity due to the fixed large-to-small teacher-student architecture. In this article, we devise a collaborative KD via multiknowledge transfer (CKD-MKT) that prompts both self-learning and collaborative learning in a unified framework. Specifically, CKD-MKT utilizes a multiple knowledge transfer framework that assembles self and online distillation strategies to effectively: 1) fuse different kinds of knowledge, which allows multiple students to learn knowledge from both individual instances and instance relations, and 2) guide each other by learning from themselves using collaborative and self-learning. Experiments and ablation studies on six image datasets demonstrate that the proposed CKD-MKT significantly outperforms recent state-of-the-art methods for KD.
Jianping Gou, Liyuan Sun 0005, Baosheng Yu, Lan Du 0002, Kotagiri Ramamohanarao, Dacheng Tao
IEEE Trans. Neural Networks Learn. Syst.3
2024 Hierarchical Multi-Attention Transfer for Knowledge Distillation
abstract
Knowledge distillation (KD) is a powerful and widely applicable technique for the compression of deep learning models. The main idea of knowledge distillation is to transfer knowledge from a large teacher model to a small student model, where the attention mechanism has been intensively explored in regard to its great flexibility for managing different teacher-student architectures. However, existing attention-based methods usually transfer similar attention knowledge from the intermediate layers of deep neural networks, leaving the hierarchical structure of deep representation learning poorly investigated for knowledge distillation. In this paper, we propose a hierarchical multi-attention transfer framework (HMAT) , where different types of attention are utilized to transfer the knowledge at different levels of deep representation learning for knowledge distillation. Specifically, position-based and channel-based attention knowledge characterize the knowledge from low-level and high-level feature representations, respectively, and activation-based attention knowledge characterize the knowledge from both mid-level and high-level feature representations. Extensive experiments on three popular visual recognition tasks, image classification, image retrieval, and object detection, demonstrate that the proposed hierarchical multi-attention transfer or HMAT significantly outperforms recent state-of-the-art KD methods.
Jianping Gou, Liyuan Sun 0005, Baosheng Yu, Shaohua Wan 0001, Dacheng Tao
ACM Trans. Multim. Comput. Commun. Appl.3
2024 Attentional Composition Networks for Long-Tailed Human Action Recognition
abstract
The problem of long-tailed visual recognition has been receiving increasing research attention. However, the long-tailed distribution problem remains underexplored for video-based visual recognition. To address this issue, in this article we propose a compositional learning based solution for video-based human action recognition. Our method, named Attentional Composition Networks (ACN), first learns verb-like and preposition-like components, then shuffles these components to generate samples for the tail classes in the feature space to augment the data for the tail classes. Specifically, during training, we represent each action video by a graph that captures the spatial-temporal relations (edges) among detected human/object instances (nodes). Then, ACN utilizes the position information to decompose each action into a set of verb and preposition representations using the edge features in the graph. After that, the verb and preposition features from different videos are combined via an attention structure to synthesize feature representations for tail classes. This way, we can enrich the data for the tail classes and consequently improve the action recognition for these classes. To evaluate the compositional human action recognition, we further contribute a new human action recognition dataset, namely NEU-Interaction (NEU-I). Experimental results on both Something-Something V2 and the proposed NEU-I demonstrate the effectiveness of the proposed method for long-tailed, few-shot, and zero-shot problems in human action recognition. Source code and the NEU-I dataset are available at https://github.com/YajieW99/ACN .
Haoran Wang 0001, Baosheng Yu, Yibing Zhan, Chunfeng Yuan, Wankou Yang
ACM Trans. Multim. Comput. Commun. Appl.3
2023 Learnable Skeleton-Aware 3D Point Cloud Sampling
abstract
Point cloud sampling is crucial for efficient large-scale point cloud analysis, where learning-to-sample methods have recently received increasing attention from the community for jointly training with downstream tasks. However, the abovementioned task-specific sampling methods usually fail to explore the geometries of objects in an explicit manner. In this paper, we introduce a new skeleton-aware learning-to-sample method by learning object skeletons as the prior knowledge to preserve the object geometry and topology information during sampling. Specifically, without labor-intensive annotations per object category, we first learn category-agnostic object skeletons via the medial axis transform definition in an unsupervised manner. With object skeleton, we then evaluate the histogram of the local feature size as the prior knowledge to formulate skeleton-aware sampling from a probabilistic perspective. Additionally, the proposed skeleton-aware sampling pipeline with the task network is thus end-to-end trainable by exploring the reparameterization trick. Extensive experiments on three popular downstream tasks, point cloud classification, retrieval, and reconstruction, demonstrate the effectiveness of the proposed method for efficient point cloud analysis.
Cheng Wen 0001, Baosheng Yu, Dacheng Tao
CVPR2
2023 Knowledge-Aware Federated Active Learning with Non-IID Data
abstract
Federated learning enables multiple decentralized clients to learn collaboratively without sharing local data. However, the expensive annotation cost on local clients remains an obstacle in utilizing local data. In this paper, we propose a federated active learning paradigm to efficiently learn a global model with a limited annotation budget while protecting data privacy in a decentralized learning manner. The main challenge faced by federated active learning is the mismatch between the active sampling goal of the global model on the server and that of the asynchronous local clients. This becomes even more significant when data is distributed non-IID across local clients. To address the aforementioned challenge, we propose Knowledge-Aware Federated Active Learning (KAFAL), which consists of Knowledge-Specialized Active Sampling (KSAS) and Knowledge-Compensatory Federated Update (KCFU). Specifically, KSAS is a novel active sampling method tailored for the federated active learning problem, aiming to deal with the mismatch challenge by sampling actively based on the discrepancies between local and global models. KSAS intensifies specialized knowledge in local clients, ensuring the sampled data is informative for both the local clients and the global model. Meanwhile, KCFU deals with the client heterogeneity caused by limited data and non-IID data distributions by compensating for each client’s ability in weak classes with the assistance of the global model. Extensive experiments and analyses are conducted to show the superiority of KAFAL over recent state-of-the-art active learning methods. Code is available at https://github.com/ycao5602/KAFAL.
Yu-Tong Cao, Ye Shi 0001, Baosheng Yu, Jingya Wang 0001, Dacheng Tao
ICCV3
2023 Domain-Specific Risk Minimization for Domain Generalization
abstract
Domain generalization (DG) approaches typically use the hypothesis learned on source domains for inference on the unseen target domain. However, such a hypothesis can be arbitrarily far from the optimal one for the target domain, induced by a gap termed ''adaptivity gap.'' Without exploiting the domain information from the unseen test samples, adaptivity gap estimation and minimization are intractable, which hinders us to robustify a model to any unknown distribution. In this paper, we first establish a generalization bound that explicitly considers the adaptivity gap. Our bound motivates two strategies to reduce the gap: the first one is ensembling multiple classifiers to enrich the hypothesis space, then we propose effective gap estimation methods for guiding the selection of a better hypothesis for the target. The other method is minimizing the gap directly by adapting model parameters using online target samples. We thus propose Domain-specific Risk Minimization (DRM). During training, DRM models the distributions of different source domains separately; for inference, DRM performs online model steering using the source hypothesis for each arriving target sample. Extensive experiments demonstrate the effectiveness of the proposed DRM for domain generalization. Code is available at: https://github.com/yfzhang114/AdaNPC.
Yifan Zhang 0004, Jindong Wang 0001, Jian Liang 0001, Zhang Zhang 0001, Baosheng Yu, Liang Wang 0001, Dacheng Tao, Xing Xie 0001
KDD5
2023 Multi-target Knowledge Distillation via Student Self-reflection
abstract
Abstract Knowledge distillation is a simple yet effective technique for deep model compression, which aims to transfer the knowledge learned by a large teacher model to a small student model. To mimic how the teacher teaches the student, existing knowledge distillation methods mainly adapt an unidirectional knowledge transfer, where the knowledge extracted from different intermedicate layers of the teacher model is used to guide the student model. However, it turns out that the students can learn more effectively through multi-stage learning with a self-reflection in the real-world education scenario, which is nevertheless ignored by current knowledge distillation methods. Inspired by this, we devise a new knowledge distillation framework entitled multi-target knowledge distillation via student self-reflection or MTKD-SSR, which can not only enhance the teacher’s ability in unfolding the knowledge to be distilled, but also improve the student’s capacity of digesting the knowledge. Specifically, the proposed framework consists of three target knowledge distillation mechanisms: a stage-wise channel distillation (SCD), a stage-wise response distillation (SRD), and a cross-stage review distillation (CRD), where SCD and SRD transfer feature-based knowledge (i.e., channel features) and response-based knowledge (i.e., logits) at different stages, respectively; and CRD encourages the student model to conduct self-reflective learning after each stage by a self-distillation of the response-based knowledge. Experimental results on five popular visual recognition datasets, CIFAR-100, Market-1501, CUB200-2011, ImageNet, and Pascal VOC, demonstrate that the proposed framework significantly outperforms recent state-of-the-art knowledge distillation methods.
Jianping Gou, Xiangshuo Xiong, Baosheng Yu, Lan Du 0002, Yibing Zhan, Dacheng Tao
Int. J. Comput. Vis.3
2023 On exploring node-feature and graph-structure diversities for node drop graph pooling
Chuang Liu 0008, Yibing Zhan, Baosheng Yu, Liu Liu 0014, Bo Du 0001, Wenbin Hu 0001, Tongliang Liu
Neural Networks3
2023 Multilevel Attention-Based Sample Correlations for Knowledge Distillation
abstract
Recently, model compression has been widely used for the deployment of cumbersome deep models on resource-limited edge devices in the performance-demanding industrial Internet of Things (IoT) scenarios. As a simple yet effective model compression technique, knowledge distillation (KD) aims to transfer the knowledge (e.g., sample relationships as the relational knowledge) from a large teacher model to a small student model. However, existing relational KD methods usually build sample correlations directly from the feature maps at a certain middle layer in deep neural networks, which tends to overfit the feature maps of the teacher model and fails to address the most important sample regions. Inspired by this, we argue that the characteristics of important regions are of great importance, and thus, introduce attention maps to construct sample correlations for knowledge distillation. Specifically, with attention maps from multiple middle layers, attention-based sample correlations are newly built upon the most informative sample regions, and can be used as an effective and novel relational knowledge for knowledge distillation. We refer to the proposed method as multilevel attention-based sample correlations for knowledge distillation (or MASCKD). We perform extensive experiments on popular KD datasets for image classification, image retrieval, and person reidentification, where the experimental results demonstrate the effectiveness of the proposed method for relational KD.
Jianping Gou, Liyuan Sun 0005, Baosheng Yu, Shaohua Wan 0001, Weihua Ou, Zhang Yi 0001
IEEE Trans. Ind. Informatics3
2023 Intra- and Inter-Class Induced Discriminative Deep Dictionary Learning for Visual Recognition
abstract
Deep dictionary learning (DDL) aims to learn dictionaries at different levels and the deepest level representations. However, existing DDL algorithms impose a$l_{1}$-norm constraint on the deepest level representations, ignoring the constraints on different level representations. Meanwhile, they fail to discover effectively the essential discrimination information. Therefore, the obtained representations are less discriminative, which degrades model performance. To tackle those issues, we propose an intra- and inter-class induced discriminative deep dictionary learning (DDDL). Specifically, both intra-class compactness and inter-class separability of layer-wise data representations are newly devised as two discriminative constraints on deep dictionary learning. In a hierarchical structure, we obtain a more informative dictionary and the class-specific representations are thus more discriminative at each layer. Due to the$l_{2}$-norm intra- and inter-class constraints of layer-wise data representation, we devise a layer-wise optimization strategy to efficiently learn the closed-form solution of the deepest representation for classification. Comprehensive experiments and analyses on several visual recognition tasks show that our DDDL model surpasses recent shallow and deep representation learning approaches.
Jianping Gou, Xia Yuan, Baosheng Yu, Zhang Yi 0001
IEEE Trans. Multim.3
2023 Understanding How Pretraining Regularizes Deep Learning Algorithms
abstract
Deep learning algorithms have led to a series of breakthroughs in computer vision, acoustical signal processing, and others. However, they have only been popularized recently due to the groundbreaking techniques developed for training deep architectures. Understanding the training techniques is important if we want to further improve them. Through extensive experimentation, Erhan et al. (2010) empirically illustrated that unsupervised pretraining has an effect of regularization for deep learning algorithms. However, theoretical justifications for the observation remain elusive. In this article, we provide theoretical supports by analyzing how unsupervised pretraining regularizes deep learning algorithms. Specifically, we interpret deep learning algorithms as the traditional Tikhonov-regularized batch learning algorithms that simultaneously learn predictors in the input feature spaces and the parameters of the neural networks to produce the Tikhonov matrices. We prove that unsupervised pretraining helps in learning meaningful Tikhonov matrices, which will make the deep learning algorithms uniformly stable and the learned predictor will generalize fast w.r.t. the sample size. Unsupervised pretraining, therefore, can be interpreted as to have the function of regularization.
Yu Yao 0005, Baosheng Yu, Chen Gong 0002, Tongliang Liu
IEEE Trans. Neural Networks Learn. Syst.2
2022 Resistance Training Using Prior Bias: Toward Unbiased Scene Graph Generation
abstract
Scene Graph Generation (SGG) aims to build a structured representation of a scene using objects and pairwise relationships, which benefits downstream tasks. However, current SGG methods usually suffer from sub-optimal scene graph generation because of the long-tailed distribution of training data. To address this problem, we propose Resistance Training using Prior Bias (RTPB) for the scene graph generation. Specifically, RTPB uses a distributed-based prior bias to improve models' detecting ability on less frequent relationships during training, thus improving the model generalizability on tail categories. In addition, to further explore the contextual information of objects and relationships, we design a contextual encoding backbone network, termed as Dual Transformer (DTrans). We perform extensive experiments on a very popular benchmark, VG150, to demonstrate the effectiveness of our method for the unbiased scene graph generation. In specific, our RTPB achieves an improvement of over 10% under the mean recall when applied to current SGG methods. Furthermore, DTrans with RTPB outperforms nearly all state-of-the-art methods with a large margin. Code is available at https://github.com/ChCh1999/RTPB
Yibing Zhan, Baosheng Yu, Liu Liu 0014, Yong Luo 0002, Bo Du 0001
AAAI3
2022 BatchFormer: Learning to Explore Sample Relationships for Robust Representation Learning
abstract
Despite the success of deep neural networks, there are still many challenges in deep representation learning due to the data scarcity issues such as data imbalance, unseen distribution, and domain shift. To address the above-mentioned issues, a variety of methods have been devised to explore the sample relationships in a vanilla way (i.e., from the perspectives of either the input or the loss function), failing to explore the internal structure of deep neural networks for learning with sample relationships. Inspired by this, we propose to enable deep neural networks themselves with the ability to learn the sample relationships from each mini-batch. Specifically, we introduce a batch transformer module or BatchFormer, which is then applied into the batch dimension of each mini-batch to implicitly explore sample relationships during training. By doing this, the proposed method enables the collaboration of different samples, e.g., the head-class samples can also contribute to the learning of the tail classes for long-tailed recognition. Furthermore, to mitigate the gap between training and testing, we share the classifier between with or without the BatchFormer during training, which can thus be removed during testing. We perform extensive experiments on over ten datasets and the proposed method achieves significant improvements on different data scarcity applications without any bells and whistles, including the tasks of long-tailed recognition, compositional zero-shot learning, domain generalization, and contrastive learning. Code is made publicly available at https://github.com/zhihou7/BatchFormer.
Zhi Hou, Baosheng Yu, Dacheng Tao
CVPR2
2022 Learning Affinity from Attention: End-to-End Weakly-Supervised Semantic Segmentation with Transformers
abstract
Weakly-supervised semantic segmentation (WSSS) with image-level labels is an important and challenging task. Due to the high training efficiency, end-to-end solutions for WSSS have received increasing attention from the community. However, current methods are mainly based on convolutional neural networks and fail to explore the global information properly, thus usually resulting in incomplete object regions. In this paper, to address the aforementioned problem, we introduce Transformers, which naturally integrate global information, to generate more integral initial pseudo labels for end-to-end WSSS. Motivated by the inherent consistency between the self-attention in Transformers and the semantic affinity, we propose an Affinity from Attention (AFA) module to learn semantic affinity from the multi-head self-attention (MHSA) in Transformers. The learned affinity is then leveraged to refine the initial pseudo labels for segmentation. In addition, to efficiently derive reliable affinity labels for supervising AFA and ensure the local consistency of pseudo labels, we devise a Pixel-Adaptive Refinement module that incorporates low-level image appearance information to refine the pseudo labels. We perform extensive experiments and our method achieves 66.0% and 38.9% mIoU on the PASCAL VOC 2012 and MS COCO 2014 datasets, respectively, significantly outperforming recent end-to-end methods and several multi-stage competitors. Code is available at https://github.com/rulixiang/afa.
Lixiang Ru, Yibing Zhan, Baosheng Yu, Bo Du 0001
CVPR3
2022 Contrastive Boundary Learning for Point Cloud Segmentation
abstract
Point cloud segmentation is fundamental in understanding 3D environments. However, current 3D point cloud segmentation methods usually perform poorly on scene boundaries, which degenerates the overall segmentation performance. In this paper, we focus on the segmentation of scene boundaries. Accordingly, we first explore metrics to evaluate the segmentation performance on scene boundaries. To address the unsatisfactory performance on boundaries, we then propose a novel contrastive boundary learning (CBL) framework for point cloud segmentation. Specifically, the proposed CBL enhances feature discrimination between points across boundaries by contrasting their representations with the assistance of scene contexts at multiple scales. By applying CBL on three different baseline methods, we experimentally show that CBL consistently improves different baselines and assists them to achieve compelling performance on boundaries, as well as the overall performance, e.g. in mIoU. The experimental results demonstrate the effectiveness of our method and the importance of boundaries for 3D point cloud segmentation. Code and model will be made publicly available at https://github.com/LiyaoTang/contrastBoundary.
Liyao Tang, Yibing Zhan, Zhe Chen 0013, Baosheng Yu, Dacheng Tao
CVPR4
2022 Discovering Human-Object Interaction Concepts via Self-Compositional Learning
Zhi Hou, Baosheng Yu, Dacheng Tao
ECCV (27)2
2022 MeshMAE: Masked Autoencoders for 3D Mesh Data Analysis
Yaqian Liang, Shanshan Zhao 0001, Baosheng Yu, Jing Zhang 0037, Fazhi He
ECCV (3)3
2022 Improving Fine-Grained Visual Recognition in Low Data Regimes via Self-boosting Attention Mechanism
Yangyang Shu, Baosheng Yu, Lingqiao Liu
ECCV (25)2
2022 Deep Dictionary Learning with an Intra-Class Constraint
abstract
In recent years, deep dictionary learning (DDL)has attracted a great amount of attention due to its effectiveness for represen-tation learning and visual recognition. However, most existing methods focus on unsupervised deep dictionary learning, failing to further explore the category information. To make full use of the category information of different samples, we pro-pose a novel deep dictionary learning model with an intra-class constraint (DDLIC) for visual classification. Specif-ically, we design the intra-class compactness constraint on the intermediate representation at different levels to encour-age the intra-class representations to be closer to each other, and eventually the learned representation becomes more dis-criminative. Unlike the traditional DDL methods, during the classification stage, our DDLIC performs a layer-wise greedy optimization in a similar way to the training stage. Experi-mental results on four image datasets show that our method is superior to the state-of-the-art methods.
Xia Yuan, Jianping Gou, Baosheng Yu, Zhang Yi 0001
ICME3
2022 Knowledge Graph enhanced Multimodal Learning for Few-shot Visual Recognition
abstract
Few-shot learning (FSL) aims to learn a classifier for novel classes with only a few labeled samples per category available. The mainstream FSL approaches fall in the meta-learning paradigm, where a meta-learner is used to learn transferable knowledge and generalize to new tasks. However, these approaches usually only leverage information from a single modality (e.g., visual image) and fail to explore the information from other modalities (e.g., the knowledge graph). Since the labeled samples are scarce in FSL, increasing the information for each example is a possible solution to improve the performance. This motivates us to develop a new meta-learning framework for few-shot visual recognition termed Knowledge Graph enhanced FSL (KGFSL), which combines the information from multiple modalities: 1) the visual information in images and 2) the rich semantics and structural information in a knowledge graph (KG). Specifically, KGFSL exploits the word embedding of the category and its relationship to other categories to improve the visual-based models. A graph convolutional network (GCN) is first introduced to learn the semantic embeddings for each node (a visual category) in KG. The visual and semantic embeddings are then aligned and combined for final prediction. Finally, the whole framework is trained in an end-to-end manner. We conduct extensive experiments on two widely-used FSL benchmarks: miniImageNet and tieredImageNet. Experimental results demonstrate the effectiveness of the multimodal information for few-shot learning, and our proposed method can significantly outperform the state-of-the-art approaches.
Mengya Han, Yibing Zhan, Baosheng Yu, Yong Luo 0002, Bo Du 0001, Dacheng Tao
MMSP3
2022 Dual-branch Density Ratio Estimation for Signed Network Embedding
abstract
Signed network embedding (SNE) has received considerable attention in recent years. A mainstream idea of SNE is to learn node representations by estimating the ratio of sampling densities. Though achieving promising performance, these methods based on density ratio estimation are limited to the issues of confusing sample, expected error, and fixed priori. To alleviate the above-mentioned issues, in this paper, we propose a novel dual-branch density ratio estimation (DDRE) architecture for SNE. Specifically, DDRE 1) consists of a dual-branch network, dealing with the confusing sample; 2) proposes the expected matrix factorization without sampling to avoid the expected error; and 3) devises an adaptive cross noise sampling to alleviate the fixed priori. We perform sign prediction and node classification experiments on four real-world and three artificial datasets, respectively. Extensive empirical results demonstrate that DDRE not only significantly outperforms the methods based on density ratio estimation but also achieves competitive performance compared with other types of methods such as graph likelihood, generative adversarial networks, and graph convolutional networks. Code is publicly available at https://github.com/WHU-SNA/DDRE.
Pinghua Xu, Yibing Zhan, Liu Liu 0014, Baosheng Yu, Bo Du 0001, Jia Wu 0001, Wenbin Hu 0001
WWW4
2022 Heatmap Regression via Randomized Rounding
abstract
Heatmap regression has become the mainstream methodology for deep learning-based semantic landmark localization, including in facial landmark localization and human pose estimation. Though heatmap regression is robust to large variations in pose, illumination, and occlusion in unconstrained settings, it usually suffers from a sub-pixel localization problem. Specifically, considering that the activation point indices in heatmaps are always integers, quantization error thus appears when using heatmaps as the representation of numerical coordinates. Previous methods to overcome the sub-pixel localization problem usually rely on high-resolution heatmaps. As a result, there is always a trade-off between achieving localization accuracy and computational cost, where the computational complexity of heatmap regression depends on the heatmap resolution in a quadratic manner. In this paper, we formally analyze the quantization error of vanilla heatmap regression and propose a simple yet effective quantization system to address the sub-pixel localization problem. The proposed quantization system induced by the randomized rounding operation 1) encodes the fractional part of numerical coordinates into the ground truth heatmap using a probabilistic approach during training; and 2) decodes the predicted numerical coordinates from a set of activation points during testing. We prove that the proposed quantization system for heatmap regression is unbiased and lossless. Experimental results on popular facial landmark localization datasets (WFLW, 300W, COFW, and AFLW) and human pose estimation datasets (MPII and COCO) demonstrate the effectiveness of the proposed method for efficient and accurate semantic landmark localization. Code is available at http://github.com/baoshengyu/H3R.
Baosheng Yu, Dacheng Tao
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Multi-Stream Interaction Networks for Human Action Recognition
abstract
Skeleton-based human action recognition has received extensive attention due to its efficiency and robustness to complex backgrounds. Though the human skeleton can accurately capture the dynamics of human poses, it fails to recognize human actions induced by the interaction between human and objects, making it is of great importance to further explore the interaction between the human and objects for human action recognition. In this paper, we devise the multi-stream interaction networks (MSIN), to simultaneously explore the dynamics of human skeleton, objects, and the interaction between human and objects. Specifically, apart from the traditional human skeleton stream, 1) the second stream explores the dynamics of object appearance from the objects surrounding the human body joints; and 2) the third stream captures the dynamics of object position in regard to the distance between the object and different human body joints. Experimental results on three popular skeleton-based human action recognition datasets, NTU RGB + D, NTU RGB + D 120, and SYSU, demonstrate the effectiveness of the proposed method, especially for recognizing the human actions with human-object interactions.
Haoran Wang 0001, Baosheng Yu, Dongyue Chen 0001
IEEE Trans. Circuits Syst. Video Technol.2
2021 Affordance Transfer Learning for Human-Object Interaction Detection
abstract
Reasoning the human-object interactions (HOI) is essential for deeper scene understanding, while object affordances (or functionalities) are of great importance for human to discover unseen HOIs with novel objects. Inspired by this, we introduce an affordance transfer learning approach to jointly detect HOIs with novel object and recognize affordances. Specifically, HOI representations can be decoupled into a combination of affordance and object representations, making it possible to compose novel interactions by combining affordance representations and novel object representations from additional images, i.e. transferring the affordance to novel objects. With the proposed affordance transfer learning, the model is also capable of inferring the affordances of novel objects from known affordance representations. The proposed method can thus be used to 1) improve the performance of HOI detection, especially for the HOIs with unseen objects; and 2) infer the affordances of novel objects. Experimental results on two datasets, HICO-DET and HOI-COCO (from V-COCO), demonstrate significant improvements over recent state-of-the-art methods for HOI detection and object affordance detection. Code is available at https://github.com/zhihou7/HOI-CL.
Zhi Hou, Baosheng Yu, Yu Qiao 0001, Xiaojiang Peng, Dacheng Tao
CVPR2
2021 Detecting Human-Object Interaction via Fabricated Compositional Learning
abstract
Human-Object Interaction (HOI) detection, inferring the relationships between human and objects from images/videos, is a fundamental task for high-level scene understanding. However, HOI detection usually suffers from the open long-tailed nature of interactions with objects, while human has extremely powerful compositional perception ability to cognize rare or unseen HOI samples. Inspired by this, we devise a novel HOI compositional learning framework, termed as Fabricated Compositional Learning (FCL), to address the problem of open long-tailed HOI detection. Specifically, we introduce an object fabricator to generate effective object representations, and then combine verbs and fabricated objects to compose new HOI samples. With the proposed object fabricator, we are able to generate large-scale HOI samples for rare and unseen categories to alleviate the open long-tailed issues in HOI detection. Extensive experiments on the most popular HOI detection dataset, HICO-DET, demonstrate the effectiveness of the proposed method for imbalanced HOI detection and significantly improve the state-of-the-art performance on rare and unseen HOI categories. Code is available at https://github.com/zhihou7/HOI-CL.
Zhi Hou, Baosheng Yu, Yu Qiao 0001, Xiaojiang Peng, Dacheng Tao
CVPR2
2021 Learning Progressive Point Embeddings for 3D Point Cloud Generation
abstract
Generative models for 3D point clouds are extremely important for scene/object reconstruction applications in autonomous driving and robotics. Despite recent success of deep learning-based representation learning, it remains a great challenge for deep neural networks to synthesize or reconstruct high-fidelity point clouds, because of the difficulties in 1) learning effective pointwise representations; and 2) generating realistic point clouds from complex distributions. In this paper, we devise a dual-generators framework for point cloud generation, which generalizes vanilla generative adversarial learning framework in a progressive manner. Specifically, the first generator aims to learn effective point embeddings in a breadth-first manner, while the second generator is used to refine the generated point cloud based on a depth-first point embedding to generate a robust and uniform point cloud. The proposed dual-generators framework thus is able to progressively learn effective point embeddings for accurate point cloud generation. Experimental results on a variety of object categories from the most popular point cloud generation dataset, ShapeNet, demonstrate the state-of-the-art performance of the proposed method for accurate point cloud generation.
Cheng Wen 0001, Baosheng Yu, Dacheng Tao
CVPR2
2021 Not All Operations Contribute Equally: Hierarchical Operation-adaptive Predictor for Neural Architecture Search
abstract
Graph-based predictors have recently shown promising results on neural architecture search (NAS). Despite their efficiency, current graph-based predictors treat all operations equally, resulting in biased topological knowledge of cell architectures. Intuitively, not all operations are equally significant during forwarding propagation when aggregating information from these operations to another operation. To address the above issue, we propose a Hierarchical Operation-adaptive Predictor (HOP) for NAS. HOP contains an operation-adaptive attention module (OAM) to capture the diverse knowledge between operations by learning the relative significance of operations in cell architectures during aggregation over iterations. In addition, a cell-hierarchical gated module (CGM) further refines and enriches the obtained topological knowledge of cell architectures, by integrating cell information from each iteration of OAM. The experimental results compared with state-of-the-art predictors demonstrate the capability of our proposed HOP.
Ziye Chen, Yibing Zhan, Baosheng Yu, Mingming Gong, Bo Du 0001
ICCV3
2021 SynFace: Face Recognition with Synthetic Data
abstract
With the recent success of deep neural networks, remarkable progress has been achieved on face recognition. However, collecting large-scale real-world training data for face recognition has turned out to be challenging, especially due to the label noise and privacy issues. Meanwhile, existing face recognition datasets are usually collected from web images, lacking detailed annotations on attributes (e.g., pose and expression), so the influences of different attributes on face recognition have been poorly investigated. In this paper, we address the above-mentioned issues in face recognition using synthetic face images, i.e., SynFace. Specifically, we first explore the performance gap between recent state-of-the-art face recognition models trained with synthetic and real face images. We then analyze the underlying causes behind the performance gap, e.g., the poor intraclass variations and the domain gap between synthetic and real face images. Inspired by this, we devise the SynFace with identity mixup (IM) and domain mixup (DM) to mitigate the above performance gap, demonstrating the great potentials of synthetic data for face recognition. Furthermore, with the controllable face synthesis model, we can easily manage different factors of synthetic face generation, including pose, expression, illumination, the number of identities, and samples per identity. Therefore, we also perform a systematically empirical analysis on synthetic face images to provide some insights on how to effectively utilize synthetic data for face recognition.
Haibo Qiu, Baosheng Yu, Dihong Gong, Zhifeng Li 0001, Wei Liu 0005, Dacheng Tao
ICCV2
2021 TGRNet: A Table Graph Reconstruction Network for Table Structure Recognition
abstract
A table arranging data in rows and columns is a very effective data structure, which has been widely used in business and scientific research. Considering large-scale tabular data in online and offline documents, automatic table recognition has attracted increasing attention from the document analysis community. Though human can easily understand the structure of tables, it remains a challenge for machines to understand that, especially due to a variety of different table layouts and styles. Existing methods usually model a table as either the markup sequence or the adjacency matrix between different table cells, failing to address the importance of the logical location of table cells, e.g., a cell is located in the first row and the second column of the table. In this paper, we reformulate the problem of table structure recognition as the table graph reconstruction, and propose an end-to-end trainable table graph reconstruction network (TGRNet) for table structure recognition. Specifically, the proposed method has two main branches, a cell detection branch and a cell logical location branch, to jointly predict the spatial location and the logical location of different cells. Experimental results on three popular table recognition datasets and a new dataset with table graph annotations (TableGraph-350K) demonstrate the effectiveness of the proposed TGRNet for table structure recognition. Code and annotations will be made publicly available at https://github.com/xuewenyuan/TGRNet.
Wenyuan Xue, Baosheng Yu, Wen Wang 0019, Dacheng Tao, Qingyong Li
ICCV2
2021 A Question Answering System for Unstructured Table Images
abstract
Question answering over tables is a very popular semantic parsing task in natural language processing (NLP). However, few existing methods focus on table images, even though there are usually large-scale unstructured tables in practice (e.g., table images). Table parsing from images is nontrivial since it is closely related to not only NLP but also computer vision (CV) to parse the tabular structure from an image. In this demo, we present a question answering system for unstructured table images. The proposed system mainly consists of 1) a table recognizer to recognize the tabular structure from an image and 2) a table parser to generate the answer to a natural language question over the table. In addition, to train the model, we further provide table images and structure annotations for two widely used semantic parsing datasets. Specifically, the test set is used for this demo, from where the users can either choose from default questions or enter a new custom question.
Wenyuan Xue, Wen Wang 0019, Qingyong Li, Baosheng Yu, Yibing Zhan, Dacheng Tao
ACM Multimedia5
2021 Contrastive Graph Poisson Networks: Semi-Supervised Learning with Extremely Limited Labels
abstract
Graph Neural Networks (GNNs) have achieved remarkable performance in the task of semi-supervised node classification. However, most existing GNN models require sufficient labeled data for effective network training. Their performance can be seriously degraded when labels are extremely limited. To address this issue, we propose a new framework termed Contrastive Graph Poisson Networks (CGPN) for node classification under extremely limited labeled data. Specifically, our CGPN derives from variational inference; integrates a newly designed Graph Poisson Network (GPN) to effectively propagate the limited labels to the entire graph and a normal GNN, such as Graph Attention Network, that flexibly guides the propagation of GPN; applies a contrastive objective to further exploit the supervision information from the learning process of GPN and GNN models. Essentially, our CGPN can enhance the learning performance of GNNs under extremely limited labels by contrastively propagating the limited labels to the entire graph. We conducted extensive experiments on different types of datasets to demonstrate the superiority of CGPN.
Sheng Wan, Yibing Zhan, Liu Liu 0014, Baosheng Yu, Shirui Pan, Chen Gong 0002
NeurIPS4
2021 Knowledge Distillation: A Survey
Jianping Gou, Baosheng Yu, Stephen J. Maybank, Dacheng Tao
Int. J. Comput. Vis.2
2021 Skeleton edge motion networks for human action recognition
Haoran Wang 0001, Baosheng Yu
Neurocomputing2
2020 Unsupervised Domain Adaptation on Reading Comprehension
abstract
Reading comprehension (RC) has been studied in a variety of datasets with the boosted performance brought by deep neural networks. However, the generalization capability of these models across different domains remains unclear. To alleviate the problem, we investigate unsupervised domain adaptation on RC, wherein a model is trained on the labeled source domain and to be applied to the target domain with only unlabeled samples. We first show that even with the powerful BERT contextual representation, a model can not generalize well from one domain to another. To solve this, we provide a novel conditional adversarial self-training method (CASe). Specifically, our approach leverages a BERT model fine-tuned on the source dataset along with the confidence filtering to generate reliable pseudo-labeled samples in the target domain for self-training. On the other hand, it further reduces domain distribution discrepancy through conditional adversarial learning across domains. Extensive experiments show our approach achieves comparable performance to supervised models on multiple large-scale benchmark datasets.
Yu Cao 0014, Baosheng Yu, Joey Tianyi Zhou
AAAI3
2020 Supreme: Fine-grained Radio Map Reconstruction via Spatial-Temporal Fusion Network
abstract
Radio map, serving as an efficient indicator of wireless environments, has been widely used in smart-city applications, including network monitoring/planning, anomaly signal detection, and indoor/outdoor localization. It is hard to maintain an update-to-date fine-grained radio map within a large area, since the radio map changes rapidly due to the internal and external factors. Previous studies usually relied on time-consuming site surveys at densely predefined reference points, leading to either coarse-grained or out-of-date radio maps. In this paper, we propose a fine-grained radio map reconstruction framework, called Supreme, based on crowd-sourced data in an image super-resolution manner. Specifically, Supreme explores spatial-temporal relationships in historical coarse-grained radio maps and builds a real-time fine-grained radio map using deep spatial-temporal reconstruction networks. Furthermore, a heterogeneous data fusion module is devised to make full use of external information. To evaluate the performance of Supreme, we conduct extensive experiments and ablation studies on a large-scale dataset with a total of six-month data collected from two university campuses. Besides, we investigate the transferability of Supreme in different locations and service networks, showing that the fine-tuned model can largely reduce the training time and achieve better performance. Experimental results demonstrate that our model outperforms state-of-the-art baselines and a case study on the localization is enhanced with marginal improvements on accuracy.
Kehan Li 0001, Jiming Chen 0001, Baosheng Yu, Zhangchong Shen, Chao Li 0062, Shibo He
IPSN3
2019 Deep Metric Learning With Tuplet Margin Loss
abstract
Deep metric learning, in which the loss function plays a key role, has proven to be extremely useful in visual recognition tasks. However, existing deep metric learning loss functions such as contrastive loss and triplet loss usually rely on delicately selected samples (pairs or triplets) for fast convergence. In this paper, we propose a new deep metric learning loss function, tuplet margin loss, using randomly selected samples from each mini-batch. Specifically, the proposed tuplet margin loss implicitly up-weights hard samples and down-weights easy samples, while a slack margin in angular space is introduced to mitigate the problem of overfitting on the hardest sample. Furthermore, we address the problem of intra-pair variation by disentangling class-specific information to improve the generalizability of tuplet margin loss. Experimental results on three widely used deep metric learning datasets, CARS196, CUB200-2011, and Stanford Online Products, demonstrate significant improvements over existing deep metric learning methods.
Baosheng Yu, Dacheng Tao
ICCV1
2019 Anchor Cascade for Efficient Face Detection
abstract
Face detection is essential to facial analysis tasks such as facial reenactment and face recognition. Both cascade face detectors and anchor-based face detectors have translated shining demos into practice and received intensive attention from the community. However, cascade face detectors often suffer from a low detection accuracy, while anchor-based face detectors rely heavily on very large neural networks pre-trained on large scale image classification datasets such as ImageNet [1], which is not computationally efficient for both training and deployment. In this paper, we devise an efficient anchor-based cascade framework called anchor cascade. To improve the detection accuracy by exploring contextual information, we further propose a context pyramid maxout mechanism for anchor cascade. As a result, anchor cascade can train very efficient face detection models with a high detection accuracy. Specifically, comparing with a popular CNN-based cascade face detector MTCNN [2], our anchor cascade face detector greatly improves the detection accuracy, e.g., from 0.9435 to 0.9704 at 1k false positives on FDDB, while it still runs in comparable speed. Experimental results on two widely used face detection benchmarks, FDDB and WIDER FACE, demonstrate the effectiveness of the proposed framework.
Baosheng Yu, Dacheng Tao
IEEE Trans. Image Process.1
2018 Correcting the Triplet Selection Bias for Triplet Loss
Baosheng Yu, Tongliang Liu, Mingming Gong, Changxing Ding, Dacheng Tao
ECCV (6)1
2018 Deep Multi-View Feature Learning for Person Re-Identification
abstract
Person re-identification aims to identify the same pedestrians across different camera views at different locations. This important yet difficult intelligent video analysis problem remains a vigorous area of research due to demands for performance improvements. Person re-identification involves two main steps: feature representation and metric learning. Handcrafted features, such as color and texture histograms, are frequently used for person re-identification, but most handcrafted features are limited by not being directly applicable to practical problems. Deep learning methods have obtained the state-of-the-art performance in a wide variety of applications, including image annotation, face recognition, and speech recognition. However, deep learning features are heavily dependent on large-scale labeling of samples. In this paper, by utilizing the Cross-view Quadratic Discriminant Analysis (XQDA) metric learning, we propose a novel scheme called deep multi-view feature learning (DMVFL), which exploits the collaboration between handcrafted and deep learning features in a simple but effective way. Furthermore, we prove that the XQDA is a robust algorithm. Extensive experiments on two challenging person re-identification data sets (VIPeR and GRID) demonstrate that DMVFL improves on current state-of-the-art methods.
Dapeng Tao, Yanan Guo 0003, Baosheng Yu, Jianxin Pang, Zhengtao Yu 0001
IEEE Trans. Circuits Syst. Video Technol.3
2016 Linear Submodular Bandits with a Knapsack Constraint
abstract
Linear submodular bandits has been proven to be effective in solving the diversification and feature-based exploration problems in retrieval systems. Concurrently, many web-based applications, such as news article recommendation and online ad placement, can be modeled as budget-limited problems. However, the diversification problem under a budget constraint has not been considered. In this paper, we first introduce the budget constraint to linear submodular bandits as a new problem called the linear submodular bandits with a knapsack constraint. We then define an alpha-approximation unit-cost regret considering that submodular function maximization is NP-hard. To solve this problem, we propose two greedy algorithms based on a modified UCB rule. We then prove these two algorithms with different regret bounds and computational costs. We also conduct a number of experiments and the experimental results confirm our theoretical analyses.
Baosheng Yu, Dacheng Tao
AAAI1
2016 Submodular Asymmetric Feature Selection in Cascade Object Detection
abstract
A cascade classifier has turned out to be effective insliding-window based real-time object detection. In acascade classifier, node learning is the key process,which includes feature selection and classifier design. Previous algorithms fail to effectively tackle the asymmetry and intersection problems existing in cascade classification, thereby limiting the performance of object detection. In this paper, we improve current feature selection algorithm by addressing both asymmetry and intersection problems. We formulate asymmetric feature selection as a submodular function maximization problem. We then propose a new algorithm SAFS with formal performance guarantee to solve this problem.We use face detection as a case study and perform experiments on two real-world face detection datasets. The experimental results demonstrate that our algorithm SAFS outperforms the state-of-art feature selection algorithms in cascade object detection, such as FFS and LACBoost.
Baosheng Yu, Dacheng Tao, Jie Yin 0001
AAAI1
2016 Per-Round Knapsack-Constrained Linear Submodular Bandits
abstract
Linear submodular bandits has been proven to be effective in solving the diversification and feature-based exploration problem in information retrieval systems. Considering there is inevitably a budget constraint in many web-based applications, such as news article recommendations and online advertising, we study the problem of diversification under a budget constraint in a bandit setting. We first introduce a budget constraint to each exploration step of linear submodular bandits as a new problem, which we call per-round knapsack-constrained linear submodular bandits. We then define an [Formula: see text]-approximation unit-cost regret considering that the submodular function maximization is NP-hard. To solve this new problem, we propose two greedy algorithms based on a modified UCB rule. We prove these two algorithms with different regret bounds and computational complexities. Inspired by the lazy evaluation process in submodular function maximization, we also prove that a modified lazy evaluation process can be used to accelerate our algorithms without losing their theoretical guarantee. We conduct a number of experiments, and the experimental results confirm our theoretical analyses.
Baosheng Yu, Dacheng Tao
Neural Comput.1