EDBT 2026 Demo / reviewers in the wild / expert
Shicai Yang
dblp:126/6822
· DBLP profile ↗
26ranked-venue papers
0as first author
23since 2021 · last 2025
0000-0002-9260-1334ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 17 since 2021Artificial intelligence and machine learning · 16 · 13 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | VQCounter: Designing Visual Prompt Queue for Accurate Open-World CountingabstractClass-agnostic counting enables enumerating arbitrary object classes beyond those seen during training. Recent studies attempted to exploit the potential of visual foundation models such as GroundingDINO. Despite the considerable progress, we observe certain shortcomings, including the limited diversity of visual prompts and suboptimal training regimen. To address these issues, we introduce VQCounter, which incorporates a visual prompt queue mechanism designed to enrich the diversity of visual prompts. A random modality switching strategy is proposed during training to strengthen both textual and visual modalities. Besides, in light of weak point supervision, a Voronoi diagram-based cost (VoronoiCost) is designed to improve Hungarian matching, leading to more stable and faster convergence. Building upon the Voronoi diagram, we also propose a novel set of more stringent evaluation metrics, which take point localization into account. Extensive experiments on the FSC-147 and CARPK datasets demonstrate that VQCounter achieves state-of-the-art performance in both zero-shot and few-shot settings, significantly outperforming existing methods across nearly all evaluations. Fanfan Ye, Yiqi Fan, Qiaoyong Zhong, Shicai Yang, Di Xie, Jie Song 0011, Mingli Song |
IJCAI | 4 |
| 2025 | Training-Free Test-Time Adaptation via Shape and Style Guidance for Vision-Language ModelsabstractTest-time adaptation with pre-trained vision-language models shows impressive zero-shot classification abilities, and training-free methods further improve the performance without any optimization burden. However, existing training-free test-time adaptation methods typically rely on entropy criteria to select the visual features and update the visual caches, while ignoring the generalizable factors, such as shape-sensitive and style-insensitive factors. In this paper, we propose a novel shape and style guidance method (SSG) for training-free test-time adaptation in vision-language models, aiming to highlight the shape-sensitive (SHS) and style-insensitive (STI) factors in addition to entropy criteria.
Specifically, SSG perturbs the raw test image with shape and style corruption operations, and measures the prediction difference between the raw and corrupted one as perturbed prediction difference (PPD). Based on the PPD measurement, SSG reweights the high-confidence visual features and corresponding predictions, aiming to highlight the effect of SHS and STI factors during the test-time procedure. Furthermore, SSG takes both PPD and entropy into consideration to update the visual cache, aiming to maintain the stored sample with high entropy and generalizable factors. Extensive experimental results on out-of-distribution and cross-domain benchmark datasets demonstrate that our proposed SSG consistently outperforms previous state-of-the-art methods while also exhibiting promising computational efficiency. Shenglong Zhou 0002, Manjiang Yin, Leiyu Sun, Shicai Yang, Di Xie |
NeurIPS | 4 |
| 2025 | Adapt Anything: Tailor Any Image Classifier Across Domains and Categories Using Text-to-Image Diffusion ModelsabstractWe study a novel problem in this paper, that is, if a modern text-to-image diffusion model can tailor any image classifier across domains and categories. Existing domain adaption works exploit both source and target data for domain alignment so as to transfer the knowledge from the labeled source data to the unlabeled target data. However, as the development of text-to-image diffusion models, we wonder if the high-fidelity synthetic data can serve as a surrogate of the source data in real world. In this way, we do not need to collect and annotate the source data for each image classification task in a one-for-one manner. Instead, we utilize only one off-the-shelf text-to-image model to synthesize images with labels derived from text prompts, and then leverage them as a bridge to dig out the knowledge from the task-agnostic text-to-image generator to the task-oriented image classifier via domain adaptation. Such a one-for-all adaptation paradigm allows us to adapt anything in the world using only one text-to-image generator as well as any unlabeled target data. Extensive experiments validate the feasibility of this idea, which even surprisingly surpasses the state-of-the-art domain adaptation works using the source data collected and annotated in real world. Weijie Chen 0006, Haoyu Wang 0016, Shicai Yang, Lei Zhang 0054, Wei Wei 0008, Yanning Zhang 0001, Luojun Lin, Di Xie, Yueting Zhuang |
IEEE Trans. Big Data | 3 |
| 2024 | Multivariate Fourier Distribution Perturbation: Domain Shifts with Uncertainty in Frequency DomainabstractDiversifying training data techniques have achieved tremendous success in Domain Generalization (DG) tasks. The key to diversifying domain data is by increasing the types of domain styles. After investigating this issue from the perspective of the Fourier transform, the domain cue is found to be implicitly encoded in the amplitude component of Fourier features, which is more indicative of domain-specific information than statistics (means and standard deviations). However, Fourier-based methods tend to augment amplitude components via linear interpolation between two samples, which limits the diversity. To break this limitation, we aim to augment novel amplitude components from a perturbation perspective, which is termed Multivariate Fourier Distribution Perturbation. Specially, we design channel-wise and pixel-wise random perturbations for in-sample and cross-sample distribution to expand the distribution scope of probabilistic feature amplitude components. Weijie Chen 0006, Shicai Yang, Yishuang Li, Wenhao Guan |
ICASSP | 3 |
| 2024 | Better Together: Data-Free Multi-Student Coevolved Distillation
Weijie Chen 0006, Yunyi Xuan, Shicai Yang, Di Xie, Luojun Lin, Yueting Zhuang |
Knowl. Based Syst. | 3 |
| 2023 | PRIME: 3D Human Pose and Body Shape Recovery with Perspective ProjectionabstractExisting monocular 3D human pose and body shape (HPS) estimation methods make the coplanar assumption and use weak perspective projection in order to simplify the problem setting for images in the wild. However, weak perspective projection inevitably introduce prediction biases. To address this issue, we propose a plug-and-play Perspective Residual Log-likehood on Monocular 3D HPS Estimation (PRIME) module to significantly improve the accuracy of monocular 3D HPS estimation with trivial sacrifice on running time. PRIME applies full perspective projection to construct 2D re-projection loss or extract mesh-alignment features. Specifically, PRIME estimates the distribution of 2D joints and scale to calculate the perspective translation with the focal length. Further, we introduce side view constrain (SVC) of 2D joints to reduce the ambiguity of 3D HPS recovery. Experimental results demonstrate the effectiveness of our method. Baobei Xu, Shukai Fang, Shicai Yang, Di Xie, Shiliang Pu |
ICASSP | 4 |
| 2023 | Unsupervised Prompt Tuning for Text-Driven Object DetectionabstractGrounded language-image pre-trained models have shown strong zero-shot generalization to various downstream object detection tasks. Despite their promising performance, the models rely heavily on the laborious prompt engineering. Existing works typically address this problem by tuning text prompts using downstream training data in a few-shot or fully supervised manner. However, a rarely studied problem is to optimize text prompts without using any annotations. In this paper, we delve into this problem and propose an Unsupervised Prompt Tuning framework for text-driven object detection, which is composed of two novel mean teaching mechanisms. In conventional mean teaching, the quality of pseudo boxes is expected to optimize better as the training goes on, but there is still a risk of overfitting noisy pseudo boxes. To mitigate this problem, 1) we propose Nested Mean Teaching, which adopts nested-annotation to supervise teacher-student mutual learning in a bi-level optimization manner; 2) we propose Dual Complementary Teaching, which employs an offline pre-trained teacher and an online mean teacher via data-augmentation-based complementary labeling so as to ensure learning without accumulating confirmation bias. By integrating these two mechanisms, the proposed unsupervised prompt tuning framework achieves significant performance improvement on extensive object detection datasets. Weizhen He, Weijie Chen 0006, Shicai Yang, Di Xie, Luojun Lin, Donglian Qi, Yueting Zhuang |
ICCV | 4 |
| 2023 | Adapt then Generalize: A Simple Two-Stage Framework for Semi-Supervised Domain GeneralizationabstractSemi-supervised domain generalization (SSDG) focuses on training a model on a single labeled source domain and several unlabeled source domains simultaneously for the purpose of generalizing to out-of-distribution domain. To prevent model overfitting on the unlabeled source domains and enhance model generalization, we propose an effective two-stage framework that disentangle the SSDG task into a burn-in stage and a mutual-training stage. In the burn-in stage, we train a domain adaptation model to reduce domain gap and produce high-accuracy pseudo labels for the unlabeled data. In the second stage, we devise two peer networks equipped with style randomization modules to take advantage of both the style and class information from the pseudo-labeled data. The peer networks are mutually teaching each other, which allows them to avoid overfitting to noisy data and ultimately improves generalization ability. Extensive experiments show that our method achieves state-of-the-art performance on various benchmarks, demonstrating the effectiveness of our method on SSDG. Zhifeng Shen, Shicai Yang, Weijie Chen 0006, Luojun Lin |
ICME | 3 |
| 2023 | Distilling Vision-Language Foundation Models: A Data-Free Approach via Prompt DiversificationabstractData-Free Knowledge Distillation (DFKD) has shown great potential in creating a compact student model while alleviating the dependency on real training data by synthesizing surrogate data. However, prior arts are seldom discussed under distribution shifts, which may be vulnerable in real-world applications. Recent Vision-Language Foundation Models, e.g., CLIP, have demonstrated remarkable performance in zero-shot out-of-distribution generalization, yet consuming heavy computation resources. In this paper, we discuss the extension of DFKD to Vision-Language Foundation Models without access to the billion-level image-text datasets. The objective is to customize a student model for distribution-agnostic downstream tasks with given category concepts, inheriting the out-of-distribution generalization capability from the pre-trained foundation models. In order to avoid generalization degradation, the primary challenge of this task lies in synthesizing diverse surrogate images driven by text prompts. Since not only category concepts but also style information are encoded in text prompts, we propose three novel Prompt Diversification methods to encourage image synthesis with diverse styles, namely Mix-Prompt, Random-Prompt, and Contrastive-Prompt. Experiments on out-of-distribution generalization datasets demonstrate the effectiveness of the proposed methods, with Contrastive-Prompt performing the best. Yunyi Xuan, Weijie Chen 0006, Shicai Yang, Di Xie, Luojun Lin, Yueting Zhuang |
ACM Multimedia | 3 |
| 2023 | Spatial-Temporal Exclusive Capsule Network for Open Set Action RecognitionabstractOpen set action recognition (OSAR) is a rising research domain that simultaneously identifies all videos from known classes and rejects videos from unknown classes. Existing methods rarely consider the open set data distribution and the spatial-temporal relations of video subsequence. Recently proposed Capsule Network (CapsNet) has shown robust performance in many fields, especially image recognition. However, the current CapsNet has not been directly applied to the OSAR task since it cannot explicitly consider the data distribution of known and unknown classes along with the spatial-temporal relations for videos. This paper proposes the Spatial-Temporal Exclusive Capsule Network (STE-CapsNet) to solve the problems in the OSAR task. The STE-CapsNet designs the temporal-spatial routing mechanism to jointly capture the spatial-temporal information of the videos. Furthermore, the exclusive capsules are learned with dot product routing mechanism to limit the data distribution of closed set and open set and reduce the open set risk for OSAR. Extensive experimental results demonstrate that our proposed approach performs favorably compared with state-of-the-art methods on three standard datasets, which verifies its effectiveness and generalization ability. Yangbo Feng, Junyu Gao 0002, Shicai Yang, Changsheng Xu |
IEEE Trans. Multim. | 3 |
| 2022 | Label Matching Semi-Supervised Object DetectionabstractSemi-supervised object detection has made significant progress with the development of mean teacher driven self-training. Despite the promising results, the label mismatch problem is not yet fully explored in the previous works, leading to severe confirmation bias during self-training. In this paper, we delve into this problem and propose a simple yet effective LabelMatch framework from two different yet complementary perspectives, i.e., distribution-level and instance-level. For the former one, it is reasonable to approximate the class distribution of the unlabeled data from that of the labeled data according to Monte Carlo Sampling. Guided by this weakly supervision cue, we introduce a re-distribution mean teacher, which leverages adaptive label-distribution-aware confidence thresholds to generate unbiased pseudo labels to drive student learning. For the latter one, there exists an overlooked label assignment ambiguity problem across teacher-student models. To remedy this issue, we present a novel label assignment mechanism for self-training framework, namely proposal self-assignment, which injects the proposals from student into teacher and generates accurate pseudo labels to match each proposal in the student model accordingly. Experiments on both MS-COCO and PASCAL-VOC datasets demonstrate the considerable superiority of our proposed framework to other state-of-the-arts. Code will be available at https://github.com/HIK-LAB/SSOD. Weijie Chen 0006, Shicai Yang, Yunyi Xuan, Jie Song 0011, Di Xie, Shiliang Pu, Mingli Song, Yueting Zhuang |
CVPR | 3 |
| 2022 | Slimmable Domain AdaptationabstractVanilla unsupervised domain adaptation methods tend to optimize the model with fixed neural architecture, which is not very practical in real-world scenarios since the target data is usually processed by different resource-limited devices. It is therefore of great necessity to facilitate architecture adaptation across various devices. In this paper, we introduce a simple framework, Slimmable Domain Adaptation, to improve cross-domain generalization with a weight-sharing model bank, from which models of different capacities can be sampled to accommodate different accuracy-efficiency trade-offs. The main challenge in this frame-work lies in simultaneously boosting the adaptation performance of numerous models in the model bank. To tackle this problem, we develop a Stochastic EnsEmble Distillation method to fully exploit the complementary knowledge in the model bank for inter-model interaction. Nevertheless, considering the optimization conflict between inter-model interaction and intra-model adaptation, we augment the existing bi-classifier domain confusion architecture into an Optimization-Separated Tri-Classifier counterpart. After optimizing the model bank, architecture adaptation is leveraged via our proposed Unsupervised Performance Evaluation Metric. Under various resource constraints, our framework surpasses other competing approaches by a very large margin on multiple benchmarks. It is also worth emphasizing that our framework can preserve the performance improvement against the source-only model even when the computing complexity is reduced to 1/64. Code will be available at https://github.com/HIK-LAB/SlimDA. Rang Meng, Weijie Chen 0006, Shicai Yang, Jie Song 0011, Luojun Lin, Di Xie, Shiliang Pu, Xinchao Wang, Mingli Song, Yueting Zhuang |
CVPR | 3 |
| 2022 | Dual-Evidential Learning for Weakly-supervised Temporal Action Localization
Junyu Gao 0002, Shicai Yang, Changsheng Xu |
ECCV (4) | 3 |
| 2022 | Attention Diversification for Domain Generalization
Rang Meng, Weijie Chen 0006, Shicai Yang, Jie Song 0011, Xinchao Wang, Lei Zhang 0038, Mingli Song, Di Xie, Shiliang Pu |
ECCV (34) | 4 |
| 2022 | Transductive Clip with Class-Conditional Contrastive LearningabstractInspired by the remarkable zero-shot generalization capacity of vision-language pre-trained model, we seek to leverage the supervision from CLIP model to alleviate the burden of data labeling. However, such supervision inevitably contains the label noise, which significantly degrades the discriminative power of the classification model. In this work, we propose Transductive CLIP, a novel framework for learning a classification network with noisy labels from scratch. Firstly, a class-conditional contrastive learning mechanism is proposed to mitigate the reliance on pseudo labels and boost the tolerance to noisy labels. Secondly, ensemble labels is adopted as a pseudo label updating strategy to stabilize the training of deep neural networks with noisy labels. This framework can reduce the impact of noisy labels from CLIP model effectively by combining both techniques. Experiments on multiple benchmark datasets demonstrate the substantial improvements over other state-of-the-art methods. Junchu Huang, Weijie Chen 0006, Shicai Yang, Di Xie, Shiliang Pu, Yueting Zhuang |
ICASSP | 3 |
| 2022 | Target-Aware Auto-Augmentation for Unsupervised Domain Adaptive Object DetectionabstractRecent researches show that data auto-augmentation strategies can enhance the performance of object detection models. However, the existing works mainly focus on in-domain generalization. There is still a blank in out-of-domain generalization. In this paper, for the first time, we propose an auto-augmentation problem under unsupervised domain adaptation scenarios. To solve this problem, we propose a simple yet effective target-aware auto-augmentation technique to search for an optimal data augmentation strategy on labeled source data, so as to boost the detection ability on the given unlabeled target data. Our method can be easily plugged into the existing domain adaptation methods. Extensive experiments have been carried out to verify the effectiveness. Weijie Chen 0006, Shicai Yang, Di Xie, Shiliang Pu |
ICASSP | 4 |
| 2022 | Simulation-and-Mining: Towards Accurate Source-Free Unsupervised Domain Adaptive Object DetectionabstractVanilla unsupervised domain adaptive (UDA) object detection typically requires the labeled source data for joint-training with the unlabeled target data, which is usually unavailable in real-world scenarios due to data privacy, leading to source data-free UDA object detection. Herein, we first analyze the phenomenon of cross-domain detection degradation varying from easy to hard samples (e.g. the objects with different scales or occlusion degrees), termed as domain generalization differentiation. In detail, the ability to detect easy samples is well transferred while the one to detect hard samples is dramatically degraded. To this end, we then revisit the existing self-training method, which is of great challenge to deal with the abundant false negatives (hard samples). Assumed that true positives (easy samples) labeled by the source model can be exploited as supervision cues. UDA is finally modeled into an unsupervised false negatives mining problem. Thus, we propose a Simulation-and-Mining (S&M) framework, which simulates false negatives by augmenting true positives and mines back false negatives alternatively and iteratively. Experimental results show the effectiveness. Weijie Chen 0006, Shicai Yang, Yunyi Xuan, Di Xie, Yueting Zhuang, Shiliang Pu |
ICASSP | 3 |
| 2022 | Learning Domain Adaptive Object Detection with Probabilistic TeacherabstractSelf-training for unsupervised domain adaptive object detection is a challenging task, of which the performance depends heavily on the quality of pseudo boxes. Despite the promising results, prior works have largely overlooked the uncertainty of pseudo boxes during self-training. In this paper, we present a simple yet effective framework, termed as Probabilistic Teacher (PT), which aims to capture the uncertainty of unlabeled target data from a gradually evolving teacher and guides the learning of a student in a mutually beneficial manner. Specifically, we propose to leverage the uncertainty-guided consistency training to promote classification adaptation and localization adaptation, rather than filtering pseudo boxes via an elaborate confidence threshold. In addition, we conduct anchor adaptation in parallel with localization adaptation, since anchor can be regarded as a learnable parameter. Together with this framework, we also present a novel Entropy Focal Loss (EFL) to further facilitate the uncertainty-guided self-training. Equipped with EFL, PT outperforms all previous baselines by a large margin and achieve new state-of-the-arts. Meilin Chen, Weijie Chen 0006, Shicai Yang, Jie Song 0011, Xinchao Wang, Lei Zhang 0038, Yunfeng Yan, Donglian Qi, Yueting Zhuang, Di Xie, Shiliang Pu |
ICML | 3 |
| 2022 | Dynamic Domain GeneralizationabstractDomain generalization (DG) is a fundamental yet very challenging research topic in machine learning. The existing arts mainly focus on learning domain-invariant features with limited source domains in a static model. Unfortunately, there is a lack of training-free mechanism to adjust the model when generalized to the agnostic target domains. To tackle this problem, we develop a brand-new DG variant, namely Dynamic Domain Generalization (DDG), in which the model learns to twist the network parameters to adapt to the data from different domains. Specifically, we leverage a meta-adjuster to twist the network parameters based on the static model with respect to different data from different domains. In this way, the static model is optimized to learn domain-shared features, while the meta-adjuster is designed to learn domain-specific features. To enable this process, DomainMix is exploited to simulate data from diverse domains during teaching the meta-adjuster to adapt to the agnostic target domains. This learning mechanism urges the model to generalize to different agnostic target domains via adjusting the model without training. Extensive experiments demonstrate the effectiveness of our proposed method. Code is available: https://github.com/MetaVisionLab/DDG Zhishu Sun, Zhifeng Shen, Luojun Lin, Yuanlong Yu 0001, Zhifeng Yang, Shicai Yang, Weijie Chen 0006 |
IJCAI | 6 |
| 2022 | Self-Supervised Noisy Label Learning for Source-Free Unsupervised Domain AdaptationabstractDomain adaptation is an important property in robot vision, which enables the neural networks pre-trained on source domains to adapt target domains automatically without any annotation efforts. During this process, source data is not always accessible due to the constraints of expensive storage overhead and data privacy protection. Therefore, the source domain pre-trained model is expected to optimize with only unlabeled target data, termed as source-free unsupervised domain adaptation. In this paper, we view this problem as a special case of noisy label learning, since the given pre-trained model can generate noisy labels for unlabeled target data via network inference. The potential semantic cues for unsupervised domain adaptation exactly lie on these noisy labels. Inspired by this problem modeling, we propose a simple yet effective Self-Supervised Noisy Label Learning method, which injects self-supervised learning to impose the intrinsic data structure and facilitate label-denoising. Extensive experiments have been conducted on diverse benchmarks to validate the effectiveness. Our method achieves state-of-the-art performance. Weijie Chen 0006, Luojun Lin, Shicai Yang, Di Xie, Shiliang Pu, Yueting Zhuang |
IROS | 3 |
| 2021 | A Free Lunch for Unsupervised Domain Adaptive Object Detection without Source DataabstractUnsupervised domain adaptation (UDA) assumes that source and target domain data are freely available and usually trained together to reduce the domain gap. However, considering the data privacy and the inefficiency of data transmission, it is impractical in real scenarios. Hence, it draws our eyes to optimize the network in the target domain without accessing labeled source data. To explore this direction in object detection, for the first time, we propose a source data-free domain adaptive object detection (SFOD) framework via modeling it into a problem of learning with noisy labels. Generally, a straightforward method is to leverage the pre-trained network from the source domain to generate the pseudo labels for target domain optimization. However, it is difficult to evaluate the quality of pseudo labels since no labels are available in target domain. In this paper, self-entropy descent (SED) is a metric proposed to search an appropriate confidence threshold for reliable pseudo label generation without using any handcrafted labels. Nonetheless, completely clean labels are still unattainable. After a thorough experimental analysis, false negatives are found to dominate in the generated noisy labels. Undoubtedly, false negatives mining is helpful for performance improvement, and we ease it to false negatives simulation through data augmentation like Mosaic. Extensive experiments conducted in four representative adaptation tasks have demonstrated that the proposed framework can easily achieve state-of-the-art performance. From another view, it also reminds the UDA community that the labeled source data are not fully exploited in the existing methods. Weijie Chen 0006, Di Xie, Shicai Yang, Shiliang Pu, Yueting Zhuang |
AAAI | 4 |
| 2021 | TransForensics: Image Forgery Localization with Dense Self-AttentionabstractNowadays advanced image editing tools and technical skills produce tampered images more realistically, which can easily evade image forensic systems and make authenticity verification of images more difficult. To tackle this challenging problem, we introduce TransForensics, a novel image forgery localization method inspired by Transformers. The two major components in our framework are dense self-attention encoders and dense correction modules. The former is to model global context and all pairwise inter-actions between local patches at different scales, while the latter is used for improving the transparency of the hidden layers and correcting the outputs from different branches. Compared to previous traditional and deep learning methods, TransForensics not only can capture discriminative representations and obtain high-quality mask predictions but is also not limited by tampering types and patch sequence orders. By conducting experiments on main bench-marks, we show that TransForensics outperforms the state-of-the-art methods by a large margin. Shicai Yang, Di Xie, Shiliang Pu |
ICCV | 3 |
| 2021 | Unsupervised object detection with scene-adaptive concept learningabstractObject detection is one of the hottest research directions in computer vision, has already made impressive progress in academia, and has many valuable applications in the industry. However, the mainstream detection methods still have two shortcomings: (1) even a model that is well trained using large amounts of data still cannot generally be used across different kinds of scenes; (2) once a model is deployed, it cannot autonomously evolve along with the accumulated unlabeled scene data. To address these problems, and inspired by visual knowledge theory, we propose a novel scene-adaptive evolution unsupervised video object detection algorithm that can decrease the impact of scene changes through the concept of object groups. We first extract a large number of object proposals from unlabeled data through a pre-trained detection model. Second, we build the visual knowledge dictionary of object concepts by clustering the proposals, in which each cluster center represents an object prototype. Third, we look into the relations between different clusters and the object information of different groups, and propose a graph-based group information propagation strategy to determine the category of an object concept, which can effectively distinguish positive and negative proposals. With these pseudo labels, we can easily fine-tune the pre-trained model. The effectiveness of the proposed method is verified by performing different experiments, and the significant improvements are achieved. Shiliang Pu, Weijie Chen 0006, Shicai Yang, Di Xie, Yunhe Pan |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2020 | Cascade region proposal and global context for deep object detection
Qiaoyong Zhong, Chao Li 0064, Di Xie, Shicai Yang, Shiliang Pu |
Neurocomputing | 5 |
| 2019 | An End-to-End Audio Classification System Based on Raw Waveforms and Mix-Training StrategyabstractAudio classification can distinguish different kinds of sounds, which is helpful for intelligent applications in daily life.However, it remains a challenging task since the sound events in an audio clip is probably multiple, even overlapping.This paper introduces an end-to-end audio classification system based on raw waveforms and mix-training strategy.Compared to human-designed features which have been widely used in existing research, raw waveforms contain more complete information and are more appropriate for multi-label classification.Taking raw waveforms as input, our network consists of two variants of ResNet structure which can learn a discriminative representation.To explore the information in intermediate layers, a multi-level prediction with attention structure is applied in our model.Furthermore, we design a mix-training strategy to break the performance limitation caused by the amount of training data.Experiments show that the mean average precision of the proposed audio classification system on Audio Set dataset is 37.2%.Without using extra training data, our system exceeds the state-of-the-art multi-level attention model. Di Xie, Shicai Yang, Shiliang Pu |
INTERSPEECH | 5 |
| 2013 | Noncooperative bovine iris recognition via SIFT
Shengnan Sun, Shicai Yang, Lindu Zhao |
Neurocomputing | 2 |