EDBT 2026 Demo / reviewers in the wild / expert
Shuyang Zhao
dblp:62/11055
· DBLP profile ↗
9ranked-venue papers
6as first author
1since 2021 · last 2023
0000-0002-6151-6460ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Video understanding and tracking · 45% Efficient and distributed learning · 26% Segmentation and scene understanding · 22% | |
| Computer graphics and multimedia
2 papers |
Audio and music processing · 54% Image and video processing · 46% |
Topics — the 7 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
active learning |
0.4 | 1 | 2020 | Active Learning for Sound Event Detection · IEEE ACM Trans. Audio Speech Lang. Process. 2020 |
Audio and music processing
sound event detection |
0.4 | 1 | 2020 | Active Learning for Sound Event Detection · IEEE ACM Trans. Audio Speech Lang. Process. 2020 |
Computer vision › Segmentation and scene understanding › image segmentation › deep learning segmentation
attention-based segmentation |
0.4 | 1 | 2019 | Learning Unsupervised Video Object Segmentation Through Visual Attention · CVPR 2019 |
Computer vision › Video understanding and tracking › video object segmentation
unsupervised video object segmentation |
0.4 | 1 | 2019 | Learning Unsupervised Video Object Segmentation Through Visual Attention · CVPR 2019 |
Computer vision › Video understanding and tracking
video object segmentation |
0.4 | 1 | 2019 | Learning Unsupervised Video Object Segmentation Through Visual Attention · CVPR 2019 |
Image and video processing › saliency detection
salient object detection |
0.4 | 1 | 2019 | Salient Object Detection With Pyramid Attention and Salient Edges · CVPR 2019 |
Computer vision › Image recognition and object detection
saliency prediction |
0.1 | 1 | 2019 | Learning Unsupervised Video Object Segmentation Through Visual Attention · CVPR 2019 |
Methods — techniques the papers use, named apart from their topics
uncertainty sampling · 0.9mismatch-first farthest-traversal · 0.9change point detection · 0.9spatiotemporal attention · 0.4pyramid attention · 0.4eye-tracking · 0.4dynamic visual attention prediction · 0.4convolutional neural network · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Data augmentation for children ASR and child-adult speaker classification using voice conversion methods
Shuyang Zhao, Mittul Singh, Abraham Woubie, Reima Karhila |
INTERSPEECH | 1 |
| 2020 | Disentangled representation learning and residual GAN for age-invariant face verification
Shuyang Zhao, Jianwu Li |
Pattern Recognit. | 1 |
| 2020 | Active Learning for Sound Event DetectionabstractThis article proposes an active learning system for sound event detection (SED). It aims at maximizing the accuracy of a learned SED model with limited annotation effort. The proposed system analyzes an initially unlabeled audio dataset, from which it selects sound segments for manual annotation. The candidate segments are generated based on a proposed change point detection approach, and the selection is based on the principle of mismatch-first farthest-traversal. During the training of SED models, recordings are used as training inputs, preserving the long-term context for annotated segments. The proposed system clearly outperforms reference methods in the two datasets used for evaluation (TUT Rare Sound 2017 and TAU Spatial Sound 2019). Training with recordings as context outperforms training with only annotated segments. Mismatch-first farthest-traversal outperforms reference sample selection methods based on random sampling and uncertainty sampling. Remarkably, the required annotation effort can be greatly reduced on the dataset where target sound events are rare: by annotating only 2% of the training data, the achieved SED performance is similar to annotating all the training data. Shuyang Zhao, Toni Heittola, Tuomas Virtanen |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2019 | Learning Unsupervised Video Object Segmentation Through Visual AttentionabstractThis paper conducts a systematic study on the role of visual attention in Unsupervised Video Object Segmentation (UVOS) tasks. By elaborately annotating three popular video segmentation datasets (DAVIS, Youtube-Objects and SegTrack V2) with dynamic eye-tracking data in the UVOS setting, for the first time, we quantitatively verified the high consistency of visual attention behavior among human observers, and found strong correlation between human attention and explicit primary object judgements during dynamic, task-driven viewing. Such novel observations provide an in-depth insight into the underlying rationale behind UVOS. Inspired by these findings, we decouple UVOS into two sub-tasks: UVOS-driven Dynamic Visual Attention Prediction (DVAP) in spatiotemporal domain, and Attention-Guided Object Segmentation (AGOS) in spatial domain. Our UVOS solution enjoys three major merits: 1) modular training without using expensive video segmentation annotations, instead, using more affordable dynamic fixation data to train the initial video attention module and using existing fixation-segmentation paired static/image data to train the subsequent segmentation module; 2) comprehensive foreground understanding through multi-source learning; and 3) additional interpretability from the biologically-inspired and assessable attention. Experiments on popular benchmarks show that, even without using expensive video object mask annotations, our model achieves compelling performance in comparison with state-of-the-arts. Wenguan Wang, Hongmei Song, Shuyang Zhao, Jianbing Shen, Sanyuan Zhao, Steven C. H. Hoi, Haibin Ling |
CVPR | 3 |
| 2019 | Salient Object Detection With Pyramid Attention and Salient EdgesabstractThis paper presents a new method for detecting salient objects in images using convolutional neural networks (CNNs). The proposed network, named PAGE-Net, offers two key contributions. The first is the exploitation of an essential pyramid attention structure for salient object detection. This enables the network to concentrate more on salient regions while considering multi-scale saliency information. Such a stacked attention design provides a powerful tool to efficiently improve the representation ability of the corresponding network layer with an enlarged receptive field. The second contribution lies in the emphasis on the importance of salient edges. Salient edge information offers a strong cue to better segment salient objects and refine object boundaries. To this end, our model is equipped with a salient edge detection module, which is learned for precise salient boundary estimation. This encourages better edge-preserving salient object segmentation. Exhaustive experiments confirm that the proposed pyramid attention and salient edges are effective for salient object detection. We show that our deep saliency model outperforms state-of-the-art approaches for several benchmarks with a fast processing speed (25fps on one GPU). Wenguan Wang, Shuyang Zhao, Jianbing Shen, Steven C. H. Hoi, Ali Borji |
CVPR | 2 |
| 2018 | Removing Ring Artifacts in Cbct Images Via Generative Adversarial NetworkabstractCone-beam computed tomography (CBCT) images often have some ring artifacts because of the inconsistent response of detector pixels. Removing ring artifacts in CBCT images without impairing the image quality is critical for the application of CBCT. In this paper, we explore this issue as an “adversarial problem” and propose a novel method to eliminate ring artifacts from CBCT images by using an image-to-image network based on Generative Adversarial Network (GAN). Through combining the generative adversarial loss and the proposed smooth loss, both of the generator and the discriminator can be trained to remove ring artifacts in CBCT images by means of image-to-image. Experimental results demonstrate that the proposed method is more effective on both simulated data and real-world CBCT images, compared with other algorithms. Shuyang Zhao, Jianwu Li, Qirun Huo |
ICASSP | 1 |
| 2017 | Active learning for sound event classification by clustering unlabeled dataabstractThis paper proposes a novel active learning method to save annotation effort when preparing material to train sound event classifiers. K-medoids clustering is performed on unlabeled sound segments, and medoids of clusters are presented to annotators for labeling. The annotated label for a medoid is used to derive predicted labels for other cluster members. The obtained labels are used to build a classifier using supervised training. The accuracy of the resulted classifier is used to evaluate the performance of the proposed method. The evaluation made on a public environmental sound dataset shows that the proposed method outperforms reference methods (random sampling, certainty-based active learning and semi-supervised learning) with all simulated labeling budgets, the number of available labeling responses. Through all the experiments, the proposed method saves 50%-60% labeling budget to achieve the same accuracy, with respect to the best reference method. Shuyang Zhao, Toni Heittola, Tuomas Virtanen |
ICASSP | 1 |
| 2017 | Training Deep Autoencoder via VLC-Genetic Algorithm
Qazi Sami Ullah Khan, Jianwu Li, Shuyang Zhao |
ICONIP (2) | 3 |
| 2017 | Generating Low-Rank Textures via Generative Adversarial Network
Shuyang Zhao, Jianwu Li |
ICONIP (3) | 1 |