VLDB 2026 Research / reviewers in the wild / expert
Yancheng Bai
dblp:53/10698
· DBLP profile ↗
21ranked-venue papers
7as first author
8since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 5 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Generative modeling · 51% Image recognition and object detection · 23% Face, body and person analysis · 11% | |
| Computer graphics and multimedia
1 paper |
Image and video processing · 100% |
Topics — the 16 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling › autoregressive model
autoregressive image generation |
1.0 | 1 | 2026 | SCALAR: Scale-wise Controllable Visual Autoregressive Learning · AAAI 2026 |
Machine learning › Generative modeling › diffusion model › controllable generation
controllable image generation |
1.0 | 1 | 2026 | SCALAR: Scale-wise Controllable Visual Autoregressive Learning · AAAI 2026 |
Computer vision › Image recognition and object detection › object detection
small object detection |
0.8 | 2 | 2020 | Multi-task Generative Adversarial Network for Detecting Small Objects in the Wild · Int. J. Comput. Vis. 2020 SOD-MTGAN: Small Object Detection via Multi-Task Generative Adversarial Network · ECCV (13) 2018 |
Machine learning › Generative modeling
generative adversarial network |
0.5 | 2 | 2020 | Finding Tiny Faces in the Wild With Generative Adversarial Network · CVPR 2018 Multi-task Generative Adversarial Network for Detecting Small Objects in the Wild · Int. J. Comput. Vis. 2020 |
Computer vision › Face, body and person analysis
face detection |
0.3 | 1 | 2018 | Finding Tiny Faces in the Wild With Generative Adversarial Network · CVPR 2018 |
Computer vision › Image recognition and object detection
object detection |
0.3 | 1 | 2018 | W2F: A Weakly-Supervised to Fully-Supervised Framework for Object Detection · CVPR 2018 |
Machine learning › Learning paradigms › semi-supervised learning
pseudo-labeling |
0.3 | 1 | 2018 | W2F: A Weakly-Supervised to Fully-Supervised Framework for Object Detection · CVPR 2018 |
Computer vision › Face, body and person analysis › face detection
small face detection |
0.3 | 1 | 2018 | Finding Tiny Faces in the Wild With Generative Adversarial Network · CVPR 2018 |
Machine learning › Generative modeling › image reconstruction
super-resolution |
0.3 | 1 | 2018 | Finding Tiny Faces in the Wild With Generative Adversarial Network · CVPR 2018 |
Computer vision › Image recognition and object detection › object detection
weakly supervised object detection |
0.3 | 1 | 2018 | W2F: A Weakly-Supervised to Fully-Supervised Framework for Object Detection · CVPR 2018 |
Image and video processing › super-resolution › image super-resolution
face super-resolution |
0.3 | 1 | 2018 | Finding Tiny Faces in the Wild With Generative Adversarial Network · CVPR 2018 |
Image and video processing
super-resolution |
0.3 | 1 | 2018 | Finding Tiny Faces in the Wild With Generative Adversarial Network · CVPR 2018 |
Machine learning › Generative modeling › conditional generative model
multimodal conditional generation |
0.3 | 1 | 2026 | SCALAR: Scale-wise Controllable Visual Autoregressive Learning · AAAI 2026 |
Robotics › Motion planning and robot control › hybrid systems
multi-modal control |
0.3 | 1 | 2026 | SCALAR: Scale-wise Controllable Visual Autoregressive Learning · AAAI 2026 |
Computer vision › Video understanding and tracking › object tracking
appearance modeling |
0.1 | 1 | 2012 | Robust tracking via weakly supervised ranking SVM · CVPR 2012 |
Computer vision › Video understanding and tracking
object tracking |
0.1 | 1 | 2012 | Robust tracking via weakly supervised ranking SVM · CVPR 2012 |
Methods — techniques the papers use, named apart from their topics
generative adversarial network · 1.4visual autoregressive model · 1.0semantic control encoding · 1.0scale-specific representation projection · 1.0multi-task learning · 0.8pseudo ground-truth excavation · 0.3multiple instance learning · 0.3weakly supervised learning · 0.1support vector machine · 0.1laplacian ranking · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SCALAR: Scale-wise Controllable Visual Autoregressive LearningabstractControllable image synthesis, which enables fine-grained control over generated outputs, has emerged as a key focus in visual generative modeling. However, controllable generation remains challenging for Visual Autoregressive (VAR) models due to their hierarchical, next-scale prediction style. Existing VAR-based methods often suffer from inefficient control encoding and disruptive injection mechanisms that compromise both fidelity and efficiency. In this work, we present SCALAR, a controllable generation method based on VAR, incorporating a Scale-wise Conditional Decoding mechanism. SCALAR leverages a pretrained image encoder to extract semantic control signal encodings, which are projected into scale-specific representations and injected into the corresponding layers of the VAR backbone. This design provides persistent and structurally aligned guidance throughout the generation process. Building on SCALAR, we develop SCALAR-Uni, a unified extension that aligns multiple control modalities into a shared latent space, supporting flexible multi-conditional guidance in a single model. Extensive experiments show that SCALAR achieves superior generation quality and control precision across various tasks. Ryan Xu, Dongyang Jin, Yancheng Bai, Rui Lan, Xu Duan, Xiangxiang Chu |
AAAI | 3 |
| 2025 | Revising Representation and Target Deviations for Accurate Human Pose EstimationabstractOwing to the normalized instance scales and robust supervision, heatmap-based human pose estimation (HPE) methods with top-down paradigm have achieved a dominant performance. However, there are two inherent deviations in the basic framework, i.e., representation and target deviations, resulting in performance bottlenecks. The representation deviation is caused by transforming various scales of instances into a unified input size, which results in performance degradation because data with different scale-related characteristics can hardly be handled via unified parameters. The target deviation is caused by exploiting a prior distribution (e.g., Gauss) to model the prediction error, which hinders sufficient network training. In this article, we propose a novel framework called DRPose to revise the abovementioned deviations. Specifically, to address the representation deviation, a scale-aware domain bridging (SDB) block is proposed to transfer feature maps from multiple scale-dependent domains into a unified intermediate domain with dynamic parameters. To address the target deviation, a differentiable coordinate decoder (DCD) is presented to adaptively adjust target distribution of heatmaps in an end-to-end manner. Extensive experiments show that the proposed method significantly improves the performance of most existing models with negligible additional cost. Beyond this, our method achieves 77.1% AP on the COCO test-dev set, outperforming prior works with similar model complexity. Zian Zhang, Yongqiang Zhang 0007, Yancheng Bai, Yin Zhang 0015, Mingli Ding, Wangmeng Zuo |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | A Study on the Adverse Impact of Synthetic Speech on Speech RecognitionabstractThe high-quality synthetic speech by TTS has been widely used in the field of human-computer interaction, bringing users better experience. However, synthetic speech is prone to be mixed with real human speech as part of the noise and recorded by the microphone, which leads to performance decrease for speech recognition. To address this issue, we propose different methods to study the adverse impact of synthetic speech on speech recognition, thereby enhancing its robustness. On the one hand, we adopt the concept of fake audio detection and incorporate an additional module into speech recognition model to differentiate between real and synthetic speech. On the other hand, we propose various methods of incorporating prompt labels from a language semantics perspective to achieve differentiation. These prompt labels provide contextual cues that help speech recognition model to better understand the difference between the two types of speech. The experimental results demonstrate the acoustic modeling of ASR is capable of distinguishing between real and synthetic speech effectively. Putting the prompt labels at the beginning achieves the best performance in a clean synthetic data scenario, while emptying the transcripts of synthetic speech obtains the best performance in a noisy synthetic data scenario. Yancheng Bai |
ICASSP | 2 |
| 2024 | R-CCF: region-aware continual contrastive fusion for weakly supervised object detection
Yongqiang Zhang 0007, Yin Zhang 0015, Zian Zhang, Yancheng Bai, Mingli Ding, Wangmeng Zuo |
Appl. Intell. | 5 |
| 2023 | Class-incremental object detection
Yongqiang Zhang 0007, Mingli Ding, Yancheng Bai |
Pattern Recognit. | 4 |
| 2023 | ThumbDet: One thumbnail image is enough for object detection
Yongqiang Zhang 0007, Yin Zhang 0015, Zian Zhang, Yancheng Bai, Wangmeng Zuo, Mingli Ding |
Pattern Recognit. | 5 |
| 2022 | One-stage object detection knowledge distillation via adversarial learning
Yongqiang Zhang 0007, Mingli Ding, Shibiao Xu, Yancheng Bai |
Appl. Intell. | 5 |
| 2021 | KGSNet: Key-Point-Guided Super-Resolution Network for Pedestrian Detection in the WildabstractIn real-world scenarios (i.e., in the wild), pedestrians are often far from the camera (i.e., small scale), and they often gather together and occlude with each other (i.e., heavily occluded). However, detecting these small-scale and heavily occluded pedestrians remains a challenging problem for the existing pedestrian detection methods. We argue that these problems arise because of two factors: 1) insufficient resolution of feature maps for handling small-scale pedestrians and 2) lack of an effective strategy for extracting body part information that can directly deal with occlusion. To solve the above-mentioned problems, in this article, we propose a key-point-guided super-resolution network (coined KGSNet) for detecting these small-scale and heavily occluded pedestrians in the wild. Specifically, to address factor 1), a super-resolution network is first trained to generate a clear super-resolution pedestrian image from a small-scale one. In the super-resolution network, we exploit key points of the human body to guide the super-resolution network to recover fine details of the human body region for easier pedestrian detection. To address factor 2), a part estimation module is proposed to encode the semantic information of different human body parts where four semantic body parts (i.e., head and upper/middle/bottom body) are extracted based on the key points. Finally, based on the generated clear super-resolved pedestrian patches padded with the extracted semantic body part images at the image level, a classification network is trained to further distinguish pedestrians/backgrounds from the inputted proposal regions. Both proposed networks (i.e., super-resolution network and classification network) are optimized in an alternating manner and trained in an end-to-end fashion. Extensive experiments on the challenging CityPersons data set demonstrate the effectiveness of the proposed method, which achieves superior performance over previous state-of-the-art methods, especially for those small-scale and heavily occluded instances. Beyond this, we also achieve state-of-the-art performance (i.e., 3.89% MR-2on the reasonable subset) on the Caltech data set. Yongqiang Zhang 0007, Yancheng Bai, Mingli Ding, Shibiao Xu, Bernard Ghanem |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | Multi-task Generative Adversarial Network for Detecting Small Objects in the Wild
Yongqiang Zhang 0007, Yancheng Bai, Mingli Ding, Bernard Ghanem |
Int. J. Comput. Vis. | 2 |
| 2019 | Corrigendum to 'Weakly-supervised Object Detection via Mining Pseudo Ground Truth Bounding-boxes' [Pattern Recognition 84 (2018) 68-81]
Yongqiang Zhang 0007, Yancheng Bai, Mingli Ding, Bernard Ghanem |
Pattern Recognit. | 2 |
| 2019 | Detecting small faces in the wild based on generative adversarial network and contextual information
Yongqiang Zhang 0007, Mingli Ding, Yancheng Bai, Bernard Ghanem |
Pattern Recognit. | 3 |
| 2019 | Learning a strong detector for action localization in videos
Yongqiang Zhang 0007, Mingli Ding, Yancheng Bai, Bernard Ghanem |
Pattern Recognit. Lett. | 3 |
| 2018 | Finding Tiny Faces in the Wild With Generative Adversarial NetworkabstractFace detection techniques have been developed for decades, and one of remaining open challenges is detecting small faces in unconstrained conditions. The reason is that tiny faces are often lacking detailed information and blurring. In this paper, we proposed an algorithm to directly generate a clear high-resolution face from a blurry small one by adopting a generative adversarial network (GAN). Toward this end, the basic GAN formulation achieves it by super-resolving and refining sequentially (e.g. SR-GAN and cycle-GAN). However, we design a novel network to address the problem of super-resolving and refining jointly. We also introduce new training losses to guide the generator network to recover fine details and to promote the discriminator network to distinguish real vs. fake and face vs. non-face simultaneously. Extensive experiments on the challenging dataset WIDER FACE demonstrate the effectiveness of our proposed method in restoring a clear high-resolution face from a blurry small one, and show that the detection performance outperforms other state-of-the-art methods. Yancheng Bai, Yongqiang Zhang 0007, Mingli Ding, Bernard Ghanem |
CVPR | 1 |
| 2018 | W2F: A Weakly-Supervised to Fully-Supervised Framework for Object DetectionabstractWeakly-supervised object detection has attracted much attention lately, since it does not require bounding box annotations for training. Although significant progress has also been made, there is still a large gap in performance between weakly-supervised and fully-supervised object detection. Recently, some works use pseudo ground-truths which are generated by a weakly-supervised detector to train a supervised detector. Such approaches incline to find the most representative parts of objects, and only seek one ground-truth box per class even though many same-class instances exist. To overcome these issues, we propose a weakly-supervised to fully-supervised framework, where a weakly-supervised detector is implemented using multiple instance learning. Then, we propose a pseudo ground-truth excavation (PGE) algorithm to find the pseudo ground-truth of each instance in the image. Moreover, the pseudo ground-truth adaptation (PGA) algorithm is designed to further refine the pseudo ground-truths from PGE. Finally, we use these pseudo ground-truths to train a fully-supervised detector. Extensive experiments on the challenging PASCAL VOC 2007 and 2012 benchmarks strongly demonstrate the effectiveness of our framework. We obtain 52.4% and 47.8% mAP on VOC2007 and VOC2012 respectively, a significant improvement over previous state-of-the-art methods. Yongqiang Zhang 0007, Yancheng Bai, Mingli Ding, Bernard Ghanem |
CVPR | 2 |
| 2018 | SOD-MTGAN: Small Object Detection via Multi-Task Generative Adversarial Network
Yancheng Bai, Yongqiang Zhang 0007, Mingli Ding, Bernard Ghanem |
ECCV (13) | 1 |
| 2018 | Weakly-supervised object detection via mining pseudo ground truth bounding-boxes
Yongqiang Zhang 0007, Yancheng Bai, Mingli Ding, Bernard Ghanem |
Pattern Recognit. | 2 |
| 2016 | Multi-Scale Fully Convolutional Network for Fast Face Detection
Yancheng Bai, Wenjing Ma, Yucheng Li 0002, Liangliang Cao, Luwei Yang |
BMVC | 1 |
| 2014 | Robust visual tracking via augmented kernel SVM
Yancheng Bai, Ming Tang 0001 |
Image Vis. Comput. | 1 |
| 2014 | Object Tracking via Robust Multitask Sparse RepresentationabstractSparse representation has been applied to the object tracking problem. Mining the self-similarities between particles via multitask learning can improve tracking performance. However, some particles may be different from others when they are sampled from a large region. Imposing all particles share the same structure may degrade the results. To overcome this problem, we propose a tracking algorithm based on robust multitask sparse representation (RMTT) in this letter. When we learn the particle representations, we decompose the sparse coefficient matrix into two parts in our algorithm. Joint sparse regularization is imposed on one coefficient matrix while element-wise sparse regularization is imposed on another matrix. The former regularization exploits self-similarities of particles while the later one considers the differences between them. Experiments on the benchmark data show the superior performance over other state-of-art algorithms. Yancheng Bai, Ming Tang 0001 |
IEEE Signal Process. Lett. | 1 |
| 2012 | Robust tracking via weakly supervised ranking SVMabstractAppearance model is a key component of tracking algorithms. Most existing approaches utilize the object information contained in the current and previous frames to construct the object appearance model and locate the object with the model in frame t + 1. This method may work well if the object appearance just fluctuates in short time intervals. Nevertheless, suboptimal locations will be generated in frame t + 1 if the visual appearance changes substantially from the model. Then, continuous changes would accumulate errors and finally result in a tracking failure. To copy with this problem, in this paper we propose a novel algorithm - online Laplacian ranking support vector tracker (LRSVT) - to robustly locate the object. The LRSVT incorporates the labeled information of the object in the initial and the latest frames to resist the occlusion and adapt to the fluctuation of the visual appearance, and the weakly labeled information from frame t + 1 to adapt to substantial changes of the appearance. Extensive experiments on public benchmark sequences show the superior performance of LRSVT over some state-of-the-art tracking algorithms. Yancheng Bai, Ming Tang 0001 |
CVPR | 1 |
| 2011 | Robust visual tracking via ranking SVMabstractIn this paper, we tackle the tracking problem in a quite other viewpoint, ranking. First, the ranking SVM is employed to learn a ranking function. Then, the ranking function ranks every instance sampled from the next frame, and the instance with the most preferred ranking score is assumed to be the object. Experiments of extensively quantitative and qualitative comparisons on public videos show the superior performance of our tracker over several state-of-the-art tracking algorithms. Yancheng Bai, Ming Tang 0001 |
ICIP | 1 |