VLDB 2026 Research / reviewers in the wild / expert
Yutong Zheng
dblp:164/0869
· DBLP profile ↗
16ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Open-World Assembly of Damaged Fragments: An Algorithm-Driven Framework for Dunhuang Manuscripts
Ming-Kun Chen, Xiaokang Zhao, Yanping Xiang, Jiaqi Dai, Zeli Tong, Langtai Cheng, Yutong Zheng |
ICIC (21) | 9 |
| 2025 | TCDNet: Texture and Color Dynamic Network for Image Harmonization
Shan Yue, Hai Huang 0001, Zhenqi Tang, Yutong Zheng |
CVM (2) | 4 |
| 2025 | DU-PMVS: Learned Patchmatch Multi-View Stereo Based on Deformable Feature Pyramid and Uncertainty Awareness ModelingabstractMulti-View Stereo is widely utilized for reconstructing the dense geometric structure of objects from multiple viewpoints. Recently, learning-based PatchMatch MVS methods have attracted significant attention due to their high efficiency and accuracy. However, existing methods neglect the constraints of convolutional kernel receptive fields and uniform depth sampling. In this paper, we propose a novel PatchMatch MVS method named DU-PMVS, which concurrently supports adaptive feature extraction and depth sampling. Specifically, we design a Deformable Feature Pyramid Extractor to capture multi-scale features for each pixel, thereby enhancing the representation of contextual information. Additionally, we propose Uncertainty Probability Distribution-guided Depth Modeling, which addresses the issue of cumulative errors in coarse-to-fine structure by exploring richer uncertainty distributions and generating more effective depth hypotheses. Experimental results on the DTU and Tanks & Temples datasets demonstrate that our DU-PMVS can reconstruct more completed point clouds with low memory, particularly in challenging regions with weak textures. Yutong Zheng, Shan Yue |
ICASSP | 1 |
| 2025 | Matching Ancient Dunhuang Manuscripts Based on Multi-dimensional Feature Fusion
Yanping Xiang, Jiaqi Dai, Ming-Kun Chen, Teer Song, Yutong Zheng |
KSEM (4) | 5 |
| 2025 | AMT-SNet: Adaptive Multi Scale Temporal-Spectral Network for Single-Channel EEG Sleep Stage Classification
Yutong Zheng, Ruiying Wang, Yuhui Du |
PRCV (18) | 1 |
| 2024 | A Reference-Based 3D Semantic-Aware Framework for Accurate Local Facial Attribute EditingabstractFacial attribute editing plays a crucial role in synthesizing realistic faces with specific characteristics while maintaining realistic appearances. Despite advancements, challenges persist in achieving precise, 3D-aware attribute modifications, which are crucial for consistent and accurate representations of faces from different angles. Current methods struggle with semantic entanglement and lack effective guidance for incorporating attributes while maintaining image integrity. To address these issues, we introduce a novel framework that merges the strengths of latent-based and reference-based editing methods. Our approach employs a 3D GAN inversion technique to embed attributes from the reference image into a tri-plane space, ensuring 3D consistency and realistic viewing from multiple perspectives. We utilize blending techniques and predicted semantic masks to locate precise edit regions, merging them with the contextual guidance from the reference image. A coarse-to-fine inpainting strategy is then applied to preserve the integrity of untargeted areas, significantly enhancing realism. Our evaluations demonstrate superior performance across diverse editing tasks, validating our framework’s effectiveness in realistic and applicable facial attribute editing. Yutong Zheng, Yen-Shuo Su, Anudeepsekhar Bolimera, Han Zhang 0048, Fangyi Chen, Marios Savvides |
IJCB | 2 |
| 2024 | Drive as Veteran: Fine-tuning of an Onboard Large Language Model for Highway Autonomous DrivingabstractDue to the limitations of network communication conditions for online calling GPT, the onboard deployment of Large Language Models for autonomous driving is in need. In this paper, we propose Drive as Veteran, a fine-tuned LLaMA-7B model with driving tasks. A training set consisting of instructions, scenario descriptions and human-annotated driving tasks is established. Through LoRA fine-tuning, the capability of generating correct driving tasks of our model is demonstrated through a numerical experiment and the comparison to GPT-3.5 is presented. We show that smaller-sized Large Language Models could be deployed onboard with fast generation speed and high accuracy, which could serve as a core component for decision-making in autonomous driving. Zhaoyan Huang, Quanfeng Liu, Yutong Zheng, Jinlong Hong, Bingzhao Gao, Hong Chen 0003 |
IV | 4 |
| 2022 | Powering Finetuning in Few-Shot Learning: Domain-Agnostic Bias Reduction with Selected SamplingabstractIn recent works, utilizing a deep network trained on meta-training set serves as a strong baseline in few-shot learning. In this paper, we move forward to refine novel-class features by finetuning a trained deep network. Finetuning is designed to focus on reducing biases in novel-class feature distributions, which we define as two aspects: class-agnostic and class-specific biases. Class-agnostic bias is defined as the distribution shifting introduced by domain difference, which we propose Distribution Calibration Module(DCM) to reduce. DCM owes good property of eliminating domain difference and fast feature adaptation during optimization. Class-specific bias is defined as the biased estimation using a few samples in novel classes, which we propose Selected Sampling(SS) to reduce. Without inferring the actual class distribution, SS is designed by running sampling using proposal distributions around support-set samples. By powering finetuning with DCM and SS, we achieve state-of-the-art results on Meta-Dataset with consistent performance boosts over ten datasets from different domains. We believe our simple yet effective method demonstrates its possibility to be applied on practical few-shot applications. Ran Tao 0013, Han Zhang 0048, Yutong Zheng, Marios Savvides |
AAAI | 3 |
| 2021 | Unsupervised Disentanglement of Linear-Encoded Facial SemanticsabstractWe propose a method to disentangle linear-encoded facial semantics from StyleGAN without external supervision. The method derives from linear regression and sparse representation learning concepts to make the disentangled latent representations easily interpreted as well. We start by coupling StyleGAN with a stabilized 3D deformable facial reconstruction method to decompose single-view GAN generations into multiple semantics. Latent representations are then extracted to capture interpretable facial semantics. In this work, we make it possible to get rid of labels for disentangling meaningful facial semantics. Also, we demonstrate that the guided extrapolation along the disentangled representations can help with data augmentation, which sheds light on handling unbalanced data. Finally, we provide an analysis of our learned localized facial representations and illustrate that the semantic information is encoded, which surprisingly complies with human intuition. The overall unsupervised design brings more flexibility to representation learning in the wild. Yutong Zheng, Ran Tao 0013, Marios Savvides |
CVPR | 1 |
| 2021 | Real-Time Semantic Segmentation of Aerial Videos Based on Bilateral Segmentation NetworkabstractIn recent years, deep learning algorithms have been widely used in semantic segmentation of aerial images. However, most of the current research in this field focus on images but not videos. In this paper, we address the problem of real-time aerial video semantic segmentation with BiSeNet[1]. Since BiSeNet is originally proposed for semantic segmentation of natural city scene images, we need a corresponding dataset to ensure the effect of transfer learning when applying it to aerial video segmentation. Therefore, we build a UAV streetscape sequence dataset (USSD) to fill the vacancy of dataset in this field and facilitate our research. Evaluation on USSD shows that BiSeNet outperforms other state-of-the-art methods. It achieves 79.26% mIoU and 93.37% OA with speed of 148.7 FPS on NVIDIA Tesla V100 for a 1920x1080 frame size input aerial video, which satisfies the demand of aerial video semantic segmentation with a competitive balance of accuracy and speed. The aerial video semantic segmentation results are provided at Our Repository. Yihao Zuo, Junli Yang, Yutong Zheng |
IGARSS | 6 |
| 2021 | CDTD: A Large-Scale Cross-Domain Benchmark for Instance-Level Image-to-Image Translation and Domain Adaptive Object Detection
Mingyang Huang, Jianping Shi, Zechun Liu, Harsh Maheshwari, Yutong Zheng, Xiangyang Xue 0001, Marios Savvides, Thomas S. Huang |
Int. J. Comput. Vis. | 6 |
| 2019 | Is Pose Really Solved? A Frontalization Study On Off-Angle Face MatchingabstractRecently, impressive results have been achieved on many large-scale face recognition benchmarks, such as IJB-A Janus and Janus CS3. These datasets were designed to test robustness to nuisance transformations simultaneously such as pose, illumination, expression etc. We present a study paper, where we find that despite this goal in evaluation, there exists a significant frontal bias in yaw pose in these datasets. Therefore, high-performance on these recent datasets is misleading and does not reflect robustness to extreme pose in yaw. Moreover, many real-world applications only allow for a single frontal enrollment in a gallery (law enforcement, immigration etc.). As we show in our study, face recognition in this highly constrained setting with extreme pose variation in the probe images remains a highly challenging problem. Traditional approaches, performing well on datasets such as IJB-A Janus, perform much worse on older but highly controlled datasets such as CMU MPIE. To aid our study, we present a simple and practical method to handle pose variation in face recognition pipelines designed to deal with extremely off-angle faces. Our approach is to ignore the half of the face with any self-occlusion. This method allows our models to be highly robust to pose, and helps us achieve state-of-the-art results on several protocols using the CMU MPIE dataset as well as very accurate results on the CFP dataset, outperforming recent efforts using the same training data. Dipan K. Pal, Chandrasekhar Bhagavatula, Yutong Zheng, Ran Tao 0013, Marios Savvides |
WACV | 3 |
| 2018 | Ring Loss: Convex Feature Normalization for Face RecognitionabstractWe motivate and present Ring loss, a simple and elegant feature normalization approach for deep networks designed to augment standard loss functions such as Softmax. We argue that deep feature normalization is an important aspect of supervised classification problems where we require the model to represent each class in a multi-class problem equally well. The direct approach to feature normalization through the hard normalization operation results in a non-convex formulation. Instead, Ring loss applies soft normalization, where it gradually learns to constrain the norm to the scaled unit circle while preserving convexity leading to more robust features. We apply Ring loss to large-scale face recognition problems and present results on LFW, the challenging protocols of IJB-A Janus, Janus CS3 (a superset of IJB-A Janus), Celebrity Frontal-Profile (CFP) and MegaFace with 1 million distractors. Ring loss outperforms strong baselines, matches state-of-the-art performance on IJB-A Janus and outperforms all other results on the challenging Janus CS3 thereby achieving state-of-the-art. We also outperform strong baselines in handling extremely low resolution face matching. Yutong Zheng, Dipan K. Pal, Marios Savvides |
CVPR | 1 |
| 2018 | Enhancing Interior and Exterior Deep Facial Features for Face Detection in the WildabstractAlthough face detection has been intensely studied for decades, it is still a challenging topic due to numerous conditions, e.g. heavy occlusions, low resolutions, extreme poses, non-face patterns that look like human faces, etc. This paper proposes a novel region-based ConvNet to address these issues. Our approach enhances the interior deep facial features and explicitly incorporates the exterior deep features. The enhanced interior features provide fine details for small faces. The exterior features capture the local information surrounding the face, supporting the detection under challenging conditions. Experiments show that our proposed components improve the baseline method significantly. Additionally, our approach consistently achieves competitive performance in four challenging databases, i.e. Wider Face, AFW, PASCAL Faces, and FDDB. We also introduce a new challenging non-face dataset 1 of 6,000 images to benchmark false positive rates for future research. Chenchen Zhu, Yutong Zheng, Khoa Luu, Marios Savvides |
FG | 2 |
| 2017 | DeepSafeDrive: A grammar-aware driver parsing approach to Driver Behavioral Situational Awareness (DB-SAW)
T. Hoang Ngan Le, Chenchen Zhu, Yutong Zheng, Khoa Luu, Marios Savvides |
Pattern Recognit. | 3 |
| 2016 | Robust hand detection in VehiclesabstractThe problems of hand detection have been widely addressed in many areas, e.g. human computer interaction environment, driver behaviors monitoring, etc. However, the detection accuracy in recent hand detection systems are still far away from the demands in practice due to a number of challenges, e.g. hand variations, highly occlusions, low-resolution and strong lighting conditions. This paper presents the Multiple Scale Faster Region-based Convolutional Neural Network (MS-FRCNN) to handle the problems of hand detection in given digital images collected under challenging conditions. Our proposed method introduces a multiple scale deep feature extraction approach in order to handle the challenging factors to provide a robust hand detection algorithm. The method is evaluated on the challenging hand database, i.e. the Vision for Intelligent Vehicles and Applications (VIVA) Challenge, and compared against various recent hand detection methods. Our proposed method achieves the state-of-the-art results with 20% of the detection accuracy higher than the second best one in the VIVA challenge. T. Hoang Ngan Le, Chenchen Zhu, Yutong Zheng, Khoa Luu, Marios Savvides |
ICPR | 3 |