VLDB 2026 Research / reviewers in the wild / expert
Yonghao Dong
dblp:316/7073
· DBLP profile ↗
10ranked-venue papers
6as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EnViT: Enhancing the Performance of Early-Exit Vision Transformers via Exit-Aware Structured Dropout-Enabled Self-DistillationabstractVision Transformers (ViTs) have gained significant attention and widespread adoption due to their impressive performance in various computer vision tasks. However, in practice, their substantial computational overhead often leads to high inference latency and increased overheads when deployed on resource-constrained edge devices like smartphones, autonomous vehicles, and robots. To address these challenges, Early Exit (EE) has emerged as a promising approach for lightweight inference on edge devices. It accelerates inference and reduces computational overhead by adaptively producing predictions through early exits based on sample complexity. Existing EE methods typically suffer from substantial accuracy decreases in late exits while providing only marginal accuracy improvements to early exits. This paper presents EnViT, an exit-aware structured dropout-enabled self-distillation approach that enhances the performance of early exits without compromising late exits. EnViT leverages structured dropout to enable self-distillation, where the full model serves as the teacher and its own virtual sub-models generated by structured dropout as students. This mechanism effectively distills knowledge from the full model to early exits and avoids performance degradation in late exits by mitigating parameter conflicts across exits during training. Evaluation on five datasets shows that our EnViT achieves accuracy improvements ranging from 0.36% to 7.92% while maintaining competitive speed-up ratios of 1.72x to 2.23x. Yonghao Dong, Qiang He 0001, Penghong Rui, Zhenzhe Zheng 0001, Zhao Li 0007, Feifei Chen 0001, Hai Jin 0001, Yun Yang 0001 |
AAAI | 1 |
| 2026 | Pedestrian Trajectory Prediction via Hierarchical Dynamics DecompositionabstractPredicting human future trajectories is crucial for various intelligent systems and applications. Previous approaches typically adopt a direct prediction strategy, which decodes trajectory features directly into future coordinates. However, they overlook different hierarchical high-order velocities, which have stronger representational abilities in dynamics. In this paper, we introduce HDDNet, a novel trajectory prediction framework that follows dynamical principles and employs a hierarchical dynamics decomposition strategy. Specifically, HDDNet models future trajectories by progressively transferring trajectory coordinates into velocity, acceleration, and jerk, up to the highest-order velocity, which sequentially represent a broader receptive field and a more compact representation of motion dynamics. Furthermore, we design a hierarchical dynamics decomposition decoder with a corresponding dynamics loss, which predicts future trajectories by sequentially refining human motions from the highest-order velocity down to the final coordinates. Compared to the traditional direct prediction strategy, our approach makes better use of dynamic information at different levels. Extensive experiments and ablation studies on the ETH-UCY, SDD and GigaTraj datasets demonstrate that our method outperforms existing state-of-the-art approaches. Yonghao Dong, Le Wang 0003, Sanping Zhou, Gang Hua 0001, Changyin Sun 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Advancing Pre-Trained Teacher: Towards Robust Feature Discrepancy for Anomaly DetectionabstractWith the wide application of knowledge distillation between an ImageNet pre-trained teacher model and a learnable student model, unsupervised anomaly detection has witnessed a significant achievement in the past few years. The success of this framework mainly relies on how to keep the feature discrepancy between the teacher and student model, in which it has two underlying sub-assumptions: (1) The teacher model can represent two separable distributions for the normal and abnormal patterns, while (2) the student model can only reconstruct the normal distribution. However, it still remains a challenging issue to maintain these ideal assumptions in practice. In this paper, we propose a simple yet effective two-stage industrial anomaly detection framework, termed AAND, which sequentially performs Anomaly Amplification and Normality Distillation to enhance the two assumptions. In the first anomaly amplification stage, we propose a novel Residual Anomaly Amplification (RAA) module to advance the pre-trained teacher encoder with synthetic anomalies. It generates adaptive residuals to amplify anomalies while maintaining the feature integrity of pre-trained model. It mainly comprises a Matching-guided Residual Gate and an Attribute-scaling Residual Generator, which can determine the residuals' proportion and characteristic, respectively. In the second normality distillation stage, we further employ a reverse distillation paradigm to train a student decoder, in which a novel Hard Knowledge Distillation (HKD) loss is built to better facilitate the reconstruction of normal patterns. Comprehensive experiments on the MvTecAD, VisA, and MvTec3D-RGB datasets show that our method achieves state-of-the-art performance. Our code is available at https://github.com/Hui-design/AAND. Canhui Tang, Sanping Zhou, Yonghao Dong, Le Wang 0003 |
IEEE Trans. Image Process. | 4 |
| 2025 | A causality-inspired single-source domain generalized method for low-slow-small threat target recognition through holographic Doppler radar
Ligen Chen, Nannan Zhu, Yonghao Dong, Yue Zhang 0060, Nian Cai |
Expert Syst. Appl. | 4 |
| 2025 | AFC-RNN: Adaptive Forgetting-Controlled Recurrent Neural Network for Pedestrian Trajectory PredictionabstractPedestrian trajectory prediction plays a crucial and fundamental role in many computer vision tasks. Most existing works utilize recurrent neural networks to extract temporal features from trajectories because their recursive structure is inherently well-suited for time series data. However, previous methods overlook the forgetting characteristics of pedestrians when modeling historical trajectories, which may cause the model to focus on the wrong positions of historical information. In this paper, we propose a simple yet effective Adaptive Forgetting-Controlled Recurrent Neural Network (AFC-RNN) for pedestrian trajectory prediction. The core idea of AFC-RNN is a novel Adaptive Forgetting Controller (AFC), which controls the forgetting degree of the historical information at each time step explicitly and adaptively. Specifically, AFC first learns memory factors for each time step based on the temporal correlation of observed trajectories using the self-attention mechanism. Then, AFC-RNN applies these memory factors to regulate the forgetting degree of observed features at each time step from RNN. Extensive experiments and ablation studies on ETH, UCY, SDD, and NBA datasets demonstrate that our method outperforms existing state-of-the-art approaches. Additionally, we provide a mathematical analysis to demonstrate the superiority of our adaptive forgetting strategy in the AFC-RNN over traditional RNNs for trajectory forgetting modeling. Yonghao Dong, Le Wang 0003, Sanping Zhou, Wei Tang 0016, Gang Hua 0001, Changyin Sun 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Recurrent Aligned Network for Generalized Pedestrian Trajectory PredictionabstractPedestrian trajectory prediction is a crucial component in computer vision and robotics, but remains challenging due to the domain shift problem. Previous studies have tried to tackle this problem by leveraging a portion of trajectory data from the target domain to fine-tune the model. However, such domain adaptation methods are impractical in real-world scenarios, as it is infeasible to collect trajectory data from all potential target domains. In this paper, we study a new task named generalized pedestrian trajectory prediction, with the aim of generalizing the model to unseen domains without accessing their trajectories. To tackle this task, we further introduce a Recurrent Aligned Network (RAN) to minimize the domain gap through domain alignment. Specifically, we devise a recurrent alignment module to effectively align the trajectory feature spaces at both time-state and time-sequence levels by the recurrent alignment strategy. Furthermore, we introduce a pre-aligned representation module to combine social interactions with the recurrent alignment strategy, which aims to consider social interactions during the alignment process instead of just target trajectories. We extensively evaluate our method and compare it with state-of-the-art methods on three widely used benchmarks. The experimental results demonstrate the superior generalization capability of our method. Our work not only fills the gap in the generalization setting for practical pedestrian trajectory prediction, but also sets strong baselines in this field. Yonghao Dong, Le Wang 0003, Sanping Zhou, Gang Hua 0001, Changyin Sun 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | End-to-end pedestrian trajectory prediction via Efficient Multi-modal Predictors
Sanping Zhou, Le Wang 0003, Liushuai Shi, Yonghao Dong, Gang Hua 0001 |
Comput. Vis. Image Underst. | 5 |
| 2024 | Bidirectional feature learning network for RGB-D salient object detectionabstractRGB-D salient object detection aims to perform the pixel-wise localization of salient objects from both RGB and depth images, whose challenge mainly comes from how to learn complementary features from each modality. Existing works often use increasingly large models for performance enhancement, which need large memory and time consumption in practice. In this paper, we propose a simple yet effective B idirectional F eature L earning Net work (BFLNet) for RGB-D salient object detection under limited memory and time conditions. To achieve accurate performance with lightweight backbone networks , an effective B idirectional F eature F usion (BFF) module is designed to merge features from both RGB and depth streams, in which the cross-modal fusions and cross-scale fusions are jointly conducted to fuse the immediate features in multiple scales and multiple modals. What is more, a simple D ual C onsistency L oss (DCL) function is designed to prompt cross-modal fusion by keeping the consistency between cross-modal target predictions. Extensive experiments on four benchmark datasets demonstrate that our method has achieved the state-of-the-art performance with high efficiency in RGB-D salient object detection. Code will be available at https://github.com/nightsky-nostar/BFLNet . Ye Niu, Sanping Zhou, Yonghao Dong, Le Wang 0003, Jinjun Wang, Nanning Zheng 0001 |
Pattern Recognit. | 3 |
| 2024 | Sparse Pedestrian Character Learning for Trajectory PredictionabstractPedestrian trajectory prediction in a first-person view has recently attracted much attention due to its importance in autonomous driving. Recent work utilizes pedestrian character information, i.e., action and appearance, to improve the learned trajectory embedding and achieves state-of-the-art performance. However, it neglects the invalid and negative pedestrian character information, which is harmful to trajectory representation and thus leads to performance degradation. To address this issue, we present a two-stream sparse-character-based network (TSNet) for pedestrian trajectory prediction. Specifically, TSNet learns the negative-removed characters in the sparse character representation stream to improve the trajectory embedding obtained in the trajectory representation stream. Moreover, to model the negative-removed characters, we propose a novel sparse character graph, including the sparse category and sparse temporal character graphs, to learn the different effects of various characters in category and temporal dimensions, respectively. Extensive experiments on two first-person view datasets, PIE and JAAD, show that our method outperforms existing state-of-the-art methods. In addition, ablation studies demonstrate different effects of various characters and prove that TSNet outperforms approaches without eliminating negative characters. Yonghao Dong, Le Wang 0003, Sanping Zhou, Gang Hua 0001, Changyin Sun 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Sparse Instance Conditioned Multimodal Trajectory PredictionabstractPedestrian trajectory prediction is critical in many vision tasks but challenging due to the multimodality of the future trajectory. Most existing methods predict multi-modal trajectories conditioned by goals (future endpoints) or instances (all future points). However, goal-conditioned methods ignore the intermediate process and instance-conditioned methods ignore the stochasticity of pedestrian motions. In this paper, we propose a simple yet effective Sparse Instance Conditioned Network (SICNet), which gives a balanced solution between goal-conditioned and instance-conditioned methods. Specifically, SICNet learns comprehensive sparse instances, i.e., representative points of the future trajectory, through a mask generated by a long short-term memory encoder and uses the memory mechanism to store and retrieve such sparse instances. Hence SICNet can decode the observed trajectory into the future prediction conditioned on the stored sparse instance. Moreover, we design a memory refinement module that refines the retrieved sparse instances from the memory to reduce memory recall errors. Extensive experiments on ETH-UCY and SDD datasets show that our method outperforms existing state-of-the-art methods. In addition, ablation studies demonstrate the superiority of our method compared with goal-conditioned and instance-conditioned approaches. Yonghao Dong, Le Wang 0003, Sanping Zhou, Gang Hua 0001 |
ICCV | 1 |