Wentao Bao

dblp:147/1365 · also Wen-Tao Bao · DBLP profile ↗
← Back
25ranked-venue papers
11as first author
15since 2021 · last 2026
0000-0003-2571-3341ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 7 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 8 first-author · 12 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorComputer networks · 1
YearPublicationVenuePosition
2026 Open Set Face Forgery Detection via Dual-Level Evidence Collection
abstract
The surge in face forgeries has increasingly undermined confidence in the authenticity of online content. As generation algorithms rapidly evolve, new fake categories will constantly emerge, severely challenging existing face forgery detection methods. Although face forgery detection has recently improved, current techniques remain largely confined to binary Real-vs-Fake classification or the recognition of known fake categories. Moreover, they fail to identify the emergence of entirely new forgery methods. In this work, we study the Open Set Face Forgery Detection (OSFFD) problem, which requires the detection model to identify novel fake categories. To enhance its real-world applicability, we reformulate the OSFFD problem and address it through uncertainty estimation. Specifically, we propose the Dual-Level Evidential face forgery Detection (DLED) approach, which estimates prediction uncertainty by extracting and integrating category-specific evidence on the spatial and frequency levels. Comprehensive experiments across diverse settings demonstrate that our proposed DLED approach achieves state-of-the-art performance. Notably, it surpasses various existing baseline models by a $20\%$ margin on average when identifying forgeries from novel fake categories. Concurrently, our DLED method yields competitive performance on the standard binary Real-versus-Fake face forgery detection task.
Zhongyi Cai, Bryce Gernon, Wentao Bao, Matthew Wright 0001, Yu Kong 0001
FG3
2026 MADiff: Motion-Aware Mamba Diffusion Models for Hand Trajectory Prediction on Egocentric Videos
abstract
Understanding human intentions and actions through egocentric videos is important on the path to embodied artificial intelligence. As a branch of egocentric vision techniques, hand trajectory prediction plays a vital role in comprehending human motion patterns, benefiting downstream tasks in extended reality and robot manipulation. However, capturing high-level human intentions consistent with reasonable temporal causality is challenging when only egocentric videos are available. This difficulty is exacerbated under camera egomotion interference and the absence of affordance labels to explicitly guide the optimization of hand waypoint distribution. In this work, we propose a novel hand trajectory prediction method dubbed MADiff, which forecasts future hand waypoints with diffusion models. The devised denoising operation in the latent space is achieved by our proposed motion-aware Mamba, where the camera wearer's egomotion is integrated to achieve motion-driven selective scan (MDSS). To discern the relationship between hands and scenarios without explicit affordance supervision, we leverage a foundation model that fuses visual and language features to capture high-level semantics from video clips. Comprehensive experiments conducted on five public datasets with the existing and our new evaluation metrics demonstrate that MADiff predicts comparably reasonable hand trajectories compared to the state-of-the-art baselines.
Junyi Ma, Xieyuanli Chen, Wentao Bao, Hesheng Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Novel Diffusion Models for Multimodal 3D Hand Trajectory Prediction
abstract
Predicting hand motion is critical for understanding human intentions and bridging the action space between human movements and robot manipulations. Existing hand trajectory prediction (HTP) methods forecast the future hand waypoints in 3D space conditioned on past egocentric observations. However, such models are only designed to accommodate 2D egocentric video inputs. There is a lack of awareness of multimodal environmental information from both 2D and 3D observations, hindering the further improvement of 3D HTP performance. In addition, these models overlook the synergy between hand movements and headset camera egomotion, either predicting hand trajectories in isolation or encoding egomotion only from past frames. To address these limitations, we propose novel diffusion models (MMTwin) for multimodal 3D hand trajectory prediction. MMTwin is designed to absorb multi-modal information as input encompassing 2D RGB images, 3D point clouds, past hand waypoints, and text prompt. Besides, two latent diffusion models, the egomotion diffusion and the HTP diffusion as twins, are integrated into MMTwin to predict camera egomotion and future hand trajectories concurrently. We propose a novel hybrid Mamba-Transformer module as the denoising model of the HTP diffusion to better fuse multimodal features. The experimental results on three publicly available datasets and our self-recorded data demonstrate that our proposed MMTwin can predict plausible future 3D hand trajectories compared to the state-of-the-art baselines, and generalizes well to unseen environments. The code and pretrained models will be released at https://github.com/IRMVLab/MMTwin.
Junyi Ma, Wentao Bao, Guanzhong Sun, Xieyuanli Chen, Hesheng Wang 0001
IROS2
2025 Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection
abstract
Action detection aims to detect (recognize and localize) human actions spatially and temporally in videos. Existing approaches focus on the closed-set setting where an action detector is trained and tested on videos from a fixed set of action categories. However, this constrained setting is not viable in an open world where test videos inevitably come beyond the trained action categories. In this paper, we address the practical yet challenging Open-Vocabulary Action Detection (OVAD) problem. It aims to detect any action in test videos while training a model on a fixed set of action categories. To achieve such an open-vocabulary capability, we propose a novel method OpenMixer that exploits the inherent semantics and localizability of large vision-language models (VLM) within the family of query-based detection transformers (DETR). Specifically, the OpenMixer is developed by spatial and temporal OpertMixer blocks (S-OMB and T-OMB), and a dynamically fused alignment (DFA) module. The three components collectively enjoy the merits of strong generalization from pretrained VLMs and end-to-end learning from DETR design. Moreover, we established OVAD benchmarks under various settings, and the experimental results show that the OpenMixer performs the best over baselines for detecting seen and unseen actions. We release the codes, models, and dataset splits at https://github.com/Cogito2012/0penMixer.
Wentao Bao, Kai Li 0012, Yuxiao Chen 0002, Deep Patel, Martin Renqiang Min, Yu Kong 0001
WACV1
2024 Prompting Language-Informed Distribution for Compositional Zero-Shot Learning
Wentao Bao, Lichang Chen, Heng Huang 0001, Yu Kong 0001
ECCV (14)1
2024 Learning to Localize Actions in Instructional Videos with LLM-Based Multi-pathway Text-Video Alignment
Yuxiao Chen 0002, Kai Li 0012, Wentao Bao, Deep Patel, Yu Kong 0001, Martin Renqiang Min, Dimitris N. Metaxas
ECCV (82)3
2024 Facial Affective Behavior Analysis with Instruction Tuning
Anh Dao, Wentao Bao, Zhen Tan 0001, Tianlong Chen 0001, Huan Liu 0001, Yu Kong 0001
ECCV (18)3
2023 3D-aware Facial Landmark Detection via Multi-view Consistent Training on Synthetic Data
abstract
Accurate facial landmark detection on wild images plays an essential role in human-computer interaction, entertainment, and medical applications. Existing approaches have limitations in enforcing 3D consistency while detecting 3D/2D facial landmarks due to the lack of multi-view in-the-wild training data. Fortunately, with the recent advances in generative visual models and neural rendering, we have witnessed rapid progress towards high quality 3D image synthesis. In this work, we leverage such approaches to construct a synthetic dataset and propose a novel multi-view consistent learning strategy to improve 3D facial landmark detection accuracy on in-the-wild images. The proposed 3D-aware module can be plugged into any learning-based landmark detection algorithm to enhance its accuracy. We demonstrate the superiority of the proposed plug-in module with extensive comparison against state-of-the-art methods on several real and synthetic datasets.
Libing Zeng, Wentao Bao, Zhong Li 0007, Yi Xu 0002, Junsong Yuan 0001, Nima Khademi Kalantari
CVPR3
2023 Uncertainty-aware State Space Transformer for Egocentric 3D Hand Trajectory Forecasting
abstract
Hand trajectory forecasting from egocentric views is crucial for enabling a prompt understanding of human intentions when interacting with AR/VR systems. However, existing methods handle this problem in a 2D image space which is inadequate for 3D real-world applications. In this paper, we set up an egocentric 3D hand trajectory forecasting task that aims to predict hand trajectories in a 3D space from early observed RGB videos in a first-person view. To fulfill this goal, we propose an uncertainty-aware state space Transformer (USST) that takes the merits of the attention mechanism and aleatoric uncertainty within the framework of the classical state-space model. The model can be further enhanced by the velocity constraint and visual prompt tuning (VPT) on large vision transformers. Moreover, we develop an annotation workflow to collect 3D hand trajectories with high quality. Experimental results on H2O and EgoPAT3D datasets demonstrate the superiority of USST for both 2D and 3D trajectory forecasting. The code and datasets are publicly released: https://actionlab-cv.github.io/EgoHandTrajPred.
Wentao Bao, Libing Zeng, Zhong Li 0007, Yi Xu 0002, Junsong Yuan 0001, Yu Kong 0001
ICCV1
2022 OpenTAL: Towards Open Set Temporal Action Localization
abstract
Temporal Action Localization (TAL) has experienced remarkable success under the supervised learning paradigm. However, existing TAL methods are rooted in the closed set assumption, which cannot handle the inevitable unknown actions in open-world scenarios. In this paper, we, for the first time, step toward the Open Set TAL (OSTAL) problem and propose a general framework Open TAL based on Evidential Deep Learning (EDL). Specifically, the OpenTAL consists of uncertainty-aware action classification, actionness prediction, and temporal location regression. With the proposed importance-balanced EDL method, classification uncertainty is learned by collecting categorical evidence majorly from important samples. To distinguish the unknown actions from background video frames, the actionness is learned by the positive-unlabeled learning. The classification uncertainty is further calibrated by leveraging the guidance from the temporal localization quality. The OpenTAL is general to enable existing TAL models for open set scenarios, and experimental results on THUMOS14 and ActivityNet1.3 benchmarks show the effectiveness of our method. The code and pre-trained models are released at https://www.rit.edu/actionlab/opental.
Wentao Bao, Qi Yu 0001, Yu Kong 0001
CVPR1
2022 Towards Open Set Video Anomaly Detection
Yuansheng Zhu, Wentao Bao, Qi Yu 0001
ECCV (34)2
2021 Gradient Frequency Modulation for Visually Explaining Video Understanding Models
Xinmiao Lin, Wentao Bao, Matthew Wright 0001, Yu Kong 0001
BMVC2
2021 DRIVE: Deep Reinforced Accident Anticipation with Visual Explanation
abstract
Traffic accident anticipation aims to accurately and promptly predict the occurrence of a future accident from dashcam videos, which is vital for a safety-guaranteed self-driving system. To encourage an early and accurate decision, existing approaches typically focus on capturing the cues of spatial and temporal context before a future accident occurs. However, their decision-making lacks visual explanation and ignores the dynamic interaction with the environment. In this paper, we propose Deep ReInforced accident anticipation with Visual Explanation, named DRIVE. The method simulates both the bottom-up and top-down visual attention mechanism in a dashcam observation environment so that the decision from the pro-posed stochastic multi-task agent can be visually explained by attentive regions. Moreover, the proposed dense anticipation reward and sparse fixation reward are effective in training the DRIVE model with our improved reinforcement learning algorithm. Experimental results show that the DRIVE model achieves state-of-the-art performance on multiple real-world traffic accident datasets. Code and pre-trained model are available at https://www.rit.edu/actionlab/drive.
Wentao Bao, Qi Yu 0001, Yu Kong 0001
ICCV1
2021 Evidential Deep Learning for Open Set Action Recognition
abstract
In a real-world scenario, human actions are typically out of the distribution from training data, which requires a model to both recognize the known actions and reject the unknown. Different from image data, video actions are more challenging to be recognized in an open-set setting due to the uncertain temporal dynamics and static bias of human actions. In this paper, we propose a Deep Evidential Action Recognition (DEAR) method to recognize actions in an open testing set. Specifically, we formulate the action recognition problem from the evidential deep learning (EDL) perspective and propose a novel model calibration method to regularize the EDL training. Besides, to mitigate the static bias of video representation, we propose a plug-and-play module to debias the learned representation through contrastive learning. Experimental results show that our DEAR method achieves consistent performance gain on multiple mainstream action recognition models and benchmarks. Code and pre-trained models are available at https://www.rit.edu/actionlab/dear.
Wentao Bao, Qi Yu 0001, Yu Kong 0001
ICCV1
2021 Multiple Instance Relational Learning for Video Anomaly Detection
abstract
Most existing video anomaly detection methods are dependent on strong supervision to achieve satisfactory performance, which could be laborious and impractical. Besides, methods using weakly supervised learning less consider the relations among event proposals. To this end, we propose an anomaly event detection method by the instance-based event proposal generation and the proposal relation learning. Specifically, the event proposals are generated by sampling temporal distributions from the multiple instance learning (MIL), while relations among the proposals are captured by graph convolutional network for anomaly localization and classification. The proposed method is free from strong frame-level supervision but only requires video-level annotations. We conduct experiments on four datasets, i.e., UCF crime, UCSD-Peds, UMN, and CityScene, and show state-of-the-art performance on anomaly recognition and detection tasks.
Xiwen Dengxiong, Wentao Bao, Yu Kong 0001
IJCNN2
2020 Group Activity Prediction with Sequential Relational Anticipation Model
Junwen Chen 0001, Wentao Bao, Yu Kong 0001
ECCV (21)2
2020 Privacy Attributes-aware Message Passing Neural Network for Visual Privacy Attributes Classification
abstract
Visual Privacy Attribute Classification (VPAC) identifies privacy information leakage via social media images. These images containing privacy attributes such as skin color, face or gender are classified into multiple privacy attribute categories in VPAC. With limited works in this task, current methods often extract features from images and simply classify the extracted feature into multiple privacy attribute classes. The dependencies between privacy attributes, e.g., skin color and face typically coexist in the same image, are usually ignored in classification, which causes performance degradation in VPAC. In this paper, we propose a novel end-to-end Privacy Attributes-aware Message Passing Neural Network (PA-MPNN) to address VPAC. Privacy attributes are considered as nodes on a graph and an MPNN is introduced to model the privacy attribute dependencies. To generate representative features for privacy attribute nodes, a class-wise encoder-decoder is proposed to learn a latent space for each attribute. An attention mechanism with multiple correlation matrices is also introduced in MPNN to learn the privacy attributes graph automatically. Experimental results on the Privacy Attribute Dataset demonstrate that our framework achieves better performance than state-of-the-art methods for visual privacy attributes classification.
Hanbin Hong, Wentao Bao, Yuan Hong 0001, Yu Kong 0001
ICPR2
2020 Object-Aware Centroid Voting for Monocular 3D Object Detection
abstract
Monocular 3D object detection aims to detect objects in a 3D physical world from a single camera. However, recent approaches either rely on expensive LiDAR devices, or resort to dense pixel-wise depth estimation that causes prohibitive computational cost. In this paper, we propose an end-to-end trainable monocular 3D object detector without learning the dense depth. Specifically, the grid coordinates of a 2D box are first projected back to 3D space with the pinhole model as 3D centroids proposals. Then, a novel object-aware voting approach is introduced, which considers both the region-wise appearance attention and the geometric projection distribution, to vote the 3D centroid proposals for 3D object localization. With the late fusion and the predicted 3D orientation and dimension, the 3D bounding boxes of objects can be detected from a single RGB image. The method is straightforward yet significantly superior to other monocular-based methods. Extensive experimental results on the challenging KITTI benchmark validate the effectiveness of the proposed method.
Wentao Bao, Qi Yu 0001, Yu Kong 0001
IROS1
2020 Uncertainty-based Traffic Accident Anticipation with Spatio-Temporal Relational Learning
abstract
Traffic accident anticipation aims to predict accidents from dashcam videos as early as possible, which is critical to safety-guaranteed self-driving systems. With cluttered traffic scenes and limited visual cues, it is of great challenge to predict how long there will be an accident from early observed frames. Most existing approaches are developed to learn features of accident-relevant agents for accident anticipation, while ignoring the features of their spatial and temporal relations. Besides, current deterministic deep neural networks could be overconfident in false predictions, leading to high risk of traffic accidents caused by self-driving systems. In this paper, we propose an uncertainty-based accident anticipation model with spatio-temporal relational learning. It sequentially predicts the probability of traffic accident occurrence with dashcam videos. Specifically, we propose to take advantage of graph convolution and recurrent networks for relational feature learning, and leverage Bayesian neural networks to address the intrinsic variability of latent relational representations. The derived uncertainty-based ranking loss is found to significantly boost model performance by improving the quality of relational features. In addition, we collect a new Car Crash Dataset (CCD) for traffic accident anticipation which contains environmental attributes and accident reasons annotations. Experimental results on both public and the newly-compiled datasets show state-of-the-art performance of our model. Our code and CCD dataset are available at https://github.com/Cogito2012/UString.
Wentao Bao, Qi Yu 0001, Yu Kong 0001
ACM Multimedia1
2020 Activity-driven Weakly-Supervised Spatio-Temporal Grounding from Untrimmed Videos
abstract
In this paper, we study the problem of weakly-supervised spatio-temporal grounding from raw untrimmed video streams. Given a video and its descriptive sentence, spatio-temporal grounding aims at predicting the temporal occurrence and spatial locations of each query object across frames. Our goal is to learn a grounding model in a weakly-supervised fashion, without the supervision of both spatial bounding boxes and temporal occurrences during training. Existing methods have been addressed in trimmed videos, but their reliance on object tracking will easily fail due to frequent camera shot cut in untrimmed videos. To this end, we propose a novel spatio-temporal multiple instance learning framework for untrimmed video grounding. Spatial MIL and temporal MIL are mutually guided to ground each query to specific spatial regions and the occurring frames of a video. Furthermore, an activity described in the sentence is captured to use the informative contextual cues for region proposals refinement and text representation. We conduct extensive evaluation on YouCookII and RoboWatch datasets, and demonstrate our method outperforms state-of-the-art methods.
Junwen Chen 0001, Wentao Bao, Yu Kong 0001
ACM Multimedia2
2020 Human scanpath prediction based on deep convolutional saccadic model
Wentao Bao, Zhenzhong Chen 0001
Neurocomputing1
2020 MonoFENet: Monocular 3D Object Detection With Feature Enhancement Networks
abstract
Monocular 3D object detection has the merit of low cost and can be served as an auxiliary module for autonomous driving system, becoming a growing concern in recent years. In this paper, we present a monocular 3D object detection method with feature enhancement networks, which we call MonoFENet. Specifically, with the estimated disparity from the input monocular image, the features of both the 2D and 3D streams can be enhanced and utilized for accurate 3D localization. For the 2D stream, the input image is used to generate 2D region proposals as well as to extract appearance features. For the 3D stream, the estimated disparity is transformed into 3D dense point cloud, which is then enhanced by the associated front view maps. With the RoI Mean Pooling layer, 3D geometric features of RoI point clouds are further enhanced by the proposed point feature enhancement (PointFE) network. The region-wise features of image and point cloud are fused for the final 2D and 3D bounding boxes regression. The experimental results on the KITTI benchmark reveal that our method can achieve state-of-the-art performance for monocular 3D object detection.
Wentao Bao, Zhenzhong Chen 0001
IEEE Trans. Image Process.1
2017 Group Lasso-Based Band Selection for Hyperspectral Image Classification
abstract
Band selection plays an important role in reducing the dimensionality of spectral response of hyperspectral images (HSIs) to avoid dimension disaster for land-cover classification. Compared with traditional dimension reduction methods, such as principal component analysis, independent component analysis, or linear discriminant analysis, band selection can help provide interpretability to later constructed models by preserving the physical meaning of selected features. In this letter, a group lasso-based band selection (GLBS) method is proposed for multilabel HSI classification. Using the group lasso algorithm, the two objectives of band selection and classification are implemented simultaneously. The performance of GLBS is fully investigated and compared with benchmark methods, and the experimental results demonstrate the superiority of GLBS.
Daiqin Yang, Wentao Bao
IEEE Geosci. Remote. Sens. Lett.2
2014 Dandelion: A locally-high-performance and globally-high-scalability hierarchical data center network
abstract
The increasing customer demand is driving modern data centers to embrace the freely-expandable network architecture. Unfortunately, state-of-the-art freely-expandable networks suffer from either the large granularity of expansion or the prohibitive implementation cost. Furthermore, a recent research showed that data center traffic tends to be highly clustered. Based on above observations, this paper proposes a freely-expandable network architecture, namely the dandelion. Dandelion is a two-level hierarchical network, where the first level aims at “high performance” and the second level aims at “high scalability”. The resulting network has two distinct advantages. First, it could arbitrarily expand with a reasonable granularity. Second, the router architecture is efficient as well as highly scalable since 1) the routing table is significantly compressed and 2) a fixed number of virtual channels per physical channel are required regardless of the network size. Finally, the traffic characteristics of four typical cloud applications are analyzed, and the generated traffic patterns are used to evaluate the proposed network architecture. Simulation results prove that the dandelion is a promising network architecture for future data centers.
Binzhang Fu, Wentao Bao, Guolong Jiang, Mingyu Chen 0001, Lixin Zhang 0002, Yidong Tao, Junfeng Zhao 0003
ICCCN3
2014 A High-Performance and Cost-Efficient Interconnection Network for High-Density Servers
Wentao Bao, Binzhang Fu, Mingyu Chen 0001, Lixin Zhang 0002
J. Comput. Sci. Technol.1