EDBT 2026 Demo / reviewers in the wild / expert
Yu Wang 0174
dblp:02/5889-174
· DBLP profile ↗
11ranked-venue papers
10as first author
11since 2021 · last 2026
0000-0001-7099-4424ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Local Semantic Signals and Inter-Class Discrepancy for Weakly Supervised Video Anomaly DetectionabstractWeakly supervised video anomaly detection (WSVAD) aims to retrieve temporal intervals containing anomalous events within untrimmed videos, leveraging video-level annotations. The existing methodologies exhibit unsatisfactory performance, primarily attributed to the absence of frame-level annotations. To alleviate this issue, we argue that representations exhibiting high similarity within local regions can offer dependable semantic knowledge. With this insight, we introduce a novel pseudo-label optimization mechanism grounded in the similarity of local features, specifically designed to direct framelevel anomaly predictions. Besides, we have noticed that the majority of videos utilized for anomaly detection originate from surveillance scenarios and are captured by stationary cameras. Consequently, the frames within a given video share consistent background properties. Motivated by this observation, we propose an intra-frame background-foreground separation strategy for each frame to extract discriminative visual representations. Given the considerable similarity in the backgrounds of most frames, the distinctions among the features of diverse frames become subtly indistinct. As a result, to ensure the rationality of predictions, we encourage maximizing the inter-class variance between normal and abnormal frames. Extensive experiments and ablation studies, encompassing both coarse-grained and fine-grained, have been conducted on XD-Violence, UCF-Crime, and ShanghaiTech benchmark datasets. The results demonstrate that our proposed method achieves substantial improvements, outperforming the current state-of-the-art approaches. Yu Wang 0174, Shengjie Zhao 0001, Jianyu Wang 0003, Xutao Chu |
IEEE Trans. Multim. | 1 |
| 2025 | Learning Event Completeness for Weakly Supervised Video Anomaly DetectionabstractWeakly supervised video anomaly detection (WS-VAD) is tasked with pinpointing temporal intervals containing anomalous events within untrimmed videos, utilizing only video-level annotations. However, a significant challenge arises due to the absence of dense frame-level annotations, often leading to incomplete localization in existing WS-VAD methods. To address this issue, we present a novel LEC-VAD, Learning Event Completeness for Weakly Supervised Video Anomaly Detection, which features a dual structure designed to encode both category-aware and category-agnostic semantics between vision and language. Within LEC-VAD, we devise semantic regularities that leverage an anomaly-aware Gaussian mixture to learn precise event boundaries, thereby yielding more complete event instances. Besides, we develop a novel memory bank-based prototype learning mechanism to enrich concise text descriptions associated with anomaly-event categories. This innovation bolsters the text’s expressiveness, which is crucial for advancing WS-VAD. Our LEC-VAD demonstrates remarkable advancements over the current state-of-the-art methods on two benchmark datasets XD-Violence and UCF-Crime. Yu Wang 0174 |
ICML | 1 |
| 2025 | Point4Bit: Post Training 4-bit Quantization for Point Cloud 3D DetectionabstractVoxel-based 3D object detectors have achieved remarkable performance in point cloud perception, yet their high computational and memory demands pose significant challenges for deployment on resource-constrained edge devices. Post-training quantization (PTQ) provides a practical means to compress models and accelerate inference; however, existing PTQ methods for point cloud detection are typically limited to INT8 and lack support for lower-bit formats such as INT4, which restricts their deployment potential. In this paper, we present Point4bit, the first general 4-bit PTQ framework tailored for voxel-based 3D object detectors. To tackle challenges in low-bit quantization, we propose two key techniques: (1) Foreground-aware Piecewise Activation Quantization (FA-PAQ), which leverages foreground structural cues to improve the quantization of sparse activations; and (2) Gradient-guided Key Weight Quantization (G-KWQ), which preserves task-critical weights through gradient-based analysis to reduce quantization-induced degradation. Extensive experiments demonstrate that Point4bit achieves INT4 quantization with minimal accuracy loss with less than 1.5\% accuracy drop. Moreover, we validate its generalization ability on point cloud classification and segmentation tasks, demonstrating broad applicability. Our method further advances the bit-width limitation of point cloud quantization to 4 bits, demonstrating strong potential for efficient deployment on resource-constrained edge devices. Jianyu Wang 0003, Yu Wang 0174, Shengjie Zhao 0001, Sifan Zhou |
NeurIPS | 2 |
| 2025 | SQL-Net: Semantic Query Learning for Point-Supervised Temporal Action LocalizationabstractPoint-supervised Temporal Action Localization (PS-TAL) detects temporal intervals of actions in untrimmed videos with a label-efficient paradigm. However, most existing methods fail to learn action completeness without instance-level annotations, resulting in fragmentary region predictions. In fact, the semantic information of snippets is crucial for detecting complete actions, meaning that snippets with similar representations should be considered as the same action category. To address this issue, we propose a novel representation refinement framework with a semantic query mechanism to enhance the discriminability of snippet-level features. Concretely, we set a group of learnable queries, each representing a specific action category, and dynamically update them based on the video context. With the assistance of these queries, we expect to search for the optimal action sequence that agrees with their semantics. Besides, we leverage some reliable proposals as pseudo labels and design a refinement and completeness module to refine temporal boundaries further, so that the completeness of action instances is captured. Finally, we demonstrate the superiority of the proposed method over existing state-of-the-art approaches on THUMOS14 and ActivityNet13 benchmarks. Notably, thanks to completeness learning, our algorithm achieves significant improvements under more stringent evaluation metrics. Yu Wang 0174, Shengjie Zhao 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | Weakly-Supervised Action Localization by Hierarchical Attention Mechanism with Multi-Scale Fusion StrategiesabstractWeakly-supervised temporal action localization focuses on locating action intervals when merely video-level supervised signals are available. Conventional methods mostly rely on the attention framework, which generates a set of scores indicating the confidence that the video snippet belongs to the foreground, the background, and the context, respectively. However, such methods fail to consider the structural properties of snippet-level features when generating attention scores, and these structural properties are critical for capturing contextual information in temporal tasks. To this end, we propose a hierarchical attention generation mechanism with multi-scale fusion strategies to model such structural information. Besides, to resolve action-context confusion issues that are quite intractable in weakly-supervised action localization tasks, metric learning is further introduced into our framework to suppress context features from approaching action features, while encouraging them to be close to background features. Finally, our model is evaluated on THUMOS14 and ActivityNet1.3 benchmarks, and the results demonstrate that the proposed approach achieves desirable performance. Yu Wang 0174, Shengjie Zhao 0001 |
ICME | 1 |
| 2024 | Action-Semantic Consistent Knowledge for Weakly-Supervised Action LocalizationabstractWeakly-supervised temporal action localization aims to detect temporal intervals of actions in arbitrarily long untrimmed videos with only video-level annotations. Owing to label sparsity, learning action consistency is intractable. In this paper, we assume that frames with similar representations in a given video should be considered as the same action. To this end, we develop a query-based contrastive learning paradigm to ensure action-semantic consistency. This mechanism encourages normalized embeddings with the same class to be pulled closer together, while embeddings from different classes are repelled apart. Besides, we design a two-branch framework, consisting of a class-aware branch and a class-agnostic branch, to learn salient features and fine-grained clues respectively. To further guarantee the action-semantic consistency of the two branches, unlike previous methods that handle each branch independently, we model the relationship between the two branches to avoid unreasonable predictions. Finally, the proposed model demonstrates superior performance over existing methods on the publicly available THUMOS-14 and ActivityNet-1.3 datasets. Substantial experiments and ablation studies also demonstrate the effectiveness of our model. Yu Wang 0174, Shengjie Zhao 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | Two-Stream Networks for Weakly-Supervised Temporal Action Localization with Semantic-Aware MechanismsabstractWeakly-supervised temporal action localization aims to detect action boundaries in untrimmed videos with only video-level annotations. Most existing schemes detect tem-poral regions that are most responsive to video-level classification, but they overlook the semantic consistency between frames. In this paper, we hypothesize that snippets with similar representations should be considered as the same action class despite the absence of supervision signals on each snippet. To this end, we devise a learnable dictionary where entries are the class centroids of the corresponding action categories. The representations of snippets identified as the same action category are induced to be close to the same class centroid, which guides the network to perceive the semantics of frames and avoid unreasonable localization. Besides, we propose a two-stream framework that in-tegrates the attention mechanism and the multiple-instance learning strategy to extract fine-grained clues and salient features respectively. Their complementarity enables the model to refine temporal boundaries. Finally, the developed model is validated on the publicly available THUMOS-14 and ActivityNet-1.3 datasets, where substantial experiments and analyses demonstrate that our model achieves remark-able advances over existing methods. Yu Wang 0174 |
CVPR | 1 |
| 2023 | Multi-Agent Trajectory Prediction With Spatio-Temporal Sequence FusionabstractAccurate trajectory prediction of surrounding agents is an important issue for building up an intelligent transportation system. Frequent interactions among agents have a major impact on their movement patterns. Current research mainly relies on agents’ spatial structure associated with the last frame of the observation to model social interactions, while paying less attention to structure information from previous moments. In addition, existing methods merely consider temporal features of a single trajectory sequence, while neglecting temporal dependencies across multiple trajectories. In this work, we endeavor to capture comprehensively social interactions among agents with the proposed Spatio-Temporal Sequence Fusion Network (STSF-Net). Specifically, we construct a spatio-temporal sequence that encodes contextual information taking explicitly spatial distributions of agents during movement into account while capturing socially temporal dependencies across multiple trajectory sequences. Besides, a social recurrent mechanism is introduced to explicitly capture temporal correlations between interactions by concerning spatial structure at each time-step. Finally, our model is evaluated on datasets covering pedestrian, vehicle, and heterogeneous multi-agent trajectories. Experimental evidence manifests that our method achieves excellent performance. Yu Wang 0174 |
IEEE Trans. Multim. | 1 |
| 2022 | Multi-Vehicle Collaborative Learning for Trajectory Prediction With Spatio-Temporal Tensor FusionabstractAccurate behavior prediction of other vehicles in the surroundings is critical for intelligent transportation systems. Common practices to reason about the future trajectory are through their historical paths. However, the impact of traffic context is ignored, which means the beneficial environment information is deserted. Although a few methods are proposed to exploit the surrounding vehicle information, they simply model the influence according to spatial relations without considering the temporal information among them. In this paper, a novel multi-vehicle collaborative learning with spatio-temporal tensor fusion model for vehicle trajectory prediction is proposed, which introduces a novel auto-encoder social convolution mechanism and a fancy recurrent social mechanism to model spatial and temporal information among multiple vehicles, respectively. Furthermore, the generative adversarial network is incorporated into our framework to handle the inherent multi-modal characteristics of the agent motion behavior. Finally, we evaluate the proposed multi-vehicle collaborative learning model on NGSIM US-101 and I-80 benchmark datasets. Experimental results demonstrate that the proposed approach outperforms the state-of-the-art for vehicle trajectory prediction. Additionally, we also present qualitative analyses of the multi-modal vehicle trajectory generation and the impacts of surrounding vehicles on trajectory prediction under various circumstances. Yu Wang 0174, Shengjie Zhao 0001, Rongqing Zhang 0001, Xiang Cheng 0001, Liuqing Yang 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2021 | Cross-Modal Representation Reconstruction for Zero-Shot ClassificationabstractZero-shot learning (ZSL) aims to recognize novel classes without training samples through transferring knowledge from seen classes, based on the assumption that both the seen and unseen classes share a latent semantic space. Previous works either focus on directly learning various mapping functions between visual space and semantic space, or searching a latent common subspace to alleviate semantic gap between different modalities. However, few methods directly learn modality-invariant representations for ZSL. In this paper, we propose a Cross-Modal Representation Reconstruction (CM-RR) framework to bridge the semantic gap between visual features and semantic attributes, as well as introducing a novel regularizer for automatically feature selection. Moreover, an iterative optimization process based on the ALM (Augmented Lagrangian Method) algorithm with the alternating direction strategy is developed to solve the proposed formulation. Extensive experiments on four benchmark datasets show the effectiveness of the proposed approach, and even the performance surpasses some deep learning based methods. Yu Wang 0174, Shenjie Zhao |
ICASSP | 1 |
| 2021 | Consistent Representation Learning Across Modalities for Zero-Shot Image RecognitionabstractZero-shot learning (ZSL) recently has drawn widespread attention due to the demand for scalability of object recognition in real scenes. Existing approaches typically focus on directly learning various mapping functions from the visual space to the semantic space or vice versa. However, these methods fail to explicitly capture consistent representations reflecting the nature of different modalities of the same object, which results in the deterioration of the domain shift problem in the ZSL context during the testing stage. To this end, a consistent representation learning mechanism across visual modalities and the associated semantic modalities based on common subspace learning is proposed in this paper. We further impose an orthogonal constraint on the subspace for informative representations, as well as ℓ21-norm regularization terms on projection matrices for automatic feature selection. Finally, an iterative process based on the ALM algorithm with an alternating direction strategy is displayed to resolve the proposed formulation. Extensive experimental results on four popular datasets demonstrate that our algorithm is promising. Yu Wang 0174, Sev-Jue Zhao |
ICME | 1 |