Houlin Wang

dblp:359/7969 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0002-7609-9348ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 SyDiM: Synergistic Diffusion-adversarial Method for socially-compliant multimodal pedestrian trajectory forecasting
abstract
Predicting pedestrian trajectories in crowded and dynamic scenes is crucial for risk aware decision making in safety-critical applications. In practice, it remains challenging because trajectory forecasting must jointly model inter-pedestrian interactions to avoid socially/physically implausible futures and capture the inherent multimodality and uncertainty of human motion. To address these challenges, we propose Synergistic Diffusion-adversarial Method (SyDiM), a pioneering generative framework that seamlessly integrates diffusion-based probabilistic modeling with adversarial training strategy in pedestrian trajectory forecasting. Our method leverages diffusion to preserve multimodality and couples it with adversarial social constraints to ensure physically and socially valid interactions. Specifically, we design a Contextual Refinement Module (CRM) within the generator that performs context alignment to capture spatio-temporal dynamics and social context, effectively encoding pedestrian motion patterns while refining group-level interactions in crowded scenes. To enforce social feasibility without sacrificing the multimodality afforded by diffusion, we design an Entity-aware Spatio-temporal Trajectory Discriminator (ESTD). By integrating pedestrian-specific embeddings, it is designed as a feasibility evaluator that uses spatio-temporal attention to identify socially implausible interactions. It then guides the diffusion generator by reweighting its trajectory proposals, effectively suppressing unrealistic trajectories while preserving diversity. To the best of our knowledge, ESTD is the first interaction-conditioned feasibility evaluator tailored for diffusion-based multi-proposal trajectory forecasting. It produces a feasibility score for each predicted trajectory, which is then used to reweight samples during training. Extensive experiments on NBA SportVU, ETH-UCY, and Stanford Drone Datasets demonstrate that our SyDiM achieves consistent performance improvements across all benchmarks. Code will be available at https://github.com/xdclby/SyDiM .
Shengwei Jia, Houlin Wang
Adv. Eng. Informatics6
2026 Knowledge Distillation-Based Spiking Neural Network for Online Video Action Understanding
abstract
Online action detection and anticipation aim to understand current or upcoming actions in video streams. In industry, current artificial neural network (ANN)-based methods suffer from prohibitive energy consumption, fundamentally limiting their deployment on resource-constrained industrial devices. To bridge the gap between theory and application practice of informatics in industrial environments, we propose a novel knowledge distillation-based spiking neural network (KDSNN), which synergistically integrates bioinspired spike-driven processing with knowledge distillation, significantly reducing the energy consumption. Specifically, KDSNN includes a pioneering spiking neural network (SNN) architecture for online action detection and anticipation, which combines well-designed hierarchical spike convolutional neural network (CNN) block and spike Transformer block to capture spike-driven information. To further improve the performance of our SNN while maintaining low energy consumption, we introduce the knowledge distillation paradigm, which aims to utilize an expert-level ANN as a teacher to guide our SNN. Based on this, we propose a novel distillation loss, which consists of feature distillation and logit distillation. Notably, to address the cross-domain feature alignment in feature distillation, the optimal transport theory is employed to realize cross-domain knowledge transfer for the first time by minimizing the Wasserstein distance between continuous features (ANNs) and discrete features (SNNs). Through extensive evaluations on THUMOS14 and EPIC-Kitchen-100 datasets, the energy consumption of our KDSNN is only 27.1% and 10.0% of the state-of-the-art ANN-based method MAT. Equally importantly, the parameter count of our KDSNN is only 37.0% and 27.7% of MAT on THUMOS14 and EPIC-Kitchen-100, respectively.
Houlin Wang, Xueqiang Han, Kuo Pang, Qixian Zhang
IEEE Trans. Ind. Informatics1
2025 Multi-scale Graph Convolutional Network for understanding human action in videos
Houlin Wang, Qing Tian 0003, Bingchun Luo, Xueqiang Han
Adv. Eng. Informatics1
2025 DSAA: Cross-modal transferable double sparse adversarial attacks from images to videos
Xueqiang Han, Houlin Wang, Qing Tian 0003
Neurocomputing3
2025 CGCN: Context graph convolutional network for few-shot temporal action localization
Houlin Wang, Xueqiang Han, Qing Tian 0003
Inf. Process. Manag.2
2025 CFENet: Context-aware Feature Enhancement Network for efficient few-shot object counting
Gangzheng Zhai, Kun Chen 0006, Houlin Wang, Shaojie Han 0001
Image Vis. Comput.4
2025 Transferable targeted adversarial attack via multi-source perturbation generation and integration
Shaojie Han 0001, Xueqiang Han, Junbin Su, Gangzheng Zhai, Houlin Wang
J. Vis. Commun. Image Represent.7
2024 Opnet: Deep Occlusion Perception Network with Boundary Awareness for Amodal Instance Segmentation
abstract
The Amodal Instance Segmentation (AIS) task aims to infer the visible and occluded regions of an object instance. Existing AIS methods typically focus on directly predicting visible and occluded regions or leveraging prior knowledge to guide predictions. However, these methods often ignore the perception of occluded views, leading to inaccurate results. To address this issue and achieve high-quality AIS, we propose a boundary-aware Occlusion Perception Network (OPNet). OPNet consists of three main components: the Dynamic Feature Augmentation Pyramid (DFAP), the Dual-path Boundary Aware Module (DBAM), and the Shape-guide Refinement Module (SRM). Specifically, DBAM employs an occlusion-perception strategy to learn discriminative features with boundary information, enabling it to distinguish occlusion from multiple views. Additionally, DFAP and SRM optimize the results by enhancing feature aggregation and imposing geometric constraints. Experiments on the D2SA, KINS, and CWALT datasets show that OPNet significantly outperforms state-of-the-art AIS methods that without prior knowledge. Code is available at https://github.com/ZitengXue/OPNet.
Ziteng Xue, Houlin Wang
ICASSP4
2024 Exploiting relation of video segments for temporal action detection
Houlin Wang, Dianlong You
Adv. Eng. Informatics1
2024 Temporal action detection in videos with generative denoising diffusion
Bingchun Luo, Houlin Wang
Knowl. Based Syst.3
2023 DL-NET: Dilation Location Network for Temporal Action Detection
abstract
Temporal Action Detection(TAD) is a challenge task in video understanding. The current methods mainly use global features for boundary matching or predefine all possible proposals, while ignoring long context information and local action boundary features, resulting in the decline of detection accuracy. To fill this gap, we propose a Dilation Location Network (DL-Net) model to generate more precise action boundaries by enhancing boundary features of actions and aggregating long contextual information in this paper. Specifically, we design the boundary feature enhancement (BFE) block, which strengthens the actions boundary feature and fuses the similar feature of the different channels by pooling and channel squeezing. Meanwhile, in action location, we design multiple dilated convolutional structures to aggregate long contextual information of time point/interval. We conduct extensive experiments on ActivityNet-1.3 and Thumos14 show that DL-Net is capable of enhancing action boundary features and aggregating long contextual information effectively.
Dianlong You, Houlin Wang
ICASSP2