Zhe Wu 0006

dblp:11/5670-6 · DBLP profile ↗
← Back
32ranked-venue papers
6as first author
23since 2021 · last 2026
0000-0002-6982-2315ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 6 first-author · 15 since 2021Artificial intelligence and machine learning · 15 · 2 first-author · 10 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 DMGINE: Day-Memory Guided Nighttime Image Enhancement for Dynamic Traffic Scenes
abstract
We introduce Daytime-Memory Guided Nighttime Image Enhancement (DMGNIE) framework, the first framework that turns long-running daytime surveillance videos of a single intersection into persistent “daytime memory” to guide nighttime image enhancement in traffic scenes. Our key insight is simple yet powerful: for a static scene, perfectly exposed daytime frames are, pixel-for-pixel, high-quality illumination prior for the same location under extreme low-light. Due to the complex lighting conditions in real-world traffic scenes, existing low-light image enhancement (LLIE) methods suffer from issues such as overexposure in highlight regions and noise amplification in low-light condition regions, which degrades the performance of downstream computer vision tasks. DMGNIE tackles these issues in two steps: (1) SegBMN, a semantic prior-based background modeling network, distills a clean, static daytime background from hours of video as scene prior guiding the enhancement of nighttime image; (2) a Foreground Localization-Guided Contrastive Learning module avoid the interference from the background prior with foreground objects during the guidance by maximizing the differences between foreground and background features. Finally, We conduct comprehensive experiments on real traffic surveillance datasets of two cities to evaluate the effectiveness. And the experimental results demonstrate that DMGNIE outperforms state-of-the-art baselines and achieves superior performance in challenging low-light conditions.
Ruizhou Liu, Zhe Wu 0006, Zimo Liu, Qingming Huang
AAAI2
2026 Cross-City Correlation Learning for Traffic Forecasting
abstract
Traffic forecasting is essential in city-level applications, where data-driven deep learning has become the most popular method. However, sufficient data in developing cities is not always accessible, posing a challenge for training effective models in scenarios with limited data. Recently, several works have promoted this issue through cross-city knowledge transfer and shown promising performances. However, existing methods can neither distinguish node divergence nor extract functional similarities between cities, which results in suboptimal performance. To overcome the limitations, we propose a Cross-city Correlation Learning (CCL) framework. Firstly, we construct a self-supervised learning model to infer accurate node-to-node and node-to-region cross-city correlations from multiple noisy labels without using any auxiliary information. Then, we achieve spatial knowledge transfer from a transfer-adaptive graph convolution network based on the learned correlations in two aspects: the learnable adjacency matrix and region-specific kernel parameters, which ensure the target models can transfer more and better utilize the knowledge from the source domain. The experiments are conducted on six real-world datasets and fully prove the effectiveness of the proposed framework.
Zhe Wu 0006, Li Su 0003, Xinfeng Zhang 0001, Yaowei Wang 0001, Qingming Huang
IEEE Trans. Intell. Transp. Syst.2
2026 High-Frequency Prioritized Sparse Attention Network for Image Restoration
abstract
Image restoration aims to restore high-quality images from degraded inputs caused by factors such as motion blur, defocus blur, and rain, where the primary difference between degraded and high-quality images lies in their high-frequency components. Despite the critical role of high frequencies in restoration, few methods explicitly prioritize computational resources for high frequencies over low frequencies. To address this issue, we propose a High-Frequency Prioritized Sparse Attention Network (HFP-SAN), a novel architecture for image restoration tasks. We explicitly prioritize high-frequency components by designing a symmetric encoder-decoder framework integrated with High-Frequency Selective Sparse Attention (HFSSA) modules while handling low-frequency components with a smaller residual network, thereby proportionally allocating computational resources based on their relative importance. HFSSA incorporates a Frequency-Selective Matching (FSM) algorithm to focus attention on strongly correlated high-frequency regions, mitigating computation on areas with weak correlations and irrelevant areas. Additionally, we introduce a dynamically adjustable high-frequency mask that guides the network to focus on the severely degraded regions, further refining restoration quality. The above designs ensure the final reconstructed image is a high-quality product. Experiments demonstrate that our HFP-SAN achieves state-of-the-art performance across multiple image restoration tasks, both quantitatively and qualitatively.
Shuting Dong, Zhe Wu 0006, Hongyang Wei, Mingzhi Chen 0003, Guanghao Li 0003, Haolong Qian, Hanyang Peng, Chun Yuan 0003
IEEE Trans. Multim.2
2026 DMutDE: Dual-View Mutual Distillation Framework for Knowledge Graph Embeddings
abstract
Knowledge graphs (KGs) have caught more and more attention in recent years. Currently, in some practical scenarios, KG embedding (KGE) models are expected to reduce their spatial complexity without losing much performance to address the challenges of storage limitations and knowledge reasoning efficiency. To achieve this, existing works use one or more large and high-performance teacher models to improve the performance of a lightweight student model via knowledge distillation (KD), thus meeting the requirements of some practical complicated applications. However, in resource-constrained scenarios, obtaining high-performance teacher models is challenging due to high training costs and significant storage requirements. Thus, enhancing the student model's performance without large teacher models is crucial. To address this issue, we propose Dual-View Mutual Distillation Framework for Knowledge Graph Embeddings (DMutDE), a distillation framework leveraging mutual learning for peer-to-peer distillation between two KGE models with different architectures. In KGE models, we notice that the way of modeling relational directed edges determines the model view of KGE model for learning KG data. Thus, integrating the model views from two different KGE models by KD into a student KGE model can improve its generalization, so as to increase its performance. To identify an effective dual-view fusion method, we design two modules in the DMutDE framework. Specifically, we design a novel soft-label fusion (SLF) module for noise filtering and response knowledge transfer. Then, we propose an entity embedding distillation (EED) module to distill structural features from each other. Finally, we conduct several comprehensive experiments on the standard open-source benchmarks to demonstrate that our framework achieves the state-of-the-art results. The code is available at https://github.com/RuizhouLiu/DMutDE.
Ruizhou Liu, Zhe Wu 0006, Yiling Wu, Zongsheng Cao, Qianqian Xu 0001, Qingming Huang
IEEE Trans. Neural Networks Learn. Syst.2
2025 Mixture-of-KAN for Multivariate Time Series Forecasting
abstract
Multivariate time series forecasting is a crucial task that predicts the future states based on historical inputs. Although current deep learning-based methods have made significant advancements, they still face the criticism of lacking interpretability. The rise of the Kolmogorov-Arnold Network (KAN) provides a new perspective to implement an efficient and interpretable deep learning-based method for forecasting time series. However, we find there are two main challenges in the application of KAN in time series forecasting: how to select the appropriate one from various KAN variants and how to train the deep KAN-based network. To this end, we propose the multi-layer mixture-of-KAN network, which achieves excellent performance while retaining KAN's ability to be transformed into a combination of symbolic functions. The core module is the mixture-of-KAN layer, which uses a mixture-of-experts structure to assign variables to best-matched KAN experts. Then, we analyze the shortcomings of parameter initialization in the original KAN and provide an effective initialization method to alleviate training instability. Extensive experimental results demonstrate that our proposed method is effective in multivariate time series forecasting. Codes are released in https://github.com/2448845600/EasyTSF.
Zhenduo Zhang, Xinfeng Zhang 0001, Yiling Wu, Zhe Wu 0006
CIKM5
2025 Extracting Global Temporal Patterns Within Short Look-Back Windows for Traffic Forecasting
abstract
With the continuous expansion of urban areas, accurate and effective traffic forecasting has become essential for intelligent urban traffic management. As traffic data inherently exhibits temporal dynamics, modeling its temporal patterns is critical to improve prediction performance. However, constrained by computational complexity, existing methods rely primarily on short-term historical data, which is typically noisy and limits the ability to capture global temporal patterns. To address this issue, we propose a novel Dual-Stream Transformer model (DSformer) that effectively captures global temporal patterns through a time-index model. To mitigate the impact of noise in short look-back windows, DSformer explicitly learns a temporal matrix that encodes structured temporal dependencies. Furthermore, we design a time-index loss that encourages similar representations for adjacent time indices, thereby reducing error propagation across time steps. In parallel, a historical-value stream is employed to model local information. Finally, a self-adaptive learning module is constructed to flexibly and accurately fuse global and local information. Extensive experiments on real-world traffic forecasting tasks across ten diverse scenarios demonstrate that our method consistently outperforms state-of-the-art baselines while maintaining competitive efficiency. The code is available at https://github.com/sky836/DSFormer.git.
Zhe Wu 0006, Li Su 0003
CIKM2
2025 Vpr-Cloak: a First Look at Privacy Cloak Against Visual Place Recognition
Shuting Dong, Mingzhi Chen 0003, Guanghao Li 0003, Zhe Wu 0006, Ming Tang 0006, Chun Yuan 0003
ICCV6
2025 Towards A Real-World Road Damage Detection Dataset
abstract
Road damage represents a serious challenge to the health of road infrastructure and driving safety, making deep learning-based image analysis for road damage detection (RDD) an important research focus. The limited diversity in road damage types, road image collection, size, environment, and imperfect damage definitions within current RDD datasets restrict the real-world applications of RDD. To address this issue, this paper constructs PCL-RDD, a new and extensive RDD dataset. It comprises 24,765 road images, 54,732 instances, and 19 types of road damage. Compared to the existing datasets that mainly include common road damage, the proposed dataset contains a variety of rare and urgent road damages. Besides, we collect road facility-related damages, which also affect traffic safety. We evaluate eight well-established object detection algorithms on the dataset, highlighting the limitations of state-of-the-art detection algorithms under complex conditions. This study contributes a significant dataset to the RDD field and can advance artificial intelligence in both city infrastructure management and environmental perception for autonomous driving. The dataset is available at https://github.com/humh-c/PCL-RDD.
Menghao Hu, Zuogan Tang, Xiaoshan Yang, Zhe Wu 0006, Zhouxin Yang, Shaocong Wu, Yaguang Song, Kui Hou, Yaowei Wang 0001
ICME4
2025 Perceptually Constrained Precipitation Nowcasting Model
abstract
Most current precipitation nowcasting methods aim to capture the underlying spatiotemporal dynamics of precipitation systems by minimizing the mean square error (MSE). However, these methods often neglect effective constraints on the data distribution, leading to unsatisfactory prediction accuracy and image quality, especially for long forecast sequences. To address this limitation, we propose a precipitation nowcasting model incorporating perceptual constraints. This model reformulates precipitation nowcasting as a posterior MSE problem under such constraints. Specifically, we first obtain the posteriori mean sequences of precipitation forecasts using a precipitation estimator. Subsequently, we construct the transmission between distributions using rectified flow. To enhance the focus on distant frames, we design a frame sampling strategy that gradually increases the corresponding weights. We theoretically demonstrate the reliability of our solution, and experimental results on two publicly available radar datasets demonstrate that our model is effective and outperforms current state-of-the-art models.
Wenzhi Feng, Xutao Li 0003, Zhe Wu 0006, Kenghong Lin, Demin Yu, Yunming Ye, Yaowei Wang 0001
ICML3
2025 A Driving-Style-Adaptive Framework for Vehicle Trajectory Prediction
abstract
Vehicle trajectory prediction serves as a critical enabler for autonomous navigation and intelligent transportation systems. While existing approaches predominantly focus on temporal pattern extraction and vehicle-environment interaction modeling, they exhibit a fundamental limitation in addressing trajectory heterogeneity originating from human driving styles. This oversight constrains prediction reliability in complex real-world scenarios. To bridge this gap, we propose the Driving-Style-Adaptive (\underline{\textbf{DSA}}) framework, which establishes the first systematic integration of heterogeneous driving behaviors into trajectory prediction models. Specifically, our framework employs a set of basis functions tailored to each driving style to approximate the trajectory patterns. By dynamically combining and adaptively adjusting the degree of these basis functions, DSA not only enhances prediction accuracy but also provides \textbf{explanations} insights into the prediction process. Extensive experiments on public real-world datasets demonstrate that the DSA framework outperforms state-of-the-art methods.
Di Wen 0005, Zhaocheng He, Zhe Wu 0006
NeurIPS5
2025 SGKGE: Semantically Guided Knowledge Graph Embeddings via Complementary Latent Representations
abstract
Knowledge graph (KG) completion is a challenging yet essential task that has attracted increasing attention in recent years. While entities in KGs typically present complex semantics (a phenomenon known as polysemy), previous works primarily focus on holistic but often inaccurate representations of entities, neglecting the diversity of their semantics. This limitation results in suboptimal representations for entities within KGs. To address this issue, we propose a new method termed semantically guided KG embeddings (SGKGE), which captures the precise semantics of entities in KGs from a semantics-guided perspective. Specifically, SGKGE first guides the learning of holistic semantics of entities through a hyperbolic manifold with learnable shared curvature and a geometric attention-fusion module, facilitating efficient reasoning. Subsequently, SGKGE captures fine-grained semantics through a set of Cartesian product Riemannian manifolds with distinct curvatures, coupled with a semantic interactions module. This approach enables SGKGE to produce more accurate entity semantics and enhance downstream applications. Experimental results demonstrate that our model achieves state-of-the-art performance on six well-established KG completion benchmarks. The release code is available at https://github.com/RuizhouLiu/SGKGE.
Ruizhou Liu, Zongsheng Cao, Zhe Wu 0006, Yiling Wu, Qianqian Xu 0001, Qingming Huang
IEEE Trans. Neural Networks Learn. Syst.3
2024 Density-Adaptive Model Based on Motif Matrix for Multi-Agent Trajectory Prediction
abstract
Multi-agent trajectory prediction is essential in autonomous driving, risk avoidance, and traffic flow control. However, the heterogeneous traffic density on interactions, which caused by physical laws, social norms and so on, is often overlooked in existing methods. When the density varies, the number of agents involved in interactions and the corresponding interaction probability change dynami-cally. To tackle this issue, we propose a new method, called Density-Adaptive Model based on Motif Matrix for Multi-Agent Trajectory Prediction (DAMM), to gain insights into multi-agent systems. Here we leverage the motif matrix to represent dynamic connectivity in a higher-order pattern, and distill the interaction information from the perspectives of the spatial and the temporal dimensions. Specifically, in spatial dimension, we utilize multi-scale feature fusion to adaptively select the optimal range of neighbors participating in interactions for each time slot. In temporal dimension, we extract the temporal interaction features and adapt a pyramidal pooling layer to generate the interaction probability for each agent. Experimental results demonstrate that our approach surpasses state-of-the-art methods on autonomous driving dataset.
Di Wen 0005, Haoran Xu 0004, Zhaocheng He, Zhe Wu 0006, Guang Tan, Peixi Peng
CVPR4
2024 Multimodal Knowledge Graph Embeddings via Lorentz-based Contrastive Learning
abstract
Multimodal knowledge graph embeddings (MKGE) have recently garnered significant attention. Unlike traditional unimodal knowledge graph embeddings, MKGE integrates both structural and multimodal knowledge to represent entities within a unified framework. However, real-world entities exhibit heterogeneity, often resulting in semantic inconsistencies where structurally similar embeddings may diverge significantly in their multimodal representations. Previous approaches primarily focus on directly fusing structural and multimodal embeddings, thus overlooking the issue of semantic-embedding inconsistency. To tackle this issue, we propose a new multimodal knowledge graph embedding method via Lorentz-based contrastive learning (LCKGE). we firstly introduce a well-designed nearest-neighbor fusion module via contrastive learning for multimodal fusion. Then, an attention-based Lorentz transformation is proposed for capturing more complex geometric information in MKGs. Furthermore, a series of comprehensive experiments are conducted to demonstrate the effectiveness of our model. We provide the code and appendix of LCKGE in https://github.com/RuizhouLiu/LCKGE
Ruizhou Liu, Zongsheng Cao, Zhe Wu 0006, Qianqian Xu 0001, Qingming Huang
ICME3
2024 Event Traffic Forecasting with Sparse Multimodal Data
abstract
With the development of deep learning, traffic forecasting technology has made significant progress and is being applied in many practical scenarios. However, various events held in cities, such as sporting events, exhibitions, concerts, etc., have a significant impact on traffic patterns of surrounding areas, causing current advanced prediction models to fail in this case. In this paper, to broaden the applicable scenarios of traffic forecasting, we focus on modeling the impact of events on traffic patterns and propose an event traffic forecasting problem with multimodal inputs. We outline the main challenges of this problem: diversity and sparsity of events, as well as insufficient data. To address these issues, we first use textual modal data containing rich semantics to describe the diverse characteristics of events. Then, we propose a simple yet effective multi-modal event traffic forecasting model that uses pre-trained text and traffic encoders to extract the embeddings and fuses the two embeddings for prediction. Encoders pre-trained on large-scale data have powerful generalization abilities to cope with the challenge of sparse data. Next, we design an efficient large language model-based event description text generation pipeline to build multi-modal event traffic forecasting datasets, ShenzhenCEC and SuzhouIEC. Experiments on two real-world datasets show that our method achieves state-of-the-art performance compared with eight baselines, reducing mean absolute error during the event peak period by 4.26%. Code is available at: https://github.com/2448845600/EventTrafficForecasting.
Zhenduo Zhang, Yiling Wu, Xinfeng Zhang 0001, Zhe Wu 0006
ACM Multimedia5
2024 Mamba-FETrack: Frame-Event Tracking via State Space Model
Ju Huang, Shiao Wang, Zhe Wu 0006, Xiao Wang 0014, Bo Jiang 0002
PRCV (12)4
2023 Recurrent Fine-Grained Self-Attention Network for Video Crowd Counting
abstract
Striking a balance between exploring the spatio-temporal correlation and controlling model complexity is vital for video-based crowd counting methods. In this paper, we propose a Recurrent Fine-Grained Self-Attention Network (RFSNet) to achieve efficient and accurate counting in video scenes via the self-attention mechanism and a recurrent fine-tuning strategy. Specifically, we design a decoder which consists of patch-wise spatial self-attention and temporal self-attention. Compared with vanilla self-attention, it effectively leverages the dependencies in spatial and temporal domain respectively, while significantly reducing computational complexity. Moreover, the RFSNet recurrently feeds the features into the decoder to enhance the spatio-temporal representations. This strategy not only simplifies the model structure and reduces the number of parameters, but also improves the quality of estimated density maps. Our RFSNet achieves state-of-the-art performance on three video crowd counting benchmarks, and outperforms other methods by more than 20% on the challenging FDST dataset.
Jifan Zhang, Zhe Wu 0006, Xinfeng Zhang 0001, Guoli Song, Yaowei Wang 0001, Jie Chen 0001
ICASSP2
2023 DFVSR: Directional Frequency Video Super-Resolution via Asymmetric and Enhancement Alignment Network
abstract
Recently, techniques utilizing frequency-based methods have gained significant attention, as they exhibit exceptional restoration capabilities for detail and structure in video super-resolution tasks. However, most of these frequency-based methods mainly have three major limitations: 1) insufficient exploration of object motion information, 2) inadequate enhancement for high-fidelity regions, and 3) loss of spatial information during convolution. In this paper, we propose a novel network, Directional Frequency Video Super-Resolution (DFVSR), to address these limitations. Specifically, we reconsider object motion from a new perspective and propose Directional Frequency Representation (DFR), which not only borrows the property of frequency representation of detail and structure information but also contains the direction information of the object motion that is extremely significant in videos. Based on this representation, we propose a Directional Frequency-Enhanced Alignment (DFEA) to use double enhancements of task-related information for ensuring the retention of high-fidelity frequency regions to generate the high-quality alignment feature. Furthermore, we design a novel Asymmetrical U-shaped network architecture to progressively fuse these alignment features and output the final output. This architecture enables the intercommunication of the same level of resolution in the encoder and decoder to achieve the supplement of spatial information. Powered by the above designs, our method achieves superior performance over state-of-the-art models on both quantitative and qualitative evaluations.
Shuting Dong, Zhe Wu 0006, Chun Yuan 0003
IJCAI3
2023 Enhanced Image Deblurring: An Efficient Frequency Exploitation and Preservation Network
abstract
Most of these frequency-based deblurring methods mainly have two major limitations: (1) insufficient exploitation of frequency information, (2) inadequate preservation of frequency information. In this paper, we propose a novel Efficient Frequency Exploitation and Preservation Network (EFEP) to address these limitations. Firstly, we propose a novel Frequency-Balanced Exploitation Encoder (FBE-Encoder) to sufficiently exploit frequency information. We insert a novel Frequency-Balanced Navigator (FBN) module in the encoder, which establishes a dynamic balance that adaptively explores and integrates the correlations between frequency features and other features presented in the network. And it also can highlight the most important regions in frequency features. Secondly, considering the limitation that frequency information is inevitably lost in deep network architectures, we present an Enhanced Selective Frequency Decoder (ESF-Decoder) that not only effectively reduces spatial information redundancy, but also fully explores the different importance of various frequency information to ensure the supplement of valid spatial information and weaken the invalid information. Thirdly, each encoder/decoder block of the EFEP consists of multiple Contrastive Residual Blocks (CRBs), which are designed to explicitly compute and incorporate feature distinctions. Powered by the above designs, our EFEP outperforms state-of-the-art models on both quantitative and qualitative evaluations.
Shuting Dong, Zhe Wu 0006, Chun Yuan 0003
ACM Multimedia2
2023 Learning Spatial-Frequency Transformer for Visual Object Tracking
abstract
Recently, some researchers have begun to adopt the Transformer to combine or replace the widely used ResNet as their new backbone network. As the Transformer captures the long-range relations between pixels well using the self-attention scheme, which complements the issues caused by the limited receptive field of CNN. Although their trackers work well in regular scenarios, they simply flatten the 2D features into a sequence to better match the Transformer. We believe these operations ignore the spatial prior of the target object, which may lead to sub-optimal results only. In addition, many works demonstrate that self-attention is actually a low-pass filter, which is independent of input features or keys/queries. That is to say, it may suppress the high-frequency component of the input features and preserve or even amplify the low-frequency information. To handle these issues, in this paper, we propose a unified Spatial-Frequency Transformer that models the Gaussian spatial Prior and High-frequency emphasis Attention (GPHA) simultaneously. To be specific, Gaussian spatial prior is generated using dual Multi-Layer Perceptrons (MLPs) and injected into the similarity matrix produced by multiplying Query and Key features in self-attention. The output will be fed into a softmax layer and then decomposed into two components, i.e., the direct and high-frequency signal. The low- and high-pass branches are rescaled and combined to achieve all-pass, therefore, the high-frequency features will be protected well in stacked self-attention layers. We further integrate the Spatial-Frequency Transformer into the Siamese tracking framework and propose a novel tracking algorithm termed SFTransT. The cross-scale fusion based SwinTransformer is adopted as the backbone, and also a multi-head cross-attention module is used to boost the interaction between search and template features. The output will be fed into the tracking head for target localization. Extensive experiments on short-term and long-term tracking benchmarks all demonstrate the effectiveness of our proposed framework. Source code will be released athttps://github.com/Tchuanm/SFTransT.git.
Chuanming Tang, Xiao Wang 0014, Yuanchao Bai, Zhe Wu 0006, Jianlin Zhang 0001, Yongmei Huang
IEEE Trans. Circuits Syst. Video Technol.4
2023 Spatial-Temporal Graph Network for Video Crowd Counting
abstract
In recent years, researchers have developed many deep-learning-based methods to count crowd numbers in static images. However, much fewer works focus on video-based crowd counting, in which the critical challenge of temporal correlation has not been well explored. This paper proposes a Spatial-Temporal Graph Network (STGN) to achieve efficient and accurate crowd counting in videos via learning pixel-wise and patch-wise relations in local spatial-temporal domains. Specifically, we design a pyramid graph module to leverage multi-scale features. In each scale, we sequentially construct three graphs: spatial-temporal pixel graph, temporal patch graph, and spatial pixel graph, in which we apply the self-attention mechanism to capture pixel-wise relation, learn structure-aware relation, and aggregate local features, respectively. Furthermore, we propose spatial-aware channel-wise attention to effectively fuse multi-scale features. To demonstrate the effectiveness of the proposed method, we conduct experiments on five crowd counting datasets, including a large-scale video crowd dataset (FDST). Moreover, the proposed model is also applied in the vehicle counting dataset (TRANCOS). The results show that the proposed model outperforms existing spatial-temporal crowd counting models and achieves state-of-the-art. The code is available athttps://github.com/wuzhe71/STGN
Zhe Wu 0006, Xinfeng Zhang 0001, Geng Tian, Yaowei Wang 0001, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.1
2023 Robust and Hierarchical Spatial Relation Analysis for Traffic Forecasting
abstract
How to model the complex spatial-temporal relation in traffic data is an important problem for precisely predicting the future status of a city traffic system. Existing traffic forecasting methods rarely consider the traffic state trend, and the robust spatial relation has not been well explored. To tackle these issues, we design a novel Robust And Hierarchical spatial Relation Analysis (RAHRA) method to calculate the local-period spatial relation, which applies temporal context information in both traffic state and trend similarities. This could capture abundant traffic patterns and learn stable and comprehensive spatial relations for accurate traffic forecasting. Furthermore, we introduce a Temporal Attention Module (TAM) to capture the temporal features and propose a Future Feature Inference Module (FFIM) to infer the future traffic information. Experiments on four real-world traffic datasets demonstrate that the proposed method outperforms the other state-of-the-art methods.
Zhe Wu 0006, Xinfeng Zhang 0001, Guoli Song, Yaowei Wang 0001, Jie Chen 0001
IEEE Trans. Intell. Transp. Syst.2
2021 Two-Stage Polishing Network for Camouflaged Object Detection
Zhe Wu 0006, Li Su 0003, Qingming Huang
ICIG (1)2
2021 Decomposition and Completion Network for Salient Object Detection
abstract
Recently, fully convolutional networks (FCNs) have made great progress in the task of salient object detection and existing state-of-the-arts methods mainly focus on how to integrate edge information in deep aggregation models. In this paper, we propose a novel Decomposition and Completion Network (DCN), which integrates edge and skeleton as complementary information and models the integrity of salient objects in two stages. In the decomposition network, we propose a cross multi-branch decoder, which iteratively takes advantage of cross-task aggregation and cross-layer aggregation to integrate multi-level multi-task features and predict saliency, edge, and skeleton maps simultaneously. In the completion network, edge and skeleton maps are further utilized to fill flaws and suppress noises in saliency maps via hierarchical structure-aware feature learning and multi-scale feature completion. Through jointly learning with edge and skeleton information for localizing boundaries and interiors of salient objects respectively, the proposed network generates precise saliency maps with uniformly and completely segmented salient objects. Experiments conducted on five benchmark datasets demonstrate that the proposed model outperforms existing networks. Furthermore, we extend the proposed model to the task of RGB-D salient object detection, and it also achieves state-of-the-art performance. The code is available at https://github.com/wuzhe71/DCN.
Zhe Wu 0006, Li Su 0003, Qingming Huang
IEEE Trans. Image Process.1
2020 Label Decoupling Framework for Salient Object Detection
abstract
To get more accurate saliency maps, recent methods mainly focus on aggregating multi-level features from fully convolutional network (FCN) and introducing edge information as auxiliary supervision. Though remarkable progress has been achieved, we observe that the closer the pixel is to the edge, the more difficult it is to be predicted, because edge pixels have a very imbalance distribution. To address this problem, we propose a label decoupling framework (LDF) which consists of a label decoupling (LD) procedure and a feature interaction network (FIN). LD explicitly decomposes the original saliency map into body map and detail map, where body map concentrates on center areas of objects and detail map focuses on regions around edges. Detail map works better because it involves much more pixels than traditional edge supervision. Different from saliency map, body map discards edge pixels and only pays attention to center areas. This successfully avoids the distraction from edge pixels during training. Therefore, we employ two branches in FIN to deal with body map and detail map respectively. Feature interaction (FI) is designed to fuse the two complementary branches to predict the saliency map, which is then used to refine the two branches again. This iterative refinement is helpful for learning better representations and more precise saliency maps. Comprehensive experiments on six benchmark datasets demonstrate that LDF outperforms state-of-the-art approaches on different evaluation metrics.
Jun Wei 0006, Shuhui Wang, Zhe Wu 0006, Chi Su, Qingming Huang, Qi Tian 0001
CVPR3
2020 Reverse Perspective Network for Perspective-Aware Object Counting
abstract
One of the critical challenges of object counting is the dramatic scale variations, which is introduced by arbitrary perspectives. We propose a reverse perspective network to solve the scale variations of input images, instead of generating perspective maps to smooth final outputs. The reverse perspective network explicitly evaluates the perspective distortions, and efficiently corrects the distortions by uniformly warping the input images. Then the proposed network delivers images with similar instance scales to the regressor. Thus the regression network doesn't need multi-scale receptive fields to match the various scales. Besides, to further solve the scale problem of more congested areas, we enhance the corresponding regions of ground-truth with the evaluation errors. Then we force the regressor to learn from the augmented ground-truth via an adversarial process. Furthermore, to verify the proposed model, we collected a vehicle counting dataset based on Unmanned Aerial Vehicles (UAVs). The proposed dataset has fierce scale variations. Extensive experimental results on four benchmark datasets show the improvements of our method against the state-of-the-arts.
Guorong Li, Zhe Wu 0006, Li Su 0003, Qingming Huang, Nicu Sebe
CVPR3
2020 Weakly-Supervised Crowd Counting Learns from Sorting Rather Than Locations
Guorong Li, Zhe Wu 0006, Li Su 0003, Qingming Huang, Nicu Sebe
ECCV (8)3
2019 Cascaded Partial Decoder for Fast and Accurate Salient Object Detection
abstract
Existing state-of-the-art salient object detection networks rely on aggregating multi-level features of pre-trained convolutional neural networks (CNNs). However, compared to high-level features, low-level features contribute less to performance. Meanwhile, they raise more computational cost because of their larger spatial resolutions. In this paper, we propose a novel Cascaded Partial Decoder (CPD) framework for fast and accurate salient object detection. On the one hand, the framework constructs partial decoder which discards larger resolution features of shallow layers for acceleration. On the other hand, we observe that integrating features of deep layers will obtain relatively precise saliency map. Therefore we directly utilize generated saliency map to recurrently optimize features of deep layers. This strategy efficiently suppresses distractors in the features and significantly improves their representation ability. Experiments conducted on five benchmark datasets exhibit that the proposed model not only achieves state-of-the-art but also runs much faster than existing models. Besides, we apply the proposed framework to optimize existing multi-level feature aggregation models and significantly improve their efficiency and accuracy.
Zhe Wu 0006, Li Su 0003, Qingming Huang
CVPR1
2019 Stacked Cross Refinement Network for Edge-Aware Salient Object Detection
abstract
Salient object detection is a fundamental computer vision task. The majority of existing algorithms focus on aggregating multi-level features of pre-trained convolutional neural networks. Moreover, some researchers attempt to utilize edge information for auxiliary training. However, existing edge-aware models design unidirectional frameworks which only use edge features to improve the segmentation features. Motivated by the logical interrelations between binary segmentation and edge maps, we propose a novel Stacked Cross Refinement Network (SCRN) for salient object detection in this paper. Our framework aims to simultaneously refine multi-level features of salient object detection and edge detection by stacking Cross Refinement Unit (CRU). According to the logical interrelations, the CRU designs two direction-specific integration operations, and bidirectionally passes messages between the two tasks. Incorporating the refined edge-preserving features with the typical U-Net, our model detects salient objects accurately. Extensive experiments conducted on six benchmark datasets demonstrate that our method outperforms existing state-of-the-art algorithms in both accuracy and efficiency. Besides, the attribute-based performance on the SOC dataset show that the proposed model ranks first in the majority of challenging scenes. Code can be found at https://github.com/wuzhe71/SCAN.
Zhe Wu 0006, Li Su 0003, Qingming Huang
ICCV1
2019 Learning Coupled Convolutional Networks Fusion for Video Saliency Prediction
abstract
Visual saliency provides important information for understanding scenes in many computer vision tasks. The existing video saliency algorithms mainly focus on predicting spatial and temporal saliency maps. However, these maps are simply fused without considering the complex dynamic scenes in videos. To overcome this drawback, we propose a deep convolutional fusion framework for video saliency prediction. The proposed model, which is based on coupled fully convolutional networks (FCNs), effectively encodes the spatiotemporal information by integrating spatial and temporal features. We demonstrate that this information is helpful for accurately fusing the spatial and temporal saliency maps according to changes in video scenes. In particular, we gradually design three different deep fusion architectures to investigate how to better utilize the spatiotemporal information. Moreover, we propose a reasonable sampling strategy for selecting suitable training sets for the coupled FCNs. Through extensive experiments, we demonstrate that our model outperforms the state-of-the-art algorithms on four public video saliency data sets.
Zhe Wu 0006, Li Su 0003, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.1
2017 Saliency detection with two-level fully convolutional networks
abstract
This paper proposes a deep architecture for saliency detection by fusing pixel-level and superpixel-level predictions. Different from the previous methods that either make dense pixellevel prediction with complex networks or region-level prediction for each region with fully-connected layers, this paper investigates an elegant route to make two-level predictions based on a same simple fully convolutional network via seamless transformation. In the transformation module, we integrate the low level features to model the similarities between pixels and superpixels as well as superpixels and superpixels. The pixel-level saliency map detects and highlights the salient object well and the superpixel-level saliency map preserves sharp boundary in a complementary way. A shallow fusion net is applied to learn to fuse the two saliency maps, followed by a CRF post-refinement module. Experiments on four benchmark data sets demonstrate that our method performs favorably against the state-of-art methods.
Li Su 0003, Qingming Huang, Zhe Wu 0006
ICME4
2016 Webpage saliency prediction with multi-features fusion
abstract
We proposed a novel model to predict human's visual attention when free-viewing webpages. Compared with natural images, webpages are usually full of salient regions such as logos, text, and faces, while few of them attract human's attention in a short sight. Moreover, webpages perform distinct viewing patterns which are quite different from the natural images. In this paper, we introduced multi-features according to our observation on webpages characters and related eye-tracking data. Further, in order to achieve a flexible adaptation to various types of webpages, we employed a machine-learning framework based on our proposed features. Experimental results demonstrate that our model outperforms other state-of-the-art methods in webpage saliency prediction.
Li Su 0003, Bo Wu 0016, Junbiao Pang, Zhe Wu 0006, Qingming Huang
ICIP6
2016 Video saliency prediction with optimized optical flow and gravity center bias
abstract
Dynamic videos are viewed fundamentally different from static images. Besides spatial features, motion feature also plays an important role as a temporal factor. Most existing video saliency models usually employ optical flow to represent the motion feature. However, optical flow often suffers from the discontinuity problem. And we also notice that human fixations in one single video frame are much sparser than that in an identical still picture. However, many spatial saliency models take each video frame as static image independently. In this paper, we predict the dynamic visual saliency by fusing spatial and temporal features. In order to construct the temporal relationships among a set of successive frames, we introduce a smoothness operator in optical flow field to obtain more accurate motion feature. Then, considering the sparse property of video saliency, we adapt the weights of the regions surrounding to the saliency gravity center in the final maps. The experiments show that our model is more consistent with humans eye-tracking benchmarks than the state-of-the-art models.
Zhe Wu 0006, Li Su 0003, Qingming Huang, Bo Wu 0016, Guorong Li
ICME1