VLDB 2026 Research / reviewers in the wild / expert
Tian Wang 0002
dblp:25/3246-2
· DBLP profile ↗
52ranked-venue papers
14as first author
40since 2021 · last 2026
0000-0001-8427-4495ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 3 first-author · 19 since 2021Artificial intelligence and machine learning · 16 · 2 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 6 first-author · 9 since 2021Security and privacy · 3 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Robust 3-D Gaussian SLAM for Humanoid Robots With Visual Enhancement in IoT-Enabled Dynamic Crowd ScenesabstractIn recent years, humanoid robots have gradually emerged as crucial “smart terminals” within the IoT, significantly expanding the application scenarios of IoT. This paper proposes a visual enhancement based robust 3D Gaussian SLAM (3DGS-SLAM) to improve the environmental perception ability of humanoid robots. First, the dynamic mask provided by YOLOv11 is optimized using composite morphological operations to obtain high-confidence static features, which are then employed for accurate estimation of the robot’s pose. Subsequently, a keyframe selection strategy combining scaled interval Kalman filtering (SIKF) with the motion model of a floating base humanoid robot is designed. It quantifies the uncertainty of pose by dynamically dividing the motion interval of the robot, and introduces a scaling factor to adaptively adjust the interval width, thereby filtering out high-quality keyframes to ensure stable pose tracking and efficient mapping. Further, a mapping strategy that combines Gaussian-Laplacian pyramid and Gaussian ellipsoid adaptive density control is developed by combining 3D Gaussian splashing. This method can not only identify Gaussian ellipsoids requiring segmentation by preserving high-frequency details in keyframes, but also control their splitting direction through gradient modulation, effectively reducing redundancy among Gaussian ellipsoids while enhancing mapping quality. The experimental results show that our method improves the average absolute trajectory error(ATE) on the BONN and TUM datasets by an average of 90.8% and 93.1% compared to ORB-SLAM3, significantly improving the localization accuracy and mapping consistency of humanoid robots in dynamic crowded scenes. Tian Wang 0002, Xuanzhen Chen, Jingwen Luo |
IEEE Internet Things J. | 2 |
| 2026 | CVC-Net: A Cross-View Consistency Network for Noise-Generalization Fault DiagnosisabstractDeep learning applications in fault diagnosis face two critical challenges. First, a significant distribution gap between source domain training data and target domain samples with unknown noise patterns. Second, labeled fault data remain scarce in practice. These issues hinder the practical deployment. This paper presents a Cross-View Consistency Network (CVC-Net) to tackle these problems through noise-generalization capabilities. The method learns robust features from limited source domain data. It maintains diagnostic accuracy with unknown noise data, without prior knowledge of target noise characteristics. CVC-Net processes temporal waveforms and Gramian Angular Field representations through specialized encoders, exploiting their asymmetric noise sensitivities. A cross-view consistency mechanism extracts fault patterns across modalities. The method integrates fault-aware prototype learning for enhanced discrimination with limited labels and employs adaptive fusion that weights view contributions based on cross-view prediction. Experimental validation shows that CVC-Net is effective in challenging scenarios. When tested on target domain with unknown noise types, CVC-Net maintains reliable performance, effectively handling noise patterns not present during source domain training. Under limited-label conditions, it outperforms existing methods in diagnostic performance. Tian Wang 0002, Hetian Feng, Jintong Wang, Jinghe Zhao, Hichem Snoussi |
IEEE Signal Process. Lett. | 2 |
| 2026 | MCFM: Multimodal Competitive Fusion Mechanism for Sentiment AnalysisabstractWith the popularity of social media, users are able to express their opinions in multiple forms, such as text, audio, and video. Traditional unimodal sentiment analysis methods can no longer meet the processing requirements of such multisource heterogeneous data, which makes multimodal sentiment analysis a research hotspot. However, most existing methods rely on simple feature splicing or weighted fusion, neglecting the differences in the reliability of different modalities and failing to fully explore the intermodality consistency and difference information. In this article, we propose a multimodal competitive fusion mechanism and construct multimodal competitive fusion model (MCFM). The model first dynamically evaluates the reliability of each modality through the competition mechanism and adaptively assigns weights accordingly. Then it decomposed the modality representations into similar and dissimilar features through modality feature decomposition, supplemented by the overlap of orthogonal traffic channel attention constraints, to achieve the collaborative learning of consistency and dissimilarity features. We evaluate the proposed model on several datasets. In our experiments, we used textual modality data from the dataset with audio modality data for the experiments. The experimental results show that MCFM has 2%–3% higher binary accuracy (ACC2) than the baseline model on the sentiment classification task (with 2% higher binary accuracy under the negative/nonnegative metrics and 3% higher binary accuracy under the positive/negative metrics), and that on the regression task, MCFM’s mean absolute error on the test dataset is 3% lower than that of the baseline model. Mali Xing, Zilang Zhai, Muqing Deng, Qianqian Cai, Hongru Ren, Tian Wang 0002 |
IEEE Trans. Comput. Soc. Syst. | 6 |
| 2026 | Taking Astray Domain Back Home for Single-Source Domain Generalizable Text-to-Image Person RetrievalabstractGiven a query sentence, text-to-image person retrieval aims to identify matched pedestrian images from a large gallery. Most of the existing methods are designed for the unified domain setting, which is operated under the assumption that the training and test data are drawn from the same distribution. However, this assumption is difficult to guarantee in real application scenes, as data is often collected from various surveillance scenarios. To this end, in this paper, we introduce the concept of single-source domain generalization into the context of text-to-image person retrieval and propose a novel task called single-source domain generalizable text-to-image person retrieval (SSDG-TIPR). This task is applicable in real-world scenarios but poses significant challenges due to the limitation of accessible training data. Intuitively, a trained model is the most familiar with the domain on which it was trained, that is, the source domain. Therefore, to handle this SSDG-TIPR task, we propose a new method to infinitely close astray features from unseen target domains to the source domain, namely, to take it home (TIME), allowing the model to handle the features in a familiar manner. The proposed TIME method comprises three main modules: the Domain Astray Leading (DAL) module, the Domain Invariant Feature Extract (DIFE) module and the Domain Home Taking (DoT) module. We evaluated TIME on 3 benchmark datasets, namely CUHK-PEDES, ICFG-PEDES and RSTPReid, and demonstrated its superior performance on 10 SSDG-TIPR sub-tasks as well as on 3 conventional TIPR sub-tasks, establishing a new state-of-the-art in both settings. Guan-Nan Dong, Zijie Wang 0003, Aichun Zhu, Yuanfei Dai, Tian Wang 0002, Hichem Snoussi |
IEEE Trans. Image Process. | 7 |
| 2026 | Weakly Supervised Temporal Action Localization With Proposal-Level Action Consistency LearningabstractExisting weakly supervised temporal action localization (WTAL) methods typically follow a decoupled classification-localization pipeline: segment-level classifiers are trained first, and their predictions are then aggregated to score proposals at inference. Under this training-inference discrepancy, proposal scoring at inference relies on an additional aggregation step, which can accumulate errors from noisy segment responses and thus undermine score reliability. Moreover, proposal scores are often directly used as confidence without explicit score-quality modeling or quality-aware evaluation, further contributing to pronounced score-quality misalignment and thus widening the classification-localization gap. To address proposal score-quality misalignment, we propose ACL-Net, a framework for proposal score calibration. At its core is a dual-axis Proposal-level Action Consistency Learning (PACL) paradigm, implemented through two complementary modules: (i) a Semantic Consistency Module (SCM) that refines proposal representations by maintaining fused class centers to enforce compact and robust same-class features; within SCM, a cross-modal consistency-driven Classification Enhancement Module (CEM) denoises the fused class centers to mitigate error accumulation under weak supervision; and (ii) a Process Consistency Module (PCM) that derives geometry-aware reference scores from relative temporal relations among overlapping proposals, guiding the model to assess proposal quality in terms of relative process completeness and improve score-quality alignment. By jointly modeling semantic and process consistency to calibrate proposal scores, ACL-Net markedly improves localization accuracy. On THUMOS14 and ActivityNet1.3, it achieves state-of-the-art performance with uniform and substantial gains across multiple established baselines, while markedly lowering the expected calibration error (ECE). Maodong Li 0006, Zhihao Wang 0002, Tian Wang 0002, Jingxiong Wang, Jian Wang 0018, Bing Li 0010 |
IEEE Trans. Image Process. | 3 |
| 2026 | Enhanced Query Attention Constrained by Bi-Directional Graphs for Human Pose Estimation NetworksabstractIn human pose estimation, formulating keypoint localization as a classification task over discretized coordinate grids has proven effective. Essentially, the 2D features of the keypoints are reduced to 1D coordinate representations. This process leads to the loss of spatial constraints among keypoints and increases the difficulty for the model to capture their structural relationships. To address this issue, we propose an enhanced query attention mechanism constrained by bidirectional graphs. The core idea is to establish the topological constraints on the 1D coordinate representations. First, two fundamental connection directions of the skeleton are defined and encoded as a pair of adjacency matrices to enhance the feature interaction capability of the graph convolutional network (GCN). Second, a GCN-guided multi-scale feature fusion framework is designed to effectively combine multi-scale visual features with structural priors, thereby enhancing the representation of keypoint spatial distributions. Finally, a dual-gate module is incorporated into a GCN-guided attention unit to construct a structured query matrix constrained by the bidirectional skeleton graphs, which helps filter out spurious joint interactions and emphasize plausible ones. Extensive experiments on Tai Chi Chuan-Pose, Animal-Pose, AP-10K, MPII, COCO, and COCO-WholeBody datasets demonstrate that the proposed method outperforms existing methods in terms of both accuracy and robustness, particularly in balancing precise local keypoint localization with global pose consistency. Yi Yang 0043, Wei Qian 0002, Tian Wang 0002, Yunlong Lv |
IEEE Trans. Image Process. | 4 |
| 2026 | STNMamba: Mamba-Based Spatial-Temporal Normality Learning for Video Anomaly DetectionabstractVideo anomaly detection (VAD) has been extensively researched due to its potential for intelligent video systems. However, most existing methods based on CNNs and transformers still suffer from substantial computational burdens and have room for improvement in learning spatial-temporal normality. Recently, Mamba has shown great potential for modeling long-range dependencies with linear complexity, providing an effective solution to the above dilemma. To this end, we propose a lightweight and effective Mamba-based network named STNMamba, which incorporates carefully designed Mamba modules to enhance the learning of spatial-temporal normality. Firstly, we develop a dual-encoder architecture, where the spatial encoder equipped with Multi-Scale Vision Space State Blocks (MS-VSSB) extracts multi-scale appearance features, and the temporal encoder employs Channel-Aware Vision Space State Blocks (CA-VSSB) to capture significant motion patterns. Secondly, a Spatial-Temporal Interaction Module (STIM) is introduced to integrate spatial and temporal information across multiple levels, enabling effective modeling of intrinsic spatial-temporal consistency. Within this module, the Spatial-Temporal Fusion Block (STFB) is proposed to fuse the spatial and temporal features into a unified feature space, and the memory bank is utilized to store spatial-temporal prototypes of normal patterns, restricting the model's ability to represent anomalies. Extensive experiments on three benchmark datasets demonstrate that our STNMamba achieves competitive performance with fewer parameters and lower computational costs than existing methods. Zhangxun Li, Mengyang Zhao 0002, Yang Liu 0246, Jiamu Sheng, Xinhua Zeng, Tian Wang 0002, Kewei Wu, Yu-Gang Jiang 0001 |
IEEE Trans. Multim. | 7 |
| 2025 | Balance Orthogonal Projection for Prompt in Continual Learning
Junjian Ren, Tian Wang 0002, Aichun Zhu, Chuanyun Wang, Nadia Bali, Hichem Snoussi |
PRCV (2) | 2 |
| 2025 | Implicit Diffusion Models for Continuous Super-Resolution
Xuhui Liu, Sicheng Gao, Bohan Zeng, Tian Wang 0002, Jianzhuang Liu, Baochang Zhang 0001 |
Int. J. Comput. Vis. | 5 |
| 2025 | Hydraulic-Supports Alignment by TD3 with Segmented Experience PoolabstractAbstract Hydraulic-supports alignment is to keep the coal mining face in line and is heavily influenced by the various geological states. The experiences produced by the moving process are unbalanced, which leads to the agent not learning important knowledge from the rare samples. This paper is the first to introduce the reinforcement learning to the hydraulic-supports alignment, and establish the Markov optimal decision model by TD3 algorithm. Aiming at the imbalance issue of the experience, this paper proposes a segmented experience pool and three sampling replay mechanisms according to the characteristics of the moving process with various geological states. Experimental results show that the improved TD3, utilizing a segmented experience pool with three different replay mechanisms, could effectively identify the optimal moving policy and achieve significant convergence in cases involving both normal movement and insufficient movement of hydraulic-supports. In contrast, the TD3 performs inadequately and struggles to find the optimal policy. Yi Yang 0043, Yapeng Dai, Tian Wang 0002, Wei Qian 0002 |
Neural Process. Lett. | 3 |
| 2025 | Grasping With Occlusion-Aware Ally Method in Complex ScenesabstractRobotic arm target grasping by vision support is a commonly used method in grasping tasks and is usually used for multi-target complex scenes. Where vision support is generally used to identify the targets and to get their positions, categories and sizes. Most robotic arm grasping tasks using target recognition methods as visual inspection ignore the relationship between target objects such as the occlusion problem between objects. This limits the targets to be grasped and makes the crawling task inefficient. We propose Grasping with Occlusion-Aware aLly (GOAL) method based on binocular stereo-vision. Firstly, occlusion relationships in the view are directly inferred and targets are segmented as well as localized. Subsequently, multi-target grasping pose estimation is performed to obtain effective grasping positions. Ultimately, validation is conducted on a high-resolution dataset using the EPSON robotic arm. Note to Practitioners—This research significantly advances the field by addressing occlusion challenges in robotic grasping, offering effective methods, a valuable dataset, and practical insights. The proposed Grasping with Occlusion-Aware aLly (GOAL) method was validated on a high-resolution dataset using the EPSON robotic arm, showcasing its applicability and efficiency in real-world scenarios. This work provides valuable contributions to practitioners in the field of robotic manipulation and grasping tasks. Lulu Li 0012, Abel Cherouat, Hichem Snoussi, Tian Wang 0002 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2025 | Adaptive Sliding Mode Synchronous Control for Complex Networks With Amplify-and-Forward RelaysabstractThis paper aims to the remote synchronous control for complex networks with amplify-and-forward (AF) relays. Firstly, the AF relay dynamic with random noise is integrated into the control signal transmission. Secondly, the complex networks incorporating an AF relay is converted into the sliding control mode (SMC), which robustly addresses unknown disturbances among the nodes to achieve synchronous states. Thirdly, an adaptive gain mechanism based on the equivalent value of sign function of SMC is proposed to estimate the disturbance amplitude, hence mitigating the chattering of SMC. Finally, the illustrative simulation demonstrated the effectiveness of proposed methodology. Yi Yang 0043, Wei Qian 0002, Tian Wang 0002, Keping Wang |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | Understanding the Dimensional Need of Noncontrastive LearningabstractNoncontrastive self-supervised learning methods offer an effective alternative to contrastive approaches by avoiding the need for negative samples to avoid representation collapse. Noncontrastive learning methods explicitly or implicitly optimize the representation space, yet they often require large representation dimensions, leading to dimensional inefficiency. To provide negative samples, contrastive learning methods often require large batch sizes, thus regarded as sample inefficient, while noncontrastive learning methods require large representation dimensions, thus regarded as dimension inefficient. Although we have some understanding of the noncontrastive learning method, theoretical analysis of such phenomenon still remains largely unexplored. We present a theoretical analysis of the dimensional need for noncontrastive learning. We investigate the transfer between upstream representation learning and downstream tasks' performance, demonstrating how noncontrastive methods implicitly increase interclass distances within the representation space and how the distance affects the model performance of evaluation performance. We prove that the performance of noncontrastive methods is affected by the output dimension and the number of latent classes, and illustrate why performance degrades significantly when the output dimension is substantially smaller than the number of latent classes. We demonstrate our findings through experiments on image classification experiments, and enrich the verification in audio, graph and text modalities. We also perform empirical evaluation for image models on extensive detection and segmentation tasks beyond classification that show satisfactory correspondence to our theorem. Zhexiao Cao, Lei Huang 0015, Tian Wang 0002, Yinquan Wang, Jingang Shi, Aichun Zhu, Tianyun Shi, Hichem Snoussi |
IEEE Trans. Cybern. | 3 |
| 2025 | Onet: Twin U-Net Architecture for Unsupervised Binary Semantic Segmentation in Radar and Remote Sensing ImagesabstractSegmenting objects from cluttered backgrounds in single-channel images, such as marine radar echoes, medical images, and remote sensing images, poses significant challenges due to limited texture, color information, and diverse target types. This paper proposes a novel solution: the Onet, an O-shaped assembly of twin U-Net deep neural networks, designed for unsupervised binary semantic segmentation. The Onet, trained with an intensity-complementary image pair and without the need for annotated labels, maximizes the Jensen-Shannon divergence (JSD) between the densely localized features and the class probability maps. By leveraging the symmetry of U-Net, Onet subtly strengthens the dependence between dense local features, global features, and class probability maps during the training process. The design of the complementary input pair aligns with the theoretical requirement that optimizing JSD needs the class probability of negative samples to accurately estimate the marginal distribution. Compared to the current leading unsupervised segmentation methods, the Onet demonstrates superior performance in target segmentation in marine radar frames and cloud segmentation in remote sensing images. Notably, we found that Onet's foreground prediction significantly enhances the signal-to-noise ratio (SNR) of targets amidst marine radar clutter. Onet's source code is publicly accessible at https://github.com/joeyee/Onet. Yi Zhou 0011, Hang Su 0006, Tian Wang 0002, Qing Hu 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | Improving Text-Based Person Retrieval by Excavating All-Round Information Beyond ColorabstractText-based person retrieval is the process of searching a massive visual resource library for images of a particular pedestrian, based on a textual query. Existing approaches often suffer from a problem of color (CLR) over-reliance, which can result in a suboptimal person retrieval performance by distracting the model from other important visual cues such as texture and structure information. To handle this problem, we propose a novel framework to Excavate All-round Information Beyond Color for the task of text-based person retrieval, which is therefore termed EAIBC. The EAIBC architecture includes four branches, namely an RGB branch, a grayscale (GRS) branch, a high-frequency (HFQ) branch, and a CLR branch. Furthermore, we introduce a mutual learning (ML) mechanism to facilitate communication and learning among the branches, enabling them to take full advantage of all-round information in an effective and balanced manner. We evaluate the proposed method on three benchmark datasets, including CUHK-PEDES, ICFG-PEDES, and RSTPReid. The experimental results demonstrate that EAIBC significantly outperforms existing methods and achieves state-of-the-art (SOTA) performance in supervised, weakly supervised, and cross-domain settings. Aichun Zhu, Zijie Wang 0003, Jingyi Xue, Xili Wan, Jing Jin 0002, Tian Wang 0002, Hichem Snoussi |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Target-Specific Domain Adaptation via Geometry-Correlation Prediction for Point Cloud
Junqiao Li, Leyan Zhu, Tian Wang 0002, Jingang Shi, Hichem Snoussi |
PRCV (4) | 3 |
| 2024 | A Multihead Attention Self-Supervised Representation Model for Industrial Sensors Anomaly DetectionabstractIndustrial sensors capture critical information for intelligent manufacturing maintenance. To promote equipment upgrading and manufacturing processes, intelligent decisions, and information learning play an important role. Although deep learning methods historically obtain excellent results, there is always a tradeoff between fine-tuning existing networks or designing models from scratch for sensor data processing. In this article, we propose the multihead attention self-supervised (MAS) representation model, which is a self-supervised learning-based sensor feature extraction network. To the best of our knowledge, this is the first time a self-supervised contrastive learning method using positive samples that represent multidimensional industry sensor data is being used for anomaly detection. We review alternative data augmentation methods proposed for better-representing sensor sequence data. We use this insight to design a new structure that adapts to the temporal characteristics of the application. We apply our method to a real-world water circulation system that uses a variety of industrial sensors. The effectiveness of the proposed MAS methods is demonstrated. Yiqun Qiao, Jinhu Lü 0001, Tian Wang 0002, Baochang Zhang 0001, Hichem Snoussi |
IEEE Trans. Ind. Informatics | 3 |
| 2023 | Few-Shot Learning with Visual Distribution Calibration and Cross-Modal Distribution AlignmentabstractPre-trained vision-language models have inspired much research on few-shot learning. However, with only a few training images, there exist two crucial problems: (1) the visual feature distributions are easily distracted by class-irrelevant information in images, and (2) the alignment between the visual and language feature distributions is difficult. To deal with the distraction problem, we propose a Selective Attack module, which consists of trainable adapters that generate spatial attention maps of images to guide the attacks on class-irrelevant image areas. By messing up these areas, the critical features are captured and the visual distributions of image features are calibrated. To better align the visual and language feature distributions that describe the same object class, we propose a cross-modal distribution alignment module, in which we introduce a vision-language prototype for each class to align the distributions, and adopt the Earth Mover's Distance (EMD) to optimize the prototypes. For efficient computation, the upper bound of EMD is derived. In addition, we propose an augmentation strategy to increase the diversity of the images and the text prompts, which can reduce overfitting to the few-shot training images. Extensive experiments on 11 datasets demonstrate that our method consistently outperforms prior arts in few-shot learning. The implementation code will be available at https://gitee.com/mindspore/models/tree/master/research/cv/SADA. Runqi Wang, Xiaoyue Duan, Jianzhuang Liu, Yuning Lu, Tian Wang 0002, Songcen Xu, Baochang Zhang 0001 |
CVPR | 6 |
| 2023 | Learning Attention from Attention: Efficient Self-Refinement Transformer for Face Super-ResolutionabstractRecently, Transformer-based architecture has been introduced into face super-resolution task due to its advantage in capturing long-range dependencies. However, these approaches tend to integrate global information in a large searching region, which neglect to focus on the most relevant information and induce blurry effect by the irrelevant textures. Some improved methods simply constrain self-attention in a local window to suppress the useless information. But it also limits the capability of recovering high-frequency details when flat areas dominate the local searching window. To improve the above issues, we propose a novel self-refinement mechanism which could adaptively achieve texture-aware reconstruction in a coarse-to-fine procedure. Generally, the primary self-attention is first conducted to reconstruct the coarse-grained textures and detect the fine-grained regions required further compensation. Then, region selection attention is performed to refine the textures on these key regions. Since self-attention considers the channel information on tokens equally, we employ a dual-branch feature integration module to privilege the important channels in feature extraction. Furthermore, we design the wavelet fusion module which integrate shallow-layer structure and deep-layer detailed feature to recover realistic face images in frequency domain. Extensive experiments demonstrate the effectiveness on a variety of datasets. Guanxin Li, Jingang Shi, Yuan Zong, Fei Wang 0037, Tian Wang 0002, Yihong Gong |
IJCAI | 5 |
| 2023 | Memory-Augmented Spatial-Temporal Consistency Network for Video Anomaly Detection
Zhangxun Li, Mengyang Zhao 0002, Xinhua Zeng, Tian Wang 0002, Chengxin Pang |
PRCV (6) | 4 |
| 2023 | DCP-NAS: Discrepant Child-Parent Neural Architecture Search for 1-bit CNNs
Yanjing Li, Sheng Xu 0007, Xianbin Cao 0001, Lian Zhuo, Baochang Zhang 0001, Tian Wang 0002, Guodong Guo |
Int. J. Comput. Vis. | 6 |
| 2023 | Tri-HGNN: Learning triple policies fused hierarchical graph neural networks for pedestrian trajectory prediction
Yanghong Liu, Tian Wang 0002 |
Pattern Recognit. | 5 |
| 2023 | Synchronous Spatiotemporal Graph Transformer: A New Framework for Traffic Data PredictionabstractModeling the spatiotemporal relationship (STR) of traffic data is important yet challenging for existing graph networks. These methods usually capture features separately in temporal and spatial dimensions or represent the spatiotemporal data by adopting multiple local spatial-temporal graphs. The first kind of method mentioned above is difficult to capture potential temporal-spatial relationships, while the other is limited for long-term feature extraction due to its local receptive field. To handle these issues, the Synchronous Spatio-Temporal grAph Transformer (S2TAT) network is proposed for efficiently modeling the traffic data. The contributions of our method include the following: 1) the nonlocal STR can be synchronously modeled by our integrated attention mechanism and graph convolution in the proposed S2TAT block; 2) the timewise graph convolution and multihead mechanism designed can handle the heterogeneity of data; and 3) we introduce a novel attention-based strategy in the output module, being able to capture more valuable historical information to overcome the shortcoming of conventional average aggregation. Extensive experiments are conducted on PeMS datasets that demonstrate the efficacy of the S2TAT by achieving a top-one accuracy but less computational cost by comparing with the state of the art. Tian Wang 0002, Jinhu Lü 0001, Aichun Zhu, Hichem Snoussi, Baochang Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Delving into the Estimation Shift of Batch Normalization in a NetworkabstractBatch normalization (BN) is a milestone technique in deep learning. It normalizes the activation using mini-batch statistics during training but the estimated population statistics during inference. This paper focuses on investigating the estimation of population statistics. We define the estimation shift magnitude of BN to quantitatively measure the difference between its estimated population statistics and expected ones. Our primary observation is that the estimation shift can be accumulated due to the stack of BN in a network, which has detriment effects for the test performance. We further find a batch-free normalization (BFN) can block such an accumulation of estimation shift. These observations motivate our design of XBNBlock that replace one BN with BFN in the bottleneck block of residual-style networks. Experiments on the ImageNet and COCO benchmarks show that XBNBlock consistently improves the performance of different architectures, including ResNet and ResNeXt, by a significant margin and seems to be more robust to distribution shift. Lei Huang 0015, Yi Zhou 0007, Tian Wang 0002, Jie Luo 0004, Xianglong Liu 0001 |
CVPR | 3 |
| 2022 | Bi-level Doubly Variational Learning for Energy-based Latent Variable ModelsabstractEnergy-based latent variable models (EBLVMs) are more expressive than conventional energy-based models. However, its potential on visual tasks are limited by its training process based on maximum likelihood estimate that requires sampling from two intractable distributions. In this paper, we propose Bi-level doubly variational learning (BiDVL), which is based on a new bi-level optimization framework and two tractable variational distributions to facilitate learning EBLVMs. Particularly, we lead a decoupled EBLVM consisting of a marginal energy-based distribution and a structural posterior to handle the difficulties when learning deep EBLVMs on images. By choosing a symmetric KL divergence in the lower level of our framework, a compact BiDVL for visual tasks can be obtained. Our model achieves impressive image generation performance over related works. It also demonstrates the significant capacity of testing image reconstruction and out-of-distribution detection. Ge Kan, Jinhu Lü 0001, Tian Wang 0002, Baochang Zhang 0001, Aichun Zhu, Lei Huang 0015, Guodong Guo, Hichem Snoussi |
CVPR | 3 |
| 2022 | Look Before You Leap: Improving Text-based Person Retrieval by Learning A Consistent Cross-modal Common ManifoldabstractThe core problem of text-based person retrieval is how to bridge the heterogeneous gap between multi-modal data. Many previous approaches contrive to learning a latent common manifold mapping paradigm following a cross-modal distribution consensus prediction (CDCP) manner. When mapping features from distribution of one certain modality into the common manifold, feature distribution of the opposite modality is completely invisible. That is to say, how to achieve a cross-modal distribution consensus so as to embed and align the multi-modal features in a constructed cross-modal common manifold all depends on the experience of the model itself, instead of the actual situation. With such methods, it is inevitable that the multi-modal data can not be well aligned in the common manifold, which finally leads to a sub-optimal retrieval performance. To overcome this CDCP dilemma, we propose a novel algorithm termed LBUL to learn a Consistent Cross-modal Common Manifold (C3 M) for text-based person retrieval. The core idea of our method, just as a Chinese saying goes, is to 'san si er hou xing', namely, to Look Before yoU Leap (LBUL). The common manifold mapping mechanism of LBUL contains a looking step and a leaping step. Compared to CDCP-based methods, LBUL considers distribution characteristics of both the visual and textual modalities before embedding data from one certain modality into C3 M to achieve a more solid cross-modal distribution consensus, and hence achieve a superior retrieval accuracy. We evaluate our proposed method on two text-based person retrieval datasets CUHK-PEDES and RSTPReid. Experimental results demonstrate that the proposed LBUL outperforms previous methods and achieves the state-of-the-art performance. Zijie Wang 0003, Aichun Zhu, Jingyi Xue, Xili Wan, Tian Wang 0002, Yifeng Li 0002 |
ACM Multimedia | 6 |
| 2022 | CAIBC: Capturing All-round Information Beyond Color for Text-based Person RetrievalabstractGiven a natural language description, text-based person retrieval aims to identify images of a target person from a large-scale person image database. Existing methods generally face a color over-reliance problem, which means that the models rely heavily on color information when matching cross-modal data. Indeed, color information is an important decision-making accordance for retrieval, but the over-reliance on color would distract the model from other key clues (e.g. texture information, structural information, etc.), and thereby lead to a sub-optimal retrieval performance. To solve this problem, in this paper, we propose to Capture All-round Information Beyond Color (CAIBC) via a jointly optimized multi-branch architecture for text-based person retrieval. CAIBC contains three branches including an RGB branch, a grayscale (GRS) branch and a color (CLR) branch. Besides, with the aim of making full use of all-round information in a balanced and effective way, a mutual learning mechanism is employed to enable the three branches which attend to varied aspects of information to communicate with and learn from each other. Extensive experimental analysis is carried out to evaluate our proposed CAIBC method on the CUHK-PEDES and RSTPReid datasets in both supervised and weakly supervised text-based person retrieval settings, which demonstrates that CAIBC significantly outperforms existing methods and achieves the state-of-the-art performance on all the three tasks. Zijie Wang 0003, Aichun Zhu, Jingyi Xue, Xili Wan, Tian Wang 0002, Yifeng Li 0002 |
ACM Multimedia | 6 |
| 2022 | Accelerating temporal action proposal generation via high performance computing
Tian Wang 0002, Shiye Lei, Youyou Jiang, Chang Choi, Hichem Snoussi, Guangcun Shan |
Frontiers Comput. Sci. | 1 |
| 2022 | ResLNet: deep residual LSTM network with longer input for action recognition
Tian Wang 0002, Huai-Ning Wu, Ce Li 0001, Hichem Snoussi, Yang Wu 0001 |
Frontiers Comput. Sci. | 1 |
| 2022 | Adaptive Optimization Method in Digital Twin Conveyor Systems via Range-Inspection ControlabstractThe automated conveyor system, as the core component in the modern manufacturing world, has gained lots of attention from researchers. To optimize the operation of the conveyor system, range-inspection control (RIC) has been considered an efficient strategy to bring this conventional technology to an intelligent level. Various algorithms have been put into use to achieve optimal control. However, the current methodologies are only focusing on control optimization, not scaled into the smart manufacturing framework. The schema of alignment and corporation between the physical and virtual spaces for the system remains an important problem. Therefore, the work in this article aims for an effective framework of implementation between the physical and virtual stations in an automated conveyor system. Since increasingly more application scenarios rely on the digital twin (DT) technology to realize the integration of physical and virtual systems, we proposed the DT automated conveyor system (DT-ACS) that constructs the road map to implement the RIC-based conveyor system under the background of a smart factory. Besides, profit-sharing-based deep Q-networks (PDQNs) have been proposed to cope with the RIC optimization problem. The robustness and efficiency of the proposed PDQN were evaluated via sets of experiments. The discussion and conclusion are presented at last accordingly.Note to Practitioners—This article aims to propose a strategy of control optimization for conveyor-based manufacturing systems under the digital twin (DT) framework. The conveyor system can be flexible to control the running flows to avoid overloading workstations. Due to the complex environment in the production line, the range that is able to be inspected and the capacity of the reserve area can be considerably diverse among the workstations. To maximally evaluate our framework, we set a comparatively complex environment for the experiments. Nevertheless, to obtain practically ideal performance under other circumstances, the parameters should be precisely tested and fine-tuned with simulation in advance. Tian Wang 0002, Jiaxiang Cheng, Yi Yang 0043, Christian Esposito 0001, Hichem Snoussi, Fei Tao 0001 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2022 | CACrowdGAN: Cascaded Attentional Generative Adversarial Network for Crowd CountingabstractCrowd counting is a valuable technology for extremely dense scenes in the transportation. Existing methods generally have higher-order inconsistencies between ground truth density maps and generated density maps. To address this issue, we incorporate an attentional discriminator to take charge of checking the density map between the generator and the ground truth. Thus, a Cascaded Attentional Generative Adversarial Network (CACrowdGAN) is proposed that enables the attentional-driven discriminator to distinguish implausible density maps and simultaneously to guide the generator to deliver fine-grained high quality density maps. The proposed CACrowdGAN consists of two components: an attentional generator and a cascaded attentional discriminator. The attentional generator has an attention module and a density module. The attention module is developed for the generator to focus on the crowd regions of the input images, while the density module is used to provide the attentional input of the discriminator. In addition, a cascaded attentional discriminator is proposed to synthesize attentional-driven fine-grained details at different crowd regions of the input image and compute a per-pixel fine-grained loss for training generator. The proposed CACrowdGAN achieves the state-of-the-art performance on five popular crowd counting datasets (ShanghaiTech, WorldEXPO’10, UCSD, UCF_CC_50 and UCF_QNRF), which demonstrates the effectiveness and robustness of the proposed approach in the complex scenes. Aichun Zhu, Yaoying Huang, Tian Wang 0002, Jing Jin 0002, Fangqiang Hu, Gang Hua 0002, Hichem Snoussi |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | IDARTS: Interactive Differentiable Architecture SearchabstractDifferentiable Architecture Search (DARTS) improves the efficiency of architecture search by learning the architecture and network parameters end-to-end. However, the intrinsic relationship between the architecture’s parameters is neglected, leading to a sub-optimal optimization process. The reason lies in the fact that the gradient descent method used in DARTS ignores the coupling relationship of the parameters and therefore degrades the optimization. In this paper, we address this issue by formulating DARTS as a bi-linear optimization problem and introducing an Interactive Differentiable Architecture Search (IDARTS). We first develop a backtracking backpropagation process, which can decouple the relationships of different kinds of parameters and train them in the same framework. The backtracking method coordinates the training of different parameters that fully explore their interaction and optimize training. We present experiments on the CIFAR10 and ImageNet datasets that demonstrate the efficacy of the IDARTS approach by achieving a top-1 accuracy of 76.52% on ImageNet without additional search cost vs. 75.8% with the state-of-the-art PC-DARTS. Runqi Wang, Baochang Zhang 0001, Tian Wang 0002, Guodong Guo, David S. Doermann |
ICCV | 4 |
| 2021 | DSSL: Deep Surroundings-person Separation Learning for Text-based Person RetrievalabstractMany previous methods on text-based person retrieval tasks are devoted to learning a latent common space mapping, with the purpose of extracting modality-invariant features from both visual and textual modality. Nevertheless, due to the complexity of high-dimensional data, the unconstrained mapping paradigms are not able to properly catch discriminative clues about the corresponding person while drop the misaligned information. Intuitively, the information contained in visual data can be divided into person information (PI) and surroundings information (SI), which are mutually exclusive from each other. To this end, we propose a novel Deep Surroundings-person Separation Learning (DSSL) model in this paper to effectively extract and match person information, and hence achieve a superior retrieval accuracy. A surroundings-person separation and fusion mechanism plays the key role to realize an accurate and effective surroundings-person separation under a mutually exclusion constraint. In order to adequately utilize multi-modal and multi-granular information for a higher retrieval accuracy, five diverse alignment paradigms are adopted. Extensive experiments are carried out to evaluate the proposed DSSL on CUHK-PEDES, which is currently the only accessible dataset for text-base person retrieval task. DSSL achieves the state-of-the-art performance on CUHK-PEDES. To properly evaluate our proposed DSSL in the real scenarios, a Real Scenarios Text-based Person Reidentification (RSTPReid) dataset is constructed to benefit future research on text-based person retrieval, which will be publicly available. Aichun Zhu, Zijie Wang 0003, Yifeng Li 0002, Xili Wan, Jing Jin 0002, Tian Wang 0002, Fangqiang Hu, Gang Hua 0002 |
ACM Multimedia | 6 |
| 2021 | Tiny-FASNet: A Tiny Face Anti-spoofing Method Based on Tiny Module
Ce Li 0001, Enbing Chang, Fenghua Liu, Shuxing Xuan, Tian Wang 0002 |
PRCV (3) | 6 |
| 2021 | An enhanced 3DCNN-ConvLSTM for spatiotemporal multimedia data analysisabstractSummary At present, human action recognition is a challenging and complex task in the field of computer vision. The combination of CNN and RNN is a common and effective network structure for this task. Especially, we use 3DCNN in CNN part and ConvLSTM in RNN part. We divide the video into multiple temporal segments by average and compress each segment into one feature map by pooling layer. Adding the pooling layer, dropout layer, and batch normalization layer into ConvLSTM is our groundbreaking work. We test our model on KTH, UCF‐11, and HMDB51 datasets and achieve a high accuracy of action recognition. Tian Wang 0002, Aichun Zhu, Hichem Snoussi, Chang Choi |
Concurr. Comput. Pract. Exp. | 1 |
| 2021 | Recent advances of single-object tracking methods: A brief survey
Tian Wang 0002, Baochang Zhang 0001, Lei Chen 0033 |
Neurocomputing | 2 |
| 2021 | Pose-Guided Inflated 3D ConvNet for action recognition in videos
Qianyu Wu, Aichun Zhu, Ran Cui, Tian Wang 0002, Fangqiang Hu, Yaping Bao, Hichem Snoussi |
Signal Process. Image Commun. | 4 |
| 2021 | RecapNet: Action Proposal Generation Mimicking Human Cognitive ProcessabstractGenerating action proposals in untrimmed videos is a challenging task, since video sequences usually contain lots of irrelevant contents and the duration of an action instance is arbitrary. The quality of action proposals is key to action detection performance. The previous methods mainly rely on sliding windows or anchor boxes to cover all ground-truth actions, but this is infeasible and computationally inefficient. To this end, this article proposes a RecapNet-a novel framework for generating action proposal, by mimicking the human cognitive process of understanding video content. Specifically, this RecapNet includes a residual causal convolution module to build a short memory of the past events, based on which the joint probability actionness density ranking mechanism is designed to retrieve the action proposals. The RecapNet can handle videos with arbitrary length and more important, a video sequence will need to be processed only in one single pass in order to generate all action proposals. The experiments show that the proposed RecapNet outperforms the state of the art under all metrics on the benchmark THUMOS14 and ActivityNet-1.3 datasets. The code is available publicly at https://github.com/tianwangbuaa/RecapNet. Tian Wang 0002, Yang Chen 0030, Zhiwei Lin 0002, Aichun Zhu, Yong Li 0025, Hichem Snoussi, Hui Wang 0001 |
IEEE Trans. Cybern. | 1 |
| 2021 | FT-MDnet: A Deep-Frozen Transfer Learning Framework for Person SearchabstractMatching manually cropped pedestrian images between queries and candidates, termed as person re-identification, has achieved significant progress with deep convolutional neural networks. Recently, a topic called ‘person search’ is proposed for the end-to-end application of re-identification technologies. It integrates object detection and person re-identification and aims to both locate and match pedestrians on a gallery of raw images. However, the design and implementation of such kind of hybrid network are difficult and computationally consuming in real practical situations. In order to fasten the design and ease the implementation, this paper proposes a deep-frozen transfer learning framework, named FT-MDnet, to extract re-identification features from a pre-trained detection network in two steps. First, using a channel-wise attention mechanism, a network called adaptive transfer learning network (ATLnet) is used to convert the sharing data of the underlying detection network to a re-identification feature map. Then, a multi-branch feature representation network called multiple descriptor network (MDnet) is proposed to extract re-identification features from the re-identification feature map. Our proposed solution has been verified on different types of mainstream detection networks, including YOLOv3, YOLOv4, Mask RCNN, and CenterNet. The experimental results show that our solution outperforms all other person search solutions by a large margin. It proves that the feature representations of detection networks are highly compatible with re-identification, and the proposed framework effectively extracts these features out. To encourage further research, we have made our framework open source. Ronghua Hu, Tian Wang 0002, Yi Zhou 0011, Hichem Snoussi, Abel Cherouat |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2021 | Online Detection of Action Start via Soft Computing for Smart CityabstractSoft computing is facing a rapid evolution thanks to the development of artificial intelligence especially the deep learning. With video surveillance technologies of soft computing, such as image processing, computer vision, and pattern recognition combined with cloud computing, the construction of smart cities could be maintained and greatly enhanced. In this article, we focus on the online detection of action start task in video understanding and analysis, which is critical to the multimedia security in smart cities. We propose a novel model to tackle this problem and achieves state-of-the-art results on the benchmark THUMOS14 data set. Tian Wang 0002, Yang Chen 0030, Hongqiang Lv, Jing Teng, Hichem Snoussi, Fei Tao 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2020 | Abnormal event detection via the analysis of multi-frame optical flow information
Tian Wang 0002, Meina Qiao, Aichun Zhu, Guangcun Shan, Hichem Snoussi |
Frontiers Comput. Sci. | 1 |
| 2020 | Exploring a rich spatial-temporal dependent relational model for skeleton-based action recognition by bidirectional LSTM-CNN
Aichun Zhu, Qianyu Wu, Ran Cui, Tian Wang 0002, Wenlong Hang, Gang Hua 0002, Hichem Snoussi |
Neurocomputing | 4 |
| 2019 | A reinforcement learning approach for UAV target searching and tracking
Tian Wang 0002, Ruoxi Qin, Yang Chen 0030, Hichem Snoussi, Chang Choi |
Multim. Tools Appl. | 1 |
| 2019 | Multiple human upper bodies detection via candidate-region convolutional neural network
Aichun Zhu, Tian Wang 0002 |
Multim. Tools Appl. | 2 |
| 2019 | Generative Neural Networks for Anomaly Detection in Crowded ScenesabstractSecurity surveillance is critical to social harmony and people's peaceful life. It has a great impact on strengthening social stability and life safeguarding. Detecting anomaly timely, effectively and efficiently in video surveillance remains challenging. This paper proposes a new approach, called S2-VAE, for anomaly detection from video data. The S2-VAE consists of two proposed neural networks: a Stacked Fully Connected Variational AutoEncoder (SF-VAE) and a Skip Convolutional VAE (SC-VAE). The SF-VAE is a shallow generative network to obtain a model like Gaussian mixture to fit the distribution of the actual data. The SC-VAE, as a key component of S2-VAE, is a deep generative network to take advantages of CNN, VAE and skip connections. Both SF-VAE and SC-VAE are efficient and effective generative networks and they can achieve better performance for detecting both local abnormal events and global abnormal events. The proposed S2-VAE is evaluated using four public datasets. The experimental results show that the S2-VAE outperforms the state-of-the-art algorithms. The code is available publicly at https://github.com/tianwangbuaa/. Tian Wang 0002, Meina Qiao, Zhiwei Lin 0002, Ce Li 0001, Hichem Snoussi, Zhe Liu 0001, Chang Choi |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2018 | Abnormal event detection via covariance matrix for optical flow based feature
Tian Wang 0002, Meina Qiao, Aichun Zhu, Yida Niu, Ce Li 0001, Hichem Snoussi |
Multim. Tools Appl. | 1 |
| 2018 | Distributed Harmonic Form ComputationabstractSpectral graph analysis based on Laplace operators has been ubiquitously applied in graph signal processing. Extending these tools to a generalization of graphs is helpful for several applications, such as network analysis, a sensor network coverage problem, and other fields, where the relationships between two or more vertices should be modeled. Such mathematical theories already exist, but the associated algorithms are still in their infancy. In this letter, we propose a new algorithm to compute harmonic forms, i.e., solutions of the Laplace equation, whose implementation is simple enough to be distributed among networks without central control center. Alban Goupil, Anas Hanaf, Tian Wang 0002 |
IEEE Signal Process. Lett. | 4 |
| 2016 | Detection of Abnormal Event in Complex Situations Using Strong Classifier Based on BP Adaboost
Tian Wang 0002, Meina Qiao, Aichun Zhu, Ce Li 0001, Hichem Snoussi |
ICIC (2) | 2 |
| 2016 | An enhancement method for X-ray image via fuzzy noise removal and homomorphic filtering
Limei Xiao, Ce Li 0001, Tian Wang 0002 |
Neurocomputing | 4 |
| 2015 | Joint Abnormal Blob Detection and Localization Under Complex Scenes
Tian Wang 0002, Keyu Lai, Ce Li 0001, Hichem Snoussi |
ICIC (1) | 1 |
| 2014 | Detection of Abnormal Visual Events via Global Optical Flow Orientation HistogramabstractThe aim of this paper is to detect abnormal events in video streams, a challenging but important subject in video surveillance. We propose a novel algorithm to address this problem. The algorithm is based on an image descriptor and a nonlinear classification method. We introduce a histogram of optical flow orientation as a descriptor encoding the moving information of each video frame. The nonlinear one-class support vector machine classification algorithm, following a learning period characterizing the normal behavior of training frames, detects abnormal events in the current frame. Further, a fast version of the detection algorithm is designed by fusing the optical flow computation with a background subtraction step. We finally apply the method to detect abnormal events on several benchmark data sets, and show promising results. Tian Wang 0002, Hichem Snoussi |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2012 | Histograms of Optical Flow Orientation for Visual Abnormal Events DetectionabstractIn this paper, we propose an algorithm to detect abnormal events based on video streams. The algorithm is based on histograms of the orientation of optical flow descriptor and one-class SVM classifier. We introduce grids of Histograms of the Orientation of Optical Flow (HOFs) as the descriptors for motion information of the monolithic video frame. The one-class SVM, after a learning period characterizing normal behaviors, detects the abnormal events in the current frame. Extensive testing on benchmark dataset corroborates the effectiveness of the proposed detection method. Tian Wang 0002, Hichem Snoussi |
AVSS | 1 |