Tongzhen Si

dblp:223/7053 · DBLP profile ↗
← Back
25ranked-venue papers
11as first author
24since 2021 · last 2026
0000-0002-1141-9718ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Computer networks · 3 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Label-knowledge guided heterogeneous-temporal graph network for multimodal intent understanding
Tongzhen Si, Penglei Li, Fazhi He
Knowl. Based Syst.1
2026 Dependability Feature Learning Based on Sample Generation for Unsupervised Text-to-Image Person Re-Identification
abstract
Text-to-image person re-identification (TIReID) aims to retrieve the target pedestrians according to specific textual descriptions. Benefiting from abundant annotated training data, current supervised TIReID methods have achieved impressive performance. However, annotating cross-modality data is extremely time-consuming, which limits their application in real-world scenarios. Several methods attempt to generate text descriptions or pseudo-labels but neglect the dependability of image-text matching relationships or identity information. To this end, we propose a Dependability Feature Learning based on Sample Generation (DFLSG) for unsupervised TIReID. First, we introduce a dependable text generation method that leverages multimodal large language models to generate diverse texts and further filtrate dependable texts for establishing image-text matching relationships. Second, we design an Error Sample Filtering Module (ESFM) to eliminate abnormal samples and obtain reliable identity labels. Furthermore, we develop a Multilevel Triplet Joint Learning (MTJL) process, which continuously optimizes the cross-modality dependable feature from center and instance views. Extensive experiments are implemented to assess the proposed DFLSG on four mainstream TIReID databases. Experimental results demonstrate that DFLSG achieves state-of-the-art performance compared with other unsupervised methods. Code will be available at: https://github.com/CLS-2001/DFLSG.
Chenglong Shao, Tongzhen Si, Hui Yuan 0001
IEEE Trans. Circuits Syst. Video Technol.2
2026 Learning Universal Attack via Model-Guided Meta-Learning for Person Reidentification
Tongzhen Si, Penglei Li, Fazhi He, Zhiquan Feng, Tao Xu 0021
IEEE Trans. Ind. Informatics1
2025 Driver Cognitive Distraction Detection based on eye movement behavior and integration of multi-view space-channel feature
Yu Qiao 0001, Tongzhen Si, Qingbei Guo
Expert Syst. Appl.4
2025 Exploring granularity-associated invariance features for text-to-image person re-identification
Chenglong Shao, Tongzhen Si
Multim. Syst.2
2025 Eye-SCAN: Eye-Movement-Attention-based Spatial Channel Adaptive Network for traffic accident prediction
Yu Qiao 0001, Tongzhen Si, Tao Xu 0021
Pattern Recognit.3
2024 Invisible Backdoor Attack against 3D Point Cloud Classifier in Graph Spectral Domain
abstract
3D point cloud has been wildly used in security crucial domains, such as self-driving and 3D face recognition. Backdoor attack is a serious threat that usually destroy Deep Neural Networks (DNN) in the training stage. Though a few 3D backdoor attacks are designed to achieve guaranteed attack efficiency, their deformation will alarm human inspection. To obtain invisible backdoored point cloud, this paper proposes a novel 3D backdoor attack, named IBAPC, which generates backdoor trigger in the graph spectral domain. The effectiveness is grounded by the advantage of graph spectral signal that it can induce both global structure and local points to be responsible for the caused deformation in spatial domain. In detail, a new backdoor implanting function is proposed whose aim is to transform point cloud to graph spectral signal for conducting backdoor trigger. Then, we design a backdoor training procedure which updates the parameter of backdoor implanting function and victim 3D DNN alternately. Finally, the backdoored 3D DNN and its associated backdoor implanting function is obtained by finishing the backdoor training procedure. Experiment results suggest that IBAPC achieves SOTA attack stealthiness from three aspects including objective distance measurement, subjective human evaluation, graph spectral signal residual. At the same time, it obtains competitive attack efficiency. The code is available at https://github.com/f-lk/IBAPC.
Linkun Fan, Fazhi He, Tongzhen Si, Wei Tang 0018, Bing Li 0010
AAAI3
2024 Driver Cognitive Distraction Detection Based on Eye Movement Behavior and Spatio-Temporal Information Fusion
Yu Qiao 0001, Tongzhen Si, Qingbei Guo
ICONIP (10)4
2024 UnifiedSC: a unified framework via collaborative optimization for multi-task person re-identification
Tongzhen Si, Fazhi He, Penglei Li
Appl. Intell.1
2024 Semantic-aware transformer with feature integration for remote sensing change detection
Penglei Li, Tongzhen Si, Chuanlong Ye, Qingbei Guo
Eng. Appl. Artif. Intell.2
2024 MFFSP: Multi-scale feature fusion scene parsing network for landslides detection based on high-resolution satellite images
Penglei Li, Yi Wang 0021, Tongzhen Si, Kashif Ullah, Wei Han 0006, Lizhe Wang 0001
Eng. Appl. Artif. Intell.3
2024 Homogeneous and Heterogeneous Optimization for Unsupervised Cross-Modality Person Reidentification in Visual Internet of Things
abstract
Cross-modality visible-infrared person reidentification (VI-ReID) has attracted widespread concern due to its scalability in 24-h video surveillance of the Visual Internet of Things (VIoT). Driven by enough annotated training data, supervised VI-ReID has achieved superior performance. However, annotating a large amount of cross-modality data is extremely time-consuming, which limits its employment in real-world scenarios. Existing several works neglect the image-level discrepancy and could not obtain reliable feature-level heterogeneous correlation. In this article, we propose a novel homogeneous and heterogeneous optimization with modality style adaptation (HHO) mechanism to eliminate intramodality and intermodality discrepancies without any label information for unsupervised VI-ReID. Specifically, we present the modality style adaptation strategy to transfer unlabeled cross-modality pedestrian styles, which not only increases the image diversity but also bridges the intermodality gap. Meanwhile, we employ the clustering algorithm to generate pseudo labels for each modality. The homogeneous feature optimization is developed to extract intramodality pedestrian features. Furthermore, we propose heterogeneous feature optimization to eliminate the intermodality discrepancy. To this end, a heterogeneous feature search (HFS) module is designed to mine reliable cross-modality signals for each identity. These reliable heterogeneous features are constrained to generate the compact feature distribution, while different identities are forced to be separated. The HHO are seamlessly integrated to learn cross-modality robust features. Abundant experiments prove the superiority of HHO, which gains superior performance.
Tongzhen Si, Fazhi He, Penglei Li, Mang Ye
IEEE Internet Things J.1
2024 Model-aware privacy-preserving with start trigger method for person re-identification
Tongzhen Si, Penglei Li, Linkun Fan, Fazhi He
Inf. Process. Manag.1
2024 MBA: Backdoor Attacks Against 3D Mesh Classifier
abstract
3D mesh classification deep neural network (3D DNN) has been widely applied in many safety-critical domains. Backdoor attack is a serious threat that occurs during the training stage. Previous backdoor attacks from 2D image and 3D point cloud domains are not suitable for 3D mesh due to data structure restrictions. Therefore, in a pioneering effort, this paper presents two types of backdoor attacks on 3D mesh. Specifically, the first attack is a Mesh Geometrical Feature guided 3D Mesh Backdoor Attack named MGF-MBA. Most 3D DNNs have to convert 3D mesh to a regular matrix (mesh geometrical feature), which is a refinement of the input 3D mesh. The 3D DNN directly learns the 3D shape from the mesh geometrical feature, which enables attackers to implant backdoor through it. Hence, the proposed MGF-MBA generates a backdoored 3D mesh under the guidance of mesh geometrical feature. The second attack is a Remeshing based 3D Mesh Backdoor Attack named ReMBA. The quality of samples backdoored by exiting backdoor attacks always decrease. Although many efforts have been made to reduce the descent in quality in return for stealthiness, the descent persists. For better stealthiness, we regard the backdoor implantation process as a way to increase the quality of backdoored sample rather than a way to reduce it. Specifically, ReMBA designs a new isotropic remeshing method that attempts to represent a 3D mesh by equilateral triangles while keeping the number of vertices, edges and faces unchanged. Numerous experimental results show that both MGF-MBA and ReMBA achieve guaranteed attack performance on 3D DNNs. Furthermore, transferability experiments demonstrate that ReMBA can even attack 3D point cloud networks with an increased ability to resist defenses.
Linkun Fan, Fazhi He, Tongzhen Si, Rubin Fan, Chuanlong Ye, Bing Li 0010
IEEE Trans. Inf. Forensics Secur.3
2023 HIGSA: Human image generation with self-attention
Fazhi He, Tongzhen Si, Yansong Duan, Xiaohu Yan
Adv. Eng. Informatics3
2023 Tri-modality consistency optimization with heterogeneous augmented images for visible-infrared person re-identification
Tongzhen Si, Fazhi He, Penglei Li, Xiaoxin Gao
Neurocomputing1
2023 Diversity feature constraint based on heterogeneous data for unsupervised person re-identification
Tongzhen Si, Fazhi He, Penglei Li, Yupeng Song, Linkun Fan
Inf. Process. Manag.1
2023 DATFuse: Infrared and Visible Image Fusion via Dual Attention Transformer
abstract
The fusion of infrared and visible images aims to generate a composite image that can simultaneously contain the thermal radiation information of an infrared image and the plentiful texture details of a visible image to detect targets under various weather conditions with a high spatial resolution of scenes. Previous deep fusion models were generally based on convolutional operations, resulting in a limited ability to represent long-range context information. In this paper, we propose a novel end-to-end model for infrared and visible image fusion via a dual attention Transformer termed DATFuse. To accurately examine the significant areas of the source images, a dual attention residual module (DARM) is designed for important feature extraction. To further model long-range dependencies, a Transformer module (TRM) is devised for global complementary information preservation. Moreover, a loss function that consists of three terms, namely, pixel loss, gradient loss, and structural loss, is designed to train the proposed model in an unsupervised manner. This can avoid manually designing complicated activity-level measurement and fusion strategies in traditional image fusion methods. Extensive experiments on public datasets reveal that our DATFuse outperforms other representative state-of-the-art approaches in both qualitative and quantitative assessments. The proposed model is also extended to address other infrared and visible image fusion tasks without fine-tuning, and the promising results demonstrate that it has good generalization ability. The source code is available athttps://github.com/tthinking/DATFuse.
Wei Tang 0018, Fazhi He, Yu Liu 0023, Yansong Duan, Tongzhen Si
IEEE Trans. Circuits Syst. Video Technol.5
2023 A Multistrategy Evolutionary Multiobjective Optimization Method for Hyperspectral Endmember Extraction
abstract
Hyperspectral endmember extraction (HEE) is an essential part of remote-sensing image processing. There have been recent attempts to model the HEE as a multiobjective optimization problem and apply multiobjective evolutionary algorithms to solve the problem. However, because of the large HEE search space, it is difficult for the current algorithms to achieve exploration–exploitation balance, and they easily stall prematurely. To address these issues, this article proposes a multistrategy evolutionary multiobjective method based on roulette wheel selection and the genetic algorithm (RWS-GA) for endmember extraction. This method designs two parallel algorithms corresponding to global exploration and local exploitation. In the RWS-GA, an improved NSGA-II method, adopting a novel method to sort individuals on the same front instead of the crowding distance, is proposed to divide individuals into superior and inferior subpopulations. Thereafter, different modified population update strategies are utilized for subpopulations based on characteristics. Pixels that appear more frequently in the population are considered to perform better to have a higher probability of forming an endmember set with other pixels. In addition, excellent individuals often exhibit a higher probability of including endmembers compared with inferior individuals. Considering the abovementioned opinions, roulette wheel selection is performed on the inferior subpopulation for global search. Meanwhile, the superior subpopulation is responsible for local search based on the genetic algorithm (GA). Furthermore, an offspring complement mechanism (OCM) is presented to prevent duplicate individuals from appearing in historical archives. Numerous comparative experiments show that the proposed method is superior to other endmember extraction methods in three real-world datasets.
Chuanlong Ye, Fazhi He, Jinkun Luo, Lyuyang Tong, Xiaoxin Gao, Tongzhen Si, Linkun Fan
IEEE Trans. Geosci. Remote. Sens.6
2023 Hybrid Contrastive Learning for Unsupervised Person Re-Identification
abstract
Unsupervised person re-identification (Re-ID) aims to learn discriminative features without human-annotated labels. Recently, contrastive learning has provided a new prospect for unsupervised person Re-ID, and existing methods primarily constrain the feature similarity among easy sample pairs. However, the feature similarity among hard sample pairs is neglected, which yields suboptimal performance in unsupervised person Re-ID. In this paper, we propose a novel Hybrid Contrastive Model (HCM) to perform the identity-level contrastive learning and the image-level contrastive learning for unsupervised person Re-ID, which adequately explores feature similarities among hard sample pairs. Specifically, for the identity-level contrastive learning, an identity-based memory is constructed to store pedestrian features. Accordingly, we define the dynamic contrast loss to identify identity information with dynamic factor for distinguishing hard/easy samples. As for the image-level contrastive learning, an image-based memory is established to store each image feature. We design the sample constraint loss to explore the similarity relationship between hard positive and negative sample pairs. Furthermore, we optimize the two contrastive learning processes in one unified framework to make use of their own advantages as so to constrain the feature distribution for extracting potential information. Extensive experiments demonstrate that the proposed HCM distinctly outperforms existing methods.
Tongzhen Si, Fazhi He, Zhong Zhang 0001, Yansong Duan
IEEE Trans. Multim.1
2023 Hybrid feature constraint with clustering for unsupervised person re-identification
Tongzhen Si, Fazhi He, Penglei Li
Vis. Comput.1
2022 Spatial-driven features based on image dependencies for person re-identification
Tongzhen Si, Fazhi He, Yansong Duan
Pattern Recognit.1
2022 LSLPCT: An Enhanced Local Semantic Learning Transformer for 3-D Point Cloud Analysis
abstract
The 3D point cloud is a common 3D data representation that has received increasing attention for remote sensing applications. However, processing 3D point cloud semantics, especially local semantic information, has always been a challenge and has attracted much attention. In this paper, we propose a novel enhanced local semantic learning transformer for 3D point cloud analysis, which aims to enhance the transformer awareness of local semantic features to handle complex point cloud tasks. First, we propose a novel transformer framework, the local semantic learning point cloud transformer (LSLPCT), which not only learns 3D point clouds the global information, but also enhances the perception of local semantic information end-to-end. Second, we design an efficient local semantic learning self-attention mechanism, namely LSL-SA, which can parallelize the perception of global contextual information and the capture of finer-grained local semantic features. Third, our proposed LSL-SA is easy to implement and can integrate existing transformers and CNN-based networks for processing various point cloud tasks. Numerous experiments in different types of point cloud tasks have been conducted, and our method performs better or is competitive with other state-of-the-art methods.
Yupeng Song, Fazhi He, Yansong Duan, Tongzhen Si, Junwei Bai
IEEE Trans. Geosci. Remote. Sens.4
2021 Dynamically occluded samples via adversarial learning for person re-identification in sensor networks
Wenmin Huang, Shuang Liu 0001, Ruiling Luo, Tongzhen Si, Zhong Zhang 0001
Ad Hoc Networks4
2019 Compact Triplet Loss for person re-identification in camera sensor networks
Tongzhen Si, Zhong Zhang 0001, Shuang Liu 0001
Ad Hoc Networks1