Dazhen Lin

dblp:32/6871 · DBLP profile ↗
← Back
31ranked-venue papers
4as first author
20since 2021 · last 2025
0000-0002-7221-5591ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 7 since 2021Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Security and privacy · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Pick and mix reliable pseudo labels for scribble-supervised medical image segmentation
Jiawei Su, Zhiming Luo, Dazhen Lin, Lihui Lin, Shaozi Li
Neurocomputing3
2025 Short video rumor detection based on causal graph
Donglin Cao, Xiong Tang, Yanghao Lin, Dazhen Lin
Inf. Sci.4
2025 Cross-modality average precision optimization for visible thermal person re-identification
Yongguo Ling, Zhiming Luo, Dazhen Lin, Shaozi Li, Min Jiang 0005, Nicu Sebe, Zhun Zhong
Pattern Recognit.3
2025 A Self-Adaptive Feature Extraction Method for Aerial-View Geo-Localization
abstract
Cross-view geo-localization aims to match the same geographic location from different view images, e.g., drone-view images and geo-referenced satellite-view images. Due to UAV cameras' different shooting angles and heights, the scale of the same captured target building in the drone-view images varies greatly. Meanwhile, there is a difference in size and floor area for different geographic locations in the real world, such as towers and stadiums, which also leads to scale variants of geographic targets in the images. However, existing methods mainly focus on extracting the fine-grained information of the geographic targets or the contextual information of the surrounding area, which overlook the robust feature for scale changes and the importance of feature alignment. In this study, we argue that the key underpinning of this task is to train a network to mine a discriminative representation against scale variants. To this end, we design an effective and novel end-to-end network called Self-Adaptive Feature Extraction Network (Safe-Net) to extract powerful scale-invariant features in a self-adaptive manner. Safe-Net includes a global representation-guided feature alignment module and a saliency-guided feature partition module. The former applies an affine transformation guided by the global feature for adaptive feature alignment. Without extra region annotations, the latter computes saliency distribution for different regions of the image and adopts the saliency information to guide a self-adaptive feature partition on the feature map to learn a visual representation against scale variants. Experiments on two prevailing large-scale aerial-view geo-localization benchmarks, i.e., University-1652 and SUES-200, show that the proposed method achieves state-of-the-art results. In addition, our proposed Safe-Net has a significant scale adaptive capability and can extract robust feature representations for those query images with small target buildings. The source code of this study is available at: https://github.com/AggMan96/Safe-Net.
Jinliang Lin, Zhiming Luo, Dazhen Lin, Shaozi Li, Zhun Zhong
IEEE Trans. Image Process.3
2024 Interpretable Short Video Rumor Detection Based on Modality Tampering
abstract
With the rapid development of social media and short video applications in recent years, browsing short videos has become the norm. Due to its large user base and unique appeal, spreading rumors via short videos has become a severe social problem. Many methods simply fuse multimodal features for rumor detection, which lack interpretability. For short video rumors, rumor makers create rumors by modifying and/or splicing different modal information, so we should consider how to detect rumors from the perspective of modality tampering. Inspired by cross-modal contrastive learning, we propose a novel short video rumor detection framework by designing two pretraining tasks: modality tampering detection and inter-modal matching, imbuing the model with the ability to detect modality tampering and employing it for downstream rumor detection tasks. In addition, we design an interpretability mechanism to make the rumor detection results more reasonable by backtracking the model’s decision-making process. The experimental results show that the method on the short video rumor dataset has an improvement of about 4.6%-12% in macro-F1 compared with other models and can explain whether the short video is a rumor or not through the perspective of modality tampering.
Kaixuan Wu, Yanghao Lin, Donglin Cao, Dazhen Lin
LREC/COLING4
2024 Leverage Causal Graphs and Rumor-Refuting Texts for Interpretable Rumor Analysis
abstract
Previous rumors detection study mostly ignores causal features in rumor texts and the interpretability of rumor classification results. Based on the phenomenon that rumors can cause changes or even loss of the original causal relationship in the truth, we can leverage causal features to help classify rumors. Meanwhile, rumor-refuting text is of great significance in curbing the spread of rumors and can interpret the results of rumor detection. To address these challenges, we propose a rumor analysis model integrating rumor detection and rumor refuting. It builds the causal graph of rumor texts and uses it to improve the performance of rumor detection. In particular, we embed the rumor-refuting text generation module in the model to realize the integration of rumor detection and rum -cation results’ analytical ability. We conducted experiments on two benchmark datasets and performed better than the state-of-the-art methods.
Donglin Cao, Dazhen Lin
ICASSP3
2024 End-To-End Spatially-Constrained Multi-Perspective Fine-Grained Image Captioning
abstract
The perspective of captions in fine-grained image captioning crucially impacts people’s perception and understanding of the image. However, existing methods often overlook this aspect, resulting in captions that struggle to accurately convey the image’s hierarchical and spatial information. In this paper, we propose an end-to-end Spatially-Constrained multi-perspective fine-grained Image Captioning (SCIC) model. SCIC initially predicts the optimal perspective for captioning the image and subsequently utilizes this optimal perspective as a constraint condition for caption generation. Furthermore, SCIC is capable of generating multi-perspective image captions based on customized perspectives. Experimental results show that our model effectively improves the state-of-the-art CIDEr score by about 14.38% and can generate multi-perspective fine-grained captions for the same image.
Chunzhen Lin, Donglin Cao, Dazhen Lin
ICASSP4
2024 Comparative evaluation of recent universal adversarial perturbations in image classification
Juanjuan Weng, Zhiming Luo, Dazhen Lin, Shaozi Li
Comput. Secur.3
2024 Learning transferable targeted universal adversarial perturbations by sequential meta-learning
Juanjuan Weng, Zhiming Luo, Dazhen Lin, Shaozi Li
Comput. Secur.3
2024 Learning multi-organ and tumor segmentation from partially labeled datasets by a conditional dynamic attention network
abstract
Summary Multi‐organ segmentation is a critical prerequisite for many clinical applications. Deep learning‐based approaches have recently achieved promising results on this task. However, they heavily rely on massive data with multi‐organ annotated, which is labor‐ and expert‐intensive and thus difficult to obtain. In contrast, single‐organ datasets are easier to acquire, and many well‐annotated ones are publicly available. It leads to the partially labeled issue: How to learn a unified multi‐organ segmentation model from several single‐organ datasets? Pseudo‐label‐based methods and conditional information‐based methods make up the majority of existing solutions, where the former largely depends on the accuracy of pseudo‐labels, and the latter has a limited capacity for task‐related features. In this paper, we propose the Conditional Dynamic Attention Network (CDANet). Our approach is designed with two key components: (1) multisource parameter generator, fusing the conditional and multiscale information to better distinguish among different tasks, and (2) dynamic attention module, promoting more attention to task‐related features. We have conducted extensive experiments on seven partially labeled challenging datasets. The results show that our method achieved competitive results compared with the advanced approaches, with an average Dice score of 75.08%. Additionally, the Hausdorff Distance is 26.31, which is a competitive result.
Lei Li 0048, Sheng Lian, Dazhen Lin, Zhiming Luo, Beizhan Wang, Shaozi Li
Concurr. Comput. Pract. Exp.3
2024 An autoencoder-based self-supervised learning for multimodal sentiment analysis
Wenjun Feng, Donglin Cao, Dazhen Lin
Inf. Sci.4
2024 Mutual learning with reliable pseudo label for semi-supervised medical image segmentation
Jiawei Su, Zhiming Luo, Sheng Lian, Dazhen Lin, Shaozi Li
Medical Image Anal.4
2024 Boosting Adversarial Transferability via Logits Mixup With Dominant Decomposed Feature
abstract
Recent research has shown that adversarial samples are highly transferable and can be used to attack other unknown black-box Deep Neural Networks (DNNs). To improve the transferability of adversarial samples, several feature-based adversarial attack methods have been proposed to disrupt neuron activation in the middle layers. However, current state-of-the-art feature-based attack methods typically require additional computation costs for estimating the importance of neurons. To address this challenge, we propose a Singular Value Decomposition (SVD)-based feature-level attack method. Our approach is inspired by the discovery that eigenvectors associated with the larger singular values decomposed from the middle layer features exhibit superior generalization and attention properties. Specifically, we conduct the attack by retaining the dominant decomposed feature that corresponds to the largest singular value (i.e., Rank-1 decomposed feature) for computing the output logits before the final softmax. These logits are later integrated with the original logits to optimize adversarial examples. Our extensive experimental results verify the effectiveness of our proposed method, which can be easily integrated into various baselines to significantly enhance the transferability of adversarial samples for disturbing normally trained CNNs and advanced defense strategies. The source code is available at Link.
Juanjuan Weng, Zhiming Luo, Shaozi Li, Dazhen Lin, Zhun Zhong
IEEE Trans. Inf. Forensics Secur.4
2023 Exploring Non-target Knowledge for Improving Ensemble Universal Adversarial Attacks
abstract
The ensemble attack with average weights can be leveraged for increasing the transferability of universal adversarial perturbation (UAP) by training with multiple Convolutional Neural Networks (CNNs). However, after analyzing the Pearson Correlation Coefficients (PCCs) between the ensemble logits and individual logits of the crafted UAP trained by the ensemble attack, we find that one CNN plays a dominant role during the optimization. Consequently, this average weighted strategy will weaken the contributions of other CNNs and thus limit the transferability for other black-box CNNs. To deal with this bias issue, the primary attempt is to leverage the Kullback–Leibler (KL) divergence loss to encourage the joint contribution from different CNNs, which is still insufficient. After decoupling the KL loss into a target-class part and a non-target-class part, the main issue lies in that the non-target knowledge will be significantly suppressed due to the increasing logit of the target class. In this study, we simply adopt a KL loss that only considers the non-target classes for addressing the dominant bias issue. Besides, to further boost the transferability, we incorporate the min-max learning framework to self-adjust the ensemble weights for each CNN. Experiments results validate that considering the non-target KL loss can achieve superior transferability than the original KL loss by a large margin, and the min-max training can provide a mutual benefit in adversarial ensemble attacks. The source code is available at: https://github.com/WJJLL/ND-MM.
Juanjuan Weng, Zhiming Luo, Zhun Zhong, Dazhen Lin, Shaozi Li
AAAI4
2023 Multimodal Rumor Detection with Causal Graph Attention Network
abstract
In the era of big data, the automatic detection and governance of rumors are of great significance in protecting the privacy and security of users. Many previous works have achieved good results in rumor detection through multimodal feature fusion, background knowledge expansion, etc. Causality in rumors is a potential textual feature representing rumors’ linguistic logic and causality semantics. The rumors’ causality often appears inconsistent with the truth, based on which we can improve the rumor detection model. However, previous work has neglected the use of the feature of causality between entities in rumor texts. To address these challenges, we propose a Causal Graph Attention Network (CGANet) for multimodal rumor detection, which constructs a causal graph for the rumor dataset to achieve better detection performance. CGANet extracts the node and causal relationship features of causal graphs through graph learning and enhances the causal semantics of multimodal features, thereby improving the detection ability of the model on false causal semantic content. Finally, we use a self-attention classification network to obtain the rumor detection results. We conducted experiments on two public benchmark datasets, Pheme and Weibo, and achieved an accuracy of 92.81% and 91.32%, respectively, better performance than the state-of-the-art methods. It proves the advantage of causal graphs in rumor detection tasks.
Donglin Cao, Dazhen Lin
ICPADS3
2023 Multimodal Causal Relations Enhanced CLIP for Image-to-Text Retrieval
Wenjun Feng, Dazhen Lin, Donglin Cao
PRCV (1)2
2023 Text Causal Discovery Based on Sequence Structure Information
Donglin Cao, Dazhen Lin
PRCV (7)3
2023 Dual-Stream Transformer With Distribution Alignment for Visible-Infrared Person Re-Identification
abstract
Visible-infrared person re-identification(VI-ReID) aims to match the person images captured by visible and infrared cameras and suffers from severe cross-modality discrepancy and intra-modality variations. Existing approaches mainly use convolution neural network (CNN)-based architectures to extract pedestrian features, which fail to capture the long-range dependencies within an image. In addition, previous works usually attempt to bridge the modality gap by using adversarial learning to generate style-consistent images or designing different feature-level metric learning constraints. However, few works consider the cross-modality disparity from the perspective of assessing overall distance distribution discrepancy. To address these problems, we design a pure Transformer-based Visible-Infrared (TransVI) network with a conventional two-stream structure, which can explicitly capture modality-specific representations and learn multi-modality sharable knowledge. TransVI can efficiently address the lack of global dependency in CNN-based architectures due to the multi-head self-attention modules in the transformer, which allows us to capture the long-range dependencies of pedestrian images. Furthermore, we introduce the Cross-Modality Dissimilarity-based Maximum Mean Discrepancy (CMD-MMD) constraint to handle the cross-modality discrepancy at the distance distribution level. Specifically, CMD-MMD leverages intra-modality distribution separability to guide inter-modality distribution separability learning, aligning pair-wise distance distributions of intra- and inter-modality for within-class and between-class, respectively. In this way, the distance distributions of intra- and inter-modality become more similar, significantly mitigating the cross-modality discrepancy and learning more modality invariant representations. Extensive experimental results on two public VI-ReID datasets confirm that our proposed framework can achieve state-of-the-art performance.
Zehua Chai, Yongguo Ling, Zhiming Luo, Dazhen Lin, Min Jiang 0005, Shaozi Li
IEEE Trans. Circuits Syst. Video Technol.4
2021 Grammar guided embedding based Chinese long text sentiment classification
abstract
Abstract Although the state‐of‐the‐art sentiment classification approaches, such as LSTM and TextCNN, have achieved a good performance on Chinese short text sentiment analysis, the Chinese long text sentiment classification is still a challenge because of the sentiment change problem and the long text structure problem. Therefore, we propose a grammar guided embedding model (GGE) and a novel Chinese long text sentiment classification framework. First, the part‐of‐speech (POS) tags are introduced as the Chinese long text grammar guided information which can help classification approaches to model the Chinese long text structure and the important structure of sentiment change. Second, we proposed a simple GGE training method which considers the combination representation of word sequence and POS sequence. Finally, we proposed a unified framework which combines our novel GGE with TextCNN. Experiment results show that after using GGE, the model outperforms the state‐of‐the‐art approaches. At the same time, we also found that the GGE achieves the model converge faster, that is, it can achieve better results than without GGE when there is only a small amount of training data. Thus, we believe that the GGE can help machines better understand human language sentiment expression structure.
Dazhen Lin, Donglin Cao, Shaozi Li
Concurr. Comput. Pract. Exp.2
2021 Rumor knowledge embedding based data augmentation for imbalanced rumor detection
Xiangyan Chen, Duoduo Zhu, Dazhen Lin, Donglin Cao
Inf. Sci.3
2019 Chinese microblog rumor detection based on deep sequence context
abstract
Summary Rumor is one of the main problems in social media, which often shows deeply and rapidly undesirable affection on the society. Although many rumor detection models consider content features and social features, all of them are based on the word independence assumption, which lacks the sequence context. Thus, if we use some words that often appear in rumors, our posts will be recognized as a rumor. To solve this problem, we propose a deep sequence context model (DSCM) for Chinese microblog rumor detection. This model considers two important factors of rumors: falsity and influence. Firstly, to learn falsity, we abolish the word independence assumption and use long short‐term memory (LSTM) units to capture bi‐direction sequence context information in content. Secondly, to learn influence, we combine the deep sequence context information with social features to learn the connection between content and social features. In our experiment, our results show that our approach outperforms several state‐of‐the‐art machine learning approaches, including term frequency and inverse document frequency (TFIDF), LSTM, and gated recurrent unit (GRU) in rumor detection.
Dazhen Lin, Donglin Cao, Shaozi Li
Concurr. Comput. Pract. Exp.1
2018 Multi-modality weakly labeled sentiment learning based on Explicit Emotion Signal for Chinese microblog
Dazhen Lin, Donglin Cao, Yanping Lv, Xiao Ke
Neurocomputing1
2018 Chinese Character CAPTCHA Recognition and performance estimation via deep neural network
Dazhen Lin, Fan Lin, Yanping Lv, Feipeng Cai, Donglin Cao
Neurocomputing1
2018 Chinese microblog users' sentiment-based traffic condition analysis
Donglin Cao, Shuru Wang, Dazhen Lin
Soft Comput.3
2018 Fuzzy cerebellar model articulation controller network optimization via self-adaptive global best harmony search algorithm
Fei Chao 0001, Dajun Zhou, Chih-Min Lin, Changle Zhou, Minghui Shi, Dazhen Lin
Soft Comput.6
2016 A spatial-temporal visual mid-level ontology for GIF sentiment analysis
abstract
With the progress of social medias, an increasing number of dynamic multimedia, such as video clips (GIF), were used in Internet. Many of them show the subjective sentiment of users. GIF sentiment analysis is important for political election predication and economic indicator evaluation. Unfortunately, such task is quite challenging because the users' sentiment hinges on spatial-temporal visual concepts. And the relationship between such concepts and overall sentiment polarity remains unknown. In this paper, dedicated to build a bridge to exploring such relationship, we proposed a SentiPair Sequence based spatial-temporal visual sentiment ontology. The ontology serves as a mid-level representation of GIF sentiment analysis. In order to analyze sentiment polarity, we first constructed a Synset Forest to define the semantic tree structure of visual sentiment concepts. Then, through the Synset Forest, we organically select and combine sentiment label elements to form a mid-level visual sentiment representation. Our experiments indicate that SentiPair outperforms Adjective Noun Pairs which is a state-of-art visual sentiment ontology for image. We also released our dataset (GSO-2015) to the research community. GSO-2015 contains more than 6,000 manually labeled GIFs and 40,000 unlabeled GIFs. Each is labeled with both sentiment and SentiPair Sequence.
Zheng Cai, Donglin Cao, Dazhen Lin, Rongrong Ji
CEC3
2016 Chinese character CAPTCHA recognition based on convolution neural network
abstract
CAPTCHAs (Completely Automated Public Turing test to tell Computers and Humans Apart) are increasingly used in many applications for machine and human identification. Compared with traditional English and digital characters based CAPTCHAs, Chinese characters contain more complicated characters which greatly enhance difficulty of automatic recognition. To solve that problem, we proposed a Convolution Neural Network (CNN) based approach. This approach greatly improves the recognition accuracy of Chinese Character CAPTCHAs with distortion, rotation and background noise. Our experiment results show that this approach achieves more than 95% accuracy for single character and 84% accuracy for three types of Chinese Character CAPTCHAs with four characters. This encouraging result indicates that deep neural network is useful in complicated structure perception of Chinese Character CAPTCHAs.
Yanping Lv, Feipeng Cai, Dazhen Lin, Donglin Cao
CEC3
2016 A cross-media public sentiment analysis system for microblog
Donglin Cao, Rongrong Ji, Dazhen Lin, Shaozi Li
Multim. Syst.3
2016 Visual sentiment topic model based microblog image sentiment analysis
Donglin Cao, Rongrong Ji, Dazhen Lin, Shaozi Li
Multim. Tools Appl.3
2010 Making intelligent business decisions by mining the implicit relation from bloggers' posts
Dazhen Lin, Shaozi Li, Donglin Cao
Soft Comput.1
2006 A Collaborative Multimedia Editing System Based on Shallow Nature Language Parsing
Donglin Cao, Dazhen Lin, Shaozi Li
CDVE2