Yumin Tian

dblp:59/6260 · DBLP profile ↗
← Back
26ranked-venue papers
8as first author
18since 2021 · last 2025
0000-0002-2131-6456ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Security and privacy · 1
YearPublicationVenuePosition
2025 Global-aware Fragment Representation Aggregation Network for image-text retrieval
Di Wang 0011, Jiabo Tian, Yumin Tian, Lihuo He
Pattern Recognit.4
2024 TMFN: A Target-oriented Multi-grained Fusion Network for End-to-end Aspect-based Multimodal Sentiment Analysis
abstract
End-to-end multimodal aspect-based sentiment analysis (MABSA) combines multimodal aspect terms extraction (MATE) with multimodal aspect sentiment classification (MASC), aiming to simultaneously extract aspect words and classify the sentiment polarity of each aspect. However, existing MABSA methods have overlooked two issues: (i) They only focus on fusing image regional information and textual words for two subtasks of MABSA. Whereas, MATE subtask relies more on global image information to assist in obtaining the quantity and attributes of aspects. Ignoring the integration with global information may affect the performance of MABSA methods. (ii) They fail to take advantage of target information. Nevertheless, the fine-grained details of targets are important for classifying sentiments of aspects. To solve these problems, we propose a Target-oriented Multi-grained Fusion Network(TMFN). It fuses text information with global coarse-grained image information for MATE subtask and with fine-grained image information for MASC subtask. In addition, a target-oriented feature alignment (TOFA) module is designed to enhance target-related information in image features with target details. In such a way, image features will contain more target emotional-related information which is beneficial to sentiment classification. Extensive experiments show that our method outperforms state-of-the-art methods on two benchmark datasets.
Di Wang 0011, Yuzheng He, Yumin Tian, Lin Zhao 0003
LREC/COLING4
2024 Alignment and Multimodal Reasoning for Remote Sensing Visual Question Answering
abstract
Recently, visual question answering for remote sensing data (RSVQA) has emerged as a prominent research area in the field of remote sensing. Transformer-based approaches have demonstrated impressive results, attributed to their superior performance in jointly modeling visual and textual modalities. However, existing Remote Sensing Visual Question Answering (RSVQA) methods often overlook the modality biases present in visual-language interactions, leading to in-accuracies in answers. To address this issue, we propose a novel Transformer-based approach aimed at mitigating modality biases in RSVQA. Specifically, we introduce a contrastive learning loss to align image and text representations before cross-modal fusion, facilitating foundational learning of visual and language representations. Subsequently, we design a cross-modal decoder to comprehensively understand the correlations between images and text. Notably, in addition to predicting answers to questions, we incorporate an extra head for regression prediction of question types. Experimental results demonstrate that our approach achieves higher accuracy in answer prediction compared to state-of-the-art (SoTA) methods, establishing a new record.
Yumin Tian, Di Wang 0011, Ke Li 0024, Lin Zhao 0003
IGARSS1
2024 Visual Selection and Multistage Reasoning for RSVG
abstract
Visual grounding of remote sensing (RSVG) is a task to locate targets indicated by referring expressions in remote sensing (RS) images. Previous approaches directly concatenate visual and language features, and stack a series of transformer encoders for cross-modal fusion. However, this fusion strategy fails to fully leverage attributes and contextual information of the targets in referring expressions, limiting the performance of existing methods. To address this issue, we propose a novel visual grounding framework for RSVG, named VSMR, which achieves accurate localization by adaptively selecting target-relevant features and performing multi-stage cross-modal reasoning. Specifically, we propose an Adaptive Feature Selection (AFS) module, which automatically selects visual features relevant to queries while suppressing background noises. A Multi-Stage Decoder (MSD) is designed to iteratively infer correlations between images and queries by leveraging abundant object attributes and contextual information in the referring expressions, thereby achieving accurate target localization. Experiments demonstrate our method is superior to other state-of-the-art (SoTA) methods, achieving accuracy of 78.24%.
Yueli Ding, Di Wang 0011, Ke Li 0024, Yumin Tian
IEEE Geosci. Remote. Sens. Lett.5
2024 Part-of-speech- and syntactic-aware graph convolutional network for aspect-level sentiment classification
Yumin Tian, Ruifeng Yue, Di Wang 0011
Multim. Tools Appl.1
2024 Gist, Content, Target-Oriented: A 3-Level Human-Like Framework for Video Moment Retrieval
abstract
Video moment retrieval (VMR) aims to locate corresponding moments in an untrimmed video via a given natural language query. While most existing approaches treat this task as a cross-modal content matching or boundary prediction problem, recent studies have started to solve the VMR problem from a reading comprehension perspective. However, the cross-modal interaction processes of existing models are either insufficient or overly complex. Therefore, we reanalyze human behaviors in the document fragment location task of reading comprehension, and design a specific module for each behavior to propose a 3-level human-like moment retrieval framework (Tri-MRF). Specifically, we summarize human behaviors such as grasping the general structures of the document and the question separately, cross-scanning to mark the direct correspondences between keywords in the document and in the question, and summarizing to obtain the overall correspondences between document fragments and the question. Correspondingly, the proposed Tri-MRF model contains three modules: 1) a gist-oriented intra-modal comprehension module is used to establish contextual dependencies within each modality; 2) a content-oriented fine-grained comprehension module is used to explore direct correspondences between clips and words; and 3) a target-oriented integrated comprehension module is used to verify the overall correspondence between the candidate moments and the query. In addition, we introduce a biconnected GCN feature enhancement module to optimize query-guided moment representations. Extensive experiments conducted on three benchmarks, TACoS, ActivityNet Captions and Charades-STA demonstrate that the proposed framework outperforms State-of-the-Art methods.
Di Wang 0011, Xiantao Lu, Quan Wang 0006, Yumin Tian, Bo Wan 0002, Lihuo He
IEEE Trans. Multim.4
2024 Deep Hierarchical Multimodal Metric Learning
abstract
Multimodal metric learning aims to transform heterogeneous data into a common subspace where cross-modal similarity computing can be directly performed and has received much attention in recent years. Typically, the existing methods are designed for nonhierarchical labeled data. Such methods fail to exploit the intercategory correlations in the label hierarchy and, therefore, cannot achieve optimal performance on hierarchical labeled data. To address this problem, we propose a novel metric learning method for hierarchical labeled multimodal data, named deep hierarchical multimodal metric learning (DHMML). It learns the multilayer representations for each modality by establishing a layer-specific network corresponding to each layer in the label hierarchy. In particular, a multilayer classification mechanism is introduced to enable the layerwise representations to not only preserve the semantic similarities within each layer, but also retain the intercategory correlations across different layers. In addition, an adversarial learning mechanism is proposed to bridge the cross-modality gap by producing indistinguishable features for different modalities. Through integration of the multilayer classification and adversarial learning mechanisms, DHMML can obtain hierarchical discriminative modality-invariant representations for multimodal data. Experiments on two benchmark datasets are used to demonstrate the superiority of the proposed DHMML method over several state-of-the-art methods.
Di Wang 0011, Aqiang Ding, Yumin Tian, Quan Wang 0006, Lihuo He, Xinbo Gao 0001
IEEE Trans. Neural Networks Learn. Syst.3
2023 Enhancing CLIP-Based Text-Person Retrieval by Leveraging Negative Samples
Yumin Tian, Di Wang 0011, Bo Wan 0002
PRCV (7)1
2023 Bi-Attention enhanced representation learning for image-text matching
abstract
Image-text matching has become a research hotspot in recent years. The key point of image-text matching is to accurately measure the similarity between an image and a sentence. However, most existing methods either focus on the inter-modality similarities between regions in image and words in text or the intra-modality similarities within image regions or words, such that they cannot well exploit detailed correlations between images and texts. Furthermore, existing methods typically train their models using a triplet ranking loss, which relies on the similarity of randomly sampled triples. Since the weights of positive and negative samples are not adjusted, it cannot provide enough gradient information for training, resulting in slow convergence and limited performance. To address the above problems, we propose an image-text matching method named Bi-Attention Enhanced Representation Learning (BAERL). It builds a self-attention learning sub-network to exploit intra-modality correlations within image regions or words and a co-attention learning sub-network to exploit inter-modality correlations between image regions and words. Then, representations obtained by two sub-networks capture holistic correlations between images and texts. In addition, BAERL uses the self-similarity polynomial loss instead of triplet ranking loss to train the model. The self-similarity polynomial loss can adaptively assign appropriate weights to different pairs based on their similarity scores so as to further improve the retrieval performance . Experiments on two benchmark datasets demonstrate the superior performance of the proposed BAERL method over several state-of-the-art methods.
Yumin Tian, Aqiang Ding, Di Wang 0011, Xuemei Luo, Bo Wan 0002, Yifeng Wang 0004
Pattern Recognit.1
2023 TETFN: A text enhanced transformer fusion network for multimodal sentiment analysis
Di Wang 0011, Xutong Guo, Yumin Tian, Lihuo He, Xuemei Luo
Pattern Recognit.3
2023 Cross-Modal Enhancement Network for Multimodal Sentiment Analysis
abstract
Multimodal sentiment analysis (MSA) plays an important role in many applications, such as intelligent question-answering, computer-assisted psychotherapy and video understanding, and has attracted considerable attention in recent years. It leverages multimodal signals including verbal language, facial gestures, and acoustic behaviors to identify sentiments in videos. Language modality typically outperforms nonverbal modalities in MSA. Therefore, strengthening the significance of language in MSA will be a vital way to promote recognition accuracy. Considering that the meaning of a sentence often varies in different nonverbal contexts, combining nonverbal information with text representations is conducive to understanding the exact emotion conveyed by an utterance. In this paper, we propose a Cross-modal Enhancement Network (CENet) model to enhance text representations by integrating visual and acoustic information into a language model. Specifically, it embeds a Cross-modal Enhancement (CE) module, which enhances each word representation according to long-range emotional cues implied in unaligned nonverbal data, into a transformer-based pre-trained language model. Moreover, a feature transformation strategy is introduced for acoustic and visual modalities to reduce the distribution differences between the initial representations of verbal and nonverbal modalities, thereby facilitating the fusion of distinct modalities. Extensive experiments on benchmark datasets demonstrate the significant gains of CENet over state-of-the-art methods.
Di Wang 0011, Shuai Liu 0009, Quan Wang 0006, Yumin Tian, Lihuo He, Xinbo Gao 0001
IEEE Trans. Multim.4
2023 Hierarchical Semantic Structure Preserving Hashing for Cross-Modal Retrieval
abstract
Cross-modal hashing has become a vital technique in cross-modal retrieval due to its fast query speed and low storage cost in recent years. Generally, most of the priors supervised cross-modal hashing methods are flat methods which are designed for non-hierarchical labeled data. They treat different categories independently and ignore the inter-category correlations. In practical applications, many instances are labeled with hierarchical categories. The hierarchical label structure provides rich information among different categories. To rationally take use of category correlations, hierarchical cross-modal hashing is proposed. However, existing methods intend to preserve instance-pairwise or class-pairwise similarities, which cannot fully explore the semantic correlations among different categories and make the learned hash codes less discriminative. In this paper, we propose a deep cross-modal hashing method named hierarchical semantic structure preserving hashing (HSSPH), which directly exploits the label hierarchy information to learn discriminative hash codes. Specifically, HSSPH learns a set of class-wise hash codes for each layer. By augmenting class-wise codes with labels, it generates layer-wise prototype codes which reflect the semantic structure of each layer. In order to enhance the discriminative ability of hash codes, HSSPH supervises the hash codes learning with both labels and semantic structures to preserve the hierarchical semantics. Besides, efficient optimization algorithms are developed to directly learn the discrete hash codes for each instance and each class. Extensive experiments on two benchmark datasets show the superiority of HSSPH over several state-of-the-art methods.
Di Wang 0011, Caiping Zhang, Quan Wang 0006, Yumin Tian, Lihuo He, Lin Zhao 0003
IEEE Trans. Multim.4
2022 True wide convolutional neural network for image denoising
Gang Liu 0006, Min Dang, Jing Liu 0007, Ruotong Xiang, Yumin Tian, Nan Luo
Inf. Sci.5
2022 Pseudo-Label Guided Collective Matrix Factorization for Multiview Clustering
abstract
Multiview clustering has aroused increasing attention in recent years since real-world data are always comprised of multiple features or views. Despite the existing clustering methods having achieved promising performance, there still remain some challenges to be solved: 1) most existing methods are unscalable to large-scale datasets due to the high computational burden of eigendecomposition or graph construction and 2) most methods learn latent representations and cluster structures separately. Such a two-step learning scheme neglects the correlation between the two learning stages and may obtain a suboptimal clustering result. To address these challenges, a pseudo-label guided collective matrix factorization (PLCMF) method that jointly learns latent representations and cluster structures is proposed in this article. The proposed PLCMF first performs clustering on each view separately to obtain pseudo-labels that reflect the intraview similarities of each view. Then, it adds a pseudo-label constraint on collective matrix factorization to learn unified latent representations, which preserve the intraview and interview similarities simultaneously. Finally, it intuitively incorporates latent representation learning and cluster structure learning into a joint framework to directly obtain clustering results. Besides, the weight of each view is learned adaptively according to data distribution in the joint framework. In particular, the joint learning problem can be solved with an efficient iterative updating method with linear complexity. Extensive experiments on six benchmark datasets indicate the superiority of the proposed method over state-of-the-art multiview clustering methods in both clustering accuracy and computational efficiency.
Di Wang 0011, Songwei Han, Quan Wang 0006, Lihuo He, Yumin Tian, Xinbo Gao 0001
IEEE Trans. Cybern.5
2021 DFER-Net: Recognizing Facial Expression In The Wild
abstract
Recently, deep convolutional neural networks have made remarkable progress in facial expression recognition. However, they often fail in the natural environment, which is due to two challenging issues. One is that facial expression data in real world often has imbalanced distribution. The other is the intra-and inter-variations caused by changes in head pose, illumination and occlusions. In this paper, a discriminative facial expression recognition network (DFER-Net) is proposed to recognize facial expression in the wild by maximizing the intra-class similarity while minimizing inter-class similarity as well as strengthening the weight of minority class. Specifically, DFER-Net adds two fully-connected layers of a novel quadruplet-mean loss and a decision layer of a novel balanced-softmax loss on the traditional deep convolutional neural network. The quadruplet-mean loss enlarges intra-class similarity and inter-class distinction with high efficiency, which enhances the discriminative power of the DFER-Net. And the balanced-softmax loss strengthens the weight of minority class, which solves the class imbalance problem. Extensive experiments on benchmark datasets show the superior performance of the proposed DFER-Net over baseline methods in facial expression recognition.
Yumin Tian, Di Wang 0011
ICIP1
2021 Zero-watermarking method for resisting rotation attacks in 3D models
Gang Liu 0006, Quan Wang 0006, Lianqin Wu, Rong Pan 0004, Bo Wan 0002, Yumin Tian
Neurocomputing6
2021 kNN-based feature learning network for semantic segmentation of point cloud data
Nan Luo, Yifeng Wang 0004, Yumin Tian, Quan Wang 0006, Chuan Jing
Pattern Recognit. Lett.4
2021 Improved Bell-LaPadula Model With Break the Glass Mechanism
abstract
The Bell-LaPadula (BLP) model is a widely used access control model for the multilevel security system. The researchers proposed many modified BLP models to express privileges that cannot be expressed by the BLP model. However, these models are not compatible with the BLP model, leading to the transportation cost-prohibitive and difficult to be practically applied. In this article, an improved BLP model incorporated the break the glass (BTG) mechanism is proposed to overcome the limitations of the standard BLP and other modified BLP models. The improved model inherits some of the advantages of BTG, such as policy dynamic modification and fine-grained access control, which gives it wide availability. Additionally, in the implementation, BTG is used as an independent function attached to the original BLP; the proposed BLP model can be easily implemented in systems where BLP models have been implemented. The results of the analysis and simulations showed that the proposed BLP model improves the ability of expressing policy of BLP and achieves fine-grained access control without compromise in security. Compared with other modified BLP models, the proposed BLP model could express policy more effectively and is compatible with the original BLP model.
Runnan Zhang, Gang Liu 0006, Hongzhaoning Kang, Quan Wang 0006, Yumin Tian
IEEE Trans. Reliab.5
2020 Online Collective Matrix Factorization Hashing for Large-Scale Cross-Media Retrieval
abstract
Cross-modal hashing has been widely investigated recently for its efficiency in large-scale cross-media retrieval. However, most existing cross-modal hashing methods learn hash functions in a batch-based learning mode. Such mode is not suitable for large-scale data sets due to the large memory consumption and loses its efficiency when training streaming data. Online cross-modal hashing can deal with the above problems by learning hash model in an online learning process. However, existing online cross-modal hashing methods cannot update hash codes of old data by the newly learned model. In this paper, we propose Online Collective Matrix Factorization Hashing (OCMFH) based on collective matrix factorization hashing (CMFH), which can adaptively update hash codes of old data according to dynamic changes of hash model without accessing to old data. Specifically, it learns discriminative hash codes for streaming data by collective matrix factorization in an online optimization scheme. Unlike conventional CMFH which needs to load the entire data points into memory, the proposed OCMFH retrains hash functions only by newly arriving data points. Meanwhile, it generates hash codes of new data and updates hash codes of old data by the latest updated hash model. In such way, hash codes of new data and old data are well-matched. Furthermore, a zero mean strategy is developed to solve the mean-varying problem in the online hash learning process. Extensive experiments on three benchmark data sets demonstrate the effectiveness and efficiency of OCMFH on online cross-media retrieval.
Di Wang 0011, Quan Wang 0006, Yaqiang An, Xinbo Gao 0001, Yumin Tian
SIGIR5
2020 Policy Evaluation and Dynamic Management Based on Matching Tree for XACML
abstract
As a widely recognized policy language of access control, the eXtensible Access Control Markup Language (XACML) is widely used with its fine-grained and easy-to-read. With the application of XACML, researchers find that the XACML based policy evaluation and policy management methods can no longer meet the current large-scale requests for efficient access and dynamic management requirements. To improve the performance of policy evaluation based on XACML, we propose a policy evaluation method based on the matching tree to search policy efficiently and avoid the extra consumption of invalid policy participation. Furthermore, we propose a policy dynamic management method based on the matching tree to reduce the scale of the policy to be disabled for management, by adding locks in the tree node and the information mapping table. Through theoretical derivation and the factors that may affect its evaluation performance, we verify the improvement of evaluation efficiency. The simulation also shows the improvement of the evaluation engine based on the matching tree compared with OuenAz.
Hongzhaoning Kang, Gang Liu 0006, Quan Wang 0006, Runnan Zhang, Zichao Zhong, Yumin Tian
TrustCom6
2020 Joint and individual matrix factorization hashing for large-scale cross-modal retrieval
Di Wang 0011, Quan Wang 0006, Lihuo He, Xinbo Gao 0001, Yumin Tian
Pattern Recognit.5
2020 LPR-Net: Recognizing Chinese license plate in complex environments
Di Wang 0011, Yumin Tian, Wenhui Geng, Lin Zhao 0003, Chen Gong 0002
Pattern Recognit. Lett.2
2019 High confidence detection for moving target in aerial video
abstract
The moving target detection and tracking in aerial video is a challenge task because of its moving background, smaller target sizes, lower resolution and limited onboard computing resources. In this study, a high confidence detection method based on background compensation and three‐frame‐difference method is designed, which can detect moving objects in a dynamic background accurately. First, the authors use local feature extraction and matching for image registration and demonstrate that speed‐up robust feature key points are suitable for the stabilisation task. Then, they estimate the global camera motion parameters using affine transformation which are obtained by the random sample consensus algorithm. Finally, they detect moving object by three‐frame‐difference method. As the detection results of the frame‐difference method generally exists ‘empty’ and noise, in order to select the two higher‐quality differential images to perform the logic AND operation, they add image quality assessment to the three‐frame‐difference method to obtain more accurate moving objects. Moreover, the edge detection algorithm and morphological processing are integrated together to further boost the overall detecting performance. The extensive empirical evaluations on aerial videos demonstrate that the proposed detector is very promising for the various challenging scenarios.
Yumin Tian, Chenhui Peng, Di Wang 0011, Bo Wan 0002
IET Image Process.1
2019 Robust joint learning network: improved deep representation learning for person re-identification
Yumin Tian, Di Wang 0011, Bo Wan 0002
Multim. Tools Appl.1
2016 Surveillance video synopsis generation method via keeping important relationship among objects
abstract
To reduce the efforts of human in browsing long surveillance videos, synopsis videos are proposed. Traditional synopsis video generation methods condense most of the activities in the video by simultaneously showing several actions, even when they originally occurred at different times. This inevitably causes ignorance of temporal relationship among objects. For example, two persons walk shoulder to shoulder and they are detected and tracked separately, but in the synopsis they never ‘met’. In this study, a trajectory mapping model is defined, whose energy function includes not only the cost caused by the synopsis video, but that of the original video. In this way, it tries to make the relationship between objects of the original video consistent with that of the synopsis. Finally, the video synopsis is generated by an energy minimisation method. Experiments show that the proposed video synopsis can reduce the spatiotemporal redundancies of the input video as much as possible. Moreover, it can keep the important relationship between objects and maintain the time consistency of important activities.
Yumin Tian, Haihong Zheng, Qichao Chen, Risan Lin
IET Comput. Vis.1
2008 Adaptive video watermarking algorithm based on MPEG-4 streams
abstract
Video digital watermarking is a topic of the current watermarking research. To better ensure the robustness and invisibility, an adaptive video watermarking algorithm is proposed based on MPEG-4 video compression principle and the human visual system model. The adaptive factor is designed according to the direct current coefficient and the number of low and intermediate frequency coefficients after discrete cosine transform (DCT) of I frame. For different image blocks, the different embedded intensity is used. The watermarking is embedded into the low frequency coefficients. Experiments show that the algorithm works well with the compatibility of human visual system, and the common attack is very robust as well.
Yumin Tian, Yun-hui Qu
ICARCV2