Lina Wei

dblp:160/6241 · DBLP profile ↗
← Back
23ranked-venue papers
7as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 The Generalized Champollion Semantic System with the Logical Order Operator for the Interpretation of Conjoined Sentence
abstract
In the field of artificial intelligence, accurately analyzing sentence semantics is crucial for many tasks, particularly in natural language processing. While Davidson event semantics offers an effective method for analyzing sentence semantics, it faces challenges when dealing with ambiguities in sentences combined with the logical word “and” , such as “John caught and ate a fish” . The ambiguity arises because it is unclear whether John ate the same fish he caught or a different one, due to the commutativity of “and” . Traditional event semantics are inadequate for addressing these ambiguities because they lack a consideration of event order. In this paper, we generalize the Champollion semantic system by the logical order operator, which resolves conflicts stemming from the commutativity of conjunctions by incorporating the logical order of events. We apply the generalized Champollion semantic system to existing research to demonstrate its potential for improving the accuracy of event semantics. Additionally, the generalized Champollion semantic system can help refine the interpretation of atomic events in sentences with “for-adverbial” clauses by preventing concurrent occurrences.
Guangjian Huang, Lina Wei, Shahbaz Hassan Wasti
ICIC2
2026 Uncertainty-constrained fusion of single-view and multi-view depth estimation for AR virtual-real occlusion
Shuai Ding 0001, Yongze Li, Lina Wei, Dapeng Chen
Neural Networks5
2025 Generalized Video Moment Retrieval
abstract
In this paper, we introduce the Generalized Video Moment Retrieval (GVMR) framework, which extends traditional Video Moment Retrieval (VMR) to handle a wider range of query types. Unlike conventional VMR systems, which are often limited to simple, single-target queries, GVMR accommodates both non-target and multi-target queries. To support this expanded task, we present the NExT-VMR dataset, derived from the YFCC100M collection, featuring diverse query scenarios to enable more robust model evaluation. Additionally, we propose BCANet, a transformer-based model incorporating the novel Boundary-aware Cross Attention (BCA) module. The BCA module enhances boundary detection and uses cross-attention to achieve a comprehensive understanding of video content in relation to queries. BCANet accurately predicts temporal video segments based on natural language descriptions, outperforming traditional models in both accuracy and adaptability. Our results demonstrate the potential of the GVMR framework, the NExT-VMR dataset, and BCANet to advance VMR systems, setting a new standard for future multimedia information retrieval research.
You Qin, Yicong Li 0004, Wei Ji 0008, Li Li 0091, Pengcheng Cai, Lina Wei, Roger Zimmermann
ICLR7
2025 MamFusion: Multi-Mamba with Temporal Fusion for Partially Relevant Video Retrieval
abstract
Partially Relevant Video Retrieval (PRVR) is a challenging task in the domain of multimedia retrieval. It is designed to identify and retrieve untrimmed videos that are partially relevant to the provided query. In this work, we investigate long-sequence video content understanding to address information redundancy issues. Leveraging the outstanding long-term state space modeling capability and linear scalability of the Mamba module, we introduce a multi-Mamba module with temporal fusion framework (MamFusion) tailored for PRVR task. This framework effectively captures the state-relatedness in long-term video content and seamlessly integrates it into text-video relevance understanding, thereby enhancing the retrieval process. Specifically, we introduce Temporal T-to-V Fusion and Temporal V-to-T Fusion to explicitly model temporal relationships between text queries and video moments, improving contextual awareness and retrieval accuracy. Extensive experiments conducted on large-scale datasets demonstrate that MamFusion achieves state-of-the-art performance in retrieval effectiveness. Code is available at the link: https://github.com/Vision-Multimodal-Lab-HZCU/MamFusion.
Xinru Ying, Jiaqi Mo, Canghong Jin, Lina Wei
ICME6
2025 Few-Shot Incremental Multi-modal Learning via Touch Guidance and Imaginary Vision Synthesis
abstract
Multimodal perception, which integrates vision and touch, is increasingly demonstrating its significance in domains such as embodied intelligence and human-computer interaction. However, in open-world scenarios, multimodal data streams face significant challenges, including catastrophic forgetting and overfitting, during few-shot class incremental learning (FSCIL), leading to a severe degradation in model performance. In this work, we propose a novel approach named Few-Shot Incremental Multi-modal Learning via Touch Guidance and Imaginary Vision Synthesis (TIFS). Our method leverages vision imagination synthesis to enhance the semantic understanding and integrates touch and vision fusion to improve the problem of modal imbalance. Specifically, we introduce a framework that employs touch-guided vision information for cross-modal contrastive learning to address the challenges of few-shot learning. Additionally, we incorporate multiple learning mechanisms, including regularization, memory mechanisms, and attention mechanisms, to mitigate catastrophic forgetting during multi-incremental step learning. Experimental results on the Touch and Go and VisGel datasets demonstrate that the TIFS framework exhibits robust continuous learning capabilities and strong generalization performance in touch-vision few-shot incremental learning tasks. Our code is available at https://github.com/Vision-Multimodal-Lab-HZCU/TIFS.
Lina Wei, Zhongsheng Lin, Canghong Jin, Hanbin Zhao, Dapeng Chen
IJCAI1
2025 Tactile-Visual Class-Continual Learning via Temporal Attention Interaction
abstract
Current high-performance multi-modal models are predominantly trained in a static learning scenario, where the model undergoes a single joint training on the entire dataset. However, in real-world applications, data are dynamically generated, and tasks evolve continuously, necessitating that multi-modal models adapt to a continual learning scenario. The core challenge in continuous learning is catastrophic forgetting problem, where training data from previous tasks are often not fully retained. As a result, when the model learns new tasks, it can only utilize the training data associated with these new tasks, leading to a significant degradation in performance on previously learned tasks. Specifically, due to the costly process of collecting touch data and the low standardization of sensor outputs, tactile-visual continual learning poses significant challenges. To effectively integrate tactile and visual information and mitigate catastrophic forgetting, this paper proposes the Tactile-Visual Class-Continual Learning via Temporal Attention Interaction (TV-CCL) model. TV-CCL maintains the instance and class-level semantic similarity between tactile and visual modalities through Tactile-Visual Multi-modal Semantic Alignment (TV-MSA). Additionally, we incorporate Tactile-Guided Visual Attention Distillation (TG-VAD) to preserve previously learned haptic-guided visual attention capabilities. Our experiments on the Touch and Go dataset demonstrate that TV-CCL significantly outperforms existing CCL methods when leveraging combined haptic and visual information. Our code is available at https://github.com/Vision-Multimodal-Lab-HZCU/TV-CCL.
Lina Wei, Zhongsheng Lin, Xinru Ying, Canghong Jin
IJCNN1
2025 Service Area Vehicle Flow Prediction Model for Highway Service Areas Based on Gravity Model Quadratic Assignment
Lai Meng, Yichu Dai, Zhengdong Fei, Canghong Jin, Lina Wei
KSEM (4)6
2025 RCD-DETR: A Lightweight Real-Time Detection Transformer for Conveyor Belt Egg Detection
abstract
Precise detection of eggs on conveyor belts in industrial automated production lines is of significant importance for reducing detection error rates and improving production efficiency. However, existing detection models face considerable challenges in detection accuracy and real-time processing when handling practical scenarios such as densely arranged eggs, complex background interference, and high-speed conveyor belts. This paper proposes a lightweight real-time detection Transformer model called RCD-DETR, consisting of three key innovative components: (1) a lightweight Reparam Context-aware Network (RCNet) that effectively balances feature extraction capability and computational efficiency; (2) a Context-Sensitive Refinement Feature Pyramid Network (CSRFPN) with enhanced contextual awareness and multi-scale feature representation capabilities, strengthening the model’s ability to recognize small and dense objects; and (3) a Dilated Reparam Bottleneck C3 (DRepC3) module that expands the receptive field range while further reducing computational resource requirements. Experimental evaluation indicates that, compared to the RT-DETR baseline model, the proposed method achieves improved detection accuracy while reducing computational complexity by 54.1%, decreasing parameter count by 51.3%, and compressing model size by 50.8%. The method achieves an excellent balance between accuracy, inference speed, and resource consumption, making it suitable for deployment on resource-constrained edge computing devices for precise real-time egg detection on conveyor belts.
Yidong Shen, Miaoyang Dai, Lina Wei
SMC5
2024 Panoptic Scene Graph Generation with Semantics-Prototype Learning
abstract
Panoptic Scene Graph Generation (PSG) parses objects and predicts their relationships (predicate) to connect human language and visual scenes. However, different language preferences of annotators and semantic overlaps between predicates lead to biased predicate annotations in the dataset, i.e. different predicates for the same object pairs. Biased predicate annotations make PSG models struggle in constructing a clear decision plane among predicates, which greatly hinders the real application of PSG models. To address the intrinsic bias above, we propose a novel framework named ADTrans to adaptively transfer biased predicate annotations to informative and unified ones. To promise consistency and accuracy during the transfer process, we propose to observe the invariance degree of representations in each predicate class, and learn unbiased prototypes of predicates with different intensities. Meanwhile, we continuously measure the distribution changes between each presentation and its prototype, and constantly screen potentially biased data. Finally, with the unbiased predicate-prototype representation embedding space, biased annotations are easily identified. Experiments show that ADTrans significantly improves the performance of benchmark models, achieving a new state-of-the-art performance, and shows great generalization and effectiveness on multiple datasets. Our code is released at https://github.com/lili0415/PSG-biased-annotation.
Li Li 0091, Wei Ji 0008, Yiming Wu 0005, Mengze Li 0001, You Qin, Lina Wei, Roger Zimmermann
AAAI6
2024 Soften to Defend: Towards Adversarial Robustness via Self-Guided Label Refinement
abstract
Adversarial training (AT) is currently one of the most effective ways to obtain the robustness of deep neural networks against adversarial attacks. However, most AT methods suffer from robust overfitting, i.e., a significant generalization gap in adversarial robustness between the training and testing curves. In this paper, we first identify a connection between robust overfitting and the excessive memorization of noisy labels in AT from a view of gradient norm. As such label noise is mainly caused by a distribution mismatch and improper label assignments, we are motivated to propose a label refinement approach for AT. Specifically, our Self-Guided Label Refinement first self-refines a more accurate and informative label distribution from over-confident hard labels, and then it calibrates the training by dynamically incorporating knowledge from self-distilled models into the current model and thus requiring no external teachers. Empirical results demonstrate that our method can simultaneously boost the standard accuracy and robust performance across multiple benchmark datasets, attack types, and architectures. In addition, we also provide a set of analyses from the perspectives of information theory to dive into our method and suggest the importance of soft labels for robust generalization.
Zhuorong Li, Daiwei Yu, Lina Wei, Canghong Jin, Yun Zhang 0011
CVPR3
2024 Subgroup total perfect codes in Cayley sum graphs
Lina Wei, Shoujun Xu, Sanming Zhou
Des. Codes Cryptogr.2
2023 A note on characterization of the induced matching extendable Cayley graphs generated by transpositions
Yong-De Feng, Yan-Ting Xie, Lina Wei, Shoujun Xu
Discret. Appl. Math.3
2023 Semi-Supervised Entity Alignment via Relation-Based Adaptive Neighborhood Matching
abstract
Many recent studies of Entity Alignment (EA) use Graph Neural Networks (GNNs) to aggregate the neighborhood features of entities and achieve better performance. However, aligned entities in real Knowledge Graphs (KGs) usually have non-isomorphic neighborhood structures due to the different data sources of KGs. Therefore, it is insufficient to simply compare the global direct neighborhood of aligned entities, which may also become a variable for the EA judgment. In this paper, we propose a Relation-based Adaptive Neighborhood Matching method (RANM), which matches larger range and higher confidence neighborhoods for aligned entities based on relation matching instead of alignment seeds.RANMfirst uses alignment seeds to construct the best relation matching set, and then performs local direct neighborhood matching and feature aggregation on the candidate alignments. To obtain high-quality entity embeddings, we design a variant attention mechanism based on heterogeneous graphs, which considers the heterogeneity of relations in KGs. We also adopt a bi-directional iterative co-training to further improve the performance. Extensive experiments on three well-known datasets show our method significantly outperforms 14 state-of-the-art methods, and is 3.01-11.5% higher than the best-performing baselines in [email protected] shows high performance on the long-tailed entities and the dataset with less alignment seeds.
Weishan Cai, Wenjun Ma, Lina Wei, Yuncheng Jiang 0001
IEEE Trans. Knowl. Data Eng.3
2021 Generalized fuzzy automata with semantic computing
Lina Wei, Guangjian Huang, Shahbaz Hassan Wasti, Muhammad Jawad Hussain
Soft Comput.1
2021 End-to-End Video Saliency Detection via a Deep Contextual Spatiotemporal Network
abstract
As an interesting and important problem in computer vision, learning-based video saliency detection aims to discover the visually interesting regions in a video sequence. Capturing the information within frame and between frame at different aspects (such as spatial contexts, motion information, temporal consistency across frames, and multiscale representation) is important for this task. A key issue is how to jointly model all these factors within a unified data-driven scheme in an end-to-end fashion. In this article, we propose an end-to-end spatiotemporal deep video saliency detection approach, which captures the information on spatial contexts and motion characteristics. Furthermore, it encodes the temporal consistency information across the consecutive frames by implementing a convolutional long short-term memory (Conv-LSTM) model. In addition, the multiscale saliency properties for each frame are adaptively integrated for final saliency prediction in a collaborative feature-pyramid way. Finally, the proposed deep learning approach unifies all the aforementioned parts into an end-to-end joint deep learning scheme. Experimental results demonstrate the effectiveness of our approach in comparison with the state-of-the-art approaches.
Lina Wei, Shanshan Zhao 0001, Omar El Farouk Bourahla, Xi Li 0001, Fei Wu 0001, Yueting Zhuang, Junwei Han 0001, Mingliang Xu 0001
IEEE Trans. Neural Networks Learn. Syst.1
2020 An approach for measuring semantic similarity between Wikipedia concepts using multiple inheritances
Muhammad Jawad Hussain, Shahbaz Hassan Wasti, Guangjian Huang, Lina Wei, Yong Tang 0001
Inf. Process. Manag.4
2020 Context-Aware Graph Label Propagation Network for Saliency Detection
abstract
Recently, a large number of existing methods for saliency detection have mainly focused on designing complex network architectures to aggregate powerful features from backbone networks. However, contextual information is not well utilized, which often causes false background regions and blurred object boundaries. Motivated by these issues, we propose an easyto-implement module that utilizes the edge-preserving ability of superpixels and the graph neural network to interact the context of superpixel nodes. In more detail, we first extract the features from the backbone network and obtain the superpixel information of images. This step is followed by superpixel pooling in which we transfer the irregular superpixel information to a structured feature representation. To propagate the information among the foreground and background regions, we use a graph neural network and self-attention layer to better evaluate the degree of saliency degree. Additionally, an affinity loss is proposed to regularize the affinity matrix to constrain the propagation path. Moreover, we extend our module to a multiscale structure with different numbers of superpixels. Experiments on five challenging datasets show that our approach can improve the performance of three baseline methods in terms of some popular evaluation metrics.
Wei Ji 0008, Xi Li 0001, Lina Wei, Fei Wu 0001, Yueting Zhuang
IEEE Trans. Image Process.3
2019 Deep Group-Wise Fully Convolutional Network for Co-Saliency Detection With Graph Propagation
abstract
A key problem in co-saliency detection is how to effectively model the interactive relationship of a whole image group and the individual perspective of each image in a united data-driven manner. In this paper, we propose a group-wise deep co-saliency detection approach to address the co-saliency object discovery problem based on the fully convolutional network (FCN). The proposed approach captures the group-wise interaction information for group images by learning a semantics-aware image representation based on a convolutional neural network, which adaptively learns the group-wise features for co-saliency detection. Furthermore, the proposed approach discovers the collaborative and interactive relationships between group-wise feature representation and single image individual feature representation, and model this in a collaborative learning framework. Then, we set up a unified deep learning scheme to jointly optimize the process of group-wise feature representation learning and the collaborative learning, leading to more reliable and robust co-saliency detection results. Finally, we present a graph Laplacian regularized nonlinear regression model for saliency refinement. Experimental results demonstrate the effectiveness of our approach in comparison with the state-of-the-art approaches.
Lina Wei, Shanshan Zhao 0001, Omar El Farouk Bourahla, Xi Li 0001, Fei Wu 0001, Yueting Zhuang
IEEE Trans. Image Process.1
2017 Graph-theoretic spatiotemporal context modeling for video saliency detection
abstract
As an important and challenging problem in computer vision, video saliency detection is typically cast as a spatiotemporal context modeling problem over consecutive frames. As a result, a key issue in video saliency detection is how to effectively capture the intrinsical properties of atomic video structures as well as their associated contextual interactions along the spatial and temporal dimensions. Motivated by this observation, we propose a graph-theoretic video saliency detection approach based on adaptive video structure discovery, which is carried out within a spatiotemporal atomic graph. Through graph-based manifold propagation, the proposed approach is capable of effectively modeling the semantically contextual interactions among atomic video structures for saliency detection while preserving spatial smoothness and temporal consistency. Experiments demonstrate the effectiveness of the proposed approach over several benchmark datasets.
Lina Wei, Xi Li 0001, Fei Wu 0001, Jun Xiao 0001
ICIP1
2017 Group-wise Deep Co-saliency Detection
abstract
In this paper, we propose an end-to-end group-wise deep co-saliency detection approach to address the co-salient object discovery problem based on the fully convolutional network (FCN) with group input and group output. The proposed approach captures the group-wise interaction information for group images by learning a semantics-aware image representation based on a convolutional neural network, which adaptively learns the group-wise features for co-saliency detection. Furthermore, the proposed approach discovers the collaborative and interactive relationships between group-wise feature representation and single-image individual feature representation, and model this in a collaborative learning framework. Finally, we set up a unified end-to-end deep learning scheme to jointly optimize the process of group-wise feature representation learning and the collaborative learning, leading to more reliable and robust co-saliency detection results. Experimental results demonstrate the effectiveness of our approach in comparison with the state-of-the-art approaches.
Lina Wei, Shanshan Zhao 0001, Omar El Farouk Bourahla, Xi Li 0001, Fei Wu 0001
IJCAI1
2016 Recognition of infant's emotions and needs from speech signals
abstract
Speech is not only a way for infants under one year of age to communicate with the outside world, but also the important information source to reflect their emotions and needs, as well as health status and mental level. In order to explore the intelligent machine technology for understanding infant's emotions and needs from speech signals, and therefore help parents in child rearing, this paper studied the signal processing and feature extraction of the above speeches. It seems that a high accuracy and reliable result couldn't be reached from the infant's speech signals only in dealing with multi kinds of emotions and needs. So an effective recognition approach considering the combined features of acoustic characteristics and rearing behaviors was proposed based on the self-adapting algorithm. Experiment results showed that most common emotions and needs of infants, such as happy, hungry, and sleepy states which reflect their typically physiological and psychological status in daily life, can be recognized correctly at a relatively high accuracy.
Hongzhi Hu, Lina Wei, Weihui Dai, Huajuan Mao
SMC3
2016 Optimal time allocation for multi-antenna wireless powered heterogeneous sensor network communications under imperfect CSI
Feng Zhao 0002, Lina Wei, Hongbin Chen 0001
Signal Process.2
2016 DeepSaliency: Multi-Task Deep Neural Network Model for Salient Object Detection
abstract
A key problem in salient object detection is how to effectively model the semantic properties of salient objects in a data-driven manner. In this paper, we propose a multi-task deep saliency model based on a fully convolutional neural network with global input (whole raw images) and global output (whole saliency maps). In principle, the proposed saliency model takes a data-driven strategy for encoding the underlying saliency prior information, and then sets up a multi-task learning scheme for exploring the intrinsic correlations between saliency detection and semantic image segmentation. Through collaborative feature learning from such two correlated tasks, the shared fully convolutional layers produce effective features for object perception. Moreover, it is capable of capturing the semantic information on salient objects across different levels using the fully convolutional layers, which investigate the feature-sharing properties of salient object detection with a great reduction of feature redundancy. Finally, we present a graph Laplacian regularized nonlinear regression model for saliency refinement. Experimental results demonstrate the effectiveness of our approach in comparison with the state-of-the-art approaches.
Xi Li 0001, Lina Wei, Ming-Hsuan Yang 0001, Fei Wu 0001, Yueting Zhuang, Haibin Ling, Jingdong Wang 0001
IEEE Trans. Image Process.3