Jun Yu 0011

dblp:50/5754-11 · DBLP profile ↗
← Back
22ranked-venue papers
7as first author
17since 2021 · last 2026
0000-0001-5711-3696ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 6 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 EduYOLO: A classroom behavior recognition framework based on high-resolution feature attention fusion
Jun Yu 0011, Shengzhao Li, Huijie Liu 0001, Qi Liu 0003, Zhiyuan Cheng 0014
Expert Syst. Appl.1
2025 Text-dominant multimodal perception network for sentiment analysis based on cross-modal semantic enhancements
Zuhe Li, Panbo Liu, Yushan Pan, Jun Yu 0011, Haoran Chen 0004, Hao Wang 0076
Appl. Intell.4
2025 Multimodal sentiment analysis based on disentangled representation learning and cross-modal-context association mining
Zuhe Li, Panbo Liu, Yushan Pan, Weiping Ding 0001, Jun Yu 0011, Haoran Chen 0004, Hao Wang 0003
Neurocomputing5
2025 Representation distribution matching and dynamic routing interaction for multimodal sentiment analysis
Zuhe Li, Zhenwei Huang, Xiaojiang He, Jun Yu 0011, Haoran Chen 0004, Chenguang Yang 0001, Yushan Pan
Knowl. Based Syst.4
2024 Hierarchical denoising representation disentanglement and dual-channel cross-modal-context interaction for multimodal sentiment analysis
Zuhe Li, Zhenwei Huang, Yushan Pan, Jun Yu 0011, Haoran Chen 0004, Di Wu 0035, Hao Wang 0003
Expert Syst. Appl.4
2024 Zero-shot discrete hashing with adaptive class correlation for cross-modal retrieval
Kailing Yong, Zhenqiu Shu, Jun Yu 0011, Zhengtao Yu 0001
Knowl. Based Syst.3
2024 Proxy-Based Graph Convolutional Hashing for Cross-Modal Retrieval
abstract
Cross-modal hashing retrieval approaches have received extensive attention owing to their storage superiority and retrieval efficiency. To achieve better retrieval performances, hashing methods seek to embed more semantic information of multi-modal data into hash codes. Existing deep cross-modal hashing methods typically learn hash functions from the similarity of paired data to generate hash codes. However, such locally-oriented learning methods often suffer from low efficiency and incomplete acquisition of semantic information. To address these challenges, this paper presents a novel deep hashing approach, called Proxy-based Graph Convolutional Hashing (PGCH), for cross-modal retrieval. Specifically, we use global similarity to construct proxy hash codes for two different modalities. This strategy of these proxy hash codes ensures that they include data points with significant distribution differences. It helps to match data from different modalities to different proxy hash codes, which can capture the global similarity of multi-modal hash codes and improve the efficiency of hash code learning. Subsequently, we employ a multi-modal contrastive loss to learn the global similarity. Furthermore, by constructing a proxy hash matrix from the proxy hash codes, we apply graph convolution to efficiently narrow the gap between different modalities, leading to a substantial improvement in retrieval performance for cross-modal retrieval tasks. The comprehensive experiments on four benchmark multimedia datasets demonstrate that our PGCH approach achieves better retrieval performances than a bundle of state-of-the-art hashing approaches.
Yibing Bai, Zhenqiu Shu, Jun Yu 0011, Zhengtao Yu 0001, Xiaojun Wu 0001
IEEE Trans. Big Data3
2023 Actor-Multi-Scale Context Bidirectional Higher Order Interactive Relation Network for Spatial-Temporal Action Localization
abstract
The key to video action detection lies in the understanding of interaction between persons and background objects in a video. Current methods usually employ object detectors to extract objects directly or use grid features to represent objects in the environment, which underestimate the great potential of multi-scale context information (e.g., objects and scenes of different sizes). How to exactly represent the multi-scale context and make full utilization of it still remains an unresolved challenge for spatial-temporal action localization. In this paper, we propose a novel Actor-Multi-Scale Context Bidirectional Higher Order Interactive Relation Network (AMCRNet) that extracts multi-scale context through multiple pooling layers with different sizes. Specifically, we develop an Interactive Relation Extraction module to model the higher-order relation between the target person and the context (e.g., other persons and objects). Along this line, we further propose a History Feature Bank and Interaction method to achieve better performance by modeling such relation across continuing video clips. Extensive experimental results on AVA2.2 and UCF101-24 demonstrate the superiority and rationality of our proposed AMCRNet.
Jun Yu 0011, Yingshuai Zheng, Shulan Ruan, Qi Liu 0003, Zhiyuan Cheng 0014
IJCAI1
2023 Online supervised collective matrix factorization hashing for cross-modal retrieval
Zhenqiu Shu, Jun Yu 0011, Donglin Zhang 0001, Zhengtao Yu 0001, Xiaojun Wu 0001
Appl. Intell.3
2023 Robust supervised matrix factorization hashing with application to cross-modal retrieval
Zhenqiu Shu, Kailing Yong, Donglin Zhang 0001, Jun Yu 0011, Zhengtao Yu 0001, Xiaojun Wu 0001
Neural Comput. Appl.4
2022 CPEE: Civil Case Judgment Prediction centering on the Trial Mode of Essential Elements
abstract
Civil Case Judgment Prediction (CCJP) is a fundamental task in the legal intelligence of the civil law system, which aims to automatically predict the judgment results on each plea of the plaintiff. Existing studies mainly focus on making judgment predictions only on a certain civil cause (e.g., the divorce dispute) by utilizing the fact descriptions and pleas of the plaintiff, which still suffer from the various causes and complicated legal essential elements in the real court. Thus, in this paper, we formalize CCJP as a multi-task learning problem and propose a CCJP method centering on the trial mode of essential elements, CPEE, which explores the practical judicial process and analyzes comprehensive legal essential elements to make judgment predictions. Specifically, we first construct three tasks (i.e., the predictions on the civil causes, law articles, and the final judgment on each plea) necessary for CCJP, that follow the judgment process and exploit the results of intermediate subtasks to make judgment predictions. Then we design a logic-enhanced network to predict the results of three tasks and conduct a comprehensive study of civil cases. Finally, owing to the interlinked and dependent relationships among each task, we adopt the cause prediction result to help predict law articles and incorporate them into final judgment prediction through a gate mechanism. Furthermore, since the existing dataset fails to provide sufficient case information, we construct a real-world CCJP dataset that contains various causes and comprehensive legal elements. Extensive experimental results on the dataset validate the effectiveness of our method.
Lili Zhao 0002, Linan Yue, Yanqing An, Yuren Zhang, Jun Yu 0011, Qi Liu 0003, Enhong Chen
CIKM5
2022 Adaptive multi-modal fusion hashing via Hadamard matrix
Jun Yu 0011, Donglin Zhang 0001, Zhenqiu Shu
Appl. Intell.1
2022 Discrete asymmetric zero-shot hashing with application to cross-modal retrieval
Zhenqiu Shu, Kailing Yong, Jun Yu 0011, Shengxiang Gao, Cunli Mao, Zhengtao Yu 0001
Neurocomputing3
2022 Specific class center guided deep hashing for cross-modal retrieval
Zhenqiu Shu, Yibing Bai, Donglin Zhang 0001, Jun Yu 0011, Zhengtao Yu 0001, Xiaojun Wu 0001
Inf. Sci.4
2021 Discrete Bidirectional Matrix Factorization Hashing for Zero-Shot Cross-Media Retrieval
Donglin Zhang 0001, Xiaojun Wu 0001, Jun Yu 0011
PRCV (2)3
2021 Learning latent hash codes with discriminative structure preserving for cross-modal retrieval
Donglin Zhang 0001, Xiaojun Wu 0001, Jun Yu 0011
Pattern Anal. Appl.3
2021 Label Consistent Flexible Matrix Factorization Hashing for Efficient Cross-modal Retrieval
abstract
Hashing methods have sparked a great revolution on large-scale cross-media search due to its effectiveness and efficiency. Most existing approaches learn unified hash representation in a common Hamming space to represent all multimodal data. However, the unified hash codes may not characterize the cross-modal data discriminatively, because the data may vary greatly due to its different dimensionalities, physical properties, and statistical information. In addition, most existing supervised cross-modal algorithms preserve the similarity relationship by constructing an n × n pairwise similarity matrix, which requires a large amount of calculation and loses the category information. To mitigate these issues, a novel cross-media hashing approach is proposed in this article, dubbed label flexible matrix factorization hashing (LFMH). Specifically, LFMH jointly learns the modality-specific latent subspace with similar semantic by the flexible matrix factorization. In addition, LFMH guides the hash learning by utilizing the semantic labels directly instead of the large n × n pairwise similarity matrix. LFMH transforms the heterogeneous data into modality-specific latent semantic representation. Therefore, we can obtain the hash codes by quantifying the representations, and the learned hash codes are consistent with the supervised labels of multimodal data. Then, we can obtain the similar binary codes of the corresponding modality, and the binary codes can characterize such samples flexibly. Accordingly, the derived hash codes have more discriminative power for single-modal and cross-modal retrieval tasks. Extensive experiments on eight different databases demonstrate that our model outperforms some competitive approaches.
Donglin Zhang 0001, Xiaojun Wu 0001, Jun Yu 0011
ACM Trans. Multim. Comput. Commun. Appl.3
2020 Fast Discrete Cross-Modal Hashing Based on Label Relaxation and Matrix Factorization
abstract
In recent years, cross-media retrieval has drawn considerable attention due to the exponential growth of multimedia data. Many hashing approaches have been proposed for the cross-media search task. However, there are still open problems that warrant investigation. For example, most existing supervised hashing approaches employ a binary label matrix, which achieves small margins between wrong labels (0) and true labels (1). This may affect the retrieval performance by generating many false negatives and false positives. In addition, some methods adopt a relaxation scheme to solve the binary constraints, which may cause large quantization errors. There are also some discrete hashing methods that have been presented, but most of them are time-consuming. To conquer these problems, we present a label relaxation and discrete matrix factorization method (LRMF) for cross-modal retrieval. It offers a number of innovations. First of all, the proposed approach employs a novel label relaxation scheme to control the margins adaptively, which has the benefit of reducing the quantization error. Second, by virtue of the proposed discrete matrix factorization method designed to learn the binary codes, large quantization errors caused by relaxation can be avoided. The experimental results obtained on two widely-used databases demonstrate that LRMF outperforms state-of-the-art cross- media methods.
Donglin Zhang 0001, Xiaojun Wu 0001, Zhen Liu 0015, Jun Yu 0011, Josef Kittler
ICPR4
2020 Cross-modal subspace learning via kernel correlation maximization and discriminative structure-preserving
Jun Yu 0011, Xiaojun Wu 0001
Multim. Tools Appl.1
2020 Learning discriminative hashing codes for cross-modal retrieval based on multi-view features
Jun Yu 0011, Xiaojun Wu 0001, Josef Kittler
Pattern Anal. Appl.1
2019 Discriminative Supervised Hashing for Cross-Modal Similarity Search
Jun Yu 0011, Xiaojun Wu 0001, Josef Kittler
Image Vis. Comput.1
2018 Semi-supervised Hashing for Semi-Paired Cross-View Retrieval
abstract
Recently, hashing techniques have gained importance in large-scale retrieval tasks because of their retrieval speed. Most of the existing cross-view frameworks assume that data are well paired. However, the fully-paired multiview situation is not universal in real applications. The aim of the method proposed in this paper is to learn the hashing function for semi-paired cross-view retrieval tasks. To utilize the label information of partial data, we propose a semi-supervised hashing learning framework which jointly performs feature extraction and classifier learning. The experimental results on two datasets show that our method outperforms several state-of-the-art methods in terms of retrieval accuracy.
Jun Yu 0011, Xiaojun Wu 0001, Josef Kittler
ICPR1