EDBT 2026 Demo / reviewers in the wild / expert
Jinmeng Wu
dblp:264/4428
· DBLP profile ↗
19ranked-venue papers
8as first author
16since 2021 · last 2026
0000-0002-0264-8025ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 5 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PointTFA$^{m}$: Multi-Modal, Training-Free Adaptation for Point Cloud Understanding
Jinmeng Wu, Youxiang Hu, Hao Zhang 0047, Basura Fernando, Yanbin Hao, Hanyu Hong |
IEEE Trans. Multim. | 1 |
| 2025 | Improving Open-vocabulary Video Visual Relation Detection with Decomposed Prompt Learning and Relation AdjustmentabstractOpen-vocabulary video visual relation detection (VidVRD) expands the scope of detecting object relations in videos to include unseen categories. It marks considerable advancement in recognizing novel relations solely by training on a base set, thus extending the frontiers of automated video understanding. However, the performance of current methods on novel predicates remains significantly inferior to that on base categories. We attribute this discrepancy to two primary factors: (1) A significant task misalignment between the Visual Relation Detection (VRD) task and the pre-trained models’ visual feature extractors, which are often designed for tasks like video-text retrieval and image-text retrieval, resulting in poor generalization to the novel set. (2) The relatively small size and limited vocabulary of open-vocabulary datasets, which create a substantial gap between base and novel predicates. Consequently, text prompts trained on the base set fail to generalize effectively to the novel set. To address these issues, we propose two improvement measures: (1) We decompose base and novel relations into actional and spatial patterns and introduce an innovative text prompt learning method that leverages the shared patterns between base and novel relations. (2) We develop a relation probability adjustment mechanism that utilizes reliable base relation predictions to adjust the probabilities of relations in novel classes by considering their overlaps in either actional or spatial contents. Experimental results on the benchmark dataset demonstrate significant performance improvements. Ming Pei, Yi Tan 0001, Yanbin Hao, Hao Zhang 0047, Jinmeng Wu, Basura Fernando, Xun Yang 0001 |
ICASSP | 5 |
| 2024 | Generative Adversarial Network-Based Jitter Distortion Correction for High Resolution Spaceborne ImagesabstractThis paper presents a Generative Adversarial Network (GAN)-based jitter distortion correction method for spaceborne images of Time Delay Integration (TDI) Charge-Coupled Device (CCD) camera. This method leverages the advantages of GANs and combines content loss, adversarial loss, and perceptual loss to effectively repair distorted images while preserving image details, which does not rely on jitter information captured by high-frequency attitude sensors, nor depends on the analysis of overlapping areas between different bands in multispectral images. The experimental results show that the proposed method achieves automated correction of geometric distortions and has shown promising restoration results on real distorted images captured by Yaogan-26 satellite and GaoFen satellite, which achieves better results than other blind restoration methods. Ying Zhu 0002, Lei Wang 0068, Lei Ma 0004, Jinmeng Wu |
IGARSS | 5 |
| 2024 | PointTFA: Training-Free Clustering Adaption for Large 3D Point Cloud Models
Jinmeng Wu, Hao Zhang 0047, Basura Fernando, Yanbin Hao, Hanyu Hong |
IJCAI | 1 |
| 2024 | JPA: A Joint-Part Attention for Mitigating Overfocusing on 3D Human Pose Estimation
Dengqing Yang, Zhenhua Tang 0001, Jinmeng Wu, Shuo Wang 0008, Lechao Cheng, Yanbin Hao |
PRCV (6) | 3 |
| 2024 | Iterative Semantic Transformer by Greedy Distillation for Community Question AnsweringabstractThe semantic matching problem consists of recognizing if the candidate text is relevant to a particular input text. Semantic similarities can be determined from human-curated knowledge, but such knowledge may not be available in every language. Instead, statistical learning techniques have been applied, but these techniques circumvent the need for manual feature engineering by using large datasets to train models to perform semantic similarity scoring between portions of text or words. The pre-trained transformer provides a further mechanism to consolidate the information throughout a sentence into single sentence-level representations, but these representations may not be optimal for the matching task. As an alternative, we propose an interactive semantic transformer based on a greedy layer-wise framework to learn a distributed similarity representation for sentence pairs. The novelty of the architecture lies in an abstract representation of the semantic similarities created by three-stage learning strategies. Model training is accomplished through a greedy layer-wise training scheme, that incorporates both supervised and unsupervised learning. The proposed model is experimentally compared to state-of-the-art approaches on three different dataset types: the library TREC, the Yahoo!, and Stack Exchange community question datasets, and results show the proposed model outperforming other approaches. Jinmeng Wu, Tingting Mu, Jeyan Thiyagalingam, Hanyu Hong, Yanbin Hao, Tianxu Zhang, John Yannis Goulermas |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2024 | Deep Progressive Asymmetric Quantization Based on Causal Intervention for Fine-Grained Image RetrievalabstractIn the field of computer vision, fine-grained image retrieval is an extremely challenging task due to the inherently subtle intra-class object variations. In addition, the high-dimensional real-valued features extracted from large-scale fine-grained image datasets slow the retrieval speed and increase the storage cost. To solve above issues, existing fine-grained image retrieval methods mainly focus on finding more discriminative local regions for generating discriminative and compact hash codes, which achieve limited fine-grained image retrieval performance due to the large quantization errors and the confounding granularities and context of discriminative parts, i.e., the correct recognition of fine-grained objects mainly attribute to the discriminative parts and their context. To learn robust causal features and reduce the quantization errors, we propose a deep progressive asymmetric quantization (DPAQ) method based on causal intervention to learn compact and robust descriptions for fine-grained image retrieval task. Specifically, we introduce a structural causal model to learn robust casual features via causal intervention for fine-grained visual recognition. Subsequently, we design a progressive asymmetric quantization layer in the feature embedding space, which can preserve the semantic information and reduce the quantization errors sufficiently. Finally, we incorporate both the fine-grained image classification and retrieval tasks into an end-to-end deep learning architecture for generating robust and compact descriptions. Experimental results on several fine-grained image retrieval datasets demonstrate that the proposed DPAQ method performs the best for fine-grained image retrieval task and surpasses the state-of-the art fine-grained hashing methods by a large margin. Lei Ma 0004, Hanyu Hong, Fanman Meng, Qingbo Wu 0001, Jinmeng Wu |
IEEE Trans. Multim. | 5 |
| 2023 | How Can Contrastive Pre-training Benefit Audio-Visual Segmentation? A Study from Supervised and Zero-shot Perspectives
Jiarui Yu, Yanbin Hao, Jinmeng Wu, Tong Xu 0001, Shuo Wang 0008, Xiangnan He 0001 |
BMVC | 4 |
| 2023 | Unsupervised Encoder-Decoder Model for Anomaly Prediction Task
Jinmeng Wu, Pengcheng Shu, Hanyu Hong, Xingxun Li, Lei Ma 0004, Yaozong Zhang, Ying Zhu 0002, Lei Wang 0068 |
MMM (2) | 1 |
| 2023 | Scribble-attention hierarchical network for weakly supervised salient object detection in optical remote sensing images
Lei Ma 0004, Hanyu Hong, Yaozong Zhang, Lei Wang 0068, Jinmeng Wu |
Appl. Intell. | 6 |
| 2023 | Car Emotion Labeling Based on Color-SSL Semi-Supervised Learning Algorithm by Color AugmentationabstractIn the era of emotional consumption, it has become a hot topic that commodities meet consumers’ emotional needs. As a necessity of life, the car also needs to meet the needs of consumers. To achieve that consumers can purchase cars according to their emotional needs, we need to label cars with emotional words. The car’s appearance is the crucial medium of emotional information transmission, especially the car’s color is an essential emotional factor. As the first impression of products, color affects people’s emotional attitude. Therefore, introducing color features into the training process of sample marking is an excellent idea for intelligent labeling of a large number of product emotions. This paper proposes a semi‐supervised learning method, Color‐SSL, based on color data augmentation to realize the label of car emotion. Color‐SSL takes FlexMatch as the framework of a semi‐supervised learning model and augments data by extracting subject color. Compared with the baseline method, the accuracy of this method improved by 3.2%, 8.3%, 8.6%, and 1.4% with 10, 50, 100, and 200 training samples and 1000 test samples. The results show that Color‐SSL obtains the best emotion‐label result (94%). In addition, this study publishes pictures of emotional car datasets with high resolution, orthogonal perspective, and uniform background. Zhuen Guo, Baoqi Liu, Kaixin Chang, Zuoya Jiang, Jinmeng Wu |
Int. J. Intell. Syst. | 8 |
| 2023 | Question-aware dynamic scene graph of local semantic representation learning for visual question answering
Jinmeng Wu, Fulin Ge, Hanyu Hong, Yu Shi 0004, Yanbin Hao, Lei Ma 0004 |
Pattern Recognit. Lett. | 1 |
| 2023 | Memory-Aware Attentive Control for Community Question Answering With Knowledge-Based Dual RefinementabstractThe question answering system in open domain enables a machine to automatically select and generate the answer for questions posed by humans in a natural language form on the website. Previous approaches seek effective ways of extracting the semantic features between question and answer, but the contextual information effects in semantic matching are still limited by short-term memory. As an alternative, we propose an internal knowledge-based end-to-end model, enhanced by an attentive memory network for both answer selection and answer generation tasks by considering the full advantages of the semantics and multifacts (i.e., timescales, topics, and context). In detail, we design a long-term memory to learn the top-$k$fine-grained similarity representations, where two memory-aware mechanisms aggregate the series of semantic word-level and sentence-level similarities to support the coarse contextual information. Furthermore, we propose a novel memory refinement mechanism with the two-dimensional of writing heads that offer an efficient approach to multiview selection of the salient word pairs. In the training stage, we adopt the transformer-based transfer learning skill to effectively pretrain the model. Experimentally, we compare the state-of-the-art approaches on four public datasets, the experimental results show that the proposed model achieves competitive performance. Jinmeng Wu, Tingting Mu, Jeyan Thiyagalingam, John Yannis Goulermas |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2022 | PolSAR-SSN: An End-to-End Superpixel Sampling Network for PolSAR Image ClassificationabstractPolarimetric synthetic aperture radar (PolSAR) image classification is one of the fundamental research areas in remote sensing. Superpixels can provide boundary constraint information and are widely used in PolSAR image interpretation. However, traditional machine learning superpixel algorithms have many limitations for PolSAR image interpretation. Pseudo-color images are usually used as the superpixel algorithm inputs, and the loss of polarimetric information will decrease the performance. In addition, the superpixel algorithms are difficult to incorporate into state-of-the-art deep learning models and cannot be trained in an end-to-end manner. In this letter, a trainable end-to-end deep superpixel network is proposed for PolSAR image classification. The inputs of the proposed method can be any low/middle-level polarimetric features of a PolSAR image and the rich polarimetric feature representation can be learned. The produced superpixels of the proposed method are more concentrated near the land cover boundaries and can significantly improve the performance of PolSAR image classification. Experimental results show that the overall accuracies of the proposed method are approximately 2.57% and 1.44% higher than traditional superpixel algorithms on two PolSAR datasets and surpass some well-known deep learning methods. Lei Wang 0068, Hanyu Hong, Yaozong Zhang, Jinmeng Wu, Lei Ma 0004, Ying Zhu 0002 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Attention in Attention: Modeling Context Correlation for Efficient Video ClassificationabstractAttention mechanisms have significantly boosted the performance of video classification neural networks thanks to the utilization of perspective contexts. However, the current research on video attention generally focuses on adopting a specific aspect of contexts (e.g., channel, spatial/temporal, or global context) to refine the features and neglects their underlying correlation when computing attentions. This leads to incomplete context utilization and hence bears the weakness of limited performance improvement. To tackle the problem, this paper proposes an efficient attention-in-attention (AIA) method for element-wise feature refinement, which investigates the feasibility of inserting the channel context into the spatio-temporal attention learning module, referred to as CinST, and also its reverse variant, referred to as STinC. Specifically, we instantiate the video feature contexts as dynamics aggregated along a specific axis with global average and max pooling operations. The workflow of an AIA module is that the first attention block uses one kind of context information to guide the gating weights calculation of the second attention that targets at the other context. Moreover, all the computational operations in attention units act on the pooled dimension, which results in quite few computational cost increase (https://github.com/haoyanbin918/Attention-in-Attention. Yanbin Hao, Shuo Wang 0008, Pei Cao 0001, Xinjian Gao, Tong Xu 0001, Jinmeng Wu, Xiangnan He 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2021 | Learning discrete class-specific prototypes for deep semantic hashing
Lei Ma 0004, Yu Shi 0004, Likun Huang, Zhenghua Huang, Jinmeng Wu |
Neurocomputing | 6 |
| 2020 | Cross-sentence Pre-trained Model for Interactive QA matchingabstractSemantic matching measures the dependencies between query and answer representations, it is an important criterion for evaluating whether the matching is successful. In fact, such matching does not examine each sentence individually, context information outside a sentence should be considered equally important to the syntactic context inside a sentence. We proposed a new QA matching model, built upon a cross-sentence context-aware architecture. An interactive attention mechanism with a pre-trained language model is proposed to automatically select salient positional answer representations that contribute more significantly to the answer relevance of a given question. In addition to the context information captured at each word position, we incorporate a new quantity of context information jump to facilitate the attention weight formulation. This reflects the amount of new information brought by the next word and is computed by modeling the joint probability between two adjacent word states. The proposed method is compared to multiple state-of-the-art ones evaluated using the TREC library, WikiQA, and the Yahoo! community question datasets. Experimental results show that the proposed method outperforms satisfactorily the competing ones. Jinmeng Wu, Yanbin Hao |
LREC | 1 |
| 2020 | Building interactive sentence-aware representation based on generative language model for community question answering
Jinmeng Wu, Tingting Mu, Jeyan Thiyagalingam, John Yannis Goulermas |
Neurocomputing | 1 |
| 2020 | Correlation Filtering-Based Hashing for Fine-Grained Image RetrievalabstractThe low storage and strong representation capabilities of hash codes for image retrievalhas made hashing technologies very popular. Several existing deep hashing methods focuson the task of general image retrieval, while neglecting the task of fine-grained image retrieval. Recently, some fine-grained hashing methods have been proposed to capture the subtle differences, which mainly utilize the single-modality visual features to solve the discriminative region localization while ignoring the semantic information. In this letter, we propose a correlation filtering hashing (CFH) method to learn discrete binary codes, which can adequately take advantage of the cross-modal correlation between the semantic information and the visual features for discriminative region localization. Specifically, we utilize a feature pyramid network to learn multi-level visual features. Subsequently, the label vector is embedded into the visual space, which can be used as a correlation filter on the feature maps to capture the latent location of objects. Finally, weperform global average pooling over the output maps and concatenate the features of different levels to produce the hash codes of query images. Extensive experiments on two fine-grained datasets show that the proposed CFH outperforms the state-of-the-art hashing methods. Lei Ma 0004, Yu Shi 0004, Jinmeng Wu |
IEEE Signal Process. Lett. | 4 |