Wentao Ma 0003

dblp:39/8088-3 · DBLP profile ↗
← Back
5ranked-venue papers in the field
1as first author
5since 2021 · last 2027
0000-0003-3059-6629ORCID · conflict

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 5 (1 first)
YearPublicationVenuePosition
2027 TailoredUAA: Uncertainty-Aware Alignment with Noisy Correspondence for Text-based Person Re-identification
Wentao Ma 0003, Shaofan Chen, Lu Liu 0023, Guolong Shi, Zhongyang Yao, Lichuan Gu
Inf. Process. Manag.2
2026 GjPest4CMR: LLMs-Assisted Morphology-Aware Fine-grained Goji Pest Image-Text Retrieval
abstract
Fine-grained agricultural pest image-text retrieval demands robust cross-modal matching capable of handling subtle appearance differences and structured multi-sentence descriptions. Yet, small-scale domain data introduces two key challenges in supervision and representation. First, existing corpora typically pair each image with a single coarse paragraph, providing incomplete supervision that omits life-stage appearances and subtle morphological traits, while high-frequency redundant fields shared across pest categories further cause semantic confounding. Second, prevailing retrieval models rely on global pooled alignment, which averages out low-occupancy but discriminative local cues such as stripe layouts, body segmentation, and wing venation, entangling near-neighbor categories. To address these issues, we propose StruMorAlign, a parameter-efficient cross-attention CLIP framework for fine-grained pest image-text retrieval. We first construct GjPest4CMR, an English goji pest image-text retrieval dataset created via LLMs-assisted re-writing with expert-in-the-loop verification, providing multi-sentence structured descriptions. On the modeling side, we design StruPD (Structured Semantic Deconfounding) to provide multi-granularity structured supervision and strengthen coarse-to-fine consistency learning on morphology-relevant local regions, addressing the first challenge; and introduce GateA2 (Gated Adaptive Alignment) self-distillation with high-momentum EMA teacher for parameter-efficient, morphology-aware cross-modal alignment, addressing the second challenge. On the GjPest4CMR dataset, StruMorAlign achieves mR 78.72% for both image-to-text and text-to-image retrieval, with I2T R@1=53.11% and T2I R@1=51.53%, outperforming the strongest baseline HarMA by +9.72 mR. Ablation studies confirm the complementary benefits of StruPD and GateA2, and qualitative analyses including t-SNE visualizations and attention heatmaps further validate the effectiveness of our approach.
Wentao Ma 0003, Yuwei Wang 0002, Lu Liu 0023, Junfei Yang, Xianbao Xu, Yuan Rao 0003
ICMR2
2024 DeepEnhancer: Temporally Consistent Focal Transformer for Comprehensive Video Enhancement
abstract
Restoring and colorizing old films is a comprehensive video enhancement task, marked by the presence of heterogeneous and structured degradations. Our DeepEnhancer addresses this challenge through a unified pipeline that combines restoration and colorization. In this workflow, we implement a bidirectional propagation strategy. Specifically, we incorporate second-order feature alignment to reduce the accumulation of inaccuracies in optical flow estimation. Simultaneously, we utilize cross-scale long-term attention mechanisms to model correlations within hidden states, thereby ensuring spatial and temporal consistency. To address the notable content loss in aging films, we introduce a temporally consistent focal transformer guided by global information. This transformer utilizes various window levels with distinct sub-window sizes to seamlessly integrate fine-grained and coarse-grained features. Comprehensive experimental results conclusively demonstrate the superiority of our model in both quantitative and qualitative comparisons when compared to existing approaches.
Lihua Chi, Wentao Ma 0003, Feng Li 0037, Jie Liu 0002
ICMR4
2023 Turning backdoors for efficient privacy protection against image retrieval violations
Qiang Liu 0004, Tongqing Zhou, Zhiping Cai, Yuan Yuan 0034, Ming Xu 0002, Jiaohua Qin, Wentao Ma 0003
Inf. Process. Manag.7
2023 Adaptive multi-feature fusion via cross-entropy normalization for effective image retrieval
Wentao Ma 0003, Tongqing Zhou, Jiaohua Qin, Xuyu Xiang, Yun Tan, Zhiping Cai
Inf. Process. Manag.1