VLDB 2026 Research / reviewers in the wild / expert
Maobin Tang
dblp:128/5026
· DBLP profile ↗
21ranked-venue papers
0as first author
17since 2021 · last 2026
0009-0005-5448-4172ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Computer networks · 3 · 3 since 2021Theory of computation · 3Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | End-to-End Knowledge Distillation for Unsupervised Domain Adaptation with Large Vision-language ModelsabstractKnowledge distillation based on large vision-language models (VLMs) has recently emerged as a significant solution to transfer knowledge from the source domain to the target domain in unsupervised domain adaptation (UDA) tasks. However, existing methods employ a two-stage training pipeline, which not only complicates the training procedure but also lacks interactions between the source and target domains, severely hindering real-time cross-domain knowledge transfer. To address these challenges, we propose End-to-End Knowledge Distillation for UDA with large VLMs (termed as EKDA). (1) EKDA employs a lightweight prompt learning mechanism to first embed the knowledge from the source domain into VLMs, and then simultaneously utilize the image encoder and text encoder of VLMs to perform knowledge distillation on the target domain, significantly reducing the domain gap. (2) EKDA designs a teacher-student alternating training strategy to implement real-time collaborative interactions across domains, enabling an end-to-end paradigm to provide accurate source domain-aware supervision for the target domain. We conduct extensive experiments on 4 widely recognized benchmark datasets including Office-31, Office-Home, VisDA-2017, and Mini-DomainNet. Experimental results demonstrate that EKDA achieves significant performance improvement over the state-of-the-art UDA approaches, while maintaining a much lower model complexity. Take Office-Home for example, EKDA has gained at least 2.7% performance improvement while reducing the learnable parameters by over 80% compared with the state-of-the-art UDA baselines. Yangtao Wang, Xingwei Deng, Yanzhao Xie, Weilong Peng, Siyuan Chen 0005, Xiaocui Li 0001, Maobin Tang, Meie Fang |
AAAI | 7 |
| 2026 | Spectral Efficiency Analysis for IRS-Assisted mmWave Massive MISO Systems with Mixed-Resolution ADCs
Weiqiang Tan, Pengling Li, Maobin Tang, Ting Liu 0013, Xiyuan Chen 0001, Chunguo Li |
INFOCOM | 3 |
| 2026 | Intra-modal consistency for image-text retrieval through soft-label distillation
Yangtao Wang, Yanzhao Xie, Siyuan Chen 0005, Weilong Peng, Maobin Tang, Meie Fang, C. L. Philip Chen, Ping Li 0016, Wensheng Zhang 0002 |
Pattern Recognit. | 6 |
| 2026 | F3-SD: Focal feature fusion with self-distillation on large vision-language models for cross-modal retrieval
Yangtao Wang, Yanzhao Xie, Xin Tan 0002, Xiaocui Li 0001, Maobin Tang, Meie Fang, Wensheng Zhang 0002 |
Pattern Recognit. | 6 |
| 2026 | Prompt-affinity multi-modal class centroids for unsupervised domain adaptionabstractIn recent years, the advancements in large vision-language models (VLMs) like CLIP have sparked a renewed interest in leveraging the prompt learning mechanism to preserve semantic consistency between source and target domains in unsupervised domain adaption (UDA). While these approaches show promising results, they encounter fundamental limitations when quantifying the similarity between source and target domain data , primarily stemming from the redundant and modality-missing class centroids . To address these limitations, we propose P rompt-affinity M ulti-modal C lass C entroids for UDA (termed as PMCC). Firstly, we fuse the text class centroids (directly generated from the text encoder of CLIP with manual prompts for each class) and image class centroids (generated from the image encoder of CLIP for each class based on source domain images) to yield the multi-modal class centroids. Secondly, we conduct the cross-attention operation between each source or target domain image and these multi-modal class centroids. In this way, these class centroids that contain rich semantic information of each class will serve as a bridge to effectively measure the semantic similarity between different domains. Finally, we design a logit bias head and employ a multi-modal prompt learning mechanism to accurately predict the true class of each image for both source and target domains. We conduct extensive experiments on 4 popular UDA datasets including Office-31, Office-Home, VisDA-2017, and DomainNet. The experimental results validate our PMCC achieves higher performance with lower model complexity than the state-of-the-art (SOTA) UDA methods. The code of this project is available at GitHub: https://github.com/246dxw/PMCC . Xingwei Deng, Yangtao Wang, Yanzhao Xie, Xiaocui Li 0001, Maobin Tang, Meie Fang, Wensheng Zhang 0002 |
Pattern Recognit. | 5 |
| 2026 | Cross-domain distillation for unsupervised domain adaptation with large vision-language models
Xingwei Deng, Yangtao Wang, Yanzhao Xie, Xin Tan 0002, Maobin Tang, Meie Fang, Wensheng Zhang 0002 |
Pattern Recognit. | 5 |
| 2026 | Adaptive message passing mechanism for graph neural networks
Yangtao Wang, Linruo Liu, Yanzhao Xie, Maobin Tang, Xiaocui Li 0001 |
Pattern Recognit. | 6 |
| 2026 | PTPD: Prototype-Guided Triplet Prompt Distillation with Vision-language models
Yanzhao Xie, Yangtao Wang, Rukai Wei, Dandan Shao, Maobin Tang, Meie Fang, Weilong Peng, Lisheng Fan, Wensheng Zhang 0002 |
Pattern Recognit. | 7 |
| 2026 | MKGPL: graph prompt learning with multi-view knowledge for few-shot recognition
Yanzhao Xie, Man Qiu, Yangtao Wang, Siyuan Chen 0005, Meie Fang, Maobin Tang, Wensheng Zhang 0002 |
Pattern Recognit. | 6 |
| 2026 | High Feature Distinguishability for Adaptive Image-text Matching with Dual-stream TransformersabstractRecently, most image-text matching (ITM) approaches have embraced a dual-stream transformer architecture to facilitate the learning and alignment of cross-modal semantic information. Despite the efficacy of this methodology in bridging the semantic disparity between images and texts, it exhibits two primary limitations. Firstly, it falls short in discriminating the nuanced similarities among features, which leads to misleading outcomes or even compromises the overall ITM process. Secondly, the conventional triplet training paradigm relies on a pre-determined, fixed margin coefficient, thereby impeding its capacity to accurately gauge the similarity relationships between positive and negative samples. In this article, we propose high feature D istinguishability for A daptive I mage-text M atching with dual-stream transformers (termed as DAIM). To address the first limitation, we design a feature discriminability module to bring similar features closer together but with a certain degree of distinction and push dissimilar features farther apart, resulting in high feature distinguishability for accurate ITM. To address the second limitation, we devise a margin optimization module to perceive the similarity distribution between positive and negative samples in real-time during training, thereby adaptively adjusting the margin coefficient to minimize the cross-modal semantic gap to the greatest extent possible. Based on this, we align the multi-level (i.e., representations from low-, middle-, and high-layer transformer encoders) semantic information of cross-modal data by adaptively optimizing the semantic distributions of positive and negative samples. We conduct extensive experiments on two commonly used benchmark datasets, including MSCOCO and Flickr30K. Experimental results verify that DAIM can achieve a higher performance (e.g., 4.7% RSUM gain on MSCOCO) than the state-of-the-art ITM methods. The open-sourced code of this project is available at: https://github.com/Hudjkfhdsjfhdjkg/DAIM.git . Yangtao Wang, Weibin Huang, Yanzhao Xie, Siyuan Chen 0005, Weilong Peng, Maobin Tang, Meie Fang, Wensheng Zhang 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2025 | Enhancing Cross-modal Semantic Consistency via Key Token Alignment for Image-text RetrievalabstractImage-text retrieval (ITR) plays a pivotal role in advancing intelligent transportation systems, facilitating efficient retrieval and utilization of multimedia data to enhance traffic management and safety significantly. However, existing ITR solutions have not effectively addressed the issues of image patch redundancy and text word redundancy, leading to erroneous image-text matching. In this paper, we propose SCTA that enhances cross-modal semantic consistency via key token alignment for ITR. Firstly, SCTA evaluates the importance of each image patch by calculating the self-attention scores within patches and cross-attention scores between patches and words. Secondly, SCTA implements aggregation operations on image and text separately, aiming to generate information-rich key image patch embeddings and text word token embeddings. Finally, SCTA completes fine-grained alignment by maximizing the similarity between patch-to-word and word-to-patch. Therefore, SCTA simultaneously addresses image patch redundancy and text word redundancy issues, enhancing semantic consistency by aligning the core semantic information between image-text pairs. Extensive experiments on multiple datasets including Flickr30K and MS-COCO verify the superior performance of SCTA compared with the SOTA fine-grained ITR methods. The code of this paper is released at GitHub: https://github.com/ICME2025ITR/SCTA. Huilong Lin, Yangtao Wang, Meie Fang, Yanzhao Xie, Xiaocui Li 0001, Weilong Peng, Siyuan Chen 0005, Maobin Tang, Ping Li 0016 |
ICME | 9 |
| 2025 | A Novel CSI Feedback Scheme for Massive MIMO Systems Using Differentiable Histogram Attention MechanismabstractAccurate channel state information (CSI) is critical for signal detection and precoding design in massive multiple-input multiple-output (MIMO) systems. However, traditional channel attention mechanisms for CSI feedback heavily rely on global pooling methods, which overlook finer-grained statistical patterns. In this paper, we propose a novel CSI feedback scheme for massive MIMO systems by utilizing a differentiable histogram attention mechanism, named DHANet, to boosts channel feature extraction and improve system performance. Specifically, DHANet replaces the traditional global pooling operations in the Squeeze-and-Excitation Block with kernel density estimation (KDE)-based differentiable histogram feature extraction, thereby enabling the capture of detailed channel-specific statistical information for more efficient CSI feedback. Moreover, the proposed mechanism can be seamlessly integrated into the existing CSI feedback architectures. Simulation results demonstrate that the DHANet achieves superior performance compared to the traditional pooling-based methods, particularly under 1/8 compression rates, highlighting its significant potential for CSI feedback in massive MIMO systems. Weiqiang Tan, Minwei Zhang, Maobin Tang, Jintao Wang 0002, Chunguo Li |
VTC2025-Fall | 3 |
| 2025 | Angle Metric Learning for Discriminative Features on Vehicle Re-IdentificationabstractABSTRACT Vehicle re‐identification (Re‐ID) facilitates the recognition and distinction of vehicles based on their visual characteristics in images or videos. However, accurately identifying a vehicle poses great challenges due to (i) the pronounced intra‐instance variations encountered under varying lighting conditions such as day and night and (ii) the subtle inter‐instance differences observed among similar vehicles. To address these challenges, the authors propose A ngle M etric learning for D iscriminative F eatures on vehicle Re‐ID (termed as AMDF), which aims to maximise the variance between visual features of different classes while minimising the variance within the same class. AMDF comprehensively measures the angle and distance discrepancies between features. First, to mitigate the impact of lighting conditions on intra‐class variation, the authors employ CycleGAN to generate images that simulate consistent lighting (either day or night), thereby standardising the conditions for distance measurement. Second, Swin Transformer was integrated to help generate more detailed features. At last, a novel angle metric loss based on cosine distance is proposed, which organically integrates angular metric and 2‐norm metric, effectively maximising the decision boundary in angular space. Extensive experimental evaluations on three public datasets including VERI‐776, VERI‐Wild, and VEHICLEID, indicate that the method achieves state‐of‐the‐art performance. The code of this project is released at https://github.com/ZnCu‐0906/AMDF . Yutong Xie 0009, Shuoqi Zhang, Lide Guo, Rukai Wei, Yanzhao Xie, Yangtao Wang, Maobin Tang, Lisheng Fan |
IET Comput. Vis. | 8 |
| 2025 | HSALC: hard sample aware label correction for medical image classification
Yangtao Wang, Yicheng Ye, Yanzhao Xie, Maobin Tang, Lisheng Fan |
Multim. Tools Appl. | 4 |
| 2024 | Image-text Retrieval with Main Semantics ConsistencyabstractImage-text retrieval (ITR) has been one of the primary tasks in cross-modal retrieval, serving as a crucial bridge between computer vision and natural language processing. Significant progress has been made to achieve global alignment and local alignment between images and texts by mapping images and texts into a common space to establish correspondences between these two modalities. However, the rich semantic content contained in each image may bring false matches, resulting in the matched text ignoring the main semantics but focusing on the secondary or other semantics of this image. To address this issue, this paper proposes a semantically optimized approach with a novel Main Semantics Consistency (MSC) loss function, which aims to rank the semantically most similar images (or texts) corresponding to the given query at the top position during the retrieval process. First, in each batch of image-text pairs, we separately compute (i) the image-image similarity, i.e., the similarity between every two images, (ii) the text-text similarity, i.e., the similarity between a group of texts (that belong to a certain image) and another group of texts (that belong to another image), and (iii) the image-text similarity, i.e., the similarity between each image and each text. Afterward, our proposed MSC effectively aligns the above image-image, image-text, and text-text similarity, since the main semantics of every two images will be highly close if their text descriptions remain highly semantically consistent. By this means, we can capture the main semantics of each image to be matched with its corresponding texts, prioritizing the semantically most related retrieval results. Extensive experiments on MSCOCO and FLICKR30K verify the superior performance of MSC compared with the SOTA image-text retrieval methods. The source code of this project is released at GitHub: https://github.com/xyi007/MSC. Yangtao Wang, Yanzhao Xie, Xin Tan 0002, Jingjing Li 0001, Xiaocui Li 0001, Weilong Peng, Maobin Tang, Meie Fang |
CIKM | 8 |
| 2024 | Channel Estimation for IRS-Assisted mmWave Massive MIMO Systems in Mixed-ADC ArchitectureabstractThe accuracy of channel state information (CSI) acquisition is nontrivial for intelligent reflection surface (IRS)-assisted massive multiple-input–multiple-output (MIMO) systems due to the IRS comprises of a large number of low-cost electromagnetic reflection units that can smartly reflect the impinging signal. In this article, we investigate the downlink channel estimation problem for IRS-assisted millimeter-wave (mmWave) massive MIMO systems, where the IRS provides effective reflected paths to enhance the coverage and spectral efficiency. To reduce the hardware costs and power consumption, the massive MIMO system is adopted the mixed-analog-to-digital converter (ADC) architecture, in which a fraction of antennas is equipped with high-resolution ADCs and a large number of antennas are deployed with low-resolution ADCs. Considered the finite-dimensional millimeter-wave channel model, we utilize the special row–column sparse characteristics of the channel matrix and propose a row-structured sparsity based on the orthogonal matching pursuit (RS-OMP) algorithm. The RS-OMP algorithm sequentially calculates the row–column support sets of the angular cascaded channel matrix and then uses the least square (LS) algorithm to reconstruct the cascaded channel matrix, which aim to achieve a lower computational complexity. Simulation results demonstrate that under the same system configuration, the proposed RS-OMP algorithm not only improves the accuracy of channel estimation but also reduces more than 75% of the pilot overhead compared to the traditional orthogonal matching pursuit algorithm. Rui Zhang 0117, Weiqiang Tan, Shidang Li, Maobin Tang |
IEEE Internet Things J. | 4 |
| 2022 | SICKNet: A Humor Detection Network Integrating Semantic Incongruity and Commonsense KnowledgeabstractHumor is a great linguistic tool to express feelings and enhance social bonding. Limited by the diversity of humor expressions and the differential understanding of listeners, automatic detection of humor text is still a difficult and important area in nature language processing. Current methods of humor detection mainly focus on fine-tuning of pre-trained language models, and rarely consider the degree of humor incongruity and knowledge distinction of contextual environments. To alleviate these challenges, we propose SICKNet, a novel multi-tasks learning network based on the incongruity theory of humor and commonsense knowledge. We first utilize the difference between set-up and punchline to detect the semantic incongruity of humor, and next use commonsense knowledge to detect the strength of humorous features. SICKNet achieves the start-of-the-art results on Reddit and TaivopJokes datasets, with accuracy rates of 76.27% and 73.64%, respectively. Our code is available at Github11https://github.com/xing-wei-zeng/SICKNet. Penglong Huang, Xingwei Zeng, Jinta Weng, Ying Gao 0003, Heyan Huang, Maobin Tang |
ICTAI | 6 |
| 2020 | A Primal-Dual Randomized Algorithm for the Online Weighted Set Multi-cover Problem
Wenbin Chen 0003, Fufang Li, Ke Qi, Miao Liu 0005, Maobin Tang |
TAMC | 5 |
| 2015 | Algorithms for the Densest Subgraph with at Least k Vertices and with a Specified Subset
Wenbin Chen 0003, Lingxi Peng, Jianxiong Wang, Fufang Li, Maobin Tang |
COCOA | 5 |
| 2014 | Solving the maximum duo-preservation string mapping problem with linear programming
Wenbin Chen 0006, Zhengzhang Chen, Nagiza F. Samatova, Lingxi Peng, Jianxiong Wang, Maobin Tang |
Theor. Comput. Sci. | 6 |
| 2013 | Inapproximability results for the minimum integral solution problem with preprocessing over ℓ∞ℓ∞ norm
Wenbin Chen 0003, Lingxi Peng, Jianxiong Wang, Fufang Li, Maobin Tang |
Theor. Comput. Sci. | 5 |