VLDB 2026 Research / reviewers in the wild / expert
Ke Mei
dblp:251/5015
· DBLP profile ↗
9ranked-venue papers
2as first author
5since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Transfer learning and domain adaptation · 38% Vision and language · 38% Segmentation and scene understanding · 25% | |
| Computer graphics and multimedia
1 paper |
Geometric modeling and processing · 44% Multimedia analysis and retrieval · 44% Audio and music processing · 13% |
Topics — the 8 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Transfer learning and domain adaptation › domain adaptation › unsupervised domain adaptation
self-training |
1.3 | 2 | 2025 | Hard-Aware Instance Adaptive Self-Training for Unsupervised Cross-Domain Semantic Segmentation · IEEE Trans. Pattern Anal. Mach. Intell. 2025 Instance Adaptive Self-training for Unsupervised Domain Adaptation · ECCV (26) 2020 |
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation |
1.3 | 2 | 2025 | Hard-Aware Instance Adaptive Self-Training for Unsupervised Cross-Domain Semantic Segmentation · IEEE Trans. Pattern Anal. Mach. Intell. 2025 Instance Adaptive Self-training for Unsupervised Domain Adaptation · ECCV (26) 2020 |
Computer vision › Vision and language › vision-language pretraining
CLIP training |
0.9 | 1 | 2025 | HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models · ICCV 2025 |
Computer vision › Vision and language › vision-language dataset
image-text dataset construction |
0.9 | 1 | 2025 | HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models · ICCV 2025 |
Computer vision › Segmentation and scene understanding › pseudo-label learning
pseudo-label generation |
0.9 | 1 | 2025 | Hard-Aware Instance Adaptive Self-Training for Unsupervised Cross-Domain Semantic Segmentation · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.9 | 1 | 2025 | Hard-Aware Instance Adaptive Self-Training for Unsupervised Cross-Domain Semantic Segmentation · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Computer vision › Vision and language
vision-language pretraining |
0.9 | 1 | 2025 | HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models · ICCV 2025 |
Geometric modeling and processing › shape matching
semantic alignment |
0.9 | 1 | 2025 | HarmonySet: A Comprehensive Dataset for Understanding Video-Music Semantic Alignment and Temporal Synchronization · CVPR 2025 |
Methods — techniques the papers use, named apart from their topics
instance adaptive self-training · 1.3region-adaptive regularization · 0.9pseudo-label augmentation · 0.9multimodal annotation · 0.9large vision-language model · 0.9human-machine collaborative annotation · 0.9consistency constraint · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HarmonySet: A Comprehensive Dataset for Understanding Video-Music Semantic Alignment and Temporal SynchronizationabstractThis paper introduces HarmonySet, a comprehensive dataset designed to advance video-music understanding. Harmony-Set consists of 48,328 diverse video-music pairs, annotated with detailed information on rhythmic synchronization, emotional alignment, thematic coherence, and cultural relevance. We propose a multi-step human-machine collaborative framework for efficient annotation, combining human insights with machine-generated descriptions to identify key transitions and assess alignment across multiple dimensions. Addition ally, we introduce a novel evaluation framework with tasks and metrics to assess the multi-dimensional alignment of video and music, including rhythm, emotion, theme, and cultural context. Our extensive experiments demonstrate that HarmonySet, along with the proposed evaluation framework, significantly improves the ability of multimodal models to capture and analyze the intricate relationships between video and music. Project page: https://harmonyset.github.io/. Zitang Zhou, Ke Mei, Fengyun Rao |
CVPR | 2 |
| 2025 | HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models
Zhixiang Wei, Guangting Wang, Xiaoxiao Ma 0006, Ke Mei, Huaian Chen, Yi Jin 0002, Fengyun Rao |
ICCV | 4 |
| 2025 | Hard-Aware Instance Adaptive Self-Training for Unsupervised Cross-Domain Semantic SegmentationabstractThe divergence between labeled training data and unlabeled testing data is a significant challenge for recent deep learning models. Unsupervised domain adaptation (UDA) attempts to solve such problem. Recent works show that self-training is a powerful approach to UDA. However, existing methods have difficulty in balancing the scalability and performance. In this paper, we propose a hard-aware instance adaptive self-training framework for UDA on the task of semantic segmentation. To effectively improve the quality and diversity of pseudo-labels, we develop a novel pseudo-label generation strategy with an instance adaptive selector. We further enrich the hard class pseudo-labels with inter-image information through a skillfully designed hard-aware pseudo-label augmentation. Besides, we propose the region-adaptive regularization to smooth the pseudo-label region and sharpen the non-pseudo-label region. For the non-pseudo-label region, consistency constraint is also constructed to introduce stronger supervision signals during model optimization. Our method is so concise and efficient that it is easy to be generalized to other UDA methods. Experiments on GTA5 $\rightarrow$→ Cityscapes, SYNTHIA $\rightarrow$→ Cityscapes, and Cityscapes $\rightarrow$→ Oxford RobotCar demonstrate the superior performance of our approach compared with the state-of-the-art methods. Chuang Zhu, Kebin Liu 0002, Wenqi Tang, Ke Mei, Jiaqi Zou, Tiejun Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | DigestPath: A benchmark dataset with challenge review for the pathological detection and segmentation of digestive-system
Qian Da, Zhongyu Li 0002, Yanfei Zuo, Chenbin Zhang, Jingxin Liu 0005, Wen Chen 0001, Jiahui Li 0005, Dou Xu, Hongmei Yi, Zhe Wang 0043, Li Zhang 0040, Xianying He, Xiaofan Zhang 0002, Ke Mei, Chuang Zhu, Weizeng Lu, LinLin Shen, Jun Shi 0006, Jun Li 0106, Sreehari S, Ganapathy Krishnamurthi, Jiangcheng Yang, Tiancheng Lin 0001, Qingyu Song 0004, Xuechen Liu 0004, Simon Graham, Raja Muhammad Saad Bashir, Canqian Yang, Shaofei Qin, Xinmei Tian 0001, Jie Zhao 0014, Dimitris N. Metaxas, Hongsheng Li 0001, Chaofu Wang, Shaoting Zhang 0001 |
Medical Image Anal. | 18 |
| 2021 | Multi-level colonoscopy malignant tissue detection with adversarial CAC-UNet
Chuang Zhu, Ke Mei, Yihao Luo, Jun Liu 0014, Ying Wang 0043, Mulan Jin |
Neurocomputing | 2 |
| 2020 | Instance Adaptive Self-training for Unsupervised Domain Adaptation
Ke Mei, Chuang Zhu, Jiaqi Zou, Shanghang Zhang |
ECCV (26) | 1 |
| 2020 | Cross-Stained Segmentation from Renal Biopsy Images Using Multi-Level Adversarial LearningabstractSegmentation from renal pathological images is a key step in automatic analyzing the renal histological characteristics. However, the performance of models varies significantly in different types of stained datasets due to the appearance variations. In this paper, we design a robust and flexible model for cross-stained segmentation. It is a novel multi-level deep adversarial network architecture that consists of three sub-networks: (i) a segmentation network; (ii) a pair of multi-level mirrored discriminators for guiding the segmentation network to extract domain-invariant features; (iii) a shape discriminator that is utilized to further identify the output of the segmentation network and the ground truth. Experimental results on glomeruli segmentation from renal biopsy images indicate that our network is able to improve segmentation performance on target type of stained images and use unlabeled data to achieve similar accuracy to labeled data. In addition, this method can be easily applied to other tasks. Ke Mei, Chuang Zhu, Jun Liu 0014, Yuanyuan Qiao 0002 |
ICASSP | 1 |
| 2020 | Multi-Scale Video Inverse Tone Mapping with Deformable AlignmentabstractInverse tone mapping(iTM) is an operation to transform low-dynamic-range (LDR) content to high-dynamic-range (HDR) content, which is an effective technique to improve the visual experience. ITM has developed rapidly with deep learning algorithms in recent years. However, the great majority of deep-learning-based iTM methods are aimed at images and ignore the temporal correlations of consecutive frames in videos. In this paper, we propose a multi-scale video iTM network with deformable alignment, which increases time consistency in videos. We first align the input consecutive LDR frames at the feature level by deformable convolutions and then simultaneously use multi-frame information to generate the HDR frame. Additionally, we adopt a multi-scale iTM architecture with a pyramid pooling module, which enables our network to reconstruct details as well as global features. The proposed network achieves better performance compared to other iTM methods on quantitative metrics and gain a significant visual improvement. Jiaqi Zou, Ke Mei, Songlin Sun |
VCIP | 2 |
| 2019 | Adaptive Frame Rate Optimization Based on Particle Swarm and Neural Network for Industrial Video StreamabstractThe emergence of a large number of video data puts forward higher requirements on traditional video transmission technology. The new streaming media technology based on HTTP dynamic adaptive streaming DASH transmission protocol has become an important research direction of video services. How to overcome the unstable characteristics of wireless links in a limited bandwidth, achieve high-quality intelligent transmission of video, and obtain optimal user quality of experience (QoE), has become an urgent problem to be solved. This paper abandons the traditional streaming media adaptive transmission method, and combines neural network and particle swarm optimization algorithm to design a new intelligent transmission scheme. The particle swarm optimization algorithm obtains the optimal transmission parameters of QoE, and the model established by neural network predicts the optimal one. The system sets parameters to ensure video service quality under limited bandwidth and large network fluctuations in the wireless network. Ke Mei |
ETFA | 3 |