VLDB 2026 Research / reviewers in the wild / expert
Shumin Han
dblp:119/8234
· DBLP profile ↗
15ranked-venue papers
4as first author
11since 2021 · last 2024
0000-0001-5602-9863ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | UV-IDM: Identity-Conditioned Latent Diffusion Model for Face UV-Texture Generationabstract3D face reconstruction aims at generating high-fidelity 3D face shapes and textures from single-view or multi-view images. However, current prevailing facial texture generation methods generally suffer from low-quality texture, identity information loss, and inadequate handling of occlusions. To solve these problems, we introduce an Identity-Conditioned Latent Diffusion Model for face UV-texture generation (UV-IDM) to generate photo-realistic textures based on the Basel Face Model (BFM). UV-IDM leverages the powerful texture generation capacity of a latent diffusion model (LDM) to obtain detailed facial textures. To preserve the identity during the reconstruction procedure, we design an identity-conditioned module that can utilize any in-the-wild image as a robust condition for the LDM to guide texture generation. UV-IDM can be easily adapted to different BFM-based methods as a high-fidelity texture generator. Furthermore, in light of the limited accessibility of most existing UV-texture datasets, we build a large-scale and publicly available UV-texture dataset based on BFM, termed BFM-UV. Extensive experiments show that our UV-IDM can generate high-fidelity textures in 3D face reconstruction within seconds while maintaining image consistency, bringing new state-of-the-art performance in facial texture generation. Hong Li 0016, Yutang Feng, Xuhui Liu, Bohan Zeng, Shanglin Li, Jianzhuang Liu, Shumin Han, Baochang Zhang 0001 |
CVPR | 9 |
| 2024 | WAVE: Warping DDIM Inversion Features for Zero-Shot Text-to-Video Editing
Yutang Feng, Sicheng Gao, Yuxiang Bao, Shumin Han, Baochang Zhang 0001, Angela Yao |
ECCV (76) | 5 |
| 2024 | Context Autoencoder for Self-supervised Representation Learning
Xiaokang Chen, Mingyu Ding, Shentong Mo, Shumin Han, Ping Luo 0002, Jingdong Wang 0001 |
Int. J. Comput. Vis. | 7 |
| 2024 | MAFormer: A transformer network with multi-scale attention fusion for visual recognition
Huixin Sun, Baochang Zhang 0001, Xianbin Cao 0001, Errui Ding, Shumin Han |
Neurocomputing | 9 |
| 2023 | Prompt Tuning Inversion for Text-Driven Image Editing Using Diffusion ModelsabstractRecently large-scale language-image models (e.g., text-guided diffusion models) have considerably improved the image generation capabilities to generate photorealistic images in various domains. Based on this success, current image editing methods use texts to achieve intuitive and versatile modification of images. To edit a real image using diffusion models, one must first invert the image to a noisy latent from which an edited image is sampled with a target text prompt. However, most methods lack one of the following: user-friendliness (e.g., additional masks or precise descriptions of the input image are required), generalization to larger domains, or high fidelity to the input image. In this paper, we design an accurate and quick inversion technique, Prompt Tuning Inversion, for text-driven image editing. Specifically, our proposed editing method consists of a reconstruction stage and an editing stage. In the first stage, we encode the information of the input image into a learnable conditional embedding via Prompt Tuning Inversion. In the second stage, we apply classifier-free guidance to sample the edited image, where the conditional embedding is calculated by linearly interpolating between the target embedding and the optimized one obtained in the first stage. This technique ensures a superior trade-off between editability and high fidelity to the input image of our method. For example, we can change the color of a specific object while preserving its original shape and background under the guidance of only a target text prompt. Extensive experiments on ImageNet demonstrate the superior editing performance of our method compared to the state-of-the-art baselines. Wenkai Dong, Xiaoyue Duan, Shumin Han |
ICCV | 4 |
| 2022 | Multi-party Privacy-Preserving Record Linkage Method Based on Trusted Execution Environment
Xuefei He, Haiping Wei, Shumin Han, Derong Shen |
WISA | 3 |
| 2022 | Efficient Multi-party Privacy-Preserving Record Linkage Based on Blockchain
Haoshan Yao, Haiping Wei, Shumin Han, Derong Shen |
WISA | 3 |
| 2021 | Student-Teacher Feature Pyramid Matching for Anomaly Detection
Guodong Wang 0006, Shumin Han, Errui Ding, Di Huang 0001 |
BMVC | 2 |
| 2021 | Layer-Wise Searching for 1-Bit Detectorsabstract1-bit detectors show great promise for resource-constrained embedded devices but often suffer from a significant performance gap compared with their real-valued counterparts. The primary reason lies in the error during binarization. This paper presents a layer-wise searching (LWS) strategy to generate 1-bit detectors that maintain a performance very close to the original real-valued model. The approach introduces angular and amplitude loss functions to increase detector capacity. At 1-bit layers, it exploits a differentiable binarization search (DBS) to minimize the angular error in a student-teacher framework. We also learn the scale factor by minimizing the amplitude loss in the same student-teacher framework. Extensive experiments show that LWS-Det outperforms state-of-the-art 1-bit detectors by a considerable margin on the PASCAL VOC and COCO datasets. For example, the LWS-Det achieves 1-bit Faster-RCNN with ResNet-34 backbone within 2.0% mAP of its real-valued counterpart on the PASCAL VOC dataset. Sheng Xu 0007, Junhe Zhao, Jinhu Lü 0001, Baochang Zhang 0001, Shumin Han, David S. Doermann |
CVPR | 5 |
| 2021 | Dual-stream Network for Visual RecognitionabstractTransformers with remarkable global representation capacities achieve competitive results for visual tasks, but fail to consider high-level local pattern information in input images. In this paper, we present a generic Dual-stream Network (DS-Net) to fully explore the representation capacity of local and global pattern features for image classification. Our DS-Net can simultaneously calculate fine-grained and integrated features and efficiently fuse them. Specifically, we propose an Intra-scale Propagation module to process two different resolutions in each block and an Inter-Scale Alignment module to perform information interaction across features at dual scales. Besides, we also design a Dual-stream FPN (DS-FPN) to further enhance contextual information for downstream dense predictions. Without bells and whistles, the proposed DS-Net outperforms DeiT-Small by 2.4\% in terms of top-1 accuracy on ImageNet-1k and achieves state-of-the-art performance over other Vision Transformers and ResNets. For object detection and instance segmentation, DS-Net-Small respectively outperforms ResNet-50 by 6.4\% and 5.5 \% in terms of mAP on MSCOCO 2017, and surpasses the previous state-of-the-art scheme, which significantly demonstrates its potential to be a general backbone in vision tasks. The code will be released soon. Mingyuan Mao, Peng Gao 0007, Renrui Zhang, Honghui Zheng, Teli Ma, Errui Ding, Baochang Zhang 0001, Shumin Han |
NeurIPS | 9 |
| 2021 | Learning From Large-Scale Noisy Web Data With Ubiquitous Reweighting for Image ClassificationabstractMany important advances of deep learning techniques have originated from the efforts of addressing the image classification task on large-scale datasets. However, the construction of clean datasets is costly and time-consuming since the Internet is overwhelmed by noisy images with inadequate and inaccurate tags. In this paper, we propose a Ubiquitous Reweighting Network (URNet) that can learn an image classification model from noisy web data. By observing the web data, we find that there are five key challenges, i.e., imbalanced class sizes, high intra-classes diversity and inter-class similarity, imprecise instances, insufficient representative instances, and ambiguous class labels. With these challenges in mind, we assume every training instance has the potential to contribute positively by alleviating the data bias and noise via reweighting the influence of each instance according to different class sizes, large instance clusters, its confidence, small instance bags, and the labels. In this manner, the influence of bias and noise in the data can be gradually alleviated, leading to the steadily improving performance of URNet. Experimental results in the WebVision 2018 challenge with 16 million noisy training images from 5000 classes show that our approach outperforms state-of-the-art models and ranks first place in the image classification task. Jia Li 0003, Yafei Song 0002, Lele Cheng, Pengcheng Yuan, Shumin Han |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2020 | Efficient private multi-party numerical records matching
Shumin Han, Derong Shen, Tiezheng Nie, Yue Kou, Ge Yu 0001 |
Frontiers Comput. Sci. | 1 |
| 2017 | Private Blocking Technique for Multi-party Privacy-Preserving Record LinkageabstractThe process of matching and integrating records that relate to the same entity from one or more datasets is known as record linkage, and it has become an increasingly important subject in many application areas, including business, government and health system. The data from these areas often contain sensitive information. To prevent privacy breaches, ideally records should be linked in a private way such that no information other than the matching result is leaked in the process, and this technique is called privacy-preserving record linkage (PPRL). With the increasing data, scalability becomes the main challenge of PPRL, and many private blocking techniques have been developed for PPRL. They are aimed at reducing the number of record pairs to be compared in the matching process by removing obvious non-matching pairs without compromising privacy. However, most of them are designed for two databases and they vary widely in their ability to balance competing goals of accuracy, efficiency and security. In this paper, we propose a novel private blocking approach for PPRL based on dynamic k -anonymous blocking and Paillier cryptosystem which can be applied on two or multiple databases. In dynamic k -anonymous blocking, our approach dynamically generates blocks satisfying k -anonymity and more accurate values to represent the blocks with varying k . We also propose a novel similarity measure method which performs on the numerical attributes and combines with Paillier cryptosystem to measure the similarity of two or more blocks in security, which provides strong privacy guarantees that none information reveals even collusion. Experiments conducted on a public dataset of voter registration records validate that our approach is scalable to large databases and keeps a high quality of blocking. We compare our method with other techniques and demonstrate the increases in security and accuracy. Shumin Han, Derong Shen, Tiezheng Nie, Yue Kou, Ge Yu 0001 |
Data Sci. Eng. | 1 |
| 2016 | Scalable Private Blocking Technique for Privacy-Preserving Record Linkage
Shumin Han, Derong Shen, Tiezheng Nie, Yue Kou, Ge Yu 0001 |
APWeb (2) | 1 |
| 2012 | An Efficient Background Reconstruction Based Coding Method for Surveillance Videos Captured by Moving CameraabstractWith the proliferation of moving surveillance cameras, how to effectively compress videos captured from them is becoming more and more important. One significant characteristic is that, these cameras always go and return cyclically within a limited area. Thus we propose to dynamically build up a background frame for each input frame from a generated panorama background and employ it for a background frame based motion compensation to improve the coding efficiency. For the background reconstruction procedure, we firstly extract limited number of feature point pairs between the robustly searched area in the decoded panorama and the current frame. Afterwards, the global motion transformation matrix is obtained to rectify the searched area into a projective plane of the current frame, and then the reconstructed background is produced. Experiments on six in-door and out-door surveillance videos show that, the background reconstruction based coding method achieves significant performance gain. Shumin Han, Xianguo Zhang, Yonghong Tian 0001, Tiejun Huang 0001 |
AVSS | 1 |