VLDB 2026 Research / reviewers in the wild / expert
Weifeng Ou
dblp:239/7289
· DBLP profile ↗
13ranked-venue papers
2as first author
11since 2021 · last 2025
0000-0002-8908-3863ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LLaVA-SpaceSGG: Visual Instruct Tuning for Open-Vocabulary Scene Graph Generation with Enhanced Spatial RelationsabstractScene Graph Generation (SGG) converts visual scenes into structured graph representations, providing deeper scene understanding for complex vision tasks. However, existing SGG models often overlook essential spatial relationships and struggle with generalization in open-vocabulary contexts. To address these limitations, we propose LLaVA-SpaceSGG, a multimodal large language model (MLLM) designed for open-vocabulary SGG with enhanced spatial relation modeling. To train it, we collect the SGG instruction-tuning dataset, named SpaceSGG. This dataset is constructed by combining publicly available datasets and synthesizing data using open-source models within our data construction pipeline. It combines object locations, object relations, and depth information, resulting in three data formats: spatial SGG description, question-answering, and conversation. To enhance the transfer of MLLMs' inherent capabilities to the SGG task, we introduce a two-stage training paradigm. Experiments show that LLaVA-SpaceSGG outperforms other open-vocabulary SGG methods, boosting recall by 8.6% and mean recall by 28.4% compared to the baseline. Our codebase, dataset, and trained models are publicly accessible on GitHub at the following URL: https://github.com/Endlinc/LLaVA-SpaceSGG. Mingjie Xu, Mengyang Wu, Yuzhi Zhao, Jason Chun Lok Li, Weifeng Ou |
WACV | 5 |
| 2024 | Motion Transfer-Driven Intra-Class Data Augmentation for Finger Vein RecognitionabstractFinger vein recognition (FVR) has emerged as a secure biometric technique because of the confidentiality of vascular bio-information. Recently, deep learning-based FVR has gained increased popularity and achieved promising performance. However, the limited size of public vein datasets has caused overfitting issues and greatly limits the recognition performance. Although traditional data augmentation can partially alleviate this data shortage issue, it cannot capture the real finger posture variations due to the rigid label-preserving image transformations, bringing limited performance improvement. To address this issue, we propose a novel motion transfer (MT) model for finger vein image data augmentation via modeling the actual finger posture and rotational movements. The proposed model first utilizes a key point detector to extract the key point and pose map of the source and drive finger vein images. We then utilize a dense motion module to estimate the motion optical flow, which is fed to an image generation module for generating the image with the target pose. Experiments conducted on three public finger vein databases demonstrate that the proposed motion transfer model can generate realistic intra-class augmented samples and effectively improve the recognition accuracy. Xiu-Feng Huang, Lai-Man Po, Weifeng Ou |
ICASSP | 3 |
| 2024 | Joint Learning of Identity and Vein Features for Enhanced Representations in Vascular BiometricsabstractVascular biometrics have shown great promise for secure authentication applications and have received increased attention in recent years. This paper proposes a novel framework for joint identity and segmentation feature learning to enrich representations and improve verification performance. The framework utilizes an encoder-decoder architecture, where the encoder is trained under metric learning supervision to extract discriminative identity features. Concurrently, the decoder is trained with vein mask segmentation supervision to extract vein pattern features. By jointly learning high-level identity features and low-level vein features in an end-to-end manner, the representations are enriched. We further design a bi-feature matching scheme utilizing score fusion to integrate both features for identity verification. Experiments conducted on public finger and palm vein datasets reveal that the proposed approach significantly improves verification accuracy, while introducing reasonable complexity overhead. Weifeng Ou, Lai-Man Po, Xiu-Feng Huang |
ICASSP | 1 |
| 2023 | CSRNet: Cascaded Selective Resolution Network for real-time semantic segmentation
Jingjing Xiong, Lai-Man Po, Wing Yin Yu, Chang Zhou 0008, Pengfei Xian, Weifeng Ou |
Expert Syst. Appl. | 6 |
| 2023 | ChildPredictor: A Child Face Prediction Framework With Disentangled LearningabstractThe appearances of children are inherited from their parents, which makes it feasible to predict them. Predicting realistic children's faces may help settle many social problems, such as age-invariant face recognition, kinship verification, and missing child identification. It can be regarded as an image-to-image translation task. Existing approaches usually assume domain information in the image-to-image translation can be interpreted by “style”, i.e., the separation of image content and style. However, such separation is improper for the child face prediction, because the facial contours between children and parents are not the same. To address this issue, we propose a new disentangled learning strategy for children's face prediction. We assume that children's faces are determined by genetic factors (compact family features, e.g., face contour), external factors (facial attributes irrelevant to prediction, such as moustaches and glasses), and variety factors (individual properties for each child). On this basis, we formulate predictions as a mapping from parents’ genetic factors to children's genetic factors, and disentangle them from external and variety factors. In order to obtain accurate genetic factors and perform the mapping, we propose a ChildPredictor framework. It transfers human faces to genetic factors by encoders and back by generators. Then, it learns the relationship between the genetic factors of parents and children through a mapping function. To ensure the generated faces are realistic, we collect a large Family Face Database to train ChildPredictor and evaluate it on the FF-Database validation set. Experimental results demonstrate that ChildPredictor is superior to other well-known image-to-image translation methods in predicting realistic and diverse child faces. Implementation codes can be found athttps://github.com/zhaoyuzhi/ChildPredictor. Yuzhi Zhao, Lai-Man Po, Qiong Yan, Wei Shen 0002, Yujia Zhang 0002, Wei Liu 0004, Chun Kit Wong, Chiu-Sing Pang, Weifeng Ou, Wing Yin Yu, Buhua Liu |
IEEE Trans. Multim. | 10 |
| 2023 | VCGAN: Video Colorization With Hybrid Generative Adversarial NetworkabstractWe propose a Video Colorization with Hybrid Generative Adversarial Network (VCGAN), an improved approach to video colorization using end-to-end learning and recurrent architecture. The VCGAN addresses two prevalent issues in the video colorization domain: Temporal consistency and the unification of colorization network and refinement network into a single architecture. To enhance colorization quality and spatiotemporal consistency, the mainstream of the generator in VCGAN is assisted by two additional networks,i.e.,global feature extractor and placeholder feature extractor, respectively. The global feature extractor encodes the global semantics of grayscale input to enhance colorization quality, whereas the placeholder feature extractor serves as a feedback connection to encode the semantics of the previous colorized frame in order to maintain spatiotemporal consistency. If changing the input for placeholder feature extractor as grayscale input, the hybrid VCGAN also has the potential to colorize single images. To improve the color consistency of far frames, we propose a dense long-term loss that minimizes the temporal disparity of every two remote frames. Trained with colorization and temporal losses jointly, VCGAN strikes a good balance between video color vividness and spatiotemporal continuity. Experimental results demonstrate that VCGAN produces higher-quality and temporally more consistent colorful videos than existing approaches. Yuzhi Zhao, Lai-Man Po, Wing Yin Yu, Yasar Abbas Ur Rehman, Mengyang Liu, Yujia Zhang 0002, Weifeng Ou |
IEEE Trans. Multim. | 7 |
| 2022 | Contrastive Spatio-Temporal Pretext Learning for Self-Supervised Video RepresentationabstractSpatio-temporal representation learning is critical for video self-supervised representation. Recent approaches mainly use contrastive learning and pretext tasks. However, these approaches learn representation by discriminating sampled instances via feature similarity in the latent space while ignoring the intermediate state of the learned representations, which limits the overall performance. In this work, taking into account the degree of similarity of sampled instances as the intermediate state, we propose a novel pretext task - spatio-temporal overlap rate (STOR) prediction. It stems from the observation that humans are capable of discriminating the overlap rates of videos in space and time. This task encourages the model to discriminate the STOR of two generated samples to learn the representations. Moreover, we employ a joint optimization combining pretext tasks with contrastive learning to further enhance the spatio-temporal representation learning. We also study the mutual influence of each component in the proposed scheme. Extensive experiments demonstrate that our proposed STOR task can favor both contrastive learning and pretext tasks and the joint optimization scheme can significantly improve the spatio-temporal representation in video understanding. The code is available at https://github.com/Katou2/CSTP. Yujia Zhang 0002, Lai-Man Po, Xuyuan Xu, Mengyang Liu, Yexin Wang, Weifeng Ou, Yuzhi Zhao, Wing Yin Yu |
AAAI | 6 |
| 2022 | Pixel Voting Decoder: A novel decoder that regresses pixel relationships for segmentation
Pengfei Xian, Lai-Man Po, Jingjing Xiong, Chang Zhou 0008, Yuzhi Zhao, Wing Yin Yu, Weifeng Ou, Yujia Zhang 0002, Xiaori Zhang |
Expert Syst. Appl. | 7 |
| 2022 | Angular Deep Supervised Vector Quantization for Image RetrievalabstractMost of the deep quantization methods adopt unsupervised approaches, and the quantization process usually occurs in the Euclidean space on top of the deep feature and its approximate value. When this approach is applied to the retrieval tasks, since the internal product space of the retrieval process is different from the Euclidean space of quantization, minimizing the quantization error (QE) does not necessarily lead to a good performance on the maximum inner product search (MIPS). To solve these problems, we treat Softmax classification as vector quantization (VQ) with angular decision boundaries and propose angular deep supervised VQ (ADSVQ) for image retrieval. Our approach can simultaneously learn the discriminative feature representation and the updatable codebook, both lying on a hypersphere. To reduce the QE between centroids and deep features, two regularization terms are proposed as supervision signals to encourage the intra-class compactness and inter-class balance, respectively. ADSVQ explicitly reformulates the asymmetric distance computation in MIPS to transform the image retrieval process into a two-stage classification process. Moreover, we discuss the extension of multiple-label cases from the perspective of quantization with binary classification. Extensive experiments demonstrate that the proposed ADSVQ has excellent performance on four well-known image data sets when compared with the state-of-the-art hashing methods. Chang Zhou 0008, Lai-Man Po, Weifeng Ou |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Fusion loss and inter-class data augmentation for deep finger vein feature learning
Weifeng Ou, Lai-Man Po, Chang Zhou 0008, Yasar Abbas Ur Rehman, Pengfei Xian, Yujia Zhang 0002 |
Expert Syst. Appl. | 1 |
| 2021 | Deep triplet residual quantization
Chang Zhou 0008, Lai-Man Po, Weifeng Ou, Pengfei Xian, Kwok-Wai Cheung 0002 |
Expert Syst. Appl. | 3 |
| 2020 | Data-level information enhancement: Motion-patch-based Siamese Convolutional Neural Networks for human activity recognition in videos
Yujia Zhang 0002, Lai-Man Po, Mengyang Liu, Yasar Abbas Ur Rehman, Weifeng Ou, Yuzhi Zhao |
Expert Syst. Appl. | 5 |
| 2019 | Face liveness detection using convolutional-features fusion of real and deep network generated face images
Yasar Abbas Ur Rehman, Lai-Man Po, Mengyang Liu, Zijie Zou, Weifeng Ou, Yuzhi Zhao |
J. Vis. Commun. Image Represent. | 5 |