VLDB 2026 Research / reviewers in the wild / expert
Wei Peng 0009
dblp:16/5560-9
· DBLP profile ↗
39ranked-venue papers
11as first author
33since 2021 · last 2026
0000-0002-2892-5764ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 7 first-author · 18 since 2021Artificial intelligence and machine learning · 19 · 6 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 9 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Ivy-Fake: A Unified Explainable Framework and Benchmark for Image and Video AIGC DetectionabstractThe rapid development of Artificial Intelligence Generated Content (AIGC) techniques has enabled the creation of high-quality synthetic content, but it also raises significant security concerns. Current detection methods face two major limitations: (1) the lack of multidimensional explainable datasets for generated images and videos. Existing open-source datasets (e.g., WildFake, GenVideo) rely on oversimplified binary annotations, which restrict the explainability and trustworthiness of trained detectors. (2) Prior MLLM-based forgery detectors (e.g., FakeVLM) exhibit insufficiently fine-grained interpretability in their step-by-step reasoning, which hinders reliable localization and explanation. To address these challenges, we introduce Ivy-Fake, the first large-scale multimodal benchmark for fake image and video detection. It consists of over 106K richly annotated training samples (images and videos) and 5,000 manually verified evaluation examples, sourced from multiple generative models and real-world datasets through a carefully designed pipeline to ensure both diversity and quality. Furthermore, we propose Ivy-xDetector, a multimodel large language model (MLLM) based on reinforcement fine-tuning (RFT), capable of producing explainable reasoning chains and achieving robust performance across multiple fake image and video detection benchmarks. Changjiang Jiang, Fengchang Yu, Wei Peng 0009, Xinbin Yuan, Yifei Bi, Zian Zhou, Chenyang Si, Caifeng Shan |
ICMR | 5 |
| 2026 | Integrating Anatomical Priors Into a Causal Diffusion Modelabstract3D brain MRI studies often examine subtle morphometric differences between cohorts that are hard to detect visually. Given the high cost of MRI acquisition, these studies could greatly benefit from image syntheses, particularly counterfactual image generation, as has been the case for applications in computer vision. However, counterfactual models struggle to produce anatomically plausible MRIs due to a lack of explicit inductive biases to preserve fine-grained anatomical details. This shortcoming arises from the training of models that optimize overall image appearance (e.g., via cross-entropy) rather than preserving subtle, yet medically relevant, local variations across subjects. To preserve subtle variations, we propose to explicitly integrate anatomical constraints at the voxel level as priors into a generative diffusion framework. Termed Probabilistic Causal Graph Model (PCGM), the approach captures anatomical constraints via a probabilistic graph module and translates those constraints into spatial binary masks of regions where subtle variations occur. The masks (encoded by a 3D ControlNet) constrain a novel counterfactual denoising UNet, whose encodings are then transferred into high-quality brain MRIs via our 3D diffusion decoder. Extensive experiments across multiple datasets demonstrate that PCGM generates structural brain MRIs of higher quality than several baseline approaches. Furthermore, we show, for the first time, that brain measurements extracted from counterfactuals (generated by PCGM) replicate the subtle effects of a disease on cortical brain regions previously reported in the neuroscience literature. This achievement is an important milestone in the use of synthetic MRIs in studies investigating subtle morphological differences. The codes are available at https://github.com/AndyCA111/PCGM. Binxu Li, Wei Peng 0009, Mingjie Li 0006, Ehsan Adeli-Mosabbeb, Kilian M. Pohl |
IEEE Trans. Medical Imaging | 2 |
| 2025 | FreeNet: Liberating Depth-Wise Separable Operations for Building Faster Mobile Vision ArchitecturesabstractIn the pursuit of efficient vision architectures, substantial efforts have been devoted to optimizing operator efficiency. Depth-wise separable operators, such as DWConv, are found cheap in both FLOPs and parameters. As a result, they are increasingly incorporated into efficient backbones, trading for deeper and wider architectures to enhance performance. However, separable operators are not really fast on devices due to the discontinuous memory access requirements. In this paper, we propose FreeNets, a family of simple and efficient backbones that free the separable operation to further accelerate the running speed. We introduce sparse sampling mixers (S2-Mixer) to supersede existing separable token mixers. The S2-Mixer samples multiple segments of partially continuous signals across spatial and channel dimensions for convolutional processing, achieving extremely fast on-device speed. The sparse sampling also enables S2-Mixer to capture long-range pixel relationships from dynamic receptive fields. Furthermore, we introduce a Shift Feed-Forward Network (ShiftFFN) as a faster alternative to existing channel mixers. It utilizes a shift neck architecture that aggregates global information to shift features, enabling faster channel mixing while incorporating global pixel information. Extensive experiments demonstrate that FreeNet offers a superior accuracy-efficiency tradeoff compared to the latest efficient models. On ImageNet-1k, FreeNet-S2 outperforms the StarNet-S4 by 0.4% in top-1 accuracy, while running around 40% faster on desktop GPU and 15% faster on Mobile GPU. Hao Yu 0015, Haoyu Chen 0001, Wei Peng 0009, Xu Cheng 0003, Guoying Zhao 0001 |
AAAI | 3 |
| 2025 | PFDiff: Training-Free Acceleration of Diffusion Models Combining Past and Future ScoresabstractDiffusion Probabilistic Models (DPMs) have shown remarkable potential in image generation, but their sampling efficiency is hindered by the need for numerous denoising steps. Most existing solutions accelerate the sampling process by proposing fast ODE solvers. However, the inevitable discretization errors of the ODE solvers are significantly magnified when the number of function evaluations (NFE) is fewer. In this work, we propose PFDiff, a novel training-free and orthogonal timestep-skipping strategy, which enables existing fast ODE solvers to operate with fewer NFE. Specifically, PFDiff initially utilizes score replacement from past time steps to predict a springboard. Subsequently, it employs this ``springboard" along with foresight updates inspired by Nesterov momentum to rapidly update current intermediate states. This approach effectively reduces unnecessary NFE while correcting for discretization errors inherent in first-order ODE solvers. Experimental results demonstrate that PFDiff exhibits flexible applicability across various pre-trained DPMs, particularly excelling in conditional DPMs and surpassing previous state-of-the-art training-free methods. For instance, using DDIM as a baseline, we achieved 16.46 FID (4 NFE) compared to 138.81 FID with DDIM on ImageNet 64x64 with classifier guidance, and 13.06 FID (10 NFE) on Stable Diffusion with 7.5 guidance scale. Code is available at https://github.com/onefly123/PFDiff. Guangyi Wang, Yuren Cai, Lijiang Li, Wei Peng 0009, Songzhi Su |
ICLR | 4 |
| 2025 | Diffusion Sampling Correction via Approximately 10 ParametersabstractWhile powerful for generation, Diffusion Probabilistic Models (DPMs) face slow sampling challenges, for which various distillation-based methods have been proposed. However, they typically require significant additional training costs and model parameter storage, limiting their practicality. In this work, we propose **P**CA-based **A**daptive **S**earch (PAS), which optimizes existing solvers for DPMs with minimal additional costs. Specifically, we first employ PCA to obtain a few basis vectors to span the high-dimensional sampling space, which enables us to learn just a set of coordinates to correct the sampling direction; furthermore, based on the observation that the cumulative truncation error exhibits an ``S"-shape, we design an adaptive search strategy that further enhances the sampling efficiency and reduces the number of stored parameters to approximately 10. Extensive experiments demonstrate that PAS can significantly enhance existing fast solvers in a plug-and-play manner with negligible costs. E.g., on CIFAR10, PAS optimizes DDIM's FID from 15.69 to 4.37 (NFE=10) using only **12 parameters and sub-minute training** on a single A100 GPU. Code is available at https://github.com/onefly123/PAS. Guangyi Wang, Wei Peng 0009, Lijiang Li, Yuren Cai, Songzhi Su |
ICML | 2 |
| 2025 | WASABI: A Metric for Evaluating Morphometric Plausibility of Synthetic Brain MRIs
Bahram Jafrasteh, Wei Peng 0009, Yimin Luo, Ehsan Adeli-Mosabbeb, Qingyu Zhao |
MICCAI (2) | 2 |
| 2025 | Generating Novel Brain Morphology by Deforming Learned Templates
Alan Q. Wang 0001, Fangrui Huang, Bailey Trang Nguyen, Wei Peng 0009, Mohammad H. Abbasi, Kilian M. Pohl, Mert R. Sabuncu, Ehsan Adeli-Mosabbeb |
MICCAI (2) | 4 |
| 2025 | Evidential Remote Physiological Measurement via Uncertainty-aware Fusion of Video and RFabstractRemote physiological measurement enables the capture of vital signals in a non-contact way, which offers significant potential for various applications. Monitoring these signals is achieved through video cameras or radio frequency (RF) sensors, with recent few methods attempting to fuse both sources to leverage complementary patterns for enhanced accuracy. However, these two modalities operate on distinct principles, where video-based methods detect subtle facial color changes from blood volume variations, while RF-based methods capture subtle body vibration due to heartbeats. In practical applications, they may encounter interference at different occasions. Treating these modalities as equally reliable in all situations can lead to suboptimal fusion. To address this issue, we propose an evidential video-RF fusion framework for robust remote physiological signal measurement. We design an uncertainty regression head for each uni-modality, which estimates uncertainty features together with the corresponding physiological signal in each branch. Then an evidential multi-modal fusion module is employed to dynamically fuse the two modalities according to their uncertainty. Extensive experiments carried on public and self-collected datasets show that the proposed method not only achieves superior fusion performance on easy data collected under well-controlled environment, it also generalizes well to unseen data which represents challenging practical conditions that one or both sensors are disturbed. Jieyi Ge, Zhaodong Sun, Wei Peng 0009, Chenhang Ying, Yuwei Chen 0005, Kui Ren 0001 |
ACM Multimedia | 3 |
| 2025 | VoRec: Enhancing Recommendation with Voronoi Diagram in Hyperbolic SpaceabstractThe sparse user-item interactions in recommender systems hinder the quality of embedding representations and degraded recommendation performance. Existing methods attempt to alleviate this sparsity issue by incorporating auxiliary information via item tags, but often neglect structured characteristics of embedding space, such as semantic distribution and logical relations. To this end, we propose VoRec, a novel framework that explores the spatial distribution of items and their associated tags to achieve accurate recommendations in hyperbolic space. Specifically, we employ the Voronoi diagram to partition hyperbolic space into logically related subspaces, based on tag distributions and relationships derived from existing tag taxonomies. In addition, we combine the Voronoi diagram with the Hyperbolic Graph Convolutional Network (HGCN) and exploit the respective advantages of the Poincaré and Lorentz models in hyperbolic space. Finally, we develop two types of Voronoi site update strategies, namely active and passive ones, to optimize the Voronoi diagram for recommendation tasks. The active strategy employs contrastive learning to guide updates to the Voronoi diagram, while the passive strategy adaptively freezes parameters based on information gain to regulate the update rate. Extensive experiments on four real-world benchmark datasets demonstrate that our proposed VoRec framework delivers substantial performance improvements, achieving an average 16.35% enhancement in Recall and NDCG metrics compared to state-of-the-art baselines. The model implementation is publicly available at: https://github.com/s35lay/VoRec. Li Li 0122, Wei Peng 0009, Songzhi Su |
SIGIR | 3 |
| 2025 | HRCUNet: Hierarchical Region Contrastive Learning for Segmentation of Breast Tumors in DCE-MRIabstractABSTRACT Segmenting breast tumors from dynamic contrast‐enhanced magnetic resonance images is a critical step in the early detection and diagnosis of breast cancer. However, this task becomes significantly more challenging due to the diverse shapes and sizes of tumors, which make it difficult to establish a unified perception field for modeling them. Moreover, tumor regions are often subtle or imperceptible during early detection, exacerbating the issue of extreme class imbalance. This imbalance can lead to biased training and challenge accurately segmenting tumor regions from the predominant normal tissues. To address these issues, we propose a hierarchical region contrastive learning approach for breast tumor segmentation. Our approach introduces a novel hierarchical region contrastive learning loss function that addresses the class imbalance problem. This loss function encourages the model to create a clear separation between feature embeddings by maximizing the inter‐class margin and minimizing the intra‐class distance across different levels of the feature space. In addition, we design a novel Attention‐based 3D Multi‐scale Feature Fusion Residual Module to explore more granular multi‐scale representations to improve the feature learning ability of tumors. Extensive experiments on two breast DCE‐MRI datasets demonstrate that the proposed algorithm is more competitive against several state‐of‐the‐art approaches under different segmentation metrics. Jiezhou He, Zhiming Luo, Wei Peng 0009, Songzhi Su, Shaozi Li |
Concurr. Comput. Pract. Exp. | 3 |
| 2024 | CC-DA: Cross-Domain Consistency Data Augmentation for 3D Tumor SegmentationabstractDeep learning-based tumor segmentation in 3D medical images faces the challenges of limited annotated data and class imbalance. In this paper, we proposed a novel Cross-domain Consistency Data Augmentation (CC-DA) for 3D tumor segmentation. Specifically, we copy the tumor from source data and apply random transformations to enhance its diversity. Then, we paste the enhanced tumor into the organ area of target data to generate a new sample. This process can alleviate class imbalance by regulating the merged tumor pixel ratio. To further enhance the generated data credibility, we proposed a domain consistency constraint that aligns the source data distribution with the target data distribution. We conduct extensive experiments on KiTS19 and LiTS17 datasets. The promising results clearly show that our CC-DA method can effectively improve the existing state-of-the-art 3D tumor segmentation performance. Jiezhou He, Zhiming Luo, Wei Peng 0009, Songzhi Su, Shaozi Li |
ICASSP | 3 |
| 2024 | Evaluating the Quality of Brain MRI Generators
Jiaqi Wu 0016, Wei Peng 0009, Binxu Li, Yu Zhang 0009, Kilian M. Pohl |
MICCAI (10) | 2 |
| 2024 | Metadata-conditioned generative models to synthesize anatomically-plausible 3D brain MRIs
Wei Peng 0009, Tomas M. Bosschieter, Jiahong Ouyang, Robert Paul, Edith V. Sullivan, Adolf Pfefferbaum, Ehsan Adeli-Mosabbeb, Qingyu Zhao, Kilian M. Pohl |
Medical Image Anal. | 1 |
| 2024 | Discovering attention-guided cross-modality correlation for visible-infrared person re-identification
Hao Yu 0015, Xu Cheng 0003, Kevin H. M. Cheng, Wei Peng 0009, Zitong Yu, Guoying Zhao 0001 |
Pattern Recognit. | 4 |
| 2024 | Data Leakage and Evaluation Issues in Micro-Expression AnalysisabstractMicro-expressions have drawn increasing interest lately due to various potential applications. The task is, however, difficult as it incorporates many challenges from the fields of computer vision, machine learning and emotional sciences. Due to the spontaneous and subtle characteristics of micro-expressions, the available training and testing data are limited, which make evaluation complex. We show that data leakage and fragmented evaluation protocols are issues among the micro-expression literature. We find that fixing data leaks can drastically reduce model performance, in some cases even making the models perform similarly to a random classifier. To this end, we go through common pitfalls, propose a new standardized evaluation protocol using facial action units with over 2000 micro-expression samples, and provide an open source library that implements the evaluation protocols in a standardized manner. Code is publicly available inhttps://github.com/tvaranka/meb. Tuomas Varanka, Yante Li, Wei Peng 0009, Guoying Zhao 0001 |
IEEE Trans. Affect. Comput. | 3 |
| 2024 | Geometric Graph Representation With Learnable Graph Structure and Adaptive AU Constraint for Micro-Expression RecognitionabstractMicro-expression recognition (MER) holds significance in uncovering hidden emotions. Most works take image sequences as input and cannot effectively explore ME information because subtle ME-related motions are easily submerged in unrelated information. Instead, the facial landmark is a lowdimensional and compact modality, which achieves lower computational cost and potentially concentrates on ME-related movement features. However, the discriminability of facial landmarks for MER is unclear. Thus, this paper investigates the contribution of facial landmarks and proposes a novel framework to efficiently recognize MEs with facial landmarks. Firstly, a geometric twostream graph network is constructed to aggregate the low-order and high-order geometric movement information from facial landmarks to obtain discriminative ME representation. Secondly, a self-learning fashion is introduced to automatically model the dynamic relationship between nodes even long-distance nodes. Furthermore, an adaptive action unit loss is proposed to reasonably build a strong correlation between landmarks, facial action units and MEs. Notably, this work provides a novel idea with much higher efficiency to promote MER, only utilizing graphbased geometric features. The experimental results demonstrate that the proposed method achieves competitive performance with a significantly reduced computational cost. Furthermore, facial landmarks significantly contribute to MER and are worth further study for high-efficient ME analysis. Jinsheng Wei, Wei Peng 0009, Guanming Lu, Yante Li, Jingjie Yan, Guoying Zhao 0001 |
IEEE Trans. Affect. Comput. | 2 |
| 2024 | Hyperbolic Uncertainty Aware Semantic SegmentationabstractSemantic segmentation (SS) aims to classify each pixel into one of the pre-defined classes. This task plays an important role in self-driving cars and autonomous drones. In SS, many works have shown that most misclassified pixels are commonly near object boundaries with high uncertainties. However, existing SS loss functions are not tailored to handle these uncertain pixels during training, as these pixels are usually treated equally as confidently classified pixels and cannot be embedded with arbitrary low distortion in Euclidean space, thereby degenerating the performance of SS. To overcome this problem, this paper designs a Hyperbolic Uncertainty Loss (HyperUL), which dynamically highlights the misclassified and high-uncertainty pixels in Hyperbolic space during training via the hyperbolic distances. The proposed HyperUL is model agnostic and can be easily applied to various neural architectures. After employing HyperUL to three recent SS models, the experimental results on Cityscapes, UAVid, and ACDC datasets reveal that the segmentation performance of existing SS models can be consistently improved. Additionally, reliable measurement of model uncertainty plays a key role in real-world applications such as autonomous controls of vehicles and drones. To meet this requirement, we propose the Hyperbolic Uncertainty Estimation method, which is easily implemented by only post-processing the generated Hyperbolic embeddings. By this approach, we can calculate the uncertainty values almost for free. Quantitative and qualitative results on Cityscapes, UAVid, and ACDC datasets verify that our proposed uncertainty estimation method usually outputs more meaningful results compared with popular MC-dropout and ensembling methods. Bike Chen, Wei Peng 0009, Xiaofeng Cao 0002, Juha Röning |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | MedSyn: Text-Guided Anatomy-Aware Synthesis of High-Fidelity 3-D CT ImagesabstractThis paper introduces an innovative methodology for producing high-quality 3D lung CT images guided by textual information. While diffusion-based generative models are increasingly used in medical imaging, current state-of-the-art approaches are limited to low-resolution outputs and underutilize radiology reports' abundant information. The radiology reports can enhance the generation process by providing additional guidance and offering fine-grained control over the synthesis of images. Nevertheless, expanding text-guided generation to high-resolution 3D images poses significant memory and anatomical detail-preserving challenges. Addressing the memory issue, we introduce a hierarchical scheme that uses a modified UNet architecture. We start by synthesizing low-resolution images conditioned on the text, serving as a foundation for subsequent generators for complete volumetric data. To ensure the anatomical plausibility of the generated samples, we provide further guidance by generating vascular, airway, and lobular segmentation masks in conjunction with the CT images. The model demonstrates the capability to use textual input and segmentation tasks to generate synthesized images. Algorithmic comparative assessments and blind evaluations conducted by 10 board-certified radiologists indicate that our approach exhibits superior performance compared to the most advanced models based on GAN and diffusion techniques, especially in accurately retaining crucial anatomical features such as fissure lines and airways. This innovation introduces novel possibilities. This study focuses on two main objectives: (1) the development of a method for creating images based on textual prompts and anatomical components, and (2) the capability to generate new images conditioning on anatomical elements. The advancements in image generation can be applied to enhance numerous downstream tasks. Yanwu Xu 0003, Li Sun 0010, Wei Peng 0009, Shuyue Jia, Katelyn Morrison, Adam Perer, Afrooz Zandifar, Shyam Visweswaran, Motahhare Eslami, Kayhan Batmanghelich |
IEEE Trans. Medical Imaging | 3 |
| 2024 | Rethinking Few-Shot Class-Incremental Learning With Open-Set Hypothesis in Hyperbolic GeometryabstractBy training first with a large base dataset, Few-Shot Class-Incremental Learning (FSCIL) aims at continually learning a sequence of few-shot learning tasks with novel classes. There are mainly two challenges in FSCIL: the overfitting issue of novel classes with limited labeled samples and the catastrophic forgetting of previously seen classes. The current protocol of FSCIL is built by mimicking the general class-incremental learning setting by building a unified framework, while the existing frameworks for FSCIL on this protocol always bias to the classes in the base dataset because the dominant performance of the deep model is decided by the size of the training dataset. Moreover, it is difficult to handle the stability-plasticity constraint in a unified FSCIL framework. To solve these issues, we rethink the configuration of FSCIL with the open-set hypothesis by reserving the possibility in the first session for incoming categories. To find a better decision boundary of close space and open space, Hyperbolic Reciprocal Point Learning module (Hyper-RPL) is built on Reciprocal Point Learning with hyperbolic neural networks. Besides, when learning novel categories from limited labeled data, we incorporate a hyperbolic metric learning (Hyper-Metric) module into the distillation-based framework to alleviate the overfitting issue and better handle the trade-off issue between the preservation of old knowledge and the acquisition of new knowledge. Finally, the comprehensive assessments of the proposed configuration and modules on three benchmark datasets are executed to validate the effectiveness, and state-of-the-art results are achieved. Yawen Cui, Zitong Yu, Wei Peng 0009, Qi Tian 0001, Li Liu 0002 |
IEEE Trans. Multim. | 3 |
| 2023 | SRBGCN: Tangent space-Free Lorentz Transformations for Graph Feature Learning
Abdelrahman Mostafa, Wei Peng 0009, Guoying Zhao 0001 |
BMVC | 2 |
| 2023 | TOPLight: Lightweight Neural Networks with Task-Oriented Pretraining for Visible-Infrared RecognitionabstractVisible-infrared recognition (VI recognition) is a challenging task due to the enormous visual difference across heterogeneous images. Most existing works achieve promising results by transfer learning, such as pretraining on the ImageNet, based on advanced neural architectures like ResNet and ViT. However, such methods ignore the neg-ative influence of the pretrained colour prior knowledge, as well as their heavy computational burden makes them hard to deploy in actual scenarios with limited resources. In this paper, we propose a novel task-oriented pretrained lightweight neural network (TOPLight) for VI recognition. Specifically, the TOPLight method simulates the domain conflict and sample variations with the proposed fake do-main loss in the pretraining stage, which guides the network to learn how to handle those difficulties, such that a more general modality-shared feature representation is learned for the heterogeneous images. Moreover, an effective fine-grained dependency reconstruction module (FDR) is developed to discover substantial pattern dependencies shared in two modalities. Extensive experiments on VI person re-identification and VI face recognition datasets demonstrate the superiority of the proposed TOPLight, which signifi-cantly outperforms the current state of the arts while de-manding fewer computational resources. Hao Yu 0015, Xu Cheng 0003, Wei Peng 0009 |
CVPR | 3 |
| 2023 | Modality Unifying Network for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) is a challenging task due to large cross-modality discrepancies and intra-class variations. Existing methods mainly focus on learning modality-shared representations by embedding different modalities into the same feature space. As a result, the learned feature emphasizes the common patterns across modalities while suppressing modality-specific and identity-aware information that is valuable for Re-ID. To address these issues, we propose a novel Modality Unifying Network (MUN) to explore a robust auxiliary modality for VI-ReID. First, the auxiliary modality is generated by combining the proposed cross-modality learner and intra-modality learner, which can dynamically model the modality-specific and modality-shared representations to alleviate both cross-modality and intra-modality variations. Second, by aligning identity centres across the three modalities, an identity alignment loss function is proposed to discover the discriminative feature representations. Third, a modality alignment loss is introduced to consistently reduce the distribution distance of visible and infrared images by modality prototype modeling. Extensive experiments on multiple public datasets demonstrate that the proposed method surpasses the current state-of-the-art methods by a significant margin. Hao Yu 0015, Xu Cheng 0003, Wei Peng 0009, Guoying Zhao 0001 |
ICCV | 3 |
| 2023 | LSOR: Longitudinally-Consistent Self-Organized Representation Learning
Jiahong Ouyang, Qingyu Zhao, Ehsan Adeli-Mosabbeb, Wei Peng 0009, Greg Zaharchuk, Kilian M. Pohl |
MICCAI (1) | 4 |
| 2023 | Generating Realistic Brain MRIs via a Conditional Diffusion Probabilistic Model
Wei Peng 0009, Ehsan Adeli-Mosabbeb, Tomas M. Bosschieter, Sanghyun Park 0004, Qingyu Zhao, Kilian M. Pohl |
MICCAI (8) | 1 |
| 2022 | Learning Optimal K-space Acquisition and Reconstruction using Physics-Informed Neural NetworksabstractThe inherent slow imaging speed of Magnetic Resonance Image (MRI) has spurred the development of various acceleration methods, typically through heuristically undersampling the MRI measurement domain known as k-space. Recently, deep neural networks have been applied to reconstruct undersampled k-space data and have shown improved reconstruction performance. While most of these methods focus on designing novel reconstruction networks or new training strategies for a given undersampling pattern, e.g., Cartesian undersampling or Non-Cartesian sampling, to date, there is limited research aiming to learn and optimize k-space sampling strategies using deep neural networks. This work proposes a novel optimization framework to learn k-space sampling trajectories by considering it as an Ordinary Differential Equation (ODE) problem that can be solved using neural ODE. In particular, the sampling of k-space data is framed as a dynamic system, in which neural ODE is formulated to approximate the system with additional constraints on MRI physics. In addition, we have also demonstrated that trajectory optimization and image reconstruction can be learned collaboratively for improved imaging efficiency and reconstruction performance. Experiments were conducted on different in-vivo datasets (e.g., brain and knee images) acquired with different sequences. Initial results have shown that our proposed method can generate better image quality in accelerated MRI than conventional undersampling schemes in Cartesian and Non-Cartesian acquisitions. Wei Peng 0009, Guoying Zhao 0001, Fang Liu 0005 |
CVPR | 1 |
| 2022 | Hyperbolic Spatial Temporal Graph Convolutional NetworksabstractSpatial-temporal graph convolutional networks (ST-GCNs) have been successfully applied for dynamic graphs representation learning, such as modeling skeleton-based human actions. However, ST-GCNs embed these non-Euclidean graph structures into Euclidean space, which is not the natural space to represent such structures as embedding them in this space incurs a large distortion. In this work, we make use of hyperbolic non-Euclidean geometry and construct compact ST-GCNs in the hyperbolic space. It can be shown that hyperbolic ST-GCNs (HST-GCNs) outperform the corresponding Euclidean counterparts. Additionally, these compact hyperbolic models can be used to increase the performance of large complex Euclidean models. Moreover, we show that the same or even better performance of large Euclidean models can be achieved by fusing the scores of smaller Euclidean models and a compact hyperbolic model. This in turn leads to reducing the total number of model parameters and hence model size. To validate the performance of these hyperbolic networks, we conducted extensive experiments on NTU RGB+D, NTU RGB+D 120 and Kinectics-Skeleton datasets for human action recognition. Abdelrahman Mostafa, Wei Peng 0009, Guoying Zhao 0001 |
ICIP | 2 |
| 2022 | Hyperbolic Deep Neural Networks: A SurveyabstractRecently, hyperbolic deep neural networks (HDNNs) have been gaining momentum as the deep representations in the hyperbolic space provide high fidelity embeddings with few dimensions, especially for data possessing hierarchical structure. Such a hyperbolic neural architecture is quickly extended to different scientific fields, including natural language processing, single-cell RNA-sequence analysis, graph embedding, financial analysis, and computer vision. The promising results demonstrate its superior capability, significant compactness of the model, and a substantially better physical interpretability than its counterpart in the euclidean space. To stimulate future research, this paper presents a comprehensive review of the literature around the neural components in the construction of HDNN, as well as the generalization of the leading deep approaches to the hyperbolic space. It also presents current applications of various tasks, together with insightful observations and identifying open questions and promising future directions. Wei Peng 0009, Tuomas Varanka, Abdelrahman Mostafa, Henglin Shi, Guoying Zhao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Micro-expression Action Unit Detection with Dual-view Attentive Similarity-Preserving Knowledge DistillationabstractEncoding facial expressions via action units (AUs) has been found to be effective in resolving the ambiguity issue among different expressions. Therefore, AU detection plays an important role for emotion analysis. While a number of AU detection methods have been proposed for common facial expressions, there is very limited study for micro-expression AU detection. Micro-expression AU detection is challenging because of the weakness of micro-expression appearance and the spontaneous characteristic leading to difficult collection, thus has small-scale datasets. In this paper, we focus on the micro-expression AU detection and expect to contribute to the community. To address above issues, a novel dual-view attentive similarity-preserving distillation method is proposed for robust micro-expression AU detection by leveraging massive facial expressions in the wild. Through such an attentive similarity-preserving distillation method, we break the domain shift problem and essential AU knowledge from common facial AUs is efficiently distilled. Furthermore, considering that the generalization ability of teacher network is important for knowledge distillation, a semi-supervised co-training approach is developed to construct a generalized teacher network for learning discriminative AU representation. Extensive experiments have demonstrated that our proposed knowledge distillation method can effectively distill and transfer the cross-domain knowledge for robust micro-expression AU detection. Yante Li, Wei Peng 0009, Guoying Zhao 0001 |
FG | 2 |
| 2021 | Intrinsic-Extrinsic Preserved GANs for Unsupervised 3D Pose TransferabstractWith the strength of deep generative models, 3D pose transfer regains intensive research interests in recent years. Existing methods mainly rely on a variety of constraints to achieve the pose transfer over 3D meshes, e.g., the need for manually encoding for shape and pose disentanglement. In this paper, we present an unsupervised approach to conduct the pose transfer between any arbitrate given 3D meshes. Specifically, a novel Intrinsic-Extrinsic Preserved Generative Adversarial Network (IEP-GAN) is presented for both intrinsic (i.e., shape) and extrinsic (i.e., pose) information preservation. Extrinsically, we propose a co-occurrence discriminator to capture the structural/pose invariance from distinct Laplacians of the mesh. Meanwhile, intrinsically, a local intrinsic-preserved loss is introduced to preserve the geodesic priors while avoiding heavy computations. At last, we show the possibility of using IEP-GAN to manipulate 3D human meshes in various ways, including pose transfer, identity swapping and pose interpolation with latent code vector arithmetic. The extensive experiments on various 3D datasets of humans, animals and hands qualitatively and quantitatively demonstrate the generality of our approach. Our proposed model produces better results and is substantially more efficient compared to recent state-of-the-art methods. Code is available: https://github.com/mikecheninoulu/Unsupervised_IEPGAN Haoyu Chen 0001, Hao Tang 0005, Henglin Shi, Wei Peng 0009, Nicu Sebe, Guoying Zhao 0001 |
ICCV | 4 |
| 2021 | Rethinking the ST-GCNs for 3D skeleton-based human action recognitionabstractThe skeletal data has been an alternative for the human action recognition task as it provides more compact and distinct information compared to the traditional RGB input. However, unlike the RGB input, the skeleton data lies in a non-Euclidean space that traditional deep learning methods are not able to use their fullest potential. Fortunately, with the emerging trend of Geometric deep learning, the spatial-temporal graph convolutional network (ST-GCN) has been proposed to deal with the action recognition problem from skeleton data. ST-GCN and its variants fit well with skeleton-based action recognition and are becoming the mainstream frameworks for this task. However, the efficiency and the performance of the task are hindered by either fixing the skeleton joint correlations or providing a computational expensive strategy to construct a dynamic topology for the skeleton. We argue that many of these operations are either unnecessary or even harmful for the task. By theoretically and experimentally analysing the state-of-the-art ST-GCNs, we provide a simple but efficient strategy to capture the global graph correlations and thus efficiently model the representation of the input graph sequences. Moreover, the global graph strategy also reduces the graph sequence into the Euclidean space, thus a multi-scale temporal filter is introduced to efficiently capture the dynamic information. With the method, we are not only able to better extract the graph correlations with much fewer parameters (only 12.6% of the current best), but we also achieve a superior performance. Extensive experiments on current largest 3D datasets, NTU-RGB+D and NTU-RGB+D 120, demonstrate the ability of our network to perform efficient and lightweight priority on this task. Wei Peng 0009, Jingang Shi, Tuomas Varanka, Guoying Zhao 0001 |
Neurocomputing | 1 |
| 2021 | A hybrid quantum-classical neural network with deep residual learningabstractInspired by the success of classical neural networks, there has been tremendous effort to develop classical effective neural networks into quantum concept. In this paper, a novel hybrid quantum-classical neural network with deep residual learning (Res-HQCNN) is proposed. We firstly analyse how to connect residual block structure with a quantum neural network, and give the corresponding training algorithm. At the same time, the advantages and disadvantages of transforming deep residual learning into quantum concept are provided. As a result, the model can be trained in an end-to-end fashion, analogue to the backpropagation in classical neural networks. To explore the effectiveness of Res-HQCNN , we perform extensive experiments for quantum data with or without noisy on classical computer. The experimental results show the Res-HQCNN performs better to learn an unknown unitary transformation and has stronger robustness for noisy data, when compared to state of the arts. Moreover, the possible methods of combining residual learning with quantum neural networks are also discussed. Yanying Liang, Wei Peng 0009, Zhu-Jun Zheng, Olli Silvén, Guoying Zhao 0001 |
Neural Networks | 2 |
| 2021 | Tripool: Graph triplet pooling for 3D skeleton-based action recognitionabstractGraph Convolutional Network (GCN) has already been successfully applied to skeleton-based action recognition. However, current GCNs in this task are lack of pooling operations such that the architectures are inherently flat, which not only increases the computational complexity but also requires larger memory space to keep the entire graph embedding. More seriously, a flat architecture forces the high-level semantic feature representations to have the same physical structure of the low-level input skeletons, which we argue is unreasonable and harmful for the final performance. To address these issues, we propose Tripool, a novel graph pooling method for 3D action recognition from skeleton data. Tripool provides to optimize a triplet pooling loss, in which both graph topology and global graph context are taken into consideration, to learn a hierarchical graph representation. The training process of graph pooling is efficient since it optimizes the graph topology by minimizing an upper bound of the pooling loss. Besides, Tripool also automatically generates an embedding matrix since the graph is changed after pooling. On one hand, Tripool reduces the computational cost by removing the redundant nodes. On the other hand it overcomes the limitation of the topology constrain for the high-level semantic representations, thus improves the final performance. Tripool can be combined with various graph neural networks in an end-to-end fashion. Comprehensive experiments on two current largest scale 3D datasets are conducted to evaluate our method. With our Tripool, we consistently get the best results in terms of various performance measures. Wei Peng 0009, Xiaopeng Hong, Guoying Zhao 0001 |
Pattern Recognit. | 1 |
| 2021 | Spatial Temporal Graph Deconvolutional Network for Skeleton-Based Human Action RecognitionabstractBenefited from the powerful ability of spatial temporal Graph Convolutional Networks (ST-GCNs), skeleton-based human action recognition has gained promising success. However, the node interaction through message propagation does not always provide complementary information. Instead, it May even produce destructive noise and thus make learned representations indistinguishable. Inevitably, the graph representation would also become over-smoothing especially when multiple GCN layers are stacked. This paper proposes spatial-temporal graph deconvolutional networks (ST-GDNs), a novel and flexible graph deconvolution technique, to alleviate this issue. At its core, this method provides a better message aggregation by removing the embedding redundancy of the input graphs from either node-wise, frame-wise or element-wise at different network layers. Extensive experiments on three current most challenging benchmarks verify that ST-GDN consistently improves the performance and largely reduce the model size on these datasets. Wei Peng 0009, Jingang Shi, Guoying Zhao 0001 |
IEEE Signal Process. Lett. | 1 |
| 2020 | Learning Graph Convolutional Network for Skeleton-Based Human Action Recognition by Neural SearchingabstractHuman action recognition from skeleton data, fuelled by the Graph Convolutional Network (GCN) with its powerful capability of modeling non-Euclidean data, has attracted lots of attention. However, many existing GCNs provide a pre-defined graph structure and share it through the entire network, which can loss implicit joint correlations especially for the higher-level features. Besides, the mainstream spectral GCN is approximated by one-order hop such that higher-order connections are not well involved. All of these require huge efforts to design a better GCN architecture. To address these problems, we turn to Neural Architecture Search (NAS) and propose the first automatically designed GCN for this task. Specifically, we explore the spatial-temporal correlations between nodes and build a search space with multiple dynamic graph modules. Besides, we introduce multiple-hop modules and expect to break the limitation of representational capacity caused by one-order approximation. Moreover, a corresponding sampling- and memory-efficient evolution strategy is proposed to search in this space. The resulted architecture proves the effectiveness of the higher-order approximation and the layer-wise dynamic graph modules. To evaluate the performance of the searched model, we conduct extensive experiments on two very large scale skeleton-based action recognition datasets. The results show that our model gets the state-of-the-art results in term of given metrics. Wei Peng 0009, Xiaopeng Hong, Haoyu Chen 0001, Guoying Zhao 0001 |
AAAI | 1 |
| 2020 | Mix Dimension in Poincaré Geometry for 3D Skeleton-based Action RecognitionabstractGraph Convolutional Networks (GCNs) have already demonstrated their powerful ability to model the irregular data, e.g., skeletal data in human action recognition, providing an exciting new way to fuse rich structural information for nodes residing in different parts of a graph. In human action recognition, current works introduce a dynamic graph generation mechanism to better capture the underlying semantic skeleton connections and thus improves the performance. In this paper, we provide an orthogonal way to explore the underlying connections. Instead of introducing an expensive dynamic graph generation paradigm, we build a more efficient GCN on a Riemann manifold, which we think is a more suitable space to model the graph data, to make the extracted representations fit the embedding matrix. Specifically, we present a novel spatial-temporal GCN (ST-GCN) architecture which is defined via the Poincaré geometry such that it is able to better model the latent anatomy of the structure data. To further explore the optimal projection dimension in the Riemann space, we mix different dimensions on the manifold and provide an efficient way to explore the dimension for each ST-GCN layer. With the final resulted architecture, we evaluate our method on two current largest scale 3D datasets, i.e., NTU RGB+D and NTU RGB+D 120. The comparison results show that the model could achieve a superior performance under any given evaluation metrics with only 40% model size when compared with the previous best GCN method, which proves the effectiveness of our model. Wei Peng 0009, Jingang Shi, Zhaoqiang Xia, Guoying Zhao 0001 |
ACM Multimedia | 1 |
| 2020 | Revealing the Invisible With Model and Data Shrinking for Composite-Database Micro-Expression RecognitionabstractComposite-database micro-expression recognition is attracting increasing attention as it is more practical for real-world applications. Though the composite database provides more sample diversity for learning good representation models, the important subtle dynamics are prone to disappearing in the domain shift such that the models greatly degrade their performance, especially for deep models. In this paper, we analyze the influence of learning complexity, including input complexity and model complexity, and discover that the lower-resolution input data and shallower-architecture model are helpful to ease the degradation of deep models in composite-database task. Based on this, we propose a recurrent convolutional network (RCN) to explore the shallower-architecture and lower-resolution input data, shrinking model and input complexities simultaneously. Furthermore, we develop three parameter-free modules (i.e., wide expansion, shortcut connection and attention unit) to integrate with RCN without increasing any learnable parameters. These three modules can enhance the representation ability in various perspectives while preserving not-very-deep architecture for lower-resolution data. Besides, three modules can further be combined by an automatic strategy (a neural architecture search strategy) and the searched architecture becomes more robust. Extensive experiments on the MEGC2019 dataset (composited of existing SMIC, CASME II and SAMM datasets) have verified the influence of learning complexity and shown that RCNs with three modules and the searched combination outperform the state-of-the-art approaches. Zhaoqiang Xia, Wei Peng 0009, Huai-Qian Khor, Xiaoyi Feng, Guoying Zhao 0001 |
IEEE Trans. Image Process. | 2 |
| 2019 | A Boost in Revealing Subtle Facial Expressions: A Consolidated Eulerian FrameworkabstractFacial Micro-expression Recognition (MER) distinguishes the underlying emotional states of spontaneous subtle facial expressions. Automatic MER is challenging because that the intensity of subtle facial muscle movement is extremely low and the duration of ME is transient.Recent works adopt motion magnification or temporal interpolation to resolve these issues. Nevertheless, existing works divide them into two separate modules due to their non-linearity. Though such operation eases the difficulty in implementation, it ignores their underlying connections and thus results in inevitable losses in both accuracy and speed. Instead, in this paper, we propose a consolidated Eulerian framework to reveal the subtle facial movements. It expands the temporal duration and amplifies the muscle movements in micro-expressions simultaneously. Compared to existing approaches, the proposed method can not only process ME clips more efficiently but also make subtle ME movements more distinguishable. Experiments on two public MER databases indicate that our model outperforms the state-of-the-art in both speed and accuracy. Wei Peng 0009, Xiaopeng Hong, Yingyue Xu, Guoying Zhao 0001 |
FG | 1 |
| 2019 | Remote Heart Rate Measurement From Highly Compressed Facial Videos: An End-to-End Deep Learning Solution With Video EnhancementabstractRemote photoplethysmography (rPPG), which aims at measuring heart activities without any contact, has great potential in many applications (e.g., remote healthcare). Existing rPPG approaches rely on analyzing very fine details of facial videos, which are prone to be affected by video compression. Here we propose a two-stage, end-to-end method using hidden rPPG information enhancement and attention networks, which is the first attempt to counter video compression loss and recover rPPG signals from highly compressed videos. The method includes two parts: 1) a Spatio-Temporal Video Enhancement Network (STVEN) for video enhancement, and 2) an rPPG network (rPPGNet) for rPPG signal recovery. The rPPGNet can work on its own for robust rPPG measurement, and the STVEN network can be added and jointly trained to further boost the performance especially on highly compressed videos. Comprehensive experiments are performed on two benchmark datasets to show that, 1) the proposed method not only achieves superior performance on compressed videos with high-quality videos pair, 2) it also generalizes well on novel data with only compressed videos available, which implies the promising potential for real-world applications. Zitong Yu, Wei Peng 0009, Xiaopeng Hong, Guoying Zhao 0001 |
ICCV | 2 |
| 2019 | Video Action Recognition Via Neural Architecture SearchingabstractDeep neural networks have achieved great success for video analysis and understanding. However, designing a high-performance neural architecture requires substantial efforts and expertise. In this paper, we make the first attempt to let algorithm automatically design neural networks for video action recognition tasks. Specifically, a spatio-temporal network is developed in a differentiable space modeled by a directed acyclic graph, thus a gradient-based strategy can be performed to search an optimal architecture. Nonetheless, it is computationally expensive, since the computational burden to evaluate each architecture candidate is still heavy. To alleviate this issue, we, for the video input, introduce a temporal segment approach to reduce the computational cost without losing global video information. For the architecture, we explore in an efficient search space by introducing pseudo 3D operators. Experiments show that, our architecture outperforms popular neural architectures, under the training from scratch protocol, on the challenging UCF101 dataset, surprisingly, with only around one percentage of parameters of its manual-design counterparts. Wei Peng 0009, Xiaopeng Hong, Guoying Zhao 0001 |
ICIP | 1 |