EDBT 2026 Demo / reviewers in the wild / expert
Weipeng Zhuo
dblp:119/0329
· DBLP profile ↗
16ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0002-1810-7071ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Computer networks · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fine-grained alignment in medical pathology vision-language models via variational distillationabstractPre-trained vision-language (V-L) models exhibit promising performance across various general-domain tasks. However, they fall short in medical pathology due to the critical need for fine-grained semantic alignment, which is essential for distinguishing subtle visual patterns across categories. This limitation is not merely due to domain gaps but stems from the inability to capture detailed, pathology-specific semantics. Previous efforts leveraging large language models (LLMs) or cross-modal training often introduce redundant or ambiguous cues, ultimately weakening generalization. To explicitly enhance fine-grained alignment, we propose a Variational Distillation framework tailored for Medical Pathology V-L models. This method introduces a dual-loop optimization mechanism that jointly distills and aligns semantic signals from both textual inputs and external LLM knowledge. Specifically, we use variational latent distributions to model semantic ambiguity and apply a KL-based loss to reduce differences between signals. This encourages the model to retain robust and generalizable features, enabling improved sensitivity to subtle semantic variations critical in pathology image understanding. During cross-modal alignment, the proposed method further amplifies modality-shared semantics while suppressing modality-specific noise and task-irrelevant factors, yielding more precise and pathology-aware image-text matching. Extensive experiments on five pathology benchmarks across three settings, including class generalization, few-shot learning, and cross-organ transfer, demonstrate that the proposed method consistently outperforms the existing approaches. Runlin Huang, Haowei Lin, Weipeng Zhuo, Yiu-Ming Cheung, Hongmin Cai, Weifeng Su |
Pattern Recognit. | 3 |
| 2025 | Improving Fairness in Skin Cancer Diagnosis Via Feature Pattern Separating and Diverse Objective OptimizationabstractMedical AI models have achieved remarkable progress in various tasks, including medical image classification and segmentation. However, a critical issue frequently overlooked in clinical applications is the fairness of these models across different subgroups. Existing data collection strategies typically prioritize balancing disease categories while neglecting the importance of subgroup balance based on factors such as age and gender. This imbalance leads to the model's inconsistent performance across subgroups, thereby limiting its clinical applicability. This issue is particularly pronounced in skin cancer diagnosis due to the entanglement and complexity of feature patterns in skin cancer images. In this article, we approach the problem from a new perspective: making the feature patterns extracted by the model more independent and enhancing the representational ability of subgroups. We propose a novel framework that enhances feature independence across channels and improves subgroup representation. Central to our design is the Feature Pattern Separating Module (FPSM), combined with dual subgroupspecific classifiers and an auxiliary subgroup-type predictor. A diverse objective optimization model was introduced to guide the joint optimization of classification accuracy and subgroup fairness. Additionally, we proposed a novel classification loss function to dynamically balance the loss between subgroups. Experiments on two public skin cancer datasets demonstrate that our method improves both overall performance and fairness across subgroups, outperforming existing methods and offering a promising solution for fairer medical AI. Weipeng Zhuo, Zhewei Su, Wentao Fan 0001, Hongmin Cai, Weifeng Su |
BIBM | 2 |
| 2025 | SDT-GNN: Streaming-Based Distributed Training Framework for Graph Neural Networks
Xin Huang 0020, Weipeng Zhuo, Minh Phu Vuong, Shiju Li 0001, Jongryool Kim, Bradley Rees, Chul-Ho Lee |
IEEE Big Data | 2 |
| 2025 | MTM: A Multi-Scale Token Mixing Transformer for Irregular Multivariate Time Series ClassificationabstractIrregular multivariate time series (IMTS) is characterized by the lack of synchronized observations across its different channels.In this paper, we point out that this channel-wise asynchrony can lead to poor channel-wise modeling of existing deep learning methods.To overcome this limitation, we propose MTM, a multi-scale token mixing transformer for the classification of IMTS.We find that the channel-wise asynchrony can be alleviated by down-sampling the time series to coarser timescales, and propose to incorporate a masked concat pooling in MTM that gradually down-samples IMTS to enhance the channel-wise attention modules.Meanwhile, we propose a novel channel-wise token mixing mechanism which proactively chooses important tokens from one channel and mixes them with other channels, to further boost the channel-wise learning of our model.Through extensive experiments on real-world datasets and comparison with state-of-the-art methods, we demonstrate that MTM consistently achieves the best performance on all the benchmarks, with improvements of up to 3.8% in AUPRC for classification. Shuhan Zhong, Weipeng Zhuo, Sizhe Song, Guanyao Li, Zhongyi Yu, Shueng-Han Gary Chan |
KDD (2) | 2 |
| 2025 | ADA: An Adaptive Augmentation Framework for Single-Source Domain Generalization in Medical Image Segmentation
Runlin Huang, Hongmin Cai, Weipeng Zhuo, Shangyan Cai, Haowei Lin, Wentao Fan 0001, Weifeng Su |
MICCAI (10) | 3 |
| 2025 | PathoPrompt: Cross-Granular Semantic Alignment for Medical Pathology Vision-Language Models
Runlin Huang, Haohui Liang, Hongmin Cai, Weipeng Zhuo, Wentao Fan 0001, Weifeng Su |
MICCAI (7) | 4 |
| 2025 | TAMI: Taming Heterogeneity in Temporal Interactions for Temporal Graph Link PredictionabstractTemporal graph link prediction aims to predict future interactions between nodes in a graph based on their historical interactions, which are encoded in node embeddings. We observe that heterogeneity naturally appears in temporal interactions, e.g., a few node pairs can make most interaction events, and interaction events happen at varying intervals. This leads to the problems of ineffective temporal information encoding and forgetting of past interactions for a pair of nodes that interact intermittently for their link prediction. Existing methods, however, do not consider such heterogeneity in their learning process, and thus their learned temporal node embeddings are less effective, especially when predicting the links for infrequently interacting node pairs. To cope with the heterogeneity, we propose a novel framework called TAMI, which contains two effective components, namely log time encoding function (LTE) and link history aggregation (LHA). LTE better encodes the temporal information through transforming interaction intervals into more balanced ones, and LHA prevents the historical interactions for each target node pair from being forgotten. State-of-the-art temporal graph neural networks can be seamlessly and readily integrated into TAMI to improve their effectiveness. Experiment results on 13 classic datasets and three newest temporal graph benchmark (TGB) datasets show that TAMI consistently improves the link prediction performance of the underlying models in both transductive and inductive settings. Our code is available at https://github.com/Alleinx/TAMI_temporal_graph. Zhongyi Yu, Jianqiu Wu, Shuhan Zhong, Weifeng Su, Chul-Ho Lee, Weipeng Zhuo |
NeurIPS | 7 |
| 2024 | A Multi-Scale Decomposition MLP-Mixer for Time Series AnalysisabstractTime series data, including univariate and multivariate ones, are characterized by unique composition and complex multi-scale temporal variations. They often require special consideration of decomposition and multi-scale modeling to analyze. Existing deep learning methods on this best fit to univariate time series only, and have not sufficiently considered sub-series modeling and decomposition completeness. To address these challenges, we propose MSD-Mixer, a M ulti- S cale D ecomposition MLP- Mixer , which learns to explicitly decompose and represent the input time series in its different layers. To handle the multi-scale temporal patterns and multivariate dependencies, we propose a novel temporal patching approach to model the time series as multi-scale patches, and employ MLPs to capture intra- and inter-patch variations and channel-wise correlations. In addition, we propose a novel loss function to constrain both the mean and the autocorrelation of the decomposition residual for better decomposition completeness. Through extensive experiments on various real-world datasets for five common time series analysis tasks, we demonstrate that MSD-Mixer consistently and significantly outperforms other state-of-the-art algorithms with better efficiency. Shuhan Zhong, Sizhe Song, Weipeng Zhuo, Guanyao Li, Yang Liu 0278, Shueng-Han Gary Chan |
Proc. VLDB Endow. | 3 |
| 2024 | Online Path Description Learning Based on IMU Signals From IoT DevicesabstractA user's movement path can be precisely and concisely described as a concatenation of straight lines having the user's turns as their end points. Learning such a path description or representation from inertial measurement unit (IMU) sensors enables various mobile and IoT applications, as it allows efficient processing of the movement path data. It is, however, non-trivial to learn a succinct yet accurate path description from IMU sensor readings in the mobile device of a moving useron the flydue to the dynamically changing behaviors and the technical difficulty in detecting the user's turns. We propose PATHLIT, a novel online path description learning system based on IMU signals. PATHLIT learns position vectors of a user from IMU sensor readings by our custom-made self-attention network model. Once each position vector is learned, PATHLIT also decides whether or not to take it as a part of the resulting path description by our efficient online algorithm developed under the minimum description length principle, which essentially detects the user's turns along the path. We conduct extensive experiments on two large datasets. The experiment results show that PATHLIT achieves superior performance over state-of-the-art algorithms by up to 50% in absolute trajectory error using only 15% of trajectory data points. Weipeng Zhuo, Shiju Li 0001, Tianlang He, Shueng-Han Gary Chan, Sangtae Ha, Chul-Ho Lee |
IEEE Trans. Mob. Comput. | 1 |
| 2023 | Run, Don't Walk: Chasing Higher FLOPS for Faster Neural NetworksabstractTo design fast neural networks, many works have been focusing on reducing the number of floating-point operations (FLOPs). We observe that such reduction in FLOPs, however, does not necessarily lead to a similar level of re-duction in latency. This mainly stems from inefficiently low floating-point operations per second (FLOPS). To achieve faster networks, we revisit popular operators and demonstrate that such low FLOPS is mainly due to frequent memory access of the operators, especially the depthwise con-volution. We hence propose a novel partial convolution (PConv) that extracts spatial features more efficiently, by cutting down redundant computation and memory access simultaneously. Building upon our PConv, we further propose FasterNet, a new family of neural networks, which attains substantially higher running speed than others on a wide range of devices, without compromising on accuracy for various vision tasks. For example, on ImageNet-lk, our tiny FasterNet-TO is 2.8×, 3.3×, and 2.4× faster than MobileViT-XXS on GPU, CPU, and ARM processors, respectively, while being 2.9% more accurate. Our large FasterNet-L achieves impressive 83.5% top-1 accuracy, on par with the emerging Swin-B, while having 36% higher inference throughput on GPU, as well as saving 37% compute time on CPU. Code is available at https://github.com/JierunChen/FasterNet. Jierun Chen, Shiu-Hong Kao, Hao He 0011, Weipeng Zhuo, Song Wen 0001, Chul-Ho Lee, Shueng-Han Gary Chan |
CVPR | 4 |
| 2023 | FIS-ONE: Floor Identification System with One Label for Crowdsourced RF SignalsabstractFloor labels of crowdsourced RF signals are crucial for many smart-city applications, such as multi-floor indoor localization, geofencing, and robot surveillance. To build a prediction model to identify the floor number of a new RF signal upon its measurement, conventional approaches using the crowdsourced RF signals assume that at least few labeled signal samples are available on each floor. In this work, we push the envelope further and demonstrate that it is technically feasible to enable such floor identification with only one floor-labeled signal sample on the bottom floor while having the rest of signal samples unlabeled. We propose FIS-ONE, a novel floor identification system with only one labeled sample. FIS-ONE consists of two steps, namely signal clustering and cluster indexing. We first build a bipartite graph to model the RF signal samples and obtain a latent representation of each node (each signal sample) using our attention-based graph neural network model so that the RF signal samples can be clustered more accurately. Then, we tackle the problem of indexing the clusters with proper floor labels, by leveraging the observation that signals from an access point can be detected on different floors, i.e., signal spillover. Specifically, we formulate a cluster indexing problem as a combinatorial optimization problem and show that it is equivalent to solving a traveling salesman problem, whose (near-)optimal solution can be found efficiently. We have implemented FIS-ONE and validated its effectiveness on the Microsoft dataset and in three large shopping malls. Our results show that FIS- ONE outperforms other baseline algorithms significantly, with up to 23 % improvement in adjusted rand index and 25% improvement in normalized mutual information using only one floor-labeled signal sample. Weipeng Zhuo, Ka Ho Chiu, Jierun Chen, Shueng-Han Gary Chan, Sangtae Ha, Chul-Ho Lee |
ICDCS | 1 |
| 2023 | Semi-supervised Learning with Network Embedding on Ambient RF Signals for Geofencing ServicesabstractIn applications such as elderly care, dementia anti-wandering and pandemic control, it is important to ensure that people are within a predefined area for their safety and well-being. We propose GEM, a practical, semi-supervised Geofencing system with network EMbedding, which is based only on ambient radio frequency (RF) signals. GEM models measured RF signal records as a weighted bipartite graph. With access points on one side and signal records on the other, it is able to precisely capture the relationships between signal records. GEM then learns node embeddings from the graph via a novel bipartite network embedding algorithm called BiSAGE, based on a Bipartite graph neural network with a novel bi-level SAmple and aggreGatE mechanism and non-uniform neighborhood sampling. Using the learned embeddings, GEM finally builds a one-class classification model via an enhanced histogram-based algorithm for in-out detection, i.e., to detect whether the user is inside the area or not. This model also keeps on improving with newly collected signal records. We demonstrate through extensive experiments in diverse environments that GEM shows state-of-the-art performance with up to 34% improvement in F-score. BiSAGE in GEM leads to a 54% improvement in F-score, as compared to the one without BiSAGE. Weipeng Zhuo, Ka Ho Chiu, Jierun Chen, Jiajie Tan, Edmund Sumpena, Shueng-Han Gary Chan, Sangtae Ha, Chul-Ho Lee |
ICDE | 1 |
| 2022 | TVConv: Efficient Translation Variant Convolution for Layout-aware Visual ProcessingabstractAs convolution has empowered many smart applications, dynamic convolution further equips it with the ability to adapt to diverse inputs. However, the static and dynamic convolutions are either layout-agnostic or computation-heavy, making it inappropriate for layout-specific applications, e.g., face recognition and medical image segmentation. We observe that these applications naturally exhibit the characteristics of large intra-image (spatial) variance and small cross-image variance. This observation motivates our efficient translation variant convolution (TVConv) for layout-aware visual processing. Technically, TVConv is composed of affinity maps and a weight-generating block. While affinity maps depict pixel-paired relationships gracefully, the weight-generating block can be explicitly over-parameterized for better training while maintaining efficient inference. Although conceptually simple, TVConv significantly improves the efficiency of the convolution and can be readily plugged into various network architectures. Extensive experiments on face recognition show that TVConv reduces the computational cost by up to 3.1 × and improves the corresponding throughput by 2.3× while maintaining a high accuracy compared to the depthwise convolution. Moreover, for the same computation cost, we boost the mean accuracy by up to 4.21%. We also conduct experiments on the optic disc/cup segmentation task and obtain better generalization performance, which helps mitigate the critical data scarcity issue. Code is available at https://github.com/JierunChen/TVConv. Jierun Chen, Tianlang He, Weipeng Zhuo, Sangtae Ha, Shueng-Han Gary Chan |
CVPR | 3 |
| 2022 | GRAFICS: Graph Embedding-based Floor Identification Using Crowdsourced RF SignalsabstractWe study the problem of floor identification for radiofrequency (RF) signal samples obtained in a crowdsourced manner, where the signal samples are highly heterogeneous and most samples lack their floor labels. We propose GRAFICS, a graph embedding-based floor identification system. GRAFICS first builds a highly versatile bipartite graph model, having APs on one side and signal samples on the other. GRAFICS then learns the low-dimensional embeddings of signal samples via a novel graph embedding algorithm named E-LINE. GRAFICS finally clusters the node embeddings along with the embeddings of a few labeled samples through a proximity-based hierarchical clustering, which eases the floor identification of every new sample. We validate the effectiveness of GRAFICS based on two large-scale datasets that contain RF signal records from 204 buildings in Hangzhou, China, and five buildings in Hong Kong. Our experiment results show that GRAFICS achieves highly accurate prediction performance with only a few labeled samples (96% in both micro- and macro-F scores) and significantly outperforms several state-of-the-art algorithms (by about 45% improvement in micro-F score and 53% in macro-F score). Weipeng Zhuo, Ka Ho Chiu, Shiju Li 0001, Sangtae Ha, Chul-Ho Lee, Shueng-Han Gary Chan |
ICDCS | 1 |
| 2022 | Tackling Multipath and Biased Training Data for IMU-Assisted BLE Proximity DetectionabstractProximity detection is to determine whether an IoT receiver is within a certain distance from a signal transmitter. Due to its low cost and high popularity, Bluetooth low energy (BLE) has been used to detect proximity based on the received signal strength indicator (RSSI). To address the fact that RSSI can be markedly influenced by device carriage states, previous works have incorporated RSSI with inertial measurement unit (IMU) using deep learning. However, they have not sufficiently accounted for the impact of multipath. Furthermore, due to the special setup, the IMU data collected in the training process may be biased, which hampers the system’s robustness and generalizability. This issue has not been studied before.We propose PRID, an IMU-assisted BLE proximity detection approach robust against RSSI fluctuation and IMU data bias. PRID histogramizes RSSI to extract multipath features and uses carriage state regularization to mitigate overfitting due to IMU data bias. We further propose PRID-lite based on a binarized neural network to substantially cut memory requirements for resource-constrained devices. We have conducted extensive experiments under different multipath environments, data bias levels, and a crowdsourced dataset. Our results show that PRID significantly reduces false detection cases compared with the existing arts (by over 50%). PRID-lite further reduces over 90% PRID model size and extends 60% battery life, with a minor compromise in accuracy (7%). Tianlang He, Jiajie Tan, Weipeng Zhuo, Maximilian Printz, Shueng-Han Gary Chan |
INFOCOM | 3 |
| 2012 | Error Modeling and Estimation Fusion for Indoor LocalizationabstractThere has been much interest in offering multimedia location-based service (LBS) to indoor users (e.g., sending video/audio streams according to user locations). Offering good LBS largely depends on accurate indoor localization of mobile stations (MSs). To achieve that, in this paper we first model and analyze the error characteristics of important indoor localization schemes, using Radio Frequency Identification (RFID) and Wi-Fi. Our models are simple to use, capturing important system parameters and measurement noises, and quantifying how they affect the accuracies of the localization. Given that there have been many indoor localization techniques deployed, an MS may receive simultaneously multiple co-existing estimations on its location. Equipped with the understanding of location errors, we then investigate how to optimally combine, or fuse, all the co-existing estimations of an MS's location. We present computationally-efficient closed-form expressions to fuse the outputs of the estimators. Simulation and experimental results show that our fusion technique achieves higher location accuracy in spite of location errors in the estimators. Weipeng Zhuo, Bo Zhang 0026, Shueng-Han Gary Chan, Edward Y. Chang |
ICME | 1 |