EDBT 2026 Demo / reviewers in the wild / expert
Weibing Zhao
dblp:239/7036
· DBLP profile ↗
22ranked-venue papers
2as first author
20since 2021 · last 2026
0000-0002-2819-990XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Computer networks · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Systems, architecture and hardware · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Physical Embedding for Radio Map Construction
Zheng Xing 0001, Liang Xie 0011, Tao Guo 0004, Qi Tan 0003, Qihua Zhou, Weibing Zhao, Ruikang Zhong, Laizhong Cui |
ICC | 7 |
| 2026 | Blind Radio Map Construction via Topology-Guided Manifold LearningabstractConstructing radio maps traditionally requires extensive site surveys with precise location labels, resulting in costly and time-consuming calibration. Conventional approaches derive labels from inertial measurement units (IMUs), but are constrained by device-level access permissions and the need for pre-installed IMU hardware. In this paper, we present a calibration-free radio map construction method that relies solely on Channel State Information (CSI) measurements, thereby obviating the need for location labels. Our key insight is to embed CSI data into a two-dimensional geographic space using a neural network, without any label information. The primary challenge is to preserve the real-world topology in the embedded space. To address this issue, we propose a novelTopology-Guided Manifold Learning (TGML)approach that learns a low-dimensional embedding through self-supervision based on its mapping to a topological map in the geographic space. Specifically, we introduce an embedding network that learns locally smoothed proximity, thus creating a compact low-dimensional representation in the latent space. We further employ a regularized transport method with a differentiable Sinkhorn distance to establish an optimal mapping between the latent and geographic spaces. We use the resulting mapping deviation to supervise the training of the embedding network, enabling continuous refinement through self-supervision. Experiments in an office environment show that TGML achieves an average localization error of 2.37m and a relative beam estimation error of 6.68%, outperforming state-of-the-art methods. Zheng Xing 0001, Jinfeng Xu 0003, Shuo Yang 0011, Ruikang Zhong, Weibing Zhao, Edith C. H. Ngai |
IEEE Internet Things J. | 8 |
| 2026 | Clustering structure identification with ordering graph
Zheng Xing 0001, Weibing Zhao |
Knowl. Based Syst. | 2 |
| 2026 | Temporal Visual Semantics-Induced Human Motion Understanding With Large Language ModelsabstractUnsupervised human motion segmentation (HMS) can be effectively achieved using subspace clustering techniques. However, traditional methods overlook the role of temporal semantic exploration in HMS. This paper explores the use of temporal vision semantics (TVS) derived from human motion sequences, leveraging the image-to-text capabilities of a large language model (LLM) to enhance subspace clustering performance. The core idea is to extract textual motion information from consecutive frames via LLM and incorporate this learned information into the subspace clustering framework. The primary challenge lies in learning TVS from human motion sequences using LLM and incorporating this information into subspace clustering. To address this, we determine whether consecutive frames depict the same motion by querying the LLM and subsequently learn temporal neighboring information based on its response. We then develop a TVS-integrated subspace clustering approach, incorporating subspace embedding with a temporal regularizer that induces each frame to share similar subspace embeddings with its temporal neighbors. Additionally, segmentation is performed based on subspace embedding with a temporal constraint that induces the grouping of each frame with its temporal neighbors. We also introduce a feedback-enabled framework that continuously optimizes subspace embedding based on the segmentation output. Experimental results demonstrate that the proposed method outperforms existing state-of-the-art approaches on four benchmark human motion datasets. Zheng Xing 0001, Weibing Zhao |
IEEE Trans. Image Process. | 2 |
| 2026 | Topology-Aware Embedding Network for Label-Free Radio Map Construction
Zheng Xing 0001, Weibing Zhao, Mengru Wu, Wenjie Liu 0017, Cheng Zeng 0002, Huijun Xing, Ruimao Zhang |
IEEE Trans. Wirel. Commun. | 2 |
| 2025 | HARM3-Fusion: Hierarchical Attentional Representation Learning of Multi-modal, Multi-temporal, and Multi-sequence Fusion for Pathological Complete Response Prediction of Head and Neck Squamous Cell Carcinoma
Jianye Wang, Zhiying Gong, Lingjie Yang, Yimeng Fan, Xiaohui Duan, Weibing Zhao |
MICCAI (10) | 10 |
| 2025 | Cervical-RG: Automated Cervical Cancer Report Generation from 3D Multi-sequence MRI via CoT-Guided Hierarchical Experts
Yimeng Fan, Zhaoyi Zhan, Zheng Xing 0001, Xiaohui Duan, Weibing Zhao |
MICCAI (5) | 12 |
| 2025 | Block-diagonal structure learning for subspace clustering
Zheng Xing 0001, Weibing Zhao |
Expert Syst. Appl. | 2 |
| 2025 | Calibration-Free Indoor Positioning via Regional Channel TracingabstractThe traditional construction of radio maps demands extensive radio measurement data accompanied by precise location labels, thus necessitating considerable calibration. This article presents a calibration-free radio map construction method that relies solely on received signal strength (RSS) measurements, obviating the need for location labels. In this method, multiple mobile users traverse an indoor environment equipped with WiFi sensors, generating RSS measurement sequences. We estimate RSS collection locations along trajectories using prior knowledge of the layout, principles of signal propagation, and user mobility patterns. The challenge of estimating these locations based purely on RSS measurements is met by conforming to the intricate indoor layout and adhering to signal propagation models. Traditional calibration-free approaches often depend on inertial measurement units (IMUs) for real-time location estimation, but such reliance on IMUs is both cumbersome and encumbered by privacy concerns. We introduce a regional channel tracing (RCT) method, employing a signal subspace model and a subspace segmental clustering algorithm to classify RSS measurements and map them to regions. This method further includes a location tagging technique that, given the region labels, estimates RSS collection locations and regional path-loss models. Empirical results in an office setting demonstrate that our RCT-based radio map achieves performance comparable to advanced IMU-assisted methods. Zheng Xing 0001, Weibing Zhao |
IEEE Internet Things J. | 2 |
| 2025 | Trajectory Map-Matching in Urban Road Networks Based on RSS MeasurementsabstractThe widespread deployment of wireless communication networks has catalyzed significant advancements in utilizing signal channs to address real-world challenges, such as vehicle trajectory reconstruction (VTR), drone trajectory planning, and network optimization. Existing methods primarily utilize time-difference-of-arrival (TDoA) measurements for vehicle localization. However, these methods require specialized decoding receivers capable of deciphering communication protocols, leading to increased application costs. received signal strength (RSS), a measure of wireless signal strength, can be recorded by any standard communication device, thus allowing RSS-based VTR to benefit from cost-effectiveness. Nevertheless, the inherently noisy and sporadic nature of RSS poses significant challenges for accurately reconstructing vehicle trajectories. This paper aims to utilize RSS measurements to reconstruct vehicle trajectories within a road network. We constrain the trajectories to comply with signal propagation rules and vehicle mobility constraints, thereby mitigating the impact of the noisy and sporadic nature of RSS data on the accuracy of trajectory reconstruction. The primary challenge involves exploiting latent spatial-temporal correlations within the noisy and sporadic RSS data while navigating the complex road network. To overcome these challenges, we develop an hidden Markov model (HMM)-based RSS embedding (HRE) technique that utilizes alternating optimization to search for the vehicle trajectory based on RSS measurements. This model effectively captures the spatial-temporal relationships among RSS measurements, while a road graph model ensures compliance with network pathways. Additionally, we introduce a maximum speed-constrained rough trajectory estimation (MSR) method to effectively guide the proposed alternating optimization procedure, ensuring that the proposed HRE method rapidly converges to a favorable local solution. The proposed method is validated using real RSS measurements from 5G NR networks in Chengdu and Shenzhen, China. The experimental results demonstrate that the proposed approach significantly outperforms state-of-the-art methods, even with limited RSS data. Zheng Xing 0001, Weibing Zhao |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | Unsupervised Action Segmentation via Fast Learning of Semantically Consistent ActomsabstractAction segmentation serves as a pivotal component in comprehending videos, encompassing the learning of a sequence of semantically consistent action units known as actoms. Conventional methodologies tend to require a significant consumption of time for both training and learning phases. This paper introduces an innovative unsupervised framework for action segmentation in video, characterized by its fast learning capability and absence of mandatory training. The core idea involves splitting the video into distinct actoms, which are then merging together based on shared actions. The key challenge here is to prevent the inadvertent creation of singular actoms that attempt to represent multiple actions during the splitting phase. Additionally, it is crucial to avoid situations where actoms associated with the same action are incorrectly grouped into multiple clusters during the merging phase. In this paper, we present a method for calculating the similarity between adjacent frames under a subspace assumption. Then, we employ a local minimum searching procedure, which effectively splits the video into coherent actoms aligned with their semantic meaning and provides us an action segmentation proposal. Subsequently, we calculate a spatio-temporal similarity between actoms, followed by developing a merging process to merge actoms representing identical actions within the action segmentation proposals. Our approach is evaluated on four benchmark datasets, and the results demonstrate that our method achieves state-of-the-art performance. Besides, our method also achieves the optimal balance between accuracy and learning time when compared to existing unsupervised techniques. Code is available at https://github.com/y66y/SaM. Zheng Xing 0001, Weibing Zhao |
AAAI | 2 |
| 2024 | A Delta-Sigma-Based Computing-In-Memory Macro Targeting Edge ComputationabstractMany applications of machine learning (ML) have been integrated into edge devices with their low communication latency. In edge computation, the reprocessing of redundant data results in considerable energy waste. The prior research utilized a digital-delta-digital-sigma computing-in-memory (CIM) scheme to mitigate this redundancy. However, the 7-bit LSB-first ADC resulting from the near-zero-mean output distribution led to excessive area and latency overhead. The following digital adder further induced power consumption and latency. We propose a digital-delta-analog-sigma CIM macro incorporating an analog sigma converter (SC) for edge computation, involving a switch-capacitor integrator with a floating inverter amplifier (FIA) and a quantizer. The increased analog swing of the sigma integrator leads to the expanded output distribution, thereby maintaining comparable accuracy with a relaxed quantizer resolution. The simulation demonstrates that our strategy contributes to a 57.5% reduction in latency, a resolution decrease of 2 bits, and better energy efficiency. These improvements can potentially enhance energy efficiency and computational speed in edge computation devices. Ka-Fai Un, Mingqiang Guo, Liang Qi 0002, Dengke Xu, Weibing Zhao, Rui Paulo Martins, Franco Maloberti, Sai-Weng Sin |
ISCAS | 6 |
| 2024 | Segmentation and Completion of Human Motion Sequence via Temporal Learning of Subspace Variety ModelabstractSubspace-based models have been extensively employed in unsupervised segmentation and completion of human motion sequence (HMS). However, existing approaches often neglect the incorporation of temporal priors embedded in HMS, resulting in suboptimal results. This paper presents a subspace variety model for HMS, along with an innovative Temporal Learning of Subspace Variety Model (TL-SVM) method for enhanced segmentation and completion in HMS. The key idea is to segment incomplete HMS into motion clusters and extracting the subspace features of each motion through the temporal learning of the subspace variety model. Subsequently, the HMS is completed based on the extracted subspace features. Thus, the main challenge is to learn the subspace variety model with temporal priors when confronted with missing entries. To tackle this, the paper develops a spatio-temporal assignment consistency (STAC) constraint for the subspace variety model, leveraging temporal priors embedded in HMS. In addition, a subspace clustering approach under the STAC constraint is proposed to learn the subspace variety model by extracting subspace features from HMS and segmenting HMS into motion clusters alternatively. The proposed subspace clustering model can also handle missing entries with theoretical guarantees. Furthermore, the missing entries of HMS are completed by minimizing the distance between each human motion frame and its corresponding subspace. Extensive experimental results, along with comparisons to state-of-the-art methods on four benchmark datasets, underscore the advantages of the proposed method. Zheng Xing 0001, Weibing Zhao |
IEEE Trans. Image Process. | 2 |
| 2024 | Block-Diagonal Guided DBSCAN ClusteringabstractCluster analysis constitutes a pivotal component of database mining, with DBSCAN being one of the most extensively employed algorithms in this domain. Nevertheless, DBSCAN is encumbered by several limitations, including challenges in processing high-dimensional datasets, a pronounced sensitivity to input parameters, and inconsistencies in generating reliable clustering outcomes. This paper presents a refined version of DBSCAN that utilizes the block-diagonal property of similarity graphs to enhance the clustering process. The core concept involves the construction of a graph that assesses the similarity among high-dimensional data points, capable of transformation into a block-diagonal form via an unknown permutation. This is followed by a cluster-ordering procedure that establishes the requisite permutation, thereby facilitating the straightforward identification of clustering structures through the recognition of diagonal blocks in the permuted graph. The principal obstacle addressed in this study is the construction of a graph that inherently possesses a block-diagonal structure, the permutation of this graph to actualize such a structure, and the autonomous identification of diagonal blocks within the permuted graph. To surmount these challenges, we initially devise a block-diagonal constrained self-representation model to create a similarity graph that exhibits a block-diagonal form post-permutation. A gradient descent-based methodology is proposed to resolve this problem effectively. Concurrently, we engineer a traversal algorithm, inspired by DBSCAN, that discerns clusters of high density within the graph and generates an enhanced cluster ordering. The attainment of a block-diagonal structure is then realized through permutation aligned with the traversal sequence, laying a robust foundation for both automated and interactive cluster analysis. Moreover, a novel split-and-refine algorithm is introduced to autonomously identify all diagonal blocks within the permuted graph, offering theoretical optimality under specific conditions. Extensive evaluations of our method across twelve rigorous real-world benchmark datasets affirm its superiority over contemporary state-of-the-art clustering techniques. Zheng Xing 0001, Weibing Zhao |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2024 | Spatio-Temporal Contextual Learning for Single Object Tracking on Point CloudsabstractSingle object tracking (SOT) is one of the most active research directions in the field of computer vision. Compared with the 2-D image-based SOT which has already been well-studied, SOT on 3-D point clouds is a relatively emerging research field. In this article, a novel approach, namely, the contextual-aware tracker (CAT), is investigated to achieve a superior 3-D SOT through spatially and temporally contextual learning from the LiDAR sequence. More precisely, in contrast to the previous 3-D SOT methods merely exploiting point clouds in the target bounding box as the template, CAT generates templates by adaptively including the surroundings outside the target box to use available ambient cues. This template generation strategy is more effective and rational than the previous area-fixed one, especially when the object has only a small number of points. Moreover, it is deduced that LiDAR point clouds in 3-D scenes are often incomplete and significantly vary from frame to another, which makes the learning process more difficult. To this end, a novel cross-frame aggregation (CFA) module is proposed to enhance the feature representation of the template by aggregating the features from a historical reference frame. Leveraging such schemes enables CAT to achieve a robust performance, even in the case of extremely sparse point clouds. The experiments confirm that the proposed CAT outperforms the state-of-the-art methods on both the KITTI and NuScenes benchmarks, achieving 3.9% and 5.6% improvements in terms of precision. Jiantao Gao, Xu Yan 0005, Weibing Zhao, Zhen Lyu, Yinghong Liao, Chaoda Zheng |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | CPU: Codebook Lookup Transformer with Knowledge Distillation for Point Cloud UpsamplingabstractPoint clouds produced by 3D scanning are typically sparse, non-uniform, and noisy. Existing upsampling techniques directly learn the mapping from a sparse point set to a dense point set, which is often under-determined and ill-posed. To reduce the uncertainty and ambiguity of the upsampling mapping, this paper proposes a generic three-stage vector-quantization framework, which incorporates a Codebook lookup Transformer and knowledge distillation for Point Cloud Upsampling, named CPU. The proposed CPU reformulates the upsampling task into a relatively determinate code prediction task within a small, discrete proxy space. Since the traditional vector-quantization methods cannot be directly applied to point cloud upsampling scenarios, we introduce a knowledge distillation training scheme that facilitates efficient codebook learning and ensures full utilization of codebook entries. Specifically, we adopt a teacher-student training paradigm to avoid model collapse during codebook learning. In the first stage, we pre-train a vanilla auto-encoder of the dense point set as the teacher model, which provides rich guidance features to ensure sufficient codebook learning. In the second stage, we train a vector-quantized auto-encoder as a student model to capture high-fidelity geometric priors into a learned codebook with the aid of distillation. In the third stage, we propose a Codebook Lookup Transformer to model the global context of the sparse point set and predict the code indices. Then the coarse features of the sparse point set can be quantized and substituted by looking up the indices in the learned codebook. Benefiting from the expressive codebook priors and the distillation training scheme, the proposed CPU outperforms state-of-the-art methods quantitatively and qualitatively. Weibing Zhao, Haiming Zhang 0001, Chaoda Zheng, Xu Yan 0005, Shuguang Cui, Zhen Li 0026 |
ACM Multimedia | 1 |
| 2023 | A 10b 700 MS/s Single-Channel 1b/Cycle SAR ADC Using a Monotonic-Specific Feedback SAR Logic With Power-Delay-Optimized Unbalanced N/P-MOS SizingabstractThis article presents a power-delay-optimized monotonic-specific successive approximation register (SAR) ADC. The SAR feedback loop, comprising the proposed unbalanced N/P-MOS sizing technique, simultaneously reduces the SAR logic delay and the power to overcome the SAR ADC’s speed bottleneck. Benefiting from this technique, the sampling rate of the prototype 10b single channel 1b/cycle SAR ADC reaches 600 and 700 MS/s at 0.9 and 0.95 V supply voltage, while consuming 1.49 and 2.02 mW in 28 nm CMOS, respectively. Moreover, the 10b ADC achieves the SNDR of 56.39 and 56.42-dB at a Nyquist rate input frequency of 600 and 700 MS/s, leading to a Walden FoM of 4.6 and 5.3 fJ/conversion-step, respectively. Mingqiang Guo, Liang Qi 0002, Weibing Zhao, Gang Xiao 0001, Rui Paulo Martins, Sai-Weng Sin |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2022 | A Robust Hybrid CT/DT 0-2 MASH DSM with Passive Noise-Shaping SAR ADCabstractThis paper presents a hybrid CT/DT0-2 multi-stage noise-shaping (MASH) delta-sigma modulator (DSM) with a passive noise-shaping successive approximation register (NSSAR) ADC as the $2^{\mathrm{n}\mathrm{d}}$ stage. The overall architecture is simple and robust. The front-end stage employs the continuous-time (CT) operation to perform coarse quantization and provide inherent anti-aliasing and easy driving. The back-end stage uses a second-order NS-SAR architecture, which excels at PVT robustness, power efficiency, and scaling friendliness. It also results in large relaxation of matching issues between the analog and the digital domains compared with conventional CT-MASH. Behavioral simulation results demonstrate the effectiveness and robustness of the proposed hybrid MASH architecture. Sai-Weng Sin, Liang Qi 0002, Weibing Zhao, Guoxing Wang, Rui Paulo Martins |
ISCAS | 4 |
| 2021 | Box-Aware Feature Enhancement for Single Object Tracking on Point CloudsabstractCurrent 3D single object tracking approaches track the target based on a feature comparison between the target template and the search area. However, due to the common occlusion in LiDAR scans, it is non-trivial to conduct accurate feature comparisons on severe sparse and incomplete shapes. In this work, we exploit the ground truth bounding box given in the first frame as a strong cue to enhance the feature description of the target object, enabling a more accurate feature comparison in a simple yet effective way. In particular, we first propose the BoxCloud, an informative and robust representation, to depict an object using the point-to-box relation. We further design an efficient box-aware feature fusion module, which leverages the aforementioned BoxCloud for reliable feature matching and embedding. Integrating the proposed general components into an existing model P2B [27], we construct a superior box-aware tracker (BAT)1. Experiments confirm that our proposed BAT outperforms the previous state-of-the-art by a large margin on both KITTI and NuScenes benchmarks, achieving a 12.8% improvement in terms of precision while running ∼20% faster. Chaoda Zheng, Xu Yan 0005, Jiantao Gao, Weibing Zhao, Wei Zhang 0001, Zhen Li 0026, Shuguang Cui |
ICCV | 4 |
| 2021 | PointLIE: Locally Invertible Embedding for Point Cloud Sampling and RecoveryabstractPoint Cloud Sampling and Recovery (PCSR) is critical for massive real-time point cloud collection and processing since raw data usually requires large storage and computation. This paper addresses a fundamental problem in PCSR: How to downsample the dense point cloud with arbitrary scales while preserving the local topology of discarded points in a case-agnostic manner (i.e., without additional storage for point relationships)? We propose a novel Locally Invertible Embedding (PointLIE) framework to unify the point cloud sampling and upsampling into one single framework through bi-directional learning. Specifically, PointLIE decouples the local geometric relationships between discarded points from the sampled points by progressively encoding the neighboring offsets to a latent variable. Once the latent variable is forced to obey a pre-defined distribution in the forward sampling path, the recovery can be achieved effectively through inverse operations. Taking the recover-pleasing sampled points and a latent embedding randomly drawn from the specified distribution as inputs, PointLIE can theoretically guarantee the fidelity of reconstruction and outperform state-of-the-arts quantitatively and qualitatively. Weibing Zhao, Xu Yan 0005, Jiantao Gao, Ruimao Zhang, Jiayan Zhang, Zhen Li 0026, Shuguang Cui |
IJCAI | 1 |
| 2020 | Towards Content-Independent Multi-Reference Super-Resolution: Adaptive Pattern Matching and Feature Aggregation
Xu Yan 0005, Weibing Zhao, Kun Yuan 0004, Ruimao Zhang, Zhen Li 0026, Shuguang Cui |
ECCV (25) | 2 |
| 2020 | Characterizing Label Errors: Confident Learning for Noisy-Labeled Image Segmentation
Minqing Zhang, Jiantao Gao, Zhen Lyu, Weibing Zhao, Qin Wang 0011, Weizhen Ding, Sheng Wang 0001, Zhen Li 0026, Shuguang Cui |
MICCAI (1) | 4 |