VLDB 2026 Research / reviewers in the wild / expert
Zhizhen Li
dblp:297/8002
· DBLP profile ↗
13ranked-venue papers
7as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 9 · 7 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MMSense: Adapting Vision-based Foundation Model for Multi-task Multi-modal Wireless SensingabstractLarge AI models have been widely adopted in wireless communications for channel modeling, beamforming, and resource optimization. However, most existing efforts remain limited to single-modality inputs and channel-specific objec- tives, overlooking the broader potential of large foundation models for unified wireless sensing. To bridge this gap, we propose MMSense, a multi-modal, multi-task foundation model that jointly addresses channel-centric, environment-aware, and human-centered sensing. Our framework integrates image, radar, LiDAR, and textual data by transforming them into vision- compatible representations, enabling effective cross-modal align- ment within a unified feature space. A modality gating mecha- nism adaptively fuses these representations, while a vision-based large language model backbone enables unified feature align- ment and instruction-driven task adaptation. Furthermore, task- specific sequential attention and uncertainty-based loss weighting mechanisms enhance cross-task generalization. Experiments on real wireless scenario datasets show that our approach outper- forms both task-specific and large-model baselines, confirming its strong generalization across heterogeneous sensing tasks. Zhizhen Li, Xuanhao Luo, Xueren Ge, Longyu Zhou, Xingqin Lin, Yuchen Liu 0001 |
ICC | 1 |
| 2026 | BiOVQL: Brain-inspired One-stage Egocentric Visual Query LocalizationabstractVisual query localization (VQL) is pivotal for constructing episodic memory from egocentric videos. However, current methods often rely on computationally intensive attention mechanisms and rigid one-shot regression, which inherently struggle to model uncertainty and exhibit limited adaptability to degraded query appearances. This contrasts sharply with the human brain’s selective encoding and iterative hypothesis verification processes for episodic memory. Inspired by the human brain’s ability to selectively filter irrelevant information and reconstruct vague memory fragments through generative inference, we propose BiOVQL, a brain-inspired one-stage VQL framework. First, inspired by the hippocampus’s selective retention mechanism, we propose the Hippocampus-like Query-guided Spatio-Temporal Compression (HQSTC) module. By leveraging a selective state space model, visual queries are treated as neuromodulators, dynamically gating the video stream to maintain a compact latent state. This guides the model to consistently focus on query-relevant visual cues, enabling query-conditioned feature compression and efficient spatio-temporal memory encoding. Second, inspired by the prefrontal cortex’s re-localization mechanisms, we propose the Prefrontal-like Generative Refinement Localization (PGRL) module. We leverage a diffusion model to reconstruct the localization process as iterative denoising from noise to certainty, which aligns well with the human visual system’s coarse-to-fine perceptual reasoning. This enhances the model’s robustness in handling spatial ambiguities and achieving precise spatio-temporal retrieval. We conducted extensive experiments on the Ego4D-VQ benchmark, demonstrating that BiOVQL achieves state-of-the-art performance with comparable computational efficiency, thus offering an efficient and brain-inspired paradigm for VQL. Yifei Cao, Guolong Wang 0001, Mingliang Hou, Jizhe Yu, Xianjie Zhang, Xiya Bu, Zhizhen Li, Yu Liu 0035 |
ICMR | 7 |
| 2026 | TrackNetV6: A Unified Framework for Lightweight and Robust Fast-Moving Tiny Ball TrackingabstractAlthough vision-based tiny ball tracking has advanced in specific sports, existing methods remain heavily coupled to domain-specific distributions, severely constraining cross-domain generalization. Concurrently, lightweight designs sacrifice representational capacity, while high-performance models incur prohibitive computational costs. To address these challenges, we propose TrackNetV6, a unified fast-moving tiny ball tracking framework that reconciles efficiency with accuracy. Central to our framework is a novel and compact decoding paradigm rooted in the Linear Multistep Method (LMM), designed to supersede conventional single-step feature fusion. This paradigm orchestrates two core components: a Cross-Scale Semantic Consensus Predictor (CSCP) that distills multi-scale features into semantic-correlation location priors, and a Prior-guided Context Corrector (PCC) that injects these priors into current-scale mappings for stable refinement. By iteratively alternating between these components, the model progressively strengthens feature representation for precise tracking. Furthermore, we introduce a Direction-aware Dynamic Fusion (DDF) module as the bottleneck layer, which explicitly models the direction-sensitive feature relationships of the fast-moving ball by synergizing the dynamic interaction between wavelet-based high-frequency cues and deep semantics. Extensive experiments across badminton, table tennis, and tennis benchmarks demonstrate that TrackNetV6 not only sets a new state-of-the-art (SOTA) but also delivers exceptional real-time inference at 183 FPS. Code will be available at https://github.com/Gi-gigi/TrackNetV6. Jizhe Yu, Xiya Bu, Yu Liu 0035, Kaiping Xu, Yifei Cao, Zhizhen Li |
ICMR | 6 |
| 2026 | Text-Conditional Visual-Language Alignment for Video CaptioningabstractVideo captioning remains a challenging task due to the diverse video content and the complex relationships between visual and textual elements. Recent efforts predominantly focus on multimodal architecture designs trained with paired video-caption data. Nonetheless, the learning paradigm suffers from the “one-to-many” corresponding problem, since one source video is mapped to multiple caption annotations. The difficulty of video captioning is further exacerbated by the poor-written captions, which mislead the captioner with irrelevant information. Essentially, the problem stems from the inadequate alignment between video and caption. In this work, we propose a Text-Conditional Alignment Transformer, which fully exploits the rich information provided by diverse labeled captions, and avoids the impacts of label ambiguity and noise. To alleviate the challenge of the “one-to-many” correspondence, we introduce Text-conditioned Video Encoding, which diversifies the video representation by emphasizing the spatial-temporal visual areas relevant to the given descriptions while filtering out redundant visual information. The refined video representation is well-aligned to match the corresponding text description, and naturally converts the “one-to-many” mapping to “one-to-one” mapping. To deal with the noisy annotations, we propose Quality-aware Caption Decoding. We first dynamically measure the qualities of different captions corresponding to the same video in a reference-free manner. Then the estimated qualities are further utilized as auxiliary signals, guiding the model to perform quality-aligned learning from noisy captions. We conduct extensive experiments on MSR-VTT, MSVD, VATEX and ActivityNet-Entities datasets, and demonstrate their consistent performance improvements compared to state-of-the-arts. Wenhui Jiang 0001, Wenbin Guan, Zhizhen Li, Yuming Fang 0001, Yuxin Peng 0001, Yang Liu 0293 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Beamforming Feedback-Driven Wireless Positioning: A Transferable Vision Transformer ApproachabstractWiFi-based indoor positioning plays a crucial role in a variety of location-based services due to its widespread avail ability and cost-effectiveness. However, most existing indoor positioning systems predominantly utilize channel state information (CSI) to learn channel characteristics and apply fingerprinting for position estimation. Unfortunately, CSI can only be extracted from a limited set of commercial WiFi devices, hindering its widespread application in practice. In this work, we introduce BFMLoc, a novel indoor positioning framework that exploits the beamforming feedback matrix (BFM), which is readily available on commercial WiFi devices. Although BFM provides broader sensing coverage, it sacrifices detailed channel information due to the data compression applied to reduce feedback overhead. To address this limitation, we explore the feasibility of using BFM derivatives for indoor positioning and propose a U-net model to reconstruct the angle-delay profiles (ADP) from the compressed BFM data, thereby enhancing positioning accuracy. A Vision Transformer (ViT) model is then developed to extract spatial features from the predicted ADP maps to perform localization. Additionally, we design a model adaptation module based on transfer learning, integrated into the overall framework. This allows the positioning model to be easily deployed and adapted to various indoor environments with minimal retraining overhead. Extensive evaluations and validation on a digital twin testbed demonstrate that our framework achieves high positioning ac curacy and enhanced robustness compared to state-of-the-art methods. Zhizhen Li, Xuanhao Luo, Mingzhe Chen, Gaolei Li, Yuchen Liu 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | TouchWave: Exploring mmWave-Based Non-Contact Fingertip-Force Sensing in Activities of Daily LivingabstractFingertip forces are important biomarkers for the detection and management of various conditions, including stroke and Parkinson's disease. This paper presents TouchWave, a non-contact sensing system designed to monitor fingertip forces during activities of daily living (ADL). TouchWave leverages under-cabinet millimeter-wave (mmWave) sensors to capture both macroscopic hand movements and subtle biomechanical cues associated with fingertip force production. A novel signal processing scheme is developed to suppress noise while preserving force-related information in the mmWave signals. Additionally, a hybrid deep neural network model is proposed to estimate highfidelity fingertip forces. A comprehensive evaluation involving 21 participants demonstrates the effectiveness of TouchWave in both controlled settings and ADL scenarios. Yuliang Fu, Rakshita Ranganath, Zhizhen Li, Yuchen Liu 0001, Ning Sui, Huining Li, Chenhan Xu |
BSN | 4 |
| 2025 | BFMLoc: Transformer-Based Indoor Positioning Leveraging Beamforming Feedback MatricesabstractWiFi-based indoor positioning plays a crucial role in a variety of location-based services due to its widespread availability and cost-effectiveness. However, most existing indoor positioning systems predominantly utilize channel state information (CSI) to learn channel characteristics and apply fingerprinting for position estimation. Unfortunately, CSI can only be extracted from a limited set of commercial WiFi devices, hindering its widespread application in practice. In this work, we introduce BFMLoc, a novel indoor positioning framework that exploits the beamforming feedback matrix (BFM), which is readily available on commercial WiFi devices. Although BFM provides broader sensing coverage, it sacrifices detailed channel information due to the compression applied to reduce feedback overhead. To address this limitation, we explore the feasibility of using BFM derivatives for indoor positioning and propose a U-net model to reconstruct the angle-delay profiles (ADP) from the compressed BFM data, thereby enhancing positioning accuracy. A Vision Transformer (ViT) model is then developed to extract spatial features from the predicted ADP maps to perform localization. Extensive evaluation results demonstrate that our framework achieves high positioning accuracy and improved robustness compared to state-of-the-art methods. Zhizhen Li, Xuanhao Luo, Mingzhe Chen, Chenhan Xu, Yuchen Liu 0001 |
ICC | 1 |
| 2025 | ALPHA: LLM-Enabled Active Learning for Human-Free Network Anomaly DetectionabstractNetwork log data analysis plays a critical role in detecting security threats and operational anomalies. Traditional log analysis methods for anomaly detection and root cause analysis rely heavily on expert knowledge or fully supervised learning models, both of which require extensive labeled data and significant human effort. To address these challenges, we propose ALPHA, the first Active Learning Pipeline for Human-free log Analysis. ALPHA integrates semantic embedding, clustering-based representative sampling, and large language model (LLM)assisted few-shot annotation to automate the anomaly detection process. The LLM annotated labels are propagated across clusters, enabling large-scale training of an anomaly detector with minimal supervision. To enhance the annotation accuracy, we propose a two-step few-shot refinement strategy that adaptively selects informative prompts based on the LLM's observed error patterns. Extensive experiments11The source code of our proposed ALPHA framework is available at https://github.com/Xuanhao-Luo/ALPHA. on real-world log datasets demonstrate that ALPHA achieves detection accuracy comparable to fully supervised methods while mitigating human efforts in the loop. ALPHA also supports interpretable analysis through LLM-driven root cause explanations in the post-detection stage. These capabilities make ALPHA a scalable and cost-efficient solution for truly automated log-based anomaly detection. Xuanhao Luo, Shivesh Madan Nath Jha, Akruti Sinha, Zhizhen Li, Yuchen Liu 0001 |
IPCCC | 4 |
| 2025 | Contextual Combinatorial Beam Management via Online Probing for Multiple Access mmWave Wireless NetworksabstractDue to the exponential increase in wireless devices and a diversification of network services, unprecedented challenges, such as managing heterogeneous data traffic and massive access demands, have arisen in next-generation wireless networks. To address these challenges, there is a pressing need for the evolution of multiple access schemes with advanced transceivers. Millimeter-wave (mmWave) communication emerges as a promising solution by offering substantial bandwidth and accommodating massive connectivities. Nevertheless, the inherent signaling directionality and susceptibility to blockages pose significant challenges for deploying multiple transceivers with narrow antenna beams. Consequently, beam management becomes imperative for practical network implementations to identify and track the optimal transceiver beam pairs, ensuring maximum received power and maintaining high-quality access service. In this context, we propose a Contextual Combinatorial Beam Management (CCBM) framework tailored for mmWave wireless networks. By leveraging advanced online probing techniques and integrating predicted contextual information, such as dynamic link qualities in spatial-temporal domain, CCBM aims to jointly optimize transceiver pairing and beam selection while balancing the network load. This approach not only facilitates multiple access effectively but also enhances bandwidth utilization and reduces computational overheads for real-time applications. Theoretical analysis establishes the asymptotically optimality of the proposed approach, complemented by extensive evaluation results showcasing the superiority of our framework over other state-of-the-art schemes in multiple dimensions. Zhizhen Li, Xuanhao Luo, Mingzhe Chen, Chenhan Xu, Shiwen Mao, Yuchen Liu 0001 |
IEEE J. Sel. Areas Commun. | 1 |
| 2024 | Context-Aware Beam Management via Online Probing in Combinatorial Multi-Armed BanditsabstractMillimeter-wave (mmWave) communication, a cor-nerstone in the evolution of next-generation wireless networks, offers substantial bandwidth and plays a crucial role in advancing wireless connectivity capabilities. Nevertheless, the inherent directionality and susceptibility to blockages pose significant challenges for a cost-effective beam management in densely deployed networks. This paper presents a Contextual Combina-torial Beam Management (CCBM) framework, leveraging both location-aware link qualities and beam correlation to tackle the joint access point (AP) and beam selection problem in mmWave networks, with a specific focus on mitigating coordination overhead and balancing the load across APs. Built upon a formulated multi-armed bandit problem, CCBM significantly reduces the uncertainty during online probing process by employing early stopping and attention-based selection mechanisms. Theoretical analysis establishes the asymptotically optimality of the proposed approach, complemented by extensive evaluation results showcasing the superiority of our framework over other state-of-the-art schemes in multiple dimensions. Zhizhen Li, Xuanhao Luo, Mingzhe Chen, Chenhan Xu, Yuchen Liu 0001 |
ICC | 1 |
| 2024 | GemNet: Analysis and Prediction of Building Materials for Optimizing Indoor Wireless NetworksabstractThis paper investigates the correlation between building material properties and indoor network coverage, encompassing both indoor Wi-Fi and outdoor 5G technologies to provide customized network services tailored to users' needs in diverse areas. We first analyze the impact of building material characteristics, with a special focus on wall materials, on the distribution of wireless signal propagation. Then, a ray-tracing-based method is introduced to synthetically generate high-quality training data that covers fine-grained network scenarios with a wide range of wall materials, extending beyond traditional materials. This dataset serves as the foundation for our proposed Global Embedding Isomorphism Network (GemNet), a machine learning framework that facilitates the prediction of optimal material parameters for customized in-building coverage. This innovation enables architects and builders to design novel, network-friendly materials, ensuring ubiquitous and on-demand network services. Extensive evaluations consistently demonstrate a re-markable prediction accuracy of 90.52% on material parameters, underscoring the framework's ability to optimize indoor wireless network planning through the lens of material engineering. Zhijin Yang, Zhizhen Li, Yi Wang 0068, Jianqing Liu, Mingzhe Chen, Yuchen Liu 0001 |
ICC | 2 |
| 2024 | Map-Driven mmWave Link Quality Prediction With Spatial-Temporal Mobility AwarenessabstractThe susceptibility of millimeter-wave (mmWave) links to blockages poses challenges for maintaining consistent high-rate performance. By predicting link quality in advance at specific locations or times of interest, proactive resource allocation techniques, such as link-quality-aware scheduling, can be employed to optimize the utilization of network resources. In this paper, we introduce a map-driven link quality prediction framework that divides the problem into long-term and short-term link quality predictions to cater to the needs of mobile computing. The first stage aims to predict a long-term radio map considering static network characteristics. We propose to separate LoS and NLoS scenarios, and build an analytical model and a regression-based approach to construct a complete link quality map in the spatial domain. Next, short-term link quality prediction is explored to anticipate future variations in link quality through a spatial-temporal attention-based prediction framework. The essence of this approach lies in capturing the spatial correlation and temporal dependency of mmWave wireless characteristics, followed by an attention mechanism to complement the dynamic link quality prediction task. On top of that, we also design a regional training mechanism with a weighted loss function to address the classical data imbalance problem of map-driven prediction. Extensive experimental and simulation results show that our integrated framework effectively captures comprehensive spatial-temporal knowledge and achieves significantly higher accuracy than other baseline prediction methods, making it a promising solution for a wide range proactive configuration tasks in mobile mmWave networks. Zhizhen Li, Mingzhe Chen, Gaolei Li, Xi Lin 0003, Yuchen Liu 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2023 | Spatial-Temporal Attention-Based mmWave Link Quality Prediction Under Dynamic BlockagesabstractMillimeter-wave (mmWave) communication is a promising technology that has become a key component of next-generation wireless networks due to its large available band-width. However, the susceptibility of mmWave link to dynamic blockages makes it challenging to maintain consistently high rate performance. Hence, it is imperative to have the knowledge of link quality in advance at the location of interest to proactively optimize the use of network resources. In this work, we propose a Spatial-Temporal Attention-based Prediction (STAP) framework to predict the link quality at arbitrary locations in the presence of dynamic blockages. Specifically, our STAP model is built to capture the spatial correlation and temporal dependency of mmWave wireless characteristics in an integrated module, followed by an attention mechanism to complement the link quality prediction task. On top of that, we also design a regional training approach with a weighted loss function to address the data imbalance problem of map-based prediction. Extensive evaluation results show that our framework effectively captures comprehensive spatial-temporal knowledge and achieves significantly higher accuracy than other baseline prediction methods. Zhizhen Li, Mingzhe Chen, Gaolei Li, Yuchen Liu 0001 |
GLOBECOM | 1 |