Ziqin Zhou

dblp:128/8157 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 8 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 HiTVideo: Hierarchical Tokenizers for Enhancing Text-to-Video Generation with Autoregressive Large Language Models
abstract
Text-to-video generation poses significant challenges due to the inherent complexity of video data, which spans both temporal and spatial dimensions. It introduces additional redundancy, abrupt variations, and a domain gap between language and vision tokens while generation. Addressing these challenges requires an effective video tokenizer that can efficiently encode video data while preserving essential semantic and spatiotemporal information, serving as a critical bridge between text and vision. Inspired by the observation in VQ-VAE-2, we propose HiTVideo, a novel approach for text-to-video generation with hierarchical tokenizers. It utilizes a 3D causal VAE with a multi-layer discrete token framework, encoding video content into hierarchically structured codebooks. Higher layers capture semantic information with higher compression, while lower layers focus on fine-grained spatiotemporal details, striking a balance between compression efficiency and reconstruction quality. Our approach efficiently encodes longer video sequences (e.g., 8 seconds, 64 frames), reducing bits per pixel (bpp) by approximately 70% compared to previous tokenizers, while maintaining competitive reconstruction quality. We explore the trade-offs between compression and reconstruction, while emphasizing the advantages of high-compressed semantic tokens in text-to-video tasks. HiTVideo aims to address the potential limitations of existing video tokenizers in text-to-video generation tasks, striving for higher compression ratios, improved token quality, and simplify LLMs modeling under language guidance, offering a scalable and promising framework for advancing text to video generation.
Ziqin Zhou, Yifan Yang 0004, Yuqing Yang 0001, Tianyu He, Houwen Peng, Qi Dai 0001, Lili Qiu, Chong Luo 0001, Lingqiao Liu
AAAI1
2026 MTAVG-Bench: A Diagnostic Benchmark for Multi-Talker Dialogue-Centric Audio-Video Generation
abstract
Yanghao Zhou, Haitian Li, Rexar Lin, Heyan Huang, Jinxing Zhou, Changsen Yuan, Tian Lan, Ziqin Zhou, Yudong Li, Jiajun Xu, Jingyun Liao, YiMing Cheng, Xuefeng Chen, Xian-Ling Mao, Yousheng Feng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yanghao Zhou, Haitian Li, Rexar Lin, Heyan Huang, Jinxing Zhou, Changsen Yuan, Tian Lan 0003, Ziqin Zhou, Jingyun Liao, YiMing Cheng, Xianling Mao, Yousheng Feng
ACL (1)8
2026 Characterization of the heterogeneity in SARS-CoV-2 fitness dynamics via graph representation learning
abstract
Understanding the heterogeneity of population-level viral fitness dynamics, which reflect the interplay between intrinsic viral properties and population immunity, is critical for pandemic preparedness. However, how these dynamics vary across diverse immune backgrounds and mutational landscapes remain poorly characterized. We present Geno-GNN, a graph representation learning approach for retrospectively characterizing the viral fitness dynamics of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2). Geno-GNN accurately predicts angiotensin-converting enzyme 2 (ACE2) binding affinity and immune escape potential across multiple external datasets. Using Geno-GNN, we identified temporal patterns in SARS-CoV-2 fitness and detected varying rates of fitness change associated with distinct immune backgrounds. Virtual mutation scanning revealed two fitness trajectories: broad immune evasion at the cost of ACE2 affinity and ACE2 affinity maintenance at or above the Wuhan-Hu-1 level along with moderate immune escape. Notably, real-world SARS-CoV-2 variants predominantly followed the latter trajectory, sustaining ACE2 affinity via fixed mutations. These findings underscore the heterogeneous, immune-contextualized nature of viral fitness dynamics and the complex evolutionary pathways of SARS-CoV-2.
Zengmiao Wang, Ziqin Zhou, Junfu Wang, Lingyue Yang, Zhirui Zhang, Weina Xu, Zeming Liu, Yuxi Ge, Liang Yang 0002, Quanyi Wang, Yunlong Cao, Yuanfang Guo, Huaiyu Tian
PLoS Comput. Biol.2
2025 CF-ViT: Cross-Feature Vision Transformer for Improving Feature Learning on Tiny Datasets
abstract
Efficient feature learning is considered indispensable for maximizing the representation of scarce information in tiny datasets. However, existing methods are often unable to fully exploit local features and contextual dependencies when dealing with tiny datasets. To overcome this shortcoming, a Cross-Feature Vision Transformer (CF-ViT) was proposed, which decouples local feature refinement from global context modeling and leverages the complementary strengths of CNNs and Transformers. Specifically, a Cross-Scale Fusion (CSF) module was introduced to integrate features from multiple scales, ensuring that cross-scale information is globally embedded. In addition, a Feature Enhancement and Reorganization (FER) module was incorporated into CF-ViT, whereby Transformer outputs are reorganized into 2D feature maps for convolution-based detail enhancement to thoroughly exploit local information. Extensive experiments have demonstrated that CF-ViT consistently surpasses baselines across 4 tiny datasets, reaching a 96.87% (KSDD) Top-1 accuracy with only 29.19 million parameters and 2.67 billion FLOPs. Moreover, a Top-1 accuracy of 85.03% is attained on a real-world tiny dataset of wood surface defect detection, exceeding all baselines. These findings underscore the effectiveness and generalization capability of CF-ViT in capturing fine-grained local details and global context, offering a promising and deployable solution for vision tasks in tiny datasets.
Chunlei Meng, Yi Liu 0027, Hongda Zhang, Yuning Chen, Bowen Liu 0017, Ziqin Zhou, Chun Ouyang 0002, Zhongxue Gan 0001, Dunzhao Wu, Zhihua Nie
SMC8
2025 Energy Efficient Data Processing: Integrated Sensing-Communication-Computation Design
abstract
In space-air-ground-sea networks, the conventional data processing designs separately considering sensing, communication and computation processes lead to severe wastes of radio, energy, and computation resources. To overcome this drawback, an integrated sensing-communication-computation design is pro-posed in this paper, which aims at realizing energy efficient data processing by jointly determining the data offloading ratio together with the sensing and offloading rates according to the processor profiles of mobile devices and servers. It is proved that the data offloading ratio is determined by the server's processor profile, while the string-pulling algorithms are designed to obtain the optimal sensing and offloading rates. Simulations are conducted to verify the effectiveness of the proposed design.
Ziqin Zhou, Xiaoyang Li 0002, Guangxu Zhu, Bingpeng Zhou, Chang Liu 0008, Kaibin Huang
VTC2025-Spring1
2025 THzCondenser: A System Design for IRS-Aided Terahertz Wideband Communications
abstract
With the access to tens of gigahertz of bandwidth, terahertz (THz) wideband communication emerges as a promising technology for the upcoming next generation mobile networks. To deal with the severe path loss and blockage of THz signals, massive multiple-input multiple-output and intelligent reflecting surface (IRS) can be jointly employed. Due to the extremely large signal bandwidth, the beams generated by the transmit hybrid beamforming may point to different directions around the target direction at different frequencies, which results in the beam splitting effect (BSE). In this paper, a new system design namelyTHzCondenseris introduced to mitigate the BSE, where the signals generated by each transmit radio frequency (RF) chain are reflected by one of the distributed IRSs in a one-to-one manner via the joint transmit and IRS beamforming design, thus creating adjustable multi-path components to achieve both high spatial multiplexing gain and array gain. Moreover, for practical scenarios when the number of transmit RF chains is more than that of IRSs, each IRS may need to reflect the signals generated by multiple RF chains in a one-to-many manner. For the above two cases, the joint beamforming design problems are efficiently solved to maximize the achievable rate. Simulations are conducted to verify the effectiveness of the proposed algorithms for mitigating the BSE and improving the achievable rate in IRS-aided THz wideband communications.
Yihang Jiang 0001, Yi Gong 0001, Ziqin Zhou, Xiaoyang Li 0002, Rui Zhang 0006
IEEE Trans. Wirel. Commun.3
2024 Unlocking the Potential of Pre-Trained Vision Transformers for Few-Shot Semantic Segmentation through Relationship Descriptors
abstract
The recent advent of pre-trained vision transformers has unveiled a promising property: their inherent capability to group semantically related visual concepts. In this paper, we explore to harnesses this emergent feature to tackle few-shot semantic segmentation, a task focused on classifying pixels in a test image with a few example data. A critical hurdle in this endeavor is preventing overfitting to the limited classes seen during training the few-shot segmentation model. As our main discovery, we find that the concept of “relationship descriptors”, initially conceived for enhancing the CLIP model for zero-shot semantic segmentation, offers a potential solution. We adapt and refine this concept to craft a relationship descriptor construction tailored for few-shot semantic segmentation, extending its application across multiple layers to enhance performance. Building upon this adaptation, we proposed a few-shot semantic segmentation framework that is not only easy to implement and train but also effectively scales with the number of support examples and categories. Through rigorous experimentation across various datasets, including PASCAL-5iand COCO-20i, we demonstrate a clear advantage of our method in diverse few-shot semantic segmentation scenarios, and a range of pre-trained vision transformer models. The findings clearly show that our method significantly outperforms current state-of-the-art techniques, highlighting the effectiveness of harnessing the emerging capabilities of vision transformers for few-shot semantic segmentation. We release the code at https://github.com/ZiqinZhou66/FewSegwithRD.git.
Ziqin Zhou, Yangyang Shu, Lingqiao Liu
CVPR1
2024 Integrating Sensing, Communication, and Power Transfer: Multiuser Beamforming Design
abstract
In the sixth-generation (6G) networks, massive low-power devices are expected to sense environment and deliver tremendous data. To enhance the radio resource efficiency, the integrated sensing and communication (ISAC) technique exploits the sensing and communication functionalities of signals, while the simultaneous wireless information and power transfer (SWIPT) techniques utilizes the same signals as the carriers for both information and power delivery. The further combination of ISAC and SWIPT leads to the advanced technology namely integrated sensing, communication, and power transfer (ISCPT). In this paper, a multi-user multiple-input multiple-output (MIMO) ISCPT system is considered, where a base station equipped with multiple antennas transmits messages to multiple information receivers (IRs), transfers power to multiple energy receivers (ERs), and senses a target simultaneously. The sensing target can be regarded as a point or an extended surface. When the locations of IRs and ERs are separated, the MIMO beamforming designs are optimized to improve the sensing performance while meeting the communication and power transfer requirements. The resultant non-convex optimization problems are solved based on a series of techniques including Schur complement transformation and rank reduction. Moreover, when the IRs and ERs are co-located, the power splitting factors are jointly optimized together with the beamformers to balance the performance of communication and power transfer. To better understand the performance of ISCPT, the target positioning problem is further investigated. Simulations are conducted to verify the effectiveness of our proposed designs, which also reveal a performance tradeoff among sensing, communication, and power transfer.
Ziqin Zhou, Xiaoyang Li 0002, Guangxu Zhu, Jie Xu 0002, Kaibin Huang, Shuguang Cui
IEEE J. Sel. Areas Commun.1
2024 Wireless Communication and Control Co-Design for System Identification
abstract
The unprecedented growth of industrial Internet of Things applications requires the evolution of wireless networked control system (WNCS). WNCSs are becoming the fundamental infrastructure technologies for critical wireless control applications due to the main benefits of the reduced deployment and maintenance cost, as well as the enhanced flexibility and safety. However, independent designs between communication and control without considering their tight interaction in conventional WNCS lead to poor overall system performance and efficiency. Co-designs are expected to achieve the target control performance while improving the wireless resource efficiency. In this paper, by considering how to allocate wireless resource while guaranteeing control performance, a co-design framework is established based on the finite-time wireless system identification (WSI) - a fundamental problem in systems theory and intelligent control. To this end, two design problems are investigated aiming at maximizing the communication throughput or minimizing the power consumption while guaranteeing the WSI performance. In the former design, the joint optimization of power and channel allocations leads to a non-convex integer combinatorial problem, which is iteratively solved by optimizing the power allocation via Lagrangian method and obtaining the optimal channel allocation via Hungarian algorithm. The minimum number of data samples for guaranteeing the WSI accuracy under confidence level is further derived by exploiting the relationship between WSI accuracy and the number of state sampling processes, which leads to the maximum throughput with respect to both the communication and control processes. In the latter design for energy-efficient WSI, by exploiting the relationship between the power consumption and channel allocation given the WSI performance requirement, the optimization problem can be simplified and solved by Hungarian algorithm. Simulations are conducted to verify the performance of the proposed solutions.
Xiaoyang Li 0002, Ziqin Zhou, Kaibin Huang, Yi Gong 0001, Qinyu Zhang 0001
IEEE Trans. Wirel. Commun.3
2023 ZegCLIP: Towards Adapting CLIP for Zero-shot Semantic Segmentation
abstract
Recently, CLIP has been applied to pixel-level zero-shot learning tasks via a two-stage scheme. The general idea is to first generate class-agnostic region proposals and then feed the cropped proposal regions to CLIP to utilize its image-level zero-shot classification capability. While effective, such a scheme requires two image encoders, one for proposal generation and one for CLIP, leading to a complicated pipeline and high computational cost. In this work, we pursue a simpler-and-efficient one-stage solution that directly extends CLIP's zero-shot prediction capability from image to pixel level. Our investigation starts with a straightforward extension as our baseline that generates semantic masks by comparing the similarity between text and patch embeddings extracted from CLIP. However, such a paradigm could heavily overfit the seen classes and fail to generalize to unseen classes. To handle this issue, we propose three simple-but-effective designs and figure out that they can significantly retain the inherent zero-shot capacity of CLIP and improve pixel-level generalization ability. Incorporating those modifications leads to an efficient zero-shot semantic segmentation system called ZegCLIP. Through extensive experiments on three public benchmarks, ZegCLIP demonstrates superior performance, outperforming the state-of-the-art methods by a large margin under both “inductive” and “transductive” zero-shot settings. In addition, compared with the two-stage method, our one-stage ZegCLIP achieves a speedup of about 5 times faster during inference. We release the code at https://github.com/ZiqinZhou66/ZegCLIP.git.
Ziqin Zhou, Yinjie Lei, Bowen Zhang 0009, Lingqiao Liu, Yifan Liu 0001
CVPR1
2023 Multi-User Beamforming Design for Integrating Sensing, Communications, and Power Transfer
abstract
To facilitate the data collection process, simultaneous wireless information and power transfer utilizes the same signal for powering the devices and delivering the information, while the integrated sensing and communication utilizes the same signal for data transmission and radar sensing. In next generation networks, the sensing, communication, and power transfer functionalities are expected to be integrated together to enhance the radio resource efficiency and enable the data collection by massive low-power devices, which leads to the new research direction namely integrating sensing, communication, and power transfer (ISCPT). The ISCPT beamforming design for multiple users is investigated in this paper to improve the sensing performance while guaranteeing the communication and power transfer requirements. The resultant non-convex optimization problem is solved by the approach based on semidefinite relaxation and rank reduction methods. Simulations are further conducted to verify the effectiveness of the proposed design.
Xiaoyang Li 0002, Xuan Yi, Ziqin Zhou, Kaifeng Han, Yi Gong 0001
WCNC3
2023 Integrated Sensing, Communication, and Computation Over-the-Air: MIMO Beamforming Design
abstract
To support the unprecedented growth of the Internet of Things (IoT) applications, tremendous data need to be collected by the IoT devices and delivered to the server for further computation. By utilizing the same signals for both radar sensing and data transmission, theintegrated sensing and communication(ISAC) technique enables simultaneous data collection and delivery in the physical layer. By exploiting the analog-wave addition property in a multi-access channel,over-the-air computation(AirComp) has been proposed as a communication approach that also enables function computation. The promising performances of ISAC and AirComp motivate the current work on developing a framework calledintegrated sensing, communication, and computation over-the-air(ISCCO). Two schemes are designed to supportmultiple-input-multiple-output(MIMO) ISCCO simultaneously, namely theseparated and sharedschemes. The separated scheme splits antenna array for radar sensing and AirComp, while all the antennas transmit a joint waveform for both radar sensing and AirComp in the shared scheme. The performance of radar sensing is evaluated by themean squared error(MSE) of the estimated target response matrix, while the MSE of the estimated function is adopted as the metric to evaluate the performance of the coupled communication and computation in AirComp. The design challenge of MIMO ISCCO lies in the joint optimization of beamformers at both the IoT devices and the server, which results in a non-convex problem. To solve this problem, an algorithmic solution based on the technique of semidefinite relaxation is proposed. The results reveal that the beamformer at each sensor needs to account for supporting dual-functional signals in the shared scheme, while dedicated beamformers for sensing and AirComp are needed to mitigate the mutual interference between the two functionalities in the separated scheme. The application of ISCCO on target location estimation is further demonstrated via simulation.
Xiaoyang Li 0002, Fan Liu 0005, Ziqin Zhou, Guangxu Zhu, Shuai Wang 0004, Kaibin Huang, Yi Gong 0001
IEEE Trans. Wirel. Commun.3
2023 Joint Sensing and Communication-Rate Control for Energy Efficient Mobile Crowd Sensing
abstract
Driven by the rapid growth of Internet of Things applications, tremendous data need to be collected by sensors and uploaded to the servers for further process. As a promising solution, mobile crowd sensing (MCS) enables controllable sensing and transmission processes of multiple types of data in a single device. Despite the appealing advantages, existing works on MCS have mostly simplified two design issues, namely joint control of sensing and transmission processes and corresponding energy consumption. To address the above issues, a single-user MCS system is considered with a typical MCS device sensing and transmitting data to a server in a given time duration. In particular, there exists a busy time interval when the device is incapable of sensing. To minimize the sensing-and-transmission energy consumption of the device, an optimization problem is formulated, where the sensing and transmission rates are jointly optimized over time subjecting to the constraints on the sensing data sizes, transmission data sizes, data casualty, and busy time of sensing. This problem is highly challenging due to the coupling between the rates as well as the existence of the busy time. To deal with this problem, we first show that it can be equivalently decomposed into two subproblems, corresponding to a search for the amount of data size that needs to be sensed before the busy time (referred to as the height), as well as the control of sensing and transmission rates given the height. Next, we show that the latter problem can be efficiently solved by using the classical string-pulling method, while an efficient algorithm is proposed to progressively find the optimal height without the exhaustive search. Moreover, the solution approach is extended to a more complex scenario where there is a finite-size buffer at the server for receiving data. Last, simulations are conducted to evaluate the performance of the proposed designs.
Ziqin Zhou, Xiaoyang Li 0002, Changsheng You, Kaibin Huang, Yi Gong 0001
IEEE Trans. Wirel. Commun.1
2022 Learning and Energy Efficient Edge Intelligence: Data Partition and Rate Control
abstract
The rapid development of artificial intelligence together with the powerful computation capabilities of the advanced edge servers make it possible to deploy learning tasks at the wireless network edge, which is dubbed as edge intelligence (EI). The communication bottleneck between the data resource and the server results in deteriorated learning performance as well as tremendous energy consumption. To tackle this challenge, we explore a new paradigm called learning-and-energy-efficient (LEE) EI, which simultaneously maximizes the learning accuracies and energy efficiencies of multiple tasks via data partition and rate control. Mathematically, this results in a multi-objective optimization problem. Moreover, the continuous varying rates introduce infinite variables, which further complicates the problem. To solve this complex problem, the number of variables is reduced to a finite level by exploiting the optimality of constant-rate transmission in each epoch, based on which a string-pulling (SP) algorithm is proposed to obtain the numerical values. The performance of the proposed joint data partition and rate control design is examined by experiments based on public datasets.
Xiaoyang Li 0002, Shuai Wang 0004, Guangxu Zhu, Ziqin Zhou, Kaibin Huang, Yi Gong 0001
ICC4
2022 Data Partition and Rate Control for Learning and Energy Efficient Edge Intelligence
abstract
The rapid development of artificial intelligence together with the powerful computation capabilities of the advanced edge servers make it possible to deploy learning tasks at the wireless network edge, which is dubbed as edge intelligence (EI). The communication bottleneck between the data resource and the server results in deteriorated learning performance as well as tremendous energy consumption. To tackle this challenge, we explore a new paradigm called learning-and-energy-efficient (LEE) EI, which simultaneously maximizes the learning accuracies and energy efficiencies of multiple tasks via data partition and rate control. Mathematically, this results in a multi-objective optimization problem. Moreover, the continuously varying communication rates introduce infinite variables, which further complicates the problem. To solve this complex problem, we consider the case with infinite server buffer capacity and one-shot data arrival at sensor. First, the number of variables is reduced to a finite level by exploiting the optimality of constant-rate transmission in each epoch. Second, the optimal solution of the multi-objective problem is found by applying the stratified sequencing or merging of objectives. By assuming higher priority of learning efficiency in stratified sequencing, the optimal data partition is derived in closed form by the Lagrange method, while the optimal rate control is proved to have the structure of directional water filling (DWF), based on which a string-pulling (SP) algorithm is proposed to obtain the numerical values. The DWF structure of rate control is also proved to be optimal in merging of objectives, which combines different objectives in a weighted manner. By exploiting the optimal rate changing properties, the SP algorithm is further extended to tackle the more challenging cases with limited server buffer capacity or bursty data arrival at sensor. The performance of the proposed joint data partition and rate control design is examined by extensive experiments based on public datasets.
Xiaoyang Li 0002, Shuai Wang 0004, Guangxu Zhu, Ziqin Zhou, Kaibin Huang, Yi Gong 0001
IEEE Trans. Wirel. Commun.4
2019 Deep point-to-subspace metric learning for sketch-based 3D shape retrieval
Yinjie Lei, Ziqin Zhou, Yulan Guo, Zijun Ma, Lingqiao Liu
Pattern Recognit.2