Zixuan Song

dblp:249/2735 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Intelligent task management via dynamic multi-region division in LEO satellite networks
Zixuan Song, Zhishu Shen, Xiaoyu Zheng 0005, Qiushi Zheng, Zheng Lei, Jiong Jin
Comput. Networks1
2026 MoDe-Track: Robust Multi-Object Tracking With Motion Decoupling in UAV Videos
abstract
Multi-Object Tracking (MOT) in Unmanned Aerial Vehicle (UAV) scenarios is characterized by frequent and abrupt camera motion, which presents two unique challenges: nonlinear motion and appearance degradation. Traditional motion models, designed for smooth and consistent motion, struggle to capture the complex background motion patterns caused by UAV movement; while appearance-based methods are easily disrupted by occlusion and blur, leading to unreliable associations. Even though dense optical flow is widely utilized to model complex motion patterns, the entanglement of background and object motion often introduces interference, limiting its effectiveness. To this end, we propose MoDe-Track, a unified framework that explicitly decouples scene motion into background and object components, and serves as an elegant integration of three robust components. Specifically, the Scene Motion Decomposition (SMD) module decouples the motion into the background and object components based on robust principal component analysis, serving as the foundation for motion compensation and feature propagation. Afterwards, the Background Motion Compensation (BMC) uses the decomposed background flow to estimate and compensate for camera motion, mitigating the effects of nonlinear motion. Finally, the Foreground-guided Feature Propagation (FFP) module uses the decoupled object flow to guide feature propagation across frames, achieving temporal consistency and enhancing robustness against occlusion and motion blur. Extensive experimental results on two benchmarks, VisDrone2019 and UAVDT, demonstrate that MoDe-Track consistently outperforms current multi-object tracking methods. We achieve 56.0% MOTA on VisDrone2019 and 56.2% MOTA on UAVDT, reaching the state-of-the-art among existing methods.
Zixuan Song, Sanping Zhou, Wei Tang 0016, Le Wang 0003
IEEE Trans. Multim.1
2026 ZlibBoost: An Efficient and Flexible Open-Source Framework for Standard Cell Characterization
abstract
As VLSI designs grow increasingly complex and transition to smaller process nodes, accurate and efficient library characterization has become essential for modern design workflows. Existing open-source tools are often constrained by limited functionality, efficiency, and accuracy, making them insufficient for today’s design challenges. This article reviews the shortcomings of current open-source tools and introduces ZlibBoost, a novel open-source framework designed to provide both flexibility and high performance. Its modular, front-end and back-end separated architecture, along with user-friendly interfaces, enables seamless customization, integration of machine learning models, and expanded simulator compatibility. A variety of key features are introduced to significantly enhance both accuracy and efficiency of library characterization. Experimental results demonstrate ZlibBoost’s capability to meet the demands of both academic research and practical applications, establishing it as a robust solution for advancing semiconductor design.
Zhengrui Chen, Chengjun Guo, Shizhang Wang, Guozhu Feng, Zixuan Song, Xunzhao Yin, Weiquan Song, Li Zhang 0021, Zheyu Yan, Cheng Zhuo
ACM Trans. Design Autom. Electr. Syst.5
2026 A Novel Gradual Inference Approach for Handling Perturbations in Sparse CrowdSensing
abstract
Sparse CrowdSensing has emerged as a promising paradigm for data collection, which utilizes mobile devices to gather partial sensing data and infer the remaining ones. Most existing studies focus on data inference. However, they often overlook the impact of unforeseen circumstances. These perturbations can lead to many issues in the data inference process such as over-correction (adjustments are excessively applied), perturbation diffusion (errors affecting related data), and data bias (overall trends in inferred data deviating from the truth). To address the aforementioned challenges, we propose a novel approach called Gradual Matrix Completion (GMC) for handling perturbations in Sparse CrowdSensing. This method distinguishes itself by inferring adjacent unsensed data and using it as a basis for ongoing optimization, gradually completing the entire data matrix. Through the gradual inference process, GMC mitigates perturbation diffusion by continuously refining feature selection, thus preserving and reconstructing essential features. Moreover, GMC concentrates on both intra-area and inter-area spatio-temporal relationships, progressively enhancing its understanding of local and global dependencies. Evaluation of the GMC approach on four diverse real-world datasets highlights its robust capability to complete perturbed data and cope with noise in Sparse CrowdSensing.
En Wang, Zixuan Song, Jing Deng 0001, Bo Yang 0002, Yongjian Yang 0001
IEEE Trans. Netw.2
2025 Invited Paper: Boosting Standard Cell Library Characterization with Machine Learning
abstract
As VLSI designs grow more complex and transition to smaller process nodes, accurate and efficient library characterization has become increasingly crucial within DTCO and STCO flows. Current open-source tools, however, are constrained to basic library characterization functions and fail to adequately meet modern design demands. In this paper, we review the existing open-source standard cell characterization tools, summarize their limitations, and introduce ZlibBoost---a new open-source framework designed to offer both flexibility and efficiency. We leverage ZlibBoost for LUT index optimization, dynamic power supply noise modeling, and machine learning-based prediction to enhance efficiency and accuracy in library characterization. Experimental results show that such a tool is helpful for both academia and industry to effectively navigate DTCO and STCO challenges.
Zhengrui Chen, Chengjun Guo, Zixuan Song, Guozhu Feng, Shizhang Wang, Li Zhang 0021, Xunzhao Yin, Zheyu Yan, Cheng Zhuo
ASP-DAC3
2025 Towards General Continuous Memory for Vision-Language Models
abstract
Language models (LMs) and their extension, vision-language models (VLMs), have achieved remarkable performance across various tasks. However, they still struggle with complex reasoning tasks that require multimodal or multilingual real world knowledge. To support such capabilities, an external memory system that can efficiently provide relevant multimodal information is essential. Existing approaches generally concatenate image and text tokens into a long sequence as memory, which, however, may drastically increase context length and even degrade performance. In contrast, we propose using continuous memory-a compact set of dense embeddings-to more effectively and efficiently represent multimodal and multilingual knowledge. Our key insight is that a VLM can serve as its own continuous memory encoder. We empirically show that this design improves performance on complex multimodal reasoning tasks. Building on this, we introduce a data-efficient and parameter-efficient method to fine-tune the VLM into a memory encoder, requiring only 1.2\% of the model’s parameters and a small corpus of 15.6K self-synthesized samples. Our approach CoMEM utilizes VLM's original capabilities to encode arbitrary multimodal and multilingual knowledge into just 8 continuous embeddings. Since the inference-time VLM remains frozen, our memory module is plug-and-play and can be flexibly integrated as needed. Extensive experiments across eight multimodal reasoning benchmarks demonstrate the effectiveness of our approach. Code and data is publicly released here https://github.com/WenyiWU0111/CoMEM.
Wenyi Wu, Zixuan Song, Kun Zhou 0002, Yifei Shao, Zhiting Hu, Biwei Huang
NeurIPS2
2025 SmartQCache: Fast and Precise Pulse Control With Near-Quantum Cache Design on FPGA
abstract
Quantum pulse serves as the machine language of superconducting quantum devices, which needs to be synthesized and calibrated for precise control of quantum operations. However, existing pulse control systems suffer from the dilemma between long synthesis latency and inaccuracy of quantum control systems. compute-in-CPU synthesis frameworks, like IBM Qiskit Pulse, involve massive redundant computation during pulse calculation, suffering from a high computational cost when handling large-scale circuits. On the other hand, field-programmable gate array (FPGA)-based synthesis frameworks, like QuMA, faces inaccurate pulse control problem. In this article, we propose both compute-in-CPU and all-in-FPGA solutions to collaboratively solve the latency and inaccuracy problem. First, we propose QPulseLib, a novel compute-in-CPU library with reusable pulses that can directly provide the pulse of a circuit pattern. To establish this library, we transform the circuit and apply convolutional operators to extract reusable patterns and precalculate their resultant pulses. Then, we develop a matching algorithm to identify such patterns shared by the target circuit. Experiments show that QPulseLib achieves$158.46\times $and$16.03\times $speedup for pulse calculation, compared to Qiskit Pulse and AccQOC. Moreover, we extend the design as a fast and precise all-in-FPGA pulse control approach using near-quantum cache design, SmartQCache. To be specific, we employ a two-level cache to hold reusable pulses of frequently-used circuit patterns. Such a design enables pulse prefetching in near-quantum peripherals, dramatically reducing the end-to-end synthesis latency. To achieve precise pulse control, SmartQCache incorporates duration optimization and pulse sequence calibration to mitigate the execution errors from imperfect hardware, crosstalk, and time shift. Experimental results demonstrate that SmartQCache achieves$294.37\times $and$145.43\times $speedup in pulse synthesis compared to Qiskit Pulse and AccQOC. It also reduces the pulse inaccuracy by$1.27\times $compared to QuMA.
Liqiang Lu, Wuwei Tian, Xinghui Jia, Zixuan Song, Siwei Tan, Jianwei Yin
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2025 A New Data Completion Perspective on Sparse CrowdSensing: Spatiotemporal Evolutionary Inference Approach
abstract
Mobile CrowdSensing (MCS) has emerged as a popular paradigm to engage mobile users in collaborative sensing tasks. However, its performance is hindered by its limited spatiotemporal range and the cost of data collection. An effective strategy is to integrate Sparse MCS with data completion, allowing for unsensed data inference. However, when confronted with situations where sensed data is excessively sparse, data inference results may be unsatisfactory due to several challenges including: 1) uneven data distribution, 2) complex spatiotemporal correlation, and 3) the presence of inference noise. To address these challenges, we propose a model named Spatiotemporal Evolutionary Inference (STEI) that achieves accurate inference of unsensed data in Sparse MCS. Specifically, we complete the unsensed data by uncovering strong local correlations in the data and gradually evolving those correlations to the global situation. In each evolution step, we thoroughly consider the impact of spatiotemporal consistency and difference. To minimize the interference of noise during the evolution process, we design an adaptive coefficient to enhance the dependence on sensed data. Finally, to validate the effectiveness of STEI, we conduct extensive qualitative and quantitative experiments using three popular datasets. The experimental results demonstrate that our approach excels in accurately inferring data, particularly in situations where the distribution of data is notably uneven.
En Wang, Zixuan Song, Mengni Wu, Bo Yang 0002, Yongjian Yang 0001, Jie Wu 0001
IEEE Trans. Mob. Comput.2
2024 Few-Shot Data Completion for New Tasks in Sparse Crowdsensing
abstract
Mobile Crowdsensing is a type of technology that utilizes mobile devices and volunteers to gather data about specific topics at large scales in real-time. However, in practice, limited participation leads to missing data, i.e., the collected data may be sparse, which makes it difficult to perform accurate analysis. A possible technique called sparse crowdsensing incorporates the sparse case with data completion, where unsensed data could be estimated through inference. However, sparse crowdsensing typically suffers from poor performance during the data completion stage due to various challenges: the sparsity of the sensed data, reliance on numerous timeslots, and uncertain spatiotemporal connections. To resolve such few-shot issues, the proposed solution uses the Correlated Data Fusion for Matrix Completion (CDFMC) approach, which leverages a small amount of objective data to retrain an auxiliary dataset-based pre-trained model that can estimate unsensed data efficiently. CDFMC is trained using a combination of the traditional Deep Matrix Factorization and the Kalman Filtering, which not only enables the efficient representation and comparison of data samples but also fuses the objective data and auxiliary data effectively. Evaluation results show that the proposed CDFMC outperforms baseline techniques, achieving high accuracy in completing unsensed data with minimal training data.
En Wang, Mijia Zhang, Bo Yang 0002, Yang Xu 0013, Zixuan Song, Yongjian Yang 0001
INFOCOM5
2023 QPulseLib: Accelerating the Pulse Generation of Quantum Circuit with Reusable Patterns
abstract
Quantum circuit serves as a popular programming model that describes the computation using a set of quantum gates, which requires generating a sequence of pulses that collect the operation of each gate for superconducting quantum devices. However, existing quantum synthesis frameworks, like IBM OpenPulse [1], involve massive redundant computation during pulse generation, suffering from a high computational cost when handling large-scale circuits. In this paper, we propose QPulseLib, a novel library with reusable pulses that can directly provide the pulse of a circuit block. To establish this library, we transform the circuit and apply convolutional operators to extract reusable patterns and pre-calculate their resultant pulses. Then, we develop a matching algorithm to identify such patterns shared by the target circuit. Experiments show that QPulseLib achieves 158.46 × and 16.03 × speedup for pulse generation, compared to OpenPulse and AccQOC [2].
Wuwei Tian, Xinghui Jia, Siwei Tan, Zixuan Song, Liqiang Lu, Jianwei Yin
ICCAD4
2023 An data augmentation method for source code summarization
Zixuan Song, Xiuwei Shang, Guanxi Li, Hui Li 0014, Shikai Guo
Neurocomputing1
2023 Code samples summarization for knowledge exchange in developer community
abstract
Abstract A question title's function is to generate readable titles and describe a problem encountered by the code. Previous studies often used an end‐to‐end sequence‐to‐sequence system to generate question title's from source code. However, long‐term dependencies are often difficult to capture, and this may result in an incomplete source code representation. To address this issue, we propose a Transformer for Generating Code Title (hereinafter referred to as TGCT) model. Specifically, the TGCT model uses the position coding mechanism to model paired relationships between source terms by applying relative position representations. Multiple self‐attention mechanism components are also used to capture long‐term dependencies of the code. Comprehensive experiments on datasets from five coding languages, namely Python, Java, JavaScript, C#, and SQL, are conducted, and the results show that TGCT outperforms state‐of‐the‐art models based on the measurements of BLEU and ROUGE in general. In addition, a cross‐sectional comparison experiment was conducted to verify the effects of different model parameters, different data set sizes, position coding mechanism, and self‐attention mechanism on model results.
Shikai Guo, Zhongyan Liu, Zixuan Song, Hui Li 0014, Rong Chen 0003
Softw. Pract. Exp.3
2023 ERMF: Edge refinement multi-feature for change detection in bitemporal remote sensing images
Zixuan Song, Rui Zhu 0015, Zeyu Wang 0009, Xiaoli Zhang 0001
Signal Process. Image Commun.1