Ziming Cheng

dblp:203/9718 · DBLP profile ↗
← Back
15ranked-venue papers
8as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Computer networks · 3 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2025 SpiritSight Agent: Advanced GUI Agent with One Look
abstract
Graphical User Interface (GUI) agents demonstrate promising potential in assisting human-computer interaction, automating human user’s navigation on digital devices. An ideal GUI agent is expected to achieve high accuracy, low latency, and compatibility for different GUI platforms. Recent vision-based approaches have shown promise by leveraging advanced Vision Language Models (VLMs). While they generally meet the requirements of compatibility and low latency, these vision-based GUI agents tend to have low accuracy due to their limitations in element grounding. To address this issue, we propose SpiritSight, a vision-based, end-to-end GUI agent that excels in GUI navigation tasks across various GUI platforms. First, we create a multi-level, large-scale, high-quality GUI dataset called GUI-Lasagne using scalable methods, empowering SpiritSight with robust GUI understanding and grounding capabilities. Second, we introduce the Universal Block Parsing (UBP) method to resolve the ambiguity problem inherited from the dynamic resolution strategy, further enhancing SpiritSight’s ability to ground GUI objects. Through these efforts, SpiritSight agent outperforms other advanced methods on diverse GUI benchmarks, demonstrating its superior capability and compatibility in GUI navigation tasks. The models and code will be made available upon publication.
Ziming Cheng, Junting Pan, Zhaohui Hou, Mingjie Zhan
CVPR2
2025 GameMLD: A Game-Sourced Motion-Language Dataset for Stylized Motion Generation
abstract
Text-guided character animation generation has emerged as a significant research area with broad applications in gaming, film, interactive media, and beyond. However, existing motion-language datasets face limitations in motion quality, stylistic diversity, and annotation depth, particularly for professional applications. In contrast to existing datasets based on motion capture or video reconstruction techniques, our dataset leverages professionally crafted game animations and employs a structured annotation framework that incorporates standardized game design terminology. The dataset contains 8,700 high-fidelity motion sequences paired with 26,100 multi-level textual descriptions, generated through our proposed annotation pipeline that combines domain expertise with large language models. Through comprehensive experiments and user studies, we demonstrate GameMLD’s advantages in motion quality, style expressiveness, and annotation quality. Additionally, we showcase its practical value by developing a text-driven character animation generation system that effectively supports game production pipelines. Our experiments with state-of-the-art motion synthesis models demonstrate significant improvements in both animation quality and style control. The GameMLD dataset and source code can be reached via this link.
Yiyu Fu, Ziming Cheng, Yihao Liao, Jiangfeiyang Wang, Ruomei Wang 0001, Guanghui Yue 0001, Chenlei Lv, Baoquan Zhao
ICME2
2025 Few-Shot Medical Image Segmentation With High-Confidence Prior Mask
abstract
Labeling large amounts of medical data is travailing, leading to the blooming of few-shot medical image segmentation, which aims to segment the foreground of a query image given a labeled support set. Almost all current models adopt the cosine distance to measure the similarity between prototypes and query features. However, the limitation of the cosine distance is exacerbated by intra-class differences and inter-class imbalances in medical image scenarios, where angle-only evaluation can induce misclassification to under- and over-segmentation. Motivated by this, we propose a High-Confidence Prior Mask-guided Network (HCPMNet), comprising a High-Confidence Mask Generator (HCPMG), a Target Region Mining (TRM) module, and a Prototype-Oriented Expansion Match (POEM) module. Our HCPMNet offers key advantages: 1) HCPMG is the first to combinatively evaluate angle and magnitude similarity, generating high-confidence priori masks that accurately and completely localize target regions. 2) TRM mines and aggregates target class information under the guidance of priori masks. 3) POEM, based on both similarity metrics, correctly matches prototypes with query features. Extensive experiments on three general medical datasets show that our HCPMNet achieves a new SoTA with great superiority.
Ziming Cheng, Jianqin Zhao, Jingjing Deng 0001, Haofeng Zhang 0001
IEEE J. Biomed. Health Informatics1
2025 Dual Interspersion and Flexible Deployment for Few-Shot Medical Image Segmentation
abstract
Acquiring a large volume of annotated medical data is impractical due to time, financial, and legal constraints. Consequently, few-shot medical image segmentation is increasingly emerging as a prominent research direction. Nowadays, Medical scenarios pose two major challenges: 1) intra-class variation caused by diversity among support and query sets; 2) inter-class extreme imbalance resulting from background heterogeneity. However, existing prototypical networks struggle to tackle these obstacles effectively. To this end, we propose a Dual Interspersion and Flexible Deployment (DIFD) model. Drawing inspiration from military interspersion tactics, we design the dual Interspersion module to generate representative basis prototypes from support features. These basis prototypes are then deeply interacted with query features. Furthermore, we introduce a fusion factor to fuse and refine the basis prototypes. Ultimately, we seamlessly integrate and flexibly deploy the basis prototypes to facilitate correct matching between the query features and basis prototypes, thus conducive to improving the segmentation accuracy of the model. Extensive experiments on three publicly available medical image datasets demonstrate that our model significantly outshines other SoTAs (2.78% higher dice score on average across all datasets), achieving a new level of performance. The code is available at: https://github.com/zmcheng9/DIFD.
Ziming Cheng, Yang Long 0001, Tao Zhou 0002, Haofeng Zhang 0001, Ling Shao 0001
IEEE Trans. Medical Imaging1
2024 MTPNet: Learning Multiple Twin-support Prototypes for Few-shot Medical Image Segmentation
abstract
Given the high annotation costs and ethical considerations associated with medical images, leveraging a limited number of annotated samples for Few-Shot Medical Image Segmentation (FSMIS) has become increasingly prevalent. However, existing models tend to focus on visible foreground support information, often overlooking extreme foreground-background imbalances. In addition, query images sometimes have slight different appearance compared to support images of the same category due to the differences in size as well as slicing angle, thus employing only support images to generate prototypes inevitably leads to matching bias. To address these challenges, we present an innovative approach through learning a Multiple Twin-support Prototypes Network (MTPNet). Our approach includes the design of the Scale Consistent Sampling (SCS) module, which adaptively adjusts the foreground and background points within the support set, thereby balancing the influence of various structural elements in the image. Additionally, the Twin-support Prototypes Extraction (TPE) module facilitates the critical interaction between query and support features to extract twin-support prototypes. This module incorporates a Backtrace Interaction Filter (BIF) to eliminate erroneous interaction prototypes. Extensive experimental validation on three widely used medical image datasets demonstrates that our method surpasses current State-of-the-arts, showcasing its potential to address key limitations in FSMIS. The code is available at https://github.com/FeifanSong/MTPNet.
Feifan Song 0004, Ziming Cheng, Lunbo Li, Haofeng Zhang 0001
BIBM2
2024 Frequency-aware Adaptive Filtering Network for Few-Shot Medical Image Segmentation
abstract
Manual annotation of massive data in the biomedical field necessitates huge costs, expertise, and privacy considerations. However, Few-Shot Medical Image Segmentation (FSMIS) offers the possibility of learning a segmentation model with excellent performance from limited medical data. In this paper, we propose a Frequency-aware Adaptive Filtering Network (FAF-Net) for FSMIS, which is the first to calibrate the correlations between support prototype and query features from a frequency perspective, addressing the problems of FSMIS. The FAF-Net mines the commonalities between the support prototype and query features to reduce the intra-class differences through an Intra-class Commonality Miner (ICM), and introduces an Adaptive Fourier Filter (AFF) to adaptively filter the spectrum of the attention map, thus balancing the foreground and background classes. Additionally, an Adaptive Threshold Learner (ATL) is incorporated to learn an adaptive threshold for each spatial location of the predicted mask from the support features, overcoming the limitations of single thresholding. Extensive experiments on four generic medical datasets showcase that our model significantly outperforms SoTA methods (exceeding SoTA by 2.13% on average).
Ziming Cheng, Tong Xin 0002, Haofeng Zhang 0001
BIBM1
2024 The Root Element of Human Poses is Radian: MCPRL is All You Need
abstract
3D Human Pose Estimation aims to determine the spatial coordinates of key anatomical landmarks on human body. Common benchmarks for this task include Human3.6M and MPI-INF-3DHP, while substantial redundancy exists. Additionally, both datasets are confined indoors due to equipment limitations, compromising diversity and generalization. To address the above issues, we first quantify dataset redundancy by introducing Generalized Radian Pruning (GRP), which employs a novel radians-based method to categorize and optimize human poses. Secondly, based on merging H36M and MPII datasets, we construct a new Manifold Cadre Poses with Radian List-systematically (MCPRL) dataset where missing outdoor scenarios, particularly dynamic collision actions are supplemented. Five strong baselines are compared and the experimental results demonstrate the effectiveness of the GRP method and the superiority of the MCPRL dataset. Our released dataset reduces data redundancy by 35%, surpasses H36M and MPII by 114% in diversity, with an average 11.19mm reduction in the MPJPE indicator. MCPRL dataset can be found at https://github.com/Rxn666/MCPRL.
Ziming Cheng, Xiangning Ruan, Qixiang Yin, Zhicheng Zhao 0001
ICME1
2024 Learning De-biased prototypes for Few-shot Medical Image Segmentation
Yazhou Zhu 0001, Ziming Cheng, Haofeng Zhang 0001
Pattern Recognit. Lett.2
2024 Few-Shot Medical Image Segmentation via Generating Multiple Representative Descriptors
abstract
Automatic medical image segmentation has witnessed significant development with the success of large models on massive datasets. However, acquiring and annotating vast medical image datasets often proves to be impractical due to the time consumption, specialized expertise requirements, and compliance with patient privacy standards, etc. As a result, Few-shot Medical Image Segmentation (FSMIS) has become an increasingly compelling research direction. Conventional FSMIS methods usually learn prototypes from support images and apply nearest-neighbor searching to segment the query images. However, only a single prototype cannot well represent the distribution of each class, thus leading to restricted performance. To address this problem, we propose to Generate Multiple Representative Descriptors (GMRD), which can comprehensively represent the commonality within the corresponding class distribution. In addition, we design a Multiple Affinity Maps based Prediction (MAMP) module to fuse the multiple affinity maps generated by the aforementioned descriptors. Furthermore, to address intra-class variation and enhance the representativeness of descriptors, we introduce two novel losses. Notably, our model is structured as a dual-path design to achieve a balance between foreground and background differences in medical images. Extensive experiments on four publicly available medical image datasets demonstrate that our method outperforms the state-of-the-art methods, and the detailed analysis also verifies the effectiveness of our designed module.
Ziming Cheng, Tong Xin 0002, Tao Zhou 0002, Haofeng Zhang 0001, Ling Shao 0001
IEEE Trans. Medical Imaging1
2023 Cross-View Language Modeling: Towards Unified Cross-Lingual Cross-Modal Pre-training
abstract
In this paper, we introduce Cross-View Language Modeling, a simple and effective pretraining framework that unifies cross-lingual and cross-modal pre-training with shared architectures and objectives.Our approach is motivated by a key observation that cross-lingual and cross-modal pre-training share the same goal of aligning two different views of the same object into a common semantic space.To this end, the cross-view language modeling framework considers both multi-modal data (i.e., image-caption pairs) and multi-lingual data (i.e., parallel sentence pairs) as two different views of the same object, and trains the model to align the two views by maximizing the mutual information between them with conditional masked language modeling and contrastive learning.We pre-train CCLM, a Crosslingual Cross-modal Language Model, with the cross-view language modeling framework.Empirical results on IGLUE, a multi-lingual multi-modal benchmark, and two multi-lingual image-text retrieval datasets show that while conceptually simpler, CCLM significantly outperforms the prior state-of-the-art with an average absolute improvement of over 10%.Moreover, CCLM is the first multi-lingual multimodal pre-trained model that surpasses the translate-test performance of representative English vision-language models by zero-shot cross-lingual transfer.1
Yan Zeng 0003, Wangchunshu Zhou, Ao Luo, Ziming Cheng, Xinsong Zhang
ACL (1)4
2022 BiBL: AMR Parsing and Generation with Bidirectional Bayesian Learning
abstract
Abstract Meaning Representation (AMR) offers a unified semantic representation for natural language sentences. Thus transformation between AMR and text yields two transition tasks in opposite directions, i.e., Text-to-AMR parsing and AMR-to-Text generation. Existing AMR studies only focus on one-side improvements despite the duality of the two tasks, and their improvements are greatly attributed to the inclusion of large extra training data or complex structure modifications which harm the inference speed. Instead, we propose data-efficient Bidirectional Bayesian learning (BiBL) to facilitate bidirectional information transition by adopting a single-stage multitasking strategy so that the resulting model may enjoy much lighter training at the same time. Evaluation on benchmark datasets shows that our proposed BiBL outperforms strong previous seq2seq refinements without the help of extra data which is indispensable in existing counterpart models. We release the codes of BiBL at: https://github.com/KHAKhazeus/BiBL.
Ziming Cheng, Zuchao Li, Hai Zhao 0001
COLING1
2020 Req2Lib: A Semantic Neural Model for Software Library Recommendation
abstract
Third-party libraries are crucial to the development of software projects. To get suitable libraries, developers need to search through millions of libraries by filtering, evaluating, and comparing. The vast number of libraries places a barrier for programmers to locate appropriate ones. To help developers, researchers have proposed automated approaches to recommend libraries based on library usage pattern. However, these prior studies can not sufficiently match user requirements and suffer from cold-start problem. In this work, we would like to make recommendations based on requirement descriptions to avoid these problems. To this end, we propose a novel neural approach called Req2Lib which recommends libraries given descriptions of the project requirement. We use a Sequence-to-Sequence model to learn the library linked-usage information and semantic information of requirement descriptions in natural language. Besides, we apply a domain-specific pre-trained word2vec model for word embedding, which is trained over textual corpus from Stack Overflow posts. In the experiment, we train and evaluate the model with data from 5,625 java projects. Our preliminary evaluation demonstrates that Req2Lib can recommend libraries accurately.
Zhensu Sun, Ziming Cheng, Pengyu Che
SANER3
2020 Channel Path Identification in mmWave Systems With Large-Scale Antenna Arrays
abstract
We consider the uplink channel estimation problem in a millimeter wave (mmWave) system with large-scale antenna arrays. Unlike many existing works which estimate the channel assuming that the number of channel paths is known a priori, we address the problem of channel estimation with an unknown number of channel paths. The spatial channel is transformed into the beamspace channel by the discrete Fourier transform (DFT). Based on the sparsity property of the beamspace channel, we propose three algorithms to estimate the number of paths, direction of arrivals (DoAs) and path gains. The first one is the Spectrum Weighted Identification of Signal Sources (SWISS) for the case when the channel statistics are unknown, which introduces a weight vector to amplify the desired signal and suppress the noise. The second one is the Neyman-Pearson criterion based-Detector (NPD) based on the Rician channel model, which adopts the Neyman-Pearson criterion to decide whether there exists a path on each DFT point. In practice, the DoAs are continuously distributed, leading to the power leakage problem. We solve this leakage problem by proposing the combined algorithm with leakage (CAL). Simulation results show that the proposed algorithms perform better than the conventional spatial smoothing.
Ziming Cheng, Meixia Tao, Pooi Yuen Kam
IEEE Trans. Commun.1
2018 SWISS: Spectrum weighted identification of signal sources for mmWave systems
abstract
This paper considers the channel estimation problem in millimeter-wave (mmWave) systems where a single-antenna user communicates with a massive multiple-input multiple-output (MIMO) base station (BS) in the uplink. Unlike many existing works which estimate the channel gain under the assumption that the number of channel paths is given a priori, we address first the problem of path-number identification. By taking the weighted discrete Fourier transform (WDFT) of the received noisy signal, we formulate an optimization problem to determine the optimum combination of DFT components in this weighted spectrum that leads to a time-domain reconstructed signal (the channel vector) that is at the minimum Euclidean distance from the received signal. Our algorithm, called SWISS (Spectrum Weighted Identification of Signal Sources), is an accurate and computationally efficient means for identifying the paths in the channel vector, providing the information needed for BS beamforming. Once the paths are identified, their individual directions-of-arrival (DoAs) and complex fading gains can be obtained easily. Simulation results for the case of no power leakage in the DFT are presented to demonstrate the effectiveness of SWISS.
Ziming Cheng, Jingyue Huang, Meixia Tao, Pooi Yuen Kam
WCNC1
2017 Low-complexity hybrid analog/digital beamforming for multicast transmission in mmwave systems
abstract
This paper studies multi-group multicast beamforming with a hybrid large-scale antenna array in millimeter wave (mmWave) communication systems. A low-complexity hybrid structure is adopted, where each RF chain is only connected to part of the antenna elements. We formulate a hybrid analog and digital beamforming design problem for multi-group multicast transmission with the objective of minimizing the total transmit power at the base station, subject to an individual signal-to-interference-plus-noise ratio constraint for each multicast group. The problem is very challenging and its global optimal solution is difficult to obtain. We first adopt alternating minimization method to design the analog and digital beamformer alternatively. Then we solve each of the analog and digital subproblems through solving a sequence of convex problems via concave-convex procedure (CCP). Each convex CCP subproblem is reformulated as a novel alternating direction method of multipliers (ADMM) form. Our ADMM reformulation enables that each updating step can be decomposed into multiple subproblems with much smaller size, which can be solved optimally in parallel with closed-form expressions. Simulation results show that our algorithm can achieve favorable performance with very low complexity compared with the state-of-art methods.
Jingyue Huang, Ziming Cheng, Erkai Chen, Meixia Tao
ICC2