Xinyu Zhang 0017

dblp:58/4582-17 · DBLP profile ↗
← Back
18ranked-venue papers
7as first author
15since 2021 · last 2026
0000-0002-6891-7718ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Security and privacy · 4 · 4 first-author · 3 since 2021Computer networks · 1
YearPublicationVenuePosition
2026 A dynamic balanced training regime for mitigating class imbalance in thyroid nodule classification
abstract
Handling class imbalance remains a major challenge in developing intelligent diagnostic systems for ultrasound-based thyroid nodule classification. Synthetic sample generation often distorts data distributions and weakens texture fidelity, a critical concern in ultrasound, where fine anatomical details carry diagnostic value. Loss reweighting may destabilize optimization under severe imbalance, while recent approaches favor architectural complexity over training-level refinement, limiting clinical deployment. These issues hinder accurate minority-class detection in screening, where precise classification is essential to avoid missed diagnoses or unnecessary interventions. In this paper, a lightweight adaptive framework extending the Dynamic Balanced Training Regime (DBTR) is proposed for class-imbalanced thyroid ultrasound classification. First, balanced subsets are constructed directly from original ultrasound images, preserving diagnostic texture without synthetic augmentation. Then, a recall-guided learning-rate adjustment is introduced to modulate the classifier head with a fixed backbone, strengthening minority-class sensitivity under speckle noise. Experiments were conducted on the imbalanced DDTI benchmark using 10-fold cross-validation and further tested on two recent independent external cohorts, TN5000 and ThyUS2Path, acquired under different clinical settings. Internal evaluation demonstrates that the proposed framework achieves the highest AUC (0.714) among the evaluated CNN architectures, while maintaining high sensitivity (0.984) and improving specificity from 0.233 to 0.250 over the Xception baseline, reaching balanced accuracy (0.617). Under external evaluation, a 5.5% relative increase in specificity was obtained on ThyUS2Path despite domain shift, while sensitivity remained preserved across both cohorts. Computational benchmarking confirmed efficient deployment, supporting the integration of intelligent diagnostic support systems into point-of-care ultrasound screening workflows.
Fatimah Ali S. Aljalis, Vincent Cheng-Siong Lee, Xinyu Zhang 0017
Expert Syst. Appl.3
2025 Are Images Indistinguishable to Humans Also Indistinguishable to Classifiers?
abstract
The ultimate goal of generative models is to perfectly capture the data distribution. For image generation, common metrics of visual quality (e.g., FID) and the perceived truthfulness of generated images seem to suggest that we are nearing this goal. However, through distribution classification tasks, we reveal that, from the perspective of neural network-based classifiers, even advanced diffusion models are still far from this goal. Specifically, classifiers are able to consistently and effortlessly distinguish real images from generated ones across various settings. Moreover, we uncover an intriguing discrepancy: classifiers can easily differentiate between diffusion models with comparable performance (e.g., U-ViTH vs. DiT-XL), but struggle to distinguish between models within the same family but of different scales (e.g., EDM2-XS vs. EDM2-XXL). Our methodology carries several important implications. First, it naturally serves as a diagnostic tool for diffusion models by analyzing specific features of generated data. Second, it sheds light on the model autophagy disorder and offers insights into the use of generated data: augmenting real data with generated data is more effective than replacing it. Third, classifier guidance can significantly enhance the realism of generated images.
Zebin You, Xinyu Zhang 0017, Hanzhong Guo, Jingdong Wang 0001, Chongxuan Li
CVPR2
2025 Let Your Video Listen to Your Music! - Beat-Aligned, Content-Preserving Video Editing with Arbitrary Music
abstract
Aligning the rhythm of visual motion in a video with a given music track is a practical need in multimedia production, yet remains an underexplored task in autonomous video editing. Effective alignment between motion and musical beats enhances viewer engagement and visual appeal, particularly in music videos, promotional content, and cinematic editing. Existing methods typically depend on labor-intensive manual cutting, speed adjustments, or heuristic-based editing techniques to achieve synchronization. While some generative models handle joint video and music generation, they often entangle the two modalities, limiting flexibility in aligning video to music beats while preserving the full visual content. In this paper, we propose a novel and efficient framework-termed MVAA (Music-Video Auto-Alignment)-that automatically edits video to align with the rhythm of a given music track while preserving the original visual content. To enhance flexibility, we modularize the task into a two-step process in our MVAA: aligning motion keyframes with audio beats, followed by rhythm-aware video inpainting. Specifically, we first insert keyframes at timestamps aligned with musical beats, then use a frame-conditioned diffusion model to generate coherent intermediate frames, preserving the original video's semantic content. Since comprehensive test-time training can be time-consuming, we adopt a two-stage strategy: pretraining the inpainting module on a small video set to learn general motion priors, followed by rapid inference-time fine-tuning for video-specific adaptation. This hybrid approach enables adaptation within ~10 minutes with one epoch on a single NVIDIA 4090 GPU using CogVideoX-5b-I2V [77] as the backbone. Extensive experiments show that our approach can achieve high-quality beat alignment and visual smoothness. User studies further validate the natural rhythmic quality of the results, confirming their effectiveness for practical music-video editing. The code is available at: zhangxinyu-xyz.github.io/MVAA
Xinyu Zhang 0017, Dong Gong, Zicheng Duan, Anton van den Hengel, Lingqiao Liu
ACM Multimedia1
2025 FedMGP: Personalized Federated Learning with Multi-Group Text-Visual Prompts
abstract
In this paper, we introduce FedMGP, a new paradigm for personalized federated prompt learning in vision-language models (VLMs). Existing federated prompt learning (FPL) methods often rely on a single, text-only prompt representation, which leads to client-specific overfitting and unstable aggregation under heterogeneous data distributions. Toward this end, FedMGP equips each client with multiple groups of paired textual and visual prompts, enabling the model to capture diverse, fine-grained semantic and instance-level cues. A diversity loss is introduced to drive each prompt group to specialize in distinct and complementary semantic aspects, ensuring that the groups collectively cover a broader range of local characteristics.During communication, FedMGP employs a dynamic prompt aggregation strategy based on similarity-guided probabilistic sampling: each client computes the cosine similarity between its prompt groups and the global prompts from the previous round, then samples s groups via a softmax-weighted distribution. This soft selection mechanism preferentially aggregates semantically aligned knowledge while still enabling exploration of underrepresented patterns—effectively balancing the preservation of common knowledge with client-specific features. Notably, FedMGP maintains parameter efficiency by redistributing a fixed prompt capacity across multiple groups, achieving state-of-the-art performance with the lowest communication parameters (5.1k) among all federated prompt learning methods. Theoretical analysis shows that our dynamic aggregation strategy promotes robust global representation learning by reinforcing shared semantics while suppressing client-specific noise. Extensive experiments demonstrate that FedMGP consistently outperforms prior approaches in both personalization and domain generalization across diverse federated vision-language benchmarks.The code will be released on https://github.com/weihao-bo/FedMGP.git.
Weihao Bo, Yanpeng Sun, Xinyu Zhang 0017, Zechao Li
NeurIPS4
2025 Plum: SNARK-Friendly Post-Quantum Signature Based on Power Residue PRFs
Xinyu Zhang 0017, Qishuang Fu, Ron Steinfeld, Joseph K. Liu, Tsz Hon Yuen, Man Ho Au
ProvSec1
2025 Symmetric Hallucination With Knowledge Transfer for Few-Shot Learning
abstract
Data hallucination or augmentation is a straightforward solution for few-shot learning (FSL), where FSL is proposed to classify a novel object under limited training samples. Common hallucination strategies use visual or textual knowledge to simulate the distribution of a given novel category and generate more samples for training. However, the diversity and capacity of generated samples through these techniques can be insufficient when the knowledge domain of the novel category is narrow. Therefore, the performance improvement of the classifier is limited. To address this issue, we propose a Symmetric data hallucination strategy with Knowledge Transfer (SHKT) that interacts with multi-modal knowledge in both visual and textual spaces. Specifically, we first calculate the relations based on semantic knowledge and select the most related categories of a given novel category for hallucination. Second, we design two parameter-free data hallucination strategies to enrich the training samples by mixing the given and selected samples in both visual and textual spaces. The generated visual and textual samples improve the visual representation and enrich the textual supervision, respectively. Finally, we connect the visual and textual knowledge through transfer calculation, which not only exchanges content from different modalities but also constrains the distribution of the generated samples during the training. We apply our method to four benchmark datasets and achieve state-of-the-art performance in all experiments. Specifically, compared to the baseline on the Mini-ImageNet dataset, it achieves 12.84% and 3.46% accuracy improvements for 1 and 5 support training samples, respectively.
Shuo Wang 0008, Xinyu Zhang 0017, Meng Wang 0001, Xiangnan He 0001
IEEE Trans. Multim.2
2024 DualRing-PRF: Post-quantum (Linkable) Ring Signatures from Legendre and Power Residue PRFs
Xinyu Zhang 0017, Ron Steinfeld, Joseph K. Liu, Muhammed F. Esgin, Dongxi Liu, Sushmita Ruj
ACISP (2)1
2024 Loquat: A SNARK-Friendly Post-quantum Signature Based on the Legendre PRF with Applications in Ring and Aggregate Signatures
Xinyu Zhang 0017, Ron Steinfeld, Muhammed F. Esgin, Joseph K. Liu, Dongxi Liu, Sushmita Ruj
CRYPTO (1)1
2024 VRP-SAM: SAM with Visual Reference Prompt
abstract
In this paper, we propose a novel Visual Reference Prompt (VRP) encoder that empowers the Segment Any-thing Model (SAM) to utilize annotated reference images as prompts for segmentation, creating the VRP-SAM model. In essence, VRP-SAM can utilize annotated reference images to comprehend specific objects and perform segmen-tation of specific objects in target image. It is note that the VRP encoder can support a variety of annotation for-mats for reference images, including point, box, scribble, and mask. VRP-SAM achieves a breakthrough within the SAM framework by extending its versatility and applicabil-ity while preserving SAM's inherent strengths, thus enhancing user-friendliness. To enhance the generalization abil-ity of VRP-SAM, the VRP encoder adopts a meta-learning strategy. To validate the effectiveness of VRP-SAM, we con-ducted extensive empirical studies on the Pascal and COCO datasets. Remarkably, VRP-SAM achieved state-of-the-art performance in visual reference segmentation with mini-mal learnable parameters. Furthermore, VRP-SAM demon-strates strong generalization capabilities, allowing it to per-form segmentation of unseen objects and enabling cross-domain segmentation. The source code and models will be available at https://github.com/syp2ysy/VRP-SAM
Yanpeng Sun, Shan Zhang 0002, Xinyu Zhang 0017, Qiang Chen 0007, Errui Ding, Jingdong Wang 0001, Zechao Li
CVPR4
2024 Multi-Stage Fusion for Event-based Multimodal Tracker
abstract
Event cameras are bio-inspired sensors with high dynamic range and time resolution, which are favorable properties for visual object tracking. There are already some methods that fuse the event modality and RGB modality with cross-domain feature integrator to achieve improved tracking performance. Researchers have developed some architectures for event modality processing or fusion, successfully boosting the tracking performance. In this work, we design a RGB-E tracker with multi-stage fusion. In the early stage, frames are enhanced with aid of events to mitigate blur or under/over-exposure degradation. During the middle stage, we utilize a fusion module for feature-level integration. At the late stage, we carry out decision-level fusion by predicting tracking boxes based on frame features, event features, and fused features, and the one with highest score is taken as the final estimation. Our design thoroughly integrate information from various levels, allowing each modality to contribute to the tracking process as much as possible. Extensive experiments demonstrate that the proposed method performs favorably against state-of-the-art RGB-E trackers in both accuracy and efficiency.
Xinyu Zhang 0017, Hefei Huang, Xu Jia 0012, Wenyue Chen, Dong Wang 0004, Shengming Li, Huchuan Lu
ICME1
2024 Event-Guided Rolling Shutter Correction with Time-Aware Cross-Attentions
abstract
Many consumer cameras with rolling shutter (RS) CMOS would suffer undesired distortion and artifacts, particularly when objects experiences fast motion. The neuromorphic event camera, with high temporal resolution events, could bring much benefit to the RS correction process. In this work, we explore the characteristics of RS images and event data for the design of the rolling shutter correction (RSC) model. Specifically, the relationship between RS images and event data is modeled by incorporating time encoding to the computation of cross-attention in transformer encoder to achieve time-aware multi-modal information fusion. Features from RS images enhanced by event data are adopted as keys and values in transformer decoder, providing source for appearance, while features from event data enhanced by RS images are adopted as queries, providing spatial transition information. By embedding the time information of the desired global shutter (GS) image into the query, the transformer with deformable attention is capable of producing the target GS image.To enhance the model's generalization ability, we propose to further self-supervise the model by cycling between time coordinate systems corresponding to RS images and GS images. Extensive evaluations over both synthetic and real datasets demonstrate that the proposed method performs favorably against state-of-the-art approaches.
Hefei Huang, Xu Jia 0012, Xinyu Zhang 0017, Shengming Li, Huchuan Lu
ACM Multimedia3
2024 Evaluation of Text-to-Video Generation Models: A Dynamics Perspective
abstract
Comprehensive and constructive evaluation protocols play an important role when developing sophisticated text-to-video (T2V) generation models. Existing evaluation protocols primarily focus on temporal consistency and content continuity, yet largely ignore dynamics of video content. Such dynamics is an essential dimension measuring the visual vividness and the honesty of video content to text prompts. In this study, we propose an effective evaluation protocol, termed DEVIL, which centers on the dynamics dimension to evaluate T2V generation models, as well as improving existing evaluation metrics. In practice, we define a set of dynamics scores corresponding to multiple temporal granularities, and a new benchmark of text prompts under multiple dynamics grades. Upon the text prompt benchmark, we assess the generation capacity of T2V models, characterized by metrics of dynamics ranges and T2V alignment. Moreover, we analyze the relevance of existing metrics to dynamics metrics, improving them from the perspective of dynamics. Experiments show that DEVIL evaluation metrics enjoy up to about 90\% consistency with human ratings, demonstrating the potential to advance T2V generation models.
Mingxiang Liao, Hannan Lu, Qixiang Ye, Wangmeng Zuo, Fang Wan 0001, Tianyu Wang 0028, Yuzhong Zhao, Jingdong Wang 0001, Xinyu Zhang 0017
NeurIPS9
2024 Neural image re-exposure
Xinyu Zhang 0017, Hefei Huang, Xu Jia 0012, Dong Wang 0004, Lihe Zhang, Bolun Zheng, Wei Zhou 0021, Huchuan Lu
Comput. Vis. Image Underst.1
2023 Dual Memory Aggregation Network for Event-Based Object Detection with Learnable Representation
abstract
Event-based cameras are bio-inspired sensors that capture brightness change of every pixel in an asynchronous manner. Compared with frame-based sensors, event cameras have microsecond-level latency and high dynamic range, hence showing great potential for object detection under high-speed motion and poor illumination conditions. Due to sparsity and asynchronism nature with event streams, most of existing approaches resort to hand-crafted methods to convert event data into 2D grid representation. However, they are sub-optimal in aggregating information from event stream for object detection. In this work, we propose to learn an event representation optimized for event-based object detection. Specifically, event streams are divided into grids in the x-y-t coordinates for both positive and negative polarity, producing a set of pillars as 3D tensor representation. To fully exploit information with event streams to detect objects, a dual-memory aggregation network (DMANet) is proposed to leverage both long and short memory along event streams to aggregate effective information for object detection. Long memory is encoded in the hidden state of adaptive convLSTMs while short memory is modeled by computing spatial-temporal correlation between event pillars at neighboring time intervals. Extensive experiments on the recently released event-based automotive detection dataset demonstrate the effectiveness of the proposed method.
Xu Jia 0012, Xinyu Zhang 0017, Yaoyuan Wang, Dong Wang 0004, Huchuan Lu
AAAI4
2021 Alpha-Refine: Boosting Tracking Performance by Precise Bounding Box Estimation
abstract
Visual object tracking aims to precisely estimate the bounding box for the given target, which is a challenging problem due to factors such as deformation and occlusion. Many recent trackers adopt the multiple-stage strategy to improve bounding box estimation. These methods first coarsely locate the target and then refine the initial prediction in the following stages. However, existing approaches still suffer from limited precision, and the coupling of different stages severely restricts the method’s transferability. This work proposes a novel, flexible, and accurate refinement module called Alpha-Refine (AR), which can significantly improve the base trackers’ box estimation quality. By exploring a series of design options, we conclude that the key to successful refinement is extracting and maintaining detailed spatial information as much as possible. Following this principle, Alpha-Refine adopts a pixel-wise correlation, a corner prediction head, and an auxiliary mask head as the core components. Comprehensive experiments on TrackingNet, LaSOT, GOT-10K, and VOT2020 benchmarks with multiple base trackers show that our approach significantly improves the base tracker’s performance with little extra latency. The proposed Alpha-Refine method leads to a series of strengthened trackers, among which the ARSiamRPN (AR strengthened SiamRPNpp) and the ARDiMP50 (AR strengthened DiMP50) achieve good efficiency-precision trade-off, while the ARDiMPsuper (AR strengthened DiMPsuper) achieves very competitive performance at a realtime speed. Code and pretrained models are available at https://github.com/MasterBin-IIAU/AlphaRefine.
Bin Yan 0004, Xinyu Zhang 0017, Dong Wang 0004, Huchuan Lu, Xiaoyun Yang
CVPR2
2019 Revocable and Linkable Ring Signature
Xinyu Zhang 0017, Joseph K. Liu, Ron Steinfeld, Veronika Kuchta, Jiangshan Yu
Inscrypt1
2018 A hybrid Markov-based model for human mobility prediction
Yuanyuan Qiao 0002, Zhongwei Si, Yanting Zhang 0001, Fehmi Ben Abdesslem, Xinyu Zhang 0017, Jie Yang 0023
Neurocomputing5
2015 Global and individual mobility pattern discovery based on hotspots
abstract
Data collected from the mobile Internet have the potential knowledge to provide important human mobility patterns. Understanding human mobility patterns is important to many location-based services, and could be used to predict users' behavior. In this paper, we concentrate on the issue of discovering human mobility patterns on both global and individual levels based on hotspots. We study the human mobility trajectories during 22 days for 3474 individuals collected at the core of a metropolitan Long Term Evolution (LTE) network in China. We employ a parameter-free method to detect hotspots, and demonstrate the effectiveness of our mobility pattern discovery algorithm by using the hotspots identified on both global and individual levels. We analyze the occurrence time distribution of these patterns and find that the global mobility patterns have higher occurrence probability in the morning, which indicates that people in a city tend to share the common commuting routes. For individual mobility patterns, there exists a strong spatiotemporal correlation property, implying that the individual mobility patterns have their own typical occurrence time depending on the pattern's context.
Jie Yang 0023, Xinyu Zhang 0017, Yuanyuan Qiao 0002, Zubair Md Fadlullah, Nei Kato
ICC2