Zhiyu Wu

dblp:144/9815 · DBLP profile ↗
← Back
25ranked-venue papers
4as first author
24since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 11 since 2021Security and privacy · 4 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Not Just What's There: Enabling CLIP to Comprehend Negated Visual Descriptions Without Fine-Tuning
abstract
Vision-Language Models (VLMs) like CLIP struggle to understand negation, often embedding affirmatives and negatives similarly (e.g., matching "no dog" with dog images). Existing methods refine negation understanding via fine-tuning CLIP’s text encoder, risking overfitting. In this work, we propose CLIPGlasses, a plug-and-play framework that enhances CLIP’s ability to comprehend negated visual descriptions. CLIPGlasses adapts a dual-stage design: a Lens module disentangles negated semantics from text embeddings, and a Frame module predicts context-aware repulsion strength, which is integrated into the modified similarity computation to penalize alignment with negated semantics, thereby reducing false positive matches. Experiments show that CLIP equipped with CLIPGlasses achieves competitive in-domain performance and outperforms state-of-the-art methods in cross-domain generalization. Its superiority is especially evident under low-resource conditions, indicating stronger robustness across domains.
Zhiyu Wu, Zejiang He
AAAI2
2026 JITServe: SLO-aware LLM Serving with Imprecise Request Information
Wei Zhang 0044, Zhiyu Wu, Yi Mu 0005, Rui Ning, Banruo Liu, Nikhil Sarda, Myungjin Lee, Fan Lai 0001
NSDI2
2026 DualFormer: A dual-branch transformer framework for frame-event tracking with complementary fusion attention and sparse spatial channel attention
Ruke Xiong, Guixi Liu, Hanlin Huang, Yisong Xiao, Zhiyu Wu
Eng. Appl. Artif. Intell.7
2026 CAMT: A novel symmetric cross-modal adaptive modulation framework for RGB-T tracking
Yisong Xiao, Guixi Liu, Hanlin Huang, Ruke Xiong, Zhiyu Wu
Neurocomputing7
2026 VT-BM3D: A collaborative filtering framework with joint optimization of structure awareness and noise characteristics
Zhiyu Wu
Signal Process.4
2026 Fine-Grained Detection of Java Cross-Library Vulnerability Propagation by Extracting Semantic Constraints From Security Patches
Fute Sun, Lei Zhang 0096, Zhiyu Wu, Tianyang Han, Min Yang 0002
IEEE Trans. Dependable Secur. Comput.3
2025 GIBRNet: A Multimodal Spatiotemporal Reasoning Network Integrating Emotion, Gaze, and Position for Gaze Interaction Behavior Recognition
Junhao Xiao 0002, Jingxing Zhong, Zhiyu Wu
CogSci4
2025 JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation
abstract
We present JanusFlow, a powerful framework that unifies image understanding and generation in a single model. JanusFlow introduces a minimalist architecture that integrates autoregressive language models with rectified flow, a state-of-the-art method in generative modeling. Our key finding demonstrates that rectified flow can be straightforwardly trained within the large language model framework, eliminating the need for complex architectural modifications. To further improve the performance of our unified model, we adopt two key strategies: (i) decoupling the understanding and generation encoders, and (ii) aligning their representations during unified training. Extensive experiments show that JanusFlow achieves comparable or superior performance to specialized models in their respective domains, while significantly outperforming existing unified approaches. This work represents a step toward more efficient and versatile vision-language models.
Yiyang Ma, Xingchao Liu, Xiaokang Chen, Wen Liu 0011, Chengyue Wu, Zhiyu Wu, Zizheng Pan, Zhenda Xie, Xingkai Yu, Liang Zhao 0026, Jiaying Liu 0001, Chong Ruan
CVPR6
2025 Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
abstract
We introduce Janus, an autoregressive framework that unifies multimodal understanding and generation. Prior research often relies on a single visual encoder for both tasks, such as Chameleon. However, due to the differing levels of information granularity required by multimodal understanding and generation, this approach can lead to suboptimal performance, particularly in multimodal understanding. To address this issue, we decouple visual encoding into separate pathways, while still leveraging a single, unified transformer architecture for processing. The decoupling not only alleviates the conflict between the visual encoder’s roles in understanding and generation, but also enhances the framework’s flexibility. For instance, both the multi-modal understanding and generation components can independently select their most suitable encoding methods. Experiments show that Janus surpasses previous unified model and matches or exceeds the performance of task-specific models. The simplicity, high flexibility, and effectiveness of Janus make it a strong candidate for next-generation unified multimodal models. The code will be made available.
Chengyue Wu, Xiaokang Chen, Zhiyu Wu, Yiyang Ma, Xingchao Liu, Zizheng Pan, Wen Liu 0011, Zhenda Xie, Xingkai Yu, Chong Ruan, Ping Luo 0002
CVPR3
2025 RecNet: Optimization for Dense Object Detection in Retail Scenarios Based on View Rectification
abstract
High-precision dense object detection in retail is crucial for automation, inventory management, and sales optimization. Our experiments revealed that detection models perform significantly better with frontal views than with oblique views, motivating the development of RecNet. RecNet utilizes a Rectify-Detect (R-D) pipeline to transform oblique views into frontal views, mitigating perspective distortion and focusing on key regions. To optimize bounding box prediction, we propose the CeIoU loss function, which focuses on high-quality boxes using a fair elite selection strategy. We also introduce the Neighbor Scattering Algorithm to address accuracy loss caused by rounding errors during the rectification process. Additionally, we present a Transform-Aware Branch that integrates transformation information into the regression branch for direct bounding box prediction. Experiments show that RecNet achieves state-of-the-art performance on SKU110K and demonstrates strong generalization on PUCPR+.
Junhao Xiao 0002, Zhiyu Wu
ICASSP5
2025 Traffic Sign Small Object Detection Algorithm Based on Lightweight Structure Design
abstract
This paper proposes a traffic sign small object detection algorithm based on lightweight structural design to address the existing problems of low detection accuracy for small objects, high complexity, and unsuitability for deployment on mobile in-vehicle devices in current traffic sign detection algorithms. Firstly, the algorithm optimizes YOLOv8n. By reducing the number of backbone network layers and replacing the large object detection layer with the small object detection layer, the number of parameters is greatly reduced, and the detection performance of small traffic signs is significantly improved. Secondly, we design a lightweight detection head (LDH), which greatly reduces the number of parameters and calculation amount of the algorithm while maintaining the original detection accuracy, thus significantly improving the inference speed of the algorithm. Finally, a slice non-parametric attention module (SimAMWithSlicing) is introduced. The module does not need to introduce additional parameters, and only generates attention weights by calculating local self-similarity of feature graphs, thus effectively enhancing the feature representation ability of small objects. Our algorithm was tested on the TT100K2021 and CCTSDB2021 datasets. Versus the original YOLOv8n, the average precision was increased by 2.6% and 0.9%, respectively, while the number of parameters was reduced by 80.3% and 81%. The experimental results demonstrate that our algorithm achieves a balance between accuracy and speed with lower parameters and computational load in traffic sign small object detection.
Shuaibo Chen, Puping An, Shiwu Zeng, Chenyu Lin, Zhiyu Wu, Zunwang Ke
IJCNN5
2025 Exploring Static Taint Analysis in LLMs: A Dynamic Benchmarking Framework for Measurement and Enhancement
abstract
LLMs offer a promising avenue to overcome the limitations of traditional taint analysis techniques, with a growing number of studies leveraging LLMs for taint analysis and its downstream applications. However, these studies lack a systematic understanding of LLMs’ taint analysis capabilities, limiting their transferability and reliability. To bridge this gap and better apply LLMs to static taint analysis, we aim to comprehensively measure and understand LLMs’ taint analysis capabilities.Using existing benchmarks is a straightforward approach, but they are unsuitable due to issues such as training data leakage, not accounting for LLMs’ features, and improper assessment criteria. Manually constructing new benchmarks is not only labor-intensive but also struggles to remain effective as LLMs evolve. To address these, we propose LLMCapLens, a dynamic benchmark generation framework to systematically measure and enhance LLMs’ capabilities. LLMCapLens models influencing factors of LLMs’ taint analysis capabilities, employing a Basic Unit-Based generation method and a lightweight dynamic taint analysis-based verification method to implement the automated generation of targeted benchmarks, ensuring both diversity and correctness. Furthermore, LLMCapLens proposes a measurement-driven, training-free, model-specific enhancement approach.We apply LLMCapLens to 10 mainstream LLMs, revealing how they perform under various influencing factors and identifying unique characteristics, such as the underlying error causes for each model. Notably, our enhancement approach significantly improves LLM performance—GPT-4 Turbo, for instance, achieved improvements across 16 out of 19 factors, with an average True Negative Rate increase of 21.29%. Finally, we validate the real-world impact of our method by applying enhanced LLMs to vulnerability detection, demonstrating a substantial improvement over prior approaches.
Lei Zhang 0006, Keke Lian, Fute Sun, Bofei Chen, Yongheng Liu, Zhiyu Wu, Yuan Zhang 0009, Min Yang 0002
ASE7
2025 FabasedVC: Enhancing Voice Conversion with Text Modality Fusion and Phoneme-Level SSL Features
abstract
In voice conversion (VC), it is crucial to preserve complete semantic information while accurately modeling the target speaker’s timbre and prosody. This paper proposes FabasedVC to achieve VC with enhanced similarity in timbre, prosody, and duration to the target speaker, as well as improved content integrity. It is an end-to-end VITS-based VC system that integrates relevant textual modality information, phoneme-level self-supervised learning (SSL) features, and a duration predictor. Specifically, we employ a text feature encoder to encode attributes such as text, phonemes, tones and BERT features. We then process the frame-level SSL features into phoneme-level features using two methods: average pooling and attention mechanism based on each phoneme’s duration. Moreover, a duration predictor is incorporated to better align the speech rate and prosody of the target speaker. Experimental results demonstrate that our method outperforms competing systems in terms of naturalness, similarity, and content integrity. We strongly recommend that readers listen to our samples.1
Zhetao Hu, Yiquan Zhou, Jiacheng Xu 0007, Zhiyu Wu
MMAsia5
2025 The ML.ENERGY Benchmark: Toward Automated Inference Energy Measurement and Optimization
abstract
As the adoption of Generative AI in real-world services grow explosively, energy has emerged as a critical bottleneck resource. However, energy remains a metric that is often overlooked, under-explored, or poorly understood in the context of building ML systems. We present the ML.ENERGY Benchmark, a benchmark suite and tool for measuring inference energy consumption under realistic service environments, and the corresponding ML.ENERGY Leaderboard, which have served as a valuable resource for those hoping to understand and optimize the energy consumption of their generative AI services. In this paper, we explain four key design principles for benchmarking ML energy we have acquired over time, and then describe how they are implemented in the ML.ENERGY Benchmark. We then highlight results from the early 2025 iteration of the benchmark, including energy measurements of 40 widely used model architectures across 6 different tasks, case studies of how ML design choices impact energy consumption, and how automated optimization recommendations can lead to significant (sometimes more than 40%) energy savings without changing what is being computed by the model. The ML.ENERGY Benchmark is open-source and can be easily extended to various customized models and application scenarios.
Jae-Won Chung, Jeff J. Ma, Oh Jun Kweon, Yuxuan Xia, Zhiyu Wu, Mosharaf Chowdhury
NeurIPS7
2025 Road Disease Detection Algorithm Based on Multiscale Feature Fusion and Receptive Field Enhancement
Zhihao Xue, Zunwang Ke, Yugui Zhang, Menghui Shen, Zhiyu Wu
PRCV (18)7
2025 Lightweight real-time discriminative Siamese deep coupling framework for robust aerial tracking
abstract
Recently, transformer-based Unmanned Aerial Vehicle (UAV) trackers have achieved notable success. However, the computationally intensive transformer model limits these trackers to static templates and shallow backbone networks, hampering their discriminative power and localization precision. Here, we propose a novel discriminative Siamese deep-coupling framework. This framework constructs a lightweight fine-grid anchor-free Siamese tracker with high spatial resolution specifically tailored for UAV scenarios, and complements its discriminative power with a targeted online discriminator. To achieve this, an efficient distractor detector is developed via knowledge transfer, enabling targeted detection of distractors that disturb the Siamese tracker. These distractors are utilized as training samples to construct a targeted online discriminator, which is deeply coupled with the Siamese tracker to enhance its discriminative power and specifically suppress hard distractors that hinder tracking performance. Additionally, a leading principal submatrix cluster sample space model and a scene-aware dynamic update strategy are developed to purify online samples and dynamically schedule the online discriminator update, significantly reducing the computational cost of the online discriminator optimization and boosting the tracker’s real-time performance. Finally, extensive experiments on eight UAV tracking benchmarks demonstrate that our tracker surpasses state-of-the-art transformer-based UAV trackers while achieving 70 FPS on CPU.
Hanlin Huang, Guixi Liu, Ruke Xiong, Zhiyu Wu
Inf. Sci.6
2024 Qiao: DIY your routing protocol in Internet-of-Things
abstract
Routers in the Internet of Things (IoT) operate in a failure-prone, multi-device integrated, and dynamically changing network topology environment. Dealing with diverse network failures, device compatibility, and flexible configurations are the challenges faced by IoT routers [21]. Existing traditional Internet solutions and recently emerging Software-Defined Networking (SDN) solutions are not capable of addressing these issues effectively. They either lack support for flexible configurations or suffer from compatibility problems. This article proposes Qiao, which a novel routing system written in a high-level language (HLL), with built-in concurrency and a robust standard library for industrial internet, that aims to tackle the interoperability issues in small to medium-sized industrial IoT networks. Qiao, which means bridge in Chinese, connecting multiple end-to-end systems like a router, provides a universal, flexible, and efficient solution. First, Qiao is developed using the high-level programming language Go, which allows it to run on any Unix-compatible device, making it highly versatile. Second, Qiao leverages the advantages of multi-threading and concurrency to provide low-latency data forwarding capabilities. Third, Qiao adopts a layered and modular design, with all its source code being open-source, enabling anyone to customize their routing protocols according to their collaborative requirements in the industrial internet.
Zhiyu Wu, Yisu Wang
CSCWD1
2024 Image-Feature Weak-to-Strong Consistency: An Enhanced Paradigm for Semi-supervised Learning
Zhiyu Wu, Jinshi Cui
ECCV (15)1
2024 AllMatch: Exploiting All Unlabeled Data for Semi-Supervised Learning
Zhiyu Wu, Jinshi Cui
IJCAI1
2024 A Wolf in Sheep's Clothing: Practical Black-box Adversarial Attacks for Evading Learning-based Windows Malware Detection in the Wild
Xiang Ling 0001, Zhiyu Wu, Bin Wang 0062, JingZheng Wu, Shouling Ji, Tianyue Luo
USENIX Security Symposium2
2023 BERT-ERC: Fine-Tuning BERT Is Enough for Emotion Recognition in Conversation
abstract
Previous works on emotion recognition in conversation (ERC) follow a two-step paradigm, which can be summarized as first producing context-independent features via fine-tuning pretrained language models (PLMs) and then analyzing contextual information and dialogue structure information among the extracted features. However, we discover that this paradigm has several limitations. Accordingly, we propose a novel paradigm, i.e., exploring contextual information and dialogue structure information in the fine-tuning step, and adapting the PLM to the ERC task in terms of input text, classification structure, and training strategy. Furthermore, we develop our model BERT-ERC according to the proposed paradigm, which improves ERC performance in three aspects, namely suggestive text, fine-grained classification module, and two-stage training. Compared to existing methods, BERT-ERC achieves substantial improvement on four datasets, indicating its effectiveness and generalization capability. Besides, we also set up the limited resources scenario and the online prediction scenario to approximate real-world scenarios. Extensive experiments demonstrate that the proposed paradigm significantly outperforms the previous one and can be adapted to various scenes.
Xiangyu Qin, Zhiyu Wu, Yanran Li, Jian Luan 0001, Bin Wang 0004, Li Wang 0114, Jinshi Cui
AAAI2
2023 LA-Net: Landmark-Aware Learning for Reliable Facial Expression Recognition under Label Noise
abstract
Facial expression recognition (FER) remains a challenging task due to the ambiguity of expressions. The derived noisy labels significantly harm the performance in real-world scenarios. To address this issue, we present a new FER model named Landmark-Aware Net (LA-Net), which leverages facial landmarks to mitigate the impact of label noise from two perspectives. Firstly, LA-Net uses landmark information to suppress the uncertainty in expression space and constructs the label distribution of each sample by neighborhood aggregation, which in turn improves the quality of training supervision. Secondly, the model incorporates landmark information into expression representations using the devised expression-landmark contrastive loss. The enhanced expression feature extractor can be less susceptible to label noise. Our method can be integrated with any deep neural network for better training supervision without introducing extra inference costs. We conduct extensive experiments on both in-the-wild datasets and synthetic noisy datasets and demonstrate that LA-Net achieves state-of-the-art performance.
Zhiyu Wu, Jinshi Cui
ICCV1
2023 Detecting backdoor in deep neural networks via intentional adversarial perturbations
Mingfu Xue, Yinghao Wu, Zhiyu Wu, Yushu Zhang 0001, Jian Wang 0038, Weiqiang Liu 0001
Inf. Sci.3
2021 SocialGuard: An adversarial example based privacy-preserving technique for social images
Mingfu Xue, Shichang Sun, Zhiyu Wu, Can He, Jian Wang 0038, Weiqiang Liu 0001
J. Inf. Secur. Appl.3
2020 Active DNN IP Protection: A Novel User Fingerprint Management and DNN Authorization Control Technique
abstract
The training process of deep learning model is costly. As such, deep learning model can be treated as an intellectual property (IP) of the model creator. However, a pirate can illegally copy, redistribute or abuse the model without permission. In recent years, a few Deep Neural Networks (DNN) IP protection works have been proposed. However, most of existing works passively verify the copyright of the model after the piracy occurs, and lack of user identity management, thus cannot provide commercial copyright management functions. In this paper, a novel user fingerprint management and DNN authorization control technique based on backdoor is proposed to provide active DNN IP protection. The proposed method can not only verify the ownership of the model, but can also authenticate and manage the user's unique identity, so as to provide a commercially applicable DNN IP management mechanism. Experimental results on CIFAR-10, CIFAR-100 and Fashion-MNIST datasets show that the proposed method can achieve high detection rate for user authentication (up to 100% in the three datasets). Illegal users with forged fingerprints cannot pass authentication as the detection rates are all 0 % in the three datasets. Model owner can verify his ownership since he can trigger the backdoor with a high confidence. In addition, the accuracy drops are only 0.52%, 1.61 % and -0.65% on CIFAR-10, CIFAR-100 and Fashion-MNIST, respectively, which indicate that the proposed method will not affect the performance of the DNN models. The proposed method is also robust to model fine-tuning and pruning attacks. The detection rates for owner verification on CIFAR-10, CIFAR-100 and Fashion-MNIST are all 100% after model pruning attack, and are 90 %, 83 % and 93 % respectively after model fine-tuning attack, on the premise that the attacker wants to preserve the accuracy of the model.
Mingfu Xue, Zhiyu Wu, Can He, Jian Wang 0038, Weiqiang Liu 0001
TrustCom2