Zhiyuan Fang

dblp:75/4027 · also Fang Zhiyuan · DBLP profile ↗
← Back
20ranked-venue papers
9as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 8 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Fate: Fasss sEsdge Inference of Mixture-of-Experts Models via Cross-Layer Gate
abstract
With the rapid growth and rising complexity of web content, edge-deployed LLMs have become essential for enhancing users' online experiences. However, sparsely-activated Mixture-of-Experts (MoE) models, which are well-suited for edge scenarios, face significant memory bottleneck challenges. Offload-based methods have been proposed to mitigate the problem, but they face difficulties with expert prediction. To promote the application of MoE models in edge scenarios, we propose Fate, an offloading system designed for MoE models to enable efficient inference in resource-constrained environments. The key insight behind Fate is that gate inputs from adjacent layers can be effectively used for expert prefetching, achieving high prediction accuracy. Furthermore, Fate employs a shallow-favoring expert caching strategy that increases the expert hit rate to 99%. Additionally, Fate integrates tailored quantization strategies for cache optimization and I/O efficiency. Experimental results show that, compared to baselines, Fate achieves up to 1.34×-5.07× prefill speedup and 1.26×-4.41× decoding speedup, while maintaining inference quality.
Zhiyuan Fang, Xingfan Yu, Yuegui Huang, Zicong Hong, Yufeng Lyu, Wuhui Chen, Yue Yu 0001, Fan Yu 0004
WWW1
2026 TransformKV: Optimizing Multi-Turn Conversational Services in LLMs via KV Cache Transformation
abstract
Multi-turn conversational systems based on large language models are increasingly being integrated into web platforms and applied across a wide range of domains. However, these systems typically combine the userߣs current query with contextual information from previous interactions, resulting in continuously expanding input prompts. This leads to a significant increase in time-to-first-token (TTFT), causing intolerable delays in web response times. To address this issue, we introduce TransformKV, which maximizes the reuse of the KV cache from previous conversations rather than recomputing, thereby reducing TTFT latency. TransformKV first identifies the specific locations that require transformation to maximize the reuse of the KV cache with minimal operations. It then efficiently transforms the KV cache for a subset of tokens by recomputing only the KV cache that impact semantics. Additionally, TransformKV further reduces TTFT latency by performing only QKV computations in certain layers while skipping other computations that contribute less to overall performance. Experimental results demonstrate that in multi-turn conversation tasks, TransformKV can reduce inference latency by up to 30%, achieving up to a 1.8× improvement in performance compared to similar approaches. Notably, as the context window size increases, the performance gains become even more pronounced.
Jiahang Zhou, Zhiyuan Fang, Yusheng Qin, Wuhui Chen, Tao Zhang 0096, Chuanfu Zhang, Zibin Zheng
IEEE Trans. Computers2
2025 Klotski: Efficient Mixture-of-Expert Inference via Expert-Aware Multi-Batch Pipeline
abstract
Mixture of Experts (MoE), with its distinctive sparse structure, enables the scaling of language models up to trillions of parameters without significantly increasing computational costs. However, the substantial parameter size presents a challenge for inference, as the expansion in GPU memory cannot keep pace with the growth in parameters. Although offloading techniques utilise memory from the CPU and disk and parallelise the I/O and computation for efficiency, the computation for each expert in MoE models is often less than the I/O, resulting in numerous bubbles in the pipeline.
Zhiyuan Fang, Yuegui Huang, Zicong Hong, Yufeng Lyu, Wuhui Chen, Yue Yu 0001, Fan Yu 0004, Zibin Zheng
ASPLOS (2)1
2025 A3GS: Arbitrary Artistic Style into Arbitrary 3D Gaussian Splatting
Zhiyuan Fang, Rengan Xie, Xuancheng Jin, Qi Ye 0001, Wei Chen 0001, Wenting Zheng, Rui Wang 0004, Yuchi Huo
ICCV1
2025 SmartOracle: Generating Smart Contract Oracle via Fine-Grained Invariant Detection
abstract
As decentralized applications (DApps) proliferate, the increased complexity and usage of smart contracts have heightened their susceptibility to security incidents and financial losses. Although various vulnerability detection tools have been developed to mitigate these issues, they often suffer poor performance in detecting vulnerabilities, as they either rely on simplistic and general-purpose oracles that may be inadequate for vulnerability detection, or require user-specified oracles, which are labor-intensive to create. In this paper, we introduce SmartOracle, a dynamic invariant detector that automatically generates fine-grained invariants as application-specific oracles for vulnerability detection. From historical transactions, SmartOracle uses pattern-based detection and advanced inference to construct comprehensive properties, and mines multi-layerlikelyinvariants to accommodate the complicated contract functionalities. After that, SmartOracle identifies smart contract vulnerabilities by hunting the violated invariants in new transactions. In the field of invariant detection, SmartOracle detects 50% more ERC20 invariants than existing dynamic invariant detection and achieves 96% precision rate. Furthermore, we build a dataset that contains vulnerable contracts from real-world security incidents. SmartOracle successfully detects 466 abnormal transactions with an acceptable precision rate 96%, involving 31 vulnerable contracts. The experimental results demonstrate its effectiveness in detecting smart contract vulnerabilities, especially those related to complicated contract functionalities.
Jianzhong Su, Jiachi Chen, Zhiyuan Fang, Xingwei Lin, Yutian Tang, Zibin Zheng
IEEE Trans. Software Eng.3
2024 Skews in the Phenomenon Space Hinder Generalization in Text-to-Image Generation
Yingshan Chang, Yasi Zhang, Zhiyuan Fang, Ying Nian Wu, Yonatan Bisk, Feng Gao 0013
ECCV (87)3
2023 End-to-end Knowledge Retrieval with Multi-modal Queries
abstract
We investigate knowledge retrieval with multimodal queries, i.e. queries containing information split across image and text inputs, a challenging task that differs from previous work on cross-modal retrieval.We curate a new dataset called ReMuQ 1 for benchmarking progress on this task.ReMuQ requires a system to retrieve knowledge from a large corpus by integrating contents from both text and image queries.We introduce a retriever model "ReViz" that can directly process input text and images to retrieve relevant knowledge in an end-to-end fashion without being dependent on intermediate modules such as object detectors or caption generators.We introduce a new pretraining task that is effective for learning knowledge retrieval with multimodal queries and also improves performance on downstream tasks.We demonstrate superior performance in retrieval on two datasets (ReMuQ and OK-VQA) under zeroshot settings as well as further improvements when finetuned on these datasets.
Man Luo 0003, Zhiyuan Fang, Tejas Gokhale, Yezhou Yang, Chitta Baral
ACL (1)2
2023 DeFiWarder: Protecting DeFi Apps from Token Leaking Vulnerabilities
abstract
Decentralized Finance (DeFi) apps have rapidly proliferated with the development of blockchain and smart contracts, whose maximum total value locked (TVL) has exceeded 100 billion dollars in the past few years. These apps allow users to interact and perform complicated financial activities. However, the vulnerabilities hiding in the smart contracts of DeFi apps have resulted in numerous security incidents, with most of them leading to funds (tokens) leaking and resulting in severe financial loss. In this paper, we summarize Token Leaking vulnerability of DeFi apps, which enable someone to abnormally withdraw funds that far exceed their deposits. Due to the massive amount of funds in DeFi apps, it is crucial to protect DeFi apps from Token Leaking vulnerabilities. Unfortunately, existing tools have limitations in addressing this vulnerability. To address this issue, we propose DeFiWarder, a tool that traces on-chain transactions and protects DeFi apps from Token Leaking vulnerabilities. Specifically, DeFiWarder first records the execution logs (traces) of smart contracts. It then accurately recovers token transfers within transactions to catch the funds flow between users and DeFi apps, as well as the relations between users based on role mining. Finally, DeFiWarder utilizes anomaly detection to reveal Token Leaking vulnerabilities and related attack behaviors. We conducted experiments to demonstrate the effectiveness and efficiency of DeFiWarder. Specifically, DeFi-Warder successfully revealed 25 Token Leaking vulnerabilities from 30 Defi apps. Moreover, its efficiency supports real-time detection of token leaking within on-chain transactions. In addition, we summarize five major reasons for Token Leaking vulnerability to assist DeFi apps in protecting their funds.
Jianzhong Su, Xingwei Lin, Zhiyuan Fang, Zhirong Zhu, Jiachi Chen, Zibin Zheng, Jiashui Wang
ASE3
2023 Hierarchical context-agnostic network with contrastive feature diversity for one-shot semantic segmentation
Zhiyuan Fang, Guangyu Gao, Zekang Zhang, Anqi Zhang 0002
J. Vis. Commun. Image Represent.1
2022 Injecting Semantic Concepts into End-to-End Image Captioning
abstract
Tremendous progresses have been made in recent years in developing better image captioning models, yet most of them rely on a separate object detector to extract regional features. Recent vision-language studies are shifting towards the detector-free trend by leveraging grid representations for more flexible model training and faster inference speed. However, such development is primarily focused on image understanding tasks, and remains less investigated for the caption generation task. In this paper, we are concerned with a better-performing detector-free image captioning model, and propose a pure vision transformer-based image captioning model, dubbed as ViTCAP, in which grid representations are used without extracting the regional features. For improved performance, we introduce a novel Concept Token Network (CTN) to predict the semantic concepts and then incorporate them into the end-to-end captioning. In particular, the CTN is built on the basis of a vision transformer, and is designed to predict the concept tokens through a classification task, from which the rich semantic information contained greatly benefits the captioning task. Compared with the previous detector-based models, ViTCAP drastically simplifies the architectures and at the same time achieves competitive performance on various challenging image captioning datasets. In particular, ViTCAP reaches 138.1 CIDEr scores on COCO-caption Karpathy-split, 93.8 and 108.6 CIDEr scores on nocaps and Google-CC captioning datasets, respectively.
Zhiyuan Fang, Xiaowei Hu 0006, Zhe Gan, Yezhou Yang, Zicheng Liu 0001
CVPR1
2022 CAVAN: Commonsense Knowledge Anchored Video Captioning
abstract
It is not merely an aggregation of static entities that a video clip carries, but also a variety of interactions and relations among these entities. Challenges still remain for a video captioning system to generate descriptions focusing on the prominent interest and aligning with the latent aspects beyond observations. In this work, we present a Commonsense knowledge Anchored Video cAptioNing(dubbed as CAVAN) approach. CAVAN exploits inferential commonsense knowledge to assist the training of video captioning model with a novel paradigm for sentence-level semantic alignment. Specifically, we acquire commonsense knowledge complementing per training caption by querying a generic knowledge atlas (ATOMIC [1]), and form the commonsense-caption entailment corpus. A BERT [2] based language entailment model trained from this corpus then serves as a commonsense discriminator for the training of video captioning model, and penalizes the model from generating semantically misaligned captions. Experimental results with ablations on MSRVTT [3], V2C [4] and VATEX [5] datasets validate the effectiveness of CAVAN and reveal that the use of commonsense knowledge benefits video caption generation.
Huiliang Shao, Zhiyuan Fang, Yezhou Yang
ICPR2
2022 Mining Unseen Classes via Regional Objectness: A Simple Baseline for Incremental Segmentation
abstract
Incremental or continual learning has been extensively studied for image classification tasks to alleviate catastrophic forgetting, a phenomenon in which earlier learned knowledge is forgotten when learning new concepts. For class incremental semantic segmentation, such a phenomenon often becomes much worse due to the semantic shift of the background class, \ie, some concepts learned at previous stages are assigned to the background class at the current training stage, therefore, significantly reducing the performance of these old concepts. To address this issue, we propose a simple yet effective method in this paper, named Mining unseen Classes via Regional Objectness (MicroSeg). Our MicroSeg is based on the assumption that \emph{background regions with strong objectness possibly belong to those concepts in the historical or future stages}. Therefore, to avoid forgetting old knowledge at the current training stage, our MicroSeg first splits the given image into hundreds of segment proposals with a proposal generator. Those segment proposals with strong objectness from the background are then clustered and assigned new defined labels during the optimization. In this way, the distribution characterizes of old concepts in the feature space could be better perceived, relieving the catastrophic forgetting caused by the semantic shift of the background class accordingly. We conduct extensive experiments on Pascal VOC and ADE20K, and competitive results well demonstrate the effectiveness of our MicroSeg. Code is available at \href{https://github.com/zkzhang98/MicroSeg}{\textcolor{orange}{\texttt{https://github.com/zkzhang98/MicroSeg}}}.
Zekang Zhang, Guangyu Gao, Zhiyuan Fang, Jianbo Jiao, Yunchao Wei
NeurIPS3
2022 DRNet: Double Recalibration Network for Few-Shot Semantic Segmentation
abstract
Few-shot segmentation aims at learning to segment query images guided by only a few annotated images from the support set. Previous methods rely on mining the feature embedding similarity across the query and the support images to achieve successful segmentation. However, these models tend to perform badly in cases where the query instances have a large variance from the support ones. To enhance model robustness against such intra-class variance, we propose a Double Recalibration Network (DRNet) with two recalibration modules, i.e., the Self-adapted Recalibration (SR) module and the Cross-attended Recalibration (CR) module. In particular, beyond learning robust feature embedding for pixel-wise comparison between support and query as in conventional methods, the DRNet further exploits semantic-aware knowledge embedded in the query image to help segment itself, which we call ‘self-adapted recalibration’. More specifically, DRNet first employs guidance from the support set to roughly predict an incomplete but correct initial object region for the query image, and then reversely uses the feature embedding extracted from the incomplete object region to segment the query image. Also, we devise a CR module to refine the feature representation of the query image by propagating the underlying knowledge embedded in the support image’s foreground to the query. Instead of foreground global pooling, we refine the response at each pixel in the query feature map by attending to all foreground pixels in the support feature map and taking the weighted average by their similarity; meanwhile, feature maps of the query image are also added back to weighted feature maps as a residual connection. Our DRNet can effectively address the intra-class variance under the few-shot setting with such two recalibration modules, and mine more accurate target regions for query images. We conduct extensive experiments on the popular benchmarks PASCAL-$5^{i}$and COCO-$20^{i}$. The DRNet with the best configuration achieves the mIoU of$\textbf {63.6}\%$and$\textbf {64.9}\%$on PASCAL-$5^{i}$and$\textbf {44.7}\%$and$\textbf {49.6}\%$on COCO-$20^{i}$for 1-shot and 5-shot settings respectively, significantly outperforming the state-of-the-arts without any bells and whistles. Code is available at:https://github.com/fangzy97/drnet.
Guangyu Gao, Zhiyuan Fang, Cen Han, Yunchao Wei, Chi Harold Liu, Shuicheng Yan
IEEE Trans. Image Process.2
2021 Compressing Visual-linguistic Model via Knowledge Distillation
abstract
Despite exciting progress in pre-training for visual-linguistic (VL) representations, very few aspire to a small VL model. In this paper, we study knowledge distillation (KD) to effectively compress a transformer based large VL model into a small VL model. The major challenge arises from the inconsistent regional visual tokens extracted from different detectors of Teacher and Student, resulting in the misalignment of hidden representations and attention distributions. To address the problem, we retrain and adapt the Teacher by using the same region proposals from Student’s detector while the features are from Teacher’s own object detector. With aligned network inputs, the adapted Teacher is capable of transferring the knowledge through the intermediate representations. Specifically, we use the mean square error loss to mimic the attention distribution inside the transformer block, and present a token-wise noise contrastive loss to align the hidden state by contrasting with negative representations stored in a sample queue. To this end, we show that our proposed distillation significantly improves the performance of small VL models on image captioning and visual question answering tasks. It reaches 120.8 in CIDEr score on COCO captioning, an improvement of 5.1 over its non-distilled counterpart; and an accuracy of 69.8 on VQA 2.0, a 0.8 gain from the baseline. Our extensive experiments and ablations confirm the effectiveness of VL distillation in both pre-training and fine-tuning stages.
Zhiyuan Fang, Xiaowei Hu 0006, Yezhou Yang, Zicheng Liu 0001
ICCV1
2021 SEED: Self-supervised Distillation For Visual Representation
Zhiyuan Fang, Lei Zhang 0001, Yezhou Yang, Zicheng Liu 0001
ICLR1
2021 HRDNet: High-Resolution Detection Network for Small Objects
abstract
Small object detection is a very challenging yet practical vision task. With deep network-based methods, the contextual information of small objects may disappear when the network goes deeper. An intuitive solution to alleviate this issue is to increase the input resolution, however, it will aggravate the large variant of object scale and introduce unbearable computation cost. To leverage the benefits of high-resolution images without bringing up new problems, we propose a High-Resolution Detection Network (HRDNet) which takes multiple resolution inputs with multi-depth backbones. Meanwhile, we propose the Multi-Depth Image Pyramid Network (MD-IPN) and Multi-Scale Feature Pyramid Network (MS-FPN). The MD-IPN maintains multiple position information using multiple depth backbones. Specifically, high-resolution input will be fed into a shallow network to reserve more positional information and reduce computational costs, while low-resolution input will be fed into a deep network to extract more semantics. By extracting various features from high to low resolutions, the MD-IPN can improve the performance of small object detection and maintain the performance of middle and large objects. Additionally, MS-FPN is introduced to align and fuse multi-scale feature groups generated by MD-IPN to reduce the information imbalance. Extensive experiments are conducted on the COCO2017 and the typical small object dataset, VisDrone 2019. Notably, our HRDNet achieves the state-of-the-art on these two datasets with significant improvements on small objects.
Ziming Liu 0003, Guangyu Gao, Lin Sun 0004, Zhiyuan Fang
ICME4
2020 ViTAA: Visual-Textual Attributes Alignment in Person Search by Natural Language
Zhe Wang 0013, Zhiyuan Fang, Jun Wang 0041, Yezhou Yang
ECCV (12)2
2020 Video2Commonsense: Generating Commonsense Descriptions to Enrich Video Captioning
abstract
Captioning is a crucial and challenging task for video understanding.In videos that involve active agents such as humans, the agent's actions can bring about myriad changes in the scene.Observable changes such as movements, manipulations, and transformations of the objects in the scene, are reflected in conventional video captioning.Unlike images, actions in videos are also inherently linked to social aspects such as intentions (why the action is taking place), effects (what changes due to the action), and attributes that describe the agent.Thus for video understanding, such as when captioning videos or when answering questions about videos, one must have an understanding of these commonsense aspects.We present the first work on generating commonsense captions directly from videos, to describe latent aspects such as intentions, effects, and attributes.We present a new dataset "Video-to-Commonsense (V2C)" that contains ∼ 9k videos of human agents performing various actions, annotated with 3 types of commonsense descriptions.Additionally we explore the use of open-ended video-based commonsense question answering (V2C-QA) as a way to enrich our captions.Both the generation task and the QA task can be used to enrich video captions.. frame frame frame CNN
Zhiyuan Fang, Tejas Gokhale, Pratyay Banerjee, Chitta Baral, Yezhou Yang
EMNLP (1)1
2019 Modularized Textual Grounding for Counterfactual Resilience
abstract
Computer Vision applications often require a textual grounding module with precision, interpretability, and resilience to counterfactual inputs/queries. To achieve high grounding precision, current textual grounding methods heavily rely on large-scale training data with manual annotations at the pixel level. Such annotations are expensive to obtain and thus severely narrow the model's scope of real-world applications. Moreover, most of these methods sacrifice interpretability, generalizability, and they neglect the importance of being resilient to counterfactual inputs. To address these issues, we propose a visual grounding system which is 1) end-to-end trainable in a weakly supervised fashion with only image-level annotations, and 2) counterfactually resilient owing to the modular design. Specifically, we decompose textual descriptions into three levels: entity, semantic attribute, color information, and perform compositional grounding progressively. We validate our model through a series of experiments and demonstrate its improvement over the state-of-the-art methods. In particular, our model's performance not only surpasses other weakly/un-supervised methods and even approaches the strongly supervised ones, but also is interpretable for decision making and performs much better in face of counterfactual classes than all the others.
Zhiyuan Fang, Shu Kong, Charless C. Fowlkes, Yezhou Yang
CVPR1
2017 Range Loss for Deep Face Recognition with Long-Tailed Training Data
abstract
Deep convolutional neural networks have achieved significant improvements on face recognition task due to their ability to learn highly discriminative features from tremendous amounts of face images. Many large scale face datasets exhibit long-tail distribution where a small number of entities (persons) have large number of face images while a large number of persons only have very few face samples (long tail). Most of the existing works alleviate this problem by simply cutting the tailed data and only keep identities with enough number of examples. Unlike these work, this paper investigated how long-tailed data impact the training of face CNNs and develop a novel loss function, called range loss, to effectively utilize the tailed data in training process. More specifically, range loss is designed to reduce overall intrapersonal variations while enlarge interpersonal differences simultaneously. Extensive experiments on two face recognition benchmarks, Labeled Faces in the Wild (LFW) [11] and YouTube Faces (YTF) [33], demonstrate the effectiveness of the proposed range loss in overcoming the long tail effect, and show the good generalization ability of the proposed methods.
Xiao Zhang 0024, Zhiyuan Fang, Yandong Wen, Zhifeng Li 0001, Yu Qiao 0001
ICCV2