EDBT 2026 Demo / reviewers in the wild / expert
Zhiyuan Fang
dblp:75/4027 · also Fang Zhiyuan
· DBLP profile ↗
20ranked-venue papers
9as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 6 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 8 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fate: Fasss sEsdge Inference of Mixture-of-Experts Models via Cross-Layer GateabstractWith the rapid growth and rising complexity of web content, edge-deployed LLMs have become essential for enhancing users' online experiences. However, sparsely-activated Mixture-of-Experts (MoE) models, which are well-suited for edge scenarios, face significant memory bottleneck challenges. Offload-based methods have been proposed to mitigate the problem, but they face difficulties with expert prediction. To promote the application of MoE models in edge scenarios, we propose Fate, an offloading system designed for MoE models to enable efficient inference in resource-constrained environments. The key insight behind Fate is that gate inputs from adjacent layers can be effectively used for expert prefetching, achieving high prediction accuracy. Furthermore, Fate employs a shallow-favoring expert caching strategy that increases the expert hit rate to 99%. Additionally, Fate integrates tailored quantization strategies for cache optimization and I/O efficiency. Experimental results show that, compared to baselines, Fate achieves up to 1.34×-5.07× prefill speedup and 1.26×-4.41× decoding speedup, while maintaining inference quality. Zhiyuan Fang, Xingfan Yu, Yuegui Huang, Zicong Hong, Yufeng Lyu, Wuhui Chen, Yue Yu 0001, Fan Yu 0004 |
WWW | 1 |
| 2026 | TransformKV: Optimizing Multi-Turn Conversational Services in LLMs via KV Cache TransformationabstractMulti-turn conversational systems based on large language models are increasingly being integrated into web platforms and applied across a wide range of domains. However, these systems typically combine the userߣs current query with contextual information from previous interactions, resulting in continuously expanding input prompts. This leads to a significant increase in time-to-first-token (TTFT), causing intolerable delays in web response times. To address this issue, we introduce TransformKV, which maximizes the reuse of the KV cache from previous conversations rather than recomputing, thereby reducing TTFT latency. TransformKV first identifies the specific locations that require transformation to maximize the reuse of the KV cache with minimal operations. It then efficiently transforms the KV cache for a subset of tokens by recomputing only the KV cache that impact semantics. Additionally, TransformKV further reduces TTFT latency by performing only QKV computations in certain layers while skipping other computations that contribute less to overall performance. Experimental results demonstrate that in multi-turn conversation tasks, TransformKV can reduce inference latency by up to 30%, achieving up to a 1.8× improvement in performance compared to similar approaches. Notably, as the context window size increases, the performance gains become even more pronounced. Jiahang Zhou, Zhiyuan Fang, Yusheng Qin, Wuhui Chen, Tao Zhang 0096, Chuanfu Zhang, Zibin Zheng |
IEEE Trans. Computers | 2 |
| 2025 | Klotski: Efficient Mixture-of-Expert Inference via Expert-Aware Multi-Batch PipelineabstractMixture of Experts (MoE), with its distinctive sparse structure, enables the scaling of language models up to trillions of parameters without significantly increasing computational costs. However, the substantial parameter size presents a challenge for inference, as the expansion in GPU memory cannot keep pace with the growth in parameters. Although offloading techniques utilise memory from the CPU and disk and parallelise the I/O and computation for efficiency, the computation for each expert in MoE models is often less than the I/O, resulting in numerous bubbles in the pipeline. Zhiyuan Fang, Yuegui Huang, Zicong Hong, Yufeng Lyu, Wuhui Chen, Yue Yu 0001, Fan Yu 0004, Zibin Zheng |
ASPLOS (2) | 1 |
| 2025 | A3GS: Arbitrary Artistic Style into Arbitrary 3D Gaussian Splatting
Zhiyuan Fang, Rengan Xie, Xuancheng Jin, Qi Ye 0001, Wei Chen 0001, Wenting Zheng, Rui Wang 0004, Yuchi Huo |
ICCV | 1 |
| 2025 | SmartOracle: Generating Smart Contract Oracle via Fine-Grained Invariant DetectionabstractAs decentralized applications (DApps) proliferate, the increased complexity and usage of smart contracts have heightened their susceptibility to security incidents and financial losses. Although various vulnerability detection tools have been developed to mitigate these issues, they often suffer poor performance in detecting vulnerabilities, as they either rely on simplistic and general-purpose oracles that may be inadequate for vulnerability detection, or require user-specified oracles, which are labor-intensive to create. In this paper, we introduce SmartOracle, a dynamic invariant detector that automatically generates fine-grained invariants as application-specific oracles for vulnerability detection. From historical transactions, SmartOracle uses pattern-based detection and advanced inference to construct comprehensive properties, and mines multi-layerlikelyinvariants to accommodate the complicated contract functionalities. After that, SmartOracle identifies smart contract vulnerabilities by hunting the violated invariants in new transactions. In the field of invariant detection, SmartOracle detects 50% more ERC20 invariants than existing dynamic invariant detection and achieves 96% precision rate. Furthermore, we build a dataset that contains vulnerable contracts from real-world security incidents. SmartOracle successfully detects 466 abnormal transactions with an acceptable precision rate 96%, involving 31 vulnerable contracts. The experimental results demonstrate its effectiveness in detecting smart contract vulnerabilities, especially those related to complicated contract functionalities. Jianzhong Su, Jiachi Chen, Zhiyuan Fang, Xingwei Lin, Yutian Tang, Zibin Zheng |
IEEE Trans. Software Eng. | 3 |
| 2024 | Skews in the Phenomenon Space Hinder Generalization in Text-to-Image Generation
Yingshan Chang, Yasi Zhang, Zhiyuan Fang, Ying Nian Wu, Yonatan Bisk, Feng Gao 0013 |
ECCV (87) | 3 |
| 2023 | End-to-end Knowledge Retrieval with Multi-modal QueriesabstractWe investigate knowledge retrieval with multimodal queries, i.e. queries containing information split across image and text inputs, a challenging task that differs from previous work on cross-modal retrieval.We curate a new dataset called ReMuQ 1 for benchmarking progress on this task.ReMuQ requires a system to retrieve knowledge from a large corpus by integrating contents from both text and image queries.We introduce a retriever model "ReViz" that can directly process input text and images to retrieve relevant knowledge in an end-to-end fashion without being dependent on intermediate modules such as object detectors or caption generators.We introduce a new pretraining task that is effective for learning knowledge retrieval with multimodal queries and also improves performance on downstream tasks.We demonstrate superior performance in retrieval on two datasets (ReMuQ and OK-VQA) under zeroshot settings as well as further improvements when finetuned on these datasets. Man Luo 0003, Zhiyuan Fang, Tejas Gokhale, Yezhou Yang, Chitta Baral |
ACL (1) | 2 |
| 2023 | DeFiWarder: Protecting DeFi Apps from Token Leaking VulnerabilitiesabstractDecentralized Finance (DeFi) apps have rapidly proliferated with the development of blockchain and smart contracts, whose maximum total value locked (TVL) has exceeded 100 billion dollars in the past few years. These apps allow users to interact and perform complicated financial activities. However, the vulnerabilities hiding in the smart contracts of DeFi apps have resulted in numerous security incidents, with most of them leading to funds (tokens) leaking and resulting in severe financial loss. In this paper, we summarize Token Leaking vulnerability of DeFi apps, which enable someone to abnormally withdraw funds that far exceed their deposits. Due to the massive amount of funds in DeFi apps, it is crucial to protect DeFi apps from Token Leaking vulnerabilities. Unfortunately, existing tools have limitations in addressing this vulnerability. To address this issue, we propose DeFiWarder, a tool that traces on-chain transactions and protects DeFi apps from Token Leaking vulnerabilities. Specifically, DeFiWarder first records the execution logs (traces) of smart contracts. It then accurately recovers token transfers within transactions to catch the funds flow between users and DeFi apps, as well as the relations between users based on role mining. Finally, DeFiWarder utilizes anomaly detection to reveal Token Leaking vulnerabilities and related attack behaviors. We conducted experiments to demonstrate the effectiveness and efficiency of DeFiWarder. Specifically, DeFi-Warder successfully revealed 25 Token Leaking vulnerabilities from 30 Defi apps. Moreover, its efficiency supports real-time detection of token leaking within on-chain transactions. In addition, we summarize five major reasons for Token Leaking vulnerability to assist DeFi apps in protecting their funds. Jianzhong Su, Xingwei Lin, Zhiyuan Fang, Zhirong Zhu, Jiachi Chen, Zibin Zheng, Jiashui Wang |
ASE | 3 |
| 2023 | Hierarchical context-agnostic network with contrastive feature diversity for one-shot semantic segmentation
Zhiyuan Fang, Guangyu Gao, Zekang Zhang, Anqi Zhang 0002 |
J. Vis. Commun. Image Represent. | 1 |
| 2022 | Injecting Semantic Concepts into End-to-End Image CaptioningabstractTremendous progresses have been made in recent years in developing better image captioning models, yet most of them rely on a separate object detector to extract regional features. Recent vision-language studies are shifting towards the detector-free trend by leveraging grid representations for more flexible model training and faster inference speed. However, such development is primarily focused on image understanding tasks, and remains less investigated for the caption generation task. In this paper, we are concerned with a better-performing detector-free image captioning model, and propose a pure vision transformer-based image captioning model, dubbed as ViTCAP, in which grid representations are used without extracting the regional features. For improved performance, we introduce a novel Concept Token Network (CTN) to predict the semantic concepts and then incorporate them into the end-to-end captioning. In particular, the CTN is built on the basis of a vision transformer, and is designed to predict the concept tokens through a classification task, from which the rich semantic information contained greatly benefits the captioning task. Compared with the previous detector-based models, ViTCAP drastically simplifies the architectures and at the same time achieves competitive performance on various challenging image captioning datasets. In particular, ViTCAP reaches 138.1 CIDEr scores on COCO-caption Karpathy-split, 93.8 and 108.6 CIDEr scores on nocaps and Google-CC captioning datasets, respectively. Zhiyuan Fang, Xiaowei Hu 0006, Zhe Gan, Yezhou Yang, Zicheng Liu 0001 |
CVPR | 1 |
| 2022 | CAVAN: Commonsense Knowledge Anchored Video CaptioningabstractIt is not merely an aggregation of static entities that a video clip carries, but also a variety of interactions and relations among these entities. Challenges still remain for a video captioning system to generate descriptions focusing on the prominent interest and aligning with the latent aspects beyond observations. In this work, we present a Commonsense knowledge Anchored Video cAptioNing(dubbed as CAVAN) approach. CAVAN exploits inferential commonsense knowledge to assist the training of video captioning model with a novel paradigm for sentence-level semantic alignment. Specifically, we acquire commonsense knowledge complementing per training caption by querying a generic knowledge atlas (ATOMIC [1]), and form the commonsense-caption entailment corpus. A BERT [2] based language entailment model trained from this corpus then serves as a commonsense discriminator for the training of video captioning model, and penalizes the model from generating semantically misaligned captions. Experimental results with ablations on MSRVTT [3], V2C [4] and VATEX [5] datasets validate the effectiveness of CAVAN and reveal that the use of commonsense knowledge benefits video caption generation. Huiliang Shao, Zhiyuan Fang, Yezhou Yang |
ICPR | 2 |
| 2022 | Mining Unseen Classes via Regional Objectness: A Simple Baseline for Incremental SegmentationabstractIncremental or continual learning has been extensively studied for image classification tasks to alleviate catastrophic forgetting, a phenomenon in which earlier learned knowledge is forgotten when learning new concepts. For class incremental semantic segmentation, such a phenomenon often becomes much worse due to the semantic shift of the background class, \ie, some concepts learned at previous stages are assigned to the background class at the current training stage, therefore, significantly reducing the performance of these old concepts. To address this issue, we propose a simple yet effective method in this paper, named Mining unseen Classes via Regional Objectness (MicroSeg). Our MicroSeg is based on the assumption that \emph{background regions with strong objectness possibly belong to those concepts in the historical or future stages}. Therefore, to avoid forgetting old knowledge at the current training stage, our MicroSeg first splits the given image into hundreds of segment proposals with a proposal generator. Those segment proposals with strong objectness from the background are then clustered and assigned new defined labels during the optimization. In this way, the distribution characterizes of old concepts in the feature space could be better perceived, relieving the catastrophic forgetting caused by the semantic shift of the background class accordingly. We conduct extensive experiments on Pascal VOC and ADE20K, and competitive results well demonstrate the effectiveness of our MicroSeg. Code is available at \href{https://github.com/zkzhang98/MicroSeg}{\textcolor{orange}{\texttt{https://github.com/zkzhang98/MicroSeg}}}. Zekang Zhang, Guangyu Gao, Zhiyuan Fang, Jianbo Jiao, Yunchao Wei |
NeurIPS | 3 |
| 2022 | DRNet: Double Recalibration Network for Few-Shot Semantic SegmentationabstractFew-shot segmentation aims at learning to segment query images guided by only a few annotated images from the support set. Previous methods rely on mining the feature embedding similarity across the query and the support images to achieve successful segmentation. However, these models tend to perform badly in cases where the query instances have a large variance from the support ones. To enhance model robustness against such intra-class variance, we propose a Double Recalibration Network (DRNet) with two recalibration modules, i.e., the Self-adapted Recalibration (SR) module and the Cross-attended Recalibration (CR) module. In particular, beyond learning robust feature embedding for pixel-wise comparison between support and query as in conventional methods, the DRNet further exploits semantic-aware knowledge embedded in the query image to help segment itself, which we call ‘self-adapted recalibration’. More specifically, DRNet first employs guidance from the support set to roughly predict an incomplete but correct initial object region for the query image, and then reversely uses the feature embedding extracted from the incomplete object region to segment the query image. Also, we devise a CR module to refine the feature representation of the query image by propagating the underlying knowledge embedded in the support image’s foreground to the query. Instead of foreground global pooling, we refine the response at each pixel in the query feature map by attending to all foreground pixels in the support feature map and taking the weighted average by their similarity; meanwhile, feature maps of the query image are also added back to weighted feature maps as a residual connection. Our DRNet can effectively address the intra-class variance under the few-shot setting with such two recalibration modules, and mine more accurate target regions for query images. We conduct extensive experiments on the popular benchmarks PASCAL-$5^{i}$and COCO-$20^{i}$. The DRNet with the best configuration achieves the mIoU of$\textbf {63.6}\%$and$\textbf {64.9}\%$on PASCAL-$5^{i}$and$\textbf {44.7}\%$and$\textbf {49.6}\%$on COCO-$20^{i}$for 1-shot and 5-shot settings respectively, significantly outperforming the state-of-the-arts without any bells and whistles. Code is available at:https://github.com/fangzy97/drnet. Guangyu Gao, Zhiyuan Fang, Cen Han, Yunchao Wei, Chi Harold Liu, Shuicheng Yan |
IEEE Trans. Image Process. | 2 |
| 2021 | Compressing Visual-linguistic Model via Knowledge DistillationabstractDespite exciting progress in pre-training for visual-linguistic (VL) representations, very few aspire to a small VL model. In this paper, we study knowledge distillation (KD) to effectively compress a transformer based large VL model into a small VL model. The major challenge arises from the inconsistent regional visual tokens extracted from different detectors of Teacher and Student, resulting in the misalignment of hidden representations and attention distributions. To address the problem, we retrain and adapt the Teacher by using the same region proposals from Student’s detector while the features are from Teacher’s own object detector. With aligned network inputs, the adapted Teacher is capable of transferring the knowledge through the intermediate representations. Specifically, we use the mean square error loss to mimic the attention distribution inside the transformer block, and present a token-wise noise contrastive loss to align the hidden state by contrasting with negative representations stored in a sample queue. To this end, we show that our proposed distillation significantly improves the performance of small VL models on image captioning and visual question answering tasks. It reaches 120.8 in CIDEr score on COCO captioning, an improvement of 5.1 over its non-distilled counterpart; and an accuracy of 69.8 on VQA 2.0, a 0.8 gain from the baseline. Our extensive experiments and ablations confirm the effectiveness of VL distillation in both pre-training and fine-tuning stages. Zhiyuan Fang, Xiaowei Hu 0006, Yezhou Yang, Zicheng Liu 0001 |
ICCV | 1 |
| 2021 | SEED: Self-supervised Distillation For Visual Representation
Zhiyuan Fang, Lei Zhang 0001, Yezhou Yang, Zicheng Liu 0001 |
ICLR | 1 |
| 2021 | HRDNet: High-Resolution Detection Network for Small ObjectsabstractSmall object detection is a very challenging yet practical vision task. With deep network-based methods, the contextual information of small objects may disappear when the network goes deeper. An intuitive solution to alleviate this issue is to increase the input resolution, however, it will aggravate the large variant of object scale and introduce unbearable computation cost. To leverage the benefits of high-resolution images without bringing up new problems, we propose a High-Resolution Detection Network (HRDNet) which takes multiple resolution inputs with multi-depth backbones. Meanwhile, we propose the Multi-Depth Image Pyramid Network (MD-IPN) and Multi-Scale Feature Pyramid Network (MS-FPN). The MD-IPN maintains multiple position information using multiple depth backbones. Specifically, high-resolution input will be fed into a shallow network to reserve more positional information and reduce computational costs, while low-resolution input will be fed into a deep network to extract more semantics. By extracting various features from high to low resolutions, the MD-IPN can improve the performance of small object detection and maintain the performance of middle and large objects. Additionally, MS-FPN is introduced to align and fuse multi-scale feature groups generated by MD-IPN to reduce the information imbalance. Extensive experiments are conducted on the COCO2017 and the typical small object dataset, VisDrone 2019. Notably, our HRDNet achieves the state-of-the-art on these two datasets with significant improvements on small objects. Ziming Liu 0003, Guangyu Gao, Lin Sun 0004, Zhiyuan Fang |
ICME | 4 |
| 2020 | ViTAA: Visual-Textual Attributes Alignment in Person Search by Natural Language
Zhe Wang 0013, Zhiyuan Fang, Jun Wang 0041, Yezhou Yang |
ECCV (12) | 2 |
| 2020 | Video2Commonsense: Generating Commonsense Descriptions to Enrich Video CaptioningabstractCaptioning is a crucial and challenging task for video understanding.In videos that involve active agents such as humans, the agent's actions can bring about myriad changes in the scene.Observable changes such as movements, manipulations, and transformations of the objects in the scene, are reflected in conventional video captioning.Unlike images, actions in videos are also inherently linked to social aspects such as intentions (why the action is taking place), effects (what changes due to the action), and attributes that describe the agent.Thus for video understanding, such as when captioning videos or when answering questions about videos, one must have an understanding of these commonsense aspects.We present the first work on generating commonsense captions directly from videos, to describe latent aspects such as intentions, effects, and attributes.We present a new dataset "Video-to-Commonsense (V2C)" that contains ∼ 9k videos of human agents performing various actions, annotated with 3 types of commonsense descriptions.Additionally we explore the use of open-ended video-based commonsense question answering (V2C-QA) as a way to enrich our captions.Both the generation task and the QA task can be used to enrich video captions.. frame frame frame CNN Zhiyuan Fang, Tejas Gokhale, Pratyay Banerjee, Chitta Baral, Yezhou Yang |
EMNLP (1) | 1 |
| 2019 | Modularized Textual Grounding for Counterfactual ResilienceabstractComputer Vision applications often require a textual grounding module with precision, interpretability, and resilience to counterfactual inputs/queries. To achieve high grounding precision, current textual grounding methods heavily rely on large-scale training data with manual annotations at the pixel level. Such annotations are expensive to obtain and thus severely narrow the model's scope of real-world applications. Moreover, most of these methods sacrifice interpretability, generalizability, and they neglect the importance of being resilient to counterfactual inputs. To address these issues, we propose a visual grounding system which is 1) end-to-end trainable in a weakly supervised fashion with only image-level annotations, and 2) counterfactually resilient owing to the modular design. Specifically, we decompose textual descriptions into three levels: entity, semantic attribute, color information, and perform compositional grounding progressively. We validate our model through a series of experiments and demonstrate its improvement over the state-of-the-art methods. In particular, our model's performance not only surpasses other weakly/un-supervised methods and even approaches the strongly supervised ones, but also is interpretable for decision making and performs much better in face of counterfactual classes than all the others. Zhiyuan Fang, Shu Kong, Charless C. Fowlkes, Yezhou Yang |
CVPR | 1 |
| 2017 | Range Loss for Deep Face Recognition with Long-Tailed Training DataabstractDeep convolutional neural networks have achieved significant improvements on face recognition task due to their ability to learn highly discriminative features from tremendous amounts of face images. Many large scale face datasets exhibit long-tail distribution where a small number of entities (persons) have large number of face images while a large number of persons only have very few face samples (long tail). Most of the existing works alleviate this problem by simply cutting the tailed data and only keep identities with enough number of examples. Unlike these work, this paper investigated how long-tailed data impact the training of face CNNs and develop a novel loss function, called range loss, to effectively utilize the tailed data in training process. More specifically, range loss is designed to reduce overall intrapersonal variations while enlarge interpersonal differences simultaneously. Extensive experiments on two face recognition benchmarks, Labeled Faces in the Wild (LFW) [11] and YouTube Faces (YTF) [33], demonstrate the effectiveness of the proposed range loss in overcoming the long tail effect, and show the good generalization ability of the proposed methods. Xiao Zhang 0024, Zhiyuan Fang, Yandong Wen, Zhifeng Li 0001, Yu Qiao 0001 |
ICCV | 2 |