EDBT 2026 Demo / reviewers in the wild / expert
Shiyu Zhao 0001
dblp:120/7474-1
· DBLP profile ↗
15ranked-venue papers
9as first author
12since 2021 · last 2025
0000-0002-4978-725XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 9 first-author · 11 since 2021Artificial intelligence and machine learning · 11 · 6 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MLLM-as-a-Judge for Image Safety without Human LabelingabstractImage content safety has become a significant challenge with the rise of visual media on online platforms. Meanwhile, in the age of AI-generated content (AIGC), many image generation models are capable of producing harmful content, such as images containing sexual or violent material. Thus, it becomes crucial to identify such unsafe images based on established safety rules. Pre-trained Multimodal Large Language Models (MLLMs) offer potential in this regard, given their strong pattern recognition abilities. Existing approaches typically fine-tune MLLMs with humanlabeled datasets, which however brings a series of drawbacks. First, relying on human annotators to label data following intricate and detailed guidelines is both expensive and labor-intensive. Furthermore, users of safety judgment systems may need to frequently update safety rules, making fine-tuning on human-based annotation more challenging. This raises the research question: Can we detect unsafe images by querying MLLMs in a zero-shot setting using a predefined safety constitution (a set of safety rules)? Our research showed that simply querying pre-trained MLLMs does not yield satisfactory results. This lack of effectiveness stems from factors such as the subjectivity of safety rules, the complexity of lengthy constitutions, and the inherent biases in the models. To address these challenges, we propose a MLLM-based method includes objectifying safety rules, assessing the relevance between rules and images, making quick judgments based on debiased token probabilities with logically complete yet simplified precondition chains for safety rules, and conducting more in-depth reasoning with cascaded chain-of-thought processes if necessary. Experiment results demonstrate that our method is highly effective for zero-shot image safety judgment tasks. Zhenting Wang, Shuming Hu, Shiyu Zhao 0001, Xiaowen Lin, Felix Juefei-Xu, Zhuowei Li 0002, Ligong Han, Harihar Subramanyam, Jianfa Chen, Lingjuan Lyu, Shiqing Ma, Dimitris N. Metaxas |
CVPR | 3 |
| 2025 | Accelerating Multimodal Large Language Models by Searching Optimal Vision Token ReductionabstractPrevailing Multimodal Large Language Models (MLLMs) encode the input image(s) as vision tokens and feed them into the language backbone, similar to how Large Language Models (LLMs) process the text tokens. However, the number of vision tokens increases quadratically as the image resolutions, leading to huge computational costs. In this paper, we consider improving MLLM’s efficiency from two scenarios, (I) Reducing computational cost without degrading the performance. (II) Improving the performance with given budgets. We start with our main finding that the ranking of each vision token sorted by attention scores is similar in each layer except the first layer. Based on it, we assume that the number of essential top vision tokens does not increase along layers. Accordingly, for Scenario I, we propose a greedy search algorithm (G-Search) to find the least number of vision tokens to keep at each layer from the shallow to the deep. Interestingly, G-Search is able to reach the optimal reduction strategy based on our assumption. For Scenario II, based on the reduction strategy from G-Search, we design a parametric sigmoid function (P-Sigmoid) to guide the reduction at each layer of the MLLM, whose parameters are optimized by Bayesian Optimization. Extensive experiments demonstrate that our approach can significantly accelerate those popular MLLMs, e.g. LLaVA, and InternVL2 models, by more than 2⇥ without performance drops. Our approach also far outperforms other token reduction methods when budgets are limited, achieving a better trade-off between efficiency and effectiveness. Shiyu Zhao 0001, Zhenting Wang, Felix Juefei-Xu, Xide Xia, Miao Liu 0007, Mingfu Liang, Dimitris N. Metaxas, Licheng Yu |
CVPR | 1 |
| 2025 | LoR-VP: Low-Rank Visual Prompting for Efficient Vision Model AdaptationabstractVisual prompting has gained popularity as a method for adapting pre-trained models to specific tasks, particularly in the realm of parameter-efficient tuning. However, existing visual prompting techniques often pad the prompt parameters around the image, limiting the interaction between the visual prompts and the original image to a small set of patches while neglecting the inductive bias present in shared information across different patches. In this study, we conduct a thorough preliminary investigation to identify and address these limitations. We propose a novel visual prompt design, introducing **Lo**w-**R**ank matrix multiplication for **V**isual **P**rompting (LoR-VP), which enables shared and patch-specific information across rows and columns of image pixels. Extensive experiments across seven network architectures and four datasets demonstrate significant improvements in both performance and efficiency compared to state-of-the-art visual prompting methods, achieving up to $6\times$ faster training times, utilizing $18\times$ fewer visual prompt parameters, and delivering a 3.1% improvement in performance. Can Jin, Shiyu Zhao 0001, Zhenting Wang, Xiaoxiao He, Ligong Han, Tong Che, Dimitris N. Metaxas |
ICLR | 4 |
| 2024 | AIDE: An Automatic Data Engine for Object Detection in Autonomous DrivingabstractAutonomous vehicle (AV) systems rely on robust perception models as a cornerstone of safety assurance. However, objects encountered on the road exhibit a long-tailed distri-bution, with rare or unseen categories posing challenges to a deployed perception model. This necessitates an expen-sive process of continuously curating and annotating data with significant human effort. We propose to leverage recent advances in vision-language and large language models to design an Automatic Data Engine (AIDE) that automati-cally identifies issues, efficiently curates data, improves the model through auto-labeling, and verifies the model through generation of diverse scenarios. This process operates it-eratively, allowing for continuous self-improvement of the model. We further establish a benchmark for open-world detection on AV datasets to comprehensively evaluate vari-ous learning paradigms, demonstrating our method's supe-rior performance at a reduced cost. Mingfu Liang, Jong-Chyi Su, Samuel Schulter, Sparsh Garg, Shiyu Zhao 0001, Ying Wu 0001, Manmohan Krishna Chandraker |
CVPR | 5 |
| 2024 | Generating Enhanced Negatives for Training Language-Based Object DetectorsabstractThe recent progress in language-based open-vocabulary object detection can be largely attributed to finding better ways of leveraging large-scale data with free-form text annotations. Training such models with a discriminative objective function has proven successful, but requires good positive and negative samples. However, the free-form nature and the open vocabulary of object descriptions make the space of negatives extremely large. Prior works randomly sample negatives or use rule-based techniques to build them. In contrast, we propose to leverage the vast knowledge built into modern generative models to automatically build negatives that are more relevant to the original data. Specifically, we use large-Language-models to generate negative text descriptions, and text-to-image diffusion models to also generate corresponding negative images. Our experimental analysis confirms the relevance of the generated negative data, and its use in language-based detectors improves performance on two complex benchmarks. Code is available at https://github.com/xiaofeng94/Gen-Enhanced-Negs. Shiyu Zhao 0001, Long Zhao 0003, Yumin Suh, Dimitris N. Metaxas, Manmohan Krishna Chandraker, Samuel Schulter |
CVPR | 1 |
| 2024 | Taming Self-Training for Open-Vocabulary Object DetectionabstractRecent studies have shown promising performance in open-vocabulary object detection (OVD) by utilizing pseudo labels (PLs) from pre-trained vision and language models (VLMs). However, teacher-student self-training, a powerful and widely used paradigm to leverage PLs, is rarely explored for OVD. This work identifies two challenges of using self-training in OVD: noisy PLs from VLMs and frequent distribution changes of PLs. To address these challenges, we propose SAS-Det that tames self-training for OVD from two key perspectives. First, we present a split-and-fusion (SAF) head that splits a standard detection into an open-branch and a closed-branch. This design can reduce noisy supervision from pseudo boxes. More-over, the two branches learn complementary knowledge from different training data, significantly enhancing performance when fused together. Second, in our view, un-like in closed-set tasks, the PL distributions in OVD are solely determined by the teacher model. We introduce a periodic update strategy to decrease the number of up-dates to the teacher, thereby decreasing the frequency of changes in PL distributions, which stabilizes the training process. Extensive experiments demonstrate SAS-Det is both efficient and effective. SAS-Det outperforms recent models of the same scale by a clear margin and achieves 37.4 AP50and 29.1 APron novel categories of the COCO and LVIS benchmarks, respectively. Code is available at https://github.com/xiaofeng94/SAS-Det. Shiyu Zhao 0001, Samuel Schulter, Long Zhao 0003, Yumin Suh, Manmohan Krishna Chandraker, Dimitris N. Metaxas |
CVPR | 1 |
| 2023 | OmniLabel: A Challenging Benchmark for Language-Based Object DetectionabstractLanguage-based object detection is a promising direction towards building a natural interface to describe objects in images that goes far beyond plain category names. While recent methods show great progress in that direction, proper evaluation is lacking. With OmniLabel, we propose a novel task definition, dataset, and evaluation metric. The task subsumes standard- and open-vocabulary detection as well as referring expressions. With more than 28K unique object descriptions on over 25K images, OmniLabel provides a challenging benchmark with diverse and complex object descriptions in a naturally open-vocabulary setting. Moreover, a key differentiation to existing benchmarks is that our object descriptions can refer to one, multiple or even no object, hence, providing negative examples in free-form text. The proposed evaluation handles the large label space and judges performance via a modified average precision metric, which we validate by evaluating strong language-based baselines. OmniLabel indeed provides a challenging test bed for future research on language-based detection. Visit the project website at https://www.omnilabel.org Samuel Schulter, Yumin Suh, Konstantinos M. Dafnis, Shiyu Zhao 0001, Dimitris N. Metaxas |
ICCV | 6 |
| 2022 | Global Matching with Overlapping Attention for Optical Flow EstimationabstractOptical flow estimation is a fundamental task in computer vision. Recent direct-regression methods using deep neural networks achieve remarkable performance improvement. However, they do not explicitly capture long-term motion correspondences and thus cannot handle large motions effectively. In this paper, inspired by the traditional matching-optimization methods where matching is introduced to handle large displacements before energy-based optimizations, we introduce a simple but effective global matching step before the direct regression and develop a learning-based matching-optimization framework, namely GMFlowNet. In GMFlowNet, global matching is efficiently calculated by applying argmax on 4D cost volumes. Additionally, to improve the matching quality, we propose patch-based overlapping attention to extract large context features. Extensive experiments demonstrate that GM-FlowNet outperforms RAFT, the most popular optimization-only method, by a large margin and achieves state-of-the-art performance on standard benchmarks. Thanks to the matching and overlapping attention, GMFlowNet obtains major improvements on the predictions for textureless regions and large motions. Our code is made publicly available at https://github.com/xiaofeng94/GMFlowNet. Shiyu Zhao 0001, Long Zhao 0003, Enyu Zhou, Dimitris N. Metaxas |
CVPR | 1 |
| 2022 | Exploiting Unlabeled Data with Vision and Language Models for Object Detection
Shiyu Zhao 0001, Samuel Schulter, Long Zhao 0003, Anastasis Stathopoulos, Manmohan Krishna Chandraker, Dimitris N. Metaxas |
ECCV (9) | 1 |
| 2021 | Deep Animation Video Interpolation in the WildabstractIn the animation industry, cartoon videos are usually produced at low frame rate since hand drawing of such frames is costly and time-consuming. Therefore, it is desirable to develop computational models that can automatically interpolate the in-between animation frames. However, existing video interpolation methods fail to produce satisfying results on animation data. Compared to natural videos, animation videos possess two unique characteristics that make frame interpolation difficult: 1) cartoons comprise lines and smooth color pieces. The smooth areas lack textures and make it difficult to estimate accurate motions on animation videos. 2) cartoons express stories via exaggeration. Some of the motions are non-linear and extremely large. In this work, we formally define and study the animation video interpolation problem for the first time. To address the aforementioned challenges, we propose an effective framework, AnimeInterp, with two dedicated modules in a coarse-to-fine manner. Specifically, 1) Segment-Guided Matching resolves the "lack of textures" challenge by exploiting global matching among color pieces that are piece-wise coherent. 2) Recurrent Flow Refinement resolves the "non-linear and extremely large motion" challenge by recur-rent predictions using a transformer-like architecture. To facilitate comprehensive training and evaluations, we build a large-scale animation triplet dataset, ATD-12K, which comprises 12,000 triplets with rich annotations. Extensive experiments demonstrate that our approach outperforms existing state-of-the-art interpolation methods for animation videos. Notably, AnimeInterp shows favorable perceptual quality and robustness for animation scenarios in the wild. The proposed dataset and code are available at https://github.com/lisiyao21/AnimeInterp/. Li Siyao, Shiyu Zhao 0001, Weijiang Yu, Wenxiu Sun, Dimitris N. Metaxas, Chen Change Loy, Ziwei Liu 0002 |
CVPR | 2 |
| 2021 | Simulation of Atmospheric Visibility ImpairmentabstractChanges in aerosol composition and its proportions can cause changes in atmospheric visibility. Vision systems deployed outdoors must take into account the negative effects brought by visibility impairment. In order to develop vision algorithms that can adapt to low atmospheric visibility conditions, a large-scale dataset containing pairs of clear images and their visibility-impaired versions (along with other annotations if necessary) is usually indispensable. However, it is almost impossible to collect large amounts of such image pairs in a real physical environment. A natural and reasonable solution is to use virtual simulation technologies, which is also the focus of this paper. In this paper, we first deeply analyze the limitations and irrationalities of the existing work specializing on simulation of atmospheric visibility impairment. We point out that many simulation schemes actually even violate the assumptions of the Koschmieder's law. Second, more importantly, based on a thorough investigation of the relevant studies in the field of atmospheric science, we present simulation strategies for five most commonly encountered visibility impairment phenomena, including mist, fog, natural haze, smog, and Asian dust. Our work establishes a direct link between the fields of atmospheric science and computer vision. In addition, as a byproduct, with the proposed simulation schemes, a large-scale synthetic dataset is established, comprising 40,000 clear source images and their 800,000 visibility-impaired versions. To make our work reproducible, source codes and the dataset have been released at https://cslinzhang.github.io/AVID/. Lin Zhang 0014, Shiyu Zhao 0001, Yicong Zhou |
IEEE Trans. Image Process. | 3 |
| 2021 | RefineDNet: A Weakly Supervised Refinement Framework for Single Image DehazingabstractHaze-free images are the prerequisites of many vision systems and algorithms, and thus single image dehazing is of paramount importance in computer vision. In this field, prior-based methods have achieved initial success. However, they often introduce annoying artifacts to outputs because their priors can hardly fit all situations. By contrast, learning-based methods can generate more natural results. Nonetheless, due to the lack of paired foggy and clear outdoor images of the same scenes as training samples, their haze removal abilities are limited. In this work, we attempt to merge the merits of prior-based and learning-based approaches by dividing the dehazing task into two sub-tasks, i.e., visibility restoration and realness improvement. Specifically, we propose a two-stage weakly supervised dehazing framework, RefineDNet. In the first stage, RefineDNet adopts the dark channel prior to restore visibility. Then, in the second stage, it refines preliminary dehazing results of the first stage to improve realness via adversarial learning with unpaired foggy and clear images. To get more qualified results, we also propose an effective perceptual fusion strategy to blend different dehazing outputs. Extensive experiments corroborate that RefineDNet with the perceptual fusion has an outstanding haze removal capability and can also produce visually pleasing results. Even implemented with basic backbone networks, RefineDNet can outperform supervised dehazing approaches as well as other state-of-the-art methods on indoor and outdoor datasets. To make our results reproducible, relevant code and data are available at https://github.com/xiaofeng94/RefineDNet-for-dehazing. Shiyu Zhao 0001, Lin Zhang 0014, Ying Shen 0005, Yicong Zhou |
IEEE Trans. Image Process. | 1 |
| 2020 | Dehazing Evaluation: Real-World Benchmark Datasets, Criteria, and BaselinesabstractOn benchmark images, modern dehazing methods are able to achieve very comparable results whose differences are too subtle for people to qualitatively judge. Thus, it is imperative to adopt quantitative evaluation on a vast number of hazy images. However, existing quantitative evaluation schemes are not convincing due to a lack of appropriate datasets and poor correlations between metrics and human perceptions. In this work, we attempt to address these issues, and we make two contributions. First, we establish two benchmark datasets, i.e., the BEnchmark Dataset for Dehazing Evaluation (BeDDE) and the EXtension of the BeDDE (exBeDDE), which had been lacking for a long period of time. The BeDDE is used to evaluate dehazing methods via full reference image quality assessment (FR-IQA) metrics. It provides hazy images, clear references, haze level labels, and manually labeled masks that indicate the regions of interest (ROIs) in image pairs. The exBeDDE is used to assess the performance of dehazing evaluation metrics. It provides extra dehazed images and subjective scores from people. To the best of our knowledge, the BeDDE is the first dehazing dataset whose image pairs were collected in natural outdoor scenes without any simulation. Second, we provide a new insight that dehazing involves two separate aspects, i.e., visibility restoration and realness restoration, which should be evaluated independently; thus, to characterize them, we establish two criteria, i.e., the visibility index (VI) and the realness index (RI), respectively. The effectiveness of the criteria is verified through extensive experiments. Furthermore, 14 representative dehazing methods are evaluated as baselines using our criteria on BeDDE. Our datasets and relevant code are available at https://github.com/xiaofeng94/BeDDE-for-defogging. Shiyu Zhao 0001, Lin Zhang 0014, Shuaiyi Huang, Ying Shen 0005, Shengjie Zhao 0001 |
IEEE Trans. Image Process. | 1 |
| 2019 | Evaluation of Defogging: A Real-World Benchmark Dataset, A New Criterion and BaselinesabstractModern defogging methods are able to achieve very comparable results whose differences are too subtle for people to qualitatively judge. On the other hand, existing quantitative evaluation methods are also not convincing due to a lack of proper datasets. In this work, we attempt to address these issues and establish a long-term lacking benchmark dataset, namely BeDDE (BEnchmark Dataset for Defogging Evaluation), for evaluating the performance of defogging algorithms. To our knowledge, BeDDE is the first real-world dataset comprising foggy images with their registered clear counterparts. Using BeDDE, we set up a new criterion for evaluating defogging methods where VSI, a full reference image quality assessment metric, is calculated and averaged on registered ROIs of all image pairs. The evaluation results of the proposed criterion correlate well with human judgements. 10 state-of-the-art defogging methods are evaluated as baselines on BeDDE. BeDDE is available online. Shiyu Zhao 0001, Lin Zhang 0014, Shuaiyi Huang, Ying Shen 0005, Shengjie Zhao 0001, Yukai Yang |
ICME | 1 |
| 2018 | A CNN-Based Depth Estimation Approach with Multi-scale Sub-pixel Convolutions and a Smoothness Constraint
Shiyu Zhao 0001, Lin Zhang 0014, Ying Shen 0005, Yongning Zhu |
ACCV (2) | 1 |