EDBT 2026 Demo / reviewers in the wild / expert
Minyi Zhao
dblp:70/2526
· DBLP profile ↗
21ranked-venue papers
13as first author
14since 2021 · last 2026
0000-0001-7720-806XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 9 first-author · 8 since 2021Artificial intelligence and machine learning · 9 · 4 first-author · 9 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 2 since 2021Computer networks · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Heterogeneous Uncertainty-Guided Composed Image Retrieval with Fine-Grained Probabilistic LearningabstractComposed Image Retrieval (CIR) enables image search by combining a reference image with modification text. Intrinsic noise in CIR triplets incurs intrinsic uncertainty and threatens model's robustness. Probabilistic learning approaches have shown promise in addressing such issues; however, they fall short for CIR due to their instance-level holistic modeling and homogeneous treatments for queries and targets. This paper introduces a Heterogeneous Uncertainty-Guided (HUG) paradigm to overcome these limitations. HUG utilizes a fine-grained probabilistic learning framework, where queries and targets are represented by Gaussian embeddings capturing detailed concepts and uncertainties. We customize heterogeneous uncertainty estimations for multi-modal queries and uni-modal targets. Given a query, we capture uncertainties not only regarding uni-modal content quality but also multi-modal coordination, followed by a provable dynamic weighting mechanism to derive the comprehensive query uncertainty. We further design uncertainty-guided objectives, including query-target holistic contrast and fine-grained contrasts with comprehensive negative sampling strategies, which effectively enhance discriminative learning. Experiments on benchmarks demonstrate HUG's effectiveness beyond state-of-the-art baselines, with faithful analysis justifying the technical contributions. Haomiao Tang, Jinpeng Wang 0002, Minyi Zhao, Guanghao Meng, Ruisheng Luo, Long Chen 0016, Shutao Xia |
AAAI | 3 |
| 2026 | Anime-2026: A Large-scale Anime Character Dataset for Anime-related AI TasksabstractAnime, as a popular medium, has attracted hundreds of millions audience, especially the young followers. In recent years, various anime-related AI tasks like anime character classification (ACC), retrieval (ACR), tag-prediction (ACTP), question answering (ACQA), and generation (ACG) have been proposed to meet the requirements of various applications. However, there is still a lack of large-scale datasets for these tasks. Such a situation is definitely not beneficial to anime-related academic research and industrial applications. In this paper, to boost anime-related AI technical research and application development, we present Anime-2026, a new and large-scale anime character dataset, which can support various anime-related AI tasks, including ACC, ACR, ACTP, ACQA and ACG etc. Anime-2026 consists of 1.5M anime character images, 14k different characters, 16k unique semantic keyword tags, 10k question-answer pairs, and 4k manually designed text queries by crowdsourcing for the ACR and ACG tasks. Furthermore, to assess the dataset, we re-implement a number of generic and anime-specific AI baseline models, and conduct extensive experiments to evaluate these models on Anime-2026. In summary, as a general benchmark dataset, Anime-2026 provides the largest free anime character resource to support future anime-related AI research and development. We expect that Anime-2026 will promote the R&D of new and more advanced models and methods of various anime-related AI tasks. The dataset is available on https://huggingface.co/datasets/miaojiemiao/Anime-2026. Shijie Xuyang, Bingzhe Yu, Minyi Zhao, Guangze Li, Jihong Guan, Shuigeng Zhou |
ICMR | 3 |
| 2025 | Towards High Robust Vision-Language Large Models: Benchmark and MethodabstractRecently, numerous benchmarks have been constructed to evaluate various general capabilities (e.g., perception and reasoning) of Vision-Language Large Models (VLLMs). However, few studies have focused on the robustness of VLLMs when dealing with altered prompts and images. To fill this gap, this paper first constructs a real-world, high-quality, and challenging benchmark, namely RBench (i.e., Robust Bench). Specifically, RBench is human-annotated, with both prompts and images being modified to enrich the difficulty, and cross-validation to ensure data quality. Then, we propose a new method, called Robustness Booster (RBoost in short), to effectively enhance the robustness of existing VLLMs by automatically generating high-value instruction-tuning training data. Extensive experiments demonstrate the vulnerability of existing VLLMs when handling altered inputs, and the superiority of our RBoost method in improving model robustness. RBench is available at https://github.com/zhaominyiz/RBench. Minyi Zhao, Wensong He, Bingzhe Yu, Yuxi Mi, Shuigeng Zhou |
ACM Multimedia | 1 |
| 2025 | HiREN: Towards higher supervision quality for better scene text image super-resolution
Minyi Zhao, Yi Xu 0003, Bingjia Li, Jihong Guan, Shuigeng Zhou |
Neurocomputing | 1 |
| 2023 | Privacy-Preserving Face Recognition Using Random Frequency ComponentsabstractThe ubiquitous use of face recognition has sparked increasing privacy concerns, as unauthorized access to sensitive face images could compromise the information of individuals. This paper presents an in-depth study of the privacy protection of face images’ visual information and against recovery. Drawing on the perceptual disparity between humans and models, we propose to conceal visual information by pruning human-perceivable low-frequency components. For impeding recovery, we first elucidate the seeming paradox between reducing model-exploitable information and retaining high recognition accuracy. Based on recent theoretical insights and our observation on model attention, we propose a solution to the dilemma, by advocating for the training and inference of recognition models on randomly selected frequency components. We distill our findings into a novel privacy-preserving face recognition method, PartialFace. Extensive experiments demonstrate that PartialFace effectively balances privacy protection goals and recognition accuracy. Code is available at: https://github.com/Tencent/TFace. Yuxi Mi, Yuge Huang, Jiazhen Ji, Minyi Zhao, Jiaxiang Wu 0001, Xingkun Xu, Shouhong Ding, Shuigeng Zhou |
ICCV | 4 |
| 2023 | Anime Character Identification and Tag Prediction by Multimodality Modeling: Dataset and ModelabstractIn recent years, some advances have been achieved in classification and object detection related to animation. However, these works do not take full advantage of the tags and text description content attached to the anime data when they are created, which restricts both the related methods and data to unimodality, consequently leading to unsatisfactory performance. In this paper, we propose a novel multimodal deep learning network for Anime character identification and tag prediction by exploiting multimodal data. Considering that in many realistic scenarios, text annotations accompanying anime may be missing, we introduce the concept of curriculum learning in transformers to enable inference with only one modality. Another challenge lies in that the existing dataset does not meet our demand for large-scale multimodal deep learning. To train the proposed network, we construct a new anime dataset Dan: mul that contains over 1.6M images spread across more than 14K categories, with an average of 24 tags per image. To the best of our knowledge, this is the first dataset specifically designed for multimodal anime character identification. With the trained network, we can identify the anime characters in images and generate the related tags. Experiments show that our method achieves state-of-the-art performance on Dan: mul in animation identification. Jiaxiang Wu 0001, Minyi Zhao, Shuigeng Zhou |
IJCNN | 3 |
| 2023 | Boosting Aspect Sentiment Quad Prediction by Data Augmentation and Self-TrainingabstractAspect Sentiment Quad Prediction (ASQP), which aims to predict the (aspect category, aspect term, opinion term, sentiment polarity) quadruple in a sentence, is an important subtask of Aspect-based Sentiment Analysis (ABSA) and has also gained much research attention. Recently, a number of works adopt generative models to do ASQP and achieve promising performance. Nevertheless, these approaches are limited by data scarcity, due to the challenge of labeling fine-grained quadruple annotations. Even data augmentation techniques do not help much because it is difficult to maintain the quality of augmented samples. Furthermore, the proposed templates of generative models are undesirable in extracting semantic and structural information, thus cannot fully exploit the knowledge of the language models. In this paper, to boost the performance of ASQP, we propose a novel method called DAST (the abbreviation of Data Augmentation and Self-Training). Specifically, DAST first generates large amounts of unlabeled data from the original training data. Then, an iterative self-training mechanism is proposed to fully utilize the unlabeled data to train the ASQP model. In each iteration, we employ a base ASQP model and a discriminator to obtain high-quality samples with pseudo labels, and further enhance the two models with the augmented data. In our method, both the potential of data augmentation and the ASQP model are fully explored thanks to the iterative mechanism. Finally, we advance the ASQP base model via a novel tree-like template that improves the representations of the relations between different sentiment elements, and a double-check mechanism that leverages multi-task learning to detect incorrect predictions. Extensive experiments on two benchmark datasets demonstrate the superiority of our method over the existing ones. Yongxin Yu, Minyi Zhao, Shuigeng Zhou |
IJCNN | 2 |
| 2023 | STIRER: A Unified Model for Low-Resolution Scene Text Image Recovery and RecognitionabstractThough scene text recognition (STR) from high-resolution (HR) images has achieved significant success in the past years, text recognition from low-resolution (LR) images is still a challenging task. This inspires the study on scene text image super-resolution (STISR) to generate super-resolution (SR) images based on the LR images, then STR is performed on the generated SR images, which eventually boosts the recognition performance. However, existing methods have two major drawbacks: 1) STISR models may generate imperfect SR images, which mislead the subsequent recognition. 2) As the STISR models are optimized for high recognition accuracy, the fidelity of SR images may be degraded. Consequently, neither the recognition performance of STR nor the fidelity of STISR is desirable. In this paper, a novel model called STIRER (the abbreviation of Scene Text Image REcovery and Recognition) is proposed to effectively and simultaneously recover and recognize LR scene text images under a unified framework. Concretely, STIRER consists of a feature encoder to obtain pixel features and two dedicated decoders to generate SR images and recognize texts respectively based on the encoded features and the raw LR images. We propose a progressive scene text swin transformer architecture as the encoder to enrich the representations of the pixel features for better recovery and recognition. Extensive experiments on two LR datasets show the superiority of our model to the existing methods on recognition performance, super-resolution fidelity and computational cost. The STIRER Code is available in https://github.com/zhaominyiz/STIRER. Minyi Zhao, Shijie Xuyang, Jihong Guan, Shuigeng Zhou |
ACM Multimedia | 1 |
| 2023 | Keyword-Based Diverse Image Retrieval by Semantics-aware Contrastive Learning and TransformerabstractIn addition to relevance, diversity is an important yet less studied performance metric of cross-modal image retrieval systems, which is critical to user experience. Existing solutions for diversity-aware image retrieval either explicitly post-process the raw retrieval results from standard retrieval systems or try to learn multi-vector representations of images to represent their diverse semantics. However, neither of them is good enough to balance relevance and diversity. On the one hand, standard retrieval systems are usually biased to common semantics and seldom exploit diversity-aware regularization in training, which makes it difficult to promote diversity by post-processing. On the other hand, multi-vector representation methods are not guaranteed to learn robust multiple projections. As a result, irrelevant images and images of rare or unique semantics may be projected inappropriately, which degrades the relevance and diversity of the results generated by some typical algorithms like top-k. To cope with these problems, this paper presents a new method called CoLT that tries to generate much more representative and robust representations for accurately classifying images. Specifically, CoLT first extracts semantics-aware image features by enhancing the preliminary representations of an existing one-to-one cross-modal system with semantics-aware contrastive learning. Then, a transformer-based token classifier is developed to subsume all the features into their corresponding categories. Finally, a post-processing algorithm is designed to retrieve images from each category to form the final retrieval result. Extensive experiments on two real-world datasets Div400 and Div150Cred show that CoLT can effectively boost diversity, and outperforms the existing methods as a whole (with a higher F1 score). Minyi Zhao, Jinpeng Wang 0002, Dongliang Liao, Huanzhong Duan, Shuigeng Zhou |
SIGIR | 1 |
| 2022 | Two-Stage Multimodality Fusion for High-Performance Text-Based Visual Question Answering
Bingjia Li, Minyi Zhao, Shuigeng Zhou |
ACCV (4) | 3 |
| 2022 | C3-STISR: Scene Text Image Super-resolution with Triple CluesabstractScene text image super-resolution (STISR) has been regarded as an important pre-processing task for text recognition from low-resolution scene text images. Most recent approaches use the recognizer's feedback as clues to guide super-resolution. However, directly using recognition clue has two problems: 1) Compatibility. It is in the form of probability distribution, has an obvious modal gap with STISR - a pixel-level task; 2) Inaccuracy. it usually contains wrong information, thus will mislead the main task and degrade super-resolution performance. In this paper, we present a novel method C3-STISR that jointly exploits the recognizer's feedback, visual and linguistical information as clues to guide super-resolution. Here, visual clue is from the images of texts predicted by the recognizer, which is informative and more compatible with the STISR task; while linguistical clue is generated by a pre-trained character-level language model, which is able to correct the predicted texts. We design effective extraction and fusion mechanisms for the triple cross-modal clues to generate a comprehensive and unified guidance for super-resolution. Extensive experiments on TextZoom show that C3-STISR outperforms the SOTA methods in fidelity and recognition performance. Code is available in https://github.com/zhaominyiz/C3-STISR. Minyi Zhao, Fan Bai 0001, Bingjia Li, Shuigeng Zhou |
IJCAI | 1 |
| 2022 | EPiDA: An Easy Plug-in Data Augmentation Framework for High Performance Text ClassificationabstractMinyi Zhao, Lu Zhang, Yi Xu, Jiandong Ding, Jihong Guan, Shuigeng Zhou. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Minyi Zhao, Lu Zhang 0060, Yi Xu 0003, Jiandong Ding, Jihong Guan, Shuigeng Zhou |
NAACL-HLT | 1 |
| 2022 | Towards Video Text Visual Question Answering: Benchmark and BaselineabstractThere are already some text-based visual question answering (TextVQA) benchmarks for developing machine's ability to answer questions based on texts in images in recent years. However, models developed on these benchmarks cannot work effectively in many real-life scenarios (e.g. traffic monitoring, shopping ads and e-learning videos) where temporal reasoning ability is required. To this end, we propose a new task named Video Text Visual Question Answering (ViteVQA in short) that aims at answering questions by reasoning texts and visual information spatiotemporally in a given video. In particular, on the one hand, we build the first ViteVQA benchmark dataset named M4-ViteVQA --- the abbreviation of Multi-category Multi-frame Multi-resolution Multi-modal benchmark for ViteVQA, which contains 7,620 video clips of 9 categories (i.e., shopping, traveling, driving, vlog, sport, advertisement, movie, game and talking) and 3 kinds of resolutions (i.e., 720p, 1080p and 1176x664), and 25,123 question-answer pairs. On the other hand, we develop a baseline method named T5-ViteVQA for the ViteVQA task. T5-ViteVQA consists of five transformers. It first extracts optical character recognition (OCR) tokens, question features, and video representations via two OCR transformers, one language transformer and one video-language transformer, respectively. Then, a multimodal fusion transformer and an answer generation module are applied to fuse multimodal information and generate the final prediction. Extensive experiments on M4-ViteVQA demonstrate the superiority of T5-ViteVQA to the existing approaches of TextVQA and VQA tasks. The ViteVQA benchmark is available in https://github.com/bytedance/VTVQA. Minyi Zhao, Bingjia Li, Wanqing Li 0007, Shijie Xuyang, Zhihang Yu, Xinkun Yu, Guangze Li, Aobotao Dai, Shuigeng Zhou |
NeurIPS | 1 |
| 2021 | Recursive Fusion and Deformable Spatiotemporal Attention for Video Compression Artifact ReductionabstractA number of deep learning based algorithms have been proposed to recover high-quality videos from low-quality compressed ones. Among them, some restore the missing details of each frame via exploring the spatiotemporal information of neighboring frames. However, these methods usually suffer from a narrow temporal scope, thus may miss some useful details from some frames outside the neighboring ones. In this paper, to boost artifact removal, on the one hand, we propose a Recursive Fusion (RF) module to model the temporal dependency within a long temporal range. Specifically, RF utilizes both the current reference frames and the preceding hidden state to conduct better spatiotemporal compensation. On the other hand, we design an efficient and effective Deformable Spatiotemporal Attention (DSTA) module such that the model can pay more effort on restoring the artifact-rich areas like the boundary area of a moving object. Extensive experiments show that our method outperforms the existing ones on the MFQE 2.0 dataset in terms of both fidelity and perceptual effect. Code is available at https://github.com/zhaominyiz/RFDA-PyTorch. Minyi Zhao, Yi Xu 0003, Shuigeng Zhou |
ACM Multimedia | 1 |
| 2003 | Optimal error protection for progressive image transmission over finite-state Markov channelsabstractThis paper presents a unified framework for addressing progressive image transmission over both memoryless and fading channels based on finite-state Markov channels (FSMC) models. The main advantage of using FSMC models is that they allow analytical derivation of optimal joint source-channel coding solutions in the form of unequal error protection without the burden of simulating wireless channels. Our analyses and experiments confirm that 1) FSMC models work for fading channels and 2) PSNR results obtained with FSMC models compare favorably with those published in the literature. Zhongmin Liu, Minyi Zhao, Zixiang Xiong |
ICC | 2 |
| 2002 | Performance evaluation and optimization of embedded image sources over noisy channelsabstractWe investigate and prove the relationships among several commonly used and new performance metrics for embedded image bit streams transmitted over noisy channels. The average of the first error-free run length (FEFRL) is proposed as a simpler performance metric. On binary symmetric channels (BSC), the average FEFRL is obtained in closed form, which greatly simplifies performance optimization. Simulation results justify the merit of the proposed technique. Minyi Zhao, William A. Pearlman, Ali N. Akansu |
IEEE Signal Process. Lett. | 1 |
| 2001 | Optimal Protection for Progressive Image Transmission over Noisy Channels: A General Approach
Minyi Zhao, Ali N. Akansu |
Data Compression Conference | 1 |
| 2000 | A New Method for Optimal Rate Allocation for Progressive Image Transmission over Noisy ChannelsabstractWe present a new method for optimal rate allocation between a progressive image coder and a channel coder for noisy channels. A mathematical model for the embedded bit streams is developed and used as a new metric for such a joint source-channel coding optimization problem. It is further shown that maximization of such a new metric is equivalent to the maximization of a conventional PSNR measure. PSNR improvements up to 0.3 dB over very noisy channels at low bit rate are achieved by using the proposed method without any additional overheads. Furthermore, the PSNR performance obtained by using the proposed method is upper bounded by those of the conventional Monte Carlo method. Computational results and numerical comparisons consistently verify the merit of the proposed technique. Minyi Zhao, A. Aydin Alatan, Ali N. Akansu |
Data Compression Conference | 1 |
| 2000 | Dynamic UEP of embedded image bit streams over noisy channelsabstractWe present a dynamic unequal error protection (UEP) framework for the embedded image bit streams over noisy channels. A general source model for M-rate UEP schemes is derived and an optimization algorithm for 2-rate dynamic UEP schemes is developed. To facilitate the design and implementation of UEP schemes, we also derived a necessary condition and an upper bound for UEP gains. Simulation results demonstrate about 0.3 dB improvements over equal error protection (EEP) schemes at the price of negligible overheads while the progressiveness of the original bit streams is still kept. Minyi Zhao, A. Aydin Alatan, Ali N. Akansu |
ICASSP | 1 |
| 2000 | Optimization of Dynamic UEP Schemes for Embedded Image Sources in Noisy ChannelsabstractWe present an optimization algorithm for 3-rate dynamic unequal error protection (UEP) schemes for embedded image bit streams transmitted over noisy channels. A theoretical upper bound for UEP gains is also given. Simulation results demonstrate about 0.3 dB improvements over the optimal equal error protection (EEP) schemes at the price of negligible overheads and off-line computation load. The progression of the source bit streams is still kept. Furthermore, it is discovered that for binary symmetric channels (BSCs) with bit error rate (BER) satisfying /spl epsiv//spl les/10/sup -1/, at most 3 protection levels are enough to squeeze out almost all the UEP gains. Minyi Zhao, Ali N. Akansu |
ICIP | 1 |
| 2000 | Unequal error protection of SPIHT encoded image bit streamsabstractA derivative of the set partitioning into hierarchical trees (SPIHT) image coding method, which generates substreams with different error-resilience properties, is proposed. By dividing the image bit stream into three classes, substreams with different immunity properties are obtained. The unequal protection of these substreams with different channel coding rates improves the overall performance of the method against channel errors. Simulation results show the superiority of the proposed method over some of the state-of-the-art methods. A. Aydin Alatan, Minyi Zhao, Ali N. Akansu |
IEEE J. Sel. Areas Commun. | 2 |