EDBT 2026 Demo / reviewers in the wild / expert
Zhao Zhou
dblp:31/8349
· DBLP profile ↗
17ranked-venue papers
5as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 8 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bionic odor encoding and olfactory bulb model for processing electronic nose data
Siyuan Deng, Yangfan Huang, Fengchun Tian, Saverio De Vito, Hantao Li, Zhao Zhou |
Neurocomputing | 7 |
| 2025 | An Exemplar-based Framework for Chinese Text RecognitionabstractThis paper introduces a novel exemplar-based framework for reading Chinese texts in natural scene or document images. We present the Deep Exemplar-based Chinese Text Recognizer, which is structured to first identify candidate characters as exemplars from each text-line, and subsequently recognize them by retrieving analogous exemplars from a database. With text-line level annotations, we design the exemplar discovery network to simultaneously recognize texts and capture individual character positions in a weak-supervision manner. The exemplar retrieval module is then crafted to identify the most similar exemplar and propagate the corresponding character label. This enables us to effectively rectify the misrecognized characters and boost the performance of scene text recognition. Experiments on four scenarios of Chinese texts demonstrate the effectiveness of our proposed framework. Zhao Zhou, Xiangcheng Du, Yingbin Zheng, Xingjiao Wu, Cheng Jin 0001 |
AAAI | 1 |
| 2025 | Achieving Ensemble-Like Performance in a Single Model: A Feature Diversification Framework for Image-Text MatchingabstractModel ensembling is a widely used technique that enhances performance in image-text matching tasks by combining multiple models, each trained with different initializations. However, the inefficiencies associated with training several models and generating outputs from them constrain their practical applicability. In this paper, we argue that while the parameters of two randomly initialized models can differ significantly, their feature distributions can be similar at certain stages. By employing a proposed technique called cross-modal realignment, we demonstrate that features derived from differently initialized models maintain similarity at the feature extraction stage and can be effectively transformed by fine-tuning a small number of parameters. These findings provide an efficient way to achieve ensemble-like performance within a single model. Specifically, we propose a Feature Diversification Framework (FDF) that emulates the outputs of multiple model initializations to generate diverse features from a common shared feature. Firstly, we introduce feature conversion methods to transform shared features into a set of distinct features. Next, a realignment training strategy is presented to optimize negative pairs for realigning these transformed features, thereby enhancing their diversification to resemble the outputs of different models. Additionally, we propose a reweighting module that assigns weights to these features, enabling a weighted fusion approach for robust feature representation. Extensive experiments on the Flickr30K and MS-COCO datasets demonstrate the effectiveness and generalizability of our framework. Zhao Zhou, Yingbin Zheng, Xiangcheng Du, Cheng Jin 0001 |
AAAI | 1 |
| 2025 | Expanding the Scope of Negatives: Boosting Image-Text Matching with Negatives Distribution Guided LearningabstractImage-text matching is a crucial task that bridges visual and linguistic modalities. Recent research typically formulates it into the problem of maximizing the margin with the truly hardest negatives to enhance the learning efficiency and avoid the poor local optima. We argue that such formulation can lead to a serious limitation, i.e., under this formulation, conventional trainers would confine their horizon within the hardest negative examples, while other negative examples offer a range of semantic differences not present in the hardest negatives. In this paper, we propose an efficient negative distribution guided training framework for image-text matching to unlock the substantial promotion space left by the above limitation. Rather than simply incorporating additional negative examples into the training objective, which could diminish both the leading role of the hardest negatives in training and the effect of a large margin learning in producing a robust matching model, our central idea is to supply the objective with distributional information on the entire set of negative examples. To be precise, we first construct the sample similarity matrix based on several pretrained models to extract the distributional information of the entire negative sample dataset. Then we encode it into a margin regularization module to smooth the similarities differences of all negatives. This enhancement facilitates the capture of fine-grained semantic differences and guides the main learning process by maximizing the margin with hard negative examples. Furthermore, we propose a hardest negative rectification module to address the instability in hardest negative selection based on predicted similarity and to correct erroneous hardest negatives. We evaluate our method in combination with several state-of-the-art image-text matching methods, and our quantitative and qualitative experiments demonstrate its significant generalizability and effectiveness. Zhao Zhou, Xiangcheng Du, Yingbin Zheng, Cheng Jin 0001 |
AAAI | 1 |
| 2025 | Unleashing the Semantic Adaptability of Controlled Diffusion Model for Image ColorizationabstractRecent data-driven image colorization methods have leveraged pre-trained Text-to-Image (T2I) diffusion models as generative prior, while still suffering from unsatisfactory and inaccurate semantic-level color control. To address these issues, we propose a Semantic Adaptation method (SeAda) that enhances the prior while considering the semantic discrepancy between color and grayscale image pairs. The SeAda employs a semantic adapter to produce refined semantic embeddings and a controlled T2I diffusion model to create reasonably colored images. Specifically, the semantic adapter transfers the embedding from grayscale to color domain, while the diffusion model utilizes the refined embedding and prior knowledge to achieve realistic and diverse results. We also design a three-staged training strategy to improve semantic comprehension and prior integration for further performance improvement. Extensive experiments on public datasets demonstrate that our method outperforms existing state-of-the-art techniques, yielding superior performance in image colorization. Xiangcheng Du, Zhao Zhou, Yingbin Zheng, Xingjiao Wu, Peizhu Gong, Cheng Jin 0001 |
IJCAI | 2 |
| 2024 | Efficient Scene Text Image Super-Resolution with Semantic GuidanceabstractScene text image super-resolution has significantly improved the accuracy of scene text recognition. However, many existing methods emphasize performance over efficiency and ignore the practical need for lightweight solutions in deployment scenarios. Faced with the issues, our work proposes an efficient framework called SGENet to facilitate deployment on resource-limited platforms. SGENet contains two branches: super-resolution branch and semantic guidance branch. We apply a lightweight pre-trained recognizer as a semantic extractor to enhance the understanding of text information. Meanwhile, we design the visual-semantic alignment module to achieve bidirectional alignment between image features and semantics, resulting in the generation of high-quality prior guidance. We conduct extensive experiments on benchmark dataset, and the proposed SGENet achieves excellent performance with fewer computational costs. LeoWu TomyEnrique, Xiangcheng Du, Kangliang Liu, Zhao Zhou, Cheng Jin 0001 |
ICASSP | 5 |
| 2024 | Fine-Grained Scene Image Classification with Modality-Agnostic AdapterabstractWhen dealing with the task of fine-grained scene image classification, most previous works lay much emphasis on global visual features when doing multi-modal feature fusion. In other words, models are deliberately designed based on prior intuitions about the importance of different modalities. In this paper, we present a new multi-modal feature fusion approach named MAA (Modality-Agnostic Adapter), trying to make the model learn the importance of different modalities in different cases adaptively, without giving a prior setting in the model architecture. More specifically, we eliminate the modal differences in distribution and then use a modality-agnostic Transformer encoder for a semantic-level feature fusion. Our experiments demonstrate that MAA achieves state-of-the-art results on benchmarks by applying the same modalities with previous methods. Besides, it is worth mentioning that new modalities can be easily added when using MAA and further boost the performance. Zhao Zhou, Xiangcheng Du, Xingjiao Wu, Yingbin Zheng, Cheng Jin 0001 |
ICME | 2 |
| 2024 | MultiColor: Image Colorization by Learning from Multiple Color Spaces
Xiangcheng Du, Zhao Zhou, Xingjiao Wu, Yingbin Zheng, Cheng Jin 0001 |
ACM Multimedia | 2 |
| 2023 | Modeling Stroke Mask for End-to-End Text ErasingabstractScene text erasing aims to wipe text regions in scene images with reasonable background. Most previous approaches employ scene text detectors to assist localization of the text regions. However, detected text boxes contain both text strokes and background clutters, and directly in-painting on the whole boxes may remain text artifacts and make regions unnatural. In this paper, we present an end-to-end network that focuses on modeling text stroke masks that provide more accurate locations to compute erased images. The network consists of two stages, i.e., a basic network with stroke generation and a refinement network with stroke awareness. The basic network predicts the text stroke masks and initial erasing results simultaneously. The refinement network receives the masks as supervision to generate natural erased results. Experiments on both synthetic and real-world scene images demonstrate the effectiveness of our framework in producing high quality erasing results. Xiangcheng Du, Zhao Zhou, Yingbin Zheng, Tianlong Ma, Xingjiao Wu, Cheng Jin 0001 |
WACV | 2 |
| 2023 | Progressive scene text erasing with self-supervision
Xiangcheng Du, Zhao Zhou, Yingbin Zheng, Xingjiao Wu, Tianlong Ma, Cheng Jin 0001 |
Comput. Vis. Image Underst. | 2 |
| 2023 | Reading Scene Text with Aggregated Temporal Convolutional EncoderabstractReading scene text in the natural image is of fundamental importance in many real-world problems. Text recognition has a profound effect on information processing by enabling automated extraction and interpretation. Recent scene text recognition methods employ the encoder-decoder framework, which constructs the encoder by obtaining the visual representations based on the last layer of the backbone network and then feeding them into a sequence model. In this article, we propose a novel encoder structure that performs the feature extractor and the sequence modeling within a unified framework. The introduced Aggregated Temporal Convolutional Encoder (ATCE) first incorporates the temporal convolutional layers to consider the long-term temporal relationship in the encoder stage. The aggregation of these temporal convolution modules is designed to utilize visual features from different levels, by augmenting the standard architecture with deeper aggregation to better fuse information across modules. We also study the impact of different attention modules in convolutional blocks for learning accurate text representations. We conduct comparisons on several scene text recognition benchmarks for both Chinese and English; the experiments demonstrate the complementary ability with different decoder variants and the effectiveness of our proposed approach. Tianlong Ma, Xiangcheng Du, Xingjiao Wu, Zhao Zhou, Yingbin Zheng, Cheng Jin 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2021 | Two-Phase Jointly Optimal Strategies and Winning Regions of the Capture-the-Flag GameabstractThis paper considers a one versus one two-phase capture-the-flag differential game with speed heterogeneity. The attacker tries to capture the flag in Phase-I (Flag-Capture) and return to the safe region in Phase-II (Flag-Return), while the defender aims to intercept the attacker. Firstly, we analyse the single-phase optimal strategies by employing the Apollonius circle. Moreover, we present two-phase jointly optimal strategies by modeling the game as constrained nonlinear optimization problems , which are solved by sequential quadratic programming (SQP) algorithm. Besides, we construct the winning regions by searching the border radius of the region when positions of defender and flag are given. It is worth noting that all results is analytical except for the SQP and search Algorithms. Finally, we present simulations to verify the proposed methods. Zhao Zhou, Jiapeng Xu, Yang Tang 0001 |
IECON | 1 |
| 2021 | Opinion separation in leader-follower coopetitive social networks
Haili Liang, Fanli Yuan, Zhao Zhou, Housheng Su |
Neurocomputing | 3 |
| 2020 | Detecting Curve Text with Local Segmentation Network and Curve ConnectionabstractCurve text or arbitrary shape text is very common in real-world scenarios. In this paper, we propose a novel framework with the local segmentation network (LSN) followed by the curve connection to detect text in horizontal, oriented and curved forms. The LSN is composed of two elements, i.e., proposal generation to get the horizontal rectangle proposals with high overlap with text and text segmentation to find the arbitrary shape text region within proposals. The curve connection is then designed to connect the local mask to the detection results. We conduct experiments using the proposed framework on two real-world curve text detection datasets and demonstrate the effectiveness over previous approaches. Zhao Zhou, Hao Ye 0005, Luhui Chen, Yingbin Zheng |
IEEE BigData | 1 |
| 2020 | Margin Guidance Network for Arbitrary-shaped Scene Text DetectionabstractSegmentation-based scene text detection approaches have been adopted to arbitrary-shaped texts and have achieved a great progress. However, false detection always easily exist when the arbitrary-shaped texts are close to each other. In this paper, we propose the Margin Guidance Network (MGN) that mainly based on the margin constraint residual module (MCRM) to address aforementioned problem. The MCRM considers the margins between multiple text instance masks to guide the training of network and improve the performance on text detection. The MCRM contains two prediction branch, the one can generate the multiple different scale of masks for a text instance and the other branch is used to generate multiple margins between the above masks. Experimental results on three public benchmarks including ICDAR2015, CTW1500 and Total-Text have demonstrated that the proposed MGN achieves the state-of-the-art results. Xin Li 0110, Xingjiao Wu, Tianlong Ma, Zhao Zhou, Luhui Chen, Liang He 0001 |
ICTAI | 4 |
| 2019 | Math Expression Image Retrieval via Attention-Based FrameworkabstractMath expression image retrieval concerns not only visual features but also high-level semantic understanding. Considering math expression image retrieval as traditional content-based image retrieval may suffer the layout misunderstanding, as math expressions with same symbols but different layouts may be interpreted as different meaning. In this paper, we propose a novel retrieval indexing framework for math expression retrieval, namely Scanner-Recognizer-Embedding (SRE) framework. The math expression images passed through SRE are projected into a low dimension semantic space. Retrieval based on embedded semantic vectors is fast and accurate. Experiments on a math expression database demonstrate that the SRE framework outperforms state-of-the-art image-based features. Caili Wu, Zhao Zhou, Hao Ye 0005, Jing Yang 0023, Liang He 0001 |
ICTAI | 2 |
| 2013 | Academic education of software engineering practices: towards planning and improving capstone courses based upon intensive coaching and team routinesabstractAcademic education of professional processes is challenged by a necessary balance of practical activities with academic reflection. In this paper we address this issue by discussing our experiences with teaching software engineering practices and their continuous improvement. By designing a graduate course we embed an intensive coaching routine based upon agile practices with research activities to leverage knowledge of students and coaches. As a concrete example of an embedded research project we conduct an experiment on the impact of two different meeting routines on the teams satisfaction with information exchange. Our results show that the intensive coaching in individual teams is shorter in nature and more appealing to the students. Our findings suggest that software engineering education can benefit from the notion of team routines and process improvement practices contributing to maturity of students and educators. Christoph J. Stettina, Zhao Zhou, Thomas Bäck, Bernhard R. Katzy |
CSEE&T | 2 |