Maoying Qiao

dblp:62/9638 · DBLP profile ↗
← Back
22ranked-venue papers
7as first author
13since 2021 · last 2025
0000-0002-0990-5506ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 12 · 4 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-authorSystems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 QuARF: Quality-Adaptive Receptive Fields for Degraded Image Perception
abstract
Advanced Deep Neural Networks (DNNs) perform well for high-quality images, but their performance dramatically decreases for degraded images. Data augmentation is commonly used to alleviate this problem, but using too much perturbed data might seriously decrease the performance on pristine images. To tackle this challenge, we take our cue from the assumption of spatial coincidence in human visual perception, i.e. multiscale and varying receptive fields are required for understanding pristine and degraded images. Correspondingly, we propose a novel plug-and-play network architecture, dubbed Quality-Adaptive Receptive Fields (QuARF), to automatically select the optimal receptive fields based on the quality of the input image. To this end, we first design a multi-kernel convolutional block, which comprises multiscale continuous receptive fields. Afterward, we design a quality-adaptive routing network to predict the significance of each kernel, based on the quality features extracted from the input image. In this way, QuARF automatically selects the optimal inference route for each image. To further boost efficiency and effectiveness, the input feature map is split into multiple groups, with each group independently learning its quality-adaptive routing parameters. We apply QuARF to a variety of DNNs and conduct experiments in both discriminative and generation tasks, including semantic segmentation, image translation, and restoration. Thorough experimental results show that QuARF significantly and robustly improves the performance for degraded images, and outperforms data augmentation in most cases.
Fei Gao 0006, Ziyun Li 0002, Wenwang Han, Maoying Qiao, Jinlan Xu, Nannan Wang 0001
AAAI6
2025 MAJoR: Visual Emotion Analysis via Multi-Attribute Joint Reasoning
abstract
Visual Emotion Analysis (VEA) seeks to anticipate individuals’ emotional reactions to visual stimuli. The subjective perception of visual emotion is an integrated impact of the appearance, scene, and objects presented in an image. It is thus significance to analysis visual emotion by incorporating diverse visual attributes. Motivated by this, in this paper, we propose a novel VEA method based on Multi-Attribute Joint Reasoning (MAJoR). Specifically, we first use a multi-stream networks to learning multi-attribute representations, including the color, brightness, scene type, and object class of an input image. Afterward, we use a Graph Convolution Network (GCN) to model the inherent relationships among such visual attributes, and to predict the emotion category. Finally, we propose a two-stage knowledge distillation strategy, to boost the performance of light-weight VEA models via MAJoR. Extensive experiments conducted on several VEA databases showcase the superiority of the proposed MAJoR model and the distilled lightweight versions, compared to state-of-the-art approaches. Our code and models are available at: https://github.com/AiArt-Gao/MAJoR.
Yuxin Fei, Jinlan Xu, Maoying Qiao, Fei Gao 0006
ICASSP3
2025 Q-Norm: Robust Representation Learning via Quality-Adaptive Normalization
Lanning Zhang, Fei Gao 0006, Ziyun Li 0002, Maoying Qiao, Jinlan Xu, Nannan Wang 0001
ICCV5
2025 EBaR: Efficient Buffer and Resetting for Single-Sample Continual Test-Time Adaptation
Maoying Qiao
ACM Multimedia2
2025 Disentangle Source and Target Knowledge for Continual Test-Time Adaptation
abstract
Continual Test-Time Adaptation (CoTTA) task is proposed to tackle the challenges of constant domain shifts during testing. The goals are twofold: 1) to preserve the knowledge from the source domain without source data and 2) to effectively extract target knowledge using unlabeled target domain data. Existing works primarily focus on either source or target knowledge, attempting to learn both in a mixed manner. This may harm the source knowledge preser-vation and target knowledge extraction. To this end, this pa-per proposes a Source and Target knowledge Disentangle Transformer (SoTa-DiT) with the prompting mechanism. Specifically, in a vision transformer (ViT), we employ source and target prompts, supervised by two groups of deliber-ately designed loss functions, to learn source and target knowledge separately. The source prompt focuses on anti-source-forgetting by extracting and preserving knowledge from the source model, while the target prompt focuses on protarget-extracting using target data contrastive learning. With comprehensive evaluations across various datasets using different ViT backbones, we demonstrate that this dual-prompt architecture of SoTa-DiT is effective and that disentangling knowledge with the prompts benefits CoTTA. As a result, SoTa-DiT significantly improves image classification accuracy under the CoTTA setting.
Maoying Qiao
WACV2
2025 S3OIL: Semi-Supervised SAR-to-Optical Image Translation via Multi-Scale and Cross-Set Matching
abstract
Image-to-image translation has achieved great success, but still faces the significant challenge of limited paired data, particularly in translatingSynthetic Aperture Radar(SAR) images to optical images. Furthermore, most existing semi-supervised methods place limited emphasis on leveraging the data distribution. To address those challenges, we propose aSemi-Supervised SAR-to-Optical Image Translation(S3OIL) method that achieves high-quality image generation using minimal paired data and extensive unpaired data while strategically exploiting the data distribution. To this end, we first introduce aCross-Set Alignment Matching(CAM) mechanism to create local correspondences between the generated results of paired and unpaired data, ensuring cross-set consistency. In addition, for unpaired data, we apply weak and strong perturbations and establish intra-setMulti-Scale Matching(MSM) constraints. For paired data, intra-modal semantic consistency (ISC) is presented to ensure alignment with the ground truth. Finally, we propose local and global cross-modal semantic consistency (CSC) to boost structural identity during translation. We conduct extensive experiments on SAR-to-optical datasets and another sketch-to-anime task, demonstrating that S3OIL delivers competitive performance compared to state-of-the-art unsupervised, supervised, and semi-supervised methods, both quantitatively and qualitatively. Ablation studies further reveal that S3OIL can ensure the preservation of both semantic content and structural integrity of the generated images. Our code is available at: https://github.com/XduShi/SOIL.
Xi Yang 0011, Ziyun Li 0002, Maoying Qiao, Fei Gao 0006, Nannan Wang 0001
IEEE Trans. Image Process.4
2024 Generating Handwritten Mathematical Expressions From Symbol Graphs: An End-to-End Pipeline
abstract
In this paper, we explore a novel challenging generation task, i.e. Handwritten Mathematical Expression Generation (HMEG) from symbolic sequences. Since symbolic sequences are naturally graph-structured data, we formulate HMEG as a graph-to-image (G2I) generation problem. Unlike the generation of natural images, HMEG requires critic layout clarity for synthesizing correct and recognizable formulas, but has no real masks available to supervise the learning process. To alleviate this challenge, we propose a novel end-to-end G2I generation pipeline (i.e. graph → layout →mask →image), which requires no real masks or nondifferentiable alignment between layouts and masks. Technically, to boost the capacity of predicting detailed relations among adjacent symbols, we propose a Less-is-More (LiM) learning strategy. In addition, we design a differentiable layout refinement module, which maps bounding boxes to pixel-level soft masks, so as to further alleviate ambiguous layout areas. Our whole model, including layout prediction, mask refinement, and image generation, can be jointly optimized in an end-to-end manner. Experimental results show that, our model can generate highquality HME images, and outperforms previous generative methods. Besides, a series of ablations study demonstrate effectiveness of the proposed techniques. Finally, we validate that our generated images promisingly boosts the performance of HME recognition models, through data augmentation. Our code and results are available at: https://github.com/AiArt-HDU/HMEG.
Yu Chen 0003, Fei Gao 0006, Yanguang Zhang, Maoying Qiao, Nannan Wang 0001
CVPR4
2024 Learning Discriminative Style Representations for Unsupervised and Few-Shot Artistic Portrait Drawing Generation
abstract
In this paper, we propose an unsupervised artistic portrait drawing generation method for few-shot datasets based on contrastive learning of style features. Firstly, we construct a discriminative style encoder with contrastive learning, improving the ability of the encoder to separate style features. Secondly, based on the dynamic codebook and momentum network, we used historical average features instead of batch instance features to prevent the problem of style bias in few-shot datasets. Finally, a conditional projection discriminator with filter response normalization is utilized to improve the discriminative ability of the discriminator and the stability of the generative adversarial network, which motivates the generator to synthesize more realistic image details. Quantitative and qualitative analysis show that the method proposed in this paper significantly improves the quality of artistic portrait drawing generation, and outperforms existing benchmarks in terms of visual effect and metrics evaluation. Our code and results are avilable at https://github.com/AiArt-HDU/Co-GAN.
Junkai Fang, Maoying Qiao, Fei Gao 0006
ICASSP5
2024 Human-Robot Interactive Creation of Artistic Portrait Drawings
abstract
In this paper, we present a novel system for Human-Robot Interactive Creation of Artworks (HRICA). Different from previous robot painters, HRICA allows a human user and a robot to alternately draw strokes on a canvas, to collaboratively create a portrait drawing through frequent interactions. The key is to enable the robot to understand human intentions, during the interactive creation process. We here formulate this as a mask-free image inpainting problem, and propose a novel method to estimate the complete version of a portrait drawing, after the human user has drawn some initial strokes. In this way, the robot can select some complementary strokes and draw them on the canvas. To train and evaluate our inpainting method, we construct a novel large-scale portrait drawing dataset, CelebLine, which composes of high-quality portrait line-drawings, with dense labels of both 2D semantic parsing masks and 3D depth maps. Finally, we develop a human-robot interactive drawing system with low-cost hardware, user-friendly interface, and interesting creation experience. Experiments show that our robot can stably cooperate with human users to create diverse styles of portrait drawings. In addition, our portrait drawing inpainting method significantly outperforms previous advanced methods. The code and dataset have been released at: https://github.com/fei-aiart/HRICA.
Fei Gao 0006, Lingna Dai, Jingjie Zhu, Mei Du, Maoying Qiao, Chenghao Xia, Nannan Wang 0001, Peng Li 0031
ICRA6
2024 AesMamba: Universal Image Aesthetic Assessment with State Space Models
abstract
Image Aesthetic Assessment (IAA) aims to objectively predict the generic or personalized evaluations, of the aesthetic or fine-grained multi-attributes, based on visual or multimodal inputs. Previously, researchers have designed diverse and specialized methods, for specific IAA tasks, based on different input-output situations. Is it possible to design a universal IAA framework applicable for the whole IAA task taxonomy? In this paper, we explore this issue, and propose a modular IAA framework, dubbed AesMamba. Specially, we use the Visual State Space Model (VMamba), instead of CNNs or ViTs, to learn comprehensive representations of aesthetic-related attributes; because VMamba can efficiently achieve both global and local effective receptive fields. Afterward, a modal-adaptive module is used to automatically produce the integrated representations, conditioned on the type of input. In the prediction module, we propose a Multitask Balanced Adaptation (MBA) module, to boost task-specific features, with emphasis on the tail instances. Finally, we formulate the personalized IAA task as a multimodal learning problem, by converting a user's anonymous subject characters to a text prompt. This prompting strategy effectively employs the semantics of flexibly selected characters, for inferring individual preferences. AesMamba can be applied to diverse IAA tasks, through flexible combination of these modules. Extensive experiments on numerous datasets, demonstrate that AesMamba consistently achieves superior or competitive performance, on all IAA tasks, in comparison with previous SOTA methods. The code has been released at https://github.com/AiArt-Gao/AesMamba Github.
Fei Gao 0006, Maoying Qiao, Nannan Wang 0001
ACM Multimedia4
2024 Graph Convolutional Neural Networks With Diverse Negative Samples via Decomposed Determinant Point Processes
abstract
Graph convolutional neural networks (GCNs) have achieved great success in graph representation learning by extracting high-level features from nodes and their topology. Since GCNs generally follow a message-passing mechanism, each node aggregates information from its first-order neighbor to update its representation. As a result, the representations of nodes with edges between them should be positively correlated and thus can be considered positive samples. However, there are more non-neighbor nodes in the whole graph, which provide diverse and useful information for the representation update. Two non-adjacent nodes usually have different representations, which can be seen as negative samples. Besides the node representations, the structural information of the graph is also crucial for learning. In this article, we used quality-diversity decomposition in determinant point processes (DPPs) to obtain diverse negative samples. When defining a distribution on diverse subsets of all non-neighboring nodes, we incorporate both graph structure information and node representations. Since the DPP sampling process requires matrix eigenvalue decomposition, we propose a new shortest-path-base method to improve computational efficiency. Finally, we incorporate the obtained negative samples into the graph convolution operation. The ideas are evaluated empirically in experiments on node classification tasks. These experiments show that the newly proposed methods not only improve the overall performance of standard representation learning but also significantly alleviate over-smoothing problems.
Wei Duan 0003, Junyu Xuan, Maoying Qiao, Jie Lu 0001
IEEE Trans. Neural Networks Learn. Syst.3
2023 Trust-Region Adaptive Frequency for Online Continual Learning
abstract
Abstract In the paradigm of online continual learning, one neural network is exposed to a sequence of tasks, where the data arrive in an online fashion and previously seen data are not accessible. Such online fashion causes insufficient learning and severe forgetting on past tasks issues, preventing a good stability-plasticity trade-off, where ideally the network is expected to have high plasticity to adapt to new tasks well and have the stability to prevent forgetting on old tasks simultaneously. To solve these issues, we propose a trust-region adaptive frequency approach, which alternates between standard-process and intra-process updates. Specifically, the standard-process replays data stored in a coreset and interleaves the data with current data, and the intra-process updates the network parameters based on the coreset. Furthermore, to improve the unsatisfactory performance stemming from online fashion, the frequency of the intra-process is adjusted based on a trust region, which is measured by the confidence score of current data. During the intra-process, we distill the dark knowledge to retain useful learned knowledge. Moreover, to store more representative data in the coreset, a confidence-based coreset selection is presented in an online manner. The experimental results on standard benchmarks show that the proposed method significantly outperforms state-of-art continual learning algorithms.
Yajing Kong, Liu Liu 0014, Maoying Qiao, Zhen Wang 0030, Dacheng Tao
Int. J. Comput. Vis.3
2022 Learning from the Dark: Boosting Graph Convolutional Neural Networks with Diverse Negative Samples
abstract
Graph Convolutional Neural Networks (GCNs) have been generally accepted to be an effective tool for node representations learning. An interesting way to understand GCNs is to think of them as a message passing mechanism where each node updates its representation by accepting information from its neighbours (also known as positive samples). However, beyond these neighbouring nodes, graphs have a large, dark, all-but forgotten world in which we find the non-neighbouring nodes (negative samples). In this paper, we show that this great dark world holds a substantial amount of information that might be useful for representation learning. Most specifically, it can provide negative information about the node representations. Our overall idea is to select appropriate negative samples for each node and incorporate the negative information contained in these samples into the representation updates. Moreover, we show that the process of selecting the negative samples is not trivial. Our theme therefore begins by describing the criteria for a good negative sample, followed by a determinantal point process algorithm for efficiently obtaining such samples. A GCN, boosted by diverse negative samples, then jointly considers the positive and negative information when passing messages. Experimental evaluations show that this idea not only improves the overall performance of standard representation learning but also significantly alleviates over-smoothing problems.
Wei Duan 0003, Junyu Xuan, Maoying Qiao, Jie Lu 0001
AAAI3
2020 Diversified Bayesian Nonnegative Matrix Factorization
Maoying Qiao, Jun Yu 0002, Tongliang Liu, Xinchao Wang, Dacheng Tao
AAAI1
2019 Adapting Stochastic Block Models to Power-Law Degree Distributions
abstract
Stochastic block models (SBMs) have been playing an important role in modeling clusters or community structures of network data. But, it is incapable of handling several complex features ubiquitously exhibited in real-world networks, one of which is the power-law degree characteristic. To this end, we propose a new variant of SBM, termed power-law degree SBM (PLD-SBM), by introducing degree decay variables to explicitly encode the varying degree distribution over all nodes. With an exponential prior, it is proved that PLD-SBM approximately preserves the scale-free feature in real networks. In addition, from the inference of variational E-Step, PLD-SBM is indeed to correct the bias inherited in SBM with the introduced degree decay factors. Furthermore, experiments conducted on both synthetic networks and two real-world datasets including Adolescent Health Data and the political blogs network verify the effectiveness of the proposed model in terms of cluster prediction accuracies.
Maoying Qiao, Jun Yu 0002, Wei Bian 0003, Qiang Li 0024, Dacheng Tao
IEEE Trans. Cybern.1
2017 Improving Stochastic Block Models by Incorporating Power-Law Degree Characteristic
abstract
Stochastic block models (SBMs) provide a statistical way modeling network data, especially in representing clusters or community structures. However, most block models do not consider complex characteristics of networks such as scale-free feature, making them incapable of handling degree variation of vertices, which is ubiquitous in real networks. To address this issue, we introduce degree decay variables into SBM, termed power-law degree SBM (PLD-SBM), to model the varying probability of connections between node pairs. The scale-free feature is approximated by a power-law degree characteristic. Such a property allows PLD-SBM to correct the distortion of degree distribution in SBM, and thus improves the performance of cluster prediction. Experiments on both simulated networks and two real-world networks including the Adolescent Health Data and the political blogs network demonstrate the validity of the motivation of PLD-SBM, and its practical superiority.
Maoying Qiao, Jun Yu 0002, Wei Bian 0003, Qiang Li 0024, Dacheng Tao
IJCAI1
2017 Diversified dictionaries for multi-instance learning
Maoying Qiao, Liu Liu 0014, Jun Yu 0002, Chang Xu 0002, Dacheng Tao
Pattern Recognit.1
2016 Conditional Graphical Lasso for Multi-label Image Classification
abstract
Multi-label image classification aims to predict multiple labels for a single image which contains diverse content. By utilizing label correlations, various techniques have been developed to improve classification performance. However, current existing methods either neglect image features when exploiting label correlations or lack the ability to learn image-dependent conditional label structures. In this paper, we develop conditional graphical Lasso (CGL) to handle these challenges. CGL provides a unified Bayesian framework for structure and parameter learning conditioned on image features. We formulate the multi-label prediction as CGL inference problem, which is solved by a mean field variational approach. Meanwhile, CGL learning is efficient due to a tailored proximal gradient procedure by applying the maximum a posterior (MAP) methodology. CGL performs competitively for multi-label image classification on benchmark datasets MULAN scene, PASCAL VOC 2007 and PASCAL VOC 2012, compared with the state-of-the-art multi-label classification algorithms.
Qiang Li 0024, Maoying Qiao, Wei Bian 0003, Dacheng Tao
CVPR2
2016 Diversified hidden Markov models for sequential labeling
abstract
Labeling of sequential data is a prevalent metaproblem in a wide range of real world applications. A first-order hidden Markov model (HMM) provides a fundamental approach for sequential labeling. However, it does not show satisfactory performance for real world problems, such as optical character recognition (OCR). Aiming at addressing this problem, important extensions of HMM have been proposed in literature. One of the common key features in these extensions is the incorporation of proper prior information. In this paper, we propose a new extension of HMM, termed diversified hidden Markov models (dHMM), with incorporating a diversity-encouraging prior. The prior is added over the state-transition probabilities and thus facilitates more dynamic sequential labelling. Specifically, the diversity is modeled with a continuous determinantal point process. An EM framework for parameter learning and MAP inference is derived, and empirical evaluation on OCR dataset verifies its effectiveness.
Maoying Qiao, Wei Bian 0003, Dacheng Tao
ICDE1
2016 Fast Sampling for Time-Varying Determinantal Point Processes
abstract
Determinantal Point Processes (DPPs) are stochastic models which assign each subset of a base dataset with a probability proportional to the subset’s degree of diversity. It has been shown that DPPs are particularly appropriate in data subset selection and summarization (e.g., news display, video summarizations). DPPs prefer diverse subsets while other conventional models cannot offer. However, DPPs inference algorithms have a polynomial time complexity which makes it difficult to handle large and time-varying datasets, especially when real-time processing is required. To address this limitation, we developed a fast sampling algorithm for DPPs which takes advantage of the nature of some time-varying data (e.g., news corpora updating, communication network evolving), where the data changes between time stamps are relatively small. The proposed algorithm is built upon the simplification of marginal density functions over successive time stamps and the sequential Monte Carlo (SMC) sampling technique. Evaluations on both a real-world news dataset and the Enron Corpus confirm the efficiency of the proposed algorithm.
Maoying Qiao, Wei Bian 0003, Dacheng Tao
ACM Trans. Knowl. Discov. Data1
2015 Diversified Hidden Markov Models for Sequential Labeling
abstract
Labeling of sequential data is a prevalent meta-problem for a wide range of real world applications. While the first-order Hidden Markov Models (HMM) provides a fundamental approach for unsupervised sequential labeling, the basic model does not show satisfying performance when it is directly applied to real world problems, such as part-of-speech tagging (PoS tagging) and optical character recognition (OCR). Aiming at improving performance, important extensions of HMM have been proposed in the literatures. One of the common key features in these extensions is the incorporation of proper prior information. In this paper, we propose a new extension of HMM, termed diversified Hidden Markov Models (dHMM), which utilizes a diversity-encouraging prior over the statetransition probabilities and thus facilitates more dynamic sequential labellings. Specifically, the diversity is modeled by a continuous determinantal point process prior, which we apply to both unsupervised and supervised scenarios. Learning and inference algorithms for dHMM are derived. Empirical evaluations on benchmark datasets for unsupervised PoS tagging and supervised OCR confirmed the effectiveness of dHMM, with competitive performance to the state-of-the-art.
Maoying Qiao, Wei Bian 0003, Dacheng Tao
IEEE Trans. Knowl. Data Eng.1
2011 3D human posture segmentation by spectral clustering with surface normal constraint
Jun Cheng 0002, Maoying Qiao, Wei Bian 0003, Dacheng Tao
Signal Process.2