EDBT 2026 Demo / reviewers in the wild / expert
Zhongqin Wu
dblp:222/5722
· DBLP profile ↗
26ranked-venue papers
0as first author
25since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 15 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Pseudo-ISP: Learning pseudo in-camera signal processing pipeline from a color image denoiser
Yue Cao 0009, Xiaohe Wu, Shuran Qi, Xiao Liu 0040, Zhongqin Wu, Wangmeng Zuo |
Neurocomputing | 5 |
| 2022 | All-In-One Image Restoration for Unknown CorruptionabstractIn this paper, we study a challenging problem in image restoration, namely, how to develop an all-in-one method that could recover images from a variety of unknown corruption types and levels. To this end, we propose an All-in-one Image Restoration Network (AirNet) consisting of two neural modules, named Contrastive-Based Degraded Encoder (CBDE) and Degradation-Guided Restoration Network (DGRN). The major advantages of AirNet are two-fold. First, it is an all-in-one solution which could recover various degraded images in one network. Second, AirNet is free from the prior of the corruption types and levels, which just uses the observed corrupted image to perform inference. These two advantages enable AirNet to enjoy better flexibility and higher economy in real world scenarios wherein the priors on the corruptions are hard to know and the degradation will change with space and time. Extensive experimental results show the proposed method outperforms 17 image restoration baselines on four challenging datasets. The code is available at https://github.com/XLearning-SCU/2022-CVPR-AirNet. Boyun Li, Xiao Liu 0040, Peng Hu 0002, Zhongqin Wu, Jiancheng Lv 0001, Xi Peng 0001 |
CVPR | 4 |
| 2022 | Retrieval-based Spatially Adaptive Normalization for Semantic Image SynthesisabstractSemantic image synthesis is a challenging task with many practical applications. Albeit remarkable progress has been made in semantic image synthesis with spatiallyadaptive normalization, existing methods usually normalize the feature activations under the coarse-level guidance (e.g., semantic class). However, different parts of a semantic object (e.g., wheel and window of car) are quite different in structures and textures, making blurry synthesis results usually inevitable due to the missing of fine-grained guidance. In this paper, we propose a novel normalization module, termed as REtrieval-based Spatially Adaptive normaLization (RESAIL), for introducing pixel level fine- grained guidance to the normalization architecture. Specifically, we first present a retrieval paradigm by finding a content patch of the same semantic class from training set with the most similar shape to each test semantic mask. Then, the retrieved patches are composited into retrieval-based guidance, which can be used by RESAIL for pixel level fine-grained modulation on feature activations, thereby greatly mitigating blurry synthesis results. Moreover, distorted ground-truth images are also utilized as alternatives of retrieval-based guidance for feature normalization, further benefiting model training and improving visual quality of generated images. Experiments on several challenging datasets show that our RESAIL performs favorably against state-of-the-arts in terms of quantitative metrics, visual quality, and subjective evaluation. The source code is available at https://github.com/Shi-Yupeng/RESAIL-For-SIS. Yupeng Shi, Xiao Liu 0040, Yuxiang Wei 0001, Zhongqin Wu, Wangmeng Zuo |
CVPR | 4 |
| 2022 | Syntax-Aware Network for Handwritten Mathematical Expression RecognitionabstractHandwritten mathematical expression recognition (HMER) is a challenging task that has many potential applications. Recent methods for HMER have achieved outstanding performance with an encoder-decoder architecture. However, these methods adhere to the paradigm that the prediction is made “from one character to another”, which inevitably yields prediction errors due to the complicated structures of mathematical expressions or crabbed handwritings. In this paper, we propose a simple and efficient method for HMER, which is the first to incorporate syntax information into an encoder-decoder network. Specifically, we present a set of grammar rules for converting the LaTeX markup sequence of each expression into a parsing tree; then, we model the markup sequence prediction as a tree traverse process with a deep neural network. In this way, the proposed method can effectively describe the syntax context of expressions, alleviating the structure prediction errors of HMER. Experiments on three benchmark datasets demonstrate that our method achieves better recognition performance than prior arts. To further validate the effectiveness of our method, we create a large-scale dataset consisting of 100k handwritten mathematical expression images acquired from ten thousand writers. The source code, new dataset††https://ai.100tal.com/dataset, and pre-trained models of this work will be publicly available. Xiao Liu 0040, Wondimu Dikubab, Zhilong Ji, Zhongqin Wu, Xiang Bai |
CVPR | 6 |
| 2022 | A Character-Level Span-Based Model for Mandarin Prosodic Structure PredictionabstractThe accuracy of prosodic structure prediction is crucial to the naturalness of synthesized speech in Mandarin text-to-speech system, but now is limited by widely-used sequence-to-sequence framework and error accumulation from previous word segmentation results. In this paper, we propose a span-based Mandarin prosodic structure prediction model to obtain an optimal prosodic structure tree, which can be converted to corresponding prosodic label sequence. Instead of the prerequisite for word segmentation, rich linguistic features are provided by Chinese character-level BERT and sent to encoder with self-attention architecture. On top of this, span representation and label scoring are used to describe all possible prosodic structure trees, of which each tree has its corresponding score. To find the optimal tree with the highest score for a given sentence, a bottom-up CKYstyle algorithm is further used. The proposed method can predict prosodic labels of different levels at the same time and accomplish the process directly from Chinese characters in an end-to-end manner. Experiment results on two real-world datasets demonstrate the excellent performance of our span-based method over all sequence-to-sequence baseline approaches. Xueyuan Chen, Changhe Song, Yixuan Zhou 0002, Zhiyong Wu 0001, Changbin Chen, Zhongqin Wu, Helen M. Meng |
ICASSP | 6 |
| 2022 | Time-Domain Audio-Visual Speech Separation on Low Quality VideosabstractIncorporating visual information is a promising approach to improve the performance of speech separation. Many related works have been conducted and provide inspiring results. However, low quality videos appear commonly in real scenarios, which may significantly degrade the performance of normal audio-visual speech separation system. In this paper, we propose a new structure to fuse the audio and visual features, which uses the audio feature to select relevant visual features by utilizing the attention mechanism. A Conv-TasNet based model is combined with the proposed attention-based multi-modal fusion, trained with proper data augmentation and evaluated with 3 categories of low quality videos. The experimental results show that our system outperforms the baseline which simply concatenates the audio and visual features when training with normal or low quality data, and is robust to low quality video inputs at inference time. Chenda Li, Jinfeng Bai, Zhongqin Wu, Yanmin Qian |
ICASSP | 4 |
| 2022 | AdvExpander: Generating Natural Language Adversarial Examples by Expanding TextabstractAdversarial examples are vital to expose vulnerability of machine learning models. Despite the success of the most popular word-level substitution-based attacks which substitute some words in the original examples, only substitution is insufficient to uncover all robustness issues of models. In this paper, we focus on perturbations beyond word-level substitution, and presentAdvExpander, a method that crafts new adversarial examples by expanding text. We first utilize linguistic rules to determine which constituents to expand and what types of modifiers to expand with. We then expand each constituent by inserting an adversarial modifier searched from a pre-trained CVAE-based generative model. To ensure that our adversarial examples are label-preserving for text matching, we also constrain the modifications with a heuristic rule. Experiments on three classification tasks verify the effectiveness of AdvExpander and the validity of our adversarial examples. AdvExpander is significantly more effective than sentence-level attack baselines and is complementary to previous word substitution-based attacks, thus promising to reveal new robustness issues. Zhihong Shao, Zhongqin Wu, Minlie Huang |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Locality-Aware Channel-Wise Dropout for Occluded Face RecognitionabstractFace recognition remains a challenging task in unconstrained scenarios, especially when faces are partially occluded. To improve the robustness against occlusion, augmenting the training images with artificial occlusions has been proved as a useful approach. However, these artificial occlusions are commonly generated by adding a black rectangle or several object templates including sunglasses, scarfs and phones, which cannot well simulate the realistic occlusions. In this paper, based on the argument that the occlusion essentially damages a group of neurons, we propose a novel and elegant occlusion-simulation method via dropping the activations of a group of neurons in some elaborately selected channel. Specifically, we first employ a spatial regularization to encourage each feature channel to respond to local and different face regions. Then, the locality-aware channel-wise dropout (LCD) is designed to simulate occlusions by dropping out a few feature channels. The proposed LCD can encourage its succeeding layers to minimize the intra-class feature variance caused by occlusions, thus leading to improved robustness against occlusion. In addition, we design an auxiliary spatial attention module by learning a channel-wise attention vector to reweight the feature channels, which improves the contributions of non-occluded regions. Extensive experiments on various benchmarks show that the proposed method outperforms state-of-the-art methods with a remarkable improvement. Jie Zhang 0071, Shiguang Shan, Xiao Liu 0040, Zhongqin Wu, Xilin Chen 0001 |
IEEE Trans. Image Process. | 5 |
| 2021 | Multi-task Learning Based Online Dialogic Instruction Detection with Pre-trained Language Models
Yang Hao 0004, Hang Li 0007, Wenbiao Ding, Zhongqin Wu, Jiliang Tang, Rosemary Luckin, Zitao Liu 0001 |
AIED (2) | 4 |
| 2021 | A Multimodal Machine Learning Framework for Teacher Vocal Delivery Evaluation
Hang Li 0007, Yang Hao 0004, Wenbiao Ding, Zhongqin Wu, Zitao Liu 0001 |
AIED (2) | 5 |
| 2021 | Solving ESL Sentence Completion Questions via Pre-trained Neural Language Models
Qiongqiong Liu, Tianqiao Liu, Jiafu Zhao, Wenbiao Ding, Zhongqin Wu, Feng Xia 0001, Jiliang Tang, Zitao Liu 0001 |
AIED (2) | 6 |
| 2021 | Automatic Task Requirements Writing Evaluation via Machine Reading Comprehension
Peilei Jia, Wenbiao Ding, Zhongqin Wu, Zitao Liu 0001 |
AIED (1) | 5 |
| 2021 | FAIEr: Fidelity and Adequacy Ensured Image Caption EvaluationabstractImage caption evaluation is a crucial task, which involves the semantic perception and matching of image and text. Good evaluation metrics aim to be fair, comprehensive, and consistent with human judge intentions. When humans evaluate a caption, they usually consider multiple aspects, such as whether it is related to the target image without distortion, how much image gist it conveys, as well as how fluent and beautiful the language and wording is. The above three different evaluation orientations can be summarized as fidelity, adequacy, and fluency. The former two rely on the image content, while fluency is purely related to linguistics and more subjective. Inspired by human judges, we propose a learning-based metric named FAIEr to ensure evaluating the fidelity and adequacy of the captions. Since image captioning involves two different modalities, we employ the scene graph as a bridge between them to represent both images and captions. FAIEr mainly regards the visual scene graph as the criterion to measure the fidelity. Then for evaluating the adequacy of the candidate caption, it high-lights the image gist on the visual scene graph under the guidance of the reference captions. Comprehensive experimental results show that FAIEr has high consistency with human judgment as well as high stability, low reference dependency, and the capability of reference-free evaluation. Sijin Wang, Ziwei Yao, Ruiping Wang 0001, Zhongqin Wu, Xilin Chen 0001 |
CVPR | 4 |
| 2021 | CTAL: Pre-training Cross-modal Transformer for Audio-and-Language RepresentationsabstractExisting approaches for audio-language taskspecific prediction focus on building complicated late-fusion mechanisms.However, these models face challenges of overfitting with limited labels and poor generalization.In this paper, we present a Cross-modal Transformer for Audio-and-Language, i.e., CTAL, which aims to learn the intra-and inter-modalities connections between audio and language through two proxy tasks from a large number of audio-and-language pairs: masked language modeling and masked cross-modal acoustic modeling.After fine-tuning our CTAL model on multiple downstream audioand-language tasks, we observe significant improvements on different tasks, including emotion classification, sentiment analysis, and speaker verification.Furthermore, we design a fusion mechanism in the fine-tuning phase, which allows CTAL to achieve better performance.Lastly, we conduct detailed ablation studies to demonstrate that both our novel cross-modality fusion component and audiolanguage pre-training methods contribute to the promising results.The code and pretrained models are available at https:// github.com/tal-ai/CTAL_EMNLP2021. Hang Li 0007, Wenbiao Ding, Tianqiao Liu, Zhongqin Wu, Zitao Liu 0001 |
EMNLP (1) | 5 |
| 2021 | Mathematical Word Problem Generation from Commonsense Knowledge Graph and EquationsabstractThere is an increasing interest in the use of mathematical word problem (MWP) generation in educational assessment.Different from standard natural question generation, MWP generation needs to maintain the underlying mathematical operations between quantities and variables, while at the same time ensuring the relevance between the output and the given topic.To address above problem, we develop an end-to-end neural model to generate diverse MWPs in real-world scenarios from commonsense knowledge graph and equations.The proposed model (1) learns both representations from edge-enhanced Levi graphs of symbolic equations and commonsense knowledge; (2) automatically fuses equation and commonsense knowledge information via a self-planning module when generating the MWPs.Experiments on an educational gold-standard set and a large-scale generated MWP set show that our approach is superior on the MWP generation task, and it outperforms the SOTA models in terms of both automatic evaluation metrics, i.e., BLEU-4, ROUGE-L, Self-BLEU, and human evaluation metrics, i.e., equation relevance, topic relevance, and language coherence.To encourage reproducible results, we make our code and MWP dataset public available at https:// github.com/tal-ai/MaKE_EMNLP2021. Tianqiao Liu, Wenbiao Ding, Hang Li 0007, Zhongqin Wu, Zitao Liu 0001 |
EMNLP (1) | 5 |
| 2021 | Local Global Relational Network for Facial Action Units RecognitionabstractMany existing facial action units (AUs) recognition approaches often enhance the AU representation by combining local features from multiple independent branches, each corresponding to a different AU. However, such multi-branch combination-based methods usually neglect potential mutual assistance and exclusion relationship between AU branches or simply employ a pre-defined and fixed knowledge-graph as a prior. In addition, extracting features from pre-defined AU regions of regular shapes limits the representation ability. In this paper, we propose a novel Local Global Relational Network (LGRNet) for facial AU recognition. LGRNet mainly consists of two novel structures, i.e., a skip-BiLSTM module which models the latent mutual assistance and exclusion relationship among local AU features from multiple branches to enhance the feature robustness, and a feature fusion&refining module which explores the complementarity between local AUs and the whole face in order to refine the local AU features to improve the discriminability. Experiments on the BP4D and DISFA AU datasets show that the proposed approach outperforms the state-of-the-art methods by a large margin. Xuri Ge, Hu Han 0001, Joemon M. Jose, Zhilong Ji, Zhongqin Wu, Xiao Liu 0040 |
FG | 6 |
| 2021 | Orthogonal Jacobian Regularization for Unsupervised Disentanglement in Image GenerationabstractUnsupervised disentanglement learning is a crucial issue for understanding and exploiting deep generative models. Recently, SeFa tries to find latent disentangled directions by performing SVD on the first projection of a pretrained GAN. However, it is only applied to the first layer and works in a post-processing way. Hessian Penalty minimizes the off-diagonal entries of the output’s Hessian matrix to facilitate disentanglement, and can be applied to multi-layers. However, it constrains each entry of output independently, making it not sufficient in disentangling the latent directions (e.g., shape, size, rotation, etc.) of spatially correlated variations. In this paper, we propose a simple Orthogonal Jacobian Regularization (OroJaR) to encourage deep generative model to learn disentangled representations. It simply encourages the variation of output caused by perturbations on different latent dimensions to be orthogonal, and the Jacobian with respect to the input is calculated to represent this variation. We show that our OroJaR also encourages the output’s Hessian matrix to be diagonal in an indirect manner. In contrast to the Hessian Penalty, our OroJaR constrains the output in a holistic way, making it very effective in disentangling latent dimensions corresponding to spatially correlated variations. Quantitative and qualitative experimental results show that our method is effective in disentangled and controllable image generation, and performs favorably against the state-of-the-art methods. Our code is available at https://github.com/csyxwei/OroJaR. Yuxiang Wei 0001, Yupeng Shi, Xiao Liu 0040, Zhilong Ji, Zhongqin Wu, Wangmeng Zuo |
ICCV | 6 |
| 2021 | CrowdRL: An End-to-End Reinforcement Learning Framework for Data LabellingabstractData labelling is very important in many database and machine learning applications. Traditional methods rely on humans (workers or experts) to acquire labels. However, the human cost is rather expensive for a large dataset. Active learning based methods only label a small set of data with large uncertainty, train a model on these labelled data, and use the trained model to label the remainder unlabelled data. However they have two limitations. First, they cannot judiciously select appropriate data (task selection) and assign the tasks to proper humans (task assignment). Moreover, they independently process task selection and task assignment, which cannot capture the correlation between them. Second, they simply infer the truth of a task based on the answers from humans and the trained model (truth inference) by independently modeling humans and models. In other words, they ignore the correlation between them (the labelled data may have noise caused by humans with biases, and the model trained by the noisy labels may bring additional biases), and thus lead to poor inference results. To address these limitations, in this paper, we propose CrowdRL, an end-to-end reinforcement learning (RL) based framework for data labelling. To the best of our knowledge, CrowdRL is the first RL framework designed for the data labelling workflow by seamlessly integrating task selection, task assignment and truth inference together. CrowdRL fully utilizes the power of heterogeneous annotators (experts and crowdsourcing workers) and machine learning models together to infer the truth, which highly improves the quality of data labelling. CrowdRL uses RL to model task assignment and task selection, and designs an agent to judiciously assign tasks to appropriate workers. CrowdRL jointly models the answers of workers, experts and models, and designs a joint inference model to infer the truths. Experimental results on real datasets show that CrowdRL outperforms state-of-the-art approaches with the same (even fewer) monetary cost while achieving 5%-20% higher accuracy. Guoliang Li 0001, Yong Wang 0088, Zitao Liu 0001, Zhongqin Wu |
ICDE | 6 |
| 2021 | Learning Fine-Grained Cross Modality Excitement for Speech Emotion RecognitionabstractSpeech emotion recognition is a challenging task because the emotion expression is complex, multimodal and fine-grained. In this paper, we propose a novel multimodal deep learning approach to perform fine-grained emotion recognition from real-life speeches. We design a temporal alignment mean-max pooling mechanism to capture the subtle and fine-grained emotions implied in every utterance. In addition, we propose a cross modality excitement module to conduct sample-specific adjustment on cross modality embeddings and adaptively recalibrate the corresponding values by its aligned latent features from the other modality. Our proposed model is evaluated on two well-known real-world speech emotion recognition datasets. The results demonstrate that our approach is superior on the prediction tasks for multimodal speech utterances, and it outperforms a wide range of baselines in terms of prediction accuracy. Further more, we conduct detailed ablation studies to show that our temporal alignment mean-max pooling mechanism and cross modality excitement significantly contribute to the promising results. In order to encourage the research reproducibility, we make the code publicly available at \url{https://github.com/tal-ai/FG_CME.git}. Hang Li 0007, Wenbiao Ding, Zhongqin Wu, Zitao Liu 0001 |
Interspeech | 3 |
| 2021 | Audio-Visual Multi-Talker Speech Recognition in a Cocktail Party
Chenda Li, Zhongqin Wu, Yanmin Qian |
Interspeech | 4 |
| 2021 | The TAL System for the INTERSPEECH2021 Shared Task on Automatic Speech Recognition for Non-Native Childrens Speech
Gaopeng Xu, Chengfei Li, Zhongqin Wu |
Interspeech | 5 |
| 2021 | Language Recognition Based on Unsupervised Pretrained Models
Zhongqin Wu, Yuting Nie |
Interspeech | 4 |
| 2021 | Structured Multi-modal Feature Embedding and Alignment for Image-Sentence RetrievalabstractThe current state-of-the-art image-sentence retrieval methods implicitly align the visual-textual fragments, like regions in images and words in sentences, and adopt attention modules to highlight the relevance of cross-modal semantic correspondences. However, the retrieval performance remains unsatisfactory due to a lack of consistent representation in both semantics and structural spaces. In this work, we propose to address the above issue from two aspects: (i) constructing intrinsic structure (along with relations) among the fragments of respective modalities, e.g., "dog → play → ball" in semantic structure for an image, and (ii) seeking explicit inter-modal structural and semantic correspondence between the visual and textual modalities. Xuri Ge, Fuhai Chen, Joemon M. Jose, Zhilong Ji, Zhongqin Wu, Xiao Liu 0040 |
ACM Multimedia | 5 |
| 2021 | UniCon: Unified Context Network for Robust Active Speaker DetectionabstractWe propose a new efficient framework, the Unified Context Network (UniCon), for robust active speaker detection (ASD). Traditional methods for ASD usually operate on each candidate's pre-cropped face track separately and do not sufficiently consider the relationships among the candidates. This potentially limits performance, especially in challenging scenarios with low-resolution faces, multiple candidates, etc. Our solution is a novel, unified framework that focuses on jointly modeling multiple types of contextual information: spatial context to indicate the position and scale of each candidate's face, relational context to capture the visual relationships among the candidates and contrast audio-visual affinities with each other, and temporal context to aggregate long-term information and smooth out local uncertainties. Based on such information, our model optimizes all candidates in a unified process for robust and reliable ASD. A thorough ablation study is performed on several challenging ASD benchmarks under different settings. In particular, our method outperforms the state-of-the-art by a large margin of about 15% mean Average Precision (mAP) absolute on two challenging subsets: one with three candidate speakers, and the other with faces smaller than 64 pixels. Together, our UniCon achieves 92.0% mAP on the AVA-ActiveSpeaker validation set, surpassing 90% for the first time on this challenging dataset at the time of submission. Project website: https://unicon-asd.github.io/. Yuanhang Zhang 0001, Susan Liang, Xiao Liu 0040, Zhongqin Wu, Shiguang Shan, Xilin Chen 0001 |
ACM Multimedia | 5 |
| 2021 | Robust Learning for Text Classification with Multi-source Noise Simulation and Hard Example Mining
Wenbiao Ding, Weiping Fu, Zhongqin Wu, Zitao Liu 0001 |
ECML/PKDD (5) | 4 |
| 2020 | Personalized Multimodal Feedback Generation in EducationabstractThe automatic feedback of school assignments is an important application of AI in education.In this work, we focus on the task of personalized multimodal feedback generation, which aims to generate personalized feedback for teachers to evaluate students' assignments involving multimodal inputs such as images, audios, and texts.This task involves the representation and fusion of multimodal information and natural language generation, which presents the challenges from three aspects: (1) how to encode and integrate multimodal inputs; (2) how to generate feedback specific to each modality; and (3) how to fulfill personalized feedback generation.In this paper, we propose a novel Personalized Multimodal Feedback Generation Network (PMFGN) armed with a modality gate mechanism and a personalized bias mechanism to address these challenges.Extensive experiments on real-world K-12 education data show that our model significantly outperforms baselines by generating more accurate and diverse feedback.In addition, detailed ablation experiments are conducted to deepen our understanding of the proposed framework. Zitao Liu 0001, Zhongqin Wu, Jiliang Tang |
COLING | 3 |