EDBT 2026 Demo / reviewers in the wild / expert
Weidong Tian 0001
dblp:40/7075-1 · also Wei-Dong Tian 0001
· DBLP profile ↗
35ranked-venue papers
16as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 18 · 9 first-author · 9 since 2021Artificial intelligence and machine learning · 14 · 6 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Q-FETA:Question-Guided Fine-Grained Feature Transport Alignment for Robust Visual Question Answering
Weidong Tian 0001, Zhong-Qiu Zhao |
ICIC (9) | 1 |
| 2025 | Cross-Resolution Deep Face Recognition via Collaborative Knowledge Distillation
Weidong Tian 0001, Zejun Gu, Zhong-Qiu Zhao |
ICIC (17) | 1 |
| 2025 | Capsule network with using shifted windows for 3D human pose estimation
Xiufeng Liu 0005, Zhong-Qiu Zhao, Weidong Tian 0001, Hongmei He |
J. Vis. Commun. Image Represent. | 3 |
| 2024 | Image Captioning with Masked Diffusion Model
Weidong Tian 0001, Wenzheng Xu, Junxiang Zhao, Zhong-Qiu Zhao |
ICIC (8) | 1 |
| 2024 | Dual-Branch Collaborative Learning for Visual Question Answering
Weidong Tian 0001, Junxiang Zhao, Wenzheng Xu, Zhong-Qiu Zhao |
ICIC (3) | 1 |
| 2024 | Contextual Feature Modulation Network for Efficient Super-Resolution
Wandi Zhang, Hao Shen 0006, Weidong Tian 0001, Zhong-Qiu Zhao |
ICIC (6) | 4 |
| 2024 | DDNet: Detection-Focused Dehazing Network
Weidong Tian 0001, Wandi Zhang, Zhong-Qiu Zhao |
ICIC (10) | 2 |
| 2024 | GSLip: A Global Lip-Reading Framework with Solid Dilated ConvolutionsabstractThe mainstream lip-reading framework employs Residual Network (ResNet) for spatial feature extraction and Temporal Convolutional Network (TCN) for constructing the temporal model. Addressing the modes of single-frame spatial and multi-frame temporal information extraction, we analyze and propose respective improvements. When extracting single-frame lip information, it is crucial to consider not only local fine-grained information but also global coarse-grained information. Global features such as lip contour, size, and movement amplitude significantly contribute to extracting local fine-grained lip features. However, traditional convolution, a key component of ResNet, is primarily utilized to extract features based on sliding windows in small regions, making ResNet less effective at capturing global features. Additionally, we observe key motion changes appearing in only a few consecutive frames in lipreading task. For multi-frame temporal information extraction, TCN is found to be inadequate in handling local continuous temporal dependencies. Therefore, we develop GSLip, a global lip-reading framework with solid dilated convolutions. The key ideas of GSLip include: 1) The Residual Global Context Network (ResGNet), which addresses the deficiency of traditional convolution in extracting global features; 2) the Continuous Temporal Convolutional Network (C-TCN), that enhances the ability to extract continuous temporal information. With the incorporation of these modules, our method has demonstrated excellent results on the Lip Reading in the Wild (LRW) and LRW-1000 datasets. Junxia Jiang, Zhong-Qiu Zhao, Weidong Tian 0001 |
IJCNN | 4 |
| 2023 | Lightweight Human Pose Estimation Based on Densely Guided Self-Knowledge Distillation
Mingyue Wu, Zhong-Qiu Zhao, Weidong Tian 0001 |
ICANN (2) | 4 |
| 2023 | Semantic Object Alignment and Region-Aware Learning for Change CaptioningabstractThe change captioning task, a downstream task of the image captioning task, is an emerging deep learning task. It aims to output sentences to describe the differences between two images, one of which is the original image and the other is the image changed from the original. Many current methods for generating disparity descriptions are based on the encoder-decoder model. The process is to first obtain the grid features using pre-trained ResNet on the image and encode the features, and then to obtain the sentences via a decoder. However, using traditional methods of encoding area features and applying them to grid features regardless can lead to performance degradation. Therefore, we propose a new model incorporating a novel design that is more robust to describe variations of image pairs, and a new addition of pseudo-region features that reduce the information missing from the grid features. Intensive experiments we have done can adequately demonstrate that our proposed model achieves state-of-the-art performance. Weidong Tian 0001, Quan Ren, Zhong-Qiu Zhao, Ruihua Tian |
IJCNN | 1 |
| 2023 | An Answer FeedBack Network for Visual Question AnsweringabstractRecent advances have explored the power of transformer architecture in Visual Question Answering(VQA). However, most of the models suffer from misalignment of multimodal features, and they focus on unimportant image regions when answering the given questions. To address this, in this paper, we propose an Answer FeedBack Network (AFBN) to focus on image region features that are more beneficial for answering questions. The generate answers of the backbone network are again inputted into the network as feedback information. Then, we propose a FeedBack module (FB) to control the answer feedback. Additionally, we adopt the consistency loss function to reconstruct the image region features. By this function, the model can ensure the same of the image region features related to the question or answer. Extensive experiments on VQA-v2 benchmark dataset show that our method achieves better performance than the state-of-the-art methods. Weidong Tian 0001, Ruihua Tian, Zhong-Qiu Zhao, Quan Ren |
IJCNN | 1 |
| 2022 | Dual Capsule Attention Mask Network with Mutual Learning for Visual Question AnsweringabstractA Visual Question Answering (VQA) model processes images and questions simultaneously with rich semantic information. The attention mechanism can highlight fine-grained features with critical information, thus ensuring that feature extraction emphasizes the objects related to the questions. However, unattended coarse-grained information is also essential for questions involving global elements. We believe that global coarse-grained information and local fine-grained information can complement each other to provide richer comprehensive information. In this paper, we propose a dual capsule attention mask network with mutual learning for VQA. Specifically, it contains two branches processing coarse-grained features and fine-grained features, respectively. We also design a novel stackable dual capsule attention module to fuse features and locate evidence. The two branches are combined to make final predictions for VQA. Experimental results show that our method outperforms the baselines in terms of VQA performance and interpretability and achieves new SOTA performance on the VQA-v2 dataset. Weidong Tian 0001, Zhong-Qiu Zhao |
COLING | 1 |
| 2022 | Lip Reading Using Deformable 3D Convolution and Channel-Temporal Attention
Jie Chai, Zhong-Qiu Zhao, Housen Zhang, Weidong Tian 0001 |
ICANN (4) | 6 |
| 2022 | Lipreading Model Based On Whole-Part Collaborative LearningabstractLipreading is a task to recognize speech content from visual information of the speaker’s lips movements. Recently, some work has focused more on how to adequately extract temporal information, and spatial information is simply used after extraction. In this paper, we focus on the full use of spatial information in lipreading tasks. The whole lip represents global spatial information, while the parts of the lip contain fine-grained spatial information. We propose the lipreading model based on whole-part collaborative learning (WPCL), which can help this model make full use of both global and fine-grained spatial information of the lip. WPCL contains two branches, which deal with the whole and the part features respectively and are trained jointly by collaborative learning. Further, in order to highlight the different importance of part features when fusing them, we propose an adaptive part features fusion module (APFF) to fusion part features. Finally, we prove our viewpoints and evaluate our WPCL by severed experiments. Experiments on LRW and CAS-VSR-W1k datasets demonstrate that our approach achieves state-of-the-art performance. Weidong Tian 0001, Housen Zhang, Zhong-Qiu Zhao |
ICASSP | 1 |
| 2022 | A Sub-captions Semantic-Guided Network for Image Captioning
Weidong Tian 0001, Jun-jun Zhu, Zhong-Qiu Zhao, Yu-Zheng Zhang |
ICIC (3) | 1 |
| 2022 | Improved YOLOv5 Network with Attention and Context for Small Object Detection
Jie Chai, Zhong-Qiu Zhao, Weidong Tian 0001 |
ICIC (3) | 5 |
| 2022 | Joint operation and attention block search for lightweight image restoration
Hao Shen 0006, Zhong-Qiu Zhao, Wenrui Liao, Weidong Tian 0001, De-Shuang Huang |
Pattern Recognit. | 4 |
| 2021 | Multi-Branch Network for Small Human Pose Estimation
Yuchen Ge, Zhong-Qiu Zhao, Weidong Tian 0001, Hai Min |
ICANN (3) | 4 |
| 2021 | Visual-Textual Semantic Alignment Network for Visual Question Answering
Weidong Tian 0001, Yuzheng Zhang, Junjun Zhu, Zhong-Qiu Zhao |
ICANN (5) | 1 |
| 2021 | VISFF: An Approach for Video Summarization Based on Feature Fusion
Weidong Tian 0001, Xiao-Yu Cheng, Zhong-Qiu Zhao |
ICIC (2) | 1 |
| 2021 | Residual Attention Block Search for Lightweight Image Super-ResolutionabstractRecently, lightweight neural networks with different manual designs have presented a promising performance in single image super-resolution (SR). However, these designs rely on too much expert experience. To address this issue, we focus on searching a lightweight block for efficient and accurate image SR. Due to the frequent use of various residual blocks and attention mechanisms in SR methods, we propose the residual attention search block (RASB) which combines an operation search block (OSB) with an attention search block (ASB). The former is used to explore the suitable operation at the proper position, and the latter is applied to discover the optimal connection of various attention mechanisms. Moreover, we build the modified residual attention network (MRAN) with stacked found blocks and a refinement module. Extensive experiments demonstrate that our MRAN achieves a better trade-off against the state-of-the-art methods in terms of accuracy and model complexity. Wenrui Liao, Zhong-Qiu Zhao, Hao Shen 0006, Weidong Tian 0001 |
ICME | 4 |
| 2020 | Real-Time Object Detection Based on Convolutional Block Attention Module
Ming-Yang Ban, Weidong Tian 0001, Zhong-Qiu Zhao |
ICIC (3) | 2 |
| 2020 | Image Super-Resolution Network Based on Prior Information Fusion
Weidong Tian 0001, Zhong-Qiu Zhao |
ICIC (3) | 2 |
| 2020 | TFPGAN: Tiny Face Detection with Prior Information and GAN
Dian Liu, Zhong-Qiu Zhao, Weidong Tian 0001 |
ICIC (3) | 3 |
| 2020 | Regenerating Image Caption with High-Level Semantics
Weidong Tian 0001, Nan-Xun Wang, Yue-Lin Sun, Zhong-Qiu Zhao |
ICIC (3) | 1 |
| 2020 | Multi-Channel Co-Attention Network for Visual Question AnsweringabstractVisual Question Answering (VQA) is to reason out correct answers based on input questions and images. Significant progresses have been made by learning rich embedding features from images and questions by bilinear models. Attention mechanisms are widely used to focus on specific visual and textual information in VQA reasoning process. However, most state-of-the-art methods concentrate on fusing the global multi-modal features, while neglect local features. Besides, the dimension is reduced excessively (from K×2048 to 2048) in general visual attention, which causes a mass of visual information loss. In this paper, we propose a novel multi-channel co-attention network (MC-CAN), which integrates multi-modal features from global level to local level. We design different multi-channel attention mechanisms separately for visual (from K×2048 to M×2048) and textual features at different level of integrations. Additionally, we further improve our proposed approach by combining it with the complementary modules such as the MLB and the Count modules. Experiments on benchmark datasets show that our approach achieves better VQA performance than other state-of-the-art methods. Weidong Tian 0001, Nanxun Wang, Zhong-Qiu Zhao |
IJCNN | 1 |
| 2020 | Cascading Top-Down Attention for Visual Question AnsweringabstractFor solving Visual Question Answering (VQA), we commonly employ images and questions simultaneously to predict answers. Some attention mechanisms should be used to focus on the most valuable information, because there are too much information extracted from images and questions. Top-Down Attention (TDA) is one of the famous attention mechanisms. For standard TDA, only important regions of the image associated with the question are highlighted. In this work, we propose a Cascading Top-Down Attention (CTDA) model. CTDA highlights the most important information collected from images and questions by a cascading attention process. First, the key words of the question, associated with the image, are highlighted by using a Question Top-Down Attention (QTDA). Then, important regions of the image, associated with the question, are highlighted by using of Image Top-Down Attention (ITDA), useless information of the images and questions can be ignored effectively. We evaluate our model on two popular VQA data sets. CTDA obtains better results than standard TDA and the other state of the art models. Weidong Tian 0001, Rencai Zhou, Zhong-Qiu Zhao |
IJCNN | 1 |
| 2020 | Extricating from GroundTruth: An Unpaired Learning Based Evaluation Metric for Image CaptioningabstractRecently, instead of pursuing high performance on classical evaluation metrics, the research focus of image captioning has shifted to generating sentences which are more vivid and stylized than human-written ones. However, there are still no applicable metrics which can judge how close the generated captions are to the human-written ones. In this paper, we propose a novel learning-based evaluation metric, namely Unpaired Image Captioning Evaluation (UICE), which can be trained to distinguish between human-written and generated captions. Unlike existing metrics, our UICE consists of two parts: the semantic alignment module measuring the semantic distance between extracted image features and caption meanings, and the syntactic discriminating module syntactically judging how human-like the candidate caption is. The semantic alignment module is implemented by mapping the image features and the word embedding into a unified tensor space. And the syntactic discriminating module is designed to be learning-based, and thereby can be trained to be stylized by users' own, fed with additional personalized corpus during the training process. Extensive experiments indicate that our metric can correctly judge the grammatical correctness of generated captions and the semantic consistency between captions and corresponding images. Zhong-Qiu Zhao, Yue-Lin Sun, Nan-Xun Wang, Weidong Tian 0001 |
IJCNN | 4 |
| 2019 | Person Re-identification Based on Feature Fusion
Qiang-Qiang Ren, Weidong Tian 0001, Zhong-Qiu Zhao |
ICIC (3) | 2 |
| 2019 | A Novel Concise Representation of Frequent Subtrees Based on Density
Weidong Tian 0001, Chuang Guo, Hongjuan Zhou, Zhong-Qiu Zhao |
ICIC (3) | 1 |
| 2019 | Coarse-to-Fine Supervised Descent Method for Face Alignment
Xijing Zhu, Zhong-Qiu Zhao, Weidong Tian 0001 |
ICIC (1) | 3 |
| 2019 | A review of image set classification
Zhong-Qiu Zhao, Shou-tao Xu, Dian Liu, Weidong Tian 0001, Zhi-Da Jiang |
Neurocomputing | 4 |
| 2018 | Cooperative Adversarial Network for Accurate Super Resolution
Zhong-Qiu Zhao, Weidong Tian 0001, Ning Ling |
ACCV (2) | 3 |
| 2018 | Improved Adaptive Incremental Error-Minimization-Based Extreme Learning Machine with Localized Generalization Error Model
Wen-wen Han, Zhong-Qiu Zhao, Weidong Tian 0001 |
ICIC (3) | 4 |
| 2018 | GridWall: A Novel Condensed Representation of Frequent Itemsets
Weidong Tian 0001, Jianqiang Mei, Hongjuan Zhou, Zhong-Qiu Zhao |
ICIC (1) | 1 |