Yuechen Wang

dblp:233/7798 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Incremental Transformer: Efficient Encoder for Incremented Text Over MRC and Conversation Tasks
abstract
Some encoder inputs such as conversation histories are frequently extended with short additional inputs like new responses. However, to obtain the real-time encoding of the extended input, existing Transformer-based encoders like BERT have to encode the whole extended input again without utilizing the existing encoding of the original input, which may be prohibitively slow for real-time applications. In this paper, we introduce Incremental Transformer, an efficient encoder dedicated for faster encoding of incremented input. It takes only added input as input but attends to cached representations of original input in lower layers for better performance. By treating questions as additional inputs of a passage, Incremental Transformer can also be applied to accelerate MRC tasks. Experimental results show tiny decline in effectiveness but significant speedup against traditional full encoder across various MRC and multi-turn conversational question answering tasks. With the help from simple distillation-like auxiliary losses, Incremental Transformer achieves a speedup of 6.2x, with a mere 2.2 point accuracy reduction in comparison to RoBERTa-Large on SQuADV1.1.
Yuechen Wang, Jiaxin Shi, Wengang Zhou 0001, Qi Tian 0001, Houqiang Li
COLING2
2025 MsRAG: Knowledge Augumented Image Captioning with Object-level Multi-source RAG
abstract
Language-Visual Large Models (LVLMs) have made significant strides in enhancing visual understanding capabilities. However, these models often struggle with knowledge-based visual tasks due to constrains in their pre-training data scope and timeliness. Existing Retrieval-Augmented Generation (RAG) methods can effectively solve the problem but primarily rely on user queries, limiting their applicability in scenarios without explicit language input. To overcome these challenges, we introduce MsRAG, a knowledge-augmented captioning framework designed to effectively retrieve and utilize external real-world knowledge, particularly in the absence of user queries, and perform dense captioning for subjects. MsRAG comprises three key components: (1) Parallel Visual Search Module. It retrieves fine-grained object-level knowledge using both online visual search engines and offline domain-knowledge databases, enhancing the robustness and richness of retrieved information. (2) Prompt Templates Pool. The prompt pool dynamically assigns appropriate prompts based on retrieved information, optimizing LVLMs' ability to leverage relevant data under complex RAG conditions. (3) Visual-RAG Alignment Module, which employs a novel visual prompting method to bridge the modality gap between textual RAG content and corresponding visual objects, enabling precise alignment of visual elements with their text-format RAG content. To validate the effectiveness of MsRAG, we conducted a series of qualitative and quantitative experiments. The evaluation results demonstrate the superiority of MsRAG over other methods.
Yuming Qiao, Yuechen Wang, Dan Meng 0001, Haonan Lu
IJCAI2
2024 Cross-Lingual Transfer for Natural Language Inference via Multilingual Prompt Translator
abstract
Based on multilingual pre-trained models, cross-lingual transfer with prompt learning has shown promising effectiveness, where soft prompt learned in a source language is transferred to target languages for downstream tasks, particularly in the low-resource scenario. To efficiently transfer soft prompt, we propose a novel framework, Multilingual Prompt Translator (MPT), where a multilingual prompt translator is introduced to properly process crucial knowledge embedded in prompt by changing language knowledge while retaining task knowledge. More concretely, we first train prompt in source language and employ translator to translate it into target prompt. Besides, we extend an external corpus as auxiliary data, on which an alignment task for predicted answer probability is designed to convert language knowledge, thereby equipping target prompt with multilingual knowledge. In few-shot settings on XNLI, MPT demonstrates superiority over baselines by remarkable improvements. MPT is more prominent compared with vanilla prompting when transferring to languages quite distinct from source language. Code is available at https://github.com/qiuxiaoyu9954/MPT.
Xiaoyu Qiu, Yuechen Wang, Jiaxin Shi, Wengang Zhou 0001, Houqiang Li
ICME2
2024 Progressive Multi-modal Conditional Prompt Tuning
abstract
Pre-trained vision-language models (VLMs) have shown remarkable generalization capabilities via prompting, which leverages VLMs as knowledge bases to extract information beneficial for downstream tasks. However, existing methods primarily employ uni-modal prompting, which only engages a uni-modal branch, failing to simultaneously adjust vision-language (V-L) features. Additionally, the one-pass forward pipeline in VLM encoding struggles to align V-L features that have a huge gap. Confronting these challenges, we propose a novel method, Progressive Multi-modal conditional Prompt Tuning (ProMPT). ProMPT exploits a recurrent structure, optimizing and aligning V-L features by iteratively utilizing image and current encoding information. It comprises an initialization and a multi-modal iterative evolution (MIE) module. Initialization is responsible for encoding images and text using a VLM, followed by a feature filter that selects text features similar to image. MIE then facilitates multi-modal prompting through class-conditional vision prompting, instance-conditional text prompting, and feature filtering. In each MIE iteration, vision prompts are obtained from filtered text features via a vision generator, promoting image features to focus more on target object during vision prompting. The encoded image features are fed into a text generator to produce text prompts that are more robust to class shifts. Thus, V-L features are progressively aligned, enabling advance from coarse to exact prediction. Extensive experiments are conducted in three settings to evaluate the efficacy of ProMPT. The results indicate that ProMPT outperforms existing methods on average across all settings, demonstrating its superior generalization and robustness. Code is available at https://github.com/qiuxiaoyu9954/ProMPT.
Xiaoyu Qiu, Hao Feng 0009, Yuechen Wang, Wengang Zhou 0001, Houqiang Li
ICMR3
2023 Text-Only Training for Visual Storytelling
abstract
Visual storytelling aims to generate a narrative based on a sequence of images, necessitating both vision-language alignment and coherent story generation. Most existing solutions predominantly depend on paired image-text training data, which can be costly to collect and challenging to scale. To address this, we formulate visual storytelling as a visual-conditioned story generation problem and propose a text-only training method that separates the learning of cross-modality alignment and story generation. Our approach specifically leverages the cross-modality pre-trained CLIP model to integrate visual control into a story generator, trained exclusively on text data. Moreover, we devise a training-free visual condition planner that accounts for the temporal structure of the input image sequence while balancing global and local visual content. The distinctive advantage of requiring only text data for training enables our method to learn from external text story data, enhancing the generalization capability of visual storytelling. We conduct extensive experiments on the VIST benchmark, showcasing the effectiveness of our approach in both in-domain and cross-domain settings. Further evaluations on expression diversity and human assessment underscore the superiority of our method in terms of informativeness and robustness.
Yuechen Wang, Wengang Zhou 0001, Zhenbo Lu, Houqiang Li
ACM Multimedia1
2022 Geometric Representation Learning for Document Image Rectification
Hao Feng 0009, Wengang Zhou 0001, Jiajun Deng, Yuechen Wang, Houqiang Li
ECCV (37)4
2022 Weakly Supervised Temporal Adjacent Network for Language Grounding
abstract
Temporal language grounding (TLG) is a fundamental and challenging problem for vision and language understanding. Existing methods mainly focus on fully supervised setting with temporal boundary labels for training, which, however, suffers expensive cost of annotation. In this work, we are dedicated to weakly supervised TLG, where multiple description sentences are given to an untrimmed video without temporal boundary labels. In this task, it is critical to learn a strong cross-modal semantic alignment between sentence semantics and visual content. To this end, we introduce a novel weakly supervised temporal adjacent network (WSTAN) for temporal language grounding. Specifically, WSTAN learns cross-modal semantic alignment by exploiting temporal adjacent network in a multiple instance learning (MIL) paradigm, with a whole description paragraph as input. Moreover, we integrate a complementary branch into the framework, which explicitly refines the predictions with pseudo supervision from the MIL stage. An additional self-discriminating loss is devised on both the MIL branch and the complementary branch, aiming to enhance semantic discrimination by self-supervising. Extensive experiments are conducted on three widely used benchmark datasets,i.e., ActivityNet-Captions, Charades-STA, and DiDeMo, and the results demonstrate the effectiveness of our approach.
Yuechen Wang, Jiajun Deng, Wengang Zhou 0001, Houqiang Li
IEEE Trans. Multim.1
2021 SignBERT: Pre-Training of Hand-Model-Aware Representation for Sign Language Recognition
abstract
Hand gesture serves as a critical role in sign language. Current deep-learning-based sign language recognition (SLR) methods may suffer insufficient interpretability and overfitting due to limited sign data sources. In this paper, we introduce the first self-supervised pre-trainable SignBERT with incorporated hand prior for SLR. Sign-BERT views the hand pose as a visual token, which is derived from an off-the-shelf pose extractor. The visual tokens are then embedded with gesture state, temporal and hand chirality information. To take full advantage of available sign data sources, SignBERT first performs self-supervised pre-training by masking and reconstructing visual tokens. Jointly with several mask modeling strategies, we attempt to incorporate hand prior in a model-aware method to better model hierarchical context over the hand sequence. Then with the prediction head added, SignBERT is fine-tuned to perform the downstream SLR task. To validate the effectiveness of our method on SLR, we perform extensive experiments on four public benchmark datasets, i.e., NMFs-CSL, SLR500, MSASL and WLASL. Experiment results demonstrate the effectiveness of both self-supervised learning and imported hand prior. Furthermore, we achieve state-of-the-art performance on all benchmarks with a notable gain.
Hezhen Hu, Weichao Zhao, Wengang Zhou 0001, Yuechen Wang, Houqiang Li
ICCV4
2021 DocTr: Document Image Transformer for Geometric Unwarping and Illumination Correction
abstract
In this work, we propose a new framework, called Document Image Transformer (DocTr), to address the issue of geometry and illumination distortion of the document images. Specifically, DocTr consists of a geometric unwarping transformer and an illumination correction transformer. By setting a set of learned query embedding, the geometric unwarping transformer captures the global context of the document image by self-attention mechanism and decodes the pixel-wise displacement solution to correct the geometric distortion. After geometric unwarping, our illumination correction transformer further removes the shading artifacts to improve the visual quality and OCR accuracy. Extensive evaluations are conducted on several datasets, and superior results are reported against the state-of-the-art methods. Remarkably, our DocTr achieves $20.02%$ Character Error Rate (CER), a $15%$ absolute improvement over the state-of-the-art methods. Moreover, it also shows high efficiency on running time and parameter count.
Hao Feng 0009, Yuechen Wang, Wengang Zhou 0001, Jiajun Deng, Houqiang Li
ACM Multimedia2
2019 Asking Clarification Questions in Knowledge-Based Question Answering
abstract
Jingjing Xu, Yuechen Wang, Duyu Tang, Nan Duan, Pengcheng Yang, Qi Zeng, Ming Zhou, Xu Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Jingjing Xu 0001, Yuechen Wang, Duyu Tang, Nan Duan 0001, Qi Zeng 0001, Ming Zhou 0001, Xu Sun 0001
EMNLP/IJCNLP (1)2
2019 Deep Self-Paced Learning for Semi-Supervised Person Re-Identification Using Multi-View Self-Paced Clustering
abstract
Semi-supervised person re-identification (Re-ID) is an extension of the existing popular Re-ID research, which only uses a small portion of labeled data, while the majority of the training samples are unlabeled. This paper approaches the problem by constructing a set of heterogeneous convolutional neural networks (CNNs) fine-tuned by utilizing the labeled training samples, and then propagating the labels to the unlabeled portion for further fine-tuning the overall system in a self-paced manner. In this work, a novel self-paced multi-view clustering is presented to generate pseudo labels for unlabeled training samples, which combines multiple heterogeneous CNNs features to cluster. In our clustering method, we introduce a self-paced regularizer to select reliable samples for fine-tuning each CNNs by minimizing ranking loss and identification loss. Specifically, we select a small portion of unlabeled training data when multiple CNNs are weak. With CNNs become stronger, more and more unlabeled samples are selected. Pseudo label estimation and CNNs training are improved simultaneously, which optimize alternatively until all the unlabeled training samples are selected. In our framework, both the optimization of multiple CNNs training and multi-view clustering on unlabeled training samples are self-paced optimizing procedure. Extensive experiments have been conducted on two large-scale Re-ID datasets to demonstrate the superiority of the proposed method.
Xiaomeng Xin, Xindi Wu, Yuechen Wang, Jinjun Wang
ICIP3
2017 Wind turbine gearbox condition monitoring based on extreme gradient boosting
abstract
Currently, supervisory control and data acquisition (SCADA) systems are deployed in most wind farms, with the low cost of data acquisition. However, SCADA data contains a lot of redundant and dirty information, which makes it difficult to observe the actual condition of Wind Turbine (WT) or WT's components directly. In this paper, a gearbox condition monitoring (CM) model of WT based on SCADA data is proposed. In the first stage, after data preprocessing, the prediction model of gearbox oil temperature is obtained based on healthy data with Extreme Gradient Boosting (XGBoost), and the absolute percentage error (APE) of oil temperature is the final observational variable. In the second stage, the Multivariate Quality Control Charts (MQCC) is used to generate the threshold to detect the fault symptoms, combined with the APE of healthy data. Afterwards, the CM framework is established, which is capable of identifying the abnormal state of gearboxes based on whether the APE of new data exceeds the threshold. Finally, the effectiveness of the gearbox CM model presented is demonstrated by examining 2 groups of WT from different wind farms in China.
Yuechen Wang, Zhiliang Zhu 0001
IECON1