Xingjiao Wu

dblp:232/1759 · DBLP profile ↗
← Back
54ranked-venue papers
9as first author
45since 2021 · last 2026
0000-0001-9146-051XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 3 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 3 first-author · 23 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MARS: Multimodal-Assisted Refined Semantic Alignment
abstract
Audio-to-image generation (AIG) faces challenges in fine-grained semantic alignment, particularly semantic semantic misalignment, and loss of visual detail. To address these issues, we proposed MARS ( M ultimodal- A ssisted R efined S emantic alignment), a novel framework leveraging a Mamba-based audio encoder to manage the complexity of long audio sequences, coupled with a fine-grained multimodal alignment strategy using visual descriptions from multimodal large language models. We enhanced semantic coherence and aesthetic quality by fine-tuning an image generator using an image aesthetic perception generator. Furthermore, we validated MARS on VGGSound and VEGAS benchmarks, comprising 37,250 and 9,500 records, respectively. The results suggest that MARS significantly outperforms existing methods, achieving average improvements of 28.73% in semantic relevance and 127.35% in aesthetic scores compared with the best AIG generation baseline. In addition, cross-domain evaluations on the AudioCaps and Clotho datasets confirmed the robustness and generalization capability of MARS , with an average improvement of 73.9% on the V2A metric. • MARS refines semantic alignment for audio-based image generation. • A Mamba-based encoder processes long audio sequences. • MLLMs supply rich visuals to boost semantics and aesthetics. • Extensive experiments prove effectiveness and robustness.
Xingjiao Wu, Tianlong Ma, Daoguo Dong, Liang He 0001
Inf. Process. Manag.2
2026 Evidence-chain-driven multimodal retrieval question answering
Anran Wu, Xingjiao Wu, Jiabao Zhao, Liang He 0001
Knowl. Based Syst.3
2026 ACRA: An adaptive chain retrieval architecture for multi-modal knowledge-Augmented visual question answering
Xingjiao Wu, Jiabao Zhao, Qin Chen 0001, Jing Yang 0023, Liang He 0001
Knowl. Based Syst.3
2025 An Exemplar-based Framework for Chinese Text Recognition
abstract
This paper introduces a novel exemplar-based framework for reading Chinese texts in natural scene or document images. We present the Deep Exemplar-based Chinese Text Recognizer, which is structured to first identify candidate characters as exemplars from each text-line, and subsequently recognize them by retrieving analogous exemplars from a database. With text-line level annotations, we design the exemplar discovery network to simultaneously recognize texts and capture individual character positions in a weak-supervision manner. The exemplar retrieval module is then crafted to identify the most similar exemplar and propagate the corresponding character label. This enables us to effectively rectify the misrecognized characters and boost the performance of scene text recognition. Experiments on four scenarios of Chinese texts demonstrate the effectiveness of our proposed framework.
Zhao Zhou, Xiangcheng Du, Yingbin Zheng, Xingjiao Wu, Cheng Jin 0001
AAAI4
2025 CL-MoE: Enhancing Multimodal Large Language Model with Dual Momentum Mixture-of-Experts for Continual Visual Question Answering
abstract
Multimodal large language models (MLLMs) have garnered widespread attention from researchers due to their remarkable understanding and generation capabilities in visual language tasks (e.g., visual question answering). However, the rapid pace of knowledge updates in the real world makes offline training of MLLMs costly, and when faced with non-stationary data streams, MLLMs suffer from catastrophic forgetting during learning. In this paper, we propose an MLLMs-based dual momentum Mixture-of-Experts (CL-MoE) framework for continual visual question answering (VQA). We integrate MLLMs with continual learning to utilize the rich commonsense knowledge in LLMs. We introduce a Dual-Router MoE (RMoE) strategy to select the global and local experts using task-level and instance-level routers, to robustly assign weights to the experts most appropriate for the task. Then, we design a dynamic Momentum MoE (MMoE) to update the parameters of experts dynamically based on the relationships between the experts and tasks/instances, so that the model can absorb new knowledge while maintaining existing knowledge. The extensive experimental results indicate that our method achieves state-of-the-art performance on 10 VQA tasks, proving the effectiveness of our approach.
Tianyu Huai, Jie Zhou 0015, Xingjiao Wu, Qin Chen 0001, Qingchun Bai, Liang He 0001
CVPR3
2025 Lark: Low-Rank Updates After Knowledge Localization for Few-Shot Class-Incremental Learning
Jinxin Shi, Jiabao Zhao, Yifan Yang 0001, Xingjiao Wu, Liang He 0001
ICCV4
2025 Multi-Type Preference Learning: Empowering Preference-Based Reinforcement Learning with Equal Preferences
abstract
Preference-Based reinforcement learning (PBRL) learns directly from the preferences of human teachers regarding agent behaviors without needing meticulously designed reward functions. However, existing PBRL methods often learn primarily from explicit preferences, neglecting the possibility that teachers may choose equal preferences. This neglect may hinder the understanding of the agent regarding the task perspective of the teacher, leading to the loss of important information. To address this issue, we introduce the Equal Preference Learning Task, which optimizes the neural network by promoting similar reward predictions when the behaviors of two agents are labeled as equal preferences. Building on this task, we propose a novel PBRL method, Multi-Type Preference Learning (MTPL), which allows simultaneous learning from equal preferences while leveraging existing methods for learning from explicit preferences. To validate our approach, we design experiments applying MTPL to four existing state-of-the-art baselines across ten locomotion and robotic manipulation tasks in the DeepMind Control Suite. The experimental results indicate that simultaneous learning from both equal and explicit preferences enables the PBRL method to more comprehensively understand the feedback from teachers, thereby enhancing feedback efficiency. Project page: https://github.com/FeiCuiLengMMbb/paper_MTPL
Ziang Liu 0019, Xingjiao Wu, Jing Yang 0023, Liang He 0001
ICRA3
2025 Unleashing the Semantic Adaptability of Controlled Diffusion Model for Image Colorization
abstract
Recent data-driven image colorization methods have leveraged pre-trained Text-to-Image (T2I) diffusion models as generative prior, while still suffering from unsatisfactory and inaccurate semantic-level color control. To address these issues, we propose a Semantic Adaptation method (SeAda) that enhances the prior while considering the semantic discrepancy between color and grayscale image pairs. The SeAda employs a semantic adapter to produce refined semantic embeddings and a controlled T2I diffusion model to create reasonably colored images. Specifically, the semantic adapter transfers the embedding from grayscale to color domain, while the diffusion model utilizes the refined embedding and prior knowledge to achieve realistic and diverse results. We also design a three-staged training strategy to improve semantic comprehension and prior integration for further performance improvement. Extensive experiments on public datasets demonstrate that our method outperforms existing state-of-the-art techniques, yielding superior performance in image colorization.
Xiangcheng Du, Zhao Zhou, Yingbin Zheng, Xingjiao Wu, Peizhu Gong, Cheng Jin 0001
IJCAI5
2025 FLIP: Adaptive Comparison Method Selection for Efficient Preference-Based Reinforcement Learning
abstract
Preference-based Reinforcement Learning (PBRL) relies on the efficient collection and use of preference data to train accurate reward functions, enabling agents to learn directly from human preferences. This process allows agents to better understand human intentions while effectively reducing biases inherent in AI systems. The pairwise comparison method gathers diverse preference data, and Seqrank expands preference datasets through transitivity, both fail to establish preference relationships across different rounds of labeling. This limitation can result in fragmented signals and slow convergence toward the optimal policy. To address this, we propose the Global Tree (GTree), a method built on the Seqrank framework that integrates trajectory preferences across multiple rounds, providing a unified representation of global preferences. Moreover, we posit that different trajectory comparison methods offer distinct advantages depending on the task and the stage of training. To fully exploit these strengths, we introduce FLIP. This adaptive strategy dynamically selects either the pairwise method or GTree based on historical performance, optimizing method use for each task and training stage. Our evaluations demonstrate that integrating cross-round preferences accelerates the convergence of the reward function, while the FLIP strategy further enhances learning efficiency and overall performance, thereby enabling agents to better understand human intentions.
Ziang Liu 0019, Xingjiao Wu, Hongxin Chen, Luwei Xiao, Jing Yang 0023
IJCNN2
2025 McGE '25: The 3rd International Workshop on Multimedia Content Generation and Evaluation: New Methods and Practice
abstract
This workshop addresses next-generation methods in multimedia research, with a focus on content generation, quality assessment, and dataset development. These three areas are foundational for advancing multimedia technologies and applications. Emerging approaches in multimedia content generation, powered by generative AI and multimodal learning, are reshaping domains such as entertainment, advertising, education, and healthcare. At the same time, robust quality assessment is essential to ensure that generated content achieves high standards of perceptual fidelity, semantic consistency, and user satisfaction, thereby determining the real-world impact of multimedia systems. Datasets remain indispensable for training and evaluating algorithms, and innovative strategies in dataset construction-ranging from augmentation and annotation to addressing issues of bias and small-sample imbalance-are driving the development of more reliable and ethical multimedia applications. By convening leading researchers and practitioners, this workshop provides a platform to explore state-of-the-art methods, share best practices, and discuss open challenges in next-generation multimedia research. The goal is to foster interdisciplinary collaboration and inspire innovative solutions that advance the creation, evaluation, and application of multimedia content, setting new benchmarks for the field and shaping the future of multimedia technologies.
Cheng Jin 0001, Mingli Song, Rui Wang 0032, Xingjiao Wu
ACM Multimedia4
2025 Temporal-Conditioned Symbolic Alignment for Controllable Text-to-Music Generation
abstract
In recent years, Text-to-Music (T2M) generation models have rapidly emerged as powerful tools in content creation across fields. While existing models have made notable progress in sound quality, instrument identification, and stylistic alignment, they still exhibit clear limitations in modeling musical structure and musicality-particularly in terms of harmonic coherence and rhythmic alignment. To address these issues, we propose a Temporal-Conditioned Symbolic Alignment for Controllable Text-to-Music Generation(TCSA), which introduces explicit local condition controls to enhance structural fidelity in music generation. Specifically, we design a music theory enrichment strategy based on GPT-2 that transforms input text into detailed descriptions with embedded music theory knowledge, from which accurate chord progressions and rhythmic patterns are extracted as generation conditions. To synchronize these local features effectively, we develop a temporal alignment feature fusion mechanism. Additionally, we propose a layer-skipping fine-tuning strategy to avoid overfitting and enable fine-grained structural modeling. Finally, we introduce a perception-driven loss function based on Mel spectrograms to optimize the harmonic consistency and structural coherence of the generated music. Experimental results demonstrate that TCSA achieves competitive generation quality while offering significantly improved controllability over musical structure, making it well-suited for professional music production and refined content creation.
Xingjiao Wu, Tianlong Ma, Tangren Yao, Wen Wu 0006, Liang He 0001
ACM Multimedia2
2025 Efficiency is the rule: Domain adaptive semantic segmentation with minimal annotations
Tianyu Huai, Junhang Zhang, Xingjiao Wu, Liang He 0001
Expert Syst. Appl.3
2025 Mitigating reasoning hallucination through Multi-agent Collaborative Filtering
Jinxin Shi, Jiabao Zhao, Xingjiao Wu, Ruyi Xu, Liang He 0001
Expert Syst. Appl.3
2025 A lightweight depth completion network with spatial efficient fusion
Zhichao Fu, Anran Wu, Zisong Zhuang, Xingjiao Wu
Image Vis. Comput.4
2024 Chinese Elderly Healthcare-Oriented Conversation: CareQA Dataset and Its Knowledge Distillation Based Generation Framework
abstract
The increasing global aging brings the substantial demand for healthcare knowledge among the elderly. Large Language Models (LLMs) based Conversation Agents (CAs) hold significant promise for addressing the elderly’s healthcare knowledge inquiries. Yet, general LLMs often fall short in providing professional and practically usable healthcare conversations due to the lack of specific knowledge, possible hallucination issues and contextual comprehension biases. To address these challenges, we first propose a cost-effective, domain-specific questioning-answering (QA) generation framework based on knowledge distillation (KD). Based on this framework, we then built CareQA, the first Chinese healthcare QA dataset specifically for the elderly, with 41,694 QA pairs spanning geriatric diseases covering multiple categories. A comprehensive benchmarking experiment, including both automated and human evaluation, is conducted to examine the usability of CareQA. The results demonstrate that the LLMs fine-tuned on CareQA perform better in answering elderly healthcare-related questions.
Xingjiao Wu, Jialiang Tong, Bangyan Li, Yuling Sun
BIBM2
2024 MindScope: Exploring Cognitive Biases in Large Language Models Through Multi-Agent Systems
abstract
Detecting cognitive biases in large language models (LLMs) is a fascinating task that aims to probe the existing cognitive biases within these models. Current methods for detecting cognitive biases in language models generally suffer from incomplete detection capabilities and a restricted range of detectable bias types. To address this issue, we introduced the ‘MindScope’ dataset, which distinctively integrates static and dynamic elements. The static component comprises 5,170 open-ended questions spanning 72 cognitive bias categories. The dynamic component leverages a rule-based, multi-agent communication framework to facilitate the generation of multi-round dialogues. This framework is flexible and readily adaptable for various psychological experiments involving LLMs. In addition, we introduce a multi-agent detection method applicable to a wide range of detection tasks, which integrates Retrieval-Augmented Generation (RAG), competitive debate, and a reinforcement learning-based decision module. Demonstrating substantial effectiveness, this method has shown to improve detection accuracy by as much as 35.10% compared to GPT-4. Codes and appendix are available at https://github.com/2279072142/MindScope.
Zhentao Xie, Jiabao Zhao, Jinxin Shi, Yanhong Bai, Xingjiao Wu, Liang He 0001
ECAI6
2024 Artistry in Pixels: FVS - A Framework for Evaluating Visual Elegance and Sentiment Resonance in Generated Images
abstract
The field of image generation models has seen substantial progress, characterized by a proliferation of diverse generative models and their associated outputs. However, there currently exists a deficiency in methodologies that can concurrently and effectively evaluate both the intrinsic quality of generated images and the alignment between image features and textual prompts. To address these challenges, we propose a novel Framework for evaluating Visual elegance and Sentiment resonance (FVS). The FVS incorporates a novel image aesthetic assessment model, specifically trained to assess the visual attractiveness of the generated images. Additionally, it evaluates the sentiment and aesthetic consistency between textual prompt and the generated image. Experimental results verify that the evaluations from our framework align more closely with human preferences. Moreover, we apply our framework to filter and construct a higher-quality training set of generated images. This curated dataset is then exploited to adapt the generative model, resulting in enhanced generation quality.
Luwei Xiao, Xingjiao Wu, Tianlong Ma, Jiabao Zhao, Liang He 0001
ICME3
2024 Fine-Grained Scene Image Classification with Modality-Agnostic Adapter
abstract
When dealing with the task of fine-grained scene image classification, most previous works lay much emphasis on global visual features when doing multi-modal feature fusion. In other words, models are deliberately designed based on prior intuitions about the importance of different modalities. In this paper, we present a new multi-modal feature fusion approach named MAA (Modality-Agnostic Adapter), trying to make the model learn the importance of different modalities in different cases adaptively, without giving a prior setting in the model architecture. More specifically, we eliminate the modal differences in distribution and then use a modality-agnostic Transformer encoder for a semantic-level feature fusion. Our experiments demonstrate that MAA achieves state-of-the-art results on benchmarks by applying the same modalities with previous methods. Besides, it is worth mentioning that new modalities can be easily added when using MAA and further boost the performance.
Zhao Zhou, Xiangcheng Du, Xingjiao Wu, Yingbin Zheng, Cheng Jin 0001
ICME4
2024 Enhancing Out-of-Distribution Generalization in VQA through Gini Impurity-guided Adaptive Margin Loss
abstract
In the Visual Question Answering (VQA) task context, most methods are influenced by language bias, resulting in poor performance on out-of-distribution data. Recently, some works attempted to use the adaptive margin loss to address this bias issue. However, these works typically consider only the frequency of answer labels when designing margin loss, leading to some samples being overly emphasized or lacking sufficient attention during model training. To address this issue, we propose a novel margin loss guided by the Gini-impurity for VQA debiasing. By comprehensively considering label distribution and instance complexity, we use Gini impurity to adjust the margin values in margin loss, balancing the attention of the model to different samples. Importantly, our method is plug-and-play and can be directly applied to any baseline. In the VQA-CP v2 task, our evaluation results across various baselines surpass the current state-of-the-art methods.
Tianyu Huai, Anran Wu, Xingjiao Wu, Wenxin Hu, Liang He 0001
ICME4
2024 MultiColor: Image Colorization by Learning from Multiple Color Spaces
Xiangcheng Du, Zhao Zhou, Xingjiao Wu, Yingbin Zheng, Cheng Jin 0001
ACM Multimedia3
2024 Charting the Uncharted: Building and Analyzing a Multifaceted Chart Question Answering Dataset for Complex Logical Reasoning Process
Anran Wu, Yujia Xia, Xingjiao Wu, Tianlong Ma, Liang He 0001
PRCV (5)4
2024 Simple contrastive learning in a self-supervised manner for robust visual question answering
Luwei Xiao, Xingjiao Wu, Liang He 0001
Comput. Vis. Image Underst.3
2024 Cross-domain document layout analysis using document style guide
Xingjiao Wu, Luwei Xiao, Xiangcheng Du, Yingbin Zheng, Xin Li 0110, Tianlong Ma, Cheng Jin 0001, Liang He 0001
Expert Syst. Appl.1
2024 K-NNDP: K-means algorithm based on nearest neighbor density peak optimization and outlier removal
Jiyong Liao, Xingjiao Wu, Yaxin Wu, Juelin Shu
Knowl. Based Syst.2
2023 A Dynamic Composite Ensemble Learning Framework for Multi-Stage Dementia Prediction
abstract
Dementia has increasingly impacted the health of older adults, necessitating precise prediction of cognitive levels to deliver targeted healthcare services and mitigate disease progression. In recent years, although there has been surging interest in research on AI-driven dementia prediction, existing methods commonly rely on background-specific datasets and utilize multi-models with equal weights for predictions. These methods limit the result to the inherent features of the datasets on the one hand and ignore the differences among various feature channels on the other hand. To address these limitations, this paper proposes a dynamic composite ensemble learning framework, including the following three novel modules. Firstly, based on the selected common feature set, we construct a top-level feature set through dynamically determined screening thresholds. Subsequently, we deploy transfer learning models for knowledge complementation to enhance model generalization. In addition, we purposefully designed a dynamic weighting module to prioritize feature channels with stronger relevance to the target task. We deploy our model on the ELSA-HCAP dataset and conduct a series of experiments to evaluate its practical effectiveness. The results demonstrate that the proposed feature engineering, dynamic weighting, and knowledge transfer modules collectively enhance the overall model performance. Moreover, the combination of these three modules in the fusion model attains optimal performance, achieving an accuracy rate of 95.76%.
Bangyan Li, Junyan Mao, Xingjiao Wu, Yuling Sun, Liang He 0001
BIBM3
2023 LoGoNet: Towards Accurate 3D Object Detection with Local-to-Global Cross- Modal Fusion
abstract
LiDAR-camera fusion methods have shown impressive performance in 3D object detection. Recent advanced multi-modal methods mainly perform global fusion, where image features and point cloud features are fused across the whole scene. Such practice lacks fine-grained region-level information, yielding suboptimal fusion performance. In this paper, we present the novel Local-to-Global fusion network (LoGoNet), which performs LiDAR-camerafusion at both local and global levels. Concretely, the Global Fusion (GoF) of LoGoNet is built upon previous literature, while we exclusively use point centroids to more precisely represent the position of voxel features, thus achieving better crossmodal alignment. As to the Local Fusion (LoF), we first divide each proposal into uniform grids and then project these grid centers to the images. The image features around the projected grid points are sampled to be fused with position-decorated point cloud features, maximally uti-lizing the rich contextual information around the proposals. The Feature Dynamic Aggregation (FDA) module is further proposed to achieve information interaction between these locally and globally fused features, thus producing more informative multi-modal features. Extensive experiments on both Waymo Open Dataset (WOD) and KITTI datasets show that LoGoNet outperforms all state-of-the-art 3D detection methods. Notably, LoGoNet ranks 1st on Waymo 3D object detection leaderboard and obtains 81.02 mAPH (L2) detection performance. It is noteworthy that, for the first time, the detection performance on three classes surpasses 80 APH (L2) simultaneously. Code will be available at https://github.com/sankin97/LoGoNet.
Xin Li 0110, Tao Ma 0002, Yuenan Hou, Botian Shi, Yuchen Yang 0003, Youquan Liu, Xingjiao Wu, Qin Chen 0001, Yikang Li 0002, Yu Qiao 0001, Liang He 0001
CVPR7
2023 DDT: Dual-branch Deformable Transformer for Image Denoising
abstract
Transformer is beneficial for image denoising tasks since it can model long-range dependencies to overcome the limitations presented by inductive convolutional biases. However, directly applying the transformer structure to remove noise is challenging because its complexity grows quadratically with the spatial resolution. In this paper, we propose an efficient Dual-branch Deformable Transformer (DDT) denoising network which captures both local and global interactions in parallel. We divide features with a fixed patch size and a fixed number of patches in local and global branches, respectively. In addition, we apply deformable attention operation in both branches, which helps the network focus on more important regions and further reduces computational complexity. We conduct extensive experiments on real-world and synthetic denoising tasks, and the proposed DDT achieves state-of-the-art performance with significantly fewer computational costs.
Kangliang Liu, Xiangcheng Du, Yingbin Zheng, Xingjiao Wu, Cheng Jin 0001
ICME5
2023 Image Layer Modeling for Complex Document Layout Generation
abstract
Document layout analysis (DLA) plays an essential role in information extraction and document understanding. At present, DLA has reached the milestone achievement; however, DLA of non-Manhattan is still challenging because of annotation data limitations. In this paper, we propose an image layer modeling method to mitigate this issue. The image layer modeling method generates document images of non-Manhattan layouts by superimposing images under pre-defined aesthetic rules. Due to the lack of evaluation benchmark for non-Manhattan layout, we have constructed a manually-labeled non-Manhattan layout fine-grained segmentation dataset. To the best of our knowledge, this is the first manually-labeled non-Manhattan layout fine-grained segmentation dataset. Extensive experimental results verify that our proposed image layer modeling method can better deal with the fine-grained segmented document of the non-Manhattan layout.
Tianlong Ma, Xingjiao Wu, Xiangcheng Du, Cheng Jin 0001
ICME2
2023 Dual-Expert Distillation Network for Few-Shot Segmentation
abstract
Few-shot segmentation has attracted growing interest owing to its value in practical applications. The primary challenge of few-shot segmentation lies in semantic information discovery, especially for query images. To tackle this issue, we propose a dual-expert distillation network (DEDN) made up of a scenario-level expert and an object-level expert to obtain semantic information from different perspectives. In DEDN, experts can learn from each other through online knowledge distillation with positive-guided Kullback-Leibler divergence. We innovate the Scenario Normalization and Object Continuity Guidance on dual experts to guarantee the various perspectives respectively. We further propose the Adaptive Weighted Fusion to adapt the trained experts to novel classes and obtain reliable fused predictions. Extensive experiments on Pascal-5i and COCO-20i show that our approach achieves state-of-the-art results.
Junhang Zhang, Zisong Zhuang, Luwei Xiao, Xingjiao Wu, Tianlong Ma, Liang He 0001
ICME4
2023 Modeling Stroke Mask for End-to-End Text Erasing
abstract
Scene text erasing aims to wipe text regions in scene images with reasonable background. Most previous approaches employ scene text detectors to assist localization of the text regions. However, detected text boxes contain both text strokes and background clutters, and directly in-painting on the whole boxes may remain text artifacts and make regions unnatural. In this paper, we present an end-to-end network that focuses on modeling text stroke masks that provide more accurate locations to compute erased images. The network consists of two stages, i.e., a basic network with stroke generation and a refinement network with stroke awareness. The basic network predicts the text stroke masks and initial erasing results simultaneously. The refinement network receives the masks as supervision to generate natural erased results. Experiments on both synthetic and real-world scene images demonstrate the effectiveness of our framework in producing high quality erasing results.
Xiangcheng Du, Zhao Zhou, Yingbin Zheng, Tianlong Ma, Xingjiao Wu, Cheng Jin 0001
WACV5
2023 Progressive scene text erasing with self-supervision
Xiangcheng Du, Zhao Zhou, Yingbin Zheng, Xingjiao Wu, Tianlong Ma, Cheng Jin 0001
Comput. Vis. Image Underst.4
2023 DRFN: A unified framework for complex document layout analysis
Xingjiao Wu, Tianlong Ma, Xiangcheng Du, Ziling Hu, Jing Yang 0023, Liang He 0001
Inf. Process. Manag.1
2023 Cross-modal fine-grained alignment and fusion network for multimodal aspect-based sentiment analysis
Luwei Xiao, Xingjiao Wu, Jie Zhou 0015, Liang He 0001
Inf. Process. Manag.2
2023 Reading Scene Text with Aggregated Temporal Convolutional Encoder
abstract
Reading scene text in the natural image is of fundamental importance in many real-world problems. Text recognition has a profound effect on information processing by enabling automated extraction and interpretation. Recent scene text recognition methods employ the encoder-decoder framework, which constructs the encoder by obtaining the visual representations based on the last layer of the backbone network and then feeding them into a sequence model. In this article, we propose a novel encoder structure that performs the feature extractor and the sequence modeling within a unified framework. The introduced Aggregated Temporal Convolutional Encoder (ATCE) first incorporates the temporal convolutional layers to consider the long-term temporal relationship in the encoder stage. The aggregation of these temporal convolution modules is designed to utilize visual features from different levels, by augmenting the standard architecture with deeper aggregation to better fuse information across modules. We also study the impact of different attention modules in convolutional blocks for learning accurate text representations. We conduct comparisons on several scene text recognition benchmarks for both Chinese and English; the experiments demonstrate the complementary ability with different decoder variants and the effectiveness of our proposed approach.
Tianlong Ma, Xiangcheng Du, Xingjiao Wu, Zhao Zhou, Yingbin Zheng, Cheng Jin 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2022 Lightweight Network Based Real-time Anomaly Detection Method for Caregiving at Home
abstract
Using data-driven technologies to support the healthcare of the elderly has been largely celebrated as an effective means. This paper focuses on the issue of using video-based sensing technologies to remotely monitor the activities and conditions of the elderly. Although it is a widely explored field, the high cost and high infrastructural requirements of most existing technologies usually challenge their effectiveness and efficiency in practical caregiving context. To address these challenges, we propose a lightweight network based real-time anomaly detection system, which consists of video-based ADL sensing and pre-processing, AI streaming aggregating and cluster computing. We examine our method by implementing and deploying it into a real-world care facility for the elderly in Shanghai China. The results show that our method has good performance in expansibility, reliability, bandwidth availability, accuracy and privacy protection.
Xingjiao Wu, Miaomiao Gong, Jin Zhao 0001, Yuling Sun
CSCWD2
2022 Homogeneous Multi-modal Feature Fusion and Interaction for 3D Object Detection
Xin Li 0110, Botian Shi, Yuenan Hou, Xingjiao Wu, Tianlong Ma, Yikang Li 0002, Liang He 0001
ECCV (38)4
2022 Multi-Channel Attentive Graph Convolutional Network with Sentiment Fusion for Multimodal Sentiment Analysis
abstract
Nowadays, with the explosive growth of multimodal reviews on social media platforms, multimodal sentiment analysis has recently gained popularity because of its high relevance to these social media posts. Although most previous studies design various fusion frameworks for learning an interactive representation of multiple modalities, they fail to incorporate sentimental knowledge into inter-modality learning. This pa-per proposes a Multi-channel Attentive Graph Convolutional Network (MAGCN), consisting of two main components: cross-modality interactive learning and sentimental feature fusion. For cross-modality interactive learning, we exploit the self-attention mechanism combined with densely connected graph convolutional networks to learn inter-modality dynamics. For sentimental feature fusion, we utilize multi-head self-attention to merge sentimental knowledge into inter-modality feature representations. Extensive experiments are conducted on three widely-used datasets. The experimental results demonstrate that the proposed model achieves competitive performance on accuracy and F1 scores compared to several state-of-the-art approaches.
Luwei Xiao, Xingjiao Wu, Wen Wu 0006, Jing Yang 0023, Liang He 0001
ICASSP2
2022 3D Clues Guided Convolution for Depth Completion
abstract
Depth completion is a task that recovers a dense depth map from a sparse depth map with the corresponding color image. Recently, the intensive depth generation guided by image clues in the color map has achieved good results. Color images can provide structural and semantic information as guidance information, but cannot provide the more important information about geometric relationships. In this paper, we propose a novel network to learn latent 3D cues from RGB images and depth images. More specifically, the network contains a 3D clues extractor and a dense depth generator. The extractor is designed to fusion and extract the 3D joint clues from the color image and sparse depth. The generator is trained with the sparse depth map and 3D clues to producing a more accurate dense depth map. Extensive experiments show that our proposed method has a significant improvement over existing image-guided methods.
Zhichao Fu, Xingjiao Wu, Xiangcheng Du, Tianlong Ma, Liang He 0001
ICIP3
2022 Document Layout Analysis Via Positional Encoding
abstract
Document layout analysis plays a vital role in computer vision research. Current document layout analysis methods mostly use pixel-based classification for document layout analysis. However, the method based on pixel classification is insufficient for maintaining the continuity of the classification area. In this paper, we propose a document layout analysis method based on positional encoding and bounding box specification. We maintain the continuity of the analysis area by constructing a document layout analysis framework based on the bounding box. In addition, we also integrate a positional encoding module in the framework to maintain the detailed information in the document layout analysis and modeling process. Experimental results prove that our proposed method has achieved state-of-the-art results.
Ejian Zhou, Xingjiao Wu, Luwei Xiao, Xiangcheng Du, Tianlong Ma, Liang He 0001
ICIP2
2022 Adaptive Multi-Feature Extraction Graph Convolutional Networks for Multimodal Target Sentiment Analysis
abstract
The multi-modal target-oriented sentiment analysis aims at predicting the sentiment polarities for target entities in a sentence by combining vision and language information. However, most existing deep learning approaches fail to extract valuable information from the visual modality and ignore the usability of syntactic dependency information embedded in the text modality. In this paper, we propose a two-stream adaptive multi-feature extraction graph convolutional networks (AME-GCN), which translates the image into a textual caption and dynamically fuses the semantic and syntactic feature from the given sentence and generated caption to model the inter/intra-modality dynamics. Extensive experiments on two multi-modal Twitter datasets show the effectiveness of the proposed model against popular textual and multi-modal approaches, demonstrating that AME-GCN is a best alternative for this task.
Luwei Xiao, Ejian Zhou, Xingjiao Wu, Tianlong Ma, Liang He 0001
ICME3
2022 Graph Convolution over the Semantic-syntactic Hybrid Graph Enhanced by Affective Knowledge for Aspect-level Sentiment Classification
abstract
Aspect-level sentiment classification (ASC), detecting and predicting the sentiment polarity of the given aspecs, has attracted increasing attention in the field of Natural Language Processing (NLP). Recent studies in ASC leveraged the graph based on the dependency tree of the context to incorporate the syntactic information and structure of a sentence for better relation extraction. Some researchers noted that existing methods ignored semantic relations or failed to consider affective dependency information, and then proposed several state-of-art methods tackling the above two limitations. However, these approaches failed to consider both informative relations simultaneously. Therefore, we explore and propose a novel solution based on semantic latent graph and SenticNet to leverage semantic and affective information. Specifically, we build a latent semantic graph based on self-attention networks to parse semantic relations within the contexts. In addition, we utilize affective knowledge from SenticNet to enhance the dependency graphs of sentences. Moreover, we use the gate mechanism to dynamically combine information from both the enhanced dependency graphs and latent semantic graphs. Experimental results on three benchmark datasets illustrate the effectiveness and state-of-the-art performance of our model.
Luwei Xiao, Zhichao Fu, Xingjiao Wu, Tianlong Ma, Liang He 0001
IJCNN5
2022 A survey of human-in-the-loop for machine learning
Xingjiao Wu, Luwei Xiao, Yixuan Sun, Junhang Zhang, Tianlong Ma, Liang He 0001
Future Gener. Comput. Syst.1
2021 LSTMVAEF: Vivid Layout via LSTM-Based Variational Autoencoder Framework
Xingjiao Wu, Wenxin Hu, Jing Yang 0023
ICDAR (2)2
2021 Document Layout Analysis via Dynamic Residual Feature Fusion
abstract
The document layout analysis (DLA) aims to split the document image into different interest regions and understand the role of each region, which has wide application such as optical character recognition (OCR) systems and document retrieval. However, it is a challenge to build a DLA system because the training data is very limited and lacks an efficient model. In this paper, we propose an end-to-end united network named Dynamic Residual Fusion Network (DRFN) for the DLA task. Specifically, we design a dynamic residual feature fusion module which can fully utilize low-dimensional information and maintain high-dimensional category information. Besides, to deal with the model overfitting problem that is caused by lacking enough data, we propose the dynamic select mechanism for efficient fine-tuning in limited train data. We experiment with two challenging datasets and demonstrate the effectiveness of the proposed module.
Xingjiao Wu, Ziling Hu, Xiangcheng Du, Jing Yang 0023, Liang He 0001
ICME1
2021 Document image layout analysis via explicit edge embedding network
Xingjiao Wu, Yingbin Zheng, Tianlong Ma, Hao Ye 0005, Liang He 0001
Inf. Sci.1
2020 Scene Text Recognition with Temporal Convolutional Encoder
abstract
Texts from scene images typically consist of several characters and exhibit a characteristic sequence structure. Existing methods capture the structure with the sequence-to-sequence models by an encoder to have the visual representations and then a decoder to translate the features into the label sequence. In this paper, we study text recognition framework by considering the long-term temporal dependencies in the encoder stage. We demonstrate that the proposed Temporal Convolutional Encoder with increased sequential extents improves the accuracy of text recognition. We also study the impact of different attention modules in convolutional blocks for learning accurate text representations. We conduct comparisons on seven datasets and the experiments demonstrate the effectiveness of our proposed approach.
Xiangcheng Du, Tianlong Ma, Yingbin Zheng, Hao Ye 0005, Xingjiao Wu, Liang He 0001
ICASSP5
2020 TCATD: Text Contour Attention for Scene Text Detection
abstract
Segmentation-based approaches have enabled state-of-the-art performance in long or curved text detection tasks. However, false detection still is a challenge when two text instances are close to each other. To address this problem, in this paper, we propose a Text Contour Attention Text Detector (TCATD), which can locate scene text with arbitrary orientation and shape accurately. Different from previous work, TCATD focus on text contour map (TC), text center intensity map (TCI) and text kernel maps (TK). The TC can introduce text contour information, the TCI can help to learn the accurate text segmentation and the TK can generate the complete shape of text instances. Besides, we propose a Text Contour Attention Module to deal with contour information. After the Text Contour Attention Module, TC, TCI and TK will be obtained. Extensive experiments on ICDAR2015, CTW1500 and Total-Text demonstrate that the proposed method achieves the state-of-the-art performance.
Ziling Hu, Xingjiao Wu, Jing Yang 0023
ICPR2
2020 CPSPNet: Crowd Counting via Semantic Segmentation Framework
abstract
Crowd counting, i.e., estimation number of the pedestrian in crowd images, is emerging as an essential research problem with the public security applications. The density-based method of crowd counting still has some challenges, such as lack of perspective information in density map and background noise. Current models often misjudge background noise as a person and the ground truth density map widely used now is not so accurate. In this paper, we present a novel approach to help generate a higher quality density map. On the one hand, we eliminate the apparent mistakes in the density map with the help of a semantic segmentation model, which provides more information about fine-granted negative samples. On the other hand, we modify the density map to make sure it maintains a natural attribute. The experimental results prove the effectiveness of our method for crowd counting models, especially in uneven distribution monitoring scenario.
Xingjiao Wu, Jing Yang 0023, Wenxin Hu
ICTAI2
2020 Margin Guidance Network for Arbitrary-shaped Scene Text Detection
abstract
Segmentation-based scene text detection approaches have been adopted to arbitrary-shaped texts and have achieved a great progress. However, false detection always easily exist when the arbitrary-shaped texts are close to each other. In this paper, we propose the Margin Guidance Network (MGN) that mainly based on the margin constraint residual module (MCRM) to address aforementioned problem. The MCRM considers the margins between multiple text instance masks to guide the training of network and improve the performance on text detection. The MCRM contains two prediction branch, the one can generate the multiple different scale of masks for a text instance and the other branch is used to generate multiple margins between the above masks. Experimental results on three public benchmarks including ICDAR2015, CTW1500 and Total-Text have demonstrated that the proposed MGN achieves the state-of-the-art results.
Xin Li 0110, Xingjiao Wu, Tianlong Ma, Zhao Zhou, Luhui Chen, Liang He 0001
ICTAI2
2020 Feature channel enhancement for crowd counting
abstract
Crowd counting, i.e. count the number of people in a crowded visual space, is emerging as an essential research problem with public security. A key in the design of the crowd counting system is to create a stable and accurate robust model, which requires to process on the feature channels of the counting network. In this study, the authors present a featured channel enhancement (FCE) block for crowd counting. First, they use a feature extraction unit to obtain the information of each channel and encodes the information of each channel. Then use a non‐linear variation unit to deal with the encoded channel information, finally, normalise the data and affixed to each channel separately. With the use of the FCE, the positive characteristic channel can be enhanced and weak or negative channel information can be suppressed. The authors successfully incorporate the FCE with two compact networks on the standard benchmarks and prove that the proposed FCE achieves promising results.
Xingjiao Wu, Shuchen Kong, Yingbin Zheng, Hao Ye 0005, Jing Yang 0023, Liang He 0001
IET Image Process.1
2020 Fast video crowd counting with a Temporal Aware Network
Xingjiao Wu, Baohan Xu, Yingbin Zheng, Hao Ye 0005, Jing Yang 0023, Liang He 0001
Neurocomputing1
2020 Counting crowds with varying densities via adaptive scenario discovery framework
Xingjiao Wu, Yingbin Zheng, Hao Ye 0005, Wenxin Hu, Tianlong Ma, Jing Yang 0023, Liang He 0001
Neurocomputing1
2019 Aggregating Rich Deep Semantic Features for Fine-Grained Place Classification
Tingyu Wei, Wenxin Hu, Xingjiao Wu, Yingbin Zheng, Hao Ye 0005, Jing Yang 0023, Liang He 0001
ICANN (3)3
2019 Adaptive Scenario Discovery for Crowd Counting
abstract
Crowd counting, i.e., estimation number of the pedestrian in crowd images, is emerging as an important research problem with the public security applications. A key component for the crowd counting systems is the construction of counting models which are robust to various scenarios under facts such as camera perspective and physical barriers. In this paper, we present an adaptive scenario discovery framework for crowd counting. The system is structured with two parallel pathways that are trained with different sizes of the receptive field to represent different scales and crowd densities. After ensuring that these components are present in the proper geometric configuration, a third branch is designed to adaptively recalibrate the pathway-wise responses by discovering and modeling the dynamic scenarios implicitly. Our system is able to represent highly variable crowd images and achieves state-of-the-art results in two challenging benchmarks.
Xingjiao Wu, Yingbin Zheng, Hao Ye 0005, Wenxin Hu, Jing Yang 0023, Liang He 0001
ICASSP1