Yujia Zhang 0001

dblp:116/1506-1 · DBLP profile ↗
← Back
24ranked-venue papers
6as first author
18since 2021 · last 2026
0000-0002-2335-7657ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 ProPy : Building interactive and efficient prompt pyramids upon CLIP for partially relevant video retrieval
Yujia Zhang 0001, Michael Kampffmeyer, Xiaoguang Zhao
Neural Networks2
2025 RefCap: Zero-shot Video Corpus Moment Retrieval Based on Refined Dense Video Captioning
abstract
Video corpus moment retrieval (VCMR) is a challenging task aimed at localizing specific segments from untrimmed videos within a vast video collection. It has long been addressed using end-to-end supervised or weakly-supervised methods, which often lack explainability and rely on laborious annotations. To address these issues, we exploit generated captions from Vision Large Language Models (VLLMs) and propose the first zero-shot training-free VCMR system, RefCap. The system consists of two decoupled stages: the Construction Stage to propose dense events and construct corresponding captions, and the Retrieval Stage to retrieve events based on text similarities between queries and captions. In the Construction Stage, we design a Sliding-Window Denoiser and a Quality-Measured Event Generator to produce high-quality dense captions, and derive the Indexing Keyword Sets to fully utilize the generated captions. In the Retrieval Stage, we propose a multi-granularity retrieval strategy integrating both sentence-level and word-level textual similarities between queries and event captions. Extensive experiments on the Charades and ActivityNet datasets demonstrate the effectiveness and competitiveness of our method. Code is available at https://github.com/BUAAPY/RefCap.
Yujia Zhang 0001, Michael Kampffmeyer, Xiaoguang Zhao
ICASSP2
2025 FAWL: Weakly-Supervised Video Corpus Moment Retrieval with Frame-Wise Auxiliary Alignment and Weighted Contrastive Learning
abstract
Video Corpus Moment Retrieval (VCMR) is a challenging task that aims to localize query-specified moments from a collection of untrimmed videos. The recent state-of-the-art method, JSG, tries to tackle this task using only video-level annotations in a weakly-supervised setting. However, the late fusion strategy of JSG suffers from insufficient alignment, and the proposal-level video representations aggregated with Gaussian Distributions result in semantic inconsistency across frames. To address these issues, we propose a novel weakly-supervised VCMR method, FAWL, which incorporates frame-wise auxiliary alignment and weighted contrastive learning. Two frame-wise auxiliary alignment tasks, namely Query-guided Saliency Alignment (QSA) and Event-aware Boundary Alignment (EBA), are designed first. During training, QSA projects video and text representations through a shared layer, and computes frame-level saliency losses for sufficient multimodal alignment. EBA guides frames to predict distances to proposal boundaries making each frame aware of the corresponding events and thus ensuring semantic consistency. To help the model distinguish proposals with similar visual semantics, we further propose the Weighted Contrastive Learning (WCL) which integrates inter-video similarities into the conventional InfoNCE loss. FAWL enjoys both improved alignment and inference efficiency, achieving new state-of-the-art performance on two challenging datasets. Code is available at https://github.com/BUAAPY/FAWL.
Yujia Zhang 0001, Xiaoguang Zhao
ICASSP2
2025 Compositional Text-Modality Completion Model for Partially Relevant Video Retrieval
abstract
Partially Relevant Video Retrieval (PRVR) is a challenging task aimed at retrieving videos based on partially relevant text queries. Previous PRVR models typically adopt the Multiple Instance Learning (MIL) framework, which are limited to annotated segments and are ineffective when dealing with unlabeled segments. To address this limitation, we approach the PRVR task from a novel missing modality completion perspective to provide supplementary textual supervision signals for better alignment. To this end, we propose a plug-and-play Text-Modality Completion (CTC) Model to generate pseudo-descriptions based on existing annotations and the intrinsic attribute of the PRVR task. Our key insight is that the semantics of unlabeled segments can be composed of intra-video and inter-video semantics. Thus, we first design a Semantic Decomposing Module to decompose features into fine-grained semantic units, with a Textual Semantic Sampling strategy to further enrich text representations. Subsequently, a Semantic Composing Module is introduced to compose intra-video and inter-video semantics based on the decomposed semantic units. The decomposed textual semantics, weighted by visual semantic similarities, serve as intra-video semantics, mapping visual representations of the same objects to their corresponding concepts. For inter-video semantics, we design a Generative Semantic Memory that accumulates shared semantics across videos by generating textual semantics given visual semantics. Finally, both intra-video and inter-video semantics are integrated to form pseudo-descriptions, which are leveraged as enriched textual supervision signals for further contrastive alignment. We evaluate the CTC model with different base PRVR models on different public datasets. The consistent improvements demonstrate the effectiveness and generalizability of our method.
Yujia Zhang 0001, Xiaoguang Zhao
ICME2
2025 Prompt-guided bidirectional deep fusion network for referring image segmentation
Junxian Wu 0001, Yujia Zhang 0001, Michael Kampffmeyer, Xiaoguang Zhao
Neurocomputing2
2025 HierGAT: hierarchical spatial-temporal network with graph and transformer for video HOI detection
Junxian Wu 0001, Yujia Zhang 0001, Michael Kampffmeyer, Shiying Sun, Xiaoguang Zhao
Multim. Syst.2
2024 Coarse-to-Fine Recurrently Aligned Transformer with Balance Tokens for Video Moment Retrieval and Highlight Detection
abstract
Video moment retrieval (MR) and highlight detection (HD) are two user-oriented video understanding tasks aimed at extracting query-dependent or highlighted moments to provide valuable content for users. While many recent works have proposed solutions for the joint task of MR and HD leveraging transformer architecture, we argue that existing approaches have not adequately aligned the video and text modalities using basic transformer encoders, and have overlooked the misalignment between irrelevant video clips and text queries. To address these issues, we introduce COREBA: a Coarse-to-Fine Recurrently Aligned Transformer with Balance Tokens. Firstly, we design a plug-and-play Coarse-to-Fine Cross-modal interaction (CFC) module, replacing the original transformer encoder to align the two modalities in a progressive manner. Secondly, we present a novel Recurrent Alignment Mechanism (RAM) to deeply align the modalities in a recurrent fashion. Thirdly, to mitigate the misalignment problem, we append text queries with learnable Balance Tokens to restrict the text information fused with irrelevant clips. Extensive experiments validate the effectiveness and superiority of our proposed method.
Yujia Zhang 0001, Shiying Sun, Feihu Zhou, Xiaoguang Zhao
IJCNN2
2024 PSAIR: A Neuro-Symbolic Approach to Zero-Shot Visual Grounding
abstract
Supervised methods for Visual Grounding often require costly annotations of paired sentences and images with ground truth boxes. Recent zero-shot approaches to visual grounding such as ReCLIP and ChatRef aim to avoid the need for costly annotation of paired sentences and images with ground truth boxes. However, these approaches leverage an inflexible detect-then-reasoning paradigm, which leads to a notable semantic information loss. Additionally, these approaches are highly dependent on the definition of predefined keywords or the potentially inconsistent reasoning capabilities of Large Language Models (LLMs). To address these limitations, we propose a neuro-symbolic visual grounding method, PSAIR that incorporates two novel mechanisms, namely Parallel Scoring and Active Information Retrieval. PSAIR equips an LLM with external encapsulated reasoning functions, which the LLM can invoke, ensuring a flexible and stable reasoning process. The proposed Parallel Scoring mechanism is then used to reframe the sequential reasoning process found in prior approaches to facilitate robustness to noise in the detection process. Subsequently, the Active Information Retrieval mechanism is designed to address the loss of semantic information by having the ability to retrieve essential visual information, resulting in a new detect-reason-retrieve paradigm. These innovations result in superior performance and robustness across three public datasets compared to recent state-of-the-art zero-shot visual grounding methods.
Yujia Zhang 0001, Michael Kampffmeyer, Xiaoguang Zhao
IJCNN2
2024 Multi-Stage Image-Language Cross-Generative Fusion Network for Video-Based Referring Expression Comprehension
abstract
Video-based referring expression comprehension is a challenging task that requires locating the referred object in each video frame of a given video. While many existing approaches treat this task as an object-tracking problem, their performance is heavily reliant on the quality of the tracking templates. Furthermore, when there is not enough annotation data to assist in template selection, the tracking may fail. Other approaches are based on object detection, but they often use only one adjacent frame of the key frame for feature learning, which limits their ability to establish the relationship between different frames. In addition, improving the fusion of features from multiple frames and referring expressions to effectively locate the referents remains an open problem. To address these issues, we propose a novel approach called the Multi-Stage Image-Language Cross-Generative Fusion Network (MILCGF-Net), which is based on one-stage object detection. Our approach includes a Frame Dense Feature Aggregation module for dense feature learning of adjacent time sequences. Additionally, we propose an Image-Language Cross-Generative Fusion module as the main body of multi-stage learning to generate cross-modal features by calculating the similarity between video and expression, and then refining and fusing the generated features. To further enhance the cross-modal feature generation capability of our model, we introduce a consistency loss that constrains the image-language similarity and language-image similarity matrices during feature generation. We evaluate our proposed approach on three public datasets and demonstrate its effectiveness through comprehensive experimental results.
Yujia Zhang 0001, Qianzhong Li, Xiaoguang Zhao, Min Tan 0001
IEEE Trans. Image Process.1
2023 Visual enhanced hierarchical network for sentence-based video thumbnail generation
Junxian Wu 0001, Yujia Zhang 0001, Xiaoguang Zhao
Appl. Intell.2
2022 Hand Acupoint Detection from Images Based on Improved HRNet
abstract
As an important component of Traditional Chinese Medicine (TCM), the acupoint therapy has achieved significant success in clinical practice. However, at present, the effect of acupoint therapy heavily depends on the skills of doctors and the acupuncture medical resources are seriously insufficient. The introduction of artificial intelligence technology in acupoint therapy can reduce the workload of doctors and ensure the consistency of acupoint operations, which is of great significance. The key process of acupoint therapy is acupoint detection. In this paper, we apply the deep learning method in automatic acupoint detection using images and propose an improved High-Resolution Network (HRNet) method for hand acupoint detection. What's more, we build a hand acupoint detection dataset and propose an evaluation metric. Experiments on the proposed dataset verify the effectiveness of the proposed method.
Shiying Sun, Hongduo Xu, Lingyao Sun, Yuanbo Fu, Yujia Zhang 0001, Xiaoguang Zhao
IJCNN5
2022 Multi-Task Learning for Pavement Disease Segmentation Using Wavelet Transform
abstract
Pavement safety is a significant part of transportation safety. In order to realize pavement inspection, many existing approaches have been proposed to detect and segment long-thin cracks while ignoring other pavement diseases. However, other diseases can also cause safety hazards if they are not maintained timely. It is essential to realize the fine-grained segmentation for different pavement diseases to monitor the exact condition of the pavement. Existing algorithms for crack segmentation cannot deal with the segmentation for different diseases with different shapes in the complex environment. To address this, we propose a novel multi-task network for fine-grained pavement disease segmentation. Specifically, the low-frequency information is first introduced to suppress the background noise. The semantic boundary branch is then designed to extract multi-scale feature maps, which can guide the generation of segmentation results for pavement diseases with different shapes. Furthermore, attention mechanisms are utilized to alleviate the effects of extreme imbalances in the number of pixels of different diseases and backgrounds. The network jointly learns the semantic boundary task and the segmentation task in an end-to-end manner to produce the final predictions, and the extensive experiments on a newly collected airport pavement dataset verify the ef-fectiveness of our approach. All the codes are available at https://github.com/wjx1198/MTSSN-WT.
Junxian Wu 0001, Yujia Zhang 0001, Xiaoguang Zhao
IJCNN2
2022 Generalized zero-shot emotion recognition from body gestures
Jinting Wu, Yujia Zhang 0001, Shiying Sun, Qianzhong Li, Xiaoguang Zhao
Appl. Intell.2
2022 Cross-modality synergy network for referring expression comprehension and segmentation
Qianzhong Li, Yujia Zhang 0001, Shiying Sun, Jinting Wu, Xiaoguang Zhao, Min Tan 0001
Neurocomputing2
2022 Beyond Crack: Fine-Grained Pavement Defect Segmentation Using Three-Stream Neural Networks
abstract
Pavement defect segmentation is a fundamental task in the field of transport infrastructure inspection. Existing methods mainly focus on detection/segmentation for long and thin cracks. However, there are many other types of defects with various sizes and shapes that are also essential to segment, which brings more challenges toward detailed road inspection. To address the above problems and provide a more comprehensive understanding of the overall road conditions, we propose a three-stream neural network that combines spatial, contextual and boundary information for fine-grained defect segmentation. Specifically, the spatial stream captures rich low-level spatial features. The contextual stream utilizes an attention mechanism and models high-level contextual relationships over local features. To further refine the segmentation results, the boundary stream encodes detailed boundaries using a global gated convolution and generates additional boundary maps. By combining the above different information, our model can effectively produce pixel-wise predictions for fine-grained road inspection. The network is trained using a dual-task loss in an end-to-end manner, and experiments were performed on three newly collected datasets, i.e., a fine-grained defect dataset and two crack datasets, which shows that the proposed method achieves favorable segmentation results on complex multi-class defects, and is also able to segment single-class cracks. Specifically, on the fine-grained dataset, it achieved state-of-the-art performance over other competing baselines (mPA of 0.54, mIoU of 0.38, Mic_F of 0.78 and Mac_F of 0.65), where each image is resized to 512$\times 512$and the processing speed is 21 FPS on average.
Yujia Zhang 0001, Junxian Wu 0001, Qianzhong Li, Xiaoguang Zhao, Min Tan 0001
IEEE Trans. Intell. Transp. Syst.1
2021 Stress Detection Using Wearable Devices based on Transfer Learning
abstract
Excessive stress will have a negative impact on people’s physical and mental health, especially for some special occupations. Because stressful stimuli can trigger a variety of physiological responses, analyzing physiological signals collected by wearable devices has become an important way to evaluate the stress state in recent years. However, the number of available subjects of a target group may be small, and collecting a large amount of data when the target group changes is costly and time-consuming. To solve this problem, we propose a stress detection framework for a small target group which uses adversarial transfer learning method to learn shared knowledge about stress between different groups. In order to verify the performance of the framework, we establish a dataset consisting of 264 ordinary college students and 32 police school students, aiming to evaluate the acute stress state of police school students under video stimuli for psychological training in the future. Comprehensive experiments show that our algorithm has achieved a significant improvement in the target group compared with the baseline methods.
Jinting Wu, Yujia Zhang 0001, Xiaoguang Zhao
BIBM2
2021 TB-Net: A Three-Stream Boundary-Aware Network for Fine-Grained Pavement Disease Segmentation
abstract
Regular pavement inspection plays a significant role in road maintenance for safety assurance. Existing methods mainly address the tasks of crack detection and segmentation that are only tailored for long-thin crack disease. However, there are many other types of diseases with a wider variety of sizes and patterns that are also essential to segment in practice, bringing more challenges towards fine-grained pavement inspection. In this paper, our goal is not only to automatically segment cracks, but also to segment other complex pavement diseases as well as typical landmarks (markings, runway lights, etc.) and commonly seen water/oil stains in a single model. To this end, we propose a three-stream boundary-aware network (TB-Net). It consists of three streams fusing the low-level spatial and the high-level contextual representations as well as the detailed boundary information. Specifically, the spatial stream captures rich spatial features. The context stream, where an attention mechanism is utilized, models the contextual relationships over local features. The boundary stream learns detailed boundaries using a global-gated convolution to further refine the segmentation outputs. The network is trained using a dual-task loss in an end-to-end manner, and experiments on a newly collected fine-grained pavement disease dataset show the effectiveness of our TB-Net.
Yujia Zhang 0001, Qianzhong Li, Xiaoguang Zhao, Min Tan 0001
WACV1
2021 Rethinking semantic-visual alignment in zero-shot object detection via a softplus margin focal loss
Qianzhong Li, Yujia Zhang 0001, Shiying Sun, Xiaoguang Zhao, Min Tan 0001
Neurocomputing2
2020 A Prototype-Based Generalized Zero-Shot Learning Framework for Hand Gesture Recognition
abstract
Hand gesture recognition plays a significant role in human-computer interaction for understanding various human gestures and their intent. However, most prior works can only recognize gestures of limited labeled classes and fail to adapt to new categories. The task of Generalized Zero-Shot Learning (GZSL) for hand gesture recognition aims to address the above issue by leveraging semantic representations and detecting both seen and unseen class samples. In this paper, we propose an end-to-end prototype-based GZSL framework for hand gesture recognition which consists of two branches. The first branch is a prototype-based detector that learns gesture representations and determines whether an input sample belongs to a seen or unseen category. The second branch is a zero-shot label predictor which takes the features of unseen classes as input and outputs predictions through a learned mapping mechanism between the feature and the semantic space. We further establish a hand gesture dataset that specifically targets this GZSL task, and comprehensive experiments on this dataset demonstrate the effectiveness of our proposed approach on recognizing both seen and unseen gestures.
Jinting Wu, Yujia Zhang 0001, Xiaoguang Zhao
ICPR2
2020 Unsupervised object-level video summarization with online motion auto-encoder
Yujia Zhang 0001, Xiaodan Liang, Dingwen Zhang, Min Tan 0001, Eric P. Xing
Pattern Recognit. Lett.1
2019 Rethinking Knowledge Graph Propagation for Zero-Shot Learning
abstract
Graph convolutional neural networks have recently shown great potential for the task of zero-shot learning. These models are highly sample efficient as related concepts in the graph structure share statistical strength allowing generalization to new classes when faced with a lack of data. However, multi-layer architectures, which are required to propagate knowledge to distant nodes in the graph, dilute the knowledge by performing extensive Laplacian smoothing at each layer and thereby consequently decrease performance. In order to still enjoy the benefit brought by the graph structure while preventing dilution of knowledge from distant nodes, we propose a Dense Graph Propagation (DGP) module with carefully designed direct links among distant nodes. DGP allows us to exploit the hierarchical graph structure of the knowledge graph through additional connections. These connections are added based on a node's relationship to its ancestors and descendants. A weighting scheme is further used to weigh their contribution depending on the distance to the node to improve information propagation in the graph. Combined with finetuning of the representations in a two-stage training approach our method outperforms state-of-the-art zero-shot learning approaches.
Michael Kampffmeyer, Yinbo Chen, Xiaodan Liang, Hao Wang 0014, Yujia Zhang 0001, Eric P. Xing
CVPR5
2019 Dilated temporal relational adversarial network for generic video summarization
Yujia Zhang 0001, Michael Kampffmeyer, Xiaodan Liang, Dingwen Zhang, Min Tan 0001, Eric P. Xing
Multim. Tools Appl.1
2019 ConnNet: A Long-Range Relation-Aware Pixel-Connectivity Network for Salient Segmentation
abstract
Salient segmentation aims to segment out attentiongrabbing regions, a critical yet challenging task and the foundation of many high-level computer vision applications. It requires semantic-aware grouping of pixels into salient regions and benefits from the utilization of global multi-scale contexts to achieve good local reasoning. Previous works often address it as two-class segmentation problems utilizing complicated multi-step procedures including refinement networks and complex graphical models. We argue that semantic salient segmentation can instead be effectively resolved by reformulating it as a simple yet intuitive pixel-pair based connectivity prediction task. Following the intuition that salient objects can be naturally grouped via semanticaware connectivity between neighboring pixels, we propose a pure Connectivity Net (ConnNet). ConnNet predicts connectivity probabilities of each pixel with its neighboring pixels by leveraging multi-level cascade contexts embedded in the image and long-range pixel relations. We investigate our approach on two tasks, namely salient object segmentation and salient instancelevel segmentation, and illustrate that consistent improvements can be obtained by modeling these tasks as connectivity instead of binary segmentation tasks for a variety of network architectures. We achieve state-of-the-art performance, outperforming or being comparable to existing approaches while reducing inference time due to our less complex approach.
Michael Kampffmeyer, Nanqing Dong, Xiaodan Liang, Yujia Zhang 0001, Eric P. Xing
IEEE Trans. Image Process.4
2018 Query-Conditioned Three-Player Adversarial Network for Video Summarization
Yujia Zhang 0001, Michael Kampffmeyer, Xiaodan Liang, Min Tan 0001, Eric P. Xing
BMVC1