Yanyi Zhang

dblp:145/4473 · DBLP profile ↗
← Back
16ranked-venue papers
8as first author
10since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Computer networks · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Temporal refinement and multi-grained matching for moment retrieval and highlight detection
Cunjuan Zhu, Yanyi Zhang
Multim. Syst.2
2024 CSCNet: Class-Specified Cascaded Network for Compositional Zero-Shot Learning
abstract
Attribute and object (A-O) disentanglement is a fundamental and critical problem for Compositional Zero-shot Learning (CZSL), whose aim is to recognize novel A-O compositions based on foregone knowledge. Existing methods based on disentangled representation learning lose sight of the contextual dependency between the A-O primitive pairs. Inspired by this, we propose a novel A-O disentangled framework for CZSL, namely Class-specified Cascaded Network (CSC-Net). The key insight is to firstly classify one primitive and then specifies the predicted class as a priori for guiding another primitive recognition in a cascaded fashion. To this end, CSCNet constructs Attribute-to-Object and Object-to- Attribute cascaded branches, in addition to a composition branch modeling the two primitives as a whole. Notably, we devise a parametric classifier (ParamCls) to improve the matching between visual and semantic embeddings. By improving the A-O disentanglement, our framework achieves superior results than previous competitive methods.
Yanyi Zhang, Qi Jia 0001, Xin Fan 0001, Yu Liu 0012
ICASSP1
2024 Not Just Object, But State: Compositional Incremental Learning without Forgetting
abstract
Most incremental learners excessively prioritize object classes while neglecting various kinds of states (e.g. color and material) attached to the objects. As a result, they are limited in the ability to model state-object compositionality accurately. To remedy this limitation, we propose a novel task called Compositional Incremental Learning (composition-IL), which enables the model to recognize a variety of state-object compositions in an incremental learning fashion. Since the lack of suitable datasets, we re-organize two existing datasets and make them tailored for composition-IL. Then, we propose a prompt-based Composition Incremental Learner (CompILer), to overcome the ambiguous composition boundary. Specifically, we exploit multi-pool prompt learning, and ensure the inter-pool prompt discrepancy and intra-pool prompt diversity. Besides, we devise object-injected state prompting which injects object prompts to guide the selection of state prompts. Furthermore, we fuse the selected prompts by a generalized-mean strategy, to eliminate irrelevant information learned in the prompts. Extensive experiments on two datasets exhibit state-of-the-art performance achieved by CompILer. Code and datasets are available at: https://github.com/Yanyi-Zhang/CompILer.
Yanyi Zhang, Binglin Qiu
NeurIPS1
2024 PMGNet: Disentanglement and entanglement benefit mutually for compositional zero-shot learning
Yu Liu 0012, Yanyi Zhang, Qi Jia 0001, Weimin Wang 0007, Nan Pu, Nicu Sebe
Comput. Vis. Image Underst.3
2023 Discrete Cosin TransFormer: Image Modeling From Frequency Domain
abstract
In this paper, we propose Discrete Cosin TransFormer (DCFormer) that directly learn semantics from DCT-based frequency domain representation. We first show that transformer-based networks are able to learn semantics directly from frequency domain representation based on discrete cosine transform (DCT) without compromising the performance. To achieve the desired efficiency-effectiveness trade-off, we then leverage an input information compression on its frequency domain representation, which highlights the visually significant signals inspired by JPEG compression. We explore different frequency domain downsampling strategies and show that it is possible to preserve the semantic meaningful information by strategically dropping the high-frequency components. The proposed DCFormer is tested on various downstream tasks including image classification, object detection and instance segmentation, and achieves state-of-the-art comparable performance with less FLOPs, and outperforms the commonly used backbone (e.g. SWIN) at similar FLOPs. Our ablation results also show that the proposed method generalizes well on different transformer backbones.
Yanyi Zhang, Hanlin Lu, Yibo Zhu 0001
WACV2
2022 TubeR: Tubelet Transformer for Video Action Detection
abstract
We propose TubeR: a simple solution for spatio-temporal video action detection. Different from existing methods that depend on either an offline actor detector or hand-designed actor-positional hypotheses like proposals or anchors, we propose to directly detect an action tubelet in a video by simultaneously performing action localization and recognition from a single representation. TubeR learns a set of tubelet-queries and utilizes a tubelet-attention module to model the dynamic spatio-temporal nature of a video clip, which effectively reinforces the model capacity compared to using actor-positional hypotheses in the spatio-temporal space. For videos containing transitional states or scene changes, we propose a context aware classification head to utilize short-term and long-term context to strengthen action classification, and an action switch regression head for detecting the precise temporal action extent. TubeR directly produces action tubelets with variable lengths and even maintains good results for long video clips. TubeR outperforms the previous state-of-the-art on commonly used action detection datasets AVA, UCF101-24 and JHMDB51-21. Code will be available on GluonCV(https://cv.gluon.ai/).
Jiaojiao Zhao, Yanyi Zhang, Xinyu Li 0003, Hao Chen 0024, Bing Shuai, Chunhui Liu 0002, Kaustav Kundu, Yuanjun Xiong, Davide Modolo, Ivan Marsic, Cees Snoek, Joseph Tighe
CVPR2
2021 Multi-Label Activity Recognition Using Activity-Specific Features and Activity Correlations
abstract
Multi-label activity recognition is designed for recognizing multiple activities that are performed simultaneously or sequentially in each video. Most recent activity recognition networks focus on single-activities, that assume only one activity in each video. These networks extract shared features for all the activities, which are not designed for multi-label activities. We introduce an approach to multi-label activity recognition that extracts independent feature descriptors for each activity and learns activity correlations. This structure can be trained end-to-end and plugged into any existing network structures for video classification. Our method outperformed state-of-the-art approaches on four multi-label activity recognition datasets. To better understand the activity-specific features that the system generated, we visualized these activity-specific features in the Charades dataset. The code will be released later.
Yanyi Zhang, Xinyu Li 0003, Ivan Marsic
CVPR1
2021 VidTr: Video Transformer Without Convolutions
abstract
We introduce Video Transformer (VidTr) with separable-attention for video classification. Comparing with commonly used 3D networks, VidTr is able to aggregate spatiotemporal information via stacked attentions and provide better performance with higher efficiency. We first introduce the vanilla video transformer and show that transformer module is able to perform spatio-temporal modeling from raw pixels, but with heavy memory usage. We then present VidTr which reduces the memory cost by 3.3× while keeping the same performance. To further optimize the model, we propose the standard deviation based topK pooling for attention (pooltopK_std), which reduces the computation by dropping non-informative features along temporal dimension. VidTr achieves state-of-the-art performance on five commonly used datasets with lower computational requirement, showing both the efficiency and effectiveness of our design. Finally, error analysis and visualization show that VidTr is especially good at predicting actions that require long-term temporal reasoning.
Yanyi Zhang, Xinyu Li 0003, Chunhui Liu 0002, Bing Shuai, Yi Zhu 0001, Biagio Brattoli, Hao Chen 0024, Ivan Marsic, Joseph Tighe
ICCV1
2021 Auxiliary diagnostic system for ADHD in children based on AI technology
abstract
Traditional diagnosis of attention deficit hyperactivity disorder (ADHD) in children is primarily through a questionnaire filled out by parents/teachers and clinical observations by doctors. It is inefficient and heavily depends on the doctor’s level of experience. In this paper, we integrate artificial intelligence (AI) technology into a software-hardware coordinated system to make ADHD diagnosis more efficient. Together with the intelligent analysis module, the camera group will collect the eye focus, facial expression, 3D body posture, and other children’s information during the completion of the functional test. Then, a multi-modal deep learning model is proposed to classify abnormal behavior fragments of children from the captured videos. In combination with other system modules, standardized diagnostic reports can be automatically generated, including test results, abnormal behavior analysis, diagnostic aid conclusions, and treatment recommendations. This system has participated in clinical diagnosis in Department of Psychology, The Children’s Hospital, Zhejiang University School of Medicine, and has been accepted and praised by doctors and patients.
Yanyi Zhang, Ming Kong 0001, Wenchen Hong, Di Xie, Chunmao Wang, Rongwang Yang
Frontiers Inf. Technol. Electron. Eng.1
2021 Real-time medical phase recognition using long-term video understanding and progress gate method
Yanyi Zhang, Ivan Marsic, Randall S. Burd
Medical Image Anal.1
2020 ADHD Intelligent Auxiliary Diagnosis System Based on Multimodal Information Fusion
abstract
The traditional medical diagnosis methods of ADHD mainly rely on scale evaluation and interview observation. The diagnosis conclusion is subjective and extremely dependent on the doctor's experience level. There is an urgent need to improve diagnosis efficiency and improve the diagnosis standard through other technical means in the clinical process. We have designed and developed the ADHD intelligent auxiliary diagnosis system with software and hardware cooperation. The system performs a set of functional test tasks, uses a camera module to capture multimodal information such as facial expressions, eye movements, limb movements, language expressions and reaction abilities of children during task completion, and uses computer vision technology to automatically extract measurable characteristics. Finally, deep learning technology is used to detect children's specific behaviors in the video, which is complementary to the existing doctor's diagnosis basis. This system was deployed in the Department of Psychology of Children's Hospital of Zhejiang University in July 2019 and has been used in actual clinical diagnosis to date. It has completed the testing and evaluation of hundreds of ADHD children.
Yanyi Zhang, Ming Kong 0001, Wenchen Hong, Fei Wu 0001
ACM Multimedia1
2017 3D activity localization with multiple sensors: poster abstract
abstract
We present a deep learning framework for fast 3D activity localization and tracking in a dynamic and crowded real world setting. Our training approach reverses the traditional activity localization approach, which first estimates the possible location of activities and then predicts their occurrence. Instead, we first trained a deep convolutional neural network for activity recognition using depth video and RFID data as input, and then used the activation maps of the network to locate the recognized activity in the 3D space. Our system achieved around 20cm average localization error (in a 4m × 5m room) which is comparable to Kinect's body skeleton tracking error (10--20cm), but our system tracks activities instead of Kinect's location of people.
Xinyu Li 0003, Yanyi Zhang, Shuhong Chen, Richard A. Farneth, Ivan Marsic, Randall S. Burd
IPSN2
2017 CAR - a deep learning structure for concurrent activity recognition: poster abstract
abstract
We introduce the Concurrent Activity Recognizer (CAR) - an efficient deep learning structure that recognizes complex concurrent teamwork activities from multimodal data. We implemented the system in a challenging medical setting, where it recognizes 35 different activities using Kinect depth video and data from passive RFID tags on 25 types of medical objects. Our preliminary results showed our system achieved an 84% average accuracy with 0.20 F1-Score.
Yanyi Zhang, Xinyu Li 0003, Shuhong Chen, Moliang Zhou, Richard A. Farneth, Ivan Marsic, Randall S. Burd
IPSN1
2017 Region-based Activity Recognition Using Conditional GAN
abstract
We present a method for activity recognition that first estimates the activity performer's location and uses it with input data for activity recognition. Existing approaches directly take video frames or entire video for feature extraction and recognition, and treat the classifier as a black box. Our method first locates the activities in each input video frame by generating an activity mask using a conditional generative adversarial network (cGAN). The generated mask is appended to color channels of input images and fed into a VGG-LSTM network for activity recognition. To test our system, we produced two datasets with manually created masks, one containing Olympic sports activities and the other containing trauma resuscitation activities. Our system makes activity prediction for each video frame and achieves performance comparable to the state-of-the-art systems while simultaneously outlining the location of the activity. We show how the generated masks facilitate the learning of features that are representative of the activity rather than accidental surrounding information.
Xinyu Li 0003, Yanyi Zhang, Yueyang Chen, Huangcan Li, Ivan Marsic, Randall S. Burd
ACM Multimedia2
2016 Privacy Preserving Dynamic Room Layout Mapping
Xinyu Li 0003, Yanyi Zhang, Ivan Marsic, Randall S. Burd
ICISP2
2016 Deep Learning for RFID-Based Activity Recognition
abstract
We present a system for activity recognition from passive RFID data using a deep convolutional neural network. We directly feed the RFID data into a deep convolutional neural network for activity recognition instead of selecting features and using a cascade structure that first detects object use from RFID data followed by predicting the activity. Because our system treats activity recognition as a multi-class classification problem, it is scalable for applications with large number of activity classes. We tested our system using RFID data collected in a trauma room, including 14 hours of RFID data from 16 actual trauma resuscitations. Our system outperformed existing systems developed for activity recognition and achieved similar performance with process-phase detection as systems that require wearable sensors or manually-generated input. We also analyzed the strengths and limitations of our current deep learning architecture for activity recognition from RFID data.
Xinyu Li 0003, Yanyi Zhang, Ivan Marsic, Aleksandra Sarcevic, Randall S. Burd
SenSys2