VLDB 2026 Research / reviewers in the wild / expert
Ziyi Liu 0001
dblp:09/3084-1
· DBLP profile ↗
15ranked-venue papers
5as first author
8since 2021 · last 2024
0000-0002-2407-9599ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 9 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Correlation Information Bottleneck: Towards Adapting Pretrained Multimodal Models for Robust Visual Question Answering
Jingjing Jiang, Ziyi Liu 0001, Nanning Zheng 0001 |
Int. J. Comput. Vis. | 2 |
| 2023 | SOAR: Scene-debiasing Open-set Action RecognitionabstractDeep learning models have a risk of utilizing spurious clues to make predictions, such as recognizing actions based on the background scene. This issue can severely degrade the open-set action recognition performance when the testing samples have different scene distributions from the training samples. To mitigate this problem, we propose a novel method, called Scene-debiasing Open-set Action Recognition (SOAR), which features an adversarial scene reconstruction module and an adaptive adversarial scene classification module. The former prevents the decoder from reconstructing the video background given video features, and thus helps reduce the background information in feature learning. The latter aims to confuse scene type classification given video features, with a specific emphasis on the action foreground, and helps to learn scene-invariant information. In addition, we design an experiment to quantify the scene bias. The results indicate that the current open-set action recognizers are biased toward the scene, and our proposed SOAR method better mitigates such bias. Furthermore, our extensive experiments demonstrate that our method outperforms state-of-the-art methods, and the ablation studies confirm the effectiveness of our proposed modules. Yuanhao Zhai 0001, Ziyi Liu 0001, Zhenyu Wu 0002, Chunluan Zhou, David S. Doermann, Junsong Yuan 0001, Gang Hua 0001 |
ICCV | 2 |
| 2023 | LiVLR: A Lightweight Visual-Linguistic Reasoning Framework for Video Question AnsweringabstractVideo Question Answering (VideoQA), aiming to correctly answer a given question based on understanding multimodal video content, is challenging due to the richness of the video content. From the perspective of video understanding, a complete VideoQA framework needs to understand the video content at different semantic levels and flexibly integrate diverse video content to distill question-related content. To this end, we propose a Lightweight Visual-Linguistic Reasoning framework named$\text{LiVLR}$. Specifically,$\text{LiVLR}$first utilizes graph-based visual and linguistic encoders to obtain multi-grained visual and linguistic representations, respectively. Subsequently, the obtained representations are integrated with the devised Diversity-aware Visual-Linguistic Reasoning module ($\text{DaVL}$).$\text{DaVL}$distinguishes different types of representations with the learnable index embedding in graph embedding. Therefore,$\text{DaVL}$can flexibly adjust the importance of different representations when generating the question-related joint representation. The proposed$\text{LiVLR}$is lightweight and shows its performance advantage on three VideoQA benchmarks, MRSVTT-QA, KnowIT VQA, and TVQA. Extensive ablation studies demonstrate the effectiveness of the key components of$\text{LiVLR}$. Jingjing Jiang, Ziyi Liu 0001, Nanning Zheng 0001 |
IEEE Trans. Multim. | 2 |
| 2022 | Learning Disentangled Classification and Localization Representations for Temporal Action LocalizationabstractA common approach to Temporal Action Localization (TAL) is to generate action proposals and then perform action classification and localization on them. For each proposal, existing methods universally use a shared proposal-level representation for both tasks. However, our analysis indicates that this shared representation focuses on the most discriminative frames for classification, e.g., ``take-offs" rather than ``run-ups" in distinguishing ``high jump" and ``long jump", while frames most relevant to localization, such as the start and end frames of an action, are largely ignored. In other words, such a shared representation can not simultaneously handle both classification and localization tasks well, and it makes precise TAL difficult. To address this challenge, this paper disentangles the shared representation into classification and localization representations. The disentangled classification representation focuses on the most discriminative frames, and the disentangled localization representation focuses on the action phase as well as the action start and end. Our model could be divided into two sub-networks, i.e., the disentanglement network and the context-based aggregation network. The disentanglement network is an autoencoder to learn orthogonal hidden variables of classification and localization. The context-based aggregation network aggregates the classification and localization representations by modeling local and global contexts. We evaluate our proposed method on two popular benchmarks for TAL, which outperforms all state-of-the-art methods. Zixin Zhu, Le Wang 0003, Wei Tang 0016, Ziyi Liu 0001, Nanning Zheng 0001, Gang Hua 0001 |
AAAI | 4 |
| 2022 | Weakly Supervised Temporal Action Localization Through Contrast Based Evaluation NetworksabstractGiven only video-level action categorical labels during training, weakly-supervised temporal action localization (WS-TAL) learns to detect action instances and locates their temporal boundaries in untrimmed videos. Compared to its fully supervised counterpart, WS-TAL is more cost-effective in data labeling and thus favorable in practical applications. However, the coarse video-level supervision inevitably incurs ambiguities in action localization, especially in untrimmed videos containing multiple action instances. To overcome this challenge, we observe that significant temporal contrasts among video snippets, e.g., caused by temporal discontinuities and sudden changes, often occur around true action boundaries. This motivates us to introduce a Contrast-based Localization EvaluAtioN Network (CleanNet), whose core is a new temporal action proposal evaluator, which provides fine-grained pseudo supervision by leveraging the temporal contrasts among snippet-level classification predictions. As a result, the uncertainty in locating action instances can be resolved via evaluating their temporal contrast scores. Moreover, the new action localization module is an integral part of CleanNet which enables end-to-end training. This is in contrast to many existing WS-TAL methods where action localization is merely a post-processing step. Besides, we also explore the usage of temporal contrast on temporal action proposal (TAP) generation task, which we believe is the first attempt with the weak supervision setting. Experiments on the THUMOS14, ActivityNet v1.2 and v1.3 datasets validate the efficacy of our method against existing state-of-the-art WS-TAL algorithms. Ziyi Liu 0001, Le Wang 0003, Qilin Zhang 0004, Wei Tang 0016, Nanning Zheng 0001, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Weakly Supervised Temporal Action Localization Through Learning Explicit Subspaces for Action and ContextabstractWeakly-supervised Temporal Action Localization (WS-TAL) methods learn to localize temporal starts and ends of action instances in a video under only video-level supervision. Existing WS-TAL methods rely on deep features learned for action recognition. However, due to the mismatch between classification and localization, these features cannot distinguish the frequently co-occurring contextual background, i.e., the context, and the actual action instances. We term this challenge action-context confusion, and it will adversely affect the action localization accuracy. To address this challenge, we introduce a framework that learns two feature subspaces respectively for actions and their context. By explicitly accounting for action visual elements, the action instances can be localized more precisely without the distraction from the context. To facilitate the learning of these two feature subspaces with only video-level categorical labels, we leverage the predictions from both spatial and temporal streams for snippets grouping. In addition, an unsupervised learning task is introduced to make the proposed module focus on mining temporal information. The proposed approach outperforms state-of-the-art WS-TAL methods on three benchmarks, i.e., THUMOS14, ActivityNet v1.2 and v1.3 datasets. Ziyi Liu 0001, Le Wang 0003, Wei Tang 0016, Junsong Yuan 0001, Nanning Zheng 0001, Gang Hua 0001 |
AAAI | 1 |
| 2021 | ACSNet: Action-Context Separation Network for Weakly Supervised Temporal Action LocalizationabstractThe object of Weakly-supervised Temporal Action Localization (WS-TAL) is to localize all action instances in an untrimmed video with only video-level supervision. Due to the lack of frame-level annotations during training, current WS-TAL methods rely on attention mechanisms to localize the foreground snippets or frames that contribute to the video-level classification task. This strategy frequently confuse context with the actual action, in the localization result. Separating action and context is a core problem for precise WS-TAL, but it is very challenging and has been largely ignored in the literature. In this paper, we introduce an Action-Context Separation Network (ACSNet) that explicitly takes into account context for accurate action localization. It consists of two branches (i.e., the Foreground-Background branch and the Action-Context branch). The Foreground-Background branch first distinguishes foreground from background within the entire video while the Action-Context branch further separates the foreground as action and context. We associate video snippets with two latent components (i.e., a positive component and a negative component), and their different combinations can effectively characterize foreground, action and context. Furthermore, we introduce extended labels with auxiliary context categories to facilitate the learning of action-context separation. Experiments on THUMOS14 and ActivityNet v1.2/v1.3 datasets demonstrate the ACSNet outperforms existing state-of-the-art WS-TAL methods by a large margin. Ziyi Liu 0001, Le Wang 0003, Qilin Zhang 0004, Wei Tang 0016, Junsong Yuan 0001, Nanning Zheng 0001, Gang Hua 0001 |
AAAI | 1 |
| 2021 | X-GGM: Graph Generative Modeling for Out-of-distribution Generalization in Visual Question AnsweringabstractEncouraging progress has been made towards Visual Question Answering (VQA) in recent years, but it is still challenging to enable VQA models to adaptively generalize to out-of-distribution (OOD) samples. Intuitively, recompositions of existing visual concepts (i.e., attributes and objects) can generate unseen compositions in the training set, which will promote VQA models to generalize to OOD samples. In this paper, we formulate OOD generalization in VQA as a compositional generalization problem and propose a graph generative modeling-based training scheme (X-GGM) to handle the problem implicitly. X-GGM leverages graph generative modeling to iteratively generate a relation matrix and node representations for the predefined graph that utilizes attribute-object pairs as nodes. Furthermore, to alleviate the unstable training issue in graph generative modeling, we propose a gradient distribution consistency loss to constrain the data distribution with adversarial perturbations and the generated distribution. The baseline VQA model (LXMERT) trained with the X-GGM scheme achieves state-of-the-art OOD performance on two standard VQA OOD benchmarks, i.e., VQA-CP v2 and GQA-OOD. Extensive ablation studies demonstrate the effectiveness of X-GGM components. Jingjing Jiang, Ziyi Liu 0001, Yifan Liu 0014, Zhixiong Nan, Nanning Zheng 0001 |
ACM Multimedia | 2 |
| 2020 | MuRF-Net: Multi-Receptive Field Pillars for 3D Object Detection from Point CloudabstractIn this paper, we propose a point cloud based 3D object detection framework that accounts for both contextual and local information by leveraging multi-receptive field pillars, named as MuRF-Net. Recently, common pipelines can be divided into a voxel-based feature encoder and an object detector. During the feature encoding steps, contextual information is neglected, which is critical for the 3D object detection task. Thus, the encoded features are not suitable to input to the subsequent object detector. To address this challenge, we propose the MuRF-Net with a multi-receptive field voxelization mechanism to capture both contextual and local information. After the voxelization, the voxelized points (pillars) are processed by a feature encoder, and a channel-wise feature reconfiguration module is proposed to combine the features with different receptive fields using a lateral enhanced fusion network. In addition, to handle the increase of memory and computational cost brought by multi-receptive field voxelization, a dynamic voxel encoder is applied taking advantage of the sparseness of the point cloud. Experiments on the KITTI benchmark for both 3D object and Bird's Eye View (BEV) detection tasks on car class are conducted and MuRF-Net achieved the state-of-the-art results compared with other voxel-based methods. Besides, the MuRF-Net can achieve nearly real-time speed with 20Hz. Xinrui Yan, Shi-tao Chen, Jinpeng Dong, Ziyi Liu 0001, Zhixiong Nan, Nanning Zheng 0001 |
IV | 7 |
| 2019 | Weakly Supervised Temporal Action Localization Through Contrast Based Evaluation NetworksabstractWeakly-supervised temporal action localization (WS-TAL) is a promising but challenging task with only video-level action categorical labels available during training. Without requiring temporal action boundary annotations in training data, WS-TAL could possibly exploit automatically retrieved video tags as video-level labels. However, such coarse video-level supervision inevitably incurs confusions, especially in untrimmed videos containing multiple action instances. To address this challenge, we propose the Contrast-based Localization EvaluAtioN Network (CleanNet) with our new action proposal evaluator, which provides pseudo-supervision by leveraging the temporal contrast in snippet-level action classification predictions. Essentially, the new action proposal evaluator enforces an additional temporal contrast constraint so that high-evaluation-score action proposals are more likely to coincide with true action instances. Moreover, the new action localization module is an integral part of CleanNet which enables end-to-end training. This is in contrast to many existing WS-TAL methods where action localization is merely a post-processing step. Experiments on THUMOS14 and ActivityNet datasets validate the efficacy of CleanNet against existing state-ofthe- art WS-TAL algorithms. Ziyi Liu 0001, Le Wang 0003, Qilin Zhang 0004, Zhanning Gao, Zhenxing Niu, Nanning Zheng 0001, Gang Hua 0001 |
ICCV | 1 |
| 2019 | Har Enhanced Weakly-Supervised Semantic Segmentation Coupled with Adversarial LearningabstractSemantic segmentation is a challenging computer visual task which needs enormous pixel-level annotation data. But collecting a large amount of pixel-level annotation data is labor intensive. To address this issue, our work focuses on weakly-supervised learning approach which combines the adversarial learning and localization ability of classification model together, in this way, data with different annotations can be fully utilized. Specifically, the adversarial learning encourages the high order spatial consistences thus offers a relatively reliable initial confidence map. And we find that the hybrid atrous rate (HAR) can improve the localization ability of the classification model, thus indicate more precise object-related regions, which serves as strong supervision information. We conduct experiments with different settings to demonstrate the effectiveness of this weakly-supervised learning approach. The results show that our approach can improve the performance of baseline adversarial learning from 73.2 to 75.1 (mIOU), which is pretty effective. Leiyuan Ma, Ziyi Liu 0001, Nanning Zheng 0001, Jianji Wang 0001 |
ICIP | 2 |
| 2019 | Action Coherence Network for Weakly Supervised Temporal Action LocalizationabstractMost prominent temporal action localization methods are of the fully-supervised type, which rely heavily on frame-level labels, which could be prohibitively expensive to annotate. Thanks to recent developments on the Weakly-supervised Temporal Action Localization (W-TAL), this alternative paradigm requires only video-level labels in training, alleviating such annotation efforts. Specifically, we present Action Coherence Network (ACN) for W-TAL, which features a new coherence loss that better supervises action boundary learning and facilitate proposal regression. In addition, a purpose-built fusion module is proposed for localization inference based on features extracted by two streams of convolutional neural network. Overall, the proposed ACN achieves state-of-the-art W-TAL performance on two challenging datasets (THU-MOS14 and ActivityNet1.2, particularly ACN attains mAP of 24.2% on THUMOS14 under IoU threshold 0.5), which is approaching some recent fully-supervised TAL methods. Yuanhao Zhai 0001, Le Wang 0003, Ziyi Liu 0001, Qilin Zhang 0004, Gang Hua 0001, Nanning Zheng 0001 |
ICIP | 3 |
| 2018 | A Novel Approach for Detecting Road Based on Two-Stream Fusion Fully Convolutional NetworkabstractRoad detection is one of the most basic tasks of autonomous driving systems. At present, researches on this issue mainly take two kinds of data as input,i.e., LIDAR point clouds and RGB images from cameras. To make best use of the advantages and bypass the disadvantages of these two kinds of data, we propose a novel network, namely two- stream fusion fully convolutional network (TSF-FCN), which can take advantage of both the accurate location information from LIDAR point clouds and rich appearance information from RGB images. One stream of this network is LIDAR stream which aggregates multi-scale contextual information from LIDAR point clouds. The other stream is RGB stream which is used for extracting features from RGB images. To fuse the two streams, the feature maps of RGB stream are converted to a bird-view representation to concatenate with that of LIDAR stream. In this way, the two kinds of data can complement each other for detecting road. To verify the efficacy of our TSF-FCN, experiments are carried on KITTI- ROAD benchmark and competitive performance is achieved compared with state-of-the-art methods. Ziyi Liu 0001, Jingmin Xin, Nanning Zheng 0001 |
Intelligent Vehicles Symposium | 2 |
| 2018 | Joint Video Object Discovery and Segmentation by Coupled Dynamic Markov NetworksabstractIt is a challenging task to extract segmentation mask of a target from a single noisy video, which involves object discovery coupled with segmentation. To solve this challenge, we present a method to jointly discover and segment an object from a noisy video, where the target disappears intermittently throughout the video. Previous methods either only fulfill video object discovery, or video object segmentation presuming the existence of the object in each frame. We argue that jointly conducting the two tasks in a unified way will be beneficial. In other words, video object discovery and video object segmentation tasks can facilitate each other. To validate this hypothesis, we propose a principled probabilistic model, where two dynamic Markov networks are coupled-one for discovery and the other for segmentation. When conducting the Bayesian inference on this model using belief propagation, the bi-directional message passing reveals a clear collaboration between these two inference tasks. We validated our proposed method in five data sets. The first three video data sets, i.e., the SegTrack data set, the YouTube-objects data set, and the Davis data set, are not noisy, where all video frames contain the objects. The two noisy data sets, i.e., the XJTU-Stevens data set, and the Noisy-ViDiSeg data set, newly introduced in this paper, both have many frames that do not contain the objects. When compared with state of the art, it is shown that although our method produces inferior results on video data sets without noisy frames, we are able to obtain better results on video data sets with noisy frames. Ziyi Liu 0001, Le Wang 0003, Gang Hua 0001, Qilin Zhang 0004, Zhenxing Niu, Ying Wu 0001, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 1 |
| 2017 | Hybrid-augmented intelligence: collaboration and cognitionabstractThe long-term goal of artificial intelligence (AI) is to make machines learn and think like human beings. Due to the high levels of uncertainty and vulnerability in human life and the open-ended nature of problems that humans are facing, no matter how intelligent machines are, they are unable to completely replace humans. Therefore, it is necessary to introduce human cognitive capabilities or human-like cognitive models into AI systems to develop a new form of AI, that is, hybrid-augmented intelligence. This form of AI or machine intelligence is a feasible and important developing model. Hybrid-augmented intelligence can be divided into two basic models: one is human-in-the-loop augmented intelligence with human-computer collaboration, and the other is cognitive computing based augmented intelligence, in which a cognitive model is embedded in the machine learning system. This survey describes a basic framework for human-computer collaborative hybrid-augmented intelligence, and the basic elements of hybrid-augmented intelligence based on cognitive computing. These elements include intuitive reasoning, causal models, evolution of memory and knowledge, especially the role and basic principles of intuitive reasoning for complex problem solving, and the cognitive learning framework for visual scene understanding based on memory and reasoning. Several typical applications of hybrid-augmented intelligence in related fields are given. Nanning Zheng 0001, Ziyi Liu 0001, Pengju Ren, Shi-tao Chen, Si-yu Yu, Jianru Xue, Badong Chen, Fei-Yue Wang 0001 |
Frontiers Inf. Technol. Electron. Eng. | 2 |