VLDB 2026 Research / reviewers in the wild / expert
Zijia Lu
dblp:241/7369
· DBLP profile ↗
9ranked-venue papers
7as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 7 first-author · 7 since 2021Artificial intelligence and machine learning · 8 · 7 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DeCafNet: Delegate and Conquer for Efficient Temporal Grounding in Long VideosabstractLong Video Temporal Grounding (LVTG) aims at identifying specific moments within lengthy videos based on user-provided text queries for effective content retrieval. The approach taken by existing methods of dividing video into clips and processing each clip via a full-scale expert encoder is challenging to scale due to prohibitive computational costs of processing a large number of clips in long videos. To address this issue, we introduce DeCafNet, an approach employing "delegate-and-conquer" strategy to achieve computation efficiency without sacrificing grounding performance. DeCafNet introduces a sidekick encoder that performs dense feature extraction over all video clips in a resource-efficient manner, while generating a saliency map to identify the most relevant clips for full processing by the expert encoder. To effectively leverage features from sidekick and expert encoders that exist at different temporal resolutions, we introduce DeCaf-Grounder, which unifies and refines them via query-aware temporal aggregation and multi-scale temporal refinement for accurate grounding. Experiments on two LTVG benchmark datasets demonstrate that DeCafNet reduces computation by up to 47% while still outperforming existing methods, establishing a new state-of-the-art for LTVG in terms of both efficiency and performance. Zijia Lu, A S. M. Iftekhar, Gaurav Mittal, Tianjian Meng, Xiawei Wang, Rohith Kukkala, Ehsan Elhamifar |
CVPR | 1 |
| 2025 | Multi-Modal Few-Shot Temporal Action Segmentation
Zijia Lu, Ehsan Elhamifar |
ICCV | 1 |
| 2024 | Error Detection in Egocentric Procedural Task VideosabstractWe present a new egocentric procedural error dataset containing videos with various types of errors as well as normal videos and propose a new framework for procedural error detection using error-free training videos only. Our framework consists of an action segmentation model and a contrastive step prototype learning module to segment actions and learn useful features for error detection. Based on the observation that interactions between hands and objects often inform action and error understanding, we propose to combine holistic frame features with relations features, which we learn by building a graph using active object detection followed by a Graph Convolutional Network. To handle errors, unseen during training, we use our contrastive step prototype learning to learn multiple prototypes for each step, capturing variations of error-free step executions. At inference time, we use feature-prototype similarities for error detection. By experiments on three datasets, we show that our proposed framework outperforms state-of-the-art video anomaly detection methods for error detection and provides smooth action and error predictions.11Code and data is available at https://github.com/robert80203/EgoPER_official Shih-Po Lee, Zijia Lu, Minh Hoai, Ehsan Elhamifar |
CVPR | 2 |
| 2024 | FACT: Frame-Action Cross-Attention Temporal Modeling for Efficient Action SegmentationabstractWe study supervised action segmentation, whose goal is to predict framewise action labels of a video. To capture tem-poral dependencies over long horizons, prior works either improve framewise features with transformer or refine frame-wise predictions with learned action features. However, they are computationally costly and ignore that frame and action features contain complimentary information, which can be leveraged to enhance both features and improve temporal modeling. Therefore, we propose an efficient Frame-Action Cross-attention Temporal modeling (FACT) framework that performs temporal modeling withframe and action features in parallel and leverage this parallelism to achieve iterative bidirectional information transfer between the features and refine them. FACT network contains (i) aframe branch to learn frame-level information with convolutions and frame features, (ii) an action branch to learn action-level depen-dencies with transformers and action tokens and (iii) cross-attentions to allow communication between the two branches. We also propose a new matching loss to ensure each action to-ken uniquely encodes an action segment, thus better captures its semantics. Thanks to our architecture, we can also lever-age textual transcripts of videos to help action segmentation. We evaluate FACT on four video datasets (two egocentric and two third-person) for action segmentation with and without transcripts, showing that it significantly improves the state-of-the-art accuracy while enjoys lower computational cost (3 times faster) than existing transformer-based methods.11Code available at github.com/ZijiaLewisLu/CVPR2024-FACT. Zijia Lu, Ehsan Elhamifar |
CVPR | 1 |
| 2024 | Self-Supervised Multi-Object Tracking with Path ConsistencyabstractIn this paper, we propose a novel concept of path consis-tency to learn robust object matching without using manual object identity supervision. Our key idea is that, to track a object through frames, we can obtain multiple different as-sociation results from a model by varying the frames it can observe, i.e., skipping frames in observation. As the differ-ences in observations do not alter the identities of objects, the obtained association results should be consistent. Based on this rationale, we generate multiple observation paths, each specifying a different set of frames to be skipped, and formulate the Path Consistency Loss that enforces the as-sociation results are consistent across different observation paths. We use the proposed loss to train our object matching model with only self-supervision. By extensive experiments on three tracking datasets (MOT17, PersonPath22, KITTI), we demonstrate that our method outperforms existing unsu-pervised methods with consistent margins on various eval-uation metrics, and even achieves performance close to su-pervised methods. Zijia Lu, Bing Shuai, Yanbei Chen, Zhenlin Xu, Davide Modolo |
CVPR | 1 |
| 2022 | Set-Supervised Action Learning in Procedural Task Videos via Pairwise Order ConsistencyabstractWe address the problem of set-supervised action learning, whose goal is to learn an action segmentation model using weak supervision in the form of sets of actions occurring in training videos. Our key observation is that videos within the same task have similar ordering of actions, which can be leveraged for effective learning. Therefore, we propose an attention-based method with a new Pairwise Ordering Consistency (POC) loss that encourages that for each common action pair in two videos of the same task, the attentions of actions follow a similar ordering. Unlike existing sequence alignment methods, which misalign actions in videos with different orderings or cannot reliably separate more from less consistent orderings, our POC loss efficiently aligns videos with different action orders and is differentiable, which enables end-to-end training. In addition, it avoids the time-consuming pseudo-label generation of prior works. Our method efficiently learns the actions and their temporal locations, therefore, extends the existing attention-based action localization methods from learning one action per video to multiple actions using our POC loss along with video-level and frame-level losses. By experiments on three datasets, we demonstrate that our method significantly improves the state of the art. We also show that our method, with a small modification, can effectively address the transcript-supervised action learning task, where actions and their ordering are available during training.11Code available at https://github.com/ZijiaLewisLu/CVPR22-POC. Zijia Lu, Ehsan Elhamifar |
CVPR | 1 |
| 2021 | Weakly-Supervised Action Segmentation and Alignment via Transcript-Aware Union-of-Subspaces LearningabstractWe address the problem of learning to segment actions from weakly-annotated videos, i.e., videos accompanied by transcripts (ordered list of actions). We propose a framework in which we model actions with a union of low-dimensional subspaces, learn the subspaces using transcripts and refine video features that lend themselves to action subspaces. To do so, we design an architecture consisting of a Union-of-Subspaces Network, which is an ensemble of autoencoders, each modeling a low-dimensional action subspace and can capture variations of an action within and across videos. For learning, at each iteration, we generate positive and negative soft alignment matrices using the segmentations from the previous iteration, which we use for discriminative training of our model. To regularize the learning, we introduce a constraint loss that prevents imbalanced segmentations and enforces relatively similar duration of each action across videos. To have a real-time inference, we develop a hierarchical segmentation framework that uses subset selection to find representative transcripts and hierarchically align a test video with increasingly refined representative transcripts. Our experiments on three datasets show that our method improves the state-of-the-art action segmentation and alignment, while speeding up the inference time by a factor of 4 to 13.1 Zijia Lu, Ehsan Elhamifar |
ICCV | 1 |
| 2019 | DFT-Net: Disentanglement of Face Deformation and Texture Synthesis for Expression EditingabstractThis paper presents a novel deep architecture DFT-Net that combines the advantages of Generative Adversarial Networks (GANs) and warp mechanisms for expression editing. Recent generative models leverage Action Units as annotations and show more flexible expression manipulation than previous approaches using other guiding information. However, those methods bring inevitable artifacts where facial components deform (e.g. eyes from open to close), for the structural defect in modeling shape variations without geometric guidance such as facial landmarks. Our approach explicitly disentangles face deformations and appearance details by constructing two parallel networks, one that learns an appearance flow for 2D warps and the other generates corresponding texture and hallucinates hidden regions such as mouth interiors. Experimental results show our method outperforms the state-of-the-art on various expression editing tasks. Jie Zhang 0071, Zijia Lu, Shiguang Shan |
ICIP | 3 |
| 2018 | Zero-Shot Facial Expression Recognition with Multi-label Label Propagation
Zijia Lu, Jiabei Zeng, Shiguang Shan, Xilin Chen 0001 |
ACCV (3) | 1 |