VLDB 2026 Research / reviewers in the wild / expert
Jiang Liu 0014
dblp:23/108-14
· DBLP profile ↗
15ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0001-7568-2454ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mesh-Gait: A Unified Framework for Gait Recognition Through Multi-Modal Representation Learning from 2D Silhouettes
Zhao-Yang Wang, Jieneng Chen, Yuxiang Guo 0001, Jiang Liu 0014, Rama Chellappa |
FG | 4 |
| 2026 | DiffProtect: Generative adversarial examples using diffusion models for facial privacy protectionabstractThe increasingly pervasive facial recognition (FR) systems raise serious concerns about personal privacy, especially for billions of users who have publicly shared their photos on social media. To address this challenge, several adversarial attack methods have been proposed to protect individuals from being identified by unauthorized FR systems with perturbed facial images. However, these approaches suffer from poor visual quality or low attack success rates, which limit their practical utility. Recently, diffusion models have achieved tremendous success in image generation. In this work, we ask: can diffusion models be used to generate adversarial examples against FR systems to improve both visual quality and attack performance? We propose DiffProtect, a novel method leveraging a diffusion autoencoder to generate semantically meaningful perturbations on FR systems. Extensive experiments demonstrate that DiffProtect produces more natural-looking encrypted images than state-of-the-art methods while achieving significantly higher attack success rates, e.g. , 24.5 % and 25.1 % absolute improvements on the CelebA-HQ and FFHQ datasets. We further evaluate the effectiveness of DiffProtect in the real world using a commercial FR API and validate its usefulness in practice through a user study. Our code is available at https://github.com/joellliu/DiffProtect . Jiang Liu 0014, Chun Pong Lau 0001, Zhongliang Guo 0001, Yuxiang Guo 0001, Zhao-Yang Wang, Rama Chellappa |
Pattern Recognit. | 1 |
| 2025 | Self-Taught Agentic Long Context UnderstandingabstractYufan Zhuang, Xiaodong Yu, Jialian Wu, Ximeng Sun, Ze Wang, Jiang Liu, Yusheng Su, Jingbo Shang, Zicheng Liu, Emad Barsoum. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yufan Zhuang, Jialian Wu, Ximeng Sun, Ze Wang 0008, Jiang Liu 0014, Yusheng Su, Jingbo Shang, Zicheng Liu 0001, Emad Barsoum |
ACL (1) | 6 |
| 2025 | SoftVQ-VAE: Efficient 1-Dimensional Continuous TokenizerabstractEfficient image tokenization with high compression ratios remains a critical challenge for training generative models. We present SoftVQ-VAE, a continuous image tokenizer that leverages soft categorical posteriors to aggregate multiple codewords into each latent token, substantially increasing the representation capacity of the latent space. When applied to Transformer-based architectures, our approach compresses 256×256 and 512×512 images using as few as 32 or 64 1-dimensional tokens. Not only does SoftVQ-VAE show consistent and high-quality reconstruction, more importantly, it also achieves state-of-the-art and significantly faster image generation results across different denoising-based generative models. Remarkably, SoftVQ-VAE improves inference throughput by up to 18x for generating 256×256 images and 55x for 512×512 images while achieving competitive FID scores of 1.78 and 2.21 for SiT-XL. It also improves the training efficiency of the generative models by reducing the number of training iterations by 2.3x while maintaining comparable performance. With its fully-differentiable design and semantic-rich latent space, our experiment demonstrates that SoftVQ-VAE achieves efficient tokenization without compromising generation quality, paving the way for more efficient generative models. Code and model are released1. Hao Chen 0102, Ze Wang 0008, Xiang Li 0106, Ximeng Sun, Fangyi Chen, Jiang Liu 0014, Jindong Wang 0001, Bhiksha Raj, Zicheng Liu 0001, Emad Barsoum |
CVPR | 6 |
| 2025 | TTT-Bench: A Benchmark for Evaluating Reasoning Ability with Simple and Novel Tic-Tac-Toe-style GamesabstractLarge reasoning models (LRMs) have demonstrated impressive reasoning capabilities across a broad range of tasks including Olympiadlevel mathematical problems, indicating evidence of their complex reasoning abilities.While many reasoning benchmarks focus on the STEM domain, the ability of LRMs to reason correctly in broader task domains remains underexplored.In this work, we introduce TTT-Bench, a new benchmark that is designed to evaluate basic strategic, spatial, and logical reasoning abilities in LRMs through a suite of four two-player Tic-Tac-Toe-style games that humans can effortlessly solve from a young age.We propose a simple yet scalable programmatic approach for generating verifiable two-player game problems for TTT-Bench.Although these games are trivial for humans, they require reasoning about the intentions of the opponent, as well as the game board's spatial configurations, to ensure a win.We evaluate a diverse set of state-of-the-art LRMs, and discover that the models that excel at hard math problems frequently fail at these simple reasoning games.Further testing reveals that our evaluated reasoning models score on average ↓ 41% & ↓ 5% lower on TTT-Bench compared to MATH 500 & AIME 2024 respectively, with larger models achieving higher performance using shorter reasoning traces, where most of the models struggle on long-term strategic reasoning situations on simple and new TTT-Bench tasks. Prakamya Mishra, Jiang Liu 0014, Jialian Wu, Zicheng Liu 0001, Emad Barsoum |
EMNLP | 2 |
| 2025 | UniGait: A Unified Transformer-based Multitask Framework for Gait Analysis in the WildabstractGait recognition is a rapidly emerging and significant area of biometrics, leveraging the unique walking patterns of individuals to perform personal identification and facilitate healthcare monitoring, such as elderly care, fall detection, etc. While existing gait recognition methods perform well in indoor, or short-range environments, their effectiveness diminishes significantly when applied to unconstrained outdoor scenarios. Challenges such as environmental turbulence, occlusion, varying viewing angles contribute to this performance drop. To address these challenges and enhance gait recognition accuracy in real-world settings, while also expanding the functionality of gait features for healthcare applications, we propose a unified multitask framework called UniGait. UniGait is designed to perform a comprehensive range of gait analysis tasks, including gait recognition and estimation of gait-related human attributes. UniGait is built upon a transformer-based architecture, which leverages the power of a cross-attention mechanism to simultaneously process multiple sub-tasks. This multitask learning approach allows the model to extract more robust gait features by jointly learning gait recognition and human attribute estimation, leading to improved overall performance. We report the results of extensive experiments and analysis on large-scale, real-world datasets collected under challenging conditions, including long-range (up to 1000 meters) and high-pitch angles (including UAV-based data). The results demonstrated state-of-the-art performance, highlighting the potential of UniGait for deployment in real-world applications, making it a valuable tool for a range of biometric and healthcare monitoring scenarios. Zhao-Yang Wang, Jiang Liu 0014, Yuxiang Guo 0001, Jieneng Chen, Rama Chellappa |
FG | 2 |
| 2025 | Unleashing Hour-Scale Video Training for Long Video-Language UnderstandingabstractRecent long-form video-language understanding benchmarks have driven progress in video large multimodal models (Video-LMMs). However, the scarcity of well-annotated long videos has left the training of hour-long Video-LMMs underexplored. To close this gap, we present VideoMarathon, a large-scale hour-long video instruction-following dataset. This dataset includes around 9,700 hours of long videos sourced from diverse domains, ranging from 3 to 60 minutes per video. Specifically, it contains 3.3M high-quality QA pairs, spanning six fundamental topics: temporality, spatiality, object, action, scene, and event. Compared to existing video instruction datasets, VideoMarathon significantly extends training video durations up to 1 hour, and supports 22 diverse tasks requiring both short- and long-term video comprehension. Building on VideoMarathon, we propose Hour-LLaVA, a powerful and efficient Video-LMM for hour-scale video-language modeling. It enables hour-long video training and inference at 1-FPS sampling by leveraging a memory augmentation module, which adaptively integrates question-relevant and spatiotemporally informative semantics from the cached full video context. In our experiments, Hour-LLaVA achieves the best performance on multiple representative long video-language benchmarks, demonstrating the high quality of the VideoMarathon dataset and the superiority of the Hour-LLaVA model. Jialian Wu, Ximeng Sun, Ze Wang 0008, Jiang Liu 0014, Yusheng Su, Hao Chen 0102, Jiebo Luo 0001, Zicheng Liu 0001, Emad Barsoum |
NeurIPS | 5 |
| 2025 | GaitContour: Efficient Gait Recognition Based on a Contour-Pose RepresentationabstractGait recognition holds the promise to robustly identify subjects based on walking patterns instead of appearance information. In recent years, this field has been dominated by learning methods based on two input formats: silhouette images and sparse keypoints. Compared to image-based approaches, keypoint-based methods can achieve significantly higher efficiency due to their sparsity. However, sparsity also results in information loss, thereby reducing performance. In this work, we propose a novel, keypoint-based Contour-Pose representation, which compactly encodes both body shape and parts information. We further propose a local-to-global architecture, called GaitContour, to leverage this novel representation and efficiently compute subject embedding in two stages. The first stage consists of a local transformer that extracts features from five different body regions. The second stage then aggregates the regional features to estimate a global human gait representation. Such a design significantly reduces the complexity of the attention operation and improves both efficiency and performance. Through large scale experiments, Gait-Contour is shown to perform significantly better than previous keypoint-based methods. Furthermore, the ContourPose representation also achieves new SoTA performances on fusion-based gait recognition methods. Yuxiang Guo 0001, Anshul Shah 0001, Jiang Liu 0014, Ayush Gupta 0001, Rama Chellappa, Cheng Peng 0008 |
WACV | 3 |
| 2025 | VM-Gait: Multi-Modal 3D Representation Based on Virtual Marker for Gait RecognitionabstractGait recognition plays a vital role in biometric applications by analyzing the unique characteristics of an individ-ual's walking pattern. Methods based on 2D representations, such as silhouettes and skeletons, are increasingly being developed to learn the shape features and joint dy-namic movements. Nevertheless, the effectiveness of 2D representation-based methods is impeded by factors such as changes in viewpoint, partial occlusion, and noisy en-vironments. 3D representation-based methods can complement 2D representation-based approaches by providing more precise dynamic body shapes and motion information, along with increased robustness against changes in viewpoint and partial occlusion. However, the complex-ity of acquiring accurate 3D representations and the chal-lenges associated with extracting dynamic topological features from sequences of 3D representations hinder the de-velopment of 3D representations-based methods. In this pa-per, we present VM-Gait, a novel multi-modal gait recognition framework that harnesses the advantages of integrating both 2D and 3D representations. Furthermore, we in-troduce a new 3D representation, Virtual Marker, into gait recognition to efficiently learn topological features from 3D representations, avoiding the computational complexi-ties inherent in directly learning from 3D representations like 3D meshes or 3D point clouds. Extensive experiments demonstrate that the proposed framework effectively learns and fuses discriminative information from different gait modalities, enhancing gait recognition performance. Zhao-Yang Wang, Jiang Liu 0014, Jieneng Chen, Rama Chellappa |
WACV | 2 |
| 2024 | HyperGait: A Video-based Multitask Network for Gait Recognition and Human Attribute Estimation at Range and AltitudeabstractGait recognition is one of the mainstream approaches for identifying individuals when face information is not available. Most previous methods achieve good performance on structured indoor walking sequences with silhouettes provided. However, when these methods are applied to unconstrained outdoor sequences, a significant reduction in performance is inevitably observed due to factors such as turbulence, occlusion, view angle, and oversized clothing. To make gait recognition methods stable and effective for real-world settings, we extend gait-only-based approaches by introducing more useful biometric information such as gender, age, height, weight, and body mass index to cooperatively work with the gait recognition module. In this paper, we propose a video-based multitasking network for gait recognition and human attribute prediction at ranges of up to 1000 meters and high-pitch angles to mutually improve the robustness and accuracy of each task. Through a series of experiments on OU-MVLP and BRIAR datasets, we show that our multitasking network outperforms previous methods and provides more useful biometric information for human identification tasks. Zhao-Yang Wang, Jiang Liu 0014, Ram Prabhakar Kathirvel, Chun Pong Lau 0001, Rama Chellappa |
IJCB | 2 |
| 2023 | Interpolated Joint Space Adversarial Training for Robust and Generalizable DefensesabstractAdversarial training (AT) is considered to be one of the most reliable defenses against adversarial attacks. However, models trained with AT sacrifice standard accuracy and do not generalize well to unseen attacks. Recent works show generalization improvement with adversarial samples under unseen threat models such as on-manifold threat model or neural perceptual threat model. However, the former requires exact manifold information while the latter requires algorithm relaxation. Motivated by these considerations, we propose a novel threat model called Joint Space Threat Model (JSTM), which exploits the underlying manifold information with Normalizing Flow, ensuring that the exact manifold assumption holds. Under JSTM, we develop novel adversarial attacks and defenses. Specifically, we propose the Robust Mixup strategy in which we maximize the adversity of the interpolated images and gain robustness and prevent overfitting. Our experiments show that Interpolated Joint Space Adversarial Training (IJSAT) achieves good performance in standard accuracy, robustness, and generalization. IJSAT is also flexible and can be used as a data augmentation method to improve standard accuracy and combined with many existing AT approaches to improve robustness. We demonstrate the effectiveness of our approach on three benchmark datasets, CIFAR-10/100, OM-ImageNet and CIFAR-10-C. Chun Pong Lau 0001, Jiang Liu 0014, Hossein Souri, Wei-An Lin, Soheil Feizi, Rama Chellappa |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | One Model to Synthesize Them All: Multi-Contrast Multi-Scale Transformer for Missing Data ImputationabstractMulti-contrast magnetic resonance imaging (MRI) is widely used in clinical practice as each contrast provides complementary information. However, the availability of each imaging contrast may vary amongst patients, which poses challenges to radiologists and automated image analysis algorithms. A general approach for tackling this problem is missing data imputation, which aims to synthesize the missing contrasts from existing ones. While several convolutional neural networks (CNN) based algorithms have been proposed, they suffer from the fundamental limitations of CNN models, such as the requirement for fixed numbers of input and output channels, the inability to capture long-range dependencies, and the lack of interpretability. In this work, we formulate missing data imputation as a sequence-to-sequence learning problem and propose a multi-contrast multi-scale Transformer (MMT), which can take any subset of input contrasts and synthesize those that are missing. MMT consists of a multi-scale Transformer encoder that builds hierarchical representations of inputs combined with a multi-scale Transformer decoder that generates the outputs in a coarse-to-fine fashion. The proposed multi-contrast Swin Transformer blocks can efficiently capture intra- and inter-contrast dependencies for accurate image synthesis. Moreover, MMT is inherently interpretable as it allows us to understand the importance of each input contrast in different regions by analyzing the in-built attention maps of Transformer blocks in the decoder. Extensive experiments on two large-scale multi-contrast MRI datasets demonstrate that MMT outperforms the state-of-the-art methods quantitatively and qualitatively. Jiang Liu 0014, Srivathsa Pasumarthi, Ben A. Duffy, Enhao Gong, Keshav Datta, Greg Zaharchuk |
IEEE Trans. Medical Imaging | 1 |
| 2022 | Segment and Complete: Defending Object Detectors against Adversarial Patch Attacks with Robust Patch DetectionabstractObject detection plays a key role in many security-critical systems. Adversarial patch attacks, which are easy to implement in the physical world, pose a serious threat to state-of-the-art object detectors. Developing reliable defenses for object detectors against patch attacks is critical but severely understudied. In this paper, we propose Segment and Complete defense (SAC), a general framework for defending object detectors against patch attacks through detection and removal of adversarial patches. We first train a patch segmenter that outputs patch masks which provide pixel-level localization of adversarial patches. We then propose a self adversarial training algorithm to robustify the patch segmenter. In addition, we design a robust shape completion algorithm, which is guaranteed to remove the entire patch from the images if the outputs of the patch segmenter are within a certain Hamming distance of the ground-truth patch masks. Our experiments on COCO and xView datasets demonstrate that SAC achieves superior robustness even under strong adaptive attacks with no reduction in performance on clean images, and generalizes well to unseen patch shapes, attack budgets, and unseen attack methods. Furthermore, we present the APRICOT-Mask dataset, which augments the APRICOT dataset with pixel-level annotations of adversarial patches. We show SAC can significantly reduce the targeted attack success rate of physical patch attacks. Our code is available at https://github.com/joellliu/SegmentAndComplete. Jiang Liu 0014, Alexander Levine 0001, Chun Pong Lau 0001, Rama Chellappa, Soheil Feizi |
CVPR | 1 |
| 2022 | Mutual Adversarial Training: Learning Together is Better Than Going AloneabstractRecent studies have shown that robustness to adversarial attacks can be transferred across deep neural networks. In other words, we can make a weak model more robust with the help of a strong teacher model. In this paper, we ask if models can “learn together” and “teach each other” to achieve better robustness instead of learning from a static teacher.We study how interactions among models affect robustness via knowledge distillation. We propose mutual adversarial training (MAT), in which multiple models are trained together and share the knowledge of adversarial examples to achieve improved robustness. MAT allows robust models to explore a larger space of adversarial samples and find more robust feature spaces and decision boundaries. Through extensive experiments on the CIFAR-10, CIFAR-100, and mini-ImageNet datasets, we demonstrate that MAT can effectively improve model robustness and outperform state-of-the-art methods under white-box attacks. In addition, we show that MAT can also mitigate the robustness trade-off among different perturbation types. Specially, we train specialist models that learn to defend a specific perturbation type and a generalist model that learns to defend multiple perturbation types by learning from the specialists, which brings as much as 13.4% accuracy gain to AT baselines against the union ofl∞,l2, andl1 attacks. Our results show the superiority of the proposed method and demonstrate that collaborative learning is an effective strategy for designing robust models. Jiang Liu 0014, Chun Pong Lau 0001, Hossein Souri, Soheil Feizi, Rama Chellappa |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2018 | Left Atrial Appendage Segmentation Using Fully Convolutional Neural Networks and Modified Three-Dimensional Conditional Random FieldsabstractThrombosis has become a global disease threatening human health. The left atrial appendage (LAA) is a major source of thrombosis in patients with atrial fibrillation (AF). Positive correlation exists between LAA volume and AF risk. LAA morphology has been suggested to influence thromboembolic risk in AF patients and to help predict thromboembolic events in low-risk patient groups. Automatic segmentation of LAA can greatly help physicians diagnose AF. In consideration of the large anatomical variations of the LAA, we proposed a robust method for automatic LAA segmentation on computed tomographic angiography (CTA) data using fully convolutional neural networks with three-dimensional (3-D) conditional random fields (CRFs). After manual localization of ROI of LAA, we adopted the FCN in natural image segmentation and transferred their learned models by fine-tuning the networks to segment each 2-D LAA slice. Subsequently, we used a modified dense 3-D CRF that accounts for the 3-D spatial information and larger contextual information to refine the segmentations of all slices. Our method was evaluated on 150 sets of CTA data using five-fold cross validation. Compared with manual annotation, we obtained a mean dice overlap of and a mean volume overlap of with a computation time of less than 40 s per volume. Experimental results demonstrated the robustness of our method in dealing with large anatomical variations and computational efficiency for adoption in a daily clinical routine.). Cheng Jin 0007, Jianjiang Feng, Heng Yu 0005, Jiang Liu 0014, Jiwen Lu, Jie Zhou 0001 |
IEEE J. Biomed. Health Informatics | 5 |