VLDB 2026 Research / reviewers in the wild / expert
Yang Liu 0084
dblp:51/3710-84
· DBLP profile ↗
29ranked-venue papers
10as first author
24since 2021 · last 2026
0000-0002-9423-9252ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 8 first-author · 12 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 9 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Structure-preserving contrastive graph clustering with dual-channel label alignmentabstractThe past few years have witnessed the rapid development of contrastive graph clustering (CGC). Although a series of achievements have been made, there still remain two challenging problems in the literature. First, previous works typically generate different views via some pre-defined graph augmentation strategies, but inappropriate augmentations may alter the latent semantics of the original data. Second, they often overlook the discriminative unsupervised information when constructing positive and negative sample pairs, resulting in compromised clustering performance. Third, some of them are restricted to only static neighborhood connections for contrastive learning, which neglect the dynamical structural relationship via robust neighboring graph learning. To cope with these issues, this paper proposes a Structure-preserving Contrastive Graph Clustering approach with Dual-channel Label Alignment (SCGC-DLA). In terms of the high-and-low frequency issues, the low-pass and hybrid graph filters are designed for generating two views of reliable augmentations, which can supply rich and complementary information to each other. Further, we construct a structure-preserving matrix, which is derived from the edge betweenness centrality (EBC) perspective design and allows us to efficiently capture the topological relationships among different embedding representations. Under the guidance of the non-dominated sorting theory, the clustering distribution information of dual-channel is used to construct high-confidence pseudo labels. Especially, the generated high-confidence pseudo labels are aligned with latent semantic labels. Finally, the overall network is guided by a self-supervised learning scheme and therefore the final clustering could be obtained. Substantial results on five benchmarks prove the robustness and effectiveness of our approach compared to several state-of-the-arts. Yan-Di Huang, Dong Huang 0001, Chang-Dong Wang 0001, Yang Liu 0084, Enbo Huang |
Neural Networks | 5 |
| 2026 | In-Context Model Predictive Generation: Open-Vocabulary Motion Synthesis From Language Models to PhysicsabstractSynthesizing human motion from textual descriptions is essential for immersive digital applications, yet existing methods face a persistent trade-off between semantic fidelity and physical realism. Large language model (LLM)-based approaches can interpret diverse open-vocabulary instructions and compose high-level action plans, but they often generate motions that violate physical constraints. Physics-aware models improve realism through simulation or control, but they struggle with semantic complexity, fine-grained instructions, and novel concepts. To address this gap, we propose In-Context Model Predictive Generation (ICMPG), a framework that integrates language-model planning with inference-time physical feedback. ICMPG reformulates motion synthesis as a Model Predictive Control (MPC)-like process with two modules. The Context-Aware Motion Generation (CAMG) module uses an LLM as a planner to decompose textual commands and generate candidate motion sequences from motion tokens. The Model Predictive Generation (MPG) module evaluates these candidates through physical simulation and semantic alignment, estimates a composite reward, and selects the best sequence to guide subsequent generation steps. Unlike open-loop generation, this closed-loop refinement enables ICMPG to adapt motions to both the input semantics and the simulated physical environment without task-specific policy retraining. Extensive experiments across standard and zero-shot open-vocabulary settings show that ICMPG generalizes robustly to diverse commands and produces motions that are more physically plausible and semantically faithful than representative baselines on the evaluated benchmarks. The framework bridges semantic interpretation and physical simulation while remaining flexible enough to incorporate different LLM backbones, enabling more versatile and controllable text-driven motion synthesis. Xiaomeng Fu, Junfan Lin, Yang Liu 0084, Yaowei Wang 0001, Guanbin Li, Liang Lin 0004, Ziliang Chen 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | Cross-modal Causal Relation Alignment for Video Question GroundingabstractVideo question grounding (VideoQG) requires models to answer the questions and simultaneously infer the relevant video segments to support the answers. However, existing VideoQG methods usually suffer from spurious cross-modal correlations, leading to a failure to identify the dominant visual scenes that align with the intended question. Moreover, vision-language models exhibit unfaithful generalization performance and lack robustness on challenging downstream tasks such as VideoQG. In this work, we propose a novel VideoQG framework named Cross-modal Causal Relation Alignment (CRA), to eliminate spurious correlations and improve the causal consistency between question-answering and video temporal grounding. Our CRA involves three essential components: i) Gaussian Smoothing Grounding (GSG) module for estimating the time interval via cross-modal attention, which is de-noised by an adaptive Gaussian filter, ii) Cross-Modal Alignment (CMA) enhances the performance of weakly supervised VideoQG by leveraging bidirectional contrastive learning between estimated video segments and QA features, iii) Explicit Causal Intervention (ECI) module for multimodal deconfounding, which involves front-door intervention for vision and backdoor intervention for language. Extensive experiments on two VideoQG datasets demonstrate the superiority of our CRA in discovering visually grounded content and achieving robust question reasoning. Codes are available at https://github.com/WissingChen/CRA-GQA. Yang Liu 0084, Binglin Chen, Jiandong Su, Yongsen Zheng, Liang Lin 0004 |
CVPR | 2 |
| 2025 | Towards Long-Horizon Vision-Language Navigation: Platform, Benchmark and MethodabstractExisting Vision-Language Navigation (VLN) methods primarily focus on single-stage navigation, limiting their effectiveness in multi-stage and long-horizon tasks within complex and dynamic environments. To address these limitations, we propose a novel VLN task, named Long-Horizon Vision-Language Navigation (LH-VLN), which emphasizes long-term planning and decision consistency across consecutive subtasks. Furthermore, to support LH-VLN, we develop an automated data generation platform NavGen, which constructs datasets with complex task structures and improves data utility through a bidirectional, multi-granularity generation approach. To accurately evaluate complex tasks, we construct the Long-Horizon Planning and Reasoning in VLN (LHPR-VLN) benchmark consisting of 3,260 tasks with an average of 150 task steps, serving as the first dataset specifically designed for the long-horizon vision-language navigation task. Furthermore, we propose Independent Success Rate (ISR), Conditional Success Rate (CSR), and CSR weight by Ground Truth (CGT) metrics, to provide fine-grained assessments of task completion. To improve model adaptability in complex tasks, we propose a novel Multi-Granularity Dynamic Memory (MGDM) module that integrates short-term memory blurring with long-term memory retrieval to enable flexible navigation in dynamic environments. Our platform, benchmark and method supply LH-VLN with a robust data generation pipeline, comprehensive model evaluation dataset, reasonable metrics, and a novel VLN model, establishing a foundational framework for advancing LH-VLN. Xinshuai Song, Yang Liu 0084, Weikai Chen 0001, Guanbin Li, Liang Lin 0004 |
CVPR | 3 |
| 2025 | Dual-Level Facilitated Multi-View Contrastive Graph ClusteringabstractMulti-view attributed graph clustering (MAGC) has recently experienced impressive attention in the graph exploration literature. Although several excellent achievements have been made, previous MVAGC approaches merely consider the homogeneous information across different views, easily resulting in the compromised results when faced with the heterogeneous graph scenarios. Further, many of them rely on the static neighborhood connection from original attributed graphs, which ignores the dynamical structural relationship for enhancing cross-view contrastive learning. To deal with these drawbacks, this paper derives a Dual-level Facilitated Multi-view Contrastive Graph Clustering (DF-MCGC) approach. Specifically, we design a hybrid graph filter by considering the homogeneity hidden in individual view. Further, the view-consistent topology invariant matrix is derived to exploit the topological relationship among different embedded representations. This design progressively helps to construct the cluster-wise sample pairs for cross-view contrastive learning. Especially, the overall network is incorporated with a dual-level self-supervised paradigm, among which the self-supervised signals can be efficiently facilitated in a mutually enhanced manner. Experiments on heterogeneous benchmarks have confirmed the advantages of our DF-MCGC approach in comparison with the advanced competitors. Zi-Ying Li, Dong Huang 0001, Chang-Dong Wang 0001, Yang Liu 0084 |
ICPADS | 6 |
| 2025 | 3DAffordSplat: Efficient Affordance Reasoning with 3D Gaussiansabstract3D affordance reasoning plays a critical role in associating human instructions with the functional regions of 3D objects, facilitating precise, task-oriented manipulations in embodied AI. However, current methods, which predominantly depend on sparse 3D point clouds, exhibit limited generalizability and robustness due to their sensitivity to coordinate variations and the inherent sparsity of the data. By contrast, 3D Gaussian Splatting (3DGS) delivers high-fidelity, real-time rendering with minimal computational overhead by representing scenes as dense, continuous distributions. This positions 3DGS as a highly effective approach for capturing fine-grained affordance details and improving recognition accuracy. Nevertheless, its full potential remains largely untapped due to the absence of large-scale, 3DGS-specific affordance datasets. To overcome these limitations, we present 3DAffordSplat, the first large-scale, multi-modal dataset tailored for 3DGS-based affordance reasoning. This dataset includes 23,672 Gaussian instances, 8,231 point cloud instances, and 6,631 manually annotated affordance labels, encompassing 21 object categories and 18 affordance types. Building upon this dataset, we introduce AffordSplatNet, a novel model specifically designed for affordance reasoning using 3DGS representations. AffordSplatNet features an innovative cross-modal structure alignment module that exploits structural consistency priors to align 3D point cloud and 3DGS representations, resulting in enhanced affordance recognition accuracy. Extensive experiments demonstrate that the 3DAffordSplat dataset significantly advances affordance learning within the 3DGS domain, while AffordSplatNet consistently outperforms existing methods across both seen and unseen settings, highlighting its robust generalization capabilities. Code, model, and video are available at https://hcplab-sysu.github.io/3DAffordSplat. Zeming Wei, Yang Liu 0084, Jingzhou Luo, Guanbin Li, Liang Lin 0004 |
ACM Multimedia | 3 |
| 2025 | Learn 3D VQA Better with Active Selection and Reannotation
Yang Liu 0084, Feng Zheng 0001 |
ACM Multimedia | 2 |
| 2025 | Quadratic Coreset Selection: Certifying and Reconciling Sequence and Token Mining for Efficient Instruction TuningabstractInstruction-Tuning (IT) was recently found the impressive data efficiency in post-training large language models (LLMs). While the pursuit of efficiency predominantly focuses on sequence-level curation, often overlooking the nuanced impact of critical tokens and the inherent risks of token noise and biases. Drawing inspiration from bi-level coreset selection, our work provides the principled view of the motivation behind selecting instructions' responses. It leads to our approach Quadratic Coreset Selection (QCS) that reconciles sequence-level and token-level influence contributions, deriving more expressive LLMs with established theoretical result. Despite the original QCS framework challenged by prohibitive computation from inverted LLM-scale Hessian matrices, we overcome this barrier by proposing a novel QCS probabilistic variant, which relaxes the original formulation through re-parameterized densities. This innovative solver is efficiently learned using hierarchical policy gradients without requiring back-propagation, achieving provable convergence and certified asymptotic equivalence to the original objective. Our experiments demonstrate QCS's superior sequence-level data efficiency and reveal how strategically leveraging token-level influence elevates the performance ceiling of data-efficient IT. Furthermore, QCS's adaptability is showcased through its successes in regular IT and challenging targeted IT scenarios, particularly in the cases of free-form complex instruction-following and CoT reasoning. They underscore QCS's potential for a wide array of versatile post-training applications. Ziliang Chen 0001, Yongsen Zheng, Zhao-Rong Lai, Zhanfu Yang, Cuixi Li, Yang Liu 0084, Liang Lin 0004 |
NeurIPS | 6 |
| 2025 | MM-OPERA: Benchmarking Open-ended Association Reasoning for Large Vision-Language ModelsabstractLarge Vision-Language Models (LVLMs) have exhibited remarkable progress. However, deficiencies remain compared to human intelligence, such as hallucination and shallow pattern matching. In this work, we aim to evaluate a fundamental yet underexplored intelligence: association, a cornerstone of human cognition for creative thinking and knowledge integration. Current benchmarks, often limited to closed-ended tasks, fail to capture the complexity of open-ended association reasoning vital for real-world applications. To address this, we present MM-OPERA, a systematic benchmark with 11,497 instances across two open-ended tasks: Remote-Item Association (RIA) and In-Context Association (ICA), aligning association intelligence evaluation with human psychometric principles. It challenges LVLMs to resemble the spirit of divergent thinking and convergent associative reasoning through free-form responses and explicit reasoning paths. We deploy tailored LLM-as-a-Judge strategies to evaluate open-ended outputs, applying process-reward-informed judgment to dissect reasoning with precision. Extensive empirical studies on state-of-the-art LVLMs, including sensitivity analysis of task instances, validity analysis of LLM-as-a-Judge strategies, and diversity analysis across abilities, domains, languages, cultures, etc., provide a comprehensive and nuanced understanding of the limitations of current LVLMs in associative reasoning, paving the way for more human-like and general-purpose AI. The dataset and code are available at https://github.com/MM-OPERA-Bench/MM-OPERA. Zimeng Huang, Jinxin Ke, Xiaoxuan Fan, Yang Liu 0084, Liu Zhonghan, Zedi Wang, Junteng Dai, Haoyi Jiang, Yuyu Zhou, Keze Wang, Ziliang Chen 0001 |
NeurIPS | 5 |
| 2025 | DAPO : Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage-Based Policy OptimizationabstractThe role of reinforcement learning (RL) in enhancing the reasoning of large language models (LLMs) is becoming increasingly significant. Despite the success of RL in many scenarios, there are still many challenges in improving the reasoning of LLMs. One key challenge is the sparse reward, which introduces more training variance in policy optimization and makes it difficult to obtain a good estimation for value function in Actor-Critic (AC) methods. To address these issues, we introduce Direct Advantage-Based Policy Optimization (DAPO), a novel step-level offline RL algorithm with theoretical guarantees for enhancing the reasoning abilities of LLMs. Unlike response-level methods (such as DPO and GRPO) that the update directions of all reasoning steps are governed by the outcome reward uniformly, DAPO employs a critic function to provide step-level dense signals for policy optimization. Additionally, the actor and critic in DAPO are trained independently, ensuring that critic is a good estimation of true state value function and avoiding the co-training instability observed in standard AC methods. We train DAPO on mathematical and code problems and then evaluate its performance on multiple benchmarks. Our results show that DAPO can effectively enhance the mathematical and code capabilities on both SFT models and RL models, demonstrating the effectiveness of DAPO. Jiacai Liu, Chaojie Wang 0001, Chris Yuhao Liu, Rui Yan 0010, Yang Liu 0084 |
NeurIPS | 7 |
| 2025 | Incentivizing LLMs to Self-Verify Their AnswersabstractLarge Language Models (LLMs) have demonstrated remarkable progress in complex reasoning tasks through both post-training and test-time scaling laws. While prevalent test-time scaling approaches are often realized by using external reward models to guide the model generation process, we find that only marginal gains can be acquired when scaling a model post-trained on specific reasoning tasks. We identify that the limited improvement stems from distribution discrepancies between the specific post-trained generator and the general reward model. To address this, we propose a framework that incentivizes LLMs to self-verify their own answers. By unifying answer generation and verification within a single reinforcement learning (RL) process, we train models that can effectively assess the correctness of their own solutions. The trained model can further scale its performance at inference time by verifying its generations, without the need for external verifiers. We train our self-verification models based on Qwen2.5-Math-7B and DeepSeek-R1-Distill-Qwen-1.5B, demonstrating their capabilities across varying reasoning context lengths. Experiments on multiple mathematical reasoning benchmarks show that our models can not only improve post-training performance but also enable effective test-time scaling. Our code is available at https://github.com/mansicer/self-verification. Fuxiang Zhang, Chaojie Wang 0001, Ce Cui, Yang Liu 0084, Bo An 0001 |
NeurIPS | 5 |
| 2025 | Cross-Modal Causal Representation Learning for Radiology Report GenerationabstractRadiology Report Generation (RRG) is essential for computer-aided diagnosis and medication guidance, which can relieve the heavy burden of radiologists by automatically generating the corresponding radiology reports according to the given radiology image. However, generating accurate lesion descriptions remains challenging due to spurious correlations from visual-linguistic biases and inherent limitations of radiological imaging, such as low resolution and noise interference. To address these issues, we propose a two-stage framework named Cross-Modal Causal Representation Learning (CMCRL), consisting of the Radiological Cross-modal Alignment and Reconstruction Enhanced (RadCARE) pre-training and the Visual-Linguistic Causal Intervention (VLCI) fine-tuning. In the pre-training stage, RadCARE introduces a degradation-aware masked image restoration strategy tailored for radiological images, which reconstructs high-resolution patches from low-resolution inputs to mitigate noise and detail loss. Combined with a multiway architecture and four adaptive training strategies (e.g., text postfix generation with degraded images and text prefixes), RadCARE establishes robust cross-modal correlations even with incomplete data. In the VLCI phase, we deploy causal front-door intervention through two modules: the Visual Deconfounding Module (VDM) disentangles local-global features without fine-grained annotations, while the Linguistic Deconfounding Module (LDM) eliminates context bias without external terminology databases. Experiments on IU-Xray and MIMIC-CXR show that our CMCRL pipeline significantly outperforms state-of-the-art methods, with ablation studies confirming the necessity of both stages. Code and models are available at https://github.com/WissingChen/CMCRL. Yang Liu 0084, Ce Wang 0001, Jiarui Zhu, Guanbin Li, Cheng-Lin Liu 0001, Liang Lin 0004 |
IEEE Trans. Image Process. | 2 |
| 2025 | ODMixer: Fine-Grained Spatial-Temporal MLP for Metro Origin-Destination PredictionabstractMetro Origin-Destination (OD) prediction is a crucial yet challenging spatial-temporal prediction task in urban computing, which aims to accurately forecast cross-station ridership for optimizing metro scheduling and enhancing overall transport efficiency. Analyzing fine-grained and comprehensive relations among stations effectively is imperative for metro OD prediction. However, existing metro OD models either mix information from multiple OD pairs from the station's perspective or exclusively focus on a subset of OD pairs. These approaches may overlook fine-grained relations among OD pairs, leading to difficulties in predicting potential anomalous conditions. To address these challenges, we learn traffic evolution from the perspective of all OD pairs and propose a fine-grained spatialtemporal MLP architecture for metro OD prediction, namely ODMixer. Specifically, our ODMixer has double-branch structure and involves the Channel Mixer, the Multi-view Mixer, and the Bidirectional Trend Learner. The Channel Mixer aims to capture short-term temporal relations among OD pairs, the Multi-view Mixer concentrates on capturing spatial relations from both origin and destination perspectives. To model long-term temporal relations, we introduce the Bidirectional Trend Learner. Extensive experiments on two large-scale metro OD prediction datasets HZMOD and SHMO demonstrate the advantages of our ODMixer. Our code is available at https://github.com/KLatitude/ODMixer Yang Liu 0084, Binglin Chen, Yongsen Zheng, Lechao Cheng, Guanbin Li, Liang Lin 0004 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2024 | Confidence-oriented Contrastive Graph ClusteringabstractContrastive clustering has recently been an emerging topic in deep unsupervised learning. Nevertheless, the previous works mostly adopt the stochastic data augmentations, which easily leads to the semantic drift problem by limited transformations. Moreover, these approaches ignore the data distribution information when generating the positive and negative pair-wise samples. In light of this, this paper proposes a simple yet effective unsupervised clustering network termed Confidence-oriented Contrastive Graph Clustering (CoCGC). Particularly, we design an end-to-end network paradigm with un-shared weights, among which a hybrid graph filter is utilized to generate two views of reliable augmentations. Guided by the non-dominated sorting theory, we further construct a confidence-oriented sample set from the latent data distribution perspective. By considering the local density and cluster distribution of the embedding representations, the discriminative sample pairs can be derived from the confidence-oriented sets in a two-view contrastive manner. Finally, a cross-view neighbor contrastive loss is devised for better exploiting the self-supervised network signals. Extensive experimental results on five benchmark datasets demonstrate the effectiveness of our method against the existing state-of-the-art deep graph clustering methods. Yan-Di Huang, Dong Huang 0001, Chang-Dong Wang 0001, Yang Liu 0084, Enbo Huang |
IJCNN | 5 |
| 2024 | Diversity Matters: User-Centric Multi-Interest Learning for Conversational Movie RecommendationabstractDiversity plays a crucial role in Recommender Systems (RSs) as it ensures a wide range of recommended items, providing users with access to new and varied options. Without diversity, users often encounter repetitive content, limiting their exposure to novel choices. While significant efforts begin to enhance recommendation diversification in static offline scenarios, relatively less attention has been given to online Conversational Recommender Systems (CRSs). However, the lack of recommendation diversity in CRSs will increasingly exacerbate over time due to the dynamic user-system feedback loop, resulting in challenges such as the Matthew effect, filter bubbles, and echo chambers. To address these issues, we propose a novel paradigm, User-Centric Multi-Interest Learning for Conversational Movie Recommendation (CoMoRec), aiming to learn multiple user interests to improve result diversity for movie recommendations. Firstly, CoMoRec automatically models various facets of user interests, including context-, graph-, and review-based interests, to explore a wide range of user potential intentions. Then, it leverages these multi-aspect user interests to accurately predict personalized and diverse movie recommendations and generate fluent and informative responses during conversations. Extensive experiments on two publicly CRS-based movie datasets show that our CoMoRec achieves a new state-of-the-art performance and the superiority of improving recommendation diversity in the CRS. Yongsen Zheng, Guohua Wang 0005, Yang Liu 0084, Liang Lin 0004 |
ACM Multimedia | 3 |
| 2024 | Progressive Multi-Iteration Registration-Fusion Co-Optimization Network for Unregistered Hyperspectral Image Super-ResolutionabstractExisting fusion-based hyperspectral image super-resolution (fusion-based HSI-SR) methods usually reconstruct high-resolution hyperspectral image (HR-HSI) by integrating the complementary information of low-resolution hyperspectral image (LR-HSI) and high-resolution multispectral image (HR-MSI). However, most of such methods rely on accurately registered images or consider registration and fusion as a two-stage task, which means that fusion must tolerate the accumulation of errors due to misregistration. In this paper, we propose a progressive multi-iteration registration-fusion co-optimization network (PMI-RFCoNet) for unregistered hyperspectral image super-resolution, which progressively refines the registration and fusion result over multiple levels to reconstruct registered HR-HSI. To achieve registration-fusion co-optimization, the registration-fusion cooptimization block (Co-RFB) is designed to iterate continuously over multiple levels. We embed the interactive registration module (IRM) and the spectral recalibration and fusion module (SRFU) in Co-RFB, which can facilitate the network utilizing spatial and spectral features at different levels to generate more accurate HR-HSI. Specifically, IRM generates deformation field based on spatial correlations captured at long distances to repair non-rigid pixel offsets, and SRFU further performs adaptive high-fidelity spectral correction and spatial information fusion on the registration results. We conduct experimental verification on four widely used datasets, and the results show that PMI-RFCoNet can flexibly cope with different types and degrees of non-rigid deformation and achieve superior performance. Code is available at https://github.com/Jiahuiqu/PMI-RFCoNet. Jiahui Qu, Xuyao Liu, Wenqian Dong, Yang Liu 0084, Tongzhen Zhang, Yang Xu 0070, Yunsong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Self-Supervised Contrastive Learning for Audio-Visual Action RecognitionabstractThe underlying correlation between audio and visual modalities can be utilized to learn supervised information for unlabeled videos. In this paper, we propose an end-to-end self-supervised framework named Audio-Visual Contrastive Learning (AVCL), to learn discriminative audio-visual representations for action recognition. Specifically, we design an attention based multi-modal fusion module (AMFM) to fuse audio and visual modalities. To align heterogeneous audio-visual modalities, we construct a novel co-correlation guided representation alignment module (CGRA). To learn supervised information from unlabeled videos, we propose a novel self-supervised contrastive learning module (SelfCL). Furthermore, we build a new audio-visual action recognition dataset named Kinetics-Sounds100. Experimental results on the Kinetics-Sounds32 and Kinetics-Sounds100 datasets demonstrate the superiority of our AVCL over the state-of-the-art methods on large-scale action recognition benchmark. Yang Liu 0084, Haoyuan Lan |
ICIP | 1 |
| 2023 | Visual Causal Scene Refinement for Video Question AnsweringabstractExisting methods for video question answering (VideoQA) often suffer from spurious correlations between different modalities, leading to a failure in identifying the dominant visual evidence and the intended question. Moreover, these methods function as black boxes, making it difficult to interpret the visual scene during the QA process. In this paper, to discover critical video segments and frames that serve as the visual causal scene for generating reliable answers, we present a causal analysis of VideoQA and propose a framework for cross-modal causal relational reasoning, named Visual Causal Scene Refinement (VCSR). Particularly, a set of causal front-door intervention operations is introduced to explicitly find the visual causal scenes at both segment and frame levels. Our VCSR involves two essential modules: i) the Question-Guided Refiner (QGR) module, which refines consecutive video frames guided by the question semantics to obtain more representative segment features for causal front-door intervention; ii) the Causal Scene Separator (CSS) module, which discovers a collection of visual causal and non-causal scenes based on the visual-linguistic causal relevance and estimates the causal effect of the scene-separating intervention in a contrastive learning manner. Extensive experiments on the NExT-QA, Causal-VidQA, and MSRVTT-QA datasets demonstrate the superiority of our VCSR in discovering visual causal scene and achieving robust video question answering. Yushen Wei, Yang Liu 0084, Hong Yan 0004, Guanbin Li, Liang Lin 0004 |
ACM Multimedia | 2 |
| 2023 | VCD: Visual Causality Discovery for Cross-Modal Question Reasoning
Yang Liu 0084, Jingzhou Luo |
PRCV (7) | 1 |
| 2023 | Urban regional function guided traffic flow prediction
Lingbo Liu, Yang Liu 0084, Guanbin Li, Liang Lin 0004 |
Inf. Sci. | 3 |
| 2023 | Cross-Modal Causal Relational Reasoning for Event-Level Visual Question AnsweringabstractExisting visual question answering methods often suffer from cross-modal spurious correlations and oversimplified event-level reasoning processes that fail to capture event temporality, causality, and dynamics spanning over the video. In this work, to address the task of event-level visual question answering, we propose a framework for cross-modal causal relational reasoning. In particular, a set of causal intervention operations is introduced to discover the underlying causal structures across visual and linguistic modalities. Our framework, named Cross-Modal Causal RelatIonal Reasoning (CMCIR), involves three modules: i) Causality-aware Visual-Linguistic Reasoning (CVLR) module for collaboratively disentangling the visual and linguistic spurious correlations via front-door and back-door causal interventions; ii) Spatial-Temporal Transformer (STT) module for capturing the fine-grained interactions between visual and linguistic semantics; iii) Visual-Linguistic Feature Fusion (VLFF) module for learning the global semantic-aware visual-linguistic representations adaptively. Extensive experiments on four event-level datasets demonstrate the superiority of our CMCIR in discovering visual-linguistic causal structures and achieving robust event-level visual question answering. Yang Liu 0084, Guanbin Li, Liang Lin 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Hybrid-Order Representation Learning for Electricity Theft DetectionabstractElectricity theft is the primary cause of electrical losses in power systems, which severely harms the economic benefits of electricity providers and threatens the safety of the power supply. However, due to the inherent complex correlation and periodicity of electricity consumption and the low efficiency of large-scale data processing, detecting anomalies in electricity consumption data accurately and efficiently remains challenging. Existing methods usually focus on first-order information and ignore the second-order representation learning that can efficiently model global temporal dependency and facilitate discriminative representation learning of electricity consumption data. In this article, we propose a novel electricity theft detection framework named hybrid-order representation learning network (HORLN). Specifically, the sequential electricity consumption data is transformed into the matrix format containing weekly consumption records. Then, an inter-and-intra week convolution block is designed to capture multiscale features in a local-to-global manner. Meanwhile, a self-dependency modeling module is proposed to learn the second-order representations from self-correlation matrices, which are finally combined with the first-order representations to predict the anomaly scores of electricity consumers. Extensive experiments on a real-world benchmark demonstrate the advantages of our HORLN over state-of-the-art methods. Yuying Zhu 0007, Lingbo Liu, Yang Liu 0084, Guanbin Li, Mingzhi Mao, Liang Lin 0004 |
IEEE Trans. Ind. Informatics | 4 |
| 2022 | TCGL: Temporal Contrastive Graph for Self-Supervised Video Representation LearningabstractVideo self-supervised learning is a challenging task, which requires significant expressive power from the model to leverage rich spatial-temporal knowledge and generate effective supervisory signals from large amounts of unlabeled videos. However, existing methods fail to increase the temporal diversity of unlabeled videos and ignore elaborately modeling multi-scale temporal dependencies in an explicit way. To overcome these limitations, we take advantage of the multi-scale temporal dependencies within videos and propose a novel video self-supervised learning framework named Temporal Contrastive Graph Learning (TCGL), which jointly models the inter-snippet and intra-snippet temporal dependencies for temporal representation learning with a hybrid graph contrastive learning strategy. Specifically, a Spatial-Temporal Knowledge Discovering (STKD) module is first introduced to extract motion-enhanced spatial-temporal representations from videos based on the frequency domain analysis of discrete cosine transform. To explicitly model multi-scale temporal dependencies of unlabeled videos, our TCGL integrates the prior knowledge about the frame and snippet orders into graph structures, i.e., the intra-/inter-snippet Temporal Contrastive Graphs (TCG). Then, specific contrastive learning modules are designed to maximize the agreement between nodes in different graph views. To generate supervisory signals for unlabeled videos, we introduce an Adaptive Snippet Order Prediction (ASOP) module which leverages the relational knowledge among video snippets to learn the global context representation and recalibrate the channel-wise features adaptively. Experimental results demonstrate the superiority of our TCGL over the state-of-the-art methods on large-scale action recognition and video retrieval benchmarks. The code is publicly available at https://github.com/YangLiu9208/TCGL. Yang Liu 0084, Keze Wang, Lingbo Liu, Haoyuan Lan, Liang Lin 0004 |
IEEE Trans. Image Process. | 1 |
| 2021 | Semantics-Aware Adaptive Knowledge Distillation for Sensor-to-Vision Action RecognitionabstractExisting vision-based action recognition is susceptible to occlusion and appearance variations, while wearable sensors can alleviate these challenges by capturing human motion with one-dimensional time-series signals (e.g. acceleration, gyroscope, and orientation). For the same action, the knowledge learned from vision sensors (videos or images) and wearable sensors, may be related and complementary. However, there exists a significantly large modality difference between action data captured by wearable-sensor and vision-sensor in data dimension, data distribution, and inherent information content. In this paper, we propose a novel framework, named Semantics-aware Adaptive Knowledge Distillation Networks (SAKDN), to enhance action recognition in vision-sensor modality (videos) by adaptively transferring and distilling the knowledge from multiple wearable sensors. The SAKDN uses multiple wearable-sensors as teacher modalities and uses RGB videos as student modalities. To preserve the local temporal relationship and facilitate employing visual deep learning models, we transform one-dimensional time-series signals of wearable sensors to two-dimensional images by designing a gramian angular field based virtual image generation model. Then, we introduce a novel Similarity-Preserving Adaptive Multi-modal Fusion Module (SPAMFM) to adaptively fuse intermediate representation knowledge from different teacher networks. Finally, to fully exploit and transfer the knowledge of multiple well-trained teacher networks to the student network, we propose a novel Graph-guided Semantically Discriminative Mapping (GSDM) module, which utilizes graph-guided ablation analysis to produce a good visual explanation to highlight the important regions across modalities and concurrently preserve the interrelations of original data. Experimental results on Berkeley-MHAD, UTD-MHAD, and MMAct datasets well demonstrate the effectiveness of our proposed SAKDN for adaptive knowledge transfer from wearable-sensors modalities to vision-sensors modalities. The code is publicly available at https://github.com/YangLiu9208/SAKDN. Yang Liu 0084, Keze Wang, Guanbin Li, Liang Lin 0004 |
IEEE Trans. Image Process. | 1 |
| 2020 | Deep Image-to-Video Adaptation and Fusion Networks for Action RecognitionabstractExisting deep learning methods for action recognition in videos require a large number of labeled videos for training, which is labor-intensive and time-consuming. For the same action, the knowledge learned from different media types, e.g., videos and images, may be related and complementary. However, due to the domain shifts and heterogeneous feature representations between videos and images, the performance of classifiers trained on images may be dramatically degraded when directly deployed to videos. In this paper, we propose a novel method, named Deep Image-to-Video Adaptation and Fusion Networks (DIVAFN), to enhance action recognition in videos by transferring knowledge from images using video keyframes as a bridge. The DIVAFN is a unified deep learning model, which integrates domain-invariant representations learning and cross-modal feature fusion into a unified optimization framework. Specifically, we design an efficient cross-modal similarities metric to reduce the modality shift among images, keyframes and videos. Then, we adopt an autoencoder architecture, whose hidden layer is constrained to be the semantic representations of the action class names. In this way, when the autoencoder is adopted to project the learned features from different domains to the same space, more compact, informative and discriminative representations can be obtained. Finally, the concatenation of the learned semantic feature representations from these three autoencoders are used to train the classifier for action recognition in videos. Comprehensive experiments on four real-world datasets show that our method outperforms some state-of-the-art domain adaptation and action recognition methods. Yang Liu 0084, Zhaoyang Lu, Jing Li 0010, Tao Yang 0006 |
IEEE Trans. Image Process. | 1 |
| 2019 | Hierarchically Learned View-Invariant Representations for Cross-View Action RecognitionabstractRecognizing human actions from varied views is challenging due to huge appearance variations in different views. The key to this problem is to learn discriminant view-invariant representations generalizing well across views. In this paper, we address this problem by learning view-invariant representations hierarchically using a novel method, referred to as joint sparse representation and distribution adaptation. To obtain robust and informative feature representations, we first incorporate a sample-affinity matrix into the marginalized Stacked Denoising Autoencoder to obtain shared features that are then combined with the private features. In order to make the feature representations of videos across views transferable, we then learn a transferable dictionary pair simultaneously from pairs of videos taken at different views to encourage each action video across views to have the same sparse representation. However, the distribution difference across views still exists because a unified subspace, where the sparse representations of one action across views are the same, may not exist when the view difference is large. Therefore, we propose a novel unsupervised distribution adaptation method that learns a set of projections that project the source and target views data into respective low-dimensional subspaces, where the marginal and conditional distribution differences are reduced simultaneously. Therefore, the finally learned feature representation is view-invariant and robust for substantial distribution difference across views even though the view difference is large. Experimental results on four multi-view datasets show that our approach outperforms the state-of-the-art approaches. Yang Liu 0084, Zhaoyang Lu, Jing Li 0010, Tao Yang 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Global Temporal Representation Based CNNs for Infrared Action RecognitionabstractInfrared human action recognition has many advantages, i.e., it is insensitive to illumination change, appearance variability, and shadows. Existing methods for infrared action recognition are either based on spatial or local temporal information, however, the global temporal information, which can better describe the movements of body parts across the whole video, is not considered. In this letter, we propose a novel global temporal representation named optical-flow stacked difference image (OFSDI) and extract robust and discriminative feature from the infrared action data by considering the local, global, and spatial temporal information together. Due to the small size of the infrared action dataset, we first apply convolutional neural networks on local, spatial, and global temporal stream respectively to obtain efficient convolutional feature maps from the raw data rather than train a classifier directly. Then these convolutional feature maps are aggregated into effective descriptors named three-stream trajectory-pooled deep-convolutional descriptors by trajectory-constrained pooling. Furthermore, we improve the robustness of these features by using the locality-constrained linear coding (LLC) method. With these features, a linear support vector machine (SVM) is adopted to classify the action data in our scheme. We conduct the experiments on infrared action recognition datasets InfAR and NTU RGB+D. The experimental results show that the proposed approach outperforms the representative state-of-the-art handcrafted features and deep learning features based methods for the infrared action recognition. Yang Liu 0084, Zhaoyang Lu, Jing Li 0010, Tao Yang 0006 |
IEEE Signal Process. Lett. | 1 |
| 2017 | Adaptive maximum margin analysis for image recognition
Qianqian Wang 0001, Quanxue Gao, Yunsong Li 0001, Yunfang Huang, Yang Liu 0084 |
Pattern Recognit. | 6 |
| 2017 | A Non-Greedy Algorithm for L1-Norm LDAabstractRecently, L1-norm-based discriminant subspace learning has attracted much more attention in dimensionality reduction and machine learning. However, most existing approaches solve the column vectors of the optimal projection matrix one by one with greedy strategy. Thus, the obtained optimal projection matrix does not necessarily best optimize the corresponding trace ratio objective function, which is the essential criterion function for general supervised dimensionality reduction. In this paper, we propose a non-greedy iterative algorithm to solve the trace ratio form of L1-norm-based linear discriminant analysis. We analyze the convergence of our proposed algorithm in detail. Extensive experiments on five popular image databases illustrate that our proposed algorithm can maximize the objective function value and is superior to most existing L1-LDA algorithms. Yang Liu 0084, Quanxue Gao, Shuo Miao, Xinbo Gao 0001, Feiping Nie 0001, Yunsong Li 0001 |
IEEE Trans. Image Process. | 1 |