VLDB 2026 Research / reviewers in the wild / expert
Tianshan Liu
dblp:235/6426
· DBLP profile ↗
41ranked-venue papers
15as first author
36since 2021 · last 2026
0000-0003-3831-8893ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 9 first-author · 20 since 2021Artificial intelligence and machine learning · 15 · 7 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Goal-Guided Prompting With Adaptive Modality Selection for Efficient Assembly Activity Anticipation in Egocentric VideosabstractWith the functions of egocentric observation and multimodal perception equipped in augmented reality (AR) devices, the next generation of smart assistants has the potential to reduce human labor and enhance execution efficiency in assembly tasks. Among diverse assembly activity understanding tasks, anticipating the near future activities is crucial yet challenging, which can assist humans or agents to actively plan and engage in interactions with the environment. However, the existing egocentric activity anticipation methods still struggle to achieve a decent trade-off between accuracy and computational efficiency, hindering them to be deployed in practical applications. To address this dilemma, in this paper, we propose a goal-guided prompting framework with adaptive modality selection (GP-AMS), for assembly activity anticipation in egocentric videos. For bridging the semantic gap between the historical observations and unobserved future activities, we inject the inferred high-level goal clues into the constructed prompts, which are further utilized to guide a pre-trained vision-language (V-L) model to compensate relevant semantics of unseen future. Moreover, a mask-and-predict strategy is adopted with two imposed constraints, i.e., casual masking and probabilistic token-dropping, to mine the intrinsic associations between the assembly activities within a specific procedure. For maintaining the benefits of exploiting multimodal information while avoiding extensively increasing the computational burdens, an adaptive modality selection strategy is designed to train a policy network, which learns to dynamically decide which modalities should be sampled for processing by the anticipation model on a per observation time-step basis. By allocating major computation to the selected indicative modalities on-the-fly, the efficiency of the overall model can be improved, thus paving the way for feasibility on real-world devices. Extensive experimental results on two public data sets validate that the proposed method yields not only consistent improvements in anticipation accuracy, but also significant savings in computation budgets. Tianshan Liu, Bing-Kun Bao |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | An Episode Memory-Guided Dual-Stage Framework for Long-Form Video Temporal GroundingabstractVideo temporal grounding (VTG) aims to localize video moments that are semantically related to a given natural language query. In spite of recent progress in short-form videos, research on VTG in long-form videos (e.g., hours long) remains highly demanded yet underexplored. Existing methods predominantly adopt sliding window-based or multi-scale anchor-based strategies to generate temporal proposals, which require time-consuming post-processing or are independent of video content, thereby limiting their performance and efficiency. To address this dilemma, in this paper, we propose an episode memory-prompted (EMP) two-stage framework for temporal grounding in long-form videos. Specifically, the first stage generates a set of dynamic episode memories, which explicitly summarize various activities occurring throughout the lengthy video. An unsupervised memory learning paradigm is formulated by imposing discriminability and diversity constraints, eliminating the reliance on additional activity-instance annotations. Then, in the second stage, based on the supplement of frame-level detailed content and the guidance of a language query, the augmented memory prompts function as anchors for efficiently regressing the refined boundaries of the target video moment. Extensive experimental results on two public long-form video data sets, i.e., MAD and Ego4d, validate that the proposed EMP framework saves more than 8.5% trainable parameters and 13.9% FLOPs, while still achieving comparable performance with existing methods. Tianshan Liu, Bing-Kun Bao, Kin-Man Lam 0001 |
IEEE Trans. Image Process. | 1 |
| 2026 | A Reciprocal Interaction Framework for Collaborative Temporal Grounding and Question Answering in Egocentric VideosabstractCollaborative Temporal Grounding and Question Answering (CTGQA) in egocentric videos enables users to inquire about past visual experiences and obtain corresponding temporal segments and answers. Existing CTGQA methods typically treat Video Temporal Grounding (VTG) and Video Question Answering (VQA) as separate tasks, overlooking their inherent semantic and temporal complementarity. As a result, VQA models often generate ambiguous answers due to the lack of precise temporal cues, while VTG models fail to fully exploit the high-level semantic information embedded in the answers. To address these limitations, we propose a Reciprocal Interaction Framework (RIF). RIF employs a two-branch interaction structure to enhance the performance of both VTG and VQA. RIF consists of two modules: Localization-Guided Answering (LGA) and Answer-Enhanced Temporal Grounding (AETG). The LGA module assists the VQA model in generating high-quality answers by highlighting relevant segments while minimizing the influence of irrelevant content. To mitigate model overconfidence, we propose a progressive feature fusion strategy that dynamically adjusts the weights of relevant segments, thus preventing localization errors. The AETG module leverages additional information embedded in the generated answer to improve VTG performance. Moreover, we employ a perplexity-based filtering strategy to ensure the reliability of the answer. Extensive experiments show that our framework performs well on the QAEGO4D and Ego4D-NLQ benchmarks. Tianshan Liu, Bing-Kun Bao |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | A Memory-Assisted Knowledge Transferring Framework with Curriculum Anticipation for Weakly Supervised Online Activity Detection
Tianshan Liu, Kin-Man Lam 0001, Bing-Kun Bao |
Int. J. Comput. Vis. | 1 |
| 2025 | MAD-DGTD: Multivariate time series Anomaly Detection based on Dynamic Graph structure learning with Time Delay
Jun Kong 0001, Meicheng Zhang, Min Jiang 0008, Tianshan Liu |
Neurocomputing | 5 |
| 2024 | Hierarchical Vertex-Wise Intensification Graph Convolution for Skeleton-Based Activity RecognitionabstractGraph convolutional networks (GCNs), which can effectively captures the spatial and temporal relationships between skeleton joints through graph topology, have shown promising performances in skeleton-based activity recognition in recent years. These methods typically learn the semantic features of the vertices of a skeleton and the associated adjacency matrix. However, how to efficiently establish relationships between vertices still remains a substantial problem. To solve this problem, we propose a novel Hierarchical Vertex-wise Intensification Graph Convolution Network (HVI-GCN) for skeleton-based action recognition. The proposed module dilates input features into higher dimensions to broaden the temporal horizon, and builds a vertex-wise topology based on self-adaptively learned attention. With the adjacency matrix, features from other positions can be collected to aid the prediction of the current position. The proposed module provides a better receptive field and semantic understanding of both the spatial and temporal domains than related methods. Experiments were mainly conducted on the at NTU-RGB-D, NTU-GRB-D 120, and NW-UCLA datasets with joint and bone integrated with motion sequences. Experimental results show that HVI-GCN can improve accuracy by up to 1.1% on the RGB-D 120 dataset. Meanwhile, the accuracy on RGB-D 60 dataset and NW-UCLA dataset can be boosted by 1.4% and 1.2%, respectively. Jun Xiao 0010, Tianshan Liu, Kin-Man Lam 0001 |
ICIP | 5 |
| 2024 | Label Text-aided Hierarchical Semantics Mining for Panoramic Activity RecognitionabstractPanoramic activity recognition is a comprehensive yet challenging task in crowd scene understanding, which aims to concurrently identify multi-grained human behaviors, including individual actions, social group activities, and global activities. Previous studies tend to capture cross-granularity activity-semantics relations from solely the video input, thus ignoring the intrinsic semantic hierarchy in label-text space. To this end, we propose a label text-aided hierarchical semantics mining (THSM) framework, which explores multi-level cross-modal associations by learning hierarchical semantic alignment between visual content and label texts. Specifically, a hierarchical encoder is first constructed to encode the visual and text inputs into semantics-aligned representations at different granularities. To fully exploit the cross-modal semantic correspondence learned by the encoder, a hierarchical decoder is further developed, which progressively integrates the lower-level representations with the higher-level contextual knowledge for coarse-to-fine action/activity recognition. Extensive experimental results on the public JRDB-PAR benchmark validate the superiority of the proposed THSM framework over state-of-the-art methods. Tianshan Liu, Kin-Man Lam 0001, Bing-Kun Bao |
ACM Multimedia | 1 |
| 2024 | Deep progressive feature aggregation network for multi-frame high dynamic range imaging
Jun Xiao 0010, Tianshan Liu, Kin-Man Lam 0001 |
Neurocomputing | 3 |
| 2024 | Patch-based tendency camera multi-constraint learning for unsupervised person re-identification
Xuefeng Tao, Jun Kong 0001, Min Jiang 0008, Tianshan Liu |
J. Vis. Commun. Image Represent. | 5 |
| 2024 | Category-Aware Saliency Enhance Learning Based on CLIP for Weakly Supervised Salient Object DetectionabstractAbstract Weakly supervised salient object detection (SOD) using image-level category labels has been proposed to reduce the annotation cost of pixel-level labels. However, existing methods mostly train a classification network to generate a class activation map, which suffers from coarse localization and difficult pseudo-label updating. To address these issues, we propose a novel Category-aware Saliency Enhance Learning (CSEL) method based on contrastive vision-language pre-training (CLIP), which can perform image-text classification and pseudo-label updating simultaneously. Our proposed method transforms image-text classification into pixel-text matching and generates a category-aware saliency map, which is evaluated by the classification accuracy. Moreover, CSEL assesses the quality of the category-aware saliency map and the pseudo saliency map, and uses the quality confidence scores as weights to update the pseudo labels. The two maps mutually enhance each other to guide the pseudo saliency map in the correct direction. Our SOD network can be trained jointly under the supervision of the updated pseudo saliency maps. We test our model on various well-known RGB-D and RGB SOD datasets. Our model achieves an S-measure of 87.6 $$\%$$ % on the RGB-D NLPR dataset and 84.3 $$\%$$ % on the RGB ECSSD dataset. Additionally, we obtain satisfactory performance on the weakly supervised E-measure, F-measure, and mean absolute error metrics for other datasets. These results demonstrate the effectiveness of our model. Yunde Zhang, Tianshan Liu, Jun Kong 0001 |
Neural Process. Lett. | 3 |
| 2024 | Structured Adversarial Self-Supervised Learning for Robust Object Detection in Remote Sensing ImagesabstractObject detection plays a crucial role in scene understanding and has extensive practical applications. In the field of remote sensing object detection, both detection accuracy and robustness are of significant concern. Existing methods heavily rely on sophisticated adversarial training strategies that tend to improve robustness at the expense of accuracy. However, detection robustness is not always indicative of improved accuracy. Therefore, in this paper, we research how to enhance robustness, while still preserving high accuracy, or even improve both simultaneously, with simple vanilla adversarial training or even in the absence thereof. In pursuit of a solution, we first conduct an exploratory investigation by shifting our attention from adversarial training, referred to as adversarial fine-tuning, to adversarial pretraining. Specifically, we propose a novel pretraining paradigm, namely structured adversarial self-supervised (SASS) pretraining, to strengthen both clean accuracy and adversarial robustness for object detection in remote sensing images. At a high level, SASS pretraining aims to unify adversarial learning and self-supervised learning into pretraining and encode structured knowledge into pretrained representations for powerful transferability to downstream detection. Moreover, to fully explore the inherent robustness of vision Transformers and facilitate their pretraining efficiency, by leveraging the recent masked image modeling (MIM) as the pretext task, we further instantiate SASS pretraining into a concise end-to-end framework, named structured adversarial MIM (SA-MIM). SA-MIM consists of two pivotal components, structured adversarial attack and structured MIM (S-MIM). The former establishes structured adversaries for the context of adversarial pretraining, while the latter introduces a structured local-sampling global-masking strategy to adapt to hierarchical encoder architectures. Comprehensive experiments on three different datasets have demonstrated the significant superiority of the proposed pretraining paradigm over previous counterparts for remote sensing object detection. More importantly, regardless of with or without adversarial fine-tuning, it enables simultaneous improvements on detection accuracy and robustness as expected, promisingly alleviating the dependence on complicated adversarial fine-tuning. Kin-Man Lam 0001, Tianshan Liu, Yui-Lam Chan, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Injecting Text Clues for Improving Anomalous Event Detection From Weakly Labeled VideosabstractVideo anomaly detection (VAD) aims at localizing the snippets containing anomalous events in long unconstrained videos. The weakly supervised (WS) setting, where solely video-level labels are available during training, has attracted considerable attention, owing to its satisfactory trade-off between the detection performance and annotation cost. However, due to lack of snippet-level dense labels, the existing WS-VAD methods still get easily stuck on the detection errors, caused by false alarms and incomplete localization. To address this dilemma, in this paper, we propose to inject text clues of anomaly-event categories for improving WS-VAD, via a dedicated dual-branch framework. For suppressing the response of confusing normal contexts, we first present a text-guided anomaly discovering (TAG) branch based on a hierarchical matching scheme, which utilizes the label-text queries to search the discriminative anomalous snippets in a global-to-local fashion. To facilitate the completeness of anomaly-instance localization, an anomaly-conditioned text completion (ATC) branch is further designed to perform an auxiliary generative task, which intrinsically forces the model to gather sufficient event semantics from all the relevant anomalous snippets for completely reconstructing the masked description sentence. Furthermore, to encourage the cross-branch knowledge sharing, a mutual learning strategy is introduced by imposing a consistency constraint on the anomaly scores of these two branches. Extensive experimental results on two public benchmarks validate that the proposed method achieves superior performance over the competing methods. Tianshan Liu, Kin-Man Lam 0001, Bing-Kun Bao |
IEEE Trans. Image Process. | 1 |
| 2024 | Distilling Privileged Knowledge for Anomalous Event Detection From Weakly Labeled VideosabstractWeakly supervised video anomaly detection (WS-VAD) aims to identify the snippets involving anomalous events in long untrimmed videos, with solely text video-level binary labels. A typical paradigm among the existing text WS-VAD methods is to employ multiple modalities as inputs, e.g., RGB, optical flow, and audio, as they can provide sufficient discriminative clues that are robust to the diverse, complicated real-world scenes. However, such a pipeline has high reliance on the availability of multiple modalities and is computationally expensive and storage demanding in processing long sequences, which limits its use in some applications. To address this dilemma, we propose a privileged knowledge distillation (KD) framework dedicated to the WS-VAD task, which can maintain the benefits of exploiting additional modalities, while avoiding the need for using multimodal data in the inference phase. We argue that the performance of the privileged KD framework mainly depends on two factors: 1) the effectiveness of the multimodal teacher network and 2) the completeness of the useful information transfer. To obtain a reliable teacher network, we propose a text cross-modal interactive learning strategy and an anomaly normal discrimination loss, which target learning task-specific cross-modal features and encourage the separability of anomalous and normal representations, respectively. Furthermore, we design both representation- and text logits-level distillation loss functions, which force the unimodal student network to distill abundant privileged knowledge from the text well-trained multimodal teacher network, in a snippet-to-video fashion. Extensive experimental results on three public benchmarks demonstrate that the proposed privileged KD framework can train a lightweight yet effective detector, for localizing anomaly events under the supervision of video-level annotations. Tianshan Liu, Kin-Man Lam 0001, Jun Kong 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Holistic-Guided Disentangled Learning With Cross-Video Semantics Mining for Concurrent First-Person and Third-Person Activity RecognitionabstractThe popularity of wearable devices has increased the demands for the research on first-person activity recognition. However, most of the current first-person activity datasets are built based on the assumption that only the human-object interaction (HOI) activities, performed by the camera-wearer, are captured in the field of view. Since humans live in complicated scenarios, in addition to the first-person activities, it is likely that third-person activities performed by other people also appear. Analyzing and recognizing these two types of activities simultaneously occurring in a scene is important for the camera-wearer to understand the surrounding environments. To facilitate the research on concurrent first- and third-person activity recognition (CFT-AR), we first created a new activity dataset, namely PolyU concurrent first- and third-person (CFT) Daily, which exhibits distinct properties and challenges, compared with previous activity datasets. Since temporal asynchronism and appearance gap usually exist between the first- and third-person activities, it is crucial to learn robust representations from all the activity-related spatio-temporal positions. Thus, we explore both holistic scene-level and local instance-level (person-level) features to provide comprehensive and discriminative patterns for recognizing both first- and third-person activities. On the one hand, the holistic scene-level features are extracted by a 3-D convolutional neural network, which is trained to mine shared and sample-unique semantics between video pairs, via two well-designed attention-based modules and a self-knowledge distillation (SKD) strategy. On the other hand, we further leverage the extracted holistic features to guide the learning of instance-level features in a disentangled fashion, which aims to discover both spatially conspicuous patterns and temporally varied, yet critical, cues. Experimental results on the PolyU CFT Daily dataset validate that our method achieves the state-of-the-art performance. Tianshan Liu, Rui Zhao 0012, Wenqi Jia 0001, Kin-Man Lam 0001, Jun Kong 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2023 | Pyramid Masked Image Modeling for Transformer-Based Aerial Object DetectionabstractTwo obstacles, the scarcity of annotated samples and the difficulty in preserving multi-scale hierarchical representations, hinder the advancement of vision Transformer-based aerial object detection. The emergence of self-supervised learning has inspired some solutions to the first issue. However, most solutions focus on single-scale features, conflicting with solving the second issue. To bridge this gap, this paper proposes a novel pyramid masked image modeling (MIM) framework, termed PyraMIM, for self-supervised pretraining in aerial scenarios. Without manual annotation, PyraMIM enables establishing pyramid representations during pretraining, which can be seamlessly adapted to downstream aerial object detection for performance improvement. Experimental results demonstrate the effectiveness and superiority of our method. Tianshan Liu, Yakun Ju, Kin-Man Lam 0001 |
ICIP | 2 |
| 2023 | Multi-stream ternary enhanced graph convolutional network for skeleton-based action recognition
Jun Kong 0001, Shengquan Wang, Min Jiang 0008, Tianshan Liu |
Neural Comput. Appl. | 4 |
| 2023 | CALTracker: Cross-Task Association Learning for Multiple Object TrackingabstractMultiple object tracking has recently achieved excellent performance based on the joint optimization of detection and re-identification tasks. However, joint optimization normally homogenizes the features in the detection and re-identification tasks, which weakens the representation of each task's inherent geometric and semantic information. Additionally, the stability of the tracking trajectory will be impacted by the feature misalignment of the associated information between different tasks. In this letter, we propose a Cross-task Association Learning Tracker (CALTracker) to trade off the inherent and associated information. We first design a Triplet Shrinkage Decoupling (TSD) module to ensure the independence of sub-task features in the optimization process, thereby minimizing the optimization conflicts caused by homogeneous features. Secondly, to improve the consistent representation of the associated information between subtasks, a Double Attention Cross-task Learning (DACL) strategy is designed to achieve cross-task feature alignment and mutual gain. Finally, extensive experimental results on MOT17 and MOT20 demonstrate the effectiveness of the proposed method over the state-of-the-art performance. Jun Kong 0001, Min Jiang 0008, Tianshan Liu |
IEEE Signal Process. Lett. | 4 |
| 2023 | Geometry-Aware Facial Expression Recognition via Attentive Graph Convolutional NetworksabstractLearning discriminative representations with good robustness from facial observations serves as a fundamental step towards intelligent facial expression recognition (FER). In this article, we propose a novel geometry-aware FER framework to boost the FER performance based on both the geometric and appearance knowledge. Specifically, we propose an encoding strategy for facial landmarks, and adopt a graph convolutional network (GCN) to fully explore the structural information of the facial components behind different expressions. A convolutional neural network (CNN) is further applied to the whole facial observation to learn the global characteristics of different expressions. The features from these two networks are fused into a comprehensive high-semantic representation, which promotes the FER reasoning from both visual and structural perspectives. Moreover, to facilitate the networks to concentrate on the most informative facial regions and components, we introduce multi-level attention mechanisms into the proposed framework, which enhance the reliability of the learned representations for effective FER. Experiments on two challenging FER benchmarks demonstrate that the attentive graph-based learning on the facial geometry boosts the FER accuracy. Furthermore, the insensitivity of the geometric information to the appearance variations also improves the generalization of the proposed framework. Rui Zhao 0012, Tianshan Liu, Zixun Huang, Daniel Pak-Kong Lun, Kin-Man Lam 0001 |
IEEE Trans. Affect. Comput. | 2 |
| 2023 | Spatial-Temporal Graphs Plus Transformers for Geometry-Guided Facial Expression RecognitionabstractFacial expression recognition (FER) is of great interest to the current studies of human-computer interaction. In this paper, we propose a novel geometry-guided facial expression recognition framework, based on graph convolutional networks and transformers, to perform effective emotion recognition from videos. Specifically, we detect and utilize facial landmarks to construct a spatial-temporal graph, based on both the landmark coordinates and local appearance, for representing a facial expression sequence. The graph convolutional blocks and transformer modules are employed to produce high-semantic emotion-related representations from the structured facial graphs, which facilitate the framework to establish both the local and non-local dependency between the vertices. Moreover, spatial and temporal attention mechanisms are introduced into graph-based learning to promote FER reasoning, via the emphasis on the most informative facial components and frames. Extensive experiments demonstrate that the proposed framework achieves promising performance for geometry-based FER and shows great generalization and robustness in real-world applications. Rui Zhao 0012, Tianshan Liu, Zixun Huang, Daniel Pak-Kong Lun, Kin-Man Lam 0001 |
IEEE Trans. Affect. Comput. | 2 |
| 2023 | Decouple and Resolve: Transformer-Based Models for Online Anomaly Detection From Weakly Labeled VideosabstractAs one of the vital topics in intelligent surveillance, weakly supervised online video anomaly detection (WS-OVAD) aims to identify the ongoing anomalous events moment-to-moment in streaming videos, trained with only video-level annotations. Previous studies tended to utilize a unified single-stage framework, which struggled to simultaneously address the issues of online constraints and weakly supervised settings. To solve this dilemma, in this paper, we propose a two-stage-based framework, namely “decouple and resolve” (DAR), which consists of two modules, i.e., temporal proposal producer (TPP) and online anomaly localizer (OAL). With the supervision of video-level binary labels, the TPP module targets fully exploiting hierarchical temporal relations among snippets for generating precise snippet-level pseudo-labels. Then, given fine-grained supervisory signals produced by TPP, the Transformer-based OAL module is trained to aggregate both the useful cues retrieved from historical observations and anticipated future semantics, for making predictions at the current time step. Both the TPP and OAL modules are jointly trained to share the beneficial knowledge in a multi-task learning paradigm. Extensive experimental results on three public data sets validate the superior performance of the proposed DAR framework over the competing methods. Tianshan Liu, Kin-Man Lam 0001, Jun Kong 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2023 | Weakly Supervised Distribution Discrepancy Minimization Learning With State Information for Person Re-IdentificationabstractWeakly supervised person re-identification (Re-ID) is appealing to handle real-world tasks by using state information that is available without manual annotation. At present, most methods perform unsupervised cross domain (UCD) learning by transferring the knowledge from the labeled source domain to the unlabeled target domain, which results in poor performance due to the severe shift. To address this problem, in this paper, we utilize the tracklet and camera information as weak supervision to propose a distribution discrepancy minimization learning (DDML) model for UCD person Re-ID. In addition to aligning data distributions from the perspective of domain adaptation learning, two losses are developed from the view of neighborhood invariance exploration to optimize matching results. Specifically, to bridge the gap between domains, we propose a camera-distribution-based (CDB) loss to align pair-wise distance distributions. Furthermore, to alleviate the biased search within the target domain, we propose a ranking-confidence-based (RCB) loss to perform the mined neighborhood for intra-camera and inter-camera separately to explore a high degree of confidence neighbor relations. Extensive experiments on three challenging datasets demonstrate that applying our method to unlabeled target domain outperforms current weakly supervised methods for person Re-ID. Jun Kong 0001, Xuefeng Tao, Min Jiang 0008, Tianshan Liu |
IEEE Trans. Multim. | 4 |
| 2022 | A Hybrid Egocentric Activity Anticipation Framework via Memory-Augmented Recurrent and One-shot Representation ForecastingabstractEgocentric activity anticipation involves identifying the interacted objects and target action patterns in the near future. A standard activity anticipation paradigm is re-currently forecasting future representations to compensate the missing activity semantics of the unobserved sequence. However, the limitations of current recursive prediction models arise from two aspects: (i) The vanilla recurrent units are prone to accumulated errors in relatively long periods of anticipation. (ii) The anticipated representations may be insufficient to reflect the desired semantics of the target activity, due to lack of contextual clues. To address these issues, we propose “HRO ”, a hybrid framework that integrates both the memory-augmented recurrent and one-shot representation forecasting strategies. Specifically, to solve the limitation (i), we introduce a memory-augmented contrastive learning paradigm to regulate the process of the recurrent representation forecasting. Since the external memory bank maintains long-term prototypical activity semantics, it can guarantee that the anticipated representations are reconstructed from the discriminative activity prototypes. To further guide the learning of the memory bank, two auxiliary loss functions are designed, based on the diversity and sparsity mechanisms, respectively. Furthermore, to resolve the limitation (ii), a one-shot transferring paradigm is proposed to enrich the forecasted representations, by distilling the holistic activity semantics after the target anticipation moment, in the offline training. Extensive experimental results on two large-scale data sets validate the effectiveness of our proposed HRO method. Tianshan Liu, Kin-Man Lam 0001 |
CVPR | 1 |
| 2022 | Angle Tokenization Guided Multi-Scale Vision Transformer for Oriented Object Detection in Remote Sensing ImageryabstractIn this paper, an angle tokenization guided multi-scale Trans-former framework is proposed for oriented object detection in remote sensing images. Different from existing detectors that are based on convolutional neural networks (CNNs), our proposed method is based on a pyramid Transformer architecture with a compact and flexible angle tokenization module (ATM) to efficiently learn the orientation knowledge for rotated geospatial objects. The Transformer structure can progressively render long-range dependencies and multi-scale spatial details required for accurate localization, while the ATM provides robust guidance on feature refinement for angle prediction, jointly achieving end-to-end orientation de-tection. To the best of our knowledge, this is the first work to adapt Vision Transformers to remote sensing oriented object detection. Experimental results demonstrate the effectiveness and superiority of our method. Tianshan Liu, Kin-Man Lam 0001 |
IGARSS | 2 |
| 2022 | Visual-semantic graph neural network with pose-position attentive learning for group activity recognition
Tianshan Liu, Rui Zhao 0012, Kin-Man Lam 0001, Jun Kong 0001 |
Neurocomputing | 1 |
| 2022 | SPRTracker: Learning Spatial-Temporal Pixel Aggregations for Multiple Object TrackingabstractRecently, multiple object tracking based on a single frame has achieved excellent performance. However, in crowded scenes, occlusion and motion blur will increase the difficulty of foreground object detection. In this letter, we propose a Spatial-temporal Pixel Resampling Tracker (SPRTracker) that introduces a novel cross-frame input framework to improve the anti-occlusion and anti-interference ability during tracking. We first propose a sampling mechanism in which Inter-frame Pixel Alignment (IPA) is designed to maintain spatial and temporal consistency in the propagation of pixel-level information between frames, aiming to improve the anti-interference ability during tracking target motion. Secondly, To improve the compensation effect of the gain information in the historical frame to the occluded target in the current frame, Similarity Reparameter Fusion (SRF) strategy is designed to fuse the features of the two frames. We evaluated the proposed method on two common benchmarks (MOT17 and MOT20), and the experimental results effectively demonstrated the superiority of our method. Jun Kong 0001, Min Jiang 0008, Tianshan Liu |
IEEE Signal Process. Lett. | 4 |
| 2022 | MOTFR: Multiple Object Tracking Based on Feature RecodingabstractThe stable continuation of trajectories among different targets has always been the key to the tracking performance of multi-object tracking (MOT) tasks. If features of the target are aggregated and classified simply, the discriminant features of the target will be ignored. This will affect the robustness of the trajectory generated by the model. Meanwhile, many popular models are keen to execute detection and feature extraction tasks in parallel. But these two tasks will conflict with each other when optimized respectively. Therefore, we propose our tracker MOTFR to solve the above problems. In this paper, we propose a Locally Shared Information Decoupling Module (LSIDM) to reduce task optimization conflicts while ensuring the necessary information sharing. Meanwhile, a feature recoding module for deep extraction of identity discriminative features is proposed, which is called the Feature Purification Module (FPM). By combining LSIDM and FPM modules, the model utilizes the discriminative appearance features to guide the optimization of detection and further improves the performance of our model. To solve the problem of targets disappearing due to various abnormal occlusion, a Short-term Trajectory Online Complement Strategy (STOCS) is proposed to realize the trajectories continuation of these targets in the tracking stage. Through sufficient experiments, we demonstrate the superior performance of our MOTFR, which guarantees high-quality detection while achieving the stability of the target trajectory. Jun Kong 0001, Ensen Mo, Min Jiang 0008, Tianshan Liu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Enhanced Attention Tracking With Multi-Branch Network for Egocentric Activity RecognitionabstractThe emergence of wearable devices has opened up new potentials for egocentric activity recognition. Although some methods integrate attention mechanisms into deep neural networks to capture fine-grained human-object interactions in a weak-supervision manner, they either ignore exploiting the temporal consistency or generate attention based on considering appearance cues only. To address these limitations, in this paper, we propose an enhanced attention-tracking method, combined with multi-branch network (EAT-MBNet), for egocentric activity recognition. Specifically, we propose class-aware attention maps (CAAMs) by employing a self-attention-based module to refine the class activation maps (CAMs). Our proposed method can enhance the semantic dependency between the activity categories and the feature maps. To highlight the discriminative features from the regions of interest across frames, we propose a flow-guided attention-tracking (F-AT) module, by simultaneously leveraging historical attention and motion patterns. Furthermore, we propose a cross-modality modeling branch based on an interactive GRU module, which captures the time-synchronized long-term relationships between the appearance and motion branches. Experimental results on four egocentric activity benchmarks demonstrate that the proposed method achieves state-of-the-art performance. Tianshan Liu, Kin-Man Lam 0001, Rui Zhao 0012, Jun Kong 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Deep Cross-Modal Representation Learning and Distillation for Illumination-Invariant Pedestrian DetectionabstractIntegrating multispectral data has been demonstrated to be an effective solution for illumination-invariant pedestrian detection, in particular, RGB and thermal images can provide complementary information to handle light variations. However, most of the current multispectral detectors fuse the multimodal features by simple concatenation, without discovering their latent relationships. In this paper, we propose a cross-modal feature learning (CFL) module, based on a split-and-aggregation strategy, to explicitly explore both the shared and modality-specific representations between paired RGB and thermal images. We insert the proposed CFL module into multiple layers of a two-branch-based pedestrian detection network, to learn the cross-modal representations in diverse semantic levels. By introducing a segmentation-based auxiliary task, the multimodal network is trained end-to-end by jointly optimizing a multi-task loss. On the other hand, to alleviate the reliance of existing multispectral pedestrian detectors on thermal images, we propose a knowledge distillation framework to train a student detector, which only receives RGB images as input and distills the cross-modal representations guided by a well-trained multimodal teacher detector. In order to facilitate the cross-modal knowledge distillation, we design different distillation loss functions for the feature, detection and segmentation levels. Experimental results on the public KAIST multispectral pedestrian benchmark validate that the proposed cross-modal representation learning and distillation method achieves robust performance. Tianshan Liu, Kin-Man Lam 0001, Rui Zhao 0012, Guoping Qiu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Unsupervised Domain Adaptation by Multi-Loss Gap Minimization Learning for Person Re-IdentificationabstractUnsupervised domain adaptation (UDA) person re-identification (ReID) faces enormous challenges due to the severe shift between the source and target domains, as well as the dramatic variations within the target domain. In this paper, to address these issues, we propose a multi-loss gap minimization learning (MGML) approach for UDA person ReID. Firstly, we introduce the part model to learn discriminative patch features and design a Patch-based Part Ignoring (PPI) loss to select reliable instances for the efficient learning of the part model. Then, given the gap that typically occurs because of the inter-domain shift and intra-domain variations, a Gap-based Minimum Camera Discrepancy (G-MCD) loss is proposed. Specifically, in terms of the inter-domain, we propose to leverage the tracklet and camera information to label each distance vector, and accordingly align pair-wise distance distributions to bridge the inter-domain gap. As for the intra-domain, to alleviate the biased search, we propose to perform the mined neighborhood for intra-camera and inter-camera separately to optimize matching results by exploring neighborhood relations more deeply. Finally, experimental results on three challenging datasets demonstrate that applying our method to unlabeled target domain outperforms current UDA methods for person ReID. Xuefeng Tao, Jun Kong 0001, Min Jiang 0008, Tianshan Liu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Structure-Enhanced Attentive Learning For Spine Segmentation From Ultrasound Volume Projection ImagesabstractAutomatic spine segmentation, based on ultrasound volume projection imaging (VPI), is of great value in clinical applications to diagnose scoliosis in teenagers. In this paper, we propose a novel framework to improve the segmentation accuracy on spine images via structure-enhanced attentive learning. Since the spine bones contain strong prior knowledge of their shapes and positions in ultrasound VPI images, we propose to encode this information into the semantic representations in an attentive manner. We first revisit the self-attention mechanism in representation learning, and then present a strategy to introduce the structural knowledge into the key representation in self-attention. By this means, the network explores both the contextual and structural information in the learned features, and consequently improves the segmentation accuracy. We conduct various experiments to demonstrate that our proposed method achieves promising performance on spine image segmentation, which shows great potential in clinical diagnosis. Rui Zhao 0012, Zixun Huang, Tianshan Liu, Frank H. F. Leung, Sai-Ho Ling, De Yang, Timothy Tin-Yan Lee, Daniel Pak-Kong Lun, Kin-Man Lam 0001 |
ICASSP | 3 |
| 2021 | Multimodal-Semantic Context-Aware Graph Neural Network for Group Activity RecognitionabstractGroup activities in videos involve visual interaction contexts in multiple modalities between actors, and co-occurrence between individual action labels. However, most of the current group activity recognition methods either model actor-actor relations based on the single RGB modality, or ignore exploiting the label relationships. To capture these rich visual and semantic contexts, we propose a multimodal-semantic context-aware graph neural network (MSCA-GNN). Specifically, we first build two visual sub-graphs based on the appearance cues and motion patterns extracted from RGB and optical-flow modalities, respectively. Then, two attention-based aggregators are proposed to refine each node, by gathering representations from other nodes and heterogeneous modalities. In addition, a semantic graph is constructed based on linguistic embeddings to model label relationships. We employ a bi-directional mapping learning strategy to further integrate the information from both multimodal visual and semantic graphs. Experimental results on two group activity benchmarks show the effectiveness of the proposed method. Tianshan Liu, Rui Zhao 0012, Kin-Man Lam 0001 |
ICME | 1 |
| 2021 | Balanced distortion and perception in single-image super-resolution based on optimal transport in wavelet domain
Jun Xiao 0010, Tianshan Liu, Rui Zhao 0012, Kin-Man Lam 0001 |
Neurocomputing | 2 |
| 2021 | Incorporating multi-level CNN and attention mechanism for Chinese clinical named entity recognition
Jun Kong 0001, Min Jiang 0008, Tianshan Liu |
J. Biomed. Informatics | 4 |
| 2021 | Diverse Features Fusion Network for video-based action recognition
Haoyang Deng, Jun Kong 0001, Min Jiang 0008, Tianshan Liu |
J. Vis. Commun. Image Represent. | 4 |
| 2021 | Dynamic Center Aggregation Loss With Mixed Modality for Visible-Infrared Person Re-IdentificationabstractVisible-infrared person re-identification (VI-ReID) is a challenging cross-modality pedestrian retrieval task which aims to match person images between the visible and infrared modality of the same identity. Existing methods usually adopt two-stream network to solve cross-modality gap, but they ignore the pixel-level discrepancy between the visible and infrared images. Some methods introduce auxiliary modalities in the network, but they lack powerful constraints on the feature distribution of multiple modalities. In this letter, we propose a Dynamic Center Aggregation (DCA) loss with mixed modality for VI-ReID. Concretely, we employ a mixed modality as a bridge between the visible and infrared modality, reducing the difference of the two modalities at the pixel-level. The mixed modality is generated by a Dual-modality Feature Mixer (DFM), which combines the features of visible and infrared images. Moreover, we dynamically adjust the relative distance across multi-modality through DCA loss, which is conducive to explore the modality-invariant feature. We evaluate the proposed method on two public available VI-ReID datasets (SYSU-MM01 and RegDB). Experimental results demonstrate that our method achieves competitive performance. Jun Kong 0001, Qibin He 0003, Min Jiang 0008, Tianshan Liu |
IEEE Signal Process. Lett. | 4 |
| 2021 | Invertible Image DecolorizationabstractInvertible image decolorization is a useful color compression technique to reduce the cost in multimedia systems. Invertible decolorization aims to synthesize faithful grayscales from color images, which can be fully restored to the original color version. In this paper, we propose a novel color compression method to produce invertible grayscale images using invertible neural networks (INNs). Our key idea is to separate the color information from color images, and encode the color information into a set of Gaussian distributed latent variables via INNs. By this means, we force the color information lost in grayscale generation to be independent of the input color image. Therefore, the original color version can be efficiently recovered by randomly re-sampling a new set of Gaussian distributed variables, together with the synthetic grayscale, through the reverse mapping of INNs. To effectively learn the invertible grayscale, we introduce the wavelet transformation into a UNet-like INN architecture, and further present a quantization embedding to prevent the information omission in format conversion, which improves the generalizability of the framework in real-world scenarios. Extensive experiments on three widely used benchmarks demonstrate that the proposed method achieves a state-of-the-art performance in terms of both qualitative and quantitative results, which shows its superiority in multimedia communication and storage systems. Rui Zhao 0012, Tianshan Liu, Jun Xiao 0010, Daniel Pak-Kong Lun, Kin-Man Lam 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | Flow-guided Spatial Attention Tracking for Egocentric Activity RecognitionabstractThe popularity of wearable cameras has opened up a new dimension for egocentric activity recognition. While some methods introduce attention mechanisms into deep learning networks to capture fine-grained hand-object interactions, they often neglect exploring the spatio-temporal relationships. Generating spatial attention, without adequately exploiting temporal consistency, will result in potentially sub-optimal performance in the video-based task. In this paper, we propose a flow-guided spatial attention tracking (F-SAT) module, which is based on enhancing motion patterns and inter-frame information, to highlight the discriminative features from regions of interest across a video sequence. A new form of input, namely the optical-flow volume, is presented to provide informative cues from moving parts for spatial attention tracking. The proposed F-SAT module is deployed to a two-branch-based deep architecture, which fuses complementary information for egocentric activity recognition. Experimental results on three egocentric activity benchmarks show that the proposed method achieves state-of-the-art performance. Tianshan Liu, Kin-Man Lam 0001 |
ICPR | 1 |
| 2020 | Deep Multi-task Learning for Facial Expression Recognition and Synthesis Based on Selective Feature SharingabstractMulti-task learning is an effective learning strategy for deep-learning-based facial expression recognition tasks. However, most existing methods take into limited consideration the feature selection, when transferring information between different tasks, which may lead to task interference when training the multi-task networks. To address this problem, we propose a novel selective feature-sharing method, and establish a multi-task network for facial expression recognition and facial expression synthesis. The proposed method can effectively transfer beneficial features between different tasks, while filtering out useless and harmful information. Moreover, we employ the facial expression synthesis task to enlarge and balance the training dataset to further enhance the generalization ability of the proposed method. Experimental results show that the proposed method achieves state-of-the-art performance on those commonly used facial expression recognition benchmarks, which makes it a potential solution to real-world facial expression recognition problems. Rui Zhao 0012, Tianshan Liu, Jun Xiao 0010, Daniel Pak-Kong Lun, Kin-Man Lam 0001 |
ICPR | 2 |
| 2020 | Multiple depth-levels features fusion enhanced network for action recognition
Shengquan Wang, Jun Kong 0001, Min Jiang 0008, Tianshan Liu |
J. Vis. Commun. Image Represent. | 4 |
| 2020 | Progressive Motion Representation Distillation With Two-Branch Networks for Egocentric Activity RecognitionabstractVideo-based egocentric activity recognition involves fine-grained spatio-temporal human-object interactions. State-of-the-art methods, based on the two-branch-based architecture, rely on pre-calculated optical flows to provide motion information. However, this two-stage strategy is computationally intensive, storage demanding, and not task-oriented, which hampers it from being deployed in real-world applications. Albeit there have been numerous attempts to explore other motion representations to replace optical flows, most of the methods were designed for third-person activities, without capturing fine-grained cues. To tackle these issues, in this letter, we propose a progressive motion representation distillation (PMRD) method, based on two-branch networks, for egocentric activity recognition. We exploit a generalized knowledge distillation framework to train a hallucination network, which receives RGB frames as input and produces motion cues guided by the optical-flow network. Specifically, we propose a progressive metric loss, which aims to distill local fine-grained motion patterns in terms of each temporal progress level. To further enforce the proposed distillation framework to concentrate on those informative frames, we integrate a temporal attention mechanism into the metric loss. Moreover, a multi-stage training procedure is employed for the efficient learning of the hallucination network. Experimental results on three egocentric activity benchmarks demonstrate the state-of-the-art performance of the proposed method. Tianshan Liu, Rui Zhao 0012, Jun Xiao 0010, Kin-Man Lam 0001 |
IEEE Signal Process. Lett. | 1 |
| 2019 | Collaborative multimodal feature learning for RGB-D action recognition
Jun Kong 0001, Tianshan Liu, Min Jiang 0008 |
J. Vis. Commun. Image Represent. | 2 |