Jingjia Huang

dblp:202/1720 · DBLP profile ↗
← Back
23ranked-venue papers
14as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 10 first-author · 10 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 7 since 2021Computer networks · 4 · 4 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Fair SSB Codebook Design for Multi-Cell mmWave MIMO Communications
abstract
For millimeter wave communications, beams used to transmit synchronization signal blocks (SSBs) affect both base station coverage and beam training overhead. We therefore consider fair SSB codebook design, formulated as an optimization problem, aiming to maximize the minimum average signal-to-interference-plus-noise ratio (SINR) across user clusters. This problem is challenging due to the non-smoothness of the objective function, arising from the optimal beam-pair selection function and the minimum operator. To address this, we propose a double-loop framework, where the outer loop constructs approximations for the selection function with iteratively reduced error, and the inner loop solves the resulting approximate problems. In each inner-loop iteration, the objective function of the approximate problem is smoothed with iteratively reduced smoothness, enabling gradient derivation. This gradient is then estimated using variance-reduced estimators based on samples from users, and the result is used to update codebooks. Following this framework, we develop both first-order (FO) and zeroth-order (ZO) oracle schemes. The FO scheme requires full channel state information samples for gradient estimation while the ZO scheme only requires SINR samples. Simulation results show that in given scenarios, both schemes achieve SINR fairness comparable to or better than that of discrete Fourier transform codebooks, but with fewer beams.
Jingjia Huang, Chenhao Qi 0001, Geoffrey Ye Li, Octavia A. Dobre
IEEE Trans. Wirel. Commun.1
2025 Accelerated Diffusion via High-Low Frequency Decomposition for Pan-Sharpening
abstract
Pan-sharpening aims to preserve the spectral information of the multi-spectral (MS) image while leveraging the high-frequency details from the guided high-resolution panchromatic (PAN) image to enhance its spatial resolution. The key challenge is how to preserve the spectral information from the MS image and the spatial details from the PAN image as much as possible. Diffusion models have achieved favorable results in image restoration and synthesis tasks but suffer from excessive computational resource and time consumption. In this paper, we design a novel and computationally efficient diffusion-based pan-sharpening network that achieves accelerated diffusion while reducing task complexity by decoupling the high and low-frequency components of the fused image. Specifically, leveraging the information-preserving characteristic of the wavelet transformation, we introduce a Wavelet-based Low-frequency Diffusion Model (WLDM). WLDM generates the low-frequency coefficient of high-resolution MS (HRMS) image from the low-resolution MS (LRMS) image. This approach significantly reduces computational resources and complexity compared to the direct restoration of the HRMS image. Furthermore, we have devised a High-frequency Information Restoration Module (HIRM) to restore the high-frequency information in the HRMS image through the interaction of high-frequency coefficients from the PAN image in three directions. Extensive experiments on three different datasets demonstrate that our method outperforms existing approaches in both quantitative metrics, qualitative metrics, and inference efficiency.
Ge Meng, Jingjia Huang, Jingyan Tu, Yingying Wang 0005, Yunlong Lin, Xiaotong Tu, Yue Huang 0001, Xinghao Ding
AAAI2
2025 Sp3ctralMamba: Physics-Driven Joint State Space Model for Hyperspectral Image Reconstruction
abstract
Hyperspectral image (HSI) reconstruction aims to restore the original 3D HSIs from the 2D hyperspectral snapshot compressive images (SCIs). The key to high-fidelity HSI reconstruction lies in designing refined spatial and spectral attention mechanisms, which are crucial for generating fine-grained representations of HSI based on the limited spatial and spectral information available in SCI. Recently, Mamba has demonstrated remarkable performance and efficiency in modeling spatial correlations. Its implicit attention mechanism generates three orders of magnitude more attention matrices than transformers, significantly raising the performance ceiling for HSI reconstruction. In this paper, we propose a novel joint SSM network named Sp3ctralMamba for HSI reconstruction. Sp3ctralMamba integrates frequency domain knowledge and physical priors to enhance reconstruction quality. Specifically, we first perform hierarchical decomposition of the 3D HSI embedding to mitigate the negative impact of distant bands on reconstruction. Next, we design a joint SSM block S3Mamba (S3MAB) to perform parallel scans of the embeddings from different bands. In addition to the conventional vanilla scan, S3MAB introduces a local scanning scheme to address the reconstruction challenges posed by the spatial sparsity of spectral information. Furthermore, a spiral scanning scheme in the frequency domain is incorporated to enhance the order correlation between different frequency signals. Finally, we introduce energy priors and structural priors to constrain the generation of spectral and spatial representations during the training process. Extensive experiments on both simulated and real datasets demonstrate that Sp3ctralMamba significantly elevates HSI reconstruction performance to a new level, surpassing SOTA methods in both quantitative and qualitative metrics.
Ge Meng, Jingyan Tu, Jingjia Huang, Yunlong Lin, Yingying Wang 0005, Xiaotong Tu, Yue Huang 0001, Xinghao Ding
AAAI3
2025 Site-Specific Fair SSB Codebook Design for mmWave MIMO-OFDM Communications
abstract
In millimeter wave communications, the initial access stage involves a trade-off between beam sweeping overhead and base station coverage. To balance this, we consider site-specific fair synchronization signal block codebook design, formulated as an optimization problem, aiming to maximize the minimum average signal-to-noise ratio (SNR) across user clusters, subject to constant modulus constraints. To solve this, we propose a double-loop framework, where the outer loop relaxes the original problem into sub-problems using an augmented Lagrangian (AL) algorithm, and the inner loop solves these sub-problems using a hybrid variance-reduced (HVR) stochastic gradient descent (SGD) algorithm. In each inner iteration, the non-smooth objective function of the sub-problem, comprising numerous user-wise SNR functions, is approximated by a differentiable function, with its gradient estimated using HVR estimators. The estimated gradient is then used to update the codebook. Simulation results demonstrate the efficiency of the proposed AL-HVR-SGD scheme.
Jingjia Huang, Chenhao Qi 0001, Octavia A. Dobre
ICC1
2025 Beam Switching Based Beam Design for High-Speed Train mmWave Communications
abstract
For high-speed train (HST) millimeter wave (mmWave) communications, the use of narrow beams with small beam coverage needs frequent beam switching, while wider beams with small beam gain leads to weaker mmWave signal strength. In this paper, we consider beam switching based beam design, which is formulated as an optimization problem aiming to minimize the number of switched beams within a predetermined railway range subject to that the receiving signal-to-noise ratio (RSNR) at the HST is no lower than a predetermined threshold. To solve this problem, we propose two sequential beam design schemes, both including two alternately-performed stages. In the first stage, given an updated beam coverage according to the railway range, we transform the problem into a feasibility problem and further convert it into a min-max optimization problem by relaxing the RSNR constraints into a penalty of the objective function. In the second stage, we evaluate the feasibility of the beamformer obtained from solving the min-max problem and determine the beam coverage accordingly. Simulation results show that compared to the first scheme, the second scheme can achieve 96.20% reduction in computational complexity at the cost of only 0.0657% performance degradation.
Jingjia Huang, Chenhao Qi 0001, Octavia A. Dobre, Geoffrey Ye Li
IEEE Trans. Wirel. Commun.1
2024 Stitching Segments and Sentences towards Generalization in Video-Text Pre-training
abstract
Video-language pre-training models have recently achieved remarkable results on various multi-modal downstream tasks. However, most of these models rely on contrastive learning or masking modeling to align global features across modalities, neglecting the local associations between video frames and text tokens. This limits the model’s ability to perform fine-grained matching and generalization, especially for tasks that selecting segments in long videos based on query texts. To address this issue, we propose a novel stitching and matching pre-text task for video-language pre-training that encourages fine-grained interactions between modalities. Our task involves stitching video frames or sentences into longer sequences and predicting the positions of cross-model queries in the stitched sequences. The individual frame and sentence representations are thus aligned via the stitching and matching strategy, encouraging the fine-grained interactions between videos and texts. in the stitched sequences for the cross-modal query. We conduct extensive experiments on various benchmarks covering text-to-video retrieval, video question answering, video captioning, and moment retrieval. Our results demonstrate that the proposed method significantly improves the generalization capacity of the video-text pre-training models.
Fan Ma, Xiaojie Jin 0004, Jingjia Huang, Linchao Zhu, Yi Yang 0001
AAAI4
2024 Progressive High-Frequency Reconstruction for Pan-Sharpening with Implicit Neural Representation
abstract
Pan-sharpening aims to leverage the high-frequency signal of the panchromatic (PAN) image to enhance the resolution of its corresponding multi-spectral (MS) image. However, deep neural networks (DNNs) tend to prioritize learning the low-frequency components during the training process, which limits the restoration of high-frequency edge details in MS images. To overcome this limitation, we treat pan-sharpening as a coarse-to-fine high-frequency restoration problem and propose a novel method for achieving high-quality restoration of edge information in MS images. Specifically, to effectively obtain fine-grained multi-scale contextual features, we design a Band-limited Multi-scale High-frequency Generator (BMHG) that generates high-frequency signals from the PAN image within different bandwidths. During training, higher-frequency signals are progressively injected into the MS image, and corresponding residual blocks are introduced into the network simultaneously. This design enables gradients to flow from later to earlier blocks smoothly, encouraging intermediate blocks to concentrate on missing details. Furthermore, to address the issue of pixel position misalignment arising from multi-scale features fusion, we propose a Spatial-spectral Implicit Image Function (SIIF) that employs implicit neural representation to effectively represent and fuse spatial and spectral features in the continuous domain. Extensive experiments on different datasets demonstrate that our method outperforms existing approaches in terms of quantitative and visual measurements for high-frequency detail recovery.
Ge Meng, Jingjia Huang, Yingying Wang 0005, Zhenqi Fu, Xinghao Ding, Yue Huang 0001
AAAI2
2024 Efficient Perceiving Local Details via Adaptive Spatial-Frequency Information Integration for Multi-focus Image Fusion
abstract
Multi-focus image fusion (MFIF) aims to combine multiple images with different focused regions into a single all-in-focus image. Existing unsupervised deep learning-based methods only fuse structural information of images in the spatial domain, neglecting potential solutions from the frequency domain exploration. In this paper, we make the first attempt to integrate spatial-frequency information to achieve high-quality MFIF. We propose a novel unsupervised spatial-frequency interaction MFIF network named SFIMFN, which consists of three key components: Adaptive Frequency Domain Information Interaction Module (AFIM), Ret-Attention-Based Spatial Information Extraction Module (RASEM), and Invertible Dual-domain Feature Fusion Module (IDFM). Specifically, in AFIM, we interactively explore global contextual information by combining the amplitude and phase information of multiple images separately. In RASEM, we design a customized transformer to encourage the network to capture important local high-frequency information by redesigning the self-attention mechanism with a bidirectional, two-dimensional form of explicit decay. Finally, we employ IDFM to fuse spatial-frequency information without information loss to generate the desired all-in-focus image. Extensive experiments on different datasets demonstrate that our method significantly outperforms state-of-the-art unsupervised methods in terms of qualitative and quantitative metrics as well as the generalization ability.
Jingjia Huang, Jingyan Tu, Ge Meng, Yingying Wang 0005, Xiaotong Tu, Xinghao Ding, Yue Huang 0001
ACM Multimedia1
2023 Clover: Towards A Unified Video-Language Alignment and Fusion Model
abstract
Building a universal Video-Language model for solving various video understanding tasks (e.g., text-video retrieval, video question answering) is an open challenge to the machine learning field. Towards this goal, most recent works build the model by stacking uni-modal and cross-modal feature encoders and train it with pair-wise contrastive pre-text tasks. Though offering attractive generality, the resulted models have to compromise between efficiency and performance. They mostly adopt different architectures to deal with different downstream tasks. We find this is because the pair-wise training cannot well align and fuse features from different modalities. We then introduce Clover-a Correlated Video-Language pre-training method-towards a universal Video-Language modelfor solving multiple video understanding tasks with neither performance nor efficiency compromise. It improves cross-modal feature alignment and fusion via a novel tri-modal alignment pre-training task. Additionally, we propose to enhance the tri-modal alignment via incorporating learning from semantic masked samples and a new pair-wise ranking loss. Clover establishes new state-of-the-arts on multiple downstream tasks, including three retrieval tasks for both zero-shot and fine-tuning settings, and eight video question answering tasks. Codes and pre-trained models will be released at https://github.com/LeeYN-43/Clover.
Jingjia Huang, Yinan Li 0005, Jiashi Feng, Xiaoshuai Sun, Rongrong Ji
CVPR1
2023 Revisiting Temporal Modeling for CLIP-Based Image-to-Video Knowledge Transferring
abstract
Image-text pretrained models, e.g., CLIP, have shown impressive general multi-modal knowledge learned from large-scale image-text data pairs, thus attracting increasing attention for their potential to improve visual representation learning in the video domain. In this paper, based on the CLIP model, we revisit temporal modeling in the context of image-to-video knowledge transferring, which is the key point for extending image-text pretrained models to the video domain. We find that current temporal modeling mechanisms are tailored to either high-level semantic-dominant tasks (e.g., retrieval) or low-level visual pattern-dominant tasks (e.g., recognition), and fail to work on the two cases simultaneously. The key difficulty lies in modeling temporal dependency while taking advantage of both high-level and low-level knowledge in CLIP model. To tackle this problem, we present Spatial-Temporal Auxiliary Network (STAN) - a simple and effective temporal modeling mechanism extending CLIP model to diverse video tasks. Specifically, to realize both low-level and high-level knowledge transferring, STAN adopts a branch structure with decomposed spatial-temporal modules that enable multi-level CLIP features to be spatial-temporally contextualized. We evaluate our method on two representative video tasks: Video-Text Retrieval and Video Recognition. Extensive experiments demonstrate the superiority of our model over the state-of-the-art methods on various datasets, including MSR-VTT, DiDeMo, LSMDC, MSVD, Kinetics-400, and Something-Something- V2. Codes will be available at https://github.com/farewellthree/STAN
Ruyang Liu, Jingjia Huang, Ge Li 0002, Jiashi Feng, Thomas H. Li
CVPR2
2023 Two-Stage Beamforming Design for High-Speed Train mmWave Communications
abstract
Millimeter wave (mmWave) communications can achieve high data-rate transmission for high-speed trains (HSTs). However, the rapid change in path loss during the fast movement of HSTs poses a significant challenge to the mm Wave beamforming design. In this paper, a two-stage beam-forming (TSB) scheme is proposed to address this challenge for downlink HST mmWave communications. In the first stage, an algorithm based on semi-definite relaxation (SDR) and alternating minimization (AM) is proposed to stabilize the instantaneous receive signal-to-noise ratio (SNR) above a predefined threshold when the HSTs travel along the railway. In the second stage, the coverage of each beam used by the base station (BS) is widened to reduce the number of beam switches. Simulation results demonstrate that the proposed scheme requires fewer BS beams to cover the same railway range than the existing schemes while keeping the instantaneous receive SNR of the HSTs above the predefined threshold.
Jingjia Huang, Chenhao Qi 0001, Octavia A. Dobre
GLOBECOM1
2023 Causality Compensated Attention for Contextual Biased Visual Recognition
Ruyang Liu, Jingjia Huang, Thomas H. Li, Ge Li 0002
ICLR2
2023 DP-INNet: Dual-Path Implicit Neural Network for Spatial and Spectral Features Fusion in Pan-Sharpening
Jingjia Huang, Ge Meng, Yingying Wang 0005, Yunlong Lin, Yue Huang 0001, Xinghao Ding
PRCV (8)1
2022 Learning Disentangled Representation for Multi-View 3D Object Recognition
abstract
3D object recognition is a hot research topic. Particularly, view-based methods, which represent a 3D object with a collection of its rendered views on the 2D domain, play an important role in this field. Currently, view-based researches tend to aggregate information from multiple views via pooling based strategies to endow the models with the characteristic of view permutation invariance, at the cost of inevitable loss of useful features. In this paper, we introduce a new method that learns a more comprehensive descriptor for a 3D object from its views while successfully keeping its robustness to the variation of view permutation. Our method disentangles the information in the set of multi-view images into a global category-related feature and a set of view-permutation related features. To unbind these two parts, an encode-decoder based disentangling architecture is proposed, which barely bring extra computations compared to the baseline model. Systematic experiments are conducted for this new method to demonstrates the effectiveness and the competitive performance based on ModelNet40, ModelNet10, and ShapeNetCore55 datasets. Codes for our paper will be released soon on “https://github.com/hjjpku/multi_view_sort”.
Jingjia Huang, Ge Li 0002, Thomas H. Li, Shan Liu 0001
IEEE Trans. Circuits Syst. Video Technol.1
2022 Learning the Global Descriptor for 3-D Object Recognition Based on Multiple Views Decomposition
abstract
The key point of view based strategies for the analysis of 3D object is to obtain a global descriptor from a collection of its rendered views on 2D images. The views are always redundantly sampled as to ensure the completeness of the information. In this paper, we bring new insight into the study of multi-view object recognition, which models an object as a View Mixture Model (VMM). We argue that each object represented by the multiple views can be decomposed into just a few latent views. Based on the VMM, we introduce a decomposition module to mine the representations of these latent views for the construction of a compact and comprehensive descriptor. After that, we further propose a view alignment module to ensure the descriptor is robust to the variation of view permutation. We evaluate our method on the ModelNet-40, ModelNet-10 and ShapeNetCore55 datasets. The experimental results show that our method can learn efficient and comprehensive representation for 3D objects, and achieves state-of-the-art performance on both the 3D object classification and retrieval tasks. Lastly, experiments are conducted for benchmarking various popular CNN backbones on the 3D object recognition task, with a view to achieving fair comparisons and promoting the future research in this area. Codes for our paper are released: “https://github.com/hjjpku/multi_view_sort”.
Jingjia Huang, Thomas H. Li, Shan Liu 0001, Ge Li 0002
IEEE Trans. Multim.1
2020 Spatial-Temporal Context-Aware Online Action Detection and Prediction
abstract
Spatial-temporal action detection in videos is a challenging problem that has attracted considerable attention in recent years. Most current approaches address action detection as an object detection problem, which utilizes successful object detection frameworks such as Faster R-CNN to operate action detection at every single frame first, and then generates action tubes by linking bounding boxes across the whole video in an offline fashion. However, unlike object detection in static images, temporal context information is vital for action detection in videos. Therefore, we propose an online action detection model that leverages the spatial-temporal context information existing in videos to perform action inference and localization. More specifically, we try to depict the spatial-temporal context pattern of actions via an encoder-decoder model that is based on a convolutional recurrent neural network. The model accepts a video snippet as input and encodes the dynamic information inside the snippet in the forward pass. During the backward pass, the decoder resolves the information for action detection with the current appearance or motion cue at each time stamp. In addition, we devise an incremental action-tube construction algorithm that enables our model to accomplish action prediction ahead of time and performs action detection in an online fashion. To evaluate the performance of our method, we conduct experiments on three popular public datasets UCF-101, UCF-Sports, and J-HMDB-21. The experimental results demonstrate that our method can achieve competitive or superior performance when compared to the state-of-the-art methods. To encourage further research, we release our project on “https://github.com.hjjpku.OATD.”
Jingjia Huang, Nannan Li 0001, Thomas H. Li, Shan Liu 0001, Ge Li 0002
IEEE Trans. Circuits Syst. Video Technol.1
2019 AttPool: Towards Hierarchical Feature Representation in Graph Convolutional Networks via Attention Mechanism
abstract
Graph convolutional networks (GCNs) are potentially short of the ability to learn hierarchical representation for graph embedding, which holds them back in the graph classification task. Here, we propose AttPool, which is a novel graph pooling module based on attention mechanism, to remedy the problem. It is able to select nodes that are significant for graph representation adaptively, and generate hierarchical features via aggregating the attention-weighted information in nodes. Additionally, we devise a hierarchical prediction architecture to sufficiently leverage the hierarchical representation and facilitate the model learning. The AttPool module together with the entire training structure can be integrated into existing GCNs, and is trained in an end-to-end fashion conveniently. The experimental results on several graph-classification benchmark datasets with various scales demonstrate the effectiveness of our method.
Jingjia Huang, Zhangheng Li, Nannan Li 0001, Shan Liu 0001, Ge Li 0002
ICCV1
2019 ARMIN: Towards a More Efficient and Light-weight Recurrent Memory Network
abstract
In recent years, memory-augmented neural networks(MANNs) have shown promising power to enhance the memory ability of neural networks for sequential processing tasks. However, previous MANNs suffer from complex memory addressing mechanism, making them relatively hard to train and causing computational overheads. Moreover, many of them reuse the classical RNN structure such as LSTM for memory processing, causing inefficient exploitations of memory information. In this paper, we introduce a novel MANN, the Auto-addressing and Recurrent Memory Integrating Network (ARMIN) to address these issues. The ARMIN only utilizes hidden state h_t for automatic memory addressing, and uses a novel RNN cell for refined integration of memory information. Empirical results on a variety of experiments demonstrate that the ARMIN is more light-weight and efficient compared to existing memory networks. Moreover, we demonstrate that the ARMIN can achieve much lower computational overhead than vanilla LSTM while keeping similar performances. Codes are available on github.com/zoharli/armin.
Zhangheng Li, Jia-Xing Zhong, Jingjia Huang, Tao Zhang 0069, Thomas H. Li, Ge Li 0002
IJCAI3
2018 SAP: Self-Adaptive Proposal Model for Temporal Action Detection Based on Reinforcement Learning
abstract
Existing action detection algorithms usually generate action proposals through an extensive search over the video at multiple temporal scales, which brings about huge computational overhead and deviates from the human perception procedure. We argue that the process of detecting actions should be naturally one of observation and refinement: observe the current window and refine the span of attended window to cover true action regions. In this paper, we propose a Self-Adaptive Proposal (SAP) model that learns to find actions through continuously adjusting the temporal bounds in a self-adaptive way. The whole process can be deemed as an agent, which is firstly placed at the beginning of the video and traverse the whole video by adopting a sequence of transformations on the current attended region to discover actions according to a learned policy. We utilize reinforcement learning, especially the Deep Q-learning algorithm to learn the agent’s decision policy. In addition, we use temporal pooling operation to extract more effective feature representation for the long temporal window, and design a regression network to adjust the position offsets between predicted results and the ground truth. Experiment results on THUMOS’14 validate the effectiveness of SAP, which can achieve competitive performance with current action detection algorithms via much fewer proposals.
Jingjia Huang, Nannan Li 0001, Tao Zhang 0069, Ge Li 0002, Tiejun Huang 0001, Wen Gao 0001
AAAI1
2018 An Active Action Proposal Method Based on Reinforcement Learning
abstract
Detecting human activities in untrimmed video is a significant yet challenging task. Existing methods usually generate temporal action proposals via searching extensively at multiple preset scales or combining a bunch of short video snippets. However, we argue that the localization of action instances should be a process of observation, refinement and determination: observe the attended temporal window, refine its position and scale, then determine whether a true action region has been accurately found. To this end, we formulate temporal action localization task as a Markov Decision Process, and propose an active temporal action proposal model based on reinforcement learning. Our model learns to localize actions in videos by automatically adjusting the position and span of temporal window via a sequence of transformations. We train an action/non-action binary classifier to determine whether a temporal window contains an action instance. Validation results on THUMOS'14 dataset show that our proposed method achieves competitive performance both in accuracy and efficiency compared with some state-of-the-art methods, while using much less proposals.
Tao Zhang 0069, Nannan Li 0001, Jingjia Huang, Jia-Xing Zhong, Ge Li 0002
ICIP3
2018 Online Action Tube Detection via Resolving the Spatio-temporal Context Pattern
abstract
At present, spatio-temporal action detection in the video is still a challenging problem, considering the complexity of the background, the variety of the action or the change of the viewpoint in the unconstrained environment. Most of current approaches solve the problem via a two-step processing: first detecting actions at each frame; then linking them, which neglects the continuity of the action and operates in an offline and batch processing manner. In this paper, we attempt to build an online action detection model that introduces the spatio-temporal coherence existed among action regions when performing action category inference and position localization. Specifically, we seek to represent the spatio-temporal context pattern via establishing an encoder-decoder model based on the convolutional recurrent network. The model accepts a video snippet as input and encodes the dynamic information of the action in the forward pass. During the backward pass, it resolves such information at each time instant for action detection via fusing the current static or motion cue. Additionally, we propose an incremental action tube generation algorithm, which accomplishes action bounding-boxes association, action label determination and the temporal trimming in a single pass. Our model takes in the appearance, motion or fused signals as input and is tested on two prevailing datasets, UCF-Sports and UCF-101. The experiment results demonstrate the effectiveness of our method which achieves a performance superior or comparable to compared existing approaches.
Jingjia Huang, Nannan Li 0001, Jia-Xing Zhong, Thomas H. Li, Ge Li 0002
ACM Multimedia1
2018 Detecting action tubes via spatial action estimation and temporal path inference
Nannan Li 0001, Jingjia Huang, Thomas H. Li, Huiwen Guo, Ge Li 0002
Neurocomputing2
2017 A Violence Detection Approach Based on Spatio-temporal Hypergraph Transition
Jingjia Huang, Ge Li 0002, Nannan Li 0001, Ronggang Wang, Wenmin Wang 0001
CAIP (2)1