VLDB 2026 Research / reviewers in the wild / expert
Thomas H. Li
dblp:213/4037
· DBLP profile ↗
70ranked-venue papers
0as first author
46since 2021 · last 2025
0000-0001-6123-1265ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 55 · 35 since 2021Artificial intelligence and machine learning · 29 · 20 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning Semantic Facial Descriptors for Accurate Face AnimationabstractFace animation is a challenging task. Existing model-based methods (utilizing 3DMMs or landmarks) often result in a model-like reconstruction effect, which doesn't effectively preserve identity. Conversely, model-free approaches face challenges in attaining a decoupled and semantically rich feature space, thereby making accurate motion transfer difficult to achieve. We introduce the semantic facial descriptors in learnable disentangled vector space to address the dilemma. The approach involves decoupling the facial space into identity and motion subspaces while endowing each of them with semantics by learning complete orthogonal basis vectors. We obtain basis vector coefficients by employing an encoder on the source and driving faces, leading to effective facial descriptors in the identity and motion subspaces. Ultimately, these descriptors can be recombined as latent codes to animate faces. Our approach successfully addresses the issue of model-based methods' limitations in high-fidelity identity and the challenges faced by model-free methods in accurate motion transfer. Extensive experiments are conducted on three challenging benchmarks (i.e. VoxCeleb, HDTF, CelebV). Comprehensive quantitative and qualitative results demonstrate that our model outperforms SOTA methods with superior identity preservation and motion transfer. Yuanqi Chen, Thomas H. Li |
ICASSP | 4 |
| 2025 | SPU+: Dimension Folding for Semantic Point Cloud UpsamplingabstractSemantic Point Cloud Upsampling (SPU) aims to reconstruct a high-resolution (dense) 3D point cloud from a low-resolution (sparse) one, ensuring that the upsampled point cloud is easily recognizable by downstream tasks. Conventional upsampling architectures typically represent point clouds using high-dimensional feature vectors. However, we observe a dimensional bottleneck, where simply increasing the feature dimensionality does not necessarily improve performance on semantic tasks. This insight motivates us to explore more effective feature representations within upsampling networks. In this paper, we propose a novel SPU method called SPU+, which introduces dimension folding as an alternative strategy for handling high-dimensional features. Specifically, SPU+ decomposes each high-dimensional feature into several g-dimensional packages, allowing interactions among packages within the feature space. Guided by the principle of maximizing feature diversity, we determine that setting the package dimension to 3 yields optimal performance. To enable convolutional operations over these 3D packages, we present a 3D Residual Graph Convolution Block (3D-RGCB) that achieves high computational efficiency. Based on 3D-RGCBs, we design an upsampling network that incorporates three structural modes: pre-mode, middle-mode, and end-mode. Additionally, for large-scale upsampling, we develop a scaling-and-shuffling strategy that adaptively adjusts the spatial size of each 3D package. Finally, we analyze the covering number of the 3D package representation and compare it to traditional high-dimensional feature representations. Experiments on publicly available datasets demonstrate not only the effectiveness of dimension folding but also the state-of-the-art performance achieved by SPU+. Code is available at: https://github.com/lizhuangzi/SPU_plus. Zhuangzi Li, Thomas H. Li, Shan Liu 0001, Ge Li 0002 |
IEEE Trans. Image Process. | 2 |
| 2025 | SDE2D: Semantic-Guided Discriminability Enhancement Feature Detector and DescriptorabstractLocal feature detectors and descriptors serve various computer vision tasks, such as image matching, visual localization, and 3D reconstruction. To address the extreme variations of rotation and light in the real world, most detectors and descriptors capture as much invariance as possible. However, these methods ignore feature discriminability and perform poorly in indoor scenes. Indoor scenes have too many weak-textured and even repeatedly textured regions, so it is necessary for the extracted features to possess sufficient discriminability. Therefore, we propose a semantic-guided method (called SDE2D) enhancing feature discriminability to improve the performance of descriptors for indoor scenes. We develop a kind of semantic-guided discriminability enhancement (SDE) loss function that uses semantic information from indoor scenes. To the best of our knowledge, this is the first deep research that applies semantic segmentation to enhance discriminability. In addition, we design a novel framework that allows semantic segmentation network to be well embedded as a module in the overall framework and provides guidance information for training. Besides, we explore the impact of different semantic segmentation models on our method. The experimental results on indoor scenes datasets demonstrate that the proposed SDE2D performs well compared with the state-of-the-art models. Ruonan Zhang 0002, Ge Li 0002, Thomas H. Li |
IEEE Trans. Multim. | 4 |
| 2024 | BT-Adapter: Video Conversation is Feasible Without Video Instruction TuningabstractThe recent progress in Large Language Models (LLM) has spurred various advancements in image-language con-versation agents, while how to build a proficient video-based dialogue system is still under exploration. Consid-ering the extensive scale of LLM and visual backbone, min-imal GPU memory is left for facilitating effective temporal modeling, which is crucial for comprehending and providing feedback on videos. To this end, we propose Branching Temporal Adapter (BT-Adapter), a novel method for ex-tending image-language pretrained models into the video domain. Specifically, BT-Adapter serves as a plug-and-use temporal modeling branch alongside the pretrained vi-sual encoder, which is tuned while keeping the backbone frozen. Just pretrained once, BT-Adapter can be seamlessly integrated into all image conversation models using this version of CLIP, enabling video conversations without the need for video instructions. Besides, we develop a unique asymmetric token masking strategy inside the branch with tailor-made training tasks for BT-Adapter, facilitating faster convergence and better results. Thanks to BT-Adapter, we are able to empower existing multimodal dialogue models with strong video understanding capabilities without incur-ring excessive GPU costs. Without bells and whistles, BT-Adapter achieves (1) state-of-the-art zero-shot results on various video tasks using thousands of fewer GPU hours. (2) better performance than current video chatbots without any video instruction tuning. (3) state-of-the-art results of video chatting using video instruction tuning, outperforming previous SOTAs by a large margin. The code has been available at https://github.com/farewellthreeIBT-Adapter. Ruyang Liu, Chen Li 0046, Yixiao Ge, Thomas H. Li, Ying Shan, Ge Li 0002 |
CVPR | 4 |
| 2024 | ScanPCGC: Learning-Based Lossless Point Cloud Geometry Compression using Sequential Slice RepresentationabstractThe efficient storage and transportation requirements of point clouds promote the development of point cloud compression algorithms. In this paper, we develop a novel point cloud geometry compression using sequential slice representation. Unlike the limited contexts in conventional codecs and other voxel-based works, sufficient contexts are provided by the previous slices, enabling more accurate modeling of the current slice distribution. To reduce the sparsity of the point cloud, we further divide each slice into patches and reorganize non-empty patches along with contexts fed into the conditional entropy model. The 3D convolution-based entropy model with residual structure is designed to exploit sufficient contexts and estimate a probability distribution of the voxels in non-empty patches. In addition to auto-regressive context, we provide a grouped context to address the serial decoding issue. Experimental results on object point cloud datasets (e.g., MPEG 8i, MVUB) demonstrate that our approaches outperform MPEG G-PCC and competitive learning-based methods. Jiangwei Deng, Yuhao An, Thomas H. Li, Shan Liu 0001, Ge Li 0002 |
ICASSP | 3 |
| 2024 | EPContrast: Effective Point-level Contrastive Learning for Large-scale Point Cloud UnderstandingabstractThe acquisition of inductive bias through pointlevel contrastive learning holds paramount significance in point cloud pre-training. However, the square growth in computational requirements with the scale of the point cloud poses a substantial impediment to the practical deployment and execution. To address this challenge, this paper proposes an Effective Pointlevel Contrastive Learning method for large-scale point cloud understanding dubbed EPContrast, which consists of AGContrast and ChannelContrast. In practice, AGContrast constructs positive and negative pairs based on asymmetric granularity embedding, while ChannelContrast imposes contrastive supervision between channel feature maps. EPContrast offers point-level contrastive loss while concurrently mitigating the computational resource burden. The efficacy of EPContrast is substantiated through comprehensive validation on S3DIS and ScanNetV2, encompassing tasks such as semantic segmentation, instance segmentation, and object detection. In addition, rich ablation experiments demonstrate remarkable bias induction capabilities under label-efficient and one-epoch training settings. Zhiyi Pan 0001, Wei Gao 0003, Thomas H. Li |
ICME | 4 |
| 2024 | Sketch-aided Interactive Fusion Point Cloud Place RecognitionabstractExisting point cloud place recognition methods ignore textureless descriptions of scenes by point clouds. This further leads to lower generalization and bottlenecks in performance improvement. To solve these problems, we propose a novel sketch-aided interactive fusion point cloud place recognition method, which involves two networks to separately deal with point clouds and sketches and an interaction feature fused module to fuse features mathematically. Specifically, this is the first time to introduce sketches to guide the point cloud place recognition task as far as we know. The sketch-aided part and the point cloud could enhance the texture structure of the scene which is omitted in only the point cloud scenario. Meanwhile, we devise an interactive feature fusion module for fusing two features, which is encouraged by square summation in math. This module reflects the communication between features as well as the non-linear influence on the fused feature without bringing dimension growth. The experiments on two datasets witness the effectiveness of the proposed method in performance improvement and generalization subjectively and objectively. Ruonan Zhang 0002, Ge Li 0002, Thomas H. Li |
ICMR | 4 |
| 2024 | StreamFlow: Streamlined Multi-Frame Optical Flow Estimation for Video SequencesabstractPrior multi-frame optical flow methods typically estimate flow repeatedly in a pair-wise manner, leading to significant computational redundancy. To mitigate this, we implement a Streamlined In-batch Multi-frame (SIM) pipeline, specifically tailored to video inputs to minimize redundant calculations. It enables the simultaneous prediction of successive unidirectional flows in a single forward pass, boosting processing speed by 44.43% and reaching efficiencies on par with two-frame networks. Moreover, we investigate various spatiotemporal modeling methods for optical flow estimation within this pipeline. Notably, we propose a simple yet highly effective parameter-efficient Integrative spatiotemporal Coherence (ISC) modeling method, alongside a lightweight Global Temporal Regressor (GTR) to harness temporal cues. The proposed ISC and GTR bring powerful spatiotemporal modeling capabilities and significantly enhance accuracy, including in occluded areas, while adding modest computations to the SIM pipeline. Compared to the baseline, our approach, StreamFlow, achieves performance enhancements of 15.45% and 11.37% on the Sintel clean and final test sets respectively, with gains of 15.53% and 10.77% on occluded regions and only a 1.11% rise in latency. Furthermore, StreamFlow exhibits state-of-the-art cross-dataset testing results on Sintel and KITTI, demonstrating its robust cross-domain generalization capabilities. The code is available [here](https://github.com/littlespray/StreamFlow). Shangkun Sun, Huaxia Li, Thomas H. Li, Wei Gao 0003 |
NeurIPS | 5 |
| 2024 | ComPoint: Can Complex-Valued Representation Benefit Point Cloud Place Recognition?abstract“Where was this place?”-figuring out the location of a point cloud scene is a challenge that has attracted researchers in recent years, under the name of point cloud place recognition. Driven by the drastic acceleration of 3D data and corresponding technique forces, research in this field has witnessed remarkable progress. However, the existing methods are stuck in a dilemma, the limited capability of feature representations needs more complicated architectures to enhance the performance further. This inspires us to envision whether there is a better representation for this task. To explore its possibility, in this paper, we propose a new framework, dubbed ComPoint, in the form of complex-valued representations for large-scale point cloud place recognition. Theoretically, the framework is guided by two proven propositions where one implies that richer information provided by the complex-valued representations of point clouds can benefit the performance. Practically speaking, ComPoint is designed with three modules, with each module highlighting its different characteristics. First, Com-Transform guarantees informative data delivery by mining informative complex-valued representations of initial point clouds. Next, Com-Perception perceives and digs deeper into complex-valued features via a series of simple-design convolution blocks, i.e.,ComplexPointConvandComplexPointFT. Then, Com-Fusion dynamic aggregates and interacts with the above features to obtain compact global complex-valued ones based on devised effective soft-balancing block in the VLAD network without involving extra memory footprint. Finally, our method is trained with proper strategies that are analyzed in-depth. The proposed method is witnessed to outperform the prior methods on four large-scale benchmarks quantitatively and qualitatively. It is also flexible plug-and-play in other approaches to improve their performance. Ruonan Zhang 0002, Ge Li 0002, Wei Gao 0003, Thomas H. Li |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Closing the Gap Between Theory and Practice During Alternating Optimization for GANsabstractSynthesizing high-quality and diverse samples is the main goal of generative models. Despite recent great progress in generative adversarial networks (GANs), mode collapse is still an open problem, and mitigating it will benefit the generator to better capture the target data distribution. This article rethinks alternating optimization in GANs, which is a classic approach to training GANs in practice. We find that the theory presented in the original GANs does not accommodate this practical solution. Under the alternating optimization manner, the vanilla loss function provides an inappropriate objective for the generator. This objective forces the generator to produce the output with the highest discriminative probability of the discriminator, which leads to mode collapse in GANs. To address this problem, we introduce a novel loss function for the generator to adapt to the alternating optimization nature. When updating the generator by the proposed loss function, the reverse Kullback-Leibler divergence between the model distribution and the target distribution is theoretically optimized, which encourages the model to learn the target distribution. The results of extensive experiments demonstrate that our approach can consistently boost model performance on various datasets and network structures. Yuanqi Chen, Shangkun Sun, Ge Li 0002, Wei Gao 0003, Thomas H. Li |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Hard Sample Matters a Lot in Zero-Shot QuantizationabstractZero-shot quantization (ZSQ) is promising for compressing and accelerating deep neural networks when the data for training full-precision models are inaccessible. In ZSQ, network quantization is performed using synthetic samples, thus, the performance of quantized models depends heavily on the quality of synthetic samples. Nonetheless, we find that the synthetic samples constructed in existing ZSQ methods can be easily fitted by models. Accordingly, quantized models obtained by these methods suffer from significant performance degradation on hard samples. To address this issue, we propose HArd sample Synthesizing and Training (HAST). Specifically, HAST pays more attention to hard samples when synthesizing samples and makes synthetic samples hard to fit when training quantized models. HAST aligns features extracted by full-precision and quantized models to ensure the similarity between features extracted by these two models. Extensive experiments show that HAST significantly outperforms existing ZSQ methods, achieving performance comparable to models that are quantized with real data. Huantong Li, Xiangmiao Wu, Fanbing Lv, Daihai Liao, Thomas H. Li, Yonggang Zhang 0003, Bo Han 0003, Mingkui Tan |
CVPR | 5 |
| 2023 | Revisiting Temporal Modeling for CLIP-Based Image-to-Video Knowledge TransferringabstractImage-text pretrained models, e.g., CLIP, have shown impressive general multi-modal knowledge learned from large-scale image-text data pairs, thus attracting increasing attention for their potential to improve visual representation learning in the video domain. In this paper, based on the CLIP model, we revisit temporal modeling in the context of image-to-video knowledge transferring, which is the key point for extending image-text pretrained models to the video domain. We find that current temporal modeling mechanisms are tailored to either high-level semantic-dominant tasks (e.g., retrieval) or low-level visual pattern-dominant tasks (e.g., recognition), and fail to work on the two cases simultaneously. The key difficulty lies in modeling temporal dependency while taking advantage of both high-level and low-level knowledge in CLIP model. To tackle this problem, we present Spatial-Temporal Auxiliary Network (STAN) - a simple and effective temporal modeling mechanism extending CLIP model to diverse video tasks. Specifically, to realize both low-level and high-level knowledge transferring, STAN adopts a branch structure with decomposed spatial-temporal modules that enable multi-level CLIP features to be spatial-temporally contextualized. We evaluate our method on two representative video tasks: Video-Text Retrieval and Video Recognition. Extensive experiments demonstrate the superiority of our model over the state-of-the-art methods on various datasets, including MSR-VTT, DiDeMo, LSMDC, MSVD, Kinetics-400, and Something-Something- V2. Codes will be available at https://github.com/farewellthree/STAN Ruyang Liu, Jingjia Huang, Ge Li 0002, Jiashi Feng, Thomas H. Li |
CVPR | 6 |
| 2023 | CAT: LoCalization and IdentificAtion Cascade Detection Transformer for Open-World Object DetectionabstractOpen-world object detection (OWOD), as a more general and challenging goal, requires the model trained from data on known objects to detect both known and unknown objects and incrementally learn to identify these unknown objects. The existing works which employ standard detection framework and fixed pseudo-labelling mechanism$(PLM)$have the following problems: (i) The inclusion of detecting unknown objects substantially reduces the model's ability to detect known ones. (ii) The$PLM$does not adequately utilize the priori knowledge of inputs. (iii) The fixed selection manner of$PLM$cannot guarantee that the model is trained in the right direction. We observe that humans subconsciously prefer to focus on all foreground objects and then identify each one in detail, rather than localize and identify a single object simultaneously, for alleviating the confusion. This motivates us to propose a novel solution called CAT: LoCalization and IdentificAtion Cascade Detection Transformer which decouples the detection process via the shared decoder in the cascade decoding way. In the meanwhile, we propose the self-adaptive pseudo-labelling mechanism which combines the model-driven with input-driven$PLM$and self-adaptively generates robust pseudo-labels for unknown objects, significantly improving the ability of CAT to retrieve unknown objects. Experiments on two benchmarks, i.e., MS-COCO and PASCAL VOC, show that our model outperforms the state-of-the-art methods. The code is publicly available at https://github.com/xiaomabufei/CAT. Shuailei Ma, Ying Wei 0007, Thomas H. Li, Fanbing Lv |
CVPR | 5 |
| 2023 | Masked Motion Encoding for Self-Supervised Video Representation LearningabstractHow to learn discriminative video representation from unlabeled videos is challenging but crucial for video analysis. The latest attempts seek to learn a representation model by predicting the appearance contents in the masked regions. However, simply masking and recovering appearance contents may not be sufficient to model temporal clues as the appearance contents can be easily reconstructed from a single frame. To overcome this limitation, we present Masked Motion Encoding (MME), a new pretraining paradigm that reconstructs both appearance and motion information to explore temporal clues. In MME, we focus on addressing two critical challenges to improve the representation performance: 1) how to well represent the possible long-term motion across multiple frames; and 2) how to obtain fine-grained temporal clues from sparsely sampled videos. Motivated by the fact that human is able to recognize an action by tracking objects' position changes and shape changes, we propose to reconstruct a motion trajectory that represents these two kinds of change in the masked regions. Besides, given the sparse video input, we enforce the model to reconstruct dense motion trajectories in both spatial and temporal dimensions. Pre-trained with our MME paradigm, the model is able to anticipate long-term and fine-grained motion details. Code is available at https://github.com/XinyuSun/MME. Peihao Chen, Liangwei Chen, Thomas H. Li, Mingkui Tan, Chuang Gan 0001 |
CVPR | 5 |
| 2023 | Improving Graph Representation for Point Cloud Segmentation via Attentive FilteringabstractRecently, self-attention networks achieve impressive performance in point cloud segmentation due to their superiority in modeling long-range dependencies. However, compared to self-attention mechanism, we find graph convolutions show a stronger ability in capturing local geometry information with less computational cost. In this paper, we employ a hybrid architecture design to construct our Graph Convolution Network with Attentive Filtering (AF-GCN), which takes advantage of both graph convolution and selfattention mechanism. We adopt graph convolutions to aggregate local features in the shallow encoder stages, while in the deeper stages, we propose a self-attention-like module named Graph Attentive Filter (GAF) to better model long-range contexts from distant neighbors. Besides, to further improve graph representation for point cloud segmentation, we employ a Spatial Feature Projection (SFP) module for graph convolutions which helps to handle spatial variations of unstructured point clouds. Finally, a graphshared down-sampling and up-sampling strategy is introduced to make full use of the graph structures in point cloud processing. We conduct extensive experiments on multiple datasets including S3DIS, ScanNetV2, Toronto-3D, and ShapeNetPart. Experimental results show our AF-GCN obtains competitive performance. Nan Zhang 0015, Zhiyi Pan 0001, Thomas H. Li, Wei Gao 0003, Ge Li 0002 |
CVPR | 3 |
| 2023 | Learning Vision-and-Language Navigation from YouTube VideosabstractVision-and-language navigation (VLN) requires an embodied agent to navigate in realistic 3D environments using natural language instructions. Existing VLN methods suffer from training on small-scale environments or unreasonable path-instruction datasets, limiting the generalization to unseen environments. There are massive house tour videos on YouTube, providing abundant real navigation experiences and layout information. However, these videos have not been explored for VLN before. In this paper, we propose to learn an agent from these videos by creating a large-scale dataset which comprises reasonable path-instruction pairs from house tour videos and pre-training the agent on it. To achieve this, we have to tackle the challenges of automatically constructing path-instruction pairs and exploiting real layout knowledge from raw and unlabeled videos. To address these, we first leverage an entropy-based method to construct the nodes of a path trajectory. Then, we propose an action-aware generator for generating instructions from unlabeled trajectories. Last, we devise a trajectory judgment pretext task to encourage the agent to mine the layout knowledge. Experimental results show that our method achieves state-of-the-art performance on two popular benchmarks (R2R and REVERIE). Code is available at https://github.com/JeremyLinky/YouTube-VLN Kunyang Lin, Peihao Chen, Diwei Huang, Thomas H. Li, Mingkui Tan, Chuang Gan 0001 |
ICCV | 4 |
| 2023 | Causality Compensated Attention for Contextual Biased Visual Recognition
Ruyang Liu, Jingjia Huang, Thomas H. Li, Ge Li 0002 |
ICLR | 3 |
| 2023 | LIO-PPF: Fast LiDAR-Inertial Odometry via Incremental Plane Pre-Fitting and Skeleton TrackingabstractAs a crucial infrastructure of intelligent mobile robots, LiDAR-Inertial odometry (LIO) provides the basic capability of state estimation by tracking LiDAR scans. The high-accuracy tracking generally involves the$k\text{NN}$search, which is used with minimizing the point-to-plane distance. The cost for this, however, is maintaining a large local map and performing$k\text{NN}$plane fit for each point. In this work, we reduce both time and space complexity of LIO by saving these unnecessary costs. Technically, we design a plane pre-fitting (PPF) pipeline to track the basic skeleton of the 3D scene. In PPF, planes are not fitted individually for each scan, let alone for each point, but are updated incrementally as the scene ‘flows’. Unlike$k\text{NN}$, the PPF is more robust to noisy and non-strict planes with our iterative Principal Component Analyse (iPCA) refinement. Moreover, a simple yet effective sandwich layer is introduced to eliminate false point-to-plane matches. Our method was extensively tested on a total number of 22 sequences across 5 open datasets, and evaluated in 3 existing state-of-the-art LIO systems. By contrast, LIO-PPF can consume only 36% of the original local map size to achieve up to 4x faster residual computing and 1.92x overall FPS, while maintaining the same level of accuracy. We fully open source our implementation at https://github.com/xingyuuchen/LIO-PPF. Peixi Wu, Ge Li 0002, Thomas H. Li |
IROS | 4 |
| 2023 | PDE-based Progressive Prediction Framework for Attribute Compression of 3D Point CloudsabstractIn recent years, the diffusion-based image compression scheme has achieved significant success, which inspires us to use diffusion theory to employ the diffusion model for point cloud attribute compression. However, the relevant existing methods cannot be used to deal with our task due to the irregular structure of point clouds. To handle this, we propose the partial differential equation (PDE) based progressive prediction framework for attribute compression of 3D point clouds. Firstly, we propose a PDE-based prediction module, which performs prediction by optimizing attribute gradients, allowing the geometric distribution of adjacent areas to be fully utilized and explaining the weighting method for prediction. Besides, we propose a low-complexity method for calculating partial derivative operations on point clouds to address the uncertainty of neighbor occupancy in three-dimensional space. In the proposed prediction framework, we design a two-layer level of detail (LOD) structure, where the attribute information in the high level is used for interpolating the low level by edge-enhancing anisotropic diffusion (EED) to infer local features from the high-level information. After the diffusion-based interpolation, we design a texture-wise prediction method making use of interpolated values and texture information. Experiment results show that our proposed framework achieves an average of 12.00% BD-rate reduction and 1.59% bitrate saving compared with Predlift (PLT) under attribute near-lossless and attribute lossless conditions, respectively. Furthermore, additional experiments demonstrate our proposed scheme has better texture preservation and subjective quality. Yiting Shao, Shan Liu 0001, Thomas H. Li, Ge Li 0002 |
ACM Multimedia | 4 |
| 2023 | Efficient Test-Time Adaptation for Super-Resolution with Second-Order Degradation and ReconstructionabstractImage super-resolution (SR) aims to learn a mapping from low-resolution (LR) to high-resolution (HR) using paired HR-LR training images. Conventional SR methods typically gather the paired training data by synthesizing LR images from HR images using a predetermined degradation model, e.g., Bicubic down-sampling. However, the realistic degradation type of test images may mismatch with the training-time degradation type due to the dynamic changes of the real-world scenarios, resulting in inferior-quality SR images. To address this, existing methods attempt to estimate the degradation model and train an image-specific model, which, however, is quite time-consuming and impracticable to handle rapidly changing domain shifts. Moreover, these methods largely concentrate on the estimation of one degradation type (e.g., blur degradation), overlooking other degradation types like noise and JPEG in real-world test-time scenarios, thus limiting their practicality. To tackle these problems, we present an efficient test-time adaptation framework for SR, named SRTTA, which is able to quickly adapt SR models to test domains with different/unknown degradation types. Specifically, we design a second-order degradation scheme to construct paired data based on the degradation type of the test image, which is predicted by a pre-trained degradation classifier. Then, we adapt the SR model by implementing feature-level reconstruction learning from the initial test image to its second-order degraded counterparts, which helps the SR model generate plausible HR images. Extensive experiments are conducted on newly synthesized corrupted DIV2K datasets with 8 different degradations and several real-world datasets, demonstrating that our SRTTA framework achieves an impressive improvement over existing methods with satisfying speed. The source code is available at https://github.com/DengZeshuai/SRTTA. Zeshuai Deng, Zhuokun Chen, Shuaicheng Niu, Thomas H. Li, Bohan Zhuang, Mingkui Tan |
NeurIPS | 4 |
| 2023 | FGPrompt: Fine-grained Goal Prompting for Image-goal NavigationabstractLearning to navigate to an image-specified goal is an important but challenging task for autonomous systems like household robots. The agent is required to well understand and reason the location of the navigation goal from a picture shot in the goal position. Existing methods try to solve this problem by learning a navigation policy, which captures semantic features of the goal image and observation image independently and lastly fuses them for predicting a sequence of navigation actions. However, these methods suffer from two major limitations. 1) They may miss detailed information in the goal image, and thus fail to reason the goal location. 2) More critically, it is hard to focus on the goal-relevant regions in the observation image, because they attempt to understand observation without goal conditioning. In this paper, we aim to overcome these limitations by designing a Fine-grained Goal Prompting (\sexyname) method for image-goal navigation. In particular, we leverage fine-grained and high-resolution feature maps in the goal image as prompts to perform conditioned embedding, which preserves detailed information in the goal image and guides the observation encoder to pay attention to goal-relevant regions. Compared with existing methods on the image-goal navigation benchmark, our method brings significant performance improvement on 3 benchmark datasets (\textit{i.e.,} Gibson, MP3D, and HM3D). Especially on Gibson, we surpass the state-of-the-art success rate by 8\% with only 1/50 model size. Peihao Chen, Jugang Fan, Jian Chen 0011, Thomas H. Li, Mingkui Tan |
NeurIPS | 5 |
| 2023 | IPFR: Identity-Preserving Face Reenactment with Enhanced Domain Adversarial Training and Multi-level Identity Priors
Ge Li 0002, Yuanqi Chen, Thomas H. Li |
PRCV (10) | 4 |
| 2023 | Frequency-Aware Self-Supervised Monocular Depth EstimationabstractWe present two versatile methods to generally enhance self-supervised monocular depth estimation (MDE) models. The high generalizability of our methods is achieved by solving the fundamental and ubiquitous problems in photometric loss function. In particular, from the perspective of spatial frequency, we first propose Ambiguity-Masking to suppress the incorrect supervision under photometric loss at specific object boundaries, the cause of which could be traced to pixel-level ambiguity. Second, we present a novel frequency-adaptive Gaussian low-pass filter, designed to robustify the photometric loss in high-frequency regions. We are the first to propose blurring images to improve depth estimators with an interpretable analysis. Both modules are lightweight, adding no parameters and no need to manually change the network structures. Experiments show that our methods provide performance boosts to a large number of existing models, including those who claimed state-of-the-art, while introducing no extra inference computation at all. Thomas H. Li, Ruonan Zhang 0002, Ge Li 0002 |
WACV | 2 |
| 2023 | Self-Supervised Monocular Depth Estimation: Solving the Edge-Fattening ProblemabstractSelf-supervised monocular depth estimation (MDE) models universally suffer from the notorious edge-fattening issue. Triplet loss, as a widespread metric learning strategy, has largely succeeded in many computer vision applications. In this paper, we redesign the patch-based triplet loss in MDE to alleviate the ubiquitous edge-fattening issue. We show two drawbacks of the raw triplet loss in MDE and demonstrate our problem-driven redesigns. First, we present a min. operator based strategy applied to all negative samples, to prevent well-performing negatives sheltering the error of edge-fattening negatives. Second, we split the anchor-positive distance and anchor-negative distance from within the original triplet, which directly optimizes the positives without any mutual effect with the negatives. Extensive experiments show the combination of these two small redesigns can achieve unprecedented results: Our powerful and versatile triplet loss not only makes our model outperform all previous SoTA by a large margin, but also provides substantial performance boosts to a large number of existing models, while introducing no extra inference computation at all. Ruonan Zhang 0002, Ji Jiang, Ge Li 0002, Thomas H. Li |
WACV | 6 |
| 2023 | Mitigating Label Noise in GANs via Enhanced Spectral NormalizationabstractLabel noise is a ubiquitous issue in GANs, which degrades the generalization ability of the discriminator and usually leads to instability when training GANs. This issue stems from both real data and generated data. Previous works either only consider one of these two sources, or are not robust enough to noisy labels. In this paper, we revisit spectral normalization in robust learning with noisy labels. Based on its pros and cons, we propose to combine spectral normalization and weight decay to regularize the discriminator, which enjoys a more robust training process. To extend to conditional GANs, we propose to balance the relative importance of marginal matching and conditional matching in the projection discriminator. The proposed Enhanced Spectral Normalization for Generative Adversarial Networks (ESNGAN) can be easily integrated into various existing GANs frameworks without excessive additional cost. The effectiveness of the proposed method is validated on the CIFAR10, LSUN Church, CelebA, and ImageNet datasets, including the unconditional image generation task and the class-conditional image generation task. We also show that the proposed method can further improve the performance of the high-resolution image generation task. Yuanqi Chen, Cece Jin, Ge Li 0002, Thomas H. Li, Wei Gao 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Semantic Point Cloud UpsamplingabstractDownsampled sparse point clouds are beneficial for data transmission and storage, but they are detrimental for semantic tasks due to information loss. In this paper, we examine an upsampling methodology that significantly reconstructs sparse clouds’ semantic representations. Specifically, we propose a novel semantic point cloud upsampling (SPU) framework for sparse point cloud classification. An SPU consists of two networks, i.e. an upsampling network and a classification network. They are skillfully unified to intensify semantic representations acting on the upsampling process. In the upsampling network, we first propose a novel graph aggregation convolution to construct hierarchical relations on sparse point clouds. To enhance stability and diversity during point upsampling, we then combine point shuffling and pre-interpolation technologies to build an enhanced upsampling module. Furthermore, we adopt the semantic prior information provided by a sparse point cloud to enhance its upsampling quality. The prior information is applied to an attention mechanism that can highlight key positions of the point cloud. We investigate different loss functions and conduct experiments on classical deep point networks, which effectively demonstrate the promising performance of our framework. Zhuangzi Li, Ge Li 0002, Thomas H. Li, Shan Liu 0001, Wei Gao 0003 |
IEEE Trans. Multim. | 3 |
| 2022 | Neural Texture Extraction and Distribution for Controllable Person Image SynthesisabstractWe deal with the controllable person image synthesis task which aims to re-render a human from a reference image with explicit control over body pose and appearance. Observing that person images are highly structured, we propose to generate desired images by extracting and distributing semantic entities of reference images. To achieve this goal, a neural texture extraction and distribution operation based on double attention is described. This operation first extracts semantic neural textures from reference feature maps. Then, it distributes the extracted neural textures according to the spatial distributions learned from target poses. Our model is trained to predict human images in arbitrary poses, which encourages it to extract disentangled and expressive neural textures representing the appearance of different semantic entities. The disentangled representation further enables explicit appearance control. Neural textures of different reference images can be fused to control the appearance of the interested areas. Experimental comparisons show the superiority of the proposed model. Code is available at https://github.com/RenYurui/Neural-Texture-Extraction-Distribution. Yurui Ren, Ge Li 0002, Shan Liu 0001, Thomas H. Li |
CVPR | 5 |
| 2022 | Attention Guided Invariance Selection for Local Feature DescriptorsabstractTo copy with the extreme variations of illumination and rotation in the real world, popular descriptors have captured more invariance recently, but more invariance makes descriptors less informative. So this paper designs a unique attention guided framework (named AISLFD) to select appropriate invariance for local feature descriptors, which boosts the performance of descriptors even in the scenes with extreme changes. Specifically, we first explore an efficient multi-scale feature extraction module that provides our local descriptors with more useful information. Besides, we propose a novel parallel self-attention module to get meta descriptors with the global receptive field, which guides the invariance selection more correctly. Compared with state-of-the-art methods, our method achieves competitive performance through sufficient experiments. Ge Li 0002, Thomas H. Li |
ICASSP | 3 |
| 2022 | Pointivae: Invertible Variational Autoencoder Framework for 3D Point Cloud GenerationabstractPoint cloud generation is a challenging task and has drawn great attention in 3D vision community. However, existing methods rarely consider to exploit local features, leading to unsatisfactory generated results that lack of high frequency. In this paper, we put forward a novel point cloud generation framework called PointIVAE, which adopts VAE based framework to construct local relations and enhance generating capability. PointIVAE contains three components, including an encoder, a flow model and a decoder. Specially, the encoder aims to aggregate neighborhood relations and provides high-quality latent codes. We then propose the invertible residual coupling stack in the flow model, in order to learn from the latent codes via an invertible manner. Based on the shape latent codes generated by the flow, the decoder converts the input noises into point clouds in an inverse way. Experimental results demonstrate that PointIVAE obtains the SOTA results in both point cloud generation and autoencoding. Ge Li 0002, Ruonan Zhang 0002, Thomas H. Li, Wei Gao 0003 |
ICIP | 4 |
| 2022 | Deep Geometry Post-Processing for Decompressed Point CloudsabstractPoint cloud compression plays a crucial role in reducing the huge cost of data storage and transmission. However, distortions can be introduced into the decompressed point clouds due to quantization. In this paper, we propose a novel learning-based post-processing method to enhance the decompressed point clouds. Specifically, a voxelized point cloud is first divided into small cubes. Then, a 3D convolutional network is proposed to predict the occupancy probability for each location of a cube. We leverage both local and global contexts by generating multi-scale probabilities. These probabilities are progressively summed to predict the results in a coarse-to-fine manner. Finally, we obtain the geometry-refined point clouds based on the predicted probabilities. Different from previous methods, we deal with decompressed point clouds with huge variety of distortions using a single model. Experimental results show that the proposed method can significantly improve the quality of the decompressed point clouds, achieving 9.30dB BDPSNR gain on three representative datasets on average. Ge Li 0002, Dingquan Li, Yurui Ren, Wei Gao 0003, Thomas H. Li |
ICME | 6 |
| 2022 | Fine-Grained Correlation Representation for Graph-Based Point Cloud Attribute CompressionabstractRecent years have witnessed remarkable success of Graph Fourier Transform (GFT) in point cloud attribute compression. A key to good compression performance of GFT is to construct the graph Laplacian matrix that accurately models signal correlation. Nevertheless, existing attribute compression methods based on GFT adopt the distance metric to define the Laplacian matrix, which does not well represent the color correlation in case of a poor relationship between geometry and color. Hence, considering point cloud color variation in space, we propose additional three kinds of fine-grained correlation representation as Laplacian matrices for patches with different texture categories. Furthermore, we utilize texture complexity features as prior to design a stage-wise decision strategy for guiding each patch to choose appropriate correlation representation in low computation complexity. Experimental results demonstrate our method achieves better compression performance compared with other platforms. Meanwhile, additional experiments also adopt Lagrangian Rate-Distortion Optimization (RDO) to choose optimal one from three correlation representations, verifying the effectiveness of our proposed stage-wise decision strategy. Ge Li 0002, Wei Gao 0003, Thomas H. Li |
ICME | 5 |
| 2022 | DKNAS: A Practical Deep Keypoint Extraction Framework Based on Neural Architecture SearchabstractKeypoint extraction including both keypoint detection and description is a fundamental step in a wide range of geometric multimedia applications. In recent years, many learning-based approaches for keypoint extraction emerge and achieve promising results. However, they usually design network architectures empirically and lack of considerations about the comprehensive performance, which leads to limited applications. In this paper, we propose a practical framework based on Neural Architecture Search (NAS) technology, DKNAS, which can search architectures automatically and maintain efficiency and effectiveness, simultaneously. To the best of our knowledge, the proposed framework is the first NAS framework for keypoint extraction. The evaluation on HPatches dataset shows that our method achieves state-of-the-art results in the metrics of repeatability, localization error, homography accuracy and matching scores. Besides, our model is applied to a traditional Simultaneous Localization and Mapping (SLAM) system, ORB-SLAM2, to replace the handcrafted keypoints. Experimental results demonstrate that the system adopting our model outperforms ORB-SLAM2 and some other deep keypoints enhanced systems. Xing Cai, Ge Li 0002, Thomas H. Li |
ICRA | 4 |
| 2022 | Learning Active Camera for Multi-Object NavigationabstractGetting robots to navigate to multiple objects autonomously is essential yet difficult in robot applications. One of the key challenges is how to explore environments efficiently with camera sensors only. Existing navigation methods mainly focus on fixed cameras and few attempts have been made to navigate with active cameras. As a result, the agent may take a very long time to perceive the environment due to limited camera scope. In contrast, humans typically gain a larger field of view by looking around for a better perception of the environment. How to make robots perceive the environment as efficiently as humans is a fundamental problem in robotics. In this paper, we consider navigating to multiple objects more efficiently with active cameras. Specifically, we cast moving camera to a Markov Decision Process and reformulate the active camera problem as a reinforcement learning problem. However, we have to address two new challenges: 1) how to learn a good camera policy in complex environments and 2) how to coordinate it with the navigation policy. To address these, we carefully design a reward function to encourage the agent to explore more areas by moving camera actively. Moreover, we exploit human experience to infer a rule-based camera action to guide the learning process. Last, to better coordinate two kinds of policies, the camera policy takes navigation actions into account when making camera moving decisions. Experimental results show our camera policy consistently improves the performance of multi-object navigation over four baselines on two datasets. Peihao Chen, Dongyu Ji, Kunyang Lin, Weiwen Hu, Wenbing Huang 0001, Thomas H. Li, Mingkui Tan, Chuang Gan 0001 |
NeurIPS | 6 |
| 2022 | Weakly-Supervised Multi-Granularity Map Learning for Vision-and-Language NavigationabstractWe address a practical yet challenging problem of training robot agents to navigate in an environment following a path described by some language instructions. The instructions often contain descriptions of objects in the environment. To achieve accurate and efficient navigation, it is critical to build a map that accurately represents both spatial location and the semantic information of the environment objects. However, enabling a robot to build a map that well represents the environment is extremely challenging as the environment often involves diverse objects with various attributes. In this paper, we propose a multi-granularity map, which contains both object fine-grained details (\eg, color, texture) and semantic classes, to represent objects more comprehensively. Moreover, we propose a weakly-supervised auxiliary task, which requires the agent to localize instruction-relevant objects on the map. Through this task, the agent not only learns to localize the instruction-relevant objects for navigation but also is encouraged to learn a better map representation that reveals object information. We then feed the learned map and instruction to a waypoint predictor to determine the next navigation goal. Experimental results show our method outperforms the state-of-the-art by 4.0% and 4.6% w.r.t. success rate both in seen and unseen environments, respectively on VLN-CE dataset. The code is available at https://github.com/PeihaoChen/WS-MGMap. Peihao Chen, Dongyu Ji, Kunyang Lin, Runhao Zeng, Thomas H. Li, Mingkui Tan, Chuang Gan 0001 |
NeurIPS | 5 |
| 2022 | Rate-Distortion Optimized Graph for Point Cloud Attribute CodingabstractRecent years have witnessed remarkable success of Graph Fourier Transform (GFT) in point cloud attribute compression. Existing researches mainly utilize geometry distance to define graph structure for coding attribute (e.g., color), which may distribute high weights to the edges connecting points across texture boundaries. In this case, these geometry-based graphs cannot model attribute differences between points adequately, thus limiting the compression efficiency of GFT. Hence, we firstly utilize attribute itself to refine the distance-based weight values by setting penalty function, which smoothens signal variations on graph and concentrates more energies in the low frequencies. Then, adjacency matrices acting as penalty function variables are transmitted to decoder with extra bit overheads. To balance the attribute smoothness on graph and the cost of coding adjacency matrices, we finally propose the graph based on Rate-Distortion (RD) optimization and find the optimal adjacency matrix. Experimental results show that our algorithm improves RD performance compared with competitive platforms. Moreover, additional experiments also analyze the gain source by evaluating the effectiveness of RD optimized graphs. Ge Li 0002, Wei Gao 0003, Thomas H. Li |
IEEE Signal Process. Lett. | 4 |
| 2022 | Learning Disentangled Representation for Multi-View 3D Object Recognitionabstract3D object recognition is a hot research topic. Particularly, view-based methods, which represent a 3D object with a collection of its rendered views on the 2D domain, play an important role in this field. Currently, view-based researches tend to aggregate information from multiple views via pooling based strategies to endow the models with the characteristic of view permutation invariance, at the cost of inevitable loss of useful features. In this paper, we introduce a new method that learns a more comprehensive descriptor for a 3D object from its views while successfully keeping its robustness to the variation of view permutation. Our method disentangles the information in the set of multi-view images into a global category-related feature and a set of view-permutation related features. To unbind these two parts, an encode-decoder based disentangling architecture is proposed, which barely bring extra computations compared to the baseline model. Systematic experiments are conducted for this new method to demonstrates the effectiveness and the competitive performance based on ModelNet40, ModelNet10, and ShapeNetCore55 datasets. Codes for our paper will be released soon on “https://github.com/hjjpku/multi_view_sort”. Jingjia Huang, Ge Li 0002, Thomas H. Li, Shan Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | PointOT: Interpretable Geometry-Inspired Point Cloud Generative Model via Optimal TransportabstractPoint cloud generative models have aroused increasing concern for their realistic generation potentialities. However, most existing methods adopt deep-neural-network (DNN) models for continuous mapping. DNN usually induces mode collapse and mixture problems without clear interpretation. Consequently, in this paper, we design a geometry-inspired point cloud generative framework called PointOT. PointOT decouples the generative model into two separate sub-tasks: manifold learning of the point cloud and distribution transformation. Then, we propose corresponding instances according to the requirements in each sub-task, where they can be established by the point cloud auto-encoder (AE) and the semi-continuous optimal transportation (SCOT) mapping, respectively. In particular, the transportation map between the source and the target distributions is discrete rather than continuous in geometric view. The learned continuous shape model of the DNN point cloud does not conform with the discrete distribution transformation. Therefore, the proposed SCOT efficiently relieves these problems by connecting the continuous-to-discrete domain. Besides, we provide theoretical explanations from a geometric view and analyze the fundamental reason for mode collapse and mixture in point cloud generative models. The proposed SCOT algorithm without the DNN model is computationally efficient and makes the original black box semi-transparent. Final experiments validate the virtue of the proposed approach, including the designed decomposition framework and the rigorous theory. Ruonan Zhang 0002, Wei Gao 0003, Ge Li 0002, Thomas H. Li |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | QINet: Decision Surface Learning and Adversarial Enhancement for Quasi-Immune Completion of Diverse Corrupted Point CloudsabstractIn point cloud completion task, most previous works fail to deal with diverse corrupted point clouds with large missing areas. Meanwhile, they are restricted by discrete point clouds lacking smooth surfaces to represent an object, and the resolution of generated point clouds is fixed once their networks are determined. In addition, the evaluation metrics are not specific for this task. Thus, we propose an innovative quasi-immune completion architecture of point cloud calledQINetin this paper, which is inspired by the artificial immunization process in biology. Specifically, to increase robustness and adaptation of the model, we conceive a mask algorithm named onion-peeling to generate diverse corrupted inputs. Meanwhile, two proposed modules are combined together to produce flexible resolution of point clouds, namely the decision surface learning and adversarial enhancement for the latent representation recovery. The first module transforms point clouds to surfaces with a continuous decision boundary function, which the second module is applied to deduce complete surface from corrupted point cloud by the cooperation of reinforcement learning and latent generative adversarial network. Besides, we evaluate the shortcomings of the existing methods and present two novel metrics to support multi-faceted comparisons. Experimental results verify that our approach can generate continuous 3D shapes with optional resolutions compared to other approaches, and achieves competitive results both quantitatively and qualitatively. Ruonan Zhang 0002, Wei Gao 0003, Ge Li 0002, Thomas H. Li |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Learning the Global Descriptor for 3-D Object Recognition Based on Multiple Views DecompositionabstractThe key point of view based strategies for the analysis of 3D object is to obtain a global descriptor from a collection of its rendered views on 2D images. The views are always redundantly sampled as to ensure the completeness of the information. In this paper, we bring new insight into the study of multi-view object recognition, which models an object as a View Mixture Model (VMM). We argue that each object represented by the multiple views can be decomposed into just a few latent views. Based on the VMM, we introduce a decomposition module to mine the representations of these latent views for the construction of a compact and comprehensive descriptor. After that, we further propose a view alignment module to ensure the descriptor is robust to the variation of view permutation. We evaluate our method on the ModelNet-40, ModelNet-10 and ShapeNetCore55 datasets. The experimental results show that our method can learn efficient and comprehensive representation for 3D objects, and achieves state-of-the-art performance on both the 3D object classification and retrieval tasks. Lastly, experiments are conducted for benchmarking various popular CNN backbones on the 3D object recognition task, with a view to achieving fair comparisons and promoting the future research in this area. Codes for our paper are released: “https://github.com/hjjpku/multi_view_sort”. Jingjia Huang, Thomas H. Li, Shan Liu 0001, Ge Li 0002 |
IEEE Trans. Multim. | 3 |
| 2021 | SSD-GAN: Measuring the Realness in the Spatial and Spectral DomainsabstractThis paper observes that there is an issue of high frequencies missing in the discriminator of standard GAN, and we reveal it stems from downsampling layers employed in the network architecture. This issue makes the generator lack the incentive from the discriminator to learn high-frequency content of data, resulting in a significant spectrum discrepancy between generated images and real images. Since the Fourier transform is a bijective mapping, we argue that reducing this spectrum discrepancy would boost the performance of GANs. To this end, we introduce SSD-GAN, an enhancement of GANs to alleviate the spectral information loss in the discriminator. Specifically, we propose to embed a frequency-aware classifier into the discriminator to measure the realness of the input in both the spatial and spectral domains. With the enhanced discriminator, the generator of SSD-GAN is encouraged to learn high-frequency content of real data and generate exact details. The proposed method is general and can be easily integrated into most existing GANs framework without excessive cost. The effectiveness of SSD-GAN is validated on various network architectures, objective functions, and datasets. Code is available at https://github.com/cyq373/SSD-GAN. Yuanqi Chen, Ge Li 0002, Cece Jin, Shan Liu 0001, Thomas H. Li |
AAAI | 5 |
| 2021 | ATVIO: Attention Guided Visual-Inertial OdometryabstractVisual-inertial odometry (VIO) aims to predict trajectory by ego- motion estimation. In recent years, end-to-end VIO has made great progress. However, how to handle visual and inertial measurements and make full use of the complementarity of cameras and inertial sensors remains a challenge. In the paper, we propose a novel attention guided deep framework for visual-inertial odometry (ATVIO) to improve the performance of VIO. Specifically, we extraordinarily concentrate on the effective utilization of the Inertial Measurement Unit (IMU) information. Therefore, we carefully design a one-dimension inertial feature encoder for IMU data processing. The network can extract inertial features quickly and effectively. Meanwhile, we should prevent the inconsistency problem when fusing inertial and visual features. Hence, we explore a novel cross-domain channel attention block to combine the extracted features in a more adaptive manner. Extensive experiments demonstrate that our method achieves competitive performance against state-of-the-art VIO methods. Ge Li 0002, Thomas H. Li |
ICASSP | 3 |
| 2021 | PIRenderer: Controllable Portrait Image Generation via Semantic Neural RenderingabstractGenerating portrait images by controlling the motions of existing faces is an important task of great consequence to social media industries. For easy use and intuitive control, semantically meaningful and fully disentangled parameters should be used as modifications. However, many existing techniques do not provide such fine-grained controls or use indirect editing methods i.e. mimic motions of other individuals. In this paper, a Portrait Image Neural Renderer (PIRenderer) is proposed to control the face motions with the parameters of three-dimensional morphable face models (3DMMs). The proposed model can generate photo-realistic portrait images with accurate movements according to intuitive modifications. Experiments on both direct and indirect editing tasks demonstrate the superiority of this model. Meanwhile, we further extend this model to tackle the audio-driven facial reenactment task by extracting sequential motions from audio inputs. We show that our model can generate coherent videos with convincing movements from only a single reference image and a driving audio stream. Our source code is available at https://github.com/RenYurui/PIRender. Yurui Ren, Ge Li 0002, Yuanqi Chen, Thomas H. Li, Shan Liu 0001 |
ICCV | 4 |
| 2021 | Structure-transformed Texture-enhanced Network for Person Image SynthesisabstractPose-guided virtual try-on task aims to modify the fashion item based on pose transfer task. These two tasks that belong to person image synthesis have strong correlations and similarities. However, existing methods treat them as two individual tasks and do not explore correlations between them. Moreover, these two tasks are challenging due to large misalignment and occlusions, thus most of these methods are prone to generate unclear human body structure and blurry fine-grained textures. In this paper, we devise a structure-transformed texture-enhanced network to generate high-quality person images and construct the relationships between two tasks. It consists of two modules: structure-transformed renderer and texture-enhanced stylizer. The structure-transformed renderer is introduced to transform the source person structure to the target one, while the texture-enhanced stylizer is served to enhance detailed textures and controllably inject the fashion style founded on the structural transformation. With the two modules, our model can generate photorealistic person images in diverse poses and even with various fashion styles. Extensive experiments demonstrate that our approach achieves state-of-the-art results on two tasks. Munan Xu, Yuanqi Chen, Shan Liu 0001, Thomas H. Li, Ge Li 0002 |
ICCV | 4 |
| 2021 | Rethinking Training Objective For Self-Supervised Monocular Depth Estimation: Semantic Cues To RescueabstractMonocular depth estimation finds a wide range of applications in modeling 3D scenes. Since it is expensive to collect ground truth labels to supervise training, plenty of works have been done in a self-supervised manner. A common practice is to train the network optimizing a photometric objective (i.e., view synthesis) due to its effectiveness. However, this training objective is sensitive to optical changes and lacks a consideration of object-level cues, which leads to sub-optimal results in some cases, e.g., artifacts in complex regions and depth discontinuities around thin structures. We summarize them as depth ambiguities. In this paper, we propose an easy yet effective architecture, introducing semantic cues into supervision to solve problems mentioned above. First through our study on the problems we Figure out that they are due to the limitation of the commonly applied photometric reconstruction training objective. Then we come up with our method using semantic cues to encode the geometry constraint behind view synthesis. The proposed novel objective is more credible towards confusing pixels, also takes an object-level perception. Experiments show that without introducing extra inference complexity, our method alleviates depth ambiguities greatly and performs comparably with state-of-the-art methods on KITTI benchmark. Keyao Li, Ge Li 0002, Thomas H. Li |
ICIP | 3 |
| 2021 | Information-Growth Attention Network for Image Super-ResolutionabstractIt is generally known that a high-resolution (HR) image contains more productive information compared with its low-resolution (LR) versions, so image super-resolution (SR) satisfies an information-growth process. Considering the property, we attempt to exploit the growing information via a particular attention mechanism. In this paper, we propose a concise but effective Information-Growth Attention Network (IGAN) that shows the incremental information is beneficial for SR. Specifically, a novel information-growth attention is proposed. It aims to pay attention to features involving large information-growth capacity by assimilating the difference from current features to the former features within a network. We also illustrate its effectiveness contrasted by widely-used self-attention using entropy and generalization analysis. Furthermore, existing channel-wise attention generation modules (CAGMs) have large informational attenuation due to directly calculating global mean for feature maps. Therefore, we present an innovative CAGM that progressively decreases feature maps' sizes, leading to more adequate feature exploitation. Extensive experiments also demonstrate IGAN outperforms state-of-the-art attention-aware SR approaches. Zhuangzi Li, Ge Li 0002, Thomas H. Li, Shan Liu 0001, Wei Gao 0003 |
ACM Multimedia | 3 |
| 2021 | Combining Attention with Flow for Person Image SynthesisabstractPose-guided person image synthesis aims to synthesize person images by transforming reference images into target poses. In this paper, we observe that the commonly used spatial transformation blocks have complementary advantages. We propose a novel model by combining the attention operation with the flow-based operation. Our model not only takes the advantage of the attention operation to generate accurate target structures but also uses the flow-based operation to sample realistic source textures. Both objective and subjective experiments demonstrate the superiority of our model. Meanwhile, comprehensive ablation studies verify our hypotheses and show the efficacy of the proposed modules. Besides, additional experiments on the portrait image editing task demonstrate the versatility of the proposed combination. Yurui Ren, Yubo Wu, Thomas H. Li, Shan Liu 0001, Ge Li 0002 |
ACM Multimedia | 3 |
| 2020 | Over-Exposure Correction via Exposure and Scene Information Disentanglement
Yuhui Cao, Yurui Ren, Thomas H. Li, Ge Li 0002 |
ACCV (4) | 3 |
| 2020 | Deep Image Spatial Transformation for Person Image GenerationabstractPose-guided person image generation is to transform a source person image to a target pose. This task requires spatial manipulations of source data. However, Convolutional Neural Networks are limited by the lack of ability to spatially transform the inputs. In this paper, we propose a differentiable global-flow local-attention framework to reassemble the inputs at the feature level. Specifically, our model first calculates the global correlations between sources and targets to predict flow fields. Then, the flowed local patch pairs are extracted from the feature maps to calculate the local attention coefficients. Finally, we warp the source features using a content-aware sampling method with the obtained local attention coefficients. The results of both subjective and objective experiments demonstrate the superiority of our model. Besides, additional results in video animation and view synthesis show that our model is applicable to other tasks requiring spatial transformation. Our source code is available at https://github.com/RenYurui/Global-Flow-Local-Attention. Yurui Ren, Xiaoming Yu, Thomas H. Li, Ge Li 0002 |
CVPR | 4 |
| 2020 | Regression Before Classification for Temporal Action DetectionabstractAction classification combined with location regression is a widely-utilized mechanism in existing temporal action detection methods. However, there exists an inconsistency problem between locations and categories of action instances in this mechanism. More specifically, while the location of the proposal has been refined by the regressor, the action classifier still uses input and loss corresponding to the outdated unrefined proposal to predict category. In this paper, we propose to eliminate this inconsistency by making two modifi-cations to the action classifier: 1) redirecting the classification loss to the refined proposal, and 2) rearranging the location regressor before the action classifier so that the feature of the refined proposal is fed to the classifier. Extensive experiments show that eliminating the inconsistency problem can significantly promote the detection performance. Our method achieves state-of-the-art performance for temporal action detection on the challenging THUMOS'14 dataset. Cece Jin, Tao Zhang 0069, Weijie Kong, Thomas H. Li, Ge Li 0002 |
ICASSP | 4 |
| 2020 | ROIMIX: Proposal-Fusion Among Multiple Images for Underwater Object DetectionabstractGeneric object detection algorithms have proven their excellent performance in recent years. However, object detection on underwater datasets is still less explored. In contrast to generic datasets, underwater images usually have color shift and low contrast; sediment would cause blurring in underwater images. In addition, underwater creatures often appear closely to each other on images due to their living habits. To address these issues, our work investigates augmentation policies to simulate overlapping, occluded and blurred objects, and we construct a model capable of achieving better generalization. We propose an augmentation method called RoIMix, which characterizes interactions among images. Proposals extracted from different images are mixed together. Previous data augmentation methods operate on a single image while we apply RoIMix to multiple images to create enhanced samples as training data. Experiments show that our proposed method improves the performance of region-based object detectors on both Pascal VOC and URPC datasets. Wei-Hong Lin, Jia-Xing Zhong, Shan Liu 0001, Thomas H. Li, Ge Li 0002 |
ICASSP | 4 |
| 2020 | Pose Refinement: Bridging the Gap Between Unsupervised Learning and Geometric Methods for Visual Odometry
Lanqing Zhang, Ge Li 0002, Thomas H. Li |
ICASSP | 3 |
| 2020 | Towards Loss Balance and Consistent Model in Self-supervised Monocular Depth EstimationabstractRecently, self-supervised methods based on Convolutional Neural Networks (CNN) have achieved remarkable success in monocular depth estimation. To obtain higher quality depth maps, some of these approaches leverage traditional schemes to compute rough depth maps as proxy labels and adopt classic regression loss functions to minimize the differences between network-predicted depth maps and proxy labels. However, at proxy labels with large depth values, these methods suffer from a loss imbalance problem. To address this limitation and further improve the network performance, this article offers three key contributions. Firstly, a novel regression loss function is proposed, which can alleviate the loss imbalance problem and better handle rough proxy labels. Secondly, a dynamic mask is designed to accelerate network convergence. Thirdly, an innovative consistency loss is introduced, which can produce a more accurate and consistent model by maintaining consistency between the produced depth maps of each input image and its mirror. The effectiveness of our contributions is demonstrated by a series of ablation studies. Extensive experiments on KITTI dataset reveal that our approach achieves state-of-the-art results. Lanqing Zhang, Xing Cai, Keyao Li, Ge Li 0002, Thomas H. Li |
ICTAI | 6 |
| 2020 | VONAS: Network Design in Visual Odometry using Neural Architecture SearchabstractThe end-to-end VO (visual odometry) is a complicated task with the property of highly temporal dependency, but the design of its deep networks lacks thorough investigation. Meanwhile, NAS (Neural architecture search) has been widely searched and applied in many computer vision fields due to its advantage in automatic network design. However, most of the existing NAS frameworks only consider single image tasks such as image classification, lacking the consideration of the video (multi-frames) tasks such as VO. Therefore, this paper explores the network design for the VO task and proposes a more general single path based one-shot NAS, named VONAS, which can model sequential information for video-related tasks. Extensive experiments prove that the network architecture is significant for the (un)supervised VO. The models obtained by VONAS are lightweight and achieve SOTA performance with good generalization. Xing Cai, Lanqing Zhang, Ge Li 0002, Thomas H. Li |
ACM Multimedia | 5 |
| 2020 | Vaccine-style-net: Point Cloud Completion in Implicit Continuous Function SpaceabstractThough recent advances in point cloud completion have shown exciting promise with learning-based methods, most of them still generate coarse point clouds with a fixed number of points (e.g. 2048). In this paper, we propose Vaccine-Style-Net, a new point cloud completion method that can produce high resolution 3D shapes with complete smooth surface. Vaccine-Style-Net performs point cloud completion in the function space of 3D surface, which represent the 3D surface as the continuous decision boundary function. Meanwhile, a reinforcement learning agent is embedded to deduce the complete 3D geometry from the incomplete point cloud. In contrast to the existing approaches, the completed 3D shapes produced by our method can be any resolution without excessive memory footprint. Moreover, to increase the diversity and adaptability of the method, we introduce two-type-free-form masks to simulate various corrupted inputs as well as a mask dataset called onion-peeling-mask (OPM). Finally, we discuss the limitations of existing evaluation metrics for shape completion tasks and explore a novel metric to supplement the existing ones. Experiments demonstrate that our method not only achieves competitive results qualitatively and quantitatively but also can produce a continuous 3D shape with any resolution. Ruonan Zhang 0002, Jing Wang 0115, Shan Liu 0001, Thomas H. Li, Ge Li 0002 |
ACM Multimedia | 5 |
| 2020 | Spatial-Temporal Context-Aware Online Action Detection and PredictionabstractSpatial-temporal action detection in videos is a challenging problem that has attracted considerable attention in recent years. Most current approaches address action detection as an object detection problem, which utilizes successful object detection frameworks such as Faster R-CNN to operate action detection at every single frame first, and then generates action tubes by linking bounding boxes across the whole video in an offline fashion. However, unlike object detection in static images, temporal context information is vital for action detection in videos. Therefore, we propose an online action detection model that leverages the spatial-temporal context information existing in videos to perform action inference and localization. More specifically, we try to depict the spatial-temporal context pattern of actions via an encoder-decoder model that is based on a convolutional recurrent neural network. The model accepts a video snippet as input and encodes the dynamic information inside the snippet in the forward pass. During the backward pass, the decoder resolves the information for action detection with the current appearance or motion cue at each time stamp. In addition, we devise an incremental action-tube construction algorithm that enables our model to accomplish action prediction ahead of time and performs action detection in an online fashion. To evaluate the performance of our method, we conduct experiments on three popular public datasets UCF-101, UCF-Sports, and J-HMDB-21. The experimental results demonstrate that our method can achieve competitive or superior performance when compared to the state-of-the-art methods. To encourage further research, we release our project on “https://github.com.hjjpku.OATD.” Jingjia Huang, Nannan Li 0001, Thomas H. Li, Shan Liu 0001, Ge Li 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Deep Spatial Transformation for Pose-Guided Person Image Generation and AnimationabstractPose-guided person image generation and animation aim to transform a source person image to target poses. These tasks require spatial manipulation of source data. However, Convolutional Neural Networks are limited by the lack of ability to spatially transform the inputs. In this paper, we propose a differentiable global-flow local-attention framework to reassemble the inputs at the feature level. This framework first estimates global flow fields between sources and targets. Then, corresponding local source feature patches are sampled with content-aware local attention coefficients. We show that our framework can spatially transform the inputs in an efficient manner. Meanwhile, we further model the temporal consistency for the person image animation task to generate coherent videos. The experiment results of both image generation and animation tasks demonstrate the superiority of our model. Besides, additional results of novel view synthesis and face image animation show that our model is applicable to other tasks requiring spatial transformation. The source code of our project is available at https://github.com/RenYurui/Global-Flow-Local-Attention. Yurui Ren, Ge Li 0002, Shan Liu 0001, Thomas H. Li |
IEEE Trans. Image Process. | 4 |
| 2019 | Graph Convolutional Label Noise Cleaner: Train a Plug-And-Play Action Classifier for Anomaly DetectionabstractVideo anomaly detection under weak labels is formulated as a typical multiple-instance learning problem in previous works. In this paper, we provide a new perspective, i.e., a supervised learning task under noisy labels. In such a viewpoint, as long as cleaning away label noise, we can directly apply fully supervised action classifiers to weakly supervised anomaly detection, and take maximum advantage of these well-developed classifiers. For this purpose, we devise a graph convolutional network to correct noisy labels. Based upon feature similarity and temporal consistency, our network propagates supervisory signals from high-confidence snippets to low-confidence ones. In this manner, the network is capable of providing cleaned supervision for action classifiers. During the test phase, we only need to obtain snippet-wise predictions from the action classifier without any extra post-processing. Extensive experiments on 3 datasets at different scales with 2 types of action classifiers demonstrate the efficacy of our method. Remarkably, we obtain the frame-level AUC score of 82.12% on UCF-Crime. Jia-Xing Zhong, Nannan Li 0001, Weijie Kong, Shan Liu 0001, Thomas H. Li, Ge Li 0002 |
CVPR | 5 |
| 2019 | BLP - Boundary Likelihood Pinpointing Networks for Accurate Temporal Action LocalizationabstractDespite tremendous progress achieved in temporal action detection, state-of-the-art methods still suffer from the sharp performance deterioration when localizing the starting and ending temporal action boundaries. Although most methods apply boundary regression paradigm to tackle this problem, we argue that the direct regression lacks detailed enough information to yield accurate temporal boundaries. In this paper, we propose a novel Boundary Likelihood Pinpointing (BLP) network to alleviate this deficiency of boundary regression and improve the localization accuracy. Given a loosely localized search interval that contains an action instance, BLP casts the problem of localizing temporal boundaries as that of assigning probabilities on each equally divided unit of this interval. These generated probabilities provide useful information regarding the boundary location of the action inside this search interval. Based on these probabilities, we introduce a boundary pinpointing paradigm to pinpoint the accurate boundaries under a simple probabilistic framework. Compared with other C3D feature based detectors, extensive experiments demonstrate that BLP significantly improves the localization performance of recent state-of-the-art detectors, and achieves competitive detection mAP on both THUMOS' 14 and ActivityNet datasets, particularly when the evaluation tIoU is high. Weijie Kong, Nannan Li 0001, Shan Liu 0001, Thomas H. Li, Ge Li 0002 |
ICASSP | 4 |
| 2019 | Boundary Information Matters More: Accurate Temporal Action Detection with Temporal Boundary NetworkabstractTemporal action detection in untrimmed videos is an important yet challenging task. How to locate complex actions accurately is still an open question due to the ambiguous boundaries between action instances and the background. Recently a newly proposed work exploits Structured Segment Networks (SSN) for temporal action detection, which models temporal structure of action instances via structured temporal pyramids, and comprises two classifiers, respectively for classifying actions and determining proposal completeness. In this paper we attempt to delve the temporal boundary information when modeling temporal structure of action instance, by introducing to SSN the structured temporal boundary attention pyramid. On top of the pyramid, we add another set of classifiers for unit-wise completeness evaluation, which enables proposal recycling for efficient action detection. Experimental results on two challenging benchmarks, THUMOS’14 and ActivityNet, indicate that our Temporal Boundary Network shows a significant performance improvement compared with SSN, and achieves a competitive performance compared with state-of-the-arts. Tao Zhang 0069, Shan Liu 0001, Thomas H. Li, Ge Li 0002 |
ICASSP | 3 |
| 2019 | StructureFlow: Image Inpainting via Structure-Aware Appearance FlowabstractImage inpainting techniques have shown significant improvements by using deep neural networks recently. However, most of them may either fail to reconstruct reasonable structures or restore fine-grained textures. In order to solve this problem, in this paper, we propose a two-stage model which splits the inpainting task into two parts: structure reconstruction and texture generation. In the first stage, edge-preserved smooth images are employed to train a structure reconstructor which completes the missing structures of the inputs. In the second stage, based on the reconstructed structures, a texture generator using appearance flow is designed to yield image details. Experiments on multiple publicly available datasets show the superior performance of the proposed network. Yurui Ren, Xiaoming Yu, Ruonan Zhang 0002, Thomas H. Li, Shan Liu 0001, Ge Li 0002 |
ICCV | 4 |
| 2019 | PDNet: Prior-Model Guided Depth-Enhanced Network for Salient Object DetectionabstractFully convolutional neural networks (FCNs) have shown outstanding performance in many computer vision tasks including salient object detection. However, there still remains two issues needed to be addressed in deep learning based saliency detection. One is the lack of tremendous amount of annotated data to train a network. The other is the lack of robustness for extracting salient objects in images containing complex scenes. In this paper, we present a new architecture-PDNet, a robust prior-model guided depth-enhanced network for RGB-D salient object detection. In contrast to existing works, in which RGB-D values of image pixels are fed directly to a network, the proposed architecture is composed of a master network for processing RGB values, and a sub-network making full use of depth cues and incorporate depth-based features into the master network. To overcome the limited size of the labeled RGB-D dataset for training, we employ a large conventional RGB dataset to pre-train the master network, which proves to contribute largely to the final accuracy. Extensive evaluations over five benchmark datasets demonstrate that our proposed method performs favorably against the state-of-the-art approaches. Chunbiao Zhu, Xing Cai, Kan Huang, Thomas H. Li, Ge Li 0002 |
ICME | 4 |
| 2019 | ARMIN: Towards a More Efficient and Light-weight Recurrent Memory NetworkabstractIn recent years, memory-augmented neural networks(MANNs) have shown promising power to enhance the memory ability of neural networks for sequential processing tasks. However, previous MANNs suffer from complex memory addressing mechanism, making them relatively hard to train and causing computational overheads. Moreover, many of them reuse the classical RNN structure such as LSTM for memory processing, causing inefficient exploitations of memory information. In this paper, we introduce a novel MANN, the Auto-addressing and Recurrent Memory Integrating Network (ARMIN) to address these issues. The ARMIN only utilizes hidden state h_t for automatic memory addressing, and uses a novel RNN cell for refined integration of memory information. Empirical results on a variety of experiments demonstrate that the ARMIN is more light-weight and efficient compared to existing memory networks. Moreover, we demonstrate that the ARMIN can achieve much lower computational overhead than vanilla LSTM while keeping similar performances. Codes are available on github.com/zoharli/armin. Zhangheng Li, Jia-Xing Zhong, Jingjia Huang, Tao Zhang 0069, Thomas H. Li, Ge Li 0002 |
IJCAI | 5 |
| 2019 | Multi-mapping Image-to-Image Translation via Learning DisentanglementabstractRecent advances of image-to-image translation focus on learning the one-to-many mapping from two aspects: multi-modal translation and multi-domain translation. However, the existing methods only consider one of the two perspectives, which makes them unable to solve each other's problem. To address this issue, we propose a novel unified model, which bridges these two objectives. First, we disentangle the input images into the latent representations by an encoder-decoder architecture with a conditional adversarial training in the feature space. Then, we encourage the generator to learn multi-mappings by a random cross-domain translation. As a result, we can manipulate different parts of the latent representations to perform multi-modal and multi-domain translations simultaneously. Experiments demonstrate that our method outperforms state-of-the-art methods. Xiaoming Yu, Yuanqi Chen, Shan Liu 0001, Thomas H. Li, Ge Li 0002 |
NeurIPS | 4 |
| 2019 | LECARM: Low-Light Image Enhancement Using the Camera Response ModelabstractLow-light image enhancement algorithms can improve the visual quality of low-light images and support the extraction of valuable information for some computer vision techniques. However, existing techniques inevitably introduce color and lightness distortions when enhancing the images. To lower the distortions, we propose a novel enhancement framework using the response characteristics of cameras. First, we discuss how to determine a reasonable camera response model and its parameters. Then, we use the illumination estimation techniques to estimate the exposure ratio for each pixel. Finally, the selected camera response model is used to adjust each pixel to the desired exposure according to the estimated exposure ratio map. Experiments show that our method can obtain enhancement results with fewer color and lightness distortions compared with the several state-of-the-art methods. Yurui Ren, Zhenqiang Ying, Thomas H. Li, Ge Li 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Exploiting the Value of the Center-dark Channel Prior for Salient Object DetectionabstractSaliency detection aims to detect the most attractive objects in images and is widely used as a foundation for various applications. In this article, we propose a novel salient object detection algorithm for RGB-D images using center-dark channel priors. First, we generate an initial saliency map based on a color saliency map and a depth saliency map of a given RGB-D image. Then, we generate a center-dark channel map based on center saliency and dark channel priors. Finally, we fuse the initial saliency map with the center dark channel map to generate the final saliency map. Extensive evaluations over four benchmark datasets demonstrate that our proposed method performs favorably against most of the state-of-the-art approaches. Besides, we further discuss the application of the proposed algorithm in small target detection and demonstrate the universal value of center-dark channel priors in the field of object detection. Chunbiao Zhu, Thomas H. Li, Shan Liu 0001, Ge Li 0002 |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2018 | SingleGAN: Image-to-Image Translation by a Single-Generator Network Using Multiple Generative Adversarial Learning
Xiaoming Yu, Xing Cai, Zhenqiang Ying, Thomas H. Li, Ge Li 0002 |
ACCV (5) | 4 |
| 2018 | Online Action Tube Detection via Resolving the Spatio-temporal Context PatternabstractAt present, spatio-temporal action detection in the video is still a challenging problem, considering the complexity of the background, the variety of the action or the change of the viewpoint in the unconstrained environment. Most of current approaches solve the problem via a two-step processing: first detecting actions at each frame; then linking them, which neglects the continuity of the action and operates in an offline and batch processing manner. In this paper, we attempt to build an online action detection model that introduces the spatio-temporal coherence existed among action regions when performing action category inference and position localization. Specifically, we seek to represent the spatio-temporal context pattern via establishing an encoder-decoder model based on the convolutional recurrent network. The model accepts a video snippet as input and encodes the dynamic information of the action in the forward pass. During the backward pass, it resolves such information at each time instant for action detection via fusing the current static or motion cue. Additionally, we propose an incremental action tube generation algorithm, which accomplishes action bounding-boxes association, action label determination and the temporal trimming in a single pass. Our model takes in the appearance, motion or fused signals as input and is tested on two prevailing datasets, UCF-Sports and UCF-101. The experiment results demonstrate the effectiveness of our method which achieves a performance superior or comparable to compared existing approaches. Jingjia Huang, Nannan Li 0001, Jia-Xing Zhong, Thomas H. Li, Ge Li 0002 |
ACM Multimedia | 4 |
| 2018 | Step-by-step Erasion, One-by-one Collection: A Weakly Supervised Temporal Action DetectorabstractWeakly supervised temporal action detection is a Herculean task in understanding untrimmed videos, since no supervisory signal except the video-level category label is available on training data. Under the supervision of category labels, weakly supervised detectors are usually built upon classifiers. However, there is an inherent contradiction between classifier and detector; i.e., a classifier in pursuit of high classification performance prefers top-level discriminative video clips that are extremely fragmentary, whereas a detector is obliged to discover the whole action instance without missing any relevant snippet. To reconcile this contradiction, we train a detector by driving a series of classifiers to find new actionness clips progressively, via step-by-step erasion from a complete video. During the test phase, all we need to do is to collect detection results from the one-by-one trained classifiers at various erasing steps. To assist in the collection process, a fully connected conditional random field is established to refine the temporal localization outputs. We evaluate our approach on two prevailing datasets, THUMOS'14 and ActivityNet. The experiments show that our detector advances state-of-the-art weakly supervised temporal action detection results, and even compares with quite a few strongly supervised methods. Jia-Xing Zhong, Nannan Li 0001, Weijie Kong, Tao Zhang 0069, Thomas H. Li, Ge Li 0002 |
ACM Multimedia | 5 |
| 2018 | Deep Pedestrian Detection Using Contextual Information and Multi-level Features
Weijie Kong, Nannan Li 0001, Thomas H. Li, Ge Li 0002 |
MMM (1) | 3 |
| 2018 | Detecting action tubes via spatial action estimation and temporal path inference
Nannan Li 0001, Jingjia Huang, Thomas H. Li, Huiwen Guo, Ge Li 0002 |
Neurocomputing | 3 |