EDBT 2026 Demo / reviewers in the wild / expert
Teng Li 0001
dblp:09/6669-1
· DBLP profile ↗
57ranked-venue papers
16as first author
22since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 35 · 10 first-author · 13 since 2021Artificial intelligence and machine learning · 25 · 4 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-authorComputer networks · 1Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unifying Modality and Scale: Visual Mamba for Feature Fusion in RGB-X Crowd CountingabstractCrowd counting has long been a crucial topic in the domains of computer vision and video surveillance. In particular, with the widespread use of thermal or depth cameras, RGB-X crowd counting has emerged as a prominent research focus. Although depth or infrared images provide complementary information, the core challenge remains in effectively unifying heterogeneous cross-modality and cross-scale information to form a comprehensive representation of crowd distributions in complex scenes. To address this problem, we propose a novel Mamba-based framework, termed UMS-VMamba-CC for multi-modal (RGB-X) crowd counting. Specifically, we design the cross-modality disentanglement fusion visual Mamba (CMDF-VMamba) that uses self-supervised learning to decompose modality-invariant and modality-specific features in spatial-frequency domains, followed by multi-modal feature aggregation via the gating mechanism. For cross-scale fusion, we design the cross-scale pyramid fusion visual Mamba (CSPF-VMamba), which adopts the bi-directional pyramid structure to incorporate low-level features into high-level representations during downsampling, while upsampling high-level features and aggregating them with low-level features through state space contextual modeling. Comprehensive experiments on multiple mainstream datasets demonstrate that the UMS-VMamba-CC framework achieves competitive performance for RGB-X crowd counting. Yaocong Hu, Mengbo Jia, Pindeng Wang, Wenbo Zhu 0002, Huanjie Tao, Tianming Ni, Teng Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 10 |
| 2026 | EMFFTrans: Efficient Multi-Scale Feature Fusion Transformer for Road Scene Semantic Segmentation
Yaocong Hu, Pindeng Wang, Mengbo Jia, Jinwen Hong, Guoyang Wan, Huicheng Yang, Bingyou Liu, Tianming Ni, Teng Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 12 |
| 2025 | An Inverse Cavity Scattering Inversion Method Based on Adaptive Neural Fuzzy Inference System
Teng Li 0001, Aoyu Zhu |
ICIC (22) | 1 |
| 2025 | Dual-Branch Enhancement and Multi-Modal Fusion for Low-Light Visible Polarization Image Object Detection in Dense Smog EnvironmentsabstractABSTRACT In scenarios with heavy smog, the accuracy of object detection in low‐light visible polarization images significantly decreases. To address this issue, we propose a dual‐branch enhancement and multi‐modal fusion network for object detection in low‐light visible polarization images in dense smog environments. Specifically, the network consists of an image enhancement stage and an object detection stage. In the image enhancement stage, a dual‐branch enhancement structure comprising greyscale feature map prediction and atmospheric light transmission network is proposed to remove noise from the images and enhance texture information, jointly generating enhanced visible polarization images. In the object detection stage, feature maps of the enhanced visible polarization images and the degree of visible polarization images are fused, and their fused texture‐enhanced feature maps are fed into the detection module for object detection. Additionally, we have collected a dataset of low‐light visible polarization images under real smog conditions. Extensive experiments demonstrate that our method can generate visually improved enhanced images and significantly increase detection accuracy and the number of detected objects in low‐light and dense smog environments. Fudong Nian, Jianguo Huang, Teng Li 0001 |
IET Image Process. | 5 |
| 2025 | Local and global self-attention enhanced graph convolutional network for skeleton-based action recognition
Zhize Wu, Long Wan, Teng Li 0001, Fudong Nian |
Pattern Recognit. | 4 |
| 2024 | Swelling-ViT: Rethink Data-Efficient Vision Transformer from Locality
Chuanrui Hu, Fudong Nian, Teng Li 0001 |
PRCV (4) | 6 |
| 2024 | NLDF: Neural Light Dynamic Fields for 3D Talking Head Generation
Niu Guanchen, Songsong Cheng, Teng Li 0001 |
PRICAI (1) | 3 |
| 2024 | TISE-LSTM: A LSTM model for precipitation nowcasting with temporal interactions and spatial extract blocksabstractPrecipitation nowcasting has a profound impact on humanity and society, especially in areas with heavy rainfall, playing a central role in alerting against rainstorm disasters. At present, numerous deep learning-based methods have been proposed and proven superior to traditional radar echo extrapolation techniques like the Recurrent Neural Networks (RNNs). Our study introduces a novel precipitation forecasting model named TISE-LSTM, which can use the real image of the past half hour to predict the radar echo image of the next hour. By integrating the TIB and SEB into TISL-LSTM, we can alleviate the inherent issue observed in existing models, which is the decrease in forecast accuracy for high radar echo regions with extended extrapolation time. On both the CIKM23017 and AHEM real radar echo datasets, the performance of TISL-LSTM (including POD, FAR, CSI and HSS) is improved compared to the second-ranked model by 7.53%, 2.32%, 8.93%, 2.6%, and 14.84%, 6.18%, 8.92%, 18.12%, respectively, when the precipitation threshold is set to 40. Moreover, our model obtained an optimal MAE, MSE, SSIM score. Predicted images and the graphs of each metric over extrapolation time both demonstrate that our model accurately forecasts regions with high radar echoes even with a one-hour extrapolation. Changyong Zheng, Yifan Tao, Lina Xun, Teng Li 0001 |
Neurocomputing | 5 |
| 2024 | Diffusion Models and Pseudo-Change: A Transfer Learning-Based Change Detection in Remote Sensing ImagesabstractRemote sensing (RS) image change detection (CD) has been a research hotspot in recent years, which plays an important role in urban planning and disaster assessment. However, since CD labels are difficult to obtain, how to utilize semantic information in RS images to improve the change prediction performance is a problem worth exploring. To solve this problem, we propose a transfer learning-based CD method that utilizes a diffusion generation model to translate high-level semantic information into low-level change information. First, we propose a pseudo-change image pair generation method that utilizes semantic labels to guide the diffusion model to generate change images. Then, the refined loss (RL) is designed to improve the model’s ability to recognize change features based on the difference between pseudo-change image pairs and unlabeled image pairs. Experimental results on WHU-CD, LEVIR-CD, and GoogleGZ-CD datasets show that the proposed method effectively transfers the semantic information into change information and finally improves the model’s feature recognition ability for change objects. Compared with recent CD and transfer learning methods, the proposed transfer learning model (TLM) achieves the best performance. The source code is available athttps://github.com/VCISwang/STCD. Jia-Xin Wang, Teng Li 0001, Sibao Chen 0001, Cheng-Jie Gu, Zhi-Hui You, Bin Luo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | AudioEar: Single-View Ear Reconstruction for Personalized Spatial AudioabstractSpatial audio, which focuses on immersive 3D sound rendering, is widely applied in the acoustic industry. One of the key problems of current spatial audio rendering methods is the lack of personalization based on different anatomies of individuals, which is essential to produce accurate sound source positions. In this work, we address this problem from an interdisciplinary perspective. The rendering of spatial audio is strongly correlated with the 3D shape of human bodies, particularly ears. To this end, we propose to achieve personalized spatial audio by reconstructing 3D human ears with single-view images. First, to benchmark the ear reconstruction task, we introduce AudioEar3D, a high-quality 3D ear dataset consisting of 112 point cloud ear scans with RGB images. To self-supervisedly train a reconstruction model, we further collect a 2D ear dataset composed of 2,000 images, each one with manual annotation of occlusion and 55 landmarks, named AudioEar2D. To our knowledge, both datasets have the largest scale and best quality of their kinds for public use. Further, we propose AudioEarM, a reconstruction method guided by a depth estimation network that is trained on synthetic data, with two loss functions tailored for ear data. Lastly, to fill the gap between the vision and acoustics community, we develop a pipeline to integrate the reconstructed ear mesh with an off-the-shelf 3D human body and simulate a personalized Head-Related Transfer Function (HRTF), which is the core of spatial audio rendering. Code and data are publicly available in https://github.com/seanywang0408/AudioEar. Bingbing Ni, Wenjun Zhang 0001, Jinxian Liu, Teng Li 0001 |
AAAI | 7 |
| 2023 | Boosting Point Clouds Rendering via Radiance MappingabstractRecent years we have witnessed rapid development in NeRF-based image rendering due to its high quality. However, point clouds rendering is somehow less explored. Compared to NeRF-based rendering which suffers from dense spatial sampling, point clouds rendering is naturally less computation intensive, which enables its deployment in mobile computing device. In this work, we focus on boosting the image quality of point clouds rendering with a compact model design. We first analyze the adaption of the volume rendering formulation on point clouds. Based on the analysis, we simplify the NeRF representation to a spatial mapping function which only requires single evaluation per pixel. Further, motivated by ray marching, we rectify the the noisy raw point clouds to the estimated intersection between rays and surfaces as queried coordinates, which could avoid spatial frequency collapse and neighbor point disturbance. Composed of rasterization, spatial mapping and the refinement stages, our method achieves the state-of-the-art performance on point clouds rendering, outperforming prior works by notable margins, with a smaller model size. We obtain a PSNR of 31.74 on NeRF-Synthetic, 25.88 on ScanNet and 30.81 on DTU. Code and data are publicly available in https://github.com/seanywang0408/RadianceMapping. Bingbing Ni, Teng Li 0001, Kai Chen 0006, Wenjun Zhang 0001 |
AAAI | 4 |
| 2023 | Frequency-Modulated Point Cloud Rendering with Easy EditingabstractWe develop an effective point cloud rendering pipeline for novel view synthesis, which enables high fidelity local detail reconstruction, real-time rendering and user-friendly editing. In the heart of our pipeline is an adaptive frequency modulation module called Adaptive Frequency Net (AFNet), which utilizes a hypernetwork to learn the local texture frequency encoding that is consecutively injected into adaptive frequency activation layers to modulate the implicit radiance signal. This mechanism improves the frequency expressive ability of the network with richer frequency basis support, only at a small computational budget. To further boost performance, a preprocessing module is also proposed for point cloud geometry optimization via point opacity estimation. In contrast to implicit rendering, our pipeline supports high-fidelity interactive editing based on point cloud manipulation. Extensive experimental results on NeRF-Synthetic, ScanNet, DTU and Tanks and Temples datasets demonstrate the superior performances achieved by our method in terms of PSNR, SSIM and LPIPS, in comparison to the state-of-the-art. Code is released at https://github.com/yizhangphd/FreqPCR. Bingbing Ni, Wenjun Zhang 0001, Teng Li 0001 |
CVPR | 5 |
| 2023 | Learning Shape Primitives via Implicit Convexity RegularizationabstractShape primitives decomposition has been an important and long-standing task in 3D shape analysis. Prior arts heavily rely on 3D point clouds or voxel data for shape primitives extraction, which are less practical in real-world scenarios. This paper proposes to learn shape primitives from multi-view images by introducing implicit surface rendering. It is challenging since implicit shapes have a high degree of freedom, which violates the simplicity property of shape primitives. In this work, a novel regularization term named Implicit Convexity Regularization (ICR) imposed on implicit primitive learning is proposed to tackle this problem. We start with the convexity definition of general 3D shapes, and then derive the equivalent expression for implicit shapes represented by signed distance functions (SDFs). Further, instead of directly constraining the output SDF values which cause unstable optimization, we alternatively impose constraint on second order directional derivatives on line segments inside the shapes, which proves to be a tighter condition for 3D convexity. Implicit primitives constrained by the proposed ICR are combined into a whole object via softmax-weighted-sum operation over all primitive SDFs. Experiments on synthetic and real-world datasets show that our method is able to decompose objects into simple and reasonable shape primitives without the need of segmentation labels or 3D data. Code and data is publicly available in https://github.com/seanywang0408/ICR. Kai Chen 0006, Teng Li 0001, Wenjun Zhang 0001, Bingbing Ni |
ICCV | 4 |
| 2023 | Carotid Lumen Diameter and Intima-Media Thickness Measurement via Boundary-Guided Pseudo-LabelingabstractThe carotid lumen diameter (CALD) and intima-media thickness (CIMT) are essential indicators for diagnosing and treating cardiovascular disease. Existing supervised methods rely on extensive labeled data, which is time-consuming and labor-intensive. This letter presents a boundary-guided pseudo-labeling (BGPL) method, which adopts prior anatomical knowledge to generate and select reliable pseudo-labels for unlabeled data to boost measurement performance. To improve the quality of pseudo-labels, we propose a fine self-attention (FSA) module and a boundary attention module (BAM) to force the feature extractor to highlight boundary information. The FSA maintains internal resolution and enhances non-linearity. The BAM injects boundary heatmaps from the pre-trained boundary regressor into the feature extractor. We specifically provide an alternate training strategy to increase the quality of pseudo-labels further. We evaluate the proposed method on challenging carotid ultrasound datasets. Experiments show that the proposed method outperforms the existing state-of-the-art algorithm. Shimeng Yang, Teng Li 0001, Yinping Lv, Shuo Li 0001 |
IEEE Signal Process. Lett. | 2 |
| 2022 | Bi-volution: A Static and Dynamic Coupled FilterabstractDynamic convolution has achieved significant gain in performance and computational complexity, thanks to its powerful representation capability given limited filter number/layers. However, SOTA dynamic convolution operators are sensitive to input noises (e.g., Gaussian noise, shot noise, e.t.c.) and lack sufficient spatial contextual information in filter generation. To alleviate this inherent weakness, we propose a lightweight and heterogeneous-structure (i.e., static and dynamic) operator, named Bi-volution. On the one hand, Bi-volution is designed as a dual-branch structure to fully leverage complementary properties of static/dynamic convolution, which endows Bi-volution more robust properties and higher performance. On the other hand, the Spatial Augmented Kernel Generation module is proposed to improve the dynamic convolution, realizing the learning of spatial context information with negligible additional computational complexity. Extensive experiments illustrate that the ResNet-50 equipped with Bi-volution achieves a highly competitive boost in performance (+2.8% top-1 accuracy on ImageNet classification, +2.4% box AP and +2.2% mask AP on COCO detection and instance segmentation) while maintaining extremely low FLOPs (i.e., [email protected] GFLOPs). Furthermore, our Bi-volution shows better robustness than dynamic convolution against various noise and input corruptions. Our code is available at https://github.com/neuralchen/Bivolution. Xiwei Hu, Xuanhong Chen, Bingbing Ni, Teng Li 0001, Yutian Liu 0004 |
AAAI | 4 |
| 2022 | Contrastive Regression for Domain Adaptation on Gaze EstimationabstractAppearance-based Gaze Estimation leverages deep neural networks to regress the gaze direction from monocular images and achieve impressive performance. However, its success depends on expensive and cumbersome annotation capture. When lacking precise annotation, the large domain gap hinders the performance of trained models on new domains. In this paper, we propose a novel gaze adaptation approach, namely Contrastive Regression Gaze Adaptation (CRGA), for generalizing gaze estimation on the target domain in an unsupervised manner. CRGA leverages the Contrastive Domain Generalization (CDG) module to learn the stable representation from the source domain and leverages the Contrastive Self-training Adaptation (CSA) module to learn from the pseudo labels on the target domain. The core of both CDG and CSA is the Contrastive Regression (CR) loss, a novel contrastive loss for regression by pulling features with closer gaze directions closer together while pushing features with farther gaze directions farther apart. Experimentally, we choose ETH-XGAZE and Gaze-360 as the source domain and test the domain generalization and adaptation performance on MPIIGAZE, RT-GENE, Gaze-Capture, EyeDiap respectively. The results demonstrate that our CRGA achieves remarkable performance improvement compared with the baseline models and also outperforms the state-of-the-art domain adaptation approaches on gaze adaptation tasks. Yangzhou Jiang, Jin Li 0057, Bingbing Ni, Wenrui Dai, Hongkai Xiong, Teng Li 0001 |
CVPR | 8 |
| 2022 | Representation-Agnostic Shape Fields
Jiancheng Yang, Linguo Li, Teng Li 0001, Bingbing Ni, Wenjun Zhang 0001 |
ICLR | 6 |
| 2022 | Structure injected weight normalization for training deep networks
Xu Yuan 0008, Xiangjun Shen, Sumet Mehta, Teng Li 0001, Shiming Ge, Zhengjun Zha |
Multim. Syst. | 4 |
| 2022 | Visible light polarization image desmogging via Cycle Convolutional Neural Network
Teng Li 0001, Yuzhou Zeng, Fudong Nian |
Multim. Syst. | 3 |
| 2022 | Reliable Contrastive Learning for Semi-Supervised Change Detection in Remote Sensing ImagesabstractWith the development of deep learning in remote sensing (RS) image change detection (CD), the dependence of CD models on labeled data has become an important problem. To make better use of the comparatively resource-saving unlabeled data, the CD method based on semi-supervised learning (SSL) is worth further study. This article proposes a reliable contrastive learning (RCL) method for semi-supervised RS image CD. First, according to the task characteristics of CD, we design the contrastive loss based on the changed areas to enhance the model’s feature extraction ability for changed objects. Then, to improve the quality of pseudo labels in SSL, we use the uncertainty of unlabeled data to select reliable pseudo labels for model training. Combining these methods, semi-supervised CD models can make full use of unlabeled data. Extensive experiments on three widely used CD datasets demonstrate the effectiveness of the proposed method. The results show that our semi-supervised approach has a better performance than related methods. The code is available athttps://github.com/VCISwang/RC-Change-Detection. Jia-Xin Wang, Teng Li 0001, Sibao Chen 0001, Jin Tang 0001, Bin Luo 0001, Richard C. Wilson 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2021 | Shape Self-Correction for Unsupervised Point Cloud UnderstandingabstractWe develop a novel self-supervised learning method named Shape Self-Correction for point cloud analysis. Our method is motivated by the principle that a good shape representation should be able to find distorted parts of a shape and correct them. To learn strong shape representations in an unsupervised manner, we first design a shape-disorganizing module to destroy certain local shape parts of an object. Then the destroyed shape and the normal shape are sent into a point cloud network to get representations, which are employed to segment points that belong to distorted parts and further reconstruct them to restore the shape to normal. To perform better in these two associated pretext tasks, the network is constrained to capture useful shape features from the object, which indicates that the point cloud network encodes rich geometric and contextual information. The learned feature extractor transfers well to downstream classification and segmentation tasks. Experimental results on ModelNet, ScanNet and ShapeNetPart demonstrate that our method achieves state-of-the-art performance among unsupervised methods. Our framework can be applied to a wide range of deep learning networks for point cloud analysis and we show experimentally that pre-training with our framework significantly boosts the performance of supervised models. Ye Chen 0006, Jinxian Liu, Bingbing Ni, Jiancheng Yang, Teng Li 0001, Qi Tian 0001 |
ICCV | 7 |
| 2021 | Distributed Policy Evaluation with Fractional Order Dynamics in Multiagent Reinforcement LearningabstractThe main objective of multiagent reinforcement learning is to achieve a global optimal policy. It is difficult to evaluate the value function with high-dimensional state space. Therefore, we transfer the problem of multiagent reinforcement learning into a distributed optimization problem with constraint terms. In this problem, all agents share the space of states and actions, but each agent only obtains its own local reward. Then, we propose a distributed optimization with fractional order dynamics to solve this problem. Moreover, we prove the convergence of the proposed algorithm and illustrate its effectiveness with a numerical example. Wei Wang 0176, Zhongtian Mao, Ruwen Jiang, Fudong Nian, Teng Li 0001 |
Secur. Commun. Networks | 6 |
| 2020 | Adversarial Domain Adaptation with Domain MixupabstractRecent works on domain adaptation reveal the effectiveness of adversarial learning on filling the discrepancy between source and target domains. However, two common limitations exist in current adversarial-learning-based methods. First, samples from two domains alone are not sufficient to ensure domain-invariance at most part of latent space. Second, the domain discriminator involved in these methods can only judge real or fake with the guidance of hard label, while it is more reasonable to use soft scores to evaluate the generated images or features, i.e., to fully utilize the inter-domain information. In this paper, we present adversarial domain adaptation with domain mixup (DM-ADA), which guarantees domain-invariance in a more continuous latent space and guides the domain discriminator in judging samples' difference relative to source and target domains. Domain mixup is jointly conducted on pixel and feature level to improve the robustness of models. Extensive experiments prove that the proposed approach can achieve superior performance on tasks with various degrees of domain shift and data complexity. Jian Zhang 0079, Bingbing Ni, Teng Li 0001, Chengjie Wang 0001, Qi Tian 0001, Wenjun Zhang 0001 |
AAAI | 4 |
| 2020 | Relative coordinates constraint for face alignment
Fudong Nian, Teng Li 0001, Bing-Kun Bao, Changsheng Xu |
Neurocomputing | 2 |
| 2020 | Arbitrary perspective crowd counting via local to global algorithm
Chuanrui Hu, Yixiang Xie, Teng Li 0001 |
Multim. Tools Appl. | 4 |
| 2020 | A Novel Adaptive Directional Interpolation Algorithm for Digital Video Resolution EnhancementabstractIn this paper, a novel digital video resolution enhancement algorithm based on adaptive directional interpolation is proposed, where the directionality of the edge structure and the nonlocal self-similarity prior within the current frame as well as its adjacent frames are both considered. First, we establish the regularization equation that conforms to the prior model of a video frame and then take the classic bicubic interpolation result as the initial estimation to iteratively solve the restoration equation, in which the edge structures and contours in low resolution (LR) input are reconstructed to estimate and refine the desired high resolution (HR) output. Experimental results show that the proposed algorithm can effectively enhance the clarity of a video frame, with satisfying subjective visual quality and PSNR value. Dong Sun 0003, Qingqing Xie, Teng Li 0001, Yixiang Lu, De Zhu, Qingwei Gao |
Wirel. Commun. Mob. Comput. | 3 |
| 2018 | Learning Semantic-Aligned Action RepresentationabstractA fundamental bottleneck for achieving highly discriminative action representation is that local motion/appearance features are usually not semantic aligned. Namely, a local feature, such as a motion vector or motion trajectory, does not possess any attribute that indicates which moving body part or operated object it is associated with. This mostly leads to global feature pooling/representation learning methods that are often too coarse. Inspired by the recent success of end-to-end (pixel-to-pixel) deep convolutional neural networks (DCNNs), in this paper, we first propose a DCNN architecture, which maps a human centric image region onto human body part response maps. Based on these response maps, we propose a second DCNN, which achieves semantic-aligned feature representation learning. Prior knowledge that only a few parts are responsible for a certain action is also utilized by introducing a group (part) sparseness prior during feature learning. The learned semantic-aligned feature not only boosts the discriminative capability of action representation, but also possesses the good nature of robustness to pose variations and occlusions. Finally, an iterative mining method is employed for learning discriminative action primitive detectors. Extensive experiments on action recognition benchmarks demonstrate a superior recognition performance of the proposed framework. Bingbing Ni, Teng Li 0001, Xiaokang Yang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2017 | Multi-Modal Knowledge Representation Learning via Webly-Supervised Relationships MiningabstractKnowledge representation learning (KRL) encodes enormous structured information with entities and relations into a continuous low-dimensional semantic space. Most conventional methods solely focus on learning knowledge representation from single modality, yet neglect the complementary information from others. The more and more rich available multi-modal data on Internet also drive us to explore a novel approach for KRL in multi-modal way, and overcome the limitations of previous single-modal based methods. This paper proposes a novel multi-modal knowledge representation learning (MM-KRL) framework which attempts to handle knowledge from both textual and visual modal web data. It consists of two stages, i.e., webly-supervised multi-modal relationship mining, and bi-enhanced cross-modal knowledge representation learning. Compared with existing knowledge representation methods, our framework has several advantages: (1) It can effectively mine multi-modal knowledge with structured textual and visual relationships from web automatically. (2) It is able to learn a common knowledge space which is independent to both task and modality by the proposed Bi-enhanced Cross-modal Deep Neural Network (BC-DNN). (3) It has the ability to represent unseen multi-modal relationships by transferring the learned knowledge with isolated seen entities and relations into unseen relationships. We build a large-scale multi-modal relationship dataset (MMR-D) and the experimental results show that our framework achieves excellent performance in zero-shot multi-modal retrieval and visual relationship recognition. Fudong Nian, Bing-Kun Bao, Teng Li 0001, Changsheng Xu |
ACM Multimedia | 3 |
| 2017 | Learning explicit video attributes from mid-level representation for video captioning
Fudong Nian, Teng Li 0001, Yan Wang 0059, Xinyu Wu 0001, Bingbing Ni, Changsheng Xu |
Comput. Vis. Image Underst. | 2 |
| 2017 | Contextual aerial image categorization using codebook
Yan Wang 0059, Xinyu Wu 0001, Yating Yin, Teng Li 0001 |
J. Vis. Commun. Image Represent. | 5 |
| 2017 | Vehicle counting in crowded scenes with multi-channel and multi-task convolutional neural networks
Maojin Sun, Yan Wang 0059, Teng Li 0001, Jing Lv |
J. Vis. Commun. Image Represent. | 3 |
| 2017 | Robust face anti-spoofing with depth information
Yan Wang 0059, Fudong Nian, Teng Li 0001, Kongqiao Wang |
J. Vis. Commun. Image Represent. | 3 |
| 2016 | Peak-Piloted Deep Network for Facial Expression Recognition
Xiangyun Zhao, Xiaodan Liang, Luoqi Liu, Teng Li 0001, Yugang Han, Nuno Vasconcelos, Shuicheng Yan |
ECCV (2) | 4 |
| 2016 | Multiple Granularity Modeling: A Coarse-to-Fine Framework for Fine-grained Action Analysis
Bingbing Ni, Vignesh R. Paramathayalan, Teng Li 0001, Pierre Moulin |
Int. J. Comput. Vis. | 3 |
| 2016 | Dimensionality reduction of data sequences for human activity recognition
Yen-Lun Chen, Xinyu Wu 0001, Teng Li 0001, Jun Cheng 0002, Yongsheng Ou, Mingliang Xu 0001 |
Neurocomputing | 3 |
| 2016 | On random hyper-class random forest for visual classification
Teng Li 0001, Bingbing Ni, Xinyu Wu 0001, Qingwei Gao, Qianmu Li, Dong Sun 0003 |
Neurocomputing | 1 |
| 2016 | Pornographic image detection utilizing deep convolutional neural networks
Fudong Nian, Teng Li 0001, Yan Wang 0059, Mingliang Xu 0001 |
Neurocomputing | 2 |
| 2016 | Robust geometric ℓp-norm feature pooling for image classification and action recognition
Teng Li 0001, Bingbing Ni, Jianbing Shen, Meng Wang 0001 |
Image Vis. Comput. | 1 |
| 2016 | Dense crowd counting from still images with convolutional neural networks
Yaocong Hu, Fudong Nian, Yan Wang 0059, Teng Li 0001 |
J. Vis. Commun. Image Represent. | 5 |
| 2016 | Efficient video copy detection using multi-modality and dynamic path search
Teng Li 0001, Fudong Nian, Xinyu Wu 0001, Qingwei Gao, Yixiang Lu |
Multim. Syst. | 1 |
| 2016 | Efficient near-duplicate image detection with a local-based binary representation
Fudong Nian, Teng Li 0001, Xinyu Wu 0001, Qingwei Gao, Feifeng Li |
Multim. Tools Appl. | 2 |
| 2016 | Multitask Low-Rank Affinity Graph for Image Segmentation and Image AnnotationabstractThis article investigates a low-rank representation--based graph, which can used in graph-based vision tasks including image segmentation and image annotation. It naturally fuses multiple types of image features in a framework named multitask low-rank affinity pursuit. Given the image patches described with multiple types of features, we aim at inferring a unified affinity matrix that implicitly encodes the relations among these patches. This is achieved by seeking the sparsity-consistent low-rank affinities from the joint decompositions of multiple feature matrices into pairs of sparse and low-rank matrices, the latter of which is expressed as the production of the image feature matrix and its corresponding image affinity matrix. The inference process is formulated as a minimization problem and solved efficiently with the augmented Lagrange multiplier method. Considering image patches as vertices, a graph can be built based on the resulted affinity matrix. Compared to previous methods, which are usually based on a single type of feature, the proposed method seamlessly integrates multiple types of features to jointly produce the affinity matrix in a single inference step. The proposed method is applied to graph-based image segmentation and graph-based image annotation. Experiments on benchmark datasets well validate the superiority of using multiple features over single feature and also the superiority of our method over conventional methods for feature fusion. Teng Li 0001, Bin Cheng 0001, Bingbing Ni, Guangcan Liu, Shuicheng Yan |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2015 | Weakly-supervised scene parsing with multiple contextual cues
Teng Li 0001, Xinyu Wu 0001, Bingbing Ni, Ke Lu 0002, Shuicheng Yan |
Inf. Sci. | 1 |
| 2015 | Crowded Scene Analysis: A SurveyabstractAutomated scene analysis has been a topic of great interest in computer vision and cognitive science. Recently, with the growth of crowd phenomena in the real world, crowded scene analysis has attracted much attention. However, the visual occlusions and ambiguities in crowded scenes, as well as the complex behaviors and scene semantics, make the analysis a challenging task. In the past few years, an increasing number of works on the crowded scene analysis have been reported, which covered different aspects including crowd motion pattern learning, crowd behavior and activity analyses, and anomaly detection in crowds. This paper surveys the state-of-the-art techniques on this topic. We first provide the background knowledge and the available features related to crowded scenes. Then, existing models, popular algorithms, evaluation protocols, and system performance are provided corresponding to different aspects of the crowded scene analysis. We also outline the available datasets for performance evaluation. Finally, some research problems and promising future directions are presented with discussions. Teng Li 0001, Meng Wang 0001, Bingbing Ni, Richang Hong, Shuicheng Yan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2015 | Joint Local and Global Consistency on Interdocument and Interword Relationships for Co-ClusteringabstractCo-clustering has recently received a lot of attention due to its effectiveness in simultaneously partitioning words and documents by exploiting the relationships between them. However, most of the existing co-clustering methods neglect or only partially reveal the interword and interdocument relationships. To fully utilize those relationships, the local and global consistencies on both word and document spaces need to be considered, respectively. Local consistency indicates that the label of a word/document can be predicted from its neighbors, while global consistency enforces a smoothness constraint on words/documents labels over the whole data manifold. In this paper, we propose a novel co-clustering method, called co-clustering via local and global consistency, to not only make use of the relationship between word and document, but also jointly explore the local and global consistency on both word and document spaces, respectively. The proposed method has the following characteristics: 1) the word-document relationships is modeled by following information-theoretic co-clustering (ITCC); 2) the local consistency on both interword and interdocument relationships is revealed by a local predictor; and 3) the global consistency on both interword and interdocument relationships is explored by a global smoothness regularization. All the fitting errors from these three-folds are finally integrated together to formulate an objective function, which is iteratively optimized by a convergence provable updating procedure. The extensive experiments on two benchmark document datasets validate the effectiveness of the proposed co-clustering method. Bing-Kun Bao, Weiqing Min, Teng Li 0001, Changsheng Xu |
IEEE Trans. Cybern. | 3 |
| 2015 | Data-Driven Affective Filtering for Images and VideosabstractIn this paper, a novel system is developed for synthesizing user-specified emotions onto arbitrary input images or videos. Other than defining the visual affective model based on empirical knowledge, a data-driven learning framework is proposed to extract the emotion-related knowledge from a set of emotion-annotated images. In a divide-and-conquer manner, the images are clustered into several emotion-specific scene subgroups for model learning. The visual affection is modeled with Gaussian mixture models based on color features of local image patches. For the purpose of affective filtering, the feature distribution of the target is aligned to the statistical model constructed from the emotion-specific scene subgroup, through a piecewise linear transformation. The transformation is derived through a learning algorithm, which is developed with the incorporation of a regularization term enforcing spatial smoothness, edge preservation, and temporal smoothness for the derived image or video transformation. Optimization of the objective function is sought via standard nonlinear method. Intensive experimental results and user studies demonstrate that the proposed affective filtering framework can yield effective and natural effects for images and videos. Teng Li 0001, Bingbing Ni, Mengdi Xu, Meng Wang 0001, Qingwei Gao, Shuicheng Yan |
IEEE Trans. Cybern. | 1 |
| 2014 | Beta Process Multiple Kernel LearningabstractIn kernel based learning, the kernel trick transforms the original representation of a feature instance into a vector of similarities with the training feature instances, known as kernel representation. However, feature instances are sometimes ambiguous and the kernel representation calculated based on them do not possess any discriminative information, which can eventually harm the trained classifier. To address this issue, we propose to automatically select good feature instances when calculating the kernel representation in multiple kernel learning. Specifically, for the kernel representation calculated for each input feature instance, we multiply it element-wise with a latent binary vector named as instance selection variables, which targets at selecting good instances and attenuate the effect of ambiguous ones in the resulting new kernel representation. Beta process is employed for generating the prior distribution for the latent instance selection variables. We then propose a Bayesian graphical model which integrates both MKL learning and inference for the distribution of the latent instance selection variables. Variational inference is derived for model learning under a max-margin principle. Our method is called Beta process multiple kernel learning. Extensive experiments demonstrate the effectiveness of our method on instance selection and its high discriminative capability for various classification problems in vision. Bingbing Ni, Teng Li 0001, Pierre Moulin |
CVPR | 2 |
| 2014 | A novel image denoising algorithm using linear Bayesian MAP estimation based on sparse representation
Dong Sun 0003, Qingwei Gao, Yixiang Lu, Zhixiang Huang, Teng Li 0001 |
Signal Process. | 5 |
| 2012 | Hidden-Concept Driven Multilabel Image Annotation and Label RankingabstractConventional semisupervised image annotation algorithms usually propagate labels predominantly via holistic similarities over image representations and do not fully consider the label locality, inter-label similarity, and intra-label diversity among multilabel images. Taking these problems into consideration, we present the hidden-concept driven image annotation and label ranking algorithm (HDIALR), which conducts label propagation based on the similarity over a visually semantically consistent hidden-concepts space. The proposed method has the following characteristics: 1) each holistic image representation is implicitly decomposed into label representations to reveal label locality: the decomposition is guided by the so-called hidden concepts, characterizing image regions and reconstructing both visual and nonvisual labels of the entire image; 2) each label is represented by a linear combination of hidden concepts, while the similar linear coefficients reveal the inter-label similarity; 3) each hidden concept is expressed as a respective subspace, and different expressions of the same label over the subspace then induce the intra-label diversity; and 4) the sparse coding-based graph is proposed to enforce the collective consistency between image labels and image representations, such that it naturally avoids the dilemma of possible inconsistency between the pairwise label similarity and image representation similarity in multilabel scenario. These properties are finally embedded in a regularized nonnegative data factorization formulation, which decomposes images representations into label representations over both labeled and unlabeled data for label propagation and ranking. The objective function is iteratively optimized by a convergence provable updating procedure. Extensive experiments on three benchmark image datasets well validate the effectiveness of our proposed solution to semisupervised multilabel image annotation and label ranking problem. Bing-Kun Bao, Teng Li 0001, Shuicheng Yan |
IEEE Trans. Multim. | 2 |
| 2011 | Contextual Bag-of-Words for Visual CategorizationabstractBag-of-words (BOW), which represents an image by the histogram of local patches on the basis of a visual vocabulary, has attracted intensive attention in visual categorization due to its good performance and flexibility. Conventional BOW neglects the contextual relations between local patches due to its Naïve Bayesian assumption. However, it is well known that contextual relations play an important role for human beings to recognize visual categories from their local appearance. This paper proposes a novel contextual bag-of-words (CBOW) representation to model two kinds of typical contextual relations between local patches, i.e., a semantic conceptual relation and a spatial neighboring relation. To model the semantic conceptual relation, visual words are grouped on multiple semantic levels according to the similarity of class distribution induced by them, accordingly local patches are encoded and images are represented. To explore the spatial neighboring relation, an automatic term extraction technique is adopted to measure the confidence that neighboring visual words are relevant. Word groups with high relevance are used and their statistics are incorporated into the BOW representation. Classification is taken using the support vector machine with an efficient kernel to incorporate the relational information. The proposed approach is extensively evaluated on two kinds of visual categorization tasks, i.e., video event and scene categorization. Experimental results demonstrate the importance of contextual relations of local patches and the CBOW shows superior performance to conventional BOW. Teng Li 0001, Tao Mei 0001, In-So Kweon, Xian-Sheng Hua 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2011 | Image Decomposition With Multilabel Context: Algorithms and ApplicationsabstractMost research on image decomposition, e.g., image segmentation and image parsing, has predominantly focused on the low-level visual clues within a single image and neglected the contextual information across images. In this paper, we present a new perspective to image decomposition piloted by the multilabel context associated with each individual image. Observing that the contextual information (i.e., local label representations of the same label are similar while those from different labels are dissimilar) exists across images, we propose to perform image decomposition in a collective way and obtain an optimal representation for each label from a set of multilabeled images. We formulate the problem as an optimization problem which maximizes inter-label difference while minimizing the intra-label difference of the target label representations and propose two ways to solve this problem. Such a contextual image decomposition has a wide variety of applications, among which two exemplary ones-multilabel image annotation and label ranking, are presented and evaluated with different classification techniques. Extensive experiments on two benchmark datasets demonstrate promising results. Teng Li 0001, Shuicheng Yan, Tao Mei 0001, Xian-Sheng Hua 0001, In-So Kweon |
IEEE Trans. Image Process. | 1 |
| 2009 | Contextual decomposition of multi-label imagesabstractMost research on image decomposition, e.g. image segmentation and image parsing, has predominantly focused on the low-level visual clues within single image and neglected the contextual information across different images. In this paper, we present a new perspective to image decomposition piloted by the multi-labels associated with individual images. Observing that the context information (i.e., local label representations of the same label are similar while those from different labels are dissimilar) exists across different images, we propose to perform image decomposition in a collective way, and then the image decomposition problem is formulated as an optimization which maximizes inter-label difference and at the same time minimizes intra-label difference of the target label representations. Such contextual image decomposition has a wide variety of applications, among which the two exemplary ones are: 1) multi-label image annotation in which the sparse coding of a query image over the bases consisting of all learned label representations naturally produces the multi-label annotation, and 2) label ranking in which the annotated labels are re-ordered according to the sparse coding coefficients on those learned label representations. It is worth noting that these two applications can be performed simultaneously via the label propagation process in sparse coding. Teng Li 0001, Tao Mei 0001, Shuicheng Yan, In-So Kweon, Chil-Woo Lee |
CVPR | 1 |
| 2009 | Measuring conceptual relation of visual words for visual categorizationabstractRepresenting image using the distribution of local features on a group of visual words is an effective method for visual categorization. Visual words can be related conceptually and the information can be incorporated to enhance the performance. However, conventional methods usually use visual words independently without considering this. This paper proposes a novel approach to measure the conceptual relation of visual words and incorporate the information into visual categorization. The conceptual relation is measured by the similarity of class distributions induced by visual words, accordingly visual words are grouped and images are represented on multiple levels. Categorization is taken using the support vector machine (SVM) with an effective kernel designed for matching multi-level representations. The proposed method is evaluated for video events categorization on the benchmark dataset and shows superior performance to conventional methods. Teng Li 0001, In-So Kweon |
ICIP | 1 |
| 2009 | Local-driven semi-supervised learning with multi-labelabstractIn this paper, we present a local-driven semi-supervised learning framework to propagate the labels of the training data (with multi-label) to the unlabeled data. Instead of using each datum as a vertex of graph, we encode each extracted local feature descriptor as a vertex, and then the labels for each vertex from the training data are derived based on the context among different training data, finally the decomposed labels on each vertex are further propagated to the unlabeled vertices based on the similarities measured according to the features extracted at each local regions. With the learnt local descriptor graph we can predict the semantic labels for not only the test local features but also the test images. The experiments on multi-label image annotation demonstrate the encouraging results from our proposed framework of semi-supervised learning. Teng Li 0001, Shuicheng Yan, Tao Mei 0001, In-So Kweon |
ICME | 1 |
| 2009 | Multi-video synopsis for video representation
Teng Li 0001, Tao Mei 0001, In-So Kweon, Xian-Sheng Hua 0001 |
Signal Process. | 1 |
| 2008 | A semantic region descriptor for local feature based image categorizationabstractRegion descriptor has proved to be very important for local feature based image categorization. Previous region descriptors are usually based on the statistics of low level features, such as intensity, edge response, and etc. In this paper a novel descriptor named local texton statistics (LTS) that explores the high level semantic statistical characteristics of image regions is presented. Perceptual information is obtained by applying Gaussian filter banks and the image regions are described by the statistics of different 'texton's. Using the bag of words as classification algorithm, experiments show that the proposed descriptor is superior to the previous popular SIFT descriptors on the Wang dataset. The combination of these two descriptors shows high performance for categorization on both the Wang dataset and the fifteen scene categories dataset. Teng Li 0001, In-So Kweon |
ICASSP | 1 |
| 2008 | Learning Optimal Compact Codebook for Efficient Object CategorizationabstractRepresentation of images using the distribution of local features on a visual codebook is an effective method for object categorization. Typically, discriminative capability of the codebook can lead to a better performance. However, conventional methods usually use clustering algorithms to learn codebooks without considering this. This paper presents a novel approach of learning optimal compact codebooks by selecting a subset of discriminative codes from a large codebook. Firstly, the Gaussian models of object categories based on a single code are learned from the distribution of local features within each image. Then two discriminative criteria, i.e. likelihood ratio and Fisher, are introduced to evaluate how each code contributes to the categorization. We evaluate the optimal codebooks constructed by these two criteria on Caltech-4 dataset, and report superior performance of object categorization compared with traditional K-means method with the same size of codebook. Teng Li 0001, Tao Mei 0001, In-So Kweon |
WACV | 1 |