Jun Li 0010

dblp:l/JunLi10 · DBLP profile ↗
← Back
41ranked-venue papers
17as first author
11since 2021 · last 2026
0000-0002-1336-2241ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 13 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 8 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multimodal 3D Monitoring and Visual Analytics via Dynamic Frequency Residual Splatting
Yongfeng Shan, Christy Jie Liang, Daming Luo, Chenxuan Zhou, Xiaoru Yuan, Jun Li 0010
PacificVis7
2026 ONIR: Object-Noted Tagging for Aerial Image Captioning generation
abstract
Automated captioning for remote sensing imagery often struggles to balance the high descriptive power of large models with the deployment feasibility of smaller ones. To bridge this gap, this paper introduces ONIR, a LLM-efficient, tag-guided framework that empowers compact language models (1-3B parameters) to achieve state-of-the-art captioning accuracy. Specifically, the proposed approach synthesizes a large-scale pseudo-caption dataset by leveraging GPT-4O on existing segmentation benchmarks. Explicit semantic tags are then extracted to train a multi-label Contrastive Language-Image Pre-Training (CLIP) encoder, providing interpretable visual guidance. To maintain parameter efficiency, the architecture incorporates a simple Multilayer Perceptron (MLP) bridge and a two-stage LoRA fine-tuning strategy. Extensive experiments on standard benchmark dataset, such as UCM and Sydney Captions, demonstrate that ONIR significantly outperforms models up to four times its size (7-13B). By combining superior performance with computational efficiency and tag-based controllability, ONIR offers a highly practical solution for real-world remote sensing applications.
Xing Zi, Tengjun Ni, Xianjing Fan, Xian Tao, Xinyi Gong, Jun Li 0010, Ali Braytee, Mukesh Prasad
J. Vis. Commun. Image Represent.6
2026 DyLite-VSR: A dynamic and lightweight video super-resolution network for highly compressed videos
abstract
Services based on various forms of video streaming deliver content to clients by compressing originally captured videos using video codecs. During this process, due to limitations such as network bandwidth and display device capabilities, the original resolution is often reduced or a higher compression rate is applied, resulting in smaller data sizes being transmitted. Consequently, the video quality experienced by end users is often degraded. To address this issue, numerous AI-based Video Super-Resolution (VSR) techniques have been proposed. However, in addition to their high architectural complexity, many of these models require a large number of input frames or rely on recurrent frameworks, which further complicate both training and inference. These factors present significant challenges for deployment in real-time video services. In this paper, we propose a lightweight VSR model that achieves high performance even on low-quality, highly compressed video content. We propose an adaptive inference method that dynamically selects between lightweight and enhanced processing modules based on the input video’s resolution, thereby improving the suitability of our approach for real-time video streaming applications. Additionally, we visualize the performance of the proposed model with Grad-CAM to demonstrate its effectiveness compared to existing methods.
Ilhwan Kwon, Jun Li 0010, Mahardhika Pratama, Mukesh Prasad
Knowl. Based Syst.2
2025 RSVLM-QA: A Benchmark Dataset for Remote Sensing Vision Language Model-based Question Answering
abstract
Visual Question Answering (VQA) in remote sensing (RS) is pivotal for interpreting Earth observation data. However, existing RS VQA datasets are constrained by limitations in annotation richness, question diversity, and the assessment of specific reasoning capabilities. This paper introduces Remote Sensing Vision Language Model Question Answering (RSVLM-QA) dataset, a new large-scale, content-rich VQA dataset for the RS domain. RSVLM-QA is constructed by integrating data from several prominent RS segmentation and detection datasets: WHU, LoveDA, INRIA, and iSAID. We employ an innovative dual-track annotation generation pipeline. Firstly, we leverage Large Language Models (LLMs), specifically GPT-4.1, with meticulously designed prompts to automatically generate a suite of detailed annotations including image captions, spatial relations, and semantic tags, alongside complex caption-based VQA pairs. Secondly, to address the challenging task of object counting in RS imagery, we have developed a specialized automated process that extracts object counts directly from the original segmentation data; GPT-4.1 then formulates natural language answers from these counts, which are paired with preset question templates to create counting QA pairs. RSVLM-QA comprises 13,820 images and 162,373 VQA pairs, featuring extensive annotations and diverse question types. We provide a detailed statistical analysis of the dataset and a comparison with existing RS VQA benchmarks, highlighting the superior depth and breadth of RSVLM-QA's annotations. Furthermore, we conduct benchmark experiments on Six mainstream Vision Language Models (VLMs), demonstrating that RSVLM-QA effectively evaluates and challenges the understanding and reasoning abilities of current VLMs in the RS domain. We believe RSVLM-QA will serve as a pivotal resource for the RS VQA and VLM research communities, poised to catalyze advancements in the field. The dataset, generation code, and benchmark models are publicly available at https://github.com/StarZi0213/RSVLM-QA.
Xing Zi, Jinghao Xiao, Yunxiao Shi, Xian Tao, Jun Li 0010, Ali Braytee, Mukesh Prasad
ACM Multimedia5
2025 Lightweight Motion-Aware Video Super-Resolution for Compressed Videos
Ilhwan Kwon, Jun Li 0010, Rajiv Ratn Shah, Mukesh Prasad
MMM (2)2
2025 Dynamic Appearance Particle Neural Radiance Field
abstract
Neural Radiance Fields (NeRFs) have shown great potential in modeling 3D scenes. Dynamic NeRFs extend this model by capturing time-varying elements, typically using deformation fields. The existing dynamic NeRFs employ a similar Eulerian representation for both light radiance and deformation fields. This leads to a close coupling of appearance and motion and lacks a physical interpretation. In this work, we propose Dynamic Appearance Particle Neural Radiance Field (DAP-NeRF), which introduces particle-based representation to model the motions of visual elements in a dynamic 3D scene. DAP-NeRF consists of the superposition of a static field and a dynamic field. The dynamic field is quantized as a collection of appearance particles, which carries the visual information of a small dynamic element in the scene and is equipped with a motion model. All components, including the static field, the visual features and the motion models of particles, are learned from monocular videos without any prior geometric knowledge of the scene. We develop an efficient computational framework for the particle-based model. We also construct a new dataset to evaluate motion modeling. Experimental results show that DAP-NeRF is an effective technique to capture not only the appearance but also the physically meaningful motions in a 3D dynamic scene. Code is available at:https://github.com/Cenbylin/DAP-NeRF.
Ancheng Lin, Yusheng Xiang, Jun Li 0010, Mukesh Prasad
IEEE Trans. Circuits Syst. Video Technol.3
2024 BDC Dataset: A Comprehensive Dataset for Automated Build Damage Classification
Xing Zi, Yunxiao Shi, Taoyuan Zhu, Kairui Jin, Xian Tao, Jun Li 0010, Karthick Thiyagarajan, Mukesh Prasad
ADMA (1)6
2024 Cooperative Markov Decision Process model for human-machine co-adaptation in robot-assisted rehabilitation
Kairui Guo, Adrian Cheng, Jun Li 0010, Rob Duffield, Steven W. Su
Knowl. Based Syst.4
2022 A Multiscale Wavelet Kernel Regularization-Based Feature Extraction Method for Electronic Nose
abstract
In the electronic nose (e-nose), a stable feature representation of the gas sensor’s response is a key step to realize subsequent odor identification algorithms. However, the noises in gas sensors hinder the acquisition of such features. In order to solve this problem, this article proposes a stable feature extraction algorithm which takes the impulse response of the e-nose system as the feature. The impulse response is estimated from a nonparametric model constrained by a multiscale wavelet kernel regularization matrix. The kernel regularization matrix equips the proposed feature extraction method with an ability in resistance to random noise. A numerical experiment proves that compared with single-scale kernel regularization, the use of multiscale wavelet kernel helps to achieve more stable and accurate impulse response estimation. Then, a field experiment is conducted to demonstrate the performance of the proposed features. This experiment aims to identify four different whiskies measured by a self-designed e-nose with four commercial gas sensors. Under the framework of transfer learning, the classification result based on the proposed features outperforms those using other considered features. The accuracy of whisky identification reaches 92.00%, showing a good potential of applying the proposed feature representations in the area of e-noses.
Taoping Liu, Wentian Zhang, Jun Li 0010, Maiken Ueland, Shari L. Forbes, Wei Xing Zheng 0001, Steven W. Su
IEEE Trans. Syst. Man Cybern. Syst.3
2021 F-Net: Fusion Neural Network for Vehicle Trajectory Prediction in Autonomous Driving
abstract
Recent research has been remarkable in recurrent neural networks (RNNs) on sequence-to-sequence problems for image caption, and promising in convolutional neural networks (CNNs) on spatial analysis problems for image detection and sematic segmentation problems. In this paper, based on recurrent neural networks and convolutional neural networks, we propose a fusion neural network architecture named F-Net to deal with vehicle trajectory prediction on highway and urban scenarios in autonomous driving applications. The novelty of the proposed method is the attention mechanism that affects effectively in the progress of both RNN and CNN feature extraction. Besides, our sufficient usage of raw sensor data protects scene texture information of environment and interaction among surrounding vehicles. Experimental results on the nuScene dataset show that our proposed method outperforms the state-of-the-art methods.
Jue Wang 0004, Ping Wang 0003, Chao Zhang 0001, Kuifeng Su, Jun Li 0010
ICASSP5
2021 Face Hallucination With Finishing Touches
abstract
Obtaining a high-quality frontal face image from a low-resolution (LR) non-frontal face image is primarily important for many facial analysis applications. However, mainstreams either focus on super-resolving near-frontal LR faces or frontalizing non-frontal high-resolution (HR) faces. It is desirable to perform both tasks seamlessly for daily-life unconstrained face images. In this paper, we present a novel Vivid Face Hallucination Generative Adversarial Network (VividGAN) for simultaneously super-resolving and frontalizing tiny non-frontal face images. VividGAN consists of coarse-level and fine-level Face Hallucination Networks (FHnet) and two discriminators, i.e., Coarse-D and Fine-D. The coarse-level FHnet generates a frontal coarse HR face and then the fine-level FHnet makes use of the facial component appearance prior, i.e., fine-grained facial components, to attain a frontal HR face image with authentic details. In the fine-level FHnet, we also design a facial component-aware module that adopts the facial geometry guidance as clues to accurately align and merge the frontal coarse HR face and prior information. Meanwhile, two-level discriminators are designed to capture both the global outline of a face image as well as detailed facial characteristics. The Coarse-D enforces the coarsely hallucinated faces to be upright and complete while the Fine-D focuses on the fine hallucinated ones for sharper details. Extensive experiments demonstrate that our VividGAN achieves photo-realistic frontal HR faces, reaching superior performance in downstream tasks, i.e., face recognition and expression classification, compared with other state-of-the-art methods.
Yang Zhang 0067, Ivor W. Tsang, Jun Li 0010, Ping Liu 0004, Xiaobo Lu, Xin Yu 0002
IEEE Trans. Image Process.3
2020 Enhancing Automated COVID-19 Chest X-ray Diagnosis by Image-to-Image GAN Translation
abstract
The severe pneumonia induced by the infection of the SARS-CoV-2 virus causes massive death in the ongoing COVID-19 pandemic. The early detection of the SARS-CoV-2 induced pneumonia relies on the unique patterns of the chest XRay images. Deep learning is a data-greedy algorithm to achieve high performance when adequately trained. A common challenge for machine learning in the medical domain is the accessibility to properly annotated data. In this study, we apply a conditional adversarial network (cGAN) to perform image to image (Pix2Pix) translation from the non-COVID-19 chest X-Ray domain to the COVID-19 chest X-Ray domain. The objective is to learn a mapping from the normal chest X-Ray visual patterns to the COVID-19 pneumonia chest X-ray patterns. The original dataset has a typical imbalanced issue because it contains only 219 COVID-19 positive images but has 1,341 images for normal chest X-Ray and 1,345 images for viral pneumonia. A U-Net based architecture is applied for the image-to-image translation to generate synthesized COVID-19 X-Ray chest images from the normal chest X-ray images. A 50-convolutional-layer residual net (ResNet) architecture is applied for the final classification task. After training the GAN model for 100 epochs, we use the GAN generator to translate 1,100 COVID-19 images from the normal X-Ray to form a balanced training dataset (3,762 images) for the classification task. The ResNet based classifier trained by the enhanced dataset achieves the classification accuracy of 97.8% compared to 96.1% in the transfer learning mode. When trained with the original imbalanced dataset, the model achieves an accuracy of 96.1% compared to 95.6% in the training from trainby-scratch model. In addition, the classifier trained by the enhanced dataset has more stable measures in precision, recall, and F1 scores across different image classes. We conclude that the GAN-based data enhancement strategy is applicable to most medical image pattern recognition tasks, and it provides an effective way to solve the common expertise dependence issue in the medical domain.
Zhaohui Liang, Jimmy Huang 0001, Jun Li 0010, Stephen Chan
BIBM3
2019 Learn to focus on objects for visual detection
Zijing Chen, Jun Li 0010, Xinhua You
Neurocomputing2
2019 Face detection and alignment method for driver on highroad based on improved multi-task cascaded convolutional networks
Yang Zhang 0067, Peihua Lv, Xiaobo Lu, Jun Li 0010
Multim. Tools Appl.4
2018 Residual MeshNet: Learning to Deform Meshes for Single-View 3D Reconstruction
abstract
This work presents a novel architecture of deep neural networks to generate meshes approximating the surface of a 3D object from a single image. Compared to existing learning-based 3D reconstruction models, our architecture is characterized by (1) deep mesh deformation stacks with residual network design, where a simple mesh is transformed to approximate the target surface and undergoes multiple deformation steps to progressively refine the result and reduce the residuals, and (2) parallel paths per deformation step, which can exponentially enrich the generated meshes using deeper structure and more model parameters. We also propose novel regularization scheme that encourages the meshes to be both globally complementary to cover the target surface and locally consistent with each other. Empirical evaluation on benchmark datasets show advantage of the proposed architecture over existing methods.
Junyi Pan, Jun Li 0010, Xiaoguang Han 0001, Kui Jia
3DV2
2018 Utilizing Information from Task-Independent Aspects via GAN-Assisted Knowledge Transfer
abstract
Observed data often have multiple labels with respect to different aspects. For example, a picture can have one label specifying the contents in terms of the object category such as aeroplane, building, cat, etc. and in the meanwhile have another label describing the image style such as photo-realistic or artistic. The central idea of this work is that any annotation of the data contains precious knowledge and is not to be foregone: an analytic task focusing on one aspect of the data can benefit from the knowledge transferred from the other aspects. We propose a passive knowledge transfer scheme for deep neural network training based on the generative adversarial nets (GANs). The adversarial training scheme encourages the nets to encode data into representations that are both discriminative for the target aspect and invariant with respect to the irrelevant aspects. We show that the scheme mixes the conditional distributions of the encoded data on the irrelevant aspects, by the theory on the link between the GAN framework and the Wasserstein metric in distribution spaces. Moreover, we empirically verified the method by i) classifying images despite influence by geometric transform and ii) recognizing the movements (geometric transform) regardless the image contents.
Lunkai Fu, Jun Li 0010, Langxiong Zhou, Zhenyuan Ma, Mukesh Prasad
IJCNN2
2018 GAN2C: Information Completion GAN with Dual Consistency Constraints
abstract
This paper proposes an information completion technique, GAN2C, by imposing dual consistency constraints (2C) to a closed loop encoder-decoder architecture based on the generative adversarial nets (GAN). When adopting deep neural networks as function approximators, GAN2C enables highly effective multi-modality image conversion with sparse observation in the target modes. For empirical demonstration and model evaluation, we show that trained deep neural networks in GAN2C can infer colors for grayscale images, as well as estimate rich 3D information of a scene by densely predicting the depths. The results of the experiments show that in both tasks GAN2C as a generic framework has been comparable to or advanced the state-of-the-art performance which are achieved by highly specialized systems. Code is available at https://github.com/AdalinZhang/GAN2C.
Lujuan Zhang, Jun Li 0010, Zhenyuan Ma, Mukesh Prasad
IJCNN2
2018 Shakeout: A New Approach to Regularized Deep Neural Network Training
abstract
Recent years have witnessed the success of deep neural networks in dealing with a plenty of practical problems. Dropout has played an essential role in many successful deep neural networks, by inducing regularization in the model training. In this paper, we present a new regularized training approach: Shakeout. Instead of randomly discarding units as Dropout does at the training stage, Shakeout randomly chooses to enhance or reverse each unit's contribution to the next layer. This minor modification of Dropout has the statistical trait: the regularizer induced by Shakeout adaptively combines , and regularization terms. Our classification experiments with representative deep architectures on image datasets MNIST, CIFAR-10 and ImageNet show that Shakeout deals with over-fitting effectively and outperforms Dropout. We empirically demonstrate that Shakeout leads to sparser weights under both unsupervised and supervised settings. Shakeout also leads to the grouping effect of the input units in a layer. Considering the weights in reflecting the importance of connections, Shakeout is superior to Dropout, which is valuable for the deep model compression. Moreover, we demonstrate that Shakeout can effectively reduce the instability of the training process of the deep architecture.
Guoliang Kang, Jun Li 0010, Dacheng Tao
IEEE Trans. Pattern Anal. Mach. Intell.2
2017 Generic Pixel Level Object Tracker Using Bi-Channel Fully Convolutional Network
Zijing Chen, Jun Li 0010, Zhe Chen 0013, Xinge You
ICONIP (1)2
2017 Dynamically Modulated Mask Sparse Tracking
abstract
Visual tracking is a critical task in many computer vision applications such as surveillance and robotics. However, although the robustness to local corruptions has been improved, prevailing trackers are still sensitive to large scale corruptions, such as occlusions and illumination variations. In this paper, we propose a novel robust object tracking technique depends on subspace learning-based appearance model. Our contributions are twofold. First, mask templates produced by frame difference are introduced into our template dictionary. Since the mask templates contain abundant structure information of corruptions, the model could encode information about the corruptions on the object more efficiently. Meanwhile, the robustness of the tracker is further enhanced by adopting system dynamic, which considers the moving tendency of the object. Second, we provide the theoretic guarantee that by adapting the modulated template dictionary system, our new sparse model can be solved by the accelerated proximal gradient algorithm as efficient as in traditional sparse tracking methods. Extensive experimental evaluations demonstrate that our method significantly outperforms 21 other cutting-edge algorithms in both speed and tracking accuracy, especially when there are challenges such as pose variation, occlusion, and illumination changes.
Zijing Chen, Xinge You, Boxuan Zhong, Jun Li 0010, Dacheng Tao
IEEE Trans. Cybern.4
2017 Deep Neural Network for Structural Prediction and Lane Detection in Traffic Scene
abstract
Hierarchical neural networks have been shown to be effective in learning representative image features and recognizing object classes. However, most existing networks combine the low/middle level cues for classification without accounting for any spatial structures. For applications such as understanding a scene, how the visual cues are spatially distributed in an image becomes essential for successful analysis. This paper extends the framework of deep neural networks by accounting for the structural cues in the visual signals. In particular, two kinds of neural networks have been proposed. First, we develop a multitask deep convolutional network, which simultaneously detects the presence of the target and the geometric attributes (location and orientation) of the target with respect to the region of interest. Second, a recurrent neuron layer is adopted for structured visual detection. The recurrent neurons can deal with the spatial distribution of visible cues belonging to an object whose shape or structure is difficult to explicitly define. Both the networks are demonstrated by the practical task of detecting lane boundaries in traffic scenes. The multitask convolutional neural network provides auxiliary geometric information to help the subsequent modeling of the given lane structures. The recurrent neural network automatically detects lane boundaries, including those areas containing no marks, without any explicit prior knowledge or secondary modeling.
Jun Li 0010, Xue Mei, Danil V. Prokhorov, Dacheng Tao
IEEE Trans. Neural Networks Learn. Syst.1
2016 Shakeout: A New Regularized Deep Neural Network Training Scheme
abstract
Recent years have witnessed the success of deep neural networks in dealing with a plenty of practical problems. The invention of effective training techniques largely contributes to this success. The so-called "Dropout" training scheme is one of the most powerful tool to reduce over-fitting. From the statistic point of view, Dropout works by implicitly imposing an L2 regularizer on the weights. In this paper, we present a new training scheme: Shakeout. Instead of randomly discarding units as Dropout does at the training stage, our method randomly chooses to enhance or inverse the contributions of each unit to the next layer. We show that our scheme leads to a combination of L1 regularization and L2 regularization imposed on the weights, which has been proved effective by the Elastic Net models in practice.We have empirically evaluated the Shakeout scheme and demonstrated that sparse network weights are obtained via Shakeout training. Our classification experiments on real-life image datasets MNIST and CIFAR-10 show that Shakeout deals with over-fitting effectively.
Guoliang Kang, Jun Li 0010, Dacheng Tao
AAAI2
2016 ReD-SFA: Relation Discovery Based Slow Feature Analysis for Trajectory Clustering
abstract
For spectral embedding/clustering, it is still an open problem on how to construct an relation graph to reflect the intrinsic structures in data. In this paper, we proposed an approach, named Relation Discovery based Slow Feature Analysis (ReD-SFA), for feature learning and graph construction simultaneously. Given an initial graph with only a few nearest but most reliable pairwise relations, new reliable relations are discovered by an assumption of reliability preservation, i.e., the reliable relations will preserve their reliabilities in the learnt projection subspace. We formulate the idea as a cross entropy (CE) minimization problem to reduce the discrepancy between two Bernoulli distributions parameterized by the updated distances and the existing relation graph respectively. Furthermore, to overcome the imbalanced distribution of samples, a Boosting-like strategy is proposed to balance the discovered relations over all clusters. To evaluate the proposed method, extensive experiments are performed with various trajectory clustering tasks, including motion segmentation, time series clustering and crowd detection. The results demonstrate that ReDSFA can discover reliable intra-cluster relations with high precision, and competitive clustering performance can be achieved in comparison with state-of-the-art.
Zhang Zhang 0001, Kaiqi Huang, Tieniu Tan, Peipei Yang, Jun Li 0010
CVPR5
2015 A Distributed Approach Toward Discriminative Distance Metric Learning
abstract
Distance metric learning (DML) is successful in discovering intrinsic relations in data. However, most algorithms are computationally demanding when the problem size becomes large. In this paper, we propose a discriminative metric learning algorithm, develop a distributed scheme learning metrics on moderate-sized subsets of data, and aggregate the results into a global solution. The technique leverages the power of parallel computation. The algorithm of the aggregated DML (ADML) scales well with the data size and can be controlled by the partition. We theoretically analyze and provide bounds for the error induced by the distributed treatment. We have conducted experimental evaluation of the ADML, both on specially designed tests and on practical image annotation tasks. Those tests have shown that the ADML achieves the state-of-the-art performance at only a fraction of the cost incurred by most existing methods.
Jun Li 0010, Xun Lin, Xiaoguang Rui, Yong Rui, Dacheng Tao
IEEE Trans. Neural Networks Learn. Syst.1
2013 A Bayesian Factorised Covariance Model for Image Analysis
Jun Li 0010, Dacheng Tao
IJCAI1
2013 Learning colours from textures by sparse manifold embedding
Jun Li 0010, Wei Bian 0003, Dacheng Tao, Chengqi Zhang
Signal Process.1
2013 A Bayesian Hierarchical Factorization Model for Vector Fields
abstract
Factorization-based techniques explain arrays of observations using a relatively small number of factors and provide an essential arsenal for multi-dimensional data analysis. Most factorization models are, however, developed on general arrays of scalar values. For a class of practical data arising from observing spatial signals including images, it is desirable for a model to consider general observations, e.g., handling a vector field and non-exchangeable factors, e.g., handling spatial connections between the columns and the rows of the data. In this paper, a probabilistic model for factorization is proposed. We adopt Bayesian hierarchical modeling and treat the factors as latent random variables. A Markov structure is imposed on the distribution of factors to account for the spatial connections. The model is designed to represent vector arrays sampled from fields of continuous domains. Therefore, a tailored observation model is developed to represent the link between the factor product and the data. The proposed technique has been shown effective in analyzing optical flow fields computed on both synthetic images and real-life videoclips.
Jun Li 0010, Dacheng Tao
IEEE Trans. Image Process.1
2013 Simple Exponential Family PCA
abstract
Principal component analysis (PCA) is a widely used model for dimensionality reduction. In this paper, we address the problem of determining the intrinsic dimensionality of a general type data population by selecting the number of principal components for a generalized PCA model. In particular, we propose a generalized Bayesian PCA model, which deals with general type data by employing exponential family distributions. Model selection is realized by empirical Bayesian inference of the model. We name the model as simple exponential family PCA (SePCA), since it embraces both the principal of using a simple model for data representation and the practice of using a simplified computational procedure for the inference. Our analysis shows that the empirical Bayesian inference in SePCA formally realizes an intuitive criterion for PCA model selection - a preserved principal component must sufficiently correlate to data variance that is uncorrelated to the other principal components. Experiments on synthetic and real data sets demonstrate effectiveness of SePCA and exemplify its characteristics for model selection.
Jun Li 0010, Dacheng Tao
IEEE Trans. Neural Networks Learn. Syst.1
2013 Exponential Family Factors for Bayesian Factor Analysis
abstract
Expressing data as linear functions of a small number of unknown variables is a useful approach employed by several classical data analysis methods, e.g., factor analysis, principal component analysis, or latent semantic indexing. These models represent the data using the product of two factors. In practice, one important concern is how to link the learned factors to relevant quantities in the context of the application. To this end, various specialized forms of the factors have been proposed to improve interpretability. Toward developing a unified view and clarifying the statistical significance of the specialized factors, we propose a Bayesian model family. We employ exponential family distributions to specify various types of factors, which provide a unified probabilistic formulation. A Gibbs sampling procedure is constructed as a general computation routine. We verify the model by experiments, in which the proposed model is shown to be effective in both emulating existing models and motivating new model designs for particular problem settings.
Jun Li 0010, Dacheng Tao
IEEE Trans. Neural Networks Learn. Syst.1
2012 Sampling Normal Distribution Restricted on Multiple Regions
Jun Li 0010, Dacheng Tao
ICONIP (1)1
2012 Segment-Based Features for Time Series Classification
abstract
In this paper, we propose an approach termed segment-based features (SBFs) to classify time series. The approach is inspired by the success of the component- or part-based methods of object recognition in computer vision, in which a visual object is described as a number of characteristic parts and the relations among the parts. Utilizing this idea in the problem of time series classification, a time series is represented as a set of segments and the corresponding temporal relations. First, a number of interest segments are extracted by interest point detection with automatic scale selection. Then, a number of feature prototypes are collected by random sampling from the segment set, where each feature prototype may include single segment or multiple ordered segments. Subsequently, each time series is transformed to a standard feature vector, i.e. SBF, where each entry in the SBF is calculated as the maximum response (maximum similarity) of the corresponding feature prototype to the segment set of the time series. Based on the original SBF, an incremental feature selection algorithm is conducted to form a compact and discriminative feature representation. Finally, a multi-class support vector machine is trained to classify the test time series. Extensive experiments on different time series datasets, including one synthetic control dataset, two sign language datasets and one gait dynamics dataset, have been performed to evaluate the proposed SBF method. Compared with other state-of-the-art methods, our approach achieves superior classification performance, which clearly validates the advantages of the proposed method.
Zhang Zhang 0001, Jun Cheng 0002, Jun Li 0010, Wei Bian 0003, Dacheng Tao
Comput. J.3
2012 A probabilistic model for image representation via multiple patterns
Jun Li 0010, Dacheng Tao, Xuelong Li 0001
Pattern Recognit.1
2012 On Preserving Original Variables in Bayesian PCA With Application to Image Analysis
abstract
Principal component analysis (PCA) computes a succinct data representation by converting the data to a few new variables while retaining maximum variation. However, the new variables are difficult to interpret, because each one is combined with all of the original input variables and has obscure semantics. Under the umbrella of Bayesian data analysis, this paper presents a new prior to explicitly regularize combinations of input variables. In particular, the prior penalizes pair-wise products of the coefficients of PCA and encourages a sparse model. Compared to the commonly used l1regularizer, the proposed prior encourages the sparsity pattern in the resultant coefficients to be consistent with the intrinsic groups in the original input variables. Moreover, the proposed prior can be explained as recovering a robust estimation of the covariance matrix for PCA. The proposed model is suited for analyzing visual data, where it encourages the output variables to correspond to meaningful parts in the data. We demonstrate the characteristics and effectiveness of the proposed technique through experiments on both synthetic and real data.
Jun Li 0010, Dacheng Tao
IEEE Trans. Image Process.1
2011 A Probabilistic Model for Discovering High Level Brain Activities from fMRI
Jun Li 0010, Dacheng Tao
ICONIP (1)1
2010 Boosted Dynamic Cognitive Activity Recognition from Brain Images
abstract
Functional Magnetic Resonance Imaging (fMRI) has become an important diagnostic tool for measuring brain haemodynamics. Previous research on analysing fMRI data mainly focuses on detecting low-level neuron activation from the ensued haemodynamic activities. An important recent advance is to show that the high-level cognitive status is recognisable from a period of fMRI records. Nevertheless, it would also be helpful to reveal dynamics of cognitive activities during the period. In this paper, we tackle the problem of discovering the dynamic cognitive activities by proposing an algorithm of boosted structure learning. We employ statistic model of random fields to represent the dynamics of the brain. To exploit the rich fMRI observations with reasonable model complexity, we build multiple models, where one model links the cognitive activities to only a fraction of the fMRI observations. We combine the simple models by using an altered AdaBoost scheme for multi-class structure learning and show theoretical justification of the proposed scheme. Empirical test shows the method effectively links the physiological and the psychological activities of the brain.
Jun Li 0010, Dacheng Tao
ICMLA1
2009 Finding representative landmarks of data on manifolds
Jun Li 0010, Pengwei Hao
Pattern Recognit.1
2008 Reliable Representation of Data on Manifolds
abstract
The manifold learning algorithms are promising data analysis tools. However, to fit an unseen point in a learned model, the point must be located in the training set, which limits its scalability. In this paper, we discuss how to select landmarks from the data to help locate the test points. Our method is for data on manifolds: the way the landmarks represent the data in the ambient space should resemble the way they represent the data on the manifold. Compared to the previous research, (i) Our test foregoes the requirement of knowing the intrinsic manifold dimension and thus is more applicable and robust. (ii) Our selection implies a provable topology preservation property. (iii) We also provide a way to improve existing landmarks. Experiments on the synthetic data and the real data have been done. The results support the proposed properties and algorithms. 1
Jun Li 0010, Pengwei Hao
BMVC1
2008 Transferring Colours to Grayscale Images by Locally Linear Embedding
abstract
In this paper, we propose a learning-based method for adding colours to grayscale images. In contrast to many previous computer-aided colourizing methods, which require intensive and accurate human intervention, our method needs only the user to provide a colourful image of the similar content as the grayscale image. We accept the “image manifold ” assumption and apply manifold learning methods to model the relations between the chromatic channels and the gray levels in the training images. Then we synthesize the objective chromatic channels using the learned relations. Experiments show that our method gives superior results to those of the previous work. 1
Jun Li 0010, Pengwei Hao, Chao Zhang 0001
BMVC1
2008 Hallucinating faces from thermal infrared images
abstract
This paper addresses the face hallucination problem of converting thermal infrared face images into photo-realistic ones. It is a challenging task because the two modalities are of dramatical difference, which makes many developed linear models inapplicable. We propose a learning-based framework synthesizing the normal face from the infrared input. Compared to the previous work, we further exploit the local linearity in not only the image spatial domain but also the image manifolds. We have also developed a measurement of the variance between an input and its prediction, thus we can apply the Markov random field model to the predicted normal face to improve the hallucination result. Experimental results show the advantage of our algorithm over the existing methods. Our algorithm can be readily generalized to solve other multi-modal image conversion problems as well.
Jun Li 0010, Pengwei Hao, Chao Zhang 0001, Mingsong Dou
ICIP1
2007 Converting Thermal Infrared Face Images into Normal Gray-Level Images
Mingsong Dou, Chao Zhang 0001, Pengwei Hao, Jun Li 0010
ACCV (2)4
2007 Hierarchical Structuring of Data on Manifolds
abstract
Manifold learning methods are promising data analysis tools. However, if we locate a new test sample on the manifold, we have to find its embedding by making use of the learned embedded representation of the training samples. This process often involves accessing considerable volume of data for large sample set. In this paper, an approach of selecting "landmark points" from the given samples is proposed for hierarchical structuring of data on manifolds. The selection is made such that if one use the Voronoi diagram generated by the landmark points in the ambient space to partition the embeded manifold, the topology of the manifold is preserved. The landmark points then are used to recursively construct a hierarchical structure of the data. Thus it can speed up queries in a manifold data set. It is a general framework that can fit any manifold learning algorithm as long as its result of an input can be predicted by the results of the neighbor inputs. Compared to the existing techniques of organizing data based on spatial partitioning, our method preserves the topology of the latent space of the data. Different from manifold learning algorithms that use landmark points to reduce complexity, our approach is designed for fast retrieval of samples. It may find its way in high dimensional data analysis such as indexing, clustering, and progressive compression. More importantly, it extends the manifold learning methods to applications in which they were previously considered to be not fast enough. Our algorithm is stable and fast, and its validity is proved mathematically.
Jun Li 0010, Pengwei Hao
CVPR1