Junbiao Pang

dblp:55/1995 · DBLP profile ↗
← Back
51ranked-venue papers
22as first author
14since 2021 · last 2026
0000-0001-8153-7229ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 29 · 10 first-author · 6 since 2021Artificial intelligence and machine learning · 19 · 8 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Efficiently Seeking Flat Minima for Better Generalization in Fine-Tuning Large Language Models and Beyond
abstract
Little research explores the correlation between the expressive ability and generalization ability of the low-rank adaptation (LoRA). Sharpness-Aware Minimization (SAM) improves model generalization for both Convolutional Neural Networks (CNNs) and Transformers by encouraging convergence to locally flat minima. However, the connection between sharpness and generalization has not been fully explored for LoRA due to the lack of tools to either empirically seek flat minima or develop theoretical methods. In this work, we propose Flat Minima LoRA (FMLoRA) and its efficient version i.e., EFMLoRA, to seek flat minima for LoRA. Concretely, we theoretically demonstrate that perturbations in the full parameter space can be transferred to the low-rank subspace. This approach eliminates the potential interference introduced by perturbations across multiple matrices in the low-rank subspace. Our extensive experiments on large language models and vision-language models demonstrate that EFMLoRA achieves optimization efficiency comparable to that of LoRA while simultaneously attaining comparable or even better performance. For example, on the GLUE dataset with RoBERTa-large, EFMLoRA outperforms LoRA and full fine-tuning by 1.0% and 0.5% on average, respectively. On vision-language models e.g., Qwen-VL-Chat, there are performance improvements of 1.5% and 1.0% on the SQA and VizWiz datasets, respectively. These empirical results also verify that the generalization of LoRA is closely related to sharpness, which is omitted by previous methods.
Jiaxin Deng, Qingcheng Zhu, Junbiao Pang, Linlin Yang 0001, Zhongqian Fu, Baochang Zhang 0001
AAAI3
2026 Accurate pixel-wise keypoint localization for rectangle symbol spotting in CAD images
Jiaxin Deng, Junbiao Pang, Zailin Dong, Mengyuan Zhu
Multim. Syst.3
2026 Precise 2D mouse pose estimation via multi-scale context and sensitive-aware loss from low illumination environment
Yubin Geng, Jiaxin Deng, Junbiao Pang
Multim. Syst.4
2026 Handling maximal value drift in heatmap-based point localization via progressive order-preserving regularization
Jiaxin Deng, Zailin Dong, Junbiao Pang
Multim. Syst.4
2026 Adaptively sampling-reusing-mixing decomposed gradients to speed up sharpness aware minimization
Jiaxin Deng, Junbiao Pang, Baochang Zhang 0001
Pattern Recognit.2
2026 Unsupervised Abnormal Stop Detection for Long-Distance Coaches With Low-Frequency GPS
abstract
In our urban life, long-distance coaches provide a convenient yet economical approach to the public. One notable problem is to discover the abnormal stop of the coaches for some reasons,i.e., illegal pick up/ drop off passengers on the way which possibly endangers the safety of passengers. It has become a pressing issue to detect the abnormal stop with low-quality GPS. In this paper, we propose an unsupervised method that helps transportation managers efficiently discover Abnormal Stop Detection (ASD) behaviors for long distance coaches. Concretely, our method converts the ASD problem into an unsupervised clustering framework in which both the normal stops and the abnormal ones are decomposed. We propose a stop duration model for low-frequency GPS data that approximates linear speed changes when the coach stops within a short time interval. Secondly, we strip the abnormal stops from the normal stop points by the low rank assumption. The proposed method is conceptually simple yet efficient, by leveraging low rank assumption to handle normal stop points, our approach enables domain experts to discover the ASD for coaches, from a case study motivated by traffic managers. Dataset and code are publicly available at:https://github.com/pangjunbiao/IPPs
Jiaxin Deng, Junbiao Pang, Muhammad Ayub Sabir, Haitao Yu 0008
IEEE Trans. Intell. Transp. Syst.3
2025 Asymptotic Unbiased Sample Sampling to Speed Up Sharpness-Aware Minimization
abstract
Sharpness-Aware Minimization (SAM) has emerged as a promising approach for effectively reducing the generalization error. However, SAM incurs twice the computational cost compared to the base optimizer (e.g., SGD). We propose Asymptotic Unbiased data sampling to accelerate SAM (AUSAM), which maintains the model's generalization capacity while significantly enhancing computational efficiency. Concretely, we probabilistically sample a subset of data points beneficial for SAM optimization based on a theoretically guaranteed criterion, i.e., the Gradient Norm of each Sample (GNS). We further approximate the GNS by evaluating the difference in loss values before and after perturbation in SAM. As a plug-and-play, architecture-agnostic method, our approach consistently accelerates SAM across various tasks and networks, i.e., classification, human pose estimation, and network quantization. On CIFAR-10/100 and Tiny-ImageNet, AUSAM achieves results comparable to SAM while providing a speedup of over 70%. By adjusting hyperparameters, AUSAM can match the speed of the base optimizer while significantly surpassing the base optimizer's performance. Compared to recent dynamic data pruning methods, AUSAM is better suited for SAM and excels in maintaining performance. Additionally, AUSAM accelerates optimization in human pose estimation and model quantization without sacrificing performance, demonstrating its broad practicality.
Jiaxin Deng, Junbiao Pang, Baochang Zhang 0001, Guodong Guo
AAAI2
2025 Decorrelating structure via adapters makes ensemble learning practical for Semi-supervised Learning
Jiaqi Wu 0013, Junbiao Pang, Qingming Huang
Eng. Appl. Artif. Intell.2
2025 Bundle fragments into a whole: Mining more complete clusters via submodular selection of interesting webpages for web topic detection
Junbiao Pang, Anjing Hu, Qingming Huang
Expert Syst. Appl.1
2025 Towards scalable topic detection on web via simulating Lévy walks nature of topics in similarity space
Junbiao Pang, Qingming Huang
Inf. Sci.1
2025 GLAD-TL: A Time-Sensitive Crowdsourced Model for Robust Detection of Fake Taxis in Urban Traffic Surveillance
abstract
Fake taxis pose serious risks in the transportation industry, potentially leading to traffic accidents and negative social consequences. A common approach compares a taxi’s GPS location with its visual location captured by surveillance cameras. However, accurate license plate number (LPN) recognition is often compromised by factors such as lighting, angle, and image quality, and some drivers deliberately avoid exposing plates to cameras. Accurately estimating each plate number is thus essential for identifying fake taxis, especially in high-traffic areas like airports and railway stations, where vehicles pass through multiple fields of view. This article introduces GLAD-TL, a crowdsourced probabilistic model with time-dependent low-rank constraints that enhances LPN accuracy by adapting to variations in camera performance. By modeling differences in camera capability and recognition task difficulty over time, GLAD-TL provides a robust and scalable solution for urban fake taxi detection without additional infrastructure. Experiments on real-world data demonstrate the model’s ability to achieve a detection accuracy of 76.10%, underscoring its potential for deployment in dense urban environments.
Fatima Ashraf, Jiale Bian, Junbiao Pang, Haitao Yu 0008, Muhammad Ayub Sabir
IEEE Trans. Comput. Soc. Syst.3
2025 Modeling Multi-Granularity Context Information Flow for Pavement Crack Detection
abstract
Pavement cracks have a highly complex spatial structure, a low contrasting background and a weak spatial continuity, posing a significant challenge to an effective crack detection method. To precisely localize crack from an image, it is critical to effectively extract and aggregate multi-granularity context, including the fine-grained local context around the cracks (in spatial-level) and the coarse-grained semantics (in semantic-level). In this paper, we apply the dilated convolution as the backbone feature extractor to model local context, then we build a context guidance module to leverage semantic context to guide local feature extraction at multiple stages. To handle label alignment between stages, we apply the Multiple Instance Learning (MIL) strategy to align the feature between two stages. In addition, to our best knowledge, we have released the largest, most complex and most challenging Bitumen Pavement Crack (BPC) dataset. The experimental results on the three crack datasets demonstrate that the proposed method performs well and outperforms the current state-of-the-art methods. On BPC, the proposed model achieved AP 88.32% with the 16.89 M parameters under the 45.36 GFlops runing speed. Datset and code are publicly available at: https://github.com/pangjunbiao/BPC-Crack-Dataset.
Junbiao Pang, Baocheng Xiong, Jiaqi Wu 0013, Qingming Huang
IEEE Trans. Intell. Transp. Syst.1
2024 Finding a Taxi With Illegal Driver Substitution Activity via Behavior Modelings
abstract
In our urban life, Illegal Driver Substitution (IDS) activity for a taxi is a grave unlawful activity in the taxi industry. Currently, the IDS activity is manually supervised by law enforcers, i.e., law enforcers empirically choose a taxi and inspect it. The pressing problem of this scheme is the dilemma between the limited number of law-enforcers and the large volume of taxis. In this paper, we propose a computational method that helps law enforcers efficiently find the taxis which tend to have the IDS activity. Firstly, our method converts the identification of the IDS activity to a supervised learning task. Secondly, two kinds of taxi driver behaviors, i.e., the Sleeping Time and Location (STL) behavior and the Pick-Up (PU) behavior are proposed. Thirdly, the multiple scale pooling on self-similarity is proposed to encode the individual behaviors into the universal features for all taxis. Finally, a Multiple Component-Multiple Instance Learning (MC-MIL) is proposed to handle the deficiency of the behavior features and to align the behavior features, simultaneously. Extensive experiments on a real-world data set shows that the proposed behavior features have a good generalization ability across different classifiers, and the proposed MC-MIL method suppresses the baseline methods.
Junbiao Pang, Muhammad Ayub Sabir, Zuyun Wang, Anjing Hu, Haitao Yu 0008, Qingming Huang
IEEE Trans. Intell. Transp. Syst.1
2021 Pavement Crack Detection Using Multi-stage Structural Feature Extraction Model
abstract
Pavement crack detection is of great significance for road maintenance. However, the complexity of road surfaces and the irregularity of cracks make it difficult to accurately detect crack regions. We propose a crack detection method based on structural features for the patch-wise crack detection. The novelty of this method lies on the fusion of the local patches in a multi-staged strategy. Deep supervision learning is further used to learn these features at each stage. The fusion features model the structural relevance among cracks. The experimental results prove the effectiveness of our method on the dataset collected from the industrial environments. Among these state-of-the-art methods we compared, our model achieved the best experimental results with an AP 86.97%.
Lijuan Duan, Junbiao Pang
ICIP3
2019 Fast and Accurately Measuring Crack Width via Cascade Principal Component Analysis
abstract
Crack width is an important indicator to diagnose the safety of constructions, e.g., asphalt road, concrete bridge. In practice, measuring crack width is a challenge task: (1) the irregular and non-smooth boundary makes the traditional method inefficient; (2) pixel-wise measurement guarantees the accuracy of a system and (3) understanding the damage of constructions from any pre-selected points is a mandatary requirement. To address these problems, we propose a cascade Principal Component Analysis (PCA) to efficiently measure crack width from images. Firstly, the binary crack image is obtained to describe the crack via the off-the-shelf crack detection algorithms. Secondly, given a pre-selected point, PCA is used to find the main axis of a crack. Thirdly, Robust Principal Component Analysis (RPCA) is proposed to compute the main axis of a crack with a irregular boundary. We evaluate the proposed method on a real data set. The experimental results show that the proposed method achieves the state-of-the-art performances in terms of efficiency and effectiveness.
Lijuan Duan, Huiling Geng, Junbiao Pang, Qingming Huang
MMAsia4
2019 Accelerating Topic Detection on Web for a Large-Scale Data Set via Stochastic Poisson Deconvolution
Jinzhong Lin, Junbiao Pang, Li Su 0003, Yugui Liu, Qingming Huang
MMM (1)2
2019 Increasing Interpretation of Web Topic Detection via Prototype Learning From Sparse Poisson Deconvolution
abstract
Organizing webpages into interesting topics is one of the key steps to understand the trends from multimodal Web data. The sparse, noisy, and less-constrained user-generated content results in inefficient feature representations. These descriptors unavoidably cause that a detected topic still contains a certain number of the false detected webpages, which further make a topic be less coherent, less interpretable, and less useful. In this paper, we address this problem from a viewpoint interpreting a topic by its prototypes, and present a two-step approach to achieve this goal. Following the detection-by-ranking approach, a sparse Poisson deconvolution is proposed to learn the intratopic similarities between webpages. To find the prototypes, leveraging the intratopic similarities, top- k diverse yet representative prototype webpages are identified from a submodularity function. Experimental results not only show the improved accuracies for the Web topic detection task, but also increase the interpretation of a topic by its prototypes on two public datasets.
Junbiao Pang, Anjing Hu, Qingming Huang, Qi Tian 0001
IEEE Trans. Cybern.1
2019 Learning to Predict Bus Arrival Time From Heterogeneous Measurements via Recurrent Neural Network
abstract
Bus arrival time prediction intends to improve the level of the services provided by transportation agencies. Intuitively, many stochastic factors affect the predictability of the arrival time, e.g., weather and local events. Moreover, the arrival time prediction for a current station is closely correlated with that of multiple passed stations. Motivated by the observations above, this paper proposes to exploit the long-range dependencies among the multiple time steps for bus arrival prediction via recurrent neural network (RNN). Concretely, RNN with long short-term memory block is used to “correct” the prediction for a station by the correlated multiple passed stations. During the correlation among multiple stations, one-hot coding is introduced to fuse heterogeneous information into a unified vector space. Therefore, the proposed framework leverages the dynamic measurements (i.e., historical trajectory data) and the static observations (i.e., statistics of the infrastructure) for bus arrival time prediction. In order to fairly compare with the state-of-the-art methods, to the best of our knowledge, we have released the largest data set for this task. The experimental results demonstrate the superior performances of our approach on this data set.
Junbiao Pang, Haitao Yu 0008, Qingming Huang
IEEE Trans. Intell. Transp. Syst.1
2019 Two Birds With One Stone: A Coupled Poisson Deconvolution for Detecting and Describing Topics From Multimodal Web Data
abstract
Organizing multimodal Web pages into hot topics is the core step to grasp trends on the Web. However, the less-constrained social media generate noisy user-generated content, which makes a detected topic be less coherent and less interpretable. In this paper, we address this problem by proposing a coupled Poisson deconvolution to jointly handle topic detection and topic description. For the topic detection, the interestingness of a topic is estimated from the similarities refined by the description of topics; for the topic description, the interestingness of topics is leveraged to describe topics. Two processes cyclically detect interesting topics and generate the multimodal description of topics. This is the innovation of this paper, which just likes killing two birds with one stone. Experiments not only show the significantly improved accuracies for the topic detection but also demonstrate the interpretable descriptions for the topic description on two public data sets.
Junbiao Pang, Qingming Huang, Qi Tian 0001
IEEE Trans. Neural Networks Learn. Syst.1
2018 A two-step approach to describing web topics via probable keywords and prototype images from background-removed similarities
Junbiao Pang, Liang Li 0003, Qingming Huang, Qi Tian 0001
Neurocomputing1
2018 Discovering Fine-Grained Spatial Pattern From Taxi Trips: Where Point Process Meets Matrix Decomposition and Factorization
abstract
As increasing volumes of urban data are being available, new opportunities arise for data-driven analysis that can lead to improvements in the lives of citizens through evidence-based policies. In particular, taxi trip is an important urban sensor that provides unprecedented insights into many aspects of a city, from economic activity, human mobility to land development. However, analyzing these data presents many challenges, e.g., sparse data for fine-grained patterns, and the regularity submerged by seemingly random data. Inspired by above challenges, we focus on Pick-Up (PU)/Drop-Off (DO) points from taxi trips, and propose a fine-grained approach to unveil a set of low spatio-temporal patterns from the regularity-discovered intensity. The proposed method is conceptually simple yet efficient, by leveraging point process to handle sparsity of points, and by decomposing point intensities into the low-rank regularity and the factorized basis patterns, our approach enables domain experts to discover patterns that are previously unattainable for them, from a case study motivated by traffic engineers.
Junbiao Pang, Zuyun Wang, Haitao Yu 0008, Qingming Huang
IEEE Trans. Intell. Transp. Syst.1
2017 Rotative maximal pattern: A local coloring descriptor for object classification and recognition
Junbiao Pang, Weigang Zhang, Laiyun Qing, Qingming Huang
Inf. Sci.1
2017 Justify role of Similarity Diffusion Process in cross-media topic ranking: an empirical evaluation
Junbiao Pang, Weigang Zhang, Qingming Huang
Multim. Tools Appl.1
2016 Webpage saliency prediction with multi-features fusion
abstract
We proposed a novel model to predict human's visual attention when free-viewing webpages. Compared with natural images, webpages are usually full of salient regions such as logos, text, and faces, while few of them attract human's attention in a short sight. Moreover, webpages perform distinct viewing patterns which are quite different from the natural images. In this paper, we introduced multi-features according to our observation on webpages characters and related eye-tracking data. Further, in order to achieve a flexible adaptation to various types of webpages, we employed a machine-learning framework based on our proposed features. Experimental results demonstrate that our model outperforms other state-of-the-art methods in webpage saliency prediction.
Li Su 0003, Bo Wu 0016, Junbiao Pang, Zhe Wu 0006, Qingming Huang
ICIP4
2016 Accelerate convolutional neural networks for binary classification via cascading cost-sensitive feature
abstract
Convolutional Neural Networks (CNNs) have delivered impressive state-of-the-art performances for many vision tasks, while the computation costs of these networks during test-time are notorious. Empirical results have discovered that CNNs have learned the redundant representations both within and across different layers. When CNNs are applied for binary classification, we investigate a method to exploit this redundancy across layers, and construct a cascade of classifiers which explicitly balances classification accuracy and hierarchical feature extraction costs. Our method cost-sensitively selects feature points across several layers from trained networks and embeds non-expensive yet discriminative features into a cascade. Experiments on binary classification demonstrate that our framework leads to drastic test-time improvements, e.g., possible 47.2x speedup for TRECVID upper body detection, 2.82x speedup for Pascal VOC2007 People detection, 3.72x for INRIA Person detection with less than 0.5% drop in accuracies of the original networks.
Junbiao Pang, Huihuang Lin, Li Su 0003, Chunjie Zhang 0001, Weigang Zhang, Lijuan Duan, Qingming Huang
ICIP1
2016 Robust latent poisson deconvolution from multiple imperfect features for web topic detection
abstract
In web topic detection, detecting “hot” topics from enormous User-Generated Content (UGC) on web data poses two main difficulties that conventional approaches can barely handle: 1) poor feature representations from noisy images and short texts; and 2) uncertain roles of modalities where visual content is either highly or weakly relevant to textual cues due to less-constrained data. In this paper, following the detection by ranking approach, we address the problem by learning a robust shared representation from multiple, noisy and complementary features, and integrating both textual and visual graphs into a k-Nearest Neighbor Similarity Graph (k-N2SG). Then Non-negative Matrix Factorization using Random walk (NMFR) is introduced to generate topic candidates. An efficient fusion of multiple graphs is then done by a Latent Poisson Deconvolution (LPD) which consists of a poisson deconvolution with sparse basis similarities for each edge. Experiments show significantly improved accuracy of the proposed approach in comparison with the state-of-the-art methods on two public data sets.
Junbiao Pang, Chunjie Zhang 0001, Liang Li 0003, Li Su 0003, Weigang Zhang, Qingming Huang, Guiping Su
ICME2
2016 Online web video topic detection and tracking with semi-supervised learning
Guorong Li, Shuqiang Jiang, Weigang Zhang, Junbiao Pang, Qingming Huang
Multim. Syst.4
2016 Robust Latent Poisson Deconvolution From Multiple Features for Web Topic Detection
abstract
Detecting “hot” topics from the enormous usergenerated content (UGC) data on web poses two main difficulties that the conventional approaches can barely handle:1) poor feature representations from noisy images or short texts, and 2) uncertain roles of modalities where the visual content is either highly or weakly relevant to the textual cues due to the less-constrained UGC. In this paper, following the detection-by-ranking approach, we address above challenges by learning a robust latent representation from multiple, noisy and a high probability of the complementary features. Both the textual features and the visual ones are encoded into a k-nearest neighbor hybrid similarity graph (HSG), where nonnegative matrix factorization using random walk is introduced to generate topic candidates. An efficient fusion of multiple HSGs is then done by a latent poisson deconvolution, which consists of a poisson deconvolution with sparse basis similarity for each edge. Experiments show significantly improved accuracy of the proposed approach in comparison with the state-of-the-art methods on two public datasets.
Junbiao Pang, Chunjie Zhang 0001, Weigang Zhang, Qingming Huang
IEEE Trans. Multim.1
2015 Set-label modeling and deep metric learning on person re-identification
Hao Liu 0019, Bingpeng Ma, Junbiao Pang, Chunjie Zhang 0001, Qingming Huang
Neurocomputing4
2015 Online dictionary learning for Local Coordinate Coding with Locality Coding Adaptors
Junbiao Pang, Chunjie Zhang 0001, Weigang Zhang, Laiyun Qing, Qingming Huang
Neurocomputing1
2015 Fusing cross-media for topic detection by dense keyword groups
Weigang Zhang, Tianlong Chen 0003, Guorong Li, Junbiao Pang, Qingming Huang, Wen Gao 0001
Neurocomputing4
2015 Image classification using boosted local features with random orientation and location selection
Chunjie Zhang 0001, Jian Cheng 0001, Yifan Zhang 0001, Jing Liu 0001, Chao Liang 0001, Junbiao Pang, Qingming Huang, Qi Tian 0001
Inf. Sci.6
2015 Local Laplacian Coding From Theoretical Analysis of Local Coding Schemes for Locally Linear Classification
abstract
Local coordinate coding (LCC) is a framework to approximate a Lipschitz smooth function by combining linear functions into a nonlinear one. For locally linear classification, LCC requires a coding scheme that heavily determines the nonlinear approximation ability, posing two main challenges: 1) the locality making faraway anchors have smaller influences on current data and 2) the flexibility balancing well between the reconstruction of current data and the locality. In this paper, we address the problem from the theoretical analysis of the simplest local coding schemes, i.e., local Gaussian coding and local student coding, and propose local Laplacian coding (LPC) to achieve the locality and the flexibility. We apply LPC into locally linear classifiers to solve diverse classification tasks. The comparable or exceeded performances of state-of-the-art methods demonstrate the effectiveness of the proposed method.
Junbiao Pang, Chunjie Zhang 0001, Weigang Zhang, Qingming Huang
IEEE Trans. Cybern.1
2015 Beyond Explicit Codebook Generation: Visual Representation Using Implicitly Transferred Codebooks
abstract
The bag-of-visual-words model plays a very important role for visual applications. Local features are first extracted and then encoded to get the histogram-based image representation. To encode local features, a proper codebook is needed. Usually, the codebook has to be generated for each data set which means the codebook is data set dependent. Besides, the codebook may be biased when we only have a limited number of training images. Moreover, the codebook has to be pre-learned which cannot be updated quickly, especially when applied for online visual applications. To solve the problems mentioned above, in this paper, we propose a novel implicit codebook transfer method for visual representation. Instead of explicitly generating the codebook for the new data set, we try to make use of pre-learned codebooks using non-linear transfer. This is achieved by transferring the pre-learned codebooks with non-linear transformation and use them to reconstruct local features with sparsity constraints. The codebook does not need to be explicitly generated but can be implicitly transferred. In this way, we are able to make use of pre-learned codebooks for new visual applications by implicitly learning the codebook and the corresponding encoding parameters for image representation. We apply the proposed method for image classification and evaluate the performance on several public image data sets. Experimental results demonstrate the effectiveness and efficiency of the proposed method.
Chunjie Zhang 0001, Jian Cheng 0001, Jing Liu 0001, Junbiao Pang, Qingming Huang, Qi Tian 0001
IEEE Trans. Image Process.4
2015 Unsupervised Web Topic Detection Using A Ranked Clustering-Like Pattern Across Similarity Cascades
abstract
Despite the massive growth of social media on the Internet, the process of organizing, understanding, and monitoring user generated content (UGC) has become one of the most pressing problems in today's society. Discovering topics on the web from a huge volume of UGC is one of the promising approaches to achieve this goal. Compared with classical topic detection and tracking in news articles, identifying topics on the web is by no means easy due to the noisy, sparse, and less- constrained data on the Internet. In this paper, we investigate methods from the perspective of similarity diffusion, and propose a clustering-like pattern across similarity cascades (SCs). SCs are a series of subgraphs generated by truncating a similarity graph with a set of thresholds, and then maximal cliques are used to capture topics. Finally, a topic-restricted similarity diffusion process is proposed to efficiently identify real topics from a large number of candidates. Experiments demonstrate that our approach outperforms the state-of-the-art methods on three public data sets.
Junbiao Pang, Fei Jia, Chunjie Zhang 0001, Weigang Zhang, Qingming Huang
IEEE Trans. Multim.1
2014 Nonlinear learning using LCC for online visual tracking
abstract
In this paper, we propose to address online visual tracking on the basis of Local Coordinate Coding (LCC), which integrates the advantages of the discriminative method and the generative method. In the discriminative module, a nonlinear function is trained using the local coordinate codes of image patches to identify the foreground patches from background. In the generative module, we introduce a similarity function that takes the spatial structures of local patches in the target into account between the candidate and holistic templates by reconstruction error. To deal with appearance change during tracking, an online update method is introduced. The proposed tracking method is evaluated on different challenging video sequences with center location error, and experimental results demonstrate the good performance of our method.
Hongwei Hu, Bo Ma 0001, Junbiao Pang
ICME4
2014 Web topic detection using a ranked clustering-like pattern across similarity cascades
abstract
In multi-media and social media communities, web topic detection poses two main difficulties that conventional approaches can barely handle: 1) there are large inter-topic variations among web topics; 2) supervised information is rare to identify the real topics. In this paper, we address these problems from the similarity diffusion perspective among objects on web, and present a clustering-like pattern across similarity cascades (SCs). SCs are a series of subgraphs generated by truncating a weighted graph with a set of thresholds, and then maximal cliques are used to describe the topic candidates. Poisson deconvolution is adopted to efficiently identify the real topics from these topic candidates. Experiments demonstrate that our approach outperforms the state-of-the-arts on two datasets. In addition, we report accuracy v.s. false positives per topic (FPPT) curves for performance evaluation. To our knowledge, this is the first complete evaluation of web topic detection at the topic-wise level, and it establishes a new benchmark for this problem.
Fei Jia, Junbiao Pang, Weigang Zhang, Guorong Li, Chunjie Zhang 0001, Qingming Huang, Yugui Liu
ICME2
2014 Image classification by non-negative sparse coding, correlation constrained low-rank and sparse decomposition
Chunjie Zhang 0001, Jing Liu 0001, Chao Liang 0001, Zhe Xue, Junbiao Pang, Qingming Huang
Comput. Vis. Image Underst.5
2014 Object categorization in sub-semantic space
Chunjie Zhang 0001, Jian Cheng 0001, Jing Liu 0001, Junbiao Pang, Chao Liang 0001, Qingming Huang, Qi Tian 0001
Neurocomputing4
2014 Beyond visual word ambiguity: Weighted local feature encoding with governing region
Chunjie Zhang 0001, Xian Xiao, Junbiao Pang, Chao Liang 0001, Yifan Zhang 0001, Qingming Huang
J. Vis. Commun. Image Represent.3
2014 Undoing the codebook bias by linear transformation with sparsity and F-norm constraints for image classification
Chunjie Zhang 0001, Chao Liang 0001, Junbiao Pang, Yifan Zhang 0001, Jing Liu 0001, Qingming Huang
Pattern Recognit. Lett.3
2013 Stochastic boosting for large-scale image classification
abstract
Boosting has been extensively used in image processing. Many work focuses on the design or the usage of boosting, but training boosting on large-scale datasets tends to be ignored. To handle the large-scale problem, we present stochastic boosting (StocBoost) that relies on stochastic gradient descent (SGD) which uses one sample at each iteration. To understand the efficacy of StocBoost, the convergence of training algorithm is theoretically analyzed. Experimental results show that StocBoost is faster than the batch ones, and is also comparable with the state-of-the-arts.
Junbiao Pang, Qingming Huang, Dan Wang 0019
ICIP1
2013 Undo the codebook bias by linear transformation for visual applications
abstract
The bag of visual words model (BoW) and its variants have demonstrate their effectiveness for visual applications and have been widely used by researchers. The BoW model first extracts local features and generates the corresponding codebook, the elements of a codebook are viewed as visual words. The local features within each image are then encoded to get the final histogram representation. However, the codebook is dataset dependent and has to be generated for each image dataset. This costs a lot of computational time and weakens the generalization power of the BoW model. To solve these problems, in this paper, we propose to undo the dataset bias by codebook linear transformation. To represent every points within the local feature space using Euclidean distance, the number of bases should be no less than the space dimensions. Hence, each codebook can be viewed as a linear transformation of these bases. In this way, we can transform the pre-learned codebooks for a new dataset. However, not all of the visual words are equally important for the new dataset, it would be more effective if we can make some selection using sparsity constraints and choose the most discriminative visual words for transformation. We propose an alternative optimization algorithm to jointly search for the optimal linear transformation matrixes and the encoding parameters. Image classification experimental results on several image datasets show the effectiveness of the proposed method.
Chunjie Zhang 0001, Yifan Zhang 0001, Shuhui Wang, Junbiao Pang, Chao Liang 0001, Qingming Huang, Qi Tian 0001
ACM Multimedia4
2012 Theoretical analysis of learning local anchors for classification
Junbiao Pang, Qingming Huang, Dan Wang 0019
ICPR1
2012 Color Maximal-Dissimilarity Pattern for pedestrian detection
Junbiao Pang, Guoyi Liu, Qingming Huang, Shuqiang Jiang
ICPR2
2012 Online selection of the best k-feature subset for object tracking
Guorong Li, Qingming Huang, Junbiao Pang, Shuqiang Jiang
J. Vis. Commun. Image Represent.3
2011 Treat samples differently: Object tracking with semi-supervised online CovBoost
abstract
Most feature selection methods for object tracking assume that the labeled samples obtained in the next frames follow the similar distribution with the samples in the previous frame. However, this assumption is not true in some scenarios. As a result, the selected features are not suitable for tracking and the “drift” problem happens. In this paper, we consider data's distribution in tracking from a new perspective. We classify the samples into three categories: auxiliary samples (samples in the previous frames), target samples (collected in the current frame) and unlabeled samples (obtained in the next frame). To make the best use of them for tracking, we propose a novel semi-supervised transfer learning approach. Specifically, we assume only target samples follow the same distribution as the unlabeled samples and develop a novel semi-supervised CovBoost method. It could utilize auxiliary samples and unlabeled samples effectively when training the best strong classifier for tracking. Furthermore, we develop a new online updating algorithm for semi-supervised CovBoost, making our tracker handle with significant variations of the tracked target and background successfully. We demonstrate the excellent performance of the proposed tracker on several challenging test videos.
Guorong Li, Qingming Huang, Junbiao Pang, Shuqiang Jiang
ICCV4
2011 Transferring Boosted Detectors Towards Viewpoint and Scene Adaptiveness
abstract
In object detection, disparities in distributions between the training samples and the test ones are often inevitable, resulting in degraded performance for application scenarios. In this paper, we focus on the disparities caused by viewpoint and scene changes and propose an efficient solution to these particular cases by adapting generic detectors, assuming boosting style. A pretrained boosting-style detector encodes a priori knowledge in the form of selected features and weak classifier weighting. Towards adaptiveness, the selected features are shifted to the most discriminative locations and scales to compensate for the possible appearance variations. Moreover, the weighting coefficients are further adapted with covariate boost, which maximally utilizes the related training data to enrich the limited new examples. Extensive experiments validate the proposed adaptation mechanism towards viewpoint and scene adaptiveness and show encouraging improvement on detection accuracy over state-of-the-art methods.
Junbiao Pang, Qingming Huang, Shuicheng Yan, Shuqiang Jiang
IEEE Trans. Image Process.1
2008 Multiple Instance Boost Using Graph Embedding Based Decision Stump for Pedestrian Detection
Junbiao Pang, Qingming Huang, Shuqiang Jiang
ECCV (4)1
2008 Pedestrian detection via logistic multiple instance boosting
abstract
Pedestrian detection in still image should handle the large appearance and pose variations arising from the articulated structure and various clothing of human bodies as well as view points. So it is difficult to design effective classifier for this problem. In this paper, we address these variations in detection via multiple instance learning, specifically logistic multiple instance boosting (LMIB). In LMIB, a example is represented as a set of instances, which implicitly encode the variations. Giving different confidence to the instances in a bag, the LMIB will automatically reduce the influence of the variations at training stage. To obtain rapid detection speed, the LMIBs are grouped into the cascaded structure. The proposed detection algorithm is tested on MIT and INRIA human datasets where promising detection results are comparable with the baseline algorithms.
Junbiao Pang, Qingming Huang, Shuqiang Jiang, Wen Gao 0001
ICIP1
2007 Monocular Tracking 3D People By Gaussian Process Spatio-Temporal Variable Model
abstract
Tracking 3D people from monocular video is often poorly constrained. To mitigate this problem, prior knowledge should be exploited. In this paper, the Gaussian process spatio-temporal variable model (GPSTVM), a novel dynamical system modeling method is proposed for learning human pose and motion priors. The GPSTVM provides a low dimensional embedding of human motion data, with a smooth density function that provides higher probability to the poses and motions close to the training data. The low dimensional latent space is optimized directly to retain the spatio-temporal structure of the high dimensional pose space. After the prior on human pose is learned, the particle filtering can be used tracking articulated human pose; particle filtering propagates over time in the embedding space, avoiding the curse of dimensionality. Experiments demonstrate that our approach tracks 3D people accurately.
Junbiao Pang, Laiyun Qing, Qingming Huang, Shuqiang Jiang, Wen Gao 0001
ICIP (5)1