Qiang Peng

dblp:09/3920 · DBLP profile ↗
← Back
63ranked-venue papers
4as first author
12since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 33 · 4 since 2021Artificial intelligence and machine learning · 16 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 3 since 2021Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3Databases, data management, data science and information retrieval · 2 · 1 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2025 A Multi-Strategy Enhanced Wolf Pack Algorithm for Three-Dimensional Path Planning of Unmanned Aerial Vehicles
abstract
ABSTRACT In this paper, a Multi‐strategy Enhanced Wolf Pack Algorithm (MSEWPA) is proposed to address the three‐dimensional (3D) path planning problem for unmanned aerial vehicles (UAVs) in complex environments. Initially, a mathematical model for 3D path planning is constructed, comprehensively considering constraints such as UAV operational efficiency, path safety risks, performance limitations, obstacle avoidance requirements, and noise limits in urban functional areas. Subsequently, the design of the MSEWPA algorithm is elaborated in detail, including the utilization of the Good Lattice Point (GLP) theory to optimize population initialization for enhanced global search capability, the integration of selection, crossover, and mutation operations from the Differential Evolution (DE) algorithm to augment the randomness of wandering, the introduction of a behavior transition factor for adaptive behavior adjustment, the incorporation of light propagation phenomena to improve random search capabilities during the running process, and the design of multiple siege strategies to guide the exploration of globally optimal solutions. To validate the robustness of the algorithm, sensitivity analysis is conducted on key parameters to determine their optimal settings, and ablation experiments are performed to verify the effectiveness of each improvement strategy. Experimental results on the CEC‐2017 benchmark test functions demonstrate that MSEWPA excels in solving complex optimization problems, achieving rapid convergence to high‐quality global optimal solutions. Furthermore, in four path planning problems of varying complexity, MSEWPA outperforms 11 other state‐of‐the‐art metaheuristic optimization algorithms, demonstrating a strong balance between global and local exploration capabilities. This provides an effective solution for UAV 3D path planning.
Qiang Peng, Renjun Zhan, Husheng Wu, Aiai Wang, Yuanda Lai
Concurr. Comput. Pract. Exp.1
2025 An Air-Ground Unmanned Swarm Collaborative Area Search Strategy Based on the Learning Wolf Pack Algorithm
abstract
Collaborative search by air-ground unmanned swarm, as a pivotal and efficient approach for intelligence gathering and disaster relief, highlights the critical role of search path planning in enhancing overall performance. Addressing the inefficiency resulting from insufficient collaboration between air and ground unmanned platforms in current research, this paper delves into the fundamental characteristics and challenges of collaborative search by air-ground unmanned swarm. This paper clarifies the objectives and constraints of path planning and introduces a method for collaborative search path planning based on the Learning Wolf Pack Algorithm (LWPA). This method initially constructs an optimization model that comprehensively considers area coverage, target detection probability, and search uncertainty. It integrates Distributed Model Predictive Control (DMPC) with the Distributed Constraint Optimization Problem (DCOP) framework, forming an architecture for real-time search path planning. To overcome the limitation of existing DCOP solution methods, which tend to get stuck in local optimal solutions, the LWPA employs a Q-learning mechanism for hierarchical learning and dynamically adjusts parameters to balance local refinement and global exploration. Experimental results demonstrate that this method offers significant advantages in improving search efficiency, coverage, and target detection rates, with an average area coverage of 99.28% and uncertainty as low as 0.86%. These results fully validate its effectiveness and superiority in search tasks in complex urban environments. Furthermore, tests on dynamic adaptability and scalability further verify the potential and value of this method in practical applications.
Qiang Peng, Husheng Wu, Renjun Zhan, Yinan Guo 0001, Jingyi Geng, Feng Wang 0048, Wenxing Fu
IEEE Trans Autom. Sci. Eng.1
2025 HighlightNet: Learning Highlight-Guided Attention Network for Nighttime Vehicle Detection
abstract
Vehicle detection at night is a crucial task in Intelligent Transportation Systems. Due to the complex lighting environment, vehicle detection at night remains a challenging task. Headlights and taillights are essential cues to identify vehicles at night. However, existing methods struggle to effectively utilize the light information of the vehicle. This paper proposes a novel highlight-guided framework to identify vehicles, named HighlightNet, by utilizing both the illumination data from the vehicle lights and the reflective properties of vehicles. The framework combines vehicle detection and highlight area recognition via dual-branch joint learning. To ensure that both branches focus on the highlighted regions, Feature Similarity Awareness Attention (FSAA) is introduced to capture the common attention regions of different branches. Highlight Region Perception (HRP) is proposed to exclude streetlights and other reflective illuminations from the FSAA output, which generates a mask map capable of differentiating the foreground from the background of highlighted areas. It improves the allocation of feature weights and adaptively modifies the distribution within the dual-branch configuration. Furthermore, to address the severe pixel imbalance between the highlighted area and the background, Adaptive Spatial Balance (ASB) loss is introduced to allocate the attention towards prospective vehicle regions while diminishing the emphasis on background regions. Extensive experiments conducted on the BDD100K-Night dataset and a newly acquired dataset specifically designed for nighttime surveillance, called the NightVehicle dataset, demonstrate that HighlightNet outperforms the state-of-the-art methods for nighttime vehicle detection.
Yu-Pei Song, Xiao Wu 0001, Wei Li 0110, Tingquan He, Dongfeng Hu, Qiang Peng
IEEE Trans. Intell. Transp. Syst.6
2025 A distance determination wolf pack algorithm for solving high-dimensional complex functions and its application
Yuanda Lai, Husheng Wu, Qiang Peng, Shi Cheng 0002
J. Supercomput.3
2024 PostureHMR: Posture Transformation for 3D Human Mesh Recovery
abstract
Human Mesh Recovery (HMR) aims to estimate the 3D human body from 2D images, which is a challenging task due to inherent ambiguities in translating 2D observations to 3D space. A novel approach called PostureHMR is pro-posed to leverage a multi-step diffusion-style process, which converts this task into a posture transformation from an SMPL T-pose mesh to the target mesh. To inject the learning process of posture transformation with the physical structure of the human body model, a kinematics-based forward process is proposed to interpolate the intermediate state with pose and shape decomposition. Moreover, a mesh-to-posture (M2P) decoder is designed, by combining the in-put of 3D and 2D mesh constraints estimated from the im-age to model the posture changes in the reverse process. It mitigates the difficulties of posture change learning directly from RGB pixels. To overcome the limitation of pixel-level misalignment of modeling results with the input image, a new trimap-based rendering loss is designed to highlight the areas with poor recognition. Experiments conducted on three widely used datasets demonstrate that the proposed approach outperforms the state-of-the-art methods.
Yu-Pei Song, Xiao Wu 0001, Zhaoquan Yuanl, Jian-Jun Qiao, Qiang Peng
CVPR5
2023 TNSEIR: A SEIR pattern-based embedding approach for temporal network
Lei Wang 0289, Yan Zhu 0007, Qiang Peng
Appl. Intell.3
2022 User-dependent interactive light field video streaming system
Qiang Peng, Eric Wang 0001, Wei Xiang 0001, Xiao Wu 0001
Multim. Tools Appl.2
2022 Correction to: User‑dependent interactive light field video streaming system
Qiang Peng, Eric Wang 0001, Wei Xiang 0001, Xiao Wu 0001
Multim. Tools Appl.2
2022 Learning-based high-efficiency compression framework for light field videos
Wei Xiang 0001, Eric Wang 0001, Qiang Peng, Pan Gao 0001, Xiao Wu 0001
Multim. Tools Appl.4
2022 SWNet: A Deep Learning Based Approach for Splashed Water Detection on Road
abstract
Adverse weather conditions seriously threaten the traffic safety, especially for rainy days with the ponding water on the road surface, which potentially result in vehicle crashes, person injuries and crash fatalities. Automatic splashed water detection based on surveillance videos is an attractive way to effectively prevent the traffic accidents. However, surveillance videos exhibit great variations with lighting changes, illumination conditions and complex backgrounds, which pose great difficulties in automatic recognition. In this paper, a novel deep learning based approach is proposed to detect the splashed water. To the best of our knowledge, this is the first work on this topic based on deep learning. An effective semantic segmentation network, called SWNet, is novelly proposed to extract the potential splashed water regions. An encoder-decoder structure is designed to capture the visual characteristics of splashed water. SWNet achieves high efficiency by reusing pooling indices and adopting the light-weight decoder. With the multi-scale feature fusion structure, SWNet integrates the coarse semantic information and detailed appearance information, which significantly boosts the accuracy and refines the edge segmentation. A weighted cross entropy loss for splashed water is adopted to cope with the unbalanced distribution between splashed water and backgrounds. Moreover, a splashed water attention module is designed to focus on the salient regions of moving vehicles and splashed water, by performing attention mechanism to integrate global contextual information in semantic segmentation. Experiments conducted on a newly collected splashed water dataset demonstrate the effectiveness and efficiency of the proposed approach, which outperforms the state-of-the-art methods.
Jian-Jun Qiao, Xiao Wu 0001, Jun-Yan He, Wei Li 0110, Qiang Peng
IEEE Trans. Intell. Transp. Syst.5
2021 Using deep belief network to demote web spam
Xu Zhuang, Yan Zhu 0007, Qiang Peng, Faisal Khurshid 0001
Future Gener. Comput. Syst.3
2021 Differentially private data publishing for arbitrarily partitioned data
Rong Wang 0006, Benjamin C. M. Fung, Yan Zhu 0007, Qiang Peng
Inf. Sci.4
2020 CODAN: Counting-driven Attention Network for Vehicle Detection in Congested Scenes
abstract
Although recent object detectors have shown excellent performance for vehicle detection, they are incompetent for scenarios with a relatively large number of vehicles. In this paper, we explore the dense vehicle detection given the number of vehicles. Existing crowd counting methods cannot directly applied for dense vehicle detection due to insufficient description of density map, and the lack of effective constraint for mining the spatial awareness of dense vehicles. Inspired by these observations, a conceptually simple yet efficient framework, called CODAN, is proposed for dense vehicle detection. The proposed approach is composed of three major components: (i) an efficient strategy for generating multi-scale density maps (MDM) is designed to represent the vehicle counting, which can capture the global semantics and spatial information of dense vehicles, (ii) a multi-branch attention module (MAM) is proposed to bridging the gap between object counting and vehicle detection framework, (iii) with the well-designed density maps as explicit supervision, an effective counting-awareness loss (C-Loss) is employed to guide the attention learning by building the pixel-level constrain. Extensive experiments conducted on four benchmark datasets demonstrate that the proposed method outperforms the state-of-the-art methods. The impressive results indicate that vehicle detection and counting can be mutually supportive, which is an important and meaningful finding.
Wei Li 0110, Zhenting Wang, Xiao Wu 0001, Ji Zhang 0027, Qiang Peng, Hongliang Li 0001
ACM Multimedia5
2020 Pull Request Prioritization Algorithm based on Acceptance and Response Probability
abstract
Pull requests (PRs) prioritization is one of the main challenges faced by integrators in pull-based development. This is especially true for large open-source projects where hundreds of pull requests are submitted daily. Indeed, managing these pull requests manually consumes time and resources and may lead to delays in the reaction (i.e., acceptance or response) to enhancements or bug fixes suggested in the codebase by contributors. We propose an approach, called AR-Prioritizer (Acceptance and Response based Prioritizer), integrating a PRs prioritization mechanism that considers these two aspects. The results of our study demonstrate that our approach can recommend top@5, top@10, and top@20 most likely to be accepted and responded pull requests with Mean Average Precision of 95.3%, 89.6%, and 79.6% and Average Recall of 40%, 65.7%, and 92.9%. Moreover, AR-Prioritizer has outperformed the baseline models with a statistical significance in prioritizing the most likely to be accepted and responded to PRs.
Muhammad Ilyas Azeem, Qiang Peng, Qing Wang 0001
QRS2
2020 Privacy-preserving high-dimensional data publishing for classification
Rong Wang 0006, Yan Zhu 0007, Chin-Chen Chang 0001, Qiang Peng
Comput. Secur.4
2020 Product image recognition with guidance learning and noisy supervision
Qing Li 0058, Xiaojiang Peng, Liangliang Cao, Wenbin Du, Yu Qiao 0001, Qiang Peng
Comput. Vis. Image Underst.7
2020 Learning fashion compatibility across categories with deep multimodal neural networks
Guang-Lu Sun, Jun-Yan He, Xiao Wu 0001, Bo Zhao 0032, Qiang Peng
Neurocomputing5
2020 Learning label correlations for multi-label image recognition with graph networks
Qing Li 0058, Xiaojiang Peng, Yu Qiao 0001, Qiang Peng
Pattern Recognit. Lett.4
2019 Generative Adversarial Networks Based Error Concealment for Low Resolution Video
abstract
In this paper, a novel deep generative model-based approach for video error concealment is proposed. Our method is comprised of completion network and two critics. The frame completion network is trained to fool the both the local and global critics, which requires completion network to conceal frame distortions with regard to overall consistency as well as in details. Specifically, mask attention convolution layer is proposed, which utilize not only the temporal information of the previous frame, but also the intact pixels of the current distorted frame to mask and re-normalize convolution features. Then, both qualitative and quantitative experiments validate the effectiveness and generality of our approach in advancing the error concealment on low resolution video.
Chongyang Xiang, Chuan Yan, Qiang Peng, Xiao Wu 0001
ICASSP4
2019 BranchGAN: Unsupervised Mutual Image-to-Image Transfer With A Single Encoder and Dual Decoders
abstract
Image-to-image translation is a fundamental task for a wide range of applications, such as image style transfer, video effect generation, cross-domain retrieval, etc. Due to the limited number of labeled data, complex scenes, abstract semantics and various involved domains, image translation remains a challenging task. Compared to the supervised approaches for image translation that need a large collection of paired images for training, the unsupervised methods can significantly reduce the training cost. In this paper, an unsupervised end-to-end generative adversarial network is proposed, namedBranchGAN, for mutual image-to-image transfer between two domains. A structure with one single encoder and dual decoders is novelly proposed to capture the cross-domain distributions and generate the images in both domains. Three factors, that is, pixel-level overall style, region semantics, and domain distinguishability are comprehensively considered to constrain the training process of the proposed model, corresponding toreconstruction loss,encoding loss, andadversarial loss, respectively. Experiments conducted on three benchmark datasets demonstrate the effectiveness of the proposed method that outperforms the unsupervised state-of-the-art approaches and has the competitive performance as the supervised method.
Yi-Fan Zhou, Runhao Jiang, Xiao Wu 0001, Jun-Yan He, Shuang Weng, Qiang Peng
IEEE Trans. Multim.6
2018 A Novel Weighted Boundary Matching Error Concealment Schema for HEVC
abstract
In this paper, a novel weighted boundary matching error concealment schema for HEVC is proposed, which is based on the CU depths and PU partitions in reference frame. Firstly, the information of CU depths in reference frames is used for lost slices. For each LCU in a lost slice, the LCUs surrounding to the co-located LCU are used to calculate summed CU-depth weight, which is used to determine the conceal order of each CU. Then, the co-located partition decision from the reference frame is adopted for PUs in each lost CU. The sequence of PUs to conceal is sorted based on the texture randomness index weight and the PU with the largest weight will be concealed next. Finally, the best estimated motion vector for the lost PU is selected for concealment. The experimental results show that our method achieves higher PSNR gains and has a better visual quality than the state-of-the-art methods.
Chuan Yan, Qiang Peng, Xiao Wu 0001
ICIP4
2018 Learning to Transfer: Generalizable Attribute Learning with Multitask Neural Model Search
abstract
As attribute leaning brings mid-level semantic properties for objects, it can benefit many traditional learning problems in multimedia and computer vision communities. When facing the huge number of attributes, it is extremely challenging to automatically design a generalizable neural network for other attribute learning tasks. Even for a specific attribute domain, the exploration of the neural network architecture is always optimized by a combination of heuristics and grid search, from which there is a large space of possible choices to be searched. In this paper, Generalizable Attribute Learning Model (GALM) is proposed to automatically design the neural networks for generalizable attribute learning. The main novelty of GALM is that it fully exploits the Multi-Task Learning and Reinforcement Learning to speed up the search procedure. With the help of parameter sharing, GALM is able to transfer the pre-searched architecture to different attribute domains. In experiments, we comprehensively evaluate GALM on 251 attributes from three domains: animals, objects, and scenes. Extensive experimental results demonstrate that GALM significantly outperforms the state-of-the-art attribute learning approaches and previous neural architecture search methods on two generalizable attribute learning scenarios.
Zhi-Qi Cheng, Xiao Wu 0001, Siyu Huang, Jun-Xiu Li, Alex Hauptmann 0001, Qiang Peng
ACM Multimedia6
2018 Personalized clothing recommendation combining user social circle and fashion style consistency
Guang-Lu Sun, Zhi-Qi Cheng, Xiao Wu 0001, Qiang Peng
Multim. Tools Appl.4
2018 Hookworm Detection in Wireless Capsule Endoscopy Images With Deep Learning
abstract
As one of the most common human helminths, hookworm is a leading cause of maternal and child morbidity, which seriously threatens human health. Recently, wireless capsule endoscopy (WCE) has been applied to automatic hookworm detection. Unfortunately, it remains a challenging task. In recent years, deep convolutional neural network (CNN) has demonstrated impressive performance in various image and video analysis tasks. In this paper, a novel deep hookworm detection framework is proposed for WCE images, which simultaneously models visual appearances and tubular patterns of hookworms. This is the first deep learning framework specifically designed for hookworm detection in WCE images. Two CNN networks, namely edge extraction network and hookworm classification network, are seamlessly integrated in the proposed framework, which avoid the edge feature caching and speed up the classification. Two edge pooling layers are introduced to integrate the tubular regions induced from edge extraction network and the feature maps from hookworm classification network, leading to enhanced feature maps emphasizing the tubular regions. Experiments have been conducted on one of the largest WCE datasets with WCE images, which demonstrate the effectiveness of the proposed hookworm detection framework. It significantly outperforms the state-of-the-art approaches. The high sensitivity and accuracy of the proposed method in detecting hookworms shows its potential for clinical application.
Jun-Yan He, Xiao Wu 0001, Yu-Gang Jiang 0001, Qiang Peng, Ramesh Jain 0001
IEEE Trans. Image Process.4
2017 Sketch Recognition with Deep Visual-Sequential Fusion Model
abstract
In this paper, a deep end-to-end network for sketch recognition, named Deep Visual-Sequential Fusion model (DVSF) is proposed to model the visual and sequential patterns of the strokes. To capture the intermediate states of sketches, a three-way representation learner is first utilized to extract the visual features. These deep features are simultaneously fed into the visual and sequential networks to capture spatial and temporal properties, respectively. More specifically, visual networks are novelly proposed to learn the stroke patterns by stacking the Residual Fully-Connected (R-FC) layers, which integrate ReLU and Tanh activation functions to achieve the sparsity and generalization ability. To learn the patterns of stroke order, sequential networks are constructed by Residual Long Short-Term Memory (R-LSTM) units, which optimize the network architecture by skip connection. Finally, the visual and sequential representations of the sketches are seamlessly integrated with a fusion layer to obtain the final results. Experiments conducted on the benchmark sketch dataset TU-Berlin demonstrate the effectiveness of the proposed method, which outperforms the state-of-the-art approaches.
Jun-Yan He, Xiao Wu 0001, Yu-Gang Jiang 0001, Bo Zhao 0032, Qiang Peng
ACM Multimedia5
2017 Feature bundling in decision tree algorithm
abstract
In empirical data modelling, a model of system is built up from a set of cases that the system has observed. Eventually, the performance of the inducted model is dominated by the quality and quantity of observations. Feature transformation methods are widely used to improve quality of knowledge ext racted from observations to build up more accurate and robust model. In the paper, a new feature transformation method named dynamical feature bundling for decision tree algorithm is proposed. Dynamical feature bundling groups a set of features in the tree induction phase and it enables decision tree algorithms to 1) make use of features in one bundle together to make collective judgments in splitting phase; 2) learn more reliable and stable knowledge from feature bundles created based on domain knowledge of experts; 3) embed feature transformation step into tree induction phase, and therefore the extra pre-process step which are necessary for static feature transformation methods is inessential. Our experiments show 2%-9% improvements of AUC value on a very imbalanced dataset. Slight improvements are also obtained on a more balanced data set.
Xu Zhuang, Yan Zhu 0007, Chin-Chen Chang 0001, Qiang Peng
Intell. Data Anal.4
2017 Automatic content understanding with cascaded spatial-temporal deep framework for capsule endoscopy videos
Honghan Chen, Xiao Wu 0001, Tao Gan, Qiang Peng
Neurocomputing4
2017 A unified score propagation model for web spam demotion algorithm
Xu Zhuang, Yan Zhu 0007, Chin-Chen Chang 0001, Qiang Peng, Faisal Khurshid 0001
Inf. Retr. J.4
2017 Perception-based adaptive quantization for transform-domain Wyner-Ziv video coding
Lei Zhang 0006, Qiang Peng, Xiao Wu 0001
Multim. Tools Appl.2
2017 Analysis of Packet-Loss-Induced Distortion in View Synthesis Prediction-Based 3D Video Coding
abstract
View synthesis prediction (VSP) is a crucial coding tool for improving compression efficiency in the next generation 3D video systems. However, VSP is susceptible to catastrophic error propagation when multi-view video plus depth (MVD) data are transmitted over lossy networks. This paper aims at accurately modeling the transmission errors propagated in the inter-view direction caused by VSP. Toward this end, we first study how channel errors gradually propagate along the VSP-based inter-view prediction path. Then, a new recursive model is formulated to estimate the expected end-to-end distortion caused by those channel losses. For the proposed model, the compound impact of the transmission distortions of both the texture video and depth map on the quality of the synthetic reference view is mathematically analyzed. Especially, the expected view synthesis distortion due to depth errors is characterized in the frequency domain using a new approach, which combines the energy densities of the reconstructed texture image and the channel errors. The proposed model also explicitly considers the disparity rounding operation invoked for the sub-pixel precision rendering of the synthesized reference view. Experimental results are presented to demonstrate that the proposed analytic model is capable of effectively modeling the channel-induced distortion for MVD-based 3D video transmission.
Pan Gao 0001, Qiang Peng, Wei Xiang 0001
IEEE Trans. Image Process.2
2017 Diversified Visual Attention Networks for Fine-Grained Object Classification
abstract
Fine-grained object classification attracts increasing attention in multimedia applications. However, it is a quite challenging problem due to the subtle interclass difference and large intraclass variation. Recently, visual attention models have been applied to automatically localize the discriminative regions of an image for better capturing critical difference, which have demonstrated promising performance. Unfortunately, without consideration of the diversity in attention process, most of existing attention models perform poorly in classifying fine-grained objects. In this paper, we propose a diversified visual attention network (DVAN) to address the problem of fine-grained object classification, which substantially relieves the dependency on strongly supervised information for learning to localize discriminative regions compared with attention-less models. More importantly, DVAN explicitly pursues the diversity of attention and is able to gather discriminative information to the maximal extent. Multiple attention canvases are generated to extract convolutional features for attention. An LSTM recurrent unit is employed to learn the attentiveness and discrimination of attention canvases. The proposed DVAN has the ability to attend the object from coarse to fine granularity, and a dynamic internal representation for classification is built up by incrementally combining the information from different locations and scales of the image. Extensive experiments conducted on CUB-2011, Stanford Dogs, and Stanford Cars datasets have demonstrated that the pro-posed DVAN achieves competitive performance compared to the state-of-the-art approaches, without using any prior knowledge, user interaction, or external resource in training and testing.
Bo Zhao 0032, Xiao Wu 0001, Jiashi Feng, Qiang Peng, Shuicheng Yan
IEEE Trans. Multim.4
2016 A low-complexity error concealment algorithm for video transmission based on non-local means denoising
abstract
In video communication, the concealment of transmission errors is important to mitigate error propagation and allow for a pleasant visual quality. The previous spatio-temporal error concealment algorithm, known as Denoised Temporal Extrapolation Refinement (DTER) used a spatial denoising algorithm to reduce the imperfectness of the temporal extrapolation. However, the computational complexity of DTER is extremely high. In this paper, a fast algorithm based on Non-Local Means (NLM) is proposed for video transmission. The proposed fast strategy mainly includes two steps. First, the search window can be narrowed down by using the characteristic that the similarity of any two image blocks will decrease with increasing Euclidean distance. Afterwards, a successive elimination method is used for cutting out some redundant points. By making use of these two steps, our proposed algorithm achieves 10~30 times faster than DTER without sacrifice in picture quality. Moreover, simulation results clearly show the proposed fast algorithm outperform other existing error concealment algorithms.
Qiang Peng
VCIP2
2016 Web video categorization using category-predictive classifiers and category-specific concept classifiers
Mehtab Afzal, Xiao Wu 0001, Honghan Chen, Yu-Gang Jiang 0001, Qiang Peng
Neurocomputing5
2016 Part-based clothing image annotation by visual neighbor retrieval
Guang-Lu Sun, Xiao Wu 0001, Qiang Peng
Neurocomputing3
2016 Detection of bird nests in overhead catenary system images for high-speed rail
Xiao Wu 0001, Ping Yuan, Qiang Peng, Chong-Wah Ngo, Jun-Yan He
Pattern Recognit.3
2016 Near-Duplicate Segments based news web video event mining
Chengde Zhang, Dianting Liu, Xiao Wu 0001, Guiru Zhao, Mei-Ling Shyu, Qiang Peng
Signal Process.6
2016 Integration of Visual Temporal Information and Textual Distribution Information for News Web Video Event Mining
abstract
News web videos exhibit several characteristics, including a limited number of features, noisy text information, and error in near-duplicate keyframes (NDK) detection. Such characteristics have made the mining of the events from news web videos a challenging task. In this paper, a novel framework is proposed to better group the associated web videos to events. First, the data preprocessing stage performs feature selection and tag relevance learning. Next, multiple correspondence analysis is applied to explore the correlations between terms and events with the assistance of visual information. Cooccurrence and visual near-duplicate feature trajectory induced from NDKs are combined to calculate the similarity between NDKs and events. Finally, a probabilistic model is proposed for news web video event mining, where both visual temporal information and textual distribution information are integrated. Experiments on the news web videos from YouTube demonstrate that the integration of visual temporal information and textual distribution information outperforms the existing methods in the news web video event mining.
Chengde Zhang, Xiao Wu 0001, Mei-Ling Shyu, Qiang Peng
IEEE Trans. Hum. Mach. Syst.4
2016 Automatic Hookworm Detection in Wireless Capsule Endoscopy Images
abstract
Wireless capsule endoscopy (WCE) has become a widely used diagnostic technique to examine inflammatory bowel diseases and disorders. As one of the most common human helminths, hookworm is a kind of small tubular structure with grayish white or pinkish semi-transparent body, which is with a number of 600 million people infection around the world. Automatic hookworm detection is a challenging task due to poor quality of images, presence of extraneous matters, complex structure of gastrointestinal, and diverse appearances in terms of color and texture. This is the first few works to comprehensively explore the automatic hookworm detection for WCE images. To capture the properties of hookworms, the multi scale dual matched filter is first applied to detect the location of tubular structure. Piecewise parallel region detection method is then proposed to identify the potential regions having hookworm bodies. To discriminate the unique visual features for different components of gastrointestinal, the histogram of average intensity is proposed to represent their properties. In order to deal with the problem of imbalance data, Rusboost is deployed to classify WCE images. Experiments on a diverse and large scale dataset with 440 K WCE images demonstrate that the proposed approach achieves a promising performance and outperforms the state-of-the-art methods. Moreover, the high sensitivity in detecting hookworms indicates the potential of our approach for future clinical application.
Xiao Wu 0001, Honghan Chen, Tao Gan, Junzhou Chen 0001, Chong-Wah Ngo, Qiang Peng
IEEE Trans. Medical Imaging6
2016 Clothing Cosegmentation for Shopping Images With Cluttered Background
abstract
In this paper, we address an important and practical problem of clothing cosegmentation (CCS): given multiple fashion model photos with natural backgrounds on e-commerce websites, to automatically and simultaneously segment all images and extract the clothing regions. However, cluttered backgrounds, variations in colors and styles, and inconsistent human poses all make it a challenging task. In this paper, a novel CCS algorithm is proposed to improve the accuracy of clothing extraction by exploiting the properties of multiple clothing images with the same apparel. First, the co-salient objects are computed by detecting the upper bodies of fashion models and transferring their locations within multiple images. Based on the coarse clothing regions determined by the upper body localization and co-salient object detection, the foreground (clothing) and background Gaussian mixture models are estimated, respectively. Finally, the clothing region in each image is extracted through energy minimization based on graph cuts iteratively. The proposed cosegmentation algorithm is mainly designed for multiple clothing images. As a byproduct, it can also be applied to single image segmentation without any modification. The experiments demonstrate that the proposed approach outperforms the state-of-the-art cosegmentation methods as well as traditional single image segmentation solution for shopping images.
Bo Zhao 0032, Xiao Wu 0001, Qiang Peng, Shuicheng Yan
IEEE Trans. Multim.3
2015 Error-resilient multi-view video coding using Wyner-Ziv techniques
Pan Gao 0001, Qiang Peng, Wei Xiang 0001
Multim. Tools Appl.2
2015 CSIFT based locality-constrained linear coding for image classification
Junzhou Chen 0001, Qing Li 0058, Qiang Peng, Kin Hong Wong
Pattern Anal. Appl.3
2015 Subjective and Objective Video Quality Assessment of 3D Synthesized Views With Texture/Depth Compression Distortion
abstract
The quality assessment for synthesized video with texture/depth compression distortion is important for the design, optimization, and evaluation of the multi-view video plus depth (MVD)-based 3D video system. In this paper, the subjective and objective studies for synthesized view assessment are both conducted. First, a synthesized video quality database with texture/depth compression distortion is presented with subjective scores given by 56 subjects. The 140 videos are synthesized from ten MVD sequences with different texture/depth quantization combinations. Second, a full reference objective video quality assessment (VQA) method is proposed concerning about the annoying temporal flicker distortion and the change of spatio-temporal activity in the synthesized video. The proposed VQA algorithm has a good performance evaluated on the entire synthesized video quality database, and is particularly prominent on the subsets which have significant temporal flicker distortion induced by depth compression and view synthesis process.
Xiangkai Liu, Yun Zhang 0002, Sudeng Hu, Sam Kwong, C.-C. Jay Kuo, Qiang Peng
IEEE Trans. Image Process.6
2014 Boosting VLAD with Supervised Dictionary Learning and High-Order Statistics
Xiaojiang Peng, Limin Wang 0002, Yu Qiao 0001, Qiang Peng
ECCV (3)4
2014 Action Recognition with Stacked Fisher Vectors
Xiaojiang Peng, Changqing Zou, Yu Qiao 0001, Qiang Peng
ECCV (5)4
2014 A Joint Evaluation of Dictionary Learning and Feature Encoding for Action Recognition
abstract
Many mid-level representations have been developed to replace traditional bag-of-words model (VQ+k-means) such as sparse coding, OMP-k with k-SVD, and fisher vector with GMM in image domain. These approaches can be split into a dictionary learning phase and a feature encoding phase which are often closely related. In this paper, we jointly evaluate the effect of these two phases for video-based action recognition. Specially, we compare several dictionary learning methods and feature encoding schemes through extensive experiments on the KTH and HMDB51 datasets. Experimental results indicate that fisher vector performs consistently better than the other encoding methods, and sparse coding is robust to different dictionaries even random weights. In addition, we observe that the advantages of sophisticated mid-level representations do not come from their specific dictionaries but the encoding mechanisms, and we can just use randomly selected exemplars as dictionaries for most of encoding methods. Finally, we achieve the state-of-the-art results on the HMDB51 and UCF101 by combining our configurations with improved dense trajectory features.
Xiaojiang Peng, Limin Wang 0002, Yu Qiao 0001, Qiang Peng
ICPR4
2014 Clothing Extraction Using Region-Based Segmentation and Pixel-Level Refinement
abstract
In this paper, we demonstrate an effective method for automatic extracting clothing object from fashion photographs, an extremely challenging problem due to the non-uniform natural backgrounds, various types of apparel and different poses of human models. This method consists of three phases: (1) coarse clothing area localization by pose estimation and super pixel segmentation, (2) region-level image segmentation, (3) pixel-level refinement using spatial information and Grab cut. Experiments on a dataset with 1000 images crawled from Taobao demonstrate that the proposed method outperforms other methods, which can extract clothing from images with complex background.
Zhao-Rui Liu, Xiao Wu 0001, Bo Zhao 0032, Qiang Peng
ISM4
2014 Motion boundary based sampling and 3D co-occurrence descriptors for action recognition
Xiaojiang Peng, Yu Qiao 0001, Qiang Peng
Image Vis. Comput.3
2014 Large Margin Dimensionality Reduction for Action Similarity Labeling
abstract
Action recognition in videos is receiving extensive research interest due to its wide applications. This task needs to assign a specific action class for each video. In this paper, we study the problem of action similarity labeling (ASLAN) that is to verify whether two action videos present the same type of action or not. We show that both Fisher vector (FV) and vector of locally aggregated descriptors (VLAD) with dense trajectory features can achieve state-of-the-art performance on the ASLAN benchmark. Our main contribution is to develop a large margin dimensionality reduction (LMDR) method to compress high-dimensional FV and VLAD. Specially, we leverage the hinge loss objective function and stochastic gradient descent to optimize the discriminative projection matrix of these vectors. Extensive experiments on the ASLAN dataset indicate that our LMDR method not only reduces the dimension significantly but also improves the verification performance.
Xiaojiang Peng, Yu Qiao 0001, Qiang Peng, Qiong-Hua Wang
IEEE Signal Process. Lett.3
2013 Exploring Motion Boundary based Sampling and Spatial-Temporal Context Descriptors for Action Recognition
abstract
Feature representation is important for human action recognition.Recently, Wang et al. [25] proposed dense trajectory (DT) based features for action video representation and achieved state-of-the-art performance on several action datasets.In this paper, we improve the DT method in two folds.Firstly, we introduce a motion boundary based dense sampling strategy, which greatly reduces the number of valid trajectories while preserves the discriminative power.Secondly, we develop a set of new descriptors which describe the spatial-temporal context of motion trajectories.To evaluate the performance of the proposed methods, we conduct extensive experiments on three benchmarks including K-TH, YouTube and HMDB51.The results show that our sampling strategy significantly reduces the computational cost of point tracking without degrading performance.Meanwhile, we achieve superior performance than the state-of-the-art methods by utilizing our spatial-temporal context descriptors.
Xiaojiang Peng, Yu Qiao 0001, Qiang Peng, Xianbiao Qi
BMVC3
2013 An Error Resilient Depth Map Coding Scheme Using Adaptive Wyner-Ziv Frame
Xiangkai Liu, Qiang Peng, Xiao Wu 0001, Lei Zhang 0006, Ling-Yu Duan
MMM (2)2
2013 Clothing Extraction by Coarse Region Localization and Fine Foreground/Background Estimation
Xiao Wu 0001, Bo Zhao 0032, Ling-Ling Liang, Qiang Peng
MMM (2)4
2013 SSIM-Based End-to-End Distortion Model for Error Resilient Video Coding over Packet-Switched Networks
Lei Zhang 0006, Qiang Peng, Xiao Wu 0001
MMM (1)2
2013 A Novel Web Video Event Mining Framework with the Integration of Correlation and Co-Occurrence Information
Chengde Zhang, Xiao Wu 0001, Mei-Ling Shyu, Qiang Peng
J. Comput. Sci. Technol.4
2012 A top-down design process oriented adaptive conceptual layout design environment based on semantic norm model
abstract
Product design is a complicated and creative process. Even for the up-to-date commercial 3D CAD tools, it is very difficult to achieve automation of the entire design process. In the article, a semantic norm model (SNM) is presented to support the top-down process in conceptual design, which is an important phase in product design and has a decisive influence on the cost of the product design. The SNM system can define virtual components in early design stage with semantics and instantiate those components in detailed design, and it bridges the gaps between conceptual design and detailed design. The SNM system is also managed by 3D constraints system to support the variational design of the product. A combined algebraic together with numerical method was used to solve SNM constraints incrementally. Based on the SNM system, an adaptive conceptual design environment is developed and the top-down oriented product design process is demonstrated.
Jianjie Wu, Xiaobing Pei, Qiang Peng
CSCWD5
2012 Source Distortion Temporal Propagation Model for Motion Compensated Video Coding Optimization
abstract
Rate-distortion optimization (RDO) is widely employed to maximize coding efficiency in hybrid video coding. Due to an extensive use of spatial-temporal predictions in video coding, a global RDO problem becomes so complex that the processing of each coding unit (e.g. a macro block) is dependent and entangled each other. Usually the RDO is simply performed for each coding unit individually and independently, thus compromising the RDO performance significantly. In this paper, we examine temporal dependent RDO by developing a novel source distortion temporal propagation (SDTP) model, which takes into account the influence of a current coding macro block to future macro blocks in its temporal propagation chain. Accordingly estimations of the influence and the corresponding Lagrange multiplier for the proposed SDTP-based RDO are obtained. Experimental results show consistently a significant coding gain of the proposed scheme over H.264/AVC.
Tianwu Yang, Ce Zhu, Xiaojiu Fan, Qiang Peng
ICME4
2012 Boosting web video categorization with contextual information from social web
Xiao Wu 0001, Chong-Wah Ngo, Yi-Ming Zhu, Qiang Peng
World Wide Web4
2011 Face Region Based Conversational Video Coding
abstract
Face regions are visual focuses in conversational video communications, thus better reconstruction quality of the regions of interest (ROI) is highly desired or necessary in the bandwidth-constrained conversational video coding. In this paper, we introduce an efficient motion based face detection method to identify face blocks in the first step, which can reduce computational complexity substantially without any loss in face detection results. Then an active contour model is applied to find face contours for more refined and compact face regions. Based on the well-located and compact face regions, facial feature priority based bit allocation is proposed for face ROI based conversational video coding. Experimental results demonstrate that the proposed face region based coding can considerably improve the coding results in the face regions, compared with two other relevant video coding schemes, in terms of objective rate-distortion performance as well as subjective visual quality.
Bing Xiong 0005, Xiaojiu Fan, Ce Zhu, Xuan Jing, Qiang Peng
IEEE Trans. Circuits Syst. Video Technol.5
2010 3D Automatic Feature Construction System for Lower Limb Alignment
abstract
Accurate measurement of lower limb alignment is vital for orthopedic surgery planning. Based on the 3D lower limb bone models reconstructed from CT images, a 3D feature construction system is developed to find the lower limb alignment. This paper presents new methods to automatically identify the anatomic axes, mechanical axes, joint center points, and reference planes of the lower extremities. As far as we know, there is no other system that can construct all these features automatically. With the above identified features, our system can calculate the related angles of the lower limb alignment accurately for orthopedic knee surgery planning. We applied our system on two groups of specimen (with 6 pairs of lower limbs), and the resulting related angles are identical to the actual angles of the lower limb alignment.
Qi Xing, Wenzhen Yang, Mark M. Theiss, Jihui Li, Qiang Peng, Jim X. Chen
CW5
2007 A Fast Transform Domain Based Algorithm for H.264/AVC Intra Prediction
abstract
Directional intra prediction is one of the new features adopted in H.264/AVC standard. In contrast to some previous coding standards, the prediction is performed in the spatial domain. Two types of intra prediction for luma - Intra16times16 and Intra4times4 are included in each profile. For high profile, there is an additional type - Intra8times8. Intra 16x16 and Intra4times4 support four and nine modes respectively. As a result, the encoder complexity is increased dramatically. In this paper we proposed a fast intra prediction algorithm using transform domain features of the target block to filter out the majority of candidate modes. Extensive simulations verify that the proposed method speeds up the intra prediction process by 67.5% on average without sacrificing the picture quality and compression ratio. A comparison between the proposed algorithm and others is also provided.
Zhengning Wang, Jun Yang 0005, Qiang Peng, Changqian Zhu
ICME3
2005 An adaptive key-frame reference picture selection algorithm for video transmission via error prone networks
abstract
When compressed video is transmitted via error prone network, it can suffer from severe degradation. To combat the effect of transmission errors, a novel adaptive key-frame reference picture selection (AKRPS) algorithm is proposed. In this algorithm, key-frames are predefined according to the round-trip delay, and the latest key-frame received ACK signal is used as the reference picture of key-frame while the other frames is encoded normally, which can prevent the temporal error propagation efficiently with keeping high coding efficiency. The required reference picture memory for AKRPS is less than 3 frames independent of the round-trip delay. A frame-layer rate control scheme is also proposed for smoothing the video quality. Simulation results, with different video sequence and H.263+ codec, show that the KRPS can achieve high error resilience at packet error rates (PER) ranging from 1% to 10% in both single-user and multi-user applications.
Tianwu Yang, Qiang Peng, Changqian Zhu
ISADS3
2005 Residual Texture Based Fast Block-Size Selection for Inter-Frame Coding in H.264/AVC
abstract
One of the new features adopted in H.264/AVC is the utilization of flexible block size ranging from 16x16 to 4x4 in inter-frame coding. The aim is to reduce the error due to fixed block size prediction within a macroblock. However, this feature requires extremely computational complexity. In this paper, we proposed a residual texture based fast block size selection algorithm for inter-frame coding. Firstly, we perform a motion estimation (ME) for a macroblock and get the residual; then we predict the block size by analyzing the residual texture. Extensive simulations verify that the proposed method speeds up the block-size selection procedure by 50% without sacrificing picture quality and compression ratio.
Zhengning Wang, Qiang Peng
PDCAT2
2004 System-on-chip design for TV-centric home networks
abstract
We design a low cost, high performance system-on-chip with video/audio/data processing engines, open standard 32-bit RISC microprocessor, and networking processor unit. We use system-level modeling technology to design and verify H.264, MPEG-4 AAC, and MP3 codecs. We adopt a transaction-level verification technique for self-checking and regression testing of the design. The SoC incorporates peripherals, such as UART, USB, 100/10 Mbps Ethernet, IEEE 1394, ATAPI, Utopia, WiFi, and SVGA. We used real-time Linux as the software. The design can be used to build an inexpensive "TV-centric" multimedia service-center for the home; it provides functions for digital TV, DVD recording, video-on-demand, interactive games, set top box, home networking, and network computer terminal.
Qiang Peng, Jin Jing
CCNC1
2000 MUST: multiple-stem analysis for identifying sequentially untestable faults
abstract
In this paper we present MUST-a multiple-stem analysis algorithm for identifying untestable faults in sequential circuits. In general, processing untestable faults is the most time-consuming part of a sequential ATPG. MUST extends the scope of the single-stem analysis done in the FIRES algorithm by identifying additional untestable faults that cannot be found by single-stem analysis. While its computational requirements are greater than those of FIRES, the run-time of MUST remains significantly lower than that used by sequential ATPG. We show that the faults identified by MUST are difficult targets for conventional ATPG programs, that can benefit by using MUST as a preprocessor and excluding the untestable faults identified by multiple stem analysis from the target faults processed by ATPG. We report experimental results obtained by our prototype implementation of MUST on ISCAS benchmarks and other circuits.
Qiang Peng, Miron Abramovici, Jacob Savir
ITC1