EDBT 2026 Demo / reviewers in the wild / expert
Zengfu Wang
dblp:22/5156
· DBLP profile ↗
159ranked-venue papers
2as first author
43since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 101 · 25 since 2021Artificial intelligence and machine learning · 48 · 12 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 4 since 2021Databases, data management, data science and information retrieval · 10 · 1 first-author · 1 since 2021Computer networks · 7 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Behavior-aware group target tracking for AAV swarms with dynamic structural evolution
Hua Lan, Yuxiang Mao, Xiwei Lan, Xiaolei Hou, Zengfu Wang, Zhunga Liu |
Signal Process. | 6 |
| 2026 | KalmanTNet - hybrid-driven Kalman filter for trajectory filtering
Zengfu Wang, Kuilong Yang, Hua Lan |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | Exploring a novel data- and parameter-efficient fine-tuning method for robust face recognition
Yin Lin, Qidong Huang, Yang Cao 0010, Zengfu Wang |
Pattern Recognit. | 8 |
| 2026 | Improving Chinese Text Recognition with Multi-Granularity Features and Vision-Language ReasoningabstractAccurate Chinese text recognition (CTR) is vital for applications such as document digitization, but remains challenging due to high inter-class visual similarity, complex hierarchical structures, and diverse visual degradations in real-world scenes. To address these challenges, we propose a robust CTR framework that synergizes structure-aware visual discrimination with cross-modal reasoning. The hybrid encoder integrates multi-scale attention modulation and global self-attention, guided by hierarchical structural supervision, to capture fine-grained structure-aware visual representations essential for distinguishing visually similar characters. Complementing this, an iterative vision-language decoder, trained with a stochastic masking strategy, learns to reconstruct text from partially observed visual and contextual cues, enabling complementary vision-language reasoning that effectively resolves visually ambiguous or degraded characters. Extensive experiments on public benchmarks demonstrate that our method achieves state-of-the-art performance, validating its effectiveness in addressing the challenges of Chinese text recognition. Zengfu Wang |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2026 | Network Slicing in MEC-Based RANs With Nonlinear Cost Rate FunctionsabstractThis paper addresses network slicing in a large-scale Multi-Access Edge Computing (MEC)-enabled Radio Access Network (RAN) comprising heterogeneous edge nodes with varying computing and storage resource capacities. These resources are dynamically allocated to slice requests and released when the service of a slice request is completed. Our objective is to optimize the resource allocation for each admitted arriving slice request, considering its demands for computing and storage resources, to maximize the long-run average Earning Before Interest and Taxes (EBIT) of the MEC slicing system. We formulate the optimization problem as a Restless Multi-Armed Bandit (RMAB)-based resource allocation problem with a nonlinear cost rate function. To solve this, we introduce a new policy called Prioritizing-the-Future-Approximated earning per request (PFA) where for each admitted slice request, we always prioritize the allocation of the resource combination that gives the highest achievable earning, considering the future effects of this allocation. PFA is designed to be scalable and applicable to large-scale networks. We numerically demonstrate the superior performance of PFA in maximizing long-run average EBIT through simulations, comparing it with two baseline policies, at various cases of parameter values. Moreover, our findings offer insights for network operators in resource allocation policy selection. Jiahe Xu 0004, Jing Fu 0001, Bige Yang, Zengfu Wang, Jingjin Wu, Xinyu Wang 0011, Moshe Zukerman |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2025 | Efficient Fine-tuning Strategies for Enhancing Face Recognition Performance in Challenging ScenariosabstractFace recognition plays a crucial role in human life, prompting numerous excellent research efforts. However, face recognition in real-world applications presents various scenarios such as occluded, overexposed and near-infrared face recognition. Due to domain discrepancy and a lack of large-scale training data, effectively transferring pre-trained face recognition models to these scenarios has become a challenge. Recently, Parameter-Efficient Fine-Tuning (PEFT) methods have shown great potential in natural language processing tasks, but their effectiveness in computer vision tasks, especially in face recognition tasks, remains under-explored. In this paper, we propose a Data-Parameter-Efficient Fine-Tuning (DPEFT) approach for the face recognition tasks, encompassing two kinds of fine-tuning strategies. With these strategies, the DPEFT method requires only an additional 2.7% learnable parameters and 20% of the training data during the training phase to achieve competitive results. Moreover, by further integrating the concept of structural re-parameterization, our approach maintains the same model architecture and parameters as the pre-trained model during inference. Extensive experimental results on both holistic and occluded face datasets demonstrate that our approach achieves performance comparable to or better than the fully fine-tuning methods, and significantly lower training costs. Our DPEFT enables the pre-trained face recognition model to adapt efficiently and effectively to a variety of scenarios, indicating its potential in practical applications. Yin Lin, Ziyang Wu, Qidong Huang, Jinshui Hu, Zengfu Wang |
ICASSP | 8 |
| 2025 | Exploring Part-Informed Visual-Language Learning for Person Re-IdentificationabstractRecently, visual-language learning (VLL) has shown great potential in enhancing visual-based person re-identification (ReID). Existing VLL-based ReID methods typically focus on image-text feature alignment at the whole-body level, while neglecting supervision on fine-grained part features, thus lacking constraints for local feature semantic consistency. To this end, we propose Part-Informed Visual-language Learning (π-VL) to enhance fine-grained visual features with part-informed language supervisions for ReID tasks. Specifically, π-VL introduces a human parsing-guided prompt tuning strategy and a hierarchical visual-language alignment paradigm to ensure within-part feature semantic consistency. The former combines both identity labels and human parsing maps to constitute pixel-level text prompts, and the latter fuses multi-scale visual features with a light-weight auxiliary head to perform fine-grained image-text alignment. As a plug-and-play and inference-free solution, our π-VL achieves performance comparable to or better than state-of-the-art methods on four commonly used ReID benchmarks. Notably, it reports 91.0% Rank-1 and 76.9% mAP on the challenging MSMT17 database, without bells and whistles. Yin Lin, Yehansen Chen, Jinshui Hu, Cong Liu 0006, Zengfu Wang |
ICME | 7 |
| 2025 | Unleashing Foundation Vision Models: Adaptive Transfer for Diverse Data-Limited Scientific DomainsabstractIn the big data era, the computer vision field benefits from large-scale datasets such as LAION-2B, LAION-400M, and ImageNet-21K, Kinetics, on which popular models like the ViT and ConvNeXt series have been pre-trained, acquiring substantial knowledge.
However, numerous downstream tasks in specialized and data-limited scientific domains continue to pose significant challenges.
In this paper, we propose a novel Cluster Attention Adapter (CLAdapter), which refines and adapts the rich representations learned from large-scale data to various data-limited downstream tasks. Specifically, CLAdapter introduces attention mechanisms and cluster centers to personalize the enhancement of transformed features through distribution correlation and transformation matrices. This enables models fine-tuned with CLAdapter to learn distinct representations tailored to different feature sets, facilitating the models' adaptation from rich pre-trained features to various downstream scenarios effectively. In addition, CLAdapter's unified interface design allows for seamless integration with multiple model architectures, including CNNs and Transformers, in both 2D and 3D contexts.
Through extensive experiments on 10 datasets spanning domains such as generic, multimedia, biological, medical, industrial, agricultural, environmental, geographical, materials science, out-of-distribution (OOD), and 3D analysis, CLAdapter achieves state-of-the-art performance across diverse data-limited scientific domains, demonstrating its effectiveness in unleashing the potential of foundation vision models via adaptive transfer.
Code is available at https://github.com/qklee-lz/CLAdapter. Qiankun Li 0004, Feng He 0008, Huabao Chen, Xin Ning 0001, Kun Wang 0056, Zengfu Wang |
NeurIPS | 6 |
| 2025 | Combinatorial-restless-bandit-based transmitter-receiver online selection of distributed MIMO radar with non-stationary channels
Yuhang Hao, Zengfu Wang, Jing Fu 0001, Xianglong Bai, Can Li 0001, Quan Pan 0001 |
Signal Process. | 2 |
| 2025 | Application of Adaptive Parallel Fast Marching Method in Automatic Submarine Cable Path PlanningabstractSubmarine optical fiber communication cables (subsequently referred to as submarine cables) form the backbone of the Internet’s infrastructure. Damage to these cables can precipitate Internet outages with far-reaching socio-economic impacts. The prevailing practice of manual cable routing is laborious, considering the thousands of kilometers these cables span. It also fails to strike an optimal balance between cost and risk due to its lack of scalability and precision. The Fast Marching Method (FMM), a non-iterative, precise numerical approach capable of solving the Eikonal equation, offers a viable alternative by optimizing the required path between source and destination, considering a summary objective function of costs and risk factors. An interpretation of its solution signifies the optimum value of the objective function between a starting point and all other points. However, the sequential nature of the FMM algorithm suffers from computational limitations and impedes direct parallelization. In this study, we introduce an Adaptive Parallel FMM (APFMM), an innovative approach utilizing adaptive domain decomposition and dynamic multi-resolution analysis. This scalable and widely applicable method can overcome the limitations of existing methods and achieve planning of high-precision, ultra-long-distance (over 14,000 km) submarine cable routes over the Earth’s surface. Simulated experiment results corroborate that APFMM effectively overcomes the computational challenges posed by the sequential FMM when dealing with large datasets. Additionally, it reduces the running time by more than 81% compared to the traditional parallel FMM. This marks a substantial advancement in facilitating efficient, automated, high-precision planning for long-distance submarine cable paths. Note to Practitioners—This paper introduces APFMM, a novel technique based on adaptive domain decomposition and multi-resolution analysis, facilitating high-precision, ultra-long-distance (over 14,000 km) submarine cable path planning. Results from simulated experiments show that APFMM not only overcomes the computational constraints associated with the sequential FMM for large datasets but also cuts the running time by more than 81% relative to the conventional parallel FMM. This breakthrough improvement holds significant implications for practitioners in submarine cable design and construction. With the use of APFMM, designers can plan and optimize cable paths more efficiently, thereby lowering cabling costs and enhancing network resilience. Furthermore, the application of APFMM is not limited to submarine cable path planning and can be employed in other domains involving large-scale data processing and complex path planning, such as electricity cables, gas pipelines, and transportation route planning. While our focus in this paper is primarily on submarine cable path planning, we anticipate practitioners extending the application of APFMM to other use cases, realizing broader utility and benefits. Xinyu Wang 0011, Zengfu Wang, Moshe Zukerman |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2025 | Joint Mixing Data Augmentation for Skeleton-Based Action RecognitionabstractSkeleton-based action recognition is beneficial for understanding human behavior in videos, and thus has received much attention in recent years as an important research area in action recognition. Current research focuses on designing more advanced algorithms to better extract spatio-temporal information from skeleton data. However, due to the small amount of data in the existing skeleton dataset and the lack of effective data augmentation methods, it is easy to lead to overfitting in model training. To address this challenge, we propose a mix-based data augmentation method, Joint Mixing Data Augmentation (JMDA), which can generally improve the effectiveness and robustness of various skeleton-based action recognition algorithms. In terms of spatial information, we introduce SpatialMix (SM), a method that projects the original 3D skeleton discrete information into a 2D space. Then, SM mixes the projected spatial information between two random samples during the training process to achieve the spatial-based mixing data augmentation. Concerning temporal information, we propose TemporalMix (TM). Leveraging the temporal continuity in skeleton data, we perform a temporal resize operation on the original skeleton data, and then merge two random samples during training to achieve the temporal-based mixed data augmentation. Additionally, we analyze the Feature Mismatch (FM) problem caused by introducing mix-based data augmentation into skeleton data. Then we propose a new data preprocessing method called Feature Alignment (FA) to effectively address this problem and improve model performance. Moreover, we propose a novel training pipeline, Joint Training Strategy (JTS), which combines multiple mix-based data augmentation methods for further improvement of model performance. Specifically, our proposed JMDA is plug-and-play and widely applicable to skeleton-based action recognition models. At the same time, the application of JMDA does not increase the model parameters and there is almost no additional training cost. We conduct extensive experiments on NTU RGB+D 60 and NTU RGB+D 120 datasets to demonstrate the effectiveness and robustness of the proposed JMDA on several mainstream skeleton-based action recognition algorithms. Linhua Xiang, Zengfu Wang |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | A Deep Reinforcement Learning-Based Whittle Index Policy for Multibeam AllocationabstractIn this paper, a non-myopic beam scheduling policy is proposed for multi-target tracking (MTT) in a phased-array radar network, seeking to minimize the discounted sum of tracking error of targets and improve the long-term tracking performance. The Whittle index policy based on the restless multiarmed bandit (RMAB) model can decompose the state space of the underlying optimization problem into independent spaces with reduced sizes. We consider the tracking error covariance (TEC) matrix as the state of each target (arm), which evolves based on the Kalman filter. However, for a real-world MTT, the exact calculation of the Whittle index in multiple dimensions is challenging. The neural network is established to achieve the feature extraction of TEC states and learn the corresponding Whittle index. The deep reinforcement learning (DRL) method is exploited to train the neural network by leveraging the threshold property of the Whittle index policy and engaging in interactions with a single target tracking environment. We propose the DRL-based Whittle index policy, namely DRLWI, aiming to solve the beam allocation problem for MTT with multi-dimensional TEC states. This approach effectively mitigates the exponential computational complexity of classical dynamic programming approaches and the low convergence rate caused by large joint state and action spaces in the simple application of DRL algorithms. Numerical results demonstrate the performance of the proposed DRLWI policy surpasses that of DRL algorithms and myopic policies. Yuhang Hao, Zengfu Wang, Jing Fu 0001, Quan Pan 0001 |
FUSION | 2 |
| 2024 | Advancing Micro-Action Recognition with Multi-Auxiliary Heads and Hybrid Loss OptimizationabstractVideo action recognition has been a hot research direction in computer vision, with most existing technologies focusing on coarse-grained macro-action recognition. However, fine-grained action recognition remains challenging. Micro-actions, characterized by high fine-grained, low-intensity, and brief, are crucial for emotion recognition and psychological assessment applications. In this paper, we build on popular video action recognition frameworks as foundation models, introducing multi-auxiliary heads and hybrid loss optimization to advance micro-action recognition. Specifically, the Frame-Level pred and Coarse-Grained Body-Action auxiliary heads work collaboratively to enhance the model and Fine-Grained Micro-Action primary head for perceiving fine-grained and capturing keyframes. Incorporating F1 loss, ArcFace loss, and weighted multi-task loss improves training stability, convergence speed, and performance. Additionally, integrating the optical flow modality enriches the model's diversity, and ensemble learning across all foundational models. Finally, our method achieves a 75.37% F1-mean on the MA-52 dataset, ranking 1st in the Micro-Action Analysis Grand Challenge in conjunction with ACM MM'24. The code is available at https://github.com/qklee-lz/ACMMM2024-MAC. Qiankun Li 0004, Xiaolong Huang 0001, Huabao Chen, Feng He 0008, Qiupu Chen, Zengfu Wang |
ACM Multimedia | 6 |
| 2024 | Enhancing Semi-Dense Feature Matching Through Probabilistic Modeling of Cascaded Supervision and Consistency
Hongchang Min, Yihong Tang, Qiankun Li 0004, Zengfu Wang |
PRCV (15) | 4 |
| 2024 | DeepSweep: Real-Time Multi-View 3D Pose Estimation Via Cross-View Deep Matching and Plane Sweeping
Wenrui Zhu, Qiankun Li 0004, Debin Liu, Zengfu Wang |
PRCV (11) | 4 |
| 2024 | A sea-land clutter classification framework for over-the-horizon radar based on weighted loss semi-supervised generative adversarial network
Zengfu Wang, Mingyue Ji, Yang Li 0055, Quan Pan 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | Chinese text recognition enhanced by glyph and character semantic information
Shilian Wu, Zengfu Wang |
Int. J. Document Anal. Recognit. | 3 |
| 2024 | Video Rescaling With Recurrent DiffusionabstractVideo rescaling helps to fit different display devices. In video rescaling systems, videos are downsampled for easier storage, transmission and preview. The downsampled videos can be upsampled with a neural network to restore the details when needed. Previous group-based video rescaling algorithms benefit from the joint downsampling and joint upsampling of multiple frames, but are restricted by the fully joint operation. In this paper, we propose a recurrent diffusion-based framework for video rescaling. We employ biased joint operation and recurrent diffusion, to make a better use of the temporal relation within different frames in each image group. We explicitly control the direction of information propagation by arranging the processing order of all frames. In biased joint operation, we concentrate on restoring one frame, i.e., the middle frame. The other frames in the group are coarsely reconstructed. Our recurrent diffusion compensates the coarse frames by gradually propagating information from the middle to borders backwardly and forwardly. The recurrent diffusion module is performed by fusing the information of adjacent frames. Biased joint operation and recurrent diffusion are jointly trained. We design several propagation variants and find that our recurrent diffusion is the best among them. It is also shown that recurrent diffusion is better than non-recurrent diffusion in terms of reconstruction quality and model size. We also adopt a high-resolution fine-tuning strategy to further improve the quality of high-resolution frames. Experimental results demonstrate the effectiveness of the proposed method in terms of visual quality, quantitative evaluations, and computational efficiency. The code will be released at https://github.com/5ofwind/RDVR. Dingyi Li, Yu Liu 0023, Zengfu Wang, Jian Yang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | One-Shot Any-Scene Crowd Counting With Local-to-Global GuidanceabstractDue to different installation angles, heights, and positions of the camera installation in real-world scenes, it is difficult for crowd counting models to work in unseen surveillance scenes. In this paper, we are interested in accurate crowd counting based on the data collected by any surveillance camera, that is to count the crowd from any scene given only one annotated image from that scene. To this end, we firstly pose crowd counting as a one-shot learning task. Through the metric-learning, we propose a simple yet effective method that firstly estimates crowd characteristics and then transfers them to guide the model to count the crowd. Specifically, to fully capture these crowd characteristics of the target scene, we devise the Multi-Prototype Learner to learn the prototypes of foreground and density from the limited support image using the Expectation-Maximization algorithm. To learn the adaptation capability for any unseen scene, estimated multi prototypes are proposed to guide the crowd counting of query images in a local-to-global way. CNN is utilized to activate the local features. And transformer is introduced to correlate global features. Extensive experiments on three surveillance datasets suggest that our method outperforms the SOTA methods in the few-shot crowd counting. Zengfu Wang |
IEEE Trans. Image Process. | 2 |
| 2023 | Mask Guided Selective Context Decoding for Handwritten Chinese Text RecognitionabstractHandwritten Chinese text recognition (HCTR) is challenging due to its thousand of characters, diverse writing styles, and ambiguous segmentation. Currently, methods based on connectionist temporal classification (CTC) are widely used, while its independent decoding nature makes it unable to leverage contextual information effectively. In contrast, auto-regression based methods benefit from contextual reasoning, but only half of the contextual information is utilized due to its inherent unidirectionally decoding nature. This article proposes a multi-modal attention-based framework for offline HCTR capable of visual and semantic reasoning. Moreover, a novel mask-guided context-selective decoder is presented to guide the network to decode with randomly selected bidirectional context, further improving the semantic reasoning ability. Extensive experiments show that the proposed method significantly outperforms previous methods. Shilian Wu, Zengfu Wang |
ICASSP | 3 |
| 2023 | Optimizing Spca-based Continual Learning: A Theoretical Approach
Chunchun Yang, Malik Tiomoko, Zengfu Wang |
ICLR | 3 |
| 2023 | Improving CTC-based Handwritten Chinese Text Recognition with Cross-Modality Knowledge Distillation and Feature AggregationabstractOffline handwritten Chinese text recognition (HCTR) models based on connectionist temporal classification (CTC) have recently achieved impressive results. Due to the conditional independence assumption and per-frame prediction characteristics, CTC-based models cannot capture semantic relationships between output tokens and leverage global visual features of characters. To solve these issues, we propose a Cross-Modality knowledge distillation approach that leverages pre-trained LM (BERT) to transfer contextual semantic information, and then design a feature aggregation module to dynamically aggregate local and global features. Experimental results on the HCTR datasets (CASIA-HWDB, ICDAR2013, HCCDOC) show that our proposed method can significantly improve the model’s performance. Shilian Wu, Zengfu Wang |
ICME | 3 |
| 2023 | Data-Efficient Masked Video Modeling for Self-supervised Action RecognitionabstractRecently, self-supervised video representation learning based on Masked Video Modeling (MVM) has demonstrated promising results for action recognition. However, existing methods face two significant challenges: (1) video actions involve a crucial temporal dimension, yet current masking strategies adopt inefficient random approaches that undermine low-density dynamic motion clues in videos; (2) pre-training requires large-scale datasets and significant computing resources (including large batch sizes and enormous iterations). To address these issues, we propose a novel method named Data-Efficient Masked Video Modeling (DEMVM) for self-supervised action recognition. Specifically, a novel masking strategy named Flow-Guided Dense Masking (FGDM) is proposed to facilitate efficient learning by focusing more on the action-related temporal clues, which applies dense masking to dynamic regions based on optical flow priors, while sparse masking to background regions. Furthermore, DEMVM introduces a 3D video tokenizer to enhance the modeling of temporal clues. Finally, Progressive Masking Ratio (PMR) and 2D initialization strategies are presented to enable the model to adapt to the characteristics of the MVM paradigm during different training stages. Extensive experiments on multiple benchmarks, UCF101, HMDB51, and Mimetics, demonstrate that our method achieves state-of-the-art performance in the downstream action recognition task with both efficient data and low computational cost. More interestingly, the few-shot experiment on the Mimetics dataset shows that DEMVM can accurately recognize actions even in the presence of context bias. Qiankun Li 0004, Xiaolong Huang 0001, Zhifan Wan, Lanqing Hu, Shuzhe Wu, Jie Zhang 0071, Shiguang Shan, Zengfu Wang |
ACM Multimedia | 8 |
| 2023 | A Length-Sensitive Language-Bound Recognition Network for Multilingual Text Recognition
Shilian Wu, Zengfu Wang |
MMM (2) | 3 |
| 2023 | SText-DETR: End-to-End Arbitrary-Shaped Text Detection with Scalable Query in Transformer
Pujin Liao, Zengfu Wang |
PRCV (9) | 2 |
| 2023 | Multi-task semi-supervised crowd counting via global to local self-correction
Zengfu Wang |
Pattern Recognit. | 2 |
| 2023 | Data Augmentation and Classification of Sea-Land Clutter for Over-the-Horizon Radar Using AC-VAEGANabstractIn the sea-land clutter classification of sky-wave over-the-horizon-radar (OTHR), the imbalanced and scarce data leads to a poor performance of the deep learning-based classification model. To solve this problem, this paper proposes an improved auxiliary classifier generative adversarial network (AC-GAN) architecture, namely auxiliary classifier variational autoencoder generative adversarial network (AC-VAEGAN). AC-VAEGAN can synthesize higher quality sea-land clutter samples than AC-GAN and serve as an effective tool for data augmentation. Specifically, a one-dimensional convolutional AC-VAEGAN architecture is designed to synthesize sea-land clutter samples. Additionally, an evaluation method combining both traditional evaluation of GAN domain and statistical evaluation of signal domain is proposed to evaluate the quality of synthetic samples. Using a dataset of OTHR sea-land clutter, both the quality of the synthetic samples and the performance of data augmentation of AC-VAEGAN are verified. Further, the effect of AC-VAEGAN as a data augmentation method on the classification performance of imbalanced and scarce sea-land clutter samples is validated. The experiment results show that the quality of samples synthesized by AC-VAEGAN is better than those synthesized by the state-of-the-art GAN-based methods, and the data augmentation method with AC-VAEGAN is able to improve the classification performance in the case of imbalanced and scarce sea-land clutter samples. Zengfu Wang, Quan Pan 0001, Yang Li 0055 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Multi-Branch Network with Ensemble Learning for Text Removal in the Wild
Yujie Hou, Zengfu Wang |
ACCV (3) | 3 |
| 2022 | Least-Squares Estimation of Keypoint Coordinate for Human Pose Estimation
Linhua Xiang, Jia Li 0026, Zengfu Wang |
PRCV (3) | 3 |
| 2022 | Video super-resolution with inverse recurrent net and hybrid local fusion
Dingyi Li, Zengfu Wang, Jian Yang 0003 |
Neurocomputing | 2 |
| 2022 | Monocular depth estimation with spatially coherent sliced network
Wen Su 0004, Haifeng Zhang 0006, Yuan Su, Jun Yu 0001, Zengfu Wang |
Image Vis. Comput. | 5 |
| 2022 | AFA: adversarial frequency alignment for domain generalized lung nodule detection
Mei Sun, Jing Zhang 0037, Cong Liu 0006, Zengfu Wang |
Neural Comput. Appl. | 6 |
| 2022 | SSR-HEF: Crowd Counting With Multiscale Semantic Refining and Hard Example FocusingabstractCrowd counting based on density maps is generally regarded as a regression task. Deep learning is used to learn the mapping between image content and crowd density distribution. Although great success has been achieved, some pedestrians far away from the camera are difficult to be detected. And the number of hard examples is often larger. Existing methods with simple Euclidean distance algorithm indiscriminately optimize the hard and easy examples so that the densities of hard examples are usually incorrectly predicted to be lower or even zero, which results in large counting errors. To address this problem, we are the first to propose the hard example focusing (HEF) algorithm for the regression task of crowd counting. The HEF algorithm makes our model rapidly focus on hard examples by attenuating the contribution of easy examples. Then higher importance will be given to the hard examples with wrong estimations. Moreover, the scale variations in crowd scenes are large, and the scale annotations are labor-intensive and expensive. By proposing a multiscale semantic refining strategy, lower layers of our model can break through the limitation of deep learning to capture semantic features of different scales to sufficiently deal with the scale variation. We perform extensive experiments on six benchmark datasets to verify the proposed method. Results indicate the superiority of our proposed method over the state-of-the-art methods. Moreover, our designed model is smaller and faster. Kewei Wang 0004, Wen Su 0004, Zengfu Wang |
IEEE Trans. Ind. Informatics | 4 |
| 2022 | Densely Enhanced Semantic Network for Conversation System in Social MediaabstractThe human–computer conversation system is a significant application in the field of multimedia. To select an appropriate response, retrieval-based systems model the matching between the dialogue history and response candidates. However, most of the existing methods cannot fully capture and utilize varied matching patterns, which may degrade the performance of the systems. To address the issue, a densely enhanced semantic network (DESN) is proposed in our work. Given a multi-turn dialogue history and a response candidate, DESN first constructs the semantic representations of sentences from the word perspective, the sentence perspective, and the dialogue perspective. In particular, the dialogue perspective is a novel one introduced in our work. The dependencies between a single sentence and the whole dialogue are modeled from the dialogue perspective. Then, the response candidate and each utterance in the dialogue history are made to interact with each other. The varied matching patterns are captured for each utterance–response pair by using a dense matching module. The matching patterns of all the utterance–response pairs are accumulated in chronological order to calculate the matching degree between the dialogue history and the response. The responses in the candidate pool are ranked with the matching degree, thereby returning the most appropriate candidate. Our model is evaluated on the benchmark datasets. The experimental results prove that our model achieves significant and consistent improvement when compared with other baselines. Zengfu Wang, Jun Yu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | Submarine Cable Network Design for Regional ConnectivityabstractThis paper optimizes path planning for a trunk-and-branch topology network in an irregular 2-dimensional manifold embedded in 3-dimensional Euclidean space with application to submarine cable network planning. We go beyond our earlier focus on the weighted costs of cables (cable laying cost, resilience, design level and repair rate) to include the cost of branching units (BUs), including material and labor, as well as submarine cable landing stations (CLSs). This optimization also includes choices of locations of BUs and CLSs. These are important issues for the economics of cable laying and significantly change the model and the optimization process. We pose the problem as a variant of the Steiner tree problem, but one in which the Steiner nodes can vary in number, while incurring a penalty. We refer to it as the weighted Steiner node problem. It differs from the Euclidean Steiner tree problem, where Steiner points are forced to have degree three; this is no longer the case, in general, when nodes incur a cost. We are able to prove that our algorithm is applicable to Steiner nodes with degree greater than three, enabling optimization of network costs in this context. The optimal solution is achieved in polynomial-time using dynamic programming. Zengfu Wang, William Moran 0001, Moshe Zukerman |
IEEE/ACM Trans. Netw. | 2 |
| 2021 | Color Multi-focus Image Fusion Using Quaternion Morphological Gradient and Improved KNN Matting
Zengfu Wang |
ICIG (1) | 3 |
| 2021 | Fine-Grained Language Identification in Scene Text ImagesabstractIdentifying the language of the text in scene images is crucial for various applications. Studies that focus on identifying the script, which is a set of letters used for writing in a given language, in scene text images already exist. However, these works do not distinguish between different languages written in the same script and are thus unable to meet the needs of many applications. To address this challenge, we study a novel task: fine-grained language identification in scene text images, which aims to distinguish languages that share the same script. The datasets that include samples in seven languages, which are Dutch, English, French, Italian, German, Spanish, and Portuguese, are constructed. Furthermore, well-designed end-to-end trainable neural networks are proposed for fine-grained language identification, where semantic information concerning the text is mined and utilized to assist the language identification. We train the networks on the synthetic dataset and evaluate them with the collected real dataset. The experimental results demonstrate that the proposed frameworks are effective. Shilian Wu, Jun Yu 0001, Zengfu Wang |
ACM Multimedia | 4 |
| 2021 | SimplePose V2: Greedy Offset-Guided Keypoint Grouping for Human Pose Estimation
Jia Li 0026, Linhua Xiang, Zengfu Wang |
PRCV (1) | 4 |
| 2021 | Crowd counting with segmentation attention convolutional neural networkabstractAbstract Deep learning occupies an undisputed dominance in crowd counting. This paper proposes a novel convolutional neural network architecture called SegCrowdNet. Despite the complex background in crowd scenes, the proposed SegCrowdNet still adaptively highlights the human head region and suppresses the non‐head region by segmentation. With the guidance of an attention mechanism, the proposed SegCrowdNet pays more attention to the human head region and automatically encodes the highly refined density map. The crowd count can be obtained by integrating the density map. To adapt the variation of crowd counts, SegCrowdNet intelligently classifies the crowd count of each image into several groups. In addition, the multi‐scale features are learned and extracted in the proposed SegCrowdNet to overcome the scale variations of the crowd. To verify the effectiveness of this proposed method, extensive experiments are conducted on four challenging datasets. The results demonstrate that the proposed SegCrowdNet achieves excellent performance compared with the state‐of‐the‐art methods. Zengfu Wang |
IET Image Process. | 2 |
| 2021 | Generative domain adaptation for chest X-ray image analysisabstractAbstract Chest X‐ray images taken under different conditions follow different distributions, preventing the models trained on a domain from generalising well on the other domain. In this paper, a generative domain adaptation (GDA) method is proposed to address this issue and facilitate the learning process for downstream analysis. GDA adapts different domains to a virtual common one where images are aligned at the appearance level. To this end, a domain shared generator is used to transform the input images and two competitive discriminators are used to adversarially supervise the transforming process. The domain discriminator drives the generator to narrow the domain gap while the fidelity discriminator forces the generator to keep the inherent information. Moreover, a specific classification or detection network is attached to the generator to supervise it in a task‐oriented manner. Experiment results on a large‐scale dataset containing 46k chest X‐ray images demonstrate that GDA outperforms representative domain adaptation methods by a large margin for both disease classification and lesion detection as well as provides useful transformed images to assist experts for diagnosis. Zhong-Hua Fu, Jing Zhang 0037, Cong Liu 0006, Zengfu Wang |
IET Image Process. | 6 |
| 2021 | OTHR multitarget tracking with a GMRF model of ionospheric parameters
Zengfu Wang, Hua Lan, Quan Pan 0001 |
Signal Process. | 2 |
| 2021 | Robust multi-focus image fusion using lazy random walks with multiscale focus measures
Zengfu Wang |
Signal Process. | 3 |
| 2021 | Monocular Depth Estimation Using Information Exchange NetworkabstractDepth estimation from single monocular image attracts increasing attention in autonomous driving and computer vision. While most existing approaches regress depth values or classify depth labels based on features extracted from limited image area, the resulting depth maps are still perceptually unsatisfying. Neither local context nor low-level semantic information is sufficient to predict depth. Learning based approaches suffer from inherent defects of supervision signals. This paper addresses monocular depth estimation with a general information exchange convolutional neural network. We maintain a high-resolution prediction throughout the network. Meanwhile, both low-resolution features capturing long-range context and fine-grained features describing local context can be refined with information exchange path stage by stage. Mutual channel attention mechanism is applied to emphasize interdependent feature maps and improve the feature representation of specific semantics. The network is trained under the supervision of improved log-cosh and gradient constraints so that the abnormal predictions have less impacts and the estimation can be consistent in high order. The results of ablation studies verify the efficiency of every proposed components. Experiments on the popular indoor and street-view datasets show competitive results compared with the recent state-of-the-art approaches. Wen Su 0004, Haifeng Zhang 0006, Quan Zhou 0004, Wenzhen Yang, Zengfu Wang |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2020 | Simple Pose: Rethinking and Improving a Bottom-up Approach for Multi-Person Pose EstimationabstractWe rethink a well-known bottom-up approach for multi-person pose estimation and propose an improved one. The improved approach surpasses the baseline significantly thanks to (1) an intuitional yet more sensible representation, which we refer to as body parts to encode the connection information between keypoints, (2) an improved stacked hourglass network with attention mechanisms, (3) a novel focal L2 loss which is dedicated to “hard” keypoint and keypoint association (body part) mining, and (4) a robust greedy keypoint assignment algorithm for grouping the detected keypoints into individual poses. Our approach not only works straightforwardly but also outperforms the baseline by about 15% in average precision and is comparable to the state of the art on the MS-COCO test-dev dataset. The code and pre-trained models are publicly available on our project page1. Jia Li 0026, Wen Su 0004, Zengfu Wang |
AAAI | 3 |
| 2020 | Weakly Supervised Local-Global Relation Network for Facial Expression RecognitionabstractTo extract crucial local features and enhance the complementary relation between local and global features, this paper proposes a Weakly Supervised Local-Global Relation Network (WS-LGRN), which uses the attention mechanism to deal with part location and feature fusion problems. Firstly, the Attention Map Generator quickly finds the local regions-of-interest under the supervision of image-level labels. Secondly, bilinear attention pooling is employed to generate and refine local features. Thirdly, Relational Reasoning Unit is designed to model the relation among all features before making classification. The weighted fusion mechanism in the Relational Reasoning Unit makes the model benefit from the complementary advantages between different features. In addition, contrastive losses are introduced for local and global features to increase the inter-class dispersion and intra-class compactness at different granularities. Experiments on lab-controlled and real-world facial expression dataset show that WS-LGRN achieves state-of-the-art performance, which demonstrates its superiority in FER. Haifeng Zhang 0006, Wen Su 0004, Jun Yu 0001, Zengfu Wang |
IJCAI | 4 |
| 2020 | Crowd counting with crowd attention convolutional neural network
Wen Su 0004, Zengfu Wang |
Neurocomputing | 3 |
| 2020 | A message passing approach for multiple maneuvering target tracking
Hua Lan, Jirong Ma, Zengfu Wang, Quan Pan 0001 |
Signal Process. | 3 |
| 2020 | A novel multi-focus image fusion method using multiscale shearing non-local guided averaging filter
Zengfu Wang |
Signal Process. | 2 |
| 2020 | Optimal Submarine Cable Path Planning and Trunk-and-Branch Tree Network Topology DesignabstractWe study the path planning of submarine cable systems with trunk-and-branch tree topology on the surface of the earth. Existing work on path planning represents the earth's surface by triangulated manifolds and takes account of laying cost of the cable including material, labor, alternative protection levels, terrain slope and survivability of the cable. Survivability issues include the risk of future cable break associated with laying the cable through sensitive and risky areas, such as, in particular, earthquake-prone regions. The key novelty of this paper is an examination and solution of the path planning of submarine cable systems with trunk-and-branch tree topology. We formulate the problem as a Steiner minimal tree problem on irregular 2D manifolds in R3. For a given Steiner topology, we propose a polynomial time computational complexity numerical method based on the dynamic programming principle. If the topology is unknown, a branch and bound algorithm is adopted. Simulations are performed on real-world three-dimensional geographical data. Zengfu Wang, Qing Wang 0022, William Moran 0001, Moshe Zukerman |
IEEE/ACM Trans. Netw. | 1 |
| 2019 | Expression-identity Fusion Network for Facial Expression RecognitionabstractResearch shows that the facial expression recognition is strongly related to the person's identity. This paper presents an expression-identity fusion network to address the great inter-subject variations in facial expression recognition. The model is designed to jointly learn identity-related features and expression-related features via two branches with the same expression image input. A bilinear module is introduced to fuse two kinds of features and learn the relationship between them. Experimental results show that identity-related features can greatly boost the performance of facial expression recognition. Our method outperforms most of the state-of-the-art. On two popular facial expression databases (CK+ and Oulu-CASIA), our method achieves 96.02% and 85.21% recognition accuracy, respectively. Haifeng Zhang 0006, Wen Su 0004, Zengfu Wang |
ICASSP | 3 |
| 2019 | Dense Semantic Matching Network for Multi-turn ConversationabstractMining the semantic information in the text to model the multi-turn conversation has attracted great interests. Previous models ignore the capturing and utilizing of the matching patterns at different levels, which causes the loss of the valuable information for the calculation of the matching degree. To address the problem, we propose a dense semantic matching network (DMN). Given a context-response pair, DMN first constructs semantic representations for the response candidate and each utterance in the context. Then, DMN models the interaction between the response candidate and each utterance in the context to generate interactive matrices. By processing the interactive matrices with dense convolutional blocks, the hierarchical matching patterns are generated for each utterance-response pair. The matching patterns of all the utterance-response pairs are finally accumulated in chronological order with bidirectional long short-term memory network. The final matching score of the response candidate and the multi-turn context is finally calculated. We evaluate the performance of our network on the benchmark dataset. The results show that our network yields a significant performance gain compared with other methods. Jun Yu 0001, Zengfu Wang |
ICDM | 3 |
| 2019 | A Novel Multi-focus Image Fusion Based on Lazy Random Walks
Zengfu Wang |
ICIG (2) | 2 |
| 2019 | Monocular Depth Estimation as Regression of Classification using Piled Residual NetworksabstractPredicting depth from single monocular image is a challenging task in scene understanding. Most existing work predicts depth by regression or classification with features extracted from local neighborhood area. However, neither regression nor classification achieves the final satisfying solution and local context can be insufficient to predict the depth. This paper innovatively addresses this problem as regression of class related features on a piled residual convolutional neural network. Our framework works at two stages. First, a well-designed deep convolutional neural network model is employed to classify the depths in difference-scale invariance space. The model utilizes all scales of context though piled residual paths. The deeper layers that capture high-level semantic features with long-range context can be directly refined using fine-grained features with local context from earlier convolutions. We then apply centered information gain loss to the model to produce intra-class compact and inter-class discriminative features. Second, to obtain depths instead of class labels, we infer depth regression with convolutional layers which model the mapping from class discriminative features to continuous depth values. Experiments on the popular indoor and outdoor datasets show competitive results compared with the recent state of the art methods. Wen Su 0004, Haifeng Zhang 0006, Jia Li 0013, Wenzhen Yang, Zengfu Wang |
ACM Multimedia | 5 |
| 2019 | 3D Singing Head for Music VR: Learning External and Internal Articulatory Synchronicity from Lyric, Audio and NotesabstractWe propose a real-time 3D singing head system to enhance the talking head on model integrity, keyframe generation and song synchronicity. The individual head appearance meshes are first obtained by matching multi-view visible images with face prior for accuracy, and then used to reconstruct entire head model by integrating with generic internal articulatory meshes for efficiency. After embedding physiology, the keyframes of each phoneme-music note correspondence are substantially synthesized from real articulation data. The song synchronicity of articulators is learned using a deep neural network to train visual co-articulation model (VCM) on parallel audio-visual data. Finally, the keyframes of adjacent phoneme-music note correspondences are blended by VCM to produce song synchronized animation. Compared to state-of-the-art baselines, our system can not only clearly distinguish phonemes and notes, but also significantly reduce the dependence on training data. Jun Yu 0001, Chang Wen Chen, Zengfu Wang |
ACM Multimedia | 3 |
| 2019 | Widening residual refine edge reserved neural network for semantic segmentation
Wen Su 0004, Zengfu Wang |
Multim. Tools Appl. | 2 |
| 2019 | Video Super-Resolution Using Non-Simultaneous Fully Recurrent Convolutional NetworkabstractVideo super-resolution (SR) aims at restoring fine details and enhancing visual experience for low-resolution (LR) videos. In this paper, we propose a very deep non-simultaneous fully recurrent convolutional network for video SR. To make full use of temporal information, we employ motion compensation, very deep fully recurrent convolutional layers and late fusion in our system. Residual connection is also employed in our recurrent structure for more accurate SR. Finally a new model ensemble strategy is used to combine our method with single-image SR method. Experimental results demonstrate that the proposed method is better than state-of-the-art SR methods on quantitative visual quality assessment. Dingyi Li, Yu Liu 0023, Zengfu Wang |
IEEE Trans. Image Process. | 3 |
| 2019 | Real-Time Traffic Sign Recognition Based on Efficient CNNs in the WildabstractBoth unmanned vehicles and driver assistance systems require solving the problem of traffic sign recognition. A lot of work has been done in this area, but no approach has been presented to perform the task with high accuracy and high speed under various conditions until now. In this paper, we have designed and implemented a detector by adopting the framework of faster R-convolutional neural networks (CNN) and the structure of MobileNet. Here, color and shape information have been used to refine the localizations of small traffic signs, which are not easy to regress precisely. Finally, an efficient CNN with asymmetric kernels is used to be the classifier of traffic signs. Both the detector and the classifier have been trained on challenging public benchmarks. The results show that the proposed detector can detect all categories of traffic signs. The detector and the classifier proposed here are proved to be superior to the state-of-the-art method. Our code and results are available online. Jia Li 0013, Zengfu Wang |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2018 | OTHR Multipath Tracking with Correlated Virtual Ionospheric HeightsabstractThis paper proposes a new virtual ionospheric height model for over-the-horizon radar (OTHR) target tracking. Considering the spatial correlation of different ionosphere site, the virtual ionospheric heights are modeled by a Gaussian Markov random field (GMRF). The priors of the GMRF model can be learned from the historical measurements from ionosondes. Given the acquired measurements of the ionosphere subregions, the virtual ionospheric heights of the unmeasured subregions are inferred based on the GMRF model. Then we present the multipath probabilistic data association for uncertain coordinate registration (MPCR) with the new virtual ionospheric height model. Numerical simulation shows that the accuracy of OTHR target tracking is improved. Zengfu Wang, Yumei Hu, Quan Pan 0001 |
FUSION | 2 |
| 2018 | Improved Adaptive Kalman Filter with Unknown Process Noise CovarianceabstractThis paper considers the joint recursive estimation of the dynamic state and the time-varying process noise covariance for a linear state space model. The conjugate prior on the process noise covariance, the inverse Wishart distribution, provides a latent variable. A variational Bayesian inference framework is then adopted to iteratively estimate the posterior density functions of the dynamic state, process noise covariance and the introduced latent variable. The performance of the algorithm is demonstrated with simulated data in a target tracking application. Jirong Ma, Hua Lan, Zengfu Wang, Xuezhi Wang 0001, Quan Pan 0001, William Moran 0001 |
FUSION | 3 |
| 2018 | A Generative Adversarial Network Based Framework for Unsupervised Visual Surface InspectionabstractVisual surface inspection is a challenging task due to the highly inconsistent appearance of the target surfaces and the abnormal regions. Most of the state-of-the-art methods are highly dependent on the labelled training samples, which are difficult to collect in practical industrial applications. To address this problem, we propose a generative adversarial network based framework for unsupervised surface inspection. The generative adversarial network is trained to generate the fake images analogous to the normal surface images. It implies that a well-trained GAN indeed learns a good representation of the normal surface images in a latent feature space. And consequently, the discriminator of GAN can serve as a naturally one-class classifier. We use the first three conventional layer of the discriminator as the feature extractor, whose response is sensitive to the abnormal regions. Particularly, a multi-scale fusion strategy is adopted to fuse the responses of the three convolution layers and thus improve the segmentation performance of abnormal detection. Various experimental results demonstrate the effectiveness of our proposed method. Wei Zhai, Yang Cao 0010, Zengfu Wang |
ICASSP | 4 |
| 2018 | Local Regression Based Hourglass Network for Hand Pose Estimation from a Single Depth ImageabstractHand pose estimation plays an important role in many applications such as human-computer interaction. With the advent of commodity depth sensors and the developments of deep learning, noticeable improvements have been made in this field recently. Nevertheless, the accuracy and robustness of existing approaches are still dissatisfying. In this paper, we propose an end-to-end local regression based hourglass network with a modified loss function to estimate the 3D pose of the hand in a depth image. We use a third order hourglass block to extract features of the hand. At the top of our network, we slice the feature map into several regions and regress the regions independently first. Then, we merge the regression results and feed them to the final regressor. Besides, we compare performances of different loss functions for the task. The results indicate that the structure of the network and the loss function designed here lead to an obvious improvement. And the proposed approach is comparable to, or superior to the state-of-the-art on a public challenging dataset. Our system can run at over 910 FPS on a single GPU, and the mean error of estimation is reduced to 12.36 mm. Jia Li 0026, Zengfu Wang |
ICPR | 2 |
| 2018 | Dense Convolutional Recurrent Neural Network for Generalized Speech AnimationabstractThis paper presents a novel automated speech animation approach named Dense Convolutional Recurrent Neural Network (DenseCRNN). The approach learns a non-linear mapping from acoustic speech to multiple articulator movements in a unified framework to which feature extraction, context encoding and multi-parameter decoding are integrated. We propose DenseCRNN based on three insights: (1) One can use a convolutional neural network incorporated with dense connectivity to extract speaker-independent features from arbitrary spoken audio effectively. (2) A bidirectional long short-term memory neural network is able to model the context information with respect to the phoneme coarticulation. (3) Multi-domain learning can be implemented to achieve better performance on account of the implicit correlation and explicit difference among outputs, where each domain is responsible for a single visual parameter. Experiments on MNGU0 dataset demonstrate our approach achieves significant improvements over state-of-the-art methods. Moreover, the proposed approach generalizes over different gender or accent, and has the capability of deploying on various character models. Zengfu Wang |
ICPR | 2 |
| 2018 | Multi-scale Semantic Segmentation Enriched Features for Pedestrian DetectionabstractPedestrian detection, as a branch of computer vision, has many significant real world applications such as autonomous driving or human behavior analysis. In this paper, we propose a convolutional neural network (CNN) based pedestrian detection framework which can be trained end-to-end. We design a feature enrichment unit to produce more representative features to improve detection performance. The feature enrichment units receive feature maps from the body network layer by layer and convey features in a backward manner. Together they produce multi-scale semantic segmentation results as extra features and merge them with feature maps of the body network. Then the merged feature maps will be fed into the detector to produce final predictions. The feature enrichment unit is easy to embed into existing convolutional neural networks based detection frameworks since it receives and produces feature maps. We use an alternating training strategy to train the network for detection and segmentation respectively and achieve considerable accuracy. The multi-scale feature enrichment units improve detection accuracy significantly as proven by the experiments. Xiaolu Xie, Zengfu Wang |
ICPR | 2 |
| 2018 | Dynamic Facial Expression Recognition Based on Trained Convolutional Neural Networks
Zengfu Wang |
PRCV (2) | 2 |
| 2018 | Learning Non-local Representation for Visual Tracking
Peng Zhang 0075, Zengfu Wang |
PRCV (4) | 2 |
| 2018 | Probability contour guided depth map inpainting and superresolution using non-local total generalized variation
Hai-Tao Zhang, Jun Yu 0001, Zengfu Wang |
Multim. Tools Appl. | 3 |
| 2018 | Real-Time 3-D Facial Animation: From Appearance to Internal ArticulatorsabstractA real-time 3-D facial animation system produces animation for appearance and internal articulators. For appearance, an anatomical model, including the skeleton, muscle, and skin, is built based on anatomical characteristics and a data-driven model is obtained by learning the mapping between texture and depth. Then, the two models are combined to produce animations with various strengths, since the anatomical model can control the animation strength directly and the data-driven model can capture the nuances of facial motion. For internal articulators, tongue tissue arrangements are obtained from medical data. Then, a nonlinear, quasi-incompressible, isotropic, hyperelastic biomechanical model is applied to describe tongue tissues and an anisotropic biomechanical model is applied to reflect the active and passive mechanical behavior of tongue muscle fibers. The tongue animation is simulated using the finite-element method for realism, while the collisions between the tongue and other articulators are simulated with a mass-spring model for efficiency. Experiments show that the system achieves high perceptual evaluation scores and quantitative improvements are demonstrated in the objective evaluation and user studies, compared with the outputs of other systems. Jun Yu 0001, Chen Jiang 0002, Rui Li 0021, Changwei Luo, Zengfu Wang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2017 | A real-time 3D head mesh modeling and expressive articulatory animation systemabstractIn view of animated human computer interfaces, this paper proposes a 3D head mesh modeling and expressive articulatory animation system. The appearance mesh model is first reconstructed from multi-view visible images using inter-regional cooperative optimization and depth super-resolution, and the universal internal articulatory mesh is then integrated with the reconstructed appearance mesh by interpolation. After establishing the head mesh model, the anatomical and biomechanical characteristics of articulators are combined to synthesize articulatory animation. The evaluations demonstrate the system can build a realistic and vivid virtual head for animated interface in real-time. Jun Yu 0001, Zengfu Wang |
ICASSP | 2 |
| 2017 | Video super-resolution using motion compensation and residual bidirectional recurrent convolutional networkabstractVideo super-resolution (SR) aims at restoring finer details and enhancing visual experience. In this paper, we propose a novel method named residual recurrent convolutional network (RRCN) for video SR. In our method, motion compensation and bidirectional residual convolutional network are combined to model the spatial and temporal non-linear mappings. To leverage sufficient amount of temporal information, we employ motion compensation, bidirectional recurrent convolutional layers and late fusion in of our network. We also apply residual connections in our recurrent structure for more accurate SR. Experimental results demonstrate the superiority of the proposed method over state-of-the-art single-image and multi-frame based SR approaches in terms of both quantitative assessment and visual quality. Dingyi Li, Yu Liu 0023, Zengfu Wang |
ICIP | 3 |
| 2017 | A deep CNN method for underwater image enhancementabstractUnderwater images often suffer from color distortion and visibility degradation due to the light absorption and scattering. Existing methods utilize various assumptions/constrains to achieve reasonable solutions for underwater image enhancement. However, these methods share the common limitation that the adopted assumptions may not work for some particular scenes. To address this problem, this paper proposes an end to end framework for underwater image enhancement, where a CNN-based network called UIE-Net is presented. The UIE-net is trained with two tasks, color correction and haze removal. This unified training approach enables learning a strong feature representation for both tasks simultaneously. For better extracting the inherent features in local patches, a pixels disrupting strategy is exploited in the proposed learning framework, which significantly improves the convergent speed and accuracy. To handle the training of UIE-net, we synthesize 200000 training images based on the physical underwater imaging model. Experiments on benchmark underwater images for cross-scenes show that UIE-net achieves superior performance over existing methods. Yang Wang 0015, Jing Zhang 0037, Yang Cao 0010, Zengfu Wang |
ICIP | 4 |
| 2017 | Depth map super-resolution using non-local higher-order regularization with classified weightsabstractHigh-order regularization in depth map super-resolution (SR) contributes to producing smoother depth map. However, assigning appropriate weights within regularization term is also important for preserving more detail information. In this paper, a novel and more adaptive depth SR model is proposed by using non-local total generalized variation (NLTGV) with classified weights. A random forest based classifier is trained to classify the pixels of depth map into four categories on the basis of several local structure features, such as gradient magnitude and texture energy extracted from color image, and then the weights within NLTGV are assigned with four groups of parameters corresponding to the four kinds of pixels. Evaluation results demonstrate that the local features make pixels have a good separability, and the classified weights can obviously improve the accuracy of depth map SR. Hai-Tao Zhang, Jun Yu 0001, Zengfu Wang |
ICIP | 3 |
| 2017 | Image guided depth enhancement via deep fusion and local linear regularizaronabstractDepth maps captured by RGB-D cameras are often noisy and incomplete at edge regions. Most existing methods assume that there is a co-occurrence of edges in depth map and its corresponding color image, and improve the quality of depth map guided by the color image. However, when the color image is noisy or richly detailed, the high frequency artifacts will be introduced into depth map. In this paper, we propose a deep residual network based on deep fusion and local linear regularization for guided depth enhancement. The presented scheme can effectively extract the correlation between depth map and color image in the deep feature space. To reduce the difficulty of training, a specific layer of network which introduces a local linear regularization constraint on the output depth is designed. Experiments on various applications, including depth denoising, super-resolution and inpainting, demonstrate the effectiveness and reliability of our proposed approach. Jing Zhang 0037, Yang Cao 0010, Zengfu Wang |
ICIP | 4 |
| 2017 | Widening residual skipped network for semantic segmentationabstractOver the past two years deep convolutional neural networks have pushed the performance of computer vision systems to soaring heights on semantic segmentation. In this study, the authors present a novel semantic segmentation method of using a deep fully convolutional neural network to achieve image segmentation results with more precise boundary localisation. The above segmentation engine is trainable, and consists of an encoder network with widening residual skipped connections and a decoder network with a pixel‐wise classification layer. Here the encoder network with widening residual skipped connections allows the combination of shallow layer features and deep layer semantic features, and the decoder network with classification layer maps the low‐resolution encoder features to full resolution image with pixel‐wise classification. The experimental results on PASCAL VOC 2012 semantic segmentation dataset and Cityscapes dataset show that the proposed method is effective and competitive. Wen Su 0004, Zengfu Wang |
IET Image Process. | 2 |
| 2017 | Simultaneously retargeting and super-resolution for stereoscopic video
Yang Cao 0010, Zengfu Wang |
Multim. Tools Appl. | 3 |
| 2017 | Image classification based on convolutional neural networks with cross-level strategy
Yu Liu 0023, Jun Yu 0001, Zengfu Wang |
Multim. Tools Appl. | 4 |
| 2017 | Creating and simulating a realistic physiological tongue model for speech production
Jun Yu 0001, Chen Jiang 0002, Zengfu Wang |
Multim. Tools Appl. | 3 |
| 2017 | A Video-Based Facial Motion Tracking and Expression Recognition System
Jun Yu 0001, Zengfu Wang |
Multim. Tools Appl. | 2 |
| 2017 | A realistic 3D articulatory animation system for emotional visual pronunciation
Lingyun Yu 0002, Jun Yu 0001, Zengfu Wang |
Multim. Tools Appl. | 3 |
| 2017 | Route Selection for Cabling Considering Cost Minimization and Earthquake Survivability Via a Semi-Supervised Probabilistic ModelabstractThis paper focuses on an important and fundamental problem of connecting two points by a cable, subject to a tradeoff between cost and earthquake survivability. In particular, we address the problem of selecting a route for laying a cable under arbitrary topography, based on earthquake data. First, we derive a semi-supervised probability density estimation model for the likelihood of earthquake disaster. Based on this probabilistic model, we generate a nearest neighbor graph. The graph represents each data point with a four-dimensional space formed by the three-dimensional undersea coordinates and the one-dimensional data of earthquake disaster level. It then forms the weight on graph between any positions. The data used in this study are all real data of undersea topography and earthquake information of the Taiwan Strait. As a result, both the undersea topology and the earthquake level can be transferred into a distance for shortest route finding. Finally, Dijkstra's algorithm is used for finding the optimal shortest route for cabling between the two given points on the graph. Extensive simulations based on a synthetic dataset and the Taiwan Strait real-world dataset corroborate the effectiveness of the proposed method. Ming-Bo Zhao, Tommy W. S. Chow, Zengfu Wang, Jun Guo 0001, Moshe Zukerman |
IEEE Trans. Ind. Informatics | 4 |
| 2016 | A fast and precise speech-triggered tongue animation system by combining parameterized model and anatomical modelabstractA 3D realistic tongue system is proposed. Firstly, the muscle geometry and fiber arrangement are specified after a tongue mesh model is constructed from medical data. Secondly, with the target of the efficiency and realism of animation, the tongue tissues, including tongue muscles, are described by combining a fast parametric model and a precise anatomical model to simulate the active and passive mechanical characteristics. The experiments demonstrate the suitability of the system for speech visualization. Jun Yu 0001, Chen Jiang 0002, Zengfu Wang |
BIBM | 3 |
| 2016 | A realistic and reliable 3D pronunciation visualization instruction system for computer-assisted language learningabstractA text-driven 3D pronunciation visualization instruction system is proposed for computer-assisted language learning. Based on a 3D articulatory mesh model including appearance and internal articulators, both finite element method and anatomical model are used to synthesize the articulatory animation of phonemes by fitting the mesh model to the detected articulatory shapes in X-ray images. Visual co-articulation is modeled with a Hidden Markov Model trained on an articulatory speech corpus. Articulatory animations corresponding to all phonemes of a learned text are concatenated by visual co-articulation model to produce the speech synchronized articulatory animation. The experiments for Mandarin Chinese show the system can increase the pronunciation accuracy of learners. Jun Yu 0001, Zengfu Wang |
BIBM | 2 |
| 2016 | The application of sum-product algorithm for data association
Hua Lan, Zengfu Wang, Quan Pan 0001 |
FUSION | 3 |
| 2016 | Fast response aggregation for depth estimation using light field cameraabstractLight field cameras have been recently shown to be very effective in applications such as multifocusing and 3D reconstruction. These cameras can provide depth cues from both defocus and correspondence in a single snapshot. In this paper, we present a fast response aggregation framework for depth estimation by jointly using defocus and correspondence cues. Different from existing approaches, we perform a fast gradient preserving filtering in a label domain, instead of in a depth domain, to efficiently compute a dense depth map. The proposed approach comprises of three steps: 1) constructing defocus and correspondence response volumes, 2) adaptively smoothing the two volumes and performing Winner-Takes-All label selections, and 3) post-processing by using nonlocal image guided averaging. With such a compact framework, currently best depth estimation results can be achieved. This compact framework is suitable for various applications such as object segmentation and surface reconstruction. Cao Yang, Jing Zhang 0037, Zengfu Wang |
ICASSP | 4 |
| 2016 | Regularized fully convolutional networks for RGB-D semantic segmentationabstractThe prospect of semantic segmentation using depth is alluring. In most of the previous work features are only combined by using simple classification strategies. The inter-feature and inter-label relationships have been ignored. This paper proposes a novel unified framework for RGB-D semantic segmentation. We use regularized fully convolutional networks whose inputs are depth map and hand-crafted features. Relationships between those features and their labels are learnt and utilized by rigorously imposing regularization in fully connected layers. The regularized fully convolutional networks can be efficiently launched using a GPU implementation at an affordable training cost. Experiments demonstrate that our regularized fully convolutional networks taking features as inputs obtain competitive results on the PASCAL VOC 2011 dataset and NYUDv2. Wen Su 0004, Zengfu Wang |
VCIP | 2 |
| 2016 | Image guided depth map superresolution using non-local total generalized variationabstractIn order to solve the problem of depth map super-resolution, this paper proposes a novel approach to obtain super-resolution result from the raw depth map captured by a 3D camera under a convex optimization framework. In our method, non-local total generalized variation (NL-TGV) is utilized to measure the smoothness of depth map, and the data-fidelity term is expressed by the Huber norm. To preserve the sharpness of depth discontinuities, the smoothing weights are decided by the combination of bilateral weight and separating probability between pixels, and the histogram distance between pixels in the raw depth map is used to adjust the combination weights for suppressing the texture-transfer. We have derived a numerical solution scheme for the optimization problem using the first order primal dual algorithm. Quantitative and qualitative evaluations on the synthesis datasets and the real datasets demonstrate that our method is as good as the state-of-the-art approaches. Hai-Tao Zhang, Zengfu Wang |
VCIP | 3 |
| 2016 | Automatic chessboard corner detection methodabstractChessboard corner detection is a necessary procedure of the popular chessboard pattern‐based camera calibration technique, in which the inner corners on a two‐dimensional chessboard are employed as calibration markers. In this study, an automatic chessboard corner detection algorithm is presented for camera calibration. In authors’ method, an initial corner set is first obtained with an improved Hessian corner detector. Then, a novel strategy that utilises both intensity and geometry characteristics of the chessboard pattern is presented to eliminate fake corners from the initial corner set. After that, a simple yet effective approach is adopted to sort the detected corners into a meaningful order. Finally, the sub‐pixel location of each corner is calculated. The proposed algorithm only requires a user input of the chessboard size, while all the other parameters can be adaptively calculated with a statistical approach. The experimental results demonstrate that the proposed method has advantages over the popular OpenCV chessboard corner detection method in terms of detection accuracy and computational efficiency. Furthermore, the effectiveness of the proposed method used for camera calibration is also verified in authors’ experiments. Yu Liu 0023, Shuping Liu, Yang Cao 0010, Zengfu Wang |
IET Image Process. | 4 |
| 2016 | Automatic tag saliency ranking for stereo images
Yang Cao 0010, Jing Zhang 0037, Zengfu Wang |
Neurocomputing | 5 |
| 2016 | Salient object detection and classification for stereoscopic images
Yang Cao 0010, Jing Zhang 0037, Zengfu Wang |
Multim. Tools Appl. | 4 |
| 2016 | Facial video coding/decoding at ultra-low bit-rate: a 2D/3D model-based approach
Jun Yu 0001, Changwei Luo, Lingyun Yu 0002, Ling-yan Li, Zengfu Wang |
Multim. Tools Appl. | 5 |
| 2016 | A Unified Scheme for Super-Resolution and Depth Estimation From Asymmetric Stereoscopic VideoabstractReconstructing a full-resolution stereoscopic video from an asymmetric stereoscopic video is a challenging task. The existing approaches require depth information, which imposes an additional challenge in data acquisition. In this paper, we propose a novel scheme that is capable of obtaining super-resolution and depth estimation simultaneously from an asymmetric stereoscopic video. The proposed scheme models the video super-resolution and stereo matching with a unified energy function. Then, we apply an alternating optimization method to minimize this energy function, which can be implemented with a two-step algorithm. In the first step we calculate the initial depth map by using a region-based cooperative optimization technique while considering the temporal consistency in video. In the second step we resolve the super-resolution problem under the guidance of the depth information. It is effective because each step benefits from the additional improvement over the previous step. We iteratively update the two steps until stable depth and super-resolution results are obtained. We have conducted a series of experiments on public stereoscopic video sequences to evaluate the performance of the proposed method. Both objective indexes and subjective visual comparisons verify that the proposed scheme can achieve satisfactory super-resolution results and high-quality depth map simultaneously. In particular, the subjective evaluation experiments on a 3-D monitor show that this scheme outperforms others and achieves the best visual sharpness. Jing Zhang 0037, Yang Cao 0010, Zhengjun Zha, Zhigang Zheng, Chang Wen Chen, Zengfu Wang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2016 | Bounds on Multiple Sensor FusionabstractWe consider the problem of fusing measurements in a sensor network, where the sensing regions overlap and data are nonnegative real numbers, possibly resulting from a count of indistinguishable discrete entities. Because of overlaps, it is generally impossible to fuse this information to arrive at an accurate value of the overall amount or count of material present in the union of the sensing regions. Here we study the computation of the range of overall values consistent with the data and provide several results. Posed as a linear programming problem, this leads to questions associated with the geometry of the sensor regions, specifically the arrangement of their nonempty intersections. We define a computational tool called the fusion polytope , based on the geometry of the sensing regions. Its properties are explored, and in particular, a topological necessary and sufficient condition for this to be in the positive orthant, a property that considerably simplifies calculations, is provided. We show that in two dimensions, inflated tiling schemes based on rectangular regions fail to satisfy this condition, whereas inflated tiling schemes based on hexagons do. William Moran 0001, Frederick R. Cohen, Zengfu Wang, Sofia Suvorova, Douglas Cochran, Tom Taylor, Peter Mark Farrell, Stephen D. Howard |
ACM Trans. Sens. Networks | 3 |
| 2016 | Optimal Cable Laying Across an Earthquake Fault Line Considering Elliptical FailuresabstractWhether it be a road, railway track, pipeline, or fiber-optic cable, achieving a low probability of “cable” damage is important for the proper operation of modern infrastructure. This paper considers the fundamental problem of how to connect two points with a cable that crosses an earthquake fault line. We first develop a model, under certain general assumptions, for the cable break probability. Then, we formulate a multiobjective optimization problem, with cable cost and probability of cable break being the two objectives. For two important and meaningful sets of cable shape alternatives, the Pareto front is determined for these two objectives. All the analytical results are verified by simulations. Zengfu Wang, Moshe Zukerman, Jonathan H. Manton, Alain Bensoussan 0001, Yu Wang 0042 |
IEEE Trans. Reliab. | 2 |
| 2015 | Video Based Face Tracking and Animation
Changwei Luo, Jun Yu 0001, Zhigang Zheng, Lingyun Yu 0002, Zengfu Wang |
ICIG (3) | 6 |
| 2015 | Real-Time Robust Video Stabilization Based on Empirical Mode Decomposition and Multiple Evaluation Criteria
Jun Yu 0001, Changwei Luo, Chen Jiang 0002, Rui Li 0021, Ling-yan Li, Zengfu Wang |
ICIG (3) | 6 |
| 2015 | Simultaneous image fusion and denoising with adaptive sparse representationabstractIn this study, a novel adaptive sparse representation (ASR) model is presented for simultaneous image fusion and denoising. As a powerful signal modelling technique, sparse representation (SR) has been successfully employed in many image processing applications such as denoising and fusion. In traditional SR‐based applications, a highly redundant dictionary is always needed to satisfy signal reconstruction requirement since the structures vary significantly across different image patches. However, it may result in potential visual artefacts as well as high computational cost. In the proposed ASR model, instead of learning a single redundant dictionary, a set of more compact sub‐dictionaries are learned from numerous high‐quality image patches which have been pre‐classified into several corresponding categories based on their gradient information. At the fusion and denoising processes, one of the sub‐dictionaries is adaptively selected for a given set of source image patches. Experimental results on multi‐focus and multi‐modal image sets demonstrate that the ASR‐based fusion method can outperform the conventional SR‐based method in terms of both visual quality and objective assessment. Yu Liu 0023, Zengfu Wang |
IET Image Process. | 2 |
| 2015 | Depth map super-resolution using stereo-vision-assisted model
Yuxiang Yang 0001, Mingyu Gao 0002, Jing Zhang 0037, Zhengjun Zha, Zengfu Wang |
Neurocomputing | 5 |
| 2015 | Dense SIFT for ghost-free multi-exposure fusion
Yu Liu 0023, Zengfu Wang |
J. Vis. Commun. Image Represent. | 2 |
| 2015 | Locating Facial Landmarks Using Probabilistic Random ForestabstractRandom forest is a useful tool for face alignment/tracking. The method of regressing local binary features learned from random forest has achieved state-of-the-art performance both in fitting accuracy and speed. Despite the great success of this method, it has certain weaknesses: the number of available local binary features is rather limited and is not optimal for face alignment; the binary features inevitably lead to serious jitter when tracking a video sequence. To address these problems, we propose learning probability features from probabilistic random forest (PRF). The proposed PRF is the same as standard random forest except that it models the probability of a sample belonging to the nodes of a tree. By using the probability features, our method significantly outperforms the state-of-the-art in terms of accuracy. It also achieves about 60 fps for locating a few facial landmarks. In addition, our method shows excellent stability in face tracking. Changwei Luo, Zengfu Wang, Shaobiao Wang, Juyong Zhang, Jun Yu 0001 |
IEEE Signal Process. Lett. | 2 |
| 2015 | A Video, Text, and Speech-Driven Realistic 3-D Virtual Head for Human-Machine InterfaceabstractA multiple inputs-driven realistic facial animation system based on 3-D virtual head for human-machine interface is proposed. The system can be driven independently by video, text, and speech, thus can interact with humans through diverse interfaces. The combination of parameterized model and muscular model is used to obtain a tradeoff between computational efficiency and high realism of 3-D facial animation. The online appearance model is used to track 3-D facial motion from video in the framework of particle filtering, and multiple measurements, i.e., pixel color value of input image and Gabor wavelet coefficient of illumination ratio image, are infused to reduce the influence of lighting and person dependence for the construction of online appearance model. The tri-phone model is used to reduce the computational consumption of visual co-articulation in speech synchronized viseme synthesis without sacrificing any performance. The objective and subjective experiments show that the system is suitable for human-machine interaction. Jun Yu 0001, Zengfu Wang |
IEEE Trans. Cybern. | 2 |
| 2014 | A distributed expectation-maximization algorithm for OTHR multipath target tracking
Hua Lan, Yan Liang 0001, Zengfu Wang, Feng Yang 0001, Quan Pan 0001 |
FUSION | 3 |
| 2014 | Synthesizing real-time speech-driven facial animationabstractWe present a real-time speech-driven facial animation system. In this system, Gaussian Mixture Models (GMM) are employed to perform the audio-to-visual conversion. The conventional GMM-based method performs the conversion frame by frame using minimum mean square error (MMSE) estimation. The method is reasonably effective. However, discontinuities often appear in the sequences of estimated visual features. To solve this problem, we incorporate previous visual features into the conversion so that the conversion procedure is performed in the manner of a Markov chain. After audio-to-visual conversion, the estimated visual features are transformed to blendshape weights to synthesize facial animation. Experiments show that our system can accurately convert audio features into visual features. The conversion accuracy is comparable to a current state-of-the-art trajectory-based approach. Moreover, our system runs in real time and outputs high quality lip-sync animations. Changwei Luo, Jun Yu 0001, Zengfu Wang |
ICASSP | 3 |
| 2014 | A new image filtering method: Nonlocal image guided averagingabstractImage guided filtering has been widely used in many image processing applications. However, it is a local filtering method and has limited propagation ability. In this paper, we propose a new image filtering method: nonlocal image guided averaging (NLGA). Derived from a nonlocal linear model, the proposed method can utilize the nonlocal similarity of the guidance image, so that it can propagate nonlocal information reliably. Consequently, NLGA can obtain a sharper filtering results in the edge regions and more smooth results in the smooth regions. It shows superiority over image guided filtering in different applications, such as image dehazing, depth map super-resolution and image denoising. Jing Zhang 0037, Yang Cao 0010, Zengfu Wang |
ICASSP | 3 |
| 2014 | A practical algorithm for automatic chessboard corner detectionabstractChessboard corner detection is a fundamental work of the popular chessboard pattern-based camera calibration technique. In this paper, a fast and robust algorithm for chessboard corner detection is presented. In our method, an initial corner set is obtained with an improved Hessian corner detector. And then, a novel strategy which takes both textural and geometrical characteristics of a chessboard into consideration is employed to eliminate fake corners in the initial corner set. The proposed algorithm only requires a user-input of the total number of chessboard inner corners, while all the other parameters can be adaptively calculated with a statistical approach. Experimental results on two public data sets demonstrate that the proposed method can outperform the most commonly used OpenCV method in terms of both detection rate and computational efficiency. Yu Liu 0023, Shuping Liu, Yang Cao 0010, Zengfu Wang |
ICIP | 4 |
| 2014 | Expressive facial animation from videosabstractWe address the issue of synthesizing real-time expressive facial animation from videos. Given video footage of a person's face, some existing methods track a few facial landmarks to drive the virtual character. Compared with these methods, ours has the following characteristics. 1) We incorporate global texture into a constrained local model to increase the accuracy of facial tracking. 2) To animate a blenshape face model, facial tracking results as well as facial expression recognition results are used to estimate blendshape weights. Experimental results demonstrate that our method is effective for producing realistic expressive facial animations. Moreover, the method does not require facial markers or complex offline pre-processing, these properties make it very easy to use for ordinary users. Changwei Luo, Chen Jiang 0002, Jun Yu 0001, Zengfu Wang |
ICIP | 4 |
| 2014 | Nighttime haze removal based on a new imaging modelabstractNighttime haze removal is important for different applications such as nighttime video surveillance in haze environment. Different from the imaging conditions in the daytime, nighttime haze images may suffer from non-uniform illumination from artificial light sources. In this paper, we proposed a novel efficient dehazing method with illumination estimation for nighttime haze condition. First, we estimate the light intensity and enhance it to obtain an illumination balanced result. Then, we process a color correction step after estimating the color characteristics of the incident light. Finally, we remove the haze by using the dark channel prior along with estimating the pointwise environmental light. Experimental results show that the proposed method can achieve both illumination balanced and haze free results. Moreover, it also has good color rendition ability. Jing Zhang 0037, Yang Cao 0010, Zengfu Wang |
ICIP | 3 |
| 2014 | Real-time control of 3D facial animationabstractFacial animation is useful in human-machine interaction, computer games and teleconferences. We propose a realtime performance-driven facial animation system for ordinary users. The system enables a user to animate an avatar by performing desired facial motions in front of a video camera. First, a constrained local model based approach is used to track facial features of a performer in the video. To increase the tracking accuracy, we propose an efficient method to build a user-specific local texture model. Next, a 3D blendshape face model is fitted to the tracked feature points. To improve the expressiveness of synthesized animations, facial expression recognition results and pre-recorded animation priors are incorporated into the fitting procedure. Finally, facial animations are created using blendshape interpolation. Experiments show that the synthetic facial motions are realistic and quite similar to the facial actions of the performer. By using an ordinary camera, our system provides the user complete control over the generated facial animations. Changwei Luo, Jun Yu 0001, Chen Jiang 0002, Rui Li 0021, Zengfu Wang |
ICME | 5 |
| 2014 | External Vision based Robust Pose Estimation System for a Quadrotor in Outdoor EnvironmentsabstractIn this paper, an external vision based robust pose estimation system for a quadrotor in outdoor environment has been proposed. This system can provide us with approximate ground truth of a quadrotor outdoors, while most of external vision based systems perform indoors. Here, we do not modify the architecture of the quadrotor or put colored blobs, LEDs on it. We only use the own features of the quadrotor. We implement a pure vision-based approach to get preliminary position of a flying quadrotor. Then only using the own features of the quadrotor, we present a novel robust pose estimation algorithm to get the accurate pose of a quadrotor. With good observed results, we get all the four rotors and calculate the pose. But when fewer than four rotors are observed, all of other external vision based systems of the quadrotor did not mention this and can not get right pose results. In this paper, we have solved this problem and got accurate pose estimation with IMU data. This system can provide us with approximate ground truth outdoors. We demonstrate in real experiments that the vision-based pose estimation system for outdoor environment can perform accurately and robustly in real time. Zengfu Wang |
ICPRAM | 3 |
| 2014 | Adaptive Noise Variance Identification in Vision-aided Motion Estimation for UAVsabstractVision location methods have been widely used in the motion estimation of unmanned aerial vehicles (UAVs).
The noise of the vision location result is usually modeled as the white gaussian noise so that this result could
be utilized as the observation vector in the kalman filter to estimate the motion of the vehicle. Since the
noise of the vision location result is affected by external environments, the variance of the noise is uncertain.
However, in previous researches the variance is usually set as a fixed empirical value, which will lower the
accuracy of the motion estimation. In this paper, a novel adaptive noise variance identification (ANVI) method
is proposed, which utilizes the special kinematic property of the UAV for frequency analysis and adaptively
identify the variance of the noise. Then, the adaptively identified variance are used in the kalman filter for
accurate motion estimation. The performance of the proposed method is assessed by simulations and field
experiments on a quadrotor system. The results illustrate the effectiveness of the method. Zengfu Wang |
ICPRAM | 3 |
| 2014 | 3D facial motion tracking by combining online appearance model and cylinder head model in particle filtering
Jun Yu 0001, Zengfu Wang |
Sci. China Inf. Sci. | 2 |
| 2014 | Background segmentation of dynamic scenes based on dual modelabstractDetecting moving objects from background in video sequences is the first step of many image applications. The background can be divided into two types according to whether the pixel values of it are variable or not: static one and dynamic one. How to correctly detect moving foreground objects from dynamic scenes is a difficult problem because of the similarity between the moving foreground and the variable background. In this study, a new method for non‐parametric background segmentation of dynamic scenes is proposed. Here the background is described by two interrelated models. One of them is called the self‐model, which concerns with the recently observed pixel values at the same position, and the other one is called the neighbourhood‐model, which is described by the pixel values of the neighbourhood. The author's method can accurately detect the dynamic background. To correctly detect pixels in the foreground as much as possible, the authors also propose an adaptive threshold for foreground decision based on the background characteristics. All of the above detection processes can be done in real time. Experimental results on public dataset demonstrate that the proposed method outperforms the state‐of‐the‐art for background segmentation in dynamic scenes. Jing Zhang 0037, Zengfu Wang |
IET Comput. Vis. | 3 |
| 2014 | Globality and locality incorporation in distance metric learning
Wei Wang 0061, Bao-Gang Hu, Zengfu Wang |
Neurocomputing | 3 |
| 2014 | A new closed loop method of super-resolution for multi-view images
Jing Zhang 0037, Yang Cao 0010, Zhigang Zheng, Chang Wen Chen, Zengfu Wang |
Mach. Vis. Appl. | 5 |
| 2013 | Joint estimation of target state and ionosphere state of over-the-horizon radar
Yan Liang 0001, Zengfu Wang, Feng Yang 0001 |
FUSION | 3 |
| 2013 | Efficient and Scalable Information Geometry Metric LearningabstractInformation Geometry Metric Learning (IGML) is shown to be an effective algorithm for distance metric learning. In this paper, we attempt to alleviate two limitations of IGML: (A) the time complexity of IGML increases rapidly for high dimensional data, (B) IGML has to transform the input low rank kernel into a full-rank one since it is undefined for singular matrices. To this end, two novel algorithms, referred to as Efficient Information Geometry Metric Learning (EIGML) and Scalable Information Geometry Metric Learning (SIGML), are proposed. EIGML scales linearly with the dimensionality, resulting in significantly reduced computational complexity. As for SIGML, it is proven to have a range-space preserving property. Following this property, SIGML is found to be capable of handling both full-rank and low-rank kernels. Additionally, the geometric information from data is further exploited in SIGML. In contrast to most existing metric learning methods, both EIGML and SIGML have closed-form solutions and can be efficiently optimized. Experimental results on various data sets demonstrate that the proposed methods outperform the state-of-the-art metric learning algorithms. Wei Wang 0061, Bao-Gang Hu, Zengfu Wang |
ICDM | 3 |
| 2013 | Multi-focus Image Fusion Based on Sparse Representation with Adaptive Sparse Domain SelectionabstractSparse representation (SR) has been widely used in many image processing applications including image fusion. As the contents vary significantly across different images, a highly redundant dictionary is always required in the sparse model, which reduces the algorithm stability and efficiency. This paper proposes a multi-focus image fusion method based on SR with adaptive sparse domain selection (SR-ASDS). Under SR-ASDS, numerous high-quality image patches are first classified into several categories according to their gradient information, and each category is applied into training a compact sub-dictionary. At the fusion process, a corresponding sub-dictionary is adaptively selected for a given pair of source image patches. Moreover, we present a general optimization framework for the merging rule design of the SR based image fusion. Numerous experiments on both clear images and the noisy ones demonstrate that the proposed method outperforms the fusion methods which use a single dictionary, in terms of several popular objective evaluation criteria. Yu Liu 0023, Zengfu Wang |
ICIG | 2 |
| 2013 | A simultaneous method for 3D video super-resolution and high-quality depth estimationabstractMixed-resolution approach serves as a feasible solution to 3D video data reduction in limited bandwidth network environments, i.e., mobile networks. In this paper, we propose a simultaneous method for video super-resolution and high-quality depth estimation of mixed-resolution 3D video. Our method tackles the problem in a joint manner: i). Depth Estimation. We calculate the initial depths by stereo matching, and then warp them according to the optical flow field. ii). Video Super-resolution. Under the guidance of the warped depth information, we resolve the super-resolution problem by a fusion method, which involves a mapping step and a nonlocal reconstruction step. We run the above two steps iteratively until obtaining the stable depth and super-resolution result. The experimental results on public 3D video sequences verify the effectiveness of our proposed method. Jing Zhang 0037, Yang Cao 0010, Zengfu Wang |
ICIP | 3 |
| 2013 | 2D/3D Model-Based Facial Video Coding/Decoding at Ultra-Low Bit-Rate
Jun Yu 0001, Zengfu Wang, Yang Cao 0010 |
MMM (2) | 2 |
| 2013 | A New Closed Loop Method of Super-Resolution for Multi-view Images
Jing Zhang 0037, Yang Cao 0010, Zhigang Zheng, Zengfu Wang |
MMM (2) | 4 |
| 2013 | Interactive social group recommendation for Flickr photos
Zhengjun Zha, Qi Tian 0001, Zengfu Wang |
Neurocomputing | 4 |
| 2013 | Digital Multi-Focusing From a Single Photograph Taken With an Uncalibrated Conventional CameraabstractThe demand to restore all-in-focus images from defocused images and produce photographs focused at different depths is emerging in more and more cases, such as low-end hand-held cameras and surveillance cameras. In this paper, we manage to solve this challenging multi-focusing problem with a single image taken with an uncalibrated conventional camera. Different from all existing multi-focusing approaches, our method does not need to include a deconvolution process, which is quite time-consuming and will cause ringing artifacts in the focused region and low depth-of-field. This paper proposes a novel systematic approach to realize multi-focusing from a single photograph. First of all, with the optical explanation for the local smooth assumption, we present a new point-to-point defocus model. Next, the blur map of the input image, which reflects the amount of defocus blur at each pixel in the image, is estimated by two steps. 1) With the sharp edge prior, a rough blur map is obtained by estimating the blur amount at the edge regions. 2) The guided image filter is applied to propagate the blur value from the edge regions to the whole image by which a refined blur map is obtained. Thus far, we can restore the all-in-focus photograph from a defocused input. To further produce photographs focused at different depths, the depth map from the blur map must be derived. To eliminate the ambiguity over the focal plane, user interaction is introduced and a binary graph cut algorithm is used. So we introduce user interaction and use a binary graph cut algorithm to eliminate the ambiguity over the focal plane. Coupled with the camera parameters, this approach produces images focused at different depths. The performance of this new multi-focusing algorithm is evaluated both objectively and subjectively by various test images. Both results demonstrate that this algorithm produces high quality depth maps and multi-focusing results, outperforming the previous approaches. Yang Cao 0010, Shuai Fang, Zengfu Wang |
IEEE Trans. Image Process. | 3 |
| 2012 | Enhanced OTHR detection using Bayesian fusion of multipath target returns
Zengfu Wang, Xuezhi Wang 0001, Yan Liang 0001, Quan Pan 0001 |
FUSION | 1 |
| 2012 | Discriminating classes collapsing for Globality and Locality Preserving ProjectionsabstractIn this paper, a novel approach, namely Globality and Locality Preserving Projections (GLPP), is proposed in the study of dimensionality reduction. The method is designed to combine the ideas behind Locality Preserving (LP), Discriminating Power (DP) and Maximally Collapsing Metric Learning (MCML), resulting in a unified model. Several distinguished features are obtained from the integration design. First, the method is able to take into account both global and local information of the data set. We introduce a new formula for calculating the conditional probabilities, which can remove the locality distortions from MCML. Second, discrimination information is applied so that a projection matrix is formed which can collapse all data points of the same class closer together, while pushing points of different classes further away. Third, the proposed method guarantees a supervised convex algorithm, which is a critical feature in data processing. Furthermore on this concern, GLPP is mapped to a Graphics Processor Unit (GPU) architecture in the implementation to be appropriate for large scale data sets. Several numerical studies are conducted on a variety of data sets. The numerical results confirm that GLPP consistently outperforms most up-to-date methods, allowing high classification accuracy, good visualization and sharply decreased consuming time. Wei Wang 0061, Bao-Gang Hu, Zengfu Wang |
IJCNN | 3 |
| 2012 | A comprehensive representation scheme for video semantic ontology and its applications in semantic concept detection
Zhengjun Zha, Tao Mei 0001, Yantao Zheng, Zengfu Wang, Xian-Sheng Hua 0001 |
Neurocomputing | 4 |
| 2011 | A New Image Super-Resolution Method in the Wavelet DomainabstractIn this paper, we propose a novel single image super-resolution method based on MAP statistical reconstruction. Our approach takes the wavelet domain Implicit Markov Random Field (IMRF) model as the prior constraint. Specifically, we apply the IMRF model to show the probability distribution of natural images in wavelet domain, and utilize the MAP theory of Bayesian estimation to construct the objective function with this model. Furthermore, we employ steepest descent method to optimize this objective function. In the experiments, we use both Peak Signal to Noise Ratio (PSNR) and visual effect to evaluate our method. The experimental results demonstrate that our method obtains the superior performance in comparison with traditional single image super-resolution approaches. Zengfu Wang |
ICIG | 2 |
| 2011 | Semi-automatic Flickr Group Suggestion
Zhengjun Zha, Qi Tian 0001, Zengfu Wang |
MMM (2) | 4 |
| 2010 | A Close-Form Iterative Algorithm for Depth Inferring from a Single Image
Yang Cao 0010, Zengfu Wang |
ECCV (5) | 3 |
| 2010 | Evaluation of histogram based interest point detector in web image classification and searchabstractLocal image feature has received increasing attention in various applications, such as web image classification and search. The process of local feature extraction consists of two main steps: interest point detection and local feature description. A wealth of interest point detectors have been proposed in last decades. Most of them measure pixel-wise differences in image intensity or color. Recently, a new type of interest point detector has been developed, which incorporates histogram-based representation into the process of interest point detection. In this paper, we evaluate this histogram-based interest point detector in the context of web image classification and search, as well as compare it against typical pixel-based detectors and heuristic grid-based detector. The evaluation is performed on two web image datasets: NUS-WIDE-OBJECT and MIRFLICKR-25000 datasets. The experimental results demonstrate that the histogram-based interest point detector outperforms the pixel-based and grid-based detectors in both web image classification and search tasks. Zhengjun Zha, Yinghai Zhao, Zengfu Wang |
ICME | 4 |
| 2010 | Model-assisted face reconstruction based on binocular stereoabstractIn this paper we propose a model-assisted binocular stereo algorithm for 3D face reconstruction. First, stereo correspondence is reliably performed between input stereo images by employing a 3D reference face model as a medium. Then, a high-quality dataset of point cloud is generated by triangulation from the stereo correspondence. Finally, an accurate and feature-preserving 3D face surface is reconstructed from the point cloud dataset by a denoising operation of bilateral filtering and surface meshing. We have compared the reconstructed faces with ground truth 3D data. The experimental results show that the proposed algorithm is effective with the average error of 2.6mm. Zengfu Wang |
VCIP | 3 |
| 2010 | 3D Modeling of Faces from Near Infrared Images Using Statistical LearningabstractThis paper proposes a statistical learning based method for 3D modeling of faces directly from Near Infrared (NIR) images. We use a specially designed camera system with active NIR illumination to capture the NIR images of faces. The NIR images captured in such a way are invariant to environmental lighting changes. The property provides more reasonable data sources for statistical learning. By using the NIR images and the depth images of some known faces, we can observe a mapping relation between the two image modalities. The mapping relation can then be used to recover depth data of an unknown face from his NIR image. To perform the learning, the images of different modalities taken from different persons are elaborately aligned to make pixel-to-pixel correspondences between images. Based on these aligned images, two face spaces corresponding to NIR and depth face images can be constructed, respectively. We then use a PCA based or kernel based scheme to perform the learning between spaces of large dimensions. Several regression algorithms with linear and nonlinear kernels are employed and evaluated to find the mapping that best describes the relation between the two face spaces. The experimental results show that the method presented in this paper is effective. It can reconstruct 3D face model directly from NIR image of a face with high accuracy and low computational costs. Stan Z. Li, Jianglong Chang, Zengfu Wang |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2010 | Visual query suggestion: Towards capturing user intent in internet image searchabstractQuery suggestion is an effective approach to bridge the Intention Gap between the users' search intents and queries. Most existing search engines are able to automatically suggest a list of textual query terms based on users' current query input, which can be called Textual Query Suggestion. This article proposes a new query suggestion scheme named Visual Query Suggestion (VQS) which is dedicated to image search. VQS provides a more effective query interface to help users to precisely express their search intents by joint text and image suggestions. When a user submits a textual query, VQS first provides a list of suggestions, each containing a keyword and a collection of representative images in a dropdown menu. Once the user selects one of the suggestions, the corresponding keyword will be added to complement the initial query as the new textual query, while the image collection will be used as the visual query to further represent the search intent. VQS then performs image search based on the new textual query using text search techniques, as well as content-based visual retrieval to refine the search results by using the corresponding images as query examples. We compare VQS against three popular image search engines, and show that VQS outperforms these engines in terms of both the quality of query suggestion and the search performance. Zhengjun Zha, Linjun Yang, Tao Mei 0001, Meng Wang 0001, Zengfu Wang, Tat-Seng Chua, Xian-Sheng Hua 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2009 | A Robust Algorithm for Color Correction between Two Stereo Images
Zengfu Wang |
ACCV (2) | 3 |
| 2009 | A Subjective Method for Image Segmentation Evaluation
Zengfu Wang |
ACCV (3) | 2 |
| 2009 | Robust Distance Metric Learning with Auxiliary Knowledge
Zhengjun Zha, Tao Mei 0001, Meng Wang 0001, Zengfu Wang, Xian-Sheng Hua 0001 |
IJCAI | 4 |
| 2009 | Visual query suggestionabstractQuery suggestion is an effective approach to improve the usability of image search. Most existing search engines are able to automatically suggest a list of textual query terms based on users' current query input, which can be called Textual Query Suggestion. This paper proposes a new query suggestion scheme named Visual Query Suggestion (VQS) which is dedicated to image search. It provides a more effective query interface to formulate an intent-specific query by joint text and image suggestions. We show that VQS is able to more precisely and more quickly help users specify and deliver their search intents. When a user submits a text query, VQS first provides a list of suggestions, each containing a keyword and a collection of representative images in a dropdown menu. If the user selects one of the suggestions, the corresponding keyword will be added to complement the initial text query as the new text query, while the image collection will be formulated as the visual query. VQS then performs image search based on the new text query using text search techniques, as well as content-based visual retrieval to refine the search results by using the corresponding images as query examples. We compare VQS with three popular image search engines, and show that VQS outperforms these engines in terms of both the quality of query suggestion and search performance. Zhengjun Zha, Linjun Yang, Tao Mei 0001, Meng Wang 0001, Zengfu Wang |
ACM Multimedia | 5 |
| 2009 | Graph-based semi-supervised learning with multiple labels
Zhengjun Zha, Tao Mei 0001, Jingdong Wang 0001, Zengfu Wang, Xian-Sheng Hua 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2008 | Joint multi-label multi-instance learning for image classificationabstractIn real world, an image is usually associated with multiple labels which are characterized by different regions in the image. Thus image classification is naturally posed as both a multi-label learning and multi-instance learning problem. Different from existing research which has considered these two problems separately, we propose an integrated multi-label multi-instance learning (MLMIL) approach based on hidden conditional random fields (HCRFs), which simultaneously captures both the connections between semantic labels and regions, and the correlations among the labels in a single formulation. We apply this MLMIL framework to image classification and report superior performance compared to key existing approaches over the MSR Cambridge (MSRC) and Corel data sets. Zhengjun Zha, Xian-Sheng Hua 0001, Tao Mei 0001, Jingdong Wang 0001, Guo-Jun Qi, Zengfu Wang |
CVPR | 6 |
| 2008 | Eye states detection by boosting Local Binary Pattern Histogram featuresabstractIn this paper, we propose a novel method for eye states detection. The detection of eye states is treated as an appearance based binary classification problem. The whole eye region is first scanned by a series of blocks with various locations and scales. The Local Binary Pattern Histogram (LBPH) is then extracted from each block to form a descriptor of the local texture. A reference template for each block is later calculated as the optimal histogram which makes the distances between it and the LBPHs of different clusters most separable. For all the blocks, the bin-wise distances between local LBPHs and the corresponding reference templates are then extracted to form a feature set for classification. A cascaded AdaBoost learning is further employed to select most discriminative features from the whole feature set and to accelerate the classification. Experimental results demonstrate that our eye states detection algorithm and can give correct and robust detection results in real-time. Cui Xu, Zengfu Wang |
ICIP | 3 |
| 2008 | Robust depth estimation for efficient 3D face reconstructionabstractThis paper proposes a learning based framework for efficient 3D face reconstruction. We transfer the 3D reconstruction into a statistical learning problem of finding appropriate mapping between texture and depth subspaces. Instead of using grayscales to directly estimate the depth, we use local binary pattern (LBP) to further encode the face texture, providing robustness for depth estimation under different illumination conditions. Then the high dimension learning problem between face subspaces is tackled by the kernel partial least squares (PLS) regression. The experimental results show that the proposed method can reconstruct 3D face from single frontal image efficiently and robustly. Zengfu Wang |
ICIP | 2 |
| 2008 | Graph-based semi-supervised learning with multi-labelabstractConventional graph-based semi-supervised learning methods predominantly focus on single label problem. However, it is more popular in real-world applications that an example is associated with multiple labels simultaneously. In this paper, we propose a novel graph-based learning framework in the setting of semi-supervised learning with multi-label. The proposed approach is characterized by simultaneously exploiting the inherent correlations among multiple labels and the label consistency over the graph. We apply the proposed framework to video annotation and report superior performance compared to key existing approaches over the TRECVID 2006 corpus. Zhengjun Zha, Tao Mei 0001, Jingdong Wang 0001, Zengfu Wang, Xian-Sheng Hua 0001 |
ICME | 4 |
| 2008 | Semantic feature extraction for accurate eye corner detectionabstractIn this paper a novel eye corner (canthus) detection method is proposed. The method is based on sematic features which are extracted from the structure and appearance characteristics of canthus. The eyelids are first fitted to construct an angle model. Based on this model, one feature is proposed to characterize the appearance difference between inner and outer canthus regions, and another feature is proposed to emphasize the role of the canthus angle bisector region. The two features are fused in a logistic regression classifier to detect accurate canthus. The effectiveness and accuracy of the proposed method are demonstrated in experiments. Cui Xu, Zengfu Wang |
ICPR | 3 |
| 2008 | Robust and precise eye detection based on locally selective projectionabstractThis paper proposes a robust and precise eye detection method based on a new projection algorithm called locally selective projection (LSP). Along each projection axis, LSP selects a pixel and uses a function calculated in the neighborhood of the pixel as response. The local selectivity of LSP makes it robust against rotation, illumination and occlusion. Moreover, the positions of selected pixels can be recorded, providing a 2D cue for image analysis. We apply LSP to eye detection. The vertical and horizontal LSPs are first used to give reliable eye candidates, then a SVM classifier is employed to verify the real eye pairs. The experiment result compared with an AdaBoost detector shows the robustness and accuracy of proposed method. Zengfu Wang |
ICPR | 2 |
| 2007 | Pattern classification with a PSO optimization based elliptical basis function neural networksabstractIn this paper, a novel model of elliptical basis function neural networks (EBFNN) based on a hybrid optimization algorithm is proposed. Firstly, a geometry analytic algorithm is applied to construct the hyper-ellipsoid units of hidden layer of the EBFNN, i.e., an initial structure of the EBFNN, which is further pruned by the particle swarm optimization (PSO) algorithm. And the shape parameters of kernel function for the hidden layer are also optimized by the PSO simultaneously. Finally, the hybrid learning algorithm (HLA) is further applied to adjust the hidden centers and the shape parameters of kernel function for the hidden layer. The experimental results demonstrated the proposed hybrid optimization algorithm for the EBFNN model is feasible and efficient, and the EBFNN is not only parsimonious but also has better generalization performance than the RBFNN. Jixiang Du, De-Shuang Huang, Zengfu Wang |
IEEE Congress on Evolutionary Computation | 3 |
| 2007 | Estimation of Markov Jump systems with mode observation one-step lagged to state measurementabstractThe estimation of Markov jump systems (MJS) is widely used in target tracking, fault detection, signal processing and digital communications. However, the above researches all assume that state measurement and additional mode observation are synchronous which means both state measurement and mode observation at each sampling time arrive at the fusion centre at the same time. The problem of estimation of MJS that mode observation is one-step lagged to its corresponding state measurement is considered. Along state-augmentation approach and the derivation of image-enhanced interacting multiple model (IE-IMM), a new generic estimation algorithm is proposed. It is shown by simulation result that the proposed algorithm is effective. Yan Liang 0001, Zengfu Wang, Yongmei Cheng, Quan Pan 0001 |
FUSION | 2 |
| 2007 | 3D Face Reconstruction from Stereo: A Model Based ApproachabstractSince the human faces are lowly textured, the conventional stereo methods based on intensity correlation can not give satisfying 3D face reconstruction results. In this paper, a model based stereo matching method is proposed. A reference 3D face is used as an intermedium for correspondence calculation. The virtual face images with known correspondences are first synthesized from the reference face. Then the known correspondences are extended to the incoming stereo face images, using face alignment and warping. The complete 3D face can thus be reconstructed from stereo images reliably. Jianglong Chang, Zhigang Zheng, Zengfu Wang |
ICIP (3) | 4 |
| 2007 | A speech rate related lip movement model for speech animation
Zengfu Wang |
INTERSPEECH | 2 |
| 2007 | A Novel Elliptical Basis Function Neural Networks Model Based on a Hybrid Learning Algorithm
Jixiang Du, Guo-Jun Zhang, Zengfu Wang |
ISNN (1) | 3 |
| 2007 | Refining video annotation by exploiting pairwise concurrent relationabstractVideo annotation is a promising and essential step for content-based video search and retrieval. Most of the state-of-the-art video annotation approaches detect multiple semantic concepts in an isolated manner, which neglect the fact that video concepts are usually correlated in semantic nature. In this paper, we propose to refine video annotation by leveraging the pairwise concurrent relation among video concepts. Such concurrent relation is explicitly modeled by a concurrent matrix and then a propagation strategy is adopted to refine the annotations. Through spreading the scores of all related concepts to each other iteratively, the detection results approach stable and optimal. In contrast with existing concept fusion methods, the proposed approach is computationally more efficient and easy to implement, not requiring to construct any contextual model. Furthermore, we show its intuitive connection with the PageRank algorithm. We conduct the experiments on TRECVID 2005 corpus and report superior performance compared to existing key approaches. Zhengjun Zha, Tao Mei 0001, Xian-Sheng Hua 0001, Guo-Jun Qi, Zengfu Wang |
ACM Multimedia | 5 |
| 2006 | A Novel Elliptical Basis Function Neural Networks Optimized by Particle Swarm Optimization
Jixiang Du, Chuan-Min Zhai, Zengfu Wang, Guo-Jun Zhang |
ISNN (1) | 3 |
| 2006 | A novel full structure optimization algorithm for radial basis probabilistic neural networks
Jixiang Du, De-Shuang Huang, Guo-Jun Zhang, Zengfu Wang |
Neurocomputing | 4 |
| 2005 | An Emotion Space Model for Recognition of Emotions in Spoken Chinese
Xuecheng Jin, Zengfu Wang |
ACII | 2 |
| 2005 | An Approach to Affective-Tone Modeling for Mandarin
Zhuangluan Su, Zengfu Wang |
ACII | 2 |
| 2005 | Globally Attractive Periodic State of Discrete-Time Cellular Neural Networks with Time-Varying Delays
Zhigang Zeng, Boshan Chen, Zengfu Wang |
ISNN (1) | 3 |
| 2005 | Global Stability of a General Class of Discrete-Time Recurrent Neural Networks
Zhigang Zeng, De-Shuang Huang, Zengfu Wang |
Neural Process. Lett. | 3 |
| 2004 | Global exponential stability of delayed Cohen-Grossberg neural networksabstractIn this paper, using a fixed-point theorem and reduction to absurdity, the authors have obtained some sufficient conditions to guarantee that Cohen-Grossberg neural networks with discrete and distributed delays are globally exponentially stable. Since the model is more general and the assumptions relax the previous assumptions in some existing works, the results presented in this paper are the improvement and extension of the existed ones. Finally, the validity and performance of the results are illustrated by two simulation examples. Zhigang Zeng, Zengfu Wang, De-Shuang Huang |
ICARCV | 2 |
| 2004 | Image representation for stereo: stripes and stripe adjacency graphabstractA new algorithm for stereo correspondence and surface reconstruction is presented in this paper. We advance the stripe as the matching primitive. The stripe is a special kind of region composed of some adjacent similar scanline segments. Each input image is segmented into stripes and then converted to a stripe adjacency graph. Our method matches stripes and stripe adjacencies globally and adaptively, with stripes being merged or split according to the disparity estimates. A pair of matched stripes constructs a surface patch, and, correspondingly, a pair of matched subgraph presents a smooth surface. Experimental results show that our algorithm is fast and effective. Changchang Wu, Zengfu Wang |
ICIG | 2 |
| 2004 | Global Convergence of Steepest Descent for Quadratic Functions
Zhigang Zeng, De-Shuang Huang, Zengfu Wang |
IDEAL | 3 |
| 2004 | Stability Analysis of Discrete-Time Cellular Neural Networks
Zhigang Zeng, De-Shuang Huang, Zengfu Wang |
ISNN (1) | 3 |
| 2004 | Pattern Recognition Based on Stability of Discrete Time Cellular Neural Networks
Zhigang Zeng, De-Shuang Huang, Zengfu Wang |
ISNN (1) | 3 |
| 2004 | Attractability and location of equilibrium point of cellular neural networks with time-varying delaysabstractThis paper presents new theoretical results on global exponential stability of cellular neural networks with time-varying delays. The stability conditions depend on external inputs, connection weights and delays of cellular neural networks. Using these results, global exponential stability of cellular neural networks can be derived, and the estimate for location of equilibrium point can also be obtained. Finally, the simulating results demonstrate the validity and feasibility of our proposed approach. Zhigang Zeng, De-Shuang Huang, Zengfu Wang |
Int. J. Neural Syst. | 3 |