VLDB 2026 Research / reviewers in the wild / expert
Jianmin Jiang
dblp:13/1729
· DBLP profile ↗
198ranked-venue papers
28as first author
38since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 95 · 11 first-author · 22 since 2021Artificial intelligence and machine learning · 73 · 4 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 2 first-authorSoftware engineering, systems software and programming languages · 11 · 5 first-author · 2 since 2021Databases, data management, data science and information retrieval · 7Systems, architecture and hardware · 4 · 2 first-author · 1 since 2021Computer networks · 4Security and privacy · 2Human-computer interaction and ubiquitous computing · 2 · 1 first-authorTheory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RPA-former: A relative position-aware transformer with temporal tokenization for motor imagery EEG decoding
Siraj Khan, Ahmed Salem 0005, Jianmin Jiang |
Neurocomputing | 3 |
| 2026 | Efficient 3D Surface Super-Resolution via Normal-Based Multimodal RestorationabstractHigh-fidelity 3D surface is essential for vision tasks across various domains such as medical imaging, cultural heritage preservation, quality inspection, virtual reality, and autonomous navigation. However, the intricate nature of 3D data representations poses significant challenges in restoring diverse 3D surfaces while capturing fine-grained geometric details at a low cost. This paper introduces an efficient multimodal normal-based 3D surface super-resolution (mn3DSSR) framework, designed to address the challenges of microgeometry enhancement and computational overhead. Specifically, we have constructed one of the largest normal-based multimodal dataset, ensuring superior data quality and diversity through meticulous subjective selection. Furthermore, we explore a new two-branch multimodal alignment approach along with a multimodal split fusion module to mitigate computational complexity while improving restoration performances. To address the limitations associated with normal-based multimodal learning, we develop novel normal-induced loss functions that facilitate geometric consistency and improve feature alignment. Extensive experiments conducted on seven benchmark datasets across four different 3D data representations demonstrate that mn3DSSR consistently outperforms state-of-the-art super-resolution methods in terms of restoration accuracy with high computational efficiency. Miaohui Wang, Yunheng Liu, Wuyuan Xie, Boxin Shi, Jianmin Jiang |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | A Visual-Linguistic Approach for Robust RGB-Thermal Tracking With Dynamic Template AdaptationabstractRGB-T tracking seeks to improve tracking robust ness in complex environments by exploiting the complementary advantages of RGB and thermal infrared (TIR) modalities. Nevertheless, current RGB-T tracking methods encounter two critical limitations stemming from their reliance on fixed template images for search region matching. First, the resemblance between the initial fixed template and the target in later frames diminishes as time goes on, causing tracking reliability to decline. Second, background clutter within template bounding boxes introduces disruptive noise that compromises tracking accuracy. To address these challenges, we propose a new referring RGB-T tracking approach that integrates visual-linguistic cues to enhance multimodal data, enabling background noise elimination in the template and dynamic template image updating through adaptive mechanisms. Additionally, to facilitate research in language guided multimodal tracking, we construct a large-scale dataset named Refer-RGBT, which comprises 1508 multimodal video pairs with synchronized RGB/TIR sequences of diverse objects. Each sequence is annotated by descriptive textual captions that detail their appearance and actions. Extensive evaluations on the Refer-RGBT dataset demonstrate that our proposed approach achieves cutting-edge performance compared to the state-of-the art (SOTA) methods. Xu Wang 0006, Huanxin Zheng, Haohong Liao, Qiudan Zhang, Lin Ma 0002, Jianmin Jiang |
IEEE Trans. Multim. | 6 |
| 2025 | EEG-Face: A Facial-Image Stimulated EEG Data-Set for Analysis of Brain Perceived MultimediaabstractOver recent years, EEG-based brain decoding of perceived multimedia is emerging to be an important multidisciplinary research area. Lack of data sets with multimedia stimuli, however, presents a significant challenge for its further advancement. In this paper, we establish a facial-image stimulated EEG dataset, named as EEG-Face, to address the challenge and provide a crucial support for relevant research, such as brain-computer interface (BCI), face recognition via brain-perceived EEGs, and multimedia content analysis via brain perception activities. As facial images not only distinguish between genders but also dive deeper into individual differences, our proposed EEG-Face provides larger scope, more focus, and greater potential for dedicated research on brain perception of human faces. As shown in Figure 1, the proposed EEG-Face essentially consists of 20,000 brain responded EEG trials stimulated with 40 individual faces, all of whom are Chinese film stars. Following the establishment of the dataset, a range of experiments over EEG-Face is carried out to demonstrate its usability and feasibility, which include: (i) neural correlation of gender perceptions; (ii) EEG-Stimulus pairing verification; and (iii) face recognition via classification of randomized EEG trials. The dataset and the codes for all reported experiments are available from: https://github.com/eeg-wx2024/EEG-Face. Wuxia Zhang, Yang Xin 0004, Shibo Lv, Xin Zhang 0056, Jianmin Jiang |
ACM Multimedia | 6 |
| 2025 | Semantic driven ViT and dual-level token fusion for weakly supervised semantic segmentation
Qiudan Zhang, Jianmin Jiang |
Neurocomputing | 3 |
| 2025 | Interpretable Optimization-Inspired Unfolding Network for Low-Light Image EnhancementabstractRetinex model-based methods have shown to be effective in layer-wise manipulation with well-designed priors for low-light image enhancement (LLIE). However, the hand-crafted priors and conventional optimization algorithm adopted to solve the layer decomposition problem result in the lack of adaptivity and efficiency. To this end, this paper proposes a Retinex-based deep unfolding network (URetinex-Net++), which unfolds an optimization problem into a learnable network to decompose a low-light image into reflectance and illumination layers. By formulating the decomposition problem as an implicit priors regularized model, three learning-based modules are carefully designed, responsible for data-dependent initialization, high-efficient unfolding optimization, and fairly-flexible component adjustment, respectively. Particularly, the proposed unfolding optimization module, introducing two networks to adaptively fit implicit priors in the data-driven manner, can realize noise suppression and details preservation for decomposed components. URetinex-Net++ is a further augmented version of URetinex-Net, which introduces a cross-stage fusion block to alleviate the color defect in URetinex-Net. Therefore, boosted performance on LLIE can be obtained in both visual quality and quantitative metrics, where only a few parameters are introduced and little time is cost. Extensive experiments on real-world low-light images qualitatively and quantitatively demonstrate the effectiveness and superiority of the proposed URetinex-Net++ over state-of-the-art methods. Wenhui Wu 0001, Jian Weng 0009, Xu Wang 0006, Wenhan Yang, Jianmin Jiang |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Denoiser-Regulated Deep Unfolding Compressed Sensing With Learnable Fixed-Point ProjectionsabstractThe family of regularization by denoising (RED) methods introduce denoising operator as the regularization term to perform compressed sensing (CS) reconstruction, which shows higher flexibility and scalability. However, traditional RED framework has strict requirements on several properties of denoiser, making it hard to design the specific denoiser and limits the quality of reconstructed images. Although some relaxation for denoisers can be made by incorporating the fixed point projection during the iteration process, the involved parameters have great impact on the effectiveness and efficiency of the algorithm, which is non-trivial to set them properly. In this paper, we propose an innovative Deep Unfolding Network framework termed FP-DUN based on the iterative process of Regularization by Denoising via Fixed-Point Projection (RED-PRO). In FP-DUN, fix-point projection module is implemented with learnable weights of neural networks, where an effective denoiser based on dual attention mechanism (DAM) is developed to capture the details of the reconstructed image. Additionally, we propose a new loss function based on fixed point constraints, which is able to overcome the over-smoothness caused by multi-stage denoising and maintain the structural details to progressively improve the reconstruction quality. By training the DUN model, the parameters for the process of fix point projection and denoiser are learned automatically. Extensive experimental results comparing with state-of-the-art CS algorithms and traditional RED-PRO approach validate the effectiveness of FP-DUN, especially on some images with complex details. Yu Zhou 0027, Wei Xie 0020, Huisi Wu, Lei Huang 0001, Sam Kwong, Jianmin Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Robust Face Recognition via Adaptive Mining and Margining of Noise and Hard SamplesabstractAt present, deep face recognition models working on millions of images are confronted with the challenge that such large-scale datasets are often corrupted with noises and mislabeled identities yet most deep models are primarily designed for clean datasets. In this paper, we propose a robust deep face recognition model by exploiting the advantage of integrating the strength of margin-based learning models with the strength of mining-based approaches to effectively mitigate the impact of noises during training. By monitoring the recognition performances at a batch level to provide optimization-oriented feedback, we introduce a noise-adaptive mining strategy to dynamically adjust the emphasis balance between hard and noise samples, enabling direct training on noisy datasets without the requirement of pre-training. With a novel anti-noise loss function, learning is empowered for direct and robust training on noisy datasets yet its effectiveness over clean datasets is still preserved, sustaining effective mining of both clean and noisy samples whilst weakening its learning intensiveness over noisy samples. Extensive experiments reveal that: (i) our proposed achieves competitive performances in comparison with representative existing SoTA models when trained with clean datasets; (ii) when trained with both real-world and synthesized noisy datasets, our proposed significantly outperforms the existing models, especially when the synthesized datasets are corrupted with both close-set and open-set noises; (iii) while the existing deep models suffer from an average performance drop of around 20% over noise-corrupted large scale datasets, our proposed still delivers accuracy rates of more than 95%. Our source codes are publicly available on GitHub. Yang Xin 0004, Yu Zhou 0027, Jianmin Jiang |
IEEE Trans. Image Process. | 4 |
| 2025 | Geometry-Aware Self-Supervised Indoor 360$^{\circ }$ Depth Estimation via Asymmetric Dual-Domain Collaborative LearningabstractBeing able to estimate monocular depth for spherical panoramas is of fundamental importance in 3D scene perception. However, spherical distortion severely limits the effectiveness of vanilla convolutions. To push the envelope of accuracy, recent approaches attempt to utilize Tangent projection (TP) to estimate the depth of$360 ^{\circ }$images. Yet, these methods still suffer from discrepancies and inconsistencies among patch-wise tangent images, as well as the lack of accurate ground truth depth maps under a supervised fashion. In this paper, we propose a geometry-aware self-supervised$360 ^{\circ }$image depth estimation methodology that explores the complementary advantages of TP and Equirectangular projection (ERP) by an asymmetric dual-domain collaborative learning strategy. Especially, we first develop a lightweight asymmetric dual-domain depth estimation network, which enables to aggregate depth-related features from a single TP domain, and then produce depth distributions of the TP and ERP domains via collaborative learning. This effectively mitigates stitching artifacts and preserves fine details in depth inference without overspending model parameters. In addition, a frequent-spatial feature concentration module is devised to simultaneously capture non-local Fourier features and local spatial features, such that facilitating the efficient exploration of monocular depth cues. Moreover, we introduce a geometric structural alignment module to further improve geometric structural consistency among tangent images. Extensive experiments illustrate that our designed approach outperforms existing self-supervised$360 ^{\circ }$depth estimation methods on three publicly available benchmark datasets. Xu Wang 0006, Ziyan He, Qiudan Zhang, You Yang 0002, Tiesong Zhao, Jianmin Jiang |
IEEE Trans. Multim. | 6 |
| 2025 | Hierarchical Uncertainty-Aware Salient Object Detection for $360 ^{\circ }$ Images via Bi-Projection Collaborative Learningabstract$360^{\circ }$salient object detection has recently received much attention for 3D scene perception owing to its omnidirectional field of view (FoV). The capability of recognizing salient objects of$360^{\circ }$images remains technically challenging due to severe spherical distortion. In this paper, we develop a hierarchical uncertainty-aware$360^{\circ }$image salient object detection methodology that explicitly explores the geometric and spatial complementary coherence of Tangent projection (TP) and Equirectangular projection (ERP) by a collaborative learning strategy. Concretely, to mitigate spherical distortion, we first intend to learn saliency-related features from less-distorted tangent images, in which a deformation-aware attention block is introduced to mitigate the geometric distortion caused by projecting a$360^{\circ }$image onto a 2D plane. However, the discrepancies among tangent images pose a new challenge to$360^{\circ }$image salient object detection. To tackle this issue and achieve accurate localization for salient objects of all sizes, we design a spatial-frequency saliency feature aggregation module to leverage fast Fourier convolution to capture global contextual information from ERP images, such that obtaining more representative saliency features. Moreover, a hierarchical uncertainty-aware bi-projection consistency learning module with strong local-global information embedding capabilities is constructed, which learns the geometric and spatial correlations between tangent images and ERP images via a collaborative learning strategy. Ultimately, salient object maps are produced for$360^{\circ }$images on the basis of the merged saliency features driven by the uncertainty. Extensive experiments show that our developed method improves${\mathrm{F}}_\beta ^{\sigma }$by an average of 31.67% compared to twenty existing advanced methods on the publicly available 360-SOD dataset. Qiudan Zhang, Kaiyu Ji, Xu Wang 0006, Zhaoqing Pan, Jianmin Jiang |
IEEE Trans. Multim. | 6 |
| 2025 | VAT: Visibility Aware Transformer for Fine-Grained Clothed Human ReconstructionabstractIn order to reconstruct 3D clothed human with accurate fine-grained details from sparse views, we propose a deep cooperating two-level global to fine-grained reconstruction framework that constructs robust global geometry to guide fine-grained geometry learning. The core of the framework is a novel visibility aware Transformer VAT, which bridges the two-level reconstruction architecture by connecting its global encoder and fine-grained decoder with two pixel-aligned implicit functions, respectively. The global encoder fuses semantic features of multiple views to integrate global geometric features. In the fine-grained decoder, visibility aware attention mechanism is designed to efficiently fuse multi-view and multi-scale features for mining fine-grained geometric features. The global encoder and fine-grained decoder are connected by a global embeding module to form a deep cooperation in the two-level framework, which provides global geometric embedding as a query guidance for calculating visibility aware attention in the fine-grained decoder. In addition, to extract highly aligned multi-scale features for the two-level reconstruction architecture, we design an image feature extractor MSUNet, which establishes strong semantic connections between different scales at minimal cost. Our proposed framework is end-to-end trainable, with all modules jointly optimized. We validate the effectiveness of our framework on public benchmarks, and experimental results demonstrate that our method has significant advantages over state-of-the-art methods in terms of both fine-grained performance and generalization. Xiaoyan Zhang 0002, Zibin Zhu, Sisi Ren, Jianmin Jiang |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | RobustFace: Adaptive Mining of Noise and Hard Samples for Robust Face RecognitionsabstractWhile margin-based deep face recognition models, such as ArcFace and AdaFace, have achieved remarkable successes over recent years, they may suffer from degraded performances when encountering training sets corrupted with noises. This is often inevitable when massively large scale datasets need to be dealt with, yet it remains difficult to construct clean enough face datasets under these circumstances. In this paper, we propose a robust deep face recognition model, RobustFace, by combining the advantages of margin-based learning models with the strength of mining-based approaches to effectively mitigate the impact of noises during trainings. Specifically, we introduce a noise-adaptive mining strategy to dynamically adjust the emphasis balance between hard and noise samples by monitoring the model's recognition performances at the batch level to provide optimization-oriented feedback, enabling direct training on noisy datasets without the requirement of pre-training. Extensive experiments validate that our proposed RobustFace achieves competitive performances in comparison with the existing SoTA models when trained with clean datasets. When trained with both real-world and synthetic noisy datasets, RobustFace significantly outperforms the existing models, especially when the synthetic noisy datasets are corrupted with both close-set and open-set noises. While the existing baseline models suffer from an average performance drop of around 40%, under these circumstances, our proposed still delivers accuracy rates of more than 90%. Yang Xin 0004, Yu Zhou 0027, Jianmin Jiang |
ACM Multimedia | 3 |
| 2024 | Effective Optimization of Root Selection Towards Improved Explanation of Deep ClassifiersabstractExplaining what part of the input images primarily contributed to the predicted classification results by deep models has been widely researched over the years and many effective methods have been reported in the literature, for which deep Taylor decomposition (DTD) served as the primary foundation due to its advantage in theoretical explanations brought in by Taylor expansion and approximation. Recent research, however, has shown that the root of Taylor decomposition could extend beyond local linearity, and thus causing DTD to fail in delivering expected performances. In this paper, we propose a universal root inference method to overcome the shortfall and strengthen the roles of DTD in explainability and interpretability of deep classifications. In comparison with the existing approaches, our proposed features in: (i) theoretical establishment of the relationship between ideal roots and the propagated relevances; (ii) exploitation of gradient descents in learning a universal root inference; and (iii) constrained optimization of its final root selection. Extensive experiments, including both quantitative and qualitative, validate that our proposed root inference is not only effective, but also delivers significantly improved performances in explaining a range of deep classifiers. We share our codes via the link: https://github.com/meetxinzhang/XAI-RootInference. Xin Zhang 0056, Shenghua Zhong, Jianmin Jiang |
ACM Multimedia | 3 |
| 2024 | Managing Traceability for Software Life Cycle Processes
Hao Wen 0008, Jianmin Jiang, Zhong Hong |
TASE | 3 |
| 2024 | Distortion-Aware Self-Supervised Indoor 360$^{\circ }$ Depth Estimation via Hybrid Projection Fusion and Structural RegularitiesabstractOwing to the rapid development of emerging 360$^{\circ }$panoramic imaging techniques, indoor 360$^{\circ }$depth estimation has aroused extensive attention in the community. Due to the lack of available ground truth depth data, it is extremely urgent to model indoor 360$^{\circ }$depth estimation in self-supervised mode. However, self-supervised 360$^{\circ }$depth estimation suffers from two major limitations. One is the distortion and network training problems caused by Equirectangular projection (ERP), and the other is that texture-less regions are quite difficult to back-propagate in self-supervised mode. Hence, to address the above issues, we introduce spherical view synthesis for learning self-supervised 360$^{\circ }$depth estimation. Specifically, to alleviate the ERP-related problems, we first propose a dual-branch distortion-aware network to produce the coarse depth map, including a distortion-aware module and a hybrid projection fusion module. Subsequently, the coarse depth map is utilized for spherical view synthesis, in which a spherically weighted loss function for view reconstruction and depth smoothing is investigated to optimize the projection distribution problem of 360$^{\circ }$images. In addition, two structural regularities of indoor 360$^{\circ }$scenes are devised as two additional supervisory signals to efficiently optimize our self-supervised 360$^{\circ }$depth estimation model, containing the principal-direction normal constraint and the co-planar depth constraint. The principal-direction normal constraint is designed to align the normal of the 360$^{\circ }$image with the direction of the vanishing points. Meanwhile, we employ the co-planar depth constraint to fit the estimated depth of each pixel through its 3D plane. Finally, a depth map is obtained for the 360$^{\circ }$image. Experimental results illustrate that our proposed method achieves superior performance than the current advanced depth estimation methods on four publicly available datasets. Xu Wang 0006, Weifeng Kong, Qiudan Zhang, You Yang 0002, Tiesong Zhao, Jianmin Jiang |
IEEE Trans. Multim. | 6 |
| 2024 | Weakly-Supervised 3D Scene Graph Generation via Visual-Linguistic Assisted Pseudo-LabelingabstractLearning to build 3D scene graphs is essential for real-world perception in a structured and rich fashion. However, previous 3D scene graph generation methods utilize a fully supervised learning manner and require a large amount of entity-level annotation data of objects and relations, which is extremely resource-consuming and tedious to obtain. To tackle this problem, we propose 3D-VLAP, a weakly-supervised 3D scene graph generation method via Visual-Linguistic Assisted Pseudo-labeling. Specifically, our 3D-VLAP exploits the superior ability of current large-scale visual-linguistic models to align the semantics between texts and 2D images, as well as the naturally existing correspondences between 2D images and 3D point clouds, and thus implicitly constructs correspondences between texts and 3D point clouds. First, we establish the positional correspondence from 3D point clouds to 2D images via camera intrinsic and extrinsic parameters, thereby achieving alignment of 3D point clouds and 2D images. Subsequently, a large-scale cross-modal visual-linguistic model is employed to indirectly align 3D instances with the textual category labels of objects by matching 2D images with object category labels. The pseudo labels for objects and relations are then produced for 3D-VLAP model training by calculating the similarity between visual embeddings and textual category embeddings of objects and relations encoded by the visual-linguistic model, respectively. Ultimately, we design an edge self-attention based graph neural network to generate scene graphs of 3D point clouds. Experiments demonstrate that our 3D-VLAP achieves comparable results with current fully supervised methods, meanwhile alleviating the data annotation pressure. Xu Wang 0006, Qiudan Zhang, Wenhui Wu 0001, Mark Junjie Li, Lin Ma 0002, Jianmin Jiang |
IEEE Trans. Multim. | 7 |
| 2023 | Weakly supervised semantic segmentation via self-supervised destruction learning
Jinlong Li 0003, Zequn Jie, Xu Wang 0006, Yu Zhou 0027, Lin Ma 0002, Jianmin Jiang |
Neurocomputing | 6 |
| 2023 | A Formal Approach for Consistency Management in UML ModelsabstractConsistency is a significant indicator to measure the correctness of a software system in its lifecycle. It is inevitable to introduce inconsistencies between different software artifacts in the software development process. In practice, developers perform consistency checking to detect inconsistencies, and apply their corresponding repairs to restore consistencies. Even if all inconsistencies can be repaired, how to preserve consistencies in the subsequent evolution should be considered. Consistency management (consistency checking and consistency preservation) is a challenging task, especially in the multi-view model-driven software development process. Although there are some efforts to discuss consistency management, most of them lack the support of formal methods. Our work aims to provide a framework for formal consistency management, which may be used in the practical software development process. A formal model, called a Structure model, is first presented for specifying the overall model-based structure of the software system. Next, the definition of consistency is given based on consistency rules. We then investigate consistency preservation under the following two situations. One is that if the initial system is inconsistent, then the consistency can be restored through repairs. The other is that if the initial system is consistent, then the consistency can be maintained through update propagation. To demonstrate the effectiveness of our approach, we finally present a case study with a prototype tool. Hao Wen 0008, Jianmin Jiang, Guofu Tang, Zhong Hong |
Int. J. Softw. Eng. Knowl. Eng. | 3 |
| 2023 | Surface Geometry Processing: An Efficient Normal-Based Detail RepresentationabstractWith the rapid development of high-resolution 3D vision applications, the traditional way of manipulating surface detail requires considerable memory and computing time. To address these problems, we introduce an efficient surface detail processing framework in 2D normal domain, which extracts new normal feature representations as the carrier of micro geometry structures that are illustrated both theoretically and empirically in this article. Compared with the existing state of the arts, we verify and demonstrate that the proposed normal-based representation has three important properties, including detail separability, detail transferability and detail idempotence. Finally, three new schemes are further designed for geometric surface detail processing applications, including geometric texture synthesis, geometry detail transfer, and 3D surface super-resolution. Theoretical analysis and experimental results on the latest benchmark dataset verify the effectiveness and versatility of our normal-based representation, which accepts 30 times of the input surface vertices but at the same time only takes 6.5% memory cost and 14.0% running time in comparison with existing competing algorithms. Wuyuan Xie, Miaohui Wang, Di Lin 0002, Boxin Shi, Jianmin Jiang |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | URetinex-Net: Retinex-based Deep Unfolding Network for Low-light Image EnhancementabstractRetinex model-based methods have shown to be effective in layer-wise manipulation with well-designed priors for low-light image enhancement. However, the commonly used handcrafted priors and optimization-driven solutions lead to the absence of adaptivity and efficiency. To address these issues, in this paper, we propose a Retinex-based deep unfolding network (URetinex-Net), which unfolds an optimization problem into a learnable network to decompose a low-light image into reflectance and illumination layers. By formulating the decomposition problem as an implicit priors regularized model, three learning-based modules are carefully designed, responsible for data-dependent initialization, high-efficient unfolding optimization, and user-specified illumination enhancement, respectively. Particularly, the proposed unfolding optimization module, introducing two networks to adaptively fit implicit priors in data-driven manner, can realize noise suppression and details preservation for the final decomposition results. Extensive experiments on real-world low-light images qualitatively and quantitatively demonstrate the effectiveness and superiority of the proposed method over state-of-the-art methods. The code is available at https://github.com/AndersonYong/URetinex-Net. Wenhui Wu 0001, Jian Weng 0009, Xu Wang 0006, Wenhan Yang, Jianmin Jiang |
CVPR | 6 |
| 2022 | Deep image inpainting via contextual modelling in ADCT domainabstractAbstract Pixel‐based generative image inpainting has been widely researched over recent years and certain level of success via deep learning of feature representations and hallucinations of missing pixel values from surrounding backgrounds have also been reported in the literature. However, existing approaches rely on context‐based attentions and progressive inferences to capture the pixel correlations yet such pixel‐based approaches often fail to adapt to the constantly varying ranges and distances among surrounding background pixels. On the other hand, the modelling cost is also increasingly expensive whenever correlations of those pixels at longer distance away are to be exploited. To resolve the problem, we implement the principle of learning and hallucinating frequency components rather than pixel values. Therefore, we can avoid the dilemma that, on one hand the wish is to exploit all correlated pixels inside the image no matter how far away they are spatially located, but on the other, the price of increasing the modelling cost incurred by those pixels far away from the missing regions has to be paid. Extensive experiments carried out verify the effectiveness of the proposed method, which outperforms the representative existing state of the arts in terms of all assessment metrics. Adhiyaman Manickam, Jianmin Jiang, Yu Zhou 0027 |
IET Image Process. | 2 |
| 2022 | Generative synthesis of logos across DCT domain
Lisha Dong, Yu Zhou 0027, Jianmin Jiang |
Neurocomputing | 3 |
| 2022 | Deep stereoscopic image saliency inspired stereoscopic image thumbnail generation
Yu Zhou 0027, Xiaotong Xiao, Qiudan Zhang, Xu Wang 0006, Jianmin Jiang |
Multim. Tools Appl. | 5 |
| 2022 | Learning Across Tasks for Zero-Shot Domain Adaptation From a Single Source DomainabstractDomain adaptation techniques learn transferable knowledge from a source domain to a target domain and train models that generalize well in the target domain. Unfortunately, a majority of the existing techniques are only applicable to scenarios that the target-domain data in the task of interest is available for training, yet this is not often true in practice. In general, human beings are experts in generalization across domains. For example, a baby can easily identify the bear from a clipart image after learning this category of animal from the photo images. To reduce the gap between the generalization ability of human and that of machines, we propose a new solution to the challenging zero-shot domain adaptation (ZSDA) problem, where only a single source domain is available and the target domain for the task of interest is not accessible. Inspired by the observation that the knowledge about domain correlation can improve our generalization ability, we explore the correlation between source domain and target domain in an irrelevant knowledge task ([Formula: see text]-task), where dual-domain samples are available. We denote the task of interest as the question task ([Formula: see text]-task) and synthesize its non-accessible target-domain as such that these two tasks have the shared domain correlation. In order to realize our idea, we introduce a new network structure, i.e., conditional coupled generative adversarial networks (CoCoGAN), by extending the coupled generative adversarial networks (CoGAN) into a conditioning model. With a pair of coupling GANs, our CoCoGAN is able to capture the joint distribution of data samples across two domains and two tasks. For CoCoGAN training in a ZSDA task, we introduce three supervisory signals, i.e., semantic relationship consistency across domains, global representation alignment across tasks, and alignment consistency across domains. Experimental results demonstrate that our method can learn a suitable model for the non-accessible target domain and outperforms the existing state of the arts in both image classification and semantic segmentation. Jianmin Jiang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Preserving similarity order for unsupervised clustering
Li Wang 0057, Jianmin Jiang |
Pattern Recognit. | 3 |
| 2022 | Adaptive Viewpoint Feature Enhancement-Based Binocular Stereoscopic Image Saliency DetectionabstractModeling 3D visual saliency has received great attention due to the development of emerging 3D display technologies. Traditional methods relying on low-level features may not be efficient in interpreting 3D visual content from high-level semantic perspective. Despite numerous efforts dedicated to this area, existing 3D visual saliency detection methods do not necessarily excel in exploring the stereoscopic image saliency driven by the intra-view and inter-view dependencies among left and right views. In this paper, we propose a visual saliency detection method for stereoscopic images grounded on adaptive viewpoint feature enhancement via binocular vision. More specifically, the correlation among left and right views is investigated through a delicately designed binocular stereoscopic saliency feature aggregation module, enabling the generation of more representative saliency features towards binocular vision. Subsequently, to further aggregate the saliency features in multiple scales, we design a progressive attention-based saliency feature pyramid extraction module to effectively integrate the features from top-level to down-level based on the network hierarchy mechanism. The saliency maps are ultimately produced for stereoscopic images by evaluating the obtained saliency features. In addition, we create a stereoscopic image saliency dataset (SIS-3D) that includes 1086 stereoscopic image pairs with various content and their corresponding human eye fixation annotations, aiming to further facilitate the research on visual saliency detection for stereoscopic images. Extensive experiments demonstrate that our proposed method improves CC by an average of 4.02% compared to representative counterparts on the newly built saliency dataset and another publicly available dataset. Qiudan Zhang, Xiaotong Xiao, Xu Wang 0006, Shiqi Wang 0001, Sam Kwong, Jianmin Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | Fine Tuning of Deep Contexts Toward Improved Perceptual Quality of In-PaintingsabstractOver the recent years, a number of deep learning approaches are successfully introduced to tackle the problem of image in-painting for achieving better perceptual effects. However, there still exist obvious hole-edge artifacts in these deep learning-based approaches, which need to be rectified before they become useful for practical applications. In this article, we propose an iteration-driven in-painting approach, which combines the deep context model with the backpropagation mechanism to fine-tune the learning-based in-painting process and hence, achieves further improvement over the existing state of the arts. Our iterative approach fine tunes the image generated by a pretrained deep context model via backpropagation using a weighted context loss. Extensive experiments on public available test sets, including the CelebA, Paris Streets, and PASCAL VOC 2012 dataset, show that our proposed method achieves better visual perceptual quality in terms of hole-edge artifacts compared with the state-of-the-art in-painting methods using various context models. Qinglong Chang, Kwok-Wai Hung, Jianmin Jiang |
IEEE Trans. Cybern. | 3 |
| 2022 | Video Saliency Prediction via Joint Discrimination and Local ConsistencyabstractWhile saliency detection on static images has been widely studied, the research on video saliency detection is still in an early stage and requires more efforts due to the challenge to bring both local and global consistency of salient objects into full consideration. In this article, we propose a novel dynamic saliency network based on both local consistency and global discriminations, via which semantic features across video frames are simultaneously extracted and a recurrent feature optimization structure is designed to further enhance its performances. To ensure that the generated dynamic salient map is more concentrated, we design a lightweight discriminator with a local consistency loss LC to identify subtle differences between predicted maps and ground truths. As a result, the proposed network can be further stimulated to produce more realistic saliency maps with smoother boundaries and simpler layer transitions. The added LC loss forces the network to pay more attention to the local consistency between continuous saliency maps. Both qualitative and quantitative experiments are carried out on three large datasets, and the results demonstrate that our proposed network not only achieves improved performances but also shows good robustness. Zheng Wang 0008, Ziqi Zhou 0002, Huchuan Lu, Qinghua Hu, Jianmin Jiang |
IEEE Trans. Cybern. | 5 |
| 2022 | Scheduling in Real-Time Mobile SystemsabstractTo guarantee the safety and security of a real-time mobile system such as an intelligent transportation system, it is necessary to model and analyze its behaviors prior to actual development. In particular, the mobile objects in such systems must be isolated from each other so that they do not collide with each other. Since isolation means two or more mobile objects must not be located in the same place at the same time, a scheduling policy is required to control and coordinate the movement of such objects. However, traditional scheduling theories are based on task scheduling which is coarse-grained and cannot be directly used for fine-grained isolation controls. In this article, we first propose a fine-grained event-based formal model called a time dependency structure and use it to model and analyze real-time mobile systems. Next, an event-based schedule is defined and the composition of schedules is discussed. Then, we investigate the schedulability of isolation—that is, checking whether a given schedule ensures the isolation relationship among mobile objects or not. After that, we present an automation approach for scheduling generation to guarantee isolation controls in real-time mobile systems. Finally, case studies and simulation experiments demonstrate the usability and effectiveness of our approach. Zhong Hong, Jianmin Jiang |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2022 | PR-RL: Portrait Relighting Via Deep Reinforcement LearningabstractIn this paper, we propose a portrait relighting method based on deep reinforcement learning (called PR-RL). Our PR-RL model could conduct portrait relighting by sequentially predicting local light editing strokes, and use strokes to conduct dodge and burn operations on the image lightness, simulating image editing by artists using brush strokes. Reinforcement learning with Deep Deterministic Policy Gradient is introduced to design our PR-RL model, defining the action (stroke parameters) in a continuous space, through which a reward can be designed to guide the agent to learn and relight a portrait image like an artist. To optimize the relighting effect, we further enable the reward to be location relevant and hence a coarse-to-fine strategy can be applied to select corresponding actions and maximize the performance of the proposed method. In comparison with the existing efforts, our proposed PR-RL method is locally effective, scale-invariant and interpretable. We apply the proposed method to tasks of portrait relighting based on both SH-lighting and reference images. The experiments show that our PR-RL method outperforms state-of-the-art methods in generating locally effective and interpretable high resolution relighting results for wild portrait images. Xiaoyan Zhang 0002, Yukai Song, Zhuopeng Li, Jianmin Jiang |
IEEE Trans. Multim. | 4 |
| 2021 | A recurrent video quality enhancement framework with multi-granularity frame-fusion and frame difference based attention
Yongkai Huo, Qiyan Lian, Shaoshi Yang, Jianmin Jiang |
Neurocomputing | 4 |
| 2021 | Unsupervised deep clustering via adaptive GMM modeling and optimization
Jianmin Jiang |
Neurocomputing | 2 |
| 2021 | Progressive Point Cloud Upsampling via Differentiable RenderingabstractIn this paper, we propose one novel progressive point cloud upsampling framework to tackle the non-uniform distribution issue during the point cloud upsampling process. Specifically, we design an Up-UNet feature expansion module which is capable of learning the local and global point features via a down-feature operator and an up-feature operator, respectively, to alleviate the non-uniform distribution issue and remove the outliers. Moreover, we design a hybrid loss function considering both the multi-scale reconstruction loss and the rendering loss. The multi-scale reconstruction loss enables each upsampling module to generate a denser point cloud, while the rendering loss via point-based differentiable rendering ensures that the proposed model preserves the point cloud structures. Extensive experimental results demonstrate that our proposed model achieves state-of-the-art performance in terms of both qualitative and quantitative evaluations. Xu Wang 0006, Lin Ma 0002, Shiqi Wang 0001, Sam Kwong, Jianmin Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2021 | A Multi-Task Collaborative Network for Light Field Salient Object DetectionabstractBeing able to predict the salient object is of fundamental importance in image processing and computer vision. With numerous approaches proposed for automatic image and video salient object detection, much less work has been dedicated to detecting and segmenting salient objects from light fields. In this article, based on the intrinsic characteristics of light fields, we carefully explore the complementary coherence among multiple cues including spatial, edge and depth information, and elaborately design a multi-task collaborative network for light field salient object detection. More specifically, the correlation mechanisms among edge detection, depth inference and salient object detection are carefully investigated to facilitate the representative saliency features. We first model the coherence among low-level features and heuristic semantic priors, as well as the edge information. Subsequently, the depth-oriented saliency features are derived from the geometry of light fields, in which the 3D convolution operation is leveraged with powerful representation capability to model the disparity correlations among multiple viewpoint images. Finally, a feature-enhanced salient object generator is developed to integrate these complementary saliency features, leading to the final salient object predictions for light fields. Quantitative and qualitative experiments demonstrate the superiority of our proposed model against the state-of-the-art methods over the public light field salient object detection datasets. Qiudan Zhang, Shiqi Wang 0001, Xu Wang 0006, Zhenhao Sun, Sam Kwong, Jianmin Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2021 | Domain Shift Preservation for Zero-Shot Domain AdaptationabstractIn learning-based image processing a model that is learned in one domain often performs poorly in another since the image samples originate from different sources and thus have different distributions. Domain adaptation techniques alleviate the problem of domain shift by learning transferable knowledge from the source domain to the target domain. Zero-shot domain adaptation (ZSDA) refers to a category of challenging tasks in which no target-domain sample for the task of interest is accessible for training. To address this challenge, we propose a simple but effective method that is based on the strategy of domain shift preservation across tasks. First, we learn the shift between the source domain and the target domain from an irrelevant task for which sufficient data samples from both domains are available. Then, we transfer the domain shift to the task of interest under the hypothesis that different tasks may share the domain shift for a specified pair of domains. Via this strategy, we can learn a model for the unseen target domain of the task of interest. Our method uses two coupled generative adversarial networks (CoGANs) to capture the joint distribution of data samples in dual-domains and another generative adversarial network (GAN) to explicitly model the domain shift. The experimental results on image classification and semantic segmentation demonstrate the satisfactory performance of our method in transferring various kinds of domain shifts across tasks. Ming-Ming Cheng, Jianmin Jiang |
IEEE Trans. Image Process. | 3 |
| 2021 | Geometry Auxiliary Salient Object Detection for Light Fields via Graph Neural NetworksabstractLight field imaging, originated from the availability of light field capture technology, offers a wide range of applications in the field of computational vision. The capability of predicting salient objects of light fields remains technologically challenging due to its complicated geometry structure. In this paper, we propose a light field salient object detection approach that formulates the geometric coherence among multiple views of light fields as graphs, where the angular/central views represent the nodes and their relations compose the edges. The spatial and disparity correlations between multiple views are effectively explored through multi-scale graph neural networks, enabling the more comprehensive understanding of light field content and more representative and discriminative saliency features generation. Moreover, a multi-scale saliency feature consistency learning module is embedded to enhance the saliency features. Finally, an accurate salient object map is produced for the light field based upon the extracted features. In addition, we establish a new light field salient object detection dataset (CITYU-Lytro) that contains 817 light fields with diverse contents and their corresponding annotations, aiming to further promote the research on light field salient object detection. Quantitative and qualitative experiments demonstrate that the proposed method performs favorably compared with the state-of-the-art methods on the benchmark datasets. Qiudan Zhang, Shiqi Wang 0001, Xu Wang 0006, Zhenhao Sun, Sam Kwong, Jianmin Jiang |
IEEE Trans. Image Process. | 6 |
| 2021 | A Brain-Media Deep Framework Towards Seeing Imaginations Inside BrainsabstractWhile current research on multimedia is essentially dealing with the information derived from our observations of the world, internal activities inside human brains, such as imaginations and memories of past events etc., could become a brand new concept of multimedia, for which we coin as “brain-media”. In this paper, we pioneer this idea by directly applying natural images to stimulate human brains and then collect the corresponding electroencephalogram (EEG) sequences to drive a deep framework to learn and visualize the corresponding brain activities. By examining the relevance between the visualized image and the stimulation image, we are able to assess the performance of our proposed deep framework in terms of not only the quality of such visualization but also the feasibility of introducing the new concept of “brain-media”. To ensure that our explorative research is meaningful, we introduce a dually conditioned learning mechanism in the proposed deep framework. One condition is analyzing EEG sequences through deep learning to extract a more compact and class-dependent brain features via exploiting those unique characteristics of human brains such as hemispheric lateralization and biological neurons myelination (neurons importance), and the other is to analyze the content of images via computing approaches and extract representative visual features to exploit artificial intelligence in assisting our automated analysis of brain activities and their visualizations. By combining the brain feature space with the associated visual feature space of those images that are candidates of the stimuli, we are able to generate a combined-conditional space to support the proposed dual-conditioned and lateralization-supported GAN framework. Extensive experiments carried out illustrate that our proposed deep framework significantly outperforms the existing relevant work, indicating that our proposed does provide a good potential for further research upon the introduced concept of “brain-media”, a new member for the big family of multimedia. To encourage more research along this direction, we make our source codes publicly available for downloading at GitHub. Jianmin Jiang, Ahmed Fares, Shenghua Zhong |
IEEE Trans. Multim. | 1 |
| 2021 | Emotion Attention-Aware Collaborative Deep Reinforcement Learning for Image CroppingabstractThis paper proposes a collaborative deep reinforcement learning model for automatic image cropping (called CDRL-IC). By modeling image cropping as a decision-making process of reinforcement learning, our model could generate optimal cropping result in a few moving and zooming steps. An image with good composition is a comprehensive result by considering the relative importance of objects and also the spatial organization of visual elements. Therefore, emotion attention information which indicates the relationship and importance between objects is applied together with contextual information of color image for image cropping. In order to sufficiently use the emotion attention map and the color image, they are processed by two collaborative agents. The two agents make their primary learning separately and then share information through an information interaction module for making joint action prediction. In order to efficiently evaluate the cropping quality in the reward function, weighted Intersection Over Union (WIoU) is designed by integrating emotion attention map in the traditional IoU. Our CDRL-IC model is tested on a variety of datasets for both image cropping and thumbnail generation. The experiments show that our CDRL-IC model outperforms state-of-the-art methods on these benchmark datasets. Xiaoyan Zhang 0002, Zhuopeng Li, Jianmin Jiang |
IEEE Trans. Multim. | 3 |
| 2020 | Accelerating Epidemiological Investigation Analysis by Using NLP and Knowledge Reasoning: A Case Study on COVID-19
Jianmin Jiang, Jing Mei, Shaochun Li |
AMIA | 4 |
| 2020 | Adversarial Learning for Zero-Shot Domain Adaptation
Jianmin Jiang |
ECCV (21) | 2 |
| 2020 | Lossy Geometry Compression Of 3d Point Cloud Data Via An Adaptive Octree-Guided NetworkabstractIn this paper, we propose a deep learning based framework for point cloud geometry lossy compression via hybrid representation of point cloud. First, the input raw 3D point cloud data is adaptively decomposed into non-overlapping local patches through adaptive Octree decomposition and clustering. Second, a framework of point cloud auto-encoder network with quantization layer is proposed for learning compact latent feature representation from each patch. Specifically, the proposed point cloud auto-encoder networks with different input size are trained for achieving optimal rate-distortion (RD) performance. Final, bitstream specifications of proposed compression systems with additional signaled meta-data and header information are designed to support parallel decoding and successive reconstruction. Experimental results shows that our proposed method can achieve 40.20% bitrate saving in average than the existing standard Geometry based Point Cloud Compression (G-PCC) codec. Xuanzheng Wen, Xu Wang 0006, Junhui Hou, Lin Ma 0002, Yu Zhou 0027, Jianmin Jiang |
ICME | 6 |
| 2020 | A Viewport Prediction Framework for Panoramic VideosabstractPanoramic video is considered to be an attractive video format, since it provides the viewers with an immersive experience, such as virtual reality (VR) gaming. However, the viewers only focus on part of panoramic video, which is referred to as viewport. Hence, the resources consumed for distributing the remaining part of the panoramic video are wasted. It is intuitive to only deliver the video data within this viewport for reducing the distribution cost. Empirically, viewports within a time interval are highly correlated, hence the historical trajectory may be used for predicting the future viewports. On the other hand, a viewer tends to sustain attention on a specific object in a panoramic video. Motivated by these findings, we propose a deep learning-based viewport Prediction scheme, namely HOP, where the Historical viewport trajectory of viewers and Object tracking are jointly exploited by the long short-term memory (LSTM) networks. Additionally, our solution is capable of predicting multiple future viewports, while a single viewport prediction was supported by the state-of-the-art contributions. Simulation results show that our proposed HOP scheme outperforms the benchmarkers by up to 33.5% in terms of the prediction error. Jinting Tang, Yongkai Huo, Shaoshi Yang, Jianmin Jiang |
IJCNN | 4 |
| 2020 | Computerized Logo Synthesis with Wavelets-Enhanced Adversarial LearningabstractWhile logo design requires creative thoughts from artistic side, computerized logo synthesis could provide significant assistance in terms of workload reduction and productivity improvements. By applying wavelet transform to decompose the input logo images into four frequency bands, we introduce two new regularization terms into the GAN-based adversarial learning towards improved logo synthesis. As the LL band of images preserve the primary content information, we apply clustering to these LL bands to generate supervisory labels to regulate the logo generation and hence the logo synthesis can be predominated by a label-guided theme. To create varieties and diversities for the synthesized logos, we further establish a second regularization term out of the HH-band and enable the learning process to simulate the creativity illustrated by logo designers. Extensive experiments are carried out and, compared with the existing state of the arts, the results show that our proposed achieves overwhelmingly better performances in terms of the inception scores. Longchun Mao, Jianmin Jiang |
ISCAS | 3 |
| 2020 | Brain-media: A Dual Conditioned and Lateralization Supported GAN (DCLS-GAN) towards Visualization of Image-evoked Brain ActivitiesabstractEssentially, the current concept of multimedia is limited to presenting what people see in their eyes. What people think inside brains, however, remains a rich source of multimedia, such as imaginations of paradise and memories of good old days etc. In this paper, we propose a dual conditioned and lateralization supported GAN (DCLS-GAN) framework to learn and visualize the brain thoughts evoked by stimulating images and hence enable multimedia to reflect not only what people see but also what people think. To reveal such a new world of multimedia inside human brains, we coin such an attempt as "brain-media". By examining the relevance between the visualized image and the stimulation image, we are able to measure the efficiency of our proposed deep framework regarding the quality of such visualization and also the feasibility of exploring the concept of "brain-media". To ensure that such extracted multimedia elements remain meaningful, we introduce a dually conditioned learning technique in the proposed deep framework, where one condition is analyzing EEGs through deep learning to extract a class-dependent and more compact brain feature space utilizing the distinctive characteristics of hemispheric lateralization and brain stimulation, and the other is to extract expressive visual features assisting our automated analysis of brain activities as well as their visualizations aided by artificial intelligence. To support the proposed GAN framework, we create a combined-conditional space by merging the brain feature space with the visual feature space provoked by the stimuli. Extensive experiments are carried out and the results show that our proposed deep framework significantly outperforms the representative existing state-of-the-arts under several settings, especially in terms of both visualization and classification of brain responses to the evoked images. For the convenience of research dissemination, we make the source code openly accessible for downloading at GitHub. Ahmed Fares, Shenghua Zhong, Jianmin Jiang |
ACM Multimedia | 3 |
| 2020 | Content-Aware Cubemap Projection for Panoramic Image via Deep Q-Learning
Xu Wang 0006, Yu Zhou 0027, Longhao Zou, Jianmin Jiang |
MMM (2) | 5 |
| 2020 | Event-based functional decomposition
Jianmin Jiang, Huibiao Zhu, Qin Li 0002, Ping Gong 0004, Zhong Hong |
Inf. Comput. | 1 |
| 2020 | SA-Net: A deep spectral analysis network for image clustering
Jianmin Jiang |
Neurocomputing | 2 |
| 2020 | Global and local sensitivity guided key salient object re-augmentation for video saliency detection
Zheng Wang 0008, Ziqi Zhou 0002, Huchuan Lu, Jianmin Jiang |
Pattern Recognit. | 4 |
| 2020 | Learning to Explore Saliency for Stereoscopic Videos Via Component-Based InteractionabstractIn this paper, we devise a saliency prediction model for stereoscopic videos that learns to explore saliency inspired by the component-based interactions including spatial, temporal, as well as depth cues. The model first takes advantage of specific structure of 3D residual network (3D-ResNet) to model the saliency driven by spatio-temporal coherence from consecutive frames. Subsequently, the saliency inferred by implicit-depth is automatically derived based on the displacement correlation between left and right views by leveraging a deep convolutional network (ConvNet). Finally, a component-wise refinement network is devised to produce final saliency maps over time by aggregating saliency distributions obtained from multiple components. In order to further facilitate research towards stereoscopic video saliency, we create a new dataset including 175 stereoscopic video sequences with diverse content, as well as their dense eye fixation annotations. Extensive experiments support that our proposed model can achieve superior performance compared to the state-of-the-art methods on all publicly available eye fixation datasets. Qiudan Zhang, Xu Wang 0006, Shiqi Wang 0001, Zhenhao Sun, Sam Kwong, Jianmin Jiang |
IEEE Trans. Image Process. | 6 |
| 2019 | A Simple Pooling-Based Design for Real-Time Salient Object DetectionabstractWe solve the problem of salient object detection by investigating how to expand the role of pooling in convolutional neural networks. Based on the U-shape architecture, we first build a global guidance module (GGM) upon the bottom-up pathway, aiming at providing layers at different feature levels the location information of potential salient objects. We further design a feature aggregation module (FAM) to make the coarse-level semantic information well fused with the fine-level features from the top-down path- way. By adding FAMs after the fusion operations in the top-down pathway, coarse-level features from the GGM can be seamlessly merged with features at various scales. These two pooling-based modules allow the high-level semantic features to be progressively refined, yielding detail enriched saliency maps. Experiment results show that our proposed approach can more accurately locate the salient objects with sharpened details and hence substantially improve the performance compared to the previous state-of-the-arts. Our approach is fast as well and can run at a speed of more than 30 FPS when processing a 300×400 image. Code can be found at http://mmcheng.net/poolnet/. Jiang-Jiang Liu 0001, Qibin Hou, Ming-Ming Cheng, Jiashi Feng, Jianmin Jiang |
CVPR | 5 |
| 2019 | Surface Reconstruction From Normals: A Robust DGP-Based Discontinuity Preservation ApproachabstractIn 3D surface reconstruction from normals, discontinuity preservation is an important but challenging task. However, existing studies fail to address the discontinuous normal maps by enforcing the surface integrability in the continuous domain. This paper introduces a robust approach to preserve the surface discontinuity in the discrete geometry way. Firstly, we design two representative normal incompatibility features and propose an efficient discontinuity detection scheme to determine the splitting pattern for a discrete mesh. Secondly, we model the discontinuity preservation problem as a light-weight energy optimization framework by jointly considering the discontinuity detection and the overall reconstruction error. Lastly, we further shrink the feasible solution space to reduce the complexity based on the prior knowledge. Experiments show that the proposed method achieves the best performance on an extensive 3D dataset compared with the state-of-the-arts in terms of mean angular error and computational complexity. Wuyuan Xie, Miaohui Wang, Mingqiang Wei, Jianmin Jiang, Harry Qin |
CVPR | 4 |
| 2019 | Learning to Explore Intrinsic Saliency for Stereoscopic VideoabstractThe human visual system excels at biasing the stereoscopic visual signals by the attention mechanisms. Traditional methods relying on the low-level features and depth relevant information for stereoscopic video saliency prediction have fundamental limitations. For example, it is cumbersome to model the interactions between multiple visual cues including spatial, temporal, and depth information as a result of the sophistication. In this paper, we argue that the high-level features are crucial and resort to the deep learning framework to learn the saliency map of stereoscopic videos. Driven by spatio-temporal coherence from consecutive frames, the model first imitates the mechanism of saliency by taking advantage of the 3D convolutional neural network. Subsequently, the saliency originated from the intrinsic depth is derived based on the correlations between left and right views in a data-driven manner. Finally, a Convolutional Long Short-Term Memory (Conv-LSTM) based fusion network is developed to model the instantaneous interactions between spatio-temporal and depth attributes, such that the ultimate stereoscopic saliency maps over time are produced. Moreover, we establish a new large-scale stereoscopic video saliency dataset (SVS) including 175 stereoscopic video sequences and their fixation density annotations, aiming to comprehensively study the intrinsic attributes for stereoscopic video saliency detection. Extensive experiments show that our proposed model can achieve superior performance compared to the state-of-the-art methods on the newly built dataset for stereoscopic videos. Qiudan Zhang, Xu Wang 0006, Shiqi Wang 0001, Shikai Li, Sam Kwong, Jianmin Jiang |
CVPR | 6 |
| 2019 | Conditional Coupled Generative Adversarial Networks for Zero-Shot Domain AdaptationabstractMachine learning models trained in one domain perform poorly in the other domains due to the existence of domain shift. Domain adaptation techniques solve this problem by training transferable models from the label-rich source domain to the label-scarce target domain. Unfortunately, a majority of the existing domain adaptation techniques rely on the availability of the target-domain data, and thus limit their applications to a small community across few computer vision problems. In this paper, we tackle the challenging zero-shot domain adaptation (ZSDA) problem, where the target-domain data is non-available in the training stage. For this purpose, we propose conditional coupled generative adversarial networks (CoCoGAN) by extending the coupled generative adversarial networks (CoGAN) into a conditioning model. Compared with the existing state of the arts, our proposed CoCoGAN is able to capture the joint distribution of dual-domain samples in two different tasks, i.e. the relevant task (RT) and an irrelevant task (IRT). We train the CoCoGAN with both source-domain samples in RT and the dual-domain samples in IRT to complete the domain adaptation. While the former provide the high-level concepts of the non-available target-domain data, the latter carry the sharing correlation between the two domains in RT and IRT. To train the CoCoGAN in the absence of the target-domain data for RT, we propose a new supervisory signal, i.e. the alignment between representations across tasks. Extensive experiments carried out demonstrate that our proposed CoCoGAN outperforms existing state of the arts in image classifications. Jianmin Jiang |
ICCV | 2 |
| 2019 | GEOCAPSNET: Ground to Aerial View Image Geo-Localization using Capsule NetworkabstractThe task of cross-view image geo-localization aims to determine the geo-location (GPS coordinates) of a query ground-view image by matching it with the GPS-tagged aerial (satellite) images in a reference dataset. Due to the dramatic changes of viewpoint, matching the cross-view images is challenging. In this paper, we propose the GeoCapsNet based on the capsule network for ground-to-aerial image geo-localization. The network first extracts features from both ground and aerial images via standard convolution layers and the capsule layers further encode the features to model the spatial feature hierarchies and enhance the representation power. Moreover, we introduce a simple and effective weighted soft-margin triplet loss with online batch hard sample mining, which can greatly improve the image retrieval accuracy. Experimental results show that our GeoCapsNet significantly outperforms the state-of-the-art approaches on two benchmark datasets. Chen Chen 0001, Yingying Zhu 0001, Jianmin Jiang |
ICME | 4 |
| 2019 | Spectral Analysis Network for Deep Representation Learning and Image ClusteringabstractDeep representation learning is a crucial procedure in multimedia analysis and attracts increasing attention. Most of the popular techniques rely on convolutional neural network and require a large amount of labeled data in the training procedure. However, it is time consuming or even impossible to obtain the label information in some tasks due to cost limitation. Thus, it is necessary to develop unsupervised deep representation learning techniques. This paper proposes a new network structure for unsupervised deep representation learning based on spectral analysis, which is a popular technique with solid theory foundations. Compared with the existing spectral analysis methods, the proposed network structure has at least three advantages. Firstly, it can identify the local similarities among images in patch level and thus more robust against occlusion. Secondly, through multiple consecutive spectral analysis procedures, the proposed network can learn more clustering-friendly representations and is capable to reveal the deep correlations among data samples. Thirdly, it can elegantly integrate different spectral analysis procedures, so that each spectral analysis procedure can have their individual strengths in dealing with different data sample distributions. Extensive experimental results show the effectiveness of the proposed methods on various image clustering tasks. Adrian Hilton 0001, Jianmin Jiang |
ICME | 3 |
| 2019 | An Attentional-LSTM for Improved Classification of Brain Activities Evoked by ImagesabstractMultimedia stimulation of brain activities is not only becoming an emerging area for intensive research, but also achieved significant progresses towards classification of brain activities and interpretation of brain understanding of multimedia content. To exploit the characteristics of EEG signals in capturing human brain activities, we propose a region-dependent and attention-driven bi-directional LSTM network (RA-BiLSTM) for image evoked brain activity classification. Inspired by the hemispheric lateralization of human brains, the proposed RA-BiLSTM extracts additional information at regional level to strengthen and emphasize the differences between two hemispheres. In addition, we propose a new attentional-LSTM by adding an extra attention gate to: (i) measure and seize the importance of channel-based spatial information, and (ii) support the proposed RA-BiLSTM to capture the dynamic correlations hidden from both the past and the future in the current state across EEG sequences. Extensive experiments are carried out and the results demonstrate that our proposed RA-BiLSTM not only achieves effective classification of brain activities on evoked image categories, but also significantly outperforms the existing state of the arts. Shenghua Zhong, Ahmed Fares, Jianmin Jiang |
ACM Multimedia | 3 |
| 2019 | Schedulability analysis for real-time mobile systems (S)abstractAutonomous driving systems are complex real-time mobile systems.To guarantee their safety and security, the mobile objects (agents) in these systems must be isolated from each other so that they do not collide with each other.Since isolation means two or more mobile objects cannot be located in the same area at the same time, a scheduling policy is required to control the movement of these mobile objects.However, traditional scheduling theories are based on task scheduling which is coarsegrained and cannot be directly used for fine-grained isolation controls.In this paper, we first propose an event-based formal model called a time dependency structure which is used to model and analyze real-time mobile systems.Then, an event-based schedule is defined.Finally, we analyze the schedulability of isolation-that is, checking whether a given schedule ensures the isolation relationship among mobile objects or not. Jianmin Jiang, Zhong Hong, Hongping Shu, Zeng Qiong |
SEKE | 3 |
| 2019 | A robust deep style transfer for headshot portraits
Meiqin Guo, Jianmin Jiang |
Neurocomputing | 2 |
| 2019 | Content-adaptive selective steganographer detection via embedding probability estimation deep networks
Mingjie Zheng 0002, Jianmin Jiang, Songtao Wu, Shenghua Zhong, Yan Liu 0004 |
Neurocomputing | 2 |
| 2019 | Video summarization via spatio-temporal deep architecture
Shenghua Zhong, Jiaxin Wu 0001, Jianmin Jiang |
Neurocomputing | 3 |
| 2019 | SFAD: Toward effective anomaly detection based on session feature similarity
Ruliang Xiao, Jiawei Su, Xin Du 0003, Jianmin Jiang, Xinhong Lin, Li Lin 0001 |
Knowl. Based Syst. | 4 |
| 2019 | Image interpolation using convolutional neural networks with deep recursive residual learning
Kwok-Wai Hung, Jianmin Jiang |
Multim. Tools Appl. | 3 |
| 2019 | SG-FCN: A Motion and Memory-Based Deep Learning Model for Video Saliency DetectionabstractData-driven saliency detection has attracted strong interest as a result of applying convolutional neural networks to the detection of eye fixations. Although a number of image-based salient object and fixation detection models have been proposed, video fixation detection still requires more exploration. Different from image analysis, motion and temporal information is a crucial factor affecting human attention when viewing video sequences. Although existing models based on local contrast and low-level features have been extensively researched, they failed to simultaneously consider interframe motion and temporal information across neighboring video frames, leading to unsatisfactory performance when handling complex scenes. To this end, we propose a novel and efficient video eye fixation detection model to improve the saliency detection performance. By simulating the memory mechanism and visual attention mechanism of human beings when watching a video, we propose a step-gained fully convolutional network by combining the memory information on the time axis with the motion information on the space axis while storing the saliency information of the current frame. The model is obtained through hierarchical training, which ensures the accuracy of the detection. Extensive experiments in comparison with 11 state-of-the-art methods are carried out, and the results show that our proposed model outperforms all 11 methods across a number of publicly available datasets. Meijun Sun, Ziqi Zhou 0002, Qinghua Hu, Zheng Wang 0008, Jianmin Jiang |
IEEE Trans. Cybern. | 5 |
| 2019 | Adaptive Bi-Weighting Toward Automatic Initialization and Model Selection for HMM-Based Hybrid Meta-Clustering EnsemblesabstractTemporal data clustering can provide underpinning techniques for the discovery of intrinsic structures, which proved important in condensing or summarizing information demanded in various fields of information sciences, ranging from time series analysis to sequential data understanding. In this paper, we propose a novel hidden Markov model (HMM)-based hybrid meta-clustering ensemble with bi-weighting scheme to solve the problems of initialization and model selection associated with temporal data clustering. To improve the performance of the ensemble techniques, the proposed bi-weighting scheme adaptively examines the partition process and hence optimizes the fusion of consensus functions. Specifically, three consensus functions are used to combine the input partitions, generated by HMM-based K -models under different initializations, into a robust consensus partition. An optimal consensus partition is then selected from the three candidates by a normalized mutual information-based objective function. Finally, the optimal consensus partition is further refined by the HMM-based agglomerative clustering algorithm in association with dendrogram-based similarity partitioning algorithm, leading to the advantage that the number of clusters can be automatically and adaptively determined. Extensive experiments on synthetic data, time series, and real-world motion trajectory datasets illustrate that our proposed approach outperforms all the selected benchmarks and hence providing promising potentials for developing improved clustering tools for information analysis and management. Yun Yang 0003, Jianmin Jiang |
IEEE Trans. Cybern. | 2 |
| 2019 | Semisupervised Regression With Optimized Rank for Matrix Data ClassificationabstractThere has been growing interest in developing more effective algorithms for matrix data classification. At present, most of the existing vector-based classifications involve vectorization process, which results in two main problems. First, the underlying structural information is disregarded. Second, the vectorization of a matrix incurs the creation of a vector with potentially very high dimensionality, which may lead to overfitting when the number of training data is small. To avoid such problems, we propose a new matrix-based regression algorithm for classification, in which the input matrices to be classified are directly used to learn two regression matrices for each order of the input matrix. To further explore the discrimination information, we add a joint ℓ2,1-norm on two regression matrices, which endows the algorithm optimized regression rank by uncovering common sparse columns in the two regression matrices. To further boost the classification performance, we incorporate a semisupervised learning process, which leverages both labeled and unlabeled data to enhance the training process. Experiments on public benchmark datasets show that our method outperforms a number of the existing state-of-the-art classification methods even when only few labeled training samples are provided. Jianguang Zhang, Jianmin Jiang, Yahong Han |
IEEE Trans. Cybern. | 2 |
| 2019 | A Context-Supported Deep Learning Framework for Multimodal Brain Imaging ClassificationabstractOver the past decade, “content-based” multimedia systems have realized success. By comparison, brain imaging and classification systems demand more efforts for improvement with respect to accuracy, generalization, and interpretation. The relationship between electroencephalogram (EEG) signals and corresponding multimedia content needs to be further explored. In this paper, we integrate implicit and explicit learning modalities into a context-supported deep learning framework. We propose an improved solution for the task of brain imaging classification via EEG signals. In our proposed framework, we introduce a consistency test by exploiting the context of brain images and establishing a mapping between visual-level features and cognitive-level features inferred based on EEG signals. In this way, a multimodal approach can be developed to deliver an improved solution for brain imaging and its classification based on explicit learning modalities and research from the image processing community. In addition, a number of fusion techniques are investigated in this work to optimize individual classification results. Extensive experiments have been carried out, and their results demonstrate the effectiveness of our proposed framework. In comparison with the existing state-of-the-art approaches, our proposed framework achieves superior performance in terms of not only the standard visual object classification criteria, but also the exploitation of transfer learning. For the convenience of research dissemination, we make the source code publicly available for downloading at GitHub (https://github.com/aneeg/dual-modal-learning). Jianmin Jiang, Ahmed Fares, Shenghua Zhong |
IEEE Trans. Hum. Mach. Syst. | 1 |
| 2019 | Isolation Modeling and Analysis Based on MobilityabstractIn a mobile system, mobility refers to a change in position of a mobile object with respect to time and its reference point, whereas isolation means the isolation relationship between mobile objects under some scheduling policies. Inspired by event-based formal models and the ambient calculus, we first propose the two types of special events, entering and exiting an ambient, as movement events to model and analyze mobility. Based on mobility, we then introduce the notion of the isolation of mobile objects for ambients. To ensure the isolation, a priority policy needs to be used to schedule the movement of mobile objects. However, traditional scheduling policies focus on task scheduling and depend on the strong hypothesis: The scheduled tasks are independent—that is, the scheduled tasks do not affect each other. In a practical mobile system, mobile objects and ambients interact with each other. It is difficult to separate a mobile system into independent tasks. We finally present an automatic approach for generating a priority scheduling policy without considering the preceding assumption. The approach can guarantee the isolation of the mobile objects for ambients in a mobile system. Experiments demonstrate these results. Jianmin Jiang, Huibiao Zhu, Qin Li 0002, Zhong Hong, Ping Gong 0004 |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2018 | An Unsupervised Deep Learning Framework via Integrated Optimization of Representation Learning and GMM-Based Modeling
Jianmin Jiang |
ACCV (1) | 2 |
| 2018 | Region level Bi-directional Deep Learning Framework for EEG-based Image Classification
Ahmed Fares, Shenghua Zhong, Jianmin Jiang |
BIBM | 3 |
| 2018 | An Adaptive Box-Normalization Stock Index Trading Strategy Based on Reinforcement Learning
Yingying Zhu 0001, Jianmin Jiang |
ICONIP (3) | 3 |
| 2018 | Revised Spatial Transformer Network towards Improved Image Super-resolutionsabstractIn this paper, we propose two approaches to generate a high resolution (HR) image from a low resolution (LR) version. It is commonly referred as single-image super-resolution(SISR). Our approaches are inspired by the spatial transformer (ST) module and Very Deep Convolutional Network (VDSR). The spatial transformer module is the neural network originally used for geometric transformations of images, while the VDSR is for image super-resolution. In the first approach, we propose to add the ST module with the VDSR to generate HR images. The use of the spatial transform with VDSR makes the network more robust to the geometric transformations. We propose the second approach to replace the convolutional neural network (CNN) used in the ST module with VDSR network. The replacement of CNN leads to improvements in the performance of ST module. The simulation results confirm that the feasibility of combing ST module and VDSR for super-resolution reconstruction, where the performance of the combination of ST module and VDSR is comparable to the VDSR alone in terms of Peak signal-to-noise ratio (PSNR) and structural Similarity Index Measurement (SSIM). Hence, the revised spatial transformer network can be used in the future for simultaneous geometric transformation and image super-resolution, which solve the practical applications of image super-resolution in real life. Hossam M. Kasem, Kwok-Wai Hung, Jianmin Jiang |
ICPR | 3 |
| 2018 | Video Restoration Using Convolutional Neural Networks for Low-Level FPGAs
Kwok-Wai Hung, Chaoming Qiu, Jianmin Jiang |
KSEM (2) | 3 |
| 2018 | Steganographer Detection based on Multiclass Dilated Residual NetworksabstractSteganographer detection task is to identify criminal users, who attempt to conceal confidential information by steganography methods, among a large number of innocent users. The significant challenge of the task is how to collect the evidences to identify the guilty user with suspicious images, which are embedded with secret messages generating by unknown steganography and payload. Unfortunately, existing methods for steganalysis were served for the binary classification. It makes them harder to classify the images with different kinds of payloads, especially when the payloads of images in test dataset have not been provided in advance. In this paper, we propose a novel steganographer detection method based on multiclass deep neural networks. In the training stage, the networks are trained to classify the images with six types of payloads. The networks can preserve even strengthen the weak stego signals from secret messages in much larger receptive filed by virtue of residual and dilated residual learning. In the inference stage, the learnt model is used to extract the discriminative features, which can capture the difference between guilty users and innocent users. A series of empirical experimental results demonstrate that the proposed method achieves good performance in spatial and frequency domains even though the embedding payload is low. The proposed method achieves a higher level of robustness of inter-steganographic algorithms and can provide a possible solution to address the payload mismatch problem Mingjie Zheng 0002, Shenghua Zhong, Songtao Wu, Jianmin Jiang |
ICMR | 4 |
| 2018 | Data Augmentation for EEG-Based Emotion Recognition with Deep Convolutional Neural Networks
Shenghua Zhong, Jianfeng Peng, Jianmin Jiang, Yan Liu 0004 |
MMM (2) | 4 |
| 2018 | Modeling mobility and communication in a unified way (S)abstractTraditional formalisms model communication and mobility in a separate way.This may cause complex name management and complex analysis for a communicating and mobile system.In this paper, following the ambient calculus [2], we first propose two types of special events, entering and exiting an ambient, as movement events and discuss the relationship of ambients based on mobility.Then a communication model is introduced based on message movement, which can represent synchronous communication, asynchronous communication and broadcasting communication in a unified way.Finally, we show that such a communication model is contained in a general event-based formal model called a dependency structure [4], [5]. Jianmin Jiang, Zhong Hong |
SEKE | 1 |
| 2018 | Decomposition and Composition of Sequence DiagramsabstractAs a semi-formal model, a UML sequence diagram can be widely useful for modelling and analyzing functional behaviors of a developing software system. However, reasoning about decomposition and composition of sequence diagrams has not been addressed adequately. The traditional decomposition method is based on the ideas from a global system to local components, and local components are recomposed into the whole system according to communication assumptions or other mechanisms. In this paper, we propose an event-based formal semantics for sequence diagrams and present a new functional decomposition approach based on sequence diagrams. The new decomposition idea is from a global system to global subsystems. Under the decomposition method global subsystems can be recomposed into the global system using the composition operation called union without necessarily considering assumptions such as communication. Jianmin Jiang, Zhong Hong |
TASE | 2 |
| 2018 | Deep learning based image Super-resolution for nonlinear lens distortions
Qinglong Chang, Kwok-Wai Hung, Jianmin Jiang |
Neurocomputing | 3 |
| 2018 | A deep-learning based feature hybrid framework for spatiotemporal saliency detection inside videos
Zheng Wang 0008, Jinchang Ren, Meijun Sun, Jianmin Jiang |
Neurocomputing | 5 |
| 2018 | Haze removal method for natural restoration of images with sky
Yingying Zhu 0001, Gaoyang Tang, Xiaoyan Zhang 0002, Jianmin Jiang, Qi Tian 0001 |
Neurocomputing | 4 |
| 2018 | Image up-sampling using deep cascaded neural networks in dual domains for images down-sampled in DCT domain
Kwok-Wai Hung, Jianmin Jiang |
J. Vis. Commun. Image Represent. | 3 |
| 2018 | Deep intensity guidance based compression artifacts reduction for depth map
Xu Wang 0006, Yun Zhang 0002, Lin Ma 0002, Sam Kwong, Jianmin Jiang |
J. Vis. Commun. Image Represent. | 6 |
| 2018 | Editorial: Industrial Internet of Things (I2oT)
Leandros Maglaras, Lei Shu 0001, Athanasios Maglaras, Jianmin Jiang, Helge Janicke, Dimitrios Katsaros 0001, Tiago Cruz 0001 |
Mob. Networks Appl. | 4 |
| 2018 | Foveated convolutional neural networks for video summarization
Jiaxin Wu 0001, Shenghua Zhong, Zheng Ma 0004, Stephen J. Heinen, Jianmin Jiang |
Multim. Tools Appl. | 5 |
| 2018 | Facial expression synthesis with direction field preservation based mesh deformation and lighting fitting based wrinkle mapping
Weicheng Xie 0001, LinLin Shen, Meng Yang 0001, Jianmin Jiang |
Multim. Tools Appl. | 4 |
| 2018 | Tensor learning and automated rank selection for regression-based video classification
Jianguang Zhang, Yanbin Liu 0003, Jianmin Jiang |
Multim. Tools Appl. | 3 |
| 2018 | Rank-Optimized Logistic Matrix Regression toward Improved Matrix Data ClassificationabstractWhile existing logistic regression suffers from overfitting and often fails in considering structural information, we propose a novel matrix-based logistic regression to overcome the weakness. In the proposed method, 2D matrices are directly used to learn two groups of parameter vectors along each dimension without vectorization, which allows the proposed method to fully exploit the underlying structural information embedded inside the 2D matrices. Further, we add a joint [Formula: see text]-norm on two parameter matrices, which are organized by aligning each group of parameter vectors in columns. This added co-regularization term has two roles-enhancing the effect of regularization and optimizing the rank during the learning process. With our proposed fast iterative solution, we carried out extensive experiments. The results show that in comparison to both the traditional tensor-based methods and the vector-based regression methods, our proposed solution achieves better performance for matrix data classifications. Jianguang Zhang, Jianmin Jiang |
Neural Comput. | 2 |
| 2018 | Bi-weighted ensemble via HMM-based approaches for temporal data clustering
Yun Yang 0003, Jianmin Jiang |
Pattern Recognit. | 2 |
| 2018 | Decomposition of UML activity diagramsabstractSummary In software engineering, UML activity diagrams in general can be useful for a modeling system functional behavior, ranging from the sequences of activities/actions from business processes within an organization or among organizations down to the detail of an algorithm. The stepwise refinement process makes activity diagrams more and more complex. To guarantee the behavior consistency and correctness under refinement, the activity diagrams must be decomposed according to the divide‐and‐conquer strategy. Traditional decomposition methods adopt manual techniques and cannot ensure the independence and completeness of the obtained subdiagrams. In this paper, a novel decomposition approach is proposed, which can automatically divide an activity diagram into atomic and correct subdiagrams (subdiagrams without abnormal behavioral problems) at the same level. When such an activity diagram specifies the whole functional behavior of a software system, the approach can in fact decompose a system into multiple atomic subsystems. Every atomic subsystem is a completely independent system. It may be independently developed, independently tested, and independently deployed. The method facilitates the management, development, and maintenance of a software system. With the help of a prototype tool, a case study demonstrates the decomposition method. Huifeng Chen, Jianmin Jiang, Zhong Hong |
Softw. Pract. Exp. | 2 |
| 2017 | DCT-based image up-sampling using anchored neighborhood regressionabstractDown-sampling in the discrete cosine transform (DCT) domain is preferable for images coded by DCT transform, such as JPEG/MJPEG/H.264, etc. Recent researches show that the truncated high-frequency DCT coefficients during the DCT down-sampling process can be estimated by learning the correlations between low-frequency and high-frequency DCT coefficients. In this paper, we propose to utilize the powerful super-resolution framework using sparse dictionaries with anchored neighborhood regression to significantly improve the accuracy of the estimated high-frequency DCT coefficients. Experimental results show that the proposed framework outperforms the state-of-the-art DCT-based up-sampling methods in terms of PSNR (0.3-1.63dB) and SSIM values for standard image datasets Set5 and Set14, while the computational time of the proposed method is 23× times faster than the state-of-the-art learning-based method using k-NN MMSE due to pre-computation of the ridge regression during the training process. Kwok-Wai Hung, Jianmin Jiang, Qinglong Chang, Xu Wang 0006 |
ICIP | 2 |
| 2017 | Study of subjective and objective quality assessment for screen content imagesabstractIn this paper, we present the results of a recent large-scale subjective study of image quality on a collection of screen contents distorted by a variety of application-relevant processes. With the development of multi-device interactive multimedia applications, metrics to predict the visual quality of screen content images (SCIs) as perceived by subjects are becoming fundamentally important. For developing the objective image quality assessment (IQA) method, there is a need for large-scale public database with diversity of distorted types and scene contents, and available subjective scores of distorted SCIs. The resulting Immersive Media Laboratory screen content image quality database (IML-SCIQD) contains 1250 distorted SCIs from 25 reference SCIs with 10 distortion types. Each image was rated by 35 human observers, and the different mean opinion scores (DMOS) were obtained after data processing. The performance comparison of 17 state-of-the-arts, publicly available IQA algorithms are evaluated on the new database. The database will be available online in our project website. Xu Wang 0006, Yingying Zhu 0001, Yun Zhang 0002, Jianmin Jiang, Sam Kwong |
ICIP | 5 |
| 2017 | Steganographer detection via deep residual networkabstractSteganographer detection problem is to identify culprit actors, who try to hide confidential information with steganography, among many innocent actors. This task has significant challenges, including various embedding steganographic algorithms and payloads, which are usually avoided in steganalysis under laboratory conditions. In this paper, we propose a novel steganographer detection model based on deep residual network. The proposed method strengthens the signal coming from secret messages, which is beneficial for the discrimination between guilty actors and innocent actors. Comprehensive experiments demonstrate that the proposed model achieves very low detection error rates in steganographer detection task. It also outperforms the classical rich model method and other CNN based method. Moreover, the model shows the robustness of inter-steganographic algorithms and inter-payloads. Mingjie Zheng 0002, Shenghua Zhong, Songtao Wu, Jianmin Jiang |
ICME | 4 |
| 2017 | Super-Resolution for Images with Barrel Lens Distortions
Mei Su 0003, Kwok-Wai Hung, Jianmin Jiang |
KSEM | 3 |
| 2017 | A Novel Blemish Detection Algorithm for Camera Quality Testing
Kwok-Wai Hung, Jianmin Jiang |
KSEM | 3 |
| 2017 | Interpretation of users' feedback via swarmed particles for content-based image retrieval
Yingying Zhu 0001, Jianmin Jiang, Wenlong Han, Qi Tian 0001 |
Inf. Sci. | 2 |
| 2017 | Semi-supervised tensor learning for image classification
Jianguang Zhang, Yahong Han, Jianmin Jiang |
Multim. Syst. | 3 |
| 2017 | A novel clustering method for static video summarization
Jiaxin Wu 0001, Shenghua Zhong, Jianmin Jiang, Yunyun Yang |
Multim. Tools Appl. | 3 |
| 2017 | Authentication Protocols for Internet of Things: A Comprehensive SurveyabstractIn this paper, a comprehensive survey of authentication protocols for Internet of Things (IoT) is presented. Specifically more than forty authentication protocols developed for or applied in the context of the IoT are selected and examined in detail. These protocols are categorized based on the target environment: (1) Machine to Machine Communications (M2M), (2) Internet of Vehicles (IoV), (3) Internet of Energy (IoE), and (4) Internet of Sensors (IoS). Threat models, countermeasures, and formal security verification techniques used in authentication protocols for the IoT are presented. In addition a taxonomy and comparison of authentication protocols that are developed for the IoT in terms of network model, specific security goals, main processes, computation complexity, and communication overhead are provided. Based on the current survey, open issues are identified and future research directions are proposed. Mohamed Amine Ferrag, Leandros Maglaras, Helge Janicke, Jianmin Jiang, Lei Shu 0001 |
Secur. Commun. Networks | 4 |
| 2017 | Event-Based Mobility Modeling and AnalysisabstractMobility is a critical issue that must be considered during the modeling and analyzing of a mobile system. At a high abstract level, event-based models can directly specify a mobile system without the introduction of additional mechanisms. In this article, we first propose two types of special events, entering and exiting an ambient, as movement events. Next, based on the movement events, we introduce the notion of a movement path and propose a feasible movement criterion (deciding whether a given movement path of a mobile object (agent) is feasible or not in terms of spatiotemporal topological relationships of ambients). Then, we investigate how a message movement--based communication model represents synchronous communication, asynchronous communication, and broadcast communication in a unified way. Finally, we use movement event sequences to discuss the exclusivity of ambients (an ambient only allows one mobile object to occupy (enter) it at any moment) and show that a priority scheduling control policy can guarantee exclusivity. Accordingly, we propose a correct movement criterion—that is, a correct movement path is feasible and satisfies the exclusivity of ambients. Case studies demonstrate these results. Jianmin Jiang, Huibiao Zhu, Qin Li 0002, Ping Gong 0004, Zhong Hong, Donghuo Chen |
ACM Trans. Cyber Phys. Syst. | 1 |
| 2017 | Semi-Supervised Image-to-Video Adaptation for Video Action RecognitionabstractHuman action recognition has been well explored in applications of computer vision. Many successful action recognition methods have shown that action knowledge can be effectively learned from motion videos or still images. For the same action, the appropriate action knowledge learned from different types of media, e.g., videos or images, may be related. However, less effort has been made to improve the performance of action recognition in videos by adapting the action knowledge conveyed from images to videos. Most of the existing video action recognition methods suffer from the problem of lacking sufficient labeled training videos. In such cases, over-fitting would be a potential problem and the performance of action recognition is restrained. In this paper, we propose an adaptation method to enhance action recognition in videos by adapting knowledge from images. The adapted knowledge is utilized to learn the correlated action semantics by exploring the common components of both labeled videos and images. Meanwhile, we extend the adaptation method to a semi-supervised framework which can leverage both labeled and unlabeled videos. Thus, the over-fitting can be alleviated and the performance of action recognition is improved. Experiments on public benchmark datasets and real-world datasets show that our method outperforms several other state-of-the-art action recognition methods. Jianguang Zhang, Yahong Han, Jinhui Tang 0001, Qinghua Hu, Jianmin Jiang |
IEEE Trans. Cybern. | 5 |
| 2017 | Visual Object Tracking With Partition Loss SchemesabstractObject tracking is a fundamental task for building vision systems of automatic transportation. Despite demonstrated success in this active research field, it is still difficult to cope with complicated appearance changes caused by background clutters, illumination change, scale variation, deformation, rotation, and occlusion, etc. Due to these challenging factors, a target bounding box tends to easily contain a disturbing context of background, which may lead to an inaccurate localization if some key parts of the foreground object share an excess of information loss. In this paper, we propose an online algorithm using the local loss features to alleviate the drift problem during tracking. An adaptive block-division appearance model is constructed to exploit patch-based loss representations by decomposing sample region sequences into a set of subblocks. The basic purpose of partition coefficients is to indicate local relevance and effectively capture spatial correlation through measuring image similarity. Namely, they can strengthen positive effects of discriminative patches within locating bounding boxes and weaken negative impacts of disturbing context possibly included in the surrounding regions. The object state estimation of a moving target is then formulated as an integrated likelihood evaluation on the ensemble loss. Experimental results on a suite of representative video sequences of realistic scenarios demonstrate the superiority of the proposed method to several state-of-the-art tracking approaches in terms of both accuracy and robustness. Xiaochun Cao, Jianmin Jiang |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2017 | A Novel Transient Wrinkle Detection Algorithm and Its Application for Expression SynthesisabstractBecause facial wrinkle is a representative feature of facial expression, automatic wrinkle detection has been an important and challenging topic for expression simulation, recognition, and animation. Recently, most works about wrinkle detection have focused on permanent wrinkles (e.g., age wrinkles), which are usually linear shapes, whereas the detection of transient wrinkles (e.g., expression wrinkles) has not been sufficiently studied because of their shape diversity and complexity. In this work, a novel algorithm for automatic detection of transient wrinkles with linear, fixed, and chaotic shapes is proposed, which largely consists of edge pair matching, active-appearance-model-based wrinkle structure location, and support-vector-machine-based wrinkle classification. The proposed wrinkle detector is applied for expression synthesis and an improved Poisson wrinkle mapping approach is proposed. Experimental results illustrate the competitiveness of the proposed wrinkle detector in detecting different transient wrinkles. Compared with state-of-the-art algorithms, the proposed approach yields complete and accurate wrinkle centers. The expression synthesized by the improved wrinkle mapping is also much more realistic. Weicheng Xie 0001, LinLin Shen, Jianmin Jiang |
IEEE Trans. Multim. | 3 |
| 2016 | Visual Orientation Inhomogeneity Based Convolutional Neural NetworksabstractThe details of oriented visual stimuli are better resolved when they are horizontal or vertical rather than oblique. This "oblique effect" has been researched and confirmed in numerous research studies, including behavioral studies and neurophysiological and neuroimaging findings. Although the "oblique effect" has influence in many fields, little research integrated it into computational models. In this paper, we try to explore this inhomogeneity of visual orientation based on Convolutional neural networks (CNNs) in image recognition. We validate that visual orientation inhomogeneity CNNs can achieve comparable performance with higher computational efficiency on various datasets. We can also get the conclusion that, compared with the cardinal information, oblique information is indeed less useful in natural color image recognition. Through the exploration of the proposed model on image recognition, we gain more understanding of the inhomogeneity of visual orientation. It also illuminates a wide range of opportunities for integrating the inhomogeneity of visual orientation with other computational models. Shenghua Zhong, Jiaxin Wu 0001, Yingying Zhu 0001, Peiqi Liu, Jianmin Jiang, Yan Liu 0004 |
ICTAI | 5 |
| 2016 | Transfer Learning Based on A+ for Image Super-Resolution
Mei Su 0003, Shenghua Zhong, Jianmin Jiang |
KSEM | 3 |
| 2016 | Combining ensemble methods and social network metrics for improving accuracy of OCSVM on intrusion detection in SCADA systems
Leandros Maglaras, Jianmin Jiang, Tiago Cruz 0001 |
J. Inf. Secur. Appl. | 2 |
| 2016 | Decomposition-based tensor learning regression for improved classification of multimedia
Jianguang Zhang, Jianmin Jiang |
J. Vis. Commun. Image Represent. | 2 |
| 2016 | Semi-supervised feature selection via hierarchical regression for web image classification
Jianguang Zhang, Yahong Han, Jianmin Jiang |
Multim. Syst. | 4 |
| 2016 | Tucker decomposition-based tensor learning for human action recognition
Jianguang Zhang, Yahong Han, Jianmin Jiang |
Multim. Syst. | 3 |
| 2016 | Video super-resolution using an adaptive superpixel-guided auto-regressive model
Kun Li 0001, Yanming Zhu 0001, Jing-Yu Yang 0002, Jianmin Jiang |
Pattern Recognit. | 4 |
| 2016 | An Optimized Higher Order CRF for Automated Labeling and Segmentation of Video ObjectsabstractIn this paper, we propose an optimized higher order conditional random field (CRF) labeling approach toward automated video object segmentation. Our approach introduces a computerized optimization scheme to fine tune the CRF-associated parameters, and hence make the labeling of segmented regions optimal in formulating the video objects. In comparison with the existing efforts using CRF, our optimized CRF has introduced a number of novel features, which can be highlighted as: 1) higher order CRF labeling is made adaptive to video content changes via a windowed dynamics; 2) fusion of multiple features is automatically optimized via fuzzy modeling of incoming video content and regression of parameters; 3) unary potential of higher order CRF labeling is modulated by the shortest path between neighboring regions to improve the effectiveness of higher order CRF labeling; and 4) making the algorithm affordable for a simpler graph-based video segmentation to reduce the overall computing cost, making the proposed algorithm more efficient without compromising on its performances. Jianmin Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2016 | A Cybersecurity Detection Framework for Supervisory Control and Data Acquisition SystemsabstractThis paper presents a distributed intrusion detection system (DIDS) for supervisory control and data acquisition (SCADA) industrial control systems, which was developed for the CockpitCI project. Its architecture was designed to address the specific characteristics and requirements for SCADA cybersecurity that cannot be adequately fulfilled by techniques from the information technology world, thus requiring a domain-specific approach. DIDS components are described in terms of their functionality, operation, integration, and management. Moreover, system evaluation and validation are undertaken within an especially designed hybrid testbed emulating the SCADA system for an electrical distribution grid. Tiago Cruz 0001, Luís Rosa 0001, Jorge Proença, Leandros Maglaras, Matthieu Aubigny, Leonid Lev, Jianmin Jiang, Paulo Simões 0001 |
IEEE Trans. Ind. Informatics | 7 |
| 2016 | Hybrid Sampling-Based Clustering Ensemble With Global and Local ConstitutionsabstractAmong a number of ensemble learning techniques, boosting and bagging are the most popular sampling-based ensemble approaches for classification problems. Boosting is considered stronger than bagging on noise-free data set with complex class structures, whereas bagging is more robust than boosting in cases where noise data are present. In this paper, we extend both ensemble approaches to clustering tasks, and propose a novel hybrid sampling-based clustering ensemble by combining the strengths of boosting and bagging. In our approach, the input partitions are iteratively generated via a hybrid process inspired by both boosting and bagging. Then, a novel consensus function is proposed to encode the local and global cluster structure of input partitions into a single representation, and applies a single clustering algorithm to such representation to obtain the consolidated consensus partition. Our approach has been evaluated on 2-D-synthetic data, collection of benchmarks, and real-world facial recognition data sets, which show that the proposed technique outperforms the existing benchmarks for a variety of clustering tasks. Yun Yang 0003, Jianmin Jiang |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2015 | A solution of dynamic VMs placement problem for energy consumption optimization based on evolutionary game theory
Zhijiao Xiao, Jianmin Jiang, Yingying Zhu 0001, Zhong Ming 0001, Shenghua Zhong, Shubin Cai |
J. Syst. Softw. | 2 |
| 2015 | Tensor rank selection for multimedia analysis
Jianguang Zhang, Yahong Han, Jianmin Jiang |
J. Vis. Commun. Image Represent. | 3 |
| 2015 | Nonrigid Structure From Motion via Sparse RepresentationabstractThis paper proposes a new approach for nonrigid structure from motion with occlusion, based on sparse representation. We address the occlusion problem based on the latest developments on sparse representation: matrix completion, which can recover the observation matrix that has high percentages of missing data and can also reduce the noises and outliers in the known elements. We introduce sparse transform to the joint estimation of 3-D shapes and motions. 3-D shape trajectory space is fit by wavelet basis to achieve better modeling of complex motion. Experimental results on datasets without and with occlusion show that our method can better estimate the 3-D shapes and motions, compared with state-of-the-art algorithms. Kun Li 0001, Jing-Yu Yang 0002, Jianmin Jiang |
IEEE Trans. Cybern. | 3 |
| 2015 | Analyzing Event-Based Scheduling in Concurrent Reactive SystemsabstractThe traditional research on scheduling focuses on task scheduling and schedulability analysis in concurrent reactive systems. In this article, we dedicate ourselves to event-based scheduling. We first formally define an event-based scheduling policy and propose the notion of the correctness of a scheduling policy in terms of weak termination. Then we investigate the correctness of the decomposition of scheduling controls and finally obtain a decentralized scheduling method. The method can automatically decompose the scheduling policies of a concurrent reactive system into atomic scheduling policies. Every atomic scheduling policy corresponds to one subsystem. Each of the subsystems is a completely independent system, which may be developed and deployed independently. An experiment demonstrates these results that may help engineers to design correct and efficient schedule policies for a concurrent reactive system. Jianmin Jiang, Huibiao Zhu, Qin Li 0002, Ping Gong 0004, Zhong Hong |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2014 | Non-rigid structure from motion via sparse representationabstractThis paper proposes a new approach for non-rigid structure from motion with occlusion, based on sparse representation. We introduce sparse transform to the joint estimation of 3D shapes and motions. 3D shape trajectory space is fit by wavelet basis to achieve better modeling of complex motion. We address the occlusion problem based on the latest developments on sparse representation: matrix completion, which can recover the observation matrix that has high percentages of missing data and can also reduce the noises and outliers in the known elements. Experimental results on datasets without and with occlusion show that our method can better estimate the 3D shapes and motions, compared with state-of-the-art algorithms. Kun Li 0001, Jing-Yu Yang 0002, Jianmin Jiang |
ICME | 3 |
| 2014 | What Can We Learn about Motion Videos from Still Images?abstractHuman action recognition from motion videos plays an important role in multimedia analysis. Different from the temporal cues of action series in motion videos, the motion tendency can also be revealed from the still images or key frames. Thus, if the action knowledge in related still images can be well adapted to the target motion videos, we would have a great chance to improve the performance of video action recognition. In this paper, we propose a framework of Still-to-Motion Adaptation (SMA) for human action recognition. Common visual features are extracted both from the related images and target videos' key frames, by which the gap between still images and videos are bridged. Meanwhile, to utilize the unlabeled training videos in target domain, we incorporate a semi-supervised process into our framework. By minimizing the difference of action prediction from still features and motion features, we formulate the still-to-motion adaptation into a joint optimization process. Experiments successfully demonstrate the effectiveness of the proposed framework and show the better performance of action recognition compared with the state-of-the-art methods. We also analyze the impact on the recognition results of target videos by knowledge adaptation from still images. Jianguang Zhang, Yahong Han, Jinhui Tang 0001, Qinghua Hu, Jianmin Jiang |
ACM Multimedia | 5 |
| 2014 | OCSVM model combined with K-means recursive clustering for intrusion detection in SCADA systemsabstractIntrusion detection in Supervisory Control and Data Acquisition (SCADA) systems is of major importance nowadays. Most of the systems are designed without cyber security in mind, since interconnection with other systems through unsafe channels, is becoming the rule during last years. The de-isolation of SCADA systems make them vulnerable to attacks, disrupting its correct functioning and tampering with its normal operation. In this paper we present a intrusion detection module capable of detecting malicious network traffic in a SCADA (Supervisory Control and Data Acquisition) system, based on the combination of One-Class Support Vector Machine (OCSVM) with RBF kernel and recursive k-means clustering. The combination of OCSVM with recursive k-means clustering leads the proposed intrusion detection module to distinguish real alarms from possible attacks regardless of the values of parameters σ and ν, making it ideal for real-time intrusion detection mechanisms for SCADA systems. The OCSVM module developed is trained by network traces off line and detect anomalies in the system real time. The module is part of an IDS (Intrusion Detection System) system developed under CockpitCI project. Leandros Maglaras, Jianmin Jiang |
QSHINE | 2 |
| 2014 | Configuration of Services Based on VirtualizationabstractVirtualization is fundamental to cloud computing. It allows abstraction centred on services and isolation of lower level functionalities and underlying hardware. Modeling, analyzing and verifying cloud systems necessarily involve virtualization and services. However, there exist few efforts to effectively formalizing virtualization in cloud computing. In this paper, based on services we present an approach for defining virtualization. We discuss some properties of service virtualization under some operations and the correctness of virtual services (virtual services without abnormal behavioral problems). Moreover, we investigate automatic configuration of a service based on virtualization, that is, given a virtualized service, how can we automatically obtain all possible correct virtual services of such a service? The configuration process is to first separate a virtualized service into atomic and correct virtual services and then merge these atomic virtual services into all possible correct virtual services of such a virtualized service. The obtained theoretical results help to formally analyze, verify and configure cloud systems. Jianmin Jiang, Huibiao Zhu, Qin Li 0002, Ping Gong 0004, Zhong Hong |
TASE | 1 |
| 2014 | Gradient-based subspace phase correlation for fast and effective image alignment
Jinchang Ren, Theodore Vlachos, Jiangbin Zheng 0001, Jianmin Jiang |
J. Vis. Commun. Image Represent. | 5 |
| 2014 | HMM-based hybrid meta-clustering ensemble for temporal data
Yun Yang 0003, Jianmin Jiang |
Knowl. Based Syst. | 2 |
| 2014 | Recognition of Chinese artists via windowed and entropy balanced fusion in classification of their authored ink and wash paintings (IWPs)
Jiachuan Sheng, Jianmin Jiang |
Pattern Recognit. | 2 |
| 2014 | Video super-resolution based on automatic key-frame selection and feature-guided variational optical flow
Yanming Zhu 0001, Kun Li 0001, Jianmin Jiang |
Signal Process. Image Commun. | 3 |
| 2013 | A spectral-multiplicity-tolerant approach to robust graph matching
Wei Feng 0005, Chi-Man Pun, Jianmin Jiang |
Pattern Recognit. | 5 |
| 2013 | Heterogeneous Delay Embedding for Travel Time and Energy Cost Prediction Via Regression AnalysisabstractIn this paper, we study travel time and energy cost prediction at any future departure time for a targeted road segment and vehicle. These two prediction tasks play an important part in the design of advanced driver-assistance systems (ADAS) that can automatically manage battery charging, energy saving, and route planning for fully electric vehicles. Compared with the fundamental problem of travel time prediction, which usually learns from the historical and current data of travel time itself, energy cost prediction is a more complex problem that involves multiple context conditions and vehicle status measured by various time-invariant and time-variant data. We define a general learning problem based on multiple time-invariant and time-variant inputs to unify these two prediction tasks. To solve the defined learning problem, we propose heterogeneous delay embedding (HDE), which extracts an informative feature space for regression analysis and aims at achieving satisfactory prediction for any future departure time. The proposed HDE first categorizes the historical and current data of a time-variant measurement into different types, then incorporates different delay settings for embedding multiple types of time-series data, and finally removes redundant information and noise from the generated features using orthogonal locality preserving projection. Experimental results demonstrate the effectiveness of the proposed method for both short- and long-term predictions of travel time and energy cost. Tingting Mu, Jianmin Jiang, Yan Wang 0084 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2013 | Context-Aware and Energy-Driven Route Optimization for Fully Electric Vehicles via CrowdsourcingabstractRoute planning for fully electric vehicles (FEVs) must take energy efficiency into account due to limited battery capacity and time-consuming recharging. In addition, the planning algorithm should allow for negative energy costs in the road network due to regenerative braking, which is a unique feature of FEVs. In this paper, we propose a framework for energy-driven and context-aware route planning for FEVs. It has two novel aspects: 1) It is context aware, i.e., the framework has access to real-time traffic data for routing cost estimation; and it is energy driven, i.e., both time and energy efficiency are accounted for; which implies a biobjective nature of the optimization. In addition, in the case of insufficient energy on board, an optimal detour via recharge points is computed. Our main contributions to address these issues can be highlighted as follows: A vehicle-to-vehicle (V2V) communication protocol is proposed to realize the context awareness, and we replace the original biobjective form of optimality with two single-objective forms and propose a constrained A* ( CA*) algorithm to find the solutions. The algorithm maintains a Pareto front while it confines its search by energy constraints. The best recharging detour can be also found using the algorithm. We first compared the performance of the CA* algorithm with other algorithms. We then evaluate the impact of the context awareness on road traffic by simulations using a realistic road network regarding different forms of optimality. Finally, we show that the CA* algorithm can effectively produce optimal recharging detours. Yan Wang 0084, Jianmin Jiang, Tingting Mu |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2013 | Automated Induction of Heterogeneous Proximity Measures for Supervised Spectral EmbeddingabstractSpectral embedding methods have played a very important role in dimensionality reduction and feature generation in machine learning. Supervised spectral embedding methods additionally improve the classification of labeled data, using proximity information that considers both features and class labels. However, these calculate the proximity information by treating all intraclass similarities homogeneously for all classes, and similarly for all interclass samples. In this paper, we propose a very novel and generic method which can treat all the intra- and interclass sample similarities heterogeneously by potentially using a different proximity function for each class and each class pair. To handle the complexity of selecting these functions, we employ evolutionary programming as an automated powerful formula induction engine. In addition, for computational efficiency and expressive power, we use a compact matrix tree representation equipped with a broad set of functions that can build most currently used similarity functions as well as new ones. Model selection is data driven, because the entire model is symbolically instantiated using only problem training data, and no user-selected functions or parameters are required. We perform thorough comparative experimentations with multiple classification datasets and many existing state-of-the-art embedding methods, which show that the proposed algorithm is very competitive in terms of classification accuracy and generalization ability. Eduardo Rodríguez-Martínez, Tingting Mu, Jianmin Jiang, John Yannis Goulermas |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2012 | Executability Analysis for Semantically Annotated Process ModelabstractSemantically annotated process model (SPM) is a process model with semantic annotations, which include the precondition and effect, labeled for its activities based on the domain ontology. Such semantic annotations can increase the model's understanding and reusability, and facilitate its implement as well as compliance analysis. However, SPM analysis is challenging, since its correctness is beyond the soundness of process model and its formal semantics needs to consider domain state change. To assuring the correctness of SPM, the executablity analysis, i.e., whether there exists activity whose precondition is falsatisfied when it is active, is essential and has also been identified as a coNP-hard problem. To tame the hardness of the executability, we define a dynamic semantics for SPM based on a model which is proposed for defining domain state transition, present an encoding method by which the formal semantics is encoded into formulae as well as the executability, and propose a procedure which can bounded model check and diagnose the executability by using SAT solver. Our method has been implemented as a tool called SPMT and its validity is also illustrated through examples. Ping Gong 0004, Jianmin Jiang, Zhi Qin Chen |
APSCC | 2 |
| 2012 | Scalable image co-segmentation using color and covariance features
Wei Feng 0005, Jiawan Zhang, Jianmin Jiang |
ICPR | 5 |
| 2012 | Optimized image super-resolution based on sparse representation
Yanming Zhu 0001, Jianmin Jiang, Kun Li 0001 |
ICPR | 2 |
| 2012 | Lexical Story Co-Segmentation of Chinese Broadcast NewsabstractWe present an unsupervised technique, namely story co-segmentation, to automatically extract the common sto-ries on the same topic within a pair of Chinese broadcast news transcripts. Unlike classical topic tracking that usu-ally relies on previously trained topic models, our method is purely data-driven and is able to simultaneously deter-mine the common stories of the input texts. Specifical-ly, we propose an iterative four-step MRF solution to the problem of story co-segmentation using lexical cues only. We first construct a sentence-level graph formulation of the input news transcripts, and initialize foreground and background labeling by lexical clustering. We then up-date both foreground and background models based on the current labeling. We formalize story co-segmentation as a Gibbs energy minimization problem that balances the optimal objectives of foreground/background likeli-hood, intra-doc coherence, and inter-doc similarity. Fi-nally, the labeling refinement is obtained by hybrid op-timization with QPBO and BP. The effectiveness of our method has been validated on real-world CCTV corpus. Index Terms: story co-segmentation, foreground and background story modeling, lexical clustering, MRF, QP- Wei Feng 0005, Xuecheng Nie, Lei Xie 0001, Jianmin Jiang |
INTERSPEECH | 5 |
| 2012 | Modeling and analyzing mixed communications in service-oriented trustworthy software
Jianmin Jiang, Ping Gong 0004, Zhong Hong, HouGuang Yue |
Sci. China Inf. Sci. | 1 |
| 2012 | Effective venue image retrieval using robust feature extraction and model constrained matching for mobile robot localization
Yue Feng 0002, Jinchang Ren, Jianmin Jiang, Martin Halvey, Joemon M. Jose |
Mach. Vis. Appl. | 3 |
| 2012 | DBN-based structural learning and optimisation for automated handwritten character recognition
Olivier Pauplin, Jianmin Jiang |
Pattern Recognit. Lett. | 2 |
| 2012 | Sparse Unsupervised Dimensionality Reduction for Multiple View DataabstractDifferent kinds of high-dimensional visual features can be extracted from a single image. Images can thus be treated as multiple view data when taking each type of extracted high-dimensional visual feature as a particular understanding of images. In this paper, we propose a framework of sparse unsupervised dimensionality reduction for multiple view data. The goal of our framework is to find a low-dimensional optimal consensus representation from multiple heterogeneous features by multiview learning. In this framework, we first learn low-dimensional patterns individually from each view, considering the specific statistical property of each view. We construct a low-dimensional optimal consensus representation from those learned patterns, the goal of which is to leverage the complementary nature of the multiple views. We formulate the construction of the low-dimensional consensus representation to approximate the matrix of patterns by means of a low-dimensional consensus base matrix and a loading matrix. To select the most discriminative features for the spectral embedding of multiple views, we propose to add anl1-norm into the loading matrix's columns and impose orthogonal constraints on the base matrix. We develop a new alternating algorithm, i.e., spectral sparse multiview embedding, to efficiently obtain the solution. Each row of the loading matrix encodes structured information corresponding to multiple patterns. In order to gain flexibility in sharing information across subsets of the views, we impose a novel structured sparsity-inducing norm penalty on the loading matrix's rows. This penalty makes the loading coefficients adaptively load shared information across subsets of the learned patterns. We call this method structured sparse multiview dimensionality reduction. Experiments on a toy benchmark image data set and two real-world Web image data sets demonstrate the effectiveness of the proposed algorithms. Yahong Han, Fei Wu 0001, Dacheng Tao, Jian Shao 0001, Yueting Zhuang, Jianmin Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2012 | Adaptive Data Embedding Framework for Multiclass ClassificationabstractThe objective of this paper is the design of an engine for the automatic generation of supervised manifold embedding models. It proposes a modular and adaptive data embedding framework for classification, referred to as DEFC, which realizes in different stages including initial data preprocessing, relation feature generation and embedding computation. For the computation of embeddings, the concepts of friend closeness and enemy dispersion are introduced, to better control at local level the relative positions of the intraclass and interclass data samples. These are shown to be general cases of the global information setup utilized in the Fisher criterion, and are employed for the construction of different optimization templates to drive the DEFC model generation. For model identification, we use a simple but effective bilevel evolutionary optimization, which searches for the optimal model and its best model parameters. The effectiveness of DEFC is demonstrated with experiments using noisy synthetic datasets possessing nonlinear distributions and real-world datasets from different application fields. Tingting Mu, Jianmin Jiang, Yan Wang 0084, John Yannis Goulermas |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2012 | Three-Dimensional Motion Estimation via Matrix CompletionabstractThree-dimensional motion estimation from multiview video sequences is of vital importance to achieve high-quality dynamic scene reconstruction. In this paper, we propose a new 3-D motion estimation method based on matrix completion. Taking a reconstructed 3-D mesh as the underlying scene representation, this method automatically estimates motions of 3-D objects. A "separating + merging" framework is introduced to multiview 3-D motion estimation. In the separating step, initial motions are first estimated for each view with a neighboring view. Then, in the merging step, the motions obtained by each view are merged together and optimized by low-rank matrix completion method. The most accurate motion estimation for each vertex in the recovered matrix is further selected by three spatiotemporal criteria. Experimental results on data sets with synthetic motions and real motions show that our method can reliably estimate 3-D motions. Kun Li 0001, Qionghai Dai, Wenli Xu, Jing-Yu Yang 0002, Jianmin Jiang |
IEEE Trans. Syst. Man Cybern. Part B | 5 |
| 2012 | Quality-of-Service Analysis of Queuing Systems with Long-Range-Dependent Network Traffic and Variable Service CapacityabstractMany high-quality measurement studies have demonstrated that wireless network traffic exhibits the noticeable Long-Range-Dependent (LRD) property. Moreover, the fading nature of wireless channels can lead to variable service capacity. Due to the inherent difficulty and high complexity of modelling the fractal-like LRD traffic, existing analytical models of queuing systems with LRD arrival processes have been primarily limited to the simplified scenarios where the service capacity is assumed to be constant. Given the time-varying nature of wireless channels in the real-world working environments, it is very important and necessary to investigate system performance in the presence of variable service capacity. To this end, this paper presents a comprehensive analytical model for queuing systems subject to LRD traffic and variable service capacity. We extend the application of a Large Deviation Principle and derive the closed-form expressions of Quality-of-Service (QoS) metrics. The accuracy of the model validated through extensive simulation experiments makes it a cost-effective evaluation tool for performance analysis of communication networks. To illustrate its applications, the model is adopted to investigate the effects of LRD traffic and variable service capacity on the system performance and resource configuration. Xiaolong Jin 0001, Geyong Min, Ross Velentzas, Jianmin Jiang |
IEEE Trans. Wirel. Commun. | 4 |
| 2011 | Formal Analysis of OWL-S Process Model by FDRabstractNowadays, SOA is considered as a promising architecture for enterprise applications integration. OWL-S, as a Semantic Web Service technology, promises to facilitate various service tasks, such as service specification, discovery, composition, etc.. It is essential to analysis the OWL-Sprocess model for the quality issue of the composition of web services before its deployment. In this work, By CSPm and FDR tool, a formal analysis method for OWL-S process model, especially the interplay between the control part and dataflow part, is proposed, and its validity is illustrated by the extended version of the Bravo Air process model. Ping Gong 0004, Jianmin Jiang |
APSCC | 2 |
| 2011 | Message Dependency-Based Adaptation of ServicesabstractMismatch patterns capture the possible differences between two service (business) protocols to adapt. For these mismatches, formal definitions are presented in this paper. And a novel technique provides support for adapting two or more services. This technique requires that messages and message dependencies are used to directly model service (business) protocols and form a novel model, called a\emph{protocol structure}. Unlike most of existing approaches that only consider typical mismatches (e.g. deadlock), this model is easily used to detect multiple mismatches at a time and to automatically generate BPEL adapters. Jianmin Jiang, Ping Gong 0004, Zhong Hong |
APSCC | 1 |
| 2011 | Service Adaptation at Message LevelabstractIn SOA, adaptation techniques aim to automatically generate adapters. However, the generation of the adapter is a complicated task and requires extra knowledge to resolve all kinds of mismatches. We propose a novel model, called a protocol structure, which is used to model services and adapters and detect the mismatches among services. Once developers present interface mappings among services, adapters can be derived from the interface mappings. Jianmin Jiang, Ping Gong 0004, Zhong Hong |
SERVICES | 1 |
| 2011 | Effective recognition of MCCs in mammograms using an improved neural classifier
Jinchang Ren, Jianmin Jiang |
Eng. Appl. Artif. Intell. | 3 |
| 2011 | Performance of hidden Markov model and dynamic Bayesian network classifiers on handwritten Arabic word recognition
Jawad Hasan Yasin AlKhateeb, Olivier Pauplin, Jinchang Ren, Jianmin Jiang |
Knowl. Based Syst. | 4 |
| 2011 | Modelling of content-aware indicators for effective determination of shot boundaries in compressed MPEG videos
Jinchang Ren, Jianmin Jiang |
Multim. Tools Appl. | 3 |
| 2011 | Offline handwritten Arabic cursive text recognition using Hidden Markov Models and re-ranking
Jawad Hasan Yasin AlKhateeb, Jinchang Ren, Jianmin Jiang, Husni Al-Muhtaseb |
Pattern Recognit. Lett. | 3 |
| 2011 | K -NN Regression to Improve Statistical Feature Extraction for Texture RetrievalabstractThis correspondence presents an iterative method based upon k -nearest neighbors ( k-NN) regression to improve the performance of statistical feature extraction for texture image retrieval. The idea exploits the fact that an ideal feature extraction system would extract similar signatures from images characterized by the same texture and different signatures from dissimilar textures. Under the assumption that conventional statistical feature extraction contributes to sufficiently good retrieval performance, the signatures of k retrieved textures are used to update the signature of the query image using the k -NN regression algorithm. Extensive experiments show significant improvements with respect to retrieval performance in comparison to conventional statistical feature extraction. Fouad Khelifi, Jianmin Jiang |
IEEE Trans. Image Process. | 2 |
| 2010 | A Boosted Manifold Learning for Automatic Face RecognitionabstractManifold learning is an effective dimension reduction method to extract nonlinear structures from high dimensional data. Recently, manifold learning has received much attention within the research communities of image analysis, computer vision and document data analysis. In this paper, we propose a boosted manifold learning algorithm towards automatic 2D face recognition by using AdaBoost to select the best possible discriminating projection for manifold learning to exploit the strength of both techniques. Experimental results support that the proposed algorithm improves over existing benchmarks in terms of stability and recognition precision rates. Chunyuan Lu, Jianmin Jiang, Guo-Can Feng |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2010 | An EDBoost algorithm towards robust face recognition in JPEG compressed domain
Chunmei Qing, Jianmin Jiang |
Image Vis. Comput. | 2 |
| 2010 | A novel feature selection approach for biomedical data classification
Yonghong Peng, Zhi Qing Wu, Jianmin Jiang |
J. Biomed. Informatics | 3 |
| 2010 | Activity-driven content adaptation for effective video summarization
Jinchang Ren, Jianmin Jiang, Yue Feng 0002 |
J. Vis. Commun. Image Represent. | 2 |
| 2010 | Analysis of the Security of Perceptual Image Hashing Based on Non-Negative Matrix FactorizationabstractIn this letter, we analyze the security of a perceptual image hashing technique based on non-negative matrix factorization which was recently proposed and reported in the literature. We theoretically demonstrate that, although the technique uses different secret keys in subsequent stages, the first key plays an essential role to secure the hashing system. We next act as an attacker and propose a technique to estimate the secret key. Extensive experiments support our theoretical analysis and validate the proposed key estimation technique. Fouad Khelifi, Jianmin Jiang |
IEEE Signal Process. Lett. | 2 |
| 2010 | Normalized Co-Occurrence Mutual Information for Facial Pose Detection Inside VideosabstractHuman faces captured inside videos are often presented with variable poses, making it difficult to recognize and thus pose detection becomes crucial for such face recognition under non-controlled environment. While existing mutual in formation (MI) primarily considers the relationship between corresponding individual pixels, we propose a normalized co occurrence mutual information in this letter to capture the information embedded not only in corresponding pixel values but also in their geographical locations. In comparison with the existing Mis, the proposed presents an essential advantage that both marginal entropy and joint entropy can be optimally exploited in measuring the similarity between two given images. When developed into a facial pose detection algorithm inside video sequences, we show, through extensive experiments, that such design is capable of achieving the best performances among all the representative existing techniques compared. Chunmei Qing, Jianmin Jiang, Zhijing Yang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2010 | Perceptual Image Hashing Based on Virtual Watermark DetectionabstractThis paper proposes a new robust and secure perceptual image hashing technique based on virtual watermark detection. The idea is justified by the fact that the watermark detector responds similarly to perceptually close images using a non embedded watermark. The hash values are extracted in binary form with a perfect control over the probability distribution of the hash bits. Moreover, a key is used to generate pseudo-random noise whose real values contribute to the randomness of the feature vector with a significantly increased uncertainty of the adversary, measured by mutual information, in comparison with linear correlation. Experimentally, the proposed technique has been shown to outperform related state-of-the art techniques recently proposed in the literature in terms of robustness with respect to image processing manipulations and geometric attacks. Fouad Khelifi, Jianmin Jiang |
IEEE Trans. Image Process. | 2 |
| 2010 | High-Accuracy Sub-Pixel Motion Estimation From Noisy Images in Fourier DomainabstractIn this paper, we propose a new method for estimating sub-pixel motion via exploiting the principle of phase correlation in the Fourier domain. The method is based on linear weighting of the height of the main peak on the one hand and the difference between its two neighboring side-peaks on the other. Using both synthetic and real data we show that the proposed method outperforms many established approaches and achieves improved accuracy even in the presence of noisy samples. Jinchang Ren, Jianmin Jiang, Theodore Vlachos |
IEEE Trans. Image Process. | 2 |
| 2009 | Analytical Modelling of the GC-Based Handover Scheme with Heavy-Tailed Call Holding TimesabstractThe Guard Channel (GC) scheme is an extensively deployed handover prioritization policy in wireless mobile networks due to its appealing properties of easy implementation and flexible control. Self-similar traffic has been demonstrated to be a ubiquitous phenomenon in modern communication networks and many measurement studies have demonstrated that the call holding times in wireless mobile networks follow heavy-tailed lognormal distributions. However, existing analytical models for the handover scheme usually assume that both new and handover calls can be modelled by traditional Markovian arrival processes and call holding times can be characterized by exponential distributions. In order to obtain a deep understanding of the performance behaviour of the GC-based handover scheme under more realistic working conditions, this paper presents a new analytical model for this scheme in the presence of self-similar new call and handover call arrivals and heavy-tailed lognormal call holding times. We derive the blocking probability of new calls and the dropping probability of handover calls. Extensive simulation experiments are conducted to validate the accuracy of the developed model. Xiaolong Jin 0001, Geyong Min, Jianmin Jiang |
GLOBECOM | 3 |
| 2009 | A fuzzy logic method of feature representation for shot boundary detectionabstractUnlike most approaches reported in the literature, our proposed algorithm is characterized by using the fuzzy logic method for feature representation of shot detection. Firstly, novel features are extracted in the compressed domain. Secondly, a fuzzy logic method is used to implement feature representation. Thirdly, multiple Support Vector Machines (SVMs) are constructed for further verification using features generated from the fuzzification process in step two. Our algorithm combines the merits of the generalization ability of the SVM and the comprehensibility of the fuzzy logic. We have carried out extensive experiments using the data from TRECVID07. Our proposed algorithm achieved very good detection results compared with TRECVID07 algorithms and its run-speed is 4 times faster than real-time video play. Stanley S. Ipson, Jianmin Jiang |
ICIP | 3 |
| 2009 | Modelling of heterogeneous wireless networks under batch arrival traffic with communication localityabstractWireless mesh networks (WMNs) have been proposed to provide rapid deployment and easy reconfiguration of wireless broadband communications. WMNs can interoperate with WiMAX, Wi-Fi, sensor, or cellular networks in the hybrid working environments to relay packets robustly among these heterogeneous networks and significantly extend the coverage of individual wireless access networks. Many recent studies have shown that the packet arrival process in wireless networks exhibits the batch arrival nature and the communication locality has an important impact on the network capacity. With the aim of obtaining an effective performance evaluation tool of wireless networks, this paper proposes an analytical model for heterogeneous wireless networks integrated by WMNs in the presence of batch arrival traffic with communication locality. The validity of the analytical model is demonstrated through extensive comparison between analytical and simulation results. Yulei Wu, Geyong Min, Guojun Wang 0001, Jianmin Jiang |
WCNC | 4 |
| 2009 | Towards Computerized Digital Preservation based on Intelligent Agents and Web Services
Xiaolong Jin 0001, Jianmin Jiang, Geyong Min |
WEBIST | 2 |
| 2009 | Shot Boundary Detection in MPEG Videos Using Local and Global IndicatorsabstractShot boundary detection (SBD) plays important roles in many video applications. In this letter, we describe a novel method on SBD operating directly in the compressed domain. First, several local indicators are extracted from MPEG macroblocks, and AdaBoost is employed for feature selection and fusion. The selected features are then used in classifying candidate cuts into five sub-spaces via pre-filtering and rule-based decision making. Following that, global indicators of frame similarity between boundary frames of cut candidates are examined using phase correlation of dc images. Gradual transitions like fade, dissolve, and combined shot cuts are also identified. Experimental results on the test data from TRECVID'07 have demonstrated the effectiveness and robustness of our proposed methodology. Jinchang Ren, Jianmin Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2009 | Hierarchical Modeling and Adaptive Clustering for Real-Time Summarization of Rush VideosabstractIn this paper, we provide detailed descriptions of a proposed new algorithm for video summarization, which are also included in our submission to TRECVID'08 on BBC rush summarization. Firstly, rush videos are hierarchically modeled using the formal language technique. Secondly, shot detections are applied to introduce a new concept of V-unit for structuring videos in line with the hierarchical model, and thus junk frames within the model are effectively removed. Thirdly, adaptive clustering is employed to group shots into clusters to determine retakes for redundancy removal. Finally, each most representative shot selected from every cluster is ranked according to its length and sum of activity level for summarization. Competitive results have been achieved to prove the effectiveness and efficiency of our techniques, which are fully implemented in the compressed domain. Our work does not require high-level semantics such as object detection and speech/audio analysis which provides a more flexible and general solution for this topic. Jinchang Ren, Jianmin Jiang |
IEEE Trans. Multim. | 2 |
| 2008 | Knowledge-Supported Segmentation and Semantic Contents Extraction from MPEG Videos for Highlight-Based Annotation, Indexing and Retrieval
Jinchang Ren, Jianmin Jiang, Stanley S. Ipson |
ICIC (1) | 3 |
| 2008 | Skin Detection from Different Color Spaces for Model-Based Face Detection
Jinchang Ren, Jianmin Jiang, Stanley S. Ipson |
ICIC (3) | 3 |
| 2008 | Metric Learning: A general dimension reduction framework for classification and visualizationabstractA new general dimension reduction framework based on similar and dissimilar metric learning is proposed in this paper which allows us to exploit the geometry of data to reduce the data dimension for classification and visualization. The general formulation can unify the existing dimension reduction algorithms within a common framework. Furthermore, this metric learning framework can be used as a general platform for developing new dimension reduction algorithms. By utilizing this framework as a tool, we propose a novel supervised dimension reduction algorithm named sub-manifold preserving analysis (SMPA) in which the intrinsic sub-manifold structure will be preserved while the margin of interclass will be separated. Experimental evidences show that performance of our proposed SMPA algorithm is better than other algorithms. Chunyuan Lu, Guo-Can Feng, Jianmin Jiang, Patrick Shen-Pei Wang |
ICPR | 3 |
| 2008 | Effective features based on normal linear structures for detecting microcalcifications in mammogramsabstractMany features have been proposed for the detection of microcalcification clusters (MCCs) or classification of benign/malignant MCCs. However, most of them were designed based on the characteristics of MCC. In this paper, 16 features, which have been commonly adopted in many applications, are examined and six new features based on the linear structure are proposed. To evaluate the effectiveness of these six features, 800 suspicious regions detected from 320 full-field mammograms are equally divided into two parts for training and testing respectively. Experiments demonstrate that the area under the receiver operating characteristic (ROC) is increased from 0.86 to 0.89 after the new features are added into the set of feature selection. In the best feature sequence selected by the sequential floating forward search (SFFS) algorithm, the new proposed features take up the half number of features in the sequence. Zhi Qing Wu, Jianmin Jiang, Yonghong Peng |
ICPR | 2 |
| 2008 | A Manifolded AdaBoost for Face Recognition
Chunyuan Lu, Jianmin Jiang, Guo-Can Feng, Chunmei Qing |
KES (1) | 2 |
| 2008 | A Block-Edge-Pattern-Based Content Descriptor in DCT DomainabstractIn this correspondence, we describe a robust and effective content descriptor based on block-edge patterns extracted in discrete cosine transform domain, which is suitable for applications in JPEG or MPEG compressed images and videos. This content descriptor is constructed by a run-length edge-block histogram with three patterns including horizontal edge, vertical edge and no edge. In comparison with existing descriptors, the proposed features: 1) low-cost computing suitable for real-time implementation and high-speed processing of compressed videos; 2) robust to orientation changes such as rotation, noise, reverse, etc.; 3) operates in compressed domain. Extensive experiments support that the proposed content descriptor is effective in describing visual content, and achieves superior performances in terms of retrieval precision and recall rates. Jianmin Jiang, Kaijin Qiu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2007 | Statistical Classification of Skin Color Pixels from MPEG Videos
Jinchang Ren, Jianmin Jiang |
ACIVS | 2 |
| 2007 | Subspace Extension to Phase Correlation Approach for Fast Image RegistrationabstractA novel extension of phase correlation to subspace correlation is proposed, in which 2-D translation is decomposed into two 1-D motions thus only 1-D Fourier transform is used to estimate the corresponding motion. In each subspace, the first two highest peaks from 1-D correlation are linearly interpolated for subpixel accuracy. Experimental results have shown both the robustness and accuracy of our method. Jinchang Ren, Theodore Vlachos, Jianmin Jiang |
ICIP (1) | 3 |
| 2007 | Symmetry in Process AlgebraabstractAn original notion of symmetry for process algebra is defined, which is based on permutation groups. Given a process which is regarded as a structure and a permutation group on it, the quotient process (reduced process) is showed to be interleaving trace equivalent and interleaving bisimulation equivalent to the original process. Furthermore, an algorithm and two examples for this symmetric reduction are presented. Jianmin Jiang, Hongping Shu |
TASE | 1 |
| 2006 | Constrained Region-Growing and Edge Enhancement Towards Automated Semantic Video Object Segmentation
Jianmin Jiang, Shuyuan Yang 0002 |
ACIVS | 2 |
| 2006 | Knowledge-discovery incorporated evolutionary search for microcalcification detection in breast cancer diagnosis
Yonghong Peng, Jianmin Jiang |
Artif. Intell. Medicine | 3 |
| 2006 | Dominant colour extraction in DCT domain
Jianmin Jiang, Ying Weng, Pengjie Li |
Image Vis. Comput. | 1 |
| 2006 | Shape-based image retrieval for JPEG-2000 compressed image databases
Jianmin Jiang, B. F. Guo, Stanley S. Ipson |
Multim. Tools Appl. | 1 |
| 2006 | A fuzzy logic approach for detection of video shot boundaries
Hui Fang 0003, Jianmin Jiang |
Pattern Recognit. | 2 |
| 2006 | Adding lossless video compression to MPEGsabstractIn this correspondence, we propose to add a lossless compression functionality into existing MPEGs by developing a new context tree to drive arithmetic coding for lossless video compression. In comparison with the existing work on context tree design, the proposed algorithm features in 1) prefix sequence matching to locate the statistics model at the internal node nearest to the stopping point, where successful match of context sequence is broken; 2) traversing the context tree along a fixed order of context structure with a maximum number of four motion compensated errors; and 3) context thresholding to quantize the higher end of error values into a single statistics cluster. As a result,the proposed algorithm is able to achieve competitive processing speed, low computational complexity and high compression performances, which bridges the gap between universal statistics modeling and practical compression techniques. Extensive experiments show that the proposed algorithm outperforms JPEG-LS by up to 24% and CALIC by up to 22%, yet the processing time ranges from less than 2 seconds per frame to 6 seconds per frame on a typical PC computing platform. Jianmin Jiang |
IEEE Trans. Multim. | 1 |
| 2005 | Pseudo-stereo Conversion from 2D Video
Jianmin Jiang |
ACIVS | 2 |
| 2005 | The Preservation of Interleaving EquivalencesabstractIn recent years, though many attempts have been made to solve whether equivalences are preserved under action refinement, this problem still needs to be further investigated. The usual approach to the preservation problem is: given some well-established equivalence notion which is not preserved under refinement, is there a way of adding some restricted conditions under consideration such that preservation of this equivalence in the restricted setting is obtained? Generally, one investigates how to restrict the concept of action refinement such that these equivalences are preserved under the restricted refinement. In this paper, in another way, we investigate how to find the class of suitable systems satisfying that the established equivalences on them are preserved under no restricted refinement. Interleaving trace equivalence and interleaving bisimulation equivalence which are not preserved under refinement are showed that they are preserved under refinement in the systems in which there are not causal independence relations or all the transitions are bundle action transitions. Jianmin Jiang |
ICECCS | 1 |
| 2005 | A shape-match based algorithm for pseudo-3D conversion of 2D videosabstractThis paper presents a shape-match based method of converting conventional 2D videos to their pseudo-3D versions to achieve stereo effect via stereopsis. While conventional 2D videos do not contain true 3D information or difficult to extract such true 3D information, the proposed algorithm is to exploit the tolerance of the human visual perception to estimate pseudo-3D parameters for 3D conversion. The proposed algorithm has the features of: (i) the original 2D video frame is regarded as the reference frame for the pseudo-stereo image pair; (ii) appropriate disparity information is extracted by shape-based matching inside a small library of true stereo image pairs; and (iii) the extracted disparity is then used to construct a right video frame to complete the pseudo-stereo conversion. Our experiments show that a certain level of stereo effect has been achieved for our test video set and some samples of such experiments are illustrated for visual inspection. Jianmin Jiang, Stanley S. Ipson |
ICIP (3) | 2 |
| 2005 | Symmetry and AutobisimulationabstractFor event structures as a major branch of concurrent models, no one notices their symmetries. It is very likely that someone thinks that symmetry reduction over an event structure can be replaced by the (largest) autobisimulation reduction. We show that, given an event structure, there are some differences between the symmetric quotient model induced by symmetry reduction and the bisimilar quotient model induced by the (largest) autobisimulation. The former is ( interleaving ) bisimulation equivalent and pomset trace equivalent to the original event structure. However, though the latter is bisimulation equivalent to the original, it and its original event structure are not likely pomset trace equivalent. Jianmin Jiang |
PDCAT | 1 |
| 2004 | A CBR Driven Genetic Algorithm for Microcalcification Cluster Detection
Jianmin Jiang, Yonghong Peng |
EKAW | 2 |
| 2004 | Combination of SVM Knowledge for Microcalcification Detection in Digital Mammograms
Jianmin Jiang |
IDEAL | 2 |
| 2004 | Chinese Named Entity Recognition Based on Multilevel Linguistic Features
Jianmin Jiang, Tong Zhang 0001 |
IJCNLP | 2 |
| 2004 | Web-based image indexing and retrieval in JPEG compressed domain
Jianmin Jiang, Andrew James Armstrong, Guo-Can Feng |
Multim. Syst. | 1 |
| 2004 | Video extraction for fast content access to MPEG compressed videosabstractAs existing video processing technology is primarily developed in the pixel domain yet digital video is stored in compressed format, any application of those techniques to compressed videos would require decompression. For discrete cosine transform (DCT)-based MPEG compressed videos, the computing cost of standard row-by-row and column-by-column inverse DCT (IDCT) transforms for a block of 8/spl times/8 elements requires 4096 multiplications and 4032 additions, although practical implementation only requires 1024 multiplications and 896 additions. In this paper, we propose a new algorithm to extract videos directly from MPEG compressed domain (DCT domain) without full IDCT, which is described in three extraction schemes: 1) video extraction in 2/spl times/2 blocks with four coefficients; 2) video extraction in 4/spl times/4 blocks with four DCT coefficients; and 3) video extraction in 4/spl times/4 blocks with nine DCT coefficients. The computing cost incurred only requires 8 additions and no multiplication for the first scheme, 2 multiplication and 28 additions for the second scheme, and 47 additions (no multiplication) for the third scheme. Extensive experiments were carried out, and the results reveal that: 1) the extracted video maintains competitive quality in terms of visual perception and inspection and 2) the extracted videos preserve the content well in comparison with those fully decompressed ones in terms of histogram measurement. As a result, the proposed algorithm will provide useful tools in bridging the gap between pixel domain and compressed domain to facilitate content analysis with low latency and high efficiency such as those applications in surveillance videos, interactive multimedia, and image processing. Jianmin Jiang, Ying Weng |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2003 | Video Extraction in Compressed DomainabstractIn this paper, we propose a video extraction algorithm directly in the compressed domain for low cost and fast content access to those compressed video data via MPEG. Extensive experiments show that such extracted images and videos not only maintain well-preserved content features, but also illustrate reasonable quality in terms of both PSNR values and visual inspection. In cases where video processing tasks do not necessarily require full resolution pixel data such as browsing, pattern recognition, and object tracking in surveillance applications, the proposed algorithm will provide superior performance in terms of computing efficiency, under the context that millions of video frames need to be accessed, yet they are stored in compressed format. Puteri Norhashimah, Hui Fang 0003, Jianmin Jiang |
AVSS | 3 |
| 2003 | JPEG compressed image retrieval via statistical features
Guo-Can Feng, Jianmin Jiang |
Pattern Recognit. | 2 |
| 2003 | On-line improvements of the rate-distortion performance in MPEG-2 rate controlabstractExisting rate-control research can be summarized into two approaches. One attempts to optimize the rate-distortion (R-D) characteristics of the input source before any decisions about the quantizer step-size assignment are made. As such, rate-control schemes out of this approach incur iterative procedures, high computational cost and encoding delays. The other approach is highlighted by direct buffer-state feedback techniques in predicting the R-D characteristics based on either probabilistic models or on historic information. Both approaches trade suboptimality in the prediction of R-D characteristics for reduced computational cost and delay. We combine the simplicity of the predictive models with the high-complexity R-D approaches to design effective rate-control algorithms to exploit the advantages of both approaches: low encoding delay and low complexity, yet optimized estimation of the R-D statistics. Extensive experiments show that improvements over the MPEG-2 scheme with gains of 0.5-1 dB per frame are achieved by the proposed schemes for a variety of test sequences when the same target bit rates are maintained. Christos Grecos, Jianmin Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2002 | A modified shape descriptor in wavelets compressed domainabstractIn order to implement shape-based image retrieval in the wavelet domain effectively and efficiently, a new method is introduced to improve the moment-based shape descriptor. Firstly, several morphological operations are used to incorporate isolated significant points of a wavelet coefficient domain into a meaningful object region, which leads to a more accurate shape description. Then, by using the invariant properties of certain moments, a compact moment representation method is also proposed to integrate shape information spread at various directional wavelet subbands. Experimental results show that the proposed algorithm outperforms the existing shape-based image retrieval method in the wavelet domain. Baofeng Guo, Jianmin Jiang |
ICIP (1) | 2 |
| 2002 | Direct content access and extraction from JPEG compressed images
Jianmin Jiang, Andrew James Armstrong, Guo-Can Feng |
Pattern Recognit. | 1 |
| 2002 | A hybrid scheme for low bit-rate coding of stereo imagesabstractIn this paper, we propose a hybrid scheme to implement an object driven, block based algorithm to achieve low bit-rate compression of stereo image pairs. The algorithm effectively combines the simplicity and adaptability of the existing block based stereo image compression techniques with an edge/contour based object extraction technique to determine appropriate compression strategy for various areas of the right image. Unlike the existing object-based coding such as MPEG-4 developed in the video compression community, the proposed scheme does not require any additional shape coding. Instead, the arbitrary shape is reconstructed by the matching object inside the left frame, which has been encoded by standard JPEG algorithm and hence made available at the decoding end for those shapes in right frames. Yet the shape reconstruction for right objects incurs no distortion due to the unique correlation between left and right frames inside stereo image pairs and the nature of the proposed hybrid scheme. Extensive experiments carried out support that significant improvements of up to 20% in compression ratios are achieved by the proposed algorithm in comparison with the existing block-based technique, while the reconstructed image quality is maintained at a competitive level in terms of both PSNR values and visual inspections. Jianmin Jiang, Eran A. Edirisinghe |
IEEE Trans. Image Process. | 1 |
| 2001 | Image spatial transformation in DCT domainabstractIn this paper, we report work on generalizing spatial relationships between the DCTs of any block and its sub-blocks, which paves the way for image processing in the JPEG compressed domain. The results reveal that DCT coefficients of any block can be directly obtained from the DCT coefficients of its sub-blocks and the inter-block relationship remains to be linear. Due to the fact that the corresponding coefficient matrix of linear combination is sparse, the computational complexity of the proposed algorithms is significantly lower than that of the existing methods. Guo-Can Feng, Jianmin Jiang |
ICIP (3) | 2 |
| 2001 | Texture-based image retrieval in wavelets compressed domainabstractWe describe an algorithm designed to retrieve and index images directly in wavelets compressed domain. Starting from compressed codes produced by wavelets-based techniques (SPHIT), texture keys are constructed by analysis of a significance map, hence, image retrieval can be conducted based on their content. The primary targets of this work can be highlighted as: (a) to be able to extract image content information in native compressed domain, i.e. without decompressing the stored images to any extent, (b) to compensate for the content information loss introduced by the lossy compression scheme by localizing the information provided by the significance map. Extensive experiments are carried out on a database of over 1000 images. Our contribution to this area can be outlined as: (a) lower requirement for storage space as no content information need to be stored separately and (b) fast response times for queries in the database. Georgios Voulgaris, Jianmin Jiang |
ICIP (2) | 2 |
| 2001 | A low cost design of rate controlled JPEG-LS near lossless image compression
Jianmin Jiang, Christos Grecos |
Image Vis. Comput. | 1 |
| 2000 | A low-cost content-adaptive and rate-controllable near-lossless image codec in DPCM domainabstractIt remains an important issue for any image codec to optimize the balance between the reconstructed image quality and the rate-controllable compression ratio. A content adaptive information loss distribution scheme is proposed to design a low-cost and rate-controllable image codec. This design exploits the principle that any extra information loss required by the rate-control should be introduced only to those areas where the local texture analysis reveals that human visual perception is less sensitive to the incurred distortion. Therefore, while rate-distortion theory may enable us to optimize PSNR measurement of quality in rate control design, the proposed technique is able to optimize the perceptual quality of the reconstructed images. Experimental results are reported to support this statement for the performance of the proposed algorithm, benchmarked by the nonrate-controlled JPEG-LS. Jianmin Jiang |
IEEE Trans. Image Process. | 1 |
| 1998 | Comparative investigation of a non-linear predictive codec versus JPEG lossless compressionabstractA non-linear predictive coding based algorithm is proposed for lossless image compression. The algorithm uses two neighbouring pixels, one left and the other top, as a pioneering block to search for the best matched blocks inside a pre-defined window. The corresponding pixels associated with the best matched blocks are then taken to produce the predictive value, together with the two pioneering pixels. Comparative investigation is carried out by experiments which show clearly that the proposed algorithm outperform JPEG lossless compression mode. Jianmin Jiang, Meiying Lo |
ICASSP | 1 |
| 1996 | Distortion equalized fuzzy competitive learning for image data vector quantizationabstractVector quantization is a popular approach to image compression as it allows images to be coded at less than one bit per pixel. This paper presents a modified fuzzy competitive learning algorithm and applies it to image data vector quantization. The proposed algorithm overcomes the neuron underutilization problem by applying both fuzzy learning and distortion equalization to the competitive learning algorithm. Experimental results on real image data shows that this approach produces a higher quality codebook than applying fuzzy learning or distortion equalization to the competitive learning algorithm individually. Darren Butler, Jianmin Jiang |
ICASSP | 2 |
| 1996 | A novel parallel design of a codec for black and white image compression
Jianmin Jiang |
Signal Process. Image Commun. | 1 |
| 1994 | Parallel Design of Q-Coders for Bilevel Image CompressionabstractA parallel algorithm is presented in this paper to implement the adaptive binary arithmetic coding for lossless bilevel image compression. Based on the sequential Q-coder, software analysis in C is carried out to establish a tree array to process 4 bits in parallel. This development of parallel Q-coder substantially improves the encoding speed of bilevel images. As a matter of fact, the parallel algorithm can also be extended theoretically to any number of bits to be processed in parallel. The implication involved will be the design of internal structure for each PE, especially the buffer size where each bit renormalized locally is to be updated by the PE at the next level before it is sent out at the top of the tree array. Jianmin Jiang |
ICPADS | 1 |