EDBT 2026 Demo / reviewers in the wild / expert
Jiachao Zhang
dblp:166/6754
· DBLP profile ↗
27ranked-venue papers
7as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 5 first-author · 9 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Federated Unsupervised Skeletal Action Recognition From Condensation to Expansion
Jinjin Gong, Binqian Xu, Jiachao Zhang, Xiangbo Shu |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2026 | Attack-Augmented Mixing-Contrastive Skeletal Representation LearningabstractContrastive learning facilitates the acquisition of informative skeleton representations for unsupervised action recognition by leveraging effective positive and negative sample pairs. However, most existing methods construct these pairs through weak or strong data augmentations, which typically rely on random appearance alterations of skeletons. While such augmentations are somewhat effective, they introduce semantic variations only indirectly and face two inherent limitations. First, simply modifying the appearance of skeletons often fails to reflect meaningful semantic variations. Second, random perturbations can unintentionally blur the boundary between positive and negative pairs, weakening the contrastive objective. To address these challenges, we propose an attack-driven augmentation framework that explicitly introduces semantic-level perturbations. This approach facilitates the generation of hard positives while guiding the model to mine more informative hard negatives. Building on this idea, we present Attack-Augmented Mixing-Contrastive Skeletal Representation Learning (A2MC), a novel framework that focuses on contrasting hard positive and hard negative samples for more robust representation learning. Within A2MC, we design an Attack-Augmentation (Att-Aug) module that integrates both targeted (attack-based) and untargeted (augmentation-based) perturbations to generate informative hard positive samples. In parallel, we propose the Positive-Negative Mixer (PNM), which blends hard positive and negative features to synthesize challenging hard negatives. These are then used to update a mixed memory bank for more effective contrastive learning. Comprehensive evaluations across three public benchmarks demonstrate that our approach, termed A2MC, achieves performance on par with or exceeding existing state-of-the-art methods. Binqian Xu, Xiangbo Shu, Jiachao Zhang, Rui Yan 0010, Guosen Xie |
IEEE Trans. Image Process. | 3 |
| 2025 | Tensor-Aggregated LoRA in Federated Fine-Tuning
Binqian Xu, Xiangbo Shu, Jiachao Zhang, Yazhou Yao, Guosen Xie, Jinhui Tang 0001 |
ICCV | 4 |
| 2025 | STPM: Spatial-Temporal Token Pruning and Merging for Complex Activity RecognitionabstractLightweight video representation techniques have advanced significantly for simple activity recognition, but they still encounter several issues when applied to complex activity recognition: 1) The presence of numerous individuals and varying spatial positions makes it difficult for traditional token pruning methods to maintain accuracy. 2) Simply discarding entire frames may result in the loss of crucial clues. 3) To maintain parallel computing, applying the same pruning rate to every frame leads to significant redundancy in frames with low information content. To this end, we propose a lightweight and novel Spatial-Temporal Token Pruning and Merging (STPM) framework, specifically designed for complex action videos where human actors occupy a small spatial resolution within video frames. Our framework considers two critical factors: semantic importance and spatial-temporal redundancy, to further reduce overhead. For semantic importance, STPM captures class-specific attention scores by learning multiple class tokens within the transformer to guide token pruning. For spatial-temporal redundancy, STPM employs an anchor graph and temporal attention to perform spatial and temporal token merging, preserving appearance and temporal cues while eliminating semantic duplication and redundancy. We conduct extensive experiments on JRDB-PAR primarily using recently introduced video transformer backbones, e.g., MViT and ViT. Our framework achieves similar results while requiring 40% less computation. Yumeng Su, Jiachao Zhang, Rui Yan 0010, Pengpeng Li 0001, Guosen Xie, Xiangbo Shu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | GPT4Ego: Unleashing the Potential of Pre-Trained Models for Zero-Shot Egocentric Action RecognitionabstractVision-Language Models (VLMs), pre-trained on large-scale datasets, have shown impressive performance in various visual recognition tasks. This advancement paves the way for notable performance in some egocentric tasks, Zero-Shot Egocentric Action Recognition (ZS-EAR), entailing VLMs zero-shot to recognize actions from first-person videos enriched in more realistic human-environment interactions. Typically, VLMs handle ZS-EAR as a global video-text matching task, which often leads to suboptimal alignment of vision and linguistic knowledge. We propose a refined approach for ZS-EAR using VLMs, emphasizing fine-grained concept-description alignment that capitalizes on the rich semantic and contextual details in egocentric videos. In this work, we introduce a straightforward yet remarkably potent VLM framework,akaGPT4Ego, designed to enhance the fine-grained alignment of concept and description between vision and language. Specifically, we first propose a new Ego-oriented Text Prompting (EgoTP$\spadesuit$) scheme, which effectively prompts action-related text-contextual semantics by evolving word-level class names to sentence-level contextual descriptions by ChatGPT with well-designed chain-of-thought textual prompts. Moreover, we design a new Ego-oriented Visual Parsing (EgoVP$\clubsuit$) strategy that learns action-related vision-contextual semantics by refining global-level images to part-level contextual concepts with the help of SAM. Extensive experiments demonstrate GPT4Ego significantly outperforms existing VLMs on three large-scale egocentric video benchmarks, i.e., EPIC-KITCHENS-100 (33.2%$\uparrow$$_{\bm {+9.4}}$), EGTEA (39.6%$\uparrow$$_{\bm {+5.5}}$), and CharadesEgo (31.5%$\uparrow$$_{\bm {+2.6}}$). In addition, benefiting from the novel mechanism of fine-grained concept and description alignment, GPT4Ego can sustainably evolve with the advancement of ever-growing pre-trained foundational models. We hope this work can encourage the egocentric community to build more investigation into pre-trained vision-language models. Guangzhao Dai, Xiangbo Shu, Rui Yan 0010, Jiachao Zhang |
IEEE Trans. Multim. | 5 |
| 2025 | Leveraging Frame- and Feature-level Progressive Augmentation for Semi-supervised Action RecognitionabstractSemi-supervised action recognition is a challenging yet prospective task due to its low reliance on costly labeled videos. One high-profile solution is to explore frame-level weak/strong augmentations for learning abundant representations, inspired by the FixMatch framework dominating the semi-supervised image classification task. However, such a solution mainly brings perturbations in terms of texture and scale, leading to the limitation in learning action representations in videos with spatiotemporal redundancy and complexity. Therefore, we revisit the creative trick of weak/strong augmentations in FixMatch and then propose, to the best of our knowledge, a novel Frame- and Feature-level augmentation FixMatch (dubbed as F 2 -FixMatch) framework to learn more abundant action representations for being robust to complex and dynamic video scenarios. Specifically, we design a new Progressive Augmentation mechanism that implements the weak/strong augmentations first at the frame level, and further implements the perturbation at the feature level, to obtain abundant four types of augmented features in broader perturbation spaces. Moreover, we present an evolved Multihead Pseudo-Labeling scheme to promote the consistency of features across different augmented versions based on the pseudo labels. We conduct extensive experiments on several public datasets to demonstrate that our F 2 -FixMatch achieves the performance gain compared with current state-of-the-art methods. The source codes of F 2 -FixMatch are publicly available at https://github.com/zwtu/F2FixMatch . Zhewei Tu, Xiangbo Shu, Rui Yan 0010, Zhenxing Liu 0001, Jiachao Zhang |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2024 | Structured residual sparsity for video compressive sensing reconstruction
Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Jiachao Zhang, Jiantao Zhou 0001, Ce Zhu |
Signal Process. | 4 |
| 2024 | Multiple Complementary Priors for Multispectral Image Compressive Sensing ReconstructionabstractCompressive sensing (CS) techniques using a few compressed measurements have drawn considerable interest in reconstructing multispectral imagery (MSI). Nonlocal-based tensor methods have been widely used for MSI-CS reconstruction, which employ the nonlocal self-similarity (NSS) property of MSI to obtain satisfactory results. However, such methods only consider the internal priors of MSI while ignoring important external image information, for example deep-driven priors learned from a corpus of natural image datasets. Meanwhile, they usually suffer from annoying ringing artifacts due to the aggregation of overlapping patches. In this article, we propose a novel approach for highly effective MSI-CS reconstruction using multiple complementary priors (MCPs). The proposed MCP jointly exploits nonlocal low-rank and deep image priors under a hybrid plug-and-play framework, which contains multiple pairs of complementary priors, namely, internal and external, shallow and deep, and NSS and local spatial priors. To make the optimization tractable, a well-known alternating direction method of multiplier (ADMM) algorithm based on the alternating minimization framework is developed to solve the proposed MCP-based MSI-CS reconstruction problem. Extensive experimental results demonstrate that the proposed MCP algorithm outperforms many state-of-the-art CS techniques in MSI reconstruction. The source code of the proposed MCP-based MSI-CS reconstruction algorithm is available at: https://github.com/zhazhiyuan/MCP_MSI_CS_Demo.git. Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Jiachao Zhang, Jiantao Zhou 0001, Xudong Jiang 0001, Ce Zhu |
IEEE Trans. Cybern. | 4 |
| 2024 | Spatiotemporal Decouple-and-Squeeze Contrastive Learning for Semisupervised Skeleton-Based Action RecognitionabstractContrastive learning has been successfully leveraged to learn action representations for addressing the problem of semisupervised skeleton-based action recognition. However, most contrastive learning-based methods only contrast global features mixing spatiotemporal information, which confuses the spatial- and temporal-specific information reflecting different semantic at the frame level and joint level. Thus, we propose a novel spatiotemporal decouple-and-squeeze contrastive learning (SDS-CL) framework to comprehensively learn more abundant representations of skeleton-based actions by jointly contrasting spatial-squeezing features, temporal-squeezing features, and global features. In SDS-CL, we design a new spatiotemporal-decoupling intra-inter attention (SIIA) mechanism to obtain the spatiotemporal-decoupling attentive features for capturing spatiotemporal specific information by calculating spatial- and temporal-decoupling intra-attention maps among joint/motion features, as well as spatial- and temporal-decoupling inter-attention maps between joint and motion features. Moreover, we present a new spatial-squeezing temporal-contrasting loss (STL), a new temporal-squeezing spatial-contrasting loss (TSL), and the global-contrasting loss (GL) to contrast the spatial-squeezing joint and motion features at the frame level, temporal-squeezing joint and motion features at the joint level, as well as global joint and motion features at the skeleton level. Extensive experimental results on four public datasets show that the proposed SDS-CL achieves performance gains compared with other competitive methods. Binqian Xu, Xiangbo Shu, Jiachao Zhang, Guangzhao Dai, Yan Song 0005 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | MUP: Multi-granularity Unified Perception for Panoramic Activity RecognitionabstractPanoramic activity recognition is required to jointly identify multi-granularity human behaviors including individual actions, group activities, and global activities in multi-person videos. Previous methods encode these behaviors hierarchically through multiple stages, which disturb the inherent co-occurrence across multi-granularity behaviors in the same scene. To this end, we propose a novel Multi-granularity Unified Perception (MUP) framework that perceives different granularity behaviors universally to explore the co-occurrence motion pattern via the same parameters in an end-to-end fashion. To be specific, the proposed framework stacks three Unified Motion Encoding (UME) blocks for modeling multiple granularity behaviors with shared parameters. UME block mines intra-relevant and cross-relevant semantics synchronously from input feature sequences via Intra-granularity Motion Embedding (IME) and Cross-granularity Motion Prototyping (CMP). In particular, IME aims to model the interactions among visual features within each granularity based on the attention mechanism. CMP aims to aggregate features across different granularities (i.e., person to group) via several learnable prototypes. Extensive experiments demonstrate that MUP outperforms the state-of-the-art methods on JRDB-PAR and has satisfactory interpretability. Meiqi Cao, Rui Yan 0010, Xiangbo Shu, Jiachao Zhang, Jinpeng Wang 0001, Guosen Xie |
ACM Multimedia | 4 |
| 2023 | Nonlocal Structured Sparsity Regularization Modeling for Hyperspectral Image DenoisingabstractThe non-local-based model for hyperspectral image (HSI) denoising first uses non-local self-similarity (NSS) prior to group similar full-band patches into three-dimensional non-local full-band groups (tensors) using a block matching (BM) operation, and then a low-rank (LR) penalty is typically applied to each non-local full-band group to reduce noise. While non-local-based methods have shown promising performance in HSI denoising, most existing methods have only considered the LR property of the non-local full-band group while ignoring the strong correlation between sparse coefficients. Moreover, such methods often result in unsatisfactory visual artifacts due to the noise sensitivity of BM operations, while requiring expensive computations. To address these limitations, this paper proposes a novel non-local structured sparsity regularization (NLSSR) approach for HSI denoising. First, to mitigate the noise sensitivity of the BM operation, we propose a graph-based domain distance scheme to index similar full-band patches to form the non-local full-band group. Second, we design an adaptive unidirectional low-rank (LR) dictionary with low complexity that takes into account the differences in intrinsic structure correlation among different modes of the non-local full-band tensor. Third, we utilize a global spectral LR prior to reduce spectral redundancy. Fourth, we develop a generalized soft-thresholding (GST) algorithm based on the alternating minimization framework to solve the NLSSR-based HSI denoising problem. We perform extensive experiments on both simulated and real data to show that the proposed NLSSR algorithm outperforms many popular or state-of-the-art HSI denoising methods in both quantitative and visual evaluations. Zhiyuan Zha, Bihan Wen, Xin Yuan 0002, Jiachao Zhang, Jiantao Zhou 0001, Yilong Lu, Ce Zhu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | A simple yet effective image stitching with computational suture zone
Jiachao Zhang, Yunbin Huang, Yanming Yu, Xiangbo Shu |
Vis. Comput. | 1 |
| 2022 | Skip-attention encoder-decoder framework for human motion prediction
Xiangbo Shu, Rui Yan 0010, Jiachao Zhang, Yan Song 0005 |
Multim. Syst. | 4 |
| 2022 | Nonconvex Structural Sparsity Residual Constraint for Image RestorationabstractThis article proposes a novel nonconvex structural sparsity residual constraint (NSSRC) model for image restoration, which integrates structural sparse representation (SSR) with nonconvex sparsity residual constraint (NC-SRC). Although SSR itself is powerful for image restoration by combining the local sparsity and nonlocal self-similarity in natural images, in this work, we explicitly incorporate the novel NC-SRC prior into SSR. Our proposed approach provides more effective sparse modeling for natural images by applying a more flexible sparse representation scheme, leading to high-quality restored images. Moreover, an alternating minimizing framework is developed to solve the proposed NSSRC-based image restoration problems. Extensive experimental results on image denoising and image deblocking validate that the proposed NSSRC achieves better results than many popular or state-of-the-art methods over several publicly available datasets. Zhiyuan Zha, Xin Yuan 0002, Bihan Wen, Jiachao Zhang, Ce Zhu |
IEEE Trans. Cybern. | 4 |
| 2021 | FLDDoS: DDoS Attack Detection Model based on Federated LearningabstractRecently, DDoS attack has developed rapidly and become one of the most important threats to the Internet. Traditional machine learning and deep learning methods can-not train a satisfactory model based on the data of a single client. Moreover, in the real scenes, there are a large number of devices used for traffic collection, these devices often do not want to share data between each other depending on the research and analysis value of the attack traffic, which limits the accuracy of the model. Therefore, to solve these problems, we design a DDoS attack detection model based on federated learning named FLDDoS, so that the local model can learn the data of each client without sharing the data. In addition, considering that the distribution of attack detection datasets is extremely imbalanced and the proportion of attack samples is very small, we propose a hierarchical aggregation algorithm based on K-Means and a data resampling method based on SMOTEENN. The result shows that our model improves the accuracy by 4% compared with the traditional method, and reduces the number of communication rounds by 40%. Jiachao Zhang, Peiran Yu, Jianzhong Zhang 0003 |
TrustCom | 1 |
| 2020 | Web-Supervised Network for Fine-Grained Visual ClassificationabstractFine-grained visual classification (FGVC) is a tough task due to its high annotation cost of the fine-grained subcategories. To build a large-scale dataset at low manual cost, straightforwardly learning from web images for FGVC has attracted broad attention. However, there exist two characteristics in the need of concerning for the web dataset: 1) Noisy images; 2) A large proportion of hard examples. In this paper, we propose a simple yet effective approach to deal with noisy images and hard examples during training. Our method is a pure web-supervised method for FGVC. Extensive experiments on three commonly used fine-grained datasets demonstrate that our approach is much superior to the state-of-the-art web-supervised methods. The data and source code of this work have been posted available at: https://github.com/NUST-Machine-Intelligence-Laboratory/WSNFG. Chuanyi Zhang, Yazhou Yao, Jiachao Zhang, Jian Zhang 0002, Zhenmin Tang |
ICME | 3 |
| 2020 | CAN-GAN: Conditioned-attention normalized GAN for face age synthesis
Chenglong Shi, Jiachao Zhang, Yazhou Yao, Yunlian Sun, Huaming Rao, Xiangbo Shu |
Pattern Recognit. Lett. | 2 |
| 2020 | From Rank Estimation to Rank Approximation: Rank Residual Constraint for Image RestorationabstractIn this paper, we propose a novel approach for the rank minimization problem, termed rank residual constraint (RRC). Different from existing low-rank based approaches, such as the well-known nuclear norm minimization (NNM) and the weighted nuclear norm minimization (WNNM), which estimate the underlying low-rank matrix directly from the corrupted observation, we progressively approximate (approach) the underlying low-rank matrix via minimizing the rank residual. Through integrating the image nonlocal self-similarity (NSS) prior with the proposed RRC model, we apply it to image restoration tasks, including image denoising and image compression artifacts reduction. Toward this end, we first obtain a good reference of the original image groups by using the image NSS prior, and then the rank residual of the image groups between this reference and the degraded image is minimized to achieve a better estimate to the desired image. In this manner, both the reference and the estimated image in each iteration are improved gradually and jointly. Based on the group-based sparse representation model, we further provide a theoretical analysis on the feasibility of the proposed RRC model. Experimental results demonstrate that the proposed RRC model outperforms many state-of-the-art schemes in both the objective and perceptual qualities. Zhiyuan Zha, Xin Yuan 0002, Bihan Wen, Jiantao Zhou 0001, Jiachao Zhang, Ce Zhu |
IEEE Trans. Image Process. | 5 |
| 2020 | A Benchmark for Sparse Coding: When Group Sparsity Meets Rank MinimizationabstractSparse coding has achieved a great success in various image processing tasks. However, a benchmark to measure the sparsity of image patch/group is missing since sparse coding is essentially an NP-hard problem. This work attempts to fill the gap from the perspective of rank minimization. We firstly design an adaptive dictionary to bridge the gap between group-based sparse coding (GSC) and rank minimization. Then, we show that under the designed dictionary, GSC and the rank minimization problems are equivalent, and therefore the sparse coefficients of each patch group can be measured by estimating the singular values of each patch group. We thus earn a benchmark to measure the sparsity of each patch group because the singular values of the original image patch groups can be easily computed by the singular value decomposition (SVD). This benchmark can be used to evaluate performance of any kind of norm minimization methods in sparse coding through analyzing their corresponding rank minimization counterparts. Towards this end, we exploit four well-known rank minimization methods to study the sparsity of each patch group and the weighted Schatten p-norm minimization (WSNM) is found to be the closest one to the real singular values of each patch group. Inspired by the aforementioned equivalence regime of rank minimization and GSC, WSNM can be translated into a non-convex weighted ℓp-norm minimization problem in GSC. By using the earned benchmark in sparse coding, the weighted ℓp-norm minimization is expected to obtain better performance than the three other norm minimization methods, i.e., ℓ1-norm, ℓp-norm and weighted ℓ1-norm. To verify the feasibility of the proposed benchmark, we compare the weighted ℓp-norm minimization against the three aforementioned norm minimization methods in sparse coding. Experimental results on image restoration applications, namely image inpainting and image compressive sensing recovery, demonstrate that the proposed scheme is feasible and outperforms many state-of-the-art methods. Zhiyuan Zha, Xin Yuan 0002, Bihan Wen, Jiantao Zhou 0001, Jiachao Zhang, Ce Zhu |
IEEE Trans. Image Process. | 5 |
| 2020 | Image Restoration Using Joint Patch-Group-Based Sparse RepresentationabstractSparse representation has achieved great success in various image processing and computer vision tasks. For image processing, typical patch-based sparse representation (PSR) models usually tend to generate undesirable visual artifacts, while group-based sparse representation (GSR) models lean to produce over-smooth effects. In this paper, we propose a new sparse representation model, termed joint patch-group based sparse representation (JPG-SR). Compared with existing sparse representation models, the proposed JPG-SR provides an effective mechanism to integrate the local sparsity and nonlocal self-similarity of images. We then apply the proposed JPG-SR to image restoration tasks, including image inpainting and image deblocking. An iterative algorithm based on the alternating direction method of multipliers (ADMM) framework is developed to solve the proposed JPG-SR based image restoration problems. Experimental results demonstrate that the proposed JPG-SR is effective and outperforms many state-of-the-art methods in both objective and perceptual quality. Zhiyuan Zha, Xin Yuan 0002, Bihan Wen, Jiachao Zhang, Jiantao Zhou 0001, Ce Zhu |
IEEE Trans. Image Process. | 4 |
| 2019 | A Comparative Study for the Nuclear Norms Minimization MethodsabstractThe nuclear norm minimization (NNM) is commonly used to approximate the matrix rank by shrinking all singular values equally. However, the singular values have clear physical meanings in many practical problems, and NNM may not be able to faithfully approximate the matrix rank. To alleviate the above-mentioned limitation of NNM, recent studies have suggested that the weighted nuclear norm minimization (WNNM) can achieve a better rank estimation than NNM, which heuristically set the weight being inverse to the singular values. However, it still lacks a rigorous explanation why WNNM is more effective than NMM in various applications. In this paper, we analyze NNM and WNNM from the perspective of group sparse representation (GSR). Concretely, an adaptive dictionary learning method is devised to connect the rank minimization and GSR models. Based on the proposed dictionary, we prove that NNM and WNNM are equivalent to ℓ1-norm minimization and the weighted ℓ1-norm minimization in GSR, respectively. Inspired by enhancing sparsity of the weighted ℓ1-norm minimization in comparison with ℓ1-norm minimization in sparse representation, we thus explain that WNNM is more effective than NMM. By integrating the image nonlocal self-similarity (NSS) prior with the WNNM model, we then apply it to solve the image denoising problem. Experimental results demonstrate that WNNM is more effective than NNM and outperforms several state-of-the-art methods in both objective and perceptual quality. Zhiyuan Zha, Bihan Wen, Jiachao Zhang, Jiantao Zhou 0001, Ce Zhu |
ICIP | 3 |
| 2019 | Simultaneous Nonlocal Self-Similarity Prior for Image DenoisingabstractNonlocal image representation has achieved great success in various image processing tasks such as image denoising, image deblurring and image deblocking. Particularly, by exploiting the image nonlo-cal self-similarity (NSS) prior, many nonlocal similar patches can be searched across the whole image for a given patch, which has significantly boosted the performance of image restoration. To the best of our knowledge, most existing methods only consider the NSS prior of the input degraded image, while few methods exploit the NSS prior from external clean image corpus. However, how to utilize the NSS priors of input degraded image and external clean image corpus simultaneously is still an open problem. In this paper, we propose a novel approach for image denoising, which exploits simultaneous nonlocal self-similarity (SNSS) by integrating the NSS priors of both the input degraded image and external clean image corpus. Firstly, we search and group nonlocal similar patches from a clean image corpus, and a group-based Gaussian Mixture Model (GMM) learning algorithm is developed to learn an external NSS prior. Then, an optimal group is selected from the best suitable Gaussian component for a group of the noisy image. By integrating the group of the noisy image and the corresponding group of the Gaussian component with a low-rank constraint, an iterative algorithm is developed to solve the proposed SNSS model. Experimental results demonstrate that the proposed SNSS-based denoising method produces superior results compared with many state-of-the-art denoising methods in both objective and perceptual quality. Zhiyuan Zha, Xin Yuan 0002, Bihan Wen, Jiachao Zhang, Jiantao Zhou 0001, Ce Zhu |
ICIP | 4 |
| 2018 | A Wavelet-GSM Approach to DemosaickingabstractWe propose a wavelet-based Gaussian scale mixture (GSM) demosaicking method. The wavelet coefficients of the proposed method corresponding to the luminance and chrominance components are reconstructed using Bayesian minimum mean square error estimation. The proposed wavelet-GSM prior exploits the correlation of neighboring wavelets coefficients to improve upon a previously proposed posterior sparsity directed demosaicking method. As a result, our proposed demosaicking method suppresses the zippering artifacts more effectively than the state of the arts. Jiachao Zhang, Andong Sheng, Keigo Hirakawa |
IEEE Signal Process. Lett. | 1 |
| 2018 | Pixel Binning for High Dynamic Range Color Image Sensor Using Square Sampling LatticeabstractWe propose a new pixel binning scheme for color image sensors. We minimized distortion caused by binning by requiring that the superpixels lie on a square sampling lattice. The proposed binning schemes achieve the equivalent of 4.42 times signal strength improvement with the image resolution loss of 5 times, higher in noise performance and in resolution than the existing binning schemes. As a result, the proposed binning has considerably less artifacts and better noise performance compared with the existing binning schemes. In addition, we provide an extension to the proposed binning scheme for performing single-shot high dynamic range image acquisition. Jiachao Zhang, Andong Sheng, Keigo Hirakawa |
IEEE Trans. Image Process. | 1 |
| 2017 | Improved Denoising via Poisson Mixture Modeling of Image Sensor NoiseabstractThis paper describes a study aimed at comparing the real image sensor noise distribution to the models of noise often assumed in image denoising designs. A quantile analysis in pixel, wavelet transform, and variance stabilization domains reveal that the tails of Poisson, signal-dependent Gaussian, and Poisson-Gaussian models are too short to capture real sensor noise behavior. A new Poisson mixture noise model is proposed to correct the mismatch of tail behavior. Based on the fact that noise model mismatch results in image denoising that undersmoothes real sensor data, we propose a mixture of Poisson denoising method to remove the denoising artifacts without affecting image details, such as edge and textures. Experiments with real sensor data verify that denoising for real image sensor data is indeed improved by this new technique. Jiachao Zhang, Keigo Hirakawa |
IEEE Trans. Image Process. | 1 |
| 2016 | A real-time GPU-based coupled fluid-structure simulation with haptic interactionabstractTime series scientific simulation on supercomputers generates huge amounts of data at each time step. In the big data era, these data become impossible to be stored anymore, so simultaneous analysis of these data is strongly demanded. In order to realize such an on-the-fly intuitive analysis of simulation, this paper showcases a multimodal visualization and steering system with visual and haptic interfaces for a real-time GPU-based coupled fluid-structure simulation. Since the nature of touching sense of human beings requires extremely fast and continuous refreshing to form authentic feeling, parallel techniques were utilized to speed up haptic updating to around one millisecond. A middle-layer interface for a haptic device was developed to realize the ease of use of this device. Furthermore, a model of palpation was proposed to allow users to touch, push and sense the dynamic fluid motion inside a deformable tube. Jiachao Zhang, Shunpei Yuasa, Shinji Fukuma, Shin-ichiro Mori |
ICIS | 1 |
| 2015 | Quantile analysis of image sensor noise distributionabstractThis paper describes a study aimed at comparing the real image sensor noise distribution to the models of noise often assumed in image denoising designs. Quantile analysis in pixel, wavelet, and variance stabilization domains reveal that the tails of Poisson, signal-dependent Gaussian, and Poisson-Gaussian models are too short to capture real sensor noise behavior. Noise model mismatch would likely result in image denoising that undersmoothes real sensor data. Jiachao Zhang, Keigo Hirakawa, Xiaodan Jin |
ICASSP | 1 |