VLDB 2026 Research / reviewers in the wild / expert
Bo Zhang 0025
dblp:36/2259-25
· DBLP profile ↗
30ranked-venue papers
7as first author
16since 2021 · last 2025
0000-0002-9795-4673ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 5 first-author · 14 since 2021Artificial intelligence and machine learning · 20 · 3 first-author · 14 since 2021Systems, architecture and hardware · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MovieDreamer: Hierarchical Generation for Coherent Long Visual SequencesabstractRecent advancements in video generation have primarily leveraged diffusion models for short-duration content. However, these approaches often fall short in modeling complex narratives and maintaining character consistency over extended periods, which is essential for long-form video production like movies. We propose MovieDreamer, a novel hierarchical framework that integrates the strengths of autoregressive models with diffusion-based rendering to pioneer long-duration video generation with intricate plot progressions and high visual fidelity. Our approach utilizes autoregressive models for global narrative coherence, predicting sequences of visual tokens that are subsequently transformed into high-quality video frames through diffusion rendering. This method is akin to traditional movie production processes, where complex stories are factorized down into manageable scene capturing. Further, we employ a multimodal script that enriches scene descriptions with detailed character information and visual style, enhancing continuity and character identity across scenes. We present extensive experiments across various movie genres, demonstrating that our approach not only achieves superior visual and narrative quality but also effectively extends the duration of generated content significantly beyond current capabilities. Canyu Zhao, Wen Wang 0015, Fan Wang 0019, Hao Chen 0041, Bo Zhang 0025, Chunhua Shen |
ICLR | 7 |
| 2024 | 3DFaceShop: Explicitly Controllable 3D-Aware Portrait GenerationabstractIn contrast to the traditional avatar creation pipeline which is a costly process, contemporary generative approaches directly learn the data distribution from photographs. While plenty of works extend unconditional generative models and achieve some levels of controllability, it is still challenging to ensure multi-view consistency, especially in large poses. In this work, we propose a network that generates 3D-aware portraits while being controllable according to semantic parameters regarding pose, identity, expression and illumination. Our network uses neural scene representation to model 3D-aware portraits, whose generation is guided by a parametric face model that supports explicit control. While the latent disentanglement can be further enhanced by contrasting images with partially different attributes, there still exists noticeable inconsistency in non-face areas when animating expressions. We solve this by proposing a volume blending strategy in which we form a composite output by blending dynamic and static areas, with two parts segmented from the jointly learned semantic field. Our method outperforms prior arts in extensive experiments, producing realistic portraits with vivid expression in natural lighting when viewed from free viewpoints. It also demonstrates generalization ability to real images as well as out-of-domain data, showing great promise in real applications. Junshu Tang, Bo Zhang 0025, Binxin Yang, Ting Zhang 0002, Dong Chen 0003, Lizhuang Ma, Fang Wen 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2023 | RODIN: A Generative Model for Sculpting 3D Digital Avatars Using DiffusionabstractThis paper presents a 3D diffusion model that automatically generates 3D digital avatars represented as neural radiance fields (NeRFs). A significant challenge for 3D diffusion is that the memory and processing costs are prohibitive for producing high-quality results with rich details. To tackle this problem, we propose the roll-out diffusion network (RODIN), which takes a 3D NeRF model represented as multiple 2D feature maps and rolls out them onto a single 2D feature plane within which we perform 3D-aware diffusion. The RODIN model brings much-needed computational efficiency while preserving the integrity of 3D diffusion by using 3D-aware convolution that attends to projected features in the 2D plane according to their original relationships in 3D. We also use latent conditioning to orchestrate the feature generation with global coherence, leading to high-fidelity avatars and enabling semantic editing based on text prompts. Finally, we use hierarchical synthesis to further enhance details. The 3D avatars generated by our model compare favorably with those produced by existing techniques. We can generate highly detailed avatars with realistic hairstyles and facial hair. We also demonstrate 3D avatar generation from image or text, as well as text-guided editability. Tengfei Wang 0002, Bo Zhang 0025, Ting Zhang 0002, Shuyang Gu, Jianmin Bao, Tadas Baltrusaitis, Jingjing Shen, Dong Chen 0003, Fang Wen 0001, Qifeng Chen 0001, Baining Guo |
CVPR | 2 |
| 2023 | Paint by Example: Exemplar-based Image Editing with Diffusion ModelsabstractLanguage-guided image editing has achieved great success recently. In this paper, we investigate exemplar-guided image editing for more precise control. We achieve this goal by leveraging self-supervised training to disentangle and re-organize the source image and the exemplar. However, the naive approach will cause obvious fusing artifacts. We carefully analyze it and propose a content bottleneck and strong augmentations to avoid the trivial solution of directly copying and pasting the exemplar image. Meanwhile, to ensure the controllability of the editing process, we design an arbitrary shape mask for the exemplar image and leverage the classifier-free guidance to increase the similarity to the exemplar image. The whole framework involves a single forward of the diffusion model without any iterative optimization. We demonstrate that our method achieves an impressive performance and enables controllable editing on in-the-wild images with high fidelity. The code and pretrained models are available at https://github.com/Fantasy-Studio/Paint-by-Example. Binxin Yang, Shuyang Gu, Bo Zhang 0025, Ting Zhang 0002, Xuejin Chen, Xiaoyan Sun 0001, Dong Chen 0003, Fang Wen 0001 |
CVPR | 3 |
| 2023 | MetaPortrait: Identity-Preserving Talking Head Generation with Fast Personalized AdaptationabstractIn this work, we propose an ID-preserving talking head generation framework, which advances previous methods in two aspects. First, as opposed to interpolating from sparse flow, we claim that dense landmarks are crucial to achieving accurate geometry-aware flow fields. Second, inspired by face-swapping methods, we adaptively fuse the source identity during synthesis, so that the network better preserves the key characteristics of the image portrait. Although the proposed model surpasses prior generation fidelity on established benchmarks, personalized fine-tuning is still needed to further make the talking head generation qualified for real usage. However, this process is rather computationally demanding that is unaffordable to standard users. To alleviate this, we propose a fast adaptation model using a metalearning approach. The learned model can be adapted to a high-quality personalized model as fast as 30 seconds. Last but not least, a spatial-temporal enhancement module is proposed to improve the fine details while ensuring temporal coherency. Extensive experiments prove the significant superiority of our approach over the state of the arts in both one-shot and personalized settings. Bowen Zhang 0010, Pan Zhang 0003, Bo Zhang 0025, HsiangTao Wu, Dong Chen 0003, Qifeng Chen 0001, Fang Wen 0001 |
CVPR | 4 |
| 2023 | Make-It-3D: High-Fidelity 3D Creation from A Single Image with Diffusion PriorabstractIn this work, we investigate the problem of creating high-fidelity 3D content from only a single image. This is inherently challenging: it essentially involves estimating the underlying 3D geometry while simultaneously hallucinating unseen textures. To address this challenge, we leverage prior knowledge from a well-trained 2D diffusion model to act as 3D-aware supervision for 3D creation. Our approach, Make-It-3D, employs a two-stage optimization pipeline: the first stage optimizes a neural radiance field by incorporating constraints from the reference image at the frontal view and diffusion prior at novel views; the second stage transforms the coarse model into textured point clouds and further elevates the realism with diffusion prior while leveraging the high-quality textures from the reference image. Extensive experiments demonstrate that our method outperforms prior works by a large margin, resulting in faithful reconstructions and impressive visual quality. Our method presents the first attempt to achieve high-quality 3D creation from a single image for general objects and enables various applications such as text-to-3D creation and texture editing. Junshu Tang, Tengfei Wang 0002, Bo Zhang 0025, Ting Zhang 0002, Ran Yi 0002, Lizhuang Ma, Dong Chen 0003 |
ICCV | 3 |
| 2023 | Old Photo Restoration via Deep Latent Space TranslationabstractWe propose to restore old photos that suffer from severe degradation through a deep learning approach. Unlike conventional restoration tasks that can be solved through supervised learning, the degradation in real photos is complex and the domain gap between synthetic images and real old photos makes the network fail to generalize. Therefore, we propose a novel triplet domain translation network by leveraging real photos along with massive synthetic image pairs. Specifically, we train two variational autoencoders (VAEs) to respectively transform old photos and clean photos into two latent spaces. And the translation between these two latent spaces is learned with synthetic paired data. This translation generalizes well to real photos because the domain gap is closed in the compact latent space. Besides, to address multiple degradations mixed in one old photo, we design a global branch with a partial nonlocal block targeting the structured defects, such as scratches and dust spots, and a local branch targeting the unstructured defects, such as noises and blurriness. We also extend the global branch with a more memory-efficient scheme, named multi-scale patch-based attention to processing high-resolution photos. Two branches are fused in the latent space, leading to improved capability to restore old photos from multiple defects. Furthermore, we apply another face refinement network to recover fine details of faces in the old photos, thus ultimately generating photos with enhanced perceptual quality. With comprehensive experiments, the proposed pipeline demonstrates superior performance over state-of-the-art methods as well as existing commercial tools in terms of visual quality for old photos restoration. Both code and models could be found at https://github.com/microsoft/Bringing-Old-Photos-Back-to-Life. Ziyu Wan, Bo Zhang 0025, Dongdong Chen 0001, Pan Zhang 0003, Dong Chen 0003, Fang Wen 0001, Jing Liao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Vector Quantized Diffusion Model for Text-to-Image SynthesisabstractWe present the vector quantized diffusion (VQ-Diffusion) model for text-to-image generation. This method is based on a vector quantized variational autoencoder (VQ-VAE) whose latent space is modeled by a conditional variant of the recently developed Denoising Diffusion Probabilistic Model (DDPM). We find that this latent-space method is well-suited for text-to-image generation tasks because it not only eliminates the unidirectional bias with existing methods but also allows us to incorporate a mask-and-replace diffusion strategy to avoid the accumulation of errors, which is a serious problem with existing methods. Our experiments show that the VQ-Diffusion produces significantly better text-to-image generation results when compared with conventional autoregressive (AR) models with similar numbers of parameters. Compared with previous GAN-based text-to-image methods, our VQ-Diffusion can handle more complex scenes and improve the synthesized image quality by a large margin. Finally, we show that the image generation computation in our method can be made highly efficient by reparameterization. With traditional AR methods, the text-to-image generation time increases linearly with the output image resolution and hence is quite time consuming even for normal size images. The VQ-Diffusion allows us to achieve a better trade-off between quality and speed. Our experiments indicate that the VQ-Diffusion model with the reparameterization is fifteen times faster than traditional AR methods while achieving a better image quality. The code and models are available at https://github.com/cientgu/VQ-Diffusion. Shuyang Gu, Dong Chen 0003, Jianmin Bao, Fang Wen 0001, Bo Zhang 0025, Dongdong Chen 0001, Lu Yuan 0001, Baining Guo |
CVPR | 5 |
| 2022 | Bringing Old Films Back to LifeabstractWe present a learning-based framework, recurrent transformer network (RTN), to restore heavily degraded old films. Instead of performing frame-wise restoration, our method is based on the hidden knowledge learned from adjacent frames that contain abundant information about the occlusion, which is beneficial to restore challenging artifacts of each frame while ensuring temporal coherency. Moreover, contrasting the representation of the current frame and the hidden knowledge makes it possible to infer the scratch position in an unsupervised manner, and such defect localization generalizes well to real-world degradations. To better resolve mixed degradation and compensate for the flow estimation error during frame alignment, we propose to leverage more expressive transformer blocks for spatial restoration. Experiments on both synthetic dataset and real-world old films demonstrate the significant superiority of the proposed RTN over existing solutions. In addition, the same framework can effectively propagate the color from keyframes to the whole video, ultimately yielding compelling restored films. The implementation and model will be released at https://github.com/raywzy/Bringing-Old-Films-Back-to-Life. Ziyu Wan, Bo Zhang 0025, Dongdong Chen 0001, Jing Liao 0001 |
CVPR | 2 |
| 2022 | StyleSwin: Transformer-based GAN for High-resolution Image GenerationabstractDespite the tantalizing success in a broad of vision tasks, transformers have not yet demonstrated on-par ability as ConvNets in high-resolution image generative modeling. In this paper, we seek to explore using pure transformers to build a generative adversarial network for high-resolution image synthesis. To this end, we believe that local attention is crucial to strike the balance between computational efficiency and modeling capacity. Hence, the proposed generator adopts Swin transformer in a style-based architecture. To achieve a larger receptive field, we propose double attention which simultaneously leverages the context of the local and the shifted windows, leading to improved generation quality. Moreover, we show that offering the knowledge of the absolute position that has been lost in window-based transformers greatly benefits the generation quality. The proposed StyleSwin is scalable to high resolutions, with both the coarse geometry and fine structures benefit from the strong expressivity of transformers. However, blocking artifacts occur during high-resolution synthesis because performing the local attention in a block-wise manner may break the spatial coherency. To solve this, we empirically investigate various solutions, among which we find that employing a wavelet discriminator to examine the spectral discrepancy effectively suppresses the artifacts. Extensive experiments show the superiority over prior transformer-based GANs, especially on high resolutions, e.g.,$1024 \times$1024. The StyleSwin, without complex training strategies, excels over StyleGAN on CelebA-HQ 1024, and achieves on-par performance on FFHQ-1024, proving the promise of using transformers for high-resolution image generation. The code and pretrained models are available at https://github.com/microsoft/StyleSwin. Bowen Zhang 0010, Shuyang Gu, Bo Zhang 0025, Jianmin Bao, Dong Chen 0003, Fang Wen 0001, Baining Guo |
CVPR | 3 |
| 2022 | Real-Time Neural Character Rendering with Pose-Guided Multiplane Images
Hao Ouyang, Bo Zhang 0025, Pan Zhang 0003, Hao Yang 0036, Jiaolong Yang, Dong Chen 0003, Qifeng Chen 0001, Fang Wen 0001 |
ECCV (32) | 2 |
| 2022 | Deep Sketch-Guided Cartoon Video InbetweeningabstractWe propose a novel framework to produce cartoon videos by fetching the color information from two input keyframes while following the animated motion guided by a user sketch. The key idea of the proposed approach is to estimate the dense cross-domain correspondence between the sketch and cartoon video frames, and employ a blending module with occlusion estimation to synthesize the middle frame guided by the sketch. After that, the input frames and the synthetic frame equipped with established correspondence are fed into an arbitrary-time frame interpolation pipeline to generate and refine additional inbetween frames. Finally, a module to preserve temporal consistency is employed. Compared to common frame interpolation methods, our approach can address frames with relatively large motion and also has the flexibility to enable users to control the generated video sequences by editing the sketch guidance. By explicitly considering the correspondence between frames and the sketch, we can achieve higher quality results than other image synthesis methods. Our results show that our system generalizes well to different movie frames, achieving better results than existing solutions. Xiaoyu Li 0002, Bo Zhang 0025, Jing Liao 0001, Pedro V. Sander |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2021 | Style-Based Point Generator With Adversarial Rendering for Point Cloud CompletionabstractIn this paper, we proposed a novel Style-based Point Generator with Adversarial Rendering (SpareNet) for point cloud completion. Firstly, we present the channel-attentive EdgeConv to fully exploit the local structures as well as the global shape in point features. Secondly, we observe that the concatenation manner used by vanilla foldings limits its potential of generating a complex and faithful shape. Enlightened by the success of StyleGAN, we regard the shape feature as style code that modulates the normalization layers during the folding, which considerably enhances its capability. Thirdly, we realize that existing point supervisions, e.g., Chamfer Distance or Earth Mover’s Distance, cannot faithfully reflect the perceptual quality of the reconstructed points. To address this, we propose to project the completed points to depth maps with a differentiable renderer and apply adversarial training to advocate the perceptual realism under different viewpoints. Comprehensive experiments on ShapeNet and KITTI prove the effectiveness of our method, which achieves state-of-the-art quantitative performance while offering superior visual quality. Chulin Xie, Chuxin Wang, Bo Zhang 0025, Hao Yang 0036, Dong Chen 0003, Fang Wen 0001 |
CVPR | 3 |
| 2021 | Prototypical Pseudo Label Denoising and Target Structure Learning for Domain Adaptive Semantic SegmentationabstractSelf-training is a competitive approach in domain adaptive segmentation, which trains the network with the pseudo labels on the target domain. However inevitably, the pseudo labels are noisy and the target features are dispersed due to the discrepancy between source and target domains. In this paper, we rely on representative prototypes, the feature centroids of classes, to address the two issues for unsupervised domain adaptation. In particular, we take one step further and exploit the feature distances from prototypes that provide richer information than mere prototypes. Specifically, we use it to estimate the likelihood of pseudo labels to facilitate online correction in the course of training. Meanwhile, we align the prototypical assignments based on relative feature distances for two different views of the same target, producing a more compact target feature space. Moreover, we find that distilling the already learned knowledge to a self-supervised pretrained model further boosts the performance. Our method shows tremendous performance advantage over state-of-the-art methods. The code is available at https://github.com/microsoft/ProDA. Pan Zhang 0003, Bo Zhang 0025, Ting Zhang 0002, Dong Chen 0003, Fang Wen 0001 |
CVPR | 2 |
| 2021 | CoCosNet v2: Full-Resolution Correspondence Learning for Image TranslationabstractWe present the full-resolution correspondence learning for cross-domain images, which aids image translation. We adopt a hierarchical strategy that uses the correspondence from coarse level to guide the fine levels. At each hierarchy, the correspondence can be efficiently computed via PatchMatch that iteratively leverages the matchings from the neighborhood. Within each PatchMatch iteration, the ConvGRU module is employed to refine the current correspondence considering not only the matchings of larger context but also the historic estimates. The proposed Co-CosNet v2, a GRU-assisted PatchMatch approach, is fully differentiable and highly efficient. When jointly trained with image translation, full-resolution semantic correspondence can be established in an unsupervised manner, which in turn facilitates the exemplar-based image translation. Experiments on diverse translation tasks show that CoCosNet v2 performs considerably better than state-of-the-art literature on producing high-resolution images. Xingran Zhou, Bo Zhang 0025, Ting Zhang 0002, Pan Zhang 0003, Jianmin Bao, Dong Chen 0003, Zhongfei Zhang, Fang Wen 0001 |
CVPR | 2 |
| 2021 | Let's See Clearly: Contaminant Artifact Removal for Moving CamerasabstractContaminants such as dust, dirt and moisture adhering to the camera lens can greatly affect the quality and clarity of the resulting image or video. In this paper, we propose a video restoration method to automatically remove these contaminants and produce a clean video. Our approach first seeks to detect attention maps that indicate the regions that need to be restored. In order to leverage the corresponding clean pixels from adjacent frames, we propose a flow completion module to hallucinate the flow of the background scene to the attention regions degraded by the contaminants. Guided by the attention maps and completed flows, we propose a recurrent technique to restore the input frame by fetching clean pixels from adjacent frames. Finally, a multi-frame processing stage is used to further process the entire video sequence in order to enforce temporal consistency. The entire network is trained on a synthetic dataset that approximates the physical lighting properties of contaminant artifacts. This new dataset and our novel framework lead to our method that is able to address different contaminants and outperforms competitive restoration approaches both qualitatively and quantitatively. Xiaoyu Li 0002, Bo Zhang 0025, Jing Liao 0001, Pedro V. Sander |
ICCV | 2 |
| 2020 | Bringing Old Photos Back to LifeabstractWe propose to restore old photos that suffer from severe degradation through a deep learning approach. Unlike conventional restoration tasks that can be solved through supervised learning, the degradation in real photos is complex and the domain gap between synthetic images and real old photos makes the network fail to generalize. Therefore, we propose a novel triplet domain translation network by leveraging real photos along with massive synthetic image pairs. Specifically, we train two variational autoencoders (VAEs) to respectively transform old photos and clean photos into two latent spaces. And the translation between these two latent spaces is learned with synthetic paired data. This translation generalizes well to real photos because the domain gap is closed in the compact latent space. Besides, to address multiple degradations mixed in one old photo, we design a global branch with a partial nonlocal block targeting to the structured defects, such as scratches and dust spots, and a local branch targeting to the unstructured defects, such as noises and blurriness. Two branches are fused in the latent space, leading to improved capability to restore old photos from multiple defects. The proposed method outperforms state-of-the-art methods in terms of visual quality for old photos restoration. Ziyu Wan, Bo Zhang 0025, Dongdong Chen 0001, Pan Zhang 0003, Dong Chen 0003, Jing Liao 0001, Fang Wen 0001 |
CVPR | 2 |
| 2020 | Cross-Domain Correspondence Learning for Exemplar-Based Image TranslationabstractWe present a general framework for exemplar-based image translation, which synthesizes a photo-realistic image from the input in a distinct domain (e.g., semantic segmentation mask, or edge map, or pose keypoints), given an exemplar image. The output has the style (e.g., color, texture) in consistency with the semantically corresponding objects in the exemplar. We propose to jointly learn the cross-domain correspondence and the image translation, where both tasks facilitate each other and thus can be learned with weak supervision. The images from distinct domains are first aligned to an intermediate domain where dense correspondence is established. Then, the network synthesizes images based on the appearance of semantically corresponding patches in the exemplar. We demonstrate the effectiveness of our approach in several image translation tasks. Our method is superior to state-of-the-art methods in terms of image quality significantly, with the image style faithful to the exemplar with semantic consistency. Moreover, we show the utility of our method for several applications. Pan Zhang 0003, Bo Zhang 0025, Dong Chen 0003, Lu Yuan 0001, Fang Wen 0001 |
CVPR | 2 |
| 2019 | Blind Geometric Distortion Correction on Images Through Deep LearningabstractWe propose the first general framework to automatically correct different types of geometric distortion in a single input image. Our proposed method employs convolutional neural networks (CNNs) trained by using a large synthetic distortion dataset to predict the displacement field between distorted images and corrected images. A model fitting method uses the CNN output to estimate the distortion parameters, achieving a more accurate prediction. The final corrected image is generated based on the predicted flow using an efficient, high-quality resampling method. Experimental results demonstrate that our algorithm outperforms traditional correction methods, and allows for interesting applications such as distortion transfer, distortion exaggeration, and co-occurring distortion correction. Xiaoyu Li 0002, Bo Zhang 0025, Pedro V. Sander, Jing Liao 0001 |
CVPR | 2 |
| 2019 | Deep Exemplar-Based Video ColorizationabstractThis paper presents the first end-to-end network for exemplar-based video colorization. The main challenge is to achieve temporal consistency while remaining faithful to the reference style. To address this issue, we introduce a recurrent framework that unifies the semantic correspondence and color propagation steps. Both steps allow a provided reference image to guide the colorization of every frame, thus reducing accumulated propagation errors. Video frames are colorized in sequence based on the colorization history, and its coherency is further enforced by the temporal consistency loss. All of these components, learned end-to-end, help produce realistic videos with good temporal stability. Experiments show our result is superior to the state-of-the-art methods both quantitatively and qualitatively. Bo Zhang 0025, Mingming He, Jing Liao 0001, Pedro V. Sander, Lu Yuan 0001, Amine Bermak, Dong Chen 0003 |
CVPR | 1 |
| 2019 | Microshift: An Efficient Image Compression Algorithm for HardwareabstractIn this paper, we propose a lossy image compression algorithm called microshift. We employ an algorithm-hardware co-design methodology, yielding a hardware-friendly compression approach with low power consumption. In our method, the image is first micro-shifted, and then the sub-quantized values are further compressed. Two methods, FAST and MRF models, are proposed to recover the bitdepth by exploiting the spatial correlation of natural images. Both methods can decompress images progressively. On an average, our compression algorithm can compress images to 1.25-bits per pixel with a resulting quality that outperforms the state-of-the-art on-chip compression algorithms in both peak signal-to-noise ratio and structual similarity. Then, we propose a hardware architecture and implement the algorithm on an FPGA. The results on the ASIC design further validate the low-hardware complexity and high-power efficiency, showing that our method is promising, particularly for low-power wireless vision sensor networks. Bo Zhang 0025, Pedro V. Sander, Chi-Ying Tsui, Amine Bermak |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | Document rectification and illumination correction using a patch-based CNNabstractWe propose a novel learning method to rectify document images with various distortion types from a single input image. As opposed to previous learning-based methods, our approach seeks to first learn the distortion flow on input image patches rather than the entire image. We then present a robust technique to stitch the patch results into the rectified document by processing in the gradient domain. Furthermore, we propose a second network to correct the uneven illumination, further improving the readability and OCR accuracy. Due to the less complex distortion present on the smaller image patches, our patch-based approach followed by stitching and illumination correction can significantly improve the overall accuracy in both the synthetic and real datasets. Xiaoyu Li 0002, Bo Zhang 0025, Jing Liao 0001, Pedro V. Sander |
ACM Trans. Graph. | 2 |
| 2017 | Gradient magnitude similarity deviation on multiple scales for color image quality assessmentabstractRecently, various image quality assessment (IQA) metrics based on gradient similarity have been developed. In this paper, we extend the work of gradient magnitude similarity deviation (GMSD) and propose a more efficient metric. First, a novel similarity index is proposed, which gives the flexibility to tune the masking parameter to more closely match the human vision system (HVS). Then, we propose a multi-scale GMSD method by incorporating scores of luminance distortion at different scales. Furthermore, a method for measuring chromatic distortions in YIQ color space based on our metric is proposed. The final IQA index, MS-GMSDc, is obtained by combining luminance and chrominance scores. Experimental results on four comprehensive datasets clearly show that, compared with 14 state-of-the-art IQA methods, our method achieves the best performance for both grayscale and chromatic image assessment. Bo Zhang 0025, Pedro V. Sander, Amine Bermak |
ICASSP | 1 |
| 2017 | Registration based retargeted image quality assessmentabstractIn recent years, a large number of image retargeting methods have been proposed. Measuring their relative quality is of significant importance, and there is still room for improvement in the effectiveness of objective retargeted image quality assessment (RIQA) metrics. In this paper, we propose a registration based RIQA metric. First, we propose to calculate the flow map using an image registration method which involves SURF point matching and halfway domain optimization. Using the computed flow map and the source image, we propose an LGI metric which contains three factors: 1) local similarity which assesses the local aspect ratio change, edge directional similarity and flow smoothness; 2) global distortion which measures the appearance change of salient objects; 3) salient information loss. Comparing with other six metrics, our LGI metric correlates the best with subjective rankings on the RetargetMe dataset. Bo Zhang 0025, Pedro V. Sander, Amine Bermak |
ICASSP | 1 |
| 2016 | Wide dynamic range PSD algorithms and their implementation for compressive imagingabstractPlanned Sensor Distortion (PSD) is a compression method that quantizes shifted signal with low bit depth. In this paper, we analyze the dynamic range loss issue in the PSD algorithm and propose two novel methods to overcome this issue: a blocked PSD, which divides the image into sub-blocks that adapt to pixel values, and an auto-reset PSD, which utilizes Markov property to recover a high dynamic range image from the modulo image. Simulation on a 3-bit depth image of indoor environment shows PSNR of 35.7dB and 35.0dB respectively after reconstruction using our algorithms. Thereafter, two different implementations for PSD are proposed, introducing shifts at either reset phase or readout phase. These circuits are then extended to be compatible with our proposed algorithms. Finally, simulation results using 0.18um GlobalFoundries process validate our designs. Spontaneous power optimization of ADC and transmission, hardware friendly feature, and the ability of high quality imaging make our compression method promising. Bo Zhang 0025, Xiaopeng Zhong, Bo Wang 0012, Pedro V. Sander, Amine Bermak |
ISCAS | 1 |
| 2016 | A background subtraction based column-parallel analog-to-information converter for motion-triggered vision sensorabstractAn analog-to-information converter (AIC) enables information quantization instead of signal quantization, which can reduce both quantization efforts and data bandwidth. Therefore, an AIC will relax data processing and transmission burden and improve overall power efficiency, making it attractive for wireless vision sensor networks. In this paper, we propose a background subtraction based column-parallel AIC for motion-triggered vision sensors. It features low-power and robust background subtraction for motion extraction as well as efficient information quantization. The AIC is implemented by integrating background subtraction with a successive-approximation-register and single-lope (SAR-SS) hybrid ADC using a new architecture. Targeting at scene interpretation applications, 6-bit background and 8-bit foreground are adopted. Simulation results show that vision data bandwidth and ADC power are reduced by 85.23% and 95.88% respectively when full-bit-depth motion images are required. The bandwidth reduction and the ADC power reduction can be further improved to 87.50% and 98.48% respectively if only binary motion images are needed. Xiaopeng Zhong, Bo Zhang 0025, Amine Bermak |
ISCAS | 2 |
| 2005 | Polygonal Shape Blending with Topological Evolutions
Ligang Liu 0001, Bo Zhang 0025, Baining Guo, Harry Shum |
J. Comput. Sci. Technol. | 2 |
| 2004 | Perceptually Based Approach for Planar Shape MorphingabstractThis paper presents an approach for establishing vertex correspondences between two planar shapes. Correspondences are established between the perceptual feature points extracted from both source and target shapes. A similarity metric between two feature points is defined using the intrinsic properties of their local neighborhoods. The optimal correspondence is found by an efficient dynamic programming technique. Our approach treats shape noise by allowing discarding small feature points, which introduces skips in the traversal of the dynamic programming graph. Our method is fast, feature preserving, and invariant to geometric transformations. We demonstrate the superiority of our approach over other approaches by experimental results. Ligang Liu 0001, Guopu Wang, Bo Zhang 0025, Baining Guo, Harry Shum |
PG | 3 |
| 2001 | Spoken dialogue management as planning and acting under uncertaintyabstractSome stochastic models like Markov decision process (MDP) are used to model the dialogue manager. MDP-based system degrades fast when uncertainty about user’s intention increases. We propose a novel dialogue model based on the partially observable Markov decision process (POMDP). We use hidden system states and user intentions as the state set, parser results and low-level information as the observation set, domain actions and dialogue repair actions as the action set. Here the low-level information is extracted from different input modals using Bayesian networks. Because of the limitation of exact algorithms, we focus on heuristic methods and their applicability in dialogue management. Bo Zhang 0025, Qingsheng Cai, Jianfeng Mao, Eric Chang, Baining Guo |
INTERSPEECH | 1 |
| 2001 | Planning and Acting under Uncertainty: A New Model for Spoken Dialogue System
Bo Zhang 0025, Qingsheng Cai, Jianfeng Mao, Baining Guo |
UAI | 1 |