Zhe Zhu

dblp:79/6387 · DBLP profile ↗
← Back
53ranked-venue papers
17as first author
30since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 30 · 10 first-author · 19 since 2021Artificial intelligence and machine learning · 10 · 4 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 5 since 2021Systems, architecture and hardware · 5 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021
YearPublicationVenuePosition
2026 BridgeShape: Latent Diffusion Schrödinger Bridge for 3D Shape Completion
abstract
Existing diffusion-based 3D shape completion methods typically use a conditional paradigm, injecting incomplete shape information into the denoising network via deep feature interactions (e.g., concatenation, cross-attention) to guide sampling toward complete shapes, often represented by voxel-based distance functions. However, these approaches fail to explicitly model the optimal global transport path, leading to suboptimal completions. Moreover, performing diffusion directly in voxel space imposes resolution constraints, limiting the generation of fine-grained geometric details. To address these challenges, we propose BridgeShape, a novel framework for 3D shape completion via latent diffusion Schrödinger bridge. The key innovations lie in two aspects: (i) BridgeShape formulates shape completion as an optimal transport problem, explicitly modeling the transition between incomplete and complete shapes to ensure a globally coherent transformation. (ii) We introduce a Depth-Enhanced Vector Quantized Variational Autoencoder (VQ-VAE) to encode 3D shapes into a compact latent space, leveraging self-projected multi-view depth information enriched with strong DINOv2 features to enhance geometric structural perception. By operating in a compact yet structurally informative latent space, BridgeShape effectively mitigates resolution constraints and enables more efficient and high-fidelity 3D shape completion. BridgeShape achieves state-of-the-art performance on 3D shape completion benchmarks, demonstrating superior fidelity at higher resolutions and for unseen object classes.
Dequan Kong, Honghua Chen, Zhe Zhu, Mingqiang Wei
AAAI3
2026 Sensitivity Analysis of Process Variables for Large-Scale Circuits based on Graph Attention Networks
Zhe Zhu, Jun Tao 0001
ISCAS1
2026 CoreEditor: Correspondence-Constrained Diffusion for Consistent 3D Editing
abstract
Text-driven 3D editing is an emerging task that focuses on modifying scenes based on text prompts. Current methods often adapt pre-trained 2D image editors to multi-view observations, using specific strategies to combine information across views. However, these approaches still struggle with ensuring consistency across views, as they lack precise control over the sharing of information, resulting in edits with insufficient visual changes and blurry details. In this paper, we propose CoreEditor, a novel framework for consistent text-to-3D editing. At the core of our approach is a novel correspondence-constrained attention mechanism, which enforces structured interactions between corresponding pixels that are expected to remain visually consistent during the diffusion denoising process. Unlike conventional wisdom that relies solely on scene geometry, we enhance the correspondence by incorporating semantic similarity derived from the diffusion denoising process. This combined support from both geometry and semantics ensures a robust multi-view editing process. Additionally, we introduce a selective editing pipeline that enables users to choose their preferred edits from multiple candidates, creating a more flexible and user-centered 3D editing process. Extensive experiments demonstrate the effectiveness of CoreEditor, showing its ability to generate high-quality 3D edits, significantly outperforming existing methods.
Zhe Zhu, Honghua Chen, Peng Li 0064, Mingqiang Wei
IEEE Trans. Vis. Comput. Graph.1
2026 Rectangling stitched images via unsupervised warping
Yun Zhang 0024, Jialing Yang, Zhe Zhu, Yukun Lai, Xinyuan Zheng
Vis. Comput.4
2025 CosCAD: Cross-Modal CAD Model Retrieval and Pose Alignment from a Single Image
Zhikun Wen, Honghua Chen, Zhe Zhu, Zeyong Wei, Liangliang Nan, Mingqiang Wei
CVM (2)3
2025 GenPC: Zero-shot Point Cloud Completion via 3D Generative Priors
abstract
Existing point cloud completion methods, which typically depend on predefined synthetic training datasets, encounter significant challenges when applied to out-of-distribution, real-world scans. To overcome this limitation, we introduce a zero-shot completion framework, termed GenPC, designed to reconstruct high-quality real-world scans by leveraging explicit 3D generative priors. Our key insight is that recent feed-forward 3D generative models, trained on extensive internet-scale data, have demonstrated the ability to perform 3D generation from single-view images in a zero-shot setting. To harness this for completion, we first develop a Depth Prompting module that links partial point clouds with image-to-3D generative models by leveraging depth images as a stepping stone. To retain the original partial structure in the final results, we design the Geometric Preserving Fusion module that aligns the generated shape with input by adaptively adjusting its pose and scale. Extensive experiments on widely used benchmarks validate the superiority and gen-eralizability of our approach, bringing us a step closer to robust real-world scan completion. Our code is available at https://github.com/liannuaa/GenPC.
Zhe Zhu, Mingqiang Wei
CVPR2
2025 TGGI: Text-Guided Object Removal in 3D Gaussian Scenes via Multi-View Image Inpainting
Liangyuan Zhang, Zhe Zhu, Lina Gong
CW2
2025 IDD: An Identity Disentanglement Framework for Deepfake Detection
Pengyao Xu, Zhe Zhu, Hongkuan Zhang
ICIC (5)3
2025 Fast Bounding Box Hierarchy
abstract
We introduce a fast algorithm designed for Bounding Box Hierarchy (BBH), a hierarchical tree structure wherein leaf nodes encapsulate sets of original bounding boxes, and non-leaf nodes represent merging operations of their children. A straightforward algorithm employs a bottom-up strategy in a brute force approach to construct the tree, entailing the iterative merging of candidate pairs with the minimum distance until reaching the root node. A pivotal challenge inherent to this brute force paradigm lies in its computational bottleneck, as determining the candidate pair with the minimum distance necessitates a global operation, rendering it highly computationally intensive. Our novel approach strategically circumvents this bottleneck by introducing the computation of approximate minimum distance pairs within local neighborhoods. By using the transformation of 2D bounding boxes into 1D space through Morton coding, the computational cost associated with identifying candidate bounding boxes for merging is significantly diminished, reducing it from O(N2) to O(logN). Compared with brute force approach which has the overall time complexity O(N3), our algorithm is only O(Nlog(N). The acceleration is critical for computation sensitive applications, particularly in embedded systems and mobile devices. Our approach also supports multi-class bounding box inputs, making it particularly useful for computer vision tasks where bounding boxes are inherently associated with class labels. We validate our algorithm on both real world data and simulated data.
Zhe Zhu, Joonsoo Kim, Arshita Gupta, Tien C. Bau
ICIP1
2025 BSGS: Bi-Stage 3D Gaussian Splatting for Camera Motion Deblurring
abstract
3D Gaussian Splatting has exhibited remarkable capabilities in 3D scene reconstruction. However, reconstructing high-quality 3D scenes from motion-blurred images caused by camera motion poses a significant challenge. The performance of existing 3DGS-based deblurring methods are limited due to their inherent mechanisms, such as extreme dependence on the accuracy of camera poses and inability to effectively control erroneous Gaussian primitives densification caused by motion blur. To solve these problems, we introduce a novel framework, Bi-Stage 3D Gaussian Splatting, to accurately reconstruct 3D scenes from motion-blurred images. BSGS contains two stages. First, Camera Pose Refinement roughly optimizes camera poses to reduce motion-induced distortions. Second, with fixed rough camera poses, Global Rigid Transformation further corrects motion-induced blur distortions. To alleviate multi-subframe gradient conflicts, we propose a subframe gradient aggregation strategy to optimize both stages. Furthermore, a space-time bi-stage optimization strategy is introduced to dynamically adjust primitive densification thresholds and prevent premature noisy Gaussian generation in blurred regions. Comprehensive experiments verify the effectiveness of our proposed deblurring method and show its superiority over the state of the arts.
Piaopiao Yu, Zhe Zhu, Mingqiang Wei
ACM Multimedia3
2025 KEREM: Enhancing Reliability and Transparency in Medical QA through LLM and Knowledge Graph Fusion
abstract
Open-domain medical question answering (QA) systems face significant challenges in achieving accurate reasoning and transparent explanations. In this study, we propose KEREM (Knowledge graph Enhanced Reasoning with Explainable Modeling), a novel framework that integrates large language models (LLMs) with structured medical knowledge graphs (KGs) to enable deep multimodal joint reasoning. KEREM supports multi-hop inference and causal chain explanations through four key modules: input processing and entity alignment, knowledge subgraph construction, joint reasoning and path control, and answer generation with natural language explanations. We evaluate KEREM on two real-world medical QA benchmarks—CMCQA and ChatDoctor-5k—where it consistently outperforms existing baselines in terms of answer accuracy, reasoning transparency, and structural consistency. Furthermore, lightweight supervised fine-tuning demonstrates KEREM’s strong contextual transferability and significantly improves its generation quality in specialized clinical settings. These findings highlight KEREM’s effectiveness in generating accurate answers and causal explanations, establishing a solid foundation for trustworthy medical QA in high-stakes clinical domains.
Shaojie Dong, Zhe Zhu, Pengyao Xu, Ke Shan, Shuwang Zhou
SMC2
2025 VarGes: Improving Variation in Co-Speech 3D Gesture Generation via StyleCLIPS
abstract
Generating expressive and diverse human gestures from audio is crucial in fields like human-computer interaction, virtual reality, and animation. While existing methods have achieved remarkable performance, they often exhibit limitations due to constrained dataset diversity and the restricted amount of information derived from audio inputs. To address these challenges, we present VarGes, a novel variationdriven framework designed to enhance co-speech gesture generation by integrating visual stylistic cues while maintaining naturalness. Our approach begins with a variation-enhanced feature extraction module, which seamlessly incorporates style-reference video data into a 3D human pose estimation network to extract StyleCLIPS, thereby enriching the input with stylistic information. Subsequently, we employ a variation-compensation style encoder, a transformer-style encoder equipped with an additive attention mechanism pooling layer, to robustly encode diverse StyleCLIPS representations and effectively manage stylistic variations. Finally, a variation-driven gesture predictor module fuses MFCC audio features with StyleCLIPS encodings via cross-attention, injecting this fused data into a cross-conditional autoregressive model to modulate 3D human gesture generation based on audio input and stylistic clues. The efficacy of our approach is validated on benchmark datasets, on which it outperforms existing methods in terms of gesture diversity and naturalness. Our code and video results are publicly available at https://github.com/mookerr/VarGES/.
Ke Mu, Yonggui Zhu, Zhe Zhu, Heyang Yan, Zhaoxin Fan
Comput. Vis. Media4
2025 3D Indoor Scene Geometry Estimation from a Single Omnidirectional Image: A Comprehensive Survey
abstract
This paper surveys the technology used in three-dimensional indoor scene geometry estimation from a single 360° omnidirectional image, which is pivotal in extracting 3D structural information from indoor environments. The technology transforms omnidirectional data into a 3D model, depicting spatial structure, object positions, and scene layout. Its significance spans various domains, including virtual reality (VR), augmented reality (AR), mixed reality (MR), game development, urban planning, and robot navigation. We begin by revisiting foundational concepts of omnidirectional imaging and detailing the problems, applications, and challenges in this field. Our review categorizes the fundamental tasks of structure recovery, depth estimation, and layout recovery. We also review pertinent datasets and evaluation metrics, providing the latest research as a reference. Finally, we summarize the field and discuss potential future trends to inform and guide further research.
Yonggui Zhu, Zhaoxin Li, Zhe Zhu
Comput. Vis. Media5
2025 PointSea: Point Cloud Completion via Self-structure Augmentation
Zhe Zhu, Honghua Chen, Mingqiang Wei
Int. J. Comput. Vis.1
2025 Toward Seamless Global 30-m Terrestrial Monitoring: Evaluating 2022 Cloud Free Coverage of Harmonized Landsat and Sentinel-2 (HLS) V2.0
abstract
Global observations at 30-m ground sampling distance (GSD) are now possible at a cadence of one-three days by combining Landsat 8 and 9 with Sentinel-2A and -2B satellites. Previous studies characterizing pixel-level Landsat-class measurement frequency used data from different sources but offered little information on observation availability after rigorous quality screening. This study examined the coverage frequency of Harmonized Landsat and Sentinel-2 (HLS) V2.0 data for 2022, the first year all four satellites data were available. These data have had quality control filtering and harmonization, and therefore reflect the spatial-temporal distribution of usable observations. On average, HLS data provide observations every 1.6 days at the global scale, and 2.2 days in the data-scarce tropical regions, regardless of cloud cover. The global mean and median cloud-free observations were 69 and 64, respectively. The frequency of good-quality observations varies geographically and seasonally due to changes in satellite swath overlap, cloud frequency, and solar illumination. High latitudes ($\gt \sim 75^{\circ }$N) exhibit the highest number of cloud-free observations between March and September. However, data are unavailable during winter months due to low solar elevation angles and boreal regions have a lower number of clear observations in the summer months. The tropical regions have the lowest number of clear observations. More frequent HLS observations could improve terrestrial monitoring. We mapped the monthly and weekly number of clear observations globally to show where HLS data could support monthly or subweekly time series applications.
Christopher S. R. Neigh, Junchang Ju, Philip W. Dabney, Bruce D. Cook, Zhe Zhu, Christopher J. Crawford, Ferran Gascon, Peter Strobl, Madhu Sridhar
IEEE Geosci. Remote. Sens. Lett.6
2024 An Encoder-Decoder Based Approach for ECG Delineation
Xiangju Kong, Shuwang Zhou, Tianlei Gao, Zhe Zhu
ICONIP (4)5
2024 Torque based Structured Pruning for Deep Neural Network
abstract
Structured pruning is a popular way of convolutional neural network (CNN) acceleration. However, current state of the art pruning techniques require modifications to the network architecture, implementation of complex gradient update rules or repetitive training and long fine-tuning stages. Our novel physics-inspired approach for structured pruning aims to solve these issues. Analogous to ‘Torque’ we apply a force that consolidates the weights of a convolutional layer around a selected pivot point during training. Using the distance-dependency nature of torque, we can encourage high density of weights in filters around this point and increase filter sparsity as we move away. Filters away from the pivot point can be pruned, resulting in a minimum loss of information. We can control the tightness of the weights by varying the hyper-parameters, thus assisting us in creating a more compact network. Our proposed technique is jointly able to perform both filter learning and filter importance sorting. Additionally, our method is easy to implement, requires no change to model architecture and needs very little to no fine-tuning. We show that our approach reaches competitive results with previous state-of-the-art by evaluating popular networks such as VGGNet and ResNet on multiple image classification tasks. Notably, our method can reduce the parameter count of VGGNet by 96% and still maintain the accuracy achieved by the full-size model without any fine-tuning. This makes our method both latency and memory efficient for hardware deployment.
Arshita Gupta, Tien C. Bau, Joonsoo Kim, Zhe Zhu, Hrishikesh Garud
WACV4
2024 Beyond a technological tool: How a platform as a meta-organization enables a successful public service project in China
Zhe Zhu, Nan Zhang 0025, Felix T. C. Tan
Inf. Manag.1
2024 Semantic Loopback Detection Method Based on Instance Segmentation and Visual SLAM in Autonomous Driving
abstract
Autonomous driving has gradually become a research hotspot in recent years, but the robustness of loopback detection in complex environments such as dynamic and weak textures needs to be improved. A semantic loopback detection method is proposed based on instance segmentation and visual SLAM to make sufficient use of semantic information in autonomous driving. The proposed method combines image segmentation and visual SLAM (Simultaneous Localization and Mapping) to construct a semantic SLAM system. What’s more, a data association method that combines semantic and geometric information is proposed to improve the traditional loopback detection method by using semantic information to increase the accuracy of loopback detection. The result of experiment on the TUM public dataset shows that the loopback detection accuracy of the improved loopback detection method is higher than that of the bag-of-words method in all four datasets, and our proposed algorithm can effectively improve the accuracy of loopback detection of the SLAM system in general.
Zhe Zhu, Juntong Yun, Manman Xu, Ying Liu 0087, Ying Sun 0004, Fazeng Li
IEEE Trans. Intell. Transp. Syst.2
2024 CSDN: Cross-Modal Shape-Transfer Dual-Refinement Network for Point Cloud Completion
abstract
How will you repair a physical object with some missings? You may imagine its original shape from previously captured images, recover its overall (global) but coarse shape first, and then refine its local details. We are motivated to imitate the physical repair procedure to address point cloud completion. To this end, we propose a cross-modal shape-transfer dual-refinement network (termed CSDN), a coarse-to-fine paradigm with images of full-cycle participation, for quality point cloud completion. CSDN mainly consists of "shape fusion" and "dual-refinement" modules to tackle the cross-modal challenge. The first module transfers the intrinsic shape characteristics from single images to guide the geometry generation of the missing regions of point clouds, in which we propose IPAdaIN to embed the global features of both the image and the partial point cloud into completion. The second module refines the coarse output by adjusting the positions of the generated points, where the local refinement unit exploits the geometric relation between the novel and the input points by graph convolution, and the global constraint unit utilizes the input image to fine-tune the generated offset. Different from most existing approaches, CSDN not only explores the complementary information from images but also effectively exploits cross-modal data in the whole coarse-to-fine completion procedure. Experimental results indicate that CSDN performs favorably against twelve competitors on the cross-modal benchmark.
Zhe Zhu, Liangliang Nan, Haoran Xie 0001, Honghua Chen, Jun Wang 0039, Mingqiang Wei, Harry Qin
IEEE Trans. Vis. Comput. Graph.1
2023 SVDFormer: Complementing Point Cloud via Self-view Augmentation and Self-structure Dual-generator
abstract
In this paper, we propose a novel network, SVDFormer, to tackle two specific challenges in point cloud completion: understanding faithful global shapes from incomplete point clouds and generating high-accuracy local structures. Current methods either perceive shape patterns using only 3D coordinates or import extra images with well-calibrated intrinsic parameters to guide the geometry estimation of the missing parts. However, these approaches do not always fully leverage the cross-modal self-structures available for accurate and high-quality point cloud completion. To this end, we first design a Self-view Fusion Network that leverages multiple-view depth image information to observe incomplete self-shape and generate a compact global shape. To reveal highly detailed structures, we then introduce a refinement module, called Self-structure Dual-generator, in which we incorporate learned shape priors and geometric self-similarities for producing new points. By perceiving the incompleteness of each point, the dual-path design disentangles refinement strategies conditioned on the structural type of each point. SVDFormer absorbs the wisdom of self-structures, avoiding any additional paired information such as color images with precisely calibrated camera intrinsic parameters. Comprehensive experiments indicate that our method achieves state-of-the-art performance on widely-used benchmarks. Code is available at https://github.com/czvvd/SVDFormer.
Zhe Zhu, Honghua Chen, Weiming Wang 0002, Harry Qin, Mingqiang Wei
ICCV1
2023 Model-guided 3D stitching for augmented virtual environment
Zhong Zhou, Zhe Zhu, Jingdi You
Sci. China Inf. Sci.4
2023 Sphere Face Model: A 3D morphable model with hypersphere manifold latent space using joint 2D/3D training
abstract
3D morphable models (3DMMs) are generative models for face shape and appearance. Recent works impose face recognition constraints on 3DMM shape parameters so that the face shapes of the same person remain consistent. However, the shape parameters of traditional 3DMMs satisfy the multivariate Gaussian distribution. In contrast, the identity embeddings meet the hypersphere distribution, and this conflict makes it challenging for face reconstruction models to preserve the faithfulness and the shape consistency simultaneously. In other words, recognition loss and reconstruction loss can not decrease jointly due to their conflict distribution. To address this issue, we propose the Sphere Face Model (SFM), a novel 3DMM for monocular face reconstruction, preserving both shape fidelity and identity consistency. The core of our SFM is the basis matrix which can be used to reconstruct 3D face shapes, and the basic matrix is learned by adopting a two-stage training approach where 3D and 2D training data are used in the first and second stages, respectively. We design a novel loss to resolve the distribution mismatch, enforcing that the shape parameters have the hyperspherical distribution. Our model accepts 2D and 3D data for constructing the sphere face models. Extensive experiments show that SFM has high representation ability and clustering performance in its shape parameter space. Moreover, it produces high-fidelity face shapes consistently in challenging conditions in monocular face reconstruction. The code will be released at https://github.com/a686432/SIR
Diqiong Jiang, Yiwei Jin, Zhe Zhu, Yun Zhang 0024, Ruofeng Tong 0001, Min Tang 0001
Comput. Vis. Media4
2023 AGConv: Adaptive Graph Convolution on 3D Point Clouds
abstract
Convolution on 3D point clouds is widely researched yet far from perfect in geometric deep learning. The traditional wisdom of convolution characterises feature correspondences indistinguishably among 3D points, arising an intrinsic limitation of poor distinctive feature learning. In this article, we propose Adaptive Graph Convolution (AGConv) for wide applications of point cloud analysis. AGConv generates adaptive kernels for points according to their dynamically learned features. Compared with the solution of using fixed/isotropic kernels, AGConv improves the flexibility of point cloud convolutions, effectively and precisely capturing the diverse relations between points from different semantic parts. Unlike the popular attentional weight schemes, AGConv implements the adaptiveness inside the convolution operation instead of simply assigning different weights to the neighboring points. Extensive evaluations clearly show that our method outperforms state-of-the-arts of point cloud classification and segmentation on various benchmark datasets. Meanwhile, AGConv can flexibly serve more point cloud analysis approaches to boost their performance. To validate its flexibility and effectiveness, we explore AGConv-based paradigms of completion, denoising, upsampling, registration and circle extraction, which are comparable or even superior to their competitors.
Mingqiang Wei, Zeyong Wei, Huajian Si, Zhilei Chen, Zhe Zhu, Jingbo Qiu, Xuefeng Yan 0001, Yanwen Guo 0001, Jun Wang 0039, Harry Qin
IEEE Trans. Pattern Anal. Mach. Intell.7
2023 A variational approach for feature-aware B-spline curve design on surface meshes
Rongyan Xu, Huaxiong Zhang, Yun Zhang 0024, Yukun Lai, Zhe Zhu
Vis. Comput.6
2022 SPCNet: Stepwise Point Cloud Completion Network
abstract
Abstract How will you repair a physical object with large missings? You may first recover its global yet coarse shape and stepwise increase its local details. We are motivated to imitate the above physical repair procedure to address the point cloud completion task. We propose a novel stepwise point cloud completion network (SPCNet) for various 3D models with large missings. SPCNet has a hierarchical bottom‐to‐up network architecture. It fulfills shape completion in an iterative manner, which 1) first infers the global feature of the coarse result; 2) then infers the local feature with the aid of global feature; and 3) finally infers the detailed result with the help of local feature and coarse result. Beyond the wisdom of simulating the physical repair, we newly design a cycle loss to enhance the generalization and robustness of SPCNet. Extensive experiments clearly show the superiority of our SPCNet over the state‐of‐the‐art methods on 3D point clouds with large missings. Code is available at https://github.com/1127368546/SPCNet .
Honghua Chen, Xuequan Lu, Zhe Zhu, Jun Wang 0039, Weiming Wang 0002, Fu Lee Wang, Mingqiang Wei
Comput. Graph. Forum4
2022 Towards natural object-based image recoloring
abstract
Existing color editing algorithms enable users to edit the colors in an image according to their own aesthetics. Unlike artists who have an accurate grasp of color, ordinary users are inexperienced in color selection and matching, and allowing non-professional users to edit colors arbitrarily may lead to unrealistic editing results. To address this issue, we introduce a palette-based approach for realistic object-level image recoloring. Our data-driven approach consists of an offline learning part that learns the color distributions for different objects in the real world, and an online recoloring part that first recognizes the object category, and then recommends appropriate realistic candidate colors learned in the offline step for that category. We also provide an intuitive user interface for efficient color manipulation. After color selection, image matting is performed to ensure smoothness of the object boundary. Comprehensive evaluation on various color editing examples demonstrates that our approach outperforms existing state-of-the-art color editing algorithms.
Mengyao Cui 0001, Zhe Zhu, Yulu Yang, Shao-Ping Lu
Comput. Vis. Media2
2022 3D Pyramid Pooling Network for Abdominal MRI Series Classification
abstract
Recognizing and organizing different series in an MRI examination is important both for clinical review and research, but it is poorly addressed by the current generation of picture archiving and communication systems (PACSs) and post-processing workstations. In this paper, we study the problem of using deep convolutional neural networks for automatic classification of abdominal MRI series to one of many series types. Our contributions are three-fold. First, we created a large abdominal MRI dataset containing 3717 MRI series including 188,665 individual images, derived from liver examinations. 30 different series types are represented in this dataset. The dataset was annotated by consensus readings from two radiologists. Both the MRIs and the annotations were made publicly available. Second, we proposed a 3D pyramid pooling network, which can elegantly handle abdominal MRI series with varied sizes of each dimension, and achieved state-of-the-art classification performance. Third, we performed the first ever comparison between the algorithm and the radiologists on an additional dataset and had several meaningful findings.
Zhe Zhu, Amber Mittendorf, Erin Shropshire, Brian C. Allen, Chad M. Miller, Mustafa R. Bashir, Maciej A. Mazurowski
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 DSNet: Deep Shadow Network for Illumination Estimation
abstract
Illumination consistency has applications to modeling and rendering in virtual reality. In 3D reconstruction and Mixed Reality(MR) fusion, the appearance of a large-scale outdoor scene may change in response to lighting and seasons, for example. Since 3D reconstruction from scratch is costly, it is helpful to be able to update existing models with recently captured photographs. However, the illumination conditions of the captured photograph can be arbitrary, making it challenging to fit to the existing model. To tackle this problem, this paper proposes a novel approach that can precisely estimate the illumination of the input image. Our Deep Shadow Network (DSNet) collaboratively utilizes illumination-based data augmentation for sun position estimation, along with a dataset of illumination-based augmented renderings. Our run-time rendering and optimization strategy is also discussed. We show that accurate simulation of illumination can improve the performance of visual applications including place recognition and long-term localization. Experimental results validate the effectiveness of the proposed approach, and show its superiority over the state-of-the-art.
Yuan Xiong, Hongrui Chen, Zhe Zhu, Zhong Zhou
VR4
2021 Efficient propagation of sparse edits on 360∘ panoramas
Yun Zhang 0024, Yukun Lai, Zhe Zhu
Comput. Graph.4
2020 A Metric for Video Blending Quality Assessment
abstract
We propose an objective approach to assess the quality of video blending. Blending is a fundamental operation in video editing, which can smooth the intensity changes of relevant regions. However blending also generates artefacts such as bleeding and ghosting. To assess the quality of the blended videos, our approach considers the illuminance consistency as a positive aspect while regard the artefacts as a negative aspect. Temporal coherence between frames is also considered. We evaluate our metric on a video blending dataset where the results of subjective evaluation are available. Experimental results validate the effectiveness of our proposed metric, and shows that this metric gives superior performance over existing video quality metrics.
Zhe Zhu, Hantao Liu, Jiaming Lu, Shi-Min Hu 0001
IEEE Trans. Image Process.1
2019 Mask Embedding for Realistic High-Resolution Medical Image Synthesis
Yinhao Ren, Zhe Zhu, Yingzhou Li, Dehan Kong, Rui Hou 0002, Lars J. Grimm, Jeffrey R. Marks, Joseph Y. Lo
MICCAI (6)2
2019 A three-stage real-time detector for traffic signs in large panoramas
abstract
Traffic sign detection is one of the key components in autonomous driving. Advanced autonomous vehicles armed with high quality sensors capture high definition images for further analysis. Detecting traffic signs, moving vehicles, and lanes is important for localization and decision making. Traffic signs, especially those that are far from the camera, are small, and so are challenging to traditional object detection methods. In this work, in order to reduce computational cost and improve detection performance, we split the large input images into small blocks and then recognize traffic signs in the blocks using another detection module. Therefore, this paper proposes a three-stage traffic sign detector, which connects a BlockNet with an RPN–RCNN detection network. BlockNet, which is composed of a set of CNN layers, is capable of performing block-level foreground detection, making inferences in less than 1 ms. Then, the RPN–RCNN two-stage detector is used to identify traffic sign objects in each block; it is trained on a derived dataset named TT100KPatch. Experiments show that our framework can achieve both state-of-the-art accuracy and recall; its fastest detection speed is 102 fps.
Ruochen Fan, Sharon X. Huang, Zhe Zhu, Ruofeng Tong 0001
Comput. Vis. Media4
2019 A Large Chinese Text Dataset in the Wild
Tailing Yuan, Zhe Zhu, Kun Xu 0003, Cheng-Jun Li, Tai-Jiang Mu, Shi-Min Hu 0001
J. Comput. Sci. Technol.2
2019 Hierarchical Convolutional Neural Networks for Segmentation of Breast Tumors in MRI With Application to Radiogenomics
abstract
Breast tumor segmentation based on dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) is a challenging problem and an active area of research. Particular challenges, similarly as in other segmentation problems, include the class-imbalance problem as well as confounding background in DCE-MR images. To address these issues, we propose a mask-guided hierarchical learning (MHL) framework for breast tumor segmentation via fully convolutional networks (FCN). Specifically, we first develop an FCN model to generate a 3D breast mask as the region of interest (ROI) for each image, to remove confounding information from input DCE-MR images. We then design a two-stage FCN model to perform coarse-to-fine segmentation for breast tumors. Particularly, we propose a Dice-Sensitivity-like loss function and a reinforcement sampling strategy to handle the class-imbalance problem. To precisely identify locations of tumors that underwent a biopsy, we further propose an FCN model to detect two landmarks located at two nipples. We finally selected the biopsied tumor based on both identified landmarks and segmentations. We validate our MHL method on 272 patients, achieving a mean Dice similarity coefficient (DSC) of 0.72 which is comparable to mutual DSC between expert radiologists. Using the segmented biopsied tumors, we also demonstrate that the automatically generated masks can be applied to radiogenomics and can identify luminal A subtype from other molecular subtypes with the similar accuracy with the analysis based on semi-manual tumor segmentation.
Jun Zhang 0018, Ashirbani Saha, Zhe Zhu, Maciej A. Mazurowski
IEEE Trans. Medical Imaging3
2018 A Comparative Study of Algorithms for Realtime Panoramic Video Blending
abstract
Unlike image blending algorithms, video blending algorithms have been little studied. In this paper, we investigate 6 popular blending algorithms-feather blending, multi-band blending, modified Poisson blending, mean value coordinate blending, multi-spline blending and convolution pyramid blending. We consider their application to blending realtime panoramic videos, a key problem in various virtual reality tasks. To evaluate the performances and suitabilities of the 6 algorithms for this problem, we have created a video benchmark with several videos captured under various conditions. We analyze the time and memory needed by the above 6 algorithms, for both CPU and GPU implementations (where readily parallelizable). The visual quality provided by these algorithms is also evaluated both objectively and subjectively. The video benchmark and algorithm implementations are publicly available1.
Zhe Zhu, Jiaming Lu, Minxuan Wang, Song-Hai Zhang, Ralph R. Martin, Hantao Liu, Shi-Min Hu 0001
IEEE Trans. Image Process.1
2018 Computational Design of Transforming Pop-up Books
abstract
We present the first computational tool to help ordinary users create transforming pop-up books. In each transforming pop-up, when the user pulls a tab, an initial flat two-dimensional (2D) pattern, i.e., a 2D shape with a superimposed picture, such as an airplane, turns into a new 2D pattern, such as a robot. Given the two 2D patterns, our approach automatically computes a 3D pop-up mechanism that transforms one pattern into the other; it also outputs a design blueprint, allowing the user to easily make the final model. We also present a theoretical analysis of basic transformation mechanisms; combining these basic mechanisms allows more flexibility of final designs. Using our approach, inexperienced users can create models in a short time; previously, even experienced artists often took weeks to manually create them. We demonstrate our method on a variety of real-world examples.
Zhe Zhu, Ralph R. Martin, Kun Xu 0003, Jiaming Lu, Shi-Min Hu 0001
ACM Trans. Graph.2
2017 Avoiding bleeding in image blending
abstract
Though elegant in mathematical formulation, gradient-domain image blending suffers from bleeding artefacts in real world applications. We propose an image blending algorithm that avoids bleeding artefacts while preserving the good properties of gradient-domain blending, such as smooth transitions between the candidate regions. Our key idea to finesse the non-smooth boundary difference calculation that causes bleeding artefacts is to use local patch differences. While most previous gradient-domain blending algorithms change one region to fit the other, to further reduce bleeding, we perform bidirectional blending so that both regions change simultaneously. Our blending algorithm is fast: when applied to image stitching, it can achieve 20 fps at 4K resolution. Source code and test images are publicly available.
Minxuan Wang, Zhe Zhu, Song-Hai Zhang, Ralph R. Martin, Shi-Min Hu 0001
ICIP2
2017 Converter side phase-to-ground fault protection of full-bridge modular multilevel converter-based bipolar HVDC
abstract
This study proposed a novel protective action for converter side phase-to-ground fault of FB-MMC-based bipolar HVDC. The faulty converter is kept in operation during fault, and its dc-link voltage reference is set to zero in order to avoid serious overvoltage of the sub-module capacitors. This protective action was verified by simulation. The overcurrent and overvoltage of converter arms are moderate when adopting this protective action.
Wenbo Yang 0004, Qiang Song 0002, Hong Rao, Shukai Xu, Zhe Zhu
IECON5
2017 An Optimization Approach for Localization Refinement of Candidate Traffic Signs
abstract
We propose a localization refinement approach for candidate traffic signs. Previous traffic sign localization approaches, which place a bounding rectangle around the sign, do not always give a compact bounding box, making the subsequent classification task more difficult. We formulate localization as a segmentation problem, and incorporate prior knowledge concerning color and shape of traffic signs. To evaluate the effectiveness of our approach, we use it as an intermediate step between a standard traffic sign localizer and a classifier. Our experiments use the well-known German Traffic Sign Detection Benchmark (GTSDB) as well as our new Chinese Traffic Sign Detection Benchmark. This newly created benchmark is publicly available,1and goes beyond previous benchmark data sets: it has over 5000 high-resolution images containing more than 14 000 traffic signs taken in realistic driving conditions. Experimental results show that our localization approach significantly improves bounding boxes when compared with a standard localizer, thereby allowing a standard traffic sign classifier to generate more accurate classification results.1http://cg.cs.tsinghua.edu.cn/ctsdb/.
Zhe Zhu, Jiaming Lu, Ralph R. Martin, Shi-Min Hu 0001
IEEE Trans. Intell. Transp. Syst.1
2016 Traffic-Sign Detection and Classification in the Wild
abstract
Although promising results have been achieved in the areas of traffic-sign detection and classification, few works have provided simultaneous solutions to these two tasks for realistic real world images. We make two contributions to this problem. Firstly, we have created a large traffic-sign benchmark from 100000 Tencent Street View panoramas, going beyond previous benchmarks. It provides 100000 images containing 30000 traffic-sign instances. These images cover large variations in illuminance and weather conditions. Each traffic-sign in the benchmark is annotated with a class label, its bounding box and pixel mask. We call this benchmark Tsinghua-Tencent 100K. Secondly, we demonstrate how a robust end-to-end convolutional neural network (CNN) can simultaneously detect and classify trafficsigns. Most previous CNN image processing solutions target objects that occupy a large proportion of an image, and such networks do not work well for target objects occupying only a small fraction of an image like the traffic-signs here. Experimental results show the robustness of our network and its superiority to alternatives. The benchmark, source code and the CNN model introduced in this paper is publicly available1.
Zhe Zhu, Dun Liang, Song-Hai Zhang, Sharon X. Huang, Baoli Li 0004, Shi-Min Hu 0001
CVPR1
2016 High efficient modeling of a diode clamped Modular Multilevel Converter for EMT simulation
abstract
Conventional Electromagnetic Transient (EMT) simulations of a Modular Multilevel Converter based HVDC (MMC-HVDC) is quite time consuming. Therefore, the sub modules are often modeled with equivalent transformation methods. In this paper, two equivalent models with much higher efficiency have been presented for a specific MMC topology, i.e. the diode clamped half bridge topology, which convenient the researches on MMC-HVDC with overhead lines. Details of the modeling technologies have been explained. Simulations have been carried out, focusing on computational efficiency, steady and transient performances. Compared to conventional EMT model, the proposed models are accurate enough, while with much higher efficiency.
Wenming Gong, Zhe Zhu, Shukai Xu, Hong Rao
IECON2
2016 3D modeling and motion parallax for improved videoconferencing
abstract
We consider a face-to-face videoconferencing system that uses a Kinect camera at each end of the link for 3D modeling and an ordinary 2D display for output. The Kinect camera allows a 3D model of each participant to be transmitted; the (assumed static) background is sent separately. Furthermore, the Kinect tracks the receiver’s head, allowing our system to render a view of the sender depending on the receiver’s viewpoint. The resulting motion parallax gives the receivers a strong impression of 3D viewing as they move, yet the system only needs an ordinary 2D display. This is cheaper than a full 3D system, and avoids disadvantages such as the need to wear shutter glasses, VR headsets, or to sit in a particular position required by an autostereo display. Perceptual studies show that users experience a greater sensation of depth with our system compared to a typical 2D videoconferencing system.
Zhe Zhu, Ralph R. Martin, Robert Pepperell, Alistair Burleigh
Comput. Vis. Media1
2016 Faithful Completion of Images of Scenic Landmarks Using Internet Images
abstract
Previous works on image completion typically aim to produce visually plausible results rather than factually correct ones. In this paper, we propose an approach to faithfully complete the missing regions of an image. We assume that the input image is taken at a well-known landmark, so similar images taken at the same location can be easily found on the Internet. We first download thousands of images from the Internet using a text label provided by the user. Next, we apply two-step filtering to reduce them to a small set of candidate images for use as source images for completion. For each candidate image, a co-matching algorithm is used to find correspondences of both points and lines between the candidate image and the input image. These are used to find an optimal warp relating the two images. A completion result is obtained by blending the warped candidate image into the missing region of the input image. The completion results are ranked according to combination score, which considers both warping and blending energy, and the highest ranked ones are shown to the user. Experiments and results demonstrate that our method can faithfully complete images.
Zhe Zhu, Hao-Zhi Huang 0001, Kun Xu 0003, Shi-Min Hu 0001
IEEE Trans. Vis. Comput. Graph.1
2015 The Goalkeeper Strategy of RoboCup MSL Based on Dual Image Source
abstract
Gatekeeping is one of the most important tasks in MSL, especially defending aerial shots. This paper proposes a goalkeeper strategy based on combining two image sources. The goalkeeper uses the Omni-vision system to adjust heading orientation of the Kinect sensor for capturing both depth and RGB images of the ball. The depth information is used to detect the spatial position of the ball and the RGB image ensures the validity of the recognition. Once opponents shoot, the proposed system uses the first few detected positions of the ball to predict the interception point by the goalkeeper based on the least square method. The accuracy and effectiveness of the proposed strategy are verified by experimentation.
Peng Zhao 0015, Shiyu Mou, Jieming Zhou, Dongbiao Sun, Zhe Zhu, Zongyi Zhang, Ye Lv, Wanjie Zhang, Binbin Li 0019, Abudouyimujiang Kader, Yiheng Su
RoboCup8
2015 Panorama completion for street views
abstract
This paper considers panorama images used for street views. Their viewing angle of 360° causes pixels at the top and bottom to appear stretched and warped. Although current image completion algorithms work well, they cannot be directly used in the presence of such distortions found in panoramas of street views. We thus propose a novel approach to complete such 360° panoramas using optimization-based projection to deal with distortions. Experimental results show that our approach is efficient and provides an improvement over standard image completion algorithms.
Zhe Zhu, Ralph R. Martin, Shi-Min Hu 0001
Comput. Vis. Media1
2014 Testing a complete control and protection system for multi-terminal MMC HVDC links using hardware-in-the-loop simulation
abstract
This paper presents the dynamic performance test of a complete control and protection system for a Multi-terminal MMC HVDC system using hardware-in-the-loop (HIL) simulation. A novel HIL bench with a cost-effective input and output interface between the control system under test and the real-time simulator is introduced. Two critical test cases, namely the start-up of the MMC connected to islanded networks and the AC fault in the bus close to the MMC substation are studied. The validity of the proposed methodology for the dynamic performance test is confirmed by comparing the results from the HIL test and the actual waveform recorded from the field, after the MMC is commissioned.
Zhe Zhu, Hong Rao
IECON1
2014 Load-Balanced Breadth-First Search on GPUs
Zhe Zhu, Jianjun Li 0010, Guohui Li 0001
WAIM1
2014 Efficient decomposition of strongly connected components on GPUs
Guohui Li 0001, Zhe Zhu, Cong Zhang 0007, Fumin Yang
J. Syst. Archit.2
2013 Team Water: The Champion of the RoboCup Middle Size League Competition 2013
Zhe Zhu, Charles Lynch, Ye Lv, Wanjie Zhang, Peng Zhao 0015, Feichan Yang, Binbin Li 0019, Jieming Zhou, Zongyi Zhang
RoboCup2
2013 What's Your Next Move: User Activity Prediction in Location-based Social Networks
abstract
Location-based social networks have been gaining increasing popularity in recent years. To increase users’ engagement with location-based services, it is important to provide attractive features, one of which is geo-targeted ads and coupons. To make ads and coupon delivery more effective, it is essential to predict the location that is most likely to be visited by a user at the next step. However, an inherent challenge in location prediction is a huge prediction space, with millions of distinct check-in locations as prediction target. In this paper we exploit the check-in category information to model the underlying user movement pattern. We propose a framework which uses a mixed hidden Markov model to predict the category of user activity at the next step and then predict the most likely location given the estimated category distribution. The advantages of modeling the category level include a significantly reduced prediction space and a precise expression of the semantic meaning of user activities. Extensive experimental results show that, with the predicted category distribution, the number of location candidates for prediction is 5.45 times smaller, while the prediction accuracy is 13.21% higher.
Hong Cheng 0001, Jihang Ye, Zhe Zhu
SDM3
2013 Predicting positive and negative links in signed social networks by transfer learning
abstract
Different from a large body of research on social networks that has focused almost exclusively on positive relationships, we study signed social networks with both positive and negative links. Specifically, we focus on how to reliably and effectively predict the signs of links in a newly formed signed social network (called a target network). Since usually only a very small amount of edge sign information is available in such newly formed networks, this small quantity is not adequate to train a good classifier. To address this challenge, we need assistance from an existing, mature signed network (called a source network) which has abundant edge sign information. We adopt the transfer learning approach to leverage the edge sign information from the source network, which may have a different yet related joint distribution of the edge instances and their class labels.
Jihang Ye, Hong Cheng 0001, Zhe Zhu, Minghua Chen 0001
WWW3
2013 3-Sweep: extracting editable objects from a single photo
abstract
We introduce an interactive technique for manipulating simple 3D shapes based on extracting them from a single photograph. Such extraction requires understanding of the components of the shape, their projections, and relations. These simple cognitive tasks for humans are particularly difficult for automatic algorithms. Thus, our approach combines the cognitive abilities of humans with the computational accuracy of the machine to solve this problem. Our technique provides the user the means to quickly create editable 3D parts---human assistance implicitly segments a complex object into its components, and positions them in space. In our interface, three strokes are used to generate a 3D component that snaps to the shape's outline in the photograph, where each stroke defines one dimension of the component. The computer reshapes the component to fit the image of the object in the photograph as well as to satisfy various inferred geometric constraints imposed by its global 3D structure. We show that with this intelligent interactive modeling tool, the daunting task of object extraction is made simple. Once the 3D object has been extracted, it can be quickly edited and placed back into photos or 3D scenes, permitting object-driven photo editing tasks which are impossible to perform in image-space. We show several examples and present a user study illustrating the usefulness of our technique.
Tao Chen 0015, Zhe Zhu, Ariel Shamir, Shi-Min Hu 0001, Daniel Cohen-Or
ACM Trans. Graph.2