EDBT 2026 Demo / reviewers in the wild / expert
Xiuming Zhang
dblp:147/2810
· DBLP profile ↗
28ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0002-4326-727XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 2 first-author · 15 since 2021Artificial intelligence and machine learning · 17 · 2 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-label Learning for Reliable Cervical Cytology Screening
Linyun Zhou, Jian Yang 0003, Xiuming Zhang, Zunlei Feng, Bingde Hu |
ICIC (29) | 5 |
| 2026 | A hypergraph-based model for tumor prognosis using local and global information fusion on H&E-stained histology images
Yanfen Cui, Zhenhui Li, Xiuming Zhang, Su Yao, Dacheng Yang, Zhishun Liu, Shiwei Luo, Guangjun Yang, Lixu Yan, Xiangtian Zhao, Yingqiu Huo, Jiahui Ma, Wenfeng He, Tao Tan 0002, Anant Madabhushi, Jinglei Tang, Zaiyi Liu, Cheng Lu 0001 |
Medical Image Anal. | 5 |
| 2026 | An Intelligent Interactive Visual Analytics System for Exploring Large and Multi-Scale Pathology ImagesabstractPathology images are crucial for cancer diagnosis and treatment. Although artificial intelligence has driven rapid advancements in pathology image analysis, the interpretation of ultra-large and multi-scale pathology images in clinical practice still heavily relies on physicians' experience. Clinicians need to repeatedly zoom in and out on individual slides to compare and assess pathological details - a process that is both time-consuming and prone to visual fatigue. The system first employs a diffusion model to perform tissue segmentation on pathology images, then calculates pathological tissue proportions and morphological metrics. Finally, through multi-scale dynamic comparison and multi-level visual evaluation, the system facilitates comprehensive and precise analysis of pathology images. The system provides clinicians with an intelligent and interactive tool for pathology image interpretation, enabling efficient visualization and precise analysis of pathological details, thereby reducing the effort require for detailed analysis. Chaoqing Xu, Xinyuan Fu, Liting Fang, Zunlei Feng, Xiuming Zhang, Can Wang 0001, Mingli Song, Wei Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2025 | Large Vision-Language Models are Generalist Solvers For Pathology TasksabstractLeveraging the powerful capabilities of large language models (LLMs), large vision-language models (LVLMs) can perform a wide variety of tasks based on input images and user instructions. However, existing pathology-focused LVLMs are limited to relatively simple tasks, such as image captioning, visual question answering, and generating brief pathology reports, which restricts their clinical applicability. To enhance the practicality of pathology LVLMs and explore their performance boundaries across pathology tasks, we curated a multi-task pathology instruction-following dataset that better aligns with clinical needs. This dataset encompasses tasks such as cancer classification and grading, molecular subtype identification, and the detection of structures like nuclei, blood vessels, nerves, and lymph nodes. Extensive experiments were conducted on this dataset to identify key factors influencing the performance of LVLMs on these pathology tasks, and optimal solutions were proposed. Our findings provide valuable insights to advance the clinical application of large vision-language models in pathology. Shengxuming Zhang, Hengrui Lou, Jing Zhang 0120, Xiuming Zhang, Mingli Song, Zunlei Feng |
ICIP | 4 |
| 2025 | L-Diffusion: Laplace Diffusion for Efficient Pathology Image SegmentationabstractPathology image segmentation plays a pivotal role in artificial digital pathology diagnosis and treatment. Existing approaches to pathology image segmentation are hindered by labor-intensive annotation processes and limited accuracy in tail-class identification, primarily due to the long-tail distribution inherent in gigapixel pathology images. In this work, we introduce the Laplace Diffusion Model, referred to as L-Diffusion, an innovative framework tailored for efficient pathology image segmentation. L-Diffusion utilizes multiple Laplace distributions, as opposed to Gaussian distributions, to model distinct components—a methodology supported by theoretical analysis that significantly enhances the decomposition of features within the feature space. A sequence of feature maps is initially generated through a series of diffusion steps. Following this, contrastive learning is employed to refine the pixel-wise vectors derived from the feature map sequence. By utilizing these highly discriminative pixel-wise vectors, the segmentation module achieves a harmonious balance of precision and robustness with remarkable efficiency. Extensive experimental evaluations demonstrate that L-Diffusion attains improvements of up to 7.16%, 26.74%, 16.52%, and 3.55% on tissue segmentation datasets, and 20.09%, 10.67%, 14.42%, and 10.41% on cell segmentation datasets, as quantified by DICE, MPA, mIoU, and FwIoU metrics. The source are available at https://github.com/Lweihan/LDiffusion. Linyun Zhou, Yang Jian, Shengxuming Zhang, Xiangtong Du, Xiuming Zhang, Jing Zhang 0120, Chaoqing Xu, Mingli Song, Zunlei Feng |
ICML | 6 |
| 2025 | Self-calibration Enhanced Whole Slide Pathology Image AnalysisabstractPathology images are considered the ``gold standard" for cancer diagnosis and treatment, with gigapixel images providing extensive tissue and cellular information. Existing methods fail to simultaneously extract global structural and local detail features for comprehensive pathology image analysis efficiently. To address these limitations, we propose a self-calibration enhanced framework for whole slide pathology image analysis, comprising three components: a global branch, a focus predictor, and a detailed branch. The global branch initially classifies using the pathological thumbnail, while the focus predictor identifies relevant regions for classification based on the last layer features of the global branch. The detailed extraction branch then assesses whether the magnified regions correspond to the lesion area. Finally, a feature consistency constraint between the global and detail branches ensures that the global branch focuses on the appropriate region and extracts sufficient discriminative features for final identification. These focused discriminative features can facilitate the discovery of novel prognostic tumor markers, from the perspective of feature uniqueness and tissue spatial distribution. Extensive experiment results demonstrate that the proposed framework can rapidly deliver accurate and explainable results for pathological grading and prognosis tasks. Haoming Luo, Xiaotian Yu, Shengxuming Zhang, Jiabin Xia, Jian Yang 0003, Yuning Sun, Xiuming Zhang, Jing Zhang 0120, Zunlei Feng |
IJCAI | 7 |
| 2025 | DenseSAM: Semantic Enhance SAM for Efficient Dense Object SegmentationabstractDense object segmentation is essential for various applications, particularly in pathology image and remote sensing image analysis. However, distinguishing numerous similar and densely packed objects in this task presents significant challenges. Several methods, including CNN- and ViT-based approaches, have been proposed to tackle these issues. Yet, models trained on limited datasets exhibit limited generalization ability. The Segment Anything Model (SAM) has recently achieved significant progress in zero-shot segmentation but relies heavily on precise positional guidance. However, providing numerous accurate location prompts in dense scenarios is time-consuming. To overcome this limitation, we conducted an in-depth exploration of the SAM mechanism and found that its strong generalization ability stems from the encoder’s edge detection capability, which is semantically independent, making location prompts essential for segmentation. This insight inspired the development of DenseSAM, which replaces location prompts with semantic guidance for automatic segmentation in dense scenarios. Specifically, it uses local details to weaken the edges of background objects, leverages global context to enhance intra-class feature similarity, while further increasing contrast with the background, and integrates a dual-head decoding process to enable lightweight automatic semantic segmentation. Extensive experiments on pathology images demonstrate that DenseSAM delivers remarkable performance with minimal training parameters, providing a cost-effective and efficient solution. Moreover, experiments on remote sensing images further validate its excellent scalability, making DenseSAM suitable for various dense object segmentation domains. The code is available at https://github.com/imAzhou/DenseSAM. Linyun Zhou, Jiacong Hu, Shengxuming Zhang, Xiangtong Du, Mingli Song, Xiuming Zhang, Zunlei Feng |
IJCAI | 6 |
| 2024 | Hundredfold Accelerating for Pathological Images Diagnosis and Prognosis through Self-reform Critical Region Focusing
Xiaotian Yu, Haoming Luo, Jiacong Hu, Xiuming Zhang, Yijun Bei, Mingli Song, Zunlei Feng |
IJCAI | 4 |
| 2024 | Loose Lesion Location Self-supervision Enhanced Colorectal Cancer Diagnosis
Tianhong Gao, Jie Song 0011, Xiaotian Yu, Shengxuming Zhang, Xiuming Zhang, Zipeng Zhong, Mingli Song, Zunlei Feng |
MICCAI (11) | 9 |
| 2023 | DiffusionRig: Learning Personalized Priors for Facial Appearance EditingabstractWe address the problem of learning person-specific facial priors from a small number (e.g., 20) of portrait photos of the same person. This enables us to edit this specific person's facial appearance, such as expression and lighting, while preserving their identity and high-frequency facial details. Key to our approach, which we dub DiffusionRig, is a diffusion model conditioned on, or “rigged by,“ crude 3D face models estimated from single in-the-wild images by an off-the-shelf estimator. On a high level, DiffusionRig learns to map simplistic renderings of 3D face models to realistic photos of a given person. Specifically, DiffusionRig is trained in two stages: It first learns generic facial priors from a large-scale face dataset and then person-specific priors from a small portrait photo collection of the person of interest. By learning the CGI-to-photo mapping with such personalized priors,DiffusionRig can “rig“ the lighting, facial expression, head pose, etc. of a portrait photo, conditioned only on coarse 3D models while preserving this person's identity and other high-frequency characteristics. Qualitative and quantitative experiments show that DiffusionRig outperforms existing approaches in both identity preservation and photorealism. Please see the project website: https://diffusionrig.github.io for the supplemental material, video, code, and data. Zheng Ding, Xuaner Cecilia Zhang, Zhihao Xia, Lars Jebe, Zhuowen Tu, Xiuming Zhang |
CVPR | 6 |
| 2023 | SunStage: Portrait Reconstruction and Relighting Using the Sun as a Light StageabstractA light stage uses a series of calibrated cameras and lights to capture a subject's facial appearance under varying illumination and viewpoint. This captured information is crucial for facial reconstruction and relighting. Unfortunately, light stages are often inaccessible: they are expensive and require significant technical expertise for construction and operation. In this paper, we present SunStage: a lightweight alternative to a light stage that captures comparable data using only a smartphone camera and the sun. Our method only requires the user to capture a selfie video outdoors, rotating in place, and uses the varying angles between the sun and the face as guidance in joint reconstruction of facial geometry, reflectance, camera pose, and lighting parameters. Despite the in-the-wild un-calibrated setting, our approach is able to reconstruct detailed facial appearance and geometry, enabling compelling effects such as relighting, novel view synthesis, and reflectance editing. Wang Yifan 0001, Aleksander Holynski, Xiuming Zhang, Xuaner Cecilia Zhang |
CVPR | 3 |
| 2023 | A Loopback Network for Explainable Microvascular Invasion ClassificationabstractMicrovascular invasion (MVI) is a critical factor for prognosis evaluation and cancer treatment. The current diagnosis of MVI relies on pathologists to manually find out cancerous cells from hundreds of blood vessels, which is time-consuming, tedious, and subjective. Recently, deep learning has achieved promising results in medical image analysis tasks. However, the unexplainability of black box models and the requirement of massive annotated samples limit the clinical application of deep learning based diagnostic methods. In this paper, aiming to develop an accurate, objective, and explainable diagnosis tool for MVI, we propose a Loopback Network (LoopNet) for classifying MVI efficiently. With the image-level category annotations of the collected Pathologic Vessel Image Dataset (PVID), LoopNet is devised to be composed binary classification branch and cell locating branch. The latter is devised to locate the area of cancerous cells, regular non-cancerous cells, and background. For healthy samples, the pseudo masks of cells supervise the cell locating branch to distinguish the area of regular non-cancerous cells and background. For each MVI sample, the cell locating branch predicts the mask of cancerous cells. Then the masked cancerous and non-cancerous areas of the same sample are input back to the binary classification branch separately. The loopback between two branches enables the category label to supervise the cell locating branch to learn the locating ability for cancerous areas. Experiment results show that the proposed LoopNet achieves 97.5% accuracy on MVI classification. Surprisingly, the proposed loopback mechanism not only enables LoopNet to predict the cancerous area but also facilitates the classification backbone to achieve better classification performance. Shengxuming Zhang, Tianqi Shi, Xiuming Zhang, Jie Lei 0002, Zunlei Feng, Mingli Song |
CVPR | 4 |
| 2022 | Space and Level Cooperation Framework for Pathological Cancer GradingabstractClinically, the pathological images are intuitive for cancer diagnosis and have been considered as the ‘gold standard’. There are two challenges for applying deep learning into the pathological images analysis: the ultra-large size and the noisy annotations. A pathological image usually contains billions of pixels, which is unsuitable for normal classification models. Furthermore, the ultra-large size and mixed cancerous cells compel the doctor to draw rough boundaries of lesion area according to the cancerous level, which brings two kinds of noisy labels: space noise (annotating inaccurate scope of cancerous area) and level noise (annotating inaccurate cancerous level). Based on the above findings, we propose the space and level cooperation framework, comprising a space-aware branch and a level-aware branch, for pathological cancer grading with noisy annotations. The space-aware branch first turns the ultra-large image into a Multilayer Superpixel (MS) graph, significantly reducing the size and preserving the global features. Then, a global-to-local rectifying strategy is adopted to solve the space noise. The level-aware branch adopts different grouped kernels and a novel grading loss function to handle level noise. Mean-while, two branches cooperate through complementing missing features of each other for handling the above two challenges. Extensive experiments demonstrate that with noisy annotations, the proposed framework achieves SOTA performance on our HCC dataset and two public datasets. Xiaotian Yu, Zunlei Feng, Xiuming Zhang, Thomas Li |
VCIP | 3 |
| 2021 | Edge-competing Pathological Liver Vessel Segmentation with Limited LabelsabstractThe microvascular invasion (MVI) is a major prognostic factor in hepatocellular carcinoma, which is one of the malignant tumors with the highest mortality rate. The diagnosis of MVI needs discovering the vessels that contain hepatocellular carcinoma cells and counting their number in each vessel, which depends heavily on experiences of the doctor, is largely subjective and time-consuming. However, there is no algorithm as yet tailored for the MVI detection from pathological images. This paper collects the first pathological liver image dataset containing $522$ whole slide images with labels of vessels, MVI, and hepatocellular carcinoma grades. The first and essential step for the automatic diagnosis of MVI is the accurate segmentation of vessels. The unique characteristics of pathological liver images, such as super-large size, multi-scale vessel, and blurred vessel edges, make the accurate vessel segmentation challenging. Based on the collected dataset, we propose an Edge-competing Vessel Segmentation Network (EVS-Net), which contains a segmentation network and two edge segmentation discriminators. The segmentation network, combined with an edge-aware self-supervision mechanism, is devised to conduct vessel segmentation with limited labeled patches. Meanwhile, two discriminators are introduced to distinguish whether the segmented vessel and background contain residual features in an adversarial manner. In the training stage, two discriminators are devised to compete for the predicted position of edges. Exhaustive experiments demonstrate that, with only limited labeled patches, EVS-Net achieves a close performance of fully supervised methods, which provides a convenient tool for the pathological liver vessel segmentation. Code is publicly available at https://github.com/wang97zh/EVS-Net. Zunlei Feng, Xinchao Wang, Xiuming Zhang, Lechao Cheng, Jie Lei 0002, Mingli Song |
AAAI | 4 |
| 2021 | Tendentious Noise-rectifying Framework for Pathological HCC Grading
Xiaotian Yu, Zunlei Feng, Thomas Kwok To Li, Xiuming Zhang, Mingli Song |
BMVC | 5 |
| 2021 | NeRV: Neural Reflectance and Visibility Fields for Relighting and View SynthesisabstractWe present a method that takes as input a set of images of a scene illuminated by unconstrained known lighting, and produces as output a 3D representation that can be rendered from novel viewpoints under arbitrary lighting conditions. Our method represents the scene as a continuous volumetric function parameterized as MLPs whose inputs are a 3D location and whose outputs are the following scene properties at that input location: volume density, surface normal, material parameters, distance to the first surface intersection in any direction, and visibility of the external environment in any direction. Together, these allow us to render novel views of the object under arbitrary lighting, including indirect illumination effects. The predicted visibility and surface intersection fields are critical to our model’s ability to simulate direct and indirect illumination during training, because the brute-force techniques used by prior work are intractable for lighting conditions outside of controlled setups with a single light. Our method outperforms alternative approaches for recovering relightable 3D scene representations, and performs well in complex lighting settings that have posed a significant challenge to prior work. Pratul P. Srinivasan, Boyang Deng, Xiuming Zhang, Matthew Tancik, Ben Mildenhall, Jonathan T. Barron |
CVPR | 3 |
| 2021 | Editing Conditional Radiance FieldsabstractA neural radiance field (NeRF) is a scene model supporting high-quality view synthesis, optimized per scene. In this paper, we explore enabling user editing of a category-level NeRF – also known as a conditional radiance field – trained on a shape category. Specifically, we introduce a method for propagating coarse 2D user scribbles to the 3D space, to modify the color or shape of a local region. First, we propose a conditional radiance field that incorporates new modular network components, including a shape branch that is shared across object instances. Observing multiple instances of the same category, our model learns underlying part semantics without any supervision, thereby allowing the propagation of coarse 2D user scribbles to the entire 3D region (e.g., chair seat). Next, we propose a hybrid network update strategy that targets specific network components, which balances efficiency and accuracy. During user interaction, we formulate an optimization problem that both satisfies the user’s constraints and preserves the original object structure. We demonstrate our editing approach on rendered views of three shape datasets and show that it outperforms prior neural editing approaches. Finally, we edit the appearance and shape of a single-view real photograph and show that the edit propagates to extrapolated novel views. Steven Liu, Xiuming Zhang, Zhoutong Zhang, Richard Zhang 0001, Jun-Yan Zhu, Bryan Russell |
ICCV | 2 |
| 2021 | NeRFactor: neural factorization of shape and reflectance under an unknown illuminationabstractWe address the problem of recovering the shape and spatially-varying reflectance of an object from multi-view images (and their camera poses) of an object illuminated by one unknown lighting condition. This enables the rendering of novel views of the object under arbitrary environment lighting and editing of the object's material properties. The key to our approach, which we call Neural Radiance Factorization (NeRFactor), is to distill the volumetric geometry of a Neural Radiance Field (NeRF) [Mildenhall et al. 2020] representation of the object into a surface representation and then jointly refine the geometry while solving for the spatially-varying reflectance and environment lighting. Specifically, NeRFactor recovers 3D neural fields of surface normals, light visibility, albedo, and Bidirectional Reflectance Distribution Functions (BRDFs) without any supervision, using only a re-rendering loss, simple smoothness priors, and a data-driven BRDF prior learned from real-world BRDF measurements. By explicitly modeling light visibility, NeRFactor is able to separate shadows from albedo and synthesize realistic soft or hard shadows under arbitrary lighting conditions. NeRFactor is able to recover convincing 3D models for free-viewpoint relighting in this challenging and underconstrained capture setup for both synthetic and real scenes. Qualitative and quantitative experiments show that NeRFactor outperforms classic and deep learning-based state of the art across various tasks. Our videos, code, and data are available at people.csail.mit.edu/xiuming/projects/nerfactor/. Xiuming Zhang, Pratul P. Srinivasan, Boyang Deng, Paul E. Debevec, William T. Freeman, Jonathan T. Barron |
ACM Trans. Graph. | 1 |
| 2020 | Perspective Plane Program Induction From a Single ImageabstractWe study the inverse graphics problem of inferring a holistic representation for natural images. Given an input image, our goal is to induce a neuro-symbolic, program-like representation that jointly models camera poses, object locations, and global scene structures. Such high-level, holistic scene representations further facilitate low-level image manipulation tasks such as inpainting. We formulate this problem as jointly finding the camera pose and scene structure that best describe the input image. The benefits of such joint inference are two-fold: scene regularity serves as a new cue for perspective correction, and in turn, correct perspective correction leads to a simplified scene structure, similar to how the correct shape leads to the most regular texture in shape from texture. Our proposed framework, Perspective Plane Program Induction (P3I), combines search-based and gradient-based algorithms to efficiently solve the problem. P3I outperforms a set of baselines on a collection of Internet images, across tasks including camera pose estimation, global structure inference, and down-stream image manipulation tasks. Jiayuan Mao, Xiuming Zhang, William T. Freeman, Josh Tenenbaum, Jiajun Wu 0001 |
CVPR | 3 |
| 2020 | Multi-Plane Program Induction with 3D Box PriorsabstractWe consider two important aspects in understanding and editing images: modeling regular, program-like texture or patterns in 2D planes, and 3D posing of these planes in the scene. Unlike prior work on image-based program synthesis, which assumes the image contains a single visible 2D plane, we present Box Program Induction (BPI), which infers a program-like scene representation that simultaneously models repeated structure on multiple 2D planes, the 3D position and orientation of the planes, and camera parameters, all from a single image. Our model assumes a box prior, i.e., that the image captures either an inner view or an outer view of a box in 3D. It uses neural networks to infer visual cues such as vanishing points, wireframe lines to guide a search-based algorithm to find the program that best explains the image. Such a holistic, structured scene representation enables 3D-aware interactive image editing operations such as inpainting missing pixels, changing camera parameters, and extrapolate the image contents. Jiayuan Mao, Xiuming Zhang, William T. Freeman, Josh Tenenbaum, Noah Snavely, Jiajun Wu 0001 |
NeurIPS | 3 |
| 2020 | Light stage super-resolution: continuous high-frequency relightingabstractThe light stage has been widely used in computer graphics for the past two decades, primarily to enable the relighting of human faces. By capturing the appearance of the human subject under different light sources, one obtains the light transport matrix of that subject, which enables image-based relighting in novel environments. However, due to the finite number of lights in the stage, the light transport matrix only represents a sparse sampling on the entire sphere. As a consequence, relighting the subject with a point light or a directional source that does not coincide exactly with one of the lights in the stage requires interpolation and resampling the images corresponding to nearby lights, and this leads to ghosting shadows, aliased specularities, and other artifacts. To ameliorate these artifacts and produce better results under arbitrary high-frequency lighting, this paper proposes a learning-based solution for the "super-resolution" of scans of human faces taken from a light stage. Given an arbitrary "query" light direction, our method aggregates the captured images corresponding to neighboring lights in the stage, and uses a neural network to synthesize a rendering of the face that appears to be illuminated by a "virtual" light source at the query location. This neural network must circumvent the inherent aliasing and regularity of the light stage data that was used for training, which we accomplish through the use of regularized traditional interpolation methods within our network. Our learned model is able to produce renderings for arbitrary light directions that exhibit realistic shadows and specular highlights, and is able to generalize across a wide variety of subjects. Our super-resolution approach enables more accurate renderings of human subjects under detailed environment maps, or the construction of simpler light stages that contain fewer light sources while still yielding comparable quality renderings as light stages with more densely sampled lights. Tiancheng Sun, Zexiang Xu, Xiuming Zhang, Sean Ryan Fanello, Christoph Rhemann, Paul E. Debevec, Yun-Ta Tsai, Jonathan T. Barron, Ravi Ramamoorthi |
ACM Trans. Graph. | 3 |
| 2020 | Portrait shadow manipulationabstractCasually-taken portrait photographs often suffer from unflattering lighting and shadowing because of suboptimal conditions in the environment. Aesthetic qualities such as the position and softness of shadows and the lighting ratio between the bright and dark parts of the face are frequently determined by the constraints of the environment rather than by the photographer. Professionals address this issue by adding light shaping tools such as scrims, bounce cards, and flashes. In this paper, we present a computational approach that gives casual photographers some of this control, thereby allowing poorly-lit portraits to be relit post-capture in a realistic and easily-controllable way. Our approach relies on a pair of neural networks---one to remove foreign shadows cast by external objects, and another to soften facial shadows cast by the features of the subject and to add a synthetic fill light to improve the lighting ratio. To train our first network we construct a dataset of real-world portraits wherein synthetic foreign shadows are rendered onto the face, and we show that our network learns to remove those unwanted shadows. To train our second network we use a dataset of Light Stage scans of human subjects to construct input/output pairs of input images harshly lit by a small light source, and variably softened and fill-lit output images of each face. We propose a way to explicitly encode facial symmetry and show that our dataset and training procedure enable the model to generalize to images taken in the wild. Together, these networks enable the realistic and aesthetically pleasing enhancement of shadows and lights in real-world portrait images. 1 Xuaner Cecilia Zhang, Jonathan T. Barron, Yun-Ta Tsai, Rohit Pandey, Xiuming Zhang, Ren Ng, David E. Jacobs |
ACM Trans. Graph. | 5 |
| 2019 | Program-Guided Image ManipulatorsabstractHumans are capable of building holistic representations for images at various levels, from local objects, to pairwise relations, to global structures. The interpretation of structures involves reasoning over repetition and symmetry of the objects in the image. In this paper, we present the Program-Guided Image Manipulator (PG-IM), inducing neuro-symbolic program-like representations to represent and manipulate images. Given an image, PG-IM detects repeated patterns, induces symbolic programs, and manipulates the image using a neural network that is guided by the program. PG-IM learns from a single image, exploiting its internal statistics. Despite trained only on image inpainting, PG-IM is directly capable of extrapolation and regularity editing in a unified framework. Extensive experiments show that PG-IM achieves superior performance on all the tasks. Xiuming Zhang, Jiayuan Mao, William T. Freeman, Josh Tenenbaum, Jiajun Wu 0001 |
ICCV | 1 |
| 2018 | Pix3D: Dataset and Methods for Single-Image 3D Shape ModelingabstractWe study 3D shape modeling from a single image and make contributions to it in three aspects. First, we present Pix3D, a large-scale benchmark of diverse image-shape pairs with pixel-level 2D-3D alignment. Pix3D has wide applications in shape-related tasks including reconstruction, retrieval, viewpoint estimation, etc. Building such a large-scale dataset, however, is highly challenging; existing datasets either contain only synthetic data, or lack precise alignment between 2D images and 3D shapes, or only have a small number of images. Second, we calibrate the evaluation criteria for 3D shape reconstruction through behavioral studies, and use them to objectively and systematically benchmark cutting-edge reconstruction algorithms on Pix3D. Third, we design a novel model that simultaneously performs 3D reconstruction and pose estimation; our multi-task learning approach achieves state-of-the-art performance on both tasks. Xingyuan Sun, Jiajun Wu 0001, Xiuming Zhang, Zhoutong Zhang, Chengkai Zhang, Tianfan Xue, Josh Tenenbaum, William T. Freeman |
CVPR | 3 |
| 2018 | Learning Shape Priors for Single-View 3D Completion And Reconstruction
Jiajun Wu 0001, Chengkai Zhang, Xiuming Zhang, Zhoutong Zhang, William T. Freeman, Josh Tenenbaum |
ECCV (11) | 3 |
| 2018 | Learning to Reconstruct Shapes from Unseen ClassesabstractFrom a single image, humans are able to perceive the full 3D shape of an object by exploiting learned shape priors from everyday life. Contemporary single-image 3D reconstruction algorithms aim to solve this task in a similar fashion, but often end up with priors that are highly biased by training classes. Here we present an algorithm, Generalizable Reconstruction (GenRe), designed to capture more generic, class-agnostic shape priors. We achieve this with an inference network and training procedure that combine 2.5D representations of visible surfaces (depth and silhouette), spherical shape representations of both visible and non-visible surfaces, and 3D voxel-based representations, in a principled manner that exploits the causal structure of how 3D shapes give rise to 2D images. Experiments demonstrate that GenRe performs well on single-view shape reconstruction, and generalizes to diverse novel objects from categories not seen during training. Xiuming Zhang, Zhoutong Zhang, Chengkai Zhang, Josh Tenenbaum, William T. Freeman, Jiajun Wu 0001 |
NeurIPS | 1 |
| 2018 | MoSculp: Interactive Visualization of Shape and TimeabstractWe present a system that visualizes complex human motion via 3D motion sculptures-a representation that conveys the 3D structure swept by a human body as it moves through space. Our system computes a motion sculpture from an input video, and then embeds it back into the scene in a 3D-aware fashion. The user may also explore the sculpture directly in 3D or physically print it. Our interactive interface allows users to customize the sculpture design, for example, by selecting materials and lighting conditions. To provide this end-to-end workflow, we introduce an algorithm that estimates a human's 3D geometry over time from a set of 2D images, and develop a 3D-aware image-based rendering approach that inserts the sculpture back into the original video. By automating the process, our system takes motion sculpture creation out of the realm of professional artists, and makes it applicable to a wide range of existing video material. By conveying 3D information to users, motion sculptures reveal space-time motion information that is difficult to perceive with the naked eye, and allow viewers to interpret how different parts of the object interact over time. We validate the effectiveness of motion sculptures with user studies, finding that our visualizations are more informative about motion than existing stroboscopic and space-time visualization methods. Xiuming Zhang, Tali Dekel, Tianfan Xue, Andrew Owens, Qiurui He 0001, Jiajun Wu 0001, Stefanie Mueller 0001, William T. Freeman |
UIST | 1 |
| 2014 | Joint encryption and compressed sensing in smart grid data transmissionabstractIn smart grid, a huge amount of privacy data need to be transmitted securely and efficiently. Compressed sensing can be used to improve the transmission efficiency by exploiting the data sparsity, while this sparsity is usually destroyed by the encryption process, making compressed sensing inapplicable. A possible workaround is to perform compressed sensing first and then encrypt the compressed data. However, this two-step workaround lowers the efficiency. In this paper, we propose a novel data transmission scheme EncryCS. Compressed sensing in EncryCS provides security and simultaneously enhances the transmission efficiency in one step. The prerequisite of compressed sensing is to construct a measurement matrix satisfying the restricted isometry property that ensures the perfect recovery of the signal. To make the matrix secret and feasible, we generate it with a pseudorandom sequence generator. By using this matrix, EncryCS transforms a large amount of plaintext data into a small amount of ciphertext data. EncryCS is proven to possess a high security. Extensive simulations are performed based on the real-world data obtained from the PowerNet project at Stanford University. The results demonstrate that EncryCS can compress data by order of magnitude, and the transmitted data can be almost perfectly recovered. Juntao Gao, Xiuming Zhang, Hao Liang 0002, Xuemin Shen |
GLOBECOM | 2 |