VLDB 2026 Research / reviewers in the wild / expert
Sharon X. Huang
dblp:293/8974 · also Sharon Xiaolei Huang, Xiaolei Huang 0001
· DBLP profile ↗
96ranked-venue papers
6as first author
31since 2021 · last 2025
0000-0003-2338-6535ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 67 · 4 first-author · 20 since 2021Artificial intelligence and machine learning · 43 · 3 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 38 · 2 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Towards In-the-wild 3D Plane Reconstruction from a Single Imageabstract3D plane reconstruction from a single image is a crucial yet challenging topic in 3D computer vision. Previous state-of-the-art (SOTA) methods have focused on training their system on a single dataset from either indoor or outdoor domain, limiting their generalizability across diverse testing data. In this work, we introduce a novel framework dubbed ZeroPlane, a Transformer-based model targeting zero-shot 3D plane detection and reconstruction from a single image, over diverse domains and environments. To enable data-driven models across multiple domains, we have curated a large-scale planar benchmark, comprising over 14 datasets and 560,000 high-resolution, dense planar annotations for diverse indoor and outdoor scenes. To address the challenge of achieving desirable planar geometry on multi-dataset training, we propose to disentangle the representation of plane normal and offset, and employ an exemplar-guided, classification-then-regression paradigm to learn plane and offset respectively. Additionally, we employ advanced backbones as image encoder, and present an effective pixel-geometry-enhanced plane embedding module to further facilitate planar reconstruction. Extensive experiments across multiple zero-shot evaluation datasets have demonstrated that our approach significantly outperforms previous methods on both reconstruction accuracy and generalizability, especially over in-the-wild data. Our code and data are available at: https://github.com/jcliu0428/ZeroPlane. Rui Yu 0002, Sili Chen, Sharon X. Huang, Hengkai Guo |
CVPR | 4 |
| 2025 | Enhancing AI-Assisted Stroke Emergency Triage with Adaptive Uncertainty Estimation
Tongan Cai, Haomiao Ni, Yuan Xue 0002, Kelvin K. Wong, John Volpi, James Z. Wang 0001, Sharon X. Huang, Stephen T. C. Wong |
MICCAI (14) | 9 |
| 2025 | RigAnyFace: Scaling Neural Facial Mesh Auto-Rigging with Unlabeled DataabstractIn this paper, we present RigAnyFace (RAF), a scalable neural auto-rigging framework for facial meshes of diverse topologies, including those with multiple disconnected components. RAF deforms a static neutral facial mesh into industry-standard FACS poses to form an expressive blendshape rig. Deformations are predicted by a triangulation-agnostic surface learning network augmented with our tailored architecture design to condition on FACS parameters and efficiently process disconnected components. For training, we curated a dataset of facial meshes, with a subset meticulously rigged by professional artists to serve as accurate 3D ground truth for deformation supervision. Due to the high cost of manual rigging, this subset is limited in size, constraining the generalization ability of models trained exclusively on it. To address this, we design a 2D supervision strategy for unlabeled neutral meshes without rigs. This strategy increases data diversity and allows for scaled training, thereby enhancing the generalization ability of models trained on this augmented data. Extensive experiments demonstrate that RAF is able to rig meshes of diverse topologies on not only our artist-crafted assets but also in-the-wild samples, outperforming previous works in accuracy and generalizability. Moreover, our method advances beyond prior work by supporting multiple disconnected components, such as eyeballs, for more detailed expression animation. Dario Kneubuehler, Maurice Chu, Ian Sachs, Haomiao Jiang, Sharon X. Huang |
NeurIPS | 6 |
| 2025 | Computer-Aided Layout Generation for Building Design: A ReviewabstractGenerating realistic building layouts for automatic building design has been studied in both computer vision and architectural domains. Traditional approaches in the latter, which are based on optimization techniques or heuristic design guidelines, can synthesize desirable layouts, but usually require post-processing and involve human interaction in the design pipeline, making them costly and time-consuming. The advent of deep generative models has significantly improved the fidelity and diversity of the generated architecture layouts, reducing the workload of designers and making the process much more efficient. This paper presents a comprehensive review of three major research topics in architectural layout design and generation: floorplan layout generation, scene layout synthesis, and generation of various other formats of building layouts. For each topic, we overview the leading paradigms, categorized either by research domains (architecture or machine learning) or by user input conditions or constraints. We then introduce commonly-adopted benchmark datasets used to verify the effectiveness of the methods, as well as corresponding evaluation metrics. Finally, we identify the well-solved problems and limitations of existing approaches, and then propose promising directions for future research. This survey has an associated project which aims to maintain the resources, at https://github.com/jcliu0428/awesome-building-layout-generation. Yuan Xue 0002, Haomiao Ni, Rui Yu 0002, Zihan Zhou 0001, Sharon X. Huang |
Comput. Vis. Media | 6 |
| 2024 | Semantic-Aware Transformation-Invariant RoI AlignabstractGreat progress has been made in learning-based object detection methods in the last decade. Two-stage detectors often have higher detection accuracy than one-stage detectors, due to the use of region of interest (RoI) feature extractors which extract transformation-invariant RoI features for different RoI proposals, making refinement of bounding boxes and prediction of object categories more robust and accurate. However, previous RoI feature extractors can only extract invariant features under limited transformations. In this paper, we propose a novel RoI feature extractor, termed Semantic RoI Align (SRA), which is capable of extracting invariant RoI features under a variety of transformations for two-stage detectors. Specifically, we propose a semantic attention module to adaptively determine different sampling areas by leveraging the global and local semantic relationship within the RoI. We also propose a Dynamic Feature Sampler which dynamically samples features based on the RoI aspect ratio to enhance the efficiency of SRA, and a new position embedding, i.e., Area Embedding, to provide more accurate position information for SRA through an improved sampling area representation. Experiments show that our model significantly outperforms baseline models with slight computational overhead. In addition, it shows excellent generalization ability and can be used to improve performance with various state-of-the-art backbones and detection methods. The code is available at https://github.com/cxjyxxme/SemanticRoIAlign. Guo-Ye Yang, George Kiyohiro Nakayama, Zi-Kai Xiao, Tai-Jiang Mu, Sharon X. Huang, Shi-Min Hu 0001 |
AAAI | 5 |
| 2024 | TI2V-Zero: Zero-Shot Image Conditioning for Text-to-Video Diffusion ModelsabstractText-conditioned image-to-video generation (TI2V) aims to synthesize a realistic video starting from a given image (e.g., a woman's photo) and a text description (e.g., “a woman is drinking water.”). Existing TI2V frameworks often require costly training on video-text datasets and spe-cific model designs for text and image conditioning. In this paper, we propose TI2V-Zero, a zero-shot, tuning-free method that empowers a pretrained text-to-video (T2V) diffusion model to be conditioned on a provided image, enabling TI2V generation without any optimization, fine-tuning, or introducing external modules. Our approach leverages a pretrained T2V diffusion foundation model as the generative prior. To guide video generation with the additional image input, we propose a “repeat-and-slide” strategy that modulates the reverse denoising process, al-lowing the frozen diffusion model to synthesize a video frame-by-frame starting from the provided image. To ensure temporal continuity, we employ a DDPM inversion strategy to initialize Gaussian noise for each newly synthesized frame and a resampling technique to help preserve visual details. We conduct comprehensive experiments on both domain-specific and open-domain datasets, where TI2V-Zero consistently outperforms a recent open-domain TI2V model. Furthermore, we show that TI2V-Zero can seam-lessly extend to other tasks such as video infilling and pre-diction when provided with more images. Its autoregressive design also supports long video generation. Haomiao Ni, Bernhard Egger 0001, Suhas Lohit, Anoop Cherian, Ye Wang 0001, Toshiaki Koike-Akino, Sharon X. Huang, Tim K. Marks |
CVPR | 7 |
| 2024 | NeRF-Enhanced Outpainting for Faithful Field-of-View ExtrapolationabstractIn various applications, such as robotic navigation and remote visual assistance, expanding the field of view (FOV) of the camera proves beneficial for enhancing environmental perception. Unlike image outpainting techniques aimed solely at generating aesthetically pleasing visuals, these applications demand an extended view that faithfully represents the scene. To achieve this, we formulate a new problem of faithful FOV extrapolation that utilizes a set of pre-captured images as prior knowledge of the scene. To address this problem, we present a simple yet effective solution called NeRF-Enhanced Outpainting (NEO) that uses extended-FOV images generated through NeRF to train a scene-specific image outpainting model. To assess the performance of NEO, we conduct comprehensive evaluations on three photorealistic datasets and one real-world dataset. Extensive experiments on the benchmark datasets showcase the robustness and potential of our method in addressing this challenge. We believe our work lays a strong foundation for future exploration within the research community. Rui Yu 0002, Zihan Zhou 0001, Sharon X. Huang |
ICRA | 4 |
| 2024 | MonoPlane: Exploiting Monocular Geometric Cues for Generalizable 3D Plane ReconstructionabstractThis paper presents a generalizable 3D plane detection and reconstruction framework named MonoPlane. Unlike previous robust estimator-based works (which require multiple images or RGB-D input) and learning-based works (which suffer from domain shift), MonoPlane combines the best of two worlds and establishes a plane reconstruction pipeline based on monocular geometric cues, resulting in accurate, robust and scalable 3D plane detection and reconstruction in the wild. Specifically, we first leverage large-scale pre-trained neural networks to obtain the depth and surface normals from a single image. These monocular geometric cues are then incorporated into a proximity-guided RANSAC framework to sequentially fit each plane instance. We exploit effective 3D point proximity and model such proximity via a graph within RANSAC to guide the plane fitting from noisy monocular depths, followed by image-level multi-plane joint optimization to improve the consistency among all plane instances. We further design a simple but effective pipeline to extend this single-view solution to sparse-view 3D plane reconstruction. Extensive experiments on a list of datasets demonstrate our superior zero-shot generalizability over baselines, achieving state-of-the-art plane reconstruction performance in a transferring setting. Our code is available at https://github.com/thuzhaowang/MonoPlane. Wang Zhao 0001, Yishu Li, Sili Chen, Sharon X. Huang, Yong-Jin Liu 0001, Hengkai Guo |
IROS | 6 |
| 2024 | 3D-Aware Talking-Head Video Motion TransferabstractMotion transfer of talking-head videos involves generating a new video with the appearance of a subject video and the motion pattern of a driving video. Current methodologies primarily depend on a limited number of subject images and 2D representations, thereby neglecting to fully utilize the multi-view appearance features inherent in the subject video. In this paper, we propose a novel 3D-aware talking-head video motion transfer network, Head3D, which fully exploits the subject appearance information by generating a visually-interpretable 3D canonical head from the 2D subject frames with a recurrent network. A key component of our approach is a self-supervised 3D head geometry learning module, designed to predict head poses and depth maps from 2D subject video frames. This module facilitates the estimation of a 3D head in canonical space, which can then be transformed to align with driving video frames. Additionally, we employ an attention-based fusion network to combine the background and other details from subject frames with the 3D subject head to produce the synthetic target video. Our extensive experiments on two public talking-head video datasets demonstrate that Head3D outperforms both 2D and 3D prior arts in the practical cross-identity setting, with evidence showing it can be readily adapted to the pose-controllable novel view synthesis task. Haomiao Ni, Yuan Xue 0002, Sharon X. Huang |
WACV | 4 |
| 2024 | Tuning Vision-Language Models With Multiple Prototypes ClusteringabstractBenefiting from advances in large-scale pre-training, foundation models, have demonstrated remarkable capability in the fields of natural language processing, computer vision, among others. However, to achieve expert-level performance in specific applications, such models often need to be fine-tuned with domain-specific knowledge. In this paper, we focus on enabling vision-language models to unleash more potential for visual understanding tasks under few-shot tuning. Specifically, we propose a novel adapter, dubbed as lusterAdapter, which is based on trainable multiple prototypes clustering algorithm, for tuning the CLIP model. It can not only alleviate the concern of catastrophic forgetting of foundation models by introducing anchors to inherit common knowledge, but also improve the utilization efficiency of few annotated samples via bringing in clustering and domain priors, thereby improving the performance of few-shot tuning. We have conducted extensive experiments on 11 common classification benchmarks. The results show our method significantly surpasses the original CLIP and achieves state-of-the-art (SOTA) performance under all benchmarks and settings. For example, under the 16-shot setting, our method exhibits a remarkable improvement over the original CLIP by 19.6%, and also surpasses TIP-Adapter and GraphAdapter by 2.7% and 2.2%, respectively, in terms of average accuracy across the 11 benchmarks. Menghao Guo 0001, Yi Zhang 0099, Tai-Jiang Mu, Sharon X. Huang, Shi-Min Hu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Conditional Image-to-Video Generation with Latent Flow Diffusion ModelsabstractConditional image-to-video (cI2V) generation aims to synthesize a new plausible video starting from an image (e.g., a person's face) and a condition (e.g., an action class label like smile). The key challenge of the cI2V task lies in the simultaneous generation of realistic spatial appearance and temporal dynamics corresponding to the given image and condition. In this paper, we propose an approach for cI2V using novel latent flow diffusion models (LFDM) that synthesize an optical flow sequence in the latent space based on the given condition to warp the given image. Compared to previous direct-synthesis-based works, our proposed LFDM can better synthesize spatial details and temporal motion by fully utilizing the spatial content of the given image and warping it in the latent space according to the generated temporally-coherent flow. The training of LFDM consists of two separate stages: (1) an unsupervised learning stage to train a latent flow auto-encoder for spatial content generation, including a flow predictor to estimate latent flow between pairs of video frames, and (2) a conditional learning stage to train a 3D-UNet-based diffusion model (DM) for temporal latent flow generation. Unlike previous DMs operating in pixel space or latent feature space that couples spatial and temporal information, the DM in our LFDM only needs to learn a low-dimensional latent flow space for motion generation, thus being more computationally efficient. We conduct comprehensive experiments on multiple datasets, where LFDM consistently outperforms prior arts. Furthermore, we show that LFDM can be easily adapted to new domains by simply finetuning the image decoder. Our code is available at https://github.com/nihaomiao/CVPR23_LFDM. Haomiao Ni, Changhao Shi, Kai Li 0012, Sharon X. Huang, Martin Renqiang Min |
CVPR | 4 |
| 2023 | Be Real in Scale: Swing for True Scale in Dual Camera ModeabstractMany mobile AR apps that use the front-facing camera can benefit significantly from knowing the metric scale of the user’s face. However, the true scale of the face is hard to measure because monocular vision suffers from a fundamental ambiguity in scale. The methods based on prior knowledge about the scene either have a large error or are not easily accessible. In this paper, we propose a new method to measure the face scale by a simple user interaction: the user only needs to swing the phone to capture two selfies while using the recently popular Dual Camera mode. This mode allows simultaneous streaming of the front camera and the rear cameras and has become a key feature in many social apps. A computer vision method is applied to first estimate the absolute motion of the phone from the images captured by two rear cameras, and then calculate the point cloud of the face by triangulation. We develop a prototype mobile app to validate the proposed method. Our user study shows that the proposed method is favored compared to existing methods because of its high accuracy and ease of use. Our method can be built into Dual Camera mode and can enable a wide range of applications (e.g., virtual try-on for online shopping, true-scale 3D face modeling, gaze tracking, and face anti-spoofing) by introducing true scale to smartphone-based XR. The code is available at https://github.com/ruiyu0/Swing-for-True-Scale. Rui Yu 0002, Jian Wang 0100, Sizhuo Ma, Sharon X. Huang, Gurunandan Krishnan |
ISMAR | 4 |
| 2023 | Synthetic Augmentation with Large-Scale Unconditional Pre-training
Jiarong Ye, Haomiao Ni, Sharon X. Huang, Yuan Xue 0002 |
MICCAI (2) | 4 |
| 2023 | Cross-identity Video Motion Retargeting with Joint Transformation and SynthesisabstractIn this paper, we propose a novel dual-branch Transformation-Synthesis network (TS-Net), for video motion retargeting. Given one subject video and one driving video, TS-Net can produce a new plausible video with the subject appearance of the subject video and motion pattern of the driving video. TS-Net consists of a warp-based transformation branch and a warp-free synthesis branch. The novel design of dual branches combines the strengths of deformation-grid-based transformation and warp-free generation for better identity preservation and robustness to occlusion in the synthesized videos. A mask-aware similarity module is further introduced to the transformation branch to reduce computational overhead. Experimental results on face and dance datasets show that TS-Net achieves better performance in video motion retargeting than several state-of-the-art models as well as its single-branch variants. Our code is available at https://github.com/nihaomiao/WACV23_TSNet. Haomiao Ni, Yihao Liu 0003, Sharon X. Huang, Yuan Xue 0002 |
WACV | 3 |
| 2023 | Semi-supervised body parsing and pose estimation for enhancing infant general movement assessment
Haomiao Ni, Yuan Xue 0002, Liya Ma, Qian Zhang 0081, Xiaoye Li, Sharon X. Huang |
Medical Image Anal. | 6 |
| 2023 | Editorial for special issue on explainable and generalizable deep learning methods for medical image computing
Guotai Wang, Shaoting Zhang 0001, Sharon X. Huang, Tom Vercauteren, Dimitris N. Metaxas |
Medical Image Anal. | 3 |
| 2023 | Skeleton-CutMix: Mixing Up Skeleton With Probabilistic Bone Exchange for Supervised Domain AdaptationabstractWe present Skeleton-CutMix, a simple and effective skeleton augmentation framework for supervised domain adaptation and show its advantage in skeleton-based action recognition tasks. Existing approaches usually perform domain adaptation for action recognition with elaborate loss functions that aim to achieve domain alignment. However, they fail to capture the intrinsic characteristics of skeleton representation. Benefiting from the well-defined correspondence between bones of a pair of skeletons, we instead mitigate domain shift by fabricating skeleton data in a mixed domain, which mixes up bones from the source domain and the target domain. The fabricated skeletons in the mixed domain can be used to augment training data and train a more general and robust model for action recognition. Specifically, we hallucinate new skeletons by using pairs of skeletons from the source and target domains; a new skeleton is generated by exchanging some bones from the skeleton in the source domain with corresponding bones from the skeleton in the target domain, which resembles a cut-and-mix operation. When exchanging bones from different domains, we introduce a class-specific bone sampling strategy so that bones that are more important for an action class are exchanged with higher probability when generating augmentation samples for that class. We show experimentally that the simple bone exchange strategy for augmentation is efficient and effective and that distinctive motion features are preserved while mixing both action and style across domains. We validate our method in cross-dataset and cross-age settings on NTU-60 and ETRI-Activity3D datasets with an average gain of over 3% in terms of action recognition accuracy, and demonstrate its superior performance over previous domain adaptation approaches as well as other skeleton augmentation strategies. Yuhe Liu, Tai-Jiang Mu, Sharon X. Huang, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 4 |
| 2023 | Forecasting User Interests Through Topic Tag Predictions in Online Health CommunitiesabstractThe increasing reliance on online communities for healthcare information by patients and caregivers has led to the increase in the spread of misinformation, or subjective, anecdotal and inaccurate or non-specific recommendations, which, if acted on, could cause serious harm to the patients. Hence, there is an urgent need to connect users with accurate and tailored health information in a timely manner to prevent such harm. This article proposes an innovative approach to suggesting reliable information to participants in online communities as they move through different stages in their disease or treatment. We hypothesize that patients with similar histories of disease progression or course of treatment would have similar information needs at comparable stages. Specifically, we pose the problem of predicting topic tags or keywords that describe the future information needs of users based on their profiles, traces of their online interactions within the community (past posts, replies) and the profiles and traces of online interactions of other users with similar profiles and similar traces of past interaction with the target users. The result is a variant of the collaborative information filtering or recommendation system tailored to the needs of users of online health communities. We report results of our experiments on two unique datasets from two different social media platforms which demonstrates the superiority of the proposed approach over the state of the art baselines with respect to accurate and timely prediction of topic tags (and hence information sources of interest). Amogh Subbakrishna Adishesha, Lily Jakielaszek, Fariha Azhar, Peixuan Zhang, Vasant G. Honavar, Fenglong Ma, Chandra Belani, Prasenjit Mitra 0001, Sharon X. Huang |
IEEE J. Biomed. Health Informatics | 9 |
| 2022 | PlaneMVS: 3D Plane Reconstruction from Multi-View StereoabstractWe present a novel framework named PlaneMVS for 3D plane reconstruction from multiple input views with known camera poses. Most previous learning-based plane reconstruction methods reconstruct 3D planes from single images, which highly rely on single-view regression and suffer from depth scale ambiguity. In contrast, we reconstruct 3D planes with a multi-view-stereo (MVS) pipeline that takes advantage of multi-view geometry. We decouple plane reconstruction into a semantic plane detection branch and a plane MVS branch. The semantic plane detection branch is based on a single-view plane detection framework but with differences. The plane MVS branch adopts a set of slanted plane hypotheses to replace conventional depth hypotheses to perform plane sweeping strategy and finally learns pixel-level plane parameters and its planar depth map. We present how the two branches are learned in a balanced way, and propose a soft-pooling loss to associate the outputs of the two branches and make them benefit from each other. Extensive experiments on various indoor datasets show that PlaneMVS significantly outperforms state-of-the-art (SOTA) single-view plane reconstruction methods on both plane detection and 3D geometry metrics. Our method even outperforms a set of SOTA learning-based MVS methods thanks to the learned plane priors. To the best of our knowledge, this is the first work on 3D plane reconstruction within an end-to-end MVS framework. Pan Ji, Nitin Bansal, Changjiang Cai, Qingan Yan, Sharon X. Huang, Yi Xu 0002 |
CVPR | 6 |
| 2022 | Deep Depth from Focus with Differential Focus VolumeabstractDepth-from-focus (DFF) is a technique that infers depth using the focus change of a camera. In this work, we propose a convolutional neural network (CNN) to find the best-focused pixels in a focal stack and infer depth from the focus estimation. The key innovation of the network is the novel deep differential focus volume (DFV). By computing the first-order derivative with the stacked features over different focal distances, DFV is able to capture both the focus and context information for focus analysis. Besides, we also introduce a probability regression mechanism for focus estimation to handle sparsely sampled focal stacks and provide uncertainty estimation to the final prediction. Comprehensive experiments demonstrate that the proposed model achieves state-of-the-art performance on multiple datasets with good generalizability and fast speed. Fengting Yang, Sharon X. Huang, Zihan Zhou 0001 |
CVPR | 2 |
| 2022 | End-to-End Graph-Constrained Vectorized Floorplan Generation with Panoptic Refinement
Yuan Xue 0002, José Pinto Duarte, Krishnendra Shekhawat, Zihan Zhou 0001, Sharon X. Huang |
ECCV (15) | 6 |
| 2022 | Unsupervised Learning of Full-Waveform Inversion: Connecting CNN and Partial Differential Equation in a Loop
Xitong Zhang, Yinpeng Chen, Sharon X. Huang, Zicheng Liu 0001, Youzuo Lin |
ICLR | 4 |
| 2022 | Asymmetry Disentanglement Network for Interpretable Acute Ischemic Stroke Infarct Segmentation in Non-contrast CT Scans
Haomiao Ni, Yuan Xue 0002, Kelvin K. Wong, John Volpi, Stephen T. C. Wong, James Z. Wang 0001, Sharon X. Huang |
MICCAI (8) | 7 |
| 2022 | Patcher: Patch Transformers with Mixture of Experts for Precise Medical Image Segmentation
Yanglan Ou, Sharon X. Huang, Stephen T. C. Wong, John Volpi, James Z. Wang 0001, Kelvin K. Wong |
MICCAI (5) | 3 |
| 2022 | Deep image synthesis from intuitive user input: A review and perspectivesabstractIn many applications of computer graphics, art, and design, it is desirable for a user to provide intuitive non-image input, such as text, sketch, stroke, graph, or layout, and have a computer system automatically generate photo-realistic images according to that input. While classically, works that allow such automatic image content generation have followed a framework of image retrieval and composition, recent advances in deep generative models such as generative adversarial networks (GANs), variational autoencoders (VAEs), and flow-based methods have enabled more powerful and versatile image generation approaches. This paper reviews recent works for image synthesis given intuitive user input, covering advances in input versatility, image generation methodology, benchmark datasets, and evaluation metrics. This motivates new perspectives on input representation and interactivity, cross fertilization between major image generation paradigms, and evaluation and comparison of generation methods. Yuan Xue 0002, Han Zhang 0010, Tao Xu 0029, Song-Hai Zhang, Sharon X. Huang |
Comput. Vis. Media | 6 |
| 2022 | DeepStroke: An efficient stroke screening framework for emergency rooms with multimodal adversarial deep learning
Tongan Cai, Haomiao Ni, Mingli Yu, Sharon X. Huang, Kelvin K. Wong, John Volpi, James Z. Wang 0001, Stephen T. C. Wong |
Medical Image Anal. | 4 |
| 2022 | Scheduling Massive Camera Streams to Optimize Large-Scale Live Video AnalyticsabstractIn smart cities, more and more government departments will make use of live analytics of videos from surveillance cameras in their tasks, such as vehicle traffic monitoring and criminal detection. Obviously, it is costly for each individual department to deploy its own infrastructure,i.e., cameras and analytics system. In this paper, we consider a scenario in which a city deploys an infrastructure and departments submit requests to access and analyze videos for their own purposes. The live analytics of massive streams is computation-intensive and the tasks might be latency-critical, which makes scheduling massive streams to optimize all tasks an essential and challenging work. We exploit an end-edge-cloud architecture and propose an adaptive system to schedule the massive camera streams and tasks, which considers all factors affecting the computation and networking resource consumption,e.g., sharing of model computation, video quality, model partition, and task placement. Particularly, the resource consumption ofFaster R-CNN + ResNet101under each partition scheme is profiled for the first time and we notice the partition must be used together with lossless compression techniques to be beneficial. Furthermore, sometimes tasks might be required to migrate because the scheduling decision made by the system changes to adapt to the changing resource supply and demand. In order to avoid the performance degradation during migration, we propose a non-destructive migration scheme and implement it in the system. Simulations demonstrate our system achieves a total utility close to the maximum and our analytics system performs better than state-of-the-art solutions. Chenghao Rong, Hui Wang 0011, Juncai Liu, Jilong Wang 0001, Sharon X. Huang |
IEEE/ACM Trans. Netw. | 6 |
| 2021 | LambdaUNet: 2.5D Stroke Lesion Segmentation of Diffusion-Weighted MR Images
Yanglan Ou, Sharon X. Huang, Kelvin K. Wong, John Volpi, James Z. Wang 0001, Stephen T. C. Wong |
MICCAI (1) | 3 |
| 2021 | A Multi-attribute Controllable Generative Model for Histopathology Image Synthesis
Jiarong Ye, Yuan Xue 0002, Peter Liu, Richard Zaino, Keith C. Cheng, Sharon X. Huang |
MICCAI (8) | 6 |
| 2021 | Detecting human - object interaction with multi-level pairwise feature networkabstractHuman–object interaction (HOI) detection is crucial for human-centric image understanding which aims to infer ⟨human, action, object⟩ triplets within an image. Recent studies often exploit visual features and the spatial configuration of a human–object pair in order to learn the action linking the human and object in the pair. We argue that such a paradigm of pairwise feature extraction and action inference can be applied not only at the whole human and object instance level, but also at the part level at which a body part interacts with an object, and at the semantic level by considering the semantic label of an object along with human appearance and human–object spatial configuration, to infer the action. We thus propose a multi-level pairwise feature network (PFNet) for detecting human–object interactions. The network consists of three parallel streams to characterize HOI utilizing pairwise features at the above three levels; the three streams are finally fused to give the action prediction. Extensive experiments show that our proposed PFNet outperforms other state-of-the-art methods on the V-COCO dataset and achieves comparable results to the state-of-the-art on the HICO-DET dataset. Tai-Jiang Mu, Sharon X. Huang |
Comput. Vis. Media | 3 |
| 2021 | Selective synthetic augmentation with HistoGAN for improved histopathology image classification
Yuan Xue 0002, Jiarong Ye, Qianying Zhou, L. Rodney Long, Sameer K. Antani, Zhiyun Xue, Carl Cornwell, Richard Zaino, Keith C. Cheng, Sharon X. Huang |
Medical Image Anal. | 10 |
| 2020 | Shape-Aware Organ Segmentation by Predicting Signed Distance MapsabstractIn this work, we propose to resolve the issue existing in current deep learning based organ segmentation systems that they often produce results that do not capture the overall shape of the target organ and often lack smoothness. Since there is a rigorous mapping between the Signed Distance Map (SDM) calculated from object boundary contours and the binary segmentation map, we exploit the feasibility of learning the SDM directly from medical scans. By converting the segmentation task into predicting an SDM, we show that our proposed method retains superior segmentation performance and has better smoothness and continuity in shape. To leverage the complementary information in traditional segmentation training, we introduce an approximated Heaviside function to train the model by predicting SDMs and segmentation maps simultaneously. We validate our proposed models by conducting extensive experiments on a hippocampus segmentation dataset and the public MICCAI 2015 Head and Neck Auto Segmentation Challenge dataset with multiple organs. While our carefully designed backbone 3D segmentation network improves the Dice coefficient by more than 5% compared to current state-of-the-arts, the proposed model with SDM learning produces smoother segmentation results with smaller Hausdorff distance and average surface distance, thus proving the effectiveness of our method. Yuan Xue 0002, Guanzhong Gong, Chao Huang 0016, Wei Fan 0001, Sharon X. Huang |
AAAI | 9 |
| 2020 | Neural Wireframe Renderer: Learning Wireframe to Image Translations
Yuan Xue 0002, Zihan Zhou 0001, Sharon X. Huang |
ECCV (26) | 3 |
| 2020 | SiamParseNet: Joint Body Parsing and Label Propagation in Infant Movement Videos
Haomiao Ni, Yuan Xue 0002, Qian Zhang 0081, Sharon X. Huang |
MICCAI (4) | 4 |
| 2020 | Synthetic Sample Selection via Reinforcement Learning
Jiarong Ye, Yuan Xue 0002, L. Rodney Long, Sameer K. Antani, Zhiyun Xue, Keith C. Cheng, Sharon X. Huang |
MICCAI (1) | 7 |
| 2020 | Toward Rapid Stroke Diagnosis with Multimodal Deep Learning
Mingli Yu, Tongan Cai, Sharon X. Huang, Kelvin K. Wong, John Volpi, James Z. Wang 0001, Stephen T. C. Wong |
MICCAI (3) | 3 |
| 2020 | Lane Detection: A Survey with New Results
Dun Liang, Shao-Kui Zhang, Tai-Jiang Mu, Sharon X. Huang |
J. Comput. Sci. Technol. | 5 |
| 2019 | Synthetic Augmentation and Feature-Based Filtering for Improved Cervical Histopathology Image Classification
Yuan Xue 0002, Qianying Zhou, Jiarong Ye, L. Rodney Long, Sameer K. Antani, Carl Cornwell, Zhiyun Xue, Sharon X. Huang |
MICCAI (1) | 8 |
| 2019 | A three-stage real-time detector for traffic signs in large panoramasabstractTraffic sign detection is one of the key components in autonomous driving. Advanced autonomous vehicles armed with high quality sensors capture high definition images for further analysis. Detecting traffic signs, moving vehicles, and lanes is important for localization and decision making. Traffic signs, especially those that are far from the camera, are small, and so are challenging to traditional object detection methods. In this work, in order to reduce computational cost and improve detection performance, we split the large input images into small blocks and then recognize traffic signs in the blocks using another detection module. Therefore, this paper proposes a three-stage traffic sign detector, which connects a BlockNet with an RPN–RCNN detection network. BlockNet, which is composed of a set of CNN layers, is capable of performing block-level foreground detection, making inferences in less than 1 ms. Then, the RPN–RCNN two-stage detector is used to identify traffic sign objects in each block; it is trained on a derived dataset named TT100KPatch. Experiments show that our framework can achieve both state-of-the-art accuracy and recall; its fastest detection speed is 102 fps. Ruochen Fan, Sharon X. Huang, Zhe Zhu, Ruofeng Tong 0001 |
Comput. Vis. Media | 3 |
| 2019 | StackGAN++: Realistic Image Synthesis with Stacked Generative Adversarial NetworksabstractAlthough Generative Adversarial Networks (GANs) have shown remarkable success in various tasks, they still face challenges in generating high quality images. In this paper, we propose Stacked Generative Adversarial Networks (StackGANs) aimed at generating high-resolution photo-realistic images. First, we propose a two-stage generative adversarial network architecture, StackGAN-v1, for text-to-image synthesis. The Stage-I GAN sketches the primitive shape and colors of a scene based on a given text description, yielding low-resolution images. The Stage-II GAN takes Stage-I results and the text description as inputs, and generates high-resolution images with photo-realistic details. Second, an advanced multi-stage generative adversarial network architecture, StackGAN-v2, is proposed for both conditional and unconditional generative tasks. Our StackGAN-v2 consists of multiple generators and multiple discriminators arranged in a tree-like structure; images at multiple scales corresponding to the same scene are generated from different branches of the tree. StackGAN-v2 shows more stable training behavior than StackGAN-v1 by jointly approximating multiple distributions. Extensive experiments demonstrate that the proposed stacked generative adversarial networks significantly outperform other state-of-the-art methods in generating photo-realistic images. Han Zhang 0010, Tao Xu 0029, Hongsheng Li 0001, Shaoting Zhang 0001, Xiaogang Wang 0001, Sharon X. Huang, Dimitris N. Metaxas |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2018 | AttnGAN: Fine-Grained Text to Image Generation With Attentional Generative Adversarial NetworksabstractIn this paper, we propose an Attentional Generative Adversarial Network (AttnGAN) that allows attention-driven, multi-stage refinement for fine-grained text-to-image generation. With a novel attentional generative network, the AttnGAN can synthesize fine-grained details at different sub-regions of the image by paying attentions to the relevant words in the natural language description. In addition, a deep attentional multimodal similarity model is proposed to compute a fine-grained image-text matching loss for training the generator. The proposed AttnGAN significantly outperforms the previous state of the art, boosting the best reported inception score by 14.14% on the CUB dataset and 170.25% on the more challenging COCO dataset. A detailed analysis is also performed by visualizing the attention layers of the AttnGAN. It for the first time shows that the layered attentional GAN is able to automatically select the condition at the word level for generating different parts of the image. Tao Xu 0029, Pengchuan Zhang, Qiuyuan Huang, Han Zhang 0010, Zhe Gan, Sharon X. Huang, Xiaodong He 0001 |
CVPR | 6 |
| 2018 | Multimodal Recurrent Model with Attention for Automated Radiology Report Generation
Yuan Xue 0002, Tao Xu 0029, L. Rodney Long, Zhiyun Xue, Sameer K. Antani, George R. Thoma, Sharon X. Huang |
MICCAI (1) | 7 |
| 2018 | An efficient algorithm for dynamic MRI using low-rank and total variation regularizations
Jiawen Yao, Zheng Xu 0005, Sharon X. Huang, Junzhou Huang |
Medical Image Anal. | 3 |
| 2017 | Integrated local binary pattern texture features for classification of breast tissue imaged by optical coherence microscopy
Sunhua Wan, Hsiang-Chieh Lee, Sharon X. Huang, Tao Xu 0029, Xianxu Zeng, Yuri Sheikine, James L. Connolly, James G. Fujimoto, Chao Zhou 0005 |
Medical Image Anal. | 3 |
| 2017 | Multi-feature based benchmark for cervical dysplasia classification evaluation
Tao Xu 0029, Han Zhang 0010, Cheng Xin, L. Rodney Long, Zhiyun Xue, Sameer K. Antani, Sharon X. Huang |
Pattern Recognit. | 8 |
| 2016 | SPDA-CNN: Unifying Semantic Part Detection and Abstraction for Fine-Grained RecognitionabstractMost convolutional neural networks (CNNs) lack midlevel layers that model semantic parts of objects. This limits CNN-based methods from reaching their full potential in detecting and utilizing small semantic parts in recognition. Introducing such mid-level layers can facilitate the extraction of part-specific features which can be utilized for better recognition performance. This is particularly important in the domain of fine-grained recognition. In this paper, we propose a new CNN architecture that integrates semantic part detection and abstraction (SPDACNN) for fine-grained classification. The proposed network has two sub-networks: one for detection and one for recognition. The detection sub-network has a novel top-down proposal method to generate small semantic part candidates for detection. The classification sub-network introduces novel part layers that extract features from parts detected by the detection sub-network, and combine them for recognition. As a result, the proposed architecture provides an end-to-end network that performs detection, localization of multiple semantic parts, and whole object recognition within one framework that shares the computation of convolutional filters. Our method outperforms state-of-theart methods with a large margin for small parts detection (e.g. our precision of 93.40% vs the best previous precision of 74.00% for detecting the head on CUB-2011). It also compares favorably to the existing state-of-the-art on finegrained classification, e.g. it achieves 85.14% accuracy on CUB-2011. Han Zhang 0010, Tao Xu 0029, Sharon X. Huang, Shaoting Zhang 0001, Ahmed M. Elgammal, Dimitris N. Metaxas |
CVPR | 4 |
| 2016 | Traffic-Sign Detection and Classification in the WildabstractAlthough promising results have been achieved in the areas of traffic-sign detection and classification, few works have provided simultaneous solutions to these two tasks for realistic real world images. We make two contributions to this problem. Firstly, we have created a large traffic-sign benchmark from 100000 Tencent Street View panoramas, going beyond previous benchmarks. It provides 100000 images containing 30000 traffic-sign instances. These images cover large variations in illuminance and weather conditions. Each traffic-sign in the benchmark is annotated with a class label, its bounding box and pixel mask. We call this benchmark Tsinghua-Tencent 100K. Secondly, we demonstrate how a robust end-to-end convolutional neural network (CNN) can simultaneously detect and classify trafficsigns. Most previous CNN image processing solutions target objects that occupy a large proportion of an image, and such networks do not work well for target objects occupying only a small fraction of an image like the traffic-signs here. Experimental results show the robustness of our network and its superiority to alternatives. The benchmark, source code and the CNN model introduced in this paper is publicly available1. Zhe Zhu, Dun Liang, Song-Hai Zhang, Sharon X. Huang, Baoli Li 0004, Shi-Min Hu 0001 |
CVPR | 4 |
| 2016 | Multimodal Deep Learning for Cervical Dysplasia Diagnosis
Tao Xu 0029, Han Zhang 0010, Sharon X. Huang, Shaoting Zhang 0001, Dimitris N. Metaxas |
MICCAI (2) | 3 |
| 2016 | Distributed Learning for Multi-Channel Selection in Wireless Network MonitoringabstractIn this paper, we address an important problem in the wireless monitoring, i.e., how to choose channels with best (or worst) qualities timely and accurately. We consider both scenarios of one or more sniffers simultaneously monitoring multiple channels in the same area. Since the channel information is initially unknown to the sniffers, we shall adopt learning methods during the monitoring to predict the channel condition by a short time of observation. We formulate this problem as a novel branch of the classic multi-armed bandit (MAB) problem, named exploration bandit problem, to achieve a trade-off between monitoring time/resource budget and the channel selection accuracy. In the multiple sniffer cases, including partly-distributed (with limited communications) and fully-distributed (without any communications) scenarios, we take communication costs and interference costs into account, and analyze how these costs affect the accuracy of channel selection. Extensive simulations are conducted and the results show that the proposed algorithms could achieve higher channel selection accuracy than other exploration bandit approaches, hence it proves the advantages of the proposed algorithms. Yuan Xue 0002, Pan Zhou 0001, Tao Jiang 0002, Shiwen Mao, Sharon X. Huang |
SECON | 5 |
| 2015 | Accelerated Dynamic MRI Reconstruction with Total Variation and Nuclear Norm Regularization
Jiawen Yao, Zheng Xu 0005, Sharon X. Huang, Junzhou Huang |
MICCAI (2) | 3 |
| 2015 | Global Contrast Based Salient Region DetectionabstractAutomatic estimation of salient object regions across images, without any prior assumption or knowledge of the contents of the corresponding scenes, enhances many computer vision and computer graphics applications. We introduce a regional contrast based salient object detection algorithm, which simultaneously evaluates global contrast differences and spatial weighted coherence scores. The proposed algorithm is simple, efficient, naturally multi-scale, and produces full-resolution, high-quality saliency maps. These saliency maps are further used to initialize a novel iterative version of GrabCut, namely SaliencyCut, for high quality unsupervised salient object segmentation. We extensively evaluated our algorithm using traditional salient object detection datasets, as well as a more challenging Internet image dataset. Our experimental results demonstrate that our algorithm consistently outperforms 15 existing salient object detection and segmentation methods, yielding higher precision and better recall rates. We also show that our algorithm can be used to efficiently extract salient object masks from Internet images, enabling effective sketch-based image retrieval (SBIR) via simple shape comparisons. Despite such noisy internet images, where the saliency regions are ambiguous, our saliency guided image retrieval achieves a superior retrieval rate compared with state-of-the-art SBIR methods, and additionally provides important target object region information. Ming-Ming Cheng, Niloy J. Mitra, Sharon X. Huang, Philip Torr 0001, Shi-Min Hu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2015 | Multimodal Entity Coreference for Cervical Dysplasia DiagnosisabstractCervical cancer is the second most common type of cancer for women. Existing screening programs for cervical cancer, such as Pap Smear, suffer from low sensitivity. Thus, many patients who are ill are not detected in the screening process. Using images of the cervix as an aid in cervical cancer screening has the potential to greatly improve sensitivity, and can be especially useful in resource-poor regions of the world. In this paper, we develop a data-driven computer algorithm for interpreting cervical images based on color and texture. We are able to obtain 74% sensitivity and 90% specificity when differentiating high-grade cervical lesions from low-grade lesions and normal tissue. On the same dataset, using Pap tests alone yields a sensitivity of 37% and specificity of 96%, and using HPV test alone gives a 57% sensitivity and 93% specificity. Furthermore, we develop a comprehensive algorithmic framework based on Multimodal Entity Coreference for combining various tests to perform disease classification and diagnosis. When integrating multiple tests, we adopt information gain and gradient-based approaches for learning the relative weights of different tests. In our evaluation, we present a novel algorithm that integrates cervical images, Pap, HPV, and patient age, which yields 83.21% sensitivity and 94.79% specificity, a statistically significant improvement over using any single source of information alone. Dezhao Song, Sharon X. Huang, Joseph Patruno, Hector Muñoz-Avila, Jeff Heflin, L. Rodney Long, Sameer K. Antani |
IEEE Trans. Medical Imaging | 3 |
| 2014 | Instrument Tracking via Online Learning in Retinal Microsurgery
Yeqing Li, Chen Chen 0003, Sharon X. Huang, Junzhou Huang |
MICCAI (1) | 3 |
| 2014 | 3D actin network centerline extraction with multiple active contours
Dimitrios Vavylonis, Sharon X. Huang |
Medical Image Anal. | 3 |
| 2014 | Feature Matching with Affine-Function Transformation ModelsabstractFeature matching is an important problem and has extensive uses in computer vision. However, existing feature matching methods support either a specific or a small set of transformation models. In this paper, we propose a unified feature matching framework which supports a large family of transformation models. We call the family of transformation models the affine-function family, in which all transformations can be expressed by affine functions with convex constraints. In this framework, the goal is to recover transformation parameters for every feature point in a template point set to calculate their optimal matching positions in an input image. Given pairwise feature dissimilarity values between all points in the template set and the input image, we create a convex dissimilarity function for each template point. Composition of such convex functions with any transformation model in the affine-function family is shown to have an equivalent convex optimization form that can be optimized efficiently. Four example transformation models in the affine-function family are introduced to show the flexibility of our proposed framework. Our framework achieves 0.0 percent matching errors for both CMU House and Hotel sequences following the experimental setup in [6]. Hongsheng Li 0001, Sharon X. Huang, Junzhou Huang, Shaoting Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2014 | SalientShape: group saliency in image collections
Ming-Ming Cheng, Niloy J. Mitra, Sharon X. Huang, Shi-Min Hu 0001 |
Vis. Comput. | 3 |
| 2013 | Object Matching Using a Locally Affine Invariant and Linear Programming TechniquesabstractIn this paper, we introduce a new matching method based on a novel locally affine-invariant geometric constraint and linear programming techniques. To model and solve the matching problem in a linear programming formulation, all geometric constraints should be able to be exactly or approximately reformulated into a linear form. This is a major difficulty for this kind of matching algorithm. We propose a novel locally affine-invariant constraint which can be exactly linearized and requires a lot fewer auxiliary variables than other linear programming-based methods do. The key idea behind it is that each point in the template point set can be exactly represented by an affine combination of its neighboring points, whose weights can be solved easily by least squares. Errors of reconstructing each matched point using such weights are used to penalize the disagreement of geometric relationships between the template points and the matched points. The resulting overall objective function can be solved efficiently by linear programming techniques. Our experimental results on both rigid and nonrigid object matching show the effectiveness of the proposed algorithm. Hongsheng Li 0001, Sharon X. Huang, Lei He 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2012 | A hierarchical image clustering cosegmentation frameworkabstractGiven the knowledge that the same or similar objects appear in a set of images, our goal is to simultaneously segment that object from the set of images. To solve this problem, known as the cosegmentation problem, we present a method based upon hierarchical clustering. Our framework first eliminates intra-class heterogeneity in a dataset by clustering similar images together into smaller groups. Then, from each image, our method extracts multiple levels of segmentation and creates connections between regions (e.g. superpixel) across levels to establish intra-image multi-scale constraints. Next we take advantage of the information available from other images in our group. We design and present an efficient method to create inter-image relationships, e.g. connections between image regions from one image to all other images in an image cluster. Given the intra & inter-image connections, we perform a segmentation of the group of images into foreground and background regions. Finally, we compare our segmentation accuracy to several other state-of-the-art segmentation methods on standard datasets, and also demonstrate the robustness of our method on real world data. Hongsheng Li 0001, Sharon X. Huang |
CVPR | 3 |
| 2012 | Simplified Labeling Process for Medical Image Segmentation
Mingchen Gao, Junzhou Huang, Sharon X. Huang, Shaoting Zhang 0001, Dimitris N. Metaxas |
MICCAI (2) | 3 |
| 2011 | Global contrast based salient region detectionabstractAutomatic estimation of salient object regions across images, without any prior assumption or knowledge of the contents of the corresponding scenes, enhances many computer vision and computer graphics applications. We introduce a regional contrast based salient object detection algorithm, which simultaneously evaluates global contrast differences and spatial weighted coherence scores. The proposed algorithm is simple, efficient, naturally multi-scale, and produces full-resolution, high-quality saliency maps. These saliency maps are further used to initialize a novel iterative version of GrabCut, namely SaliencyCut, for high quality unsupervised salient object segmentation. We extensively evaluated our algorithm using traditional salient object detection datasets, as well as a more challenging Internet image dataset. Our experimental results demonstrate that our algorithm consistently outperforms 15 existing salient object detection and segmentation methods, yielding higher precision and better recall rates. We also show that our algorithm can be used to efficiently extract salient object masks from Internet images, enabling effective sketch-based image retrieval (SBIR) via simple shape comparisons. Despite such noisy internet images, where the saliency regions are ambiguous, our saliency guided image retrieval achieves a superior retrieval rate compared with state-of-the-art SBIR methods, and additionally provides important target object region information. Ming-Ming Cheng, Guo-Xin Zhang, Niloy J. Mitra, Sharon X. Huang, Shi-Min Hu 0001 |
CVPR | 4 |
| 2011 | Optimal object matching via convexification and compositionabstractIn this paper, we propose a novel object matching method to match an object to its instance in an input scene image, where both the object template and the input scene image are represented by groups of feature points. We relax each template point's discrete feature cost function to create a convex function that can be optimized efficiently. Such continuous and convex functions with different regularization terms are able to create different convex optimization models handling objects undergoing (i) global transformation, (ii) locally affine transformation, and (iii) articulated transformation. These models can better constrain each template point's transformation and therefore generate more robust matching results. Unlike traditional object or feature matching methods with “hard” node-to-node results, our proposed method allows template points to be transformed to any location in the image plane. Such a property makes our method robust to feature point occlusion or mis-detection. Our extensive experiments demonstrate the robustness and flexibility of our method. Hongsheng Li 0001, Junzhou Huang, Shaoting Zhang 0001, Sharon X. Huang |
ICCV | 4 |
| 2011 | A 3D Laplacian-driven parametric deformable modelabstract3D parametric deformable models have been used to extract volumetric object boundaries and they generate smooth boundary surfaces as results. However, in some segmentation cases, such as cerebral cortex with complex folds and creases, and human lung with high curvature boundary, parametric deformable models often suffer from over-smoothing or decreased mesh quality during model deformation. To address this problem, we propose a 3D Laplacian-driven parametric deformable model with a new internal force. Derived from a Mesh Laplacian, the internal force exerted on each control vertex can be decomposed into two orthogonal vectors based on the vertex's tangential plane. We then introduce a weighting function to control the contributions of the two vectors based on the model mesh's geometry. Deforming the new model is solving a linear system, so the new model can converge very efficiently. To validate the model's performance, we tested our method on various segmentation cases and compared our model with Finite Element and Level Set deformable models. Tian Shen, Sharon X. Huang, Hongsheng Li 0001, Shaoting Zhang 0001, Junzhou Huang |
ICCV | 2 |
| 2011 | Finding VIPs - A visual image persons search using a content property reasoner and web ontologyabstractWe present a semantic based search tool, VIPs, i.e. Visual Image Persons Search, on the domain of VIPs, i.e. very important people. Our tool explores the possibilities of content based image search supported by ontological reasoning. Our framework integrates information from both image processing algorithms and semantic knowledge bases to perform interesting queries that would otherwise be impossible. We describe a novel property reasoner that is able to translate low level image features into semantically relevant object properties. Finally, we demonstrate interesting searches supported by our framework on the domain of people, the majority of whom are movie celebrities, using the properties translated by our system as well as existing ontologies available on the web. Sharon X. Huang, Jeff Heflin |
ICME | 2 |
| 2011 | 3D Segmentation of Rodent Brain Structures Using Hierarchical Shape Priors and Deformable Models
Shaoting Zhang 0001, Junzhou Huang, Mustafa Gökhan Uzunbas, Tian Shen, Foteini Delis, Sharon X. Huang, Nora D. Volkow, Panayotis K. Thanos, Dimitris N. Metaxas |
MICCAI (3) | 6 |
| 2011 | Approximately Global Optimization for Robust Alignment of Generalized ShapesabstractIn this paper, we introduce a novel method to solve shape alignment problems. We use gray-scale "images" to represent source shapes, and propose a novel two-component Gaussian Mixture (GM) distance map representation for target shapes. This asymmetric representation is a flexible image-based representation which is able to represent different kinds of shape data, including continuous contours, unstructured sparse point sets, edge maps, and even gray-scale gradient maps. Using this representation, a new energy function based on a novel two-component Gaussian Mixture distance model is proposed. The new energy function was empirically evaluated to be a more robust shape dissimilarity metric that can be computed efficiently. Such high efficiency is essential for global optimization methods. We adopt and modify one of them, the Particle Swarm Optimization (PSO), to effectively estimate the global optimum of the new energy function. Differently from the original PSO, several new strategies were employed to make the optimization more robust and prevent it from converging prematurely. The overall performance of the proposed framework as well as the properties of each algorithmic component were evaluated and compared with those of some state-of-the-art methods. Extensive experiments and comparison performed on generalized 2D and 3D shape data demonstrate the robustness and effectiveness of the method. Hongsheng Li 0001, Tian Shen, Sharon X. Huang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2011 | Active Volume Models for Medical Image SegmentationabstractIn this paper, we propose a novel predictive model, active volume model (AVM), for object boundary extraction. It is a dynamic "object" model whose manifestation includes a deformable curve or surface representing a shape, a volumetric interior carrying appearance statistics, and an embedded classifier that separates object from background based on current feature information. The model focuses on an accurate representation of the foreground object's attributes, and does not explicitly represent the background. As we will show, however, the model is capable of reasoning about the background statistics thus can detect when is change sufficient to invoke a boundary decision. When applied to object segmentation, the model alternates between two basic operations: 1) deforming according to current region of interest (ROI), which is a binary mask representing the object region predicted by the current model, and 2) predicting ROI according to current appearance statistics of the model. To further improve robustness and accuracy when segmenting multiple objects or an object with multiple parts, we also propose multiple-surface active volume model (MSAVM), which consists of several single-surface AVM models subject to high-level geometric spatial constraints. An AVM's deformation is derived from a linear system based on finite element method (FEM). To keep the model's surface triangulation optimized, surface remeshing is derived from another linear system based on Laplacian mesh optimization (LMO). Thus efficient optimization and fast convergence of the model are achieved by solving two linear systems. Segmentation, validation and comparison results are presented from experiments on a variety of 2-D and 3-D medical images. Tian Shen, Hongsheng Li 0001, Sharon X. Huang |
IEEE Trans. Medical Imaging | 3 |
| 2011 | Markup SVG - An Online Content-Aware Image Abstraction and Annotation ToolabstractSuppose you want to effectively search through millions of images, train an algorithm to perform image and video object recognition, or research the complex patterns and relationships that exist in our visual world. A common and essential component for any of these tasks is a large annotated image dataset. However, obtaining labeled image data is a complex and tedious task that requires methods for annotating and structuring content. Therefore, we developed a comprehensive online tool and data structure, Markup SVG, that simplifies the collection of annotated image data by leveraging state-of-the-art image processing techniques. As the core data structure of our tool, we adopt scalable vector graphics (SVG), an extensible and versatile language built upon XML. Given the extensibility of our framework, we are able to encode low-level image features, high-level semantics, and further define interactions with the data to assist the user with image annotation. We also demonstrate the ability to merge multiple online and offline datasets into our system in an effort to standardize image collection and its data representation. Lastly, we present our modular design; each component acts as a plug-in to our system. We developed several novel components and algorithms to highlight the possibilities of semi-supervised segmentation and automatic annotation within our proposed framework. Further, our modular design provides the necessary capabilities to incorporate future image features, methods, or algorithms. Our results show that our tool is able to greatly simplify the process of obtaining large annotated image collections in an online collaborative platform. Sharon X. Huang, Gang Tan |
IEEE Trans. Multim. | 2 |
| 2010 | A parallel cellular automata with label priors for interactive brain tumor segmentationabstractWe present a novel method for 3D brain tumor volume segmentation based on a parallel cellular automata framework. Our method incorporates prior label knowledge gathered from user seed information to influence the cellular automata decision rules. Our proposed method is able to segment brain tumor volumes quickly and accurately using any number of label classifications. Exploiting the inherent parallelism of our algorithm, we adopt this method to the Graphics Processing Unit (GPU). Additionally, we introduce the concept of individual label strength maps to visualize the improvements of our method. As we demonstrate in our quantitative and qualitative results, the key benefits of our system are accuracy, robustness to complex structures, and speed. We compute segmentations nearly 45× faster than conventional CPU methods, enabling user feedback at interactive rates. Tian Shen, Sharon X. Huang |
CBMS | 3 |
| 2010 | Object matching with a locally affine-invariant constraintabstractIn this paper, we present a new object matching algorithm based on linear programming and a novel locally affine-invariant geometric constraint. Previous works have shown possible ways to solve the feature and object matching problem by linear programming techniques. To model and solve the matching problem in a linear formulation, all geometric constraints should be able to be exactly or approximately reformulated into a linear form. This is a major difficulty for this kind of matching algorithms. We propose a novel locally affine-invariant constraint which can be exactly linearized and requires a lot fewer auxiliary variables than the previous work does. The key idea behind it is that each point can be exactly represented by an affine combination of its neighboring points, whose weights can be solved easily by least squares. The resulting overall objective function can then be solved efficiently by linear programming techniques. Our experimental results on both rigid and non-rigid object matching show the advantages of the proposed algorithm. Hongsheng Li 0001, Sharon X. Huang, Lei He 0001 |
CVPR | 3 |
| 2010 | Actin Filament Segmentation Using Spatiotemporal Active-Surface and Active-Contour Models
Hongsheng Li 0001, Tian Shen, Dimitrios Vavylonis, Sharon X. Huang |
MICCAI (1) | 4 |
| 2010 | RepFinder: finding approximately repeated scene elements for image editingabstractRepeated elements are ubiquitous and abundant in both manmade and natural scenes. Editing such images while preserving the repetitions and their relations is nontrivial due to overlap, missing parts, deformation across instances, illumination variation, etc. Manually enforcing such relations is laborious and error-prone. We propose a novel framework where user scribbles are used to guide detection and extraction of such repeated elements. Our detection process, which is based on a novel boundary band method, robustly extracts the repetitions along with their deformations. The algorithm only considers the shape of the elements, and ignores similarity based on color, texture, etc. We then use topological sorting to establish a partial depth ordering of overlapping repeated instances. Missing parts on occluded instances are completed using information from other instances. The extracted repeated instances can then be seamlessly edited and manipulated for a variety of high level tasks that are otherwise difficult to perform. We demonstrate the versatility of our framework on a large set of inputs of varying complexity, showing applications to image rearrangement, edit transfer, deformation propagation, and instance replacement. Ming-Ming Cheng, Niloy J. Mitra, Sharon X. Huang, Shi-Min Hu 0001 |
ACM Trans. Graph. | 4 |
| 2009 | Global optimization for alignment of generalized shapesabstractIn this paper, we introduce a novel algorithm to solve global shape registration problems. We use gray-scale “images” to represent source shapes, and propose a novel two-component Gaussian Mixtures (GM) distance map representation for target shapes. Based on this flexible asymmetric image-based representation, a new energy function is defined. It proves to be a more robust shape dissimilarity metric that can be computed efficiently. Such high efficiency is essential for global optimization methods. We adopt one of them, the Particle Swarm Optimization (PSO), to effectively estimate the global optimum of the new energy function. Experiments and comparison performed on generalized shape data including continuous shapes, unstructured sparse point sets, and gradient maps, demonstrate the robustness and effectiveness of the algorithm. Hongsheng Li 0001, Tian Shen, Sharon X. Huang |
CVPR | 3 |
| 2009 | Active volume models for 3D medical image segmentationabstractIn this paper, we propose a novel predictive model for object boundary, which can integrate information from any sources. The model is a dynamic “object” model whose manifestation includes a deformable surface representing shape, a volumetric interior carrying appearance statistics, and an embedded classifier that separates object from background based on current feature information. Unlike Snakes, Level Set, Graph Cut, MRF and CRF approaches, the model is “self-contained” in that it does not model the background, but rather focuses on an accurate representation of the foreground object's attributes. As we will show, however, the model is capable of reasoning about the background statistics thus can detect when is change sufficient to invoke a boundary decision. The shape of the 3D model is considered as an elastic solid, with a simplex-mesh (i.e. finite element triangulation) surface made of thousands of vertices. Deformations of the model are derived from a linear system that encodes external forces from the boundary of a Region of Interest (ROI), which is a binary mask representing the object region predicted by the current model. Efficient optimization and fast convergence of the model are achieved using the Finite Element Method (FEM). Other advantages of the model include the ease of dealing with topology changes and its ability to incorporate human interactions. Segmentation and validation results are presented for experiments on noisy 3D medical images. Tian Shen, Hongsheng Li 0001, Sharon X. Huang |
CVPR | 4 |
| 2009 | Learning with dynamic group sparsityabstractThis paper investigates a new learning formulation called dynamic group sparsity. It is a natural extension of the standard sparsity concept in compressive sensing, and is motivated by the observation that in some practical sparse data the nonzero coefficients are often not random but tend to be clustered. Intuitively, better results can be achieved in these cases by reasonably utilizing both clustering and sparsity priors. Motivated by this idea, we have developed a new greedy sparse recovery algorithm, which prunes data residues in the iterative process according to both sparsity and group clustering priors rather than only sparsity as in previous methods. The proposed algorithm can recover stably sparse data with clustering trends using far fewer measurements and computations than current state-of-the-art algorithms with provable guarantees. Moreover, our algorithm can adaptively learn the dynamic group structure and the sparsity number if they are not available in the practical applications. We have applied the algorithm to sparse recovery and background subtraction in videos. Numerous experiments with improved performance over previous methods further validate our theoretical proofs and the effectiveness of the proposed algorithm. Junzhou Huang, Sharon X. Huang, Dimitris N. Metaxas |
ICCV | 2 |
| 2009 | Document Analysis Support for the Manual Auditing of ElectionsabstractRecent developments have resulted in dramatic changes in the way elections are conducted, both in the United States and around the world. Well-publicized flaws in the security of electronic voting systems have led to a push for the use of verifiable paper records in the election process. In this paper, we describe the application of document analysis techniques to facilitate the manual auditing of elections,both to assure the reliability of the final outcome as well as to help reconcile the differences that may arise between repeated scans of the same ballot. We show how techniques developed for document duplicate detection can be applied to this problem, and present experimental results that demonstrate the efficacy of our approach. Related issues concerning machine support for the auditing of elections are also discussed. Daniel P. Lopresti, Xiang Sean Zhou, Sharon X. Huang, Gang Tan |
ICDAR | 3 |
| 2009 | Actin Filament Tracking Based on Particle Filters and Stretching Open Active Contour Models
Hongsheng Li 0001, Tian Shen, Dimitrios Vavylonis, Sharon X. Huang |
MICCAI (1) | 4 |
| 2009 | 3D Medical Image Segmentation by Multiple-Surface Active Volume Models
Tian Shen, Sharon X. Huang |
MICCAI (1) | 2 |
| 2008 | Commentary Paper on "Video Surveillance for Biometrics: Long-Range Multi-biometric System"abstractThis paper introduces a long-range multi-biometric system by integrating hierarchical camera system design, human detection and tracking, face detection and pose estimation, NIR laser illumination, and iris acquisition by a PTU and long focal length zoom lens. The system proposed by the paper is sound. Hierarchical design and error-correction steps are introduced in an attempt to make the system robust in real world scenarios. The main concern is two folds: the low resolution of the acquired iris images, and the processing time between frames. Low resolution iris images may lead to significantly decreased accuracy in biometric-based recognition. And the long processing time (due to detection, tracking, etc.) may prevent the system from being applicable to acquiring iris images of moving subjects. Sharon X. Huang |
AVSS | 1 |
| 2008 | Web-Based Multi-Observer Segmentation Evaluation ToolabstractMulti-observer segmentation evaluation is useful in the imaging community. We have developed web-based software for automatic performance evaluation of multiple image segmentations which is based on the Baysian decision framework. It computes a probabilistic estimate of the true segmentation (ground truth map) and performance measures for the individual segmentations (sensitivity and specificity). The strength of the tool is that it integrates the two kinds of prior knowledge of segmentations: the truth prior (the prior probability) and the observer prior (the performance measures of observers), which can generate more accurate evaluations. Yaoyao Zhu, Sharon X. Huang, Daniel P. Lopresti, L. Rodney Long, Sameer K. Antani, Zhiyun Xue, George R. Thoma |
CBMS | 2 |
| 2008 | Simultaneous image transformation and sparse representation recoveryabstractSparse representation in compressive sensing is gaining increasing attention due to its success in various applications. As we demonstrate in this paper, however, image sparse representation is sensitive to image plane transformations such that existing approaches can not reconstruct the sparse representation of a geometrically transformed image. We introduce a simple technique for obtaining transformation-invariant image sparse representation. It is rooted in two observations: 1) if the aligned model images of an object span a linear subspace, their transformed versions with respect to some group of transformations can still span a linear subspace in a higher dimension; 2) if a target (or test) image, aligned with the model images, lives in the above subspace, its pre-alignment versions would get closer to the subspace after applying estimated transformations with more and more accurate parameters. These observations motivate us to project a potentially unaligned target image to random projection manifolds defined by the model images and the transformation model. Each projection is then separated into the aligned projection target and a residue due to misalignment. The desired aligned projection target is then iteratively optimized by gradually diminishing the residue. In this framework, we can simultaneously recover the sparse representation of a target image and the image plane transformation between the target and the model images. We have applied the proposed methodology to two applications: face recognition, and dynamic texture registration. The improved performance over previous methods that we obtain demonstrates the effectiveness of the proposed approach. Junzhou Huang, Sharon X. Huang, Dimitris N. Metaxas |
CVPR | 2 |
| 2008 | Tag Separation in Cardiac Tagged MRI
Junzhou Huang, Sharon X. Huang, Dimitris N. Metaxas, Leon Axel |
MICCAI (2) | 3 |
| 2008 | Active Volume Models with Probabilistic Object Boundary Prediction Module
Tian Shen, Yaoyao Zhu, Sharon X. Huang, Junzhou Huang, Dimitris N. Metaxas, Leon Axel |
MICCAI (1) | 3 |
| 2008 | Metamorphs: Deformable Shape and Appearance ModelsabstractThis paper presents a new deformable modeling strategy that is aimed at integrating shape and appearance in a unified space. If we think of traditional deformable models as "active contours" or "evolving curve fronts," the new deformable shape and appearance models that we propose are "deforming disks or volumes." Each model not only has boundary shape but also interior appearance. The model shape is implicitly embedded in a higher dimensional space of distance transforms and is thus represented by a distance map "image." This way, both the shape and the appearance of the model are defined in the pixel space. A common deformation scheme, that is, the free-form deformations (FFDs), parameterizes warping deformations of the volumetric space in which the model is embedded, hence simultaneously deforming both model boundary and interior. When applied to segmentation, a metamorphs model can be initialized by covering a seed region far from the object boundary, and then the model efficiently evolves and converges to an optimal solution. The model dynamics are derived in a unified variational framework that consists of edge-based and region-based energy terms, both of which are differentiable with respect to the common set of FFD parameters. As the model deforms, its interior appearance statistics are adaptively learned and, then, toward the next-step deformation, the model examines not only edge information but also its exterior region statistics to ensure that it only expands to new territory with consistent appearance statistics. The Metamorphs formulation also allows natural merging and competition of multiple models. We demonstrate the robustness of metamorphs by using both natural and medical images that have high noise levels, intensity inhomogeneity, and complex texture. Sharon X. Huang, Dimitris N. Metaxas |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2007 | Robust Click-Point Linking: Matching Visually Dissimilar Local RegionsabstractThis paper presents robust click-point linking: a novel localized registration framework that allows users to interactively prescribe where the accuracy has to be high. By emphasizing locality and interactivity, our solution is faithful to how the registration results are used in practice. Given a user-specified point, the click-point linking provides a single point-wise correspondence between a data pair. In order to link visually dissimilar local regions, a correspondence is sought by using only geometrical context without comparing the local appearances. Our solution is formulated as a maximum likelihood estimation (MLE) without estimating a domain transformation explicitly. A spatial likelihood of Gaussian mixture form is designed to capture geometrical configurations between the point-of-interest and a hierarchy of global-to-local 3D landmarks that are detected using machine learning and entropy based feature detectors. A closed-form formula is derived to specify each Gaussian component by exploiting geometric in-variances under specific group of domain transformation via RANSAC-like random sampling. A mean shift algorithm is applied to robustly and efficiently solve the local MLE problem, replacing the standard consensus step of the RANSAC. Two transformation groups of pure translation and scaling/translation are considered in this paper. We test feasibility of the proposed approach with 16 pairs of whole-body CT data, demonstrating the effectiveness. Kazunori Okada, Sharon X. Huang |
CVPR | 2 |
| 2007 | Optimization and Learning for Registration of Moving Dynamic TexturesabstractWe address the problem of registering a sequence of images in a moving dynamic texture video. This involves optimization with respect to camera motion, the average image, and the dynamic texture model. This problem is highly ill-posed and almost impossible to have good solutions without priors. In this paper, we introduce powerful priors for this problem, based on two simple observations: 1) registration should simplify the dynamic texture model while preserving all useful information. It motivates us to compute a prior for the dynamic texture by marginalizing over specific dynamics in the space of all stable auto-regressive sequences; 2) the statistics of derivative filter responses in the average image can be significantly changed by registration, and better registration should lead to a sharper average image. This offers us the prior of requiring the derivative distribution of the estimated average image to be close to that learned from the input image sequence. With these priors, a new registration approach is proposed by marginalizing over the "nuisance" variables under a Bayesian framework. And superior motion estimation results are obtained by jointly optimizing over the registration parameters, the average image, and the dynamic texture model. Experimental results on real video sequences of moving dynamic textures show convincing performance of the proposed approach. Junzhou Huang, Sharon X. Huang, Dimitris N. Metaxas |
ICCV | 2 |
| 2007 | Adaptive Metamorphs Model for 3D Medical Image Segmentation
Junzhou Huang, Sharon X. Huang, Dimitris N. Metaxas, Leon Axel |
MICCAI (1) | 2 |
| 2006 | Patch-Based Texture Edges and Segmentation
Lior Wolf, Sharon X. Huang, Ian Martin, Dimitris N. Metaxas |
ECCV (2) | 2 |
| 2006 | Hybrid Deformable Models for Medical Segmentation and RegistrationabstractDeformable models have had great successes over the past 20 years in medical applications. We have recently developed new classes of deformable models which we term hybrid deformable models to automate the model initialization process and make improvements in segmentation and registration. In this paper we present several hybrid deformable methods we have been developing for segmentation and registration. These methods include metamorphs, a novel shape and texture integration deformable model framework and the integration of deformable models with graphical models and learning methods. We first present a framework for the robust segmentation and tracking of the heart from tagged MRI images and second applications involving brain tumor segmentation as well as brain and cardiac shape registration Dimitris N. Metaxas, Sharon X. Huang, Rui Huang 0001, Ting Chen 0001, Leon Axel |
ICARCV | 3 |
| 2006 | Automatic Hot Spot Detection and Segmentation in Whole Body FDG-PET ImagesabstractWe present a system for automatic hot spots detection and segmentation in whole body FDG-PET images. The main contribution of our system is threefold. First, it has a novel body-section labeling module based on spatial hidden-Markov models (HMM); this allows different processing policies to be applied in different body sections. Second, the competition diffusion (CD) segmentation algorithm, which takes into account body-section information, converts the binary thresholding results to probabilistic interpretation and detects hot-spot region candidates. Third, a recursive intensity mode-seeking algorithm finds hot spot centers efficiently, and given these centers, a clinically meaningful protocol is proposed to accurately quantify hot spot volumes. Experimental results show that our system works robustly despite the large variations in clinical PET images. Haiying Guan, Toshiro Kubota, Sharon X. Huang, Xiang Sean Zhou, Matthew Turk 0001 |
ICIP | 3 |
| 2006 | Shape Registration in Implicit Spaces Using Information Theory and Free Form DeformationsabstractWe present a novel, variational and statistical approach for shape registration. Shapes of interest are implicitly embedded in a higher-dimensional space of distance transforms. In this implicit embedding space, registration is formulated in a hierarchical manner: the Mutual Information criterion supports various transformation models and is optimized to perform global registration; then, a B-spline-based Incremental Free Form Deformations (IFFD) model is used to minimize a Sum-of-Squared-Differences (SSD) measure and further recover a dense local nonrigid registration field. The key advantage of such framework is twofold: (1) it naturally deals with shapes of arbitrary dimension (2D, 3D, or higher) and arbitrary topology (multiple parts, closed/open) and (2) it preserves shape topology during local deformation and produces local registration fields that are smooth, continuous, and establish one-to-one correspondences. Its invariance to initial conditions is evaluated through empirical validation, and various hard 2D/3D geometric shape registration examples are used to show its robustness to noise, severe occlusion, and missing parts. We demonstrate the power of the proposed framework using two applications: one for statistical modeling of anatomical structures, another for 3D face scan registration and expression tracking. We also compare the performance of our algorithm with that of several other well-known shape registration algorithms. Sharon X. Huang, Nikos Paragios, Dimitris N. Metaxas |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2005 | Efficient Learning by Combining Confidence-Rated Classifiers to Incorporate Unlabeled Medical Data
Weijun He, Sharon X. Huang, Dimitris N. Metaxas, Xiaoyou Ying |
MICCAI | 2 |
| 2004 | MetaMorphs: Deformable Shape and Texture Models
Sharon X. Huang, Dimitris N. Metaxas, Ting Chen 0001 |
CVPR (1) | 1 |
| 2004 | Learning Coupled Prior Shape and Appearance Models for Segmentation
Sharon X. Huang, Dimitris N. Metaxas |
MICCAI (1) | 1 |
| 2004 | High Resolution Acquisition, Learning and Transfer of Dynamic 3D Facial ExpressionsabstractAbstract Synthesis and re‐targeting of facial expressions is central to facial animation and often involves significant manual work in order to achieve realistic expressions, due to the difficulty of capturing high quality dynamic expression data. In this paper we address fundamental issues regarding the use of high quality dense 3‐D data samples undergoing motions at video speeds, e.g. human facial expressions. In order to utilize such data for motion analysis and re‐targeting, correspondences must be established between data in different frames of the same faces as well as between different faces. We present a data driven approach that consists of four parts: 1) High speed, high accuracy capture of moving faces without the use of markers, 2) Very precise tracking of facial motion using a multi‐resolution deformable mesh, 3) A unified low dimensional mapping of dynamic facial motion that can separate expression style, and 4) Synthesis of novel expressions as a combination of expression styles. The accuracy and resolution of our method allows us to capture and track subtle expression details. The low dimensional representation of motion data in a unified embedding for all the subjects in the database allows for learning the most discriminating characteristics of each individual's expressions as that person's “expression style”. Thus new expressions can be synthesized, either as dynamic morphing between individuals, or as expression transfer from a source face to a target face, as demonstrated in a series of experiments. Categories and Subject Descriptors (according to ACM CCS): I.3.7 [Computer Graphics]: Animation; I.3.5 [Computer Graphics]: Curve, surface, solid, and object representations; I.3.3 [Computer Graphics]: Digitizing and scanning; I.2.10 [Artificial intelligence]: Motion ; I.2.10 [Artificial intelligence]: Representations, data structures, and transforms; I.2.10 [Artificial intelligence]: Shape; I.2.6 [Artificial intelligence]: Concept learning Yang Wang 0001, Sharon X. Huang, Chan-Su Lee, Song Zhang 0002, Dimitris Samaras, Dimitris N. Metaxas, Ahmed M. Elgammal, Peisen Huang |
Comput. Graph. Forum | 2 |
| 2003 | Computing Layered Surface Representations: An Algorithm for Detecting and Separating Transparent OverlaysabstractThe biological visual system possesses the ability to compute layered surface representations, in which one surface is represented as being viewed through another. This ability is remarkable because, in scenes involving transparency, the link between surface topology and image topology is greatly complicated by the collapse of the photometric contributions of two distinct surfaces onto image intensity. Previous analysis of transparency has focused largely on the role of different kinds of junctions. Although junctions are important, they are not sufficient to predict layered surface structure. We present an algorithm that propagates local junction information by searching for chains of polarity-preserving junctions with consistent 'sidedness', and then propagates the transparency labeling into interior regions. The algorithm outputs a layered representation specifying: (i) the distinct surfaces, (ii) their depth ordering, and (iii) their surface attributes. We demonstrate the results of the algorithm on a number of images - both synthetic and real. We end by considering implications for related domains, such as shading. Manish Singh 0001, Sharon X. Huang |
CVPR (2) | 2 |
| 2003 | Establishing Local Correspondences towards Compact Representations of Anatomical Structures
Sharon X. Huang, Nikos Paragios, Dimitris N. Metaxas |
MICCAI (2) | 1 |