VLDB 2026 Research / reviewers in the wild / expert
Shi-Min Hu 0001
dblp:h/ShiMinHu · also Shimin Hu 0001
· DBLP profile ↗
332ranked-venue papers
50as first author
83since 2021 · last 2026
0000-0001-7507-6542ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 256 · 33 first-author · 61 since 2021Artificial intelligence and machine learning · 34 · 16 since 2021Applied, interdisciplinary, general and emerging computing · 32 · 17 first-author · 6 since 2021Software engineering, systems software and programming languages · 18 · 3 since 2021Systems, architecture and hardware · 15 · 5 since 2021Security and privacy · 7 · 3 since 2021Human-computer interaction and ubiquitous computing · 6 · 2 first-author · 3 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Foreword to the Special Section on CAD/Graphics 2023
Joaquim Jorge 0001, Shi-Min Hu 0001, Paul L. Rosin, Yiyu Cai |
Comput. Graph. | 2 |
| 2026 | Spatial Multiple Importance Sampling for Real-Time Irradiance ProbesabstractReal-time global illumination rendered with low variance remains a persistent challenge. Many engines employ irradiance probes as a relatively cheap technique, but constrained computational budgets often lead to flickering artifacts in rendered images. In this paper, we propose spatial multiple importance sampling, which reuses ray-surface intersection data among real-time irradiance probes to significantly reduce flicker caused by variance, enabling efficient computation in complex scenes under limited ray tracing budgets. Moreover, our approach incorporates a probe selection mechanism to enhance reuse efficiency and a visibility estimation method to mitigate bias. Experimental results demonstrate that our method significantly reduces variance at a fixed ray tracing cost, delivering high-quality, stable outputs in real-time scenarios. Tuo Chen, Zi-Heng Zhou, Lingqi Yan 0001, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2026 | SeparateGen: Semantic Component-Based 3D Character Generation From Single ImagesabstractCreating detailed 3D characters from a single image remains challenging due to the difficulty in separating semantic components during generation. Existing methods often produce entangled meshes with poor topology, hindering downstream applications like rigging and animation. We introduce SeparateGen, a novel framework that generates high-quality 3D characters by explicitly reconstructing them as distinct semantic components (e.g., body, clothing, hair, shoes) from a single, arbitrary-pose image. SeparateGen first leverages a multi-view diffusion model to generate consistent multi-view images in a canonical A-pose. Then, a novel component-aware reconstruction model, SC-LRM, conditioned on these multi-view images, adaptively decomposes and reconstructs each component with high fidelity. To train and evaluate SeparateGen, we contribute SC-Anime, the first large-scale dataset of 7,580 anime-style 3D characters with detailed component-level annotations. Extensive experiments demonstrate that SeparateGen significantly outperforms state-of-the-art methods in both reconstruction quality and multi-view consistency. Furthermore, our component-based approach effectively resolves mesh entanglement issues, enabling seamless rigging and asset reuse. SeparateGen thus represents a step towards generating high-quality, application-ready 3D characters from a single image. The SC-Anime dataset and our code will be publicly released. Dong-Yang Li, Yi-Long Liu, Zi-Xian Liu, Yan-Pei Cao 0001, Menghao Guo 0001, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2025 | High-Accuracy Fractured Object Reassembly Under Arbitrary Poses
Qun-Ce Xu, Yan-Pei Cao 0001, Weihao Cheng 0002, Tai-Jiang Mu, Ying Shan, Yongliang Yang 0002, Shi-Min Hu 0001 |
CVM (2) | 7 |
| 2025 | Adaptive Parameter Selection for Tuning Vision-Language ModelsabstractVision-language models (VLMs) like CLIP have been widely used in various specific tasks. Parameter-efficient fine-tuning (PEFT) methods, such as prompt and adapter tuning, have become key techniques for adapting these models to specific domains. However, existing approaches rely on prior knowledge to manually identify the locations requiring fine-tuning. Adaptively selecting which parameters in VLMs should be tuned remains unexplored. In this paper, we propose CLIP with Adaptive Selective Tuning (CLIP-AST), which can be used to automatically select critical parameters in VLMs for fine-tuning for specific tasks. It opportunely leverages the adaptive learning rate in the optimizer and improves model performance without extra parameter overhead. We conduct extensive experiments on 13 benchmarks, such as ImageNet, Food101, Flowers102, etc, with different settings, including few-shot learning, base-to-novel class generalization, and out-of-distribution. The results show that CLIP-AST consistently outperforms the original CLIP model as well as its variants and achieves state-of-the-art (SOTA) performance in all cases. For example, with the 16-shot learning, CLIP-AST surpasses GraphAdapter and PromptSRC by 3.56% and 2.20% in average accuracy on 11 datasets, respectively. Code will be publicly available. Yi Zhang 0099, Yi-Xuan Deng, Menghao Guo 0001, Shi-Min Hu 0001 |
CVPR | 4 |
| 2025 | RBench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning EvaluationabstractReasoning stands as a cornerstone of intelligence, enabling the synthesis of existing knowledge to solve complex problems. Despite remarkable progress, existing reasoning benchmarks often fail to rigorously evaluate the nuanced reasoning capabilities required for complex, real-world problemsolving, particularly in multi-disciplinary and multimodal contexts. In this paper, we introduce a graduate-level, multi-disciplinary, EnglishChinese benchmark, dubbed as Reasoning Bench (RBench), for assessing the reasoning capability of both language and multimodal models. RBench spans 1,094 questions across 108 subjects for language model evaluation and 665 questions across 83 subjects for multimodal model testing. These questions are meticulously curated to ensure rigorous difficulty calibration, subject balance, and cross-linguistic alignment, enabling the assessment to be an Olympiad-level multidisciplinary benchmark. We evaluate many models such as o1, GPT-4o, DeepSeek-R1, etc. Experimental results indicate that advanced models perform poorly on complex reasoning, especially multimodal reasoning. Even the top-performing model OpenAI o1 achieves only 53.2% accuracy on our multimodal evaluation. Data and code are made publicly available athttps://evalmodels.github.io/rbench/ Menghao Guo 0001, Yi Zhang 0099, Jiaxi Song, Haoyang Peng, Yi-Xuan Deng, Xinzhi Dong, Kiyohiro Nakayama, Zhengyang Geng, Chen Wang 0049, Bolin Ni, Yongming Rao, Houwen Peng, Han Hu 0001, Gordon Wetzstein, Shi-Min Hu 0001 |
ICML | 17 |
| 2025 | RBench-V: A Primary Assessment for Visual Reasoning Models with Multimodal OutputsabstractThe rapid advancement of native multi-modal models and omni-models, exemplified by GPT-4o, Gemini and o3 with their capability to process and generate content across modalities such as text and images, marks a significant milestone in the evolution of intelligence. Systematic evaluation of their multi-modal output capabilities in visual thinking process (a.k.a., multi-modal chain of thought, M-CoT) becomes critically important. However, existing benchmarks for evaluating multi-modal models primarily focus on assessing multi-modal inputs and text-only reasoning process while neglecting the importance of reasoning through multi-modal outputs. In this paper, we present a benchmark, dubbed as RBench-V, designed to assess models’ vision-indispensable reasoning. To conduct RBench-V, we carefully hand-pick 803 questions covering math, physics, counting and games. Unlike problems in previous benchmarks, which typically specify certain input modalities, RBench-V presents problems centered on multi-modal outputs, which require image manipulation, such as generating novel images and constructing auxiliary lines to support reasoning process. We evaluate numerous open- and closed-source models on RBench-V, including o3, Gemini 2.5 pro, Qwen2.5-VL, etc. Even the best-performing model, o3, achieves only 25.8% accuracy on RBench-V, far below the human score of 82.3%, which shows current models struggle to leverage multi-modal reasoning. Data and code are available at https://evalmodels.github.io/rbenchv. Menghao Guo 0001, Xuanyu Chu, Qianrui Yang, Zhe-Han Mo, Yiqing Shen 0005, Pei-lin Li, Xinjie Lin 0001, Jinnian Zhang, Xin-Sheng Chen, Yi Zhang 0099, Kiyohiro Nakayama, Zhengyang Geng, Houwen Peng, Han Hu 0001, Shi-Min Hu 0001 |
NeurIPS | 15 |
| 2025 | SDLKF: Signed Distance Linear Kernel Function for surface reconstruction
Haoxiang Chen 0004, Xiao-Lei Li, Tai-Jiang Mu, Qun-Ce Xu, Shi-Min Hu 0001 |
Comput. Graph. | 5 |
| 2025 | FastMAE: Efficient Masked Autoencoder with Offline TokenizerabstractMasked autoencoders (MAEs) have recently achieved great success in computer vision. They can automatically extract representations from unlabeled data and improve the performance of various downstream tasks. However, training an MAE model requires substantial resources, which limits their accessibility to many academic institutions: often laboratories in universities lack the necessary resources. This issue significantly hinders the development of this field. In this paper, we propose FastMAE, an efficient MAE approach. Inspired by the idea of offline tokenizers in natural language processing, FastMAE presents a novel way to build an offline vision tokenizer, which can provide high-level semantics in an efficient way. Benefiting from the offline tokenizer, FastMAE becomes an efficient vision learner. Our experiments demonstrate that FastMAE can achieve 83.6% accuracy with ViT-B in only 18.8 h on 8 NVIDIA Tesla-V100 GPUs, which is 31.3× faster than the original MAE, providing a resource friendly baseline for the computer vision community. Moreover, it also achieves comparable performance to state-of-the-art methods. We hope our research will attract more people to engage in MAE-related research and that we can advance its development together. Menghao Guo 0001, Chen Wang 0049, Wei Liu 0005, Shi-Min Hu 0001 |
Comput. Vis. Media | 4 |
| 2025 | Diffusion Models for 3D Generation: A SurveyabstractDenoising diffusion models have demonstrated tremendous success in modeling data distributions and synthesizing high-quality samples. In the 2D image domain, they have become the state-of-the-art and are capable of generating photo-realistic images with high controllability. More recently, researchers have begun to explore how to utilize diffusion models to generate 3D data, as doing so has more potential in real-world applications. This requires careful design choices in two key ways: identifying a suitable 3D representation and determining how to apply the diffusion process. In this survey, we provide the first comprehensive review of diffusion models for manipulating 3D content, including 3D generation, reconstruction, and 3D-aware image synthesis. We classify existing methods into three major categories: 2D space diffusion with pretrained models, 2D space diffusion without pretrained models, and 3D space diffusion. We also summarize popular datasets used for 3D generation with diffusion models. Along with this survey, we maintain a repository https://github.com/cwchenwang/awesome-3d-diffusion to track the latest relevant papers and codebases. Finally, we pose current challenges for diffusion models for 3D generation, and suggest future research directions. Chen Wang 0049, Hao-Yang Peng, Ying-Tian Liu, Jiatao Gu, Shi-Min Hu 0001 |
Comput. Vis. Media | 5 |
| 2025 | Remote Sensing Tuning: A SurveyabstractLarge models have accelerated the development of intelligent interpretation in remote sensing. Many remote sensing foundation models (RSFM) have emerged in recent years, sparking a new wave of deep learning in this field. Fine-tuning techniques serve as a bridge between remote sensing downstream tasks and advanced foundation models. As RSFMs become more powerful, fine-tuning techniques are expected to lead the next research frontier in numerous critical remote sensing applications. Advanced fine-tuning techniques can reduce the data and computational resource requirements during the downstream adaptation process. Current fine-tuning techniques for remote sensing are still in their early stages, leaving a large space for optimization and application. To elucidate the current development and future trends of remote sensing fine-tuning techniques, this survey offers a comprehensive overview of recent research. Specifically, this survey summarizes the applications and innovations of each work and categorizes recent remote sensing fine-tuning techniques into six types: adapter-based, prompt-based, reparameterization-based, hybrid methods, partial tuning, and improved tuning. In the final section, this survey suggests nine areas worth exploring in this field. Remote sensing fine-tuning methods in this survey can be found at https://github.com/DongshuoYin/Remote-Sensing-Tuning-A-Survey. Dongshuo Yin, Ting-Feng Zhao, Deng-Ping Fan, Shutao Li 0001, Bo Du 0001, Xian Sun 0001, Shi-Min Hu 0001 |
Comput. Vis. Media | 7 |
| 2025 | Preface
Shi-Min Hu 0001, Piotr Didyk, Junhui Hou |
J. Comput. Sci. Technol. | 1 |
| 2025 | Implicit Bonded Discrete Element Method with Manifold OptimizationabstractThis article proposes a novel simulation approach that combines implicit integration with the Bonded Discrete Element Method (BDEM) to achieve faster, more stable, and more accurate fracture simulation. The new method leverages the efficiency of implicit schemes in dynamic simulation and the versatility of BDEM in fracture modeling. Specifically, an optimization-based integrator for BDEM is introduced and combined with a manifold optimization approach to accelerate the solution process of the quaternion-constrained system. Our comparative experiments indicate that our method offers better scale consistency and more realistic collision effects than finite element method and material point method fragmentation approaches. Additionally, our method achieves a computational speedup of 2.1 to 9.8 times over explicit BDEM methods. Jia-Ming Lu, Geng-Chen Cao, Chenfeng Li, Shi-Min Hu 0001 |
ACM Trans. Graph. | 4 |
| 2025 | Reliable Iterative Dynamics: A Versatile Method for Fast and Robust SimulationabstractSimulating stiff materials has long posed formidable challenges for traditional physics-based solvers. Explicit time integration schemes demand prohibitively small time steps, while implicit methods necessitate an excessive number of iterations to converge, often yielding visually objectionable transient configurations in the early iterations, severely limiting their real-time applicability. Position-based dynamics techniques can efficiently simulate stiff constraints but are inherently restricted to constraint-based formulations, curtailing their versatility. We present “Reliable Iterative Dynamics” (RID), a novel iterative solver that introduces a dual descent framework with theoretical guarantees for visual reliability at each iteration, while maintaining fast and stable convergence even for extremely stiff systems. Our core innovation is an iterative method that circumvents the need for numerous iterations or small time steps to handle stiff materials robustly. Experimental evaluations demonstrate our method’s ability to handle a wide range of materials, from soft to infinitely rigid, while producing visually reliable results even with large time steps and minimal iterations. The versatile formulation allows seamless integration with diverse simulation paradigms like the finite element method, material point method, smoothed particle hydrodynamics, and incremental potential contact for applications ranging from elastic body simulations to fluids and collision handling. Jia-Ming Lu, Shi-Min Hu 0001 |
ACM Trans. Graph. | 2 |
| 2025 | Fast Galerkin Multigrid Method for Unstructured MeshesabstractWe present a novel multigrid solver framework that significantly advances the efficiency of physical simulation for unstructured meshes. While multi-grid methods theoretically offer linear scaling, their practical implementation for deformable body simulations faces substantial challenges, particularly on GPUs. Our framework achieves up to 6.9× speedup over traditional methods through an innovative combination of matrix-free vertex block Jacobi smoothing with a Full Approximation Scheme (FAS), enabling both piecewise constant and linear Galerkin formulations without the computational burden of dense coarse matrices. Our approach demonstrates superior performance across varying mesh resolutions and material stiffness values, maintaining consistent convergence even under extreme deformations and challenging initial configurations. Comprehensive evaluations against state-of-the-art methods confirm our approach achieves lower simulation error with reduced computational cost, enabling simulation of tetrahedral meshes with over one million vertices at approximately one frame per second on modern GPUs. Jia-Ming Lu, Tailing Yuan, Zhe-Han Mo, Shi-Min Hu 0001 |
ACM Trans. Graph. | 4 |
| 2025 | One Model to Rig Them All: Diverse Skeleton Rigging with UniRigabstractThe rapid evolution of 3D content creation, encompassing both AI-powered methods and traditional workflows, is driving an unprecedented demand for automated rigging solutions that can keep pace with the increasing complexity and diversity of 3D models. We introduce UniRig , a novel, unified framework for automatic skeletal rigging that leverages the power of large autoregressive models and a bone-point cross-attention mechanism to generate both high-quality skeletons and skinning weights. Unlike previous methods that struggle with complex or non-standard topologies, UniRig accurately predicts topologically valid skeleton structures thanks to a new Skeleton Tree Tokenization method that efficiently encodes hierarchical relationships within the skeleton. To train and evaluate UniRig, we present Rig-XL , a new large-scale dataset of over 14,000 rigged 3D models spanning a wide range of categories. UniRig significantly outperforms state-of-the-art academic and commercial methods, achieving a 215% improvement in rigging accuracy and a 194% improvement in motion accuracy on challenging datasets. Our method works seamlessly across diverse object categories, from detailed anime characters to complex organic and inorganic structures, demonstrating its versatility and robustness. By automating the tedious and time-consuming rigging process, UniRig has the potential to speed up animation pipelines with unprecedented ease and efficiency. Project Page: https://zjp-shadow.github.io/works/UniRig/ Jia-Peng Zhang, Cheng-Feng Pu, Menghao Guo 0001, Yan-Pei Cao 0001, Shi-Min Hu 0001 |
ACM Trans. Graph. | 5 |
| 2025 | SN$^{2}$2eRF: A Framework for Neural Radiance Fields Given Sparse and Noisy PosesabstractNeural Radiance Fields (NeRFs) have shown impressive capabilities in synthesizing photorealistic novel views. However, their application to room-size scenes is limited by the requirement of several hundred views with accurate poses for training. To address this challenge, we propose SN$^{2}$2eRF, a framework which can reconstruct the neural radiance field with significantly fewer views and noisy poses by exploiting multiple priors. Our key insight is to leverage both multi-view and monocular priors to constrain the optimization of NeRF in the setting of sparse and noisy pose inputs. Specifically, we extract and match key points to constrain pose optimization and use Ray Transformer with a monocular depth estimator to provide dense depth prior for geometry optimization. Benefiting from these priors, our approach achieves state-of-the-art accuracy in novel view synthesis for indoor room scenarios. Haoxiang Chen 0004, Tai-Jiang Mu, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | Semantic-Aware Transformation-Invariant RoI AlignabstractGreat progress has been made in learning-based object detection methods in the last decade. Two-stage detectors often have higher detection accuracy than one-stage detectors, due to the use of region of interest (RoI) feature extractors which extract transformation-invariant RoI features for different RoI proposals, making refinement of bounding boxes and prediction of object categories more robust and accurate. However, previous RoI feature extractors can only extract invariant features under limited transformations. In this paper, we propose a novel RoI feature extractor, termed Semantic RoI Align (SRA), which is capable of extracting invariant RoI features under a variety of transformations for two-stage detectors. Specifically, we propose a semantic attention module to adaptively determine different sampling areas by leveraging the global and local semantic relationship within the RoI. We also propose a Dynamic Feature Sampler which dynamically samples features based on the RoI aspect ratio to enhance the efficiency of SRA, and a new position embedding, i.e., Area Embedding, to provide more accurate position information for SRA through an improved sampling area representation. Experiments show that our model significantly outperforms baseline models with slight computational overhead. In addition, it shows excellent generalization ability and can be used to improve performance with various state-of-the-art backbones and detection methods. The code is available at https://github.com/cxjyxxme/SemanticRoIAlign. Guo-Ye Yang, George Kiyohiro Nakayama, Zi-Kai Xiao, Tai-Jiang Mu, Sharon X. Huang, Shi-Min Hu 0001 |
AAAI | 6 |
| 2024 | Multi-Dimensional and Message-Guided Fuzzing for Robotic Programs in Robot Operating SystemabstractAn increasing number of robotic programs are implemented based on Robot Operating System (ROS), which provides many practical tools and libraries for robot development. To improve robot reliability and security, several recent approaches apply fuzzing to ROS programs for bug detection. However, these approaches still have some main limitations, including inefficient test case generation, ineffective program feedback and weak generality/automation. Jia-Ju Bai, Haoxuan Song, Shi-Min Hu 0001 |
ASPLOS (2) | 3 |
| 2024 | Theoretically Achieving Continuous Representation of Oriented Bounding BoxesabstractConsiderable efforts have been devoted to Oriented Ob-ject Detection (OOD). However, one lasting issue regarding the discontinuity in Oriented Bounding Box (OBB) rep-resentation remains unresolved, which is an inherent bot-tleneck for extant OOD methods. This paper endeavors to completely solve this issue in a theoretically guaranteed manner and puts an end to the ad-hoc efforts in this di-rection. Prior studies typically can only address one of the two cases of discontinuity: rotation and aspect ratio, and often inadvertently introduce decoding discontinuity, e.g. Decoding Incompleteness (DI) and Decoding Ambi-guity (DA) as discussed in literature. Specifically, we pro-pose a novel representation method called Continuous OBB (COBB), which can be readily integrated into existing de-tectors e.g. Faster-RCNN as a plugin. It can theoreti-cally ensure continuity in bounding box regression which to our best knowledge, has not been achieved in literature for rectangle-based object representation. For fairness and transparency of experiments, we have developed a modu-larized benchmark based on the open-source deep learning framework Jittor's detection toolbox JDetfor OOD evaluation. On the popular DOTA dataset, by integrating Faster-RCNN as the same baseline model, our new method out-performs the peer method Gliding Vertex by 1.13% mAP50(relative improvement 1.54%), and 2.46% mAP75(relative improvement 5.91%), without any tricks. Zi-Kai Xiao, Guo-Ye Yang, Xue Yang 0005, Tai-Jiang Mu, Junchi Yan, Shi-Min Hu 0001 |
CVPR | 6 |
| 2024 | Exploring Regional Clues in CLIP for Zero-Shot Semantic SegmentationabstractCLIP has demonstrated marked progress in visual recognition due to its powerful pre-training on large-scale image-text pairs. However, it still remains a critical challenge: how to transfer image-level knowledge into pixel-level understanding tasks such as semantic segmentation. In this paper, to solve the mentioned challenge, we analyze the gap between the capability of the CLIP model and the requirement of the zero-shot semantic segmentation task. Based on our analysis and observations, we propose a novel method for zero-shot semantic segmentation, dubbed CLIP-RC (CLIP with Regional Clues), bringing two main insights. On the one hand, a region-level bridge is necessary to provide fine-grained semantics. On the other hand, over-fitting should be mitigated during the training stage. Benefiting from the above discoveries, CLIP-RC achieves state-of-the-art performance on various zero-shot semantic segmentation benchmarks, including PASCAL VOC, PASCAL Context, and COCO-Stuff 164K. Code will be available at https://github.com/Jittor/JSeg. Yi Zhang 0099, Menghao Guo 0001, Miao Wang 0004, Shi-Min Hu 0001 |
CVPR | 4 |
| 2024 | Recovering Complete Actions for Cross-dataset Skeleton Action RecognitionabstractDespite huge progress in skeleton-based action recognition, its generalizability to different domains remains a challenging issue.
In this paper, to solve the skeleton action generalization problem, we present a recover-and-resample augmentation framework based on a novel complete action prior. We observe that human daily actions are confronted with temporal mismatch across different datasets, as they are usually partial observations of their complete action sequences. By recovering complete actions and resampling from these full sequences, we can generate strong augmentations for unseen domains. At the same time, we discover the nature of general action completeness within large datasets, indicated by the per-frame diversity over time. This allows us to exploit two assets of transferable knowledge that can be shared across action samples and be helpful for action completion: boundary poses for determining the action start, and linear temporal transforms for capturing global action patterns. Therefore, we formulate the recovering stage as a two-step stochastic action completion with boundary pose-conditioned extrapolation followed by smooth linear transforms. Both the boundary poses and linear transforms can be efficiently learned from the whole dataset via clustering. We validate our approach on a cross-dataset setting with three skeleton action datasets, outperforming other domain generalization approaches by a considerable margin. Yujiang Li, Tai-Jiang Mu, Shi-Min Hu 0001 |
NeurIPS | 4 |
| 2024 | EVSplitting: An Efficient and Visually Consistent Splitting Algorithm for 3D Gaussian SplattingabstractThis paper presents EVSplitting, an efficient and visually consistent splitting algorithm for 3D Gaussian Splatting (3DGS). It is designed to make operating 3DGS as easy and effective as other 3D explicit representations, readily for industrial productions. The challenges of above target are: 1) The huge number and complex attributes of 3DGS make it tough to explicitly operate on 3DGS in a real-time and learning-free manner; 2) The visual effect of 3DGS is very difficult to maintain during explicit operations and 3) The anisotropism of Gaussian always leads to blurs and artifacts. As far as we know, no prior work can address these challenges well. In this work, we introduce a direct and efficient 3DGS splitting algorithm to solve them. Specifically, we formulate the 3DGS splitting as two minimization problems that aim to ensure visual consistency and reduce Gaussian overflow across boundary (splitting plane), respectively. Firstly, we impose conservations on the zero-, first- and second-order moments of the weighted Gaussian distribution to guarantee visual consistency. Secondly, we reduce the boundary overflow with a special constraint on the aforementioned conservations. With these conservations and constraints, we derive a closed-form solution for the 3DGS splitting problem. This yields an easy-to-implement, plug-and-play, efficient and fundamental tool, benefiting various downstream applications of 3DGS. Qi-Yuan Feng, Geng-Chen Cao, Haoxiang Chen 0004, Qun-Ce Xu, Tai-Jiang Mu, Ralph R. Martin, Shi-Min Hu 0001 |
SIGGRAPH Asia | 7 |
| 2024 | DIScene: Object Decoupling and Interaction Modeling for Complex Scene Generation
Xiao-Lei Li, Haoxiang Chen 0004, Tai-Jiang Mu, Shi-Min Hu 0001 |
SIGGRAPH Asia | 5 |
| 2024 | LR-Miner: Static Race Detection in OS Kernels by Mining Locking Rules
Tuo Li 0005, Jia-Ju Bai, Gui-Dong Han, Shi-Min Hu 0001 |
USENIX Security Symposium | 4 |
| 2024 | OpBench: an operator-level GPU benchmark for deep learning
Qingwen Gu, Zheng-Ning Liu, Kaicheng Cao, Songhai Zhang, Shi-Min Hu 0001 |
Sci. China Inf. Sci. | 6 |
| 2024 | Preface
Shi-Min Hu 0001, Andrei Sharf |
J. Comput. Sci. Technol. | 1 |
| 2024 | Tuning Vision-Language Models With Multiple Prototypes ClusteringabstractBenefiting from advances in large-scale pre-training, foundation models, have demonstrated remarkable capability in the fields of natural language processing, computer vision, among others. However, to achieve expert-level performance in specific applications, such models often need to be fine-tuned with domain-specific knowledge. In this paper, we focus on enabling vision-language models to unleash more potential for visual understanding tasks under few-shot tuning. Specifically, we propose a novel adapter, dubbed as lusterAdapter, which is based on trainable multiple prototypes clustering algorithm, for tuning the CLIP model. It can not only alleviate the concern of catastrophic forgetting of foundation models by introducing anchors to inherit common knowledge, but also improve the utilization efficiency of few annotated samples via bringing in clustering and domain priors, thereby improving the performance of few-shot tuning. We have conducted extensive experiments on 11 common classification benchmarks. The results show our method significantly surpasses the original CLIP and achieves state-of-the-art (SOTA) performance under all benchmarks and settings. For example, under the 16-shot setting, our method exhibits a remarkable improvement over the original CLIP by 19.6%, and also surpasses TIP-Adapter and GraphAdapter by 2.7% and 2.2%, respectively, in terms of average accuracy across the 11 benchmarks. Menghao Guo 0001, Yi Zhang 0099, Tai-Jiang Mu, Sharon X. Huang, Shi-Min Hu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Learning Virtual View Selection for 3D Scene Semantic Segmentationabstract2D-3D joint learning is essential and effective for fundamental 3D vision tasks, such as 3D semantic segmentation, due to the complementary information these two visual modalities contain. Most current 3D scene semantic segmentation methods process 2D images "as they are", i.e., only real captured 2D images are used. However, such captured 2D images may be redundant, with abundant occlusion and/or limited field of view (FoV), leading to poor performance for the current methods involving 2D inputs. In this paper, we propose a general learning framework for joint 2D-3D scene understanding by selecting informative virtual 2D views of the underlying 3D scene. We then feed both the 3D geometry and the generated virtual 2D views into any joint 2D-3D-input or pure 3D-input based deep neural models for improving 3D scene understanding. Specifically, we generate virtual 2D views based on an information score map learned from the current 3D scene semantic segmentation results. To achieve this, we formalize the learning of the information score map as a deep reinforcement learning process, which rewards good predictions using a deep neural network. To obtain a compact set of virtual 2D views that jointly cover informative surfaces of the 3D scene as much as possible, we further propose an efficient greedy virtual view coverage strategy in the normal-sensitive 6D space, including 3-dimensional point coordinates and 3-dimensional normal. We have validated our proposed framework for various joint 2D-3D-input or pure 3D-input based deep neural models on two real-world 3D scene datasets, i.e., ScanNet v2 and S3DIS, and the results demonstrate that our method obtains a consistent gain over baseline models and achieves new top accuracy for joint 2D and 3D scene semantic segmentation. Code is available at https://github.com/smy-THU/VirtualViewSelection. Tai-Jiang Mu, Ming-Yuan Shen, Yukun Lai, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 4 |
| 2024 | SPATA: Effective OS Bug Detection with Summary-Based, Alias-Aware, and Path-Sensitive Typestate AnalysisabstractThe operating system (OS) is the cornerstone for computer systems. It manages hardware and provides fundamental service for user-level applications. Thus, detecting bugs in OSes is important to improve the reliability of computer systems. Static typestate analysis is a common technique for detecting various types of bugs, but it is often inaccurate or unscalable for large-size OS code, due to imprecision of identifying alias relationships as well as high costs of typestate tracking, path-feasibility validation, and inter-procedural analysis. In this article, 1 we present SPATA, a novel summary-based, alias-aware, and path-sensitive typestate analysis framework to detect OS bugs. To identify precise alias relationships in the OS code, SPATA performs a path-based alias analysis based on control-flow paths and access paths. With these alias relationships, SPATA reduces the costs of typestate tracking and path-feasibility validation, to accelerate path-sensitive typestate analysis for accurate bug detection. Moreover, SPATA uses an alias-summary-based analysis to accelerate inter-procedural bug detection, without time-consuming alias analysis across functions. We have evaluated SPATA on the Linux kernel and three popular IoT OSes, and it finds 651 real bugs with a false-positive rate of 18%. Besides, our alias-summary-based analysis achieves a 6.7x speedup in bug detection compared to non-summary-based analysis. Tuo Li 0005, Jia-Ju Bai, Yulei Sui, Shi-Min Hu 0001 |
ACM Trans. Comput. Syst. | 4 |
| 2024 | CharacterGen: Efficient 3D Character Generation from Single Images with Multi-View Pose CanonicalizationabstractBNRist, Department of Computer Science and Technology, Tsinghua University, China In the field of digital content creation, generating high-quality 3D characters from single images is challenging, especially given the complexities of various body poses and the issues of self-occlusion and pose ambiguity. In this paper, we present CharacterGen, a framework developed to efficiently generate 3D characters. CharacterGen introduces a streamlined generation pipeline along with an image-conditioned multi-view diffusion model. This model effectively calibrates input poses to a canonical form while retaining key attributes of the input image, thereby addressing the challenges posed by diverse poses. A transformer-based, generalizable sparse-view reconstruction model is the other core component of our approach, facilitating the creation of detailed 3D models from multi-view images. We also adopt a texture-back-projection strategy to produce high-quality texture maps. Additionally, we have curated a dataset of anime characters, rendered in multiple poses and views, to train and evaluate our model. Our approach has been thoroughly evaluated through quantitative and qualitative experiments, showing its proficiency in generating 3D characters with high-quality shapes and textures, ready for downstream applications such as rigging and animation. Hao-Yang Peng, Jia-Peng Zhang, Menghao Guo 0001, Yan-Pei Cao 0001, Shi-Min Hu 0001 |
ACM Trans. Graph. | 5 |
| 2024 | Mesh Neural Networks Based on Dual Graph PyramidsabstractDeep neural networks (DNNs) have been widely used for mesh processing in recent years. However, current DNNs can not process arbitrary meshes efficiently. On the one hand, most DNNs expect 2-manifold, watertight meshes, but many meshes, whether manually designed or automatically generated, may have gaps, non-manifold geometry, or other defects. On the other hand, the irregular structure of meshes also brings challenges to building hierarchical structures and aggregating local geometric information, which is critical to conduct DNNs. In this paper, we present DGNet, an efficient, effective and generic deep neural mesh processing network based on dual graph pyramids; it can handle arbitrary meshes. First, we construct dual graph pyramids for meshes to guide feature propagation between hierarchical levels for both downsampling and upsampling. Second, we propose a novel convolution to aggregate local features on the proposed hierarchical graphs. By utilizing both geodesic neighbors and euclidean neighbors, the network enables feature aggregation both within local surface patches and between isolated mesh components. Experimental results demonstrate that DGNet can be applied to both shape analysis and large-scale scene understanding. Furthermore, it achieves superior performance on various benchmarks, including ShapeNetCore, HumanBody, ScanNet and Matterport3D. Code and models will be available at https://github.com/li-xl/DGNet. Xiang-Li Li, Zheng-Ning Liu, Tuo Chen, Tai-Jiang Mu, Ralph R. Martin, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2024 | LC-NeRF: Local Controllable Face Generation in Neural Radiance Fieldabstract3D face generation has achieved high visual quality and 3D consistency thanks to the development of neural radiance fields (NeRF). However, these methods model the whole face as a neural radiance field, which limits the controllability of the local regions. In other words, previous methods struggle to independently control local regions, such as the mouth, nose, and hair. To improve local controllability in NeRF-based face generation, we propose LC-NeRF, which is composed of a Local Region Generators Module (LRGM) and a Spatial-Aware Fusion Module (SAFM), allowing for geometry and texture control of local facial regions. The LRGM models different facial regions as independent neural radiance fields and the SAFM is responsible for merging multiple independent neural radiance fields into a complete representation. Finally, LC-NeRF enables the modification of the latent code associated with each individual generator, thereby allowing precise control over the corresponding local region. Qualitative and quantitative evaluations show that our method provides better local controllability than state-of-the-art 3D-aware face generation methods. A perception study reveals that our method outperforms existing state-of-the-art methods in terms of image quality, face consistency, and editing effects. Furthermore, our method exhibits favorable performance in downstream tasks, including real image editing and text-driven facial image editing. Wenyang Zhou, Lin Gao 0004, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2023 | Long Range Pooling for 3D Large-Scale Scene UnderstandingabstractInspired by the success of recent vision transformers and large kernel design in convolutional neural networks (CNNs), in this paper, we analyze and explore essential reasons for their success. We claim two factors that are critical for 3D large-scale scene understanding: a larger receptive field and operations with greater non-linearity. The former is responsible for providing long range contexts and the latter can enhance the capacity of the network. To achieve the above properties, we propose a simple yet effective long range pooling (LRP) module using dilation max pooling, which provides a network with a large adaptive receptive field. LRP has few parameters, and can be readily added to current CNNs. Also, based on LRP, we present an entire network architecture, LRPNet, for 3D understanding. Ablation studies are presented to support our claims, and show that the LRP module achieves better results than large kernel convolution yet with reduced computation, due to its non-linearity. We also demonstrate the superiority of LRPNet on various benchmarks: LRPNet performs the best on ScanNet and surpasses other CNN-based methods on S3DIS and Matterport3D. Code will be avalible at https://github.com/li-xl/LRPNet. Xiang-Li Li, Menghao Guo 0001, Tai-Jiang Mu, Ralph R. Martin, Shi-Min Hu 0001 |
CVPR | 5 |
| 2023 | DiffFacto: Controllable Part-Based 3D Point Cloud Generation with Cross DiffusionabstractWhile the community of 3D point cloud generation has witnessed a big growth in recent years, there still lacks an effective way to enable intuitive user control in the generation process, hence limiting the general utility of such methods. Since an intuitive way of decomposing a shape is through its parts, we propose to tackle the task of controllable part-based point cloud generation. We introduce DiffFacto, a novel probabilistic generative model that learns the distribution of shapes with part-level control. We propose a factorization that models independent part style and part configuration distributions, and present a novel cross diffusion network that enables us to generate coherent and plausible shapes under our proposed factorization. Experiments show that our method is able to generate novel shapes with multiple axes of control. It achieves state-of-the-art part-level generation quality and generates plausible and coherent shape while enabling various downstream editing applications such as shape interpolation, mixing, and transformation editing. Please visit our project webpage at https://difffacto.github.io/ George Kiyohiro Nakayama, Mikaela Angelina Uy, Shi-Min Hu 0001, Ke Li 0011, Leonidas J. Guibas |
ICCV | 4 |
| 2023 | Visual attention networkabstractWhile originally designed for natural language processing tasks, the self-attention mechanism has recently taken various computer vision areas by storm. However, the 2D nature of images brings three challenges for applying self-attention in computer vision: (1) treating images as 1D sequences neglects their 2D structures; (2) the quadratic complexity is too expensive for high-resolution images; (3) it only captures spatial adaptability but ignores channel adaptability. In this paper, we propose a novel linear attention named large kernel attention (LKA) to enable self-adaptive and long-range correlations in self-attention while avoiding its shortcomings. Furthermore, we present a neural network based on LKA, namely Visual Attention Network (VAN). While extremely simple, VAN achieves comparable results with similar size convolutional neural networks (CNNs) and vision transformers (ViTs) in various tasks, including image classification, object detection, semantic segmentation, panoptic segmentation, pose estimation, etc. For example, VAN-B6 achieves 87.8% accuracy on ImageNet benchmark, and sets new state-of-the-art performance (58.2 PQ) for panoptic segmentation. Besides, VAN-B2 surpasses Swin-T 4 mIoU (50.1 vs. 46.1) for semantic segmentation on ADE20K benchmark, 2.6 AP (48.8 vs. 46.2) for object detection on COCO dataset. It provides a novel method and a simple yet strong baseline for the community. The code is available at https://github.com/Visual-Attention-Network . Menghao Guo 0001, Chengze Lu, Zheng-Ning Liu, Ming-Ming Cheng, Shi-Min Hu 0001 |
Comput. Vis. Media | 5 |
| 2023 | Message from the Editor-in-ChiefabstractI would like to take this opportunity to thank everyone who Shi-Min Hu 0001 |
Comput. Vis. Media | 1 |
| 2023 | Message from the Editor-in-ChiefabstractI would like to take this opportunity to thank everyone who has helped to make Computational Visual Media a success in its ninth year of 2023. In particular, my Shi-Min Hu 0001 |
Comput. Vis. Media | 1 |
| 2023 | Preface
Shi-Min Hu 0001, Amit Bermano, Ruizhen Hu |
J. Comput. Sci. Technol. | 1 |
| 2023 | Beyond Self-Attention: External Attention Using Two Linear Layers for Visual TasksabstractAttention mechanisms, especially self-attention, have played an increasingly important role in deep feature representation for visual tasks. Self-attention updates the feature at each position by computing a weighted sum of features using pair-wise affinities across all positions to capture the long-range dependency within a single sample. However, self-attention has quadratic complexity and ignores potential correlation between different samples. This article proposes a novel attention mechanism which we call external attention, based on two external, small, learnable, shared memories, which can be implemented easily by simply using two cascaded linear layers and two normalization layers; it conveniently replaces self-attention in existing popular architectures. External attention has linear complexity and implicitly considers the correlations between all data samples. We further incorporate the multi-head mechanism into external attention to provide an all-MLP architecture, external attention MLP (EAMLP), for image classification. Extensive experiments on image classification, object detection, semantic segmentation, instance segmentation, image generation, and point cloud analysis reveal that our method provides results comparable or superior to the self-attention mechanism and some of its variants, with much lower computational and memory costs. Menghao Guo 0001, Zheng-Ning Liu, Tai-Jiang Mu, Shi-Min Hu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Multiway Non-Rigid Point Cloud Registration via Learned Functional Map SynchronizationabstractWe present SyNoRiM, a novel way to jointly register multiple non-rigid shapes by synchronizing the maps that relate learned functions defined on the point clouds. Even though the ability to process non-rigid shapes is critical in various applications ranging from computer animation to 3D digitization, the literature still lacks a robust and flexible framework to match and align a collection of real, noisy scans observed under occlusions. Given a set of such point clouds, our method first computes the pairwise correspondences parameterized via functional maps. We simultaneously learn potentially non-orthogonal basis functions to effectively regularize the deformations, while handling the occlusions in an elegant way. To maximally benefit from the multi-way information provided by the inferred pairwise deformation fields, we synchronize the pairwise functional maps into a cycle-consistent whole thanks to our novel and principled optimization formulation. We demonstrate via extensive experiments that our method achieves a state-of-the-art performance in registration accuracy, while being flexible and efficient as we handle both non-rigid and multi-body cases in a unified framework and avoid the costly optimization over point-wise permutations by the use of basis function maps. Tolga Birdal, Zan Gojcic, Leonidas J. Guibas, Shi-Min Hu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Skeleton-CutMix: Mixing Up Skeleton With Probabilistic Bone Exchange for Supervised Domain AdaptationabstractWe present Skeleton-CutMix, a simple and effective skeleton augmentation framework for supervised domain adaptation and show its advantage in skeleton-based action recognition tasks. Existing approaches usually perform domain adaptation for action recognition with elaborate loss functions that aim to achieve domain alignment. However, they fail to capture the intrinsic characteristics of skeleton representation. Benefiting from the well-defined correspondence between bones of a pair of skeletons, we instead mitigate domain shift by fabricating skeleton data in a mixed domain, which mixes up bones from the source domain and the target domain. The fabricated skeletons in the mixed domain can be used to augment training data and train a more general and robust model for action recognition. Specifically, we hallucinate new skeletons by using pairs of skeletons from the source and target domains; a new skeleton is generated by exchanging some bones from the skeleton in the source domain with corresponding bones from the skeleton in the target domain, which resembles a cut-and-mix operation. When exchanging bones from different domains, we introduce a class-specific bone sampling strategy so that bones that are more important for an action class are exchanged with higher probability when generating augmentation samples for that class. We show experimentally that the simple bone exchange strategy for augmentation is efficient and effective and that distinctive motion features are preserved while mixing both action and style across domains. We validate our method in cross-dataset and cross-age settings on NTU-60 and ETRI-Activity3D datasets with an average gain of over 3% in terms of action recognition accuracy, and demonstrate its superior performance over previous domain adaptation approaches as well as other skeleton augmentation strategies. Yuhe Liu, Tai-Jiang Mu, Sharon X. Huang, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 5 |
| 2023 | Sampling Equivariant Self-Attention Networks for Object Detection in Aerial ImagesabstractObjects in aerial images show greater variations in scale and orientation than in other images, making them harder to detect using vanilla deep convolutional neural networks. Networks with sampling equivariance can adapt sampling from input feature maps to object transformation, allowing a convolutional kernel to extract effective object features under different transformations. However, methods such as deformable convolutional networks can only provide sampling equivariance under certain circumstances, as they sample by location. We propose sampling equivariant self-attention networks, which treat self-attention restricted to a local image patch as convolution sampling by masks instead of locations, and a transformation embedding module to improve the equivariant sampling further. We further propose a novel randomized normalization module to enhance network generalization and a quantitative evaluation metric to fairly evaluate the ability of sampling equivariance of different models. Experiments show that our model provides significantly better sampling equivariance than existing methods without additional supervision and can thus extract more effective image features. Our model achieves state-of-the-art results on the DOTA-v1.0, DOTA-v1.5, and HRSC2016 datasets without additional computations or parameters. Guo-Ye Yang, Xiang-Li Li, Zi-Kai Xiao, Tai-Jiang Mu, Ralph R. Martin, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 6 |
| 2023 | Real-Time Globally Consistent 3D Reconstruction With Semantic PriorsabstractMaintaining global consistency continues to be critical for online 3D indoor scene reconstruction. However, it is still challenging to generate satisfactory 3D reconstruction in terms of global consistency for previous approaches using purely geometric analysis, even with bundle adjustment or loop closure techniques. In this article, we propose a novel real-time 3D reconstruction approach which effectively integrates both semantic and geometric cues. The key challenge is how to map this indicative information, i.e., semantic priors, into a metric space as measurable information, thus enabling more accurate semantic fusion leveraging both the geometric and semantic cues. To this end, we introduce a semantic space with a continuous metric function measuring the distance between discrete semantic observations. Within the semantic space, we present an accurate frame-to-model semantic tracker for camera pose estimation, and semantic pose graph equipped with semantic links between submaps for globally consistent 3D scene reconstruction. With extensive evaluation on public synthetic and real-world 3D indoor scene RGB-D datasets, we show that our approach outperforms the previous approaches for 3D scene reconstruction both quantitatively and qualitatively, especially in terms of global consistency. Shi-Sheng Huang, Haoxiang Chen 0004, Hongbo Fu 0001, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2023 | On Rotation Gains Within and Beyond Perceptual Limitations for Seated VRabstractHead tracking in head-mounted displays (HMDs) enables users to explore a 360-degree virtual scene with free head movements. However, for seated use of HMDs such as users sitting on a chair or a couch, physically turning around 360-degree is not possible. Redirection techniques decouple tracked physical motion and virtual motion, allowing users to explore virtual environments with more flexibility. In seated situations with only head movements available, the difference of stimulus might cause the detection thresholds of rotation gains to differ from that of redirected walking. Therefore we present an experiment with a two-alternative forced-choice (2AFC) design to compare the thresholds for seated and standing situations. Results indicate that users are unable to discriminate rotation gains between 0.89 and 1.28, a smaller range compared to the standing condition. We further treated head amplification as an interaction technique and found that a gain of 2.5, though not a hard threshold, was near the largest gain that users consider applicable. Overall, our work aims to better understand human perception of rotation gains in seated VR and the results provide guidance for future design choices of its applications. Chen Wang 0049, Song-Hai Zhang, Yizhuo Zhang 0001, Stefanie Zollmann, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2023 | Recursive-NeRF: An Efficient and Dynamically Growing NeRFabstractView synthesis methods using implicit continuous shape representations learned from a set of images, such as the Neural Radiance Field (NeRF) method, have gained increasing attention due to their high quality imagery and scalability to high resolution. However, the heavy computation required by its volumetric approach prevents NeRF from being useful in practice; minutes are taken to render a single image of a few megapixels. Now, an image of a scene can be rendered in a level-of-detail manner, so we posit that a complicated region of the scene should be represented by a large neural network while a small neural network is capable of encoding a simple region, enabling a balance between efficiency and quality. Recursive-NeRF is our embodiment of this idea, providing an efficient and adaptive rendering and training approach for NeRF. The core of Recursive-NeRF learns uncertainties for query coordinates, representing the quality of the predicted color and volumetric intensity at each level. Only query coordinates with high uncertainties are forwarded to the next level to a bigger neural network with a more powerful representational capability. The final rendered image is a composition of results from neural networks of all levels. Our evaluation on public datasets and a large-scale scene dataset we collected shows that Recursive-NeRF is more efficient than NeRF while providing state-of-the-art quality. The code will be available at https://github.com/Gword/Recursive-NeRF. Wenyang Zhou, Hao-Yang Peng, Dun Liang, Tai-Jiang Mu, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2023 | Adaptive Optimization Algorithm for Resetting Techniques in Obstacle-Ridden EnvironmentsabstractRedirected Walking (RDW) algorithms aim to impose several types of gains on users immersed in Virtual Reality and distort their walking paths in the real world, thus enabling them to explore a larger space. Since collision with physical boundaries is inevitable, a reset strategy needs to be provided to allow users to reset when they hit the boundary. However, most reset strategies are based on simple heuristics by choosing a seemingly suitable solution, which may not perform well in practice. In this article, we propose a novel optimization-based reset algorithm adaptive to different RDW algorithms. Inspired by the approach of finite element analysis, our algorithm splits the boundary of the physical world by a set of endpoints. Each endpoint is assigned a reset vector to represent the optimized reset direction when hitting the boundary. The reset vectors on the edge will be determined by the interpolation between two neighbouring endpoints. We conduct simulation-based experiments for three RDW algorithms with commonly used reset algorithms to compare with. The results demonstrate that the proposed algorithm significantly reduces the number of resets. Song-Hai Zhang, Chia-Hao Chen, Fu Zheng, Yongliang Yang 0002, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2022 | Path-sensitive and alias-aware typestate analysis for detecting OS bugsabstractOperating system (OS) is the cornerstone for modern computer systems. It manages devices and provides fundamental service for user-level applications. Thus, detecting bugs in OSes is important to improve reliability and security of computer systems. Static typestate analysis is a common technique for detecting different types of bugs, but it is often inaccurate or unscalable for large-size OS code, due to imprecision of identifying alias relationships as well as high costs of typestate tracking and path-feasibility validation. Tuo Li 0005, Jia-Ju Bai, Yulei Sui, Shi-Min Hu 0001 |
ASPLOS | 4 |
| 2022 | CIRCLE: Convolutional Implicit Reconstruction and Completion for Large-Scale Indoor Scene
Haoxiang Chen 0004, Tai-Jiang Mu, Shi-Min Hu 0001 |
ECCV (32) | 4 |
| 2022 | NeRF-SR: High Quality Neural Radiance Fields using SupersamplingabstractWe present NeRF-SR, a solution for high-resolution (HR) novel view synthesis with mostly low-resolution (LR) inputs. Our method is built upon Neural Radiance Fields (NeRF) that predicts per-point density and color with a multi-layer perceptron. While producing images at arbitrary scales, NeRF struggles with resolutions that go beyond observed images. Our key insight is that NeRF benefits from 3D consistency, which means an observed pixel absorbs information from nearby views. We first exploit it by a super-sampling strategy that shoots multiple rays at each image pixel, which further enforces multi-view constraint at a sub-pixel level. Then, we show that NeRF-SR can further boost the performance of super-sampling by a refinement network that leverages the estimated depth at hand to hallucinate details from related patches on only one HR reference image. Experiment results demonstrate that NeRF-SR generates high-quality results for novel view synthesis at HR on both synthetic and real-world datasets without any external information. Project page: https://cwchenwang.github.io/NeRF-SR Chen Wang 0049, Xian Wu 0004, Song-Hai Zhang, Yu-Wing Tai, Shi-Min Hu 0001 |
ACM Multimedia | 6 |
| 2022 | Context-Sensitive and Directional Concurrency Fuzzing for Data-Race Detection
Zu-Ming Jiang, Jia-Ju Bai, Kangjie Lu, Shi-Min Hu 0001 |
NDSS | 4 |
| 2022 | SegNeXt: Rethinking Convolutional Attention Design for Semantic SegmentationabstractWe present SegNeXt, a simple convolutional network architecture for semantic segmentation. Recent transformer-based models have dominated the field of se- mantic segmentation due to the efficiency of self-attention in encoding spatial information. In this paper, we show that convolutional attention is a more efficient and effective way to encode contextual information than the self-attention mech- anism in transformers. By re-examining the characteristics owned by successful segmentation models, we discover several key components leading to the perfor- mance improvement of segmentation models. This motivates us to design a novel convolutional attention network that uses cheap convolutional operations. Without bells and whistles, our SegNeXt significantly improves the performance of previous state-of-the-art methods on popular benchmarks, including ADE20K, Cityscapes, COCO-Stuff, Pascal VOC, Pascal Context, and iSAID. Notably, SegNeXt out- performs EfficientNet-L2 w/ NAS-FPN and achieves 90.6% mIoU on the Pascal VOC 2012 test leaderboard using only 1/10 parameters of it. On average, SegNeXt achieves about 2.0% mIoU improvements compared to the state-of-the-art methods on the ADE20K datasets with the same or fewer computations. Menghao Guo 0001, Chengze Lu, Qibin Hou, Zheng-Ning Liu, Ming-Ming Cheng, Shi-Min Hu 0001 |
NeurIPS | 6 |
| 2022 | DLOS: Effective Static Detection of Deadlocks in OS Kernels
Jia-Ju Bai, Tuo Li 0005, Shi-Min Hu 0001 |
USENIX ATC | 3 |
| 2022 | Attention mechanisms in computer vision: A surveyabstractHumans can naturally and effectively find salient regions in complex scenes. Motivated by this observation, attention mechanisms were introduced into computer vision with the aim of imitating this aspect of the human visual system. Such an attention mechanism can be regarded as a dynamic weight adjustment process based on features of the input image. Attention mechanisms have achieved great success in many visual tasks, including image classification, object detection, semantic segmentation, video understanding, image generation, 3D vision, multimodal tasks, and self-supervised learning. In this survey, we provide a comprehensive review of various attention mechanisms in computer vision and categorize them according to approach, such as channel attention, spatial attention, temporal attention, and branch attention; a related repository https://github.com/MenghaoGuo/Awesome-Vision-Attentions is dedicated to collecting related work. We also suggest future directions for attention mechanism research. Menghao Guo 0001, Tian-Xing Xu, Jiang-Jiang Liu 0001, Zheng-Ning Liu, Peng-Tao Jiang, Tai-Jiang Mu, Song-Hai Zhang, Ralph R. Martin, Ming-Ming Cheng, Shi-Min Hu 0001 |
Comput. Vis. Media | 10 |
| 2022 | Message from the Editor-in-ChiefabstractI would like to take this opportunity to thank everyone who helped to make Computational Visual Media a success in its seventh year of 2021. In particular, my thanks go Shi-Min Hu 0001 |
Comput. Vis. Media | 1 |
| 2022 | Preface
Shi-Min Hu 0001, Paul L. Rosin, Tian-Jia Shao |
J. Comput. Sci. Technol. | 1 |
| 2022 | User-Guided Deep Human Image Matting Using Arbitrary TrimapsabstractImage matting is widely studied for accurate foreground extraction. Most algorithms, including deep-learning based solutions, require a carefully edited trimap. Recent works attempt to combine the segmentation stage and matting stage in one CNN model, but errors occurring at the segmentation stage lead to unsatisfactory matte. We propose a user-guided approach for practical human matting. More precisely, we provide a good automatic initial matting and a natural way of interaction that reduces the workload of drawing trimaps and allows users to guide the matting in ambiguous situation. We also combine the segmentation and matting stage in an end-to-end CNN architecture and introduce a residual-learning module to support convenient stroke-based interaction. The proposed model learns to propagate the input trimap and modify the deep image features, which can efficiently correct the segmentation errors. Our model supports arbitrary forms of trimaps from carefully edited to totally unknown maps. Our model also allows users to choose from different foreground estimations according to their preference. We collected a large human matting dataset consisting of 12K real-world human images with complex background and human-object relations. The proposed model is trained on the new dataset with a novel trimap generation strategy that enables the model to tackle different test situations and highly improves the interaction efficiency. Our method outperforms other state-of-the-art automatic methods and achieve competitive accuracy when high-quality trimaps are provided. Experiments indicate that our interactive matting strategy is superior to separately estimating the trimap and alpha matte using two models. Xiaonan Fang 0001, Song-Hai Zhang, Tao Chen 0015, Xian Wu 0004, Ariel Shamir, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 6 |
| 2022 | Subdivision-based Mesh Convolution NetworksabstractConvolutionalneural networks (CNNs) have made great breakthroughs in two-dimensional (2D) computer vision. However, their irregular structure makes it hard to harness the potential of CNNs directly on meshes. A subdivision surface provides a hierarchical multi-resolution structure in which each face in a closed 2-manifold triangle mesh is exactly adjacent to three faces. Motivated by these two observations, this article presents SubdivNet , an innovative and versatile CNN framework for three-dimensional (3D) triangle meshes with Loop subdivision sequence connectivity. Making an analogy between mesh faces and pixels in a 2D image allows us to present a mesh convolution operator to aggregate local features from nearby faces. By exploiting face neighborhoods, this convolution can support standard 2D convolutional network concepts, e.g., variable kernel size, stride, and dilation. Based on the multi-resolution hierarchy, we make use of pooling layers that uniformly merge four faces into one and an upsampling method that splits one face into four. Thereby, many popular 2D CNN architectures can be easily adapted to process 3D meshes. Meshes with arbitrary connectivity can be remeshed to have Loop subdivision sequence connectivity via self-parameterization, making SubdivNet a general approach. Extensive evaluation and various applications demonstrate SubdivNet’s effectiveness and efficiency. Shi-Min Hu 0001, Zheng-Ning Liu, Menghao Guo 0001, Junxiong Cai, Tai-Jiang Mu, Ralph R. Martin |
ACM Trans. Graph. | 1 |
| 2022 | A Neural Galerkin Solver for Accurate Surface ReconstructionabstractTo reconstruct meshes from the widely-available 3D point cloud data, implicit shape representation is among the primary choices as an intermediate form due to its superior representation power and robustness in topological optimizations. Although different parameterizations of the implicit fields have been explored to model the underlying geometry, there is no explicit mechanism to ensure the fitting tightness of the surface to the input. We present in response, NeuralGalerkin, a neural Galerkin-method-based solver designed for reconstructing highly-accurate surfaces from the input point clouds. NeuralGalerkin internally discretizes the target implicit field as a linear combination of a set of spatially-varying basis functions inferred by an adaptive sparse convolution neural network. It then solves differentiably for a variational problem that incorporates both positional and normal constraints from the data in closed form within a single forward pass, highly respecting the raw input points. The reconstructed surface extracted from the implicit interpolants is hence very accurate and incorporates useful inductive biases benefiting from the training data. Extensive evaluations on various datasets demonstrate our method's promising reconstruction performance and scalability. Haoxiang Chen 0004, Shi-Min Hu 0001 |
ACM Trans. Graph. | 3 |
| 2022 | Hybrid Static-Dynamic Analysis of Data Races Caused by Inconsistent Locking Discipline in Device DriversabstractData races are often hard to detect in device drivers. According to our study of Linux driver patches that fix data races, about 39% of patches involve a pattern that we callinconsistent locking discipline. Specifically, if a variable is accessed within two concurrently executed functions, the sets of locks held around each access are disjoint, at least one of the locksets is non-empty, and at least one of the involved accesses is a write, then a data race may occur. In this paper, we present a hybrid static-dynamic analysis approach, named SDILP, to detect data races caused by inconsistent locking discipline in device drivers. SDILP has a dynamic lockset analysis to detect data races at runtime, and a static lockset analysis to detect more data races based on the dynamic-analysis results. It also performs a static taint analysis to reduce the number of variable accesses monitored by the dynamic analysis. Compared to our previous dynamic approach DILP (Chen et al., 2019), introducing static analysis allows SDILP to achieve better performance and find more data races. We evaluate SDILP on 12 drivers in Linux 5.4, and find 117 real data races, 50 of which have been confirmed by driver developers. Jia-Ju Bai, Qiu-Liang Chen, Zu-Ming Jiang, Julia Lawall, Shi-Min Hu 0001 |
IEEE Trans. Software Eng. | 5 |
| 2022 | Context-Consistent Generation of Indoor Virtual Environments Based on Geometry ConstraintsabstractIn this article, we propose a system that can automatically generate immersive and interactive virtual reality (VR) scenes by taking real-world geometric constraints into account. Our system can not only help users avoid real-world obstacles in virtual reality experiences, but also provide context-consistent contents to preserve their sense of presence. To do so, our system first identifies the positions and bounding boxes of scene objects as well as a set of interactive planes from 3D scans. Then context-consistent virtual objects that have similar geometric properties to the real ones can be automatically selected and placed into the virtual scene, based on learned object association relations and layout patterns from large amounts of indoor scene configurations. We regard virtual object replacement as a combinatorial optimization problem, considering both geometric and contextual consistency constraints. Quantitative and qualitative results show that our system can generate plausible interactive virtual scenes that highly resemble real environments, and have the ability to keep the sense of presence for users in their VR experiences. Yu He 0001, Ying-Tian Liu, Yi-Han Jin, Song-Hai Zhang, Yukun Lai, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2022 | Simulating Fractures With Bonded Discrete Element MethodabstractAlong with motion and deformation, fracture is a fundamental behaviour for solid materials, playing a critical role in physically-based animation. Many simulation methods including both continuum and discrete approaches have been used by the graphics community to animate fractures for various materials. However, compared with motion and deformation, fracture remains a challenging task for simulation, because the material's geometry, topology and mechanical states all undergo continuous (and sometimes chaotic) changes as fragmentation develops. Recognizing the discontinuous nature of fragmentation, we propose a discrete approach, namely the Bonded Discrete Element Method (BDEM), for fracture simulation. The research of BDEM in engineering has been growing rapidly in recent years, while its potential in graphics has not been explored. We also introduce several novel changes to BDEM to make it more suitable for animation design. Compared with other fracture simulation methods, the BDEM has some attractive benefits, e.g., efficient handling of multiple fractures, simple formulation and implementation, and good scaling consistency. But it also has some critical weaknesses, e.g., high computational cost, which demand further research. A number of examples are presented to demonstrate the pros and cons, which are then highlighted in the conclusion and discussion. Jia-Ming Lu, Chenfeng Li, Geng-Chen Cao, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2021 | MultiBodySync: Multi-Body Segmentation and Motion Estimation via 3D Scan SynchronizationabstractWe present MultiBodySync, a novel, end-to-end trainable multi-body motion segmentation and rigid registration framework for multiple input 3D point clouds. The two non-trivial challenges posed by this multi-scan multibody setting that we investigate are: (i) guaranteeing correspondence and segmentation consistency across multiple input point clouds capturing different spatial arrangements of bodies or body parts; and (ii) obtaining robust motion-based rigid body segmentation applicable to novel object categories. We propose an approach to address these issues that incorporates spectral synchronization into an iterative deep declarative network, so as to simultaneously recover consistent correspondences as well as motion segmentation. At the same time, by explicitly disentangling the correspondence and motion segmentation estimation modules, we achieve strong generalizability across different object categories. Our extensive evaluations demonstrate that our method is effective on various datasets ranging from rigid parts in articulated objects to individually moving objects in a 3D scene, be it single-view or full point clouds. Code at https://github.com/huangjh-pub/multibody-sync. He Wang 0010, Tolga Birdal, Minhyuk Sung, Federica Arrigoni, Shi-Min Hu 0001, Leonidas J. Guibas |
CVPR | 6 |
| 2021 | DI-Fusion: Online Implicit 3D Reconstruction With Deep PriorsabstractPrevious online 3D dense reconstruction methods struggle to achieve the balance between memory storage and surface quality, largely due to the usage of stagnant underlying geometry representation, such as TSDF (truncated signed distance functions) or surfels, without any knowledge of the scene priors. In this paper, we present DI-Fusion (Deep Implicit Fusion), based on a novel 3D representation, i.e. Probabilistic Local Implicit Voxels (PLIVoxs), for online 3D reconstruction with a commodity RGB-D camera. Our PLIVox encodes scene priors considering both the local geometry and uncertainty parameterized by a deep neural network. With such deep priors, we are able to perform online implicit 3D reconstruction achieving state-of-the-art camera trajectory estimation accuracy and mapping quality, while achieving better storage efficiency compared with previous online 3D reconstruction approaches. Our implementation is available at https://www.github.com/huangjh-pub/di-fusion. Shi-Sheng Huang, Haoxuan Song, Shi-Min Hu 0001 |
CVPR | 4 |
| 2021 | TCP-Fuzz: Detecting Memory and Semantic Bugs in TCP Stacks with Fuzzing
Yonghao Zou, Jia-Ju Bai, Jielong Zhou, Jianfeng Tan, Chenggang Qin, Shi-Min Hu 0001 |
USENIX ATC | 6 |
| 2021 | Static Detection of Unsafe DMA Accesses in Device Drivers
Jia-Ju Bai, Tuo Li 0005, Kangjie Lu, Shi-Min Hu 0001 |
USENIX Security Symposium | 4 |
| 2021 | Detection Thresholds with Joint Horizontal and Vertical Gains in Redirected JumpingabstractRedirected jumping (RDJ) is a locomotion technique that allows users to explore a virtual space that is larger than the available physical space by imperceptibly manipulating users' virtual viewpoints according to different gains. In previous redirected jumping work, different types of gains were imposed separately, without considering the possible interaction effects of horizontal and vertical gains on the jumping distance perception. To figure out how humans perceive distance manipulation when more than one gain is used, in this paper, we explored joint horizontal and vertical gains that manipulate horizontal and vertical distances at the same time during two-legged takeoff jumping in the virtual space. We estimated and analyzed horizontal and vertical detection thresholds by conducting a user study, fitting the data to two-dimensional psychometric functions, and visualizing the fitted 3D plots. We provided quantitative insights into the effects of joint gains on detection thresholds, where the imperceptible range for one gain can be affected by the variation of the other gain. Finally, we designed redirected jumping-based games as applications with joint horizontal and vertical gains and demonstrated the effectiveness of the redirected jumping technique. Yijun Li 0006, De-Rong Jin, Miao Wang 0004, Frank Steinicke, Shi-Min Hu 0001, Qinping Zhao |
VR | 6 |
| 2021 | LinkNet: 2D-3D linked multi-modal network for online semantic segmentation of RGB-D videos
Junxiong Cai, Tai-Jiang Mu, Yukun Lai, Shi-Min Hu 0001 |
Comput. Graph. | 4 |
| 2021 | PCT: Point cloud transformerabstractThe irregular domain and lack of ordering make it challenging to design deep neural networks for point cloud processing. This paper presents a novel framework named Point Cloud Transformer (PCT) for point cloud learning. PCT is based on Transformer, which achieves huge success in natural language processing and displays great potential in image processing. It is inherently permutation invariant for processing a sequence of points, making it well-suited for point cloud learning. To better capture local context within the point cloud, we enhance input embedding with the support of farthest point sampling and nearest neighbor search. Extensive experiments demonstrate that the PCT achieves the state-of-the-art performance on shape classification, part segmentation, semantic segmentation, and normal estimation tasks. Menghao Guo 0001, Junxiong Cai, Zheng-Ning Liu, Tai-Jiang Mu, Ralph R. Martin, Shi-Min Hu 0001 |
Comput. Vis. Media | 6 |
| 2021 | Can attention enable MLPs to catch up with CNNs?abstractIn the first week of May 2021, researchers from four different institutions: Google, Menghao Guo 0001, Zheng-Ning Liu, Tai-Jiang Mu, Dun Liang, Ralph R. Martin, Shi-Min Hu 0001 |
Comput. Vis. Media | 6 |
| 2021 | Message from the Editor-in-Chiefabstractan annual award for the best papers published in Computational Visual Media.After carefully considering the papers published in 2020, the Editorial Board chose the paper Learning local shape descriptors for computing non-rigid dense correspondence [1] as the winner of the Best Paper Award; while three other papers: What and where: A context-based recommendation system for object insertion [2], A survey on deep geometry learning: From a representation perspective [3], and Temporal scatterplots [4], have won the Honorable Mention Awards. Shi-Min Hu 0001 |
Comput. Vis. Media | 1 |
| 2021 | ClusterSLAM: A SLAM backend for simultaneous rigid body clustering and motion estimationabstractWe present a practical backend for stereo visual SLAM which can simultaneously discover individual rigid bodies and compute their motions in dynamic environments. While recent factor graph based state optimization algorithms have shown their ability to robustly solve SLAM problems by treating dynamic objects as outliers, their dynamic motions are rarely considered. In this paper, we exploit the consensus of 3D motions for landmarks extracted from the same rigid body for clustering, and to identify static and dynamic objects in a unified manner. Specifically, our algorithm builds a noise-aware motion affinity matrix from landmarks, and uses agglomerative clustering to distinguish rigid bodies. Using decoupled factor graph optimization to revise their shapes and trajectories, we obtain an iterative scheme to update both cluster assignments and motion estimation reciprocally. Evaluations on both synthetic scenes and KITTI demonstrate the capability of our approach, and further experiments considering online efficiency also show the effectiveness of our method for simultaneously tracking ego-motion and multiple objects. Sheng Yang 0007, Yukun Lai, Shi-Min Hu 0001 |
Comput. Vis. Media | 5 |
| 2021 | Jittor-GAN: A fast-training generative adversarial network model zoo based on Jittor
Wenyang Zhou, Shi-Min Hu 0001 |
Comput. Vis. Media | 3 |
| 2021 | Preface
Shi-Min Hu 0001, Connelly Barnes, Changhe Tu |
J. Comput. Sci. Technol. | 1 |
| 2021 | Hierarchical Generation of Human Pose With Part-Based Layer RepresentationabstractHuman pose transfer has been becoming one of the emerging research topics in recent years. However, state-of-the-art results are still far from satisfactory. One main reason is that these end-to-end methods are often blindly trained without the semantic understanding of its content. In this paper, we propose a novel method for human pose transfer with consideration of the semantic part-based representation of a human. In particular, we propose to segment the human body into multiple parts, and each of them represents a semantic region of a human. With the proposed part-based layer generators, a high-quality result is guaranteed for each local semantic region. We design a three-stage hierarchical framework to fuse local representations into the final result in a coarse-to-fine manner, which provides adaptive attention for global consistency and local details, respectively. Via exploiting spatial guidance from 3D human model through the framework, our method can naturally handle the ambiguity of self-occlusions which always causes artifacts in previous methods. With semantic-aware and spatial-aware representations, our method outperforms previous approaches quantitatively and qualitatively in better handling self-occlusions, fine detail preservation/synthesis and a higher resolution result. Xian Wu 0004, Chen Li 0031, Shi-Min Hu 0001, Yu-Wing Tai |
IEEE Trans. Image Process. | 3 |
| 2021 | ChoreoMaster: choreography-oriented music-driven dance synthesisabstractDespite strong demand in the game and film industry, automatically synthesizing high-quality dance motions remains a challenging task. In this paper, we present ChoreoMaster, a production-ready music-driven dance motion synthesis system. Given a piece of music, ChoreoMaster can automatically generate a high-quality dance motion sequence to accompany the input music in terms of style, rhythm and structure. To achieve this goal, we introduce a novel choreography-oriented choreomusical embedding framework, which successfully constructs a unified choreomusical embedding space for both style and rhythm relationships between music and dance phrases. The learned choreomusical embedding is then incorporated into a novel choreography-oriented graph-based motion synthesis framework, which can robustly and efficiently generate high-quality dance motions following various choreographic rules. Moreover, as a production-ready system, ChoreoMaster is sufficiently controllable and comprehensive for users to produce desired results. Experimental results demonstrate that dance motions generated by ChoreoMaster are accepted by professional artists. Jin Lei, Song-Hai Zhang, Shi-Min Hu 0001 |
ACM Trans. Graph. | 7 |
| 2021 | MoCap-solver: a neural solver for optical motion capture dataabstractIn a conventional optical motion capture (MoCap) workflow, two processes are needed to turn captured raw marker sequences into correct skeletal animation sequences. Firstly, various tracking errors present in the markers must be fixed ( cleaning or refining ). Secondly, an agent skeletal mesh must be prepared for the actor/actress, and used to determine skeleton information from the markers ( re-targeting or solving ). The whole process, normally referred to as solving MoCap data, is extremely time-consuming, labor-intensive, and usually the most costly part of animation production. Hence, there is a great demand for automated tools in industry. In this work, we present MoCap-Solver, a production-ready neural solver for optical MoCap data. It can directly produce skeleton sequences and clean marker sequences from raw MoCap markers, without any tedious manual operations. To achieve this goal, our key idea is to make use of neural encoders concerning three key intrinsic components: the template skeleton, marker configuration and motion, and to learn to predict these latent vectors from imperfect marker sequences containing noise and errors. By decoding these components from latent vectors, sequences of clean markers and skeletons can be directly recovered. Moreover, we also provide a novel normalization strategy based on learning a pose-dependent marker reliability function, which greatly improves system robustness. Experimental results demonstrate that our algorithm consistently outperforms the state-of-the-art on both synthetic and real-world datasets. Yupan Wang, Song-Hai Zhang, Sen-Zhe Xu 0001, Shi-Min Hu 0001 |
ACM Trans. Graph. | 6 |
| 2021 | Supervoxel Convolution for Online 3D Semantic SegmentationabstractOnline 3D semantic segmentation, which aims to perform real-time 3D scene reconstruction along with semantic segmentation, is an important but challenging topic. A key challenge is to strike a balance between efficiency and segmentation accuracy. There are very few deep-learning-based solutions to this problem, since the commonly used deep representations based on volumetric-grids or points do not provide efficient 3D representation and organization structure for online segmentation. Observing that on-surface supervoxels, i.e., clusters of on-surface voxels, provide a compact representation of 3D surfaces and brings efficient connectivity structure via supervoxel clustering, we explore a supervoxel-based deep learning solution for this task. To this end, we contribute a novel convolution operation (SVConv) directly on supervoxels. SVConv can efficiently fuse the multi-view 2D features and 3D features projected on supervoxels during the online 3D reconstruction, and leads to an effective supervoxel-based convolutional neural network, termed as Supervoxel-CNN , enabling 2D-3D joint learning for 3D semantic prediction. With the Supervoxel-CNN , we propose a clustering-then-prediction online 3D semantic segmentation approach. The extensive evaluations on the public 3D indoor scene datasets show that our approach significantly outperforms the existing online semantic segmentation systems in terms of efficiency or accuracy. Shi-Sheng Huang, Tai-Jiang Mu, Hongbo Fu 0001, Shi-Min Hu 0001 |
ACM Trans. Graph. | 5 |
| 2021 | Fast and accurate spherical harmonics productsabstractSpherical Harmonics (SH) have been proven as a powerful tool for rendering, especially in real-time applications such as Precomputed Radiance Transfer (PRT). Spherical harmonics are orthonormal basis functions and are efficient in computing dot products. However, computations of triple product and multiple product operations are often the bottlenecks that prevent moderately high-frequency use of spherical harmonics. Specifically state-of-the-art methods for accurate SH triple products of order n have a time complexity of O ( n 5 ), which is a heavy burden for most real-time applications. Even worse, a brute-force way to compute k -multiple products would take O ( n 2 k ) time. In this paper, we propose a fast and accurate method for spherical harmonics triple products with the time complexity of only O ( n 3 ), and further extend it for computing k -multiple products with the time complexity of O ( kn 3 + k 2 n 2 log ( kn )). Our key insight is to conduct the triple and multiple products in the Fourier space, in which the multiplications can be performed much more efficiently. To our knowledge, our method is theoretically the fastest for accurate spherical harmonics triple and multiple products. And in practice, we demonstrate the efficiency of our method in rendering applications including mid-frequency relighting and shadow fields. Hanggao Xin, Zhiqian Zhou, Di An, Lingqi Yan 0001, Kun Xu 0003, Shi-Min Hu 0001, Shing-Tung Yau |
ACM Trans. Graph. | 6 |
| 2021 | High-Quality Textured 3D Shape Reconstruction with Cascaded Fully Convolutional NetworksabstractWe present a learning-based approach to reconstructing high-resolution three-dimensional (3D) shapes with detailed geometry and high-fidelity textures. Albeit extensively studied, algorithms for 3D reconstruction from multi-view depth-and-color (RGB-D) scans are still prone to measurement noise and occlusions; limited scanning or capturing angles also often lead to incomplete reconstructions. Propelled by recent advances in 3D deep learning techniques, in this paper, we introduce a novel computation- and memory-efficient cascaded 3D convolutional network architecture, which learns to reconstruct implicit surface representations as well as the corresponding color information from noisy and imperfect RGB-D maps. The proposed 3D neural network performs reconstruction in a progressive and coarse-to-fine manner, achieving unprecedented output resolution and fidelity. Meanwhile, an algorithm for end-to-end training of the proposed cascaded structure is developed. We further introduce Human10, a newly created dataset containing both detailed and textured full-body reconstructions as well as corresponding raw RGB-D scans of 10 subjects. Qualitative and quantitative experimental results on both synthetic and real-world datasets demonstrate that the presented approach outperforms existing state-of-the-art work regarding visual quality and accuracy of reconstructed models. Zheng-Ning Liu, Yan-Pei Cao 0001, Zheng-Fei Kuang, Leif Kobbelt, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2021 | Prominent Structures for Video Analysis and EditingabstractWe present prominent structures in video, a representation of visually strong, spatially sparse and temporally stable structural units, for use in video analysis and editing. With a novel quality measurement of prominent structures in video, we develop a general framework for prominent structure computation, and an efficient hierarchical structure alignment algorithm between a pair of videos. The prominent structural unit map is proposed to encode both binary prominence guidance and numerical strength and geometry details for each video frame. Even though the detailed appearance of videos could be visually different, the proposed alignment algorithm can find matched prominent structure sub-volumes. Prominent structures in video support a wide range of video analysis and editing applications including graphic match-cut between successive videos, instant cut editing, finding transition portals from a video collection, structure-aware video re-ranking, visualizing human action differences, etc. Miao Wang 0004, Xiaonan Fang 0001, Ariel Shamir, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2021 | Effects of virtual environment and self-representations on perception and physical performance in redirected jumpingabstractRedirected jumping (RDJ) allows users to explore virtual environments (VEs) naturally by scaling a small real-world jump to a larger virtual jump with virtual camera motion manipulation, thereby addressing the problem of limited physical space in VR applications. Previous RDJ studies have mainly focused on detection threshold estimation. However, the effect VE or selfrepresentation (SR) has on the perception or performance of RDJs remains unclear. In this paper, we report experiments to measure the perception (detection thresholds for gains, presence, embodiment, intrinsic motivation, and cybersickness) and physical performance (heart rate intensity, preparation time, and actual jumping distance) of redirected forward jumping under six different combinations of VE (low and high visual richness) and SRs (invisible, shoes, and human-like). Our results indicated that the detection threshold ranges for horizontal translation gains were significantly smaller in the VE with high rather than low visual richness. When different SRs were applied, our results did not suggest significant differences in detection thresholds, but it did report longer actual jumping distances in the invisible body case compared with the other two SRs. In the high visual richness VE, the preparation time for jumping with a human-like avatar was significantly longer than that with other SRs. Finally, some correlations were found between perception and physical performance measures. All these findings suggest that both VE and SRs influence users' perception and performance in RDJ and must be considered when designing locomotion techniques. Yijun Li 0006, Miao Wang 0004, De-Rong Jin, Frank Steinicke, Shi-Min Hu 0001, Qinping Zhao |
Virtual Real. Intell. Hardw. | 5 |
| 2021 | Locomotion perception and redirection
Miao Wang 0004, Songhai Zhang, Shi-Min Hu 0001 |
Virtual Real. Intell. Hardw. | 3 |
| 2020 | Morphing and Sampling Network for Dense Point Cloud Completionabstract3D point cloud completion, the task of inferring the complete geometric shape from a partial point cloud, has been attracting attention in the community. For acquiring high-fidelity dense point clouds and avoiding uneven distribution, blurred details, or structural loss of existing methods' results, we propose a novel approach to complete the partial point cloud in two stages. Specifically, in the first stage, the approach predicts a complete but coarse-grained point cloud with a collection of parametric surface elements. Then, in the second stage, it merges the coarse-grained prediction with the input point cloud by a novel sampling algorithm. Our method utilizes a joint loss function to guide the distribution of the points. Extensive experiments verify the effectiveness of our method and demonstrate that it outperforms the existing methods in both the Earth Mover's Distance (EMD) and the Chamfer Distance (CD). Minghua Liu, Lu Sheng, Sheng Yang 0007, Shi-Min Hu 0001 |
AAAI | 5 |
| 2020 | ClusterVO: Clustering Moving Instances and Estimating Visual Odometry for Self and SurroundingsabstractWe present ClusterVO, a stereo Visual Odometry which simultaneously clusters and estimates the motion of both ego and surrounding rigid clusters/objects. Unlike previous solutions relying on batch input or imposing priors on scene structure or dynamic object models, ClusterVO is online, general and thus can be used in various scenarios including indoor scene understanding and autonomous driving. At the core of our system lies a multi-level probabilistic association mechanism and a heterogeneous Conditional Random Field (CRF) clustering approach combining semantic, spatial and motion information to jointly infer cluster segmentations online for every frame. The poses of camera and dynamic objects are instantly solved through a sliding-window optimization. Our system is evaluated on Oxford Multimotion and KITTI dataset both quantitatively and qualitatively, reaching comparable results to state-of-the-art solutions on both odometry and dynamic trajectory recovery. Sheng Yang 0007, Tai-Jiang Mu, Shi-Min Hu 0001 |
CVPR | 4 |
| 2020 | Lidar-Monocular Visual Odometry using Point and Line FeaturesabstractWe introduce a novel lidar-monocular visual odometry approach using point and line features. Compared to previous point-only based lidar-visual odometry, our approach leverages more environment structure information by introducing both point and line features into pose estimation. We provide a robust method for point and line depth extraction, and formulate the extracted depth as prior factors for point-line bundle adjustment. This method greatly reduces the features' 3D ambiguity and thus improves the pose estimation accuracy. Besides, we also provide a purely visual motion tracking method and a novel scale correction scheme, leading to an efficient lidar-monocular visual odometry system with high accuracy. The evaluations on the public KITTI odometry benchmark show that our technique achieves more accurate pose estimation than the state-of-the-art approaches, and is sometimes even better than those leveraging semantic information. Shi-Sheng Huang, Tai-Jiang Mu, Hongbo Fu 0001, Shi-Min Hu 0001 |
ICRA | 5 |
| 2020 | Message from the ISMAR 2020 Science and Technology Program Chairs
Shi-Min Hu 0001, Denis Kalkofen, Jonathan Ventura, Stefanie Zollmann |
ISMAR | 1 |
| 2020 | Message from the ISMAR 2020 Science and Technology Program Chairs and TVCG Guest EditorsabstractIn this special issue of IEEE Transactions on Visualization and Computer Graphics (TVCG), we are pleased to present the TVCG papers from the 19th IEEE International Symposium on Mixed and Augmented Reality (ISMAR 2020), which had been originally planned to hold in Recife/Porto de Galinhas, Brazil. In order to preserve the safety and well-being of all participants under the global pandemic of COVID-19, ISMAR 2020 will be held as a virtual conference between November 9 and 13, 2020. ISMAR continues the over twenty year long tradition of IWAR, ISMR, and ISAR, and is undoubtedly the premier conference for Mixed and Augmented Reality in the world. Shi-Min Hu 0001, Denis Kalkofen, Jonathan Ventura, Stefanie Zollmann |
ISMAR | 1 |
| 2020 | Transitioning360: Content-aware NFoV Virtual Camera Paths for 360° Video PlaybackabstractDespite the increasing number of head-mounted displays, many 360° VR videos are still being viewed by users on existing 2D displays. To this end, a subset of the 360° video content is often shown inside a manually or semi-automatically selected normal-field-of-view (NFoV) window. However, during the playback, simply watching an NFoV video can easily miss concurrent off-screen content. We present Transitioning360, a tool for 360° video navigation and playback on 2D displays by transitioning between multiple NFoV views that track potentially interesting targets or events. Our method computes virtual NFoV camera paths considering content awareness and diversity in an offline preprocess. During playback, the user can watch any NFoV view corresponding to a precomputed camera path. Moreover, our interface shows other candidate views, providing a sense of concurrent events. At any time, the user can transition to other candidate views for fast navigation and exploration. Experimental results including a user study demonstrate that the viewing experience using our method is more enjoyable and convenient than previous methods. Miao Wang 0004, Yijun Li 0006, Christian Richardt, Shi-Min Hu 0001 |
ISMAR | 5 |
| 2020 | Fuzzing Error Handling Code using Context-Sensitive Software Fault Injection
Zu-Ming Jiang, Jia-Ju Bai, Kangjie Lu, Shi-Min Hu 0001 |
USENIX Security Symposium | 4 |
| 2020 | A Divergence-free Mixture Model for Multiphase FluidsabstractAbstract We present a novel divergence free mixture model for multiphase flows and the related fluid‐solid coupling. The new mixture model is built upon a volume‐weighted mixture velocity so that the divergence free condition is satisfied for miscible and immiscible multiphase fluids. The proposed mixture velocity can be solved efficiently by adapted single phase incompressible solvers, allowing for larger time steps and smaller volume deviations. Besides, the drift velocity formulation is corrected to ensure mass conservation during the simulation. The new approach increases the accuracy of multiphase fluid simulation by several orders. The capability of the new divergence‐free mixture model is demonstrated by simulating different multiphase flow phenomena including mixing and unmixing of multiple fluids, fluid‐solid coupling involving deformable solids and granular materials. Yuntao Jiang, Chenfeng Li, Shujie Deng, Shi-Min Hu 0001 |
Comput. Graph. Forum | 4 |
| 2020 | Jittor: a novel deep learning framework with meta-operators and unified graph execution
Shi-Min Hu 0001, Dun Liang, Guo-Ye Yang, Wenyang Zhou |
Sci. China Inf. Sci. | 1 |
| 2020 | S4Net: Single stage salient-instance segmentationabstractIn this paper, we consider salient instance segmentation. As well as producing bounding boxes, our network also outputs high-quality instance-level segments as initial selections to indicate the regions of interest. Taking into account the category-independent property of each target, we design a single stage salient instance segmentation framework, with a novel segmentation branch. Our new branch regards not only local context inside each detection window but also the surrounding context, enabling us to distinguish instances in the same scope even with partial occlusion. Our network is end-to-end trainable and is fast (running at 40 fps for images with resolution 320 × 320). We evaluate our approach on a publicly available benchmark and show that it outperforms alternative solutions. We also provide a thorough analysis of our design choices to help readers better understand the function of each part of our network. Source code can be found at https://github.com/RuochenFan/S4Net. Ruochen Fan, Ming-Ming Cheng, Qibin Hou, Tai-Jiang Mu, Jingdong Wang 0001, Shi-Min Hu 0001 |
Comput. Vis. Media | 6 |
| 2020 | Message from the Editor-in-ChiefabstractThe current impact factor of Computational Visual Media is 1.82 according to Scopus CiteScore.Following the success of the past four years, Tsinghua University Press will continue to sponsor an annual award for the best papers published in Computational Visual Media.After carefully considering the papers published in 2019, the Editorial Board chose the paper Unsupervised natural image patch learning [1] as the winner of the Best Paper Award; while three other papers: Salient object detection: A survey [2], Reconstructing piecewise planar scenes with multi-view regularization [3], and ShadowGAN: Shadow synthesis for virtual objects with conditional adversarial networks [4], have won the Honorable Mention Awards. Shi-Min Hu 0001 |
Comput. Vis. Media | 1 |
| 2020 | Preface
Shi-Min Hu 0001, Ying He 0001, Belén Masiá |
J. Comput. Sci. Technol. | 1 |
| 2020 | Shallow2Deep: Indoor scene modeling by single image understanding
Yinyu Nie, Shihui Guo, Jian Chang 0001, Xiaoguang Han 0001, Shi-Min Hu 0001, Jian J. Zhang 0001 |
Pattern Recognit. | 6 |
| 2020 | Temporally Coherent Video Harmonization Using Adversarial NetworksabstractCompositing is one of the most important editing operations for images and videos. The process of improving the realism of composite results is often called harmonization. Previous approaches for harmonization mainly focus on images. In this paper, we take one step further to attack the problem of video harmonization. Specifically, we train a convolutional neural network in an adversarial way, exploiting a pixel-wise disharmony discriminator to achieve more realistic harmonized results and introducing a temporal loss to increase temporal consistency between consecutive harmonized frames. Thanks to the pixel-wise disharmony discriminator, we are also able to relieve the need of input foreground masks. Since existing video datasets which have ground-truth foreground masks and optical flows are not sufficiently large, we propose a simple yet efficient method to build up a synthetic dataset supporting supervised training of the proposed adversarial network. The experiments show that training on our synthetic dataset generalizes well to the real-world composite dataset. In addition, our method successfully incorporates temporal consistency during training and achieves more harmonious visual results than previous methods. Hao-Zhi Huang 0001, Sen-Zhe Xu 0001, Junxiong Cai, Wei Liu 0005, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 5 |
| 2020 | Deep Portrait Image Completion and ExtrapolationabstractGeneral image completion and extrapolation methods often fail on portrait images where parts of the human body need to be recovered -a task that requires accurate human body structure and appearance synthesis. We present a twostage deep learning framework for tackling this problem. In the first stage, given a portrait image with an incomplete human body, we extract a complete, coherent human body structure through a human parsing network, which focuses on structure recovery inside the unknown region with the help of full-body pose estimation. In the second stage, we use an image completion network to fill the unknown region, guided by the structure map recovered in the first stage. For realistic synthesis the completion network is trained with both perceptual loss and conditional adversarial loss.We further propose a face refinement network to improve the fidelity of the synthesized face region. We evaluate our method on publicly-available portrait image datasets, and show that it outperforms other state-of-the-art general image completion methods. Our method enables new portrait image editing applications such as occlusion removal and portrait extrapolation. We further show that the proposed general learning framework can be applied to other types of images, e.g. animal images. Xian Wu 0004, Ruilong Li, Jian-Cheng Liu, Jue Wang 0001, Ariel Shamir, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 7 |
| 2020 | A Metric for Video Blending Quality AssessmentabstractWe propose an objective approach to assess the quality of video blending. Blending is a fundamental operation in video editing, which can smooth the intensity changes of relevant regions. However blending also generates artefacts such as bleeding and ghosting. To assess the quality of the blended videos, our approach considers the illuminance consistency as a positive aspect while regard the artefacts as a negative aspect. Temporal coherence between frames is also considered. We evaluate our metric on a video blending dataset where the results of subjective evaluation are available. Experimental results validate the effectiveness of our proposed metric, and shows that this metric gives superior performance over existing video quality metrics. Zhe Zhu, Hantao Liu, Jiaming Lu, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 4 |
| 2020 | A moving least square reproducing kernel particle method for unified multiphase continuum simulationabstractIn physically based-based animation, pure particle methods are popular due to their simple data structure, easy implementation, and convenient parallelization. As a pure particle-based method and using Galerkin discretization, the Moving Least Square Reproducing Kernel Method (MLSRK) was developed in engineering computation as a general numerical tool for solving PDEs. The basic idea of Moving Least Square (MLS) has also been used in computer graphics to estimate deformation gradient for deformable solids. Based on these previous studies, we propose a multiphase MLSRK framework that animates complex and coupled fluids and solids in a unified manner. Specifically, we use the Cauchy momentum equation and phase field model to uniformly capture the momentum balance and phase evolution/interaction in a multiphase system, and systematically formulate the MLSRK discretization to support general multiphase constitutive models. A series of animation examples are presented to demonstrate the performance of our new multiphase MLSRK framework, including hyperelastic, elastoplastic, viscous, fracturing and multiphase coupling behaviours etc. Xiao-Song Chen, Chenfeng Li, Geng-Chen Cao, Yun-Tao Jiang, Shi-Min Hu 0001 |
ACM Trans. Graph. | 5 |
| 2020 | Noise-Resilient Reconstruction of Panoramas and 3D Scenes Using Robot-Mounted Unsynchronized Commodity RGB-D CamerasabstractWe present a two-stage approach to first constructing 3D panoramas and then stitching them for noise-resilient reconstruction of large-scale indoor scenes. Our approach requires multiple unsynchronized RGB-D cameras, mounted on a robot platform, which can perform in-place rotations at different locations in a scene. Such cameras rotate on a common (but unknown) axis, which provides a novel perspective for coping with unsynchronized cameras, without requiring sufficient overlap of their Field-of-View (FoV). Based on this key observation, we propose novel algorithms to track these cameras simultaneously. Furthermore, during the integration of raw frames onto an equirectangular panorama, we derive uncertainty estimates from multiple measurements assigned to the same pixels. This enables us to appropriately model the sensing noise and consider its influence, so as to achieve better noise resilience, and improve the geometric quality of each panorama and the accuracy of global inter-panorama registration. We evaluate and demonstrate the performance of our proposed method for enhancing the geometric quality of scene reconstruction from both real-world and synthetic scans. Sheng Yang 0007, Beichen Li 0005, Yan-Pei Cao 0001, Hongbo Fu 0001, Yukun Lai, Leif Kobbelt, Shi-Min Hu 0001 |
ACM Trans. Graph. | 7 |
| 2020 | Poisson Vector Graphics (PVG)abstractThis paper presents Poisson vector graphics (PVG), an extension of the popular diffusion curves (DC), for generating smooth-shaded images. Armed with two new types of primitives, called Poisson curves and Poisson regions, PVG can easily produce photorealistic effects such as specular highlights, core shadows, translucency and halos. Within the PVG framework, the users specify color as the Dirichlet boundary condition of diffusion curves and control tone by offsetting the Laplacian of colors, where both controls are simply done by mouse click and slider dragging. PVG distinguishes itself from other diffusion based vector graphics for 3 unique features: 1) explicit separation of colors and tones, which follows the basic drawing principle and eases editing; 2) native support of seamless cloning in the sense that PCs and PRs can automatically fit into the target background; and 3) allowed intersecting primitives (except for DC-DC intersection) so that users can create layers. Through extensive experiments and a preliminary user study, we demonstrate that PVG is a simple yet powerful authoring tool that can produce photo-realistic vector graphics from scratch. Fei Hou 0001, Qian Sun 0003, Zheng Fang 0008, Yong-Jin Liu 0001, Shi-Min Hu 0001, Hong Qin 0001, Aimin Hao, Ying He 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2020 | Semantic Labeling and Instance Segmentation of 3D Point Clouds Using Patch Context Analysis and Multiscale ProcessingabstractWe present a novel algorithm for semantic segmentation and labeling of 3D point clouds of indoor scenes, where objects in point clouds can have significant variations and complex configurations. Effective segmentation methods decomposing point clouds into semantically meaningful pieces are highly desirable for object recognition, scene understanding, scene modeling, etc. However, existing segmentation methods based on low-level geometry tend to either under-segment or over-segment point clouds. Our method takes a fundamentally different approach, where semantic segmentation is achieved along with labeling. To cope with substantial shape variation for objects in the same category, we first segment point clouds into surface patches and use unsupervised clustering to group patches in the training set into clusters, providing an intermediate representation for effectively learning patch relationships. During testing, we propose a novel patch segmentation and classification framework with multiscale processing, where the local segmentation level is automatically determined by exploiting the learned cluster based contextual information. Our method thus produces robust patch segmentation and semantic labeling results, avoiding parameter sensitivity. We further learn object-cluster relationships from the training set, and produce semantically meaningful object level segmentation. Our method outperforms state-of-the-art methods on several representative point cloud datasets, including S3DIS, SceneNN, Cornell RGB-D and ETH. Shi-Min Hu 0001, Junxiong Cai, Yukun Lai |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2020 | Message from the ISMAR 2020 Science and Technology Program Chairs and TVCG Guest EditorsabstractIn this special issue ofIEEE Transactions on Visualization and Computer Graphics (TVCG), we are pleased to present theTVCGpapers from the 19th IEEE International Symposium on Mixed and Augmented Reality (ISMAR 2020), which had been originally planned to hold in Recife/Porto de Galinhas, Brazil. In order to preserve the safety and well-being of all participants under the global pandemic of COVID-19, ISMAR 2020 will be held as a virtual conference between November 9 and 13, 2020. ISMAR continues the over twenty year long tradition of IWAR, ISMR, and ISAR, and is undoubtedly the premier conference for Mixed and Augmented Reality in the world. Shi-Min Hu 0001, Denis Kalkofen, Jonathan Ventura, Stefanie Zollmann |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2020 | Photorealistic Audio-driven Video PortraitsabstractVideo portraits are common in a variety of applications, such as videoconferencing, news broadcasting, and virtual education and training. We present a novel method to synthesize photorealistic video portraits for an input portrait video, automatically driven by a person's voice. The main challenge in this task is the hallucination of plausible, photorealistic facial expressions from input speech audio. To address this challenge, we employ a parametric 3D face model represented by geometry, facial expression, illumination, etc., and learn a mapping from audio features to model parameters. The input source audio is first represented as a high-dimensional feature, which is used to predict facial expression parameters of the 3D face model. We then replace the expression parameters computed from the original target video with the predicted one, and rerender the reenacted face. Finally, we generate a photorealistic video portrait from the reenacted synthetic face sequence via a neural face renderer. One appealing feature of our approach is the generalization capability for various input speech audio, including synthetic speech audio from text-to-speech software. Extensive experimental results show that our approach outperforms previous general-purpose audio-driven video portrait methods. This includes a user study demonstrating that our results are rated as more realistic than previous methods. Miao Wang 0004, Christian Richardt, Ze-Yin Chen, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2020 | HeteroFusion: Dense Scene Reconstruction Integrating Multi-SensorsabstractWe present a novel approach to integrate data from multiple sensor types for dense 3D reconstruction of indoor scenes in realtime. Existing algorithms are mainly based on a single RGBD camera and thus require continuous scanning of areas with sufficient geometric features. Otherwise, tracking may fail due to unreliable frame registration. Inspired by the fact that the fusion of multiple sensors can combine their strengths towards a more robust and accurate self-localization, we incorporate multiple types of sensors which are prevalent in modern robot systems, including a 2D range sensor, an inertial measurement unit (IMU), and wheel encoders. We fuse their measurements to reinforce the tracking process and to eventually obtain better 3D reconstructions. Specifically, we develop a 2D truncated signed distance field (TSDF) volume representation for the integration and ray-casting of laser frames, leading to a unified cost function in the pose estimation stage. For validation of the estimated poses in the loop-closure optimization process, we train a classifier for the features extracted from heterogeneous sensors during the registration progress. To evaluate our method on challenging use case scenarios, we assembled a scanning platform prototype to acquire real-world scans. We further simulated synthetic scans based on high-fidelity synthetic scenes for quantitative evaluation. Extensive experimental evaluation on these two types of scans demonstrate that our system is capable of robustly acquiring dense 3D reconstructions and outperforms state-of-the-art RGBD and LiDAR systems. Sheng Yang 0007, Beichen Li 0005, Minghua Liu, Yukun Lai, Leif Kobbelt, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2020 | Photographic style transferabstractImage style transfer has attracted much attention in recent years. However, results produced by existing works still have lots of distortions. This paper investigates the CNN-based artistic style transfer work specifically and finds out the key reasons for distortion coming from twofold: the loss of spatial structures of content image during content-preserving process and unexpected geometric matching introduced by style transformation process. To tackle this problem, this paper proposes a novel approach consisting of a dual-stream deep convolution network as the loss network and edge-preserving filters as the style fusion model. Our key contribution is the introduction of an additional similarity loss function that constrains both the detail reconstruction and style transfer procedures. The qualitative evaluation shows that our approach successfully suppresses the distortions as well as obtains faithful stylized results compared to state-of-the-art methods. Li Wang 0105, Xiaosong Yang, Shi-Min Hu 0001, Jian J. Zhang 0001 |
Vis. Comput. | 4 |
| 2019 | DCNS: Automated Detection Of Conservative Non-Sleep Defects in the Linux KernelabstractFor waiting, the Linux kernel offers both sleep-able and non-sleep operations. However, only non-sleep operations can be used in atomic context. Detecting the possibility of execution in atomic context requires a complete inter-procedural flow analysis, often involving function pointers. Developers may thus conservatively use non-sleep operations even outside of atomic context, which may damage system performance, as such operations unproductively monopolize the CPU. Until now, no systematic approach has been proposed to detect such conservative non-sleep (CNS) defects. Jia-Ju Bai, Julia Lawall, Wende Tan, Shi-Min Hu 0001 |
ASPLOS | 4 |
| 2019 | Example-Guided Style-Consistent Image Synthesis From Semantic LabelingabstractExample-guided image synthesis aims to synthesize an image from a semantic label map and an exemplary image indicating style. We use the term "style" in this problem to refer to implicit characteristics of images, for example: in portraits "style" includes gender, racial identity, age, hairstyle; in full body pictures it includes clothing; in street scenes it refers to weather and time of day and such like. A semantic label map in these cases indicates facial expression, full body pose, or scene segmentation. We propose a solution to the example-guided image synthesis problem using conditional generative adversarial networks with style consistency. Our key contributions are (i) a novel style consistency discriminator to determine whether a pair of images are consistent in style; (ii) an adaptive semantic consistency loss; and (iii) a training data sampling strategy, for synthesizing style-consistent results to the exemplar. We demonstrate the efficiency of our method on face, dance and street view synthesis tasks. Miao Wang 0004, Guo-Ye Yang, Ruilong Li, Runze Liang, Song-Hai Zhang, Peter Hall 0001, Shi-Min Hu 0001 |
CVPR | 7 |
| 2019 | S4Net: Single Stage Salient-Instance SegmentationabstractWe consider an interesting problem---salient instance segmentation. Other than producing approximate bounding boxes, our network also outputs high-quality instance-level segments. Taking into account the category-independent property of each target, we design a single stage salient instance segmentation framework, with a novel segmentation branch. Our new branch regards not only local context inside each detection window but also its surrounding context, enabling us to distinguish the instances in the same scope even with obstruction. Our network is end-to-end trainable and runs at a fast speed (40 fps when processing an image with resolution 320 x 320). We evaluate our approach on a public available benchmark and show that it outperforms other alternative solutions. We also provide a thorough analysis of the design choices to help readers better understand the functions of each part of our network. The source code can be found at https://github.com/RuochenFan/S4Net. Ruochen Fan, Ming-Ming Cheng, Qibin Hou, Tai-Jiang Mu, Jingdong Wang 0001, Shi-Min Hu 0001 |
CVPR | 6 |
| 2019 | Pose2Seg: Detection Free Human Instance SegmentationabstractThe standard approach to image instance segmentation is to perform the object detection first, and then segment the object from the detection bounding-box. More recently, deep learning methods like Mask R-CNN perform them jointly. However, little research takes into account the uniqueness of the "human" category, which can be well defined by the pose skeleton. Moreover, the human pose skeleton can be used to better distinguish instances with heavy occlusion than using bounding-boxes. In this paper, we present a brand new pose-based instance segmentation framework for humans which separates instances based on human pose, rather than proposal region detection. We demonstrate that our pose-based framework can achieve better accuracy than the state-of-art detection-based approach on the human instance segmentation problem, and can moreover better handle occlusion. Furthermore, there are few public datasets containing many heavily occluded humans along with comprehensive annotations, which makes this a challenging problem seldom noticed by researchers. Therefore, in this paper we introduce a new benchmark "Occluded Human (OCHuman)", which focuses on occluded humans with comprehensive annotations including bounding-box, human pose and instance masks. This dataset contains 8110 detailed annotated human instances within 4731 images. With an average 0.67 MaxIoU for each person, OCHuman is the most complex and challenging dataset related to human instance segmentation. Through this dataset, we want to emphasize occlusion as a challenging problem for researchers to study. Song-Hai Zhang, Ruilong Li, Paul L. Rosin, Zixi Cai, Dingcheng Yang, Hao-Zhi Huang 0001, Shi-Min Hu 0001 |
CVPR | 9 |
| 2019 | ClusterSLAM: A SLAM Backend for Simultaneous Rigid Body Clustering and Motion EstimationabstractWe present a practical backend for stereo visual SLAM which can simultaneously discover individual rigid bodies and compute their motions in dynamic environments. While recent factor graph based state optimization algorithms have shown their ability to robustly solve SLAM problems by treating dynamic objects as outliers, the dynamic motions are rarely considered. In this paper, we exploit the consensus of 3D motions among the landmarks extracted from the same rigid body for clustering and estimating static and dynamic objects in a unified manner. Specifically, our algorithm builds a noise-aware motion affinity matrix upon landmarks, and uses agglomerative clustering for distinguishing those rigid bodies. Accompanied by a decoupled factor graph optimization for revising their shape and trajectory, we obtain an iterative scheme to update both cluster assignments and motion estimation reciprocally. Evaluations on both synthetic scenes and KITTI demonstrate the capability of our approach, and further experiments considering online efficiency also show the effectiveness of our method for simultaneous tracking of ego-motion and multiple objects. Sheng Yang 0007, Yukun Lai, Shi-Min Hu 0001 |
ICCV | 5 |
| 2019 | Probabilistic Projective Association and Semantic Guided Relocalization for Dense ReconstructionabstractWe present a real-time dense mapping system which uses the predicted 2D semantic labels for optimizing the geometric quality of reconstruction. With a combination of Convolutional Neural Networks (CNNs) for 2D labeling and a Simultaneous Localization and Mapping (SLAM) system for camera trajectory estimation, recent approaches have succeeded in incrementally fusing and labeling 3D scenes. However, the geometric quality of the reconstruction can be further improved by incorporating such semantic prediction results, which is not sufficiently exploited by existing methods. In this paper, we propose to use semantic information to improve two crucial modules in the reconstruction pipeline, namely tracking and loop detection, for obtaining mutual benefits in geometric reconstruction and semantic recognition. Specifically for tracking, we use a novel probabilistic projective association approach to efficiently pick out candidate correspondences, where the confidence of these correspondences is quantified concerning similarities on all available short-term invariant features. For the loop detection, we incorporate these semantic labels into the original encoding through Randomized Ferns to generate a more comprehensive representation for retrieving candidate loop frames. Evaluations on a publicly available synthetic dataset have shown the effectiveness of our approach that considers such semantic hints as a reliable feature for achieving higher geometric quality. Sheng Yang 0007, Zheng-Fei Kuang, Yan-Pei Cao 0001, Yukun Lai, Shi-Min Hu 0001 |
ICRA | 5 |
| 2019 | TZC: Efficient Inter-Process Communication for Robotics Middleware with Partial SerializationabstractInter-process communication (IPC) is one of the core functions of modern robotics middleware. We propose an efficient IPC technique called TZC (Towards Zero-Copy). As a core component of TZC, we design a novel algorithm called partial serialization. Our formulation can generate messages that can be divided into two parts. During message transmission, one part is transmitted through a socket and the other part uses shared memory. The part within shared memory is never copied or serialized during its lifetime. We have integrated TZC with ROS and ROS2 and find that TZC can be easily combined with current open-source platforms. By using TZC, the overhead of IPC remains constant when the message size grows. In particular, when the message size is 4MB (less than the size of a full HD image), TZC can reduce the overhead of ROS IPC from tens of milliseconds to hundreds of microseconds and can reduce the overhead of ROS2 IPC from hundreds of milliseconds to less than 1 millisecond. We also demonstrate the benefits of TZC by integrating it with TurtleBot2 to be used in autonomous driving scenarios. We show that by using TZC, the braking distance can be 16% shorter than with ROS. Yu-Ping Wang 0001, Wende Tan, Xu-Qiang Hu, Dinesh Manocha, Shi-Min Hu 0001 |
IROS | 5 |
| 2019 | Fuzzing Error Handling Code in Device Drivers Based on Software Fault InjectionabstractDevice drivers remain a main source of runtime failures in operating systems. To detect bugs in device drivers, fuzzing has been commonly used in practice. However, a main limitation of existing fuzzing approaches is that they cannot effectively test error handling code. Indeed, these fuzzing approaches require effective inputs to cover target code, but much error handling code in drivers is triggered by occasional errors (such as insufficient memory and hardware malfunctions) that are not related to inputs. In this paper, based on software fault injection, we propose a new fuzzing approach named FIZZER, to test error handling code in device drivers. At compile time, FIZZER uses static analysis to recommend possible error sites that can trigger error handling code. During driver execution, by analyzing runtime information, it automatically fuzzes error-site sequences for fault injection to improve code coverage. We evaluate FIZZER on 18 device drivers in Linux 4.19, and in total find 22 real bugs. The code coverage is increased by over 15% compared to normal execution without fuzzing. Zu-Ming Jiang, Jia-Ju Bai, Julia Lawall, Shi-Min Hu 0001 |
ISSRE | 4 |
| 2019 | Effective Static Analysis of Concurrency Use-After-Free Bugs in Linux Device Drivers
Jia-Ju Bai, Julia Lawall, Qiu-Liang Chen, Shi-Min Hu 0001 |
USENIX ATC | 4 |
| 2019 | Detecting Data Races Caused by Inconsistent Lock Protection in Device DriversabstractData races are often hard to detect in device drivers, due to the non-determinism of concurrent execution. According to our study of Linux driver patches that fix data races, more than 38% of patches involve a pattern that we call inconsistent lock protection. Specifically, if a variable is accessed within two concurrently executed functions, the sets of locks held around each access are disjoint, at least one of the locksets is non-empty, and at least one of the involved accesses is a write, then a data race may occur.In this paper, we present a runtime analysis approach, named DILP, to detect data races caused by inconsistent lock protection in device drivers. By monitoring driver execution, DILP collects the information about runtime variable accesses and executed functions. Then after driver execution, DILP analyzes the collected information to detect and report data races caused by inconsistent lock protection. We evaluate DILP on 12 device drivers in Linux 4.16.9, and find 25 real data races. Qiu-Liang Chen, Jia-Ju Bai, Zu-Ming Jiang, Julia Lawall, Shi-Min Hu 0001 |
SANER | 5 |
| 2019 | Learning Explicit Smoothing Kernels for Joint Image FilteringabstractAbstract Smoothing noises while preserving strong edges in images is an important problem in image processing. Image smoothing filters can be either explicit (based on local weighted average) or implicit (based on global optimization). Implicit methods are usually time‐consuming and cannot be applied to joint image filtering tasks, i.e., leveraging the structural information of a guidance image to filter a target image. Previous deep learning based image smoothing filters are all implicit and unavailable for joint filtering. In this paper, we propose to learn explicit guidance feature maps as well as offset maps from the guidance image and smoothing parameter that can be utilized to smooth the input itself or to filter images in other target domains. We design a deep convolutional neural network consisting of a fully‐convolution block for guidance and offset maps extraction together with a stacked spatially varying deformable convolution block for joint image filtering. Our models can approximate several representative image smoothing filters with high accuracy comparable to state‐of‐the‐art methods, and serve as general tools for other joint image filtering tasks, such as color interpolation, depth map upsampling, saliency map upsampling, flash/non‐flash image denoising and RGB/NIR image denoising. Xiaonan Fang 0001, Miao Wang 0004, Ariel Shamir, Shi-Min Hu 0001 |
Comput. Graph. Forum | 4 |
| 2019 | A Rigging-Skinning Scheme to Control Fluid SimulationabstractAbstract Inspired by skeletal animation, a novel rigging‐skinning flow control scheme is proposed to animate fluids intuitively and efficiently. The new animation pipeline creates fluid animation via two steps: fluid rigging and fluid skinning. The fluid rig is defined by a point cloud with rigid‐body movement and incompressible deformation, whose time series can be intuitively specified by a rigid body motion and a constrained free‐form deformation, respectively. The fluid skin generates plausible fluid flows by virtually fluidizing the point‐cloud fluid rig with adjustable zero‐ and first‐order flow features and at fixed computational cost. Fluid rigging allows the animator to conveniently specify the desired low‐frequency flow motion through intuitive manipulations of a point cloud, while fluid skinning truthfully and efficiently converts the motion specified on the fluid rig into plausible flows of the animation fluid, with adjustable fine‐scale effects. Besides being intuitive, the rigging‐skinning scheme for fluid animation is robust and highly efficient, avoiding completely iterative trials or time‐consuming nonlinear optimization. It is also versatile, supporting both particle‐ and grid‐ based fluid solvers. A series of examples including liquid, gas and mixed scenes are presented to demonstrate the performance of the new animation pipeline. Jiaming Lu, Xiao-Song Chen, Xiao Yan 0004, Chenfeng Li, Ming C. Lin, Shi-Min Hu 0001 |
Comput. Graph. Forum | 6 |
| 2019 | Deep point-based scene labeling with depth mapping and geometric patch feature encoding
Junxiong Cai, Tai-Jiang Mu, Yukun Lai, Shi-Min Hu 0001 |
Graph. Model. | 4 |
| 2019 | Message from the Editor-in-ChiefabstractFollowing the success of the past three years, Tsinghua University Press will continue to sponsor an annual award for the best paper published in Computational Visual Media. After carefully considering the 30 papers published in 2018, the Editorial Board chose the paper: Photometric stereo for strong specular highlights [1] as the winner of the Best Paper Award, while three other papers: Knowledge graph construction with structure and parameter learning for indoor scene design [2], Robust edge-preserving surface mesh polycube deformation [3], and Transferring pose and augmenting background for deep human image parsing and its applications [4], have won Honorable Mention Awards. Shi-Min Hu 0001 |
Comput. Vis. Media | 1 |
| 2019 | Preface
Shi-Min Hu 0001, Hongbo Fu 0001, Marcus A. Magnor |
J. Comput. Sci. Technol. | 1 |
| 2019 | A Large Chinese Text Dataset in the Wild
Tailing Yuan, Zhe Zhu, Kun Xu 0003, Cheng-Jun Li, Tai-Jiang Mu, Shi-Min Hu 0001 |
J. Comput. Sci. Technol. | 6 |
| 2019 | Debugopt: Debugging fully optimized natively compiled programs using multistage instrumentation
Gang Tan, Hao Li 0056, Xiaolong Bai, Yu-Ping Wang 0001, Shi-Min Hu 0001 |
Sci. Comput. Program. | 6 |
| 2019 | Deep Online Video Stabilization With Multi-Grid Warping Transformation LearningabstractVideo stabilization techniques are essential for most hand-held captured videos due to high-frequency shakes. Several 2D, 2.5D and 3D-based stabilization techniques have been presented previously, but to our knowledge, no solutions based on deep neural networks had been proposed to date. The main reason for this omission is shortage in training data as well as the challenge of modeling the problem using neural networks. In this paper, we present a video stabilization technique using a convolutional neural network. Previous works usually propose an offline algorithm that smoothes a holistic camera path based on feature matching. Instead, we focus on low-latency, real-time camera path smoothing, that does not explicitly represent the camera path, and does not use future frames. Our neural network model, called StabNet, learns a set of mesh-grid transformations progressively for each input frame from the previous set of stabalized camera frames, and creates stable corresponding latent camera paths implicitly. To train the network, we collect a dataset of synchronized steady and unsteady video pairs via a specially designed hand-held hardware. Experimental results show that our proposed online method performs comparatively to traditional offline video stabilization methods without using future frames, while running about 10× faster. More importantly, our proposed StabNet is able to handle low-quality videos such as night-scene videos, watermarked videos, blurry videos and noisy videos, where existing methods fail in feature extraction or matching. Miao Wang 0004, Guo-Ye Yang, Jin-Kun Lin, Song-Hai Zhang, Ariel Shamir, Shao-Ping Lu, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 7 |
| 2019 | Two-Layer QR CodesabstractA quick-response code (QR code) is a two-dimensional code akin to a barcode that encodes a message of limited length. In this paper, we present a variant of QR code, a two-layer QR code. Its two-layer structure can display two alternative messages when scanned from two different directions. We propose a method to generate such two-layer QR codes encoding two given messages in a few seconds. We also demonstrate the robustness of our method on both synthetic and fabricated examples. All source code will be made publicly available (https://github.com/yuantailing/two-layer-qrcode). Tailing Yuan, Yili Wang 0003, Kun Xu 0003, Ralph R. Martin, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 5 |
| 2019 | Write-a-video: computational video montage from themed textabstractWe present Write-A-Video , a tool for the creation of video montage using mostly text-editing. Given an input themed text and a related video repository either from online websites or personal albums, the tool allows novice users to generate a video montage much more easily than current video editing tools. The resulting video illustrates the given narrative, provides diverse visual content, and follows cinematographic guidelines. The process involves three simple steps: (1) the user provides input, mostly in the form of editing the text, (2) the tool automatically searches for semantically matching candidate shots from the video repository, and (3) an optimization method assembles the video montage. Visual-semantic matching between segmented text and shots is performed by cascaded keyword matching and visual-semantic embedding, that have better accuracy than alternative solutions. The video assembly is formulated as a hybrid optimization problem over a graph of shots, considering temporal constraints, cinematography metrics such as camera movement and tone, and user-specified cinematography idioms. Using our system, users without video editing experience are able to generate appealing videos. Miao Wang 0004, Shi-Min Hu 0001, Shing-Tung Yau, Ariel Shamir |
ACM Trans. Graph. | 3 |
| 2019 | Message from the ISMAR 2019 Science and Technology Program Chairs and TVCG Guest EditorsabstractIn this special issue ofIEEE Transactions on Visualization and Computer Graphics (TVCG), we are pleased to present theTVCGpapers from the 18th IEEE International Symposium on Mixed and Augmented Reality (ISMAR 2019), held October 14–18 in Beijing, China. ISMAR continues the over 20-year long tradition of IWAR, ISMR, and ISAR, and is undoubtedly the premier conference for mixed and augmented reality in the world. Joseph L. Gabbard, Jens Grubert, Shi-Min Hu 0001, Stefanie Zollmann |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2018 | AutoPA: automatically generating active driver from original passive driver codeabstractOriginal device drivers are often passive in common operating systems, and they should correctly handle synchronization when concurrently invoked by multiple external threads. However, many concurrency bugs have occurred in drivers due to incautious synchronization. To solve concurrency problems, active driver is proposed to replace original passive driver. An active driver has its own thread and does not need to handle synchronization, thus the occurrence probability of many concurrency bugs can be effectively reduced. But previous approaches of active driver have some limitations. The biggest limitation is that original passive driver code needs to be manually rewritten. In this paper, we propose a practical approach, AutoPA, to automatically generate efficient active driver from original passive driver code. AutoPA uses function analysis and code instrumentation to perform automated driver generation, and it uses an improved active driver architecture to reduce performance degradation. We have evaluated AutoPA on 20 Linux drivers. The results show that AutoPA can automatically and successfully generate usable active drivers from original driver code. And generated active drivers can work normally with or without the synchronization primitives in original driver code. To check the effect of AutoPA on driver reliability, we perform fault injection testing on the generated active drivers, and find that all injected concurrency faults are well tolerated and the drivers can work normally. And the performance of generated active drivers is not obviously degraded compared to original passive drivers. Jia-Ju Bai, Yu-Ping Wang 0001, Shi-Min Hu 0001 |
CGO | 3 |
| 2018 | Learning to Reconstruct High-Quality 3D Shapes with Cascaded Fully Convolutional Networks
Yan-Pei Cao 0001, Zheng-Ning Liu, Zheng-Fei Kuang, Leif Kobbelt, Shi-Min Hu 0001 |
ECCV (9) | 5 |
| 2018 | Associating Inter-image Salient Instances for Weakly Supervised Semantic Segmentation
Ruochen Fan, Qibin Hou, Ming-Ming Cheng, Gang Yu 0002, Ralph R. Martin, Shi-Min Hu 0001 |
ECCV (9) | 6 |
| 2018 | DSAC: Effective Static Analysis of Sleep-in-Atomic-Context Bugs in Kernel Modules
Jia-Ju Bai, Yu-Ping Wang 0001, Julia Lawall, Shi-Min Hu 0001 |
USENIX ATC | 4 |
| 2018 | A Temporally Adaptive Material Point Method with Regional Time SteppingabstractAbstract Spatially and temporally adaptive algorithms can substantially improve the computational efficiency of many numerical schemes in computational mechanics and physics‐based animation. Recently, a crucial need for temporal adaptivity in the Material Point Method (MPM) is emerging due to the potentially substantial variation of material stiffness and velocities in multi‐material scenes. In this work, we propose a novel temporally adaptive symplectic Euler scheme for MPM with regional time stepping (RTS), where different time steps are used in different regions. We design a time stepping scheduler operating at the granularity of small blocks to maintain a natural consistency with the hybrid particle/grid nature of MPM. Our method utilizes the Sparse Paged Grid (SPGrid) data structure and simultaneously offers high efficiency and notable ease of implementation with a practical multi‐threaded particle‐grid transfer strategy. We demonstrate the efficacy of our asynchronous MPM method on various examples including elastic objects, granular media, and fluids. Yu Fang 0010, Yuanming Hu, Shi-Min Hu 0001, Chenfanfu Jiang |
Comput. Graph. Forum | 3 |
| 2018 | Controllable Dendritic Crystal Simulation Using Orientation FieldabstractAbstract Real world dendritic growths show charming structures by their exquisite balance between the symmetry and randomness in the crystal formation. Other than the variety in the natural crystals, richer visual appearance of crystals can benefit from artificially controlling of the crystal growth on its growing directions and shapes. In this paper, by introducing one extra dimension of freedom, i.e. the orientation field, into the simulation, we propose an efficient algorithm for dendritic crystal simulation that is able to reproduce arbitrary symmetry patterns with different levels of asymmetry breaking effect on general grids or meshes, including spreading on curved surfaces and growth in 3D. Flexible artistic control is also enabled in a unified manner by exploiting and guiding the orientation field in the visual simulation. We show the effectiveness of our approach by various demonstrations of simulation results. Bo Ren 0003, Ming C. Lin, Shi-Min Hu 0001 |
Comput. Graph. Forum | 4 |
| 2018 | Deep Video Stabilization Using Adversarial NetworksabstractAbstract Video stabilization is necessary for many hand‐held shot videos. In the past decades, although various video stabilization methods were proposed based on the smoothing of 2D, 2.5D or 3D camera paths, hardly have there been any deep learning methods to solve this problem. Instead of explicitly estimating and smoothing the camera path, we present a novel online deep learning framework to learn the stabilization transformation for each unsteady frame, given historical steady frames. Our network is composed of a generative network with spatial transformer networks embedded in different layers, and generates a stable frame for the incoming unstable frame by computing an appropriate affine transformation. We also introduce an adversarial network to determine the stability of apiece of video. The network is trained directly using the pair of steady and unsteady videos. Experiments show that our method can produce similar results as traditional methods, moreover, it is capable of handling challenging unsteady video of low quality, where traditional methods fail, such as video with heavy noise or multiple exposures. Our method runs in real time, which is much faster than traditional methods. Sen-Zhe Xu 0001, Miao Wang 0004, Tai-Jiang Mu, Shi-Min Hu 0001 |
Comput. Graph. Forum | 5 |
| 2018 | MPM simulation of interacting fluids and solidsabstractAbstract The material point method (MPM) has attracted increasing attention from the graphics community, as it combines the strengths of both particle‐ and grid‐based solvers. Like the smoothed particle hydrodynamics (SPH) scheme, MPM uses particles to discretize the simulation domain and represent the fundamental unknowns. This makes it insensitive to geometric and topological changes, and readily parallelizable on a GPU. Like grid‐based solvers, MPM uses a background mesh for calculating spatial derivatives, providing more accurate and more stable results than a purely particle‐based scheme. MPM has been very successful in simulating both fluid flow and solid deformation, but less so in dealing with multiple fluids and solids, where the dynamic fluid‐solid interaction poses a major challenge. To address this shortcoming of MPM, we propose a new set of mathematical and computational schemes which enable efficient and robust fluid‐solid interaction within the MPM framework. These versatile schemes support simulation of both multiphase flow and fully‐coupled solid‐fluid systems. A series of examples is presented to demonstrate their capabilities and performance in the presence of various interacting fluids and solids, including multiphase flow, fluid‐solid interaction, and dissolution. Xiao Yan 0004, Chenfeng Li, Xiao-Song Chen, Shi-Min Hu 0001 |
Comput. Graph. Forum | 4 |
| 2018 | Message from the Editor-in-ChiefabstractI would like to take Shi-Min Hu 0001 |
Comput. Vis. Media | 1 |
| 2018 | Preface
Shi-Min Hu 0001, Cewu Lu, Ariel Shamir |
J. Comput. Sci. Technol. | 1 |
| 2018 | Automated and reliable resource release in device drivers based on dynamic analysis
Jia-Ju Bai, Yu-Ping Wang 0001, Shi-Min Hu 0001 |
J. Syst. Softw. | 3 |
| 2018 | Hyper-Lapse From Multiple Spatially-Overlapping VideosabstractHyper-lapse video with high speed-up rate is an efficient way to overview long videos, such as a human activity in first-person view. Existing hyper-lapse video creation methods produce a fast-forward video effect using only one video source. In this paper, we present a novel hyper-lapse video creation approach based on multiple spatially-overlapping videos. We assume the videos share a common view or location, and find transition points where jumps from one video to another may occur. We represent the collection of videos using a hyper-lapse transition graph; the edges between nodes represent possible hyper-lapse frame transitions. To create a hyper-lapse video, a shortest path search is performed on this digraph to optimize frame sampling and assembly simultaneously. Finally, we render the hyper-lapse results using video stabilization and appearance smoothing techniques on the selected frames. Our technique can synthesize novel virtual hyper-lapse routes, which may not exist originally. We show various application results on both indoor and outdoor video collections with static scenes, moving objects, and crowds. Miao Wang 0004, Jun-Bang Liang, Song-Hai Zhang, Shao-Ping Lu, Ariel Shamir, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 6 |
| 2018 | BiggerSelfie: Selfie Video Expansion With Hand-Held CameraabstractSelfie photography from the hand-held camera is becoming a popular media type. Although being convenient and flexible, it suffers from low camera motion stability, small field of view, and limited background content. These limitations can annoy users, especially, when touring a place of interest and taking selfie videos. In this paper, we present a novel method to create what we call a BiggerSelfie that deals with these shortcomings. Using a video of the environment that has partial content overlap with the selfie video, we stitch plausible frames selected from the environment video to the original selfie frames and stabilize the composed video content with a portrait-preserving constraint. Using the proposed method, one can easily obtain a stable selfie video with expanded background content by merely capturing some background shots. We show various results and several evaluations to demonstrate the applicability of our method. Miao Wang 0004, Ariel Shamir, Guo-Ye Yang, Jin-Kun Lin, Shao-Ping Lu, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 6 |
| 2018 | A Comparative Study of Algorithms for Realtime Panoramic Video BlendingabstractUnlike image blending algorithms, video blending algorithms have been little studied. In this paper, we investigate 6 popular blending algorithms-feather blending, multi-band blending, modified Poisson blending, mean value coordinate blending, multi-spline blending and convolution pyramid blending. We consider their application to blending realtime panoramic videos, a key problem in various virtual reality tasks. To evaluate the performances and suitabilities of the 6 algorithms for this problem, we have created a video benchmark with several videos captured under various conditions. We analyze the time and memory needed by the above 6 algorithms, for both CPU and GPU implementations (where readily parallelizable). The visual quality provided by these algorithms is also evaluated both objectively and subjectively. The video benchmark and algorithm implementations are publicly available1. Zhe Zhu, Jiaming Lu, Minxuan Wang, Song-Hai Zhang, Ralph R. Martin, Hantao Liu, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 7 |
| 2018 | Detecting and Removing Visual Distractors for Video Aesthetic EnhancementabstractPersonal videos often contain visual distractors, which are objects that are accidentally captured and can distract viewers from focusing on the main subjects. We propose a method to automatically detect and localize these distractors through learning from a manually labeled dataset. To achieve spatially and temporally coherent detection, we propose extracting features at the temporal-superpixel level using a traditional supporting vector machine based learning framework. We also experiment with end-to-end learning using convolutional neural networks, which achieves slightly higher performance than other methods. The classification result is further refined in a postprocessing step based on graph-cut optimization. Experimental results show that our method achieves an accuracy of 81% and a recall of 86%. We demonstrate several ways of removing the detected distractors to improve the video quality, including video hole filling, video frame replacement, and camera path replanning. The user study results show that our method can significantly improve the aesthetic quality of videos. Xian Wu 0004, Ruilong Li, Jue Wang 0001, Zhao-Heng Zheng, Shi-Min Hu 0001 |
IEEE Trans. Multim. | 6 |
| 2018 | Effective Detection of Sleep-in-atomic-context Bugs in the Linux KernelabstractAtomic context is an execution state of the Linux kernel in which kernel code monopolizes a CPU core. In this state, the Linux kernel may only perform operations that cannot sleep, as otherwise a system hang or crash may occur. We refer to this kind of concurrency bug as a sleep-in-atomic-context (SAC) bug. In practice, SAC bugs are hard to find, as they do not cause problems in all executions. In this article, we propose a practical static approach named DSAC to effectively detect SAC bugs in the Linux kernel. DSAC uses three key techniques: (1) a summary-based analysis to identify the code that may be executed in atomic context, (2) a connection-based alias analysis to identify the set of functions referenced by a function pointer, and (3) a path-check method to filter out repeated reports and false bugs. We evaluate DSAC on Linux 4.17 and find 1,159 SAC bugs. We manually check all the bugs and find that 1,068 bugs are real. We have randomly selected 300 of the real bugs and sent them to kernel developers. 220 of these bugs have been confirmed, and 51 of our patches fixing 115 bugs have been applied. Jia-Ju Bai, Julia Lawall, Shi-Min Hu 0001 |
ACM Trans. Comput. Syst. | 3 |
| 2018 | Real-time High-accuracy Three-Dimensional Reconstruction with Consumer RGB-D CamerasabstractWe present an integrated approach for reconstructing high-fidelity three-dimensional (3D) models using consumer RGB-D cameras. RGB-D registration and reconstruction algorithms are prone to errors from scanning noise, making it hard to perform 3D reconstruction accurately. The key idea of our method is to assign a probabilistic uncertainty model to each depth measurement, which then guides the scan alignment and depth fusion. This allows us to effectively handle inherent noise and distortion in depth maps while keeping the overall scan registration procedure under the iterative closest point framework for simplicity and efficiency. We further introduce a local-to-global, submap-based, and uncertainty-aware global pose optimization scheme to improve scalability and guarantee global model consistency. Finally, we have implemented the proposed algorithm on the GPU, achieving real-time 3D scanning frame rates and updating the reconstructed model on-the-fly. Experimental results on simulated and real-world data demonstrate that the proposed method outperforms state-of-the-art systems in terms of the accuracy of both recovered camera trajectories and reconstructed models. Yan-Pei Cao 0001, Leif Kobbelt, Shi-Min Hu 0001 |
ACM Trans. Graph. | 3 |
| 2018 | Computational Design of Transforming Pop-up BooksabstractWe present the first computational tool to help ordinary users create transforming pop-up books. In each transforming pop-up, when the user pulls a tab, an initial flat two-dimensional (2D) pattern, i.e., a 2D shape with a superimposed picture, such as an airplane, turns into a new 2D pattern, such as a robot. Given the two 2D patterns, our approach automatically computes a 3D pop-up mechanism that transforms one pattern into the other; it also outputs a design blueprint, allowing the user to easily make the final model. We also present a theoretical analysis of basic transformation mechanisms; combining these basic mechanisms allows more flexibility of final designs. Using our approach, inexperienced users can create models in a short time; previously, even experienced artists often took weeks to manually create them. We demonstrate our method on a variety of real-world examples. Zhe Zhu, Ralph R. Martin, Kun Xu 0003, Jiaming Lu, Shi-Min Hu 0001 |
ACM Trans. Graph. | 6 |
| 2018 | PhotoRecomposer: Interactive Photo Recomposition by CroppingabstractWe present a visual analysis method for interactively recomposing a large number of photos based on example photos with high-quality composition. The recomposition method is formulated as a matching problem between photos. The key to this formulation is a new metric for accurately measuring the composition distance between photos. We have also developed an earth-mover-distance-based online metric learning algorithm to support the interactive adjustment of the composition distance based on user preferences. To better convey the compositions of a large number of example photos, we have developed a multi-level, example photo layout method to balance multiple factors such as compactness, aspect ratio, composition distance, stability, and overlaps. By introducing an EulerSmooth-based straightening method, the composition of each photos is clearly displayed. The effectiveness and usefulness of the method has been demonstrated by the experimental results, user study, and case studies. Xiting Wang, Song-Hai Zhang, Shi-Min Hu 0001, Shixia Liu |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2018 | Real-Time High-Fidelity Surface Flow SimulationabstractSurface flow phenomena, such as rain water flowing down a tree trunk and progressive water front in a shower room, are common in real life. However, compared with the 3D spatial fluid flow, these surface flow problems have been much less studied in the graphics community. To tackle this research gap, we present an efficient, robust and high-fidelity simulation approach based on the shallow-water equations. Specifically, the standard shallow-water flow model is extended to general triangle meshes with a feature-based bottom friction model, and a series of coherent mathematical formulations are derived to represent the full range of physical effects that are important for real-world surface flow phenomena. In addition, by achieving compatibility with existing 3D fluid simulators and by supporting physically realistic interactions with multiple fluids and solid surfaces, the new model is flexible and readily extensible for coupled phenomena. A wide range of simulation examples are presented to demonstrate the performance of the new approach. Bo Ren 0003, Tailing Yuan, Chenfeng Li, Kun Xu 0003, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2017 | Avoiding bleeding in image blendingabstractThough elegant in mathematical formulation, gradient-domain image blending suffers from bleeding artefacts in real world applications. We propose an image blending algorithm that avoids bleeding artefacts while preserving the good properties of gradient-domain blending, such as smooth transitions between the candidate regions. Our key idea to finesse the non-smooth boundary difference calculation that causes bleeding artefacts is to use local patch differences. While most previous gradient-domain blending algorithms change one region to fit the other, to further reduce bleeding, we perform bidirectional blending so that both regions change simultaneously. Our blending algorithm is fast: when applied to image stitching, it can achieve 20 fps at 4K resolution. Source code and test images are publicly available. Minxuan Wang, Zhe Zhu, Song-Hai Zhang, Ralph R. Martin, Shi-Min Hu 0001 |
ICIP | 5 |
| 2017 | Picking Up My Tab: Understanding and Mitigating Synchronized Token Lifting and Spending in Mobile Payment
Xiaolong Bai, Zhe Zhou 0001, XiaoFeng Wang 0001, Zhou Li 0001, Xianghang Mi, Nan Zhang 0018, Tongxin Li 0002, Shi-Min Hu 0001, Kehuan Zhang |
USENIX Security Symposium | 8 |
| 2017 | Extracting Sharp Features from RGB-D ImagesabstractAbstract Sharp edges are important shape features and their extraction has been extensively studied both on point clouds and surfaces. We consider the problem of extracting sharp edges from a sparse set of colour‐and‐depth (RGB‐D) images. The noise‐ridden depth measurements are challenging for existing feature extraction methods that work solely in the geometric domain (e.g. points or meshes). By utilizing both colour and depth information, we propose a novel feature extraction method that produces much cleaner and more coherent feature lines. We make two technical contributions. First, we show that intensity edges can augment the depth map to improve normal estimation and feature localization from a single RGB‐D image. Second, we designed a novel algorithm for consolidating feature points obtained from multiple RGB‐D images. By utilizing normals and ridge/valley types associated with the feature points, our algorithm is effective in suppressing noise without smearing nearby features. Yan-Pei Cao 0001, Shi-Min Hu 0001 |
Comput. Graph. Forum | 4 |
| 2017 | Saliency-aware Real-time Volumetric Fusion for Object ReconstructionabstractAbstract We present a real‐time approach for acquiring 3D objects with high fidelity using hand‐held consumer‐level RGB‐D scanning devices. Existing real‐time reconstruction methods typically do not take the point of interest into account, and thus might fail to produce clean reconstruction results of desired objects due to distracting objects or backgrounds. In addition, any changes in background during scanning, which can often occur in real scenarios, can easily break up the whole reconstruction process. To address these issues, we incorporate visual saliency into a traditional real‐time volumetric fusion pipeline. Salient regions detected from RGB‐D frames suggest user‐intended objects, and by understanding user intentions our approach can put more emphasis on important targets, and meanwhile, eliminate disturbance of non‐important objects. Experimental results on real‐world scans demonstrate that our system is capable of effectively acquiring geometric information of salient objects in cluttered real‐world scenes, even if the backgrounds are changing. Sheng Yang 0007, Minghua Liu, Hongbo Fu 0001, Shi-Min Hu 0001 |
Comput. Graph. Forum | 5 |
| 2017 | Message from the Editor-in-ChiefabstractI would like to take this opportunity to thank everyone who has helped to make Computational Visual Media a success in its second year of 2016. In particular, my thanks go to the authors, the reviewers, and the Editorial Board members, as well as the staff of Tsinghua University Press and Springer. Your combined efforts have helped to ensure that all four issues for 2016 were published on schedule, before the end of the year. 31 papers were published in 4 issues in 2016, including regular papers and papers recommended to us by the CVM conference and Pacific Graphics. The acceptance rate for regular papers was 37.5%. Shi-Min Hu 0001 |
Comput. Vis. Media | 1 |
| 2017 | Preface
Shi-Min Hu 0001, Niloy J. Mitra, Yizhou Yu |
J. Comput. Sci. Technol. | 1 |
| 2017 | An Optimization Approach for Localization Refinement of Candidate Traffic SignsabstractWe propose a localization refinement approach for candidate traffic signs. Previous traffic sign localization approaches, which place a bounding rectangle around the sign, do not always give a compact bounding box, making the subsequent classification task more difficult. We formulate localization as a segmentation problem, and incorporate prior knowledge concerning color and shape of traffic signs. To evaluate the effectiveness of our approach, we use it as an intermediate step between a standard traffic sign localizer and a classifier. Our experiments use the well-known German Traffic Sign Detection Benchmark (GTSDB) as well as our new Chinese Traffic Sign Detection Benchmark. This newly created benchmark is publicly available,1and goes beyond previous benchmark data sets: it has over 5000 high-resolution images containing more than 14 000 traffic signs taken in realistic driving conditions. Experimental results show that our localization approach significantly improves bounding boxes when compared with a standard localizer, thereby allowing a standard traffic sign classifier to generate more accurate classification results.1http://cg.cs.tsinghua.edu.cn/ctsdb/. Zhe Zhu, Jiaming Lu, Ralph R. Martin, Shi-Min Hu 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2017 | A unified particle system framework for multi-phase, multi-material visual simulationsabstractWe introduce a unified particle framework which integrates the phase-field method with multi-material simulation to allow modeling of both liquids and solids, as well as phase transitions between them. A simple elasto-plastic model is used to capture the behavior of various kinds of solids, including deformable bodies, granular materials, and cohesive soils. States of matter or phases , particularly liquids and solids, are modeled using the non-conservative Allen-Cahn equation. In contrast, materials---made of different substances---are advected by the conservative Cahn-Hilliard equation. The distributions of phases and materials are represented by a phase variable and a concentration variable, respectively, allowing us to represent commonly observed fluid-solid interactions. Our multi-phase, multi-material system is governed by a unified Helmholtz free energy density. This framework provides the first method in computer graphics capable of modeling a continuous interface between phases. It is versatile and can be readily used in many scenarios that are challenging to simulate. Examples are provided to demonstrate the capabilities and effectiveness of this approach. Jian Chang 0001, Ming C. Lin, Ralph R. Martin, Jian J. Zhang 0001, Shi-Min Hu 0001 |
ACM Trans. Graph. | 6 |
| 2017 | Pairwise Force SPH Model for Real-Time Multi-Interaction ApplicationsabstractIn this paper, we present a novel pairwise-force smoothed particle hydrodynamics (PF-SPH) model to enable simulation of various interactions at interfaces in real time. Realistic capture of interactions at interfaces is a challenging problem for SPH-based simulations, especially for scenarios involving multiple interactions at different interfaces. Our PF-SPH model can readily handle multiple types of interactions simultaneously in a single simulation; its basis is to use a larger support radius than that used in standard SPH. We adopt a novel anisotropic filtering term to further improve the performance of interaction forces. The proposed model is stable; furthermore, it avoids the particle clustering problem which commonly occurs at the free surface. We show how our model can be used to capture various interactions. We also consider the close connection between droplets and bubbles, and show how to animate bubbles rising in liquid as well as bubbles in air. Our method is versatile, physically plausible and easy-to-implement. Examples are provided to demonstrate the capabilities and effectiveness of our approach. Ralph R. Martin, Ming C. Lin, Jian Chang 0001, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2017 | PlenoPatch: Patch-Based Plenoptic Image ManipulationabstractPatch-based image synthesis methods have been successfully applied for various editing tasks on still images, videos and stereo pairs. In this work we extend patch-based synthesis to plenoptic images captured by consumer-level lenselet-based devices for interactive, efficient light field editing. In our method the light field is represented as a set of images captured from different viewpoints. We decompose the central view into different depth layers, and present it to the user for specifying the editing goals. Given an editing task, our method performs patch-based image synthesis on all affected layers of the central view, and then propagates the edits to all other views. Interaction is done through a conventional 2D image editing user interface that is familiar to novice users. Our method correctly handles object boundary occlusion with semi-transparency, thus can generate more realistic results than previous methods. We demonstrate compelling results on a wide range of applications such as hole-filling, object reshuffling and resizing, changing object depth, light field upscaling and parallax magnification. Jue Wang 0001, Eli Shechtman, Zi-Ye Zhou, Jiaxin Shi, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2017 | High-speed video generation with an event camera
David Marshall 0001, Luping Shi, Shi-Min Hu 0001 |
Vis. Comput. | 5 |
| 2016 | Traffic-Sign Detection and Classification in the WildabstractAlthough promising results have been achieved in the areas of traffic-sign detection and classification, few works have provided simultaneous solutions to these two tasks for realistic real world images. We make two contributions to this problem. Firstly, we have created a large traffic-sign benchmark from 100000 Tencent Street View panoramas, going beyond previous benchmarks. It provides 100000 images containing 30000 traffic-sign instances. These images cover large variations in illuminance and weather conditions. Each traffic-sign in the benchmark is annotated with a class label, its bounding box and pixel mask. We call this benchmark Tsinghua-Tencent 100K. Secondly, we demonstrate how a robust end-to-end convolutional neural network (CNN) can simultaneously detect and classify trafficsigns. Most previous CNN image processing solutions target objects that occupy a large proportion of an image, and such networks do not work well for target objects occupying only a small fraction of an image like the traffic-signs here. Experimental results show the robustness of our network and its superiority to alternatives. The benchmark, source code and the CNN model introduced in this paper is publicly available1. Zhe Zhu, Dun Liang, Song-Hai Zhang, Sharon X. Huang, Baoli Li 0004, Shi-Min Hu 0001 |
CVPR | 6 |
| 2016 | HFS: Hierarchical Feature Selection for Efficient Image Segmentation
Ming-Ming Cheng, Yun Liu 0011, Qibin Hou, Jiawang Bian, Philip Torr 0001, Shi-Min Hu 0001, Zhuowen Tu |
ECCV (3) | 6 |
| 2016 | Staying Secure and Unprepared: Understanding and Mitigating the Security Risks of Apple ZeroConfabstractWith the popularity of today's usability-oriented designs, dubbed Zero Configuration or ZeroConf, unclear are the security implications of these automatic service discovery, "plug-and-play" techniques. In this paper, we report the first systematic study on this issue, focusing on the security features of the systems related to Apple, the major proponent of ZeroConf techniques. Our research brings to light a disturbing lack of security consideration in these systems' designs: major ZeroConf frameworks on the Apple platforms, including the Core Bluetooth Framework, Multipeer Connectivity and Bonjour, are mostly unprotected and popular apps and system services, such as Tencent QQ, Apple Handoff, printer discovery and AirDrop, turn out to be completely vulnerable to an impersonation or Man-in-the-Middle (MitM) attack, even though attempts have been made to protect them against such threats. The consequences are serious, allowing a malicious device to steal the user's SMS messages, email notifications, documents to be printed out or transferred to another device. Most importantly, our study highlights the fundamental security challenges underlying ZeroConf techniques: in the absence of any pre-configured secret across different devices, authentication has to rely on Apple-issued public-key certificate, which however cannot be properly verified due to the difficulty in finding a unique, nonsensitive and widely known identity of a human user to bind her to her certificate. To address this issue, we developed a suite of new techniques, including a conflict detection approach and a biometric technique that enables the user to speak out her certificate through 6 distinct, rare but pronounceable words to let those who know her voice verify her certificate. We performed a security analysis on the new protection and evaluated its usability and effectiveness using two user studies involving 60 participants. Our research shows that the new protection fits well with the existing ZeroConf systems such as AirDrop. It is well received by users and also providing effective defense even against recently proposed speech synthesis attacks. Xiaolong Bai, Luyi Xing, Nan Zhang 0018, XiaoFeng Wang 0001, Xiaojing Liao, Tongxin Li 0002, Shi-Min Hu 0001 |
IEEE Symposium on Security and Privacy | 7 |
| 2016 | Testing Error Handling Code in Device Drivers Using Characteristic Fault Injection
Jia-Ju Bai, Yu-Ping Wang 0001, Shi-Min Hu 0001 |
USENIX ATC | 4 |
| 2016 | Appearance Harmonization for Single Image Shadow RemovalabstractAbstract Shadow removal is a challenging problem and previous approaches often produce de‐shadowed regions that are visually inconsistent with the rest of the image. We propose an automaticshadow region harmonizationapproach that makes the appearance of a de‐shadowed region (produced using any previous technique) compatible with the rest of the image. We use a shadow‐guided patch‐based image synthesis approach that reconstructs the shadow region using patches sampled from non‐shadowed regions. This result is then refined based on the reconstruction confidence to handle unique textures. Qualitative comparisons over a wide range of images, and a quantitative evaluation on a benchmark dataset show that our technique significantly improves upon the state‐of‐the‐art. Li-Qian Ma, Jue Wang 0001, Eli Shechtman, Kalyan Sunkavalli, Shi-Min Hu 0001 |
Comput. Graph. Forum | 5 |
| 2016 | Structure guided interior scene synthesis via graph matching
Shi-Sheng Huang, Hongbo Fu 0001, Shi-Min Hu 0001 |
Graph. Model. | 3 |
| 2016 | Message from the Editor-in-Chief
Shi-Min Hu 0001 |
Comput. Vis. Media | 1 |
| 2016 | Mining and checking paired functions in device drivers using characteristic fault injection
Jia-Ju Bai, Yu-Ping Wang 0001, Hu-Qiu Liu, Shi-Min Hu 0001 |
Inf. Softw. Technol. | 4 |
| 2016 | Preface
Shi-Min Hu 0001, Ligang Liu 0001, Ralph R. Martin |
J. Comput. Sci. Technol. | 1 |
| 2016 | PF-Miner: A practical paired functions mining method for Android kernel in error paths
Hu-Qiu Liu, Yu-Ping Wang 0001, Jia-Ju Bai, Shi-Min Hu 0001 |
J. Syst. Softw. | 4 |
| 2016 | PAST: accurate instrumentation on fully optimized programabstractInstrumentation is a powerful technique for monitoring, profiling, debugging, logging and tracing the software. In order to determine the instrumentation location, the user needs to know where the current executed location is in the source code. Previous instrumentation approaches rely on debugging information to find the location in the source code. For fully optimized programs, debugging information is not complete, which limits the application of those approaches. In this paper, we present pattern-based abstract syntax tree (PAST) instrumentation, an ideal instrumentation methodology that accurately instruments the fully optimized program. The instrumentation location is specified in an intuitive way that matches the source code at the abstract syntax tree level. The program can be instrumented either at the compile time using the ordinary compiling or when it is running using the just-in-time compiling. Experimental results show that PAST can accurately instrument the target program. There is negligible run time overhead when the running program is instrumented without any operation. We have implemented PAST on both x86-32 and x86-64 to show that PAST is easily portable across different architecture. Copyright © 2015 John Wiley & Sons, Ltd. Shi-Min Hu 0001 |
Softw. Pract. Exp. | 3 |
| 2016 | Efficient, Edge-Aware, Combined Color Quantization and DitheringabstractIn this paper, we present a novel algorithm to simultaneously accomplish color quantization and dithering of images. This is achieved by minimizing a perception-based cost function, which considers pixel-wise differences between filtered versions of the quantized image and the input image. We use edge aware filters in defining the cost function to avoid mixing colors on the opposite sides of an edge. The importance of each pixel is weighted according to its saliency. To rapidly minimize the cost function, we use a modified multi-scale iterative conditional mode (ICM) algorithm, which updates one pixel a time while keeping other pixels unchanged. As ICM is a local method, careful initialization is required to prevent termination at a local minimum far from the global one. To address this problem, we initialize ICM with a palette generated by a modified median-cut method. Compared with previous approaches, our method can produce high-quality results with a fewer visual artifacts but also requires significantly less computational effort. Hao-Zhi Huang 0001, Kun Xu 0003, Ralph R. Martin, Fei-Yue Huang, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 5 |
| 2016 | Multiphase SPH simulation for interactive fluids and solidsabstractThis work extends existing multiphase-fluid SPH frameworks to cover solid phases, including deformable bodies and granular materials. In our extended multiphase SPH framework, the distribution and shapes of all phases, both fluids and solids, are uniformly represented by their volume fraction functions. The dynamics of the multiphase system is governed by conservation of mass and momentum within different phases. The behavior of individual phases and the interactions between them are represented by corresponding constitutive laws, which are functions of the volume fraction fields and the velocity fields. Our generalized multiphase SPH framework does not require separate equations for specific phases or tedious interface tracking. As the distribution, shape and motion of each phase is represented and resolved in the same way, the proposed approach is robust, efficient and easy to implement. Various simulation results are presented to demonstrate the capabilities of our new multiphase SPH framework, including deformable bodies, granular materials, interaction between multiple fluids and deformable solids, flow in porous media, and dissolution of deformable solids. Xiao Yan 0004, Yun-Tao Jiang, Chenfeng Li, Ralph R. Martin, Shi-Min Hu 0001 |
ACM Trans. Graph. | 5 |
| 2016 | Robust background identification for dynamic video editingabstractExtracting background features for estimating the camera path is a key step in many video editing and enhancement applications. Existing approaches often fail on highly dynamic videos that are shot by moving cameras and contain severe foreground occlusion. Based on existing theories, we present a new, practical method that can reliably identify background features in complex video, leading to accurate camera path estimation and background layering. Our approach contains a local motion analysis step and a global optimization step. We first divide the input video into overlapping temporal windows, and extract local motion clusters in each window. We form a directed graph from these local clusters, and identify background ones by finding a minimal path through the graph using optimization. We show that our method significantly outperforms other alternatives, and can be directly used to improve common video editing applications such as stabilization, compositing and background reconstruction. Xian Wu 0004, Hao-Tian Zhang, Jue Wang 0001, Shi-Min Hu 0001 |
ACM Trans. Graph. | 5 |
| 2016 | Support Substructures: Support-Induced Part-Level Structural RepresentationabstractIn this work we explore a support-induced structural organization of object parts. We introduce the concept of support substructures, which are special subsets of object parts with support and stability. A bottom-up approach is proposed to identify such substructures in a support relation graph. We apply the derived high-level substructures to part-based shape reshuffling between models, resulting in nontrivial functionally plausible model variations that are difficult to achieve with symmetry-induced substructures by the state-of-the-art methods. We also show how to automatically or interactively turn a single input model to new functionally plausible shapes by structure rearrangement and synthesis, enabled by support substructures. To the best of our knowledge no single existing method has been designed for all these applications. Shi-Sheng Huang, Hongbo Fu 0001, Ling-Yu Wei, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2016 | Faithful Completion of Images of Scenic Landmarks Using Internet ImagesabstractPrevious works on image completion typically aim to produce visually plausible results rather than factually correct ones. In this paper, we propose an approach to faithfully complete the missing regions of an image. We assume that the input image is taken at a well-known landmark, so similar images taken at the same location can be easily found on the Internet. We first download thousands of images from the Internet using a text label provided by the user. Next, we apply two-step filtering to reduce them to a small set of candidate images for use as source images for completion. For each candidate image, a co-matching algorithm is used to find correspondences of both points and lines between the candidate image and the input image. These are used to find an optimal warp relating the two images. A completion result is obtained by blending the warped candidate image into the missing region of the input image. The completion results are ranked according to combination score, which considers both warping and blending energy, and the highest ranked ones are shown to the user. Experiments and results demonstrate that our method can faithfully complete images. Zhe Zhu, Hao-Zhi Huang 0001, Kun Xu 0003, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2016 | Fast SPH simulation for gaseous fluids
Bo Ren 0003, Xiao Yan 0004, Chenfeng Li, Ming C. Lin, Shi-Min Hu 0001 |
Vis. Comput. | 6 |
| 2015 | Cracking App Isolation on Apple: Unauthorized Cross-App Resource Access on MAC OS~X and iOSabstractOn modern operating systems, applications under the same user are separated from each other, for the purpose of protecting them against malware and compromised programs. Given the complexity of today's OSes, less clear is whether such isolation is effective against different kind of cross-app resource access attacks (called XARA in our research). To better understand the problem, on the less-studied Apple platforms, we conducted a systematic security analysis on MAC OS~X and iOS. Our research leads to the discovery of a series of high-impact security weaknesses, which enable a sandboxed malicious app, approved by the Apple Stores, to gain unauthorized access to other apps' sensitive data. More specifically, we found that the inter-app interaction services, including the keychain, WebSocket and NSConnection on OS~X and URL Scheme on the MAC OS and iOS, can all be exploited by the malware to steal such confidential information as the passwords for iCloud, email and bank, and the secret token of Evernote. Further, the design of the app sandbox on OS~X was found to be vulnerable, exposing an app's private directory to the sandboxed malware that hijacks its Apple Bundle ID. As a result, sensitive user data, like the notes and user contacts under Evernote and photos under WeChat, have all been disclosed. Fundamentally, these problems are caused by the lack of app-to-app and app-to-OS authentications. To better understand their impacts, we developed a scanner that automatically analyzes the binaries of MAC OS and iOS apps to determine whether proper protection is missing in their code. Running it on hundreds of binaries, we confirmed the pervasiveness of the weaknesses among high-impact Apple apps. Since the issues may not be easily fixed, we built a simple program that detects exploit attempts on OS~X, helping protect vulnerable apps before the problems can be fully addressed. Luyi Xing, Xiaolong Bai, Tongxin Li 0002, XiaoFeng Wang 0001, Kai Chen 0012, Xiaojing Liao, Shi-Min Hu 0001, Xinhui Han |
CCS | 7 |
| 2015 | Complete Runtime Tracing for Device Drivers Based on LLVMabstractDevice drivers often suffer from much more bugs than the kernel, so testing device drivers becomes more and more important and necessary. In software testing, runtime tracing is an important technique to monitor real executing procedures of the program. Meanwhile, runtime information can also assist the programmer to make more accurate analysis of the program, like verifying the correctness of code execution and detecting bugs. However, due to kernel-mode execution and high complexity of kernel code, completely tracing drivers is hard, which causes real execution paths can not be clearly identified. In order to provide more powerful support for software testing of device drivers, we propose a method named Driver Trace, to do complete runtime tracing at the function level. Driver Trace utilizes instrumentation technique for runtime tracing, which is implemented based on LLVM compiler infrastructure. When the target driver works, Driver Trace records complete runtime information of function calls, like function names, return values and parameter pointers, and the information is recorded in a log file for future analysis. We have successfully implemented Driver Trace on 10 real device drivers in Linux 3.16.4 and made the evaluation as well. The experimental results show that Driver Trace provides an effective method of runtime tracing for device drivers with the modest overhead. Moreover, using an automated analysis of the runtime information recorded by Driver Trace, we also find 6 violations about resource usages in these 10 device drivers. Jia-Ju Bai, Hu-Qiu Liu, Yu-Ping Wang 0001, Shi-Min Hu 0001 |
COMPSAC | 4 |
| 2015 | Pairminer: mining for paired functions in Kernel extensionsabstractDrivers use kernel extension functions to manage devices, and there are often many rules on how they should be used. Among the rules, utilization of paired functions, which means that the functions must be called in pairs between two different functions, is extremely complex and important. However, such pairing rules are not well documented, and these rules can be easily violated by programmers when they unconsciously ignore or forget about them. Therefore it is useful to develop a tool to automatically extract paired functions in the kernel source and detect incorrect usages. We put forward a method called PairMiner in this paper. Heuristic and statistical mechanisms are adopted to associate with the special structure of drivers’ source code, to find out paired functions between relative operations, and then to detect violations with extracted paired functions. In the experiment evaluation, we have successfully found 1023 paired functions in Linux 3.10.10. The utility of PairMiner was evaluated by analyzing the source code of Linux 2.6.38 and 3.10.10. PairMiner located 265 bugs about paired function violations in 2.6.38 which have been fixed in 3.10.10. We also have identified 1994 paired function violations which have not yet been fixed in 3.10.10. We have reported some violations as potential bugs with emails to the developers, 27 developers have replied the emails and 20 bugs have been confirmed so far, 2 violations are confirmed as false positive. Hu-Qiu Liu, Jia-Ju Bai, Yu-Ping Wang 0001, Zhe Bian, Shi-Min Hu 0001 |
ISPASS | 5 |
| 2015 | Automated resource release in device driversabstractDevice drivers require system resources to control hardware and provide fundamental services for applications. The acquired resources must be explicitly released by drivers. Otherwise, these resources will never be reclaimed by the operating system, and they are not available for other programs any more, causing hard-to-find system problems. We study on Linux driver mailing lists, and find many applied patches handle improper resource-release operations, especially in error handling paths. In order to improve current resource management and avoid resource-release omissions in device drivers, we propose a novel approach named AutoRR, which can automatically and safely release resources based on specification-mining techniques. During execution, we maintain a resource-state table by recording the runtime information of function calls. If the driver fails to release acquired resources during execution, AutoRR will report violations and call corresponding releasing functions with the recorded runtime information to release acquired resources. To fully and safely release acquired resources, a dynamic analysis of resource dependency and allocation hierarchy is also performed, which can avoid dead resources and double frees. AutoRR works in both normal execution and error handling paths for reliable resource management. We implement AutoRR with LLVM, and evaluate it on 8 Ethernet drivers in Linux 3.17.2. The evaluation shows that the overhead of AutoRR is very low, and it has successfully fixed 18 detected resource-release omission violations without side effects. Our work shows a feasible way of using specification-mining results to avoid related violations. Jia-Ju Bai, Yu-Ping Wang 0001, Hu-Qiu Liu, Shi-Min Hu 0001 |
ISSRE | 4 |
| 2015 | WebC: toward a portable framework for deploying legacy code in web browsers
Gang Tan, Xiaolong Bai, Shi-Min Hu 0001 |
Sci. China Inf. Sci. | 4 |
| 2015 | 3D indoor scene modeling from RGB-D data: a surveyabstract3D scene modeling has long been a fundamental problem in computer graphics and computer vision. With the popularity of consumer-level RGB-D cameras, there is a growing interest in digitizing real-world indoor 3D scenes. However, modeling indoor 3D scenes remains a challenging problem because of the complex structure of interior objects and poor quality of RGB-D data acquired by consumer-level sensors. Various methods have been proposed to tackle these challenges. In this survey, we provide an overview of recent advances in indoor scene modeling techniques, as well as public datasets and code libraries which can facilitate experiments and evaluation. Yukun Lai, Shi-Min Hu 0001 |
Comput. Vis. Media | 3 |
| 2015 | Message from the Editor-in-Chief
Shi-Min Hu 0001 |
Comput. Vis. Media | 1 |
| 2015 | Panorama completion for street viewsabstractThis paper considers panorama images used for street views. Their viewing angle of 360° causes pixels at the top and bottom to appear stretched and warped. Although current image completion algorithms work well, they cannot be directly used in the presence of such distortions found in panoramas of street views. We thus propose a novel approach to complete such 360° panoramas using optimization-based projection to deal with distortions. Experimental results show that our approach is efficient and provides an improvement over standard image completion algorithms. Zhe Zhu, Ralph R. Martin, Shi-Min Hu 0001 |
Comput. Vis. Media | 3 |
| 2015 | Preface
Shi-Min Hu 0001, Leif Kobbelt |
J. Comput. Sci. Technol. | 1 |
| 2015 | Global Contrast Based Salient Region DetectionabstractAutomatic estimation of salient object regions across images, without any prior assumption or knowledge of the contents of the corresponding scenes, enhances many computer vision and computer graphics applications. We introduce a regional contrast based salient object detection algorithm, which simultaneously evaluates global contrast differences and spatial weighted coherence scores. The proposed algorithm is simple, efficient, naturally multi-scale, and produces full-resolution, high-quality saliency maps. These saliency maps are further used to initialize a novel iterative version of GrabCut, namely SaliencyCut, for high quality unsupervised salient object segmentation. We extensively evaluated our algorithm using traditional salient object detection datasets, as well as a more challenging Internet image dataset. Our experimental results demonstrate that our algorithm consistently outperforms 15 existing salient object detection and segmentation methods, yielding higher precision and better recall rates. We also show that our algorithm can be used to efficiently extract salient object masks from Internet images, enabling effective sketch-based image retrieval (SBIR) via simple shape comparisons. Despite such noisy internet images, where the saliency regions are ambiguous, our saliency guided image retrieval achieves a superior retrieval rate compared with state-of-the-art SBIR methods, and additionally provides important target object region information. Ming-Ming Cheng, Niloy J. Mitra, Sharon X. Huang, Philip Torr 0001, Shi-Min Hu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2015 | Simultaneous Camera Path Optimization and Distraction Removal for Improving Amateur VideoabstractA major difference between amateur and professional video lies in the quality of camera paths. Previous work on video stabilization has considered how to improve amateur video by smoothing the camera path. In this paper, we show that additional changes to the camera path can further improve video aesthetics. Our new optimization method achieves multiple simultaneous goals: 1) stabilizing video content over short time scales; 2) ensuring simple and consistent camera paths over longer time scales; and 3) improving scene composition by automatically removing distractions, a common occurrence in amateur video. Our approach uses an L(1) camera path optimization framework, extended to handle multiple constraints. Two passes of optimization are used to address both low-level and high-level constraints on the camera path. The experimental and user study results show that our approach outputs video that is perceptually better than the input, or the results of using stabilization only. Jue Wang 0001, Ralph R. Martin, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 5 |
| 2015 | PatchTable: efficient patch queries for large datasets and applicationsabstractThis paper presents a data structure that reduces approximate nearest neighbor query times for image patches in large datasets. Previous work in texture synthesis has demonstrated real-time synthesis from small exemplar textures. However, high performance has proved elusive for modern patch-based optimization techniques which frequently use many exemplar images in the tens of megapixels or above. Our new algorithm, PatchTable, offloads as much of the computation as possible to a pre-computation stage that takes modest time, so patch queries can be as efficient as possible. There are three key insights behind our algorithm: (1) a lookup table similar to locality sensitive hashing can be precomputed, and used to seed sufficiently good initial patch correspondences during querying, (2) missing entries in the table can be filled during pre-computation with our fast Voronoi transform, and (3) the initially seeded correspondences can be improved with a precomputed k-nearest neighbors mapping. We show experimentally that this accelerates the patch query operation by up to 9× over k-coherence, up to 12× over TreeCANN, and up to 200× over PatchMatch. Our fast algorithm allows us to explore efficient and practical imaging and computational photography applications. We show results for artistic video stylization, light field super-resolution, and multi-image editing. Connelly Barnes, Li-ming Lou, Xian Wu 0004, Shi-Min Hu 0001 |
ACM Trans. Graph. | 5 |
| 2015 | Magic decorator: automatic material suggestion for indoor digital scenesabstractAssigning textures and materials within 3D scenes is a tedious and labor-intensive task. In this paper, we present Magic Decorator , a system that automatically generates material suggestions for 3D indoor scenes. To achieve this goal, we introduce local material rules , which describe typical material patterns for a small group of objects or parts, and global aesthetic rules , which account for the harmony among the entire set of colors in a specific scene. Both rules are obtained from collections of indoor scene images. We cast the problem of material suggestion as a combinatorial optimization considering both local material and global aesthetic rules. We have tested our system on various complex indoor scenes. A user study indicates that our system can automatically and efficiently produce a series of visually plausible material suggestions which are comparable to those produced by artists. Kun Xu 0003, Yizhou Yu, Tian-Yi Wang, Shi-Min Hu 0001 |
ACM Trans. Graph. | 5 |
| 2015 | Fast multiple-fluid simulation using Helmholtz free energyabstractMultiple-fluid interaction is an interesting and common visual phenomenon we often observe. In this paper, we present an energy-based Lagrangian method that expands the capability of existing multiple-fluid methods to handle various phenomena, such as extraction, partial dissolution, etc. Based on our user-adjusted Helmholtz free energy functions, the simulated fluid evolves from high-energy states to low-energy states, allowing flexible capture of various mixing and unmixing processes. We also extend the original Cahn-Hilliard equation to be better able to simulate complex fluid-fluid interaction and rich visual phenomena such as motion-related mixing and position based pattern. Our approach is easily integrated with existing state-of-the-art smooth particle hydrodynamic (SPH) solvers and can be further implemented on top of the position based dynamics (PBD) method, improving the stability and incompressibility of the fluid during Lagrangian simulation under large time steps. Performance analysis shows that our method is at least 4 times faster than the state-of-the-art multiple-fluid method. Examples are provided to demonstrate the new capability and effectiveness of our approach. Jian Chang 0001, Bo Ren 0003, Ming C. Lin, Jian J. Zhang 0001, Shi-Min Hu 0001 |
ACM Trans. Graph. | 6 |
| 2015 | Active Exploration of Large 3D Model RepositoriesabstractWith broader availability of large-scale 3D model repositories, the need for efficient and effective exploration becomes more and more urgent. Existing model retrieval techniques do not scale well with the size of the database since often a large number of very similar objects are returned for a query, and the possibilities to refine the search are quite limited. We propose an interactive approach where the user feeds an active learning procedure by labeling either entire models or parts of them as "like" or "dislike" such that the system can automatically update an active set of recommended models. To provide an intuitive user interface, candidate models are presented based on their estimated relevance for the current query. From the methodological point of view, our main contribution is to exploit not only the similarity between a query and the database models but also the similarities among the database models themselves. We achieve this by an offline pre-processing stage, where global and local shape descriptors are computed for each model and a sparse distance metric is derived that can be evaluated efficiently even for very large databases. We demonstrate the effectiveness of our method by interactively exploring a repository containing over 100 K models. Lin Gao 0004, Yan-Pei Cao 0001, Yukun Lai, Hao-Zhi Huang 0001, Leif Kobbelt, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2015 | A response time model for abrupt changes in binocular disparity
Tai-Jiang Mu, Jia-Jia Sun, Ralph R. Martin, Shi-Min Hu 0001 |
Vis. Comput. | 4 |
| 2014 | Runtime Checking for Paired Functions in Device DriversabstractDevice drivers usually invoke functions to allocate resources for managing hardware devices and communicating with the kernel, and these resources should be released by functions when the work is finished. Thus allocating functions and releasing functions must be invoked in pairs. However, many developers ignore this vital rule, and some allocated resources are not released in time, which may cause resource related problems like deadlocks and memory leak. For improving the resource management of device drivers, we propose an approach named Pair Dyn to check these paired functions during runtime. When the driver runs, Pair Dyn records the runtime information of allocating functions such as key parameters and return value, and dynamically detects whether the relevant releasing functions are invoked to free allocated resources during runtime. Before the driver exits, Pair Dyn automatically attempts to invoke the related releasing functions which are lacked in runtime, in order to free the allocated resources of the operation system. We have implemented Pair Dyn with the LLVM compiler infrastructure, and make the evaluation with four real device drivers in Linux version 3.10.1. The experimental result shows that with the low extra overhead, Pair Dyn can provide effective runtime checking for allocate-release paired functions. Moreover, 9 potential bugs are found in the four drivers, which are all fixed automatically before exiting. Finally, no manual modification of the source code is needed with Pair Dyn. Jia-Ju Bai, Hu-Qiu Liu, Yu-Ping Wang 0001, Shi-Min Hu 0001 |
APSEC (1) | 4 |
| 2014 | BP-Miner: Mining Paired Functions from the Binary Code of Drivers for Error HandlingabstractKernel extension functions are provided as interfaces for drivers to manage devices and resources, and there are many implicit rules about their usages. One of the most important rules is that many functions should be called in pairs. That is to say, when an error occurs in a function, the driver should call related functions to handle it and release the acquired resources before returning, and we name these functions between normal execution paths and error handling paths as paired functions. However, many developers are unaware of them, which causes lots of bugs. Therefore, it is highly significant to automatically extract paired functions and detect violations for drivers. This paper proposes an efficient tool named BP-Miner, which can extract paired functions from binary code of driver modules and detect violations for error handling in drivers with extracted paired functions. BP-Miner constructs control flow graph (CFG) based on basic blocks of binary code, and locates potential execution paths to extract paired functions. We have evaluated BP-Miner with Linux drivers 2.6.38 and 3.13.0-rc7. 76 bugs are reported by BP-Miner in 2.6.38 which have been fixed in the current latest version 3.13.0-rc7. BP-Miner spends about 90 minutes handling 3653 module files for 3.13.0-rc7, and 859 violations have been detected with 1167 extracted paired functions. As it works on the binary code, it can be utilized to check close-source drivers. Hu-Qiu Liu, Jia-Ju Bai, Yu-Ping Wang 0001, Shi-Min Hu 0001 |
APSEC (1) | 4 |
| 2014 | PF-Miner: A New Paired Functions Mining Method for Android Kernel in Error PathsabstractDrivers are significant components of the operating systems(OSs), and they run in kernel mode. Generally, drivers have many errors to handle, and the functions called in the normal execution paths and error handling paths are in pairs, which are named as paired functions. However, some developers do not handle the errors completely as they forget about or are unaware of releasing the acquired resources, thus memory leaks and other potential problems can be easily introduced. Therefore, it is highly valuable to automatically extract paired functions for these problems and detect violations for the programmers. This paper proposes an efficient tool named PF-Miner, which can automatically extract paired functions and detect violations between normal execution paths and error handling paths from the source code of C program with the data mining and statistical methods. We have evaluated PF-Miner on different versions of Android kernel 2.6.39 and 3.10.0, and 81 bugs reported by PF-Miner in 2.6.39 have been fixed before the latest version 3.10.0. PF-Miner only needs about 150 seconds to analyze the source code of 3.10.0, and 983 violations have been detected from 546 paired functions that have been extracted. We have reported the top 51 violations as potential bugs to the developers, and 15 bugs have been confirmed. Hu-Qiu Liu, Yu-Ping Wang 0001, Lingbo Jiang, Shi-Min Hu 0001 |
COMPSAC | 4 |
| 2014 | Interactive Image-Guided Modeling of Extruded ShapesabstractAbstract A recent trend in interactive modeling of 3D shapes from a single image is designing minimal interfaces, and accompanying algorithms, for modeling a specific class of objects. Expanding upon the range of shapes that existing minimal interfaces can model, we present an interactive image‐guided tool for modeling shapes made up of extruded parts. An extruded part is represented by extruding a closed planar curve, called base, in the direction orthogonal to the base. To model each extruded part, the user only needs to sketch the projected base shape in the image. The main technical contribution is a novel optimization‐based approach for recovering the 3D normal of the base of an extruded object by exploring both geometric regularity of the sketched curve and image contents. We developed a convenient interface for modeling multi‐part shapes and a method for optimizing the relative placement of the parts. Our tool is validated using synthetic data and tested on real‐world images. Yan-Pei Cao 0001, Zhao Fu, Shi-Min Hu 0001 |
Comput. Graph. Forum | 4 |
| 2014 | Learning Natural Colors for Image RecoloringabstractAbstract We present a data‐driven method for automatically recoloring a photo to enhance its appearance or change a viewer's emotional response to it. A compact representation called a RegionNet summarizes color and geometric features of image regions, and geometric relationships between them. Correlations between color property distributions and geometric features of regions are learned from a database of well‐colored photos. A probabilistic factor graph model is used to summarize distributions of color properties and generate an overall probability distribution for color suggestions. Given a new input image, we can generate multiple recolored results which unlike previous automatic results, are both natural and artistic, and compatible with their spatial arrangements. Hao-Zhi Huang 0001, Song-Hai Zhang, Ralph R. Martin, Shi-Min Hu 0001 |
Comput. Graph. Forum | 4 |
| 2014 | Structure Aware Visual CryptographyabstractAbstract Visual cryptography is an encryption technique that hides a secret image by distributing it between some shared images made up of seemingly random black‐and‐white pixels. Extended visual cryptography (EVC) goes further in that the shared images instead represent meaningful binary pictures. The original approach to EVC suffered from low contrast, so later papers considered how to improve the visual quality of the results by enhancing contrast of the shared images. This work further improves the appearance of the shared images by preserving edge structures within them using a framework of dithering followed by a detail recovery operation. We are also careful to suppress noise in smooth areas. Ralph R. Martin, Jiwu Huang, Shi-Min Hu 0001 |
Comput. Graph. Forum | 4 |
| 2014 | Automatic semantic modeling of indoor scenes from low-quality RGB-D data using contextual informationabstractWe present a novel solution to automatic semantic modeling of indoor scenes from a sparse set of low-quality RGB-D images. Such data presents challenges due to noise, low resolution, occlusion and missing depth information. We exploit the knowledge in a scene database containing 100s of indoor scenes with over 10,000 manually segmented and labeled mesh models of objects. In seconds, we output a visually plausible 3D scene, adapting these models and their parts to fit the input scans. Contextual relationships learned from the database are used to constrain reconstruction, ensuring semantic compatibility between both object models and parts. Small objects and objects with incomplete depth information which are difficult to recover reliably are processed with a two-stage approach. Major objects are recognized first, providing a known scene structure. 2D contour-based model retrieval is then used to recover smaller objects. Evaluations using our own data and two public datasets show that our approach can model typical real-world indoor scenes efficiently and robustly. Yukun Lai, Ralph R. Martin, Shi-Min Hu 0001 |
ACM Trans. Graph. | 5 |
| 2014 | Multiple-Fluid SPH Simulation Using a Mixture ModelabstractThis article presents a versatile and robust SPH simulation approach for multiple-fluid flows. The spatial distribution of different phases or components is modeled using the volume fraction representation, the dynamics of multiple-fluid flows is captured by using an improved mixture model, and a stable and accurate SPH formulation is rigorously derived to resolve the complex transport and transformation processes encountered in multiple-fluid flows. The new approach can capture a wide range of real-world multiple-fluid phenomena, including mixing/unmixing of miscible and immiscible fluids, diffusion effect and chemical reaction, etc. Moreover, the new multiple-fluid SPH scheme can be readily integrated into existing state-of-the-art SPH simulators, and the multiple-fluid simulation is easy to set up. Various examples are presented to demonstrate the effectiveness of our approach. Bo Ren 0003, Chenfeng Li, Xiao Yan 0004, Ming C. Lin, Javier Bonet, Shi-Min Hu 0001 |
ACM Trans. Graph. | 6 |
| 2014 | BiggerPicture: data-driven image extrapolation using graph matchingabstractFilling a small hole in an image with plausible content is well studied. Extrapolating an image to give a distinctly larger one is much more challenging---a significant amount of additional content is needed which matches the original image, especially near its boundaries. We propose a data-driven approach to this problem. Given a source image, and the amount and direction(s) in which it is to be extrapolated, our system determines visually consistent content for the extrapolated regions using library images. As well as considering low-level matching, we achieve consistency at a higher level by using graph proxies for regions of source and library images. Treating images as graphs allows us to find candidates for image extrapolation in a feasible time. Consistency of subgraphs in source and library images is used to find good candidates for the additional content; these are then further filtered. Region boundary curves are aligned to ensure consistency where image parts are joined using a photomontage method. We demonstrate the power of our method in image editing applications. Miao Wang 0004, Yukun Lai, Ralph R. Martin, Shi-Min Hu 0001 |
ACM Trans. Graph. | 5 |
| 2014 | A practical algorithm for rendering interreflections with all-frequency BRDFsabstractAlgorithms for rendering interreflection (or indirect illumination) effects often make assumptions about the frequency range of the materials' reflectance properties. For example, methods based on Virtual Point Lights (VPLs) perform well for diffuse and semi-glossy materials but not so for highly glossy or specular materials; the situation is reversed for methods based on ray tracing. In this article, we present a practical algorithm for rendering interreflection effects with all-frequency BRDFs. Our method builds upon a spherical Gaussian representation of the BRDF, based on which a novel mathematical development of the interreflection equation is made. This allows us to efficiently compute one-bounce interreflection from a triangle to a shading point, by using an analytic formula combined with a piecewise linear approximation. We show through evaluation that this method is accurate for a wide range of BRDFs. We further introduce a hierarchical integration method to handle complex scenes (i.e., many triangles) with bounded errors. Finally, we have implemented the present algorithm on the GPU, achieving rendering performance ranging from near interactive to a few seconds per frame for various scenes with different complexity. Kun Xu 0003, Yan-Pei Cao 0001, Li-Qian Ma, Zhao Dong 0001, Rui Wang 0003, Shi-Min Hu 0001 |
ACM Trans. Graph. | 6 |
| 2014 | SalientShape: group saliency in image collections
Ming-Ming Cheng, Niloy J. Mitra, Sharon X. Huang, Shi-Min Hu 0001 |
Vis. Comput. | 4 |
| 2014 | Parametric meta-filter modeling from a single example pair
Shi-Sheng Huang, Guo-Xin Zhang, Yukun Lai, Johannes Kopf 0001, Daniel Cohen-Or, Shi-Min Hu 0001 |
Vis. Comput. | 6 |
| 2014 | Stereoscopic image completion and depth recovery
Tai-Jiang Mu, Ju-Hong Wang, Song-Pei Du, Shi-Min Hu 0001 |
Vis. Comput. | 4 |
| 2013 | A Data-Driven Approach to Realistic Shape MorphingabstractAbstract Morphing between 3D objects is a fundamental technique in computer graphics. Traditional methods of shape morphing focus on establishing meaningful correspondences and finding smooth interpolation between shapes. Such methods however only take geometric information as input and thus cannot in general avoid producing unnatural interpolation, in particular for large‐scale deformations. This paper proposes a novel data‐driven approach for shape morphing. Given a database with various models belonging to the same category, we treat them as data samples in the plausible deformation space. These models are then clustered to form local shape spaces of plausible deformations. We use a simple metric to reasonably represent the closeness between pairs of models. Given source and target models, the morphing problem is casted as a global optimization problem of finding a minimal distance path within the local shape spaces connecting these models. Under the guidance of intermediate models in the path, an extended as‐rigid‐as‐possible interpolation is used to produce the final morphing. By exploiting the knowledge of plausible models, our approach produces realistic morphing for challenging cases as demonstrated by various examples in the paper. Lin Gao 0004, Yukun Lai, Qixing Huang, Shi-Min Hu 0001 |
Comput. Graph. Forum | 4 |
| 2013 | Advanced graph model for tainted variable tracking
Yu-Ping Wang 0001, Shi-Min Hu 0001 |
Sci. China Inf. Sci. | 4 |
| 2013 | Preface of special issue on computational visual media
Shi-Min Hu 0001, Ralph R. Martin |
Graph. Model. | 1 |
| 2013 | Efficient synthesis of gradient solid textures
Guo-Xin Zhang, Yukun Lai, Shi-Min Hu 0001 |
Graph. Model. | 3 |
| 2013 | Preface
Shi-Min Hu 0001, Daniel Thalmann, Ruofeng Tong 0001 |
J. Comput. Sci. Technol. | 1 |
| 2013 | Motion-Aware Gradient Domain Video CompositionabstractFor images, gradient domain composition methods like Poisson blending offer practical solutions for uncertain object boundaries and differences in illumination conditions. However, adapting Poisson image blending to video presents new challenges due to the added temporal dimension. In video, the human eye is sensitive to small changes in blending boundaries across frames and slight differences in motions of the source patch and target video. We present a novel video blending approach that tackles these problems by merging the gradient of source and target videos and optimizing a consistent blending boundary based on a user-provided blending trimap for the source video. Our approach extends mean-value coordinates interpolation to support hybrid blending with a dynamic boundary while maintaining interactive performance. We also provide a user interface and source object positioning method that can efficiently deal with complex video sequences beyond the capabilities of alpha blending. Tao Chen 0015, Jun-Yan Zhu, Ariel Shamir, Shi-Min Hu 0001 |
IEEE Trans. Image Process. | 4 |
| 2013 | Mixed-Domain Edge-Aware Image ManipulationabstractThis paper presents a novel approach to edge-aware image manipulation. Our method processes a Gaussian pyramid from coarse to fine, and at each level, applies a nonlinear filter bank to the neighborhood of each pixel. Outputs of these spatially-varying filters are merged using global optimization. The optimization problem is solved using an explicit mixed-domain (real space and DCT transform space) solution, which is efficient, accurate, and easy-to-implement. We demonstrate applications of our method to a set of problems, including detail and contrast manipulation, HDR compression, nonphotorealistic rendering, and haze removal. Xian-Ying Li, Yan Gu 0001, Shi-Min Hu 0001, Ralph R. Martin |
IEEE Trans. Image Process. | 3 |
| 2013 | Aesthetic Image Enhancement by Dependence-Aware Object RecompositionabstractThis paper proposes an image-enhancement method to optimize photograph composition by rearranging foreground objects in the photograph. To adjust objects' positions while keeping the original scene content, we first perform a novel structure dependence analysis on the image to obtain the dependencies between all background regions. To determine the optimal positions for foreground objects, we formulate an optimization problem based on widely used heuristics for aesthetically pleasing pictures. Semantic relations between foreground objects are also taken into account during optimization. The final output is produced by moving foreground objects, together with their dependent regions, to optimal positions. The results show that our approach can effectively optimize photographs with single or multiple foreground objects without compromising the original photograph content. Miao Wang 0004, Shi-Min Hu 0001 |
IEEE Trans. Multim. | 3 |
| 2013 | 3-Sweep: extracting editable objects from a single photoabstractWe introduce an interactive technique for manipulating simple 3D shapes based on extracting them from a single photograph. Such extraction requires understanding of the components of the shape, their projections, and relations. These simple cognitive tasks for humans are particularly difficult for automatic algorithms. Thus, our approach combines the cognitive abilities of humans with the computational accuracy of the machine to solve this problem. Our technique provides the user the means to quickly create editable 3D parts---human assistance implicitly segments a complex object into its components, and positions them in space. In our interface, three strokes are used to generate a 3D component that snaps to the shape's outline in the photograph, where each stroke defines one dimension of the component. The computer reshapes the component to fit the image of the object in the photograph as well as to satisfy various inferred geometric constraints imposed by its global 3D structure. We show that with this intelligent interactive modeling tool, the daunting task of object extraction is made simple. Once the 3D object has been extracted, it can be quickly edited and placed back into photos or 3D scenes, permitting object-driven photo editing tasks which are impossible to perform in image-space. We show several examples and present a user study illustrating the usefulness of our technique. Tao Chen 0015, Zhe Zhu, Ariel Shamir, Shi-Min Hu 0001, Daniel Cohen-Or |
ACM Trans. Graph. | 4 |
| 2013 | A metric of visual comfort for stereoscopic motionabstractWe propose a novel metric of visual comfort for stereoscopic motion, based on a series of systematic perceptual experiments. We take into account disparity, motion in depth, motion on the screen plane, and the spatial frequency of luminance contrast. We further derive a comfort metric to predict the comfort of short stereoscopic videos. We validate it on both controlled scenes and real videos available on the internet, and show how all the factors we take into account, as well as their interactions, affect viewing comfort. Last, we propose various applications that can benefit from our comfort measurements and metric. Song-Pei Du, Belén Masiá, Shi-Min Hu 0001, Diego Gutierrez |
ACM Trans. Graph. | 3 |
| 2013 | Inverse image editing: recovering a semantic editing history from a before-and-after image pairabstractWe study the problem of inverse image editing , which recovers a semantically-meaningful editing history from a source image and an edited copy. Our approach supports a wide range of commonly-used editing operations such as cropping, object insertion and removal, linear and non-linear color transformations, and spatially-varying adjustment brushes. Given an input image pair, we first apply a dense correspondence method between them to match edited image regions with their sources. For each edited region, we determine geometric and semantic appearance operations that have been applied. Finally, we compute an optimal editing path from the region-level editing operations, based on predefined semantic constraints. The recovered history can be used in various applications such as image re-editing, edit transfer, and image revision control. A user study suggests that the editing histories generated from our system are semantically comparable to the ones generated by artists. Shi-Min Hu 0001, Kun Xu 0003, Li-Qian Ma, Bi-Ye Jiang, Jue Wang 0001 |
ACM Trans. Graph. | 1 |
| 2013 | PatchNet: a patch-based image representation for interactive library-driven image editingabstractWe introduce PatchNets , a compact, hierarchical representation describing structural and appearance characteristics of image regions, for use in image editing. In a PatchNet, an image region with coherent appearance is summarized by a graph node, associated with a single representative patch, while geometric relationships between different regions are encoded by labelled graph edges giving contextual information. The hierarchical structure of a PatchNet allows a coarse-to-fine description of the image. We show how this PatchNet representation can be used as a basis for interactive, library-driven, image editing. The user draws rough sketches to quickly specify editing constraints for the target image. The system then automatically queries an image library to find semantically-compatible candidate regions to meet the editing goal. Contextual image matching is performed using the PatchNet representation, allowing suitable regions to be found and applied in a few seconds, even from a library containing thousands of images. Shi-Min Hu 0001, Miao Wang 0004, Ralph R. Martin, Jue Wang 0001 |
ACM Trans. Graph. | 1 |
| 2013 | Qualitative organization of collections of shapes via quartet analysisabstractWe present a method for organizing a heterogeneous collection of 3D shapes for overview and exploration. Instead of relying on quantitative distances, which may become unreliable between dissimilar shapes, we introduce aqualitativeanalysis which utilizes multiple distance measures but only in cases where the measures can be reliably compared. Our analysis is based on the notion ofquartets, each defined by two pairs of shapes, where the shapes in each pair are close to each other, but far apart from the shapes in the other pair. Combining the information from many quartets computed across a shape collection using several distance measures, we create a hierarchical structure we callcategorization treeof the shape collection. This tree satisfies the topological (qualitative) constraints imposed by the quartets creating an effective organization of the shapes. We present categorization trees computed on various collections of shapes and compare them to ground truth data from human categorization. We further introduce the concept ofdegree of separationchart for every shape in the collection and show the effectiveness of using it for interactive shapes exploration. Shi-Sheng Huang, Ariel Shamir, Chao-Hui Shen, Hao (Richard) Zhang, Alla Sheffer, Shi-Min Hu 0001, Daniel Cohen-Or |
ACM Trans. Graph. | 6 |
| 2013 | Cubic mean value coordinatesabstractWe present a new method for interpolating both boundary values and gradients over a 2D polygonal domain. Despite various previous efforts, it remains challenging to define a closed-form interpolant that produces natural-looking functions while allowing flexible control of boundary constraints. Our method builds on an existing transfinite interpolant over a continuous domain, which in turn extends the classical mean value interpolant. We re-derive the interpolant from the mean value property of biharmonic functions, and prove that the interpolant indeed matches the gradient constraints when the boundary is piece-wise linear. We then give closed-form formula (as generalized barycentric coordinates) for boundary constraints represented as polynomials up to degree 3 (for values) and 1 (for normal derivatives) over each polygon edge. We demonstrate the flexibility and efficiency of our coordinates in two novel applications, smooth image deformation using curved cage networks and adaptive simplification of gradient meshes. Xian-Ying Li, Shi-Min Hu 0001 |
ACM Trans. Graph. | 3 |
| 2013 | Sketch2Scene: sketch-based co-retrieval and co-placement of 3D modelsabstractThis work presents Sketch2Scene , a framework that automatically turns a freehand sketch drawing inferring multiple scene objects to semantically valid, well arranged scenes of 3D models. Unlike the existing works on sketch-based search and composition of 3D models, which typically process individual sketched objects one by one, our technique performs co-retrieval and co-placement of 3D relevant models by jointly processing the sketched objects. This is enabled by summarizing functional and spatial relationships among models in a large collection of 3D scenes as structural groups . Our technique greatly reduces the amount of user intervention needed for sketch-based modeling of 3D scenes and fits well into the traditional production pipeline involving concept design followed by 3D modeling. A pilot study indicates that it is promising to use our technique as an alternative but more efficient tool of standard 3D modeling for 3D scene construction. Kun Xu 0003, Hongbo Fu 0001, Wei-Lun Sun, Shi-Min Hu 0001 |
ACM Trans. Graph. | 5 |
| 2013 | Anisotropic spherical GaussiansabstractWe present a novel anisotropic Spherical Gaussian (ASG) function, built upon the Bingham distribution [Bingham 1974], which is much more effective and efficient in representing anisotropic spherical functions than Spherical Gaussians (SGs). In addition to retaining many desired properties of SGs, ASGs are also rotationally invariant and capable of representing all-frequency signals. To further strengthen the properties of ASGs, we have derived approximate closed-form solutions for their integral, product and convolution operators, whose errors are nearly negligible, as validated by quantitative analysis. Supported by all these operators, ASGs can be adapted in existing SG-based applications to enhance their scalability in handling anisotropic effects. To demonstrate the accuracy and efficiency of ASGs in practice, we have applied ASGs in two important SG-based rendering applications and the experimental results clearly reveal the merits of ASGs. Kun Xu 0003, Wei-Lun Sun, Zhao Dong 0001, Danyong Zhao, Run-Dong Wu, Shi-Min Hu 0001 |
ACM Trans. Graph. | 6 |
| 2013 | PoseShop: Human Image Database Construction and Personalized Content SynthesisabstractWe present PoseShop--a pipeline to construct segmented human image database with minimal manual intervention. By downloading, analyzing, and filtering massive amounts of human images from the Internet, we achieve a database which contains 400 thousands human figures that are segmented out of their background. The human figures are organized based on action semantic, clothes attributes, and indexed by the shape of their poses. They can be queried using either silhouette sketch or a skeleton to find a given pose. We demonstrate applications for this database for multiframe personalized content synthesis in the form of comic-strips, where the main character is the user or his/her friends. We address the two challenges of such synthesis, namely personalization and consistency over a set of frames, by introducing head swapping and clothes swapping techniques. We also demonstrate an action correlation analysis application to show the usefulness of the database for vision application. Tao Chen 0015, Li-Qian Ma, Ming-Ming Cheng, Ariel Shamir, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2013 | Semiregular Solid Texturing from 2D Image ExemplarsabstractSolid textures, comprising 3D particles embedded in a matrix in a regular or semiregular pattern, are common in natural and man-made materials, such as brickwork, stone walls, plant cells in a leaf, etc. We present a novel technique for synthesizing such textures, starting from 2D image exemplars which provide cross-sections of the desired volume texture. The shapes and colors of typical particles embedded in the structure are estimated from their 2D cross-sections. Particle positions in the texture images are also used to guide spatial placement of the 3D particles during synthesis of the 3D texture. Our experiments demonstrate that our algorithm can produce higher quality structures than previous approaches; they are both compatible with the input images, and have a plausible 3D nature. Song-Pei Du, Shi-Min Hu 0001, Ralph R. Martin |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2013 | Changing Perspective in Stereoscopic ImagesabstractTraditional image editing techniques cannot be directly used to edit stereoscopic ("3D") media, as extra constraints are needed to ensure consistent changes are made to both left and right images. Here, we consider manipulating perspective in stereoscopic pairs. A straightforward approach based on depth recovery is unsatisfactory: Instead, we use feature correspondences between stereoscopic image pairs. Given a new, user-specified perspective, we determine correspondence constraints under this perspective and optimize a 2D warp for each image that preserves straight lines and guarantees proper stereopsis relative to the new camera. Experiments verify that our method generates new stereoscopic views that correspond well to expected projections, for a wide range of specified perspective. Various advanced camera effects, such as dolly zoom and wide angle effects, can also be readily generated for stereoscopic image pairs using our method. Song-Pei Du, Shi-Min Hu 0001, Ralph R. Martin |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2013 | View-Dependent Multiscale Fluid SimulationabstractFluid flows are highly nonlinear and nonstationary, with turbulence occurring and developing at different length and time scales. In real-life observations, the multiscale flow generates different visual impacts depending on the distance to the viewer. We propose a new fluid simulation framework that adaptively allocates computational resources according to the viewer's position. First, a 3D empirical mode decomposition scheme is developed to obtain the velocity spectrum of the turbulent flow. Then, depending on the distance to the viewer, the fluid domain is divided into a sequence of nested simulation partitions. Finally, the multiscale fluid motions revealed in the velocity spectrum are distributed nonuniformly to these view-dependent partitions, and the mixed velocity fields defined on different partitions are solved separately using different grid sizes and time steps. The fluid flow is solved at different spatial-temporal resolutions, such that higher frequency motions closer to the viewer are solved at higher resolutions and vice versa. The new simulator better utilizes the computing power, producing visually plausible results with realistic fine-scale details in a more efficient way. It is particularly suitable for large scenes with the viewer inside the fluid domain. Also, as high-frequency fluid motions are distinguished from low-frequency motions in the simulation, the numerical dissipation is effectively reduced. Chenfeng Li, Bo Ren 0003, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2013 | Poisson CoordinatesabstractHarmonic functions are the critical points of a Dirichlet energy functional, the linear projections of conformal maps. They play an important role in computer graphics, particularly for gradient-domain image processing and shape-preserving geometric computation. We propose Poisson coordinates, a novel transfinite interpolation scheme based on the Poisson integral formula, as a rapid way to estimate a harmonic function on a certain domain with desired boundary values. Poisson coordinates are an extension of the Mean Value coordinates (MVCs) which inherit their linear precision, smoothness, and kernel positivity. We give explicit formulas for Poisson coordinates in both continuous and 2D discrete forms. Superior to MVCs, Poisson coordinates are proved to be pseudoharmonic (i.e., they reproduce harmonic functions on n-dimensional balls). Our experimental results show that Poisson coordinates have lower Dirichlet energies than MVCs on a number of typical 2D domains (particularly convex domains). As well as presenting a formula, our approach provides useful insights for further studies on coordinates-based interpolation and fast estimation of harmonic functions. Xian-Ying Li, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2013 | Timeline Editing of Objects in VideoabstractWe present a video editing technique based on changing the timelines of individual objects in video, which leaves them in their original places but puts them at different times. This allows the production of object-level slow motion effects, fast motion effects, or even time reversal. This is more flexible than simply applying such effects to whole frames, as new relationships between objects can be created. As we restrict object interactions to the same spatial locations as in the original video, our approach can produce highquality results using only coarse matting of video objects. Coarse matting can be done efficiently using automatic video object segmentation, avoiding tedious manual matting. To design the output, the user interactively indicates the desired new life spans of objects, and may also change the overall running time of the video. Our method rearranges the timelines of objects in the video whilst applying appropriate object interaction constraints. We demonstrate that, while this editing technique is somewhat restrictive, it still allows many interesting results. Shao-Ping Lu, Song-Hai Zhang, Shi-Min Hu 0001, Ralph R. Martin |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2013 | Change Blindness ImagesabstractChange blindness refers to human inability to recognize large visual changes between images. In this paper, we present the first computational model of change blindness to quantify the degree of blindness between an image pair. It comprises a novel context-dependent saliency model and a measure of change, the former dependent on the site of the change, and the latter describing the amount of change. This saliency model in particular addresses the influence of background complexity, which plays an important role in the phenomenon of change blindness. Using the proposed computational model, we are able to synthesize changed images with desired degrees of blindness. User studies and comparisons to state-of-the-art saliency models demonstrate the effectiveness of our model. Li-Qian Ma, Kun Xu 0003, Tien-Tsin Wong, Bi-Ye Jiang, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2013 | Flow Field ModulationabstractThe nonlinear and nonstationary nature of Navier-Stokes equations produces fluid flows that can be noticeably different in appearance with subtle changes. In this paper, we introduce a method that can analyze the intrinsic multiscale features of flow fields from a decomposition point of view, by using the Hilbert-Huang transform method on 3D fluid simulation. We show how this method can provide insights to flow styles and help modulate the fluid simulation with its internal physical information. We provide easy-to-implement algorithms that can be integrated with standard grid-based fluid simulation methods and demonstrate how this approach can modulate the flow field and guide the simulation with different flow styles. The modulation is straightforward and relates directly to the flow's visual effect, with moderate computational overhead. Bo Ren 0003, Chenfeng Li, Ming C. Lin, Theodore Kim, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2013 | Internet visual media processing: a survey with graphics and vision applications
Shi-Min Hu 0001, Tao Chen 0015, Kun Xu 0003, Ming-Ming Cheng, Ralph R. Martin |
Vis. Comput. | 1 |
| 2012 | Efficient Solid Texture Synthesis Using Gradient Solids
Guo-Xin Zhang, Yukun Lai, Shi-Min Hu 0001 |
CVM | 3 |
| 2012 | Visual storylines: Semantic visualization of movie sequence
Tao Chen 0015, Aidong Lu, Shi-Min Hu 0001 |
Comput. Graph. | 3 |
| 2012 | Data-Driven Object Manipulation in ImagesabstractAbstract We present a framework for interactively manipulating objects in a photograph using related objects obtained from internet images. Given an image, the user selects an object to modify, and provides keywords to describe it. Objects with a similar shape are retrieved and segmented from online images matching the keywords, and deformed to correspond with the selected object. By matching the candidate object and adjusting manipulation parameters, our method appropriately modifies candidate objects and composites them into the scene. Supported manipulations include transferring texture, color and shape from the matched object to the target in a seamless manner. We demonstrate the versatility of our framework using several inputs of varying complexity, for object completion, augmentation, replacement and revealing. Our results are evaluated using a user study. Chen Goldberg, Tao Chen 0015, Ariel Shamir, Shi-Min Hu 0001 |
Comput. Graph. Forum | 5 |
| 2012 | An optimization approach for extracting and encoding consistent maps in a shape collectionabstractWe introduce a novel approach for computing high quality point-to-point maps among a collection of related shapes. The proposed approach takes as input a sparse set of imperfect initial maps between pairs of shapes and builds a compact data structure which implicitly encodes an improved set of maps between all pairs of shapes. These maps align well with point correspondences selected from initial maps; they map neighboring points to neighboring points; and they provide cycle-consistency, so that map compositions along cycles approximate the identity map. The proposed approach is motivated by the fact that a complete set of maps between all pairs of shapes that admits nearly perfect cycle-consistency are highly redundant and can be represented by compositions of maps through a single base shape. In general, multiple base shapes are needed to adequately cover a diverse collection. Our algorithm sequentially extracts such a small collection of base shapes and creates correspondences from each of these base shapes to all other shapes. These correspondences are found by global optimization on candidate correspondences obtained by diffusing initial maps. These are then used to create a compact graphical data structure from which globally optimal cycle-consistent maps can be extracted using simple graph algorithms. Experimental results on benchmark datasets show that the proposed approach yields significantly better results than state-of-the-art data-driven shape matching methods. Qixing Huang, Guo-Xin Zhang, Lin Gao 0004, Shi-Min Hu 0001, Adrian Butscher, Leonidas J. Guibas |
ACM Trans. Graph. | 4 |
| 2012 | Structure recovery by part assemblyabstractThis paper presents a technique that allows quick conversion of acquired low-quality data from consumer-level scanning devices to high-quality 3D models with labeled semantic parts and meanwhile their assembly reasonably close to the underlying geometry. This is achieved by a novel structure recovery approach that is essentially local to global and bottom up, enabling the creation of new structures by assembling existing labeled parts with respect to the acquired data. We demonstrate that using only a small-scale shape repository, our part assembly approach is able to faithfully recover a variety of high-level structures from only a single-view scan of man-made objects acquired by the Kinect system, containing a highly noisy, incomplete 3D point cloud and a corresponding RGB image. Chao-Hui Shen, Hongbo Fu 0001, Shi-Min Hu 0001 |
ACM Trans. Graph. | 4 |
| 2012 | Interactive images: cuboid proxies for smart image manipulationabstractImages are static and lack important depth information about the underlying 3D scenes. We introduceinteractive imagesin the context of man-made environments wherein objects are simple and regular, share various non-local relations (e.g., coplanarity, parallelism, etc.), and are often repeated. Our interactive framework creates partial scene reconstructions based on cuboid-proxies with minimal user interaction. It subsequently allows a range of intuitive image edits mimicking real-world behavior, which are otherwise difficult to achieve. Effectively, the user simply provides high-level semantic hints, while our system ensures plausible operations by conforming to the extracted non-local relations. We demonstrate our system on a range of real-world images and validate the plausibility of the results using a user study. Youyi Zheng, Xiang Chen 0001, Ming-Ming Cheng, Kun Zhou 0001, Shi-Min Hu 0001, Niloy J. Mitra |
ACM Trans. Graph. | 5 |
| 2012 | Fisheye Video CorrectionabstractVarious types of video can be captured with fisheye lenses; their wide field of view is particularly suited to surveillance video. However, fisheye lenses introduce distortion, and this changes as objects in the scene move, making fisheye video difficult to interpret. Current still fisheye image correction methods are either limited to small angles of view, or are strongly content dependent, and therefore unsuitable for processing video streams. We present an efficient and robust scheme for fisheye video correction, which minimizes time-varying distortion and preserves salient content in a coherent manner. Our optimization process is controlled by user annotation, and takes into account a wide set of measures addressing different aspects of natural scene appearance. Each is represented as a quadratic term in an energy minimization problem, leading to a closed-form solution via a sparse linear system. We illustrate our method with a range of examples, demonstrating coherent natural-looking video output. The visual quality of individual frames is comparable to those produced by state-of-the-art methods for fisheye still photograph correction. Chenfeng Li, Shi-Min Hu 0001, Ralph R. Martin, Chiew-Lan Tai |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2012 | ImageAdmixture: Putting Together Dissimilar Objects from GroupsabstractWe present a semiautomatic image editing framework dedicated to individual structured object replacement from groups. The major technical difficulty is element separation with irregular spatial distribution, hampering previous texture, and image synthesis methods from easily producing visually compelling results. Our method uses the object-level operations and finds grouped elements based on appearance similarity and curvilinear features. This framework enables a number of image editing applications, including natural image mixing, structure preserving appearance transfer, and texture mixing. Ming-Ming Cheng, Jiaya Jia, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2011 | Global contrast based salient region detectionabstractAutomatic estimation of salient object regions across images, without any prior assumption or knowledge of the contents of the corresponding scenes, enhances many computer vision and computer graphics applications. We introduce a regional contrast based salient object detection algorithm, which simultaneously evaluates global contrast differences and spatial weighted coherence scores. The proposed algorithm is simple, efficient, naturally multi-scale, and produces full-resolution, high-quality saliency maps. These saliency maps are further used to initialize a novel iterative version of GrabCut, namely SaliencyCut, for high quality unsupervised salient object segmentation. We extensively evaluated our algorithm using traditional salient object detection datasets, as well as a more challenging Internet image dataset. Our experimental results demonstrate that our algorithm consistently outperforms 15 existing salient object detection and segmentation methods, yielding higher precision and better recall rates. We also show that our algorithm can be used to efficiently extract salient object masks from Internet images, enabling effective sketch-based image retrieval (SBIR) via simple shape comparisons. Despite such noisy internet images, where the saliency regions are ambiguous, our saliency guided image retrieval achieves a superior retrieval rate compared with state-of-the-art SBIR methods, and additionally provides important target object region information. Ming-Ming Cheng, Guo-Xin Zhang, Niloy J. Mitra, Sharon X. Huang, Shi-Min Hu 0001 |
CVPR | 5 |
| 2011 | Preserving detailed features in digital bas-relief making
Zhe Bian, Shi-Min Hu 0001 |
Comput. Aided Geom. Des. | 2 |
| 2011 | Painting patches: Reducing flicker in painterly re-rendering of video
Song-Hai Zhang, Qiang Tong 0001, Shi-Min Hu 0001, Ralph R. Martin |
Sci. China Inf. Sci. | 3 |
| 2011 | Sketch guided solid texturing
Guo-Xin Zhang, Song-Pei Du, Yukun Lai, Tianyun Ni, Shi-Min Hu 0001 |
Graph. Model. | 5 |
| 2011 | ISRA-Based Grouping: A Disk Reorganization Approach for Disk Energy Conservation and Disk Performance EnhancementabstractReducing disk energy consumption and improving disk performance in high-performance computer systems are increasingly pressing issues for reasons of disk economy and efficiency. To achieve these goals, we define the concept of Immediate Successor Relationship Amount (ISRA) to represent the successor relationship of data blocks, and propose an ISRA-based grouping algorithm for disk reorganization, based on an undirected graph. We group data blocks that experience frequent successive accesses, then sort them using a merge-sort-like algorithm to determine the position of every group as well as the new position of every block within those groups. We evaluate our approach in terms of disk seek time and disk energy consumption, using Disksim and the log energy model. The results show clearly that both disk seek time and the energy needs can be reduced by about 50 percent. Xue-Liang Liao, Yu-Ping Wang 0001, Shi-Min Hu 0001 |
IEEE Trans. Computers | 4 |
| 2011 | Online Video Stream Abstraction and StylizationabstractThis paper gives an automatic method for online video stream abstraction, producing a temporally coherent output video stream, in a style with large regions of constant color and highlighted bold edges. Our system includes two novel components. Firstly, to provide coherent and simplified output, we segment frames, and use optical flow to propagate segmentation information from frame to frame; an error control strategy is used to help ensure that the propagated information is reliable. Secondly, to achieve coherent and attractive coloring of the output, we use a color scheme replacement algorithm specifically designed for an online video stream. We demonstrate real-time performance for CIF videos, allowing our approach to be used for live communication and other related applications. Song-Hai Zhang, Xian-Ying Li, Shi-Min Hu 0001, Ralph R. Martin |
IEEE Trans. Multim. | 3 |
| 2011 | A geometric study of v-style pop-ups: theories and algorithmsabstractPop-up books are a fascinating form of paper art with intriguing geometric properties. In this paper, we present a systematic study of a simple but common class of pop-ups consisting of patches falling into four parallel groups, which we call v-style pop-ups. We give sufficient conditions for a v-style paper structure to be pop-uppable. That is, it can be closed flat while maintaining the rigidity of the patches, the closing and opening do not need extra force besides holding two patches and are free of intersections, and the closed paper is contained within the page border. These conditions allow us to identify novel mechanisms for making pop-ups. Based on the theory and mechanisms, we developed an interactive tool for designing v-style pop-ups and an automated construction algorithm from a given geometry, both of which guaranteeing the pop-uppability of the results. Xian-Ying Li, Yan Gu 0001, Shi-Min Hu 0001 |
ACM Trans. Graph. | 4 |
| 2011 | Adaptive partitioning of urban facadesabstractAutomatically discovering high-level facade structures in unorganized 3D point clouds of urban scenes is crucial for applications like digitalization of real cities. However, this problem is challenging due to poor-quality input data, contaminated with severe missing areas, noise and outliers. This work introduces the concept of adaptive partitioning to automatically derive a flexible and hierarchical representation of 3D urban facades. Our key observation is that urban facades are largely governed by concatenated and/or interlaced grids. Hence, unlike previous automatic facade analysis works which are typically restricted to globally rectilinear grids, we propose to automatically partition the facade in an adaptive manner, in which the splitting direction, the number and location of splitting planes are all adaptively determined. Such an adaptive partition operation is performed recursively to generate a hierarchical representation of the facade. We show that the concept of adaptive partitioning is also applicable to flexible and robust analysis of image facades. We evaluate our method on a dozen of LiDAR scans of various complexity and styles, and the image facades from the eTRIMS database and the Ecole Centrale Paris database. A series of applications that benefit from our approach are also demonstrated. Chao-Hui Shen, Shi-Sheng Huang, Hongbo Fu 0001, Shi-Min Hu 0001 |
ACM Trans. Graph. | 4 |
| 2011 | Interactive hair rendering and appearance editing under environment lightingabstractWe present an interactive algorithm for hair rendering and appearance editing under complex environment lighting represented as spherical radial basis functions (SRBFs). Our main contribution is to derive a compact 1D circular Gaussian representation that can accurately model the hair scattering function introduced by [Marschner et al. 2003]. The primary benefit of this representation is that it enables us to evaluate, at run-time, closed-form integrals of the scattering function with each SRBF light, resulting in efficient computation of both single and multiple scatterings. In contrast to previous work, our algorithm computes the rendering integrals entirely on the fly and does not depend on expensive pre-computation. Thus we allow the user to dynamically change the hair scattering parameters, which can vary spatially. Analyses show that our 1D circular Gaussian representation is both accurate and concise. In addition, our algorithm incorporates the eccentricity of the hair. We implement our algorithm on the GPU, achieving interactive hair rendering and simultaneous appearance editing under complex environment maps for the first time. Kun Xu 0003, Li-Qian Ma, Bo Ren 0003, Rui Wang 0003, Shi-Min Hu 0001 |
ACM Trans. Graph. | 5 |
| 2010 | Feature aligned quad dominant remeshing using iterative local updates
Yukun Lai, Leif Kobbelt, Shi-Min Hu 0001 |
Comput. Aided Des. | 3 |
| 2010 | Optimization approach for 3D model watermarking by linear binary programming
Yu-Ping Wang 0001, Shi-Min Hu 0001 |
Comput. Aided Geom. Des. | 2 |
| 2010 | Instant Propagation of Sparse Edits on Images and VideosabstractAbstract The ability to quickly and intuitively edit digital contents has become increasingly important in our everyday life. We propose a novel method for propagating a sparse set of user edits (e.g., changes in color, brightness, contrast, etc.) expressed as casual strokes to nearby regions in an image or video with similar appearances. Existing methods for edit propagation are typically based on optimization, whose computational cost can be prohibitive for large inputs. We re‐formulate propagation as a function interpolation problem in a high‐dimensional space, which we solve very efficiently using radial basis functions. While simple to implement, our method significantly improves the speed and space cost of existing methods, and provides instant feedback of propagation results even on large images and videos. Shi-Min Hu 0001 |
Comput. Graph. Forum | 3 |
| 2010 | Harmonic Field Based Volume Model Construction from Triangle Soup
Chao-Hui Shen, Guo-Xin Zhang, Yukun Lai, Shi-Min Hu 0001, Ralph R. Martin |
J. Comput. Sci. Technol. | 4 |
| 2010 | EditorialabstractThis volume of Computer Animation and Virtual Worlds (CAVW) contains a selection of papers submitted to CASA 2010, the 23rd International Conference on Computer Animation and Social Agents. CASA is one of the premier international conferences in the field of computer animation and social agents, organized under the auspices of the Computer Graphics Society (CGS). It has been founded in 1988, and, over the last years, it has been organized in Europe: Geneva (2002, 2004, 2006), Hasselt (2007), Amsterdam (2009); in USA: Philadelphia (1998, 2000), New Jersey (2003); and in Asia: Seoul (2001, 2008), Hong Kong (2005). This year, CASA 2010 was organized in Saint-Malo, France from the 30th of May to the 2nd of June 2010. The organization was done by Bunraku, an INRIA Project-team in common with CNRS, INSA of Rennes, University of Rennes 1, and Ecole Normale Supérieure de Cachan. The CASA 2010 edition received 104 submissions from 27 countries and 6 continents. Each submission received at least 3 reviews and 32 among them were selected to appear in this special issue of Computer Animation and Virtual Worlds. We thank all the authors who have submitted their work to this conference allowing us to present this nice and diversified program. We thank also the International Program Committee members and the additional external reviewers for the time and energy they have invested in the reviewing process. A particular thank goes to the INRIA conference support team, especially to Edith Blin-Guyot and Steeve Tessier, for their support in organizing the conference, taking care of financial, material, and organizational matters. CASA 2010 has been sponsored by INRIA, by the IRIS European Network of Excellence (Integrating Research in Interactive Storytelling), by GDR IG (Groupement de Recherche Informatique Graphique), by Fondation Michel Métivier, by the Brittany Regional Council, by the University of Rennes 1, by the Ecole Normale Supérieure de Cachan, and by the Biometrics Company. The 32 papers presented in this special issue are divided into several categories: Cartoon and Sketch-based animation techniques, Stylized animation, Deformable models, Meshes, Physically based animation, Motion Analysis and Synthesis, Steering and Crowds, Facial expression, Social Agents, and finally Augmented Reality. Stéphane Donikian, Elisabeth André, Shi-Min Hu 0001, Daniel Thalmann |
Comput. Animat. Virtual Worlds | 3 |
| 2010 | RepFinder: finding approximately repeated scene elements for image editingabstractRepeated elements are ubiquitous and abundant in both manmade and natural scenes. Editing such images while preserving the repetitions and their relations is nontrivial due to overlap, missing parts, deformation across instances, illumination variation, etc. Manually enforcing such relations is laborious and error-prone. We propose a novel framework where user scribbles are used to guide detection and extraction of such repeated elements. Our detection process, which is based on a novel boundary band method, robustly extracts the repetitions along with their deformations. The algorithm only considers the shape of the elements, and ignores similarity based on color, texture, etc. We then use topological sorting to establish a partial depth ordering of overlapping repeated instances. Missing parts on occluded instances are completed using information from other instances. The extracted repeated instances can then be seamlessly edited and manipulated for a variety of high level tasks that are otherwise difficult to perform. We demonstrate the versatility of our framework on a large set of inputs of varying complexity, showing applications to image rearrangement, edit transfer, deformation propagation, and instance replacement. Ming-Ming Cheng, Niloy J. Mitra, Sharon X. Huang, Shi-Min Hu 0001 |
ACM Trans. Graph. | 5 |
| 2010 | Popup: automatic paper architectures from 3D modelsabstractPaper architectures are 3D paper buildings created by folding and cutting. The creation process of paper architecture is often labor-intensive and highly skill-demanding, even with the aid of existing computer-aided design tools. We propose an automatic algorithm for generating paper architectures given a user-specified 3D model. The algorithm is grounded on geometric formulation of planar layout for paper architectures that can be popped-up in a rigid and stable manner, and sufficient conditions for a 3D surface to be popped-up from such a planar layout. Based on these conditions, our algorithm computes a class of paper architectures containing two sets of parallel patches that approximate the input geometry while guaranteed to be physically realizable. The method is demonstrated on a number of architectural examples, and physically engineered results are presented. Xian-Ying Li, Chao-Hui Shen, Shi-Sheng Huang, Shi-Min Hu 0001 |
ACM Trans. Graph. | 5 |
| 2010 | Metric-Driven RoSy Field Design and RemeshingabstractDesigning rotational symmetry fields on surfaces is an important task for a wide range of graphics applications. This work introduces a rigorous and practical approach for automatic N-RoSy field design on arbitrary surfaces with user-defined field topologies. The user has full control of the number, positions, and indexes of the singularities (as long as they are compatible with necessary global constraints), the turning numbers of the loops, and is able to edit the field interactively. We formulate N-RoSy field construction as designing a Riemannian metric such that the holonomy along any loop is compatible with the local symmetry of N-RoSy fields. We prove the compatibility condition using discrete parallel transport. The complexity of N-RoSy field design is caused by curvatures. In our work, we propose to simplify the Riemannian metric to make it flat almost everywhere. This approach greatly simplifies the process and improves the flexibility such that it can design N-RoSy fields with single singularity and mixed-RoSy fields. This approach can also be generalized to construct regular remeshing on surfaces. To demonstrate the effectiveness of our approach, we apply our design system to pen-and-ink sketching and geometry remeshing. Furthermore, based on our remeshing results with high global symmetry, we generate Celtic knots on surfaces directly. Yukun Lai, Miao Jin, Xuexiang Xie, Ying He 0001, Jonathan Palacios, Eugene Zhang, Shi-Min Hu 0001, Xianfeng Gu |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2009 | Robust principal curvatures using feature adapted integral invariantsabstractPrincipal curvatures and principal directions are fundamental local geometric properties. They are well defined on smooth surfaces. However, due to the nature as higher order differential quantities, they are known to be sensitive to noise. A recent work by Yang et al. combines principal component analysis with integral invariants and computes robust principal curvatures on multiple scales. Although the freedom of choosing the radius r gives results on different scales, in practice it is not an easy task to choose the most appropriate r for an arbitrary given model. Worse still, if the model contains features of different scales, a single r does not work well at all. In this work, we propose a scheme to automatically assign appropriate radii across the surface based on local surface characteristics. The radius r is not constant and adapts to the scale of local features. An efficient, iterative algorithm is used to approach the optimal assignment and the partition of unity is incorporated to smoothly combine the results with different radii. In this way, we can achieve a better balance between the robustness and the accuracy of feature locations. We demonstrate the effectiveness of our approach with robust principal direction field computation and feature extraction. Yukun Lai, Shi-Min Hu 0001, Tong Fang |
Symposium on Solid and Physical Modeling | 2 |
| 2009 | Rapid and effective segmentation of 3D models using random walks
Yukun Lai, Shi-Min Hu 0001, Ralph R. Martin, Paul L. Rosin |
Comput. Aided Geom. Des. | 2 |
| 2009 | Simulating Gaseous Fluids with Low and High SpeedsabstractAbstract Gaseous fluids may move slowly, as smoke does, or at high speed, such as occurs with explosions. High‐speed gas flow is always accompanied by low‐speed gas flow, which produces rich visual details in the fluid motion. Realistic visualization involves a complex dynamic flow field with both low and high speed fluid behavior. In computer graphics, algorithms to simulate gaseous fluids address either the low speed case or the high speed case, but no algorithm handles both efficiently. With the aim of providing visually pleasing results, we present a hybrid algorithm that efficiently captures the essential physics of both low‐ and high‐speed gaseous fluids. We model the low speed gaseous fluids by a grid approach and use a particle approach for the high speed gaseous fluids. In addition, we propose a physically sound method to connect the particle model to the grid model. By exploiting complementary strengths and avoiding weaknesses of the grid and particle approaches, we produce some animation examples and analyze their computational performance to demonstrate the effectiveness of the new hybrid method. Chenfeng Li, Shi-Min Hu 0001, Brian A. Barsky |
Comput. Graph. Forum | 3 |
| 2009 | Edit Propagation on Bidirectional Texture FunctionsabstractAbstract We propose an efficient method for editing bidirectional texture functions (BTFs) based on edit propagation scheme. In our approach, users specify sparse edits on a certain slice of BTF. An edit propagation scheme is then applied to propagate edits to the whole BTF data. The consistency of the BTF data is maintained by propagating similar edits to points with similar underlying geometry/reflectance. For this purpose, we propose to use view independent features including normals and reflectance features reconstructed from each view to guide the propagation process. We also propose an adaptive sampling scheme for speeding up the propagation process. Since our method needn't any accurate geometry and reflectance information, it allows users to edit complex BTFs with interactive feedback. Kun Xu 0003, Jiaping Wang, Xin Tong 0001, Shi-Min Hu 0001, Baining Guo |
Comput. Graph. Forum | 4 |
| 2009 | Generalized Discrete Ricci FlowabstractAbstract Surface Ricci flow is a powerful tool to design Riemannian metrics by user defined curvatures. Discrete surface Ricci flow has been broadly applied for surface parameterization, shape analysis, and computational topology. Conventional discrete Ricci flow has limitations. For meshes with low quality triangulations, if high conformality is required, the flow may get stuck at the local optimum of the Ricci energy. If convergence to the global optimum is enforced, the conformality may be sacrificed. This work introduces a novel method to generalize the traditional discrete Ricci flow. The generalized Ricci flow is more flexible, more robust and conformal for meshes with low quality triangulations. Conventional method is based on circle packing, which requires two circles on an edge intersect each other at an acute angle. Generalized method allows the two circles either intersect or separate from each other. This greatly improves the flexibility and robustness of the method. Furthermore, the generalized Ricci flow preserves the convexity of the Ricci energy, this ensures the uniqueness of the global optimum. Therefore the algorithm won't get stuck at the local optimum. Generalized discrete Ricci flow algorithms are explained in details for triangle meshes with both Euclidean and hyperbolic background geometries. Its advantages are demonstrated by theoretic proofs and practical applications in graphics, especially surface parameterization. Yongliang Yang 0002, Ren Guo, Feng Luo 0002, Shi-Min Hu 0001, Xianfeng Gu |
Comput. Graph. Forum | 4 |
| 2009 | A Shape-Preserving Approach to Image ResizingabstractAbstract We present a novel image resizing method which attempts to ensure that important local regions undergo a geometric similarity transformation, and at the same time, to preserve image edge structure. To accomplish this, we define handles to describe both local regions and image edges, and assign a weight for each handle based on an importance map for the source image. Inspired by conformal energy, which is widely used in geometry processing, we construct a novel quadratic distortion energy to measure the shape distortion for each handle. The resizing result is obtained by minimizing the weighted sum of the quadratic distortion energies of all handles. Compared to previous methods, our method allows distortion to be diffused better in all directions, and important image edges are well‐preserved. The method is efficient, and offers a closed form solution. Guo-Xin Zhang, Ming-Ming Cheng, Shi-Min Hu 0001, Ralph R. Martin |
Comput. Graph. Forum | 3 |
| 2009 | Editor's note
Shi-Min Hu 0001 |
Sci. China Ser. F Inf. Sci. | 1 |
| 2009 | Video-based running water animation in Chinese painting style
Song-Hai Zhang, Tao Chen 0015, Yi-Fei Zhang, Shi-Min Hu 0001, Ralph R. Martin |
Sci. China Ser. F Inf. Sci. | 4 |
| 2009 | Evaluation for Small Visual Difference Between Conforming Meshes on Strain Field
Zhe Bian, Shi-Min Hu 0001, Ralph R. Martin |
J. Comput. Sci. Technol. | 2 |
| 2009 | Guest Editorial Solid and Physical ModelingabstractThis special section contains improved and extended versions of six selected presentations at the ACM Symposium on Solid and Physical Modeling and Applications (ACM SPM), held in Beijing, China, from June 4 to 6, 2007. Shi-Min Hu 0001, Bruno Lévy 0001, Dinesh Manocha |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2009 | Stripification of Free-Form Surfaces With Global Error Bounds for Developable ApproximationabstractDevelopable surfaces have many desired properties in the manufacturing process. Since most existing CAD systems utilize tensor-product parametric surfaces including B-splines as design primitives, there is a great demand in industry to convert a general free-form parametric surface within a prescribed global error bound into developable patches. In this paper, we propose a practical and efficient solution to approximate a rectangular parametric surface with a small set ofC0-joint developable strips. The key contribution of the proposed algorithm is that, several optimization problems are elegantly solved in a sequence that offers a controllable global error bound on the developable surface approximation. Experimental results are presented to demonstrate the effectiveness and stability of the proposed algorithm. Yong-Jin Liu 0001, Yukun Lai, Shi-Min Hu 0001 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2009 | Sketch2Photo: internet image montageabstractWe present a system that composes a realistic picture from a simple freehand sketch annotated with text labels. The composed picture is generated by seamlessly stitching several photographs in agreement with the sketch and text labels; these are found by searching the Internet. Although online image search generates many inappropriate results, our system is able to automatically select suitable photographs to generate a high quality composition, using a filtering scheme to exclude undesirable images. We also provide a novel image blending algorithm to allow seamless image composition. Each blending result is given a numeric score, allowing us to find an optimal combination of discovered images. Experimental results show the method is very successful; we also evaluate our system using the results from two user studies. Tao Chen 0015, Ming-Ming Cheng, Ariel Shamir, Shi-Min Hu 0001 |
ACM Trans. Graph. | 5 |
| 2009 | Automatic and topology-preserving gradient mesh generation for image vectorizationabstractGradient mesh vector graphics representation, used in commercial software, is a regular grid with specified position and color, and their gradients, at each grid point. Gradient meshes can compactly represent smoothly changing data, and are typically used for single objects. This paper advances the state of the art for gradient meshes in several significant ways. Firstly, we introduce a topology-preserving gradient mesh representation which allows an arbitrary number of holes . This is important, as objects in images often have holes, either due to occlusion, or their 3D structure. Secondly, our algorithm uses the concept of image manifolds, adapting surface parameterization and fitting techniques to generate the gradient mesh in a fully automatic manner. Existing gradient-mesh algorithms require manual interaction to guide grid construction, and to cut objects with holes into disk-like regions. Our new algorithm is empirically at least 10 times faster than previous approaches. Furthermore, image segmentation can be used with our new algorithm to provide automatic gradient mesh generation for a whole image . Finally, fitting errors can be simply controlled to balance quality with storage. Yukun Lai, Shi-Min Hu 0001, Ralph R. Martin |
ACM Trans. Graph. | 2 |
| 2009 | Efficient affinity-based edit propagation using K-D treeabstractImage/video editing by strokes has become increasingly popular due to the ease of interaction. Propagating the user inputs to the rest of the image/video, however, is often time and memory consuming especially for large data. We propose here an efficient scheme that allows affinity-based edit propagation to be computed on data containing tens of millions of pixels at interactive rate (in matter of seconds). The key in our scheme is a novel means for approximately solving the optimization problem involved in edit propagation, using adaptive clustering in a high-dimensional, affinity space. Our approximation significantly reduces the cost of existing affinity-based propagation methods while maintaining visual fidelity, and enables interactive stroke-based editing even on high resolution images and long video sequences using commodity computers. Kun Xu 0003, Shi-Min Hu 0001, Tian-Qiang Liu |
ACM Trans. Graph. | 4 |
| 2009 | A New Watermarking Method for 3D Models Based on Integral InvariantsabstractIn this paper, we propose a new semi-fragile watermarking algorithm for the authentication of 3D models based on integral invariants. A watermark image is embedded by modifying the integral invariants of some of the vertices. In order to modify the integral invariants, the positions of a vertex and its neighbors are shifted. To extract the watermark, all the vertices are tested for the embedded information, and this information is combined to recover the watermark image. The number of parts of the watermark image that can be recovered will determine the authentication decision. Experimental tests show that this method is robust against normal use modifications introduced by rigid transformations, format conversions, rounding errors, etc., and can be used to test for malicious attacks such as mesh editing and cropping. An additional contribution of this paper is a new algorithm for computing two kinds of integral invariants. Yu-Ping Wang 0001, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2009 | Vectorizing Cartoon AnimationsabstractWe present a system for vectorizing 2D raster format cartoon animations. The output animations are visually flicker free, smaller in file size, and easy to edit. We identify decorative lines separately from colored regions. We use an accurate and semantically meaningful image decomposition algorithm, supporting an arbitrary color model for each region. To ensure temporal coherence in the output, we reconstruct a universal background for all frames and separately extract foreground regions. Simple user-assistance is required to complete the background. Each region and decorative line is vectorized and stored together with their motions from frame to frame. The contributions of this paper are: 1) the new trapped-ball segmentation method, which is fast, supports nonuniformly colored regions, and allows robust region segmentation even in the presence of imperfectly linked region edges, 2) the separate handling of decorative lines as special objects during image decomposition, avoiding results containing multiple short, thin oversegmented regions, and 3) extraction of a single patch-based background for all frames, which provides a basis for consistent, flicker-free animations. Song-Hai Zhang, Tao Chen 0015, Yi-Fei Zhang, Shi-Min Hu 0001, Ralph R. Martin |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2008 | Comparing Small Visual Differences between Conforming Meshes
Zhe Bian, Shi-Min Hu 0001, Ralph R. Martin |
GMP | 2 |
| 2008 | Fast mesh segmentation using random walksabstract3D mesh models are now widely available for use in various applications. The demand for automatic model analysis and understanding is ever increasing. Mesh segmentation is an important step towards model understanding, and acts as a useful tool for different mesh processing applications, e.g. reverse engineering and modeling by example. We extend a random walk method used previously for image segmentation to give algorithms for both interactive and automatic mesh segmentation. This method is extremely efficient, and scales almost linearly with increasing number of faces. For models of moderate size, interactive performance is achieved with commodity PCs. It is easy-to-implement, robust to noise in the mesh, and yields results suitable for downstream applications for both graphical and engineering models. Yukun Lai, Shi-Min Hu 0001, Ralph R. Martin, Paul L. Rosin |
Symposium on Solid and Physical Modeling | 2 |
| 2008 | An incremental approach to feature aligned quad dominant remeshingabstractIn this paper we present a new algorithm which turns an unstructured triangle mesh into a quad-dominant mesh with edges aligned to the principal directions of the underlying geometry. Instead of computing a globally smooth parameterization or integrating curvature lines along a tangent vector field, we simply apply an iterative relaxation scheme which incrementally aligns the mesh edges to the principal directions. The quad-dominant mesh is eventually obtained by dropping the not-aligned diagonals from the triangle mesh. A post-processing stage is introduced to further improve the results. The major advantage of our algorithm is its conceptual simplicity since it is merely based on elementary mesh operations such as edge collapse, flip, and split. The resulting meshes exhibit a very good alignment to surface features and rather uniform distribution of mesh vertices. This makes them very well-suited, e.g., as Catmull-Clark Subdivision control meshes. Yukun Lai, Leif Kobbelt, Shi-Min Hu 0001 |
Symposium on Solid and Physical Modeling | 3 |
| 2008 | Fairing wireframes in industrial surface designabstractWireframe is a modeling tool widely used in industrial geometric design. The term wireframe refers to two sets of curves, with the property that each curve from one set intersects with each curve from the other set. Akin to the mu-, v-isocurves in a tensor-product surface, the two sets of curves in a wireframe span an underlying surface. In many industrial design activities, wireframes are usually set up and adjusted by the designers before the whole surfaces are reconstructed. For adjustment, the fairness of wireframe has a direct influence on the quality of the underlying surface. Wireframe fairing is significantly different from fairing individual curves in that intersections should be preserved and kept in the same order. In this paper, we first present a technique for wireframe fairing by fixing the parameters during fairing. The limitation of fixed parameters is further released by an iterative gradient descent optimization method with step-size control. Experimental results show that our solution is efficient, and produces reasonably fairing results of the wireframes. Yukun Lai, Yong-Jin Liu 0001, Shi-Min Hu 0001 |
Shape Modeling International | 4 |
| 2008 | Solid and physical modeling
Shi-Min Hu 0001, Bruno Lévy 0001, Dinesh Manocha |
Comput. Aided Des. | 1 |
| 2008 | Solid and Physical Modeling
Shi-Min Hu 0001, Bruno Lévy 0001, Dinesh Manocha |
Comput. Aided Geom. Des. | 1 |
| 2008 | Shrinkability Maps for Content-Aware Video ResizingabstractAbstract A novel method is given for content‐aware video resizing, i.e. targeting video to a new resolution (which may involve aspect ratio change) from the original. We precompute a per‐pixel cumulative shrinkability map which takes into account both the importance of each pixel and the need for continuity in the resized result. (If both x and y resizing are required, two separate shrinkability maps are used, otherwise one suffices). A random walk model is used for efficient offline computation of the shrinkability maps. The latter are stored with the video to create a multi‐sized video, which permits arbitrary‐sized new versions of the video to be later very efficiently created in real‐time, e.g. by a video‐on‐demand server supplying video streams to multiple devices with different resolutions. These shrinkability maps are highly compressible, so the resulting multi‐sized videos are typically less than three times the size of the original compressed video. A scaling function operates on the multi‐sized video, to give the new pixel locations in the result, giving a high‐quality content‐aware resized video. Despite the great efficiency and low storage requirements for our method, we produce results of comparable quality to state‐of‐the‐art methods for content‐aware image and video resizing. Yi-Fei Zhang, Shi-Min Hu 0001, Ralph R. Martin |
Comput. Graph. Forum | 2 |
| 2008 | Spherical Piecewise Constant Basis Functions for All-Frequency Precomputed Radiance TransferabstractThis paper presents a novel basis function, called spherical piecewise constant basis function (SPCBF), for precomputed radiance transfer. SPCBFs have several desirable properties: rotatability, ability to represent all-frequency signals, and support for efficient multiple product. By smartly partitioning the illumination sphere into a set of subregions, and associating each subregion with an SPCBF valued 1 inside the region and 0 elsewhere, we precompute the light coefficients using the resulting SPCBFs. Efficient rotation of the light representation in SPCBFs is achieved by rotating the domain of SPCBFs. We run-time approximate the BRDF and visibility coefficients using the set of SPCBFs for light, possibly rotated, through fast lookup of summed-area-table (SAT) and visibility distance table (VDT), respectively. SPCBFs enable new effects such as object rotation in all-frequency rendering of dynamic scenes and on-the-fly BRDF editing under rotating environment lighting. With graphics hardware acceleration, our method achieves real-time frame rates. Kun Xu 0003, Yun-Tao Jia, Hongbo Fu 0001, Shi-Min Hu 0001, Chiew-Lan Tai |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2008 | Shape Deformation Using a Skeleton to Drive Simplex TransformationsabstractThis paper presents a novel skeleton-based method for deforming meshes (using an approximate skeleton, rather than a precise medial axis). The significant difference from previous skeleton-based methods is that the latter use the skeleton to control movement of vertices whereas we use it to control the simplices defining the model. By doing so, errors that occur near joints in other methods can be spread over the whole mesh, using an optimization process, resulting in smooth transitions near joints of the skeleton. By controlling simplices, our method has the advantage that no vertex weights need to be defined on the bones, which is a tedious requirement in previous skeleton-based methods. Our method can also easily be extended to control deformation by moving a few chosen line segments or vertices embedded in the object, rather than a skeleton. Shi-Min Hu 0001, Ralph R. Martin, Yongliang Yang 0002 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2008 | Optimal Surface Parameterization Using Inverse Curvature MapabstractMesh parameterization is a fundamental technique in computer graphics. Our paper focuses on solving the problem of finding the best discrete conformal mapping that also minimizes area distortion. Firstly, we deduce an exact analytical differential formula to represent area distortion by curvature change in the discrete conformal mapping, giving a dynamic Poisson equation. Our result shows the curvature map is invertible. Furthermore, we give the explicit Jacobi matrix of the inverse curvature map. Secondly, we formulate the task of computing conformal parameterizations with least area distortions as a constrained nonlinear optimization problem in curvature space. We deduce explicit conditions for the optima. Thirdly, we give an energy form to measure the area distortions, and show it has a unique global minimum. We use this to design an efficient algorithm, called free boundary curvature diffusion, which is guaranteed to converge to the global minimum. This result proves the common belief that optimal parameterization with least area distortion has a unique solution and can be achieved by free boundary conformal mapping. Major theoretical results and practical algorithms are presented for optimal parameterization based on the inverse curvature map. Comparisons are conducted with existing methods and using different energies. Novel parameterization applications are also introduced. Yongliang Yang 0002, Feng Luo 0002, Shi-Min Hu 0001, Xianfeng Gu |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2007 | Developable Strip Approximation of Parametric Surfaces with Global Error BoundsabstractDevelopable surfaces have many desired properties in manufacturing process. Since most existing CAD systems utilize parametric surfaces as the design primitive, there is a great demand in industry to convert a parametric surface within a prescribed global error bound into developable patches. In this work we propose a simple and efficient solution to approximate a general parametric surface with a minimum set of C0-joint developable strips. The key contribution of the proposed algorithm is that, several global optimization problems are elegantly solved in a sequence that offers a controllable global error bound on the developable surface approximation. Experimental results are presented to demonstrate the effectiveness and stability of the proposed algorithm. Yong-Jin Liu 0001, Yukun Lai, Shi-Min Hu 0001 |
PG | 3 |
| 2007 | Principal curvatures from the integral invariant viewpoint
Helmut Pottmann, Johannes Wallner 0001, Yongliang Yang 0002, Yukun Lai, Shi-Min Hu 0001 |
Comput. Aided Geom. Des. | 5 |
| 2007 | Real-time homogenous translucent material editingabstractAbstract This paper presents a novel method for real‐time homogenous translucent material editing under fixed illumination. We consider the complete analytic BSSRDF model proposed by Jensen et al. [ JMLH01 ], including both multiple scattering and single scattering. Our method allows the user to adjust the analytic parameters of BSSRDF and provides high‐quality, real‐time rendering feedback. Inspired by recently developed Precomputed Radiance Transfer (PRT) techniques, we approximate both the multiple scattering diffuse reflectance function and the single scattering exponential attenuation function in the analytic model using basis functions, so that re‐computing the outgoing radiance at each vertex as parameters change reduces to simple dot products. In addition, using a non‐uniform piecewise polynomial basis, we are able to achieve smaller approximation error than using bases adopted in previous PRT‐based works, such as spherical harmonics and wavelets. Using hardware acceleration, we demonstrate that our system generates images comparable to [ JMLH01 ]at real‐time frame‐rates. Kun Xu 0003, Shi-Min Hu 0001 |
Comput. Graph. Forum | 5 |
| 2007 | 3D Morphing Using Strain Field Interpolation
Shi-Min Hu 0001, Ralph R. Martin |
J. Comput. Sci. Technol. | 2 |
| 2007 | Editing the topology of 3D models by sketchingabstractWe present a method for modifying the topology of a 3D model with user control. The heart of our method is a guided topology editing algorithm. Given a source model and a user-provided target shape, the algorithm modifies the source so that the resulting model is topologically consistent with the target. Our algorithm permits removing or adding various topological features (e.g., handles, cavities and islands) in a common framework and ensures that each topological change is made by minimal modification to the source model. To create the target shape, we have also designed a convenient 2D sketching interface for drawing 3D line skeletons. As demonstrated in a suite of examples, the use of sketching allows more accurate removal of topological artifacts than previous methods, and enables creative designs with specific topological goals. Qian-Yi Zhou, Shi-Min Hu 0001 |
ACM Trans. Graph. | 3 |
| 2007 | Robust Feature Classification and EditingabstractSharp edges, ridges, valleys, and prongs are critical for the appearance and an accurate representation of a 3D model. In this paper, we propose a novel approach that deals with the global shape of features in a robust way. Based on a remeshing algorithm which delivers an isotropic mesh in a feature-sensitive metric, features are recognized on multiple scales via integral invariants of local neighborhoods. Morphological and smoothing operations are then used for feature region extraction and classification into basic types such as ridges, valleys, and prongs. The resulting representation of feature regions is further used for feature-specific editing operations. Yukun Lai, Qian-Yi Zhou, Shi-Min Hu 0001, Johannes Wallner 0001, Helmut Pottmann |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2007 | Topology Repair of Solid Models Using SkeletonsabstractWe present a method for repairing topological errors on solid models in the form of small surface handles, which often arise from surface reconstruction algorithms. We utilize a skeleton representation that offers a new mechanism for identifying and measuring handles. Our method presents two unique advantages over previous approaches. First, handle removal is guaranteed not to introduce invalid geometry or additional handles. Second, by using an adaptive grid structure, our method is capable of processing huge models efficiently at high resolutions. Qian-Yi Zhou, Shi-Min Hu 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2007 | Handling degenerate cases in exact geodesic computation on triangle meshes
Yong-Jin Liu 0001, Qian-Yi Zhou, Shi-Min Hu 0001 |
Vis. Comput. | 3 |
| 2006 | Skeleton-Based Shape Deformation Using Simplex Transformations
Shi-Min Hu 0001, Ralph R. Martin |
Computer Graphics International | 2 |
| 2006 | Robust principal curvatures on multiple scales
Yongliang Yang 0002, Yukun Lai, Shi-Min Hu 0001, Helmut Pottmann |
Symposium on Geometry Processing | 3 |
| 2006 | Feature sensitive mesh segmentationabstractSegmenting meshes into natural regions is useful for model understanding and many practical applications. In this paper, we present a novel, automatic algorithm for segmenting meshes into meaningful pieces. Our approach is a clustering-based top-down hierarchical segmentation algorithm. We extend recent work on feature sensitive isotropic remeshing to generate a mesh hierarchy especially suitable for segmentation of large models with regions at multiple scales. Using integral invariants for estimation of local characteristics, our method is robust and efficient. Moreover, statistical quantities can be incorporated, allowing our approach to segment regions with different geometric characteristics or textures. Yukun Lai, Qian-Yi Zhou, Shi-Min Hu 0001, Ralph R. Martin |
Symposium on Solid and Physical Modeling | 3 |
| 2006 | A sweepline algorithm for Euclidean Voronoi diagram of circles
Donguk Kim 0001, Lisen Mu, Deok-Soo Kim, Shi-Min Hu 0001 |
Comput. Aided Des. | 5 |
| 2006 | Surface fitting based on a feature sensitive parametrization
Yukun Lai, Shi-Min Hu 0001, Helmut Pottmann |
Comput. Aided Des. | 2 |
| 2006 | Rendering Soft Shadows using Multilayered Shadow FinsabstractAbstract Generating soft shadows in real time is difficult. Exact methods (such as ray tracing, and multiple light source simulation) are too slow, while approximate methods often overestimate the umbra regions. In this paper, we introduce a new algorithm based on the shadow map method to quickly and highly accurately render soft shadows produced by a light source. Our method builds inner and outer translucent fins on objects to represent the penumbra area inside and outside hard shadows, respectively. The fins are traced into multilayered light space maps to store illuminance adjustment to shadows. The viewing space illuminance buffer is then calculated using those maps. Finally, by blending illuminance and shading, a scene with highly accurate soft shadow effects is produced. Our method does not suffer from umbra overestimation. Physical relations between light, objects and shadows demonstrate the soundness of our approach. Xiao-Hua Cai, Yun-Tao Jia, Shi-Min Hu 0001, Ralph R. Martin |
Comput. Graph. Forum | 4 |
| 2006 | Special issue on SPM 05
Leif Kobbelt, Vadim Shapiro, Mario Botsch, Frédéric Cazals, Daniel Cohen-Or, Hugues Hoppe, Shi-Min Hu 0001, Bert Jüttler, Myung-Soo Kim, James F. O'Brien |
Graph. Model. | 7 |
| 2006 | Geometry and Convergence Analysis of Algorithms for Registration of 3D Shapes
Helmut Pottmann, Qixing Huang, Yongliang Yang 0002, Shi-Min Hu 0001 |
Int. J. Comput. Vis. | 4 |
| 2006 | Surface mosaics
Yukun Lai, Shi-Min Hu 0001, Ralph R. Martin |
Vis. Comput. | 2 |
| 2006 | Spherical harmonics scaling
Jiaping Wang, Kun Xu 0003, Kun Zhou 0001, Stephen Lin 0001, Shi-Min Hu 0001, Baining Guo |
Vis. Comput. | 5 |
| 2005 | Geometric texture synthesis and transfer via geometry imagesabstractIn this paper, we present an automatic method which can transfer geometric textures from one object to another, and can apply a manually designed geometric texture to a model. Our method is based on geometry images as introduced by Gu et al. The key ideas in this method involve geometric texture extraction, boundary consistent texture synthesis, discretized orientation and scaling, and reconstruction of synthesized geometry. Compared to other methods, our approach is efficient and easy-to-implement, and produces results of high quality. Yukun Lai, Shi-Min Hu 0001, D. X. Gu, Ralph R. Martin |
Symposium on Solid and Physical Modeling | 2 |
| 2005 | Special section on geometric modeling and processing
Shi-Min Hu 0001, Helmut Pottmann |
Comput. Aided Des. | 1 |
| 2005 | A second order algorithm for orthogonal projection onto curves and surfaces
Shi-Min Hu 0001, Johannes Wallner 0001 |
Comput. Aided Geom. Des. | 1 |
| 2005 | Fast degree elevation and knot insertion for B-spline curves
Qixing Huang, Shi-Min Hu 0001, Ralph R. Martin |
Comput. Aided Geom. Des. | 2 |
| 2005 | Video completion using tracking and fragment merging
Yun-Tao Jia, Shi-Min Hu 0001, Ralph R. Martin |
Vis. Comput. | 2 |
| 2004 | Special issue on geometric modeling and processing
Shi-Min Hu 0001, Helmut Pottmann |
Comput. Aided Geom. Des. | 1 |
| 2004 | Morphing based on strain field interpolationabstractAbstract Strain fields provide a method of deformation measurement based on physics. Using these as a tool, we can analyze deformation of objects in a measurable way. We have developed a new morphing technique based on strain field interpolation. Shape shaking and squeezing, which often happen when using linear interpolation for morphing, do not arise in our approach. We have also developed a new method to create isomorphic meshes from corresponding objects in two images. Meshes generated by this method have much fewer triangles than other methods, which greatly decreases calculation loads in the morphing process. Copyright © 2004 John Wiley & Sons, Ltd. Shi-Min Hu 0001, Ralph R. Martin |
Comput. Animat. Virtual Worlds | 2 |
| 2003 | Interactive Modeling of Tree BarkabstractThere exist many computer graphics techniques which could achieve high quality tree generation. However, only few works focus on realistic modeling of tree bark. Difficulties lie in the complex appearance of the bark surfaces from a single image. We address three main issues here: feature specification; height field assignment; and texture correction. For feature specification, we use texton channel analysis to specify a variant of common bark features, inluding ironbark, vertical and horizontal fractures, tessellation, furrowed cork, and lenticels. For height field assignment, we develop an intuitive and easy-to-use user interface (UI). Here similarity-based texture editing is used for assigning height fields within a texton channel mask. For texture correction, we use the modeled height fields to eliminate the underlying lighting effects in a captured texture. Our modeling system is image-based: it takes as input a bark image and produces as output a textured height field representing a bark sample. We demonstrate that out method is an effective and easy-to-use technique to interactively model a variety of photo realistic bark surfaces. Lifeng Wang 0001, Ligang Liu 0001, Shi-Min Hu 0001, Baining Guo |
PG | 4 |
| 2003 | Approximate merging of B-spline curves via knot adjustment and constrained optimization
Chiew-Lan Tai, Shi-Min Hu 0001, Qixing Huang |
Comput. Aided Des. | 2 |
| 2003 | Special issue on Pacific Graphics 2002
Shi-Min Hu 0001, Sabine Coquillart, Harry Shum |
Graph. Model. | 1 |
| 2003 | Adaptive tree similarity learning for image retrieval
Yong Rui, Shi-Min Hu 0001, Jia-Guang Sun 0001 |
Multim. Syst. | 3 |
| 2003 | View-dependent displacement mappingabstractSignificant visual effects arise from surface mesostructure, such as fine-scale shadowing, occlusion and silhouettes. To efficiently render its detailed appearance, we introduce a technique called view-dependent displacement mapping (VDM) that models surface displacements along the viewing direction. Unlike traditional displacement mapping, VDM allows for efficient rendering of self-shadows, occlusions and silhouettes without increasing the complexity of the underlying surface mesh. VDM is based on per-pixel processing, and with hardware acceleration it can render mesostructure with rich visual appearance in real time. Lifeng Wang 0001, Xin Tong 0001, Stephen Lin 0001, Shi-Min Hu 0001, Baining Guo, Harry Shum |
ACM Trans. Graph. | 5 |
| 2002 | An extension algorithm for B-splines by curve unclamping
Shi-Min Hu 0001, Chiew-Lan Tai, Song-Hai Zhang |
Comput. Aided Des. | 1 |
| 2002 | A constructive approach to solving 3-D geometric constraint systems using dependence analysis
Yan-Tao Li, Shi-Min Hu 0001, Jia-Guang Sun 0001 |
Comput. Aided Des. | 2 |
| 2002 | Two Accelerating Techniques for 3D Reconstruction
Shi-Xia Liu, Shi-Min Hu 0001, Jia-Guang Sun 0001 |
J. Comput. Sci. Technol. | 2 |
| 2001 | Optimal Adaptive Learning for Image RetrievalabstractLearning-enhanced relevance feedback is one of the most promising and active research directions in content-based image retrieval. However, the existing approaches either require prior knowledge of the data or entail high computation costs, making them less practical. To overcome these difficulties and motivated by the successful history of optimal adaptive filters, we present a new approach to interactive image retrieval. Specifically, we cast the image retrieval problem in the optimal filtering framework, which does not require prior knowledge of the data, supports incremental learning, is simple to implement and achieves better performance than state-of-the-art approaches. To evaluate the effectiveness and robustness of the proposed approach, extensive experiments have been carried out on a large heterogeneous image collection with 17,000 images. We report promising results on a wide variety of queries. Yong Rui, Shi-Min Hu 0001 |
CVPR (1) | 3 |
| 2001 | On the Numerical Redundancies of Geometric Constraint SystemsabstractDetermining redundant constraints is a critical task for geometric constraint solvers, since it dramatically affects the solution speed, accuracy, and stability. The paper attempts to determine the numerical redundancies of three-dimensional geometric constraint systems via a disturbance method. The constraints are translated into some unified forms and added to a constraint system incrementally. The redundancy of a constraint can then be decided by disturbing its value. We also prove that graph reduction methods can be used to accelerate the determination process. Yan-Tao Li, Shi-Min Hu 0001, Jia-Guang Sun 0001 |
PG | 2 |
| 2001 | An Effective Feature-Preserving Mesh Simplification Scheme Based on Face ConstrictionabstractA novel mesh simplification scheme that uses the face constriction process is presented. By introducing a statistical measure that can distinguish triangles having vertices of high local roughness from triangles in flat regions into our weight-ordering equation, along with other heuristics, our scheme can better preserve visually important features in the original mesh. To improve the shape quality of triangles, we adopt nonlinear face area sensitivity in the weight ordering. A learning and feedback mechanism is also utilized to enhance user controllability. The computations are simple, making our scheme time-effective and easy to implement. In addition to comparing our scheme with other mesh simplification algorithms empirically, we compare their performances by establishing a unifying ground among three basic simplification processes: decimate vertex, collapse edge, and constrict face. This unification allows us to analyze the intrinsic merits and demerits of simplification algorithms to help users make better selections. Shi-Min Hu 0001, Jia-Guang Sun 0001, Chiew-Lan Tai |
PG | 2 |
| 2001 | Modifying the shape of NURBS surfaces with geometric constraints
Shi-Min Hu 0001, Youfu Li 0001 |
Comput. Aided Des. | 1 |
| 2001 | Approximate merging of a pair of Bézier curves
Shi-Min Hu 0001, Rou-Feng Tong, Jia-Guang Sun 0001 |
Comput. Aided Des. | 1 |
| 2001 | Reconstruction of curved solids from engineering drawings
Shi-Xia Liu, Shi-Min Hu 0001, Yujian Chen, Jia-Guang Sun 0001 |
Comput. Aided Des. | 2 |
| 2001 | Conversion between triangular and rectangular Bézier patches
Shi-Min Hu 0001 |
Comput. Aided Geom. Des. | 1 |
| 2001 | Degree reduction of B-spline curves
Jun-Hai Yong, Shi-Min Hu 0001, Jia-Guang Sun 0001, Xing-Yu Tan |
Comput. Aided Geom. Des. | 2 |
| 2001 | CIM Algorithm for Approximating Three-Dimensional Polygonal Curves
Jun-Hai Yong, Shi-Min Hu 0001, Jia-Guang Sun 0001 |
J. Comput. Sci. Technol. | 2 |
| 2001 | Direct manipulation of FFD: efficient explicit solutions and decomposible multiple point constraints
Shi-Min Hu 0001, Hui Zhang 0013, Chiew-Lan Tai, Jia-Guang Sun 0001 |
Vis. Comput. | 1 |
| 2000 | A Matrix-Based Approach to Reconstruction of 3D Objects from Three Orthographic ViewsabstractPresents a matrix-based technique for reconstructing solids with quadric surfaces from three orthographic views. First, the relationship between a conic and its orthographic projections is developed using matrix theory. We then address the problem of finding the theoretical minimum number of views that are necessary for reconstructing an object with quadric surfaces. Next, we reconstruct the conic edges by finding their matrix representations in 3D space. This effectively constructs a model corresponding to the three views. Finally, volume information is searched within the wireframe model to form the final solids. The novelty of our algorithm is in the use of the matrix representation of conics to assist in the 3D reconstruction, which increases both the efficiency and the reliability of the proposed approach. Shi-Xia Liu, Shi-Min Hu 0001, Jia-Guang Sun 0001, Chiew-Lan Tai |
PG | 2 |
| 2000 | Bisection algorithms for approximating quadratic Bézier curves by G1 arc splines
Jun-Hai Yong, Shi-Min Hu 0001, Jia-Guang Sun 0001 |
Comput. Aided Des. | 2 |
| 1999 | A note on approximation of discrete data by G1 arc splines
Jun-Hai Yong, Shi-Min Hu 0001, Jia-Guang Sun 0001 |
Comput. Aided Des. | 2 |
| 1998 | A type of triangular ball surface and its properties
Shi-Min Hu 0001, Guojin Wang, Jianguang Sun |
J. Comput. Sci. Technol. | 1 |
| 1996 | Properties of two types of generalized ball curves
Shi-Min Hu 0001, Guo-Zhao Wang, Tong-Guang Jin |
Comput. Aided Des. | 1 |
| 1996 | Conversion of a tringular Bézier patch into three rectangular Bézier patches
Shi-Min Hu 0001 |
Comput. Aided Geom. Des. | 1 |
| 1996 | Generalized Subdivision of Bézier Surfaces
Shi-Min Hu 0001, Guo-Zhao Wang, Tong-Guang Jin |
CVGIP Graph. Model. Image Process. | 1 |
| 1996 | A subdivision scheme for rational triangular Bézier surfaces
Shi-Min Hu 0001 |
J. Comput. Sci. Technol. | 1 |